geohot
34ef448331
enable rangeify const folding
2025-09-15 11:50:13 +08:00
George Hotz and GitHub
1353250b6c
tags on bufferize are the tensor tags ( #12180 )
2025-09-15 11:46:03 +08:00
George Hotz and GitHub
60d7db093e
delete bufferized consts + output noops ( #12163 )
...
* bring const folding to rangeify
* comment that
2025-09-15 11:07:44 +08:00
qazal and GitHub
525c20dc7e
viz: remove unused runtime_stats feature ( #12177 )
2025-09-15 02:53:05 +03:00
qazal and GitHub
75ff9b7a9a
viz: add buffer lifetime to tooltip ( #12175 )
2025-09-15 02:33:50 +03:00
chenyu and GitHub
15b166ce6d
bump test_module_runs to 30 seconds ( #12174 )
...
25 seconds sometimes
2025-09-14 16:48:40 -04:00
943236ef74
move cast pat out of symbolic_simple ( #11945 )
...
* move pat
* move it here
* rm extra check
---------
Co-authored-by: Sieds Lykles <[email protected] >
2025-09-14 21:39:48 +02:00
Steven Shi and GitHub
25b1bc8eff
added top k sampling to examples/mamba ( #12061 )
2025-09-14 15:27:34 -04:00
Shun Usami and GitHub
34a05b31fe
Fix advanced tensor indexing setitem ( #12128 )
...
* Add failure test case for advanced tensor indexing setitem
* Fix advanced tensor indexing setitem when permuted
* Reduce line count
* Revert unnecessary change
* Combine two lines into one
2025-09-14 15:22:40 -04:00
chenyu and GitHub
d09c0f28c5
increase test_module_runs ( #12173 )
...
timed out on ci windows llvm
2025-09-14 15:19:21 -04:00
chenyu and GitHub
12a910f1d2
update torch 2.8 ( #12172 )
...
support _reshape_alias. something is wrong with one case of unfold
2025-09-14 15:19:03 -04:00
chenyu and GitHub
98ecab7563
remove ml_dtypes ( #12169 )
2025-09-14 14:20:05 -04:00
qazal and GitHub
02054b53fe
remove tests that pre date the uop spec ( #12168 )
...
* remove tests that pre date the uop spec
* const src
* for RANGEIFY=1
* update with bind
* remove import
2025-09-14 18:47:42 +03:00
qazal and GitHub
1591e4f66b
update outbufs selection in test_linearizer [pr] ( #12166 )
2025-09-14 13:46:49 +03:00
nimlgen and GitHub
d1ae30f7ef
hcq: do not spam with errors in -m device ( #12150 )
...
* hcq: do not spam with errors in -m device
* um?
* um?
* nn
* helps?
* um?
* no gc?
* fix
2025-09-14 10:56:59 +03:00
George Hotz and GitHub
d5bc27797b
fix some multitensor on rangeify ( #12162 )
...
* fix some multitensor on rangeify
* rangeify multi hacks
* copy on const
2025-09-14 14:31:57 +08:00
Meng Zhuo and GitHub
4b7904eca9
add cpu support for riscv64 ( #12136 )
2025-09-14 11:40:58 +08:00
George Hotz and GitHub
bcafa72b7f
use tags instead of graph_rewrite_map in rangeify ( #12110 )
...
* use tags instead of graph_rewrite_map in rangeify
* new style, add realize
* metadata works
* simple failure
* fix
* loops
* stuff becomes a NOOP when you remove it
* stuff becomes a NOOP when you remove it
* tags on bufferize
* bmnist works
* locals don't work
* shippable
* fix some tests
* simpler map_realize
* remove const hack
* debuggable test
* broke
* assign test
* straight up bug
* wooo it passes
* sink shouldn't be there
* fix ops
* bmnist
* kv cache ish
* Set RANGEIFY context variable to 0
* should work normal
* better
* types
* hacks to fix test_symbolic
* pm_add_buffers
* tests should pass
2025-09-14 11:39:01 +08:00
chenyu and GitHub
d2316ba91a
don't validate output in sdxl with fakeweights ( #12160 )
...
NULL backend passed validation before because both desired and actual went through NULL backend
2025-09-13 21:47:51 -04:00
nimlgen and GitHub
b1d1816f43
device: fix envvars ( #12159 )
2025-09-13 23:38:09 +03:00
nimlgen and GitHub
19d9d29b7e
device: compilers in tinygrad.device ( #12151 )
...
* hcq: do not spam with errors in -m device
* -m tinygrad p2
* fix
* ugh
* comp in ckey
* fix
* one more
* print defaults
* xx
2025-09-13 21:45:29 +03:00
qazal and GitHub
6410dcb7c2
viz: less verbose render loop ( #12158 )
...
* define visible once
* move y offsets to one place
2025-09-13 19:04:37 +03:00
nimlgen and GitHub
92df52d79a
make method_cache account for compiler ( #12156 )
...
* make method_cache account for compiler
* sorry
2025-09-13 17:00:11 +03:00
chenyu and GitHub
0c392089d9
update mypy ( #12155 )
2025-09-13 09:48:38 -04:00
qazal and GitHub
fbca6183ad
do not launch BEAM when opts_to_apply exists [pr] ( #12152 )
2025-09-13 14:57:46 +03:00
George Hotz and GitHub
b2a95d32bb
check clSetKernelArg ( #12149 )
2025-09-13 17:24:55 +08:00
George Hotz and GitHub
0695e322a8
fix android cpu device ( #12148 )
2025-09-13 15:42:04 +08:00
Sieds Lykles and GitHub
e3a3764917
delete fold_unrolled_divs ( #12146 )
2025-09-13 03:09:36 +02:00
Sieds Lykles and GitHub
51ed6e94b2
AxisType __repr__ method ( #12145 )
2025-09-13 01:15:38 +02:00
Sieds Lykles and GitHub
0757a9a819
add pytest-timeout of 3 min per item ( #12144 )
...
* add pytest-timeout with timeout of 3 min
* func_only
2025-09-13 00:48:41 +02:00
Sieds Lykles and GitHub
2fc0bd150b
Arange overflow raises error and one_hot upcast ( #11975 )
...
* add error
* to_dtype
* shorten line
* add test
* upcast one hot dim im overflows
2025-09-13 00:18:25 +02:00
chenyu and GitHub
aac3dceaf6
merge two PYTHON backend ci job ( #12143 )
...
* merge two PYTHON backend ci job
and mark anything that takes > 10 in test_ops slow
* two more
2025-09-12 17:36:46 -04:00
a12d0933c1
fix vec dtype in fast idiv ( #12080 )
...
* fix
* add vec dtypes to fuzzer
* add vec=False
---------
Co-authored-by: Sieds Lykles <[email protected] >
2025-09-12 23:00:43 +02:00
chenyu and GitHub
25091951ba
update test/models ( #12142 )
...
minor fix and run more stuff in tinygrad for speed
2025-09-12 16:43:28 -04:00
Sieds Lykles and GitHub
62376c8b2b
update store load noop pattern to use Invalid ( #12141 )
...
* update pattern
* add test
2025-09-12 22:25:53 +02:00
chenyu and GitHub
647965fb09
test_train cleanup ( #12140 )
...
* test_train cleanup
remove skipIf due to buffer sizes, runs locally
* those are slow
2025-09-12 13:21:30 -04:00
chenyu and GitHub
0fad07c684
viz serve default path ( #12139 )
...
`python tinygrad/viz/serve.py` shows last session instead of an empty page
2025-09-12 18:32:44 +03:00
nimlgen and GitHub
81e33b8439
system: cpu memory mappings are uncached ( #12137 )
...
* system: cpu memory mappings is uncached
* adm amd
2025-09-12 13:28:25 +03:00
qazal and GitHub
68b0ad05a4
viz: format tuple tags ( #12135 )
...
* viz: format tuple tags
* use python repr
2025-09-12 11:36:53 +03:00
qazal and GitHub
e80c8a7548
merge TestIndexing with TestSchedule + remove duplicate tests ( #12134 )
...
* merge TestIndexing with TestSchedule
* remove the arange_copy tests
* no FUSE_ARANGE import
2025-09-12 10:35:14 +03:00
Sieds Lykles and GitHub
b5a3b8de20
remove where on gated load if gates are the same ( #12129 )
...
* add rules
* add tests
2025-09-12 06:52:35 +02:00
George Hotz and GitHub
a2f502b89e
fix rangeify=1 ops on GPU ( #12130 )
2025-09-12 11:17:37 +08:00
George Hotz and GitHub
0766616962
isolate the const hacks in the old kernelize ( #12126 )
...
* isolate the const hacks in the old kernelize
* if rangeify, don't waste time
2025-09-12 08:35:35 +08:00
Sieds Lykles and GitHub
1f3950a484
Invalid idx ( #12067 )
...
* merge index_dtype_3
* new lowering with Invalid idx
* remove that dtype from range
* finish merge
* annotate better
* indentation
* dont need that anymore
* always process replay for openpilot
* more uop_given_valid for idx
* valid past index_child
* fix bug preventing load getting an alt value
* add track_match_stats back in in shapetracker and remove cache
* get_valid_idx -> get_valid and get_idx
* fix heuristics with new idx
* split line
* fix typo
* fix signature
* dont skip idx if stride is 0
the idx may still be invalid
* lower const with new valid
* delete to_indexed_uops
* update shapetracker test
* delete axis_is_masked
* add cache back
* move around comment
* fix get_valid bug
* move invalid fold to symbolic so its earlier
* cleanup
* update applying padto to new idx
* add unit tests
* cleanup
* fold line
* improve spec
* dont try to render Invalid as a float
* more consistent invalid index
* update some tests
* Fold index with true cond
* skip test
* vconst min max if Invalid in arg
* fix signature of UOp.const
* add test for min/max of Invalid CONST/VCONST
* add InvalidType to as_const signature
* is Invalid to isinstance
* Add InvalidType to ConstLike
* index gate is a where gate
* make that a metaclass
* fix heurisics for new idx
* mypy happy
2025-09-12 01:42:02 +02:00
chenyu and GitHub
544eb2c402
clean up test_scatter_reduce ( #12125 )
2025-09-11 16:36:58 -04:00
chenyu and GitHub
9ad6a56d17
smaller test_simple_reduce ( #12124 )
2025-09-11 15:45:38 -04:00
chenyu and GitHub
e5ef9ec5b1
remove IGNORE_OOB=0 in ci tests ( #12117 )
2025-09-11 15:05:04 -04:00
chenyu and GitHub
3a83b56da5
fix test_dequantization_mxfp4 ( #12123 )
...
* fix test_dequantization_mxfp4
* assert_allclose
* rtol
2025-09-11 14:22:06 -04:00
chenyu and GitHub
520e2e0727
actually run unit tests in ci MacOS (unit) ( #12122 )
...
* actually run unit tests in ci MacOS (unit)
* that's always wrong
2025-09-11 13:32:30 -04:00
nimlgen and GitHub
acb700fc26
ci: fix ptx env ( #12120 )
2025-09-11 12:42:15 -04:00
chenyu and GitHub
20cd7177de
delete test_bert_fuse_arange ( #12121 )
...
* delete test_bert_fuse_arange
it's the default now and we are not interested in FUSE_ARANGE=0 version
* remove -v
2025-09-11 12:35:51 -04:00
chenyu and GitHub
b07f962058
split metal model tests ( #12119 )
...
* split metal model tests
* llama too
2025-09-11 12:20:12 -04:00
chenyu and GitHub
66593f135f
remove duplicated test_real_world ( #12118 )
...
included in the test/models right below
2025-09-11 11:57:14 -04:00
qazal and GitHub
e76211fcbc
viz: specify all rect styles in parent ( #12115 )
...
* viz: specify all rect styles in parent
Visually a no-op, but it's easier to reason about when the rect's coloring comes from `g` parent that holds UOp data.
* this stays
2025-09-11 13:48:59 +03:00
nimlgen and GitHub
400ad93892
ci: gate boost paths for macos only ( #12114 )
2025-09-11 12:48:34 +03:00
George Hotz and GitHub
3ef0e5e01e
rangeify: use Ops.REALIZE and not Ops.CONTIGUOUS if it's added by system ( #12111 )
...
* rangeify: use Ops.REALIZE and not Ops.CONTIGUOUS if it's added by system
* fix contig + BufferizeOpts
* no outerworld
2025-09-11 11:56:59 +08:00
b1tg and GitHub
52ebed991e
change dtype promo lattice when fp8s is supported ( #12088 )
...
* change dtype promo lattice when fp8s is supported
* no device check
* int64 + uint64 => fp8
2025-09-10 22:09:11 -04:00
George Hotz and GitHub
d4eba5800d
rangeify cost function infrastructure ( #12091 )
...
* one call to hc opt
* does that pass?
* add cost function to rangeify
* test
* more test
* gate thread
* bufferize has shape
* ish
* match old behavior
* no ci there
2025-09-11 07:19:53 +08:00
qazal and GitHub
78610b681e
viz: light up children ( #12107 )
...
* viz: light up children
* keep tag coloring
2025-09-11 01:28:01 +03:00
Sieds Lykles and GitHub
3989f5b559
Revert "Simplify valid in symbolic ( #12104 )" ( #12108 )
...
This reverts commit 73d479a016 .
2025-09-10 23:36:40 +02:00
Sieds Lykles and GitHub
73d479a016
Simplify valid in symbolic ( #12104 )
...
* cleanup cast_folding
* from sym to symbolic
* no more sym in dtype lowering
* move around simplify_valid
* update test
2025-09-10 23:26:19 +02:00
chenyu and GitHub
e306650d39
remove GPUDevice ( #12106 )
2025-09-10 16:35:00 -04:00
George Hotz and GitHub
d8a7a1c9c7
BUFFERIZE shape should be each range, not the product ( #12105 )
...
* BUFFERIZE shape should be each range, not the product
* fix tests
* resolve
2025-09-11 04:02:24 +08:00
Sieds Lykles and GitHub
3730172c10
cleanup cast_folding ( #12101 )
...
* cleanup cast_folding
* from sym to symbolic
* no more sym in dtype lowering
2025-09-10 21:30:20 +02:00
chenyu and GitHub
0e266f376c
ops_gpu -> ops_cl ( #12103 )
2025-09-10 15:15:48 -04:00
chenyu and GitHub
0599e86186
replace hardcoded GPU in llama debug msg ( #12102 )
2025-09-10 13:56:40 -04:00
qazal and GitHub
5a84d86db7
viz: fix buffer tooltip offset ( #12100 )
...
* fixup offsets
* add buffer num to tooltip
2025-09-10 20:12:20 +03:00
nimlgen and GitHub
fb96394ff5
auto-select available compilers ( #12094 )
...
* device: auto select compilers
* fix
* metal+opencl
* nv/cuda
* test without ptx
* ptx
* fix tests
* fix
* fix test
* rename
* test + cleaner
* xx
* ops
* better test
* win?
* um?
* types
* debug
* win??
* sep rung
* wtf?
* debug
* skip win
* revert this
* types
2025-09-10 19:52:01 +03:00
chenyu and GitHub
bb67829e99
raise KernelOptError in TC _apply_tc_opt ( #12099 )
...
currently getting
```
2025-09-10 13:18:19
File "/home/chenyu/tinygrad/tinygrad/codegen/opt/search.py", line 149, in beam_search
2025-09-10 13:18:19
acted_lins: list[Scheduler] = flatten([get_kernel_actions(lin, include_0=False).values() for lin,_ in beam])
2025-09-10 13:18:19
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
2025-09-10 13:18:19
File "/home/chenyu/tinygrad/tinygrad/codegen/opt/search.py", line 107, in get_kernel_actions
2025-09-10 13:18:19
lin2.apply_opt(a)
2025-09-10 13:18:19
File "/home/chenyu/tinygrad/tinygrad/codegen/opt/postrange.py", line 169, in apply_opt
2025-09-10 13:18:19
ret = self._apply_tc_opt(use_tensor_cores, cast(int, opt.axis), tc_select, tc_opt)
2025-09-10 13:18:19
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
2025-09-10 13:18:19
File "/home/chenyu/tinygrad/tinygrad/codegen/opt/postrange.py", line 235, in _apply_tc_opt
2025-09-10 13:18:19
idx = self.rngs.index(a)
2025-09-10 13:18:19
^^^^^^^^^^^^^^^^^^
2025-09-10 13:18:19
ValueError: UOp(Ops.RANGE, dtypes.index, arg=(1002, <AxisType.REDUCE: 6>), src=(
2025-09-10 13:18:19
UOp(Ops.CONST, dtypes.index, arg=15, src=()),)) is not in list
```
2025-09-10 12:32:19 -04:00
George Hotz and GitHub
84b249ef0e
move simplify reduce out of devectorizer ( #12098 )
2025-09-10 21:24:57 +08:00
qazal and GitHub
5d66a2d885
viz: refactor range clipping ( #12097 )
2025-09-10 16:23:46 +03:00
George Hotz and GitHub
9789337722
early reduce simplify ( #12046 )
...
* early reduce simplify
* min changes
* need that
* that goes in simplify
* no more arange reduce opt
2025-09-10 21:02:46 +08:00
nimlgen and GitHub
21e6926a6a
HostLLVMCompiler -> CPULLVMCompiler ( #12096 )
2025-09-10 14:04:16 +03:00
nimlgen and GitHub
551560b87c
do not use getenv('PTX') in tests ( #12095 )
...
* test without ptx
* fix tests
* fix test
* linters
2025-09-10 14:04:07 +03:00
Sieds Lykles and GitHub
0e420e68b4
delete axis_is_masked ( #12092 )
2025-09-10 05:26:19 +02:00
George Hotz and GitHub
ef53a6fc19
one call to hc opt ( #12074 )
...
* one call to hc opt
* does that pass?
* Clean up postrange.py by removing comments
2025-09-10 11:18:18 +08:00
Sieds Lykles and GitHub
499f50483b
x | !x -> True ( #12090 )
2025-09-10 03:26:01 +02:00
Sieds Lykles and GitHub
5b73076e48
assert benchmark times ( #12042 )
...
* assert jitted times in openpilot
* better error
* better error
* add ASSERT_MIN_STEP_TIME to more models
* t is step_times
* update benchmark times
* update times
2025-09-09 23:40:02 +02:00
b1tg and GitHub
58d13a6e3e
remove redundant check ( #12087 )
2025-09-09 15:15:39 -04:00
qazal and GitHub
71fcb23d4a
viz: cleanup renderDag ( #12086 )
2025-09-09 19:19:45 +03:00
b1tg and GitHub
82e955fe79
fix inf bug in float_to_fp8 ( #12085 )
2025-09-09 12:02:56 -04:00
b1tg and GitHub
14faf7a5c0
AutoCastType tests for fp8s/bf16 ( #12084 )
2025-09-09 11:33:01 -04:00
qazal and GitHub
5e76eff26d
viz: pre fetch workers ( #12083 )
...
* viz: pre fetch workers
* move check
2025-09-09 15:56:39 +03:00
qazal and GitHub
5fde033794
viz: prune worker payload ( #12082 )
2025-09-09 14:45:13 +03:00
nimlgen and GitHub
1c6c42715f
unify cpu and llvm ( #11982 )
...
* try unify cpu and llvm
* fixes
* fix
* ops
* no llvm
* fix
* rm
* lvmm is ot
* oops
* override
* no llvm
* ignore
* skip llvm
* ooops
2025-09-09 13:54:44 +03:00
qazal and GitHub
50cc7175cb
viz: use complete progress helper ( #12081 )
...
* viz: use complete progress helper
* min diff
* rename show to start
2025-09-09 11:00:52 +03:00
Sieds Lykles and GitHub
239091d111
numba>=0.55 for uv resolution ( #12079 )
...
* force numba version
* update comment
2025-09-09 01:43:32 +02:00
chenyu and GitHub
2bd1fff79c
ci GPU misc cleanups ( #12078 )
2025-09-08 16:47:29 -04:00
chenyu and GitHub
1781d5bced
remove PYTHONPATH in test.yml ( #12077 )
...
set globally already
2025-09-08 15:41:47 -04:00
nimlgen and GitHub
9182948951
remove llvm_bf16_cast ( #12075 )
2025-09-08 20:51:15 +03:00
chenyu and GitHub
11213398b9
reorder amdremote in test yml ( #12073 )
2025-09-08 13:43:04 -04:00
nimlgen and GitHub
ebbcdd6577
cpu: use suppress_finalizing ( #12071 )
2025-09-08 18:28:09 +03:00
qazal and GitHub
73ca0e870c
viz: index visible rects ( #12070 )
2025-09-08 17:37:17 +03:00
chenyu and GitHub
d40f5b766b
default BEAM_PADTO to 0 ( #12069 )
...
seems incorrect, disable by default now
2025-09-08 10:17:03 -04:00
Sieds Lykles and GitHub
75b58fe2d3
move simplify_valid pat to sym ( #12065 )
...
* move simplify_valid pat to sym
* fix expectedfailure
2025-09-08 07:01:26 +02:00
chenyu and GitHub
56861852be
enable IMAGE for test_mnist and test_mnist_backward ( #12064 )
...
passes now
2025-09-07 09:06:39 -04:00
nimlgen and GitHub
ef71acc88a
hcq: cleanup fileio iface ( #12063 )
...
* hcq: cleanup fileio iface
* typo
* _
2025-09-07 15:43:27 +03:00
nimlgen and GitHub
35ddfc3d39
change default cpu_count ( #12062 )
2025-09-06 23:30:20 +03:00
nimlgen and GitHub
97187bf8b6
cleanup win and arch checks ( #12060 )
...
* cleanup win and arch checks
* stupid mypy
2025-09-06 23:08:46 +03:00
Sieds Lykles and GitHub
f326df8ae8
add type: ignore ( #12059 )
2025-09-06 21:17:35 +02:00
George Hotz and GitHub
c66935f7b9
only run hcopts once ( #12053 )
...
* only run hcopts once
* same?
2025-09-06 11:14:52 -07:00
qazal and GitHub
801be5f7b9
viz: memory graph cleanups ( #12057 )
...
* delete the total nbytes tooltip
* split pixel rescaling from layout
2025-09-06 19:44:53 +03:00
nimlgen and GitHub
10ac427aaa
cpu threading ( #11951 )
...
* start cpu threading
* fix
* fix2
* fix
* hacks?
* threads
* minor
* no dsp
* dsp 2
* n
* more
* test
* xm
* cleaner
* readable
* f
* reorder
* when no threads
* rangeify
* typos
* not needed
* reapply
* remoev this
* linter
* fixed cpu count in ci
* fix
* fixes
* rm
* typo
* sort based on speed
* test if test works in ci
* Revert "test if test works in ci"
This reverts commit 1f05edb531 .
* do not pad thread
2025-09-06 16:13:43 +03:00
nimlgen and GitHub
2b1844da27
cpu: support several threads in runtime ( #12055 )
2025-09-06 13:29:31 +03:00
nimlgen and GitHub
f37b836618
factor out _globalizable_rngs ( #12054 )
2025-09-06 13:29:23 +03:00
nimlgen and GitHub
1630c87d0e
run optimize_local_size only when locals supported ( #12056 )
2025-09-06 13:29:09 +03:00
Jordan Chalupka and GitHub
48ec5efad9
only run autogen tests on change ( #12049 )
...
* only run autogen tests on change
* example change
* rm example change
2025-09-05 23:53:01 -07:00
Sieds Lykles and GitHub
581b2388c2
add dtypes.index ( #12015 )
...
* add dtypes.index
* cast shape, stride and mask to dtypes.index in view.create
* move pm_lower_index_dtype to ops
* DEFINE_VAR is dtype.index by default
* merge var_val_using_str
* remove int from commutative
* fix test_rewrite_map
* change that to dtypes.index
* change some int to index
* shorten those
* remove old cast in renderer
* cleanup
* change that back
* add comment
* delete comment
* just delete those
* view doesnt have to cast anymore
* adjust comment
2025-09-06 06:03:44 +02:00
Sieds Lykles and GitHub
c6c16b2946
var_vals uses str for var (#12011 )
...
* var_vals is str,int
* remove imports
* remove print
* fix test
* change var_vals in hcq
* update test_hcq
* fix multitensor _device_num var
* fix syminfer test
* shorten line
* p.vars stays list[Variable]
* shorten line
* vars is back to tuple[Variable, ...]
* change var_vals in extra
* change var_vals from shapetracker
* var_vals is str:int
* fix signature
2025-09-06 04:16:12 +02:00
geohot
8658a97197
hotfix: name the shift rewrite better + no ctx there
2025-09-05 19:01:59 -07:00
George Hotz and GitHub
6ef3270fc8
fix opt gate ( #12050 )
2025-09-05 18:59:54 -07:00
geohot
66c5206b42
hotfix: minimal scheduler copy
2025-09-05 18:24:00 -07:00
geohot
478e758755
Revert "fix scheduler copy ( #12048 )"
...
This reverts commit 51b7c40788 .
2025-09-05 18:21:55 -07:00
George Hotz and GitHub
51b7c40788
fix scheduler copy ( #12048 )
...
* fix scheduler copy
* hand coded opt only runs once
2025-09-05 17:17:49 -07:00
George Hotz and GitHub
0123c394e5
early simplfy_merge_adjacent ( #12045 )
...
* do simplify_merge_adjacent before schedule
* do simplify_merge_adjacent before schedule
* disable that slow test
2025-09-05 16:39:20 -07:00
George Hotz and GitHub
8423c06144
delete unused bufs_from_lin ( #12044 )
2025-09-05 16:08:28 -07:00
George Hotz and GitHub
38dcadf07b
delete kernel.py ( #12040 )
...
* delete kernel.py
* delete that file
* rip and tear
* don't test search
* imports
* fix torch frontend
* not a part of regen
2025-09-05 15:52:07 -07:00
George Hotz and GitHub
ee4f696086
delete more tests ( #12043 )
...
* delete more tests
* delete and simplify
* flaky on windows
* a few more, those remained
2025-09-05 15:31:30 -07:00
George Hotz and GitHub
12c7b1bb01
cleanup lin tests without Kernel ( #12041 )
...
* cleanup lin tests without Kernel
* no kernel.py there
* remove that test
2025-09-05 15:13:14 -07:00
Sieds Lykles and GitHub
8435d2d23b
fix openpilot speed regeression ( #12039 )
...
* set local_size=None if special.arg[0]=='i'
* add cast back
2025-09-06 00:05:45 +02:00
George Hotz and GitHub
e00858a2c3
only POSTOPT ( #12038 )
2025-09-05 14:46:33 -07:00
George Hotz and GitHub
433581f8ed
make POSTOPT=2 the default ( #12034 )
...
* make POSTOPT=2 the default
* more matching tc
* fix winograd
* fix that test
* add matvec to Scheduler
* flip tc sort order
* similar speed
* fix beam on image
* disable slow tests
* slow
2025-09-05 14:34:05 -07:00
chenyu and GitHub
3b41a04b96
remove test_openpilot in test_onnx ( #12037 )
...
openpilot is tested in compile3
2025-09-05 16:20:03 -04:00
Sieds Lykles and GitHub
290521f68e
add check for z3>=4.12.4 ( #12035 )
2025-09-05 20:33:26 +02:00
George Hotz and GitHub
870f63d9cc
add WARP axistype, fix postopt bugs ( #12033 )
...
* postopt is 83% match
* warp is bright CYAN
* beautiful mnist beam works
* fix shutdown bug
2025-09-05 10:36:55 -07:00
chenyu and GitHub
4c2d4f683a
lower universal_test_unary cos domain ( #12032 )
...
flaky
2025-09-05 12:19:44 -04:00
chenyu and GitHub
a340723bf1
SKIP_SLOW_TEST=1 for nv CI ( #12031 )
2025-09-05 11:52:02 -04:00
chenyu and GitHub
ce7163e9b4
clean up skip slow tests in PYTHON ( #12028 )
...
skip with SKIP_SLOW_TEST and decorators
2025-09-05 11:35:26 -04:00
qazal and GitHub
f08299d2ec
viz: small profiler resizing improvements ( #12026 )
...
* switch to ResizeObserver
* set a fixed size for device-list
* less
* height from devices
* int
* side rect, more const
2025-09-05 18:29:03 +03:00
chenyu and GitHub
5dcc4c7f1b
skip test_linalg in windows unit test ( #12030 )
2025-09-05 11:28:40 -04:00
George Hotz and GitHub
f8e2dd4dd1
investigate opts mismatches ( #12020 )
2025-09-05 07:40:29 -07:00
chenyu and GitHub
e0da644171
lower sample count in test_multinomial ( #12027 )
2025-09-05 10:10:28 -04:00
chenyu and GitHub
9b6f1b86cb
add Tensor.maximum in test_dtype_alu ( #12025 )
...
works except nan
2025-09-05 09:48:39 -04:00
nimlgen and GitHub
3e1c04bcdf
jit: noopt for copy buffers ( #12023 )
2025-09-05 16:04:35 +03:00
qazal and GitHub
ab413ce72f
viz: give tooltips a max-width ( #12022 )
...
* viz: give tooltips a max-width
* better
2025-09-05 14:25:38 +03:00
qazal and GitHub
f461ccf407
exclude op2 nan lt in test_dtype_alu ( #12024 )
...
failure: https://github.com/tinygrad/tinygrad/actions/runs/17490320000/job/49679581331?pr=12022#step:6:125
2025-09-05 14:14:22 +03:00
nimlgen and GitHub
4fcea8493d
viz: add label to tooltip ( #12021 )
2025-09-05 13:06:33 +03:00
George Hotz and GitHub
2b5a73ac65
improve test_linearizer ( #12016 )
...
* improve test_linearizer
* tweaks
* simpler
* get_prg
* that one doesn't have to return
* fix postopt bugs
* fix rng
2025-09-04 20:44:05 -07:00
chenyu and GitHub
7f3df6ea21
exclude nan in test_dtype_alu lt ( #12019 )
2025-09-04 23:38:37 -04:00
Sieds Lykles and GitHub
f5404ca53c
Divmod combine - associative variations ( #12017 )
...
* add rule and test
* more rules and tests
* add all four variations
* fix test
* test fixed!
* adjust commment
* add new variations
* disable intel tensor core ops count test for bigger_matmul_half
2025-09-05 03:44:02 +02:00
chenyu and GitHub
677220ae7e
test_tesnor_data to unit/ ( #12013 )
2025-09-04 19:58:27 -04:00
George Hotz and GitHub
431666da74
POSTOPT=2 work ( #12012 )
...
* POSTOPT=2 work
* bugfixes
* add chain in one place
* tensor cores match
* better hcopt check
* match from old
* Change POSTOPT ContextVar value to 0
* we didn't need to check that
2025-09-04 16:55:56 -07:00
George Hotz and GitHub
30eb42a69e
fix POSTOPT pad ( #11999 )
...
* fix POSTOPT=1
* fix some tests
* Revert "fix some tests"
This reverts commit 8ee058e206 .
* fix padding restrictions
* cuda has two tensor cores
* Set POSTOPT ContextVar to 0 in helpers.py
2025-09-04 14:28:58 -07:00
qazal and GitHub
da61b40604
some viz tests don't need track_rewrites ( #12010 )
2025-09-04 23:59:32 +03:00
qazal and GitHub
be364a1adb
viz: add default tracing group ( #12009 )
...
This enables seeing rewrites in unit tests like `VIZ=1 python3 test/test_uop_graph.py TestUOpGraph.test_in_bounds_access_gated_local` that call graph_rewrite directly.
`@track_rewrites` keeps existing as an optional helper to organize larger traces.
2025-09-04 23:29:56 +03:00
chenyu and GitHub
52166fd7eb
smaller test_ops inputs ( #12007 )
2025-09-04 16:22:33 -04:00
chenyu and GitHub
dc8501af30
clean up wino tests ( #12008 )
...
removed the one that tests hcopt and added one for backward kernel counts
2025-09-04 16:14:55 -04:00
chenyu and GitHub
8c720e8760
less iterations for symbolic double for loops ( #12006 )
2025-09-04 15:09:17 -04:00
George Hotz and GitHub
70ce29b630
test pyrender ( #12005 )
...
* test pyrender
* make them print
* switch to pyrendered
2025-09-04 11:48:40 -07:00
George Hotz and GitHub
560df206cc
split tc test ( #12003 )
...
* split tc test
* split hand coded opts
* remove some skipped tests
* skips on emulated
2025-09-04 11:47:56 -07:00
qazal and GitHub
4996bb668b
load all traces before asserting in test_viz ( #12004 )
2025-09-04 21:34:48 +03:00
George Hotz and GitHub
9dee724fc4
make EMULATE a context var ( #12002 )
...
* make EMULATE a context var
* fix test amx
2025-09-04 11:15:43 -07:00
George Hotz and GitHub
09106e4aae
refactor and split test_linearizer ( #12001 )
...
* refactor and split test_linearizer
* forget that file
* imports
* remove from docs
* test gen float4
2025-09-04 10:53:07 -07:00
chenyu and GitHub
fb71d1e5fd
delete some test_search tests ( #11998 )
...
TC_SEARCH_OVER_SHAPE was removed so should the tests
2025-09-04 11:19:49 -04:00
chenyu and GitHub
ca7574cb2d
ci set PYTHONPATH for all ( #11997 )
2025-09-04 10:06:04 -04:00
nimlgen and GitHub
e213b85810
cpu: add thread_id to worker ( #11995 )
2025-09-04 14:58:13 +03:00
qazal and GitHub
35f37a64a9
viz: remove useless ctx.save and restore calls ( #11996 )
...
It's a UI no-op since we always set the styles right before drawing.
2025-09-04 14:56:41 +03:00
Sieds Lykles and GitHub
572a3c15c6
Move Ops.SPECIAL arg to src ( #11918 )
...
* initial moving bound to src
* arg to src
* remove import
* fixup linearizer
* arg to src
* fix test_uop_graph
* fix more tests
* fix python renderer
* get const value from const uop
* ssimplify uop estimates
* fix webgpu locals
* fix old test
* gate Ops.SPECIAL in linearizer
* use ssimplify() for local/global_size
* remove toposort gate_parents_instead_of_self
* fix rendering in comment
* cleanup
* rename and add comments
* add BottomUpGate with test
2025-09-04 09:31:44 +02:00
George Hotz and GitHub
5cf42dc4db
add Scheduler to replace Kernel with POSTOPT=2 ( #11924 )
...
* ** simple kernel to replace Kernel for postopt
* support old
* fix beam
* beaming
* beam on old
* bring tensor cores back
* raise
* postbeam
* test ops passes on mac
* skip that
* postopt default
* gate that
* fix tensor cores
* a few test fixes
* dsp fix
* tc fix
* loop
* support swap
* test_gemv
* fix beam for variable
* test opts from high level stuff
* range annoying
* compile slow
* metal slow
* better beam
* no POSTBEAM
* fix nolocals
* hc opt mostly works
* put that back
* lil
* some work
* fix that
* POSTOPT 2
* fix tests
* no postopt 2
* work
* back
* padded tensors cores
* shift_to
* postopt 0 passes?
* write PADTO
* fix padded tensor cores
* compare hcopt
* 18000 lines
* should pass tests
* fix rangeify
* put types back
2025-09-03 19:23:30 -07:00
chenyu and GitHub
b13e071463
move test_winograd to unit test ( #11993 )
2025-09-03 21:47:32 -04:00
chenyu and GitHub
edc8b99853
more tests that pass PTX now ( #11992 )
2025-09-03 21:18:14 -04:00
chenyu and GitHub
ed2f45712b
remove skip PTX in test_arange ( #11991 )
...
all passes now
2025-09-03 20:45:19 -04:00
George Hotz and GitHub
a5f2b4872a
use_tensor_cores is a heuristic ( #11989 )
...
* use_tensor_cores is a heuristic
* context
2025-09-03 17:05:10 -07:00
George Hotz and GitHub
63e930fec3
apply_tensor_cores is a heuristic ( #11988 )
...
* apply_tensor_cores is a heuristic
* delete extra_opts
2025-09-03 16:39:33 -07:00
chenyu and GitHub
d0e739453e
update many einsum tests ( #11981 )
...
correct the exception testing, and raise ValueError instead of assert when checking args
2025-09-03 15:40:20 -04:00
George Hotz and GitHub
55e4bdd353
split_uop is a method ( #11984 )
2025-09-03 10:46:17 -07:00
ttomsa and GitHub
1877eddde4
broadcast for upat ( #11940 )
2025-09-03 10:04:23 -07:00
George Hotz and GitHub
5ed262982a
remove some tc hacks from BEAM ( #11980 )
...
* remove some tc hacks from BEAM
* cosmetic changes
* revert that
2025-09-03 09:59:10 -07:00
6d53cac457
dtype fuzz: log need input > 0 ( #11979 )
...
Co-authored-by: b1tg <[email protected] >
2025-09-03 12:10:42 -04:00
Jordan Chalupka and GitHub
68e83b850f
nbytes should raise an exception when size is unlimited ( #11928 )
...
* nbytes should raise an exception when size is unlimited
* adding a test
2025-09-03 07:06:20 -07:00
Sieds Lykles and GitHub
86e908db57
cast parents of int64 alu to int32 if possible ( #11977 )
...
* add overflows helper
* add rules
* x -> y
* check overflow of u too
* cleaner
* use alu instead of replace to preserve vectorization
* just one rule
* add test
2025-09-03 11:05:04 +02:00
Sieds Lykles and GitHub
033184b3cb
parse_valid with non const rhs ( #11957 )
...
* const to using vmin/vmax
* add test
* convert to int
* remove left over part of and
2025-09-03 08:08:46 +02:00
Sieds Lykles and GitHub
53eff8970a
add Ops.GEP to _min_max ( #11976 )
2025-09-03 07:07:54 +02:00
Sieds Lykles and GitHub
d1d0960e6e
remove intermediate cast using bounds - weaker pattern ( #11974 )
2025-09-03 06:24:40 +02:00
Sieds Lykles and GitHub
8a2846b31a
assert embedding input is integer dtype ( #11963 )
...
* cast embedding input
* raise error if not using int for index embedding
2025-09-03 01:44:26 +02:00
wozeparrot and GitHub
d16cc6c012
feat: resume ckpt ( #11970 )
2025-09-02 15:47:48 -07:00
George Hotz and GitHub
1b73993521
pyrender to render uops ( #11968 )
...
* pyrender to render uops
* new pyrender style
* pyrender works
* list str
* store render
2025-09-02 15:44:01 -07:00
chenyu and GitHub
e921fb44ee
clean up testnvidia env ( #11969 )
2025-09-02 18:29:00 -04:00
chenyu and GitHub
69dd1817d0
raise RuntimeError in merge_dicts instead of assert [pr] ( #11965 )
2025-09-02 17:18:44 -04:00
qazal and GitHub
f750c15965
viz: add python marker ( #11952 )
...
* viz: add python marker
* remove duplicate
2025-09-02 23:44:00 +03:00
George Hotz and GitHub
550cf2ca7f
tests from postopt ( #11964 )
...
* tests from postopt
* reraise is fine
2025-09-02 13:34:17 -07:00
qazal and GitHub
b977ec0813
viz: axes domains cleanup ( #11962 )
2025-09-02 19:30:45 +03:00
nimlgen and GitHub
897254ad6c
ci: add dev<->cpu copy speeds ( #11959 )
2025-09-02 15:22:44 +03:00
George Hotz and GitHub
74040663bf
make ptrdtype a UOp property ( #11955 )
2025-09-01 16:35:43 -07:00
George Hotz and GitHub
0dfca4e74b
add failing test for rangeify setitem ( #11954 )
2025-09-01 16:24:35 -07:00
wozeparrot and GitHub
7c21271a5f
feat: end_lr envvar ( #11953 )
2025-09-01 14:53:07 -07:00
chenyu and GitHub
6a40216724
correct bf16 fuzz input in test_dtype_alu ( #11933 )
...
it was using float16 inputs, now it's uint16 then convert to bf16
2025-09-01 10:52:26 -04:00
chenyu and GitHub
965ea59b16
test_dtype_alu use AMD_LLVM from helpers ( #11950 )
2025-09-01 10:03:17 -04:00
a9f07c31bc
fix amd llvm sqrt ( #11936 )
...
* fix amd llvm sqrt
* lint
---------
Co-authored-by: b1tg <[email protected] >
Co-authored-by: chenyu <[email protected] >
2025-09-01 09:31:14 -04:00
qazal and GitHub
0a53e72f70
viz: fix trace duration in python test decoder ( #11949 )
2025-09-01 14:32:25 +03:00
qazal and GitHub
27c9ed5a84
viz: more consistent naming of events ( #11948 )
...
* s/shapes/events in test_viz
* s/bufs/events in the memory packer
2025-09-01 14:16:47 +03:00
qazal and GitHub
c7bb561ef9
remu: add v_rsq_f32_e32 instruction ( #11947 )
...
https://github.com/tinygrad/tinygrad/pull/11936 introduces a change to
the AMD LLVM renderer that outputs this instruction. Adding both 32 and
64 bit variants.
2025-09-01 11:29:31 +03:00
Sieds Lykles and GitHub
d9560a631c
remove cast between ints if safe ( #11946 )
2025-09-01 05:56:49 +02:00
Sieds Lykles and GitHub
a19d689481
fix vec dtype _min_max ( #11944 )
2025-09-01 03:24:07 +02:00
Sieds Lykles and GitHub
f32f3464d6
Can safe cast from certain ints to floats ( #11941 )
...
* add rule
* add some tests
* prevent infinite loop with bfloat16
* add some ints to double and float can_safe_cast
* add tests
2025-09-01 00:51:24 +02:00
Sieds Lykles and GitHub
1c6e43c203
Double cast is one cast if intermediate cast is safe ( #11939 )
...
* add rule
* add some tests
* prevent infinite loop with bfloat16
* prevent more infinite rewrite
2025-09-01 00:36:29 +02:00
wozeparrot and GitHub
7e68045fb2
feat: small llama3 training ( #11829 )
2025-08-31 13:41:47 -07:00
nimlgen and GitHub
020abe0556
hcq: finalize without synchronization when in error state ( #11872 )
...
* hcq: finalize without synchronization when in error state
* ooops
* fix
* fix
* fix
2025-08-31 18:39:13 +03:00
qazal and GitHub
2004c9757d
tracing: add default clock ( #11935 )
2025-08-31 18:24:44 +03:00
c1eeb3b99c
only skip AMD_LLVM ( #11934 )
...
Co-authored-by: b1tg <[email protected] >
2025-08-31 18:15:47 +03:00
75d380a77c
fix transcendentals in python renderer ( #11932 )
...
* fix transcendentals in python renderer
* add test
---------
Co-authored-by: b1tg <[email protected] >
2025-08-31 09:37:17 -04:00
Sieds Lykles and GitHub
61e4dc6ad5
render special arg in cstyle if arg is UOp ( #11931 )
2025-08-31 07:01:29 +02:00
Sieds Lykles and GitHub
d3252ccd85
fix special vmax when arg is UOp ( #11930 )
2025-08-31 06:54:39 +02:00
qazal and GitHub
0bacd9fc9b
viz: give disassembly its own node ( #11927 )
2025-08-31 00:28:52 +03:00
chenyu and GitHub
af89be317e
relax rtol for bfloat16 test_dtype_alu ( #11926 )
2025-08-30 17:16:08 -04:00
George Hotz and GitHub
632c2fb119
lowerer works on rangeifed + print exception ( #11925 )
2025-08-30 12:05:44 -07:00
qazal and GitHub
c27b99d68f
viz: refactor to indexed rewrite traces ( #11923 )
2025-08-30 20:01:47 +03:00
qazal and GitHub
9aff00a6ea
switch viz command line args to pathlib ( #11922 )
2025-08-30 18:13:47 +03:00
qazal and GitHub
c86ee5bfaf
viz: canonicalize device name colors ( #11921 )
2025-08-30 18:12:30 +03:00
nimlgen and GitHub
a4f05ebd1a
ci: rebuild gpuocelot with boost libs ( #11920 )
2025-08-30 17:24:19 +03:00
qazal and GitHub
bf0d055b39
viz: color by name ( #11919 )
2025-08-30 16:04:58 +03:00
Sieds Lykles and GitHub
0bc34c000f
simplify range mod its own upper bound ( #11917 )
...
* add rules
* add tests
2025-08-30 08:37:35 +02:00
chenyu and GitHub
561318fea7
Tensor.cos in test_stype_alu ( #11916 )
...
* Tensor.cos in test_stype_alu
* need this fix anyway
2025-08-29 20:26:36 -04:00
0838021753
remove np from beautiful_cifar ( #10988 )
...
* remove np from beautiful_cifar
* remove np from cifar
* rename variable and rename tensor.arrange to just tensor.randperm
---------
Co-authored-by: chenyu <[email protected] >
2025-08-29 19:34:16 -04:00
nimlgen and GitHub
cf9d8c8142
ci: pin boost for macos runners ( #11910 )
2025-08-30 01:38:06 +03:00
nimlgen and GitHub
c6e342cdac
mockgpu: no hang if gpuocelot failed ( #11915 )
2025-08-30 00:44:49 +03:00
chenyu and GitHub
26d03a86a1
test_symbolic_ops.py cleanup ( #11895 )
2025-08-29 17:11:59 -04:00
b2cc06218a
python bfloat16 ( #11912 )
...
* python bf16
* _to_torch_storage_type
---------
Co-authored-by: b1tg <[email protected] >
2025-08-29 15:18:02 -04:00
George Hotz and GitHub
afad7d0cd1
remove dtype from range, it will be dtypes.index soon [pr] ( #11914 )
...
* remove dtype from range, it will be dtypes.index soon [pr]
* a few more
2025-08-29 09:52:07 -07:00
qazal and GitHub
30e72d5820
multi device and copy tracing for NULL device ( #11913 )
...
* add device name to NULL programs
* trace transfers
2025-08-29 15:31:00 +03:00
qazal and GitHub
d8e1e4dc61
tracing: show NULL programs ( #11911 )
2025-08-29 14:09:33 +03:00
nimlgen and GitHub
75678b2cbe
amd: retire pm4 xcc sync ( #11835 )
...
* amd: aql default when several xccs
* amd: retire om4 xcc sync
* remove more
* more
* more
2025-08-29 09:56:27 +03:00
George Hotz and GitHub
394c2d1db1
update Kernel API in tests + move optimize_local_size ( #11907 )
2025-08-28 15:12:47 -07:00
nimlgen and GitHub
fa695ac1ce
ci: mac gpuocelot ( #11906 )
...
* gm
* fix?
* ops
* imp
* xx
* add file
2025-08-28 23:29:43 +03:00
George Hotz and GitHub
b9b438c516
small updates from postopt ( #11903 )
...
* tests from postopt
* modernize
* skip lin tests
* that's fixed?
* skip, not failure
2025-08-28 12:34:52 -07:00
nimlgen and GitHub
bb55a3001f
nv: flush reset message ( #11897 )
2025-08-28 22:17:20 +03:00
nimlgen and GitHub
e8289c75b1
ci: do not reinstall existing pkgs in macos ( #11900 )
2025-08-28 21:20:15 +03:00
chenyu and GitHub
134cf56904
update cache name for gpuocelot ( #11896 )
2025-08-28 13:11:10 -04:00
ea1be2e4cd
[bounty] Remove using reshape to register symbolic shape ( #11771 )
...
* Modify tests and start work towards removing symbolic reshape
* Refactor symbolic reshape
* fix small error
* much cleaner + fix more tests
* Can remove this now
* Update test_symbolic_ops and test_tiny
* Couple more tests
* Unused import
* More tests and add EXPAND to Tensor.empty
* Fix test beam search
* all int
* Fix rangeify by adding shrink
* Remove OOB check and so fix test_symbolic_jit
* test_symbolic_jit doesn't need OOB Context anymore either
* Should remove that test now
* Cleanups part 1
* fix linters
* Final cleanups
* Don't reassign inside for loop
---------
Co-authored-by: chenyu <[email protected] >
2025-08-28 12:30:49 -04:00
qazal and GitHub
53853ae49b
viz: switch to Path2D ( #11892 )
2025-08-28 18:58:16 +03:00
nimlgen and GitHub
874c1db4af
am: init support for aql ( #11888 )
2025-08-28 18:41:46 +03:00
17ecaf4682
Add test_variable_empty ( #11889 )
...
* Add test_variable_empty
* Move test and add TODO
---------
Co-authored-by: chenyu <[email protected] >
2025-08-28 11:38:27 -04:00
Nino Risteski and GitHub
54be477152
rope cache optim for jit prune in llm.py ( #11678 )
...
* rope cache optim for jit prune
* rope test
* tests in test attention
* Revert "rope test"
This reverts commit 69ede543d0 .
* lint
2025-08-28 08:31:29 -07:00
quortus and GitHub
5f8fe9a331
Replace ASSIGN with STORE in test_linearizer ( #11821 )
2025-08-28 07:33:20 -07:00
4e8370309c
Support onnx If OP ( #11648 )
...
* start
* tiny clean up
* whoops, didn't mean to accidentally fix this
* fix .to(device), kinda hacky and this fix makes it slower?
* merge properly
* FINALLY figured out slowness, also hack pylint for now
* add DEBUGONNX print for subgraph
* oops
* WOOOOOOOO SHAPE CACHE 50% SPEED INCREASE
* small fix, but maybe all deterministic Tensor creation in fp should be cached
* cache condition
* sliiiightly cleaner
* better abstraction?
* remove sam from model_benchmark
* remove shape cache speed up for now
* less lines
* isinstance fix
---------
Co-authored-by: chenyu <[email protected] >
2025-08-28 10:17:35 -04:00
George Hotz and GitHub
6d6f0dada7
support for tuple ranges ( #11890 )
...
* support for tuple ranges
* breaks it
2025-08-28 07:02:31 -07:00
nimlgen and GitHub
60dd9a162c
memory: tiny tlsf cleanup ( #11887 )
2025-08-28 14:07:18 +03:00
chenyu and GitHub
beb5982165
FUSE_ATTENTION ( #11884 )
2025-08-27 19:59:17 -04:00
George Hotz and GitHub
cb5295168d
postrange boilerplate work ( #11881 )
2025-08-27 15:22:59 -07:00
George Hotz and GitHub
fd579433bc
pre expander shouldn't go in gpudims ( #11880 )
2025-08-27 14:52:24 -07:00
nimlgen and GitHub
44816218b5
memplan: fix large buffers planning ( #11878 )
...
* memplan: fix large buffers planning
* fix
* fix dsp
2025-08-27 23:54:27 +03:00
nimlgen and GitHub
4006366752
Revert "memplan: fix large buffers planning ( #11876 )" ( #11877 )
...
This reverts commit 7f90497efc .
2025-08-27 22:36:14 +03:00
nimlgen and GitHub
7f90497efc
memplan: fix large buffers planning ( #11876 )
...
* memplan: fix large buffers planning
* fix
2025-08-27 22:04:15 +03:00
George Hotz and GitHub
e4afdf9ea1
improve DEBUG=2 string with TB/s and TFLOPS [pr] ( #11875 )
2025-08-27 11:42:41 -07:00
Jordan Chalupka and GitHub
e9789d8a70
Add mxfp4 support ( #11873 )
...
* bump ggml url
* map mxfp4 to tensor
* tests
2025-08-27 10:56:56 -07:00
qazal and GitHub
884eb53e89
tracing: fix types ( #11871 )
...
* tracing: fix types
* /profiler isn't a thing
* return list
2025-08-27 15:50:43 +03:00
Sieds Lykles and GitHub
d39365809a
add ctx to z3_renderer arg ( #11867 )
...
* add ctx to z3_renderer arg
* update symbolic fuzzer
* rewrite u1,u2,u3
* update fuzz_fast_idiv
* remove imports
2025-08-27 03:38:15 +02:00
George Hotz and GitHub
24c00a4061
darken hex on viz ( #11865 )
...
* darken hex on viz
* more readable
2025-08-26 15:57:50 -07:00
qazal and GitHub
f38e4af226
viz: add custom zoom filter ( #11861 )
2025-08-27 01:30:29 +03:00
nimlgen and GitHub
62df6c39af
amd: correct handling of relocations ( #11863 )
...
* amd: correct handling of relocations
* ops
* add
2025-08-27 01:26:45 +03:00
George Hotz and GitHub
d261458ecd
add colors to range ( #11860 )
2025-08-26 14:32:12 -07:00
Sieds Lykles and GitHub
7dfc7e4abc
uops_to_z3 helper( #11859 )
2025-08-26 22:58:05 +02:00
chenyu and GitHub
1bbb578afd
named expression for POW and MAX gradient ( #11858 )
2025-08-26 16:03:03 -04:00
chenyu and GitHub
7028cb4167
clean up TestBitcastConstFolding ( #11856 )
2025-08-26 15:26:47 -04:00
George Hotz and GitHub
d4154e0349
split devectorizing of buf/index ( #11855 )
2025-08-26 12:05:48 -07:00
George Hotz and GitHub
b268755d51
small changes from postopt ( #11854 )
2025-08-26 11:56:16 -07:00
Sieds Lykles and GitHub
a3aeef45cc
associative variation of where branch-merging ( #11851 )
...
* add rule and test
* change comment
2025-08-26 19:27:05 +02:00
chenyu and GitHub
aabe7756be
fix type in fold_bitcast [pr] ( #11853 )
2025-08-26 13:22:30 -04:00
Jordan Chalupka and GitHub
4785cd959a
[TYPED=1] cvar should allow dtype as a tuple ( #11770 )
...
* cvar dtype:DType|tuple[DType, ...]|None=None
* fmt
* add a test
* list typeguard as a dep for CI
* extra step to install mypy
* fix venv
* ci fixes
* mv typeguard to testing install group
* simpler TYPED=1 test
* add typeguard to lint group
2025-08-26 12:49:51 -04:00
qazal and GitHub
b111076301
viz: fixup click on overlay rect ( #11850 )
2025-08-26 19:25:42 +03:00
1dd613cb89
test float_to_bf16 round-to-even behavior ( #11849 )
...
Co-authored-by: b1tg <[email protected] >
2025-08-26 12:16:10 -04:00
409399c609
fix nan in float_to_bf16 ( #11843 )
...
Co-authored-by: b1tg <[email protected] >
2025-08-26 11:42:25 -04:00
qazal and GitHub
43d5d66d34
viz: add UOp ports to edges ( #11847 )
...
* viz: add UOp ports to edges
* one edge label
* g.tag styling
* replace with NodeList
2025-08-26 18:31:52 +03:00
chenyu and GitHub
f28f613f85
improved float_to_bf16 ( #11848 )
...
round instead of truncate
2025-08-26 11:14:06 -04:00
nimlgen and GitHub
afe14ccbfa
amd: aql default when several xccs ( #11832 )
2025-08-26 15:16:36 +03:00
qazal and GitHub
3674c0754e
viz: small uop click changes ( #11846 )
...
* also highlight self
* can always unselect by clicking outside
* less layout
2025-08-26 14:56:13 +03:00
qazal and GitHub
f2a3c27372
viz: g.edges() once ( #11845 )
2025-08-26 13:29:59 +03:00
qazal and GitHub
b0df3e62a8
viz: light up srcs and paths on UOp click ( #11844 )
...
* viz: light up srcs and paths on UOp click
* safari doesn't have context-stroke
* safari also has a bug
* safari acceptance
2025-08-26 09:03:09 +03:00
qazal and GitHub
6236749867
viz: move rect styles to classes ( #11842 )
...
* viz: move rect styles to classes
* add rect
2025-08-26 07:55:34 +03:00
qazal and GitHub
81ffa07439
viz: pass through nodes without a link ( #11841 )
2025-08-26 07:00:43 +03:00
Sieds Lykles and GitHub
265d287615
add decomp for !x&!y -> !(x|y) ( #11836 )
2025-08-26 05:21:06 +02:00
chenyu and GitHub
337e979a59
call dtypes.as_const in Tensor(list) ( #11840 )
2025-08-25 22:08:26 -04:00
George Hotz and GitHub
215818379b
new (post) group for reduce ( #11837 )
...
* new (post) group for reduce
* fixes
* leave if
* fix locals
* size
* no vectorized buf
* image fixes
* don't track that
* fix ptx
* name buffer with reduce range
* remove unused in lowerer
* yay DEFINE_REG refactor
2025-08-25 18:03:00 -07:00
chenyu and GitHub
ac3449b0c8
truncate_fp16 cleanup ( #11838 )
...
native `@` is default
2025-08-25 19:03:41 -04:00
qazal and GitHub
e146418f65
hotfix: profiler content-type is application/octet-stream ( #11831 )
2025-08-25 15:56:42 +03:00
qazal and GitHub
a1f6823060
viz: memory layout in client side ( #11830 )
...
* viz: memory layout in client side
* update test_viz
2025-08-25 14:49:33 +03:00
George Hotz and GitHub
a6dbb09058
changes for postrange ( #11828 )
2025-08-24 17:37:07 -07:00
George Hotz and GitHub
27701ef823
add locals support to rangeify ( #11826 )
2025-08-24 14:03:12 -07:00
Sieds Lykles and GitHub
a286a1a6f7
Fast idiv try removing factors of two before cast ( #11824 )
...
* try removing factors of two
* dont return if None
* add test
2025-08-24 20:04:25 +02:00
geohot
a03b930339
hotfix: green v2 in docs
2025-08-24 10:25:14 -07:00
George Hotz and GitHub
6540bb32a6
move into codegen late [pr] ( #11823 )
2025-08-24 10:23:25 -07:00
nimlgen and GitHub
bba088ef11
amd aql queue ( #11708 )
...
* amd aql queue
* xcc
* fiz
* aql better
* llvm
* no for aql
* wrap
* is_sql
* am support
* complete
* fix
* mypy
* minor
2025-08-24 19:53:00 +03:00
George Hotz and GitHub
1fa09d9ede
BLOCK_REORDER is context var, heuristic cleanups [pr] ( #11819 )
...
* BLOCK_REORDER is context var, heuristic cleanups [pr]
* split get opt and do opt
* oops, should be on
2025-08-24 09:41:34 -07:00
qazal and GitHub
8b18cc2a94
viz memory layout cleanup ( #11820 )
...
* rename to dtype_size
* cleanr memory shape creator
2025-08-24 19:37:31 +03:00
Sieds Lykles and GitHub
dd69114573
Revert "Better div nesting ( #11811 )" ( #11818 )
...
This reverts commit 952f729b07 .
2025-08-24 18:11:24 +02:00
nimlgen and GitHub
e19f901330
amd: rptr/wptr in create_queue ( #11817 )
2025-08-24 18:03:45 +03:00
nimlgen and GitHub
d71444857e
amd: apply relocs for kernel_code_entry_byte_offset for AMD_LLVM ( #11816 )
...
* amd: apply relocs for kernel_code_entry_byte_offset for AMD_LLVM
* fix
2025-08-24 17:48:40 +03:00
George Hotz and GitHub
44bc7dc73d
remove KernelInfo from GROUP_REDUCE ( #11814 )
2025-08-23 19:55:41 -07:00
George Hotz and GitHub
229adfb7c3
Revert "remove KernelInfo from gpudims ( #11809 )" ( #11813 )
...
This reverts commit 846753f343 .
2025-08-23 19:37:10 -07:00
Sieds Lykles and GitHub
952f729b07
Better div nesting ( #11811 )
...
* remove check
* use fold_divmod_congruence instead of simplify
* adjust tests
* shorten line
2025-08-24 04:17:40 +02:00
Sieds Lykles and GitHub
e652062f92
tweak divmod_folding condition ( #11810 )
2025-08-24 02:59:02 +02:00
George Hotz and GitHub
846753f343
remove KernelInfo from gpudims ( #11809 )
...
* remove KernelInfo from gpudims
* that's good in there
2025-08-23 16:32:45 -07:00
Sieds Lykles and GitHub
07d4ed7e4c
one more symbolic add variation ( #11807 )
2025-08-24 01:15:04 +02:00
qazal and GitHub
759ebea4eb
viz: reflect timeline API boundary in names ( #11808 )
...
* define shapes once
* depth isn't an event property
* update server naming
2025-08-24 02:12:12 +03:00
George Hotz and GitHub
132f09fab7
global/locals from AxisType in range ( #11806 )
2025-08-23 15:49:17 -07:00
qazal and GitHub
0d86288bd7
viz: calculate timeline fixed points in client side ( #11805 )
...
* viz: calculate timeline fixed points in client side
* 26 bytes / event
* math
2025-08-24 01:44:40 +03:00
George Hotz and GitHub
a75da49951
use AxisType for UPCAST/UNROLL ( #11800 )
...
* use AxisType for UPCAST/UNROLL
* fixes
* fix the bug
* fix hack
* bad test
* flaky test
2025-08-23 14:44:48 -07:00
qazal and GitHub
2407fecdae
viz bytepack format ( #11792 )
...
* viz bytepack format
Training a 1B llama yields ~20M profiler events.
With JSON serialization, the browser tries to load 6GB to memory. This OOMs since each tab is limited to <3-4GB memory usage. Using a packed format, we only need ~600MB.
**Design decisions:**
- Timestamps are in microseconds relative to start time. They're stored in u32, which can express up to ~1 hr of trace events.
- Strings (kernel names, metadata, etc) are deduped.
- Buffer sizes are in u64 nbytes.
More optimization possible:
- The string lookup is a JSON dumped array, we can compress this.
- Can store less for memory by moving the layout to client.
**Results**
| | Events | JSON | bytepack |
|----------------|---------|-------------|-------------|
| DP=8 llama 1B train (`command: [1]`) | 24M | 5.8GB | 640MB |
| examples/beautiful_mnist.py | 16K | 3.7MB | 745KB |
| examples/gpt2.py | 55K | 12.54MB | 1.40MB |
`[1]`: `VIZ=1 FAKEDATA=1 OFFLOAD_OPTIM=1 DP=8 BS=8 GRADIENT_ACC_STEPS=2 BLOCK_REORDER=0 LR=3e-4 TRAIN_ON_VAL=1 DEFAULT_FLOAT=bfloat16 OPTIM_DTYPE=bfloat16 LLAMA3_SIZE=1B WARMUP_STEPS=36 DECAY_STEPS=360 SEQLEN=8192 PYTHONPATH=. AMD=1 AMD_LLVM=0 MODEL=llama3 python3 examples/mlperf/model_train.py`
* python reference decoder
* 27 bytes / event, 1hr hard limit
2025-08-23 23:50:21 +03:00
qazal and GitHub
b12d1d866c
count bytes per kernel in test_viz ( #11801 )
...
Currently at ~100 bytes/kernel with JSON.
2025-08-23 23:35:27 +03:00
Sieds Lykles and GitHub
6a50ab6b87
adjust idiv min_max ( #11802 )
...
* change div min_max
* add tests
2025-08-23 22:25:51 +02:00
chenyu and GitHub
9d4cccd0f9
test_dtype_alu cleanups ( #11799 )
2025-08-23 15:11:17 -04:00
George Hotz and GitHub
aefabaf774
add AxisType to range ( #11798 )
...
* add AxisType to range
* missed them
* fix that test
* fix that test
2025-08-23 11:15:00 -07:00
qazal and GitHub
b975830424
add profile loader helper in test_viz ( #11797 )
2025-08-23 19:20:29 +03:00
chenyu and GitHub
7123df3928
Use Tensor.logaddexp to implement Tensor.softplus ( #11796 )
...
instead of piecewise linear, numerical is handled by logaddexp. jax does this and i think it's more elegant than torch's approach
2025-08-23 11:52:29 -04:00
qazal and GitHub
aaea6b97ad
viz memory: compute nbytes ( #11795 )
...
* viz memory: compute nbytes
* local map
2025-08-23 17:34:07 +03:00
qazal and GitHub
58653b5eae
viz: store memory scale ( #11794 )
2025-08-23 16:19:44 +03:00
chenyu and GitHub
fb8ee02424
Tensor.logaddexp ( #11793 )
2025-08-23 09:15:00 -04:00
Sieds Lykles and GitHub
5a6817d5f8
Fix z3 rendering of floats in indexing ( #11740 )
...
* Fix floating point comparison in indexing
* wrap in noop
* update tests
* improve rules for loading and comparing floats
* add test cast to bool
2025-08-23 05:56:19 +02:00
chenyu and GitHub
4267c45db3
non-supported dtype in transcendental ( #11754 )
...
* non-supported dtype in transcendental
`CPU=1 python3 test/test_dtype_alu.py TestDTypeALU.test_bfloat16_unary` works
* test
* works on real mac
2025-08-22 23:13:45 -04:00
chenyu and GitHub
e39b25cd36
upcast float exp to at least float32 ( #11758 )
...
* upcast float exp to at least float32
* unlucky seed
2025-08-22 20:16:34 -04:00
nimlgen and GitHub
b057a90d49
memory: rename is_huge_page -> is_page ( #11786 )
2025-08-22 20:08:58 +03:00
qazal and GitHub
38f0fa7bde
viz: only send trace duration ( #11789 )
...
* viz: only send trace duration
* can unwrap
2025-08-22 20:00:48 +03:00
qazal and GitHub
1c81ec9248
viz: rename to start/end timestamp ( #11788 )
2025-08-22 19:47:49 +03:00
qazal and GitHub
9ff03680ba
viz: store relative timestamps ( #11787 )
...
* viz: store relative timestamps
* err
* update test
2025-08-22 19:30:21 +03:00
nimlgen and GitHub
698392334f
system: message for eaccess as well ( #11785 )
2025-08-22 18:21:32 +03:00
geohotstan and GitHub
1e679bd789
fix max_unpool2d inf ( #11784 )
...
* start
* add regression test for maxunpool2d
2025-08-22 08:31:24 -04:00
George Hotz and GitHub
9832599c9e
test_vmap + permute isn't a sint ( #11783 )
...
* test_vmap + permute isn't a sint
* order
2025-08-21 22:39:35 -07:00
George Hotz and GitHub
bb8de51e5f
remove unused early cleanups + contig w range [pr] ( #11780 )
...
* remove unused early cleanups [pr]
* contiguous with range
* woah, this works
2025-08-21 20:04:45 -07:00
chenyu and GitHub
91a4de4ca7
fix getitem with inf in tensor ( #11781 )
2025-08-21 21:55:32 -04:00
George Hotz and GitHub
66e9d54eed
RANGEIFY=2 is partial contig ( #11777 )
2025-08-21 16:53:58 -07:00
Jordan Chalupka and GitHub
8de6db15ac
exclude .git from ruff ( #11773 )
2025-08-21 15:37:50 -07:00
George Hotz and GitHub
5954a0975f
fix some assigns on rangeify ( #11774 )
...
* fix some assigns
* llvm test
* more tests
* upd test
2025-08-21 15:15:54 -07:00
qazal and GitHub
2e0eb88549
viz: add metadata to UOp tracing ( #11772 )
...
* viz: add metadata to UOp tracing
* place after tag
* optional field
* err, refcount of root must be 0
2025-08-22 00:18:45 +03:00
George Hotz and GitHub
d6f9606e93
small cleanups to rangeify ( #11769 )
2025-08-21 11:15:09 -07:00
bd4a9473b0
Multihost exception handling ( #11729 )
...
Co-authored-by: wozeparrot <[email protected] >
2025-08-21 13:51:49 -04:00
George Hotz and GitHub
a2c7b807e0
don't bufferize 0s ( #11766 )
2025-08-21 10:10:56 -07:00
nimlgen and GitHub
9eff7cd1d8
am: support 64bit discovery ( #11768 )
2025-08-21 18:28:13 +03:00
56cd47a159
fix amd llvm bf16 tc ( #11713 )
...
* fix amd llvm bf16 tc
* is_cdna
---------
Co-authored-by: b1tg <[email protected] >
Co-authored-by: chenyu <[email protected] >
2025-08-21 09:33:28 -04:00
George Hotz and GitHub
a044648111
rangeify load cleanups + multi support ( #11765 )
...
* use the old buf_uop + cleanups
* simpler handling of load
* everything needed for multi too
2025-08-20 20:55:49 -07:00
George Hotz and GitHub
9f94c25a25
fix symbolic usage. use shrink, not reshape ( #11762 )
...
* fix test_var
* revert those things
* fix the ones in test tiny
* use better syntax
* it's the same, but that's clearer
* fix pad
2025-08-20 18:35:42 -07:00
chenyu and GitHub
5276fbc9c5
fix gather with inf values ( #11760 )
...
(mask * x) is wrong because 0*inf is nan. i feel we have a lot of those still...
2025-08-20 20:35:40 -04:00
wozeparrot and GitHub
b979162c5d
llama3 eval train ( #11706 )
2025-08-20 19:56:35 -04:00
chenyu and GitHub
dbd3b67657
clamp GRAD_CLIP_NORM in llama ( #11761 )
2025-08-20 19:55:50 -04:00
George Hotz and GitHub
9635592141
** rangeify, try 3 ( #11683 )
...
* ** rangeify, try 3
* bring that over
* bufferize, don't use contig tag
* work
* ish
* fix rangeify
* flash attention is back
* fix rangeify tests
* stuff passes
* fix test_log_softmax
* more stuff passes
* progress children
* new endrange solution
* progress
* progress counter
* basic assign
* contigs only
* symbolic in schedule
* unbind_kernel
* late children
* ops fixed
* beautiful mnist is close
* that seems to work
* mnist works
* improve names
* fix bmnist
* no pcontig
* testing backward
* work
* clone movement ops
* new_range helper
* MBLOCK/MERGE
* ops tests pass
* revert mblock stuff
* cleanups...but it breaks ops
* remove reindex
* hack for relu
* disable the hacks
* more hacks
* upd
* mostly works with cleanups disabled
* ndr
* ops tests pass
* terrible hacks for indexing to work
* context mismatch
* pcontig
* split pcontig v contig
* z3 trunc
* null
* no fuse in rangeify
* ops test passes
* lnorm
* fix assign
* nd rangeify
* both should work
* tests for rangeify
* cleanups
* stores pass the pointer through
* disable pcontig for now
* PARTIAL_CONTIG is a flag
2025-08-20 14:22:44 -07:00
chenyu and GitHub
d7553721d1
clean up test_dtype_alu ( #11757 )
...
remove the check that looks into schedule, only test if output matches
2025-08-20 14:36:18 -04:00
chenyu and GitHub
5f08a3e928
hotfix: cast half to float in Tensor.tolist ( #11755 )
...
workaround for python < 3.12
2025-08-20 12:18:35 -04:00
qazal and GitHub
de4cb722a4
viz: add metadata and var_vals tracing ( #11753 )
...
* viz: add metadata and var_vals tracing
* add test_trace_metadata
* set TRACEMETA=1
2025-08-20 18:39:51 +03:00
nimlgen and GitHub
6589c9e643
hcq: better errors for ifaces ( #11751 )
...
* hcq: better errors for ifaces
* fix linter
* typo
* space
2025-08-20 17:50:51 +03:00
chenyu and GitHub
be7b0b6970
TRANSCENDENTAL_SUPPORTED_DTYPES->TRANSCENDENTAL_DTYPES ( #11752 )
2025-08-20 10:29:36 -04:00
ttomsa and GitHub
220a2a88d7
a*(1/b) -> a/b on LLVM, CPU ( #11743 )
...
* add fdiv rewrite
* :)
* use float_lop
* use reciprocal()
* revert
* move to decompositions
2025-08-20 09:35:10 -04:00
George Hotz and GitHub
12ab3f8b06
correct row_count in process replay ( #11748 )
2025-08-19 22:21:07 -07:00
George Hotz and GitHub
8af8808c61
cleanup tests, bump caches ( #11746 )
2025-08-19 21:21:07 -07:00
George Hotz and GitHub
00391db628
no ast for mem estimate ( #11744 )
...
* no ast for mem estimate
* skip for webgpu
2025-08-19 20:18:45 -07:00
chenyu and GitHub
dd413e1208
remove a Ops.REDUCE check in reduce_collapse [pr] ( #11734 )
2025-08-19 19:21:28 -04:00
ttomsa and GitHub
70c3f1fb29
x.where(False, True) -> !x ( #11738 )
...
* add pat
* add test
2025-08-19 19:08:16 -04:00
George Hotz and GitHub
1d307f568c
move device tests to test/device + test cleanups ( #11735 )
...
* move device tests to test/device
* test speedups
* test device
* linalg to unit
* upd
* so pytest just works
* more divide and skip
* speed
* test devectorize
* add pillow
2025-08-19 16:02:20 -07:00
wozeparrot and GitHub
bcc7623025
feat: bump version to 0.11.0 ( #11736 )
2025-08-19 17:08:56 -04:00
qazal and GitHub
8c987b3293
DISABLE_FAST_IDIV is a context var [pr] ( #11733 )
2025-08-19 23:30:50 +03:00
George Hotz and GitHub
bf467c623d
changes from rangeify + better NullRenderer ( #11732 )
...
* changes from rangeify + better NullRenderer
* fix test
2025-08-19 12:51:54 -07:00
chenyu and GitHub
02353588cb
small getitem cleanup ( #11730 )
2025-08-19 12:25:58 -04:00
chenyu and GitHub
712a5c651a
minor Tensor.triu cleanup ( #11728 )
...
less confusing dtype
2025-08-19 08:07:38 -04:00
nimlgen and GitHub
9c9e337c78
amd: parse soc enums ( #11727 )
...
* amd: parse soc enums
* remove from mock
* fix
* minimal amd_gpu
2025-08-19 15:06:09 +03:00
qazal and GitHub
57ad69160a
viz: inline memory shape spec ( #11725 )
2025-08-19 08:03:29 +03:00
chenyu and GitHub
c5b52e9321
onnx RotaryEmbedding cleanup ( #11724 )
2025-08-18 23:34:42 -04:00
George Hotz and GitHub
31619774a9
Revert "Revert "fix the misused cast in amd llvm tc ( #11711 )" ( #11715 )" ( #11723 )
...
This reverts commit ca28db5a97 .
2025-08-18 19:44:35 -07:00
2ea54d7337
improve syntax of UPats using f [pr] ( #11717 )
...
Co-authored-by: chenyu <[email protected] >
2025-08-18 20:49:45 -04:00
chenyu and GitHub
b67345caa3
use truncate in onnx read_int64 [pr] ( #11720 )
2025-08-18 20:49:35 -04:00
qazal and GitHub
50e789e290
hotfix: add device to decompositions ctx ( #11721 )
...
fast_idiv requires it for checking if a dtype is supported. Without
this, codegen creates non reproducible output without a complete
os.environ. since `is_dtype_supported` will open devices based on the
env var unless the device is specified by the caller.
2025-08-19 03:31:16 +03:00
George Hotz and GitHub
4b3fcb4064
Revert "REDUCE_AXIS keepdim=False ( #11311 )" ( #11718 )
...
This reverts commit b518a7378a .
2025-08-18 13:28:53 -07:00
George Hotz and GitHub
67d0ba5bd8
new ops from rangeify ( #11716 )
2025-08-18 13:13:11 -07:00
geohot
4afa0b86bb
hotfix: ls -lh on wheel size
2025-08-18 11:52:59 -07:00
George Hotz and GitHub
ca28db5a97
Revert "fix the misused cast in amd llvm tc ( #11711 )" ( #11715 )
...
This reverts commit 799a637b03 .
2025-08-18 11:51:28 -07:00
chenyu and GitHub
c10e4c4e20
print wheel build size ( #11714 )
2025-08-18 14:29:47 -04:00
b518a7378a
REDUCE_AXIS keepdim=False ( #11311 )
...
* progress
* fix tests
* fix tests
* remove hack for test_symfold
* fix test_conv.py on llvm
* hack test_cache_speed
* lint
* remove hack for helper_linearizer_opt
* tests
* fix DSP
* clean up
* remove hack for kernelize.py
* hack for test/test_multitensor.py TestMultiTensor.test_matmul_shard_none
* clean
* uop.r need reshape?
* lower_store cause fail
* fix lower?
* avoid contiguous hack
* 2134
* conv2d count
* remove unused
* hack lower
* reduced and clean up
* fix TestMultiTensor.test_matmul_shard_none
* src sync + fix TestMultiTensor.test_matmul_shard_none
* remove excluded in mop
---------
Co-authored-by: b1tg <[email protected] >
Co-authored-by: George Hotz <[email protected] >
Co-authored-by: nimlgen <[email protected] >
2025-08-18 10:09:17 -07:00
61884f2057
add cstyle renderer to the NULL device ( #11709 )
...
Co-authored-by: b1tg <[email protected] >
2025-08-18 09:52:22 -07:00
18db8fa311
Allow choosing leaders in multinode reduce ( #11506 )
...
Co-authored-by: wozeparrot <[email protected] >
2025-08-18 12:43:20 -04:00
799a637b03
fix the misused cast in amd llvm tc ( #11711 )
...
Co-authored-by: b1tg <[email protected] >
2025-08-18 09:15:34 -07:00
qazal and GitHub
fef97547f9
viz: preset the final timestamp ( #11712 )
2025-08-18 17:51:21 +03:00
chenyu and GitHub
c30a113b2a
support bf16 and fp8 in Tensor.tolist ( #11704 )
...
memoryview does not support it, but casting works fine so cast is fine
2025-08-17 15:11:13 -04:00
nimlgen and GitHub
1c62a3833b
am: add versioned_header to load_fw ( #11702 )
...
* am: add versioned_header to load_fw
* fix mypy
2025-08-17 20:11:57 +03:00
qazal and GitHub
eb3c918c5b
viz: s/area/height ( #11703 )
2025-08-17 19:20:01 +03:00
qazal and GitHub
d762edd694
viz: define tracks in python ( #11701 )
...
* viz: defines tracks in python
* update unittests
* figuring it out
* works
* diff cleanup
* math
* y axis is back
2025-08-17 18:19:13 +03:00
qazal and GitHub
eeeea29171
viz: device list refactor ( #11700 )
...
* viz: device list refactor
* paddingTop/padding-top
2025-08-17 15:08:54 +03:00
George Hotz and GitHub
9366a23eb0
test backward in test_tiny ( #11697 )
...
* test backward in test_tiny
* empty
2025-08-16 20:29:39 -07:00
chenyu and GitHub
4666df71c1
fix test_fuse_and_tc_opt ( #11699 )
2025-08-16 21:10:53 -04:00
geohotstan and GitHub
3d7c35d615
add fuse and tc opt bug repro ( #11695 )
...
* FINALLY HAVE A SMALL REPRO OH BOY
* show failure in CI
* cleaner?
* 1 possible fix
* Revert "1 possible fix"
This reverts commit 9e0fd215dd .
2025-08-16 18:24:49 -04:00
nimlgen and GitHub
d1224a7c4a
am: check both signatures ( #11694 )
...
* am: check both signatures
* fix
2025-08-16 20:01:07 +03:00
qazal and GitHub
58c8991fa4
add Ops.REWRITE_ERROR ( #11689 )
2025-08-16 00:56:53 +03:00
qazal and GitHub
ec4fccb1da
viz: pass through RewriteNotReady ( #11690 )
2025-08-16 00:33:59 +03:00
qazal and GitHub
e954decb44
viz: pass UOp.st errors ( #11688 )
2025-08-16 00:07:56 +03:00
nimlgen and GitHub
bf0c45fd16
system: resource_resize might be unavail ( #11680 )
2025-08-15 22:03:23 +03:00
George Hotz and GitHub
4ab9fb2edd
explicit fixed point rewrite ( #11685 )
...
* explicit fixed point rewrite
* local cache
* fix that
2025-08-15 11:08:41 -07:00
chenyu and GitHub
5d6963c968
RuntimeError for unsupported dtype in PYTHON ( #11686 )
2025-08-15 13:59:27 -04:00
nimlgen and GitHub
b970cd6895
am: fix psp ring completion ( #11679 )
...
* am: psp ring timeout + fix 0 fence_value
* no sleep
2025-08-15 20:15:49 +03:00
qazal and GitHub
c8ba48b223
show rewrite errors in viz ( #11684 )
2025-08-15 19:09:47 +03:00
George Hotz and GitHub
560984fd8d
small changes from rangeify ( #11682 )
...
* small changes from rangeify
* const like thing
* ksym
2025-08-15 08:45:52 -07:00
chenyu and GitHub
d0d39885c3
onnx in tinygrad ( #11675 )
2025-08-14 19:57:21 -04:00
wozeparrot and GitHub
71260a5ea4
feat: only bench openpilot 0.9.9 models ( #11664 )
2025-08-14 19:27:18 -04:00
chenyu and GitHub
4ddefbccb4
update setup packages ( #11674 )
...
sorted, and added missing 'tinygrad.frontend' and 'tinygrad.runtime.autogen.nv'
2025-08-14 19:24:57 -04:00
chenyu and GitHub
48c4033ae1
fix pylint for onnx ( #11673 )
...
* fix pylint for onnx
* too long
2025-08-14 18:48:02 -04:00
chenyu and GitHub
e9d0027591
llama MP realize weight after shard ( #11672 )
...
* llama MP realize weight after shard
prevents memory spike on device 0
* empty weight for FAKEDATA
2025-08-14 16:17:46 -04:00
nimlgen and GitHub
4176b24264
amd: support xcc in regs ( #11670 )
...
* amd: support xcc in regs
* mockamd
* typong
2025-08-14 21:20:11 +03:00
Sieds Lykles and GitHub
f399d0d75d
Render mod in terms of idiv ( #11668 )
...
* Render mod in terms of idiv
* cvar -> var
2025-08-14 19:59:39 +02:00
nimlgen and GitHub
d747eeed32
amd logs parser based on device ( #11669 )
2025-08-14 19:49:33 +03:00
geohotstan and GitHub
1e904155e3
Add Onnx Huggingface to test/models/test_onnx.py ( #11468 )
...
* BOOM
* cache extra/huggingface/models/
* why max buffer size is not 0
* override MAX_BUFFER_SIZE
* less models
* remove more models and change cache dir to already cached dir
* only metal
* less is more?
* remove check ops
* why is this not setting the ENVVAR
* ughhhhh just test in models
* only cpu and gpu
* only cpu actually
* just override it idk
* final
* move extra dependencies up top
* simplification
* fix print
* make README better
* revert ops_disk fix for now
* clean up test_onnx
* remove testing fashion clip model cuz sloooowwwwww
* actually let METAL run this
* fix comment mistake
* fix download path in run_models
* does this work?
* cleanup setup and teardown
* contextvar like this?
* prove model is cached
* do I need to increment DOWNLOAD_CACHE_VERSION?
* see if cached with incremented DOWNLOAD_CACHE_VERSION
* use warnings to see if the model exists
* revert DOWNLOAD_CACHE_VERSION stuff and clean up
* add retry to download
* nit
2025-08-14 11:16:41 -04:00
Sieds Lykles and GitHub
06beeb6e13
Nest div even if factor is negative ( #11666 )
2025-08-14 13:58:59 +02:00
Sieds Lykles and GitHub
661e9a2d5d
div_and_mod_folding refactor ( #11585 )
...
* divmod const folding is its own function
* split nested mod optimization out of div and mod folding
* make `fold_binary_numerator` its own function
* factor out `fold_divmod_congruence`
* check sign of numerator
* add tests
* assert int on vmin and vmax
* add type: ignore
* factor out more rules
* remove div_and_mod_folding
* cached_property to property
* remove import
* add returns
* restore old order
* check sign of x.vmin and newx.vmin
* check more signs
* add some test that would have caught bugs
* better test if the div simplified
* shorten line
* replace terms_factors_const with pop_const
* move that back
* minor cleanup
* remove comments
* some cleanup
2025-08-14 11:52:42 +02:00
chenyu and GitHub
0fc43c2e54
fix test_const_tensor_index index ( #11660 )
...
index should be ints
2025-08-13 19:50:16 -04:00
chenyu and GitHub
4fe19eec72
Ops.TRUNC ( #11659 )
2025-08-13 18:40:48 -04:00
qazal and GitHub
eb10a9c76a
viz: always left align timeline values ( #11658 )
2025-08-13 23:55:28 +03:00
George Hotz and GitHub
22bdf48cdd
render ranges in viz, name gbufs with sizes. changes from rangeify ( #11656 )
...
* render ranges in viz, name gbufs with sizes. changes from rangeify
* fix unit test dtypes
2025-08-13 12:46:16 -07:00
George Hotz and GitHub
9b4da590bb
remove need for cast_vec ( #11653 )
...
* remove need for cast_vec
* fix amdllvm
2025-08-13 12:09:47 -07:00
e2873a3a41
[bounty] Muon optim ( #11414 )
...
* newton schulz
* add muon + move newton schulz to tensor
* compact newton schulz
* better tests
* cleanup
* add comments for muon
* cleanup
* add export with tests
* match muon optim with test optim
* cleanup
* unsed import
* correct comment
* whitespace
* move export
* muon test fix
* match reference impl + tests
* remove export by moving muon device
* add credit
* cleanup
* remove print
* spacing
* spacing
* comma
* cleanup
* removal
* fix tests + optim momentum
* consistent is not/ not
* more consistency
* fix test
* cleanup
* fix the nones
* remove comment
* cast
* comment
* comment
* muon teeny test
* muon flag beautiful mnist
* set steps
* steps as hyperparam
* match default test steps
* name
* large cleanup
* dont care about steps
* nesterov false default
* match each other impl
* steps
* switch nest
* swap defaults
* update docstring
* add no nesterov test
* ban fuse_optim
* prints
* classical momentum
* alternative condition
* recon
* pre + post wd
* false default
* detach
* signature changes
* context
* swap order
* big cleanup
* 0 step instead
* parity
* remove fuse
* remove fused
* better paper
* assert message
* correct shape check + eps
* multidim
* add eps
* cleanup
* correct assert message
* lint
* better tests
* naming
* ns_steps,ns_params
* update docstring
* docstring
* match sgd and muon together
* sandwich
* add back fused
* parity
---------
Co-authored-by: George Hotz <[email protected] >
2025-08-13 14:27:55 -04:00
chenyu and GitHub
94e6d84e32
rewrite Tensor.round to not use cast int ( #11654 )
2025-08-13 13:51:08 -04:00
George Hotz and GitHub
d2521d828a
transcendental+idiv+threefry are uop decompositions ( #11636 )
...
* transcendental+idiv+threefry are uop decompositions [pr]
* threefry decomp
* fix randomness tests
* fix webgpu
* unneeded now
* fix
* move prematcher
* all cast should probably be cast_vec
2025-08-13 09:37:12 -07:00
geohotstan and GitHub
cf7224ce3e
fully lint onnx.py ( #11647 )
...
* mypy
* ruff ruff ruff
2025-08-13 08:22:06 -07:00
geohotstan and GitHub
925555b62a
Fix onnx Domain bug ( #11650 )
2025-08-13 08:20:50 -07:00
Sieds Lykles and GitHub
67df617fe1
add launch bounds to ptx ( #11646 )
2025-08-13 13:05:39 +02:00
qazal and GitHub
88f95e9f59
viz: minor fixups for firefox ( #11645 )
...
* fix circle attr
* set fill color
2025-08-13 12:59:28 +03:00
qazal and GitHub
6f88eac0fc
viz: refactor node and edge tagging ( #11644 )
2025-08-13 12:41:01 +03:00
qazal and GitHub
8140bf9778
viz: create layout once ( #11643 )
...
* start
* work
* works
* diff cleanup
2025-08-13 09:24:58 +03:00
chenyu and GitHub
3fb79bb43a
minor onnx cleanups ( #11642 )
2025-08-13 01:05:19 -04:00
chenyu and GitHub
e9e5a08a04
simplify onnx cubic ( #11641 )
...
we can drop the double where and abs since we know which ranges the inputs map into
2025-08-12 19:57:31 -04:00
George Hotz and GitHub
18cdbec447
split decompositions pass ( #11638 )
...
* split decompositions pass
* fix ptx
* pack load store early
* restore that
2025-08-12 12:56:05 -07:00
chenyu and GitHub
0d8a0d7a96
update test_multi_const_folding_tensor to include pow ( #11635 )
...
pow folds now
2025-08-12 13:35:37 -04:00
Sieds Lykles and GitHub
4d6e407eb0
Extend fast_idiv to negative ints ( #11632 )
...
* fast idiv for signed ints
* Add rule and test
* fix tests
* redo fuzz_fast_idiv to do negative ints as well
* adjust comments
* remove unused imports
2025-08-12 19:34:49 +02:00
qazal and GitHub
17adbe86d8
hotfix: do not default to capturing args in track_rewrites ( #11634 )
2025-08-12 20:01:24 +03:00
ad9dec25b3
combine onnx parser and onnx ( #11485 )
...
* start
* more
* fix onnx_runner test
* pass
* patch for disk and add domains from huggingface
* simpler docs
* revert domain changes
* rerun ci
* revert onnx ops test change
* add fix from strenum stuff
* correct way
* revert correct way to leave the fix for another PR
* test segfault
* Revert "test segfault"
This reverts commit 4e1aaf41e7 .
* remove some unnecessary documentation
* test segfault again
* Revert "test segfault again"
This reverts commit 56fc5f03e7 .
* try gemini suggested patch for sys._getframe
* keep trying with gemini
* revert not working gemini suggestions and try faulthandler
* remove pythonfaulthandler
* trigger CI a few times
* minimize diff
---------
Co-authored-by: chenyu <[email protected] >
2025-08-12 12:56:39 -04:00
Sieds Lykles and GitHub
4c3982c44e
Take sign out of mod ( #11631 )
...
* Add rule and test
* fix tests
2025-08-12 18:44:36 +02:00
qazal and GitHub
e28605e324
rename profile point event fields [pr] ( #11633 )
2025-08-12 19:11:21 +03:00
nimlgen and GitHub
8a7be0a747
metal: workaround for transfers sync issue ( #11622 )
...
* metal: workaround for transfers sync issue
* metal tracsfer sync is broken
* hm
* rm it?
* keep it
2025-08-12 16:16:34 +03:00
qazal and GitHub
efe8b5611d
move ProfilePointEvent out of device.py [pr] ( #11630 )
...
Generic profiling events exist in helpers so they can be imported from
everywhere in tinygrad.
2025-08-12 09:58:32 +03:00
chenyu and GitHub
0d7075f2de
assign should broadcast input tensor ( #11629 )
...
fixed test_assign_broadcast
2025-08-11 23:36:35 -04:00
Joshua Kissoon and GitHub
c44760c89d
torch backend: fix arange, add linalg.cross, add tests ( #11628 )
2025-08-11 23:34:41 -04:00
George Hotz and GitHub
ca41b5e38b
skip_0 in graph rewrite [pr] ( #11627 )
...
* skip_0 in graph rewrite [pr]
* no track_rewrites on test
* use dict instead of set
2025-08-11 18:29:04 -07:00
Sardor and GitHub
ca7a641442
fix bugs at examples/yolov3.py ( #11614 )
...
* Update load_weight. Give valid model url
* Fix bug in iou function
2025-08-11 21:14:47 -04:00
chenyu and GitHub
0c97d6de1b
don't round pow output for int pow int ( #11625 )
...
also added atol=0 and big pows for the tests
2025-08-11 20:57:47 -04:00
chenyu and GitHub
d623f6d850
support int Tensor pow to const non-negative int ( #11624 )
...
matches torch
2025-08-11 19:50:19 -04:00
chenyu and GitHub
857a830dcc
fix test_arange_float_step ( #11623 )
2025-08-11 16:58:42 -04:00
chenyu and GitHub
0806677b51
rewrite sort idx ( #11613 )
2025-08-11 16:20:56 -04:00
George Hotz and GitHub
700c11597b
switch contextvars.ContextVar to _ContextVar ( #11621 )
2025-08-11 12:20:09 -07:00
ae0c3cfff6
change clang -march flag to -mcpu on arm ( #10970 )
...
Co-authored-by: wozeparrot <[email protected] >
2025-08-11 13:38:48 -04:00
geohotstan and GitHub
27bcb9fd1c
Support cubic mode for ONNX Resize OP ( #11612 )
...
* start
* add reference
* this is so much slower
* this makes sense but differs from official impl, but results are still correct..?
* add a comment
* Just keep it simple for now since I don't fully get it yet
* address comments
* correct
* teeny clean up
* another small comment improvement lol
2025-08-11 11:49:30 -04:00
nimlgen and GitHub
d2bb1bcb97
cloud: a bit better err handling ( #11616 )
...
* cloud: err propagation to client
* fix
* print exc
* linter
* excs
* fix
* hm
* flaky
2025-08-11 15:51:22 +03:00
qazal and GitHub
6a232ccdac
viz: add tiny range drawing helper ( #11620 )
...
* viz: add tiny range drawing helper
* less
2025-08-11 15:15:43 +03:00
qazal and GitHub
e768773e13
viz: use colors helper ( #11618 )
2025-08-11 13:10:15 +03:00
qazal and GitHub
7d6c0a8cc7
viz: refactor progress msg ( #11617 )
2025-08-11 13:01:36 +03:00
chenyu and GitHub
630edcffd8
remove .float calls in olmoe ( #11610 )
...
still matches torch
2025-08-10 20:33:22 -04:00
chenyu and GitHub
a67e0917c3
list indexing can normalize in python ( #11609 )
...
* list indexing can normalize in python
list index does not need to be normalized in tensor
* update those
2025-08-10 20:02:38 -04:00
chenyu and GitHub
1181ec0cd2
few more tensor indexing test cases ( #11608 )
2025-08-10 18:56:42 -04:00
George Hotz and GitHub
996c907c0b
rewrite not ready + children machinery ( #11607 )
...
* rewrite not ready + children machinery
* it doesn't like track rewrites
2025-08-10 15:28:30 -07:00
Sieds Lykles and GitHub
1875bc69f9
Late rewrite rules for CMPLT ( #11591 )
...
* add rules
* more rules
* fix comment spelling
* remove two rules
2025-08-10 22:18:13 +02:00
nimlgen and GitHub
5403a4aeaf
null dev: support offset on buffers ( #11606 )
...
* null dev: support offset on buffers
* nolimit
2025-08-10 21:58:37 +03:00
geohotstan and GitHub
b0dab6a4cd
onnx Resize OP clean up ( #11603 )
...
* start
* slight clean up
2025-08-10 14:10:39 -04:00
Sieds Lykles and GitHub
10540414cd
Add Ops.CMPEQ ( #10431 )
...
* Add op
* add to Groupop.ALU
* fix spec
* fix ptx
* temporary pickle by name to see process replay
* add Ops.EQ to binary ops
* Actuall rename properly
* add test to assert CMPEQ is being used
* Ops.CMPEQ is automatic cast to bool
* add Ops.CMPEQ to llvm
* add Ops.CMPEQ to llvm
2025-08-10 13:13:16 +02:00
chenyu and GitHub
f7aa1b85fe
minor sort cleanups ( #11602 )
2025-08-10 01:51:23 -04:00
chenyu and GitHub
dfb702ef33
fix sort for small dim ( #11601 )
...
* fix sort for small dim
* fixed test_sort_empty
2025-08-10 01:17:41 -04:00
chenyu and GitHub
ef17af85c6
remove .float call in llama logit ( #11598 )
...
* remove .float call in llama logit
* bfloat item
2025-08-10 00:02:18 -04:00
chenyu and GitHub
dd3d2eb36c
add training llama3 test in ci ( #11599 )
2025-08-09 22:35:39 -04:00
chenyu and GitHub
3e64467322
remove freqs_cis contiguous in llama ( #11597 )
2025-08-09 21:11:12 -04:00
chenyu and GitHub
7338ffead0
small beautiful_mnist update ( #11596 )
...
gather is fast now. there's a conv/bw kernel that only gets fast with BEAM, but whole thing runs < 5 seconds now regardless
2025-08-09 19:51:14 -04:00
chenyu and GitHub
45baec1aab
model parallel llama ( #11588 )
...
MP=8 GRADIENT_ACC_STEPS=3 BS=1 DEFAULT_FLOAT=bfloat16 OPTIM_DTYPE=bfloat16 LLAMA3_SIZE=70B SEQLEN=512 PYTHONPATH=. MODEL=llama3 python3 examples/mlperf/model_train.py
2025-08-09 16:54:27 -04:00
nimlgen and GitHub
09bc377da3
search: print runtime failures on debug ( #11593 )
2025-08-09 23:01:19 +03:00
nimlgen and GitHub
14f99ff1a1
amd: doorbell_cpu_addr is not used ( #11592 )
...
* amd: doorbell_cpu_addr is not used
* hm
2025-08-09 20:03:21 +03:00
Sieds Lykles and GitHub
01c770c77b
Fix z3 float cast in indexing ( #11590 )
...
* adjust dtype of z3_renderer and add rule for cast
* dtypes.bool is also cast noop
* add regression test
* make embedding smaller
* even smaller test
2025-08-09 17:59:23 +02:00
Sieds Lykles and GitHub
10d388499d
Refactor optional.py ( #11578 )
...
* move fast_idiv to transcendental
* move optional.py
* adjust comment
* change import
* mypy needs this?
2025-08-09 17:35:05 +02:00
nimlgen and GitHub
20e46a175c
do not use disk with usb ( #11119 )
...
* not use disk with usb
* better name
2025-08-09 11:58:02 +03:00
qazal and GitHub
53179953fc
viz: factor out memory graph render ( #11586 )
2025-08-08 20:18:11 +03:00
qazal and GitHub
8ce72d3fad
simpler disassembly table spec ( #11583 )
...
* simpler disassembly table spec
* update ui
* move to scalar/vec render
2025-08-08 17:59:26 +03:00
qazal and GitHub
44a222a9b2
viz: move resource usage summary to server ( #11582 )
2025-08-08 17:08:28 +03:00
qazal and GitHub
793ace530e
update amd_uop_matmul.py import ( #11581 )
...
Using this for testing SQTT
2025-08-08 17:07:35 +03:00
chenyu and GitHub
b232c60def
benchmark openpilot 0.9.9 ( #11575 )
...
* benchmark openpilot 0.9.9
not sure what to do with the 0.9.7 ones with IMAGE=2 and validate
* name
2025-08-08 01:26:14 -04:00
qazal and GitHub
16f0edbe90
pass opts arg in get_program process replay [pr] ( #11571 )
...
* fix ptx process replay
* keyword arg
* renderer is also optional [pr]
* test_linearizer fixup
* name function order is args,ret,kwargs
* can use opts_to_apply
* pass through p.applied_opts
* sink_arg
* now it opens devices too
2025-08-08 03:05:09 +03:00
qazal and GitHub
960cc6533a
pass through name function args in track_rewrites ( #11572 )
2025-08-08 02:28:52 +03:00
wozeparrot and GitHub
1826004ef9
feat: add tinyos builder link ( #11570 )
2025-08-07 17:42:18 -04:00
George Hotz and GitHub
82be8abfd2
move opt under codegen ( #11569 )
2025-08-07 14:19:17 -07:00
chenyu and GitHub
702e38dc19
remove FUSE_ARANGE_UINT ( #11567 )
...
also add IGNORE_OOB=1 to bert runs. lowered BS on tinybox to 90 since 96 oom during eval without reset
2025-08-07 16:49:06 -04:00
George Hotz and GitHub
6ed2dfd187
delete the arange dim mismatch restriction ( #11568 )
...
* delete the arange dim mismatch restriction
* skip that test race
2025-08-07 13:46:17 -07:00
wozeparrot and GitHub
7ae4335127
feat: generate blend index ( #11566 )
2025-08-07 14:20:28 -04:00
chenyu and GitHub
594cbdc66f
skip AM ResNet50 benchmark ( #11565 )
...
hanging with FUSE_ARANGE?
2025-08-07 14:07:01 -04:00
chenyu and GitHub
aa1a6f2132
support threshold in Tensor.softplus ( #11564 )
...
fix gradient for large input
2025-08-07 13:43:18 -04:00
7ee3770961
FUSE_ARANGE=1 ( #11427 )
...
* FUSE_ARANGE=1
* fix test
---------
Co-authored-by: George Hotz <[email protected] >
2025-08-07 13:32:34 -04:00
geohot
4dfcfb1ae5
Revert "Revert "viz: align-center checkbox ( #11555 )""
...
This reverts commit c52facfd29 .
2025-08-07 08:15:57 -07:00
geohot
7e42427a7b
Revert "Revert "viz: remove color for unbind step ( #11554 )""
...
This reverts commit 5650c7b86c .
2025-08-07 08:15:51 -07:00
geohot
dc765fbeb7
Revert "viz: timeline perf ( #11533 )"
...
This reverts commit 031f26632b .
2025-08-07 08:08:51 -07:00
geohot
5650c7b86c
Revert "viz: remove color for unbind step ( #11554 )"
...
This reverts commit 1e205775bd .
2025-08-07 08:08:50 -07:00
geohot
c52facfd29
Revert "viz: align-center checkbox ( #11555 )"
...
This reverts commit 91ec093464 .
2025-08-07 08:08:50 -07:00
geohot
974cfbe76d
Revert "viz: add support for colored tooltip text ( #11556 )"
...
This reverts commit b3f7ea6f93 .
2025-08-07 08:08:49 -07:00
geohot
3bf0db80ef
Revert "viz: pick the largest rect for proxy fillColor ( #11558 )"
...
This reverts commit 76079bc7f2 .
2025-08-07 08:08:48 -07:00
George Hotz and GitHub
9764c6cdee
fix mismatch reduce, try 2 ( #11560 )
...
* fix mismatch reduce, try 2
* fix heuristic
* delete that test
* don't start allowing ones
2025-08-07 07:57:58 -07:00
qazal and GitHub
76079bc7f2
viz: pick the largest rect for proxy fillColor ( #11558 )
2025-08-07 16:40:17 +03:00
nimlgen and GitHub
4f29a2c441
fix flaky test on macos ( #11557 )
2025-08-07 15:55:35 +03:00
qazal and GitHub
b3f7ea6f93
viz: add support for colored tooltip text ( #11556 )
2025-08-07 15:04:43 +03:00
qazal and GitHub
91ec093464
viz: align-center checkbox ( #11555 )
2025-08-07 14:22:02 +03:00
qazal and GitHub
1e205775bd
viz: remove color for unbind step ( #11554 )
2025-08-07 14:16:21 +03:00
nimlgen and GitHub
031f26632b
viz: timeline perf ( #11533 )
...
* viz: timeline perf
* progress
* fast
* less lines
* less lines
* less lines
* fix chrome
2025-08-07 13:16:17 +03:00
George Hotz and GitHub
a1aa5670aa
Revert "fix mismatch reduce ( #11547 )" ( #11549 )
...
This reverts commit 49d21a9055 .
2025-08-06 22:43:15 -07:00
George Hotz and GitHub
49d21a9055
fix mismatch reduce ( #11547 )
...
* fix mismatch reduce
* cleanups
* fix shape
* fix mypy
* resolve
2025-08-06 21:12:51 -07:00
George Hotz and GitHub
21570545d3
move view pushing to codegen, try 2 ( #11534 )
...
* move view pushing to codegen, try 2
* fix up some linearizer tests
* fix test search
* fix test schedule
* delete that test
* fix test arange
* fix a few tests
* update tests
* push views
* ebs cleanup
* fix local/reg
* test and lint
* fix more tests
* test cleanups
* skipped that one
2025-08-06 15:58:38 -07:00
wozeparrot and GitHub
2d5bdc939d
faster llama3 dataloader ( #11540 )
2025-08-06 18:25:57 -04:00
George Hotz and GitHub
80d9cced07
more test cleanups ( #11544 )
...
* more test cleanups
* revert that
2025-08-06 15:05:21 -07:00
George Hotz and GitHub
6fd1332763
update some tests for less Kernel ( #11543 )
...
* update some tests for less Kernel
* get_program update
2025-08-06 14:19:59 -07:00
George Hotz and GitHub
09dc7af8e9
move bind to big graph ( #11539 )
...
* move bind to big graph
* fix tests
* unbind inside kernel only
* merge views
* fix multitensor
* failure text change
2025-08-06 13:27:51 -07:00
George Hotz and GitHub
7c5e115747
test_mismatch_reduce ( #11538 )
2025-08-06 10:02:14 -07:00
George Hotz and GitHub
4fe11725c6
pass through sink arg, update linearizer test ( #11536 )
...
* pass through sink arg, update linearizer test
* get_program help
* bump line count
* use new api
2025-08-06 09:48:48 -07:00
George Hotz and GitHub
bfebb5c37b
do store in the replace_buffers ( #11535 )
2025-08-06 08:42:45 -07:00
geohotstan and GitHub
1163292759
move onnx_parser into onnx ( #11530 )
2025-08-06 10:46:27 -04:00
George Hotz and GitHub
7b16fadd87
load view late + simpler rewrite ( #11525 )
...
* add the load view later
* simpler replace buffers
* rewrite name
2025-08-06 06:55:11 -07:00
nimlgen and GitHub
930d8dae0c
hcq: lazy prof signal allocation ( #11531 )
2025-08-06 15:28:11 +03:00
nimlgen and GitHub
eafc7fda12
upd perfetto ( #11528 )
2025-08-06 14:00:34 +03:00
nimlgen and GitHub
1afb290027
ci: fix runner in nv ( #11527 )
2025-08-06 10:38:04 +03:00
qazal and GitHub
61dae0685c
viz: show total mem in tooltip ( #11526 )
2025-08-06 06:51:26 +03:00
George Hotz and GitHub
cf66df0ea6
put load early to make pointers match ( #11524 )
2025-08-05 20:04:32 -07:00
George Hotz and GitHub
92175626e3
prereqs: move views to codegen ( #11522 )
2025-08-05 19:27:58 -07:00
chenyu and GitHub
c9225d22ce
only disable flaky test_jit_multidev_xfer ( #11523 )
2025-08-05 22:17:25 -04:00
George Hotz and GitHub
f58fd3143d
cleanup fix_kernel ( #11520 )
...
* cleanup fix_kernel
* early load buffer
* early meta ops
* move those to fix_kernel_ops
* fix tests
* remote metal was flaky
* Revert "fix tests"
This reverts commit a27019383d .
* that hack broke things
* fine for ptx
2025-08-05 18:38:43 -07:00
George Hotz and GitHub
067daee5be
pin torch to 2.7.1 ( #11519 )
2025-08-05 15:58:57 -07:00
George Hotz and GitHub
b39f43c46a
optimize in rewrite, try 2 ( #11518 )
...
* changes
* fix test uops
* optimize in rewrite, try 2
2025-08-05 15:52:53 -07:00
geohot
07b0df0d86
hotfix: test tensor dims start at 1
2025-08-05 15:40:24 -07:00
George Hotz and GitHub
4dabdf7c6d
Revert "optimize in rewrite ( #11516 )" ( #11517 )
...
This reverts commit 3b777a9e05 .
2025-08-05 15:39:07 -07:00
George Hotz and GitHub
3b777a9e05
optimize in rewrite ( #11516 )
...
* changes
* fix test uops
* dim shouldn't be 0
* huh, why did that one not save
2025-08-05 15:33:26 -07:00
nimlgen and GitHub
ec676eddfa
nv: move base address higher ( #11514 )
2025-08-05 22:42:53 +03:00
qazal and GitHub
7703f8b805
viz: skip flops info if estimates is symbolic ( #11513 )
2025-08-05 22:12:52 +03:00
nimlgen and GitHub
fc4e713d1c
jit graph split tests ( #11507 )
...
* jit graph split tests
* fix
* one more test
* more tests
* fix
* xm
* rmeote
2025-08-05 21:32:37 +03:00
George Hotz and GitHub
c57fde51f9
move swizzler to opt ( #11509 )
2025-08-05 11:31:30 -07:00
chenyu and GitHub
ace8e9a706
fix test_conv2d_winograd ( #11511 )
2025-08-05 12:15:46 -04:00
chenyu and GitHub
223aaa0492
clean up more conv tests ( #11510 )
2025-08-05 12:15:30 -04:00
Garret Castro and GitHub
76e62a1c23
extract conv layer test logic ( #11488 )
...
* refactor: extract conv layer test logic
* tuple is unnecessary
* integrate _test_conv logic into all conv tests
* fix linter, forgot dilation
* undo winograd extraction
adds too many if statements for a single case
2025-08-05 11:15:54 -04:00
8b8bd6c534
make einsum generate same kernels ( #11508 )
...
Co-authored-by: b1tg <[email protected] >
2025-08-05 11:12:52 -04:00
uuuvn and GitHub
011ef8fa9d
Fix incorrect jit current batch devs reset ( #11505 )
...
`current_batch_devs = []` (in `flush_batch()`) happens between
`new_batched_devs = ...` and `current_batch_devs = new_batched_devs` =>
doesn't actually reset anything leading to things not jitting properly
which 2xs remote bert step time (should have similar effects on any
non-hcq backend)
2025-08-05 08:16:16 +03:00
chenyu and GitHub
f02720ca2d
fix fuse gate_contiguous unique ( #11504 )
2025-08-04 23:43:31 -04:00
George Hotz and GitHub
7f6acfb0d5
give define global and friends a shape ( #11502 )
...
* give define global and friends a shape
* ignore negative size
* ptx fix
2025-08-04 19:09:39 -07:00
chenyu and GitHub
83385e7abc
update gradient src in ramp.py ( #11499 )
...
that's simplified now
2025-08-04 18:58:03 -04:00
qazal and GitHub
846a2826ab
viz: remove TracingKey.fmt ( #11482 )
...
* viz: remove TracingKey.fmt
* remove from test too
2025-08-05 00:00:03 +03:00
chenyu and GitHub
01d44e8f16
tiny reduce_gradient cleanup [pr] ( #11498 )
2025-08-04 16:12:53 -04:00
chenyu and GitHub
8a11af01ed
remove broken paperswithcode links in doc ( #11497 )
2025-08-04 13:12:33 -04:00
4f0ee4e982
BPE tokenizer ( #11415 )
...
* BPE works
* refactor tok
* oops
* basic tests
* fix eval
* smaller diff
* fix error
* proper vocab decoding
* use regex for splitting
* escape ucatrange
* full compat
---------
Co-authored-by: George Hotz <[email protected] >
2025-08-04 09:52:38 -07:00
06af9f9236
fix double exception + add name,loc in error msg ( #11487 )
...
Co-authored-by: b1tg <[email protected] >
2025-08-04 13:41:23 +03:00
nimlgen and GitHub
4877aa965a
ast seems to probe nv as well ( #11494 )
2025-08-04 11:47:07 +03:00
chenyu and GitHub
e0106b6b25
1/(x*c) -> (1/c)*(1/x) ( #11491 )
...
example: 2*(2*a).reciprocal() -> a.reciprocal()
# TODO: bounds for reciprocal
# TODO: should z3 work?
2025-08-03 23:35:46 -04:00
qazal and GitHub
5870352fe1
viz: factorize llvm-mca call ( #11490 )
2025-08-04 00:31:23 +03:00
chenyu and GitHub
dbc7807c61
enable WEBGPU tests with buffer limit ( #11489 )
...
TestSample still fails?
2025-08-03 13:02:44 -07:00
nimlgen and GitHub
8f374ee1f7
nv: print devfmr in gsp logs ( #11484 )
2025-08-03 15:12:53 +03:00
chenyu and GitHub
823f1a01db
move cast around expand backward to tensor.py ( #11483 )
2025-08-02 23:03:54 -04:00
chenyu and GitHub
0ce0f51010
generic double cast folding ( #11481 )
...
b.cast(a).cast(b) -> b if a preserves all values in b
2025-08-02 19:26:37 -04:00
qazal and GitHub
72e0d1d0dc
viz: profile the compiler in TINY device ( #11457 )
...
* viz: profile the compiler in TINY device
* leanup
2025-08-03 02:03:20 +03:00
chenyu and GitHub
66be747908
few more dtype cast convinience methods ( #11480 )
2025-08-02 15:47:09 -04:00
chenyu and GitHub
e22e5da9a5
move some test_dtype tests to unit ( #11479 )
2025-08-02 15:25:00 -04:00
nimlgen and GitHub
da0b955be4
hcq: cpu can be graphed ( #11474 )
...
* hcq: cpu can be graphed
* ops
* new jit decisions
* fix test
* fix remote
* cleaner
* fix
2025-08-02 21:01:19 +03:00
chenyu and GitHub
f7965f85aa
Revert "feat: faster index building ( #11462 )" ( #11478 )
...
This reverts commit 3a4deb08d2 .
2025-08-02 12:50:48 -04:00
kevvz and GitHub
ef7e01cadf
Fix SVD shape bug + Fix batched SVD bug ( #11477 )
...
* failing test case
* fix
* better test
* space
2025-08-02 09:47:41 -07:00
6ecaf8e7b2
refactor: use less index and simplify reduce axes check [pr] ( #11476 )
...
* use output_shape/full_shape
* simple final_reduces check
---------
Co-authored-by: b1tg <[email protected] >
2025-08-02 09:44:51 -07:00
wozeparrot and GitHub
3a4deb08d2
feat: faster index building ( #11462 )
...
* feat: faster index building
* feat: correct training samples
2025-08-02 11:50:18 -04:00
nimlgen and GitHub
8cc2d64edb
amd: reuse create_queues for usb iface ( #11473 )
2025-08-02 14:40:46 +03:00
chenyu and GitHub
9e8e6b45ab
grad acc train llama ( #11467 )
...
* grad acc train llama
* log step time
2025-08-01 15:54:50 -04:00
chenyu and GitHub
7ad7329257
data parallel train llama ( #11466 )
2025-08-01 12:13:51 -04:00
nimlgen and GitHub
9f2182f92f
cpu: start threading ( #11324 )
...
* cpu: threading
* syncs
* llvm
* fix
* opt
* fx
* fix
* missed sync
* one line less
* cleaner
* fix
2025-08-01 15:35:07 +03:00
qazal and GitHub
c7ae1bd474
viz: more consistent border styling ( #11464 )
2025-08-01 09:31:06 +03:00
George Hotz and GitHub
8ff03806e8
add llama layers ( #11460 )
...
* add llama layers
* add contig bw for speed
2025-07-31 16:28:04 -07:00
qazal and GitHub
719827b95d
viz: add flops / mem bw to device programs ( #11459 )
...
* viz: add flops / mem bw to device programs
* better spacing style
2025-08-01 02:12:30 +03:00
chenyu and GitHub
3f742a5a7c
comma space lab models benchmark ( #11461 )
2025-07-31 19:06:18 -04:00
geohot
474ee9daa5
hotfix: add contiguous_backward to llama
2025-07-31 15:07:12 -07:00
qazal and GitHub
fa66d9772d
viz: show const node when it's root ( #11456 )
2025-08-01 01:01:58 +03:00
qazal and GitHub
056dabda5a
viz: refactor to color scheme ( #11455 )
2025-08-01 00:17:50 +03:00
nimlgen and GitHub
e5b6149dfb
more typing in drivers ( #11454 )
...
* more typing in drivers
* rm
2025-07-31 23:26:33 +03:00
qazal and GitHub
bad3cf5731
viz: add LLVM machine code analysis ( #11421 )
...
* start
* works everywhere
* add viz api
* utilization table
* reg pressure ui
* use llvm-mca
* llvm-mca ui
* work
* cleanup
* cycle through, defaults are enough
* x86 pending
* x86 nops
* get mcpu/mtriple from autogen
* cleanup server diff
* move parser to python
* normalize to pct of max
* segments legend
* imports
* also monospace
* max comes from the total per instruction
* base on the value
2025-08-01 01:59:26 +08:00
chenyu and GitHub
e847677e8a
use AxisType in search instead of colors ( #11452 )
2025-07-31 13:07:33 -04:00
nimlgen and GitHub
75c2c42def
suppress exceptions only during finalization ( #11451 )
...
* suppress exceptions only during finalization
* fix
* fix typing
* fix more warns
* fix
* better?
* Revert "better?"
This reverts commit a068aa5793 .
* mm?
* no as e
2025-07-31 13:57:12 +03:00
wozeparrot and GitHub
24dd0d52ed
feat: test remove to cpu ( #11444 )
2025-07-30 20:18:56 -07:00
c3cfcb50cb
Add linalg_det and test for torch backend ( #11405 )
...
* add linalg_det and test
* space
---------
Co-authored-by: chenyu <[email protected] >
2025-07-30 22:04:44 -04:00
cba3655de5
Add Test for Setitem ( #10559 )
...
* init
* update
* better
* failing test
* works
* Delete test file
* clean
* lint
* simplify variable name
* rm contigious, rm int dtype, and add assertEqual
---------
Co-authored-by: chenyu <[email protected] >
2025-07-30 22:03:41 -04:00
wozeparrot and GitHub
6252f7770e
feat: fake data ( #11447 )
2025-07-30 17:18:20 -07:00
chenyu and GitHub
e300451f3a
update llama3 ( #11446 )
...
`LR=1e-4 TRAIN_ON_VAL=1 DEFAULT_FLOAT=bfloat16 FUSE_ARANGE=1 JITBEAM=2 OPTIM_DTYPE=bfloat16 LLAMA3_SIZE=1B WARMUP_STEPS=36 DECAY_STEPS=360 SEQLEN=512 PYTHONPATH=. AMD=1 AMD_LLVM=0 MODEL=llama3 python3 examples/mlperf/model_train.py` trained to 7
2025-07-30 19:34:21 -04:00
wozeparrot and GitHub
5fb975351a
feat: flag for training on val ( #11441 )
2025-07-30 14:29:45 -07:00
chenyu and GitHub
4ca430e5bf
fix search dedup ( #11439 )
...
it should check against pre real_axis axis in actions, not real_axis.
2025-07-30 17:24:16 -04:00
wozeparrot and GitHub
d3da20eca6
feat: bump mlperf workflow timeout to 6 hours ( #11440 )
2025-07-30 14:12:12 -07:00
wozeparrot and GitHub
825b6a2505
feat: llama3 dataloader ( #11340 )
2025-07-30 13:27:55 -07:00
qazal and GitHub
af357b5dc8
disable TRACK_MATCH_STATS in BEAM workers [pr] ( #11437 )
2025-07-30 23:22:08 +03:00
George Hotz and GitHub
7c2d2eff86
check tensor core dims ( #11436 )
...
* check elements_per_thread in tensorcore [pr]
* check tc dims
2025-07-30 13:06:59 -07:00
nimlgen and GitHub
5fc5bb5237
ci: clear processes ( #11434 )
...
* unified hcq_smi for managment
* fix
* fix
* no reset for amd
2025-07-30 22:15:18 +03:00
George Hotz and GitHub
4f26a9ad32
check elements_per_thread in tensorcore [pr] ( #11435 )
2025-07-30 11:55:48 -07:00
nimlgen and GitHub
4b4ba5454c
ci: move driver start higher ( #11431 )
2025-07-30 10:48:38 +03:00
George Hotz and GitHub
1bef2d80c1
unrolls are all in the same scope ( #11429 )
...
* unrolls are all in the same scope
* fix that import
2025-07-29 16:55:37 -07:00
chenyu and GitHub
204da24cfc
increase driverbenchmark timeout-minutes to 15 ( #11428 )
2025-07-29 19:45:05 -04:00
chenyu and GitHub
d5fc6af4a2
remove unused ShapeTracker.consecutive [pr] ( #11426 )
2025-07-29 18:36:19 -04:00
George Hotz and GitHub
49a2583584
real new lowerer ( #11419 )
...
* real new lowerer
* fix group for reduce
* skip missing ranges
* fix wmma and unroll/contract
* real fix for wmma
* disable that test
* fix if gate
* simpler
* flash attention fusion works
* no end barriers
* still broken
* flash attention finally works
2025-07-29 15:35:51 -07:00
chenyu and GitHub
0e5d8d5c3c
remove tests that used .to_uop() ( #11425 )
...
* remove tests that used .to_uop()
* import
2025-07-29 15:52:16 -04:00
nimlgen and GitHub
c88e401d0e
ci: fix typos in h machine benchmarks ( #11423 )
2025-07-29 22:11:47 +03:00
chenyu and GitHub
90a5a312eb
simplify ShapeTracker in UOp.const [pr] ( #11424 )
2025-07-29 15:04:06 -04:00
chenyu and GitHub
398594029b
spec checks arg of VIEW are ShapeTracker ( #11422 )
2025-07-29 14:05:12 -04:00
geohot
1f1f99c287
hotfix: add DEBUG=3 to driver CI
2025-07-29 11:03:47 -07:00
George Hotz and GitHub
50fae54175
global local dims in gpudims [pr] ( #11420 )
2025-07-29 10:39:03 -07:00
chenyu and GitHub
9bc413f104
remove ShapeTracker.to_uop [pr] ( #11418 )
2025-07-29 13:29:37 -04:00
George Hotz and GitHub
ba2c4df125
dont render cast ptrs standalone ( #11417 )
...
* dont render cast ptrs standalone
* barrier cleanups
2025-07-29 09:24:26 -07:00
nimlgen and GitHub
d38d285489
ci: add h machines ( #11416 )
...
* ci: add h machines
* more
* fix names
* names not collide
* 20
* 10
2025-07-29 19:21:51 +03:00
2568bc0d99
ci: add caching for apt packages ( #11162 )
...
* add caching for apt packages
* remove 'inputs' from apt cache key, use outputs instead of env
* remove unnecessary mkdir for partial
---------
Co-authored-by: George Hotz <[email protected] >
2025-07-29 09:04:56 -07:00
George Hotz and GitHub
03909f2772
permute locals for HL uop matmul ( #11412 )
...
* permute locals for HL uop matmul
* parens fix that
* permutes
* 20 TFLOPS
2025-07-29 08:19:59 -07:00
nimlgen and GitHub
e0c9747684
amd: fix typo in has_scratch_base_registers for mi350 ( #11413 )
2025-07-29 10:30:06 +03:00
George Hotz and GitHub
735ad5f10d
kernel4 and 5 in uops ( #11411 )
...
* move simplify views to merge views
* add amd kernel 4
* Revert "move simplify views to merge views"
This reverts commit 1e07dff384 .
* k4 in python
* kernel4 written in uops
* k5 support
* cleanups
2025-07-28 19:35:48 -07:00
George Hotz and GitHub
fddc645668
HL=2 top matmul ( #11406 )
...
* HL=2 top matmul
* top colored
2025-07-28 12:32:38 -07:00
nimlgen and GitHub
c7b4ab86e4
fix llvm tc on mi350 ( #11404 )
2025-07-28 21:37:43 +03:00
chenyu and GitHub
9f7c72ff8f
remove UOp.valid method [pr] ( #11402 )
...
only used in add_buffer_ops
2025-07-28 11:29:08 -04:00
chenyu and GitHub
b22a34331b
remove const valid in fixup_ast [pr] ( #11401 )
2025-07-28 11:07:59 -04:00
qazal and GitHub
7737cbb2a0
viz: tabulate runtime stats ( #11400 )
2025-07-28 15:56:39 +03:00
chenyu and GitHub
ab6a27f627
remove a branch in UOp.r [pr] ( #11398 )
2025-07-27 18:00:01 -04:00
052191eae4
Remote multihost (p2p with infiniband verbs) ( #9746 )
...
Co-authored-by: wozeparrot <[email protected] >
2025-07-27 14:44:32 -07:00
qazal and GitHub
a22417cc75
viz: fix bug with wrong program links ( #11396 )
2025-07-28 02:52:06 +08:00
nimlgen and GitHub
a5371f514b
cpu: copies in profile ( #11392 )
...
* cpu: copies in profile
* fix
* rename to tiny?
2025-07-27 20:56:27 +03:00
George Hotz and GitHub
8c10085459
assert shape on lowerer store [pr] ( #11395 )
...
* assert shape on lowerer store [pr]
* fix ptx
2025-07-27 10:41:57 -07:00
qazal and GitHub
6174cfa828
viz: only show match counts greater than 0 ( #11394 )
2025-07-28 00:25:00 +08:00
qazal and GitHub
3466a220de
viz: disassembly viewer ( #11393 )
...
* test
* CPU=1 disasm works
* METAL=1 disasm works
* fix that
* work
* can unwrap
* work p2
* don't crash
2025-07-27 18:44:28 +03:00
qazal and GitHub
3bb232eb29
viz: query path in rewrite steps ( #11391 )
2025-07-27 14:51:47 +03:00
b7ef73babd
fix wmma ptx ( #11389 )
...
Co-authored-by: b1tg <[email protected] >
2025-07-26 23:28:35 -07:00
8dfcdb123d
less wmma args ( #11385 )
...
* less wmma args
* scalar
* ops_python
* mypy
* lint
* dedup
* helper wmma_args
---------
Co-authored-by: b1tg <[email protected] >
Co-authored-by: George Hotz <[email protected] >
2025-07-26 21:24:05 -07:00
George Hotz and GitHub
dfeee63d30
uop matmul work ( #11388 )
...
* uop matmul work
* works with locals
2025-07-26 21:23:55 -07:00
George Hotz and GitHub
3923e78061
no_vectorized_acc keeps single DEFINE_REG ( #11387 )
...
* no_vectorized_acc keeps single DEFINE_REG
* fix ptx, skip flaky test
2025-07-26 11:44:09 -07:00
qazal and GitHub
4866ad57da
viz: add runtime stats ( #11383 )
...
* viz: add runtime stats
* lint
* better
* flat
2025-07-26 20:40:46 +03:00
George Hotz and GitHub
2c70eaf18c
fix load / barrier ( #11386 )
...
* fix load / barrier
* cleanups
* fix CI
2025-07-26 10:27:37 -07:00
nimlgen and GitHub
65673e68ca
hcq: do not import during __del__ ( #11384 )
...
* hcq: do not import during __del__
* ignore
2025-07-26 13:58:55 +03:00
George Hotz and GitHub
466ab5a3f2
store/load not pass through index ( #11381 )
...
* noop
* fix noop
* store cat is NOOP
* store dtype is void
* stores aren't passed through anymore
* meh, skip those for ptx
* correct ptx skip
* hl runs
2025-07-25 21:01:47 -07:00
George Hotz and GitHub
0a5f37946b
unused permute arg on r ( #11379 )
2025-07-25 19:52:37 -07:00
George Hotz and GitHub
48562cb2db
full shape simpler ( #11376 )
2025-07-25 18:27:48 -07:00
chenyu and GitHub
3d68feb67d
minor onnx Gather cleanup ( #11375 )
...
removed a type ignore and one error code skip
2025-07-25 21:08:08 -04:00
chenyu and GitHub
88c338bfcc
add kernelize to keccak for each data block ( #11370 )
...
* add kernelize to keccak for each data block
test_long works now. this prevents internal uops from growing propotional to data length and eventually too deep
* this?
* hash stuff
* gate test
* mv
2025-07-25 16:07:20 -04:00
chenyu and GitHub
dab07bcad9
use next instead of full list in UOp._device [pr] ( #11369 )
...
prevents exponential fan out
2025-07-25 10:04:29 -04:00
nimlgen and GitHub
1bb1f1aee8
hcq: fix race in _at_profile_finalize ( #11368 )
2025-07-25 14:14:02 +03:00
George Hotz and GitHub
490a93902c
define reg doesn't have init anymore ( #11365 )
...
* define reg doesn't have init anymore
* remove that
* no special logic for dr
* fix amd uop matmul
2025-07-24 19:15:49 -07:00
George Hotz and GitHub
9da3f72495
identity store for DEFINE_REG ( #11363 )
...
* identity store for DEFINE_REG
* identity store for DEFINE_REG
* noop continue
2025-07-24 16:41:29 -07:00
chenyu and GitHub
cc795c6656
simplify keccak pad mask code ( #11362 )
2025-07-24 19:24:10 -04:00
chenyu and GitHub
c0c4bc9d7c
use int32 for keccak reorder_indexes ( #11360 )
...
it's used for tensor indexing, so int32 instead of uint64 is slightly faster
2025-07-24 15:54:50 -04:00
George Hotz and GitHub
0602b22086
kernel spec ( #11359 )
...
* kernel spec
* ops.VIEW
* work
2025-07-24 12:45:38 -07:00
qazal and GitHub
519f1d13cc
viz: generic stuff from gpu counters ui ( #11358 )
...
* viz: generic stuff from gpu counters ui
* move pointer
* pre fetch
* move timeout
2025-07-24 20:29:24 +03:00
nimlgen and GitHub
3b3de8df61
hcq: graphed copies ( #11302 )
...
* fast copies p2
* upd and fix
* graph supports
* fixes
* fixes
* fixes
* fix
* fix
* fix mockgpu
* fix alignment
* smaller in ci
2025-07-24 17:36:19 +03:00
nimlgen and GitHub
3046ead6e8
jit: graph reports ei support ( #11356 )
2025-07-24 16:35:10 +03:00
nimlgen and GitHub
bf12041910
hcq: mapping of cpu to all hcq devices ( #11354 )
...
* hcq: mapping of cpu to all hcq devices
* fix kfd
* nv
* simpler
* cleaner
* correct skip
* fix ifaces
* system fixes
* mypy
2025-07-24 12:52:38 +03:00
chenyu and GitHub
82e6de7fc6
more keccak reference tests ( #11329 )
2025-07-23 22:06:39 -04:00
George Hotz and GitHub
b0dc97d1f7
write out kernel 3 in uops ( #11352 )
...
* write out kernel 3 in uops
* matmul is correct
* gemm passes spec
* bugfix to match speed
* cleanups
2025-07-23 17:32:38 -07:00
chenyu and GitHub
5b570196e4
support DEV= to specify device ( #11351 )
2025-07-23 17:40:55 -04:00
76a2ddbd78
Move remote tests out of onnx ( #11310 )
...
Co-authored-by: wozeparrot <[email protected] >
2025-07-23 13:25:55 -07:00
George Hotz and GitHub
7f0a41df4d
move optional out of devectorize [pr] ( #11350 )
...
* move optional out of devectorize [pr]
* fast idiv
2025-07-23 11:26:05 -07:00
nimlgen and GitHub
0f374e10d2
cpu: use mmap for allocations ( #11349 )
...
* cpu: use mmap for allocations
* ops
* fix mypy
2025-07-23 20:30:18 +03:00
George Hotz and GitHub
ae07a93814
simple block barrier ( #11341 )
...
* simple block barrier
* simple
2025-07-23 10:14:11 -07:00
chenyu and GitHub
86e7504111
mypy check extra/onnx.py ( #11348 )
...
instead of running test with 3.10, add onnx to mypy which would have caught StrEnum regression. Several type annotation failed mypy now that does not affect running the code and were skipped for now
2025-07-23 12:42:59 -04:00
chenyu and GitHub
960da9319d
Remove StrEnum in onnx for python 3.10 ( #11345 )
...
some training tests failed looks like parsing error?
2025-07-23 11:52:25 -04:00
qazal and GitHub
478a355325
gate PRINT_MATCH_STATS behind graph_rewrite tracking ( #11344 )
2025-07-23 16:32:43 +03:00
nimlgen and GitHub
ca09c180dc
cpu: remove del spam ( #11343 )
...
* cpu: remove del spam
* fix
2025-07-23 12:02:37 +03:00
nimlgen and GitHub
304eb9cecb
allocate less memory in am tests ( #11342 )
2025-07-23 11:11:26 +03:00
George Hotz and GitHub
e14b4fefa5
ranges on store ( #11334 )
...
* ranges on store
* fix store spec
* fix that
* fix gates
* fix tests
* fix ptx
2025-07-22 21:00:50 -07:00
George Hotz and GitHub
c65b5aab62
small things from endrange ( #11339 )
...
* small things from endrange
* store
2025-07-22 19:45:37 -07:00
George Hotz and GitHub
53339e62f7
no gate store anymore ( #11338 )
...
* no gate store anymore
* fix up spec
2025-07-22 18:41:15 -07:00
chenyu and GitHub
7a9a5cfd28
isolate test/external/external_test_am.py ( #11335 )
...
seems to be the one crashing, also remove -n=auto for that
2025-07-22 19:02:20 -04:00
George Hotz and GitHub
fcbd0e4de3
assigns are no longer used [pr] ( #11333 )
2025-07-22 15:35:07 -07:00
George Hotz and GitHub
09431d4ad1
make DEFINE_REG behave like the others ( #11273 )
...
* simpler define reg
* cast
* PTRCAT define_acc
* cleanups
* fix uops stats
* fix linearizer tests
* llvm
* define reg sets const
* define reg sets const
* no assign
* collapse that
* fix test_max_pool2d_bigger_stride_dilation
* use index, fix webgpu
* devec
* fix tests
* fix webgpu
* fix llvm
* threads for python
* fix ops_python
* only for reg
* acc_half is real now in the emulator
* fix llvm
* fix webgpu init
* fix wgpu test
* fix some tests
* fix ptx
* fix ptx bool acc
* cleanups
* broken, meh. will fix with ENDRANGE
* line count
2025-07-22 13:53:56 -07:00
chenyu and GitHub
4535908679
update keccak test_long ( #11331 )
...
it should compare with arg "shake_128"
2025-07-22 16:08:01 -04:00
nimlgen and GitHub
3faa352dcc
am: bump version after mm changes ( #11328 )
2025-07-22 21:54:10 +03:00
George Hotz and GitHub
affd83961c
small changes from define_reg ( #11327 )
...
* small changes from define_reg
* fix webgpu
2025-07-22 11:11:48 -07:00
nimlgen and GitHub
53b3d87456
am: use 4-lvl pdir ( #11326 )
2025-07-22 20:58:15 +03:00
chenyu and GitHub
2d7c28de6a
clean up dup lambdas in helper_test_exception ( #11325 )
2025-07-22 12:21:57 -04:00
chenyu and GitHub
c6aa8e58ca
fix TestDropoutProbabilityEdgeCases ( #11322 )
2025-07-22 11:13:56 -04:00
chenyu and GitHub
fb42c84365
merge TestRollEdgeCases into test_ops ( #11321 )
2025-07-22 10:55:57 -04:00
chenyu and GitHub
1d8b3e9d1c
movementop only Tensor.roll ( #11317 )
...
* movementop only Tensor.roll
* fixed
2025-07-22 10:34:15 -04:00
chenyu and GitHub
a41140241b
truncate unsigned const in cstyle ( #11318 )
...
it can be a warning or a hard error in clang
PTX and PYTHON also need fix, skipping for now
2025-07-22 08:02:12 -04:00
qazal and GitHub
6668d6d241
fix word_wrap with newlines in input string [pr] ( #11319 )
2025-07-22 12:03:13 +03:00
qazal and GitHub
0c4e19f270
hotfix: disable process replay in REMOTE=1 tests ( #11320 )
...
* hotfix: disable process replay in REMOTE=1 tests
* comment
2025-07-22 10:41:58 +03:00
George Hotz and GitHub
3b674df34b
generic changes from define_reg_2 ( #11315 )
...
* generic changes from define_reg_2
* fix for ptx
* ugh, that one
2025-07-21 15:14:06 -07:00
chenyu and GitHub
6e9506e6fd
Tensor.roll supports dims=None ( #11313 )
2025-07-21 17:29:23 -04:00
George Hotz and GitHub
108aac8af4
use AddrSpace instead of local ( #11314 )
...
* use AddrSpace instead of local
* addrspace in test
2025-07-21 14:00:06 -07:00
chenyu and GitHub
d3a93185a6
clean up test_roll ( #11312 )
2025-07-21 16:00:50 -04:00
George Hotz and GitHub
532b52fcef
store has a dtype, like assign ( #11309 )
...
* store has a dtype, like assign
* fix upat
* fix test
2025-07-21 12:50:01 -07:00
445ff8de56
ONNX onnx_parser and buffer_parse clean up ( #11000 )
...
* start
* remove onnx.load from compile4 and move np to dropout
* clean up and enable test
* clean up
* move WebGPU ONNX test into MacOS (WebGPU)
* leave test in ONNX (CPU)
* fix raw_data init None, and simplify onnx_runner test a little?
* THESE TESTS ARE SO UGLY UGHH
* need to really think about how to structure the test
* wow LLMs are quite something
* not always on disk now
* also add external data loading test
* cleaner tests
* minimize diff and add const folding tests
* add external data loading too
* whoops add webgpu back.. but why was it not needed in the first place?
* better comment
* move webgpu test to macos(webgpu)?
* llm english so much better than me wow
* trigger CI to check flakiness
---------
Co-authored-by: chenyu <[email protected] >
2025-07-21 15:10:25 -04:00
George Hotz and GitHub
842184a1ab
rename kernelize to schedule, try 2 ( #11305 )
2025-07-21 11:18:36 -07:00
George Hotz and GitHub
7e8f5dde74
matmul style is still reshape ( #11308 )
2025-07-21 11:14:57 -07:00
George Hotz and GitHub
41de76a7fd
put assign and store next to each other [pr] ( #11306 )
2025-07-21 11:07:35 -07:00
nimlgen and GitHub
de2df92551
hcq: use devices instead of ids in HCQGraph ( #11303 )
...
* hcq: use devices instead of ids in HCQGraph
* fiz
2025-07-21 20:03:12 +03:00
wozeparrot and GitHub
30ce16a424
feat: failing test for long keccak ( #11292 )
2025-07-21 12:49:23 -04:00
uuuvn and GitHub
178dbf3f66
Remote scheduler changes ( #11177 )
2025-07-21 09:29:44 -07:00
वेदांत and GitHub
e368628736
Add amin support to Tensor operations in Torch backend ( #11290 )
...
* intiger div mod fix
* Revert "intiger div mod fix"
This reverts commit d5d2f201bf .
* feat arg_min support
* tets update
* test fix
2025-07-21 09:14:08 -04:00
qazal and GitHub
5eb54e2499
viz: close event streams before profiler render ( #11300 )
2025-07-21 15:42:31 +03:00
nimlgen and GitHub
cc3c1e4c14
hcq: move cpu to hcq ( #11262 )
...
* hcq: move cpu to hcq
* import time
* upd
* fix
* windows support
* hm
* cleaner
* fix timer
* fix timing
* std is ns
* skip profiler
* mypy
* cleaner
* cleanups
* after merge
* default is back
2025-07-21 15:10:38 +03:00
nimlgen and GitHub
816c01c2d4
hcq: default copy_queue_t=None ( #11297 )
2025-07-21 14:45:20 +03:00
qazal and GitHub
6520a7fcb6
viz: factorize event stream ( #11298 )
2025-07-21 14:42:00 +03:00
nimlgen and GitHub
9c533e5c38
hcq: cpu prereq ( #11296 )
2025-07-21 13:35:18 +03:00
nimlgen and GitHub
e87a42e243
hcq: prepare for windows ( #11293 )
...
* hcq: prepare for windows
* comments
2025-07-21 13:08:56 +03:00
nimlgen and GitHub
df3ba0a7c0
autogen: fix imports in libusb ( #11294 )
2025-07-21 13:04:27 +03:00
nimlgen and GitHub
dd6a2d432f
hcq: default timestamp metrics is ns ( #11295 )
2025-07-21 12:56:30 +03:00
wozeparrot and GitHub
53345ef4e2
feat: make ops_disk work on block devices ( #11291 )
2025-07-20 14:39:50 -07:00
qazal and GitHub
3002c63b1e
process replay: optionally pass tinygrad import error ( #11289 )
...
* process replay: optionally pass tinygrad import error
* gate all tinygrad internals
* s/getenv/os.getenv pre import
* diff
2025-07-20 22:57:56 +03:00
chenyu and GitHub
9e3a593313
minor kernel.py cleanups [pr] ( #11286 )
2025-07-20 10:15:31 -04:00
quortus and GitHub
5f17927a87
Shorten UOp.load method ( #11285 )
2025-07-20 13:48:04 +03:00
chenyu and GitHub
54924f9969
type remove Union and Optional [pr] ( #11283 )
...
use `|` for consistency
2025-07-19 14:05:52 -04:00
nimlgen and GitHub
2f72be5055
nv_smi: init basic insmod/rmmod/reset cmds ( #11282 )
2025-07-19 15:43:03 +03:00
qazal and GitHub
577e581943
fix typo in sqtt/readme ( #11281 )
2025-07-19 15:10:24 +03:00
nimlgen and GitHub
188ed38315
replace from_mv with lightweight mv_address ( #11280 )
2025-07-19 13:50:51 +03:00
1a25e27f32
Do not produce out of spec intermediate UOp in gated LOAD/STORE folding ( #11207 )
...
Co-authored-by: chenyu <[email protected] >
2025-07-18 15:42:55 -04:00
chenyu and GitHub
ec3efd2919
move upcast before reduce ( #11250 )
...
* move upcast before reduce
upcast goes to end of global+local+upcast
* r_196_32_4_24_8
2025-07-18 14:42:15 -04:00
chenyu and GitHub
be2f4336e6
use onnx 1.18.0 in DSP test ( #11279 )
2025-07-18 14:09:23 -04:00
nimlgen and GitHub
9a88bd841c
hcq: refactor into peer_groups ( #11277 )
...
* hcq: refactor into peer_groups
* fix fors
* fixes
* ooops
* mypy
* tiny fixes
2025-07-18 16:34:18 +03:00
nimlgen and GitHub
f432eef708
hcq: rename CPU -> KICK in graph for kickoff signal ( #11278 )
2025-07-18 15:54:35 +03:00
quortus and GitHub
52bbd9900b
[pr] Stable tensor order in _find_all_tensors_for_uops ( #11276 )
...
* Use dict for all_tensors to get stable tensor order in _find_all_tensors_for_uops
* Rerun tests
2025-07-18 13:12:01 +03:00
chenyu and GitHub
c5a5d74642
Revert "image_dot of 2 half inputs returns half ( #11007 )" ( #11274 )
...
This reverts commit fa8e08f922 .
2025-07-17 17:34:18 -04:00
fa8e08f922
image_dot of 2 half inputs returns half ( #11007 )
...
* cast after sum
* comment out skipif
* minor fix
* only test IMAGE
* IMAGE is supported now
* simpler
* simplerr
* only cast if dtype is None
* dont need to change base_imaeg_type
* only cast when dtype is half
* add explicit test
* actually no, workflow seems better
* actually, keep both
* move test
* fix indent
---------
Co-authored-by: Utkarsh Gill <[email protected] >
2025-07-17 13:47:22 -07:00
geohotstan and GitHub
536b254df4
Bump onnx to 1.18.0 ( #11266 )
...
* bump
* thou hast implement functions
* hacked in domain support
* some clean ups
* hack quantize_onnx_test too
* add helper lol, why onnx tests why
* better dispatcher, but need tests and better naming
* flaky ci
* change some names
* small clean ups
* make it easier to clean up tests once ORT supports 1.18.0
* nits
* fix bug of Softmax_1 being registered in onnx_ops
* need a default value
* resolve_const is better name
* fix OnnxRunner.to
* use proper domain names
2025-07-17 15:35:41 -04:00
qazal and GitHub
1606491b1c
viz: refactor to generic shape spec ( #11272 )
2025-07-17 20:25:15 +03:00
nimlgen and GitHub
cfb229473f
hcq: refactor buffer mapping ( #11271 )
...
* hcq: refactor buffer mapping
* fix
* fix mypy
2025-07-17 15:16:49 +03:00
qazal and GitHub
e68af3b336
disable flaky assert in test_cpu_profile ( #11270 )
2025-07-17 06:50:39 +03:00
chenyu and GitHub
60ffe00172
remove Kernel.first_reduce [pr] ( #11269 )
2025-07-16 18:30:14 -04:00
chenyu and GitHub
522dc72f08
remove Kernel.local_dims [pr] ( #11268 )
...
* remove Kernel.local_dims [pr]
also not needed
* fix test_matvec
2025-07-16 17:46:19 -04:00
chenyu and GitHub
d8c783f65f
remove Kernel.global_dims [pr] ( #11267 )
...
all reference to global used axis_types, so we don't need number of global helper that was used to locate GLOBAL
2025-07-16 17:16:49 -04:00
uuuvn and GitHub
6f0ddcc24c
Remote cross-host graph ( #11229 )
2025-07-16 13:27:54 -07:00
nimlgen and GitHub
6aa20c607d
nv: graceful shutdown to cold state ( #11265 )
2025-07-16 19:49:35 +03:00
chenyu and GitHub
59b52d49d7
remove .global_dims that are for locating GLOBAL [pr] ( #11264 )
2025-07-16 11:19:31 -04:00
chenyu and GitHub
e6c016ddd0
move check axis < shape_len to real_axis [pr] ( #11263 )
...
ensure output of real_axis is always valid
2025-07-16 10:15:44 -04:00
quortus and GitHub
924bc7c9ae
Fix test_uop_spec ( #11259 )
2025-07-16 11:02:31 +03:00
chenyu and GitHub
c8e5c4d7c3
insert_before -> insert_at [pr] ( #11257 )
...
more precise
2025-07-15 17:44:34 -04:00
wozeparrot and GitHub
b32d9321fb
feat: more keccak cleanup + more explicit shape ( #11256 )
2025-07-15 13:57:47 -07:00
chenyu and GitHub
9f79079cbe
update KernelInfo dims to return list of dims [pr] ( #11255 )
...
local dims are not contiguous once upcast sits between local and groupreduce
2025-07-15 15:01:39 -04:00
chenyu and GitHub
629fa21b6b
remove final range in heuristic [pr] ( #11251 )
...
all dims are based on AxisType now
2025-07-15 11:39:15 -04:00
chenyu and GitHub
d7adc24083
remove Kernel.first_upcast [pr] ( #11248 )
...
first_reduce does not need a default now
2025-07-15 10:21:34 -04:00
nimlgen and GitHub
197d345804
nv: print rpc msg with DEBUG>=3 ( #11247 )
2025-07-15 16:39:58 +03:00
chenyu and GitHub
034e51bd36
remove first_reduce used for locate real_axis [pr] ( #11245 )
...
LOCAL goes to the last of (GLOBAL+LOCAL)+1
GROUP goes to right before first REDUCE
2025-07-15 09:19:38 -04:00
chenyu and GitHub
0e2422d216
Kernel.axes_of helper [pr] ( #11243 )
...
look up dim based on AxisType
2025-07-14 22:17:43 -04:00
chenyu and GitHub
968f6b2a2e
remove hasattr(self, 'axis_types') checks in dims property [pr] ( #11242 )
...
no needed anymore
2025-07-14 20:59:51 -04:00
leopf and GitHub
557ca7d757
testing SimpleTokenizer against OASST1 ( #11214 )
2025-07-14 17:09:31 -07:00
wozeparrot and GitHub
5878b189b8
don't const fold shape changing bitcast ( #11236 )
2025-07-14 16:42:16 -07:00
chenyu and GitHub
b6662096cb
remove more first_reduce [pr] ( #11239 )
2025-07-14 19:13:44 -04:00
chenyu and GitHub
eb8e17ef59
remove most of the first_upcast [pr] ( #11238 )
2025-07-14 16:54:24 -04:00
qazal and GitHub
c78b1cbae7
viz profiler cleanups ( #11234 )
...
* move all render calls to zoom callback
* cleanup the naming
* require transform arg
2025-07-14 19:06:33 +03:00
chenyu and GitHub
36ce883c7d
update heuristic to use k.upcastable_dims and k.unrollable_dims [pr] ( #11233 )
...
idea is to make it behave the same regardless of axis order and with empty 1s in shape.
not quite fully remove all first_upcast yet because some conditions used already upcasted size which need a separate benchmark to remove.
2025-07-14 11:10:30 -04:00
qazal and GitHub
c0c695dd89
viz: remove extra transform ( #11232 )
2025-07-14 16:51:47 +03:00
chenyu and GitHub
da219199f5
minor hcopt cleanup [pr] ( #11231 )
2025-07-14 09:36:25 -04:00
nimlgen and GitHub
756ba1a5f9
nv: support ampere in nvpci ( #11230 )
2025-07-14 15:35:44 +03:00
uuuvn and GitHub
b2cc6cfa1b
JIT_BATCH_SIZE is a ContextVar ( #11228 )
2025-07-14 14:03:45 +03:00
nimlgen and GitHub
c4a920d95c
nv: use last signature ( #11227 )
2025-07-14 13:00:39 +03:00
nimlgen and GitHub
a830d37881
nv: check wpr2 is inited ( #11226 )
2025-07-14 11:46:14 +03:00
chenyu and GitHub
0387bb9630
clean up image upcast in hcopt [pr] ( #11220 )
...
GLOBAL+LOCAL for upcast
GROUP_REDUCE+REDUCE for unroll
2025-07-13 18:06:43 -04:00
chenyu and GitHub
85ddd72038
simpler grouptop in hcopt ( #11219 )
...
* simpler grouptop in hcopt
keep the only perf relevant conditions and the rest is handled by try except
* update openpilot read image count
2025-07-13 16:06:09 -04:00
qazal and GitHub
40847ca29c
viz: prune out of screen rects ( #11217 )
2025-07-13 21:49:59 +03:00
chenyu and GitHub
674dc28505
remove Kernel.full_unupcasted_shape [pr] ( #11215 )
...
decomp to shape_len and first_upcast to get the last upcast-able dim
2025-07-13 13:56:23 -04:00
chenyu and GitHub
9575cf6c6e
shave more hcopt [pr] ( #11213 )
...
start to use AxisType for conditions
2025-07-13 12:43:58 -04:00
Alisher Zhubanyshev and GitHub
4ef6b46b34
hcq: reduce launch overhead ( #11193 )
...
* nv: improve mmio creation speed
* add memoryview test
* fix indents
* move mv bench to `test_helpers`, remove comparison
2025-07-13 19:25:50 +03:00
nimlgen and GitHub
1cc2b3f845
nv: use wait_cond ( #11212 )
2025-07-13 19:25:20 +03:00
nimlgen and GitHub
6cce3a5d58
generic wait_cond ( #11210 )
...
* generic wait_cond
* fix linter
* fix linter
2025-07-13 16:59:21 +03:00
chenyu and GitHub
e11ccf2342
update float4 condition in hcopt ( #11211 )
...
don't need all upcast candidates to be upcast-able, only check the actual one
2025-07-13 09:51:45 -04:00
nimlgen and GitHub
55c54d9745
nv: sync after gpfifo setup ( #11209 )
2025-07-13 14:40:11 +03:00
chenyu and GitHub
d90d837013
clean up hcopt [pr] ( #11205 )
...
removed one condition that's always true
2025-07-12 23:10:27 -04:00
chenyu and GitHub
2b48b961be
fix a few broken AMX tests ( #11204 )
2025-07-12 21:42:38 -04:00
wozeparrot and GitHub
667c7a9fa6
clean: keccak cleanups + explicit shapes ( #11202 )
2025-07-12 18:17:14 -07:00
chenyu and GitHub
a0438012af
remove Kernel.get_program [pr] ( #11203 )
2025-07-12 20:50:29 -04:00
George Hotz and GitHub
d67c8e7b42
local metal on metal in uop syntax ( #11185 )
...
* local metal on metal in uop syntax
* TODO: just put the axis_info in the kernelinfo
* local
* amd_matmul works @ 28 TFLOPS
* clean up matmul
* kernel8 works
* remove that
* locals
* axistype innovation
* work
* cleanup
* kernel3 regs
* cleanup kernel3
* work
* why is it broken
* no beam
* reenable
* permutes
2025-07-12 16:31:19 -07:00
uuuvn and GitHub
40da5f0c81
fix silent mypy failure in ci ( #11201 )
...
Example: https://github.com/tinygrad/tinygrad/actions/runs/16215577171/job/45784110543?pr=11177#step:7:20
Caused by footguny exception in how `set -e` works:
```bash
python -m mypy --strict-equality --lineprecision-report . && cat lineprecision.txt
```
Will fail (and have non-zero exit code if run in interactive mode) but
because there is `&&` it won't count as script-terminating failure in a
script with `set -e` and instead as a test (similar to how fail of a
command in if condition won't count as a script-terminating failure
despite having non-zero exit code)
2025-07-12 15:12:25 -04:00
chenyu and GitHub
73caa5dd1b
remove Kernel.membufs [pr] ( #11200 )
2025-07-12 14:48:47 -04:00
5ce278b245
OnnxRunner file as input ( #10789 )
...
* file path as input and have parse be in OnnxRunner.__init__
* modelproto_to_onnxrunner -> modelproto_to_runner
* whoops, fix import
* oh flakiness again, is it because it's getting gc-ed?
* small changes
* CI flaky so just move compile4 fix in
* copy typing of onnx_load
* actually can just import onnx_load instead of onnx.load
* fix external_benchmark_openpilot
* fix onnx_runner test to use onnx_helper
* rerun CI
* try run_modelproto
* spam CI a few times
* revert run_modelproto since that's flaky also
* no external onnx_load usage except onnx.py
* cursor tab complete is evil. Snuck a darn sorted in. But does order change result? Why?
* model_benchmark 193s -> 80s, add OnnxRunner.to()...
* minimize diff and clean up
* device can be None, weird but eh
---------
Co-authored-by: chenyu <[email protected] >
2025-07-12 14:27:46 -04:00
nimlgen and GitHub
110cff3f2e
fix device arg to Tensor.randn ( #11194 )
...
* fix device arg to Tensor.randn
* simpler test
* self.assertEqual
2025-07-12 13:51:59 -04:00
chenyu and GitHub
6283d50224
DEPRECATED_linearize -> to_program [pr] ( #11198 )
2025-07-12 13:46:20 -04:00
George Hotz and GitHub
770a558585
lil cleanups from uop branch [pr] ( #11197 )
2025-07-12 09:46:28 -07:00
George Hotz and GitHub
5625e1904b
axis types in KernelInfo ( #11196 )
...
* axis types in KernelInfo [pr]
* simpler lowerer
* fix tests
2025-07-12 09:36:20 -07:00
nimlgen and GitHub
ea7f2f779c
hcq: p2p nv-amd ( #11195 )
...
* hcq: p2p between diff devices
* fix
2025-07-12 18:53:34 +03:00
qazal and GitHub
6a9f059b21
viz: early convert to cpu time ( #11192 )
2025-07-12 17:19:41 +03:00
chenyu and GitHub
12b04efd69
remove a TODO prod(k.full_shape[k.first_upcast:]) ( #11191 )
...
IMAGE=2 test/test_ops.py works now
2025-07-12 10:16:56 -04:00
nimlgen and GitHub
6f5250d158
nv: fix typing in rpc_rm_control ( #11189 )
2025-07-12 16:09:42 +03:00
qazal and GitHub
c0a5490c72
viz: minor profiler cleanup ( #11190 )
2025-07-12 14:18:24 +03:00
chenyu and GitHub
fdcc25e392
some noop hand_coded_optimizations cleanup [pr] ( #11188 )
2025-07-12 00:09:23 -04:00
chenyu and GitHub
1ad852a892
break up Kernel.reshape_and_permute [pr] ( #11187 )
2025-07-11 18:08:08 -04:00
d11b20129d
DMARef infra ( #10753 )
...
Co-authored-by: wozeparrot <[email protected] >
2025-07-11 14:09:47 -07:00
chenyu and GitHub
b072be0e2d
hotfix whisper main script ( #11184 )
2025-07-11 12:34:00 -04:00
qazal and GitHub
0b7e9b5db7
viz: bugfix for multiple rewrites with the same name ( #11182 )
2025-07-11 18:26:12 +03:00
nimlgen and GitHub
f9e4c4e57a
nv: nvpci blackwell support ( #11127 )
...
* nv: start 5090
* gsp init 5090
* mmu
* works
* after merge
* clenaer
* rwk
* x
* fx
* finish?
* fix
* unrelated
* fix
* commenbt
2025-07-11 17:02:09 +03:00
qazal and GitHub
1d85323572
viz: absolute scaling of memory graph ( #11181 )
2025-07-11 16:39:11 +03:00
nimlgen and GitHub
c7f6b617b4
nv: do not hardcode lv0 pd size ( #11180 )
2025-07-11 16:26:18 +03:00
nimlgen and GitHub
27922c986a
nv: generic mmu impl ( #11179 )
2025-07-11 16:26:09 +03:00
qazal and GitHub
d3ec63a5c3
viz: add base class for unittests ( #11178 )
2025-07-11 13:58:03 +03:00
qazal and GitHub
b791ea117d
viz: enable scrolling in profiler ( #11169 )
...
* viz: add scrollbar to profiler
* using margin fixes the layout bug
* s/profiler.clientHeight/profiler.scrollHeight, it's important
* closer
* scrolling on the device list also works
2025-07-11 11:30:13 +03:00
chenyu and GitHub
b219e47bef
remove Kernel.upcasted_axis [pr] ( #11175 )
2025-07-10 23:19:21 -04:00
George Hotz and GitHub
ccd382bc6f
use axis_types more [pr] ( #11172 )
...
* use axis_types more
* fix local shape
* simpler clause
* fix local shape
2025-07-10 15:05:13 -07:00
nimlgen and GitHub
fb278c6a02
do not recreate Compiled.profile_events in helper_collect_profile ( #11171 )
2025-07-10 23:55:12 +03:00
George Hotz and GitHub
5c5eb92ed4
tc unroll after upcast [pr] ( #11170 )
2025-07-10 13:43:50 -07:00
George Hotz and GitHub
05613c8cac
use shape str for tensor cores upcast/reduce [pr] ( #11168 )
...
* use shape str for tensor cores upcast/reduce [pr]
* reduce axis count isn't fixed
2025-07-10 13:10:58 -07:00
nimlgen and GitHub
cc6ed30f4f
nv: relative lv addressing in NVPageTableEntry ( #11164 )
2025-07-10 22:35:50 +03:00
chenyu and GitHub
439d033af9
update the README matmul example ( #11167 )
...
don't call rand and numpy to show that it's indeed one kernel
2025-07-10 14:47:29 -04:00
qazal and GitHub
bde80c0cdf
record GraphEvents in metal graph ( #11145 )
...
* record GraphEvents in metal graph
* add TestProfiler.test_graph, revert old stuff
* move profile capture to MetalGraph
* comment
* don't double record graph command buffers
* wait_check
* explicit delete
2025-07-10 21:32:06 +03:00
George Hotz and GitHub
8ce3d5906b
use shape_str for tensor cores ( #11165 )
2025-07-10 09:10:36 -07:00
nimlgen and GitHub
581397110f
nv: use classes in GSP_IP ( #11163 )
2025-07-10 17:47:12 +03:00
nimlgen and GitHub
705de6b8a6
nv: parse sizes of ctx buffers ( #11161 )
2025-07-10 17:46:48 +03:00
qazal and GitHub
dcc9704b6b
viz: profile RewriteSteps in TINY device ( #11125 )
...
* viz: profile RewriteSteps in TINY device
* use TracingKey with category
* split by whitespace
* add tracing.py
* work
* tracing_key
* TRACK_MATCH_STATS=3, can this be in defaults?
* fallback name
* work
* javascript
* measure text is slow
* checkout
* profile graph_rewrite/graph_rewrite_map
* change that
* no as
* finally
* work
* linking works
2025-07-10 17:45:57 +03:00
Pyry Kovanen and GitHub
32117402dd
metal: fix incorrect _free on interpreter exit ( #11158 )
2025-07-10 14:01:30 +03:00
qazal and GitHub
3d610f6d2b
viz: small ui cleanup ( #11157 )
...
* viz: small ui cleanup
* 2
2025-07-10 11:43:36 +03:00
chenyu and GitHub
7db07e5f2c
don't narrow range of CAST on bool/unsigned ( #11156 )
2025-07-09 22:20:09 -04:00
George Hotz and GitHub
e154a66f43
unroll axis 0 in tensor core ( #11155 )
...
* unroll is 0 in tc [pr]
* flip order of upcast/reduce in tensor core
* Revert "flip order of upcast/reduce in tensor core"
This reverts commit e564e38bcd .
2025-07-09 17:28:23 -07:00
George Hotz and GitHub
b7742ad9e4
migrate to string swizzle [pr] ( #11154 )
2025-07-09 16:57:53 -07:00
George Hotz and GitHub
4156baee93
break swizzle into three chunks [pr] ( #11153 )
...
* break swizzle into three chunks [pr]
* test failed
2025-07-09 15:30:34 -07:00
George Hotz and GitHub
ca2dc95433
swizzle in tc can't be none [pr] ( #11152 )
2025-07-09 14:44:23 -07:00
George Hotz and GitHub
53ae153404
tc should be in opt ( #11148 )
...
* tc should be in opt [pr]
* fix import
2025-07-09 14:12:21 -07:00
wozeparrot and GitHub
6697d0089d
initial gfx950 kfd support ( #11151 )
...
* feat: initial gfx950 support
* fix: lint
2025-07-09 13:45:16 -07:00
George Hotz and GitHub
262054be52
gfx950 tc support ( #11150 )
2025-07-09 13:30:42 -07:00
nimlgen and GitHub
b6981404ed
memory: use page shifts in memory manager ( #11149 )
...
* memory: use page shifts in memory manager
* fix
2025-07-09 22:05:00 +03:00
qazal and GitHub
5c1d215b41
viz: add Graph stream ( #11144 )
...
* viz: stack an event for the entire batch
* multi
* whitespace
* work
* multi graph, Graph gets its own row
2025-07-09 20:56:46 +03:00
George Hotz and GitHub
22305260e0
move tc to tc.py [pr] ( #11147 )
2025-07-09 10:55:56 -07:00
George Hotz and GitHub
2893feb9f6
cleanups for kernel.py ( #11143 )
...
* cleanups for kernel.py
* fixups
2025-07-08 18:10:25 -07:00
George Hotz and GitHub
b11ca104e9
axis cleanups [pr] ( #11142 )
2025-07-08 17:07:26 -07:00
chenyu and GitHub
7ce9e45474
mypy onnx_parser ( #11141 )
2025-07-08 19:50:28 -04:00
George Hotz and GitHub
a1b8f3e64f
delete info from kernel [pr] ( #11139 )
...
* delete info from kernel [pr]
* update kernel info
* delete info
2025-07-08 15:53:13 -07:00
George Hotz and GitHub
359bed74f8
axis type tracking [pr] ( #11137 )
...
* axis type tracking [pr]
* keep update_info
* keep legacy colors
* update tests to apply_opt
2025-07-08 14:16:25 -07:00
chenyu and GitHub
dada3f5bf3
skip some new onnx tests ( #11135 )
...
these fails on master with latest onnx
2025-07-08 16:12:48 -04:00
chenyu and GitHub
ffcc557986
lint onnx and onnx_parser ( #11134 )
2025-07-08 15:28:35 -04:00
George Hotz and GitHub
3238d21cd1
add finalized to kernel [pr] ( #11132 )
...
* add finalized to kernel [pr]
* add copy
2025-07-08 11:06:17 -07:00
geohot
289a411f5f
hotfix: remove unused GBARRIER, CONTIGUOUS color is GBARRIER
2025-07-08 10:31:06 -07:00
nimlgen and GitHub
43650169f4
nv: switch headers to 570.144 to match gsp ( #11131 )
2025-07-08 20:29:06 +03:00
quortus and GitHub
790b05ab12
[pr] Unify CONTIGUOUS and GBARRIER ( #11121 )
...
* Unify CONTIGUOUS and GBARRIER
* Simplify rules
2025-07-08 10:27:23 -07:00
nimlgen and GitHub
b516fe71b4
nv: return real struct in _alloc_boot_struct ( #11130 )
2025-07-08 20:04:43 +03:00
qazal and GitHub
3dfc0ff887
move cpu_profile and shared ProfileEvents from device.py to helpers [pr] ( #11126 )
...
* move cpu_profile and shared ProfileEvents to helpers [pr]
* TestProfiler.test_cpu_profile
* update test_viz.py
* TestProfiler.test_profile_multiops ordering, it's different streams now
2025-07-08 12:14:03 +03:00
George Hotz and GitHub
397826f0b4
add a test for 1B llm ( #11124 )
...
* add a test for 1B llm
* fix mbs
* add apps to release
2025-07-07 18:47:25 -07:00
George Hotz and GitHub
f7d4638e05
start LLM app, tons of clean up required. target is 200 line ollama ( #11068 )
...
* start LLM app, tons of clean up required. target is 200 line ollama
* kind of works
* simpler
* add k/v cache
* with SYM=1, it loops
* no rope cache
* simpler
* more cleanups
* cleanups
* works
* argparse and comments
* from gguf
* generate is a function
* no copy from cpu
* fix max context pass in
* test
* improve test
* ai2_arc
* fix 8B, use less ram
* 136 lines
2025-07-07 17:09:46 -07:00
chenyu and GitHub
341a686799
Tensor.diagonal ( #11122 )
...
only implemented main diagonal for 2-D tensors. with diagonal and qr, we can get determinant
2025-07-07 16:21:26 -04:00
584fd6af5a
Fix division by zero and mask bug in add views ( #11088 )
...
* merge view infinite loop test
* adjust condition in `x//d -> x//(-d)*-1`
* Fix division by zero in add views
* adjust offset end
* fix typo in comment
* add target to test_merge_views_variable
* fix view incorrectly being masked
* ssimplify strides and offset of the new view to canonicalize
* remove print in test
---------
Co-authored-by: qazal <[email protected] >
2025-07-07 10:05:47 -07:00
nimlgen and GitHub
71377cd233
nv: parse falcon app descs ( #11118 )
2025-07-07 18:14:14 +03:00
nimlgen and GitHub
9a573a1d99
nv: finalize nvdev ( #11117 )
...
* nv: finalize nvdev
* typo
2025-07-07 16:31:59 +03:00
nimlgen and GitHub
fa59c05282
nv: import flags from system ( #11115 )
...
* nv: import flags from system
* not used
2025-07-07 14:46:49 +03:00
a1a146a499
adding enable_gqa in SDPA ( #11097 )
...
Co-authored-by: wozeparrot <[email protected] >
2025-07-06 23:25:33 -07:00
nimlgen and GitHub
b73e89110e
nv: align allocations for perf ( #11114 )
2025-07-06 22:32:11 +03:00
chenyu and GitHub
7468959f4b
Tensor.argsort ( #11112 )
2025-07-06 13:56:35 -04:00
kevvz and GitHub
b7af9cf849
clean svd tests, set full_matrices false in torch backend ( #11113 )
...
* clean tests, set full_matrices false
* add more shape asserts
2025-07-06 13:55:49 -04:00
qazal and GitHub
a556f50668
viz: small ui fixes ( #11110 )
...
* share styling of ctx-list and metadata
* scrollbar-gutter: stable prevents layout shift when changing steps
* margin-left makes left side unaligned
2025-07-06 17:05:36 +03:00
chenyu and GitHub
ba88ec3ad0
pipe linalg svd to torch ( #11109 )
...
and found a bug in svd
2025-07-06 08:37:25 -04:00
chenyu and GitHub
845a4d32bc
Tensor.diag ( #11108 )
...
also updated Tensor.eye to use it
2025-07-05 23:03:02 -04:00
ttomsa and GitHub
4905af4ae0
remove invalid int div test ( #11106 )
...
* rm test
* also rm this
2025-07-05 18:57:55 -04:00
qazal and GitHub
a4aa769c0a
fix: type checking for track_rewrites key [pr] ( #11104 )
...
* fix: type checking for track_rewrites key [pr]
* also for cpu_profile
* func.__name__ to start
2025-07-05 20:11:21 +03:00
qazal and GitHub
81781dc12b
viz: renames and spacing changes to tracing ( #11102 )
2025-07-05 18:40:39 +03:00
qazal and GitHub
7619bf35e7
cleanup: remove disabled TestIndexingOrdering ( #11101 )
...
* cleanup: remove disabled TestIndexingOrdering
* don't import kernelize internals
2025-07-05 18:14:37 +03:00
qazal and GitHub
4fcfaa0ef7
viz: switch to TracingKey ( #11100 )
...
* viz: switch to TracingKey
* tuple
* order is name, keys, fmt
* add test_tracing_key
2025-07-05 17:46:18 +03:00
qazal and GitHub
458be950d9
viz: add TINY device ( #11095 )
...
* viz: add TINY device
* replace Any with a proper type
* reorder
* diff
* rename
* space
* from diff
* multiple keys
2025-07-05 16:54:55 +03:00
nimlgen and GitHub
4dccb2ea49
am_smi: increase kill retries ( #11099 )
2025-07-05 16:23:50 +03:00
chenyu and GitHub
39b4d72687
remove flatten and reshape in sparse_categorical_crossentropy [pr] ( #11093 )
...
not needed, directly operating on the classes dim is fine
2025-07-04 15:15:27 -04:00
nimlgen and GitHub
577afc9f05
hcq: remove redunt syncs and fix typing ( #11096 )
...
Before this patch the code could issues reduntdant syncs because of
the typing issue. Current tests should cover all correctness checks.
2025-07-04 21:49:47 +03:00
qazal and GitHub
41aa54eb5a
viz: resolve all graph references in python ( #11087 )
...
* viz: resolve all graph references in python
* it just maps things to the index
* always map the name
* key on the uop
* diff
* close
2025-07-04 20:35:25 +03:00
qazal and GitHub
3d8569f6d8
hotfix: infinite loop in tracking pattern matcher ( #11094 )
...
* failing test
* fix that
* given matchers
2025-07-04 19:55:26 +03:00
qazal and GitHub
a783211fc7
viz: allow end_time=None in trace events ( #11092 )
2025-07-04 17:48:17 +03:00
0xSG and GitHub
17119b0f23
hip_ioctl: platform.machine added ( #11084 )
2025-07-04 17:20:24 +03:00
nimlgen and GitHub
6656aa162c
nv: enable huge pages ( #11091 )
2025-07-04 17:17:24 +03:00
nimlgen and GitHub
01f3c4f44d
memory: simpler paddr allocation logic ( #11090 )
...
* memory: new paddr allocation logic
* am fix
* am refactrros
* fix
* mypy
* use it
* am
2025-07-04 17:00:36 +03:00
qazal and GitHub
f6d55d9272
viz: pickle UPat location ( #11086 )
2025-07-04 13:09:00 +03:00
qazal and GitHub
2403f126ed
move printable out of UPat [pr] ( #11085 )
...
* move printable out of UPat [pr]
* print_match_stats
2025-07-04 12:31:11 +03:00
qazal and GitHub
988540f401
support capturing cpu_profile on error ( #11078 )
...
* support capturing cpu_profile on error
* spacing
* pylint complains
2025-07-04 11:53:12 +03:00
chenyu and GitHub
a2f5a54458
move sparse_categorical_crossentropy to test_ops ( #11083 )
...
also flattened the tests
2025-07-03 21:40:54 -04:00
chenyu and GitHub
7c8ccb0267
sparse_categorical_crossentropy cleanup [pr] ( #11082 )
2025-07-03 18:32:52 -04:00
nimlgen and GitHub
e02ee8ef1b
nv: cleanups from 5090 ( #11081 )
2025-07-04 00:08:47 +03:00
George Hotz and GitHub
e9a01dd04a
Revert "Fix division by zero in add views ( #11075 )" ( #11080 )
...
This reverts commit 19f07e72f6 .
2025-07-03 11:39:44 -07:00
Sieds Lykles and GitHub
19f07e72f6
Fix division by zero in add views ( #11075 )
2025-07-03 11:37:59 -07:00
chenyu and GitHub
678cabc6f2
use argfix in Tensor.stack ( #11077 )
...
works for multiple Tensor args or single tuple/list of Tensors, but not the mixed
2025-07-03 12:15:11 -04:00
qazal and GitHub
b695e8c4d6
viz: remove support for naming with self ( #11076 )
2025-07-03 17:29:14 +03:00
Sieds Lykles and GitHub
53985297bd
add test, fix rewrite rule and raise error on division by zero ( #11073 )
2025-07-03 08:25:06 -04:00
nimlgen and GitHub
2d138c6cf1
am: factor out init_sw ( #11070 )
2025-07-03 11:01:17 +03:00
quortus and GitHub
a937ac80dc
Replace ASSIGN with STORE in UPat compiler ( #11065 )
2025-07-02 19:15:43 -07:00
George Hotz and GitHub
d049639221
little setitem test ( #11064 )
...
* setitem has one less realize, why broken
* put realize back
2025-07-02 15:10:24 -07:00
quortus and GitHub
17d85b9793
Refactor STORE implementation in ops_python ( #11060 )
2025-07-02 14:29:12 -07:00
George Hotz and GitHub
3b85534df0
outerworld range test [pr] ( #11059 )
...
* outerworld range test [pr]
* bound range
* grad acc test
* more tests
* 5 steps is fine
2025-07-02 14:28:44 -07:00
chenyu and GitHub
425d5f55c4
generate kernel dataset and upload artifact ( #11063 )
2025-07-02 17:21:25 -04:00
chenyu and GitHub
09cc64eea7
remove const 0 clause in "UOp with size 0 is zero" [pr] ( #11061 )
2025-07-02 16:36:40 -04:00
chenyu and GitHub
4d57437a67
add timeout to benchmark_search and mlperf action ( #11058 )
...
default timeout is 6 hours which is too long and occupies a box
2025-07-02 14:17:34 -04:00
nimlgen and GitHub
6067568087
nv: remove hardcoded CTRL_CMD_VASPACE_COPY_SERVER_RESERVED_PDES ( #11057 )
2025-07-02 20:41:10 +03:00
qazal and GitHub
ad155f5454
print inputs to get_program in process replay [pr] ( #11051 )
...
* print inputs to get_program in process replay [pr]
* colors
* keep dataclass default escapes
* Revert "keep dataclass default escapes"
This reverts commit c6db7e8a7a .
* note for ast_repr
* add that back
2025-07-02 20:20:01 +03:00
Ignacio Sica and GitHub
a22aa77c82
cleanup opts_to_apply ( #11055 )
...
* fix kernelinfo init in fixup_ast
* opts_to_apply None
2025-07-02 20:03:19 +03:00
qazal and GitHub
a919b8325b
add test_kernel_info ( #11054 )
...
* add test_kernel_info
* reorder
2025-07-02 19:48:12 +03:00
3b041d188f
[bounty] Singular Value Decomposition ( #10875 )
...
* inital commit
* add qr + expand svd to full matrix
* add odd number support
* add linalg tests
* qr supports dims of arbitrary size
* add qr tests
* svd supports dims of arbitrary size
* small cleanip
* improvements over svd batch handling
* improve linalg tests
* make u_pad match q shape
* add nonfull matrix tests
* little less verbose nonfull svd test
* added dtypes on svd + return vt instead of vt
* lint
* more lint
* lint + set seed
* small fix
* small lint
* lint
* add int casting to indices and shapes
* remove int from shape tuple in svd
* small cleanup
* add return types
* reuse inverse_permute
* refactoring
* whitespace
* remove regularization term to prevent bad outputs on ill conditioned matrices
* remove seed
* refactor
* lint
* refactor
* spacing
* remove clone
* line reduction
* smarter heuristic for iterations_per_round
* add big test
* lint
* turns out no constant needed?
* wrap tests
* some small matrices need the constant
* remove realize
---------
Co-authored-by: George Hotz <[email protected] >
2025-07-02 09:06:03 -07:00
Ignacio Sica and GitHub
fc42c3063e
use kernel info ( #11049 )
...
* use kernel info
* keep api
* revert change in comment
2025-07-02 08:42:32 -07:00
e992ed10dc
WebGPU on Windows ( #10890 )
...
* WebGPU on Windows
* Fix dawn-python install
* New test
* pydeps
* Minor fix
* Only install dawn-python on windows webgpu
---------
Co-authored-by: George Hotz <[email protected] >
2025-07-02 08:38:45 -07:00
nimlgen and GitHub
e67a6d2310
nv: tiny cleanups ( #11053 )
2025-07-02 18:37:32 +03:00
chenyu and GitHub
4626e9c172
is_numpy_ndarray helper [pr] ( #11050 )
2025-07-02 09:12:53 -04:00
qazal and GitHub
452b22c9b6
fix process replay diff in PYTHON device [pr] ( #11052 )
...
* fix process replay diff in PYTHON device [pr]
The PYTHON backend pickles and encodes UOps, the encoded binary can't be
directly diffed in process replay.
* note
2025-07-02 11:06:46 +03:00
8ebf0abaae
ONNX external_test_onnx_backend use PYTHON device for model ( #10915 )
...
* try
* ruff check --fix
* no skip test
* hmmmmmmm I don't get this D:
* run CI again
* why is PYTHON device faster than CPU?
* run ci again and fix lint
* actually doesn't PYTHON device make sense here?
* see cpu speed again
* Revert "see cpu speed again"
This reverts commit 1e366f2256 .
* trigger CI
* pretty good
---------
Co-authored-by: chenyu <[email protected] >
2025-07-01 12:11:17 -04:00
qazal and GitHub
8b0871ac31
viz: test for no lockup on infinite loop ( #11041 )
...
* viz: add test infinite loop fallback
* assert
* continue til the end
* work
* bring that back
* fallback to nop
2025-07-01 17:44:20 +03:00
fcbefde8f5
fix DiskDevice reuse ( #11039 )
...
* fix DiskDevice reuse
* fix mypy and DiskDevice.count
* mypy
* add test
---------
Co-authored-by: b1tg <[email protected] >
2025-07-01 10:29:21 -04:00
geohot
5628e2054c
hotfix: if no ranges, return None
2025-06-30 18:07:56 -07:00
George Hotz and GitHub
0597735f28
remove TC=3 not porting this ( #11045 )
2025-06-30 15:12:49 -07:00
geohot
cccfe6b422
hotfix: test_no_inf_loop_bottom_up
2025-06-30 14:21:45 -07:00
George Hotz and GitHub
752c76ceb7
tc3 shape expand [pr] ( #11043 )
...
* tc3 shape expand [pr]
* remove unused stuff in lowerer
2025-06-30 13:38:14 -07:00
George Hotz and GitHub
539b17fcbf
expand local shape so shapes work [pr] ( #11042 )
2025-06-30 13:03:31 -07:00
nimlgen and GitHub
9ea7deb515
hcq: select_iface shared ( #11033 )
...
* hcq: select_iface shared
* errs
* sorry
* upprt
2025-06-30 21:12:39 +03:00
qazal and GitHub
013085da7d
viz: only path "/" serves the UI ( #11037 )
...
The dict used to exist for /profiler and main localhost:8000, we don't
need it anymore.
2025-06-30 19:10:33 +03:00
George Hotz and GitHub
b829331219
infinite loop detect in fixed_point_rewrite [pr] ( #11038 )
2025-06-30 08:57:29 -07:00
Nino Risteski and GitHub
bc15e98f5c
clean up unused imports in examples and update CI linting ( #11024 )
...
* clean up unused imports in examples
* enable unused import checking in examples
* lint
* ignore F541 and F841 - focus on unused imports only
* clean up
* restore tinygrad.frontend.torch for TINY_BACKEND
* tiny change
2025-06-30 08:21:27 -07:00
George Hotz and GitHub
cb531dba42
detect infinite loop in graph rewrite [pr] ( #11036 )
2025-06-30 08:15:13 -07:00
qazal and GitHub
710d734ce7
viz: don't need PICKLE_BUFFER=0 in capture ( #11031 )
2025-06-30 16:20:04 +03:00
qazal and GitHub
2ea4737930
viz: fix newlines breaking label colors ( #11030 )
...
* viz: fix newlines breaking label colors
* TestViz.test_colored_label
* TestWordWrap
2025-06-30 13:39:44 +03:00
George Hotz and GitHub
5911b71404
early support for bidirectional pattern matcher ( #11027 )
...
* early support for bidirectional pattern matcher
* expose it and add a test
* no bottom up arg there
* disable flaky test
2025-06-29 16:54:07 -07:00
George Hotz and GitHub
ec1d97191d
minor cleanup to lowerer [pr] ( #11026 )
...
* minor cleanup to lowerer [pr]
* add that rule to sym
2025-06-29 11:01:29 -07:00
Piyush and GitHub
454bc3393d
redundant code ( #11014 )
2025-06-29 09:06:10 -07:00
qazal and GitHub
19b11cb778
hotfix: check canvas exists before access ( #11022 )
2025-06-29 14:44:14 +03:00
chenyu and GitHub
126fcf4129
clean up AMD_LLVM in tests ( #11021 )
2025-06-28 22:45:47 -04:00
qazal and GitHub
cb6a66ea84
viz: remove per schedule renderMemoryGraph ( #11019 )
...
replaced with per device Buffer viz https://github.com/tinygrad/tinygrad/pull/10960
2025-06-28 22:09:38 +03:00
qazal and GitHub
4c8d2a0383
buffer viz ( #10960 )
...
* add mem_layout
* ui
* cleanup
* work
* debugLine work and expander
* tooltip style
* real expand device
* wheel does one thing
* diff
* shows llama oom
* add y axis
* mypy chill
* work
* unittests for the memory layout
2025-06-28 21:50:32 +03:00
qazal and GitHub
e3d024afa0
viz: split into scale, shapes, axes last ( #11018 )
...
* viz: split into scale, shapes, axes last
* set zoom on render
2025-06-28 19:10:58 +03:00
qazal and GitHub
508bc68078
viz: small fixups from memory graph ( #11017 )
...
* don't need div.id
* tooltip z-index
2025-06-28 16:34:14 +03:00
qazal and GitHub
fc3e509822
viz: new canvas on first render ( #11016 )
2025-06-28 16:04:51 +03:00
chenyu and GitHub
c14c9a8eff
llama3 grad clip ( #11003 )
2025-06-27 19:14:12 -04:00
nimlgen and GitHub
e53673a0b2
amd: sdma queue overrun fix ( #11012 )
...
* amd: sdma queue overrun fix
* add ()
* fix
* bug
* this is correct
2025-06-28 01:42:03 +03:00
chenyu and GitHub
f2548afeb5
bert grad clipping start with const 0 ( #11008 )
...
saved the init kernels
2025-06-27 18:02:23 -04:00
chenyu and GitHub
a6485d00c8
very tiny generate_dataset ( #11013 )
...
one minute to gen on my mac
2025-06-27 17:10:45 -04:00
qazal and GitHub
382fa6a325
viz: support axis colors in UOp nodes ( #11009 )
...
* work
* javascript
* optional defaultColor
* fine
2025-06-27 23:02:55 +03:00
qazal and GitHub
44257f25e4
bump line count to 14600 ( #11010 )
2025-06-27 22:48:14 +03:00
George Hotz and GitHub
be53ef4f0a
rename DEFINE_ACC -> DEFINE_REG ( #11006 )
...
* rename DEFINE_ACC -> DEFINE_REG
* add CMPEQ to groupops
2025-06-27 11:09:25 -07:00
George Hotz and GitHub
05c35d0db8
reorder ops and add comments ( #11005 )
2025-06-27 10:52:14 -07:00
George Hotz and GitHub
5a1911b7c4
apply the global dims late ( #11002 )
...
* apply the global dims late [pr]
* late gpudims
* tests passing
* remove the random local_dims inc
* simpler
2025-06-27 09:54:34 -07:00
qazal and GitHub
4ef10c57f9
remove unused test helper ( #10999 )
2025-06-27 13:48:48 +03:00
qazal and GitHub
a39343e39f
viz: move timeline layout to python ( #10998 )
...
* viz: move timeline layout to python
* DevEvent has a device and a name
2025-06-27 13:06:00 +03:00
George Hotz and GitHub
b4eb876d5a
kernel.py no longer permutes reduce axis [pr] ( #10968 )
...
* kernel.py no longer permutes reduce axis [pr]
* delete tests that handcode uops
* regen of sops is broken...
* put import back
* just remove that
* disable those tests
2025-06-26 17:44:58 -07:00
chenyu and GitHub
6ab5a5cb6c
llama3 mlperf train ( #10983 )
...
work in progress. now it can overfit small examples and vram roughly matches
2025-06-26 20:24:27 -04:00
George Hotz and GitHub
856759c79c
add halide example ( #10980 )
...
* add halide example
* upd halide gemm
* partial works
* touchups
2025-06-26 16:14:57 -07:00
qazal and GitHub
1127302c46
move perfetto to extra ( #10994 )
...
* move perfetto to extra
* update TestViz and fix tests
* remove perfetto.html from viz directory
* work
* mypy
2025-06-27 01:53:54 +03:00
qazal and GitHub
712980e167
fix extract_dataset + add tests to CI ( #10995 )
...
* fix extract_dataset + tests
* add CI
* sops.gz itself is same as master
* yml + gzip -c + ge
* don't commit that
* bump limit to 1000
* axis=7
* test_tiny
2025-06-27 01:51:36 +03:00
chenyu and GitHub
4572e65f0f
remove duplicated move_early logic in UOp.r [pr] ( #10993 )
2025-06-26 18:33:54 -04:00
Ignacio Sica and GitHub
579194f523
remove some linearize calls from tests 2 [pr] ( #10992 )
...
* refactor count_float4 to take uops as input instead of kernel
* remove some calls to linearize in test_linearizer
* remove some more calls
* remove one more call
2025-06-26 18:22:27 -03:00
50936b4a18
ONNX real float16 ( #10694 )
...
* squash commits
* temp fix for const tensor
* actually realizing float16 can only happen in raw_data
* .float -> cast(float) to rerun CI
---------
Co-authored-by: chenyu <[email protected] >
2025-06-26 14:05:12 -04:00
qazal and GitHub
73484b0803
viz: generic shape tooltip/click handlers + renames ( #10990 )
...
* viz: generic tooltip
* assign kernel
* labelParts/label
* rect with a fillColor
* line
2025-06-26 19:14:04 +03:00
qazal and GitHub
7f79c1388f
viz: update y offset calculation ( #10987 )
...
* viz: update y offset calculation
* don't rescale padding
2025-06-26 12:05:20 +03:00
chenyu and GitHub
49bba2f0a0
improve test_nll_loss ( #10986 )
...
build target and weight tensors outside so it tests backward too.
2025-06-26 02:46:55 -04:00
chenyu and GitHub
0612acfc70
improve Tensor.cross_entropy ( #10985 )
...
separate when Y is prob vs indices and check shapes for indices. also fix higher dim cases
2025-06-26 01:39:48 -04:00
chenyu and GitHub
8751d47985
CosineAnnealingLRWithWarmup ( #10981 )
2025-06-25 17:45:21 -04:00
Ignacio Sica and GitHub
21f1c4cc09
remove some linearize calls from tests [pr] ( #10978 )
...
* remove some linearize calls from tests
speed_compare_cuda_ptx
test_uop_spec
test_linearizer
test_uops
test_winograd
* more clear assert message
2025-06-25 12:37:17 -07:00
chenyu and GitHub
efad567ebd
ruff check whole examples/mlperf/ ( #10979 )
2025-06-25 12:57:48 -04:00
Sieds Lykles and GitHub
15e60caf09
add Ops.EQ ( #10976 )
2025-06-25 11:25:10 -04:00
Ignacio Sica and GitHub
98d2cde293
revert tc_group feature ( #10971 )
2025-06-24 20:58:13 -07:00
George Hotz and GitHub
306dbc76f6
early view simplify ( #10974 )
...
* shape const if it has a device [pr]
* early view simplify
2025-06-24 20:52:45 -07:00
77fff73295
fix viz vscode link on windows ( #10972 )
...
Co-authored-by: b1tg <[email protected] >
2025-06-25 06:47:59 +03:00
George Hotz and GitHub
9d995c2a4d
shape const if it has a device [pr] ( #10969 )
2025-06-24 16:22:54 -07:00
George Hotz and GitHub
cf60ccac6a
support new const lowering ( #10967 )
...
* support new const lowering
* delete invalid linearizer failure tests
2025-06-24 15:21:41 -07:00
geohot
8a65720528
hotfix: disable test_tensor_core_opts_group test on real metal
2025-06-24 15:21:33 -07:00
nimlgen and GitHub
1c45b9f7fb
start nvpci ( #10521 )
...
* start nvpci
* talk to fsp
* boot args
* riscv core bootted
* q
* agen
* got gsp init msg
* some fixes
* set registry, stuck aft lockdown(
* start ga/ad port
* gsp init on ada
* more classes allocated
* more
* mm
* fixes and progress
* no huge pages for now
* mm seems workin, but switch to 512mb page for simplicity
* working state
* not cleaned
* claned
* nvd=1
* start gr ctx
* compute
* clean 1
* cleanup 2
* cleanup 3
* cleaner 4
* cleaner 6
* add iface to nv
* save before reboot
* merged into NV
* moveout mm
* post merge
* cleaner 7
* merge and rebase
* pciiface abstraction + reset
* download fw from web
* print logs
* minor changes + p2p
* cleaner 8
* cleaner 9
* cleaner 10
* delete
* delete this as well
* linter 1
* oops
* priv_client -> priv_root
* fix mypy
* mypy?
* mypy?
* small changes
* shorter
* ops
* remove this
* do not allocate paddr for reserve
* nodiff
* unified script
* ops
* dif ver
* add lock
* setup
2025-06-25 00:37:34 +03:00
uuuvn and GitHub
c8d0f68763
Weaker renderer validation in remote ( #10964 )
...
```
training bert
training on ['REMOTE:0', 'REMOTE:1', 'REMOTE:2', 'REMOTE:3', 'REMOTE:4', 'REMOTE:5']
Traceback (most recent call last):
File "/home/uuuvn/src/tinygrad/examples/mlperf/model_train.py", line 1300, in <module>
with Profiling(enabled=getenv("PYPROFILE")): globals()[nm]()
^^^^^^^^^^^^^^^
File "/home/uuuvn/src/tinygrad/examples/mlperf/model_train.py", line 975, in train_bert
for x in GPUS: Device[x]
~~~~~~^^^
File "/home/uuuvn/src/tinygrad/tinygrad/device.py", line 22, in __getitem__
def __getitem__(self, ix:str) -> Compiled: return self.__get_canonicalized_item(self.canonicalize(ix))
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/home/uuuvn/src/tinygrad/tinygrad/device.py", line 28, in __get_canonicalized_item
ret = [cls for cname, cls in inspect.getmembers(importlib.import_module(f'{base}.runtime.ops_{x}')) \
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/home/uuuvn/src/tinygrad/tinygrad/runtime/ops_remote.py", line 417, in __init__
if not renderer[0].startswith("tinygrad.renderer.") or not renderer[1].endswith("Renderer"): raise RuntimeError(f"bad renderer {renderer}")
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
RuntimeError: bad renderer ('tinygrad.runtime.ops_null', 'NullRenderer', ())
```
2025-06-24 14:15:09 -07:00
George Hotz and GitHub
c2f5f0f198
more robust reduce_gradient ( #10965 )
2025-06-24 14:09:33 -07:00
George Hotz and GitHub
8743ca40e2
force reduce to be in axis order ( #10837 )
...
* force reduce to be in axis order
* disable rule causing loop
* disable that rule
* no ra there
* only move non reduce
* fix tests
2025-06-24 13:00:16 -07:00
chenyu and GitHub
ffb032e31d
test_diagonal touchup ( #10962 )
2025-06-24 15:51:19 -04:00
7f9958b632
Fix torch.linalg.diagonal crash due to invalid shrink in to_movement_ops ( #10945 )
...
* fix as_strided shrink bug breaking torch.linalg.diagonal on tinygrad backend
* cleanup
* generic fix
* tests
* cmp with diagonal too
* oops
* move tests
* fix test
* remove unnecessary import
* fix assert
* compare against numpy
---------
Co-authored-by: Utkarsh Gill <[email protected] >
2025-06-24 15:36:06 -04:00
nimlgen and GitHub
26ddf8d714
amd: rename dev_iface -> iface to match nv ( #10959 )
2025-06-24 20:22:19 +03:00
chenyu and GitHub
bfa87f3490
clean up binary_crossentropy_logits ( #10958 )
2025-06-24 12:23:40 -04:00
qazal and GitHub
2ccddfc0ca
viz: match canvas fontsize ( #10957 )
...
it's 10px https://developer.mozilla.org/en-US/docs/Web/API/CanvasRenderingContext2D/font?utm_source=chatgpt.com .
2025-06-24 19:07:06 +03:00
qazal and GitHub
de4b9bf53b
add opts_to_apply option to AST KernelInfo ( #10950 )
...
* proposal: add option to override opts in the get_program API
* update test_linearizer_rewrite
* state in uops
* update process_replay and names
* empty isn't none
* fix process replay
2025-06-24 18:55:39 +03:00
chenyu and GitHub
18e264a449
Tensor.logsigmoid ( #10955 )
2025-06-24 11:16:14 -04:00
Ignacio Sica and GitHub
f15247d2d2
remove outdated index masking in lowerer [pr] ( #10953 )
...
* add assert to check idx is never replaced with const 0
* remove outdated index masking
2025-06-24 07:53:30 -07:00
cc32394b32
support copyin/copyout/is_allocated for subbuffers ( #10869 )
...
* support copyin/copyout/is_allocated for subbuffers
* simple
* clean up
* rm underlying_buf
* add function is_initialized
* add tests
* better test_subbuffer_copy_in_out
* fix allocator
---------
Co-authored-by: b1tg <[email protected] >
Co-authored-by: George Hotz <[email protected] >
2025-06-24 07:49:04 -07:00
chenyu and GitHub
35504c938e
torch.clip(x,y) -> x.clip(y) in test_ops ( #10954 )
...
* torch.clip(x,y) -> x.clip(y) in test_ops
* test_binary_crossentropy_logits_pos_weights
2025-06-24 10:22:19 -04:00
Fang-Pen Lin and GitHub
86d458533f
Add pos_weight for binary_crossentropy_logits ( #10855 )
...
* Add pos_weight for binary_crossentropy_logits
* Remove debug code
* Code style
* Code style
* Rename
2025-06-24 09:42:37 -04:00
Sieds Lykles and GitHub
61dad3740f
fix min_max and add test ( #10952 )
2025-06-24 09:33:26 -04:00
qazal and GitHub
ab8c5d04ab
viz: convert to function_name in server [pr] ( #10951 )
...
* viz: convert to function_name in server [pr]
* it exists
2025-06-24 13:59:37 +03:00
nimlgen and GitHub
c0d9cf09e0
system: flock ( #10949 )
...
* system: flock
* imports
* xx
2025-06-24 11:33:49 +03:00
nimlgen and GitHub
5202970feb
system: move memory_barrier to System ( #10948 )
...
* system: move memory_barrier to System
* fixed
2025-06-24 11:09:43 +03:00
qazal and GitHub
f41c28a048
update test_tensor_uop_representation comments [pr] ( #10946 )
...
These comments can update to match new tinygrad.
2025-06-24 10:47:09 +03:00
qazal and GitHub
7a5e4e0bf1
fix unittests process replay [pr] ( #10947 )
2025-06-24 10:30:23 +03:00
geohot
7d560dbd75
hotfix: corealize in the tiny mnist test
2025-06-23 17:41:16 -07:00
230ad3a460
[bounty] Don't use numpy inside hlb_cifar10 training loop ( #10777 )
...
* Don't use numpy inside hlb_cifar10 training loop
* Lint it
* jit it
* Drop the last half-batch
* Use gather for random_crop and reuse perms
* Wrap train_cifar in FUSE_ARANGE context
* No need to pass FUSE_ARANGE=1 to hlb_cifar10.py
* Add cutmix to jittable augmentations
* Remove .contiguous() from fetch_batches
* Fix indexing boundary
---------
Co-authored-by: Irwin1138 <[email protected] >
2025-06-23 17:24:56 -07:00
George Hotz and GitHub
383010555f
delete linearize and to_program from kernel.py ( #10943 )
2025-06-23 17:04:05 -07:00
George Hotz and GitHub
0f89660ce4
Revert "change clang -march flag to -mcpu on arm ( #10841 )" ( #10942 )
...
This reverts commit 897e42fd1b .
2025-06-23 16:48:28 -07:00
956a8391a5
minor cleanup on test_tensor_core_opts tests ( #10924 )
...
* minor cleanup on test_tensor_core_opts tests
Tests now notify when skipped
Before, they silently skipped if backend didn't had half precision and
accumulation
Also cleaned up atol and rtol setup
* refactor test_tensor_core_opts_group
---------
Co-authored-by: George Hotz <[email protected] >
2025-06-23 16:30:21 -07:00
ttomsa and GitHub
897e42fd1b
change clang -march flag to -mcpu on arm ( #10841 )
...
* change clang -march flag to -mcpu with fp16 disassembly test
* fix
* add capstone to macos dependencies
* just check no cast in test
* rm import
* woops
* lets check
* move check
* llvm init before cpu chcek
* try this
* bump autogen llvm version
* also update libclang?
* revert
* add comment
* skip llvm test and add comment
* linter
2025-06-23 16:28:48 -07:00
Sieds Lykles and GitHub
772cd02ad2
Perform index validation on load/store, not on the index ( #10849 )
...
* move index validation to load/stores
* add name
* add linearizer_failure
* add validate_store with implicit gates
* linearizer_failure_58 is fixed!
* add test_uop_graph test
* rename cond to gate
* test gated load/stores
* use or_casted()
2025-06-23 16:25:05 -07:00
geohot
ae4d2d71b4
bump line count to 14500
2025-06-23 15:32:27 -07:00
Harsh Natuskar and GitHub
79d7cdd9ba
Fix device ( #10929 )
...
* fix: pkg
* better
* added test
* less lines
2025-06-23 15:30:19 -07:00
George Hotz and GitHub
e15754db28
remove (some) kernelize from llama and test schedule speed ( #10939 )
...
* remove kernelize from llama
* 405B
* space
2025-06-23 15:07:31 -07:00
chenyu and GitHub
3699d1d3ba
hotfix llama3 temperature is float ( #10938 )
2025-06-23 15:20:56 -04:00
uuuvn and GitHub
4e2c9e36c7
Remote multihost (p2p transfer) ( #10601 )
2025-06-23 11:47:29 -07:00
chenyu and GitHub
42b1c9625b
skip test TestKiTS19Dataset::test_training_set ( #10936 )
...
flaky
2025-06-23 14:27:24 -04:00
9e9fd44987
refactor test/external/external_llama_eval.py ( #10567 )
...
Co-authored-by: wozeparrot <[email protected] >
2025-06-23 10:43:20 -07:00
chenyu and GitHub
785b4ea8ac
optim flatten().shape[0] is numel ( #10935 )
2025-06-23 13:11:19 -04:00
qazal and GitHub
ac39f27ae6
viz: non blocking UOp tracing ( #10913 )
...
* viz: non blocking UOp tracing
* u.arg
* no if Ops.KENREL
* drop replace
* switch to weakref.WeakKeyDictionary
* back
* remove ram usage skips, viz works here
* cache on reconstruct
2025-06-23 19:59:28 +03:00
Ignacio Sica and GitHub
b8d09a1dae
tc with group/grouptop ( #10903 )
2025-06-23 09:58:41 -07:00
qazal and GitHub
9944c2c02d
viz: show time taken on hover ( #10934 )
2025-06-23 19:00:40 +03:00
geohot
1e99a7f1c9
hotfix: don't viz the indexing rewrites
2025-06-23 08:20:26 -07:00
chenyu and GitHub
f9b59924f1
OPTIM_DTYPE to specify dtype for optim params ( #10925 )
...
one more flag
2025-06-23 10:32:03 -04:00
qazal and GitHub
7820aeca8e
update codegen process replay to use get_program [pr] ( #10921 )
...
* update codegen process replay to get_program [pr]
* precommit
* try str replace
* +to_function_name
* fixup tc
* local2.sh
* fix openpilot NOLOCALS
* new local.sh
* correct merge
* beam cache
* back
* revert beam thing
* adding opts_override and name_override makes output of get_program
reproducible
* min diff
2025-06-23 17:31:41 +03:00
nimlgen and GitHub
eceb7a00d2
nv: rename iface mem functions ( #10931 )
2025-06-23 16:34:51 +03:00
qazal and GitHub
4e864bd304
fix: getenv("NOLOCALS")/NOLOCALS context var ( #10927 )
...
OptOps shouldn't rely on os.environ.
2025-06-23 11:23:59 +03:00
alpharush and GitHub
22f9696522
Fix/hcqfuzz harnesss bug ( #10923 )
...
* update command so extra module is found
* fix empty range in randrange errors
* lint
2025-06-23 11:22:30 +03:00
qazal and GitHub
f037f85532
s/getenv("TC")/USE_TC context var ( #10922 )
2025-06-23 00:39:45 +03:00
qazal and GitHub
9201224e0b
viz: remove Kernel check [pr] ( #10920 )
...
* viz: remove Kernel check [pr]
* TestVizIntegration
* test/unit allows opening of devices
* kernel -> Kernel
2025-06-22 20:47:54 +03:00
nimlgen and GitHub
3ccdb2356b
system: factor out PCIIfaceBase ( #10917 )
...
* system: factor out PCIIfaceBase
* linter
* typing
2025-06-22 20:03:14 +03:00
George Hotz and GitHub
b09c47366f
opt transforms the ast into an optimized ast ( #10900 )
...
* opt transforms the ast into an optimized ast
* fix get_kernel order and to_function_name
* function_name property
* update docs
* copy from kernel.py
* improve docs
* ci didn't trigger?
2025-06-22 09:41:26 -07:00
qazal and GitHub
ffddf165f8
viz: color by kernel names in profiler ( #10919 )
...
* viz: color by kernel names in profiler
* ellipsis stays in bounds
2025-06-22 18:07:52 +03:00
nimlgen and GitHub
36536ef6f0
nv: minor changes from nvpci ( #10918 )
2025-06-22 18:04:39 +03:00
geohotstan and GitHub
4ab7d792cc
ONNX improve dtype fallback ( #10800 )
...
* fix
* add early verbose demo test
* is this how to write tests :s
* is definition drift even a thing? gemini says it is
* clean up
* better
* even better
* try add to CI
* doesn't work quite yet
* much more work to be done
* whoops
* partition the test heh
* skipif
* some nits for better names
* add webgpu test for onnxrunner
* fix reference links
* flush for now
2025-06-21 19:29:45 -04:00
chenyu and GitHub
0480139def
log_perplexity metrics ( #10912 )
2025-06-21 10:44:47 -04:00
nimlgen and GitHub
0e7bd9fd03
factor out generic MemoryManager ( #10910 )
...
* allocator -> memory
* just moveout it
* mm is abstracted
* need entry abstraction
* fix
* mypy
2025-06-21 16:18:33 +03:00
qazal and GitHub
c7ec913210
viz: cleanup unit tests ( #10909 )
...
* cleanup test_viz
* tree view
2025-06-21 12:35:09 +03:00
chenyu and GitHub
1373071f19
simplify logcumsumexp ( #10908 )
...
clarify and remove some flatten and squeeze/unsqueeze
2025-06-20 22:56:42 -04:00
George Hotz and GitHub
fa52bdb50f
applied_opts is in the optimized ast] ( #10906 )
2025-06-20 18:56:23 -07:00
chenyu and GitHub
2d9c61e39e
test more dims in test_logsumexp and test_logcumsumexp ( #10907 )
...
refactoring squeeze and unsqueeze is easy to get wrong
2025-06-20 21:42:18 -04:00
Nino Risteski and GitHub
3771cc0f77
fix test logcumsumexp broken devectorize=0 ( #10880 )
...
* fix test logcumsumexp numerical
* lint
* Use dtypes.min instead of -1e4
2025-06-20 20:54:50 -04:00
George Hotz and GitHub
7636d2cdc5
flip order of get_program args ( #10905 )
2025-06-20 17:23:23 -07:00
George Hotz and GitHub
1ce63f8d04
move functions to view and update docs [pr] ( #10904 )
...
* move functions to view and update docs [pr]
* move quantize
2025-06-20 16:47:58 -07:00
George Hotz and GitHub
b41e0563a3
move stuff to kernelize folder ( #10902 )
...
* move stuff to kernelize folder
* oops, forgot that
2025-06-20 16:10:20 -07:00
George Hotz and GitHub
d399a4587d
move mem estimate to ProgramSpec [pr] ( #10901 )
2025-06-20 15:54:28 -07:00