George Hotz and GitHub
5766865193
Merge branch 'master' into no_merge_views
2025-08-15 08:20:46 -07:00
geohot
addd19d5e1
cleanups
2025-08-14 19:11:46 -07:00
chenyu and GitHub
d0d39885c3
onnx in tinygrad ( #11675 )
2025-08-14 19:57:21 -04:00
geohot
66b92ffc82
one at a time
2025-08-14 16:43:15 -07:00
wozeparrot and GitHub
71260a5ea4
feat: only bench openpilot 0.9.9 models ( #11664 )
2025-08-14 19:27:18 -04:00
chenyu and GitHub
4ddefbccb4
update setup packages ( #11674 )
...
sorted, and added missing 'tinygrad.frontend' and 'tinygrad.runtime.autogen.nv'
2025-08-14 19:24:57 -04:00
geohot
35116959ea
test mnist passes
2025-08-14 16:24:32 -07:00
chenyu and GitHub
48c4033ae1
fix pylint for onnx ( #11673 )
...
* fix pylint for onnx
* too long
2025-08-14 18:48:02 -04:00
geohot
ffa08e9c94
rangeify works again
2025-08-14 15:04:41 -07:00
geohot
c735855dc0
test_plus works
2025-08-14 13:54:13 -07:00
geohot
a5d3b54f47
work
2025-08-14 13:52:14 -07:00
geohot
6131c0aad3
work
2025-08-14 13:17:49 -07:00
chenyu and GitHub
e9d0027591
llama MP realize weight after shard ( #11672 )
...
* llama MP realize weight after shard
prevents memory spike on device 0
* empty weight for FAKEDATA
2025-08-14 16:17:46 -04:00
geohot
46caa43733
k splitting
2025-08-14 11:21:56 -07:00
nimlgen and GitHub
4176b24264
amd: support xcc in regs ( #11670 )
...
* amd: support xcc in regs
* mockamd
* typong
2025-08-14 21:20:11 +03:00
Sieds Lykles and GitHub
f399d0d75d
Render mod in terms of idiv ( #11668 )
...
* Render mod in terms of idiv
* cvar -> var
2025-08-14 19:59:39 +02:00
nimlgen and GitHub
d747eeed32
amd logs parser based on device ( #11669 )
2025-08-14 19:49:33 +03:00
geohotstan and GitHub
1e904155e3
Add Onnx Huggingface to test/models/test_onnx.py ( #11468 )
...
* BOOM
* cache extra/huggingface/models/
* why max buffer size is not 0
* override MAX_BUFFER_SIZE
* less models
* remove more models and change cache dir to already cached dir
* only metal
* less is more?
* remove check ops
* why is this not setting the ENVVAR
* ughhhhh just test in models
* only cpu and gpu
* only cpu actually
* just override it idk
* final
* move extra dependencies up top
* simplification
* fix print
* make README better
* revert ops_disk fix for now
* clean up test_onnx
* remove testing fashion clip model cuz sloooowwwwww
* actually let METAL run this
* fix comment mistake
* fix download path in run_models
* does this work?
* cleanup setup and teardown
* contextvar like this?
* prove model is cached
* do I need to increment DOWNLOAD_CACHE_VERSION?
* see if cached with incremented DOWNLOAD_CACHE_VERSION
* use warnings to see if the model exists
* revert DOWNLOAD_CACHE_VERSION stuff and clean up
* add retry to download
* nit
2025-08-14 11:16:41 -04:00
George Hotz and GitHub
4fd4e13fcf
Merge branch 'master' into no_merge_views
2025-08-14 08:07:52 -07:00
geohot
ab4ccf56a7
no
2025-08-14 08:07:06 -07:00
Sieds Lykles and GitHub
06beeb6e13
Nest div even if factor is negative ( #11666 )
2025-08-14 13:58:59 +02:00
Sieds Lykles and GitHub
661e9a2d5d
div_and_mod_folding refactor ( #11585 )
...
* divmod const folding is its own function
* split nested mod optimization out of div and mod folding
* make `fold_binary_numerator` its own function
* factor out `fold_divmod_congruence`
* check sign of numerator
* add tests
* assert int on vmin and vmax
* add type: ignore
* factor out more rules
* remove div_and_mod_folding
* cached_property to property
* remove import
* add returns
* restore old order
* check sign of x.vmin and newx.vmin
* check more signs
* add some test that would have caught bugs
* better test if the div simplified
* shorten line
* replace terms_factors_const with pop_const
* move that back
* minor cleanup
* remove comments
* some cleanup
2025-08-14 11:52:42 +02:00
geohot
332630ddb5
threefry one kernel
2025-08-13 19:54:09 -07:00
geohot
e5eae3f524
assign becomes store
2025-08-13 19:11:59 -07:00
geohot
cae3616a68
cleanups
2025-08-13 19:02:53 -07:00
geohot
3aa80e7176
rangify bmnist
2025-08-13 18:44:43 -07:00
geohot
b1e2fb9afd
rangeify in
2025-08-13 18:28:19 -07:00
geohot
59bfab8a9b
sym
2025-08-13 17:49:54 -07:00
geohot
b5d7d339f4
no range arg
2025-08-13 17:43:17 -07:00
chenyu and GitHub
0fc43c2e54
fix test_const_tensor_index index ( #11660 )
...
index should be ints
2025-08-13 19:50:16 -04:00
George Hotz and GitHub
b7c195bf7e
Merge branch 'master' into no_merge_views
2025-08-13 16:21:41 -07:00
chenyu and GitHub
4fe19eec72
Ops.TRUNC ( #11659 )
2025-08-13 18:40:48 -04:00
qazal and GitHub
eb10a9c76a
viz: always left align timeline values ( #11658 )
2025-08-13 23:55:28 +03:00
George Hotz and GitHub
8592fba874
Merge branch 'master' into no_merge_views
2025-08-13 12:46:52 -07:00
George Hotz and GitHub
22bdf48cdd
render ranges in viz, name gbufs with sizes. changes from rangeify ( #11656 )
...
* render ranges in viz, name gbufs with sizes. changes from rangeify
* fix unit test dtypes
2025-08-13 12:46:16 -07:00
George Hotz and GitHub
9b4da590bb
remove need for cast_vec ( #11653 )
...
* remove need for cast_vec
* fix amdllvm
2025-08-13 12:09:47 -07:00
geohot
10ffd7e17b
simpler
2025-08-13 11:42:42 -07:00
e2873a3a41
[bounty] Muon optim ( #11414 )
...
* newton schulz
* add muon + move newton schulz to tensor
* compact newton schulz
* better tests
* cleanup
* add comments for muon
* cleanup
* add export with tests
* match muon optim with test optim
* cleanup
* unsed import
* correct comment
* whitespace
* move export
* muon test fix
* match reference impl + tests
* remove export by moving muon device
* add credit
* cleanup
* remove print
* spacing
* spacing
* comma
* cleanup
* removal
* fix tests + optim momentum
* consistent is not/ not
* more consistency
* fix test
* cleanup
* fix the nones
* remove comment
* cast
* comment
* comment
* muon teeny test
* muon flag beautiful mnist
* set steps
* steps as hyperparam
* match default test steps
* name
* large cleanup
* dont care about steps
* nesterov false default
* match each other impl
* steps
* switch nest
* swap defaults
* update docstring
* add no nesterov test
* ban fuse_optim
* prints
* classical momentum
* alternative condition
* recon
* pre + post wd
* false default
* detach
* signature changes
* context
* swap order
* big cleanup
* 0 step instead
* parity
* remove fuse
* remove fused
* better paper
* assert message
* correct shape check + eps
* multidim
* add eps
* cleanup
* correct assert message
* lint
* better tests
* naming
* ns_steps,ns_params
* update docstring
* docstring
* match sgd and muon together
* sandwich
* add back fused
* parity
---------
Co-authored-by: George Hotz <[email protected] >
2025-08-13 14:27:55 -04:00
chenyu and GitHub
94e6d84e32
rewrite Tensor.round to not use cast int ( #11654 )
2025-08-13 13:51:08 -04:00
geohot
5489be812c
random works
2025-08-13 09:47:02 -07:00
geohot
e3d8185ba4
ignore that
2025-08-13 09:43:08 -07:00
George Hotz and GitHub
cbf85fbfd0
Merge branch 'master' into no_merge_views
2025-08-13 09:40:44 -07:00
George Hotz and GitHub
d2521d828a
transcendental+idiv+threefry are uop decompositions ( #11636 )
...
* transcendental+idiv+threefry are uop decompositions [pr]
* threefry decomp
* fix randomness tests
* fix webgpu
* unneeded now
* fix
* move prematcher
* all cast should probably be cast_vec
2025-08-13 09:37:12 -07:00
geohotstan and GitHub
cf7224ce3e
fully lint onnx.py ( #11647 )
...
* mypy
* ruff ruff ruff
2025-08-13 08:22:06 -07:00
geohotstan and GitHub
925555b62a
Fix onnx Domain bug ( #11650 )
2025-08-13 08:20:50 -07:00
Sieds Lykles and GitHub
67df617fe1
add launch bounds to ptx ( #11646 )
2025-08-13 13:05:39 +02:00
qazal and GitHub
88f95e9f59
viz: minor fixups for firefox ( #11645 )
...
* fix circle attr
* set fill color
2025-08-13 12:59:28 +03:00
qazal and GitHub
6f88eac0fc
viz: refactor node and edge tagging ( #11644 )
2025-08-13 12:41:01 +03:00
qazal and GitHub
8140bf9778
viz: create layout once ( #11643 )
...
* start
* work
* works
* diff cleanup
2025-08-13 09:24:58 +03:00
chenyu and GitHub
3fb79bb43a
minor onnx cleanups ( #11642 )
2025-08-13 01:05:19 -04:00
chenyu and GitHub
e9e5a08a04
simplify onnx cubic ( #11641 )
...
we can drop the double where and abs since we know which ranges the inputs map into
2025-08-12 19:57:31 -04:00
George Hotz and GitHub
18cdbec447
split decompositions pass ( #11638 )
...
* split decompositions pass
* fix ptx
* pack load store early
* restore that
2025-08-12 12:56:05 -07:00
chenyu and GitHub
0d8a0d7a96
update test_multi_const_folding_tensor to include pow ( #11635 )
...
pow folds now
2025-08-12 13:35:37 -04:00
Sieds Lykles and GitHub
4d6e407eb0
Extend fast_idiv to negative ints ( #11632 )
...
* fast idiv for signed ints
* Add rule and test
* fix tests
* redo fuzz_fast_idiv to do negative ints as well
* adjust comments
* remove unused imports
2025-08-12 19:34:49 +02:00
qazal and GitHub
17adbe86d8
hotfix: do not default to capturing args in track_rewrites ( #11634 )
2025-08-12 20:01:24 +03:00
ad9dec25b3
combine onnx parser and onnx ( #11485 )
...
* start
* more
* fix onnx_runner test
* pass
* patch for disk and add domains from huggingface
* simpler docs
* revert domain changes
* rerun ci
* revert onnx ops test change
* add fix from strenum stuff
* correct way
* revert correct way to leave the fix for another PR
* test segfault
* Revert "test segfault"
This reverts commit 4e1aaf41e7 .
* remove some unnecessary documentation
* test segfault again
* Revert "test segfault again"
This reverts commit 56fc5f03e7 .
* try gemini suggested patch for sys._getframe
* keep trying with gemini
* revert not working gemini suggestions and try faulthandler
* remove pythonfaulthandler
* trigger CI a few times
* minimize diff
---------
Co-authored-by: chenyu <[email protected] >
2025-08-12 12:56:39 -04:00
Sieds Lykles and GitHub
4c3982c44e
Take sign out of mod ( #11631 )
...
* Add rule and test
* fix tests
2025-08-12 18:44:36 +02:00
qazal and GitHub
e28605e324
rename profile point event fields [pr] ( #11633 )
2025-08-12 19:11:21 +03:00
nimlgen and GitHub
8a7be0a747
metal: workaround for transfers sync issue ( #11622 )
...
* metal: workaround for transfers sync issue
* metal tracsfer sync is broken
* hm
* rm it?
* keep it
2025-08-12 16:16:34 +03:00
qazal and GitHub
efe8b5611d
move ProfilePointEvent out of device.py [pr] ( #11630 )
...
Generic profiling events exist in helpers so they can be imported from
everywhere in tinygrad.
2025-08-12 09:58:32 +03:00
chenyu and GitHub
0d7075f2de
assign should broadcast input tensor ( #11629 )
...
fixed test_assign_broadcast
2025-08-11 23:36:35 -04:00
Joshua Kissoon and GitHub
c44760c89d
torch backend: fix arange, add linalg.cross, add tests ( #11628 )
2025-08-11 23:34:41 -04:00
geohot
11d65cb002
test_gemm works
2025-08-11 18:58:40 -07:00
geohot
5f0816ef69
simpler
2025-08-11 18:41:32 -07:00
George Hotz and GitHub
b8b28e1135
Merge branch 'master' into no_merge_views
2025-08-11 18:29:15 -07:00
George Hotz and GitHub
ca41b5e38b
skip_0 in graph rewrite [pr] ( #11627 )
...
* skip_0 in graph rewrite [pr]
* no track_rewrites on test
* use dict instead of set
2025-08-11 18:29:04 -07:00
Sardor and GitHub
ca7a641442
fix bugs at examples/yolov3.py ( #11614 )
...
* Update load_weight. Give valid model url
* Fix bug in iou function
2025-08-11 21:14:47 -04:00
chenyu and GitHub
0c97d6de1b
don't round pow output for int pow int ( #11625 )
...
also added atol=0 and big pows for the tests
2025-08-11 20:57:47 -04:00
geohot
9d46bc2939
endrange
2025-08-11 17:19:02 -07:00
chenyu and GitHub
d623f6d850
support int Tensor pow to const non-negative int ( #11624 )
...
matches torch
2025-08-11 19:50:19 -04:00
geohot
2b7957e765
map_expand
2025-08-11 15:22:56 -07:00
geohot
4102e46370
cache has issues
2025-08-11 14:35:23 -07:00
chenyu and GitHub
857a830dcc
fix test_arange_float_step ( #11623 )
2025-08-11 16:58:42 -04:00
George Hotz and GitHub
6da4784c66
Merge branch 'master' into no_merge_views
2025-08-11 13:23:14 -07:00
chenyu and GitHub
0806677b51
rewrite sort idx ( #11613 )
2025-08-11 16:20:56 -04:00
George Hotz and GitHub
700c11597b
switch contextvars.ContextVar to _ContextVar ( #11621 )
2025-08-11 12:20:09 -07:00
ae0c3cfff6
change clang -march flag to -mcpu on arm ( #10970 )
...
Co-authored-by: wozeparrot <[email protected] >
2025-08-11 13:38:48 -04:00
geohot
2feeb8c8a6
cleanups
2025-08-11 08:57:30 -07:00
geohotstan and GitHub
27bcb9fd1c
Support cubic mode for ONNX Resize OP ( #11612 )
...
* start
* add reference
* this is so much slower
* this makes sense but differs from official impl, but results are still correct..?
* add a comment
* Just keep it simple for now since I don't fully get it yet
* address comments
* correct
* teeny clean up
* another small comment improvement lol
2025-08-11 11:49:30 -04:00
nimlgen and GitHub
d2bb1bcb97
cloud: a bit better err handling ( #11616 )
...
* cloud: err propagation to client
* fix
* print exc
* linter
* excs
* fix
* hm
* flaky
2025-08-11 15:51:22 +03:00
qazal and GitHub
6a232ccdac
viz: add tiny range drawing helper ( #11620 )
...
* viz: add tiny range drawing helper
* less
2025-08-11 15:15:43 +03:00
qazal and GitHub
e768773e13
viz: use colors helper ( #11618 )
2025-08-11 13:10:15 +03:00
qazal and GitHub
7d6c0a8cc7
viz: refactor progress msg ( #11617 )
2025-08-11 13:01:36 +03:00
chenyu and GitHub
630edcffd8
remove .float calls in olmoe ( #11610 )
...
still matches torch
2025-08-10 20:33:22 -04:00
chenyu and GitHub
a67e0917c3
list indexing can normalize in python ( #11609 )
...
* list indexing can normalize in python
list index does not need to be normalized in tensor
* update those
2025-08-10 20:02:38 -04:00
chenyu and GitHub
1181ec0cd2
few more tensor indexing test cases ( #11608 )
2025-08-10 18:56:42 -04:00
geohot
04fa825a26
careful w the cache
2025-08-10 15:52:41 -07:00
geohot
706188ad16
bugfix
2025-08-10 15:48:07 -07:00
geohot
fbe9909d90
update for master
2025-08-10 15:43:27 -07:00
George Hotz and GitHub
16d2d9daac
Merge branch 'master' into no_merge_views
2025-08-10 15:39:37 -07:00
George Hotz and GitHub
996c907c0b
rewrite not ready + children machinery ( #11607 )
...
* rewrite not ready + children machinery
* it doesn't like track rewrites
2025-08-10 15:28:30 -07:00
geohot
48ca6d888d
was dumb
2025-08-10 14:36:35 -07:00
geohot
b7ea16f161
localish fa
2025-08-10 14:30:25 -07:00
geohot
cc34518a52
RewriteNotReady
2025-08-10 13:56:38 -07:00
Sieds Lykles and GitHub
1875bc69f9
Late rewrite rules for CMPLT ( #11591 )
...
* add rules
* more rules
* fix comment spelling
* remove two rules
2025-08-10 22:18:13 +02:00
geohot
76a97e04b0
cleanups
2025-08-10 12:24:16 -07:00
geohot
0d64aa1f1e
this does work...but with a global
2025-08-10 12:00:55 -07:00
nimlgen and GitHub
5403a4aeaf
null dev: support offset on buffers ( #11606 )
...
* null dev: support offset on buffers
* nolimit
2025-08-10 21:58:37 +03:00
geohotstan and GitHub
b0dab6a4cd
onnx Resize OP clean up ( #11603 )
...
* start
* slight clean up
2025-08-10 14:10:39 -04:00
geohot
9979730f3f
children stuff that doesn't work
2025-08-10 10:57:54 -07:00
Sieds Lykles and GitHub
10540414cd
Add Ops.CMPEQ ( #10431 )
...
* Add op
* add to Groupop.ALU
* fix spec
* fix ptx
* temporary pickle by name to see process replay
* add Ops.EQ to binary ops
* Actuall rename properly
* add test to assert CMPEQ is being used
* Ops.CMPEQ is automatic cast to bool
* add Ops.CMPEQ to llvm
* add Ops.CMPEQ to llvm
2025-08-10 13:13:16 +02:00
chenyu and GitHub
f7aa1b85fe
minor sort cleanups ( #11602 )
2025-08-10 01:51:23 -04:00
chenyu and GitHub
dfb702ef33
fix sort for small dim ( #11601 )
...
* fix sort for small dim
* fixed test_sort_empty
2025-08-10 01:17:41 -04:00
chenyu and GitHub
ef17af85c6
remove .float call in llama logit ( #11598 )
...
* remove .float call in llama logit
* bfloat item
2025-08-10 00:02:18 -04:00
chenyu and GitHub
dd3d2eb36c
add training llama3 test in ci ( #11599 )
2025-08-09 22:35:39 -04:00
chenyu and GitHub
3e64467322
remove freqs_cis contiguous in llama ( #11597 )
2025-08-09 21:11:12 -04:00
chenyu and GitHub
7338ffead0
small beautiful_mnist update ( #11596 )
...
gather is fast now. there's a conv/bw kernel that only gets fast with BEAM, but whole thing runs < 5 seconds now regardless
2025-08-09 19:51:14 -04:00
chenyu and GitHub
45baec1aab
model parallel llama ( #11588 )
...
MP=8 GRADIENT_ACC_STEPS=3 BS=1 DEFAULT_FLOAT=bfloat16 OPTIM_DTYPE=bfloat16 LLAMA3_SIZE=70B SEQLEN=512 PYTHONPATH=. MODEL=llama3 python3 examples/mlperf/model_train.py
2025-08-09 16:54:27 -04:00
nimlgen and GitHub
09bc377da3
search: print runtime failures on debug ( #11593 )
2025-08-09 23:01:19 +03:00
nimlgen and GitHub
14f99ff1a1
amd: doorbell_cpu_addr is not used ( #11592 )
...
* amd: doorbell_cpu_addr is not used
* hm
2025-08-09 20:03:21 +03:00
Sieds Lykles and GitHub
01c770c77b
Fix z3 float cast in indexing ( #11590 )
...
* adjust dtype of z3_renderer and add rule for cast
* dtypes.bool is also cast noop
* add regression test
* make embedding smaller
* even smaller test
2025-08-09 17:59:23 +02:00
Sieds Lykles and GitHub
10d388499d
Refactor optional.py ( #11578 )
...
* move fast_idiv to transcendental
* move optional.py
* adjust comment
* change import
* mypy needs this?
2025-08-09 17:35:05 +02:00
geohot
7ddcb8632f
simpler
2025-08-09 08:14:02 -07:00
nimlgen and GitHub
20e46a175c
do not use disk with usb ( #11119 )
...
* not use disk with usb
* better name
2025-08-09 11:58:02 +03:00
geohot
e268eb2d5c
tform ffn
2025-08-08 18:29:48 -07:00
geohot
ee06481036
ranges
2025-08-08 18:18:27 -07:00
geohot
38c9b5ed2c
conv hack
2025-08-08 14:36:48 -07:00
qazal and GitHub
53179953fc
viz: factor out memory graph render ( #11586 )
2025-08-08 20:18:11 +03:00
geohot
7249a711c2
half contig
2025-08-08 08:43:41 -07:00
qazal and GitHub
8ce72d3fad
simpler disassembly table spec ( #11583 )
...
* simpler disassembly table spec
* update ui
* move to scalar/vec render
2025-08-08 17:59:26 +03:00
qazal and GitHub
44a222a9b2
viz: move resource usage summary to server ( #11582 )
2025-08-08 17:08:28 +03:00
qazal and GitHub
793ace530e
update amd_uop_matmul.py import ( #11581 )
...
Using this for testing SQTT
2025-08-08 17:07:35 +03:00
chenyu and GitHub
b232c60def
benchmark openpilot 0.9.9 ( #11575 )
...
* benchmark openpilot 0.9.9
not sure what to do with the 0.9.7 ones with IMAGE=2 and validate
* name
2025-08-08 01:26:14 -04:00
qazal and GitHub
16f0edbe90
pass opts arg in get_program process replay [pr] ( #11571 )
...
* fix ptx process replay
* keyword arg
* renderer is also optional [pr]
* test_linearizer fixup
* name function order is args,ret,kwargs
* can use opts_to_apply
* pass through p.applied_opts
* sink_arg
* now it opens devices too
2025-08-08 03:05:09 +03:00
qazal and GitHub
960cc6533a
pass through name function args in track_rewrites ( #11572 )
2025-08-08 02:28:52 +03:00
geohot
efdf08f3e2
global rangeify
2025-08-07 15:12:00 -07:00
geohot
9a2f55425b
global rangeify
2025-08-07 15:08:45 -07:00
wozeparrot and GitHub
1826004ef9
feat: add tinyos builder link ( #11570 )
2025-08-07 17:42:18 -04:00
George Hotz and GitHub
9ed409d0f8
Merge branch 'master' into no_merge_views
2025-08-07 14:42:07 -07:00
George Hotz and GitHub
82be8abfd2
move opt under codegen ( #11569 )
2025-08-07 14:19:17 -07:00
chenyu and GitHub
702e38dc19
remove FUSE_ARANGE_UINT ( #11567 )
...
also add IGNORE_OOB=1 to bert runs. lowered BS on tinybox to 90 since 96 oom during eval without reset
2025-08-07 16:49:06 -04:00
George Hotz and GitHub
6ed2dfd187
delete the arange dim mismatch restriction ( #11568 )
...
* delete the arange dim mismatch restriction
* skip that test race
2025-08-07 13:46:17 -07:00
wozeparrot and GitHub
7ae4335127
feat: generate blend index ( #11566 )
2025-08-07 14:20:28 -04:00
chenyu and GitHub
594cbdc66f
skip AM ResNet50 benchmark ( #11565 )
...
hanging with FUSE_ARANGE?
2025-08-07 14:07:01 -04:00
chenyu and GitHub
aa1a6f2132
support threshold in Tensor.softplus ( #11564 )
...
fix gradient for large input
2025-08-07 13:43:18 -04:00
7ee3770961
FUSE_ARANGE=1 ( #11427 )
...
* FUSE_ARANGE=1
* fix test
---------
Co-authored-by: George Hotz <[email protected] >
2025-08-07 13:32:34 -04:00
geohot
4dfcfb1ae5
Revert "Revert "viz: align-center checkbox ( #11555 )""
...
This reverts commit c52facfd29 .
2025-08-07 08:15:57 -07:00
geohot
7e42427a7b
Revert "Revert "viz: remove color for unbind step ( #11554 )""
...
This reverts commit 5650c7b86c .
2025-08-07 08:15:51 -07:00
geohot
dc765fbeb7
Revert "viz: timeline perf ( #11533 )"
...
This reverts commit 031f26632b .
2025-08-07 08:08:51 -07:00
geohot
5650c7b86c
Revert "viz: remove color for unbind step ( #11554 )"
...
This reverts commit 1e205775bd .
2025-08-07 08:08:50 -07:00
geohot
c52facfd29
Revert "viz: align-center checkbox ( #11555 )"
...
This reverts commit 91ec093464 .
2025-08-07 08:08:50 -07:00
geohot
974cfbe76d
Revert "viz: add support for colored tooltip text ( #11556 )"
...
This reverts commit b3f7ea6f93 .
2025-08-07 08:08:49 -07:00
geohot
3bf0db80ef
Revert "viz: pick the largest rect for proxy fillColor ( #11558 )"
...
This reverts commit 76079bc7f2 .
2025-08-07 08:08:48 -07:00
George Hotz and GitHub
9764c6cdee
fix mismatch reduce, try 2 ( #11560 )
...
* fix mismatch reduce, try 2
* fix heuristic
* delete that test
* don't start allowing ones
2025-08-07 07:57:58 -07:00
qazal and GitHub
76079bc7f2
viz: pick the largest rect for proxy fillColor ( #11558 )
2025-08-07 16:40:17 +03:00
nimlgen and GitHub
4f29a2c441
fix flaky test on macos ( #11557 )
2025-08-07 15:55:35 +03:00
qazal and GitHub
b3f7ea6f93
viz: add support for colored tooltip text ( #11556 )
2025-08-07 15:04:43 +03:00
qazal and GitHub
91ec093464
viz: align-center checkbox ( #11555 )
2025-08-07 14:22:02 +03:00
qazal and GitHub
1e205775bd
viz: remove color for unbind step ( #11554 )
2025-08-07 14:16:21 +03:00
nimlgen and GitHub
031f26632b
viz: timeline perf ( #11533 )
...
* viz: timeline perf
* progress
* fast
* less lines
* less lines
* less lines
* fix chrome
2025-08-07 13:16:17 +03:00
George Hotz and GitHub
a1aa5670aa
Revert "fix mismatch reduce ( #11547 )" ( #11549 )
...
This reverts commit 49d21a9055 .
2025-08-06 22:43:15 -07:00
George Hotz and GitHub
49d21a9055
fix mismatch reduce ( #11547 )
...
* fix mismatch reduce
* cleanups
* fix shape
* fix mypy
* resolve
2025-08-06 21:12:51 -07:00
geohot
b8791e962c
don't merge views, mops in kernel
2025-08-06 17:23:31 -07:00
George Hotz and GitHub
21570545d3
move view pushing to codegen, try 2 ( #11534 )
...
* move view pushing to codegen, try 2
* fix up some linearizer tests
* fix test search
* fix test schedule
* delete that test
* fix test arange
* fix a few tests
* update tests
* push views
* ebs cleanup
* fix local/reg
* test and lint
* fix more tests
* test cleanups
* skipped that one
2025-08-06 15:58:38 -07:00
wozeparrot and GitHub
2d5bdc939d
faster llama3 dataloader ( #11540 )
2025-08-06 18:25:57 -04:00
George Hotz and GitHub
80d9cced07
more test cleanups ( #11544 )
...
* more test cleanups
* revert that
2025-08-06 15:05:21 -07:00
George Hotz and GitHub
6fd1332763
update some tests for less Kernel ( #11543 )
...
* update some tests for less Kernel
* get_program update
2025-08-06 14:19:59 -07:00
George Hotz and GitHub
09dc7af8e9
move bind to big graph ( #11539 )
...
* move bind to big graph
* fix tests
* unbind inside kernel only
* merge views
* fix multitensor
* failure text change
2025-08-06 13:27:51 -07:00
George Hotz and GitHub
7c5e115747
test_mismatch_reduce ( #11538 )
2025-08-06 10:02:14 -07:00
George Hotz and GitHub
4fe11725c6
pass through sink arg, update linearizer test ( #11536 )
...
* pass through sink arg, update linearizer test
* get_program help
* bump line count
* use new api
2025-08-06 09:48:48 -07:00
George Hotz and GitHub
bfebb5c37b
do store in the replace_buffers ( #11535 )
2025-08-06 08:42:45 -07:00
geohotstan and GitHub
1163292759
move onnx_parser into onnx ( #11530 )
2025-08-06 10:46:27 -04:00
George Hotz and GitHub
7b16fadd87
load view late + simpler rewrite ( #11525 )
...
* add the load view later
* simpler replace buffers
* rewrite name
2025-08-06 06:55:11 -07:00
nimlgen and GitHub
930d8dae0c
hcq: lazy prof signal allocation ( #11531 )
2025-08-06 15:28:11 +03:00
nimlgen and GitHub
eafc7fda12
upd perfetto ( #11528 )
2025-08-06 14:00:34 +03:00
nimlgen and GitHub
1afb290027
ci: fix runner in nv ( #11527 )
2025-08-06 10:38:04 +03:00
qazal and GitHub
61dae0685c
viz: show total mem in tooltip ( #11526 )
2025-08-06 06:51:26 +03:00
George Hotz and GitHub
cf66df0ea6
put load early to make pointers match ( #11524 )
2025-08-05 20:04:32 -07:00
George Hotz and GitHub
92175626e3
prereqs: move views to codegen ( #11522 )
2025-08-05 19:27:58 -07:00
chenyu and GitHub
c9225d22ce
only disable flaky test_jit_multidev_xfer ( #11523 )
2025-08-05 22:17:25 -04:00
George Hotz and GitHub
f58fd3143d
cleanup fix_kernel ( #11520 )
...
* cleanup fix_kernel
* early load buffer
* early meta ops
* move those to fix_kernel_ops
* fix tests
* remote metal was flaky
* Revert "fix tests"
This reverts commit a27019383d .
* that hack broke things
* fine for ptx
2025-08-05 18:38:43 -07:00
George Hotz and GitHub
067daee5be
pin torch to 2.7.1 ( #11519 )
2025-08-05 15:58:57 -07:00
George Hotz and GitHub
b39f43c46a
optimize in rewrite, try 2 ( #11518 )
...
* changes
* fix test uops
* optimize in rewrite, try 2
2025-08-05 15:52:53 -07:00
geohot
07b0df0d86
hotfix: test tensor dims start at 1
2025-08-05 15:40:24 -07:00
George Hotz and GitHub
4dabdf7c6d
Revert "optimize in rewrite ( #11516 )" ( #11517 )
...
This reverts commit 3b777a9e05 .
2025-08-05 15:39:07 -07:00
George Hotz and GitHub
3b777a9e05
optimize in rewrite ( #11516 )
...
* changes
* fix test uops
* dim shouldn't be 0
* huh, why did that one not save
2025-08-05 15:33:26 -07:00
nimlgen and GitHub
ec676eddfa
nv: move base address higher ( #11514 )
2025-08-05 22:42:53 +03:00
qazal and GitHub
7703f8b805
viz: skip flops info if estimates is symbolic ( #11513 )
2025-08-05 22:12:52 +03:00
nimlgen and GitHub
fc4e713d1c
jit graph split tests ( #11507 )
...
* jit graph split tests
* fix
* one more test
* more tests
* fix
* xm
* rmeote
2025-08-05 21:32:37 +03:00
George Hotz and GitHub
c57fde51f9
move swizzler to opt ( #11509 )
2025-08-05 11:31:30 -07:00
chenyu and GitHub
ace8e9a706
fix test_conv2d_winograd ( #11511 )
2025-08-05 12:15:46 -04:00
chenyu and GitHub
223aaa0492
clean up more conv tests ( #11510 )
2025-08-05 12:15:30 -04:00
Garret Castro and GitHub
76e62a1c23
extract conv layer test logic ( #11488 )
...
* refactor: extract conv layer test logic
* tuple is unnecessary
* integrate _test_conv logic into all conv tests
* fix linter, forgot dilation
* undo winograd extraction
adds too many if statements for a single case
2025-08-05 11:15:54 -04:00
8b8bd6c534
make einsum generate same kernels ( #11508 )
...
Co-authored-by: b1tg <[email protected] >
2025-08-05 11:12:52 -04:00
uuuvn and GitHub
011ef8fa9d
Fix incorrect jit current batch devs reset ( #11505 )
...
`current_batch_devs = []` (in `flush_batch()`) happens between
`new_batched_devs = ...` and `current_batch_devs = new_batched_devs` =>
doesn't actually reset anything leading to things not jitting properly
which 2xs remote bert step time (should have similar effects on any
non-hcq backend)
2025-08-05 08:16:16 +03:00
chenyu and GitHub
f02720ca2d
fix fuse gate_contiguous unique ( #11504 )
2025-08-04 23:43:31 -04:00
George Hotz and GitHub
7f6acfb0d5
give define global and friends a shape ( #11502 )
...
* give define global and friends a shape
* ignore negative size
* ptx fix
2025-08-04 19:09:39 -07:00
chenyu and GitHub
83385e7abc
update gradient src in ramp.py ( #11499 )
...
that's simplified now
2025-08-04 18:58:03 -04:00
qazal and GitHub
846a2826ab
viz: remove TracingKey.fmt ( #11482 )
...
* viz: remove TracingKey.fmt
* remove from test too
2025-08-05 00:00:03 +03:00
chenyu and GitHub
01d44e8f16
tiny reduce_gradient cleanup [pr] ( #11498 )
2025-08-04 16:12:53 -04:00
chenyu and GitHub
8a11af01ed
remove broken paperswithcode links in doc ( #11497 )
2025-08-04 13:12:33 -04:00
4f0ee4e982
BPE tokenizer ( #11415 )
...
* BPE works
* refactor tok
* oops
* basic tests
* fix eval
* smaller diff
* fix error
* proper vocab decoding
* use regex for splitting
* escape ucatrange
* full compat
---------
Co-authored-by: George Hotz <[email protected] >
2025-08-04 09:52:38 -07:00
06af9f9236
fix double exception + add name,loc in error msg ( #11487 )
...
Co-authored-by: b1tg <[email protected] >
2025-08-04 13:41:23 +03:00
nimlgen and GitHub
4877aa965a
ast seems to probe nv as well ( #11494 )
2025-08-04 11:47:07 +03:00
chenyu and GitHub
e0106b6b25
1/(x*c) -> (1/c)*(1/x) ( #11491 )
...
example: 2*(2*a).reciprocal() -> a.reciprocal()
# TODO: bounds for reciprocal
# TODO: should z3 work?
2025-08-03 23:35:46 -04:00
qazal and GitHub
5870352fe1
viz: factorize llvm-mca call ( #11490 )
2025-08-04 00:31:23 +03:00
chenyu and GitHub
dbc7807c61
enable WEBGPU tests with buffer limit ( #11489 )
...
TestSample still fails?
2025-08-03 13:02:44 -07:00
nimlgen and GitHub
8f374ee1f7
nv: print devfmr in gsp logs ( #11484 )
2025-08-03 15:12:53 +03:00
chenyu and GitHub
823f1a01db
move cast around expand backward to tensor.py ( #11483 )
2025-08-02 23:03:54 -04:00
chenyu and GitHub
0ce0f51010
generic double cast folding ( #11481 )
...
b.cast(a).cast(b) -> b if a preserves all values in b
2025-08-02 19:26:37 -04:00
qazal and GitHub
72e0d1d0dc
viz: profile the compiler in TINY device ( #11457 )
...
* viz: profile the compiler in TINY device
* leanup
2025-08-03 02:03:20 +03:00
chenyu and GitHub
66be747908
few more dtype cast convinience methods ( #11480 )
2025-08-02 15:47:09 -04:00
chenyu and GitHub
e22e5da9a5
move some test_dtype tests to unit ( #11479 )
2025-08-02 15:25:00 -04:00
nimlgen and GitHub
da0b955be4
hcq: cpu can be graphed ( #11474 )
...
* hcq: cpu can be graphed
* ops
* new jit decisions
* fix test
* fix remote
* cleaner
* fix
2025-08-02 21:01:19 +03:00
chenyu and GitHub
f7965f85aa
Revert "feat: faster index building ( #11462 )" ( #11478 )
...
This reverts commit 3a4deb08d2 .
2025-08-02 12:50:48 -04:00
kevvz and GitHub
ef7e01cadf
Fix SVD shape bug + Fix batched SVD bug ( #11477 )
...
* failing test case
* fix
* better test
* space
2025-08-02 09:47:41 -07:00
6ecaf8e7b2
refactor: use less index and simplify reduce axes check [pr] ( #11476 )
...
* use output_shape/full_shape
* simple final_reduces check
---------
Co-authored-by: b1tg <[email protected] >
2025-08-02 09:44:51 -07:00
wozeparrot and GitHub
3a4deb08d2
feat: faster index building ( #11462 )
...
* feat: faster index building
* feat: correct training samples
2025-08-02 11:50:18 -04:00
nimlgen and GitHub
8cc2d64edb
amd: reuse create_queues for usb iface ( #11473 )
2025-08-02 14:40:46 +03:00
chenyu and GitHub
9e8e6b45ab
grad acc train llama ( #11467 )
...
* grad acc train llama
* log step time
2025-08-01 15:54:50 -04:00
chenyu and GitHub
7ad7329257
data parallel train llama ( #11466 )
2025-08-01 12:13:51 -04:00
nimlgen and GitHub
9f2182f92f
cpu: start threading ( #11324 )
...
* cpu: threading
* syncs
* llvm
* fix
* opt
* fx
* fix
* missed sync
* one line less
* cleaner
* fix
2025-08-01 15:35:07 +03:00
qazal and GitHub
c7ae1bd474
viz: more consistent border styling ( #11464 )
2025-08-01 09:31:06 +03:00
George Hotz and GitHub
8ff03806e8
add llama layers ( #11460 )
...
* add llama layers
* add contig bw for speed
2025-07-31 16:28:04 -07:00
qazal and GitHub
719827b95d
viz: add flops / mem bw to device programs ( #11459 )
...
* viz: add flops / mem bw to device programs
* better spacing style
2025-08-01 02:12:30 +03:00
chenyu and GitHub
3f742a5a7c
comma space lab models benchmark ( #11461 )
2025-07-31 19:06:18 -04:00
geohot
474ee9daa5
hotfix: add contiguous_backward to llama
2025-07-31 15:07:12 -07:00
qazal and GitHub
fa66d9772d
viz: show const node when it's root ( #11456 )
2025-08-01 01:01:58 +03:00
qazal and GitHub
056dabda5a
viz: refactor to color scheme ( #11455 )
2025-08-01 00:17:50 +03:00
nimlgen and GitHub
e5b6149dfb
more typing in drivers ( #11454 )
...
* more typing in drivers
* rm
2025-07-31 23:26:33 +03:00
qazal and GitHub
bad3cf5731
viz: add LLVM machine code analysis ( #11421 )
...
* start
* works everywhere
* add viz api
* utilization table
* reg pressure ui
* use llvm-mca
* llvm-mca ui
* work
* cleanup
* cycle through, defaults are enough
* x86 pending
* x86 nops
* get mcpu/mtriple from autogen
* cleanup server diff
* move parser to python
* normalize to pct of max
* segments legend
* imports
* also monospace
* max comes from the total per instruction
* base on the value
2025-08-01 01:59:26 +08:00
chenyu and GitHub
e847677e8a
use AxisType in search instead of colors ( #11452 )
2025-07-31 13:07:33 -04:00
nimlgen and GitHub
75c2c42def
suppress exceptions only during finalization ( #11451 )
...
* suppress exceptions only during finalization
* fix
* fix typing
* fix more warns
* fix
* better?
* Revert "better?"
This reverts commit a068aa5793 .
* mm?
* no as e
2025-07-31 13:57:12 +03:00
wozeparrot and GitHub
24dd0d52ed
feat: test remove to cpu ( #11444 )
2025-07-30 20:18:56 -07:00
c3cfcb50cb
Add linalg_det and test for torch backend ( #11405 )
...
* add linalg_det and test
* space
---------
Co-authored-by: chenyu <[email protected] >
2025-07-30 22:04:44 -04:00
cba3655de5
Add Test for Setitem ( #10559 )
...
* init
* update
* better
* failing test
* works
* Delete test file
* clean
* lint
* simplify variable name
* rm contigious, rm int dtype, and add assertEqual
---------
Co-authored-by: chenyu <[email protected] >
2025-07-30 22:03:41 -04:00
wozeparrot and GitHub
6252f7770e
feat: fake data ( #11447 )
2025-07-30 17:18:20 -07:00
chenyu and GitHub
e300451f3a
update llama3 ( #11446 )
...
`LR=1e-4 TRAIN_ON_VAL=1 DEFAULT_FLOAT=bfloat16 FUSE_ARANGE=1 JITBEAM=2 OPTIM_DTYPE=bfloat16 LLAMA3_SIZE=1B WARMUP_STEPS=36 DECAY_STEPS=360 SEQLEN=512 PYTHONPATH=. AMD=1 AMD_LLVM=0 MODEL=llama3 python3 examples/mlperf/model_train.py` trained to 7
2025-07-30 19:34:21 -04:00
wozeparrot and GitHub
5fb975351a
feat: flag for training on val ( #11441 )
2025-07-30 14:29:45 -07:00
chenyu and GitHub
4ca430e5bf
fix search dedup ( #11439 )
...
it should check against pre real_axis axis in actions, not real_axis.
2025-07-30 17:24:16 -04:00
wozeparrot and GitHub
d3da20eca6
feat: bump mlperf workflow timeout to 6 hours ( #11440 )
2025-07-30 14:12:12 -07:00
wozeparrot and GitHub
825b6a2505
feat: llama3 dataloader ( #11340 )
2025-07-30 13:27:55 -07:00
qazal and GitHub
af357b5dc8
disable TRACK_MATCH_STATS in BEAM workers [pr] ( #11437 )
2025-07-30 23:22:08 +03:00
George Hotz and GitHub
7c2d2eff86
check tensor core dims ( #11436 )
...
* check elements_per_thread in tensorcore [pr]
* check tc dims
2025-07-30 13:06:59 -07:00
nimlgen and GitHub
5fc5bb5237
ci: clear processes ( #11434 )
...
* unified hcq_smi for managment
* fix
* fix
* no reset for amd
2025-07-30 22:15:18 +03:00
George Hotz and GitHub
4f26a9ad32
check elements_per_thread in tensorcore [pr] ( #11435 )
2025-07-30 11:55:48 -07:00
nimlgen and GitHub
4b4ba5454c
ci: move driver start higher ( #11431 )
2025-07-30 10:48:38 +03:00
George Hotz and GitHub
1bef2d80c1
unrolls are all in the same scope ( #11429 )
...
* unrolls are all in the same scope
* fix that import
2025-07-29 16:55:37 -07:00
chenyu and GitHub
204da24cfc
increase driverbenchmark timeout-minutes to 15 ( #11428 )
2025-07-29 19:45:05 -04:00
chenyu and GitHub
d5fc6af4a2
remove unused ShapeTracker.consecutive [pr] ( #11426 )
2025-07-29 18:36:19 -04:00
George Hotz and GitHub
49a2583584
real new lowerer ( #11419 )
...
* real new lowerer
* fix group for reduce
* skip missing ranges
* fix wmma and unroll/contract
* real fix for wmma
* disable that test
* fix if gate
* simpler
* flash attention fusion works
* no end barriers
* still broken
* flash attention finally works
2025-07-29 15:35:51 -07:00
chenyu and GitHub
0e5d8d5c3c
remove tests that used .to_uop() ( #11425 )
...
* remove tests that used .to_uop()
* import
2025-07-29 15:52:16 -04:00
nimlgen and GitHub
c88e401d0e
ci: fix typos in h machine benchmarks ( #11423 )
2025-07-29 22:11:47 +03:00
chenyu and GitHub
90a5a312eb
simplify ShapeTracker in UOp.const [pr] ( #11424 )
2025-07-29 15:04:06 -04:00
chenyu and GitHub
398594029b
spec checks arg of VIEW are ShapeTracker ( #11422 )
2025-07-29 14:05:12 -04:00
geohot
1f1f99c287
hotfix: add DEBUG=3 to driver CI
2025-07-29 11:03:47 -07:00
George Hotz and GitHub
50fae54175
global local dims in gpudims [pr] ( #11420 )
2025-07-29 10:39:03 -07:00
chenyu and GitHub
9bc413f104
remove ShapeTracker.to_uop [pr] ( #11418 )
2025-07-29 13:29:37 -04:00
George Hotz and GitHub
ba2c4df125
dont render cast ptrs standalone ( #11417 )
...
* dont render cast ptrs standalone
* barrier cleanups
2025-07-29 09:24:26 -07:00
nimlgen and GitHub
d38d285489
ci: add h machines ( #11416 )
...
* ci: add h machines
* more
* fix names
* names not collide
* 20
* 10
2025-07-29 19:21:51 +03:00
2568bc0d99
ci: add caching for apt packages ( #11162 )
...
* add caching for apt packages
* remove 'inputs' from apt cache key, use outputs instead of env
* remove unnecessary mkdir for partial
---------
Co-authored-by: George Hotz <[email protected] >
2025-07-29 09:04:56 -07:00
George Hotz and GitHub
03909f2772
permute locals for HL uop matmul ( #11412 )
...
* permute locals for HL uop matmul
* parens fix that
* permutes
* 20 TFLOPS
2025-07-29 08:19:59 -07:00
nimlgen and GitHub
e0c9747684
amd: fix typo in has_scratch_base_registers for mi350 ( #11413 )
2025-07-29 10:30:06 +03:00
George Hotz and GitHub
735ad5f10d
kernel4 and 5 in uops ( #11411 )
...
* move simplify views to merge views
* add amd kernel 4
* Revert "move simplify views to merge views"
This reverts commit 1e07dff384 .
* k4 in python
* kernel4 written in uops
* k5 support
* cleanups
2025-07-28 19:35:48 -07:00
George Hotz and GitHub
fddc645668
HL=2 top matmul ( #11406 )
...
* HL=2 top matmul
* top colored
2025-07-28 12:32:38 -07:00
nimlgen and GitHub
c7b4ab86e4
fix llvm tc on mi350 ( #11404 )
2025-07-28 21:37:43 +03:00
chenyu and GitHub
9f7c72ff8f
remove UOp.valid method [pr] ( #11402 )
...
only used in add_buffer_ops
2025-07-28 11:29:08 -04:00
chenyu and GitHub
b22a34331b
remove const valid in fixup_ast [pr] ( #11401 )
2025-07-28 11:07:59 -04:00
qazal and GitHub
7737cbb2a0
viz: tabulate runtime stats ( #11400 )
2025-07-28 15:56:39 +03:00
chenyu and GitHub
ab6a27f627
remove a branch in UOp.r [pr] ( #11398 )
2025-07-27 18:00:01 -04:00
052191eae4
Remote multihost (p2p with infiniband verbs) ( #9746 )
...
Co-authored-by: wozeparrot <[email protected] >
2025-07-27 14:44:32 -07:00
qazal and GitHub
a22417cc75
viz: fix bug with wrong program links ( #11396 )
2025-07-28 02:52:06 +08:00
nimlgen and GitHub
a5371f514b
cpu: copies in profile ( #11392 )
...
* cpu: copies in profile
* fix
* rename to tiny?
2025-07-27 20:56:27 +03:00
George Hotz and GitHub
8c10085459
assert shape on lowerer store [pr] ( #11395 )
...
* assert shape on lowerer store [pr]
* fix ptx
2025-07-27 10:41:57 -07:00
qazal and GitHub
6174cfa828
viz: only show match counts greater than 0 ( #11394 )
2025-07-28 00:25:00 +08:00
qazal and GitHub
3466a220de
viz: disassembly viewer ( #11393 )
...
* test
* CPU=1 disasm works
* METAL=1 disasm works
* fix that
* work
* can unwrap
* work p2
* don't crash
2025-07-27 18:44:28 +03:00
qazal and GitHub
3bb232eb29
viz: query path in rewrite steps ( #11391 )
2025-07-27 14:51:47 +03:00
b7ef73babd
fix wmma ptx ( #11389 )
...
Co-authored-by: b1tg <[email protected] >
2025-07-26 23:28:35 -07:00
8dfcdb123d
less wmma args ( #11385 )
...
* less wmma args
* scalar
* ops_python
* mypy
* lint
* dedup
* helper wmma_args
---------
Co-authored-by: b1tg <[email protected] >
Co-authored-by: George Hotz <[email protected] >
2025-07-26 21:24:05 -07:00
George Hotz and GitHub
dfeee63d30
uop matmul work ( #11388 )
...
* uop matmul work
* works with locals
2025-07-26 21:23:55 -07:00
George Hotz and GitHub
3923e78061
no_vectorized_acc keeps single DEFINE_REG ( #11387 )
...
* no_vectorized_acc keeps single DEFINE_REG
* fix ptx, skip flaky test
2025-07-26 11:44:09 -07:00
qazal and GitHub
4866ad57da
viz: add runtime stats ( #11383 )
...
* viz: add runtime stats
* lint
* better
* flat
2025-07-26 20:40:46 +03:00
George Hotz and GitHub
2c70eaf18c
fix load / barrier ( #11386 )
...
* fix load / barrier
* cleanups
* fix CI
2025-07-26 10:27:37 -07:00
nimlgen and GitHub
65673e68ca
hcq: do not import during __del__ ( #11384 )
...
* hcq: do not import during __del__
* ignore
2025-07-26 13:58:55 +03:00
George Hotz and GitHub
466ab5a3f2
store/load not pass through index ( #11381 )
...
* noop
* fix noop
* store cat is NOOP
* store dtype is void
* stores aren't passed through anymore
* meh, skip those for ptx
* correct ptx skip
* hl runs
2025-07-25 21:01:47 -07:00
George Hotz and GitHub
0a5f37946b
unused permute arg on r ( #11379 )
2025-07-25 19:52:37 -07:00
George Hotz and GitHub
48562cb2db
full shape simpler ( #11376 )
2025-07-25 18:27:48 -07:00
chenyu and GitHub
3d68feb67d
minor onnx Gather cleanup ( #11375 )
...
removed a type ignore and one error code skip
2025-07-25 21:08:08 -04:00
chenyu and GitHub
88c338bfcc
add kernelize to keccak for each data block ( #11370 )
...
* add kernelize to keccak for each data block
test_long works now. this prevents internal uops from growing propotional to data length and eventually too deep
* this?
* hash stuff
* gate test
* mv
2025-07-25 16:07:20 -04:00
chenyu and GitHub
dab07bcad9
use next instead of full list in UOp._device [pr] ( #11369 )
...
prevents exponential fan out
2025-07-25 10:04:29 -04:00
nimlgen and GitHub
1bb1f1aee8
hcq: fix race in _at_profile_finalize ( #11368 )
2025-07-25 14:14:02 +03:00
George Hotz and GitHub
490a93902c
define reg doesn't have init anymore ( #11365 )
...
* define reg doesn't have init anymore
* remove that
* no special logic for dr
* fix amd uop matmul
2025-07-24 19:15:49 -07:00
George Hotz and GitHub
9da3f72495
identity store for DEFINE_REG ( #11363 )
...
* identity store for DEFINE_REG
* identity store for DEFINE_REG
* noop continue
2025-07-24 16:41:29 -07:00
chenyu and GitHub
cc795c6656
simplify keccak pad mask code ( #11362 )
2025-07-24 19:24:10 -04:00
chenyu and GitHub
c0c4bc9d7c
use int32 for keccak reorder_indexes ( #11360 )
...
it's used for tensor indexing, so int32 instead of uint64 is slightly faster
2025-07-24 15:54:50 -04:00
George Hotz and GitHub
0602b22086
kernel spec ( #11359 )
...
* kernel spec
* ops.VIEW
* work
2025-07-24 12:45:38 -07:00
qazal and GitHub
519f1d13cc
viz: generic stuff from gpu counters ui ( #11358 )
...
* viz: generic stuff from gpu counters ui
* move pointer
* pre fetch
* move timeout
2025-07-24 20:29:24 +03:00
nimlgen and GitHub
3b3de8df61
hcq: graphed copies ( #11302 )
...
* fast copies p2
* upd and fix
* graph supports
* fixes
* fixes
* fixes
* fix
* fix
* fix mockgpu
* fix alignment
* smaller in ci
2025-07-24 17:36:19 +03:00
nimlgen and GitHub
3046ead6e8
jit: graph reports ei support ( #11356 )
2025-07-24 16:35:10 +03:00
nimlgen and GitHub
bf12041910
hcq: mapping of cpu to all hcq devices ( #11354 )
...
* hcq: mapping of cpu to all hcq devices
* fix kfd
* nv
* simpler
* cleaner
* correct skip
* fix ifaces
* system fixes
* mypy
2025-07-24 12:52:38 +03:00
chenyu and GitHub
82e6de7fc6
more keccak reference tests ( #11329 )
2025-07-23 22:06:39 -04:00
George Hotz and GitHub
b0dc97d1f7
write out kernel 3 in uops ( #11352 )
...
* write out kernel 3 in uops
* matmul is correct
* gemm passes spec
* bugfix to match speed
* cleanups
2025-07-23 17:32:38 -07:00
chenyu and GitHub
5b570196e4
support DEV= to specify device ( #11351 )
2025-07-23 17:40:55 -04:00
76a2ddbd78
Move remote tests out of onnx ( #11310 )
...
Co-authored-by: wozeparrot <[email protected] >
2025-07-23 13:25:55 -07:00
George Hotz and GitHub
7f0a41df4d
move optional out of devectorize [pr] ( #11350 )
...
* move optional out of devectorize [pr]
* fast idiv
2025-07-23 11:26:05 -07:00
nimlgen and GitHub
0f374e10d2
cpu: use mmap for allocations ( #11349 )
...
* cpu: use mmap for allocations
* ops
* fix mypy
2025-07-23 20:30:18 +03:00
George Hotz and GitHub
ae07a93814
simple block barrier ( #11341 )
...
* simple block barrier
* simple
2025-07-23 10:14:11 -07:00
chenyu and GitHub
86e7504111
mypy check extra/onnx.py ( #11348 )
...
instead of running test with 3.10, add onnx to mypy which would have caught StrEnum regression. Several type annotation failed mypy now that does not affect running the code and were skipped for now
2025-07-23 12:42:59 -04:00
chenyu and GitHub
960da9319d
Remove StrEnum in onnx for python 3.10 ( #11345 )
...
some training tests failed looks like parsing error?
2025-07-23 11:52:25 -04:00
qazal and GitHub
478a355325
gate PRINT_MATCH_STATS behind graph_rewrite tracking ( #11344 )
2025-07-23 16:32:43 +03:00
nimlgen and GitHub
ca09c180dc
cpu: remove del spam ( #11343 )
...
* cpu: remove del spam
* fix
2025-07-23 12:02:37 +03:00
nimlgen and GitHub
304eb9cecb
allocate less memory in am tests ( #11342 )
2025-07-23 11:11:26 +03:00
George Hotz and GitHub
e14b4fefa5
ranges on store ( #11334 )
...
* ranges on store
* fix store spec
* fix that
* fix gates
* fix tests
* fix ptx
2025-07-22 21:00:50 -07:00
George Hotz and GitHub
c65b5aab62
small things from endrange ( #11339 )
...
* small things from endrange
* store
2025-07-22 19:45:37 -07:00
George Hotz and GitHub
53339e62f7
no gate store anymore ( #11338 )
...
* no gate store anymore
* fix up spec
2025-07-22 18:41:15 -07:00
chenyu and GitHub
7a9a5cfd28
isolate test/external/external_test_am.py ( #11335 )
...
seems to be the one crashing, also remove -n=auto for that
2025-07-22 19:02:20 -04:00
George Hotz and GitHub
fcbd0e4de3
assigns are no longer used [pr] ( #11333 )
2025-07-22 15:35:07 -07:00
George Hotz and GitHub
09431d4ad1
make DEFINE_REG behave like the others ( #11273 )
...
* simpler define reg
* cast
* PTRCAT define_acc
* cleanups
* fix uops stats
* fix linearizer tests
* llvm
* define reg sets const
* define reg sets const
* no assign
* collapse that
* fix test_max_pool2d_bigger_stride_dilation
* use index, fix webgpu
* devec
* fix tests
* fix webgpu
* fix llvm
* threads for python
* fix ops_python
* only for reg
* acc_half is real now in the emulator
* fix llvm
* fix webgpu init
* fix wgpu test
* fix some tests
* fix ptx
* fix ptx bool acc
* cleanups
* broken, meh. will fix with ENDRANGE
* line count
2025-07-22 13:53:56 -07:00
chenyu and GitHub
4535908679
update keccak test_long ( #11331 )
...
it should compare with arg "shake_128"
2025-07-22 16:08:01 -04:00
nimlgen and GitHub
3faa352dcc
am: bump version after mm changes ( #11328 )
2025-07-22 21:54:10 +03:00
George Hotz and GitHub
affd83961c
small changes from define_reg ( #11327 )
...
* small changes from define_reg
* fix webgpu
2025-07-22 11:11:48 -07:00
nimlgen and GitHub
53b3d87456
am: use 4-lvl pdir ( #11326 )
2025-07-22 20:58:15 +03:00
chenyu and GitHub
2d7c28de6a
clean up dup lambdas in helper_test_exception ( #11325 )
2025-07-22 12:21:57 -04:00
chenyu and GitHub
c6aa8e58ca
fix TestDropoutProbabilityEdgeCases ( #11322 )
2025-07-22 11:13:56 -04:00
chenyu and GitHub
fb42c84365
merge TestRollEdgeCases into test_ops ( #11321 )
2025-07-22 10:55:57 -04:00
chenyu and GitHub
1d8b3e9d1c
movementop only Tensor.roll ( #11317 )
...
* movementop only Tensor.roll
* fixed
2025-07-22 10:34:15 -04:00
chenyu and GitHub
a41140241b
truncate unsigned const in cstyle ( #11318 )
...
it can be a warning or a hard error in clang
PTX and PYTHON also need fix, skipping for now
2025-07-22 08:02:12 -04:00
qazal and GitHub
6668d6d241
fix word_wrap with newlines in input string [pr] ( #11319 )
2025-07-22 12:03:13 +03:00
qazal and GitHub
0c4e19f270
hotfix: disable process replay in REMOTE=1 tests ( #11320 )
...
* hotfix: disable process replay in REMOTE=1 tests
* comment
2025-07-22 10:41:58 +03:00
George Hotz and GitHub
3b674df34b
generic changes from define_reg_2 ( #11315 )
...
* generic changes from define_reg_2
* fix for ptx
* ugh, that one
2025-07-21 15:14:06 -07:00
chenyu and GitHub
6e9506e6fd
Tensor.roll supports dims=None ( #11313 )
2025-07-21 17:29:23 -04:00
George Hotz and GitHub
108aac8af4
use AddrSpace instead of local ( #11314 )
...
* use AddrSpace instead of local
* addrspace in test
2025-07-21 14:00:06 -07:00
chenyu and GitHub
d3a93185a6
clean up test_roll ( #11312 )
2025-07-21 16:00:50 -04:00
George Hotz and GitHub
532b52fcef
store has a dtype, like assign ( #11309 )
...
* store has a dtype, like assign
* fix upat
* fix test
2025-07-21 12:50:01 -07:00
445ff8de56
ONNX onnx_parser and buffer_parse clean up ( #11000 )
...
* start
* remove onnx.load from compile4 and move np to dropout
* clean up and enable test
* clean up
* move WebGPU ONNX test into MacOS (WebGPU)
* leave test in ONNX (CPU)
* fix raw_data init None, and simplify onnx_runner test a little?
* THESE TESTS ARE SO UGLY UGHH
* need to really think about how to structure the test
* wow LLMs are quite something
* not always on disk now
* also add external data loading test
* cleaner tests
* minimize diff and add const folding tests
* add external data loading too
* whoops add webgpu back.. but why was it not needed in the first place?
* better comment
* move webgpu test to macos(webgpu)?
* llm english so much better than me wow
* trigger CI to check flakiness
---------
Co-authored-by: chenyu <[email protected] >
2025-07-21 15:10:25 -04:00
George Hotz and GitHub
842184a1ab
rename kernelize to schedule, try 2 ( #11305 )
2025-07-21 11:18:36 -07:00
George Hotz and GitHub
7e8f5dde74
matmul style is still reshape ( #11308 )
2025-07-21 11:14:57 -07:00
George Hotz and GitHub
41de76a7fd
put assign and store next to each other [pr] ( #11306 )
2025-07-21 11:07:35 -07:00
nimlgen and GitHub
de2df92551
hcq: use devices instead of ids in HCQGraph ( #11303 )
...
* hcq: use devices instead of ids in HCQGraph
* fiz
2025-07-21 20:03:12 +03:00
wozeparrot and GitHub
30ce16a424
feat: failing test for long keccak ( #11292 )
2025-07-21 12:49:23 -04:00
uuuvn and GitHub
178dbf3f66
Remote scheduler changes ( #11177 )
2025-07-21 09:29:44 -07:00
वेदांत and GitHub
e368628736
Add amin support to Tensor operations in Torch backend ( #11290 )
...
* intiger div mod fix
* Revert "intiger div mod fix"
This reverts commit d5d2f201bf .
* feat arg_min support
* tets update
* test fix
2025-07-21 09:14:08 -04:00
qazal and GitHub
5eb54e2499
viz: close event streams before profiler render ( #11300 )
2025-07-21 15:42:31 +03:00
nimlgen and GitHub
cc3c1e4c14
hcq: move cpu to hcq ( #11262 )
...
* hcq: move cpu to hcq
* import time
* upd
* fix
* windows support
* hm
* cleaner
* fix timer
* fix timing
* std is ns
* skip profiler
* mypy
* cleaner
* cleanups
* after merge
* default is back
2025-07-21 15:10:38 +03:00
nimlgen and GitHub
816c01c2d4
hcq: default copy_queue_t=None ( #11297 )
2025-07-21 14:45:20 +03:00
qazal and GitHub
6520a7fcb6
viz: factorize event stream ( #11298 )
2025-07-21 14:42:00 +03:00
nimlgen and GitHub
9c533e5c38
hcq: cpu prereq ( #11296 )
2025-07-21 13:35:18 +03:00
nimlgen and GitHub
e87a42e243
hcq: prepare for windows ( #11293 )
...
* hcq: prepare for windows
* comments
2025-07-21 13:08:56 +03:00
nimlgen and GitHub
df3ba0a7c0
autogen: fix imports in libusb ( #11294 )
2025-07-21 13:04:27 +03:00
nimlgen and GitHub
dd6a2d432f
hcq: default timestamp metrics is ns ( #11295 )
2025-07-21 12:56:30 +03:00
wozeparrot and GitHub
53345ef4e2
feat: make ops_disk work on block devices ( #11291 )
2025-07-20 14:39:50 -07:00
qazal and GitHub
3002c63b1e
process replay: optionally pass tinygrad import error ( #11289 )
...
* process replay: optionally pass tinygrad import error
* gate all tinygrad internals
* s/getenv/os.getenv pre import
* diff
2025-07-20 22:57:56 +03:00
chenyu and GitHub
9e3a593313
minor kernel.py cleanups [pr] ( #11286 )
2025-07-20 10:15:31 -04:00
quortus and GitHub
5f17927a87
Shorten UOp.load method ( #11285 )
2025-07-20 13:48:04 +03:00
chenyu and GitHub
54924f9969
type remove Union and Optional [pr] ( #11283 )
...
use `|` for consistency
2025-07-19 14:05:52 -04:00
nimlgen and GitHub
2f72be5055
nv_smi: init basic insmod/rmmod/reset cmds ( #11282 )
2025-07-19 15:43:03 +03:00
qazal and GitHub
577e581943
fix typo in sqtt/readme ( #11281 )
2025-07-19 15:10:24 +03:00
nimlgen and GitHub
188ed38315
replace from_mv with lightweight mv_address ( #11280 )
2025-07-19 13:50:51 +03:00
1a25e27f32
Do not produce out of spec intermediate UOp in gated LOAD/STORE folding ( #11207 )
...
Co-authored-by: chenyu <[email protected] >
2025-07-18 15:42:55 -04:00
chenyu and GitHub
ec3efd2919
move upcast before reduce ( #11250 )
...
* move upcast before reduce
upcast goes to end of global+local+upcast
* r_196_32_4_24_8
2025-07-18 14:42:15 -04:00
chenyu and GitHub
be2f4336e6
use onnx 1.18.0 in DSP test ( #11279 )
2025-07-18 14:09:23 -04:00
nimlgen and GitHub
9a88bd841c
hcq: refactor into peer_groups ( #11277 )
...
* hcq: refactor into peer_groups
* fix fors
* fixes
* ooops
* mypy
* tiny fixes
2025-07-18 16:34:18 +03:00
nimlgen and GitHub
f432eef708
hcq: rename CPU -> KICK in graph for kickoff signal ( #11278 )
2025-07-18 15:54:35 +03:00
quortus and GitHub
52bbd9900b
[pr] Stable tensor order in _find_all_tensors_for_uops ( #11276 )
...
* Use dict for all_tensors to get stable tensor order in _find_all_tensors_for_uops
* Rerun tests
2025-07-18 13:12:01 +03:00
chenyu and GitHub
c5a5d74642
Revert "image_dot of 2 half inputs returns half ( #11007 )" ( #11274 )
...
This reverts commit fa8e08f922 .
2025-07-17 17:34:18 -04:00
fa8e08f922
image_dot of 2 half inputs returns half ( #11007 )
...
* cast after sum
* comment out skipif
* minor fix
* only test IMAGE
* IMAGE is supported now
* simpler
* simplerr
* only cast if dtype is None
* dont need to change base_imaeg_type
* only cast when dtype is half
* add explicit test
* actually no, workflow seems better
* actually, keep both
* move test
* fix indent
---------
Co-authored-by: Utkarsh Gill <[email protected] >
2025-07-17 13:47:22 -07:00
geohotstan and GitHub
536b254df4
Bump onnx to 1.18.0 ( #11266 )
...
* bump
* thou hast implement functions
* hacked in domain support
* some clean ups
* hack quantize_onnx_test too
* add helper lol, why onnx tests why
* better dispatcher, but need tests and better naming
* flaky ci
* change some names
* small clean ups
* make it easier to clean up tests once ORT supports 1.18.0
* nits
* fix bug of Softmax_1 being registered in onnx_ops
* need a default value
* resolve_const is better name
* fix OnnxRunner.to
* use proper domain names
2025-07-17 15:35:41 -04:00
qazal and GitHub
1606491b1c
viz: refactor to generic shape spec ( #11272 )
2025-07-17 20:25:15 +03:00
nimlgen and GitHub
cfb229473f
hcq: refactor buffer mapping ( #11271 )
...
* hcq: refactor buffer mapping
* fix
* fix mypy
2025-07-17 15:16:49 +03:00
qazal and GitHub
e68af3b336
disable flaky assert in test_cpu_profile ( #11270 )
2025-07-17 06:50:39 +03:00
chenyu and GitHub
60ffe00172
remove Kernel.first_reduce [pr] ( #11269 )
2025-07-16 18:30:14 -04:00
chenyu and GitHub
522dc72f08
remove Kernel.local_dims [pr] ( #11268 )
...
* remove Kernel.local_dims [pr]
also not needed
* fix test_matvec
2025-07-16 17:46:19 -04:00
chenyu and GitHub
d8c783f65f
remove Kernel.global_dims [pr] ( #11267 )
...
all reference to global used axis_types, so we don't need number of global helper that was used to locate GLOBAL
2025-07-16 17:16:49 -04:00
uuuvn and GitHub
6f0ddcc24c
Remote cross-host graph ( #11229 )
2025-07-16 13:27:54 -07:00
nimlgen and GitHub
6aa20c607d
nv: graceful shutdown to cold state ( #11265 )
2025-07-16 19:49:35 +03:00
chenyu and GitHub
59b52d49d7
remove .global_dims that are for locating GLOBAL [pr] ( #11264 )
2025-07-16 11:19:31 -04:00
chenyu and GitHub
e6c016ddd0
move check axis < shape_len to real_axis [pr] ( #11263 )
...
ensure output of real_axis is always valid
2025-07-16 10:15:44 -04:00
quortus and GitHub
924bc7c9ae
Fix test_uop_spec ( #11259 )
2025-07-16 11:02:31 +03:00
chenyu and GitHub
c8e5c4d7c3
insert_before -> insert_at [pr] ( #11257 )
...
more precise
2025-07-15 17:44:34 -04:00
wozeparrot and GitHub
b32d9321fb
feat: more keccak cleanup + more explicit shape ( #11256 )
2025-07-15 13:57:47 -07:00
chenyu and GitHub
9f79079cbe
update KernelInfo dims to return list of dims [pr] ( #11255 )
...
local dims are not contiguous once upcast sits between local and groupreduce
2025-07-15 15:01:39 -04:00
chenyu and GitHub
629fa21b6b
remove final range in heuristic [pr] ( #11251 )
...
all dims are based on AxisType now
2025-07-15 11:39:15 -04:00
chenyu and GitHub
d7adc24083
remove Kernel.first_upcast [pr] ( #11248 )
...
first_reduce does not need a default now
2025-07-15 10:21:34 -04:00
nimlgen and GitHub
197d345804
nv: print rpc msg with DEBUG>=3 ( #11247 )
2025-07-15 16:39:58 +03:00
chenyu and GitHub
034e51bd36
remove first_reduce used for locate real_axis [pr] ( #11245 )
...
LOCAL goes to the last of (GLOBAL+LOCAL)+1
GROUP goes to right before first REDUCE
2025-07-15 09:19:38 -04:00
chenyu and GitHub
0e2422d216
Kernel.axes_of helper [pr] ( #11243 )
...
look up dim based on AxisType
2025-07-14 22:17:43 -04:00
chenyu and GitHub
968f6b2a2e
remove hasattr(self, 'axis_types') checks in dims property [pr] ( #11242 )
...
no needed anymore
2025-07-14 20:59:51 -04:00
leopf and GitHub
557ca7d757
testing SimpleTokenizer against OASST1 ( #11214 )
2025-07-14 17:09:31 -07:00
wozeparrot and GitHub
5878b189b8
don't const fold shape changing bitcast ( #11236 )
2025-07-14 16:42:16 -07:00
chenyu and GitHub
b6662096cb
remove more first_reduce [pr] ( #11239 )
2025-07-14 19:13:44 -04:00
chenyu and GitHub
eb8e17ef59
remove most of the first_upcast [pr] ( #11238 )
2025-07-14 16:54:24 -04:00
qazal and GitHub
c78b1cbae7
viz profiler cleanups ( #11234 )
...
* move all render calls to zoom callback
* cleanup the naming
* require transform arg
2025-07-14 19:06:33 +03:00
chenyu and GitHub
36ce883c7d
update heuristic to use k.upcastable_dims and k.unrollable_dims [pr] ( #11233 )
...
idea is to make it behave the same regardless of axis order and with empty 1s in shape.
not quite fully remove all first_upcast yet because some conditions used already upcasted size which need a separate benchmark to remove.
2025-07-14 11:10:30 -04:00
qazal and GitHub
c0c695dd89
viz: remove extra transform ( #11232 )
2025-07-14 16:51:47 +03:00
chenyu and GitHub
da219199f5
minor hcopt cleanup [pr] ( #11231 )
2025-07-14 09:36:25 -04:00
nimlgen and GitHub
756ba1a5f9
nv: support ampere in nvpci ( #11230 )
2025-07-14 15:35:44 +03:00
uuuvn and GitHub
b2cc6cfa1b
JIT_BATCH_SIZE is a ContextVar ( #11228 )
2025-07-14 14:03:45 +03:00
nimlgen and GitHub
c4a920d95c
nv: use last signature ( #11227 )
2025-07-14 13:00:39 +03:00
nimlgen and GitHub
a830d37881
nv: check wpr2 is inited ( #11226 )
2025-07-14 11:46:14 +03:00
chenyu and GitHub
0387bb9630
clean up image upcast in hcopt [pr] ( #11220 )
...
GLOBAL+LOCAL for upcast
GROUP_REDUCE+REDUCE for unroll
2025-07-13 18:06:43 -04:00
chenyu and GitHub
85ddd72038
simpler grouptop in hcopt ( #11219 )
...
* simpler grouptop in hcopt
keep the only perf relevant conditions and the rest is handled by try except
* update openpilot read image count
2025-07-13 16:06:09 -04:00
qazal and GitHub
40847ca29c
viz: prune out of screen rects ( #11217 )
2025-07-13 21:49:59 +03:00
chenyu and GitHub
674dc28505
remove Kernel.full_unupcasted_shape [pr] ( #11215 )
...
decomp to shape_len and first_upcast to get the last upcast-able dim
2025-07-13 13:56:23 -04:00
chenyu and GitHub
9575cf6c6e
shave more hcopt [pr] ( #11213 )
...
start to use AxisType for conditions
2025-07-13 12:43:58 -04:00
Alisher Zhubanyshev and GitHub
4ef6b46b34
hcq: reduce launch overhead ( #11193 )
...
* nv: improve mmio creation speed
* add memoryview test
* fix indents
* move mv bench to `test_helpers`, remove comparison
2025-07-13 19:25:50 +03:00
nimlgen and GitHub
1cc2b3f845
nv: use wait_cond ( #11212 )
2025-07-13 19:25:20 +03:00
nimlgen and GitHub
6cce3a5d58
generic wait_cond ( #11210 )
...
* generic wait_cond
* fix linter
* fix linter
2025-07-13 16:59:21 +03:00
chenyu and GitHub
e11ccf2342
update float4 condition in hcopt ( #11211 )
...
don't need all upcast candidates to be upcast-able, only check the actual one
2025-07-13 09:51:45 -04:00
nimlgen and GitHub
55c54d9745
nv: sync after gpfifo setup ( #11209 )
2025-07-13 14:40:11 +03:00
chenyu and GitHub
d90d837013
clean up hcopt [pr] ( #11205 )
...
removed one condition that's always true
2025-07-12 23:10:27 -04:00
chenyu and GitHub
2b48b961be
fix a few broken AMX tests ( #11204 )
2025-07-12 21:42:38 -04:00
wozeparrot and GitHub
667c7a9fa6
clean: keccak cleanups + explicit shapes ( #11202 )
2025-07-12 18:17:14 -07:00
chenyu and GitHub
a0438012af
remove Kernel.get_program [pr] ( #11203 )
2025-07-12 20:50:29 -04:00
George Hotz and GitHub
d67c8e7b42
local metal on metal in uop syntax ( #11185 )
...
* local metal on metal in uop syntax
* TODO: just put the axis_info in the kernelinfo
* local
* amd_matmul works @ 28 TFLOPS
* clean up matmul
* kernel8 works
* remove that
* locals
* axistype innovation
* work
* cleanup
* kernel3 regs
* cleanup kernel3
* work
* why is it broken
* no beam
* reenable
* permutes
2025-07-12 16:31:19 -07:00
uuuvn and GitHub
40da5f0c81
fix silent mypy failure in ci ( #11201 )
...
Example: https://github.com/tinygrad/tinygrad/actions/runs/16215577171/job/45784110543?pr=11177#step:7:20
Caused by footguny exception in how `set -e` works:
```bash
python -m mypy --strict-equality --lineprecision-report . && cat lineprecision.txt
```
Will fail (and have non-zero exit code if run in interactive mode) but
because there is `&&` it won't count as script-terminating failure in a
script with `set -e` and instead as a test (similar to how fail of a
command in if condition won't count as a script-terminating failure
despite having non-zero exit code)
2025-07-12 15:12:25 -04:00
chenyu and GitHub
73caa5dd1b
remove Kernel.membufs [pr] ( #11200 )
2025-07-12 14:48:47 -04:00
5ce278b245
OnnxRunner file as input ( #10789 )
...
* file path as input and have parse be in OnnxRunner.__init__
* modelproto_to_onnxrunner -> modelproto_to_runner
* whoops, fix import
* oh flakiness again, is it because it's getting gc-ed?
* small changes
* CI flaky so just move compile4 fix in
* copy typing of onnx_load
* actually can just import onnx_load instead of onnx.load
* fix external_benchmark_openpilot
* fix onnx_runner test to use onnx_helper
* rerun CI
* try run_modelproto
* spam CI a few times
* revert run_modelproto since that's flaky also
* no external onnx_load usage except onnx.py
* cursor tab complete is evil. Snuck a darn sorted in. But does order change result? Why?
* model_benchmark 193s -> 80s, add OnnxRunner.to()...
* minimize diff and clean up
* device can be None, weird but eh
---------
Co-authored-by: chenyu <[email protected] >
2025-07-12 14:27:46 -04:00
nimlgen and GitHub
110cff3f2e
fix device arg to Tensor.randn ( #11194 )
...
* fix device arg to Tensor.randn
* simpler test
* self.assertEqual
2025-07-12 13:51:59 -04:00
chenyu and GitHub
6283d50224
DEPRECATED_linearize -> to_program [pr] ( #11198 )
2025-07-12 13:46:20 -04:00
George Hotz and GitHub
770a558585
lil cleanups from uop branch [pr] ( #11197 )
2025-07-12 09:46:28 -07:00
George Hotz and GitHub
5625e1904b
axis types in KernelInfo ( #11196 )
...
* axis types in KernelInfo [pr]
* simpler lowerer
* fix tests
2025-07-12 09:36:20 -07:00
nimlgen and GitHub
ea7f2f779c
hcq: p2p nv-amd ( #11195 )
...
* hcq: p2p between diff devices
* fix
2025-07-12 18:53:34 +03:00
qazal and GitHub
6a9f059b21
viz: early convert to cpu time ( #11192 )
2025-07-12 17:19:41 +03:00
chenyu and GitHub
12b04efd69
remove a TODO prod(k.full_shape[k.first_upcast:]) ( #11191 )
...
IMAGE=2 test/test_ops.py works now
2025-07-12 10:16:56 -04:00
nimlgen and GitHub
6f5250d158
nv: fix typing in rpc_rm_control ( #11189 )
2025-07-12 16:09:42 +03:00
qazal and GitHub
c0a5490c72
viz: minor profiler cleanup ( #11190 )
2025-07-12 14:18:24 +03:00
chenyu and GitHub
fdcc25e392
some noop hand_coded_optimizations cleanup [pr] ( #11188 )
2025-07-12 00:09:23 -04:00
chenyu and GitHub
1ad852a892
break up Kernel.reshape_and_permute [pr] ( #11187 )
2025-07-11 18:08:08 -04:00
d11b20129d
DMARef infra ( #10753 )
...
Co-authored-by: wozeparrot <[email protected] >
2025-07-11 14:09:47 -07:00
chenyu and GitHub
b072be0e2d
hotfix whisper main script ( #11184 )
2025-07-11 12:34:00 -04:00
qazal and GitHub
0b7e9b5db7
viz: bugfix for multiple rewrites with the same name ( #11182 )
2025-07-11 18:26:12 +03:00
nimlgen and GitHub
f9e4c4e57a
nv: nvpci blackwell support ( #11127 )
...
* nv: start 5090
* gsp init 5090
* mmu
* works
* after merge
* clenaer
* rwk
* x
* fx
* finish?
* fix
* unrelated
* fix
* commenbt
2025-07-11 17:02:09 +03:00
qazal and GitHub
1d85323572
viz: absolute scaling of memory graph ( #11181 )
2025-07-11 16:39:11 +03:00
nimlgen and GitHub
c7f6b617b4
nv: do not hardcode lv0 pd size ( #11180 )
2025-07-11 16:26:18 +03:00
nimlgen and GitHub
27922c986a
nv: generic mmu impl ( #11179 )
2025-07-11 16:26:09 +03:00
qazal and GitHub
d3ec63a5c3
viz: add base class for unittests ( #11178 )
2025-07-11 13:58:03 +03:00
qazal and GitHub
b791ea117d
viz: enable scrolling in profiler ( #11169 )
...
* viz: add scrollbar to profiler
* using margin fixes the layout bug
* s/profiler.clientHeight/profiler.scrollHeight, it's important
* closer
* scrolling on the device list also works
2025-07-11 11:30:13 +03:00
chenyu and GitHub
b219e47bef
remove Kernel.upcasted_axis [pr] ( #11175 )
2025-07-10 23:19:21 -04:00
George Hotz and GitHub
ccd382bc6f
use axis_types more [pr] ( #11172 )
...
* use axis_types more
* fix local shape
* simpler clause
* fix local shape
2025-07-10 15:05:13 -07:00
nimlgen and GitHub
fb278c6a02
do not recreate Compiled.profile_events in helper_collect_profile ( #11171 )
2025-07-10 23:55:12 +03:00
George Hotz and GitHub
5c5eb92ed4
tc unroll after upcast [pr] ( #11170 )
2025-07-10 13:43:50 -07:00
George Hotz and GitHub
05613c8cac
use shape str for tensor cores upcast/reduce [pr] ( #11168 )
...
* use shape str for tensor cores upcast/reduce [pr]
* reduce axis count isn't fixed
2025-07-10 13:10:58 -07:00
nimlgen and GitHub
cc6ed30f4f
nv: relative lv addressing in NVPageTableEntry ( #11164 )
2025-07-10 22:35:50 +03:00
chenyu and GitHub
439d033af9
update the README matmul example ( #11167 )
...
don't call rand and numpy to show that it's indeed one kernel
2025-07-10 14:47:29 -04:00
qazal and GitHub
bde80c0cdf
record GraphEvents in metal graph ( #11145 )
...
* record GraphEvents in metal graph
* add TestProfiler.test_graph, revert old stuff
* move profile capture to MetalGraph
* comment
* don't double record graph command buffers
* wait_check
* explicit delete
2025-07-10 21:32:06 +03:00
George Hotz and GitHub
8ce3d5906b
use shape_str for tensor cores ( #11165 )
2025-07-10 09:10:36 -07:00
nimlgen and GitHub
581397110f
nv: use classes in GSP_IP ( #11163 )
2025-07-10 17:47:12 +03:00
nimlgen and GitHub
705de6b8a6
nv: parse sizes of ctx buffers ( #11161 )
2025-07-10 17:46:48 +03:00
qazal and GitHub
dcc9704b6b
viz: profile RewriteSteps in TINY device ( #11125 )
...
* viz: profile RewriteSteps in TINY device
* use TracingKey with category
* split by whitespace
* add tracing.py
* work
* tracing_key
* TRACK_MATCH_STATS=3, can this be in defaults?
* fallback name
* work
* javascript
* measure text is slow
* checkout
* profile graph_rewrite/graph_rewrite_map
* change that
* no as
* finally
* work
* linking works
2025-07-10 17:45:57 +03:00
Pyry Kovanen and GitHub
32117402dd
metal: fix incorrect _free on interpreter exit ( #11158 )
2025-07-10 14:01:30 +03:00
qazal and GitHub
3d610f6d2b
viz: small ui cleanup ( #11157 )
...
* viz: small ui cleanup
* 2
2025-07-10 11:43:36 +03:00
chenyu and GitHub
7db07e5f2c
don't narrow range of CAST on bool/unsigned ( #11156 )
2025-07-09 22:20:09 -04:00
George Hotz and GitHub
e154a66f43
unroll axis 0 in tensor core ( #11155 )
...
* unroll is 0 in tc [pr]
* flip order of upcast/reduce in tensor core
* Revert "flip order of upcast/reduce in tensor core"
This reverts commit e564e38bcd .
2025-07-09 17:28:23 -07:00
George Hotz and GitHub
b7742ad9e4
migrate to string swizzle [pr] ( #11154 )
2025-07-09 16:57:53 -07:00
George Hotz and GitHub
4156baee93
break swizzle into three chunks [pr] ( #11153 )
...
* break swizzle into three chunks [pr]
* test failed
2025-07-09 15:30:34 -07:00
George Hotz and GitHub
ca2dc95433
swizzle in tc can't be none [pr] ( #11152 )
2025-07-09 14:44:23 -07:00
George Hotz and GitHub
53ae153404
tc should be in opt ( #11148 )
...
* tc should be in opt [pr]
* fix import
2025-07-09 14:12:21 -07:00
wozeparrot and GitHub
6697d0089d
initial gfx950 kfd support ( #11151 )
...
* feat: initial gfx950 support
* fix: lint
2025-07-09 13:45:16 -07:00
George Hotz and GitHub
262054be52
gfx950 tc support ( #11150 )
2025-07-09 13:30:42 -07:00
nimlgen and GitHub
b6981404ed
memory: use page shifts in memory manager ( #11149 )
...
* memory: use page shifts in memory manager
* fix
2025-07-09 22:05:00 +03:00
qazal and GitHub
5c1d215b41
viz: add Graph stream ( #11144 )
...
* viz: stack an event for the entire batch
* multi
* whitespace
* work
* multi graph, Graph gets its own row
2025-07-09 20:56:46 +03:00
George Hotz and GitHub
22305260e0
move tc to tc.py [pr] ( #11147 )
2025-07-09 10:55:56 -07:00
George Hotz and GitHub
2893feb9f6
cleanups for kernel.py ( #11143 )
...
* cleanups for kernel.py
* fixups
2025-07-08 18:10:25 -07:00
George Hotz and GitHub
b11ca104e9
axis cleanups [pr] ( #11142 )
2025-07-08 17:07:26 -07:00
chenyu and GitHub
7ce9e45474
mypy onnx_parser ( #11141 )
2025-07-08 19:50:28 -04:00
George Hotz and GitHub
a1b8f3e64f
delete info from kernel [pr] ( #11139 )
...
* delete info from kernel [pr]
* update kernel info
* delete info
2025-07-08 15:53:13 -07:00
George Hotz and GitHub
359bed74f8
axis type tracking [pr] ( #11137 )
...
* axis type tracking [pr]
* keep update_info
* keep legacy colors
* update tests to apply_opt
2025-07-08 14:16:25 -07:00
chenyu and GitHub
dada3f5bf3
skip some new onnx tests ( #11135 )
...
these fails on master with latest onnx
2025-07-08 16:12:48 -04:00
chenyu and GitHub
ffcc557986
lint onnx and onnx_parser ( #11134 )
2025-07-08 15:28:35 -04:00
George Hotz and GitHub
3238d21cd1
add finalized to kernel [pr] ( #11132 )
...
* add finalized to kernel [pr]
* add copy
2025-07-08 11:06:17 -07:00
geohot
289a411f5f
hotfix: remove unused GBARRIER, CONTIGUOUS color is GBARRIER
2025-07-08 10:31:06 -07:00
nimlgen and GitHub
43650169f4
nv: switch headers to 570.144 to match gsp ( #11131 )
2025-07-08 20:29:06 +03:00
quortus and GitHub
790b05ab12
[pr] Unify CONTIGUOUS and GBARRIER ( #11121 )
...
* Unify CONTIGUOUS and GBARRIER
* Simplify rules
2025-07-08 10:27:23 -07:00
nimlgen and GitHub
b516fe71b4
nv: return real struct in _alloc_boot_struct ( #11130 )
2025-07-08 20:04:43 +03:00
qazal and GitHub
3dfc0ff887
move cpu_profile and shared ProfileEvents from device.py to helpers [pr] ( #11126 )
...
* move cpu_profile and shared ProfileEvents to helpers [pr]
* TestProfiler.test_cpu_profile
* update test_viz.py
* TestProfiler.test_profile_multiops ordering, it's different streams now
2025-07-08 12:14:03 +03:00
George Hotz and GitHub
397826f0b4
add a test for 1B llm ( #11124 )
...
* add a test for 1B llm
* fix mbs
* add apps to release
2025-07-07 18:47:25 -07:00
George Hotz and GitHub
f7d4638e05
start LLM app, tons of clean up required. target is 200 line ollama ( #11068 )
...
* start LLM app, tons of clean up required. target is 200 line ollama
* kind of works
* simpler
* add k/v cache
* with SYM=1, it loops
* no rope cache
* simpler
* more cleanups
* cleanups
* works
* argparse and comments
* from gguf
* generate is a function
* no copy from cpu
* fix max context pass in
* test
* improve test
* ai2_arc
* fix 8B, use less ram
* 136 lines
2025-07-07 17:09:46 -07:00
chenyu and GitHub
341a686799
Tensor.diagonal ( #11122 )
...
only implemented main diagonal for 2-D tensors. with diagonal and qr, we can get determinant
2025-07-07 16:21:26 -04:00
584fd6af5a
Fix division by zero and mask bug in add views ( #11088 )
...
* merge view infinite loop test
* adjust condition in `x//d -> x//(-d)*-1`
* Fix division by zero in add views
* adjust offset end
* fix typo in comment
* add target to test_merge_views_variable
* fix view incorrectly being masked
* ssimplify strides and offset of the new view to canonicalize
* remove print in test
---------
Co-authored-by: qazal <[email protected] >
2025-07-07 10:05:47 -07:00
nimlgen and GitHub
71377cd233
nv: parse falcon app descs ( #11118 )
2025-07-07 18:14:14 +03:00
nimlgen and GitHub
9a573a1d99
nv: finalize nvdev ( #11117 )
...
* nv: finalize nvdev
* typo
2025-07-07 16:31:59 +03:00
nimlgen and GitHub
fa59c05282
nv: import flags from system ( #11115 )
...
* nv: import flags from system
* not used
2025-07-07 14:46:49 +03:00
a1a146a499
adding enable_gqa in SDPA ( #11097 )
...
Co-authored-by: wozeparrot <[email protected] >
2025-07-06 23:25:33 -07:00
nimlgen and GitHub
b73e89110e
nv: align allocations for perf ( #11114 )
2025-07-06 22:32:11 +03:00
chenyu and GitHub
7468959f4b
Tensor.argsort ( #11112 )
2025-07-06 13:56:35 -04:00
kevvz and GitHub
b7af9cf849
clean svd tests, set full_matrices false in torch backend ( #11113 )
...
* clean tests, set full_matrices false
* add more shape asserts
2025-07-06 13:55:49 -04:00
qazal and GitHub
a556f50668
viz: small ui fixes ( #11110 )
...
* share styling of ctx-list and metadata
* scrollbar-gutter: stable prevents layout shift when changing steps
* margin-left makes left side unaligned
2025-07-06 17:05:36 +03:00
chenyu and GitHub
ba88ec3ad0
pipe linalg svd to torch ( #11109 )
...
and found a bug in svd
2025-07-06 08:37:25 -04:00
chenyu and GitHub
845a4d32bc
Tensor.diag ( #11108 )
...
also updated Tensor.eye to use it
2025-07-05 23:03:02 -04:00
ttomsa and GitHub
4905af4ae0
remove invalid int div test ( #11106 )
...
* rm test
* also rm this
2025-07-05 18:57:55 -04:00
qazal and GitHub
a4aa769c0a
fix: type checking for track_rewrites key [pr] ( #11104 )
...
* fix: type checking for track_rewrites key [pr]
* also for cpu_profile
* func.__name__ to start
2025-07-05 20:11:21 +03:00
qazal and GitHub
81781dc12b
viz: renames and spacing changes to tracing ( #11102 )
2025-07-05 18:40:39 +03:00
qazal and GitHub
7619bf35e7
cleanup: remove disabled TestIndexingOrdering ( #11101 )
...
* cleanup: remove disabled TestIndexingOrdering
* don't import kernelize internals
2025-07-05 18:14:37 +03:00
qazal and GitHub
4fcfaa0ef7
viz: switch to TracingKey ( #11100 )
...
* viz: switch to TracingKey
* tuple
* order is name, keys, fmt
* add test_tracing_key
2025-07-05 17:46:18 +03:00
qazal and GitHub
458be950d9
viz: add TINY device ( #11095 )
...
* viz: add TINY device
* replace Any with a proper type
* reorder
* diff
* rename
* space
* from diff
* multiple keys
2025-07-05 16:54:55 +03:00
nimlgen and GitHub
4dccb2ea49
am_smi: increase kill retries ( #11099 )
2025-07-05 16:23:50 +03:00
chenyu and GitHub
39b4d72687
remove flatten and reshape in sparse_categorical_crossentropy [pr] ( #11093 )
...
not needed, directly operating on the classes dim is fine
2025-07-04 15:15:27 -04:00
nimlgen and GitHub
577afc9f05
hcq: remove redunt syncs and fix typing ( #11096 )
...
Before this patch the code could issues reduntdant syncs because of
the typing issue. Current tests should cover all correctness checks.
2025-07-04 21:49:47 +03:00
qazal and GitHub
41aa54eb5a
viz: resolve all graph references in python ( #11087 )
...
* viz: resolve all graph references in python
* it just maps things to the index
* always map the name
* key on the uop
* diff
* close
2025-07-04 20:35:25 +03:00
qazal and GitHub
3d8569f6d8
hotfix: infinite loop in tracking pattern matcher ( #11094 )
...
* failing test
* fix that
* given matchers
2025-07-04 19:55:26 +03:00
qazal and GitHub
a783211fc7
viz: allow end_time=None in trace events ( #11092 )
2025-07-04 17:48:17 +03:00
0xSG and GitHub
17119b0f23
hip_ioctl: platform.machine added ( #11084 )
2025-07-04 17:20:24 +03:00
nimlgen and GitHub
6656aa162c
nv: enable huge pages ( #11091 )
2025-07-04 17:17:24 +03:00
nimlgen and GitHub
01f3c4f44d
memory: simpler paddr allocation logic ( #11090 )
...
* memory: new paddr allocation logic
* am fix
* am refactrros
* fix
* mypy
* use it
* am
2025-07-04 17:00:36 +03:00
qazal and GitHub
f6d55d9272
viz: pickle UPat location ( #11086 )
2025-07-04 13:09:00 +03:00
qazal and GitHub
2403f126ed
move printable out of UPat [pr] ( #11085 )
...
* move printable out of UPat [pr]
* print_match_stats
2025-07-04 12:31:11 +03:00
qazal and GitHub
988540f401
support capturing cpu_profile on error ( #11078 )
...
* support capturing cpu_profile on error
* spacing
* pylint complains
2025-07-04 11:53:12 +03:00
chenyu and GitHub
a2f5a54458
move sparse_categorical_crossentropy to test_ops ( #11083 )
...
also flattened the tests
2025-07-03 21:40:54 -04:00
chenyu and GitHub
7c8ccb0267
sparse_categorical_crossentropy cleanup [pr] ( #11082 )
2025-07-03 18:32:52 -04:00
nimlgen and GitHub
e02ee8ef1b
nv: cleanups from 5090 ( #11081 )
2025-07-04 00:08:47 +03:00
George Hotz and GitHub
e9a01dd04a
Revert "Fix division by zero in add views ( #11075 )" ( #11080 )
...
This reverts commit 19f07e72f6 .
2025-07-03 11:39:44 -07:00
Sieds Lykles and GitHub
19f07e72f6
Fix division by zero in add views ( #11075 )
2025-07-03 11:37:59 -07:00
chenyu and GitHub
678cabc6f2
use argfix in Tensor.stack ( #11077 )
...
works for multiple Tensor args or single tuple/list of Tensors, but not the mixed
2025-07-03 12:15:11 -04:00
qazal and GitHub
b695e8c4d6
viz: remove support for naming with self ( #11076 )
2025-07-03 17:29:14 +03:00
Sieds Lykles and GitHub
53985297bd
add test, fix rewrite rule and raise error on division by zero ( #11073 )
2025-07-03 08:25:06 -04:00
nimlgen and GitHub
2d138c6cf1
am: factor out init_sw ( #11070 )
2025-07-03 11:01:17 +03:00
quortus and GitHub
a937ac80dc
Replace ASSIGN with STORE in UPat compiler ( #11065 )
2025-07-02 19:15:43 -07:00
George Hotz and GitHub
d049639221
little setitem test ( #11064 )
...
* setitem has one less realize, why broken
* put realize back
2025-07-02 15:10:24 -07:00
quortus and GitHub
17d85b9793
Refactor STORE implementation in ops_python ( #11060 )
2025-07-02 14:29:12 -07:00
George Hotz and GitHub
3b85534df0
outerworld range test [pr] ( #11059 )
...
* outerworld range test [pr]
* bound range
* grad acc test
* more tests
* 5 steps is fine
2025-07-02 14:28:44 -07:00
chenyu and GitHub
425d5f55c4
generate kernel dataset and upload artifact ( #11063 )
2025-07-02 17:21:25 -04:00
chenyu and GitHub
09cc64eea7
remove const 0 clause in "UOp with size 0 is zero" [pr] ( #11061 )
2025-07-02 16:36:40 -04:00
chenyu and GitHub
4d57437a67
add timeout to benchmark_search and mlperf action ( #11058 )
...
default timeout is 6 hours which is too long and occupies a box
2025-07-02 14:17:34 -04:00
nimlgen and GitHub
6067568087
nv: remove hardcoded CTRL_CMD_VASPACE_COPY_SERVER_RESERVED_PDES ( #11057 )
2025-07-02 20:41:10 +03:00
qazal and GitHub
ad155f5454
print inputs to get_program in process replay [pr] ( #11051 )
...
* print inputs to get_program in process replay [pr]
* colors
* keep dataclass default escapes
* Revert "keep dataclass default escapes"
This reverts commit c6db7e8a7a .
* note for ast_repr
* add that back
2025-07-02 20:20:01 +03:00
Ignacio Sica and GitHub
a22aa77c82
cleanup opts_to_apply ( #11055 )
...
* fix kernelinfo init in fixup_ast
* opts_to_apply None
2025-07-02 20:03:19 +03:00
qazal and GitHub
a919b8325b
add test_kernel_info ( #11054 )
...
* add test_kernel_info
* reorder
2025-07-02 19:48:12 +03:00
3b041d188f
[bounty] Singular Value Decomposition ( #10875 )
...
* inital commit
* add qr + expand svd to full matrix
* add odd number support
* add linalg tests
* qr supports dims of arbitrary size
* add qr tests
* svd supports dims of arbitrary size
* small cleanip
* improvements over svd batch handling
* improve linalg tests
* make u_pad match q shape
* add nonfull matrix tests
* little less verbose nonfull svd test
* added dtypes on svd + return vt instead of vt
* lint
* more lint
* lint + set seed
* small fix
* small lint
* lint
* add int casting to indices and shapes
* remove int from shape tuple in svd
* small cleanup
* add return types
* reuse inverse_permute
* refactoring
* whitespace
* remove regularization term to prevent bad outputs on ill conditioned matrices
* remove seed
* refactor
* lint
* refactor
* spacing
* remove clone
* line reduction
* smarter heuristic for iterations_per_round
* add big test
* lint
* turns out no constant needed?
* wrap tests
* some small matrices need the constant
* remove realize
---------
Co-authored-by: George Hotz <[email protected] >
2025-07-02 09:06:03 -07:00
Ignacio Sica and GitHub
fc42c3063e
use kernel info ( #11049 )
...
* use kernel info
* keep api
* revert change in comment
2025-07-02 08:42:32 -07:00
e992ed10dc
WebGPU on Windows ( #10890 )
...
* WebGPU on Windows
* Fix dawn-python install
* New test
* pydeps
* Minor fix
* Only install dawn-python on windows webgpu
---------
Co-authored-by: George Hotz <[email protected] >
2025-07-02 08:38:45 -07:00
nimlgen and GitHub
e67a6d2310
nv: tiny cleanups ( #11053 )
2025-07-02 18:37:32 +03:00
chenyu and GitHub
4626e9c172
is_numpy_ndarray helper [pr] ( #11050 )
2025-07-02 09:12:53 -04:00
qazal and GitHub
452b22c9b6
fix process replay diff in PYTHON device [pr] ( #11052 )
...
* fix process replay diff in PYTHON device [pr]
The PYTHON backend pickles and encodes UOps, the encoded binary can't be
directly diffed in process replay.
* note
2025-07-02 11:06:46 +03:00
8ebf0abaae
ONNX external_test_onnx_backend use PYTHON device for model ( #10915 )
...
* try
* ruff check --fix
* no skip test
* hmmmmmmm I don't get this D:
* run CI again
* why is PYTHON device faster than CPU?
* run ci again and fix lint
* actually doesn't PYTHON device make sense here?
* see cpu speed again
* Revert "see cpu speed again"
This reverts commit 1e366f2256 .
* trigger CI
* pretty good
---------
Co-authored-by: chenyu <[email protected] >
2025-07-01 12:11:17 -04:00
qazal and GitHub
8b0871ac31
viz: test for no lockup on infinite loop ( #11041 )
...
* viz: add test infinite loop fallback
* assert
* continue til the end
* work
* bring that back
* fallback to nop
2025-07-01 17:44:20 +03:00
fcbefde8f5
fix DiskDevice reuse ( #11039 )
...
* fix DiskDevice reuse
* fix mypy and DiskDevice.count
* mypy
* add test
---------
Co-authored-by: b1tg <[email protected] >
2025-07-01 10:29:21 -04:00
geohot
5628e2054c
hotfix: if no ranges, return None
2025-06-30 18:07:56 -07:00
George Hotz and GitHub
0597735f28
remove TC=3 not porting this ( #11045 )
2025-06-30 15:12:49 -07:00
geohot
cccfe6b422
hotfix: test_no_inf_loop_bottom_up
2025-06-30 14:21:45 -07:00
George Hotz and GitHub
752c76ceb7
tc3 shape expand [pr] ( #11043 )
...
* tc3 shape expand [pr]
* remove unused stuff in lowerer
2025-06-30 13:38:14 -07:00
George Hotz and GitHub
539b17fcbf
expand local shape so shapes work [pr] ( #11042 )
2025-06-30 13:03:31 -07:00
nimlgen and GitHub
9ea7deb515
hcq: select_iface shared ( #11033 )
...
* hcq: select_iface shared
* errs
* sorry
* upprt
2025-06-30 21:12:39 +03:00
qazal and GitHub
013085da7d
viz: only path "/" serves the UI ( #11037 )
...
The dict used to exist for /profiler and main localhost:8000, we don't
need it anymore.
2025-06-30 19:10:33 +03:00
George Hotz and GitHub
b829331219
infinite loop detect in fixed_point_rewrite [pr] ( #11038 )
2025-06-30 08:57:29 -07:00
Nino Risteski and GitHub
bc15e98f5c
clean up unused imports in examples and update CI linting ( #11024 )
...
* clean up unused imports in examples
* enable unused import checking in examples
* lint
* ignore F541 and F841 - focus on unused imports only
* clean up
* restore tinygrad.frontend.torch for TINY_BACKEND
* tiny change
2025-06-30 08:21:27 -07:00
George Hotz and GitHub
cb531dba42
detect infinite loop in graph rewrite [pr] ( #11036 )
2025-06-30 08:15:13 -07:00
qazal and GitHub
710d734ce7
viz: don't need PICKLE_BUFFER=0 in capture ( #11031 )
2025-06-30 16:20:04 +03:00
qazal and GitHub
2ea4737930
viz: fix newlines breaking label colors ( #11030 )
...
* viz: fix newlines breaking label colors
* TestViz.test_colored_label
* TestWordWrap
2025-06-30 13:39:44 +03:00
George Hotz and GitHub
5911b71404
early support for bidirectional pattern matcher ( #11027 )
...
* early support for bidirectional pattern matcher
* expose it and add a test
* no bottom up arg there
* disable flaky test
2025-06-29 16:54:07 -07:00
George Hotz and GitHub
ec1d97191d
minor cleanup to lowerer [pr] ( #11026 )
...
* minor cleanup to lowerer [pr]
* add that rule to sym
2025-06-29 11:01:29 -07:00
Piyush and GitHub
454bc3393d
redundant code ( #11014 )
2025-06-29 09:06:10 -07:00
qazal and GitHub
19b11cb778
hotfix: check canvas exists before access ( #11022 )
2025-06-29 14:44:14 +03:00
chenyu and GitHub
126fcf4129
clean up AMD_LLVM in tests ( #11021 )
2025-06-28 22:45:47 -04:00
qazal and GitHub
cb6a66ea84
viz: remove per schedule renderMemoryGraph ( #11019 )
...
replaced with per device Buffer viz https://github.com/tinygrad/tinygrad/pull/10960
2025-06-28 22:09:38 +03:00
qazal and GitHub
4c8d2a0383
buffer viz ( #10960 )
...
* add mem_layout
* ui
* cleanup
* work
* debugLine work and expander
* tooltip style
* real expand device
* wheel does one thing
* diff
* shows llama oom
* add y axis
* mypy chill
* work
* unittests for the memory layout
2025-06-28 21:50:32 +03:00
qazal and GitHub
e3d024afa0
viz: split into scale, shapes, axes last ( #11018 )
...
* viz: split into scale, shapes, axes last
* set zoom on render
2025-06-28 19:10:58 +03:00
qazal and GitHub
508bc68078
viz: small fixups from memory graph ( #11017 )
...
* don't need div.id
* tooltip z-index
2025-06-28 16:34:14 +03:00
qazal and GitHub
fc3e509822
viz: new canvas on first render ( #11016 )
2025-06-28 16:04:51 +03:00
chenyu and GitHub
c14c9a8eff
llama3 grad clip ( #11003 )
2025-06-27 19:14:12 -04:00
nimlgen and GitHub
e53673a0b2
amd: sdma queue overrun fix ( #11012 )
...
* amd: sdma queue overrun fix
* add ()
* fix
* bug
* this is correct
2025-06-28 01:42:03 +03:00
chenyu and GitHub
f2548afeb5
bert grad clipping start with const 0 ( #11008 )
...
saved the init kernels
2025-06-27 18:02:23 -04:00
chenyu and GitHub
a6485d00c8
very tiny generate_dataset ( #11013 )
...
one minute to gen on my mac
2025-06-27 17:10:45 -04:00
qazal and GitHub
382fa6a325
viz: support axis colors in UOp nodes ( #11009 )
...
* work
* javascript
* optional defaultColor
* fine
2025-06-27 23:02:55 +03:00
qazal and GitHub
44257f25e4
bump line count to 14600 ( #11010 )
2025-06-27 22:48:14 +03:00
George Hotz and GitHub
be53ef4f0a
rename DEFINE_ACC -> DEFINE_REG ( #11006 )
...
* rename DEFINE_ACC -> DEFINE_REG
* add CMPEQ to groupops
2025-06-27 11:09:25 -07:00
George Hotz and GitHub
05c35d0db8
reorder ops and add comments ( #11005 )
2025-06-27 10:52:14 -07:00
George Hotz and GitHub
5a1911b7c4
apply the global dims late ( #11002 )
...
* apply the global dims late [pr]
* late gpudims
* tests passing
* remove the random local_dims inc
* simpler
2025-06-27 09:54:34 -07:00
qazal and GitHub
4ef10c57f9
remove unused test helper ( #10999 )
2025-06-27 13:48:48 +03:00
qazal and GitHub
a39343e39f
viz: move timeline layout to python ( #10998 )
...
* viz: move timeline layout to python
* DevEvent has a device and a name
2025-06-27 13:06:00 +03:00
George Hotz and GitHub
b4eb876d5a
kernel.py no longer permutes reduce axis [pr] ( #10968 )
...
* kernel.py no longer permutes reduce axis [pr]
* delete tests that handcode uops
* regen of sops is broken...
* put import back
* just remove that
* disable those tests
2025-06-26 17:44:58 -07:00
chenyu and GitHub
6ab5a5cb6c
llama3 mlperf train ( #10983 )
...
work in progress. now it can overfit small examples and vram roughly matches
2025-06-26 20:24:27 -04:00
George Hotz and GitHub
856759c79c
add halide example ( #10980 )
...
* add halide example
* upd halide gemm
* partial works
* touchups
2025-06-26 16:14:57 -07:00
qazal and GitHub
1127302c46
move perfetto to extra ( #10994 )
...
* move perfetto to extra
* update TestViz and fix tests
* remove perfetto.html from viz directory
* work
* mypy
2025-06-27 01:53:54 +03:00
qazal and GitHub
712980e167
fix extract_dataset + add tests to CI ( #10995 )
...
* fix extract_dataset + tests
* add CI
* sops.gz itself is same as master
* yml + gzip -c + ge
* don't commit that
* bump limit to 1000
* axis=7
* test_tiny
2025-06-27 01:51:36 +03:00
chenyu and GitHub
4572e65f0f
remove duplicated move_early logic in UOp.r [pr] ( #10993 )
2025-06-26 18:33:54 -04:00
Ignacio Sica and GitHub
579194f523
remove some linearize calls from tests 2 [pr] ( #10992 )
...
* refactor count_float4 to take uops as input instead of kernel
* remove some calls to linearize in test_linearizer
* remove some more calls
* remove one more call
2025-06-26 18:22:27 -03:00
50936b4a18
ONNX real float16 ( #10694 )
...
* squash commits
* temp fix for const tensor
* actually realizing float16 can only happen in raw_data
* .float -> cast(float) to rerun CI
---------
Co-authored-by: chenyu <[email protected] >
2025-06-26 14:05:12 -04:00
qazal and GitHub
73484b0803
viz: generic shape tooltip/click handlers + renames ( #10990 )
...
* viz: generic tooltip
* assign kernel
* labelParts/label
* rect with a fillColor
* line
2025-06-26 19:14:04 +03:00
qazal and GitHub
7f79c1388f
viz: update y offset calculation ( #10987 )
...
* viz: update y offset calculation
* don't rescale padding
2025-06-26 12:05:20 +03:00
chenyu and GitHub
49bba2f0a0
improve test_nll_loss ( #10986 )
...
build target and weight tensors outside so it tests backward too.
2025-06-26 02:46:55 -04:00
chenyu and GitHub
0612acfc70
improve Tensor.cross_entropy ( #10985 )
...
separate when Y is prob vs indices and check shapes for indices. also fix higher dim cases
2025-06-26 01:39:48 -04:00
chenyu and GitHub
8751d47985
CosineAnnealingLRWithWarmup ( #10981 )
2025-06-25 17:45:21 -04:00
Ignacio Sica and GitHub
21f1c4cc09
remove some linearize calls from tests [pr] ( #10978 )
...
* remove some linearize calls from tests
speed_compare_cuda_ptx
test_uop_spec
test_linearizer
test_uops
test_winograd
* more clear assert message
2025-06-25 12:37:17 -07:00
chenyu and GitHub
efad567ebd
ruff check whole examples/mlperf/ ( #10979 )
2025-06-25 12:57:48 -04:00
Sieds Lykles and GitHub
15e60caf09
add Ops.EQ ( #10976 )
2025-06-25 11:25:10 -04:00
Ignacio Sica and GitHub
98d2cde293
revert tc_group feature ( #10971 )
2025-06-24 20:58:13 -07:00
George Hotz and GitHub
306dbc76f6
early view simplify ( #10974 )
...
* shape const if it has a device [pr]
* early view simplify
2025-06-24 20:52:45 -07:00
77fff73295
fix viz vscode link on windows ( #10972 )
...
Co-authored-by: b1tg <[email protected] >
2025-06-25 06:47:59 +03:00
George Hotz and GitHub
9d995c2a4d
shape const if it has a device [pr] ( #10969 )
2025-06-24 16:22:54 -07:00
George Hotz and GitHub
cf60ccac6a
support new const lowering ( #10967 )
...
* support new const lowering
* delete invalid linearizer failure tests
2025-06-24 15:21:41 -07:00
geohot
8a65720528
hotfix: disable test_tensor_core_opts_group test on real metal
2025-06-24 15:21:33 -07:00
nimlgen and GitHub
1c45b9f7fb
start nvpci ( #10521 )
...
* start nvpci
* talk to fsp
* boot args
* riscv core bootted
* q
* agen
* got gsp init msg
* some fixes
* set registry, stuck aft lockdown(
* start ga/ad port
* gsp init on ada
* more classes allocated
* more
* mm
* fixes and progress
* no huge pages for now
* mm seems workin, but switch to 512mb page for simplicity
* working state
* not cleaned
* claned
* nvd=1
* start gr ctx
* compute
* clean 1
* cleanup 2
* cleanup 3
* cleaner 4
* cleaner 6
* add iface to nv
* save before reboot
* merged into NV
* moveout mm
* post merge
* cleaner 7
* merge and rebase
* pciiface abstraction + reset
* download fw from web
* print logs
* minor changes + p2p
* cleaner 8
* cleaner 9
* cleaner 10
* delete
* delete this as well
* linter 1
* oops
* priv_client -> priv_root
* fix mypy
* mypy?
* mypy?
* small changes
* shorter
* ops
* remove this
* do not allocate paddr for reserve
* nodiff
* unified script
* ops
* dif ver
* add lock
* setup
2025-06-25 00:37:34 +03:00
uuuvn and GitHub
c8d0f68763
Weaker renderer validation in remote ( #10964 )
...
```
training bert
training on ['REMOTE:0', 'REMOTE:1', 'REMOTE:2', 'REMOTE:3', 'REMOTE:4', 'REMOTE:5']
Traceback (most recent call last):
File "/home/uuuvn/src/tinygrad/examples/mlperf/model_train.py", line 1300, in <module>
with Profiling(enabled=getenv("PYPROFILE")): globals()[nm]()
^^^^^^^^^^^^^^^
File "/home/uuuvn/src/tinygrad/examples/mlperf/model_train.py", line 975, in train_bert
for x in GPUS: Device[x]
~~~~~~^^^
File "/home/uuuvn/src/tinygrad/tinygrad/device.py", line 22, in __getitem__
def __getitem__(self, ix:str) -> Compiled: return self.__get_canonicalized_item(self.canonicalize(ix))
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/home/uuuvn/src/tinygrad/tinygrad/device.py", line 28, in __get_canonicalized_item
ret = [cls for cname, cls in inspect.getmembers(importlib.import_module(f'{base}.runtime.ops_{x}')) \
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/home/uuuvn/src/tinygrad/tinygrad/runtime/ops_remote.py", line 417, in __init__
if not renderer[0].startswith("tinygrad.renderer.") or not renderer[1].endswith("Renderer"): raise RuntimeError(f"bad renderer {renderer}")
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
RuntimeError: bad renderer ('tinygrad.runtime.ops_null', 'NullRenderer', ())
```
2025-06-24 14:15:09 -07:00
George Hotz and GitHub
c2f5f0f198
more robust reduce_gradient ( #10965 )
2025-06-24 14:09:33 -07:00
George Hotz and GitHub
8743ca40e2
force reduce to be in axis order ( #10837 )
...
* force reduce to be in axis order
* disable rule causing loop
* disable that rule
* no ra there
* only move non reduce
* fix tests
2025-06-24 13:00:16 -07:00
chenyu and GitHub
ffb032e31d
test_diagonal touchup ( #10962 )
2025-06-24 15:51:19 -04:00
7f9958b632
Fix torch.linalg.diagonal crash due to invalid shrink in to_movement_ops ( #10945 )
...
* fix as_strided shrink bug breaking torch.linalg.diagonal on tinygrad backend
* cleanup
* generic fix
* tests
* cmp with diagonal too
* oops
* move tests
* fix test
* remove unnecessary import
* fix assert
* compare against numpy
---------
Co-authored-by: Utkarsh Gill <[email protected] >
2025-06-24 15:36:06 -04:00
nimlgen and GitHub
26ddf8d714
amd: rename dev_iface -> iface to match nv ( #10959 )
2025-06-24 20:22:19 +03:00
chenyu and GitHub
bfa87f3490
clean up binary_crossentropy_logits ( #10958 )
2025-06-24 12:23:40 -04:00
qazal and GitHub
2ccddfc0ca
viz: match canvas fontsize ( #10957 )
...
it's 10px https://developer.mozilla.org/en-US/docs/Web/API/CanvasRenderingContext2D/font?utm_source=chatgpt.com .
2025-06-24 19:07:06 +03:00
qazal and GitHub
de4b9bf53b
add opts_to_apply option to AST KernelInfo ( #10950 )
...
* proposal: add option to override opts in the get_program API
* update test_linearizer_rewrite
* state in uops
* update process_replay and names
* empty isn't none
* fix process replay
2025-06-24 18:55:39 +03:00
chenyu and GitHub
18e264a449
Tensor.logsigmoid ( #10955 )
2025-06-24 11:16:14 -04:00
Ignacio Sica and GitHub
f15247d2d2
remove outdated index masking in lowerer [pr] ( #10953 )
...
* add assert to check idx is never replaced with const 0
* remove outdated index masking
2025-06-24 07:53:30 -07:00
cc32394b32
support copyin/copyout/is_allocated for subbuffers ( #10869 )
...
* support copyin/copyout/is_allocated for subbuffers
* simple
* clean up
* rm underlying_buf
* add function is_initialized
* add tests
* better test_subbuffer_copy_in_out
* fix allocator
---------
Co-authored-by: b1tg <[email protected] >
Co-authored-by: George Hotz <[email protected] >
2025-06-24 07:49:04 -07:00
chenyu and GitHub
35504c938e
torch.clip(x,y) -> x.clip(y) in test_ops ( #10954 )
...
* torch.clip(x,y) -> x.clip(y) in test_ops
* test_binary_crossentropy_logits_pos_weights
2025-06-24 10:22:19 -04:00
Fang-Pen Lin and GitHub
86d458533f
Add pos_weight for binary_crossentropy_logits ( #10855 )
...
* Add pos_weight for binary_crossentropy_logits
* Remove debug code
* Code style
* Code style
* Rename
2025-06-24 09:42:37 -04:00
Sieds Lykles and GitHub
61dad3740f
fix min_max and add test ( #10952 )
2025-06-24 09:33:26 -04:00
qazal and GitHub
ab8c5d04ab
viz: convert to function_name in server [pr] ( #10951 )
...
* viz: convert to function_name in server [pr]
* it exists
2025-06-24 13:59:37 +03:00
nimlgen and GitHub
c0d9cf09e0
system: flock ( #10949 )
...
* system: flock
* imports
* xx
2025-06-24 11:33:49 +03:00
nimlgen and GitHub
5202970feb
system: move memory_barrier to System ( #10948 )
...
* system: move memory_barrier to System
* fixed
2025-06-24 11:09:43 +03:00
qazal and GitHub
f41c28a048
update test_tensor_uop_representation comments [pr] ( #10946 )
...
These comments can update to match new tinygrad.
2025-06-24 10:47:09 +03:00
qazal and GitHub
7a5e4e0bf1
fix unittests process replay [pr] ( #10947 )
2025-06-24 10:30:23 +03:00
geohot
7d560dbd75
hotfix: corealize in the tiny mnist test
2025-06-23 17:41:16 -07:00
230ad3a460
[bounty] Don't use numpy inside hlb_cifar10 training loop ( #10777 )
...
* Don't use numpy inside hlb_cifar10 training loop
* Lint it
* jit it
* Drop the last half-batch
* Use gather for random_crop and reuse perms
* Wrap train_cifar in FUSE_ARANGE context
* No need to pass FUSE_ARANGE=1 to hlb_cifar10.py
* Add cutmix to jittable augmentations
* Remove .contiguous() from fetch_batches
* Fix indexing boundary
---------
Co-authored-by: Irwin1138 <[email protected] >
2025-06-23 17:24:56 -07:00
George Hotz and GitHub
383010555f
delete linearize and to_program from kernel.py ( #10943 )
2025-06-23 17:04:05 -07:00
George Hotz and GitHub
0f89660ce4
Revert "change clang -march flag to -mcpu on arm ( #10841 )" ( #10942 )
...
This reverts commit 897e42fd1b .
2025-06-23 16:48:28 -07:00
956a8391a5
minor cleanup on test_tensor_core_opts tests ( #10924 )
...
* minor cleanup on test_tensor_core_opts tests
Tests now notify when skipped
Before, they silently skipped if backend didn't had half precision and
accumulation
Also cleaned up atol and rtol setup
* refactor test_tensor_core_opts_group
---------
Co-authored-by: George Hotz <[email protected] >
2025-06-23 16:30:21 -07:00
ttomsa and GitHub
897e42fd1b
change clang -march flag to -mcpu on arm ( #10841 )
...
* change clang -march flag to -mcpu with fp16 disassembly test
* fix
* add capstone to macos dependencies
* just check no cast in test
* rm import
* woops
* lets check
* move check
* llvm init before cpu chcek
* try this
* bump autogen llvm version
* also update libclang?
* revert
* add comment
* skip llvm test and add comment
* linter
2025-06-23 16:28:48 -07:00
Sieds Lykles and GitHub
772cd02ad2
Perform index validation on load/store, not on the index ( #10849 )
...
* move index validation to load/stores
* add name
* add linearizer_failure
* add validate_store with implicit gates
* linearizer_failure_58 is fixed!
* add test_uop_graph test
* rename cond to gate
* test gated load/stores
* use or_casted()
2025-06-23 16:25:05 -07:00
geohot
ae4d2d71b4
bump line count to 14500
2025-06-23 15:32:27 -07:00
Harsh Natuskar and GitHub
79d7cdd9ba
Fix device ( #10929 )
...
* fix: pkg
* better
* added test
* less lines
2025-06-23 15:30:19 -07:00
George Hotz and GitHub
e15754db28
remove (some) kernelize from llama and test schedule speed ( #10939 )
...
* remove kernelize from llama
* 405B
* space
2025-06-23 15:07:31 -07:00
chenyu and GitHub
3699d1d3ba
hotfix llama3 temperature is float ( #10938 )
2025-06-23 15:20:56 -04:00
uuuvn and GitHub
4e2c9e36c7
Remote multihost (p2p transfer) ( #10601 )
2025-06-23 11:47:29 -07:00
chenyu and GitHub
42b1c9625b
skip test TestKiTS19Dataset::test_training_set ( #10936 )
...
flaky
2025-06-23 14:27:24 -04:00
9e9fd44987
refactor test/external/external_llama_eval.py ( #10567 )
...
Co-authored-by: wozeparrot <[email protected] >
2025-06-23 10:43:20 -07:00
chenyu and GitHub
785b4ea8ac
optim flatten().shape[0] is numel ( #10935 )
2025-06-23 13:11:19 -04:00
qazal and GitHub
ac39f27ae6
viz: non blocking UOp tracing ( #10913 )
...
* viz: non blocking UOp tracing
* u.arg
* no if Ops.KENREL
* drop replace
* switch to weakref.WeakKeyDictionary
* back
* remove ram usage skips, viz works here
* cache on reconstruct
2025-06-23 19:59:28 +03:00
Ignacio Sica and GitHub
b8d09a1dae
tc with group/grouptop ( #10903 )
2025-06-23 09:58:41 -07:00
qazal and GitHub
9944c2c02d
viz: show time taken on hover ( #10934 )
2025-06-23 19:00:40 +03:00
geohot
1e99a7f1c9
hotfix: don't viz the indexing rewrites
2025-06-23 08:20:26 -07:00
chenyu and GitHub
f9b59924f1
OPTIM_DTYPE to specify dtype for optim params ( #10925 )
...
one more flag
2025-06-23 10:32:03 -04:00
qazal and GitHub
7820aeca8e
update codegen process replay to use get_program [pr] ( #10921 )
...
* update codegen process replay to get_program [pr]
* precommit
* try str replace
* +to_function_name
* fixup tc
* local2.sh
* fix openpilot NOLOCALS
* new local.sh
* correct merge
* beam cache
* back
* revert beam thing
* adding opts_override and name_override makes output of get_program
reproducible
* min diff
2025-06-23 17:31:41 +03:00
nimlgen and GitHub
eceb7a00d2
nv: rename iface mem functions ( #10931 )
2025-06-23 16:34:51 +03:00
qazal and GitHub
4e864bd304
fix: getenv("NOLOCALS")/NOLOCALS context var ( #10927 )
...
OptOps shouldn't rely on os.environ.
2025-06-23 11:23:59 +03:00
alpharush and GitHub
22f9696522
Fix/hcqfuzz harnesss bug ( #10923 )
...
* update command so extra module is found
* fix empty range in randrange errors
* lint
2025-06-23 11:22:30 +03:00
qazal and GitHub
f037f85532
s/getenv("TC")/USE_TC context var ( #10922 )
2025-06-23 00:39:45 +03:00
qazal and GitHub
9201224e0b
viz: remove Kernel check [pr] ( #10920 )
...
* viz: remove Kernel check [pr]
* TestVizIntegration
* test/unit allows opening of devices
* kernel -> Kernel
2025-06-22 20:47:54 +03:00
nimlgen and GitHub
3ccdb2356b
system: factor out PCIIfaceBase ( #10917 )
...
* system: factor out PCIIfaceBase
* linter
* typing
2025-06-22 20:03:14 +03:00
George Hotz and GitHub
b09c47366f
opt transforms the ast into an optimized ast ( #10900 )
...
* opt transforms the ast into an optimized ast
* fix get_kernel order and to_function_name
* function_name property
* update docs
* copy from kernel.py
* improve docs
* ci didn't trigger?
2025-06-22 09:41:26 -07:00
qazal and GitHub
ffddf165f8
viz: color by kernel names in profiler ( #10919 )
...
* viz: color by kernel names in profiler
* ellipsis stays in bounds
2025-06-22 18:07:52 +03:00
nimlgen and GitHub
36536ef6f0
nv: minor changes from nvpci ( #10918 )
2025-06-22 18:04:39 +03:00
geohotstan and GitHub
4ab7d792cc
ONNX improve dtype fallback ( #10800 )
...
* fix
* add early verbose demo test
* is this how to write tests :s
* is definition drift even a thing? gemini says it is
* clean up
* better
* even better
* try add to CI
* doesn't work quite yet
* much more work to be done
* whoops
* partition the test heh
* skipif
* some nits for better names
* add webgpu test for onnxrunner
* fix reference links
* flush for now
2025-06-21 19:29:45 -04:00
chenyu and GitHub
0480139def
log_perplexity metrics ( #10912 )
2025-06-21 10:44:47 -04:00
nimlgen and GitHub
0e7bd9fd03
factor out generic MemoryManager ( #10910 )
...
* allocator -> memory
* just moveout it
* mm is abstracted
* need entry abstraction
* fix
* mypy
2025-06-21 16:18:33 +03:00
qazal and GitHub
c7ec913210
viz: cleanup unit tests ( #10909 )
...
* cleanup test_viz
* tree view
2025-06-21 12:35:09 +03:00
chenyu and GitHub
1373071f19
simplify logcumsumexp ( #10908 )
...
clarify and remove some flatten and squeeze/unsqueeze
2025-06-20 22:56:42 -04:00
George Hotz and GitHub
fa52bdb50f
applied_opts is in the optimized ast] ( #10906 )
2025-06-20 18:56:23 -07:00
chenyu and GitHub
2d9c61e39e
test more dims in test_logsumexp and test_logcumsumexp ( #10907 )
...
refactoring squeeze and unsqueeze is easy to get wrong
2025-06-20 21:42:18 -04:00
Nino Risteski and GitHub
3771cc0f77
fix test logcumsumexp broken devectorize=0 ( #10880 )
...
* fix test logcumsumexp numerical
* lint
* Use dtypes.min instead of -1e4
2025-06-20 20:54:50 -04:00
George Hotz and GitHub
7636d2cdc5
flip order of get_program args ( #10905 )
2025-06-20 17:23:23 -07:00
George Hotz and GitHub
1ce63f8d04
move functions to view and update docs [pr] ( #10904 )
...
* move functions to view and update docs [pr]
* move quantize
2025-06-20 16:47:58 -07:00
George Hotz and GitHub
b41e0563a3
move stuff to kernelize folder ( #10902 )
...
* move stuff to kernelize folder
* oops, forgot that
2025-06-20 16:10:20 -07:00
George Hotz and GitHub
d399a4587d
move mem estimate to ProgramSpec [pr] ( #10901 )
2025-06-20 15:54:28 -07:00
George Hotz and GitHub
92678e59ee
move kernel to opt ( #10899 )
2025-06-20 15:22:28 -07:00
nimlgen and GitHub
bb0299b9e5
system: shared pci logic ( #10894 )
...
* moveout pci logic
* fixes
* oops
* types
* more type
* one style
* thi is imp
2025-06-21 00:09:49 +03:00
nimlgen and GitHub
c83fdc50d1
nv: driver iface ( #10895 )
...
* nv: driver iface
* fixes
* ops
* not used anymore
* fix mypy
* too long
* fix
* fixed
* mypy
* ugh, it's misc
* rename to NVK
2025-06-20 22:36:08 +03:00
George Hotz and GitHub
fc9f883870
if upat returns self, it's none ( #10898 )
...
* if upat returns self, it's none
* fix pm tests
2025-06-20 12:11:19 -07:00
qazal and GitHub
4f179b9ddb
viz: gate launch behind a ContextVar [pr] ( #10892 )
2025-06-20 17:30:32 +03:00
chenyu and GitHub
3f29c7edda
minor onnx dropout cleanup ( #10891 )
...
we should consider removing numpy random and test it similar to test_randomness, unless how seed works is part of spec?
2025-06-20 10:18:34 -04:00
simone-pietro and GitHub
e94ac6e20c
Cast ptr to int in test_from_mv_to_mv ( #10876 )
...
* Cast ptr to int in test_from_mv_to_mv
* Add type hints for from_mv
2025-06-20 14:52:34 +03:00
qazal and GitHub
000eb30f04
viz: remove prev profiler file ( #10888 )
...
The new profiler is integrated in the main VIZ tab.
Will also delete perfetto.html after matching [final features](https://github.com/tinygrad/tinygrad/pull/10763#issuecomment-2980543715 ) soon.
2025-06-19 23:05:46 +03:00
chenyu and GitHub
62a540066e
remove DEBUG=2 in mi300x bert setup ( #10886 )
...
seems fine now, not sure what the issue was
2025-06-19 13:28:53 -04:00
Nino Risteski and GitHub
5a56710ff4
small fix replacing download_file with fetch ( #10877 )
...
* imported a missing os and replaced download_file with fetch from tg helpers
* use fetch directly
* Remove if not os.path.isfile
2025-06-19 12:12:09 -04:00
chenyu and GitHub
8d721a4ead
add 405B params to llama3.py ( #10884 )
...
tested with `python examples/llama3.py --model /raid/weights/llama31_405b/ --size 405B --shard 8 --benchmark` on tinyamd2
2025-06-19 11:45:37 -04:00
chenyu and GitHub
a3dae51085
lower test_gemm_8192 on red ( #10883 )
2025-06-19 10:01:25 -04:00
simone-pietro and GitHub
36f01411a2
Pass list to block_reorder in test_loads ( #10881 )
2025-06-19 09:49:45 -04:00
chenyu and GitHub
f377cc19cd
use AM for bert ( #10882 )
...
have triained 3 runs and all seem fine
2025-06-19 09:48:54 -04:00
borgwang and GitHub
06ea74bf2c
fix-typos ( #10879 )
2025-06-19 09:13:31 -04:00
qazal and GitHub
ac891b78f8
skip UOp del when python is shutting down [pr] ( #10847 )
2025-06-19 15:31:40 +03:00
simone-pietro and GitHub
58252e3c49
Change type hint for init_c_struct_t and to_struct [pr] ( #10878 )
...
* Change type hint for init_c_struct_t
* Change type hint for to_struct
2025-06-19 13:22:44 +03:00
qazal and GitHub
00d0071b36
simpler viz naming [pr] ( #10874 )
...
* simpler viz naming [pr]
* n2
2025-06-19 12:10:47 +03:00
qazal and GitHub
5839542fc8
viz: one name arg in track_rewrites [pr] ( #10873 )
...
* viz: one name arg in track_rewrites [pr]
* other test
2025-06-19 03:34:56 +03:00
George Hotz and GitHub
18593c9800
one less rewrite on schedule [pr] ( #10872 )
...
* one less rewrite on schedule [pr]
* verify in ebs
2025-06-18 17:06:17 -07:00
uuuvn and GitHub
e7a26211d2
Queue remote transfers on source ( #10871 )
...
https://github.com/tinygrad/tinygrad/pull/10601#issuecomment-2985624147
I personally don't see how that is a good standalone pr, but whatever
2025-06-18 16:08:44 -07:00
uuuvn and GitHub
a9f3632c4f
SessionKey is a dataclass ( #10870 )
2025-06-18 15:07:31 -07:00
wozeparrot and GitHub
bdbf121285
fix: contigous -> contiguous ( #10868 )
2025-06-18 13:09:51 -07:00
qazal and GitHub
344a220b87
s/lb_refcount/uop_refcount [pr] ( #10865 )
2025-06-18 21:48:04 +03:00
simone-pietro and GitHub
f59df04998
Generalize type hint for get_single_element [pr] ( #10866 )
...
* Generalize type hint for get_single_element
* Improve wording in assert
2025-06-18 13:13:04 -04:00
chenyu and GitHub
d71bb6a7b2
remove comma 0.9.4 from benchmark ( #10867 )
2025-06-18 12:43:59 -04:00
chenyu and GitHub
b70c7d3631
bert grad accumulation ( #10863 )
...
* bert grad accumulation
* realize grad
2025-06-18 12:17:07 -04:00
simone-pietro and GitHub
56fe5b60a9
Cast int to str for render_cast ( #10864 )
...
* Add type hint for render_cast
* Revert "Add type hint for render_cast"
This reverts commit 33858eb711 .
* Cast int to str for render_cast
2025-06-18 10:55:27 -04:00
simone-pietro and GitHub
d8cea1a279
Change rate to int in test_tqdm ( #10848 )
2025-06-18 08:40:43 -04:00
simone-pietro and GitHub
0735224ac2
Pass PythonRenderer instance to full_rewrite ( #10859 )
2025-06-18 08:39:27 -04:00
qazal and GitHub
96509daaba
enable copy folding tests [pr] ( #10862 )
2025-06-18 13:05:35 +03:00
qazal and GitHub
84d568d0cc
do not import grouper internals in test_schedule [pr] ( #10861 )
...
* fix import
* fix test_multitoutput_ast
* fix test_recursive_swizzle
* test_alu_after_copy
* remove that test
2025-06-18 12:47:57 +03:00
qazal and GitHub
8b879b0314
merge TestTensorUOpSpec with the other spec unittests [pr] ( #10860 )
...
* merge TestTensorUOpSpec with the other spec unittests [pr]
* rename to test_uop_spec
2025-06-18 12:12:08 +03:00
qazal and GitHub
a5f2bb614a
remove validate_kernel, it is asserting implementation details [pr] ( #10858 )
2025-06-18 11:42:36 +03:00
George Hotz and GitHub
cba6e15937
split grouper and kernelize [pr] ( #10854 )
2025-06-17 17:54:20 -07:00
George Hotz and GitHub
75503955bf
simple schedule test [pr] ( #10853 )
2025-06-17 16:19:27 -07:00
chenyu and GitHub
075a74cf25
add global_batch_size to mlperf bert ( #10852 )
...
global_batch_size = grad_acc_steps * batch_size. no-op change to prep grad acc for bert
2025-06-17 17:54:15 -04:00
uuuvn and GitHub
a51f18f8f9
CI flakiness ( #10851 )
...
https://github.com/tinygrad/tinygrad/actions/runs/15718103629/job/44292845140?pr=10753#step:4:161
2025-06-17 14:46:30 -07:00
qazal and GitHub
e77cd81662
time viz ( #10763 )
...
* work
* basic stuff
* work
* also reset
* moving through time
* cleanup
* proper zoom
* add livereload.js
pip install livereload
livereload tinygrad/viz
* minor
* fixed width, remove viewbox
* bit of flexbox magic
* show pid/tid
* merge loops
* min-height
* redo some layout stuff
* create cell groups
* text is hard
* javascript Math.min causes "Maximum call stack size"
bert repro: VIZ=1 PYTHONPATH=. DEFAULT_FLOAT=HALF BS=66 GPUS=6 BERT_LAYERS=2 FUSE_ARANGE=1 MODEL=bert python3 examples/mlperf/model_train.py
* fix recursion issue
* no viz/server changes
* fix test_viz
* everything is a g
* text is easy
* no it's still hard
* livereload+notes
* height: 100% fixes the device bug
* start canvas work
* base canvas
* take chrome's stuff
* serve chrome's thing
* fetch traces from get_profile
* remove junk
* remove some more
* bring everything back again
* dispatch resize events
* base ticks
* hook d3.zoom
* zoom on the x axis
* bring filter back, makes ctrl+drag possible
* remove junk
* Revert "remove junk"
This reverts commit 4987e7bec1 .
* draws something, the zooms aren't right
* move to canvas
* fix zooming
* Revert "Revert "remove junk""
This reverts commit 5aac2034fb .
* space key resets zoom
* Divide timelines by device on y axis
* Show kernel names when the width allows
* Clicking on kernel opens it in the kernel graph
* remove livereload.js
* reset diff
* base diff:
- fetch traceEvents
- displayGraph
- flexbox layout
- rest of canvas
* rescale in-place is faster, d3's rescaleX creates a copy
* less
* aesthetics
* map names
* first viz is profiler
* this will work when i make canvas once
* initial cleanups
* factor out of loop
* refactor + only show devices with events
* properly align program rects
* cleaner tick lines
* padding
* listen for resize
* simple zoom
* space more
* i always end up making zoom globl
* how is this ever allowed
* clicking works again
* back button goes back to the same zoom level
* work
* more work
* bring that back
* coloring work and simplify
* black
* keep perfetto button for comparison
* better
* ph===X
* simplify history stuff
* temp: handcoded
* test: flamegraph style leveling
* Revert "temp: handcoded"
This reverts commit bdcd538e88 .
* disable flamegraph
* group by pid
* factor y and height out of render
* now flamegraph is easy
* livereload stuff
* remove that
* less
2025-06-17 19:39:34 +03:00
qazal and GitHub
9e2cb7522a
viz: define launch_viz when tracking is enabled ( #10846 )
2025-06-17 19:38:02 +03:00
Bhavya Gada and GitHub
3a474ef5b7
move bitwise_and/bitwise_or/bitwise_xor to MathTrait [pr] ( #10794 )
...
* move bitwise and, or, xor to MathTrait
* refactor
2025-06-17 09:19:43 -07:00
George Hotz and GitHub
531d143780
bring back old sharded rand behavior ( #10842 )
2025-06-16 17:23:47 -07:00
George Hotz and GitHub
a493eb396c
fix view add 0 ( #10840 )
2025-06-16 16:46:12 -07:00
geohot
b5ce227850
Revert "hotfix: remove setrecursionlimit"
...
This reverts commit acfc81642a .
2025-06-16 16:01:42 -07:00
geohot
acfc81642a
hotfix: remove setrecursionlimit
2025-06-16 15:31:34 -07:00
George Hotz and GitHub
00c46e7077
print rules count in match stats ( #10839 )
2025-06-16 14:56:27 -07:00
George Hotz and GitHub
e2907360b7
multi is one PM [pr] ( #10838 )
...
* multi is one PM [pr]
* disable flaky tests
2025-06-16 14:52:47 -07:00
b1fefb76dd
More conditions for (x//c1+a)//c2 -> (x+a*c1)//(c1*c2) ( #10834 )
...
* add rule and test
* typo
---------
Co-authored-by: chenyu <[email protected] >
2025-06-16 16:34:52 -04:00
uuuvn and GitHub
18d936f981
Remote multihost ( #10598 )
2025-06-16 13:18:56 -07:00
George Hotz and GitHub
0629e45332
remove cpu graph ( #10836 )
...
* remove cpu graph, it's different from the others
* remote was blacklisting CPUGraph
* remove cpugraph from dsp
2025-06-16 11:40:58 -07:00
Sieds Lykles and GitHub
deb6af0638
Remove incorrect rule for x%-d -> (x%d)*-1 ( #10832 )
...
* fix rule and add test
* combine tests
2025-06-16 11:37:44 -04:00
Sieds Lykles and GitHub
946243dbb2
Change z3 cdiv to euclidian division ( #10833 )
...
* change z3_cdiv
* shorter
2025-06-16 11:04:51 -04:00
qazal and GitHub
2c6fd5bf81
viz: rename to svgZoom ( #10831 )
2025-06-16 13:05:22 +03:00
qazal and GitHub
e8ec3f544b
viz: helper for keeping state in browser history ( #10830 )
2025-06-16 12:26:10 +03:00
chenyu and GitHub
e5d5ae55f9
smaller inputs for test_sort and test_topk ( #10829 )
2025-06-16 00:21:15 -04:00
nimlgen and GitHub
c0329148c7
am: check va is aligned to page size ( #10815 )
...
* am: check va is aligned to page size
* swap them
* is this faster
2025-06-15 22:51:09 +03:00
Sieds Lykles and GitHub
ac27c46104
fix UPat get_location after mathtraits refactor ( #10814 )
...
* fix UPat get_location
* fold line
2025-06-15 12:47:55 -07:00
George Hotz and GitHub
5dc1bc6070
switch get_kernel -> get_program [pr] ( #10817 )
...
* switch get_kernel -> get_program [pr]
* fix tests
2025-06-15 12:26:50 -07:00
George Hotz and GitHub
a36b09a715
universal device import [pr] ( #10818 )
2025-06-15 12:01:02 -07:00
George Hotz and GitHub
cc5e4e54b8
move type verify to codegen [pr] ( #10816 )
2025-06-15 12:00:52 -07:00
George Hotz and GitHub
27cf836958
split ocelot out for autogen, fix CI ( #10819 )
...
* split ocelot out for autogen, fix CI
* mac ocelot
2025-06-15 11:37:23 -07:00
Ahmed Harmouche and GitHub
c380efc220
Support aarch64 linux on webgpu ( #10802 )
2025-06-14 14:57:18 -04:00
Sieds Lykles and GitHub
37d3ca152e
Adapt >> for division by power of two to all ints ( #10803 )
...
* Change divison by power of two to always use shift
* Change test to test int instead of uint
* simplify condition
* add old rule back with comment
* remove import
* use sresolve instead of simplify
* use keyword in simplify instead of sresolve
* webgpu cast y to uint
* remove comment
* explicitly set dtype in wgsl
* without simplify
* undo simplify kwarg
* change test to test both int32 and uint32
2025-06-14 14:55:51 -04:00
chenyu and GitHub
652db5702b
move test_conv_shapetracker and some test_search util into unit test ( #10812 )
2025-06-14 13:29:32 -04:00
George Hotz and GitHub
754667093f
remove IGNORE stuff ( #10796 )
...
* remove IGNORE stuff, was this even tested? [pr]
* delete IGNORE op
2025-06-14 09:59:45 -07:00
leopf and GitHub
118a09ddcf
xor self folding ( #10806 )
...
* xor folding
* tests + z3 bitwise xor
2025-06-14 10:01:17 -04:00
qazal and GitHub
8e6ac18436
viz: make sidebar list responsive to keyboard smashing ( #10811 )
...
* expanded is a static style
* only draw the list once
* identify with ids
* state isn't used here anymore
* only toggle states
* less
2025-06-14 13:52:27 +03:00
chenyu and GitHub
8c28b5d833
move dtype spec tests into unit test ( #10808 )
...
* move dtype spec tests into unit test
can clean up more after the split
* skip CI test_backward_sum_acc_dtype
2025-06-13 22:21:22 -04:00
chenyu and GitHub
7a6df0a161
remove .relu() call in several conv tests in test_ops ( #10807 )
...
* remove .relu() call in several conv tests in test_ops
testing negative parts double the effectiveness. keep the relu between two convs and the tests that explicitly test relu
* relax tol
2025-06-13 17:10:16 -04:00
nimlgen and GitHub
b6e574fcdf
am: smu 14.0.3 is smu 14.0.2 ( #10714 )
2025-06-13 23:07:56 +03:00
chenyu and GitHub
7d5c769c6b
fix compile4 ( #10797 )
2025-06-12 22:28:56 -04:00
wozeparrot and GitHub
c01b20fd83
amd: more verbose out of memory error ( #10798 )
2025-06-12 19:06:58 -07:00
geohotstan and GitHub
806b68c2b3
Add fallback dtype to ONNX ( #10788 )
...
* start
* still need the float16 workaround in
* tiny nit for correctness
* idk hacks, I need to understand this device stuff better
* no-op?
* remove that assert for true nooooooop
* add fallback_context
2025-06-12 20:39:21 -04:00
George Hotz and GitHub
dcd1928f29
tensor cores for gfx1200 [pr] ( #10795 )
2025-06-12 16:33:29 -07:00
qazal and GitHub
a113c5e3ae
viz: update browser test to properly shutdown [pr] ( #10793 )
...
Using `await page.evaluate` can cause non deterministic `TargetCloseError`
exceptions if it cannot find the elements on the page, Puppeteer
doesn't cleanly stop when `browser.close()` is called.
[Failing CI](https://github.com/tinygrad/tinygrad/actions/runs/15596803685/job/43928961323?pr=10763#step:9:61 )
2025-06-12 17:58:42 +03:00
Dan German and GitHub
24e7aed74b
ramp.py: correct UOp and Ops import path from tinygrad.uop to tinygrad.uop.ops ( #10791 )
2025-06-12 10:07:03 -04:00
qazal and GitHub
c066baea65
viz: enter key only expands steps ( #10792 )
...
It shouldn't be changing any step or context state. Those are handled
explicitly by the arrow keys (or clicking).
2025-06-12 16:00:14 +03:00
qazal and GitHub
822e2dcb20
viz: back button returns to the kernel graph ( #10790 )
...
* create space
* viz: back button returns to the kernel graph
2025-06-12 15:19:48 +03:00
chenyu and GitHub
4242b9874e
remove AMD_LLVM=0 in mlperf and search ci ( #10785 )
...
tinybox updated to llvm 20
2025-06-11 21:10:31 -04:00
wozeparrot and GitHub
eb739bb96a
hotfix: lower threshold ( #10786 )
2025-06-11 19:36:20 -04:00
wozeparrot and GitHub
53edd49a33
feat: bump to llvm20 ( #10784 )
2025-06-11 16:04:18 -07:00
chenyu and GitHub
7d8939908f
AMD_LLVM=0 for resnet cron ( #10780 )
...
similar pf on llvm19 and fine on 20
2025-06-11 16:28:40 -04:00
qazal and GitHub
a6af8db4d3
viz work from the profiler ( #10781 )
...
* inline ansistrip
* refactor to changeStep + explicitly set expandSteps
2025-06-11 23:20:41 +03:00
Sieds Lykles and GitHub
10b61157b9
Support symbolic slice with no start [pr] ( #10775 )
...
* add symbolic slice with no start
* reshape the test
* step must be int
* just add a cast...
* more cast...
2025-06-11 16:00:38 -04:00
chenyu and GitHub
d465ef4acb
AMD_LLVM=0 for sdxl search ( #10779 )
...
hangs with llvm19 but seems fine with llvm20
2025-06-11 14:56:45 -04:00
uuuvn and GitHub
0d45e1a3ec
Explicitly use CUDA_KERNEL_NODE_PARAMS v1 ( #10776 )
2025-06-11 16:27:50 +03:00
George Hotz and GitHub
a38947b4bb
move symbolic and transcendental to uop [pr] ( #10771 )
2025-06-10 20:51:22 -07:00
chenyu and GitHub
81e296d7b8
remove Tensor.test() in retinanet ( #10770 )
...
test was removed
2025-06-10 22:14:57 -04:00
chenyu and GitHub
25304c3dd0
default AMD_LLVM=1 ( #10253 )
2025-06-10 18:19:21 -04:00
George Hotz and GitHub
9d0383634d
bump cache and include full python version [pr] ( #10768 )
...
* bump cache and include full python version [pr]
* stupid windows
* really stupid windows
2025-06-10 15:07:30 -07:00
chenyu and GitHub
612cdf5146
move fuzz_shape_ops to run with other fuzzer ( #10767 )
...
* move fuzz_shape_ops to run with other fuzzer
* don't skip CPU
2025-06-10 17:43:04 -04:00
chenyu and GitHub
5e7ad70aae
don't run linearize().uop tests in get_action_space test ( #10766 )
...
* don't run linearize().uop tests in get_action_space test
this part takes 2 minutes in CI and has nothing to do with action space. also not sure if the "for some reason" comment is still relevant
* -n=auto test/models
2025-06-10 17:23:53 -04:00
52c49dd4f3
fix onnx ci ( #10762 )
...
Co-authored-by: b1tg <[email protected] >
2025-06-10 14:28:40 -04:00
qazal and GitHub
9e1d1ebc52
print tag in UOp [pr] ( #10755 )
2025-06-10 21:16:07 +03:00
chenyu and GitHub
14fa62c61d
move high level tests to unit ( #10760 )
...
either no need a backend, or running on one to check suffice
2025-06-10 12:55:44 -04:00
George Hotz and GitHub
0fbf3f5554
Revert "Revert "Update autogen ci runner to ubuntu 24.04 ( #10736 )" ( #10757 )" ( #10758 )
...
This reverts commit a6dba9b9d9 .
2025-06-10 09:32:27 -07:00
George Hotz and GitHub
a6dba9b9d9
Revert "Update autogen ci runner to ubuntu 24.04 ( #10736 )" ( #10757 )
...
This reverts commit 1d15374c7a .
2025-06-10 09:31:51 -07:00
uuuvn and GitHub
1d15374c7a
Update autogen ci runner to ubuntu 24.04 ( #10736 )
...
For `kfd.AMDKFD_IOC_EXPORT_DMABUF`
2025-06-10 08:33:02 -07:00
Adrian Wijaya and GitHub
78b9c30640
move idiv to MathTraits [pr] ( #10748 )
2025-06-10 08:32:09 -07:00
Sieds Lykles and GitHub
0daa4c6ed0
Add DType.min and DType.max properties ( #10749 )
...
* add properties
* cleaner test
* remove added newline
2025-06-10 08:31:34 -07:00
qazal and GitHub
5d9c274924
keep UOp tags if sources are replaced ( #10754 )
...
* keep UOp tags in unified_rewrite
* add failing test, print tag if defined
* remove the repr change
2025-06-10 08:30:14 -07:00
qazal and GitHub
3de4c9839f
viz: display UOp tags ( #10751 )
...
* viz: display UOp tags
* g.tag
2025-06-10 16:02:23 +03:00
nimlgen and GitHub
800d1796d5
am_smi: kill process group ( #10750 )
2025-06-10 15:23:39 +03:00
qazal and GitHub
5bd4ad2e8b
viz: remove unused arg ( #10747 )
2025-06-10 12:00:09 +03:00
George Hotz and GitHub
413e223d6e
Revert "remove cpu graph, it's different from the others ( #10743 )" ( #10745 )
...
This reverts commit 3d64a98432 .
2025-06-09 22:40:48 -07:00
George Hotz and GitHub
3d64a98432
remove cpu graph, it's different from the others ( #10743 )
...
* remove cpu graph, it's different from the others
* remote was blacklisting CPUGraph
2025-06-09 22:17:10 -07:00
George Hotz and GitHub
245b1d3a46
move add/mul to MathTrait [pr] ( #10741 )
...
* move add to MathTrait [pr]
* both add and mul
2025-06-09 21:48:55 -07:00
George Hotz and GitHub
c28eceaf44
move to mathtraits.py ( #10742 )
2025-06-09 21:17:35 -07:00
George Hotz and GitHub
acf72872b3
move view left to the outer graph prereqs + testing ( #10725 )
...
* move view left to the outer graph
* global view right
* dont need that one
* remove comment
* test kernelize
* simple
* split onnx, test sdxl null
* fix testing
* ugh, wrong one
* Update test.yml
2025-06-09 20:43:25 -07:00
chenyu and GitHub
b7198fdcfd
linearizer failure from wino fuse arange cifar ( #10739 )
2025-06-09 23:10:19 -04:00
George Hotz and GitHub
58eebdb507
don't reassign metadata to the same uop + ignore oob in pr [pr] ( #10737 )
2025-06-09 18:43:39 -07:00
chenyu and GitHub
364b903850
minor cleanups in linearize.py [pr] ( #10735 )
2025-06-09 19:49:19 -04:00
George Hotz and GitHub
81ef879da3
non recursive top_down_rewrite ( #10729 )
...
* non recursive top_down_rewrite
* nicer algorithm
* rewrite bottom up also
* only top down is broken?
* simpler iterative algo
* no recursion errors
* top down and bottom up
* unified rewrite
* simpler rewrite
* clean up comments
* move that comment
2025-06-09 16:33:04 -07:00
chenyu and GitHub
53cbd4254b
suppress filter_too_much on test_float_cast_to_unsigned ( #10733 )
...
falky, already done in test_float_cast_to_unsigned_overflow and test_float_cast_to_unsigned_underflow
2025-06-09 18:30:04 -04:00
George Hotz and GitHub
916bbd5c6b
fixed point rewrite [pr] ( #10732 )
2025-06-09 14:46:20 -07:00
chenyu and GitHub
55cdbb9a20
fix mask in expand into symbolic size ( #10730 )
...
failed before when old size is 1 and it expands into symbolic size, because `resolve(s != ns, False)` is False and it does not expand the mask
2025-06-09 17:33:22 -04:00
wozeparrot and GitHub
926b11381c
failing test for symbolic expand after pad ( #10727 )
...
* feat: failing test for symbolic expand after pad
* feat: mark test as failing
2025-06-09 16:55:21 -04:00
chenyu and GitHub
49f999d919
update _reshape_mask for symbolic shape expand ( #10726 )
...
* don't merge shape symbolic reshape symbolic
* proper fix
2025-06-09 16:35:02 -04:00
wozeparrot and GitHub
27dd97f688
support variable shape none slice in getitem ( #10724 )
2025-06-09 11:53:02 -07:00
Ignacio Sica and GitHub
afd5140a09
remove no longer used IndexContext acc_num var ( #10720 )
2025-06-09 14:06:59 -04:00
George Hotz and GitHub
f84c320548
better external_benchmark_schedule [pr] ( #10722 )
2025-06-09 10:26:11 -07:00
George Hotz and GitHub
6270c0eac0
default ignore oob to 0 ( #10660 )
2025-06-09 10:25:43 -07:00
24d328e313
onnx parser ( #10435 )
...
* onnx parser
* fix compile, lint
* onnx.load -> onnx_load
* compatible with ModelProto
* fix test external_test_onnx_ops.py
* fix tests
* fix signed int
* reduce to 261 lines
* fix TypeProto.Optional
* debug for _parse_message, add TypeProto.Sequence, cleanup
* onnx_load from Tensor
* remove BufferedReader
* 174 lines and reduce tensor copy
* cleanup
* use onnx_load in external_model_benchmark.py
* fix qcom test
* [onnx] parser support external data
---------
Co-authored-by: b1tg <[email protected] >
Co-authored-by: chenyu <[email protected] >
2025-06-09 12:44:28 -04:00
Sieds Lykles and GitHub
cfa65bea05
Subtract 1 from Variable upper bound ( #10715 )
2025-06-09 09:25:53 -07:00
geohot
ef58ab340a
hotfix: remove n=auto from REMOTE=1 test
2025-06-09 09:19:36 -07:00
qazal and GitHub
419a1286f2
viz: share cacheKey [pr] ( #10717 )
2025-06-09 17:48:29 +03:00
chenyu and GitHub
35523dc35f
move BLOCK_REORDER to caller [pr] ( #10711 )
...
so block_reorder tests won't fail with flag set to 0
2025-06-08 23:26:01 -04:00
chenyu and GitHub
bb34c28b36
debug flag for linearize block_reorder [pr] ( #10710 )
2025-06-08 22:26:06 -04:00
chenyu and GitHub
d93a0bee6b
mlperf ci uses its own cache ( #10705 )
...
not to interfere with regular cache which is used by benchmark
2025-06-08 19:43:32 -04:00
qazal and GitHub
8cdf6e4d1e
viz memory graph tiny fixes [pr] ( #10709 )
...
* sched_sink is a step
* offset for yaxis
* clear existing
* scale offset
2025-06-09 01:10:12 +03:00
George Hotz and GitHub
81b9c04574
move high level stuff to unit tests [pr] ( #10708 )
...
* move high level stuff to unit tests [pr]
* process replay on unit tests
* fix pr, less compute
* set omp num threads
* set 200MB buffer size limit
* delete junk
* fix tests
* faster
* move test_indexing to unit
* faster
2025-06-08 14:05:56 -07:00
nimlgen and GitHub
171580e9ec
am: fix reg update ( #10707 )
2025-06-08 21:45:55 +03:00
George Hotz and GitHub
4305f532d9
clean up apt stuff ( #10706 )
...
* clean up apt stuff
* single apt install
* fixes
* fix opencl + ldconfig
2025-06-08 11:06:09 -07:00
George Hotz and GitHub
4e2c3560b4
smaller tests are faster tests [pr] ( #10704 )
...
* remove del spam from CI
* more
* preconstruct default buffer spec
* ignore those errors
* check exception
* more exception check
* skip stuff
* smaller tests mean faster tests
* a few more
2025-06-08 10:54:19 -07:00
George Hotz and GitHub
67a1c92fc0
remove del spam from CI ( #10699 )
...
* remove del spam from CI
* more
* preconstruct default buffer spec
* ignore those errors
* check exception
* more exception check
* skip stuff
2025-06-08 10:14:30 -07:00
George Hotz and GitHub
32141ec867
make apt CI faster ( #10702 )
2025-06-08 09:43:39 -07:00
chenyu and GitHub
4f535641f7
add one huggingface_onnx test to mac benchmark ci ( #10700 )
...
this crashed for me on onnx parser pr but seems fine for the author. see if ci mac is fine
2025-06-08 12:26:12 -04:00
George Hotz and GitHub
32e9949052
rename lazydata to uop ( #10698 )
2025-06-08 08:42:22 -07:00
uuuvn and GitHub
8e3f337075
Skip flaky test in ci ( #10696 )
...
`test_data_parallel_resnet_train_step` is already skipped on LLVM/CPU:
```python
@unittest.skipIf(CI and REAL_DEV in ("CUDA", "NV", "LLVM", "CPU"), "slow, and flaky on LLVM/CPU")
@unittest.skipIf(REAL_DEV == "WEBGPU" and not OSX, "WEBGPU Vulkan can only run kernels with up to 10 buffers")
def test_data_parallel_resnet_train_step(self):
```
It looks like `test_data_parallel_resnet` (no `_train_step`) is flaky in a similar way:
https://github.com/tinygrad/tinygrad/actions/runs/15472667248/job/43560773882?pr=10642#step:9:64
2025-06-08 08:24:09 -07:00
geohot
3ece2e4bb5
hotfix: remove accel from extra
2025-06-08 08:20:34 -07:00
qazal and GitHub
1ad8062591
more generic naming in VIZ [pr] ( #10695 )
...
* note
* rename kernel to ctx
* rename uop things to currentStep + expandSteps
* already destructured
* some things that were called ctx are steps
* still a kernel
2025-06-08 15:37:39 +03:00
qazal and GitHub
c70486908e
viz: clicking a KERNEL node can open codegen rewrite ( #10683 )
...
* work
* now it doesn't have 20% slowdown
* label like this
* closer
* ansiStrip
* remove
* better
* id is faster
* fix that
2025-06-08 13:11:03 +03:00
George Hotz and GitHub
48eb7d76b1
use ALLOW_DEVICE_USAGE context variable instead of MainProcess check ( #10693 )
...
* use DISALLOW_DEVICE_OPEN context variable instead of MainProcess check
* device usage can be disallowed
2025-06-08 00:07:40 -07:00
geohotstan and GitHub
dedff0e96c
fix run huggingface onnx debug ( #10679 )
2025-06-08 00:59:20 -04:00
George Hotz and GitHub
8c76250d31
speed up a few tests ( #10692 )
2025-06-07 20:39:25 -07:00
chenyu and GitHub
e80870e27c
BasicBlock2 -> BasicBlock [pr] ( #10691 )
2025-06-07 23:33:51 -04:00
George Hotz and GitHub
7ff175c022
cache a venv to avoid pip usage ( #10689 )
...
* try built in pip caching
* try venv
* export venv
* set VIRTUAL_ENV
* revert that
* venv key
* fix
* ci cache hit?
* fix windows
2025-06-07 20:13:41 -07:00
ihar and GitHub
40c1479267
added unit tests for 'argfix' ( #10678 )
2025-06-07 22:17:10 -04:00
ihar and GitHub
74b849b5e1
remove unnecessary 'argfix' because 'view' is an alias to 'reshape'. all functionality must be inside 'reshape' ( #10677 )
...
* remove unnecessary 'argfix' because 'view' is an alias to 'reshape'. all functionality must be inside 'reshape'
* added the same set of unit tests for 'view' as for 'reshape' since 'view' is just an alias for 'reshape'
* improved tests for 'view' op
2025-06-07 22:15:31 -04:00
chenyu and GitHub
e88fe41d37
update vits vctk model to use download from huggingface ( #10688 )
...
google drive points to a warning page that does not work
2025-06-07 20:47:28 -04:00
Sieds Lykles and GitHub
c29a56dd51
Fix whisper OOB ( #10685 )
...
* fix whisper and test
* remove import
2025-06-07 20:23:50 -04:00
George Hotz and GitHub
53ed64e133
ci speed work 1 ( #10676 )
...
* skip a few slow tests
* use a venv for python packages
* create venv
* no user, it's in venv
* ignore venv
* venv
* new cache key
* try that
* this
* version the python cache
2025-06-07 16:33:11 -07:00
George Hotz and GitHub
db01c5a08a
ramp.py file from stream ( #10686 )
2025-06-07 14:58:21 -07:00
Sieds Lykles and GitHub
2f605eadf7
fix oob ( #10666 )
2025-06-07 11:32:03 -04:00
qazal and GitHub
cb61774ab6
move shared viz fields out of serve.py [pr] ( #10684 )
...
* move shared viz fields out [pr]
* update javascript
* update test_viz
2025-06-07 17:18:18 +03:00
qazal and GitHub
b515d796fb
inline viz get_name [pr] ( #10682 )
...
* inline viz get_name [pr]
* changing name_fxn makes this simpler
* waitUntil dom
2025-06-07 11:16:16 +03:00
qazal and GitHub
86a19e19e8
cleanup bits of viz [pr] ( #10681 )
2025-06-07 09:18:12 +03:00
wozeparrot and GitHub
e3805171e2
feat: variable bs bitcast ( #10674 )
2025-06-06 17:21:53 -07:00
George Hotz and GitHub
54db1f8ee8
prevent huge waste of multi ram ( #10669 )
...
* prevent huge waste of multi ram
* fix ram usage
* only define var
* add resolve
* fix tests
* fix cifar training
* remove that logic
* fix test without long
2025-06-06 17:17:21 -07:00
b68b7dbc2a
test winograd is close to normal conv [pr] ( #10557 )
...
Co-authored-by: chenyu <[email protected] >
2025-06-06 19:11:49 -04:00
nimlgen and GitHub
85cea23557
nv: original bw qmd ( #10672 )
...
* nv: original bw qmd
* forgot
2025-06-07 01:43:22 +03:00
George Hotz and GitHub
5ef7c5923f
docs: remove unused METAL_XCODE env var ( #10421 )
2025-06-06 18:39:54 -04:00
Sidharth N. Babu and GitHub
ef14dfb277
compile fixes ( #10442 )
2025-06-06 18:38:37 -04:00
eb7305e6a4
Tensor.keccak("sha3_256") ( #7186 )
...
Co-authored-by: George Hotz <[email protected] >
Co-authored-by: George Hotz <[email protected] >
Co-authored-by: wozeparrot <[email protected] >
2025-06-06 15:24:05 -07:00
nimlgen and GitHub
346b8542da
nv: fix inval from gpu_get_id_info_v2 ( #10670 )
2025-06-07 00:54:32 +03:00
chenyu and GitHub
bdede4924e
fix odd number in get_test_global_size ( #10671 )
...
factor might not be a integer if input global_size has an odd number in it
2025-06-06 17:31:35 -04:00
George Hotz and GitHub
bf4ffc054c
mstack replaces scheduler complexity ( #10654 )
...
* mstack replaces scheduler complexity
* leave that one
* contiguous
* work
* upd
* minimal failing test
* simpler
* attention is broken
* fix transformer
* failing tests
* real fix for llama
* kv cache test
* jit multi assign test
* better tests
* comment
* fix jit issue
* traverse after buf_uop
2025-06-06 11:31:41 -07:00
George Hotz and GitHub
7f0f97aa76
new test_multitensor tests ( #10667 )
...
* new test_multitensor tests
* cleanup scheduler
2025-06-06 10:26:28 -07:00
qazal and GitHub
5170f387b3
remove UOp.metaop [pr] ( #10664 )
...
* little simpler UOp.const_like [pr]
* remove UOp.metaop
* bind
* remove
* min diff
* that comment is fine
2025-06-06 16:21:48 +03:00
4a6d84c4c3
hotfix llama start_pos vmax is max_context-1 ( #10659 )
...
* hotfix llama start_pos vmax is max_context-1
fixed `IGNORE_OOB=0 python3 examples/llama3.py --size 1B --benchmark --temperature 0`
* hotfix: multitensor transformer test tests kv cache
---------
Co-authored-by: George Hotz <[email protected] >
2025-06-06 00:41:25 -04:00
geohot
5eb6e1e65a
Revert "hotfix: multitensor transformer test tests kv cache"
...
This reverts commit ad9f88419a .
2025-06-05 21:15:34 -07:00
geohot
ad9f88419a
hotfix: multitensor transformer test tests kv cache
2025-06-05 21:08:57 -07:00
George Hotz and GitHub
8325c4f192
tests for multi assign ( #10658 )
...
* tests for multi assign
* transformer tests
* add that assert
2025-06-05 20:56:40 -07:00
wozeparrot and GitHub
0d86f8d375
fix failed threefry ( #10646 )
2025-06-05 17:17:42 -07:00
chenyu and GitHub
e67642d430
update doc example for multinomial ( #10657 )
...
also added many `s` for consistency
2025-06-05 20:16:52 -04:00
Eitan Turok and GitHub
61352b8aa2
Add some more docs ( #10634 )
...
* more docs
* Add multinomial to ops
* better doc
2025-06-05 19:40:37 -04:00
qazal and GitHub
884b6cf288
remove gbarrier on const ( #10656 )
2025-06-06 02:36:52 +03:00
chenyu and GitHub
ff1aad7b69
fix const float pow to int tensor ( #10655 )
...
was incorrectly casted into int
2025-06-05 19:15:12 -04:00
George Hotz and GitHub
6619f17e26
force store to be contiguous ( #10652 )
2025-06-05 15:42:54 -07:00
wozeparrot and GitHub
37e1ef1be3
feat: cleanup old AM processes ( #10653 )
2025-06-05 15:41:00 -07:00
George Hotz and GitHub
baba274a76
minimal mstack pr to fix allreduce ( #10649 )
...
* minimal mstack pr to fix allreduce
* fix webgpu
2025-06-05 15:14:53 -07:00
George Hotz and GitHub
4c315f8e17
MSTACK little non-functional changes ( #10648 )
2025-06-05 13:20:22 -07:00
79d04d1baf
AMD_LLVM: support mfma for mi300x ( #10625 )
...
* amd llvm: support mfma for mi300x
* don't pass self
* refactor wmma render
* arch as lambda arg
---------
Co-authored-by: b1tg <[email protected] >
2025-06-05 15:55:44 -04:00
chenyu and GitHub
46811d0d3c
minor external_model_benchmark cleanup ( #10644 )
2025-06-05 14:13:28 -04:00
qazal and GitHub
26afbc954f
delete redundant tests from test_schedule [pr] ( #10643 )
2025-06-05 20:08:39 +03:00
chenyu and GitHub
80ebce421d
remove metal buffer limit in external_model_benchmark [pr] ( #10642 )
...
not needed anymore
2025-06-05 13:00:51 -04:00
qazal and GitHub
28c4997236
check for matching shape order in fused reduce ( #10641 )
...
* failing test
* shapes match with ones removed
2025-06-05 19:37:22 +03:00
qazal and GitHub
1190062812
prevent grouper can_chase while fusing arange [pr] ( #10623 )
2025-06-05 18:50:21 +03:00
uuuvn and GitHub
69f7778985
refactor renderer launch bounds [pr] ( #10617 )
2025-06-05 08:38:04 -07:00
qazal and GitHub
8c5ea00522
push permutes through fused reduces ( #10628 )
...
* fix pushing reshapes through reduceops
* reduceop_view_right should assert on ndims mismatch
* update that, view.reshape asserts it
2025-06-05 16:14:04 +03:00
qazal and GitHub
8db0ba1161
simpler swizzle_reducop + comments [pr] ( #10638 )
2025-06-05 13:54:49 +03:00
qazal and GitHub
ed37f29184
remove unused lib directory from viz setup [pr] ( #10639 )
2025-06-05 13:54:31 +03:00
chenyu and GitHub
f6d7db25b7
simpler unbind_view [pr] ( #10636 )
2025-06-05 01:03:27 -04:00
chenyu and GitHub
d0969f5a1f
cleanup multi tests ( #10635 )
2025-06-05 00:28:44 -04:00
qazal and GitHub
571c0296a9
linearizer failure from FUSE_ARANGE default diff ( #10629 )
...
* start with test_arange_sum
* test_arange_avgpool2d
* device.renderer.supports_float4
2025-06-04 19:11:52 +03:00
qazal and GitHub
5056d21b29
add failing TestSchedule.test_arange_sum [pr] ( #10627 )
2025-06-04 17:23:59 +03:00
9acaa6bc9a
Fix button layout in viz UI for safari ( #10621 )
...
Co-authored-by: Utkarsh Gill <[email protected] >
Co-authored-by: qazal <[email protected] >
2025-06-04 15:33:22 +03:00
Xingyu and GitHub
7a1bfb668d
Implement linalg_eigh function for tensor eigenvalue decomposition in torch backend ( #10612 )
...
* Implement private _linalg_eigh function for tensor eigenvalue decomposition in torch backend
* Add unit test for linalg.eigh function in TestTorchBackend
This test verifies the eigenvalue decomposition of a 2x2 tensor using the linalg.eigh function, ensuring the computed eigenvalues and reconstructed tensor match the expected results.
2025-06-04 07:59:50 -04:00
qazal and GitHub
7114b6ab31
viz browser tests ( #10626 )
...
* viz browser tests
* expect failure if js/ isn't included
* back green
2025-06-04 14:58:24 +03:00
Fang-Pen Lin and GitHub
b0913295d2
Add missing js files in python package data for viz ( #10624 )
2025-06-04 10:49:43 +03:00
wozeparrot and GitHub
4d1686f767
clean: becnhmark -> benchmark ( #10620 )
2025-06-03 19:28:18 -07:00
chenyu and GitHub
18e9ec3ea1
add wino cifar to search benchmark ( #10615 )
...
* add wino cifar to search benchmark
* FUSE_OPTIM=1
* revert those
2025-06-03 20:38:43 -04:00
Bhavya Gada and GitHub
bafd0c30d7
fix some minor typos and grammar ( #10619 )
2025-06-03 15:55:25 -07:00
nimlgen and GitHub
4381b54543
am: disable page migration ( #10608 )
...
* am: disable page migration
* fixed
* enable
* fxi
* typ
* fix check
2025-06-03 18:51:28 +03:00
chenyu and GitHub
1c1f578490
DISABLE_COMPILER_CACHE in sdxl search ( #10614 )
2025-06-03 09:22:25 -04:00
qazal and GitHub
ce9f12dc13
reorder cast before masking constants ( #10609 )
...
* failing test from fuzzer
* .numpy() handles bfloat16 better
* const->view->cast becomes const->cast->view
* update TestMovedConstFolding.test_cast_padded
2025-06-03 15:44:03 +03:00
qazal and GitHub
910cabb081
add kernel count to grouper process replay differ [pr] ( #10611 )
2025-06-03 15:21:27 +03:00
chenyu and GitHub
26dee71bc1
hotfix don't overwrite acc dtype in scatter_reduce ( #10606 )
...
dtype is inferred by individul reduce
2025-06-02 21:17:01 -04:00
ihar and GitHub
ba02a6331e
removed unnecessary 'isinstance(data, UOp)' check ( #10605 )
2025-06-02 20:58:14 -04:00
nimlgen and GitHub
07de095b27
am: more info on PFs ( #10602 )
...
* am: more info on PFs
* fix
2025-06-02 23:48:40 +03:00
qazal and GitHub
b8fb2ba829
rename to finalize_gbarrier [pr] ( #10596 )
2025-06-02 12:55:31 +03:00
Ahmed Harmouche and GitHub
650404a143
[webgpu] Proper shared mem size for packed types ( #10585 )
...
* Proper shared mem size in webgpu
* Add test
* Refactor test
2025-06-01 20:18:33 -04:00
qazal and GitHub
00822603ec
allow stacking of VIEW UOps [pr] ( #10532 )
...
* allow stacking of VIEW UOps [pr]
* merge_views is first
* simpler
* loc for pr, this needs a helper
* keep
* diff [pr]
* formatting
2025-06-01 23:27:04 +03:00
qazal and GitHub
3cc73a0172
simpler process replay main loop [pr] ( #10588 )
...
* simpler process replay main loop [pr]
* use logging
* default to 1
2025-06-01 15:03:21 +03:00
qazal and GitHub
dc882d3d7d
merge process replay and viz captures [pr] ( #10581 )
...
* refactoring
* test script
* work
* more work
* diff
* repr splits lines correctly
* that
* add location
* add location
* also don't need name_override
* k.copy
* [pr]
* name_override 2
* err
2025-06-01 12:30:10 +03:00
qazal and GitHub
1f8a8721e9
remove test_unaligns_idxs, UOps don't have order like this [pr] ( #10587 )
2025-06-01 12:16:14 +03:00
ihar and GitHub
c45936c4fc
replaced '.upper()' which is never needed with '.lower()' which were duplicated ( #10586 )
2025-05-31 20:58:42 -04:00
ihar and GitHub
88f38d3fcc
remove '_metaop' because it is an old wrapper around 'UOp.metaop' with no additional functionality anymore ( #10583 )
2025-05-31 14:06:39 -04:00
chenyu and GitHub
77c7989fa0
remove a MUL rewrite rule for wgsl ( #10582 )
...
tests are fine without it
2025-05-31 14:05:49 -04:00
Ahmed Harmouche and GitHub
35eb4d357a
[webgpu] Fix atomic shared mem load inside loop ( #10530 )
...
* Disable shared mem atomics on webgpu
* allow_any_len in load pattern matcher to fix temp load inside loop
2025-05-31 09:29:02 -04:00
qazal and GitHub
6af4b02374
use plain dict and list in grouper [pr] ( #10580 )
2025-05-31 13:09:59 +03:00
chenyu and GitHub
4ab3391e6f
set -o pipefail for mlperf run_and_time (#10577 )
...
also run the 5.1 script in ci cron job
2025-05-30 16:36:44 -04:00
chenyu and GitHub
baf482d314
copy mlperf stuff to 5.1 ( #10576 )
...
5.0 is finalized, new changes go to 5.1
2025-05-30 16:12:39 -04:00
nimlgen and GitHub
883bb4541c
am: reserve address space ( #10564 )
...
* am: reserve address space
* f
* cc
* errno
* fix
* always has cpu mapping
2025-05-30 19:31:03 +03:00
qazal and GitHub
e0305e54fc
remove custom merge_views rewrite rule for buffer ops [pr] ( #10574 )
2025-05-30 15:27:13 +03:00
qazal and GitHub
de9597a8a9
cleanup kernel.py ShapeTracker replacement [pr] ( #10573 )
2025-05-30 15:06:01 +03:00
qazal and GitHub
5b59728c75
refactor LOAD(DEFINE_GLOBAL, VIEW) in kernels to LOAD(VIEW(DEFINE_GLOBAL)) ( #10541 )
...
* changes to core tinygrad
* fixups pt1
TC=3
docs/abstractions2.py
IMAGE=2
test_quantize_dsp
test_schedule
* more tests
* green now
* images stay images
2025-05-30 14:27:58 +03:00
chenyu and GitHub
116ffc4e92
cstyle strips paren for AND and OR ( #10560 )
2025-05-30 07:09:05 -04:00
qazal and GitHub
bbf05110a2
use kernelize in TestLinearizer.test_indexing_multireduce [pr] ( #10571 )
2025-05-30 11:27:09 +03:00
qazal and GitHub
7051bf3fd5
fixup hardcoded asts ptr dtype and constants [pr] ( #10570 )
...
* fixup hardcoded asts ptr dtype and constants [pr]
* use kernelize for test_kernel_count
2025-05-30 09:38:32 +03:00
qazal and GitHub
066196415f
UOp.valid and const_like work with just shapes [pr] ( #10569 )
...
* UOp.valid and const_like work with just shapes [pr]
* pm_quant left
* pm_quant
2025-05-30 08:55:06 +03:00
wozeparrot and GitHub
5e3c4a8431
fix: comma testsig ( #10568 )
2025-05-29 19:00:07 -07:00
Eitan Turok and GitHub
c07f13c438
Docs for masked_fill ( #10558 )
...
* add docs
* fix doc examples
* add to docs
* fix typo
2025-05-29 03:49:02 -07:00
George Hotz and GitHub
b3b43a82c4
remove Tensor.no_grad, it's meaningless now [pr] ( #10556 )
2025-05-28 22:20:02 -07:00
George Hotz and GitHub
e4e7b5d7e1
continue work on beautiful cifar ( #10555 )
2025-05-28 21:42:01 -07:00
George Hotz and GitHub
e140f8f0d8
linearizer test_failure_61 ( #10552 )
...
* enumerate cases of Tensors in the JIT
* optional fused optimizers
* add fused optimizer test
* move that there
* ugh
* work on beautiful_cifar
* speed close to hlb_cifar
* test_failure_61
* just the failure
2025-05-28 21:30:50 -07:00
George Hotz and GitHub
871df1436a
more beautiful cifar ( #10551 )
...
* enumerate cases of Tensors in the JIT
* optional fused optimizers
* add fused optimizer test
* move that there
* ugh
* work on beautiful_cifar
* speed close to hlb_cifar
* schedule to corealize all
* one line sched step
* less lines
2025-05-28 20:48:20 -07:00
George Hotz and GitHub
ee12e801a3
optional fused optimizers ( #10549 )
...
* enumerate cases of Tensors in the JIT
* optional fused optimizers
* add fused optimizer test
* move that there
* ugh
2025-05-28 13:50:30 -07:00
ae02a1e232
[bounty] Z3 symbolic fuzzer [pr] ( #10514 )
...
* First version, caught a bug?
* Nicely print failure to reproduce
* Remove that
* Put the assert back
* Change fuzzing to use testing_unit so it has z3
* Test key to match
* Add rule
* Add test
* Add test for edge case 0
* Merge patterns
* update comment
* consistent whitespace
* whitespace
* add condition
* add test
* update comment
* use Variable
* fuzzer using z3_renderer
* Cleaned up printing and debugging
* working new fuzzer
* change some comments and printing
* more formatting
* fuzz failures in seperate file
* fix fstring
* more tests
* naming
* remove added line
* remove comment
* print number of skipped expressions
* use self.assertEqual
---------
Co-authored-by: chenyu <[email protected] >
2025-05-28 16:28:37 -04:00
chenyu and GitHub
74cf5dbd9e
mlperf system updates ( #10550 )
...
standardized processor and accelerator names
2025-05-28 16:15:46 -04:00
George Hotz and GitHub
98f3d1c26d
enumerate cases of Tensors in the JIT ( #10548 )
2025-05-28 11:51:27 -07:00
nimlgen and GitHub
d1d9e729fd
am_smi: mem usage ( #10547 )
2025-05-28 16:53:31 +03:00
chenyu and GitHub
23e41f523a
sdxl also run with cached search ( #10546 )
2025-05-28 06:51:56 -04:00
chenyu and GitHub
fffdc4d31c
workflow to run sdxl with search ( #10543 )
2025-05-27 17:25:41 -04:00
qazal and GitHub
d1f0043331
use store_val helper in test_schedule asserts [pr] ( #10540 )
2025-05-27 21:48:06 +03:00
George Hotz and GitHub
5b268121d4
remove becomes map ( #10533 )
...
* remove becomes map
* add comment and delete dead code
* multi is a view
2025-05-27 11:47:11 -07:00
qazal and GitHub
271110bb5a
s/src[0]/buf in lowerer.py [pr] ( #10539 )
2025-05-27 21:08:54 +03:00
qazal and GitHub
f0042629d1
replace arg in merge_view [pr] ( #10537 )
2025-05-27 21:00:24 +03:00
qazal and GitHub
617ecc1a7b
fixup grouper process replay [pr] ( #10538 )
2025-05-27 20:46:55 +03:00
George Hotz and GitHub
a07caaca0d
handle stride 0 variable reshape ( #10536 )
2025-05-27 10:00:24 -07:00
George Hotz and GitHub
0515622d95
the schedule graph is the tensor graph ( #10534 )
...
* the schedule graph is the tensor graph
* gate type_verify on debug
* relax that spec
* unmasked check is okay
2025-05-27 09:23:57 -07:00
qazal and GitHub
142f6ba873
move merge_views to grouper swizzler [pr] ( #10531 )
2025-05-27 16:33:26 +03:00
qazal and GitHub
c03e9c8995
fixup typing for @diskcache [pr] ( #10529 )
2025-05-27 15:30:49 +03:00
George Hotz and GitHub
41e3d07d7f
view gradient is tricky ( #10528 )
...
* view gradient is tricky
* explicit
2025-05-26 22:28:30 -07:00
George Hotz and GitHub
ab4ca5da29
remove gradient nonsense [pr] ( #10527 )
...
* remove gradient nonsense [pr]
* grads to base
2025-05-26 19:09:59 -07:00
chenyu and GitHub
76eb130d8c
hotfix: BenchEvent MLPERF_RUN is mlperf_run ( #10526 )
2025-05-26 20:19:37 -04:00
uuuvn and GitHub
c29c46853f
Very basic mock sqtt ( #10512 )
...
This mockgpu sqtt emulation will just ignore basically everything and end
up with a 0x1000 size trace full of zeroes, but just testing for things
like register rename is better than nothing i guess
2025-05-26 14:38:28 -07:00
chenyu and GitHub
51dc7eedb0
correct use AM for resnet run_and_time ( #10524 )
2025-05-26 15:33:11 -04:00
chenyu and GitHub
c1919ad55f
use AM for resnet run_and_time ( #10523 )
2025-05-26 14:50:49 -04:00
geohot
e9bb2052cf
hotfix: update readme
2025-05-26 10:28:16 -07:00
6d07087fe1
remove contiguous from MSELECT 2 ( #10522 )
...
* remove contiguous from MSELECT
* test_shrink_on_shard_axis
---------
Co-authored-by: George Hotz <[email protected] >
2025-05-26 19:19:01 +03:00
geohotstan and GitHub
602a145f8f
Add Tensor.unfold ( #10518 )
...
* yoinked 10272
* eitanturok's fixes
* hmmm should size be sint?
* add test
2025-05-26 11:15:44 -04:00
qazal and GitHub
9169dcfb49
do not create kernels with more inputs than the backend allows ( #10510 )
...
* work
* no itertools + top down pass
* clean viz
* python can do that
* webgpu
* gbarrier of gbarrier is gbarrier
* device can be tuple
* bug in toposort
* failing test for gated toposort
* contiguous of gbarrier is gbarrier
* check for binops
* Revert "check for binops"
This reverts commit 53e3cdf720 .
* viz + match on gbarrier, self exists by default
* alt
* green now
* cleanup
2025-05-26 18:02:03 +03:00
nimlgen and GitHub
deb369417c
am_smi: print device usage ( #10520 )
...
* am_smi: print device usage
* tiny comments
2025-05-26 17:17:56 +03:00
chenyu and GitHub
2d50efb92b
set -e on mlperf run_and_time scripts (#10519 )
2025-05-26 09:22:30 -04:00
Sieds Lykles and GitHub
478c76f4b7
More div conditions ( #10432 )
...
* add condition
* add test
* use Variable
2025-05-26 07:36:05 -04:00
Sieds Lykles and GitHub
c6c7882bdf
bugfix: seperate rule for x//d<-c ( #10148 )
...
* Add rule
* Add test
* Add test for edge case 0
* Merge patterns
* update comment
* consistent whitespace
* whitespace
* update comment
2025-05-26 07:35:41 -04:00
chenyu and GitHub
2eeea373af
add BENCHMARK_LOG for mlperf resnet cron ( #10516 )
2025-05-25 22:00:29 -04:00
a1f64af92d
ci: setup llvm for amdremote ( #10507 )
...
Co-authored-by: b1tg <[email protected] >
2025-05-25 21:52:27 -04:00
geohotstan and GitHub
fd9f236a82
move test over ( #10508 )
2025-05-25 21:51:51 -04:00
wozeparrot and GitHub
7c81f9f95e
fix: gate mlperf workflow ( #10515 )
2025-05-25 17:06:21 -07:00
Panagiotis Kourouklidis and GitHub
4941486cb0
Add method to return field masks for AMDReg ( #10511 )
2025-05-25 14:47:20 -07:00
nimlgen and GitHub
88c5864bf3
nv: do not hardcode sass version ( #10513 )
2025-05-25 22:41:15 +03:00
geohot
941cbd3471
hotfix: amd works on arch linux w/o rocm
2025-05-24 16:47:13 -07:00
nimlgen and GitHub
d90ddcc365
nv: blackwell support ( #10487 )
...
* nv: blackwell support
* fixes
* hm
* h
* fixes
* mypy
* xx
* yy
* arr
* revert
* oops
* unrelated
2025-05-24 18:23:53 +03:00
chenyu and GitHub
dc6309242d
WallTimeEvent for mlperf ci ( #10506 )
2025-05-24 10:56:03 -04:00
qazal and GitHub
dd5601af68
readable COPY(VIEW) reordering [pr] ( #10505 )
...
* readable COPY(VIEW) reordering [pr]
* assert that
* spec
* resolve
* Revert "resolve"
This reverts commit f5629fbef8 .
* arg
2025-05-24 17:08:58 +03:00
Ahmed Harmouche and GitHub
bbb6deff53
Increase op limit in test_index_mnist to pass on webgpu ( #10504 )
...
* Increase op limit to enable mnist indexing on webgpu
* Only relax op_limit on WebGPU
2025-05-24 09:37:31 -04:00
nimlgen and GitHub
c472ab636c
nv: use regcount from meta ( #10503 )
2025-05-24 14:14:33 +03:00
qazal and GitHub
82b444796d
fix display of kernel args in viz [pr] ( #10502 )
2025-05-24 14:09:52 +03:00
qazal and GitHub
a9d0bf5c4c
proper error for device mismatch ( #10500 )
...
* failing test
* use bufs
* buf_uop
* not on cpu
2025-05-24 12:17:41 +03:00
qazal and GitHub
fc1300f5e3
top down create_kernels + delete "replace assign sources" ( #10478 )
...
* rebase from #10468
* fixup metadata 2
* that too
* comments for metadata
* remove_gbarrier is not needed anymore
* skip that
* break metadata more
* delete more metadata fixups
* err, fix kernelize diamond
* unskip metadata
* new_map
* roots
* replace metadata of roots
* check empty
* replace globals is better
2025-05-24 09:50:06 +03:00
George Hotz and GitHub
9eee5ae276
its copying the dataset every time ( #10498 )
...
* its copying the dataset every time
* add comment
* expect failure
* todo
2025-05-23 21:25:53 -07:00
George Hotz and GitHub
4467b52721
remove all self copies after Tensor.clone fix ( #10494 )
2025-05-23 19:04:20 -07:00
George Hotz and GitHub
b58f2d4544
fix tests ( #10493 )
2025-05-23 18:38:07 -07:00
6b8eb5fec2
split mlperf to its own red benchmark run ( #10492 )
...
* Add mmapeak implementation for 7900 XTX
* Change identation
* Use a template instead of multiple assebly files
* Fix output formatting
* Reduce register file bank conflicts
* More accurate measurement for quick instructions
* Add support for gfx1201
* RDNA4 wmma requires less VGRPs
* RDNA4 does not have s_cmpk instructions
* Add v_wmma_i32_16x16x32_iu4 for gfx1201
* Add sparse wmma instructions
* split to tinybox red MLPerf Benchmark
---------
Co-authored-by: Panagiotis Kourouklidis <[email protected] >
2025-05-23 17:12:41 -07:00
e21836952d
mmapeak implementation for 7900 XTX ( #10417 )
...
* Add mmapeak implementation for 7900 XTX
* Change identation
* Use a template instead of multiple assebly files
* Fix output formatting
* Reduce register file bank conflicts
* More accurate measurement for quick instructions
* Add support for gfx1201
* RDNA4 wmma requires less VGRPs
* RDNA4 does not have s_cmpk instructions
* Add v_wmma_i32_16x16x32_iu4 for gfx1201
* Add sparse wmma instructions
---------
Co-authored-by: George Hotz <[email protected] >
2025-05-23 16:26:12 -07:00
George Hotz and GitHub
0a313d98a0
add rocm 6.4 support ( #10491 )
...
* add rocm 6.4 support
* update to newer amdcomgr, assert lang is right
* fix aux-triple
2025-05-23 16:20:54 -07:00
wozeparrot and GitHub
a18963d9e7
feat: use tinygrad useragent ( #10488 )
2025-05-23 15:44:40 -07:00
George Hotz and GitHub
0ebd440872
add mselect op ( #10453 )
...
* add mselect op
* more work
* that shouldn't be contiguous
* remove junk
* it segfaults...
* more correct
* test fail
* inserting a contiguous fixes it
* fix children in mselect
* complain
* error
* push RESHAPE through MSELECT
* no copy arg, use mselect
2025-05-23 14:11:37 -07:00
George Hotz and GitHub
bf2a0907be
gate the mockdsp behind MOCKDSP=1 [pr] ( #10486 )
2025-05-23 11:44:02 -07:00
3ca5680920
Test remote in benchmark ( #10304 )
...
hlb cifar is fast so added it, can add bert too if you think it's ok
6 real gpus to test multigraph and transfers + accuracy validation
should probably be added to tinystats too, i don't know how though
Co-authored-by: chenyu <[email protected] >
2025-05-23 12:12:57 -04:00
qazal and GitHub
7a762f01ab
s/shape_spec/ast_spec [pr] ( #10485 )
2025-05-23 15:43:54 +03:00
qazal and GitHub
127a7c8aee
assert AST views only exist in the edges ( #10484 )
...
* assert AST views only exist in the edges
* valid without device
2025-05-23 15:27:09 +03:00
qazal and GitHub
e491168685
add metadata note + whitespace fixup [pr] ( #10483 )
...
* add metadata note + whitespace fixup [pr]
* TestSchedule.test_kernelize_diamond
2025-05-23 14:37:45 +03:00
chenyu and GitHub
c5acb4e06e
run mlperf resnet daily ( #10482 )
...
Runs at 08:05 UTC (12:05 AM Pacific Time)
2025-05-23 07:16:20 -04:00
Sieds Lykles and GitHub
ce6ebfb8ee
verify rewrites in test_uop_symbolic ( #10430 )
...
* verify rewrites in test_uop_symbolic
* use global context
2025-05-23 06:57:29 -04:00
qazal and GitHub
52e8b69d98
create_kernels only matches on GBARRIER and ASSIGN [pr] ( #10480 )
2025-05-23 11:11:57 +03:00
George Hotz and GitHub
1e4d63e06e
uops can have multiple metadata ( #10479 )
...
* uops can have multiple metadata
* fixups
2025-05-22 21:35:02 -07:00
George Hotz and GitHub
283586bb96
insert GBARRIER into graph ( #10468 )
...
* insert contiguous into graph
* exclude contiguous from kernels
* and copy
* not needed on copy
* gbarrier
* gbarrier closer
* gb
* gb
* fix double realize logic bug
* remove gbarrier
* del that
* uop tags
* tag
* fix setitem, flaky
* no ctx there
* flip rewrite
* revert order until metadata is fixed
2025-05-22 20:53:36 -07:00
George Hotz and GitHub
d2bb50d75b
graph_rewrite_map in the other order [pr] ( #10476 )
...
* graph_rewrite_map in the other order [pr]
* reversed to preserve behavior
2025-05-22 20:22:07 -07:00
George Hotz and GitHub
9fc01c1e03
support for uop tags ( #10477 )
...
* support for uop tags [pr]
* test uop tags
2025-05-22 19:53:48 -07:00
chenyu and GitHub
8cc2dff4d8
only float Tensors have gradient [pr] ( #10475 )
2025-05-22 21:02:11 -04:00
George Hotz and GitHub
147f7747f2
remove the map from create_schedule_with_vars [pr] ( #10472 )
2025-05-22 15:58:25 -07:00
George Hotz and GitHub
6d5f87a18a
lshift/rshift reverse is broken [pr] ( #10467 )
2025-05-22 13:01:48 -07:00
Mike Ashcroft and GitHub
209d4401f8
Merge SimpleMathTrait and MathTrait ( #10463 )
2025-05-22 11:47:22 -07:00
George Hotz and GitHub
0d39bb5de1
rename to get_kernelize_map ( #10465 )
2025-05-22 11:44:44 -07:00
Xingyu and GitHub
1e0a59aca4
fix: handle buffer size calculation in to_movement_ops and add scalar assignment test in torch_backend ( #10464 )
2025-05-22 10:54:13 -07:00
George Hotz and GitHub
577a0b4cfa
openpilot compile4 (wip) ( #10407 )
...
* openpilot compile4
* add copies
* remove junk
2025-05-22 10:47:34 -07:00
George Hotz and GitHub
ab591fa4dd
make schedule explicit about kernels [pr] ( #10462 )
2025-05-22 09:32:16 -07:00
geohot
c46edbf262
hotfix: add note to relu
2025-05-22 09:13:38 -07:00
George Hotz and GitHub
c6cbf0145a
check that arg on copy is only used on multi [pr] ( #10461 )
2025-05-22 09:08:43 -07:00
Ignacio Sica and GitHub
f69722dc2a
refactor cuda disassemble ( #10449 )
2025-05-22 08:58:24 -07:00
qazal and GitHub
5c4cfbc22c
remove merge_views from kernel grouping rewrite [pr] ( #10457 )
2025-05-22 18:36:54 +03:00
nimlgen and GitHub
035dffb00c
nv: refactor qmd from ctypes ( #10459 )
...
* nv: refactor qmd from ctypes
* shorter
* imports
* x
* fix prefetch
2025-05-22 17:20:11 +03:00
wozeparrot and GitHub
12285e926a
fix: apply ip version fixes during AMDIP creation ( #10454 )
2025-05-22 10:14:48 +03:00
Ignacio Sica and GitHub
5e6b96a1be
align 16 in ptx, metal, cuda and amd ( #10450 )
2025-05-21 14:38:54 -07:00
nimlgen and GitHub
570cb89652
amd: handle all exceptions ( #10448 )
...
* amd: handle all exceptions
* linter
2025-05-21 16:51:44 +03:00
nimlgen and GitHub
475a7583b3
usbgpu: tiny changes ( #10445 )
2025-05-21 16:20:35 +03:00
qazal and GitHub
7720c1aef1
hotfix: remove viz_sz.py [pr] ( #10446 )
2025-05-21 14:17:42 +03:00
chenyu and GitHub
7bfb20757c
fix tensor int floor div ( #10327 )
...
* fix tensor int floor div
* test_float_floordiv_scalar
2025-05-21 06:46:54 -04:00
Sieds Lykles and GitHub
2b4375f36d
Correct divmod folding behind flag ( #10433 )
...
* add flag
* add test
* remove import
2025-05-21 06:46:13 -04:00
qazal and GitHub
df4cbb69e9
move fuzz_schedule.py to extra [pr] ( #10444 )
2025-05-21 10:07:24 +03:00
chenyu and GitHub
29624af872
skip commavq in external_model_benchmark ( #10439 )
...
precision issue with different onnxruntime version
2025-05-21 01:45:33 -04:00
George Hotz and GitHub
03e7a99ca8
add edge cases found by codex [pr] ( #10423 )
...
* add edge cases found by codex [pr]
* another test
* more edgecases
* docs
* instructions
* fine, add that one
* nan cases
* roll failures
* inv prob
* more failing tests
* err, that's failing
* more tests
* more failures
* uop verif
* failures
* webgpu
2025-05-20 14:53:18 -07:00
nimlgen and GitHub
2895198c36
am: download regs ( #10419 )
...
* am: download regs
* x
* linter
* mypy
* after merge
* raise
* fixed name
* fix
* xx
* remove
* missing reg
* missing reg
* move to online
* ops
2025-05-20 18:59:56 +03:00
nimlgen and GitHub
965f9e0696
amd: refactor amdreg ( #10427 )
2025-05-20 15:23:28 +03:00
nimlgen and GitHub
0b65c367f5
hotfix: make vfio not default ( #10429 )
...
* to validate
* disable vfio causing blockingio err
2025-05-20 15:08:49 +03:00
nimlgen and GitHub
cfa5c1cac6
am: disable idle d3 ( #10428 )
2025-05-20 13:50:58 +03:00
nimlgen and GitHub
252c1dc737
am: close flock in fini ( #10426 )
2025-05-20 13:15:48 +03:00
George Hotz and GitHub
ceb9d94eab
Update AGENTS.md
2025-05-19 17:59:59 -07:00
geohot
9389edf7ac
hotfix: add AGENTS.md
2025-05-19 17:48:42 -07:00
uuuvn and GitHub
ec9955c956
Use REAL_DEV for test skips ( #10420 )
...
This should fix remote cpu tests flakiness (segfaults were in
`test_data_parallel_resnet_train_step` which is skipped on cpu but wasn't
skipped on remote cpu)
2025-05-19 17:32:14 -07:00
nimlgen and GitHub
9a199ccd81
am: try to modprobe vfio ( #10418 )
...
* am: try to modprobe vfio
* fix
2025-05-19 23:46:50 +03:00
chenyu and GitHub
67d1364106
update LOGMLPERF in red resnet run_and_time ( #10416 )
2025-05-19 13:23:33 -04:00
Sieds Lykles and GitHub
db09676250
Dont simplify gate in gate, fix FUSE_ARANGE=1 python test/test_ops.py TestOps.test_scatter_add ( #10411 )
...
* substitute out index
* Add test
* change comment
2025-05-19 13:16:21 -04:00
chenyu and GitHub
116d9e6306
run mlperf resnet on red box ( #10413 )
...
also made push to `update_mlperf` branch trigger
2025-05-19 12:48:36 -04:00
geohot
f1fe1f93c1
hotfix: 14000 lines
2025-05-19 09:40:53 -07:00
qazal and GitHub
90eb3c0e5d
add MobileNetV2 benchmark to comma CI ( #10250 )
...
* add MobileNetV2 to comma CI
* symlink imagenet
* also the signature
* comment that out
* need imagenetmock
* same train and test set
* quantize on CPU=1
* verbose
* need __hexagon_divsf3
* 0x858d6c15
* quant cpu + CC=clang-19
2025-05-19 18:22:50 +03:00
qazal and GitHub
f9a5ad24c5
faster viz to_program [pr] ( #10410 )
...
* faster viz to_program [pr]
* Callable
2025-05-19 12:27:49 +03:00
qazal and GitHub
cc8dda1d75
move multi_map to grouper rewrite pass ( #10409 )
...
* move multi_map to grouper rewrite pass
* delete that
2025-05-19 10:44:06 +03:00
George Hotz and GitHub
b06291077c
no amdgpu kernel driver ( #10408 )
...
* no amdgpu kernel driver
* don't test hip
* lower req
2025-05-18 20:52:39 -07:00
geohot
4b1f1a47bb
hotfix: allow ModuleNotFoundError in metal llvm import
2025-05-18 20:46:31 -07:00
chenyu and GitHub
485e80da69
run_and_time for resnet ci ( #10405 )
2025-05-18 23:39:57 -04:00
qazal and GitHub
d1eeb19437
count viz javascript in lines ( #10403 )
...
* count viz javascript in lines
* don't count }
* it's javascript
* share with autogen
2025-05-18 19:34:00 -07:00
qazal and GitHub
260d194523
merge insert_fuse and do_fuse [pr] ( #10406 )
2025-05-19 04:44:36 +03:00