geohot
b4ab6de416
opt work
2025-08-01 16:56:47 -07:00
geohot
3d31b0b5f6
axistype in range
2025-08-01 16:22:54 -07:00
geohot
5bab842337
beautiful mnist
2025-08-01 15:12:17 -07:00
geohot
427f773bc2
all test ops pass
2025-08-01 11:53:56 -07:00
geohot
45a09207f9
test ops work (except empty)
2025-08-01 11:46:15 -07:00
geohot
494e951e90
nicer
2025-08-01 10:23:11 -07:00
geohot
c4410e91fd
rendering more
2025-08-01 09:18:14 -07:00
geohot
905019a4ec
test gemm passes
2025-08-01 09:07:44 -07:00
geohot
149c3f8fe9
test plus passes
2025-07-31 22:39:42 -07:00
geohot
5706e2d845
something
2025-07-31 22:37:26 -07:00
geohot
c66d2082c6
kerneless will replace kernel and lowerer
2025-07-31 21:46:20 -07:00
George Hotz and GitHub
8ff03806e8
add llama layers ( #11460 )
...
* add llama layers
* add contig bw for speed
2025-07-31 16:28:04 -07:00
qazal and GitHub
719827b95d
viz: add flops / mem bw to device programs ( #11459 )
...
* viz: add flops / mem bw to device programs
* better spacing style
2025-08-01 02:12:30 +03:00
chenyu and GitHub
3f742a5a7c
comma space lab models benchmark ( #11461 )
2025-07-31 19:06:18 -04:00
geohot
474ee9daa5
hotfix: add contiguous_backward to llama
2025-07-31 15:07:12 -07:00
qazal and GitHub
fa66d9772d
viz: show const node when it's root ( #11456 )
2025-08-01 01:01:58 +03:00
qazal and GitHub
056dabda5a
viz: refactor to color scheme ( #11455 )
2025-08-01 00:17:50 +03:00
nimlgen and GitHub
e5b6149dfb
more typing in drivers ( #11454 )
...
* more typing in drivers
* rm
2025-07-31 23:26:33 +03:00
qazal and GitHub
bad3cf5731
viz: add LLVM machine code analysis ( #11421 )
...
* start
* works everywhere
* add viz api
* utilization table
* reg pressure ui
* use llvm-mca
* llvm-mca ui
* work
* cleanup
* cycle through, defaults are enough
* x86 pending
* x86 nops
* get mcpu/mtriple from autogen
* cleanup server diff
* move parser to python
* normalize to pct of max
* segments legend
* imports
* also monospace
* max comes from the total per instruction
* base on the value
2025-08-01 01:59:26 +08:00
chenyu and GitHub
e847677e8a
use AxisType in search instead of colors ( #11452 )
2025-07-31 13:07:33 -04:00
nimlgen and GitHub
75c2c42def
suppress exceptions only during finalization ( #11451 )
...
* suppress exceptions only during finalization
* fix
* fix typing
* fix more warns
* fix
* better?
* Revert "better?"
This reverts commit a068aa5793 .
* mm?
* no as e
2025-07-31 13:57:12 +03:00
wozeparrot and GitHub
24dd0d52ed
feat: test remove to cpu ( #11444 )
2025-07-30 20:18:56 -07:00
c3cfcb50cb
Add linalg_det and test for torch backend ( #11405 )
...
* add linalg_det and test
* space
---------
Co-authored-by: chenyu <[email protected] >
2025-07-30 22:04:44 -04:00
cba3655de5
Add Test for Setitem ( #10559 )
...
* init
* update
* better
* failing test
* works
* Delete test file
* clean
* lint
* simplify variable name
* rm contigious, rm int dtype, and add assertEqual
---------
Co-authored-by: chenyu <[email protected] >
2025-07-30 22:03:41 -04:00
wozeparrot and GitHub
6252f7770e
feat: fake data ( #11447 )
2025-07-30 17:18:20 -07:00
chenyu and GitHub
e300451f3a
update llama3 ( #11446 )
...
`LR=1e-4 TRAIN_ON_VAL=1 DEFAULT_FLOAT=bfloat16 FUSE_ARANGE=1 JITBEAM=2 OPTIM_DTYPE=bfloat16 LLAMA3_SIZE=1B WARMUP_STEPS=36 DECAY_STEPS=360 SEQLEN=512 PYTHONPATH=. AMD=1 AMD_LLVM=0 MODEL=llama3 python3 examples/mlperf/model_train.py` trained to 7
2025-07-30 19:34:21 -04:00
wozeparrot and GitHub
5fb975351a
feat: flag for training on val ( #11441 )
2025-07-30 14:29:45 -07:00
chenyu and GitHub
4ca430e5bf
fix search dedup ( #11439 )
...
it should check against pre real_axis axis in actions, not real_axis.
2025-07-30 17:24:16 -04:00
wozeparrot and GitHub
d3da20eca6
feat: bump mlperf workflow timeout to 6 hours ( #11440 )
2025-07-30 14:12:12 -07:00
wozeparrot and GitHub
825b6a2505
feat: llama3 dataloader ( #11340 )
2025-07-30 13:27:55 -07:00
qazal and GitHub
af357b5dc8
disable TRACK_MATCH_STATS in BEAM workers [pr] ( #11437 )
2025-07-30 23:22:08 +03:00
George Hotz and GitHub
7c2d2eff86
check tensor core dims ( #11436 )
...
* check elements_per_thread in tensorcore [pr]
* check tc dims
2025-07-30 13:06:59 -07:00
nimlgen and GitHub
5fc5bb5237
ci: clear processes ( #11434 )
...
* unified hcq_smi for managment
* fix
* fix
* no reset for amd
2025-07-30 22:15:18 +03:00
George Hotz and GitHub
4f26a9ad32
check elements_per_thread in tensorcore [pr] ( #11435 )
2025-07-30 11:55:48 -07:00
nimlgen and GitHub
4b4ba5454c
ci: move driver start higher ( #11431 )
2025-07-30 10:48:38 +03:00
George Hotz and GitHub
1bef2d80c1
unrolls are all in the same scope ( #11429 )
...
* unrolls are all in the same scope
* fix that import
2025-07-29 16:55:37 -07:00
chenyu and GitHub
204da24cfc
increase driverbenchmark timeout-minutes to 15 ( #11428 )
2025-07-29 19:45:05 -04:00
chenyu and GitHub
d5fc6af4a2
remove unused ShapeTracker.consecutive [pr] ( #11426 )
2025-07-29 18:36:19 -04:00
George Hotz and GitHub
49a2583584
real new lowerer ( #11419 )
...
* real new lowerer
* fix group for reduce
* skip missing ranges
* fix wmma and unroll/contract
* real fix for wmma
* disable that test
* fix if gate
* simpler
* flash attention fusion works
* no end barriers
* still broken
* flash attention finally works
2025-07-29 15:35:51 -07:00
chenyu and GitHub
0e5d8d5c3c
remove tests that used .to_uop() ( #11425 )
...
* remove tests that used .to_uop()
* import
2025-07-29 15:52:16 -04:00
nimlgen and GitHub
c88e401d0e
ci: fix typos in h machine benchmarks ( #11423 )
2025-07-29 22:11:47 +03:00
chenyu and GitHub
90a5a312eb
simplify ShapeTracker in UOp.const [pr] ( #11424 )
2025-07-29 15:04:06 -04:00
chenyu and GitHub
398594029b
spec checks arg of VIEW are ShapeTracker ( #11422 )
2025-07-29 14:05:12 -04:00
geohot
1f1f99c287
hotfix: add DEBUG=3 to driver CI
2025-07-29 11:03:47 -07:00
George Hotz and GitHub
50fae54175
global local dims in gpudims [pr] ( #11420 )
2025-07-29 10:39:03 -07:00
chenyu and GitHub
9bc413f104
remove ShapeTracker.to_uop [pr] ( #11418 )
2025-07-29 13:29:37 -04:00
George Hotz and GitHub
ba2c4df125
dont render cast ptrs standalone ( #11417 )
...
* dont render cast ptrs standalone
* barrier cleanups
2025-07-29 09:24:26 -07:00
nimlgen and GitHub
d38d285489
ci: add h machines ( #11416 )
...
* ci: add h machines
* more
* fix names
* names not collide
* 20
* 10
2025-07-29 19:21:51 +03:00
2568bc0d99
ci: add caching for apt packages ( #11162 )
...
* add caching for apt packages
* remove 'inputs' from apt cache key, use outputs instead of env
* remove unnecessary mkdir for partial
---------
Co-authored-by: George Hotz <[email protected] >
2025-07-29 09:04:56 -07:00
George Hotz and GitHub
03909f2772
permute locals for HL uop matmul ( #11412 )
...
* permute locals for HL uop matmul
* parens fix that
* permutes
* 20 TFLOPS
2025-07-29 08:19:59 -07:00