nimlgen and GitHub
6b35220622
cpu hcq2 ( #17503 )
...
* cpu hcq2
* temp
* slop
* test with backpressure
* x
* x
* x
* x
* x
* x
* Dx
* save reverts
* um?
* x
* x
* call from py
* x?
* x
* submitters gone
* x
* x
* z
* Dx
* Dx
* x
* x
* fixes
* repl
* x
* f
* for now keep hcqbuffer
2026-08-14 15:06:03 +03:00
George Hotz and GitHub
b1859805b1
remove Ops.BIND ( #17511 )
...
* remove Ops.BIND
* param arg
* simplify that
* simplify
* cleaner
* props, not functions
* param and buffer can share
2026-08-13 23:52:16 -07:00
geohot
95ca5081fe
hotfix: update extra/runbook_digitalocean_mi350x
2026-08-13 17:26:39 -07:00
wozeparrot and GitHub
1b7f040984
fa: paas through window ( #17517 )
2026-08-13 14:45:57 +08:00
qazal and GitHub
39d144546e
fix mxfp4 mem estimate ( #17515 )
...
* add mem estimates
* rename
* move
2026-08-13 14:48:15 +09:00
qazal and GitHub
04c271ac41
simplify digitalocean_mi350x ( #17502 )
...
* simplify digitalocean_mi350x
* no hardcoded rocm path
2026-08-12 16:00:04 +09:00
George Hotz and GitHub
e11df72e0f
notes from digitalocean_mi350x ( #17494 )
...
* notes from digitalocean_mi350x
* cleanup
* revert non-doc changes on digitalocean_mi350x branch
2026-08-11 13:22:11 -07:00
nimlgen and GitHub
55e4f9d4f3
hcq2: proper unmap ( #17489 )
2026-08-11 15:41:28 +03:00
nimlgen and GitHub
e29606f07e
hcq2: copy kernel ( #17480 )
...
* hcq2: copy with kernel
* test
* x
2026-08-10 17:28:46 +03:00
qazal and GitHub
44f1f45cd5
llama: custom silu kernels ( #17462 )
...
* start by copying the C
* uop kernel
* cleanup tests
* estimates is part of SPEC
2026-08-10 16:43:01 +09:00
nimlgen and GitHub
8c8b43de62
hcq2: fix beam ( #17467 )
...
* fix beam
* x
2026-08-09 16:53:47 +03:00
nimlgen and GitHub
e17c21e102
hcq2: timings ( #17464 )
...
* hcq2: timings
* Dx
* x
* x
* x
* x
* align
* x
2026-08-08 22:00:32 +03:00
wozeparrot and GitHub
1827ec57f7
gptoss: fix sharded invalids ( #17450 )
2026-08-07 08:42:02 -07:00
wozeparrot and GitHub
1fd6b1035f
fa: swa support ( #17367 )
2026-08-06 08:07:30 -07:00
qazal and GitHub
9636dd1a25
test MXFP4 llama without hipcc ( #17435 )
...
* test MXFP4 llama without hipcc
* first pythonpath then dev
2026-08-06 17:31:40 +09:00
qazal and GitHub
f258708d7d
llama: custom quantize_mxfp4+transpose kernel (codex) ( #17434 )
...
* llama: custom quantize_mxfp4+transpose kernel (codex)
* rename to cpp
* inline
* cleanup
* lds load_bf16x4
* more tests, add Estimates
2026-08-06 16:13:28 +09:00
geohot
a8a8030bc9
add benchmark_llm script
2026-08-05 15:59:51 -07:00
nimlgen and GitHub
874d33128b
hcq2 benchmark ( #17235 )
...
* hcq2 in ci?
* fix
* traning
* x
* x
* x
* recover
* debug
* impler
* x
* x
* x
* hcq2: group input scatter plans by destination
* hcq2: simplify input scatter tables
* x
2026-08-05 10:00:42 +03:00
wozeparrot and GitHub
80d2073a11
fa: fix dq hazard with D=64 ( #17393 )
2026-08-04 09:38:01 -07:00
George Hotz and GitHub
c21a552f3d
llm: bugfixes + warmup ( #17384 )
2026-08-03 18:23:14 -07:00
wozeparrot and GitHub
3331944547
gptoss: fix moe routing ( #17377 )
2026-08-03 11:16:44 -07:00
nimlgen and GitHub
7c1ce50f63
hcq2: epoch ( #17376 )
...
* hcq2: epoch
* x
* minor
2026-08-03 16:37:25 +03:00
chenyu and GitHub
b502fc1367
more const arg -> val ( #17350 )
2026-08-01 02:18:10 -04:00
nimlgen and GitHub
1095bbe409
hcq2: fix ib reuse ( #17330 )
...
* hcq2: initialize IB reuse counters at link
* x
2026-07-31 18:43:23 +03:00
chenyu and GitHub
8e2f175542
const(dtype, b) -> const(b, dtype) [PR] ( #17328 )
...
prep for dtype removal
2026-07-31 09:46:37 -04:00
qazal and GitHub
f7964acb64
llama with MXFP4 ( #17321 )
...
* mxfp4 in llama
* less
* name
2026-07-31 18:25:25 +09:00
qazal and GitHub
0a3325f9c2
add mxfp4 quantize and layout kernels ( #17320 )
2026-07-31 14:50:18 +09:00
qazal and GitHub
a8c1e89500
fp4 asm gemm 6+ pflops ( #17315 )
...
* fp4 gemm
* better kernargs structure
* move to .s files
* work
* work
* p2
* style
* use .py
* move to dsl
* cleanup
* add MFMA_SCALE_X2_ENCODING
* cleanup mfma
* fma docs
* gemm_mxfp4
* more cleanup
* move
* move to cdna_asm_gemm
* change
* rm
* change
* mx
2026-07-31 14:25:23 +09:00
George Hotz and GitHub
d65ea465ed
cleanup gemm fragment + add store unshard ( #17313 )
...
* cleanup gemm fragment + add store unshard
* multi
* fix
2026-07-30 20:53:28 -07:00
wozeparrot and GitHub
b5a2a5666a
gptoss moe routing ( #17284 )
2026-07-30 07:49:14 -07:00
sirhcm and GitHub
060f447db6
qcom: match cl for SP_CS_INSTR_SIZE ( #17289 )
2026-07-29 18:47:02 -04:00
George Hotz and GitHub
138676ab81
improve fragment example + index unshard (kimi) ( #17288 )
...
* fix dtypes in fragment example
* match tilelang
* flip locals
* fix index on unshard
* test fixes
* kimi needs more taste
2026-07-29 15:38:38 -07:00
nimlgen and GitHub
d4ba8b6e0f
hcq2: use stack ( #17286 )
2026-07-29 22:09:34 +03:00
George Hotz and GitHub
b30c7e00d4
support 2d on UNSHARD (kimi) ( #17285 )
...
* support 2d on UNSHARD
* fixes
* Fix test and spec
* single barrier
* 2d sharding works for devices too
* cleanups
* no _rewrap
2026-07-29 12:01:59 -07:00
George Hotz and GitHub
52c9e5a99e
rename LOOP -> WEAK and STRONGLOOP -> LOOP ( #17283 )
2026-07-29 10:38:36 -07:00
George Hotz and GitHub
bd296a7359
enable alloc_fragment support with UNSHARD (kimi) ( #17272 )
...
* enable alloc_fragment support with UNSHARD (kimi)
* cleaner with implicit barrier
* cleanups
* cleaner
* strongloop
* dcount cleanups
2026-07-29 09:46:45 -07:00
George Hotz and GitHub
451120c6e1
make .barrier implicit (kimi) ( #17275 )
...
* make .barrier implicit (kimi)
* simplier
* lil
* remove tinygrad stock barriers
* readable
* lil
2026-07-28 22:34:57 -07:00
George Hotz and GitHub
57ae1bc7a7
rename MULTI to UNSHARD ( #17267 )
...
* rename MULTI to UNSHARD
* comment updates (glm)
* rename method to unshard
2026-07-28 16:51:41 -07:00
755dfb243b
rename CPU_COUNT to NUM_CPU_THREADS with cgroup awareness ( #17263 )
...
Rename CPU_COUNT to NUM_CPU_THREADS so it can be overridden via env var.
Default uses _get_cpu_count() which respects cgroup limits:
- os.process_cpu_count() on Python 3.13+
- /sys/fs/cgroup/cpu.max on cgroup v2
- /sys/fs/cgroup/cpu/cpu.cfs_quota_us on cgroup v1
- os.sched_getaffinity(0) fallback
Use NUM_CPU_THREADS.value in the dataloader instead of cpu_count(),
and update export_model.py and all renderer references.
Co-authored-by: teeny-runner <runner@teeny>
2026-07-28 15:33:57 -07:00
nimlgen and GitHub
97a2265362
hcq2: amd indirect ( #17220 )
...
* ind
* mock
2026-07-26 21:36:01 +03:00
chenyu and GitHub
a60b5f77ac
fix torch backend out= into a view ( #17210 )
2026-07-25 21:04:42 -04:00
nimlgen and GitHub
3946df787d
hcq2, cpu is hcq2-ish ( #17197 )
...
* m
* i
* x
* x
* Df
* x
* x
2026-07-26 02:38:15 +03:00
chenyu and GitHub
076b37e1ae
failing batch norm test ( #17206 )
...
* failing batch norm test
running stats does not schedule in training now since there's no reader
* not that
2026-07-25 18:46:49 -04:00
chenyu and GitHub
f902513355
derive torch backend dispatch from the aten schema ( #17203 )
2026-07-25 14:21:38 -04:00
chenyu and GitHub
983ad3bd95
fix torch backend batchnorm backward ( #17201 )
2026-07-25 13:15:09 -04:00
chenyu and GitHub
9f78504304
checked cast in torch backend unwrap ( #17199 )
2026-07-25 12:34:08 -04:00
chenyu and GitHub
8a10892f5a
fix torch backend as_strided ( #17195 )
...
0 means 0 offset
2026-07-25 02:24:17 -04:00
chenyu and GitHub
4c58b260fb
less wrong calculate_storage_offset ( #17194 )
...
initially for speed, then realized it's just wrong
2026-07-25 01:53:54 -04:00
wozeparrot and GitHub
c9b60caf8c
gptoss: moe gemm kernels ( #17178 )
2026-07-24 07:25:15 -07:00
nimlgen and GitHub
0ecef210bb
hcq2 cleanup 2 ( #17177 )
...
* hcq uops
* x
2026-07-24 13:04:21 +03:00