George Hotz and GitHub
c21a552f3d
llm: bugfixes + warmup ( #17384 )
2026-08-03 18:23:14 -07:00
chenyu and GitHub
67dc02d7e7
bitcast in python for _bits_to_rand [PR] ( #17383 )
...
* bitcast in python for _bits_to_rand [PR]
const in mixin would be weak only without width, so not bitcast
2026-08-03 21:14:58 -04:00
George Hotz and GitHub
3cb786f447
llm: update test_llm_server tests ( #17382 )
2026-08-03 16:06:17 -07:00
geohot
87289a7410
hotfix: revert test_scalar_alu_index, violates spec
2026-08-03 15:25:51 -07:00
George Hotz and GitHub
c2625c78cb
scalar ALU index fix + llm: preserve_thinking ( #17381 )
...
* cstyle: scalar ALU index fix, serve: preserve_thinking, test: fix Handler import
- cstyle.py: return scalar directly when ALU buffer has 1 element
- cli.py: add preserve_thinking param to FallbackTemplate.render
- serve.py: pass preserve_thinking=True when rendering chat completions
- test_llm_server.py: fix import to use Handler from llm.serve
* real fix
2026-08-03 15:10:21 -07:00
chenyu and GitHub
33755a3465
improve threefry codegen [pr] ( #17379 )
...
decomp uint64 can handle part of it
2026-08-03 17:24:25 -04:00
YassineYousfi and GitHub
104ee90ccf
usb: wait for PCIe link after power on ( #17380 )
2026-08-03 14:19:05 -07:00
chenyu and GitHub
a2385ae21d
MAX_LINE_COUNT=26000 ( #17378 )
...
oh well
2026-08-03 15:37:58 -04:00
wozeparrot and GitHub
3331944547
gptoss: fix moe routing ( #17377 )
2026-08-03 11:16:44 -07:00
nimlgen and GitHub
7c1ce50f63
hcq2: epoch ( #17376 )
...
* hcq2: epoch
* x
* minor
2026-08-03 16:37:25 +03:00
qazal and GitHub
be5f62d269
llama: refactor amax stuff and skip in fp4 ( #17375 )
2026-08-03 20:05:52 +09:00
nimlgen and GitHub
e22935c758
hcq2: inputs table ( #17374 )
...
* revert this
* x
* simpler
* fix
2026-08-03 13:40:43 +03:00
chenyu and GitHub
314df72b5f
deflake test_hcq with MOCKGPU ( #17370 )
...
for MOCKGPU we compare with e2e wall time which would be device agnostic
2026-08-03 13:22:59 +03:00
chenyu and GitHub
23c7813f44
no hard coded dtype int for rangeify debuf [PR] ( #17373 )
2026-08-02 20:43:04 -04:00
George Hotz and GitHub
05bc7c6994
fix llm reasoning and Linear import ( #17372 )
2026-08-02 12:31:52 -07:00
chenyu and GitHub
09dabfe05e
minor pm_float_decomp cleanup [PR] ( #17371 )
...
make the rule order independently correct
2026-08-02 15:17:33 -04:00
chenyu and GitHub
59df317b12
use UOp.const to create new consts [PR] ( #17368 )
...
replace arg won't work with ConstArg
2026-08-02 13:31:55 -04:00
wozeparrot and GitHub
e14cadb1fb
gptoss: set ASM_GEMM ( #17363 )
2026-08-02 06:52:36 -07:00
chenyu and GitHub
0258c7fefc
minor argstr and alloc stack cleanup [PR] ( #17361 )
2026-08-01 22:37:06 -04:00
George Hotz and GitHub
fb607fb990
faster devectorizer with one line ( #17360 )
2026-08-01 13:06:35 -07:00
George Hotz and GitHub
15c936db01
merge devectorize + indexing ( #17354 )
...
* external benchmark schedule in 5 sec (codex slop)
* prune
* real?
* delete
2026-08-01 11:40:15 -07:00
wozeparrot and GitHub
98b700bad1
gptoss optim fixes ( #17356 )
2026-08-01 09:10:06 -07:00
chenyu and GitHub
8e524ca467
CONST related cleanups [pr] ( #17352 )
...
ConstFloat(nan) != nan should be False, and some Invalid bool cleanups
2026-08-01 09:54:28 -04:00
nimlgen and GitHub
665822ab34
hcq2: faster replace ( #17353 )
...
* hcq2: use runtime device for submit params
* hcq2: parameterize buffers in linear time
2026-08-01 16:32:04 +03:00
chenyu and GitHub
6c0ec39279
UOp.is_invalid [PR] ( #17351 )
...
helper to prep ConstArg
2026-08-01 02:36:32 -04:00
chenyu and GitHub
b502fc1367
more const arg -> val ( #17350 )
2026-08-01 02:18:10 -04:00
qazal and GitHub
161783d8f7
add _device_num back to ast.variables (kimi) ( #17327 )
2026-08-01 14:56:50 +09:00
George Hotz and GitHub
5a1c641f79
more arg -> val ( #17349 )
...
* more arg -> val
* kimi
* more
2026-07-31 22:44:24 -07:00
Adeeb Shihadeh and GitHub
20b8ecff50
support new USB vendor ID ( #17348 )
2026-07-31 21:45:06 -07:00
George Hotz and GitHub
9082ecef5d
use .val to access the value of Ops.CONST ( #17347 )
2026-07-31 19:51:42 -07:00
George Hotz and GitHub
099d69ff7d
ci: split macos unit test into metal and mock runners ( #17346 )
2026-07-31 19:41:18 -07:00
George Hotz and GitHub
a88f832f0c
remove UOp.val ( #17345 )
2026-07-31 18:58:38 -07:00
chenyu and GitHub
850989115d
__int__ and __float__ work for weak ( #17342 )
2026-07-31 19:56:37 -04:00
sirhcm and GitHub
15d515299e
heuristics: try multiple TC axes ( #17341 )
2026-07-31 19:25:18 -04:00
sirhcm and GitHub
85ced44db6
tc: don't allow reduce over output dims ( #17340 )
2026-07-31 16:48:51 -04:00
chenyu and GitHub
277433259e
fix sym_infer for CAST ( #17338 )
2026-07-31 14:38:38 -04:00
chenyu and GitHub
8dc225e28e
dtype_from_uop(INS) is None [PR] ( #17337 )
2026-07-31 14:33:15 -04:00
chenyu and GitHub
ad24750487
const(value, dtype) -> const(value).cast(dtype) in tests ( #17335 )
2026-07-31 13:24:31 -04:00
qazal and GitHub
a11ee26bb8
viz: prep for faster cli DEBUG=3 ( #17334 )
...
* move data
* split
2026-08-01 02:16:30 +09:00
sirhcm and GitHub
b95bd5b2a5
don't reset chestnut in benchmark ( #17333 )
2026-07-31 12:39:04 -04:00
nimlgen and GitHub
1095bbe409
hcq2: fix ib reuse ( #17330 )
...
* hcq2: initialize IB reuse counters at link
* x
2026-07-31 18:43:23 +03:00
chenyu and GitHub
7f4dbb8090
remove shape= from UOp.const [PR] ( #17331 )
...
inlined to const_like
2026-07-31 11:18:59 -04:00
chenyu and GitHub
8e2f175542
const(dtype, b) -> const(b, dtype) [PR] ( #17328 )
...
prep for dtype removal
2026-07-31 09:46:37 -04:00
nimlgen and GitHub
155b84ee80
hcq2 faster schedule ( #17324 )
...
* avoid quadratic STACK dtype promotion
* build HCQ patch stacks directly
* pack HCQ command buffers linearly
* remove HCQ command buffer simplification
2026-07-31 16:08:18 +03:00
wozeparrot and GitHub
4f5cadd15d
gptoss ci ( #17325 )
2026-07-31 05:51:58 -07:00
nimlgen and GitHub
0c4bfaeb48
coalesce ints ( #17323 )
...
* merge ints
* fix z3 validation of coalesced loads
* fix uint vector names in CUDA and Metal
2026-07-31 15:49:56 +03:00
qazal and GitHub
6d2700f0b7
failing test for unbound _device_num err in BEAM ( #17326 )
...
* min failing test
* switch to cpu
* err
2026-07-31 20:24:44 +09:00
wozeparrot and GitHub
93b74c75fc
gptoss: grouped moe ( #17322 )
2026-07-31 03:25:32 -07:00
qazal and GitHub
f7964acb64
llama with MXFP4 ( #17321 )
...
* mxfp4 in llama
* less
* name
2026-07-31 18:25:25 +09:00
qazal and GitHub
0a3325f9c2
add mxfp4 quantize and layout kernels ( #17320 )
2026-07-31 14:50:18 +09:00