qazal and GitHub
a1263fadf3
fused_qkv_rope in UOp try 2 ( #17619 )
...
* fused_qkv_rope in UOp try 2
* dont need that
* less
2026-08-20 12:28:34 +09:00
chenyu and GitHub
e6324d1e1c
test updates from weak const branch ( #17618 )
2026-08-19 22:56:25 -04:00
sirhcm and GitHub
2067133732
cpu: link with rt ( #17608 )
2026-08-19 17:55:08 -04:00
chenyu and GitHub
fc214da417
test updates for weak const change ( #17606 )
2026-08-19 15:56:01 -04:00
chenyu and GitHub
b8cc74ecf8
no float in tensor shape [pr] ( #17605 )
2026-08-19 15:40:16 -04:00
7064e76bc8
fix roll on zero-sized tensors ( #17603 )
...
Signed-off-by: Bennett <[email protected] >
Co-authored-by: Bennett <[email protected] >
2026-08-19 15:31:21 -04:00
chenyu and GitHub
a4fadcf606
fix TestDevCopySpeeds SIZE ( #17602 )
...
SIZE should be int
2026-08-19 14:53:17 -04:00
chenyu and GitHub
c218b4842d
fold_bitcast should truncate its input [pr] ( #17601 )
2026-08-19 14:39:36 -04:00
chenyu and GitHub
bd6e70ac15
delete stale tests ( #17596 )
2026-08-19 11:18:05 -04:00
chenyu and GitHub
9550378704
finish casted_consts migration [PR] ( #17595 )
2026-08-19 10:43:55 -04:00
chenyu and GitHub
b3e2f17b24
update NULL tests that depends on strong dtype CONST ( #17594 )
2026-08-19 10:26:38 -04:00
chenyu and GitHub
e8ba214b56
casted CONST migration for x86 [pr] ( #17592 )
...
* casted CONST migration for x86 [pr]
* style
2026-08-19 09:38:09 -04:00
chenyu and GitHub
68b4407fe3
casted CONST migration for cstyle [pr] ( #17587 )
2026-08-19 09:01:40 -04:00
qazal and GitHub
d539aaf752
Revert "fused_qkv_rope in UOp ( #17591 )" ( #17593 )
...
This reverts commit 8c2bf02d17 .
2026-08-19 21:42:55 +09:00
qazal and GitHub
8c2bf02d17
fused_qkv_rope in UOp ( #17591 )
...
* llama: 4% faster fused_qkv_rope
* prep
* add uop kernel, has_hipcc is cached
* less
2026-08-19 18:07:30 +09:00
George Hotz and GitHub
c31038ff37
use KernelCountException when kernel count is being compared ( #17584 )
2026-08-18 16:06:03 -07:00
chenyu and GitHub
49778d9a48
start renderer casted const migration [pr] ( #17582 )
...
before rendering, rewrite strong typed const to casted weak const and have renderer adopt the new UOp. starting with PYTHON
2026-08-18 17:43:27 -04:00
chenyu and GitHub
a1366e2f6c
alu(long, weakint) can do math in int too [pr] ( #17579 )
...
* alu(long, weakint) can do math in int too [pr]
* remove
2026-08-18 09:08:35 -04:00
George Hotz and GitHub
8d2cc64b69
llm: refactor delta attention ( #17564 )
...
* refactor delta attention
* cleanups
* bugfixes
* stack
* recurrent w chunk_size 1
* revert that
* extra test
2026-08-17 19:24:03 -07:00
chenyu and GitHub
34c9b9d434
add back cast where rule [pr] ( #17572 )
2026-08-17 15:27:20 -04:00
b1tg and GitHub
b757437f64
llm: respect expert_gating_func ( #17458 )
...
* llm: respect expert_gating_func
* test
* enum
* clean
2026-08-17 12:19:33 -07:00
b1tg and GitHub
2776c5b369
fix call arg indexing in shard scheduling ( #17519 )
2026-08-17 09:57:50 -07:00
nimlgen and GitHub
58edff61d9
hcq2: one submitter ( #17556 )
...
* hcq2: c submitter
* x
* x
* x
* simpler
* simpler
* x
* x
* Dx
* revrt
* Dx
* x
* fst
* fix
2026-08-17 16:08:19 +03:00
chenyu and GitHub
42714e1399
update a few is CONST check to check device None [pr] ( #17563 )
...
* update a few is CONST check to check device None [pr]
* clone
2026-08-17 07:33:16 -04:00
chenyu and GitHub
138fb4a783
delete dead DType.scalar [PR] ( #17561 )
2026-08-16 21:12:17 -04:00
chenyu and GitHub
057a18a07c
fix emulated long cast to double ( #17559 )
2026-08-16 20:52:42 -04:00
George Hotz and GitHub
e688e07758
add max_shape/max_numel to mixins + pad_to ( #17553 )
2026-08-15 20:05:10 -07:00
nimlgen and GitHub
97022960ae
device: fix remap ( #17549 )
2026-08-16 01:01:36 +03:00
chenyu and GitHub
417563ca20
fix webgpu is_nan [pr] ( #17551 )
2026-08-15 16:31:19 -04:00
chenyu and GitHub
5ca87f1bac
fix cast to weak twice [pr] ( #17548 )
...
also no gradient for weak target
2026-08-15 12:39:54 -04:00
chenyu and GitHub
26cbadd69a
no pm_fold_cast_const in full_rewrite_to_sink [PR] ( #17547 )
2026-08-15 10:18:59 -04:00
chenyu and GitHub
fae893753b
no pm_fold_cast_const in get_kernel_graph [pr] ( #17546 )
...
* no pm_fold_cast_const in get_kernel_graph [pr]
* maybe
2026-08-15 09:57:05 -04:00
chenyu and GitHub
539a03343a
no casted const from sub and div [pr] ( #17543 )
2026-08-15 07:47:40 -04:00
qazal and GitHub
a57569349c
renumber invalids before callify ( #17542 )
...
* renumber invalids before callify
* change
* Revert "change"
This reverts commit 6f4df1541e79721a85ee3f5801114454f264c973.
* renumber in tensor
* scope renumber_invalid_outputs
* cleanup
2026-08-15 18:00:54 +09:00
qazal and GitHub
5c43a89fb1
precompile_backward tests for sched_cache ( #17544 )
...
* work
* back
* work
* keep +
2026-08-15 15:25:31 +09:00
qazal and GitHub
e6f5bb9c09
simple test for Invalid clone cache miss regression ( #17541 )
...
* simple test for Invalid clone cache miss regression
* xfail
* _
2026-08-15 11:00:21 +09:00
chenyu and GitHub
4b0525e594
no pm_fold_cast_const in UOp.simplify and hcq2 [pr] ( #17540 )
2026-08-14 21:39:47 -04:00
chenyu and GitHub
13c381b0c0
remove pm_fold_cast_const from initial symbolic [pr] ( #17535 )
...
interestingly it gives more accurate numerics when composing const like log10
2026-08-14 14:14:40 -04:00
chenyu and GitHub
ac7067ac60
fix deconstruct_function for python 3.11 ( #17534 )
2026-08-14 13:38:34 -04:00
chenyu and GitHub
80169c6758
remove where push cast to branches from sym [pr] ( #17533 )
...
* remove where push cast to branches from sym [pr]
not really needed and one less place that generates casted weak const when it's not needed
* fix
2026-08-14 12:57:56 -04:00
chenyu and GitHub
89ab344c42
fix assign into bitcast with no explicit realize ( #17531 )
2026-08-14 09:23:01 -04:00
nimlgen and GitHub
6b35220622
cpu hcq2 ( #17503 )
...
* cpu hcq2
* temp
* slop
* test with backpressure
* x
* x
* x
* x
* x
* x
* Dx
* save reverts
* um?
* x
* x
* call from py
* x?
* x
* submitters gone
* x
* x
* z
* Dx
* Dx
* x
* x
* fixes
* repl
* x
* f
* for now keep hcqbuffer
2026-08-14 15:06:03 +03:00
George Hotz and GitHub
b1859805b1
remove Ops.BIND ( #17511 )
...
* remove Ops.BIND
* param arg
* simplify that
* simplify
* cleaner
* props, not functions
* param and buffer can share
2026-08-13 23:52:16 -07:00
qazal and GitHub
faba071b1d
don't enter CALL body in assign fixups ( #17527 )
...
* fix python time regression in mxfp4
* this saves even more time
* s_nop test
* itertools count
* cleanup
2026-08-14 15:05:04 +09:00
qazal and GitHub
16c5ff2490
add external_benchmark_all2all.py ( #17507 )
...
* add external_benchmark_all2all.py
* mv
* more minimal
* less
* fix space
2026-08-13 15:50:43 +09:00
George Hotz and GitHub
e103fb2a10
more lil llm improvements ( #17514 )
...
* more lil llm improvements
* default float
2026-08-12 20:13:20 -07:00
George Hotz and GitHub
2297118541
lil llm improvements ( #17513 )
2026-08-12 19:29:32 -07:00
geohot
ff0cb28c21
skip slow whisper tests
2026-08-12 13:12:37 -07:00
qazal and GitHub
ed8297a102
kerenl opts test from nan in llama 8b ( #17510 )
...
* all2all
* nan
* remove that
* less
* has_local
* only the nan change here
* use nice getitem syntax for INDEX
* work
* remove
* even simpler
2026-08-13 04:07:30 +09:00
Raine and GitHub
de04781b36
simplify equivalent const max ( #17505 )
...
* add const max folds
* add regression test
* move
2026-08-12 08:39:29 -07:00