qazal and GitHub
05e4727feb
arange stack regression test ( #17250 )
...
* simple failing test
* commend out stack
* a little smaller
* work
2026-07-28 20:00:22 +09:00
qazal and GitHub
e3b3eea1e2
test for extra copy in allreduce_cast ( #16740 )
...
* test for extra copy in allreduce_cast
* simpler
* cleanup
2026-07-28 19:02:54 +09:00
chenyu and GitHub
1380d6cc5a
pm_long_decomp doesn't depend on operand dtype [pr] ( #17249 )
...
* pm_long_decomp doesn't depend on operand dtype [pr]
the const the rules returned won't have strong dtype
* fix
2026-07-28 01:15:04 -04:00
chenyu and GitHub
4b7022e8f4
Revert "64-bit UOp.variable support ( #17246 )" ( #17248 )
...
This reverts commit ab8fb191b2 .
2026-07-28 00:00:26 -04:00
sirhcm and GitHub
ab8fb191b2
64-bit UOp.variable support ( #17246 )
2026-07-27 23:56:09 -04:00
chenyu and GitHub
f837ca3587
don't match const dtype in UPat [pr] ( #17247 )
2026-07-27 22:50:08 -04:00
chenyu and GitHub
654d475d37
update pm_long_decomp [pr] ( #17245 )
...
* update pm_long_decomp [pr]
instead of reading dtype from UOp (which might be const that's going to be weak), just pass output dtype in tag
* cleanup tag
2026-07-27 22:28:50 -04:00
chenyu and GitHub
550225f603
clean up pm_lower_weak [pr] ( #17243 )
2026-07-27 19:11:39 -04:00
chenyu and GitHub
0bb36c9989
make Tensor(None) weakfloat ( #17241 )
...
match other float consts
2026-07-27 16:45:28 -04:00
chenyu and GitHub
896afad9bf
don't match strong typed const in UPat [pr] ( #17240 )
2026-07-27 16:19:01 -04:00
nimlgen and GitHub
8b9ef157d1
run_linear in external_test_gpu_crash ( #17239 )
2026-07-27 22:00:11 +03:00
b1tg and GitHub
bdbb1d702f
fix shard axis through symbolic reshape ( #17238 )
...
* fix shard axis through symbolic reshape
* bind
2026-07-27 11:32:15 -04:00
nimlgen and GitHub
8eaeede96d
hcq2: do not cache beam ( #17234 )
2026-07-27 17:16:24 +03:00
chenyu and GitHub
818a892ebc
flip from_py to use weak dtypes [pr] ( #17229 )
2026-07-27 10:07:00 -04:00
qazal and GitHub
a3ca1e55a5
polish the viz readme ( #17236 )
...
* polish the viz readme
* style
* no epilog=
2026-07-27 22:03:13 +09:00
wozeparrot and GitHub
056974468e
gptoss: split no-wd params ( #17233 )
2026-07-27 02:58:43 -07:00
qazal and GitHub
19c4d736f2
validate json output of viz.cli in CI ( #17232 )
...
* validate viz.cli --json always prints valid JSON
* highest debug level
* jq empty we don't need a print
* gate that import
2026-07-27 15:44:24 +09:00
chenyu and GitHub
f45fc4c566
test updates from weak flip ( #17231 )
2026-07-27 01:27:42 -04:00
qazal and GitHub
95e3b0066f
webgpu failing test for duplicate PARAM in CALL [pr] ( #17230 )
...
* webgpu failing test for duplicate PARAM in CALL [pr]
* typo
2026-07-27 14:06:06 +09:00
chenyu and GitHub
0f98212e80
skip test_float_to_fp8e4m3_extreme_values ( #17228 )
...
fp8 overflow behavior changed in torch 2.13.0, skipped the test for now
2026-07-26 21:01:51 -04:00
chenyu and GitHub
456b5b5060
fix python_alu inf ( #17227 )
2026-07-26 20:04:12 -04:00
chenyu and GitHub
afb25a624c
simpler minimum and copysign [PR] ( #17226 )
2026-07-26 19:54:22 -04:00
nimlgen and GitHub
165f0626f8
speedy hcq2 ( #17225 )
...
* faster
* x
2026-07-27 01:31:56 +03:00
nimlgen and GitHub
97a2265362
hcq2: amd indirect ( #17220 )
...
* ind
* mock
2026-07-26 21:36:01 +03:00
chenyu and GitHub
4b6760539b
fix _prepare_jit_inputs for weak [pr] ( #17222 )
2026-07-26 14:03:27 -04:00
chenyu and GitHub
94dad3d261
clean up mixin cos and exp [PR] ( #17221 )
2026-07-26 13:13:05 -04:00
chenyu and GitHub
a8d51097dc
realize weak is no-op [pr] ( #17219 )
...
None device and weak dtype are both virtual
2026-07-26 11:52:56 -04:00
79c07a334c
fix ValueError in UOp.axis for shard reshape crossing boundary ( #16547 )
...
Co-authored-by: George Hotz <[email protected] >
2026-07-26 08:38:10 -07:00
C T and GitHub
d70a134845
fix nvrtc_check helper used for jitlink call ( #16362 )
2026-07-26 08:20:39 -07:00
George Hotz and GitHub
960430a5e5
Revert "nv: set lower interleave level to reduce GPU hogging (ai slop) ( #15518 )" ( #17218 )
...
This reverts commit ac12914506 .
2026-07-26 08:15:16 -07:00
ac12914506
nv: set lower interleave level to reduce GPU hogging (ai slop) ( #15518 )
...
Fixes #10773 . The NV backend makes desktop systems unusable (cursor lag,
video drops) because the channel group runs at the default HIGH interleave
level, monopolizing the GPU and starving the display compositor.
Sets tsgInterleaveLevel to LOW (0) by default so the GPU scheduler can
preempt compute work for display refresh. Dedicated compute machines can
restore full priority with NV_INTERLEAVE=2.
Also adds SET_INTERLEAVE_LEVEL to the mock GPU driver's pass-through list.
Co-authored-by: Yasko C <[email protected] >
Co-authored-by: Claude Opus 4.6 (1M context) <[email protected] >
2026-07-26 08:12:05 -07:00
chenyu and GitHub
5d1aa84901
SHR/SHL are Broadcastable [pr] ( #17216 )
2026-07-26 02:11:03 -04:00
chenyu and GitHub
acc2374b6f
minor normalize cleanup [pr] ( #17215 )
...
and add a few no-op clones in test_ops to prep for weak flip
2026-07-26 00:39:53 -04:00
chenyu and GitHub
a96974e70c
slight weak behavior tweak and cleanups [pr] ( #17214 )
...
* slight weak behavior tweak and cleanups [pr]
* ruff
2026-07-25 22:51:24 -04:00
chenyu and GitHub
eb889053bf
weak frompy prerequisite [PR] ( #17213 )
2026-07-25 22:35:45 -04:00
chenyu and GitHub
a60b5f77ac
fix torch backend out= into a view ( #17210 )
2026-07-25 21:04:42 -04:00
nimlgen and GitHub
3946df787d
hcq2, cpu is hcq2-ish ( #17197 )
...
* m
* i
* x
* x
* Df
* x
* x
2026-07-26 02:38:15 +03:00
chenyu and GitHub
076b37e1ae
failing batch norm test ( #17206 )
...
* failing batch norm test
running stats does not schedule in training now since there's no reader
* not that
2026-07-25 18:46:49 -04:00
chenyu and GitHub
492dc6d5fb
delete _index_to_concrete_int [pr] ( #17205 )
...
staying weak is okay
2026-07-25 16:05:08 -04:00
chenyu and GitHub
74c2121d99
promo (uint64, int) -> weakfloat like JAX [pr] ( #17204 )
2026-07-25 15:20:42 -04:00
chenyu and GitHub
f902513355
derive torch backend dispatch from the aten schema ( #17203 )
2026-07-25 14:21:38 -04:00
chenyu and GitHub
ee2ccb1f24
put weakint in dtypes.weaks [pr] ( #17202 )
...
* put weakint in dtypes.weaks [pr]
* custom_add_var
2026-07-25 14:00:34 -04:00
wozeparrot and GitHub
f0117e98df
refactor mlperf optim ( #17200 )
2026-07-25 10:37:11 -07:00
chenyu and GitHub
983ad3bd95
fix torch backend batchnorm backward ( #17201 )
2026-07-25 13:15:09 -04:00
chenyu and GitHub
9f78504304
checked cast in torch backend unwrap ( #17199 )
2026-07-25 12:34:08 -04:00
qazal and GitHub
732e6bd52c
add one line viz mention ( #17198 )
...
* add one line viz mention
* changes
* edit
2026-07-26 00:18:41 +09:00
chenyu and GitHub
8a10892f5a
fix torch backend as_strided ( #17195 )
...
0 means 0 offset
2026-07-25 02:24:17 -04:00
chenyu and GitHub
4c58b260fb
less wrong calculate_storage_offset ( #17194 )
...
initially for speed, then realized it's just wrong
2026-07-25 01:53:54 -04:00
chenyu and GitHub
d923263a5b
simpler lower_weakint_node [PR] ( #17193 )
...
works since weakint is in promo lattice properly now
2026-07-25 00:31:04 -04:00
sirhcm and GitHub
9fdaa4bff2
standardize Program class ( #17189 )
2026-07-25 00:17:38 -04:00