geohot
7b19732a7a
gpt fixes
2026-07-17 00:55:50 +00:00
geohot
ee19fd0b6a
fixes
2026-07-16 17:45:26 -07:00
geohot
ce5ae31f5d
work
2026-07-16 17:30:14 -07:00
geohot
22bdc7b6c2
rm that
2026-07-16 17:18:34 -07:00
geohot
e13e6b3752
add tool calling support to llm
2026-07-16 17:09:35 -07:00
chenyu and GitHub
88826a6f35
no weak dtype for randn_like either ( #17055 )
2026-07-16 18:29:12 -04:00
chenyu and GitHub
3bfd62e915
fix 0 size tolist to match numpy ( #17054 )
2026-07-16 17:42:52 -04:00
nimlgen and GitHub
709babb97c
system: remove sibling functions of PCIDevice ( #17052 )
2026-07-17 00:15:32 +03:00
George Hotz and GitHub
d8b83daac6
set tc_upcast_axes to None when done with it ( #17053 )
...
* set tc_upcast_axes to None when done with it
* no tag needed
2026-07-16 14:15:21 -07:00
stylishvoid and GitHub
c74149c973
avoid repeated parsing and toposort in _valid_priority [PR] ( #17049 )
...
* avoid repeated parsing and toposort in _valid_priority
* use backward_slice_with_self instead
2026-07-16 16:24:08 -04:00
chenyu and GitHub
6fa0b2b19e
materialize weak dtype casts to default ( #17051 )
...
in clone and _buffer
2026-07-16 16:12:33 -04:00
George Hotz and GitHub
4d8c3d3fc9
add test_hgemm to test_tiny ( #17050 )
...
* add test_hgemm to test_tiny
* dsp skip
2026-07-16 13:12:10 -07:00
George Hotz and GitHub
2b1146b3f4
further clean up wmma ( #17048 )
...
* further clean up wmma
* comment
2026-07-16 11:43:23 -07:00
chenyu and GitHub
f6a92d0a16
sum_acc_dtype(weak) is weak ( #17047 )
...
also no explicit weak for rand
2026-07-16 14:32:37 -04:00
George Hotz and GitHub
61e104bdfb
use UOp.wmma everywhere ( #17045 )
...
* use UOp.wmma everywhere
* fix
2026-07-16 10:40:48 -07:00
chenyu and GitHub
5a4156c5d1
bitcast and element_size raise for weak dtypes ( #17046 )
2026-07-16 13:07:45 -04:00
nimlgen and GitHub
7eb197b1bb
nv: always wait for reset ( #17043 )
...
* nv: always wait for reset
* x
2026-07-16 16:35:12 +03:00
chenyu and GitHub
dba8b6b505
allow weak alu operands ( #17044 )
2026-07-16 09:33:20 -04:00
nimlgen and GitHub
e33e96415f
hcq2: tiny cleanupg ( #17042 )
2026-07-16 16:14:54 +03:00
810d8732f9
fix n^2 in limit_bufs by memoizing reachable loads [PR] ( #17017 )
...
* fix n^2 in limit_bufs by memoizing reachable loads [pr]
* Update test_schedule.py
---------
Co-authored-by: Jacob Kitchen <[email protected] >
2026-07-15 23:54:04 -07:00
1c74e044a4
search /usr/lib/wsl/lib first for linux ( #17027 )
...
Co-authored-by: George Hotz <[email protected] >
2026-07-15 23:28:17 -07:00
George Hotz and GitHub
8b0dd870ce
use wmma helper ( #17038 )
2026-07-15 23:25:17 -07:00
qazal and GitHub
783042d216
viz: graph stays in place when sidebars resize ( #17037 )
...
* viz: sidebars can resize independent of main graph
* both sidebars
* fix device-list
* more work
* no variables
* raw 15%
* fix custom view
* minor detail
2026-07-16 11:47:22 +09:00
chenyu and GitHub
e8d3047a50
dtype_from_uop cleanup [PR] ( #17036 )
2026-07-15 21:52:21 -04:00
chenyu and GitHub
6b7fee7d9f
minor lower_alu_dtype cleanup [PR] ( #17034 )
2026-07-15 17:42:23 -04:00
chenyu and GitHub
be075b200a
weak dtypes in dtype_from_uop [PR] ( #17032 )
...
* weak dtypes in dtype_from_uop [PR]
* no weak in spec_program
* weak const fold tests
2026-07-15 16:54:31 -04:00
chenyu and GitHub
3ffb4dc4bc
unify lower index in lower_alu_dtype [PR] ( #17033 )
...
will work for weak types too
2026-07-15 16:38:42 -04:00
nimlgen and GitHub
d6fddb066f
usb: keep only custom ( #17029 )
...
* usb: keep only custom
* mockgpu by gpt
* gpt said sorry
* revert
* reset
* fix
* flash
2026-07-15 22:40:52 +03:00
wozeparrot and GitHub
0d30f97584
mlperf: make v6.1 dir ( #17031 )
2026-07-15 10:40:55 -07:00
chenyu and GitHub
c23d8188e1
remove _ensure_float [pr] ( #17030 )
...
do this cast late. allow `SQRT(int)`
2026-07-15 11:19:20 -04:00
chenyu and GitHub
0d19970edc
least_upper_dtype in dtype_from_uop [PR] ( #17028 )
2026-07-15 09:27:53 -04:00
chenyu and GitHub
ebe26420a7
update where Invalid rules [pr] ( #17026 )
...
fixed TestInvalidTensor.test_tensor_index
2026-07-15 00:00:30 -04:00
wozeparrot and GitHub
06169f5013
gptoss: small fixes ( #17025 )
2026-07-14 20:40:23 -07:00
chenyu and GitHub
47ddf94f17
remove InvalidType lt and gt ( #17023 )
...
not really used
2026-07-14 21:59:15 -04:00
sirhcm and GitHub
c9baa2ef79
use pattern matcher in contiguous_view_offset [PR] ( #17022 )
2026-07-14 19:37:36 -04:00
nimlgen and GitHub
4257939e50
remove copyin/copyout from Buffer ( #17020 )
...
* remove copyin/copyout from Buffer
* x
* x
* x
* x
2026-07-14 19:47:22 +03:00
qazal and GitHub
939f28d571
fused qkv rope custom kernel ( #17021 )
...
* work
* fused qkv_norm
* work
* speed
* not that yet
* test cleanup
* just clone
* remove .realize()
* cleanup tests
2026-07-15 01:08:42 +09:00
chenyu and GitHub
82fbca43c5
fix Tensor(np) dtype and support fp8 safetensor ( #17019 )
2026-07-14 09:31:21 -04:00
chenyu and GitHub
872225e47d
update dtype tests for small dtypes ( #17016 )
2026-07-14 08:07:00 -04:00
qazal and GitHub
edfef062ed
skip viz.cli -t in null device ( #17018 )
2026-07-14 19:24:58 +09:00
chenyu and GitHub
55bb251130
add pm_manual_bf16_cast to Metal [pr] ( #17015 )
...
mitigate metal compiler bug for
`as_type<half>( (bfloat)(const) )`
2026-07-13 21:53:47 -04:00
sirhcm and GitHub
a9fbc7db7b
expect _offset support, CL and WEBGPU are outliers ( #17014 )
2026-07-13 18:53:32 -04:00
chenyu and GitHub
9ce96c2628
fix subnormal in test_dtype ( #17013 )
...
* fix subnormal in test_dtype
should fix flaky test/backend/test_dtype.py::TestFp8e4m3::test_casts_from
* better
2026-07-13 18:53:13 -04:00
chenyu and GitHub
c898dfe150
remove UOp.contiguous override [PR] ( #17012 )
...
also cleaned up max_shard_shape
2026-07-13 16:10:58 -04:00
chenyu and GitHub
681a5e0cfd
remove UOp cast and bitcast override [PR] ( #17011 )
2026-07-13 14:23:11 -04:00
chenyu and GitHub
0410c9325d
make test/null follow the SPEC ( #17010 )
2026-07-13 14:01:41 -04:00
nimlgen and GitHub
e4bdc529c4
hcq2 ci ( #17008 )
...
* hcq2 ci
* x
2026-07-13 19:29:08 +03:00
nimlgen and GitHub
4536a57f79
hcq rename map ( #17009 )
...
* hcq rename map
* x
2026-07-13 19:23:12 +03:00
qazal and GitHub
62ad646d1c
llama: gemm/fa backward speedups (gpt 5.6) ( #17007 )
...
* fp8 atb gemm speedup
* work
* revert
* fa bw faster
2026-07-14 00:27:00 +09:00
nimlgen and GitHub
4d2becddf8
hcq2: spec=2 ( #17006 )
...
* hcq2: spec=2
* hcq: isolate HCQ spec rules
* chq
* move
2026-07-13 18:13:33 +03:00