geohot
048f510b51
cleanups
2026-07-21 18:29:13 -07:00
geohot
cacba3f4d5
cleanups
2026-07-21 18:11:20 -07:00
geohot
01c6f396b1
upd
2026-07-21 17:36:05 -07:00
geohot
d51003bb61
LOOP is srcless RANGE (kimi)
2026-07-21 17:18:17 -07:00
chenyu and GitHub
92f9c850b4
fix pow(int, float) ( #17126 )
...
* fix pow(int, float)
* onnx
2026-07-21 18:48:24 -04:00
chenyu and GitHub
b1060ca708
don't promote dtype in _pad_constant [pr] ( #17125 )
2026-07-21 18:13:50 -04:00
chenyu and GitHub
f19a2ad771
single where mixin [pr] ( #17118 )
...
* single where mixin [pr]
no shape broadcasting in ufix and _broadcasted anymore
* QCOM vectorized bool is broken
2026-07-21 17:38:36 -04:00
Armand du Parc Locmaria and GitHub
ef37830d13
allow freeing buffers when pickling/unpickling ( #16799 )
...
* allow pickling out of band buffers
* also need to release when loading
* test peak ram
* lint
* sync before yielding next buffer for backends with async copy in
* skip on mock devices
* reason
* or always bytearray, always free?
* Revert "or always bytearray, always free?"
This reverts commit a017bb68742985a5b7431e0b4e973c2997c92b6a.
* one less copy
2026-07-21 16:11:23 -04:00
b764599d87
add Ops.LOOP + conditional Ops.END (kimi) ( #17117 )
...
* add Ops.LOOP + conditional Ops.END (kimi)
* c
* x
---------
Co-authored-by: George Hotz <[email protected] >
2026-07-21 22:54:14 +03:00
chenyu and GitHub
46b82d4755
don't auto cast cond for WHERE ( #17115 )
...
no or_casted all WHEREs with single mixin, matched torch
2026-07-21 13:00:07 -04:00
chenyu and GitHub
76dade5a11
implicit broadcast gradient based on shape only [pr] ( #17114 )
...
fixed gradient for shape () UOp, enabled unify WHERE mixin
2026-07-21 12:51:54 -04:00
chenyu and GitHub
f64f96ec59
broadcast_axes [PR] ( #17112 )
...
prerequisite to simplify broadcasting logic and make it implicit
2026-07-21 11:53:12 -04:00
sirhcm and GitHub
f3a5337825
correct spelling of coalesce ( #17103 )
2026-07-20 21:44:04 -04:00
chenyu and GitHub
95f5c85bf3
some realize and corealize for slow tests ( #17106 )
2026-07-20 21:43:23 -04:00
chenyu and GitHub
13ca9bd8a6
remove dtypes.index again ( #17104 )
...
also reverted some dtype change, the split made things needlessly complicated
2026-07-20 20:30:04 -04:00
sirhcm and GitHub
980748ccfc
add multiple_of to ParamArg ( #17101 )
2026-07-20 20:11:54 -04:00
George Hotz and GitHub
f7ce7f330d
llm: minor fixes + tests ( #17099 )
...
* llm: minor fixes + tests
* error
2026-07-20 14:31:20 -07:00
chenyu and GitHub
4b8db13e01
rdna int8 wmma ( #17098 )
...
nice to fix _wmma_name, also more generic tests
2026-07-20 17:00:50 -04:00
Pol Puigdemont Plana and GitHub
ef77963cfd
derivative of logsumexp is independent of max ( #17088 )
...
same as #7009 but for logsumexp and logcumsumexp.
fwd+bwd kernel count 5 -> 3 for both. gradients unchanged
(ties, -inf masks, torch-compared at grad_atol=1e-7).
2026-07-20 06:52:16 -07:00
qazal and GitHub
1cf8f2f68c
llama: inplace amax update ( #17064 )
...
* llama: inplace amax update
* remove amax_out return
* work
* fit
* work
* work
* keep
* diff cleanup
2026-07-20 15:05:41 +09:00
chenyu and GitHub
ac3f56a1a2
more shift tests ( #17083 )
2026-07-19 16:05:13 -04:00
chenyu and GitHub
89117d8b9e
use real shift in l2i decomp [pr] ( #17080 )
...
works for variable shift distace too, also fixed signed arithmetic fill
2026-07-19 13:15:06 -04:00
chenyu and GitHub
9970a0aad0
fix Tensor << Tensor for x86 ( #17082 )
...
* fix Tensor << Tensor for x86
* torch
2026-07-19 12:31:05 -04:00
chenyu and GitHub
0146a30125
improve cast to unsign min_max [pr] ( #17078 )
2026-07-18 21:58:41 -04:00
George Hotz and GitHub
b53cd35cff
llm: make tokenizer fast (kimi) ( #17077 )
...
* llm: make tokenizer fast
* simpler
* re.escape + qcom mypy fix
2026-07-18 17:31:59 -07:00
nimlgen and GitHub
232529ce88
hcq2: simpler sync ( #17069 )
...
* x
* y
* n
2026-07-18 16:27:44 +03:00
chenyu and GitHub
47629f4bcf
more weak dtype materialization raise ( #17071 )
2026-07-17 23:15:14 -04:00
chenyu and GitHub
f315df29a0
no weak Tensor from and to real buffer ( #17067 )
...
* no weak Tensor from and to real buffer
creation, assign, safe_save
* is_numpy_ndarray to tensor
* one more
2026-07-17 16:09:10 -04:00
George Hotz and GitHub
86a6ad8ed2
llm: split cli.py into serve.py with the HTTP server ( #17065 )
...
* llm: split cli.py into serve.py with the HTTP server
* min edit
2026-07-17 10:45:37 -07:00
George Hotz and GitHub
3ee2baf71d
llm: add tool calling support (kimi) ( #17061 )
...
* llm: add tool calling support
* simpler
* cls
* gpt cleanup
* more gpt cleanups
* tests for tools calling
2026-07-17 10:20:01 -07:00
qazal and GitHub
7dd3422c63
llama: replace two stage amax with atomics ( #17063 )
...
* atomic amax in c kernels
* quantize fp8 UOp kernel
* diff
2026-07-17 19:27:10 +09:00
sirhcm and GitHub
6f1176ea90
benchmarks: test usbgpu copy speeds on comma ( #17060 )
2026-07-17 02:02:21 -04:00
George Hotz and GitHub
46172bb7c7
llm: add optional jinja template support (kimi) ( #17058 )
...
* add jinja template support (kimi)
* fix tests
* lil
* more crap to fallback
2026-07-16 19:02:58 -07:00
chenyu and GitHub
88826a6f35
no weak dtype for randn_like either ( #17055 )
2026-07-16 18:29:12 -04:00
chenyu and GitHub
3bfd62e915
fix 0 size tolist to match numpy ( #17054 )
2026-07-16 17:42:52 -04:00
chenyu and GitHub
6fa0b2b19e
materialize weak dtype casts to default ( #17051 )
...
in clone and _buffer
2026-07-16 16:12:33 -04:00
George Hotz and GitHub
4d8c3d3fc9
add test_hgemm to test_tiny ( #17050 )
...
* add test_hgemm to test_tiny
* dsp skip
2026-07-16 13:12:10 -07:00
chenyu and GitHub
f6a92d0a16
sum_acc_dtype(weak) is weak ( #17047 )
...
also no explicit weak for rand
2026-07-16 14:32:37 -04:00
George Hotz and GitHub
61e104bdfb
use UOp.wmma everywhere ( #17045 )
...
* use UOp.wmma everywhere
* fix
2026-07-16 10:40:48 -07:00
chenyu and GitHub
5a4156c5d1
bitcast and element_size raise for weak dtypes ( #17046 )
2026-07-16 13:07:45 -04:00
chenyu and GitHub
dba8b6b505
allow weak alu operands ( #17044 )
2026-07-16 09:33:20 -04:00
810d8732f9
fix n^2 in limit_bufs by memoizing reachable loads [PR] ( #17017 )
...
* fix n^2 in limit_bufs by memoizing reachable loads [pr]
* Update test_schedule.py
---------
Co-authored-by: Jacob Kitchen <[email protected] >
2026-07-15 23:54:04 -07:00
chenyu and GitHub
e8d3047a50
dtype_from_uop cleanup [PR] ( #17036 )
2026-07-15 21:52:21 -04:00
chenyu and GitHub
be075b200a
weak dtypes in dtype_from_uop [PR] ( #17032 )
...
* weak dtypes in dtype_from_uop [PR]
* no weak in spec_program
* weak const fold tests
2026-07-15 16:54:31 -04:00
nimlgen and GitHub
d6fddb066f
usb: keep only custom ( #17029 )
...
* usb: keep only custom
* mockgpu by gpt
* gpt said sorry
* revert
* reset
* fix
* flash
2026-07-15 22:40:52 +03:00
chenyu and GitHub
c23d8188e1
remove _ensure_float [pr] ( #17030 )
...
do this cast late. allow `SQRT(int)`
2026-07-15 11:19:20 -04:00
chenyu and GitHub
0d19970edc
least_upper_dtype in dtype_from_uop [PR] ( #17028 )
2026-07-15 09:27:53 -04:00
chenyu and GitHub
ebe26420a7
update where Invalid rules [pr] ( #17026 )
...
fixed TestInvalidTensor.test_tensor_index
2026-07-15 00:00:30 -04:00
chenyu and GitHub
47ddf94f17
remove InvalidType lt and gt ( #17023 )
...
not really used
2026-07-14 21:59:15 -04:00
sirhcm and GitHub
c9baa2ef79
use pattern matcher in contiguous_view_offset [PR] ( #17022 )
2026-07-14 19:37:36 -04:00