 chenyuandGitHub
|
ac3f56a1a2
|
more shift tests (#17083)
|
2026-07-19 16:05:13 -04:00 |
|
 chenyuandGitHub
|
89117d8b9e
|
use real shift in l2i decomp [pr] (#17080)
works for variable shift distace too, also fixed signed arithmetic fill
|
2026-07-19 13:15:06 -04:00 |
|
 chenyuandGitHub
|
9970a0aad0
|
fix Tensor << Tensor for x86 (#17082)
* fix Tensor << Tensor for x86
* torch
|
2026-07-19 12:31:05 -04:00 |
|
 chenyuandGitHub
|
0146a30125
|
improve cast to unsign min_max [pr] (#17078)
|
2026-07-18 21:58:41 -04:00 |
|
 George HotzandGitHub
|
b53cd35cff
|
llm: make tokenizer fast (kimi) (#17077)
* llm: make tokenizer fast
* simpler
* re.escape + qcom mypy fix
|
2026-07-18 17:31:59 -07:00 |
|
 nimlgenandGitHub
|
232529ce88
|
hcq2: simpler sync (#17069)
* x
* y
* n
|
2026-07-18 16:27:44 +03:00 |
|
 chenyuandGitHub
|
47629f4bcf
|
more weak dtype materialization raise (#17071)
|
2026-07-17 23:15:14 -04:00 |
|
 chenyuandGitHub
|
f315df29a0
|
no weak Tensor from and to real buffer (#17067)
* no weak Tensor from and to real buffer
creation, assign, safe_save
* is_numpy_ndarray to tensor
* one more
|
2026-07-17 16:09:10 -04:00 |
|
 George HotzandGitHub
|
86a6ad8ed2
|
llm: split cli.py into serve.py with the HTTP server (#17065)
* llm: split cli.py into serve.py with the HTTP server
* min edit
|
2026-07-17 10:45:37 -07:00 |
|
 George HotzandGitHub
|
3ee2baf71d
|
llm: add tool calling support (kimi) (#17061)
* llm: add tool calling support
* simpler
* cls
* gpt cleanup
* more gpt cleanups
* tests for tools calling
|
2026-07-17 10:20:01 -07:00 |
|
 qazalandGitHub
|
7dd3422c63
|
llama: replace two stage amax with atomics (#17063)
* atomic amax in c kernels
* quantize fp8 UOp kernel
* diff
|
2026-07-17 19:27:10 +09:00 |
|
 sirhcmandGitHub
|
6f1176ea90
|
benchmarks: test usbgpu copy speeds on comma (#17060)
|
2026-07-17 02:02:21 -04:00 |
|
 George HotzandGitHub
|
46172bb7c7
|
llm: add optional jinja template support (kimi) (#17058)
* add jinja template support (kimi)
* fix tests
* lil
* more crap to fallback
|
2026-07-16 19:02:58 -07:00 |
|
 chenyuandGitHub
|
88826a6f35
|
no weak dtype for randn_like either (#17055)
|
2026-07-16 18:29:12 -04:00 |
|
 chenyuandGitHub
|
3bfd62e915
|
fix 0 size tolist to match numpy (#17054)
|
2026-07-16 17:42:52 -04:00 |
|
 chenyuandGitHub
|
6fa0b2b19e
|
materialize weak dtype casts to default (#17051)
in clone and _buffer
|
2026-07-16 16:12:33 -04:00 |
|
 George HotzandGitHub
|
4d8c3d3fc9
|
add test_hgemm to test_tiny (#17050)
* add test_hgemm to test_tiny
* dsp skip
|
2026-07-16 13:12:10 -07:00 |
|
 chenyuandGitHub
|
f6a92d0a16
|
sum_acc_dtype(weak) is weak (#17047)
also no explicit weak for rand
|
2026-07-16 14:32:37 -04:00 |
|
 George HotzandGitHub
|
61e104bdfb
|
use UOp.wmma everywhere (#17045)
* use UOp.wmma everywhere
* fix
|
2026-07-16 10:40:48 -07:00 |
|
 chenyuandGitHub
|
5a4156c5d1
|
bitcast and element_size raise for weak dtypes (#17046)
|
2026-07-16 13:07:45 -04:00 |
|
 chenyuandGitHub
|
dba8b6b505
|
allow weak alu operands (#17044)
|
2026-07-16 09:33:20 -04:00 |
|
 
|
810d8732f9
|
fix n^2 in limit_bufs by memoizing reachable loads [PR] (#17017)
* fix n^2 in limit_bufs by memoizing reachable loads [pr]
* Update test_schedule.py
---------
Co-authored-by: Jacob Kitchen <[email protected]>
|
2026-07-15 23:54:04 -07:00 |
|
 chenyuandGitHub
|
e8d3047a50
|
dtype_from_uop cleanup [PR] (#17036)
|
2026-07-15 21:52:21 -04:00 |
|
 chenyuandGitHub
|
be075b200a
|
weak dtypes in dtype_from_uop [PR] (#17032)
* weak dtypes in dtype_from_uop [PR]
* no weak in spec_program
* weak const fold tests
|
2026-07-15 16:54:31 -04:00 |
|
 nimlgenandGitHub
|
d6fddb066f
|
usb: keep only custom (#17029)
* usb: keep only custom
* mockgpu by gpt
* gpt said sorry
* revert
* reset
* fix
* flash
|
2026-07-15 22:40:52 +03:00 |
|
 chenyuandGitHub
|
c23d8188e1
|
remove _ensure_float [pr] (#17030)
do this cast late. allow `SQRT(int)`
|
2026-07-15 11:19:20 -04:00 |
|
 chenyuandGitHub
|
0d19970edc
|
least_upper_dtype in dtype_from_uop [PR] (#17028)
|
2026-07-15 09:27:53 -04:00 |
|
 chenyuandGitHub
|
ebe26420a7
|
update where Invalid rules [pr] (#17026)
fixed TestInvalidTensor.test_tensor_index
|
2026-07-15 00:00:30 -04:00 |
|
 chenyuandGitHub
|
47ddf94f17
|
remove InvalidType lt and gt (#17023)
not really used
|
2026-07-14 21:59:15 -04:00 |
|
 sirhcmandGitHub
|
c9baa2ef79
|
use pattern matcher in contiguous_view_offset [PR] (#17022)
|
2026-07-14 19:37:36 -04:00 |
|
 nimlgenandGitHub
|
4257939e50
|
remove copyin/copyout from Buffer (#17020)
* remove copyin/copyout from Buffer
* x
* x
* x
* x
|
2026-07-14 19:47:22 +03:00 |
|
 qazalandGitHub
|
939f28d571
|
fused qkv rope custom kernel (#17021)
* work
* fused qkv_norm
* work
* speed
* not that yet
* test cleanup
* just clone
* remove .realize()
* cleanup tests
|
2026-07-15 01:08:42 +09:00 |
|
 chenyuandGitHub
|
82fbca43c5
|
fix Tensor(np) dtype and support fp8 safetensor (#17019)
|
2026-07-14 09:31:21 -04:00 |
|
 chenyuandGitHub
|
872225e47d
|
update dtype tests for small dtypes (#17016)
|
2026-07-14 08:07:00 -04:00 |
|
 chenyuandGitHub
|
55bb251130
|
add pm_manual_bf16_cast to Metal [pr] (#17015)
mitigate metal compiler bug for
`as_type<half>( (bfloat)(const) )`
|
2026-07-13 21:53:47 -04:00 |
|
 sirhcmandGitHub
|
a9fbc7db7b
|
expect _offset support, CL and WEBGPU are outliers (#17014)
|
2026-07-13 18:53:32 -04:00 |
|
 chenyuandGitHub
|
9ce96c2628
|
fix subnormal in test_dtype (#17013)
* fix subnormal in test_dtype
should fix flaky test/backend/test_dtype.py::TestFp8e4m3::test_casts_from
* better
|
2026-07-13 18:53:13 -04:00 |
|
 chenyuandGitHub
|
681a5e0cfd
|
remove UOp cast and bitcast override [PR] (#17011)
|
2026-07-13 14:23:11 -04:00 |
|
 chenyuandGitHub
|
0410c9325d
|
make test/null follow the SPEC (#17010)
|
2026-07-13 14:01:41 -04:00 |
|
 nimlgenandGitHub
|
e4bdc529c4
|
hcq2 ci (#17008)
* hcq2 ci
* x
|
2026-07-13 19:29:08 +03:00 |
|
 nimlgenandGitHub
|
4536a57f79
|
hcq rename map (#17009)
* hcq rename map
* x
|
2026-07-13 19:23:12 +03:00 |
|
 George HotzandGitHub
|
dde2e736e5
|
fix disable_gc decorator reentrancy (#16999)
|
2026-07-12 15:32:54 -07:00 |
|
 George HotzandGitHub
|
03ecad9486
|
full removal of dtype.vec (#16996)
* full removal of dtype.vec
* fix typo
|
2026-07-12 09:25:50 -07:00 |
|
 chenyuandGitHub
|
5a0751076d
|
simpler vconst_like (#16990)
also always use const_like in symbolic, fixed a crash
|
2026-07-12 08:21:57 -04:00 |
|
 qazalandGitHub
|
cae6696d75
|
llama: split current and next amax state (#16993)
|
2026-07-12 18:52:33 +09:00 |
|
 chenyuandGitHub
|
047a467bf9
|
delete UOp._stack and UOp.vectorize [PR] (#16988)
|
2026-07-11 15:26:55 -04:00 |
|
 chenyuandGitHub
|
9d47014fd8
|
first class STACK [PR] (#16986)
|
2026-07-11 13:55:32 -04:00 |
|
 George HotzandGitHub
|
afeb5c708f
|
x86 simplification (#16983)
* simplify x86
* more extras
* simpler
* work
* fixes
* should pasS
* cmt-n
* delete more
* and more
|
2026-07-11 08:11:05 -07:00 |
|
 qazalandGitHub
|
75a4bfddc9
|
fp8 gemm tests including fused scales (#16980)
* fp8 gemm tests matching fused scales
* work
* diff
|
2026-07-11 12:33:21 +09:00 |
|
 chenyuandGitHub
|
df50e0814c
|
explicit error for unbound Variable in program (#16971)
also allow Tensor(UOp, dtype)
|
2026-07-10 16:23:59 -04:00 |
|