 wozeparrotandGitHub
|
a3d59faef6
|
llama: don't save weight (#16252)
|
2026-05-18 17:05:45 -07:00 |
|
 qazalandGitHub
|
18b102f355
|
llama: also use 7.1 comgr, update startup_walltime.sh (#16253)
|
2026-05-19 08:59:02 +09:00 |
|
 qazalandGitHub
|
98b8a2b407
|
llama: use hipcc 7.1 version (#16250)
|
2026-05-19 08:09:57 +09:00 |
|
 chenyuandGitHub
|
dcee90aa3f
|
remove requires_grad use in extra/examples (#16238)
except the ones fed into optimizer
|
2026-05-16 18:40:26 -04:00 |
|
 qazalandGitHub
|
ebcb7b7cc0
|
fp8 gemm tests with scale args (#16231)
* update atol
* update fp8 path
* more work
* update profile.sh
|
2026-05-16 20:47:58 +09:00 |
|
 wozeparrotandGitHub
|
159694347e
|
llama: fix running flat_llama (#16224)
|
2026-05-15 20:16:48 -07:00 |
|
 chenyuandGitHub
|
07a172dbbb
|
remove noop requires_grad_ calls (#16213)
|
2026-05-15 13:31:10 -04:00 |
|
 chenyuandGitHub
|
409bb0c9ad
|
requires_grad cannot be None (#16212)
final goal is to remove requires_grad, first change the default to True, and don't allow None
|
2026-05-15 02:01:04 -04:00 |
|
 wozeparrotandGitHub
|
b4d267dfd4
|
llama: only save when small (#16208)
|
2026-05-14 17:46:29 -07:00 |
|
 wozeparrotandGitHub
|
88ac2ac1fd
|
llama: cleanups (#16189)
|
2026-05-13 17:08:06 -07:00 |
|
 wozeparrotandGitHub
|
e97f2c1114
|
llama: only gemm + fa custom kernel (#16180)
* llama: tie store to grad directly
* llama: set mp flags
* llama: non fused grad fp8 quantize path
|
2026-05-12 21:03:49 -07:00 |
|
 wozeparrotandGitHub
|
e9359d9e7d
|
more llama mp fixes (#16151)
* llama: SPLIT_W13
* llama: fix with no fused kernels
* llama: cast to bf16 on non asm_gemm patH
* llama: new mp flags
|
2026-05-11 21:29:23 -07:00 |
|
 wozeparrotandGitHub
|
026688f03f
|
llama: move to correct dir (#16118)
|
2026-05-08 19:42:16 -07:00 |
|
 qazalandGitHub
|
a9a87ad8fd
|
viz/cli: less flags (#16076)
* viz/cli: merge -s and -i flags
* only -t
* merge parser
* fix
|
2026-05-08 00:22:40 +09:00 |
|
 wozeparrotandGitHub
|
730fa66bf3
|
llama speed 6 (#16071)
|
2026-05-06 20:51:03 -07:00 |
|
 wozeparrotandGitHub
|
ab6218bc92
|
llama mp fixes (#16050)
|
2026-05-05 15:35:32 -07:00 |
|
 wozeparrotandGitHub
|
528d35e306
|
llama speed 4 (#15993)
|
2026-04-30 17:14:41 -07:00 |
|
 wozeparrotandGitHub
|
0080489abe
|
llama: use env vars (#15978)
|
2026-04-29 12:37:15 -07:00 |
|
 wozeparrotandGitHub
|
ef09071073
|
llama: speed 2 (#15960)
|
2026-04-28 20:44:37 -07:00 |
|
 wozeparrotandGitHub
|
5e861cd2c4
|
llama: move llama kernels to llama_kernels (#15952)
|
2026-04-27 22:48:53 -07:00 |
|
 qazalandGitHub
|
9a23de7d27
|
viz/cli: unify profile and rewrites, -s ALL default (#15931)
* work
* workg
* better
* cleanup
* better defaults
* --ls
* better
* work
* update llama
* update
|
2026-04-25 22:31:24 +09:00 |
|
 wozeparrotandGitHub
|
4b908b6e2c
|
llama: fused ce loss (#15920)
|
2026-04-24 20:01:24 -07:00 |
|
 wozeparrotandGitHub
|
9d134a2848
|
llama: fix fakedata timing (#15905)
|
2026-04-23 21:37:03 -07:00 |
|
 wozeparrotandGitHub
|
d3cbd781d9
|
llama: use fused norm mul quantize for w13 (#15878)
|
2026-04-22 21:27:41 -07:00 |
|
 wozeparrotandGitHub
|
87378331e8
|
llama: fused mul quantize fp8 (#15863)
|
2026-04-21 20:58:37 -07:00 |
|
 qazalandGitHub
|
f9655af2a3
|
viz/cli: move to tinygrad (#15835)
* move cli
* update imports
* cleanup the readme
* edit
* work
* details
* python -m tinygrad.viz.cli
* do not execv in non tty
* option
* lint
* simpler
* gemm pmc
|
2026-04-21 13:35:10 +09:00 |
|
 wozeparrotandGitHub
|
f28ea84de2
|
llama: fused silu fp8 amax (#15798)
* llama: combined w13
* llama: fused swiglu+fp8
* llama: fix amax interleaving
* llama: don't need seperate matmul
|
2026-04-19 12:03:55 +08:00 |
|
 wozeparrotandGitHub
|
06343092c8
|
llama: combined w13 (#15803)
|
2026-04-17 22:27:31 -07:00 |
|
 qazalandGitHub
|
a227dbece1
|
viz/cli: reconstruct DEBUG output (#15791)
* work
* work
* ext
* padding
* at time
* work
* reorder
* less flags
* num_rows
* feedback
* pmc
|
2026-04-17 18:27:58 +03:00 |
|
 wozeparrotandGitHub
|
9e60e4a7e7
|
llama: native fp8 (#15733)
|
2026-04-16 22:16:05 -07:00 |
|
 wozeparrotandGitHub
|
3721c60bef
|
llama: bs 16 (#15737)
|
2026-04-14 19:52:03 -07:00 |
|
 wozeparrotandGitHub
|
480ad264a4
|
llama: per device amax (#15735)
|
2026-04-14 19:01:17 -07:00 |
|
 wozeparrotandGitHub
|
2b8d303f75
|
allreduce in precast dtype (#15689)
|
2026-04-13 20:24:12 -07:00 |
|
 qazalandGitHub
|
054d78e6ff
|
fix llama profile.sh NULL source (#15685)
|
2026-04-11 22:56:05 +09:00 |
|
 wozeparrotandGitHub
|
457508d5a0
|
llama: save more 2 (#15681)
|
2026-04-11 01:03:36 -07:00 |
|
 wozeparrotandGitHub
|
590464c8d8
|
llama: only support wqkv path + cleanups (#15680)
* llama: only support wqkv path + cleanups
* llama: missing transpose
|
2026-04-11 07:39:27 +08:00 |
|
 wozeparrotandGitHub
|
55bcd7cc9e
|
llama amax outside (#15670)
|
2026-04-09 23:08:03 -07:00 |
|
 chenyuandGitHub
|
839d37b7bc
|
update median_step_time in model_train.py (#15649)
BENCHMARK=5 used to pick the 4th largest, not the middle one
|
2026-04-08 09:53:59 -04:00 |
|
 qazalandGitHub
|
39a029ec55
|
remove ASM_GEMM context var (#15645)
|
2026-04-08 18:02:40 +09:00 |
|
 wozeparrotandGitHub
|
70dbd35023
|
llama: move custom_kernel into flat_llama (#15643)
|
2026-04-08 00:19:14 -07:00 |
|
 qazalandGitHub
|
890286e8d6
|
update llama profile.sh (#15633)
* update llama profile.sh
* BENCHMARK 5
|
2026-04-08 03:18:45 +09:00 |
|
 wozeparrotandGitHub
|
810d7c00cd
|
llama: unify scripts (#15628)
|
2026-04-06 20:28:08 -07:00 |
|
 
|
7e54992bf6
|
fp8 llama (#15588)
Co-authored-by: qazal <[email protected]>
|
2026-04-04 18:24:57 -07:00 |
|
 qazalandGitHub
|
f7aed180e4
|
viz/cli: add Other row in profiler (#15600)
|
2026-04-04 22:40:53 +09:00 |
|
 wozeparrotandGitHub
|
5b2a3251c4
|
mlperf system json for mi350 (#15575)
|
2026-04-01 15:30:33 -07:00 |
|
 qazalandGitHub
|
09f60d80fd
|
llama: fix FP8=1 FAKEDATA=1 (#15564)
|
2026-04-01 20:53:03 +09:00 |
|
 wozeparrotandGitHub
|
8b5b9a0e90
|
llama: run_and_time (#15533)
|
2026-03-31 15:46:16 -07:00 |
|
 qazalandGitHub
|
8feb8edc68
|
gemm/asm: add fp8 support to cdna asm_gemm (#15542)
* work
* hmm, mixins
* rhs_transposed
* also fix the dtype
* check for hipcc
* Exception
* select dev
* default
|
2026-03-31 19:32:54 +09:00 |
|
 sirhcmandGitHub
|
adbfd82d1d
|
DEV is ContextVar, setting Device.DEFAULT is deprecated (#15508)
|
2026-03-30 17:10:49 -04:00 |
|
 wozeparrotandGitHub
|
0c3e438229
|
llama: mllog (#15502)
|
2026-03-28 11:18:25 -07:00 |
|