chenyu and GitHub
12f28ac9d4
catch runtime error in search._time_program ( #3106 )
...
return inf if search encountered runtime errors.
2024-01-12 21:53:13 -05:00
chenyu and GitHub
f018a55ea1
update NumNode.__hash__ to be hash(self.b) ( #3105 )
...
with this, `a:=NumNode(x) == b` implies `hash(a) == hash(b)`
2024-01-12 19:46:21 -05:00
chenyu and GitHub
c3c35f9142
flag to profile mixtral - 1.7 tok/s now ( #3104 )
2024-01-12 18:54:27 -05:00
chenyu and GitHub
e078e2d060
add half @ half to mac benchmark ( #3103 )
2024-01-12 16:38:41 -05:00
Francis Lam and GitHub
ddbdb52f77
wmma: enable METAL half tensor cores and clean up cstyle ( #3095 )
...
* wmma: enable METAL half tensor cores and clean up cstyle
* revert simple_matmul rand changes and break line in tensor
* added metal fp16->fp32 tensor core
2024-01-12 16:25:28 -05:00
chenyu and GitHub
f96fc6e9d4
fix gpt2 with empty prompt take 2 ( #3102 )
...
logits would be empty so need to replace that with ones before sampling, also cannot reshape with -1 when there's 0 in other axes
2024-01-12 14:46:36 -05:00
chenyu and GitHub
ca46d3541b
Revert "fix gpt2 with empty prompt" ( #3101 )
2024-01-12 14:27:41 -05:00
chenyu and GitHub
1d7f01bc6d
fix gpt2 with empty prompt ( #3100 )
...
logits would be empty so need to replace that with ones before sampling, also cannot reshape with -1 when there's 0 in other axes
2024-01-12 14:18:17 -05:00
SnakeOnex and GitHub
0c49d38ba7
replace with tensor op ( #3099 )
2024-01-12 14:13:40 -05:00
chenyu and GitHub
f3a50b4e40
fix broadcasted logic if there's 0 in shapes ( #3097 )
...
* fix broadcasted logic if there's 0 in shapes
should always expand into 0, not the other way around. fixed matmul with 0 in input shapes.
for forwards for now though, backward is more involved and would need to change 0 size shortcuts
* fix tests
2024-01-12 13:32:43 -05:00
SnakeOnex and GitHub
025fbf4e80
One hot in tensor.py ( #3093 )
...
* onehot in Tensor.py
* one_hot tests
* works for all shapes, not just 1
* pylint
* not a static method
* moved around, num_classes mandatory
* pylint
* pylint
* space & moving
* formatting
* moved tests
2024-01-12 13:31:18 -05:00
chenyu and GitHub
7086d77db1
bugfix do not reset shapetracker of 0 size lazybuffer ( #3096 )
...
it might be coming from an expand, and resetting results incorrect stride. caught by interpreted backend
2024-01-11 23:22:52 -05:00
Yixiang Gao and GitHub
13e872b53f
add mutigpu support for llama attention ( #3064 )
...
* add llama attention test for multigpu
* test fails
* kv cache trying to shrink on sharded axis
* mask None works for scale dot product
* kv cache seems to be working but scale dot product breaks
* scaled dot product works, but the last linear layer failed
* running into the reshape case where it could be wrong for multigpu
* making sure it was the reshape
* adding contiguous doesn't solve
* need to shard more properly
* remove reshape test
* minor adjustment to scale dot product attention test
* weights are sharded wrong
* continue fix new weight sharding
* clean up
* fix attention when start_pos is 0
* remove print
* add TODOs for the best mutigpu interface
2024-01-11 16:31:02 -08:00
chenyu and GitHub
dcf7ecaaff
update jit type annotation post lazy rewrite ( #3091 )
2024-01-11 15:49:30 -05:00
chenyu and GitHub
0fe6904351
use device from LinearizerOptions in kernel search ( #3090 )
...
* use device from LinearizerOptions in kernel search
removed all Device.DEFAULT in search.py
* pass device string for parallel pickle
* device for interpreted backends in LinearizerOptions
2024-01-11 14:46:03 -05:00
chenyu and GitHub
93e3f952aa
use BEAM=2 instead of BEAM=4 in cuda ci gpt2 ( #3089 )
...
BEAM=2 is faster and less search time. investigating why BEAM2+BEAM4 is slower than BEAM2 alone
2024-01-11 13:21:06 -05:00
chenyu and GitHub
f502c9b08f
minor cleanup of View.reshape ( #3088 )
...
* minor cleanup of View.reshape
removed some redundant logic
* new_strides
* revert that
2024-01-11 13:05:54 -05:00
chenyu and GitHub
f40299c3fe
remove the third merging state in view._merge_dims ( #3085 )
...
no logic depends on state == 0 or state == 2
2024-01-11 12:07:43 -05:00
chenyu and GitHub
7f9590d357
hotfix disable flaky mac runner wino cifar ( #3087 )
2024-01-11 11:57:05 -05:00
Yixiang Gao and GitHub
adcc844755
cat works ( #3086 )
2024-01-11 08:25:20 -08:00
chenyu and GitHub
cdeab9ad97
mem_estimate is always int, not symbolic ( #3083 )
...
* mem_estimate is always int, not symbolic
op_estimate can be symbolic, but mem_estimate is always int, thus we don't need to sym_infer it.
fixed some long lines too. update_stats is a very big function
* operator does not need underscores
2024-01-10 23:39:51 -05:00
Francis Lam and GitHub
162fa61a32
wmma: clean up device specific tensor core code ( #3081 )
2024-01-10 21:03:09 -05:00
chenyu and GitHub
d218d13885
minor cleanups of lazy.py ( #3080 )
2024-01-10 20:17:56 -05:00
chenyu and GitHub
56dda33fc6
Tensor.expand resolves the new_shape before shortcut return ( #3078 )
...
similar to how reshape is done. also updated shrink shortcut criteria to read similar to pad
2024-01-10 14:29:15 -05:00
Yixiang Gao and GitHub
6842476ca6
better test demonstration ( #3077 )
...
* a better test demonstration
* fix white space
2024-01-10 10:50:52 -08:00
chenyu and GitHub
507e0afba0
fix onehot and jit in examples/transformer ( #3073 )
...
trained to 0.999 in < 6 seconds on M1 Max consistently
2024-01-10 02:22:41 -05:00
chenyu and GitHub
4342fccc83
filter_strides -> canonicalize_strides ( #3072 )
2024-01-10 01:06:48 -05:00
chenyu and GitHub
023f5df0e9
simpler idxs_to_idx ( #3071 )
2024-01-10 00:30:10 -05:00
George Hotz and GitHub
2495ca95c7
early gate the graph ( #3070 )
2024-01-09 20:17:13 -08:00
George Hotz and GitHub
ff0d6e4551
jit autorealizes output ( #3069 )
2024-01-09 20:10:22 -08:00
geohot
ae83733431
hotfix: examples/transformer.py
2024-01-09 19:28:09 -08:00
chenyu and GitHub
145718a90f
unbind view or shapetracker also returns var_val ( #3067 )
...
* unbind view or shapetracker also returns var_val
4% faster for llama compile time
* one line less
* unbound_views
2024-01-09 21:45:05 -05:00
ef3aa6d7fb
update gh actions ( #3033 )
...
* update checkout actions
* update upload artifact
* update setup python
---------
Co-authored-by: George Hotz <[email protected] >
2024-01-09 17:52:22 -08:00
George Hotz and GitHub
3f80c1a098
speedtweaks3: apply shouldn't use the tensor constructor ( #3065 )
...
* speedtweaks3: apply shouldn't use the tensor constructor
* replace 0 size with CONST, not 0 in shape
2024-01-09 17:42:33 -08:00
geohot
0abe72b677
hotfix: use is for enum compare, a few more
2024-01-09 16:53:13 -08:00
geohot
b2b5849f74
hotfix: use is for enum compare
2024-01-09 16:47:27 -08:00
George Hotz and GitHub
ac3f246c11
cached size ( #3060 )
...
* cached size
* simplify simplify
* 0 doesn't have base
* fix test
* cleaner cache
* hmm, metal is flaky on this...might be real(ish) but useless as test
* short circuit reshape/expand properly
* better reshape bypass
2024-01-09 16:37:37 -08:00
Yixiang Gao and GitHub
73b72b8de2
test scaled dot product attention ( #3063 )
...
* add test
* add initial test for scaled dot product attention
* test pass for scaled dot product attention
2024-01-09 14:30:57 -08:00
chenyu and GitHub
55ac2a2cf7
Tensor.cat with 0 shape tensors ( #3062 )
...
* Tensor.cat with 0 shape tensors
supported both 0 in cat axis (for a subset of input), or 0 in non-cat axis (all needs to be 0)
* no shp
2024-01-09 16:54:06 -05:00
chenyu and GitHub
f0d7ad8aaa
fix gpt2 attention with start_pos = 0 ( #3061 )
...
* fix gpt2 attention with start_pos size 1
test cases taken from ll_transformer branch
* fix interpreted
2024-01-09 16:14:55 -05:00
George Hotz and GitHub
39b91131bc
Speed tweaks ( #3059 )
...
* base doesn't have to be a function
* no double fetch
* pop, don't check
* make the gc happy
* avoid hasattr
* cache canonicalize
* remove assert, faster base
* don't redefine that every time
2024-01-09 11:34:17 -08:00
geohot
bf6281f316
hotfix: remove useless slow assert from ShapeTracker
2024-01-09 10:56:36 -08:00
George Hotz and GitHub
4b687af98f
explicit lazybuffer caching ( #3058 )
2024-01-09 10:52:37 -08:00
George Hotz and GitHub
2c6f2e899d
No extra vars call ( #3054 )
...
* remove unused reciprocal
* comment
* remove unneeded call to vars
* free speedup
v0.8.0
2024-01-09 09:52:58 -08:00
Yixiang Gao and GitHub
259bf9bffc
add multigpu test for RMSNorm ( #3056 )
...
* need all gather
* add two multigpu test scenarios for RMSNorm
2024-01-09 09:52:51 -08:00
chenyu and GitHub
dab8214103
unit tests for Device.canonicalize ( #3055 )
2024-01-09 12:47:20 -05:00
George Hotz and GitHub
374f7659a7
remove unused reciprocal ( #3053 )
...
* remove unused reciprocal
* comment
2024-01-09 08:59:04 -08:00
Yixiang Gao and GitHub
a686663657
make Embedding device aware for multigpu ( #3051 )
...
* make Embedding device aware for multigpu
* split line instead of igore because that's cheating
* add test incomplete
* add test complete
* remove comment
* fix white space
* remove nn.Embedding
2024-01-08 20:09:26 -08:00
chenyu and GitHub
19298e7a3f
Device._buffers -> Device._devices ( #3052 )
...
backend devices used to be called buffers
2024-01-08 21:30:38 -05:00
chenyu and GitHub
4f4e8634b8
use in_features directly in nn.Linear.__init__ bound check ( #3050 )
...
* use in_features directly in nn.Linear.__init__ bound check
get rid of the unnecessary check of isinstance int
* that is always int
* long lines
2024-01-08 19:32:35 -05:00