Compare commits

...
Author SHA1 Message Date
sirhcm 56189c1ec4 onnxruntime respects NUM_CPU_THREADS
Platform Tests / MacOS (unit) (pull_request) Waiting to run
Platform Tests / MacOS (unit, mock) (pull_request) Waiting to run
Platform Tests / MacOS (DEV=METAL) (1) (pull_request) Waiting to run
Platform Tests / MacOS (DEV=METAL) (2) (pull_request) Waiting to run
Platform Tests / MacOS (DEV=CPU:CLANG) (pull_request) Waiting to run
Platform Tests / MacOS (DEV=CPU:LLVM) (pull_request) Waiting to run
Platform Tests / MacOS (DEV=CPU:LVP) (pull_request) Waiting to run
Platform Tests / MacOS (DEV=WEBGPU) (pull_request) Waiting to run
Platform Tests / Windows (DEV=CPU:CLANG) (pull_request) Waiting to run
Platform Tests / Windows (DEV=CPU:LLVM) (pull_request) Waiting to run
Platform Tests / Windows (DEV=CPU:X86) (pull_request) Waiting to run
Platform Tests / Windows (DEV=WEBGPU) (pull_request) Waiting to run
Unit Tests / Linters (pull_request) Successful in 2m0s
Unit Tests / Models (pull_request) Successful in 1m30s
Unit Tests / ONNX (CPU) Tests (pull_request) Successful in 1m51s
Unit Tests / Linux (DSP) (pull_request) Successful in 1m41s
Unit Tests / Docs (pull_request) Successful in 2m57s
Unit Tests / Test LLM (pull_request) Successful in 1m56s
Unit Tests / Fuzzing (pull_request) Successful in 2m25s
Unit Tests / Torch Backend Training (pull_request) Successful in 3m21s
Unit Tests / Null Tests (pull_request) Successful in 3m16s
Unit Tests / Python Backend (pull_request) Successful in 3m20s
Unit Tests / Unit Tests (pull_request) Successful in 3m22s
Check Line Counts / Check PR Branch status (pull_request_target) Successful in 14s
Unit Tests / openpilot Compile Tests (pull_request) Successful in 3m2s
Unit Tests / CL IMAGE Tests (pull_request) Successful in 3m12s
Unit Tests / AMD ASM IDE (pull_request) Successful in 2m34s
Check Line Counts / Core Library Line Difference (pull_request_target) Skipped
Unit Tests / SPEC=2 (1) (pull_request) Successful in 4m6s
Unit Tests / hcq2 (pull_request) Successful in 2m39s
Unit Tests / Torch Backend Tests (pull_request) Successful in 4m24s
Unit Tests / Linux (DEV=CPU:X86) (pull_request) Successful in 3m17s
Unit Tests / SPEC=2 (2) (pull_request) Successful in 3m54s
Unit Tests / Linux (DEV=CPU:LVP) (pull_request) Successful in 3m31s
Unit Tests / Optimization Tests (pull_request) Successful in 3m57s
Unit Tests / Compile-only (DEV=NULL:NAK:sm_120) (pull_request) Successful in 1m41s
Unit Tests / Linux (DEV=CL) (pull_request) Successful in 3m44s
Unit Tests / Linux (DEV=CPU:LLVM) (pull_request) Successful in 3m37s
Unit Tests / Linux (DEV=WEBGPU) (pull_request) Successful in 3m51s
Unit Tests / Compile-only (DEV=NULL:IR3:a630) (pull_request) Successful in 2m20s
Unit Tests / Linux (DEV=CPU:CLANG) (pull_request) Successful in 4m22s
Unit Tests / Linux (amdllvm gfx1100) (pull_request) Successful in 3m27s
Unit Tests / Linux (am) (pull_request) Successful in 3m45s
Unit Tests / Linux (amdllvm gfx1201) (pull_request) Successful in 3m16s
Unit Tests / Linux (amd gfx1100) (pull_request) Successful in 3m48s
Unit Tests / Linux (amd gfx1201) (pull_request) Successful in 3m47s
Unit Tests / Linux (amdllvm gfx950) (pull_request) Successful in 3m13s
Unit Tests / Linux (ptx) (pull_request) Successful in 3m7s
Unit Tests / Linux (nv) (pull_request) Successful in 3m51s
Unit Tests / Linux (amd gfx950) (pull_request) Successful in 4m20s
Unit Tests / Compile-only (DEV=NULL:QCOMCL:a630) (pull_request) Successful in 3m48s
2026-08-27 16:41:08 -07:00
sirhcmandGitHub f06832bf6f test llm with --no_chat_template (#17785)
Autogen / In-tree Autogen (macos) (push) Waiting to run
Benchmarks / Mac pytest (push) Waiting to run
Benchmarks / LLM (DEV=AMD) (push) Waiting to run
Benchmarks / LLM (DEV=METAL) (push) Waiting to run
Benchmarks / LLM (DEV=NV) (push) Waiting to run
Benchmarks / HLB-CIFAR10 (DEV=AMD) (push) Waiting to run
Benchmarks / HLB-CIFAR10 (DEV=METAL) (push) Waiting to run
Benchmarks / HLB-CIFAR10 (DEV=NV) (push) Waiting to run
Benchmarks / MLPerf (AMD) (push) Waiting to run
Benchmarks / MLPerf (NV) (push) Waiting to run
Benchmarks / Stable Diffusion (DEV=AMD) (push) Waiting to run
Benchmarks / Stable Diffusion (DEV=METAL) (push) Waiting to run
Benchmarks / Stable Diffusion (DEV=NV) (push) Waiting to run
Benchmarks / Multi-GPU Benchmarks (DEV=AMD) (push) Waiting to run
Benchmarks / Multi-GPU Benchmarks (DEV=NV) (push) Waiting to run
Benchmarks / Tests (DEV=AMD) (push) Waiting to run
Benchmarks / Tests (DEV=METAL) (push) Waiting to run
Benchmarks / Tests (DEV=NV) (push) Waiting to run
Benchmarks / UsbGPU Benchmark (push) Waiting to run
Benchmarks / openpilot 0.11.0 compile3 dmonitoring (DEV=QCOM) (push) Waiting to run
Benchmarks / openpilot 0.11.2 compile3 dmonitoring (DEV=QCOM) (push) Waiting to run
Benchmarks / openpilot 0.11.0 compile3 policy (DEV=QCOM) (push) Waiting to run
Benchmarks / openpilot 0.11.2 compile3 supercombo (DEV=QCOM) (push) Waiting to run
Benchmarks / openpilot 0.11.0 compile3 vision (DEV=QCOM) (push) Waiting to run
Benchmarks / openpilot 0.11.0 compile3 dmonitoring (DEV=QCOM:IR3) (push) Waiting to run
Benchmarks / openpilot 0.11.2 compile3 dmonitoring (DEV=QCOM:IR3) (push) Waiting to run
Benchmarks / openpilot 0.11.0 compile3 policy (DEV=QCOM:IR3) (push) Waiting to run
Benchmarks / openpilot 0.11.2 compile3 supercombo (DEV=QCOM:IR3) (push) Waiting to run
Benchmarks / openpilot 0.11.0 compile3 vision (DEV=QCOM:IR3) (push) Waiting to run
Benchmarks / DSP Benchmark (push) Waiting to run
Benchmarks / UsbGPU Benchmark (comma) (push) Waiting to run
Benchmarks / PCI Driver Benchmark (DEV=AMD) (push) Waiting to run
Benchmarks / PCI Driver Benchmark (DEV=NV) (push) Waiting to run
Benchmarks / LLVM Speed (push) Waiting to run
Platform Tests / MacOS (unit) (push) Waiting to run
Platform Tests / MacOS (unit, mock) (push) Waiting to run
Platform Tests / MacOS (DEV=METAL) (1) (push) Waiting to run
Platform Tests / MacOS (DEV=METAL) (2) (push) Waiting to run
Platform Tests / MacOS (DEV=CPU:CLANG) (push) Waiting to run
Platform Tests / MacOS (DEV=CPU:LLVM) (push) Waiting to run
Platform Tests / MacOS (DEV=CPU:LVP) (push) Waiting to run
Platform Tests / MacOS (DEV=WEBGPU) (push) Waiting to run
Platform Tests / Windows (DEV=CPU:CLANG) (push) Waiting to run
Platform Tests / Windows (DEV=CPU:LLVM) (push) Waiting to run
Platform Tests / Windows (DEV=CPU:X86) (push) Waiting to run
Platform Tests / Windows (DEV=WEBGPU) (push) Waiting to run
Unit Tests / Models (push) Successful in 1m36s
Unit Tests / Linux (DSP) (push) Successful in 1m50s
Unit Tests / Test LLM (push) Successful in 2m0s
Unit Tests / Linters (push) Successful in 2m2s
Unit Tests / Fuzzing (push) Successful in 2m16s
Unit Tests / hcq2 (push) Failing after 2m29s
Unit Tests / Docs (push) Successful in 2m57s
Unit Tests / Python Backend (push) Successful in 3m15s
Unit Tests / openpilot Compile Tests (push) Successful in 3m16s
Unit Tests / AMD ASM IDE (push) Successful in 3m12s
Unit Tests / Null Tests (push) Successful in 3m21s
Unit Tests / CL IMAGE Tests (push) Successful in 3m22s
Unit Tests / Unit Tests (push) Successful in 3m24s
Unit Tests / Torch Backend Training (push) Successful in 3m26s
Unit Tests / Linux (DEV=CPU:X86) (push) Successful in 3m22s
Unit Tests / Linux (DEV=CPU:LVP) (push) Successful in 3m53s
Unit Tests / SPEC=2 (2) (push) Successful in 3m57s
Unit Tests / Linux (DEV=CPU:LLVM) (push) Successful in 3m57s
Unit Tests / SPEC=2 (1) (push) Successful in 4m6s
Unit Tests / Linux (amdllvm gfx1100) (push) Successful in 3m59s
Unit Tests / Linux (amdllvm gfx1201) (push) Successful in 3m58s
Unit Tests / Optimization Tests (push) Successful in 4m14s
Unit Tests / Compile-only (DEV=NULL:NAK:sm_120) (push) Successful in 1m45s
Unit Tests / Linux (DEV=CL) (push) Successful in 4m17s
Unit Tests / Torch Backend Tests (push) Successful in 4m22s
Unit Tests / Linux (DEV=WEBGPU) (push) Successful in 4m24s
Unit Tests / Linux (am) (push) Successful in 4m23s
Unit Tests / Linux (amd gfx1100) (push) Successful in 4m25s
Unit Tests / Linux (amd gfx1201) (push) Successful in 4m24s
Unit Tests / ONNX (CPU) Tests (push) Failing after 4m32s
Unit Tests / Compile-only (DEV=NULL:IR3:a630) (push) Successful in 2m25s
Deploy Docs / deploy (push) Successful in 5m1s
Unit Tests / Linux (DEV=CPU:CLANG) (push) Successful in 5m5s
Unit Tests / Linux (amdllvm gfx950) (push) Successful in 3m32s
Unit Tests / Linux (ptx) (push) Successful in 3m22s
Unit Tests / Linux (nv) (push) Successful in 4m17s
Unit Tests / Linux (amd gfx950) (push) Successful in 5m3s
Unit Tests / Compile-only (DEV=NULL:QCOMCL:a630) (push) Successful in 4m28s
Autogen / In-tree Autogen (push) Successful in 12m36s
2026-08-27 19:01:55 -04:00
3 changed files with 8 additions and 5 deletions
+4 -4
View File
@@ -390,10 +390,10 @@ jobs:
run: |
parallel --link --tagstring '[{1}]' '{2}' \
::: llama 'llama q4' qwen3.5 qwen \
::: $'echo "What\'s a male chicken called? Answer with only one word." | python3 -m tinygrad.llm --model llama3.2:1b | tee /dev/stderr | grep -i rooster' \
$'echo "What\'s a male chicken called? Answer with only one word." | python3 -m tinygrad.llm --model llama3.2:1b-q4 | tee /dev/stderr | grep -i rooster' \
$'echo "What\'s a male chicken called? Answer with only one word." | python3 -m tinygrad.llm --model qwen3.5:0.8b | tee /dev/stderr | grep -i rooster' \
$'echo "What\'s a female chicken called? Answer with only one word." | python3 -m tinygrad.llm --model qwen3:0.6b | tee /dev/stderr | grep -i hen'
::: $'echo "What\'s a male chicken called? Answer with only one word." | python3 -m tinygrad.llm --no_chat_template --model llama3.2:1b | tee /dev/stderr | grep -i rooster' \
$'echo "What\'s a male chicken called? Answer with only one word." | python3 -m tinygrad.llm --no_chat_template --model llama3.2:1b-q4 | tee /dev/stderr | grep -i rooster' \
$'echo "What\'s a male chicken called? Answer with only one word." | python3 -m tinygrad.llm --no_chat_template --model qwen3.5:0.8b | tee /dev/stderr | grep -i rooster' \
$'echo "What\'s a female chicken called? Answer with only one word." | python3 -m tinygrad.llm --no_chat_template --model qwen3:0.6b | tee /dev/stderr | grep -i hen'
# NOTE: qwen is dumb and only knows about female chickens
# ****** Models Tests ******
+2
View File
@@ -1,10 +1,12 @@
from tinygrad import Tensor
from tinygrad.helpers import NUM_CPU_THREADS
from tinygrad.tensor import _to_np_dtype
from tinygrad.nn.onnx import OnnxRunner, OnnxValue
import numpy as np
import onnxruntime as ort
ort_options = ort.SessionOptions()
ort_options.log_severity_level = 3
ort_options.intra_op_num_threads = NUM_CPU_THREADS.value
def get_example_inputs(graph_inputs:dict[str, OnnxValue], config={}):
"""
+2 -1
View File
@@ -145,6 +145,7 @@ def main():
parser.add_argument("--serve", nargs='?', type=int, const=8000, metavar="PORT", help="Run OpenAI compatible API (optional port, default 8000)")
parser.add_argument("--warmup", action="store_true", help="warmup the JIT")
parser.add_argument("--benchmark", nargs='?', type=int, const=20, metavar="COUNT", help="Benchmark tok/s (optional count, default 20)")
parser.add_argument("--no_chat_template", action="store_true", help="Don't use the model's chat template, always use the fallback template")
args = parser.parse_args()
# load the model
@@ -160,7 +161,7 @@ def main():
# use the model's chat template if jinja2 is available (enables model-specific formatting)
template: jinja2.Template|FallbackTemplate = FallbackTemplate(tok)
if (ct := kv.get('tokenizer.chat_template')) is not None:
if not args.no_chat_template and (ct := kv.get('tokenizer.chat_template')) is not None:
try:
import jinja2
env = jinja2.Environment()