forked from tinygrad/tinygrad
Rename CPU_COUNT to NUM_CPU_THREADS so it can be overridden via env var. Default uses _get_cpu_count() which respects cgroup limits: - os.process_cpu_count() on Python 3.13+ - /sys/fs/cgroup/cpu.max on cgroup v2 - /sys/fs/cgroup/cpu/cpu.cfs_quota_us on cgroup v1 - os.sched_getaffinity(0) fallback Use NUM_CPU_THREADS.value in the dataloader instead of cpu_count(), and update export_model.py and all renderer references. Co-authored-by: teeny-runner <runner@teeny>
Each model should be a clean single file. They are imported from the top level `models` directory It should be capable of loading weights from the reference imp. We will focus on these 5 models: # Resnet50-v1.5 (classic) -- 8.2 GOPS/input # Retinanet # 3D UNET (upconvs) # RNNT # BERT-large (transformer) They are used in both the training and inference benchmark: https://mlcommons.org/en/training-normal-21/ https://mlcommons.org/en/inference-edge-30/ And we will submit to both. NOTE: we are Edge since we don't have ECC RAM