CoreFlow-1.9M
A 1.9 MB int8 language model that decodes 3,400+ tokens/second on a single CPU core โ trained from scratch in pure PyTorch, served by a dependency-free C11 runtime. No Hugging Face wrappers, no llama.cpp, no GPU.
Measurements
Ryzen 5 3400G, single thread, 15 s sustained โ every number links to a JSON record in the GitHub repo:
| Metric | Result |
|---|---|
| Single-core decode | 3,405 tok/s ยท p50 0.261 ms ยท p99 0.704 ms |
| 4 streams ร 4 tokens in flight (60 s) | 5,997 tok/s |
| Model size on disk | 1,965,250 B (1.91 M params, int8) |
| Working set vs L3 | 49.8% of 4 MB L3 โ cache-resident by design |
| C vs PyTorch perplexity | 26.8407 vs 26.8188 โ 0.082% rel. delta |
| Token-level parity vs PyTorch | 99.41% match |
| L3 eviction penalty (evict_read) | 3,405 โ 1,681 tok/s (2.03ร) |
Files
model_p3.i8โ int8 weights + embedded vocabulary, self-contained runtime format (1,965,250 bytes)val.binโ 240 KB validation slice used by the benchmark command
Run
Requires Windows x64 and MSVC Build Tools 2022 (C11). Clone the GitHub repo for the 5-command quickstart:
phase0\build_p0.bat
phase0\phase0.exe phase0\model_p3.i8 info
phase0\phase0.exe phase0\model_p3.i8 bench data\val.bin 15.0 65
Design
Size the model to the cache: keep the int8 weight footprint at or below half of L3 (2.09 MB working set against 4 MB L3) and decode becomes a cache-streaming problem instead of a DRAM problem. Conv-heavy hybrid architecture (4 gated short-conv blocks + 2 grouped-query attention blocks) keeps the KV cache tiny; hand-written AVX2 int8 kernels do the rest. Full rationale in BRIEF.md.
Limitations
- 1.9 M parameters: output quality is weak. This is a reproducible inference design study, not a chatbot.
- Windows x64 only today; Linux/macOS build is tracked as issue #2 on GitHub.
License
MIT