CoreFlow-1.9M

A 1.9 MB int8 language model that decodes 3,400+ tokens/second on a single CPU core โ€” trained from scratch in pure PyTorch, served by a dependency-free C11 runtime. No Hugging Face wrappers, no llama.cpp, no GPU.

Measurements

Ryzen 5 3400G, single thread, 15 s sustained โ€” every number links to a JSON record in the GitHub repo:

Metric Result
Single-core decode 3,405 tok/s ยท p50 0.261 ms ยท p99 0.704 ms
4 streams ร— 4 tokens in flight (60 s) 5,997 tok/s
Model size on disk 1,965,250 B (1.91 M params, int8)
Working set vs L3 49.8% of 4 MB L3 โ€” cache-resident by design
C vs PyTorch perplexity 26.8407 vs 26.8188 โ€” 0.082% rel. delta
Token-level parity vs PyTorch 99.41% match
L3 eviction penalty (evict_read) 3,405 โ†’ 1,681 tok/s (2.03ร—)

Files

  • model_p3.i8 โ€” int8 weights + embedded vocabulary, self-contained runtime format (1,965,250 bytes)
  • val.bin โ€” 240 KB validation slice used by the benchmark command

Run

Requires Windows x64 and MSVC Build Tools 2022 (C11). Clone the GitHub repo for the 5-command quickstart:

phase0\build_p0.bat
phase0\phase0.exe phase0\model_p3.i8 info
phase0\phase0.exe phase0\model_p3.i8 bench data\val.bin 15.0 65

Design

Size the model to the cache: keep the int8 weight footprint at or below half of L3 (2.09 MB working set against 4 MB L3) and decode becomes a cache-streaming problem instead of a DRAM problem. Conv-heavy hybrid architecture (4 gated short-conv blocks + 2 grouped-query attention blocks) keeps the KV cache tiny; hand-written AVX2 int8 kernels do the rest. Full rationale in BRIEF.md.

Limitations

  • 1.9 M parameters: output quality is weak. This is a reproducible inference design study, not a chatbot.
  • Windows x64 only today; Linux/macOS build is tracked as issue #2 on GitHub.

License

MIT

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support