Reproducers¶
The scripts behind the headline numbers, with the raw JSON they produced.
This is benchmarks/reproducers/README.md, included at build time.
The scripts behind Pion's headline numbers, so you can check them rather than take them on trust.
Read the prerequisites before running. Most of the cost is a model download, and five of these are timing measurements — a run on a loaded machine produces a wrong number, not an error. Close everything else first.
| Script | Backs | Needs |
|---|---|---|
cross_process_ttft.py |
Llama-3.2-1B time to first token from a separate process: 24× at a 2,049-token prefix (1,558 → 64.7 ms), and the prefix-length regime 1.6× → 24× | Apple Silicon, MLX, Llama-3.2-1B-Instruct-4bit, a Pion server with --kvcache |
sweep_qwen3_5_warm_ttft.py |
Qwen3.5-4B hybrid warm TTFT: 24.77× / 36.0× / 29.5× at 2K / 4K / 8K prefix | Apple Silicon, MLX, Qwen3.5-4B-MLX-4bit, a Pion server with --kvcache |
bench_qwen3_5_warm_ttft.py |
single-point version of the above | same |
stage1_hybrid_recall_bench.py |
hybrid retrieval cache recall | Apple Silicon, MLX, Llama-3.2-1B-Instruct-4bit, a Pion server |
stage0_hybrid_kv_injection.py |
the chunk-keyed cache's first evidence | Apple Silicon, MLX, Llama-3.2-1B-Instruct-4bit |
If a number does not reproduce¶
Open an issue with your hardware, the model build and the raw output. A number that only reproduces on our machine is a number we want to know about — several of these were wrong the first time.