Where Pion does not help¶
The cases below are the ones where installing Pion costs you time and gives nothing back. Read this before you install.
This section exists because the demo that ships with this project once printed 0.86× — slower than doing nothing — and we shipped it that way for a while.
- It caches prefill, not decode. If your bottleneck is tokens-per-second once generation starts, this changes nothing.
- The reuse has to be real, and the prefix has to be long. Time to first
token from a separate process, only the prefix length varying: 34 tokens →
1.6×. 268 → 6.0×. 1,035 → 16.4×. 2,049 → 24×. The ratio holds up at short
prefixes; the saving does not — at 34 tokens it is about 19 ms, not worth a
second process. The demo defaults to 2,048, and
--prefix-tokens 33shows how little there is to save. - The sparse selector is NIAH-class. Factual QA over dense paragraphs needs a query-aware selector and collapses to 0.38×.
-w Nis N independent keyspaces, not one shared keyspace. The server refuses to start with-w N > 1unless you pass--independent-workers.- No TLS. Loopback by default; terminate TLS at a proxy.
- No eviction.
--maxmemoryrefuses writes with-OOMonce the process reaches the limit, but nothing is ever evicted to make room. - Cluster mode is single-node. Slot migration and replication exist; production multi-node does not.
- One maintainer, working evenings.
The numbers above are reproducible from the reproducers.