Skip to content

Where Pion does not help

The cases below are the ones where installing Pion costs you time and gives nothing back. Read this before you install.

This section exists because the demo that ships with this project once printed 0.86× — slower than doing nothing — and we shipped it that way for a while.

  • It caches prefill, not decode. If your bottleneck is tokens-per-second once generation starts, this changes nothing.
  • The reuse has to be real, and the prefix has to be long. Time to first token from a separate process, only the prefix length varying: 34 tokens → 1.6×. 268 → 6.0×. 1,035 → 16.4×. 2,049 → 24×. The ratio holds up at short prefixes; the saving does not — at 34 tokens it is about 19 ms, not worth a second process. The demo defaults to 2,048, and --prefix-tokens 33 shows how little there is to save.
  • The sparse selector is NIAH-class. Factual QA over dense paragraphs needs a query-aware selector and collapses to 0.38×.
  • -w N is N independent keyspaces, not one shared keyspace. The server refuses to start with -w N > 1 unless you pass --independent-workers.
  • No TLS. Loopback by default; terminate TLS at a proxy.
  • No eviction. --maxmemory refuses writes with -OOM once the process reaches the limit, but nothing is ever evicted to make room.
  • Cluster mode is single-node. Slot migration and replication exist; production multi-node does not.
  • One maintainer, working evenings.

The numbers above are reproducible from the reproducers.