PION.STATS and INFO — the value receipt¶
The server keeps a per-worker ledger of what the prefix cache actually did
for you, and INFO reports resolved values — the real port, RSS, uptime and
key counts — rather than placeholders.
The server keeps a per-worker ledger of what the prefix cache actually did for you:
PION.STATS → map of 16 fields (RESP3 %-map; RESP2 flat array)
PION.STATS RESET → +OK, counters zeroed (uptime is not a counter)
INFO → the same numbers under a `# Pion` section
| Field | Meaning |
|---|---|
kvprefix_hits / kvprefix_misses |
KV.PREFIX.LOOKUP answers |
kvprefix_tokens_served |
prefix tokens whose prefill was skipped (K-side layer-0 token count at hit time, or the TOKENS the client named on KV.PREFIX.LOOKUP) |
kvprefix_bytes_served |
V.FETCH payload bytes delivered |
prefill_seconds_avoided |
cumulative prefill time skipped, measured + estimated |
prefill_seconds_avoided_measured |
the part backed by client-reported PREFILL_MS only |
kvprefix_hits_measured |
hits credited with a reported time |
semantic_hits / semantic_misses |
AI.SEMANTIC_CACHE GET and every other cache_get caller |
moe_hits / moe_misses |
MOE.EXPERT.FETCH tier hits |
vector_queries |
FT.SEARCH + FT.HYBRID answered |
Two kinds of number, kept apart on purpose. PionPromptCache times its local
cold prefill and sends it as KV.PREFIX.REGISTER ... PREFILL_MS <ms>; every
later hit on that prefix is credited exactly that time — a receipt, not an
estimate. A prefix registered without PREFILL_MS (a hand-rolled client, or a
server restart, since the reported time is deliberately not persisted) is
credited tokens × 555 µs, the per-token cost measured on
Llama-3.2-1B-Instruct-4bit on an M-series Mac (2,022 tokens: 1,218 ms cold,
95 ms warm). Larger models cost more per token, so the estimate is conservative
for them, and it is labelled an estimate wherever it appears.
The counters are per worker — they live in the worker that answered — so a
connection pool spanning -w N workers reads each worker's own receipt.
Formatting happens on demand in the reply; recording is a counter increment on
an already-slow-path command, so nothing on the KV fast path changed.
Liveness, crash and status files¶
On by default. Every server writes two files into its working directory:
| File | Written | Contents |
|---|---|---|
pion-<port>.crash.log |
append, on start and on any catchable death | one line per event |
pion-<port>.status |
rewritten in place at 1 Hz | fixed 512-byte key=value record |
./pion-server --profile vector -w 1 -p 1974 # both files, default paths
./pion-server --crash-log /var/log/pion.log ... # explicit crash-log path
./pion-server --status-file /run/pion.status ... # explicit status path
./pion-server --no-crash-log ... # disable both
./pion-server --rss-warn-pct 60 ... # warn earlier (default 70; >100 = off)
The two death shapes¶
Catchable (SIGSEGV/SIGBUS/SIGILL/SIGFPE/SIGABRT/SIGTERM/SIGINT/SIGQUIT/SIGXCPU/SIGHUP) — the handler appends a line, plus a backtrace for a fault, and then re-raises the default disposition, so the exit status and any core dump are exactly what they would have been:
PION START: pid=9659 version=0.815+7f42d4d port=1974 workers=1 unix_s=1785086858 total_ram_mb=16384
PION EXIT: signal SIGTERM (15) pid=9659 port=1974 uptime_s=17 rss_mb=67 rss_peak_mb=67 rss_pct=0 hb_age_s=0 ticks=283
A clean shutdown, a bind failure included, writes PION EXIT: clean shutdown.
Uncatchable — macOS jetsam and the Linux OOM killer send SIGKILL, which no handler in any language can observe. There is by construction no in-process trace. That is what the status file is for: it is the last thing the process said about itself, and it is written before the kill rather than after.
$ cat pion-1974.status
pion_status_version=1
state=running ← still "running" though the process is gone == killed from outside
signal=0
signal_name=-
pid=9659
version=0.815+7f42d4d
port=1974
workers=1
start_unix_s=1785086858
uptime_s=48
heartbeat_unix_s=1785086906
rss_bytes=423706624 ← RSS as of ≤1 s before death — this is what confirms/denies jetsam
rss_peak_bytes=423706624
total_ram_bytes=17179869184
rss_pct=2
ticks=749 ← event-loop ticks; frozen value distinguishes "was serving" from "was wedged"
Reading it:
| Evidence | Verdict |
|---|---|
PION EXIT: signal … present |
in-process death; the signal names the cause |
No exit line + state=running + high rss_pct |
OS memory kill (jetsam / OOM). Confirm via §2's postmortem |
No exit line + state=running + low rss_pct |
external kill -9, or the host went down |
| No exit line + a Mojo stack trace on stderr | a runtime abort; read the server's stderr log |
ticks unchanged across several seconds while the process lives |
wedged event loop, not a memory problem |
A one-shot warning fires when RSS crosses --rss-warn-pct of physical RAM, so
the operator hears about pressure before the kernel reaps the process:
PION WARN: rss_mb=11468 is 70% of total_ram_mb=16384 — OS may kill this process (jetsam/OOM) without a catchable signal
To refuse writes before that point instead of dying, set --maxmemory (see
doc/configuration.md).
Cost¶
The heartbeat rides the existing 64-tick housekeeping block of every event
loop — one non-atomic counter bump plus a clock_gettime; the RSS sample,
status write and CAS happen once per second, performed by whichever worker wins
the CAS. Nothing is added to the request path. Because the counter is per
worker, a status file whose ticks stop advancing while the process is alive
localises a hung worker.
The last command executed is not tracked: recording it would put a store on the zero-allocation fast path for every request.
FT.SEARCH against an index that does not exist returns an error, not an empty
array:
-ERR no such index 'my_docs' - not created on this server (FT.CREATE first);
if the server restarted, its index was not persisted
A client cannot distinguish "the store lost its index" from "the query matched
nothing" when both are *0. This matters most after a restart: the WAL replays
keys, but FT.CREATE is not a WAL entry, so a restarted server without a saved
index file serves KV traffic while every vector query would otherwise return
nothing.
Scope:
- Fires only when no index exists at all — this worker has never seen
FT.CREATE, has no nodes, and the shared cross-worker view is empty too. - A created-but-empty index still returns
*0. An empty result on a live index is a legitimate answer, not an error. - A server holds one index; querying an index another
FT.OPTIMIZEdisplaced errors and names both indexes (seedoc/vector_engine.md§ One index per server). - Both the KNN and
BM25sub-commands are covered, and pipelined clients stay frame-synchronised.
Client-side corollary: a retriever that catches ConnectionRefusedError and
returns [] will still hide a dead server. Distinguish "no results" from "cannot
reach the store" at your call site.
Everything else an operator needs: Running in production.