Skip to content

PION.STATS and INFO — the value receipt

The server keeps a per-worker ledger of what the prefix cache actually did for you, and INFO reports resolved values — the real port, RSS, uptime and key counts — rather than placeholders.

The server keeps a per-worker ledger of what the prefix cache actually did for you:

PION.STATS            → map of 16 fields (RESP3 %-map; RESP2 flat array)
PION.STATS RESET      → +OK, counters zeroed (uptime is not a counter)
INFO                  → the same numbers under a `# Pion` section
Field Meaning
kvprefix_hits / kvprefix_misses KV.PREFIX.LOOKUP answers
kvprefix_tokens_served prefix tokens whose prefill was skipped (K-side layer-0 token count at hit time, or the TOKENS the client named on KV.PREFIX.LOOKUP)
kvprefix_bytes_served V.FETCH payload bytes delivered
prefill_seconds_avoided cumulative prefill time skipped, measured + estimated
prefill_seconds_avoided_measured the part backed by client-reported PREFILL_MS only
kvprefix_hits_measured hits credited with a reported time
semantic_hits / semantic_misses AI.SEMANTIC_CACHE GET and every other cache_get caller
moe_hits / moe_misses MOE.EXPERT.FETCH tier hits
vector_queries FT.SEARCH + FT.HYBRID answered

Two kinds of number, kept apart on purpose. PionPromptCache times its local cold prefill and sends it as KV.PREFIX.REGISTER ... PREFILL_MS <ms>; every later hit on that prefix is credited exactly that time — a receipt, not an estimate. A prefix registered without PREFILL_MS (a hand-rolled client, or a server restart, since the reported time is deliberately not persisted) is credited tokens × 555 µs, the per-token cost measured on Llama-3.2-1B-Instruct-4bit on an M-series Mac (2,022 tokens: 1,218 ms cold, 95 ms warm). Larger models cost more per token, so the estimate is conservative for them, and it is labelled an estimate wherever it appears.

The counters are per worker — they live in the worker that answered — so a connection pool spanning -w N workers reads each worker's own receipt. Formatting happens on demand in the reply; recording is a counter increment on an already-slow-path command, so nothing on the KV fast path changed.

Liveness, crash and status files

On by default. Every server writes two files into its working directory:

File Written Contents
pion-<port>.crash.log append, on start and on any catchable death one line per event
pion-<port>.status rewritten in place at 1 Hz fixed 512-byte key=value record
./pion-server --profile vector -w 1 -p 1974        # both files, default paths
./pion-server --crash-log /var/log/pion.log ...    # explicit crash-log path
./pion-server --status-file /run/pion.status ...   # explicit status path
./pion-server --no-crash-log ...                   # disable both
./pion-server --rss-warn-pct 60 ...                # warn earlier (default 70; >100 = off)

The two death shapes

Catchable (SIGSEGV/SIGBUS/SIGILL/SIGFPE/SIGABRT/SIGTERM/SIGINT/SIGQUIT/SIGXCPU/SIGHUP) — the handler appends a line, plus a backtrace for a fault, and then re-raises the default disposition, so the exit status and any core dump are exactly what they would have been:

PION START: pid=9659 version=0.815+7f42d4d port=1974 workers=1 unix_s=1785086858 total_ram_mb=16384
PION EXIT: signal SIGTERM (15) pid=9659 port=1974 uptime_s=17 rss_mb=67 rss_peak_mb=67 rss_pct=0 hb_age_s=0 ticks=283

A clean shutdown, a bind failure included, writes PION EXIT: clean shutdown.

Uncatchable — macOS jetsam and the Linux OOM killer send SIGKILL, which no handler in any language can observe. There is by construction no in-process trace. That is what the status file is for: it is the last thing the process said about itself, and it is written before the kill rather than after.

$ cat pion-1974.status
pion_status_version=1
state=running            ← still "running" though the process is gone == killed from outside
signal=0
signal_name=-
pid=9659
version=0.815+7f42d4d
port=1974
workers=1
start_unix_s=1785086858
uptime_s=48
heartbeat_unix_s=1785086906
rss_bytes=423706624      ← RSS as of ≤1 s before death — this is what confirms/denies jetsam
rss_peak_bytes=423706624
total_ram_bytes=17179869184
rss_pct=2
ticks=749                ← event-loop ticks; frozen value distinguishes "was serving" from "was wedged"

Reading it:

Evidence Verdict
PION EXIT: signal … present in-process death; the signal names the cause
No exit line + state=running + high rss_pct OS memory kill (jetsam / OOM). Confirm via §2's postmortem
No exit line + state=running + low rss_pct external kill -9, or the host went down
No exit line + a Mojo stack trace on stderr a runtime abort; read the server's stderr log
ticks unchanged across several seconds while the process lives wedged event loop, not a memory problem

A one-shot warning fires when RSS crosses --rss-warn-pct of physical RAM, so the operator hears about pressure before the kernel reaps the process:

PION WARN: rss_mb=11468 is 70% of total_ram_mb=16384 — OS may kill this process (jetsam/OOM) without a catchable signal

To refuse writes before that point instead of dying, set --maxmemory (see doc/configuration.md).

Cost

The heartbeat rides the existing 64-tick housekeeping block of every event loop — one non-atomic counter bump plus a clock_gettime; the RSS sample, status write and CAS happen once per second, performed by whichever worker wins the CAS. Nothing is added to the request path. Because the counter is per worker, a status file whose ticks stop advancing while the process is alive localises a hung worker.

The last command executed is not tracked: recording it would put a store on the zero-allocation fast path for every request.


FT.SEARCH against an index that does not exist returns an error, not an empty array:

-ERR no such index 'my_docs' - not created on this server (FT.CREATE first);
    if the server restarted, its index was not persisted

A client cannot distinguish "the store lost its index" from "the query matched nothing" when both are *0. This matters most after a restart: the WAL replays keys, but FT.CREATE is not a WAL entry, so a restarted server without a saved index file serves KV traffic while every vector query would otherwise return nothing.

Scope:

  • Fires only when no index exists at all — this worker has never seen FT.CREATE, has no nodes, and the shared cross-worker view is empty too.
  • A created-but-empty index still returns *0. An empty result on a live index is a legitimate answer, not an error.
  • A server holds one index; querying an index another FT.OPTIMIZE displaced errors and names both indexes (see doc/vector_engine.md § One index per server).
  • Both the KNN and BM25 sub-commands are covered, and pipelined clients stay frame-synchronised.

Client-side corollary: a retriever that catches ConnectionRefusedError and returns [] will still hide a dead server. Distinguish "no results" from "cannot reach the store" at your call site.


Everything else an operator needs: Running in production.