Skip to content

CLI flags

Every flag pion-server accepts — 70 of them, in the groups pion-server --help prints. This page is generated from the help printer in src/main.mojo at build time, so a flag is documented here exactly when the binary knows it. Defaults and profiles are explained in Configuration and profiles; the security-relevant flags (--bind, --requirepass*, --tenant) in the security model.

pion-server [flags]

Server

Flag What it does
-p, --port N listen port (default 1974; binary protocol on N+1 with --kvcache)
-w, --workers N OS-thread workers (default 1; capped at 4 on Apple Silicon). N > 1 gives each worker a PRIVATE keyspace and requires --independent-workers — see below.
--independent-workers acknowledge that -w N > 1 runs N independent keyspaces: a connection is bound to whichever worker won accept(), so a write acked on one connection is invisible to a read on another. Safe only when every client pins one connection (or shards keys itself). Pooled clients will read stale nils.
--profile PROF kv | vector | ai | full — KV-only / +HNSW / +AI / everything
--ns-prefix STR multi-tenant namespace gate for KV.PREFIX. / V. (single-tenant if empty)
--requirepass STR require AUTH <password> before serving any command (no auth if empty) NOTE: visible in ps; prefer the two below
--requirepass-file P read the password from file P (trailing newline stripped) or set PION_REQUIREPASS in the environment
--tenant NAME=PW per-tenant password; key-prefixes NAME's keyspace (repeatable; requires --requirepass) [gh #101]
--bind ADDR interface to listen on (default 127.0.0.1; 0.0.0.0 when a password is set). Applies to the RESP port, port+1, replication and gossip.
--no-wal disable WAL append (benchmark only; SIGKILL loses writes)
--wal-size MB WAL segment size (default 256); rotates instead of dropping
--wal-max-segments N sealed WAL segments to keep (default 32); 0 = never rotate
--wal-full-policy P refuse (default) | drop — what to do with keyspace writes once the WAL is full. refuse replies -MISCONF rather than ACKing writes that cannot be persisted.
--blob-threshold N values >= N bytes go to the file-backed blob tier (default 1048576)
--no-blob-tier keep large values on the anonymous heap (the pre-blob-tier behaviour)
--huge-pages / --no-huge-pages request 2 MB pages (Linux)
--affinity / --no-affinity pin workers to cores
-v, --version print version and exit
-h, --help this help

Diagnostics

Flag What it does
--crash-log PATH start/exit breadcrumb log (default pion-<port>.crash.log)
--status-file PATH 1 Hz liveness record: pid/uptime/RSS/ticks (default pion-<port>.status)
--no-crash-log disable both breadcrumb files
--rss-warn-pct N warn once when RSS crosses N% of physical RAM (default 70; >100 off)
--maxmemory SIZE refuse memory-growing writes (-OOM) above this RSS; bytes, k/kb/m/mb/g/gb or N% of RAM (default 0 = off; no eviction)

(supervised serving with auto-restart: scripts/pion-supervise.sh -- <server args>)

I/O backend

Flag What it does
--epoll epoll (Linux; best at P=1 / w=1 benchmarks)
--iouring io_uring (Linux default, best for production multi-connection)
--sqpoll io_uring SQPOLL — kernel-side polling, eliminates enter() syscall
--xdp AF_XDP zero-copy kernel bypass (Linux 5.4+, P=1 only)
--xdp-iface IF NIC for XDP attach (default eth0; alias: --xdp-interface)

Vector

Flag What it does
--dim N embedding dimension (default 1536)
--max-elements N HNSW capacity per worker (default 600,000)
--polarquant block-INT4 PolarQuant (recall ~0.96, -17% QPS vs INT8)
--turboquant block-INT3 + QJL TurboQuant (recall ~0.95, -39% QPS vs INT8)
--nanoquant block-INT2 NanoQuant (EXPERIMENTAL: recall ~0.46)
--gpu enable GPU search path (Metal on macOS)

Externalized attention / KV cache (Mac)

Flag What it does
--kvcache enable KV.PREFIX. / ATTEND. / V. / AI.ROUTE. / RAG.* + binary protocol on port+1
--moe-cache PATH enable MOE.EXPERT.* tier (path to MoE safetensors directory)
--moe-cache-mib N MoE tier LRU cache budget in MiB (default 1024)
--metal-attention native Metal SDPA, fp32 (Apple Silicon, no Python sidecar)
--metal-attention-fp16 fp16 kernel — matches vanilla mlx-lm precision
--cuda-attention native CUDA SDPA (Linux GPU). Routes ATTEND.PREFIX.{STORE,QUERY,QUERY_SPARSE,QUERY_SPARSE_AUTO} through cuda_kernels.cu
--fa-window N sliding-window SDPA: scan only last N tokens; 0 = full attention. Safe on Mistral SWA / Longformer; lossy on dense Llama / Gemma / GPT.

AI gateway / inference / embedding

Flag What it does
-G, --ai-gateway enable AI gateway (semantic cache + AI.* surface)
-M, --max-engine enable MAX engine path
--flare FLARE mid-generation retrieval (caps to -w 1)
--emb-enabled enable external embedding backend (default Ollama at 11434)
--emb-model NAME external embedding model (default nomic-embed-text)
--emb-host HOST external embedding host (default 127.0.0.1)
--emb-port N external embedding port (default 11434)
--emb-dim N external embedding width (default 768; MUST match the model)
--emb-query-prefix S instruction prefix for query embeds (asymmetric retrievers)
--emb-doc-prefix S instruction prefix for FT.ADDTEXT document embeds
--llm-enabled enable external LLM backend (MAX Serve / OpenAI-compat)
--llm-port N LLM backend port (default 8000)
--llm-model NAME LLM model id
--inference start in-process inference sidecar (PyTorch MiniLM-L6-v2)
--inference-socket P socket path (default /tmp/pion-<uid>-<port>.inference.sock)
--inference-emb-model NAME embedding model for sidecar
--inference-llm-model NAME LLM model for sidecar
--auto-embed auto-launch the embedding sidecar (default on)
--no-auto-embed disable the auto sidecar (use Ollama or --emb-enabled)
--no-auto-detect skip Ollama auto-probe at startup
--nle-embed macOS Apple NLEmbedding (512-dim, ANE-accelerated, no Python)

Cluster / replication

Flag What it does
--cluster enable cluster mode (gossip + CLUSTER NODES)
--cluster-host HOST advertised host for this node
--cluster-nodes LIST comma-separated host:port,... (including self)
--cluster-replica run as replica of --cluster-primary-*
--cluster-primary-host HOST primary's advertised host
--cluster-primary-port N primary's main port (replication uses primary_port + 10000)
--gossip-ping-ms N gossip PING interval (default 1000)

Generated at build time from src/main.mojo · _print_help() — edit the source, not this page.