MOE.EXPERT.* — expert paging¶
Expert weights for MoE models live in a tiered cache (per-worker RAM LRU →
SSD → network) and are served over the wire, so a model larger than physical
memory runs at cache-hit latency. Enable with --moe-cache DIR --moe-cache-mib N.
Expert weights for MoE models live in a tiered cache (per-worker RAM LRU → SSD → network) and are served over the wire, so models larger than physical memory run at interactive cache-hit latency.
./pion-server --moe-cache /path/to/experts --moe-cache-mib 8192 -w 1
redis-cli -p 1974 MOE.EXPERT.LOAD <model> <path> # idempotent by basename
redis-cli -p 1974 MOE.EXPERT.FETCH <model> <layer> <expert> [NS <traffic-class>]
redis-cli -p 1974 MOE.EXPERT.HIST <model> [NS <name>|NSLIST] # access histograms
redis-cli -p 1974 MOE.EXPERT.PRUNE <model> <keep-set> # HIST-guided
- Measured: Gemma-4-26B-A4B (51.6 GB bf16), Phi-3.5-MoE (23.6 GB INT4), and Mixtral 8x7B (INT4) all run on a 16 GB Mac; ~5 ms cache-hit serving regardless of backing tier (60× vs cold network fetch).
- HIST-guided pruning: per-distribution access histograms (≤8 namespaces/model) drive prune sets that are safe across traffic mixtures; 25% prune reclaims ~10.4 GB on Gemma 4 (quality is workload-bound).
- Multi-model: 4 concurrent models per worker over a shared LRU. Four architectures validated (stacked bf16, per-expert INT4, stacked INT4).
Commands¶
Expert weights served from a tiered cache (per-worker RAM LRU → SSD → network)
so a MoE model larger than device RAM runs at interactive cache-hit latency.
Without --moe-cache every command answers -UNAVAILABLE naming the flag,
rather than a plausible empty result.
| Command | Pion path | Reply | Notes |
|---|---|---|---|
MOE.EXPERT.LOAD <model_id> <dir> |
SLOW | +OK / -UNAVAILABLE |
Loads a manifest and opens the tier |
MOE.EXPERT.FETCH <model_id> <layer> <expert> |
SLOW | bulk blob / -UNAVAILABLE |
~5 ms on a cache hit regardless of backing tier |
MOE.EXPERT.PREFETCH <model_id> <layer> <expert>... |
SLOW | +OK |
Asynchronous warm; returns before the fetch completes |
MOE.EXPERT.PIN <model_id> <layer> <expert> |
SLOW | +OK |
Protects an expert from LRU eviction |
MOE.EXPERT.UNPIN <model_id> <layer> <expert> |
SLOW | +OK |
|
MOE.EXPERT.INFO <model_id> |
SLOW | JSON / -UNAVAILABLE |
Per-model manifest |
| MOE.EXPERT.STATS | SLOW | JSON | Always available, even with the tier disabled — {"stage":1,"enabled":false,...} |
MOE.EXPERT.HIST <model_id> |
SLOW | JSON / -UNAVAILABLE |
Per-distribution access histograms |
MOE.EXPERT.PRUNE <model_id> <layer> <expert> [on] |
SLOW | +OK / -ERR |
HIST namespaces are per-distribution but PRUNE is global. A prune driven by one narrow histogram wrecks perplexity on other distributions — use the union-safe client policy |