Skip to content

MOE.EXPERT.* — expert paging

Expert weights for MoE models live in a tiered cache (per-worker RAM LRU → SSD → network) and are served over the wire, so a model larger than physical memory runs at cache-hit latency. Enable with --moe-cache DIR --moe-cache-mib N.

Expert weights for MoE models live in a tiered cache (per-worker RAM LRU → SSD → network) and are served over the wire, so models larger than physical memory run at interactive cache-hit latency.

./pion-server --moe-cache /path/to/experts --moe-cache-mib 8192 -w 1
redis-cli -p 1974 MOE.EXPERT.LOAD <model> <path>      # idempotent by basename
redis-cli -p 1974 MOE.EXPERT.FETCH <model> <layer> <expert> [NS <traffic-class>]
redis-cli -p 1974 MOE.EXPERT.HIST <model> [NS <name>|NSLIST]  # access histograms
redis-cli -p 1974 MOE.EXPERT.PRUNE <model> <keep-set>          # HIST-guided
  • Measured: Gemma-4-26B-A4B (51.6 GB bf16), Phi-3.5-MoE (23.6 GB INT4), and Mixtral 8x7B (INT4) all run on a 16 GB Mac; ~5 ms cache-hit serving regardless of backing tier (60× vs cold network fetch).
  • HIST-guided pruning: per-distribution access histograms (≤8 namespaces/model) drive prune sets that are safe across traffic mixtures; 25% prune reclaims ~10.4 GB on Gemma 4 (quality is workload-bound).
  • Multi-model: 4 concurrent models per worker over a shared LRU. Four architectures validated (stacked bf16, per-expert INT4, stacked INT4).

Commands

Expert weights served from a tiered cache (per-worker RAM LRU → SSD → network) so a MoE model larger than device RAM runs at interactive cache-hit latency. Without --moe-cache every command answers -UNAVAILABLE naming the flag, rather than a plausible empty result.

Command Pion path Reply Notes
MOE.EXPERT.LOAD <model_id> <dir> SLOW +OK / -UNAVAILABLE Loads a manifest and opens the tier
MOE.EXPERT.FETCH <model_id> <layer> <expert> SLOW bulk blob / -UNAVAILABLE ~5 ms on a cache hit regardless of backing tier
MOE.EXPERT.PREFETCH <model_id> <layer> <expert>... SLOW +OK Asynchronous warm; returns before the fetch completes
MOE.EXPERT.PIN <model_id> <layer> <expert> SLOW +OK Protects an expert from LRU eviction
MOE.EXPERT.UNPIN <model_id> <layer> <expert> SLOW +OK
MOE.EXPERT.INFO <model_id> SLOW JSON / -UNAVAILABLE Per-model manifest
MOE.EXPERT.STATS SLOW JSON Always available, even with the tier disabled — {"stage":1,"enabled":false,...}
MOE.EXPERT.HIST <model_id> SLOW JSON / -UNAVAILABLE Per-distribution access histograms
MOE.EXPERT.PRUNE <model_id> <layer> <expert> [on] SLOW +OK / -ERR HIST namespaces are per-distribution but PRUNE is global. A prune driven by one narrow histogram wrecks perplexity on other distributions — use the union-safe client policy