LangGraph · AutoGen · LlamaIndex · exo¶
Four small Apache-2.0 packages that plug Pion into frameworks people already use. Each README is included at build time.
pion-langgraph — agent state persistence¶
LangGraph checkpoint saver backed by Pion — deterministic low-latency agent-state persistence.
PionSaver implements LangGraph's BaseCheckpointSaver against a running Pion server. Uses only RESP2-stable commands (HSET / HGET / HGETALL, single-key ZADD / ZREVRANGE) — no RedisJSON, no Lua scripting, no Sentinel. Works against any Pion deployment, single-worker or sharded.
Install¶
pip install -e pion-langgraph/
# optional: dev tools + LangGraph for the test suite
pip install -e pion-langgraph/[dev]
Requires a running Pion server:
Usage¶
from pion_langgraph import PionSaver
saver = PionSaver(host="127.0.0.1", port=1974, key_prefix="lgcp")
graph = workflow.compile(checkpointer=saver)
result = graph.invoke({"input": "hello"})
# Saver persists every superstep to Pion; resume on next invocation
# is a single HGETALL round-trip (~30 µs hot path).
Docs¶
- Main project:
../README.md
License¶
Apache-2.0 — see LICENSE in this directory. Pion's client packages are
deliberately permissive so they can be vendored into any stack; the Pion server
is Apache-2.0 too, with one closed binary library for its tuned vector kernels —
see the top-level LICENSE.
Which parts of Pion are under which licence: doc/licensing.md.
pion-autogen — semantic agent memory¶
AutoGen memory backend powered by Pion's HNSW vector index — semantic recall for multi-agent systems.
PionMemoryStore implements AutoGen Core's Memory interface against a running Pion server, so any AutoGen agent can persist and recall facts with sub-millisecond cosine-similarity lookup. No RedisJSON dependency; uses the wire-compatible RESP2 path.
Install¶
pip install -e pion-autogen/
# optional: dev tools + AutoGen for the test suite
pip install -e pion-autogen/[dev]
Requires a running Pion server on the host/port you pass:
Usage¶
from pion_autogen import PionMemoryStore
mem = PionMemoryStore(host="127.0.0.1", port=1974, index_name="agent_a",
dimensions=384, embed_provider="ollama")
await mem.add("the user lives in Bratislava and prefers metric units")
hits = await mem.query("what city is the user in?")
print(hits[0].content) # → "the user lives in Bratislava ..."
Supports embed_provider = "mock" (deterministic, for tests), "openai", or "ollama". Embedding dim defaults to 384 (MiniLM-L6-v2) when paired with Pion's --auto-embed sidecar.
Docs¶
- Main project:
../README.md - Vector engine internals:
../doc/vector_engine.md - Substrate this backend uses:
AI.MEMORY.*+FT.SEARCH(see../doc/codebase_search.md)
License¶
Apache-2.0 — see LICENSE in this directory. Pion's client packages are
deliberately permissive so they can be vendored into any stack; the Pion server
is Apache-2.0 too, with one closed binary library for its tuned vector kernels —
see the top-level LICENSE.
Which parts of Pion are under which licence: doc/licensing.md.
pion-llamaindex — vector store for RAG¶
LlamaIndex VectorStore backed by Pion's HNSW index — sub-millisecond RAG retrieval.
PionVectorStore plugs into LlamaIndex's vector-store interface so any LlamaIndex pipeline (RAG, agents, query engines) can use Pion's quantization-aware HNSW (INT4 / INT3 / INT2 variants) for retrieval without any LlamaIndex protocol changes.
Install¶
pip install -e pion-llamaindex/
# optional: dev tools + LlamaIndex for the test suite
pip install -e pion-llamaindex/[dev]
Requires a running Pion server:
Usage¶
from pion_llamaindex import PionVectorStore
from llama_index.core import VectorStoreIndex, StorageContext, Document
store = PionVectorStore(host="127.0.0.1", port=1974,
index_name="docs", dimensions=1536)
storage = StorageContext.from_defaults(vector_store=store)
idx = VectorStoreIndex.from_documents([Document(text="...")], storage_context=storage)
print(idx.as_query_engine().query("what is X?"))
Pion handles the embedding-similarity search (cosine, INT8 SIMD); LlamaIndex still owns chunking, prompt assembly, and the LLM call. Default dimensions=1536 matches OpenAI's text-embedding-3-small; override for other embedders.
Docs¶
- Main project:
../README.md - Vector engine internals + quantization variants:
../doc/vector_engine.md
License¶
Apache-2.0 — see LICENSE in this directory. Pion's client packages are
deliberately permissive so they can be vendored into any stack; the Pion server
is Apache-2.0 too, with one closed binary library for its tuned vector kernels —
see the top-level LICENSE.
Which parts of Pion are under which licence: doc/licensing.md.
pion-exo — Mac cluster inference hook¶
Pion's side of both hooks is finished and tested; the other side is not. exo exposes no attention-hook API upstream, and the vllm-pion v1 connector branch is unmerged — so these are integration substrate, not something you can drop into a running exo cluster today.
# exo integration -- distributed inference with V offloading
# gpu_attention routes through PionPromptCache(stage2=True),
# which uses ATTEND.PREFIX.STORE/QUERY against Pion's native Metal SDPA.
# Start Pion with: ./pion-server --kvcache --metal-attention -w 1
from pion_exo import PionAttentionHook # pip install -e pion-exo/
hook = PionAttentionHook(pion_host="192.168.1.100", mode="gpu_attention")
# vllm-mlx integration -- attention backend with Pion V offloading
from pion_vllm_mlx import PionPromptCache # pip install -e pion-vllm-mlx/
Pion attention hook for exo distributed inference — V offloading and Stage-2 Metal SDPA on Apple Silicon clusters.
Two operating modes:
| Mode | What it does | When to use |
|---|---|---|
v_offload |
Offloads the V tensor of every attention layer to Pion's V-store (INT8 / turbo4 / fp16). Frees Mac unified memory for activations + KVQ. | Long-context generation on a single Mac, or 4-tier shared corpora across a small cluster. |
gpu_attention |
All of v_offload, plus the warm forward runs Metal SDPA directly via pion-vllm-mlx's PionPromptCache Stage-2 path. |
Production cluster serving with cross-host prefix sharing and the 326×-warm-TTFT (Gemma-4-E2B 64K NIAH) substrate. |
Install¶
pip install -e pion-exo/
# optional: gpu_attention mode pulls pion-vllm-mlx
pip install -e ../pion-vllm-mlx/
Requires a running Pion server with V-store + KV-prefix enabled:
For Mac-cluster topologies (Stage-3 peers=[...]) every node must run its own Pion server on port 1974.
Usage¶
from pion_exo import PionAttentionHook
# Single-host V offload — drops V cache to Pion, keeps K + activations in Mac RAM.
hook = PionAttentionHook(pion_host="127.0.0.1", pion_port=1974, mode="v_offload")
# Mac cluster — sticky-route sessions across nodes by blake2b(session_id) mod N.
hook = PionAttentionHook(
peers=[("mac-mini-1", 1974), ("mac-mini-2", 1974), ("mac-mini-3", 1974)],
mode="gpu_attention",
)
The hook is driven directly — on_prefill per layer during prefill, then
on_decode_attention per layer per step:
hook.on_prefill(session_id, layer_id, K=K, V=V) # push K/V to Pion
out = hook.on_decode_attention(session_id, layer_id, # attention output
Q=Q, K_local=K, top_k=64)
hook.drop_session(session_id) # release when done
Wiring it into exo¶
exo does not currently expose an attention-hook API — there is no
register_attention_hook or equivalent extension point upstream (checked
2026-08-21 against exo-explore/exo). Two routes, neither of which lives in
this package yet:
- Monkey-patch exo's attention path from outside, the way
pion-vllm-mlxdoes for mlx-lm inmlx_lm_patch.py. Same shape: intercept the attention call, delegate to the hook. This needs no upstream change. - Land a hook point upstream. Cleaner and durable, but it is a PR to someone else's project on their timeline.
Until one of those exists, the hook is usable — and tested — by a caller that
drives it directly, which is what the tests in tests/ do.
Docs¶
- Main project:
../README.md - Stage-2 lane design:
../doc/shared_kv_cache.md
License¶
Apache-2.0 — see LICENSE in this directory. Pion's client packages are
deliberately permissive so they can be vendored into any stack; the Pion server
is Apache-2.0 too, with one closed binary library for its tuned vector kernels —
see the top-level LICENSE.
Which parts of Pion are under which licence: doc/licensing.md.