Skip to content

LangGraph · AutoGen · LlamaIndex · exo

Four small Apache-2.0 packages that plug Pion into frameworks people already use. Each README is included at build time.

pion-langgraph — agent state persistence

LangGraph checkpoint saver backed by Pion — deterministic low-latency agent-state persistence.

PionSaver implements LangGraph's BaseCheckpointSaver against a running Pion server. Uses only RESP2-stable commands (HSET / HGET / HGETALL, single-key ZADD / ZREVRANGE) — no RedisJSON, no Lua scripting, no Sentinel. Works against any Pion deployment, single-worker or sharded.

Install

pip install -e pion-langgraph/
# optional: dev tools + LangGraph for the test suite
pip install -e pion-langgraph/[dev]

Requires a running Pion server:

./pion-server --kvcache -w 1   # binds 127.0.0.1:1974

Usage

from pion_langgraph import PionSaver

saver = PionSaver(host="127.0.0.1", port=1974, key_prefix="lgcp")
graph = workflow.compile(checkpointer=saver)
result = graph.invoke({"input": "hello"})
# Saver persists every superstep to Pion; resume on next invocation
# is a single HGETALL round-trip (~30 µs hot path).

Docs

License

Apache-2.0 — see LICENSE in this directory. Pion's client packages are deliberately permissive so they can be vendored into any stack; the Pion server is Apache-2.0 too, with one closed binary library for its tuned vector kernels — see the top-level LICENSE.

Which parts of Pion are under which licence: doc/licensing.md.

pion-autogen — semantic agent memory

AutoGen memory backend powered by Pion's HNSW vector index — semantic recall for multi-agent systems.

PionMemoryStore implements AutoGen Core's Memory interface against a running Pion server, so any AutoGen agent can persist and recall facts with sub-millisecond cosine-similarity lookup. No RedisJSON dependency; uses the wire-compatible RESP2 path.

Install

pip install -e pion-autogen/
# optional: dev tools + AutoGen for the test suite
pip install -e pion-autogen/[dev]

Requires a running Pion server on the host/port you pass:

./pion-server --kvcache -w 1   # binds 127.0.0.1:1974

Usage

from pion_autogen import PionMemoryStore

mem = PionMemoryStore(host="127.0.0.1", port=1974, index_name="agent_a",
                       dimensions=384, embed_provider="ollama")
await mem.add("the user lives in Bratislava and prefers metric units")
hits = await mem.query("what city is the user in?")
print(hits[0].content)   # → "the user lives in Bratislava ..."

Supports embed_provider = "mock" (deterministic, for tests), "openai", or "ollama". Embedding dim defaults to 384 (MiniLM-L6-v2) when paired with Pion's --auto-embed sidecar.

Docs

License

Apache-2.0 — see LICENSE in this directory. Pion's client packages are deliberately permissive so they can be vendored into any stack; the Pion server is Apache-2.0 too, with one closed binary library for its tuned vector kernels — see the top-level LICENSE.

Which parts of Pion are under which licence: doc/licensing.md.

pion-llamaindex — vector store for RAG

LlamaIndex VectorStore backed by Pion's HNSW index — sub-millisecond RAG retrieval.

PionVectorStore plugs into LlamaIndex's vector-store interface so any LlamaIndex pipeline (RAG, agents, query engines) can use Pion's quantization-aware HNSW (INT4 / INT3 / INT2 variants) for retrieval without any LlamaIndex protocol changes.

Install

pip install -e pion-llamaindex/
# optional: dev tools + LlamaIndex for the test suite
pip install -e pion-llamaindex/[dev]

Requires a running Pion server:

./pion-server --profile vector -w 1   # binds 127.0.0.1:1974

Usage

from pion_llamaindex import PionVectorStore
from llama_index.core import VectorStoreIndex, StorageContext, Document

store = PionVectorStore(host="127.0.0.1", port=1974,
                        index_name="docs", dimensions=1536)
storage = StorageContext.from_defaults(vector_store=store)
idx = VectorStoreIndex.from_documents([Document(text="...")], storage_context=storage)
print(idx.as_query_engine().query("what is X?"))

Pion handles the embedding-similarity search (cosine, INT8 SIMD); LlamaIndex still owns chunking, prompt assembly, and the LLM call. Default dimensions=1536 matches OpenAI's text-embedding-3-small; override for other embedders.

Docs

License

Apache-2.0 — see LICENSE in this directory. Pion's client packages are deliberately permissive so they can be vendored into any stack; the Pion server is Apache-2.0 too, with one closed binary library for its tuned vector kernels — see the top-level LICENSE.

Which parts of Pion are under which licence: doc/licensing.md.

pion-exo — Mac cluster inference hook

Pion's side of both hooks is finished and tested; the other side is not. exo exposes no attention-hook API upstream, and the vllm-pion v1 connector branch is unmerged — so these are integration substrate, not something you can drop into a running exo cluster today.

# exo integration -- distributed inference with V offloading
# gpu_attention routes through PionPromptCache(stage2=True),
# which uses ATTEND.PREFIX.STORE/QUERY against Pion's native Metal SDPA.
# Start Pion with: ./pion-server --kvcache --metal-attention -w 1
from pion_exo import PionAttentionHook    # pip install -e pion-exo/
hook = PionAttentionHook(pion_host="192.168.1.100", mode="gpu_attention")

# vllm-mlx integration -- attention backend with Pion V offloading
from pion_vllm_mlx import PionPromptCache           # pip install -e pion-vllm-mlx/

Pion attention hook for exo distributed inference — V offloading and Stage-2 Metal SDPA on Apple Silicon clusters.

Two operating modes:

Mode What it does When to use
v_offload Offloads the V tensor of every attention layer to Pion's V-store (INT8 / turbo4 / fp16). Frees Mac unified memory for activations + KVQ. Long-context generation on a single Mac, or 4-tier shared corpora across a small cluster.
gpu_attention All of v_offload, plus the warm forward runs Metal SDPA directly via pion-vllm-mlx's PionPromptCache Stage-2 path. Production cluster serving with cross-host prefix sharing and the 326×-warm-TTFT (Gemma-4-E2B 64K NIAH) substrate.

Install

pip install -e pion-exo/
# optional: gpu_attention mode pulls pion-vllm-mlx
pip install -e ../pion-vllm-mlx/

Requires a running Pion server with V-store + KV-prefix enabled:

./pion-server --kvcache --metal-attention -w 1

For Mac-cluster topologies (Stage-3 peers=[...]) every node must run its own Pion server on port 1974.

Usage

from pion_exo import PionAttentionHook

# Single-host V offload — drops V cache to Pion, keeps K + activations in Mac RAM.
hook = PionAttentionHook(pion_host="127.0.0.1", pion_port=1974, mode="v_offload")

# Mac cluster — sticky-route sessions across nodes by blake2b(session_id) mod N.
hook = PionAttentionHook(
    peers=[("mac-mini-1", 1974), ("mac-mini-2", 1974), ("mac-mini-3", 1974)],
    mode="gpu_attention",
)

The hook is driven directly — on_prefill per layer during prefill, then on_decode_attention per layer per step:

hook.on_prefill(session_id, layer_id, K=K, V=V)          # push K/V to Pion
out = hook.on_decode_attention(session_id, layer_id,      # attention output
                               Q=Q, K_local=K, top_k=64)
hook.drop_session(session_id)                             # release when done
Wiring it into exo

exo does not currently expose an attention-hook API — there is no register_attention_hook or equivalent extension point upstream (checked 2026-08-21 against exo-explore/exo). Two routes, neither of which lives in this package yet:

  1. Monkey-patch exo's attention path from outside, the way pion-vllm-mlx does for mlx-lm in mlx_lm_patch.py. Same shape: intercept the attention call, delegate to the hook. This needs no upstream change.
  2. Land a hook point upstream. Cleaner and durable, but it is a PR to someone else's project on their timeline.

Until one of those exists, the hook is usable — and tested — by a caller that drives it directly, which is what the tests in tests/ do.

Docs

License

Apache-2.0 — see LICENSE in this directory. Pion's client packages are deliberately permissive so they can be vendored into any stack; the Pion server is Apache-2.0 too, with one closed binary library for its tuned vector kernels — see the top-level LICENSE.

Which parts of Pion are under which licence: doc/licensing.md.