A CVM-based LLM router. Operators plug in upstream LLM providers (Anthropic, OpenAI, …) and Muxll serves their models over the CVM JSON-RPC interface — MCP-over-Nostr — so any Nostr-native client can call them. Streamed responses ride CEP-41 open-ended streams.
This repository currently holds a proof of concept: chat completions (with and without streaming) and model listing. Payments, quotas, and provider config files are intentionally out of scope for now.
Nostr client ──CVM/MCP JSON-RPC──▶ Muxll CVM server ──▶ pi-ai ──▶ upstream LLM
(Nostr relays) McpServer + NostrServerTransport
- pi-ai (
@earendil-works/pi-ai) is the unified provider layer: one streaming API across many providers, with built-in model catalogs. - CVM SDK (
@contextvm/sdk) exposes the MCP server over Nostr viaNostrServerTransport, with CEP-41 open streams enabled for token-by-token streaming. - MCP SDK (
@contextvm/mcp-sdk) provides theMcpServerthat registers the RPC tools.
Two MCP tools are exposed (called via the standard tools/call method):
Returns every model across configured providers. Shape follows OpenAI
GET /models (plus router-useful capability fields).
// arguments
{
"model": "anthropic/claude-opus-4-5", // provider/id tag (bare id = first match)
"messages": [
{ "role": "system", "content": "You are helpful." },
{ "role": "user", "content": "Hello" }
// assistant turns may carry `tool_calls`; `tool` turns carry
// `tool_call_id` + `content` (the result) to close a tool loop.
// user `content` may be an array of {type:"text"} / {type:"image_url"}
// parts for multimodal input (data-URL or http(s) images).
],
"tools": [ // optional, OpenAI function-tool shape
{ "type": "function",
"function": { "name": "get_weather",
"description": "Get current weather",
"parameters": { "type": "object",
"properties": { "location": { "type": "string" } } } } }
],
"tool_choice": "auto", // optional: "auto" | "none" | "required" | {type:"function",function:{name}} (normalized per provider)
"stream": false, // true → stream chat.completion.chunk objects over CEP-41
"temperature": 0.7, // optional
"max_tokens": 1024 // optional
}
// structuredContent (non-streaming) — OpenAI CreateChatCompletionResponse
{
"id": "chatcmpl-…", "object": "chat.completion", "created": 1730000000,
"model": "claude-opus-4-5",
"choices": [
{ "index": 0,
"message": { "role": "assistant", "content": "Hi there!" },
"finish_reason": "stop" }
],
"usage": {
"prompt_tokens": 12, "completion_tokens": 3, "total_tokens": 15,
"prompt_tokens_details": { "cached_tokens": 0 },
"cost": { "input": 0, "output": 0, "cacheRead": 0, "cacheWrite": 0, "total": 0.0000045 }
}
}When tools are provided, a tool-requesting turn comes back with
finish_reason: "tool_calls", message.content: null, and a
message.tool_calls array ({id, type:"function", function:{name, arguments}});
streamed as delta.tool_calls fragments ({index, id, type, function:{name}}
then {index, function:{arguments}}). Send each result back as a tool role
message to continue.
When stream: true, the client must include an MCP progressToken. Each CEP-41
chunk frame carries one JSON-serialized OpenAI chat.completion.chunk object
(delta.content for text, delta.reasoning_content for thinking, then a final
chunk with finish_reason + usage). The stream is the response — the
closing JSON-RPC tool result carries only metadata (id, model,
finish_reason, usage), never the assembled text. cost rides inside
usage (OpenRouter's usage.cost convention) so it also appears in the final
stream chunk. A bridge to OpenAI SSE wraps each chunk as
data: <chunk>\n\n plus a trailing data: [DONE].
Requires Bun ≥ 1.2.
bun install
cp .env.example .env # then edit: set provider keys and (optionally) a stable server keybun startOn startup the server prints its Nostr public key and relays. Clients target
that pubkey over the configured relays. Set MUXLL_ANNOUNCED=1 to publish
public discovery announcements.
@muxll/cli is the reference command-line client over @muxll/client, built on
@earendil-works/pi-tui.
Point it at a running server's pubkey and relays (env or flags):
export MUXLL_SERVER_PUBKEY=<pubkey printed by `bun start`>
bun run cli models
bun run cli chat --model anthropic/claude-sonnet-4-5 # full-screen TUI (Ctrl+C to quit)
bun run cli chat --model anthropic/claude-sonnet-4-5 "hi" # one streamed turn to stdoutchat with no prompt opens an interactive full-screen chat — streamed markdown
replies and a status bar with live usage (prompt/completion tokens, cost,
turns). chat with a prompt prints one streamed turn to stdout, for scripting.
MUXLL_RELAYS (default wss://relay.contextvm.org) and MUXLL_MODEL are
optional; --pubkey/--relays/--model/-m override per invocation.
Integration tests run the full server + client over an in-process mock Nostr
relay using pi-ai's fauxProvider (no real API keys or network needed):
bun run test # runs `bun test packages/*/tests`
bunx tsc --noEmit # typecheck the whole workspaceBare bun test (or bun test packages/) is too broad: it also scans the
vendored docs/ tree and follows the workspace node_modules symlinks into
Bun's package store, running dependency test suites. bun run test scopes to
packages/*/tests, which holds every project test.
A Bun workspace monorepo (see tsconfig.json paths — packages import each
other by @muxll/* name and resolve to source, so dev needs no build step):
packages/
core/ @muxll/core OpenAI wire shapes (zod schemas + types), shared by server + clients
server/ @muxll/server the CVM/Nostr adapter operators run
src/main.ts env-driven startup: relays, signer, provider registry
src/server.ts McpServer, tool registration, CEP-41 transport wiring
src/wire.ts OpenAI ↔ pi-ai translation (transport-agnostic)
tests/ end-to-end: models.list, chat.complete (both modes), error path
client/ @muxll/client reference CVM client (transport + tool callers + CEP-41 stream parser)
tests/ drives the server over a mock relay via MuxllClient
cli/ @muxll/cli command-line client over @muxll/client (models list + streaming chat REPL)
tests/ parseArgs + streamTurn driven over a mock relay
docs/ vendored references (cordn, routstr-core, yalr, cvm + pi docs)
— read-only, excluded from build and tests
The shared wire shapes (@muxll/core) are the seam: the server validates
inputs and builds outputs against them, clients build inputs and parse outputs
against them. @muxll/client is the programmatic surface the CLI (@muxll/cli),
HTTP proxy, and web app build on, and the one the test harness drives.
Proof of concept. Deliberately deferred until the core is validated:
- Provider config file — currently env-var-only built-ins (
builtinModels()readsANTHROPIC_API_KEY,OPENAI_API_KEY, …). Add a models.json loader for custom base URLs / providers. - Stream cancellation — a client CEP-41
abortdoes not yet cancel the upstream provider stream. - Structured stream chunks —
text_delta,thinking_delta, and tool-call deltas are all streamed aschat.completion.chunkdeltas (delta.content,delta.reasoning_content,delta.tool_calls). - OpenAI feature gaps — structured outputs (
response_format) and sampling params (top_p,stop,seed, penalties) are supported per-API by pi-ai but not yet surfaced (need theonPayloadseam). Function/tool calling is fully wired (tools,tool_choicenormalized per target API,toolrole); multimodal image input is wired (usercontent acceptstext+image_urlparts, data-URL or http(s)). - Pricing & quotas — out of scope per the design brief.
See AGENTS.md for contributor/agent instructions, and
docs/idea.md for the original design note.