enigma api exposes your local coding agents over an HTTP API that any OpenAI client library can call. Point your existing OpenAI SDK at http://127.0.0.1:8000/v1 and every request runs through a real agent on your machine - with its skills, session continuity, and the login you already use. Tools and your MCP servers are opt-in per request ("enable_tools": true), so plain chat stays fast.
It is not a network proxy to Anthropic. Each request spawns the local agent CLI in headless mode and translates its output into OpenAI (/v1/chat/completions) and Anthropic (/v1/messages) responses. Nothing new to authenticate: it reuses your active account.
One server, many agents
A single server can back several agents at once - the request’s model field picks which one:
claude-sonnet-5,opus,haiku, anyclaude-*id -> Claude Codecodex(orcodex/<model>) -> Codexopencode(oropencode/<provider/model>) -> OpenCodekimi(orkimi/<model alias>) -> Kimi Code
--tool sets the default backend for requests whose model does not name one. /v1/models lists only the agents actually installed on your machine. Claude Code is verified end-to-end; the Codex, OpenCode and Kimi Code adapters are wired from their headless modes (codex exec --json, opencode run and kimi -p --output-format stream-json) rather than run end-to-end.
Native and loopback-bound
The server is built into enigma (node:http, dependency-free) and binds to 127.0.0.1 only, so it is reachable from your machine and nowhere else. It runs on demand while enigma api is open, and stops on Ctrl+C.
Run it
enigma api
Start the API server. --port <n> overrides the port (else the apiPort config, default 8000); --api-key <k> (or the ENIGMA_API_KEY env var) requires a bearer token on every /v1 route; --tool <t> selects which agent backs it (Claude Code by default).
Use it from an OpenAI client
from openai import OpenAI
client = OpenAI(base_url="http://127.0.0.1:8000/v1", api_key="not-needed")
resp = client.chat.completions.create(
model="claude-opus-5",
messages=[{"role": "user", "content": "Explain this repo's architecture."}],
)
print(resp.choices[0].message.content)
The default model is claude-opus-5 when a request names none (or a non-Claude id). Streaming (stream=true), the Anthropic /v1/messages shape, and standard OpenAI parameters work as well.
Images
Send images the same way you would with OpenAI or Anthropic - the API forwards them to Claude Code’s vision:
resp = client.chat.completions.create(
model="claude-opus-5",
messages=[{"role": "user", "content": [
{"type": "text", "text": "What is in this image?"},
{"type": "image_url", "image_url": {"url": "data:image/png;base64,iVBORw0KGgo..."}},
]}],
)
Both a data: URL (base64) and an https:// URL work. The Anthropic shape’s native image content blocks (/v1/messages) are supported too. The dashboard Playground also has an image picker. Images are applied by Claude Code; agents without local vision support get the text only.
Endpoints
POST /v1/chat/completions- OpenAI chat completions (streaming and non-streaming).POST /v1/messages- Anthropic Messages API shape (streaming and non-streaming).GET /v1/models- the model catalog (anyclaude-*id or alias is also accepted).GET /v1/sessions,DELETE /v1/sessions/{id}- list and drop tracked sessions.GET /health,GET /- service status and endpoint index.
Accounts, profiles and packs
Every request can run under a specific account, profile, or pack - the same contexts you use elsewhere in enigma. Set a default for the whole server with a flag, or override it per request in the body:
enigma api --account work
--account <name> runs under that account’s login, --profile <name> uses the profile’s account mapping for the backend, and --pack <id> runs inside a pack’s isolated context (e.g. Helio - its skills, commands, sub-agents and MCP, seeded with the pack’s account).
Per request, add "account", "profile", or "pack" to the JSON body (they are ignored by standard OpenAI clients, so they are safe extensions):
{
"model": "claude-sonnet-5",
"pack": "helio",
"messages": [{ "role": "user", "content": "Scope this target." }]
}
A request’s context wins over the server default, so one server can serve several accounts or packs at once. GET /health reports the server’s default context.
You can also set a persisted default that enigma api uses when no flag is given - from the dashboard Local API page (Server defaults) or the CLI:
enigma config api-account work
Persist the default account / profile / pack for the API server (none clears it). A per-run flag or a per-request field still overrides it.
Account rotation
Instead of one fixed account, enigma can spread requests across several. Rotation applies to requests that name no account, profile or pack:
| Strategy | Picks |
|---|---|
off (default) |
the default account / profile / pack above, else the active account |
round-robin |
the next account in turn |
least-used |
the account with the least token use in its current 5-hour window (from your transcripts when real usage stats are on, plus what the server sent since) |
fill-first |
the first account until it hits a usage limit, then the next |
random |
any account at random |
With any strategy on, a request that hits a usage limit or a failed login is retried on the next account, and the limited one sits out until the reset time Claude reported (otherwise 2 minutes, doubling up to 30). Once every account is cooling down the server answers 429. A streamed reply fails over too: the opening text is held back until it is clear it is a real answer, not a limit notice. A warm session stays on the account it started on.
enigma config api-rotation round-robin
api-rotation sets the strategy, api-account-pool lists the accounts it may use in order (empty = every account of the backend), and --rotation overrides the strategy for one run. A context flag (--account, --profile, --pack) turns rotation off for that run.
A rotated response carries an x-enigma-account header naming the account that answered, and GET /health (with the API key, when one is set) shows the strategy, the pool, which accounts are cooling down and until when.
Who chooses. By default a caller may still pick its own account per request. To let enigma decide alone, turn that off - a request that names an account, profile or pack is then refused with 403:
enigma config api-client-context off
Whether requests may choose their own account, profile or pack.
Tools
Tools are off by default so plain chat requests stay fast and OpenAI-compatible. Pass "enable_tools": true in the request body to let Claude Code use its tools (Read, Write, Bash, …); when enabled, it honors your global permission-bypass setting so headless runs never stall on a prompt.
With tools off, Claude Code also skips your MCP servers: nothing connects to them and their tool schemas stay out of the prompt, which is most of the per-request latency and token cost on a machine with many servers configured. A request that needs an MCP-backed tool must set "enable_tools": true - that loads your MCP servers as usual. This applies to Claude Code only; the other backends load their MCP servers either way.
Sessions
There are two ways to hold a conversation, and both are always available.
Stateless (default). Send the full message history on every request, with no session_id. Each request runs in isolation from a cold start - the plain OpenAI-compatible behavior.
Stateful warm sessions. Set a session_id (a UUID) and send only the newest user message each turn. enigma keeps one warm Claude Code process per session that holds the thread, so you never resend the history and only the first turn pays the startup cost. This mirrors OpenAI’s Responses API (previous_response_id) and AWS Bedrock AgentCore’s per-session compute. Claude Code only; another backend given a session_id falls back to a stateless resume that carries the whole transcript.
import uuid
sid = str(uuid.uuid4())
# Turn 1 - introduce a fact
client.chat.completions.create(
model="claude-opus-5",
messages=[{"role": "user", "content": "My name is Ada."}],
extra_body={"session_id": sid},
)
# Turn 2 - send ONLY the new message; the warm process still remembers turn 1
client.chat.completions.create(
model="claude-opus-5",
messages=[{"role": "user", "content": "What is my name?"}],
extra_body={"session_id": sid},
)
The contract:
session_idmust be a UUID. Anything else is rejected with400before any stream is opened.- Turns of one session are serialized: a second concurrent turn on the same id gets
409. Different sessions run in parallel. - A turn that produces no result within the turn timeout (10 minutes) gets
504; the session is reset and the next turn resumes it. - The response echoes
session_id- reuse it. Streaming (stream=true) works in both modes. - A
session_idis bound to the context (account / profile / pack) that first ran it; reusing it under a different context is refused, so start a fresh UUID per context. - Sessions idle past a TTL, or beyond a live-process cap, are evicted to free memory. The transcript survives on disk, so a later turn on that id transparently resumes it.
Live sessions are listed under GET /v1/sessions and can be dropped with DELETE /v1/sessions/{id}.
Configure the default port
enigma config api-port 8080
Persist the port enigma api binds by default. A --port flag still overrides it per run.
Dashboard playground
The local dashboard (enigma dashboard) has a Playground tab to test the API without leaving the browser: pick the format (OpenAI or Anthropic), a backend/model, optionally an auth key, type a message, and send. Run it in-process (no enigma api process needed) or over HTTP against a running server with its auth key. It shows the response, token usage, and the equivalent curl command. HTTP mode only calls loopback servers.
Security
The server is loopback-only. Add an API key with --api-key or ENIGMA_API_KEY if other local processes should not reach it; /v1/models, /health and / stay open, every other route requires the bearer token. Requests never leave your machine except as the backing agent’s own normal traffic.