Local API (OpenAI-compatible)

Serve your local coding agents (Claude Code, and Codex/OpenCode where installed) over one OpenAI-compatible HTTP API - route per request by model, with your existing auth, skills and sessions, and tools and MCP on request.

enigma api exposes your local coding agents over an HTTP API that any OpenAI client library can call. Point your existing OpenAI SDK at http://127.0.0.1:8000/v1 and every request runs through a real agent on your machine - with its skills, session continuity, and the login you already use. Tools and your MCP servers are opt-in per request ("enable_tools": true), so plain chat stays fast.

It is not a network proxy to Anthropic. Each request spawns the local agent CLI in headless mode and translates its output into OpenAI (/v1/chat/completions) and Anthropic (/v1/messages) responses. Nothing new to authenticate: it reuses your active account.

One server, many agents

A single server can back several agents at once - the request’s model field picks which one:

  • claude-sonnet-5, opus, haiku, any claude-* id -> Claude Code
  • codex (or codex/<model>) -> Codex
  • opencode (or opencode/<provider/model>) -> OpenCode
  • kimi (or kimi/<model alias>) -> Kimi Code

--tool sets the default backend for requests whose model does not name one. /v1/models lists only the agents actually installed on your machine. Claude Code is verified end-to-end; the Codex, OpenCode and Kimi Code adapters are wired from their headless modes (codex exec --json, opencode run and kimi -p --output-format stream-json) rather than run end-to-end.

Native and loopback-bound

The server is built into enigma (node:http, dependency-free) and binds to 127.0.0.1 only, so it is reachable from your machine and nowhere else. It runs on demand while enigma api is open, and stops on Ctrl+C.

Run it

enigma api

Start the API server. --port <n> overrides the port (else the apiPort config, default 8000); --api-key <k> (or the ENIGMA_API_KEY env var) requires a bearer token on every /v1 route; --tool <t> selects which agent backs it (Claude Code by default).

$ enigma api
$ enigma api --port 8080
$ enigma api --api-key mysecret
$ enigma api --tool claude

Use it from an OpenAI client

from openai import OpenAI

client = OpenAI(base_url="http://127.0.0.1:8000/v1", api_key="not-needed")
resp = client.chat.completions.create(
    model="claude-opus-5",
    messages=[{"role": "user", "content": "Explain this repo's architecture."}],
)
print(resp.choices[0].message.content)

The default model is claude-opus-5 when a request names none (or a non-Claude id). Streaming (stream=true), the Anthropic /v1/messages shape, and standard OpenAI parameters work as well.

Images

Send images the same way you would with OpenAI or Anthropic - the API forwards them to Claude Code’s vision:

resp = client.chat.completions.create(
    model="claude-opus-5",
    messages=[{"role": "user", "content": [
        {"type": "text", "text": "What is in this image?"},
        {"type": "image_url", "image_url": {"url": "data:image/png;base64,iVBORw0KGgo..."}},
    ]}],
)

Both a data: URL (base64) and an https:// URL work. The Anthropic shape’s native image content blocks (/v1/messages) are supported too. The dashboard Playground also has an image picker. Images are applied by Claude Code; agents without local vision support get the text only.

Endpoints

  • POST /v1/chat/completions - OpenAI chat completions (streaming and non-streaming).
  • POST /v1/messages - Anthropic Messages API shape (streaming and non-streaming).
  • GET /v1/models - the model catalog (any claude-* id or alias is also accepted).
  • GET /v1/sessions, DELETE /v1/sessions/{id} - list and drop tracked sessions.
  • GET /health, GET / - service status and endpoint index.

Accounts, profiles and packs

Every request can run under a specific account, profile, or pack - the same contexts you use elsewhere in enigma. Set a default for the whole server with a flag, or override it per request in the body:

enigma api --account work

--account <name> runs under that account’s login, --profile <name> uses the profile’s account mapping for the backend, and --pack <id> runs inside a pack’s isolated context (e.g. Helio - its skills, commands, sub-agents and MCP, seeded with the pack’s account).

$ enigma api --account work
$ enigma api --profile team
$ enigma api --pack helio

Per request, add "account", "profile", or "pack" to the JSON body (they are ignored by standard OpenAI clients, so they are safe extensions):

{
  "model": "claude-sonnet-5",
  "pack": "helio",
  "messages": [{ "role": "user", "content": "Scope this target." }]
}

A request’s context wins over the server default, so one server can serve several accounts or packs at once. GET /health reports the server’s default context.

You can also set a persisted default that enigma api uses when no flag is given - from the dashboard Local API page (Server defaults) or the CLI:

enigma config api-account work

Persist the default account / profile / pack for the API server (none clears it). A per-run flag or a per-request field still overrides it.

$ enigma config api-account work
$ enigma config api-pack helio
$ enigma config api-account none

Account rotation

Instead of one fixed account, enigma can spread requests across several. Rotation applies to requests that name no account, profile or pack:

Strategy Picks
off (default) the default account / profile / pack above, else the active account
round-robin the next account in turn
least-used the account with the least token use in its current 5-hour window (from your transcripts when real usage stats are on, plus what the server sent since)
fill-first the first account until it hits a usage limit, then the next
random any account at random

With any strategy on, a request that hits a usage limit or a failed login is retried on the next account, and the limited one sits out until the reset time Claude reported (otherwise 2 minutes, doubling up to 30). Once every account is cooling down the server answers 429. A streamed reply fails over too: the opening text is held back until it is clear it is a real answer, not a limit notice. A warm session stays on the account it started on.

enigma config api-rotation round-robin

api-rotation sets the strategy, api-account-pool lists the accounts it may use in order (empty = every account of the backend), and --rotation overrides the strategy for one run. A context flag (--account, --profile, --pack) turns rotation off for that run.

$ enigma config api-rotation round-robin
$ enigma config api-account-pool add work
$ enigma api --rotation fill-first

A rotated response carries an x-enigma-account header naming the account that answered, and GET /health (with the API key, when one is set) shows the strategy, the pool, which accounts are cooling down and until when.

Who chooses. By default a caller may still pick its own account per request. To let enigma decide alone, turn that off - a request that names an account, profile or pack is then refused with 403:

enigma config api-client-context off

Whether requests may choose their own account, profile or pack.

$ enigma config api-client-context off
$ enigma config api-client-context on

Tools

Tools are off by default so plain chat requests stay fast and OpenAI-compatible. Pass "enable_tools": true in the request body to let Claude Code use its tools (Read, Write, Bash, …); when enabled, it honors your global permission-bypass setting so headless runs never stall on a prompt.

With tools off, Claude Code also skips your MCP servers: nothing connects to them and their tool schemas stay out of the prompt, which is most of the per-request latency and token cost on a machine with many servers configured. A request that needs an MCP-backed tool must set "enable_tools": true - that loads your MCP servers as usual. This applies to Claude Code only; the other backends load their MCP servers either way.

Sessions

There are two ways to hold a conversation, and both are always available.

Stateless (default). Send the full message history on every request, with no session_id. Each request runs in isolation from a cold start - the plain OpenAI-compatible behavior.

Stateful warm sessions. Set a session_id (a UUID) and send only the newest user message each turn. enigma keeps one warm Claude Code process per session that holds the thread, so you never resend the history and only the first turn pays the startup cost. This mirrors OpenAI’s Responses API (previous_response_id) and AWS Bedrock AgentCore’s per-session compute. Claude Code only; another backend given a session_id falls back to a stateless resume that carries the whole transcript.

import uuid

sid = str(uuid.uuid4())

# Turn 1 - introduce a fact
client.chat.completions.create(
    model="claude-opus-5",
    messages=[{"role": "user", "content": "My name is Ada."}],
    extra_body={"session_id": sid},
)
# Turn 2 - send ONLY the new message; the warm process still remembers turn 1
client.chat.completions.create(
    model="claude-opus-5",
    messages=[{"role": "user", "content": "What is my name?"}],
    extra_body={"session_id": sid},
)

The contract:

  • session_id must be a UUID. Anything else is rejected with 400 before any stream is opened.
  • Turns of one session are serialized: a second concurrent turn on the same id gets 409. Different sessions run in parallel.
  • A turn that produces no result within the turn timeout (10 minutes) gets 504; the session is reset and the next turn resumes it.
  • The response echoes session_id - reuse it. Streaming (stream=true) works in both modes.
  • A session_id is bound to the context (account / profile / pack) that first ran it; reusing it under a different context is refused, so start a fresh UUID per context.
  • Sessions idle past a TTL, or beyond a live-process cap, are evicted to free memory. The transcript survives on disk, so a later turn on that id transparently resumes it.

Live sessions are listed under GET /v1/sessions and can be dropped with DELETE /v1/sessions/{id}.

Configure the default port

enigma config api-port 8080

Persist the port enigma api binds by default. A --port flag still overrides it per run.

$ enigma config api-port 8080
$ enigma api

Dashboard playground

The local dashboard (enigma dashboard) has a Playground tab to test the API without leaving the browser: pick the format (OpenAI or Anthropic), a backend/model, optionally an auth key, type a message, and send. Run it in-process (no enigma api process needed) or over HTTP against a running server with its auth key. It shows the response, token usage, and the equivalent curl command. HTTP mode only calls loopback servers.

Security

The server is loopback-only. Add an API key with --api-key or ENIGMA_API_KEY if other local processes should not reach it; /v1/models, /health and / stay open, every other route requires the bearer token. Requests never leave your machine except as the backing agent’s own normal traffic.