AI APIs
Test your agent against a fake OpenAI, Anthropic or Gemini.
The same host-swap idea as the MCP twins, for the model APIs: point your agent CLI or SDK at a protocol-exact stand-in and it just works — no tokens, no network, and the same answer every run. The value isn't the model's intelligence; it's exercising the SDK and API layer — retries, streaming, tool-calls, usage accounting, error handling — deterministically.
Your app's model name and headers stay exactly as they are. You change one thing: the base URL.
Fastest: point a CLI at it
One env var in front of a CLI you already run. Verified end-to-end — the command below really does drive Claude Code against the twin.
ANTHROPIC_BASE_URL=https://acme-api-anthropic.sandbox.restaged.dev \
ANTHROPIC_API_KEY=sk-restaged \
claude -p "hello"
OPENAI_BASE_URL=https://acme-api-openai.sandbox.restaged.dev/v1 \
OPENAI_API_KEY=sk-restaged \
codex "hello"
The reply is a deterministic simulation from Restaged, so the transcript is the same every time — what you're testing is that your tool speaks the API correctly, offline.
Ship it: one env var in your app
Every official SDK reads a base-URL env var, so pointing your app at the twin in CI or a test suite is a config change, not a code change.
# OpenAI SDK (Python / JS) — no code change, only the env
export OPENAI_BASE_URL=https://acme-api-openai.sandbox.restaged.dev/v1
export OPENAI_API_KEY=sk-restaged
# Anthropic SDK — no code change
export ANTHROPIC_BASE_URL=https://acme-api-anthropic.sandbox.restaged.dev
export ANTHROPIC_API_KEY=sk-restaged
# Google GenAI SDK — no code change
export GOOGLE_GEMINI_BASE_URL=https://acme-api-gemini.sandbox.restaged.dev
export GEMINI_API_KEY=restaged
The base URLs
One host per vendor, under the twin zone:
OpenAI https://acme-api-openai.sandbox.restaged.dev/v1
Anthropic https://acme-api-anthropic.sandbox.restaged.dev
Gemini https://acme-api-gemini.sandbox.restaged.dev
Swap acme for any world:
acme, cascade-works, chinook, helios-analytics, northwind. The world label is
cosmetic here — the AI twins synthesise responses rather than read the world's data — so any of
them serves all three vendors identically.
What you can rely on
- Deterministic & offline. A given request always returns the same id, content and token usage — no network, no cost, no rate limits.
- Any model, any key. The model you send is accepted and echoed back; the API key is accepted and ignored. Nothing to configure, nothing to leak.
- The real surfaces. OpenAI is full — chat, the Responses API, embeddings, images, audio, files, batches, Assistants, and Realtime over a WebSocket. Anthropic and Gemini cover messages / generateContent with streaming, tools and structured output.
- Synthetic content, on purpose. Replies are simulated, not intelligent — this tests the plumbing, not the answers. Need a specific reply, tool-call or error for a test? A world can script the twins server-side, so your app's request never changes.
Unimplemented corners return the vendor's own error shape rather than pretending — the same honesty as the rest of Restaged.