Self-hosted AI inference gateway
Every AI provider, behind one endpoint
Tera Router is a unified gateway for your custom AI providers. Speak OpenAI, Anthropic, or Responses dialects — route across accounts with fallback, keys, guardrails, and usage analytics built in.
$ curl http://localhost:8080/v1/chat/completions \ -H "Authorization: Bearer tr_live_…" \ -d '{ "model": "glm-4.7", "messages": [{ "role": "user", "content": "hi" }] }'← 200 OK · routed via glm (account #2) · 1,204 tok- providers in the catalog
- 20+
- providers in the catalog
- API dialects, one gateway
- 3
- API dialects, one gateway
- accounts per provider
- N×
- accounts per provider
- streaming with usage metering
- SSE
- streaming with usage metering
One gateway · three dialects
Your tools keep their language
Point any tool at Tera Router and keep its native protocol. The transform engine translates between dialects, so a request made in OpenAI Chat can land on an Anthropic-dialect provider.
OpenAI Chat Completions
POST /v1/chat/completions
The de-facto standard. Any OpenAI SDK, LangChain pipeline, Cursor config, or curl script works as-is — just swap the base URL.
Anthropic Messages
POST /v1/messages
Native Messages API including /v1/messages/count_tokens, so Claude Code and Anthropic SDKs connect without a shim.
OpenAI Responses
POST /v1/responses
The newer Responses API surface for agents and stateful workflows, routed to Responses-native providers like Codex.
How it works
Connect. Route. Observe.
Connect accounts once, then let the planner handle the rest: priority order, cooldowns, fallback across providers, and full metering of what each key consumes.
- 01
Connect providers
Pick from the built-in catalog — API keys, OAuth sessions, or key-less local endpoints. Multiple accounts per provider, each with its own base URL and priority.
openai · anthropic · glm · ollama …
- 02
Route with fallback
Every request enters one endpoint. The planner resolves accounts in priority order, cools down failures, and falls back across accounts and providers automatically.
glm-4.7 → kimi-k2 → claude-sonnet (fallback)
- 03
Observe everything
Token usage is metered per request — including SSE streams — and lands in the admin dashboard: per key, per account, per model.
GET /admin/usage → per-key analytics
Provider catalog
Bring every account you already have
A built-in catalog of presets — base URL, dialect, and auth mode included — plus custom endpoints for anything OpenAI- or Anthropic-compatible.
API key
16Bring your own key — catalog presets ship the base URL and dialect.
- OpenAI
- Anthropic
- Gemini
- DeepSeek
- xAI (Grok)
- GLM
- Kimi
- MiniMax
- Mistral
- Groq
- Cohere
- Perplexity
- OpenRouter
- NVIDIA NIM
- Azure OpenAI
- Ollama Cloud
Subscription / OAuth
4Connect existing coding-plan sessions and route them like any other account.
- Claude Code
- OpenAI Codex
- GitHub Copilot
- Kilo Code
Self-hosted & key-less
2Authenticates by network position — no credentials, no account rows.
- Ollama Local
- vLLM
Custom
2Any OpenAI- or Anthropic-compatible endpoint, with per-account base URLs.
- OpenAI-compatible
- Anthropic-compatible
Under the hood
Built for running, not just proxying
Routing is the headline — the gateway also carries the operational load: access control, policy, and telemetry, all in one deploy.
Multi-account routing
Several accounts per provider with priority ordering, cooldowns after failures, and automatic fallback attempts.
Dialect translation
The transform engine rewrites between OpenAI Chat, Anthropic Messages, and Responses — request in one dialect, land on another.
Per-key allowlists
Every gateway key can be restricted to specific models, so a teammate or agent only reaches what it should.
Guardrails
A policy engine screens requests before traffic leaves your box — your rules, enforced centrally.
Usage metering
Per-request token accounting (streaming included) feeding per-key, per-account, per-model analytics.
Streaming first
SSE pass-through for chat and messages with usage extracted from the stream — no buffering, no lost metrics.
Self-hosted & light
A single Go service you own: one container next to your stack. No SaaS middleman, no per-token markup.
Any SDK works
OpenAI and Anthropic client libraries, IDEs, and agent frameworks all connect by pointing at one base URL.
Get started
Two commands to a unified endpoint
One container next to your stack, one base URL in every tool. From clone to first routed request in minutes.
1Run the gateway
$ git clone montara-project/tera-router$ cd deploy && docker compose up -d2Point your SDK
client = OpenAI( base_url="http://localhost:8080/v1", api_key="tr_live_…",)Create a gateway key in the admin dashboard, attach a model allowlist if you want to scope it, and ship your first request.
Stop juggling SDKs, keys, and quotas
Give every tool the same base URL and one key. Tera Router picks the account, handles the dialect, and shows you exactly what was spent.