OpenAI Chat Completions
POST /v1/chat/completions
The de-facto standard. Any OpenAI SDK, LangChain pipeline, Cursor config, or curl script works as-is — just swap the base URL.
Self-hosted AI inference gateway
Tera Router is a unified gateway for your custom AI providers. Speak OpenAI, Anthropic, or Responses dialects — route across accounts with fallback, keys, guardrails, and usage analytics built in.
$ curl http://localhost:8080/v1/chat/completions \ -H "Authorization: Bearer tr_live_…" \ -d '{ "model": "glm-4.7", "messages": [{ "role": "user", "content": "hi" }] }'← 200 OK · routed via glm (account #2) · 1,204 tokOne gateway · three dialects
Point any tool at Tera Router and keep its native protocol. The transform engine translates between dialects, so a request made in OpenAI Chat can land on an Anthropic-dialect provider.
POST /v1/chat/completions
The de-facto standard. Any OpenAI SDK, LangChain pipeline, Cursor config, or curl script works as-is — just swap the base URL.
POST /v1/messages
Native Messages API including /v1/messages/count_tokens, so Claude Code and Anthropic SDKs connect without a shim.
POST /v1/responses
The newer Responses API surface for agents and stateful workflows, routed to Responses-native providers like Codex.
How it works
Connect accounts once, then let the planner handle the rest: priority order, cooldowns, fallback across providers, and full metering of what each key consumes.
Pick from the built-in catalog — API keys, OAuth sessions, or key-less local endpoints. Multiple accounts per provider, each with its own base URL and priority.
openai · anthropic · glm · ollama …
Every request enters one endpoint. The planner resolves accounts in priority order, cools down failures, and falls back across accounts and providers automatically.
glm-4.7 → kimi-k2 → claude-sonnet (fallback)
Token usage is metered per request — including SSE streams — and lands in the admin dashboard: per key, per account, per model.
GET /admin/usage → per-key analytics
Provider catalog
A built-in catalog of presets — base URL, dialect, and auth mode included — plus custom endpoints for anything OpenAI- or Anthropic-compatible.
Bring your own key — catalog presets ship the base URL and dialect.
Connect existing coding-plan sessions and route them like any other account.
Authenticates by network position — no credentials, no account rows.
Any OpenAI- or Anthropic-compatible endpoint, with per-account base URLs.
Under the hood
Routing is the headline — the gateway also carries the operational load: access control, policy, and telemetry, all in one deploy.
Several accounts per provider with priority ordering, cooldowns after failures, and automatic fallback attempts.
The transform engine rewrites between OpenAI Chat, Anthropic Messages, and Responses — request in one dialect, land on another.
Every gateway key can be restricted to specific models, so a teammate or agent only reaches what it should.
A policy engine screens requests before traffic leaves your box — your rules, enforced centrally.
Per-request token accounting (streaming included) feeding per-key, per-account, per-model analytics.
SSE pass-through for chat and messages with usage extracted from the stream — no buffering, no lost metrics.
A single Go service you own: one container next to your stack. No SaaS middleman, no per-token markup.
OpenAI and Anthropic client libraries, IDEs, and agent frameworks all connect by pointing at one base URL.
Get started
One container next to your stack, one base URL in every tool. From clone to first routed request in minutes.
1Run the gateway
$ git clone montara-project/tera-router$ cd deploy && docker compose up -d2Point your SDK
client = OpenAI( base_url="http://localhost:8080/v1", api_key="tr_live_…",)Self-host it
deploy/docker-compose.yaml runs the gateway and dashboard as one image on port 8080.
MIT licensed
The whole monorepo is MIT — read it, change it, run it in production.
Bring your own stack
No account, no seat licence, no per-token markup — the gateway just routes.
Create a gateway key in the admin dashboard, attach a model allowlist if you want to scope it, and ship your first request.
Give every tool the same base URL and one key. Tera Router picks the account, handles the dialect, and shows you exactly what was spent.
MIT licensed · self-hosted · docs.terarouter.xyz