Self-hosted AI inference gateway

Every AI provider, behind one endpoint

Tera Router is a unified gateway for your custom AI providers. Speak OpenAI, Anthropic, or Responses dialects — route across accounts with fallback, keys, guardrails, and usage analytics built in.

terminal
$ curl http://localhost:8080/v1/chat/completions \    -H "Authorization: Bearer tr_live_…" \    -d '{        "model": "glm-4.7",        "messages": [{ "role": "user", "content": "hi" }]      }'← 200 OK · routed via glm (account #2) · 1,204 tok
providers in the catalog
20+
providers in the catalog
API dialects, one gateway
3
API dialects, one gateway
accounts per provider
N×
accounts per provider
streaming with usage metering
SSE
streaming with usage metering

One gateway · three dialects

Your tools keep their language

Point any tool at Tera Router and keep its native protocol. The transform engine translates between dialects, so a request made in OpenAI Chat can land on an Anthropic-dialect provider.

OpenAI Chat Completions

POST /v1/chat/completions

The de-facto standard. Any OpenAI SDK, LangChain pipeline, Cursor config, or curl script works as-is — just swap the base URL.

Anthropic Messages

POST /v1/messages

Native Messages API including /v1/messages/count_tokens, so Claude Code and Anthropic SDKs connect without a shim.

OpenAI Responses

POST /v1/responses

The newer Responses API surface for agents and stateful workflows, routed to Responses-native providers like Codex.

How it works

Connect. Route. Observe.

Connect accounts once, then let the planner handle the rest: priority order, cooldowns, fallback across providers, and full metering of what each key consumes.

  1. 01

    Connect providers

    Pick from the built-in catalog — API keys, OAuth sessions, or key-less local endpoints. Multiple accounts per provider, each with its own base URL and priority.

    openai · anthropic · glm · ollama …

  2. 02

    Route with fallback

    Every request enters one endpoint. The planner resolves accounts in priority order, cools down failures, and falls back across accounts and providers automatically.

    glm-4.7 → kimi-k2 → claude-sonnet (fallback)

  3. 03

    Observe everything

    Token usage is metered per request — including SSE streams — and lands in the admin dashboard: per key, per account, per model.

    GET /admin/usage → per-key analytics

Provider catalog

Bring every account you already have

A built-in catalog of presets — base URL, dialect, and auth mode included — plus custom endpoints for anything OpenAI- or Anthropic-compatible.

API key

16

Bring your own key — catalog presets ship the base URL and dialect.

  • OpenAI
  • Anthropic
  • Gemini
  • DeepSeek
  • xAI (Grok)
  • GLM
  • Kimi
  • MiniMax
  • Mistral
  • Groq
  • Cohere
  • Perplexity
  • OpenRouter
  • NVIDIA NIM
  • Azure OpenAI
  • Ollama Cloud

Subscription / OAuth

4

Connect existing coding-plan sessions and route them like any other account.

  • Claude Code
  • OpenAI Codex
  • GitHub Copilot
  • Kilo Code

Self-hosted & key-less

2

Authenticates by network position — no credentials, no account rows.

  • Ollama Local
  • vLLM

Custom

2

Any OpenAI- or Anthropic-compatible endpoint, with per-account base URLs.

  • OpenAI-compatible
  • Anthropic-compatible

Under the hood

Built for running, not just proxying

Routing is the headline — the gateway also carries the operational load: access control, policy, and telemetry, all in one deploy.

Multi-account routing

Several accounts per provider with priority ordering, cooldowns after failures, and automatic fallback attempts.

Dialect translation

The transform engine rewrites between OpenAI Chat, Anthropic Messages, and Responses — request in one dialect, land on another.

Per-key allowlists

Every gateway key can be restricted to specific models, so a teammate or agent only reaches what it should.

Guardrails

A policy engine screens requests before traffic leaves your box — your rules, enforced centrally.

Usage metering

Per-request token accounting (streaming included) feeding per-key, per-account, per-model analytics.

Streaming first

SSE pass-through for chat and messages with usage extracted from the stream — no buffering, no lost metrics.

Self-hosted & light

A single Go service you own: one container next to your stack. No SaaS middleman, no per-token markup.

Any SDK works

OpenAI and Anthropic client libraries, IDEs, and agent frameworks all connect by pointing at one base URL.

Get started

Two commands to a unified endpoint

One container next to your stack, one base URL in every tool. From clone to first routed request in minutes.

1Run the gateway

bash
$ git clone montara-project/tera-router$ cd deploy && docker compose up -d

2Point your SDK

python
client = OpenAI(  base_url="http://localhost:8080/v1",  api_key="tr_live_…",)

Create a gateway key in the admin dashboard, attach a model allowlist if you want to scope it, and ship your first request.

Stop juggling SDKs, keys, and quotas

Give every tool the same base URL and one key. Tera Router picks the account, handles the dialect, and shows you exactly what was spent.