Skip to content
Automatic routing, RPM-based

One endpoint. Automatic model routing.

Send an OpenAI-compatible request with no model field. Vechgate picks an eligible model for the task, falls back when a provider fails, and records which model actually ran.

99.9%
availability target
Monthly gateway availability, subject to load testing.
<100 ms
gateway overhead
P95, excluding provider latency.
No tokens
billing unit
Subscriptions cap requests per minute, not tokens.
request.sh
curl https://api.vechgate.dev/v1/chat/completions \
  -H "Authorization: Bearer $VECHGATE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"messages": [{"role": "user", "content": "Summarize this release."}]}'
  • Groq
  • OpenAI
  • Anthropic
  • Google Gemini

Why Vechgate

One gateway between your app and every model

Automatic routing, a single compatible endpoint, and billing you can audit. These are the parts the product spec commits to.

01 / routing

The gateway chooses the model

Complexity, required capabilities, latency history, provider health, context size, and your plan all feed one decision. The lowest-cost eligible model wins.

02 / endpoint

OpenAI-compatible endpoint

Point an existing client at the gateway base URL and keep the request shape. There is no provider-specific SDK to maintain.

03 / catalog

Automatic by default, override on demand

Standard requests never name a model. On a premium plan you can send a model field; it is validated against the catalog and billed as a fixed surcharge.

04 / security

Keys, boundaries, and rate limits

API keys are hashed and shown once. Provider credentials stay encrypted. Per-account and per-key limits are enforced before any provider work starts.

05 / visibility

Every request leaves a record

Routing mode, selected provider and model, latency, retries, and the price applied are written to an immutable usage record you can inspect.

06 / billing

Priced in requests per minute

Each tier caps throughput by RPM. Token counts are tracked for internal observability and margin control, never billed to you.

Model catalog

Models the router can reach

The registry maps every provider to its models, capabilities, and context window. Standard requests route automatically; premium plans can override.

Text generation

Chat and completion models served through the same endpoint.

GQ

Llama 3.3 70B

groq/llama-3.3-70b

JSONStreaming128K ctx
OA

GPT-4o mini

openai/gpt-4o-mini

VisionToolsJSONStreaming128K ctx
GG

Gemini 2.0 Flash

gemini/gemini-2.0-flash

VisionToolsJSONStreaming1M ctx
DS

DeepSeek Chat

deepseek/deepseek-chat

ToolsJSONStreaming64K ctx
OA

GPT-4o

openai/gpt-4o

VisionToolsJSONStreaming128K ctxPremium override
AN

Claude Sonnet 4

anthropic/claude-sonnet-4

VisionToolsJSONStreaming200K ctxPremium override
GQ

Llama 3.1 8B Instant

groq/llama-3.1-8b-instant

JSONStreaming128K ctx
GQ

Llama 3.2 11B Vision

groq/llama-3.2-11b-vision

VisionJSONStreaming128K ctx
GQ

Llama 3.2 90B Vision

groq/llama-3.2-90b-vision

VisionToolsJSONStreaming128K ctx
GQ

Llama 3 70B

groq/llama3-70b-8192

JSONStreaming8K ctx
GQ

Llama 3 8B

groq/llama3-8b-8192

JSONStreaming8K ctx
GQ

Gemma 2 9B

groq/gemma2-9b-it

JSONStreaming8K ctx
GQ

Llama 4 Scout 17B

groq/llama-4-scout-17b-16e

VisionToolsJSONStreaming128K ctx
GQ

Llama 4 Maverick 17B

groq/llama-4-maverick-17b-128e

VisionToolsJSONStreaming128K ctx
GQ

Qwen QwQ 32B

groq/qwen-qwq-32b

ToolsJSONStreaming128K ctx
GQ

DeepSeek R1 Distill Llama 70B

groq/deepseek-r1-distill-llama-70b

JSONStreaming128K ctx
GQ

Llama Guard 4 12B

groq/llama-guard-4-12b

ToolsJSON131K ctx
OA

GPT-4.1

openai/gpt-4.1

VisionToolsJSONStreaming1M ctxPremium override
OA

GPT-4.1 mini

openai/gpt-4.1-mini

VisionToolsJSONStreaming1M ctx
OA

GPT-4.1 nano

openai/gpt-4.1-nano

ToolsJSONStreaming1M ctx
OA

GPT-5

openai/gpt-5

VisionToolsJSONStreaming400K ctxPremium override
OA

GPT-5 mini

openai/gpt-5-mini

VisionToolsJSONStreaming400K ctx
OA

GPT-5 nano

openai/gpt-5-nano

ToolsJSONStreaming400K ctx
OA

o1

openai/o1

VisionToolsJSONStreaming200K ctxPremium override
OA

o1-mini

openai/o1-mini

JSONStreaming128K ctx
OA

o3

openai/o3

VisionToolsJSONStreaming200K ctxPremium override
OA

o3-mini

openai/o3-mini

ToolsJSONStreaming200K ctx
OA

o4-mini

openai/o4-mini

VisionToolsJSONStreaming200K ctx
GG

Gemini 2.0 Flash Lite

gemini/gemini-2.0-flash-lite

VisionToolsJSONStreaming1M ctx
GG

Gemini 2.5 Flash

gemini/gemini-2.5-flash

VisionToolsJSONStreaming1M ctx
GG

Gemini 2.5 Flash Lite

gemini/gemini-2.5-flash-lite

VisionToolsJSONStreaming1M ctx
GG

Gemini 2.5 Pro

gemini/gemini-2.5-pro

VisionToolsJSONStreaming1M ctxPremium override
GG

Gemini 1.5 Flash

gemini/gemini-1.5-flash

VisionToolsJSONStreaming1M ctx
GG

Gemini 1.5 Flash 8B

gemini/gemini-1.5-flash-8b

VisionJSONStreaming1M ctx
GG

Gemini 1.5 Pro

gemini/gemini-1.5-pro

VisionToolsJSONStreaming2M ctxPremium override
GG

Gemma 3 4B

gemini/gemma-3-4b-it

VisionJSONStreaming128K ctx
GG

Gemma 3 12B

gemini/gemma-3-12b-it

VisionJSONStreaming128K ctx
GG

Gemma 3 27B

gemini/gemma-3-27b-it

VisionToolsJSONStreaming128K ctx
AN

Claude Opus 4

anthropic/claude-opus-4

VisionToolsJSONStreaming200K ctxPremium override
AN

Claude Sonnet 4.5

anthropic/claude-sonnet-4.5

VisionToolsJSONStreaming200K ctxPremium override
AN

Claude Haiku 4.5

anthropic/claude-haiku-4.5

VisionToolsJSONStreaming200K ctx
AN

Claude 3.7 Sonnet

anthropic/claude-3-7-sonnet

VisionToolsJSONStreaming200K ctxPremium override
AN

Claude 3.5 Sonnet

anthropic/claude-3-5-sonnet

VisionToolsJSONStreaming200K ctxPremium override
AN

Claude 3.5 Haiku

anthropic/claude-3-5-haiku

ToolsJSONStreaming200K ctx
AN

Claude 3 Haiku

anthropic/claude-3-haiku

VisionJSONStreaming200K ctx
AN

Claude 3 Opus

anthropic/claude-3-opus

VisionToolsJSONStreaming200K ctxPremium override
DS

DeepSeek V3.1

deepseek/deepseek-v3.1

ToolsJSONStreaming128K ctx
DS

DeepSeek Reasoner

deepseek/deepseek-reasoner

JSONStreaming128K ctxPremium override
DS

DeepSeek R1 Distill 32B

deepseek/deepseek-r1-distill-32b

JSONStreaming128K ctx
DS

DeepSeek Coder

deepseek/deepseek-coder

ToolsJSONStreaming128K ctx
DS

DeepSeek Prover V2

deepseek/deepseek-prover-v2

JSONStreaming128K ctxPremium override
DS

DeepSeek VL2

deepseek/deepseek-vl2

VisionJSONStreaming32K ctx

Embeddings

Vector models for search and retrieval workloads.

OA

Embedding 3 Small

openai/text-embedding-3-small

Embeddings8K ctx
OA

Embedding 3 Large

openai/text-embedding-3-large

Embeddings8K ctx
GG

Embedding 004

gemini/text-embedding-004

Embeddings2K ctx
Under the hood

Every routing decision is inspectable

The response reports the actual model, the routing mode, the provider, and the latency, so you can see what ran without guessing.

  1. 01Inspect. Complexity, capabilities, and context size.
  2. 02Route. Pick the lowest-cost eligible model.
  3. 03Execute. Run inference against the chosen provider.
  4. 04Record. Write the actual model, latency, and price.
Read the API reference
response.json
{
  "model": "provider/actual-model",
  "gateway": {
    "routing": "auto",
    "provider": "provider-name",
    "latency_ms": 850,
    "premium": false
  }
}

Start free at 5 requests per minute

The free tier uses the same endpoint and the same routing behavior as production. Upgrade when the traffic does.

Operating rules

Built on rules, not promises

These hold on every request. They are the product's business rules, written down rather than implied.

01

Automatic by default

Automatic routing is the default and mandatory mode for standard requests.

02

No silent swap

A premium override is never quietly downgraded; an unmet override fails with a clear error.

03

RPM is the cap

The subscription caps requests per minute; customers are never billed per token.