01 / routing
The gateway chooses the model
Complexity, required capabilities, latency history, provider health, context size, and your plan all feed one decision. The lowest-cost eligible model wins.
Send an OpenAI-compatible request with no model field. Vechgate picks an eligible model for the task, falls back when a provider fails, and records which model actually ran.
curl https://api.vechgate.dev/v1/chat/completions \ -H "Authorization: Bearer $VECHGATE_API_KEY" \ -H "Content-Type: application/json" \ -d '{"messages": [{"role": "user", "content": "Summarize this release."}]}'
Why Vechgate
Automatic routing, a single compatible endpoint, and billing you can audit. These are the parts the product spec commits to.
01 / routing
Complexity, required capabilities, latency history, provider health, context size, and your plan all feed one decision. The lowest-cost eligible model wins.
02 / endpoint
Point an existing client at the gateway base URL and keep the request shape. There is no provider-specific SDK to maintain.
03 / catalog
Standard requests never name a model. On a premium plan you can send a model field; it is validated against the catalog and billed as a fixed surcharge.
04 / security
API keys are hashed and shown once. Provider credentials stay encrypted. Per-account and per-key limits are enforced before any provider work starts.
05 / visibility
Routing mode, selected provider and model, latency, retries, and the price applied are written to an immutable usage record you can inspect.
06 / billing
Each tier caps throughput by RPM. Token counts are tracked for internal observability and margin control, never billed to you.
Model catalog
The registry maps every provider to its models, capabilities, and context window. Standard requests route automatically; premium plans can override.
Chat and completion models served through the same endpoint.
groq/llama-3.3-70b
openai/gpt-4o-mini
gemini/gemini-2.0-flash
deepseek/deepseek-chat
openai/gpt-4o
anthropic/claude-sonnet-4
groq/llama-3.1-8b-instant
groq/llama-3.2-11b-vision
groq/llama-3.2-90b-vision
groq/llama3-70b-8192
groq/llama3-8b-8192
groq/gemma2-9b-it
groq/llama-4-scout-17b-16e
groq/llama-4-maverick-17b-128e
groq/qwen-qwq-32b
groq/deepseek-r1-distill-llama-70b
groq/llama-guard-4-12b
openai/gpt-4.1
openai/gpt-4.1-mini
openai/gpt-4.1-nano
openai/gpt-5
openai/gpt-5-mini
openai/gpt-5-nano
openai/o1
openai/o1-mini
openai/o3
openai/o3-mini
openai/o4-mini
gemini/gemini-2.0-flash-lite
gemini/gemini-2.5-flash
gemini/gemini-2.5-flash-lite
gemini/gemini-2.5-pro
gemini/gemini-1.5-flash
gemini/gemini-1.5-flash-8b
gemini/gemini-1.5-pro
gemini/gemma-3-4b-it
gemini/gemma-3-12b-it
gemini/gemma-3-27b-it
anthropic/claude-opus-4
anthropic/claude-sonnet-4.5
anthropic/claude-haiku-4.5
anthropic/claude-3-7-sonnet
anthropic/claude-3-5-sonnet
anthropic/claude-3-5-haiku
anthropic/claude-3-haiku
anthropic/claude-3-opus
deepseek/deepseek-v3.1
deepseek/deepseek-reasoner
deepseek/deepseek-r1-distill-32b
deepseek/deepseek-coder
deepseek/deepseek-prover-v2
deepseek/deepseek-vl2
Vector models for search and retrieval workloads.
openai/text-embedding-3-small
openai/text-embedding-3-large
gemini/text-embedding-004
The response reports the actual model, the routing mode, the provider, and the latency, so you can see what ran without guessing.
{
"model": "provider/actual-model",
"gateway": {
"routing": "auto",
"provider": "provider-name",
"latency_ms": 850,
"premium": false
}
}The free tier uses the same endpoint and the same routing behavior as production. Upgrade when the traffic does.
Operating rules
These hold on every request. They are the product's business rules, written down rather than implied.
Automatic routing is the default and mandatory mode for standard requests.
A premium override is never quietly downgraded; an unmet override fails with a clear error.
The subscription caps requests per minute; customers are never billed per token.