API reference
One endpoint set, compatible with OpenAI-style clients. Send a request without a model field and the gateway routes it; add a model field on a premium plan to override.
https://api.vechgate.dev/v101 / request and response
A standard request, start to finish
No model field means automatic routing. The response carries the actual model and the routing metadata.
curl https://api.vechgate.dev/v1/chat/completions \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"messages": [
{ "role": "user", "content": "Write a Python function to sort a list." }
]
}'{
"id": "chatcmpl_xxxxxxxxx",
"object": "chat.completion",
"created": 1791028800,
"model": "provider/actual-model",
"choices": [
{
"index": 0,
"message": { "role": "assistant", "content": "Here is the explanation..." },
"finish_reason": "stop"
}
],
"usage": {
"prompt_tokens": 120,
"completion_tokens": 240,
"total_tokens": 360
},
"gateway": {
"routing": "auto",
"provider": "provider-name",
"latency_ms": 850,
"premium": false
}
}Gateway metadata can be disabled for strict client compatibility (PRD section 5.3).
02 / premium override
Request one specific model
The model field is validated against the catalog, priced before execution, and billed as a fixed surcharge per request.
curl https://api.vechgate.dev/v1/chat/completions \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "provider/model-name",
"messages": [
{ "role": "user", "content": "Analyze this complex codebase." }
]
}'A request without the entitlement is rejected with a clear error, never routed to a different model.
03 / endpoints
What the gateway exposes
| Method | Path | Purpose |
|---|---|---|
| POST | /v1/chat/completions | Main inference endpoint |
| GET | /v1/models | Available capabilities and permitted overrides |
| POST | /v1/embeddings | Generate embeddings |
| POST | /v1/responses | Unified response interface, if supported |
| GET | /v1/usage | Retrieve API usage |
| GET | /v1/balance | Retrieve account balance |
| GET | /v1/health | Gateway health |
| GET | /v1/pricing | Retrieve pricing information |
04 / errors
How failures are reported
| Response | When |
|---|---|
| HTTP 401 | Invalid API key |
| HTTP 429 | RPM exceeded, returns retry metadata |
| HTTP 400 | Unsupported capability |
| HTTP 503 | All eligible providers unavailable |
| HTTP 402 | Insufficient balance, rejected before execution |
| Error | Premium model not allowed, explicit entitlement error |
Failed requests follow a documented policy on whether they count against RPM. Streaming errors after headers are sent are reported in the stream error format.
Create a key and send the first request
The free tier includes 5 requests per minute against the same endpoint used in production.