1M context
Several models on /oss support context windows up to
1M tokens. The gateway advertises real per-model windows
and routes large requests to a provider whose window actually fits.
Honest, per-model windows
Model discovery advertises a context trio truthfully, per model:
{
"context_window": 1000000,
"context_window_tokens": 1000000,
"max_context_tokens": 1000000
}
The default discovery floor is 200K, but long-context models advertise their real maximum so clients enable long-context paths. The window is per provider × model — the same model can be 1M tokens on one provider and 64K on another — so OpenGateway routes to a provider whose window meets the request.
If a request exceeds the resolved model's maximum, OpenGateway returns a clean 4xx rather than letting the upstream truncate silently.
Anthropic 1M, headerless
On modern Anthropic-wire models, 1M context is GA and headerless — no beta flag
is required. The legacy header anthropic-beta: context-1m-2025-08-07
is retired, but OpenGateway accepts and ignores it — it will
never 400 on it — so older Claude Code/SDK configs keep working:
curl "https://api.opengateway.one/oss/v1/messages" \ -H "Authorization: Bearer $OPENGATEWAY_API_KEY" \ -H "anthropic-version: 2023-06-01" \ -H "anthropic-beta: context-1m-2025-08-07" \ -H "Content-Type: application/json" \ -d '{ "model": "claude-sonnet-4-7", "max_tokens": 512, "messages": [{ "role": "user", "content": "Summarize this repo." }] }'
OpenAI clients
OpenAI Chat/Responses have no context header — clients rely on the model's
advertised window. OpenGateway routes to a large-window provider and clamps
max_tokens / max_output_tokens sanely. Codex
additionally honors a client-side model_context_window that it
truncates to — set it to match the model you target (see
Codex setup).