Lanes
A lane is a URL prefix that selects which catalog and routing rules a request uses. Append the standard OpenAI/Anthropic path after it. Every lane serves both wire protocols.
/oss — curated open-weight catalog
The default public lane, serving open-weight models only: curated coding flagships (Kimi, MiniMax, Qwen, DeepSeek, GLM and GPT-OSS) with live provider routing. The catalog is curated independently of upstream config, so model discovery is reliable for Claude Code Desktop, IDE, and CLI.
| Protocol | Endpoint |
|---|---|
| OpenAI Chat | POST /oss/v1/chat/completions |
| OpenAI Responses | POST /oss/v1/responses |
| Anthropic Messages | POST /oss/v1/messages |
| Embeddings | POST /oss/v1/embeddings |
| Discovery | GET /oss/v1/models |
Friendly aliases
The curated catalog exposes discovery aliases for IDE and Desktop pickers.
Each entry's display_name leads with the real
underlying model — e.g. GLM 5.2 · agentic coding — so the
names stay honest. Claude Code's picker (which only surfaces
claude* ids) is populated from these entries. For CLI and
scripts, prefer friendly ids like glm-5.2, kimi-k2.6,
and qwen3-coder. Call GET /oss/v1/models for the
freshest list.
| Open model | CLI id | Best for |
|---|---|---|
| GLM 5.2 | glm-5.2 | Agentic coding (1M context) |
| Kimi K2.7 Code | moonshotai/Kimi-K2.7-Code:preferred | Long-context agentic coding |
| Qwen3-Coder 480B | qwen3-coder | Agentic coding |
| DeepSeek V4-Pro | deepseek-v4-pro | Max reasoning |
| MiniMax M3 | MiniMaxAI/MiniMax-M3:preferred | Million-token coding |
| GPT-OSS 120B | gpt-oss-120b | Fast reasoning |
| GPT-OSS 20B | gpt-oss-20b | Fast & low-cost |
Raw ids & routing policies
You can also call any catalog model by its raw org/model id and
pin provider-selection behavior with a suffix:
| Suffix | Meaning |
|---|---|
:preferred | The gateway's preferred provider for that model. |
:fastest | Lowest-latency live provider. |
:cheapest | Lowest-cost live provider. |
:<provider> | Pin a specific provider by name. |
moonshotai/Kimi-K2.7-Code:preferred Qwen/Qwen3-Coder-480B-A35B-Instruct:fastest openai/gpt-oss-120b:cheapest
For cost control, route bulk or low-stakes work to :cheapest and reserve :fastest or :preferred for latency- and quality-sensitive requests.
/hf — full passthrough
Specify any Hugging Face model id and OpenGateway validates
and routes it to a provider. Discovery is served from the cron-built catalog
cache. Same wire protocols as /oss; passthrough is opt-in by
specifying an id and never changes the curated default surface. Ids are
validated and upstream identity is always scrubbed from responses and errors.
These public docs cover open-weight lanes only. Stay on /oss for curated, reliable open-weight work.