Docs / Concepts / Effort levels

Effort levels

Reasoning models expose a "how hard should I think" dial, and every API spells it differently. OpenGateway normalizes a single internal enum and translates it to whatever the upstream model expects — regardless of which inbound API you used.

The normalized enum

text
off · minimal · low · medium · high · max

The default is roughly medium. If the upstream model does not support an effort knob, the gateway drops it silently — it never hard-fails a coding client over an unsupported parameter.

How each API signals effort

TierOpenAI ChatOpenAI ResponsesAnthropic (4.6+)
Offreasoning_effort: "none"reasoning.effort: "none"omit thinking
Minimal"minimal""minimal"output_config.effort: "low"
Low"low""low""low"
Medium"medium""medium""medium"
High"high""high""high"
Max"xhigh""xhigh""xhigh" / "max"
Legacy thinking is deprecated

Anthropic's legacy thinking: {type:"enabled", budget_tokens} is deprecated. Modern clients send thinking: {type:"adaptive"} with output_config.effort. OpenGateway accepts both inbound and buckets the legacy form into the normalized enum.

Setting effort

typescript · openai chat
await client.chat.completions.create({
  model: "claude-sonnet-4-7",
  messages: [{ role: "user", content: "Find the bug." }],
  reasoning_effort: "high",
});
typescript · openai responses
await client.responses.create({
  model: "claude-haiku-4-8",
  input: "Diagnose this stack trace.",
  reasoning: { effort: "high" },
});
typescript · anthropic messages
await client.messages.create({
  model: "claude-sonnet-4-7",
  max_tokens: 1024,
  thinking: { type: "adaptive" },
  output_config: { effort: "high" },
  messages: [{ role: "user", content: "Review this diff." }],
});

Cross-API translation

The point of the normalized enum is that effort survives a protocol mismatch between the inbound API and the upstream model:

  • Claude Code (Anthropic thinking) → routed to an OpenAI-wire model ⇒ emitted as reasoning_effort.
  • Codex (OpenAI Responses reasoning.effort) → routed to an Anthropic-wire model ⇒ emitted as thinking / output_config.effort.

The branded SDK exposes a single effort option that maps to whichever wire your call uses.