anthropic / Chat
Claude Opus 5.5
anthropic/claude-opus-5-5Claude Opus 5.5 with medium thinking (default)
Catalog status
Enabled
Research checked
2026-10-07
Price / performance
What it costs to get this result
FrontierCode v1.1 (Main) · Source-reported results, not a composite ranking. Prices use the average of TokenBazaar input and output rates per million tokens; actual spend depends on usage. Test configurations differ.
Blended price / 1M tokens · log scale
Compare sourced models on this benchmark
Benchmark results
Results are source-attributed, not TokenBazaar measurements. Different tests and thinking variants are not interchangeable.
17 of 17 results
FrontierCode v1.1 (Main)
Coding · Vendor-reported
Test conditions and source
Provider launch evaluation with adaptive thinking at max effort and production safeguards; not the default medium catalog configuration.
Tested model: anthropic/claude-opus-5-5
Anthropic · Claude Opus 5.5 launch evaluation · Checked 2026-10-06
CursorBench 4.0
Coding · Vendor-reported
Test conditions and source
Provider launch evaluation with adaptive thinking at max effort and production safeguards; not the default medium catalog configuration.
Tested model: anthropic/claude-opus-5-5
Anthropic · Claude Opus 5.5 launch evaluation · Checked 2026-10-06
AutomationBench
Agentic · Vendor-reported
Test conditions and source
Provider launch evaluation with adaptive thinking at max effort and production safeguards; not the default medium catalog configuration.
Tested model: anthropic/claude-opus-5-5
Anthropic · Claude Opus 5.5 launch evaluation · Checked 2026-10-06
Humanity’s Last Exam (with tools)
Reasoning · Vendor-reported
Test conditions and source
Provider launch evaluation with adaptive thinking at max effort and production safeguards; not the default medium catalog configuration.
Tested model: anthropic/claude-opus-5-5
Anthropic · Claude Opus 5.5 launch evaluation · Checked 2026-10-06
Terminal-Bench-Science 0.1
Agentic · Vendor-reported
Test conditions and source
Provider launch evaluation with adaptive thinking at max effort and production safeguards; not the default medium catalog configuration.
Tested model: anthropic/claude-opus-5-5
Anthropic · Claude Opus 5.5 launch evaluation · Checked 2026-10-06
OSWorld 2.1 (partial)
Agentic · Vendor-reported
Test conditions and source
Provider launch evaluation with adaptive thinking at max effort and production safeguards; not the default medium catalog configuration.
Tested model: anthropic/claude-opus-5-5
Anthropic · Claude Opus 5.5 launch evaluation · Checked 2026-10-06
Chartography (with tools)
Reasoning · Vendor-reported
Test conditions and source
Provider launch evaluation with adaptive thinking at max effort and production safeguards; not the default medium catalog configuration.
Tested model: anthropic/claude-opus-5-5
Anthropic · Claude Opus 5.5 launch evaluation · Checked 2026-10-06
Terminal-Bench 4.0
Agentic · Vendor-reported
Test conditions and source
Launch report explicitly states xhigh effort, standard error ±2.6 points; production safeguards. Not a medium or fast-mode measurement.
Tested model: anthropic/claude-opus-5-5
Anthropic · Claude Opus 5.5 launch evaluation · Checked 2026-10-06
Artificial Analysis Intelligence Index
Reasoning · Third-party report
Test conditions and source
Independent AA run, 'Claude Opus 5.5 (Max, Default Fallback)' configuration — AA's max-effort-with-fallback harness, not the catalog medium default. Not comparable to the Xhigh config (56) or High config (54) also reported by AA.
Tested model: anthropic/claude-opus-5-5
Artificial Analysis · Claude Opus 5.5 release page · Checked 2026-10-06
Humanity's Last Exam
Reasoning · Third-party report
Test conditions and source
Independent AA evaluation, reported for the top-scoring (max-effort/fallback) AA configuration used for the Intelligence Index; AA states this is the new best score on this eval, ahead of Fable 5.1's prior 59.1%. Effort level for this specific sub-score is implied (Intelligence Index default config), not separately confirmed for every effort tier.
Tested model: anthropic/claude-opus-5-5
Artificial Analysis · 'Claude Opus 5.5 takes the top spot' article · Checked 2026-10-06
SciCode
Coding · Third-party report
Test conditions and source
Independent AA evaluation under the Intelligence Index max-effort/fallback configuration; AA reports this surpasses Fable 5.1's prior best of 63.1%.
Tested model: anthropic/claude-opus-5-5
Artificial Analysis · 'Claude Opus 5.5 takes the top spot' article · Checked 2026-10-06
AA-Briefcase v1.1 (Elo)
Agentic · Third-party report
Test conditions and source
AA's private frontier knowledge-work evaluation using the open-source Stirrup reference agent harness; max-effort/fallback configuration (the configuration AA headlines for Opus 5.5). Not the Xhigh-specific Briefcase score (1768) also published by AA.
Tested model: anthropic/claude-opus-5-5
Artificial Analysis · 'Claude Opus 5.5 takes the top spot' article · Checked 2026-10-06
Terminal-Bench 4.0
Agentic · Third-party report
Test conditions and source
Explicitly the Xhigh effort configuration ('Claude Opus 5.5 (Xhigh, Default Fallback)') on AA's independently-run Terminal-Bench 4.0 harness — matches the requested Opus 5.5 xhigh condition exactly, not max/fast/default. Distinct from Anthropic's own self-reported xhigh figure of 66.4% already on file (different methodology/harness) and from AA's narrative max-effort figure of 59.6% quoted in AA's launch article.
Tested model: anthropic/claude-opus-5-5
Artificial Analysis · Opus 5.5 (Xhigh) vs Opus 5 (High) comparison · Checked 2026-10-06
Terminal-Bench-Science 0.1
Agentic · Third-party report
Test conditions and source
Explicitly the xhigh effort configuration ('Claude Opus 5.5 (xhigh)'); AA also separately reports 59% at max effort for the same model on this benchmark — kept distinct and not merged. Independent AA harness, not Anthropic's own 58.7% (max effort) figure already on file.
Tested model: anthropic/claude-opus-5-5
Artificial Analysis · Terminal-Bench-Science 0.1 leaderboard launch post (LinkedIn) · Checked 2026-10-06
CursorBench 4.0
Coding · Third-party report
Test conditions and source
Row explicitly labeled 'Opus 5.5 Extra High' (xhigh) at 56.0%, $6.98/task. Kept distinct from the existing 57.8% entry, which is Anthropic's own self-reported max-effort figure from the launch page, and from CursorBench's own 'Opus 5.5 Max' row (57.8%, $13.43/task) — same number as Anthropic's but a separately sourced/priced data point.
Tested model: anthropic/claude-opus-5-5
Cursor · CursorBench 4.0 public leaderboard · Checked 2026-10-06
FrontierCode v1.1 (Main)
Coding · Third-party report
Test conditions and source
Cognition FrontierCode 1.1; Claude Opus 5.5; medium effort; claude-code harness; Main subset (100 tasks). Published weighted rubric score (new_score × 100), not the all-or-nothing pass rate; solution-source violations score zero. Source field new_score=0.5464. Publisher evaluation, not a TokenBazaar measurement; provider effort configuration is not a guarantee of the API's fixed thinking budget.
Tested model: anthropic/claude-opus-5-5
Cognition · FrontierCode 1.1 original leaderboard · Checked 2026-10-07
FrontierCode v1.1 (Extended)
Coding · Third-party report
Test conditions and source
Cognition FrontierCode 1.1; Claude Opus 5.5; medium effort; claude-code harness; Extended subset (150 tasks). Published weighted rubric score (new_score × 100), not the all-or-nothing pass rate; solution-source violations score zero. Source field new_score=0.6527. Publisher evaluation, not a TokenBazaar measurement; provider effort configuration is not a guarantee of the API's fixed thinking budget.
Tested model: anthropic/claude-opus-5-5
Cognition · FrontierCode 1.1 original leaderboard · Checked 2026-10-07
Model specifications
Context window
1M
Published provider limit
Maximum output
128K
Published provider limit
Knowledge cutoff
Jun 2026
Upstream model
Input and output modalities
Provider-reported formats; API compatibility is detailed below.
Input
Output
- Exact model ID
- anthropic/claude-opus-5-5
- Model type
- Chat
- Context window
- 1,000,000 tokens
- Maximum output
- 128,000 tokens
- Knowledge cutoff
- Jun 2026
A context window is the total conversation budget, not a separate maximum input allowance; generated output and reasoning can use that budget. Published provider limits are not independently tested TokenBazaar request limits.
Versions and thinking levels
Choose from 8 enabled versions. The exact ID determines the version and its price; benchmark scores do not carry over between variants.
Features and tool support
TokenBazaar’s exposed interface, not every feature advertised by the provider. “Available” describes the implemented interface, not a successful test of every input or tool.
Reasoning configuration
Adaptive thinking; reasoning and answer share the output budget.
Selected effort: medium (default catalog version). Reasoning level is not a measured intelligence or speed score.
Coding and agent workflows
Text/code generation and customer-managed tool loops.
Coding benchmark results do not establish a hosted terminal, sandbox, autonomous browser or guaranteed task success.
Streaming answers
AvailableIncremental text through TokenBazaar’s chat endpoint; Claude also has native Messages access.
Customer-defined function tools
LimitedFunction schemas and tool results are forwarded to Claude. The chat bridge uses automatic tool selection; your application authorizes and executes every tool.
Forced or named tool selection
Not exposedThe chat bridge selects tools automatically. Native Messages does not expose forced or named tool selection; do not assume a tool will be called every turn.
Images and PDFs
LimitedThe chat bridge accepts image/PDF content. Provider input modalities are listed separately; file size, content and model limits still apply.
Hosted web search / browsing
Not exposedNo provider-hosted web search or web-fetch tool is exposed. Customer-owned search can be implemented as an authorized function tool.
Hosted code execution / computer use
Not exposedNo built-in execution environment or computer-control tool is provided. Code generation is not code execution.
Hosted file search / persistent assistants
Not exposedNo hosted file-search index, Assistants endpoint or persistent provider agent is exposed.
Batch / fine-tuning / cached-price discounts
Not exposedNo public batch or fine-tuning endpoint, or separate cached-token discount, is offered here.
Sources and verification
Provider specifications describe the upstream model, not a guarantee of every feature through TokenBazaar. Prices come from the enabled catalog; benchmark results belong to the exact tested model.
- Provider model specifications
Vendor-reported
Last checked - Anthropic · Claude Opus 5.5 launch evaluation
Vendor-reported
Last checked - Artificial Analysis · Claude Opus 5.5 release page
Third-party report
Last checked - Artificial Analysis · 'Claude Opus 5.5 takes the top spot' article
Third-party report
Last checked - Artificial Analysis · Opus 5.5 (Xhigh) vs Opus 5 (High) comparison
Third-party report
Last checked - Artificial Analysis · Terminal-Bench-Science 0.1 leaderboard launch post (LinkedIn)
Third-party report
Last checked - Cursor · CursorBench 4.0 public leaderboard
Third-party report
Last checked - Cognition · FrontierCode 1.1 original leaderboard
Third-party report
Last checked
Research notes (1)
- Published provider limits; maximum-capacity requests have not been independently exercised through TokenBaazar. Thinking-level variants use the same provider model; reasoning and answer tokens share the output budget.