Compare model

anthropic / Chat

Claude Opus 5.5

anthropic/claude-opus-5-5

Claude Opus 5.5 with medium thinking (default)

Catalog status

Enabled

Research checked

2026-10-07

Price / performance

What it costs to get this result

FrontierCode v1.1 (Main) · Source-reported results, not a composite ranking. Prices use the average of TokenBazaar input and output rates per million tokens; actual spend depends on usage. Test configurations differ.

Claude Opus 5.554.4% · 43 sourced models

Blended price / 1M tokens · log scale

● Current model● Other modelsDashed line: price/score frontier among plotted models
Compare sourced models on this benchmark

Benchmark results

17 reported results

Results are source-attributed, not TokenBazaar measurements. Different tests and thinking variants are not interchangeable.

17 of 17 results

FrontierCode v1.1 (Main)

Coding · Vendor-reported

Test conditions and source

Provider launch evaluation with adaptive thinking at max effort and production safeguards; not the default medium catalog configuration.

Tested model: anthropic/claude-opus-5-5

Anthropic · Claude Opus 5.5 launch evaluation · Checked 2026-10-06

54.4%

CursorBench 4.0

Coding · Vendor-reported

Test conditions and source

Provider launch evaluation with adaptive thinking at max effort and production safeguards; not the default medium catalog configuration.

Tested model: anthropic/claude-opus-5-5

Anthropic · Claude Opus 5.5 launch evaluation · Checked 2026-10-06

57.8%

AutomationBench

Agentic · Vendor-reported

Test conditions and source

Provider launch evaluation with adaptive thinking at max effort and production safeguards; not the default medium catalog configuration.

Tested model: anthropic/claude-opus-5-5

Anthropic · Claude Opus 5.5 launch evaluation · Checked 2026-10-06

40%

Humanity’s Last Exam (with tools)

Reasoning · Vendor-reported

Test conditions and source

Provider launch evaluation with adaptive thinking at max effort and production safeguards; not the default medium catalog configuration.

Tested model: anthropic/claude-opus-5-5

Anthropic · Claude Opus 5.5 launch evaluation · Checked 2026-10-06

67.7%

Terminal-Bench-Science 0.1

Agentic · Vendor-reported

Test conditions and source

Provider launch evaluation with adaptive thinking at max effort and production safeguards; not the default medium catalog configuration.

Tested model: anthropic/claude-opus-5-5

Anthropic · Claude Opus 5.5 launch evaluation · Checked 2026-10-06

58.7%

OSWorld 2.1 (partial)

Agentic · Vendor-reported

Test conditions and source

Provider launch evaluation with adaptive thinking at max effort and production safeguards; not the default medium catalog configuration.

Tested model: anthropic/claude-opus-5-5

Anthropic · Claude Opus 5.5 launch evaluation · Checked 2026-10-06

81.8%

Chartography (with tools)

Reasoning · Vendor-reported

Test conditions and source

Provider launch evaluation with adaptive thinking at max effort and production safeguards; not the default medium catalog configuration.

Tested model: anthropic/claude-opus-5-5

Anthropic · Claude Opus 5.5 launch evaluation · Checked 2026-10-06

89%

Terminal-Bench 4.0

Agentic · Vendor-reported

Test conditions and source

Launch report explicitly states xhigh effort, standard error ±2.6 points; production safeguards. Not a medium or fast-mode measurement.

Tested model: anthropic/claude-opus-5-5

Anthropic · Claude Opus 5.5 launch evaluation · Checked 2026-10-06

66.4%

Artificial Analysis Intelligence Index

Reasoning · Third-party report

Test conditions and source

Independent AA run, 'Claude Opus 5.5 (Max, Default Fallback)' configuration — AA's max-effort-with-fallback harness, not the catalog medium default. Not comparable to the Xhigh config (56) or High config (54) also reported by AA.

Tested model: anthropic/claude-opus-5-5

Artificial Analysis · Claude Opus 5.5 release page · Checked 2026-10-06

58index

Humanity's Last Exam

Reasoning · Third-party report

Test conditions and source

Independent AA evaluation, reported for the top-scoring (max-effort/fallback) AA configuration used for the Intelligence Index; AA states this is the new best score on this eval, ahead of Fable 5.1's prior 59.1%. Effort level for this specific sub-score is implied (Intelligence Index default config), not separately confirmed for every effort tier.

Tested model: anthropic/claude-opus-5-5

Artificial Analysis · 'Claude Opus 5.5 takes the top spot' article · Checked 2026-10-06

61.4%

SciCode

Coding · Third-party report

Test conditions and source

Independent AA evaluation under the Intelligence Index max-effort/fallback configuration; AA reports this surpasses Fable 5.1's prior best of 63.1%.

Tested model: anthropic/claude-opus-5-5

Artificial Analysis · 'Claude Opus 5.5 takes the top spot' article · Checked 2026-10-06

66.9%

AA-Briefcase v1.1 (Elo)

Agentic · Third-party report

Test conditions and source

AA's private frontier knowledge-work evaluation using the open-source Stirrup reference agent harness; max-effort/fallback configuration (the configuration AA headlines for Opus 5.5). Not the Xhigh-specific Briefcase score (1768) also published by AA.

Tested model: anthropic/claude-opus-5-5

Artificial Analysis · 'Claude Opus 5.5 takes the top spot' article · Checked 2026-10-06

1822Elo

Terminal-Bench 4.0

Agentic · Third-party report

Test conditions and source

Explicitly the Xhigh effort configuration ('Claude Opus 5.5 (Xhigh, Default Fallback)') on AA's independently-run Terminal-Bench 4.0 harness — matches the requested Opus 5.5 xhigh condition exactly, not max/fast/default. Distinct from Anthropic's own self-reported xhigh figure of 66.4% already on file (different methodology/harness) and from AA's narrative max-effort figure of 59.6% quoted in AA's launch article.

Tested model: anthropic/claude-opus-5-5

Artificial Analysis · Opus 5.5 (Xhigh) vs Opus 5 (High) comparison · Checked 2026-10-06

60%

Terminal-Bench-Science 0.1

Agentic · Third-party report

Test conditions and source

Explicitly the xhigh effort configuration ('Claude Opus 5.5 (xhigh)'); AA also separately reports 59% at max effort for the same model on this benchmark — kept distinct and not merged. Independent AA harness, not Anthropic's own 58.7% (max effort) figure already on file.

Tested model: anthropic/claude-opus-5-5

Artificial Analysis · Terminal-Bench-Science 0.1 leaderboard launch post (LinkedIn) · Checked 2026-10-06

62%

CursorBench 4.0

Coding · Third-party report

Test conditions and source

Row explicitly labeled 'Opus 5.5 Extra High' (xhigh) at 56.0%, $6.98/task. Kept distinct from the existing 57.8% entry, which is Anthropic's own self-reported max-effort figure from the launch page, and from CursorBench's own 'Opus 5.5 Max' row (57.8%, $13.43/task) — same number as Anthropic's but a separately sourced/priced data point.

Tested model: anthropic/claude-opus-5-5

Cursor · CursorBench 4.0 public leaderboard · Checked 2026-10-06

56%

FrontierCode v1.1 (Main)

Coding · Third-party report

Test conditions and source

Cognition FrontierCode 1.1; Claude Opus 5.5; medium effort; claude-code harness; Main subset (100 tasks). Published weighted rubric score (new_score × 100), not the all-or-nothing pass rate; solution-source violations score zero. Source field new_score=0.5464. Publisher evaluation, not a TokenBazaar measurement; provider effort configuration is not a guarantee of the API's fixed thinking budget.

Tested model: anthropic/claude-opus-5-5

Cognition · FrontierCode 1.1 original leaderboard · Checked 2026-10-07

54.64%

FrontierCode v1.1 (Extended)

Coding · Third-party report

Test conditions and source

Cognition FrontierCode 1.1; Claude Opus 5.5; medium effort; claude-code harness; Extended subset (150 tasks). Published weighted rubric score (new_score × 100), not the all-or-nothing pass rate; solution-source violations score zero. Source field new_score=0.6527. Publisher evaluation, not a TokenBazaar measurement; provider effort configuration is not a guarantee of the API's fixed thinking budget.

Tested model: anthropic/claude-opus-5-5

Cognition · FrontierCode 1.1 original leaderboard · Checked 2026-10-07

65.27%

Model specifications

Context window

1M

Published provider limit

Maximum output

128K

Published provider limit

Knowledge cutoff

Jun 2026

Upstream model

Input and output modalities

Provider-reported formats; API compatibility is detailed below.

Input

Text
Image

Output

Text
Exact model ID
anthropic/claude-opus-5-5
Model type
Chat
Context window
1,000,000 tokens
Maximum output
128,000 tokens
Knowledge cutoff
Jun 2026

A context window is the total conversation budget, not a separate maximum input allowance; generated output and reasoning can use that budget. Published provider limits are not independently tested TokenBazaar request limits.

Versions and thinking levels

Choose from 8 enabled versions. The exact ID determines the version and its price; benchmark scores do not carry over between variants.

Claude Opus 5.5Selected₹192.00 in / ₹960.00 out per 1M tokens
Claude Opus 5.5 Fast₹384.00 in / ₹1920.00 out per 1M tokens
Claude Opus 5.5 Fast High₹384.00 in / ₹1920.00 out per 1M tokens
Claude Opus 5.5 Fast Low₹384.00 in / ₹1920.00 out per 1M tokens
Claude Opus 5.5 Fast Max₹384.00 in / ₹1920.00 out per 1M tokens
Claude Opus 5.5 High₹192.00 in / ₹960.00 out per 1M tokens
Claude Opus 5.5 Low₹192.00 in / ₹960.00 out per 1M tokens
Claude Opus 5.5 Max₹192.00 in / ₹960.00 out per 1M tokens

Features and tool support

TokenBazaar’s exposed interface, not every feature advertised by the provider. “Available” describes the implemented interface, not a successful test of every input or tool.

Reasoning configuration

Adaptive thinking; reasoning and answer share the output budget.

Selected effort: medium (default catalog version). Reasoning level is not a measured intelligence or speed score.

Coding and agent workflows

Text/code generation and customer-managed tool loops.

Coding benchmark results do not establish a hosted terminal, sandbox, autonomous browser or guaranteed task success.

Streaming answers

Available

Incremental text through TokenBazaar’s chat endpoint; Claude also has native Messages access.

Customer-defined function tools

Limited

Function schemas and tool results are forwarded to Claude. The chat bridge uses automatic tool selection; your application authorizes and executes every tool.

Forced or named tool selection

Not exposed

The chat bridge selects tools automatically. Native Messages does not expose forced or named tool selection; do not assume a tool will be called every turn.

Images and PDFs

Limited

The chat bridge accepts image/PDF content. Provider input modalities are listed separately; file size, content and model limits still apply.

Hosted web search / browsing

Not exposed

No provider-hosted web search or web-fetch tool is exposed. Customer-owned search can be implemented as an authorized function tool.

Hosted code execution / computer use

Not exposed

No built-in execution environment or computer-control tool is provided. Code generation is not code execution.

Hosted file search / persistent assistants

Not exposed

No hosted file-search index, Assistants endpoint or persistent provider agent is exposed.

Batch / fine-tuning / cached-price discounts

Not exposed

No public batch or fine-tuning endpoint, or separate cached-token discount, is offered here.

Sources and verification

Provider specifications describe the upstream model, not a guarantee of every feature through TokenBazaar. Prices come from the enabled catalog; benchmark results belong to the exact tested model.

Research notes (1)
  • Published provider limits; maximum-capacity requests have not been independently exercised through TokenBaazar. Thinking-level variants use the same provider model; reasoning and answer tokens share the output budget.