anthropic / Chat
Claude Opus 5.5 Max
anthropic/claude-opus-5-5-xhighClaude Opus 5.5 with extra-high thinking — deepest reasoning
Catalog status
Enabled
Research checked
2026-10-07
Price / performance
What it costs to get this result
FrontierCode v1.1 (Main) · Source-reported results, not a composite ranking. Prices use the average of TokenBazaar input and output rates per million tokens; actual spend depends on usage. Test configurations differ.
Blended price / 1M tokens · log scale
Compare sourced models on this benchmark
Benchmark results
Results are source-attributed, not TokenBazaar measurements. Different tests and thinking variants are not interchangeable.
5 of 5 results
Terminal-Bench 4.0
Agentic · Third-party report
Test conditions and source
Explicitly the Xhigh effort configuration ('Claude Opus 5.5 (Xhigh, Default Fallback)') on AA's independently-run Terminal-Bench 4.0 harness — matches the requested Opus 5.5 xhigh condition exactly, not max/fast/default. Distinct from Anthropic's own self-reported xhigh figure of 66.4% already on file (different methodology/harness) and from AA's narrative max-effort figure of 59.6% quoted in AA's launch article.
Tested model: anthropic/claude-opus-5-5-xhigh
Artificial Analysis · Opus 5.5 (Xhigh) vs Opus 5 (High) comparison · Checked 2026-10-06
Terminal-Bench-Science 0.1
Agentic · Third-party report
Test conditions and source
Explicitly the xhigh effort configuration ('Claude Opus 5.5 (xhigh)'); AA also separately reports 59% at max effort for the same model on this benchmark — kept distinct and not merged. Independent AA harness, not Anthropic's own 58.7% (max effort) figure already on file.
Tested model: anthropic/claude-opus-5-5-xhigh
Artificial Analysis · Terminal-Bench-Science 0.1 leaderboard launch post (LinkedIn) · Checked 2026-10-06
CursorBench 4.0
Coding · Third-party report
Test conditions and source
Row explicitly labeled 'Opus 5.5 Extra High' (xhigh) at 56.0%, $6.98/task. Kept distinct from the existing 57.8% entry, which is Anthropic's own self-reported max-effort figure from the launch page, and from CursorBench's own 'Opus 5.5 Max' row (57.8%, $13.43/task) — same number as Anthropic's but a separately sourced/priced data point.
Tested model: anthropic/claude-opus-5-5-xhigh
Cursor · CursorBench 4.0 public leaderboard · Checked 2026-10-06
FrontierCode v1.1 (Main)
Coding · Third-party report
Test conditions and source
Cognition FrontierCode 1.1; Claude Opus 5.5; xhigh effort; claude-code harness; Main subset (100 tasks). Published weighted rubric score (new_score × 100), not the all-or-nothing pass rate; solution-source violations score zero. Source field new_score=0.5142. Publisher evaluation, not a TokenBazaar measurement; provider effort configuration is not a guarantee of the API's fixed thinking budget.
Tested model: anthropic/claude-opus-5-5-xhigh
Cognition · FrontierCode 1.1 original leaderboard · Checked 2026-10-07
FrontierCode v1.1 (Extended)
Coding · Third-party report
Test conditions and source
Cognition FrontierCode 1.1; Claude Opus 5.5; xhigh effort; claude-code harness; Extended subset (150 tasks). Published weighted rubric score (new_score × 100), not the all-or-nothing pass rate; solution-source violations score zero. Source field new_score=0.6353. Publisher evaluation, not a TokenBazaar measurement; provider effort configuration is not a guarantee of the API's fixed thinking budget.
Tested model: anthropic/claude-opus-5-5-xhigh
Cognition · FrontierCode 1.1 original leaderboard · Checked 2026-10-07
Model specifications
Context window
1M
Published provider limit
Maximum output
128K
Published provider limit
Knowledge cutoff
Jun 2026
Upstream model
Input and output modalities
Provider-reported formats; API compatibility is detailed below.
Input
Output
- Exact model ID
- anthropic/claude-opus-5-5-xhigh
- Model type
- Chat
- Context window
- 1,000,000 tokens
- Maximum output
- 128,000 tokens
- Knowledge cutoff
- Jun 2026
A context window is the total conversation budget, not a separate maximum input allowance; generated output and reasoning can use that budget. Published provider limits are not independently tested TokenBazaar request limits.
Versions and thinking levels
Choose from 8 enabled versions. The exact ID determines the version and its price; benchmark scores do not carry over between variants.
Features and tool support
TokenBazaar’s exposed interface, not every feature advertised by the provider. “Available” describes the implemented interface, not a successful test of every input or tool.
Reasoning configuration
Adaptive thinking; reasoning and answer share the output budget.
Selected effort: xhigh. Reasoning level is not a measured intelligence or speed score.
Coding and agent workflows
Text/code generation and customer-managed tool loops.
Coding benchmark results do not establish a hosted terminal, sandbox, autonomous browser or guaranteed task success.
Streaming answers
AvailableIncremental text through TokenBazaar’s chat endpoint; Claude also has native Messages access.
Customer-defined function tools
LimitedFunction schemas and tool results are forwarded to Claude. The chat bridge uses automatic tool selection; your application authorizes and executes every tool.
Forced or named tool selection
Not exposedThe chat bridge selects tools automatically. Native Messages does not expose forced or named tool selection; do not assume a tool will be called every turn.
Images and PDFs
LimitedThe chat bridge accepts image/PDF content. Provider input modalities are listed separately; file size, content and model limits still apply.
Hosted web search / browsing
Not exposedNo provider-hosted web search or web-fetch tool is exposed. Customer-owned search can be implemented as an authorized function tool.
Hosted code execution / computer use
Not exposedNo built-in execution environment or computer-control tool is provided. Code generation is not code execution.
Hosted file search / persistent assistants
Not exposedNo hosted file-search index, Assistants endpoint or persistent provider agent is exposed.
Batch / fine-tuning / cached-price discounts
Not exposedNo public batch or fine-tuning endpoint, or separate cached-token discount, is offered here.
Sources and verification
Provider specifications describe the upstream model, not a guarantee of every feature through TokenBazaar. Prices come from the enabled catalog; benchmark results belong to the exact tested model.
- Provider model specifications
Vendor-reported
Last checked - Anthropic · Claude Opus 5.5 launch evaluation
Vendor-reported
Last checked - Artificial Analysis · Claude Opus 5.5 release page
Third-party report
Last checked - Artificial Analysis · 'Claude Opus 5.5 takes the top spot' article
Third-party report
Last checked - Artificial Analysis · Opus 5.5 (Xhigh) vs Opus 5 (High) comparison
Third-party report
Last checked - Artificial Analysis · Terminal-Bench-Science 0.1 leaderboard launch post (LinkedIn)
Third-party report
Last checked - Cursor · CursorBench 4.0 public leaderboard
Third-party report
Last checked - Cognition · FrontierCode 1.1 original leaderboard
Third-party report
Last checked
Research notes (1)
- Published provider limits; maximum-capacity requests have not been independently exercised through TokenBaazar. Thinking-level variants use the same provider model; reasoning and answer tokens share the output budget.