Models and Pricing

Compare models

Choose two or three models, including exact thinking and fast versions. Copy the URL to share your selection.

ComparisonClaude Opus 5.5
anthropic/claude-opus-5-5
Choose model 2
TypeChat—
TokenBaazar price · 50% off₹192.00 in / ₹960.00 out per 1M tokens—
Market price₹384.00 in / ₹1920.00 out per 1M tokens—
Context window (provider)1,000,000 tokens—
Input token limit (provider)Within the context window—
Maximum output128,000 tokens—
Knowledge cutoffJun 2026—
Vendor input modalitiesText, Image—
Vendor output modalitiesText—
TokenBaazar endpointsPOST /v1/chat/completions · POST /v1/messages—
TokenBaazar inputText, Image/PDF content through the chat bridge—
TokenBaazar outputText—
Benchmark evidence17 source-linked results—
FrontierCode v1.1 (Main)
54.4%

Provider launch evaluation with adaptive thinking at max effort and production safeguards; not the default medium catalog configuration.

Anthropic · Claude Opus 5.5 launch evaluation
54.64%

Cognition FrontierCode 1.1; Claude Opus 5.5; medium effort; claude-code harness; Main subset (100 tasks). Published weighted rubric score (new_score × 100), not the all-or-nothing pass rate; solution-source violations score zero. Source field new_score=0.5464. Publisher evaluation, not a TokenBazaar measurement; provider effort configuration is not a guarantee of the API's fixed thinking budget.

Cognition · FrontierCode 1.1 original leaderboard
—
CursorBench 4.0
57.8%

Provider launch evaluation with adaptive thinking at max effort and production safeguards; not the default medium catalog configuration.

Anthropic · Claude Opus 5.5 launch evaluation
56%

Row explicitly labeled 'Opus 5.5 Extra High' (xhigh) at 56.0%, $6.98/task. Kept distinct from the existing 57.8% entry, which is Anthropic's own self-reported max-effort figure from the launch page, and from CursorBench's own 'Opus 5.5 Max' row (57.8%, $13.43/task) — same number as Anthropic's but a separately sourced/priced data point.

Cursor · CursorBench 4.0 public leaderboard
—
AutomationBench
40%

Provider launch evaluation with adaptive thinking at max effort and production safeguards; not the default medium catalog configuration.

Anthropic · Claude Opus 5.5 launch evaluation
—
Humanity’s Last Exam (with tools)
67.7%

Provider launch evaluation with adaptive thinking at max effort and production safeguards; not the default medium catalog configuration.

Anthropic · Claude Opus 5.5 launch evaluation
—
Terminal-Bench-Science 0.1
58.7%

Provider launch evaluation with adaptive thinking at max effort and production safeguards; not the default medium catalog configuration.

Anthropic · Claude Opus 5.5 launch evaluation
62%

Explicitly the xhigh effort configuration ('Claude Opus 5.5 (xhigh)'); AA also separately reports 59% at max effort for the same model on this benchmark — kept distinct and not merged. Independent AA harness, not Anthropic's own 58.7% (max effort) figure already on file.

Artificial Analysis · Terminal-Bench-Science 0.1 leaderboard launch post (LinkedIn)
—
OSWorld 2.1 (partial)
81.8%

Provider launch evaluation with adaptive thinking at max effort and production safeguards; not the default medium catalog configuration.

Anthropic · Claude Opus 5.5 launch evaluation
—
Chartography (with tools)
89%

Provider launch evaluation with adaptive thinking at max effort and production safeguards; not the default medium catalog configuration.

Anthropic · Claude Opus 5.5 launch evaluation
—
Terminal-Bench 4.0
66.4%

Launch report explicitly states xhigh effort, standard error ±2.6 points; production safeguards. Not a medium or fast-mode measurement.

Anthropic · Claude Opus 5.5 launch evaluation
60%

Explicitly the Xhigh effort configuration ('Claude Opus 5.5 (Xhigh, Default Fallback)') on AA's independently-run Terminal-Bench 4.0 harness — matches the requested Opus 5.5 xhigh condition exactly, not max/fast/default. Distinct from Anthropic's own self-reported xhigh figure of 66.4% already on file (different methodology/harness) and from AA's narrative max-effort figure of 59.6% quoted in AA's launch article.

Artificial Analysis · Opus 5.5 (Xhigh) vs Opus 5 (High) comparison
—
Artificial Analysis Intelligence Index
58index

Independent AA run, 'Claude Opus 5.5 (Max, Default Fallback)' configuration — AA's max-effort-with-fallback harness, not the catalog medium default. Not comparable to the Xhigh config (56) or High config (54) also reported by AA.

Artificial Analysis · Claude Opus 5.5 release page
—
Humanity's Last Exam
61.4%

Independent AA evaluation, reported for the top-scoring (max-effort/fallback) AA configuration used for the Intelligence Index; AA states this is the new best score on this eval, ahead of Fable 5.1's prior 59.1%. Effort level for this specific sub-score is implied (Intelligence Index default config), not separately confirmed for every effort tier.

Artificial Analysis · 'Claude Opus 5.5 takes the top spot' article
—
SciCode
66.9%

Independent AA evaluation under the Intelligence Index max-effort/fallback configuration; AA reports this surpasses Fable 5.1's prior best of 63.1%.

Artificial Analysis · 'Claude Opus 5.5 takes the top spot' article
—
AA-Briefcase v1.1 (Elo)
1822Elo

AA's private frontier knowledge-work evaluation using the open-source Stirrup reference agent harness; max-effort/fallback configuration (the configuration AA headlines for Opus 5.5). Not the Xhigh-specific Briefcase score (1768) also published by AA.

Artificial Analysis · 'Claude Opus 5.5 takes the top spot' article
—
FrontierCode v1.1 (Extended)
65.27%

Cognition FrontierCode 1.1; Claude Opus 5.5; medium effort; claude-code harness; Extended subset (150 tasks). Published weighted rubric score (new_score × 100), not the all-or-nothing pass rate; solution-source violations score zero. Source field new_score=0.6527. Publisher evaluation, not a TokenBazaar measurement; provider effort configuration is not a guarantee of the API's fixed thinking budget.

Cognition · FrontierCode 1.1 original leaderboard
—

No overall score, winner or cost-versus-score frontier is calculated without comparable, attributable benchmark evidence. Different test configurations appear separately. Vendor capabilities are distinct from TokenBaazar API support.

— means no published value or no matching capability. Provider token limits are not tested gateway guarantees; TokenBaazar media controls and response formats are listed separately.

Claude Opus 5.5

Sources and verification

Provider specifications describe the upstream model, not a guarantee of every feature through TokenBazaar. Prices come from the enabled catalog; benchmark results belong to the exact tested model.

Research notes (1)
  • Published provider limits; maximum-capacity requests have not been independently exercised through TokenBaazar. Thinking-level variants use the same provider model; reasoning and answer tokens share the output budget.