anthropic / Chat
Claude Fable 5.1
anthropic/claude-fable-5-1Claude Fable 5.1 with medium thinking (default)
Catalog status
Enabled
Research checked
2026-10-07
Price / performance
What it costs to get this result
FrontierCode v1.1 (Extended) · Source-reported results, not a composite ranking. Prices use the average of TokenBazaar input and output rates per million tokens; actual spend depends on usage. Test configurations differ.
Blended price / 1M tokens · log scale
Compare sourced models on this benchmark
Benchmark results
Results are source-attributed, not TokenBazaar measurements. Different tests and thinking variants are not interchangeable.
33 of 33 results
SWE-bench Pro
Coding · Vendor-reported
Test conditions and source
Provider model evaluated with adaptive thinking at max effort; average over five trials. This is not the catalog default medium-effort configuration.
Tested model: anthropic/claude-fable-5-1
Anthropic Fable 5.1 system card · sections 8.1–8.11 · Checked 2026-10-06
DeepSWE v1.1
Coding · Vendor-reported
Test conditions and source
Provider model evaluated with adaptive thinking at max effort; average over five trials. This is not the catalog default medium-effort configuration.
Tested model: anthropic/claude-fable-5-1
Anthropic Fable 5.1 system card · sections 8.1–8.11 · Checked 2026-10-06
SWE-bench Multilingual
Coding · Vendor-reported
Test conditions and source
Provider model evaluated with adaptive thinking at max effort; average over five trials. This is not the catalog default medium-effort configuration.
Tested model: anthropic/claude-fable-5-1
Anthropic Fable 5.1 system card · sections 8.1–8.11 · Checked 2026-10-06
SWE-bench Multimodal
Coding · Vendor-reported
Test conditions and source
Provider model evaluated with adaptive thinking at max effort; average over five trials. This is not the catalog default medium-effort configuration.
Tested model: anthropic/claude-fable-5-1
Anthropic Fable 5.1 system card · sections 8.1–8.11 · Checked 2026-10-06
CursorBench 3.2.0
Coding · Vendor-reported
Test conditions and source
System card section 8.8 reports Cursor’s production agent harness at max effort; historical 3.2.0 test, not CursorBench 4.0.
Tested model: anthropic/claude-fable-5-1
Anthropic Fable 5.1 system card · sections 8.1–8.11 · Checked 2026-10-06
ProgramBench
Coding · Vendor-reported
Test conditions and source
Section 8.11.1: 166 filtered tasks; hidden test pass rate, mini-swe-agent without the six-hour time limit; max effort. Not interchangeable with the unfiltered public evaluation.
Tested model: anthropic/claude-fable-5-1
Anthropic Fable 5.1 system card · sections 8.1–8.11 · Checked 2026-10-06
FrontierSWE v2
Coding · Third-party report
Test conditions and source
Published current mean reward 0.5653 over 170 runs, displayed as a percentage; not the older supplied 56.3 result. Effort is not established by this leaderboard.
Tested model: anthropic/claude-fable-5-1
Proximal · FrontierSWE public leaderboard · Checked 2026-10-06
LiveCodeBench (Vals)
Coding · Third-party report
Test conditions and source
Vals evaluation: max compute effort, temperature 1, maximum output 128,000; reported uncertainty ±0.86 percentage points. Not the default medium version.
Tested model: anthropic/claude-fable-5-1
Vals AI · Claude Fable 5.1 evaluation · Checked 2026-10-06
CursorBench 4.0
Coding · Third-party report
Test conditions and source
Cursor’s live table explicitly lists Fable 5.1 Max at 51.8; Extra High is a separate 51.6 result. Max must not be equated with the catalog xhigh variant.
Tested model: anthropic/claude-fable-5-1
CursorBench public evaluation · Checked 2026-10-06
CursorBench 4.0
Coding · Third-party report
Test conditions and source
Row explicitly labeled 'Fable 5.1 Extra High' (xhigh) at 51.6%, $13.01/task, confirmed directly on Cursor's live table. Distinct from the already-recorded 51.8% 'Fable 5.1 Max' row and the 49.2% 'Fable 5.1 High' row on the same leaderboard — effort levels are not interchangeable.
Tested model: anthropic/claude-fable-5-1
Cursor · CursorBench 4.0 public leaderboard · Checked 2026-10-06
ARC-AGI-1
Reasoning · Vendor-reported
Test conditions and source
Section 8.16: ARC Prize verified semi-private dataset, max effort. Not the medium-effort default catalog version.
Tested model: anthropic/claude-fable-5-1
Anthropic Fable 5.1 system card · sections 8.1–8.11 · Checked 2026-10-06
ARC-AGI-2
Reasoning · Vendor-reported
Test conditions and source
Section 8.16: ARC Prize verified semi-private dataset, max effort. Not the medium-effort default catalog version.
Tested model: anthropic/claude-fable-5-1
Anthropic Fable 5.1 system card · sections 8.1–8.11 · Checked 2026-10-06
GMMLU
Multilingual · Vendor-reported
Test conditions and source
Section 8.18.1: 42 languages, adaptive thinking max effort, single trial, no tools or custom system prompt. Not the medium-effort default catalog version.
Tested model: anthropic/claude-fable-5-1
Anthropic Fable 5.1 system card · sections 8.1–8.11 · Checked 2026-10-06
MILU
Multilingual · Vendor-reported
Test conditions and source
Section 8.18.2: 11 languages, adaptive thinking max effort, single trial. Not the medium-effort default catalog version.
Tested model: anthropic/claude-fable-5-1
Anthropic Fable 5.1 system card · sections 8.1–8.11 · Checked 2026-10-06
HealthBench (raw)
Healthcare · Vendor-reported
Test conditions and source
Section 8.17: max effort, five trials, Opus 4.8 grader; safety classifiers and Opus 5 refusal fallback. Raw score, not length adjusted. Not the medium-effort default catalog version.
Tested model: anthropic/claude-fable-5-1
Anthropic Fable 5.1 system card · sections 8.1–8.11 · Checked 2026-10-06
HealthBench (length adjusted)
Healthcare · Vendor-reported
Test conditions and source
Section 8.17: same max-effort evaluation, length-adjusted score. Not the medium-effort default catalog version.
Tested model: anthropic/claude-fable-5-1
Anthropic Fable 5.1 system card · sections 8.1–8.11 · Checked 2026-10-06
HealthBench Professional (raw)
Healthcare · Vendor-reported
Test conditions and source
Section 8.17: five trials, max effort, Opus 4.8 grader, safety classifiers and Opus 5 fallback. Raw score. Not the medium-effort default catalog version.
Tested model: anthropic/claude-fable-5-1
Anthropic Fable 5.1 system card · sections 8.1–8.11 · Checked 2026-10-06
HealthBench Professional (length adjusted)
Healthcare · Vendor-reported
Test conditions and source
Section 8.17: same max-effort evaluation, length-adjusted score. Not the medium-effort default catalog version.
Tested model: anthropic/claude-fable-5-1
Anthropic Fable 5.1 system card · sections 8.1–8.11 · Checked 2026-10-06
BenchCAD Vision2Code (without tools)
Coding · Vendor-reported
Test conditions and source
Section 8.14.2: random 1,000-file subset, average five runs, adaptive thinking max effort, no tools. Not the medium-effort default catalog version.
Tested model: anthropic/claude-fable-5-1
Anthropic Fable 5.1 system card · sections 8.1–8.11 · Checked 2026-10-06
BenchCAD Vision2Code (with tools)
Coding · Vendor-reported
Test conditions and source
Section 8.14.2: random 1,000-file subset, average five runs, adaptive thinking max effort, container and image cropping tool. Not the medium-effort default catalog version.
Tested model: anthropic/claude-fable-5-1
Anthropic Fable 5.1 system card · sections 8.1–8.11 · Checked 2026-10-06
GDPval-AA v2
Knowledge work · Vendor-reported
Test conditions and source
Section 8.15.3: AA independent agentic professional work evaluation, max effort. Historical v2, not v2.1. Not the medium-effort default catalog version.
Tested model: anthropic/claude-fable-5-1
Anthropic Fable 5.1 system card · sections 8.1–8.11 · Checked 2026-10-06
AA-Briefcase (system card)
Knowledge work · Vendor-reported
Test conditions and source
Section 8.15.4: independent AA long-horizon work evaluation, max effort; historical system-card snapshot. Not the medium-effort default catalog version.
Tested model: anthropic/claude-fable-5-1
Anthropic Fable 5.1 system card · sections 8.1–8.11 · Checked 2026-10-06
FrontierCode v1.1 (Extended)
Coding · Third-party report
Test conditions and source
Cognition FrontierCode 1.1; Claude Fable 5.1; medium effort; claude-code harness; Extended subset (150 tasks). Published weighted rubric score (new_score × 100), not the all-or-nothing pass rate; solution-source violations score zero. Source field new_score=0.636. Publisher evaluation, not a TokenBazaar measurement; provider effort configuration is not a guarantee of the API's fixed thinking budget.
Tested model: anthropic/claude-fable-5-1
Cognition · FrontierCode 1.1 original leaderboard · Checked 2026-10-07
FrontierCode v1.1 (Main)
Coding · Third-party report
Test conditions and source
Cognition FrontierCode 1.1; Claude Fable 5.1; medium effort; claude-code harness; Main subset (100 tasks). Published weighted rubric score (new_score × 100), not the all-or-nothing pass rate; solution-source violations score zero. Source field new_score=0.5091. Publisher evaluation, not a TokenBazaar measurement; provider effort configuration is not a guarantee of the API's fixed thinking budget.
Tested model: anthropic/claude-fable-5-1
Cognition · FrontierCode 1.1 original leaderboard · Checked 2026-10-07
Terminal-Bench 4.0
Agentic · Vendor-reported
Test conditions and source
Section 8.6: Claude Code --bare, maximum thinking effort; 15 trials/task (990 trials), SE ±1.6–2 points. Published Fable 5.1 evaluation, not the default medium catalog configuration or a TokenBazaar run; benchmark tools are not hosted API capabilities.
Tested model: anthropic/claude-fable-5-1
Anthropic · Fable 5.1 system card · Checked 2026-10-07
Terminal-Bench-Science 0.1
Agentic · Vendor-reported
Test conditions and source
Section 8.7: Claude Code --bare, maximum thinking effort; 10 trials/task (700 trials), SE ±3.5–4.5 points. Published Fable 5.1 evaluation, not the default medium catalog configuration or a TokenBazaar run; benchmark tools are not hosted API capabilities.
Tested model: anthropic/claude-fable-5-1
Anthropic · Fable 5.1 system card · Checked 2026-10-07
Humanity’s Last Exam (with tools)
Reasoning · Vendor-reported
Test conditions and source
Table 8.1.A and section 8.12.1: web search, web fetch and code execution; thinking auto, 1M token cap, Opus 4.6 grader and source blocklist. Published Fable 5.1 evaluation, not the default medium catalog configuration or a TokenBazaar run; benchmark tools are not hosted API capabilities.
Tested model: anthropic/claude-fable-5-1
Anthropic · Fable 5.1 system card · Checked 2026-10-07
Humanity’s Last Exam (without tools)
Reasoning · Vendor-reported
Test conditions and source
Table 8.1.A: no tools; standard adaptive thinking at max effort, default sampling, five trials. Published Fable 5.1 evaluation, not the default medium catalog configuration or a TokenBazaar run; benchmark tools are not hosted API capabilities.
Tested model: anthropic/claude-fable-5-1
Anthropic · Fable 5.1 system card · Checked 2026-10-07
AutomationBench
Agentic · Vendor-reported
Test conditions and source
Table 8.1.A: adaptive thinking at max effort, default sampling, five trials; production safeguards enabled. Published Fable 5.1 evaluation, not the default medium catalog configuration or a TokenBazaar run; benchmark tools are not hosted API capabilities.
Tested model: anthropic/claude-fable-5-1
Anthropic · Fable 5.1 system card · Checked 2026-10-07
Chartography (with tools)
Multimodal · Vendor-reported
Test conditions and source
Section 8.14.1: adaptive thinking at max effort; container and image-cropping tool, five runs. Published Fable 5.1 evaluation, not the default medium catalog configuration or a TokenBazaar run; benchmark tools are not hosted API capabilities.
Tested model: anthropic/claude-fable-5-1
Anthropic · Fable 5.1 system card · Checked 2026-10-07
Chartography (without tools)
Multimodal · Vendor-reported
Test conditions and source
Section 8.14.1: adaptive thinking at max effort, no tools, five runs. Published Fable 5.1 evaluation, not the default medium catalog configuration or a TokenBazaar run; benchmark tools are not hosted API capabilities.
Tested model: anthropic/claude-fable-5-1
Anthropic · Fable 5.1 system card · Checked 2026-10-07
OSWorld 2.0 (partial)
Agentic · Vendor-reported
Test conditions and source
Section 8.14.3: August 2026 tasks with subsequent fixes; 1080p, maximum 500 action steps, maximum reasoning effort, five runs, Opus 4.8 grader. Not comparable to OSWorld 2.1 or earlier task releases. Published Fable 5.1 evaluation, not the default medium catalog configuration or a TokenBazaar run; benchmark tools are not hosted API capabilities.
Tested model: anthropic/claude-fable-5-1
Anthropic · Fable 5.1 system card · Checked 2026-10-07
OSWorld 2.0 (strict)
Agentic · Vendor-reported
Test conditions and source
Section 8.14.3: August 2026 tasks with subsequent fixes; 1080p, maximum 500 action steps, maximum reasoning effort, five runs, Opus 4.8 grader. Not comparable to OSWorld 2.1 or earlier task releases. Published Fable 5.1 evaluation, not the default medium catalog configuration or a TokenBazaar run; benchmark tools are not hosted API capabilities.
Tested model: anthropic/claude-fable-5-1
Anthropic · Fable 5.1 system card · Checked 2026-10-07
Model specifications
Context window
1M
Published provider limit
Maximum output
128K
Published provider limit
Knowledge cutoff
Jun 2026
Upstream model
Input and output modalities
Provider-reported formats; API compatibility is detailed below.
Input
Output
- Exact model ID
- anthropic/claude-fable-5-1
- Model type
- Chat
- Context window
- 1,000,000 tokens
- Maximum output
- 128,000 tokens
- Knowledge cutoff
- Jun 2026
A context window is the total conversation budget, not a separate maximum input allowance; generated output and reasoning can use that budget. Published provider limits are not independently tested TokenBazaar request limits.
Versions and thinking levels
Choose from 4 enabled versions. The exact ID determines the version and its price; benchmark scores do not carry over between variants.
Features and tool support
TokenBazaar’s exposed interface, not every feature advertised by the provider. “Available” describes the implemented interface, not a successful test of every input or tool.
Reasoning configuration
Adaptive thinking; reasoning and answer share the output budget.
Selected effort: medium (default catalog version). Reasoning level is not a measured intelligence or speed score.
Coding and agent workflows
Text/code generation and customer-managed tool loops.
Coding benchmark results do not establish a hosted terminal, sandbox, autonomous browser or guaranteed task success.
Streaming answers
AvailableIncremental text through TokenBazaar’s chat endpoint; Claude also has native Messages access.
Customer-defined function tools
LimitedFunction schemas and tool results are forwarded to Claude. The chat bridge uses automatic tool selection; your application authorizes and executes every tool.
Forced or named tool selection
Not exposedThe chat bridge selects tools automatically. Native Messages does not expose forced or named tool selection; do not assume a tool will be called every turn.
Images and PDFs
LimitedThe chat bridge accepts image/PDF content. Provider input modalities are listed separately; file size, content and model limits still apply.
Hosted web search / browsing
Not exposedNo provider-hosted web search or web-fetch tool is exposed. Customer-owned search can be implemented as an authorized function tool.
Hosted code execution / computer use
Not exposedNo built-in execution environment or computer-control tool is provided. Code generation is not code execution.
Hosted file search / persistent assistants
Not exposedNo hosted file-search index, Assistants endpoint or persistent provider agent is exposed.
Batch / fine-tuning / cached-price discounts
Not exposedNo public batch or fine-tuning endpoint, or separate cached-token discount, is offered here.
Sources and verification
Provider specifications describe the upstream model, not a guarantee of every feature through TokenBazaar. Prices come from the enabled catalog; benchmark results belong to the exact tested model.
- Provider model specifications
Vendor-reported
Last checked - Anthropic Fable 5.1 system card · sections 8.1–8.11
Vendor-reported
Last checked - CursorBench public evaluation
Third-party report
Last checked - Proximal · FrontierSWE public leaderboard
Third-party report
Last checked - Vals AI · Claude Fable 5.1 evaluation
Third-party report
Last checked - Google Gemini 4 Argon comparison chart
Vendor-reported
Last checked - Cognition · FrontierCode 1.1 original leaderboard
Third-party report
Last checked
Research notes (2)
- Published provider limits; maximum-capacity requests have not been independently exercised through TokenBaazar. Thinking-level variants use the same provider model; reasoning and answer tokens share the output budget.
- Benchmark configuration is stated per result. Provider max-effort tests are not measurements of the default medium or xhigh catalog versions. ProgramBench 87.6 is the vendor’s filtered 166-task test, not a generic public score.