google / Chat
Gemini 3.8 Flash
google/gemini-3.8-flashLatest-generation Gemini Flash model for fast coding, reasoning, and agentic workflows.
Catalog status
Enabled
Research checked
2026-10-07
Price / performance
What it costs to get this result
FrontierCode v1.1 (Extended) · Source-reported results, not a composite ranking. Prices use the average of TokenBazaar input and output rates per million tokens; actual spend depends on usage. Test configurations differ.
Blended price / 1M tokens · log scale
Compare sourced models on this benchmark
Benchmark results
Results are source-attributed, not TokenBazaar measurements. Different tests and thinking variants are not interchangeable.
6 of 6 results
GPQA Diamond
Reasoning · Third-party report
Test conditions and source
Vals AI independent evaluation, default model configuration, dated post-Sept 2, 2026 GA launch. Published source text: GPQA Diamond Vals AI | 94.44% | 4 of 23 | 4 of 138 | Leader Gemini 3.1 Pro | 1.01 pts behind
Tested model: google/gemini-3.8-flash
AIEvals – Gemini 3.8 Flash benchmark results (Vals AI methodology) · Checked 2026-10-06
SWE-bench Verified
Coding · Third-party report
Test conditions and source
Vals AI independent evaluation, default configuration. Published source text: SWE-bench (Verified) Vals AI | 80.00% | 18 of 23 | 24 of 88
Tested model: google/gemini-3.8-flash
AIEvals – Gemini 3.8 Flash benchmark results (Vals AI methodology) · Checked 2026-10-06
Terminal-Bench 2.1
Agentic · Third-party report
Test conditions and source
Terminus 2 default agent harness, high effort, Sept 2026 GA launch per Google model card / Google evals-methodology page. Published source text: Terminal-Bench 2.1 89.4% (Terminus 2 harness)
Tested model: google/gemini-3.8-flash
Wait Which Model – Gemini 3.8 Flash · Checked 2026-10-06
Humanity's Last Exam
Reasoning · Third-party report
Test conditions and source
High effort/thinking setting, Sept 2026 GA launch. Published source text: Humanity's Last Exam 54.9%
Tested model: google/gemini-3.8-flash
Wait Which Model – Gemini 3.8 Flash · Checked 2026-10-06
FrontierCode v1.1 (Extended)
Coding · Third-party report
Test conditions and source
Cognition FrontierCode 1.1; Gemini 3.8 Flash; medium effort; chisel harness; Extended subset (150 tasks). Published weighted rubric score (new_score × 100), not the all-or-nothing pass rate; solution-source violations score zero. Source field new_score=0.5345. Publisher evaluation, not a TokenBazaar measurement; provider effort configuration is not a guarantee of the API's fixed thinking budget.
Tested model: google/gemini-3.8-flash
Cognition · FrontierCode 1.1 original leaderboard · Checked 2026-10-07
FrontierCode v1.1 (Main)
Coding · Third-party report
Test conditions and source
Cognition FrontierCode 1.1; Gemini 3.8 Flash; medium effort; chisel harness; Main subset (100 tasks). Published weighted rubric score (new_score × 100), not the all-or-nothing pass rate; solution-source violations score zero. Source field new_score=0.4119. Publisher evaluation, not a TokenBazaar measurement; provider effort configuration is not a guarantee of the API's fixed thinking budget.
Tested model: google/gemini-3.8-flash
Cognition · FrontierCode 1.1 original leaderboard · Checked 2026-10-07
Model specifications
Input token limit
1.05M
Published provider limit
Maximum output
65.54K
Published provider limit
Input and output modalities
Provider-reported formats; API compatibility is detailed below.
Input
Output
- Exact model ID
- google/gemini-3.8-flash
- Model type
- Chat
- Input token limit
- 1,048,576 tokens
- Maximum output
- 65,536 tokens
Input token limits describe provider input capacity. Published provider limits are not independently tested TokenBazaar request limits.
Versions and thinking levels
One enabled version is currently listed for this model.
Features and tool support
TokenBazaar’s exposed interface, not every feature advertised by the provider. “Available” describes the implemented interface, not a successful test of every input or tool.
Reasoning configuration
No independent reasoning-mode guarantee is recorded for this exact model.
See enabled versions and provider sources for supported controls.
Coding and agent workflows
Text/code generation and customer-managed tool loops.
Coding benchmark results do not establish a hosted terminal, sandbox, autonomous browser or guaranteed task success.
Streaming answers
AvailableIncremental text through TokenBazaar’s chat endpoint; Claude also has native Messages access.
Customer-defined function tools
LimitedOnly customer-defined function tools are forwarded. Model/protocol compatibility applies; your application authorizes and executes every tool.
Forced or named tool selection
LimitedThe chat route accepts auto, none and required. A named-tool choice is not forwarded; exact model compatibility is not independently tested.
Images and PDFs
Not verifiedThe chat bridge accepts image/PDF content. Provider input modalities are listed separately; file size, content and model limits still apply.
Hosted web search / browsing
Not exposedNo provider-hosted web search or web-fetch tool is exposed. Customer-owned search can be implemented as an authorized function tool.
Hosted code execution / computer use
Not exposedNo built-in execution environment or computer-control tool is provided. Code generation is not code execution.
Hosted file search / persistent assistants
Not exposedNo hosted file-search index, Assistants endpoint or persistent provider agent is exposed.
Batch / fine-tuning / cached-price discounts
Not exposedNo public batch or fine-tuning endpoint, or separate cached-token discount, is offered here.
Sources and verification
Provider specifications describe the upstream model, not a guarantee of every feature through TokenBazaar. Prices come from the enabled catalog; benchmark results belong to the exact tested model.
- Provider model specifications
Vendor-reported
Last checked - AIEvals – Gemini 3.8 Flash benchmark results (Vals AI methodology)
Third-party report
Last checked - Wait Which Model – Gemini 3.8 Flash
Third-party report
Last checked - Cognition · FrontierCode 1.1 original leaderboard
Third-party report
Last checked
Research notes (1)
- Published provider limits; maximum-capacity requests have not been independently exercised through TokenBaazar. Thinking-level variants use the same provider model; reasoning and answer tokens share the output budget.