Compare model

typesafe / Decision

Jev Latest

typesafe/jev-latest

Typed decisions with probabilities — classification, scoring and routing

Catalog status

Enabled

Research checked

2026-10-06

Benchmark results

3 reported results

Results are source-attributed, not TokenBazaar measurements. Different tests and thinking variants are not interchangeable.

3 of 3 results

Typed decisions — LocalLLaMA/typed-decisions frozen test split

Accuracy · Third-party report

Test conditions and source

Independent third-party eval (not vendor-reported), frozen split, 2,000-sample bootstrap CI; zero-shot classification/scoring accuracy. Published Jev 1.13.0 checkpoint; the rolling latest API alias may change and is not independently pinned by TokenBazaar. Published source text: Independent benchmark ran Jev 1.13.0 on a frozen 400-case / 2,000-decision test split of LocalLLaMA/typed-decisions on 2026-09-20, reporting Accuracy 0.740 (95% CI 0.721–0.759).

Tested model: typesafe/jev-latest

DecisionEval — Jev by TypeSafe benchmark page · Checked 2026-10-06

0.74accuracy (0-1)

Typed decisions — calibration

Accuracy · Third-party report

Test conditions and source

Lower is better; independent third-party eval, frozen split dated 2026-09-20. Published Jev 1.13.0 checkpoint; the rolling latest API alias may change and is not independently pinned by TokenBazaar. Published source text: Same independent run reports Brier score (against gold probability distribution) of 0.148, 95% CI [0.139, 0.156].

Tested model: typesafe/jev-latest

DecisionEval — Jev by TypeSafe benchmark page · Checked 2026-10-06

0.148Brier score

Belebele (122 languages)

Accuracy · Third-party report

Test conditions and source

Zero-shot, jev-1.13.0, one frozen template per dataset, full evaluation split across 122 languages of Belebele. Published Jev 1.13.0 checkpoint; the rolling latest API alias may change and is not independently pinned by TokenBazaar. Published source text: Paper abstract: 'Jev reaches 95-99% accuracy on IMDB, SST-2, HellaSwag and ARC and 86.7% on Belebele across 122 languages.'

Tested model: typesafe/jev-latest

arXiv — Evaluating and Benchmarking the System One Model Jev · Checked 2026-10-06

86.7%

Model specifications

Input and output modalities

Provider-reported formats; API compatibility is detailed below.

Input

State and typed questions

Output

Probabilistic decisions
Exact model ID
typesafe/jev-latest
Model type
Decision

Exact token limits are not published in the reviewed sources, so no estimated limits are shown.

Versions and thinking levels

One enabled version is currently listed for this model.

Jev LatestSelected₹2.02 per 1M tokens

Features and tool support

TokenBazaar’s exposed interface, not every feature advertised by the provider. “Available” describes the implemented interface, not a successful test of every input or tool.

Input through TokenBazaar

Not verified

Consult the integration documentation; no additional input formats are claimed here.

Output through TokenBazaar

Not verified

Output compatibility has not been independently verified for this exact model.

Function-tool conversations

Not exposed

This is a dedicated media or utility endpoint, not a conversational function-tool interface.

Sources and verification

Provider specifications describe the upstream model, not a guarantee of every feature through TokenBazaar. Prices come from the enabled catalog; benchmark results belong to the exact tested model.

Research notes (1)
  • This is a decision model, not a general chat model. The moving latest ID is not a pinned benchmark version.