typesafe / Decision
Jev Latest
typesafe/jev-latestTyped decisions with probabilities — classification, scoring and routing
Catalog status
Enabled
Research checked
2026-10-06
Benchmark results
Results are source-attributed, not TokenBazaar measurements. Different tests and thinking variants are not interchangeable.
3 of 3 results
Typed decisions — LocalLLaMA/typed-decisions frozen test split
Accuracy · Third-party report
Test conditions and source
Independent third-party eval (not vendor-reported), frozen split, 2,000-sample bootstrap CI; zero-shot classification/scoring accuracy. Published Jev 1.13.0 checkpoint; the rolling latest API alias may change and is not independently pinned by TokenBazaar. Published source text: Independent benchmark ran Jev 1.13.0 on a frozen 400-case / 2,000-decision test split of LocalLLaMA/typed-decisions on 2026-09-20, reporting Accuracy 0.740 (95% CI 0.721–0.759).
Tested model: typesafe/jev-latest
DecisionEval — Jev by TypeSafe benchmark page · Checked 2026-10-06
Typed decisions — calibration
Accuracy · Third-party report
Test conditions and source
Lower is better; independent third-party eval, frozen split dated 2026-09-20. Published Jev 1.13.0 checkpoint; the rolling latest API alias may change and is not independently pinned by TokenBazaar. Published source text: Same independent run reports Brier score (against gold probability distribution) of 0.148, 95% CI [0.139, 0.156].
Tested model: typesafe/jev-latest
DecisionEval — Jev by TypeSafe benchmark page · Checked 2026-10-06
Belebele (122 languages)
Accuracy · Third-party report
Test conditions and source
Zero-shot, jev-1.13.0, one frozen template per dataset, full evaluation split across 122 languages of Belebele. Published Jev 1.13.0 checkpoint; the rolling latest API alias may change and is not independently pinned by TokenBazaar. Published source text: Paper abstract: 'Jev reaches 95-99% accuracy on IMDB, SST-2, HellaSwag and ARC and 86.7% on Belebele across 122 languages.'
Tested model: typesafe/jev-latest
arXiv — Evaluating and Benchmarking the System One Model Jev · Checked 2026-10-06
Model specifications
Input and output modalities
Provider-reported formats; API compatibility is detailed below.
Input
Output
- Exact model ID
- typesafe/jev-latest
- Model type
- Decision
Exact token limits are not published in the reviewed sources, so no estimated limits are shown.
Versions and thinking levels
One enabled version is currently listed for this model.
Features and tool support
TokenBazaar’s exposed interface, not every feature advertised by the provider. “Available” describes the implemented interface, not a successful test of every input or tool.
Input through TokenBazaar
Not verifiedConsult the integration documentation; no additional input formats are claimed here.
Output through TokenBazaar
Not verifiedOutput compatibility has not been independently verified for this exact model.
Function-tool conversations
Not exposedThis is a dedicated media or utility endpoint, not a conversational function-tool interface.
Sources and verification
Provider specifications describe the upstream model, not a guarantee of every feature through TokenBazaar. Prices come from the enabled catalog; benchmark results belong to the exact tested model.
- Vendor model documentation
Vendor-reported
Last checked - DecisionEval — Jev by TypeSafe benchmark page
Third-party report
Last checked - arXiv — Evaluating and Benchmarking the System One Model Jev
Third-party report
Last checked
Research notes (1)
- This is a decision model, not a general chat model. The moving latest ID is not a pinned benchmark version.