Compare models
Choose two or three models, including exact thinking and fast versions. Copy the URL to share your selection.
| Comparison | Jev Latesttypesafe/jev-latest | Choose model 2 |
|---|---|---|
| Type | Decision | — |
| TokenBaazar price · 50% off | ₹2.02 per 1M tokens | — |
| Market price | ₹4.03 per 1M tokens | — |
| Input token limit (provider) | Within the context window | — |
| Vendor input modalities | State and typed questions | — |
| Vendor output modalities | Probabilistic decisions | — |
| TokenBaazar endpoints | POST /v1/decisions | — |
| Benchmark evidence | 3 source-linked results | — |
| Typed decisions — LocalLLaMA/typed-decisions frozen test split | 0.74accuracy (0-1) Independent third-party eval (not vendor-reported), frozen split, 2,000-sample bootstrap CI; zero-shot classification/scoring accuracy. Published Jev 1.13.0 checkpoint; the rolling latest API alias may change and is not independently pinned by TokenBazaar. Published source text: Independent benchmark ran Jev 1.13.0 on a frozen 400-case / 2,000-decision test split of LocalLLaMA/typed-decisions on 2026-09-20, reporting Accuracy 0.740 (95% CI 0.721–0.759). DecisionEval — Jev by TypeSafe benchmark page | — |
| Typed decisions — calibration | 0.148Brier score Lower is better; independent third-party eval, frozen split dated 2026-09-20. Published Jev 1.13.0 checkpoint; the rolling latest API alias may change and is not independently pinned by TokenBazaar. Published source text: Same independent run reports Brier score (against gold probability distribution) of 0.148, 95% CI [0.139, 0.156]. DecisionEval — Jev by TypeSafe benchmark page | — |
| Belebele (122 languages) | 86.7% Zero-shot, jev-1.13.0, one frozen template per dataset, full evaluation split across 122 languages of Belebele. Published Jev 1.13.0 checkpoint; the rolling latest API alias may change and is not independently pinned by TokenBazaar. Published source text: Paper abstract: 'Jev reaches 95-99% accuracy on IMDB, SST-2, HellaSwag and ARC and 86.7% on Belebele across 122 languages.' arXiv — Evaluating and Benchmarking the System One Model Jev | — |
No overall score, winner or cost-versus-score frontier is calculated without comparable, attributable benchmark evidence. Different test configurations appear separately. Vendor capabilities are distinct from TokenBaazar API support.
— means no published value or no matching capability. Provider token limits are not tested gateway guarantees; TokenBaazar media controls and response formats are listed separately.
Jev Latest
Sources and verification
Provider specifications describe the upstream model, not a guarantee of every feature through TokenBazaar. Prices come from the enabled catalog; benchmark results belong to the exact tested model.
- Vendor model documentation
Vendor-reported
Last checked - DecisionEval — Jev by TypeSafe benchmark page
Third-party report
Last checked - arXiv — Evaluating and Benchmarking the System One Model Jev
Third-party report
Last checked
Research notes (1)
- This is a decision model, not a general chat model. The moving latest ID is not a pinned benchmark version.