WEDNESDAY, SEPTEMBER 23, 2026|No. 16065
Technology · AI

New Benchmark 'JevBench' Evaluates AI Models on Intelligence, Calibration, Speed, and Cost

A new benchmark called JevBench has been released, offering a reproducible way to evaluate typed decision models across intelligence, calibration, speed, and cost.

A visual representation of artificial intelligence data processing and analysis.
A visual representation of artificial intelligence data processing and analysis. · Photo by Growtika on Unsplash
1 sources
Pipeline ingest
3 reads
Positive / Neutral / Negative
0 countries
Related coverage

JevBench Score (Intelligence, Calibration, Speed, Cost — 25 % each)

OfficialIntelligence above chance, Calibration, Speed, Cost — 25 % each, geometric mean; below 50 Intelligence receives a growing near-chance penalty. Change the weighting ↓

Intel.Calib.SpeedCost$/1k dec. 01. 1Jev 1.13.0 74.4 I 86 · C 83 · S 83 · K 52 · $0.040 02. 2SemIf (Qwen3.5-4B) 73.1 I 79 · C 73 · S 84 · K 59 · ~$0.022 est. 03. 3djev (Maisa, diffusion-gemma)73.0 I 83 · C 65 · S 91 · K 58 · $0.026 ann. 04. 4Winnow-12B Q871.2 I 82 · C 72 · S 82 · K 53 · ~$0.037 est. 05. 5reflex 4B70.3 I 80 · C 75 · S 68 · K 60 · ~$0.022 est. 06. 6jqv68.6 I 79 · C 79 · S 75 · K 47 · ~$0.056 est. 07. 7decision-machine-168.3 I 62 · C 70 · S 93 · K 54 · $0.035 08. 8decider-35b-a3b67.6 I 80 · C 72 · S 81 · K 45 · ~$0.067 est. 09. 9open-alternative-jev (Qwen3.5-4B)67.0 I 64 · C 63 · S 83 · K 60 · ~$0.022 est. 10. 10system-one-open 66.6 I 70 · C 57 · S 77 · K 65 · ~$0.015 est. 11. 11OpenJev (razorback16) 66.4 I 79 · C 65 · S 83 · K 45 · ~$0.066 est. 12. 12SimpleJev Qwen3.8-27B66.3 I 85 · C 81 · S 71 · K 39 · ~$0.104 est. 13. 13ZeroEntropy zerank-266.0 I 63 · C 76 · S 79 · K 50 · $0.047 14. 14GPT-5.6 Luna (low) 65.9 I 95 · C 90 · S 78 · K 28 · $0.242 15. 15openjev-sglang 65.3 I 83 · C 77 · S 77 · K 36 · ~$0.131 est. 16. 16Qwen3-Reranker-4B63.8 I 64 · C 67 · S 79 · K 49 · $0.050 17. 17reflex-27b63.3 I 86 · C 86 · S 67 · K 32 · ~$0.181 est. 18. 18LitJev62.7 I 82 · C 84 · S 67 · K 34 · ~$0.163 est. 19. 19kev 0.6B62.5 I 52 · C 51 · S 76 · K 76 · ~$0.0063 est. 20. 20SimpleJev Qwen3.6-35B-A3B62.5 I 80 · C 67 · S 75 · K 38 · ~$0.116 est. 21. 21djev62.4 I 81 · C 93 · S 75 · K 27 · ~$0.274 est. 22. 22jev-local61.8 I 71 · C 69 · S 69 · K 43 · ~$0.077 est. 23. 23decider-2b61.7 I 61 · C 47 · S 83 · K 61 · ~$0.020 est. 24. 24Bespoke Nimble 9B60.5 I 78 · C 65 · S 79 · K 33 · ~$0.166 est. 25. 25Gemini 3.1 Flash-Lite 60.1 I 86 · C 68 · S 82 · K 27 · $0.264 26. 26OpenJev60.0 I 88 · C 70 · S 76 · K 28 · ~$0.255 est. 27. 27kev 4B59.7 I 65 · C 42 · S 76 · K 62 · ~$0.019 est. 28. 28DeepSeek V4.1 Flash 57.5 I 94 · C 97 · S 72 · K 17 · $0.594 29. 29kev 8B56.4 I 69 · C 44 · S 75 · K 44 · ~$0.073 est. 30. 30Open-Jev 9B55.0 I 71 · C 63 · S 72 · K 28 · ~$0.249 est. 31. 31system-one 54.8 I 70 · C 37 · S 84 · K 41 · ~$0.089 est. 32. 32jeff54.4 I 47 · C 65 · S 63 · K 77 · ~$0.0060 est. 33. 33Laya54.4 I 46 · C 62 · S 71 · K 86 · ~$0.0029 est. 34. 34Open-Jev 2B51.3 I 61 · C 55 · S 73 · K 28 · ~$0.249 est. 35. 35OpenDecision40.6 I 41 · C 56 · S 80 · K 75 · ~$0.0066 est. 36. 36openJev Verdict 1.438.9 I 39 · C 74 · S 78 · K 82 · ~$0.0039 est. 37. 37openJev Verdict38.1 I 40 · C 51 · S 77 · K 83 · ~$0.0037 est. 38. 38kev 0.5B33.2 I 38 · C 47 · S 77 · K 76 · ~$0.0063 est. 39. 39GLiNER2 large29.6 I 40 · C 24 · S 62 · K 73 · ~$0.0077 est. 40. 40smalljev semantic-v927.4 I 35 · C 59 · S 80 · K 58 · ~$0.025 est. 41. 41GLiNER224.0 I 36 · C 24 · S 72 · K 83 · ~$0.0037 est. 42. 42open-jev-deberta-v3-large 23.1 I 32 · C 66 · S 66 · K 74 · ~$0.0073 est. 43. 43GLiNER2.5 multi16.6 I 28 · C 56 · S 68 · K 82 · ~$0.0039 est. 44. 44GLiNER2.5 small13.8 I 26 · C 47 · S 78 · K 82 · ~$0.0039 est. 45. 45Mixedbread mxbai-rerank-base-v20.8 I 7 · C 83 · S 88 · K 68 · $0.012 46. 46BAAI bge-reranker-v2-m30.7 I 6 · C 84 · S 90 · K 73 · $0.0077 47. 47Alibaba GTE Reranker ModernBERT-base0.3 I 5 · C 77 · S 91 · K 70 · $0.010 48. 48Certo v10.0 I 0 · C 82 · S 94 · K 100 · ~$0.0010 est. 49. classifier.dev (fast tier)† (honorable mention) 83.6 I 85 · C 78 · S 88 · K 84 · ~$0.0033 est. 50. Qwen3.8 27B (partial run) 24.8 I 67 · C 92 · S 61 · K 0 · ~$2.669 est. 51. Needle 3, options as tools (partial run) 1.1 I 14 · C – · S 53 · K 65 · ~$0.014 est. 52. Needle 3 (partial run) 0.1 I 5 · C – · S 60 · K 59 · ~$0.024 est.

020406080100

Score = Intelligence0.25 × Calibration0.25 × Speed0.25 × Cost0.25(each 0–100; geometric mean; below 50 Intelligence, × (I / 50)²)

  • Jev (TypeSafe, closed)
  • Jev rebuild (open, or open source planned)
  • Instruction model, JSON schema
  • Small tool-calling model
  • Service built on Jev
  • Zero-shot classifier (not a Jev rebuild)
  • Closed decision model (API only, not Jev)
  • Shown, not ranked — honorable mention (runs another entrant's model) · partial run

⏱ Latency of self-hosted and demo endpoints is adjusted ×2 (+0.15 s on our own servers) to approximate production load — an assumption, not a measurement; raw measurements are in the table and the repo.I, C, S, K = Intelligence, Calibration, Speed, Cost; est./ann. = estimated/announced cost; † = see note.Legend and notes

  • ~ est. = no measured bill; priced like a large inference provider ( how costs are estimated).
  • ann. = the provider’s announced price, not yet charged.
  • Names link to each project.
  • A label-only system has no calibration (–, counted as 0).
  • djev (Maisa, diffusion-gemma): The measured endpoint was Maisa's hosted API in free preview; the cost uses its announced price ($0.035 per million input tokens, output free), and nothing was charged. The self-hostable djev-dev runtime is Apache-2.0 and applies a structured one-step inference method to Google's Apache-2.0 diffusiongemma-26B-A4B-it checkpoint; it adds no separately trained djev weights. Probabilities are djev's own (its docs call them experimental and uncalibrated).
  • Winnow-12B Q8: The submitted Q8_0 GGUF ran through the pinned author's TypeSafe-compatible /v1/systemone server with 8,192 context, four resident decision branches, Q8 KV, and full GPU offload. The private training corpus was not released. The author's checksum-based audit reports zero exact public-item overlap, but that claim cannot be independently reproduced; our scan found no exact public state or instruction text in the released artifacts. Cost uses the $0.05/M-input hosted Gemma 3 12B reference, not free/100.
  • reflex 4B: The author's reflex-serve: Qwen3.5-4B with the published LoRA and its per-primitive calibration file; the state is encoded once and each question read from the label logits. Run serially on our GPU; the author discloses that the 231 public items were used four times as a development gate.
  • jqv: A stock Qwen3-32B with no decision training: the state is prefilled once, each question is an isolated branch and the answer is read from the option-letter logits, with one fitted temperature (3.02, 400 MMLU validation items). Re-run in v1.2.8 on our own GPU from the now-public serving code (Octalab-Inc/jqv 0189b67), so all 534 decisions including the held-out hard items were asked; this full run replaces the v1.2.7 partial row, which had been measured on the submitter's machine. Cost is the base model's public per-token tariff, not free.
  • decision-machine-1: A closed-weights decision model behind a production API that serves TypeSafe's wire format, so the unchanged typesafe adapter ran it. Run on a free test key (30 requests a minute, 2.2 s between requests); the provider states the inference infrastructure is the same as for paid keys. Cost is the public paid tariff, $0.04 per million input tokens (output free), times the input tokens the API reported.
  • decider-35b-a3b: The author's TypeSafe-compatible server and published FP8 weights, run serially on our H100 NVL. The exhaustive startup batch warmup was skipped; each required serial shape captured lazily before its measured request. Self-host latency receives the standard ×2 + 0.15 s adjustment. Cost uses the closest hosted 35B-A3B input tariff and is not the temporary rental charge.
  • open-alternative-jev: With the options in reverse order (A. no, B. yes) the same model scored 21 % instead of 72 % on yes/no answer-judging items — small models are very sensitive to option order.
  • SimpleJev Qwen3.8-27B: Author's no-login shared demo, model id recorded verbatim, one request at a time at or below its 2 RPS limit. SimpleJev reads answer-token logits and returns the complete distribution; it does not generate an answer. Speed uses the public-demo x2 load adjustment; cost uses a hosted size-class input price and is not free/100.
  • ZeroEntropy zerank-2: Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.
  • Qwen3-Reranker-4B: Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.
  • reflex-27b: The frozen public Qwen3.8-27B checkpoint through reflex at the requested pinned commit, with two option orders averaged and temperature 1. No adapter or fitted calibration file. Run serially on our H100 NVL. Self-host latency receives the standard ×2 + 0.15 s adjustment; cost uses the exact base model's public hosted input tariff.
  • LitJev: The author's reproduction of Jev's decision layer on an off-the-shelf model, in its default configuration: Qwen3.8-27B, scores read from the output head, no training and no calibration file (its README says probabilities are not calibrated by default). Run serially on our GPU through an SSH tunnel, because its server binds to localhost; the request still crosses the internet and gets the ×2 + 0.15 s adjustment.
  • kev 0.6B: Self-hosted from the author's repository at commit 20fa626 through its native TypeSafe-compatible /v1/systemone server, BF16 on an RTX 3090; measured serially from Sandy over the internet. The author labels this checkpoint a research preview.
  • SimpleJev Qwen3.6-35B-A3B: Author's no-login shared demo, model id recorded verbatim, one request at a time at or below its 2 RPS limit. SimpleJev reads answer-token logits and returns the complete distribution; it does not generate an answer. Speed uses the public-demo x2 load adjustment; cost uses a hosted size-class input price and is not free/100.
  • djev: Experimental full-generation path over the same DiffusionGemma checkpoint as djev-dev: thinking was enabled and the model could generate up to 8,192 tokens before returning its distribution. Current djev-dev itself hard-codes enable_thinking=false, diffusion_max_steps=1 and read_only=true, so this is not a switch in its published typed API. It is substantially slower/costlier, and 72/534 requests exhausted the output budget without a parseable distribution; those are failures. Cost uses measured tokens and a same-size hosted reference, not the H200 rental bill.
  • jev-local: The author's local Jev-compatible server in its default full configuration: a frozen Qwen3.5-9B scores each option by its mean log-probability (one forward pass per option, no generation, no decision training). Run serially on our GPU. It re-reads the state once per option; if its reported token count covers one pass only, a per-token hosted price would be higher than this estimate.
  • decider-2b: The author's TypeSafe-compatible server and published weights (Qwen3.5-2B-Base with a trained one-pass decision readout), run serially on our GPU. Self-host latency gets the standard ×2 + 0.15 s adjustment.
  • Bespoke Nimble 9B: Re-run in v1.2.8 at Bespoke Labs' request after they raised the serving prompt limit from 2,048 to 8,192 tokens (bespokelabsai/nimble PR #4). Same recipe as the v1.1.3 run — the published LoRA merged into Qwen3.5-9B with the author's PEFT safe-merge, served with SGLang and the author's Jev-compatible API — now from current nimble main; the adapter weights are unchanged. Hard-tier accuracy rose from 43.6 % to 65.5 %, yet the score fell: the long hard items that used to fail at once are now answered and priced (so Cost fell), and this pod was in Canada while the v1.1.3 run's was in Sweden, so part of the lower Speed is network distance from our server in Germany. This complete run replaces the earlier row; its old score is kept in the artifact under superseded_rows.
  • OpenJev: OpenJev's real typed-API thinking switch at think=512, using its own /v1/systemone server over BF16 DiffusionGemma. The thought is generated first, then native probability reads are taken after it. All 534 requests returned valid distributions. Cost counts the server's billed input and thought output tokens.
  • kev 4B: Self-hosted from the author's repository at commit 20fa626 through its native TypeSafe-compatible /v1/systemone server, BF16 on an RTX 3090; measured serially from Sandy over the internet. The author labels this checkpoint a research preview.
  • kev 8B: Self-hosted from the author's repository at commit 20fa626 through its native TypeSafe-compatible /v1/systemone server, BF16 on an RTX 3090; measured serially from Sandy over the internet. The author labels this checkpoint a research preview.
  • Open-Jev 9B: The author's pinned LoRA adapter, trained scalar decision head and calibration temperature, served by the author's Open-Jev server with prefix caching off, batch size 1 and 4,096-token limit. Serial requests were measured from Sandy over an SSH tunnel to the H100. Self-host latency receives the standing x2 + 0.15 s adjustment. Cost uses the exact Qwen3.5-9B hosted input tariff for 9B and the same conservative same-family proxy for the unlisted 2B; neither receives an automatic 100. Exact normalized comparison found no JevBench public task state or instruction in the 79,116-row public training projection.
  • jeff: Self-hosted from its GitHub repo with server defaults, on our CPU (the author recommends a GPU, e.g. an L4), through the same TypeSafe-compatible API as Jev.
  • Laya: The English checkpoint (repo root), run on our CPU through its own laya package. Its budget is 512 tokens per question, so long hard-tier states are cut by the package itself.
  • Open-Jev 2B: The author's pinned LoRA adapter, trained scalar decision head and calibration temperature, served by the author's Open-Jev server with prefix caching off, batch size 1 and 4,096-token limit. Serial requests were measured from Sandy over an SSH tunnel to the H100. Self-host latency receives the standing x2 + 0.15 s adjustment. Cost uses the exact Qwen3.5-9B hosted input tariff for 9B and the same conservative same-family proxy for the unlisted 2B; neither receives an automatic 100. Exact normalized comparison found no JevBench public task state or instruction in the 79,116-row public training projection.
  • OpenDecision: A zero-shot NLI classifier behind a TypeSafe-compatible server, not a trained decision model: it scores each option as an entailment hypothesis with ModernBERT-large-zeroshot-v2.0. Its choice path runs several NLI passes over the same state, which the reported token count does not include, so a per-token hosted price would be higher than the estimate here. Pre-registered for our CPU in v1.2.7, run on our GPU because the CPU was far too slow.
  • openJev Verdict 1.4: Same public weights as the earlier Verdict row, run through the author's fixed v1.4 engine. That engine auto-loads the calibrator for every option count, frames candidate labels as NLI sentences and uses a 512-token context budget. Run locally on our CPU, serially.
  • openJev Verdict: The openJev-verdict-2.0 Hugging Face repo ships no weights; its config is byte-identical to heman10x/rlcd-modernbert-151m, whose published weights we ran with the author's engine. The 'verdict2-base' checkpoint behind the README's numbers is not downloadable yet (Git LFS 404); we will run it once it is.
  • kev 0.5B: Self-hosted from the author's repository at commit 20fa626 through its native TypeSafe-compatible /v1/systemone server, BF16 on an RTX 3090; measured serially from Sandy over the internet. This is the v0.1 release.
  • GLiNER2 large: The large checkpoint of Fastino's earlier GLiNER2 family, same documented mapping as the GLiNER2 row: the question goes in front of the text and the probabilities are the model's own single-label softmax over the labels, read out in full. A general schema classifier, not a Jev rebuild.
  • smalljev semantic-v9: The public semantic-v9 LoRA and native heads over MiniCPM5-2B-Base, through the mapping frozen before the run. It has a typed Python contract but no TypeSafe-compatible HTTP route. The released training recipe explicitly hill-climbed against JevBench's public shape and source families; this allowed public benchmark-directed development is disclosed. Cost is $0.04/M measured input tokens, not free/100.
  • GLiNER2: A general schema classifier, not a Jev rebuild. The question goes in front of the text; the probabilities are GLiNER2's own single-label softmax over the labels, read out in full (mapping fixed before the run).
  • GLiNER2.5 multi: The multilingual GLiNER2.5 checkpoint (287M), same family and same documented mapping as the GLiNER2 row. JevBench items are English only, so its multilingual training is not exercised here.
  • GLiNER2.5 small: The small GLiNER2.5 checkpoint (74M), same family and same documented mapping as the GLiNER2 row: the question goes in front of the text and the probabilities are the model's own single-label softmax over the labels, read out in full. A general schema classifier, not a Jev rebuild.
  • Mixedbread mxbai-rerank-base-v2: Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.
  • BAAI bge-reranker-v2-m3: Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.
  • Alibaba GTE Reranker ModernBERT-base: Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.
  • Certo v1: The public Certo v1 checkpoint through the author's DecisionModel, serially on our rented GPU. The question instruction is prepended to the state because Certo exposes state + runtime options but no separate question field; the published 64-token state and 48-token option limits are unchanged. The model card says v1 does not yet transfer to arbitrary natural-language prose. Cost is an estimate from same-size hosted encoders times the checkpoint's retained input tokens, not free/100.
  • classifier.dev (fast tier): Its own benchmark page says the fast tier is Jev. Free for us; the price is its published Pro plan ($20/month for 200,000 fast classifications a day) at full use, $0.0033 per 1,000 decisions.

Weighting: Intelligence : Calibration : Speed : Cost

Official default

JevBench Score25:25:25:25 · defaultBalanced, no calibration33:0:33:33Emphasis on Accuracy60:0:20:20Emphasis on Speed20:0:60:20Emphasis on Cost20:0:20:60Custom

Intelligence25 %Calibration25 %Speed25 %Cost25 %

The official JevBench Score weights the four axes 25 % each and takes their geometric mean. The other buttons are the earlier views (Balanced 33:33:33 and the three “Emphasis on” weightings, which leave Calibration out), recomputed the same way. Any of them is your view, recomputed in your browser from the published axis scores — not the published score.

Explore by task difficulty

All tasks is the published default. Choose a scope to see how the ranking changes by difficulty. Hard only uses all 220 hard-tier decisions and their measured Intelligence, Calibration, Speed and Cost.

All tasksEasy + MediumEasy onlyHard only

Tier mapping: Easy = easy; Medium = standard. Easy scopes change Intelligence only. Hard only measures all four axes on the same hard-tier subset; systems without a hard-tier run are shown as partial and are not ranked.

What the run says (JevBench Score)

  • Jev 1.13.0 (TypeSafe AI) leads with 74.4: Intelligence 85.7, Calibration 82.7, Speed 83.3, Cost 52.0 ($0.040 per 1,000 decisions).
  • classifier.dev scores 83.6 — higher than anything in the ranking — but is not ranked: it runs Jev (TypeSafe), so ranking it would put the same model in the list twice, once at the model's own price and once at the service's. It keeps every number it earned under Honorable mentions — services built on another entrant's model.
  • Open rebuilds of Jev appeared within days. The best of them, SemIf, formerly OpenJev (Qwen3.5-4B, TheoLeeCJ), is #2 at 73.1 — 1.3 points behind: more speed and a lower (estimated) price, less intelligence and calibration.
  • GPT-5.6 Luna (low reasoning effort) has the highest Intelligence (95.3) but places #14: its cost score is 28.5 ($0.242 per 1,000 decisions), and the geometric mean does not let accuracy buy that back.
  • Qwen3.8 27B, Needle 3, options as tools, Needle 3 did not answer every tier — each for the reason in its † note; they are shown below the ranking as partial runs, without a rank.

Axes, tiers, latency and cost

Sort by any column; values the run could not produce always sort last. Hover a cost for how it was priced, a latency for the endpoint. Names link to each project.

⏱ Latency of self-hosted and demo endpoints is adjusted ×2 (+0.15 s on our own servers) to approximate production load — an assumption, not a measurement; raw measurements are in the table and the repo.💲 $ per 1,000 decisions, not per 1,000 tokens — one decision ≈ 950 input tokens.

Rank#SystemJevBench Scoreofficial↓Intelligence25 %Calibration25 %Speed25 %Cost25 %$ / 1,000 decisionsnot tokensEasy72 dec. · 14 %Standard96 dec. · 28 %Judge146 dec. · 28 %Hard220 dec. · 30 %Latencyp50 · p95, raw → adjustedEndpoint
1by TypeSafe AIJev 1.13.074.485.782.783.352.0$0.040100.0%99.0%94.5%74.1%0.65 s rawp95 0.72 s rawproduction API
2by Theodore Lee (TheoLeeCJ)SemIfformerly OpenJev (Qwen3.5-4B, TheoLeeCJ73.179.072.683.759.5~$0.022 est.100.0%97.9%95.2%59.5%0.20 s raw→ 0.55 s adjustedp95 0.32 s raw → 0.78 sour RunPod GPU
3by Maisa (David Villalón)djevMaisa, diffusion-gemma73.082.765.491.457.6$0.026 announced100.0%97.9%93.2%69.5%0.24 s rawp95 0.31 s rawproduction API
4by Eldan RingWinnow-12B Q871.282.072.082.352.9~$0.037 est.100.0%96.9%91.1%70.9%0.23 s raw→ 0.60 s adjustedp95 0.41 s raw → 0.98 sour RunPod GPU
5by kshetrajna12reflex 4B70.380.175.268.059.7~$0.022 est.100.0%94.8%97.3%63.2%1.80 s raw→ 3.75 s adjustedp95 2.05 s raw → 4.26 sour RunPod GPU
6by hjmurmur (Octalab)jqvQwen3-32B zero-shot68.679.379.074.647.5~$0.056 est.100.0%95.8%92.5%64.5%0.75 s raw→ 1.64 s adjustedp95 0.97 s raw → 2.10 sour RunPod GPU
7by milliseconds.ai (Baptiste Laget)decision-machine-1milliseconds.ai68.362.170.492.953.7$0.035100.0%76.0%89.7%46.8%0.17 s rawp95 0.30 s rawproduction API
8by Mapikadecider-35b-a3b67.679.671.580.845.3~$0.067 est.100.0%96.9%91.1%65.5%0.29 s raw→ 0.73 s adjustedp95 0.49 s raw → 1.14 sour RunPod GPU
9by IkerMoelopen-alternative-jevQwen3.5-4B, IkerMoel67.064.063.283.559.6~$0.022 est.100.0%84.4%74.7%56.8%0.21 s raw→ 0.56 s adjustedp95 0.32 s raw → 0.80 sour RunPod GPU
10by mithalounisystem-one-openGemma 4 E2B LoRA on an L466.669.556.777.064.8~$0.015 est.100.0%93.8%87.7%49.1%0.65 s raw→ 1.30 s adjustedp95 0.77 s raw → 1.54 sauthor's demo server
11by razorback16 / CodivOpenJevDiffusionGemma 26B-A4B NVFP4, razorback1666.479.264.883.245.5~$0.066 est.100.0%95.8%91.1%65.5%0.24 s raw→ 0.63 s adjustedp95 0.31 s raw → 0.76 sour RunPod GPU
12by Featherless AISimpleJev Qwen3.8-27B66.384.781.171.239.5~$0.104 est.100.0%96.9%93.2%75.0%1.01 s raw→ 2.03 s adjustedp95 1.88 s raw → 3.76 sauthor's demo server
13by ZeroEntropyZeroEntropy zerank-266.063.076.579.049.8$0.047100.0%79.2%88.4%47.3%0.13 s raw→ 0.40 s adjustedp95 1.50 s raw → 3.15 sour RunPod GPU
14by OpenAIGPT-5.6 Lunalow reasoning effort65.995.389.877.528.5$0.242100.0%97.9%96.6%94.5%0.97 s rawp95 1.82 s rawproduction API
15by ekzhangopenjev-sglangQwen3.6-35B-A3B on SGLang65.383.477.477.136.5~$0.131 est.100.0%95.8%95.2%71.4%0.68 s raw→ 1.36 s adjustedp95 0.73 s raw → 1.45 sauthor's demo server
16by QwenQwen3-Reranker-4B63.864.067.078.749.2$0.050100.0%79.2%87.7%50.0%0.13 s raw→ 0.41 s adjustedp95 1.56 s raw → 3.27 sour RunPod GPU
17by kshetrajna12reflex-27bQwen3.8-27B63.385.886.267.532.3~$0.181 est.100.0%95.8%95.9%75.9%1.89 s raw→ 3.93 s adjustedp95 2.21 s raw → 4.57 sour RunPod GPU
18by Zhengxu YuLitJevQwen3.8-27B62.782.483.566.733.6~$0.163 est.100.0%97.9%88.4%73.2%2.03 s raw→ 4.20 s adjustedp95 2.46 s raw → 5.06 sour RunPod GPU
19by Jared Palmerkev 0.6Bresearch preview62.551.951.175.676.1~$0.0063 est.100.0%81.3%66.4%40.0%0.59 s raw→ 1.33 s adjustedp95 0.97 s raw → 2.09 sour RunPod GPU
20by Featherless AISimpleJev Qwen3.6-35B-A3B62.579.567.175.038.1~$0.116 est.100.0%93.8%93.2%66.4%0.85 s raw→ 1.70 s adjustedp95 0.93 s raw → 1.86 sauthor's demo server
21by David Villalon / Maisadjevthinking62.480.892.775.226.9~$0.274 est.95.8%99.0%80.1%77.7%0.43 s raw→ 1.00 s adjustedp95 1.45 s raw → 3.05 sour RunPod GPU
22by us (GitHub)jev-localQwen3.5-9B61.870.868.769.243.3~$0.077 est.100.0%84.4%89.0%59.1%1.05 s raw→ 2.24 s adjustedp95 2.62 s raw → 5.38 sour RunPod GPU
23by Mapikadecider-2b61.761.246.683.261.0~$0.020 est.100.0%85.4%77.4%47.3%0.26 s raw→ 0.67 s adjustedp95 0.28 s raw → 0.72 sour RunPod GPU
24by Bespoke LabsBespoke Nimble 9B60.577.965.378.733.4~$0.166 est.100.0%94.8%89.0%65.5%0.39 s raw→ 0.93 s adjustedp95 0.65 s raw → 1.46 sour RunPod GPU
25by GoogleGemini 3.1 Flash-Lite60.185.668.181.827.4$0.264100.0%99.0%93.2%75.0%0.76 s rawp95 0.88 s rawproduction API
26by razorback16OpenJevthinking, BF1660.088.069.676.127.8~$0.255 est.100.0%100.0%94.5%78.2%0.46 s raw→ 1.08 s adjustedp95 1.08 s raw → 2.31 sour RunPod GPU
27by Jared Palmerkev 4Bresearch preview59.764.842.075.761.8~$0.019 est.100.0%91.7%85.6%42.3%0.55 s raw→ 1.25 s adjustedp95 0.99 s raw → 2.13 sour RunPod GPU
28by DeepSeekDeepSeek V4.1 Flashthinking default57.594.396.771.616.8$0.59498.6%99.0%93.2%95.0%1.42 s rawp95 4.89 s rawproduction API
29by Jared Palmerkev 8Bresearch preview56.469.444.274.944.0~$0.073 est.100.0%92.7%90.4%47.3%0.59 s raw→ 1.33 s adjustedp95 1.15 s raw → 2.45 sour RunPod GPU
30by Zefan Cai (@Zefan_Cai)Open-Jev 9BZefan Cai55.071.263.372.028.1~$0.249 est.100.0%90.6%81.5%60.9%0.75 s raw→ 1.66 s adjustedp95 1.81 s raw → 3.77 sour RunPod GPU
31by Sean Goedeckesystem-oneQwen3-8B, Sean Goedecke54.870.336.884.441.5~$0.089 est.100.0%90.6%91.8%50.0%0.17 s raw→ 0.48 s adjustedp95 0.30 s raw → 0.76 sour RunPod GPU
32by Logan MarkewichjeffLogan Markewich, GLiFormer 400M54.446.964.663.576.6~$0.0060 est.100.0%76.0%61.6%37.7%0.94 s raw→ 2.03 s adjustedp95 10.97 s raw → 22.09 sour CPU
33by Convai InnovationsLayaConvai Innovations, ModernBERT-large 421M54.445.862.571.186.2~$0.0029 est.94.4%72.9%69.2%34.1%0.79 s raw→ 1.72 s adjustedp95 2.20 s raw → 4.54 sour CPU
34by Zefan Cai (@Zefan_Cai)Open-Jev 2BZefan Cai51.361.055.173.528.1~$0.249 est.100.0%79.2%88.4%42.7%0.66 s raw→ 1.48 s adjustedp95 1.45 s raw → 3.05 sour RunPod GPU
35by Deepan WadhwaOpenDecisionModernBERT-large zero-shot40.640.856.179.975.3~$0.0066 est.87.5%62.5%71.2%33.2%0.34 s raw→ 0.83 s adjustedp95 0.54 s raw → 1.24 sour RunPod GPU
36by Hemant (heman10x)openJev Verdict 1.438.938.674.178.182.4~$0.0039 est.86.1%67.7%56.2%37.7%0.31 s raw→ 0.78 s adjustedp95 0.92 s raw → 2.00 sour CPU
37by Hemant (heman10x)openJev Verdictheman10x, ModernBERT-base 151M38.139.851.376.783.1~$0.0037 est.86.1%65.6%61.0%38.2%0.28 s raw→ 0.71 s adjustedp95 1.45 s raw → 3.04 sour CPU
38by Jared Palmerkev 0.5B33.238.247.477.076.1~$0.0063 est.95.8%52.1%71.2%30.9%0.43 s raw→ 1.01 s adjustedp95 0.92 s raw → 1.99 sour RunPod GPU
39by FastinoGLiNER2 large29.640.124.361.773.3~$0.0077 est.98.6%62.5%61.0%36.4%1.10 s raw→ 2.34 s adjustedp95 14.49 s raw → 29.13 sour CPU
40by Aditya (isHeSatoshi)smalljev semantic-v927.435.158.979.857.9~$0.025 est.97.2%68.8%40.4%38.2%0.41 s raw→ 0.98 s adjustedp95 0.46 s raw → 1.07 sour RunPod GPU
41by FastinoGLiNER2Fastino, gliner2.5-base24.035.623.771.883.1~$0.0037 est.97.2%66.7%45.9%36.4%0.31 s raw→ 0.78 s adjustedp95 4.15 s raw → 8.46 sour CPU
42by Kotoba Labsopen-jev-deberta-v3-largelocal CPU23.131.966.466.074.0~$0.0073 est.100.0%49.0%53.4%36.4%1.77 s raw→ 3.69 s adjustedp95 3.35 s raw → 6.85 sour CPU
43by FastinoGLiNER2.5 multiFastino, 287M16.627.756.167.882.4~$0.0039 est.90.3%51.0%43.8%37.7%0.43 s raw→ 1.01 s adjustedp95 8.18 s raw → 16.50 sour CPU
44by FastinoGLiNER2.5 smallFastino, 74M13.825.647.277.882.4~$0.0039 est.83.3%47.9%50.0%33.2%0.11 s raw→ 0.38 s adjustedp95 2.10 s raw → 4.35 sour CPU
45by MixedbreadMixedbread mxbai-rerank-base-v20.86.783.187.567.9$0.01244.4%33.3%26.7%40.0%0.07 s raw→ 0.29 s adjustedp95 0.23 s raw → 0.62 sour RunPod GPU
46by BAAIBAAI bge-reranker-v2-m30.76.383.889.573.4$0.007743.1%36.5%8.9%36.8%0.03 s raw→ 0.22 s adjustedp95 0.18 s raw → 0.51 sour RunPod GPU
47by Alibaba-NLPAlibaba GTE Reranker ModernBERT-base0.34.676.890.669.6$0.01033.3%39.6%30.1%33.6%0.05 s raw→ 0.25 s adjustedp95 0.10 s raw → 0.35 sour RunPod GPU
48by AltSlate LabsCerto v10.00.082.094.0100.0~$0.0010 est.27.8%30.2%21.9%31.8%0.02 s raw→ 0.19 s adjustedp95 0.03 s raw → 0.21 sour RunPod GPU
Honorable mentions — services built on another entrant's model — shown, not ranked: A service that runs another entrant's model is listed with all of its scores and axes, but is not ranked against the models.
by mrmps (@michael_chomsky)classifier.devfast tierhonorable mention · not ranked83.685.177.987.684.3~$0.0033 est.100.0%99.0%97.3%70.5%0.39 s rawp95 0.45 s rawproduction API
Partial runs — shown, not ranked: a tier attempted for fewer than 95 % of its decisions.
by Qwen / ChutesQwen3.8 27BChutes TEEpartial run · not ranked24.867.492.161.30.0~$2.669 est.98.6%99.0%95.3%21.4%5.75 s rawp95 12.97 s rawproduction API
by Cactus ComputeNeedle 3, options as toolspost-hoc adapter modepartial run · not ranked1.113.5none (label only)52.865.3~$0.014 est.66.7%31.3%34.2%3.78 s raw→ 7.71 s adjustedp95 33.64 s raw → 67.42 sour CPU
by Cactus ComputeNeedle 3Cactus, 2-bit, local CPUpartial run · not ranked0.14.6none (label only)59.958.7~$0.024 est.47.2%16.7%31.5%7.7%1.69 s raw→ 3.52 s adjustedp95 14.36 s raw → 28.88 sour CPU

† Notes on 39 marked systems — how each was run

  • djev: The measured endpoint was Maisa's hosted API in free preview; the cost uses its announced price ($0.035 per million input tokens, output free), and nothing was charged. The self-hostable djev-dev runtime is Apache-2.0 and applies a structured one-step inference method to Google's Apache-2.0 diffusiongemma-26B-A4B-it checkpoint; it adds no separately trained djev weights. Probabilities are djev's own (its docs call them experimental and uncalibrated).
  • Winnow-12B Q8: The submitted Q8_0 GGUF ran through the pinned author's TypeSafe-compatible /v1/systemone server with 8,192 context, four resident decision branches, Q8 KV, and full GPU offload. The private training corpus was not released. The author's checksum-based audit reports zero exact public-item overlap, but that claim cannot be independently reproduced; our scan found no exact public state or instruction text in the released artifacts. Cost uses the $0.05/M-input hosted Gemma 3 12B reference, not free/100.
  • reflex 4B: The author's reflex-serve: Qwen3.5-4B with the published LoRA and its per-primitive calibration file; the state is encoded once and each question read from the label logits. Run serially on our GPU; the author discloses that the 231 public items were used four times as a development gate.
  • jqv: A stock Qwen3-32B with no decision training: the state is prefilled once, each question is an isolated branch and the answer is read from the option-letter logits, with one fitted temperature (3.02, 400 MMLU validation items). Re-run in v1.2.8 on our own GPU from the now-public serving code (Octalab-Inc/jqv 0189b67), so all 534 decisions including the held-out hard items were asked; this full run replaces the v1.2.7 partial row, which had been measured on the submitter's machine. Cost is the base model's public per-token tariff, not free.
  • decision-machine-1: A closed-weights decision model behind a production API that serves TypeSafe's wire format, so the unchanged typesafe adapter ran it. Run on a free test key (30 requests a minute, 2.2 s between requests); the provider states the inference infrastructure is the same as for paid keys. Cost is the public paid tariff, $0.04 per million input tokens (output free), times the input tokens the API reported.
  • decider-35b-a3b: The author's TypeSafe-compatible server and published FP8 weights, run serially on our H100 NVL. The exhaustive startup batch warmup was skipped; each required serial shape captured lazily before its measured request. Self-host latency receives the standard ×2 + 0.15 s adjustment. Cost uses the closest hosted 35B-A3B input tariff and is not the temporary rental charge.
  • open-alternative-jev: With the options in reverse order (A. no, B. yes) the same model scored 21 % instead of 72 % on yes/no answer-judging items — small models are very sensitive to option order.
  • SimpleJev Qwen3.8-27B: Author's no-login shared demo, model id recorded verbatim, one request at a time at or below its 2 RPS limit. SimpleJev reads answer-token logits and returns the complete distribution; it does not generate an answer. Speed uses the public-demo x2 load adjustment; cost uses a hosted size-class input price and is not free/100.
  • ZeroEntropy zerank-2: Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.
  • Qwen3-Reranker-4B: Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.
  • reflex-27b: The frozen public Qwen3.8-27B checkpoint through reflex at the requested pinned commit, with two option orders averaged and temperature 1. No adapter or fitted calibration file. Run serially on our H100 NVL. Self-host latency receives the standard ×2 + 0.15 s adjustment; cost uses the exact base model's public hosted input tariff.
  • LitJev: The author's reproduction of Jev's decision layer on an off-the-shelf model, in its default configuration: Qwen3.8-27B, scores read from the output head, no training and no calibration file (its README says probabilities are not calibrated by default). Run serially on our GPU through an SSH tunnel, because its server binds to localhost; the request still crosses the internet and gets the ×2 + 0.15 s adjustment.
  • kev 0.6B: Self-hosted from the author's repository at commit 20fa626 through its native TypeSafe-compatible /v1/systemone server, BF16 on an RTX 3090; measured serially from Sandy over the internet. The author labels this checkpoint a research preview.
  • SimpleJev Qwen3.6-35B-A3B: Author's no-login shared demo, model id recorded verbatim, one request at a time at or below its 2 RPS limit. SimpleJev reads answer-token logits and returns the complete distribution; it does not generate an answer. Speed uses the public-demo x2 load adjustment; cost uses a hosted size-class input price and is not free/100.
  • djev: Experimental full-generation path over the same DiffusionGemma checkpoint as djev-dev: thinking was enabled and the model could generate up to 8,192 tokens before returning its distribution. Current djev-dev itself hard-codes enable_thinking=false, diffusion_max_steps=1 and read_only=true, so this is not a switch in its published typed API. It is substantially slower/costlier, and 72/534 requests exhausted the output budget without a parseable distribution; those are failures. Cost uses measured tokens and a same-size hosted reference, not the H200 rental bill.
  • jev-local: The author's local Jev-compatible server in its default full configuration: a frozen Qwen3.5-9B scores each option by its mean log-probability (one forward pass per option, no generation, no decision training). Run serially on our GPU. It re-reads the state once per option; if its reported token count covers one pass only, a per-token hosted price would be higher than this estimate.
  • decider-2b: The author's TypeSafe-compatible server and published weights (Qwen3.5-2B-Base with a trained one-pass decision readout), run serially on our GPU. Self-host latency gets the standard ×2 + 0.15 s adjustment.
  • Bespoke Nimble 9B: Re-run in v1.2.8 at Bespoke Labs' request after they raised the serving prompt limit from 2,048 to 8,192 tokens (bespokelabsai/nimble PR #4). Same recipe as the v1.1.3 run — the published LoRA merged into Qwen3.5-9B with the author's PEFT safe-merge, served with SGLang and the author's Jev-compatible API — now from current nimble main; the adapter weights are unchanged. Hard-tier accuracy rose from 43.6 % to 65.5 %, yet the score fell: the long hard items that used to fail at once are now answered and priced (so Cost fell), and this pod was in Canada while the v1.1.3 run's was in Sweden, so part of the lower Speed is network distance from our server in Germany. This complete run replaces the earlier row; its old score is kept in the artifact under superseded_rows.
  • OpenJev: OpenJev's real typed-API thinking switch at think=512, using its own /v1/systemone server over BF16 DiffusionGemma. The thought is generated first, then native probability reads are taken after it. All 534 requests returned valid distributions. Cost counts the server's billed input and thought output tokens.
  • kev 4B: Self-hosted from the author's repository at commit 20fa626 through its native TypeSafe-compatible /v1/systemone server, BF16 on an RTX 3090; measured serially from Sandy over the internet. The author labels this checkpoint a research preview.
  • kev 8B: Self-hosted from the author's repository at commit 20fa626 through its native TypeSafe-compatible /v1/systemone server, BF16 on an RTX 3090; measured serially from Sandy over the internet. The author labels this checkpoint a research preview.
  • Open-Jev 9B: The author's pinned LoRA adapter, trained scalar decision head and calibration temperature, served by the author's Open-Jev server with prefix caching off, batch size 1 and 4,096-token limit. Serial requests were measured from Sandy over an SSH tunnel to the H100. Self-host latency receives the standing x2 + 0.15 s adjustment. Cost uses the exact Qwen3.5-9B hosted input tariff for 9B and the same conservative same-family proxy for the unlisted 2B; neither receives an automatic 100. Exact normalized comparison found no JevBench public task state or instruction in the 79,116-row public training projection.
  • jeff: Self-hosted from its GitHub repo with server defaults, on our CPU (the author recommends a GPU, e.g. an L4), through the same TypeSafe-compatible API as Jev.
  • Laya: The English checkpoint (repo root), run on our CPU through its own laya package. Its budget is 512 tokens per question, so long hard-tier states are cut by the package itself.
  • Open-Jev 2B: The author's pinned LoRA adapter, trained scalar decision head and calibration temperature, served by the author's Open-Jev server with prefix caching off, batch size 1 and 4,096-token limit. Serial requests were measured from Sandy over an SSH tunnel to the H100. Self-host latency receives the standing x2 + 0.15 s adjustment. Cost uses the exact Qwen3.5-9B hosted input tariff for 9B and the same conservative same-family proxy for the unlisted 2B; neither receives an automatic 100. Exact normalized comparison found no JevBench public task state or instruction in the 79,116-row public training projection.
  • OpenDecision: A zero-shot NLI classifier behind a TypeSafe-compatible server, not a trained decision model: it scores each option as an entailment hypothesis with ModernBERT-large-zeroshot-v2.0. Its choice path runs several NLI passes over the same state, which the reported token count does not include, so a per-token hosted price would be higher than the estimate here. Pre-registered for our CPU in v1.2.7, run on our GPU because the CPU was far too slow.
  • openJev Verdict 1.4: Same public weights as the earlier Verdict row, run through the author's fixed v1.4 engine. That engine auto-loads the calibrator for every option count, frames candidate labels as NLI sentences and uses a 512-token context budget. Run locally on our CPU, serially.
  • openJev Verdict: The openJev-verdict-2.0 Hugging Face repo ships no weights; its config is byte-identical to heman10x/rlcd-modernbert-151m, whose published weights we ran with the author's engine. The 'verdict2-base' checkpoint behind the README's numbers is not downloadable yet (Git LFS 404); we will run it once it is.
  • kev 0.5B: Self-hosted from the author's repository at commit 20fa626 through its native TypeSafe-compatible /v1/systemone server, BF16 on an RTX 3090; measured serially from Sandy over the internet. This is the v0.1 release.
  • GLiNER2 large: The large checkpoint of Fastino's earlier GLiNER2 family, same documented mapping as the GLiNER2 row: the question goes in front of the text and the probabilities are the model's own single-label softmax over the labels, read out in full. A general schema classifier, not a Jev rebuild.
  • smalljev semantic-v9: The public semantic-v9 LoRA and native heads over MiniCPM5-2B-Base, through the mapping frozen before the run. It has a typed Python contract but no TypeSafe-compatible HTTP route. The released training recipe explicitly hill-climbed against JevBench's public shape and source families; this allowed public benchmark-directed development is disclosed. Cost is $0.04/M measured input tokens, not free/100.
  • GLiNER2: A general schema classifier, not a Jev rebuild. The question goes in front of the text; the probabilities are GLiNER2's own single-label softmax over the labels, read out in full (mapping fixed before the run).
  • GLiNER2.5 multi: The multilingual GLiNER2.5 checkpoint (287M), same family and same documented mapping as the GLiNER2 row. JevBench items are English only, so its multilingual training is not exercised here.
  • GLiNER2.5 small: The small GLiNER2.5 checkpoint (74M), same family and same documented mapping as the GLiNER2 row: the question goes in front of the text and the probabilities are the model's own single-label softmax over the labels, read out in full. A general schema classifier, not a Jev rebuild.
  • Mixedbread mxbai-rerank-base-v2: Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.
  • BAAI bge-reranker-v2-m3: Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.
  • Alibaba GTE Reranker ModernBERT-base: Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.
  • Certo v1: The public Certo v1 checkpoint through the author's DecisionModel, serially on our rented GPU. The question instruction is prepended to the state because Certo exposes state + runtime options but no separate question field; the published 64-token state and 48-token option limits are unchanged. The model card says v1 does not yet transfer to arbitrary natural-language prose. Cost is an estimate from same-size hosted encoders times the checkpoint's retained input tokens, not free/100.

PAN's pipeline reviewed approximately 1 open sources for this article. No human editor reviewed this article before publication.

Related Reads

Show on timeline →