JevBench Score (Intelligence, Calibration, Speed, Cost — 25 % each)
OfficialIntelligence above chance, Calibration, Speed, Cost — 25 % each, geometric mean; below 50 Intelligence receives a growing near-chance penalty. Change the weighting ↓
Intel.Calib.SpeedCost$/1k dec. 01. 1Jev 1.13.0 74.4 I 86 · C 83 · S 83 · K 52 · $0.040 02. 2SemIf (Qwen3.5-4B) 73.1 I 79 · C 73 · S 84 · K 59 · ~$0.022 est. 03. 3djev (Maisa, diffusion-gemma)† 73.0 I 83 · C 65 · S 91 · K 58 · $0.026 ann. 04. 4Winnow-12B Q8† 71.2 I 82 · C 72 · S 82 · K 53 · ~$0.037 est. 05. 5reflex 4B† 70.3 I 80 · C 75 · S 68 · K 60 · ~$0.022 est. 06. 6jqv† 68.6 I 79 · C 79 · S 75 · K 47 · ~$0.056 est. 07. 7decision-machine-1† 68.3 I 62 · C 70 · S 93 · K 54 · $0.035 08. 8decider-35b-a3b† 67.6 I 80 · C 72 · S 81 · K 45 · ~$0.067 est. 09. 9open-alternative-jev (Qwen3.5-4B)† 67.0 I 64 · C 63 · S 83 · K 60 · ~$0.022 est. 10. 10system-one-open 66.6 I 70 · C 57 · S 77 · K 65 · ~$0.015 est. 11. 11OpenJev (razorback16) 66.4 I 79 · C 65 · S 83 · K 45 · ~$0.066 est. 12. 12SimpleJev Qwen3.8-27B† 66.3 I 85 · C 81 · S 71 · K 39 · ~$0.104 est. 13. 13ZeroEntropy zerank-2† 66.0 I 63 · C 76 · S 79 · K 50 · $0.047 14. 14GPT-5.6 Luna (low) 65.9 I 95 · C 90 · S 78 · K 28 · $0.242 15. 15openjev-sglang 65.3 I 83 · C 77 · S 77 · K 36 · ~$0.131 est. 16. 16Qwen3-Reranker-4B† 63.8 I 64 · C 67 · S 79 · K 49 · $0.050 17. 17reflex-27b† 63.3 I 86 · C 86 · S 67 · K 32 · ~$0.181 est. 18. 18LitJev† 62.7 I 82 · C 84 · S 67 · K 34 · ~$0.163 est. 19. 19kev 0.6B† 62.5 I 52 · C 51 · S 76 · K 76 · ~$0.0063 est. 20. 20SimpleJev Qwen3.6-35B-A3B† 62.5 I 80 · C 67 · S 75 · K 38 · ~$0.116 est. 21. 21djev† 62.4 I 81 · C 93 · S 75 · K 27 · ~$0.274 est. 22. 22jev-local† 61.8 I 71 · C 69 · S 69 · K 43 · ~$0.077 est. 23. 23decider-2b† 61.7 I 61 · C 47 · S 83 · K 61 · ~$0.020 est. 24. 24Bespoke Nimble 9B† 60.5 I 78 · C 65 · S 79 · K 33 · ~$0.166 est. 25. 25Gemini 3.1 Flash-Lite 60.1 I 86 · C 68 · S 82 · K 27 · $0.264 26. 26OpenJev† 60.0 I 88 · C 70 · S 76 · K 28 · ~$0.255 est. 27. 27kev 4B† 59.7 I 65 · C 42 · S 76 · K 62 · ~$0.019 est. 28. 28DeepSeek V4.1 Flash 57.5 I 94 · C 97 · S 72 · K 17 · $0.594 29. 29kev 8B† 56.4 I 69 · C 44 · S 75 · K 44 · ~$0.073 est. 30. 30Open-Jev 9B† 55.0 I 71 · C 63 · S 72 · K 28 · ~$0.249 est. 31. 31system-one 54.8 I 70 · C 37 · S 84 · K 41 · ~$0.089 est. 32. 32jeff† 54.4 I 47 · C 65 · S 63 · K 77 · ~$0.0060 est. 33. 33Laya† 54.4 I 46 · C 62 · S 71 · K 86 · ~$0.0029 est. 34. 34Open-Jev 2B† 51.3 I 61 · C 55 · S 73 · K 28 · ~$0.249 est. 35. 35OpenDecision† 40.6 I 41 · C 56 · S 80 · K 75 · ~$0.0066 est. 36. 36openJev Verdict 1.4† 38.9 I 39 · C 74 · S 78 · K 82 · ~$0.0039 est. 37. 37openJev Verdict† 38.1 I 40 · C 51 · S 77 · K 83 · ~$0.0037 est. 38. 38kev 0.5B† 33.2 I 38 · C 47 · S 77 · K 76 · ~$0.0063 est. 39. 39GLiNER2 large† 29.6 I 40 · C 24 · S 62 · K 73 · ~$0.0077 est. 40. 40smalljev semantic-v9† 27.4 I 35 · C 59 · S 80 · K 58 · ~$0.025 est. 41. 41GLiNER2† 24.0 I 36 · C 24 · S 72 · K 83 · ~$0.0037 est. 42. 42open-jev-deberta-v3-large 23.1 I 32 · C 66 · S 66 · K 74 · ~$0.0073 est. 43. 43GLiNER2.5 multi† 16.6 I 28 · C 56 · S 68 · K 82 · ~$0.0039 est. 44. 44GLiNER2.5 small† 13.8 I 26 · C 47 · S 78 · K 82 · ~$0.0039 est. 45. 45Mixedbread mxbai-rerank-base-v2† 0.8 I 7 · C 83 · S 88 · K 68 · $0.012 46. 46BAAI bge-reranker-v2-m3† 0.7 I 6 · C 84 · S 90 · K 73 · $0.0077 47. 47Alibaba GTE Reranker ModernBERT-base† 0.3 I 5 · C 77 · S 91 · K 70 · $0.010 48. 48Certo v1† 0.0 I 0 · C 82 · S 94 · K 100 · ~$0.0010 est. 49. classifier.dev (fast tier)† (honorable mention) 83.6 I 85 · C 78 · S 88 · K 84 · ~$0.0033 est. 50. Qwen3.8 27B (partial run) 24.8 I 67 · C 92 · S 61 · K 0 · ~$2.669 est. 51. Needle 3, options as tools (partial run) 1.1 I 14 · C – · S 53 · K 65 · ~$0.014 est. 52. Needle 3 (partial run) 0.1 I 5 · C – · S 60 · K 59 · ~$0.024 est.
020406080100
Score = Intelligence0.25 × Calibration0.25 × Speed0.25 × Cost0.25(each 0–100; geometric mean; below 50 Intelligence, × (I / 50)²)
- Jev (TypeSafe, closed)
- Jev rebuild (open, or open source planned)
- Instruction model, JSON schema
- Small tool-calling model
- Service built on Jev
- Zero-shot classifier (not a Jev rebuild)
- Closed decision model (API only, not Jev)
- Shown, not ranked — honorable mention (runs another entrant's model) · partial run
⏱ Latency of self-hosted and demo endpoints is adjusted ×2 (+0.15 s on our own servers) to approximate production load — an assumption, not a measurement; raw measurements are in the table and the repo.I, C, S, K = Intelligence, Calibration, Speed, Cost; est./ann. = estimated/announced cost; † = see note.Legend and notes
- ~ est. = no measured bill; priced like a large inference provider ( how costs are estimated).
- ann. = the provider’s announced price, not yet charged.
- Names link to each project.
- A label-only system has no calibration (–, counted as 0).
- djev (Maisa, diffusion-gemma): The measured endpoint was Maisa's hosted API in free preview; the cost uses its announced price ($0.035 per million input tokens, output free), and nothing was charged. The self-hostable djev-dev runtime is Apache-2.0 and applies a structured one-step inference method to Google's Apache-2.0 diffusiongemma-26B-A4B-it checkpoint; it adds no separately trained djev weights. Probabilities are djev's own (its docs call them experimental and uncalibrated).
- Winnow-12B Q8: The submitted Q8_0 GGUF ran through the pinned author's TypeSafe-compatible /v1/systemone server with 8,192 context, four resident decision branches, Q8 KV, and full GPU offload. The private training corpus was not released. The author's checksum-based audit reports zero exact public-item overlap, but that claim cannot be independently reproduced; our scan found no exact public state or instruction text in the released artifacts. Cost uses the $0.05/M-input hosted Gemma 3 12B reference, not free/100.
- reflex 4B: The author's reflex-serve: Qwen3.5-4B with the published LoRA and its per-primitive calibration file; the state is encoded once and each question read from the label logits. Run serially on our GPU; the author discloses that the 231 public items were used four times as a development gate.
- jqv: A stock Qwen3-32B with no decision training: the state is prefilled once, each question is an isolated branch and the answer is read from the option-letter logits, with one fitted temperature (3.02, 400 MMLU validation items). Re-run in v1.2.8 on our own GPU from the now-public serving code (Octalab-Inc/jqv 0189b67), so all 534 decisions including the held-out hard items were asked; this full run replaces the v1.2.7 partial row, which had been measured on the submitter's machine. Cost is the base model's public per-token tariff, not free.
- decision-machine-1: A closed-weights decision model behind a production API that serves TypeSafe's wire format, so the unchanged typesafe adapter ran it. Run on a free test key (30 requests a minute, 2.2 s between requests); the provider states the inference infrastructure is the same as for paid keys. Cost is the public paid tariff, $0.04 per million input tokens (output free), times the input tokens the API reported.
- decider-35b-a3b: The author's TypeSafe-compatible server and published FP8 weights, run serially on our H100 NVL. The exhaustive startup batch warmup was skipped; each required serial shape captured lazily before its measured request. Self-host latency receives the standard ×2 + 0.15 s adjustment. Cost uses the closest hosted 35B-A3B input tariff and is not the temporary rental charge.
- open-alternative-jev: With the options in reverse order (A. no, B. yes) the same model scored 21 % instead of 72 % on yes/no answer-judging items — small models are very sensitive to option order.
- SimpleJev Qwen3.8-27B: Author's no-login shared demo, model id recorded verbatim, one request at a time at or below its 2 RPS limit. SimpleJev reads answer-token logits and returns the complete distribution; it does not generate an answer. Speed uses the public-demo x2 load adjustment; cost uses a hosted size-class input price and is not free/100.
- ZeroEntropy zerank-2: Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.
- Qwen3-Reranker-4B: Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.
- reflex-27b: The frozen public Qwen3.8-27B checkpoint through reflex at the requested pinned commit, with two option orders averaged and temperature 1. No adapter or fitted calibration file. Run serially on our H100 NVL. Self-host latency receives the standard ×2 + 0.15 s adjustment; cost uses the exact base model's public hosted input tariff.
- LitJev: The author's reproduction of Jev's decision layer on an off-the-shelf model, in its default configuration: Qwen3.8-27B, scores read from the output head, no training and no calibration file (its README says probabilities are not calibrated by default). Run serially on our GPU through an SSH tunnel, because its server binds to localhost; the request still crosses the internet and gets the ×2 + 0.15 s adjustment.
- kev 0.6B: Self-hosted from the author's repository at commit 20fa626 through its native TypeSafe-compatible
/v1/systemoneserver, BF16 on an RTX 3090; measured serially from Sandy over the internet. The author labels this checkpoint a research preview. - SimpleJev Qwen3.6-35B-A3B: Author's no-login shared demo, model id recorded verbatim, one request at a time at or below its 2 RPS limit. SimpleJev reads answer-token logits and returns the complete distribution; it does not generate an answer. Speed uses the public-demo x2 load adjustment; cost uses a hosted size-class input price and is not free/100.
- djev: Experimental full-generation path over the same DiffusionGemma checkpoint as djev-dev: thinking was enabled and the model could generate up to 8,192 tokens before returning its distribution. Current djev-dev itself hard-codes enable_thinking=false, diffusion_max_steps=1 and read_only=true, so this is not a switch in its published typed API. It is substantially slower/costlier, and 72/534 requests exhausted the output budget without a parseable distribution; those are failures. Cost uses measured tokens and a same-size hosted reference, not the H200 rental bill.
- jev-local: The author's local Jev-compatible server in its default full configuration: a frozen Qwen3.5-9B scores each option by its mean log-probability (one forward pass per option, no generation, no decision training). Run serially on our GPU. It re-reads the state once per option; if its reported token count covers one pass only, a per-token hosted price would be higher than this estimate.
- decider-2b: The author's TypeSafe-compatible server and published weights (Qwen3.5-2B-Base with a trained one-pass decision readout), run serially on our GPU. Self-host latency gets the standard ×2 + 0.15 s adjustment.
- Bespoke Nimble 9B: Re-run in v1.2.8 at Bespoke Labs' request after they raised the serving prompt limit from 2,048 to 8,192 tokens (bespokelabsai/nimble PR #4). Same recipe as the v1.1.3 run — the published LoRA merged into Qwen3.5-9B with the author's PEFT safe-merge, served with SGLang and the author's Jev-compatible API — now from current nimble main; the adapter weights are unchanged. Hard-tier accuracy rose from 43.6 % to 65.5 %, yet the score fell: the long hard items that used to fail at once are now answered and priced (so Cost fell), and this pod was in Canada while the v1.1.3 run's was in Sweden, so part of the lower Speed is network distance from our server in Germany. This complete run replaces the earlier row; its old score is kept in the artifact under superseded_rows.
- OpenJev: OpenJev's real typed-API thinking switch at think=512, using its own /v1/systemone server over BF16 DiffusionGemma. The thought is generated first, then native probability reads are taken after it. All 534 requests returned valid distributions. Cost counts the server's billed input and thought output tokens.
- kev 4B: Self-hosted from the author's repository at commit 20fa626 through its native TypeSafe-compatible
/v1/systemoneserver, BF16 on an RTX 3090; measured serially from Sandy over the internet. The author labels this checkpoint a research preview. - kev 8B: Self-hosted from the author's repository at commit 20fa626 through its native TypeSafe-compatible
/v1/systemoneserver, BF16 on an RTX 3090; measured serially from Sandy over the internet. The author labels this checkpoint a research preview. - Open-Jev 9B: The author's pinned LoRA adapter, trained scalar decision head and calibration temperature, served by the author's Open-Jev server with prefix caching off, batch size 1 and 4,096-token limit. Serial requests were measured from Sandy over an SSH tunnel to the H100. Self-host latency receives the standing x2 + 0.15 s adjustment. Cost uses the exact Qwen3.5-9B hosted input tariff for 9B and the same conservative same-family proxy for the unlisted 2B; neither receives an automatic 100. Exact normalized comparison found no JevBench public task state or instruction in the 79,116-row public training projection.
- jeff: Self-hosted from its GitHub repo with server defaults, on our CPU (the author recommends a GPU, e.g. an L4), through the same TypeSafe-compatible API as Jev.
- Laya: The English checkpoint (repo root), run on our CPU through its own
layapackage. Its budget is 512 tokens per question, so long hard-tier states are cut by the package itself. - Open-Jev 2B: The author's pinned LoRA adapter, trained scalar decision head and calibration temperature, served by the author's Open-Jev server with prefix caching off, batch size 1 and 4,096-token limit. Serial requests were measured from Sandy over an SSH tunnel to the H100. Self-host latency receives the standing x2 + 0.15 s adjustment. Cost uses the exact Qwen3.5-9B hosted input tariff for 9B and the same conservative same-family proxy for the unlisted 2B; neither receives an automatic 100. Exact normalized comparison found no JevBench public task state or instruction in the 79,116-row public training projection.
- OpenDecision: A zero-shot NLI classifier behind a TypeSafe-compatible server, not a trained decision model: it scores each option as an entailment hypothesis with ModernBERT-large-zeroshot-v2.0. Its choice path runs several NLI passes over the same state, which the reported token count does not include, so a per-token hosted price would be higher than the estimate here. Pre-registered for our CPU in v1.2.7, run on our GPU because the CPU was far too slow.
- openJev Verdict 1.4: Same public weights as the earlier Verdict row, run through the author's fixed v1.4 engine. That engine auto-loads the calibrator for every option count, frames candidate labels as NLI sentences and uses a 512-token context budget. Run locally on our CPU, serially.
- openJev Verdict: The openJev-verdict-2.0 Hugging Face repo ships no weights; its config is byte-identical to heman10x/rlcd-modernbert-151m, whose published weights we ran with the author's engine. The 'verdict2-base' checkpoint behind the README's numbers is not downloadable yet (Git LFS 404); we will run it once it is.
- kev 0.5B: Self-hosted from the author's repository at commit 20fa626 through its native TypeSafe-compatible
/v1/systemoneserver, BF16 on an RTX 3090; measured serially from Sandy over the internet. This is the v0.1 release. - GLiNER2 large: The large checkpoint of Fastino's earlier GLiNER2 family, same documented mapping as the GLiNER2 row: the question goes in front of the text and the probabilities are the model's own single-label softmax over the labels, read out in full. A general schema classifier, not a Jev rebuild.
- smalljev semantic-v9: The public semantic-v9 LoRA and native heads over MiniCPM5-2B-Base, through the mapping frozen before the run. It has a typed Python contract but no TypeSafe-compatible HTTP route. The released training recipe explicitly hill-climbed against JevBench's public shape and source families; this allowed public benchmark-directed development is disclosed. Cost is $0.04/M measured input tokens, not free/100.
- GLiNER2: A general schema classifier, not a Jev rebuild. The question goes in front of the text; the probabilities are GLiNER2's own single-label softmax over the labels, read out in full (mapping fixed before the run).
- GLiNER2.5 multi: The multilingual GLiNER2.5 checkpoint (287M), same family and same documented mapping as the GLiNER2 row. JevBench items are English only, so its multilingual training is not exercised here.
- GLiNER2.5 small: The small GLiNER2.5 checkpoint (74M), same family and same documented mapping as the GLiNER2 row: the question goes in front of the text and the probabilities are the model's own single-label softmax over the labels, read out in full. A general schema classifier, not a Jev rebuild.
- Mixedbread mxbai-rerank-base-v2: Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.
- BAAI bge-reranker-v2-m3: Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.
- Alibaba GTE Reranker ModernBERT-base: Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.
- Certo v1: The public Certo v1 checkpoint through the author's DecisionModel, serially on our rented GPU. The question instruction is prepended to the state because Certo exposes state + runtime options but no separate question field; the published 64-token state and 48-token option limits are unchanged. The model card says v1 does not yet transfer to arbitrary natural-language prose. Cost is an estimate from same-size hosted encoders times the checkpoint's retained input tokens, not free/100.
- classifier.dev (fast tier): Its own benchmark page says the fast tier is Jev. Free for us; the price is its published Pro plan ($20/month for 200,000 fast classifications a day) at full use, $0.0033 per 1,000 decisions.
Weighting: Intelligence : Calibration : Speed : Cost
Official default
JevBench Score25:25:25:25 · defaultBalanced, no calibration33:0:33:33Emphasis on Accuracy60:0:20:20Emphasis on Speed20:0:60:20Emphasis on Cost20:0:20:60Custom
Intelligence25 %Calibration25 %Speed25 %Cost25 %
The official JevBench Score weights the four axes 25 % each and takes their geometric mean. The other buttons are the earlier views (Balanced 33:33:33 and the three “Emphasis on” weightings, which leave Calibration out), recomputed the same way. Any of them is your view, recomputed in your browser from the published axis scores — not the published score.
Explore by task difficulty
All tasks is the published default. Choose a scope to see how the ranking changes by difficulty. Hard only uses all 220 hard-tier decisions and their measured Intelligence, Calibration, Speed and Cost.
All tasksEasy + MediumEasy onlyHard only
Tier mapping: Easy = easy; Medium = standard. Easy scopes change Intelligence only. Hard only measures all four axes on the same hard-tier subset; systems without a hard-tier run are shown as partial and are not ranked.
What the run says (JevBench Score)
- Jev 1.13.0 (TypeSafe AI) leads with 74.4: Intelligence 85.7, Calibration 82.7, Speed 83.3, Cost 52.0 ($0.040 per 1,000 decisions).
- classifier.dev scores 83.6 — higher than anything in the ranking — but is not ranked: it runs Jev (TypeSafe), so ranking it would put the same model in the list twice, once at the model's own price and once at the service's. It keeps every number it earned under Honorable mentions — services built on another entrant's model.
- Open rebuilds of Jev appeared within days. The best of them, SemIf, formerly OpenJev (Qwen3.5-4B, TheoLeeCJ), is #2 at 73.1 — 1.3 points behind: more speed and a lower (estimated) price, less intelligence and calibration.
- GPT-5.6 Luna (low reasoning effort) has the highest Intelligence (95.3) but places #14: its cost score is 28.5 ($0.242 per 1,000 decisions), and the geometric mean does not let accuracy buy that back.
- Qwen3.8 27B, Needle 3, options as tools, Needle 3 did not answer every tier — each for the reason in its † note; they are shown below the ranking as partial runs, without a rank.
Axes, tiers, latency and cost
Sort by any column; values the run could not produce always sort last. Hover a cost for how it was priced, a latency for the endpoint. Names link to each project.
⏱ Latency of self-hosted and demo endpoints is adjusted ×2 (+0.15 s on our own servers) to approximate production load — an assumption, not a measurement; raw measurements are in the table and the repo.💲 $ per 1,000 decisions, not per 1,000 tokens — one decision ≈ 950 input tokens.
| Rank# | System | JevBench Scoreofficial↓ | Intelligence25 % | Calibration25 % | Speed25 % | Cost25 % | $ / 1,000 decisionsnot tokens | Easy72 dec. · 14 % | Standard96 dec. · 28 % | Judge146 dec. · 28 % | Hard220 dec. · 30 % | Latencyp50 · p95, raw → adjusted | Endpoint |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | by TypeSafe AIJev 1.13.0 | 74.4 | 85.7 | 82.7 | 83.3 | 52.0 | $0.040 | 100.0% | 99.0% | 94.5% | 74.1% | 0.65 s rawp95 0.72 s raw | production API |
| 2 | by Theodore Lee (TheoLeeCJ)SemIfformerly OpenJev (Qwen3.5-4B, TheoLeeCJ | 73.1 | 79.0 | 72.6 | 83.7 | 59.5 | ~$0.022 est. | 100.0% | 97.9% | 95.2% | 59.5% | 0.20 s raw→ 0.55 s adjustedp95 0.32 s raw → 0.78 s | our RunPod GPU |
| 3 | by Maisa (David Villalón)djev†Maisa, diffusion-gemma | 73.0 | 82.7 | 65.4 | 91.4 | 57.6 | $0.026 announced | 100.0% | 97.9% | 93.2% | 69.5% | 0.24 s rawp95 0.31 s raw | production API |
| 4 | by Eldan RingWinnow-12B Q8† | 71.2 | 82.0 | 72.0 | 82.3 | 52.9 | ~$0.037 est. | 100.0% | 96.9% | 91.1% | 70.9% | 0.23 s raw→ 0.60 s adjustedp95 0.41 s raw → 0.98 s | our RunPod GPU |
| 5 | by kshetrajna12reflex 4B† | 70.3 | 80.1 | 75.2 | 68.0 | 59.7 | ~$0.022 est. | 100.0% | 94.8% | 97.3% | 63.2% | 1.80 s raw→ 3.75 s adjustedp95 2.05 s raw → 4.26 s | our RunPod GPU |
| 6 | by hjmurmur (Octalab)jqv†Qwen3-32B zero-shot | 68.6 | 79.3 | 79.0 | 74.6 | 47.5 | ~$0.056 est. | 100.0% | 95.8% | 92.5% | 64.5% | 0.75 s raw→ 1.64 s adjustedp95 0.97 s raw → 2.10 s | our RunPod GPU |
| 7 | by milliseconds.ai (Baptiste Laget)decision-machine-1†milliseconds.ai | 68.3 | 62.1 | 70.4 | 92.9 | 53.7 | $0.035 | 100.0% | 76.0% | 89.7% | 46.8% | 0.17 s rawp95 0.30 s raw | production API |
| 8 | by Mapikadecider-35b-a3b† | 67.6 | 79.6 | 71.5 | 80.8 | 45.3 | ~$0.067 est. | 100.0% | 96.9% | 91.1% | 65.5% | 0.29 s raw→ 0.73 s adjustedp95 0.49 s raw → 1.14 s | our RunPod GPU |
| 9 | by IkerMoelopen-alternative-jev†Qwen3.5-4B, IkerMoel | 67.0 | 64.0 | 63.2 | 83.5 | 59.6 | ~$0.022 est. | 100.0% | 84.4% | 74.7% | 56.8% | 0.21 s raw→ 0.56 s adjustedp95 0.32 s raw → 0.80 s | our RunPod GPU |
| 10 | by mithalounisystem-one-openGemma 4 E2B LoRA on an L4 | 66.6 | 69.5 | 56.7 | 77.0 | 64.8 | ~$0.015 est. | 100.0% | 93.8% | 87.7% | 49.1% | 0.65 s raw→ 1.30 s adjustedp95 0.77 s raw → 1.54 s | author's demo server |
| 11 | by razorback16 / CodivOpenJevDiffusionGemma 26B-A4B NVFP4, razorback16 | 66.4 | 79.2 | 64.8 | 83.2 | 45.5 | ~$0.066 est. | 100.0% | 95.8% | 91.1% | 65.5% | 0.24 s raw→ 0.63 s adjustedp95 0.31 s raw → 0.76 s | our RunPod GPU |
| 12 | by Featherless AISimpleJev Qwen3.8-27B† | 66.3 | 84.7 | 81.1 | 71.2 | 39.5 | ~$0.104 est. | 100.0% | 96.9% | 93.2% | 75.0% | 1.01 s raw→ 2.03 s adjustedp95 1.88 s raw → 3.76 s | author's demo server |
| 13 | by ZeroEntropyZeroEntropy zerank-2† | 66.0 | 63.0 | 76.5 | 79.0 | 49.8 | $0.047 | 100.0% | 79.2% | 88.4% | 47.3% | 0.13 s raw→ 0.40 s adjustedp95 1.50 s raw → 3.15 s | our RunPod GPU |
| 14 | by OpenAIGPT-5.6 Lunalow reasoning effort | 65.9 | 95.3 | 89.8 | 77.5 | 28.5 | $0.242 | 100.0% | 97.9% | 96.6% | 94.5% | 0.97 s rawp95 1.82 s raw | production API |
| 15 | by ekzhangopenjev-sglangQwen3.6-35B-A3B on SGLang | 65.3 | 83.4 | 77.4 | 77.1 | 36.5 | ~$0.131 est. | 100.0% | 95.8% | 95.2% | 71.4% | 0.68 s raw→ 1.36 s adjustedp95 0.73 s raw → 1.45 s | author's demo server |
| 16 | by QwenQwen3-Reranker-4B† | 63.8 | 64.0 | 67.0 | 78.7 | 49.2 | $0.050 | 100.0% | 79.2% | 87.7% | 50.0% | 0.13 s raw→ 0.41 s adjustedp95 1.56 s raw → 3.27 s | our RunPod GPU |
| 17 | by kshetrajna12reflex-27b†Qwen3.8-27B | 63.3 | 85.8 | 86.2 | 67.5 | 32.3 | ~$0.181 est. | 100.0% | 95.8% | 95.9% | 75.9% | 1.89 s raw→ 3.93 s adjustedp95 2.21 s raw → 4.57 s | our RunPod GPU |
| 18 | by Zhengxu YuLitJev†Qwen3.8-27B | 62.7 | 82.4 | 83.5 | 66.7 | 33.6 | ~$0.163 est. | 100.0% | 97.9% | 88.4% | 73.2% | 2.03 s raw→ 4.20 s adjustedp95 2.46 s raw → 5.06 s | our RunPod GPU |
| 19 | by Jared Palmerkev 0.6B†research preview | 62.5 | 51.9 | 51.1 | 75.6 | 76.1 | ~$0.0063 est. | 100.0% | 81.3% | 66.4% | 40.0% | 0.59 s raw→ 1.33 s adjustedp95 0.97 s raw → 2.09 s | our RunPod GPU |
| 20 | by Featherless AISimpleJev Qwen3.6-35B-A3B† | 62.5 | 79.5 | 67.1 | 75.0 | 38.1 | ~$0.116 est. | 100.0% | 93.8% | 93.2% | 66.4% | 0.85 s raw→ 1.70 s adjustedp95 0.93 s raw → 1.86 s | author's demo server |
| 21 | by David Villalon / Maisadjev†thinking | 62.4 | 80.8 | 92.7 | 75.2 | 26.9 | ~$0.274 est. | 95.8% | 99.0% | 80.1% | 77.7% | 0.43 s raw→ 1.00 s adjustedp95 1.45 s raw → 3.05 s | our RunPod GPU |
| 22 | by us (GitHub)jev-local†Qwen3.5-9B | 61.8 | 70.8 | 68.7 | 69.2 | 43.3 | ~$0.077 est. | 100.0% | 84.4% | 89.0% | 59.1% | 1.05 s raw→ 2.24 s adjustedp95 2.62 s raw → 5.38 s | our RunPod GPU |
| 23 | by Mapikadecider-2b† | 61.7 | 61.2 | 46.6 | 83.2 | 61.0 | ~$0.020 est. | 100.0% | 85.4% | 77.4% | 47.3% | 0.26 s raw→ 0.67 s adjustedp95 0.28 s raw → 0.72 s | our RunPod GPU |
| 24 | by Bespoke LabsBespoke Nimble 9B† | 60.5 | 77.9 | 65.3 | 78.7 | 33.4 | ~$0.166 est. | 100.0% | 94.8% | 89.0% | 65.5% | 0.39 s raw→ 0.93 s adjustedp95 0.65 s raw → 1.46 s | our RunPod GPU |
| 25 | by GoogleGemini 3.1 Flash-Lite | 60.1 | 85.6 | 68.1 | 81.8 | 27.4 | $0.264 | 100.0% | 99.0% | 93.2% | 75.0% | 0.76 s rawp95 0.88 s raw | production API |
| 26 | by razorback16OpenJev†thinking, BF16 | 60.0 | 88.0 | 69.6 | 76.1 | 27.8 | ~$0.255 est. | 100.0% | 100.0% | 94.5% | 78.2% | 0.46 s raw→ 1.08 s adjustedp95 1.08 s raw → 2.31 s | our RunPod GPU |
| 27 | by Jared Palmerkev 4B†research preview | 59.7 | 64.8 | 42.0 | 75.7 | 61.8 | ~$0.019 est. | 100.0% | 91.7% | 85.6% | 42.3% | 0.55 s raw→ 1.25 s adjustedp95 0.99 s raw → 2.13 s | our RunPod GPU |
| 28 | by DeepSeekDeepSeek V4.1 Flashthinking default | 57.5 | 94.3 | 96.7 | 71.6 | 16.8 | $0.594 | 98.6% | 99.0% | 93.2% | 95.0% | 1.42 s rawp95 4.89 s raw | production API |
| 29 | by Jared Palmerkev 8B†research preview | 56.4 | 69.4 | 44.2 | 74.9 | 44.0 | ~$0.073 est. | 100.0% | 92.7% | 90.4% | 47.3% | 0.59 s raw→ 1.33 s adjustedp95 1.15 s raw → 2.45 s | our RunPod GPU |
| 30 | by Zefan Cai (@Zefan_Cai)Open-Jev 9B†Zefan Cai | 55.0 | 71.2 | 63.3 | 72.0 | 28.1 | ~$0.249 est. | 100.0% | 90.6% | 81.5% | 60.9% | 0.75 s raw→ 1.66 s adjustedp95 1.81 s raw → 3.77 s | our RunPod GPU |
| 31 | by Sean Goedeckesystem-oneQwen3-8B, Sean Goedecke | 54.8 | 70.3 | 36.8 | 84.4 | 41.5 | ~$0.089 est. | 100.0% | 90.6% | 91.8% | 50.0% | 0.17 s raw→ 0.48 s adjustedp95 0.30 s raw → 0.76 s | our RunPod GPU |
| 32 | by Logan Markewichjeff†Logan Markewich, GLiFormer 400M | 54.4 | 46.9 | 64.6 | 63.5 | 76.6 | ~$0.0060 est. | 100.0% | 76.0% | 61.6% | 37.7% | 0.94 s raw→ 2.03 s adjustedp95 10.97 s raw → 22.09 s | our CPU |
| 33 | by Convai InnovationsLaya†Convai Innovations, ModernBERT-large 421M | 54.4 | 45.8 | 62.5 | 71.1 | 86.2 | ~$0.0029 est. | 94.4% | 72.9% | 69.2% | 34.1% | 0.79 s raw→ 1.72 s adjustedp95 2.20 s raw → 4.54 s | our CPU |
| 34 | by Zefan Cai (@Zefan_Cai)Open-Jev 2B†Zefan Cai | 51.3 | 61.0 | 55.1 | 73.5 | 28.1 | ~$0.249 est. | 100.0% | 79.2% | 88.4% | 42.7% | 0.66 s raw→ 1.48 s adjustedp95 1.45 s raw → 3.05 s | our RunPod GPU |
| 35 | by Deepan WadhwaOpenDecision†ModernBERT-large zero-shot | 40.6 | 40.8 | 56.1 | 79.9 | 75.3 | ~$0.0066 est. | 87.5% | 62.5% | 71.2% | 33.2% | 0.34 s raw→ 0.83 s adjustedp95 0.54 s raw → 1.24 s | our RunPod GPU |
| 36 | by Hemant (heman10x)openJev Verdict 1.4† | 38.9 | 38.6 | 74.1 | 78.1 | 82.4 | ~$0.0039 est. | 86.1% | 67.7% | 56.2% | 37.7% | 0.31 s raw→ 0.78 s adjustedp95 0.92 s raw → 2.00 s | our CPU |
| 37 | by Hemant (heman10x)openJev Verdict†heman10x, ModernBERT-base 151M | 38.1 | 39.8 | 51.3 | 76.7 | 83.1 | ~$0.0037 est. | 86.1% | 65.6% | 61.0% | 38.2% | 0.28 s raw→ 0.71 s adjustedp95 1.45 s raw → 3.04 s | our CPU |
| 38 | by Jared Palmerkev 0.5B† | 33.2 | 38.2 | 47.4 | 77.0 | 76.1 | ~$0.0063 est. | 95.8% | 52.1% | 71.2% | 30.9% | 0.43 s raw→ 1.01 s adjustedp95 0.92 s raw → 1.99 s | our RunPod GPU |
| 39 | by FastinoGLiNER2 large† | 29.6 | 40.1 | 24.3 | 61.7 | 73.3 | ~$0.0077 est. | 98.6% | 62.5% | 61.0% | 36.4% | 1.10 s raw→ 2.34 s adjustedp95 14.49 s raw → 29.13 s | our CPU |
| 40 | by Aditya (isHeSatoshi)smalljev semantic-v9† | 27.4 | 35.1 | 58.9 | 79.8 | 57.9 | ~$0.025 est. | 97.2% | 68.8% | 40.4% | 38.2% | 0.41 s raw→ 0.98 s adjustedp95 0.46 s raw → 1.07 s | our RunPod GPU |
| 41 | by FastinoGLiNER2†Fastino, gliner2.5-base | 24.0 | 35.6 | 23.7 | 71.8 | 83.1 | ~$0.0037 est. | 97.2% | 66.7% | 45.9% | 36.4% | 0.31 s raw→ 0.78 s adjustedp95 4.15 s raw → 8.46 s | our CPU |
| 42 | by Kotoba Labsopen-jev-deberta-v3-largelocal CPU | 23.1 | 31.9 | 66.4 | 66.0 | 74.0 | ~$0.0073 est. | 100.0% | 49.0% | 53.4% | 36.4% | 1.77 s raw→ 3.69 s adjustedp95 3.35 s raw → 6.85 s | our CPU |
| 43 | by FastinoGLiNER2.5 multi†Fastino, 287M | 16.6 | 27.7 | 56.1 | 67.8 | 82.4 | ~$0.0039 est. | 90.3% | 51.0% | 43.8% | 37.7% | 0.43 s raw→ 1.01 s adjustedp95 8.18 s raw → 16.50 s | our CPU |
| 44 | by FastinoGLiNER2.5 small†Fastino, 74M | 13.8 | 25.6 | 47.2 | 77.8 | 82.4 | ~$0.0039 est. | 83.3% | 47.9% | 50.0% | 33.2% | 0.11 s raw→ 0.38 s adjustedp95 2.10 s raw → 4.35 s | our CPU |
| 45 | by MixedbreadMixedbread mxbai-rerank-base-v2† | 0.8 | 6.7 | 83.1 | 87.5 | 67.9 | $0.012 | 44.4% | 33.3% | 26.7% | 40.0% | 0.07 s raw→ 0.29 s adjustedp95 0.23 s raw → 0.62 s | our RunPod GPU |
| 46 | by BAAIBAAI bge-reranker-v2-m3† | 0.7 | 6.3 | 83.8 | 89.5 | 73.4 | $0.0077 | 43.1% | 36.5% | 8.9% | 36.8% | 0.03 s raw→ 0.22 s adjustedp95 0.18 s raw → 0.51 s | our RunPod GPU |
| 47 | by Alibaba-NLPAlibaba GTE Reranker ModernBERT-base† | 0.3 | 4.6 | 76.8 | 90.6 | 69.6 | $0.010 | 33.3% | 39.6% | 30.1% | 33.6% | 0.05 s raw→ 0.25 s adjustedp95 0.10 s raw → 0.35 s | our RunPod GPU |
| 48 | by AltSlate LabsCerto v1† | 0.0 | 0.0 | 82.0 | 94.0 | 100.0 | ~$0.0010 est. | 27.8% | 30.2% | 21.9% | 31.8% | 0.02 s raw→ 0.19 s adjustedp95 0.03 s raw → 0.21 s | our RunPod GPU |
| Honorable mentions — services built on another entrant's model — shown, not ranked: A service that runs another entrant's model is listed with all of its scores and axes, but is not ranked against the models. | |||||||||||||
| by mrmps (@michael_chomsky)classifier.dev†fast tierhonorable mention · not ranked | 83.6 | 85.1 | 77.9 | 87.6 | 84.3 | ~$0.0033 est. | 100.0% | 99.0% | 97.3% | 70.5% | 0.39 s rawp95 0.45 s raw | production API | |
| Partial runs — shown, not ranked: a tier attempted for fewer than 95 % of its decisions. | |||||||||||||
| by Qwen / ChutesQwen3.8 27BChutes TEEpartial run · not ranked | 24.8 | 67.4 | 92.1 | 61.3 | 0.0 | ~$2.669 est. | 98.6% | 99.0% | 95.3% | 21.4% | 5.75 s rawp95 12.97 s raw | production API | |
| by Cactus ComputeNeedle 3, options as toolspost-hoc adapter modepartial run · not ranked | 1.1 | 13.5 | none (label only) | 52.8 | 65.3 | ~$0.014 est. | 66.7% | 31.3% | 34.2% | — | 3.78 s raw→ 7.71 s adjustedp95 33.64 s raw → 67.42 s | our CPU | |
| by Cactus ComputeNeedle 3Cactus, 2-bit, local CPUpartial run · not ranked | 0.1 | 4.6 | none (label only) | 59.9 | 58.7 | ~$0.024 est. | 47.2% | 16.7% | 31.5% | 7.7% | 1.69 s raw→ 3.52 s adjustedp95 14.36 s raw → 28.88 s | our CPU |
† Notes on 39 marked systems — how each was run
- † djev: The measured endpoint was Maisa's hosted API in free preview; the cost uses its announced price ($0.035 per million input tokens, output free), and nothing was charged. The self-hostable djev-dev runtime is Apache-2.0 and applies a structured one-step inference method to Google's Apache-2.0 diffusiongemma-26B-A4B-it checkpoint; it adds no separately trained djev weights. Probabilities are djev's own (its docs call them experimental and uncalibrated).
- † Winnow-12B Q8: The submitted Q8_0 GGUF ran through the pinned author's TypeSafe-compatible /v1/systemone server with 8,192 context, four resident decision branches, Q8 KV, and full GPU offload. The private training corpus was not released. The author's checksum-based audit reports zero exact public-item overlap, but that claim cannot be independently reproduced; our scan found no exact public state or instruction text in the released artifacts. Cost uses the $0.05/M-input hosted Gemma 3 12B reference, not free/100.
- † reflex 4B: The author's reflex-serve: Qwen3.5-4B with the published LoRA and its per-primitive calibration file; the state is encoded once and each question read from the label logits. Run serially on our GPU; the author discloses that the 231 public items were used four times as a development gate.
- † jqv: A stock Qwen3-32B with no decision training: the state is prefilled once, each question is an isolated branch and the answer is read from the option-letter logits, with one fitted temperature (3.02, 400 MMLU validation items). Re-run in v1.2.8 on our own GPU from the now-public serving code (Octalab-Inc/jqv 0189b67), so all 534 decisions including the held-out hard items were asked; this full run replaces the v1.2.7 partial row, which had been measured on the submitter's machine. Cost is the base model's public per-token tariff, not free.
- † decision-machine-1: A closed-weights decision model behind a production API that serves TypeSafe's wire format, so the unchanged typesafe adapter ran it. Run on a free test key (30 requests a minute, 2.2 s between requests); the provider states the inference infrastructure is the same as for paid keys. Cost is the public paid tariff, $0.04 per million input tokens (output free), times the input tokens the API reported.
- † decider-35b-a3b: The author's TypeSafe-compatible server and published FP8 weights, run serially on our H100 NVL. The exhaustive startup batch warmup was skipped; each required serial shape captured lazily before its measured request. Self-host latency receives the standard ×2 + 0.15 s adjustment. Cost uses the closest hosted 35B-A3B input tariff and is not the temporary rental charge.
- † open-alternative-jev: With the options in reverse order (A. no, B. yes) the same model scored 21 % instead of 72 % on yes/no answer-judging items — small models are very sensitive to option order.
- † SimpleJev Qwen3.8-27B: Author's no-login shared demo, model id recorded verbatim, one request at a time at or below its 2 RPS limit. SimpleJev reads answer-token logits and returns the complete distribution; it does not generate an answer. Speed uses the public-demo x2 load adjustment; cost uses a hosted size-class input price and is not free/100.
- † ZeroEntropy zerank-2: Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.
- † Qwen3-Reranker-4B: Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.
- † reflex-27b: The frozen public Qwen3.8-27B checkpoint through reflex at the requested pinned commit, with two option orders averaged and temperature 1. No adapter or fitted calibration file. Run serially on our H100 NVL. Self-host latency receives the standard ×2 + 0.15 s adjustment; cost uses the exact base model's public hosted input tariff.
- † LitJev: The author's reproduction of Jev's decision layer on an off-the-shelf model, in its default configuration: Qwen3.8-27B, scores read from the output head, no training and no calibration file (its README says probabilities are not calibrated by default). Run serially on our GPU through an SSH tunnel, because its server binds to localhost; the request still crosses the internet and gets the ×2 + 0.15 s adjustment.
- † kev 0.6B: Self-hosted from the author's repository at commit 20fa626 through its native TypeSafe-compatible
/v1/systemoneserver, BF16 on an RTX 3090; measured serially from Sandy over the internet. The author labels this checkpoint a research preview. - † SimpleJev Qwen3.6-35B-A3B: Author's no-login shared demo, model id recorded verbatim, one request at a time at or below its 2 RPS limit. SimpleJev reads answer-token logits and returns the complete distribution; it does not generate an answer. Speed uses the public-demo x2 load adjustment; cost uses a hosted size-class input price and is not free/100.
- † djev: Experimental full-generation path over the same DiffusionGemma checkpoint as djev-dev: thinking was enabled and the model could generate up to 8,192 tokens before returning its distribution. Current djev-dev itself hard-codes enable_thinking=false, diffusion_max_steps=1 and read_only=true, so this is not a switch in its published typed API. It is substantially slower/costlier, and 72/534 requests exhausted the output budget without a parseable distribution; those are failures. Cost uses measured tokens and a same-size hosted reference, not the H200 rental bill.
- † jev-local: The author's local Jev-compatible server in its default full configuration: a frozen Qwen3.5-9B scores each option by its mean log-probability (one forward pass per option, no generation, no decision training). Run serially on our GPU. It re-reads the state once per option; if its reported token count covers one pass only, a per-token hosted price would be higher than this estimate.
- † decider-2b: The author's TypeSafe-compatible server and published weights (Qwen3.5-2B-Base with a trained one-pass decision readout), run serially on our GPU. Self-host latency gets the standard ×2 + 0.15 s adjustment.
- † Bespoke Nimble 9B: Re-run in v1.2.8 at Bespoke Labs' request after they raised the serving prompt limit from 2,048 to 8,192 tokens (bespokelabsai/nimble PR #4). Same recipe as the v1.1.3 run — the published LoRA merged into Qwen3.5-9B with the author's PEFT safe-merge, served with SGLang and the author's Jev-compatible API — now from current nimble main; the adapter weights are unchanged. Hard-tier accuracy rose from 43.6 % to 65.5 %, yet the score fell: the long hard items that used to fail at once are now answered and priced (so Cost fell), and this pod was in Canada while the v1.1.3 run's was in Sweden, so part of the lower Speed is network distance from our server in Germany. This complete run replaces the earlier row; its old score is kept in the artifact under superseded_rows.
- † OpenJev: OpenJev's real typed-API thinking switch at think=512, using its own /v1/systemone server over BF16 DiffusionGemma. The thought is generated first, then native probability reads are taken after it. All 534 requests returned valid distributions. Cost counts the server's billed input and thought output tokens.
- † kev 4B: Self-hosted from the author's repository at commit 20fa626 through its native TypeSafe-compatible
/v1/systemoneserver, BF16 on an RTX 3090; measured serially from Sandy over the internet. The author labels this checkpoint a research preview. - † kev 8B: Self-hosted from the author's repository at commit 20fa626 through its native TypeSafe-compatible
/v1/systemoneserver, BF16 on an RTX 3090; measured serially from Sandy over the internet. The author labels this checkpoint a research preview. - † Open-Jev 9B: The author's pinned LoRA adapter, trained scalar decision head and calibration temperature, served by the author's Open-Jev server with prefix caching off, batch size 1 and 4,096-token limit. Serial requests were measured from Sandy over an SSH tunnel to the H100. Self-host latency receives the standing x2 + 0.15 s adjustment. Cost uses the exact Qwen3.5-9B hosted input tariff for 9B and the same conservative same-family proxy for the unlisted 2B; neither receives an automatic 100. Exact normalized comparison found no JevBench public task state or instruction in the 79,116-row public training projection.
- † jeff: Self-hosted from its GitHub repo with server defaults, on our CPU (the author recommends a GPU, e.g. an L4), through the same TypeSafe-compatible API as Jev.
- † Laya: The English checkpoint (repo root), run on our CPU through its own
layapackage. Its budget is 512 tokens per question, so long hard-tier states are cut by the package itself. - † Open-Jev 2B: The author's pinned LoRA adapter, trained scalar decision head and calibration temperature, served by the author's Open-Jev server with prefix caching off, batch size 1 and 4,096-token limit. Serial requests were measured from Sandy over an SSH tunnel to the H100. Self-host latency receives the standing x2 + 0.15 s adjustment. Cost uses the exact Qwen3.5-9B hosted input tariff for 9B and the same conservative same-family proxy for the unlisted 2B; neither receives an automatic 100. Exact normalized comparison found no JevBench public task state or instruction in the 79,116-row public training projection.
- † OpenDecision: A zero-shot NLI classifier behind a TypeSafe-compatible server, not a trained decision model: it scores each option as an entailment hypothesis with ModernBERT-large-zeroshot-v2.0. Its choice path runs several NLI passes over the same state, which the reported token count does not include, so a per-token hosted price would be higher than the estimate here. Pre-registered for our CPU in v1.2.7, run on our GPU because the CPU was far too slow.
- † openJev Verdict 1.4: Same public weights as the earlier Verdict row, run through the author's fixed v1.4 engine. That engine auto-loads the calibrator for every option count, frames candidate labels as NLI sentences and uses a 512-token context budget. Run locally on our CPU, serially.
- † openJev Verdict: The openJev-verdict-2.0 Hugging Face repo ships no weights; its config is byte-identical to heman10x/rlcd-modernbert-151m, whose published weights we ran with the author's engine. The 'verdict2-base' checkpoint behind the README's numbers is not downloadable yet (Git LFS 404); we will run it once it is.
- † kev 0.5B: Self-hosted from the author's repository at commit 20fa626 through its native TypeSafe-compatible
/v1/systemoneserver, BF16 on an RTX 3090; measured serially from Sandy over the internet. This is the v0.1 release. - † GLiNER2 large: The large checkpoint of Fastino's earlier GLiNER2 family, same documented mapping as the GLiNER2 row: the question goes in front of the text and the probabilities are the model's own single-label softmax over the labels, read out in full. A general schema classifier, not a Jev rebuild.
- † smalljev semantic-v9: The public semantic-v9 LoRA and native heads over MiniCPM5-2B-Base, through the mapping frozen before the run. It has a typed Python contract but no TypeSafe-compatible HTTP route. The released training recipe explicitly hill-climbed against JevBench's public shape and source families; this allowed public benchmark-directed development is disclosed. Cost is $0.04/M measured input tokens, not free/100.
- † GLiNER2: A general schema classifier, not a Jev rebuild. The question goes in front of the text; the probabilities are GLiNER2's own single-label softmax over the labels, read out in full (mapping fixed before the run).
- † GLiNER2.5 multi: The multilingual GLiNER2.5 checkpoint (287M), same family and same documented mapping as the GLiNER2 row. JevBench items are English only, so its multilingual training is not exercised here.
- † GLiNER2.5 small: The small GLiNER2.5 checkpoint (74M), same family and same documented mapping as the GLiNER2 row: the question goes in front of the text and the probabilities are the model's own single-label softmax over the labels, read out in full. A general schema classifier, not a Jev rebuild.
- † Mixedbread mxbai-rerank-base-v2: Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.
- † BAAI bge-reranker-v2-m3: Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.
- † Alibaba GTE Reranker ModernBERT-base: Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.
- † Certo v1: The public Certo v1 checkpoint through the author's DecisionModel, serially on our rented GPU. The question instruction is prepended to the state because Certo exposes state + runtime options but no separate question field; the published 64-token state and 48-token option limits are unchanged. The model card says v1 does not yet transfer to arbitrary natural-language prose. Cost is an estimate from same-size hosted encoders times the checkpoint's retained input tokens, not free/100.




