# Benchmark comparison

Latest comparison · 2026-09-27 · Scores (%)

| Model | JevBench | Kev | OpenJev text | Nimble | VitaminC | MASSIVE-en | Average |
|---|---:|---:|---:|---:|---:|---:|---:|
| JEV-27B | 88.70 | 83.75 | 73.89 | 92.91 | 77.46 | 87.71 | 84.07 |
| TypeSafe Jev | 87.18 | 85.52 | 72.96 | 91.84 | 78.46 | 87.14 | 83.85 |
| NeoHorse-Jev-4B | 75.73 | 81.92 | 58.74 | 87.23 | 77.13 | 85.43 | 77.70 |
| Open-Jev-9B | 77.13 | 77.87 | 65.39 | 80.50 | 68.28 | 84.86 | 75.67 |
| Kev-4B | 73.71 | 81.47 | 54.75 | 73.40 | 76.46 | 85.71 | 74.25 |
| Laya English | 55.82 | 61.30 | 40.07 | 45.04 | 78.63 | 68.57 | 58.24 |

JEV-27B and TypeSafe Jev are AutoTrust AI evaluation results from the current merged comparison dated September 27, 2026.

NeoHorse-Jev, Open-Jev, Kev and Laya English use TokenRhythm’s published “Text Decision Benchmarks” table in the NeoHorse-Jev-4B model card on Hugging Face. The source is pinned to revision `b50e043e22e0e41e7fc0c244e4daa707b8124930`. The chart uses the JevBench, Kev, OpenJev text, Nimble, VitaminC and MASSIVE-en columns from that table.

The 84.07% highlighted value is the unweighted mean of JEV-27B’s six benchmark scores.

[Published baseline table](https://huggingface.co/TokenRhythm/NeoHorse-Jev-4B/blob/b50e043e22e0e41e7fc0c244e4daa707b8124930/README.md) · [Chart data](benchmark-comparison.json)

[Download JEV-27B](https://huggingface.co/autotrust/JEV-27B)

Comparison generated: 2026-09-26T18:18:13.975765+00:00
Source SHA-256: `c4f4702c3f98fe588905303008a9d44ea73bc811830f0c03c7510ed02a723ccd`
