AutoTrust

The JEV-27B experience

JEV-27B.
See it in action.

Explore interactive demos and ten recorded workflows across games, browser tasks, and code.

Recorded experiments with self-hosted JEV-27B. Edited highlights, with on-screen text and instrumental music.

Download reel ↓

Across six benchmarks

Six-model comparison · September 27, 2026

Six models compared across JevBench, Kev, OpenJev text, Nimble, VitaminC, and MASSIVE-en. JEV-27B is highlighted in gold with an 84.07% six-benchmark average. Each panel uses its labeled axis range.
Open full size ↗ · Download SVG ↓

JevBench, Kev, and OpenJev text use averages across task families, splits, and components respectively. Nimble, VitaminC, and MASSIVE-en report per-example accuracy. These results use separate evaluation settings from the scorecard below. Read the evaluation notes ↗

A closer look

The model scorecard

Open full size ↗
AutoTrust JEV-27B model scorecard. System 1 fidelity: mean KL 0.017, yes/no AUROC 0.995, top-1 agreement 90.5%, rating error 0.098. Calibration error 0.0009, option-shuffle flips 2.9%, unseen-task KL 0.104. System 2 HumanEval pass at 1: 78.0%. See full-size image for methods and qualifications.
Provided by AutoTrust. Metrics, evaluation settings, and comparison notes are included in the scorecard.

Your turn to explore.

Try an interactive workflow, or take a closer look at the model.

Reel soundtrack: “Tech Live” by Kevin MacLeod (incompetech.com), licensed under Creative Commons Attribution 4.0. Edited excerpt. Third-party notices.