Jev on German legal text
2,596 yes-or-no judgments on real items from German authorities and federal courts, checked against a jury of three chat models. Engine test of 20.09.2026.
{ "agreement": 0.974, "judgments": 2596, "median_s_per_item": 0.43, "usd_per_1000_judgments": 0.0077, "decided_without_person": 0.939, "model": "Jev 1.13" }
Answered before the others finish reading.
Median seconds per item · scrolling runs the clock| Name | Status | Time | Waterfall |
|---|---|---|---|
| POST systemoneJev 1.13 | 200 | 0.43 s | |
| POST chat/completionsGPT-5.6 Luna | 200 | 2.49 s | |
| POST chat/completionsGLM 5.3 Flash | 200 | 13.67 s | |
| POST chat/completionsgpt-oss-120b | 200 | 14.89 s | |
Cost per 1,000 judgments
USD · agreement with the other two jurors- Jev 1.130.0077
- GPT-5.6 Luna0.05490.966
- GLM 5.3 Flash0.08630.985
- gpt-oss-120b0.02380.968
Probabilities you can route work on.
Above 0.9: 200 judgments, the jury calls 198 true. Brier score 0.023, against 0.113 for always guessing the base rate.
Drag the two lines. Between them, everything goes to a person.
{ "fire_at": 0.80, "fired": 248, "jury_rejects": 10, "silent_at": 0.30, "silent": 2190, "jury_positives_missed": 14, "to_a_person": 158 }
93.9% decided without a person.
Calibration bins. Counts while dragging come from these, in steps of 0.1; at 0.80 and 0.30 they match the run.
| Jev probability | judgments | jury true | share |
|---|---|---|---|
| 0.0 to 0.1 | 1,927 | 1 | 0.00 |
| 0.1 to 0.2 | 185 | 4 | 0.02 |
| 0.2 to 0.3 | 78 | 9 | 0.12 |
| 0.3 to 0.4 | 40 | 9 | 0.23 |
| 0.4 to 0.5 | 28 | 10 | 0.36 |
| 0.5 to 0.6 | 34 | 21 | 0.62 |
| 0.6 to 0.7 | 30 | 22 | 0.73 |
| 0.7 to 0.8 | 27 | 23 | 0.85 |
| 0.8 to 0.9 | 48 | 40 | 0.83 |
| 0.9 to 1.0 | 200 | 198 | 0.99 |
Fifteen questions, asked in German.
caughtmissedfalse alarm
| Question | One square per jury positive, then false alarms | Caught | False alarms |
|---|---|---|---|
| Gerichtsentscheidungcourt decision | 61/63 | 0 | |
| Datenschutzdata protection | 56/57 | 4 | |
| Datenpanne / Sicherheitbreach or security | 29/38 | 2 | |
| Handlungsbedarfaction required | 30/37 | 3 | |
| KIartificial intelligence | 33/34 | 1 | |
| Termin / Personalieevent or appointment | 21/25 | 0 | |
| EU-DigitalrechtEU digital law | 17/20 | 3 | |
| Öffentliche Stellenpublic bodies | 15/18 | 6 | |
| Betroffenenrechtedata subject rights | 13/15 | 2 | |
| Gesundheithealth | 8/9 | 0 | |
| Videovideo | 6/6 | 6 | |
| Bußgeld / Sanktionfine or sanction | 5/5 | 4 | |
| Werbung / Trackingadvertising or tracking | 5/5 | 1 | |
| Beschäftigteemployees | 3/3 | 0 | |
| Drittlandthird country | 2/2 | 3 |
Headers.
How it was measured, what it does not show, and where every figure lives.
Request
- x-sources
- 200 items from eleven public feeds: three state data protection authorities, the DSK, the EDPB, the BSI, four federal courts and the European Commission. 185 analysed.
- x-questions
- Fifteen German yes-or-no questions per item. Contact data stripped before sending.
- x-reference
- The answer at least two of three models give: GPT-5.6 Luna, GLM 5.3 Flash, gpt-oss-120b. 59 of 2,655 contested judgments left out.
- x-model
- Jev 1.13 by TypeSafe, early access, through OpenRouter.
- x-run-date
- 20.09.2026
Limits
- x-limit
- A model jury is not human ground truth. The human audit is next.
- x-limit
- One run on one day. Fifteen PDF-only items were skipped; one item was refused by a content filter.
- x-limit
- Rare questions rest on few positives: two for third countries, three for employees.
- x-limit
- The jurors disagree with each other too: their own agreement ranges from 0.966 to 0.985.
Initiator: where every figure lives
| Figure | Value | Initiator |
|---|---|---|
| agreement | 0.974 | results.json/results.agreement |
| precision · recall | 0.897 · 0.902 | results.json/results.precision |
| median answer per item | 0.43 s | results.json/results.latency_s.median |
| cost per 1,000 judgments | 0.0077 USD | results.json/results.usd_per_1000_judgments |
| decided without a person | 93.9% | results.json/results.policy_bands.decided_share_pct |
| German · English items | 0.975 · 0.959 | results.json/results.by_language |
| Brier score | 0.0228 | results.json/results.brier |
normlotse.next()
"Which change concerns which client."
For external data protection, security and AI officers. Not built yet; this was its engine test.