# Jev on German legal text

Engine test of 2026-09-20, published by Lexbeam Software. Jev 1.13 (TypeSafe, a System One model) against a
jury of three chat models on 2,596 yes-or-no judgments over 185 real items from German
authorities and federal courts.

## Results

- Agreement with the jury: **0.974** (precision 0.897, recall 0.902; 337 positives,
  304 caught, 35 false alarms, 33 missed).
- Median answer per item: **0.43 s** (p95 0.65 s). The most consistent juror, GLM 5.3 Flash, needs 13.67 s.
- Cost: **0.0077 USD per 1,000 judgments**, eleven times below that juror.
- Brier score 0.0228, against 0.113 for always guessing the base rate. Above 0.9, 198 of 200
  judgments are true by the jury.
- With the policy bands (fire at 0.80 and above, silent at 0.30 and below, review between):
  93.9% decided without a person, 6.1% (158) to review; 248 fired, of which
  10 the jury rejects; 2,190 silent, of which 14 were jury positives.
- German items: 0.975 agreement (n 2,375); English items: 0.959 (n 221).

## Speed and cost of every judge

| judge | median s per item | agreement with the other two | USD per 1,000 judgments |
|---|---:|---:|---:|
| Jev 1.13 | 0.43 |  | 0.0077 |
| GPT-5.6 Luna | 2.49 | 0.966 | 0.0549 |
| GLM 5.3 Flash | 13.67 | 0.985 | 0.0863 |
| gpt-oss-120b | 14.89 | 0.968 | 0.0238 |

## The fifteen questions (at threshold 0.5)

| question (German) | gloss | jury positives | caught | false alarms |
|---|---|---:|---:|---:|
| Gerichtsentscheidung | court decision | 63 | 61 | 0 |
| Datenschutz | data protection | 57 | 56 | 4 |
| Datenpanne / Sicherheit | breach or security | 38 | 29 | 2 |
| Handlungsbedarf | action required | 37 | 30 | 3 |
| KI | artificial intelligence | 34 | 33 | 1 |
| Termin / Personalie | event or appointment | 25 | 21 | 0 |
| EU-Digitalrecht | EU digital law | 20 | 17 | 3 |
| Öffentliche Stellen | public bodies | 18 | 15 | 6 |
| Betroffenenrechte | data subject rights | 15 | 13 | 2 |
| Gesundheit | health | 9 | 8 | 0 |
| Video | video | 6 | 6 | 6 |
| Bußgeld / Sanktion | fine or sanction | 5 | 5 | 4 |
| Werbung / Tracking | advertising or tracking | 5 | 5 | 1 |
| Beschäftigte | employees | 3 | 3 | 0 |
| Drittland | third country | 2 | 2 | 3 |

## Calibration bins

| Jev probability | judgments | jury true | share |
|---|---:|---:|---:|
| 0.0 to 0.1 | 1,927 | 1 | 0.00 |
| 0.1 to 0.2 | 185 | 4 | 0.02 |
| 0.2 to 0.3 | 78 | 9 | 0.12 |
| 0.3 to 0.4 | 40 | 9 | 0.23 |
| 0.4 to 0.5 | 28 | 10 | 0.36 |
| 0.5 to 0.6 | 34 | 21 | 0.62 |
| 0.6 to 0.7 | 30 | 22 | 0.73 |
| 0.7 to 0.8 | 27 | 23 | 0.85 |
| 0.8 to 0.9 | 48 | 40 | 0.83 |
| 0.9 to 1.0 | 200 | 198 | 0.99 |

## Method

- Sources: 200 items from eleven public feeds: three state data protection authorities (Baden-Württemberg, Bavaria, Saxony), the DSK, the EDPB, the BSI, four federal courts (BGH, BAG, BVerwG, BVerfG), the European Commission. 185 analysed.
- Fifteen German yes-or-no questions per item; contact data stripped before sending.
- Reference: the answer at least two of GPT-5.6 Luna, GLM 5.3 Flash and gpt-oss-120b give; 59 of
  2,655 contested judgments left out.
- Model: Jev 1.13 by TypeSafe, early access, through OpenRouter.

## Limits

- A model jury is not human ground truth. The human audit is next.
- One run on one day. Fifteen PDF-only items were skipped; one item was refused by a content filter.
- Rare questions rest on few positives.
- Aggregate results only.

## Next

Normlotse: which new decision concerns which client, for external data protection, information security
and AI officers. In preparation, early access at 79 EUR a month: https://normlotse.de/ . This was its engine test.

Measurements, not legal advice. Machine-readable: https://normlotse.de/results.json
