Normlotse Talk to Lexbeam

Jev on German legal text

2,596 yes-or-no judgments on real items from German authorities and federal courts, checked against a jury of three chat models. Engine test of 20.09.2026.

{
  "agreement":
0.974,  "judgments": 2596,
  "median_s_per_item": 0.43,
  "usd_per_1000_judgments": 0.0077,
  "decided_without_person": 0.939,
  "model": "Jev 1.13"
}

Answered before the others finish reading.

Median seconds per item · scrolling runs the clock
NameStatusTimeWaterfall
POST systemoneJev 1.132000.43 s
POST chat/completionsGPT-5.6 Luna2002.49 s
POST chat/completionsGLM 5.3 Flash20013.67 s
POST chat/completionsgpt-oss-120b20014.89 s

Cost per 1,000 judgments

USD · agreement with the other two jurors
  • Jev 1.130.0077
  • GPT-5.6 Luna0.05490.966
  • GLM 5.3 Flash0.08630.985
  • gpt-oss-120b0.02380.968

Probabilities you can route work on.

Above 0.9: 200 judgments, the jury calls 198 true. Brier score 0.023, against 0.113 for always guessing the base rate.

Drag the two lines. Between them, everything goes to a person.

{
  "fire_at": 0.80,
  "fired": 248,
  "jury_rejects": 10,
  "silent_at": 0.30,
  "silent": 2190,
  "jury_positives_missed": 14,
  "to_a_person": 158
}

93.9% decided without a person.

Calibration bins. Counts while dragging come from these, in steps of 0.1; at 0.80 and 0.30 they match the run.
Calibration bins
Jev probabilityjudgmentsjury trueshare
0.0 to 0.11,92710.00
0.1 to 0.218540.02
0.2 to 0.37890.12
0.3 to 0.44090.23
0.4 to 0.528100.36
0.5 to 0.634210.62
0.6 to 0.730220.73
0.7 to 0.827230.85
0.8 to 0.948400.83
0.9 to 1.02001980.99

Fifteen questions, asked in German.

caughtmissedfalse alarm
QuestionOne square per jury positive, then false alarmsCaughtFalse alarms
Gerichtsentscheidungcourt decision61/630
Datenschutzdata protection56/574
Datenpanne / Sicherheitbreach or security29/382
Handlungsbedarfaction required30/373
KIartificial intelligence33/341
Termin / Personalieevent or appointment21/250
EU-DigitalrechtEU digital law17/203
Öffentliche Stellenpublic bodies15/186
Betroffenenrechtedata subject rights13/152
Gesundheithealth8/90
Videovideo6/66
Bußgeld / Sanktionfine or sanction5/54
Werbung / Trackingadvertising or tracking5/51
Beschäftigteemployees3/30
Drittlandthird country2/23

Headers.

How it was measured, what it does not show, and where every figure lives.

Request
x-sources
200 items from eleven public feeds: three state data protection authorities, the DSK, the EDPB, the BSI, four federal courts and the European Commission. 185 analysed.
x-questions
Fifteen German yes-or-no questions per item. Contact data stripped before sending.
x-reference
The answer at least two of three models give: GPT-5.6 Luna, GLM 5.3 Flash, gpt-oss-120b. 59 of 2,655 contested judgments left out.
x-model
Jev 1.13 by TypeSafe, early access, through OpenRouter.
x-run-date
20.09.2026
Limits
x-limit
A model jury is not human ground truth. The human audit is next.
x-limit
One run on one day. Fifteen PDF-only items were skipped; one item was refused by a content filter.
x-limit
Rare questions rest on few positives: two for third countries, three for employees.
x-limit
The jurors disagree with each other too: their own agreement ranges from 0.966 to 0.985.
Initiator: where every figure lives
FigureValueInitiator
agreement0.974results.json/results.agreement
precision · recall0.897 · 0.902results.json/results.precision
median answer per item0.43 sresults.json/results.latency_s.median
cost per 1,000 judgments0.0077 USDresults.json/results.usd_per_1000_judgments
decided without a person93.9%results.json/results.policy_bands.decided_share_pct
German · English items0.975 · 0.959results.json/results.by_language
Brier score0.0228results.json/results.brier

normlotse.next()

"Which change concerns which client."

For external data protection, security and AI officers. Not built yet; this was its engine test.