Clerkbench

The Luxembourg legal-agent benchmark

Every run is scored against a pre-registered rubric anchored on the corpus text, with mechanical citation verification, severity gates, and the full audit trail published. No preference judging, no reference-answer similarity — and no penalty is ever awarded on an unverified claim of fabrication.

LB-001Marchand c/ Rivastone Cleantech S.A. — licenciement avec effet immédiat
case v1.0.0 · rubric v1.0.0 · framework v1.0.0
ConfigurationQuality /100Critical-failure rateVerified citationsCost / runLatencyInput tokensnDate
lite2
vertex-trial:gemini-3.6-flash
90.7± 2.10%100%$0.30272s233k32026-08-03Full run →

Quality (0–100) aggregates eleven substance axes and is capped by severity gates: any critical failure caps a rep at 39. The critical-failure rate is the share of repeated runs with at least one critical event — a reliability measure, reported separately and never averaged into Quality. Cost, latency and tokens are reported alongside, never folded in: a fast wrong answer must not outrank a slow right one. * = price-table entry missing, cost estimated or unavailable.