The Luxembourg legal-agent benchmark
Every run is scored against a pre-registered rubric anchored on the corpus text, with mechanical citation verification, severity gates, and the full audit trail published. No preference judging, no reference-answer similarity — and no penalty is ever awarded on an unverified claim of fabrication.
| Configuration | Quality /100 | Critical-failure rate | Verified citations | Cost / run | Latency | Input tokens | n | Date | |
|---|---|---|---|---|---|---|---|---|---|
lite2 vertex-trial:gemini-3.6-flash | 90.7± 2.1 | 0% | 100% | $0.30 | 272s | 233k | 3 | 2026-08-03 | Full run → |
Quality (0–100) aggregates eleven substance axes and is capped by severity gates: any critical failure caps a rep at 39. The critical-failure rate is the share of repeated runs with at least one critical event — a reliability measure, reported separately and never averaged into Quality. Cost, latency and tokens are reported alongside, never folded in: a fast wrong answer must not outrank a slow right one. * = price-table entry missing, cost estimated or unavailable.