Today for AI
HOT RADAR
llmHEAT 10.7°

AI CLUSTERED EVENT · 10/8/2026

Harvey Releases LAB-AA v1.1 Benchmark: Introducing Hallucination-Gated All-Pass Rate for Legal Agents

1 reports archived1 independent sourcesupdated 10/8/2026, 08:00:00
Synthesis & Latest Updates
1 Sources Cross-Validated

Harvey and Artificial Analysis have released Harvey LAB-AA v1.1, updating the scoring methodology for AI agents performing real-world legal work. The new version introduces the 'Hallucination-Gated All-Pass Rate' as a headline metric, requiring deliverables to satisfy all rubric criteria while containing no material hallucinations. Based on 120 private legal tasks covering areas like M&A and litigation, the benchmark uses a three-judge panel to grade each criterion, aiming to raise reliability standards for agentic legal applications.

LATEST/Harvey and Artificial Analysis have released Harvey LAB-AA v1.1, updating the scoring methodology for AI agents performing real-world legal work. The new version introduces the 'Hallucination-Gated All-Pass Rate' as a headline metric, requiring deliverables to satisfy all rubric criteria while containing no material hallucinations. Based on 120 private legal tasks covering areas like M&A and litigation, the benchmark uses a three-judge panel to grade each criterion, aiming to raise reliability standards for agentic legal applications.

TIMELINECoverage timeline

Total 1 reports · Latest first
  1. Artificial AnalysisT1·68 pts

    Harvey Releases LAB-AA v1.1 Benchmark: Introducing Hallucination-Gated All-Pass Rate for Legal Agents

    Original: Announcing Harvey LAB-AA v1.1: adding hallucination checks to raise the bar for agentic legal work

    • Harvey LAB-AA v1.1 introduces the 'Hallucination-Gated All-Pass Rate', crediting tasks only if they are hallucination-free and meet all criteria.
    • The benchmark comprises 120 private legal tasks spanning specialized areas such as M&A, tax, and litigation.
    • A three-judge panel mechanism is employed to independently grade each rubric criterion, enhancing evaluation consistency and accuracy.