Artificial AnalysisT1·68 pts
Harvey Releases LAB-AA v1.1 Benchmark: Introducing Hallucination-Gated All-Pass Rate for Legal Agents
Original: Announcing Harvey LAB-AA v1.1: adding hallucination checks to raise the bar for agentic legal work
- Harvey LAB-AA v1.1 introduces the 'Hallucination-Gated All-Pass Rate', crediting tasks only if they are hallucination-free and meet all criteria.
- The benchmark comprises 120 private legal tasks spanning specialized areas such as M&A, tax, and litigation.
- A three-judge panel mechanism is employed to independently grade each rubric criterion, enhancing evaluation consistency and accuracy.