A New Bar for Legal AI on Harvey LAB

Rekursor's rubric-blind system achieved a 34% all-pass rate on a held-out sample of 50 public Harvey LAB tasks designed to represent the benchmark's practice areas and task types.

Amit Tandon
Founder, Rekursor5 min read

Legal AI is often measured by how much it gets right. Law firms also need to know whether it completed the entire assignment correctly.

Harvey's Legal Agent Benchmark (LAB) measures that harder standard. A task achieves all-pass only when the deliverable satisfies every criterion in a detailed rubric covering the required facts, citations, structure and analytical moves.

On a held-out sample of 50 tasks from the public benchmark, Rekursor's system, built around GLM-5.2, passed 94.9% of individual criteria. It satisfied every criterion on 17 of 50 tasks, producing a 34% all-pass rate.

To our knowledge, as of July 21, 2026, that is the highest all-pass rate yet reported across published LAB-derived evaluations.

The evaluation was rubric-blind. During generation, the system received the task instructions and supplied source materials, but not the grading rubric, answer key, prior score, grader feedback or criterion-level targets.

Results at a glance

Published all-pass results across LAB-derived evaluations

Published results from different task cohorts, graders and harnesses—not a controlled head-to-head comparison.

Sources: Artificial Analysis, Harvey LAB-AA leaderboard (Kimi K3: 26.7%; Claude Fable 5: 14.2%; Fable configuration used adaptive reasoning at maximum effort with an Opus 4.8 fallback): https://artificialanalysis.ai/evaluations/harvey-lab-aa · Harvey and Applied Compute, “Training a Legal Agent With Applied Compute” (12.6%): https://www.harvey.ai/blog/training-a-legal-agent-with-applied-compute

MeasureRekursor result
Evaluated tasks50
All-pass17/50
All-pass rate34.0%
Individual criteria passed2,029/2,137
Criterion pass rate94.9%
Foundation modelGLM-5.2
GraderGPT-5.4-mini
Rubric access during generationNone
Held out from development and task-specific tuning of the reported configurationAll 50
Task-specific tuning during scored cohortsNone

The available comparison points—Kimi K3's 26.7% and Claude Fable 5's 14.2% on Artificial Analysis's private LAB-AA evaluation, and Harvey and Applied Compute's 12.6% on Harvey's 180-task held-out evaluation—used different task cohorts, graders and harnesses. Because all-pass requires every criterion to pass, a single borderline criterion can change the outcome of an entire task. These figures provide important published context, but not a controlled head-to-head ranking.

These are system results, not a new claim about the standalone GLM-5.2 model. The result measures what becomes possible when a capable foundation model operates inside the purpose-built Rekursor system for complex professional work.

The 50 tasks were selected through a stratified process designed to reflect the practice-area and task-type distribution of the public LAB corpus. Selection was completed before the scored evaluation and was not based on expected or observed system performance. The individual tasks were not selected through a simple random draw.

Held-out means that none of the 50 tasks, outputs, grades or criterion-level feedback was used to develop, tune or select the reported Rekursor configuration. Because LAB is public, we cannot rule out prior exposure of the underlying foundation model to benchmark materials.

Harvey's published work treats grading as part of the evaluation system, not an afterthought. Its Applied Compute release analyzed grader alignment, used a GPT-5 Mini-family grader and evaluated work against the rubric criteria that define successful completion.

We used the LAB criterion-level grading formulation and exact Harvey LAB rubric_criterion prompt, with one independent grader call for each criterion. All outputs were frozen before grading, so grader decisions could not influence the scored work.

Harvey's Applied Compute evaluation batched four criteria in each GPT-5 Mini call. Rekursor instead used one criterion per call. A rubric-blind system has to identify what the work requires without being shown the evaluator's targets—the deployment condition that matters in real legal work.

The result was achieved within a minimal dollars-per-task operating envelope, not through a massive-compute or high-volume sampling campaign. Exact costs and inference topology remain proprietary.

The model is only one layer of the system

The model market is moving quickly. A model that leads one month may be displaced the next, and legal performance remains uneven across practice areas and task types.

This creates a strategic question for firms and AI companies alike: where should durable capability live?

Better weights, longer context and retrieval all matter. But dependable professional performance also depends on the system in which the model operates.

The 34% result is evidence for the agentic system layer. Rekursor is designed to operate around powerful models and extract more reliable professional work from them. The foundation model remains essential; the surrounding intelligence determines how much of its latent capability reaches the deliverable.

This distinction matters commercially. A law firm should not have to rebuild its AI strategy every time the model leaderboard changes. The durable asset is the firm's controlled intelligence layer: the systems that connect models to matter context, institutional standards, evaluated work and human oversight.

A note on methodology and disclosure

We are publishing the measured result, evaluation method and boundaries of the claim while keeping the system architecture and implementation proprietary. Additional validation materials can be reviewed privately with prospective partners.

From a public benchmark to firm-owned intelligence

Rekursor is opening a limited number of matched evaluations and pilot engagements with law firms and legal AI teams. Prospective partners can evaluate the system on their own tasks and standards or run a controlled head-to-head comparison on a common harness. We would love to talk.

A public benchmark is a proving ground, not the prize. LAB's rubrics stand in for something every firm already owns: partners' standards, reviewed work product and the accumulated judgment of a practice. Rekursor's architecture is designed to operate against those firm-owned assets under controlled governance. The 34% reported here was achieved without firm-specific knowledge; a firm's standards, reviewed work and accumulated experience are a potentially important source of further performance gains.


Rekursor develops proprietary AI infrastructure for continual learning and high-reliability professional work, with legal services as its first deployment domain. Supporting evidence is available through private technical diligence.