Twitter/X

Harvey open-sourced the Legal Agent Benchmark

Brief

Harvey Research centralizes Harvey's open research effort and technical strategy: they published the Legal Agent Benchmark (1,200 tasks across 24+ practice areas) and a research home at harvey.ai/research. Gabe Pereyra frames a legal foundation-model series inspired by Cursor's Composer with two concrete aims—deliver affordable, secure frontier intelligence across products and enable law firms to build and own specialized models—and designs an agentic system to manage tools, sub-agents, and human or frontier-model advisors for long-horizon matters that take months and dozens of associates. Harvey highlights partner-driven technical advances (Baseten compaction for large data rooms; FireworksAI using frontier-as-advisor to match frontier performance; AppliedCompute cost/performance gains on review tables; TrajectoryLabs & NVIDIA on sovereign continual learning). Harvey Labs, led by Niko Grupen and Julio Pereyra, is hiring and open to acquisitions to scale post-training, agent, and data efforts.

Why it matters

Harvey open-sourced the Legal Agent Benchmark: the largest benchmark for long-horizon legal work with 1,200 tasks spanning 24+ practice areas and published a research hub at harvey.ai/research.

Key details

  • Harvey's model strategy (announced by Gabe Pereyra) launches a legal foundation model series inspired by Cursor's Composer with two explicit goals: (1) deliver frontier intelligence across Harvey's product surface affordably and securely, and (2) enable law firms to build and own specialized models.
  • The model/agent design targets complex client matters that span months and involve dozens of associates; the agentic system will control legal tech tools, sub-agents, and consult frontier models or humans—functioning like a senior associate.
  • Harvey reports collaborative technical results with partners: Baseten (novel compaction strategies for large data rooms), FireworksAI (matching frontier by using frontier as an advisor), AppliedCompute (improving performance and reducing cost on large-scale review tables), and TrajectoryLabs & NVIDIA (sovereign continual learning).
  • Harvey Labs is the internal research group run by Niko Grupen (multi-agent RL at Google Brain) and Julio Pereyra (clerked and worked in BigLaw); the group is scaling, hiring across post-training/agent/data stacks, and open to acquiring teams/neolabs.
Source evidence

Introducing Harvey Research:

We've shared our model strategy.

We've open-sourced Legal Agent Benchmark, the largest benchmark for long-horizon legal work spanning 1,200 tasks across 24+ practice areas.

And we've collaborated on research with leading neolabs and inference providers like @baseten, @trajectorylabs, @LangChain, @FireworksAI_HQ, @appliedcompute, and @EngramLab.

Now we have a home base for it.

Live at: harvey.ai/research

Link

Harvey Research

Harvey is pushing the frontier of legal intelligence through open research. Our research agenda spans evaluation and benchmarking, post-training, harness optimization, verification, and agent...
harvey.ai

Gabe Pereyra (@gabepereyra)

Model strategy for @harvey:

We are working on the first model in our legal foundation model series, inspired by @cursor_ai's Composer. Two goals:

  1. Allow us to serve frontier intelligence across our product surface areas at an affordable price and a strong security posture.
  2. Create the foundations for law firms to build their own specialized models and own their own intelligence.

The model series will focus on complex client matters that span months and take dozens of associates. The agentic system will learn to control legal tech tools, sub agents and ask for help from frontier models or human partners, much like a senior associate.

We’ve open sourced benchmarks for evaluating our initial post training work that represents work done by associates and in-house lawyers. We are scaling these significantly using synthetic and human pipelines as well as building private evals for firms.

Open sourcing this data has allowed us to quickly validate the feasibility of post training open weight models for legal work. With our research partners we’ve already shown promising results post training open source models to approach frontier performance:

  1. @baseten - novel compaction strategies for analyzing large data rooms.
  2. @FireworksAI_HQ - matching frontier performance by using frontier as an advisor.
  3. @appliedcompute - improving performance and reducing cost of large scale review tables.
  4. @trajectorylabs & @nvidia - sovereign continual learning over client matters.

We plan to continue to invest heavily in working with research partners and open sourcing our data, models and research as much as possible. We believe open research in legal will be important to building trust in the frontier ecosystem.

We are also scaling our research team. Harvey Labs is our internal research group, responsible for pushing the frontier of legal intelligence and working closely with labs, research partners, and academia to bring the frontier of agent research into Harvey.

Labs is run by @nikogrupen and @ItsJulioPereyra - Niko worked on multi-agent RL at Google Brain and Julio clerked and worked in BigLaw. We believe this pairing is crucial for building frontier legal AI systems. Together they have already made significant progress in scaling our data and training efforts.

The long term goal of Harvey Labs is to contribute to the research and infrastructure required for the legal industry to create a frontier ecosystem. We believe that the best version of legal super intelligence is one where each law firm, enterprise and government owns their own specialized version.

We are hiring for Harvey Labs across the post training, agent and data stack and open to acquiring talented teams / neolabs in this space. If interested please DM me.

— https://nitter.net/gabepereyra/status/2067324200801452105#m