Twitter/X

Gabe Pereyra and EngramLab built a synthetic law firm, Calderwood & Harkness…

Brief

Gabe Pereyra and EngramLab are open-sourcing a synthetic law firm, Calderwood & Harkness, built to let models learn a firm's institutional knowledge without exposing confidential client data. The dataset includes work product from 46 clients and 266 matters (100M+ tokens), targets multi-hop reasoning across deals, finds generic agents costly and incomplete, and will be publicly released.

Why it matters

Gabe Pereyra and EngramLab built a synthetic law firm, Calderwood & Harkness, containing work product (versions, emails, drafts) from 46 clients and 266 client matters to avoid training on confidential client data.

Key details

  • The dataset exceeds 100M tokens and is designed to test multi-hop, cross-matter reasoning (e.g., "how did we structure the reps and warranties"), where generic agent approaches are found to be expensive and not exhaustive.
  • The team will open-source the synthetic firm dataset to separate the research problem from data-privacy and deployment challenges and to share further results soon.
Cleaned source text

We are working with @EngramLab to train models that can understand a law firm’s knowledge.

A huge source of differentiation for a law firm comes from all the past work the firm has done. When an associate is working on a new deal they have access to decades of similar deals the firm has done and a big part of the job is knowing how to effectively leverage that knowledge.

However, a large amount of this data is either client data or derived from client data which means you can’t naively train all of it into a single model because this would mean breaking confidentially. In order to separate the research problem from the data privacy and deployment problem we built a synthetic law firm.

@ItsJulioPereyra wrote an awesome article describing the dataset, tasks and approach to creating a synthetic version of the type of data you would expect to find at a large law firm. The firm has 46 clients, 266 client matters and each client matter contains work product, versions, emails, drafts, etc that you might find when searching the DMS of a firm.

This dataset allows us to explore how well agents can reason across large corpuses of knowledge (100M+ tokens) and require complicated multi-hop reasoning. For example, answering a query like “in similar deals, how did we structure the reps and warranties” requires the model to understand what makes deals similar and then search over all past deals to find examples.

We find that generic agent approaches for this type of reasoning are both expensive and not exhaustive and there is a lot of room for improvement. Excited to open source this dataset and also share more results soon

Julio Pereyra (@ItsJulioPereyra)

Article

LAB: Law Firm Knowledge

We are open-sourcing our next LAB expansion: a synthetic law firm, Calderwood & Harkness (“C&H” or “the firm”). Built in collaboration with @engramlab, the firm contains work product from more than

— https://nitter.net/ItsJulioPereyra/status/2085772997944803682#m