We are working with @EngramLab to train models that can understand a law firm’s knowledge.
A huge source of differentiation for a law firm comes from all the past work the firm has done. When an associate is working on a new deal they have access to decades of similar deals the firm has done and a big part of the job is knowing how to effectively leverage that knowledge.
However, a large amount of this data is either client data or derived from client data which means you can’t naively train all of it into a single model because this would mean breaking confidentially. In order to separate the research problem from the data privacy and deployment problem we built a synthetic law firm.
@ItsJulioPereyra wrote an awesome article describing the dataset, tasks and approach to creating a synthetic version of the type of data you would expect to find at a large law firm. The firm has 46 clients, 266 client matters and each client matter contains work product, versions, emails, drafts, etc that you might find when searching the DMS of a firm.
This dataset allows us to explore how well agents can reason across large corpuses of knowledge (100M+ tokens) and require complicated multi-hop reasoning. For example, answering a query like “in similar deals, how did we structure the reps and warranties” requires the model to understand what makes deals similar and then search over all past deals to find examples.
We find that generic agent approaches for this type of reasoning are both expensive and not exhaustive and there is a lot of room for improvement. Excited to open source this dataset and also share more results soon
Julio Pereyra (@ItsJulioPereyra)
Article
LAB: Law Firm Knowledge
We are open-sourcing our next LAB expansion: a synthetic law firm, Calderwood & Harkness (“C&H” or “the firm”). Built in collaboration with @engramlab, the firm contains work product from more than
— https://nitter.net/ItsJulioPereyra/status/2085772997944803682#m