I’ve been saying for a while that we need better benchmarks for continual learning/memory — I think this is a really exciting step towards this!
We’ve built an entire synthetic law firm (𝐂𝐚𝐥𝐝𝐞𝐫𝐰𝐨𝐨𝐝 & 𝐇𝐚𝐫𝐤𝐧𝐞𝐬𝐬 ⚖️), modeled after real legal work – a 𝘱𝘦𝘳𝘴𝘪𝘴𝘵𝘦𝘯𝘵 work environment where agents do many tasks over the 𝘴𝘢𝘮𝘦 underlying context. Most agent benchmarks are built in this world where a model is dropped into an independent task instance and has to do its best / explore as quickly as possible (e.g. "here's a repo, fix this bug"). But we want to move towards a world where models can build on their experience. A senior engineer who has a mental "map of the codebase" can pinpoint issues much more quickly and effectively. In the same way, lawyers build up experience, learning what arguments succeed in front of regulators, playbooks for dealing with certain cases, etc.
It's the 𝘥𝘪𝘧𝘧𝘦𝘳𝘦𝘯𝘵𝘪𝘢𝘵𝘪𝘰𝘯 between law firms that makes the actual practice of law interesting. Every model (and law school student) knows a lot about law from studying the textbooks, but our goal is to build models that can compound and augment the rich, internal/private knowledge that makes firm A more successful than firm B.
now it's finally possible to understand and build towards that! there's a lot more work left to do, but we've been learning a lot from @ItsJulioPereyra @nikogrupen @gabepereyra to bring these models closer to real world tasks :)
Julio Pereyra (@ItsJulioPereyra)
Article
LAB: Law Firm Knowledge
We are open-sourcing our next LAB expansion: a synthetic law firm, Calderwood & Harkness (“C&H” or “the firm”). Built in collaboration with @engramlab, the firm contains work product from more than
— https://nitter.net/ItsJulioPereyra/status/2085772997944803682#m