Research for AI labs

OpenAgreements maintains a 50-state legal knowledge base, with practice guides, contract templates, and contract review checklists, and uses the knowledge base as a research environment for legal AI. In ordinary maintenance work, we observe and document concrete failure modes in contract drafting.

Our evaluation work replicates and measures concrete failure modes. The principal failure mode we document is what we call an institutional-knowledge leakage failure mode: some frontier models leak internal-facing analysis from the knowledge base into external-facing deliverables.

Our dataset work is meant to address the identified failure modes. We capture structured expert-correction data, including diffs, rationales, and the authorities considered and not applied. We capture these expert-correction data naturally as part of our Git-based maintenance workflow for the OpenAgreements legal knowledge base. The expert corrections frequently involve manually reconciling disparate and potentially conflicting authority in the knowledge base. Such reconciliation work is among the task types that an independent leaderboard has found to be the most difficult.

The page that was previously here has been separated into the evals and dataset write-ups above.