Research for AI labs
OpenAgreements maintains a 50-state legal knowledge base, with practice guides, contract templates, and contract review checklists, and uses the knowledge base as a research environment for legal AI. In ordinary maintenance work, we observe and document concrete failure modes in contract drafting.
Our evaluation work replicates and measures concrete failure modes. The principal failure mode we document is what we call an institutional-knowledge leakage failure mode: some frontier models leak internal-facing analysis from the knowledge base into external-facing deliverables.
Our dataset work is meant to address the identified failure modes. We capture structured expert-correction data, including diffs, rationales, and the authorities considered and not applied. We capture these expert-correction data naturally as part of our Git-based maintenance workflow for the OpenAgreements legal knowledge base. The expert corrections frequently involve manually reconciling disparate and potentially conflicting authority in the knowledge base. Such reconciliation work is among the task types that an independent leaderboard has found to be the most difficult.
Steven Obiajulu · OpenAgreementsEvals: Institutional-knowledge leakage in frontier AI models on realistic legal draftingAn evaluation of three frontier models on multi-state contract drafting inside a 367-document institutional knowledge base, with per-judge records, human adjudication, and upstream LAB harness contributions.
Steven Obiajulu · OpenAgreementsEvals: Role labels: where party-favoritism lives in LLM contract fairness ratingsAn evaluation of four frontier models on 105 mirror-constructed clause pairs, the same text with the party roles swapped, isolating whether fairness-rating tilt comes from the contract text or from the role labels.
Steven Obiajulu · OpenAgreementsDataset: Expert-correction traces from maintaining an open legal knowledge baseStructured, attributable expert corrections of AI legal drafting, captured from a real maintenance workflow: machine-recorded diffs, dictated rationale, primary-source checks, and the authorities considered and not applied.
The page that was previously here has been separated into the evals and dataset write-ups above.