An outside, fixed-fee measurement of one AI workflow that reads documents, against a test set and pass marks frozen in writing before the test runs. Fixed fee. Independent. Principal-led (the founder does the work himself, no rotating team).
For teams building or buying document AI · legal technology · insurance · model risk and AI governance
Does your AI workflow meet a bar set before the test?
An outside, fixed-fee measurement of one agreed AI workflow that answers questions from documents, for example finding a named clause in a contract. Before any run, the test set, the pass marks, and a set of planted wrong answers are written down, so the scoring is checked against cases with a known result.
The written scope sets what you receive: every item and score, an error map by category and pattern, a signed report of what held and what could not be established, and a standalone program your team can run offline to recompute the reported scores. The fee does not change with what the measurement finds.
Test items, pass marks, and planted wrong answers locked first.
Your workflow answers every item in the frozen set.
Misses sorted by category and pattern, with the limits stated.
A standalone program, delivered with the report, recomputes each reported score.
What has been run. This method has been run on CUAD, a public set of real contracts drawn from SEC filings, with the pass marks fixed first and planted wrong answers included. A public write-up will be posted on this page after it passes verification.
Scope. Each test covers the agreed workflow and its frozen test set, and the signed report states exactly what was measured. Anything beyond that is set in writing before work starts.
Founding engagements. We are scheduling a small number of founding engagements, and we welcome teams willing to have the method run on their own documents. Each start date is set in writing at scoping, after our current verification round.
Reserve a founding engagementEmail the document type, the question the workflow answers, and who owns the right answers. You get an answer from the analyst, not a form queue.