AI Accuracy Proof

AI answers need evidence.
We build the test.

An outside, fixed-fee measurement of one AI workflow that reads documents, against a test set and pass marks frozen in writing before the test runs. Fixed fee. Independent. Principal-led (the founder does the work himself, no rotating team).

For teams building or buying document AI · legal technology · insurance · model risk and AI governance

The method

How the test works

AI accuracy

Does your AI workflow meet a bar set before the test?

An outside, fixed-fee measurement of one agreed AI workflow that answers questions from documents, for example finding a named clause in a contract. Before any run, the test set, the pass marks, and a set of planted wrong answers are written down, so the scoring is checked against cases with a known result.

The written scope sets what you receive: every item and score, an error map by category and pattern, a signed report of what held and what could not be established, and a standalone program your team can run offline to recompute the reported scores. The fee does not change with what the measurement finds.

1. Freeze

Test items, pass marks, and planted wrong answers locked first.

2. Run

Your workflow answers every item in the frozen set.

3. Map

Misses sorted by category and pattern, with the limits stated.

4. Recount

A standalone program, delivered with the report, recomputes each reported score.

What has been run. This method has been run on CUAD, a public set of real contracts drawn from SEC filings, with the pass marks fixed first and planted wrong answers included. A public write-up will be posted on this page after it passes verification.

Scope. Each test covers the agreed workflow and its frozen test set, and the signed report states exactly what was measured. Anything beyond that is set in writing before work starts.

Founding engagements. We are scheduling a small number of founding engagements, and we welcome teams willing to have the method run on their own documents. Each start date is set in writing at scoping, after our current verification round.

Reserve a founding engagement
Common questions

Answered before you ask

How is it priced?
One fixed fee for one workflow and one signed report, set in writing before work starts. The fee does not change with what the measurement finds.
What do you need from us?
Access to the workflow (an endpoint or its recorded answers); your documents, questions, and right answers, or a person who agrees on them with us; a decision owner; and data-use terms. How your files are handled is set in the written scope before anything is shared, so please do not email confidential documents with a first message.
What have you measured so far?
Four AI models measured on questions drawn from the public CUAD contract set, described above on this page, with the pass marks fixed before each run. A public write-up will be posted on this page after it passes verification.
When is it delivered?
The schedule is set in the written scope before work starts. The ten-business-day line on this site applies to the smaller evaluation engagements.

An AI workflow to measure?

Email the document type, the question the workflow answers, and who owns the right answers. You get an answer from the analyst, not a form queue.