BUILD N BLOOM · EVIDENCE DISCLOSURE · 27 SEPTEMBER 2026

What our public-data component tests establish

On 26–27 September 2026, we recorded 100 model outputs across 80 distinct public-data tasks: 60 pilot outputs and 40 continuation outputs, including 20 retested tasks. These experiments examined individual model components using Sol 6 at low reasoning. They did not test an installed client workflow or measure human time saved, client capacity or financial results.

These results do not support an approval-ready accuracy claim. Correct quotation did not establish correct interpretation. Qualified adjudication and complete human-versus-assisted timing remain outstanding.

Results, with their denominators

TestObserved resultLimit
FinQA pilot13/15 numerical or boolean matchesTwo inconsistent references remain in the denominator. Scoring was exploratory; calculations and source support await qualified review.
ContractNLI batch pilot13/20 classification matches; 9/20 classification plus exact evidenceBalanced labels; reference matching is not professional acceptance.
Original-case isolated retest10/20 classification matches; 6/20 with exact evidenceRevised prompt and presentation did not improve the reference score. This is development evidence, not an independent sample.
Document-disjoint continuation12/20 classification matches; 4/20 with exact evidenceTwenty documents excluded from the selected original documents; public training contamination remains possible.
Literal quotation check40/40 continuation outputs passedChecks faithful copying only. It does not establish that the conclusion follows from the quotation.
Qasper and QMSum25 outputs returnedCorrectness, completeness and professional acceptance remain ungraded.

What we learned

The isolated evidence-check prompt did not improve results. Inspected disagreements involved consent exceptions and permitted-use rights: the model's strict reading differed from the benchmark labels. Those disagreements need qualified adjudication and the publisher's annotation guidance; neither the model output nor a reference label alone establishes professional advice.

These experiments did not exercise production retrieval, permissions, handoffs, recovery or approval controls. The failures do not prove those controls work. No client result or practitioner endorsement is claimed.

What makes a capacity benefit measurable

Compare equivalent accepted cases at an agreed quality standard. Count preparation, review, correction, coordination, exceptions, setup and oversight by role. Then check demand and remaining team bottlenecks. Financial value requires useful redeployment, a real avoided cash expense or additional accepted and collected contribution after costs. Hourly billing requires separate analysis.

Dataset publishers

FinQA · ContractNLI · Qasper · QMSum. Publisher links identify the datasets; they do not imply affiliation or endorsement. No source dataset or generated professional advice is redistributed here.

Read the Firm Capacity Blueprint scope · Start with your own capacity check