BUILD N BLOOM · EVIDENCE DISCLOSURE · 27 SEPTEMBER 2026
What our public-data component tests establish
On 26–27 September 2026, we recorded 100 model outputs across 80 distinct public-data tasks: 60 pilot outputs and 40 continuation outputs, including 20 retested tasks. These experiments examined individual model components using Sol 6 at low reasoning. They did not test an installed client workflow or measure human time saved, client capacity or financial results.
These results do not support an approval-ready accuracy claim. Correct quotation did not establish correct interpretation. Qualified adjudication and complete human-versus-assisted timing remain outstanding.
Results, with their denominators
| Test | Observed result | Limit |
|---|---|---|
| FinQA pilot | 13/15 numerical or boolean matches | Two inconsistent references remain in the denominator. Scoring was exploratory; calculations and source support await qualified review. |
| ContractNLI batch pilot | 13/20 classification matches; 9/20 classification plus exact evidence | Balanced labels; reference matching is not professional acceptance. |
| Original-case isolated retest | 10/20 classification matches; 6/20 with exact evidence | Revised prompt and presentation did not improve the reference score. This is development evidence, not an independent sample. |
| Document-disjoint continuation | 12/20 classification matches; 4/20 with exact evidence | Twenty documents excluded from the selected original documents; public training contamination remains possible. |
| Literal quotation check | 40/40 continuation outputs passed | Checks faithful copying only. It does not establish that the conclusion follows from the quotation. |
| Qasper and QMSum | 25 outputs returned | Correctness, completeness and professional acceptance remain ungraded. |
What we learned
The isolated evidence-check prompt did not improve results. Inspected disagreements involved consent exceptions and permitted-use rights: the model's strict reading differed from the benchmark labels. Those disagreements need qualified adjudication and the publisher's annotation guidance; neither the model output nor a reference label alone establishes professional advice.
These experiments did not exercise production retrieval, permissions, handoffs, recovery or approval controls. The failures do not prove those controls work. No client result or practitioner endorsement is claimed.
What makes a capacity benefit measurable
Compare equivalent accepted cases at an agreed quality standard. Count preparation, review, correction, coordination, exceptions, setup and oversight by role. Then check demand and remaining team bottlenecks. Financial value requires useful redeployment, a real avoided cash expense or additional accepted and collected contribution after costs. Hourly billing requires separate analysis.
Dataset publishers
FinQA · ContractNLI · Qasper · QMSum. Publisher links identify the datasets; they do not imply affiliation or endorsement. No source dataset or generated professional advice is redistributed here.
Read the Firm Capacity Blueprint scope · Start with your own capacity check