The Prove-it Lab
The scan estimates how much of a role's work is within reach of AI models. Before committing people and money, the next question is whether it works on your work. The Lab answers it independently: a current frontier model, run once on your own task instances, graded against acceptance criteria you agreed before it saw a single one.
How it runs
- Choose the tasks. A short list from your scan, usually where the hours are largest, and 20 to 50 real instances of each, with names and identifying details removed by you before anything is shared.
- Write the criteria first. For each task, what an acceptable result must do, written with you and dated before any attempt. A criterion added afterwards is marked as added afterwards.
- One attempt, unedited. The model does each instance once, with the brief and the material, and its output is kept as it came.
- A verdict per task: passes, passes with named checks, or fails, with the grade against each criterion, what went wrong, where a person has to review, and what running it would take (systems, data access, controls, licensed sign-off).
- The payback beside it: the hours and cost within reach for the roles behind those tasks, from your scan.
- A retest at each new model release or each quarter, on the same instances and criteria, so you see what moved.
A verdict is one model, one pass, on one date, and every page of the report says so. The Lab tests tasks, never people: the report names roles and tasks, never who holds them. It never says a job can be done by AI.
Your material
- You remove names and identifying details before you share anything. We check again and return anything that still identifies a person.
- Your material is used for your engagement only. It is never used to train any model, never shown to anyone outside the engagement, and never published.
- It is kept in an engagement folder outside every code repository and database we run, and deleted on a date stated in the agreement, at the latest 90 days after the report. We confirm the deletion in writing.
- The model runs through an AI provider's service whose terms do not allow it to train on what is sent. The report names the model and the date.
- The public site never shows any of it. If you allow it, the report may later be cited only as counts with every identifying detail removed.
What you get
A report (PDF and slides) with each task's verdict, the unedited outputs, the grades, the error analysis and the payback, and a private page for the length of the engagement. The method is the same as the Work Bench, where common tasks are tested the same way in public.
Taking part
The first engagements are by arrangement, for a deal team before a bid or a company in its first hundred days: write to info@stratussc.com with the roles and tasks you want tested. The agreement is one page and says the deletion date before anything is shared.