Sheet 07 · ART-09
Pilot measurement plan
Defines what will be measured, against what baseline, by whom, and how good the evidence will be — agreed before the pilot starts so the result cannot be reverse-engineered afterward.
Artifact details
- Format
- Worksheet
- Owner
- Business owner with the AI Program Manager
- When it is used
- Completed with the pilot charter; baseline captured before the tool is introduced.
- Retention
- Life of the pilot plus two years.
Contributors: Participants, Data & Reporting where a system measure is needed
Downloads
The blank template is a standalone worksheet with every field and its guidance, ready to fill in or print. Printing uses this page’s print stylesheet, which drops the navigation and the download controls.
Required fields
The guidance matters more than the field name. A field labelled “risks” with no guidance gets filled in with “standard AI risks” and the artifact stops being worth reading.
| Field | Guidance | Required |
|---|---|---|
| Measure | One line, tied to the stated problem. | Required |
| Baseline and how it was captured | Before the tool arrives. The step most pilots skip. | Required |
| Target | What would make this worth scaling. | Required |
| Evidence level | Measured, estimate, or self-reported — decided up front. | Required |
| Harm indicators | What would tell you to stop, not just what would tell you it worked. | Required |
| Collection method and owner | Who collects it, from where, how often. | Required |
| Threshold for scale | The number that has to be met, written before you know the answer. | Required |
Worked example
Filled in against the sample record set, so the template can be judged on what a real entry looks like rather than on its headings.
PIL-003 — Standards knowledge search, measurement plan
- Citation accuracy: baseline not applicable (new capability). Target 95% on a fixed 100-question test set built by Data & Reporting. Measured. Threshold for scale: 95%. Current: 91%.
- Time to locate a current standard: baseline 6.5 minutes from timed observation of 20 lookups. Target under 2 minutes. Measured. Current: 1.8 minutes.
- Weekly active users: target 60% of 60 licensed pilot users. Measured from tool reporting. Current: 41 of 60.
Harm indicators and honesty note
- Answers citing superseded revisions presented as current — counted separately from general accuracy, because this failure mode misleads rather than merely missing.
- Users acting on an answer without opening the cited source — sampled through follow-up interviews.
- The 95% threshold was set in April, before anyone knew the tool would land at 91%. That is what makes the current 'not yet' credible rather than negotiable.
Blank template
The same fields with nothing in them. Print this page, or use the download above to get a standalone file.
Measure
One line, tied to the stated problem.
Baseline and how it was captured
Before the tool arrives. The step most pilots skip.
Target
What would make this worth scaling.
Evidence level
Measured, estimate, or self-reported — decided up front.
Harm indicators
What would tell you to stop, not just what would tell you it worked.
Collection method and owner
Who collects it, from where, how often.
Threshold for scale
The number that has to be met, written before you know the answer.