Skip to content

Back to the artifact library

Sheet 07 · ART-09

Pilot measurement plan

Defines what will be measured, against what baseline, by whom, and how good the evidence will be — agreed before the pilot starts so the result cannot be reverse-engineered afterward.

Artifact details

Format
Worksheet
Owner
Business owner with the AI Program Manager
When it is used
Completed with the pilot charter; baseline captured before the tool is introduced.
Retention
Life of the pilot plus two years.

Contributors: Participants, Data & Reporting where a system measure is needed

Downloads

The blank template is a standalone worksheet with every field and its guidance, ready to fill in or print. Printing uses this page’s print stylesheet, which drops the navigation and the download controls.

Required fields

The guidance matters more than the field name. A field labelled “risks” with no guidance gets filled in with “standard AI risks” and the artifact stops being worth reading.

Fields in ART-09, with the guidance given to whoever fills them in.
FieldGuidanceRequired
MeasureOne line, tied to the stated problem.Required
Baseline and how it was capturedBefore the tool arrives. The step most pilots skip.Required
TargetWhat would make this worth scaling.Required
Evidence levelMeasured, estimate, or self-reported — decided up front.Required
Harm indicatorsWhat would tell you to stop, not just what would tell you it worked.Required
Collection method and ownerWho collects it, from where, how often.Required
Threshold for scaleThe number that has to be met, written before you know the answer.Required

Worked example

Filled in against the sample record set, so the template can be judged on what a real entry looks like rather than on its headings.

PIL-003 — Standards knowledge search, measurement plan

  • Citation accuracy: baseline not applicable (new capability). Target 95% on a fixed 100-question test set built by Data & Reporting. Measured. Threshold for scale: 95%. Current: 91%.
  • Time to locate a current standard: baseline 6.5 minutes from timed observation of 20 lookups. Target under 2 minutes. Measured. Current: 1.8 minutes.
  • Weekly active users: target 60% of 60 licensed pilot users. Measured from tool reporting. Current: 41 of 60.

Harm indicators and honesty note

  • Answers citing superseded revisions presented as current — counted separately from general accuracy, because this failure mode misleads rather than merely missing.
  • Users acting on an answer without opening the cited source — sampled through follow-up interviews.
  • The 95% threshold was set in April, before anyone knew the tool would land at 91%. That is what makes the current 'not yet' credible rather than negotiable.

Blank template

The same fields with nothing in them. Print this page, or use the download above to get a standalone file.

Measure

One line, tied to the stated problem.

Baseline and how it was captured

Before the tool arrives. The step most pilots skip.

Target

What would make this worth scaling.

Evidence level

Measured, estimate, or self-reported — decided up front.

Harm indicators

What would tell you to stop, not just what would tell you it worked.

Collection method and owner

Who collects it, from where, how often.

Threshold for scale

The number that has to be met, written before you know the answer.