Skip to content

Back to the pilot tracker

Sheet 05 · PIL-002

Specification cross-check

Specifications drift from firm standards between revisions, and QA catches conflicts late enough that fixes mean reissued sheets.

Record summary

Stage
Active
Risk pathway
Structured review
Data classification
Confidential
Participants
9

Current status

Active in two offices, week 9 of 16. False positives above threshold; scope tightening in progress with Data & Reporting.

Next decision
Continue, modify scope, or pause at the week-16 review
Where it is decided
Biweekly pilot check-in, then AI Governance Committee 2026-10-06

The record

Expected outcome
Conflicts surface one review cycle earlier, with a false-positive rate low enough that reviewers keep using it.
Business owner
Alan Petrov, QA Lead, Structural Practice Group
Program coordinator
AI Program Manager — coordinates and records; does not approve
Participating teams
Structural Practice Group
Tool
Cadence Search — Cadence Data Systems TL-02
Originating use case
Specification cross-check against internal standards UC-003
Related decisions
DEC-2026-03
Start date
11 May 2026
Decision date
06 Oct 2026 (in 34 days)
Cost to date
$6,200 — licensing, vendor fees, and third-party support. Participant time is not included and is the larger cost.

Measures

Defined before the pilot started. That is what makes a result a finding rather than an opinion formed afterwards.

PIL-002 measures, targets set at pilot start, with the evidence behind each current value.
MeasureTargetCurrentEvidence
Conflicts identified per reviewAt or above the manual baseline of 3.14.4 across 26 reviewsMeasured
False-positive rateUnder 20%31%Measured
Reviewer willingness to keep using itMajority yes5 of 9 yes, 4 conditional on fewer false positivesSelf-reported

What the pilot has surfaced

Open questions and lessons sit here rather than in a closing report, because they are what the next pilot needs and the next pilot starts before this one ends.

Open questions

  • Can the false-positive rate reach 20% by tightening the comparison scope, or is the ceiling structural?
  • Does the tool miss conflict types that manual review catches? The missed-conflict analysis is not finished.

Lessons learned

  • Reviewers tolerate false positives up to a point, then stop opening the report. The threshold matters more than the average accuracy.
  • Most false positives came from superseded revisions that were never retired from the library. A data problem, not a model problem.

Guardrails and dependencies

Guardrails

  • Advisory output only — the tool never edits a specification
  • A licensed engineer adjudicates every flagged difference
  • Limited to two offices during the pilot

Dependencies

  • Standards library revision tagging (owner: Data & Reporting)
  • Cadence Search conditional approval remains valid through EXC-02