Business Glossary Build Kit — QA Review (ChatGPT)
Tag: S-2026-06-07-glossary-kit-qa-review Type: report Author(s): ChatGPT (independent QA), commissioned by Paul Date of source: 2026-06-07 Date ingested: 2026-06-07 Authority weight: medium — an independent AI QA review; useful and well-structured, but its specific recommendations require expert judgement before adoption (one was adapted on review). Raw file: /_raw_sources/S-2026-06-07-glossary-kit-qa-review.md
What it claims
An independent QA of the CLM Business Glossary Build Kit. Overall rating 8.8/10, verdict “Approve for Pilot Deployment with Enhancements”, with the identified issues framed as enhancement opportunities rather than material defects. Strengths cited: separation of AI vs human responsibilities, audit-trail design across T-01–T-04, governance guardrails, the conflict-resolution workflow, template maturity, consistent confidence scoring and practical operating guidelines. Seven priority recommendations: (1) define confidence thresholds; (2) introduce a source-authority hierarchy; (3) create a definition-quality standard; (4) expand the regulatory taxonomy; (5) strengthen bulk-approval controls; (6) add a golden-source strategy; (7) add AI model/prompt/skill version traceability. Plus prompt-level (OCR guidance, similarity thresholds, authority references, regulatory hierarchy, version control, measurable QA scoring), template-level (similarity & authority scores in T-03; expanded T-05; AI confidence, source authority and version fields in T-06) and test-data (homonyms, acronyms, obsolete terms, regulatory/policy conflicts, near-synonyms) enhancements, and a new Stage 7 (governance, publication & change control).
Notable quotes
“AI accelerates. Humans approve.” (executive summary framing). “The identified issues are enhancement opportunities rather than material defects.”
What’s speculative vs. asserted
Asserted: the scores, strengths and the list of recommendations. Speculative / requiring judgement: the specific numeric confidence thresholds (≥95% streamlined / 80–94% review / <80% manual) — proposed as if AI confidence were a calibrated probability, which it is not. This recommendation was adapted rather than applied literally (see the changes-applied source).
Topics this feeds
CLM Glossary Acceleration Squad — the QA drove the v1.1 upgrade of the build kit.
Open questions raised
Whether to adopt numeric confidence gates (resolved: adapted to categorical bands with a calibration caveat, per Paul). Validation of the expanded regulatory references by the bank’s Regulatory/Risk SME.