Paul Miles — CLM Glossary Build Kit Approach & Work Breakdown

Tag: S-2026-06-05-clm-glossary-build-kit-approach Type: own-writing Author(s): Paul Miles (drafted with AI assistance) Date of source: 2026-06-05 Date ingested: 2026-06-05 Authority weight: high — Paul’s own design document for an active client engagement Raw file: /_raw_sources/S-2026-06-05-clm-glossary-build-kit-approach.md

What it claims

Defines the work breakdown for a “build kit” equipping the CLM glossary squads: 6 tool-agnostic prompts (Harvest & Source Registration; Cluster & De-duplicate; Conflict Consolidation; Policy/Process/Regulatory Tagging; Master Assembly & Status; QA & Approval Readiness), 7 skill.md files (a shared _guardrails.md carrying the 11 GR controls from the Taxonomy Knowledge Tool HLD plus six task skills), 6 Excel templates (Source Register & Audit Log; Harvest Staging; Clustering & Disposition; Conflict/SME Workshop Pack; Reference Lists; Master Business Glossary extending the client’s Taxonomy Template V0.1 with all 19 existing columns retained and ~15 added, including a status lifecycle ending in “Ready to be approved”/“Approved” and link placeholders for technical metadata and DQ metrics), and one client-ready Squad Operating Guidelines document. Components map 1:1 to the squad deck’s six-step delivery flow and are designed for reuse in the future Taxonomy Knowledge Tool (prompts → agent tasks, skill.md → catalogue entries, templates → governed extract/load formats). Six delivery risks are recorded: clustering accuracy, soft-tag overreach where reference lists are incomplete, consolidated-definition hallucination, source-file parse failures, Excel/SharePoint concurrency at the master, and prompt portability across LLM tooling.

Notable quotes

“Every prompt is paired with a skill.md and a template, so the same components drop into the future Taxonomy Knowledge Tool” (§1, design principle).

What’s speculative vs. asserted

Asserted: the component inventory, the mapping to the squad delivery flow, the field treatment of Taxonomy Template V0.1. Speculative: that throughput assumptions (~250 terms/week) hold — explicitly conditioned on clustering accuracy; that prompts behave consistently across M365 Copilot and other LLMs (pilot recommended).

Topics this feeds

CLM Glossary Acceleration Squad — the build kit operationalises the squad’s delivery method.

Open questions raised

Whether the client supplies complete Policy/Process/Regulatory reference lists (drives hard- vs soft-tag balance); actual harvested term count (re-baselines throughput).