Intelligine Group — “AI in financial services: the 2026 governance benchmark” (February 2026)

Tag: S-2026-02-intelligine-governance-benchmark Type: report (independent consulting-firm primary-data benchmark) Author(s): Raghav Ram, PhD — Managing Partner, Intelligine Group (site footer also renders the name “Raghav Ramabadran, PhD”; discrepancy noted, not resolved) Date of source: 2026-02 (page byline “February 2026”; underlying interviews conducted September 2025 – January 2026) Date ingested: 2026-09-14 Authority weight: low — a single independent management/technology consulting firm reporting its own primary engagement data (n=42), self-described methodology, not peer-reviewed, full dataset/instrument unpublished; figures are the firm’s own and not independently verifiable. Directionally credible because the sample is named and substantial (42 institutions, >$41tn AUM/in-force) and the interviews were with the sitting Chief AI Officer or equivalent at each firm. Raw file: S-2026-02-intelligine-governance-benchmark.md. Primary page (fetched in full 2026-09-14): intelliginegroup.com/article/governance-benchmark.

Ingested 2026-09-14 (daily regulatory-intelligence scan — practitioner-research source line). Selected as the single practitioner signal this run because the four regulator sources (EBA, FCA, BIS/Basel, EU AI Office) returned no net-new in-window item and the other major 2026 FS AI-governance surveys are already in the vault. ⚠️ Date caveat: the byline is February 2026 — outside the strict past-7-day window the scan targets — but the benchmark is net-new to the vault and directly on-point. Authority low (see above).

What it claims

Intelligine Group’s benchmark presents primary data from 42 banks and insurers (24 banks — 11 G-SIBs, 13 regional/national; 18 insurers — 11 life, 7 P&C; aggregate assets under management or in force exceeding US$41 trillion) on how the operating model around enterprise AI is instrumented, governed and escalated. The dataset is drawn from structured interviews with the sitting Chief AI Officer or equivalent role at each institution, conducted September 2025 – January 2026, supplemented by review of the governance artefacts in production at the time of each interview. The stated objective is operational rather than technological — “what the operating model actually does, not which models or vendors institutions have adopted.”

The headline finding is that “the dispersion in operating-model maturity across the 42 institutions is significantly larger than the dispersion in technology stack. The technology choices are converging. The operating models are diverging.” Firms reporting the cleanest operating economics are not those with the most advanced stack but those with the most disciplined operating model.

The benchmark locates the discipline gap in three artefacts:

  1. Kill criteria. Only 19 of 42 firms have at least one kill criterion specified at the unit-economic level and continuously instrumented; the other 23 specify kill criteria only at program-spend level, or at unit-economic level but not continuously instrumented. Among the 19, the median portfolio stayed within 12% of the approved cost envelope over the most recent fiscal year; among the 23, the median portfolio exceeded the approved envelope by 68% (tail case: a regional bank at 340% of the original envelope).

  2. Escalation paths. Only 7 of 42 firms have a written escalation path that can move from a kill-criterion breach to a decision inside 72 hours without a board calendar event; 24 require a quarterly board review at the earliest; the remaining 11 have no written escalation path and would improvise at the moment of breach.

  3. Audit posture. The sample is converging on a model that separates the runtime audit trail (what the system did, per-call, retained for a regulator-set period) from the governance artefact (what the system was permitted to do, at policy level, with kill criteria and escalation path fixed at the last rewrite), reconciled and audited quarterly. Cleanest reconciliation occurred where the two artefacts were maintained by separate teams reporting through separate lines; most painful where the same team wrote the policy and recorded the compliance evidence.

On portfolio tiering: 31 of 42 firms report a formal tiering scheme (23 three-tier, 6 four-tier, 2 five-or-more). The modal three-tier scheme keys governance to proximity to a regulated decision — tier 1 (finalises a regulated decision): documented model risk management, instrumented kill criteria at unit-economic and quality level, board-visible escalation, full audit trail; tier 2 (materially influences): documented MRM, unit-economic kill criteria, escalation below board, summary audit trail; tier 3 (outside the regulated surface): unit-economic kill criteria, workflow-owner escalation, no formal audit trail required.

The benchmark closes with four board recommendations — every workload above tier three to carry a continuously instrumented unit-economic kill criterion; the escalation path to be written, calendar-independent and capable of moving inside 72 hours; the runtime audit trail maintained by a team separate from the policy-writing team; and the strategy’s rewrite trigger to sit below the board on a movable calendar — noting “none of the four recommendations is novel.”

Notable quotes

  • “The technology choices are converging. The operating models are diverging.” (§ headline finding)
  • “We had the kill criteria. We had the instrumentation. What we did not have was a path from the breach to a decision that did not require a board agenda item, and the agenda was set six weeks in advance.” (Interview, Chief AI Officer, regional bank, Q4 2025)
  • On audit posture: the model “separates the runtime audit trail from the governance artifact”. (§05)

What’s speculative vs. asserted

  • Asserted (Intelligine’s own primary data): the n=42 sample composition and >$41tn aggregate; the headline operating-model-vs-stack dispersion finding; 19/42 continuously-instrumented unit-economic kill criteria and the 12%-vs-68% cost-envelope association; 7/42 sub-72-hour calendar-independent escalation, 24 quarterly-board, 11 none; 31/42 formal tiering (23 three-tier); the audit-trail/governance-artefact separation observation.
  • Presented as association / operator judgement, not proven causation: the link between continuously instrumented kill criteria and staying within the cost envelope (the firm reads discipline → economics, but this is cross-sectional interview data, not a controlled comparison); the claim that separate-line audit ownership produces cleaner reconciliation.
  • Authority / verification caveats: single-firm self-reported benchmark; methodology self-described and full instrument/dataset unpublished; figures not independently verifiable; global G-SIB/large-insurer-weighted with no UK/EU cut, so EU/UK read-across is directional; author-name discrepancy unresolved.
  • Ingesting-agent inference (not the source’s claim): the mapping to Paul’s service lines, and the read that the runtime-trail-vs-governance-artefact separation is a Three Lines of Defence for AI articulation, are the wiki’s assessment.

Topics this feeds

  • AI Governance Maturity Gap — an operating-model cut that reframes the 2026 maturity gap as a dispersion in governance discipline (kill criteria, escalation paths, audit-trail separation) that is wider than the dispersion in technology, quantified with unit-economic outcomes.
  • Model Risk Management and Agentic AI — the three-tier scheme keys documented MRM and instrumented kill criteria to proximity to a regulated decision; a practitioner articulation of tiered model controls for AI/agentic workloads [inference].
  • Three Lines of Defence for AI — the separation of the runtime audit trail from the governance artefact, on separate reporting lines, is a 3LoD-style independence argument for AI evidence [inference].
  • Service Line — Independent Governance Assurance — the escalation-path and audit-reconciliation findings are precisely the evidence-and-assurance gap this service line addresses.

Open questions raised

  • Do the kill-criterion / escalation-path / audit-separation findings hold in a UK/EU-specific sample, or are they skewed by the US-weighted G-SIB composition?
  • Is the 12%-vs-68% cost-envelope gap driven by kill-criteria discipline itself, or by a confound (more disciplined firms differ on many dimensions)? The source treats it as evidence for instrumented kill criteria; the causal claim is unproven.
  • How does Intelligine’s three-tier “proximity to a regulated decision” scheme map to the EU AI Act high-risk classification and to firms’ existing model-risk tiers (SR 11-7 / SS1/23)?