Harness — “New Harness Report Reveals Enterprise Confidence in AI Agents Isn’t Backed by Real Controls” / The State of Agent DLC 2026 (10 September 2026)
Tag: S-2026-09-10-harness-state-of-agent-dlc-2026 Type: report (vendor-commissioned survey; press release read in full, report PDF not read) Author(s): Harness (fieldwork by Sapio Research, July 2026; quotes: Keith Mann, Field CTO and Head of Research; Trevor Stuart, SVP and GM) Date of source: 2026-09-10 Date ingested: 2026-09-17 Authority weight: medium — third-party fieldwork (n=700, five countries, screened for large enterprises with agents at least in live PoC) with a stated methodology, but commissioned and framed by a vendor that sells the missing controls; questionnaire and data tables unseen; not sector-specific Raw file: /_raw_sources/S-2026-09-10-harness-state-of-agent-dlc-2026.md
What it claims
Harness argues that AI agents “don’t behave like deterministic software” and therefore need controls “built for that variability”, while most organisations still rely on controls “built for deterministic software”. The survey’s central finding is a “confidence gap”: confidence in each control domain “lands in the mid-70s”, but the control that would back it up is in place “for less than half of organizations, and in some cases fewer than one in five”.
The headline pairs: inventory — 77% confident they have a complete inventory of every agent, MCP server and LLM, 44% run active discovery tooling; testing — 74% confident testing would catch a production-impacting failure, 19% have a gate that automatically blocks every bad release; kill switch — 76% believe they could disable a misbehaving agent within 15 minutes, 33% have an instant kill switch; security — 75% say agents are secure end to end, yet that group had incidents at 88% versus 87% overall; spend — 74% claim a complete per-agent spend picture, 60% overran budget last quarter.
On change management: 42% run prompt edits through the same pipeline as code, 34% have a dedicated configuration system for AI behaviour; 53% of agent-related changes go through any standard pipeline, 37% run fewer than half through one; more than four in ten decide case by case whether to trust a change, and of those promoting agent changes to production only 58% check every change against a fixed, repeatable standard. 58% report more production incidents per 100 changes since deploying agents; “roughly 7 in 8” had at least one tangible agent-related issue this year.
Recommendations: treat the agent lifecycle “as its own discipline” (evals, security, inventory, rollback built for agent behaviour); replace ad hoc reviews with a fixed, repeatable standard; adopt progressive rollout (canary/blue-green) for agent changes.
Method: 700 technology professionals (software engineering, IT operations/infrastructure, technology leadership) at enterprises with 1,000+ employees, 100+ developers and >$100M revenue in the US, UK, France, Germany and India; Sapio Research, July 2026; participation required agents in production, pilot or live PoC — “the figures reflect how far committed adopters have progressed rather than how widespread adoption is”.
Notable quotes
- “Confidence lands in the mid-70s of those surveyed in each case, but the control that would back it up is in place for less than half of organizations, and in some cases fewer than one in five.” (section 1)
- “Organizations need to test whether a control actually holds up against an agent’s variability, not assume it does because the control exists.” — Keith Mann
- “Teams moved fast to build and release agents, and are now circling back to ask how to actually govern what they’ve already shipped.” — Trevor Stuart
What’s speculative vs. asserted
Asserted (as survey findings): all percentages above, the sample description and the fieldwork window. These are reliable as what the survey found, subject to the caveats below.
Vendor framing / interpretation: “confidence gap” as the organising narrative; the claim that deterministic-software controls are the wrong model; the three recommendations (which map to Harness products — Agent DLC, AI Evals, inventory/discovery, rollback).
Not stated (gaps): sector or country breakdowns; questionnaire wording (what “complete inventory” or “instant kill switch” meant to respondents); how “security incident” was defined; whether governance/risk-function respondents were included (the screen names engineering/IT/technology leadership only); any regulation or standard.
Vault inference (not a source claim): the pairs are a ready-made checklist for an independent assurance review — each “confidence” statement is a first-line assertion and each “control in place” figure is the evidence test; the numbers are directionally consistent with the OneTrust 2026 governance-maturity report ingested 14 Sep [S-2026-09-14-onetrust-ai-ready-governance-report-2026] but the two surveys differ in respondent type and cannot be combined.
Topics this feeds
- AI Governance Platforms — supplies buyer-side evidence that the inventory, gate and kill-switch capabilities the control-plane vendors sell are, by respondents’ own account, absent in most adopters; the 44% discovery / 33% kill-switch / 19% gate figures are the demand-side counterpart of the supply-side cluster recorded 15–17 Sep.
- AI Governance Maturity Gap — a further quantified instance of the assertion-versus-evidence gap; engineering-leadership sample, cross-sector.
Open questions raised
- What is the FS-only cut, if any, in the full report? (PDF not read.)
- How did respondents interpret “complete inventory” — does it include third-party SaaS-embedded agents and MCP servers exposed by vendors?
- Would a regulator or auditor accept a canary rollout as an “appropriate human oversight” or change-control measure for an agent under EU AI Act Art. 14 / DORA ICT change management — or is a pre-deployment gate required for high-risk uses regardless?
- Is the 88%-vs-87% incident parity evidence that “secure end to end” self-assessments are uninformative, or an artefact of how “incident” was defined?