IBM — “New in watsonx Orchestrate: cross-platform agent discovery, custom evaluation and AgentOps Agent goes GA” (September 2026)

Tag: S-2026-09-03-ibm-orchestrate-cross-platform-discovery-agentops-ga Type: article (vendor “What’s New” announcement, IBM, published 3 Sep 2026; fetched in full) Author(s): Gauri Mathur, Product Marketing, watsonx Orchestrate (IBM) Date of source: 2026-09-03 (announcement; capabilities GA 17 Aug and 31 Aug 2026) Date ingested: 2026-09-15 Authority weight: medium — primary vendor announcement; GA dates and named capabilities are reliable as statements of what IBM says it ships; all efficacy and “govern any agent” claims are self-interested and unverified; no standard, customer or evidence model named Raw file: /_raw_sources/S-2026-09-03-ibm-orchestrate-cross-platform-discovery-agentops-ga.md

What it claims

IBM says watsonx Orchestrate shipped three generally available capabilities in August 2026, announced together on 3 September.

First, a new AI Gateway capability scans connected third-party platforms, discovers the agents running on them, and imports them so they are managed “from the same control plane you already use for native agents”. IBM lists four practical outcomes: a single inventory of native and external agents; policies set at the Gateway level that “apply to every agent connected through the AI Gateway”; discovery of duplicates across known registries with retirement of redundant agents; and connection of imported agents to business channels. Discovery and registration of Amazon Bedrock agents is GA (31 Aug); Azure AI Foundry and Google Vertex AI integrations “follow at the end of September”. IBM frames the point as: where an agent was built should stop determining whether it can be governed.

Second, two evaluation capabilities: Trace Inspector (GA 17 Aug) shows the full execution path of an agent run — each step, which tools were called, and “the exact point where behavior diverged” — so that debugging becomes “reading the record instead of reconstructing it”; and Custom LLM-as-a-Judge (GA 31 Aug) lets teams write their own natural-language evaluation criteria at build time, so that (IBM’s example) a claims-processing agent and a customer-onboarding agent are judged on different things.

Third, the AgentOps Agent (previewed in July; GA 31 Aug, native Orchestrate agents only) runs the “full improvement cycle” conversationally: plain-language observability discussion, automatic test-case creation, evaluation via simulated user conversations, root-cause analysis when an evaluation fails, and optimisation using GEPA and ACE — GEPA “generates and tests improved instructions”, ACE “builds a curated playbook of rules” — with the gain verified “before you promote anything”. Coverage “extends beyond native agents in September” to compatible external LangGraph agents using IBM’s SDK, including agents on AWS AgentCore.

IBM positions the sequence as June (single control plane) → July (post-deployment improvement preview) → August (GA, plus externally built agents), under the banner “Build anywhere, govern in one place”: enterprises “are not going to standardize on a single agent building platform”, and need “one place to see every agent, one set of controls that applies to all of them and one accountable owner for each”.

Notable quotes

  • “An agent built on one vendor’s platform is often invisible to another vendor’s governance tooling. Watsonx Orchestrate treats external agents as first class citizens, so the question of where an agent was built stops determining whether you can govern it.” (section 1)
  • “Debugging becomes a matter of reading the record instead of reconstructing it.” (section 2, Trace Inspector)
  • “Both verify the gain before you promote anything.” (section 3, GEPA/ACE)
  • “What they do need is one place to see every agent, one set of controls that applies to all of them and one accountable owner for each.” (closing section)

What’s speculative vs. asserted

Asserted (as vendor statements): the three capabilities and their GA dates; Bedrock-only discovery at GA; native-agent-only AgentOps at GA; Gateway-level policy application; Trace Inspector per-run traces; custom natural-language judge criteria.

Forward-looking (vendor roadmap): Azure AI Foundry and Vertex AI discovery “at the end of September”; AgentOps coverage of external LangGraph/AgentCore agents “in September”.

Vendor marketing / unverified: that Gateway-level policies constitute governance of external agents; that GEPA/ACE optimisation gains are “verified”; that trace records suffice to “read” rather than reconstruct behaviour. No customer, benchmark or independent evaluation is cited.

Not stated (gaps): any regulatory standard; the relationship between this Orchestrate discovery and watsonx.governance’s AI Asset Discovery (Jul 2026, which already named Bedrock and Azure AI Foundry as discovery sources — see S-2026-07-09-ibm-asset-discovery); whether Trace Inspector records are immutable, retained or exportable; who validates the AgentOps Agent’s rewritten instructions.

Vault inference (not a source claim): the AgentOps Agent’s automated rewriting of agent instructions is a model-change event under SS1/23 / SR 11-7-style change control; “Custom LLM-as-a-Judge” is an LLM assessing an LLM and would itself need validation.

Topics this feeds

  • AI Governance Platforms — MQ-Leader vendor extends its same-vendor pipeline (discovery → evidence → cross-platform discovery/evaluation/self-optimisation) into the agent-orchestration product line; fourth product in the Aug–Sep “standalone agent governance” cluster (Okta Agent SSO, Broadcom AgentMinder, Dataiku Agent Management).
  • IBM — third shipped agent-governance capability set recorded for IBM; first on the Orchestrate line rather than watsonx.governance.

Open questions raised

  • Are there now two IBM agent-discovery mechanisms (watsonx.governance AI Asset Discovery; Orchestrate AI Gateway discovery) with overlapping platform coverage — one architecture, or parallel offerings that a buyer must reconcile?
  • Does a “policy set at the Gateway level” for an imported Bedrock agent actually intercept that agent’s tool calls, or govern only traffic routed through Orchestrate? The source does not say.
  • Who validates instruction rewrites produced by the AgentOps Agent’s GEPA/ACE loop before promotion, and is the “verified gain” evidence a validator could use?
  • Is a Custom LLM-as-a-Judge evaluation acceptable as validation evidence under SS1/23 / SR 11-7 without an independent (non-LLM or independently validated) benchmark?