You Need All Nine Layers to Govern a Multi- Agent Portfolio

What each layer does, who should own it, and which vendors actually deliver
September 30, 2026
·
10 min Read

Why a Portfolio Is a Different Problem Than an Agent

Most AI governance conversations start with a compliance checklist or a model risk score, and almost always with one agent in mind. That's the wrong starting point. Companies aren't heading toward one agent. They're heading toward dozens, built on different frameworks, running on different platforms, and increasingly talking to each other rather than only to people.

The problem grows faster than the portfolio does. Agent count rises in a straight line, but the connections between agents multiply. Ten agents can produce up to 45 possible interactions. A hundred agents can produce nearly 5,000. Ten agents on one platform is manageable. Ten agents spread across Salesforce, ServiceNow, Workday and a few custom builds, handing work to each other, is a different problem. Most companies are already in the second situation, planned or not.

Two protocols made this real rather than theoretical. MCP, from Anthropic, standardizes how an agent reaches tools and data. A2A standardizes how one agent finds and delegates to another, whatever framework or vendor built it. Google started A2A and donated it to the Linux Foundation in mid-2025, with AWS, Cisco, Microsoft, Salesforce, SAP and ServiceNow as founding members. By April 2026 it had more than 150 supporting organizations and had shipped inside Microsoft Copilot Studio, Azure AI Foundry and Amazon Bedrock AgentCore. Agents talking to each other across platforms is infrastructure now, not a future problem, and the single-agent checklist no longer covers it.

That's what this piece is for. Not a plan for after something breaks, but the full set of layers you need in place, and who should own each one, before the portfolio gets ahead of you.

One question runs through every layer: who has the authority to decide, not just who has the tools to act. Is this a business call, an IT call or an engineering call? Getting that wrong is a governance failure on its own, no matter what you buy. Most companies hand every agent-related layer to whoever built the agents, usually engineering or IT, because they're the ones who can run the console. That's the right instinct at the bottom of the stack and the wrong one at the top.

How to read the vendor tables

Each layer ends with the same table. The rating tells you how much weight to give each vendor once your agents span more than one platform:

  • Lead — holds up across platforms. Start here.
  • Strong — real capability, narrower fit or a specific use case.
  • Not enough alone — useful, but blind outside its own boundary. Pair it with a Lead.

One disclosure first. ForceEquals builds at Layers 7, 8 and 9, and we appear in those tables. Weigh our read of those layers accordingly, and use the three questions in Layer 7 to check it yourself rather than taking our word for it. We don't sell into the layers below 7, and we've rated the vendors there the way we'd want someone to rate us.

Two things to know before you read the tables. First, this market is consolidating fast. Three leading identity vendors were acquired in a single three-month window in 2026, so treat these as a snapshot rather than a permanent map. Second, a lot of what calls itself "agent guardrails" or "agentic governance" is really engineering-configured containment: sandboxing, gateways and identity scoping, described in business language. The middle column is written to cut through that.

1. Foundational Considerations

Owner: Business (executive leadership, risk, legal). Before any technical control, you need agreement on three things: risk appetite, accountability, and scope. Who owns the consequences when an agent gets it wrong? How much damage are you willing to absorb? Which use cases can act on their own, and which stay advisory? This is the one layer where the authority has to come from the top, however the rest of the stack is organized.

Vendor What it actually does Weight
Credo AI Policy Packs translate NIST AI RMF, EU AI Act, and ISO 42001 into controls that don't assume a single platform underneath them Lead — framework-agnostic by design
OneTrust Extends the privacy/GRC suite many enterprises already run, with automatic risk re-classification when models, data, or agents change Strong — best if you don't want a second GRC system just for AI
IBM watsonx.governance The enterprise incumbent: AI use-case inventory, risk scoring, and model documentation inside an existing IBM governance footprint Strong — the default in IBM shops; heavier lift than the two above elsewhere
Monitaur Model risk management carried over from regulated industries: policy-to-proof workflow, use-case inventory, a controls library, and evidence capture, with insurance customers including Progressive, Unum, The Hanover, and Verisk Strong — deepest fit where an MRM program already exists; narrower outside regulated industries

Also worth a look: Trustible, Fairly AI, Relyance AI.

2. Data Governance

Owner: IT/data engineering, with business data owners setting sensitivity. Agents are only as trustworthy as what they can see and touch. With one agent the question is simple: can this agent see this data? With many agents across many platforms there's a harder one. Several agents, each staying inside its own permitted scope, can together reconstruct something none of them should see alone. No single platform catches that, because no single platform can see the other agents.

Vendor What it actually does Weight
Collibra Shared system of record across AI, data, and risk teams; lineage from source datasets through training, inference, deployment, and usage, with model and agent registries. Its 2026 Raito acquisition added data-access control Lead — lineage doesn't stop at one platform's boundary
Immuta Enforcement layer inside warehouses, lakes, and analytics environments; dynamic attribute-based access control on identity, role, purpose, and sensitivity Lead — agnostic to which agent on which platform is querying
Atlan Active metadata with bidirectional sync into Snowflake, Databricks, dbt, and Slack; AI asset registry and agent-ready metadata context. Gartner MQ Leader 2026 Strong — the live-sync alternative to Collibra's more static catalog
Promethium Captures lineage at query execution via its Trust Harness, so every SQL call an agent generates gets lineage automatically Strong — layers into Collibra, Atlan, or Unity Catalog rather than replacing them
Kiteworks Governs file, email, and AI-agent data exchange from one place; AI Data Gateway logs prompts and outputs, enforces least-privilege per MCP call Strong — useful when data leaves through many channels, not just queries
Microsoft Purview Data discovery, classification, and DLP across the Microsoft estate, now extended to AI prompts and agent activity Strong — the incumbent answer in Microsoft-heavy shops; weaker once data lives outside that estate
AWS Lake Formation / DataZone, Google Dataplex Cloud-native catalog, classification, and fine-grained data access inside their own estates Not enough alone — real capability, scoped to one cloud; blind to agents reaching the same data from elsewhere
Databricks Unity Catalog One permission and lineage layer across data, models, tools, functions, and agents, with row- and column-level controls and audit logs covering every table, model, and tool an agent touches. Permissions are defined once and hold whether the agent runs in a notebook, in SQL, or as a deployed app Lead — genuinely agent-aware rather than a warehouse catalog with agents bolted on; strongest where the lakehouse is already the data platform
Snowflake Horizon Native governance inside the Snowflake estate Not enough alone — fine if most data sits in one platform, blind to agents outside it

Also worth a look: BigID, Securiti AI.

3. The LLM Layer

Owner: Engineering/ML. The model itself: selection, evaluation, versioning, and finding failure modes before production does. Models fail in ways that are hard to see from the outside and hard to reproduce on demand, which is why evaluation infrastructure decides whether a pilot ever becomes something you can run.

Model choice belongs in the governance conversation, not only the build conversation. Which providers are approved for which use cases. What happens when a model is deprecated on the provider's schedule rather than yours. Whether prompts and data leave your boundary. What the provider's own safety evaluations actually cover. These are policy decisions with real consequences, and they're increasingly made by whoever is building the agent. A team that ships a working agent in two weeks picked a model in the first hour, usually without anyone signing off on it.

Once you have many agents, what you evaluate changes. Whether one response was accurate matters less than whether the whole chain was. The sequence of calls across several agents and models, often from different vendors, is what produced the outcome. That favors platforms built for distributed tracing over ones built to score a single model's output.

Vendor What it actually does Weight
The model providers — OpenAI, Anthropic, Google, xAI (SpaceXAI), Meta, Mistral The models themselves, plus the governance-relevant material around them: model cards, published safety evaluations, deprecation schedules, data-handling terms, and enterprise controls Lead — the layer's actual subject; which models are approved, and on whose deprecation timeline, is a governance decision before it's a procurement one
Arize Distributed tracing lineage from Phoenix; OpenTelemetry-native with broad framework coverage Lead — built for trajectories, not single calls
Galileo Enterprise multi-agent trace support, plus small purpose-built evaluator models (97% cost reduction versus GPT-4-based eval, 152ms latency) Lead — cheapest path to eval at production volume
Braintrust Trace-to-eval workflow: production failures become eval cases, CI gates block release Strong — best when you span frameworks rather than standardizing on one
Databricks MLflow 3.0 Universal tracing that monitors and benchmarks any model or agent, on Databricks or across clouds and on-prem, with built-in and custom LLM judges, human-in-the-loop feedback, and a prompt registry versioned alongside model artifacts Strong — unusual among platform-native tools in claiming reach beyond its own estate; a real alternative to the specialists if the lakehouse is already there
Arthur Model monitoring (drift, bias, explainability) extended to agents via Agent Discovery & Governance, Dec 2025: agent/sub-agent/tool discovery, threshold alerting, policy attestation Strong — monitoring-first heritage, same category as Fiddler
Fiddler Drift detection plus explainability; agentic observability added via its April 2026 Lumeus acquisition Strong — acquired capability still integrating as of mid-2026
Datadog LLM Observability LLM traces correlated against existing infrastructure and APM data Strong — right call when you want one pane with your infra, not a separate eval tool
LangSmith and other framework-native eval Deep evaluation inside its own framework Not enough alone — can't see across frameworks it wasn't built for

Also worth a look: Langfuse (the open-source default), Weights & Biases Weave, Patronus AI.

4. Identity & Access

Owner: IT/security engineering. Agents need their own identities, separate from the people who built them. This is where the one-to-many problem hits hardest. If every platform issues its own service accounts, you get a separate identity silo for every platform-and-agent combination, none of them visible to each other.

That blind spot set off a wave of acquisitions. Cisco announced its purchase of Astrix Security in May 2026, SailPoint closed on Entro Security on June 29, 2026, and Cyera signed a letter of intent for Oasis Security in July 2026 for roughly $1 billion. All three buyers were identity platforms built for human employees, suddenly being asked agent questions they weren't designed to answer.

What works is one identity layer that follows the agent across platforms, rather than one scoped to wherever the agent happens to run.

Vendor What it actually does Weight
SailPoint Unifies human, machine, and agent identity in one governance plane; lifecycle governance from creation to retirement Lead — one plane, not one per platform
PlainID Continuous real-time authorization across the full agent flow — prompt, data retrieval, tool and MCP invocation, output — with zero standing privileges Lead — spans the flow regardless of which platform started the call
Aembit Workload identity enforced at the point of access; secretless, just-in-time credentials via MCP Identity Gateway Lead — platform-agnostic by construction
CyberArk (Secure AI Agents) Privileged access plus the Venafi machine-identity portfolio; inside Palo Alto Networks after a $25B close in Feb 2026 Strong — deepest PAM heritage, now part of a platform roadmap
Clutch Security Non-human identity discovery and governance Strong — most prominent NHI specialist still independent
Token Security, Britive, Keycard, Corsha, Defakto, P0 Security, Natoma Various cuts at non-human and agent identity, pitching cross-platform neutrality Strong — viable independents; expect continued acquisition over 2–3 years
Microsoft Entra Agent ID Agent identities with blueprints, Conditional Access for agents, lifecycle governance, and third-party agent registration via SDK or workload identity federation — including agents on AWS Bedrock and n8n Lead — the incumbent every Entra customer already has, and more portable than its name suggests; deepest advantage inside the Microsoft estate
Okta Identity incumbent extending human IAM patterns to agents, with cross-app access brokering for agents calling third-party services Strong — the default where Okta is already the identity layer
AWS IAM / Verified Permissions, Google Cloud IAM Mature cloud authorization applied to agents as workloads, including scoping which agent may invoke which model or tool Not enough alone — deep inside one cloud, absent outside it; AWS can require a specific guardrail on inference through IAM conditions
SPIFFE/SPIRE Open workload identity standard several vendors above build on Not enough alone — substrate, not a product
Platform-native agent identity Identity scoped to one platform's own boundary Not enough alone — blind the moment an agent crosses into another platform

Worth being precise: PlainID governs access, meaning whether an identity can reach a tool or a dataset. It doesn't give the business a threshold to tune. That's why it sits here rather than in Layer 7.

5. Security

Owner: IT/security engineering. With identity in place you can secure the attack surface. With many agents, that surface moves to the seams between them. One agent's tool calls are a known risk. A handoff from a low-trust agent to a high-trust one, or an MCP call crossing an org boundary, is newer and much less inspected. The Cloud Security Alliance found in 2026 that 82% of organizations have unknown AI agents running in their environment, and 65% have already had an agent-related incident.

The platform vendors are serious at this layer, and it's worth saying why that still isn't the whole answer. AWS, Google, Microsoft and Salesforce all ship real enforcement for agents running on their own infrastructure: filtering, injection detection, policy at the gateway, audit trails. Each is good. Each stops at its own boundary. Run agents in three clouds and you're operating three enforcement models with three policy languages and no shared view of any of them. That's the case for buying at least one control that spans them.

Databricks goes furthest toward the layers above. Unity AI Gateway's hard spend caps are the closest any platform vendor gets to a business-level limit. It's still a platform-team setting rather than something a finance owner moves against live escalations, but the instinct is right.

Vendor What it actually does Weight
NVIDIA Open Agent Safety Platform Free, open-source, two parts: OpenShell runs on CPUs and sets hard boundaries on agent action; Sentry runs on network silicon and quarantines a runaway agent in milliseconds. Announced Sept 28, 2026, in response to sandbox-escape incidents at OpenAI, Anthropic, Meta, and Google Lead — hardware-anchored containment no other vendor occupies (see note below)
Zscaler AI security through the Zero Trust Exchange cloud proxy, plus an MCP Gateway shipped early 2026 Lead — sits above any single platform's own governance
Palo Alto Prisma AIRS The broadest platform here: AI Runtime Firewall, agent identity verification and action blocking, model scanning, automated red teaming, posture management, and an AI Gateway positioned as an enterprise-wide AI control plane. Palo Alto also now owns CyberArk (Layer 4) Lead — widest coverage of any single security vendor; still enforcement, not business policy
PointGuard AI MCP Security Gateway plus Agent Mission Control: zero-trust authorization, sandboxing, ring isolation, kill switches, pre-execution action validation at sub-millisecond latency Lead — explicitly cross-platform
WitnessAI Network-level, no endpoint client; discovers agents and MCP servers, enforces org-wide tool allow-lists, intent-based (not keyword) inspection at runtime Lead — governs human AI use and agent use from one control plane
Lakera / Aim Security Runtime behavioral inspection agnostic to origin platform; now under Check Point and Cato respectively Strong — capability intact, roadmaps now set by acquirers
Kosmoy Kernel-enforced sandboxing ("Action Capsule") with per-task credentials and a kill switch, plus one gateway enforcing RBAC, budgets, and logging across LLM, MCP, and A2A calls Strong — containment plus spend control in one place
Airia Runtime enforcement at the execution layer — intercepts before the tool call fires or the email sends — plus discovery and auto-generated compliance documentation (EU AI Act, NIST, ISO 42001, HIPAA, SOC 2) Strong — enforcement and evidence together
Lyzr Open Controller Permission gate between eval and action: agents that pass evaluation still need permission to act Strong — narrow but well-targeted
ServiceNow AI Control Tower Cross-platform discovery and inventory of agents, models, and MCP servers across AWS, Azure, Google Cloud, SAP, Oracle, Workday and 25 more systems, with risk scoring, least-privilege enforcement, prompt-injection blocking, a real-time kill switch, and an AI Gateway controlling MCP transactions Strong — the widest discovery reach of anything here; enforcement is coarse by design (block, restrict, shut down), which places it at this layer rather than higher up
Databricks Unity AI Gateway Runtime governance for agents, models, MCP services, and skills, announced June 2026: hard spend caps, smart routing, contextual service policies, and PII and prompt-injection guardrails, with unified tracing across agent activity Strong — extends Unity Catalog from governing data to governing what agents do; key controls were still in beta at announcement
AWS Bedrock Guardrails + AgentCore Policy Input and output filtering, prompt-attack detection, and automated reasoning for hallucination reduction, enforced at the AgentCore gateway perimeter: policy intercepts tool calls and evaluates them before execution, with an audit trail and no code changes Strong — genuinely good enforcement, entirely inside AWS. Its InvokeGuardrailChecks API is detect-only by design: it returns scores, and the application decides whether to allow, block, retry, or escalate
Google Model Armor / Vertex AI safety controls Prompt and response screening, injection detection, and policy enforcement across Vertex-hosted models and agents Strong — same shape as AWS: solid within Google Cloud, scoped to it
Salesforce Einstein Trust Layer Data masking, zero-retention model calls, toxicity screening, and audit logging around Agentforce agents Strong — protects what runs on Salesforce; no view of agents elsewhere
MCP gateway cluster — Kong, Portkey, Cloudflare AI Gateway, TrueFoundry, Aurascape, Runlayer Traffic-layer control on MCP calls Strong — a crowded category worth evaluating as a category, not one entrant

Also worth a look: Zenity, HiddenLayer, Straiker, Noma Security, Pillar Security, Prompt Security (SentinelOne), Protect AI (JFrog), Robust Intelligence (Cisco), Microsoft Defender for AI.

One note on the NVIDIA release: it ships as a reference design rather than a finished product, with Cisco, Microsoft, Oracle, CoreWeave, Dell, HPE, Lenovo, Arm and Intel expected to build products on top of it. It stops rogue behavior at the machine boundary but has nothing to say about business thresholds. It complements Layer 7 rather than replacing it.

6. Building Agents

Owner: Engineering. Architecture and design: how agents are built, what tools they call, how reasoning chains are structured. This layer picks up a new job once you have a portfolio. Which protocols a team builds on is now a governance decision, not only an architecture one. If a team doesn't build on MCP and A2A, nothing above this layer can see what its agents do with each other. An agent that only talks through a platform's proprietary channels is invisible to every cross-platform control you buy, however good that control is.

Vendor What it actually does Weight
Claude Agent SDK (Anthropic) Agent framework from the originator of MCP, with the protocol native rather than bolted on Lead — among the most-used ways teams build agents today; strongest MCP fidelity
OpenAI Agents SDK / AgentKit Agent framework and visual builder tied to the OpenAI model line Lead — the default where OpenAI is already the model provider
Google ADK / Vertex Agent Builder, AWS Bedrock AgentCore, Microsoft Agent Framework Hyperscaler agent frameworks, each tied to its own cloud and model estate Lead — whoever owns the model and the cloud tends to own the framework
LangChain / LangGraph The framework-agnostic option, model and cloud independent Lead — the neutral choice when you don't want the framework picking your provider
Databricks Agent Bricks Describe the agent's purpose, connect the datasets, and it handles build, deployment, and monitoring on top of the governed Unity Catalog layer Strong — fastest path where the data already sits in the lakehouse; inherits governance rather than reimplementing it
Salesforce Agentforce, ServiceNow Agent building inside the system of record the agent will act on Strong — fastest path where the data already lives; ties the agent to that platform
Agent "teammate" products — Claude Cowork, ChatGPT Work, Microsoft Copilot Cowork, Grok Bot Autonomous agents that sign in to company software and run multi-step work until they need approval, deployed by business users rather than engineers Lead — a fast-growing source of agents, and the least likely to pass through any build review
AI coding environments — Cursor, Claude Code, OpenAI Codex, Grok Build Not agent frameworks and not governance products. This is where agents now get written, and the reason a small team can replace a platform module in a couple of weeks Not a control point — but the reason agents increasingly appear outside any platform, built by teams that never filed an architecture review
Covasant (CAMS) Build-orchestrate-govern platform: universal agent registry, drag-and-drop orchestration with A2A support, dev → QA → staging → production gates Strong — its kill switch and PII guardrails are Layer 5 functions bundled in; what it sells is pipeline and architecture
NeuralTrust (TrustGate) Agent Gateway wired in at construction: routes calls to models, scopes each agent's tool access per user, traces agent-to-agent handoffs Strong — build-time choice, distinct from its TrustGuard runtime layer

Also worth a look: CrewAI, Temporal, n8n, Camunda.

Two developments have widened who builds agents, and both land on this layer. AI coding environments like Cursor, Claude Code, Codex and Grok Build have turned a working agent into a two-week project for a small team, rather than a quarter-long one for a platform practice. And the agent "teammate" products let a business user deploy an agent that signs into company systems and works until it needs approval, with no engineering involvement at all. xAI's Grok Bot runs its agents on a shared virtual machine holding their logins and files, which the company itself has characterized in terms of blast radius. Agents in that second category rarely pass through an architecture review, and they're exactly the ones Layers 7 through 9 have to reach.

The Line at Layer 7: Built for Operators, Not Engineers

Everything from Layer 7 up has to be usable by a business person. That's a product requirement, not a preference, and it changes what counts as good for the rest of the stack.

Layers 1 through 6 get bought by people who read traces, write policy in YAML and live in a console. Layers 7 through 9 get bought for a CFO, a renewals leader, an HR director, a compliance officer. The same capability behind an engineering interface isn't the same product to them. It's a product they'll need an engineer to run, which means they don't really own the layer.

That leads to four requirements.

Visual, not query-based. A business owner should be able to see what's happening: which thresholds are firing, where escalations are trending, where a workflow is stalling. If understanding the portfolio takes a query, a trace ID or knowing which platform an agent runs on, the layer has quietly gone back to engineering no matter whose name is on it.

Simple enough to use without training. Not less capable, just simpler to operate. Changing a threshold, approving an exception or adding a rule that spans agents should be as direct as anything else the person does in a business application.

It brings you the insight. The system has to tell the business owner what deserves attention instead of waiting to be asked. No executive is going to sit and watch a feed for anomalies. The pattern forming across agents, the threshold firing too often, the escalation that keeps recurring because a capability is missing: the platform raises those in business terms, or nobody sees them.

Built for how business teams actually work. Actions and changes need approvals, escalations, delegation and a record of who decided what and why. Putting a business-friendly view on top of an engineering tool doesn't get you there, because the workflows underneath are still engineering workflows.

This is why the redeploy test in the next section matters. A product can meet every technical requirement at Layers 7 through 9 and still fail, because the person with the authority to use it can't.

7. Agent Control & Guardrails

Owner: Business. The runtime rules: what an agent can do without approval, what needs a human checkpoint, spend and rate limits, escalation paths. A threshold isn't an engineering setting, even though engineering builds the thing that enforces it. Whether a $2.8 million contract should go to the CFO instead of auto-approving at $2 million is a finance question. Whether an HR onboarding agent needs sign-off above 500 records is an HR question. Both take policy authority and process knowledge that engineering doesn't have and shouldn't be asked to supply. Engineering builds the mechanism. The business sets the number.

One decision, many agents

The one-to-many problem isn't only about who decides. It's about how a decision actually reaches the agents once it's made. Say Legal updates a policy, maybe a new contract clause or a changed regulatory threshold, and it applies to 30 of your 98 agents, spread across however many platforms and teams built them. The usual path: Legal writes the policy, then someone tracks down the 30 engineers behind those agents, explains the change to each one, and hopes all 30 interpret it the same way, build it correctly and test it. That's 30 versions of what should have been one decision, and 30 chances for drift, delay or misreading. That's not governance. It's a game of telephone with compliance risk attached.

What Legal needs is to make the change once, in a system built for legal expertise rather than engineering expertise, and have the platform work out which agents it applies to and put it in force across all of them at the same time. Then they need to see the result in one place: how the policy is performing across all 30 agents' escalations, not spread over 30 dashboards. That's what lets them tighten a threshold that fires too often, loosen one that's too cautious, and keep adjusting until the policy works, instead of shipping it and learning three months later that a third of the implementations drifted.

The cross-agent gap: individually correct, collectively wrong

There's a different failure here, and it's easy to miss. Every agent can work exactly as designed, pass every guardrail and show no drift, and the business outcome still fails, because no agent's guardrails were ever set up to see the other agents.

An example. Agent 1 handles a service request and does it well, but the underlying issue is still open. The case is handled, not resolved. Agent 5 runs the renewal on schedule, doesn't get what it needs from a customer who's still frustrated about that open issue, and cancels the account, following perfectly reasonable renewal logic. Neither agent broke a rule. Neither drifted. Neither would look like a problem on its own dashboard. What the business wanted was "solve the issue, then renew," a sequence that lives between the two agents rather than inside either one.

So guardrails can't only be per-agent thresholds like "don't approve above $X." Some have to span a group of agents that share a customer, a contract or a case. The rule the business needs isn't a property of the renewal agent at all. It's this: don't let any agent in the customer-lifecycle group take a final action on a customer while another agent in that group still has something open for them. Writing that rule is business work. A CX or renewals leader knows the sequence matters, while an engineer building the renewal agent has no way to know Agent 1 exists. And it can't be built by editing Agent 5. It has to be enforced somewhere that can see both agents against the same customer at the same time.

Spend: knowing versus doing something about it

Cost gets treated as one topic when it's two, and conflating them hides where the real lever is.

Knowing what your agents cost is a reporting problem. Per-agent, per-model, per-workflow spend is a dashboard, and most of the platforms in this piece deliver some version of it, ForceEquals included. It's useful and it's table stakes. It is not control.

Doing something about it splits across the stack. Some of the levers are engineering's: which model a given task runs on (Layer 3), how much data gets pulled into context (Layer 2), how the agent is architected and how many calls a task takes (Layer 6). Those are real, and they're where most cost conversations stop.

The lever nobody counts is the business one, and it's usually the largest. The cheapest call is the one that never happens. A business owner looking at an expensive agent isn't limited to asking for a cheaper model. They can change what the agent runs on at all. Run this process for customer type A, never for customer type B, and escalate customer type C to a human for approval. That single rule can take more cost out than any model swap, and nobody in engineering can write it, because it depends on which customers are worth the spend. That judgment lives in the business.

That makes spend control a guardrail, defined at this layer and tuned at Layer 8 like any other threshold. Scope it, watch what it costs and what it returns, widen or narrow it. Treated as an engineering optimization, cost becomes a quarterly project. Treated as a guardrail the business owns, it's something you adjust in an afternoon.

Worth being careful here, because the obvious move is the wrong one. Cutting spend hard looks like a win on the cost line and often isn't: the calls you stopped making were producing revenue too. The setting that minimizes cost is rarely the setting that maximizes result, and the only way to find the point in between is to move the guardrail, watch both numbers, and move it again.

The real-time tuning diagnostic

There's one test a business unit can run today. Take an auto-approval rule that's already live and change the criteria. Can the business owner make that change themselves and immediately watch how it performs against live escalations? Or does it mean finding the engineers who built the agent, filing a request, and waiting through a code change, a redeploy and a release before anyone can tell whether the new criteria was right?

If it's the second, that was never a guardrail. It was application logic with a guardrail's name on it. A real guardrail lives outside the agent's code, which is exactly why it can change without touching the agent.

Put plainly: if changing an approval threshold means redeploying agent code, this layer has already failed, however well the guardrail was designed. The redeploy isn't an inconvenience to work around. It's proof the control plane and the agent were never separated.

Three questions to ask any vendor at this layer:

  1. Can one person who isn't an engineer make one change, in one place, and have it reach every agent it should?
  2. Can that same person see the combined effect without engineering assembling the picture for them?
  3. Can rules that span a group of agents be written and tuned the same way?

Individual agents, and vendors built around per-agent or per-platform configuration, fail all three by design — no matter how good their guardrail mechanics are.

Vendor What it actually does Weight
ForceEquals Single control plane above the agent portfolio (Agentforce, ServiceNow, Workday, Agent 365, custom builds) rather than inside any one of them. Runtime guardrails set and tuned at every scope — individual agent, group of agents, workflow, business unit, or whole organization — so a rule can be as narrow or as broad as the policy behind it. Breaches route to the right human and resolve as closed-loop actions: approval, denial, rollback, or a change request back into the guardrail itself Lead — built around all three tests above
Onyx Guardian Agent enforcing policies written in natural language rather than code Strong — narrows the gap between who writes the policy and who owns the outcome, though policies are authored by security teams rather than business owners directly
Kosmoy, Lyzr, Airia, PointGuard AI, PlainID, NeuralTrust gateway Containment, access scoping, and permission gating (detailed in Layers 4–6) Not enough alone — all fail the redeploy test; they enforce what engineering configured, they don't give the business a tunable surface
Guardrail features inside one agent-building platform Per-platform thresholds and approvals Not enough alone — blind to everything outside their own walls

How thin this layer is. Going through the current market, across gateway vendors, sandboxing vendors, identity vendors, agent security platforms and several products sold explicitly as agentic guardrails, nearly everything turns out to be enforcement that engineering configures rather than a surface the business can tune. ServiceNow's AI Control Tower is the most visible attempt, and it's a useful illustration of the pattern: an incumbent extending what it already had. ServiceNow knew asset inventory and IT service management, so what arrived is discovery, risk scoring, security enforcement, and escalations routed into support queues — strong work at Layers 2 and 5, and the wrong shape for this one, where the escalation has to reach a finance or renewals leader and the threshold has to move without a ticket. That's a finding, not a gap in the research. It isn't a knock on those products either. They do valuable work at their own layer. It's a warning about buying them for this one.

8. Agent Operations

Owner: Business. The day-to-day: observability, drift detection, incident response, performance against SLAs. Same ownership logic as Layer 7, but the direction flips. Layer 7 pushes one decision out to many agents. Layer 8 pulls many agents into a single view. That view is the only place some kinds of drift ever show up, and the only way a person can manage a portfolio this size without pulling their hair out.

Take a workflow that runs across several agents: a contract gets handled by a Salesforce contract agent, then a Workday onboarding agent, then a sync agent reconciling the two. Each one, on its own, may be working inside its guardrails. But drift in a workflow like this usually isn't visible at any single agent. It shows up in the pattern. The contract agent approves slower this week. The sync agent's error rate rises right after. The onboarding agent gets hit with exceptions that trace back to both. No individual dashboard tells that story.

Be clear about what escalations are, because it changes how the business should handle them. They aren't failures. An agent that hits a guardrail, surfaces an edge case or flags something ambiguous is doing its job: deferring instead of guessing. Treat escalations as an incident log to clear and you'll clear them and learn nothing. Treat them as training signal and every one of them improves the next 39 agents, not just the one that raised it.

This layer is one-to-many in five distinct ways:

  1. One place to pay attention — a business owner watching one feed instead of 40 separate queues and dashboards.
  2. One place to capture decisions — every escalation resolved and recorded consistently, instead of scattered across 40 tools, inboxes, or nowhere at all.
  3. One place to trigger action back to the right agent — without the business owner needing to know which of 40 systems that agent lives in.
  4. One place to ask any question — across every agent at once, rather than reconstructing an answer from 40 separate silos and the people who know how to query each one.
  5. One place to learn — patterns, edge cases, and lessons accumulating in a single record that every agent benefits from, instead of 40 histories nobody reads together.

Miss any of the first three and the rest don't matter much. Watching closely without a consistent record of decisions gets you well-observed chaos. A clean record that never routes back into action is documentation of a problem nobody fixed. And without the last two, the portfolio never gets smarter. It just gets watched.

The business team, not IT or engineering, should run guardrail control from that one place. The value isn't only visibility. It's making a small adjustment, watching how it performs against real escalations across the workflow, and adjusting again, with the person who understands the process doing it rather than explaining the intent to an engineer and waiting for a release. You rarely get a threshold right the first time, so what matters is how fast you can see the effect of a change and move again.

Human-in-the-loop is a dial, not a setting

Every escalation is a human-in-the-loop interaction, and how often they fire is something the business should manage rather than accept as given. This is the part most easily missed about this layer. It isn't only about catching problems. It's how a portfolio earns its way toward more autonomy.

Start conservative. An agent handles a category of work, a person reviews most of it, and a record builds up. Once that record shows the agent making the right call consistently in a given situation, the business owner raises the threshold: fewer interventions and more autonomy, in the specific area where trust has been earned. Where the record shows the opposite, they tighten it. Trust gets extended where the evidence supports it, rather than granted wholesale at launch or withheld forever out of caution.

This only works if the dial belongs to the business. A threshold that takes an engineering cycle to move won't move. It sits wherever it launched, and the portfolio never delivers the efficiency it was bought for. That's the worst of both worlds: agents doing the work, people reviewing all of it indefinitely, and nobody with both the authority and the tools to change the ratio.

Run well, this is where the business drives results directly: tightening where errors appear, loosening where performance is proven, and reading the impact of each move against live escalations, with no engineering attention on any of it.

Vendor What it actually does Weight
ForceEquals Unified portfolio view: one feed across platforms, cross-platform anomaly detection, and an ask-anything interface over the portfolio's decision history. Every human-in-the-loop interaction and escalation runs through one consistent model, so the business manages them the same way across every agent. Intelligent routing sends each one to the right expertise and authority dynamically, resolved from the specific combination of agent, workflow, guardrail, decision, and action required — not a fixed queue — so the right person responds fast, from one place, without a dashboard that assumes SQL or trace-ID fluency Lead — built for all five of the above
Wayfound Business users supervise, evaluate, and optimize agent fleets from one dashboard; ingests telemetry regardless of source platform Lead — vendor-agnostic by design, strongest in Salesforce/Agentforce-heavy shops
Mezmo AURA Autonomous incident response: investigates and remediates production incidents across the agent portfolio rather than only flagging them Strong — a step past observability, narrower than a full ops plane
ServiceNow AI Control Tower Inventory and monitoring across agents on 30+ systems, with real-time alerting and cost measurement; escalation runs through ServiceNow's IT service management workflow Not enough alone — an IT operations view, routed to IT queues and IT owners; the decision record is a ticket, and tuning runs through a change process
Dynatrace, New Relic Enterprise observability and APM extended to AI workloads: traces, dashboards, alerting, and historical analysis across the stack Not enough alone — tells you what happened and when, but stops at the alert; no intervention, no decision record, no route back to the agent
Per-agent or per-platform observability Deep telemetry on one agent or one platform Not enough alone — fails all three counts; the failure mode lives in the gaps between agents

A note on the observability vendors. Dynatrace and New Relic are genuinely good at what they do, and if you already run one, the agent telemetry will land there. But watching is not the same as operating. They produce dashboards, alerts, and history — someone still has to notice the alert, decide what to do, make the change somewhere else, and remember why. This layer is about intervening and resolving, not reporting. An alert that nobody can act on from the same place is a record of a problem, not a response to one.

Two we left out. Kore.ai is a conversational AI and contact center platform — one of the platforms this layer governs, not a layer above them. PointGuard AI's Agent Mission Control is containment and access control, so it sits in Layer 5.

9. Agent Change Management

Owner: Cross-functional. The business decides, engineering ships. This layer works differently from 7 and 8. The business doesn't own it outright. Business and engineering each own half, and the handoff between them is where the governance risk lives. Deciding a guardrail needs to change, or an agent needs a new capability, is business judgment. Building it has always been engineering work, and that's the assumption a new class of purpose-built change management tools is starting to break. The more of the path from decision to shipped change a system can carry itself, the less the handoff can drop, and the smaller the piece that has to wait in an engineering queue.

This is the same one-to-many idea, applied to learning instead of policy or escalation. Every agent's failures, near misses and edge cases are lessons. But if those lessons sit in 40 separate silos, each team only learns from its own agent. The pattern that only appears across the whole portfolio, the fix that would help a dozen other agents hitting the same edge case, never surfaces, because nobody is positioned to see it. Learning stays local even though the agents never were.

This is also where new ideas come from. A recurring escalation is often a sign that a new agent, or a new capability on an existing one, is needed, and the business sees that pattern before anyone else. That signal is only useful if there's one place to capture the idea and the context behind it: which agents were involved, which customers or cases, what the impact was, and where it ranks against everything else waiting to be built. Without that, the insight surfaces in a queue, gets resolved once, and disappears.

A further level of automation becomes possible here, but only if Layers 7, 8 and 9 run as one system rather than three tools. Building a business case and an approval flow without a person doing it by hand requires already knowing which guardrail was breached (Layer 7), the escalation pattern that surfaced it (Layer 8), and what else is competing to get built (Layer 9). Split across three tools, that becomes three manual handoffs, each one a place for the idea to stall. Run together, it becomes a single path from insight to a project that's ready to build.

Vendor What it actually does Weight
ForceEquals Applies guardrail rule updates automatically, driven by escalation intelligence rather than a ticket and a release cycle. Turns insights and improvement opportunities into context-rich projects on its own: expanding the escalations, agents, customers, and business impact behind an idea, then producing requirements, the business case, stakeholder alignment, prototypes, and production-ready code — all staged and waiting the moment approval clears Lead — the only vendor found closing both halves, and the only one automating the path between them
SailPoint / Astrix Lifecycle governance: creation-to-retirement service accounts, short-lived just-in-time credentials by default Strong — makes decommissioning a policy default rather than manual cleanup, but credential-scoped only
Guardian-agent and observability vendors Solve the business-decision half: escalation, approval Not enough alone — stop at the handoff
Jira, Asana, Monday.com Solve the engineering half: delivery Not enough alone — receive the work with no governance context attached

The open gap. Very few vendors close the loop between those last two rows, and that's where governance intent gets lost. A CFO approves an exception in a meeting, and six months later nobody can reconstruct why the guardrail changed. Nothing we found closes it the way ForceEquals does.

On the handoff itself. Where work does still need engineering, ForceEquals turns the escalation into a change request and delivers it into whatever tooling the team already uses, whether that's Jira, Asana or Monday.com, traceable back to the business decision behind it. Useful plumbing, but the smaller part of the story now that most of the path ahead of it is automated.

Compliance & Audit — The Cross-Cutting Layer

Owner: Business/Legal/Compliance. Compliance isn't a final step at the top. It cuts through all nine layers, and with many agents the burden of proof gets harder, not just bigger. "Show the regulator what this agent did" becomes "show the regulator how this chain of agents across these platforms reached this outcome: which agent handed off to which, who approved what, and what changed as a result." An audit tool that only sees one platform's logs can't answer that, however complete those logs are.

Vendor What it actually does Weight
ForceEquals Cross-agent, cross-transaction, cross-decision querying: which agents touched a transaction across platforms, what each decided, who exercised human-in-the-loop judgment at which step, and what downstream changes resulted Lead — traces an approval through to the guardrail change it triggered, where most tooling logs the decision but not its consequences
Collibra Cross-platform lineage from source data through inference and usage Lead — the data-side half of the same question
Credo AI Agent Registry spanning internal and third-party agents, risk assessments, human-oversight intervention points Strong — no runtime enforcement yet: no gateway, no inline guardrails, no containment, per its own 2026 roadmap disclosure
Vanta First major GRC vendor with a dedicated ISO 42001 module, and ISO 42001-certified itself Strong — pragmatic route for teams extending SOC 2/ISO 27001 machinery into AI
Holistic AI EU AI Act, NIST AI RMF, and ISO 42001 coverage together, plus Runtime Agentic Monitoring added April 2026 Strong — moved from documentation into runtime evidence
Modulos AI governance platform that has itself completed ISO/IEC 42001 conformity assessment (CertX, May 2026) Strong — the pick when the platform itself must be certified and self-hosted
Kosmoy (runtime evidence) Gateway logs proving a policy was actually enforced, not just documented Strong — closes the exact gap Credo AI's roadmap flags
IBM OpenPages Enterprise GRC incumbent, with AI risk and model governance folded into an existing risk register Strong — the default where OpenPages already runs the risk program
Monitaur Links written policy to captured evidence across the model lifecycle, built around the audit and filing expectations of regulated industries Strong — strongest where a regulator will actually ask; model-centric rather than agent-centric
Single-platform audit logs Thorough logging within one platform's boundary Not enough alone — can't reconstruct a decision chain that crossed platforms

Also worth a look: Drata, AuditBoard, LogicGate, Relyance AI.

One requirement runs through this whole layer: compliance and legal have to be able to query the portfolio themselves, without routing every question through engineering to pull and interpret the data for them.

Control Ownership at a Glance

Layer Primary Owner Why
1. Foundational Considerations Business (exec leadership) Risk appetite and accountability are leadership decisions by definition
2. Data Governance IT/Data Engineering + business data owners Implementation is technical; sensitivity classification is a business call
3. LLM Layer Engineering/ML Model selection and evaluation is technical work
4. Identity & Access IT/Security Engineering Credential and access scoping is a security engineering discipline
5. Security IT/Security Engineering Attack-surface defense is a security engineering discipline
6. Building Agents Engineering Architecture and protocol choices are engineering decisions
7. Agent Control & Guardrails Business Thresholds and sequencing rules require policy authority and process expertise engineering doesn't have — and the system enforcing them has to be usable by business teams, not technical ones
8. Agent Operations Business Judging urgency and impact takes business context, not telemetry literacy — and the ongoing work here is business work: moving the human-in-the-loop dial as trust is earned, routing escalations to the right authority, and tuning thresholds against live results without an engineering cycle
9. Agent Change Management Cross-functional Business decides what should change, and the handoff to engineering is where intent gets lost — so the goal is to automate as much of the path as possible: guardrail updates applied straight from escalation intelligence, and ideas developed into requirements, business cases, and ready code before anything reaches a queue
Compliance & Audit Business/Legal/Compliance The function answering to regulators needs to query the agent portfolio directly, not through engineering

The dividing line here isn't arbitrary. Layers 1 through 6 fail if IT and engineering don't own them. Layers 7 through 9 fail differently — they fail if IT and engineering are asked to own them alone, because the authority those layers require doesn't live in those functions. That's why the tooling question at Layers 7 through 9 isn't just "does it work," it's "can the person with actual decision authority use it without an engineer in the loop" — visually, proactively, in a system built for how business teams actually work. That second question filters out most of the current market at Layer 7 specifically: plenty of vendors clear "does it work," very few clear "can the business run it directly."

Read the Market by Layer, Not by Logo

Look back at the coverage map near the top of this piece, now that the layers mean something. The bars stack up across the bottom six layers and thin out above them, and that isn't an accident of who we happened to include. There's a pattern in the tables worth naming, because it changes how you should shop at each level of the stack.

Layers 1, 2, 4 and 5 belong to incumbents, and that's the right answer there. Risk management, data governance, identity and security are decades-old disciplines. Collibra already knew lineage. SailPoint and Okta already knew identity lifecycle. Palo Alto, Zscaler and Microsoft already knew traffic inspection and access control. Agents gave them a new object to point existing expertise at, not a new problem to solve from scratch. Extending a mature discipline to a new kind of actor is real work, but it's extension. At these layers the incumbent is usually the safe pick, and the specialists worth buying tend to be the ones the incumbents are acquiring.

Layers 3 and 6 belong to the labs and the clouds, with one carve-out. The model providers own Layer 3 by definition, and they increasingly own Layer 6 too, because whoever makes the model and runs the cloud ends up making the framework. Anthropic, OpenAI, Google, Microsoft and AWS are all shipping agent SDKs alongside their models. The carve-out is evaluation, which went to specialists like Arize, Galileo, Braintrust, Arthur and Fiddler, because tracing a chain of calls across several agents and models from different vendors had no natural incumbent to inherit it. At Layer 6 the governance question is less about which vendor than whether what you build on speaks MCP and A2A cleanly.

Layers 7, 8 and 9 have no native incumbent, because nothing here existed before. There's no prior discipline to extend, which hasn't stopped incumbents from trying. ServiceNow has pushed AI Control Tower across 30-plus enterprise systems, but what arrived is what ServiceNow already knew how to build: discovery, inventory, risk scoring and escalations routed into IT service management. Real capability, aimed at the wrong owner. A guardrail that spans a group of agents sharing a customer. An escalation feed routed to a CFO by business role rather than trace ID. A threshold a renewals leader moves as trust is earned. None of these is a version of something that was already being sold. That's why the vendors appearing at these layers are purpose-built, and why the reengineered entrants are easy to spot: containment and access control from Layer 5, with business-facing language on top and an engineering console underneath.

That inverts the usual buying instinct. At Layers 1 through 5, incumbency signals depth. At Layers 7 through 9, it often signals a product designed for a different problem. Buy the incumbent at the bottom of the stack. At the top, ask what the product was built to do before agents made this a category.

One caveat over all of it: this market is consolidating in real time. Three leading identity vendors changed hands inside a single three-month window in 2026: Cisco and Astrix, SailPoint and Entro, Cyera and Oasis. Palo Alto closed on CyberArk in February of the same year. Several vendors in the tables above will belong to someone else by the time you finish an evaluation, so confirm roadmap and support commitments directly rather than trusting any point-in-time comparison, this one included.

NVIDIA's Open Agent Safety Platform, announced in September 2026, shows how unsettled even the mature layers still are. A company with NVIDIA's resources shipped free, open-source containment because nobody had solved the hardware-level boundary problem on their own.

Where This Is Heading

Two forces are pushing agent building outward, away from any single platform, and both raise what the nine layers have to reach.

Multi-vendor is the steady state, not a transition. It's tempting to assume the sprawl is temporary, and that a portfolio spread across Salesforce, ServiceNow, Workday and custom builds will eventually consolidate onto one platform. It won't. Agents get built where the work and the data already live, by the teams closest to that work. The service agent belongs near the service platform, the onboarding agent near the HR system, the finance agent near the ledger. That's not disorder to be cleaned up later. It's the sensible outcome of building agents where they're needed, and it means any governance solution has to see and act across a landscape it doesn't own.

That's structurally harder for a platform vendor than it sounds, and the difficulty isn't only technical. A platform's governance tooling is built around its own execution model, its own logs and its own identity system, meaning the things it controls. Extending equal fidelity to a competitor's agents means investing in visibility into a product it would rather you replace. Some will do it well anyway, because customers will insist. But the incentive runs against neutrality, and it's worth asking any platform vendor a direct question: what exactly can you see and change on the agents you didn't build?

There's a second force pushing the same direction: repatriation. Teams are starting to look at expensive, seat-priced platform modules and ask whether an agent could do the same job. Sometimes the answer is yes, and the replacement gets built in a couple of weeks rather than a couple of quarters. We've spoken with two Salesforce customers who have already done exactly this with CPQ: the module swapped for a purpose-built agent that handles their actual quoting rules, costs a fraction of the per-user license, and is easier to change when those rules change. Not every module survives that comparison, and the ones with deep compliance, tax or integration requirements often should stay put. But enough of them don't survive it to matter.

The economics are hard to argue with. Platform licensing charges per user for functionality a company may use a fraction of, while the cost of building a focused agent keeps falling. The moment a business unit can justify replacing a line item with something it controls, it will try, and every successful attempt makes the next one easier to approve. The result isn't fewer platforms. It's more places where agents get built, by more teams, with less central coordination than any of these vendors assume. Democratized building, on a cost argument nobody in finance is going to override.

For the nine layers, that raises the bar in both directions. Governance has to reach agents built by teams that never filed an architecture review, and it has to do that without slowing down the thing that made the build worth doing. A control plane that only covers agents built on approved platforms by approved teams will govern a shrinking share of the portfolio every quarter.

The Bottom Line

Governance isn't one control you add. It's a dependency chain. Weak data governance undermines your guardrails. Sloppy identity makes security theoretical.

That's the part worth deciding before you need it. Layers 1 through 6 get you agents that work. Layers 7 through 9 are what let you run them: tuning thresholds as trust is earned, catching the failures that live between agents rather than inside them, and turning what you learn into the next change without waiting on a release cycle. Most organizations build the bottom six first and discover the top three under pressure, after something has already gone wrong. The portfolio is easier to govern while it's still small enough that governing it feels premature.

Turning pilot deployments into scaled, sustained results takes all nine layers. There's no compromise version of this and no partial credit. A portfolio missing any one of them will work right up until it doesn't.

We Got Things Wrong

We missed vendors. We put some in the wrong layer, and we were probably too generous in places and too hard in others. This landscape changes faster than any single read of it can keep up with, and we'd rather be corrected than consistent.

So tell us. If your product is in the wrong layer, if we described it inaccurately, if we left you out entirely, or if you think a whole category is missing, we want to hear it and we'll update the piece. Write to me directly at marc@forceequals.ai.

The broader point is that this industry doesn't yet share a vocabulary for any of this. The same word means three different things depending on who's selling it, and buyers are left comparing products that aren't comparable. A common set of layers and a shared understanding of what belongs where makes every conversation in this market more useful, for buyers deciding what they need and for vendors explaining what they actually built. We'd rather help build that than be right about our version of it.

Marc Chabot

Co-Founder & CEO, ForceEquals

marc@forceequals.ai

‍