AI Agent Governance, Trust, and Control
Part of Self-Improving Defense: AI Cyber Defense Against AI Attacks — read the full guide.
Who is accountable when an AI defense makes the wrong call?
Whoever approved the action is accountable, and that's the main reason every action should start out gated. The AI defense proposes, a person approves and the record shows who approved what, on which evidence. It works pretty much like a junior analyst's recommendation that a senior engineer signs off on. If you've chosen to ungate an action type, the accountable person is whoever approved lifting that gate, so gate changes should be recorded like any other approval.
The roles around this are shifting. A lot of teams now separate the person who writes a policy or skill from the person who approves it (author is not approver) so nobody can create a change and wave it through themselves. That usually ends up with something like an AI SecOps Manager, who owns rollout, governance and service levels for the program.
The approval numbers are worth knowing too. Across Simbian deployments, 95% of the responses Simbian proposes are approved by the customer's own analysts. Their call, not ours. In the other 5%, we get additional context that improves our next response. That number describes who made the decision. It isn't a guarantee of any outcome and the decision stays with your team.
Which actions should an AI defense take without human approval?
The ones you have decided are safe, one action type at a time. A sensible default is to start with every action gated and lift gates as confidence builds, so closing a false positive can probably run on its own fairly early while isolating a production server waits for you, and the line moves as the track record grows.
The UK NCSC suggested a useful way to score this in a September 2026 blog post, One does not simply defend agentically. It rates how risky an automated defensive action is on five dimensions, each scored 0 to 4, which are potency, scope, criticality, rollout confidence and recoverability.
At the bottom of the potency scale the AI only gives a human advice it can explain. Investigating an alert to a verdict sits roughly there, which is why that part can usually run from day one while response actions climb the scale. Something easy to reverse and narrow in scope that doesn't touch anything critical, like closing a duplicate alert, is a good candidate to run without approval. Something that hits many systems or touches identity or production endpoints or is hard to undo should generally stay gated longer.
Two signals help most when deciding to lift a gate. One is how often your analysts override the proposed response for that alert type. The other is whether the action can be reversed at all. Low override rates on reversible actions are a fair reason to ungate, and anything acting against an identity or an endpoint (suspending an account, revoking a privilege) usually takes longer to earn it. Even then an ungated action still has a human on the loop who can see the record and reverse it.
What should a reviewer check before approving an AI's change to its own skills?
Before approving, look at why the change got raised and how many investigations it touches, the past cases that show the gap, the exact before and after, and whatever it contradicts. A proposal to change the defense's own skills shouldn't ever arrive as a bare request, so all of that should already be on the page:
- Why it was raised: which gap it addresses, and roughly what share of investigations it will change.
- The evidence: links back to the past investigations that showed the gap, each one with a short note on how the change would've changed that investigation
- The before and after: the exact diff, so you see what gets written.
- Conflicts: anything already in context that the proposal contradicts, shown side by side.
Some proposals you can reject straight away. Blanket rules like "ignore all PowerShell alerts". Allowlists that switch off investigation. Anything generalized from one incident, and temporary exceptions with no expiry on them. The expiry is the one people skip. A red team exercise gets added as "expected activity" for two weeks, the exercise ends, nobody removes the entry, and from then on real alerts on that asset quietly get discounted. A well designed system shouldn't propose verdict-automation rules or allowlists in the first place.
When a proposal conflicts with existing context there are usually three ways to go: keep what's there, accept the proposal, or edit and merge. If you reject it, say why in a comment, because the next proposal can read earlier rejections and adjust. After approval the change should get validated once it's applied, date-stamped so stale facts get flagged later on, and covered by regression testing. A conflicted request is never applied until a human resolves the conflict.
Would you trust an AI-generated verdict without seeing the evidence?
No, and you shouldn't have to. A verdict should come with its evidence, the queries that ran, process IDs, logs pulled and the reasoning connecting them, so an analyst checks it the same way they'd check a colleague. Each query should be re-runnable in the source tool so the evidence gets checked against your SIEM and not against the AI's summary of your SIEM. A good verdict also says why the investigation decided it was done, not just what it looked at.
In-context learning helps with this. What the defense knows lives in readable skills and context, outside the model weights, so it can point at the exact context behind a verdict. Say a fact like "this service account runs encoded PowerShell nightly" shaped the verdict, the investigation can show that fact and link to the request that added it.
Early on, try a simple spot-check. Hand an analyst only the evidence the verdict shows and see if they land on the same call. If they don't, the evidence isn't doing its job yet however confident the verdict sounds.
Can you stop, reverse, and reconstruct what an AI defense did?
You should be able to do all three and if you can't, the defense generally isn't ready for production. Stopping comes down to gates and permissions you control. Reversing means actions can be rolled back and there's a record of the state things were in before. Reconstructing needs a full record of every action and query and the reasoning behind each, plus who approved it and which identity and permissions the agent was acting under, kept in a log the agent can't edit itself.
In practice that means four properties on every action the defense takes. It's sandboxed, so everything runs isolated. It's permission-controlled, every API call and every access. It's human-gated, any output can require approval first. And it's yours to extend, you can add skills and context and steer it.
Same goes for changes to the defense itself. Every change to what it knows should keep its full history (who raised it, who approved it, when) with a before and after diff, so a bad change can be tracked down and undone.
Blast radius is a useful thing to ask about when evaluating: what's the most damage this agent could do before someone stops it? If nobody can say, the permissions are probably too broad or there aren't enough gates. Test the stop before you need it.
What does AI agent governance look like in security operations?
AI agent governance in a SOC mostly looks like the governance you already apply to privileged human access, adjusted for agents that act fast and at scale. Each agent should have a named owner, a defined scope and explicit permissions, plus a review cadence, an expiry on any temporary privileges and some tested way to switch it off. In a SOC those translate roughly like this:
| Governance element | What it means in a SOC |
|---|---|
| Owner | A named person accountable for the agent's behavior |
| Scope | Which alert types, systems, and data the agent may touch |
| Permissions | Which actions run alone and which wait for approval |
| Review | Regular review of overrides, rejected proposals, and errors |
| Change control | Every change to the agent's skills shown as a diff and approved |
| Kill path | A tested way to stop the agent and reverse what it did |
Several frameworks are landing on similar ideas. Forrester's AEGIS framework lays out 39 controls across six domains (approval gates for irreversible actions are one of them) and maps them to the NIST AI RMF, ISO 42001, the EU AI Act and MITRE ATLAS. In May 2026 CISA, the NSA, Australia's ACSC and partner agencies published joint guidance on careful adoption of agentic AI services, and the UK NCSC's scoring of automated defensive actions gives you a way to decide which actions need a human first.
The organizational side matters about as much as the controls. Governance of security agents rarely sits with the SOC alone. The AI SecOps Manager usually runs it day to day, but risk appetite tends to get set together with IT, legal, privacy and compliance, and AEGIS itself calls for a governance board drawn from those teams.
Is an AI investigation deterministic or non-deterministic?
The reasoning is non-deterministic, meaning the same input won't always give the same output, and the policy around it has to be deterministic. Language models are probabilistic so two runs of one investigation can take slightly different paths, teams that drop AI steps into rigid workflows often see that as inconsistent results. Fair concern.
You deal with it by putting whatever has to be consistent into policy that's enforced deterministically. In Simbian that policy is written as skills in plain language, and covers four kinds of rules: risk tolerance and escalation, investigation playbooks, asset criticality and identity tiers, change-freeze and regulatory rules. The model reasons inside those. It doesn't get to decide if a change freeze applies.
The output is what you hold to a consistent standard. Every case documented the same way, same evidence fields, and override rates by alert type tell you if the verdicts are consistent where it counts.
Can attackers manipulate an AI defense?
Yes, they can try. The main ways in are pretty well known by now. One is prompt injection, instructions hidden in data the AI reads like an email body or a file name or a log line. Another is poisoned context, where an attacker gets false "normal" facts into what the defense believes about your environment. There's evasion too, shaping malicious activity to look like something the defense learned is normal (encoded PowerShell on a host where that's expected, say). And then the human gate, if reviewers approve proposals without actually reading them.
Each of those needs its own control. The model layer needs hardening against injection, which is what Simbian's TrustedLLM™ is for. It's private and prompt-injection hardened, hardened against data poisoning in the war lab, and no customer data goes into training.
Evasion is handled by scoping. A learned exception says when it does not apply, wrong time of day or wrong parent process, and anything outside that scope gets investigated in full. Context only changes through tracked requests that a person approves, writes stay scoped to your tenant, and a conflicted request is never applied. Review has to be real as well, which means watching for rubber-stamping, keeping proposals small enough to actually review and attaching evidence to each change.
No AI defense is immune to manipulation, same as no human analyst is immune to social engineering. The aim is making manipulation visible and hard to make stick.

.png&w=3840&q=75)