Self-Improving Defense: AI Cyber Defense Against AI Attacks
Key takeaways
- Self-improving defense is AI cyber defense that gets better without anyone rewriting its rules. It learns general defensive skill in a war lab where both the attack and the defense are authored, applies it to what is specific to your environment, and proposes changes to itself that a human approves. In October 2026 testing, no frontier model on its own passed the Cyber Defense Benchmark.
- AI-powered attacks change two things. Volume goes up, and the specific attack stops repeating, so there is often no signature, playbook, or patch waiting for it. Techniques still recur, which is why reasoning about method works where matching does not.
- A self-improving defense learns in context, not by retraining on your data. Your data stays in your tenant, a better model can be swapped in the day it ships, and every change to the defense is shown as a diff and approved by a person.
- Trust is an action boundary. Every action starts gated and you lift gates one action type at a time, with an audit record of what was done and why.
- Improvement has to be measured, not asserted. Coverage of replayed MITRE ATT&CK techniques, time from detection to response, and how often analysts approve the proposed response are more useful than alert counts.
On this page
AI Cyber Defense Fundamentals
What is self-improving defense?
Self-improving defense is AI cyber defense that gets better at defending without anyone rewriting its rules. It writes its own skills, chooses which tool to query and when, and proposes changes to itself that a human approves before they apply. General skill comes from a war lab, and everything specific to you comes from your own environment.
The full argument for the category is on our self-improving defense page. This guide stays on the mechanics: how a defense like this learns, who approves what it learns, and how you would check that it is actually improving. Most of it comes back to a four-node loop across the AI Pentest, AI Threat Hunt, AI SOC, and AI Detection Engineering Agents.
How is self-improving defense different from adaptive security?
Adaptive security usually adjusts how strongly a defense reacts, while self-improving defense changes how the defense does its work. Gartner's adaptive security architecture, written up by Neil MacDonald and Peter Firstbrook, describes four stages (predict, prevent, detect, respond) running continuously, and most adaptive security products apply that idea by tuning thresholds, risk scores, or access decisions as new signals come in.
That approach generally assumes a threat shows up often enough to build a baseline and score deviations against it, which works fine for things like risky logins or unusual data movement. It works a lot less well when an attacker builds something new for one target and only uses it once, there is no history in that case for the thresholds to learn from.
A self-improving defense changes its own method, and only after a person approves the change. When investigations of the same alert type keep stalling on the same gap, such as a check that is missing or a query that keeps failing, the defense writes the fix to its own skills and submits it for review. Once it is approved, the next investigation runs differently because of it. The thresholds can stay where they are, its the investigation that gets better.
One practical note, "adaptive security" is also used as a company name in the security awareness market, so a lot of what you find under that term is about phishing training rather than security architecture.
What is the difference between self-improving defense and self-learning AI?
In security, self-learning AI usually means anomaly detection. The system studies your environment until it knows what normal looks like and then flags whatever deviates from that. Darktrace built its brand on the term "Self-Learning AI", and the same general approach shows up across network detection and user behavior analytics tools.
The cost of that approach is the learning period. Self-learning products commonly need weeks to months of baselining before they are fully effective, and noise during that window is a common complaint in user reviews. Anomaly detection also tells you something is unusual, it does not tell you whether the unusual thing is an attack, so somebody still has to investigate.
Self-improving defense starts from the investigation. It arrives already trained on how to investigate, so it produces value from the first case, and what it learns from your environment changes how it investigates and responds next time. The three terms are often used interchangeably, so it helps to compare them on what each one actually changes:
| Adaptive security | Self-learning AI | Self-improving defense | |
|---|---|---|---|
| What it learns | Risk signals and context | What normal looks like in your environment | How to investigate and respond, plus what is specific to you |
| What changes over time | Thresholds, scores, access decisions | The baseline | The defense's own skills and context |
| Time to first value | Depends on policy tuning | Often months of baselining | Usually the first case |
| Who approves a change | Policy owner | Usually automatic | A human approves each proposed change |
| Main output | A risk decision | An anomaly flag | An investigated verdict and response |
Strictly speaking, picking up a pattern nobody handed you is self-learning. It becomes self-improving when that learning demonstrably changes the next outcome and the loop keeps closing, a correction becomes context, the context changes the next verdict, and the next verdict feeds the next correction.
Is self-improving defense the same as the self-improving AI people are worried about?
No. The self-improving AI that worries researchers is usually recursive self-improvement, where a model changes its own weights or training without anyone checking, and each version builds a more capable successor. Anthropic and the Cloud Security Alliance have both written about the security risks of that idea, and the concern is reasonable.
Self-improving defense doesn't change any model's weights. What gets better is the skills, the context and the system around the model, through in-context learning. When the defense finds a gap that keeps coming back, it writes up a proposed change, shows the before and after as a diff, flags anything that conflicts with what it already knows and then waits for a human to approve or reject it. Once an objective reaches full marks it moves into a regression suite, so improving one thing doesn't quietly break something else that already worked.
The short version for a board or an auditor is "self-improving, not self-driving." The defense finds its own blind spot and writes the fix, a person decides whether the fix ships.
What is the difference between proactive and reactive cybersecurity?
Reactive security is basically responding to whatever fires, an alert comes in, somebody investigates it and the team responds. Proactive security goes looking before anything fires at all, through threat hunting and penetration testing, and by checking if your detections would actually catch a given MITRE ATT&CK technique.
Most teams know they should do more of the proactive work. Most teams rarely get to it. In a lot of SOCs that comes down to two queues that never empty, alerts keep coming in faster than anyone can investigate them and patches keep getting published faster than anyone can apply them, so the hours that were meant for hunting and testing go to the backlog instead.
Self-improving defense treats reactive and proactive work as one loop, so the two stop fighting over the same people. A SOC investigation that finds a real threat turns into a hunt for the same activity somewhere else. Whatever the hunt finds becomes a detection. If alerts keep piling up on one application, that can kick off a pentest of that application. The proactive work ends up happening as a side effect of the reactive work and nobody is waiting around for a free week.
What is preemptive cybersecurity?
Preemptive cybersecurity is Gartner's term for security that acts before an attack lands. In a September 18, 2025 press release, Gartner described preemptive solutions as ones that "use advanced AI and machine learning (ML) to anticipate and neutralize threats before they materialize", and predicted they would account for half of IT security spending by 2030. Gartner also uses "autonomous cyber immune system" for what it calls the ultimate evolution of the idea.
Preemptive cybersecurity and self-improving defense overlap in what they are trying to do. Both act on exposure and early signals before an attacker gets to their goal, for example closing off an attack path that a pentest proved was reachable before anybody gets the chance to exploit it.
The difference is what each term describes. Preemptive cybersecurity is an analyst category that groups many kinds of products, such as exposure management, deception, and predictive threat intelligence. Self-improving defense describes a running loop, one that investigates, responds, tests, and changes its own method with human approval. A product can be preemptive without improving itself, and Gartner's definition is framing here, not an endorsement of any vendor.
What happens after an AI SOC closes the alert?
In a lot of AI SOC tools, not much. The alert gets investigated, a verdict comes back, a response is taken or recommended and the ticket closes. That's a real step up from an alert nobody ever looked at,but the environment is exactly the same the next morning and the same alert will probably fire again.
AI SOC is one discipline within self-improving defense, the part that turns alerts into verdicts, and the loop is about what happens after the verdict. Self-improving defense closes the loop, not just the ticket. One alert can end up improving a few different parts of your security depending on what it turned out to be:
- A real threat: you get the response, and you also get a tighter detection that should catch the same activity earlier next time, plus the path the attacker came in through gets closed.
- A false positive: that detection gets tuned so the same false alarm stops coming back.
- Lots of alerts clustering on one application usually means the application gets pentested, and if an exploit is proven it turns into a detection rule and a patch ticket.
- When companies like yours are getting breached, a hunt runs on that hypothesis against your own data to see if it already happened to you.
When buyers say AI SOC is "not enough" this is mostly what they mean. Clearing a queue faster is incremental. A queue where each item leaves the environment a little harder to attack, one fewer exposed path or one less noisy rule for example, is a different kind of result.
If no alert fires, does that mean you are safe?
No. When your detection stack is quiet it just means nothing matched a rule, which isn't the same thing as nothing happening. A quiet night could mean there was no attack at all, or the attacker did something none of your rules cover, or the rule that should've fired is broken.
It's a bit like the night watchman problem. He tells you in the morning nothing happened and you just have to take his word for it, you can't tell if he was awake. Detections that haven't fired for months are the same, some of them are fine and some broke quietly when a log source moved or a field got renamed or an upgrade took out an integration, and nobody noticed because nothing fired. Going by Simbian's own 2026 analysis, roughly 20% of SOC detection rules stop firing within six months, and the reasons are usually boring ones like telemetry drift and schema changes.
Most security programs quietly make four assumptions that don't really hold up, and self-improving defense starts by rejecting them: that an alert is a threat, that no alert means you are safe, that a CVE is the same as an exploit, and that patching every CVE makes the code secure. Each of these is a reason to go and check with a hunt or a test.
Can AI really defend against AI-powered attacks?
Yes, with a condition. Pointing a general-purpose model at your SIEM is not enough because that model was never trained to defend, and on Simbian's Cyber Defense Benchmark no frontier model passes. Models got good at attacking because attack attempts grade themselves (you either got in or you didn't), nothing grades a defense the same way.
You can measure this. Simbian's Cyber Defense Benchmark (arXiv 2604.19533) has frontier models investigate real attack logs, and as of October 2026 none of them pass, passing meaning at least 50% coverage on every MITRE ATT&CK tactic. Of the 30 models tested across 1,314 runs, the best was Opus 5 at 45.1% coverage.
What closes the gap is three things together. First, defensive training data that has to be manufactured, because it largely does not exist in the real world. Second, the context of your own environment: your assets, identities, processes, and the decisions your team has already made. Third, a loop that improves the defense from what each case teaches, with a human approving every change to the defense itself. AI cyber defense works against AI-powered attacks when it has all three, and it generally struggles when it has only a model.
When attackers and defenders both use AI, who wins?
Neither side wins permanently, but defenders have advantages they often under-use. Dave Chismon, the UK NCSC's CTO for Architecture, put the attacker's advantage bluntly in a September 2026 blog post: "Defenders simply cannot put AI to work in the same way attackers can." Attackers can let a model try things freely. Defenders act on production systems where a wrong move takes down something real, and someone has to answer for every action a defender's AI takes.
That is a fair point, and it is why the defender's edge has to come from somewhere else. Defenders own the environment, so they can know their assets, identities, and normal behavior far better than any outsider. They own the context, such as which service account runs encoded PowerShell every night and why. And they own the approval gate, so they decide which actions run on their own and which wait for a person.
A self-improving defense also improves from two directions at once. Your environments show us where the war lab was thin, and every stronger frontier model upgrades our attacker in the lab for free. So the same model progress that helps attackers makes the defense's training harder and better. The defender who uses that, and keeps a person in control of what the defense is allowed to do, is in a much better position than one waiting for the next signature.
How AI Cyber Attacks Work
What is an AI-powered cyberattack?
An AI-powered cyberattack, often just called an AI cyber attack, is any attack where AI gets used to plan, build or run some stage of it, for example the reconnaissance, writing a phishing lure or cloning a voice, developing the exploit, or moving through the network afterwards. It's not the same thing as an attack on an AI system (prompt injection against a chatbot is the usual example), which security researchers tend to call adversarial AI. People mix the two up a lot and they need pretty different defenses.
AI changes two things about attacks and the order matters here. The first is volume, one attacker with a model can probe a lot of targets all day long, so both of the defender's queues (alerts and patches) grow along with it. CrowdStrike's 2026 Global Threat Report, published February 24, 2026, found operations by AI-enabled adversaries went up 89% year over year. IBM's 2026 Cost of a Data Breach Report from July 29, 2026 found one in four malicious breaches was AI-enabled, up 56% on the year before, at an average cost of $6 million.
The second thing is the nature of the attack. Once it's cheap to build a new attack, somebody can build one just for you and use it once, so the specific attack stops repeating. A lot of the time that means there is no signature to match and no playbook for it, and maybe not even a patch yet. Techniques still recur (which is why MITRE ATT&CK hasn't stopped being useful) but the exact artifacts defenders used to lean on often only show up once.
How are hackers using AI in cyberattacks today?
Mostly to do what attackers have always done, only faster and on more targets, and a few things that used to need a skilled person. In one documented 2025 campaign AI did 80 to 90% of the work. The documented cases so far mostly fall into reconnaissance, social engineering, malware and exploitation.
Two cases are worth knowing about. In November 2025 Anthropic said it disrupted what it called the first reported AI-orchestrated cyber espionage campaign, AI handled 80 to 90% of that campaign against roughly thirty targets and humans only stepped in at a handful of decision points. Then in 2026 Sysdig wrote up an attacker who used an AI agent to get from a public CVE to an internal database in four pivots. Roughly, the main uses map to the defender's side like this:
| How attackers use AI | Documented example | What changes for the defender |
|---|---|---|
| Reconnaissance at scale | Anthropic's 2025 case, where AI ran the recon across roughly thirty targets | More probing, which means more alerts that look routine |
| Tailored phishing, voice cloning, deepfake video | Personalized lures and cloned voices turn up in fraud reported to the FBI | Judging an email by its writing mostly stops working |
| Self-modifying malware | GTIG's PROMPTFLUX, experimental malware that asks Gemini to rewrite it | File signatures go stale fast |
| Exploit development | UK AISI found Claude Mythos Preview solved 73% of expert-level capture-the-flag tasks (April 2026) | Less time between a disclosure and the first exploit |
| Running the operation | AI doing most of the steps in an espionage campaign | Attacks move at machine speed and across a lot of hosts |
Not every attacker uses AI for everything. Plenty of what happens in the real world is still commodity tooling and stolen credentials. But the trend only goes one way and right now its making attacks cheaper faster than it makes defense cheaper.
Are autonomous AI cyberattacks real yet?
Partly. AI-orchestrated attacks have been documented, fully end-to-end attacks with no human at all are still not common. Anthropic's November 2025 report had AI running most of an espionage campaign with people at a few decision points. Sysdig's 2026 write-up had an AI agent taking an intrusion from a CVE all the way to a database.
The UK NCSC gave a useful counterweight in its May 2025 assessment of AI's impact on the cyber threat through 2027: "The development of fully automated, end-to-end advanced cyber attacks is unlikely to 2027. Skilled cyber actors will need to remain in the loop." Against well-defended targets that still seems about right. In April 2026 the UK AI Security Institute said Claude Mythos Preview finished its 32-step simulated corporate network attack in 3 of 10 attempts, and in the same breath pointed out there were no active defenders or defensive tooling in that range.
So the fully autonomous attacker is arriving slowly. Whether a human still presses a button somewhere matters less than people think. The thing that already changes defense is the per-target attack, built for you and used once.
How can AI be used in phishing attacks?
AI makes phishing more personal, more convincing and a lot cheaper to run at scale. Attackers use language models to write lures in the target's own context (their role, their vendors, some project they posted about on LinkedIn), clone voices for phone-based social engineering, and make deepfake video to pass as an executive.
The FBI's 2025 Internet Crime Report came out April 6, 2026 with an AI section for the first time: 22,364 complaints, losses of nearly $893 million. Microsoft's Digital Defense Report 2026 (October 1, 2026) found 93% of voice phishing attacks kept the victim on the line long enough to start the social engineering.
Cost changed too. In one IBM X-Force experiment, a model wrote a convincing phishing email from five prompts in about five minutes. The human team needed roughly 16 hours for the same thing.
That's a big reason 'spot the typo' training doesn't do much anymore. An AI-written email usually has no typos, no odd phrasing and a believable reason to click, so judging the email itself keeps getting harder. What still works is looking at what happened after the email arrived. Did anyone click. Did a credential get used from somewhere new. Did a process run on the endpoint, did the account start touching things it never touched before. The question that matters is "did this phish lead to a compromise", and your EDR and identity logs can answer it however good the writing was.
Can AI create malware that evades detection?
Partly. AI malware makes evasion cheaper, but the core behaviors of malware usually stay detectable. Polymorphic malware (code that changes its own form to get around file signatures) was around long before generative AI, what AI changes is that making a fresh variant for every target now costs almost nothing.
The Google Threat Intelligence Group (GTIG) documented PROMPTFLUX in November 2025, experimental malware that calls Gemini to modify itself and avoid detection. Going the other way, Arctic Wolf Labs looked at more than 22,000 AI-assisted malware samples in March 2026 and found their core execution, persistence and command-and-control behaviors were still detectable.
Taken together the lesson is pretty simple. A detection that depends on what the file looks like won't last long against AI-generated variants, while one that depends on what the malware does, how it persists and what it talks to and touches, holds up much better since the malware still has to do those things to get anywhere.
Why can't signature-based detection stop AI-generated attacks?
Because a signature is a description of something somebody already saw, and an AI-generated attack built for one target usually hasn't been seen by anyone. A signature matches a known file hash or string or pattern. If the attacker spins up a fresh variant for every victim there's nothing to match. GTIG's PROMPTFLUX calls Gemini to rewrite itself, so its hash doesn't stay the same long enough to sign.
The other two things a bigger team reaches for hit the same wall. No playbook, because nobody has responded to this exact attack before. Maybe no patch either, if the flaw was only found this morning. That's why adding analysts helps less than it used to, the old shortcuts don't apply and more hands won't bring them back.
Reasoning about method still works. A playbook encodes a decision (if you see this, do that) and a new attack breaks that decision since it doesn't look like 'this'. An investigation encodes method: how to tell what an attacker is after from partial evidence, which pivot is worth the next twenty minutes, when there's enough to call it. Attackers change what they build, they haven't changed what an investigation is. Method transfers, decisions don't, and that's what a self-improving defense gets trained on. For detection it means writing rules on behavior, like a MITRE ATT&CK technique, which a fresh variant can't change.
What does an AI that can find zero-days mean for defenders?
Mostly that the gap between a flaw existing and that flaw being exploited keeps getting shorter, so defenders need options that don't wait on a patch. Anthropic announced Claude Mythos Preview in April 2026, and the UK AI Security Institute reported it succeeded 73% of the time on expert-level capture-the-flag tasks and was the first model to finish AISI's 32-step corporate network attack range. AISI was careful to add that the range had no active defenders. So the result says more about weakly defended networks than about hardened ones.
Simbian's own mid-2026 analysis found the window from a vulnerability being disclosed to it being exploited dropped from roughly 30 days to roughly 5 over the previous six months, and for most enterprises a patch cycle measured in weeks just doesn't fit inside that.
Without a patch, a few things still help:
- Close the path: if the vulnerable system can't be patched today put a compensating control in front of it, such as an access-control change or some segmentation, so it's unreachable or at least worth less to an attacker.
- Hunt for prior use: check whether the technique was already used against you before anyone knew about the flaw.
- Watch for the technique: detect the behavior an exploit of that class tends to produce, so a new variant can get caught even when there's no signature for it yet.
A pentest that proves whether the flaw is actually reachable in your environment helps too. It tells you which of the many new CVEs deserve the first hour, because a CVE isn't really an exploit until someone shows it can be used against you.
How fast do you need to respond to an AI-speed attack?
Faster than a human-paced queue usually allows. CrowdStrike's 2026 Global Threat Report put the average eCrime breakout time, meaning initial access to lateral movement, at 29 minutes, and the fastest one they saw was 27 seconds.
That's the attacker's clock. The defender's clock usually starts much later, when someone actually picks up the alert, and if it's sitting in a long queue that can be hours after it fired. Breakout in 29 minutes against a two hour wait for an analyst means the attacker has already moved by the time anybody opens the case.
Simbian investigates and responds to every threat within minutes, on average under 4 minutes from detection to response. That doesn't mean every action runs without a person. Every action starts gated and you decide which action types are fine to run on their own, closing a false positive for example, and which ones wait for approval, like isolating a production server. Either way the investigation itself runs at machine speed, so a gated action is already investigated and just waits on a decision.
How do you defend against AI-powered cyberattacks?
Start with the basics, AI-powered attacks still go after the same weaknesses. Phishing-resistant MFA, patching what's exposed on time, least-privilege access, decent logging and segmentation all still matter. After its Mythos Preview evaluation the UK AI Security Institute gave the same advice: regular security updates, strong access controls, secure configuration and comprehensive logging.
What changes is how AI cyber defense and the security team work on top of those controls:
- Investigate every alert: AI-built attacks are often made to look unremarkable on purpose, so they tend to end up in the low-severity tail that severity-based triage skips over.
- Respond in minutes: CrowdStrike put the 2026 average eCrime breakout time at 29 minutes, the investigation needs to start the moment the alert lands.
- Detect behavior: write detections around what attackers actually do. Signatures age out fast against generated variants.
- Close the path, not just the ticket: once a real threat is handled, tighten the detection and change whatever access let it happen in the first place.
- Test continuously: keep proving which exposures are actually reachable, all year and not only when the auditors are due.
- Hunt on hypotheses: when companies in your sector get breached, go look in your own data for the same activity.
- Keep humans in control: gate the actions, keep an audit record and have a person approve every change made to the defense itself.
The common thread is a defense that gets a little better from each case it handles.
How can you tell if an attack was AI-generated?
Often you can't, and it usually matters less than people think. A phishing email or a piece of malware rarely has anything in it that reliably proves a model wrote it, and the tools that claim to detect AI-written text generally aren't reliable enough to base a security decision on.
Deepfake audio and video are a partial exception since there is specialised detection for them, although people asked to spot a deepfake by eye or by ear generally do only a little better than chance.
There are a few signals that point toward automation, though none of them proves it by itself. Speed, for one, with lots of steps across lots of hosts in very little time. Breadth is another, probing every exposed service at the same time. Per-target variation can be a tell too, when each victim gets a slightly different lure or payload.
What actually matters is the behavior. A credential used from an unusual place, a process injecting into another one, data getting staged for exfiltration, all of those need the same investigation whether a person or a model planned them. Defending against the behavior covers you either way, and working out whether AI was involved mostly helps the incident report.
How Self-Improving Defense Works
Why are LLMs good at offense but bad at defense?
Mostly because offense makes its own training data and defense doesn't. You can measure the gap, as of October 2026 no frontier model passes Simbian's Cyber Defense Benchmark. An attack attempt grades itself, you got in or you didn't, and that answer is instant and free and nobody argues with it. So anyone can generate as much attack training data as they want and plenty of people already have, which goes a long way to explaining why models got good at attacking so fast.
Defense doesn't have anything like that signal. To learn from a defense that worked you'd need to know what the attacker was really after, every path they took including the ones they gave up on, and which defensive action actually made the difference. Side by side it looks something like this:
| Offense: it grades itself | Defense: nothing grades it |
|---|---|
| The question is "did you get in?", yes or no | The question is "what stopped it?" |
| The answer costs nothing and nobody disputes it | The attacker's goal was never stated, and recon and theft look alike at first |
| Unlimited training data can be generated | Abandoned paths leave no record, only the one that worked is visible |
| So models are good at attacking | Nothing labels which action mattered, so no model was trained to defend |
The Cyber Defense Benchmark shows what that means in practice. In the October 2026 results the leading model clears the 50% bar on only 7 of 13 MITRE ATT&CK tactics. Newer models aren't automatically better at defense either, five recent releases scored lower than the model they replaced, Sonnet 5 for example came in 17.0 points under Sonnet 4.6.
Does a self-improving defense learn my environment, or a vendor's generic baseline?
Both, and they're kept apart. General defensive skill, like how to follow an attacker or when an investigation is really finished, comes from training in a war lab. Anything specific to your company lives in the Context Lake™, an organization-wide persistent memory holding four kinds of knowledge: your assets and what each is worth, your identities and who they belong to, your processes and runbooks, and decisions your team already made.
Think about a surgeon. Surgeons learn on other patients, then operate on you using what they learned plus your scans and history. Skill carries over, the specifics are yours. The defense is the same, it applies general skill to your specifics, so value generally starts with the first case and not after months of baselining.
Specifics matter more than people expect. Take one alert, say personal data going to an unsanctioned cloud storage bucket. The right answer depends on where you are:
- A regulated bank: page the on-call team and freeze actions during the quarter-end change window.
- A mid-market SaaS company: isolate the endpoint and notify the team in Slack, with no human gate, because that team has already lifted the gate for that action type.
- An MSSP tenant: isolation is allowed, but disabling accounts is the customer's call, so it gets escalated.
A generic baseline would most likely give all three the same answer. Your context is what makes the answer yours, it stays yours, and we never train on your data.
How does an AI defense get better over time without retraining on your data?
Through in-context learning. Your data stays in your tenant and gets read at run time when an investigation needs it, it never goes into a model's weights. When the defense learns something it learns it as a skill or a bit of context the next investigation reads, not as a new version of the model. People confuse fine-tuning and in-context learning a lot so here's the difference:
| Fine-tuning | In-context learning | |
|---|---|---|
| Your data | Has to go into the model's weights | Stays in your tenant, read at run time |
| A better model ships | Retrain, revalidate, redeploy | You get it the day it lands |
| Changing its behavior | Another training run | Takes effect on the next case |
| "Why did it decide that?" | Opaque weights | Points at the exact context it used |
| Model choice | Usually tied to the model you tuned | Any frontier model, swappable |
Every case can teach the next one. If the same kind of investigation keeps getting stuck on the same gap, the defense writes up the improvement itself, attaches the past investigations that show the problem and sends it for approval, and a person approves it before anything changes. Once an improvement has proven itself it can go into a regression suite so it keeps getting tested as other things change around it.
Here's a concrete one. On a production shadow tenant every query the agent wrote in XQL (the query language of Palo Alto Cortex XSIAM) failed at first, because it had never seen that language. Before anyone opened an engineering ticket the agent was asked to find the root cause, and what it came back with was a skill for the language. Over two iterations query success went from 0% to roughly 70%, and nobody retrained a model to get there.
Does it matter whether a defensive AI learns from synthetic data or real attack telemetry?
Yes. To learn how to defend, labelled synthetic data generally teaches more. To know what normal looks like in your environment, real telemetry is still the better record. Synthetic data has some well known weaknesses in machine learning (it can miss the messy distribution of a real environment, and its noise can look too clean) which is why a good defense still reads your real telemetry at run time.
Real attack telemetry has a gap when it comes to learning defense. It only records the path that worked. The paths an attacker tried and dropped leave nothing behind, and those are exactly what a defender needs, since an AI investigator that can't recognize a dead end tends to spend most of its time in one. Microsoft's Defender research team made a similar point in May 2026, labelling real attack logs "requires not only labeling malicious activities, but also fully reconstructing attack scenarios." Real incident data also hardly ever labels which defensive action stopped the attack.
Defensive training data does not exist in nature, so we manufacture it. We build a synthetic company with its own people, machines and daily traffic, then have one AI attack it while another defends. Because we authored both sides we know the attacker's goal, every path it tried, and which defensive action stopped it. The product ships into 300+ enterprise environments, and wherever it falls short of what your team would have done, that shortfall tells us where the lab was thin. We fix the lab and ship again, the way a vendor fixes a bug from a bug report. Your telemetry never enters the lab.
We never train on your data. A bug report is not your data, and neither is this. What comes back from production is a shortfall report, where the product fell short of what your team would've done, and not your logs.
Two things keep the lab getting better. Your environments show us where it was thin, and every stronger frontier model upgrades our attacker for free. Since specific attacks stop repeating, what the lab teaches matters a lot, and what it teaches is method. Method transfers, decisions don't.
What is a cyber range, and can it train defensive AI?
A cyber range is a controlled environment that simulates networks, systems and traffic so people can train and test without touching production. NIST's "The Cyber Range: A Guide" (NICE, September 2023) calls them places for cybersecurity training, exercises and testing. Most are built for human teams practicing incident response, or for testing tools against known scenarios. A range can help train defensive AI, usually just as a test bed though, since most of them only record whether the attack got stopped.
Ranges get used to test AI more and more. AISI's 32-step "The Last Ones" range from its April 2026 evaluation of Claude Mythos Preview is one, built to measure how far an AI attacker gets. AISI estimates a human needs about 20 hours on it and says its ranges "lack security features that are often present, such as active defenders and defensive tooling." Hack The Box launched HTB AI Range in December 2025 to benchmark AI agents on offense and defense.
To actually train a defense, a range would have to record what a defender needs to learn from. Most don't. A typical range scores whether the attack was stopped and the dead ends go unrecorded.
Simbian's war lab looks like a cyber range from outside, it isn't one. Attacker and defender are both AI, and because we author both sides every path is labelled, dead ends included. A range mostly trains people on scenarios. The war lab manufactures labelled defensive training data at a scale no human exercise calendar keeps up with.
Can an AI defense handle attacks it has never seen before?
Often, yes, as long as it reasons about behavior and doesn't just match patterns. AI threat detection based on anomalies and behavioral analytics can flag that something is new, but flagging it isn't the same as handling it. Somebody still has to figure out what the new thing is, whether it matters and what to do.
A defense that is designed to think, not just follow, investigates the new thing the way an experienced analyst would. It asks what the activity is reaching for, pulls evidence from whichever tools can answer that, and decides if there's enough to make the call. That works on a never-seen attack for the same reason it works for a senior analyst, a loader nobody has seen before still has to persist, still has to talk to something and still has to touch data, and an investigator can check each of those. Arctic Wolf Labs found exactly that in March 2026 when it went through more than 22,000 AI-assisted malware samples, the core execution, persistence and command-and-control behaviors were still detectable.
The real limit is that no defense can prove a negative. It can show that it looked, what it checked and why it concluded what it did, but it can't prove nothing happened in places it didn't look, which is why hunting and testing keep running alongside investigation.
How do pentest findings become better detections?
A pentest finding becomes a better detection once someone hunts for the proven path, checks it against what the SOC actually saw and turns it into a tested rule, all mapped to the same MITRE ATT&CK technique. What happens in most organizations instead: the finding becomes a ticket, the ticket sits in a backlog, detection engineering maybe gets to it weeks later. Meanwhile the environment stays exposed and often nobody checks if the SOC would have caught someone exploiting it.
The four-node loop closes that gap, each node answers one question and passes the result on:
| Node | The question | What it produces |
|---|---|---|
| AI Pentest Agent | What could happen? | A proven, reachable attack path |
| AI Threat Hunt Agent | Did it happen? | Evidence of whether the technique was already used here |
| AI SOC Agent | Did we respond? | Whether the activity was detected and handled correctly |
| AI Detection Engineering Agent | Can we catch it next time? | A candidate detection, tested against the path that found the gap and approved by a person before it ships |
Every finding maps to the same ATT&CK technique ID, so one scoreboard shows all four answers. Pentest and SOC share findings both ways. A SOC judgment can change pentest scope and severity, a pentest finding tells the SOC what to watch for. Because the loop runs the same way every time, coverage builds from one engagement to the next. The AI pentest guide goes deeper on the testing side.
What is the difference between red, blue, and purple teams?
Red teams attack, blue teams defend. Purple teams (in most organizations more a practice than a separate team) turn what the red team did into better detection. The names come from military exercises and mainly describe goals and outputs:
| Red team | Blue team | Purple team | |
|---|---|---|---|
| Goal | Get in the way a real attacker would | Detect, investigate, and respond | Turn what the red team did into better detection |
| Typical cadence | Periodic engagements | Continuous | Scheduled exercises |
| Main output | A report of paths and findings | Investigations and responses | Detections improved and retested |
Wiz put it well in an August 2026 guide: purple teaming is "a collaborative validation loop: emulate a realistic procedure, observe what the defensive stack sees, improve the control, and retest." Cadence is the weak spot. The loop runs once per exercise, a few times a year if you're lucky, and attackers don't care about anyone's exercise calendar.
Self-improving defense runs the purple loop whenever a technique needs testing, one technique at a time, red and blue on the same MITRE ATT&CK map. Breach and attack simulation tools also make testing more frequent. The question that matters is whether each finding ends up as a tested detection. Keep the human red team, people come up with creative attacks a loop may never think of.
How do you stop detections from silently failing?
Mostly by keeping an eye on the three things that quietly degrade most SOCs, and fixing them before a gap turns into a missed attack. Simbian's own 2026 analysis found roughly 20% of SOC detection rules stop firing within six months. Detections hardly ever fail loudly, they just stop firing and nobody notices until an audit or an incident turns it up.
The first is the detection rules themselves. Rules go noisy, or break, or never existed for a technique that matters, and a defense that reads every verdict can usually see which rules are noisy or missing and propose tuned or new ones written in your SIEM's own query language (KQL, SPL, XQL and so on).
The second is data-pipeline drift, where a log source goes quiet or a field changes shape. One renamed field after an overnight SIEM upgrade can silence 40 detections with no error and no alert.
Integrations are the third. A connector breaks after a vendor update or a threat-intel feed lapses, enrichment comes back empty and verdicts keep shipping without it.
Teams that catch these early commonly keep a fire-rate baseline for each rule, replay a known-bad test event through the pipeline now and then, and alert when a field they rely on starts coming back empty.
Coverage decays. Simbian repairs. The repairs go through the same governance as everything else, detection changes get logged with the verdict that triggered them and a person approves them before they ship, and they can be rolled back afterwards.
AI Agent Governance, Trust, and Control
Who is accountable when an AI defense makes the wrong call?
Whoever approved the action is accountable, and that's the main reason every action should start out gated. The AI defense proposes, a person approves and the record shows who approved what, on which evidence. It works pretty much like a junior analyst's recommendation that a senior engineer signs off on. If you've chosen to ungate an action type, the accountable person is whoever approved lifting that gate, so gate changes should be recorded like any other approval.
The roles around this are shifting. A lot of teams now separate the person who writes a policy or skill from the person who approves it (author is not approver) so nobody can create a change and wave it through themselves. That usually ends up with something like an AI SecOps Manager, who owns rollout, governance and service levels for the program.
The approval numbers are worth knowing too. Across Simbian deployments, 95% of the responses Simbian proposes are approved by the customer's own analysts. Their call, not ours. In the other 5%, we get additional context that improves our next response. That number describes who made the decision. It isn't a guarantee of any outcome and the decision stays with your team.
Which actions should an AI defense take without human approval?
The ones you have decided are safe, one action type at a time. A sensible default is to start with every action gated and lift gates as confidence builds, so closing a false positive can probably run on its own fairly early while isolating a production server waits for you, and the line moves as the track record grows.
The UK NCSC suggested a useful way to score this in a September 2026 blog post, One does not simply defend agentically. It rates how risky an automated defensive action is on five dimensions, each scored 0 to 4, which are potency, scope, criticality, rollout confidence and recoverability.
At the bottom of the potency scale the AI only gives a human advice it can explain. Investigating an alert to a verdict sits roughly there, which is why that part can usually run from day one while response actions climb the scale. Something easy to reverse and narrow in scope that doesn't touch anything critical, like closing a duplicate alert, is a good candidate to run without approval. Something that hits many systems or touches identity or production endpoints or is hard to undo should generally stay gated longer.
Two signals help most when deciding to lift a gate. One is how often your analysts override the proposed response for that alert type. The other is whether the action can be reversed at all. Low override rates on reversible actions are a fair reason to ungate, and anything acting against an identity or an endpoint (suspending an account, revoking a privilege) usually takes longer to earn it. Even then an ungated action still has a human on the loop who can see the record and reverse it.
What should a reviewer check before approving an AI's change to its own skills?
Before approving, look at why the change got raised and how many investigations it touches, the past cases that show the gap, the exact before and after, and whatever it contradicts. A proposal to change the defense's own skills shouldn't ever arrive as a bare request, so all of that should already be on the page:
- Why it was raised: which gap it addresses, and roughly what share of investigations it will change.
- The evidence: links back to the past investigations that showed the gap, each one with a short note on how the change would've changed that investigation
- The before and after: the exact diff, so you see what gets written.
- Conflicts: anything already in context that the proposal contradicts, shown side by side.
Some proposals you can reject straight away. Blanket rules like "ignore all PowerShell alerts". Allowlists that switch off investigation. Anything generalized from one incident, and temporary exceptions with no expiry on them. The expiry is the one people skip. A red team exercise gets added as "expected activity" for two weeks, the exercise ends, nobody removes the entry, and from then on real alerts on that asset quietly get discounted. A well designed system shouldn't propose verdict-automation rules or allowlists in the first place.
When a proposal conflicts with existing context there are usually three ways to go: keep what's there, accept the proposal, or edit and merge. If you reject it, say why in a comment, because the next proposal can read earlier rejections and adjust. After approval the change should get validated once it's applied, date-stamped so stale facts get flagged later on, and covered by regression testing. A conflicted request is never applied until a human resolves the conflict.
Would you trust an AI-generated verdict without seeing the evidence?
No, and you shouldn't have to. A verdict should come with its evidence, the queries that ran, process IDs, logs pulled and the reasoning connecting them, so an analyst checks it the same way they'd check a colleague. Each query should be re-runnable in the source tool so the evidence gets checked against your SIEM and not against the AI's summary of your SIEM. A good verdict also says why the investigation decided it was done, not just what it looked at.
In-context learning helps with this. What the defense knows lives in readable skills and context, outside the model weights, so it can point at the exact context behind a verdict. Say a fact like "this service account runs encoded PowerShell nightly" shaped the verdict, the investigation can show that fact and link to the request that added it.
Early on, try a simple spot-check. Hand an analyst only the evidence the verdict shows and see if they land on the same call. If they don't, the evidence isn't doing its job yet however confident the verdict sounds.
Can you stop, reverse, and reconstruct what an AI defense did?
You should be able to do all three and if you can't, the defense generally isn't ready for production. Stopping comes down to gates and permissions you control. Reversing means actions can be rolled back and there's a record of the state things were in before. Reconstructing needs a full record of every action and query and the reasoning behind each, plus who approved it and which identity and permissions the agent was acting under, kept in a log the agent can't edit itself.
In practice that means four properties on every action the defense takes. It's sandboxed, so everything runs isolated. It's permission-controlled, every API call and every access. It's human-gated, any output can require approval first. And it's yours to extend, you can add skills and context and steer it.
Same goes for changes to the defense itself. Every change to what it knows should keep its full history (who raised it, who approved it, when) with a before and after diff, so a bad change can be tracked down and undone.
Blast radius is a useful thing to ask about when evaluating: what's the most damage this agent could do before someone stops it? If nobody can say, the permissions are probably too broad or there aren't enough gates. Test the stop before you need it.
What does AI agent governance look like in security operations?
AI agent governance in a SOC mostly looks like the governance you already apply to privileged human access, adjusted for agents that act fast and at scale. Each agent should have a named owner, a defined scope and explicit permissions, plus a review cadence, an expiry on any temporary privileges and some tested way to switch it off. In a SOC those translate roughly like this:
| Governance element | What it means in a SOC |
|---|---|
| Owner | A named person accountable for the agent's behavior |
| Scope | Which alert types, systems, and data the agent may touch |
| Permissions | Which actions run alone and which wait for approval |
| Review | Regular review of overrides, rejected proposals, and errors |
| Change control | Every change to the agent's skills shown as a diff and approved |
| Kill path | A tested way to stop the agent and reverse what it did |
Several frameworks are landing on similar ideas. Forrester's AEGIS framework lays out 39 controls across six domains (approval gates for irreversible actions are one of them) and maps them to the NIST AI RMF, ISO 42001, the EU AI Act and MITRE ATLAS. In May 2026 CISA, the NSA, Australia's ACSC and partner agencies published joint guidance on careful adoption of agentic AI services, and the UK NCSC's scoring of automated defensive actions gives you a way to decide which actions need a human first.
The organizational side matters about as much as the controls. Governance of security agents rarely sits with the SOC alone. The AI SecOps Manager usually runs it day to day, but risk appetite tends to get set together with IT, legal, privacy and compliance, and AEGIS itself calls for a governance board drawn from those teams.
Is an AI investigation deterministic or non-deterministic?
The reasoning is non-deterministic, meaning the same input won't always give the same output, and the policy around it has to be deterministic. Language models are probabilistic so two runs of one investigation can take slightly different paths, teams that drop AI steps into rigid workflows often see that as inconsistent results. Fair concern.
You deal with it by putting whatever has to be consistent into policy that's enforced deterministically. In Simbian that policy is written as skills in plain language, and covers four kinds of rules: risk tolerance and escalation, investigation playbooks, asset criticality and identity tiers, change-freeze and regulatory rules. The model reasons inside those. It doesn't get to decide if a change freeze applies.
The output is what you hold to a consistent standard. Every case documented the same way, same evidence fields, and override rates by alert type tell you if the verdicts are consistent where it counts.
Can attackers manipulate an AI defense?
Yes, they can try. The main ways in are pretty well known by now. One is prompt injection, instructions hidden in data the AI reads like an email body or a file name or a log line. Another is poisoned context, where an attacker gets false "normal" facts into what the defense believes about your environment. There's evasion too, shaping malicious activity to look like something the defense learned is normal (encoded PowerShell on a host where that's expected, say). And then the human gate, if reviewers approve proposals without actually reading them.
Each of those needs its own control. The model layer needs hardening against injection, which is what Simbian's TrustedLLM™ is for. It's private and prompt-injection hardened, hardened against data poisoning in the war lab, and no customer data goes into training.
Evasion is handled by scoping. A learned exception says when it does not apply, wrong time of day or wrong parent process, and anything outside that scope gets investigated in full. Context only changes through tracked requests that a person approves, writes stay scoped to your tenant, and a conflicted request is never applied. Review has to be real as well, which means watching for rubber-stamping, keeping proposals small enough to actually review and attaching evidence to each change.
No AI defense is immune to manipulation, same as no human analyst is immune to social engineering. The aim is making manipulation visible and hard to make stick.
Proving Your AI Cyber Defense Gets Better
How do you measure whether your security is getting better?
Measure coverage, speed and how often your team agrees with what the defense does. Not how many alerts it processed. Volume is still what most teams report though, in the 2026 SANS SOC Survey the number of incidents handled was the top SOC metric for the tenth year running, and a team handling more incidents can be getting worse at stopping attacks.
Three things tell you more:
- Coverage: of the attack techniques that matter to you, how many would really be detected, checked by replaying them
- Time from detection to response (often reported as MTTR): how long a real threat runs before something is done.
- Approval: how often your analysts accept the proposed response, by alert type.
Coverage only means something when you can see what moved it. At one Simbian customer MITRE ATT&CK detection coverage went from 33% to 56% to 83% over three loop cycles. Cycle one: the pentest side tested a set of techniques, the hunt looked for prior use in the logs, the SOC caught some and detection engineering shipped rules for the gaps. Cycle two retested those plus new ones. Cycle three ran evasion variants against the new rules. It went up because each cycle closed what the last one found.
Simbian frames it as objective, evaluation and evolution. You set an objective, investigate every alert of some type to a verdict for example. The defense scores itself against that on your own data and shows which cases fell short, then proposes changes to close the gap, and a person approves them. When an objective hits full marks it moves into a regression suite so it keeps getting tested.
Can you prove an AI defense still works after the model or tools change?
You can, as long as you actually test after every change. Newer models aren't automatically better at defense. On the Cyber Defense Benchmark (October 2026 edition) five recent releases scored under the model they replaced. Sonnet 5 was 17.0 points below Sonnet 4.6, Opus 4.8 was 11.8 points below Opus 4.6. Several of them just stopped investigating sooner (Opus 4.8 stopped after about 34 turns where Opus 4.6 ran about 52) which made runs cheaper and left part of the attack unread.
So don't swap models on release day. Re-benchmark every release on defensive work before trusting it in production. Simbian ships no model of its own. It measures every model it runs and builds around what each one gets wrong.
Everything around the model changes too. Tools get upgraded, schemas shift, connectors break. Two things keep a defense provably working through that. A regression suite, where objectives that already score well stay under test so one improvement can't quietly cost you somewhere else. And self-repair for the plumbing, which notices when a connector or schema changes and re-learns it, with review and rollback before anything ships.
What is the difference between CTEM and the SOC?
CTEM finds where you're exposed, the SOC detects when someone uses that exposure. Gartner's February 22, 2024 press release defines continuous threat exposure management (CTEM) as "a pragmatic and systemic approach organizations can use to continually evaluate the accessibility, exposure and exploitability of digital and physical assets," and predicted organizations prioritizing security investment based on a CTEM program would see a two-thirds reduction in breaches by 2026. A CTEM program is usually described in five stages: scoping, discovery, prioritization, validation and mobilization.
They answer different questions and usually sit in different teams with different tools:
| CTEM | SOC | |
|---|---|---|
| Question it answers | Where could we be attacked? | Is someone attacking us right now? |
| Main inputs | Asset inventory, vulnerabilities, attack paths, validation tests | Alerts, logs, endpoint and identity telemetry |
| Typical output | A prioritized list of exposures to fix | Investigations, verdicts, and responses |
| Cadence | Cycles of scoping and validation | Continuous |
Some vendors argue CTEM should become the SOC. The two really do answer different questions though, and in practice the problem is they rarely share a scoreboard, a CTEM program will often validate that an attack path is real and nobody checks if the SOC would catch someone using it. Putting both on the same MITRE ATT&CK map fixes that, every exposure and every detection maps to a technique so you can see where you're exposed and also blind.
How will breach and attack simulation improve threat detection and response?
Breach and attack simulation (BAS) improves detection when somebody acts on what it finds. BAS and automated pentest tools such as Cymulate, Picus and Pentera run known attack techniques safely against your environment and report which ones your controls blocked or detected. That's useful, it swaps guessing about control coverage for an actual test.
Whether anything improves still depends on what happens after the report. A finding that the EDR missed a technique only turns into better detection if someone writes or tunes the rule and reruns the test to confirm it worked. In a lot of teams that handoff is the slow part, and BAS platforms need ongoing care to keep their scenarios current. In its Market Guide for Adversarial Exposure Validation Gartner puts BAS and automated penetration testing under one category, which it describes as providing "consistent, continuous and automated evidence of the feasibility of an attack."
BAS usually stops at the report. A loop keeps going, the missed technique becomes a hunt for whether it was already used, then a new or tightened detection and a retest, all on the same MITRE ATT&CK technique map as everything else.
Is a purple team exercise worth it if you've already done red teaming?
Usually yes, as long as the output is a change to your detections and not just another report. A red team exercise tells you how someone got in. A purple team exercise runs the same MITRE ATT&CK techniques while the defenders watch, so the team sees what each tool caught and what it missed and can fix gaps while it's fresh.
Cadence is usually the problem. Exercises happen once or twice a year, attackers don't keep that schedule. A detection fixed in spring may have drifted by autumn, a log source moved, some field got renamed in a SIEM upgrade, someone edited the rule.
A continuous purple loop handles that, testing one technique at a time all year and retesting after each fix. Keep the exercises for the creative work humans are best at and let the loop do retesting, that way the lessons from each exercise actually stay fixed.
What do defensive AI benchmarks actually tell you?
A defensive AI benchmark tells you how a model does at AI cyber defense in a controlled test. Useful, but it's not the same as how a product does in your environment. Simbian's Cyber Defense Benchmark has models investigate real attack logs with deterministic ground truth, across 13 of the 14 MITRE ATT&CK tactics and 105 attack procedures, each environment holding more than 100,000 events. To pass, a model needs at least 50% coverage on every tactic. As of October 2026 none do, and of 30 models the leader reaches 45.1%. Other public defensive benchmarks exist as well, like CyberSOCEval from Meta and CrowdStrike, and Microsoft's ExCyTIn-Bench.
SentinelLabs made a fair critique of the defensive benchmarks it reviewed (CyberSOCEval and ExCyTIn-Bench among them). In January 2026 it argued "Benchmarks are Measuring Tasks, not Workflows": a benchmark can check whether a model finds evidence in logs, a real SOC also runs on context, business judgment and handoffs between people and tools.
So use both kinds of evidence. Benchmark the model to know its ceiling and blind spots before trusting it. Then measure the deployment on your own alerts, replayed techniques, verdicts compared with your analysts, override rates over time. A benchmark score with an edition date is a published claim anyone on your team can go and check.
Is MITRE ATT&CK coverage a meaningful measure of your detection?
It's meaningful as a map and often misleading as a percentage. MITRE ATT&CK gives every technique an ID so findings from pentests, hunts and investigations all land in the same coordinate system, which is valuable. The trouble is the coverage heat map, where a technique turns green because a rule exists for it whether or not that rule would catch a real attacker.
MITRE's Center for Threat-Informed Defense tackles this with its Summiting the Pyramid work on detection robustness, which it describes as "how difficult it is for adversaries to evade a detection." A rule matching one tool's command line is easy to get around. A rule watching the underlying behavior is a lot harder to slip past, and the heat map can't tell those two apart.
A more useful version of coverage is measured by replaying techniques, evasion variants included, and counting what gets detected, cycle over cycle. That's how one Simbian customer's coverage moved from 33% to 56% to 83% over three loop cycles, each cycle tested against the rules the previous one wrote. Coverage measured like that tells you whether you're getting better, coverage counted off a rule inventory mostly tells you how many rules you've got.
Evaluating and Adopting Self-Improving Defense
If every vendor uses the same frontier models, what makes one AI defense better than another?
In AI cyber defense the harness, the training data and the context make the difference, because the model itself is turning into a commodity. In an April 2026 post a group of Forrester analysts including Jeff Pollard put the question plainly: does your value proposition survive when frontier model access becomes ordinary?
The model doesn't decide the outcome on its own. The same model usually behaves differently depending on the harness around it, what context it gets and when it's allowed to stop, and the Cyber Defense Benchmark shows how much stopping early by itself can cost. The model is one cylinder. The harness is the car.
What tends to separate one AI defense from another:
- What trained it: labelled defensive data, or just general model training plus prompts.
- What it remembers: whether it builds a lasting model of your environment and your team's decisions, and if you can see and control that memory.
- What it can prove: published benchmark results with an edition date, plus measured results in your own environment.
How can you check a vendor's claim that its AI stops AI attacks?
By asking for three things you can check, what the AI was trained on and how its defensive data got labelled, a dated public benchmark on real attack logs, and evidence from your own environment. A demo isn't one of them.
Start with training. Real-world telemetry rarely records which action stopped an attack, so a vendor should be able to explain where its labels came from.
Then a published benchmark on real attack logs (the Cyber Defense Benchmark, CyberSOCEval or ExCyTIn-Bench, for example) with the edition and date stated. An unpublished internal number, or a capture-the-flag style score built for offensive tasks, doesn't say much about defense.
And evidence in your own environment, things like techniques replayed against your stack, verdicts compared side by side with your analysts, override rates by alert type, and an export of the reasoning behind whichever verdict you pick. Mature systems can export the reasoning behind any verdict and immature ones usually can't.
"We protect you from AI attacks" is a positioning statement coming from any vendor, us included. The evidence behind it is what turns it into something you can rely on.
How do you test an AI defense against realistic attacks before you trust it?
Replay MITRE ATT&CK techniques, evasion variants included, against your own stack, and run it on your own live alerts with every action gated so you can compare what it concludes with what your analysts conclude. A proof of value in a vendor's lab mostly tells you about the vendor's lab, a test in your environment tells you how it handles your tools, your data quality and your noise.
A first month typically goes roughly like this:
- Week 1: the agents connect to your tools and onboard themselves, discovering tables, fields and schemas.
- Week 2: investigations run on live alerts and every action still waits for approval.
- Weeks 3 to 4: you compare its verdicts against your own analysts' verdicts on the same alerts.
- After that: gates get lifted one action type at a time, as the record supports it
The replayed techniques show what gets detected and how it's investigated, the live alerts show how it copes with your actual noise. You generally want both before relying on it.
Does an AI defense require you to centralize your data first?
No, a well built one reads your data where it already lives. Gartner lists it among the questions to ask AI SOC agent vendors: does the solution require data centralization, or can it operate in any environment? Needing a new data lake before investigations can even start often turns a deployment into a data project that runs for months.
Simbian reads from more than 100 of the tools you already run (SIEMs, EDR, XDR, cloud and identity systems, ticketing) and reasons across them as one picture. It doesn't need new sensors or agents deployed, a data migration, or your detection stack replaced or re-tuned first.
Data readiness still matters though. An investigation needs at least the alert source, endpoint telemetry, identity logs and some asset context to get to a confident verdict, and gaps there show up as weaker investigations. That's a reason to know where your telemetry has holes, not a reason to move it all into one place first.
Which SOC use case should you trust an AI agent with first?
Start where volume is high and the actions are reversible, which for most teams means investigating alerts to a verdict. Investigating an alert doesn't change anything in your environment, so a wrong verdict costs an analyst some review time and not an outage, and with the volume the verdict comparison in weeks three and four gives you a decent sample to judge it on.
Keep the response side small at first too, starting with whatever's easiest to undo. A low override rate on an alert type, held over time, is the signal that its response can run on its own. Phishing triage and endpoint alert validation are common next steps since the evidence is concrete and the actions (quarantining an email, say) are easy to reverse.
Expansion usually follows results. One customer started with the SOC and added pentesting and threat hunting within 12 months.
Where will new analysts learn the craft once AI takes the queue?
Mostly from supervising and correcting the agents, which turns out to be a decent apprenticeship. A new analyst's first year shifts toward reviewing finished investigations, fixing the wrong ones and approving or rejecting proposed actions. The tier 1 queue used to be where people learned what normal and malicious look like. Once agents take the tier 1 and tier 2 queues and the playbook work, people move up to building, running and governing the agents, and the learning moves up with them.
In most teams that ends up as three roles. The AI SecOps Analyst supervises the agents, reviews their investigations, approves actions and works whatever the agents send up, and reading a finished investigation then deciding whether it's right tends to teach the craft faster than clearing a queue ever did. The AI Skill Manager often comes from senior analysts and SOAR engineers and encodes the team's knowledge into skills the agents run, with no coding needed. Then there's the AI SecOps Manager, who owns rollout, governance and service levels across every SecOps program, not only the SOC.
In a September 2026 post Forrester's Jess Burn and Jeff Pollard asked the same question from the other side: "Who develops the experts these systems will continue to require?" Probably the people reviewing and correcting the agents every day. Their corrections are what make the defense better and doing that work well is how a new analyst becomes a senior one.
