How Self-Improving Defense Works
Part of Self-Improving Defense: AI Cyber Defense Against AI Attacks — read the full guide.
Why are LLMs good at offense but bad at defense?
Mostly because offense makes its own training data and defense doesn't. You can measure the gap, as of October 2026 no frontier model passes Simbian's Cyber Defense Benchmark. An attack attempt grades itself, you got in or you didn't, and that answer is instant and free and nobody argues with it. So anyone can generate as much attack training data as they want and plenty of people already have, which goes a long way to explaining why models got good at attacking so fast.
Defense doesn't have anything like that signal. To learn from a defense that worked you'd need to know what the attacker was really after, every path they took including the ones they gave up on, and which defensive action actually made the difference. Side by side it looks something like this:
| Offense: it grades itself | Defense: nothing grades it |
|---|---|
| The question is "did you get in?", yes or no | The question is "what stopped it?" |
| The answer costs nothing and nobody disputes it | The attacker's goal was never stated, and recon and theft look alike at first |
| Unlimited training data can be generated | Abandoned paths leave no record, only the one that worked is visible |
| So models are good at attacking | Nothing labels which action mattered, so no model was trained to defend |
The Cyber Defense Benchmark shows what that means in practice. In the October 2026 results the leading model clears the 50% bar on only 7 of 13 MITRE ATT&CK tactics. Newer models aren't automatically better at defense either, five recent releases scored lower than the model they replaced, Sonnet 5 for example came in 17.0 points under Sonnet 4.6.
Does a self-improving defense learn my environment, or a vendor's generic baseline?
Both, and they're kept apart. General defensive skill, like how to follow an attacker or when an investigation is really finished, comes from training in a war lab. Anything specific to your company lives in the Context Lake™, an organization-wide persistent memory holding four kinds of knowledge: your assets and what each is worth, your identities and who they belong to, your processes and runbooks, and decisions your team already made.
Think about a surgeon. Surgeons learn on other patients, then operate on you using what they learned plus your scans and history. Skill carries over, the specifics are yours. The defense is the same, it applies general skill to your specifics, so value generally starts with the first case and not after months of baselining.
Specifics matter more than people expect. Take one alert, say personal data going to an unsanctioned cloud storage bucket. The right answer depends on where you are:
- A regulated bank: page the on-call team and freeze actions during the quarter-end change window.
- A mid-market SaaS company: isolate the endpoint and notify the team in Slack, with no human gate, because that team has already lifted the gate for that action type.
- An MSSP tenant: isolation is allowed, but disabling accounts is the customer's call, so it gets escalated.
A generic baseline would most likely give all three the same answer. Your context is what makes the answer yours, it stays yours, and we never train on your data.
How does an AI defense get better over time without retraining on your data?
Through in-context learning. Your data stays in your tenant and gets read at run time when an investigation needs it, it never goes into a model's weights. When the defense learns something it learns it as a skill or a bit of context the next investigation reads, not as a new version of the model. People confuse fine-tuning and in-context learning a lot so here's the difference:
| Fine-tuning | In-context learning | |
|---|---|---|
| Your data | Has to go into the model's weights | Stays in your tenant, read at run time |
| A better model ships | Retrain, revalidate, redeploy | You get it the day it lands |
| Changing its behavior | Another training run | Takes effect on the next case |
| "Why did it decide that?" | Opaque weights | Points at the exact context it used |
| Model choice | Usually tied to the model you tuned | Any frontier model, swappable |
Every case can teach the next one. If the same kind of investigation keeps getting stuck on the same gap, the defense writes up the improvement itself, attaches the past investigations that show the problem and sends it for approval, and a person approves it before anything changes. Once an improvement has proven itself it can go into a regression suite so it keeps getting tested as other things change around it.
Here's a concrete one. On a production shadow tenant every query the agent wrote in XQL (the query language of Palo Alto Cortex XSIAM) failed at first, because it had never seen that language. Before anyone opened an engineering ticket the agent was asked to find the root cause, and what it came back with was a skill for the language. Over two iterations query success went from 0% to roughly 70%, and nobody retrained a model to get there.
Does it matter whether a defensive AI learns from synthetic data or real attack telemetry?
Yes. To learn how to defend, labelled synthetic data generally teaches more. To know what normal looks like in your environment, real telemetry is still the better record. Synthetic data has some well known weaknesses in machine learning (it can miss the messy distribution of a real environment, and its noise can look too clean) which is why a good defense still reads your real telemetry at run time.
Real attack telemetry has a gap when it comes to learning defense. It only records the path that worked. The paths an attacker tried and dropped leave nothing behind, and those are exactly what a defender needs, since an AI investigator that can't recognize a dead end tends to spend most of its time in one. Microsoft's Defender research team made a similar point in May 2026, labelling real attack logs "requires not only labeling malicious activities, but also fully reconstructing attack scenarios." Real incident data also hardly ever labels which defensive action stopped the attack.
Defensive training data does not exist in nature, so we manufacture it. We build a synthetic company with its own people, machines and daily traffic, then have one AI attack it while another defends. Because we authored both sides we know the attacker's goal, every path it tried, and which defensive action stopped it. The product ships into 300+ enterprise environments, and wherever it falls short of what your team would have done, that shortfall tells us where the lab was thin. We fix the lab and ship again, the way a vendor fixes a bug from a bug report. Your telemetry never enters the lab.
We never train on your data. A bug report is not your data, and neither is this. What comes back from production is a shortfall report, where the product fell short of what your team would've done, and not your logs.
Two things keep the lab getting better. Your environments show us where it was thin, and every stronger frontier model upgrades our attacker for free. Since specific attacks stop repeating, what the lab teaches matters a lot, and what it teaches is method. Method transfers, decisions don't.
What is a cyber range, and can it train defensive AI?
A cyber range is a controlled environment that simulates networks, systems and traffic so people can train and test without touching production. NIST's "The Cyber Range: A Guide" (NICE, September 2023) calls them places for cybersecurity training, exercises and testing. Most are built for human teams practicing incident response, or for testing tools against known scenarios. A range can help train defensive AI, usually just as a test bed though, since most of them only record whether the attack got stopped.
Ranges get used to test AI more and more. AISI's 32-step "The Last Ones" range from its April 2026 evaluation of Claude Mythos Preview is one, built to measure how far an AI attacker gets. AISI estimates a human needs about 20 hours on it and says its ranges "lack security features that are often present, such as active defenders and defensive tooling." Hack The Box launched HTB AI Range in December 2025 to benchmark AI agents on offense and defense.
To actually train a defense, a range would have to record what a defender needs to learn from. Most don't. A typical range scores whether the attack was stopped and the dead ends go unrecorded.
Simbian's war lab looks like a cyber range from outside, it isn't one. Attacker and defender are both AI, and because we author both sides every path is labelled, dead ends included. A range mostly trains people on scenarios. The war lab manufactures labelled defensive training data at a scale no human exercise calendar keeps up with.
Can an AI defense handle attacks it has never seen before?
Often, yes, as long as it reasons about behavior and doesn't just match patterns. AI threat detection based on anomalies and behavioral analytics can flag that something is new, but flagging it isn't the same as handling it. Somebody still has to figure out what the new thing is, whether it matters and what to do.
A defense that is designed to think, not just follow, investigates the new thing the way an experienced analyst would. It asks what the activity is reaching for, pulls evidence from whichever tools can answer that, and decides if there's enough to make the call. That works on a never-seen attack for the same reason it works for a senior analyst, a loader nobody has seen before still has to persist, still has to talk to something and still has to touch data, and an investigator can check each of those. Arctic Wolf Labs found exactly that in March 2026 when it went through more than 22,000 AI-assisted malware samples, the core execution, persistence and command-and-control behaviors were still detectable.
The real limit is that no defense can prove a negative. It can show that it looked, what it checked and why it concluded what it did, but it can't prove nothing happened in places it didn't look, which is why hunting and testing keep running alongside investigation.
How do pentest findings become better detections?
A pentest finding becomes a better detection once someone hunts for the proven path, checks it against what the SOC actually saw and turns it into a tested rule, all mapped to the same MITRE ATT&CK technique. What happens in most organizations instead: the finding becomes a ticket, the ticket sits in a backlog, detection engineering maybe gets to it weeks later. Meanwhile the environment stays exposed and often nobody checks if the SOC would have caught someone exploiting it.
The four-node loop closes that gap, each node answers one question and passes the result on:
| Node | The question | What it produces |
|---|---|---|
| AI Pentest Agent | What could happen? | A proven, reachable attack path |
| AI Threat Hunt Agent | Did it happen? | Evidence of whether the technique was already used here |
| AI SOC Agent | Did we respond? | Whether the activity was detected and handled correctly |
| AI Detection Engineering Agent | Can we catch it next time? | A candidate detection, tested against the path that found the gap and approved by a person before it ships |
Every finding maps to the same ATT&CK technique ID, so one scoreboard shows all four answers. Pentest and SOC share findings both ways. A SOC judgment can change pentest scope and severity, a pentest finding tells the SOC what to watch for. Because the loop runs the same way every time, coverage builds from one engagement to the next. The AI pentest guide goes deeper on the testing side.
What is the difference between red, blue, and purple teams?
Red teams attack, blue teams defend. Purple teams (in most organizations more a practice than a separate team) turn what the red team did into better detection. The names come from military exercises and mainly describe goals and outputs:
| Red team | Blue team | Purple team | |
|---|---|---|---|
| Goal | Get in the way a real attacker would | Detect, investigate, and respond | Turn what the red team did into better detection |
| Typical cadence | Periodic engagements | Continuous | Scheduled exercises |
| Main output | A report of paths and findings | Investigations and responses | Detections improved and retested |
Wiz put it well in an August 2026 guide: purple teaming is "a collaborative validation loop: emulate a realistic procedure, observe what the defensive stack sees, improve the control, and retest." Cadence is the weak spot. The loop runs once per exercise, a few times a year if you're lucky, and attackers don't care about anyone's exercise calendar.
Self-improving defense runs the purple loop whenever a technique needs testing, one technique at a time, red and blue on the same MITRE ATT&CK map. Breach and attack simulation tools also make testing more frequent. The question that matters is whether each finding ends up as a tested detection. Keep the human red team, people come up with creative attacks a loop may never think of.
How do you stop detections from silently failing?
Mostly by keeping an eye on the three things that quietly degrade most SOCs, and fixing them before a gap turns into a missed attack. Simbian's own 2026 analysis found roughly 20% of SOC detection rules stop firing within six months. Detections hardly ever fail loudly, they just stop firing and nobody notices until an audit or an incident turns it up.
The first is the detection rules themselves. Rules go noisy, or break, or never existed for a technique that matters, and a defense that reads every verdict can usually see which rules are noisy or missing and propose tuned or new ones written in your SIEM's own query language (KQL, SPL, XQL and so on).
The second is data-pipeline drift, where a log source goes quiet or a field changes shape. One renamed field after an overnight SIEM upgrade can silence 40 detections with no error and no alert.
Integrations are the third. A connector breaks after a vendor update or a threat-intel feed lapses, enrichment comes back empty and verdicts keep shipping without it.
Teams that catch these early commonly keep a fire-rate baseline for each rule, replay a known-bad test event through the pipeline now and then, and alert when a field they rely on starts coming back empty.
Coverage decays. Simbian repairs. The repairs go through the same governance as everything else, detection changes get logged with the verdict that triggered them and a person approves them before they ship, and they can be rolled back afterwards.

