Evaluating and Adopting Self-Improving Defense
Part of Self-Improving Defense: AI Cyber Defense Against AI Attacks — read the full guide.
If every vendor uses the same frontier models, what makes one AI defense better than another?
In AI cyber defense the harness, the training data and the context make the difference, because the model itself is turning into a commodity. In an April 2026 post a group of Forrester analysts including Jeff Pollard put the question plainly: does your value proposition survive when frontier model access becomes ordinary?
The model doesn't decide the outcome on its own. The same model usually behaves differently depending on the harness around it, what context it gets and when it's allowed to stop, and the Cyber Defense Benchmark shows how much stopping early by itself can cost. The model is one cylinder. The harness is the car.
What tends to separate one AI defense from another:
- What trained it: labelled defensive data, or just general model training plus prompts.
- What it remembers: whether it builds a lasting model of your environment and your team's decisions, and if you can see and control that memory.
- What it can prove: published benchmark results with an edition date, plus measured results in your own environment.
How can you check a vendor's claim that its AI stops AI attacks?
By asking for three things you can check, what the AI was trained on and how its defensive data got labelled, a dated public benchmark on real attack logs, and evidence from your own environment. A demo isn't one of them.
Start with training. Real-world telemetry rarely records which action stopped an attack, so a vendor should be able to explain where its labels came from.
Then a published benchmark on real attack logs (the Cyber Defense Benchmark, CyberSOCEval or ExCyTIn-Bench, for example) with the edition and date stated. An unpublished internal number, or a capture-the-flag style score built for offensive tasks, doesn't say much about defense.
And evidence in your own environment, things like techniques replayed against your stack, verdicts compared side by side with your analysts, override rates by alert type, and an export of the reasoning behind whichever verdict you pick. Mature systems can export the reasoning behind any verdict and immature ones usually can't.
"We protect you from AI attacks" is a positioning statement coming from any vendor, us included. The evidence behind it is what turns it into something you can rely on.
How do you test an AI defense against realistic attacks before you trust it?
Replay MITRE ATT&CK techniques, evasion variants included, against your own stack, and run it on your own live alerts with every action gated so you can compare what it concludes with what your analysts conclude. A proof of value in a vendor's lab mostly tells you about the vendor's lab, a test in your environment tells you how it handles your tools, your data quality and your noise.
A first month typically goes roughly like this:
- Week 1: the agents connect to your tools and onboard themselves, discovering tables, fields and schemas.
- Week 2: investigations run on live alerts and every action still waits for approval.
- Weeks 3 to 4: you compare its verdicts against your own analysts' verdicts on the same alerts.
- After that: gates get lifted one action type at a time, as the record supports it
The replayed techniques show what gets detected and how it's investigated, the live alerts show how it copes with your actual noise. You generally want both before relying on it.
Does an AI defense require you to centralize your data first?
No, a well built one reads your data where it already lives. Gartner lists it among the questions to ask AI SOC agent vendors: does the solution require data centralization, or can it operate in any environment? Needing a new data lake before investigations can even start often turns a deployment into a data project that runs for months.
Simbian reads from more than 100 of the tools you already run (SIEMs, EDR, XDR, cloud and identity systems, ticketing) and reasons across them as one picture. It doesn't need new sensors or agents deployed, a data migration, or your detection stack replaced or re-tuned first.
Data readiness still matters though. An investigation needs at least the alert source, endpoint telemetry, identity logs and some asset context to get to a confident verdict, and gaps there show up as weaker investigations. That's a reason to know where your telemetry has holes, not a reason to move it all into one place first.
Which SOC use case should you trust an AI agent with first?
Start where volume is high and the actions are reversible, which for most teams means investigating alerts to a verdict. Investigating an alert doesn't change anything in your environment, so a wrong verdict costs an analyst some review time and not an outage, and with the volume the verdict comparison in weeks three and four gives you a decent sample to judge it on.
Keep the response side small at first too, starting with whatever's easiest to undo. A low override rate on an alert type, held over time, is the signal that its response can run on its own. Phishing triage and endpoint alert validation are common next steps since the evidence is concrete and the actions (quarantining an email, say) are easy to reverse.
Expansion usually follows results. One customer started with the SOC and added pentesting and threat hunting within 12 months.
Where will new analysts learn the craft once AI takes the queue?
Mostly from supervising and correcting the agents, which turns out to be a decent apprenticeship. A new analyst's first year shifts toward reviewing finished investigations, fixing the wrong ones and approving or rejecting proposed actions. The tier 1 queue used to be where people learned what normal and malicious look like. Once agents take the tier 1 and tier 2 queues and the playbook work, people move up to building, running and governing the agents, and the learning moves up with them.
In most teams that ends up as three roles. The AI SecOps Analyst supervises the agents, reviews their investigations, approves actions and works whatever the agents send up, and reading a finished investigation then deciding whether it's right tends to teach the craft faster than clearing a queue ever did. The AI Skill Manager often comes from senior analysts and SOAR engineers and encodes the team's knowledge into skills the agents run, with no coding needed. Then there's the AI SecOps Manager, who owns rollout, governance and service levels across every SecOps program, not only the SOC.
In a September 2026 post Forrester's Jess Burn and Jeff Pollard asked the same question from the other side: "Who develops the experts these systems will continue to require?" Probably the people reviewing and correcting the agents every day. Their corrections are what make the defense better and doing that work well is how a new analyst becomes a senior one.

