Proving Your AI Cyber Defense Gets Better

Part of Self-Improving Defense: AI Cyber Defense Against AI Attacks — read the full guide.

How do you measure whether your security is getting better?

Measure coverage, speed and how often your team agrees with what the defense does. Not how many alerts it processed. Volume is still what most teams report though, in the 2026 SANS SOC Survey the number of incidents handled was the top SOC metric for the tenth year running, and a team handling more incidents can be getting worse at stopping attacks.

Three things tell you more:

  • Coverage: of the attack techniques that matter to you, how many would really be detected, checked by replaying them
  • Time from detection to response (often reported as MTTR): how long a real threat runs before something is done.
  • Approval: how often your analysts accept the proposed response, by alert type.

Coverage only means something when you can see what moved it. At one Simbian customer MITRE ATT&CK detection coverage went from 33% to 56% to 83% over three loop cycles. Cycle one: the pentest side tested a set of techniques, the hunt looked for prior use in the logs, the SOC caught some and detection engineering shipped rules for the gaps. Cycle two retested those plus new ones. Cycle three ran evasion variants against the new rules. It went up because each cycle closed what the last one found.

Simbian frames it as objective, evaluation and evolution. You set an objective, investigate every alert of some type to a verdict for example. The defense scores itself against that on your own data and shows which cases fell short, then proposes changes to close the gap, and a person approves them. When an objective hits full marks it moves into a regression suite so it keeps getting tested.

Can you prove an AI defense still works after the model or tools change?

You can, as long as you actually test after every change. Newer models aren't automatically better at defense. On the Cyber Defense Benchmark (October 2026 edition) five recent releases scored under the model they replaced. Sonnet 5 was 17.0 points below Sonnet 4.6, Opus 4.8 was 11.8 points below Opus 4.6. Several of them just stopped investigating sooner (Opus 4.8 stopped after about 34 turns where Opus 4.6 ran about 52) which made runs cheaper and left part of the attack unread.

So don't swap models on release day. Re-benchmark every release on defensive work before trusting it in production. Simbian ships no model of its own. It measures every model it runs and builds around what each one gets wrong.

Everything around the model changes too. Tools get upgraded, schemas shift, connectors break. Two things keep a defense provably working through that. A regression suite, where objectives that already score well stay under test so one improvement can't quietly cost you somewhere else. And self-repair for the plumbing, which notices when a connector or schema changes and re-learns it, with review and rollback before anything ships.

What is the difference between CTEM and the SOC?

CTEM finds where you're exposed, the SOC detects when someone uses that exposure. Gartner's February 22, 2024 press release defines continuous threat exposure management (CTEM) as "a pragmatic and systemic approach organizations can use to continually evaluate the accessibility, exposure and exploitability of digital and physical assets," and predicted organizations prioritizing security investment based on a CTEM program would see a two-thirds reduction in breaches by 2026. A CTEM program is usually described in five stages: scoping, discovery, prioritization, validation and mobilization.

They answer different questions and usually sit in different teams with different tools:

CTEM SOC
Question it answers Where could we be attacked? Is someone attacking us right now?
Main inputs Asset inventory, vulnerabilities, attack paths, validation tests Alerts, logs, endpoint and identity telemetry
Typical output A prioritized list of exposures to fix Investigations, verdicts, and responses
Cadence Cycles of scoping and validation Continuous

Some vendors argue CTEM should become the SOC. The two really do answer different questions though, and in practice the problem is they rarely share a scoreboard, a CTEM program will often validate that an attack path is real and nobody checks if the SOC would catch someone using it. Putting both on the same MITRE ATT&CK map fixes that, every exposure and every detection maps to a technique so you can see where you're exposed and also blind.

How will breach and attack simulation improve threat detection and response?

Breach and attack simulation (BAS) improves detection when somebody acts on what it finds. BAS and automated pentest tools such as Cymulate, Picus and Pentera run known attack techniques safely against your environment and report which ones your controls blocked or detected. That's useful, it swaps guessing about control coverage for an actual test.

Whether anything improves still depends on what happens after the report. A finding that the EDR missed a technique only turns into better detection if someone writes or tunes the rule and reruns the test to confirm it worked. In a lot of teams that handoff is the slow part, and BAS platforms need ongoing care to keep their scenarios current. In its Market Guide for Adversarial Exposure Validation Gartner puts BAS and automated penetration testing under one category, which it describes as providing "consistent, continuous and automated evidence of the feasibility of an attack."

BAS usually stops at the report. A loop keeps going, the missed technique becomes a hunt for whether it was already used, then a new or tightened detection and a retest, all on the same MITRE ATT&CK technique map as everything else.

Is a purple team exercise worth it if you've already done red teaming?

Usually yes, as long as the output is a change to your detections and not just another report. A red team exercise tells you how someone got in. A purple team exercise runs the same MITRE ATT&CK techniques while the defenders watch, so the team sees what each tool caught and what it missed and can fix gaps while it's fresh.

Cadence is usually the problem. Exercises happen once or twice a year, attackers don't keep that schedule. A detection fixed in spring may have drifted by autumn, a log source moved, some field got renamed in a SIEM upgrade, someone edited the rule.

A continuous purple loop handles that, testing one technique at a time all year and retesting after each fix. Keep the exercises for the creative work humans are best at and let the loop do retesting, that way the lessons from each exercise actually stay fixed.

What do defensive AI benchmarks actually tell you?

A defensive AI benchmark tells you how a model does at AI cyber defense in a controlled test. Useful, but it's not the same as how a product does in your environment. Simbian's Cyber Defense Benchmark has models investigate real attack logs with deterministic ground truth, across 13 of the 14 MITRE ATT&CK tactics and 105 attack procedures, each environment holding more than 100,000 events. To pass, a model needs at least 50% coverage on every tactic. As of October 2026 none do, and of 30 models the leader reaches 45.1%. Other public defensive benchmarks exist as well, like CyberSOCEval from Meta and CrowdStrike, and Microsoft's ExCyTIn-Bench.

SentinelLabs made a fair critique of the defensive benchmarks it reviewed (CyberSOCEval and ExCyTIn-Bench among them). In January 2026 it argued "Benchmarks are Measuring Tasks, not Workflows": a benchmark can check whether a model finds evidence in logs, a real SOC also runs on context, business judgment and handoffs between people and tools.

So use both kinds of evidence. Benchmark the model to know its ceiling and blind spots before trusting it. Then measure the deployment on your own alerts, replayed techniques, verdicts compared with your analysts, override rates over time. A benchmark score with an edition date is a published claim anyone on your team can go and check.

Is MITRE ATT&CK coverage a meaningful measure of your detection?

It's meaningful as a map and often misleading as a percentage. MITRE ATT&CK gives every technique an ID so findings from pentests, hunts and investigations all land in the same coordinate system, which is valuable. The trouble is the coverage heat map, where a technique turns green because a rule exists for it whether or not that rule would catch a real attacker.

MITRE's Center for Threat-Informed Defense tackles this with its Summiting the Pyramid work on detection robustness, which it describes as "how difficult it is for adversaries to evade a detection." A rule matching one tool's command line is easy to get around. A rule watching the underlying behavior is a lot harder to slip past, and the heat map can't tell those two apart.

A more useful version of coverage is measured by replaying techniques, evasion variants included, and counting what gets detected, cycle over cycle. That's how one Simbian customer's coverage moved from 33% to 56% to 83% over three loop cycles, each cycle tested against the rules the previous one wrote. Coverage measured like that tells you whether you're getting better, coverage counted off a rule inventory mostly tells you how many rules you've got.

Sign up for Simbian's Newsletter

By submitting this form, you agree to our Privacy Policy.

Ask AI about Simbian