Loading...
Loading...

Anthropic shipped Claude Opus 5 this week — and on the Cyber Defense Benchmark the king is back. Against 20 other frontier models on real Windows attack campaigns — deterministic ground truth, a provider-agnostic harness, fully agentic threat hunting — Opus 5 retakes #1 at 45.1% coverage.
But the headline number isn't the interesting part. Every frontier lab claims the top of some leaderboard on launch day. Three things in our data are worth more than the ranking: Opus 5 did the defensive work without refusing a single time, the offense-vs-defense capability gap inverts depending on which job you ask for, and Anthropic now owns the entire top-tier token-efficient frontier for defense.
On coverage — the fraction of malicious events a model finds and submits across a full attack kill-chain — Opus 5 edges out Opus 4.6 and pushes open-weight leader GLM 5.2 out of the top 5.

| Rank | Model | Coverage | Type |
|---|---|---|---|
| 1 | Claude Opus 5 | 45.1% | Proprietary |
| 2 | Opus 4.6 | 44.5% | Proprietary |
| 3 | GPT 5.6 Sol | 41.0% | Proprietary |
| 4 | GPT 5.5 | 37.4% | Proprietary |
| 5 | Sonnet 4.6 | 36.9% | Proprietary |
| 6 | GLM 5.2 | 35.0% | Open-weight |
It also repairs the Opus line's own dip. Coverage sagged at Opus 4.7 (30.9%) and Opus 4.8 (32.7%) — the middle of the pack on defense — while Opus 4.6 quietly held the ceiling at 44.5%. Opus 5 restores the top of the board and nudges it forward. It's a small step in points (45.1 vs 44.5), but it's the first Opus release in three that moves the frontier the right direction on defense instead of trading capability for cost.
Worth keeping in perspective: #1 is still 45.1%. The best defensive model on the market finds fewer than half the malicious events in a realistic attack. More on that below — it's the whole reason this benchmark exists.
Here's the result that actually changes what you can deploy. Across all 26 campaigns and 1,353 agent turns, Opus 5 refused nothing. We checked the served-model metadata on every single request: all 1,353 turns were answered by claude-opus-5 itself, with zero empty responses, zero safety stops, and zero errors.

Why this matters: the most capable model isn't the one that tops a leaderboard once — it's the one that will actually do defensive security work every time you ask. Anthropic's own most-capable model, Fable 5, won't. The moment you ask it to search noisy logs for signs of an attack, it back-drops to Opus 4.8 — silently swapping in a smaller model because the request pattern-matches to "offensive-looking" cyber use case. For a SOC, an intermittent safety refusal on threat hunting is a non-starter; you can't run a detection pipeline on a model that sometimes declines to look.
Opus 5 just does the work — and Anthropic reports it as the least prompt-injectable frontier model they've measured, which is exactly the property you want in an agent that reads attacker-controlled log data all day. That combination — high coverage, no refusals, hard to inject — is the real unlock, and it's the kind of thing you only learn by running the model on the actual job.
The reason a defensive benchmark has to exist at all is that offensive capability doesn't predict defensive value — and the gap between US frontier models and the open-weight frontier looks completely different depending on which job you measure.

On offense — the kind of exploit-and-attack tasks that safety frameworks require labs to measure — the UK's AISI/CAISI assessment puts US frontier models at 76.2% versus the open-weight frontier (GLM 5.2) at 24.4%. That's a 3.1× advantage, measured with safeguards disabled to probe raw capability.
On defense — our benchmark, safeguards on, the way you'd actually run it — that same US-vs-open-weight gap collapses to 43.5% vs 35.0%: only 1.24×. The frontier's huge lead is on breaking in, not on catching the break-in.
That asymmetry has a direct consequence for defenders, and it's why we'll give kudos to Anthropic here: by improving their classifier to separate cyber offense from cyber defense, they let their strongest model do defensive work (Opus 5, 0 refusals) instead of reflexively refusing anything security-shaped. Without that separation, a defender who needs a non-refusing model is pushed toward open-weight only — which, on the defense number above, is the weaker choice. Getting the offense/defense distinction right is what keeps the best defensive tool available to defenders.
(Open-weight frontier here is GLM 5.2; Kimi 3's open-weight release isn't available yet. Offense figures are UK AISI/CAISI's, safeguards off; defense figures are ours, safeguards on.)
Coverage isn't free — every point of it costs output tokens, which is what you actually pay for higher (per token) and wait on longer. Plot coverage against output tokens per investigation and the efficient frontier — the best coverage you can buy at each output token budget — is almost entirely Anthropic.

From the mid-tier up, the frontier is Opus 4.7 → Opus 4.8 → Opus 4.6 → Opus 5, climbing to the highest coverage on the board. GPT 5.6 Terra and a couple of budget models hold the cheap tail, but every high-coverage point belongs to Anthropic. If you're buying defensive coverage, you're buying it from one vendor's frontier.
And here's the punchline that ties the whole post together: this is the exact opposite of the offense frontier. On offensive-cyber capability-per-token, the efficient frontier is dominated by OpenAI. Offense and defense don't just have different gaps between labs — they have different winners. You cannot read a model's defensive value off its offensive reputation, or off a system card. You have to measure the job you're actually going to run.
In the age of autonomous, agentic cyberattacks, response must happen at machine speed. When choosing an AI-powered, machine-speed cyber defense solution for your Security Operations Center (SOC), reliability, measurable performance, and the continuous learning of both agent workflows and environment-specific context should be top priorities.
Select a cybersecurity company that leads and contributes to AI security research specifically in the SOC defense space. Offensive security benchmarks, such as CyberGym and ExploitBench, track an entirely different business use case.
The platform should be self-evolving and agentic, improving through every interaction with its environment and every piece of feedback from human security analysts. These capabilities should be validated through rigorous, SOC-specific cyber defense benchmarks.
Currently, Simbian is the only provider offering a scalable, deterministic benchmark focused specifically on defensive SOC tasks and threat hunting. Book a demo.
Q: Is Claude Opus 5 the best LLM for cyber defense? On the Cyber Defense Benchmark, yes — it's #1 at 45.1% coverage across 21 models, ahead of Opus 4.6 (44.5%), GPT 5.6 Sol (41.0%), GPT 5.5 (37.4%), and Sonnet 4.6 (36.9%). But "best" still means finding fewer than half of a real attack's events, so the model alone isn't a SOC.
Q: Did Opus 5 refuse any of the security tasks? No. Across all 26 attack campaigns and 1,353 agent turns, every request was served by Opus 5 itself, with zero refusals, zero empty responses, and zero errors. That's a real contrast with Fable 5, which back-drops to Opus 4.8 when asked to hunt threats in logs.
Q: Why is offensive-security capability a bad proxy for defense? Because the two skills diverge. US frontier models lead the open-weight frontier 3.1× on offensive tasks (UK AISI/CAISI, safeguards off) but only 1.24× on defensive threat hunting (our benchmark, safeguards on). Finding and exploiting a vulnerability is a different job from catching an intrusion in noisy logs — and system cards mostly grade the former.
Q: How is coverage measured? Coverage is the fraction of malicious events across a full, causal attack kill-chain that the agent both finds and submits — 105 procedures, real Windows attack logs, deterministic ground truth. Full methodology is in the technical report and on the Cyber Defense Benchmark page.
Q: New models ship every week — how is this benchmark kept current? It runs as a standing pipeline. New frontier models and tiers are added as they ship (GPT 5.6, Gemini 3.6, now Opus 5), and the same environment doubles as an RL gym that keeps our defensive harness improving between releases. The full leaderboard and methodology are always on the Cyber Defense Benchmark page.
Q: What is the Cyber Defense Benchmark? It's Simbian's provider-agnostic benchmark that scores LLMs on defensive threat hunting — running each model as an agent against real Windows attack kill-chains (105 procedures) with deterministic ground truth, and measuring coverage (the malicious events a model finds and submits), cost, and investigation time. Unlike offensive evals such as CyberGym or ExploitBench, it grades the job a SOC actually does: finding an attack buried in noisy logs.
Q: What is AI threat hunting? AI threat hunting is using an LLM agent to proactively search security logs for the malicious events of an attack, instead of waiting for an alert to fire. The Cyber Defense Benchmark measures exactly this: how much of a real Windows attack kill-chain a model finds and submits. Even the best model to date, Claude Opus 5, covers only 45.1%, so AI accelerates human hunters rather than replacing them.
Q: Is Claude Opus 5 good for cybersecurity? For defensive SOC work, it's the strongest LLM cybersecurity option we've measured: #1 on the Cyber Defense Benchmark at 45.1% coverage with zero refusals across 26 campaigns. But no model clears half the events in a real attack, so Opus 5 is a force-multiplier inside an AI SOC — not a replacement for the surrounding detection system or the human analyst.