Loading...
Loading...
Claude Haiku 5.5 scores 9.3% coverage on Simbian's Cyber Defense Benchmark of agentic threat hunting, 28th of 30 models and 8.3 points below Claude Haiku 4.5. Anthropic's newest small model runs a threat hunt for about a cent, then stops after roughly 10% of its turn budget. Claude Opus 5 leads the same benchmark at 45.1%, for $3.84 per run.
Three weeks ago, a startup called TypeSafe shipped Jev, a model that returns only a typed decision with a probability attached, for about four cents per million input tokens. On October 6, OpenAI opened its Decisions API, running on GPT-6 Luna, to all developers in public beta. Anthropic shipped Claude Haiku 5.5 the next day.
Neither Anthropic's launch post nor its system card mentions Jev. The launch benchmarks Haiku 5.5 against GPT-6 Luna, and the model docs pitch it "for high-volume, latency-sensitive tasks such as classification, extraction, and routing". That is the decision-model job description. We read Haiku 5.5 as Anthropic's entry in that race, so we asked whether it can hunt. It can, briefly. Then it stops, and Anthropic's own system card measures the same habit.
We ran it on our Cyber Defense Benchmark (arXiv 2604.19533), which scores fully agentic AI threat hunting. The figures below are from the October 2026 leaderboard of 30 models:
Every model ran under the same harness against the same 26 campaigns, and Haiku 5.5 ran three times per campaign through Amazon Bedrock. The comparison model is Claude Opus 5, the benchmark leader. The newer Claude Opus 5.5 refuses threat-hunting work and could not get past the first turn.
Haiku 5.5 finds 9.3% of the malicious activity across 26 attack campaigns, which places it 28th of the 30 models on the board.
| Model | Coverage | Rank (of 30) |
|---|---|---|
| Claude Haiku 5.5 | 9.3% | 28 |
| GPT 5.6 Luna | 15.7% | 24 |
| Claude Haiku 4.5 | 17.6% | 23 |
| GPT-6 Luna | 21.6% | 19 |
| Claude Opus 5 | 45.1% | 1 |

GPT-6 Luna, the model behind OpenAI's Decisions API, finds 2.3× as much, and GPT 5.6 Luna, a generation older, still finds 1.7× as much. Haiku 4.5, the model this release replaces, finds nearly twice as much.
If Haiku 4.5 sits anywhere in your detection pipeline today, the upgrade is a downgrade on defense.
Haiku 5.5 spends about 10% of its turn budget on a threat hunt before deciding it is finished, while Claude Opus 5 spends nearly all of it. In a SOC, an early stop looks exactly like a clean hunt. Every turn it skips is a query it never runs, so whatever the attacker did in the unread part of the kill chain stays unread.
We have seen this before. Claude Opus 4.8 used only 60% of its budget and lost coverage against Opus 4.6 in the same way. Early stopping is the likeliest reason for Haiku 5.5's score, since a hunt that ends after a handful of turns reads very few logs.
When it does search, it mostly works from memory. Just over half of Haiku 5.5's searches, 50.9%, come from recall. They name mimikatz, certutil, procdump, and the other attacker tools every detection course teaches. When one of those searches lands, the hunt ends soon after. In the model's own words: "The hunt is complete from my side."
Anthropic has measured the same habit. In the Claude Haiku 5.5 system card, its ExploitBench harness includes an "AutoNudge" arm. "If the model voluntarily stopped short of the budget," the harness injects a keep-trying prompt, and with the nudge, Haiku 5.5's ExploitBench Cap% rises from 38% to 49%. Anthropic's prompting guide for Haiku 5.5 says the same about long agent prompts at low effort: "it sometimes stops early and hands the task back to the user." Raising effort from low to medium "roughly halved early stopping."
A hunt report reads the same whether the model ran 5 queries or 50, so we report turns spent next to coverage for every model we test.
Haiku 5.5 costs about one cent per benchmark run on average, the cheapest result in the cohort, against $3.84 for Claude Opus 5.
| Model | Mean cost per benchmark run |
|---|---|
| Claude Haiku 5.5 | $0.01 |
| GPT 5.6 Luna | $0.08 |
| GPT-6 Luna | $0.10 |
| Claude Haiku 4.5 | $0.34 |
| Claude Opus 5 | $3.84 |

Two things drive that number. The first is Claude Haiku pricing itself: Anthropic lists Haiku 5.5 at $0.10 per million input tokens and $0.50 per million output tokens for prompts up to 100,000 tokens. The second is the early stop, which keeps every hunt short.
Divide coverage by cost and Haiku 5.5 returns more per dollar than any model we have measured. That ratio flatters a model that finds less than a tenth of the attack.
It's also the most careful of the five models in this comparison. 82.1% of the events Haiku 5.5 flags are actually malicious, where the other four sit in a tight band between 54% and 63%. (Precision here is the share of flagged events that match ground truth.) That is what you would expect from a model that barely flags anything.
For your analysts, that means a quiet console that is right when it speaks and silent on more than 90% of what the attacker did.
Claude Haiku 5.5 almost never finds a threat that Claude Opus 5 misses. Opus 5 finds nearly everything either Haiku model finds, so the ladder of Haiku 5.5, Haiku 4.5, and Opus 5 comes close to nesting.
Each cheaper model in the Anthropic Haiku line sees a smaller slice of the same ground the larger model already covers. That matters if you hoped to pair a cheap model with an expensive one for wider coverage, since on this benchmark the pair adds little beyond Opus 5 on its own.
Most LLM security benchmarks measure attacking or patching, and Haiku 5.5's own release shows the cost of that. Anthropic's system card reports four cyber evaluations (ExploitBench, CyScenarioBench, a binary exploitation benchmark, and ExploitGym), and all four center on exploitation. On those, Haiku 5.5 "significantly outperformed Haiku 4.5." On our defensive benchmark, with a different harness, it trails Haiku 4.5 by 8.3 points.
Offense went up. Defense went down.
Whatever this release was tuned for, cyber defense wasn't near the top of the list. We can't tell from the outside whether that was a deliberate trade for speed and price or a side effect of training hard for coding.
The rest of the field follows the same pattern. CyberGym from UC Berkeley asks a model to reproduce 1,507 known vulnerabilities. CyberGym-E2E adds discovery and patching. Google's Gemini 4 Argon release cites CWE-bench, a remediation benchmark where Argon ties for first at 68%.
None of these LLM benchmarks asks a model to take 100,000 events, find the 5% that belong to an attacker, and work out what they touched. Alert investigation and threat hunting, the work your SOC does every day, stay unmeasured, so LLM cybersecurity claims about defense still rest on offense scores.
That is why we built the Cyber Defense Benchmark to double as an RL environment. Every hunt is scored deterministically against ground truth, which gives model builders something to hill-climb on for alert investigation and threat hunting.
No. Sampling Haiku 5.5 16 times (pass@16) lifts its Cyber Defense Benchmark coverage from 9.3% to 17.0%, still far short of the 45.1% Claude Opus 5 reaches in one run. Spending test-time compute on more attempts helps a little, but it can't fix a model that quits early and searches from memory, because 16 short hunts still cover much of the same obvious ground.
| Run | Coverage |
|---|---|
| Haiku 5.5, pass@1 | 9.3% |
| Haiku 5.5, pass@16 | 17.0% |
| Claude Opus 5, pass@1 | 45.1% |

The pass@k metric comes from code generation, where a problem counts as solved if any of k attempts solves it. On code, a test suite tells you which attempt passed. A threat hunt has no test suite. To collect what 16 hunts found, you have to run all 16 and triage every flag they raise, at 16 times the cost, and you still end up with less than one Opus 5 run.
That matches research on the limits of inference scaling. Stroebl, Kapoor, and Narayanan show that when the checker is imperfect, "no amount of inference scaling of weaker models can enable them to match the single-sample accuracy of a sufficiently strong model."
More compute doesn't buy a capability the model lacks. Haiku 5.5 posts 83.7 on SWE-bench Multilingual, per its system card, and that tells you nothing about whether it will find lateral movement in 100,000 noisy logs. You have to match the model to the task and to how hard the task is.
Claude Haiku 5.5 is a partial Jev competitor. It can do the classify, extract, and route job that TypeSafe's Jev is built for, and it can also start a multi-step threat hunt, which Jev isn't built to do. It stops that hunt long before the work is done, so in a SOC it belongs on classification, with investigations left to a model that keeps going.
Its real advantage is cheap, confident, shallow work, close to what Anthropic built it for: "high-volume, latency-sensitive tasks." It also ran every one of our benchmark runs without a single refusal or safety-classifier block.
The weaknesses are just as clear. Haiku 5.5 is opportunistic, and its results vary widely from run to run. It is a strong candidate for reinforcement learning with verifiable rewards (RLVR) fine-tuning on defensive tasks, though nobody has measured that lift yet. Until then, here is where it fits in your SOC:
Q: Is Claude Haiku 5.5 good for cybersecurity? Claude Haiku 5.5 is good for narrow, high-volume security steps, but not for threat hunting. On Simbian's Cyber Defense Benchmark it finds 9.3% of attacker activity across 26 campaigns, 28th of 30 models, and stops after about 10% of its turn budget. It is cheap and precise, so it suits routing, enrichment, and first-pass triage with a stronger model behind it.
Q: Is Claude Haiku 4.5 better than Haiku 5.5 for security? Yes, on defense: Claude Haiku 4.5 finds 17.6% of attacker activity on Simbian's Cyber Defense Benchmark, where Haiku 5.5 finds 9.3%, 8.3 points lower. Haiku 5.5 beats Haiku 4.5 on Anthropic's offensive cyber evaluations, so the answer depends on which side of the job you need. The strongest Claude model we have tested on defense is Claude Opus 5, at 45.1%.
Q: Is Claude Haiku 5.5 better than GPT-6 Luna? Not for threat hunting. On Simbian's Cyber Defense Benchmark, OpenAI's GPT-6 Luna, the model behind its Decisions API, finds 21.6% of attacker activity across 26 campaigns, 2.3 times Haiku 5.5's 9.3%. Haiku 5.5 is cheaper per hunt, about $0.01 against $0.10 on average. Neither comes close to Claude Opus 5 at 45.1%.
Q: Why not just upgrade to the newest model? Because on defense, Haiku 5.5 scores 8.3 points below Haiku 4.5 on the Cyber Defense Benchmark, and Claude Opus 4.8 lost coverage against Opus 4.6 the same way, using only 60% of its turn budget. Defensive skill can drop between releases, so measure each release on your own defensive tasks before it replaces the model you run today.
Q: Why does threat hunting need its own benchmark? Because published LLM security benchmarks measure vulnerability work, while a SOC's job starts after an attacker is through. A defensive benchmark also catches refusals. Analyzing a real intrusion, Hugging Face found that "Guardrails on Opus tripped every time we tried to analyze the attack logs." Claude Opus 5.5 refuses our threat-hunting runs, while Haiku 5.5 ran all of them clean.
Q: How much does Claude Haiku 5.5 cost? Anthropic lists Claude Haiku 5.5 at $0.10 per million input tokens and $0.50 per million output tokens for prompts up to 100,000 tokens, and $0.50 and $2.50 above that. On Simbian's Cyber Defense Benchmark, a full threat hunt costs about $0.01 on average, against $3.84 for Claude Opus 5.
Haiku 5.5 is a good classifier that can also hunt, briefly. Use it where a cent and a confident answer are enough, and check every step you hand it against ground truth before it touches an investigation. Simbian, which ships no model of its own, measures every model it runs and builds around what each one gets wrong. Our AI Threat Hunt Agent is built the same way, with your team's tribal knowledge kept in a Context Lake™. Book a Demo to watch it finish a hunt on your data.