Loading...
Loading...
Grok 4.7 cybersecurity performance, measured: 37.6% coverage on Simbian's Cyber Defense Benchmark — rank 5, 3.7 points below Grok 4.6's 41.3%. SpaceXAI ships it at the same list price as Grok 4.6 and it runs a hunt twice as fast, but it finds less. The mechanism is measurable. 96% of its turns spend zero reasoning tokens, and by 20 prior turns of conversation history it reasons zero on every call.
Elon Musk said on Sep 14 that "Grok 4.7 should be roughly on par with Opus 5.0" [NEEDS SOURCE: link the post — a quoted competitor CEO with no retrievable source is the highest-liability sentence in the piece]. Sorry, but compared to Grok 4.6 it's a regression for cyber defense.
We ran Grok 4.7 on our Cyber Defense Benchmark (arXiv 2604.19533). It measures fully agentic threat hunting:
Both models ran under identical conditions: reasoning_effort at SpaceXAI's (xAI's) documented default of high, 50 queries, a 10-row result limit, parallelism 1, the same 26 seeds, and a byte-identical hunt path. This is an A/B. Nothing differs between the two runs except the model.
| Metric | Grok 4.7 | Grok 4.6 | Better |
|---|---|---|---|
| Cyber Defense Benchmark rank | 5 | 3 | Grok 4.6 |
| Coverage (malicious events found and submitted) | 37.6% | 41.3% | Grok 4.6 |
| Reasoning tokens per hunt | 217 | 23,182 | Grok 4.6 |
| Turns spending zero reasoning tokens | 96% | 4% | Grok 4.6 |
| Investigation time per hunt | 3 min | 7 min | Grok 4.7 |
| Cost per hunt | $1.4 | $1.6 | Grok 4.7 |
| Single cold cyber question, reasoning tokens | 2,900 | 1,939 | Grok 4.7 |
For cyber defense, Grok 4.6 beats Grok 4.7 on every quality metric, and Grok 4.7 wins only on speed, price and single-question sharpness. Faster and slightly cheaper. It finds less.
[OPERATOR CONFIRM — Igor: the draft's false-positive row (34.8% vs 31.3%) is cut pending a source. The benchmark's own data files carry no false-positive, precision or FP-rate column for any model; the only flag metric is Flags % Mean, which is 4.08% for Grok 4.6 (cyber-defense-benchmark.csv:4, model_summary.csv:4). Either supply the derivation and N so it can be defined in-post and added to the layer0 CDB file, or leave it cut.]

Not one of the 27 frontier models on Simbian's Cyber Defense Benchmark clears 50% coverage of a real attack kill-chain. The leader, Claude Opus 5, reaches 45.1%. Grok 4.7 reaches 37.6%. Coverage is the fraction of malicious events a model finds and submits across a full attack kill-chain.
| Rank | Model | Type | Coverage | Cost | Time |
|---|---|---|---|---|---|
| 1 | Opus 5 | Proprietary | 45.1% | $3.8 | 16 min |
| 2 | Opus 4.6 | Proprietary | 44.5% | $2.7 | 7 min |
| 3 | Grok 4.6 | Proprietary | 41.3% | $1.6 | 7 min |
| 4 | GPT 5.6 Sol | Proprietary | 41.0% | $4.0 | 12 min |
| 5 | Grok 4.7 | Proprietary | 37.6% | $1.4 | 3 min |
| 6 | GPT 5.5 | Proprietary | 37.4% | $4.5 | 8 min |
| 7 | Sonnet 4.6 | Proprietary | 36.9% | $2.2 | 8 min |
| 8 | GLM 5.2 | Open-weight | 35.0% | $2.0 | 8 min |
| 9 | DeepSeek 4 Pro 0813 | Open-weight | 34.0% | $1.0 | 11 min |
| 10 | Kimi K3 | Open-weight | 33.0% | $2.4 | 9 min |
| 11 | Opus 4.8 | Proprietary | 32.7% | $1.7 | 4 min |
| 12 | Opus 4.7 | Proprietary | 30.9% | $1.6 | 4 min |
| 13 | DeepSeek 4 Flash 0731 | Open-weight | 29.7% | $0.2 | 7 min |
| 14 | Qwen 3.8 Max | Open-weight | 29.5% | $1.6 | 9 min |
| 15 | Gemini 3.7 Flash | Proprietary | 29.1% | $0.8 | 8 min |
| 16 | Kimi 2.7 Code | Open-weight | 26.7% | $1.5 | 4 min |
| 17 | GPT 5.6 Terra | Proprietary | 22.5% | $0.8 | 3 min |
| 18 | Sonnet 5 | Proprietary | 19.9% | $1.9 | 6 min |
| 19 | Qwen 3.7 | Open-weight | 19.1% | $0.2 | 2 min |
| 20 | Gemini 3.6 Flash | Proprietary | 18.7% | $1.3 | 3 min |
| 21 | Haiku 4.5 | Proprietary | 17.6% | $0.3 | 3 min |
| 22 | GPT 5.6 Luna | Proprietary | 15.7% | $0.1 | 2 min |
| 23 | Minimax 3 | Open-weight | 15.5% | $0.3 | 3 min |
| 24 | Gemini 3 Flash | Proprietary | 14.6% | $0.2 | 1 min |
| 25 | Qwen 3.8 27B | Open-weight | 10.6% | $0.2 | 45 min |
| 26 | Nemotron 3 Super | Open-weight | 9.1% | $0.1 | 3 min |
| 27 | Gemini 3.5 Flash-Lite | Proprietary | 8.7% | $0.03 | <1 min |
Across the full LLM cybersecurity leaderboard, best to worst spans 45.1% to 8.7%, a 5× range on the same benchmark. No model choice closes a gap that starts at 45.1%. That is why AI threat hunting in production is a harness problem before it is a model problem.

| Where the money goes | Measurement |
|---|---|
| Input share of the bill | 87–93% |
| Output share of the bill | 7–13% |
| Cache miss rate, Grok 4.6 → 4.7 | 5.8% → 11.0% |
| Fresh input share of spend, 4.6 → 4.7 | 17% → 31% |
| Saved by halving output | −6% |
| Added back by cache misses | +11% |
Grok 4.7 runs a hunt twice as fast as Grok 4.6 but costs only about 12% less. An agent re-reads a growing transcript every turn, so input is 87–93% of the bill and output only 7–13%. List pricing is not the variable here. SpaceXAI serves Grok 4.7 at the same $2 per million input and $6 per million output as Grok 4.6, so every difference below is token volume and cache behaviour. The newer model halved its output, which is what makes it fast, and that accounts for about 6 points of the saving.
What actually moved the bill was caching. The cache miss rate went from 5.8% to 11.0%, nearly double, pushing fresh input from 17% to 31% of spend. Those extra misses added 11% back, almost twice what the shorter output saved. Whether that is a property of the model or simply a new release being too popular to keep prefixes warm, we cannot tell from the outside.
Grok 4.7 spends 217 reasoning tokens per hunt against Grok 4.6's 23,182 (107× less thinking) while writing 61% more visible text. Reasoning tokens are billed at the same price as output tokens and they can be separated out. Just over two-thirds of Grok 4.6's output per hunt is thinking (23,182 of 33,720 tokens). For Grok 4.7 it is one-eightieth, 217 of 17,163.
| Output per hunt | Grok 4.7 | Grok 4.6 |
|---|---|---|
| Reasoning tokens | 217 | 23,182 |
| Total output tokens | 17,163 | 33,720 |
| Reasoning share of output | 1.3% | 68.8% |
| Visible text written | +61% vs 4.6 | baseline |
Grok 4.7 does the talking without the thinking.

Grok 4.6's reasoning per turn climbs from 79.7 tokens in the first ten turns of a hunt to 1,088.6 in turns 41–50, while Grok 4.7's stays flat between 0.0 and 13.8. Mean reasoning tokens per turn across all 26 campaigns, in ten-turn blocks:
| Turns | Grok 4.6 | Grok 4.7 |
|---|---|---|
| 1–10 | 79.7 | 13.8 |
| 11–20 | 133.5 | 1.2 |
| 21–30 | 263.0 | 1.0 |
| 31–40 | 537.9 | 0.0 |
| 41–50 | 1,088.6 | 5.7 |
Grok 4.6 accelerates. Its last ten turns average roughly 14× the reasoning of its first ten, which is exactly what you want as evidence piles up. Grok 4.7 goes flat. In 20 of 26 campaigns it has already spent 90% of its thinking by its second turn, and 96% of its turns spend no reasoning at all.
Asked a single cyber question cold, Grok 4.7 reasons more than Grok 4.6. It spends 2,900 tokens against 1,939.
| Single question, asked cold | Grok 4.7 | Grok 4.6 | Better |
|---|---|---|---|
| Cyber (Sysmon triage) | 2,900 | 1,939 | Grok 4.7 |
| Math (3-digit / digit-sum) | 1,993 | 1,206 | Grok 4.7 |
| Logic (five houses) | 842 | 998 | Grok 4.6 |
The three prompts, verbatim, so you can reproduce them:
MATH: "A 3-digit number equals 11 times the sum of its digits. Find every such number and prove the list is complete."
LOGIC: "Five houses in a row, each a different colour and owner. The Brit lives in the red house. The green house is immediately left of the white. The Dane drinks tea. Who owns the fish? State assumptions if underdetermined."
CYBER: "Sysmon EID 1: rundll32.exe with no args, parent winword.exe; then EID 8 CreateRemoteThread into notepad.exe; then EID 3 to 3.133.92.167:443. Enumerate the ATT&CK techniques, order the kill chain, and say which event you pivot on first and why."
The variable is conversation history. Replaying real hunt transcripts at increasing depth across 4 seeds, median reasoning tokens:
| Prior turns in the window | Grok 4.6 | Grok 4.7 | 4.7 calls with zero reasoning |
|---|---|---|---|
| 0 | 138 | 147 | 0% |
| 2 | 47 | 0 | 67% |
| 6 | 304 | 0 | 67% |
| 20 | 996 | 0 | 100% |
At zero history the two are indistinguishable. Grok 4.6 thinks harder the deeper into a hunt it gets, reaching 996 median reasoning tokens at 20 prior turns. Grok 4.7 reasons zero on 100% of calls at the same depth. It stops once it has its own transcript to read, which is the entire premise of an agent. The task is fine and the LLM harness is fine. What breaks is what the model does with its own transcript.
Agentic AI security keeps hitting this wall because labs optimise for efficiency proxies that do not include investigative persistence, and persistence is the defensive capability. Grok 4.6 supported our earlier finding that coding benchmarks are the best publicly available proxy for cyber defense. Grok tapped into Cursor training data from real developer-agent interactions, exactly the kind of feedback that makes models better at working through complex technical environments. With 4.7 we expected the long-running training to compound that.
SpaceXAI's release notes say Grok 4.7 was trained with a longer reinforcement-learning run weighted toward problems that take many hours, and that it is better at verifying its own work and managing longer context. On SpaceXAI's own benchmarks that holds up. Terminal-Bench 4.0 goes from 20.3% to 38.0%, CursorBench 4.0 from 40.4% to 46.3%. It does not hold for cyber defense threat hunting, where the deeper into a hunt it goes, the less it thinks.
SpaceXAI also makes a direct cyber-defense claim. The release notes say Grok 4.7 "balances strong cyber defense capabilities with low refusal rates for legitimate use" and "shows the highest safety on HackerBench v0.3," allowing only 3.3% of risky dual-use prompts through; the model card adds that it shows "small cybersecurity-relevant capability gains over Grok 4.6." HackerBench measures whether the model refuses the wrong requests. Not whether it can find an intrusion. On the benchmark that measures the second thing, Grok 4.7 went backwards.
And this over-optimization is not one lab's mistake. Opus 4.8 spends 36% fewer investigation turns per hunt than Opus 4.6 — 33.5 against 52.4 mean turns, and coverage fell from 44.5% to 32.7%. Grok 4.7 kept its full turn budget and cut the thinking inside each turn. Coverage fell again.
| Efficiency dial the lab turned | Model pair | What it cut | Coverage |
|---|---|---|---|
| Fewer investigation turns per hunt | Opus 4.6 → Opus 4.8 | −36% turns | 44.5% → 32.7% |
| Less deliberation inside each turn | Grok 4.6 → Grok 4.7 | −107× reasoning tokens | 41.3% → 37.6% |
Two labs, two different efficiency dials, the same casualty. Each lab is optimising a proxy that scores well on its own evals. Threat hunting is where you find out what got traded away. Persistence and deliberation per step are the cyber defense capability.
Grok 4.7's regression has three practical consequences for AI threat hunting in a SOC, because "pick a better model" is not a strategy.
An LLM harness, not the model, is what catches silent agentic failure. A model that stops reasoning still returns well-formed JSON. One that never submits still runs successful queries. One 4.7 hunt did exactly that. It scored zero. A model that lies to you looks identical to a working one on a generic benchmark. So does one that bypasses your production database guardrails.
None of those are caught by looking at outputs. They are caught by a harness that scores against deterministic ground truth, tracks per-turn behaviour, and fails loudly. If your automation trusts whatever the model returns, you are shipping those failures straight into your SOC.
Grok 4.7 is sharper than Grok 4.6 on a single cyber question and worse across a long investigation, and both readings are true at once. Which one matters depends entirely on the job:
General benchmarks cannot tell you this. They do not run models end to end against adversarial evidence. A domain benchmark does. That is how the Grok 4.7 gap showed up at all.
The most useful model is the one that gets better inside your environment. Knowledge of historical investigation data. Which service accounts legitimately touch LSASS. Which nightly job looks like exfiltration every Tuesday. That is what separates a real analyst from a capable stranger.
This is why we build the AI SOC Agent around a Context Lake™. It is the accumulated, queryable memory of an environment (prior investigations, confirmed benign patterns, analyst decisions) that an agent reads before it starts hunting. A model that begins each hunt cold has to rediscover your environment every time. One that starts with your context does not, and it improves as your environment teaches it.
Model releases will keep going sideways on defensive work, as this one did. What you build around them does not go sideways with them.
Grok 4.7 loses coverage the longer a threat-hunting task runs: sharp on a short prompt, coasting across a long agentic defense run.
Is that your experience with Grok 4.7?
Q: Is Grok 4.7 good for cybersecurity? For defensive SOC work, no. It is a regression. Grok 4.7 scores 37.6% coverage on Simbian's Cyber Defense Benchmark, rank 5, against Grok 4.6's 41.3% at rank 3. It is sharper than Grok 4.6 on a single cold cyber question, so it works as a one-shot triage classifier. For a long agentic investigation it is the wrong choice. 96% of its turns spend zero reasoning tokens, and by 20 prior turns of history it reasons zero on every call.
Q: SpaceXAI says Grok 4.7 has strong cyber defense capabilities — is that true? Only for refusal behaviour. It is not a claim about defensive work. SpaceXAI's cyber-defense claim rests on HackerBench v0.3, its own benchmark for risky and malicious cyber tasks, where Grok 4.7 allows only 3.3% of risky dual-use prompts through. That measures whether the model refuses the wrong requests. It does not measure whether the model can find an intrusion. On Simbian's Cyber Defense Benchmark, Grok 4.7 scores 37.6% coverage against Grok 4.6's 41.3%. Low refusal rate and high defensive capability are different properties. The vendor measured one of them.
Q: Is Grok 4.7 cheaper than Grok 4.6? Only by about 12% per hunt, and not because it is cheaper per token. SpaceXAI serves Grok 4.7 at the same list price as Grok 4.6: $2 per million input tokens and $6 per million output. Grok 4.7 halves its output, which is what makes it twice as fast, but an agent re-reads a growing transcript every turn, so input is 87–93% of the bill. The shorter output accounts for roughly 6 points of saving, while the cache miss rate nearly doubling from 5.8% to 11.0% pushed cost back up.
Q: How is coverage measured? Coverage is the fraction of malicious events across a full, causal attack kill-chain that the agent both finds and submits, across 105 procedures of real Windows attack logs scored against deterministic ground truth. Full methodology is in the technical report.
Q: Is Grok 4.7 better than Grok 4.6 for cybersecurity? No. On Simbian's Cyber Defense Benchmark it scores 37.6% coverage against Grok 4.6's 41.3%, a 3.7-point regression that drops it two places, from rank 3 to rank 5. It is faster (3 minutes vs 7) and marginally cheaper ($1.4 vs $1.6 per hunt). But it finds fewer malicious events. For a single cold cyber question 4.7 is the sharper model. Across a long investigation Grok 4.6 is clearly better.
Q: What is the Cyber Defense Benchmark? It is Simbian's open benchmark for agentic threat hunting, published as arXiv 2604.19533. An LLM agent is dropped into roughly 100,000 real Windows security logs per campaign, at a realistic 95% noise ratio, with no alert to follow and no Q&A hints. From there it must plan, query security tools, and reconstruct a multi-stage attack. It covers 26 campaigns across 13 of 14 MITRE ATT&CK tactics and 105 attack procedures. Scoring is deterministic against ground truth rather than an LLM judge. The headline metric is coverage: the fraction of malicious events the model finds and submits.
Q: Which LLM is best for cybersecurity? On the current leaderboard, Claude Opus 5 leads at 45.1% coverage, followed by Opus 4.6 at 44.5%, Grok 4.6 at 41.3%, and GPT 5.6 Sol at 41.0%. GLM 5.2 is the best open-weight model at 35.0%. But no model out of 27 clears the 50% coverage threshold, so no raw LLM is good enough to run a SOC alone. On every LLM cybersecurity benchmark we run, the best LLM for cybersecurity is the one wrapped in a harness that catches its failures.
Q: Does Grok 4.7 hold context better than Grok 4.6? Not in a way that helps defensive work. SpaceXAI states that 4.7 was trained on longer-running problems, holds context better, and checks its own work more carefully. Our measurement is the opposite: conversation history is the exact variable that breaks it. Replaying real hunt transcripts, it matches 4.6 at zero prior turns (147 vs 138 median reasoning tokens). Then it collapses. At 20 prior turns it reasons zero tokens on 100% of calls, while Grok 4.6 rises to 996.
Q: Why do LLMs stop reasoning during long agentic runs? Because recent post-training rewards efficiency on proxies that do not include investigative persistence. Two labs show the same casualty through different dials. Opus 4.8 spends 36% fewer investigation turns per hunt than Opus 4.6, 33.5 against 52.4, and coverage fell from 44.5% to 32.7%. Grok 4.7 kept its full turn budget and cut the reasoning inside each turn, so 96% of its turns spend zero reasoning tokens. What both labs traded away was deliberation per step, sustained across a growing transcript. That is the defensive capability itself.
Q: What is agentic AI security? Agentic AI security has two meanings, and this post measures the second. It can mean securing AI agents themselves: prompt injection, tool-chain exposure, over-broad agent identity. It also means using AI agents to do security work: planning, querying tools, pivoting on findings, and reaching a verdict without executing a pre-written SOAR playbook. The measurement here shows why the agent loop is the hard part, not the model itself. A model that silently stops reasoning still returns well-formed output. So agentic AI security depends on a harness that scores against ground truth and fails loudly.
Q: Can any LLM pass the Cyber Defense Benchmark? No. Across 27 frontier models from Anthropic, OpenAI, SpaceXAI, Google, and open-weight providers, none has cleared the 50% coverage threshold. The leader sits at 45.1%. By contrast, frontier LLMs routinely score far higher on offensive-security benchmarks, because offense has a clean success signal (did the exploit land) and defense does not. Defense is the harder problem, and it is the one a SOC actually pays for.