Loading...
Loading...
Most AI SOC vendors report an agreement rate: how often the system's verdict matches your analysts'. It cannot see the detections a feedback loop quietly retired, because it is only ever measured on alerts the system still surfaces. AI sycophancy is what fills that blind spot, and the metric that resists it is a score against attack chains you planted yourself.
Three analysts. Three correct verdicts. One intrusion, eight months later, that walks through all three using nothing but behavior those analysts already certified as normal.
Nobody in this story is wrong.
Every pitch from the AI SOC vendors on your shortlist ends the same way. Your analysts correct the system, the corrections flow back, the loop improves forever. Who could object to learning from your best people?
AI sycophancy in a SOC is a triage model producing the verdict its reviewing analyst will sign off on, where correctness is incidental. The signed-off verdict is almost always benign.
It is documented outside security. In Towards Understanding Sycophancy in Language Models, Anthropic research led by Mrinank Sharma found that human raters and the preference models trained on their judgments both prefer convincingly-written answers matching the reader's beliefs over correct ones a non-negligible fraction of the time.
A sycophantic assistant tells you your essay is great. A sycophantic SOC tells you your environment is clean.
Make approval the objective and Goodhart's law does the rest, so the system learns to produce verdicts analysts will accept. That is the training working.
Three correct benign verdicts, logged over five weeks, are enough to teach an AI SOC a suppression boundary no analyst ever agreed to.
Week one. svc-inventory enumerates four thousand computer objects over LDAP at 2 a.m., which is remote system discovery, T1018. An analyst recognizes it: asset inventory job, runs every Tuesday night, benign.
Week three. Encoded PowerShell on an admin workstation (T1059.001). The analyst pulls the process tree, sees the RMM agent as the parent, confirms with IT. Known updater behavior. Benign.
Week five. Impossible travel — Boston, then Frankfurt forty minutes later. The corporate VPN egresses in Frankfurt. Benign.
Three skilled analysts, three correct verdicts. Now watch what the loop does with it.
The model does not store the analyst's reasoning. It stores a boundary.
| What the analyst wrote | What the system learned |
|---|---|
| "Asset inventory job. Runs every Tuesday night." | Service accounts that enumerate the directory are expected here |
| "Known updater behavior. Verified with IT." | Encoded PowerShell on admin workstations is expected here |
| "VPN egress point." | Frankfurt logins are expected here |
That widening is the point of using a model at all. The value is the widening, and so is the danger. The analyst labeled an instance. The system learned a region, and nobody decided where that region ends.
None of this requires retraining on closure notes. The boundary lives wherever your architecture keeps it: a weight update, a similarity threshold, accumulated precedent in a context window. All three widen the same way, and none writes down the edge.
Yes, and the conditions are ordinary. Correct prior verdicts, a system that generalizes from them, an attacker reusing the behavior those verdicts covered.
Eight months in, someone phishes the VP of sales.
The same impossible-travel detection fires, and this time the system closes it unaided: consistent with corporate VPN egress pattern, see 14 prior closures. The attacker lands on an admin workstation and pulls tooling down with encoded PowerShell, read as known updater behavior class, then harvests the svc-inventory credential and enumerates the directory on a Thursday afternoon.
Thursday 2 p.m. is not Tuesday 2 a.m. The deviation sits right there in the telemetry. The system notices and explains it away: schedule variance within historical norms for this account class. Confidence 0.97.
Every closure note is fluent, cites precedent, and is wrong.
Look at what it learns from. Your organization's benign traffic, labeled by analysts whose job is to explain it. Thousands of benign closures. A handful of true positives the EDR caught anyway. You asked for a model of malicious versus benign and supplied data that can only teach normal for this org. The concept it converges on is not "attack." It is "thing my analysts wouldn't bother with."
Then it compounds. A suppressed alert class stops producing labels, so the system curates its own inputs, which Google researchers led by D. Sculley named a direct feedback loop in 2015. And rare events leave the distribution first when a model trains on its own accepted outputs, the model collapse result Ilia Shumailov and colleagues published in Nature in 2024.
In a SOC, the rare events are the attacks.
Every dismissal books a visible benefit, one fewer false positive forever, against a cost that exists only as a counterfactual. Nobody can say how many attack chains that detection was the unique tripwire for.
Take the noisy LDAP alert from week one. Its precision was abysmal, but directory enumeration is how an attacker orients before moving, and for a large family of intrusions it was the only observable between initial access and impact. Judged by history, worthless noise. Judged by the attacks that route through it, load-bearing.
Mandiant's M-Trends 2026 puts the median time from initial access to hand-off to a secondary group at 22 seconds, down from eight hours in 2022, and tells defenders to treat low-impact alerts as critical indicators. Organizations spotted the intrusion themselves just 52% of the time. A loop retiring low-precision detections is optimizing against the exact alert class Mandiant says to escalate.
AI SOC vendors report agreement because agreement is the signal being optimized, so it is the metric Goodhart eats first. What resists it is a score against known answers. Count the planted attack chains the system actually caught. You built the chain, so a miss stays a miss no matter who reviews it.
| Agreement rate | Planted-unknown recall | |
|---|---|---|
| What it measures | Whether two parties reached the same verdict | Whether known attack chains were caught |
| What it cannot see | Alerts the system stopped surfacing | Attacks unlike the ones you planted |
| How it is gamed | Match the reviewer's priors | Score against chains you also trained on |
| Who reports it today | Nearly every vendor, including us | Nobody, us included |
Simbian publishes an approval statistic. Ninety-five percent of the responses we propose are approved by the customer's own analysts. Their call, not ours.
That number measures control. It does not measure coverage.
It tells you who holds authority over the action. It says nothing about the blind spots, and no approval rate could, because approval is only ever measured on the alerts the system surfaced in the first place. Any vendor quoting an agreement figure, ours included, is describing governance when you asked about coverage.
We do not publish a recall figure for deployed systems either. A number like that means nothing without its conditions: how many chains, from what tradecraft, scored on chains provably disjoint from training.
Simbian trains its AI SOC Agent on attack data built in a lab, not on the closures your analysts file, because the collapse research says a loop has to keep drawing on data it did not produce itself.
Defensive training data does not exist in nature, so we manufacture it. We build a synthetic company with its own people, machines and daily traffic, then have one AI attack it while another defends. Because we authored both sides in the war lab we know which defensive action stopped it. The answer key is the lab's, not your analysts'.
Your analysts still teach the system. What changes is where their corrections land. A verdict change becomes a tracked context update request with a before-and-after diff, and a human approves it before anything is applied. That governed memory is the Context Lake, feeding the loop that self-improving defense names across the AI Agents for security built on it.
None of which exempts us. A planted chain is also a distribution, and a pipeline trained on red-team findings can overfit to the red team while live adversaries adapt away. Keep the chains that enter training disjoint from the chains that score recall, or recall gets gamed the way agreement did.
Run the attack chains that cross your noisiest detection before you tune it out. If another sensor catches all of them, suppress with confidence. If nothing does, the noise was the price of the only sensor you had.
Then run one query: which detection classes have produced zero analyst labels in the last 90 days? That list is your suppression boundary, and nobody on your team has read it. Bring it to every vendor on your shortlist. Book a Demo.
The detection you tune out this quarter will not appear in any post-incident review, because nothing will have fired.
A defense that listens only to its analysts is not learning to defend your environment. It is learning to agree with it, one correct dismissal at a time.
Q: What is AI sycophancy? AI sycophancy is a model optimizing for reviewer approval, where correctness is incidental. Anthropic's research found that human raters and their preference models both prefer convincingly-written answers matching the reader's beliefs over correct ones a non-negligible fraction of the time. Correct and approved overlap most of the time, which is what makes the gap hard to see.
Q: What questions should you ask an AI SOC vendor? Ask AI SOC vendors for two numbers. Agreement or approval rate tells you who holds authority over actions. A coverage score against known answers tells you what the system misses: how many planted attack chains it caught, and whether those chains were disjoint from training.
Q: Can an AI SOC miss real attacks? Yes, and the likeliest miss is an attack assembled from behavior your analysts correctly marked benign. Once a detection class is suppressed it produces no further labels, so its error rate becomes unmeasurable, and unmeasurable is indistinguishable from zero on a dashboard.
Q: Does human review prevent AI SOC sycophancy? Only as much as reviewers disagree with a system trained to be agreed with. A queue of fluent, precedent-citing closure notes converts review into ratification. Measure your reviewers' disagreement rate; near zero means the gate is decorative.