
TL;DR
🔬 LLM Security Test: Three open-weight models beat Claude Opus 4.8; watch the full session on demand.
🛡️ Paperclip AI Flaws: One agent import runs code on the host, and the exploit is already scripted.
🤖 Self-Improving Defense: Simbian scores itself against your objective and proposes fixes you approve.
📊 ROI Calculator: Model three years of SOC, hunting, and pentest cost on your own numbers.
🛠️ CrowdStrike Falcon: Every Falcon detection investigated to a verdict, with containment behind your approval gates.
🧪 Automated Pentesting: Full reliance on automated pentesting fell from 29% to 9% in one year.
📰 Industry Buzz: Five September stories, from AI agents hacking retailers to a record Patch Tuesday.

The LLM Security Test: Three Open-Weight Models Beat Claude — Now On Demand
We scored 27 models on 100,000+ real attack events, with no alert or hint to start from. GLM 5.2, DeepSeek 4 Pro, and Kimi K3 each beat Claude Opus 4.8. None of the 27 reached the pass mark. Ambuj Kumar and Alankrit Chona walk through the scoreboard and the test to run on your own logs before you pick a model.

Paperclip AI Flaws: Importing One Agent Was Enough to Own the Host
Paperclip is an open-source control plane for teams of AI agents. Import a malicious agent, start it, and the attacker’s command runs on the server (CVE-2026-41679, CVSS 10.0, no pre-existing account needed). A second flaw (CVSS 9.6) reaches developer laptops through DNS rebinding. Rapid7 has packaged the server attack as a Metasploit module, so anyone can now aim it at every exposed instance, faster than a team triaging by hand can respond. Defense has to run at machine speed too, with people approving the actions that matter. Patch to v2026.416.0 or later.
Read the full article on The Hacker News.

Self-Improving Defense: Security That Gets Better at Beating Your Enemy, Every Run
You set the objective, say stopping data exfiltration or keeping the plant running. Simbian defends it across the tools you already run and scores its own work against what you wrote. Where it falls short, it proposes a fix to its own skills with the evidence, and nothing changes until you approve. On one customer-authored objective, the score went from 30% to 90%. Before it reaches your environment, it has already fought thousands of AI attacks in our war lab.
Explore Self-Improving Defense.

New: The Cybersecurity ROI Calculator
The Cybersecurity ROI Calculator models three years of SOC, threat hunting, and penetration testing cost, with and without Simbian. It takes ten inputs. Every default is visible, payback month included, and breach-cost savings are left out. If you’d rather work from real figures, our team will replace the defaults with readings from your stack and walk your sponsor through the result.

Integration Spotlight: Simbian + CrowdStrike Falcon
What Simbian does through the Falcon API:
- Investigates every Falcon detection to a verdict, tracing process trees and registry changes across the attack chain.
- Isolates compromised endpoints through Falcon network containment, behind the approval gates you set.
- Hunts for the same pattern across all endpoints with Falcon Event Search and Spotlight data.
- Pulls in SIEM, IAM, and threat intel context before an analyst opens the alert.
95% of the responses Simbian proposes are approved by your own analysts.
See the CrowdStrike integration.

Automated Penetration Testing Works. Full Autonomy Doesn’t.
In 2024, practitioners called automated pentesting a Nessus scan with a dashboard. Today agents handle reconnaissance, OWASP baseline coverage, exploit chaining, and retesting on their own. They still can’t judge business logic or authorization chains, or tell you what they missed. Cobalt’s 2026 Pulse Report found full reliance on automation fell from 29% to 9% in a year, with 47% preferring a hybrid. What holds teams back is accountability: no major framework accepts a fully automated test on its own, and DORA requires testers to carry professional indemnity insurance.

1. AI agents hacked online retailers for about $25 a target: Gambit Security traced one operator using three open-source AI frameworks to find, exploit, and skim online stores. From Sept 10 to 15, at least 27 companies were compromised, including a major US airline, and more than 600,000 card records came from two of them. At one retailer, the agent’s cleanup step dropped 180 database tables, backups included (Gambit’s interim findings). Source
2. Two Citrix NetScaler zero-days, three days to patch: CVE-2026-88771 and CVE-2026-88772 give an unauthenticated attacker code execution on NetScaler ADC and Gateway. Citrix confirmed exploitation, and CISA set a Sept 30 deadline for federal agencies. Shadowserver counts 23,000+ exposed instances. Citrix says its IoCs "might be of limited forensic value," so preserve evidence before you patch. Source
3. ShinyHunters breached the FBI’s jobs portal: The FBI declared a "cyber security incident" after staff names, addresses, job titles, and Social Security numbers were stolen. The attackers say they got in through an Oracle PeopleSoft server behind FBIJobs.gov. They aren’t asking for ransom, only a correction to a May FBI bulletin about the group. Source
4. A record 974 fixes in one Patch Tuesday: Microsoft’s September release includes two exploited zero-days, both Windows privilege escalations to System, and ZDI counts 20 fixes that could be wormable. Tenable’s Satnam Narang: "AI-assisted vulnerability discovery in 2026 is creating larger haystacks, but it isn’t finding more needles." Source
5. Attackers chained JFrog Artifactory flaws to plant backdoors: Wiz saw multiple threat actors chain CVE-2026-42018 and CVE-2026-42016 to turn an anonymous token into admin access. A third flaw, CVE-2026-82329, was exploited within days of its Aug 28 patch. Attackers stole cluster keys, added their own SSH keys, and installed malicious plugins. All three flaws are in CISA’s KEV catalog. Source
