How AI Cyber Attacks Work
Part of Self-Improving Defense: AI Cyber Defense Against AI Attacks — read the full guide.
What is an AI-powered cyberattack?
An AI-powered cyberattack, often just called an AI cyber attack, is any attack where AI gets used to plan, build or run some stage of it, for example the reconnaissance, writing a phishing lure or cloning a voice, developing the exploit, or moving through the network afterwards. It's not the same thing as an attack on an AI system (prompt injection against a chatbot is the usual example), which security researchers tend to call adversarial AI. People mix the two up a lot and they need pretty different defenses.
AI changes two things about attacks and the order matters here. The first is volume, one attacker with a model can probe a lot of targets all day long, so both of the defender's queues (alerts and patches) grow along with it. CrowdStrike's 2026 Global Threat Report, published February 24, 2026, found operations by AI-enabled adversaries went up 89% year over year. IBM's 2026 Cost of a Data Breach Report from July 29, 2026 found one in four malicious breaches was AI-enabled, up 56% on the year before, at an average cost of $6 million.
The second thing is the nature of the attack. Once it's cheap to build a new attack, somebody can build one just for you and use it once, so the specific attack stops repeating. A lot of the time that means there is no signature to match and no playbook for it, and maybe not even a patch yet. Techniques still recur (which is why MITRE ATT&CK hasn't stopped being useful) but the exact artifacts defenders used to lean on often only show up once.
How are hackers using AI in cyberattacks today?
Mostly to do what attackers have always done, only faster and on more targets, and a few things that used to need a skilled person. In one documented 2025 campaign AI did 80 to 90% of the work. The documented cases so far mostly fall into reconnaissance, social engineering, malware and exploitation.
Two cases are worth knowing about. In November 2025 Anthropic said it disrupted what it called the first reported AI-orchestrated cyber espionage campaign, AI handled 80 to 90% of that campaign against roughly thirty targets and humans only stepped in at a handful of decision points. Then in 2026 Sysdig wrote up an attacker who used an AI agent to get from a public CVE to an internal database in four pivots. Roughly, the main uses map to the defender's side like this:
| How attackers use AI | Documented example | What changes for the defender |
|---|---|---|
| Reconnaissance at scale | Anthropic's 2025 case, where AI ran the recon across roughly thirty targets | More probing, which means more alerts that look routine |
| Tailored phishing, voice cloning, deepfake video | Personalized lures and cloned voices turn up in fraud reported to the FBI | Judging an email by its writing mostly stops working |
| Self-modifying malware | GTIG's PROMPTFLUX, experimental malware that asks Gemini to rewrite it | File signatures go stale fast |
| Exploit development | UK AISI found Claude Mythos Preview solved 73% of expert-level capture-the-flag tasks (April 2026) | Less time between a disclosure and the first exploit |
| Running the operation | AI doing most of the steps in an espionage campaign | Attacks move at machine speed and across a lot of hosts |
Not every attacker uses AI for everything. Plenty of what happens in the real world is still commodity tooling and stolen credentials. But the trend only goes one way and right now its making attacks cheaper faster than it makes defense cheaper.
Are autonomous AI cyberattacks real yet?
Partly. AI-orchestrated attacks have been documented, fully end-to-end attacks with no human at all are still not common. Anthropic's November 2025 report had AI running most of an espionage campaign with people at a few decision points. Sysdig's 2026 write-up had an AI agent taking an intrusion from a CVE all the way to a database.
The UK NCSC gave a useful counterweight in its May 2025 assessment of AI's impact on the cyber threat through 2027: "The development of fully automated, end-to-end advanced cyber attacks is unlikely to 2027. Skilled cyber actors will need to remain in the loop." Against well-defended targets that still seems about right. In April 2026 the UK AI Security Institute said Claude Mythos Preview finished its 32-step simulated corporate network attack in 3 of 10 attempts, and in the same breath pointed out there were no active defenders or defensive tooling in that range.
So the fully autonomous attacker is arriving slowly. Whether a human still presses a button somewhere matters less than people think. The thing that already changes defense is the per-target attack, built for you and used once.
How can AI be used in phishing attacks?
AI makes phishing more personal, more convincing and a lot cheaper to run at scale. Attackers use language models to write lures in the target's own context (their role, their vendors, some project they posted about on LinkedIn), clone voices for phone-based social engineering, and make deepfake video to pass as an executive.
The FBI's 2025 Internet Crime Report came out April 6, 2026 with an AI section for the first time: 22,364 complaints, losses of nearly $893 million. Microsoft's Digital Defense Report 2026 (October 1, 2026) found 93% of voice phishing attacks kept the victim on the line long enough to start the social engineering.
Cost changed too. In one IBM X-Force experiment, a model wrote a convincing phishing email from five prompts in about five minutes. The human team needed roughly 16 hours for the same thing.
That's a big reason 'spot the typo' training doesn't do much anymore. An AI-written email usually has no typos, no odd phrasing and a believable reason to click, so judging the email itself keeps getting harder. What still works is looking at what happened after the email arrived. Did anyone click. Did a credential get used from somewhere new. Did a process run on the endpoint, did the account start touching things it never touched before. The question that matters is "did this phish lead to a compromise", and your EDR and identity logs can answer it however good the writing was.
Can AI create malware that evades detection?
Partly. AI malware makes evasion cheaper, but the core behaviors of malware usually stay detectable. Polymorphic malware (code that changes its own form to get around file signatures) was around long before generative AI, what AI changes is that making a fresh variant for every target now costs almost nothing.
The Google Threat Intelligence Group (GTIG) documented PROMPTFLUX in November 2025, experimental malware that calls Gemini to modify itself and avoid detection. Going the other way, Arctic Wolf Labs looked at more than 22,000 AI-assisted malware samples in March 2026 and found their core execution, persistence and command-and-control behaviors were still detectable.
Taken together the lesson is pretty simple. A detection that depends on what the file looks like won't last long against AI-generated variants, while one that depends on what the malware does, how it persists and what it talks to and touches, holds up much better since the malware still has to do those things to get anywhere.
Why can't signature-based detection stop AI-generated attacks?
Because a signature is a description of something somebody already saw, and an AI-generated attack built for one target usually hasn't been seen by anyone. A signature matches a known file hash or string or pattern. If the attacker spins up a fresh variant for every victim there's nothing to match. GTIG's PROMPTFLUX calls Gemini to rewrite itself, so its hash doesn't stay the same long enough to sign.
The other two things a bigger team reaches for hit the same wall. No playbook, because nobody has responded to this exact attack before. Maybe no patch either, if the flaw was only found this morning. That's why adding analysts helps less than it used to, the old shortcuts don't apply and more hands won't bring them back.
Reasoning about method still works. A playbook encodes a decision (if you see this, do that) and a new attack breaks that decision since it doesn't look like 'this'. An investigation encodes method: how to tell what an attacker is after from partial evidence, which pivot is worth the next twenty minutes, when there's enough to call it. Attackers change what they build, they haven't changed what an investigation is. Method transfers, decisions don't, and that's what a self-improving defense gets trained on. For detection it means writing rules on behavior, like a MITRE ATT&CK technique, which a fresh variant can't change.
What does an AI that can find zero-days mean for defenders?
Mostly that the gap between a flaw existing and that flaw being exploited keeps getting shorter, so defenders need options that don't wait on a patch. Anthropic announced Claude Mythos Preview in April 2026, and the UK AI Security Institute reported it succeeded 73% of the time on expert-level capture-the-flag tasks and was the first model to finish AISI's 32-step corporate network attack range. AISI was careful to add that the range had no active defenders. So the result says more about weakly defended networks than about hardened ones.
Simbian's own mid-2026 analysis found the window from a vulnerability being disclosed to it being exploited dropped from roughly 30 days to roughly 5 over the previous six months, and for most enterprises a patch cycle measured in weeks just doesn't fit inside that.
Without a patch, a few things still help:
- Close the path: if the vulnerable system can't be patched today put a compensating control in front of it, such as an access-control change or some segmentation, so it's unreachable or at least worth less to an attacker.
- Hunt for prior use: check whether the technique was already used against you before anyone knew about the flaw.
- Watch for the technique: detect the behavior an exploit of that class tends to produce, so a new variant can get caught even when there's no signature for it yet.
A pentest that proves whether the flaw is actually reachable in your environment helps too. It tells you which of the many new CVEs deserve the first hour, because a CVE isn't really an exploit until someone shows it can be used against you.
How fast do you need to respond to an AI-speed attack?
Faster than a human-paced queue usually allows. CrowdStrike's 2026 Global Threat Report put the average eCrime breakout time, meaning initial access to lateral movement, at 29 minutes, and the fastest one they saw was 27 seconds.
That's the attacker's clock. The defender's clock usually starts much later, when someone actually picks up the alert, and if it's sitting in a long queue that can be hours after it fired. Breakout in 29 minutes against a two hour wait for an analyst means the attacker has already moved by the time anybody opens the case.
Simbian investigates and responds to every threat within minutes, on average under 4 minutes from detection to response. That doesn't mean every action runs without a person. Every action starts gated and you decide which action types are fine to run on their own, closing a false positive for example, and which ones wait for approval, like isolating a production server. Either way the investigation itself runs at machine speed, so a gated action is already investigated and just waits on a decision.
How do you defend against AI-powered cyberattacks?
Start with the basics, AI-powered attacks still go after the same weaknesses. Phishing-resistant MFA, patching what's exposed on time, least-privilege access, decent logging and segmentation all still matter. After its Mythos Preview evaluation the UK AI Security Institute gave the same advice: regular security updates, strong access controls, secure configuration and comprehensive logging.
What changes is how AI cyber defense and the security team work on top of those controls:
- Investigate every alert: AI-built attacks are often made to look unremarkable on purpose, so they tend to end up in the low-severity tail that severity-based triage skips over.
- Respond in minutes: CrowdStrike put the 2026 average eCrime breakout time at 29 minutes, the investigation needs to start the moment the alert lands.
- Detect behavior: write detections around what attackers actually do. Signatures age out fast against generated variants.
- Close the path, not just the ticket: once a real threat is handled, tighten the detection and change whatever access let it happen in the first place.
- Test continuously: keep proving which exposures are actually reachable, all year and not only when the auditors are due.
- Hunt on hypotheses: when companies in your sector get breached, go look in your own data for the same activity.
- Keep humans in control: gate the actions, keep an audit record and have a person approve every change made to the defense itself.
The common thread is a defense that gets a little better from each case it handles.
How can you tell if an attack was AI-generated?
Often you can't, and it usually matters less than people think. A phishing email or a piece of malware rarely has anything in it that reliably proves a model wrote it, and the tools that claim to detect AI-written text generally aren't reliable enough to base a security decision on.
Deepfake audio and video are a partial exception since there is specialised detection for them, although people asked to spot a deepfake by eye or by ear generally do only a little better than chance.
There are a few signals that point toward automation, though none of them proves it by itself. Speed, for one, with lots of steps across lots of hosts in very little time. Breadth is another, probing every exposed service at the same time. Per-target variation can be a tell too, when each victim gets a slightly different lure or payload.
What actually matters is the behavior. A credential used from an unusual place, a process injecting into another one, data getting staged for exfiltration, all of those need the same investigation whether a person or a model planned them. Defending against the behavior covers you either way, and working out whether AI was involved mostly helps the incident report.

.png&w=3840&q=75)