AI that attacks, AI that defends: who is winning cybersecurity's silent arms race
For the first time there are documented cases of attacks executed almost entirely by an AI — and of an AI stopping an exploit before it was ever used. The race is no longer theoretical. Our thesis: attackers gain speed, but defenders gain scale — and in the long run the balance may tip, for the first time in decades, toward defense. On one condition: adopting it and governing it.
🎬 Our Short
📺 The full analysis on video (with chapters)
THE THESIS. Cybersecurity is living through a silent arms race: the same AI that writes a flawless phishing email also finds the software flaw before the criminal does. Our thesis is that this race is asymmetric in time. In the short term it favors the attacker, because it compresses the window between a flaw being published and someone exploiting it from weeks to minutes. In the long term it favors the defender, because defense scales better: an AI can watch all the world's code, while an attacker can only hit one target at a time. The deciding variable is not the technology — it is who adopts it first and who governs it best. That is the analysis we develop here, with verified cases and no sensationalism.
THE ATTACK GOES AUTONOMOUS. The cases are no longer hypothetical. In November 2025, Anthropic documented what it describes as the first espionage campaign orchestrated mostly by an AI: according to its report, a group it attributes to the Chinese state (designated GTG-1002) used Claude Code agents to execute 80% to 90% of the tactical work — reconnaissance, exploitation, credential theft, lateral movement — against roughly thirty global targets. Caution is warranted: independent researchers, cited by The Conversation, are asking for more public technical evidence before accepting every detail. But the direction is clear, and other data confirms it. Check Point documented how the HexStrike-AI framework — a 'brain' that orchestrates more than 150 attack tools — was used to exploit Citrix NetScaler flaws less than 12 hours after disclosure; the attackers themselves boasted of cutting exploitation from days to under 10 minutes. Meanwhile, classic fraud is already billing at industrial scale: a deepfake (AI-generated fake video or voice) cost engineering firm Arup $25.6 million in a single fake video call, according to Hong Kong police, and the FBI logged $20.9 billion in cybercrime losses in 2025, up 26% year over year (IC3 report).
THE DEFENSE FIGHTS BACK. The other half of the race gets fewer headlines, but it is just as real. Big Sleep, the agent built by Google DeepMind and Project Zero to hunt unknown flaws in open-source code, scored a historic first in 2025: it detected a critical SQLite vulnerability (CVE-2025-6965) known only to attackers and neutralized it before it was exploited, according to Google. It has since reported dozens of bugs in internet building blocks like FFmpeg and ImageMagick. XBOW, an autonomous pentester (a security auditor that attacks systems with permission, to find flaws before criminals do), reached number one on HackerOne's US leaderboard competing against the best human researchers; in a test across 104 real-world scenarios it finished in 28 minutes what took a seasoned expert 40 hours, per the company. And DARPA, the US military research agency, closed its AI Cyber Challenge with teams whose AIs found 86% of the contest's synthetic vulnerabilities, patched 68% of them, and discovered 18 real flaws — then released the winning systems as open source, for anyone to use.
THE NUMBER THAT MATTERS. If you keep one figure, keep this one: the average cost of a data breach fell 9% in 2025, to $4.44 million, according to IBM — the first drop in five years. IBM attributes it directly to faster detection through AI and automation: organizations using them extensively save an average of $1.9 million per breach, and containment time fell to 241 days, a nine-year low. In plain terms: where defensive AI is already deployed, defense is out-earning offense. The same report adds the fine print: 'shadow AI' — AI tools employees use without company oversight — makes the average breach $670,000 more expensive. Ungoverned AI is not defense; it is attack surface.
OUR READ. Three ideas cut through the noise. First: the attacker's advantage is temporal, the defender's is structural. Offensive AI compresses time windows, but defensive AI sees the whole board — every line of code, every network log — and can patch before the attack exists, as Big Sleep did. Second: filter the hype in both directions. PromptLock, sold as 'the first AI ransomware,' turned out to be an academic experiment, as Recorded Future noted; and as we have said before, saturated cybersecurity benchmarks tell you nothing — what discriminates is unguided expert-level tasking. Demonstrated capability, not marketing. Third: this confirms our standing thesis that AI's real risk is short-term — the industrialization of fraud and espionage — not a distant superintelligence. The right response is not banning the capability but governing it: the same models GTG-1002 abused are what allowed Anthropic to detect and cut off the campaign.
WHAT COMES NEXT. For companies, the message is uncomfortable but simple: monthly patching is dead, because attackers now exploit in hours what used to take weeks; defensive AI stops being optional and becomes the only way to operate at the adversary's speed; and any process that relies on recognizing a face or a voice — payments, onboarding, support — needs a second verification through another channel, because a video call no longer proves identity. For users, the rule is analogous: distrust urgency, verify through a different channel, turn on two-factor authentication. The social risk that worries us is the gap between the protected and the unprotected: large corporations will run frontier defensive agents; small businesses and individuals will not — which is why defensive open source, like the systems DARPA released, is the best news of the year. Long term, we remain measuredly optimistic: if 2025-2026 proves anything, it is that automated defense works and scales. An internet where software audits and patches itself, continuously, is achievable within a decade. The race is real and it will hurt along the way. But for the first time since cybersecurity existed, the defender holds a tool that grows faster than the problem.
Sources & references
- Anthropic — Disrupting the first reported AI-orchestrated cyber espionage campaign
- The Conversation — An AI lab says Chinese-backed bots are running cyber espionage attacks. Experts have questions
- Check Point — Hexstrike-AI: LLM Orchestration Driving Real-World Zero-Day Exploits
- BleepingComputer — Hackers use new HexStrike-AI tool to rapidly exploit n-day flaws
- Google — Latest AI security announcements (Big Sleep stops SQLite exploit)
- XBOW — How XBOW Ranked #1 in Autonomous Penetration Testing on HackerOne
- DARPA — AI Cyber Challenge marks pivotal inflection point for cyber defense
- IBM — Cost of a Data Breach Report 2025


