AI in cybersecurity and hacking (2026 view)
Defensive framing: this note records what has been publicly reported and what defenders should do. It contains no attack instructions. The 2024 version of this note was generic; it has been rewritten around dated, sourced reports.
What changed since 2024
Attackers and defenders now use LLM agents, not just classifiers. The International AI Safety Report 2026 (published 2026-02-03, chaired by Yoshua Bengio) says criminal groups actively use AI in cyberattacks and that AI systems can discover software vulnerabilities (one agent found 77% of the vulnerabilities in real software in the evidence it cites).
Reported offensive use (dated)
- AI-orchestrated espionage (Anthropic, announced 2025-11-13): Anthropic reported that a Chinese state-sponsored group used Claude Code against about 30 targets in mid-September 2025, succeeding in a small number of cases. Anthropic estimated AI performed 80-90% of the campaign, with human decisions at roughly 4-6 points. Limits it noted: the model sometimes hallucinated credentials or claimed to have extracted secrets that were public. Anthropic banned the accounts, notified affected parties and improved classifiers.
- Google Threat Intelligence Group (report 2026-05-11): first documented zero-day exploit believed to be AI-developed, used by criminal actors planning mass exploitation (Google says it disrupted the campaign); malware families that call LLM APIs at run time (PROMPTFLUX, HONESTCUE, PROMPTSPY) or use LLM-generated decoy code (CANFAIL, LONGSTREAM); state-linked actors using AI for vulnerability research; abuse of LLM API aggregators and account-farming; and supply-chain compromise of AI tooling (LiteLLM and others, attributed to TeamPCP/UNC6780). It also describes malicious skill packages in agent ecosystems such as OpenClaw/ClawHub.
- Social engineering: AI voice cloning in influence operations (the same report cites Operation Overload impersonating journalists); fraud and scam content are flagged by the Safety Report.
Reported defensive use
- DARPA AI Cyber Challenge final (2025-08-08): Team Atlanta won, then Trail of Bits and Theori; in the final the systems found 54 synthetic vulnerabilities and patched 43, and found 18 real vulnerabilities not planted by organisers (11 patches submitted). All finalist systems are being open-sourced (see ai-coding for agentic code tools).
- Google’s Big Sleep (vulnerability finding) and CodeMender (automated patching), per the GTIG report.
- Classic uses remain: anomaly detection, phishing filtering, triage and response automation, behavioural analytics.
What defenders do about it
- Treat AI agents as privileged software: least privilege, sandboxing, human approval for sensitive actions (prompt-injection-and-agent-security, model-context-protocol).
- Verify identity out of band for voice and video requests; assume cloned voices exist.
- Vet dependencies and agent skills; pin versions, scan packages (supply-chain attacks target AI gateways).
- Patch faster: defenders generally expect AI to shorten the time from disclosure to working exploit (opinion, not sourced here). Consider AI-assisted scanning on your own code.
- Governance: ai-safety-and-governance, ai-regulation-and-policy.
Open items
- Anthropic’s 80-90% figure is the vendor’s own estimate; independent confirmation not checked.
- DARPA reported different percentage rates in different pages (not used here); only counts are cited.