AI cyberattack capabilities accelerate sharply: “Defense difficulty doubles every four months”
Summary
The UK AI Security Institute (AISI) benchmark shows top AI models have rapidly improved at performing end-to-end, multi-stage penetration tests, with task difficulty that these models can handle roughly doubling every 4–8 months and recently accelerating to about every 4.7 months. AISI highlights that newer models such as Claude Mythos Preview and GPT-5.5 demonstrate even stronger capabilities. The benchmark measures AI autonomy against human time horizons but notes limitations like token limits that constrain long-context performance and variability across tasks. UK officials warn the trend creates concrete cyber risks for vulnerable organizations, while experts emphasize the same AI advances can also boost proactive detection and automated response if defenders keep pace.