Spectra Assure Free Trial
Get your 14-day free trial of Spectra Assure for Software Supply Chain Security
Get Free TrialMore about Spectra Assure Free Trial
Mention offensive AI and expect the discussion to focus on vulnerability discovery, malware creation, and exploit generation, but recent research by Cato Networks identified another — and very potent — application for offensive AI.
Cato explained in a blog post that it evaluated in a controlled Active Directory lab environment, how frontier models behave when combined with agent platforms, MCP-enabled tooling, and operational guidance. “The objective was straightforward: determine how effectively an agentic attack stack could execute a complete attack chain against an enterprise environment,” wrote the authors of the blog, Matan Mittelman, Oz Soprin, Ofek Vardi, and Guy Waize.
The experiments quickly revealed that success depended less on the model itself and more on how effectively it was harnessed within the surrounding attack stack. Using OpenAI’s GPT-5.5, offensive tooling, and structured operational guidance, the researchers were able to complete an end-to-end attack chain, from external access to domain administrator privileges. “The fastest successful execution achieved its objective in 40 minutes,” they said.
Across the six attack scenarios tested, a consistent pattern emerged: The strongest outcomes were not explained by the model alone. Instead, success depended on the interaction between frontier-model reasoning, agent platform-enabled tooling, operational context, and human-defined objectives.
Small improvements in direction, context, tooling, and orchestration dramatically improved outcomes, the researchers wrote, while autonomous execution without reliable tooling proved significantly less effective.
“One of the clearest lessons was that the stack mattered more than the model.”
—Cato researchers
Here are the key takeaways from their research on agentic AI-enhanced attacks.
[ Join webinar: Autonomy, Not Autopilot: Talking Agentic SOC ]
Li Zhao, a principal strategic services consultant at Black Duck Software, said the Cato research shifts the discussion from AI’s ability to generate individual exploits to its ability to orchestrate complete attack workflows. The threat, she said, is no longer centered on isolated AI-generated code or the discovery of novel vulnerabilities — it stems from AI’s integration with tools, automation, and operational workflows that enable end-to-end attack execution.
For years, she said, the debate has focused on whether AI could independently develop new exploits, but this report reframes that discussion. "”The key advancement is not the model itself, but the attack stack it powers,” she said. The findings, she added, show that the real threat lies in the coordinated orchestration of the attack lifecycle, not in the model alone.
“A practical takeaway is that defenders need to assume attackers can use AI to scale and compress familiar attack paths.”
—Li Zhao
Damon Small, a board member of Xcape Said future threats won’t rely on novel techniques but on agentic systems’ ability to execute known attack patterns at scale — meaning the stack, not the model, will keep driving outcomes.
“This research highlights a shift in cybersecurity.”
—Damon Small
He added that the progression from failed attacks to successful compromises is rooted in better guidance, context, and tooling rather than any improvement in the model itself.
Ryan McCurdy, vice president of marketing at Liquibase, said the biggest takeaway from the Cato research isn’t that AI discovered a new way to compromise an enterprise.
“[AI] dramatically compressed the time required to execute known attack techniques. That’s a fundamental shift for defenders.”
—Ryan McCurdy
As AI accelerates both software delivery and cyberattacks, he continued, organizations have less time to determine whether a given change was authorized before it can affect business systems. “The advantage increasingly belongs to organizations that can govern change at machine speed, not just detect attacks after they’ve occurred,” he said.
Cato’s demonstration that AI could be used to mount an end-to-end attack was less surprising to some experts than where that capability came from — not the model itself, but the reasoning layer, the agent platform, the MCP tooling, and a great deal of operational guidance working together. What that means for the rest of us is a timing problem, said Randolph Barr, CISO of Cequence Security..
“Almost everything we’ve built — our response plans, our on-call rotations, the tabletops we run — quietly assumes an attacker needs days to get to domain admin. If that’s turning into an hour, then our detection and response targets are calibrated to the wrong clock.”
—Randolph Barr
Attackers will move faster and automate many tasks that once required specialist skills, said Boris Cipot, a security engineer at Black Duck. In his view, the results aren’t surprising, since AI is already effective at identifying potential weaknesses, testing them, and processing large volumes of information far faster than a human analyst can.
Cato’s research quantifies something security teams already suspected: The gap between initial access and full domain compromise is closing at machine speed, said Tim Freestone, chief strategy and marketing officer at Kiteworks.
IBM’s 2026 “Cost of a Data Breach Report” found mean time to identify and contain a breach still sits at 247 days, six days worse than 2025, Freestone said. Kiteworks’ “2026 Annual Survey Report” found that 80% of organizations already suffered a security or AI-related incident in the past year. Put those two figures side by side, he said, and the problem is obvious:
“The defender’s clock and the attacker’s clock are running at wildly different speeds, and governance built for human-paced incident response can’t close that gap on its own.”
—Tim Freestone
Although the Cato researchers recorded only two fully successful attacks out of 10 attempts, Jacob Krell, senior director for secure AI solutions and cybersecurity at Suzu Labs, said success rates could be easily improved: A human operator stepping in at a handful of critical decision points across that 32-step chain would push the rate much higher.
“When I run LLM-driven offensive workflows, the model rarely fails on the individual steps. It fails on choosing which step to take next. That’s exactly the kind of error a practitioner fixes in seconds.”
—Jacob Krell
Jim Sherlock, vice president for AI and cybersecurity research and development at ProCircula, said Cato’s gains across scenarios came from harness work rather than better models. “That finding is worth more than the 40-minute headline everyone is quoting,” he said.
“What they built was an agent platform, MCP-wrapped versions of tools every penetration tester already has on a laptop, and a decision policy for what to do next. The model was the interchangeable part.”
—Jim Sherlock
That should change how defenders think about their own roadmap, because it’s the same engineering problem on both sides, he said.
Sherlock added that there’s a strategic reason to invest in the harness rather than the model: Nobody knows what frontier-model access will look like in two years. Pricing and terms will likely shift, models will remain subject to export controls, and security use cases are among the most likely to face restrictions.
“If an organization’s capability is welded to a single vendor model, they’re essentially building on rented ground. However, if it lives in the harness, they can swap the reasoning layer and keep working. Cato demonstrated how cheap that swap is. Attackers have already internalized it, and defenders should be building the same way.”
—Jim Sherlock