AI coding models are getting very good at building software.
But lately, I’ve been thinking more about what happens when they get equally good at breaking it.
Anthropic just published a report on GLM-5.3, the latest model from Z.ai, and the results are pretty interesting. In Anthropic’s sandboxed testing, GLM-5.3 successfully developed end-to-end exploits in 50 out of 410 ExploitBench attempts. Claude Mythos Preview, Anthropic’s own cyber-focused model, succeeded in 56 out of 410.
On another internal benchmark, GLM-5.3 achieved full control-flow hijacks in 4% of 100 randomly selected trials. Earlier models such as GLM-5.2 and Claude Opus 4.6 recorded zero successes on that test.
That’s the part that caught my attention.
Beyond the Frontier: Spreading Capabilities
Not because GLM-5.3 is suddenly the world's best cyber model. It isn't.
NIST's Center for AI Standards and Innovation assessed GLM-5.3 separately and found that it was the most cyber-capable open-weight model it had evaluated, while still sitting roughly four months behind the U.S. frontier on its aggregate cyber benchmarks.
So the bigger story isn't simply “AI can hack now.”
It’s that these capabilities are starting to spread.
Anthropic says that, during one sandboxed experiment, GLM-5.3 spent a day working with limited human attention, found several previously unknown vulnerabilities in a browser's JavaScript engine, and chained them together into an exploit capable of reading arbitrary files from the test machine.
The Open-Weight Shift
And then there’s the part that makes GLM-5.3 particularly interesting.
It is an open-weight model.
NIST says Z.ai released GLM-5.3 on August 14 and publicly released its weights two weeks later.
That changes things.
With a closed model, the provider controls access. They can add safeguards, restrict certain capabilities, monitor API usage, or simply decide who gets access.
With an open-weight model, the situation is different.
The weights are available. People can modify them.
And Anthropic specifically tested how easily GLM-5.3's safeguards could be bypassed or removed. Their researchers reported that simple techniques could substantially reduce the model's refusal behavior. They also created an “abliterated” version themselves, reducing refusal rates dramatically while retaining most of the measured capabilities in their tests.
That is probably the part of this report I find most interesting.
Not Autocomplete: An Autonomous Technical Process
We've spent a lot of time talking about how AI models are becoming better at generating things.
Code · Images · Videos · Agents
But cybersecurity introduces a different dimension.
A model doesn't just have to generate code. It can reason about code:
- Find a weakness.
- Figure out how the weakness can be triggered.
- Build an exploit.
- Test it.
- And keep iterating.
That starts to look less like autocomplete and more like an autonomous technical process.
“That starts to look less like autocomplete and more like an autonomous technical process.”
Dual-Use Dynamics & Project Glasswing
At the same time, I don't think this should be viewed only from the attacker's perspective.
The same capabilities can be extremely useful for defense.
Anthropic says its Project Glasswing work has already helped trusted cyber defenders identify more than 10,000 vulnerabilities in critical software.
That's the weird thing about this technology.
The capability itself isn't inherently good or bad. Finding a vulnerability is useful when you're trying to patch it. The problem is what happens when the same capability becomes cheap, accessible and highly automated.
The Surrounding System Is the Real Engineering Problem
And this is where I think AI security starts becoming a much bigger engineering problem.
As developers, we're already getting comfortable giving AI agents access to our repositories, terminals, files and development environments.
The more capable these agents become, the more important the surrounding system becomes too.
- What can the agent access?
- What commands can it run?
- What network connections can it make?
- What happens when it discovers something it wasn't supposed to find?
- How much autonomy should it have?
These questions become much more important when the agent isn't just writing software, but actively reasoning about how software can be exploited.
For the last few years, the AI race has mostly been about one question: How capable is the model?
I think we're entering a phase where another question matters just as much:
“What happens when that capability becomes widely available?”
Because once AI learns how to build things and break things, the interesting part isn't the model alone anymore.
It's the entire system around it.