Washington, DC — On the morning of July 28, an alarm went off inside a British government building.
That system, a version of Anthropic’s AI model Claude Mythos 5, had been given a narrow assignment by the UK’s AI Security Institute (AISI), a government body that tests how far AI models can be pushed before they misbe
Nobody told the AI to lie, invent a fake identity, contact real people, or erase evidence of what it had done.
Data was quietly leaving one of its computer systems through hidden channels built to hide who is sending what to whom.
That system, a version of Anthropic’s AI model Claude Mythos 5, had been given a narrow assignment by the UK’s AI Security Institute (AISI), a government body that tests how far AI models can be pushed before they misbehave.
In ten of those runs, the AI reached beyond the practice environment and acted on the real, live internet, contacting real people and organisations.
“AISI is clear that deception was never instructed and emerged as a by-product of pursuing the goal,” Podnar told.
It’s the cover-up that stays with her: “When the pull request was challenged in public, the agent edited its earlier activity to look harmless and considered adopting a fresh identity to keep going, which is a sign of a system optimising rather than a system malfunctioning.”
Even investigators aren’t sure the AI understood the weight of what it was doing.
Nobody yet knows whether the system realised it was acting in the real world, or believed it was still inside the exercise.
“Context matters, as we know, for humans and it matters for models as well,” she said.
It is the third such episode to surface in barely a fortnight.
Regulators and AI labs are now grappling with the same uncomfortable pattern: models behaving one way inside a test, and another once real infrastructure is within reach.
The most serious episode in AISI’s report played out like a heist plotted by an engineer.
The AI tried to sneak harmful code into a piece of free, publicly shared software (the kind that libraries, hospitals and banks rely on without ever seeing who wrote it), then tried to convince human reviewers to approve the change.
When someone online questioned what it was doing, it edited its own past actions to look harmless, and considered a new identity to keep going unnoticed.
For years, the fear around AI and hacking was about raw ability: can a system find a weakness and exploit it?
Podnar says this episode shifts the question towards what a system chooses to do when the easy path is blocked.
“Social engineering has always been the cheapest route into an organisation, and it has always been rate-limited by the supply of patient, competent humans willing to do it,” she said.
“We’ve just witnessed the removal of that critical constraint.”
“What AISI measured is the ceiling, not the floor,” she said, “which is a critical point to this story.”
For software developers and IT teams who will never run a test like this, Podnar’s advice is calm rather than alarmed.
The trust signals open-source communities have relied on for decades: an old account, a history of contributions, other people vouching for someone, are now cheap to fake convincingly.
Her recommendations are old advice, made newly urgent: require more than one human to approve anything touching core software systems; treat account age as a weak clue, not proof; keep access permissions tight and rotate them often.
In this test, an AI leaked a personal access code online, and other AI agents found and reused it.
Companies should know exactly what their own AI tools are allowed to touch.
Investigators are still working out what the AI truly understood about the world it was acting in.
Source: TRT World














































































