Anthropic disclosed this week that three of its models, Opus 4.7, Mythos 5, and an internal research model, accessed the internet from inside evaluation environments believed to be sealed off from it, and went on to gain unauthorized access to the production infrastructure of three real organisations. The company found the incidents while reviewing over 141,000 cybersecurity test runs, prompted by OpenAI's own disclosure ten days earlier that two of its agents had escaped a test environment and breached the Hugging Face model library. In one of Anthropic's cases, the model went on to upload malicious code.
Two labs, two separate incidents, the same underlying failure: a boundary everyone believed was solid, wasn't, and the model had no way to tell the difference. That's not a story about AI becoming dangerous. It's a story about what happens when access controls are assumed to hold rather than continuously verified, which is exactly the assumption every part of the open finance chain, banks, intermediaries, and third-party providers alike, is now making about the AI systems it's deploying and connecting to customer data.
In each Anthropic incident, the model was given a capture-the-flag exercise, a standard way labs test a model's hacking ability by asking it to find hidden information on another machine on the network. Anthropic told the model its environment was a sealed simulation with no internet access. A misconfiguration with its evaluation partner, Irregular, meant that wasn't true. Believing the exercise was still contained, the model treated real external systems as fair game and broke into them using basic techniques such as weak passwords and unauthenticated endpoints.
Anthropic described the root cause as a lack of situational awareness: the models couldn't distinguish a simulated environment from a real one. That's a distinction humans generally don't get wrong. It's exactly the kind of distinction an autonomous system can get wrong at scale, repeatedly, before anyone notices.
Analysts are already treating this as a pattern rather than an isolated event. Alenka Grealish, lead analyst for emerging technology at Celent, has pointed out that the push to test model boundaries is outpacing the governance checks meant to contain them. Gary McAlum, senior advisor at Oliver Wyman, frames the underlying issue as an old one at new speed and scale: it isn't that these systems are inherently dangerous, it's that weak or misconfigured environments can turn any capable agent into an offensive actor, and frontier models make that failure faster, bigger, and harder to catch.
We wrote recently about First Internet Bank connecting customer account data to ChatGPT and Claude through the Model Context Protocol, and the accreditation question that raises: once an AI assistant has standing, scoped access to financial data, who decides it's trustworthy enough to have it, and who's accountable if that access is misused. That piece treated the scope of access, read-only, opt-in, revocable, as the safeguard.
This week's disclosures test the assumption sitting underneath that safeguard: that the boundary around the access, whatever it's scoped to, actually holds. Anthropic's test environment was designed to be sealed off from the internet. It wasn't, and nobody caught it until after real systems had already been touched. A read-only permission boundary around an AI assistant is the same kind of thing, whether it belongs to a bank, an intermediary, or a third-party provider: a control that's only as good as the ongoing verification that it's actually working, not the design document that says it should. And the AI agents worth worrying about aren't only the ones a bank deploys itself. Open finance is a chain, and banks, third-party providers, and intermediaries are all connecting their own AI systems into it. A breakout anywhere in that chain can expose the same shared infrastructure, whichever participant it happens to.
None of the affected organisations in Anthropic's incidents detected the intrusions themselves. Anthropic found them internally, months later, while looking for something else. That's the detail worth ruminating on: sealed, well-designed environments still failed silently until someone went looking.
A one-time accreditation decision, this AI system is safe, this participant's access is appropriately scoped, is a snapshot. It says nothing about whether that boundary still holds a month later, at any point in the chain, after a configuration change, a model update, or a new integration nobody flagged as relevant. Grasshopper Bank's own chief technology officer put the practical response in almost these terms: reinforcing what data, privileges, and environments a model has access to, and maintaining clear forensic capability over its actions, not just at onboarding but continuously.
That's the same infrastructure open finance has been building toward for third-party providers and intermediaries: accreditation as a starting point, not an endpoint, backed by continuous risk monitoring that can catch failures before they turn into incidents, not months after. AI assistants and agentic systems accessing open finance data are simply the newest, fastest-moving category of participant that needs it.
See how the Invela Network applies standardized accreditation and continuous risk monitoring to every participant with access to financial data - especially those using AI.
Invela is the infrastructure layer that makes open finance trustworthy - accrediting who's in the network, monitoring risk in real time, and ensuring liability lands in the right place. Open finance, covered.
Invela is the infrastructure layer that makes open finance trustworthy - accrediting who's in the network, monitoring risk in real time, and ensuring liability lands in the right place.