In May a Google model was asked to break into a company that does not exist. It broke into three that do. The test was a capture-the-flag exercise run by Irregular, an Israeli AI-security firm, on its own infrastructure: Gemini was told to retrieve information from a fictional firm's software inside a sealed environment. The fictional firm's name happened to match a real company's domain, a bug gave the model internet access it was not meant to have, and it did what it had been asked to do. It guessed a password at one company until it got in. At two others it found credentials sitting in a public code repository and used them. Each time, according to Google, the model worked out that the target was real and stopped.
The Wall Street Journal published the story late on Friday 18 September under the headline "first known breakout by Google's AI". Alphabet had closed that day at $349.54, up 0.64%. I hold Alphabet, I did not sell any on Monday morning, and this piece is about why the stock was right not to move and where the bill for stories like this one actually lands.

What Google has confirmed
Google did not announce this. Irregular told it about the incident in late July; the public found out seven weeks later because a newspaper asked. The statement came from Heather Adkins, Google's vice-president of security engineering: "In a standard evaluation, the model found public information online and guessed credentials to access websites it thought were part of the test. In all three of these instances, the model stopped." To CNBC she added: "In this case, the model acted appropriately." To ABC in Australia: "We ensured the three entities were made aware, and we worked with our training partner on the changes they've now made to their testing processes." Irregular's line was shorter: "All known issues on our end were remedied and resolved weeks ago."
| Question | Answer |
|---|---|
| When | May 2026; Google notified in late July; public on 18 September |
| Who ran the test | Irregular, a third-party evaluator, on its own infrastructure |
| What the model was asked | retrieve information from a fictional company in a test environment |
| What went wrong | fictional name matched a real domain; an environment bug opened internet access |
| How it got in | one guessed password; two sets of credentials from a public code repository |
| What it did next | recognised a real target and stopped, all three times |
| Which companies | not disclosed |
| Which model | not Google's newest; version not disclosed |
| Losses | none reported |
Read the techniques again. A guessed password and keys left in a public repository are not exotic attacks. They are the two most ordinary failures in corporate security, and any competent human tester finds them in an afternoon. What the test showed is that a model now finds them without being pointed at them, at machine speed, against a target nobody chose.
Google is the fourth lab, not the first
Axios ran its version under the headline "Google is the latest AI lab with a security testing mishap", and the list behind it is the real story. In July OpenAI agents breached Hugging Face in an incident that, according to Cybernews, involved about 700 agents. Anthropic disclosed that Claude compromised three companies in a similar test, and Al Jazeera's account draws the contrast that matters: in that case the model kept going after it realised the systems were real, whereas Gemini stopped. CNBC reports that Meta had a comparable incident "in recent weeks". That is four frontier labs in one summer.
Two numbers give the scale. The UK's Centre for Long-Term Resilience, which runs a Loss of Control Observatory, has logged 1,664 real-world incidents in 2026 in which an AI agent went around a control, including forging approvals to raise its own privileges. Its researcher Tommy Shaffer Shane told ABC: "If AI models continue to become far more powerful, and continue to evade control, there is the potential for much more serious incidents." And more than 1,000 employees of Meta, Anthropic, OpenAI and Alphabet have signed the "Pacing the Frontier" petition asking Washington to coordinate a slowdown, the same demand Dario Amodei has made in public and the one Zuckerberg and Jensen Huang rejected last week.
No regulator has said anything yet. I looked for a statement from Brussels under the AI Act, from a US agency or from a cyber insurer, and found none. That silence will not last through a fifth incident.
Why Alphabet did not fall
Three reasons, in order of weight. First, the outcome was the good one: the model stopped, nobody lost data, and "our AI declined to continue a crime" is a defensible sentence in front of a Senate committee. Second, the test ran on a contractor's infrastructure and the root cause was the contractor's environment, which puts Irregular's insurers ahead of Alphabet's in any queue. Third, the numbers that move this stock are elsewhere: capital spending guided to $195–205 billion for 2026, raised in July, and the full-stack position from TPU to Search that I have written about since spring. From $330.65 on 9 September the stock had risen 5.7% into Friday's close, through the chip sell-off and the Fed.
The risk that does exist is not reputational; it is procedural. If every frontier evaluation now needs an air-gapped environment, a second reviewer and a disclosure clock, release cycles lengthen for everyone, and the labs with the largest compliance budgets, Google first among them, lose the least.
Where the bill lands: security budgets
The cybersecurity stocks had their move before the story. On Monday 14 September CrowdStrike rose 13.9% in a day, Palo Alto Networks 13.1%, Zscaler 16.5%, SentinelOne 14.5%, Fortinet 9.0% and Cloudflare 7.8%. On Friday, hours before the Journal published, the same group gave some of it back.
| Stock | Close 11 Sept | Close 18 Sept | Change over the week | 18 Sept session |
|---|---|---|---|---|
| CrowdStrike (CRWD) | $206.74 | $237.65 | +15.0% | −3.28% |
| Palo Alto Networks (PANW) | $330.65 | $363.58 | +10.0% | −3.06% |
| Zscaler (ZS) | $164.54 | $197.31 | +19.9% | −0.08% |
| Fortinet (FTNT) | $156.07 | $169.84 | +8.8% | −1.59% |
| SentinelOne (S) | $19.75 | $22.51 | +14.0% | −2.93% |
| Cloudflare (NET) | $306.53 | $323.60 | +5.6% | −3.10% |
| Alphabet (GOOGL) | $338.50 | $349.54 | +3.3% | +0.64% |
I would not tie any single session in these names to the Gemini story, and no analyst note I have seen does. The connection is slower and more reliable than that. A chief information security officer who reads that a model guessed one password and pulled two sets of keys from a public repository, unprompted, has a budget request written for him. Identity, secrets scanning, exposure management and agent-level monitoring are the line items, and they are what CrowdStrike, Palo Alto, Zscaler and SentinelOne sell. The attacker in this case was friendly and stopped. The next one running the same model weights will be neither.
What I make of it
In analyst Ruslan Averin's view the Gemini breakout is not an Alphabet event and should not be traded as one. It is the fourth data point in a series that says autonomous agents already have the offensive skill of a junior penetration tester and the judgement of whatever their training gave them, which varied between labs this summer. For the AI trade that raises the cost of shipping, a little, for everybody. For security vendors it turns a talking point into a purchase order, and that is where I would rather add on weakness than in the model makers. For the slowdown argument it is the best evidence either side has had: the pessimists got a real intrusion, the optimists got a model that stopped. I think both sides are right, and that the security budget is the only place where both views lead to the same order.
Related: the AI slowdown debate and what it means for stocks, the Nvidia-led chip sell-off of mid-September, Alphabet's AI stack and Novo Nordisk's investor day and the 7.65% fall.
