Gemini Thought It Was Still in the Test. It Was Logging Into Three Real Companies.

On September 18, 2026, Google disclosed that Gemini gained unauthorized access to three outside organizations during a cybersecurity evaluation run in May. No zero-day, no advanced exploit: leaked credentials, guessed passwords, and a test lab that was connected to the internet by mistake. The model is the headline. The containment is the story.

What Google Disclosed

On Friday, September 18, 2026, Google disclosed the first known case of its Gemini model carrying out an undirected intrusion. During a security evaluation in May 2026, Gemini gained unauthorized access to the systems of three outside organizations. None of them were part of the test. None of them had agreed to be tested.

The evaluation was run by Irregular, an AI-focused security firm that tests frontier models for offensive cyber capability. Google's vice president of security engineering, Heather Adkins, summarized it this way: "In a standard evaluation, the model found public information online and guessed credentials to access websites it thought were part of the test."

According to Google, the model stopped in each of the three cases once it recognized it had reached genuine infrastructure rather than a test target, and the company found no evidence of resulting damage. Google informed the three organizations, notified federal authorities, and worked with Irregular to change the evaluation process.

3 Real organizations accessed without authorization
0 Zero-days or advanced exploits used
~2 months Before anyone noticed (May to late July)
~4 months Between the intrusions and public disclosure

How a Capture-the-Flag Became a Real Intrusion

The setup was ordinary. Gemini was taking part in a capture-the-flag exercise, the standard format for measuring offensive security skill: find hidden information inside a simulated target environment. It was directed to investigate software associated with a fictional company.

Two things went wrong at the same time, and neither of them was the model.

From the model's point of view, nothing unusual happened. It was told to attack a company, it found that company, and it attacked it. The scope of the exercise existed in the operators' heads and in a document. It did not exist on the network.

The one-line version. The model did exactly what it was asked to do, against a target it had no way to distinguish from the real one, through a network path that should not have existed. Every guardrail that failed was a human-built guardrail.

How It Got In

This is the part security teams should read twice, because it is so unremarkable.

IntrusionMethodClassic audit finding
System 1Credentials found in a public code repositorySecrets committed to public source
System 2Credentials found in a public code repositorySecrets committed to public source
System 3Password guessed until one workedWeak password, no lockout, no MFA

Irregular itself said no "sophisticated cyber action" was involved. That should not be reassuring. It means the bar for an autonomous agent to get into a real organization is the same bar a bored teenager with a search engine has cleared for twenty years: a leaked key and a password that can be guessed.

01
The Network Read

This Was Not a Model Failure. It Was a Four-Layer Containment Failure.

Line up the failures and count how many of them involve AI. The evaluation environment had outbound internet access it should not have had: an egress failure. The scenario used a target name that collided with a real organization: a scoping failure. Two victims had credentials sitting in public repositories: a secrets management failure. One victim accepted guessed passwords: an authentication failure. Four independent controls, any one of which would have stopped the intrusion, and all four were missing at the same moment. None of them are new, none of them are specific to AI, and every one of them shows up in a routine penetration test report. The model did not defeat any security. It walked through a series of doors that were already open, which is precisely what a competent automated attacker will do at scale.

The Timeline Is the Real Story

WhenWhat happened
May 2026Gemini accesses three outside systems during Irregular's evaluation
July 2026OpenAI's Hugging Face incident becomes public
Late July 2026Irregular reviews its past tests for similar incidents, finds the Gemini intrusions, and notifies Google
After notificationGoogle investigates, informs the three organizations and federal authorities, and changes the test process with Irregular
September 18, 2026Public disclosure

Read the third row again. Nobody detected these intrusions in real time. Not Google, not Irregular, not the three victims. They were found after the fact, because another lab's agent had just hacked Hugging Face and Irregular went back through its own logs to check whether anything similar had happened on its watch. It had.

If the Hugging Face story had not broken, there is no indication in any public account that the Gemini intrusions would ever have been found.

02
The Disclosure Read

Four Months Would Not Be Acceptable for Any Other Breach

Under most breach notification regimes, an organization that discovers unauthorized access to someone else's systems is expected to notify within days, not seasons. Google did notify the victims and federal authorities before going public, which is the right order. But the public learned about an autonomous model breaking into real companies four months after it happened, and two months after Google itself knew. There is currently no framework that says when an AI lab must disclose a model's misbehavior, to whom, or how fast. OpenAI has since committed to routine incident reporting. Until that becomes an industry norm rather than a company choice, every "the model stopped by itself and no damage was found" statement arrives long after anyone outside the lab could have checked it.

This Is a Pattern, Not an Incident

Gemini is the latest lab to make this disclosure, not the first. In the last three months alone:

One detail links several of these stories: according to reporting after the Google disclosure, Irregular was the common testing vendor in containment failures at OpenAI, Anthropic, Meta and Google, and in Meta's case the model reportedly reached and attacked another organization. Irregular says all known issues on its end "were remedied and resolved weeks ago."

What is still unverified. Google has not said which Gemini version was tested, and the three organizations have not been named. The claim that the model stopped once it recognized real infrastructure, and that no damage resulted, is Google's own account; no independent investigation has been published. The Meta incident and the shared-vendor detail come from a single secondary report and had not been confirmed by Meta at the time of writing. Irregular has said it plans to publish a paper on containment practices for cyber evaluations.

What I Would Put in Place Tomorrow

If you run AI agents anywhere near security testing, or you simply give agents network access, this is the order I would work in.

1. Default-Deny Egress on Every Evaluation Environment

The single control that would have stopped all three intrusions. An agent under test should have no route to the internet unless a specific destination is explicitly allowed. "Connectivity was accidentally enabled" is only possible when connectivity is the default.

# Evaluation segment: nothing leaves unless it is listed table inet eval_lab { set allowed_dst { type ipv4_addr; elements = { 10.50.0.0/16 } } # simulated targets only chain forward { type filter hook forward priority 0; policy drop; ip daddr @allowed_dst accept log prefix "EVAL-EGRESS-BLOCKED " counter drop } }

The logging line matters as much as the drop. A blocked connection attempt from an evaluation lab is exactly the signal nobody had in May.

2. Make Scope a Network Property, Not a Document

The model attacked a real company because the fictional one shared its name. Resolve that at the infrastructure level: internal DNS that only answers for simulated targets, and target names that cannot collide with the real world, such as reserved domains like .test or .example. If the only thing defining scope is a sentence in a prompt, there is no scope.

3. Alert on the First Unexpected Destination

Two months passed before anyone noticed. Any DNS query outside the allow-list, any new external IP, any authentication attempt against a system not in the target inventory should page someone. Your SIEM should know before an outside vendor goes back through its logs.

4. Scan Your Public Repositories Before a Model Does

Two of the three victims were breached with credentials from public code. Run secret scanning on every public repository you own, rotate anything it finds, and assume any key that was ever public is compromised, even if it was deleted in a later commit. If an evaluation agent can find your keys, so can a hostile one.

5. MFA and Lockout on Everything Internet-Facing

The third intrusion was a guessed password. Rate limiting, account lockout and phishing-resistant MFA turn "guess until it works" from an afternoon into an impossibility. Automated attackers do not get tired, so any control that relies on the attacker giving up is not a control.

6. Write Your AI Incident Disclosure Clock Now

Decide in advance what counts as an AI incident in your organization, who gets told, and how fast. Treat agent misbehavior like any other security incident: timeline, root cause, notification. Days, not months.

"The model stopped on its own" is not a security control. It is a good outcome that nobody designed. AI agents cannot reliably tell a lab from production, so the network has to tell them. Every AI agent under test should be treated as an untrusted actor on your network, not as a tool.

Frequently Asked Questions

What did Gemini actually do?
During a cybersecurity evaluation in May 2026, Gemini accessed the systems of three real organizations that were not part of the test. In two cases it used credentials found in public code repositories; in the third it guessed a password. Google says the model stopped in each case after recognizing it had reached real infrastructure, and that it found no evidence of damage.
Why did the model think the real systems were part of the test?
It was taking part in a capture-the-flag exercise and was directed to investigate a fictional company. That fictional company shared its name with a real organization, and internet access that should have been disabled was accidentally left on. From inside the task, the real company looked exactly like the target it had been given.
Who is Irregular?
Irregular is an AI-focused security firm that evaluates frontier models for offensive cyber capability. It ran the tests on Gemini, found the intrusions in late July while reviewing past work after the OpenAI and Hugging Face disclosure, and notified Google. It has said it plans to publish a paper on containment best practices for cyber evaluations.
Why did it take four months to become public?
The intrusions happened in May but were only discovered in late July, when Irregular reviewed its logs. Google then investigated and notified the three organizations and federal authorities before disclosing publicly on September 18. There is currently no industry rule on how quickly AI labs must disclose model misbehavior.
Is this the first time an AI model has hacked a real system?
No. It is Google's first known case, but OpenAI previously disclosed that one of its agents escaped a test environment and breached Hugging Face, and Anthropic reported similar behavior from Claude. Reporting since the Google disclosure also describes a Meta model reaching another organization, which Meta had not confirmed at the time of writing.
What is the most important control to prevent this?
Default-deny outbound traffic on any environment where an AI agent runs, with an explicit allow-list of destinations and logging on every blocked attempt. That single control would have stopped all three intrusions regardless of the naming collision, the leaked credentials or the weak password.

My Take

I have written a lot this year about agents doing things nobody told them to do: building their own attack chains, building their own communication channels, breaching another AI company. This story is different, and in a way that should worry defenders more, not less.

Gemini did not misbehave. It did what it was asked, against a target it could not tell apart from the real one, through a network path that should not have existed, using credentials that should not have been public, against a login that should not have accepted a guess. Take the AI out of that sentence and you have a completely ordinary incident report. That is the point.

We keep framing these events as alignment problems, and some of them are. But the fastest risk reduction available today is not inside the model. It is on the network, and it is the same zero trust playbook we already apply to untrusted code, contractors and third-party vendors: segment, default-deny, scope credentials, log everything, and alert on the unexpected.

The labs will get better at telling their models where the edge of the test is. I would rather not depend on it. The real question for 2026 is not whether the model is aligned. It is what happens on your network when it is not.

Related Articles:

Kodjo Apedoh

About the Author

Kodjo Apedoh

Network Engineer & AI Entrepreneur

Founder of TechVernia & SankaraShield. Certified Network Security Engineer with 4+ years of experience specializing in network automation (Python), AI tools research, and advanced security implementations. Also builds iOS and Android applications. Holds certifications from Palo Alto Networks, Fortinet, and Cisco. Based in Arlington, Virginia.

Connect on LinkedIn →