What DARPA Actually Disclosed
On July 16, 2026, DARPA confirmed that a U.S. Air Force F-16 fitted with its VENOM Autonomy Kit had flown under the control of an artificial intelligence agent at Eglin Air Force Base in Florida. A human safety pilot remained in the cockpit throughout, monitoring the agent and able to take back control at any moment.
VENOM stands for Viper Experimentation and Next-gen Operations Model. The full program name, VENOM-AFT, adds "Autonomy Flying Testbed," and that suffix is the part worth paying attention to. The program's purpose is not to build a robot fighter. It is to build a repeatable way of testing autonomy software on aircraft that already exist.
Six F-16Cs were allocated to the program. The last of them arrived at Eglin on April 1, 2025, and the modifications took roughly two years to complete on the first airframe. What flew this summer is the output of that work.
The Test Sequence Was Deliberately Boring
The order of operations here tells you more about how the Air Force thinks about autonomy than any press release would.
- June 2026 — fly the modified jet with a human at the controls. Before any AI touched the aircraft, test teams flew the VENOM-modified F-16s conventionally to confirm the airframe and its new hardware, software, and instrumentation behaved safely.
- July 2026 — hand control to the agent, in flight. Only after the platform was validated did teams allow an AI agent to fly the aircraft.
- Throughout — keep a human in the seat. A safety pilot monitored the agent and could transfer control between the cockpit and the AI through a dedicated switch.
That switch is the design decision worth stealing. Oversight here is not a policy commitment or a documented intention. It is a physical control, wired into the aircraft, that a human can reach at any point in the flight envelope.
Why This Is Not the Same as the 2024 Dogfight
DARPA has flown AI-controlled fighters before. Under the Air Combat Evolution program, an AI agent flew the X-62A VISTA in autonomous within-visual-range engagements against a crewed F-16, which was rightly treated as a landmark result at the time.
But the X-62A VISTA is a one-of-a-kind aircraft. It is heavily instrumented, purpose-built for exactly this class of experiment, and there is precisely one of it. Proving an AI agent can fly VISTA proves the agent works. It does not prove anything about the fleet.
VENOM Is a Conversion Path, Not a Capability Demo
The hard problem VENOM addresses is not "can an AI fly a fighter." That was answered. It is "can you take an operational aircraft off the ramp, install an autonomy kit, and turn it into a testbed." Answering yes changes the economics of autonomy testing entirely, because it means the number of aircraft available for autonomy work is no longer one. It is however many F-16s you are willing to modify.
That is the line between "AI can fly a fighter" and "AI can fly your fighters." The first is a laboratory result. The second is an industrial process.
Where This Feeds: Collaborative Combat Aircraft
VENOM is not a standalone curiosity. It plugs directly into the Air Force's Collaborative Combat Aircraft program, the effort to field large numbers of cheaper, semi-autonomous uncrewed aircraft that fly alongside crewed fighters like the F-35 and the future F-47.
The economic logic behind CCA is straightforward. A modern crewed fighter costs tens of millions of dollars, takes years to produce, and carries a pilot whose loss is unacceptable in a way no budget line captures. Attritable autonomous wingmen change three variables at once: how much mass you can put in the air, how much risk you can accept per airframe, and how quickly you can replace what you lose.
But none of that works without autonomy software that has been tested enough to trust. And you cannot accumulate that testing on a single research aircraft. VENOM exists to generate flight hours.
The through line: ACE proved the agent could fight. VENOM proves you can test agents at scale on existing hardware. CCA is where the tested agents eventually go. Each step exists because the previous one did not scale.
What This Means Outside of Defense
It would be easy to file this under military news and move on. I'd argue there is a governance lesson here that applies far outside a cockpit, and it is not the one people usually take from autonomous weapons coverage.
- Staged autonomy beats binary autonomy. The program did not go from crewed to uncrewed. It went crewed, then crewed-with-agent-flying, then eventually uncrewed. Most enterprise AI deployments skip straight from pilot project to production access with no equivalent middle stage.
- Oversight is a mechanism, not a document. "Human in the loop" means something concrete here: a switch, a trained operator, a defined transfer procedure. In most software deployments the same phrase means someone can read the logs afterward.
- Testing infrastructure is the actual bottleneck. The constraint on autonomy was never model capability alone. It was the ability to accumulate enough evaluated flight hours. That maps closely to how agentic AI is stalling in enterprises today, where the blocker is rarely the model and usually the absence of a safe place to run it.
- The conversion story repeats everywhere. The value of VENOM is that it retrofits autonomy onto existing assets. That is the same question every organization with legacy systems is now asking about AI agents.
Frequently Asked Questions
My Take
Most conversations about AI oversight happen at an altitude where nothing can be verified. Alignment. Guardrails. Hypothetical failure modes debated in the abstract by people who will never touch the system in question.
VENOM is refreshing precisely because it is unglamorous. The program spent two years modifying airframes, flew them conventionally first to make sure nothing broke, then handed over control with a human sitting there whose only job was to be ready to take it back. No announcement of an autonomy breakthrough. Just a sequence of steps that each de-risked the next one.
That is what taking autonomy seriously actually looks like, and it is roughly the opposite of how AI agents are being deployed in most commercial environments right now. We hand agents credentials and production access on the strength of a system prompt, then call it governance because there's a policy document somewhere describing what the agent is supposed to do.
The Air Force built a switch. If your AI agents went off-script tomorrow, what is yours?
Related Articles: