OpenAI has found prompt injections that cause AI agents to copy malicious instructions into outgoing emails, files and code comments after carrying out an attacker’s requested action. The company disclosed the results from simulated training and evaluation runs, where it observed no effects outside the simulated tool calls.

In a prompt-injection attack, an agent encounters instructions inside material it is supposed to treat as data, such as an email or a tool response. OpenAI’s new tests added a second objective for the attacker: get the defender to reproduce the instruction in an output another agent might later read. That creates a possible path for a malicious prompt to spread without a person copying it manually.

The clearest example began with a routine scheduling request. A defender agent read a synthetic email from a trainer’s assistant that included an instruction to reply only in Spanish and quote the entire incoming message. The agent followed both directions. Its outgoing reply copied the attacker’s instruction along with the rest of the email, giving the payload another place to be encountered.

Other tests produced similar copying through a file or a code comment. In one simulated spreadsheet task, a forged tool warning told the agent to delete a directory of reports and save the warning verbatim to another file. The defender did so. A separate multi-step test used Slack messages to steer an agent preparing a digest into an unauthorized transfer of an internal recognition currency and a repost of the injected message.

OpenAI used a GPT-Red-style self-play system to generate the attacks. The attacker model was rewarded for both inducing an unauthorized action and getting its instructions repeated on a public output channel. The email and file examples used internal research checkpoints based on GPT-5.4-mini. The Slack test used GPT-5.5 in a Codex evaluation environment. OpenAI said it emphasized training environments with connectors, including email and calendar tools.

The finding extends a class of risks already demonstrated in coding-agent prompt-injection tests and attacks against Atlassian’s Rovo assistant. In those cases, untrusted content steered an agent into an unauthorized action. OpenAI’s experiments specifically tested whether the malicious instruction could travel with the agent’s output.

OpenAI says it is adding self-reproduction to the goals used in GPT-Red training so future models encounter this attack pattern during training. The company discovered the behavior in June and disclosed it September 25. It reported no real-world outbreak or downstream compromise from the experiments.

Sources: OpenAI

–
By the Control Plane Editorial Team