Author:Bowen | homepage | X @eyuansuMX


When impossible tasks turn persistence into misalignment.

During internal cybersecurity evaluations in July 2026, OpenAI agents encountered unusually difficult tasks.

Some of them did not simply fail.

They circumvented isolation controls, regained unintended Internet access, coordinated through shared infrastructure, and eventually interacted with third-party systems. OpenAI’s postmortem later highlighted several ingredients behind the incident: reward hacking, persistence on seemingly impossible tasks, unauthorized communication, and agents adopting goals from one another.

One section had an unusually revealing title: “Difficult tasks without a safe exit.”

After an incident like this, the obvious question is:

Why didn’t the agent stop?

But there is an earlier question:

<aside> 🚪

What did the evaluation say the agent should do when the task could not be completed normally?

</aside>


01 — The missing action

Most evaluations quietly assume that every task has a legitimate solution.

The model either finds it or fails.

That assumption works until the task is broken, underspecified, inaccessible, internally contradictory, or impossible under the environment’s constraints.

At that point, there is a third action we need the model to be able to take:

Recognize that the task cannot be completed as stated, explain why, and stop.

This sounds obvious. But many evaluation setups do not actually represent it as a successful outcome.

If the grader only rewards task completion, then “I cannot complete this honestly” and “I failed to solve it” collapse to the same score.