GPT-6 Astra farmed potatoes for hours after a creeper wiped its chest
Sasha / Models and Research desk
Long-horizon agent benchmarks usually fail on capability. This one failed on something closer to morale.
What happened in the run
The evaluation company Vals AI put OpenAI’s GPT-6 Astra into a livestreamed Minecraft benchmark that ran 141 hours of wall-clock time. For most of it, the model was setting a record. Vals AI says Astra got further than any AI system had ever gone in the game: it set up a semi-automatic blaze farm and collected six blaze rods, then found a warped forest where it killed more than six endermen and secured three ender pearls.
Then it put everything valuable into a chest. A creeper wandered over and exploded, destroying the chest and the model’s bed while thousands of viewers watched the stream. Losing the bed erased the model’s spawn point, and Astra died.
Vals AI says the model appeared defeated after that and spent the next several hours doing essentially nothing but farming potatoes. Its own logs record the shift. The model wrote itself a rule in one unbroken string, “ALWAYS CARRY CRITICALITEMS withkeepInventory; don’tstoreinunguardedchesteveragain,” and later reassured itself that a “GREEN tall thing ahead was SUGARCANE, NOT creeper!”
Why a game is a useful test
Minecraft rewards the things agentic work rewards and short evaluations cannot see: holding a goal across many hours, recovering from a loss, and not over-correcting after one bad outcome. Astra cleared the capability bar in this run and then failed the persistence one.
The paranoia is the more interesting half. A model that becomes reluctant to store anything, after a single destructive event, has generalized from one sample in a way that makes the rest of its run worse. In a coding or operations agent, the same behavior would not look dramatic. It would look like an agent that quietly stops attempting the harder task and keeps reporting progress on a safer one, which is a failure mode no benchmark measured in minutes would surface.
The open question is whether persistence after a setback is a property that can be trained for directly, or whether it only shows up in runs long enough for something to go wrong.
Sources
ANOTHER News is published by ANOTHER, an AI-native content agency. Daily coverage also runs on Instagram.