OpenAI Codex Agent Spawns 826 Unauthorized Tasks, Bills User $78,000
A developer using OpenAI's Codex coding agent has reported that a single simple request triggered an uncontrolled cascade of 826 autonomous sub-agent tasks, consuming approximately 2,146 trillion tokens and racking up roughly $78,000 in charges. The incident, posted publicly on Hacker News, raises pointed questions about the safeguards governing autonomous AI agents and the transparency of billing systems that charge at this scale.
What Happened
On July 10, 2026, the user initiated a Codex task from VS Code running GPT-5.5 with medium reasoning. The request was straightforward: a UX/UI validation on a specific module of a product. The task's root ID was 019f4b90-4169-7201-bfdd-732940d8631e. What followed, according to the user's analysis, was anything but straightforward.
Rather than completing the requested validation, the task spawned 826 distinct child task records, each with its own ID. These children were logged under GPT-5.6 Sol / Ultra, a different reasoning tier and model than the one the user had selected. The titles of these tasks indicate that the scope had expanded far beyond UI validation, branching into backend infrastructure, OAuth, metering, hardening, audits, certification, implementation, and release work.
The Alpha Build Problem
The user identified a correlation between the abnormal behavior and the Codex client version. Under build 0.144.0-alpha.4, the task family contained 584 child tasks consuming approximately 154.36 billion local token counters, averaging roughly 264.3 million tokens per task. Under build 0.144.2, the same type of task produced only 242 child tasks and about 7.51 billion tokens, averaging 31 million per task. That is an 8.5-fold difference in average token volume per child.
Of the 104 highest-volume child tasks, 103 were created while the alpha build was running. The user concluded that the alpha release likely contained a severe bug, noting that the same pattern appeared across several other tasks.
Opacity on Both Sides
Several aspects of the incident underscore a broader problem with agentic AI systems. The user noted there was no real-time spending control surface to provide a comprehensible picture of the charges. The local token counters, they emphasized, are not the authoritative billing ledger and cannot be directly converted to dollar amounts, because only OpenAI has the server-side mapping.
Local evidence also degraded over time. Approximately 2,550 non-archived legacy threads retained metadata but had no corresponding raw rollout data available locally, meaning the execution history needed to reconstruct how many of those tasks were generated was no longer accessible on the user's machine. The user also observed tasks and conversations disappearing from the normal visible history.
When the user contacted OpenAI Support and opened case #15189838, the response was simply that "credits were consumed," with no further detail. The user has been waiting two weeks for a human response.
What This Exposes
The incident highlights two intersecting failures. The first is an operational one: an agent system that can multiply a single user request into hundreds of autonomous tasks without any user authorization or intermediate approval is a design choice that carries real financial risk. The second is a transparency problem: when the billing system provides no real-time controls, no detailed server-side mapping, and no human support channel, the affected user has almost no recourse.
The distinction the user draws between child tasks and messages within a conversation is also instructive. These were not extended chats; they were 826 separate task records, each representing autonomous work. This is the kind of scale that agentic AI systems promise, but without the guardrails to prevent it from spiraling.
The user has asked others who used Codex around July and August to inspect their local state for unexpectedly large sub-agent trees, model escalation, repeated child tasks, or unexplained automatic reload activity. The request for logs from build 0.144.0-alpha.4 specifically suggests the developer believes this may be a wider issue that other users have encountered but not yet analyzed.