An Anonymous Model Claims Frontier-Level Coding Performance at $1.50 Per Task

Terminal-Bench 4.0 is one of the harder benchmarks in use right now. Hosted by Stanford, the Harbor framework, and the Laude Institute, it consists of 66 tasks spanning software engineering, machine learning, system administration, and security. As of early September 2026, GPT-6 Astra leads with a 58.2% resolution rate at a cost of roughly $3,300 per full evaluation run. Claude Fable 5.1 sits at 57.9% for about $6,200. These are expensive numbers for a benchmark that most teams cannot afford to run repeatedly.

On September 16, a model called Union Alpha appeared on OpenRouter and Cloudflare AI Gateway with no developer attribution and a claim that caught attention: roughly 51% on Terminal-Bench 4.0 at approximately $1.50 per task. That figure comes from a developer report, not an official leaderboard entry, and should be treated as a single data point rather than a verified score. But if it holds up, the cost difference is stark. Union Alpha would be achieving near-frontier coding performance at less than one-tenth of a percent of what the leading models cost per evaluation run.

What Is Known About the Model

Union Alpha is multimodal, accepting text and image inputs and returning text output. The context window is 262,144 tokens. It supports tool calling and structured JSON output. Throughput is 21 tokens per second with a P50 latency of 11.33 seconds. The model is offered through OpenRouter, Cloudflare AI Gateway, and the coding agent OpenCode, with a one-week free trial period that started on launch day.

The provider field on all three platforms reads only "Stealth." No lab, no model card, no weights. The original reporting comes from PC Watch, a Japanese technology publication under the Impress media group, which placed the story in the context of East Asian AI development. AsiaAI.FYI, a newsletter that translates Japanese and Chinese tech news for Western readers, picked it up and framed it as a challenge to the pricing structures of established players.

A data retention policy accompanies the launch: prompts and completions may be retained by the provider but are not used for training. All other use is governed by OpenRouter's Stealth Model Terms. During the free preview, no credits are charged. What happens to pricing after the preview ends has not been announced.

The Terminal-Bench 4.0 Context

Terminal-Bench has gone through several versions, each with a distinct task set. Version 4.0, released on August 28, 2026, reduced the task count from 74 in version 3.0 to 66, removing 8 tasks and revising 20 others. The benchmark covers software engineering, machine learning, science, operations, security, hardware, and media tasks. Each task requires an agent to perform real terminal work, not just answer questions.

The leaderboard currently shows GPT-6 Astra at the top with Codex as the agent framework, followed by Claude Fable 5.1 with Claude Code. GLM-5.3 from Zhipu AI sits at 41.8%, and GPT-5.6 Sol at 37.3%. These scores represent maximum effort configurations with full token budgets. The cost column reflects the total expense of running the full benchmark suite, which ranges from a few hundred dollars for lighter models to over $9,000 for Sonnet 5.

A developer report circulating after Union Alpha's launch claimed approximately 51% resolution at about $1.50 per task. If accurate, that would place it between GPT-6 Astra and Claude Fable 5.1 in accuracy while costing roughly $100 for a full run instead of $3,000 to $6,000. The benchmark community has not yet validated this claim, and Terminal-Bench data explicitly warns against using it in training corpora.

Why the Cost Structure Matters

The existing Terminal-Bench leaderboard reveals a pattern. Top-performing models consume billions of tokens per evaluation run. GPT-6 Astra uses 1.5 billion tokens. Claude Fable 5.1 uses 2.7 billion. Sonnet 5, which scores only 12.4%, consumes 21.6 billion tokens. Token consumption directly drives cost, and the relationship is not linear.

A model that achieves 51% at $1.50 per task would need to be dramatically more token-efficient than the current leaders, or the provider is subsidizing the cost during the preview period, or both. The free trial complicates the picture further. During the one-week preview, the cost is zero for all users. Any claims about cost-per-task during this period are theoretical projections based on the provider's pricing, not actual charges.

For developers evaluating models for production coding workflows, the relevant question is not what Union Alpha costs during a free trial. It is what it will cost after the preview ends, and whether the performance holds up on independent evaluations outside the developer report.

The Broader Trend of Stealth Releases

Union Alpha is the latest in a series of anonymous models appearing on OpenRouter in 2026. Pony Alpha appeared in February before Zhipu AI confirmed it as part of GLM-5. Hunter Alpha surfaced in March and sparked speculation about DeepSeek's involvement. Ox Alpha launched in August with a one-million-token context window and scored 80% on DeepSWE coding benchmarks before technical analysis pointed to Zhipu AI's GLM-5.3 family.

Each stealth release follows a similar playbook: appear anonymously on OpenRouter, offer free access during a preview period, claim frontier-level performance, and let the community run evaluations. The anonymity generates buzz. The free access drives adoption. The benchmark claims attract developer attention. Whether the model is eventually claimed by a known lab, remains permanently anonymous, or quietly disappears varies by case.

What is consistent across these releases is the pricing pressure they create. If an anonymous model can claim 51% on Terminal-Bench 4.0 at a fraction of the cost of named alternatives, it forces the conversation about what developers are actually paying for when they use frontier models. Brand recognition, support, and reliability have value. But so does a 95% cost reduction on benchmark runs.

What to Watch

Three things will determine whether Union Alpha matters beyond the initial hype cycle. First, independent benchmark validation. The 51% claim needs verification from developers who are not the provider. Second, post-preview pricing. The free trial tells you nothing about what the model will cost in production. Third, provider identity. A model that remains permanently anonymous limits its usefulness for teams that need vendor accountability, support contracts, and compliance guarantees.

For now, the model is free to try through OpenRouter, Cloudflare AI Gateway, and OpenCode. The 262K context window and multimodal support make it worth testing for coding and research workflows, if only to see whether the performance claims hold up on your own tasks. The Terminal-Bench 4.0 results, once independently verified, will tell the real story.