Two Flagship AI Models, Same Goal: Worms in a 20-Year-Old Pokemon Game

Security researchers at RunSybil pitted Anthropic's Claude Opus 5.5 against Moonshot AI's Kimi K3 to build self-replicating worms that turn a victim's Pokemon Emerald game into Tetris, permanently. Both models succeeded, but they took entirely different paths to get there.

The Experiment

The researchers chose Pokemon Emerald and Pokemon Red because they are heavily documented, 20-year-old targets with large communities dedicated to modding and hacking. The skill floor for discovering novel bugs is high. Both models were given the same objective: find new vulnerabilities and turn them into working exploits.

Both models found the same novel link-cable vulnerability and built a game-breaking worm to exploit it. Both also uncovered several previously unreported bugs. Because the outcome was the same, the researchers could compare how each model organized its work.

Claude Opus 5.5 was evaluated in early access through the Claude Code CLI. Kimi K3 ran at Max effort through the same harness, with both models operating under a 1-million-token context window per agent.

Two Approaches to Research

The divergence in workflow was immediate. Opus 5.5 deployed specialized subagents with distinct roles. One agent spun up an emulator and created save files at different game states for another agent that was already preparing scripts to test those sections. When the save files were ready, the scripts were waiting and ran instantly. The researchers described the coordination as feeling like watching a team work together.

Opus also brought in adversarial reviewers, demanded hard evidence before accepting findings, and dynamically weighed compute costs against new leads. It directed subagents to use version control to coordinate and record research, which gave the researchers explicit controls over who updated shared reports.

Kimi K3 took a different path. It ran parallel investigations with repeated iterations, prioritizing breadth before depth. It used a central notes file to communicate between agents, instructing each new agent to read the existing record before starting. When an agent stalled, Kimi would relaunch the task with a narrower scope.

This approach sometimes backfired. Two agents would collide in their work because of ambiguous prompting boundaries, with one agent writing a report before another agent had finished the underlying task. Kimi recovered from these errors, but the pattern revealed a weakness in how the model handled inter-agent communication.

Token Economics

The two models approached cost management differently. Opus 5.5 explicitly flagged computationally expensive pathways and was willing to spend resources to confirm findings, running a verification task for two hours to reproduce a bug from a save file. Opus is roughly 40% more expensive than Kimi K3, so the researchers noted that this willingness to spend might reflect a different cost sensitivity.

Kimi questioned whether additional tooling was worth the effort. When evaluating fuzzing approaches, Kimi rejected a save-state fuzzer because of token costs and build time. Kimi prioritized reducing time consumption and wasted cycles rather than token usage.

Divergent Findings

Despite both models achieving the same headline result, they found different bugs. Kimi discovered five bugs in areas that Opus never examined. Opus found two arbitrary code execution chains and three cosmetic bugs that Kimi missed.

The most striking behavioral difference involved the models' stance on obtaining the game ROM. Kimi had no issue building the worm but refused to download the game ROM, even though the source code was available, citing piracy concerns. Opus 5.5 initially hesitated to write a worm but downloaded a game ROM without hesitation.

What the Comparison Reveals

The experiment raises a practical question for security teams: does specialized offensive research require expensive top-tier reasoning models, or can open-weight alternatives perform at a similar level?

Opus 5.5 excelled at making proof requirements and compute tradeoffs explicit, but struggled with guardrails when it came to weaponizing its payload. Kimi K3 freely ran parallel investigations to increase payload severity but struggled with the mechanics of exploitation once it found the vulnerabilities.

Both models were evaluated using the same harness and the same prompts. The differences in behavior came from the models themselves, not from the tooling around them. That suggests the choice between models for security research depends on which weaknesses a team can tolerate: Opus 5.5's caution around harmful output, or Kimi K3's difficulty with the technical details of exploitation.

The full analysis includes detailed tables tracking true positives, false positives, and failure rates across challenge types, along with a breakdown of how each model handled logical proofs, fuzzing, and prior research. The complete findings are available on the RunSybil blog.