AI agent testing has entered a new phase where every action can be traced, replayed, and verified against a known state. Sparta's Worlds platform offers exactly that: a sealed twin of the systems an agent touches, allowing teams to run agents against isolated environments and get exact, replayable proof of what happened. The platform just entered public beta, backed by a16z, and connects to Stripe, Zendesk, and other common tools.

A sealed twin for every agent run

Worlds creates what the site calls a "world" — a working copy of a system your agent acts on. Same endpoints, same errors, same state machine, but no real money and no real customers. A world starts from a seed: the bundled sample account, or a pseudonymized copy of your own. It is ready in milliseconds and identical every time. The agent runs against this world using its configured SDK or HTTP client. No code changes are needed.

When the agent finishes, the platform diffs the environment and tells you what really changed. Every refund, credit, and cancellation is recorded in dollars, and whose money moved is clear. You can replay any run byte for byte. This means if an agent mistakenly charges the wrong customer, you can reconstruct the exact sequence of requests that led to it.

Mark each state, diff every change

A mark records the world's state at a given moment. Every diff is measured from a mark, so what your agent did is never mixed up with what was already there. You start with a fresh mark, let the agent act, and then read the diff. This separation is the core of how Worlds guarantees replayability without requiring deterministic models.

The dashboard labels scripted samples for billing, native support, and combined workflows. Teams can use these to learn the controls, then connect their own agent. The replay guarantees apply to the same world inputs and request sequence; they do not make a model deterministic. The point is that given the same starting state and the same inputs, the environment changes in a known, measurable way.

Break it, then fix it

Worlds lets you inject failures deliberately. Rate limits, 500 errors, lock timeouts, slow responses, expired keys, and dropped webhooks can all be triggered. This tests whether your retry logic handles edge cases correctly. For example, you can find out whether your agent refunds twice when a webhook fires after a timeout. The platform also lets you advance the clock — 45 days, for instance — and let renewals bill, prorations settle, cancellations resolve, and webhooks fire in order. Then read the diff again to see what changed over that period.

From mocks to controllable worlds

Vendor test modes and sandboxes have long been the standard for agent evaluation. They offer isolation, but often lack controllable failures and a logical clock that moves on your schedule. Worlds gives each test its own starting state, controllable failures, and a logical clock within the connector's documented scope. A world runs beside your agent, not inside it. When your agent says a task is done, Worlds diffs the environment and tells you what really changed — not just what the agent reported.

The platform also includes conversation simulators. These exercise what an agent says to a customer, and let you check the records its actions change — including changes its reply never mentions. This closes a gap where an agent might confirm a task is complete while the underlying data tells a different story.

What this means for agent teams

For teams building AI agents that touch billing systems, customer support tools, or payment platforms, Worlds provides a way to prove what the agent actually did, not just what it said. The ability to reset, seed, and break on purpose means safety checks can be systematic rather than ad hoc. The logical clock and mark system mean you can test time-dependent behaviors — renewals, cancellations, prorations — without waiting weeks in real time.

Stripe and Zendesk are in the beta, with connectors for Shopify and Salesforce available on request. The platform is free during the internal testing phase. For any team that runs agents in production or near-production, the question becomes: how do you know the agent didn't just the move money to the wrong account, or miss a cancellation flag? Worlds answers that with a diff you can replay and verify.