Google has published benchmark results for a Kubernetes feature that captures the complete running state of a pod and restores it on demand, reporting startup latency reductions of up to 89 percent. A 70 billion parameter model loads in 37 seconds and an 8 billion parameter model in 15 seconds — a dramatic improvement over the initialization time that dominates large model deployment.
How Pod Snapshots Work
The feature is checkpoint and restore, not caching. A snapshot captures everything a workload had running: open file descriptors, threads, CPU registers, and memory, along with the container root filesystem, EmptyDir volumes, and tmpfs mounts. When a new replica starts from that snapshot, it never runs the initialization process that loads the model — which is where most of the startup time goes on large models.
The technical foundation is gVisor, Google's container runtime. Pods must run in GKE Sandbox for whole-pod snapshots to work. Autopilot clusters already include gVisor, while standard clusters require a node pool with it enabled. An agent on each node manages the snapshot lifecycle, a control plane controller clears obsolete snapshots, and Cloud Storage holds the data.
Two custom resources handle configuration. PodSnapshotStorageConfig points at the storage bucket. PodSnapshotPolicy selects pods by label, sets the trigger to either workload or manual, and configures retention with a last-access timeout and a cap on snapshots per group.
Real-World Impact at Codeway
Google's customer example is Codeway, whose Retake platform had previously built a custom caching layer for compiled artifacts that reduced startup to roughly one minute. With Pod snapshots, that dropped to eight seconds. The team now starts H100 instances for specific jobs and shuts them down when the work is done — a pattern that would have been impractical with cold-start latencies measured in minutes.
The Tradeoffs Practitioners Are Raising
Developer reaction has focused less on the capture mechanism than on what happens during restore. A senior DevOps and MLOps engineer noted that snapshot invalidation may prove a harder platform problem than capture itself. Model digest, CUDA and driver versions, GPU type and topology, and runtime configuration all become part of the compatibility key. Secrets, DNS, and downstream connections need explicit rehydration after restore.
Google's documentation addresses some of this. GKE builds a hash from the pod's essential runtime fields — called the distilled pod spec — and embeds it in the snapshot. A pod restoring from it must produce an identical hash. The target node must have an identical machine series and CPU architecture, and the gVisor kernel and GPU driver versions must match what was captured. When no compatible snapshot exists, the pod starts normally through a documented fallback.
A rootfs-only scope relaxes these rules. Because process memory is not restored, snapshots can cross machine families, including to E2 instances. But the rehydration burden falls to the application. Encryption keys and certificates created before a snapshot must be recreated afterward. Environment variables live in application memory where gVisor cannot reliably find and replace them. External connections are terminated, persistent volumes are not checkpointed, and user-added iptables or nftables rules are not restored.
The Hidden Operational Cost
One consequence that has drawn attention is that upgrading a node pool can invalidate existing snapshots if the gVisor kernel or GPU driver version changes. When that happens, the pod starts normally and the performance benefit disappears — silently, without error. Restore is also not instantaneous despite the headline numbers. gVisor's kernel comes back within a few seconds, the application begins running, and memory continues loading in the background.
Hardware support is narrower than the general framing suggests. Whole-pod snapshots do not work on E2 machine types, multi-GPU pods are supported only on L4 GPUs, and GPU sharing through Multi-Instance GPU is not available.
The Agent Sandbox Connection
The feature has particular implications for AI agent workloads. GKE Agent Sandbox reached general availability in May, with a warm pool that Google says allocates up to 300 sandboxes per second per cluster and routes 90 percent of them within 200 milliseconds. Pod snapshots are used to suspend idle agents rather than hold compute warm. A related open-source project called Agent Substrate explores the same suspend-and-resume multiplexing at higher density, though Google notes it is not ready for production.
A file in Cloud Storage holds the complete memory of a running workload — which, in the agent sandbox context, is memory that may have executed untrusted, model-generated code. Access control depends on Workload Identity Federation and IAM bindings for each pod's service account, which Google acknowledges can take time to propagate.
The feature is described as workload-agnostic, applicable to Java applications, game servers, and legacy monoliths alongside AI inference. The decisions it leaves to teams are significant: which node pools run gVisor, which bucket holds snapshots and who can read them, how long snapshots persist, and what the workload must refresh when it resumes from a state it did not expect to be frozen in.