Anthropic and OpenAI are shopping for smaller AI data center deals, a shift from the massive facilities they have pursued over the past year, as both companies race to get computing capacity online fast enough to meet demand.
Both labs have signed enormous agreements for facilities ranging from hundreds of megawatts to gigawatts of capacity. But sources familiar with the conversations told CNBC that they are also pursuing much smaller deployments in the 20 to 30 megawatt range, enough to serve production workloads without waiting for a mega-project to come online.
Smaller deals, faster delivery
Anthropic has held discussions about agreements in the 20 to 30 MW range across the United Kingdom and the Nordics, according to four people familiar with the conversations. OpenAI had been exploring similar smaller-capacity opportunities in the Nordics, two of the sources said. A separate source said both companies have also looked at U.S. deployments at that scale.
"We're building a diversified compute portfolio to meet growing demand for AI around the world," an OpenAI spokesperson told CNBC. "Different workloads need different infrastructure, so we have conversations with a range of partners and assess opportunities based on our requirements, performance, reliability, timing and cost."
Anthropic did not comment.
The appeal is speed. "Securing a few megawatts at an existing powered site can be more practical than waiting for a much larger block in one location," said Jabez Tan, head of research at Structure Research. "For workloads that can operate across separate sites, a collection of smaller deployments can add up to substantial capacity."
The training-to-inference shift
The move toward smaller facilities reflects a broader change in where AI computing power is going. Training a large model requires many chips working closely together on huge datasets. Inference, the process of serving those models to end users once they are built, can be split across smaller clusters handling separate requests in different locations.
That split matters as more AI capacity moves from building models to running them. A report by real estate company JLL projects that the share of global data center capacity used for inference will overtake training by 2027. In 2025, inference made up 9% of total data center workloads, compared to 14% for training. By 2030, inference is expected to use 37% of capacity, while training drops to 13%.
In February, Nvidia announced plans to collaborate with several data center stakeholders to study smaller-scale facilities designed specifically for distributed inference.
What the big deals still look like
The smaller deals do not replace the headline-grabbing commitments. Anthropic signed a roughly $45 billion cloud deal with Nscale in August to rent around 460 MW of compute capacity at a data center development in West Virginia. OpenAI surpassed the original 10 GW commitment to its Stargate AI infrastructure project in April and has since committed to developing a further 3 GW in Georgia and 8 GW in Ohio.
But large projects face growing friction. Local communities in the U.S. increasingly push back against data center developments, and in much of Europe, available land and power are in short supply. Smaller deployments sidestep some of that pressure by fitting into existing powered sites.
Neoclouds pivot as well
Crusoe, the U.S. company that built a large data center complex in Texas used by OpenAI, is now investing in smaller facilities. The Wall Street Journal reported on Thursday that these builds will be faster and cheaper than larger projects, which have faced delays across the country. Crusoe did not respond to a request for comment.
The company also announced a $3.9 billion funding round on Thursday, valuing it at $30.9 billion post-money. Crusoe is one of several neoclouds whose business has boomed during the AI buildout.
The pattern across the industry is the same: companies that signed long-term, large-scale leases are now layering smaller, faster-to-deploy capacity on top. The reason is straightforward. Training happens once. Serving happens continuously, and the demand for serving is growing faster than anyone can build a single massive facility. Spreading inference workloads across multiple smaller sites gets them running sooner, and the economics of inference, where requests are independent and parallelizable, suit that architecture naturally.