Nebius, the Amsterdam-listed AI cloud infrastructure company spun out of Yandex, has acquired Inferize — an AI inference optimization startup that was barely ten months old at the time of the deal. According to Globes reporting, the acquisition adds a team specializing in making large language model inference faster and cheaper, plugging directly into Nebius’s ambition to compete with the hyperscalers on GPU-powered AI compute. The financial terms of the deal were not disclosed.
The move is a telling signal about where the real bottleneck in AI infrastructure sits right now. Training a model gets the headlines, but inference — serving live predictions at scale, at low latency, and without burning through margins — is where the economics of AI products are actually won or lost. That is precisely the problem Inferize was built to solve, and it is the problem Nebius needs solved as it scales its cloud GPU business. For more context on the venture capital appetite surrounding early-stage AI infrastructure plays, see our coverage of startup mega-rounds that have defined recent funding cycles.

What Inferize Actually Built
Inferize was founded in 2024 and, despite its short runway, had already developed tooling aimed at reducing the computational overhead of running inference workloads on GPU clusters. The startup’s core focus was inference efficiency — compressing the time and hardware cost required to serve model outputs at production scale. That kind of optimization work is increasingly valuable as enterprises push AI workloads beyond pilots and into high-volume deployments where per-token costs compound quickly.
The startup operated out of Israel, part of a dense cluster of AI and deep-tech companies that has continued attracting both local and international investment. Nebius’s decision to acquire rather than build internally suggests the team and technology represented a meaningful shortcut — buying proven inference expertise rather than assembling it from scratch over 12 to 18 months. Nebius did not specify headcount for the acquired team, but the deal is understood to bring the Inferize engineers fully onto the Nebius platform.
Why This Matters for Nebius’s Competitive Position
Nebius went public on Nasdaq in late 2024 after completing its separation from Yandex’s Russian assets, positioning itself as a pure-play AI infrastructure provider for the European and global markets. The company has been on an aggressive buildout trajectory, expanding GPU cluster capacity and signing partnerships to serve AI developers and enterprises that want an alternative to AWS, Google Cloud, or Azure. Acquiring inference optimization capability tightens the stack it can offer customers — not just raw compute, but smarter compute.

The competitive pressure here is real. CoreWeave, Lambda Labs, and a growing roster of GPU cloud providers are all racing to differentiate on performance and price, while the hyperscalers pour billions into custom silicon. Nebius’s bet is that software-layer inference optimization, combined with its GPU infrastructure, creates a more efficient and cost-competitive offering than raw horsepower alone. This acquisition is a small but concrete step in that direction. The broader dynamic — large, well-capitalized AI platforms absorbing young, technically sharp startups before they can raise a Series A — is one investors are watching closely, as seen in the SoftBank-OpenAI capital flows that have reshaped how money moves through the AI ecosystem. For Inferize’s founders, a ten-month path from founding to acquisition is a remarkable outcome; for Nebius, it is a targeted technical bet on the layer of AI infrastructure that enterprises are increasingly willing to pay a premium to optimize.
