IBM and Together AI have signed a multi-year, $240 million agreement aimed at expanding the infrastructure available for running open-source artificial intelligence models at enterprise scale.
Under the agreement, IBM plans to deploy a large cluster of NVIDIA HGX B300 systems on IBM Cloud, with availability expected in the first quarter of 2027. Together AI will use the infrastructure to provide open-source model inference services to its customers.
The deployment represents another example of cloud providers racing to build infrastructure specifically around AI inference as enterprises move beyond experimentation and begin putting more models into production.
Building Infrastructure Around AI Inference
The planned environment will be IBM Cloud’s first dedicated large-scale inference cluster using NVIDIA HGX B300 systems.
It will also use NVIDIA Spectrum-X Ethernet networking, creating an infrastructure stack designed around the performance and networking demands of production AI workloads.
NVIDIA says the HGX B300 platform can deliver significantly greater AI factory output than previous generations.
For Together AI, the additional capacity could help address one of the increasingly important economics questions surrounding enterprise AI: how much it costs to generate and process tokens at scale.
Together AI’s platform covers inference, training, fine-tuning and agentic AI workflows. The company says its inference business currently serves approximately 400 trillion tokens per month.
Open Models Move Deeper Into The Enterprise
The agreement also highlights growing competition between open-source and proprietary AI models.
Together AI has built its business around giving developers and enterprises access to open models alongside infrastructure optimized to run them efficiently.
As organizations deploy more AI applications, the infrastructure underneath those models can become almost as important as the models themselves. Compute availability, networking, latency and cost all influence whether an AI deployment can move economically from a pilot into production.
IBM’s involvement gives Together AI another route for delivering that infrastructure to large organizations.
IBM Expands Its NVIDIA Infrastructure Strategy
The deal also builds on IBM’s broader relationship with NVIDIA.
IBM Cloud already supports NVIDIA GPU infrastructure, while the companies have been expanding their work across areas including data analytics, unstructured data processing, cloud infrastructure and enterprise AI services.
The Together AI deployment pushes that relationship further into inference infrastructure.
Instead of simply providing GPU capacity, IBM is positioning its cloud as an environment where AI providers can deploy dedicated infrastructure designed around increasingly large production workloads.
That distinction could become more important as AI infrastructure shifts from supporting model training alone toward handling continuous inference generated by enterprise applications and AI agents.
The Bottom Line
The $240 million agreement shows how quickly the AI infrastructure market is moving toward production-scale inference.
For IBM, the Together AI deal gives IBM Cloud a major open-source AI workload built on NVIDIA’s latest infrastructure. For Together AI, it provides additional compute capacity as the company attempts to make open models more competitive with proprietary alternatives on performance and cost.
The bigger story is that enterprise AI competition is increasingly moving below the model layer. As organizations deploy more AI applications and agents, the providers that can deliver reliable compute at the lowest practical cost per token could become just as important as the companies building the models themselves.

