6 Sources
[1]
Together AI embraces the competition with $240M IBM Cloud deal
With power, datacenter capacity, and other supply chain constraints, AI infrastructure is in short supply, and service providers will rent all the compute they can get, even if it means working with competing cloud providers. Together AI is the latest example. This week the service provider announced a $240 million deal to run its open weights inference platform on a "large cluster" of Nvidia GPUs housed in IBM Cloud. "Together AI selected IBM with Nvidia because of their innovative product roadmaps and their ability to deliver GPU capacity at the pace required for rapid AI scaling and lowest token cost," IBM's announcement says. Translation: Together AI tapped IBM because Big Blue had the capacity it needed when it needed it. Together AI sits toward the top of the AI inference ecosystem. Its business model largely revolves around renting compute from cloud or neocloud providers, many of which themselves rent floorspace and capacity from bit barn operators and operate competing services. More recently, the company has begun deploying GPU compute in datacenters in Maryland, Memphis, and Sweden, but like most AI service providers, its main value proposition remains making it easier for users to consume the hardware for applications like inference, fine-tuning, and training. For inference, this largely boils down to an OpenAI-compatible API endpoint. The service provider isn't particularly picky about what hardware that API endpoint runs on top of either, so long as price performance is favorable. For example, Together AI will be deploying services atop SambaNova's heterogeneous compute platform built in collaboration with Intel and leveraging Nvidia GPUs for prefill processing. Those systems went live in Vector Core Compute's new AI bit barn earlier this year. In this case, Together AI is getting Nvidia B300s, no doubt because that's what IBM had to offer them. Announced in early 2025, the HGX B300 platform isn't Nvidia's most powerful offering. Unlike the 72-GPU rack systems that CEO Jensen Huang likes to show off every opportunity he gets, the HGX platform is much more conventional and can be deployed in traditional air-cooled datacenters. The 14-15 kW boxes each contain eight B300 GPUs that are interconnected by a combination of NVLink in the box, and Nvidia's Spectrum-X Ethernet fabrics between them. The deployment, IBM says, is its first large scale use of B300s for inference applications. The compute capacity is expected to come online in the first quarter of 2027. ÂŽ
[2]
IBM, Together AI ink $240 million deal for Nvidia-powered AI inference cluster
Aug 11 (Reuters) - IBM (IBM.N), opens new tab and startup Together AI have signed a $240 million multi-year agreement to build a large-scale artificial intelligence cluster on IBM Cloud using Nvidia (NVDA.O), opens new tab systems, the companies said on Tuesday. The cluster will provide inference for open-source AI models, which have gained traction as businesses seek to rein in AI â costs and weigh concerns about cybersecurity incidents involving models from Anthropic, OpenAI and Meta (META.O), opens new tab. Inference, the process of running trained AI models to generate responses, has become one of the largest drivers of demand for computing capacity, prompting cloud providers and chipmakers to spend billions of dollars expanding â AI infrastructure. The cluster on IBM Cloud will use Nvidia's HGX B300 systems, which link the chipmaker's newer Blackwell processors, and its Spectrum-X Ethernet networking gear. â Nvidia has said the Blackwell chips are optimized for AI inference. San Francisco-based Together AI's platform lets companies â train and run AI workloads on open models such as DeepSeek, MiniMax and Kimi â at lower costs than closed systems. It was last valued at $8.3 billion in July. Reporting by Anhata Rooprai in Bengaluru; Editing by Tasim Zahid Our Standards: The Thomson Reuters Trust Principles., opens new tab
[3]
IBM bets $240m on cheap, open-source inference to take on the hyperscalers
A multi-year deal with the startup Together AI will put Nvidia Blackwell systems on IBM Cloud, and the wager is that enterprises now care more about the cost of running AI than the prestige of the model doing the running. IBM has decided that the money in artificial intelligence is no longer only in building the cleverest model, but in running it cheaply, and it has put $240m behind that conviction. The company has signed a multi-year agreement with Together AI, a San Francisco startup, to stand up a large-scale AI inference cluster on IBM Cloud, and the pitch is squarely aimed at enterprises trying to trim their AI bills. The cluster will run on Nvidia HGX B300 systems, built on the Blackwell architecture that Nvidia markets as tuned specifically for inference, knitted together with the company's Spectrum-X Ethernet networking, which is the same recipe that a wave of specialists have been chasing as they bet that AI's real profits lie in cheap inference rather than in ever-larger training runs. Together AI is an interesting partner to pick. Valued at $8.3bn as of July, its platform lets companies train and run workloads on open-source models, including DeepSeek, MiniMax and Kimi, and it sells itself as a cheaper, more flexible alternative to the closed, proprietary systems that dominate the headlines. It reportedly serves around 400 trillion tokens a month, a figure that gives some sense of how much raw inference now sloshes through these pipes, and of why a legacy vendor such as IBM would want a share of that traffic running on its own cloud. Inference, the business of actually answering queries once a model is trained, has quietly become one of the largest drivers of demand for computing capacity, which is why the layer has attracted a scramble of money and talent. Nebius recently paid $643m to absorb a 20-person team working on inference optimisation, a price that only makes sense if you believe that shaving cost per token is where the margins now live. The other half of the story is the slow enterprise drift towards open source. Businesses want to cut what they spend on AI, open models are gaining genuine traction inside large organisations. And there is a security dimension too, since cybersecurity worries about closed models from Anthropic, OpenAI and Meta are cited as one reason some firms would rather run something they can inspect and host themselves. None of that is a purely technical preference, because for a bank or a hospital the ability to keep sensitive data on infrastructure it controls is often the whole point. For IBM, hosting cheap, open-source inference is a way to pick a fight it might actually win. The company was never going to out-muscle Amazon, Microsoft or Google on sheer cloud scale, so positioning IBM Cloud as the place to run open models affordably lets it compete on economics rather than size, and it dovetails neatly with a wider European appetite for infrastructure that is not locked to a single American giant. That appetite is showing up elsewhere on the map. The push to build inference capacity that enterprises and governments can trust and control has spawned outfits such as TensorX, which raised âŹ8m to build sovereign AI inference for Europe on Nvidia Blackwell, a reminder that the same silicon underpinning IBM's bet is being wired into a broader argument about who owns the plumbing. There is a whiff of gold rush about all this, and IBM knows it. The open-versus-closed contest in enterprise AI has often been framed as a debate about capability. Yet the more telling battle is now about price, and a $240m cluster full of Blackwell chips is IBM's way of saying it would rather sell the picks and shovels than the promise. Whether cheap, open inference proves as sticky and as profitable as its backers hope is the wager the whole sector is quietly making, and a $240m cluster is one of its larger stakes yet.
[4]
IBM and Together AI sign $240 million AI inference deal
IBM $IBM and Together AI signed a multi-year $240 million agreement on Tuesday to build a large-scale AI inference cluster on IBM Cloud using Nvidia $NVDA hardware, the companies said. The cluster is built around Nvidia HGX B300 systems paired with Nvidia Spectrum-X $TWTR Ethernet networking gear. IBM said it is the first dedicated, large-scale cluster built for inference on IBM Cloud using those systems. Availability is expected in the first quarter of 2027, the company said. Together AI will use the cluster to run inference on open-source AI models. The San Francisco-based company operates a platform through which developers and enterprises can build and deploy AI workloads, and it reports serving 400 trillion tokens monthly. "Enterprises want the performance of the best frontier models without the closed-model price tag, and that only works if the infrastructure underneath is fast and reliable at scale," Together AI CEO Vipul Ved Prakash said in a statement. "This cluster lets us bring production-grade inference to more companies, faster." Alan Peacock, general manager of IBM Cloud, said in a statement that IBM and Nvidia are "delivering scalable, economical, enterprise-grade AI infrastructure" to help Together AI accelerate innovation. Together AI raised an $800 million Series C financing round at an $8.3 billion valuation, the company said. It added that it selected IBM and Nvidia based on their product roadmaps and ability to deliver GPU capacity at the pace required for AI scaling at low token cost. IBM's push into open-source AI infrastructure follows a separate commitment to the open-source ecosystem. Earlier this year, IBM announced Project Lightwell, a $5 billion commitment with Red Hat to help enterprises secure open-source software using AI tools and more than 20,000 engineers. That initiative centers on a trusted enterprise clearinghouse that uses AI to vet and verify patches across large portions of the open-source ecosystem, with pilot participants including Bank of America $BAC, Goldman Sachs $GS, and JPMorganChase. The IBM and Together AI deal is part of a broader collaboration between IBM and Nvidia that also spans GPU-native data analytics, unstructured data extraction, on-premises and cloud infrastructure, and consulting services, IBM said.
[5]
IBM inks $240M infrastructure deal with AI-optimized cloud operator Together AI
IBM Corp. will provide infrastructure to artificial intelligence startup Together AI Inc. as part of a $240 million deal announced today. The partnership comes a few weeks after the latter company raised $800 million in funding. The consortium that provided the capital included Nvidia Corp., whose chips IBM will use to run Together AI's workloads. San Francisco-based Together AI operates a cloud platform optimized for machine learning. The platform runs on infrastructure that the company rents from partners such as Amazon Web Services Inc. Developers can use it to train custom AI models, fine-tune existing ones and perform inference. IBM disclosed today that the platform processes about 400 trillion tokens per month for customers. The infrastructure deal announced today will see the tech giant deploy a "large cluster" of HGX B300 systems for Together AI. The hardware will become available through IBM Cloud in the first half of 2027. The HGX B300 is a motherboard equipped with eight Blackwell Ultra chips. The Blackwell Ultra was Nvidia's most advanced graphics card until the introduction of the Rubin in January. It can provide 15 petaflops of performance when processing data in the NVFP4 format. The Blackwell Ultra comprises two chiplets linked together by a custom interconnect. They host 160 computing modules called streaming multiprocessors, which each comprise more than 130 cores. They also contain memory optimized to store intermediate results, temporary data that AI models generate while processing prompts. The HGX B300 uses eight ConnectX-8 SuperNIC chips to move data between its Blackwell Ultra accelerators. A BlueField-3 data processing unit, or DPU, coordinates traffic to and from external systems. Hardware makers use the HGX B300 to build AI servers that also include contain components. According to Nvidia, the machines include at least two terabytes of memory and a pair of central processing units with 56 cores apiece. Those CPUs connect to two PCIe switches that enable users to attach several terabytes NVME flash memory. Together AI will use IBM's HGX B300 cluster to host workloads for customers of its public cloud. In particular, it will run inference workloads that use open-source AI models. Together AI offers five different inference services as part of its platform. One is optimized for media generation models that run in containers. Another service, Provisioned Throughput, is charged by the token to help customers optimize their costs. Together AI's other inference offerings provide serverless infrastructure, dedicated machines and a batch processing environment that trades off some performance for a price reduction. "This cluster lets us bring production-grade inference to more companies, faster, and it's a big step in our push to make open-source AI the obvious choice for enterprises," said Together AI Chief Executive Officer Vipul Ved Prakash.
[6]
IBM Lands $240 Million AI Deal with Together AI to Scale Inference - IBM (NYSE:IBM)
The companies aim to deploy a large cluster of NVIDIA HGX B300 systems on IBM Cloud, with availability expected in the first quarter of 2027. * IBM stock is showing upward movement. What's pushing IBM stock higher? IBM and Together AI Partner To Scale AI Inference This collaboration aids in enhancing AI workloads for enterprises, positioning IBM to better serve its clients during a mixed market backdrop. The partnership aims to create the first large-scale inference cluster on IBM Cloud, designed to facilitate fast and efficient production for AI applications. Together AI will use the infrastructure for open-source AI model inference. The deployment will be the first dedicated large-scale inference cluster on IBM Cloud using HGX B300 systems and NVIDIA Spectrum-X Ethernet networking, designed to significantly boost AI processing capacity. IBM Technical Outlook: Key Levels and Momentum Currently, IBM is trading at $239.38, which is about 7.4% above its 20-day simple moving average (SMA) of $222.62. However, the stock is approximately 7.1% below its 50-day SMA of $257.31, indicating a bearish trend as the 20-day SMA is below the 50-day SMA, a sign of potential weakness. The Relative Strength Index (RSI) is at 49.90, suggesting a neutral momentum state, indicating that the stock is neither overbought nor oversold. This positioning allows for potential movement in either direction, depending on market sentiment and upcoming catalysts. * Key Resistance: $258.50 -- Nearby level where rebounds can stall. * Key Support: $199.00 -- Level where buyers previously stepped in. Technology Sector Performance and IBM Relative Strength The Technology sector is currently ranked eight out of 11 sectors, reflecting a mid-tier performance with a decline of 0.19%. Over the past 30 days, the sector has gained 2.58%, but it has seen a more modest increase of 5.15% over the last 90 days. While IBM is outperforming its sector today by about 1.49 percentage points, the overall trend suggests the sector is facing headwinds. This context is important for traders as it highlights the stock's relative strength against a backdrop of broader sector weakness. IBM Earnings Preview and Analyst Price Targets IBM is slated to provide its next financial update on Oct. 21(estimated). * EPS Estimate: $2.89 (Up from $2.65) * Revenue Estimate: $16.85 billion (Up from $16.33 billion) * Valuation: P/E of 21.0x (Indicates fair valuation) Analyst Consensus & Recent Actions: The stock carries a Buy rating with an average price forecast of $256.12. Recent analyst moves include: * Citigroup: Buy (Lowers target to $245 on July 24) * Susquehanna: Neutral (Lowers target to $225 on July 24) * Morgan Stanley: Equal-Weight (Lowers target to $190 on July 23) How IBM Ranks On Value, Growth, Quality and Momentum Below is the Benzinga Edge scorecard for IBM, highlighting its strengths and weaknesses compared to the broader market: - Value: 27.78 -- Stock is trading at a steep premium relative to peers. - Growth: 40.42 -- Indicates moderate growth potential. - Quality: 80.22 -- Reflects a strong balance sheet and operational efficiency. - Momentum: 18.95 -- Stock is underperforming the broader market. The Verdict: IBM's Benzinga Edge signal reveals a mixed profile, with strong quality metrics but weak momentum. This suggests a focus on stability and potential for growth, albeit with challenges in market performance. Top ETFs Holding IBM Stock Significance: Because IBM carries such a heavy weight in these funds, any significant inflows or outflows for these ETFs will likely force automatic buying or selling of the stock. IBM Stock Rises As Investors React To AI Deal IBM shares were up 2.27% at $238.58 at the time of publication on Tuesday, according to Benzinga Pro data. Photo via Shutterstock This content was partially produced with the help of AI tools and was reviewed and published by Benzinga editors. Market News and Data brought to you by Benzinga APIs To add Benzinga News as your preferred source on Google, click here.
Share
Copy Link
IBM and Together AI signed a $240 million multi-year agreement to build a large-scale AI inference cluster on IBM Cloud using Nvidia HGX B300 systems. The cluster will run open-source AI models including DeepSeek and MiniMax, targeting enterprises seeking cost-effective AI infrastructure. The deployment marks IBM's strategic bet on affordable inference over closed proprietary systems.
IBM Cloud
1
and Together AI2
have signed a $240 million multi-year agreement to build a large-scale AI inference cluster on IBM Cloud using Nvidia HGX B300 systems3
. The partnership positions IBM to compete on economics rather than scale against hyperscalers like Amazon, Microsoft and Google. Together AI CEO Vipul Ved Prakash stated that enterprises want the performance of frontier models without the closed-model price tag, and this only works if the infrastructure underneath is fast and reliable at scale4
. The cluster is expected to come online in the first quarter of 20271
.
Source: The Next Web
The cluster will provide AI inference for open-source AI models including DeepSeek, MiniMax and Kimi
2
, which have gained traction as businesses seek to rein in AI costs and address cybersecurity concerns involving models from Anthropic, OpenAI and Meta2
. Together AI operates a platform that lets companies train and run AI workloads on open weights models at lower costs than closed systems2
. The San Francisco-based company was valued at $8.3 billion in July2
and reportedly serves around 400 trillion tokens monthly3
, giving a sense of how much raw inference now flows through these systems.
Source: SiliconANGLE
The deployment will use Nvidia HGX B300 systems built on Nvidia Blackwell architecture, paired with Spectrum-X networking gear
4
. The HGX B300 platform contains eight Blackwell Ultra chips per system and can be deployed in traditional air-cooled datacenters1
. These 14-15 kW boxes interconnect GPUs using NVLink within the box and Nvidia's Spectrum-X Ethernet fabrics between them1
. IBM stated this marks its first dedicated large-scale cluster built for inference applications using B300 systems4
. The HGX B300 is a motherboard equipped with eight Blackwell Ultra chips that can provide 15 petaflops of performance when processing data in the NVFP4 format5
.IBM's partnership represents a strategic bet that AI's real profits lie in cheap inference rather than ever-larger training runs
3
. Together AI selected IBM and Nvidia based on their product roadmaps and ability to deliver GPU capacity at the pace required for AI scaling at low token cost4
. The move lets IBM compete on economics rather than size, positioning IBM Cloud as the place to run open models affordably3
. This dovetails with a wider European appetite for infrastructure that is not locked to a single American giant and supports data sovereignty concerns3
.Related Stories
Inference has become one of the largest drivers of demand for computing capacity, prompting cloud providers and chipmakers to spend billions expanding AI infrastructure
2
. Together AI will use the cluster to host AI workloads for customers through five different inference services including media generation models, Provisioned Throughput charged by token, serverless infrastructure, dedicated machines and batch processing5
. The platform processes about 400 trillion tokens per month for customers5
. This partnership follows Together AI's recent $800 million Series C funding round that included Nvidia among investors4
.
Source: The Register
The deal aligns with IBM's broader push into open-source AI infrastructure through initiatives like Project Lightwell, a $5 billion commitment with Red Hat to help enterprises secure open-source software using AI tools and more than 20,000 engineers
4
. That initiative centers on a trusted enterprise clearinghouse using AI to vet and verify patches across the open-source ecosystem, with pilot participants including Bank of America, Goldman Sachs and JPMorgan Chase4
. The IBM and Together AI deal is part of a broader collaboration between IBM and Nvidia spanning GPU-native data analytics, unstructured data extraction, on-premises and cloud infrastructure, and consulting services4
. For banks and hospitals, the ability to keep sensitive data on infrastructure they control through fine-tuning and sovereign AI approaches is often the whole point3
.Summarized by
Navi
[1]
21 Feb 2025â˘Business and Economy

20 Oct 2025â˘Technology

01 Jul 2026â˘Startups

1
Technology

2
Science and Research

3
Technology
