2 Sources
[1]
Atlas Cloud optimizes AI inference service to boost GPU throughput - SiliconANGLE
Atlas Cloud optimizes AI inference service to boost GPU throughput Cloud infrastructure startup Atlas Cloud today launched a highly optimized artificial intelligence inference service that it says dramatically reduces the computational requirements of even the most demanding AI workloads. The new
[2]
Atlas Cloud Launches High-Efficiency AI Inference Platform, Outperforming DeepSeek
Developed with SGLang, Atlas Inference surpasses leading AI companies in throughput and cost, running DeepSeek V3 & R1 faster than DeepSeek themselves. NEW YORK, May 28, 2025 (Newswire.com) - Atlas Cloud, the all-in-one AI competency center for training and deploying AI models, today announced the
Share
Copy Link
Atlas Cloud introduces Atlas Inference, a highly optimized AI inference service that significantly boosts GPU throughput and reduces computational requirements for AI workloads.
Atlas Cloud, a cloud infrastructure startup specializing in AI workloads, has launched a highly optimized artificial intelligence inference service called Atlas Inference. This new offering promises to dramatically reduce the computational requirements of even the most demanding AI workloads, potentially revolutionizing the economics of AI deployment
1
.
Source: SiliconANGLE
Atlas Inference, co-developed with SGLang, an AI inference engine, claims to deliver 2.1 times greater throughput for AI workloads compared to equivalent services offered by industry giants such as Amazon Web Services and Nvidia
1
. The platform's ability to process 54,500 input tokens and 22,500 output tokens per second per node significantly outperforms current industry standards2
.In a notable achievement, Atlas Inference's 12-node cluster outperformed DeepSeek Ltd.'s reference implementation for the DeepSeek V3 model while using only two-thirds of the server's computational capacity. This impressive feat was accompanied by an 80% reduction in operational expenses
1
.The exceptional performance of Atlas Inference is attributed to four key innovations:
1
.Atlas Inference boasts linear scaling behavior across nodes, which automates the expansion and contraction of GPU clusters in real-time. This feature optimizes infrastructure costs and provides a more cost-effective solution for businesses deploying AI models
1
.Jerry Tang, CEO of Atlas Cloud, emphasized the platform's potential to change the economics of AI deployment: "Our platform's ability to process 54,500 input tokens and 22,500 output tokens per second per node means businesses can finally make high-volume LLM services profitable"
2
.Related Stories
Atlas Inference is designed to work with standard hardware and supports custom models, offering customers complete flexibility. Organizations can upload fine-tuned models and keep them isolated on dedicated GPUs, making the platform ideal for those requiring brand-specific voice or domain expertise
2
.Yineng Zhang, Core Developer at SGLang, believes that Atlas Inference represents a significant leap forward for AI inference: "What we built here may become the new standard for GPU utilization and latency management. We believe this will unlock capabilities previously out of reach for the majority of the industry regarding throughput and efficiency"
2
.The launch of Atlas Inference could have far-reaching implications for the AI industry, potentially enabling more businesses to profitably deploy and run large language models. As AI continues to play an increasingly crucial role in various sectors, innovations like Atlas Inference may accelerate the adoption and implementation of AI technologies across industries.
Summarized by
Navi
29 Jul 2025•Technology

11 Jun 2026•Technology
27 Mar 2025•Technology

1
Science and Research

2
Technology
3
Technology