2 Sources
[1]
Majestic Labs ditches the GPU to beat Nvidia's memory wall
Majestic Labs has unveiled a server that drops the GPU for Arm cores and up to 128TB of cheap LPDDR6 memory. It claims to shame a rack of Nvidia chips on memory and power. None of it has shipped or been independently tested. The AI hardware conversation is stuck on one word: compute. A startup out of Tel Aviv wants to change the word to memory. Majestic Labs, founded in 2023 by former Google and Meta engineers, has unveiled a server it says can do the work of a rack of Nvidia GPUs. It does so by attacking a different bottleneck. The pitch, reported by TechRadar, is that pairing pricey GPUs with scarce high-bandwidth memory has become a dead end for AI inference. Running a model is often limited by how much fast memory you can reach, not raw compute. So Majestic ditched the GPU. Its server, Prometheus, swaps graphics chips for what it calls Ignite AI Processing Units. Each unit blends Arm cores with RISC-V vector and tensor engines. Up to 12 sit in one server, sharing a single pool of 8TB to 128TB of LPDDR6. That is the cheap memory found in phones, not the costly high-bandwidth memory that GPUs depend on. A different way to hit the memory wall The trick is the wiring. Instead of bolting memory onto each GPU package, Majestic pools it through custom aggregation chiplets. Copper cables up to a metre long connect them. The result, it claims, is one coherent pool far larger than a GPU box can address. The comparison it reaches for is stark. An Nvidia DGX B300 with eight Blackwell GPUs carries 2.3TB of high-bandwidth memory. Majestic says Prometheus offers more than 50 times as much fast memory, at 1.7 times the interconnect bandwidth. One rack, it claims, matches 25 of Nvidia's Vera Rubin racks for fast memory, at a fraction of the power. The software story is friendlier than the hardware sounds. Prometheus is built to open standards and supports PyTorch, vLLM and OpenAI's Triton. Models built for GPUs are meant to run on it without changes. Majestic says it has already taken orders from large enterprises, neoclouds and hyperscalers. Striking numbers, no shipped hardware Every figure here carries the same asterisk. It is Majestic's own, ahead of independent testing, and nothing has shipped. The company has about 40 staff across Tel Aviv and Los Angeles. It raised $100m late last year, a modest sum against what its rivals command. There are physical questions too. A 128TB pool built from 2GB LPDDR6 dies would need roughly 64,000 of them. That implies more than a hundred aggregation chiplets in a single server. It is a lot of parts to keep coherent, and TechRadar notes buyers will likely wait for independent benchmarks before switching. What Majestic joins is a growing line of startups attacking Nvidia from odd angles. Optical chips, edge silicon for inference and open networking gear each pick a different weak point. The common thread is that Nvidia's dominance is now a target from several sides at once, even its software moat. The memory-wall pitch is the newest of them. It is also the least proven. Majestic has drawn a picture of a rack that shames a room of GPUs on memory and power. Next year, when hardware ships and someone else runs the benchmarks, the picture either holds or it does not.
[2]
Majestic Labs unveils Prometheus: GPU-free server for AI model serving
Majestic Labs has pulled the wraps off Prometheus, a new AI inference server platform that leans on shared memory instead of the usual Nvidia-style, GPU-heavy setup. The company says select customers are slated to get shipments in 2027. The startup was founded in Tel Aviv in 2023 by former Google and Meta engineers Ofer Shacham, Sha Rabii, and Masumi Reynders. Majestic Labs says it has raised $100 million. Its pitch for Prometheus centers on the AI "memory wall": up to 12 Ignite processors, each combining Arm cores with RISC-V vector and tensor engines, all tied into a shared pool of 8 TB to 128 TB of unified LPDDR6 memory. The other big part of the argument is cost and power. Majestic Labs says Prometheus skips the expensive, hard-to-get HBM route and uses LPDDR6 with custom aggregation chiplets and copper links up to 1 metre long, creating a single coherent fast-memory pool. According to the company, that could bring costs down by 10 to 50 times while also cutting power use. If you're running large-model inference, this is a platform to keep an eye on. Still, some caution is warranted. Majestic Labs says Prometheus offers more than 50 times the fast-memory capacity of Nvidia's 2.3 TB DGX B300, about 1.7 times the interconnect bandwidth, and that one rack can match 25 Nvidia Vera Rubin racks. None of that has been independently verified yet, and buyers are going to want real proof across PyTorch, vLLM, and Triton, along with evidence that the software is mature and that it fits into existing MLOps workflows. For now, Prometheus isn't something you can download, and Majestic Labs says shipments to select customers are planned for 2027.
Share
Copy Link
Tel Aviv startup Majestic Labs has unveiled Prometheus, a GPU-free server that uses up to 128TB of LPDDR6 memory and Arm cores to tackle AI's memory wall. The company claims one rack matches 25 Nvidia Vera Rubin racks at lower power, but hardware hasn't shipped or been independently tested yet.
Majestic Labs, a Tel Aviv startup founded in 2023 by former Google and Meta engineers Ofer Shacham, Sha Rabii, and Masumi Reynders, has unveiled Prometheus, an AI server that abandons GPUs entirely
1
2
. The company's bold pitch targets what it calls the AI memory wall—the bottleneck created when running large AI models is limited by available fast memory rather than raw compute power. With $100 million in funding raised late last year, Majestic Labs now has about 40 staff across Tel Aviv and Los Angeles working to challenge Nvidia's dominance from an unconventional angle1
2
.Prometheus represents a fundamental departure from conventional AI inference server platform design. Instead of pairing expensive GPUs with scarce high-bandwidth memory, the AI server uses what Majestic calls Ignite AI Processing Units
1
. Each Ignite unit combines Arm cores with RISC-V vector and tensor engines. Up to 12 of these units sit in a single server, all connected to one shared pool of 8TB to 128TB of LPDDR6 memory—the same affordable memory found in smartphones, not the costly high-bandwidth memory that GPUs depend on1
2
.The shared memory architecture relies on custom aggregation chiplets connected through copper cables up to one metre long, creating what Majestic claims is a single coherent fast-memory pool far larger than any GPU box can address
1
. Building a 128TB pool from 2GB LPDDR6 dies would require roughly 64,000 of them, implying more than a hundred aggregation chiplets working in concert within a single server1
.Majestic Labs positions Prometheus as a direct challenger to Nvidia's flagship systems. An Nvidia DGX B300 with eight Blackwell GPUs carries 2.3TB of high-bandwidth memory. Prometheus offers more than 50 times as much fast memory capacity, according to the company, along with 1.7 times the interconnect bandwidth
1
2
. The comparison extends to rack-level deployments: Majestic claims one Prometheus rack matches 25 of Nvidia's Vera Rubin racks for fast memory, at a fraction of the power consumption1
2
.On cost and power efficiency, the startup says Prometheus could reduce expenses by 10 to 50 times while also cutting power use significantly
2
. For AI model serving workloads where inference is often memory-bound rather than compute-bound, these figures suggest a meaningful shift in economics.Related Stories
Despite the radical hardware departure, Majestic Labs has built Prometheus to open standards. The platform supports PyTorch, vLLM, and OpenAI's Triton, and models developed for GPUs are designed to run on Prometheus without modifications
1
2
. This compatibility matters for enterprises already invested in GPU-based workflows. The company reports it has already secured orders from large enterprises, neoclouds, and hyperscalers, though specific customers haven't been named1
.Buyers will need evidence that the software stack is mature and fits seamlessly into existing MLOps workflows before committing at scale
2
. Select customers are slated to receive shipments in 2027, giving the company roughly two years to prove the platform in production environments2
.Every performance figure Majestic Labs has shared carries the same caveat: none has been independently verified, and no hardware has shipped
1
2
. Keeping 64,000 memory dies coherent across more than a hundred aggregation chiplets presents significant engineering challenges. Buyers are likely to wait for independent benchmarks before switching infrastructure1
.Majestic Labs joins a growing line of startups attacking Nvidia from unconventional angles—optical chips, edge silicon for inference, open networking gear. Each picks a different weakness in Nvidia's architecture. What makes Prometheus distinct is its focus on the AI memory wall, betting that inference workloads are increasingly starved for accessible memory rather than compute cycles
1
. Whether this GPU-free server can deliver on its bold claims will become clear when hardware reaches customers and third-party testing begins in earnest over the next two years.Summarized by
Navi
[1]
10 Nov 2025•Startups

04 Dec 2025•Technology

01 Jun 2026•Technology

1
Technology

2
Science and Research

3
Technology
