Gimlet Labs Raises $300M to Solve AI's Growing Power Crisis with Multi-Silicon Inference Cloud

2 Sources

Share

Gimlet Labs secured $300 million in Series B funding led by Andreessen Horowitz at a $3 billion valuation. The startup addresses AI inference bottlenecks with its disaggregated inference platform that orchestrates heterogeneous compute systems, delivering up to 10X throughput gains while optimizing intelligence per watt across GPUs and specialized accelerators.

Gimlet Labs Secures $300 Million to Address AI Inference Capacity Crunch

Gimlet Labs has raised $300 million in Series B funding at a $3 billion valuation, led by Andreessen Horowitz with participation from Arm Holdings, Samsung Ventures, Microsoft's M12 fund, and over a dozen other investors

2

. The round brings the company's total outside funding to $392 million

2

. This significant investment comes as AI inference faces unprecedented resource constraints, with GPU capacity leasing at record prices and providers like OpenAI and Anthropic constrained by compute fleet expansion rates

1

.

Source: Andreessen Horowitz

Source: Andreessen Horowitz

The Infrastructure Crisis Driving Innovation in AI Inference

AI inference has become one of the fastest-growing markets in capitalism's history, yet the industry is running critically short of nearly every physical input required to serve it: powered land, turbines, transformers, data centers, GPUs, advanced-node wafers, and high-bandwidth memory

1

. Five U.S. hyperscalers alone are expected to spend $1 trillion in capex next year, while NVIDIA has mobilized over $500 billion with capital partners for AI factories

1

. Users experience this scarcity directly through tightening rate limits and slower responses as providers sacrifice latency for throughput to extract more tokens from the same watts

1

.

Multi-Silicon Inference Cloud Delivers 10X Throughput Gains

Gimlet Labs has built what Andreessen Horowitz calls the first multi-silicon inference cloud, designed to produce more intelligence per watt

1

. The disaggregated inference platform automatically breaks up large language models into constituent modules and deploys each component on the chip architecture that best aligns with its hardware requirements

2

. For each workload, Gimlet determines an execution plan balancing latency, throughput, and cost, routing different models and tools onto different processors or separating prefill and decode phases

1

. The system already delivers up to 10X gains in throughput and interactivity on frontier models within the same power envelope

1

.

Heterogeneous Compute Systems Replace One-Size-Fits-All Architecture

The diversity of AI applications has exploded beyond what single-architecture systems can efficiently handle. A voice assistant prioritizes latency, batch processing maximizes throughput, research agents optimize cost per token, and coding agents must balance all three

1

. Even a single model call splits into compute-intensive prefill and memory-bound decode operations

1

. Gimlet's platform supports multiple disaggregation approaches, including the common PD disaggregation that runs prefill and decode phases on separate chips, as well as more granular methods where developers use lightweight drafter models alongside frontier LLMs for refinement

2

.

AI Agents and Custom Compiler Optimize Inference Optimization

Gimlet reduces the complexity of implementing disaggregation workflows through AI agents and a custom compiler that optimize each LLM module for its target chip architecture

2

. The AI agents explore multiple design approaches to find the best way of adapting LLM code to specific chips, then run tests to ensure correctness

2

. The compiler applies both generic and chip-specific optimizations to customer models

2

. This orchestration extends into physical data centers, integrating GPUs, CPUs, and purpose-built accelerators into one capacity pool while managing differences in networking, power, and cooling requirements

1

.

Source: SiliconANGLE

Source: SiliconANGLE

Frontier Labs and Hyperscalers Already Driving Billions in Orders

Gimlet has achieved scale that matters in AI inference infrastructure, counting both a frontier lab and a hyperscaler among its customers

1

. The company stated in March that its customer base includes one of the world's largest cloud providers and a top three AI lab

2

. Gimlet has received billions of dollars worth of customer orders

2

. The platform is available as a serverless service and as a managed service that enterprises can deploy on their own AI inference infrastructure

2

.

Expansion Plans Target Custom Hardware and Serverless Capacity

The Series B funding will help Gimlet grow its serverless edition's infrastructure capacity by adding several hundred megawatts of computing power

2

. The company is also expanding into custom hardware development with an inference-optimized server design that eliminates the traditional motherboard and can operate outside conventional data centers

2

. This move addresses the physical constraints of large language model inference workloads as inference demand compounds at software speed while power plants, data centers, and semiconductor fabs cannot keep pace

1

. Watch for Gimlet's motherboard-free servers to enable deployment in edge locations where traditional data center infrastructure proves impractical, potentially unlocking new use cases for latency-sensitive applications.

Today's Top Stories

© 2026 TheOutpost.AI All rights reserved