Infinity raises $15M to run AI inference on any chip, challenging Nvidia's software dominance

3 Sources

Share

San Francisco-based AI infrastructure startup Infinity has raised $15 million at a $100 million valuation to break Nvidia's software lock-in. Its AI agent Ignition automates the creation of low-level code needed to run AI models on any chip architecture, reducing what typically takes months or years to just hours or days. The company already generates millions in annual revenue from chip partnerships.

Infinity Secures Funding to Challenge Nvidia's AI Chip Dominance

AI infrastructure startup Infinity announced a $15 million seed round at a $100 million valuation on Monday, backed by Touring Capital, Principal VC, and researchers from OpenAI and Anthropic

1

2

. The San Francisco-based company is building software to make it easier for AI chips to run AI inference workloads, directly targeting the software moat that has kept Nvidia untouchable in the AI accelerator market.

The startup is already generating millions of dollars in annual recurring revenue from chip design partnerships and employs 26 people

3

. Infinity will use the capital to expand its engineering team, scale its automated research platform, and accelerate partnerships with chipmakers including d-Matrix.

Breaking the CUDA Lock-In With Automated Code Generation

Source: SiliconANGLE

Source: SiliconANGLE

Nvidia's dominance rests on more than hardware performance. Its real advantage is Nvidia CUDA, the software layer built over nearly two decades that allows developers to write applications in Python and have them run on Nvidia hardware by default

2

. The major AI frameworks PyTorch and TensorFlow sit on top of CUDA, giving Nvidia an estimated 80% share of the data-center AI accelerator market.

Rival chips from AMD, Qualcomm, and AWS often match Nvidia on raw power but lack the mature software ecosystem. Most startups wouldn't have the resources or expertise to write their own kernels—the low-level software that operates chips—and port their applications to non-Nvidia hardware

1

. This is where Infinity aims to make any AI chip inference-ready through automation.

How Ignition Automates Universal Inference Across Chip Architectures

Infinity's core technology is Ignition, an AI agent that generates, tests, and optimizes the low-level code generation needed to run models efficiently on different processors

3

. The system writes kernels, builds debuggers and profilers, and orchestrates the entire inference pipeline while continuously learning and improving itself. Human engineers provide high-level direction while the agent handles the tedious implementation work.

Ignition adapts to different chip architectures regardless of proprietary designs, effectively creating a CUDA-level software stack in hours or days rather than months or years

1

. To work with proprietary instruction sets, Infinity builds decompilers that translate executable code back into high-level source code that can be manipulated using standard programming languages

3

.

Impressive Performance Claims From Early Partnerships

The results Infinity reports are striking, though they come from the company itself and have not yet been independently validated. Working with chip maker d-Matrix, Ignition achieved 92% of the chip's theoretical peak performance within 10 hours of first touching the hardware

2

. Within 10 days, three frontier models—Qwen3, Qwen3.5, and Gemma4—were running end-to-end on the d-Matrix Corsair accelerator

3

.

In another test, Infinity increased inference throughput for the Qwen3-8B model by 34% compared with vLLM, a widely used open-source inference framework, after just one day of automated optimization on a single Nvidia H100 GPU

3

. In one case study, the company lifted a model's output from approximately 1,400 to more than 20,000 tokens per second in a single day

2

.

The Vision Behind Automated Invention

Source: TechCrunch

Source: TechCrunch

Jeremy Nixon, founder and CEO of Infinity, is a former Google Brain researcher who created AGI House, a San Francisco hacker network that has spawned hundreds of startups

2

. Nixon told TechCrunch he launched the company because he was obsessed with "automated invention"—the belief that AI systems can function as a meta-technology

1

.

He previously invented a machine learning algorithm called Omega that created new machine learning algorithms and automatically evaluated them in a feedback loop. That success led him to apply the same approach to hardware, believing automated systems could generate the low-level code needed to help run chips more effectively

1

. "The next era of AI will be defined not just by who makes the best chip, but by who can make any chip run state-of-the-art models at blazing speeds," Nixon said

2

.

Performance-Based Revenue Model Aligns Incentives

Unlike traditional software licensing, Infinity operates on a performance-based revenue model that takes a cut of the speed and cost gains it delivers rather than charging upfront fees

2

. Optimization agreements generally give Infinity approximately 20% of the savings on additional computing purchases, calculated from gains in AI model performance against an agreed baseline

3

. This approach directly links the company's compensation to the value it creates, measuring changes in tokens per second

1

.

What This Means for the AI Hardware Landscape

As inference races to become two-thirds of all AI compute spending this year, cheaper and more flexible ways to run inference workloads are worth a fortune

2

. Infinity joins a wave of startups attempting to break CUDA lock-in and chip away at Nvidia's market dominance product by product

1

.

The company is in talks with other major chip and cloud companies beyond its current partnership with d-Matrix

1

. If Ignition delivers on its promise of universal inference across chip architectures, it could lower barriers for new hardware entrants and give enterprises more flexibility in their AI infrastructure choices. The caveats remain real—this is a seed-stage firm with one public chip partner and self-reported benchmarks—but the potential impact on the AI hardware ecosystem is substantial

2

.

Today's Top Stories

© 2026 TheOutpost.AI All rights reserved