Stanford Medicine researchers developed TranscriptFormer, an AI model trained on 112 million cells from 12 species that maps gene expression patterns across organisms. The model can identify cell types in new species, distinguish healthy from diseased cells, and may eventually guide designing new cell-based treatments for complex diseases.

AI models transform how scientists study cells across species

Stanford Medicine researchers have built two groundbreaking AI models that are reshaping cell biology research. The first model, universal cell embedding, laid the foundation for TranscriptFormer, a second-generation system trained on data from 112 million cells spanning 12 species—from single-celled yeast to humans

1

2

. This cross-species mapping of cell biology opens possibilities for understanding diseases and developing new treatments that were previously impossible.

Source: Phys.org

Source: Phys.org

Stephen Quake, professor of bioengineering at Stanford Medicine and co-lead author alongside Jure Leskovec, professor of computer science, and Theo Karaletsos of the Chan Zuckerberg Initiative, describes the work as "the beginning of, we hope, a whole new field and a decade of work." The research appears in two papers published in Nature and Science

1

3

.

How TranscriptFormer learns gene expression patterns like language

While genomes contain life's instructions, modern biologists focus on gene expression—which genes each cell actively uses. A pancreatic beta cell expresses insulin genes, B cells produce antibody genes, and skin cells activate genes for hair or pigment. These gene expression patterns distinguish cell types and reveal differences between healthy vs. diseased cells across species

2

.

Recent years have seen scientists build atlases—databases defining cells based on gene expression patterns. But the data on tens of thousands of genes across hundreds of millions of cells overwhelms human analysis. "How do we think about all these genes at once? The human brain can't," Quake explains. "These models help us do that. The algorithms are ways for us to understand very complex data that humans can't wrap our minds around"

1

.

Source: News-Medical

Source: News-Medical

TranscriptFormer's training mirrors large language models like ChatGPT. LLMs learn by predicting missing words in text; TranscriptFormer predicts missing gene expression values. The team trained it on cell atlases from mice, rabbits, chickens, zebrafish, fruit flies, African clawed frogs, malaria parasites, sea urchins, sponges, and the roundworm Caenorhabditis elegans

2

3

.

Universal space reveals evolutionary relationships and cell functions

Quake describes TranscriptFormer as creating a "universal space"—a mathematical framework where cells from every organism on Earth can be positioned and compared. This enables researchers to explore evolutionary relationships that were previously speculative

1

.

Sponge researchers wondered which sponge cell most resembles neurons in other animals. TranscriptFormer revealed that choanocytes share striking similarity with neurons in roundworms and frogs, illuminating neuron evolution and choanocyte function. Meanwhile, the sponge's neuroid cell—named for neuron-like features—actually showed gene expression patterns matching frog glands, suggesting a digestive rather than neural role

2

3

.

Practical applications from disease detection to treatment design

When given gene expression data from species outside its training set, TranscriptFormer can identify cell types, demonstrating genuine understanding rather than memorization. The team confirmed it distinguishes healthy from diseased cells, a capability crucial for medical applications

2

3

.

The model's most ambitious potential lies in designing new cell-based treatments. Because it understands what makes cells functional, TranscriptFormer could eventually imagine and guide creation of therapeutic cells that don't currently exist. "We hope it's going to be a very powerful tool for discovery," Quake notes. "This is a tool to help skilled biologists think about complex problems"

3

.

Yanay Rosen, a doctoral student in computer science at Stanford University, and Yusuf Roohani, now an associate director of machine learning at the Arc Institute, served as co-first authors on the Nature paper. The Chan Zuckerberg Initiative, BioHub, Defense Advanced Research Projects Agency, National Science Foundation, and multiple corporate partners funded the research

3

.

Today's Top Stories

© 2026 TheOutpost.AI All rights reserved