Health & ScienceHealth & Science 5 min read

Stanford TranscriptFormer Maps Cells Across 12 Species

Stanford Medicine highlighted TranscriptFormer on 14 September 2026 via a EurekAlert news release describing research articles in Nature and Science, with manuscript timing notes pointing to 8 July

PC

PromptCrates Editorial

Staff Writer

0 0
Stanford TranscriptFormer Maps Cells Across 12 Species

Stanford Medicine highlighted TranscriptFormer on 14 September 2026 via a EurekAlert news release describing research articles in Nature and Science, with manuscript timing notes pointing to 8 July 2026 for media contact context. The project pairs a universal cell embedding with TranscriptFormer, trained on 112 million cells across twelve species—human, yeast, mouse, rabbit, chicken, zebrafish, fruit fly, African clawed frog, malaria parasite, sea urchin, sponge, and C. elegans. Co-leads include bioengineer Stephen Quake, computer scientist Jure Leskovec, and Theo Karaletsos of the Chan Zuckerberg Initiative. The models are trained like large language models, but on gene-expression values, to create a shared mathematical space for comparing cells across organisms.

Why a universal cell space matters for biology

Single-cell atlases have exploded in volume while remaining fragmented by species, assay, and lab convention. A model that embeds cells from sponges and humans into one comparable space is an attempt to make evolutionary and disease questions askable with the same geometry. Instead of treating each organism’s atlas as a silo, TranscriptFormer aims to let researchers ask whether a newly measured cell type looks more like a neuron, a digestive gland cell, or something else—even when the species was not in prior textbooks for that comparison.

The sponge findings make the abstract claim concrete. Choanocytes from sponges resembled neurons from roundworms and frogs in the embedding space, while sponge neuroid cells sat closer to frog gland cells associated with digestion than to neurons. That split challenges casual assumptions that any “neuro-like” sponge cell must sit near animal neurons. It also shows why a cross-species foundation model is more than a clustering convenience: it can surface homology hypotheses that specialists can then test with wet-lab methods.

Co-first Nature authors Yanay Rosen, a Stanford computer science PhD student, and Yusuf Roohani of the Arc Institute underscore the collaboration pattern: CS methods meeting biomedical questions under CZI, DARPA, NSF, and other funding. Health-science readers can place this work beside our coverage of DeepMind AlphaGenome’s DNA variant atlas and regulatory threads such as the FDA classification of cardiovascular ML notification.

From unseen species to diseased versus healthy cells

Beyond evolutionary mapping, the team highlights identification of cell types in species the model has not seen and discrimination between healthy and diseased cells. Those capabilities matter for translational pipelines where rare samples, non-model organisms, or patient biopsies arrive faster than curated labels. A shared embedding that generalizes can propose labels for human review rather than forcing every new dataset through a custom classifier from scratch.

The longer-horizon claim is more speculative and more interesting: a path toward designing functional cells that do not yet exist. That language sits at the boundary of regenerative medicine and generative biology. It does not mean clinicians can order custom cell types from a prompt tomorrow. It does mean foundation models over expression space are being framed as design tools, not only as annotation tools—an ambition that will attract both funding and safety scrutiny as capabilities grow.

Medical AI governance already wrestles with how far generative systems should go in clinical and device contexts. Our reporting on the ARPA-H ADVOCATE cardiovascular AI program and the FDA generative AI medical device discussion paper shows regulators and research agencies preparing frameworks while labs publish new modalities. TranscriptFormer is a research model for cells, not a cleared diagnostic, but its “design cells that do not exist” framing will inevitably enter those conversations.

Limits editors should keep in the frame

Medical imaging and genomics already taught hospitals that foundation models travel faster than clinical governance. Cell-expression models will face a similar gap: researchers may share embeddings long before hospitals trust generative cell design claims. That lag is healthy if it forces validation studies; it is harmful if product marketing outruns the papers. For now, TranscriptFormer belongs in the research toolkit conversation—next to atlas projects and variant models—rather than in bedside decision support.

Cross-species embeddings can hallucinate similarities that reflect training composition rather than biology. Independent labs will need to probe failure modes: batch effects, assay shifts, and whether disease separation holds outside the authors’ cohorts. Publication in Nature and Science raises the evidentiary bar, yet media releases compress years of methods into a few metaphors. Readers should treat the sponge neuron and gland mappings as published scientific claims to verify in the papers, not as settled taxonomy.

Documented facts for this draft follow the 14 September Stanford Medicine / EurekAlert package. TranscriptFormer and a preceding universal cell embedding cover 112 million cells across twelve named species; training mimics LLM methods on expression values; Quake, Leskovec, and Karaletsos co-lead; Rosen and Roohani are Nature co-first authors; sponge choanocytes and neuroid cells show distinct nearest neighbors; and applications include unseen-species cell typing, healthy-versus-diseased separation, and aspirational cell design.

Primary source: EurekAlert Stanford Medicine TranscriptFormer release.

health-scienceStanfordcell biologyfoundation models

Related articles