
Causal Models Need Causal Data - Xaira’s X-Cell model (Bo Wang & Ci Chu)
Latent Space21 July 2026Watch on YouTube
Description
Can a model predict how a cell responds to a genetic perturbation it has never seen? Xaira Therapeutics' new virtual-cell model, X-Cell, is a 4.9-billion-parameter diffusion language model trained on X-Atlas/Pisces — the largest genome-wide CRISPRi Perturb-seq dataset ever built, spanning 25.6 million single cells across 16 biological contexts. Bo Wang (Chief AI Scientist) and Ci Chu (Chief Discovery Officer) explain why observational atlases can describe biology but can't predict what happens when you intervene, why they abandoned autoregression for a diffusion "editing" approach, and how a model trained on immortalized cell lines predicted perturbation responses in primary T cells from real donors. Plus the counterintuitive scaling result: X-Cell scales like an LLM on training loss, but generalization is bottlenecked by data diversity, not compute — a conversation about why, in AI for science, the hard part isn't the model, it's the data. Bios Bo Wang is Chief AI Scientist at Xaira Therapeutics and a co-senior author of X-Cell. He is also an Associate Professor at the University of Toronto, a CIFAR AI Chair at the Vector Institute, and Chief AI Scientist at University Health Network, where his lab pioneered scGPT — one of the first single-cell foundation models — and BioReason. His work centers on foundation models that learn the underlying biology of the cell; X-Cell is initialized from scGPT's own encoder weights. Ci Chu is Chief Discovery Officer at Xaira Therapeutics and a co-senior author of X-Cell. He leads the high-throughput biology behind Xaira's data-generation engine — including the industrialized Perturb-seq platform that produced X-Atlas/Pisces — and previously led functional genomics and high-content phenotyping work at insitro. His throughline is pairing large-scale, high-quality experimental data with the most capable models. Timestamps 00:00:00 Intro 00:02:07 Guest intros & Xaira's three-platform overview 00:07:12 Why biology lags behind protein design — the data bottleneck 00:13:07 What X-Cell does — perturbation prediction explained 00:16:05 History of virtual cell modeling — from differential equations to scGPT 00:22:42 Perturb-seq at scale — building the PISCES dataset 00:36:32 Spatial transcriptomics and future modalities 00:44:00 X-Cell architecture — why diffusion beats autoregression for gene expression 00:54:25 Ablation results — what actually moves the needle 00:57:29 X-Cell generalization — beating linear baselines in unseen cell types 01:07:10 Single vs. combinatorial perturbations and platform expansion 01:16:07 Academia vs. industry in the agentic AI era 01:25:14 Open science and what academic labs should focus on 01:32:20 Magic wand questions: protein measurement and temporal sequencing "This is the first time that someone can put together not just one, but seven genome-wide Perturb-seq campaigns together." (0:26) "We are nowhere near the same kind of massive, high-quality data [in cell modeling] that we have in protein design... it's a data limitation issue." (0:9:46) "Instead of typing, think of diffusion language models as editing: you iteratively generate a sentence from a very vague, rough start and then refine it." (0:43:11) "I tell my students: in the era of agentic AI, the role of the scientist is shifting from coding to debugging the AI's outputs." (1:23:07) "You can't have the cake and eat it—currently, to sequence a cell, you have to kill it. The real 'magic wand' for virtual cells would be temporal sequencing of live cells." (1:28:43)
What you'll learn
- X-Cell is a 4.9-billion-parameter diffusion model that predicts how cells respond to genetic perturbations it has never encountered.
- The model was trained on X-Atlas/Pisces, the largest genome-wide CRISPRi Perturb-seq dataset containing 25.6 million cells across 16 biological contexts.
- In cell modeling, data diversity is a bigger bottleneck than computational power, unlike the trend in other AI domains.
- X-Cell uses a diffusion approach instead of autoregression, where you iteratively refine from a rough initial state toward gene expression predictions.
- The model can predict perturbation responses in primary T cells from real donors, despite being trained on immortalized cell lines.