All videos
0:00 / 0:00
research

Causal Models Need Causal Data - Xaira’s X-Cell model (Bo Wang & Ci Chu)

Latent Space21 July 2026Watch on YouTube

Description

Can a model predict how a cell responds to a genetic perturbation it has never seen? Xaira Therapeutics' new virtual-cell model, X-Cell, is a 4.9-billion-parameter diffusion language model trained on X-Atlas/Pisces — the largest genome-wide CRISPRi Perturb-seq dataset ever built, spanning 25.6 million single cells across 16 biological contexts. Bo Wang (Chief AI Scientist) and Ci Chu (Chief Discovery Officer) explain why observational atlases can describe biology but can't predict what happens when you intervene, why they abandoned autoregression for a diffusion "editing" approach, and how a model trained on immortalized cell lines predicted perturbation responses in primary T cells from real donors. Plus the counterintuitive scaling result: X-Cell scales like an LLM on training loss, but generalization is bottlenecked by data diversity, not compute — a conversation about why, in AI for science, the hard part isn't the model, it's the data. Bios Bo Wang is Chief AI Scientist at Xaira Therapeutics and a co-senior author of X-Cell. He is also an Associate Professor at the University of Toronto, a CIFAR AI Chair at the Vector Institute, and Chief AI Scientist at University Health Network, where his lab pioneered scGPT — one of the first single-cell foundation models — and BioReason. His work centers on foundation models that learn the underlying biology of the cell; X-Cell is initialized from scGPT's own encoder weights. Ci Chu is Chief Discovery Officer at Xaira Therapeutics and a co-senior author of X-Cell. He leads the high-throughput biology behind Xaira's data-generation engine — including the industrialized Perturb-seq platform that produced X-Atlas/Pisces — and previously led functional genomics and high-content phenotyping work at insitro. His throughline is pairing large-scale, high-quality experimental data with the most capable models. Timestamps 00:00:00 Intro 00:02:07 Guest intros & Xaira's three-platform overview 00:07:12 Why biology lags behind protein design — the data bottleneck 00:13:07 What X-Cell does — perturbation prediction explained 00:16:05 History of virtual cell modeling — from differential equations to scGPT 00:22:42 Perturb-seq at scale — building the PISCES dataset 00:36:32 Spatial transcriptomics and future modalities 00:44:00 X-Cell architecture — why diffusion beats autoregression for gene expression 00:54:25 Ablation results — what actually moves the needle 00:57:29 X-Cell generalization — beating linear baselines in unseen cell types 01:07:10 Single vs. combinatorial perturbations and platform expansion 01:16:07 Academia vs. industry in the agentic AI era 01:25:14 Open science and what academic labs should focus on 01:32:20 Magic wand questions: protein measurement and temporal sequencing "This is the first time that someone can put together not just one, but seven genome-wide Perturb-seq campaigns together." (0:26) "We are nowhere near the same kind of massive, high-quality data [in cell modeling] that we have in protein design... it's a data limitation issue." (0:9:46) "Instead of typing, think of diffusion language models as editing: you iteratively generate a sentence from a very vague, rough start and then refine it." (0:43:11) "I tell my students: in the era of agentic AI, the role of the scientist is shifting from coding to debugging the AI's outputs." (1:23:07) "You can't have the cake and eat it—currently, to sequence a cell, you have to kill it. The real 'magic wand' for virtual cells would be temporal sequencing of live cells." (1:28:43)

What you'll learn

  • X-Cell is a 4.9-billion-parameter diffusion model that predicts how cells respond to genetic perturbations it has never encountered.
  • The model was trained on X-Atlas/Pisces, the largest genome-wide CRISPRi Perturb-seq dataset containing 25.6 million cells across 16 biological contexts.
  • In cell modeling, data diversity is a bigger bottleneck than computational power, unlike the trend in other AI domains.
  • X-Cell uses a diffusion approach instead of autoregression, where you iteratively refine from a rough initial state toward gene expression predictions.
  • The model can predict perturbation responses in primary T cells from real donors, despite being trained on immortalized cell lines.

Frequently asked questions

What is the difference between observational atlases and predictive models in cell biology?
Observational atlases can describe what biology looks like but cannot predict what happens when you intervene, such as with genetic perturbations. X-Cell adds predictive capacity by being trained on Perturb-seq data, where cells are actually perturbed.
Why did Xaira choose diffusion over autoregression for X-Cell?
With diffusion, you think of 'editing' rather than 'typing': you start from a vague, rough state and iteratively refine it. This approach works better than autoregression for modeling gene expression in cellular response.
What is X-Atlas/Pisces and why is it important for X-Cell?
X-Atlas/Pisces is the largest genome-wide CRISPRi Perturb-seq dataset ever built, containing 25.6 million cells from 16 biological contexts. It serves as training data for X-Cell and represents the combination of seven genome-wide Perturb-seq campaigns.
How does X-Cell generalize to cell types not present in the training data?
X-Cell successfully predicts perturbation responses in primary T cells from real donors, despite being trained on immortalized cell lines, demonstrating that sufficient data diversity in the training set enables strong generalization.

Topics