What we
work on
Biology has rules, but most of them have never been written down. We build machine learning systems that learn those rules from data, then test what they predict. The lab grew out of manifold learning for single-cell data (MAGIC, PHATE, MELD). Today its center is foundation models that treat biology as a language.
Biology as a language
We turn cells, tissues, and genomes into text so that language models can read, write, and reason about them.
A cell’s gene expression, ranked from most to least expressed, becomes a sentence. Once biology is written this way, a language model can annotate cells, generate new ones, answer questions about them, and connect them to everything written about biology in natural language. Cell2Sentence introduced the idea. C2S-Scale took it to 27 billion parameters with Google. We are now extending it from single cells to tissues and patients, and to new modalities such as DNA methylation.
Virtual cells & causal perturbation
Predicting what a drug or a gene knockout will do to a cell before anyone runs the experiment.
Most biological questions are causal: what happens if we intervene? We build methods that separate true treatment effects from confounding variation (CINEMA-OT, MELD) and models that simulate perturbations in silico. C2S-Scale’s virtual screen found that silmitasertib amplifies antigen presentation only in an immune-active context, a prediction that was then confirmed in human cells. We are working out when such predictions transfer across cell types, donors, and diseases.
Foundation models of the brain
Models trained on thousands of hours of brain activity that predict clinical traits and simulate neural dynamics.
BrainLM was the first foundation model for fMRI. It was trained on 6,700 hours of recordings and can decode clinical variables and simulate how activity changes under different conditions. With the Wu Tsai Institute, we are building the next generation of neuroscience foundation models and an AI neuroscientist that runs analyses through conversation. Related work models behavior and psychiatric biomarkers.
AI agents for science
Teams of language-model agents that plan analyses, pull data, generate hypotheses, and design experiments.
Foundation models become far more useful when they can act. We build multi-agent systems that divide scientific work into roles and subtasks (MoRSE), agents that harmonize thousands of public datasets for aging research, and benchmarks that measure how well models handle scientific inverse design. The goal is to shorten the path from question to validated result.
Foundations of generative models & intelligence
New generative models for sequences, sets, and continuous dynamics, and a theory of where reasoning comes from.
Biology stretches our models, so we build new ones. Our work includes discrete diffusion joined with causal language models (CaDDi), diffusion over permutations, insertion-based generation, and neural operators that learn integral and differential equations from data. We also ask what makes a system intelligent. Intelligence at the Edge of Chaos showed that models trained on data at the boundary between order and chaos become better reasoners.
Software & models
Everything is on GitHub and Hugging Face.
- Cell2SentencePython library to train and use cell-sentence LLMs↗
- C2S-Scale-Gemma-2-27B27B single-cell LLM, with Google↗
- C2S-Scale-Gemma-2-2BCompact C2S model for a single GPU↗
- CINEMA-OTCausal inference for single-cell perturbations↗
- BrainLMFoundation model for fMRI recordings↗
- MAGICData diffusion for single-cell imputation↗