Transcript
Intro
0:00 · And what really blew my mind away is when I saw the model make prediction just print out the heat map of the generation changes look at the actual raw data and line up the linear baseline prediction the ground truth and XL prediction all together it's visually very clear to see that XL prediction is much more [music] similar to ground truth than the linear baseline this is a wow moment I was talking about in the beginning this is the first time that someone [music] can put together not just one prosy but seven genome web perturbsy campaigns together. Something that uh jumped out to us biologists right away [music] is that some of the partations are context universal.
0:41 · Hi, I'm R.J. Haniki, CTO of Mirror OMIX.
0:44 · This is Brandon Anderson who builds RNA therapeutics at Atomic AI and this is the latent space AI for science podcast.
0:53 · One of the themes that has run through the podcast is how the lab and experimentation and the real world have probably the biggest impact and have the most relevance to whether uh something is AI for science or something like B2B SAS. Uh we're really happy to have in the studio with us today Bo Wang and Cichu from Zera Therapeutics. At Zera, they're building with a bunch of other people a AI drug discovery platform.
1:27 · They're using high throughput experimentation system to collect very large data sets and then training AI models that can predict the way that your cells in your body will respond to drugs and therapeutics. Really happy to have you. Big fan of your work. Um why don't you two introduce yourselves to the listeners?
1:50 · Hello everyone, my name is Bowen. I'm SVP and head of biomedical AI at Zera Therapeutic. Joined Zara about eight months ago and before that I was associate professor at the University of Toronto in Canada. And um I'm such uh my first name is incredibly difficult to pronounce unless you speak Mandarin. So I go by Chu uh as in Chewbacca or Pikachu. I think [laughter] your favorite fictional character. I'm the SVP of AI enabled discovery at Zera. I joined about more than two years ago when it was still in stealth mode. Uh and here I lead the high throughput uh biology group generating the kind of data that will feed our AI models and also think about their applications. Um before this I uh spent about a decade in uh at the intersection of AI and uh uh big data and biology. uh as I previously I worked at in Citro leading uh the invitual discovery platform there and before that I was at Verily which spun out of Google X.
Guest intros & Xaira's three-platform overview
2:48 · Okay.
2:48 · So you you are at Zera the company which is on the prao frontier of confusing names and mega rounds. Um so uh Zera is I think kind of came out of stealth like a few years ago and just really big org kind of out of nothing.
3:06 · So, um I'm curious if you can explain a little bit about what is Zara's mission, what is their thesis statement, like what is, you know, special about Zera and, um, you know, kind of where you're going in the future. Yeah, Zara is a uh AI uh enabled drug discovery company and at the core of our mission uh we're using AI platforms to generate better therapeutics uh to advance patient care and so we will be making drugs uh using different AI capabilities. There are three main AI platforms that we're building here. Uh the first one is protein design work that spun out of uh our co-founder uh Dr. David Baker's group from Udub. Uh a lot of the current generation of protein designers are here in the company. So there the thinking is to use advanced AI technology to develop um uh molecules against previously unable targets. The second AI platform I guess we'll spend a lot of time talking about today is the one that uh B and I have been working on for quite some time and just released a pre-print on. That's the virtual cell or foundation model of biology work. there the hope is to build a AI model to predict biology exactly like you said uh and predict what genes and drug molecules will affect cell biology and the third piece which we're beginning to build now is patient representation models and the goal there is to have AI models that can understand which patients will respond to which therapeutics so hopefully together these platform technology will help us make uh better drugs faster and uh with a higher success rate than previous technologies to transform what is used to be artisal tri and era in the past into more and more into an engineering discipline.
4:51 · Yeah, I think what sets Zera different is not just the one billion uh [laughter] around but [clears throat] also um I think Zera is one of the very few AI native companies for drug discovery that works from end to end of all sections of drug discovery from as early as you know target ID and the protein design small molecules and two phase one to three clinical trials. We aim to use AI to accelerate every part of the drug discovery so that not only we increase the success rate of uh developing drugs but also greatly you know reduce the cycle time so that we can have you know new drugs instead of every 10 20 years so hopefully we can have the cycle time so we have more useful drugs for patients.
5:42 · That's really interesting. I know there's a lot of interest right now in that third thing some maybe called translation from [clears throat] the lab to the clinic. Um where are the bottlenecks? So you have these three models. What are the bottlenecks that you're addressing and and sort of like how are you doing that? Why are you doing it that way?
6:00 · There is a AI native company. Uh almost every parts of the sections of drug discovery we trying to use AI to revolutionize how we develop drugs. So the early part we build causal foundation models uh or sometimes we call virtual cell uh proteins we have uh you know state-of-the-arts protein engineering models and we have also patient representation learning uh models and I think what Zara is trying to do is not only we develop AI models but also we create the right data sets to empower these models and I think what's really make me excited to work at Zera is we always aim to connect three AI models together instead of letting them work individually by their own. So when we design virtual cell models, we look for connections to that can we find targets that is easier to to apply the protein engineering models and then even when we design the cellular causal models, can we connect to patient representations? what are the right patient data to connect the cellular models so that we have something to show uh clinical utilities. So I think what really make me excited is um before I joined there I'm kind of a professor in computational uh biology department or computer science department where we mostly working on computers we look at the data look at arrays uh etc. But once coming to Zera, what really excites me is that I get to talk to people like Chu, lots of uh drug hunters, you know, extremely experienced drug hunters to really understand their pinpoint. So when we design AI models, we think about questions that really excites biologists. So later maybe we can talk about how one of the rewarding signals I received after we developed Excel is that like the wow moments from biologists that this is the first time biologist actually find the model can predict exactly how these unseen cell lines uh kind of respond to different perturbations. So that's kind of the part really excites me is the integration of kind of dry lab or AI models to wet lab or the biology or even eventually to the clinical side with this clinical model. Um I know you guys are aiming to you know take a drug all the way to to FDA approval um and beyond. Where do we stand now? I don't know if you're able to talk about this, but like are you able to collect data from clinical trials and tie that back yet?
Why biology lags behind protein design — the data bottleneck
8:41 · As B said, I think if you think about drug discovery process, it's easy, right? You just need to find the right target, make the right [laughter] molecule, and find the right patients to give them to. Of course, each of those steps are incredibly difficult to get right. And so far, like I said just now, it relies a lot of on trial and error and guess work. And the main issue I think is that we don't have the right biological data really to power the training of a predictive model. And in protein design space I think that's where we have seen the most rapid progress so far.
9:16 · That's partially because we have a lot of data high quality data over 70 years curated by the entire community.
9:23 · People deposit protein structures into a database right called PDB. We also have a lot of um sequence data right collected over the years from different genomes that can help inform the model as well. Um and it's these high quality data that are uh collected and accumulated that ushered in this revolution in proven design and alpha fold and other folding models. in the other domains such as clinical model prediction such as virtual cell we are nowhere near the same kind of massive data that are high quality and I think it's mainly a data limitation issue so to your question that's where we're in very invested in generating these data particularly causal data in cell biology in a lab and that's I think what made it possible to innovate on the algorithm side as well to usher in virtual cell models on the patient side it's a Very interesting question. Perhaps that's one of the hardest data to get because um getting access to highquality patient samples is difficult in itself. Getting it matched to the right clinical annotation so that you can actually learn the difference the bridge between molecular data and clinical response.
10:35 · That's even harder. And you might be able to do that uh across different uh disease severities, but it will be harder to collect the right data to predict which drug treatment will or will not uh respond in a particular patient or not. And so that takes a lot of thought and a lot of careful curation to generate data out of. And so we're beginning to go into that area, but hopefully we'll be able to share more soon.
11:02 · Awesome. Maybe we should switch gear now. you you just released uh XL. Why don't you guys describe I I'll butcher it.
11:10 · Uh Excel is Zara's first virtual cell models. It is a AI model that can predict the response to genetic prohibitions. Uh certainly we can extend it to other type of interventions such as drug prohibitions, uh chemical prohibitions, etc. So can you just describe for the nonbiologists in the that are listening what what is a perturbation? What do you mean by that?
11:36 · In our cells um uh when Bo talk about genetic perturbations our cell human cell typically have 20,000 genes. Um uh not all cells express every gene equally. That's why your eye cell, your skin cell, your heart cell, even though they share the same genome, they function very differently. A lot of that's determined by you know selective gene expression that determine the type and the state of the cell. So um what we do is to build a model that you can encilico uh ablate certain genes from the cell that is a encyclilical partation that's to say if I reduce the expression of this gene in the cell what is the implication for the rest of the cells what's the biological consequence this you basically turn the knob down on one gene that's right and then that what happens to all the other genes in that cell correct and the hope is of course to predict the effect on all the other genes but maybe even more things than gene expression such as the function of the cell.
12:37 · Okay.
12:37 · And that's important that that's therapeutically relevant because a lot of drugs are inhibitors and they function through exactly that turning down the activity of a protein or gene.
12:48 · And so if we can start with gene partition prediction the hope is that we can also go to pathway inhibition prediction so on and so forth. So a pathway is just a set of genes that all kind of talk to each other by this this gene expresses a protein. That protein um has some impact on another gene and so forth and so on. There's this longchain reaction of of genes and proteins. And then so that that's called a pathway. And so if you interrupt that or somehow change it then that has an impact on the the larger phenotype of the cell, what the cell looks like, does etc.
What X-Cell does — perturbation prediction explained
13:23 · That's exactly right.
13:24 · Yeah.
13:24 · Yeah. So uh you have what you called a virtual cell or you're you're creating a virtual cell and virtual cells are very popular these days. A lot of people are interested in this concept but I think your approach is somewhat unique or separate from other people are doing. Can you explain what do broadly people mean when they say virtual cells?
13:44 · What are some of the distinct other strategies and then like what is your specific strategy that you're going for?
13:50 · Certainly virtual cell is a very high level term to describe a AI model that is able to predict or describe what cell looks like or predict the cell expressions or cell functions after certain interventions. It's a very high level concept. It it was first of all it was not a novel idea. uh we had virtual cell uh project almost 20 years ago but back then sometime we call it virtual cell 1.0 Oh, is that uh people trying to uh derive differential equations to trying to use mathematics to describe what's the response for certain uh pathway interventions as you just mentioned and uh by fitting uh these equations to different observations and uh largely speaking that was a failed attempt in the sense that the biology is just way too complicated to write in a few predefined set of differential equations.
14:51 · uh moving forward uh with the rise of language models, I think that the idea of using AI models to mimic how CE responds to different interventions by datadriven approach start to get popular and uh I think three years ago almost just four months after Chad was released our lab at University of Toronto published one of the early foundation model of single cell genomics called it SCGBT.
15:21 · uh you can kind of uh interpret as a GPT like a model for single cells and uh it quickly become very popular in the sense that for the first time we have a foundation model that is able to tackle different downstream tasks using the same model such as we can use the same model to integrate different batches of single cell RNIC we can use the same model to predict multiomic uh integrations let's define those things so batches uh integrate different batches of RNA [snorts] you seek. So, uh you have different equipment. You're all collecting maybe different labs.
15:56 · Yeah. Different labs, different time of day, correct?
15:59 · Different phase of the moon, whatever.
16:01 · And those actually have a big impact on the data that you collect. And so that there's a big problem of how do I even compare this data set to that data set when there's all this other differences that have nothing to do with gene expression and just how I measured it.
History of virtual cell modeling — from differential equations to scGPT
16:17 · We call that batch effect. we certainly want to remove the batch fact while preserving the cell types which are more important biology we want to uh reserve.
16:26 · So this is sort of like analogous to the tank problem in image classifiers for example is a the sort of the models pick up on these crazy spurious features which have nothing to do with what you actually care about underlying biology.
16:38 · The certainly the the core idea of integrating different batches is to uh kind of keep the biological signals while removing the batch effect. And uh before these foundation models what what happens in single cell domain is that uh for every task biologist have to choose the so-called specialist state-of-the-arts approaches and uh um with with foundation models such as SGV or gene formers uh what we hope to bring is that one model that solve all the tasks in single cells and with the popularity of foundation model lots of researchers come together under CI chakra institute and we published a perspective paper uh at journal cell to coin for the first time coin the term virtual cell almost virtual cell 2.0 Oh, in the sense that let's use datadriven approaches. Uh if we cannot describe, let's learn it. So that's the idea of virtual cells. So that gem spin can we build language model or language type of model to predict what the cell types looks like how the cell respond to different uh interventions and eventually we kind of we can replace all the cellular experiments by simply running simulations on computer without even running the actual experiments.
18:00 · Maybe for a bit more context you can think about this is so a virtual cell is just a general concept. Um but you think cells have 20,000 genes in them and in most human cells I think what roughly you know four to 5,000 are usually active at any given time or expressed at reasonable levels. So you look at a normal cell you might have four to 5,000 genes doing things. Um and so your your question is in many cases the way medicine works is you you know you target a protein or you target some sort of you know something which makes proteins more common or less common or they stop the protein from doing something and your goal is given this you know some number of genes which are in a cell every cell has a different composition of genes. What is going to change? you know, will some some pathway die off? Will some pathway grow? And how does this, you know, from this you could predict how a medicine is going to work by just understanding how changing one specific gene or some cluster of genes could change everything? Is that is that a correct understanding?
19:08 · Yeah, that's a correct high level understanding about virtual cell. What's happening for this field is that we are lacking a concrete definition of virtual sales [laughter] and uh people almost equate foundation model with virtual sale. I but in my view virtual sale is probably a much broader concept than just foundation models. Uh foundation models mostly provide a reliable uh semantic meaningful representations of sales. But I think virtual cell is more dynamic in the sense that can we build AI models even predict the the development of different uh the cell states across different times or even can can we even describe the spatial changes at the cells uh different cellular resolutions. In my understanding is that we are really at the early stage to develop such comprehensive virtual cell models and the foundation model is really just the starting point. AI models always begin with the data. You are building a high throughput experiment or have built and are continuing to develop a high throughput experimentation system. Can you that sounds uh really cool and really complicated. Can you tell us what that entails like? What are you doing?
20:24 · what are the experiments that you're running? How does that inform uh the building of an AI model? And why do this rather than pick up the CLX gene database which is a collection of gene expression data that has been aggregated over the public data sets?
20:41 · Yeah, great question. I want to pick up where Bo left off. I think Bo said something pretty profound going from a representation model, the foundation model of biology to a virtual cell. And the key difference there is um perturbation prediction or dynamic uh dynamic processes in biology. That's a causal concept. For that I think we need causal data. Uh and if you look at cell by gene that's a fantastic data set that curated at the in the beginning more than 33 million cells now a lot more than that. And at the time when SGBT was trained on that uh data set coming out of Bose lab in Toronto that was mostly a observational profiling data set. It's a descriptive data uh not causal and mostly profiling healthy human uh donors. And so the model that was trained on this data set is very very good at doing descriptive tasks such as harmonizing across batch effects uh removing effects from different labs, different technologies. But I think both us and many others in the field have found that these models that are trained on descriptive data do not yet outperform linear models on causal tasks, perturbational tasks, what we call counterfactual tasks. If I did this to the cell, then what would happen?
22:03 · That makes intuitive sense to a biologist because the correlation data in the descriptive data set can be fit with many many possible causal structures. In a very simplistic case, let's say you observe gene A, B, C all go up and down together in your descriptive data set. You can infer that A regulate B and C. That's why when A goes up, B and C also go up. You might also say that B regulate A and C and that would be perfectly reasonable as well. You could also say that A regulate B and C is completely regulated by something different. You see the problem there. And there's n number of way to fit a causal regulatory network into this group of data. Fundamentally, we believe observational data are underpowered to learn causality truly.
Perturb-seq at scale — building the PISCES dataset
22:50 · And uh this is why we realized pretty early on that we need to really start training co building causal data set to train a causal model. So what are the ways to do that? Um I think the the field has come of age to do these at scale uh technique that we call hyprobiology and there are many ways to generate these causal data at scale. The technique that we have focused on is something called perturbse seek. So um for the listeners who are not uh familiar with that technology, it combines high throughput pulled crisper perturbation together with single cell RNA technology to build 2D data sets.
23:27 · Let me let me break that down.
23:29 · Yeah, go [laughter] please.
23:31 · So we just talked about in a cell there are at least 20,000 analytes to measure. These are the genes. These are both the features to measure. These are also the liver to perturb the cells with. So for uh clarity let's call them perturbations and uh gene expressions for perturbation uh on perturbation on one axis and the features on that you measure that describe the cell on the other axis.
23:58 · Perturbs is a technique that uh leverages the latest uh breakthrough in uh lab biology crisper cast 9. These are bacterially derived enzymes that allows you to disrupt gene expression in mamleian cells in human cells for example and we can do so in one at a time fashion. So I can take out one gene at a time. Of course that would be incredibly difficult to scale if I want to do all 20,000 gene expression knockout in one single experiment.
24:30 · Probably need a huge factory a lot of robots to do that. Or you can do them in a pulled fashion. And [clears throat] I love pulled experiments. These are hyperscalable. So we have lab tricks that allow us to disrupt one gene per cell, but do all 20,000 genes across many many cells in one single pulled experiment perfectly scrambled. So there's no batch effects.
24:53 · There's no play-to-play diffication throughput. basically use some sort of commutatorial trick to first perturb all the different genes in different combinations and then you can read them out and do some math on it and you basically pull out a whole bunch of different experiments in one experiment.
25:12 · Correct. It requires barcoding technology and that barcode is actually achieved by directly reading out uh what kind of crisper guide RNA is present in which cell. So for crisper castine this bacterally derived uh machinery to work in million cells you just have to deliver two things to each cell. You have to deliver the protein the castine protein that does does the job and you have to deliver an address barcode encoded by a short piece of RNA called uh uh guide RNA and the guide RNA tells the protein where to go in the cell purely via was quick base pairing uh ATCG. So it matches a part of the gene.
25:53 · It's sufficiently long to say this will match the correct gene and then that guides it to connect to the right and and and reduce the expression of that particular gene in the cell.
26:05 · Correct. We designed this guy to go to the promoter part of a gene. That's the beginning stretch of every gene before the transcription starts. And if we bring the cast 9 protein into there armed with the rightector the silencer that promoter will get shut off and that gene will never be transcribed out of again. So effectively we tune down expression level of that gene and so all you have to know is figure out which guide RNA is in which cell and that can be done using uh genomic readouts.
26:32 · That's the barcode and you can then infer which gene is being silenced in which cell. So that's the way you scale throughput on the prohibation side. On the readout side, it's a 2D data set, right? So we just talk about one of the dimensions. On the readout side, we leverage single cell RNA seek technologies. So these are also uh recent technologies in the last decade that have been scaled that can let you read out expression level of all 20,000 genes simultaneously from each cell. So armed with both high throughput crisper perturbation and high throughput singles RN6 technologies all of a sudden we can generate these 2D data sets where we systematically perturb or knock out knock down every single gene in the human genome in the cell type and we read out this impact on every other genes in the in the same cells. So we generate these 2D rich data set not that different than the size the type of PTB data that trained alpha models right if you think about that that's hundreds of thousands of protein entries if those are the rows columns are the XYZ coordinate of every single amino acid that's also a 2D data set and I think it's these type of rich 2D data sets that power the training of foundation models of biology I find it really fun how you have turned a fairly straightforward assay in using I get this is NGS sequencing right next generation sequencing very high throughput you've used this to scale um a simple pertivation response which is individually maybe not all that interesting to this massive scale of over a basically a arbitrary number of cells I think you did 25 million or something so it's actually a lot more than that so 25 million is what came out of the most stringent quality filtering it's actually as much of a scientific challenge to figure out how to do crisper and single RNA seek as it is an engineering challenge in the first part of the experiment. Often times we have to harvest tens if not hundreds of millions of cells and they go through various quality uh funnels to arrive to give Bow and team the highest quality data at the end. That's incredibly difficult to do because as you can imagine all of these techniques have been published by academia before um and they work very well in small scale experiments but when you think about scaling them to a genomewide perturbation we're talking about handling hundreds of millions of cells.
29:00 · Techniques that are published in academia used to be all about handling fresh cells. Cells are still alive and that may be okay if your entire experiment takes only an hour or two.
29:10 · It's not quite easy to handle cells across a 14-hour day. That's hundreds of millions of cells. And so by the end of the day, uh I used to joke with my team, you can easily detect stress signals from the cells and from your scientists in the lab. [laughter] Yeah.
29:26 · And quickly we realized that's not the way to do these data generation. Uh machine learning is uh very quality dependent and we want to give the the highest quality data to our AI teams. So we're putting a lot of engineering thought and industrialize the whole workflow step by step. Introduce chemical fixations so that we lock the state of the cells in at the beginning of this experiment but figure out ways that it doesn't disrupt all of the biology uh molecular biology steps afterwards. It doesn't impact data quality so that we can do all of these data generation in a timeshifted operational manner uh that's very um uh not prone to batch effects. One thing that you didn't mention is that you're using some sort of stem cells. Um, and so and obviously like you don't have, you know, like brain cells or or or blood cells or if you did then you would have a big combinatorial effect on that.
30:18 · So how are you why are you convinced that working on stem cells um which are you know my understanding is that there are actually blood cells that have been sort of the stem cell behavior has been unlocked on them and and that causes some some sort of stress on the cell as well. So you have these like sort of not quite blood cells that are stressed and then how why are we convinced that that is a good proxy for a brain cell or whatever you're studying.
30:52 · Yeah, not quite. So we um uh didn't actually start with stem cells. That was more more of a later development. Okay.
31:00 · when we started data generation. So we put out um the method that I talk about as well as the first two data sets which is the world's largest perturbic data release at the time last June in the preprint. We call a data set XLS Orion that was actually generated from two cell lines. A lot of this field's early work started with cell lines. These are cancer cell line.
31:21 · Uh one of them is a cancer cell line, the other is just a cell line.
31:24 · These are immortalized uh cells. Some of them are derived from cancers, hence cancer cell lines. Others are just derived from primary cells but have been immortalized many times grown for many many years uh in various labs. People start with these uh cell lines in the beginning as you can imagine because those are easy to do. It's easy to scale, easy to grow a lot of cells out of. Turns out the ability to grow millions of cells is actually critical for doing these large experiments. Um so we started there first and they actually still uh capture the characteristics of the cell types that are derived from colactto cancer uh as well as um uh uh hematopoitic uh cells but later on uh in the most recent preprint we actually expanded to many more cell types. Now some of these are still cell lines our T- cells we chose to use cell lines but some of these have now gone into primary cells. So we did one experiment in IPSC these are induced plur potent stem cells and another experiment in uh and we think this is the most ambitious and coolest uh screen that we've done to date. This is a pen differentiation multi- cell type stem cell project. So effectively we differentiate iPSC into 10 different cell types in one single experiment without restriction and we did a genome scale partation across them. So you can imagine instead of just generating 10,000 different biological experiments we did 10,000 by 10 cell types was almost a library on library experiment. Why are we doing this? We think that in the beginning phase of data collection as both said I think we're just in the early days of um uh virtual cell building context and diversity and richness of the data matters. It's not just the total number of cells or total number of sequencing reads. It's about bits per dollar and information content. So we want to scale not only in the genetic perturbation landscape but we also want to scale across biological context so that we can give our AI teams the best rich data set to build a uh generalizable model on. Is there any thinking about so you say context but obviously these cells in these experiments have been sort of de uh I forget the term but they they've been separated from their cohorts right is there thinking about using spatial transcrytoics or other you know sort of technologies imaging based technologies to build models um with perturbations but with in in the context of the cells that it lives near.
33:56 · Great question. So we're thinking about that in a couple ways. Um uh number one that's actually exactly why we want to build a virtual cell model in the first place. Uh you might think that well you can already do exhaustive screening in these cell lines why do you still need a model you can just do the experiment and generate the data. Certainly, if your query is just about cell biology in cell lines, you're right. We don't need a model, right? At least if for genetic screening, we can just do this experiment. But you're also correct that often times good targets, biological insights are not about cell lines. These are about primary cells, about cells in their native physiological context in organs or even multiorgan coming together and have some emerging properties. A lot of um emological disease are that way. You cannot do exhaustive high throughput experimentation in animal systems or in organs or in all of these complex translational models. You can do some experiments and these are expensive and high stake. the ability to build a model that can be trained on massive data where it is possible to scale and be trained in a way that can be fine-tuned and transferred to make high quality causal predictions in these complex models so that we can go into the lab and have the highest quality hypothesis possible to validate. I think that's the whole point about building a virtual cell model. But from AI side, I think you're absolutely right that I believe the future virtual cell model should be able to incorporate multiple modalities not just RNA expressions. Spatial spatial single cell RN stick is already a a popular tech technology even for SGBT. We actually have a extended version. We call it SGBT spatial that is specifically designed for spatial single cells. And uh uh we also have papers on early attempts to trying to take the H& images trying to predict the gene expressions. Uh there's already some signals you can find. So eventually what I predict is that uh virtual cell model will be able to integrate not only RNC can integrate more functionally uh uh related for example proteomics or other regulatory um side of uh OMIX such as taxic to overall combine all your descriptive omix data sets to predict the the future states of the cellular uh functions. I think that's probably the the future for virtual sale model.
36:23 · We had uh Ron Alpha and Dan Bear from Noetic as guests uh recently and viewers who want to hear a little bit more about that. I think they they go we go in quite in depth there. So if you want to background you can go to that but can you explain a little bit about what um spatial transcripttoics and spatial proteomics are?
Spatial transcriptomics and future modalities
36:40 · So maybe uh a bit of a history lesson here. Before we had singles RNA seek we had RNA seek and before that we have microarray technologies. What RNA seeker microarray used to do is take a chunk of my tissue, grind it all up, put in a blender, imagine make a smoothie out of it, and take the all of the RNA from different cells in that piece of tissue and measure all of their expression levels. It is great for the first time you can measure gene expression all 20,000 at a at a time. And we used to do them, we used to have to do them one at a time, but it is not great in that we don't know which RNA came from which cell. And this is particularly a problem if you're dealing with a multisellular piece of tissue. You want to attribute RNA to the immune cell, to the skin cell, to the fibroblast, to the keratinocytes, but you can't because you grind everything up in a smoothie. What single cell technology allow you to do is analyze them cell by cell. So now I can attribute RNA gene expression to the cell that they originate from. But there's still a problem. I don't know spatially where the signal come from.
37:45 · And for many disease it matters right in imunoncology for example you want to know when TE-C cells are close to a tumor cells or uh when a T- cell is not able to penetrate the solid tumor what is the difference between them or when a T- cell is attacking the tumor cell when the T- cell is not uh what is the difference about about them and for that you need spatial information you need to observe cell in situ in their context and so now there are different technologies that solve that problem.
38:15 · Essentially, take that chunk of tissue.
38:18 · I don't have to grind it up anymore. I just make a cross-section, lay it down in a piece of slide and I can measure its morphology using standard techniques like H& staining. I can then also measure many protein expression using multiplex if assays in flororesence assays. Ultimately, I can also look at the gene expression up to genomide in all of these um uh cells in their native [clears throat] spatial coordinates by using some of the latest spatial OMX assays. So, you have the XY coordinates of every cell, but also all of the molecular analyze that we talked about earlier. And that's an exciting new direction for genomics field.
38:58 · Personally, you can imagine the the spatial omics adds more difficulty to AI modeling because you have instead of looking at the in individual cells, you have to look at the neighboring niche cells to better kind of learn the representation that is spatially cohesive. That is the challenge uh the current spatial foundation model are facing. But that context is going to be crucial for correct I mean understanding let's say cancer where immune the interaction of immune cells and cancer cells and nonimmune or cancer cells crucial for spatial aware biomarkers uh to predict some of the clinical response I think that would be extremely important uh to build such models getting back to Excel this you know presumably can inform a spatial model as well right is you you can you have like one cell in one place if you can imagine okay I can just throw away the coordinates and just get do inferences on one cell at a time and now I can create I can create a more complicated model that does that but it also knows who its neighbors are saying you're absolutely right but the current version we were releasing we're not dealing with spatial omics however definitely our ongoing work and uh the next version of Excel will be able to infer the spatially aware representations for different cells.
40:23 · I see we've talked about the data collection a bit. Let's talk about the architecture. Get some red meat for the uh for the AI engineers listening in.
40:32 · Sure.
40:32 · Let's get to the history of uh virtual cell modeling particular virtual cell 2.0. I think our SGV kind of sets the foundation for most of uh foundation models of single cells is that we adopted uh kind of auto reggressive training uh extremely similar to how chibbt is trained on languages right we use uh it's next word uh next token predictions so we mimics the way how chibb is trained on languages to train a single cell foundation model all on sales by doing that they have to assume inherent order of genes, right? The way we uh assume the order of genes is by attention mechanism. There's many other methods that using different orders of genes. Some some as simple as just rank the genes based on the expression values. There's also more complicated uh kind of methods to rank different genes.
41:26 · But inherently you have to have assume a order of genes.
41:29 · And just to be clear, so when you talk about genes, those are intrinsically ordered, right? They're a sentence spelled out in ATGC, right? So genes themselves have this the nucleotides and there's this long chain and that makes a lot of sense to have an order to them. But what we're talking about is something different. That's the expression data, the expression levels.
41:53 · So the expression level means how it's just a count for each gene of how many of these genes did I see when I was measuring. for DNA sequences the order of ATG make total sense to us right but for expression data they're literally just mattresses so it's really hard to assume uh inherent order of genes even if you shuffle the the order of genes I think the viology doesn't change much however because of the way uh kind of language model is trained everybody has kind of preset tricks to train such a model so it's easy to adopt that's how all the foundation model are started for single else and then I quickly realized that um with diffusion language models we actually don't need to assume the order of genes uh instead we can have a birectional diffusion process to generate such um long highdimensional gene expression data sets. So just to think about it what's the difference between auto reggressive training versus diffusion language models is that kind of you can think of diff auto reggressive training as typing there for example I like coffee you have to type I and they like and coffee there's inherent orders but diffusion language mode you can treat it as editing you iteratively generate a sentence from a very vague rough uh sentence and then you can iteratively refine it so same thing with jinx questions you can generate a very rough representation of the gene expressions and then iteratively from noisy representation to more refined representations. So you kind of iteratively edit the gene expression predictions until it minimize the losses. So this is a very different philosophy to generatively uh predict the response after perturbation and turns out it actually fits more to a single cell iron sick. So that's why we switch it from SGTV like model to the current Excel model which using diffusion language models.
43:57 · When I think of transformers like they're fundamentally objects which operate on sets. Uh the community spends a lot of time trying to make them things which have some sort of causal ordering to them. uh but if you just naively take a transformer it's it's a set operation right so given that why think about this in terms of diffusion or um you know autogressive LLMs why not have your initial prediction strategy be something like take a just a set of genes each of which has its own kind of one hot encoded um identity and then use that as sort of a prediction that seems like a much more natural architecture to me and I it's not you know your work a lot of people work on things like this and I have been somewhat confused why there's this bias in the community about this that's that's a so I think what what you were referring to is more related to representation learning where you can take sets of genes and trying to project to low dimensional latent space but what we care about for building uh generative modeling for for virtual cell because you want to predict the dynamics of cells so we want to have a generative models so that's why we mostly using decoder only architectures in order to generate the full transcripttoics instead of just predict a predefined small set of genes because you want to model the whole gene gene gene regulatory networks which are extremely kind of highdimensional right so just to be clear input is genes plus a perturbation output is new gene expression levels is that gene expression levels plus perturbation is input output is cells like for each cell.
X-Cell architecture — why diffusion beats autoregression for gene expression
45:36 · Correct. Correct. That is correct. Yeah.
45:38 · Right.
45:38 · Okay. And so that the way that I think about um the way I think about diffusion language models and you can correct me here because I don't know a lot about them but the way I think about them is they're like BERT but you do it o over and over again. Is that is that kind of a good Yeah, that is a rough understanding of how diffusion language model works.
45:59 · Yeah.
45:59 · So, so you just apply the diffusion. The diffusion process is like basically unmasking or editing correct once over and over again. The the similar to how like a image diffusion model kind of refineses the image over and over again. In this case, I'm using B. So, it is a transformer basically. It is a transformer but it's like repeatedly updating the the sort of sentence in this case which is a bunch of expression levels over and over again.
46:31 · That is correct. Actually in our paper we show that as the number of diffusion steps goes on the act the loss function uh keep decreasing that the the fitness of the prediction to the ground truth keep increasing. So this which means the model start to understand how iteratively refine the predictions.
46:52 · I see we're talking about diffusion versus auto reggressive. What what there I noticed in the paper there's a bunch of discussion of preconditioning using a whole bunch of stuff. Can you want to talk a little bit about that?
47:04 · Another major [clears throat] innovation we made in Excel is the way we incorporate prior knowledge into the model. So incorporating biological priors has always been a good idea in biology in general because uh biologists spend you know decades to to understand uh some of the biologist already. How do we uh tell the model some of the prior knowledge some metadata about the sales?
47:29 · Before Excel, what people do is they they trying to incorporate a single type of prior for example gear uh using gene regular network as a prior to predict the prohibitions. SGB sometimes trying to incorporate PBI as a prior as well.
47:46 · Excel to my knowledge is one of the first models that trying to incorporate a extremely diverse sets of biological priors. So in our preprint we incorporate five types of pri uh priors including literatures as simple as just ask chatb tell me everything about this gene and then we embed as the output uh as the embedding to gene pt exactly that's a gene pt and we also uh incorporate ppi protein protein interaction networks we also uh incorporate dev map which is uh cancer related essential gene informations uh morphology informations We even trying to uh incorporate SGBT embeddings which is basically cell types. So with a set of prior knowledge as conditions to the model, the model start to have more accuracy in terms of context specific predictions. And what's more interesting to us is that by looking at the weights of different priors, we can actually understand which prior knowledge are more important to these particular cell types. So it adds more interpretability to the models. So we find that combining diffusion language model plus a very diverse sets of privileges, Excel does much better in generalizing to unseen uh context. So this is some of the AI innovations we made for Excel. Do you now need to provide all of that context in order for the model to work or that those are like preconditioning that it can also do without if you want.
49:23 · We don't need to incorporate these prior knowledge anymore because these are already learnable parameters inside the models. However, what you suggest is more promptable or in context learning for virtual sales. We can do that as well. basically by adding more conditions into the prior knowledges so that to prompt the model to predict towards certain directions. In other words, you took it you the model now takes advantage of the learning using the the priors that you provided during training and doesn't need them but it but has some advantage because you provide them during training but you can even get more advantage if you are able to provide those prior during inference.
50:04 · Yeah.
50:04 · Okay.
50:04 · Wow. Nice. Yeah, how much does that matter? I mean, whenever I see big machine learning papers with tons of things thrown in, I'm always wondering where's like the big alpha and where's the little alpha? How much are you know, is this just some are these adding this little bit of incremental performance boost or I mean are all these actually crucial to general generalization.
50:24 · So there's multiple factors we have to consider. How much contribution the data contributed? How much of the contribution the AI architectures contributed? Even for the architecture what's the delta from switching to auto regressive uh training to defa what's the delta from the prior knowledges certainly all of these needs very specific ablation studies from empirical experience we find that the qualities the amount of the data sets matter the most. This is why we were extremely excited to to to publish the Pisces data sets which has 16 different cell types and uh across 25 million cells and it's genomewide you have kind of a huge tensor if you really think about from computer perspective genomewide perturbation genomewide transcripttoics plus number of cells plus times uh number of conditions so it's a massive tensors and those and because of the post screening in uh technology we don't have batch effects. So so you don't you don't need the model to climb the heel of batch effect. So that's already advantage. So we find that train on prohibition data sets high quality prohibition data sets already gives a big boost to the models. We also did ablation that if we train all the virtual cell models out there including state uh cell to send original SGV on the same data sets what's the delta we are observing we we report the results there as well we find that switching from auto regressive training to uh diffusion language models give a significant uh improvements over some of the harder tasks particularly generalized to unseen tasks And the pime knowledge more or less condition specific. Uh for certain cell types some of the pry knowledge make a huge difference but for certain cell types the delta seems to be marginal. Um we are thinking about you know how to better incorporate the the prime knowledge. We still believe that let the model know a big chunk of existing biology should be helpful but maybe it's the way we incorporate the pan knowledge through cross cross attention limited the the scope of the metadata but I think it's certainly a a research topic but overall if we have to give an order my order would be the quality among scale of the data sets and then the the architecture and then the prior knowledge But certainly this is only applies to our Excel. I'm I'm sure there's different choices of architecture have different ranks of contributions.
53:13 · First of all, this is really fascinating, very cool model. I hope everyone has a chance to look at the paper. Um there's a lot of obviously a lot of resources that were put into doing this. I don't know if you guys can disclose how much. It's a lot of money.
53:27 · Whatever it [clears throat] was operating a wet lab, probably very complicated training runs. Um I think there's a couple four billion parameter model. Is that right? Or 4.9 billion.
53:38 · Yeah.
53:38 · Five billion parameter model. So much larger model. Um probably took a lot of GPUs to train. Um what's the lift that you get from this effort versus let's just put the money into like wet lab work and the sort of traditional pipeline um that that you know basically has been the status quo up until now.
54:00 · Biology is a multiscale um uh discipline. There are cells or there are DNA sequences on the most fundamental level. There are cells, there are multisellular pieces of tissues, uh co- cultures, you have tissues, you have animal systems and finally you have human. I think we would like to be able to do causal prediction towards the right of the spectrum. ultimately do color prediction in human know what drugs will work in uh which patients but that's very difficult to collect higher data on and so the whole vision of virtual cell is to generate data where it is possible so that we can transfer the causality prediction towards the right towards the more translational the more complex systems the uh certainly you can mine the data already we generated a lot of data as both said seven screens 16 different biological um context text uh genome scale perturbation there's a lot of good ideas in that already there's a figure that we put out in the preprint that just look into inactivation of T- cells we already saw some you know we saw known biology TCR complex we also saw some pitive new biology which we're very excited to validate in the lab some of that were um actually also caught out in a very recent uh screen last December published from Alex Marson lab also in the Bay Area so very excited to see that um but the hope is to not just mine the existing data. The hope is that the model can generalize and we will be able to do incilical experiment into the future. Nobody knows before um how much data and what kind of data are needed to do that with the whole field is waiting for the demonstration that the model can beat linear baseline in partation prediction and it can generalize out of context not just within a cell line you have training data on but out of that context that's where that's why you need a model so what's very exciting for us is that in this preprint we saw that generalization capability a few demonstrations We first did in T- cells. We actually generated the data expressly for this purpose. We generated a resting T- cell perturbation screen. So these are T- cells in their baseline condition not activated. And now we have a activated T- cell perturb.
Ablation results — what actually moves the needle
56:25 · So just T- cell activation means I'm going I'm trying to kill something.
56:29 · No, this these are regulatory T- cells.
56:31 · Yes, we activate their receptors so that they're starting to proliferate. Okay.
56:36 · Um they become more active, they can do their physical job and we only critically we only train the model on the resting T- cell and we told the model hey this is how the active T- cell look like now go and predict what all of the perturbation are going to do in this active T- cell. And the model have not seen how perturbation work in active T cells. And we set up a couple of rigorous tests. One linear baseline took the perturbational delta in the resting case. Just transpose that linearly onto the active T- cell and that's our linear baseline. Essentially think about this as a combinatoral perturbation prediction problem. One of the perturbation is activation of the cell. The other is all of the genomic perturbations. Can I just linearly add the two effects together and that would be a linear baseline and second we apply other models um from the field and last but we applied uh Excel critically XL has not seen active cell T- cells and it's able to make accurate prediction not only on the known biology the TCR complex predicting their effect accurately that these are going to inactivate T- cells which is exactly what we would expect to see but also it predicted the puditive T- cell inactivators that we found in the screen correctly as well. So that's very exciting to us and that suggests the possibility that we might be able to use these virtual cell models completely out of context in an unseen context and predict new biology. And so we're very excited to follow up on those heads and validate them in the lab. Just very briefly a couple other cases that we saw exciting generalizing capability of this model. Remember we did a multi-ell type differentiated IPSC experiment. There we specifically held out one cell type from training. Uh so the model has not seen that cell type train on the other cell types as well as the rest of the data sets. The model made very good prediction across thousands of genes, thousands of perturbations in that unseen cell type. So again suggesting the model's ability to generalize out of cell type. And the last experiment I think we're very excited is that we train this on T- cell cell line but there was just very recently a primary T- cell proseek published uh from Alex Marson's lab. That's a impressive amount of work. It's not easy to do this scale screening in primary cells. Um very few labs have that kind of capabilities.
X-Cell generalization — beating linear baselines in unseen cell types
59:06 · Much easier to do that in um T- cell lines. Again, the model is able to generalize out of cell lines into primary cells and make accurate predictions there.
59:15 · So, they [clears throat] actually perturbed primary cells not not cell lines, primary T- cells harvested from donors. Um, uh, and we were a for multiple donors and XL trained on just one T- cell line is able to make predictions in across multiple donors from primary T- cell experiments.
59:34 · This is a validation of the whole theory, right? that you can train on these slightly weird cells and that it will be good be because you know you're covering the the domain well enough or whatever it is that you're able to actually predict in real real cells that come directly from real people.
59:52 · That's right. Yeah, I think building virtual cell is not to replace biological experiments as you mentioned. What we trying to do really the holy grail of virtual cell is to have a model to generalize to unseen context that is harder or even impossible to to to conduct biological experiments on.
1:00:11 · Right? So far, excel focusing on cellular cell lines and eventually want to extend to more complicated biological systems such as animals, organoid and eventually as we mentioned before to patients to human uh human biology, right? And uh bear in mind a few numbers 90% of disease has no cure and uh most of the drug uh failed at the phase three clinical trials on patients and the success rate of phase three trials is as low as 5 to 10%.
1:00:44 · And phase three means the final stage on the patient uh trials.
1:00:48 · So that's when you generalize from toxicity in phase two to efficacy in phase three, right? from a small cohort smallhore into a much larger cohort.
1:00:58 · Oh, sorry. Toxicity is one, right? Small cohort and so a lot cohort. So the generalization problem okay this drug I've very carefully selected my patients and it works pretty well and now I get a bunch more patients and suddenly it doesn't work very well and that's the big problem that you're the certainly the promise of virtual cell is that can we build such a model that learn all the causality biology so that can be grounded to predict the response on eventually on patients so that we can for certain drugs we can select the right patient s to conduct the clinical trials on right so this is a long-term vision but we are already see some early hopes that excel train on diverse set of causal data sets can already generalize to some some unseen cell types so certainly there's a lot experiment to be done to validate this model so even continuously fine-tune this model but I think the uh we certainly see some early hopes so you were talking about linear models and this brings up this famous or infamous arc challenge about you know perturbivation and there's this been this theme about complicated foundation models oftentimes not beating linear baselines and I'd like to get your take about that as terms of what is this you know first of all is this different I mean I think some of maybe your own models might also have had trouble beating linear baselines in the past um is there something different about um your current data strategy or where you're going and where's the field going and what is the role of foundation models versus and virtual cell models versus you know these simple baselines and yeah a few things first of all those benchmarks as you mentioned are conducted on replo data sets which are very small data sets and uh the metrics people report are mostly MAE certainly you can imagine if the because single cell data set are so sparse the average profiles of of the cells certainly you kind of you can imagine it's a great minim uh kind of local optimum to minimize the maees this is why sometimes the average profile of cells has lower mas even than technical replicates which are considered of ground truth for for um kind of prohibition experiments so that itself shows that that metric is not reliable however Most of these benchmarks are still comparing foundation models that trend on static expression data set such as sale by genes SGB is often benchmarked against internally we also find that when it comes to MAE sometimes SGB kind of fail to outperform linear models just because of the reasons I just stated and what's sets Excel different from these static expression models such as SV or gene formers is that we actually instead of tren train on gene expression data sets. We train on causal data sets. We train on massive amount of genomewide perturbation data sets so that it learns better about the dynamics of the interventions. And in our preprint, we extensively compare with linear model as well. And uh as Chu mentioned linear model totally failed to extend to unseen cell types. You can quickly imagine why.
1:04:25 · And I believe that foundation model or or other more complicated AI models that train on the right data will outperform these linear models in harder tasks particularly in generalization tasks and uh that's why I keep mentioning the right data set with the right AI model will lead to huge improvements but I think the the field still needs to see more biological validations to be more convinced ing that the virtual sale direction is the right one.
1:04:58 · Yeah, I think the field suffer from a lack of a consistent and unifially accepted benchmarks. What gets measured will get improved and um in our paper we measured um I think the me one of the metrics that we put a lot of um uh thought into and uh saw the model really shine is uh metrics around gene expression changes. Uh so you know Pearson delta uh the similarity between predicted changes and ground truth changes upon the perturbation that's very hard to cheat on um you have to really get the changes right and what really blew blew my mind away is when I saw the model make prediction just print out the heat map of the genion changes look at the actual raw data and line up the linear baseline prediction the ground truth and XL prediction all together it's visually very clear to see that XL prediction is very much more similar to ground truth than than the linear baseline. This is a wow moment I was talking about in the beginning.
1:05:58 · That's right. And it's not hard to understand why when we put so this is the first time that someone can put together not just one perturb but seven genomewide perturbsy campaigns together.
1:06:08 · Something that uh jumped out to us biologists right away is that some of the perturbations are context universal meaning that the genes do the same thing regardless of the cell types. you experiment in. Might not be surprising to you that these are your housekeeping genes, right? Um, of course, they do the same thing every cell. And then there are all of these other clusters of genes that have very contextspecific functions. They do different things in different cells. Again, not hard to imagine why. In iPSC's in our stem cells, we saw developmentally relevant genes, genes that are important for neuronal differentiation. They only light up in iPSC experiments, of course.
1:06:46 · Right? That makes sense. So you think about biology, it's so complicated. You have to capture these, you know, context universal perturbation effects. You also have to somehow learn the context dependent partational effects. It's not hard to then see why a very sophisticated nonlinear model is able to capture and learn all of those biology much better. Are your pertivations always single gene pertivations or do you have more? Because my understanding of regulatory networks is often times um sometimes it can be a single gene does a ton of things. Um for example, I think males are just differentiated due to one gene being enabled at like day seven of embryo development or something and that differentiates everything is this one gene. But then sometimes you have large networks of genes which all are very redundant which allows for um more subtle feedback mechanism and so on. So I could imagine a lot of single gene pertivations as being kind of irrelevant. Um and that you might want to start having a more cominatorial um strategy here.
Single vs. combinatorial perturbations and platform expansion
1:07:52 · Yeah, that's a that's a great question and Bo and I have thought about this a lot. Actually, it's interesting that you brought up um reproductive biology in I study a lot of uh in female cells the compensation uh the dosage compensation mechanisms there. Female cells have two X chromosomes. Male cells have one X chromosome to match the dosage uh output from the X chromosomes. Strategy that the mamian cells employ is one gene that produce a RNA that does not encode for any protein just a non-coding RNA. that RNA wraps around the uh one of the female X chromosomes and uh turns down most of the gene expression from that chromosome shove it away in a corner of the uh nucleus and it's called a bar body. It is never heard from again and so absolutely agree with you one gene can do a lot but in biology you also have redundancy you have compensation you have all kinds of mechanisms where knocking down one gene is not sufficient to always see a phenotype. What if four genes redundantly do the same thing, right? Taking out one is not going to be sufficient. So where we started with um one cell type at a time, loss of function, single gene perturbation and look at only RNA expression as the output, we're expanding the platform along all of those axes. So that's what we do today to build a scaffold of the data for training models like Excel. We are now beginning to grow in all three axis of the platform. Going beyond trescopomics alone to look at multimodal data. going beyond just one gene perturbation alone to look at also pathway activation and inactivations turning on and off entire cascade of gene uh uh chain reactions and uh also going beyond just cell lines monocultures into more and more complex translationally relevant systems into primary cells into organoids and doing even direct invivo perturbation screens.
1:09:50 · So we believe with all of that expansion the data will be all the more exciting to train models on. This is also why we incorporate PPI networks as the prior knowledge into our model and and although the model right now are trend on single gene pertations but once the model is trend you can actually predict combinatorial perturbations just on the model in only silicico right so in the sense that you can just perturb the tokens of two genes at the same time and see what's the response certainly without training the actual combinatorial prohibition data set the accuracy may not be there but at least with the existence of such models we can start to generate hypothesis using incilical prohibitions. How do you see the role of the scientists changing in the age of AI and you you may because you you're not using you're your focus is not language models themselves and agentic science and things like that then you may have a different slightly different take you're building you know these very specific models um but still one how are you and your students able to maintain such a high pace and I suspect it may have something to do with generative AI partly but also um and how do you see the role of the scientist the academic um changing?
1:11:09 · Yeah, that's a great question. So my official split of time is 80% on Zara, 20% on my uh uh university affiliations, but turns out the reality is 100% on Zara, 100% on [laughter] this.
1:11:23 · Um you invented a time machine. That's the answer. So the way I trying to keep up is I mean we um my lab use lots of agentic AI trying to monitor all the AI papers every day and uh which every week we have lab meetings we trying to discuss different uh topics in AI for biology, AI for healthcare etc.
1:11:48 · And it's it's to be very frank even as a professor I find it's extremely hard to catch up. The pace of AI is just so incredibly fast and uh to the point that sometimes I feel anxiety waking up says oh my god this paper already so many people published and what happened to our existing unpublished work and suddenly you can imagine students probably face 10x anxiety. So sometimes I trying to encourage students to really kind of using different tools trying to stay focused you know find is a niche area that uh we become an experts on right but with the era of generative AI nowic AI I find the way people at least academics does science are extremely different now overall most of professors or students in academics start to be very struggling in terms of fundings in terms of the pace of publications and that's why you can see lots of major breakthroughs come from industry right like Alpha for example so how academics survive or even thrive at such era of agentic AI is certainly something everybody is thinking about we see lots of faculties left university and join industry for simply for the reasons of resources right um if you are doing research on AI do you have enough GPUs at your school is the first question you should ask when a student join a professor's lab the first question they often ask is how many GPUs do you have right so certainly in that sense industry has a major advantage over academic labs but I think what academics labs have advantages on is really the kind of the pace of innovations and also the niche areas this specific academic lab can be extremely expert on. So also having the the freedom of thinking sometimes also make you easier to innovate on ideas that maybe industry people didn't even think about. Overall, I think the whole field needs to be a lot more innovative to catch up and uh hopefully with the help of different tools and and hopefully the government start to invest more into academic because I still deeply believe that academic is the main source of innovation for the whole field and particularly when it comes to biotech. So um hopefully we we see more investment into academics so that we stay afloat.
1:14:33 · I agree with you but but why why do you think that it why why shouldn't the money just go to industry? Why why should the government put any money into academia?
1:14:43 · I still believe the power of you know academic freedom. uh this is actually the the original motivation we have such uh the existence of academic professors who not only we teach but also we do research and uh there's also benefits of teaching and research at the same time in the sense that when you teach a subject you actually have to become the experts on and that force you to keep updating your knowledge base and then find easy ways to convey your knowledge in to students and by doing that you actually start to innovate on different ideas and and uh myself for example I teach uh a big class in University of Toronto about deep learning and uh neuronet networks it's a gigantic class with 600 students every year and by teaching that it forces me to update update the slides lectures every year and uh myself reading lots of materials trying to update myself so that I can find ways to convey some all the knowledge into uh to students right so that's how I keep updating myself whenever there's trans I remember vividly we have GP2 we update the lecture then then different language models how the multi- thread GPU communication is used in training larger scale neural networks etc so I forced myself just update the models and by talking to different students we really generate lots of novel ideas to apply is a cutting edge AI models to very specific niche areas in biology or in healthcare. Um maybe it's unique to Canadian academic system. Um by being a professor in academic we also have access to lots of healthcare data sets which are very hard for industry to access uh due to many legal reasons or regulatory reasons and uh that's why you see some of the papers we publish through academic hospitals in Canada where we developed some of the state-of-the-arts foundation model for ultrasound images. So that's what I mean that there's certain level of freedom of uh academic thinking that really drives lots of innovations and uh I still believe maybe it's biased but myself still believe that having certain level of freedom of academic thinking will leads to lots of kind of innovation that is unsinkable in industries.
Academia vs. industry in the agentic AI era
1:17:17 · I 100% agree with uh Bo there. So you know just thinking about the lab uh uh workflow that we do a lot of these are building upon innovations that were first pioneered in academia as well.
1:17:29 · Crisper of course was um discovered in academia. Crisper applied to mleian higher screening also demonstrated in academia first. Singles RNA seek this kind of droplet encapsulated singles RNA seek first demonstrating academia then different companies try to build it up into commercial offerings. putting all of these together to do perturbs at first also pioneered in academia right Chris box lab Jonathan Weissman's lab eval lab many of the pioneers in academia and then I think these especially on the lab side this innovation takes so long and the discovery process can be so accidental right that it's perhaps not ideal for uh pure industry to take on but once they show early promise scaling them [clears throat] and robustifying them and generating data that's not only massive but high quality especially for AI scale I think that's something that can be very well done in industry both the mindset as well as the kind of uh the resources uh that we can support uh that sometimes can be hard for academia labs to match zer has been very generous with releasing your data sets and your models you've been seem to be very committed to open science given some of the things you were just talking about and you know the discussion about like what is the best most important data strategies for virtual cells or understanding human biology.
1:18:59 · Where do you think academia should go next now and next? Um since you know Zara has probably a budget comparable to you know probably dozens or hundreds of biolabs right now. What do you think that if you are an academic, a professor, especially in a wet lab, um what would you want to be focusing on?
1:19:20 · First of all, myself is a deep believer of open science. That's why all the models we talk about here data uh are open sourced. You can find all the data weights from my lab GitHub.
1:19:32 · It's very thorough.
1:19:33 · Yeah.
1:19:33 · Thank you. And the reason I believe open science is that as ch mentioned most of the time academic lab we started an idea and we prototype it.
1:19:42 · It's not scalable. It's not even a good product and industry can take it to scale things up. This is also why SGBT quickly become one of the most widely used single cell foundation model in former companies. So that's very encouraging to us and this is why uh Zera is also start to open source some of the data sets some of the models.
1:20:05 · Part of the reason is that we believe virtual sale is such an early field and uh it doesn't help to withhold certain data sets or certain models because it's so early. A better win-win situation is everybody gets to in this field start to contribute data together start to contribute models together to exchange ideas so that this field can move move forward in a much faster pace. We see successful examples in protein space right because of the availability of open source just as in PDBs therefore we have models such as alpha fold rosetta fold again alpha folder to also open source therefore we can quickly iterate different models that's why you see kind of a booming situation in protein space we want to do the same thing for virtual cells let's put all the data set together let's have the same standard protocols to generate high quality data sets let's put all the resources together to generate next generation of virtual cell models. Uh when it comes to academic labs chun can comment on wet lab how wet lab academics can survive.
1:21:10 · But from dry lab perspective I do encourage all the you know dry lab AI researchers in universities start to collaborate with industries so that they can get more resources to develop their own ideas. And uh with the era of Egent AI, now everybody can code. So it's more important to have a right taste about your projects so that you don't just let just burn tokens without purpose. Right?
1:21:41 · So we want more academic students, academic professors to have higher taste of research so that we we we make the right utility of HN AIS. As a professor, it's your job is to have taste obviously, but as a student, how do you develop taste as in a world where uh so much of the thinking scientific process could be essentially outsourced to an LLM or you know some there's not a world where you're forced to bang your head against something and learn taste the hard way. This is why we need academic training where you get into a field you know nothing about uh and hopefully after you graduate become the expert about this particular topic in the world right this is why you have to go through different programs to talk to your peers talk to your professors to get a idea about what's a good research taste to to begin with and but more importantly I always teach students that the best way to learn something is to just to program it. By programming, you kind of know what's the details hidden in in all the mathematical equations in the paper which often you emit. But in the era of agentic AI since uh slight different in the sense that we used to spend lots of time coding a little bit time just debugging but now we let the agent do most of the coding but we spend most of time debugging which seems to be in definitely interesting to me and we had lots of discussion with the students in the lab what's the best way to spot uh bugs by by AIS u so How do you find places where AI is particularly good at also how do you find places AI are still limited at is kind of you need lots of China era as well and in the end you still need to kind of validate your model using real world evidence right so that's why collaboration with wet labs collaboration with clinical teams to validate your model provide a feedback signals to to your taste as as we discussed is certainly a very important training program.
1:23:58 · On the wet lab side um academic has extremely important roles to play. I think we'll enter um a field of um an era of symbiotic innovation and cross-pollination of ideas just like in AI field. I think we have great ideas coming out of academia still all the time but industry now increasingly are contributing new ideas on architecture on all of that as well in the web app side certainly industry seems to be able to scale this type of data generation quite well um but biology is so much more than just cell-based perturbse seek beyond RNA seek we would like to measure many other analytes right um proteins metabolites It's lipids um uh protein protein interactions. How do we do that at scale uh beyond uh individual cells?
1:24:49 · We would like to be able to measure cell interaction spatial um cell in their native context or even whole animal level invivo pertubations. Again, how do we do that at scale with lots of great innovation coming out of academia?
1:25:02 · Actually, just one great paper last week. And so, how do we connect all of these together? I think we have many years of work ahead of us to fully crack data generation for all biology and that I think we need the scale the industrialization the innovation from industry we also need that from academia I think will together um move this field to the next level one question that we've been trying to ask everyone is in your field which you could say maybe is AI and you know sort of high throughput experimentation um or however you want to define And if you could wave your magic wand and have a bottleneck removed for you or a pro a key problem solved, what would that be?
Open science and what academic labs should focus on
1:25:46 · Protein.
1:25:46 · Okay.
1:25:47 · If there is a way to do protein sequencing or high through protein measurement, um the same skill that we can do genomics, that will be amazing.
1:25:56 · [clears throat] I'm training genomics field, but if I can do that, I would incorporate that technology in a heartbeat. Um RNA is amazing. It foreshadows which proteins are going to get made but protein by and large are the functional units in a cell. Not only does their abundance matter, their post-transational modification matter, their localization in a cell matter. If we can measure all of those things, their confirmational states, their modifications, their abundances, their localizations at scale, single cell or even spatially, I think such data sets will be incredibly useful to train the next generation of financial models. I know there's a lot of innovations in that direction. Can't wait to see that become uh come of age. My hope is I hope to see a breakthrough in sequencing technology. Not just the reduced cost but sequencing technology that can sequence the same cells at different time points. I think this is much lacking right now because you have in order to sequence the cell you have to kill the cell right. So can we have a technology that can measure the cell states at different time points for the same set of cells. I think that will bring a very different dimension to the data set so that we can start to measure the temporal dynamics of cells. So far everything we measure, everything we model is extremely static. So can we have a technology that measure different cell response at different time points for the same set of cells will unlock massive opportunity to kind of model the dynamics of cells. To me, that is the real virtue cell. That's a really interesting idea. I never would have thought about that. Wow. Would you be okay with even just partial like small snippets of genes or maybe three prime regions of a small number of transcripts? Start with a small uh small set of gene panels to begin with, right?
1:27:51 · But eventually if if since we're talking about the magic one [laughter] here, eventually if we can have a system that observe how cell evolve at different time points and we have enough data to actually model such development. I think that would be real virtual cell modeling. there's uh small attempts in um oh sorry earlier attempts in just um for example sucking out portions of the cells taking almost biopsies from the cell to do you know um a fraction of the cytoplasm measurements so that might be similar to the idea you talk about there's also work from uh public lab to have the cells secrete um little vesicles and they harvest that in the cell in the cell culture media to measure what the cells are you know um producing longitudinal But there hasn't been technology that can let you measure the entire cell transcriptto home while still keeping a cell. You can't have the cake and eat it. Yeah. [laughter] Cool. Uh yeah, thank you for taking the time to chat with us. It's been great.
You can't have the cake and eat it—currently, to sequence a cell, you have to kill it. The real 'magic wand' for virtual cells would be temporal sequencing of live cells.
1:28:53 · Learned a lot. It was a lot of really interesting discussions and I think and I especially really appreciate your commitment to open science and all of the cool models and um data you've released. Is there any last like thoughts you have or anything you'd like the audience to know um follow up with?
1:29:08 · Overall, I think virtual sale is such a new and fastm moving field. We hope to have more and more people join us and uh our Excel paper is out and uh we look forward to receiving your comments and feedbacks and also we're hiring. Yeah, we are always looking for talented engineers, technologist, biologist, drug hunters, AI uh scientist, computational biologist. So look on zera.com, look for uh the open roles. We'll be happy to chat with you.
1:29:37 · Thank you. Thank you very much.
1:29:38 · Great. Thank you.