What happens when AI stops merely *reading* DNA and starts learning how to *write* it? Eric Nguyen, CEO and co-founder of Radical Numerics and a key figure b...

Transcript

What Is a Genome Language Model?

0:00 · The design side is going to get more capable.

0:02 · The defensive side needs to try to get ahead.

0:05 · So I think inherently there is this arms race style dynamic that the defensive side has been far, far lagging. And so what we want to do is bring the defensive side to par essentially. We felt it was important as a lab that a team that was both building the design capabilities is actually also best suited for building the defense capabilities because they're basically the same models.

0:28 · A model that is good at generating, turns out is also very good at discriminating or predicting if a sequence is pathogenic or not.

0:36 · For us, we as a company thought it was very important to have a dual mandate.

0:41 · It's this idea of essentially being cognizant and feeling responsible for the capabilities that we're enabling on the design side.

0:49 · So if we're going to create models that can design function into sequences, we believe and see a gap in companies being able to safeguard that technology.

1:00 · Welcome to Lane Space. I'm Brandon.

1:02 · I build RNA therapeutics at Atomic AI.

1:05 · I'm joined by my co-host RJ Honake, CTO and co-founder of Miraomics.

1:09 · Today, it's a pleasure to have with us Eric Gwynn, CEO and co-founder of Radical Numerics.

1:16 · Eric started-- got his PhD in Chris Ray's group.

1:20 · He spent a lot of time thinking about how to do long context genomic models before long context or genomic models were cool.

1:28 · He was the first author and I think basically visionary behind the Evo generative model, one of the first generative genomics platforms developed Evo 2, which naturally led into Radical Numerics. Thank you for being here.

1:44 · Did I miss anything?

1:45 · That sounds great. Cool.

1:47 · (Laughter) Welcome. Thank you.

1:49 · So Eric, let's talk about Omni and the blog posts that you guys did about the benchmarking. But I want to hear first, OK, what is a genetic language model?

2:01 · Why do I care? What does it do? And then let's talk about the top line results from the blog post.

2:08 · So a genome language model, or GLM, is a large language model trained on DNA sequences. So very much like natural language and chat bots you see, but not trained on words or natural language, but on the raw fabric of life, which is these sequence of letters that make up DNA. And we ourselves, our company, our team, is known for creating the first generative genomics models, which are models trained on DNA, not just to read, but also write, meaning able to generate new sequences of DNA.

2:42 · And we felt this was an area that was overlooked and that if AI could read and write DNA, it could change a lot, you know, science of discovery and understanding of human health and how to treat it. And so we felt that it was a big opportunity to train AI on the genome.

3:01 · What kind of things can you potentially do with a model like this?

3:05 · Great to start for us when we first started working on DNA models.

3:10 · We worked on this model called Hyena DNA, which is a large language model, but it used a convolution instead of a tension.

3:18 · So a little more technical details.

3:20 · DNA has this property that, well, it's very long, right?

3:24 · At the time, these large language models had limited constraints on context, right?

3:30 · Being able to fit long sequences. And so we were looking for a more efficient algorithm to be able to handle something like DNA.

3:38 · And so we came up with this, what we call the Hyena operator uses convolutions.

3:44 · Long story short, it let us process longer sequences, in this case up to a million, and at the time was the largest context for a language model.

3:54 · And what we did with it was essentially used it to read DNA, predict function.

3:59 · So given a sequence of DNA, string of characters, we predict its regulatory function, its effect on a genome. And this is interesting to scientists because a lot of the DNA in our bodies, perhaps people are less aware, but actually we don't know a lot that much about our genome. We know it obviously encodes the information for making us us and all the different complexities and potential diseases.

4:28 · At the same time, the grammar rules about how the combination of those letters are sort of formed, what they encode function and traits is not fully understood.

4:38 · And so the hope was using these DNA models, language models to be able to map some function from the raw DNA sequence.

From Reading DNA to Writing It

4:46 · And so our first generation models was able to show that yes, we can train AI to be able to read and understand to some degree, DNA sequences, and especially what we call the longer range interactions, meaning over sequences, if you use chatbots, for example, if you've heard the phrase context rot, the longer the input you put into a language model, it starts to deteriorate.

5:11 · And so being able to pick up long range information and sort of patterns, motifs, grammar over long sequences was what we were trying to accomplish.

5:21 · And so we showcase in that first generation of hyena DNA that that was possible over a million context. And then really what started the field now known as generative genomics was the model called Evo.

5:35 · And Evo, we should try to showcase there was this idea of not just reading DNA, but being able to generate it. And so we wanted to accelerate essentially how biologists and scientists have learned from biology and particular genomics.

5:50 · And we felt like this whole field of generative AI being applied to language, great, accelerated the, obviously our understanding and ability to manipulate the natural language. But here's this other language DNA in the genome that we don't understand.

6:07 · And it's barely being applied with AI in our pit at the time a few years ago.

6:12 · Models that we saw were really small, short context.

6:16 · So they can only pick up small patterns and limited context.

6:20 · And none of them generated DNA. So they all just would read.

6:24 · And we felt that the idea of generation was so powerful and transformative in natural language. What if we could bring that to biology and DNA particular?

6:34 · What can you accomplish that you can't do in a lab?

6:37 · Right? So what does that unlock for you if you could do that very well?

6:41 · I think one of the first things that we showcased that got folks sort of intrigued by the potential of this was a CRISPR-Cas system. So it's an enzyme that's able to cut DNA itself.

6:58 · And I think what was particularly enabled by the EVO models was the ability to generate over not just one modality or one type of sequence, but spanning multiple modalities and spanning multiple scales.

7:12 · So CRISPR-Cas, it's a molecule made up of both RNA and proteins. And so at the time you hadn't really seen models that can generate multiple modalities. They had protein language models that can generate proteins.

7:30 · Sometimes you had RNA models that can generate RNA, but you didn't have a single system to sort of co-design. We showcased that a single DNA model, sort of the foundation of both of those, right?

7:42 · From DNA you can get RNA and proteins that you can design a single system to generate and also function in the real world.

The First AI-Generated Genome

7:50 · So we asked EVO, we showcased it a bunch of natural CRISPR-Cas systems and essentially asked it, can you make a new one?

7:57 · And we're able to sample from that model that we trained.

8:01 · And indeed we showcased that EVO was able to discover a new CRISPR-Cas system and folks were intrigued by it and cover of Science Magazine.

8:09 · And later you get to give a TED talk about the work, which is interesting because obviously the audience is very general and so trying to make, what is CRISPR-Cas system? What is DNA and how do you generate it?

8:22 · Why would you generate it? All sorts of fun topics.

8:25 · And I think the more even intriguing, exciting thing that scientists eventually just last year showcase what you can do with a generative DNA model was to generate the first genome from scratch using AI.

8:38 · So this is something not possible by humans, right?

8:41 · Humans usually, you can think of it like copy and paste parts of other genomes or other DNA, put it into something else, but they would just take out small motifs that they know the function and they understand the rules, but to build something from scratch, you're gonna ground up, had not been done before at the genome level.

9:01 · And so EVO turns out was able to generate a functional genome.

9:05 · And in this case, it was known as a bacteriophage, also known as a virus.

9:09 · And this was a key turning point for, I think, the scientific community and for us as a company, a radical numerics, because we felt this was such an indicative manifestation to create a whole organism, not existing in nature, but also the potential harm that that means as well.

9:28 · If you can control, if you can manipulate the fabric of life itself, control its function, what kind of implications does that mean?

9:36 · What are you enabling into the world?

9:38 · And so we actually got a lot of feedback, a lot of comments, a lot of outreach from folks, both excited and concerned about this capability, this kind of capability and just the trajectory. This is the early stages, the first thing one can sort of project and imagine what this could lead to.

9:57 · And so we felt as a company, it was important to not only push on the biological design capabilities of these models, but also the ability to use them as defensive tools for the potential of misuse and biological risk that emerges.

10:11 · And I think indeed a lot of companies, a lot of Frontier Labs, are also being concerned about this emerging risk AI models be capable of designing biological sequences. And at the same time, it's being mostly attacked from like a natural language standpoint, like safeguards and things like clod, if you talk about viruses, it'll just like shut you down, which is great.

10:34 · I think it's to some degree, you need safeguards at the natural language level, but I think what you also need clearly is safeguards at the biological sequence level too. So you need models that not just can understand language the trajectory of your chat, but to understand the substrate itself is the next step in ultimate limit.

10:55 · If you can have models that can understand if the sequence is pathogenic or virus, that's the level of defensive capabilities that you want.

11:04 · And then being able to push that out into surveillance systems, national security, of being able to monitor emerging sequences in environment, this is what that kind of capability makes possible. And then we think bringing this frontier technology to that community as well is also important, just as important as using this for human health, which is what we primarily focus on.

Why Omni Is Different From Evo 2

11:28 · There was this evolution that you mentioned, there's the Hyena DNA, and then there's Evo, Evo2, and now Omi. Can you just talk about a little bit about Evo and Evo2 who struggled to beat specialized models across many tasks, whereas in the blog posts you talk about how, across a wide range of tasks, that Omi is actually able to outperform them now. So can we just talk a little bit about that?

11:56 · So yeah, Evo was intriguing to folks in many ways, showcased the potential for applying to multiple types of modalities, but it's still in many ways underperformed sort of the specialist DNA models, especially on human genomics.

12:15 · And so although Evo was competitive, it still wasn't state of the art or kind of pushing the needle. And so some parts of the community thought like, why use a giant LOM when I can use these smaller, more specialized models?

12:29 · And so what we wanted to do with Omi was amongst many things, but one of the first things was to showcase this idea of mid-training and post-training or broadly alignment. So the way we think about language models and bio and genomics in particular, is that mostly we've only seen base models trained.

12:48 · So they're pre-trained, but they're basically unaligned.

12:52 · So in the analogous space for natural language, it's like you're doing all the pre-training, but to make it actually useful in the real world and answer questions that users actually want and that is in the form factor, they find actually informative, there's a bunch of alignment and post-training and mid-training done to get the models to be production ready and actually useful.

13:17 · And so we felt Evo was just showcasing the potential of that pre-training, but Omi is a step of actually making it useful for folks like scientists.

13:26 · And so we spent a lot of time on alignment and mid and post-training, which is essentially showcasing tasks in the form that people generally would want to understand about genomics, given a wild type and a mutation sequence, help me find the causal variant. These types of questions and form factors for how you might want to analyze genomics doesn't just emerge necessarily easily on its own from pre-training.

13:54 · Pre-training is this next token predictions task or infilling.

13:58 · And basically I think of it as like the raw pattern making ability that you're teaching it is in the pre-training, but then taking those learned embeddings or features and pointing at specific tasks or a bunch of tasks really and aligning it, meaning have it show you the output in a way that's meaningful to you, takes a little bit of teasing and manipulating that so far, like the front-natur labs are the ones that drive that research in the natural language community.

14:28 · And so we wanted to bring a lot of that research and more to genomics.

14:33 · And so I think Omni is really just a preview to showcase that potential.

14:37 · And I think once you do that, even just a little bit, we were surprised that it did start being state of the art and pushing the boundaries, not just being affected of multiple tasks just broadly like Eva was, but actually pushing the frontier of each of those area of variant effects prediction, causal mutations for disease, it could start actually being useful for human genomics.

15:02 · And so we're really excited to share that with folks.

15:05 · And it was just a preview in that sense because we were still actively training and incorporating additional techniques into the model, like additional modalities, but I think we were basically too excited and we wanted to get this in the hands of folks faster. And some of the feedback we got in early interest, lots of hospital systems, nonprofits that have tons of genetic data.

15:30 · And for example, they know there's some kind of condition or symptom for the patient, but they can't figure out which parts of the DNA are causing it.

15:40 · So these got these VUSs or variants of unknown significance that we're extremely excited to apply these models to and actually help diagnose a lot of these patients is one example of a real use application.

15:53 · I'd like to talk more about the applications in a bit, but I am curious just from a technical standpoint, what does it look like?

16:00 · What does it mean to align a genomics model?

16:03 · I can imagine with large language model, there's a sort of natural chain of thought, RL, there's like a clear paradigm here.

16:09 · I'm curious, what does it look like for a genomic model, which is not maybe better described as something like fine tuning in a different language?

16:18 · Yeah, I mean, I think in many ways, one could describe it as fine tuning, but then introducing, I'd say, sort of the key components are the right structure of the inputs. So feeding them in a certain sequence so that the model is aware that a certain task is being asked of it.

16:35 · So there's a mix of special tokens to basically, you can think of it as like, if you're gonna do disease prediction for disease A, expect this special token, right?

16:45 · Just kind of like a way to prompt it.

16:48 · If you expect it to do design, have another special token, and then showcase the examples, kind of like in a chain of thought manner, meaning showcase a sequence of desired outputs and the trajectory of it.

17:01 · This is a little vague sort of intentionally because it's part of our secret sauce that we're still developing.

17:09 · And I think over time, we wanna showcase more and more of it.

17:13 · But in many ways, it does mimic a lot of the natural language community.

17:17 · A lot of it is fine tuning, but really it's also carving out specific datasets that we want it to focus on and then structuring the questions or tasks in specific ways, as opposed to pre-training. Pre-training is really just feeding it in and everything and really, and just doing next token prediction or mask infilling, if you're doing mass image modeling. And it has no sense of like this Q and A type structure where you have a question, a prompt, and then an output.

17:46 · And you could think of mid and post training as starting to showcase, given this type of input, I expect this type of output, whether it's score, prediction score, or design is largely the mid and post training.

17:58 · And post training also includes things like reinforcement learning too, but I think the bigger steps are introducing structure of like question and answering.

18:08 · Do the models have multiple heads that are task specific, or are you training one set of heads or whatever that can answer multiple questions at the same time?

18:20 · Just change the input tokens or whatever.

18:24 · Yeah, broadly, it's a little, you can, I think we're flexible on this.

18:28 · The idea, yeah, sometimes you can use different heads, but the idea for us is to unify more so.

18:38 · And so I think early experiments, we did have different heads, but in some cases the different heads do better. In some cases the unify, a single head, a single model does better. And so I think we're flexible on that, but I think broadly the direction that we are moving toward is a single. And the reason for a special motivation for that is that we're trying to unlock a lot of generalization.

19:05 · And I think the more unifying we're able to make these models, I think that's when you see more emerging capabilities happen.

19:14 · And in large part, that's what motivated the DNA work.

19:18 · We felt like a lot of models were specialized into other modalities, a little more downstream from DNA. So like RNA or proteins or molecules.

19:28 · We largely felt DNA is the foundation and that from DNA, you can learn a good deal of other modalities, potentially all of them. And I think other modalities is a sort of like additional context that you're showing the model.

19:45 · That's how I kind of view it philosophically in my head.

19:49 · But yeah, the idea that single models unifying across modalities, scales is what the lab builds toward.

19:56 · Can you walk through a few of the tasks that you talk about in the blog and just explain and remembering to narrate for the listener only audience, but talk about some of these top line results and maybe dig into them a little bit.

20:17 · Sure.

20:17 · One of the areas that we thought was really interesting for showcasing in this particular release for Omni was on understanding variants and their effects, which are essentially in DNA, a change in a position, changing the letter of one of the other three letters in your genome in DNA, sometimes can cause a disease and sometimes it doesn't do anything. Actually many times it doesn't do anything, but there are specific areas in your genome that if you have a different variant or a different letter there, it can cause disease.

20:54 · And in many cases, because the combinations of these changes in the genome, over 3 billion letters, right? It's so vast that for clinicians and scientists, we actually know only a very small portion of which variants are causal to disease or not.

21:11 · And so there's these benchmarks from folks who collect variants, different hospital systems and clinics, and some are known and some are still unknown.

21:22 · Folks have created some benchmarks from ClinVar or TraChim to basically the idea is given a mutation in a DNA, can you tell if it's gonna cause a disease or not?

21:33 · And so this is a very good setup for DNA models because they're probabilistic. And so when you do make a change, it basically can modify its confidence or their probability of predicting a next letter.

21:47 · And in this case, we've sort of leveraged that predictability of these models.

21:53 · They've essentially seen and been trained on so much DNA, in particular human DNA, they kind of understand what's common in such and usually common or conserved across different other folks usually typically means healthier.

22:07 · And so if it's less common, you can think of it this way, it's less common, the model can pick it up and sort of predict that it's potentially pathogenic or disease causing. And so we've taken some of these benchmarks and when you introduce a variant or a mutation, there's different types of mutations in variants.

22:26 · So sometimes you can delete a letter altogether, you can just flip it, you can remove big portions of the DNA, but these are generally single variants.

22:35 · And in this case, this is where previous DNA models really struggled, especially on humans.

22:42 · And so we showcased that not only is it competitive or capable for humans, but in many cases, it's state of the art and actually most of these cases, it's state of the art. And I think the exciting part is that areas where other models that were currently on the frontier, they still were lagging behind quite a bit in terms of where in the genome. So in the genome, there's coding and non-coding regions, like protein areas, protein regions.

23:10 · Meaning areas that are actually the coding, the structure of a protein versus other areas that do other things like regulate what genes are expressed.

23:20 · Yeah, and so these non-coding regions are largely regulatory, they kind of control how much or when to use a certain gene or turn them on.

23:28 · And in many cases, these non-coding regions, these regulatory regions, variants, their mutations there are much harder to predict if they cause disease or not.

23:38 · And I think what's exciting thing about the new generation of models we're building with Omni is that that's where we shine, especially.

23:46 · The models are able to pick up mutations and be able to distinguish if it's disease causing in these non-coding and especially long range areas.

23:55 · So I think this is particularly exciting to a lot of geneticists that have struggled to use traditional bioinformatic tools or statistical methods because they've largely focused on coding regions, which is only about 1.5, 2% of the genome.

Predicting Disease-Causing Mutations

24:09 · And turns out many, if not most of the diseases are in these non-coding regions.

24:14 · And so there's been a real desire to build models that can actually pick up these variants of disease causing variants in these non-coding regions.

24:23 · I have several questions. First, just while we're here, for the listeners, there's this column on this, a benchmark chart called Borsoli, reference number four in the blog post. I think for maybe some historical context, could you talk about what this column represents and maybe this also helped give context for the Omni column on the right.

24:48 · Yeah, great point. So what we show in this benchmark here is really taking some of the representative models or the strongest models in the deep learning side and also in the traditional methods. So we have Evo2 is the latest previous genomic model that our team had worked on. And then Borso is, it's also a DNA model, but a very different kind. Essentially it's a supervised model that predicts from DNA functional genomic tracks. So it too is inherently multimodal, but it's not a language model.

25:18 · So it doesn't predict like a next token prediction.

25:21 · It goes from a DNA sequence directly to a functional genomic track like chromatex flexibility or gene expression.

25:27 · And these tracks have been annotated extensively by people writing their dissertations and whatever.

25:34 · Yeah, right. So the big difference there is that it's a supervised task, right?

25:38 · So it requires labeled outputs. And in our case, these language models, they do not require labeled outputs, right? So you're doing raw pre-training on unannotated sequences. So why is that desirable?

25:50 · Well, it's a lot more data, a lot more genomic data that's not annotated.

25:54 · Actually most of it, pretty much in many ways, almost all of it is not annotated.

25:59 · And so being able to learn from an unsupervised manner, hugely desirable, right?

26:04 · For us, we wanted to showcase the benefit of pre-training on raw genomic sequences and comparing it to state-of-the-art models in other spaces in DNA.

26:16 · And so your prediction is when you say it's unsupervised, how does the unsupervised property work? Like how do you convert the output of whatever your model is to an actionable like ranker, score, whatever?

26:31 · What I talked about before is, to train the omnimodal is pre-trained so it's unsupervised, but when you're pointing at a specific task, there is a supervised step. So it's just taking a smaller dataset that is labeled.

26:45 · But essentially what we're doing is using the likelihood scores.

26:49 · So the raw outputs of the language model, which basically you can think of it like a probability for predicting what the next letter is.

26:57 · We can essentially showcase the probability score, the likelihood score for the mutation versus the wild type. So that's seen in the reference genome versus the mutation in this particular case. And then we'll have two scores.

27:12 · And then you can think of it as like using a ratio of the two to showcase basically how different are you from, how different is this mutation from a normal or baseline basically. And then that's, you could think of it as like a surprise factor that the model is able to use and leverage and then we can use that to essentially score an actual prediction for disease or not.

27:35 · That make sense? Yeah, so omnimodal regressive is that, or is it diffusion or something else that you can't tell me?

27:43 · (Laughing) Yeah, this time we're not describing the exact makeup, but Evo was autoregressive. I think that was the first large scale autoregressive.

27:55 · And so I think for us, we don't tie ourselves down to a specific training objective.

28:02 · We use every tool in the box, toolbox essentially.

28:05 · But for the specific benchmark, you're going along and you're just using the likelihood distribution of the tokens.

28:13 · And some tokens are, the model thinks these are unlikely and that is probably because some evolutionary constraint.

28:22 · This doesn't show up often. And because it doesn't show often across genomes, it is probably going to cause problems and people will not survive, so on.

28:32 · So you think that is basically your ranking metric or something.

28:38 · Yeah, it's one interpretation of how the model is thinking about it.

28:44 · And very similar to in natural language, you can describe the same kind of paradigm.

28:49 · And there is additional case in mid and post training to leverage more than that, I suppose, because we can teach it specific structure and benchmarks so that it can build on top of what you just described, which is like what's common in nature, but also because it's a specific task for disease variant prediction, then the model has additional training, introduced during mid training to showcase, to add additional learning power essentially.

29:18 · What are some examples of that? Again, probably secret sauce to some extent, but can you give just a gist of what that looks like?

29:24 · What are the kinds of things you would throw in there?

29:26 · We would actually throw in the score to itself.

29:29 · So, you know, that I mentioned a ratio, it's a ratio of wild type risk mutation.

29:34 · I would say that's more of a zero shot method where you don't even have to do any mid training. And that's what Evo2 is doing in particular in this column.

29:43 · So Evo2 is not fine tuned essentially.

29:45 · There's a EVE column, which people are basically fine tuning, using them scores from Evo2, that's from GoodFire. And that's also, you know, folks that we greatly expect and they kind of showcase that these models are able to be stated art when you can fine tune them as well, not just zero shot.

30:02 · And then our model is introducing sort of a step about that, not just fine tuning, but also introducing structure into, by structure I mean the format of these benchmarks into the model itself, that Q and A style formatting during mid training, which is what gives us an extra boost even.

30:19 · I see. And extra boost, but I think the other benefit too, that, you know, we didn't emphasize too much in the blog, but I think it's really convenient for practical use for scientists is that you don't, you're doing this without taking the embeddings and then tapping on ahead and then, you know, doing some regression, which is the extra step, it's extra hurdle. Can you imagine chat to PT, if like every time you asked a question, you had to like fine tune it for a certain domain.

30:47 · We've done it so that the model is flexible during mid training to be trained on many tasks at once. And so the, that fine tuning, that last step of training the embeddings doesn't have to be done.

30:58 · It's out of the box at that point, you just prompt in a certain format and it will, that sort of format tells it which task you're gonna do.

31:07 · And then we'll output the answers in that, in the desired format, basically.

31:11 · How careful were you in designing this post training scheme to avoid kind of data leakage with the ClinVar, Cray-GEM, RNA-GEM and so on, like these data sets?

31:22 · Like how confident are you that there is no data leakage either like accidental or something upstream and that, because I would not be surprised if a lot of these sequences showed up also in your training data, even in a unsupervised sort of way.

31:39 · Yeah, the short answer is we're extremely cognizant of the risk of data leakage and extremely hard to not mislead or be careful.

31:49 · And so we have bioinformaticians that are able to basically comb through the data and curate, dedupe and align sequences to make sure that things that are similar potentially to what's in the benchmark are not there.

32:02 · If they are there, we remove it. And so, yes, we actually have steps to QC the data quite extensively.

32:09 · Yeah, cool. Yeah, I guess maybe before we move on, I think it's really cool seeing that there are these supervised methods, which previously several of these numbers were, let's say, within the error bars, if not just straight up getting what came before them. Now you have significantly improved upon that.

32:26 · Yeah, we're super excited. And I think for the longest time, there's this area of this other method called CAD, which has been state of the art.

32:33 · And state of the art for a reason, which is sort of by design, they'll take the best methods and kind of do an ensemble, right?

32:40 · So they'll take another, even if the best method is another previous model, they'll mix it with like an SVM and just like throw the kitchen sink at it.

32:48 · And so you can see why it would be the best, right?

32:50 · And so that was the bar for us. We're like, if they're gonna throw the kitchen sink at it, like we're not gonna cherry pick one model and say we're better than that.

32:59 · We need to be what's possible, humanly possible now, like across everything.

33:03 · And so, yeah, our researchers were setting their sights on that to see if they can actually improve performance across every method.

33:15 · I'd be interested to see there's some discussion of chain of thought.

33:20 · And that broke my brain a little bit when I was first hearing about that. I'm really interested to hear about what that even means.

33:32 · I had to pour through the blog post to really understand that.

33:35 · Yeah, yeah. I think this is really just a taste of where we think the design capabilities can move toward and way more usable for folks.

33:45 · So chain of thought stems from natural language community.

33:50 · I believe Jason Way at OpenAI showcased the first examples.

33:54 · And really what the breakthrough there was showcasing that these language models performed better when you just show your work, essentially.

34:02 · You showed the steps of how you came to a conclusion or an argument.

34:06 · And it turns out, even if they were like simple steps, but it just gave the model a chance, maybe it's sort of like philosophically who knows exactly why it works.

34:17 · But essentially feeding more tokens in and giving it more scratch space to think.

34:22 · And so people just think this is the beginning of reasoning for these language models, this ability to kind of get to an answer by thinking to itself, by itself.

34:35 · And so it seemed quite successful in language, very successful.

34:39 · And that's why you have a lot of agents that just spent tons of tokens, right?

34:44 · Just showing its work, right? In many ways it stems from this chain of thought paradigm.

34:50 · And in biology, we saw very little of that.

34:53 · We started to see some of that in protein design a little.

34:58 · And so we wanted to push that and showcase that, well, at first explore it, is that possible in DNA? Like what does that even mean in genomics?

35:07 · Because you don't really have words that describe its thinking.

35:11 · So how do you take that same paradigm and introduce it to a DNA language model?

35:15 · And so what we did was a simpler version in so many ways.

35:19 · We had this dataset of RNA aptamers.

35:21 · So just think of it as these desired sequences with some kind of fitness score associated with them. So we took this large dataset that had RNA input and a fitness score. The fitness score go high, it's good, right?

35:34 · It's the simplest version. It's a big dataset.

35:37 · And so what we wanted to showcase was that if we show the model progressively better RNAs in a series of steps with its score, right?

35:45 · So you have like low scores first and then you gradually move up the chain.

35:49 · Can the model continue that trajectory on its own?

35:52 · And then in the final step, does it self optimize to a point where it's like the best score it can get? That was the experiment.

36:00 · Can we do that? And so we took a dataset, a large dataset of aptamers.

36:04 · We held out a portion of like the best performing ones and we showed it only the lower ones, but then we ranked it, right?

36:12 · So we showcase lower scores with the RNA aptamers and then progressively got higher and then asked the model to just like continue with that pattern.

36:20 · And it turns out it was able to recapitulate some of those higher scores that we had not shown it yet. We were actually in the process of validating the wet lap right now. So we didn't get to show it here, but we wanted to know, right?

36:35 · Actually, can it not just do this in silico, which it can, it showcased that it was able to continue this trajectory and create design plausible aptamers with higher fitness scores. And now we think this is, you know, obviously, if this works in the lab, we think this is a hugely, hugely valuable paradigm that can be pretty much applied to every other type of biological sequence.

36:58 · That's usually what you have. You have sequence, you have some kind of fitness score, desired output. And if we can get models to eventually learn that structure and, you know, basically just show a series of progressively stronger sequences, the model can then predict the rest. That's a very powerful paradigm.

37:14 · Honestly, a bit surprised about this, that specifically this task saw a strong improvement. Maybe my personal bias is coming in here, but RNA is somewhat notorious for not having good co-evolution data in terms of like concerning structures for viral genomes, oftentimes there is strong evolutionary pressure, but for mammalian, usually RNA does not have evolutionary pressure.

37:43 · And I think the community has seen that very, very clearly the predominant pressure is like RNA will code, you know, will carry information, coding information.

37:56 · So, I mean, I'm wondering like genomes carry lots of different information.

38:02 · They code for proteins, they have regulatory elements and, you know, different types of genomes have different types of structure.

38:11 · So I'm wondering, where do you think this capability might have emerged in this language model?

38:18 · That's a good question. And honestly, I'm not sure.

38:21 · Like we're surprised too. One, because the model is pre-trained on DNA and it's really just like mid-trained on RNA, very, very, you know, in a small way.

38:33 · Yeah, I think what your intuition about the DNA having a lot of evolutionary effects or information is probably where. And so I think this is hinting at the idea of why we think it's so important to pre-train on DNA genomes and genomes first, and then sort of add additional modalities on top because you get a lot of transfer and you want the modality generalization is the thing that we're working toward.

38:57 · You know, what I didn't talk about as a company for the company as well, is this idea of a building toward general biological intelligence where we are unifying a lot of the different so-called like languages or modalities of biology.

Can DNA Models “Reason”?

39:11 · At the end of the day, they stem from DNA.

39:13 · And I think people have not exploded that fact as much.

39:17 · It's usually really specialized, the domain specific modality specific models and not leveraging a lot of inherent shared structure from other modalities.

39:26 · And so one example of that is like, you know, when people talk about virtual cells, for example, there's a little bit tangent, but, you know, they tend to focus on just RNA and, you know, transcripts, and they want to generalize to describing an entire cell.

39:42 · But obviously a cell is a lot more things than that.

39:45 · In my mind, if you want to learn a system, you want to learn from all the signals or sensors of that world or that system, you know, if it's a cell and you want to fuse, you want to understand the DNA, you want to understand the metabolomics, the epigenomics, proteomics, and that's when you get closer to like, quote unquote, a virtual cell. And in our minds, we don't even want to stop at just the cell, but we want to fuse all of these sensors across all of biology.

40:14 · Will it get us to a super intelligence understands every component of my new show of biology? Who knows? But I am confident that this type of paradigm will get us a hell of a lot further than we are now. Like that's my bars.

40:27 · Like, can you make something far more useful than now?

40:30 · So I'm curious, are you focusing on eukaryotic cells?

40:33 · Are you focusing on like human genomes?

40:36 · Have you gone so far as to do viral genomes?

40:39 · I mean, there's a lot of DNA viruses, but it seems like possible that RNA, that there's a lot of RNA sequence virus sequences out there.

40:48 · And I'm not sure fundamentally they would be much different in terms of training.

40:53 · I'm curious, like what's the scope of that, if you can talk about it?

40:58 · Yeah, absolutely. We are interested in all domains of life.

41:01 · So here we focused on humans in particular, because we thought this was an area that of previous models, Evo and Evo2, were not as strong and sort of got a lot of feedback from folks asking, what are these models useful for?

41:16 · They can't understand human genomics because it's too complex of grammar and rules and DNA. It's too noisy, it's too, there should be repeat characters and all that stuff. So we wanna showcase, we think this is actually useful and it can be applied to humans. And it's sort of the most complex of the complex in some ways.

41:36 · But I think there's opportunity to apply these models generally to every form of life. So we absolutely are interested in prokaryotes and viral.

41:46 · Viral in particular, we care about it, especially for biodefence and biosecurity especially. And I think there are also lots of therapeutic applications that we can learn from microbial life, maybe obviously for some folks.

42:00 · In particular, I mentioned that folks had used Evo to generate the first AI genome, a bacteriophage. Turns out you can use bacteriophages potentially for AMR or antimicrobial resistance. If you have a superbug bacteria infection, which in the world is about 2 million deaths from bacteria infections still, the idea of using viruses, designed viruses to target specific bacteria has been done for a long time, particularly in Eastern Europe.

42:26 · And there's a potential to make a new class of antimicrobials that are not like antibiotics, but very similar that can be used just like it. And so I think we're gravitating toward things that are high impact and the potential to save lives. And so we don't stop at just one type of genome.

42:47 · I think we're interested in anything that's beneficial as you humans.

42:51 · Maybe going back to my question about SOLEX and RNA, predicting kind of a chain of thoughts of RNA evolution. I'm curious, was this model trained on RNA sequences or sequences which have some, might have evolutionary pressure on RNA structure?

43:09 · Only during mid training.

43:10 · Only during mid training.

43:11 · So pre-training is all just genomes and DNA.

43:14 · And so the only time we introduced RNA was for this specific task where, and only RNA from this dataset. So not even outside.

43:24 · I see.

43:26 · So this really was something along the, there is something encoding RNA structure in this model to some degree, maybe, or either that or implicitly. Yeah, implicitly, I would say implicitly.

43:40 · I mean, cause I'd say, you know, sequences, as we know from proteins, like implicitly should learn structure from the sequence.

43:49 · Yeah, and so we don't add in 2D or 3D information at this point, but we absolutely plan to.

43:55 · (Both Laughing) So just so I understand, first of all, the chain of thought ideas, this is a demonstration of it, but the idea is that anything that you can get sort of a training set that has a sequentially better measurement of some sort is maybe a candidate for this technique.

44:16 · Yeah. And so can you just describe for this particular experiment just so we can understand how we're mapping chain of thought to this dataset and what the datasets have to look like? Can you just describe how is the, for this, I know this wasn't your dataset, but how is the data collected in such a way that you could map accurately from sort of fitness or whatever to a particular phase or part of the dataset?

44:44 · Yeah, so I'm less familiar with how the data was actually sent us or generated from the experimental point viewpoint, but they are validated from a wet lab in the real world when it was collected.

44:57 · So it has some kind of fitness score, I believe through-- I can maybe provide a bit of context if you want.

45:04 · I mean, so the idea here is you just generate a bunch of random sequences and then you take those sequences and so you have like an aptamer structure, which is essentially like a switch with RNA, which sort of, when something binds to it, it will do something like cleave off a sequence.

Toward General Biological Intelligence

45:20 · And you can use an NGS, like Big Generation Sequencing readout, I'm very high throughput. So you create lots of these different sequences.

45:28 · I guess in this case, they were targeting a specific HIV, like protein or genome or something.

45:36 · And if it binds, you basically get the signal of like you get more reads of that.

45:41 · And so the more sequences which are floating around kind of like the more fitness, the more likely it is to bind and then you take those and then you mutate them again and you kind of iterate on this.

45:54 · Right, so you have this iterative experiment where you're progressively using the Petri dish to basically identify the most fit things.

46:02 · So and then what you're doing here is you're basically doing this same experiment in Silico.

46:08 · And they're actually doing very high throughput.

46:11 · Like I think there's like 10 to the 11 or something sequences, some like really high number of kind of sequences explored in parallel for this experiment.

46:18 · But the key is the biology of the experiment is actually doing the filtering, right?

46:23 · Yes, yes. And so you can imagine other types of experiments where you could apply the same kind of idea where you're doing this like progressive refinement of something which is very common in biological lab work.

46:36 · And so if you're capturing those intermediate states and you can maybe feed them into the model and do that kind of thing.

46:43 · Absolutely, yeah. So another example is for the antimicrobial resistance.

46:48 · I think that's something we're very interested in as well.

46:51 · And the ability to selectively target specific bacteria strains or kill bacteria strains. Yeah, that's measured basically a score zero to one of how effectively that is done. And so showing progressively more effective phage genomes and their associated scores for effectiveness fits that paradigm very well.

47:11 · There's another design task that we're working with a national lab to do this with as well. And an area that is quite different for us, but it's on designing proteins to extract rare earth minerals.

47:25 · And so it turns out that you don't just care about proteins that could bind to something, but you want it to be selective.

47:36 · You want it to bind to one type of rare earth mineral.

47:40 · And so you have scores associated with how much affinity or binding affinity for each type of mineral. We wanna essentially do the similar exercise with rare earth minerals and show it a series of progressively desirable scores, not just for binding to this, but like lower binding scores for other ones.

47:58 · You can selectively do it. So I think the creativity in which you can showcase sequence and desirable sequence with some kind of fitness or functional output is a relatively intuitive way for people to design just prompt by just prompt engineering essentially, which I think is very exciting to explore more.

48:18 · Following up on the rare earth mineral extraction, I find that is an interesting use case for this. In fact, maybe one word you probably wouldn't have a comparative advantage compared to some other techniques because it seems like your strengths are probably in larger scale design across like entire organisms, but focusing on individual proteins, that may be much more of a structural task.

48:43 · And I'm curious, like going to talk about ClinVar, the argument here is that the model now is a really good statistical representation of what type of mutations are common or not common. And I think in order to do design as a, if you want to design structures, I think you want to understand structure.

49:02 · If you want to understand disease, I think that is I think more natural, for many diseases is much more natural in terms of like a population genomic sort of way.

49:12 · So I'm curious, like where do you think your strongest competitive advantage is, in using this strategy? And do you think that your model understands structure in addition to function?

49:24 · Yeah, I'd say what drew us to this particular application rare earths and our strengths in general, why we thought it might be suited for it, is two things.

49:35 · One is I think in areas where context matters.

49:39 · So yeah, you're right. So a protein design task where you're just designing maybe not naturally where we see ourselves competitively advantaged, but in the rare earths case, our hypothesis is that context matters.

49:54 · And what matters here possibly is certain microorganisms with proteins that have the function that we desire, we can potentially prompt and provide us context for, this is the neighborhood in which, tell Omni, this is the neighborhood of genomes or microorganisms that you should search for new proteins. So it's really sort of like a mining exercise.

50:21 · And what we can showcase to our models that other folks can for like protein structure models, is that we can feed in non-coding regions before that gene of interest or that protein of interest, and then ask for the model to provide variance essentially. So like mutate this protein, but know that you're in this microorganism, but show me different variants that you've seen in nature or combine different things that you've seen in nature, given this context.

50:48 · And I think that is one reason why we can make very evolutionary diverse sequences and potentially phages. Folks have used Evo models to do this with toxin antitoxins and a similar technique. They've prompted on things upstream from the proteins of interest, and then asked the model to kind of generate a bunch of plausible other ones.

51:08 · And I think this case for rare earths, that's super exciting because then we can come up with plausible variants and then test them relatively simply for what they bind to and selectively bind to, which I think is really interesting for this partner who cares for example, about securing the supply chain of rare earths for the US strategically. And so we thought that's absolutely worth something that is worth supporting. So your point is not just you're designing a protein, but you're designing an organism which generates a protein and this protein has an action.

51:40 · But it's that the way to design this is you need to understand how through the tree of life interactions of proteins with rare earths has, or with certain minerals have occurred.

51:56 · Yeah, I'd say in this particular case for rare earths, we're not really interested in the organism, like designing the whole organism, but I think the organism does tell us about what proteins are plausible and the way they evolve, not just around the protein itself, but also the regulatory elements around it.

52:15 · And it helps us narrow and provide additional context.

52:18 · I guess I think of the genome or DNA broadly as sort of the imprint of the physical world into DNA. And so there's parts of the DNA that we want, but the things around those parts that we want also tells a little bit about the context of where it came from and how it came to be in its function.

52:36 · And so I just think of it, can we come back to that word, but context, I think context matters in many of these applications, or at least in some of them.

52:47 · (Laughs) Yeah, I think context matters a lot in biology.

52:50 · I think maybe one of our big bottlenecks is the lack of context and how we as humans mostly approach biology in terms of a very engineering, like let's isolate individual systems and systems biology approaches very hard to get any sort of meaningful quantitative predictive power.

53:09 · So maybe going off on a tangent, but I mean, I'm curious, going to context and thinking about how context scales to an organism, you're talking about right now two million length context, right? Yeah.

53:23 · I think the human genome is roughly 1000 times longer.

53:29 · But even like a lot of say bacterial genomes, if you're trying to engineer them, are quite a bit longer than that. So how do you leverage something which has a long, but still finite context compared to do synthetic biology across like large organisms?

53:48 · And it may actually be for context, how did the Evo2 bacteriophage design work, which was probably much more than two million for that as well?

53:57 · Yeah, so I could expect the Evo bacteriophage just a little bit, because it's actually a separate group that worked on that, but that context was actually pretty short. Actually, the reason why they started with phages because it's amongst the shortest genomes. And so I believe it was something around 6,000 base pairs.

Designing Biology for Rare Earth Extraction

54:14 · Oh, wow. That's really short.

54:16 · Yeah, yeah. So extremely short. Fibrosis are insane.

54:18 · They're incredibly efficient. It impact a lot in there.

54:21 · Yeah. Yeah. But yeah, no, I think your question about how do you get longer context with something smaller, is a very key question that we, I think broadly the AI community is constantly trying to fix and be creative about.

54:35 · And so, I mean, I think that's a large part why we as a company are in AI research lab first, because we think the innovation needs to be constantly pushed.

54:45 · It's not a space where we can just grab open source models and expect that many of the tasks that we care about are just gonna be solved.

54:54 · We wanna continually push the envelope.

54:57 · And so context is one of the key researchers that we drive in.

55:01 · I would say that's probably how we got our name, because we worked on long context before it was a thing, I guess, like 2023, 2022, long context.

55:11 · You and your team, I mean, your collaborators have a long history of these state space models, doing, pushing context, like what seemed insane at the time.

55:24 · I mean, now I think routine of books, the big slabs, but at the time was like orders of magnitude longer.

55:30 · I don't know if you wanna talk about that a bit.

55:33 · And also I think one, maybe one thing which I think is really fascinating is how in for both this model and other ones, you really go on a first principle way, like diving into the architectures of how GPUs work and designing models which are, you know, exploiting the architecture of GPU in addition to, I mean, it's not just, oh, we're building longer. It's like, what can we do special, given the compute constraints we have to push the boundary?

56:00 · Yeah, absolutely. If you wanna distract our researchers at radical numeric, this is how you nerd snipe them. You talk about, you bring up long context and kernels, GPUs, then they're like, wait, did someone say kernels?

56:13 · And they start trying to figure out, you know, how to make things fast.

56:17 · Yeah, long contexts were a special place in my heart because that's what I focused on in my PhD at Stanford. And that's how we started thinking about DNA and, you know, going back a little bit for fun, we were looking at working on language models in general, and then we noticed our models were good at long context.

56:34 · And so that's the progression of like how we started working in the space was like, oh, these models seem to be really efficient on long context.

56:43 · And then with Michael Pali, my lab mate, he worked on the first design of AppHaina, this convolutional architecture. And then we started thinking like, let's push this further. Let's see what new applications open up if we really lean into long context. And we asked, what's the longest sequence out there?

57:01 · And eventually, unsurprisingly, we landed on DNA.

57:04 · We were like, DNA's gotta be the longest, three billion base pairs.

57:09 · And we started thinking like, okay, what's being done there?

57:14 · Like what kind of context lanes are people doing there?

57:17 · They're doing super short. They're doing like one or 2000 base pairs or tokens at a time. It's like way smaller than what one would want for DNA.

57:26 · I mean, most human transcripts are like 3K or so.

57:29 · So that's not even like us, that's not even what you need to represent like a protein or most of the time.

57:35 · And so it's clearly a need there and overlooked.

57:38 · And we saw it as a way to initially, like let's see if we can do something that we had no idea if it was gonna work and no idea who would want it.

57:45 · And so we just started tinkering around.

57:48 · And it turns out out of the box, relatively out of the box, it was doing pretty well at reading DNA. And then it led to, okay, people keep asking about like DNA.

57:56 · They didn't really care about our language stuff as much.

57:59 · And they would say like, can you do longer?

58:01 · What can you do with it? They always ask, what can you do with it?

58:05 · We didn't know really.

58:07 · And then for some reason we'll actually on this idea of like, well, no one's writing DNA, can we get it to write DNA? And that was a really simple question, but in hindsight, it almost seems obvious that, yeah, you would wanna write DNA and design it. But when I was first pitching the idea of Evo to folks, I spent six months, which I guess in bio words, not that long, but I spent six months going around saying like, hey, if I generate DNA, like, would you find that useful?

58:34 · Like, what would you do with it? Would you back us?

58:37 · Like, would you wanna be a part of this?

58:39 · And crazy enough, most, almost every scientist at Stanford I talked to thought it was a stupid idea. It was, I was like, this would be so cool, like generating DNA, how much would we accelerate the field?

58:52 · And then people would say like, what would you do with it?

58:55 · I'm like, I don't know. And then I would get that comment or comments like, that's not possible. Like we as humans don't understand the rules.

59:03 · How could you expect an AI to learn it?

59:05 · You can't even tell if it's right or wrong.

59:08 · Like you can't tell the AI, yes, that's right or wrong.

59:11 · How can you expect it to learn it?

59:13 · Or there's too many repeat characters or DNA is too noisy of a distribution.

59:17 · Like there's no real rules and there's just a bunch of junk in there.

59:21 · I heard all the reasons and I was just so stubborn about it.

59:24 · There's gotta be a use case from being able to generate DNA.

59:27 · I just, it just feels right. And I didn't know what it was.

59:31 · So it was, it really was an experiment of like what happens.

59:34 · And so when we first trained Evo, I remember we had no idea if it was gonna work.

59:38 · We had no idea, we didn't know what thing was gonna emerge.

59:42 · We just thought, let's just train a big one.

59:44 · Which is kind of ridiculous. But somehow folks that are, they're like, sure.

59:48 · Why not? Let's see what happens. Let's pay some money for the GPUs and let these crazy kids train the model. And I remember when the first result came back and it kind of like gave us chills over like, oh, maybe there's something going on.

1:00:02 · Which was somebody took the model checkpoint, the first one and threw it at a protein gym. One of the protein benchmarks.

1:00:08 · And turned out to be competitive with protein specific models.

1:00:12 · And we were like, okay, that is pretty surprising because we never told the model what's proteins versus not. And there's actually, it was, you know, there's also protein and DNA, right? So there was one piece, I think it was just surprising that it was actually competitive with protein specific models.

1:00:29 · And this was like the first experience.

1:00:31 · And then they just kind of kept on coming one after another.

1:00:34 · And it was like, oh, this competitive here, oh, it's state of the art on RNA and our DNA. And we started seeing like, oh, it can learn across different modalities, like not just DNA. And then we felt like, okay, this, there's something there.

1:00:47 · And so it kind of went from there.

1:00:49 · And then, you know, folks at Nvidia were like, let's back Evo too, let's make this even bigger. And then Greg Brockman from OpenAI was like, I'll take a break from OpenAI and take us four months of article and like help these crazy kids out.

Why Long Context Matters for DNA

1:01:03 · And, you know, then we're Slack messaging Greg Brockman at 3 a.m.

1:01:06 · trying to debug our code, which is wild.

1:01:09 · Yeah, so it just, it went, the trajectory was very surprising in many ways.

1:01:13 · But at the same time, you still got a lot of feedback like, you know, what are these models good for? What are they, you know, what are they useful for in the real world? And so that's really motivated us to start a company.

1:01:25 · We thought what we showed was just really just a taste from like an academic flavor.

1:01:30 · In a similar way, the language models, when they first too, you know, if natural language came out, people asked some of the questions like, what are these things good for? Oh, cool, it can write some jokes for me.

1:01:42 · Is this gonna lead to like, in all, you know, all-composing AGI that can automate everything? They did not think that, right?

1:01:49 · They thought it's like a toy, it's got emerging capabilities and they would extrapolate the potential. And in many ways, I think what we're seeing here is even more exciting or reminiscent of that trajectory for DNA.

1:02:00 · But Evo2 even was showing the bacteriophage and that wasn't a persuasive, I mean, that's almost a scary example, right?

1:02:09 · As you mentioned. So the people didn't see that, like that isn't a light bulb moment for people.

1:02:16 · Yeah, so I mean, and some people, right?

1:02:18 · I think either fair or, you know, rough critiques, you can even, you know, play devil's advocate and say what it's generating is kind of, you know, pretty close to nature and you're kind of recapitulating just kind of a small variance.

1:02:34 · And so there's, I think there's a lot of ways to critique and kind of, you know, minimize the potential, which can be fair argument.

1:02:43 · Like, so I think at this point, we think what is more important is not just pushing on the science and like, cool, this can be done and like kind of leave it there as a sort of thought experiment. But like, how do we actually use this to improve human understanding of disease, improve treatments, make better rare earth mineral extractors? How do we actually make this useful is what we care about now as a company, in addition to some of these more scientific questions.

1:03:12 · So maybe that's a good segue. There's, that brings up two things for me.

1:03:17 · One is the Meck and Turt stuff that's in the end of the blog post and then also the biosafety stuff. Those are both, I think, important applications.

1:03:26 · So can we, let's do Meck and Turt first.

1:03:29 · Can you talk a little bit about, and this is actually kind of building on the work that Evo did as well, I think. I remember in the Evo paper, there was some Meck and Turt work in which they were discussing using the, reversing the question and using the model to extract insights about biology.

1:03:48 · And so, and this has actually become a common theme with I think maybe biology more than any other domain is that people, because it's a scientific question and it's learning patterns about the world, you can actually say, okay, well, what patterns did you learn?

1:04:05 · Yeah, so I think Meck and Turt is an emerging field in bio that we're extremely excited about. And admittedly we are on the early side, I'd say.

1:04:14 · So we're building up that team, but so far what we've showcased and been excited about is looking, really analyzing the embeddings and some of the activations in the model. And so this idea of Meck and Turt for bio is borrowing a lot from the natural language community currently, you could think of what the models are doing is compressing a bunch of information it's seen, right? And so in this compression, it's really basically distilling it down into the key components, key patterns that helps it understand the data or learns the distribution of the data.

1:04:48 · And so what we're trying to do is probe the models to see what did the model distill into its weights and its activations or sort of like the outputs.

1:04:57 · And so for us, we started off with a lot of the outputs of the model, so the activations, and we wanted to see what kind of structure, what kind of visualizations can we see that help us understand some of the complexities of DNA, which is a ton, right? And so some of these complexities can range around GC content, certain motifs of transcription factors.

1:05:18 · There's a bunch of regulatory types of patterns and motifs that the model, we believe, has to pick up to be able to do its tasks, right?

1:05:26 · To understand whether a disease is caused by a variant.

1:05:29 · And generally, it's going to compress all these different motifs and distill them into the model weights. And so our job is to then find these and see if we can distill a pattern or structure that can be generalized to other cases where we don't understand the patterns, right? So that's sort of largely the goal.

1:05:47 · So in this case, GC means the two nucleotides are the fraction of those in a sequence.

1:05:54 · Exactly, yeah. So the fraction of the G and C letters in a genome, which is amongst the more simpler things, but also simple things as repeats, number of retifs.

1:06:03 · I think transcription factor, TF motifs is another one that is especially interesting for folks because eventually we're going to get to a case where we can design transcription factors, meaning we'll have certain transcription factor patterns via like ChIP-seq, you know, modalities.

1:06:19 · Transcriptive factors are, well, you can go ahead.

1:06:23 · Transcriptive factors are molecules that combine to DNA and they would alter essentially the gene expression pattern.

1:06:30 · And so it has this regulatory effect that doesn't modify the DNA itself, but can modify sort of the effects of DNA in the products that DNA makes.

1:06:38 · It can have a lot of implications on, well, pretty much everything in your body.

The Surprising Origin Story of Evo

1:06:43 · So it can modify disease dates, it can modify, I guess that's a lot of aging-related research is around transcription factor design.

1:06:51 · And so we think being able to understand some of the motifs via DNA, but also additional modalities will eventually let us be able to design transcription factor patterns as well. And so I think this is very exciting.

1:07:04 · In many ways, the combinatorial space of learning these transcription factors, like what binds and where they bind and what effects it causes is just far too vast to be able to do this in a manual way. And so we want to take a data-driven approach to learn some of these motifs.

1:07:21 · The complexity here is partly because the transcription factors themselves, coding genes, and so that you can, those can regulate each other.

1:07:31 · And so that you have this, that's where that combinatorial effect comes.

1:07:36 · So like just narrating for the listeners only, we're kind of marching through these increasingly complex and higher level factors all the way from GC content.

1:07:46 · We started now, we're looking at disease, which is maybe the most complex thing or you are pointing towards in this analysis.

1:07:55 · Yeah, so I think what we, our first idea for previewing this was, and what we're working toward this broadly, is this idea of mapping the manifold of disease, right?

1:08:06 · So manifold, there's like the sort of representation space of what disease looks like to a model in terms of the output scores or embeddings.

1:08:17 · We believe that there's a lot more structure that can be gleaned from understanding some of these outputs. And so, mapping this manifold or the landscape of what the structure of disease, and obviously there's many diseases.

1:08:31 · And so I think would be a big, exciting area of research for us to actually drive motivation for meconterbine bio. I think we're just really scratching the surface because if you can get this much of, glean this much of insight potentially from just DNA, which is in my mind, just one of the sensors that you want to ultimately

1:08:53 · fuse into modeling biology, then being able to do something similar across all modalities is something like, by all modalities, I mean protein, RNA, epigenomics, the attack, chromatin accessibility and methylation patterns, all these other different types of sort of molecular phenotypes around DNA just presents such a huge opportunity that this guy is kind of late in front of us that is all green space, green, green field.

1:09:22 · No, I haven't seen anybody do this level of sophisticated techniques from machine learning, deep learning into what I think is gonna be the most important impactful area of understanding and applications for AI.

1:09:36 · That's what we're excited about. And I think we're just showing a preview, very simple preview from the DNA only models, but in our next generation of models, which will be increasingly multimodal, we're talking dozens, it's a very exciting moment for us.

1:09:53 · Sidebar on that, is language, natural language, one of the most?

1:09:58 · Not yet. Oh yeah. Yeah, but it will be.

1:10:00 · So I'd say there's a lot of questions about how to fuse that with bio, but I actually think it's pretty, will be relatively straightforward.

1:10:11 · I think the more tricky part for us is actually how to fuse the biological signals more so. Fusing language, there's a lot of examples with that with like, with image and video space. So we feel pretty good about that and we've done some early experiments with language. And I think that will make it extra accessible for folks when you connect it to language. But I think the part that recipe that folks still are trying to figure out is how to do this across biological modalities.

1:10:39 · It just seems like a natural way to be able to do chain of thought, right?

1:10:43 · Exactly, yeah.

1:10:44 · Absolutely. And I think, especially when you start having chain of thoughts from orchestrating tools, things like cloud science, I think the idea of incorporating language, I think is already in a lot of people's minds.

1:10:57 · If you don't have any questions, let's talk about the biosecurity stuff.

1:11:04 · For us, we as a company thought it was very important to have a dual mandate, what we call a dual mandate. And it's this idea of essentially being cognizant and feeling responsible or wanting to feel responsible for the capabilities that we're enabling on the design side. So if we're going to create models that can design function into sequences, we believe and see a gap in companies being able to safeguard that technology and make sure that it's used responsibly increasingly more.

1:11:36 · I mean, folks have, I was just at a panel last night, a panel on AI scientists, agents that can do scientific discovery.

1:11:45 · And one of the last questions was, what are some of the biggest risks or doomsday scenarios with AI learning about science?

1:11:53 · And every one of them talked about biological weapons.

1:11:56 · And at the same time, I was curious because I was like, okay, so then what are any of these folks doing about that? And basically, I didn't hear anything about that.

Can AI Discover New Biology?

1:12:06 · And these companies, I won't say the names, and I look at the companies, they don't have big efforts in those spaces. So anyways, we felt it was important as a lab, in a lab that a team that was both building the design capabilities is actually also best suited for building the defense capabilities, because they're basically the same models. A model that is good at generating, turns out is also very good at discriminating or predicting if a sequence is pathogenic or not.

1:12:36 · So we felt it was not just from a principle standpoint, necessary to work on both biosecurity and the design, but that it was strategically, it just made sense as well. And so we felt this resonated with our team, but also the broader community, folks in the US government and abroad even, that there was a clear gap in need for a type of entity to exist to do this.

1:12:59 · And so yeah, we felt it was necessary to build it into our mission.

1:13:03 · And so for us, what we have done and plan to do, we look at it, or I should say, biodefense biosecurity sort of has three or four different pillars of strategy for biodefense. The first one is around detection and surveillance, broadly can you detect from the environment if a sequence has a pathogen in it, something that can cause disease, let's say from swabs at a airport, a nasal swab, or this sewage system, right, collecting samples.

1:13:30 · The next is attribution, which is once you've detected a danger, can you figure out where it came from? Is it natural?

1:13:38 · Is it from a rental country abroad?

1:13:40 · Is it engineered by a human from a specific lab in a country?

1:13:44 · And that helps you figure out what to do about it, right?

1:13:47 · So this next idea is around countermeasures.

1:13:50 · So once you've detected, figure out where it's from, what do you do about it?

1:13:55 · Can you make a counter agent? Can you make it antiviral or antimicrobial?

1:13:59 · And then fourth one's generally around deterrence, but that's more of like a government kind of level thing. But yeah, so we focus primarily on the first three.

1:14:09 · So we create tools that can, given a sequence, detect if it's pathogenic, but also what we felt was missing from the community was not just detect, you know, broadly if it's pathogenic, but characterize the heck out of it, meaning what parts of the sequence are dangerous, what genes, for example, attributing where it came from, being able to not just look at its, you know, sequence and match it to a database, but be able to attribute its signatures.

1:14:37 · And then also a big component to make this effective in the first place.

1:14:42 · The community, the biodefense community broadly focuses on sequence matching.

1:14:47 · So they'll take a sequence and they'll basically align it to a known database and see, say, have I seen this before?

1:14:54 · Does it match, you know, this list of known pathogens?

1:14:57 · But I think what's emerging and a concern for a lot of labs is, well, one, new stuff, right? If it's not on your list, and two, things that were intentionally obfuscated to not be detected in sequence space, meaning the letters matching up exactly, but also function space, right?

1:15:15 · Because basically models that we're enabling now, they will be able to be function aware or structure aware. And so that means for concretely, you can have a sequence that has the same function, like a pathogen, but actually look different in terms of the letters.

1:15:31 · And you can imagine basically what we have on the screen here still is this Meckinterp thing, and there's this sort of manifold that the model constructs internally that is kind of coding for these, among other things, functions.

1:15:45 · So you can imagine how it would be able to say, oh, well, that's, you know, maybe genetically quite different or at least somewhat different, but it still has a similar function.

1:15:56 · Exactly, exactly. So what we talk about in our defense blog is this idea of things that can function similarly, they can start having separation in terms of what the sequence looks like while maintaining the same functionality, right?

1:16:14 · So this can happen in nature sort of naturally, but also what these AI models allow you to do is also intentionally do that as well.

1:16:23 · So have the same function or functional capability, but have diverse letters, they say, diverse spelling, but describe the same thing basically.

1:16:32 · And so in this case, this work from Microsoft called paraphrases, so it's like, you know, kind of rewiring things, where they showcase that you can, for example, use protein language models to essentially keep the same structure, which structure implies similar function, but then change the spelling, right?

1:16:51 · And not just that, they wanted to test that if you have this capability and you send this through existing detection systems, would it break the system?

1:17:01 · Like would it actually detect it or not?

1:17:04 · And I think one of the interesting facts that maybe the general public, but most biologists know, is that there's these DNA synthesis companies, right?

1:17:13 · Where you can basically send a design of sequences and get back a DNA molecule like it's an Amazon package.

1:17:21 · Like you just send it off and they'll send you the physical DNA of the sign.

1:17:26 · They'll manufacture it for you. And this runs the scientific community, right?

1:17:30 · It's the pipeline that allows people to do research and understand biology and make drugs and everything. So it's prevalent and it's public.

1:17:38 · And so I think one of the concerns for folks is, well, and one of the concerns for us when we first started working on this was in a world, for example, where agents are prolific online, presumably billions and trillions of patients building and taking all sorts of actions on the internet, it's wild that they don't have any tools to basically tell it if it's making anything dangerous or not.

1:18:01 · And so that was literally our first motivation of like, we should probably make something that can detect if something is dangerous or not, right?

1:18:09 · So you guys filter in the, for these manufacturers that they can say, what am I making here? Is this dangerous or whatever?

1:18:16 · And maybe you could have an exception if you were like some licensed lab or something and I'm doing something dangerous, I know I'm doing it.

1:18:22 · Please let me do it anyway.

1:18:24 · All sorts of cases. So that's one scenario.

1:18:27 · To be fair, many of the DNA synthesis companies have detection tools, but I would strongly hypothesize that they're not AI based and fairly, they're probably not robust. Yeah, they're all pattern matching mostly.

1:18:41 · Sorry, maybe I should ask this question later.

1:18:45 · I'm just curious about, we've all, everyone who is working in bio has tried to use Fable and it's a legal, literally everything you type in.

1:18:55 · My website for example. Yeah, yeah.

1:18:58 · You can't do anything in Fable without.

1:19:01 · But in terms of people like scientists exploring synthetic biology and creating new sequences and designing new sequences, it seems like it'd be very hard to get, to have an ROC curve, which you can live on that doesn't like impede novel scientific research, for legitimate purposes.

1:19:22 · How do you avoid, even with an F1 of point nine, which you're not even close to right now, I think that still could easily, if there are billions of sequences, sequence generated a day or at least like a year.

1:19:38 · I mean, I think you could really have a lot of, it seems like a very hard balance to.

1:19:43 · Yeah, it's a tough question, right?

1:19:46 · So I think broadly, the way we look at it is, if we thought about like, how do we 100% stop the dangerous design, I think it's a harder question to ask.

1:19:56 · I think the question we asked is, on the capabilities design side, there's plenty of folks pushing the frontier of that.

1:20:05 · When we look at the defense side, do we see frontier technology being applied there?

1:20:11 · And the answer to that was no, right?

1:20:13 · So we saw this huge gap and we wanted to sort of aid, come to its defense, I said, come to its aid to give it a boost, right? So I think that perspective, it's an easy choice for us to say, let's push, let's bring that to a better, head to head match against the design side. Is it gonna solve everything?

1:20:34 · Well, I think that's what we're gonna aspire to be, but realistically, there's always gonna be cases where it can get around, right?

1:20:42 · And I think that's one, that's a big motivation for why we think, safety for language models and chatbots is one layer, but also you're right, there may be things that kind of just get out there, get past that anyways. And so what do you do about those cases?

The Biosecurity Problem

1:21:00 · It's already past the chatbots, right?

1:21:02 · It's already aided someone into making dangerous sequences.

1:21:05 · So it's out there.

1:21:07 · I think the cool thing about our tools is that what we're building is tools for the folks that care about things that's already out there, out in the environment, that's made it somewhere. And now there's this whole ecosystem that we wanna build into that does the surveillance, that does the attribution, that does the countermeasures. We wanna boost that community, right?

1:21:26 · And build stronger tools for that space.

1:21:28 · Does it have to be perfect to be useful there?

1:21:31 · I don't think so. I think we can be helpful and move the bio defense community forward in terms of bringing AI technology and AI frontier technology to their aid without being perfect in it. So that's kind of the way we're gonna-- Maybe another thing to say.

1:21:45 · So threat model, it's not even clear what threat model you're actually trying to defend against at this point, but the capabilities don't currently exist at all.

1:21:55 · So you provide something and now this can be worked on in terms of a larger regulatory framework or government or nonprofit, whatever, larger framework.

1:22:06 · It now provides you a tool that the community can build upon.

1:22:10 · Even if it's not perfect, it provides a starting point.

1:22:13 · And if you don't have a tool, then you can't do anything.

1:22:17 · Yeah, the bar is the improvement, right?

1:22:19 · So the bar is like, where are you at now?

1:22:21 · Can we move the needle? Can we move beyond sequence-based matching, alignment-matched matching? Absolutely, I think there's tons.

1:22:28 · And I think the interest is only growing.

1:22:30 · I think we're hitting a lot of chatter at the regulatory side, from different politicians and different folks in think tanks.

1:22:38 · It does seem like folks are mobilizing.

1:22:39 · So we're optimistic that this gets out and is more top of mind for folks to actually take action as opposed to just talking about it.

1:22:46 · Because right now, it does feel like there's a lot of talk.

1:22:49 · And one of the reasons why we felt like let's build tools and put it out and give access to people. Because there's been a lot of talk about AI companies saying, we should do bio defense. But then, okay, what does that mean?

1:23:00 · What are you gonna do about it? And so that was our approach.

1:23:03 · So I can just steelman this for a minute, though.

1:23:06 · Can you go to the other diagram where you have the, and if you click on them, then they show other similar, similar compounds or similar genomes.

1:23:16 · So, arguably, just steelmanning the opposing viewpoint.

1:23:20 · If you are on the frontier, which it seems like you are, and you certainly strive to be, you're moving the frontier of the attack and the defense at the same time.

1:23:31 · So like, right now, what you're doing clearly helps.

1:23:34 · But that maybe if you're on the frontier, then that doesn't really matter in the long term. So how do you think about that?

1:23:42 · I mean, I guess you just have to nerf what people are doing or something.

1:23:47 · Yeah, the mindset of what was hoping to get folks to start thinking about it as less of a, oh, let's make a bio defense tool and like call it a day.

1:23:57 · It's basically an arms race, right?

1:23:59 · Similar to the cybersecurity community, you're gonna make better technology, especially with something like Fable, that can potentially attack.

1:24:08 · And that means you're gonna have this back and forth.

1:24:12 · The design side's gonna get more capable.

1:24:14 · The defensive side needs to try to get ahead, right?

1:24:18 · And then that just motivates other folks to do past that too, right, on the design side. So I think inherently there is this arms race style dynamic that the way it appears to us is that the defensive side has been far, far lagging.

1:24:33 · And so what we wanna do is bring the defensive side closer to par essentially.

1:24:38 · So that's the mindset I picture or the framing I think about it.

1:24:42 · The cybersecurity analogy is interesting, but I think it differs in some key points.

1:24:47 · So first of all, I think that with cybersecurity with a sufficiently strong model, you might actually be able to close all loopholes which are not sociological.

1:24:57 · There are certain ones which will always be hard, always be ways of getting around things, but you could in principle catch every single exploit.

Why Today’s DNA Screening May Fail

1:25:07 · And I think that might be possible in the future.

1:25:10 · And then you can patch them. And as long as people say updated, you're secure.

1:25:15 · We have fixed genomes, right? So you can't patch a human.

1:25:19 · I mean, so once something, so I think that in some sense the attack services or the way that you defending a set is much higher or much harder, but then maybe the converse is that it seems much less likely that someone would have the incentive to go on offense to the same degree. And also the barrier to entry to success, even if you can print out arbitrary DNA and the process of going from that to making a successful virus, especially one which doesn't kill the person designing it is actually quite large.

1:25:50 · So I guess maybe I'm curious, what is the single, from your opinion, what is the single biggest threat that we actually have?

1:25:59 · What would the thing that would keep you asleep or keeps you awake at night?

1:26:04 · Is there something in particular or is this like, you just think this is something we need to build and let's build it?

1:26:13 · One, it's hard to get in the mindset of a person wanting to design a bioweapon.

1:26:18 · So we're not trying to necessarily think of all the potential people or scenarios that bad actors might work on. What I think broadly, what I worry about is we're lowering the bar for how much expertise is needed and the speed at which folks can engineer these kinds of things. And that means the volume is gonna just exponentially increase at some point.

1:26:43 · And so there's a mix of, yes, there's intentional worries because there's certainly state actors that have had biological programs, very, very, very large ones. And one of our advisors on a company has physically seen these facilities and decommissioned them. And so we've heard a lot of stories about states actually being motivated to create such weapons.

1:27:08 · That is one concern. And I think during peacetime, it's less scary.

1:27:12 · During wartime, it's particularly scary.

1:27:15 · I think the unintentional ones are also things that, in my mind, potentially more likely in the near term, where folks do try to generate things and control function and inadvertently things that maybe they tried to make a certain thing to understand and to study, but it got out, right?

1:27:32 · Because these things can be hard to contain, for example, and they just have a leakage. So I think those scenarios potentially seem the most likely.

1:27:42 · Does it keep me up at night? Not necessarily.

1:27:45 · I'm much more of an overall an optimist.

1:27:47 · I think building this type of technology is ultimately a game that you weigh out the pros and cons and the costs and benefits.

1:27:55 · I think the benefits far outweigh the potential harm.

1:27:58 · And so that's why we work on it. Because ultimately, we do think it's going to be an engine for discovery and human health improvement.

1:28:07 · But at the same time, we just felt that the defensive side was sort of losing this arms race. And so we work on it as well and try to push the frontier.

1:28:18 · But I'd say overall, I think as a community, as researchers, but also folks in the policy side, I think the community is resilient enough and has the ability to mobilize to get ahead of it. And so I'm very optimistic about it.

1:28:34 · We have two questions that we like to ask every guest.

1:28:39 · So the first one is, if you could, by fiat, remove a bottleneck that is important to you, what would that be?

1:28:50 · Interesting.

1:28:52 · Could I get two answers on this one?

1:28:54 · Sure.

1:28:55 · The first one is kind of a cop-up, because every AI lab says this, like-- GPUs.

1:29:00 · GPUs.

1:29:01 · (Laughter) We've gotten the answer a few times.

1:29:03 · We can use GPUs.

1:29:06 · The other one, which I think is a little more philosophical, I think in this space, in what we're trying to do, aspire to do, is reinvent how scientists do their work in this space. And I think one hurdle we run into is folks who-- and you see this in many domains-- folks who are-- the more expertise you have in something, the more pessimistic you become about that space.

1:29:29 · And I think you especially see this in bio, where you know a disease area or modality so well, and then someone introduces something else new, and you're like, oh, but what about this and this and that?

1:29:41 · And they're very pessimistic, and probably rightfully so.

1:29:45 · I think what I've noticed at the company, what we strive to do, and the folks that we try to bring in, are domain experts that do know that field, but also are still dreamers, meaning they still do have that imagination and desire to change how things are done. I think that barrier, we see that a lot in the field.

1:30:05 · I think if we embrace that more, we can see a lot more progress and step change that I would love to see.

1:30:11 · Brings us to the last question, which is, yeah, is there something that you want the audience to take away, a single message?

1:30:18 · Yeah. I think one message is, I think folks have had, especially into the AI research community, felt like there was a choice they had to make sometimes to either work on the frontier of AI technology, and that was like consumer-related apps or enterprise-related apps. And they just had to work on chatbots.

1:30:41 · And that's the cutting-edge technology.

1:30:45 · But also, I think when people think about the true potential of what AI can do, and I think a lot of it is about improving human health, understanding our biology.

1:30:54 · But people have felt like they have had to choose.

1:30:57 · And if I work in that space, I can't work in the frontier of AI.

1:31:00 · And what I would like people to take away is that you don't have to choose.

1:31:04 · And I think we can work on things you truly care about that you think will push humanity and work on cutting-edge technology.

1:31:10 · And that's what we're trying to build at Radical Miracles.

1:31:13 · Yeah, and I mean, clearly, and I encourage people to read the blog posts, the architecture, the Meck and Turf. There's a lot of innovation that is going into building these models. And I firmly believe that biology is really on the forefront of AI.

1:31:29 · Awesome. Yeah. Glad you feel that way.

1:31:31 · Amazing. Thanks for joining.

1:31:33 · Yeah, thank you for-- That was fun.

1:31:35 · Making the long trip. That's a lot of way.

1:31:38 · Anytime. Thank you. Precious invite.

1:31:41 · I had a blast. Great. Thank you. Thank you.

1:31:44 · (Music Playing)