Transcript

Hook

0:00 · Can we create a simulation of 8 billion people living on earth? I think that's quite interesting and that really is the vision and once you get to that kind of state the kind of questions that you can help answer for the society also start to change from my perspective and for me it's questions like can we help solve climate change.

0:18 · If you look at climate change as a problem space, this is what we like social scientists would often call the wicked problems problem where you have many actors with competing incentives for trying to make a very complex decision a coordinating coordination decision. Can simulation help us solve that?

0:37 · Before we get into today's episode, I just have a small message for listeners.

0:41 · Thank you. We would not be able to bring you the AI engineering, science, and entertainment content that you so clearly want if you didn't choose to also click in and tune into our content.

0:50 · We've been approached by sponsors on an almost daily basis. But fortunately, enough of you actually subscribe to us to keep all this sustainable without ads and we want to keep it that way. But I just have one favor to ask all of you.

1:03 · The single most powerful, completely free thing you can do is to click that subscribe button. It's the only thing I'll ever ask of you. And it means absolutely everything to me and my team that works so hard to bring the Inspace to you each and every week. If you do it, I promise you we'll never stop working to make the show even better.

1:21 · Now, let's get into it.

Introduction and Joon’s Path from Art to AI

1:25 · Today, we have June in the podcast.

1:27 · Excited to kick this one off. Very exciting company. I want to kick off and ask you the question, you know, talk us through the story of your life. How have you gotten here? Right. Yeah, for sure.

1:38 · Uh, so really excited to be here. A story of of my life. So I was born in Korea. Um, and I lived there for good 11 years or so of my life. And then my family moved to Boston. Uh, so we moved uh when I was 11. And my parents were doctors. So they were basically going through their post-docctoral studies. My dad was a surgeon. So he was doing uh his sabbatical years actually at the Boston Children's Hospital. So I grew up there. um not too close to tech actually.

2:08 · I was very much like uh you know music artsy painting like that kind of guy. I I actually got into painting a little bit later uh in high school. Uh but that's what I used to do. And then I grew up mostly in the east coast after Korea. So I lived good number of years in New Hampshire and then I went to college in Pennsylvania. And I got into more of this tech scene uh in college.

2:33 · Uh so I was originally trained to be an artist. I actually thought that would be my actually professional career. So I it wasn't a hobby was actually like hey let's make a living out of this and then gradually I got really interested in this idea of hey the greatest artist often creates their own medium and the best medium that we had available today was actually in computation.

2:52 · So I decided to go deeper into that and one things uh led to another and obviously we can go deeper into this but I decided that research was something that gradually uh that I got got interested in and here I am. So there's obviously a lot that you packed into the research component. You had one of the best papers of 2023 which was the generative agents paper commonly known as the smallville paper.

Smallville, Generative Agents, and the Origins of Simulation

3:20 · Yeah.

3:20 · Feel free to call back to anything else that you mentioned, but most people would have heard of you from this obviously. Do you have any statistics of how many people uh have like read it?

3:31 · Archive gives you something, right? Some some stats.

3:33 · Yeah, it's a good question. How many people have read it? I'm actually not sure. I know this that I mean we do keep track of the number of citations uh which I know is going up uh quite fast, but the readership the Google Google Scholar has 72,000. it it made a bigger hit and it was actually a pretty instrumental paper. It was like one that got cited so many times.

3:57 · It is frequently like when people ask what is the best paper of the year like best paper you've read recently it's that's this one.

4:02 · I thought the the memory component was pretty underrated you know like very good early memory system but yeah one of the biggest papers you know.

4:10 · Yeah.

4:10 · Yeah. Yeah. So maybe I can uh talk a little bit about how this particular paper came together. Uh so when I got into research it was back in 2020 when I started my PhD program at Stanford and that was the year uh when we were about to get GP3.5 uh GP3 to be available. So we already had GPD2 and you could sense that there's this new class of models that was just becoming available in the market and the team got very intrigued

4:38 · and the general consensus was well is this motor actually going to be useful for anything? It's really strange that these models are not trained to do any particular task but we decided to take a bat. So a large group of scholars at Stanford uh and it was actually led by one of my co-founders Percy Leang came together who coined foundation models who coined the term foundation model we wrote this paper uh where that term came from called opportunities and risks of foundation model and during that process

5:07 · really the thing that I started to think deeply about was here is a model that is fundamentally new in our ecosystem the reason why this was new was it wasn't again trained to do anything in particular But it was its premise was it could do anything and everything. It was like a stem cell if you were to take a biology analogy. And I got really interested in this idea that well if we were to really think about what are the killer applications that this particular technology would enable what would that be?

5:36 · Many of my colleagues were using this for simpler classification simple generations. Interesting that these models can do that but from an interaction perspective not that interesting. We've known how to do that for many decades. And what we came down to was these models are actually trained on this very broad data from the web, right? So these are human behavioral data. It's the is social media, Wikipedia, all these kind of data. So if you poke at the right angle, then you could see human behavior that would just pop out. That's actually quite realistic and we've never seen that before.

6:08 · So that got us really interested. the exercise that we decided to do and this is something that we uh this particular group of uh colleagues that I have uh myself Michael Bernstein PC Young uh who ended up becoming my co-founder at Simile we sat down and we played this game that we call the time machine game imagine we were to get on a time machine and fast forward 10 years and look back

“Let’s Just Create a World” and the Future of Personal Agents

6:33 · what would have been the single application that would have mattered that would be the most interesting and inspiring and Well, we thought, well, what if we can just recreate the world that we live in? I mean, it's really hard to get more ambitious than that. Like, let's just create a world. And that's where we started. And initially, we had this paper that was a precursor to the generative agent's paper called social similac. Correct.

6:55 · Before you go further, was there like a were there other candidates for the most ambitious thing in the time time machine exercise?

7:01 · Exercise. Yeah. Yeah. What could have been?

7:04 · What were the next you know, what was number two or number three?

7:06 · Okay. if you're a member.

7:07 · Uh so there is a close second that we were considering which basically ended up becoming more of these um automation tools but especially the the vision around really personalized agents that actually do things for you and that's also happening. it's also happening but it was sort of interesting for us right in that the reason why uh

7:29 · we decided to go with the idea of simulation one I mean I I was a huge science you know science fiction nerd um and this idea of creating simulation I was personally really just fascinated I I love the idea it's it's really cool to see like a game town like this and just see these agents live in it but at the same time my bet was if you were to create a really amazing personal assistant out of this technology. What you actually need first is an amazing motor of your users.

7:58 · So for instance, I told the motor, hey, can you go buy late dinner for me? And it orders how I am pizza and I do not like pineapples on my pizza. Then it totally failed. The way for it to not make that mistake is only by having a deep understanding of who I am. And I gave a very simple and dumb example here but you can imagine how this core understanding of people is instrumental. This is how for instance if we have our family closest friend they have a good mental model of who we are. That's the basis of our social connection.

8:30 · So our bet also was this technology around simulation creating accurate representation of people ought to preede the more complex agents that would automate the world that we live in. So that was the bet. So for so but that was a very close second and I'm still very much fascinated by it. I think there's a lot of interesting work that's going around.

8:53 · My hot take actually here though is I don't think we've actually seen a true personal assistant that's actually useful uh in ways that actually meets the ambition of that particular line of work. I think there are early applications that are obviously interesting and if you talk to even chip nowadays or claw they obviously know a lot about us. So a lot of the generation it's doing I do think it's much more tailored but I think the ambition is quite large in that field and I don't think we quite have the all the right ingredients just yet.

9:24 · So like open claw all these clients personal agents like what what do you want to see from them that you're that they don't currently have?

9:32 · I do think it's slowly getting there but I do generally want them to have much deeper understanding of the person. uh right now you look at the models I mean open claw what's it's basically leveraging is basically markdown file and I think it's quite clever right so if you look at the generative agent paper this actually was the same intuition that we had where initially when we were creating the memory architecture for the generative agents and this is like back in 2022 so we

9:58 · didn't really quite have the idea of even agentive architecture or the term agent but the intuition that we shared with some of the work that's coming out today was we initially thought well do we want to make the memory into let's say knowledge graph do we want to train a bespoke model all these kind of things

10:19 · and what we decided to do was no no just forget about all this these language models are actually quite good at modeling text and understanding and reasoning about text so just put everything in a markdown file or text file you're done I thought that was quite interesting that we could do that and there's a lot of strength in doing that But also there is limitation. It's the way you retrieve and make sense of data that's extremely large. It takes a lot of work. So I think that technology is getting better.

10:47 · I also do however think uh there are certain things you just cannot shape just by prompting the model. So some to some degree you do need to touch the parameters of the model itself. So there's these kind of work that I do think does need to happen and obviously it is happening. The question is how far can we take it? How do we source data and how do you also create an ecosystem where the people are continuously feeding data to this model?

11:11 · So it's learning about game.

11:13 · What's the intuition between why you need to do it in the model?

Social Physics and Behavioral Foundation Models

11:16 · My intuition behind the actual when do you train or even post- train a model versus just a model is if the motor has to learn the underlying physics of the world that it's operating in. So it has to learn new social physics. The places where it doesn't have to train is it already has the physics. We trust the physics. It already has the base statistics but it's just trying to react to an environment. Then I think you can just prompt your way into getting the you know actions out of it.

11:46 · I don't think the model has yet at least the models that are out in the open has yet learned the complete mapping of social physics of humanity. And this actually is one of the core thesis of simile, right? And one of the core reason why that is the case is if you look at the data that the model was trained on, these models were trained on the web data like whatever was available in the web.

12:11 · And these are really interesting data sets, but they are fundamentally the selfexposed attitudinal data with some behavioral data that sprinkle around here and there. And it has yet to learn really deep behavioral nature of people. Not just what people say they do online, but they what they actually do in real life.

12:33 · And this is actually one of the sort of what I would consider to be the dark knowledge of humanity that we haven't quite captured. And it's these kind of data that would also need to get factored into the model creation.

12:44 · You call it behavior foundation model.

12:46 · Uh there's a good oneliner here, but outside of that, what type of data do you need? What are you changing on the model level? How do you go about actually modeling you know doing a behavior foundation model?

12:58 · We think about data in three buckets. So one bucket is actually we interview data for instance it's quite interesting like qualitative rich qualitative data is interesting. It's not behavioral but we would literally ask people hey tell me the story of your life.

13:15 · Yeah that's what we're doing here.

13:16 · Exactly.

13:16 · The question that you all asked at the beginning of this interview literally is the question we also ask and obviously you know we ask our participants to go a little bit deeper uh than how far I went. Maybe I can actually give more of my life story in le of this but the reason why that data is interesting is by learning about this very long tale information about people you actually get a lot of texture around this model like this this person as a model.

13:45 · So even understanding their childhood memory or even their trauma, their first love, these kind of things quite informative in ways that's really hard to predict. So that's one. Then there's sort of two tranches of what I would consider to be the behavioral data. So one behavioral data actually is observational. So these might actually be like transaction data or these might be data that you can get by scripping the web, right? So you can imagine why these data would be these data sets would be interesting, right? is they give you the base statistics of people's behavior.

14:18 · But then there's the last category of data and that I personally think is perhaps the most important which is the the data that basically describes the cause and mechanism the wise of people and some of this is covered by the interview data the qualitative because people talk about why they made certain decisions but really where you get to see the most behavioral aspect of this actually is in randomiz randomized control trials like RCTs. Imagine you basically have the same setup but you have a few different variables that you're trying to tweak.

14:49 · Can you actually get realistic human behavior out of it in ways where oh imagine you had to make uh imagine you had this particular uh option? Imagine you're even trying to choose whether you're going to drink coffee or not. The day you drank coffee versus the day you didn't drink coffee. Does your behavior change? That's a data set that describes a cause or mechanism.

15:11 · This actually is quite important in actually modeling people. The reason why this is important is often times when people come to us uh or not just to us but the reason why people are interested in simulation actually isn't because they want to predict the future. If you're win against if you're trying to win against the stock market predicting the future is interesting but most people most decision makers what they want to know is how can we shape the future. It doesn't really help you to hear that your sales is going to tank in two quarters.

Prediction vs. Simulation: How Do You Shape the Future?

15:41 · They're just going to say, "Wow, that sucks." What they want to know is, "Well, what do we need to do now to avoid that future?" That's caus mechanism. And this is also very hard data to come by, right? Because the world is our ground truth, but it happens once. So in a very controlled setup where everything is equal except for one variable, this kind of data set almost rarely happens. So this is the reason why this data set data set is both hard to come by but also quite important if you're trying to model human behavior.

16:12 · So the the behavior I think is the hardest um data set to acquire what is out there what is possible even because like you're not going to know a lot of details about my life. I don't even have data for myself on like I want to analyze my own um health or habits and I just don't log everything. So how can you have that data?

16:36 · So we actually run a lot of randomized control trials.

16:39 · Yeah. But you put people in the lab, they watch them sleep or what?

16:43 · So we do actually care a lot about the consent process. So people know that we are like we invite them to be a member of this community to both share data and also have theirelves represented in different forms. Um but we bring a lot of people to the lab uh or virtual lab where we design experiments that would actually pose them real behavioral decisions. And often in these kind of experimental setup what makes the difference between what is attitudinal versus behavioral is if the stake in your decision is real.

17:17 · That's ultimately what makes it behavioral. So in these kind of setups uh we are inspired by our colleagues in social sciences, psychology and so forth. So when they run studies what the kind of techniques they uh utilize is imagine there's a online store that you're inviting people to come by and then whatever they purchase in this experiment they actually get that item delivered for instance like these are the kind of things that makes the stakes real.

17:41 · So we run a lot of these experiments and we also do partner with firms um and also right now we also have customers who are quite excited to at least give us a glimpse of the kind of behaviors that their users exhibit so that we can get a little bit deeper understanding of how people behave in these different platforms. I think on the customer side they have a lot of data about their users who has bought they have the action data. Can you kind of walk us through an example of what does someone come to you for?

18:12 · What questions would they want solved in the process of do you customize a model for them? Do you have something off the shelf? What does that look like?

How Simile Models Real People and Populations

18:22 · Today, uh when people leverage our models, it's often to better understand the population of their interest. So, usually the start of the relationship, uh we basically come together and hear about what population they want us to model, right? Um, so it might be that if you're a CPG company that's selling to all of the US, it might be fairly straightforward. You want to model the gem popup of the US. But at the same time, if there's a vertical or if there's a market that they're trying to go into, imagine, uh, they want to better understand, let's say, people in their 20s and 30s living in California, that's a much more specific population.

18:56 · So we hear about these population and we go recruit these people uh with consent uh and with incentives and we basically collect some of their data and create a model of these people and then what our product allows you to do is basically query them. Uh so it can take as input a filter that is a description of the population that you want to talk to just like the one I just mentioned and an environment. Environment can literally be a survey questions. It can be behavioral experiments. It can be AB testing.

19:27 · Often times the core use cases are things like concept testing uh to start with. But also, you know, people sometimes want to do focus group or one of the sort of fun use cases that we also serve is actually even modeling things like the earnings call for public companies. So these are the use cases that we often start with.

19:45 · Concept testing. Is that an established term? I've never heard of concept testing.

19:49 · Yeah. Uh so it basically has to do with they have let's say different messaging, different products, different ideas.

19:55 · It's like a marketing exercise. Yeah.

19:56 · Okay. Got it. Got it. Politics.

19:59 · We do uh have a strategic partnership with Gallup and of course Gallup is deep into policy space and so forth. Right now we have not worked deeply with politics like that area just yet. However, I'm curious if there is demand or if they really would have different needs that somehow fundamentally don't mix with your existing uh users or people.

20:23 · I think there's certainly demand. Yeah.

20:24 · But we are very much mindful of how this technology gets adopted and the societal impact that we'll end up having with this technology. And I do see politics as an area where a company has to be particularly thoughtful about the way they operate and make impact. So this is where we also want to make sure that we form enough of guard rail and perspective on how to leverage this technology before we go on to serve markets like the politics. I'll give people an example. Um one of my favorite shows is the West Wing.

20:54 · I don't know if if people have watched uh one of the key story lines is like the president has uh multiple sclerosis but they haven't they need to figure out how to disclose it.

21:04 · So they run a poll with a fake governor and ask people to respond on the poll and they try to make decisions based on the results of that poll on like how well they'll be received like where how should we play this and I'm like well you know I think those kind of counterfactual things I would actually use a simulation for this if I could trust it for sure. Yeah.

21:25 · In that show, how did it go?

21:27 · In that show, it basically was like kind of like a foregone conclusion. They were like, "We know it's bad. We just don't know how bad." And then the poll came back. It was like, "It's really bad." And then they just did it anyway.

21:36 · Part of it is it's a show, right? So, you're you're maximizing drama.

21:41 · How bad could it be? Oh, it's horrible.

21:43 · Yeah.

21:43 · And and to some extent I think that is part of the the trick of the or the challenge or with being a customer of yours which is that if I know it's if I roughly know and can in it what the effect is going to be do I need you what sensitivity of of effect do I need in order to make a decision right so for example if I um I my my approval approval rating is 50% and I they have this negative piece news item comes out and it drops to 30.

22:15 · Yeah.

22:15 · If it drops to 20, if it drops to 40, do I care? No. It I know it drops. It's negative. So, when do I care about simulations?

22:23 · You do something that's clearly bad, that's not popular, and people don't like you like Yeah. I mean, it's a simulation.

22:30 · Well, so there are a couple of things.

22:32 · Uh one actually obviously is um there are use cases where like every day for instance developers, designers, uh policy makers, uh marketers every single day they create assets, they create new products and turns out is actually um many of the decisions in hindsight is sort of obvious. Yes, of course this is bad. But we still run those studies because understanding the magnitude and understanding how acute something is is actually quite difficult.

23:03 · Even if uh we feel like of course like this makes sense. I mean this is the reason why we make so many mistakes. Like every time somebody goes online and say something that has huge backlash, you look at that and like what an idiot. However, it's tough. That's one. There's also another aspect here which is again this is the reason why simulation is actually different from prediction in simulation in the ideal case scenario.

23:26 · So what simulation is trying to show is it's trying to show each step of the way or each step that we need to take to get to a certain outcome. Right? So in the most advanced simulations sometimes the next step that we're suggesting might actually be quite counterintuitive. the analogy that I sometimes give and I ground it in a more realistic example but you know as I mentioned I'm a huge fan of science fiction and I don't know how many of the

23:56 · audience members have read like things like the foundation series as small we've mentioned psycho history number number of times okay fantastic so I I might actually be talking to the to the right crew if you read foundation series literally the first act is there's a group of scientists who have found out that oh our galactic empire is going to collapse and we're going to have 30,000 years of unrest. And they basically run psycho history, the simulator that tries to teach them, okay, how can we keep this unrest to,000 years?

24:25 · And they plan this out and the first step of that plan is to get the scientists who say, "Okay, this is coming exiled into this random place in this, you know, galaxy."

24:41 · Terminus.

24:42 · Exactly.

24:42 · And that's so counterintuitive.

24:45 · Like what a strange move that you literally sent the group of scientists who was raising voice around the potential collapse of Galactic Empire into nowhere. How is that the right first move? Well, it turns out in this particular simulation that actually was the move. It's these kind of things, right? And the reason why these kind of reasoning is possible is because you're showing the step function or each step that results in a particular outcome.

25:10 · So really what simulation allows you to do in its highest form is you give it not a problem or question like what would people answer to the survey. That's not what we do. What we tell it is here is a goal that we have in the context of foundation. We want to keep the unrest to a thousand years. what is the path that we need to take now to get to that particular future and that's what simulation allows you to do. Now translating that into real market.

25:39 · Imagine you're a automobile company and you're about to release a uh EV and you're trying to understand well how do we market EV uh to make sure that our stock price goes up. But what if the answer comes down that well you can market your EV in XYZ way but that might change people's perception around the cars that's not EV and actually make your overall sales to go down. not very intuitive.

26:05 · Especially all you're trying to optimize is EV sale and that's the only thing you're tracking then that might actually result in a completely wrong solution or at least different solution than what you would have expected whether it's right or wrong.

26:19 · Yeah, that's the power of simulation.

26:21 · For listeners, we covered a similar topic with ML Parin from Shopify where they are working on Sim Jim. I don't know if he ever talked to you about it. Uh it's very similar.

26:30 · The goal is increase conversion but then the the the journey is very unusual.

26:34 · Journey is unusual.

26:35 · Yeah.

26:35 · the he's actually trying to look for interventions on a shopping trajectory which is similar to what you're saying like it's not about the attitudinal is your your your word for it.

26:46 · It's about behavior.

26:48 · It's about and that's exactly the difference, right? It's like not about the near-term direction about but it's more about like how do you affect multiple turns of interactions, right?

Evaluating Simulations, Digital Twins, and 85% Accuracy

26:57 · You had a good quote at the start about this as well. It's not about people wanting to know the outcome. It's about how they can change it. Change the way to get there. something out there. But I want to take it back to how do we know this is grounded? Like how do you run evals? How do you test that simulations come through? Basically, if I was to do the same thing that you described with say your favorite LLM, Opus, GPT56, have some agent to map out these things.

27:24 · How different are the answers we would get if I give it the same goal, the same objective, make a decent system. You're saying that you need to change the model weights. you have your own solution to this, but how far off are we and how do you check if it's grounded? Uh you have some interesting stuff on your site that actually points to how you run really valves, but if you could take us through that side, you know, I think that's one of the big concerns that people have.

27:47 · They're like LLM hallucinate, you're just hallucinating layer after layer, right?

27:52 · The way we do this and this is actually the the paper that we worked on after the generative agents paper that really became the at least for simile and also the field of simulation and synthetic panels really became the foundation.

28:05 · Yeah, this is the paper uh the paper is called generation simulations of thousand people. Here's what we've done for this paper. We actually brought thousand people that's representatively sample from the US to a virtual lab and what we basically have done was we spent two hours collecting fairly wide ranging data.

28:24 · In this particular study, we focus a lot on this interview data uh that was uh whose script was taken from this project called American Voices Project and then we would also pair it with a lot of behavioral data and so forth whatever we can collect within two hours and then we would actually send these people away for a couple of weeks and during that time I would use this data to create their digital twins and I would bring the humans participants back after two weeks and have them complete a battery of surveys experiment experiments, behavior studies.

28:54 · So we actually have the list here which basically included things like the behavior economic games. We would run literally like big fight personality test, general social survey. We would also go ahead and run the randomized control trials that were published on PNAS and we would have their digital twins predict how the source individuals would have acted in these studies and surveys. And this is where we basically could replicate people's behaviors and attitudes 85% as accurately as people would replicate their own.

29:26 · So that actually was the first really paper that gave this validated results that we can actually model individuals in an accurate way. And what we ended up finding now of course in AI space uh so this paper came out at the end of 2024 AI space a year and a half two years that's a lifetime.

29:48 · Yeah.

29:48 · I just uh for listeners who are not seeing the YouTube I just want to say like the the headline figure is 85% accuracy like which is a big improvement over all the other meas methods that you that you showed but the part that was actually particularly striking to us uh especially as we improved this technology even further was the generative agents model generative AI models like CHP claw that's coming out it does give you the right foundation however what they do not consider is the

30:20 · true attitudal and behavioral aspect of people especially in the population that you care about. So what these models are really really good at today is they're trying to basically become the super rational objective machines right so you go get their data from places like maror scale you talk to professional programmers scientist to create model

30:42 · that's amazing at reasoning that's what they do similar actually doesn't care about any of this the models that we're talking about here what we're trying to create are models that are as dumb as I am right so if I makes those mistakes the motor has to make the same kind of mistake.

30:57 · Oh, that's very hard.

30:58 · That's very hard.

30:58 · You're solving more of VX paradox.

31:00 · That's exactly. And this is actually completely different kind of data and training objective. This is also where we actually see quite a bit of discrepancy in the performance in human behavior prediction between the frontier models and simulating getting created in the space where in some cases the model performance of frontier models go all the way down to 20 30%. Especially if you go into that more niche population on topics that our customers would actually care about on more gem pop it might be around 50 to 60%.

31:30 · So it's not very robust like you wouldn't want to make your decision off off of these kind of these kind of findings. If you can bring that up to 85% that is ultimately what people end up getting very excited about.

31:42 · Yeah. Do we want to keep going on the paper uh routes?

Post-Training Models to Reproduce Human Behavior

31:46 · Yeah

31:46 · for sure. Uh so the last one uh was sort of an interesting one. So this uh paper was the follow-up paper that we had uh to the thousand agents paper where basically the idea was now can we augment the models even further and actually post train a model based on a lot of randomized control trials. So this was an interesting one. The data is always the most interesting part of modeling in many ways. The data that we got here was there's this uh there's this platform called open science foundation. So some uh the audience might be familiar with this.

32:19 · Uh there has been especially in the social sciences over the past 5 years or so.

32:25 · There has been this concern around replicability of studies, right? So it's a bit of a crisis uh the scientists acknowledged where we rerun the study and we don't actually see the same finding. It's it's tough. And the the reason why that was often the case was there's basically the survival bias where the papers that get published often need to maintain what we call the p value of less than 0.05 in the experiments that we ran.

32:50 · That basically suggests that only there's only 5% chance that the results that we saw is false positive. But the tricky part was all the papers that were not published and there's still a 5% chance that whatever we publish is actually totally just randomly generated like there's 5% chance that hey this effect is not real but it just happened to be real because of the sampling bias. So because of that what scientists started to do was they started to pre-register their studies.

33:20 · So before running an experiment they would go to this platform and say here is the data here is the population that we're collecting and here's the hypothesis and they would just say here is our hypothesis like this is what we believe and you cannot retroactively change those hypothesis. This is what actually gives us more scientific statistical confidence that whatever effect that you ended up seeing is actually true.

33:43 · So that ended up creating this really interesting platform where there's one platform that has now contains tens of thousands of real world experiments and hypothesis and a lot of these are actually really high quality like professionally designed behavior studies and rand randomized control trials.

34:00 · So we actually got the data and the studies from this platform and basically used that to make a point and obviously this particular motor is not uh something that we're serving commercially because this obviously was a part of the open science but this particular data set may helped us make a point that by collecting a lot of these randomized control trials that are really well designed we can make significant improvement in models capability to predict human behaviors. So that's what this paper was about.

34:32 · Is this stuff done on a individual level? Like do I need to tune the model per individual per company? Is there foundation model changes and then some slight postraining? Anything you can share there?

34:44 · So uh this particular model actually was trained uh the data we actually had at the level of individuals but this particular model actually was trained. We experimented with both and this is actually what we end up doing at Sim 2. We always train two uh distinct model.

34:58 · One is what we call the population level model. The other is what we call the individual level model. And both actually take very similar input which is the description of a sub population or individual and a stimuli. In this particular work uh we've done the same here. The results that we are reporting are much more geared towards individuals because we do actually think that is harder task in many ways but that's what we have done. You seen anything on the questions that humans can solve that models can't solve?

35:29 · So like the currently it's you know I live 5 minutes walk away from a car wash. It's a 10-minute drive. Should I walk or drive?

35:38 · The model will say oh walk to the car wash and you know you don't have your car.

35:43 · Uh is anything like this a problem in simulation? you would assume like very simple for human to think about but if the model is saying you should walk to the car wash you know any anything here it's less uh what can we solve but I think it's more about what biases or mistakes do people make that models miss like for instance imagine that you are

36:08 · you know like the when I was instead of Stanford I lived in Palo Alto so it's about I would say 40-minute walk from the campus You ask the model, "Okay, let's go home. What can I what can I do?" It would likely call an Uber or, you know, give me, you know, the bus time. But for for the longest time, I actually really liked walking back. And the reason why I wanted to do that was not for efficiency. It actually really helped me think. And I like to walk for, you know, half an hour, 40 minutes or so a day.

36:36 · Uh, where I just get to, you know, you know, just think about ideas, research, just get lost in my thoughts.

36:44 · That's very human activity. Unless the model has seen that and actually understands the importance of that activity, it would actually miss these kind of features. So that actually I think is fundamentally what we're trying to model. Like what is fundamentally human might not be the most efficient thing to do, might not be the right thing to do, but things that make us who we are.

37:06 · I'm curious if uh there are some data sets that you really want that would materially help you. one version of this may be interesting. Which is more valuable to you to acquire as a data set? All of LinkedIn, all of Twitter, all of Facebook.

37:20 · Uh, you know, to be honest, it's it's a little bit hard to rank. Uh, in part because, you know, there's there's this product saying where no feedback is wrong because it teaches you something about your users. Doesn't matter what kind of feedback.

37:33 · I think it's a little bit like that.

37:35 · So, just whatever is bigger.

37:36 · What about a different domain? Say it was what about all of Amazon data?

37:40 · Shopping data, right?

37:41 · Shopping data. So Amazon data is interesting in that it's very much behavioral. Although like what people do on social media, you could sort of squint and say that is also behavioral, but the transaction data is always interesting. It is also most commonly available. However, if we were to look at purely social media, like if if you really, you know, if I were, you know, if I had to really pick, Facebook likely is interesting because I actually do think it is most sort of a default version of people because you go to LinkedIn, it's very much professional environment.

38:12 · So people put up their you know, you know, they have their guards up, right? And that still is interesting because that is true human attitude and behavior but it is not your base state. Uh you go to Twitter and Twitter people have their own crazy personas. Uh or depending on who you are like my Twitter profile and you know persona is very much initially was I was very much an academic. Hey I'm here to share my studies.

38:36 · Now uh I share uh things that's related to simile but Facebook is one of those more private space where people just connect with their friends. In that way I actually do think it shows you a little bit more about who that person is. So if I had to pick I likely pick uh Facebook.

38:53 · Yeah.

38:53 · And you're interested in like the whole person and their background and philosophy. I I guess is it too clinical or too machine learning oriented to just say this is just ways to inject varants and biases. The broad question I guess is like is this any better than a randomized like combinatorial explosion version. Uh so we have a link to the 10-centent uh billion persona paper where they basically did not do any of the groundwork that you are doing.

39:20 · Yeah, they just sort of did like a cross matrix of here's all the professions in the world. Here's all the people possible backgrounds in the world. Do a dot product across all of them and that's it. That that's your prompt for a billion people.

39:34 · Yeah, this will do something. I don't know if it'll do what you do, but it gets you some way some percent of the way there.

39:40 · So, this actually was an interesting paper. Like what I admired about this paper when it came out was the scale and obviously you do gradually want to be able to simulate really large societies and infractions. So the scale is definitely admirable. Um it is relying heavily on the known statistics that went into training the model. So to the extent that you believe that statistics is correct, this is actually not a bad way to go about this.

40:07 · But the thesis here and this is something that we also have seen in the market like if this works then we actually have solve simulation right because I survey like okay 5% of the the the US population is in construction

40:24 · the other 5% is in medicine whatever right and then you just keep going down the list and then you do the other side 5% has like you know the big five personality of like neurotic or whatever that's it that's it so if you believe that the underlying data set and the platform that we're leveraging has all the right statistics then this actually will have solved it. You're at that point merely retrieving the knowledge that is already embedded in the model in the model parameters.

40:48 · That's not unfortunately what we see where there's such detailed and also niche knowledge about people that if you just take one example it might feel very mundane but it's actually quite rich when you put together that you actually do need to do a lot of bespoke data collection to better understand people and this is also you know I think what makes this particular uh job fun which is you want

41:13 · to deeply understand people and the process of deeply understanding them actually requires There's a lot of attention to the details and you do need to pay attention to and pay respect to the daily lives that people lead.

Scaling Laws and Simulating 8 Billion People

41:26 · I want to talk about scaling simulations. So what can't we simulate?

41:33 · What can we simulate? And how does scaling affect this? So how big are the models? What if we go from you know 8b like couple hundred million like 100 billion parameters trillion? Do we get scaling any interesting emergence like at a certain scale at a certain amount of training you uncover anything unusual and any learnings from that? What we are seeing is at simile so we do post train our own model. The thing that we're actually seeing is the early glimpse of scaling law in simulations.

42:02 · The more data about humans and more compute you ingest, you actually start to get predictive and predictable gains of the model performance in simulating and predicting people.

42:14 · We need a scaling, you know, it's scaling well whenever you find it, it's it's a beautiful thing. Uh, and we're starting to see the glimpse of it, which is quite exciting.

42:22 · But if you talk about the ambition of simulation as as a whole, it's not merely about building a model. It's about building a model then creating the agents that become the individuals in a much larger ecosystem. So you're basically creating this multi- aent simulation. Down the line you want these multi- aent simulation to also live in a very rich environment. Right?

42:42 · What we are really trying to get to at that point is hey can we actually create I let's do a time machine game again and 5 years 10 years into the future can we create a simulation of 8 billion people living on earth. I think that's quite interesting and that really is the vision and once you get to that kind of state the kind of questions that you can help answer for the society also start to change from my perspective the answers are fundamentally about emergence of the emerging behavior of

43:13 · society and large groups of people so for instance the kind of questions that I get excited by and maybe this is a bit you know I have my you know academic side of me and and for me it's questions Like can we help solve climate change?

43:28 · If you look at climate change as a problem space, this is what we like social scientists would often call the wicked problems problem where you have many actors with competing incentives who are trying to make a very complex decision at coordinating that coordination decision very difficult to really solve in real life which is also the reason why we couldn't solve it. Can simulation help us solve that? Another one is can we actually understand the signals for collapsing democracy or can we understand or can we uncover the origin story of the monetary system.

44:03 · These are societal questions that we never really had a good way of answering if we can create simulations of our society. You have to believe that these are the kind of problems that we can solve. So that's really the ambition of this field. And you know, I also think yes, I mean, I think there's a Nobel Prize to be won there, which wouldn't be surprising. And I think there's some amazing societal impact that we can have to help people make better decisions.

44:27 · Nobel Prize in economics.

44:28 · In economics.

44:29 · I see. I see. We're rooting for you to write that paper one of these days. But um you know, one of the scholars that I was deeply inspired by um when I was coming into the space of simulation actually is the scholar uh named Thomas Shelling. uh selling point.

From Schelling to Society-Scale Agent Simulations

44:47 · Uh so the canonical example of the of the work that he's done was he was one of the creators of agent based modeling. So this was like in the 1970s and ' 80s.

44:58 · It's very early days but this was truly one of the first exemplers of simulations and one of the canonical model from that time and of course many of these simulations are trying to tackle the societal problems that's most relevant for their era. It was called the model of segregation.

45:13 · So racial segregation was a big topic uh that uh we cared about and what they've done was they actually created this grid world where they had red dots and blue dots and these dots were back in the day like they were the agents and they had a simple rule that governed their behavior. If certain percentage of your neighbors are of different color and if that goes above certain threshold then you move to a new location at random.

45:44 · One of the striking finding of this paper or this Asian-based model was for the longest time people thought the segregation within society was caused by explicit and overth racism. But if you look at this model, people's preference towards living with people of the same color, that preference can be very minute, but the very small difference actually causes the society to segregate completely over time. This was very counterintuitive for a lot of people.

46:15 · And this actually this particular work ended up informing housing policies.

46:19 · mixed income housing for instance got really inspired by this kind of work and Thomas Shilling ends up winning the Nobel Prize for having laid the groundwork for very early versions of simulations. The opportunity that I do see here in the more scientific terms uh is agent-based models for the longest um had impact in the in the 1980s, '90s to some extent early 2000s, but it has now sort of gotten forgotten by the community a little bit because as you can imagine, red dots and blue dots is not really a rich description of people.

46:54 · But with the emergence of things like generative AI and in particular generative agents, we do have an opportunity to create these kind of agent based models that are high fidelity enough to help us make really complex decisions and that's the opportunity that I see. If naturally works then yes and that is the kind of work that will result in a lower price.

47:15 · Yeah.

47:15 · For what it's worth, uh, you know, I grew up in Singapore. 80% of Singapore is in public housing and public housing has, uh, enforced racial quotas for exactly that reason, which is really interesting. Um, okay. So, we talk about scaling, we talk about all these, uh, the the the sort of agent possible applications.

The Cost and Economics of Simulating the World

47:36 · I'm scared about the cost.

47:38 · Uh, if you even let's just keep it to the US, not 8 billion people. Yeah. But uh how much does it cost to model so many hundreds of millions of people?

47:49 · Often times today obviously we don't start at that scale this stage of the of industry and simulation as technology but we can actually get to our users extremely rich and meaningful insights even by modeling thousands tens of thousands of people. And today what we do is every week we are collecting data on the scale of tens of thousands people with data and we actually have panel partnerships that gets us to tens of millions of people globally. So that's what we do today.

48:18 · And just as a a side note once you've collected one person for one study can you reuse that same person for all the subsequent studies?

48:26 · That's exactly right. Okay. The beauty of this model and these agents is the fact that they are domain agnostic. that what you're really trying to understand is what is the fundamental nature of these people what's their social physics and obviously there are a lot of a lot about people that does change over time like even like even things like uh how many times have you gone to have you been to like CVS the past week obviously that will change but there's so many

48:50 · traits about people that are also known to never change like your risk tolerance doesn't really change over time it's very consistent um so it's these kind of things that we're trying to But the scale we are operating is right now hundreds or um tens of thousands to hundreds of thousands and in many of the core use cases that we uh we are deployed in and this is more than enough population uh to cover those. Really at that point what you care about is less the number of people but more do you have the right sub population of interest cover it.

49:23 · And this is also the reason why people want a larger sample. It's not because they actually want uh stronger statistical guarantees. It's more that can they actually filter down to any population of their interest.

49:35 · However, you can also imagine in 10 years if we truly believe that the compute is going to scale that we'll have much more availability for compute and our ambition for simulation is also going to scale accordingly. I mean there's definitely a reason for us to create an entire data center worth of simulations or in my hunch here is I do think in the

50:02 · next some number of years we will start creating simulations that will actually cost as much as training a foundation model but perhaps it's going to be so valuable to the society that it would be a no-brainer I mean right now even today like we are training bunch of new foundation model just so we can say we trained on and we spend tens of millions. But if we can create a simulation at the level of society that would actually solve climate change, I would run that today. I would raise the money right now just to run that.

50:33 · Amazing. Um I guess the followup question is does it also compound if you let the simulations talk to each other or do they already do that today? They don't, right? As far as as far as I understand, it depends on what kind of simulation you're trying to run. Um, in the multi- aent simulation setup, the agents do talk to each other, right? Which is exactly Smallville, right?

50:52 · That's right.

50:53 · But a lot of times, for example, in e-commerce, you're just by yourself, so there's no point talking. Um, but there's always levels, right? Like you decide what you will buy based on what other people around you buy and talk about, right?

51:07 · It depends. Again, I'm coming at this from a cost point of view. I'm like, oh my god. Like if there's like some combinatorial thing of like thousands of people talking to thousands of people then that 1 million X's might cost. I have a very different view of the cost side. Like running these studies in reality is actually a lot more expensive, right? Running any study like this is you got to have people do it.

51:29 · You got to sign people up. It's it's very expensive and sometimes like not feasible to actually run the study.

51:37 · But the outcome or the decisions you make are very expensive on them, right?

51:42 · So spend X million on something that you know the overall process cost 100 million might as well right there's there's a lot of value to be had there it's a small cost but I'm excited on the cost side actually to some extent and obviously when you deploy technology you often want to deploy in a way where you can replace existing budget or you can basically make things more efficient and that is the best way to deploy.

52:07 · However, the way you capture the long-term value of the technology actually is making an argument that no, it's actually the upside that by making this better decision using simulation, you have saved yourself or made yourself hundreds of millions or or even billions of dollars and that's a case to be made.

52:29 · Random tangent question. So if you're doing a lot of inference, a lot of model multi- aent stuff, are you at the point where it makes sense to, you know, train a model that's, you know, very sparse, you're expecting to do multi-million dollar runs. Are you thinking about this in model architecture standpoint or inference efficiency or, you know, you're still at the research phase of it works, it works. We're not super there yet.

52:57 · Efficiency we actually do think quite a bit about. I mean this is technology that is deployed now in some of the largest enterprise companies in the world and we do process significant number of queries uh that are trying to you know assimilate the populations in the world. Uh so efficiency is a consistent thing.

53:17 · Obviously we don't want to overoptimize too early. So I wouldn't say like this is the the the higher bit right now but this is definitely something that we we think pretty carefully about.

Real-World Use Cases, Synthetic Populations, and the Market

53:28 · Yeah. Are there other other case studies? So you talked about CVS, talked about Gallup, Deote, Wellfront.

53:35 · Wealthfront is an interesting one. Um because one of the things they're trying to do, they were one of the first customers that wanted to actually do product testing that goes beyond just asking people what they think about let's say behavior experiments and so forth. So there really what we had to do was reason about multimodal input. So images, but also you can also imagine like these agents traversing through Figma mockups or websites.

53:59 · So some of the things that our agents can also do is you can be given a domain like or like a website URL and actually go use it for a while. It's these kind of things and Wolf was one of the first uh customers and that was very excited about this possibility.

54:16 · Well, have people been asking like is there any demand that we have not covered like UI testing, right? I want to try a new I want to ship a new feature, test the UI, simulate how people will do it. Any any interesting things that you're seeing demand for today?

54:33 · A lot of the demand does come from basically like the places where people have historically used human panels. We can basically now replace with agents uh and the synthetic populations and this is obviously not replacing human panel.

54:50 · uh in many ways the simulation that simile is building is grounded. So the way that I think about this is we are trying to represent humanity at scale and in that way the use cases are what we would expect but it's the scale of deployment that surprises me. M turns out there are so many decisions that people make every day in these organizations groups and we want to be

55:14 · able to say we listen to people we have consulted our users but in reality that is rarely the case because getting to people and actually asking them many questions it's it's difficult it's both costly uh time consuming but most importantly people are just not available if I had to answer thousand

55:34 · survey questions for this one particular uh vendor even if I wanted to do that like I would never do it and that's very much the case what simulation can do is ensure that the voices of people is always represented in rooms where the decisions for them is made right so all the stakeholders of this particular product launch ideally they're consulted that's what this technology really is trying to enable in my mind that means it's skewed towards more consumer f focus right?

56:06 · Like anything with a wide enough customer base where you do benefit from the diversity that you represent. What are some rough statistics just for people who are not familiar with this market in general?

56:17 · What's the market size that I'm sure you have some like rough numbers? Obviously market size is like a vague question.

56:24 · Yeah.

56:24 · But like how much do people spend?

56:26 · So market research is a hundred billion dollar industry.

56:29 · Yeah.

56:29 · Um but the thing about simulation is simulation is not a tool for market research. simulation is a tool for human decision-m.

56:38 · So the the question around what is a TAM here is actually quite tricky, right?

56:42 · Because it's easy to say, well, market research TAM is roughly 100 million or 100 billion. Uh so is that a TAM? And not really, right? Because in many ways, you're trying to inform all human decision- making. You're trying to basically inform every decisions that are made about human for humans. What is a T for that? It's really unclear. And I I I I'll be honest like you know I have a scientific background. I have a research background.

57:08 · So I didn't come into the field actually calculating oh what is the time for human decision-m but I just had to assume well if we can inform every decision that is made about human for human that has to be big some something valuable.

57:22 · Exactly.

57:22 · I mean to some extent you know you are a unicorn founder now and you you have to care as a CEO. Uh but like I I I do think like yeah we go into these boardrooms with people that you're quoting millions of dollars of contracts for like you have to say well well here's what you spend on humans and here's what we save you and it's 85% similar and certainly the the value case is something that we care deeply about like what is the value that we actually uh provide to the users and the decision makers but this is also where like you know as a a founder it I think valuation

57:54 · only tells one very superficial aspect of the story and I try not to think too much about valuation in general because that's not what also motivates the team or certainly doesn't you know I'm I again the interesting thing about researchers is we are happy living in academia getting paid next to I mean we get paid okay I mean we don't get paid that much I mean as a researcher if you're in academia but it's the impact and it's the it's the value that we can

58:23 · provide to the individuals in the society that really drives us and in that way ultimately what drives us is the impact. Does the simulation we provide have a real impact in people's decision-m in ways that progresses our society forward? If the answer is yes, then yes. I mean that has to be great business and we see that in numbers and we do care deeply about that upside story but that's the higher bit.

The Future of Simulation, Painting, and UBI

58:50 · Do you have any timeline predictions? So we talked about scaling laws of simulations you brought up. Okay, maybe one day we can simulate how to solve climate change.

59:00 · Uh where are we now? If that's not the end state, what is an end state? And what what does progress look like? You know, so what I sometimes tell people is simulation as industry. It feels a lot like where Gypty 3.5, Gypty 4 was uh for

59:16 · the AGI saga which basically is we have now technology that is powerful enough to do real damage on the verticals that we are tackling at the same time there's a lot of progress that is yet to come and that's I think where this is so the way I see it I do think there will continue to be breakthroughs both in data and obviously in algorithm And there will be much more aggressive scaling that will also happen over the next few years. But I think that's roughly sort of where we are.

59:50 · I think that was about the the the rough set of topics. Anything else that we should have asked you or you you wish people asked you more about about simile? You know I think the what's for me what's actually quite fascinating fascinating about simulation it is very impactful technology but actually is also very interesting technology both

1:00:15 · in terms of like what it means for human society our philosophy and the way I sometimes interpret simulation is so going back to my background I actually as I mentioned earlier I started my career as a painter. Uh it was a professional pursuit. Uh and I actually did or painting uh for figures. So I got my training originally in sort of the realism studios and that's what I spent a lot of my uh years uh doing.

1:00:45 · Simulation is a lot like painting, right? And the best paintings teach you something deep about the subject that you're trying to represent. And it is always not a perfect representation. it. No painting is perfect. There's always some small differences and discrepancy. But what it does is it tries to highlight the thing that matters the most about the subject.

1:01:11 · The essential essententral essence.

1:01:12 · Yes. He you uh he's brought up some of your work.

1:01:16 · Just nice to put it up.

1:01:17 · Yeah.

1:01:17 · So these are some of the works. So this is actually from my personal website that I maintain when I was still a researcher. I think a lot of people will say like, you know, like a Picasso, like anything postmodern is like very much focused on the essence.

1:01:31 · Yes.

1:01:32 · Right. Um Yeah. But I don't know if any any one of these evokes something that you like to tell the story of.

1:01:37 · No, it's it's one of those things where, you know, each of these paintings, drawings, whatever may be, it is trying to surface something about the subject that you feel deeply about onto the surface. you know, when I was a painter uh and artist, the topic that I cared really deeply about actually was uh the more more mundane aspect of human lives.

1:02:01 · This actually shows up in some of the some of the work that I've done uh where like I did this entire sort of study of a a rural town where I basically went around and took photos of people for not really doing anything special, but just living their everyday lives. I thought that was the most interesting thing. I'm somebody who has this perspective where you know the world is oriented around this fractal shape and you have two choice to understand the fractal shape.

1:02:32 · You either go outward and try to explore as much as you can to understand the broader shape of the fractal or you go inward because you know the outward resembles the inward uh shapes and understanding the mundane aspect of it was very much that simulation has a lot of this right you're trying to understand even the most mundane aspect of people when put together teaches you something really deep about that individual and the society.

1:02:57 · So I think that's what's interesting about simulation sort of the way the same way that AGI helped us better understand or really think critically about humanity and human intelligence. Simulation is really an exercise of understanding more about human society and our collective lives. So that I find to be particularly interesting.

1:03:19 · Yeah.

1:03:19 · Now you're reminding me that some of the best biographers, uh, documentarians and even photographers, they're taking a photo of you. But before I take a photo of you, I must spend I must like follow you for a week just to understand you, you know, uh, which some artists uh, some do. Part of your work uh there's a very famous book called working. I don't know if you've uh, been referred to it before.

1:03:41 · Yeah.

1:03:41 · uh it's very very famous like you know to the point of having a Wikipedia page about this kind of like really in-depth understanding and interview of people as they about them about their lives which seems mundane but is told in a very uh compelling way. Yeah. 1970s as well.

1:03:57 · Okay. It was an amazing decade.

1:04:02 · Actually before closing question uh you said that you started similar question right? If we do that now 10 years down, what what can we simulate?

1:04:12 · What would you simulate if like if you've made significant process? Are there any questions outside of the ones that we brought up? Anything that you think is most impactful? Anything that you would go vision 10 years out?

1:04:26 · In many ways, as I mentioned, I I am somebody who is very much impact driven. So the what would actually inspire me is I would want to ask 10 years later what what would actually be the most important societal question that we as a society have to ask. I would love to tackle that. Like for instance, do we need UBI? That could be an interesting one.

1:04:46 · Oo, has anyone done that?

1:04:48 · Well, I mean, you know, we're thinking about it.

1:04:50 · Can I get access? Can we just OpenAI? This is like just trivia now like opening or I think Sam Alman actually funded a study on this in Africa and the answer was no.

1:05:00 · The answer was no. But what was it something about the implementation?

1:05:04 · Yeah.

1:05:05 · But but this is the thing.

1:05:07 · See when Samman funded this particular uh he spent $14 million quite a bit. But this is the thing. This is the reason why you want to run a simulation. You spend 5 years, $40 million on this one study and have one finding. But if you can run simulation many many times instantly, then that's the value.

1:05:30 · I feel like that one you could have done in a simulation. Like if you can do the housing study, you can do the UBI one. Like I mean come on.

1:05:36 · I think sometimes people will spend the money because they want to verify what you think, right? Like sometimes you just want to is it actually is it actually right? Like you got to test it.

Are We Already Living in a Simulation?

1:05:46 · Okay. Closing question. What are the chances we are in a simulation right now?

1:05:50 · It's a fun question and I I start at some point I just answer yeah we're definitely in a simulation. But what I do uh feel however is uh whether we are in a simulation or not that I don't think that makes our experience any less real and I think that's fundamentally like what I believe in. Maybe we live in a simulation maybe not but it's real to us. Yeah.

1:06:12 · Yeah. For me I don't really care.

1:06:13 · Yeah. Unless you die and you wake up in like the level higher.

1:06:18 · That would be interesting.

1:06:19 · I I feel like you wouldn't care. You know, once you die, then you find out you like I worry about it when I die.

1:06:27 · I think the other thing that I Okay, so I like the mathematical answer to this, which is like the uh sheer number of possibilities that you are in a simulation far outweigh the sheer number of possibilities that you're not. Yes. uh except for the simplest answer which is uh it is computationally very expensive to have you be a civilization.

1:06:47 · Um okay great you've been very generous with your time. Congrats on all your success. Uh you know I met you just after your smallville paper and had no idea that you could build like such an enormous company and then now you're like well it's a hundred billion dollar market but that's just where we're starting. So this is uh very exciting. $100 billion market was not the TM that was only part if you're thinking too small.

1:07:11 · Well, I I do believe that I made you my final note here might be again I love science fiction. You look at any advanced civilization in science fictions there's two twin pillar uh technology one's AGI in some form and the other is simulation. So I think the the market's pretty big here.

Building Simile and Hiring

1:07:30 · Yeah. Tell us about the company. You guys just raised a lot. You're half a research lab, half a company. I guess you're hiring. Where are you based?

1:07:38 · Yeah, so we're based in Mission Rock. Uh so not too far away from where we are right now. So we're in SF. Uh but we're also by Coastal. So we have our uh team uh I I would say our headquarter is in in SF and we have a lot of our technical talent in SF and we do have a smaller office that just opened up actually in New York.

1:07:55 · We are as a company an interesting one in that today obviously there are AI neolabs and then there are AI product companies similarly truly is both so this is a company that was founded by four co-founders myself Michael Bernstein Pong Laney Ellen uh

1:08:11 · Michael Percy and I are all researchers so of course Michael was one of the co-authors of the imageet kickstarter the AI revolution back in 2013 has been instrumental in human center AI peri coined the term foundation model and obviously so you know one of the great of the AI researchers today and Laney is my business counterpart where she lends all the fastest growing AI native companies from their C to AMV but we

1:08:35 · have this DNA at the company where the vision of the technology that we're creating is continuously developing that we are getting people who were basically my labmates we are right now about 60 or so people 15% almost 20% of the company population

1:08:55 · actually are just my labmates from my person's lab and we're it's actually quite fun because many of them then had gone on to open AAI uh Google Gemini and these places and so it's been a few years since we really got together and had a chance to work together but now they're coming back and really building out this vision that I find to be quite exciting and that excitement is shared

1:09:17 · so there is that motion at simile where we are group of researchers trying to do something that no one is working on that we find to be the most impactful potentially but at the same time this is again technology that can make impact today. So we have an amazing group of engineers, product people and designers uh who are sitting here with us basically trying to imagine what does it look like to help people understand what simulation can do and make real world decisions with this having both and then deploying it to some of the largest customers in the world today.

1:09:49 · It feels quite unique. Yeah, it's very compelling. One part of it was this is the call to action like who are you hiring? You've done part of it which is you you've got a very talented group.

1:10:01 · Who are you hiring?

1:10:02 · Like what uh what roles?

1:10:04 · So honestly at this point we are hiring a small uh section. Uh we are always excited to bring on uh amazing research talent. Um so if you're interested in working with you know our lab mates, we're always welcoming of amazing uh researchers. But also we uh hire uh amazing engineers and that some of whom I like I respect the most. Many of them actually come from places where we have personal connections with. So many of the members are from Figma, notion, Harvey and so forth.

1:10:36 · Uh but also more broadly from the companies that we as a team heavily admired. So engineers both on the product side infr side uh we're all looking for those hires.

1:10:47 · Well, lots of people I think you make a really good case. So, um, thanks and, uh, we'll see you in the simulation.

1:10:53 · Amazing. See you all there.