Transcript

Intro

0:00 · A fun story that you recently shared on X as well is how you were part of the team that built this internal Google bot that was Chad GPT but a year before Chad GPT that caught up like you know wildfires.

0:12 · It felt more than a research [music] project for Codeex you built it in Rust and at the time the model was not on distribution for Rust.

0:20 · Turns out it was quite clear that Rust as a language would actually be quite good for agents fairly quickly if we decided to put some effort into it. When someone joins the Codeex team, what do you tell them? How do things get done here?

0:32 · The thing that they hear the most about when they have a question is like, "Have you asked Codex?" It still surprises new starters that you can basically ask it anything.

0:40 · Building was the fun part, but then maintenance was the painful.

0:43 · So maintenance is really sort of like a tax that you pay over time just to keep things running. But where I think it changes is like a lot of it is just going to be automated.

0:51 · So you still have the concept of code review.

0:53 · The role of code review is changing. And the role of code review now is like I think Codex is one of the [music] most popular AI coding harnesses today. But how did it all start? Many of you will know today's guest Tibo [music] from his generous and pretty frequent Codex usage resets. He was also there when Codex as a product started and has led the broader Codex team since. Today we cover how Codex started and why it was built in Rust and made open source. How coderies are changing inside the Codex team and open AAI.

1:22 · What it means when maintenance and rearchitecting are getting ridiculously cheap. what the merge of codeex into chat GPC looked like [music] and the many underappreciated engineering challenges of this project. If you want to understand how teams inside of OpenAI plan review and ship software, this episode is for you. This episode is presented by Turbopuffer, a ridiculously scalable, fast and cheap hybrid search engine built on top of object storage by an engineering team that I've really grown to like after spending time with them.

1:49 · Turbopuffer is the tool that companies like Entropic, Notion Cognition, and Harvey all use to connect their AI products to massive amounts of unstructured data. When I've talked with engineers who use Turbopuffer, the theme that always comes up is reliability and performance at scale. The reasons for this have everything to do with Turbopuffer's architecture. Turbopuffer uses only object storage for state and MVME SSDs with memory cache for compute.

2:14 · Data in Turbopuffer is organized into namespaces. You can think of a namespace as a database table or a search index or an S3 prefix depending on the world you come from. When a namespace is not being queried, it stays on cheap object storage with no associated compute cost.

2:30 · When a namespace is active, Turbopuffer pulls it up into hot caching tiers, so queries are very fast. This design fundamentally makes it effortless to scale to hundreds of millions of namespaces. If you're building a multi-tenant AI product, every user and their agent can have their own dedicated search index without any overhead. And each namespace can hold hundreds of millions of documents without any special configuration. You can scale Turbopuffer virtually without limit. And the performance, reliability, and operating model all stay the same. If you need to connect AI to lots of data, Turbopuffer should be your first choice.

3:03 · Check it out at turbopuffer.com/pragmatic. Tibo, welcome to the podcast. So good to have you here. Thank you for having me. It's so good to see you again.

3:11 · It's good to do this. Last time we did it in person. Now, now we're doing our video. First, I I wanted to ask you, how did you get into tech? When did you first know that you want to work with computers?

3:25 · It's a good question. It was a long long time ago. Um my my parents actually decided to move out of Brussels where I was born and just just thought it was great to um just buy a small house and refurbish it. But it was in the middle of the middle of a village with not much going on. I think there was like roughly 200 people living there. Not many that I felt like I wanted to talk to or you know could make friends with.

3:54 · And so I kind of got stuck uh this is like very early like 8 8 years old. I kind of got stuck cuz like you know computers and you know it's like early days of like the the for me the internet and you know that was my way to learn about things and so just the rest is just like you know came from that. Um I sort of like I owe it to my parents to you know have moved into the middle of nowhere and then you know I had no choice but to get interested in computers. Once you you finished high school like you went on and went to university, right? Actually studying it properly.

4:26 · Yes.

4:26 · Uh I I studied mathematics, applied mathematics at university. I I went there quite quite early. Um and so I I graduated early as well. Like I thought for a long time um that I would actually not make it and that I would drop out. I had like small companies and small consulting business like uh like while I was studying uh I was like working for banks.

4:51 · I was working for I was like um very interested in supply chain and applied mathematics problems and I sort of like selling that and learning a lot through that. Eventually ended up in the startup world in Belgium. I did that for for a little while and then moved to London to work uh initially at Google and then Deep Mind and then now you know moved uh to be here at OpenAI like this California. I love the California weather. We can talk about that. Uh it's been very good.

5:19 · Right after university you started you you've founded a startup right? You had the startup bug in you or the entrepreneur entrepreneurial bug.

5:27 · Yeah.

5:27 · So this this startup was all about uh pharmaceutical uh supply chain um looking at the supply chain for uh clinical trials and like try to optimize and decide like hey you know should you produce more medicine where should you send it where should dispatch it like how do you avoid waste and through that making clinical trials more efficient and this was using traditional like nonML techniques um like traditional

5:52 · like more optimization solving Monte Carlo simulations these kinds things stochastic multi-stage optimization problem really and we also applied it on steel industry and we applied it to electrical grid as well in Europe um it's like anything that sort

6:08 · of had the shape of like an optimization problem we sort of like get interested in and you know to this day like this this this company still exists and I think they do some of the most interesting work uh still but it's changing a lot uh you know with with modern AI for sure but it's interesting because you kind of said like oh yeah that wasn't ML it was just the traditional stuff and then you go into like Monte Carlo simulation and optimization and this algorithm I get a sense that you kind of just went deep

6:32 · right that it was like okay like here's a problem space like how can I use mathematics stuff that I learned stuff that I didn't learn to just go deeper and deeper do I sense that correctly yeah that's that's why I was obsessed with applied mathematics is just really this idea of you have theoretical mathematics or you have theoretical science and physics and like there you just you you do it because there's

6:54 · something to be discovered and something beautiful about it and it's all about patterns and pushing the frontier but you don't necessarily always know like how you're going to apply it and then there was like the real world right this like you know there's all these cool problems that just lie around and I was like very interested in seeing like you know how can I make the world better and so like how do I apply like you know sophisticated mathematics you know to just optimize the world around me and that was like a lot of the thesis behind that startup yeah and then after a startup you ended up at Google and first at Google London it was in 2015 and I remember 2015

Working at Google

7:25 · Google is a really really competitive place to get into like maybe as competitive as open AI is today in terms of the industry or terms of prestige.

7:33 · You worked on maps initially and then you moved over to deep mind. Can you talk a little bit what what you worked on and and then why did you move on from a already really interesting space that you clearly loved you know like optimization logistics and all these things?

7:46 · Yes.

7:46 · I I didn't I didn't start on Google Maps. I started on a project that was meant to make the web faster. uh and make to me to uh make websites faster especially on mobile. At the time you know Google was kind of so like seeing the transition from desktop to mobile and like more and more traffic going to like you know mobile phones and so like wanted to get ahead of that. So funded like a number of a number of initiatives and projects. Um I was working on one of them.

8:15 · This was like really really fun because it was a small group um within actually the ads organization. It was meant to sort of like you know offset the the loss uh for the ad revenue loss because of the shift of traffic to mobile and worked on it for roughly two years. Um and then it was cancelled and although it was like the most fun I've had on, you know, solving hard technical uh challenges.

8:38 · I learned a lot from not having product market fit, not having the right users, not having the right feedback loop, not trusting your product manager when they say the project is going well when in fact it's not going well at all. And then you know one day it's just like this VP flew in uh from California and then it was just like oh yeah it's like you know we're canceling this project um you know unfortunately you only have you know hundreds of users and this is clearly not Google scale and then uh it's unbelievable but people were surprised um and I think there's a

9:07 · lesson there um that I that I carry with me of course is you know just always always question always go to like always you know deeply think about the impact that you're having but also like the importance of the overall project that you're contributing. And then I moved into Google Maps. Google Maps was super fun. Worked on reviews. And then after roughly a year, I couldn't ignore like Deep Mind. It was just it was this special place. Uh headquartered in London. So many great things were happening.

9:34 · This was like really the early days, you know, with rumblings of um things like AlphaGo and they just seemed to be doing extraordinary things and you know, just really tackling the very very hardest problems that you can tackle. And like with my background I was obviously drawn to that started there like I worked on uh a lot of like the research infrastructure research tooling.

9:54 · This is a theme that I carried on for almost a decade and it's like this is very much also the like how I approach things is how can I build tooling and products that help make others more efficient and bring a lot of utility to them. Initially I was doing this for research and then like over time you know I got like into thinking about things in a much more like more general and general and general way you know eventually like you know ending up where where I am now.

10:19 · Yeah.

10:19 · And a fun story that you recently shared on X as well is how you were part of the team that built this internal Google bot that was you know if you want to say similar to Chad GPT but a year before Chad GPT can you can you talk about that? That's a that that is a new story. I haven't heard it before.

10:38 · This was part of deep mind. There were like multiple efforts as well. There was like brain as well that was separate at the time. They had their own efforts on large language models, but it was definitely something that was being explored. It was not the main thrust of of deep mind. Deep mind was like very much worried um and and busy like thinking about grand challenges and games and you know thinking about RL not in the language sense. And so there was like this group um that was pushing on large language models and you know thinking about you know think what what if what if large text corpuses are everything.

11:10 · What if uh you just pushed language to its maximum and you just scaled language models like you know would that be enough to get to general intelligence? That was like a hot debate at the time and then one one group decided to just really push on that and then it it felt really natural like you know as I was building tooling you know with others for for research is like you know obviously you're like hey you know what can we do with this model like how do we present it you know to the researcher like how can they sort of like you know debug the inputs outputs and eventually you sort of like end up with you know like a chat system.

11:41 · So we built that internally. We had a lot of fun. Uh initially the models were like you know kind of like almost like a little bit absurd like you know not very coherent. Uh not super useful but it was a lot of fun um to sort of like tinker with them that caught up like you know like wildfires like you know this application is just sort of like you know everyone uh was kind of like u sharing little conversations within deep mind. It felt more than like a research project or like a research a project for researchers. And so then there was this desire over time to like launch it as an external product.

12:13 · But Deepine was just like not set up, you know, there was like the the right way to launch products at Google. There's like, you know, the whole machinery of like, you know, how you do that. Um, you know, the whole like blessed production stack. You obviously very very optimized over the years to do things well. Um, but also very very hard as an environment to truly innovate.

12:32 · And then I I I wanted to ask what made you you know look around or or or maybe even consider open AI but I feel you partially answered this question just just putting myself back into your shoes like you're you know if it's it's 2024 or 2023 you're inside of Google who are publishing amazing papers

What drew Tibo to OpenAI

12:49 · doing really good research you're doing super fun stuff right that pushing the limits of what what's been done before it's inside a company where you already moved you know for people who are feeling kind of comfortable or good about where they are right now which I imagine you must have been like what made you still explore all right like what else might be there?

13:05 · Yeah, I was I was very comfortable uh at my it's it's a good place, but really I had a I had a desire to, you know, meet great people, but also join a mission that I truly believed in and that, you know, I felt like the people were true to the mission and cared deeply about impacting the world in in a very in a deeply positive

13:30 · way, but also in a direct way, not not being like, oh yeah, it's just like, you know, we just do this work over here and then it's like it's the job of someone else to figure out you know how to how to make this useful. It's like I wanted to join a group where you know like all the parameters were sort of like considered together where you know research and product were like really co-designing. Open was just like crushing it. Uh I thought chatbt was like you know taking off.

13:52 · I was like I was uh I met a couple people from OpenAI and then I was like wait what you know you only have like 20 people working on Chad like that is that is an insanely small number that that must be like extremely empowering like you know how does that work? How do you manage to maintain you know a product with that

14:11 · level of scale um and with that level of autonomy with you know only only 20 engineers and then you know as I kind of dug and dug and dug and it's like it was just a an amazing group of people amazing mission you know super talented super driven and like it was it was like drew me in and then I joined pre-ereasoning uh efforts immediately

14:30 · like typical openi fashions like I joined it was like oh yeah you know like there's this thing going on like you know we're going to launch reasoning models like you know it's like some new paradigm and then you know start sprinting on that and like you know like a month later like the company launched 01 uh 01 01 preview and that was

14:46 · exhilarating to be part of I wanted to be part of like a place that moves fast cares about impact would be in tune with the world and you know just really listen um and sort of like that's also to me like you know what I've carried with me like when when building codecs when building products is like having a community listen to the community just

15:06 · really focus on like a really intense feedback loop uh and then building something that is just like you know you just really want to care about it and like you know care about the utility of it that it provides to the world and then of course you started to work pretty quickly on codeex. So you joined in 2024. Can you take us back what the thinking back there when you joined was about AI or LMS and and code? I I know there was this ASWE um effort back then.

The early days of Codex

15:35 · The we talked about it in the deep dive as well that we did in the pragmatic engineer the autonomous software engineer AS3. Yeah, that's what it's it was pronounced internally. Um we we don't we don't have an A3 effort anymore like you know it's it's it's codeex. Um but really for me it was I I joined I started building infrastructure for research like my a lot of what I did before was large scale um data storage analysis and then tools to understand training runs. I I I did a lot of different things over my years.

16:07 · Um, but it was always about building for others and making them faster and just really caring about, you know, fundamentally doing that well and then through tooling and infrastructure making new things possible. And so when I joined OpenAI, I was like with the same idea and then with the re with with one preview and like you know some some of some some of the later models, it was very clear that we had to use the models themselves to help us go faster.

16:31 · And so I just really got obsessed with this idea of what were the limitations, how were we going to use those models for research itself. So got together with um other folks in research. We started training models. We started building little agents. Those were truly the precursor to to to Codex.

16:49 · And like this was like we were training internal models to be very uh proficient on the Python codebase of OpenAI and then uh very proficient with you know having like good good taste in architecture, good taste in you know like code style.

17:08 · It was like Python only and then the idea was like you know we would sort of like use that to build infrastructure very quickly and you know help researchers code faster as well. uh and then you know and then we would move faster and then over time when you just kind like push that and simplify it to its core you're making a lot of uh you know we found like you could make a lot of progress very quickly and then learn very quickly and then Greg and Sam are

17:32 · you know people with they're immensely supportive and also uh Greg was very adamant that you know we would we would not just focus on ourselves but we would also focus on benefiting uh the world and so he just sort of encouraged that we would be thinking about this not just as a tool for OpenAI itself but also as something that we would actually make into a product and this is when uh we

17:55 · merged this research effort with this AS3 effort um and we started building one thing and then that led to a sprint which was like the initial cloud codecs that we launched which didn't really have PMF because it was like a little bit too high friction and then we also launched the codeex CLI and we continued to push but it was always this idea of hey how do we get models to really help here.

18:17 · You mentioned that first you started to build this model to train on the Python code and and actually help build in for better but then you made this interesting decision where for codeex you built it in Rust and at the time the the model was not on distribution for Rust right it wasn't as good as in Rust and it was in Python or Typescript why did you make that kind of

Why Codex was built in Rust

18:39 · a decision was was it kind of like did you expect that it'll catch up or or you figure that performance is more important or because it was very counterintuitive most of the other harnesses built were actually not built in Rust. They were built on distribution on TypeScript or Python or something else.

18:53 · Yes.

18:53 · From first principles like we very early on we were thinking about the product interface and the agent as different things. So it was very important um to build the core of the agent in a way that was robust, that was secure as well, that was um you know engineered for efficiency and scale and having worked

19:25 · through projects over the years that go from hey this is a fun thing to like, hey, we need to scale this to the scale of like the largest data center um is the ear decisions early on are like really turn out to be quite important as

19:41 · long as you don't sacrifice too much of the velocity and so it's like it's a it's a trade-off but we had very prolific and amazing Rust developers our internal models were not bad at Rust um and then you get a lot of uh validation as well at compile time it's like you know statically verified and all these things and that is great for agents too so turns out you know it was quite clear that you know Rust as a language would actually be quite good for agents fairly quickly if we decided to put some effort into it. But primarily we were focused on correctness and we were focused on efficiency as well.

20:13 · Interesting. So you're saying you know it's it's worth in your case it was worth thinking ahead of where you want this thing to be and for example thing like a language choice. Obviously with agents you can rewrite a bunch of stuff and easier than in the past, but it's still like you can say save yourself reworking by putting in the right I guess scaffolding or or or well the you

20:33 · know the the baseline of of what you're building on right I think we could have been successful if we had written it in Typescript or you know maybe even Python and then it would have fine and then you know we would have rewritten it at some point but having a very clean separation between the agent itself which can exist irrespective of the product.

20:50 · Um it was a very important principle and if you write everything in the same codebase in the same language it's like inevitably you're going to be a little bit sloppy and um you're going to intertwine things more than you should and then it's going to prevent further innovation after that. And so that was that was very important like the rust boundary in a sense like was very useful for that.

21:11 · One interesting decision that you made which is unique across all of the major labs is having this built-in open source right the CLI is open source the SDK and the app server are all open source when and why did you decide that it's not a

Why Codex is open source

21:26 · given especially you know there used to be jokes about open AI having things closed but this this is actually the opposite where like this is open whereas like some competitors would would ship closed source harnesses which again I I I think it's very easy to understand why you would want something closed source why did you want it open source There was something really cool about the idea of having the code open source because fundamentally what you're building is you're building a coding agent.

21:52 · And so we were sort of like thinking about well if you have that you know you're obviously going to point it at itself and you know maybe you know you can build a community of you know contributors that use it to improve it and then you know you can learn a lot from that. Also, it felt at the time is like, you know, very clear to us that if we were going to be successful, open source itself would change and the role of code itself would change and so being part of that community seemed important instead of divorced from it.

22:21 · I think, you know, it's it's it's hard to solve problems if you don't sort of like witness them yourself. Uh, and then the the other thing was just it still feels like early, but it was very early at the time. Um, it felt like we would have some ideas for how to solve things well.

22:39 · Um, and we were co-designing these, you know, with with with with the training and and and and the research and it's it's all about expressing like the capabilities of model in like the most flexible and the best way, but also we didn't have all the answers and sort of being very open about, hey, this is what a good harness looks like. This is how we think about it.

23:01 · We did like a couple of like very technical like deep dives and blog posts and we talked about it a lot and we thought you know hey it's just like the world is vast out there's like you know crazy smart people it's like you know we're we're going to get inspired by other open source project as well and so let's just make this a level playing field and sort of like encourage a lot of tinkering um and exploration at this stage.

23:26 · Now this has been now you know like a a year later a year and a half later which is a very long time in right now in this AI time frame but looking back or taking the experience what are the benefits you've seen the kind of engineering benefits the engineering team's benefits from being open source and just honestly what are things that are kind of hard about being open source right like there must be downsides like just try trying to get an honest take on both sides yeah there there there are definitely downsides it it it comes at a cost Right.

23:56 · Um the the benefits are it's almost something to build in the open. Uh it's awesome to have like a small a small repo as well. Like whenever we hire uh someone and they join the Codex team, it's like they've seen the repo before. They've they've looked at PRs. They're like onboarding is done.

24:14 · You know, it's it's it's done. Yeah.

24:15 · It's like in onboarding is just like you use Codex to look at the repo, you know, with you and you ask them questions, but it's like it's not it's not a secret issue that you can get productive right away. We get a lot of good contributions. Although we get like you know an a tsunami of like random stuff as well.

24:30 · Obviously you and everyone else right open source is changing. I think this is one of the examples.

24:35 · That's right. And then to to me it just and and to a lot of the team it just brings a lot of energy to just be part of the community and like be directly contributing um not just saying that we care about the community but actually doing things that you know you can see it's it's costing us effort right. U we don't have to do it. the the downsides are, you know, it's it's separate from the rest of our code. Um so, you know, sometimes we have to draw like artificial boundaries and, you know, work across multiple repos.

25:00 · Um when we're working on something particularly exciting, um and uh you know, we're building it in the open then, you know, at times we find that uh you know, others copy it, you know, before we have the time to release it. And it's like it's just a little bit sad. Um, but also it's like it's part of the game, you know. It's it's like you're building in the open. It's like, you know, that's that's sort of like the contract that you signed is like, you know, you can copy it.

25:25 · Uh, we have a very permissive license as well, but it does sting a little bit when you're working on something and you're like, you know, and then uh the the third thing is just just like everyone else is like, you know, we are overwhelmed with, you know, random contributions and, you know, we have to deal with that additional tax. Um, but then that pushes us to, you know, also like try and solve for it, right? which I think is good.

25:47 · And on top of the open source, one thing that surprised me about Codeex, and I didn't even know about it until recently, it's not tied to the OpenAI models, you can use other models with Codeex, you know, like putting myself if in a vendor's shoe, it it might not be very obvious because again, all the other vendors I look at when they do a a CLI, it's kind of use it with our models. again what made you decide to be this permissive about you know using or allowing to use your harness with with other models?

Codex plays nice with other models: why?

26:18 · It it felt quite natural if if you are part of this community and building an excellent coding harness is like why would you couple it to your model? That that that felt like sort of like quite disappointing to make that decision. So it didn't it didn't feel right. Um and in general it's like I think you know it's like I kind of try to make decisions that I'm like yes you know just like I can just sort of like explain it you know it is correct. It's the same reasoning with you know it is open source in the first place.

26:49 · It would have been trivial for anyone to fork it and then add support for another thing.

26:57 · But then but then you're just encouraging people to just like you know go and use that fork and then now suddenly you have overhead and the only reason you have a fork is because you know you wanted to change like 10 lines of code to add support for like another model provider that feels very silly. So like you know why not just support it in the first place. The other thing is we benefit a lot from like being able to just give optionality.

27:18 · So, you know, it's like maybe today, you know, you you you love using OpenAI models um and you know, you're super productive with them, but like tomorrow there's a new model that comes out, you want to try that.

27:32 · Why force you to go and completely change your setup just to try a new model and then we benefit from the feedback that we didn't get uh which is like, you know, maybe there's something that you liked about that model. Maybe it actually didn't work well. But it's sort of like being nice to our users and to the community is like you know feels like the right thing to do here.

27:49 · Um and then you know we also you we also try like other models right so and you know we try them in the same harness and you know it's just all all good and then this is all also often like um this optionality is very important to companies that we work with. This is something that, you know, we we absolutely lean into.

28:10 · This last point, I think, you know, as any serious company, you want to have optionality and you want to use a tool that gives you that optionality.

28:16 · But I kind of appreciate it cuz I feels to me like it's kind of honest like look like it forces the whole company to be to compete the best in everywhere in the model layer and the harness layer with open source with with choosable models and it kind of like doesn't doesn't allow you to like kick back and say like, "All right, we're done. We we can we can we can hang back for a little bit for now."

28:34 · Yeah.

28:34 · I want us to win users by having, you know, the best models, the most efficient models, the best product, and then, you know, if we do all of these things, it's like we're going to have a good time. If we sort of like force you to use the product because, you know, this one thing, it's just like then I don't think that will attract, you know, the the very best people to work on this product either. And like, you know, it's like we we're doing our best work here.

28:56 · We care a lot about the experience. It should feel delightful, you know, like it doesn't irrespect of the model that powers it, it should feel delightful. I love the idea of winning based on merit, not based on lockin. And this is a perfect time to mention our season sponsor, Entire, who also play by the same rules. Like it or not, Git is becoming a bottleneck for modern agent heavy software development. Devs are creating more code with agents. These agents are pushing more code. Many devs are running more parallel agents. These are pushing even more code. GitHub is clearly struggling to keep up and has frequent outages.

29:27 · So what's the solution? Entire was founded by GitHub's last coza and he rebuilt git hosting for the agentic era from scratch. Entire was built to be very fast and to have your repos regionally close to you to reduce latency allowing for fleets of agents to push in parallel. Some numbers they published. Entire can handle 418 pushes per second. That's up to 89 times faster than every competitor on the market.

29:52 · When GitHub is down, you can still keep working and you don't even need to migrate from GitHub. You just sign up to entire and the platform mirrors your repo. And one more neat thing, have you ever wondered what prompt resulted in this specific code being generated? I find that the prompt and conversation with the agent carries more information than the PR itself, at least for me.

30:12 · Entire captures all the prompt history with your agent right in the repo easy to check back and has a pretty innovative UI to show all of this. If you're looking for Git hosting that works even when GitHub is down, head to entire.io/pragmatic, io/pragmatic, install the CLI, and mirror your repo with a click. I've already done it. Oh, and did I mention that it works with any agent and it's open source? I'd also like to mention our season sponsor, Anticys. Tibo talked about how the experience of the software you use should feel delightful. Delightful includes no annoying bugs.

30:41 · But when you're using agents to write your code, how do you avoid chipping bugs?

30:46 · Reviewing every line of code is becoming a challenge with the amount of code that agents generate, which is why antithesis goes well beyond code review. Antithesis runs your whole system in a hostile simulation. This simulation includes both targeted testing and fuss testing.

31:01 · By running this simulation, it finds every bug before your users do. And because the simulation is fully deterministic, it doesn't only find bugs, it gives you a perfect reproduction of every issue, which makes it much easier to fix issues. The first thing I thought when I heard about antithesis is that automated bug discovery and fully deterministic testing sounds like science fiction, but it's actually hardcore engineering under the hood. Jane Street, Fly.io, and the Etscd community ship agent written code with full confidence because they know it's been verified by antithesis.

31:30 · To see more case studies and details, head to antithesis.com/pragmatic.

31:36 · And with this, let's get back to Tibo and why competition between tools is great. Yeah. And I think as as an engineer like I always see that whenever there's competition as someone who's using tools it's it's always amazing like I remember like when Microsoft had with Jet Brains the IDU wars and then there's the clouds battling with each other with all the features and now of course we have the harnesses we have the models and as a as a as a user it's great because now we have more more choice they just develop faster I guess our voice gets heard a bit better. So it's it's great to hear.

32:06 · Speaking of the harness, can you tell me how it works today in the sense of like when I start a codeex task, does it run always on my machine? Does it choose the cloud? Does it use a sandbox? And how do I control this or or know this or how much should I know about this as as an engineer?

How the harness works

32:28 · Yes.

32:28 · So, by default, it runs uh in uh it runs sandboxed. Um it everything that if there is like an an a command that should run with additional permissions outside of the sandbox, it will ask uh you as a user for permission, but everything every tool execution happens within the sandbox by default and it runs entirely on the machine um your local machine. And this has been the case for you know more than a year now.

33:02 · But it is something that is evolving and shifting where like you can select to run this um in the cloud which then runs in like a managed VM uh where it's the same VM that you get through chatk work um and you can you can sort of like inspect it but like it runs in a kata container it's like a secure environment and so everything runs inside of that VM

33:27 · and it doesn't run on your machine and then the only thing that happens on your machine is like the the your input and then the streaming back of the output and so that you know obviously then uh is like much nicer on your CPU and and your machine and you can scale much much um much more and this is like just a step um it's going to be much more

33:48 · seamless in the future um to like you know use cloud machines and then you know maybe have a combination of like partial execution on your laptop partial execution on on on cloud machines and really The thing that we're thinking about that is very natural is as models just get better and uh more capable.

34:08 · They can leverage so much more compute and many more resources than are available on your local machine and so it would be a constraint at some point to just limit execution on your local machine.

34:21 · Yeah.

34:21 · One thing that is great about it running locally, and I think the reason I I love it when it runs locally. Of course, it's a pain because if I'm doing some work, it's like, you know, I have several agents, it's it's eating CPU. If I want to close my laptop, I I I cannot kind of leave it like half open, right?

34:37 · When I was in one of the offices of an AI company, I I had it half open and they're like, "Are you running agents?"

34:42 · I'm like, "Yeah, I have one running." He's like, "I get it."

34:44 · But the reason the reason I do it because I have my local tools, I have my local Postgress database. I have my my this this and that. How are you thinking about the cloud is amazing but it doesn't have this setup or it's just a pain to set it up. are you thinking or or are you experimenting with you know making the these setups and I'm kind of reminded of a topic that we talked about prea which is cloud development environments and like 2022 23 they're hot and then we talked about AI more but yes I think outside of large tech

35:15 · companies like cloud dev boxes really never took off because there's a very big upfront cost uh and then you need to pay like a maintenance cost as well and you just you know don't benefit from it as like a a solo developer or like a small team with the level of capabilities that we have in agents now is like the setup almost is free, right?

35:33 · So like this the setup cost and this maintenance cost is like if if your agent is capable of doing it, you know, it should just do it for you. So for example, if you're saying like hey, you know, I have like I have my local um SQLite or I have a local server and MCPS and whatnot. It's like how hard is it to actually configure exactly the same setup and keep it in sync on a cloud dev box? Well, maybe it's not that hard if the model just does it for you.

35:58 · And so I think we're going to see a resurgence of, you know, fully cloud orchestrated uh machines which then frees you from your laptop, right? It's like one thing that we've seen a ton of success with with chat work is like it's just available on your mobile. I start my day just dictating a bunch of tasks into it.

36:18 · uh next to the coffee and it just does it. It has access to my calendar. It has access to my my email. Uh it has access to Slack and it's just so awesome to just be able to walk around and you know get stuff done without having to you know carry my laptop everywhere and I think it's the same things like you know we shipped like Codex remote where you know execution is like still happening on your laptop but it would be wonderful if you know you didn't have to keep your laptop open. Can you tell me a bit on how in the past how did you improve codecs?

Harness and model improvements

36:48 · Because I remember when I first used codeex this was one of the early versions you know like you could talk to it did stuff but for example I said like all right make this change and it did that change and I had unit tests and it didn't run it and then later a few months later I don't know exactly when it just started to run it automatically.

37:06 · were these things did you improve the you know the the script that runs you know the instructions I'm not sure how exactly you call the you know the the bootstrapping script or whatever that is is it improving the model like as a dev how can I imagine you making each version better between the harness and then between the model and like what's the connection between the two yeah this is a good question so the uh the harness in a sense is always a little bit ahead of the model oh really how so?

37:36 · Oh, what what I mean by that is that you you have the model, it's capable of certain things, but then you set it up with like a couple of crutches so that it can actually do the thing um to a level of reliability and um in in in a way um that is like efficient and also with the behavior that you expect as a user.

37:58 · And sort of like that's the role of the harness, right? is like you know provide guard rails like safety, make it more efficient, make it more like steerable, controllable and then the harness usually is also responsible for you know what we call like the the developer um message which is sort of like infected in the context uh at the start of of each turn. And so that affects obviously like the the purpose of that to affect like the behavior of the agent throughout uh throughout the turn.

38:29 · A lot of what you have is like the result of of the harness and the model initially like maybe you're like oh it doesn't run tests. So you know you have to remind it to run tests and then you know we train a battle model that is uh just you know capable of like better reflecting on what is it that you really want when you ask for something uh and then you know you don't actually have to tell it anymore. So over time what we see is like the system uh the the developer message shrinks and then the harness also shrinks inside of the codeex team.

39:00 · Do you have specific goals? Do you say like all right now the the codeex as a hardness and and model to combine is not very good at this or it's kind of doing silly mistakes or here or how can I imagine how as the engineering team how you're

39:17 · working on the next you know version of of codeex is because the thing that I don't really get as a as a dev is like okay there's a model which to me is this magical thing which will get better of course I'm sure you have some feedback channels but you also have the harness which is the tools that you're building like that's probably what the team is responsible for how do set even your goals right like in traditional software you'll be like we will build this feature and you build that feature because you know how to do it but it feels a bit more fuzzy to me this this development process.

39:41 · Yeah

39:41 · it is and it's why we we co-design you know most things and it's it's a process where it's a collaboration between research and the engineering team like primarily building the the core the core agent harness. It's always

39:56 · uh it's always a question of like okay we see today that you know we are very good at this but we're not very good at this and you know we have a desire to do like another thing because it would be a very cool products feature and then you know we always like look at it it's like okay this should this be like a harness change or should this be a model change and if it's a model change like how soon can we have it can we have it in a month can we have it in you know three months six months and we sort of like work through that and then depending on you know how soon we can just fix it in the model at which level of training.

40:26 · Then we might decide to not even do something in the harness at all and not and just wait for for for the model to to solve it. You know, it's it's agents all the way, right? So we use agents to analyze like a lot of the feedback to like, you know, come up with themes, you know, to just help us have these conversations and decide on priorities. But we analyze it across all of coding.

40:46 · We analyze it across like all of like know the other domains like finance, coms, marketing, you know, all the things where our users are using these agents nowadays and there's like you know subcategories within those and then we roughly know like you know how well we perform and then we're always pushing the frontier and there's a thing that is interesting is like as we make you know as our pre-training model gets better as we make the overall model better like the whole thing lifts up but then there are sometimes things that we pay a little bit more attention to. You mentioned you know he analyzes agents all all the way.

41:18 · Can we talk about the the software development life cycle on codeex in the sense of whenever a new engineer joins a team any team it's like okay how are things done here and you know back preai it would have been you join the company like Uber or Google and they would tell you that cool the way it works is we have an idea or the PM has an idea we make a plan we get together we do some estimations we break up the work we code the work we do tests we do code reviews

The SDLC behind Codex

41:44 · we release we do feature flags and then you know we we were on call that that you know that's how it used to be when someone joins the codeex team you know they they've clearly been contributing to the open source part but what what do you tell them how do things get get done here if they're like a brand total newbie I introduce them to great people um and then the thing that they hear the most about like when they have a question is like have you asked Codex um and Codex

42:11 · like is just by default at openi is like plugged into everything so it has access to slack it has access to all the documents, access to all the code, and it still surprises uh new starters that you can basically ask it anything. Uh and it will very often just like come up with like a really good response. Um and so the easiest way to understand the state of a project or who's working on something or why a decision was made is like critics knows about it all internally.

42:41 · Um and so you just you just use all of that. We do a lot of work in you know for that reason we do a lot of work in public channels. Um we open up uh documents with like you know fairly broad uh permissions and so that you know everyone has access to this information as well and so that you know your agent can go through things and like you know reason through things and then you know we have a couple of other things that are just really uh very helpful for uh team productivity and team collaboration that we haven't released yet but are are going to come like some of it at def day.

43:09 · All of that just sort of like makes you very grounded and in tune with the rest of the team. Um, and allows you to like, you know, just very very quickly like understand the state of things and and and produce things yourself. The general recommendation is just like, hey, care about the user, uh, care about the coherence of the product, care about the models and where they're going. Uh, if you're doing something and you know you're building like this 10,000 lines of code crutch to work around the model flaws, like you know, you're probably doing the wrong thing. So, we have a set of principles.

43:38 · Um but it's just really um sort of like a team culture and ethos at this point and you know it's just very much so like carries on you know when people join it's just like through the rest of the team just like you know sort like teaching the ropes and then when I have an idea I think it's a good idea I I talk it through with codeex maybe I talk it through with some my colleagues like here's a cool new feature I'm going to build as my first first contribution or first major contribution at to codeex how do I go

44:06 · about that obviously I I code it down with codeex I obviously test it and make sure that it works from there on what's the process do you still have the concept of code review or AI code review of verification of rolling out of verifying of stage rolled out you know the things because codex itself it goes out to millions of people like I just crossed a big 20 million active user mark but if it's chat GPT then it also goes out to like even a lot bigger number of people yes um but it's it's surprisingly like a

44:36 · similar process whether you ship on codeex or chat even though Chibd goes out to like a billion a billion you know active users and growing you can ship a PR um you know you can make a change and you know get it shipped like the next day or like even the same day um and it just goes out to a billion users and it's fine we

44:56 · just really instill a sense of ownership and care so you're like people are very empowered to make changes even large changes the general thing that is being asked is like sort of like evidence that it's going to be wellreceived D evidence that is like a worthy addition uh evidence that you know it's like it is worth maintaining over time but also like the cost of maintenance is like just really as you know gotten done significantly as well. So we we we think about these things slightly differently than you know say like two years ago or 3 years ago. The other thing as well is like you know we automate as much as possible.

45:27 · So like a lot of like the process of like code review and deploys and you know catching regressions is like you know all of that is like pretty much automated. Uh and so like you know you get to just focus on just really the idea and you know how it's going to help our users and you care about you know the coherence of it all and so like the overall power of the agent um and making things better and like we don't we have a long long list of things that you know we sort of like aspire to do and haven't gotten to yet.

45:54 · And then there's like the sort of like the northstar direction um which is a delightful simple to use personal AGI that you know knows everything about you like you know that it needs to know has access to the right resources can take like you know sometimes risky actions on your behalf but then you know you get like the push notification and then you know you can verify that and it's like a thing that you know you deeply understand as a user but also it knows about your schedule it knows about your goals it can be proactive and it should be like extremely natural it should be something

46:23 · that you can control through like natural language, voice, you know, like maybe it should understand, you know, your emotions like if it has like a camera feed, it should be the most natural thing on earth. It's like it should not be like a thing with 10, you know, different buttons and configurations. It's like AGI should be simple to use. Now, you kind of mentioned just briefly the review the code review, but I wanted to go back to it. you you worked at Google on a on a product used by you know like hundreds of millions which is Google maps and Google is very well known for their culture of very strict code reviews.

Code reviews at Codex

46:54 · They have I think two layers of code reviews. There's a language correctness review and they they've taken I think they've really perfected it across the industry for for a long time and they do believe that it it works and they they use it. How do you think that part is changing specifically the human review? Because for a very long time until maybe a year or two ago,

47:12 · I would have said you code review has all these benefits that knowledge sharing the second pair of eyes removing the bus factor because now someone else understands and when that person is out that they can jump in conversations are happening about architecture not just not just the code but now there's you know there's a lot more code uh and what was the value of code review? What in what cases?

47:35 · And so on your team, because you guys are so ahead of this, where do you see humans still being or developers being involved in the review stage valuable? And and where is it fine? Did you find it find it fine to uh hand it off to an agent?

47:52 · Yeah, the the the role of code review is changing. One one of the early projects that I did on Codex was like working with research on developing a code review model um that was going to be to a level where it can spot

48:08 · mistakes in logic and reasoning. uh to a degree where it would require humans like you know multiple like potentially multiple hours to capture the same level of mistake because it requires like really digging like you know three four levels deep into like the dependencies and like you know understand that maybe the documentation actually was wrong and like the implementation of like this third party dependencies like different from what you expected and so therefore your invarants are not upheld um and

48:34 · these things it's just like you know unless you're an expert in that library you wouldn't know uh and therefore you have a bug and so we developed like these uh code review models and you know we we we released them and now they're like the same level of like capability and like ability to spot these mistakes by doing like you know deep verification are like just part of the mainline models like when we benchmark them it's like they're like super human in code review and this is not just true for correctness.

48:59 · This is also true for security for example where you're capable of like reasoning across like you know very very complex things and then you know coming up with like hey you know you have a critical security vulnerability here which is now mandatory across like all of OpenAI pull requests like we block pull requests from merging if you know we flag them with like a a security issue and this is

49:19 · like all automatic and the role of code review now is like I think it was always about correctness it was always about you know ensuring that things worked but it was also sort of like a little ritual for information exchange and you know bringing people on the same page and like you know encouraging like a discussion which ideally would have happened before but sometimes it just only happens like around the code because once it merged it just actually runs in production it's doing stuff and then you have to maintain it so there's like this social aspect to it as well it's I think all of it is changing like the correctness the cyber the the

49:49 · security is like I think that will be automated really what we see and I see is there's a sort of um really discussion around the intent that takes place around the poll request. It's like what are you even trying to do? Um and is that a right thing to attempt to do?

50:04 · I think you can have that discussion outside of the poll request. It doesn't have to be around code.

50:08 · So, so maybe this helps crystallize the you know like where a discussion needs to happen versus versus where we we did it because maybe we didn't have the type the type of tooling that we have right now. Yeah, I think this is going to change and it was like a forcing function because you know you have to have that discussion where like it's good to have that discussion before you merge it and it becomes production code but I I think there are other ways to

50:31 · have these discussions and you know design things together and make sure that the intent is good uh and then the code doesn't matter as much and it's interesting because when I think back of all my code reviews like of course I have like memories where like it was great we we had a good discussion or I learned something really interesting but a bunch of times honestly It was such a pain in the ass.

50:50 · Like I I was trying to get my stuff. You're paying. Hey, could you remove my code? And like, no, right now I'm busy. No, I really need this to unblock me.

50:58 · And then you context switch. And then I feel it's always been like good and bad, right? So I I feel whatever we do there will be always upsides and and downside, but there now they're just moving. So I guess one upside is as an engineer you might have to not give your attention to just kind of basic stuff that doesn't need your input per se.

51:17 · Yes, it saves time and I think progressively what we're going to see is also like you have an agreement on you know the box and the overall contract of what it's supposed to do and then you know what is inside the box as long as you have like strict guarantees in terms of resource utilization, data access, um security, these kinds of things.

51:39 · It's like what happens inside the box is you know it could be literally anything. it's like don't really need to care and like really what you need to agree on is like what does the box actually do and what are the invariants that must be satisfied and I think that is then worthy you know having like a really good conversation on you know maybe assisted by by your favorite uh agent but then once you have that and you have that understanding it's just like changing anything within the box is like you know doesn't require for discussion and it's like you know just really preserves your attention the cost of maintenance has gone down

Maintenance and architecture

52:10 · you know maintenance is always such a hot topic whenever we build thing uh inside of all these companies like Google, Uber, even startups like building was the fun part but then maintenance was the painful and that's when we learned like okay it was not we're building it etc inside of codeex and open AAI what do you see maintenance becoming cheaper changing in terms of instead of what you're building what the ambition is the I guess custom tooling

52:35 · those kind of things maintenance is really like sort of like a tax that you pay over time just to keep things running and it's It's it's it's always been necessary. It will continue to be necessary, but where I think it changes is like a lot of it is just going to be automated. So, you know, it's like okay, you have you have this third party dependencies like you need to upgrade the version number.

52:53 · It's like a like [laughter] you can fully automate this you know um if you have good change log and you know and the code is well documented and like you know and and the model can just like reason through it like you know you can just like blast through your codebase do it uh in a couple of hours and you know previously you would have like punted on it because it's not the most fun thing to do but it's actually really important for your business. So it's like really important for your project, you know, especially for security vulnerabilities.

53:22 · You want to stay up to date, right? You want to apply, you know, all these patches.

53:26 · Um I think that's just going to be fully automated. So a large part of like maintenance, it just kind of comes for free, right? And then um I think it's it's awesome to also think about before like you know when you wanted to just completely react re you have to do like a new architecture because you're trying to make space for like a new you know

53:46 · different kind of trade-offs or you have a new understanding of like the workload or you're trying to fit a new feature and like suddenly you realize like your current system is just very limiting and you need to completely rearchitecture it that was like a really really costly endeavor right so you know like sometimes like multiple years and I think this is also like super super accelerated now.

54:03 · So you like the cost of mistakes uh you know I would say like you know is going down but then at the same time the good old rules I would say of software engineer of like you know having good abstractions like really help like you know is going back to this like having the box with invariance like you know if you sort of like draw the right shape you're going to be able to change things much more quickly within the box and like not affect the rest of the of the services or the rest of your

54:28 · infrastructure and I think it's important it's important to design for very quick iteration and I remember when I talked with Peter Shamberger that was before he joined OpenAI but about open claw and how he thinks about it like you know he told me that he doesn't read the code but he kept thinking about like I could see that he's holding the architecture in his head and he was telling me how he rearchitects a lot and he thinks about how to make it modular how to allow 100 contributors to each build their thing without stepping on each other's toes.

54:57 · So I'm hearing what you're saying that this this care this this planning this this structuring has become maybe just a lot more important to like which which which was which which was something back in the day you know it was like the architect or the staff engineer or experienced folks were doing this thing and other engineers around them were kind of building this you know smaller parts but it sounds like now all engineers need to be aware of when

55:22 · you're building your software right and plan for it yeah and and the GP models are getting better and better at this as well of like you know thinking about long-term maintenance and like good architecture and like this is like a natural sort of like next step right. It's like not just about code quality in the sense of like oh is this code clean within this file but like you know is this like is the architecture actually correct to reduce maintenance burden over time and like you know make space for like future u

55:45 · product or feature extensions or changes and just really this act of like you know engineering over time that's uh kind of like something that models are starting to become capable of like thinking about very well I think it's just kind of fascinating to understand that the software that we're building is just going through the life cycle much much faster. Right?

56:03 · Like you know before you had you know you were you were scaling you were starting it you know maybe as like a small team of you know yourself maybe a couple of engineers and then you would add engineers like slowly and then you know maybe after a year you know it's like if it's very very successful you would have 50 engineers on it or like a 100 engineers on it.

56:18 · You would have time to see it coming. you would have time to see like you know the humans on board and you sort of like you know you can think about the documentation all of that stuff but now it's just sort of like that explosion of like you know suddenly you have like a 100 agents contributing to this thing is like you know that can happen like you know in a weekend and so you know you're just going through it at you know major major speed compared to before okay but how do you and and the folks at OpenAI like deal with this like does it not mess with your mind like you know

How AI tools expand what engineers can do

56:49 · what I mean in the sense of like you you you've been in this business for quite some time now like like decades or or well over and there was a pace that we kind of got used to and obviously is it's now a lot faster but how do you get your head around the fact that a it's faster b the stuff that you've been doing a year ago right now you're not

57:09 · doing because now the model is is good at it and you know like how do you kind of reconcile that because there I'm sure there's stuff that you've been really good at uh related to software that now you can hand off to the agent do you not get a little bit of sting you know we talked about it's stinging for your features to be implemented open source.

57:23 · But it can also sting that I've been really good at like I don't know refactoring or or or right now it might be architecture but maybe the model will be really good at that and now I'm like uh okay damn like I'm glad but also like uh it would have been nice for me to do that.

57:36 · Yeah, I think there's like a craft aspect to it. Um which occasionally I still you know pull up an editor and like write some code. Um, and it's just like it it it it feels nice and and and it's sort of like I have fond memories of like late nights sitting [laughter] in uh in Vim and you know, just like cranking it out, you know, drinking um Coke Zero and uh yeah, just not having to think about anything else other than like the problem in front of me.

58:05 · But really, I think it's um it's all about being in the flow and and solving problems. And what I find is like you know folks here and also like everyone I talk to is just like adapting very quickly and I think if you if you have a mindset where it's all about code is a tool to solve problems and you can solve so many more problems. It's like before you wanted to benchmark something and you weren't quite sure where you were going to net out.

58:32 · It's like you can just do it. it's it's going to take you like no more than 30 seconds, you know, to launch something in the background and, you know, get proper numbers and be able to do like a better trade-off. It makes you it should make you a better engineer if you just really care about, you know, the outcome and the system working well.

58:51 · And so what it allows us to do at OpenAI, it allows us to run, you know, our inference much more efficiently. It allows us to, you know, get like much more like effective compute and, you know, deploy that to the world. And so like everyone's just like very focused on that and solving important problems at the speed that was not possible before. And like I I haven't yet, you know, encountered someone who's like, "Oh, that's not that's not good. Um that's not fun."

59:15 · Do I understand correctly that it sounds like if you have ambitious problems, if you have way more problems than what you can solve today or tomorrow or the next week, sounds like this is not really a problem because when you know you get more efficient somewhere, you keep going. Which which is a lot of startups, right? like startups are always way more ambitious than than what they're able to do.

59:34 · I don't we're not we're not out of problems for sure, right? So, and and I don't think we will be for a while. Um we have a long long road ahead of us in terms of like mathematical breakthroughs, scientific breakthroughs, you know, making the world a better place like just really building for humans and solving the most important problems that everyone is is facing and just doing it in a a deeply human way. that's that's what we're here for. Also, just going back to coding and like you know these late nights, it's like I think there's like it's also like maybe like a glamorous version of it.

1:00:07 · Just like I also had very a lot of late nights where I was trying to refactor something. Um and you know just like I would be like three hours deep into the refactor and then realize like actually this is a dead end. Uh and I must restart from scratch and it was like very frustrating. Um, and so it's like there was like there are like these very very fun times, but there's also the time where it's like it doesn't compile and you're just like, why is this not compiling yet?

1:00:31 · Like I'm sure you had the time where you you go you go later, it's now super late, you need to go to bed cuz you need to you need to get some sleep and then you can't really sleep and you have this thing where like you you have some some task that is halfway and it upsets you sometimes. I remember dreaming about the code as well.

1:00:50 · And I guess one thing I don't really have these days when I'm working on my uh software for my business is I don't really have something that is halfway cuz I can just tell it do this and then I can leave it at a state where it's kind of like you know done either finished it's either working or it's I have proof that it's failed. But it's interesting because you know everything's sped up right.

1:01:10 · Yeah.

1:01:10 · I may maybe like I I do have like what what a lot of people do and I do myself is like you know I have like sometimes like bigger questions that I'm asking myself like and I you know from conversations I've had during the day or like I haven't yet you know just like had the time to just look into it and so I you know I will send off codecs to just like look at it overnight and then I'm very excited to then wake up and look at the results and so you know it's always like an exciting morning. Well, I feel there's an art to doing longunning tasks. And of course, you can use the slashgoal which will go and and and run.

1:01:44 · You know, that's also something that was recently added like a few months ago, right? The /go goal command to codeex.

1:01:48 · Yeah.

1:01:48 · And back to, you know, maybe like the harness is a crutch, right? Is uh uh slash goal was like necessary to allow like no to keep the the model like on track on like a singular goal for a very long uh period of time. And it's like it allows the model to literally run for days or or weeks if it's like a really really hard problem. But with the new generation of models like what we're seeing is like you know you don't [clears throat] need SL goal anymore.

1:02:15 · You don't need a harness around it. You can just tell the model like you know hey go and work for a week and you know it will actually do it.

1:02:21 · Speaking of hard problems and the fact that you're not out of them. One of the interesting things that you shipped from the outside it I would say it it was you know as an engineer it was moderately interesting uh is the what you call the merge which is codeex appeared inside of chat GPC and the reason I say that as engineers it's kind of moderately interesting because we've been using codeex like yeah it's there you can now open it in the chat GPT app great like I just went there and I just immediately went to Codex cuz I don't I don't really use Chad GBT in the app per se but I

The Merge: ChatGPT + Codex

1:02:53 · talk with uh folks at Open AAI and people in your team and you know they were telling me like there was a lot of preparation going on a lot of engineering challenges. Can you give a sense of how big this project was, what you needed to do and why was it difficult to pull off and and how you know how did codeex and other tools help you get it done in in ways that would have been hard before cuz since you've launched the merge the the numbers that you keep sharing of how many people use codeex it's like it's going up way faster than before.

1:03:22 · So I assume I assume there's a big scale uh problem you've solved here.

1:03:26 · A lot of things were the challenging with the merge is first of all completely different stacks. Chbt is like fully uh managed cloud-based like you know you run everything on our uh on our systems. We we store things like traditional traditional way of like building things.

1:03:49 · Um built for scale, built for for efficiency. Codex fully local. And so the merge is just really like how do you get the same uh the same benefits and the same capabilities from this local coding agent and then build a product around it and build it in a way where it can benefit like a much much broader pool uh of people which is which is also

1:04:12 · why you know all of us joined OpenAI is like to benefit like this very very broad population across the world and so it was like a very exciting journey of like figuring out like how do we build a cloud version of this that in essence is

1:04:28 · capable of like very very much the same things um but is also built in a way where you know we can serve it to like tens and hundred millions of hundreds of millions of users in a way that is still like you know efficient so that we can include it all the way into the plus plan work is essentially like running uh the full codeex harness in uh a cloud

1:04:51 · like together with like a a cloud computer it's It's a very powerful machine actually like people have sort of picked up on it and showed like you know what you can do like you know if if you are creative with the prompt is that you know you can get um you can get chatbt to like you know train another model in there you know wow [laughter] there are some pretty wild things uh you know you can get it to install blender and you know do like 3D modeling it's

1:05:16 · like very permissive it has like internet access it's like a powerful machine and then codex just works on it and this is like what we ship through cas work. a lot of system uh challenges. We the team did it very quickly.

1:05:29 · Obviously like Codex helped you know to make it more efficient look at and and build a lot of the infrastructure and then you know help resolve a lot of the little differences as well that you know had been occurring between between Codex and CHBT like merging plugins

1:05:46 · architecture um you know merging library and like so like really really working towards like unified system which is really the goal is like you shouldn't feel like you know you can do something in Codex that you can't do in TAB or vice versa like what we're trying to build is like one unified product that gives you access uh to the same intelligence but in the way that you want to use it. And so it was very fun as well because Codex throughout the whole journey also acted as a journalist to sort of like document all the steps and the debates and the discussions that the teams were having.

1:06:15 · And it was very animated debate you know of how we should do it and how we should name the thing and you know when to introduce it in what way and like what to merge into what. there were like many different permutations considered and so there's like a very fun journalistic element to it where we like we have a full recounting that Codex did over time uh

1:06:36 · and uh yeah it's just it's kind of become known as well as like the toggle arc of uh OpenAI where you know we introduced like the work toggle um which there was also a lot of debate around of like you know whether this was like the right thing and then you know just like we kind of grew to to just really like it um but over time we're going to merge things further.

1:06:55 · So it's like we're really headed into this direction of like full unification and you know we kind of view this as like um a temporary state uh where you know you have like you have better stronger capabilities when you're in work mode but over time we're bringing this you know all all the way to like you know everyone that uses CHP. And how do you personally use codeex? Like what's your what's what's your working setup in terms of agents, in terms of task, in terms of what you what what you manage uh with it.

How Tibo uses Codex and ChatGPT

1:07:24 · And related to this, I asked Peter Stainberger what I should ask about you and he said like you I I you need to ask him how do you deal with the fact that you're involved in all these projects?

1:07:38 · Uh your calendar is like Tetris, but usually you show up pretty cheerful.

1:07:43 · My calendar is fine. Um um and it's just I am capable of doing so many more things nowadays because I have the technology like CEX and I actually shifted a lot of like my my my work on uh on mobile using tab work where whenever I have something that I want to take note of I just like fire that off.

1:08:09 · I use dictation a lot. Um, whenever I have a question, instead of like writing it down to look into later or delegating to someone, I just like fire it off in charge of your work and I get like a report. It has like a whole bunch of like custom skills and um custom uh instructions where it's now like very very tailored to like you know produce the kinds of reports and slide decks and uh code explorations you know in the style that I can consume effectively.

1:08:36 · And so every time I'm like between meetings or like you know you'll kind of like see me like you know I was just like dictating to my phone. As I said before it's just like we do a lot of work in public channels. We have like a lot in Slack. We have a lot in in in notion and Google Docs as well. And so there's pretty much like there's no question really that I feel I cannot ask that you know Codex will be able to sort of like do at least a first pass of thinking through whether it is like public sentiment on a feature um looking at production logs for you know how much

1:09:06 · usage we have on a certain thing making a list of things that we should deprecate because they're not getting traction uh understanding what a certain team is up to. It's like any question I have I can get an answer to like you know within 30 minutes. And so that's how I use it. I use it for everything.

1:09:20 · It's like my personal uh agent in in like all the ways. And then oftentimes on on weekends as well, I do some like code explorations or like I build some prototypes and I have fun like sort of like imagining the future of the product in some ways. And I do that with others on on on on the teams. It's not always the same team. And it's just like in one day I can build things that I sort of like I had it in my system, right?

1:09:44 · It's like it's like I woke up one day I was just like we should explore what it means to build this and then I can just sort of express all of that and get like something in front of people in a day so that they can think through it and criticize it and hopefully get inspired by it. It's like by no means you know we need to ship it but it's more like okay I flush it out of my system and then you know I go on and like you know do other things. So it's just like so I know it's such a magical time and it's like so empowering.

1:10:12 · And as closing, what would your advice be for a software engineer/ AAI engineer, someone who builds software who would want to get the skill set and the experience to have the opportunity to work at a place like the Codeex team, like OpenAI or like an AI startup. So like you know just become this really great builder with with these tools because the question that comes up is often like should I start with the theory? How important are the basics? Should I just get really good at using the tools?

1:10:41 · Yeah.

1:10:41 · I think there are two things that are important is a deep deep curiosity for how things work um and an ability to like you know train yourself to understand things very quickly and so it's it's it is the case that things will continue to change but people that do extremely well at OpenAI are like you know people that just sort of like are able to like gro uh a system quickly and like you know also dive into like a new code base and sort of like know make sense of it. Um but obviously like all of that is helped with agents um nowadays, right?

Advice for engineers who want to work in AI

1:11:13 · So just like there's so much information that you need to absorb and like you know being able to understand and reason through it and a lot of that is asking good questions really about you know how do things work and just like going into like the five W's which I think you know you can just kind of keep digging and digging and digging and you know you're learning very very fast through that. The other thing is being in tune with the community or you know the people that you're trying to solve a problem for.

1:11:37 · It's like not everything is like solving a direct problem. Sometimes you're solving a problem that will be useful, you know, to like another group of people in the pursuit of like solving a problem for humans. But just being crisp about the taste or the needs or the requirements uh and being able to think clearly and like you know exercising through this clarity of thought feels really important to me. Like if you can't explain what you're trying to achieve, if you can't explain your intent, if you don't have a tie to a community, if you don't have the taste, it's it's um it's going to be much harder to do great work.

1:12:07 · Awesome, TB. Well, thanks a bunch for this conversation. This was awesome.

1:12:10 · Thanks for having me.

1:12:12 · I've always wanted to get together with Tibo, and I'm glad that we finally made it happen. I appreciated how Tibo talked about not just the upsides of open source, but also the downsides. most notably how competitors can copy features you [music] are just working on in the open right now and then ship it right before release and just how [music] much this stings. Plus, you get a lot of low-quality contributions that you still need to somehow deal with.

1:12:34 · Another interesting one was Tibo was saying how the harness [music] is always a step ahead of the model. From the inside, the Codex team see their job as building clutches for the model with the harness, [music] the tools, and the setup instruction. And then the next version of the model will be trained to need fewer of these clutches. I'll be honest, as a dev, this sounds a little demotivating that the stuff I build in the next version of the model, it'll [music] just know and we can get rid of it. Plus, I do suspect that it's not just about building these clutches, but also building tools that models will use.

1:13:06 · And it's not like the next version of the model will reinvent an MCB protocol or scales or plugins. At least I hope not. I also enjoyed hearing what the merge merging chat GPC and codeex look like from the inside. It was merging a previously fully local coding agent, Codeex, [music] into a managed cloud-based stack and doing it efficient enough so that it can be included in OpenAX $20 per month plan when $20 is not all that much in terms of compute purchase.

1:13:31 · It was pretty amusing to hear how CODC itself acted as a journalist of the whole project [music] as it was present in all the Slack conversations and all the documents and so it could capture all the important debates and decisions. I'm not going to lie, [music] this part felt a little bit of a big brother feel to it where the AI is always watching, but it could well become the new normal in startups in the future. [music] I've not yet decided how I feel about this. And finally, I appreciated Tibo's advice for engineers to succeed.

1:13:57 · Be curious, [music] understand symptoms quickly, and be in tune with the group you are building for. It's reassuring to hear from Tibo as well how much the fundamentals [music] still matter. Do check out the show notes below for deep dives on how codecs, clock code, and cursor were built and other related topics. If you like what you heard, please hit a rating on a podcast player that [music] you're using. It means a lot to me and to the show. Thanks, and I'll see you in the next