Transcript

Intro

0:00 · Can you measure an engineer's individual productivity?

0:03 · There's been this whole push towards individual output, but teams are still what matter, team output.

0:09 · You mentioned how you're seeing two camps. There's the AI pills folks who get it and then the people who seem to like they just hate AI.

0:17 · See, the problem is that neither side is making it [music] up. They are seeing really scary trends. They're grappling with real hard problems that are getting worse. What would it take for you to be comfortable shipping code without you reading it and understanding it? Cuz that's engineering.

0:33 · How do you see the role of good skilled engineering managers and engineering directors change? [music] It's just easier now than it's ever been to pick it back up, fill in the blanks. And it's always been the case that [music] why are there firmly two camps within software engineering when it comes to AI? [music] those hating its effects and those who are AI pill. Charity majors emphasizes with both camps and thinks they are talking alongside one another.

1:01 · Charity is a co-founder and CEO of Honeycom, previously worked at Facebook and [music] Parse and is one of my favorite voices in engineering. Today we discuss what it [music] would take for us engineers to ship code we have never read and why this is more of a when question not an if question. Why are reliability quietly getting worse across the [music] industry and why it will take some time to recover. Career advice in this age of AI? Why middle manager should consider going back to being IC and why junior engineers will be okay.

1:27 · [music] If you want to hear from someone who was skeptical about AI in 2025 but has changed her mind based on the evidence, this episode is for you. In today's episode, Charity will say, spoiler alert, that the question is not if we will stop reading code written by AI, but when. And we should take lessons from ops and QA on how they prove that software that others wrote works in prod. And she's got a very good point.

1:48 · As any ops engineer or SR will tell you, that's how software has always been written by unreliable agents from their point of view. That is software engineers like me, your colleagues, or you. And let's face it, you probably haven't read all the code in your codebase either. This is where I need to mention our presenting sponsor, Antithesis. Antiysus verifies software written by unreliable agents. It runs your whole system in a hostile simulation and roots out the bugs for you. It does this by using an approach called deterministic simulation testing or DST. Antithesis is turbocharger testing by running your whole system under aggressive fault injection.

2:21 · Imagine Anticys is hundreds or thousands of versions of the Mario game running each instance aggressively trying to break the game with increasingly weird input combinations. If it finds a breakage, this is where the determinism comes in. Instead of you having to try to reproduce a tricky bug you saw in production, antithesis can provide you with a perfect deterministic replay of anything it finds every time. With antithesis, you can specify properties at the whole system level and antithesis will actively try to disprove them. So you can be confident that if your system holds up in antithesis, it will hold up in production.

2:51 · Head over to antithesis.com/pragmatic to learn more. Charity, it's so nice to do this in person.

How Parse led to Honeycomb

2:58 · You're in my city. This is amazing.

3:01 · So today I want to kick off with AI. But before we kick off with AI, I just want to make it kind of clear for people who don't know you that, you know, you're not an AI hater or an AI lover. You actually built a lot of cool stuff pre- AI, right? Starting at we just saw Lynon Labs. Was that your first job?

3:16 · My first job at Lynon Lab right across the street.

3:19 · Right across the street. We were just talking about that. So you were building Second Life.

3:22 · Yeah, we were building Second Life.

3:24 · Yeah.

3:24 · And then from there on one of the the the big hits was Parse the developer tool which was beloved by developers backend for all for best backend for mobile services. I I used to use it and then what happened? Facebook bought you.

3:37 · Facebook bought it. Yeah. It was my first great lesson in most acquisitions fail. Most were terrible. This one failed. They shut it down. Um, but ultimately I'm very grateful to have had the experience because if it wasn't for that, I've always been a startup kid. And so nobody knew my name. And it wasn't until I was leaving Facebook that investors were like, "Oh, would you like some money?" And that's that's how we started Honeycomb.

4:01 · And then you saw stuff at Facebook, right? It inspired you that.

4:04 · Yeah.

4:04 · Yeah. Facebook. There was a tool called Scuba. And so we were in a weird position. We were building on AWS, Ruby on Rails, all this stuff. And then we got to use the internal Facebook tools. And Facebook had this tool called Scuba. And it was we were experiencing hockey stick growth. It was just like we had over a million mobile apps hosted on Parse by the time I left.

4:26 · Yeah.

4:27 · And every single week a new one would break. It would hit the top 10 on iTunes or something and out of nowhere it would just be like ah. And these apps needle in the hay stack, you know, and it went from would take hours or weeks. We have to get lucky. We finally find cuz it's not it might be one app that's spamming the logs, but that might not be the reason. They might all be backed up behind the reason, you know. Uh we started getting our data sets into Scuba and finding them. It just went from being a really hard engineering problem with a lot of luck to just being like report problem. Click click click.

4:58 · Oh, there it is. And it was just mindb blown like you just that was a huge problem for our entire existence and then it was solved with scuba. And then when you started honeycom so was this a bit inspiration that you wanted to build something that feels like scuba did?

5:15 · I I just the idea of I was planning to go be an engineering manager an engineer at Slack or Stripe or something and I was just like I would be so much less powerful as an engineer without this. And so, you know, the grand plan in the beginning, I'm just like, well, all startups fail. So, you know, we'll fail, but I'll go sit in a corner and write go code for a year or two and then then I'll open source it and I can take it with me wherever I go.

5:43 · Then that that's how honeycom started.

5:45 · [laughter] That's how honeycomb started.

5:46 · And we'll we'll we'll get back to like observability or honeycomb. But be before we do now with with AI, you know, it's changing everything.

5:53 · But I kind of had a bit of a blast from the past, which is one of the first places we connected was in 2020. So almost 5 years ago or so uh someone submitted a question to both my blog and your blog and it was the question was like can you measure individual developer productivity? Now I wrote an answer and you wrote an answer. And I wanted to ask you that was five years ago. No AI, no nothing. Today someone shoots you a question saying, "Hey Charity, can you measure one of an engineer's individual productivity?

The limits of individual productivity metrics

6:22 · You know, they're using AI tools and all the all this stuff." What would you tell them?

6:28 · I would tell God, I don't even remember what I said. I remember that blog post, but we we both agreed by the way that it was it was that you can measure some dimensions and they're not going to give you the full thing and they will for example not tell you how a team is doing if someone is is actually really key part of the team and that as long as you measure individual things. We both agreed that you need to be in the details to know and and as a good manager or a good team lead you will know.

6:56 · You will know but you have to have data to back it up. It's like color in a painting on the wall and and is it goodart's law that yes it's goodart's law. So like never go well it's this thing that matters right you need to actually understand but you need it to not just be your opinion that was tossed off because you you have an opinion about some per you know we're we're all we have biases we we are selective you

7:23 · know you need to look at the picture I also believe that you know there's been this whole push towards individual output but teams are still what matter team output and honestly if there's one thing that I am encouraged and excited about with the AI movement. I think it's forcing us all to ask ourselves early and often.

7:45 · What does good look like? What does good mean? What does productivity mean? What would better look like? What would great look like? You know, and these questions are hard. I think it's telling that we all jumped so fast to speed.

8:01 · Yeah.

8:01 · Oh, fast. We can do it fast.

8:04 · same thing faster, you know, just boom. And I've come to feel like that is a very immature description of what better is. Yeah. Just just today I saw the entropic team posted a podcast with Spotify's head of engineering or or VP of engineering. I'm not sure which one in which they talk that wow Spotify with cloud code, they're shipping 4,500 changes per day per week.

8:28 · I'm not sure which one, but they talked about speed and I was kind of thinking like my experience has been different cuz I I struggled to publish any like some of my my episodes did not go on Spotify cuz it was down.

8:39 · Yeah.

8:40 · And yeah, they were talking about speed, but we're not talking about quality. We're not talking about more functionality, better functionality, or just things that people want.

8:50 · And in a comment, some people were asking like, "Okay, so what exactly does that mean that they're shipping more frequently?"

8:56 · Yeah. Do customers really want the buttons on their app to move around all the time? I don't think they do.

9:02 · Yeah, it's an interesting one.

9:04 · It's the easiest thing to measure.

9:05 · Let's jump back to last year in 2025.

How Charity’s perspective on AI has evolved

9:08 · You wrote a blog post right at the end of the year looking back saying that 2025 for AI was what 2010 was for the cloud. Can can we talk about before we go into like what this year but like last year like how was your perspective? Of course, you're working at observability company. AI will will give you lots of like business as well, but you said it went mainstream, right? Last year.

9:32 · Yeah.

9:32 · In March of 2025, Fred Habar and I gave a keynote at SRCON. We gave the closing talk and it's Fred and I standing in front of a the term vibe coding had just been invented.

9:43 · Oh, yes.

9:44 · And we were like, you guys should try vibe coding. Pause. Groans. Audible groans. Just like people laughing like, haha. Our big pitch was that people should learn AI because you can complain better if you learn it, which is legit.

9:58 · I mean, I really mean it. Um, but at the time, I think I still saw it as a really big feature or like a bigger than a programming language like the cloud, but not like generational, you know, not changing everything.

10:15 · And and I think that was accurate. Um, for me it was November of 2025 when they released Opus 4.5. But actually wrote about this recently a couple blog posts back about how in retrospect you could see it coming sooner. You could see and it wasn't actually the models, it was the harnesses, it was all the tooling.

10:37 · And it was people starting to say that around July they were like this is coming faster than you think and this is what it's going to look like. And as those people who were saying it, they were once who were playing with either cloth code or maybe pi or open code. So the harnesses, you're right, they were getting better at the tooling that you know it went from just being kind of a shell script that would try again to like they built a lot of stuff around it and then you know the opus thing kind of it was a weird time at the beginning.

11:02 · It's been a weird time every every time for a long time but the early months of this year it felt like everyone around me was just trying it again and changing their mind. everyone.

11:13 · Yeah. But I think we were just talking about right before we started recording that both you and me respect people who do change their mind.

11:21 · And I don't think we were wrong to be skeptical that the first time. It's a pretty extraordinary claim that AI is going to write code about as well as the median software engineer can in you know for limited bounds of that.

11:34 · Well, especially because if we look back at the history of software engineering, this claim has happened again and again.

11:39 · Yes. you know, neural nets should have been doing something magical.

11:43 · I there's a sticker in your pack that says we already have a programming language that lets you cobalt is the punch line. So, I don't think we were wrong to be skeptical and and also don't forget no code and low code.

11:54 · Oh, yeah.

11:55 · I mean, we know it turned out to be a joke, but the the promise was the same. And we were skeptical and we were right. And now we're skeptical again.

12:03 · We were right and we were wrong. What I was saying in that piece though was I think we were right to be skeptical that time. But now I see the same thing playing out with would you be willing to ship it some code that you didn't read. There's no point in arguing about if it will happen or when it will happen. Talk about what it would take.

12:23 · Mhm.

12:24 · What would it take for you to be comfortable shipping code without you reading it and understanding it? Cuz that is that's engineering.

12:31 · Yeah.

12:31 · And it it goes back to like, you know, my my gut reflex would have been saying, "Oh, no. I would not do that because I've been used to that."

12:38 · However, you're right. You know, it would if if I could have a way to, for example, I could see the change, I could I could tell that this was tested in like a harness or or something. Same way where, for example, preai, if there was a I had a team member who said, "I vouched for this and I've I've hammered it and I trust that person."

12:58 · So like you're right there there's these things which are of course would never but I have thought that AI or something can do anything like that but if it could you're enough right or for example if you and the AI would would both do it in tandem for a few months and you would be like you would get to how much are they catching how much am I catching is it about the same is it more is it less and you're training it and it's getting better whether it takes 5 days or 5 years or whatever I think it's pretty clear that directionally that's where we're going and the other thing I would say is this is good for us.

13:29 · If you spent much time with the Phoenix architecture stuff that Chad Feller has been writing, you you you have been quoting. Yeah.

13:37 · Ah, I've been quoting liberally. I I should probably let you get to it in your your own order, but I just feel like anyone who's ever done a painful rewrite should be on board with this.

13:48 · Yeah.

13:48 · And but here here's a quote from from from Chaff Fowler. Immutable infrastructure, stateless services, containers, blue deployments, infra infrastructure as a code. These ideas all share common premise. Never fix a running thing. Replace it. AI pushes this premise beyond infrastructure and into application code itself. When rewriting is cheap, editing in place becomes risky. Mutation accumulates entropy. Replacements resets it.

Rewriting code vs. editing code

14:13 · Yes, code is cash. This is a very interesting idea because you've compared cha chaff compared and and you you've al of course uh shared this that when we look at how infrastructure changed before you know like specifically a server you need to be configured it I think we call it like pets servers having pets versus and at some point we stopped configuring individually we stopped like fixing individual machines we just like throw it away and have a new thing And with

14:44 · code, we've always been used to the history of the profession, you know, 60 plus years or maybe a bit longer is that we edit code. And are are are you thinking this might because of the economics of it? I mean, if you if you think about it, you could generate 10,000 variants of a function faster than you could write it once. And so when you start thinking about it that way, it's like, well, okay, we're going to need a lot of evals.

15:10 · We're going to need a lot of tests, but the generation is so cheap that it it really I think it it it forces us in that direction. And I think that the the expensiveness of writing code and maintaining code and the expense of software has always been bound up in its maintenance and those lines of code.

15:34 · The reason that we we trust something is because we've been using it. because we we know we we like there's this deep thing about production. It's like, well, it's trusted. We know. And I've been I know that as well as anyone. And I will also say this. Anyone who's ever done a hard database migration should have some real humility about our ability to extrapolate those contracts, store them.

16:01 · Like I am not one of those people who's like this is we're going to generate all code. I don't know how much code. I believe that we can go some distance in that direction and it will be good for us. I don't know how far we can go. I believe we can go farther than we are now. I just man the last project I did at pars so we had spent like six months writing the original Ruby on Rails API.

16:24 · Yep.

16:24 · Spent two years rewriting it in Golang.

16:27 · Wow.

16:28 · Yeah.

16:28 · It was it was it was and what was was it two years because new stuff being kept being added that you need to some extent and also Golang was a pretty immature language at the time. And we had to write, you know, the MongoDB drivers and like the all the other bunch of things and also just like when you're writing in Ruby and MongoDB and JavaScript and everything is, you know, there's no type safety and and it's just painful just, you know, and this strangler figs that they do where you build the architecture outside the architecture and you literally find the

16:59 · contracts with your users by breaking them one after the other. Like that just does not seem like the ideal artifact. We should be able to store them somewhere. We should be able to have architecture diagrams that we can review and discuss that generate that code to spec.

17:16 · This is very interesting because some of these ideas they've been around decades ago speci specifically you know if we had Grady BH as a third person sitting here the idea of like hey we can have architecture diagrams that translate to code UML started there. I think Grady

17:33 · would disagree that like he never wanted it to to go there but irrational software back in the 90s they said hey you'll define UML it generates code it will be beautiful now it wasn't beautiful because I guess some complexity and turns out the generating code was still expensive and and reviewing it but I wonder if some of these ideas now might be just feasible that's my hope that's my hope I mean I I I'm just barely old enough that my first job I was like 17 at university. I was assisted.

18:02 · I remember when, you know, I didn't I wasn't really aware of what was going on. I was just a kid. But yeah, I remember how stressful it was and how people were agonizing about how we'll never be able to get that information back. And everyone adapted just fine.

18:19 · They I I think I read the systems that, you know, they built the systems that replaced them, but not as in replace them and worked them out of a job. built the systems and they spent their time writing code instead of like running updates by hand on every server in the closet.

18:35 · And I guess this is an interesting one because clearly like the CIS admin role and profession has been it doesn't exist today. It it's kind of legislated has been eliminated. However, the people who were CIS admins, they did understand the operating systems, they understood hardware. Yes, they were in a really good position to adopt and a lot of them just became either software engineers, product managers. I know someone who became a tech sales person.

18:59 · Yeah. Yeah.

19:00 · So, it's almost like uh like and I will hold that our generation of engineers still the best debuggers. I'm glad that people don't all have to learn about CPU and memory and all this stuff, but like there's value in knowing that stuff. It comes in handy. I think there's some analogies there to the generation of code stuff.

Production as a stage of development

19:20 · Also, you know, you you took a bunch of inspiration in your recent writing about both CIS admins, but also QA. And you wrote something interesting. You said lines of code are not the ideal artifact to review. And I'll quote a little bit from you. The tools to do this don't exist yet, but many of the ideas do exist. Most come from operations and QA, two domains that software engineering has historically been been rather snobbish about.

19:41 · Should we revisit our relationship to to QA and and ops where I I I I I I feel we always put ourselves as software engineers here and ops and QA somewhere and maybe maybe time to eat some humble pie.

19:57 · Ops equals toil, right? Yeah. I think it's time. I mean, ops and QA have always been more concerned with what is. Software engineering has always been much more concerned with what how should it be.

20:11 · Mhm. So ops and QA have always been more concerned about validating, about correctness, about does it work as a well does it work to start with?

20:22 · Does it work? Yeah. Yeah. I mean it it's always weird to me just how much software engineers really seem to believe that the world exists in the repo and it doesn't.

20:33 · It's production. You know, the code has part of the information. Some of it it's very necessary. We need that. But like I know some software engineers who and okay some some some places don't even let software engineers look at production just like how I know a lot of people are very upset about AI and but there the things that get me up very excited genuinely excited about AI are that it is pushing the discipline in directions we have desperately needed to go for a very long time. Production is not what happens after development.

21:03 · It is a stage of development. And you've been saying this consistently for preI. I'm just going to say it for those who don't because I remember we've I think we also bonded a little bit over. There was this thing called trending on Twitter when it was still Twitter and it was tech Twitter. Everyone was there who mattered and there was a trend going it's Friday don't deploy some something.

21:27 · There was maybe a hashtag even like like I'm not sure don't deploy Friday or something like that. And the point was uh it was well-meaning. It said like look when you deploy often there's an outage and on the weekend we don't want to do so there was saying every every Friday it went viral saying don't deploy on Fridays and you came in and you said you know what you should be able to deploy anytime without fear because you

21:49 · should be able to just know you know however that might be CI/CD and then on top of this you were like no like you should actually just not even have a user acceptant testing environment at UAT you should just deploy to production like and test in production right as soon as you merge merge, it should go be going out like you should have to stop the train to make your code not go into production as soon as you've merged. Absolutely.

22:12 · And one more interesting thing is is you had a long train of thought about like like AI and and what it could be is one thing you said is our brains are not built for validation. Almost everyone I talked to including um Andres Hayesburg he said that look like it's very clear that code generation is cheap. we are generating more code and the bottleneck for human engineers is for code review

Code reviews

22:34 · and everyone's trying to figure out how do we make code review easier how do we build nicer tools Uber has built amazing tools to like try to like surface important code reviews but everyone's pushing like all right let's do more code review as an engineer I'll I'll be honest like I I never liked doing a code review when there's very little to do and it's with someone I care about I'll entertain it it's more of a coaching opportunity then right but But but especi as soon as there's an AI, it's kind of like I don't know. I I don't really care. Like I'm just being honest here.

23:04 · Like do you care when I don't I've never So the pro one of the problems is that I think code review means so many things to so many people in so many places. And so a lot there's a lot of projection going on. A lot of people are if you say that you don't want code review, you're saying you don't want to talk to your co-workers, you don't want to mentor juniors, you don't want to, you know, which is not true. we've just bundled so many things into this like, you know, it's like hugely overloaded. [snorts] Hugely overloaded. And some of those things are really good. Some of those things um could be done better in other ways.

23:36 · You know, some of those things are very cultural, very specific. My friend David Pole, who um I worked with at Parse and he's now working at GitHub on poll requests.

23:47 · Amazing.

23:48 · The mafia.

23:50 · Yeah, exactly. Uh he's like to me the code review is when we decide do we want this in our product or not. I'm like well that is a that's a great it's a great discussion that is what humans are good at. We should talk about is this mental model coherent? Should we add this? Should we not like love that architectures? You know but like the code is not necessarily a great artifact for all of those. So should we be talking to people? Uh yes.

24:20 · Is the code review the right form factor? Maybe. But I I think that the emotional reaction that's so when people are getting to the like the validation in my book is at the very bottom of the list.

24:32 · I I'd like to like touch like stay here a bit more. Can you break out the the parts because it feels me code review overloaded but the parts of code review or or the things that you have seen are good things and maybe we don't need to do as code review and the things that are just like just have never been that good and maybe we just need to throw it away.

24:49 · Yeah.

24:49 · I mean I think do we want this in our product is that is great. I mean ideally you'd talk about that before you write the code for it but you know whatever. And you know is this is this API design? You know, those are great conversations. Reading for syntax and bugs and that sort of thing. It's not evil, but it feels like it could be.

25:10 · It's a it's a teaching opportunity if that's the best teaching opportunity you have. And I guess some folks at some point maybe you need them, but it doesn't feel high. This like a great use of anyone's it feels the only time where it's useful is if someone joins a team and initially it can be a little bit of feedback, especially when there's like not nothing is written down. There's no guidance.

25:29 · There's no linting rules that would give you that.

25:31 · Well, see, that's again, yes, we can fill in the cracks if we haven't built the guard rails. We can fill in the cracks all kind of ways with with our own time. But there are so many things I think that we never think to extract out of the process of building and validating software. So, we rely on us.

25:52 · So, I am a huge fan of Intercom, now Finn, their engineering um or I have been forever. Like, I noticed their CTO a decade ago had this saying, shipping is your company's heartbeat. And I love that they ship a Ruby monolith like 10, 15 minutes, hundreds of times a day. That is not trivial.

26:15 · It's not trivial thing to do, right? So they're kind of a high water mark in my mind right now for teams that were founded preAI have a lot of engineering discipline who have become AI native and they wrote a great post about how they do PRs that are AI validated and the bar for them is very high. It's like they have all the wisdom of their most senior engineers looking at every single diff and that is fantastic.

26:42 · Which means that you don't have to worry about remembering and looking and nitpicking and all the things that we're not good at anyway and they can talk about is this the direction we want to go. Is this the right is this the right path?

26:55 · You've also written that nondeterministic systems require more entering discipline not less. So like what is the thing about the these non-deterministic system we're specifically AI right we're talking about AI let's just just name it that is we see that AI does amplify both discipline and lack of discipline why do we need more and when you say discipline what specifics are we talking about well I mean tests and eval right like if

Non-deterministic systems

27:22 · if we're treating the code like a trusted artifact and we're you know trying to predict everything with our human brains and everything, then we're writing the tests that we can predict that it might break, you know, and then anytime it anytime the system breaks, we like try and write a test for that, but that is not an especially high bar.

27:39 · And and so I think the sort of um the the behavioral tests or the I don't remember the where it starts with C, but the QA folks have these suite of tests where it captures there's also smoke tests.

27:54 · Yeah, there's so many. There there can be like performance tests. There can be low tests there. There can be Yeah, there can be like just kind of fuss testing as well.

28:03 · Something that's like if okay, if I'm not going to read this code, how do I know it's going to perform within boundaries of the last code that I generated? That is conformance testing.

28:14 · Conformance testing just as important for lots of workloads as, you know, absolute performance. Is it just not changing too much? And so we I think we're going to need the trust has to go somewhere, right? If you're debiting from this trust account in the creation of the code, it has to get built up somewhere else. And I feel like one of the things that I'm really excited about in the coming months is just I actually really like thinking about it less as AI and more as deterministic and nondeterministic systems that have to play nicely together because determinism is not going anywhere.

28:45 · It's incredibly valuable and we have to learn to make AI kind of boring. You know, it's a non-deterministic tool, which means that it is all over the place, but it's so valuable, but it's all over the place. So, we have to learn how to give it carved pathways and places where we kind of corral it, where we use it in the way that it it's a superpower and not in the way that like erodess our foundations.

29:10 · This is interesting as Martin Fowler a year ago when he was on the podcast the thing that he talked about is how the biggest change with AI is the non-determinism and when we think back in the history of software it's always been deterministic safe same for neural nets but that was most of us software engineers didn't really touch too much of it because it just wasn't that useful for us but we've been used to that when we programmed it you know it just happened the same way unit tests were easy because you just run them once you don't really run them twice cuz why would you and I wonder if

29:42 · This is we need to just realize how big of a deal this change is and that the any business that employs us like you know they they want software that works the same way. We we we just had a a recent post on hacker news. Uh there's this ATS application tracking system scoring system that hacker rank outsource which scores your resume. And so software engineer just like and and you can run it locally. It's open source. You can use a local model. I think they recommend Gemma Google's small model.

30:12 · And when you run it like 100 times, it will score the same resume anywhere from like 66 points to 99 points. And typically most companies have 85 set as the bar. And you're like, "Hang on. So, we've turned what is what they were advertising as a tool to help your recruitment. We just we just proved that it's just a coin flip." Like, that's bad.

30:33 · Yeah.

30:33 · And we have to be able to say that it's bad. AI is not the right tool for every use case, you know? And I think I think every company is going through this in microcosm. And something I was saying to folks just earlier today, we we've been doing this series of conversations on our AI norms and values. And it was like a year ago I don't trust us. Like a year ago if we were like yes we should use AI. No we we didn't we didn't know enough.

30:56 · We've gone on such a journey over the past year and we know so much more now that like if one of my co-workers is like AI is the wrong tool for this job. I'm like I trust you. You know you got to get worse before you can get better.

Sensible uses of AI

31:11 · So tell me about where you are right now with your how inside of Honeycom how you're thinking about AI. how you're thinking about how to think about AI and and what what what values you came up with that works right now for you.

31:24 · Yeah. It starts with just acknowledging that the bar has gone up for all of us. That's what happens um when we get powerful new tools.

31:31 · Has the bar gone up or has the the you know the floor gone up?

31:37 · That is a great question. Maybe yes, maybe. Yeah, I don't know. We're we're definitely in a sort of wandering in the wilderness phase. Um but you can't not wander or you will be left behind. You know we acknowledge that the bar is going up for all of this and that the only viable way to define that bar is better outcomes and asking ourselves like is this good?

31:58 · Is this better? What does good look like? Another another thing that we point out is just there is no human in the loop. You own the loop. The loop is yours. The loop is mine. It would not exist if this was not for me. So I am the owner, right? There's no, "Oh, Claude said this, so." No, no, no. It's your work. You own it.

32:17 · Charity just talked about owning the loop. Owning the loop also means controlling what every agent inside of that loop is allowed to do. Which brings us to our season sponsor, Work OS. Today, agents are increasingly able to act on their own. And the old off model was never designed for that. Who is this agent? What's it allowed to touch? On whose behalf? You really don't want to get answers to these questions wrong.

32:39 · Work OS is built exactly to solve this problem. Work OS is fine grade authorization FGA designed for how agents actually operate plus SSO and skim and not just user off with agents bolted on after the fastest growing AI companies. Entropic OpenAI cursor perplexity already trust works. Check it out at work o.com. I also want to talk about built kite the CI orchestration platform trusted by cursor openai entropic uber kama and more. Charity talked about owning the loop, but here's a challenge. Thanks to AI, your agents are writing a lot more code.

33:11 · To trust this code, every change that an agent makes still has to be built, tested, and proven safe before it ships. So, obviously, you need CI more than ever.

33:21 · But when agents are pushing 5, 10, or 50 times the commit volume to your pipelines, faster CI runners won't be enough to keep up with it. Shaving 30 seconds off a single build is meaningless when the queue is 100 plus jobs deep. What you really want is a CI system that gets faster as the volume grows and CI that offers instant parallelization to give you unlimited concurrency and to intelligently route changes at runtime. This is what Buildkite does and why global software leaders at every level continues to rely on it.

33:48 · The same architecture that observed the scale of Shopify and Uber a decade ago now runs about 1.4 billion job minutes a week across Cursor, Meta, Reddit, and Snowflake. While the rest of the CI world are crackling under the weight or rearchitecting their platform, build kite continues to reliably grow.

34:05 · Agents run on your infrastructure or on build kite. Any cloud, any chip, your secrets, your scale. Every artifact and log is captured. So when something fails, either you or your agents have immediate insight for why. As you're injuring the context you give to your agents, think about how you verify what they hand back. If your system is buckling under the increased volume, head to buildkai.com/pragmatic.

34:27 · 30-day all access trial, no credit card, and an actual human engineer on standby. His name is Ola, and he's very helpful. And with this, let's get back to charity and communication norms with AI. I think there was this frenzy of, "Oh my god, I could do this. Oh my god, it's so cool."

34:42 · And I I I know you have also become very weary of this slot. Very I I just don't even read it anymore. As soon as I can tell, as soon as I recognize this might have been AI, it's like trash.

34:55 · Here's a baseline. Uh, you cannot send anyone something you haven't read. And in fact, if it would take them longer to read it than it took you to make it, it's probably slot. That's really disrespectful actually. And I think like just like asking someone to like you're asking anytime I give you something, I'm asking for your time and attention. And if I'm giving you something that I don't even know what's in it and I and I'm putting it on you, it costs you instead of me, that is [clears throat] not good.

35:24 · I also think that even before that it's like I've noticed as I start working on these norms and and values I'm noticing myself as I start to ask someone a question without trying to look up the answer. Oo I shouldn't do that or if I'm giving someone something that I kind of generated and I'm like oo you know it's so part of it is just self-awareness. It's interesting because everything you talked about it reminds me of when a new joiner would join a team, a junior engineer, a new grad.

35:55 · Either they they had emotional intelligence or they picked up on really quickly that for example you go and ask a senior of their time once you put in a little bit of work and you start to respect their time as well. And obviously it doesn't start like that. We don't want them but there's this balance and I almost feel it's the same thing.

36:10 · We're like look like respect your colleagues, respect fellow humans. if you are communicating with them, make sure that you're not wasting their attention cuz now I guess attention is we're we're kind of running low. Like we have all of these all of these a like a bunch of people have a bunch of agents doing but the point is that's kind of the currency and as long as you respect that it it doesn't matter like I I think we're not talking about don't use AI for this or that like you use it as much as you want or what make yourself more efficient just don't degrade because it it really degrades those personal skills, right?

36:41 · You can use AI as a shortcut to help you not have to think too much. And you can use AI to help you think more deeply and more rigorously.

36:51 · And both of those use cases have their place. But when it comes to your core job function, we primarily want the second one, right? And especially if you're involving someone else and you're asking them to review or you know and this is not absolutist like there are people who English is a second language and they use it. people who like neurode divergent and and that is again that is still being respectful you know it's so it's not like like you said it's not no AI but it's like make reasonable asks of

37:19 · each other and and you know we don't need to reinvent a new bar for quality or respect because we have great bars already for quality and respect we just need to apply for a while there I think that there was a bit of oh my god this is so cool do you see what this cool thing can do and I think we're all just like so over well the reason I I really respect that you came from the the cis you know this the cis dev background you also you're

The two AI camps

37:46 · very invol these are all folks who have been pretty skeptical of AI and and you you mentioned how you're seeing two camps two very clear camps there's like kind of the the AI pill folks who like get it and then the people who seem to like they just hate AI and you're you said that you're not seeing these two camps have any sort of way to go between any feedback loop? Can can we talk about what you're seeing and like maybe you know like where you see some of these camps forming?

38:17 · See the problem is that neither side is making it up. Like they are seeing really scary trends. They're see they're grappling with real hard problems that are getting worse, you know. And on the enthusiast side, it's like they're acutely conscious that it's it's a bit of a race and that we need to push ourselves out of our comfort zone and they see other companies moving faster, catching up, leapfrogging.

38:47 · They're really worried about, you know, we're falling behind. And and the first thing I don't want to make it sound like false equivalents because well, there are elements of this that are true. Lots I think every company is more one or more the other, but like they're not wrong. They're not wrong. These we've never seen technological change this fast.

39:05 · We're on the inside of an exponential curve, which is very rare and it never usually lasts that long, but it's still happening. you know, um, things that are happening that shock us and we would be wise to prepare for them. So, like that's real. That's real.

39:21 · And, and these folks are usually at most companies, usually they are the small minority and they are constantly feeling outman. One of the things that's ironic though is that both of these sides feel like they are the tiny minority and they're outman and they're being suppressed and they are standing up for what is truth and valor in the face of the big AI folks or the big skeptics.

39:41 · But the other side, so and this often starts to come down to the group that is on call and the group that is not.

39:50 · Yep.

39:50 · Because the people who the buck stops with them, they are seeing melting mental models. They're seeing slop.

39:58 · They're seeing all their hard work just dissolve and they're see and they don't see any end in sight. And so ju just to be clear, we're seeing that the people who are on call for a lot of these systems, they're seeing more incidents. They're seeing carelessness being caused by it. They're actually seeing that since that group started to use more AI, our systems are getting way worse.

40:17 · Way worse. Yeah. And and it's that's very real and not making it up.

40:21 · No, no, no. Actually, I was just talking to someone inside of Meta. Uh there's been this big drama where people have been saw your post.

40:28 · Uh so not just my post since then and I haven't written about this since and I'm not sure when this podcast came out. I might have not talked about it is inside of Meta. uh they track SE zeros which is the highest severity. Well, you remember Sevzeros. There has been a flurry of Sevzer, so many of them. And you you cannot hide like this is, you know, meta. Like this is black or white. And the past about two months, it's been crazy. And and just so it happens, it's happening inside of Instagram.

40:57 · It's happening inside of WhatsApp where the the trust and safety, the basically the the reliability folks have been axed, removed. So it's impossible to deny the connection as well that there of course it's not a direct one but again and each each one has as a postmortem but meta has not had this badge for closer to a decade. Yeah, move fast and break things.

41:24 · And you put two plus two together. And when I told this story at a conference, people came up to me and they said, I'm so glad you talked about this cuz my company, different company, often VC funded or publicly traded like same is happening. People are like whispering to me like like we are not metal, but the same thing is happening.

41:42 · Same thing is happening.

41:42 · And and you know what they all told me?

41:44 · They told me I thought it's just us or I thought it's us and then my buddy who works at this other company. And suddenly it's like, "Oh, it's all of us."

41:52 · No, it's all of us. Yeah. No, it's the real thing. And the intercom folks, you know, what I love about them is they publish the real gnarly stuff, right?

42:04 · Yeah. They don't color it out.

42:05 · They don't color it out. And they showed that for 18 months, reliability and code quality went down and it had just started to possibly be going back up. But it's it's still not there where it was. And they're very honest about and they're honest about it.

42:22 · Finally. So on this is the thing like stop like spitting in my and telling me that you know like it's just this is my thing.

42:32 · It's like we need to hear the wins. We need to hear what's we need to hear about what's possible. We need to hear what's exciting. But you got to couple it with the costs. You got to couple it with a is it worth it? You got to couple it with what are we doing? What is happening? And I feel like part of the reason both of these sides are getting so frustrated is because they're not those they're not connecting at all.

42:56 · And so the people who are seeing really incredible there are some really incredible things happening in software right now like with rewrites and with you know automating away like real toil and like not a single person that I've talked to would give it up. Yeah.

43:13 · Which is amazing. like they don't they get so excited. Nobody wants to take it away, but half of the people are seeing the wins and they're not connecting it to the cost, which makes them think that that their co-workers are just fucknuts who are just like, "Oh, they just don't want to lose their jobs. They're just afraid of getting automated out of existence. They're just blah blah blah blah blah." Like, no, dude.

43:35 · You you be on call and then see how you feel. you know, and and there's a mirror effect kind of happening where the folks who are on call, who are responsible for this stuff, they don't actually believe that these winds are real. They think they're all cooked because they're not hearing the quiet part said out loud that, "Yeah, we're seeing this wind, but this is what it cost. We're still cleaning this up. We're still And so that that's my that's my beg to everyone who loves GG's podcast and listens to this is tell the whole story.

44:09 · talk about the costs. We're all do we're all in it together.

44:13 · Yeah, cuz you're right like this technology is not going anywhere. It it will make really big positive change at a bunch of it's here, but it's not magic.

44:21 · It's it's not magic. And I I think this is what you said in in make AI boring again. Another great article of yours. What you said is AI is just technology.

44:31 · It's just technology.

44:32 · And you were arguing that let's just realize it's technology. It's a tool and let's learn to use it. Well, now one other thing you said which is very interesting is software will be the killer app with AI.

Why AI works so well for building software

44:46 · Yeah, I think which is very unique. Let's talk a little bit about that.

44:49 · Software is made of logic and language.

44:51 · AI is made of logic and language. And because of that, we can bake in guard rails. We can bake in checks. we can bake in validation that we I don't know how we do that in other parts of our lives or other applications. And so it totally makes sense to me that software is what AI is best at. I mean you you see like in the courts they're starting to get lawsuits for for the court is suing lawyers who are submitting briefs that have hallucinated crap in them.

45:21 · How do you check for that? You know with the same I we have structured data. we have you know a whole and I just don't know how you account for that in the same way it it might also mean that whatever will work outside of the software industry for AI it will be a subset of what will work in the second basically if we can do something with AI if we can automate a process or something you might be able to do it in other industries but but maybe not but if we cannot do it good

45:52 · luck you will not be able to do it because we we have the domain where you can validate stuff we have we have incredible training training data on on code that compiles, right?

46:01 · Yes.

46:01 · Like in in a place bunch of places, you might have like training data like with with magazines, you might have like lowquality magazines or whatnot, right?

46:08 · You see what I mean?

46:09 · I mean, back to your point about humans like their determinism. They like things to happen the same way.

46:15 · And it's very interesting because as I think of it, you know, one of my businesses is writing. I write a newsletter that is is I I like to think it's good and it's worth reading.

46:24 · It is. And I would have said if you asked me what is AI really good at? Now obviously it's good at coding but before that it was good at writing. It was like my mind was blown that it can actually control the language. When all the newer models come out I do this test where I say like all right like you know write an article in the style of the pragmatic engineer and every single time I can tell it's a generic because because it's it's repetitive it has this thing. So my point is AI is actually not as good as writing pros as it's a lot better in writing code.

46:54 · Way better at writing. when I asked to write code like I often I'm like yeah this is something I could have written whereas when I asked it to write words I'm like I would have never written this and it has training data on me so who knows this might prove that software is the best fit I think it is software is a simplified version of of language for a purpose

47:13 · yeah I you know at first um everybody was like trying to come up with ways to be more efficient and write with AI and everything and I sunk a lot of cycles into that and I have decided not to sink anymore Because writing is thinking on paper and there's no shortcut for doing that thinking. Anything that I write, it's not content, you know? It's not content where it's just like, well, generate me couple thousand, which I'm not shaming anyone who generates content, but that's not what I'm trying to do.

47:42 · I'm trying to think through hard and interesting problems and share them with people. And I don't think AI is the appropriate tool to use for that. I use it for structure. I I'll be like, "Hey, read this and give me feedback and stuff." But don't worry.

47:56 · So, I think we should not forget that as we improve our, you know, our our skills, our capability, our experience, our our thoughts, we do become more valuable. And I I have this idea and this this might be a flawed idea, but I think it's I think it will be correct that, you know, 5 years from now, how will people be hired? Now, of course, we know the tools will be better and all that, but in the end, I think it'll be like this. someone's sitting here and I'm going to be interviewing at you. I'm going to be trying to get into your company, honey, probably Honeycom, right?

48:26 · And we will be having a conversation and you will judge me based on how I respond and the more I have spent thinking and bettering myself, the more valuable I will be to you because you will have all these candidates and some of them will have outsourced or other things to AI and they will have a blank because that thing is off. Guess who you will want to work with, right? I am so excited about leaning into the parts of being human together.

48:52 · I don't like the feeling of chatting all day back and forth between agents and people on Slack. Like it feels way too similar. It's just gross. Honeycomb is a fully distributed company, which was never. We always wanted to have a hybrid model, but the office has not come back. And and I and I feel all kinds of ways about this because I love not leaving the house.

49:15 · But at the same time, I I crave this more full like I'm so glad you're here. It's so nice to see you.

49:23 · We were just talking how it it is different. We we've done a podcast uh remote and it was a decent one, but but this is more enjoyable.

49:30 · Yes. And so part of what I hope we do is just remember that we're in charge of the machines. They serve us and this is still what matters.

49:40 · Yeah.

49:40 · I I want to pull back to back some to something different to talk a bit more about ops and and DevOps and give one of your spicy stakes. So now that we have AI, we can actually just, you know, badmouth some of the other thing or or just be real. Let's talk about DevOps.

DevOps

49:53 · Just can we go back a little bit in time? You were there. Why was it created? And in the end there was this massive DevOps movement in the 2010s. Do you think it succeeded? Do you think it failed? So before DevOps, we needed a DevOps because there was devs and ops and there was the proverbial wall that code got thrown over, right?

50:15 · And ops were the people who were in charge of the IT. They deployed, they managed the servers, they said the Linux version, handcrafted Linux, yeah, you know, pluggable storage models and everything. That was always a bad idea because it's split brain. Half of you are writing the code and the other half are understanding it. I would argue that you can't really understand the code you write unless you're operating it. So, you know, the DevOps movement did a lot of good trying to knit back together that sort of original original sin.

50:43 · And you know, around the time that I was a CIS admin, there was this big push. All right, ops people learn to code. And great, I'm glad that happened. Everyone who works with computers should be writing code. I feel like the wave after that was a little less successful which is like okay software engineers time to learn to understand your code in production.

51:05 · But I also think that in my mind 20 years of DevOps was really about one thing trying to create one feedback loop that connected people writing code to that code in production. And it failed. I mean it failed to this day.

51:25 · like it they're done they're done by they're two different domains you know there are some people who I mean it's and I I'll show you this this diagram that that that you drew we now added agents we'll we'll put it on the so viewers can see it but this is your I I think it's a really nice uh draw up of how there is no feedback loop like the office people or often times we call it platform teams they manage the inflayer

51:48 · engineers deploy there and so to be clear I think that's actually good and fine and healthy I think that there are separations of concerns where you can't expect anyone to do everything and the nice separation of concern is do I own am I responsible for the stability of the things that you put code on or am I responsible for the code that I put on the thing right that is a nice seam because you want the

52:12 · infrastructure to be stable be like to protect itself to be resilient and all these things and you want your code like to be oriented towards is every single user having a good experience. You could have one of those things be true and the other not be true. Like they're they are decoupable.

52:29 · And and actually this is like even the most modern companies I I often refer to entropic as this company which operates in a very different way to most companies. They're very successful despite doing a lot of different things.

52:40 · However, internally they have platform teams. They had the cloud platform teams and then they have applied AI who which is more of the kind of the feature teams, the integration and the two I talked to both of them. They just have a very different outlook. They have a very different view on even basic stuff like will software engineers be obsolete. The the people on the platform team were like no we're working really hard and on the apply they're like well maybe it will happen.

53:05 · Yeah. That does not surprise me one tiny iota.

53:08 · No but but so so this company Antropic that started with a blank page they arrived at the same place.

53:14 · Yeah.

53:14 · Yeah. No, I think it's the right separation of concern. And I'm not trying to erase it, but I think that to be a good engineer, you need fast feedback loops. And and this is part and parcel with the whole, oh, the source of truth is the code. If that's where you live, if you live in the land of how it should theoretically work, no. And I think that with agents, they're breaking that, right?

53:38 · They're breaking that and they're forcing another thing on the observability trip is a lot of people if you say like what is observability they'll be like ah well there's three pillars there's metrics logs and traces we talked about this last time metrics and logs I would say are system exhaust they're the exhaust pipe they're and and they're never going away because every team runs a ton of

54:00 · third party software they didn't write it they don't own it they just have to run it and it's outputting [ __ ] yeah and and you want and you just got to put it somewhere you observe that you see was and then you do stuff with it.

54:12 · Yeah.

54:12 · And you know, you should put it somewhere cheap. There's a ton of it. It's not super high value, but you definitely need it, right? And you can't do anything about you. Just take it and put it somewhere. Then there's your code. There's your crown jewels, the the code that makes you a company.

54:30 · And for for that code, your telemetry should be a product decision. It should be you store it once with all all the connective tissue because the value of rich data goes up not linearly not even exponentially combinatorally.

54:48 · If you have a wide event or a trace with 29 bits of data and you add a 30th that 30th is more valuable than all the others multi like it is just so powerful and with non-deterministic software you know right up front you can't predict what it's going to do. You have to c like that is a product decision to capture that trace. So, so let's talk specifically about modern observability and like companies that are, you know, like either building AI related code or just complicated code that they're generating.

Modern observability

55:19 · You know, in the old world again like I'm just being, you know, observably 101 back in the day. The way I would have written the code is you write the code and you think like, hm, something funny might be going on here. Let me do a log or an info or a warn.

55:33 · And then I would also try to maybe if we're printing this in production, I realize like okay, well I guess it's crashing and we don't have any logs there. So I guess it's some other part.

55:41 · Let me do put a tool that will like log everything and now I have a bunch of stuff. Now this is the old the the very simplest way of thinking in kind of a modern business where I'm like I know this is high value stuff. What are ways that I can go about that's actually maybe a bit like more practical than cuz I I I just will use super basic one.

56:00 · Auto instrumentation has gotten so good in recent years. If you're using open telemetry and everyone should be using open telemetry. All of the common patterns like all of the models are trained on them. So it is literally faster and easier to build with instrumentation than than not to.

56:18 · And with instrumentation to just once once I have the code in a compile step or an extra step, it just adds it to the right lines.

56:25 · This is what's important, right? It's part of just developer intent, right?

56:29 · this is how you declare your intent and that's how you check up on your intent in production. It's honestly gotten so much simpler. And you know, I don't I don't fault developers or anyone else for not kind of closing that loop with with DevOps because the fact is it was it was prohibitively hard and timeconuming and difficult because you know you're old school software engineer and you're you sit down write some code you're like ah here I should instrument it and look at it in production. So you're like, "Okay, I've got a bit of data and I want to do something with it."

57:02 · All right. Is it a metric, a log, a trace, an exception, an error, a profiling, you know, just like, okay, if it's a metric, is it a is it a counter? Is it a gauge? Is it, you know, just like all down takes? So, and then well, what type of data is it?

57:20 · Is it going to have high cardality? Is it going to be a needed to worry about you can blow it? Yeah. It's just like if it's a log line, do which log level do I do? Do I append it to like it's just you could double, triple, quadruple the amount of time that you spent writing the code trying to instrument it and then still wouldn't be done. Like you deploy it and then it's like, okay, I know the name of the thing that I added, but how do I find it? How do I display it? How do I create a dashboard? It's just like that was prohibitively that was really hard.

57:49 · But now we can bring all of this to you right in your development environment. It is easier and faster to instrument with telemetry than without it. And you don't have to leave your development environment to go and get it. You know, you could have the agent like we've built some really cool [ __ ] at Honeycomb where it'll just it'll be like, "Oh, hey, that thing that you wrote, you know, maybe you want to look at this."

58:14 · And you can you can control how it is. you can you know but it's it's right there and that's how it should be it should be part of your development loop can do we talk about what spans are because uh I'll quote Eric Eric Riddok who recently wrote LinkedIn the basic idea of observability for applications is don't use logs or metrics just put it all in spans what are spans spans are bits of of a of a trace

58:39 · I mean a trace is just structured log with some fancy fields right and so the span is a subset of the trace that makes up the entire duration. And I don't know if you've followed any of this, but like the default building block has been the transaction for as long as the web has been around.

58:58 · Yeah.

58:59 · That doesn't work anymore with specifically with AI.

59:03 · Yeah.

59:03 · We just we just ship something called timeline that is like that sits on top of spans. So, you know, if if you you know, if you if you run something like intercom, you have got a chat thing and and a customer's like, I'm conversation going on.

59:17 · Yeah.

59:17 · A customer's like, I'm complaining. You're like, okay. So, you spin up an agent, supervisor agent that spins up more agents, and each of them calls APIs. Each of them calls like storage backends and stuff. Then they return and then the customer has another that could span hours, right? And you need to be able to zoom out and visualize the whole thing. It's super cool. And so this is a new primitive that you came up for these use cases where there's a conversation or or like a an LM is involved and and you you have like it's like a metatrace.

59:50 · Okay. Yeah. So I guess this a trace of traces.

59:52 · So we need these new building blocks actually just you be able to work with.

59:57 · Yeah.

59:57 · Interesting. So I guess this is something to keep in mind like any any any engineer who's like building on top of of LMS who is an AI engineer now as as we know. It's either that or you've just got all these tabs open with traces and you're just copy pasting IDs from one to the next.

1:00:11 · Yeah. Or if if you're a large enough company, you might have built your own, but we know that's it's doable, but it's painful.

1:00:20 · It's it's doable. It's painful. I'm really looking forward to seeing over the next few months or year or whatever just the the marriage of tests and tele and evals from a telemetry perspective with agents and AI agents being around a lot of them are are now very useful to connect to observability stores you you can go and and do stuff. However, one question that comes up is, well, agents have a finite context window and with observability, you can really easily overload that.

Handling context overload

1:00:51 · What are approaches you've seen of of agents either using honeycoms or or or some other data sources to like make them productive? Have you seen some patterns?

1:01:03 · There's a lot of trash data out there. Um and and a lot of traditional telemetry data, metrics, logs, traces, what was all it tends to fill up your context window with crap when the most important part of the data is again the relationships between the data.

1:01:21 · So if you can and in fact one of the AI SRE startups posted this great piece a couple months ago about how they they see the agents that they deploy in the wild bypass the observability data most of the time and they go upstream to find richer intact telemetry data.

1:01:41 · So that's what I would say either you give your agents the but it's it's the relationships that matter right because that's what actually helps the AI make decisions and when it comes to observability I cannot not mention your book observably engineering and you have a second edition can you tell me why you felt the need to write it and what's new in it oh man uh the whole thing is new so

What’s new in Observability Engineering’s 2nd edition

1:02:09 · O'Reilly any anytime a book is considered successful and if the topic is still relevant they'll ask if you want to write a second edition So, it's not really um but I was really excited to write it. The first book, I don't want to say I wasn't proud of it. You like you're not like your children in your books, you're not supposed to like say anything bad about them, you know, cuz uh it's fine.

1:02:31 · They're but it was written 2019 to 2021. The definition of observability meant one thing when we started and another by the time we ended. And there was at no point where I was like, "Oh, this book is great. Let's ship it." It was just like, "Oh god, I can't do this anymore.

1:02:47 · Just like please take it and I hope that's enough." Now it feels like the definition of observability is more settled. Um, it's everything else in the world that's like changing and and crazy and all. So, I think it's a good book. I hope it can help a bunch of folks. I It's got six parts. So, the first part is and I wrote parts one and six. first part is just kind of like grappling with what does it mean to run deterministic and nondeterministic systems you know and then you know my co-authors uh Liz and Austin and George the part two and

1:03:16 · three is how how do you instrument your code and how do you understand it and there are parallel tracks for doing this with or without AI um and a couple great guest columns from uh from Jeremy and then parts four and five are we have a whole lineup of guest authors and use cases and deep dives. Hansen Ho did one on front end and um interesting and mobile. Uh we've got some great ones on CI/CD.

1:03:48 · Click House did one on Color Storage. Some really really stellar things. Uh there's a chapter from Kesha at Finn on how they use iteratively to like do observability.

1:04:00 · Oh, so so this is this is a brand new book. It's not A a lot of second editions are like, "Oh, we added like, you know, two chapters."

1:04:07 · This is an entire rewrite and it's twice as long. The first one was 250 pages. This one is 600 pages.

1:04:12 · Okay. So, I'm interested now. I'm going to get this book.

1:04:15 · And the part six, it's my baby. And it was originally supposed to be three chapters for observability engineering teams, and it turned into, it's a third of the book. It's 200 pages, but it's it's topics for observability governance for for leaders. And it starts with an open letter to CTO's telling them why all their big AI goals are blocked behind their ability to make sense of their system, you know? And then we talk about, you know, software delivery for no buzzwords, any just systems theory, right? Just if you like Danella Danella Meadows stuff, then you will like it.

1:04:50 · And then and then stuff and then there's a chapter on how to quantify the impact of observability for your finance. How to how how to treat observability as an investment versus a cost center and when you should use observability as a cost center and when you should treat it like an investment because it inherits the type of software that you're observing, you know. And there's a great guest chapter from Rick Clark on staff plus principal distinguished engineers who are trying to drive massive change without authority. How do you do that? H and how is observably vital to that?

1:05:20 · And then there's a chapter on build versus buy versus open source.

1:05:29 · I mean it it sounds to me that anyone who is inside or wants to be inside a platform engineering team, maybe you be an engineer or a leader.

1:05:37 · You're in charge of you probably want to read this book. And at the end there's a chapter that is possibly one of my favorites that which is it's called the art and science of vendor partnerships and it's just talking about how we can't build all the software that we need and great vendor partnerships are ones where you have influence over their road map and they trust you to you know do these things and like talking about how most transformations fail. The ones that succeed succeed because someone on the inside has trust and credibility.

1:06:08 · People believe when you say something, it is true. You know, it cuts through bureaucracy like a hot knife through butter. When it comes to partnering with, you know, the sales or of another company, you do not have trust incredibly. You you work to build trust through reciprocity.

1:06:26 · You learn just how much you can trust them over time, right? Uh but the best vendor relationships are the ones where you genuinely you feel like their successes are your successes. Your successes are their successes. You're happy to see each other because each of you are delighted because you know you're getting something from the it feels like you are two different teams working at the same big company. That is rare. Doesn't usually happen and that's fine. Most vendor relationships are ones where you shake hands, you exchange money and services and that's fine. But I think in an era of AI, these are durable skills.

1:06:59 · These are durable skills for very senior engineers who care about impact.

1:07:05 · Senior engineers and also engineering leaders and anyone who wants to become an engineering leader cuz I guess like I mean both of us have been in in engineering leadership like you've been in much higher positions than I have.

1:07:16 · But I think it's fair to say that the way for you to get to that CTO role, that head of engineering, that director of engineering is to do the work for six or eight or or 6 months, a year to year and a half. And to do so, you need to know these things. I feel observe the engineering might be underelling this book. I'll be honest, the title, but I'm I I'm I'm also going to get it and I'm I'm I'll probably think of ways to share a bit more. But thank you for writing.

1:07:43 · Thanks to you all to your to all your co-authors. But speaking of leadership, I'd love to talk about a little bit of engineering leadership because there's a a lot of things are changing, but I loved one of your very recent takes on leadership. And I'm going to quote you.

What effective leadership looks like

1:07:58 · The most effective leaders are kind, caring humans and skilled business operators. The second most effective leaders are terrible humans and skilled business operators. And after that comes any anyone everyone else. There are plenty of good, kind humans who are sloppy operators and bad at business because being good at business is very hard. And you said this in relation to what happened at Twitter/X, referring to as Elon as a terrible human, but a skilled business operator.

1:08:26 · Yeah, I well I I don't know that I would call him a skilled business operator, but my point was that Twitter had 16 years to figure it out and everyone [clears throat] could see that they were not figuring it out. and whatever else he figuring around the business specifically the business yeah building products you know reaching folks and you could argue that X has gotten better or worse but you can't argue that he is running it with 20% as many people Yep.

1:08:56 · And it's working and it's working and some of that you know 30 engineers on core on the core product and another 30 and but like 60 engineers there were 1700 before you know and you could argue and I think it would be true that it's some of the work that those engineers did that allow but like this is the point if we don't do it ourselves meaning hold ourselves to a high standard build with efficiency constantly be like trying to get better if we don't do it ourselves someone will come and do it to us.

1:09:28 · And and this is what what you also said you close saying if we want to remain in leadership, if we want to set the culture and the tone and take the ethical senses that we believe in, we first have to win at the business. And I I think this is like especially now that there's so many changes happening and in technology changes. There will be whirlwinds. Business will go up and down. I guess it's a reminder that like you want to keep your eyes on the prize, which is especially if you're a leader.

1:09:49 · the 2010s there was so much money slloshing around in Silicon Valley and time started to get tough and all of these companies canceled their DEI programs and blah blah blah. Yeah, they never believed in that that they were just trying to buy people off, you know, and and that is very telling to me and I have taken a lot of lessons away from that which is just that it's it's not enough to be a good person. I believe that people who are kind and care about people can and and usually do do better than sociopaths in the same roles, but only if they are good at business.

1:10:23 · Learn learn the business, stay close to it.

1:10:24 · You got to with AI now that coding has become cheap, now that engineers are are running agents, how do you see the role of of good skilled engineering managers and engineering directors change? What what has changed? Well, the first thing that's changed is I think everyone has to gets to be hands-on specifically to generate some code to ship to production to some extent.

Engineering management: what is changing?

1:10:53 · Know what it feels like to submit a diff to get a PR through, you know, you should should know what it feels like.

1:10:59 · It's just easier now than it's ever been to pick it back up, to fill in the blanks, you know, and it's always been the case that leaders were better if they had a hand in it. And now it's just it's just there's no excuse not to. Um teams are getting smaller in general. I think this should be a good thing. If we can figure out how to own more surface area, it should be a good thing.

1:11:22 · I worry that the way it's happening is is it's being done by CEOs who are like, "Oh, well this other company is doing it or it's magic or we're going to do layoffs or like it." And I really dislike the the anti-management tone. So like no argument that power tends to drift towards managers over time and needs to get pushed back to engineers. No argument there's a tendency to have too many managers.

1:11:49 · You know the bureaucracy kind of like generates a sort of you know it's it's easier to say yes than it is to say no. And so these things happen so they need to be pushed back from time to time. But I believe that management and middle management is deeply essential and I look forward to seeing how that works out for them not having any bit but like the role of a manager middle management in my view is sense making and context giving and cuz like I

1:12:19 · don't believe in a world where engineers are just given tasks here's your JI go do the things AI can do that I want people who understand what we're trying to do understand how we're trying to do it or who are there to help us figure out how we're going to do it. And you can't engage emotionally, creatively, collaboratively without an understanding. And that understanding is incredibly difficult to build and it's fragile. It never lasts very long.

1:12:46 · For those of us listening who are middle managers, it's been a tough few years because what they're seeing is there's a push to have fewer of them. A lot of their colleagues if they're in unlucky places they were made redundant and many of them have struggled to get similar positions. We're talking director positions. We're talking head of engineering, senior senior engine manager that that role is is disappearing faster than ever. I think directors might still be there.

1:13:13 · For folks who are in this position and they they do like middle management, they they do believe they're good at it.

1:13:22 · What do you think tactics could be to give them a bit more career options?

1:13:27 · Tactically, I would say go back to BNIC for a while, even if you know it's not what you want to do.

1:13:33 · If you're at all capable, if you're not capable of it, then I would try to work.

1:13:36 · You've got to get AI on your resume. You just have to like and if and this is a huge career risk. If you're working somewhere where you're not getting these skills, that is a massive risk. I would do whatever I could. And and this is very interesting that you're saying get get AI in your career because I remember about a year a year year and a half ago I started to pay attention to like okay this is happening and I remember a year ago I wrote an article about how to become an AI engineer and I talk with engineers who just like at their workplace start to do AI and now they're AI engineers.

1:14:05 · Next thing I'm hearing right now is the people who have like two to three years of AI engineering experience are so in demand. I'm I'm doing a research research on a job market and they're like, "This is the best job market ever." However, you know, the people who were like, "Okay, I have none, but I want to get it." They and let's say they're out of a job.

1:14:25 · They're struggling because no one's giving them the benefit of a doubt.

1:14:27 · It is really hard. And I'm not saying it's right, but it's how it is.

1:14:30 · And and I guess the reason the reason we're ringing this alarm bell is is we know this change has not been asked fast. So do it now because later do it now. The next time you go out for a job interview, anyone, you're going to be asked and you're going to be filtered out if you don't have it. And and the delta between those who are just getting started, most who've been doing it, it was here for a little while and it's very easy to get started. Now it's here and it's but it's opening up. The longer it goes, the more the harder it will be to catch. You just got to get you just got to get some.

1:15:01 · Let's talk about directors. Yeah, directors are usually the ones who they have been in management for like 10 years usually and there's a real feeling of fear often of like god tech has changed a lot in 10 years and and this is where I would say your body like the way we experience anxiety and the way we experience excitement is physiologically almost the same. Like I used to play piano, right?

1:15:31 · And before a performance, I'd be like, I'm excited. I am so excited to do this, you know, because I'm like trembling and sweat. But like the difference is agency. If you sit back and wait for the water to come to you, you're just going to be freaking out. But if you run towards the waves, if you're like just like run towards try it, you know, if you have a job now and you're a director and you're afraid of it, it's always seen as kind of noble when managers want to go back to being IC's. I think it's very well respected. Own it.

1:16:02 · Run towards the waves. Own it. Be part of the wave, the frontier of people who are like, I'm so Just tell yourself, doesn't have to be true. I'm so excited to be an IC again. It's never been easier to go back and try. I'm going to do it and then I'm going to talk about my experience and tell everyone else about just you got to own it. Don't wait.

1:16:25 · And then let's talk about junior engineers. Obviously, it's it's a harder time to get started as a junior. But how do you think about the value that they bring?

Junior engineers

1:16:34 · The hardest thing about quantifying the value of junior engineers is if we don't know how to quantify the value of any engineer. It's all vibes. You know, it's so interesting because I I feel like we're over here doing all this hand ringing about will juniors be okay? Will they ever learn the basics? When but like my my friend Boris who has a new observability startup and he talks to these high school college kids all the time. He's like they are cooking.

1:16:57 · They are they don't know what the software development life cycle is but they are just like off to the they are doing so much cool [ __ ] I believe that the kids are going to be okay. We just have to hire them. We just have to give them a shot. They're going to come up with a lot of the conclusions and the ways and the hows that are going to be things that we wouldn't have thought of because uh but we just have to hire them. We just have to be willing to give them a shot.

1:17:19 · Yeah.

1:17:19 · This week in SF, I've talked with a bunch of founders uh young startups and they've been telling me the stories of this open source contributor who was outstanding. So they wanted to hire him or her. Turns out it was 17-year-old kid. They still hired and now they're telling me like, "Oh my gosh, the things they do." So I think when you're saying the kids kids are going to be fine, just give them a chat and give them a shot.

1:17:40 · Even if it's uh an internship, Yes.

1:17:43 · I I feel more com because internship is is low risk, low duration.

1:17:48 · Yeah.

1:17:48 · And and even if that person doesn't work out with an internship under their belt.

1:17:53 · Yeah.

1:17:54 · So much better for everyone.

1:17:55 · Totally.

1:17:56 · One question that came up when when I asked that you're going to be in the show, what I should ask? They said AI fatigue. like there someone someone asked like can you please ask charity as an engineer if I'm starting to get just really really drained of this. Have you had this? Do you see people having it?

AI fatigue

1:18:12 · And what is a good way to you know just deal with it? We know it's here. We know it's here to stay. But still I mean my follow-up question would be like which variety of AI fatigue?

1:18:21 · Okay, tell us some other varieties you know cuz cuz some for some people when they say AI AI fatigue they're talking about receiving swap. Some people are talking about all the hype and the the Oh, have you heard the the phrase or the term uh doom trolling?

1:18:40 · No.

1:18:40 · Uh Cal I think Cal Newport is I think his name. He's a computer he's an AI researcher professor on the east coast and it's his term for what the CEO of anthropic and open AI keep doing about oh my god this might be the end of blah blah blah. And she's he's like, "It's just doom trolling and they shouldn't they they need to stop it because they're stressing everyone the [ __ ] out."

1:19:02 · Yeah.

1:19:03 · And stop because it's it's just not responsible, you know? So, like, yeah, I think there's a lot of fatigue around that. I think that a lot of people their family members are afraid, you know, it's just it's always before in the history of technology, it's been something cool or fun or this will make the iPhone, it'll make your life better and now it's just like fear. It's pretty crappy. So, there's that. There's there's the fatigue of like I found myself being off social media because I'm just so tired of all of the AI slop posts.

1:19:34 · It's just like I'm not interested. There are a lot of different varieties here and yes, we are all feeling it. Um so I guess I would repeat my call for us to remember that we are in control. We are in charge. I think the universal nature of the frustration means that this is a great time to propose experiments where we take back control. Maybe you and your team agree we don't actually want any more AI generated uh PR descriptions.

1:20:01 · We don't maybe we none of us use AI on Wednesdays. Maybe we take a week, you know, just like take control back, try something, propose something.

1:20:12 · I guess because change is so big, experimenting is has never been easier. And and I guess most businesses, most directors, most leaders would welcome teams saying, you know, we're going to try out because their answer will probably be, I mean, you're in this position. Your answer, I guess, will be sure.

1:20:28 · Better yet, don't even tell me. Come and tell me what worked afterwards and what didn't and what you learn.

1:20:33 · What didn't and then other teams can learn from that. Right there. I think sometimes people are waiting for top down permission but like we don't know what permission to give until it it works so much better when it's bottoms up when people are just try just take control of your time and your calendar and I I guess maybe we just forgot that there have been major changes in the industry. I remember the iPhone change and I remember the people when when the iPhone came out, iPhone and Android, so smartphones. The people who were the most kick-ass iOS engineers, you know who they were?

1:21:03 · They were typically like 18 or 19 year old kids who went into this and they tried it out. And guess what? Two years later, they were the domain experts. The staff engineer was a 22year-old and then the entry- level engineer was a 40-year-old. And again, not always, but my my point is in the when there's such big change, you can actually become an expert by very little time by you taking just taking charge taking taking charge and also no one's really going to tell you no because no one knows what's working.

1:21:31 · Exactly. Exactly. There's some liberty there.

1:21:34 · So, as as closing, just to go back to a little bit of of being human and slowing down, what are one or two books that gave you something? Ooh, I really got a lot out of Catastrophe Ethics.

Book recommendations

1:21:50 · It's not I haven't seen it mentioned many places and I think it might be I think real philosophy nerds would be like uh that's kind of a pop book you know and and I think the people who are not real philosophy books are like that's kind of a lot of philosophy but you know he's he's a bioethicist I think Travis Reer I ed catastrophe ethics and he talks about how the puzzle of modern life was that it feels like everything we're implicated every choice we make are you going to use milk while you know the cows are tortured are you going to use almond milk while water is a problem.

1:22:20 · Well, you swim like a little hormones and it's just like there is no whatever you do, you are hurting someone and it feels like the problems are so large that none of our decisions really matter and that tension like what? And then he kind of walks through traditional ethical frameworks like utilitarianism and stuff and just shows how there is no recipe anyone can follow.

1:22:46 · That doesn't lead you to some really stupid and and he's like this is just no gods, no masters. We are which doesn't mean that everything's relative. doesn't mean what it means is that the way to live an ethical life of integrity is you need to educate yourself about the world. You know, you need you need to be you need to know things, right? And then listen inside, you know, where where are you drawn?

1:23:16 · What suffering really speaks to you or what what caused you, you know, cuz cuz no one can tell you what matters. You have to decide what matters. And so that that introspection and it's so at odds with the sort of performative rage, you know, and that which I'm just so exhausted. All right, so that's one. Number two, this is a book that I've recommended a couple times, but I'm just going to keep recommending it because it's so good.

1:23:42 · It's by Adam Becker and it's called More Everything Forever and he is a journalist based in San Francisco. He has a philosophy undergrad and a PhD in astrophysics. And he just demolishes all of the AI religion, the singularity and the effect of altruism and accelerationism and the whole like what if we could have infinite growth foreverism. And he's like the heat death of the inter of the universe, you guys.

1:24:11 · Literally the only thing we know about in about exponential growth is that it must end. It must end in an S-curve or in a crash. must end. And he's got this dry sense of humor. And there are a couple times where he's just like describing some of the very real things.

1:24:29 · It's just like why do Oxford ethicists want this? He's talking about like taking over star systems and it's just ridiculous. And he also he gets in this a whack. He's just like talks about these people who are working so hard on life extension. And he he's like, "These are a bunch of sad little boys who miss their daddy." And I was just like, "Oh my god, it is the oldest fear of humanity is the fear of death." And you just see it. You cannot unsee it. So yeah, those are my two. They're both so good.

1:25:00 · Well, Charity, thank you so much. This finally made it happen.

1:25:03 · Finally. It's good time.

1:25:05 · I always really really enjoy talking with Charity. I hope you also liked it.

1:25:09 · I appreciated how charity talks about the trust account. If we are debiting trust from the creation of code because AI wrote it and no human read it, then that trust needs to be refilled somewhere else. Testing evals and guard rails are all ways to add more trust that we lost by using AI. I also appreciated how she talked with empathy about both AI camps. [music] The enthusiasts or AI pill folks are seeing the practical wins while those operating production systems see the slop. Neither side is wrong, but they should talk to each other more.

1:25:35 · So if you see wins with AI, share with the broader team, but also talk about it when it creates more work, reduces reliability, or when it degrades quality. And for those of us feeling anxious about all this change, especially directors and managers, I'll leave you with Charity's advice. Anxiety and excitement are psychologically almost the same. But the difference between them is agency. So instead of waiting for change to come to you, take charge however you can and make changes yourself.

1:26:01 · Do check out the show notes below for related to pragmatic engineer deep dives on how AI is changing software engineering and for another discussion with charity on observability. And I can very much recommend her book Observ Engineering second edition. If you enjoyed this podcast, [music] please do subscribe on your favorite podcast platform and on YouTube. A special thank you if you also leave a rating on the show. Thanks [music] and see you in the next