Is AI Making You an Average Engineer?
Summary
In this episode of The Agentic Review podcast, Liz Fong-Jones, Technical Fellow at honeycomb.io and Co-Author of Observability Engineering, joins Itamar Friedman and Nnenna Ndukwe for an open conversation about what happens when your team’s pull request volume more than doubles thanks to AI. Drawing on her two-plus decades in site reliability engineering, Liz shares how Honeycomb went from roughly 30 to 70 PRs on a peak weekday, and the honest tradeoffs that came with it.
this episode’s guest
Liz Fong-Jones is a Technical Fellow at honeycomb.io and Co-Author of Observability Engineering. With more than two decades in site reliability engineering and developer advocacy, including 11 years at Google and nearly 8 at Honeycomb, Liz has become one of the most recognizable voices shaping how modern engineering teams think about production systems, feedback loops, and reliability culture. Her work sits at the intersection of engineering and go-to-market, helping forward-thinking engineering leaders move their organizations faster without sacrificing the guardrails that keep production safe.
Key takeaways
- Why doubling PR volume isn’t automatically a win
- How to defend your codebase from regression to the mean
- Why observability, tests, and feature flags finally get built
- How to shrink incidents from hours to minutes
- A new model for training junior engineers
Chapters
- The 70 PR Reality Check
- Regression to the Mean
- "Eat Your Vegetables" Gets Easy
- Feature Flags as the Real Superpower
- The Airline Training Analogy
- Hot Take: Software Engineering Isn't Everything
Transcript
[00:00:00] Liz Fong-Jones: My hot take is keep the language models for the coding. Like, don’t try applying them to every single field just because you think that it’s like software engineering.
[00:00:09] Itamar Friedman: Welcome to The Agentic Review, the podcast where we explore what good code really means in the age of AI software development.
[00:00:17] Nnenna Ndukwe: I’m Nnenna Ndukwe, Developer Relations Lead.
[00:00:20] Itamar Friedman: And I’m Itamar Friedman, the co-founder and CEO of Qodo.
[00:00:23] Nnenna Ndukwe: So, let’s get into it.
[00:00:29] Nnenna Ndukwe: Before we jump in, follow or subscribe to The Agentic Review wherever you’re listening – Spotify, Apple Podcasts or YouTube – so you don’t miss the next conversation.
[00:00:38] Itamar Friedman: Today, we’re joined by Liz Fong-Jones, technical fellow at Honeycomb and co-author of Observability Engineering. Liz has spent more than two decades in site observability engineering and developer advocacy, including 11 years at Google and nearly 8 years at Honeycomb. Congrats.
[00:00:55] Liz Fong-Jones: Thank you.
[00:00:56] Nnenna Ndukwe: Her recent report on Honeycomb’s move from roughly 30 to 70 pull requests on a peak weekday is unusually candid, right, about both the gains and the costs of an AI-first strategy. So, Liz, welcome to the show.
[00:01:11] Liz Fong-Jones: Thank you for having me on. We’ve been talking about doing this for a while, and I’m glad that we’ve made it happen.
[00:01:15] Itamar Friedman: Did we miss something? An introduction that you’d like to enhance?
[00:01:20] Liz Fong-Jones: No, I think your introduction was great. I think the one thing that I would add is that, yes, I come from an ops background, as it were, but I think that everyone can benefit from understanding how we kind of close feedback loops, regardless of whether you come from an OPS or software development background.
[00:01:37] Itamar Friedman: So, just make sure that I understood. You’re saying, like, at this, like, regularly this week, 70 pull requests a day, like, it sounds like either a productivity breakthrough or system failure, or what should we know?
[00:01:51] Liz Fong-Jones: Why not both, right? Like, I think certainly Honeycomb’s philosophy has been systems failure is inevitable. Our job is to help make it as quick as possible to recover rather than trying to prevent failures from happening. I think preventing failures from happening is a lost cause. So, you know, yes, we did increase our development velocity. We have substantially more pull requests, not just, you know, being sent out for review, but actually being reviewed and landed in production. And also, you know, we are seeing more pressure on our operations. We are seeing more incidents. And I think that that is a natural part of the process, and it’s something that we have to figure out: what are our back pressure belts?
[00:02:33] Itamar Friedman: Interesting. Would you say, like that considering everything you shared, is AI making you and the team better engineers? Same? Different? Less good? Obviously, you’re pushing for velocity, and there is a philosophy that you can’t not have any incidents. So, let’s exploit the velocity and yet is it also impacting how good code quality you’re producing?
[00:03:03] Liz Fong-Jones: So, I think there’s several questions in what you asked. I think one thing that I would like to address right off the bat is that you don’t have to have all 70 of those pull requests be adding features, right? You can choose to use some of that new capacity that you’ve added to invest in your platform engineering practice and to invest in your controls in order to then make everything else safer. Right. I don’t think that it’s a reasonable thing to say, you know, let’s just push the gas pedal down, slop all the way. I think we have to look at how we are spending that additional bandwidth and whether we are increasing our total amount of technical debt or decreasing it. And then I think the other factor around code quality. Yes, there is some amount of concern, I would say, over our codebase quality regressing to the mean, if that makes sense. Right. These things are statistical output machines; in general, they are, to some degree, influenced by the quality of what’s already there. But also they tend to output, in the absence of any other guidance, statistically average code. And if your organization has better-than-average practices, right. Like, you have to deliberately defend your codebase to prevent it from regressing to the min. So, I think that’s a kind of open challenge, and you can see it most prominently in the form of these slop-laden PR comments, right, or comments in the code. Like, where it’s 5 lines of cloddish to explain something that everyone in the organization already knows, but somehow the AI finds it novel and, you know, like something that has to absolutely be noted down in a 5-line PR comment. Right. So, I think that’s the most obvious manifestation, but I think there are more subtle things as well that are potentially getting missed.
[00:04:43] Itamar Friedman: Yeah, I think it’s an interesting idea that, like, maybe you were saying that healthy organization to begin with, that with good practices, are moving faster, but they’re being pulled to the min, like converting to the min and they need to make sure that they put the best practices and guidance and care to not be pulled to the median. And I wonder maybe dysfunctional organization are actually becoming more dangerous, or the other way around, like, they’re also pulled to the mean, and then basically everyone, no matter where you started, go to the center.
[00:05:13] Liz Fong-Jones: Yeah, I know that’s really interesting. Certainly, I think if you don’t have any guardrails at all and you start shipping faster, you know, congratulations, you’re pushing the gas pedal, and you’re not steering at all, and you’re going all over the place. So, I’m not sure that I would agree with your hypothesis that below-median organizations are being pulled up to the median. Except maybe, you know, the AI is writing tests where people might not have written tests before. Maybe that’s when I don’t know. It’s a little bit of a challenge, I would say.
[00:05:40] Nnenna Ndukwe: I would love to know more about, like, how you see. What do you think is most effective about protecting, I guess, the quality of your code to not, I guess, regress, yeah, regress the quality and yeah, defend the practices and the quality that you have with the introduction of AI?
[00:06:02] Liz Fong-Jones: An interesting observation that I’ve been hearing around the conference circuit the past couple of months has been the idea that we’re having to do things for AI developers that were previously not legible at all, that were implicit assumptions in our organization and that writing these things down is to everyone’s benefit, both the new joiners in your organization as well as to the AI’s benefit. When you actually write down, “These are our policies. This is how we do,” you know, for instance, table-driven testing in our organization, this is what we believe that the optimal size of a comment is, right? Like, of course, you know, the human probably has a good intuition for what the optimal size of a comment is, but you can still make it explicit. Right? Like, so I think making your standards written down is a positive side effect of this. And I think a lot of this is encouraging us to do things that we probably should have been doing anyways. And I think, you know, we’re bound to eventually get to observability in this conversation. But I think, you know, testing, observability, these are all things that we needed to be doing for a while. And now, actually, there is a very, very compelling business motivation for ensuring that there are baked in as part of your quality process rather than just relying on, oh, the humans that we hire produce good code. How do you make it not luck? How do you make it systematic? How do you actually, you know, add, write down those standards, enforce the standards, whether it be by human review or by automated review or by some combination thereof?
[00:07:27] Nnenna Ndukwe: In this pursuit of, I guess, maintaining or improving standards, we have this opportunity now, if we choose to see it this way, like you said, with the introduction of AI, like, okay, well now you can add tests if you weren’t adding tests before, and that’s some type of improvement. Well, what would you say? How should leaders decide between just being able to get more features out there to customers versus improving optimizing internal systems? I think I know what your answer is based on how you’ve been kind of talking about things thus far, but we’d love to hear.
[00:08:03] Liz Fong-Jones: I think it’s both. Anyone who has run an engineering organization has seen the phenomenon of if you neglect your platform, that in the short term you’ll be able to push out a lot of features fast, and then eventually you’ll be so bogged down that every change you introduce adds to new bugs. Every bug you fix adds to new bugs. Right? Like, so you do have to make that foundational investment. But I would say the cost of making that investment is not as high as you would think. Right? You know, if you just tell the bot to include tests with every PR, and then, you know, maybe you might have to tell it, “No, actually, I meant meaningful tests.” Right? Like, not, you know, I call them tautological tests. You really don’t want to have tautological tests. Right? Like, the bot will happily, you know, work to rule and do exactly what you told it rather than what you meant. But, you know, the cost of adding these hygiene things is lower than it’s ever been. The same thing is true for observability and instrumentation. For years, years, I tell you, Charity Majors and I have been telling people, “You know, add instrumentation to your code, add spans, add attributes that are wide and rich events.” Right? Like, no one would listen to us. Everyone was saying, “Oh, it’s too hard. We’re already logging all of this stuff. We just want to push the feature out.” It turns out that if you just included a line in your agents file that says, “add spans and add meaningful attributes,” it will add them in the course of writing the PR. No additional typing needed. Right. Like, it’s gone from me trying to tell people, “Eat your vegetables. Eat your vegetables,” to “There’s vegetables included in everything now.” Right? Like, it’s great.
[00:09:37] Itamar Friedman: Yeah, totally. It still, like, resonates or, like, in my head, like, thinking about where we started saying that you cannot prevent incidents and so on. I think, like, it’s still, you know, if we have, like, an outage in our platform, it’s still a human that is in charge of it, right? And if now I’m adding double the amount of lines of code and features, probably also, I’m wondering if there’s double amount of incidents, and we can just live with that because we have, like, good observability or so and maybe, actually, the severity also matters a lot. Like, if I’m not looking on a code, then maybe, you know, it will be more severe incidents or more severe technical debt. Right? So, I’m wondering like, that philosophy. I see why it makes sense in the age of humans writing code. But I wonder, like, if it’s now time to say, “Okay, we need to improve and having less incidents – 10x less incidents for every line of code in order to really push the frontier, the harnessing of our AI, in order to move us faster with coding.”
[00:10:56] Liz Fong-Jones: So, you’re correct that it’s not sustainable for your existing pool of humans to handle twice as many incidents. So, the number of incidents is something that definitely is a concern to us. But I think one thing that we can do to mitigate that is, as you say, the severity. So, if every change that goes out – I should mention feature flags here, right – like, if every change that goes out has feature flags attached, you can shorten the duration of the incident to “Oh my goodness. We have to go and, you know, root cause this, and then we have to fix forward or roll back.” Right? Like, if an incident takes 1 hour or 2 hours, right? And then you can get the incident to shrink to 5 or 10 minutes just by flipping the feature flag. I think that reduces the amount of wear and tear that it creates on the engineer. And also, we are starting to see people be able to do some degree of AIOps. I’ve been very, very negative on the concept of AIOps in the past, but I think in limited ways, kind of when you’re using them to contain the blast radius, it actually can be very, very effective, right, to automatically identify which feature and then automatically handle doing the feature flag rollback. Because you know, it should be safe to roll back a feature flag, right? Like, so you’re not making things worse and only escalate to a human if the feature flag rollback fails, right, or it doesn’t solve the problem. So, I think just because it’s an incident doesn’t mean that a human has to handle the incident; an AI can handle the incident, and it can mitigate it for you. And yes, you as a human do still have to look at it, but maybe it lodges a ticket for you to look at during business hours, and you’re not awake at 2 am doing this. I think Charlie published something last week saying something about an out-of-hours page is more like a heart attack than diabetes, right? Like, we should treat it that way. Like, you should not become chronically used to 2:00 am pages.
[00:12:48] Itamar Friedman: Yeah. So, what I’m hearing is like AI is making it easier to eat our vegetables. So, it’s like we can almost force or get that review recommendation, or, to begin with, to have it injected within the lines of code even before we get a review from AI that we have observability in place, there is a feature flag in place, that there’s some core element in our architecture that should not change or should, and also make it maybe even easier for AI to collect later on the context in order to help us to resolve maybe automatically or semi-automatically. And then, even if there are incidents, then we can solve them in 5 minutes and an hour or two, maybe. Last thing that I feel like we didn’t say directly, but even, like, maybe estimate what is the blast radius of a change, like, try to estimate a severity. I almost feel like it’s an underrated killer feature of AI because it’s for humans. Especially if you’re not, like, a principal engineer. Right? It’s hard to imagine the entire system in the 15 minutes that you’re under pressure to review; like, it’s hard to crawl around the library. But Google does it really well. Once you get a document, you probably do it better than Google, but crawling all the documents, Google does better. So, actually understanding the blast radius and when you have all that vegetables and AI, etc., then you’re saying, “Yeah, we’re equipped to have the incidents but not high severity and treat them well and have the guardrail not to be pushed to the minion.” That’s how I recap my understanding of the insight so far.
[00:14:32] Liz Fong-Jones: Yeah, I think that’s completely fair. If we can make it so that the most junior engineer on the team, number one, is not getting paged at 2 am very often. And number two, even when they are, they’re not, you know, in a sink-or-swim environment because they’re adequately supported by the tooling.
[00:14:48] Nnenna Ndukwe: I’m very curious to hear you’ve written about how code review catches bugs and creates shared understanding. That’s like one of the, I guess, principles or elements that I shared in a course that I created about AI code review, or like the history of code review is an opportunity for learning, shared learning on teams among engineers. And one thing that worries me is what happens when agents are handling more and more or automating more of the catching bugs and resolving them. Where does the human naturally play a part in this shared understanding, this educational experience or element of working alongside other engineers? What do you think happens with the future when it comes to that?
[00:15:40] Liz Fong-Jones: So, one thing that we’ve experimented with doing is doing reviews, peer reviews sooner, doing peer reviews more often at the design stage or at the, you know, hey, is my specification to the agent correct? Is my description of the feature correct? Are we designing the overall shape of the system well enough? So, humans are having those conversations. We are not, you know, building software factories that delegate all of those decisions to AI because we think that it’s important for engineers to have ownership of the design decisions. And then when things actually become translated into code by the machines, and then they’re automatically linted for bugs by the machines. Right? Like, the questions the humans are asking are: Is this faithful to the original design? Are there potentially things that the agents did not find that we wish that they would have found? Right? Like, so that does get to designing software factory, right? Like that you are not necessarily completely going hands-off, right? You’re still saying, checking: does this have fidelity with the design that we as humans put together and we agreed upon? Are there categories of things that we should be checking for that we’re not checking for? Right? So, you know, I don’t think that removing things from the category of what a reviewer should look at is necessarily a problem. So, for instance, we have linters, right? And linters can programmatically catch issues and say, “Hey, that’s not formatted according to our organization standards.” It shouldn’t take a human to do that, right? Like you can have a deterministic robot go and say, “This is not up to our standards.” Automatically reformat it, or maybe this needs a design call, right? Like a human judgment or an agent judgment to decide which of these two patterns it should be. But right, like it’s a waste of your human’s time to be acting as human linters. And therefore we take the linting away. So, I think that is also true of trivial bugs. I think that searching this codebase to find, you know, hey, did someone put a less than where they should have done less than or equal? And that’s, say, you know, off by one, right? Like, edge, you know, boundary condition bug. I don’t think that’s a tremendously useful use of people’s time. We shouldn’t be thinking about the overall design and not, you know, each individual line of code: does it have one of these lurking bugs that’s an off-by-one? So, I think when you take toil, right, like we talk in the world of site reliability engineering, toil, right? Like work that is repetitive adds no value, right, like code review. Part of code review is valuable, and part of code review is toil. And we should be taking the toil out of code review so that we, as humans, are focusing more on the valuable parts of code review, right? I’m not saying don’t do code review. And also, like, we are now starting to have this conversation with our organization and Fin, which is now part of Salesforce. Fin is further along with us on this in terms of saying, you know, if someone is making a one-line change to a config file and changing one digit in that config file, it’s a waste of time to spin up a full human review cycle for that. The bots can handle that, right? So, even, you know, moving beyond, you know, which parts of code review should be automated away, we can also talk about which code needs to be code reviewed and which code does not need to be code reviewed. Right? Like, when you try to have humans pay attention to everything, it’s assigning no value to the humans, right? Like, it’s saying everything is important; therefore, nothing is important, right? So, we need to actually make these judgment calls about limited human bandwidth to say, “This is what your important role is, and this is where we should be focusing our efforts where they have the most value.”
[00:19:09] Itamar Friedman: That’s beautiful. Like, challenging it a little bit. I do think we’re going to a world where being under pressure to deploy faster, higher velocity, be competitive, and then maybe we will almost completely rely on AI code review agents, AI quality agents, to work together with the coding agent in tandem and do the review to begin with. And then, it really resonates with me that we’re pushing the engineers to review the spec, the architecture. What are we asking to happen? And then the tandem, AI agents or swarm, can work together. And if we’re going to that future, I’m kind of like the educational part. I think it’s so interesting that you were talking about now. I’m thinking like our old generation – I’ll put myself there – we were like, at least see one, do one, teach one. At least do it once. I was almost afraid as a student that I was being asked to see one and then immediately after, push something to production. It’s just my first time at work or whatever. And then teach one and almost feel like we’re moving to a world where it’s like there’s no seeing; it’s already doing, and you don’t look at the code, you don’t learn through doing the hard work, etc. And you don’t even try to teach or anything. Everything is codified, everything is remembered, and then you need to learn from there. I’m kind of like mumbling here because I think eventually there’s going to be a new generation that will have a different way of learning that is not like see one, do one, teach one. It’s going to be a little bit different; it’s like do, do, do and learn from the ‘doing’ directly. I even heard yesterday that some teams in financial institution that are realizing that it was so surprising for me that they’re saying that they understand that if they put enough guardrails, if they invest using AI more than coding agents, etc., maybe very soon AI could do as good as the median developer in code review and things like that. And then they need to work differently.
[00:21:28] Liz Fong-Jones: I think we’re already there. I think we are already there with AI catching more bugs than a median dev can. But here’s my answer to what you’ve just said. In the airline industry, when you go through your training in the simulator, if you’re, you know, an aspiring first officer, there are no training flights. There are no airframes that are dedicated just for aspiring first officers to fly with no one sitting in the back. The first time that a first officer flies, they are flying a real plane with real paying passengers. They have the captain in a seat next to them who is, you know, a training captain who has done this before. But right, like, I don’t think it is a requirement to practice in non-production environments first before you practice in production. I think that, yes, it is our responsibility to supervise people, but I think the learning comes from people making mistakes and people learning from them. And I would argue that the faster the feedback loop is, of you having the opportunity to make many kinds of novel mistakes and learn from them, the faster you will learn. I am a lot less concerned with, you know, has a developer struggled through writing Python syntax than I am concerned about how quickly can they learn from deploying Python software, let’s say, to production and then realizing what their mistakes are and then fixing them and instructing the agents better next time. So, I think you are robbing yourself of the opportunity to learn quickly in production if you are focusing on everything having to be perfect before it goes to production. I have to practice, practice, practice writing Python code. You know, I have to write assembly code before I can write in Python. Or, you know, I have to either use a slide rule before I can use a calculator. Right? Like, we’ve hashed these decisions out. Like we’ve realized what’s essential to know and beyond that, what just has to be learned from experience.
[00:23:19] Itamar Friedman: Yeah. Amazing example. So, basically we’re saying that you learn through the doing of the end result, to some extent, of what you were trying to achieve and what it did achieve, rather than the actual, you know, like hard work, romantic work of writing the code. And then, like, somehow to learn from that.
[00:23:38] Liz Fong-Jones: Exactly right. Like, if it takes you 2 months to write this change and then you push it to production and it fails, you’ve had one learning in 2 months. You’ve maybe had other learnings at the high thumb Shunix, but fundamentally you’ve had one learning in two months. If you are pushing to production every day, you’re learning, you know, 60 times as fast.
[00:23:57] Itamar Friedman: Amazing. And then I think we’re left with one thing. If I’m thinking, from the tech lead, from the manager’s perspective, it’s that we just need to make sure that that new faster loop that maybe actually reaches to higher altitude and velocity, there’s still less opportunity to, like, severe incident. Right? Because maybe I did one, only one, you know, one time learning throughout these two months, and now I can do like 20 learnings, but maybe the opportunity to fall big time, and the first 10 is higher. But I think that’s, A, actually, I could also help with some other tooling, Honeycomb included, etc that will help us, like, reduce the risk of Amasian incidents. And that’s the future. And I think it’s hard to, what I tried to convey and what I think we’re, like, aligning here, is that it is a mindset change. I think you’re presenting it really amazingly.
[00:24:57] Liz Fong-Jones: And cultural change.
[00:24:58] Itamar Friedman: Yeah. It is a cultural change that: I see one, I do one, I teach one, and then it’s like, instead, learn from doing, immediately close the loop, learn 20 times because of that. And then we need to put the systems that make sure that during that learning we have less incidents. And that’s, like, I think a major takeaway that I’m feeling the market is having right now.
[00:25:20] Liz Fong-Jones: Exactly. Right. Like, you know, in the airline industry, you are paired with a training captain for your first 20 or 30 flights. Right? It’s not, you know, they carry you with a brand-new captain and brand-new first officer and off you go. Like, it’s very much supervised, but there is also the opportunity to, you know, yes, they might let you make a mistake, maybe one that doesn’t compromise safety, but they’ll let you make the mistake, and then they’ll tell you about it afterwards. Right? And you can debrief afterwards. And then if something, as you say, goes severely wrong, right, like, the captain immediately takes control and says, “No, you’re not doing that.” Right? Like, you know, that’s right. Like, and now you’ve learned a very powerful lesson.
[00:25:52] Itamar Friedman: Yeah. I wonder, by the way, like, a philosophy, or in that analogy, who is the human and who is the AI?
[00:26:02] Liz Fong-Jones: No. Here I’m saying that both of these people are humans, right? Like, I’m talking about the problem of how do we train juniors?
[00:26:08] Itamar Friedman: You mean in software development.
[00:26:10] Liz Fong-Jones: How do we train juniors? How do we ensure that? What skills should we be encouraging CS students to have? I’ve talked to several university students in the Northwestern University program in Vancouver, and I was like, “Hey, you need to have some additional value beyond what the model is doing. That value is your ability to learn in the long term and not just in the short term.” Right? Beyond that, humans have a massive amount of memory, and we should be leveraging that. Right? That’s your superpower, the writing of the code, yeah, an agent can do that better than you or even I could.
[00:26:47] Nnenna Ndukwe: I think bringing up the students is something that I wanted to bring up. So, I’m glad that you did. And that is, you mentioned the power of reviewing specs now and thinking more about system design. And so, it also sounds like you’re saying that, from what I think, that’s going to be more important now for junior engineers to be thinking like and be participating in those types of conversations. I remember working – I spent 4 years working at O’Reilly as an engineer, and one of the most exciting things, things I loved doing, was having the architectural discussions about ways in which all of the services that we owned were going to change based on what was in the product roadmap that we needed to build. I really loved talking about architecture, making decisions and recording that in ADRs – maybe just because I’m a systems person – but it allowed everyone to have a voice, I think, and to think about the overall system and how we expected it to evolve and to be prepared for that. And I would love – that’s what I imagine those types of conversations could be like earlier, in the way that engineers can collaborate with each other on the system design and the specs. And yeah, I wanted to hear your perspective.
[00:28:10] Liz Fong-Jones: I very much agree. And you probably only had that systems design review and ADRs maybe once a month before, right? Like you were not having ADRs every day, every week, right?
[00:28:21] Nnenna Ndukwe: Right.
[00:28:22] Liz Fong-Jones: But now that engineering velocity is less constrained by the writing of code, you can be, you know, you can and should be doing that ADR process for every major change that you’re making. And those major changes that you’re making are now happening every week, every day, and not, you know, every month or two months, right? Like, so the speed at which we evolve our systems and also get the feedback to know, hey, that decision was a bad one, or hey, that decision worked really well, right? Like, you can do that much more. You can accelerate that and do that much more quickly. And I think in terms of the role that juniors should play here, right? Like, yes, you know, the junior should get some practice with writing some code and, you know, compare it to what the AI outputs. But I think one of the constraints that we have now is decomposing broader features into smaller tasks that an AI can safely accomplish, right? And I think that that is an example of, yes, I could abdicate, you know, that decision and say, Opus or Fable, go do it. But I think that actually a junior is going to do a better job with that, right? Like, a junior is going to do a better job of task decomposition and also learn something from it.
[00:29:24] Nnenna Ndukwe: Amazing. Well said. I want to make sure that we also get to talk a lot more about observability. And I want to start by asking, what is it that you focus on as a technical fellow at Honeycomb? For everyone who’s listening, if you could just share more about that.
[00:29:42] Liz Fong-Jones: Yeah. So, I’ve been at Honeycomb almost 8 years now, and my sole job is to force-multiply head income. My sole job is to help us accomplish our mission of making production legible to every software developer and every agent. And sometimes, for me, that means talking to customers. Sometimes, that means coming on podcasts like this. Sometimes, it means doing hands-on development effort on some of our – well, not quite hands-on, but hands operating the robots, right – like hands-on work on our codebase, delivering features or kind of delivering prototypes that our customers have asked for. So, it’s a very multifaceted role, and it basically boils down to being at the intersection of our engineering and go-to-market teams and focusing kind of on the persona of the forward-thinking engineering leader who wants to move their production, their organization faster.
[00:30:36] Nnenna Ndukwe: Thank you for that backdrop.
[00:30:38] Itamar Friedman: By the way, we say, like, a hands-on, and I’m wondering, like, if we’re moving more and more towards ADRs and writing code. Like, I was skeptical a little bit a year or two ago, but if it’s moving from hands-on to voice-on, practically. Like, use it directly, and you’re like, I hear a lot of ticking in the background, and some people are doing it really fast. But still, I think like the direct access to our brains is speaking, right. So, it’s like reducing one another layer that is maybe redundant.
[00:31:07] Liz Fong-Jones: Oh, I actually have profound disagreements with people who think voice is the future. My profound disagreement is that the context of an agent really, really, really strongly influences what it does. If you are concise and controlled in what you give the agent, your agent will give you concise and controlled output. If you ramble in your input, you are going to get rambly output, and it might not do what you intended. I think maybe there is a layer of, you know, you voice-transcribe something and then you review what the generated prompt is, and you can make it more concise before you send it to the higher model. Right? So, I can definitely envision, you know, maybe you talk to Haiku, and Haiku makes it concise, and then you send the Haiku-revised prompt into Opus or Sonnet. But I don’t think talking directly to Opus or Fable without something in between is a good idea. Yes. To your analogy about, you know, being hands-on, going back to my aircraft chronology. Like, you know, right, the student, first officer has to know how to operate the autopilot. They have to understand how does the airplane behave because you are not directly flying the airplane. It is all fly-by-wire. Right? So, I think, right, like, we are kind of, as software developers, moving into this era of we are no longer directly controlling this particular flap. Like, we are moving towards figuring out how do we trust the computer to execute our intentions. So, I think, you know, being imprecise with what you tell the robots can be dangerous.
[00:32:32] Itamar Friedman: I think that a lot of the usability depends on the training that the Foundation Labs are doing. Actually, they’re, to some extent, controlling some; they have a big influence about the future of UX/UI. Like, I think one clear example is that ChatGPT started to work well because GPT-3.5 was fine-tuned for instruction following. Right. And I do see a future where, just like my take, okay, we will see that I don’t have mathematical proof, is that I do see a future where this creativity on block, at least in how Foundation Labs sees it from voice chat, including professional ones, would be an important use case for them, so they will kind of like, train the models to deal with it to some extent.
[00:33:28] Liz Fong-Jones: Or design systems around the models. Right? Like, that multi-step approach that I was talking about. Right. Like, I think, with today’s models, directly voice dictating to Fable is probably a little bit wasteful in tokens and accuracy. Yeah, I think we’re in agreement there. It just boils down to, you know, how do we most effectively use the modalities that we currently have access to?
[00:33:50] Itamar Friedman: Yeah. Completely.
[00:33:51] Liz Fong-Jones: But speaking of modalities, though, right, like, I love living in the future where I can use my phone to communicate with Claude Code running on my Mac, right, and it works, right? Like, I can type a few words with my thumbs and get, you know, 100 lines of perfect code. I would never write 100 lines of code with my thumbs.
[00:34:11] Itamar Friedman: Yeah, completely.
[00:34:12] Liz Fong-Jones: But I think I am accurate and concise in what I say, and I think that’s what I typed with my thumbs, and I think that’s why that works, at least for me.
[00:34:21] Itamar Friedman: Yeah, I think the architecture you’re suggesting actually might be already implemented. I think there is kind of a regression model on top. Like, it’s not completely as you suggest, but I think that there is a direction towards there. I actually feel that I don’t want to name names of products, but I think, like, until 2 months ago, one of the leading products did not work well with voice, and the other did. And it was very clear that the one that worked – well, there was a regression model even maybe a local one, that kind of, like, understand, oh, did you finish the sentence? Is it high probability you finished it? Is it clear enough? How can I take what’s concise with it? That’s, like, I think definitely where we’re going towards. I think that’s the vision.
[00:35:05] Liz Fong-Jones: Yeah. The other thing that I’ve seen, not related to even coding, right, like, ever since Granola was rolled out at my organization, so many people are just dictating, or, right, like, are just, you know, going just top of mind, you know, saying what’s on their mind about a subject, and Granola will condense it into something that they can work with to actually, you know, do technical communications. Right? Like, so I think that’s been a huge unblocker, as you say, to creativity, to get these ideas out of people’s heads where they might not, you know, have the patience to go and type it out in perfect words. Because people have this tendency we want to edit as we’re writing rather than just get all of the material out there.
[00:35:44] Nnenna Ndukwe: I think this is considered one of the hot takes that I really wanted us to get to with you, Liz, which is that voice mode is not the future. This is what you’re saying. This is what your opinion is right now. Do you have any other hot takes or opinions that you’ve shared before has not landed well with people, and you’re just like, I don’t care? This is how I feel about AI and software engineering right now.
[00:36:12] Liz Fong-Jones: Yeah, sure. I think that we as software engineers and software developers have a huge amount of hubris about our field. Our viewpoint is that everything is just software engineering. The same way you’ll see arrogant mathematicians being like, “Oh, everything is mathematics. Physics is mathematics.” Or you’ll see physicists being like, you know, chemistry is just physics or biology is just physics, right? So, we’ve seen this trend of software developers proclaiming that everything is just software, and they don’t know it yet. And therefore, you can apply language models to everything. And I think that is an immense mistake. Just because something works for software development does not mean that it’s going to work for all fields. Because I would argue software development is special in that we have the ability to rapidly test our hypotheses without, largely speaking, people being injured or hurt. And we have special rules for those industries where we won’t be injured or hurt, right? We have deterministic tests. We have the ability to do version control and robot rollbacks. The only compilation test for a legal document is what an angry judge thinks, right? I am very, very, very skeptical about attempts to just wholesale say, well, you know, the law is close enough to code, right? Like, you know, it’s got a set of rules. Surely that should, you know, mean that we can, you know, do legal analysis or write legal documents with language models. So, yeah, my hot take is: keep the language models for the coding. Like, don’t try applying them to every single field just because you think that it’s like software engineering. Software engineering is uniquely special in terms of our ability to put guardrails on it that other fields are not.
[00:37:39] Nnenna Ndukwe: Amazing. Amazing point. I think that’s like probably 5 hot takes. If people really think about it, in one. And I love it. I’m here for it. I have one question about software factories for you, and that is: what kind of power do you think observability can have with everyone’s goal of achieving autonomous agents with software that’s writing software and doing everything that it’s supposed to do for the most part? How do you see observability playing a huge role in that?
[00:38:13] Liz Fong-Jones: Just because something compiles doesn’t mean that it doesn’t have bugs. The only way you find out about the bugs is in production. So, if your agents are just slinging code and they have zero visibility into what’s actually happening in production, you could be serving 100% errors, and the agents would never notice, and they would just keep shipping. Right? Like, so the way that we avoid shipping broken code on top of broken code on top of broken code is not just on the test screen, but actually: how is this actually doing in production? Are we introducing performance regressions? Are we serving errors? What property do those errors or slow requests share in common? How can we go back and change the code to prevent that from happening both in the short and long term? Right? Like, so I think observability plays that very valuable role of providing the feedback loop to your software factory: is it actually working? Is it actually delivering value and not just, you know, lines of code that are sitting in your trunk?
[00:39:03] Nnenna Ndukwe: What portion of that would you say? Like, one thing that I always think about with generative AI is the reactive nature, I guess, of the way in which people are building tools, or generative AI is a part of the features or the experience of it. But there is a proactive nature of artificial intelligence and machine learning that I feel like we almost need to get back to building that into. We need to build those into products more than I think that we are currently doing. And that’s part of how I perceive observability being incredibly powerful, more so moving forward in the future. So, how much of what you’re saying is reactive versus proactive?
[00:39:49] Liz Fong-Jones: There are a million and one startups pursuing the reactive angle, and there are very few startups pursuing the proactive angle. And I think that is a mistake. I think, you know, anytime a startup calls themselves an AI SRE, and then they focus only on instant response. I think it is underselling our profession as site reliability engineers. Right? So, I think that, to your point, there is a very valuable role to be filled not just by, you know, hey, something broke, root cause it, and fix it and do the feature flag flip which we talked about earlier. Right? But also, how can we proactively identify weak points in our systems using that observability data? Can we identify things before they go wrong? Can we fix them before they become issues? And then also, you know, audit for those design patterns elsewhere in our codebase and prevent new instances of the bad design pattern from creeping in? Right? Like, so I think that’s kind of the broader systems-level view, and this is something that we are actively working on at Honeycomb: not just the, you know, immediate sleuthing through your incident, but also kind of the broader architect persona. Right? Like, well, what would a principal SRE be seeing in the system? And how can we surface those insights to you based off of the kind of collective memory of your team and of your telemetry?
[00:41:02] Nnenna Ndukwe: Wonderfully put.
[00:41:03] Itamar Friedman: Yeah, I love it. I think one perception some people might have is that we have the infinity loop of software development.
[00:41:11] Liz Fong-Jones: Oh, the DevOps infinity loop. Yes. Everyone is sick of seeing it.
[00:41:15] Itamar Friedman: And as if this is like, okay, stage, you move to another one, and every time there is a certain person that is getting the, you know, stick and passing it to another. But I think what you’re describing is more like, let’s say a few people working together in a spiral that is progressing and then that person that is in charge, for example, of reliability, it’s not like being scoped to a specific part. It’s actually working continuously. Like, so when there’s design, let’s already inject reliability into it. When you’re in the PR process, if AI needs some help to check out the reliability, maybe you’ll already come there and not just when there is an incident. And I think that’s the future. Like seeing that multiplayer game of multidisciplinary developers and agents working together in that spiral.
[00:42:06] Liz Fong-Jones: And also, why does it have to be a single person or single specialist who does it? Why can’t we empower everyone to be able to do it? Right? I think, you know, we have decided as an industry, collectively, except for some of the laggards, that, you know, it is a mistake to have dedicated software engineers in test. Right? You might have, you know, software engineers who develop testing systems. But broadly speaking, you as a software dev don’t get to say writing the test is someone else’s job. Right? I think the same thing needs to be true of reliability as well: you know, ideally your SREs should not be responsible for the reliability. Right? Like, we should be designing the systems that ensure reliability, and it’s up to the individual teams to ensure that their products are reliable.
[00:42:46] Nnenna Ndukwe: Okay, so where can listeners follow your research, your writing, your work contributions? Yeah. Where can we find you?
[00:42:55] Liz Fong-Jones: Well, you can certainly pick up a copy of my book. Also, there is the Honeycomb blog. I highly recommend the Honeycomb blog because it features voices besides just my own and charities from engineers all over Honeycomb. And I’m on various social media except for Twitter. So, you can find me usually as lizthegrey on any of these sites.
[00:45:15] Nnenna Ndukwe: Thank you so much for joining us and thank you for listening. If this conversation gave you something very useful to take back to your team, follow or subscribe to The Agentic Review on Spotify, Apple Podcasts or YouTube and share the episode with someone responsible for the quality of AI-generated code. Thank you so much, Liz. This has been a wonderful, dynamic conversation. I think that we really needed this, and it’s very timely.
[00:43:41] Itamar Friedman: Thank you.
[00:43:42] Liz Fong-Jones: Thank you for having me.
[00:43:43] Itamar Friedman: And a shout-out to Nnenna when you mentioned the code review webinar seminar that you recorded, deploying AI, right? You can find it there.
[00:43:52] Liz Fong-Jones: Yes. That is true.
[00:43:53] Itamar Friedman: Thank you.
[00:43:55] Outro: If today’s conversation challenged how you think about AI and code quality, that’s the point. At Qodo, we believe that independent, context-aware code rev, context-aware as guardrails, is how engineering teams maintain standards at scale. If you’re leading an enterprise team and want to see how intelligent AI code review can reinforce governance, visibility and accountability in your workflow, visit qodo.ai to learn how we help teams turn AI program productivity into production-ready quality. And if you enjoyed this episode, subscribe, share it with your engineering leadership circle and leave us a review. Until next time. Keep human in the loop and keep shipping.
About the hosts