Building the AI Software Factory: Quality and DX with Amir Rustamzadeh
Summary
Building a reliable software factory requires more than just generating code. This conversation focuses on the importance of quality assurance and the evolution of developer experience in the era of autonomous agents. Learn how to implement rigorous entry and exit gates to ensure that AI productivity does not come at the cost of technical debt.
this episode’s guest
Amir Rustamzadeh is a Partner at Firestreak Ventures, a fund focused on the frontier of developer tools and AI infrastructure. He previously spent four years at Cypress, where he founded their developer experience function and created the DX engineer job title. Today, he backs companies like Lovable, Modal, Daytona, and Hayes Labs while maintaining a hands-on approach to agentic engineering.
Key takeaways
- The transition from spec-driven development to a model where documentation serves as the primary interface for both humans and agents,
- Why the industry historically struggled with testing and how AI presents an opportunity to finally achieve rigorous quality assurance,
- Strategies for building a software factory by breaking down the development life cycle into repeatable, governed phases
- The importance of setting up strict entry and exit gates for long-horizon agentic tasks to ensure confident outputs
- Navigating the build versus buy decision by evaluating the long-term maintenance burden of internal developer tools
Chapters
- The evolution of developer experience
- Retooling the engineering workforce
- The case for docs-driven development
- Constructing the AI software factory
- Implementing quality gates for agents
- The true cost of building internal tools
Transcript
[00:00:01] Amir: I think we were asking too much of people. We were saying, “Hey, ship fast. It better be good, and you better maintain your test, and your CI better always be passing.”
[00:00:11] Itamar: Welcome to The Agentic Review, the podcast where we explore what good code really means in the age of AI software development.
[00:00:19] Nnenna: I’m Nnenna Ndukwe, Developer Relations Lead.
[00:00:22] Itamar: And I’m Itamar Friedman, the cofounder and CEO of Qodo.
[00:00:25] Nnenna: So, let’s get into it.
[00:00:31] Itamar: Today, we’re joined by Amir Rustamzadeh, partner at Firestreak Ventures, a fund at the frontier of developer tools and AI infra, where he combines hands-on product instincts with AI native engineering skills to help founders build durable companies like Qodo. He spent four years at Cypress before he began backing the next generation of tools.
[00:00:54] Nnenna: At Cypress, he founded their developer experience function and created the DX engineer job title at a time when nobody had it, then watched Vercel, Netlify, and the rest of the industry follow. He now backs companies like Lovable, Modal, Daytona, and Hayes Labs. And he’s deep into agentic engineering himself, running over 1,200 end-to-end tests on his own cloud projects.
[00:01:20] Itamar: So, full transparency, his portfolio includes Qodo, and he got skin in the game, and in this conversation, more ways than anyone else. Amir, welcome to the show. Without further ado, let us know what did we miss.
[00:01:32] Amir: No. You got it all. You captured the full sense of it all, and thanks for having me on.
[00:01:37] Itamar: Awesome. So, let’s start with a bit about your background. I think you coined the term developer experience almost at Cypress at least and built the function from nothing. I think you said that you kept searching on LinkedIn and looking like anyone that held that title. And looking back, what are you actually solving for back in the days, and what do you feel it was missing? And maybe you can take it also what it means for us today, like, in the age of AI.
[00:02:04] Amir: Yeah. So, look, at that point in time, there was a bit of a renaissance of developer-focused companies. These are kind of the companies we now take for granted today, but you know, not long ago, they were kind of the pioneering and the kind of the cool companies at the time, building stuff for developers like no one had ever thought to do before. You know, historically, building for developers wasn’t necessarily, like, you know, something you would start a company around. It was always relegated to the large companies like Microsoft or something like that, where they kind of owned the language, the IDE, the place you run the code. I mean, you know, don’t even get me started on, like, open source, right? And so we kind of entered this era, I would say, man, probably past 2010, where we start to see this stuff start to change. We got more open source activity, and I think that led to a lot of new companies being formed around those open source projects. And so it was a great time. I think it was a very fun time to build developer products and developer tools. But at the same time, I think the industry had not really toiled with what it meant to build an org that builds stuff for developers. We were kind of trying out the cookie-cutter playbooks of, like, this is what a PM looks like that hands things to engineers, and this is what marketing looks like. It was just all cookie-cutter. And at Cypress, you know, we were growing really quickly, and we kind of had to figure all these things out as we went along. But I quickly realized, man, I feel like the old ways of doing things just doesn’t translate into what a modern developer-focused product should look like. I don’t think you can go and grab run-of-the-mill product managers to go build, you know, to go figure out, like, what the engineers should build next, especially when the engineers were, like, the end ICP for what you were building. So, I knew we had to think about it differently. At the same time, we had to think differently about how you market developer products, right? A developer, or I would say, a product evangelist have always existed in tech, but there were always, kind of, functions of classical marketing departments. And those departments function in vastly different ways, I would say, modern developer relations or developer advocacy departments operate today. Wherein in that world, you know, developer experience, you’re kind of doing things for the long term. You’re making bets today so you can reap the rewards later. These are activities you do to build community, to create an ecosystem, and these things compound over time, and you get to reap the benefits of it much later. But traditional marketing approaches heavily relied on very tactical things where they’re easily measurable and easily attributable. And, you know, if you talk to any DevRel person now, this is one of their biggest frustrations. So, we had to think differently about product, think differently about how we go to market, and then we also had to think a little bit differently about how those two things integrate back into the engineering work we had to do. So, I felt that, you know, to some degree, there was a compression of roles across all these things where we could have really solid engineers that were very good marketers as well. They were very good storytellers. They really wanted to talk about the things they were designing. And because they had a hand in developing the product, they were much more authentic when they were out there in the wild talking to hundreds of developers in front of large audiences. And so, I like that authentic approach. But the other part of it was just out of sheer necessity. You know? Like, we were forced to almost kind of invent this because our competitors were the large companies. It was people like Microsoft and Amazon, and, you know, they’re super well funded. And they could decide to do whatever they would like, you know, relative to your market position and then swarm planet Earth with their global developer advocacy teams. They could be anywhere all the time always. And so you couldn’t play the same game as them. You had to play differently. You just didn’t have as much money as them. So, these were kind of the underpinnings of why I thought, hey, we kind of need to create, you know, a very good task force, which became the DX engineering department. And every person in that team kind of was like a member of the Avengers. You know? Like, one person was very good at, like, dealing with the AngularJS world. They were super tied into that, and they were respected in that community as well. And some of them were actually popular open source maintainers themselves. So, we would respect them because we were using their projects internally before we even hired on. So, anyway, these were kind of some of the reasons.
[00:06:36] Itamar: Yeah. So, Nnenna is in the DX department. Do you feel like you’re in the Avengers? What do you think about that? That resonates with you?
[00:06:41] Nnenna: Well, absolutely. I mean, I feel like there’s just so much that there is a multifaceted nature to this kind of role, and you know, depending on the priorities, of course, a company might index a bit more in one area than another and might just have certain natural skill sets that are more focused on one area than another. But I find it to be incredibly powerful and important and necessary, especially for just developer tools in general. Different mindsets, product mindset, you know, a developer mindset, empathy, and a builder, you know, it’s a lot, and I love it.
[00:07:18] Itamar: Yeah. So, we’re talking about the fact that necessity brought a new role where kind of being able to tell the story and have that marketing instinct, but being rooted in the engineering, but also understanding, like, product puts you in the front, like, and connect you all along and bring from all the development life cycle, let’s say, of a product that is selling to developer. And now I’m thinking, is there any lessons here for the age of the era of AI? Like, this is not necessarily related to companies that are building dev tools. Any organization, they have engineering organization, we’re seeing maybe the same, like, process as everybody basically DX. Now your audience is not just, you know, your clients. Your audience is also your agents. You need to be able to tell them the story of what you’re trying to achieve, give them the context, knowing what good looks like. So, I’m kind of, like, wondering, do you think, like, there’s any, like, learning there to what engineering managers need to think about creating a new role? And maybe I’ll connect it to what actually good engineering looks like, which, like, previous or existing engineering teams inspired you. I think you talked about NASA, JPL, SpaceX, etc. So, basically, I’d love to hear, you know, from your point of view, what good engineering looks like, what good engineering is gonna look like, and what does it mean, like, are we expected to see, like, a new role that is helping to push the engineering to the next level?
[00:08:57] Amir: Sure. Well, look, I think everyone’s role is kind of changing at the moment. Everything’s in flux. Maybe the titles might not change, but I think the responsibilities and scope of work will definitely change and have changed. So, you know, we were talking about the DX engineer role. And as I mentioned, this was kind of a compression of multiple roles. We were effectively forcing, you know, one individual to be more full-stack. You know, they had to kind of carry things from start to finish. And that made hiring for the role very hard because everybody was very, you know, used to doing one thing. You know? I’m just an engineer. All I do is code. No one sees me. No one talks to me. You know, that kind of thing. The industry just favors bucketizing people into, you know, you are this, you are that, and don’t steer away from that. And the org structures are set up around that, and the company cultures are set up around that, which is frankly the hardest part. But now with all these new, more intelligent tools, I think the expectations are very different. I think we can now expect people to accomplish more in the same amount of time that’s given to them. And with that, also has to come more autonomy and more agency to, like, the individual, you know, team member, where they can do, you know, more tasks or kind of, I would say, not necessarily more task in volume individually, but more varied activity, where, “Hey, I didn’t just ship a feature, I didn’t just make one PR today, you know, I’m happy about that, and thankfully, there’s some tests around it or something like that.” They’re thinking holistically about, like, oh, this feature is a part of a holistic strategy for the company, and I understand that. I’m a contributor to that. I understand the end goal here. I understand how it impacts other things that we’re doing, other initiatives. So, I think the individual engineer is just becoming much more patched in to kind of the goals of the company, and I think that’s great. But, of course, that also comes with more responsibility, and I think leaders need to understand that. And I think we have, you know, this isn’t gonna happen overnight either. We’re kind of onboarding, and you know, ramping up the entire industry and professionals that have worked for years doing things in a certain way. So, I think that’s the pressure now, which is how do we kind of retool and retask, you know, everybody to be more full-stack?
[00:11:15] Itamar: Yeah. I’m wondering, basically, like we’re saying, like, everyone becoming a little bit more full-stack. Will that help quality or not? Like, I think you both have, like, experience from Cypress, like, being a company helping in quality, right? And also as an investor looking on the field and how it’s evolving, like, right now, like, it’s mostly chat. We moved from tab, tab, tab to chat, chat, chat, and now we’re moving to agentic workflow, right? And, basically, like, with that evolution, people can do more with those chat and agentic workflow. And I wonder if it’s helping quality or reducing quality. Like, on one hand, for example, there are mistakes when you pass information from one human to another. And if it’s ever seen in people, like, brain and sorry, in the same person end-to-end, then maybe there’s less mistakes. On the other hand, let’s admit it. A lot of decisions and information is now actually moving into sessions with the agents into, like decisions that are made during the reasoning of the coding, etc. So, basically, I think, like, code quality was always a big thing, like a trillion-dollar problem, Cypress being part of it. But I’m wondering if you’re seeing it as someone experienced, also investors, something that is gonna be even a bigger problem that we have to build new solutions or actually, like, the fact that a person can do end-to-end is gonna reduce the quality issues.
[00:12:45] Amir: Yeah. I think with the power we have now, with everything, I think the problem is gonna be even more pronounced because we’re trying to create more things faster, right? And, you know, you mentioned something, which is, you know, is quality gonna go down with this kind of new state of affairs or up, or, you know, is quality gonna go down, for example, if you have one person doing more things? I think it will actually go up because that individual is not too distant from a lot of the decision-making that had to happen, right? They fully internalized it. They understand kind of the essence of what’s trying to get done, which are sometimes not the things that get captured in documents or or any anything else, right? Like, everyone’s much more aligned. I’ll give you an example of this view at Cypress when we were kind of really trying to better integrate kind of this DX engineering approach to the rest of the company. So, like, one of our first projects was designing kind of this network interception API. And, you know, we had designed it using kind of this doc-driven development. We kind of try to do a lot of that, where we try to, you know, effectively say, we’re gonna document this feature how we want the user to, like, experience this API, and then we’re gonna work backwards with the implementation. That’s actually how we ended up solving the problem. But before we were doing this, and this is, you know, more widespread and common in the industry, you know, we try to go kind of the specking and then, you know, direct implementation and, like, is the spec good enough? Was it sufficient? Did it really understand the technical hurdles of making this thing? And then, you know, we had an engineer go out, and you know, it was a pretty, you know, difficult thing to implement. And this person went out and technically did it. They technically did it, and you know, he came back, and we were reviewing it, and it was super complicated to use. Like, I don’t even think he even knew how to use it properly, and he just, like, got it done just barely. And then we went to document it, and we were like, wow, this is super hard to explain. Like, we took something that was supposed to be simple, and it just feels icky, wrong, the quality doesn’t seem to be there. It’s not us. We weren’t giving the kind of experience that we wanted. And at that point, it was kind of like, “Alright, never again. Never again.” Like, we’re building very tough things to implement, and so we always gotta work backwards from, like, the experience we want. And, you know, this is kind of the classic Amazon style, if you will, where, you know, Jeff Bezos is, like, famous for saying, you know, we’re gonna write the press release first, and then everyone works backwards from a press release. And I think you can do that for everything pretty much. Because if you don’t end up with the final experience you want, you did it wrong. Go back. Go back. Why ship it?
[00:15:25] Itamar: Yeah. That’s interesting. Do you feel like, basically, what you’re advocating for is instead of spec-driven development to do docs-driven development and to some extent, and maybe that’s actually today more possible than ever than when you folks try to then, Nnenna, I think, like, this is something I speak to you, right, like, you’re all day long, like, working with open source and different tools, like, ExClaw etc. Like, I think both of you are doing hands-on. How do you think about that approach of docs-driven development?
[00:15:56] Nnenna: Yeah. I mean, my first thought is, like, I could see how that is actually more very agent native, or at least it can be consumed in a way that could be a context that’s applied across different pipelines using agents to execute on different tasks. So, that was initially what I thought of, like, okay, I can see this being a very forward-thinking type of motion. What about you, Amir?
[00:16:19] Amir: Yeah. I mean, I guess, you know, there is no silver bullet of, like, oh, it should be this way, or it should be some other way. I think you kind of have to pick and choose how you wanna go about it. I think for some things, creating a standard spec, you know, that’s just, it outlines what needs to get built, and then you just go build that. I think that’s good enough for a lot of stuff. But when you’re shipping stuff to people, and it’s like a marketable aspect of your business and product, I think the docs-driven approach gives you one additional thing, which is when you ship something, it almost ships with the narrative of what you’re trying to put out there in the world. And that’s kind of the key thing. Like, you know, I think we’ve all gone out and built really cool things that we think is really cool. Then, when we go and try to, like, tell people about it, we have to go through this, like, tough experience of, like, you know, the narrative piece, the storytelling. Oh, this is the problem you have, and like, you tried this, and you tried that. It didn’t work. But if you do it this way, it will work better. And let me tell you why. Like, you have to kind of pull all that together, and that’s how people align with you and adopt what you’re building. And it actually allows you to be a little bit more opinionated when you have a narrative piece to things. And when it comes to developer tooling, opinions are gold because most people don’t have opinions. And when you’re someone that actually has one, and it’s rooted in something, you know, tangible and material, people will opt into that. They will anchor themselves to that. And more importantly, if you’re trying to get adoption and awareness, they will almost champion what you’re doing because now your opinion is their opinion, and they will fight the good fight for you every day. They’ll become kind of a catalyst for your growth.
[00:17:58] Itamar: Yeah. I’m gonna give my try, like, predicting the very near future of how we’re gonna go with, so I’ll use the old model view control. I think, like, the view is the docs. The control is the code and the LLMs. Like, today, like, it’s very few. It’s like almost every software is code and LLM calls. That’s the code. That’s control. And the view is the docs. And I think there will be a representation of the model, which is kind of intermediate, which is mostly actually for the agents themselves. I’ll explain. Sometimes in the docs, yes, there is, like, here is what APIs look like, and here is how, for example, the contract or here are the capabilities, here are the behaviors, but if you, like, put too much information there, it’s too much for humans. Like, actually, docs need to be the essence of what you need to make it work. But the LLMs, actually, the more, the merrier. Like, you don’t wanna push everything in the context. It’s a different issue, you know, what do you know how to fetch to put in the context? But you need all of these, like, structural elements to be quite established, so it could run with really high quality. So, I actually believe that the future, let’s call it just for the sake of, you know, shameless plug, Coda wiki, how it looks like, it’s basically okay, you have your code base, and your usage of whatever libraries and AI, etc. Above that, there is, like, a model of, like, all the repo and their connections and the contracts and the model inside and the full graph, which is built for the agent. And every time there is a change, a code review agent can tell you, “Oh, you did a change according to your intent or not.” And above that, it’s like a document that is being built with a human interface in the…
[00:19:52] Amir: That’s right. That’s right.
[00:19:53] Itamar: In the app. And I think it’s like I wish to think that we are humans better at some things, but we also need to admit that we’re worse at other things. Like, you don’t go to a library. You go to Google. It’s better than you searching. So, the same thing here, we want to exploit the fact that models can absorb so much context in seconds. So, you want to put that information for them. But for us, we need to have, like, a searchable, easy way. So, I do think, like, docs as docs-driven development is gonna be the future. Underneath, under the hood, there’s actually a full spec that is actually written for the agent, something like that.
[00:20:27] Amir: I totally agree with that. We need, like, a distilled view always for the human. And I’ve built things where I ask the models to go, like, build me great docs, you know. Please spit it out. Fancy. And it never really hits the mark. And I think one of the key essences of that is that, yes, it knows the features you built, it knows kind of, like, why this project exists and what the general goal is, but it never really kind of pins down the story of it all. And if you start with that and then you feed them all, you start your project with that story component, I figured that the models would, like, keep remembering that. So, every time you try to, like, discuss something new with it, it will bring that up, and it will try to align itself to that. And so it kind of goes back to, like, the, you know, don’t burden the human you just mentioned. Like, that’s the critical part. I’ve had to, like, create docs for me, and I’m like, oh, I’m never gonna read this. I’m gonna just hit the ask assistant button on the docs, and I’m just gonna have it distill for me. And maybe that’s the answer, frankly. Maybe that’s what we should aim for. But I think even us as humans being able to distill, you know, the narrative of what we want the world to think about what we’ve built is very important to spectrum and development or just, you know, agentic engineering in general, right? I think that’s gonna be the core of it. Because if you think about it, when you’re onboarding a new engineer or a PM or something like that, they go through all sorts of hurdles to figure out what this company is and what this company is about and who are these people about, but when we talk to models, I feel like we don’t do that upfront work for a lot of things, right? We’re just so specific about, like, build me a magical product, you know, and there’s just so much about the essence of maybe a company that gets lost in all of that. But there’s definitely something to, like, these things can’t burden the humans, you know, in the way they output stuff.
[00:22:18] Itamar: I wanna zoom out and ask you two, each one of you, a question. Zooming out, think about director of engineering, principal engineer, CTO, I know it’s different roles, what do you feel from the field that there are challenges around quality? I think you’re probably talking as an investor, etc, and Nnenna, like, all day long, like, you’re in with the community. Like, today, one or top three things that those managers are thinking about quality in the age of AI, in the age of vibe coding turning into vibe engineering, or grounded engineering, like, how do you think about that?
[00:22:58] Amir: Yeah. So, I think we’ve kind of crossed the chasm of people actually using these tools, so we don’t have to talk about that. I feel like in the last year or two, it’s been like, will these things get adopted at all? So, we can go past that. So, now we’re fully in the mode of, like, how do we embed this, engrain this into everything. And I think when it comes to, you know, engineering leaders, what I’m seeing is everyone’s thinking about, and this is the problem I’m personally fascinated by, like, kind of eternally, which is software factories. How do we, kind of, build the machine that builds the machine? And in solving that problem, you have to kind of deduce the whole software life cycle into these key areas, and you have to rethink all of them because you probably are doing some facet of it already, and now you have to map it into your company’s factory for building software. So, maybe we can break that down. So everyone’s kind of approaching it in their own special way. Everyone’s journey is just to kind of, you know, attack this problem, as I’ve seen, it has just been, like, very different. But maybe let’s take it a little bit of a detour, bring it back to, like, the quality piece, and then we’ll bring it back to software factories. So, the thing was, in the past, we tried to tell people, please test your software. Please write as many tests as you possibly can. Please use CI and actually run it. You know, please do all that, and people did their best to whatever degree they could. Maybe on some days, it was a bad day. They didn’t really do it. But what happened was there was no super hardcore rigor around a lot of this stuff, even though the people that maybe did it a lot wanna say they were super rigorous, but I think the industry as a whole, we did not really do it rigorously. And I don’t think it’s people’s fault necessarily. I think we were asking too much of people. We were saying, “Hey, ship fast. It better be good, and you better maintain your test, and your CI better always be passing.” Right?
[00:24:52] Itamar: Yeah. You can’t get fast, cheap, and good, right?
[00:24:55] Amir: Yeah. You were asking a lot.
[00:24:57] Itamar: Until today. Until today.
[00:24:58] Amir: Yeah. Yeah. Until today. Yeah. We’re saying, “Hey, get it done by Friday.” Right? And it was a lot. It was a lot. It was like, you know, we’re telling the world to, like, please work out. Please eat your broccoli. Please have enough protein. Right? And it’s like, well, you know, I went hard last night, so I don’t know if I can. And it was a tall order. And a key portion of that was just people not writing tests. And I think the industry as a whole didn’t really know how to do it. We just told engineers, “Write tests. Good luck.” And there was not a whole lot of education around doing that kind of stuff. Like, it’s definitely not taught in computer science programs, right? Like, and probably it only gets brought up so they grade your test to make sure you fulfilled, like, you know, whatever the professor wanted. So, there was a lack of education in that part. But, you know, we pushed it on the industry anyway, and people did their best. And, you know, billions, you know, lines of code of test were written. And the thing was, some people ran them. And I think over time, many people just stopped losing, stopped trusting it, they lost confidence in it because it just became this huge bottleneck for a lot of folks. You know, many were happy enough. They got by. Their CI runtime ballooned a lot because, hey, I’m waiting for, like, the billing to test to run.
[00:26:16] Itamar: I hate that.
[00:26:17] Amir: I know. I know. It was a huge bottleneck and drag on everything. But we did it anyway because it was the right thing to do, and some of us tried to take a lot of shortcuts. So, you know, and then we were telling people, okay, you’ve lost confidence in your test suite, you know, what else are you hanging on to? Well, it’s like, well, I have decent code coverage. You know? Like, well, we did some visual regression testing, so, like, that’s decent. So, people try to use these other kinds of tangential things to try to enforce quality, but we never really want it as an industry. Like, we never really nailed that. And because we were in the software world, it was very easy to sidestep it. Like, I started my career in the hardware world, where it’s like, that is life, you know, like, you cannot take shortcuts or, like, people thought. But in software, we’re like, hey, I’m just shipping a wine app. You know? Like, maybe it’s not that important that the test passed today. So, but now, where we’re now, like, have this opportunity to think about what a software factory is and a core component of any factory is, like, quality assurance. And now we get to rethink this stuff. And let’s break that down. I think code review, of course, I think is the first line of defense for everything. We’re slinging code. So, yes, we should have reasonable code review processes in place, and I found them to be really useful in my own workflows. I mean, it’s crazy how much stuff it catches that’s actually real, too. And I run them both in the kind of my inner loop workflow and my outer loop workflow as well. I try to run as much as I can in the inner loop before I kinda get it out into some PR or something. That was one piece. And then so the next level up from that was, like, let’s generate tests. Oh, I don’t have to write tests anymore because the AI can write tests? And so right now, AI is slinging code on the test front. It’s creating a bunch of tests, and I feel great as a developer. I’m like, wow, I could have never written all these things myself. I couldn’t even come up with all the test cases, no matter how hard I tried, right? That was, like, one of the key issues. But, hey, at least I have a little bit more confidence, and my AI is definitely soothing me in that department. I don’t know how good it is yet, but it’s definitely doing something. And the other part of that, which is kind of, like, the thing we all ignore in this industry as much as possible until something breaks, is security. Most developers have no clue how to manage security at all. Like, they’re never really taught what that means. And I think now, it’s kind of in complement to code review, we’re also getting better security reviews, and these models are, of course, getting much better at that. So, if you kind of deduce all of that, like, we have kind of these building blocks for building a software factory. We have coreview systems. We have kind of improved ways of writing tests and hopefully validating them, and now we have better security tests as well. And I’ve kind of invested into all these buckets as well as part of my strategy. So, I’ll pause there for the moment. Yeah.
[00:29:11] Itamar: Yeah. Sure. I love to dig in more on the AI factory, but Nnenna, like, I’d love to hear your thoughts. What are these things at the top, one, two, three, like, like, challenges around software development, specifically, could be around quality for CTOs, technical leaders, etc?
[00:29:26] Nnenna: Yeah. I mean, a few things that I’m hearing engineering leaders complain about or maybe more so they’re worried about and wanna prepare and solve for for what they think will happen at the rate that as a result of the way in which the life cycle is changing and the rate at which code is being produced and needs to still get through the pipeline to reap the benefits of the speed from, like, cogeneration. I think it’s about how do you actually consistently control and enforce that quality, and it starts to feel like there needs to be more determinism or at least the feeling of deterministic quality gates or tools that are being leveraged. And there’s also a fear around, I think, the visibility. There’s about proof that all of the goals around actually integrating AI are being paid off in some way, and that they can actually show that. And so, then that kind of begs the question, or that reintroduces, well, like, is the code looking and functioning the way that it’s supposed to or the way that we intend, while we also reap the gains of, like, increased developer productivity. So, proof and yeah, those are some of the things I’m seeing.
[00:30:47] Itamar: I remembered two stories I heard from customers or people out there at IX. One case, like, I heard one person saying, like, they are now generating 30x more code than they had, like, the same amount of time last year. It sounded to me a lot like, but actually feasible, looking at myself as a CEO, how much code I can generate these days, despite being the CEO, and it felt feasible. But after, like, a good glass of wine, it was exposed that actually the amount of code that shipped to production was 30% more. So, what’s happening in between, like, 30x more generated to 30% more being shipped is a lot about, like, throwing PRs to the garbage, not passing through gates, and all that needing to be refactored or whatever. And that’s actually related to the second point. Like, I felt that you shared that. It’s another thing that I’m hearing more from customers and from the market is that maybe it’s not the same engineering technology and user interface between what helps us to generate code, and what helps us to verify code. Like, the whole idea, the whole reason that we are capable to do cogeneration because these LLMs are statistical creatures that, to some extent, work like the human brain and enable that creativity, and solving, and understanding, and reasoning and all that. But then on the gating, on the gateway and gatekeeping, maybe we need, like, more rigorous. And what you’re saying is that maybe more determinism. So, I guess it won’t be fully deterministic because otherwise, like, I think these two won’t be able to catch up. But probably, like, what I’m hearing from the market is that there is a huge, like, demand for I don’t want a code review or code governance or whatever to be exactly the same as coding. I need it to be more deterministic, more rigorous, more report-building, and all that. That’s definitely, like, resonating with what I’m hearing. And maybe that’s back to you, Amir, like, on the, yeah, a factory. How do you imagine, like, a factory coming to life, like, let’s say, 2027?
[00:33:03] Amir: Yeah. Yeah. Well, I think one thing we get that’s pretty nice now, which we couldn’t have before, is feedback systems as part of all this. We’ve never really had that in software development. You know, everything was pretty linear. You build something. You pass it through some linear process, and then, hopefully, it goes through all your checks. But, you know, as you were saying, hey, maybe the code review experience is gonna look very different from the coding agent’s experience, right? Like, that part’s important because what I find is that my code review output should feed back into how I do coding in the first place. And, like, it should help me actually write those better specs because, hey, I just wrote the spec. I implemented it. Now I’m bringing code review on it, but, wow, it is catching a lot of things. At some point, I’m like, okay, it’s catching so much that maybe there was something wrong upstream. Like, maybe I didn’t spec thing. Maybe I didn’t think about this problem long enough, and the code review is really highlighting that for me. I mean, it does that a lot. It surfaces kinds of scenarios and cases that I didn’t think about at all. No matter how much of a conversation I had with the agent, and no matter how much I asked the agent to go do this particular research or that particular research. So, I think feedback systems are gonna be critical to that, both at the code review point and, of course, also at the security layer as well. And I think there has to be a good governance story of how all these things merge into the broader processes of spec generation and on all the guardrails as well.
[00:34:39] Itamar: It’s interesting to double click on why did you know, already, like, when the code-generating coding agents are working, to begin with, they didn’t generate the right code or a high-quality code? I think, like, maybe we can, you know, put the framework of there is the context and the harnessing of AI in that context, right? And I think, like, maybe the coding agents are missing because, a, they have a different context, and, b, maybe they’re harnessing that context differently. And what I’m hearing from you is that, okay, if this is the case, then, a, maybe I need to close the loop between these tools and make it one big machine, yin and yang. And the second thing, I think I heard you discuss a little bit in other blogs and podcasts about building your own context because of this. Like, you want to build your own context to feed the coding agent to begin with, to do the right thing. Like, what do you think about that?
[00:35:36] Amir: Yeah. I mean, its context management has been helpful. It hasn’t fully solved, I think, some of these issues because I still see the agents definitely getting tunnel vision sometimes, you know, maybe I don’t know if it’s the models or the harness. I think it’s improved even in the last month, frankly. But I’ve definitely seen it go down paths where I’m like, whoa, whoa, whoa, what are you doing now? Like, I thought we discussed this. You know, I thought we were pretty clear and aligned on, like, what to do here. So, yeah, happy to share with you how I’ve kind of managed it.
[00:36:10] Itamar: Yeah. Please do.
[00:36:12] Amir: So, I mean, I have kind of, like, this broad system and setup I’ve built to help me with this, where no matter what I’m trying to build, I kind of have a method to break it down into kind of this phased implementation rollout. And each phase has all these gates. It has entry gates, and it has exit gates. And then for each gate, it also has its own, like, way to test each of them as it’s going as well. So, I can’t go from phase one to phase two if the entry and exit gates are not set up properly. And I think more and more people are kind of adopting this method. Mine, when I built it, it was kind of out of sheer necessity because I was trying to have it do more long-horizon work, and I don’t think you can do it without kind of this kind of structure. It is super annoying, though. Like, it’s rigorous, but yeah, it definitely, like, it forces the agent to work much longer than maybe it probably needs sometimes. But I’m just trying to get a confident output on that part. And by the way, some of the gating for an implementation phase to be done is heavily reliant on a kind of custom code review. So, like, each phase has its own review instructions as part of it. So, everything is bespoke to the particular thing you’re building. It’s not like I’m doing a generic code review on that process. So, it’s worked out really well, and I think some of these coding harnesses are now having kind of this goal mode now. I mean, like, I’m pretty Codex-pilled frankly at this point, right? Yeah. All love for Claude and all that. But it’s been great, and I’m getting, I think, that’s driving a lot better results.
[00:37:54] Itamar: Yeah. I’m hearing a lot of taste matters, what you’re saying. What about you, Nnenna?
[00:37:59] Nnenna: Yeah. No. First of all, definitely a Codex fan as well. I use goal mode all the time, but I did wanna add to the whole context management thing, but that is actually another one of the biggest things that I hear engineering leaders are kind of struggling with. One of their pain points is that they understand the value of context on different levels, for like the individual developer completing a task, to, you know, like a whole team being able to access certain types of context that is working in their favor across the software development life cycle. And it’s like, but how do you actually implement this? Like, what does that actually look like? If that’s architecturally, if that’s, you know, a product to leverage, how do you actually integrate this to become successful? So, that’s an interesting point.
[00:38:47] Itamar: So, we’re getting toward the end, and I heard a lot about, like, managing your context, building those loops, creating that governance and control observability about the AI factory. And I have maybe a hard question for you. Build versus buy. Where do you think you’re going with it? As an investor? Like, right, it really matters. And even in any, like, would you build more of your tools? Are you using tools out there? Or what other people are doing these days?
[00:39:17] Nnenna: I think one of the metrics to determine if you wanna build or buy should be, if I build this thing, what amount of, like, engineering resources would I be allocating to building this thing, and what is the opportunity cost of that? Am I taking them away from contributing to the core product? I think that is a major question that could actually really help to get you closer to an answer.
[00:39:41] Amir: Yeah. I mean, focus on the on your core business has always kind of been the key, you know, argument against building versus buying. You think a few things are different now. One, you know, in the past, like, you kind of saw this in the zero interest rate phenomena era, where people overhired and just had too much money. So, a lot of companies, actually, you know, of course, people were sliding their corporate credit cards, buying whatever tool or option was out there. But what we saw was a lot of companies just had so many engineers that they had so much free time, apparently, that they would they wanted to just build everything. And, of course, when it comes to developer-focused things, engineers are like, what do you mean? I can just build tools for myself and get paid for it. Like, that sounds awesome. So, you saw a lot of that. But then, the second the zero rate phenomenon died down, all that went away. All that went away. Because now it’s like, oh, now you gotta get rid of all these engineers that you probably didn’t need on hand because we’re overhired, and now no one’s there to maintain any of the stuff they made that we supposedly rely on. So, this kind of burden of maintenance always remains when you’re trying to build something. And, you know, do you wanna burn your tokens on that? That’s the kind of key questions that think leaders have to ask. Like, are you the best set of people to maintain that thing? And, look, some companies, they’re just culturally set up to build. You know, if you look at companies like Uber or, you know, of that nature, they always love building things from the ground up, and I get that to some degree. But I think for the vast portions of companies out there, like, they should be buying instead of building, and just kind of focus on the core task at hand. Regardless of how enticing it is to kind of, you know, whip it up yourself, you know, in a week or so, it is very attractive, but I think people should be pretty reluctant to do that.
[00:41:29] Itamar: Yeah. What I’m seeing in the market is that there’s a new thing. Instead of build versus buy, there’s build to buy. Like, you are enticed to build something in a week. And if it does work, you’re actually happy because, like you mentioned, like, you wanna manage your context. You wanna build those loops. You wanna, you know, make the product, the solution you’re building in your taste. It matters. And if you’re if you could do that with one or two, three people for a week or two and then they probably need to keep pushing on it and maintaining it, why not? But the worst case scenario, which I’m not saying, like, a one week or two week or a month of two developers is allowed, but the worst case scenario when you’re going to buy, you know your shit. Sorry for my French. Like, you know what to ask, what you didn’t manage to do, what you have to have in the solution, but you couldn’t do it yourself because there’s a lot of tokens to be burned in engineering and harnessing and and flavor and taste that was put together. So, that’s what I feel.
[00:42:23] Amir: Like, one thing about this is that, and you probably see this all the time. I’m sure that companies building their own code review tools is probably one of the most enticing things they would want to do. So, you said build to buy. And I think sometimes you have to build it to appreciate the hell why you need to maybe buy it. And, you know, I’ve tried actually building my own code review agents as well, and I’ve been in the long run. I’m sorry. I’m sorry. But I did it for the sake of kind of intellectual curiosity, of, like, oh, like, what you have to get right to this. And I think you very quickly realize why, you know, and I still don’t know your secret sauce on this front of, like, how you’re able to surface certain, you know, bugs that, you know, no matter how hard I tried, some of these my micro review agents couldn’t find. And I do that sometimes as an investor to, you know, because you said build to buy. Sometimes, what I do is I try to build so I can figure out if I should invest.
[00:43:21] Itamar: Yeah. Build to invest.
[00:43:22] Amir: Build to invest. Yeah. Because if I can cook it up super quickly and it’s pretty good, then, yeah, maybe that’s not a hard enough problem to solve for the market, and I shouldn’t invest in that. But when I hit, when I try to go after some of these problems and you’re quickly humbled by the difficulty of it, then you’re like, okay, yeah, people should be definitely, you know, seeing what’s out there in the market.
[00:43:47] Itamar: Yeah.
[00:43:48] Nnenna: I would just like to add about the Uber piece because you said that they are very comfortable with building. I wrote an article about them a couple of months ago because they had built their own AI code review tool. And because at that point, it was the fall of 2025, the capabilities of the tools that they were testing to potentially buy were just not there to match the constraints that they had for, you know, enterprise engineering work. So, they went, and they built their own. And I do want, like, I want an update so badly. They published an article about it then. I want an update on how that is going now, with just how much that has changed in the market, in the space, with not only just like LLMs, but, you know, the features, capabilities, on-prem deployments that are available with these tools, and if that’s something they would continue investing in.
[00:44:37] Amir: Yeah. Yeah. So, if someone from Uber is listening, please follow me right now.
[00:44:41] Itamar: Yeah. So, just before we ask where to find you, etc, last one, what do you think the industry is not talking loudly enough about? Probably you’re thinking about, like, next year, that isn’t talked about right now. And, like, I don’t know. Like, are there areas of governance or future problems that, you know, like, what people are gonna be fired because of next year that they need to think about this year? Love to hear.
[00:45:07] Amir: Yeah. I mean, it’s almost like everything, right? Because I feel like we’re at the moment where we’re forced to rethink everything. I think one topic that I see now come up over and over again, and I’m glad we’re discussing it, is the stuff we started with at the beginning of our conversation, which is, like, what are the roles gonna look like long term? Some of us that are kind of at the edge of it all, you know, we think we have a decent idea of it. And some of you know, at startup player, we’re kind of building these companies from the ground up with that in mind. But I think the real test of all this and the real kind of societal level conversation we need to have is, like, what does, you know, an engineer’s role really look like moving forward? What are the expectations? How does that map to other, you know, departments in a company? What do other departments even look like? And these things are all overlapping, and I think, you know, as we have kind of industry leaders listening to our discussion here today, I think that’s the conversation they need to not be shy about, which I’ve seen people are pretty shy about, because it does touch a nerve on a lot of things. And I think we can figure it out, but we need to have more of a discourse around it. And I see many podcasts out there that are kind of poking at the topic, and we should do more and more of that.
[00:46:21] Itamar: Yeah. That’s why we’re here. So, where can we find you? And, Nnenna, where can we find you?
[00:46:26] Amir: You can find me on X. Look up my name, Amir, and I have a very easy email. It’s [email protected]. People are pretty surprised by the domain name, so you can begin testing.
[00:46:39] Nnenna: Yeah. When I saw that, I was like, how? Like, I just need to know how you acquired that.
[00:46:44] Amir: But it was available. I was the first owner.
[00:46:46] Nnenna: Oh, Amazing.
[00:46:48] Amir: Yeah. Yeah.
[00:46:49] Itamar: Nnenna, where to reach you out?
[00:46:50] Nnenna: I am definitely always on X. That’s where I yap the most. So, @nnennahacks, my first name and hacks is where I’m always talking about these exact conversations that we’re having here. So, glad that we could do this.
[00:47:03] Itamar: Awesome. Thank you so much, Amir. That was really, really great.
[00:47:06] Amir: Awesome. Awesome. Thank you so much for having me.
[00:47:08] Outro: If today’s conversation challenged how you think about AI and code quality, that’s the point. At Qodo, we believe that independent context-aware code review with rules as guardrails is how engineering teams maintain standards at scale. If you’re leading an enterprise team and want to see how intelligent AI code review can reinforce governance, visibility, and accountability in your workflow, visit qodo.ai to learn how we help teams turn AI productivity into production-ready quality. And if you enjoyed this episode, subscribe, share it with your engineering leadership circle, and leave us a review. Until next time, keep human in the loop. And keep shipping.
About the hosts