What Every Software Factory Needs That No Model Will Provide
I spend most of my weeks talking with CTOs and engineering leaders, from consumer apps, and e-commerce, via chip makers and automotive to financial institutions, defense industries and global banks. For months, every one of those conversations has started in the same place: the engineering team, supported by their CTO/CIOs, are eager to unlock their software factory potential with an agentic SDLC. Each of these are in different stages in their journey of building their software factory, and none of them need me to convince them it is the right move.
The conversation gets interesting when we reach the part they have underinvested in, although realizing it is a key and perhaps the hardest piece of that factory, and it happens to be a piece no model (current or future), agent or straightforward harness can supply on its own.
The software factory, three years from now
Within three to five years, I expect more than 90% of the companies in the world to run a software factory.

Each one will come together the way enterprises built on cloud: a handful of components engineered into a system with the company’s own logic, shaped around its own specifications. Every software factory will run two kinds of workloads: generative workloads that build, and quality workloads that challenge, verify, and govern what gets built and how. The two have to run in tandem, because a factory that builds fast without checking its own output only produces problems faster.
With every factory built from its own combination of pieces, the CTO’s job becomes designing it: deciding what to build, what to buy, and then ensuring those pieces work well together.
Where today’s factories break
So far, most of the investment has gone into the first workload. Enterprises are spending heavily on coding agents, generative harnesses, and models, thus generation got dramatically faster. The trouble is that every time one development part gets automated, the bottleneck moves to the next one. Once generation sped up, review and quality became the constraint. The factory keeps getting faster at the step it just “fixed” or automated, while quality stays invisible until something breaks in production.
Part of the problem is how generative agents deliver their work when operating without a dedicated quality layer. A coding agent takes a task and returns a work-package: pull requests spanning multiple components and even across several repositories, produced faster than any human could have produced, and now we humans need to review these. If you are a code owner at the agentic-SDLC phase, then you are probably painfully spending significant time figuring out which work-packages exist (and why!?), what they do, how their PRs relate, what their blast radius is, and what technical debt they introduce. Many report that this effort is frustrating and exhausting. On many of the features I have built with agents, pushing the stack of PRs to production took more time than developing them, and I hear the same story from nearly every team.
Analysts are seeing the same risk. Gartner projects that by 2028, half of the enterprises that adopt autonomous coding agents without real-time guardrails will decommission or severely restrict them after a quality or security incident.¹ That number stands out to me because it describes a design problem. An enterprise that has to pull its agents back after an incident loses the speed the factory was built to deliver. Avoiding that outcome means designing quality into the factory from day one, and that starts with the people running it.
Sources
¹ Gartner, How to Ensure Quality in AI-Generated Code, Joachim Herschmann, 21 July 2026. GARTNER is a registered trademark and service mark of Gartner, Inc. and/or its affiliates in the U.S. and internationally and is used herein with permission. All rights reserved.
Building quality into every step
As more of the SDLC runs on automation, the engineer’s role changes with it. Engineers spend less time writing each line and more time steering and governing how the agents produce code, which changes their needs from the quality tools and workloads dramatically. They will want to steer it, control it, and observe it. Above all, they need to trust it, because they can no longer read every line it approves. That is why the quality part of the factory has to be independent and dedicated to that job. The system that generates code should not be the system that decides whether it ships.
Quality also has to be present at each step of the factory, from the moment an agent starts writing code to the moment a change merges. Today most teams rely on a single checkpoint at the very end, and that checkpoint works a lot like an aspirin. It eases the pain of a pile of weak PRs, and the next morning the pile is back, because nothing upstream changed. The goal is a factory that produces strong PRs to begin with, and that takes two roles working together. The first is the coding agent’s best friend, which injects context and adversarial thinking while code is being written so the work arrives clean. The second is the gatekeeper, the high-precision review that serves as the last check before anything ships.
The gatekeeper’s job is changing too. When fifteen related PRs arrive as one package, reviewing them one at a time misses how they fit together. The review that holds up looks at the whole package, including its relationships, blast radius, and review order, with an engineer working alongside agents. Because that takes focused time, whoever picks up a package claims it, so the team and the agents can see who is on it.
The factory runs on wisdom
Both of those roles depend on something I believe most leaders are missing. The most critical element of the agentic SDLC is the organization’s tribal knowledge, codified so agents can use it. The LLM is one component. No LLM on its own is going to build SAP or Salesforce. Software like that is built and maintained on years of judgment about how the system works, what the team accepts, and what it rejects.
That judgment lives in senior engineers’ heads, in review threads, and in decisions nobody wrote down. Capturing it in a form agents can use is an immense engineering problem, and it takes dozens of dedicated engineers across research, algorithms, and several engineering disciplines.
We call the result a wisdom base. A database stores facts, and a knowledge base connects those facts into a graph. A wisdom base goes a step further and adds experience to that graph: which rules developers trust, measured by how often their findings get accepted; which suggestions, from Qodo and from other developers, a team takes or rejects; and how the software itself changes, PR by PR.
The agent, the harness, and the model will all turn over at least once more, while the wisdom base keeps compounding through every one of those changes. That is what makes it the moat.
Few enterprises will build it themselves, for the same reason almost none train their own foundation model. The economics mirror early cloud, when Datadog argued that a company spending $100K on cloud should spend around $15K on observability to keep that $100K from turning into outages and rework. The software factory works the same way: for every dollar spent on agents that generate code, a smaller share should go to the layer that keeps their output trustworthy.
Seeing the whole factory
This is why I think about Qodo as a Roman building. Three pillars, the coding agent’s best friend, the gatekeeper, and the work-packages gateway, all stand on the wisdom base. On top sits the roof, which gives leaders visibility and control over pillars.
The roof is the part that surprised me most. I mean personally when using Qodo. A few weeks ago I opened our own Software Map in the morning and saw that part of our system had changed overnight. The team had stood up a new environment to test our on-prem deployment, and there was no way I could have learned that without Qodo. A software factory needs that view, updated as fast as agents change the code.

How Qodo 3.0 builds toward this vision
Qodo 3.0, releasing today, adds to every part of that building, starting where today’s software factories break.
PR Triage, launching in Research Preview, goes after the bottleneck around agent output. It shows a tech lead every work package across repos and git providers, titled by the feature being built, with its PRs, blast radius, review difficulty, waiting time, and the order to review them.
For the best friend pillar, your team’s rules and skills, mined from your own codebase and review history, reach coding agents through the Agentic Toolbox while they write. For the gatekeeper, review now runs as a swarm of specialized agents working each change together.
For the wisdom base, PR Insights (releasing in Research Preview) shows what Qodo has learned from your team’s review decisions, so your team can see it and correct it.
For the roof, the Software Map shows what every change touches, and native quality metrics give leaders a number to bring to the board. All of it runs where enterprise code lives, including Gerrit, on-prem and air-gapped environments, and open source models like NVIDIA Nemotron.
Build it on wisdom
Every enterprise I talk to will build its own software factory, and I’m sure that is the right call. The hardest part to build is the wisdom that makes the quality workflow trustworthy. That is what we have spent the last two years building.
Build your factory on a foundation that already knows how to judge the work.
GARTNER is a registered trademark and service mark of Gartner, Inc. and/or its affiliates in the U.S. and internationally and is used herein with permission. All rights reserved.