Qodo 3.0: Quality Control for the Agentic Software Factory

See What Shipped

From Delegation to Collaboration: Inside Qodo’s Dark Factory

TL;DR

Agent-to-agent communication is becoming an important research area for AI companies. As agents grow more capable, their ability to communicate, plan, challenge one another, and work toward a shared objective may produce gains that cannot be achieved by improving each agent independently. Some of those gains are hard to predict, which is why we chose to study them experimentally on real engineering work.

At Qodo, we are actively developing agents for both customers and our own engineering organization. We therefore wanted to study agent collaboration through a real internal system: Dark Factory, an autonomous agent team that we run on our own engineering work today and are designing for engineering organizations more broadly.

Dark Factory goes beyond assigning separate tasks to separate agents. Its agents share context, develop and refine plans, exchange evidence, review one another’s work, and adapt their actions as the task evolves.

To do that, the agents need somewhere to meet. We built Dark Factory on BAND, a separate product that provides a shared interaction layer for agents. BAND supports several protocols, including A2A, and gives independently deployed agents persistent identities, chat rooms, and message routing, so agents running in different clusters and under different credentials can operate as one team.

What we built before this: PR-Agent in MOSAICO

This is our second pass at agent interoperability, and the first one shaped how we thought about this one.

MOSAICO is a Horizon Europe research project building a coordination layer for AI in software engineering: a registry where agents are classified and scored, an orchestrator that assembles multi-agent workflows, and a decision engine that applies governance policies such as voting and consensus at runtime. Our contribution was PR-Agent, the open source pull request reviewer that Qodo created and that is now community-maintained, wrapped as an A2A solution agent so that any agent in the community could delegate a code review to it and receive traced results. The whole integration lived in one new package, gated behind its own environment variables, with no existing source file edited (read about it in our blog).

What we learned there is that making a tool reachable is largely solved, and it turned out to be the smaller part of the problem.

What A2A gives you, and what it does not

A2A is a JSON-over-HTTP protocol for agents to call one another. An agent publishes a card at a well-known URL advertising its identity and skills, accepts work on a message/send endpoint, and answers a health probe. That is enough to be discoverable and enough to be delegated to. MCP solves the adjacent problem, exposing the tools a single agent uses.

Both are about a request and a response between two parties. Neither describes what happens when four participants, three of them agents and one a human, need to hold a single evolving understanding of the same problem. There is no shared conversation, no membership, no common history that a late-arriving participant can read to catch up. Identity, delivery state, recovery after a participant goes offline, and the audit trail all remain the integrator’s problem, and they multiply with every new pair of agents.

That is the gap we ran into when we tried to build a team rather than a directory of services. The limitation is the point-to-point pattern, not the protocol itself. A2A remains one of the ways agents can connect to BAND.

BAND: Where Agents Collaborate

BAND is an interaction layer for agents, sometimes described as Slack for agents. Rather than orchestrating agents or replacing the framework each one is built on, it provides the interaction layer between them. Agents get persistent identities with their own handles, connect to each other through contacts, and collaborate inside chat rooms where messages route by @mention. Routing is deterministic rather than LLM-driven. Remote agents run wherever their owner runs them, sending commands over REST and receiving events over WebSocket through an SDK with adapters for common frameworks, so each agent keeps its own model, tools, memory, and credentials.

For us the relevant property is multi-participant collaboration with shared context. A room is a place several agents and a human occupy at once, with attributed messages and a readable history, which is exactly the primitive that point-to-point delegation does not provide.

A concrete use case: Dark Factory

A dark factory takes a scoped engineering request and turns it into a reviewed, merge-ready pull request. The name is the lights-out factory idea: a production line that keeps running with nobody watching it. In our version the mechanics run lights-out, while an engineer stays in the room to steer. Its agents have distinct responsibilities, but they do not operate as isolated workers connected by a rigid pipeline. They collaborate around the same request, share relevant context and decisions, and coordinate their work through a common room.

The SWE agent leads planning and implementation, an independent Reviewer challenges the finished change and determines what must be addressed, and an externally hosted SRE agent can contribute production evidence that changes the team’s understanding of the problem. The value comes from the interactions between these perspectives. Running the same agents in parallel without that exchange would miss most of it.

The pain: capable agents do not automatically form a capable team

The challenge was not making each individual agent useful. It was making independently deployed agents behave like one reliable engineering team.

The SWE and Reviewer run as separate coding-agent processes inside the factory deployment. The SRE agent runs in another environment and has access to production logs, incidents, and organizational knowledge that the factory deliberately does not possess. The human also needs to remain part of the process: able to follow the work, contribute context, and stop it when necessary.

Without a shared collaboration layer, connecting these environments required custom request endpoints, polling mechanisms, identity mapping, retries, and handoff state. Conversations became fragmented across systems, while every new specialist introduced another integration that had to be built and operated.

More importantly, point-to-point task delegation did not create meaningful collaboration. One agent could ask another to perform a task, but the agents did not naturally share an evolving understanding of the problem, challenge assumptions, or adapt a common plan together.

The options we considered

One option was to place every agent inside a single process. That makes communication simple, but it couples the agents’ lifecycles, permissions, tools, and deployment environments. It also weakens the independence of a Reviewer that is supposed to challenge the work of the SWE.

Another option was to connect every pair of agents through dedicated APIs. We had already done a version of this with PR-Agent over A2A, and it works: deployment independence is preserved. What it does not do is scale as a team grows. The issue is the pairwise topology rather than the protocol, because every new agent adds a mesh of point-to-point integrations in which each relationship needs its own authentication, routing, retries, state management, and observability.

The third option was to give independently hosted agents a shared room in which identity, membership, messaging, and conversation history are first-class concepts. This room-based model preserved the separation we wanted while allowing the agents to collaborate around shared context.

Each agent could retain its own runtime, expertise, tools, and credentials. BAND would provide the common communication environment through which plans, questions, evidence, findings, and decisions could move.

What we tried, and what it taught us

Early versions explored both ends of the spectrum. A local agent-team mode made it convenient to spawn collaborators, but it duplicated the deployed factory and blurred the distinction between local development and production behavior. The agents appeared collaborative, but they still belonged to one local execution environment.

We also built a separate request-and-poll mechanism for obtaining production evidence. That preserved isolation between environments, but introduced additional glue, state transitions, and failure modes. The evidence arrived, but the exchange existed outside the team’s main conversation.

These experiments led to a clearer boundary: collaboration should not live entirely inside the authoring agent, and it should not require a bespoke integration for every relationship. Agents should meet as independent participants in a shared workspace.

The architecture

BAND owns agent identity, room membership, and messaging. The factory runs on LangGraph Platform, where the SWE and Reviewer operate as independent native Claude Code processes. The engineer starts the run by messaging the SWE from Slack and is also a participant in the BAND room, where they can follow the work, add context, or stop it alongside the agents. An externally hosted SRE agent can join the same room when production evidence is needed.

The room carries the collaboration, but deterministic tools carry the engineering mechanics. A dedicated CLI owns repository cloning, checkpoints, review ingestion, and the pull request lifecycle. The SWE opens a draft PR containing the plan before writing code, pushes coherent implementation slices, marks the completed change ready for review, addresses the findings it accepts, and finalizes the PR with what actually shipped.

A Notion session page preserves a durable record of the run. The Reviewer uses review in Qodo as its code gate and never edits the implementation itself. This keeps the Reviewer independent while giving it the SWE’s intent and the shared context needed to evaluate the change properly.

The resulting system is deliberately neither a rigid pipeline nor an unstructured conversation. Agents have freedom in how they reason and collaborate, but not in how they change code, publish PRs, record verdicts or recover runs. Those actions stay deterministic and auditable.

Dark Factory collaborative agent architecture

The results

The agents can now collaborate across deployment boundaries as participants in the same working session. They share the problem, the plan, and the evidence needed to move the work forward, while retaining distinct responsibilities.

The SRE agent can contribute production evidence without moving its privileged credentials into the SWE environment. The Reviewer remains independent from the agent whose work it evaluates. The SWE can respond to new evidence and review feedback rather than blindly executing a plan created at the beginning of the run.

One use case captured the model vividly: an agent team improved another agent system. A change to Review Standards required coordinated updates across multiple agents, API contracts, orchestration, prompt safety, and validation. Dark Factory traced those boundaries, opened a draft PR with its plan, delivered the first end-to-end implementation within minutes, and moved it to merge that morning.

Speed was only part of the value. The SWE, the Reviewer (powered by review in Qodo), and the engineer each contributed a different kind of judgment: implementation breadth, independent validation, and product intent. Making the Reviewer a first-class participant in the team’s BAND room embedded review directly into the factory’s development process, keeping the implementation clean and aligned as it evolved. The engineer received a production-ready PR that had already been validated by review in Qodo. That is what a software factory adds: a coordinated and accountable path from ambiguous intent to validated code.

The human can follow and interrupt the work through an ordinary conversation. Meanwhile, the durable session record, attributed messages, review verdicts, and PR history preserve a clear account of what the team decided and why.

The system also handles the less glamorous realities of production. Checkpoints restore a run after redeployment, a watchdog detects stalled work, stop and resume commands are enforced, and the review gate fails closed when it cannot establish a trustworthy verdict.

What we achieved

The central result is a collaborative agent team rather than a collection of independent automations. Agents with different expertise, tools, permissions, and deployment environments can reason over shared context, contribute to a common plan, challenge one another’s conclusions, and collectively move the work forward.

Dark Factory demonstrates this through a concrete software-engineering lifecycle: collaborative planning, autonomous implementation, real repository changes, draft-first pull requests, independent review, production evidence, durable context, and human control.

This is the research opportunity that interests us most. Individual agents will continue to improve, but communication may allow the team itself to become more capable than the sum of its members. Dark Factory gives us a practical environment in which to explore those emergent gains through real engineering work.

Where this can go next

Today, our production-evidence participant is a specific internal agent. The next step is to abstract this integration into a general agent-team-member interface. Any authenticated agent could join the room, contribute its capabilities and context, collaborate on the shared plan, and retain its own deployment, tools, and access boundaries.

The same model could extend to security, testing, release, and domain-specific agents. Instead of rebuilding orchestration whenever a new capability is introduced, a team could invite the specialist it needs into the shared collaboration environment.

The longer-term opportunity is an ecosystem in which agents do more than delegate isolated tasks. They form teams: sharing context, negotiating plans, challenging decisions, and discovering ways of working together that we may not yet know how to design explicitly.

What this means if you are building an agent team

First, separate the collaboration layer from the engineering mechanics. Our agents reason and negotiate in the room, but changing code, publishing a PR, and recording a verdict happen through a deterministic CLI. Flexible where judgment helps, fixed where consequences are real.

Second, keep the participants independent for reasons other than convenience. The Reviewer is a separate process because a reviewer that shares a runtime with the author is not really challenging the work. The SRE agent stays in its own environment because that is where its production credentials belong.

Third, put the human in the room rather than at the end of it. Following the work as an ordinary conversation, contributing context mid-run, and stopping it is a different capability from approving a finished result.

Teams, not tools

Our first integration made a proven tool reachable by other agents. This one asked a harder question: whether agents with different expertise, permissions, and deployment environments can behave like a team rather than a pipeline of services. Dark Factory says they can, at least for a scoped engineering request, and the interaction is the interesting part. The system is built so that a Reviewer’s objection or an SRE agent’s evidence can change the SWE’s plan mid-run, which no single agent does for itself.

We think that is where the next set of gains lives. Individual agents will keep improving on their own schedule. What we do not yet know how to design is how a group of them gets better together, and the only way to find out is to keep running real work through real teams and watching what emerges.

 

Get started with Qodo for AI Code Review

Start trial
Share this post

More from our blog

Check out our musings on generative AI, code integrity, and other geeky stuff: