Qodo 3.0: Quality Control for the Agentic Software Factory

See What Shipped

Does Jev make AI code review more efficient?

An extra AI check earns its place by verifying behavior the existing workflow missed. I evaluated Jev, TypeSafe’s structured-judgment model, alongside selected automated checks and Qodo code review.

The improvement was marginal: one additional required test gap across 24 evaluation cases, with no additional confirmed defects. Jev also raised unsupported concerns in three cases, including one that changed a verification request into a code-fix request.

What did Jev add after Qodo?

Controlled Release Lab demonstrates guarded rollouts. Its gate should stop a rollout when one workspace’s queries become slow while aggregate latency stays healthy. Case c020 implemented that behavior correctly. Its submitted tests covered healthy behavior, aggregate failure, and invalid observations; they missed query-specific isolation.

The independent probe slowed selected workspace requests to 900ms while aggregate p95 remained 20ms. P95 is the 95th percentile of recorded latencies. Two assertions isolate the behavior:

expect(result.reasonCodes).toContain('QUERY_LATENCY_HOLD');
expect(result.reasonCodes).not.toContain('LATENCY_HOLD');

The first confirms the query gate held the rollout; the second excludes an aggregate hold as the explanation. A stronger test also detected a seeded defect that exempted workspace queries.

Qodo’s six findings addressed other problems; adjudication identified none as this required gap. The combined workflow requested it across three passes. Existing blockers left the top-level recommendation unchanged. Counting only recommendation changes would miss that contribution.

Jev judged submitted verification against an explicit obligation; my analysis connected its typed response to the missing check.

How did I evaluate what Jev added to code review?

The revised protocol compared selected checks, checks plus Qodo, checks plus Jev, and all three. Cases paired correct code with adequate tests, correct code with a missing test, and defective code missed by submitted tests. Independent executable checks established labels and stayed outside reviewer inputs.

There were six pilot and 24 evaluation cases. I pinned the Jev model jev-1.13.0 and reused complete Qodo CLI 1.0.3 reviews after identity checks. Qodo had repository access; Jev received fixed source and test evidence. The workflow comparison used different evidence access; it can’t establish a model ranking under equal conditions. Selected suites, typecheck, and build didn’t represent all CI.

Qodo identified eight seeded defects and 15 of 16 required gaps; both Jev workflows requested all 16. The original protocol supplied verification obligations in prompts. Jev returned choices and confidence, without a written reasoning trace. Adjudication connects those answers to evidence. They’re requests for known checks; they don’t establish unprompted discovery or general precision. Broader Qodo integration concerns remained in the review even when they fell outside the named requirement.

Evidence flows from source code and tests through automated checks, Qodo review, and Jev judgments to a maintainer recommendation.
How evidence reaches the recommendation.

What unnecessary work did Jev request?

Unsupported concerns affected c013, c014, and c028.

Case c028 required idle-client expiry to free released clients’ entries while preserving active concurrency permits. Its submitted test exercised more than 1,024 released clients over at least 55 seconds and retained four occupied permits. A separate saved 20-requests-per-second probe also passed.

In the combined workflow, Jev reported a violation and missing coverage in all three passes, with violation confidence crossing 0.8. Under the fixed policy, sufficiently confident violations request fixes; lower-confidence concerns or missing coverage request verification. Existing blockers remain.

TypeSafe derives confidence from the answer’s probability distribution. A confidence of 0.8 doesn’t establish 80% correctness or a safe change.

The recommendation policy preserves existing blockers and maps confident violations to fixes and coverage concerns to verification.
How judgments become maintainer actions.

Qodo’s verification request became a fix request unsupported within the named requirement. Fix requests rose from 18 to 21: two reclassified defects Qodo already identified, plus c028. None was an additional confirmed defect.

Final recommendations across checks alone, checks plus Qodo, checks plus Jev, and the combined workflow.
Recommendations across the four workflows.

Across 24 cases, Jev added one required gap, with zero additional confirmed defects, three unsupported concerns, and one ambiguous requirement concern, c023. These categories overlap with legitimate findings and shouldn’t be added into a success rate.

Across 24 cases, Jev added one required gap, zero confirmed defects, three unsupported concerns, and one ambiguous concern; categories overlap.
Useful additions and unnecessary requests.

What did additional verification cost?

The first evaluation couldn’t complete oversized query requests. The revision retained source relationships: the workspace baseline originates upstream and passes through several functions before reaching the gate. A final conditional alone would hide that propagation. Each packet retained complete selected declarations, relevant dependencies, and submitted tests.

Query cases used five packets per workflow per pass; others used one. Inputs were frozen before repeated collection. The primary collection completed 252 calls across three passes and used approximately 3.88 million reported input tokens, including pilots.

For the 24 evaluation cases, combined-workflow medians were 0.62 seconds for ordinary cases and 3.26 seconds summed over five HTTP requests for a query case pass. The latter’s median input was 102,626 tokens.

These recorded durations and usage exclude preparation, adjudication, human work, and reused Qodo time. Qodo token usage was unavailable. Developer-time savings weren’t measured.

Combined-workflow medians were 0.62 seconds for ordinary cases and 3.26 seconds across five query requests, with 102,626 median query input tokens.
Recorded latency and token usage.

Does this contribution justify adding Jev?

The evidence supports further coverage evaluation; net benefit remains unknown. All 30 cases were previously seen. This was a revised-protocol replication that reused Qodo reviews, changed several context dimensions, and shared an analyst for construction and adjudication. Passes and packets were repeated observations, not independent cases.

The matched 21-case comparison had identical final recommendations in both versions. The new gap came from the formerly incomplete query family.

Before giving extra judgments influence over approval, I want requests bound to named behaviors, traceable evidence, and preserved blockers. Fresh cases and measured maintainer effort should include useful verification and unsupported requests. A recommendation didn’t execute a fix, approve a change, or establish production behavior.

The evidence and offline replay preserve c020’s useful missing assertion and c028’s unsupported escalation. Together, they explain what the extra judgment contributed and what work it created.

Get started with Qodo for AI Code Review

Start trial
Share this post

More from our blog

Check out our musings on generative AI, code integrity, and other geeky stuff: