How Cheap Is Good Enough? Four LLMs as opencode Analyst Subagents
Right now everybody seems to be short on tokens. Either the capacity is not there, or the price is, and often both at once. So you reach for smaller and cheaper models and hope the quality holds. The obvious question: how much less is acceptable?
We compared four LLMs in the analyst role: Qwen3.8-27B, DeepSeek V4 Flash 0731, Qwen3.6-35B-A3B-FP8 and Nemotron 3.5 Lightning. We choose these contemporary models because they can be run either locally or remotely for a low price. All are said to be top of their respective class. Each answered 21 questions drawn from a live coding session then asked to the analyst subagent.
TL;DR: Nemotron ranked last in the combined score and produced unusable answers on 5 of 21 questions (24%). Qwen3.8-27B and DeepSeek V4 Flash 0731 always provided useful answers. Qwen3.6-35B-A3B-FP8 failed 2 of 21 times and landed mid-tier.

The Setup: An Agent Team That Reads Code
The analyst is part of a set of AI agents designed to take over certain roles in a developer team. Each team has an orchestrator that delegates work to specialists: a coder that writes code, an analyst that investigates the codebase and a reviewer that checks quality. For this task, we added a special role: a judge that scores outputs. Each role is backed by an LLM you configure.
The analyst is the read-only researcher. It gets dispatched with a question like “Does the customer framework support a second HTTP port, or do we need a standalone server?” and is expected to read source files, query a codebase Q&A tool, and return a report with file paths, line numbers, and verified claims. It does not write code, it gathers facts for the orchestrator or the coder.
This is a job where speed matters: the orchestrator might dispatch five analyst subagents in parallel, and the user is waiting. It is also a job where being wrong is expensive: a fabricated file path or an inverted conclusion sends the coder in the wrong direction. So before configuring anything cheap into this role, we wanted to know how much quality the cheap tier actually costs.
Where the Questions Come From
The 21 questions are not synthetic or benchmark-style. They come from a real opencode session on a real TNG project: a device management platform built in Java with Spring Boot, Vaadin, and a custom framework. The codebase is a multi-module Maven project with ~50K lines of Java, a Vaadin web GUI, and an integration layer. Additionally the project depends on three libraries that often carry functionality and have to be understood as well as the main codebase to implement features correctly. In the opencode session in question we added a new module to the project that had to tightly integrate with the existing modules.
The analyst was therefore needed to understand the framework and project structure and answer questions towards it, like:
- “Is there a DummyAgent class in the codebase, or is it just the LoadTest mock infrastructure?”
- “How does the framework serialize command responses? Is there a double-wrapping problem?”
- “Does the state collector call
getStateswith or without a handle ID? Will@AgentRefthrow on a handle-less request?” - “What enum values does
DeviceOnlineStatehave? Does it includeOFFLINE?” - “How does the backend invoke device commands? Is it a single polymorphic command or multiple sub-commands?”
These are exactly the kind of questions an analyst subagent faces in production: deep, specific, requiring code reading and framework understanding. The answers were originally produced by Nemotron 3.5 Lightning, but rather than taking those original answers as the baseline, we re-ran all 21 questions with Nemotron in the same evaluation setup as the candidates. This ensures a fair comparison: same codebase state, same tool access, same environment. We then re-ran the same 21 questions with Qwen3.8-27B, DeepSeek-V4-Flash-0731, and Qwen3.6-35B-A3B-FP8 as the analyst model, keeping all other conditions identical.
Two Waves
The 21 questions naturally split into two groups:
- Wave 1 (12 questions) can be answered from the codebase alone.
- Wave 2 (9 questions) additionally require a design spec document that was present in the original session. These questions ask the analyst to verify reviewer claims against the spec, e.g. “The reviewer says the spec’s command design is wrong. Is it?”
We run Wave 1 first, then restored the spec and ran Wave 2. This separation tests whether models can handle both pure codebase investigation and spec-vs-code cross-referencing.
Sandboxing
The agents run in an omac sandbox. Without it, Qwen and DeepSeek V4 Flash 0731 will try to search the HDD for already existing answers, sometimes finding them in the original repo. Cheating, technically. We also clone the git repo from a previous state because the agents would use the reflog to get to the answers instead of working for them. They also like to read the benchmark results of previous runs for answers, which is why we remove work files before each run. We let a review agent go through the answers afterwards to check for more cheating attempts.
Judging Methodology
We cannot just eyeball the answers, we need reproducible, bias-controlled scoring. We use the following methodology.
Blinding
Each model’s answer is labeled with a random letter (A, B, C, or D) per question, using a deterministic shuffle. The mapping from letters to model names is stored in a separate file the judge never sees. For each question, the judge sees four answers labeled A, B, C, D, with no indication of which model produced which.
The Judge
We dispatch a judge subagent, a separate LLM instance configured with read-only access to the codebase, for each of the 21 questions. The judge model is DeepSeek V4 Pro (provided by skainet), chosen because it is not one of the candidate models being evaluated, not the orchestrator model of the judging session, and overall a top-performing model. Each judge receives:
- The original question text.
- All four answers, anonymized as A, B, C, D.
- A scoring rubric.
- Instructions to verify at least 3 load-bearing claims per answer against the actual source code.
The judge then returned structured JSON with scores, a ranking, and a qualitative comparison.
The Rubric
Each answer was scored 1 to 5 on six dimensions:
Table 1: scoring dimensions
| Dimension | What it measures |
|---|---|
| Accuracy | Are the claims correct? Can they be verified against the codebase? Fabricated evidence gets a 1. |
| Evidence | Does the answer cite file paths and line numbers, or just assert things? |
| Completeness | Does it address all parts of the question? |
| Usefulness | Can the orchestrator act on this, or must it re-dispatch? 1 = re-dispatch, 5 = ready for coder handoff. |
| Conciseness | Is it the shortest answer that covers the same ground? Padding and tangents lower the score. |
| Session Length | Did the agent burn through context, or was it efficient with its tool calls? |
The overall score is the mean of all six dimensions. The critical metric for us is Usefulness: a score of 1 or 2 means the answer is unusable, the orchestrator would have to throw it away and re-dispatch the question. That’s wasted time and wasted tokens.
What “Unusable” Means in Practice
When the judge scores an answer’s usefulness as 1 or 2, it means one of several things:
The answer is a shallow summary that mischaracterizes the architecture. Nemotron’s Q1 (2,239 chars) called an agent “a Spring WebFlux module” when it’s actually a servlet/Spring-MVC app, and claimed the FeatureSetsInvocationHandler enables “error injection, false information” when it’s a transparent proxy. It gave no path:line citations at all.
The answer fabricates a framework annotation that doesn’t exist. Nemotron’s Q6 invented
@Agent(name = "deviceAgent", requestBase = "/device_agent/")and recommended the mock register with it. The real architecture uses a single Agent class withrequestBase = "accountant". A coder following this advice would create a completely wrong agent registration.The answer cites almost no evidence despite claiming it did. Nemotron’s Q20 claimed “all findings are documented with exact file paths and line numbers” but provided only two citations in the entire answer. It also stated “the real agent uses the class name GetStates”, no such command exists in the agent module.
The answer is off-topic entirely. Qwen3.6-35B’s Q7, the worst single answer in the evaluation (overall 1.83), traced the framework dispatch path and discussed @Signature overloading and serialization, but never addressed the actual question: whether the framework’s standard commands (create, read, update, delete) can be overridden. The orchestrator would need to re-dispatch from scratch.
In all these cases, the answer is worse than no answer: it looks plausible enough to mislead the orchestrator or coder, but sends them in the wrong direction.
Results
Unusable Output Rate
This is the metric that matters for a model swap decision. Usefulness below 2 means the orchestrator must re-dispatch.
Table 2: unusable output per model
| Model | Unusable answers | Rate | Avg Usefulness |
|---|---|---|---|
| Nemotron 3.5 Lightning | 5 of 21 | 24% | 3.05 |
| Qwen3.8-27B | 0 of 21 | 0% | 4.52 |
| DeepSeek V4 Flash 0731 | 0 of 21 | 0% | 4.86 |
| Qwen3.6-35B-A3B-FP8 | 2 of 21 | 10% | 3.86 |
Nemotron fails on 24% of questions. The orchestrator would have to re-dispatch roughly a quarter of all analyst tasks, and it wouldn’t necessarily know which ones, which is where the expensive mistakes happen. Qwen3.8 and DeepSeek V4 Flash 0731 never produce unusable output. DeepSeek V4 Flash 0731 scored 5/5 on usefulness for 18 of 21 questions, it’s not just usable, it’s ready for the coder to act on immediately.
Fabricated Evidence
The nastiest failure mode we saw is fabricated evidence: citing a real file path with line numbers, but attributing behavior to the code that simply isn’t there. Nemotron did this once: it claims a method it connected to a mock-mode flag when the code at those lines has nothing to do with it. The three other models never fabricated evidence. This failure is hard to catch because the answer looks well-sourced: it has file paths and line numbers, but the interpretation is invented. A coder acting on this answer would chase a connection that isn’t there.
Overall Scores

DeepSeek V4 Flash 0731 leads the field at 4.48 average overall, winning 11 of 21 questions (52%). Qwen3.8-27B is a close second at 4.37 with 8 wins, and notably, both models hit a perfect 5.00 on their best questions, showing they can produce excellent work when the problem suits them. Qwen3.6-35B-A3B-FP8 sits in a distant third (3.87, 2 wins); its worst answer (1.83) is the lowest in the entire evaluation. Nemotron 3.5 Lightning never won a single question, peaking at 4.17 and bottoming out at 2.67, still 0.31 below DeepSeek V4 Flash 0731’s average of 4.48.
Per-Dimension Breakdown

Three patterns stand out:
DeepSeek V4 Flash 0731 dominates accuracy, evidence, completeness, and usefulness, scoring near-perfect on all four. It backs every claim with file paths and line numbers, addresses every part of the question, and produces answers ready for coder handoff.
Qwen3.8 is a close second: 4.37 overall versus DeepSeek V4 Flash 0731’s 4.48. Its evidence score of 4.33 (up from 2.67 in the contaminated run) puts its citations nearly on par with DeepSeek V4 Flash 0731’s, and it matches DeepSeek V4 Flash 0731 and beats Nemotron on session efficiency. The two top models are much closer than this gap suggests.
Qwen3.8 leads conciseness (4.14 vs DeepSeek V4 Flash 0731’s 3.57), but by a narrower margin than run-to-run noise would suggest. Its answers run roughly two-thirds of DeepSeek V4 Flash 0731’s length, not a token blowout.
Head-to-Head
Table 3: head-to-head wins per pair
| Pair | Wins | Losses | Ties |
|---|---|---|---|
| DeepSeek V4 Flash 0731 vs Nemotron | 20 | 1 | 0 |
| DeepSeek V4 Flash 0731 vs Qwen3.8 | 12 | 7 | 2 |
| DeepSeek V4 Flash 0731 vs Qwen3.6 | 16 | 3 | 2 |
| Qwen3.8 vs Nemotron | 21 | 0 | 0 |
| Qwen3.8 vs Qwen3.6 | 16 | 4 | 1 |
| Qwen3.6 vs Nemotron | 18 | 2 | 1 |
DeepSeek V4 Flash 0731 beats Nemotron head-to-head on 20 of 21 questions, and Qwen3.8 beats Nemotron on all 21. The interesting result is the top pair: DeepSeek V4 Flash 0731 vs Qwen3.8 is 12-7 with two ties, a real advantage, but not a blowout. Both top models win against Nemotron across the board.
The Cost of Verbosity
DeepSeek V4 Flash 0731’s principal weakness is length. It writes a lot:
Table 4: answer length per model
| Model | Avg chars/answer | Total output |
|---|---|---|
| Nemotron | 4,306 | 90,420 |
| Qwen3.8-27B | 7,450 | 156,444 |
| DeepSeek V4 Flash 0731 | 11,280 | 236,888 |
| Qwen3.6-35B | 9,736 | 204,461 |
DeepSeek V4 Flash 0731 produces 1.5× more text than Qwen3.8 per answer. Its sessions (including tool calls and intermediate reasoning) average 112,037 tokens vs Qwen3.8’s 115,118, so for total session cost the two are effectively even. Qwen3.6-35B’s sessions run 90,957 tokens on average, the most efficient of the four.
Worth highlighting: Qwen3.8-27B has a comparable session length to DeepSeek V4 Flash 0731. Earlier Qwen models were criticized for their session verbosity. This issue seems to be reduced. Also, the final answer is only 66% the length of the DeepSeek V4 Flash 0731 answers while delivering comparable quality. That matters because the answer is the text handed back to the main session, taking up context length from the orchestrator.
Conclusion
Nemotron 3.5 Lightning is not good enough for the analyst role: it produces unusable output on 24% of questions, fabricated evidence once, and won 0 of 21 head-to-head comparisons in the four-model run. Its speed advantage, real as it is, doesn’t compensate for the orchestrator having to re-dispatch a quarter of its tasks. It is also the most token-hungry of all four models (175,911 tokens on average per session, including tool calls and intermediate reasoning). The cheapest option in the line-up costs the most tokens. Of course it does.
The three alternatives are clearly ahead:
DeepSeek V4 Flash 0731 is the highest-quality analyst. It won 11 of 21 questions, never produced unusable output, never fabricated evidence, and scored near-perfect on accuracy and evidence. Its weakness is verbosity, it writes 1.5× more than Qwen3.8 per answer, but if investigation quality matters more than token cost, it’s the pick.
Qwen3.8-27B is the close runner-up, not the distant second of earlier runs. It never produced unusable output, won 8 of 21 questions, lost 7-12 to DeepSeek V4 Flash 0731 pairwise with two ties, and its evidence score of 4.33 shows it now cites file paths and line numbers almost as consistently as DeepSeek V4 Flash 0731. It matches DeepSeek V4 Flash 0731 on session cost. If token efficiency is your priority, the gap to the top is small enough that either choice is defensible.
Qwen3.6-35B-A3B-FP8 is a mid-tier option. It ranks third overall (3.87), clearly better than Nemotron but behind the top two: it won only 2 of 21 questions and produced 2 unusable answers, including the worst single answer in the evaluation (Q7, overall 1.83). A reasonable budget choice, just not in the same league as the top two.
So, how cheap is good enough? In this role, on this codebase: down to the Qwen tier, yes. Below that, the re-dispatches eat what you saved.
What this run didn’t measure: other roles, other codebases, longer sessions. 21 questions on one Java project is a narrow slice. Swapping the model is a one-line config change in opencode; the downstream impact shows up in hours saved or lost per session.
Have fun rationing your tokens…