Skip to content

Case study

One week of optimizing my OpenCode workflow

How I turned token-hygiene conversations into OpenCode rules for focused discovery, smaller changes, and evidence-driven verification.

7 min read
  • AI Infrastructure
  • Platform Engineering
  • Performance
Less noise. Better context. Scattered repository files and logs narrow into a search, patch, diff, and verify workflow.

An agent can read a lot of code without getting any closer to the right change.

That distinction became the focus of my OpenCode work this week. I wanted a workflow where every search, file read, and test had a reason. Before giving the agent more context, I wanted it to ask whether that context would change its next engineering decision.

I did not change OpenCode’s inference engine or benchmark a new model. I changed the instructions governing how my coding agent works. The result is a token-hygiene protocol in my active AGENTS.md, built from conversations about where coding agents spend context and how to keep that spending useful.

The objective is straightforward: maximum correct engineering work per token.

The word correct matters. An agent that saves tokens by skipping the evidence is a bad optimization.

Treat context as a resource budget

From an infrastructure perspective, this is a resource-allocation problem. A service needs enough memory to do its job, but allocating more memory does not fix a bad access pattern. A coding agent needs enough context to reason about a change, but loading unrelated files does not make that change safer.

My first conversation explored four sources of waste: context bloat, repeated reasoning, low-value output, and irrelevant information left in the working conversation. These are different problems. Trimming the final answer does little about an agent that already read thousands of unnecessary lines.

The proposed workflow therefore started before implementation. Identify the task, locate the relevant code, read the necessary region, and expand only when the evidence is insufficient.

That last condition is important. I am not trying to make the agent incurious. I am trying to give its curiosity a purpose.

Turn advice into operating instructions

The first discussion was broad. It covered repository maps, ignore rules, log filtering, checkpoints, model routing, and dividing responsibilities between agents. It was useful as an exploration, but it was not a record of completed changes.

The next step was to narrow that discussion into an 80/20 protocol. That protocol now appears in my active OpenCode instructions. It tells the agent to search before reading, avoid high-noise files by default, make small patches, inspect its diff, and run targeted tests before broader validation.

It also gives the agent a stopping condition. Once it has enough evidence to implement and verify the change safely, it should stop exploring. Finding another interesting file is not progress on the task.

The rule I keep coming back to is this:

Will this information materially change my next engineering decision?

If the answer is no, there needs to be a better reason to retrieve it.

Make discovery answer a specific question

The protocol replaces open-ended repository reading with a sequence of questions. Where is the relevant implementation? What does it depend on? What is the smallest safe change? What evidence would prove that change works?

A task involving token validation, for example, does not automatically justify reading every authentication file. Start by locating the relevant symbol. Read its implementation and the dependencies that matter to the behavior. Expand when a caller, configuration value, or test exposes an unanswered question.

This is also why dependency directories, build output, generated clients, and large lockfiles are excluded from routine reading. They are not forbidden. They need to be relevant to the problem before they enter the conversation.

The distinction prevents a useful default from becoming a dangerous absolute. If the failure is in generated code, that code is evidence. If the task is a dependency issue, the lockfile may be exactly what the agent needs.

Review the change, then prove the behavior

I also made the edit loop explicit. Prefer a patch to a rewrite. After editing, inspect git diff instead of rereading the entire file by habit.

The diff is the primary review artifact for the change itself. It makes accidental formatting, unrelated edits, duplicated logic, and interface changes easier to spot. It does not prove runtime correctness. Tests and other behavioral checks still have to do that job.

The protocol starts validation near the change, then expands when the risk warrants it. A focused test can provide useful feedback before a broad suite produces a large amount of unrelated output.

For an authentication change, an illustrative sequence might be:

bash
git diff
pytest tests/test_auth.py -q

These commands illustrate the workflow. They are not measurements from this week or evidence that a particular test suite passed.

A local test is a starting point, not permission to ignore integration behavior. Changes to shared interfaces, configuration, or cross-module assumptions need checks beyond the edited function.

Keep output small without hiding failures

Tool output needs the same discipline as file retrieval. A successful command usually does not need its entire log copied into the conversation. A failure needs the assertion, traceback, and surrounding evidence that explain what happened.

My protocol asks for small summaries on success and relevant output on failure, with expansion when necessary. The full log can remain available outside the conversation.

There is a reliability trap here. Truncating output is not the same as preserving the command’s result. If a test is piped through tail, the shell configuration determines whether the pipeline reports the test’s failure or only the filter’s success. A clean-looking excerpt is not proof of a successful run.

The requirement is to preserve the exit status and enough evidence to diagnose the failure. Smaller output is useful only after those requirements are met.

The protocol applies the same reasoning to narration. I want the agent to report discoveries, decisions, blockers, and validation results. I do not need an announcement before every routine search.

Keep the state that the next session needs

Long tasks create another source of waste: rediscovery. If the agent forgets which hypothesis was ruled out or which test already passed, it can repeat work without improving the result.

The protocol directs long-running work toward a compact .agent/current-task.md containing the goal, findings, changed files, validation, remaining work, and constraints. It also tells the agent to consult concise repository guidance when that guidance exists.

This is a policy in my configuration, not a claim that every repository now has those files. The first conversation proposed a larger collection of maps and decision records. The practical rule is simpler: preserve the information that prevents repeated investigation, and keep it small enough to use.

A checkpoint should describe the current state. It should not reproduce the transcript. It also needs maintenance, because an outdated summary can send the next session in the wrong direction.

What I can actually claim after one week

The evidence for this post is the progression in my saved conversations and the protocol present in my active instructions. The broad discussion became a concrete set of rules governing discovery, editing, output, and verification.

Those records do not establish a percentage reduction in token usage, a faster completion time, or a lower error rate. They also do not show that model routing, caching, or a multi-agent architecture was implemented. Those were ideas in the wider discussion.

The configuration change is real. Its performance effect still needs measurement.

A useful comparison would keep tasks and model settings comparable, then record token usage, completion time, repeated reads, verification results, and whether the final change was correct. Lower consumption would count as an improvement only if the work still met the same standard.

The tradeoff is how much evidence is enough

The risk of this approach is under-reading. A narrow function excerpt can hide an important caller. A filtered log can omit the first failure. A short plan can overlook a dependency.

That is why the protocol allows scope to expand when the evidence is insufficient. It also requires broader validation when regression risk warrants it. I prefer those conditions to a fixed limit on searches or file reads. A small bug and a cross-service migration should not have the same discovery budget.

Instructions have a cost, too. A growing AGENTS.md can become the context bloat it was meant to prevent. The rules need to remain coherent, relevant, and free of repeated advice.

The takeaway

My useful change this week was turning a broad optimization discussion into explicit working rules. The agent now has instructions about what to retrieve, what to preserve, how to review its changes, and when to stop.

That is where I would start before changing models or adding more agents. Look at the workflow first. Make each read answer a question. Make each edit serve the task. Make each correctness claim depend on evidence.

Give the model the minimum evidence needed for the next correct decision, and make it retrieve more when that minimum is not enough.

Related