Martin McRoy - VP of Engineering at Thread AI
This is the fourth post in Facts, Rules, Runtime, our series on the engineering system around agentic software development. Post 3: How a Harness Learns from Engineering Judgment followed the signals that help us improve the harness. This post follows an improvement into everyday use: how a decision made in one review becomes available to other agents, across projects and tools.
Fixing a pull request resolves the work in front of us. Carrying that correction forward takes another step. The team has to capture the judgment behind it, make it part of the workflow, and bring it into the next task where it matters. Otherwise, the same knowledge stays with the reviewer, and every new run depends on someone remembering to supply it.
Engineering Judgment Becomes Part of the Workflow
We use commands and skills to carry the recurring parts of our engineering process into agent work. The way we investigate a change, carry out a review, or prepare work for verification lives in a shared workflow. Engineers can invoke that workflow directly, and the harness can use it as part of a larger task. Both start from the same reviewed instructions.
The rules those workflows rely on live in Markdown guides alongside the code or in a shared collection. We update that guidance as our understanding changes. A review decision can clarify a convention; a change in architecture can retire an old assumption. Once adopted, the revised guidance reaches relevant tasks without an engineer having to reconstruct it from past conversations.
The evidence captured in Part 3 helps us decide what to change. AI processes the message queue and proposes updates from recurring review feedback and other engineering signals. Engineers review those proposals and decide whether the correction belongs in a guide, in the workflow itself, or in how the harness selects the guidance. We also look for instructions that can be combined or removed as the process evolves.
A correction carries into later work
This gives review work a longer life. The judgment behind a correction becomes something the team can use again, while later reviews tell us whether it helped. The verification layer from Part 2 still checks the resulting work. Better instructions improve what the agent starts with; verification establishes whether the work meets its requirements.
One Improvement Reaches More Than One Agent
Once a workflow is useful, we share it as a pack. The pack keeps the commands and skills connected to their guidance and to the conditions under which they apply. A team adopting the pack gets the working approach and the context needed to use it.
This matters when engineers use different agent tools or work across several repositories. We maintain the shared approach in one place and sync it into the tools that consume it. Claude, Gemini, and Codex have their own entry points, but improvements flow from the same reviewed source. Provider differences remain part of the maintained pack instead of becoming a separate setup task for every engineer.
Shared workflows also draw on guidance specific to the repository. That lets teams reuse a review process while keeping their own architecture and conventions close to the code. A change to a local guide can improve how the workflow operates in that repository without changing the workflow for everyone else.
The harness handles delivery and sync, with validation to check that the configuration arrives complete and consistent. Engineers maintain the source and review changes there. The result is a common way to distribute improvements across the team, with less dependence on personal copies of instructions or manual setup.
The Guidance Fits the Work
As the team's knowledge grows, giving every task every guide makes the useful material harder to find. Routing brings the relevant workflows and guidance into the task. The conditions for using a capability travel with its pack, so adopting a shared workflow also brings the information the harness uses to select it.
We use static routing for familiar boundaries such as directories and file patterns, LLM routing to interpret natural-language descriptions of the work, and classifier-based routing for defined decisions about what a task needs. A classifier such as Jev can support those decisions. The choice depends on what makes the guidance relevant to the work.
Shared knowledge, focused on the task
For the engineer, the impact is less context to assemble by hand. The harness brings the team's baseline rules together with the guidance relevant to the task, and records what it selected. When an agent misses a convention, we can see whether the guidance needs to change or whether it failed to reach the task at all. That distinction keeps us from adding more instructions to solve a selection problem.
Improvements Move at the Right Pace
Guidance changes often. We keep those changes reviewable without turning every edit into a release of the entire harness. Repository guides evolve with the code and use the repository's version history. Shared packs and shared guides have versions that teams adopt deliberately, so a change made for one project does not immediately alter every other project's workflow.
This gives us room to keep learning while controlling how changes spread. A local clarification can reach the next run with the code. A shared workflow change can be reviewed and adopted across repositories at their own pace. When a workflow and its guidance need to change together, we deliver them together.
Adopt improvements without changing work in flight
Work already in progress keeps the instructions it started with. Later work receives the revisions its repository has adopted. The run record tells us which versions were used, so we can connect a change in behavior to a change in guidance, compare the review evidence, and return to an earlier version when needed.
That is why we treat agent configuration as infrastructure. It carries the team's decisions into execution. Review gives those decisions a shared source, delivery makes them available, and routing brings them into the work where they apply. Each improvement has a path beyond the conversation that produced it.
What Comes Next
The harness now has a way to turn engineering judgment into instructions that reach future work. The final post follows that work through unattended execution: how the harness coordinates tasks, recovers from interruptions, and brings an engineer back in when a decision needs human judgment.