Spots

Building the Harness: A Production-Minded AI Agent Architecture in .NET

This is the implementation companion to the LinkedIn article The Model Isn’t Your Agent. The Harness Is. That article makes the architectural argument that a production agent is much more than the model providing its reasoning. Here, we’re going to build that idea. Most introductory agent examples begin with the model. Create a client, send it a prompt, give it a function or two, and start looping until the model stops requesting tools.

That is useful for learning an SDK. It

That is useful for learning an SDK. It is not the architecture I would want governing an agent capable of doing real work. Instead, let’s start from the outside and work inward. The system we’re going to build has a simple rule: The model can decide what it wants to do. The harness determines what it can do. The model can decide what it wants to do. The harness determines what it can do. That means identity, policy, budgets, tools, evidence, and verification belong outside the model. Our first architecture looks like this: There is deliberately no arrow that goes directly from the model to an external system. That boundary is the architecture. Let’s resist the temptation to create an OpenAI, Anthropic, or Azure client.

Our application should not begin by caring which

Our application should not begin by caring which model performs the work. It should begin with the contract for performing agent work. The request can stay simple: And our result should tell us more than whatever text the model happened to generate: Already, we have made an architectural decision. The output of our agent is not just text. It includes evidence and execution metrics because we eventually want to answer two different questions: What did the agent conclude? Why should we trust what happened? That distinction will matter later. Give the Agent an Identity Before the harness does anything, we need to know who is acting. AgentId identifies the agent performing the work. SubjectId identifies the human or workload on whose behalf the work is occurring. Those identities are deliberately separate.

An enterprise implementation might eventually map them to

An enterprise implementation might eventually map them to Entra identities, workload identities, claims, delegated authorization, or another identity provider. For now, the important thing is that identity is part of the execution contract rather than hidden somewhere inside configuration. Next we define authority. Now the agent’s permissions exist independently of anything the model was told in its prompt. That is important enough to make explicit. Instructions describe desired behavior. Policy defines permitted behavior. Instructions describe desired behavior. Policy defines permitted behavior. The model may receive both, but only one should decide whether an action actually occurs. Build an Execution Context Every agent invocation needs a bounded environment describing what is true for that particular execution. Then we can assemble the information governing one execution:

The model isn’t receiving unlimited execution and being

The model isn’t receiving unlimited execution and being asked to behave responsibly. The harness is creating a finite envelope in which reasoning can occur. That gives us another useful principle: Autonomy should be expressed as a budget, not an absence of limits. Autonomy should be expressed as a budget, not an absence of limits.

A production implementation could make these budgets considerably

A production implementation could make these budgets considerably richer. Token limits, wall-clock deadlines, network requests, compute consumption, model spend, delegated-agent count, retry count, or resource-specific quotas could all belong here. The mechanism matters less than making the boundary explicit. Put the Model Behind an Interface Now we can finally introduce a model. But instead of coupling the harness to a particular provider, we’ll define what we need from one. A model decision can produce one of several outcomes: This is intentionally provider-neutral.

An OpenAI implementation might translate a tool call

An OpenAI implementation might translate a tool call into InvokeTool. An Anthropic implementation could do the same. A local model could satisfy the same contract. Eventually, we could introduce a router: Now model selection becomes policy.

A cheap model might handle classification. A stronger

A cheap model might handle classification. A stronger model might handle complex planning. A failed verification step might cause the router to escalate. The rest of the application does not need to know. The model becomes an execution dependency rather than the architecture itself. The model becomes an execution dependency rather than the architecture itself. That is exactly where I want it. Now we reach the boundary where reasoning becomes action. The harness owns a registry of available tools: Here is where the architecture becomes interesting. Suppose the model requests: The model has made a decision, but nothing has happened yet. Before execution, the harness evaluates policy. If deploy-production is not in the allowed set, the answer is no. Not another prompt telling the model that production is dangerous. The tool does not execute.

A denied action should fail at the authority

A denied action should fail at the authority boundary, not depend on the model agreeing that it was denied. A denied action should fail at the authority boundary, not depend on the model agreeing that it was denied. That difference is the heart of the harness. Assemble the Execution Loop Now we have enough pieces to create a simplified implementation. There is nothing particularly magical here. The harness is orchestration code. Its job is not to be intelligent. Its job is to make intelligence governable. Prove That the Boundary Works Let’s create an identity for a coding agent. Its policy allows repository operations: The model receives a task: Fix the failing integration test and verify the solution. During reasoning, it decides that deploying the application would be useful and requests: Our authorization check runs: The model requested the action, but the action did not happen.

That is not an alignment success. It is

That is not an alignment success. It is an architecture success.

News

Building the Harness: A Production-Minded AI Agent Architecture in .NET

This is the implementation companion to the LinkedIn article The Model Isn’t Your Agent.

@spots #dev
Source: Dev.to
See more like this