Early versions of AI agents can be remarkably small: a prompt, a model call, one or two tools, and a loop that continues until the model returns a final answer.
That simplicity is useful. It allows teams to validate a use case quickly and determine whether a model can select the right tool or break a task into meaningful steps.
But it also hides the hardest part of agent engineering: what happens between the moment a model proposes an action and the moment that action is executed inside a real system?
Something must interpret the output, check permissions, invoke the tool, handle failures, record the result, rebuild the context, and decide whether the task should continue. The execution may also need to be paused, resumed, observed, or audited.
That is the role of the agentic harness.
What is an agentic harness?
An agentic harness is the execution scaffolding that surrounds a language model and turns it into an agent capable of performing work. It drives model and tool calls, manages context and conversation state, applies approval policies, and can keep an agent progressing through a multi-step task.
The model provides the ability to interpret a request, reason over the available information, and propose a next action. The harness turns that proposal into a controlled execution step.
A typical agent loop follows this pattern:
- Build the model input from the request, instructions, current state, and authorized context.
- Call the model.
- Interpret the output.
- Return a final response if the task is complete.
- Execute a tool if the model requests one.
- Transfer the task if another agent or specialized component should take over.
- Add the result to the state of the run.
- Call the model again until a stopping condition is reached.
This pattern is implemented explicitly in several agent frameworks: the runner calls the model, classifies its output, executes requested tools or handoffs, appends the results, and starts another iteration until it receives a final output or reaches an execution limit.
The harness does not replace the model, nor does it necessarily reason on the model's behalf. It coordinates the interactions between the model, tools, state, policies, and surrounding services. It also enforces boundaries that the model itself must not be able to bypass.
The relationship can be summarized simply:
The model proposes an action.
The harness interprets it and orchestrates what happens next.What the harness actually manages
The model-and-tool loop is only the core of the system. To support a real application, the harness must connect several other responsibilities.
1. Context construction
At every step, the model needs the right information: the agent's instructions, the current user request, session history, tool results, memory, documents, artifacts, or intermediate business state.
Passing everything to the model creates noise, increases latency, and consumes more tokens. Passing too little prevents the agent from completing its task correctly.
The harness therefore assembles the context required for the current step. It can combine session-scoped instructions, tools, memory, task state, conversation history, and other context providers before invoking the model.
Context construction is not just a retrieval problem. It is also an authorization problem. The information included in a model call should reflect the identity, permissions, and scope of the current execution.
2. Loop orchestration
After every model call, the harness determines what should happen next:
- return a final answer;
- execute another tool;
- perform a handoff;
- start a new iteration;
- request external input;
- or stop the run.
The loop must remain bounded. A maximum number of steps, time limit, budget, cancellation signal, or business rule can terminate the execution.
Without those mechanisms, an agent may repeat an ineffective action, accumulate errors, or consume resources without producing a useful result. Agent frameworks therefore expose controls such as maximum turn counts or bounded reinvocation rules.
3. Tool execution
A tool call produced by a model is not yet an executed action. It is a structured request that another system must process.
The execution layer may need to:
- identify the requested tool;
- validate its parameters;
- verify the caller's identity and permissions;
- determine whether an approval is required;
- launch the tool in the appropriate environment;
- handle timeouts and failures;
- apply retry rules;
- record the result;
- return that result to the model in a usable format.
This separation is fundamental. A model may request an action, but it should not be solely responsible for deciding whether that action is authorized.
The harness coordinates the request and the applicable controls. The runtime, sandbox, policy engine, or external service then performs the corresponding operation.
4. State, checkpoints, and recovery
A multi-step task produces state: intermediate results, decisions, files, artifacts, validations, or business progress.
If the execution fails after its fourth step, restarting the entire task from the beginning may be inefficient or unsafe. A durable execution system should be able to persist state and establish checkpoints.
The runtime can then resume from a known point instead of replaying the full sequence blindly.
This requires more than storing a conversation transcript. The system must know which actions have already completed, which outputs were committed, and what information is required for the next step.
5. Policies and human control
Some actions can be automated. Others carry enough risk to require additional controls. Reading a document and initiating a payment should not be treated as equivalent operations.
The execution system can enforce rules concerning:
- accessible tools;
- permitted data;
- user identity;
- cost or token limits;
- approval requirements;
- execution scopes;
- and emergency termination.
A natural-language instruction can guide the model. A policy enforced by the execution system constrains which actions can actually take place.
This distinction matters because prompts influence behavior, while system-level controls determine what the agent is technically allowed to do.
In a billing-support example, "never promise a refund" is a behavioral instruction; separate permissions determine whether the agent may read payment data or create a case. A later response can still violate the instruction, which is why production evaluations matter.
6. Observability
A final answer is not enough to understand an agent's behavior.
Teams may need to inspect the complete run:
- model calls;
- tools selected;
- parameters submitted;
- results received;
- failures and retries;
- state transitions;
- latency;
- token usage;
- cost;
- and resources consumed.
This visibility helps engineers diagnose failures, compare configurations, evaluate a model change, and verify that execution policies were applied.
In an agentic system, observability must cover the path from the initial request to the final result—not just the last model response.
Observability also makes it possible to evaluate runs after launch. A response that sounds helpful may still break a business rule; the trace helps a team inspect what happened, save the case, test a revision, and review the result before releasing a change.
Model, agent, harness, runtime, and SDK are different layers
These terms describe different responsibilities, even though their boundaries may vary across frameworks.
The model
The model transforms an input into an output: text, structured data, or a request to use a tool.
It is a capability used by the agent, not the entire agentic system.
The agent
The agent associates an objective with an execution configuration: instructions, a model, and, depending on the use case, tools, context, memory, or behavioral rules.
It represents what the system is expected to accomplish.
The harness
The harness orchestrates how that work unfolds.
It connects the model to tools, state, policies, context, and stopping conditions. It interprets the model's outputs and coordinates the next execution step.
The runtime
The runtime is the engine and environment in which the harness and the agent's code execute.
Depending on the architecture, it may also provide workers, state persistence, event transport, artifact storage, and access to external services. As the underlying engine, the runtime orchestrates agents, tools, callbacks, state changes, and interactions with services such as models and storage.
The boundary between harness and runtime is not perfectly uniform across the industry. Some frameworks include runtime capabilities within their definition of the harness. Others treat orchestration logic and operational infrastructure as distinct layers.
The distinction remains useful:
Harness: how the work is orchestrated.
Runtime: where and through which services it is executed.The SDK
The SDK is the development surface exposed to engineering teams.
It may:
- contain a local agent runner;
- provide direct access to a model API;
- expose classes for defining agents and tools;
- communicate with a remote runtime;
- or combine several of these functions.
The word "SDK" alone does not reveal where the agent loop runs.
This distinction is explicit in existing product architectures. An agent SDK may run the loop inside the developer's process; a client SDK may provide API access while leaving the loop to the developer; and a managed-agent API may run both the agent and its sandbox remotely.
In Cominty, the SDK provides a typed API client and the primitives used to define agents, tools, policies, hooks, and other customization components. The production loop itself runs on Cominty's infrastructure.
Embedded harness or managed harness?
There are two broad ways to operate an agentic harness.
Embedded harness
In an embedded architecture, the agent loop runs inside the application or within a service operated directly by the engineering team.
This gives the team direct control over the process. It also means that the team must operate whichever supporting capabilities its use case requires:
- state and session storage;
- recovery;
- tool execution;
- secrets;
- isolation;
- concurrency;
- policy enforcement;
- and observability.
A local harness can be an appropriate choice when the execution is short-lived, the environment is tightly controlled, or the team explicitly wants to own the complete runtime.
Managed harness
In a managed architecture, the application submits a run to a remote service. The harness and runtime then take responsibility for progressing the execution.
Depending on the implementation, the client can follow events, receive the result, or provide additional input without maintaining the process responsible for the agent loop.
A managed-agent platform can provide the harness, runtime, tool execution layer, stateful sessions, and sandbox infrastructure as an integrated service.
Cominty follows this managed architecture.
An agent can be defined through the graphical interface, API, or SDK. When a production run begins, its execution continues on Cominty workers independently of the client process that submitted it.
How Cominty fits into this architecture
Cominty brings together three complementary layers within one Managed Agents platform.
A definition and control SDK
Developers use the Cominty SDK to describe an agent and its capabilities:
- instructions;
- tools;
- policies;
- hooks;
- and customization components.
The SDK also acts as a typed client for communicating with the platform, starting a run, and consuming its result or event stream.
The SDK is not a local runner that requires the client application to remain active for the duration of the task. Once submitted, the run is handled by Cominty.
This is consistent with Cominty's product model: engineering teams build through the SDK, while the platform operates the infrastructure behind a unified streaming API.
A configurable agentic harness
Cominty provides the agent loop and its orchestration mechanisms.
Teams do not start from an empty loop. They customize an existing harness through components and hooks that reflect their product and business logic.
The harness:
- constructs the working context;
- interprets model outputs;
- orchestrates model and tool calls;
- coordinates policy enforcement;
- and determines whether the run should continue or terminate.
The runtime performs the operational work and persists the resulting state changes.
Cominty therefore provides more than a model-access SDK. The platform supplies the orchestration layer on which the agent executes, while leaving developers in control of what differentiates the agent.
A fully managed runtime
The harness and the agent's code run on Cominty workers.
The runtime manages:
- sessions;
- execution state;
- checkpoints;
- recovery;
- access to model providers;
- and end-to-end observability.
Cominty's product documentation describes this operational layer as durable orchestration with state, checkpoints and retries, sandboxed execution, and observability covering tool inputs and outputs, latency, and cost.
Cominty Managed Agents — an example of this architecture (opens in a new tab)
The architecture can be represented as follows:
The Managed Agent is the operational result of this architecture: an agent definition, its customization components, a harness, a runtime, persistent state, and an execution environment delivered as one service.
In other words, Cominty is not simply an SDK, a remote sandbox, or a model gateway. It is a Managed Agents platform composed of:
A development SDK
+ a configurable agentic harness
+ a managed agent runtimeExample: processing a billing support ticket
Consider an agent responsible for preparing a response to a customer who reports being charged twice.
In this illustrative case, the agent can consult permitted billing records and the refund policy, then open a ticket for a billing specialist. Its instructions say not to promise a refund before that review.
The business application submits the ticket and starts a run. It does not need to keep a local agent loop active until the task finishes.
The Cominty runtime loads the agent definition, its current state, and the applicable policies.
The harness builds the context and determines the next model call. The runtime sends the request to the configured model provider. The model may then request access to a tool that retrieves billing information.
A model gateway can apply the configured provider policy to each model call and route to a fallback when the primary provider is unavailable.
The harness interprets the request and coordinates the applicable controls. The runtime routes the permitted billing lookup to its connected service and records the result in the run state. If the agent also needs to execute code or a code-based skill, that work can run in a separate sandbox.
That result is passed back into the loop. The harness reconstructs the context, the runtime performs the next model call, and the model produces a proposed response based on the billing information returned by the tool.
After consulting the refund policy, the agent can create the billing ticket and draft a reply grounded in the payment record, policy, and completed action—without promising an unapproved refund.
Throughout the execution, the client application can consume the event stream, display progress, or wait for the final result. If the client process stops after submitting the task, the run can continue on Cominty workers.
Each layer has a distinct responsibility:
- the model interprets the ticket and proposes actions;
- the harness orchestrates and controls the sequence;
- the runtime performs and persists the execution;
- the sandbox isolates code and code-based skill execution when that environment is needed;
- the SDK connects the business application to the platform.
If a later run nevertheless promises a refund, an online evaluation can flag that response for review. The saved case can then be used to test revised instructions and compare the result before the change is released.
The architectural decision behind every agent
Choosing an agent architecture is not only about selecting a model.
Engineering teams must also decide:
- where the agent loop runs;
- who owns the execution state;
- how tools are isolated;
- where permissions are enforced;
- how a run recovers from interruption;
- who operates the execution workers;
- and how the system is observed and evaluated.
A team can build and operate this entire layer itself.
Alternatively, it can retain control over its agents, tools, policies, and business logic while relying on a managed harness and runtime for execution.
That is where Cominty is positioned: as a Managed Agents platform that combines a development SDK, a configurable agentic harness, and a managed runtime.
Developers build the parts that differentiate their agents. Cominty operates the execution layer required to run them durably, with explicit controls and end-to-end visibility.
In summary
A model produces an output. An agent pursues an objective. The harness turns successive model outputs into a structured execution. The runtime provides the environment and operational services that allow that execution to continue, recover, and be observed. The SDK gives developers a way to define, customize, and control the system.
Within Cominty, these components belong to the same platform while retaining distinct responsibilities:
SDK: define, customize, and control.
Harness: orchestrate and govern.
Runtime: execute, persist, and operate.
Managed Agent: combine these layers into a usable service.Moving from an experimental model loop to a governed execution system is not, by itself, the whole journey to production. But it is an essential step toward building agentic systems that can operate reliably within a production environment.
