The Harnesses
How harnesses connect models, prompts, skills, and tools
You ask an AI to fix a bug. It reads the project, edits a file, runs a test, sees a failure, and tries again. The model is part of this process. But something also has to open the files, run the commands, and bring the results back.
That surrounding software is the harness. After learning about agents and skills, it helps to look at the system that makes them work together.
What Is a Harness?
An agent harness is the software that runs an AI agent. It prepares the model's context, exposes tools, handles tool requests, and manages the steps of a task. Depending on the application, it can also manage permissions, saved progress, retries, and limits.
People sometimes use agent, framework, and harness loosely. In this chapter, we use these distinctions:
| Part | What it does | Example |
|---|---|---|
| Model | Produces a response or a request to use a tool | Suggests reading a file |
| Harness | Runs the surrounding process | Checks access, reads the file, returns its contents |
| Agent | The working system pursuing a goal, using a model and a harness | Investigates and fixes a bug |
| Prompt | Gives instructions for the task or behavior | "Fix this bug and report the tests you ran" |
| Skill | Packages reusable instructions and optional resources | A bug-fixing procedure with a testing checklist |
| Tool | Performs a specific operation | Reads a file or runs a test |
A harness can be a small loop in your own application or part of a larger coding assistant. It does not have to include multiple agents or a complicated framework. Anthropic describes this distinction between an agent and its harness in its guide to evaluating agents.
An agent harness runs the agent's work. An evaluation harness runs test tasks and measures how well that system performs. They can work together, but they have different jobs.
How the Loop Works
Imagine the task is to improve an empty search results page. A typical loop looks like this:
- Prepare context. Include the request, project instructions, relevant skill instructions, and available tool definitions.
- Call the model. It might answer directly or request a tool, such as reading the search component.
- Check and execute. The harness validates the request and checks permissions before the tool runs.
- Return the observation. The file contents, command output, or error go back into the model's context.
- Continue or stop. The model uses that evidence for its next step. The harness also enforces limits and can pause for user input.
Writing "run the tests" in a response does not run anything. A real tool has to execute the command, and its result has to come back. That feedback lets an agent react to what actually happened. See Anthropic's explanation of agent workflows.
Try the Same Task with a Different Setup
Change one setting, then step through the run. Watch what changes in the context, the tool request, and the final report.
Inside the harness
A scripted simulation. No AI calls, files, or commands run here. Real agents can take different paths.
Task
Fix the empty search results page.
Same model in every run
Changing a setting resets the run. A skill helps with the method; it is not required for every task.
Execution trace
ReadyThe harness prepares context
The model requests a tool
The harness checks access
Tool results return to the model
The loop checks the work
The run stops and reports
Ready
These are example outcomes, not rules about what every model will do. A model may ask a useful question even with a vague prompt. Clear instructions make your expectations easier to follow and check.
Why Prompts Matter Inside a Harness
A harness gives the model a way to act. Your prompts explain what useful action looks like. Without a clear goal, an agent can spend many steps doing work you did not need.
The current request is only one part of the instructions. An application may also supply a system prompt, project guidance, tool descriptions, and instructions from a loaded skill. The harness decides how these enter the context; the model follows the instruction hierarchy used by that system.
Unclear task
Fix search. Make it better.
A task the agent can check
When search has no matches, show a helpful empty state. Keep matching behavior unchanged. Follow the existing UI patterns and use translations for new text. Run the relevant tests. Report what changed, which checks passed, and anything you could not verify.
You do not need to describe every click or command. Define the outcome, relevant constraints, and what counts as done. Let the agent choose routine steps when that flexibility helps.
Tool descriptions are also part of prompting. A tool with clear inputs, outputs, and limits is easier for the model to use correctly. Adding more tools without explaining them can create confusion.
Where Skills Fit
A prompt usually describes the current task. A skill captures a method you want to reuse across tasks: how to review a change, investigate a bug, or prepare a report.
In the Agent Skills format, a skill has a SKILL.md file and can include scripts, references, and assets. A compatible harness makes skills discoverable. It can provide their names and descriptions first, load the selected skill's instructions when relevant, and read supporting files only when needed. This is called progressive disclosure.
The harness still needs the tools and environment that the skill expects. A skill describing a browser test does not install a browser. A script may need dependencies. Instructions may need changes when moved to another harness with different tools. Loading a skill also does not train the model or guarantee expertise.
Use the task prompt for what is specific today. Put a repeated procedure in a skill. Keep shared project rules in the instruction files your harness supports. This avoids copying a large, slightly different checklist into every request.
A skill says to run browser tests, but the harness has no browser tool. What should happen?
Instructions and Permissions Have Different Jobs
"Do not edit files outside this project" is a useful instruction. A filesystem boundary that blocks those writes is an enforced control. Reliable setups use both.
Permissions, sandboxing, and approval checks must be implemented by the harness, its tools, or the underlying environment. A prompt cannot enforce them by itself. The VS Code agent security documentation shows concrete examples of these controls.
The same distinction matters for information coming back from tools. A web page or repository file can contain text that looks like an instruction. Treat that text as task data unless it comes from a trusted instruction source. Keep access limited to what the work needs.
Which control actually prevents a tool from writing outside an allowed directory?
What Makes a Harness Useful?
More autonomy is useful only when the system can check its work and recover from problems. For a long task, look at how the harness handles:
- Context: Keep relevant instructions and evidence available. Summarizing old steps can save space, but it can also lose details.
- Progress: Save decisions and unfinished work so the next session can continue. Saved state still needs to be loaded into context.
- Failures: Return errors clearly, use bounded retries, and avoid repeating actions with side effects blindly.
- Limits: Stop at a time, cost, or step budget, or when the work needs user input. Stopping is not proof of success.
- Verification: Inspect actual files, test results, or other outcomes. A confident final message is not enough.
When comparing setups, try the same real tasks and inspect both the outcomes and the execution traces. Keep the model fixed when you want to study a harness change. Change one prompt, skill, or tool setting at a time, and repeat tasks because individual runs vary.