Advanced Strategies

The Harnesses

How harnesses connect models, prompts, skills, and tools

You ask an AI to fix a bug. It reads the project, edits a file, runs a test, sees a failure, and tries again. The model is part of this process. But something also has to open the files, run the commands, and bring the results back.

That surrounding software is the harness. After learning about agents and skills, it helps to look at the system that makes them work together.

What Is a Harness?

An agent harness is the software that runs an AI agent. It prepares the model's context, exposes tools, handles tool requests, and manages the steps of a task. Depending on the application, it can also manage permissions, saved progress, retries, and limits.

People sometimes use agent, framework, and harness loosely. In this chapter, we use these distinctions:

PartWhat it doesExample
ModelProduces a response or a request to use a toolSuggests reading a file
HarnessRuns the surrounding processChecks access, reads the file, returns its contents
AgentThe working system pursuing a goal, using a model and a harnessInvestigates and fixes a bug
PromptGives instructions for the task or behavior"Fix this bug and report the tests you ran"
SkillPackages reusable instructions and optional resourcesA bug-fixing procedure with a testing checklist
ToolPerforms a specific operationReads a file or runs a test

A harness can be a small loop in your own application or part of a larger coding assistant. It does not have to include multiple agents or a complicated framework. Anthropic describes this distinction between an agent and its harness in its guide to evaluating agents.

Two uses of the same word

An agent harness runs the agent's work. An evaluation harness runs test tasks and measures how well that system performs. They can work together, but they have different jobs.

How the Loop Works

Imagine the task is to improve an empty search results page. A typical loop looks like this:

  1. Prepare context. Include the request, project instructions, relevant skill instructions, and available tool definitions.
  2. Call the model. It might answer directly or request a tool, such as reading the search component.
  3. Check and execute. The harness validates the request and checks permissions before the tool runs.
  4. Return the observation. The file contents, command output, or error go back into the model's context.
  5. Continue or stop. The model uses that evidence for its next step. The harness also enforces limits and can pause for user input.

Writing "run the tests" in a response does not run anything. A real tool has to execute the command, and its result has to come back. That feedback lets an agent react to what actually happened. See Anthropic's explanation of agent workflows.

Try the Same Task with a Different Setup

Change one setting, then step through the run. Watch what changes in the context, the tool request, and the final report.

Inside the harness

A scripted simulation. No AI calls, files, or commands run here. Real agents can take different paths.

Task

Fix the empty search results page.

Model

Same model in every run

Change the setup

Changing a setting resets the run. A skill helps with the method; it is not required for every task.

Execution trace

Ready
  1. The harness prepares context

  2. The model requests a tool

  3. The harness checks access

  4. Tool results return to the model

  5. The loop checks the work

  6. The run stops and reports

Ready

These are example outcomes, not rules about what every model will do. A model may ask a useful question even with a vague prompt. Clear instructions make your expectations easier to follow and check.

Why Prompts Matter Inside a Harness

A harness gives the model a way to act. Your prompts explain what useful action looks like. Without a clear goal, an agent can spend many steps doing work you did not need.

The current request is only one part of the instructions. An application may also supply a system prompt, project guidance, tool descriptions, and instructions from a loaded skill. The harness decides how these enter the context; the model follows the instruction hierarchy used by that system.

Unclear task

Fix search. Make it better.

A task the agent can check

When search has no matches, show a helpful empty state. Keep matching behavior unchanged. Follow the existing UI patterns and use translations for new text. Run the relevant tests. Report what changed, which checks passed, and anything you could not verify.

You do not need to describe every click or command. Define the outcome, relevant constraints, and what counts as done. Let the agent choose routine steps when that flexibility helps.

Tool descriptions are also part of prompting. A tool with clear inputs, outputs, and limits is easier for the model to use correctly. Adding more tools without explaining them can create confusion.

Where Skills Fit

A prompt usually describes the current task. A skill captures a method you want to reuse across tasks: how to review a change, investigate a bug, or prepare a report.

In the Agent Skills format, a skill has a SKILL.md file and can include scripts, references, and assets. A compatible harness makes skills discoverable. It can provide their names and descriptions first, load the selected skill's instructions when relevant, and read supporting files only when needed. This is called progressive disclosure.

The harness still needs the tools and environment that the skill expects. A skill describing a browser test does not install a browser. A script may need dependencies. Instructions may need changes when moved to another harness with different tools. Loading a skill also does not train the model or guarantee expertise.

Use the task prompt for what is specific today. Put a repeated procedure in a skill. Keep shared project rules in the instruction files your harness supports. This avoids copying a large, slightly different checklist into every request.

A skill says to run browser tests, but the harness has no browser tool. What should happen?

Instructions and Permissions Have Different Jobs

"Do not edit files outside this project" is a useful instruction. A filesystem boundary that blocks those writes is an enforced control. Reliable setups use both.

Permissions, sandboxing, and approval checks must be implemented by the harness, its tools, or the underlying environment. A prompt cannot enforce them by itself. The VS Code agent security documentation shows concrete examples of these controls.

The same distinction matters for information coming back from tools. A web page or repository file can contain text that looks like an instruction. Treat that text as task data unless it comes from a trusted instruction source. Keep access limited to what the work needs.

Which control actually prevents a tool from writing outside an allowed directory?

What Makes a Harness Useful?

More autonomy is useful only when the system can check its work and recover from problems. For a long task, look at how the harness handles:

  • Context: Keep relevant instructions and evidence available. Summarizing old steps can save space, but it can also lose details.
  • Progress: Save decisions and unfinished work so the next session can continue. Saved state still needs to be loaded into context.
  • Failures: Return errors clearly, use bounded retries, and avoid repeating actions with side effects blindly.
  • Limits: Stop at a time, cost, or step budget, or when the work needs user input. Stopping is not proof of success.
  • Verification: Inspect actual files, test results, or other outcomes. A confident final message is not enough.

When comparing setups, try the same real tasks and inspect both the outcomes and the execution traces. Keep the model fixed when you want to study a harness change. Change one prompt, skill, or tool setting at a time, and repeat tasks because individual runs vary.

Try this on a task you already do0/5 complete