Skip to slide
Chapter 1 · What Is an Agent Harness?
02 / 142

CHAPTER 01 · What Is an Agent Harness? · 1 / 6

Concept explanation

When you use Claude Code or OpenAI's Codex in your terminal, it feels like one thing: a clever program that understands your code. It is actually two very separate things glued together.

The first thing is the model. It runs on Anthropic's or OpenAI's servers. It takes in a big blob of text and produces more text. It cannot touch your files. It cannot run a command. It has no memory of yesterday. Left alone, it can only talk.

The second thing is the harness. It runs on your machine (or in a managed sandbox). It is the CLI you launched. Every turn, the harness gathers up your instructions, your project's files, the conversation so far, and the list of tools the model is allowed to use, then sends all of that to the model. When the model says "run npm test," the harness is what actually runs npm test, captures the output, and hands it back. The model reasons; the harness acts.

A good analogy: the model is a brilliant consultant on the phone who has never seen your office. The harness is the assistant in the room who reads documents aloud to the consultant, writes down what the consultant says to do, walks over and does it, and reports back what happened. The consultant is smart, but without the assistant, nothing in the room actually changes. And a careless assistant (one who reads the wrong documents, or does dangerous things without checking) makes even a brilliant consultant useless.

This split is why the Mervin Praison breakdown of Claude Code insists that "Claude Code is not a CLI that calls Claude. It is an agentic harness." The model reasons; the harness mediates every action. That separation is exactly what makes the system feel magical while staying debuggable, because once you can name the layers, you can reason about where things go wrong.

← → arrow keys work too