Let’s talkLet’s talk

September 30th, 2026

Model or agent? A plain-English guide

Karolina Kosiorowska

Karolina Kosiorowska

9 mins

Model or agent? A plain-English guide

What separates a model from an agent, what the harness around the model does, and what system prompts, reasoning effort and model size change.

A model is a component, and an agent is a system built around it. The model only generates text. Everything else, like calling tools, keeping track of progress and deciding when the task is finished, is done by the software around the model. That software is called the harness. It explains why the same model can work well in one product and poorly in another.

Models and agents are easy to confuse because both get called AI. A chat window answers your question. An agent opens files, runs code and fixes the bug for you, and often it uses the exact same model underneath.

Below are the terms that come up once you look past the model: harness, system prompt, reasoning, effort and model size.

#The model: text in, text out

A large language model (LLM) like GPT, Claude or Gemini is a very large file of numbers. During training it read a huge amount of text and learned one skill: given some text, guess what comes next. You send it text, and it sends text back, one small piece of a word at a time.

That one skill goes a long way. It is enough to write an email, explain a tax rule or write a Python function. On its own, though, a model has hard limits:

  • It can't act. It can write the command to delete a file, but it can't run it.
  • It has no memory. Each request starts from zero.
  • It only knows what it was trained on. It can't look anything up.
  • It answers once and stops. It doesn't check whether its answer worked.

So when a chat app seems to "remember" what you said five messages ago, the software around the model is sending the whole conversation back with every new message. The model reads all of it fresh each time.

#The agent: a model in a loop

For a long time, "agent" was a buzzword with a hundred meanings. In September 2025, developer Simon Willison proposed a definition:

An LLM agent runs tools in a loop to achieve a goal.

Tools are things the model can ask to use: search the web, read a file, run a command. The loop means it doesn't stop after one answer.

The model looks at the goal and says, in text, "I want to search for X." It doesn't search. It only writes a request to search. The harness reads that request, runs the search and gives the results back. The model reads them and decides the next step. Maybe another search. Maybe "I'm done, here's the answer."

Anthropic's guide Building Effective Agents draws a useful line. In a workflow, a developer sets the steps in advance. In an agent, the model picks the steps as it goes. That lets it handle messier tasks, and it also gives it more ways to go wrong.

#The harness: everything that isn't the model

LangChain put the idea of a harness in one line in The Anatomy of an Agent Harness:

Agent = Model + Harness. If you're not the model, you're the harness.

The word comes from horses. A horse is strong, but without a harness it can't pull a cart. The harness around a model does the same job: it points the model's ability at real work.

A typical harness:

  • Runs the loop. It sends text to the model, reads the reply and decides whether to continue.
  • Executes tools, often in a closed-off space (a "sandbox") so a mistake can't break your computer.
  • Manages memory. A model can only read a limited amount of text at once (the "context window"), so the harness chooses what the model sees each turn.
  • Sets the rules: which tools are allowed, what needs a human's approval, when to stop.
  • Checks the work, for example by running the tests before the agent says it's finished.

Claude Code, Cursor, ChatGPT's agent mode and Devin are all harnesses. Several of them can run the same model and still behave very differently, because the harness around it is different.

In February 2026, LangChain tested their coding agent on a benchmark called Terminal Bench 2.0. They kept the model the same (GPT-5.2-Codex) and changed only the harness. They added better instructions, a self-check before finishing, and a way to spot when the agent kept repeating the same failed idea. The score went from 52.8% to 66.5%. That moved them from just outside the top 30 on the leaderboard into the top 5.

#The system prompt: instructions the model reads first

The system prompt is a block of instructions that the harness puts before your message in every conversation. It tells the model what role it plays, what it can and can't do, and which tools it has. You usually never see it, but the model reads it first, every time.

A system prompt might say:

  • "You are a support assistant for Company X. Only answer questions about our products."
  • "Today's date is 25 September 2026." (The model doesn't know the date on its own.)
  • "Always ask before deleting anything."

This is one of the main reasons the same model behaves differently in different products. Anthropic publishes the system prompts it uses in the Claude apps. They don't apply when developers call the model directly through the API. Ask Claude the same question in the app and through the API, and you are talking to the same model with different instructions.

For agents, the system prompt is also where each tool gets described. A lot of harness work is writing better instructions.

#Reasoning and effort: how long to think

A standard chat model starts writing its answer straight away. For a hard maths problem or a tricky bug, that leaves no room to plan, so an early mistake carries through to the end. Researchers found that showing the model worked examples, or simply adding "Let's think step by step" to the prompt, made it write out its steps before the answer. On maths and logic problems, the results got much better.

Reasoning models were trained to do this on their own, and to do it better. Before the final answer, they break the problem into steps, try an approach, notice a mistake and correct it. Under the hood it's still next-word prediction. The model just gets to write a rough draft first.

Thinking costs time and money, because providers charge per token (a small piece of a word). So most providers let you choose how much the model should think. OpenAI calls this setting reasoning effort and Anthropic calls it effort. Both use levels like low, medium and high.

EffortGood forTrade-off
LowQuick questions, summaries, simple editsFast and cheap, but can miss details
MediumMost everyday tasksA balance of speed and quality
HighHard bugs, maths, planning a big changeSlower and more expensive, but more careful

Higher isn't always better. Maximum effort on "what's the capital of Poland?" just wastes time. Good harnesses match the effort to the task. In the LangChain experiment above, part of the gain came from a "reasoning sandwich": very high effort for planning, lower effort for the routine middle, and very high effort again to check the work.

#Model size: what "billions of parameters" means

You've probably seen headlines like "a 70-billion-parameter model". Parameters are the numbers inside the model that were adjusted during training. MIT Technology Review describes them as the dials and levers of a planet-sized pinball machine. Each one is tiny, but together they hold everything the model learned.

GPT-3 had 175 billion parameters in 2020. The same MIT article estimates at least a trillion for Google's Gemini 3, though Google isn't saying. Most companies no longer publish exact numbers.

More parameters usually means more knowledge and better handling of complex ideas. It also means a slower and more expensive model. But size is no longer the whole story. In Meta's own benchmarks, the chat version of Llama 3 with 8 billion parameters beat Llama 2 with 70 billion. The difference is mostly training. Llama 3 was trained on 15 trillion tokens and Llama 2 on only 2 trillion. Meta also got better at turning a raw model into a chat assistant.

Size still sets the speed and the cost, which is why model families usually come in small, medium and large. Inside an agent, a small, fast model often handles the simple jobs, like sorting emails or pulling out a date. The big model is saved for the hard thinking.

One warning about words: "parameters" also means the settings you send with a request, like temperature (how random the answers are). Those don't change the model, only how it's used for that one request. Effort is one of these settings.

#Putting it all together

Let's follow one task: "The login page is broken. Find the bug and fix it."

  1. The harness starts a session with the system prompt: you're a coding assistant, here are your tools, always run the tests before you finish. It sets effort to high, because debugging is hard.
  2. The model reasons for a moment, then asks to read login.js.
  3. The harness sends back the file. The model spots a suspicious line and asks to run the tests.
  4. The harness runs them in a sandbox. Two tests fail, and the results go back to the model.
  5. The model writes a fix. The harness applies it and runs the tests again. This time they all pass.
  6. The model says it's finished, and the harness shows you a summary and the changed lines.

The model did all the thinking. The harness opened the file, ran the tests and kept track of every step. Without it, the model could only have described a possible fix and hoped it was right.

So when you compare AI tools, ask two questions: which model does it use, and what has been built around it? The second answer often explains more.


AKENA is an engineering studio working across AI, blockchain, and software infrastructure. We build the systems AI agents run on: agent-facing RPC and MCP endpoints, on-chain data pipelines, and the payment, identity, and metering layers in front of them. If you have picked a model and the agent built on it still falls short of the job, we should talk.

Let’s talk

Bring us your problem, we’ll help design the system. No hype, just engineering.