Let’s talkLet’s talk

October 1st, 2026

Everyone runs loops. Here is what's inside one.

Michał Śmiarowski

Michał Śmiarowski

10 mins

Everyone runs loops. Here is what's inside one.

A loop does the job you do by hand when you prompt an agent. Writing one means handing that job to a script with three parts, state in files, a prompt and a check. This is how each part works, and when a plain prompt is the better choice.

Everyone now says they run agents in loops: overnight, on their pull requests, sometimes on loops that write loops. Ask what a loop is and you get a different answer each time. This post builds one from start to finish. The example is a job most codebases have at some point: replacing a library. Say forty files in a project still use an old HTTP client. You want an agent to move them all to the new one. By the end, the loop is two short text files, a five-line prompt and about fifteen lines of shell.

#TL;DR

  • A loop does the job you now do by hand. Every time you read an agent's answer, decide it isn't done and send another message, you are acting as the loop. Writing one means handing that job to a script.
  • That script has three parts: files to remember what's been done, a prompt to run again, and a check to decide when to stop. The check works best as a program, because a model will tell you it's done when it isn't.
  • Inside each turn, the agent already goes round its own small loop of tool calls. That loop comes with the tool, and you never write it.

#The inner loop and the outer loop

The inner loop is one turn of the agent. Say we ask it to migrate just one file, src/api/users.ts. Each row is one step the agent takes:

StepAgent actionFile or commandWhat happened
1readsrc/api/users.tsfinds the calls to the old client
2editsrc/api/users.tsswaps the old client for the new one
3runpnpm test users1 test fails: missing auth header
4editsrc/api/users.tspasses the header through
5runpnpm test usersall tests pass
6reply"Migrated users.ts, tests pass."

A turn lasts until the model decides it's finished, no matter how many files that is. Here that happened after one file, because we asked for one. If you ask for the whole directory, one turn might cover six files, or it might stop after two and say it's done.

Nothing in the agent's code told it to run the tests twice, in steps 3 and 5. The model chose each step after it read the result of the last one. Phil Schmid put it in two sentences: "The loop is hardcoded. What the model does inside the loop is not." You don't write this loop. It comes with Claude Code, Cursor or whatever agent you use.

The outer loop exists because the model saying it's finished and the work being finished are two different things. The agent has replied, and one file is done with thirty-nine to go. Now something has to decide whether the migration is finished, and if it isn't, start the agent again. The agent can't see any of this. It gets a prompt, does its turn and stops. The simplest outer loop is the one Geoffrey Huntley named Ralph in mid-2025: the same prompt, run again and again, each time in a fresh session. It is what most people mean when they say they run a loop.

#When you prompt by hand, you are the outer loop

You write a message and the agent does its turn. You read the result, decide it isn't done, and write the next message: "three files still import the old client, do those too." You also keep the goal in your head between messages. Writing a loop means taking yourself out of that job.

Prompting by handA loop
Who starts the next runyoua script, a timer, or /goal
Who decides it's good enoughyoutests, a build, or another model
Where the memory livesyour head and the chatfiles on disk, such as a plan, a spec and git

#Building the outer loop in three steps

By hand, you would probably ask for all forty files in one message. The agent does some of them and says it's done. Then you find the files it skipped or the tests it broke, and you ask again. The outer loop takes over that checking and asking. It needs files that remember which of the forty are done, a prompt that migrates the next one, and a check that knows when all forty are done.

Step 1: write the spec and the plan

Huntley's original Ralph is one line of bash:

while :; do cat PROMPT.md | claude-code ; done

It looks like a joke, but it works because every run starts with an empty context, and models get worse as the context grows. In a June 2026 study, models with a long context often gave up early or settled for an answer they weren't sure of. A fresh start avoids that. The cost is that everything the next run needs has to be on disk.

The spec, specs/migration.md, is what you would tell a colleague before they started. It says how the new client is used and what must not change:

Replace every import of old-http-client in src/api with ~/lib/http.

- client.get(url, { params }) becomes http.get(url, { query })
- The new client throws on non-2xx responses. Catch HttpError where
  the old code checked res.status.
- Keep the auth header: pass headers: authHeaders() on every call.
- Don't change the signatures of exported functions.
- Don't edit test files.

The plan, fix_plan.md, is the to-do list. One command creates it. It finds every file that still uses the old client and writes each one as an unchecked box:

grep -rl "old-http-client" src/api | sed 's/^/- [ ] /' > fix_plan.md

After a few runs it looks like this:

- [x] src/api/users.ts
- [x] src/api/orders.ts
- [ ] src/api/payments.ts
- [ ] src/api/invoices.ts
...
- [ ] payments.ts retries on 503, check the new client does too

Both files are plain Markdown. Nothing makes the agent follow the checkboxes except the prompt. It ticks a file when it's done, and it adds a line when it finds something new, like the retry note at the bottom. You write the spec and create the plan once, and after that the agent keeps the plan up to date. A grep can build this plan because the job is mechanical. For less mechanical work, write the plan by hand, or ask the agent to draft it and read it before the loop starts.

Step 2: write the prompt

The prompt runs once per pass, so it describes one pass and not the whole project:

Read specs/migration.md and fix_plan.md.
Pick the next unchecked file in fix_plan.md. Migrate only that file.
Run the tests for that file. If they fail, fix the code, not the tests.
Tick the file in fix_plan.md and add anything new you found as a new item.
Commit with a message that names the file.

Every line has a job. The first gives back the context that the restart threw away. The second keeps each pass small enough to review. The third is there because an agent that can't make a test pass will often change the test instead of the code. The last two leave the files ready for the next run, which starts knowing nothing.

The prompt shouldn't decide when the loop stops. You could build a loop that ends when the agent writes "DONE", but then the agent decides when its own work is finished.

Step 3: write the check and the loop

Huntley's one-liner has no way to stop. while : runs until you press Ctrl-C or the invoice arrives. You could let the agent decide when it's done, but then it is grading its own work. A program is better, because a process that returns 0 or 1 can't be talked into anything. For the migration, the check is a short script:

#!/bin/sh
# Done when no file uses the old client, the tests pass,
# and no test file has been touched.
! grep -rq "old-http-client" src/api &&
  pnpm test &&
  git diff --quiet main -- '*.test.ts'

And the outer loop becomes:

for i in $(seq 1 60); do
  claude -p "$(cat PROMPT.md)" \
    --permission-mode acceptEdits \
    --allowedTools Read Edit "Bash(pnpm test:*)" \
    --max-budget-usd 1 \
    >> loop.log
  ./check.sh && break
done

Sixty runs at most, and the loop ends early only when the script says so, whatever the agent writes in loop.log.

Without these flags, claude -p has no permission to edit files or run commands. --allowedTools gives the agent only what the prompt needs, and --max-budget-usd limits what one run can spend. The script also never reads fix_plan.md, because an agent can tick a box for a file it didn't finish. When the loop ends, read the diff and not only the log. The log shows what the agent says it did, and the diff shows what it did.

An agent that can't meet a condition can also delete it from check.sh, because Edit lets it change any file. To catch that, save a hash of check.sh, the spec and the prompt outside the repo before the first run, and have the loop stop if it changes.

#Ready-made outer loops

You don't have to write this loop by hand every time, because agents now come with loops of their own. Anthropic's Claude Code has two built in. With them, the tool runs the loop and keeps the state, and you only write the prompt and the condition.

/goal runs until a condition holds. The whole migration fits in one line:

/goal no file in src/api imports old-http-client, pnpm test exits 0, and no test file is modified

Here your one sentence is the prompt and the check at the same time. After every turn, a small, fast model reads the conversation, meaning your messages, the agent's replies and the output of the commands it ran, and decides whether the condition holds. It doesn't run commands itself, so write the condition in a way that the agent has to prove with a command's output.

That makes the check a model again. It is a different model from the agent, so the agent isn't grading itself. But it still judges the conversation, and the agent can write a convincing one. The goal also runs in one session, so the context grows with every turn instead of starting fresh.

/loop runs on a timer instead:

/loop 10m check CI on my open pull request and fix anything that failed

Here the outer loop is a clock, and the check is up to you or the prompt. Recurring tasks expire after seven days, so a forgotten loop can't run forever.

#When a prompt is the better choice

All of this depends on a program that can say "done". A lot of work has no such program. Nobody can write an exit code for "the onboarding flow feels clear" or "the error messages make sense to a new user". If you put a loop on work like that, it will stop when the model decides it's finished or when your limit runs out, and neither tells you anything about quality.

That kind of work is better done by prompting by hand. The same goes for work where you don't know yet exactly what you want, because a loop can only follow the goal it got at the start. Even when work has a check, the loop trusts that check completely. If your tests miss a case, the loop will accept any code that passes them.

AKENA is an engineering studio working across AI, blockchain, and software infrastructure. We build the systems AI agents run on: agent-facing RPC and MCP endpoints, on-chain data pipelines, and the payment, identity, and metering layers in front of them. If you're handing long-running work to agents and need checks they can't talk their way past, we should talk.

Let’s talk

Bring us your problem, we’ll help design the system. No hype, just engineering.