Loop engineering is the practice of designing the loop that prompts your AI coding agent for you — what wakes it, what job it gets, how its work is checked, and when it stops — instead of prompting the agent by hand, turn by turn. The engineer's job moves from writing prompts to writing loops.

Where the term came from

The phrase took off in June 2026. Boris Cherny, who leads Claude Code at Anthropic, said in a talk that he no longer prompts Claude — his job is to write loops (The New Stack). Days later Peter Steinberger, creator of OpenClaw, posted that developers shouldn't be prompting coding agents anymore but designing loops that prompt them. Addy Osmani wrote it up and gave it a structure in an essay, "Loop engineering", and Anthropic published "Getting started with loops" on the Claude blog at the end of June. By July it was trade-press news (ADTmag), and in August researchers published the first study of it, "Loop Engineering: Building Blocks, Adoption, and Impact".

Prompt, context, harness, loop

The arXiv paper frames loop engineering as the next level in a progression developers have already climbed:

  • Prompt engineering — phrasing one request well.
  • Context engineering — giving the agent the right files, docs and history.
  • Harness engineering — the agent's environment: tools, sandbox, permissions, state.
  • Loop engineering — the cycle around it: when a run starts, what it does, how it's checked, when it stops, and when a human is pulled in.

A useful way to tell the last two apart, from Data Science Dojo: the harness is infrastructure that rarely changes; the loop is policy you tune constantly. A production agent needs both.

The parts of a loop

Every write-up describes the same handful of parts, under slightly different names:

  • A trigger. Loops start on a schedule or on an event — a push, a new issue, an error alert — rather than when someone types.
  • A job. One task per run, narrow enough to finish.
  • State. What earlier runs did and decided, kept in files or a database so the next run doesn't start from zero.
  • A verifier. A separate check on the work — tests, a second agent, a human. Osmani's version: one sub-agent has the idea, a different one checks it.
  • A stop condition. Something a machine can evaluate: tests pass, the checklist is complete, the budget is spent.
  • An escalation point. The moment the loop hands a decision to a person.
  • A budget. Tokens and time per run, so a loop that goes wrong costs a little, not a lot.

Osmani's essay lists the building blocks on the tooling side: scheduled automations for discovery and triage, worktrees so parallel agents don't collide, skills that write down project knowledge, connectors (MCP) to reach other systems, and sub-agents for maker/checker separation.

The hard part is knowing when to stop

Generating code in a loop is cheap; judging it is not. That's the point most writing on loop engineering lands on: once an agent can run all night, the bottleneck is the verifier. A loop told to "improve the code" never knows when it's finished, and an agent asked to grade its own output tends to approve it.

So a good loop needs a stop condition someone wrote down in advance — a definition of done the loop can check. That is exactly what a good spec is: the change, the files, and acceptance criteria a test can evaluate. It's why loop engineering and spec-driven development are two halves of one idea; see loop engineering and spec-driven development.

Where loops pay off — and where they don't

The arXiv study mined thousands of public repositories and found autonomous loops in only a small fraction of them, mostly doing pull-request review and issue triage. Its authors also note that most claims about loops' value still rest on anecdotes. The criticism is fair: The Register called it the latest buzzword that still needs humans in the loop, and sceptics call it a while-loop around an LLM call.

Both things are true. A loop is simple to start and hard to make trustworthy. Loops pay off on narrow, repeatable questions with a checkable answer: did this change open a security hole, can this retry charge a customer twice, does this new endpoint have a test. They disappoint when the job is vague, the verifier is the same agent that did the work, or nobody reads the output.

Loop engineering on your own product

Most loop-engineering examples aim the loop at a task: build this feature, fix this issue, keep going until the tests pass. The same machinery can be aimed at a whole product: loops that keep watching the code and production, find what should change, and propose it. That's what makes software self-improving.

Tekk is loop engineering done for you on exactly that problem. It runs twelve loops on your codebase — Security, Reliability, Backend, Payments, Performance, Testing, React, Code quality, AI engineering, Observability, Alerts and Product analytics. Each one:

  • wakes on a trigger — your code changed, Sentry fired, or part of its area went unexamined too long;
  • reads a brief — what changed, what it proposed before, what you declined and why;
  • works one method from its shelf of skills, end to end;
  • stops at one proposal — or an honest "nothing found";
  • escalates to you — you greenlight or decline, and declines steer the next run;
  • closes when your coding agent's PR merges with Closes TEK-123 and the spec's checklist is done — and that merge is the next change the loops read.

Tekk never writes or ships code; your coding agent does, from a spec you approved. If you'd rather build the loops yourself, here's how to run loops on your product either way.

Part of the Loop Engineering guide. Its other half is spec-driven development: the spec is how a loop knows it's done.