A self-improving agent gets better at its own job by changing itself — its prompts, skills, memory or code. Self-improving software is a product that gets better because loops keep finding and proposing improvements to it, which a person approves and a coding agent builds. One improves the AI; the other improves your product.

Self-improving agents

In current usage, a self-improving agent is one that uses feedback from its own work to modify itself, so the next attempt goes better. The idea is old; it became practical once agents could read and edit code. A few well-known examples:

  • SICA, "A Self-Improving Coding Agent" (arXiv, 2025) — a coding agent that edits its own Python codebase between benchmark runs. Its authors report it improving from 17% to 53% on a subset of SWE-bench Verified.
  • Darwin Gödel Machine (Sakana AI and collaborators, 2025) — keeps an archive of agent variants, rewrites their code, and keeps the descendants that score better; reported gains on SWE-bench from 20% to 50%.
  • Skill-based self-improvement — lighter versions log corrections and lessons to files the agent reads next time, and promote recurring lessons into permanent memory. OpenClaw's community "self-improving-agent" skill works this way.
  • The research field — Stanford teaches it as a course, CS329A: Self-Improving AI Agents, covering verifiers, test-time compute, reinforcement learning, tool use and memory.

What these share: the thing being improved is the agent, and the measure of improvement is a benchmark or the agent's own task success.

Self-improving software

Self-improving software turns the same machinery outward. The agent stays the same; the product it watches keeps getting better. The pattern:

  1. Loops watch the product — its code, and signals from production like errors.
  2. Each loop owns one area (security, payments, reliability…) and wakes when something in that area changes.
  3. A finding becomes a proposal with a clear definition of done.
  4. A person approves or declines it.
  5. A coding agent builds the approved work; the merge feeds the next round.

The measure of improvement isn't a benchmark. It's fewer security holes, fewer double charges, faster pages, an AI feature that answers better.

The differences that matter

Self-improving agentSelf-improving software
What changesThe agent: its code, prompts, skills, memoryYour product's code
SignalIts own task results, benchmarksYour code changes, production errors, your approvals and declines
Who decidesUsually the agent, against a scoreYou, on every change
Main riskOptimising the benchmark instead of the goal; a modification you can't easily inspectNoise — too many proposals — if loops are too broad or forget what you declined
Who it's forPeople building agentsPeople building products

They can stack. A product's own AI features — a support agent, a retrieval pipeline — can be watched by a loop that proposes how to make that agent better. That's a self-improving product whose agent improves too, with a person approving each step.

How Tekk does self-improving software

Tekk improves your product, not itself — and only with your greenlight. It runs twelve loops on your codebase: Security, Reliability, Backend, Payments, Performance, Testing, React, Code quality, AI engineering, Observability, Alerts and Product analytics. The AI engineering loop is the one that looks at your own agents — harness, prompts, tools, context, memory, retrieval, evals and cost — and proposes the highest-leverage upgrade.

Each loop wakes on your changes or a Sentry alert, brings you one proposal, and waits. You accept it (it becomes a spec on your board) or decline it with a reason (the loops learn). Your coding agent builds the spec through MCP, the merged PR closes it, and that merge is the next change the loops read. Tekk itself never writes or ships code.

The step-by-step version: how to make your product self-improving.

Part of the Loop Engineering guide. Its other half is spec-driven development: the spec is how a loop knows it's done.