A self-improving agent gets better at its own job by changing itself — its prompts, skills, memory or code. Self-improving software is a product that gets better because loops keep finding and proposing improvements to it, which a person approves and a coding agent builds. One improves the AI; the other improves your product.
Self-improving agents
In current usage, a self-improving agent is one that uses feedback from its own work to modify itself, so the next attempt goes better. The idea is old; it became practical once agents could read and edit code. A few well-known examples:
- SICA, "A Self-Improving Coding Agent" (arXiv, 2025) — a coding agent that edits its own Python codebase between benchmark runs. Its authors report it improving from 17% to 53% on a subset of SWE-bench Verified.
- Darwin Gödel Machine (Sakana AI and collaborators, 2025) — keeps an archive of agent variants, rewrites their code, and keeps the descendants that score better; reported gains on SWE-bench from 20% to 50%.
- Skill-based self-improvement — lighter versions log corrections and lessons to files the agent reads next time, and promote recurring lessons into permanent memory. OpenClaw's community "self-improving-agent" skill works this way.
- The research field — Stanford teaches it as a course, CS329A: Self-Improving AI Agents, covering verifiers, test-time compute, reinforcement learning, tool use and memory.
What these share: the thing being improved is the agent, and the measure of improvement is a benchmark or the agent's own task success.
Self-improving software
Self-improving software turns the same machinery outward. The agent stays the same; the product it watches keeps getting better. The pattern:
- Loops watch the product — its code, and signals from production like errors.
- Each loop owns one area (security, payments, reliability…) and wakes when something in that area changes.
- A finding becomes a proposal with a clear definition of done.
- A person approves or declines it.
- A coding agent builds the approved work; the merge feeds the next round.
The measure of improvement isn't a benchmark. It's fewer security holes, fewer double charges, faster pages, an AI feature that answers better.
The differences that matter
| Self-improving agent | Self-improving software | |
|---|---|---|
| What changes | The agent: its code, prompts, skills, memory | Your product's code |
| Signal | Its own task results, benchmarks | Your code changes, production errors, your approvals and declines |
| Who decides | Usually the agent, against a score | You, on every change |
| Main risk | Optimising the benchmark instead of the goal; a modification you can't easily inspect | Noise — too many proposals — if loops are too broad or forget what you declined |
| Who it's for | People building agents | People building products |
They can stack. A product's own AI features — a support agent, a retrieval pipeline — can be watched by a loop that proposes how to make that agent better. That's a self-improving product whose agent improves too, with a person approving each step.
How Tekk does self-improving software
Tekk improves your product, not itself — and only with your greenlight. It runs twelve loops on your codebase: Security, Reliability, Backend, Payments, Performance, Testing, React, Code quality, AI engineering, Observability, Alerts and Product analytics. The AI engineering loop is the one that looks at your own agents — harness, prompts, tools, context, memory, retrieval, evals and cost — and proposes the highest-leverage upgrade.
Each loop wakes on your changes or a Sentry alert, brings you one proposal, and waits. You accept it (it becomes a spec on your board) or decline it with a reason (the loops learn). Your coding agent builds the spec through MCP, the merged PR closes it, and that merge is the next change the loops read. Tekk itself never writes or ships code.
The step-by-step version: how to make your product self-improving.
Part of the Loop Engineering guide. Its other half is spec-driven development: the spec is how a loop knows it's done.
