How it works
Loops and their skills
The twelve loops Tekk can run on your project, the skills each one carries, and how a run picks the one method it will work with.
A loop is not one big prompt. Each loop is a shelf of skills: a short file that says what the loop is and what counts as a finding in its domain, plus two to five skills, each a self-contained method for one kind of question inside that domain. A run opens exactly one skill and works it end to end.
This page lists what is on each shelf. The skills are written in the open Agent Skills format — a SKILL.md per method, with reference material bundled beside it where a method needs one — and they are the same files a Tekk run mounts when it works on your code.
Work your board from your agent
The four skills for driving Tekk itself — sweep the proposal inbox, blitz specs into parallel coding sessions, triage the board against git, and write specs that can actually close — install as one plugin that also connects the Tekk MCP server:
/plugin marketplace add guccikudo92615/tekk-skills
/plugin install tekk@tekk
For Codex, Cursor and other agents, take the skills and add the MCP server as described in the MCP server docs:
npx skills add guccikudo92615/tekk-skills -a codex
Those four are how you operate Tekk. The review skills documented below are a different thing: they are the methodology the loops apply to your code, they run inside Tekk rather than in your terminal, and they are not distributed separately.
What the loops add on top of any single method is everything around it: deciding which loop is worth waking and when, assembling the brief that says what changed and what you have already declined, choosing the skill against that skill’s own track record on your project, and turning a finding into a spec on your board.
How a run chooses a skill
When a loop wakes, nothing pre-selects its method. The run is handed its brief — what changed since it last ran, what it and every other loop proposed recently and what you did with each, how each of its skills has performed on your project, and a snapshot of your product’s measured scale — and then it opens the shelf. Two things weigh the choice:
- What the change is actually about. A diff that touches a webhook receiver is a question for
retry-safetyorwebhooks-and-fulfillment, notrender-performance. - The track record. A skill whose proposals keep landing has earned another turn. One whose proposals you keep declining is not owed a turn just because it has gone longest without one.
The run opens with a short bearings statement — what this run is about, which skill it is opening and why, and what it will check — then reads the code with that one method. If reading changes its mind it may switch once. Every run records the skill it actually worked, so the next run knows which of the loop’s methods has really been exercised.
One run produces at most one proposal, and a run that finds nothing is a complete run — it is recorded as an honest negative so the loop does not re-audit the same ground tomorrow.
The loops
Twelve loops are available today. Each one is opt-in per workspace and works on your repository alone; where a loop gets sharper with a connected tool, the table says so.
Security
Audits recent code changes for security holes — auth gaps, exposed data, injection, tenant leaks — and proposes fixes. It audits like an attacker who wants free service or another tenant’s data, not like a checklist, and holds every finding to a traced request rather than a pattern match.
| Skill | What it audits |
|---|---|
auth-and-access | Who is allowed to do what — ownership and role checks on every route, IDOR, check-then-act races, entitlement enforcement — and the signup, login, reset and session flows that hand out identity in the first place. |
tenant-isolation | Whether one customer can reach another’s data: row-level security coverage, the service-role bypass, tenant scope taken from the session rather than the request, and leaks through joins, caches, exports and admin paths. |
money-and-webhooks | Whether someone can get paid features without paying, and whether inbound provider webhooks can be forged, replayed, or pointed at another account. |
injection-and-input | Untrusted input reaching a dangerous sink — SQL, shell, SSRF, path traversal, unsafe deserialization — and the agent-era version of the same bug: prompt injection, unverified tool authorization, model output trusted downstream. |
secrets-and-crypto | How credentials are stored, exposed and rotated — keys in code, secrets in logs, error bodies and client bundles — plus password hashing, token randomness and the supply-chain surface underneath. |
Reliability
Audits your code for the ways it breaks under real-world failure — a dependency that times out with no deadline, a retry that runs a charge twice, a swallowed error that looks like success, an unbounded operation that exhausts memory — and proposes the fix. Sharper with Sentry connected; works without it.
| Skill | What it audits |
|---|---|
dependency-boundary | Every call that leaves the process — HTTP, database, cache, queue — for deadlines, retry policy, backoff and blast-radius containment. |
retry-safety | State changes that can run more than once — a redelivered webhook, a re-run job, a double-submit — for the dedupe that makes the second execution harmless. |
resource-limits | What runs out under load: pool starvation, a connection held across an await, unbounded queries and accumulators, uploads with no size limit, leaked handles. |
failure-handling | What the code does once something has already failed — empty catches, errors logged and continued past on a write, a failure that returns success, half-applied multi-step work. |
Backend
Audits everything behind your API for the bugs a founder can’t see — unvalidated input, a retry that charges twice, two requests racing into a duplicate, a swallowed failure, a query with no index, a migration that could break the next deploy — and proposes the fix.
| Skill | What it audits |
|---|---|
api-correctness | The trust boundary of a request handler: input validated where it enters, errors that fail loudly, and the defaults that decide behaviour under load — pagination caps, timeouts, body limits, rate limits, CORS. |
idempotency-and-jobs | Anything that can happen twice or interleave — non-idempotent writes under retry, check-then-act races, multi-step writes with no transaction, queue and cron consumers with no retry, dead-letter or lock discipline. |
query-and-index | How the app reads — N+1 queries, filters on unindexed columns, unbounded scans, over-fetching, and ORM calls that emit something pathological. |
schema-and-migrations | The shape of the data and how it changes — migrations that must survive a non-atomic deploy, constraints that keep bad rows out, data types, pool configuration. |
Payments
Audits your billing code for the bugs that lose money — a paid signup that provisions nothing, a retry that double-charges, access that’s never revoked on cancel, a drifted price id — and proposes the fix.
| Skill | What it audits |
|---|---|
webhooks-and-fulfillment | Whether every money-critical provider event is handled and fulfilment runs exactly once — at-least-once delivery with no dedupe, out-of-order updates, the trial path that provisions nothing. |
entitlements | What the customer can actually do after the money moves — every grant with a matching revoke, cancellation and failed renewal reaching a defined state, access decided from real subscription state rather than a stale cache. |
billing-config | The configuration that decides what gets sold — price and plan ids split between env and code, test-mode versus live-mode keys, plan-to-feature maps that drift from the provider. |
Performance
Hunts slow queries, slow endpoints and latency spikes and proposes targeted speedups — and the infrastructure spend that buys nothing: a poll that should be a webhook, an always-on service, needless egress. Sharper with Sentry connected.
| Skill | What it audits |
|---|---|
query-latency | How the app reads under load — a query inside a loop, a per-row lookup that should be one batched query, unindexed queries on a hot path, pagination that scans the whole table. |
hot-path-blocking | What a request waits on — synchronous I/O or heavy CPU in a handler, an external call with no timeout, work that should be queued, a response built without a cache. |
infra-spend | The machine bill — a cron where an event would do, a schedule far tighter than its inputs change, an always-on service that could be on-demand, paid services still wired but unused. |
algorithmic-cost | The code itself — a nested loop turning linear into quadratic, a list scanned where a map would index it, a whole collection loaded to count it, pure work redone instead of memoized. |
Testing
Finds risky changed behaviour with no tests, tests that fail at random, and tests that cost more time than the confidence they buy — and proposes the fix.
| Skill | What it audits |
|---|---|
coverage | Changed behaviour and the tests it does not have — new branches and error paths, bug fixes with no regression test, contracts whose tests no longer assert them. |
flakiness | Non-determinism — fixed sleeps instead of waits, state shared across tests, order dependence, real clocks, unseeded randomness, real network I/O — and a stabilization or an honest quarantine. |
suite-speed | Wall-clock the suite does not need to spend — real I/O that should be faked, heavy per-test setup that could be shared, serial-but-independent cases — keeping every check. |
React
Checks React components for bugs and wasted renders — stale state, effect misuse, needless re-renders — and proposes fixes.
| Skill | What it audits |
|---|---|
state-and-effects | Where state lives and when effects run — index keys, missing or unstable effect dependencies, state that should be derived, effects doing an event handler’s job, async races. |
render-performance | What a component costs to render and how widely it re-renders others — unstable props and context values, over-broad subscriptions, long unvirtualised lists, updates that block typing. |
markup-safety | What the rendered markup does to the user — unsanitised HTML injected into the DOM, and controls without an accessible name, keyboard path or correct semantics. |
Code quality
Cleans up what an AI-written codebase accretes — dead code and pointless wrappers, duplicate helpers, modules that grew into grab-bags, and references left behind by a rename — without changing what the app does.
| Skill | What it audits |
|---|---|
architecture | The module level — layering violations, new coupling and circular imports, logic duplicated instead of reused, vendor specifics leaking through a seam. |
propagation | What a rename left behind — every place still using the old form of a value that changed identity: env vars, constants, config keys, API routes, and the docs that cite them. |
deslop | The statement level — pass-through wrappers, single-use helpers, dead and speculative code, comments restating the code, near-duplicate logic. |
AI engineering
Audits how your product builds with LLMs — agent harness, prompts and tools, context, memory, retrieval, evals and what it all costs to run — and proposes the highest-leverage upgrade.
| Skill | What it audits |
|---|---|
harness | The agent loop itself — whether the “agent” is a real harness with tool use, state and error recovery, and whether its prompts, tools and untrusted-input handling hold up in production. |
context-engineering | What fills the model’s context and whether any of it is dead weight — memory written but never read, prompts demanding facts nothing supplies, blocks that grew past their usefulness. |
retrieval | How knowledge gets into the model’s context — whether retrieval is the right shape at all, and whether it surfaces the right material or merely the nearest. Ships a catalog of retrieval methods. |
evals | Whether you can change a prompt, model or retrieval setting and know it improved things, and whether production misbehaviour is diagnosable. |
llm-spend | What the product pays to run its LLM calls — agent turn tax, context bloat, missing prompt caching, an oversized model for a bounded task, retries with no ceiling. |
Observability
Finds parts of production you can’t see — unwatched services, silent jobs, vanished crashes, swallowed errors, untraced critical paths — and proposes the missing check, log, or span.
| Skill | What it audits |
|---|---|
uptime-and-heartbeats | Unattended work with nothing watching it — a service with no health endpoint or a shallow always-green one, a scheduled job with no check-in. |
crash-reporting | Every application entrypoint for error reporting — a server, worker or client runtime with no error-tracking SDK, an init that misses unhandled rejections, an SDK with no release or environment. |
log-quality | What the code leaves behind when it fails — an error branch with no record, unstructured messages with no correlation id, levels that bury real failures, secrets written into a log line. |
tracing-and-metrics | Whether flight-critical paths — payments, auth, outbound calls, queue work — emit spans and metrics that record success, failure and latency. |
Alerts
Finds critical paths with no alert — payment failures, broken webhooks, auth errors — and proposes the exact Sentry rule to catch them. Needs Sentry connected.
| Skill | What it audits |
|---|---|
coverage-gaps | The critical paths that exist in the code minus the alert rules that exist in the project, and the missing rule with a concrete condition and threshold. |
rule-quality | Alert rules that already exist — thresholds too tight to survive normal traffic or too loose to ever fire, duplicate rules, alerts with no owner or route, coverage that drifted after the code moved. |
Product analytics
Checks whether analytics is even wired up and, if not, sets you up — what to track, where, and a tracking plan — so you can see your funnel. Sharper with PostHog connected.
| Skill | What it audits |
|---|---|
event-coverage | The product’s real user journeys — signup, the activation moment, the core repeated action, conversion — and whether anything captures them, with the specific capture calls and where they go. |
event-hygiene | Whether what is captured can be used — an SDK really initialized and firing, consistent machine-parseable names, the properties a question needs, no PII sent as a property, a tracking plan rather than a pile. |
Where the seams are
Several skills sound alike from the outside, and the shelves are deliberate about who owns what:
- Retries.
retry-safety(reliability) owns the dedupe that makes a second execution harmless;idempotency-and-jobs(backend) owns the transaction and lock discipline around jobs and webhook handlers;webhooks-and-fulfillment(payments) owns the money-critical events specifically. - Queries.
query-and-index(backend) is about correctness and the index that should exist;query-latency(performance) is about what the same query costs on a hot path. - Webhooks and money.
money-and-webhooks(security) asks whether an attacker can forge or replay one;webhooks-and-fulfillment(payments) asks whether an honest one is handled exactly once. - Accessibility.
markup-safety(react) keeps the accessibility half of the former UX loop.
A finding on the wrong shelf is still raised; the seam decides which loop’s standard it is held to, not whether you hear about it.
Last updated September 13, 2026