code-anything.com
Log inStart free

Agents

Agents that build, check and keep watch

One agent builds your app and repairs its own mistakes before you see them. Others you switch on yourself keep checking what you shipped. This page is the honest account of what each one does, what it may do alone, and where you stay in control.

// build repair: always on, every plan · standing agents: yours to install

Start here

What an “agent” actually means here

The word is doing two jobs on this site. Separating them is the whole point of this page.

An agent, in the general sense, is a model that is given a goal, a set of tools, and permission to use them in a loop until the goal is met or a budget runs out. On code-anything that description covers two quite different things, and conflating them is how people end up either over-trusting the automation or under-using it.

The first is the agent that builds your app. When you describe a change, it reads the codebase, plans, writes files inside a disposable sandbox, runs commands, and then checks its own output — typecheck, your test script, the production build, and a scan of the running dev server for runtime errors. If a check fails it gets the failure output back and repairs, up to three rounds inside one plan, replanning entirely if the plan rather than the code was wrong. Only a build that is green and in scope is accepted; anything else is held for you to look at, with the verdict and the reason attached. This agent is always on. You do not switch it on, configure it, or pay extra for it, and it runs on every build on every plan. The mechanics are set out on sandbox verification.

The second is a standing agent — an automation you install on a specific project that keeps running after the build has ended. A standing agent has a trigger (a cadence, or an event) and an action (check that a URL is up, or run a prompt against the project). These are the agents on the project’s Agents tab. Nothing installs one for you — they are yours to add, pause and delete — and a health-check agent and a prompt agent differ enormously in what they are allowed to touch. The rest of this page is mostly about them.

Everything below is what the system does today. Where a behaviour is opt-in, it says so. Where something needs a human, it says who and when. If you want the surrounding product — preview, deploy, repository, environment — start at the platform overview or the documentation.

Always on

Automatic repair inside every build

This is the real autonomous correction, and it happens whether or not you ever open the Agents tab.

INSIDE EVERY BUILD · NOTHING TO SWITCH ONtypechecktsc --noEmitteststhe repo suitebuildnpm run buildpreviewdev.log scanrepair rounds — up to 3 inside one planheld for your reviewacceptedgreen and in scope
The arrow that matters points backwards. A failed check is not your problem yet — it is the agent’s next turn, up to three rounds inside one plan.

Four gating checks, in order

Typecheck, then the test script, then the production build, then a scan of the dev-server log for runtime and compile errors. Each check records whether it passed and why it failed, so a past build stays inspectable.

Up to three repair rounds

A failing check hands the agent the failure output. It repairs and re-checks, up to three rounds inside one plan — and if that still does not get it green, it takes one full replan rather than retrying the same edit.

A behavioural and a visual pass

A second reviewer with none of the build’s history judges the diff against what you asked for, and clicks through the running preview. A separate visual pass crawls the rendered pages. Both can seed one more repair.

Accepted only when green and in scope

A build is accepted when the checks pass and the diff matches what was asked for. Anything else is held for your review with the verdict and the reason attached — it is never shipped broken.

verify — the gating checks, in order
typecheck   npm run typecheck (or tsc --noEmit)
tests       npm test --if-present
build       npm run build --if-present
preview     scan the dev-server log for runtime errors
✗ typecheck failed — repairing (round 1 of 3)
✓ typecheck
✓ tests
✓ build
✓ preview — no runtime errors in dev log
verdict: accept

What “held for review” means

A build ends with one of three verdicts: accepted, held for review, or rejected. Held means the agent finished but is not confident enough to call it done — a check it could not get green, or a diff that drifted outside what you asked for. You get the verdict, the reason, and the list of files it touched, and you decide.

A rejected run also fires a lifecycle event, so if you have a standing agent subscribed to run_rejected it can act on the failure. With a repository connected, accepted work reaches you as a pull request rather than a commit on your default branch.

How verification works
  • The four gating checks are the same on every plan; the build tier decides how many extra repair rounds the behavioural and visual passes get — none on a prototype build, one each on a standard build, two each at the highest tier
  • A JavaScript project with no TypeScript config, or a project with no test script, records that check as skipped rather than inventing a pass
  • If the behavioural or visual harness cannot be reached, that pass records as skipped and does not block your build — the four deterministic checks still gate it
  • A separate security scan runs when you Go Live, and a P0 finding blocks the publish outright

Opt-in

What runs continuously versus what runs on demand

Standing agents come in three shapes. They differ less in how they are triggered than in what they are permitted to touch.

Health-check agents

One bounded GET against your preview URL or your live URL, on a cadence you set. They read; they never write. They cost nothing and cannot change your project — and a project that has never been previewed or deployed has no URL to point one at yet.

Scheduled prompt agents

A standing instruction the agent runs against the project on a cadence — a nightly sweep, a weekly catalogue pass. Each run is a real build that can change your code, and it spends one runtime credit.

Event-triggered agents

The same prompt action, fired by something happening rather than by the clock: a failed deploy, a rejected run, or a business event your app reports through a signed endpoint.

AgentHow it firesWhat it does when it firesCost per run
Health checkOn a cadence — every 5 minutes at the fastestOne GET against your preview or live URL, capped at 8 seconds. Reads the status and the first 16 KB of the body. Changes nothing.Free — not metered
Scheduled promptOn a cadence — every 60 minutes at the fastestRuns your standing instruction as a full build against the project, verified like any other build.1 runtime credit
Lifecycle eventWhen a deploy fails, or a build run is rejectedSame as a scheduled prompt, but fired by the failure rather than by the clock.1 runtime credit
Business eventWhen your app reports one of seven commerce events over a signed endpointSame as a scheduled prompt. Idle until your app is actually wired to report the event.1 runtime credit

Health checks and prompt agents are genuinely different products sharing one scheduler — one observes, the other acts.

  • Every standing agent can be paused and resumed without deleting it, and its cadence and target are visible on the card
  • Preset agents exist for a storefront — sourcing, repricing, fulfilment, support, low-stock reorder — and installing one starts it right away, behind a confirmation that says exactly that
  • An event agent whose event nothing fires yet sits idle rather than pretending to work; the interface only offers lifecycle events that something really emits

Cadence

What is checked, and how often it is checked

There is no fixed monitoring interval to quote. There is a heartbeat, and a due-ness rule.

ONE HOUR · WHO IS ACTUALLY DUEheartbeatevery minutehealth checkevery 5 minhealth checkevery 15 minprompt agentevery 60 min0 min3060 min
The heartbeat is per minute; your agent’s cadence is whatever you set. The dots are where the two coincide — which is why there is no single monitoring interval to quote.

A scheduled trigger calls the control plane once a minute. That call does not run everything — it loads the enabled agents and asks the scheduler which are due. An interval agent is due when it has never run, or when the gap since its last run has reached its own cadence. So the heartbeat is per-minute; how often your agent runs is whatever cadence you gave it, floored at five minutes for a check and sixty for a prompt.

A health check does exactly one thing: a GET against the URL, with an eight-second deadline, reading the status and up to the first 16 KB of the body. That bound matters. It is why a check is free, why a dead target cannot stall the heartbeat for every other project, and why a check has no way to modify anything. The same heartbeat also sweeps stale build runs — a run still marked running after twenty minutes that has also gone fifteen minutes without writing a step is closed out as failed, so a build killed mid-flight does not sit in your list forever pretending to be in progress.

What is not checked is worth stating plainly. There is no synthetic user journey, no response-time budget, no uptime percentage, and no third-party status ingestion. A health check answers one question — is this URL serving something that is not obviously broken — and answers it honestly.

Failure modes

What happens when a check fails

Every run is recorded with a status of ok, alert or error. These are the conditions that produce an alert, and the severity each one carries.

ConditionHow it is detectedSeverityWhat the alert says
The URL does not respond at allConnection error, or no answer within the 8-second probe timeoutCritical“preview unreachable” / “deploy unreachable”, with the connection error attached
The URL answers with a server errorHTTP status 500 or aboveCritical“returned HTTP 5xx”, with the URL and the exact status
The URL answers with another failure statusAny other non-OK HTTP status — 404, 403, 502-adjacent redirect loopsWarning“returned HTTP 4xx”, with the URL and the exact status
The page loads but is showing an errorHTTP 200, but the first 16 KB of the body matches a known error signature — a Vite overlay, “internal server error”, “failed to compile”, “cannot find module”, a Next.js client-side exception pageWarning“is serving an error”, quoting the first matching line of the page
The agent itself brokeA prompt agent timed out against its 120-second wall clock, or threwCritical“failed to run”, with the reason — the agent is recorded as errored, never silently skipped
The agent could not finish its workA prompt agent stopped short — out of runtime credits, over the spend cap, or a run it could not get greenWarning“could not complete”, with the reason it gave — the agent pauses itself rather than failing quietly

Severity is derived from the failure, not configured — there are no thresholds for you to tune.

ONE GET · 8 SECONDS · THE FIRST 16 KBGETyour live URLreads only — never writesHTTP 200looks finefailed to compile16 KB — the scan stops herewarningis serving an error
A page can answer 200 and still be broken. This is the only failure the status code cannot tell you about — so the check reads the head of the body and quotes the line back to you.

The alert is recorded first

Before anything is delivered anywhere, the alert is written to the project alongside the run that produced it. That record is the source of truth, and it survives whatever happens to any delivery attempt.

It surfaces on the Agents tab

Open alerts appear at the top of the project’s Agents tab, newest first, marked by severity, each with the failing agent, the title and the detail. Dismissing one acknowledges it and clears it from the list; the record itself is kept.

And it can be emailed to the owner

Where alert email is configured for the workspace, the alert also goes to the address on the project owner’s profile — severity in the subject line, the failure detail in the body, and a link back to the project’s Agents tab. Delivery is best-effort and never blocks or fails the check.

what an alert email carries
[Alert] preview returned HTTP 502
🚨 CRITICAL

preview returned HTTP 502

GET https://your-project.preview.example → 502

Open the project's agents: https://code-anything.com/projects/…/agents

— code-anything agents (automated monitoring)

Alerts are delivered to the project owner’s email and shown in the app. There is no Slack, webhook or paging integration for alerts today — if you need one, say so on the enterprise page.

Autonomy

What an agent may do alone, and what needs your approval

Every tool call an agent makes is evaluated before it runs. The answer is allow, ask, or refuse — and some things are refused no matter who is asking.

EVERY TOOL CALL, EVALUATED BEFORE IT RUNSPOLICYreads and sandbox writesread_file edit_file execallowedthe sandbox is disposabledestructive commandsgit push --force rm -rfyou are askedunanswered for 3 min = deniedprotected paths.git .env *.pem id_rsarefusedapproval cannot unlock it3 min
The three outcomes are not three settings. One passes, one waits on you and closes itself if you never answer, and one cannot be unlocked by anybody — including you.
ActionWho decidesWhy
Read a file, search the tree, list files, screenshot the previewOn its ownPure reads. Allowed in every mode, including read-only runs.
Write, edit or patch a file inside the disposable sandboxOn its ownThe sandbox is throwaway, and every accepted build is checkpointed to version history first.
Run an ordinary shell command in the sandboxOn its ownInstalls, builds, test runs — bounded by the run’s step and spend budget.
Generate an image or a video into the projectOn its ownTreated as a file write; metered like any other model work.
Run a destructive commandAsks firstForce-push, rm -rf, curl piped into a shell, find -delete, git reset --hard, git clean -f, or a production deploy command. Escalated for sign-off even inside the sandbox — the run blocks until you answer.
Write to a protected pathNever — refusedThe .git and .github trees, any .env file, certificates and private keys, SSH identities, AWS credentials, named secrets files. Hard-denied in every mode; approval cannot unlock it.
Ship a change to your repositoryYou merge itWith a repo connected, the edit path opens a pull request you review and merge. Nothing is committed to your default branch behind your back.
Install a standing prompt agentYou opt inNothing installs one for you. Presets that arrive with a new project come switched off; one you install from the Agents tab starts immediately, behind a confirmation naming its cadence, its credit cost, and the fact that it edits code unattended.
Answer an escalated tool callYou have 3 minutesAn approval request that goes unanswered for three minutes is auto-denied and the run continues with that call refused — it never defaults to allowing the action.

An escalated call blocks the run until you answer it, and is denied — never quietly allowed — if you do not.

The honest warning about prompt agents

A prompt agent is the one genuinely unattended writer in the system. When it fires it runs a real build against your project with nobody reviewing it first, and it can overwrite edits you made by hand in the editor since its last run. That is exactly what it is for — and exactly why nothing installs one for you, why turning one on takes a confirmation that spells this out in those words, and why the product calls it advanced rather than offering it as a default.

A health-check agent has none of this exposure. It cannot write, cannot run commands, and cannot call the model at all.

Why your app cannot claim its own build broke

Business events arrive from your app through a per-project endpoint, signed with a token held in the encrypted vault and valid inside a five-minute window. Lifecycle events — deploy failed, run rejected — are our own observations and are fired in-process only.

The asymmetry is deliberate. If an app could assert its own lifecycle events over the wire, a compromised storefront could trigger unattended repair builds simply by lying about its health. Secrets and connections live on the project’s Environment surface, never in the agent’s reach.

Runbook

How to investigate an incident, in order

Six surfaces, each answering the next question. Nothing here needs a support ticket.

  1. 1

    Start at the open alertsagents tab

    The project’s Agents tab lists every alert that has not been acknowledged, newest first, colour-coded by severity. Each one carries the title, the body — the URL, the status, the matched error line — and the agent that raised it.

  2. 2

    Read that agent’s run historyhistory

    Open History on the agent and you get its last twenty runs with a status of ok, alert or error, a one-line summary and a timestamp. That tells you whether this is a first failure or the third in an hour.

  3. 3

    Open the build behind itverify + trace

    Every build has a Verify surface showing the gate as it ran — which check passed, which failed, and the accept-or-hold verdict — and a per-run trace: a span timeline with the model used, the tokens, the cost, and each check’s pass or fail attached to it. Nothing about the gate is hidden from you.

  4. 4

    Identify the release that changed thingsdeploy log

    Repository → Deploy keeps the deploy history: every Go Live, what shipped, when, and whether it failed. It is the fastest way to line a failure up against a release.

  5. 5

    Read the build that introduced itchanges

    Repository → Changes shows the recent runs with their verdict — accepted, held for review, or rejected — the reason, and the files each one touched. A held build tells you what the agent itself refused to ship.

  6. 6

    Rewind, or roll the canary backrecover

    Repository → Changes carries both restore paths: a list of build checkpoints you can rewind the sandbox to, and a durable version history that outlives the sandbox entirely. Both take a safety snapshot of the current state first, so the restore is itself reversible; both refuse while a build or a deploy is in flight; and where the project has a provisioned database, its data is snapshotted and restored alongside the code. If the release went out as a delivery flight, its gate rolls the canary cohort back on its own — by default the moment it records a single error, or once enough of the cohort rejects it.

There is no single rollback button, and we would rather say so than imply one. Recovery is the version history and rewind on Repository → Changes, and — for a staged release — the delivery flight’s own gate. Both are described step by step in the documentation.

Limits

The limits agents run inside

Real numbers from the running system. Where a limit exists to protect you rather than us, it says which.

LimitValueWhy it exists
HeartbeatEvery minuteA scheduled trigger calls the control plane once a minute; the engine then decides which agents are actually due.
Health-check cadence floorEvery 5 minutesPoliteness towards the target you are probing. You can set anything slower.
Prompt-agent cadence floorEvery 60 minutesEach prompt run is a metered build, so the floor caps a single agent at roughly 24 runs a day.
Health-check probe timeout8 secondsOne slow target can never stall the heartbeat for everything else.
Prompt-run wall clock120 secondsA hung build is aborted and the agent is recorded as errored rather than left hanging.
Prompt agents running at once3 per tickHealth checks stay fully parallel; only the expensive prompt builds are queued.
Run history kept per agentLast 20 runsStatus, summary and timestamp for each, in the History drawer on the Agents tab.
Open alerts listedUp to 20 unacknowledgedDismissing an alert acknowledges it; the row itself is kept.
Unanswered approval requestAuto-denied after 3 minutesA blocked run cannot wait forever — so it fails the call closed rather than open.
Model steps inside one work itemCeiling of 200A runaway backstop, not a spend limit. Spend is governed separately.
Stuck-loop detection3 identical calls, or 5 steps with no progressThe same tool call three times running, or five consecutive steps that change nothing, ends the attempt instead of burning your budget.

Cadence floors are ours, not the scheduler’s — the engine could fire every minute; a five-minute floor is politeness towards the URL you are probing.

Per-run budgets, by plan

PlanSpend ceiling per runModel steps per runBuilds at once
Free$0.50301
Starter$2.00603
Pro$5.001006
Team$10.0016012

Plan defaults for the run guard. Where the account’s credit balance is known, spend is bounded by what that balance can afford instead, and the stuck-loop detectors above become the practical ceiling. A run that hits a limit stops and reports rather than continuing to spend.

  • A prompt agent whose account is out of runtime credits pauses itself and records why, instead of failing silently
  • A prompt trigger is charged once per firing — a retried heartbeat cannot double-charge the same run
  • Standing workspace monitoring is a paid feature from the Pro plan upward; automatic build repair is on every plan, free included
Compare plans

FAQ

Questions about agents

The things people ask before they hand anything to an automation.

Two different things, and it is worth separating them. The first is the agent that builds your app: it plans, edits files in a disposable sandbox, verifies its own work and repairs what failed before you ever see the result. That one is always on and runs inside every build. The second is a standing agent — an automation you install on a project that keeps running after the build ends, either checking a URL on a cadence or running a prompt on a schedule or an event.

Build with an agent that checks its own work.

Automatic repair runs on every build, on every plan. Standing agents are yours to add when you want something watched — nothing starts watching until you say so.