Agents
Agents that build, check and keep watch
One agent builds your app and repairs its own mistakes before you see them. Others you switch on yourself keep checking what you shipped. This page is the honest account of what each one does, what it may do alone, and where you stay in control.
// build repair: always on, every plan · standing agents: yours to install
Start here
What an “agent” actually means here
The word is doing two jobs on this site. Separating them is the whole point of this page.
An agent, in the general sense, is a model that is given a goal, a set of tools, and permission to use them in a loop until the goal is met or a budget runs out. On code-anything that description covers two quite different things, and conflating them is how people end up either over-trusting the automation or under-using it.
The first is the agent that builds your app. When you describe a change, it reads the codebase, plans, writes files inside a disposable sandbox, runs commands, and then checks its own output — typecheck, your test script, the production build, and a scan of the running dev server for runtime errors. If a check fails it gets the failure output back and repairs, up to three rounds inside one plan, replanning entirely if the plan rather than the code was wrong. Only a build that is green and in scope is accepted; anything else is held for you to look at, with the verdict and the reason attached. This agent is always on. You do not switch it on, configure it, or pay extra for it, and it runs on every build on every plan. The mechanics are set out on sandbox verification.
The second is a standing agent — an automation you install on a specific project that keeps running after the build has ended. A standing agent has a trigger (a cadence, or an event) and an action (check that a URL is up, or run a prompt against the project). These are the agents on the project’s Agents tab. Nothing installs one for you — they are yours to add, pause and delete — and a health-check agent and a prompt agent differ enormously in what they are allowed to touch. The rest of this page is mostly about them.
Everything below is what the system does today. Where a behaviour is opt-in, it says so. Where something needs a human, it says who and when. If you want the surrounding product — preview, deploy, repository, environment — start at the platform overview or the documentation.
Always on
Automatic repair inside every build
This is the real autonomous correction, and it happens whether or not you ever open the Agents tab.
Four gating checks, in order
Typecheck, then the test script, then the production build, then a scan of the dev-server log for runtime and compile errors. Each check records whether it passed and why it failed, so a past build stays inspectable.
Up to three repair rounds
A failing check hands the agent the failure output. It repairs and re-checks, up to three rounds inside one plan — and if that still does not get it green, it takes one full replan rather than retrying the same edit.
A behavioural and a visual pass
A second reviewer with none of the build’s history judges the diff against what you asked for, and clicks through the running preview. A separate visual pass crawls the rendered pages. Both can seed one more repair.
Accepted only when green and in scope
A build is accepted when the checks pass and the diff matches what was asked for. Anything else is held for your review with the verdict and the reason attached — it is never shipped broken.
typecheck npm run typecheck (or tsc --noEmit)
tests npm test --if-present
build npm run build --if-present
preview scan the dev-server log for runtime errors✗ typecheck failed — repairing (round 1 of 3)
✓ typecheck
✓ tests
✓ build
✓ preview — no runtime errors in dev log
verdict: acceptWhat “held for review” means
A build ends with one of three verdicts: accepted, held for review, or rejected. Held means the agent finished but is not confident enough to call it done — a check it could not get green, or a diff that drifted outside what you asked for. You get the verdict, the reason, and the list of files it touched, and you decide.
A rejected run also fires a lifecycle event, so if you have a standing agent subscribed to run_rejected it can act on the failure. With a repository connected, accepted work reaches you as a pull request rather than a commit on your default branch.
- The four gating checks are the same on every plan; the build tier decides how many extra repair rounds the behavioural and visual passes get — none on a prototype build, one each on a standard build, two each at the highest tier
- A JavaScript project with no TypeScript config, or a project with no test script, records that check as skipped rather than inventing a pass
- If the behavioural or visual harness cannot be reached, that pass records as skipped and does not block your build — the four deterministic checks still gate it
- A separate security scan runs when you Go Live, and a P0 finding blocks the publish outright
Opt-in
What runs continuously versus what runs on demand
Standing agents come in three shapes. They differ less in how they are triggered than in what they are permitted to touch.
Health-check agents
One bounded GET against your preview URL or your live URL, on a cadence you set. They read; they never write. They cost nothing and cannot change your project — and a project that has never been previewed or deployed has no URL to point one at yet.
Scheduled prompt agents
A standing instruction the agent runs against the project on a cadence — a nightly sweep, a weekly catalogue pass. Each run is a real build that can change your code, and it spends one runtime credit.
Event-triggered agents
The same prompt action, fired by something happening rather than by the clock: a failed deploy, a rejected run, or a business event your app reports through a signed endpoint.
| Agent | How it fires | What it does when it fires | Cost per run |
|---|---|---|---|
| Health check | On a cadence — every 5 minutes at the fastest | One GET against your preview or live URL, capped at 8 seconds. Reads the status and the first 16 KB of the body. Changes nothing. | Free — not metered |
| Scheduled prompt | On a cadence — every 60 minutes at the fastest | Runs your standing instruction as a full build against the project, verified like any other build. | 1 runtime credit |
| Lifecycle event | When a deploy fails, or a build run is rejected | Same as a scheduled prompt, but fired by the failure rather than by the clock. | 1 runtime credit |
| Business event | When your app reports one of seven commerce events over a signed endpoint | Same as a scheduled prompt. Idle until your app is actually wired to report the event. | 1 runtime credit |
Health checks and prompt agents are genuinely different products sharing one scheduler — one observes, the other acts.
- Every standing agent can be paused and resumed without deleting it, and its cadence and target are visible on the card
- Preset agents exist for a storefront — sourcing, repricing, fulfilment, support, low-stock reorder — and installing one starts it right away, behind a confirmation that says exactly that
- An event agent whose event nothing fires yet sits idle rather than pretending to work; the interface only offers lifecycle events that something really emits
Cadence
What is checked, and how often it is checked
There is no fixed monitoring interval to quote. There is a heartbeat, and a due-ness rule.
A scheduled trigger calls the control plane once a minute. That call does not run everything — it loads the enabled agents and asks the scheduler which are due. An interval agent is due when it has never run, or when the gap since its last run has reached its own cadence. So the heartbeat is per-minute; how often your agent runs is whatever cadence you gave it, floored at five minutes for a check and sixty for a prompt.
A health check does exactly one thing: a GET against the URL, with an eight-second deadline, reading the status and up to the first 16 KB of the body. That bound matters. It is why a check is free, why a dead target cannot stall the heartbeat for every other project, and why a check has no way to modify anything. The same heartbeat also sweeps stale build runs — a run still marked running after twenty minutes that has also gone fifteen minutes without writing a step is closed out as failed, so a build killed mid-flight does not sit in your list forever pretending to be in progress.
What is not checked is worth stating plainly. There is no synthetic user journey, no response-time budget, no uptime percentage, and no third-party status ingestion. A health check answers one question — is this URL serving something that is not obviously broken — and answers it honestly.
Failure modes
What happens when a check fails
Every run is recorded with a status of ok, alert or error. These are the conditions that produce an alert, and the severity each one carries.
| Condition | How it is detected | Severity | What the alert says |
|---|---|---|---|
| The URL does not respond at all | Connection error, or no answer within the 8-second probe timeout | Critical | “preview unreachable” / “deploy unreachable”, with the connection error attached |
| The URL answers with a server error | HTTP status 500 or above | Critical | “returned HTTP 5xx”, with the URL and the exact status |
| The URL answers with another failure status | Any other non-OK HTTP status — 404, 403, 502-adjacent redirect loops | Warning | “returned HTTP 4xx”, with the URL and the exact status |
| The page loads but is showing an error | HTTP 200, but the first 16 KB of the body matches a known error signature — a Vite overlay, “internal server error”, “failed to compile”, “cannot find module”, a Next.js client-side exception page | Warning | “is serving an error”, quoting the first matching line of the page |
| The agent itself broke | A prompt agent timed out against its 120-second wall clock, or threw | Critical | “failed to run”, with the reason — the agent is recorded as errored, never silently skipped |
| The agent could not finish its work | A prompt agent stopped short — out of runtime credits, over the spend cap, or a run it could not get green | Warning | “could not complete”, with the reason it gave — the agent pauses itself rather than failing quietly |
Severity is derived from the failure, not configured — there are no thresholds for you to tune.
The alert is recorded first
Before anything is delivered anywhere, the alert is written to the project alongside the run that produced it. That record is the source of truth, and it survives whatever happens to any delivery attempt.
It surfaces on the Agents tab
Open alerts appear at the top of the project’s Agents tab, newest first, marked by severity, each with the failing agent, the title and the detail. Dismissing one acknowledges it and clears it from the list; the record itself is kept.
And it can be emailed to the owner
Where alert email is configured for the workspace, the alert also goes to the address on the project owner’s profile — severity in the subject line, the failure detail in the body, and a link back to the project’s Agents tab. Delivery is best-effort and never blocks or fails the check.
[Alert] preview returned HTTP 502🚨 CRITICAL
preview returned HTTP 502
GET https://your-project.preview.example → 502
Open the project's agents: https://code-anything.com/projects/…/agents
— code-anything agents (automated monitoring)Alerts are delivered to the project owner’s email and shown in the app. There is no Slack, webhook or paging integration for alerts today — if you need one, say so on the enterprise page.
Autonomy
What an agent may do alone, and what needs your approval
Every tool call an agent makes is evaluated before it runs. The answer is allow, ask, or refuse — and some things are refused no matter who is asking.
| Action | Who decides | Why |
|---|---|---|
| Read a file, search the tree, list files, screenshot the preview | On its own | Pure reads. Allowed in every mode, including read-only runs. |
| Write, edit or patch a file inside the disposable sandbox | On its own | The sandbox is throwaway, and every accepted build is checkpointed to version history first. |
| Run an ordinary shell command in the sandbox | On its own | Installs, builds, test runs — bounded by the run’s step and spend budget. |
| Generate an image or a video into the project | On its own | Treated as a file write; metered like any other model work. |
| Run a destructive command | Asks first | Force-push, rm -rf, curl piped into a shell, find -delete, git reset --hard, git clean -f, or a production deploy command. Escalated for sign-off even inside the sandbox — the run blocks until you answer. |
| Write to a protected path | Never — refused | The .git and .github trees, any .env file, certificates and private keys, SSH identities, AWS credentials, named secrets files. Hard-denied in every mode; approval cannot unlock it. |
| Ship a change to your repository | You merge it | With a repo connected, the edit path opens a pull request you review and merge. Nothing is committed to your default branch behind your back. |
| Install a standing prompt agent | You opt in | Nothing installs one for you. Presets that arrive with a new project come switched off; one you install from the Agents tab starts immediately, behind a confirmation naming its cadence, its credit cost, and the fact that it edits code unattended. |
| Answer an escalated tool call | You have 3 minutes | An approval request that goes unanswered for three minutes is auto-denied and the run continues with that call refused — it never defaults to allowing the action. |
An escalated call blocks the run until you answer it, and is denied — never quietly allowed — if you do not.
The honest warning about prompt agents
A prompt agent is the one genuinely unattended writer in the system. When it fires it runs a real build against your project with nobody reviewing it first, and it can overwrite edits you made by hand in the editor since its last run. That is exactly what it is for — and exactly why nothing installs one for you, why turning one on takes a confirmation that spells this out in those words, and why the product calls it advanced rather than offering it as a default.
A health-check agent has none of this exposure. It cannot write, cannot run commands, and cannot call the model at all.
Why your app cannot claim its own build broke
Business events arrive from your app through a per-project endpoint, signed with a token held in the encrypted vault and valid inside a five-minute window. Lifecycle events — deploy failed, run rejected — are our own observations and are fired in-process only.
The asymmetry is deliberate. If an app could assert its own lifecycle events over the wire, a compromised storefront could trigger unattended repair builds simply by lying about its health. Secrets and connections live on the project’s Environment surface, never in the agent’s reach.
Runbook
How to investigate an incident, in order
Six surfaces, each answering the next question. Nothing here needs a support ticket.
- 1
Start at the open alertsagents tab
The project’s Agents tab lists every alert that has not been acknowledged, newest first, colour-coded by severity. Each one carries the title, the body — the URL, the status, the matched error line — and the agent that raised it.
- 2
Read that agent’s run historyhistory
Open History on the agent and you get its last twenty runs with a status of ok, alert or error, a one-line summary and a timestamp. That tells you whether this is a first failure or the third in an hour.
- 3
Open the build behind itverify + trace
Every build has a Verify surface showing the gate as it ran — which check passed, which failed, and the accept-or-hold verdict — and a per-run trace: a span timeline with the model used, the tokens, the cost, and each check’s pass or fail attached to it. Nothing about the gate is hidden from you.
- 4
Identify the release that changed thingsdeploy log
Repository → Deploy keeps the deploy history: every Go Live, what shipped, when, and whether it failed. It is the fastest way to line a failure up against a release.
- 5
Read the build that introduced itchanges
Repository → Changes shows the recent runs with their verdict — accepted, held for review, or rejected — the reason, and the files each one touched. A held build tells you what the agent itself refused to ship.
- 6
Rewind, or roll the canary backrecover
Repository → Changes carries both restore paths: a list of build checkpoints you can rewind the sandbox to, and a durable version history that outlives the sandbox entirely. Both take a safety snapshot of the current state first, so the restore is itself reversible; both refuse while a build or a deploy is in flight; and where the project has a provisioned database, its data is snapshotted and restored alongside the code. If the release went out as a delivery flight, its gate rolls the canary cohort back on its own — by default the moment it records a single error, or once enough of the cohort rejects it.
There is no single rollback button, and we would rather say so than imply one. Recovery is the version history and rewind on Repository → Changes, and — for a staged release — the delivery flight’s own gate. Both are described step by step in the documentation.
Limits
The limits agents run inside
Real numbers from the running system. Where a limit exists to protect you rather than us, it says which.
| Limit | Value | Why it exists |
|---|---|---|
| Heartbeat | Every minute | A scheduled trigger calls the control plane once a minute; the engine then decides which agents are actually due. |
| Health-check cadence floor | Every 5 minutes | Politeness towards the target you are probing. You can set anything slower. |
| Prompt-agent cadence floor | Every 60 minutes | Each prompt run is a metered build, so the floor caps a single agent at roughly 24 runs a day. |
| Health-check probe timeout | 8 seconds | One slow target can never stall the heartbeat for everything else. |
| Prompt-run wall clock | 120 seconds | A hung build is aborted and the agent is recorded as errored rather than left hanging. |
| Prompt agents running at once | 3 per tick | Health checks stay fully parallel; only the expensive prompt builds are queued. |
| Run history kept per agent | Last 20 runs | Status, summary and timestamp for each, in the History drawer on the Agents tab. |
| Open alerts listed | Up to 20 unacknowledged | Dismissing an alert acknowledges it; the row itself is kept. |
| Unanswered approval request | Auto-denied after 3 minutes | A blocked run cannot wait forever — so it fails the call closed rather than open. |
| Model steps inside one work item | Ceiling of 200 | A runaway backstop, not a spend limit. Spend is governed separately. |
| Stuck-loop detection | 3 identical calls, or 5 steps with no progress | The same tool call three times running, or five consecutive steps that change nothing, ends the attempt instead of burning your budget. |
Cadence floors are ours, not the scheduler’s — the engine could fire every minute; a five-minute floor is politeness towards the URL you are probing.
Per-run budgets, by plan
| Plan | Spend ceiling per run | Model steps per run | Builds at once |
|---|---|---|---|
| Free | $0.50 | 30 | 1 |
| Starter | $2.00 | 60 | 3 |
| Pro | $5.00 | 100 | 6 |
| Team | $10.00 | 160 | 12 |
Plan defaults for the run guard. Where the account’s credit balance is known, spend is bounded by what that balance can afford instead, and the stuck-loop detectors above become the practical ceiling. A run that hits a limit stops and reports rather than continuing to spend.
- A prompt agent whose account is out of runtime credits pauses itself and records why, instead of failing silently
- A prompt trigger is charged once per firing — a retried heartbeat cannot double-charge the same run
- Standing workspace monitoring is a paid feature from the Pro plan upward; automatic build repair is on every plan, free included
Keep reading
Where to go next
FAQ
Questions about agents
The things people ask before they hand anything to an automation.
Build with an agent that checks its own work.
Automatic repair runs on every build, on every plan. Standing agents are yours to add when you want something watched — nothing starts watching until you say so.