Solution · Fix a failing test
Fix the failing test and open a PR
Connect your repo and point at the broken test. The agent reproduces the failure, finds the cause, fixes it, re-runs the suite green, and opens a pull request.
For developers · connect a repo → reviewed pull request
At a glance
Is this the right pattern for you?
The short version, before the detail — who this is written for, what you start from, and what exists at the end.
| This solution | |
|---|---|
| Written for | Developers with an existing codebase |
| You start from | A GitHub repository and a change you want made |
| You end up with | A scoped, verified pull request awaiting your review |
| The path runs | Connect → index → edit → verify → pull request |
| Built on | Sandbox verification · Code intelligence · GitHub PR agent |
| Plan needed | Pro — GitHub repositories and pull requests are a Pro capability |
No timings are quoted anywhere on this page, because how long a build takes depends entirely on what you asked for.
The challenge
Why this is usually hard
Worth understanding before the how — because the reason this is difficult is not the reason most people assume.
A red test in CI is rarely difficult and always expensive. The cost is not the fix; it is the interruption. You stop what you were doing, get the failure reproducing locally, work out whether the bug is in the code or in the test that covers it, make a change, then re-run everything to confirm you have not traded one failure for two. By the time the suite is green you have lost the thread of whatever you were actually working on.
The reproduce-and-trace loop is the slow part, and it is slow in a particular way: it is mostly waiting. Waiting for a suite to run, reading a stack trace, changing one thing, waiting again. That loop is a poor use of a person who was holding a different problem in their head five minutes ago, and it is the reason red tests get muted, skipped or left for the next sprint more often than anyone admits.
The important distinction in this pattern is between silencing a test and fixing a cause. A test can be made green by weakening its assertion, and that is worse than leaving it red — it converts a visible problem into an invisible one. The version worth having reproduces the failure, works out where the fault actually lives, changes that, and then proves the whole suite still passes.
How it works
From a connected repo to a reviewed PR
Every stage in order, including the ones that happen without you asking.
- 1
Connect your repository
$ connect repoPoint the agent at your GitHub repository. It reads and indexes the code, and works against a clone — your branches are not touched, and the connection can be scoped or revoked at any time.
- 2
Name the failing test, or describe the symptom
$ fix the failing checkout testTell it which test is red. If you only know the symptom — the checkout total is wrong when a coupon is removed — that works too, because the codebase is indexed semantically as well as by symbol, so the relevant test and the logic it exercises can be found from a description.
- 3
Reproduce the failure in a sandbox
The suite is run in an isolated sandbox clone to reproduce the failure first. This is the step that separates a real fix from a plausible guess: nothing is changed until the failing behaviour has actually been observed.
- 4
Trace whether the fault is in the test or the code
Symbol-aware indexing maps the test to the logic it covers, so the question "is this test wrong, or is the code wrong?" is answered against the real call graph rather than assumed. A test asserting outdated behaviour and a genuine bug get different treatment.
- 5
Apply a minimal fix and re-run everything
verifyThe change is kept as small as the cause allows, then typecheck, tests and build are re-run against the clone to confirm the suite is green and nothing else regressed. A fix that breaks two other tests is not a fix, and it does not get offered.
- 6
A pull request explains the cause, not just the change
$ open prWhat arrives is a pull request describing what was actually wrong and why the change addresses it. That reasoning is the part you need in order to review a bug fix quickly — a diff without a diagnosis is just a guess you have to re-derive.
Why code-anything
What you get out of the box
Not a feature list — the specific things this pattern removes from your side of the work.
Green before you look at it
The fix is only offered once the suite, typecheck and build pass against the sandbox clone. You are reviewing a proposal that has already cleared the gates, not investigating whether it runs.
A root cause, not a muted assertion
The agent traces whether the fault is in the test or in the code it covers and addresses the actual cause. Making a suite green by weakening it turns a visible problem into a silent one, which is the outcome this is designed to avoid.
Delivered as a reviewable pull request
The fix arrives as a standard GitHub pull request with the reasoning written down, so you can check the diagnosis as well as the diff before merging.
The interruption happens elsewhere
The slow reproduce-and-trace loop runs in the sandbox rather than on your machine. The main cost of a red test has always been the context switch, and this is where it is recovered.
A minimal change, not a rewrite
The change is scoped to the cause rather than expanded into opportunistic refactoring, which keeps the diff small enough to review properly and easy to reason about if it needs reverting.
Nothing runs against your repository
Reproduction, the fix attempt and every re-run happen in an isolated clone. Your repository sees exactly one thing: a pull request you chose to merge.
In practice
What it looks like
The second and third lines carry the value. Reproducing the failure before changing anything is what makes the result a fix rather than a plausible edit, and tracing the cause is what decides whether the test or the code was wrong.
The pull request title in that flow describes a cause, not a symptom. That is the artefact you want at the end of a bug hunt: a small change, a green suite, and a written explanation you can check in a couple of minutes.
$ connect repo acme/checkout
$ fix the failing checkout test and open a PR
✓ reproduced the failure in checkout.test.ts
✓ traced the cause to a stale total when a coupon is removed
✓ applied a minimal fix, re-ran the suite
✓ typecheck · tests · build all pass
→ opened a pull request "Fix checkout total on coupon removal"What you get
- The failing test reproduced in an isolated clone before anything was changed
- A traced root cause, with the test-versus-code question actually answered
- A minimal fix scoped to that cause rather than an opportunistic refactor
- The full suite, typecheck and build passing against the change
- A pull request explaining the diagnosis as well as the diff
- A repository that only changes when you merge
What it involves
The honest shape of the path
Stage by stage: what the platform does, and what is genuinely still asked of you.
| Stage | What actually happens | What is asked of you |
|---|---|---|
| Connect | You authorise a GitHub repository. Nothing is written to your branches. | One authorisation, scoped and revocable. |
| Index | The repository is indexed semantically, so behaviour can be found without a file name, and by symbol, so the agent knows what a definition is used by. | Nothing. This runs on connection, before any change is proposed. |
| Edit | The change is made inside an isolated sandbox clone of your repository. | A description of the change in your own terms — not a list of files. |
| Verify | Typecheck, tests and build run against the clone. If a gate fails, the agent keeps working rather than handing you broken code. | Nothing, though the value of this step is bounded by how real your own checks are. |
| Pull request | A scoped pull request is opened with a summary of what changed. | A review, and the merge — which is the only moment your repository changes. |
The stage names above are the product’s own: connect, index, edit, verify, pull request. How long a pass takes depends on the repository and the change, so no figure is quoted here.
Connecting a GitHub repository and opening pull requests is a Pro-plan capability, so this path needs Pro. Sandbox verification — typecheck, tests and build running before you see a change — is included on every plan. See the full plan comparison.
What you own
What you are left holding
On a codebase that already exists and already matters, the interesting question is not how fast a change can be produced but what is true about it by the time you are asked to look at it. Here is exactly what the arrangement guarantees, and what it deliberately leaves to you.
What is true before you see it
- Every edit happened in an isolated sandbox clone, never in your repository
- Your own typecheck, tests and build were run against the change and passed
- The work is scoped to what you asked for rather than mixed with unrelated tidying
- The pull request carries a written summary of what changed and why
What stays your decision
- Whether to merge — a pull request is a proposal, and your repository is unchanged until you accept it
- The agent never pushes to your branches and nothing is merged on your behalf
- You can leave a pull request unmerged at no cost beyond the time spent reading it
- The GitHub connection is yours to scope or revoke whenever you want
FAQ
Common questions
Keep looking
Other reference patterns
Every solution is a worked end-to-end example. If this one is not quite your situation, one of these probably is.
Add dark mode
Connect your GitHub repo and ask for dark mode. The agent indexes the code, makes the change in a sandbox, verifies it builds, and opens a pull request for you to review.
For developersJS to TypeScript
Connect your repo and name the components to convert. The agent adds real types, fixes the fallout, verifies the build, and opens a pull request you can review.
For developersOr the other way of working — describe an app in plain English and get a live, hosted result:
Bakery website
Describe your shop in plain English and watch a real website build itself in a live preview — complete with a contact form that emails you every enquiry.
For anyoneBooking app
Describe the booking flow you want and get a working app — a calendar customers book into and an email confirmation sent on every appointment, all wired up for you.
For anyoneStripe dashboard
Describe the numbers your team cares about and get a private dashboard — revenue, growth and customers — with sign-in, built without touching a line of code.
For anyoneBuild fix failing tests the way it should have been.
Connect a repository and describe the change. Every edit happens in an isolated clone, clears your own typecheck, tests and build, and arrives as a pull request you review.