Normally a turn ends when the model stops talking. /goal overrides that: at the point the
session would finish, an evaluator judges whether your condition has been met. If it hasn't,
the reason is fed back and the agent keeps going automatically.
Use it when "done" is a state you can describe rather than a single reply. Conditions that work well look like:
all tests in packages/core passno TypeScript errors remain in src/GET /health returns 200/goal all tests in packages/core pass
| Command | Effect |
|---|---|
/goal <condition> |
set a goal and start working it |
/goal status |
current goal, budget used, last evaluator reason |
/goal pause |
stop auto-continuing; the goal is remembered |
/goal resume |
resume auto-continuing |
/goal edit <text> |
rewrite the condition without losing progress |
/goal append <text> |
add a requirement to the condition, keeping what's there |
/goal clear |
remove the goal entirely |
edit replaces the condition. When you spot something while the agent is already working —
an edge case, an extra check — append adds to it instead:
/goal append and cover the empty-input case with a test
+ is shorthand for the same thing: /goal +and cover the empty-input case with a test.
The extended condition is persisted immediately and the addition is pushed to the agent as its next message, so it lands on the current burst rather than waiting for the turn to end. If the goal was paused or already achieved, appending re-activates it.
This holds in the editor panel as well as the terminal: a message sent while the agent is working is delivered to the turn already running, not queued behind it.
add and also are deliberately not accepted — /goal add pagination to the users table
reads as a new condition, and guessing wrong would silently bury it inside the old one.
Note
The appended text becomes part of the condition the evaluator judges, so the goal cannot complete until the addition is satisfied too. An addition the evaluator ignored would let the session finish with your late requirement unmet.
A goal that can never be satisfied would otherwise run forever, so each burst is bounded on three axes. Whichever is hit first ends the burst.
| Limit | Default | Config key |
|---|---|---|
| Turns | 75 | goal.maxTurns |
| Duration | 45 minutes | goal.maxDurationMinutes |
| Context growth | 500,000 tokens | goal.maxTokens |
Limits are checked between turns, never mid-turn — work already running is not cut off part way through. In practice duration is the axis that fires: a turn that edits, builds and runs tests takes minutes, so the clock usually runs out well before turn 75.
Set any subset in your config; unset axes keep their default.
{
"goal": {
"maxTurns": 150,
"maxDurationMinutes": 120,
"maxTokens": 1000000
}
}
Limits are recorded on the goal when it is created, so a burst cannot change its own budget
halfway through. A new /goal — or /goal resume on an existing one — picks up whatever the
config says at that moment, which is what you want after raising a limit you just hit.
Note
maxTokensbounds how far the context window grows during a burst, not what the burst costs you. Every turn re-sends the conversation, so tokens actually billed run well above this figure; and when compaction shrinks the context, the measured growth falls back toward zero. Treat it as a runaway-context guard, not a spend cap.
Reaching a limit does not discard the goal. The agent is sent a wrap-up prompt, reports
where it got to, and stops. Running /goal resume starts a fresh burst with the budget reset.
Note
Budgets reset at the start of each new run, not per turn. Resuming a paused goal gives you a full budget again.
At the point the session would end, an evaluator model reads the conversation and returns one of three verdicts:
The evidence bar scales to the condition. A mechanical condition ("tests pass") is met only when the actual output appears in the conversation. A qualitative one ("the design is white and classic") is met once the agent has done the work and shown concrete specifics of it — it is not held to proof it cannot produce.
If the evaluator model cannot be reached, the failure is reported in the transcript with the model and the cause — not left in a log file. A goal that silently never completes is indistinguishable from a broken product, and the cause is usually neither the condition nor the work.
Where a tier enforces a background model for goal evaluation and that model fails at call time, evaluation falls back to your session model for the rest of the burst and the goal keeps going. Only if the fallback also fails does the goal stop being judgeable, and the notice says so.
If a turn fails on something a further turn cannot fix, the goal is cleared and the cause is named, rather than the burst retrying the same failure until a safety limit:
Everything else — rate limits, overloaded servers, 5xx — is transient, and the goal keeps working through it. A model that is unavailable also leaves the goal alone: the session already falls back to your tier's default model, so the next turn runs somewhere else instead of repeating the failure.
If several turns in a row make no tool calls, the agent is restating progress rather than making it. The loop stops, prints a warning, and hands control back — with the goal still set. Evaluation resumes on your next message, so a hint ("look at the CSS file") is enough to continue.
This is distinct from Impossible, which retires the goal outright. A stall is usually recoverable; an impossible condition is not.
While the loop is stopped the goal reads as waiting for you rather than active — in
/goal status, in the TUI sidebar, and on the editor status chip. The goal is still set and
nothing is running, and "active" alone would suggest otherwise.
The evaluator sees the conversation, not your repository. A condition it cannot observe from the transcript — "the deploy is healthy" — cannot be judged reliably. Prefer conditions the agent can demonstrate in-session by running something.
| Instead of | Write |
|---|---|
make it better |
no eslint warnings in src/ |
fix the bug |
the failing test in auth.test.ts passes |
refactor this |
no function in parser.ts exceeds 50 lines |
A condition is good when a third party could look at the session and agree it is or isn't met.
/goal pays off wherever "done" takes several turns and the agent would otherwise stop at the
first plausible answer.
| Use case | Effect | Why |
|---|---|---|
| Migration | Much better | The strongest fit. Mechanical work across many files runs past a single turn; a goal keeps it going instead of leaving half the call sites converted. |
| Feature Development | Much better | "Until the endpoint returns the right shape and its tests pass" is a checkable condition, so the agent doesn't declare success at the first compile. |
| Testcase creation | Better | Give it a coverage condition — "until every branch in parser.ts is covered" — and it keeps adding cases rather than writing three and stopping. |
| Bug-Fixer | Better | Works, but /bug-fix is the better tool: same engine, with a root-cause and regression-test condition already written. |
| Design / Architecture / Planning | Little | Planning ends in a document, not a testable state. There's rarely a condition an evaluator can judge. |
| Reviewer | Little | A review is one pass over a fixed diff. Nothing to persist toward. |
The deciding question is whether you can name a finish line someone else could verify. If you
can, /goal helps; if the output is prose, it mostly doesn't.
/bug-fix <description> is /goal with a pre-written condition: root-cause → fix →
regression-test → verify. Reach for it instead of hand-writing a debugging goal./boost adds verification passes to a single turn. It is per-turn quality; /goal is
multi-turn persistence. They compose, but both consume model calls — see /loop
for the scheduling equivalent.The goal is stored on the session record, not in the message stream. It survives compaction, is restored when you resume the session, and is deleted with it.