/goal

Keep working until a condition is met, instead of stopping after one answer

What it does

Normally a turn ends when the model stops talking. /goal overrides that: at the point the session would finish, an evaluator judges whether your condition has been met. If it hasn't, the reason is fed back and the agent keeps going automatically.

Use it when "done" is a state you can describe rather than a single reply. Conditions that work well look like:

/goal all tests in packages/core pass

Sub-commands

Command Effect
/goal <condition> set a goal and start working it
/goal status current goal, budget used, last evaluator reason
/goal pause stop auto-continuing; the goal is remembered
/goal resume resume auto-continuing
/goal edit <text> rewrite the condition without losing progress
/goal append <text> add a requirement to the condition, keeping what's there
/goal clear remove the goal entirely

Adding to a goal mid-run

edit replaces the condition. When you spot something while the agent is already working — an edge case, an extra check — append adds to it instead:

/goal append and cover the empty-input case with a test

+ is shorthand for the same thing: /goal +and cover the empty-input case with a test.

The extended condition is persisted immediately and the addition is pushed to the agent as its next message, so it lands on the current burst rather than waiting for the turn to end. If the goal was paused or already achieved, appending re-activates it.

This holds in the editor panel as well as the terminal: a message sent while the agent is working is delivered to the turn already running, not queued behind it.

add and also are deliberately not accepted — /goal add pagination to the users table reads as a new condition, and guessing wrong would silently bury it inside the old one.

Note

The appended text becomes part of the condition the evaluator judges, so the goal cannot complete until the addition is satisfied too. An addition the evaluator ignored would let the session finish with your late requirement unmet.

Safety limits

A goal that can never be satisfied would otherwise run forever, so each burst is bounded on three axes. Whichever is hit first ends the burst.

Limit Default Config key
Turns 75 goal.maxTurns
Duration 45 minutes goal.maxDurationMinutes
Context growth 500,000 tokens goal.maxTokens

Limits are checked between turns, never mid-turn — work already running is not cut off part way through. In practice duration is the axis that fires: a turn that edits, builds and runs tests takes minutes, so the clock usually runs out well before turn 75.

Raising them

Set any subset in your config; unset axes keep their default.

{
  "goal": {
    "maxTurns": 150,
    "maxDurationMinutes": 120,
    "maxTokens": 1000000
  }
}

Limits are recorded on the goal when it is created, so a burst cannot change its own budget halfway through. A new /goal — or /goal resume on an existing one — picks up whatever the config says at that moment, which is what you want after raising a limit you just hit.

Note

maxTokens bounds how far the context window grows during a burst, not what the burst costs you. Every turn re-sends the conversation, so tokens actually billed run well above this figure; and when compaction shrinks the context, the measured growth falls back toward zero. Treat it as a runaway-context guard, not a spend cap.

Reaching a limit does not discard the goal. The agent is sent a wrap-up prompt, reports where it got to, and stops. Running /goal resume starts a fresh burst with the budget reset.

Note

Budgets reset at the start of each new run, not per turn. Resuming a paused goal gives you a full budget again.

How completion is judged

At the point the session would end, an evaluator model reads the conversation and returns one of three verdicts:

The evidence bar scales to the condition. A mechanical condition ("tests pass") is met only when the actual output appears in the conversation. A qualitative one ("the design is white and classic") is met once the agent has done the work and shown concrete specifics of it — it is not held to proof it cannot produce.

When evaluation itself fails

If the evaluator model cannot be reached, the failure is reported in the transcript with the model and the cause — not left in a log file. A goal that silently never completes is indistinguishable from a broken product, and the cause is usually neither the condition nor the work.

Where a tier enforces a background model for goal evaluation and that model fails at call time, evaluation falls back to your session model for the rest of the burst and the goal keeps going. Only if the fallback also fails does the goal stop being judgeable, and the notice says so.

Errors that clear the goal

If a turn fails on something a further turn cannot fix, the goal is cleared and the cause is named, rather than the burst retrying the same failure until a safety limit:

Everything else — rate limits, overloaded servers, 5xx — is transient, and the goal keeps working through it. A model that is unavailable also leaves the goal alone: the session already falls back to your tier's default model, so the next turn runs somewhere else instead of repeating the failure.

When the agent stops making progress

If several turns in a row make no tool calls, the agent is restating progress rather than making it. The loop stops, prints a warning, and hands control back — with the goal still set. Evaluation resumes on your next message, so a hint ("look at the CSS file") is enough to continue.

This is distinct from Impossible, which retires the goal outright. A stall is usually recoverable; an impossible condition is not.

While the loop is stopped the goal reads as waiting for you rather than active — in /goal status, in the TUI sidebar, and on the editor status chip. The goal is still set and nothing is running, and "active" alone would suggest otherwise.

The evaluator sees the conversation, not your repository. A condition it cannot observe from the transcript — "the deploy is healthy" — cannot be judged reliably. Prefer conditions the agent can demonstrate in-session by running something.

Writing good conditions

Instead of Write
make it better no eslint warnings in src/
fix the bug the failing test in auth.test.ts passes
refactor this no function in parser.ts exceeds 50 lines

A condition is good when a third party could look at the session and agree it is or isn't met.

Where it improves the result

/goal pays off wherever "done" takes several turns and the agent would otherwise stop at the first plausible answer.

Use case Effect Why
Migration Much better The strongest fit. Mechanical work across many files runs past a single turn; a goal keeps it going instead of leaving half the call sites converted.
Feature Development Much better "Until the endpoint returns the right shape and its tests pass" is a checkable condition, so the agent doesn't declare success at the first compile.
Testcase creation Better Give it a coverage condition — "until every branch in parser.ts is covered" — and it keeps adding cases rather than writing three and stopping.
Bug-Fixer Better Works, but /bug-fix is the better tool: same engine, with a root-cause and regression-test condition already written.
Design / Architecture / Planning Little Planning ends in a document, not a testable state. There's rarely a condition an evaluator can judge.
Reviewer Little A review is one pass over a fixed diff. Nothing to persist toward.

The deciding question is whether you can name a finish line someone else could verify. If you can, /goal helps; if the output is prose, it mostly doesn't.

Related

Where it lives

The goal is stored on the session record, not in the message stream. It survives compaction, is restored when you resume the session, and is deleted with it.