▞ jazzdocsblogpersonas
github

The agent loop

This page explains what happens between “you press enter” and “you get an answer” — especially on runs that take a long time.

Source: packages/core/src/agent/execution/agent-loop.ts

An agent run is a bounded loop: ask the model, execute whatever tools it asked for, feed the results back, repeat until it stops asking for tools. The loop itself is twenty lines. Everything that makes it survive a long autonomous run is in the guards around it.


One full iteration

flowchart TD
    START(["Iteration i of 100"]) --> COMPACT

    COMPACT["<b>Compact if needed</b><br/>tokens &gt; 80% of window?<br/>→ summarize, then resume"]
    COMPACT --> PRESSURE["<b>Build budget pressure</b><br/>just compacted → 'continue the task'<br/>else i/80 ≥ 70% → 'consolidate'<br/>i/80 ≥ 90% → 'finish now'<br/><i>ephemeral — not stored</i>"]
    PRESSURE --> ASK["<b>Ask the model</b><br/>messages + tools + reasoning effort"]

    ASK --> INT{"Interrupted?<br/>(double-Esc)"}
    INT -->|yes| STOP(["Break — keep partial output"])
    INT -->|no| RECORD["<b>Record usage</b><br/>tokens, cost,<br/>calibrate token counter"]

    RECORD --> APPEND["<b>Append assistant message</b><br/>content + reasoning + tool_calls"]
    APPEND --> TRIM["<b>Trim</b> to the token budget<br/>(turn-aware — never splits<br/>a tool call from its result)"]

    TRIM --> BRANCH{"Did the model<br/>request tools?"}
    BRANCH -->|"no"| FINAL["<b>Final response</b><br/>present, notify, exit loop"]
    BRANCH -->|"yes"| TOOLPHASE

    TOOLPHASE["<b>Tool phase</b><br/>see below"]
    TOOLPHASE --> NEXT(["Iteration i+1"])

    classDef guard fill:#f9a03f,stroke:#b3541e,color:#1a1a1a
    classDef work fill:#4f9d9d,stroke:#2f6d6d,color:#ffffff
    class COMPACT,PRESSURE,TRIM guard
    class ASK,TOOLPHASE work

The loop runs at most DEFAULT_MAX_ITERATIONS (100) times. It exits early on a final response (no tool calls) or a user interrupt.


The tool phase

flowchart TD
    IN(["Model requested N tool calls"]) --> TRACK["<b>Track</b> each call as<br/><code>name:arguments</code><br/>keep a rolling window of 10"]

    TRACK --> MELT{"<b>Meltdown?</b><br/>unique keys / 10 &lt; 40%"}
    MELT -->|yes| INJECT["Inject recovery message:<br/>'stop, summarize, try a<br/>different strategy'<br/>reset the window"]
    MELT -->|no| EXEC
    INJECT --> EXEC

    EXEC["<b>Execute</b><br/>≤10 concurrent, each forked<br/>3-min default timeout<br/>approval-gated per tool"]
    EXEC --> VALIDATE{"Every call<br/>got a result?"}
    VALIDATE -->|no| FAIL(["Hard fail — this is a bug,<br/>not a recoverable state"])
    VALIDATE -->|yes| FORMAT["<b>Format results for context</b><br/>compress before storing"]

    FORMAT --> QUEUE{"User typed<br/>while we worked?"}
    QUEUE -->|yes| PUSH["Append their message —<br/>steer mid-run"]
    QUEUE -->|no| OUT
    PUSH --> OUT(["Next iteration"])

    classDef guard fill:#f9a03f,stroke:#b3541e,color:#1a1a1a
    classDef bad fill:#c1443c,stroke:#7d2b26,color:#ffffff
    class MELT,INJECT,FORMAT guard
    class FAIL bad

Missing results are a hard failure

If any requested tool call comes back without a result, the run fails loudly. It would be easy to paste in a placeholder and continue — but a message history where a tool_calls entry has no matching tool message is invalid to most providers, and the resulting error appears three iterations later somewhere unrelated. Failing at the source keeps the bug findable.

Mid-run steering

Tool execution is slow, and users type while it happens. Anything queued during the tool phase is appended as a user message before the next iteration, so you can redirect a long run without killing it.


Guard 1 — budget pressure

An agent with 100 iterations and no sense of time will happily spend all 100 on research and produce nothing. So Jazz tells it where it stands:

timeline
    title Iteration budget (100 iterations)
    section 1–69 · Free rein
        No pressure : explore, research, spawn sub-agents
    section 70–89 · 70% warning
        "Begin consolidating results. Stop spawning new research subagents." : agent shifts to synthesis
    section 90–100 · 90% critical
        "Write your final output NOW. Use what you have collected so far." : agent lands the plane

The important detail: the pressure message is ephemeral. It’s appended to the array sent to the model for that one call and never pushed into the stored conversation:

const budgetMsg = buildBudgetPressureMessage(iterationIndex + 1, maxIterations);
const messagesForLLM = budgetMsg ? [...state.currentMessages, budgetMsg] : state.currentMessages;

If it were stored, iteration 98 would carry eight escalating “FINISH NOW” messages, each one costing tokens and confusing the transcript that later gets summarized. This way the nudge steers the run without polluting its history.

Sibling guards — cost, tokens, and wall-clock time

Three more budgets check in the same place, right after each iteration’s tool phase finishes, and are configured via maxCostUSD, maxTokens, and maxDurationMs (see Configuration reference):

  • maxCostUSD re-prices the run’s accumulated tokens (plus any sub-agent spend rolled up via childCostUSD) against the model’s models.dev pricing. Skipped entirely when pricing is unknown — it never guess-aborts a run it cannot verify the spend of.
  • maxTokens just sums totalPromptTokens + totalCompletionTokens. No pricing lookup, so it still enforces on a local/unpriced model where maxCostUSD cannot.
  • maxDurationMs compares wall-clock elapsed time (Date.now() - runMetrics.startedAt) against the budget. Unlike the other two, it also gets an ephemeral pressure message inside runIteration itself — buildTimeBudgetPressureMessage — at 50%, 80%, and 90% elapsed, mirroring the iteration-budget nudge above.

All three share the same timing as the iteration budget: checked between iterations, not a preemptive interrupt. A single expensive iteration — a costly tool call, or a sub-agent delegation that itself runs for a while — can push the total past the cap before the next check trips. A run stopped this way reports which cap fired on the response: costCapped / tokenCapped / durationCapped.

This is deliberately different from --timeout, which lives outside the loop entirely — the CLI races the whole run against a deadline (packages/core/src/utils/run-deadline.ts) and kills it with no warning to the agent. --timeout is the hard outer safety net; maxDurationMs is the warned, graceful budget.


Guard 2 — meltdown detection

The classic agent failure isn’t a crash, it’s a groove: the same search, the same file read, forever, until the budget is gone.

const keys = window.map((tc) => `${tc.name}:${tc.arguments}`);
const uniqueness = new Set(keys).size / windowSize;
return uniqueness < 0.4;

The key is composite — tool name plus arguments. This distinction is the whole design:

Recent callsUnique keysVerdict
web_search("effect-ts") × 101/10 = 10%🔴 Stuck — same query over and over
web_search(q1)web_fetch(u1)web_search(q2) → …10/10 = 100%🟢 Productive — that’s research
read_file(a) read_file(b) … 10 distinct files10/10 = 100%🟢 Productive — that’s reading a codebase
execute_command() × 6, read_file(x) × 42/10 = 20%🔴 Stuck

Keying on tool name alone would flag the second and third rows as meltdowns, which are exactly the behaviors you want. When a meltdown does trip, Jazz injects a message telling the agent to stop, summarize what it has, and either output or try a fundamentally different approach — then clears the window so it gets a fair chance.

Unlike budget pressure, this message is stored. It’s a real event in the run’s history and the agent should keep remembering that its last approach didn’t work.


Streaming vs batch

The loop is shared. Only how it talks to the model and renders output differs, expressed as a CompletionStrategy:

StreamingBatch
Model calltoken-by-token streamsingle response
Renderinglive, as tokens arrivemarkdown rendered at the end
Thinking indicatoryesno
Interruptibleyes — double-Esc races tool executionno
Used byterminal TUI, --events bridgesjazz run piped output, CI

Adding a third mode means implementing one interface, not forking the loop. That’s the reason for the indirection.


Finalization

After the loop, win or lose:

flowchart LR
    EXIT(["Loop exits"]) --> WHY{"How?"}
    WHY -->|"final answer"| OK["Normal"]
    WHY -->|"hit 100 iterations"| LIMIT["Warn: iteration limit"]
    WHY -->|"empty response"| EMPTY["Warn: empty response"]
    WHY -->|"interrupted"| INT["Keep partial output"]

    OK --> METRICS
    LIMIT --> METRICS
    EMPTY --> METRICS
    INT --> METRICS

    METRICS["<b>Write metrics</b><br/>forked fiber, awaited on release —<br/>never blocks your answer,<br/>never lost on exit"]
    METRICS --> COST["<b>Cost</b><br/>own tokens × model price<br/>+ sub-agent cost"]
    COST --> RESP(["AgentResponse<br/>content · usage · costUSD · messages"])

    classDef core fill:#4f9d9d,stroke:#2f6d6d,color:#ffffff
    class METRICS,COST core

Two details worth noting:

  • Metrics are written on a forked fiber that the loop’s release step awaits. You get your answer immediately, and the telemetry still lands even if the process is shutting down.
  • Cost includes sub-agent spend. A run whose own tokens are unpriced (a local model) but which spawned a priced sub-agent still reports a figure — otherwise the number would silently understate what you spent.

Reading a run in the logs

Log lineMeans
Sending LLM requestTop of an iteration — includes iteration number, message count, tool count
Agent decided to use toolsTool phase starting, with the chosen tool names
Meltdown detected — injecting recovery signalGuard 2 fired; the agent was looping
Compacting contextCrossed 80% of the window; a summary is being produced
Tool timeout: <name>A tool exceeded its timeout; returned as a failed result, not a crash
Agent provided final responseLoop exiting normally
Missing tool results for some tool callsBug — please open an issue with the log

machine-readable: /docs/internals/agent-loop.md · /llms.txt · /llms-full.txt