How to enable context compaction for long-running tasks
An agent session accumulates tokens with every message. For tasks that take many turns — processing large files, iterating on a plan, or running inside a persistent capsule — the conversation history can approach the model's context window. Context compaction automatically condenses the message history so the session can continue without hitting a hard limit.
The relevant manifest options are:
| Option | Controls |
|---|---|
| context.max_tokens | Token budget for the session; required to enable compaction |
| inference.compaction.threshold | Fraction of context.max_tokens that triggers compaction |
| inference.compaction.model | Model used for the compaction call (optional override) |
| inference.compaction.system_prompt | System prompt override for the compaction call (optional; hook picks its own default when unset) |
| inference.compaction.system_prompt_file | Same override, loaded from a file next to the manifest (optional; mutually exclusive with system_prompt) |
| inference.compaction.dump_summaries | When true, appends one JSON line per committed compaction to out/compaction-summaries.jsonl (optional; default false) |
Step 1 — create murmur.yaml with context.max_tokens and the compaction artifact
Compaction requires two things: context.max_tokens set to match your model's actual context window, and murmur-hook-compact declared as a hook artifact. Create a murmur.yaml file:
name: my-agent
version: "0.1.0"
context:
max_tokens: 1000000
artifacts:
- name: murmur-driver-anthropic
version: "1.0.0"
runtime: driver
- name: murmur-hook-compact
version: "1.0.0"
runtime: hook
inference:
transport: http
endpoint: https://api.anthropic.com
model: claude-sonnet-5
api_key: ${ANTHROPIC_API_KEY}
driver:
artifact: murmur-driver-anthropic
name: my-agent
version: "0.1.0"
context:
max_tokens: 200000
artifacts:
- name: murmur-driver-openai
version: "1.0.0"
runtime: driver
- name: murmur-hook-compact
version: "1.0.0"
runtime: hook
inference:
transport: http
endpoint: https://api.openai.com
model: o3-mini-high
api_key: ${OPENAI_API_KEY}
driver:
artifact: murmur-driver-openai
name: my-agent
version: "0.1.0"
context:
max_tokens: 1000000
artifacts:
- name: murmur-driver-deepseek
version: "1.0.0"
runtime: driver
- name: murmur-hook-compact
version: "1.0.0"
runtime: hook
inference:
transport: http
endpoint: https://api.deepseek.com
model: deepseek-r1
api_key: ${DEEPSEEK_API_KEY}
driver:
artifact: murmur-driver-deepseek
context.max_tokens is a manifest value — the runtime does not query the API to discover the context window. Use the number from your model's documentation; most current frontier models support between 400,000 and 1,000,000 tokens.
Using runtime: hook ensures the model never sees murmur-hook-compact as a callable tool — the runtime invokes it directly at fixed lifecycle points.
Step 2 — install dependencies
Both context.max_tokens and the compaction hook must be present for compaction to activate. Either one alone is not enough.
mur install
Different ways to install artifacts
mur install needs to know where to fetch artifacts from. You have two options:
Option A — configure a registry source in ~/.murmur/config.yaml:
registry:
default: official
sources:
- name: official
type: github
repo: <owner>/<repo>
token: "${GITHUB_TOKEN}"
Then install by artifact name and version:
mur install <artifact-name@version>
Option B — pass a full GitHub reference and skip configuration entirely:
mur install github:<username>/<repo>@<tag>
See Installing artifacts to learn more.
To install the compaction artifact directly without going through the manifest:
mur install murmur-hook-compact@1.0.0
Compaction is not built into the Murmur runtime — it is externalized as an artifact by design. Everything in Murmur is composable, including something as fundamental as how context is condensed. murmur-hook-compact is the default implementation: it summarizes the message history using the configured model and replaces it with a compact representation. But compaction strategy is not one-size-fits-all. A coding agent might benefit from keeping the full tool call history and compacting only prose; a research agent might maintain a structured memory store rather than a rolling summary. You can build any of these as a hook artifact, package it once, and swap it in by changing a single line in the manifest. The runtime does not care which artifact provides compaction — only that one is declared with runtime: hook and responds to the compaction lifecycle event.
Step 3 — set the compaction threshold
The threshold controls when compaction fires. It is a fraction of context.max_tokens. The default is 0.98 — compaction fires when the session has consumed 98% of the token budget.
For long-running tasks where you want compaction to kick in earlier and leave headroom for recovery, add compaction.threshold under inference:
inference:
transport: http
endpoint: https://api.anthropic.com
model: claude-sonnet-5
api_key: ${ANTHROPIC_API_KEY}
driver:
artifact: murmur-driver-anthropic
compaction:
threshold: 0.85
inference:
transport: http
endpoint: https://api.openai.com
model: o3-mini-high
api_key: ${OPENAI_API_KEY}
driver:
artifact: murmur-driver-openai
compaction:
threshold: 0.85
inference:
transport: http
endpoint: https://api.deepseek.com
model: deepseek-r1
api_key: ${DEEPSEEK_API_KEY}
driver:
artifact: murmur-driver-deepseek
compaction:
threshold: 0.85
With threshold: 0.85, compaction fires when the session has consumed 85% of context.max_tokens.
To use a smaller, faster model for compaction calls (saving cost while keeping your primary model for inference):
inference:
compaction:
threshold: 0.85
model: claude-haiku-4-5
inference:
compaction:
threshold: 0.85
model: gpt-4o-mini-2024-07-18
inference:
compaction:
threshold: 0.85
model: deepseek-v4-flash
To steer what the compaction call preserves, set inference.compaction.system_prompt. The string is passed to the compaction hook verbatim — no trimming, no length limit, no templating:
inference:
compaction:
threshold: 0.85
system_prompt: |
task = X, currently editing Y, already tried Z.
Preserve this nuance when summarizing.
For a longer, versioned instruction set, keep it in its own file and point at it with system_prompt_file instead. The path is resolved relative to the manifest directory and read when the session launches, and the file's contents are passed to the hook verbatim:
inference:
compaction:
threshold: 0.85
system_prompt_file: compaction-instructions.md
Setting both system_prompt and system_prompt_file on the same compaction: block is a manifest error — pick one. mur deploy run uploads the referenced file alongside the manifest, so a remote run resolves it the same way.
If left unset, compaction-event.system-prompt arrives at the hook as none, and the hook falls back to its own built-in default prompt. model and the two prompt-source fields are independent — setting one does not require or affect the other.
Step 4 — run and confirm the configuration
mur run
At startup the runtime writes a generated MURMUR.md to the session workdir. Its Capsule section reports compaction status — check it to confirm compaction was configured correctly:
grep "Context budget" workdir/<session_id>/MURMUR.md
When both context.max_tokens and a hook bound to on-compaction are staged, the status reads compaction configured:
- Context budget: 1000000 tokens (compaction configured)
If either is missing — context.max_tokens is unset, or no on-compaction hook is staged — it reads compaction not configured:
- Context budget: 1000000 tokens (compaction not configured)
The runtime selects the compaction hook by its on-compaction binding, not by name, so there is no artifact-name setting to get wrong: any staged hook with that binding satisfies the check.
Step 5 — verify compaction ran using the trace
Run a task that you expect to exceed the threshold, then check the trace:
mur trace show
Different ways to identify a session
mur trace show with no argument reads the most recent session:
mur trace show
To name another one, pass an ordinal counting back from the newest (@2), the last 4 or more
characters of its ID (3e4b), the full ID, or a path to its trace.jsonl:
mur trace show @2
mur trace show 3e4b
mur trace show ses_6801f81dd28b4a9daf434e8324c4793e
mur trace show path/to/trace.jsonl
Use --workdir <path> if your session directories are not under ./workdir. Every command
that names a session takes the same addresses — see
Session addresses.
Other trace exploration commands
mur trace has four subcommands for exploring session output:
mur trace show — print the full trace for a session to the terminal.
mur trace steps — show a turn-by-turn summary of what the agent did in a session. Pass --verbose to include a truncated summary of each tool's input.
mur trace diff — compare the traces of two sessions side by side, or with no arguments the two most recent. Useful for spotting behavioural regressions between runs.
mur trace report — generate a structured summary report from a session's trace. Covers token usage, tool calls, latency, and other session-level metrics.
If compaction fired:
── Compaction ───────────────────────────────────
fired: yes at turn 1 (12,325 → 5,488 tokens)
If it did not fire:
── Compaction ───────────────────────────────────
fired: no
Compaction does not consume a turn slot — inference.max_turns counts inference calls, not compaction events. The model continues from where it left off with the condensed history.
Step 6 — inspect the compaction summaries
trace.jsonl's compaction event records only the token counts, not the summary text the model
produced. To persist the summary itself — for understanding what context the agent retained
versus discarded, or debugging cases where the agent appears to "forget" earlier work after
compaction — set dump_summaries: true:
inference:
compaction:
threshold: 0.85
dump_summaries: true
Each committed compaction then appends one line to out/compaction-summaries.jsonl:
{"turn":17,"tokens_before":81501,"tokens_after":334,"summary":"1. THE BUG: ..."}
See out/compaction-summaries.jsonl for the full field reference. The file is only created once the first compaction actually commits, and a compaction the tool-call-pairing safety net rejects writes nothing.
Compaction failure modes
Four situations can prevent compaction from replacing the context, and the runtime treats them differently:
No hook bound to on-compaction. This is not a failure — it means compaction was never
configured to begin with. The runtime logs no hook returned replace-context to
logs/bootstrap.log, writes a compaction_declined line to trace.jsonl with
reason: "no_hook_replacement", and the session continues with the uncompacted history.
A bound hook returned a replacement the runtime rejected. A compacted history whose tool calls
and tool results no longer pair up would make the next request malformed, so the runtime discards
the replacement whole and keeps the original, larger history. Non-fatal: it logs
compacted result has an unresolved tool_call to logs/bootstrap.log, writes a
compaction_declined line with reason: "unresolved_tool_call", and the session continues over
budget. Recurring unresolved_tool_call declines mean the hook is dropping one half of a
call/result pair.
A bound hook ran and returned an error, with no refused run-inference call behind it. This
is treated as a failure, and it ends the
session: there is no fallback compactor behind a declared compaction hook, so continuing would
mean another inference turn on a context already known to be over budget. out/result.txt
records the error, the session's trace.jsonl records session_end with exit_status: "failed",
OTel emits session_end as "failed" (if observability.otel_endpoint is configured), and — if
the session has a task_id — the final SSE status event reports state: "failed". The agent
loop does not attempt another turn after this.
A bound hook returned an error after a spend ceiling refused its run-inference call. This
ends the session as a spend stop, not as a failure. The hook's call crossed
inference.max_session_tokens or
spend.machine_tokens_per_day, so nothing was sent:
| Surface | What it records |
|---|---|
trace.jsonl |
One spend_ceiling_reached line with origin: "hook:<name>"; task_end and session_end with exit_status: "spend_ceiling_reached" |
out/result.txt |
stopped: spend ceiling reached: … |
| OTel | session_end as "spend_ceiling_reached" |
| SSE | A final status event with state: "failed" and the refusal text as its message |
If your compaction hook can fail (for example, the model it calls for summarization is unreachable), account for the fact that this ends the session rather than silently skipping compaction.
Summary
| Manifest setting | Effect |
|---|---|
context.max_tokens: N |
Sets the token budget; required to enable compaction |
murmur-hook-compact with runtime: hook (or any hook bound to on-compaction) |
Provides the compaction hook; required to enable compaction |
inference.compaction.threshold: 0.85 |
Compaction fires when session tokens reach 85% of the budget |
inference.compaction.model: claude-haiku-4-5 |
Uses a different (typically cheaper) model for compaction calls |
inference.compaction.dump_summaries: true |
Appends every committed compaction's summary to out/compaction-summaries.jsonl; default false |
No hook bound to on-compaction |
Non-fatal — session continues with the uncompacted history; compaction_declined written to the trace |
| A bound hook returns a replacement with an unpaired tool call | Non-fatal — the replacement is discarded and the session continues with the uncompacted history; compaction_declined written to the trace |
A bound hook returns an error with no refused run-inference call behind it |
Fatal — the session ends as failed; see Compaction failure modes |
A bound hook returns an error after a spend ceiling refused its run-inference call |
Fatal — the session ends with exit_status: "spend_ceiling_reached", not failed; see Compaction failure modes |
| Token count after compaction | Reset to the count of the new (compacted) history, not to zero |