An AI coding agent can spend most of its working memory remembering terminal output it could simply read again. In a new preprint, researchers found that deliberately dropping old tool results—without asking another model to summarize them—cut average model-usage costs by roughly half in some coding benchmarks while matching or improving task success.
The method is called CliffCompaction. It is not a smarter foundation model, a new training run or a general memory system. It is a rule for deciding what a long-running agent keeps in its prompt. The Carnegie Mellon University and Bosch preprint posted September 22 reports results on software-engineering, terminal and CUDA-kernel tasks. Those are demanding tests, but they are still benchmarks: the work does not show that forgetting improves every kind of agent.
Most of the context was machinery, not insight
A coding agent works in a loop. It reads files, searches repositories, runs commands, edits code and tests the result. Every interaction can be appended to the conversation sent back to the model. That history helps the agent remember what it tried, but it also becomes expensive: previous tokens have to remain available on later calls, and eventually the model reaches its context limit.
In the researchers’ analysis of GLM 5.1 on Terminal-Bench 2.0, tool results accounted for 56% of context tokens and tool calls another 28%. The proposed compactor attacks those two piles. Tool outputs longer than 500 characters are removed. Tool calls become short signatures that preserve the command or file path. The task, system instructions and most recent turns remain intact; older model reasoning is shortened.
CliffCompaction does not write a clever diary for the agent. It throws away bulky old receipts while keeping enough information to fetch them again. A missing file read is visible, so the agent can reopen the file. A fluent summary can be more dangerous: it may sound complete even when it quietly omitted the detail that matters.
The name describes what happens to prompt length. Context grows normally until it crosses a threshold, then falls sharply. At the next cliff, the previous compacted block is discarded rather than compressed again. That avoids a familiar failure mode: a summary of a summary can preserve the shape of a decision while slowly losing the evidence behind it.
Forgetting can force a useful reread
The trade is precision for recall. CliffCompaction keeps recent material verbatim, so what survives is faithful. But material from earlier windows can disappear altogether. If the agent still needs it, the preserved tool signature lets it issue the command again.
That visible gap may be a feature. In a SWE-bench analysis with Kimi K2.7 and a 16,000-token budget, strategies that dropped tool output caused more rereading. CliffCompaction added about five rereads per task relative to full context. Model-written summaries triggered fewer rereads, apparently because their plausible prose convinced the agent that it already knew enough. Rereading costs tokens, but it also returns the agent to the current source of truth.
On Terminal-Bench 2.0, Kimi K2.6 resolved 61.42% of tasks with CliffCompaction at 16,000 tokens, compared with 59.16% using full context and 55.45% using the benchmark scaffold’s summarization at the same compact budget. The reported average model-usage cost fell from $0.40 for full context to $0.19. GLM 5.1 also improved at 16,000 tokens, from 49.83% resolved with full context to 54.33%, while average model-usage cost fell from $0.54 to $0.27.
Those differences are not proof that smaller prompts make models universally more capable. They show that stale or low-value history can interfere with a coding agent, and that a cheaper representation can outperform one particular full-history run under matched benchmark conditions. At a very tight 8,000-token budget, performance fell for both model families.
The memory bill includes more than tokens
Long prompts also occupy the model’s key-value cache, the intermediate representation that lets a system reuse earlier computation. Compaction invalidates part of that cache and requires a new “prefill,” which is why constantly rewriting the prompt can erase theoretical savings. CliffCompaction waits until a threshold is reached, then makes one large cut and leaves the prompt stable again.
The strongest result is also the least general. On 50 KernelBench problems, the agent repeatedly optimized CUDA kernels across sessions that would exceed a million tokens of accumulated history. With Kimi K2.7, 400 compacted steps produced a reported 3.58-times geometric-mean speedup over the original kernels. The experiment rewards iterative code improvement in an environment where files persist and can be reread—the exact conditions that make dropping conversational history unusually forgiving.
The authors acknowledge that benefits depend on the scaffold and task, and they did not compare against systems that train a model to manage context or maintain a separate external memory. Their cost analysis also uses provider prices and cache assumptions that can change. A cheaper prompt is not the same thing as lower total engineering cost, especially if extra rereads add latency.
The practical question is no longer whether an agent should remember everything. It is which facts must be carried forward, which artifacts can be reopened, and which missing context will announce itself before it causes damage. The next useful test is outside code: a long-running research or operations agent where decisions depend on evidence that cannot always be reconstructed from the filesystem.
Keep exploring
- This AI rewrites its own playbook — a different route to improvement that changes an exploration policy while keeping the model fixed.
- AI search now has to choose what to retrieve—and what to train — why information selection is becoming part of capability.
- AI’s power problem is becoming a scheduling problem — another case where system design matters as much as the model.
AI-assisted. Sources checked.




