kompact — where to compact, and how much
Study · the loop’s two settings
There is a ceiling, and it is not ours.
kompact answers a compaction at compactAtPercent of the window. Claude Code
auto-compacts at its own threshold, and it is the engine, so it goes first.
That settles the question before any measurement. A trigger at or above the engine’s means kompact is installed, logging, and never answers a single compaction: the worst shape of failure, because nothing looks broken. So the usable range sits below it, and the only question left is how far below — which is a question about how much one more turn can cost.
The trigger has to leave room for one more turn to land
8 points of headroom5.97pp the largest single message0 of 12,452 messages would overrun it
The clamp enforces this rather than trusting a settings file: a configured trigger at or above the engine’s is refused and pulled back under it. A value that silently disables the thing you installed is not a value worth honouring.
The grid
Every trigger that can run, against every floor worth trying.
Each cell replays real sessions through the shipped rule and counts the engine summaries the loop deferred. More is better, and zero means the loop never ran.
| trigger | 3pp | 4pp | 5pp | 7pp | 10pp | 15pp |
|---|---|---|---|---|---|---|
| 50% | 13 | 5 | 2 | 1 | 1 | 0 |
| 55% | 14 | 11 | 6 | 3 | 1 | 0 |
| 58% | 14 | 10 | 9 | 5 | 1 | 0 |
| 60% | 14 | 11 | 9 | 7 | 1 | 0 |
| 62% | 14 | 12 | 10ships | 7 | 3 | 0 |
| 64% | 12 | 10 | 10 | 8 | 4 | 0 |
| 66% | 12 | 11 | 10 | 8 | 2 | 0 |
| 68% | 13 | 11 | 10 | 7 | 5 | 0 |
| 69% | 13 | 12 | 10 | 8 | 5 | 0 |
bun eval/trigger-page.ts --measure re-runs it on yours.Two things fall out of it. Raising the trigger helps, because waiting lets more compactible material accumulate and each pass clears the floor more easily — but the gain plateaus well below the ceiling, and every point above the plateau spends headroom for nothing. Raising the floor does not help at all. It is not a dial between thrift and safety: past a point it is an off switch.
The floor
A stricter floor does not buy safety. It buys nothing.
The intuition is that a higher bar means fewer, better-earned passes. What it means in practice is no passes — and the lowest bar measured is the one that reclaims the most.
At the trigger that ships, what each floor reclaims
- 3pp97.714 deferred
- 4pp91.512 deferred
- 5pp ships82.310 deferred
- 7pp64.77 deferred
- 10pp33.13 deferred
- 15pp0.00 deferred
Why does it collapse so fast?
Because a pass reclaims single-digit points of the window, not tens. Compaction only ever touches tool inputs and outputs, and those are a little over half a transcript; the rest is prose it never reads. A bar set above what a pass can physically reclaim is refused every time, and a refusal hands that compaction to the model summary.
So why have a floor at all?
It is the hysteresis band. A taken pass leaves the session at least that far below the trigger, so it has to grow back through those points before another compaction can be asked for. Without it, compacting on every turn is possible rather than merely discouraged. The floor’s job is spacing, not thrift.
Then why not ship the lowest one?
Honestly: it might be right, and nothing here has shown otherwise. It reclaims more, defers more, and leaves the same share of receipts behind. What it also does is compact a session more often for smaller gains each time, and how often a session should be interrupted is not something any of these measurements can answer. The setting stays where it is until that is measured rather than argued, and this paragraph exists so the gap is not silent.
What went wrong
The first version of this study recommended disabling the loop.
It is on this page because the failure is more useful than the finding, and because a study that reports only the run that worked is not a study.
The first sweep ran on the labelled corpus, which is the reproducible one: every call in it carries the message index that first quoted its output verbatim, so it can answer the question that matters most — whether a dropped output was one the next ten messages went on to need. That corpus carries tool output and nothing else.
So the window it fills is made entirely of the most compactible material there is. A pass that reclaims a quarter of that window reclaims a small single-digit share of a real one. The sweep reported a far stricter floor as costing a few per cent of what is freed, and recommended it. Measured against whole transcripts with honest token accounting, the same floor takes no passes at all.
What caught it was this repository’s own snapshot, which had recorded a floor sweep on
real transcripts long before and disagreed flatly. The check took one grep. The
recommendation had already been written down.
What survives from it?
The trigger axis, because a trigger is a position in the session rather than a quantity of window, and both measurements agree on its direction. And the continuity column, which nothing else can produce: across every setting in that grid, not one dropped output was one whose first reuse fell within the next ten messages.
What changed so it cannot happen again?
eval/trigger.ts prints the limitation above its own table now, rather than
keeping it in a comment nobody reads at the moment of believing a number. And the eval
scripts read the shipped defaults from the hook instead of keeping their own copies, so a
published figure and the plugin cannot drift apart.
What changed
The trigger moved. The floor did not.
One setting, a short distance, for one more deferred model summary and one more session that loops at all.
The honest size of it: the gain is a single deferred summary across five sessions, and those sessions are one person’s. The direction is consistent across both studies and every column of the grid; the magnitude is thin, and the ceiling it stops below is a documented default rather than anything measured here. If Claude Code’s threshold is configured differently on your machine, this is the setting to move with it.
Reproduce it: bun eval/trigger-page.ts --measure re-runs the grid against your own
transcripts and rewrites both figures above. bun eval/passes.ts --at N --floor N is
one cell of it. bun eval/trigger.ts is the labelled-corpus study, limitation and
all.