kompact — where to compact, and how much

Study · the loop’s two settings

There is a ceiling, and it is not ours.

kompact answers a compaction at compactAtPercent of the window. Claude Code auto-compacts at its own threshold, and it is the engine, so it goes first.

That settles the question before any measurement. A trigger at or above the engine’s means kompact is installed, logging, and never answers a single compaction: the worst shape of failure, because nothing looks broken. So the usable range sits below it, and the only question left is how far below — which is a question about how much one more turn can cost.

The trigger has to leave room for one more turn to land

8 points of headroom5.97pp the largest single message0 of 12,452 messages would overrun it

Every message of the 5 largest sessions on one machine, priced as points of a 200k‑token window: median 0.03, 90th percentile 0.32, 99th 1.41, largest 5.97. Only 2 of 12,452 cost more than 5 points, so a trigger at 62% has room for the worst turn in the corpus to land before the engine reaches 70%. The engine’s 70% is Claude Code’s documented default, not something measured here: it is the assumption the whole setting rests on, and it is on the ledger as one.

The clamp enforces this rather than trusting a settings file: a configured trigger at or above the engine’s is refused and pulled back under it. A value that silently disables the thing you installed is not a value worth honouring.

The grid

Every trigger that can run, against every floor worth trying.

Each cell replays real sessions through the shipped rule and counts the engine summaries the loop deferred. More is better, and zero means the loop never ran.

Engine summaries deferred, by trigger and floor
trigger3pp4pp5pp7pp10pp15pp
50%1352110
55%14116310
58%14109510
60%14119710
62%141210ships730
64%121010840
66%121110820
68%131110750
69%131210850
Engine summaries deferred over 5 real sessions on one machine, at every trigger below the engine’s own and every floor worth trying. Darker is more deferred. A 0 is a setting where the loop never takes a pass at all: the plugin is installed, it logs, and it answers nothing. Every trigger at 15pp is that, which is why the floor did not go up; whether it should go down is the open question below. bun eval/trigger-page.ts --measure re-runs it on yours.

Two things fall out of it. Raising the trigger helps, because waiting lets more compactible material accumulate and each pass clears the floor more easily — but the gain plateaus well below the ceiling, and every point above the plateau spends headroom for nothing. Raising the floor does not help at all. It is not a dial between thrift and safety: past a point it is an off switch.

The floor

A stricter floor does not buy safety. It buys nothing.

The intuition is that a higher bar means fewer, better-earned passes. What it means in practice is no passes — and the lowest bar measured is the one that reclaims the most.

At the trigger that ships, what each floor reclaims

  1. 3pp97.714 deferred
  2. 4pp91.512 deferred
  3. 5pp ships82.310 deferred
  4. 7pp64.77 deferred
  5. 10pp33.13 deferred
  6. 15pp0.00 deferred
Points of a 200k‑token window reclaimed across 5 sessions, and the engine summaries each floor deferred, at a 62% trigger. The lowest floor measured reclaims the most and defers the most, and the share of surviving results that is a truncation note is the same at 3pp as at 5pp. What 5pp buys is a wider band a session has to grow back through before it can be compacted again — a judgement about how often to interrupt a session rather than a measurement. It is the open question on this page.

Why does it collapse so fast?

Because a pass reclaims single-digit points of the window, not tens. Compaction only ever touches tool inputs and outputs, and those are a little over half a transcript; the rest is prose it never reads. A bar set above what a pass can physically reclaim is refused every time, and a refusal hands that compaction to the model summary.

So why have a floor at all?

It is the hysteresis band. A taken pass leaves the session at least that far below the trigger, so it has to grow back through those points before another compaction can be asked for. Without it, compacting on every turn is possible rather than merely discouraged. The floor’s job is spacing, not thrift.

Then why not ship the lowest one?

Honestly: it might be right, and nothing here has shown otherwise. It reclaims more, defers more, and leaves the same share of receipts behind. What it also does is compact a session more often for smaller gains each time, and how often a session should be interrupted is not something any of these measurements can answer. The setting stays where it is until that is measured rather than argued, and this paragraph exists so the gap is not silent.

What went wrong

The first version of this study recommended disabling the loop.

It is on this page because the failure is more useful than the finding, and because a study that reports only the run that worked is not a study.

The first sweep ran on the labelled corpus, which is the reproducible one: every call in it carries the message index that first quoted its output verbatim, so it can answer the question that matters most — whether a dropped output was one the next ten messages went on to need. That corpus carries tool output and nothing else.

So the window it fills is made entirely of the most compactible material there is. A pass that reclaims a quarter of that window reclaims a small single-digit share of a real one. The sweep reported a far stricter floor as costing a few per cent of what is freed, and recommended it. Measured against whole transcripts with honest token accounting, the same floor takes no passes at all.

What caught it was this repository’s own snapshot, which had recorded a floor sweep on real transcripts long before and disagreed flatly. The check took one grep. The recommendation had already been written down.

What survives from it?

The trigger axis, because a trigger is a position in the session rather than a quantity of window, and both measurements agree on its direction. And the continuity column, which nothing else can produce: across every setting in that grid, not one dropped output was one whose first reuse fell within the next ten messages.

What changed so it cannot happen again?

eval/trigger.ts prints the limitation above its own table now, rather than keeping it in a comment nobody reads at the moment of believing a number. And the eval scripts read the shipped defaults from the hook instead of keeping their own copies, so a published figure and the plugin cannot drift apart.

What changed

The trigger moved. The floor did not.

One setting, a short distance, for one more deferred model summary and one more session that loops at all.

The honest size of it: the gain is a single deferred summary across five sessions, and those sessions are one person’s. The direction is consistent across both studies and every column of the grid; the magnitude is thin, and the ceiling it stops below is a documented default rather than anything measured here. If Claude Code’s threshold is configured differently on your machine, this is the setting to move with it.

Reproduce it: bun eval/trigger-page.ts --measure re-runs the grid against your own transcripts and rewrites both figures above. bun eval/passes.ts --at N --floor N is one cell of it. bun eval/trigger.ts is the labelled-corpus study, limitation and all.