kompact — the evidence

05What you repeat

The same file, arriving again.

Deciding what to keep means reading every tool call. The same reading says what your sessions spend on rediscovering things they already knew.

128 of the 602 files read were read more than once. Between them: 1,088 reads and 619k tokens, 91.1% of every token this machine spent reading a file.

  1. docs/index.html217k233 reads · 5 sess
  2. waveshare7b_3D/build_7b_case.py31k102 reads · 1 sess
  3. eval/RESULTS.md31k35 reads · 5 sess
  4. test/published-figures.test.ts24k43 reads · 3 sess
  5. README.md20k40 reads · 2 sess
  6. src/compact.ts17k36 reads · 2 sess
  7. plans/curried-cooking-abelson.md13k47 reads · 1 sess
  8. laya-compact/.impeccable.md12k3 reads · 2 sess
  9. eval/render-check.ts10k19 reads · 4 sess
  10. index.html10k5 reads · 2 sess
  11. src/features.ts10k8 reads · 4 sess
  12. .impeccable.md9k7 reads · 3 sess
  13. types/claude-code.d.ts9k15 reads · 1 sess
  14. hooks/laya-compact.ts9k17 reads · 3 sess
The files rediscovered most, over 8 sessions and 4,659 tool calls on one machine. Each read is the file arriving in context again. This is what the repetition cost, counted from the transcripts — not what anything would save: nothing here has measured whether writing a skill, or a line in CLAUDE.md, changes a single one of these reads, and a skill does not stop you needing to read a file. One operator’s transcripts, and a large share of the top row is this site being edited. bun eval/repetition.ts runs it on yours.

Repetition is 41% of the tokens and 19% of the time. The time is the smaller half and the less fixable: of 374 min of waiting across those sessions, 327 min was a tool running (4,576 calls), 38 min was a person deciding (33 calls), 10 min was a subagent working (50 calls). Gaps longer than 120 s are counted as 120 s, because a result arriving an hour later is a session left open rather than a tool that ran for an hour.

What this was built to find, and did not

It was built to propose skills. The plugin counts the shapes it sees, on your machine, and npm run propose ranks them and drafts a SKILL.md for any row you pick. On this corpus that fails, and it fails in a way worth publishing: the shapes that repeat are the verbs, and the verbs are not where the cost is.

69 shapes that repeated, of 2,000 recorded over 40 sessions and 6,585 tool calls

The 10 costliest of the 69 shapes that repeated

KindSignaturetimessessionscallscharsest.
the same commandsed -n4183418828,7691246.8
the same commandpython3 -c118311894,954213.0
the same commandcat <path>40240112,749152.7
the same commandgrep -n | head -<n>8528565,499150.5
the same run of toolsRead → Edit → Edit393117—117.0
the same run of toolsBash(python3 -) → Bash(python3 -c) → Bash(python3 -)28284—84.0
the same run of toolsBash(python3 -c) → Bash(python3 -) → Bash(python3 -)25275—75.0
the same run of toolsRead → Read → Edit24372—72.0
the same run of toolsBash(grep -n) → Bash(sed -n) → Bash(sed -n)24272—72.0
the same run of toolsBash(grep -nE) → Bash(sed -n) → Bash(python3 -)19257—57.0

est. = total tool calls + total output characters / 1000, the ranking chosen for this. A model of effort, not a measurement of time; every input to it is on the row, so any row can be recomputed by hand. Greyed rows are generic shell verbs.

Of 2,000 shapes recorded, 69 repeated enough to propose, and 13 of the 17 repeated command shapes are generic shell verbs. The highest-ranked thing this found is sed -n.

13 of the 17 are verbs like sed, grep or cat. No skill helps with those. The 4 that fall outside that list (git diff | head -<n> and git add && git commit && git log and git status --short and git add && git commit && git log | head -<n>) are not workflows either; they are simply verbs the list was not written to catch, and adding them to it to keep a sentence tidy is the kind of tuning this page exists to avoid. The best row that reads like an actual workflow, Bash(python3 -) → Bash(python3 -c) → Bash(python3 -), is ranked 6th, because the ranking rewards total work and a generic verb runs more often than a workflow does.

shell · the report, from the corpus above

npm run propose:demoRun: eval/fixtures/signals.json69 shapes recorded; 69 seen 3+ times in 2+ sessions.​What repeating costs you    #  times sess  calls     chars    est.  what    1    418    3    418   828,769  1246.8  the same command: sed -n    2    118    3    118    94,954   213.0  the same command: python3 -c    3     40    2     40   112,749   152.7  the same command: cat <path>    4     85    2     85    65,499   150.5  the same command: grep -n | head -<n>    5     39    3    117         0   117.0  the same run of tools: Read → Edit → Edit    6     28    2     84         0    84.0  the same run of tools: Bash(python3 -) → Bash(python3 -c) → Bash(python3 -)    7     25    2     75         0    75.0  the same run of tools: Bash(python3 -c) → Bash(python3 -) → Bash(python3 -)    8     24    3     72         0    72.0  the same run of tools: Read → Read → Edit    9     24    2     72         0    72.0  the same run of tools: Bash(grep -n) → Bash(sed -n) → Bash(sed -n)   10     19    2     57         0    57.0  the same run of tools: Bash(grep -nE) → Bash(sed -n) → Bash(python3 -)   11     14    2     42         0    42.0  the same run of tools: Bash(python3 -) → Bash(python3 -) → Bash(python3 -c)   12     14    2     42         0    42.0  the same run of tools: Bash(python3 -c) → Bash(python3 -c) → Bash(python3 -)​  est. = total tool calls + total output characters / 1000. A model of effort,  not a measurement of time, and every input to it is on the row.​Scripts you keep re-writing — a tool, not a skill    #  times sess  calls     chars    est.  what   13     14    2     14    22,692    36.7  the same ad-hoc script: tool:python:json,json.load,load,open   14     17    2     17    11,391    28.4  the same ad-hoc script: tool:python:json,json.load,load,sys   15      4    2      4     6,755    10.8  the same ad-hoc script: tool:python:dumps,json,json.dumps,json.load,load,open   16      3    2      3     1,882     4.9  the same ad-hoc script: tool:python:is_available,torch,torch.cuda.is_available   17      4    2      4       372     4.4  the same ad-hoc script: tool:python:Image,Image.open,PIL,open​  Same ad-hoc script (a python -c, a heredoc) written again from scratch,  grouped by what it does rather than its filename. These want a saved tool,  not a skill: write it once, then call it.​Nothing written. `npm run propose -- --write 1,3` drafts those rows into .kompact/proposals/.

Real output, captured from eval/propose.ts when this page was built, over the same committed corpus as the table above. Not a mock-up. npm run propose is the same report over what the recorder saw on your machine, which is nothing until the plugin has been installed a while.

Watch 3,228 commands land on the shapes they repeat

The classifier works and the counts are real. What is not established is the premise underneath: these sessions are an assistant’s tool use, and what an assistant repeats is reading files, editing them, and reaching for the shell. What that cost to find out, and what this table is not, is on the evidence page.

06Cost

281 ms to score 3,591 tool calls.

Compaction stops being a model call and becomes arithmetic. These are whole session logs, so the totals span more than one context window, so the ratio is what transfers.

per-session results on this machine
SessionCalls Tokens beforeAfter FreedScoring
206289213,0032,111,8141,435,06232.0%247 ms
25e65eab473273,749224,48418.0%26 ms
agent-ab4cc54943,40039,4519.1%3 ms
agent-afa3fd6683,02275,7688.7%5 ms
Total3,5912,511,9851,774,76529.3%281 ms

Modest on purpose. The absolute cut this replaced frees 95.3% of the labelled corpus and keeps 8.9% of the outputs that were reused later: the first row of the policy table below, and the reason a third is the honest number. Measured on one machine’s own transcripts, which grow as work continues. The per-session column is the figure that carries.

What it replaces, timed. Over 11 real auto‑compactions on this machine, Claude Code’s own summariser took a median 133 s to produce (65 s to 224 s), each cutting the transcript by about 95%. kompact answers the same compaction with a pass measured in milliseconds. This is the summary’s cost to produce, read from Claude Code’s own durationMs — which times the whole compaction request, not only the pause on screen, on one machine. What the assistant loses by not getting that narrative is a different question and stays not verified.

And what each keeps alive. The built-in summary cut the window to a median 35k tokens; kompact’s own passes here (3, median 183 ms) left 262k — it fires earlier and compresses less, so more of the transcript stays verbatim. It does not compact less often; each pass is milliseconds, and the slow summary is left only for the residue kompact cannot drop.

Get this table for your own machine first

Clone the repository and run npm run dry-run. The commands are in section 09, Install. It reads ~/.claude/projects, scores every tool call in every transcript it finds, and prints the table above for your sessions instead of these. It writes nothing, installs no plugin, and needs no key: it is the same compact the hook calls, run against your own history. If the number it prints is not worth having, you have lost a minute.

Expect somewhere in 8.7% to 32.0%, the range four sessions covered, which is not a distribution. What moved it was tool mix and length: the 32.0% session is 3,003 calls of mostly commands whose output is never quoted afterwards, and the two at the bottom are 49 and 66 calls long, where there is almost nothing to free. At four sessions that is a reading of these four and not a rule. Nothing here has measured how much of the spread is noise. The dry run tells you which of those you are.

Why it frees a third and not four fifths

The scorer produces a ranking, and a ranking transfers across sessions better than an absolute probability cut: 0.789 held out one session at a time against 0.879 on the narrower corpus, so it degrades rather than collapses. A cut does not travel at all, because every session has its own mix of tools and so its own distribution: a Bash-heavy session scores low throughout, and a fixed cut sweeps all of it. Barely 11% of tool outputs are ever reused verbatim (247 of 2,239), so a well-calibrated scorer seldom passes 0.5 even for the ones that matter.

Calls are therefore dropped lowest-score-first only until the reduction target is met, with the keep threshold acting as a floor that nothing is dropped above. Simulated per session on the labelled corpus, scored leave-one-session-out:

the policy sweep
PolicyOutput freed Reused outputs keptReused characters kept
Absolute cut at 0.595.3%8.9%16.9%
Budget 0.5, floor 0.2 (shipped)23.4%77.3%92.2%
Budget 0.6, floor 0.228.3%72.1%91.4%
Budget 0.7, floor 0.233.9%68.8%90.2%

What the cap costs, priced. The 33.6% above includes it; so does this. Of the 247 outputs that were genuinely reused later, 5 are longer than the cap, and shortening them removes 76,466 characters those later steps had quoted, 6.2% of every reused character. Counting the wrong drops beside it, 88.1% of reused characters survive the shipped defaults and 11.9% do not. The gain and the loss are measured the same way, because for a while only the gain was.

Freeing nearly everything is easy and nearly worthless. Run through the shipped code rather than this simulation, the same defaults give 33.6% freed and 79.4% kept. The sweep has no notion of two shipped rules: a call recording a change is never dropped, and a kept result is capped at 24,000 characters. The cap is most of the gap: 27 of the 2,239 calls are longer than it, and shortening those 27 alone frees 20.7% of every character in the corpus. Calibrating on your own sessions is an accuracy upgrade rather than a prerequisite, because the policy adapts to your distribution on its own.

What the neural sidecar would have cost to run

The neural model is not in this product any more. There is no option that turns it on, and nothing shipped imports it. What it cost to run is a large part of why. The eval/ scripts still measure all of this against a live sidecar, so every row below is reproducible rather than remembered. One router process loads every checkpoint and routes per request, so memory and start-up belong to the router rather than to any one checkpoint.

what the sidecar would have cost to run
Scoring 3,591 calls, the four sessions aboveTime Memory heldVRAM Cold start
Built-in scorer281 msnonenonenone
Neural sidecar, fastest checkpoint, on a GPU4.9 s3.2 GB5.8 GB6.1 s

17× the time, for a lower AUC, on an RTX 4060 Laptop. Per request the fastest checkpoint answers eight questions in 653 ms and the slowest in 1,889 ms. The VRAM is the router with all three checkpoints resident, which is what the benchmark starts; one checkpoint alone is about 1.4 GB. Fine-tuning moves the AUC; it does not move this table.

This row has been wrong twice, and the second time is the more useful story. It first read 23.1 s and 122×, from a projection that divided requests by the number of questions in a request, but a request carries one call's two questions, so there is one request per call and the projection had halved itself. It then read 90.7 s and 319× for weeks, because the benchmark behind it was measuring a sidecar started with LAYA_DEVICE=cpu while the card sat idle. The script printed the idle card two lines above its own table and nobody read it. On the GPU the same work takes 4.9 s, so the magnitude was wrong by a factor of eighteen. eval/sidecar-bench.ts now reads the sidecar's own /health, which reports its device in one field, and refuses to print publishable rows from a CPU one.

One number was never the cost of running this model. It was the cost of one deployment of it. laya-serve routes /health and /v1/systemone and nothing else, so over HTTP every call is one state in one forward pass in one round trip, which is the slowest thing the model can do; the library packs many states into shared passes. Measured on the same card, the wire costs about 0.3 ms of that round trip and batching roughly halves what is left. The evaluation page prints the whole ladder. Even at its fastest the sidecar loses this comparison, which is the point of publishing the ladder rather than the one number that flattered our own conclusion.

07Where it loses

The model this replaced fails quietly, which is worse than failing.

The English checkpoint reads 512 tokens and discards the rest with HTTP 200 and no warning. eval/truncation.ts asks one question (the text says the deployment was rolled back) of four states. Every one of them states the fact verbatim. Only the position of the sentence changes.

the truncation test
StateCharacters Tokens readAnswer
The sentence alone85470.8472
First, plus filler8,1405120.9152
First, plus 5× filler40,3605120.9152
Last, after 5× filler40,3605120.2191

Rows two and three are bit-identical: 32,220 characters were discarded and nothing in the response says so. Row four is the bill for that. It is the same 40,360 characters and the same sentence, moved to the end where the cut removes it, and the answer falls from 0.9152 to 0.2191 for a fact the text still states verbatim: from the right side of any threshold to the wrong one. A caller cannot tell rows three and four apart: same HTTP 200, same shape, same silence. Whether an answer is about your input or about the first 512 tokens of it is settled by where the evidence happened to fall. Both upstream projects send the whole conversation, up to 25,000 tokens, as one state.

This table used to be three rows that no script produced. The one claim here with no code behind it, on a page arguing that every figure is generated. It is npm run eval:truncation now, and the figures it prints are these. The ones it replaced are retracted: they showed the cut but not what the cut costs.

What the other 20.6% costs

“79.4% of reused outputs kept” invites you to supply your own answer for the rest, so here it is, priced. (The sweep above says 77.3% and this says 20.6%, which sum to 97.9 rather than 100. The gap is the two rules the sweep has no notion of: a call recording a change is never dropped, and neither is a result nothing can produce again. Run through the shipped code the same defaults give 33.6% freed and 79.4% kept, which is the pair the masthead quotes.) Held-out scores, shipped policy, per session: 51 of the 247 reused outputs are dropped: 1.24 a session, and only one of them kept a head.

what a wrong drop costs
How it would be recoveredOutputs CharactersProduced by
Re-run the command, side effects unknown4742,389Bash
Read the file again453,680Read
Not repeatable at all00—

Only one of the 51 kept a head, because when a result scores below the floor the call usually does too. Forty-seven are command output that a later tool input quoted verbatim, the shape this scorer is worst at, since most command output is never referred to again and the minority that is gets swept along with it. The bottom row used to read 4 AskUserQuestion results, a human’s answer, which no amount of re-running brings back. It reads zero because a result nothing can produce again is no longer the scorer's to drop. This is recovery cost, and it over-counts: it includes reuses that had already happened when a real compaction fired. The next section measures what is actually at risk.

Compacting where the engine would actually fire, 28 of 32 sessions lost nothing at all. Across the rest, 5 outputs were dropped that a later step went back to: 0.16 a session, and 6.9% of what was genuinely still needed. Nothing here replays an assistant against the compacted transcript, so whether that changed the work stays not verified.

How that was counted, and what the looser count says

Every figure above asks whether an output was reused somewhere later, which counts reuses that had already happened by the time a compaction fired and could never have been lost. eval/outcome.ts asks it in the right order: find where the engine would actually compact (62% of a session's tool output), decide only the calls present at that moment, then count the reuses that come afterwards whose source was dropped.

what the work loses
Compacting at 62%, across 32 sessionsCount
Calls present when it fires1,071
Reused only after that point72
…and dropped anyway5
Sessions that lose nothing at all28 of 32

0.16 lost outputs a session, and 6.9% of what was genuinely still needed. This is the necessary condition for the work to suffer, not proof that it did: nothing here replays an assistant against the compacted transcript, and that remains unverified.

Five more things it does not do

A summary compresses harder. Freeing a third of the tokens is less than replacing the conversation with a few thousand tokens of prose. What you buy with the difference is not re-running tools to recover what a summary threw away.

One feature does most of the work. Output size alone scores 0.878 with an eighth of the spread (sd 0.010 against 0.078) and a sixtieth of the variance. The other coefficients earn their place on characters freed at a given safety level, not on ranking.

It has never seen a Grep. The corpus contains no Grep and no Glob call at all, so that feature never fired during fitting and its weight is 0.0: zero by absence, not by measurement. A search falls to the reference level, alongside WebFetch and Agent, about which this scorer also knows nothing. Thirteen features were defined; twelve were fitted.

The shipped coefficients are one person’s sessions. A different tool mix scores differently and nothing detects that for you, so npm run calibrate refits on your own transcripts and reports whether it helped.

It never shortens what a tool was asked. Every lever here acts on output. eval/mass.ts says output is 46.8% of a transcript and tool input is 43.6%. A Write carries the whole file it wrote and a Bash its heredoc, and nothing shortens any of it. eval/inputs.ts priced capping it and the trade is poor: capping every input at 16,000 characters frees 1.9% of the transcript for 1.3% of the passages later steps quoted back, against 20.7% for 6.2% on the output side. Reuse inside an input sits at median depth 0.50, so a head cap samples it rather than trims it, which is the same reason head-and-tail truncation lost. The full table includes the one tool that would be worth capping and why a default built on 26 calls from one machine would be a default fitted to this operator.

08Verification

Seven claims checked, five not.

The minimum Claude Code version — not verified

Everything here was run on 2.1.281, which has function hooks. Which release first shipped them has not been established, so “2.1.281 or newer” would be a claim from a sample of one. On a build without them the plugin installs and never runs.

An assistant still finishing the job — not verifiedblocks a 1.0

The measurement above shows how often the information would no longer be there. Whether that changes what the assistant does needs the same session replayed with and without compaction, which needs a live A/B.

That a proposed skill is worth having — not verified

The repetition in section 05, What you repeat is measured, and the drafts validate as skills. Whether writing one saves anyone anything is not measured at all: it needs the same work done twice, once with the skill and once without. Thirteen of the 17 repeated command shapes are generic shell verbs and the other four are git incantations, which is evidence against rather than for. The ad-hoc scripts it found are a better case, and get their own table above.

What deferring the summary costs — not verified

Six passes run before the model summary does, and how much room each one buys is measured. What the assistant loses by not getting a narrative summary for six compactions is not: this keeps verbatim tool calls and prose, which is a different thing from a story about them. maxPasses exists so the summary happens eventually, and its value is a judgement, not a result.

The engine calling session.compact live — not verifiedblocks a 1.0

The plugin loads and turn.complete fires in a real session, but the engine declines to compact below ~31% full and the model cannot invoke /compact itself. Covered by the real-session test; the engine-side handoff is not.

And the 7 claims that are verified

Beats the model it replaces

1,063 labelled calls, 18 sessions, 10 grouped splits, paired. eval/repeat.ts

No orphaned tool call survives

Run over a real session from disk. An orphan is rejected by the API and would break the session compaction was meant to save.

Prose is never touched

Same real-session test: only tool calls and results are dropped or truncated.

A dead scorer cannot break a session

Falls back to the built-in summary; a single failed request keeps its call.

Works for Codex CLI, against a fixture

jev-compact’s own parser, scorer and HTTP client driven over its recorded rollout fixture against our server. Not run against a live Codex: that plugin needs Codex ≥ 0.155 and this machine has 0.131.

What a compaction costs the work

Compaction simulated where the engine would fire it, then the reuses that come afterwards counted: 5 of 72, 0.16 a session, 28 of 32 sessions losing nothing. eval/outcome.ts

The sidecar truncates silently

One question asked of four states that all state the fact verbatim: two differing by 32,220 characters answer bit-identically, and moving the sentence past the 512-token cut takes the answer from 0.9152 to 0.2191. eval/truncation.ts

What repeats, and why the premise under it is not established

That is not the finding this was built to get. The classifier itself works: the counts are real, the signatures group work that belongs together, and three separate design ideas turned out to be wrong only because the data said so: 94% of command shapes were seen exactly once until a pipeline stopped keeping every stage's arguments, sequences never formed at all until the window slid across batches instead of inside one, and the 2,000-row cap cut the rarest kinds to nothing until each kind was pruned against itself.

What is not established is the premise underneath it. These sessions are an assistant's tool use, and what an assistant repeats is reading files, editing them, and reaching for the shell.

Two things this table is not. It is not from a live install: the recorder only sees sessions from the moment it is switched on, so these counts come from replaying existing transcripts through the recorder's own functions (eval/signals-fixture.ts) rather than from the hook having run. And it carries no examples. The signatures are shapes and safe to publish, but the samples kept beside them are raw text, and the first fixture built without stripping them passed all eight credential checks while carrying another project's brief and a path with a session id in it. On your own machine the samples are the point, because they are how a wrong grouping becomes visible. In anything published they have no place.

Nothing recorded leaves the machine, and nothing is written into a skills directory: drafts land in .kompact/proposals/, which Claude Code does not read. Set recordSignals to true to record.

The gaps, in the order they matter

Everything above this line is measured. Everything below it is not yet, and is written in the future tense on purpose.

The questions a sceptic asks

Isn't this just a heuristic with extra steps?

It is a heuristic with error bars, which is the difference. Thirteen coefficients (twelve features and an intercept) fitted on labelled data, evaluated on sessions held out whole, and calibrated well enough that the threshold means a probability rather than a rank; section 01, Evidence carries the error itself. And most of the signal really is one feature: “big outputs get reused” alone reaches 0.878. That is three coefficients of the twelve, since size is banded. The other nine earn their place on characters freed at a fixed safety level, not on ranking, which is the answer to the question.

Why did the neural model lose?

It was asked to read facts as prose that the features encode directly, and it was never trained on this task. The model’s own documentation says its base checkpoints sit near chance on typed-decision workflows and should be treated as a fast base to specialise. That is what was measured, and the wording mattered more than the checkpoint: on typed-decisions, asking what the assistant should do next scores 0.719, and asking what is true of the text scores 0.416, below the coin flip, which is worse than not scoring at all. That is the reverse of what that documentation suggests, which is the finding rather than an embarrassment: the checkpoint is better at a decision than at an entailment.

The checkpoint named for English is also not the best at English. On that same wording it scores 0.480 (below the coin flip) where the multilingual one scores 0.667, and the two ranges do not overlap: English tops out at 0.502, multilingual bottoms out at 0.638. English wins only on entailment, by 0.021, which is inside one standard deviation, and its best config still loses to multilingual's best. It also reads half the context, 512 tokens against 1024, and runs about three times slower. None of that is why the sidecar loses: typed‑decisions reads 1024 too and reaches 0.719, but it does mean there is no configuration in which the English checkpoint is the one to pick.

Will it delete something I need?

Sometimes, and section 07, Where it loses prices it exactly. The asymmetry is what the whole design turns on: a wrong keep costs context, a wrong drop costs work, so a dropped result keeps its first 300 characters, a call that recorded a change is never dropped, and every failure mode falls back to Claude Code's own summary rather than to a half-compacted transcript.

Do I have to calibrate?

No: spending a budget lowest-score-first adapts on its own, where an absolute probability cut could not survive a distribution it was not tuned on. So npm run calibrate is an upgrade rather than a prerequisite, and it is still worth running on your own sessions, for the reason in the ledger: these coefficients come from one person's.

How is this different from the built-in /compact?

The built-in asks a model to summarise the session and replaces it with prose, which compresses far harder than a third. This keeps what survives verbatim, so a file you read is still the file, not a description of it. It is not a better compressor; it is a different trade, and it hands the session back to the built-in summary whenever it cannot do its job.