kompact — the evidence
05What you repeat
The same file, arriving again.
Deciding what to keep means reading every tool call. The same reading says what your sessions spend on rediscovering things they already knew.
128 of the 602 files read were read more than once. Between them: 1,088 reads and 619k tokens, 91.1% of every token this machine spent reading a file.
docs/index.html217k233 reads · 5 sesswaveshare7b_3D/build_7b_case.py31k102 reads · 1 sesseval/RESULTS.md31k35 reads · 5 sesstest/published-figures.test.ts24k43 reads · 3 sessREADME.md20k40 reads · 2 sesssrc/compact.ts17k36 reads · 2 sessplans/curried-cooking-abelson.md13k47 reads · 1 sesslaya-compact/.impeccable.md12k3 reads · 2 sesseval/render-check.ts10k19 reads · 4 sessindex.html10k5 reads · 2 sesssrc/features.ts10k8 reads · 4 sess.impeccable.md9k7 reads · 3 sesstypes/claude-code.d.ts9k15 reads · 1 sesshooks/laya-compact.ts9k17 reads · 3 sess
CLAUDE.md, changes a single one of these reads, and a skill does not stop you
needing to read a file. One operator’s transcripts, and a large share of the top row is
this site being edited. bun eval/repetition.ts runs it on yours.Repetition is 41% of the tokens and 19% of the time. The time is the smaller half and the less fixable: of 374 min of waiting across those sessions, 327 min was a tool running (4,576 calls), 38 min was a person deciding (33 calls), 10 min was a subagent working (50 calls). Gaps longer than 120 s are counted as 120 s, because a result arriving an hour later is a session left open rather than a tool that ran for an hour.
What this was built to find, and did not
It was built to propose skills. The plugin counts the shapes it sees, on your machine, and
npm run propose ranks them and drafts a SKILL.md for any row
you pick. On this corpus that fails, and it fails in a way worth publishing: the shapes that
repeat are the verbs, and the verbs are not where the cost is.
69 shapes that repeated, of 2,000 recorded over 40 sessions and 6,585 tool calls
The 10 costliest of the 69 shapes that repeated
| Kind | Signature | times | sessions | calls | chars | est. |
|---|---|---|---|---|---|---|
| the same command | sed -n | 418 | 3 | 418 | 828,769 | 1246.8 |
| the same command | python3 -c | 118 | 3 | 118 | 94,954 | 213.0 |
| the same command | cat <path> | 40 | 2 | 40 | 112,749 | 152.7 |
| the same command | grep -n | head -<n> | 85 | 2 | 85 | 65,499 | 150.5 |
| the same run of tools | Read → Edit → Edit | 39 | 3 | 117 | — | 117.0 |
| the same run of tools | Bash(python3 -) → Bash(python3 -c) → Bash(python3 -) | 28 | 2 | 84 | — | 84.0 |
| the same run of tools | Bash(python3 -c) → Bash(python3 -) → Bash(python3 -) | 25 | 2 | 75 | — | 75.0 |
| the same run of tools | Read → Read → Edit | 24 | 3 | 72 | — | 72.0 |
| the same run of tools | Bash(grep -n) → Bash(sed -n) → Bash(sed -n) | 24 | 2 | 72 | — | 72.0 |
| the same run of tools | Bash(grep -nE) → Bash(sed -n) → Bash(python3 -) | 19 | 2 | 57 | — | 57.0 |
est. = total tool calls + total output characters / 1000, the ranking chosen for this. A model of effort, not a measurement of time; every input to it is on the row, so any row can be recomputed by hand. Greyed rows are generic shell verbs.
Of 2,000 shapes recorded,
69 repeated enough to propose, and
13 of the 17 repeated command shapes are generic
shell verbs. The highest-ranked thing this found is
sed -n.
13 of the 17 are verbs like
sed, grep or cat. No skill helps with those. The 4 that fall outside that list
(git diff | head -<n> and git add && git commit && git log and git status --short and git add && git commit && git log | head -<n>) are not
workflows either; they are simply verbs the list was not written to catch, and adding them
to it to keep a sentence tidy is the kind of tuning this page exists to avoid. The best row that reads like an actual workflow, Bash(python3 -) → Bash(python3 -c) → Bash(python3 -), is ranked 6th, because the ranking
rewards total work and a generic verb runs more often than a workflow does.
shell · the report, from the corpus above
npm run propose:demoRun: eval/fixtures/signals.json69 shapes recorded; 69 seen 3+ times in 2+ sessions.What repeating costs you # times sess calls chars est. what 1 418 3 418 828,769 1246.8 the same command: sed -n 2 118 3 118 94,954 213.0 the same command: python3 -c 3 40 2 40 112,749 152.7 the same command: cat <path> 4 85 2 85 65,499 150.5 the same command: grep -n | head -<n> 5 39 3 117 0 117.0 the same run of tools: Read → Edit → Edit 6 28 2 84 0 84.0 the same run of tools: Bash(python3 -) → Bash(python3 -c) → Bash(python3 -) 7 25 2 75 0 75.0 the same run of tools: Bash(python3 -c) → Bash(python3 -) → Bash(python3 -) 8 24 3 72 0 72.0 the same run of tools: Read → Read → Edit 9 24 2 72 0 72.0 the same run of tools: Bash(grep -n) → Bash(sed -n) → Bash(sed -n) 10 19 2 57 0 57.0 the same run of tools: Bash(grep -nE) → Bash(sed -n) → Bash(python3 -) 11 14 2 42 0 42.0 the same run of tools: Bash(python3 -) → Bash(python3 -) → Bash(python3 -c) 12 14 2 42 0 42.0 the same run of tools: Bash(python3 -c) → Bash(python3 -c) → Bash(python3 -) est. = total tool calls + total output characters / 1000. A model of effort, not a measurement of time, and every input to it is on the row.Scripts you keep re-writing — a tool, not a skill # times sess calls chars est. what 13 14 2 14 22,692 36.7 the same ad-hoc script: tool:python:json,json.load,load,open 14 17 2 17 11,391 28.4 the same ad-hoc script: tool:python:json,json.load,load,sys 15 4 2 4 6,755 10.8 the same ad-hoc script: tool:python:dumps,json,json.dumps,json.load,load,open 16 3 2 3 1,882 4.9 the same ad-hoc script: tool:python:is_available,torch,torch.cuda.is_available 17 4 2 4 372 4.4 the same ad-hoc script: tool:python:Image,Image.open,PIL,open Same ad-hoc script (a python -c, a heredoc) written again from scratch, grouped by what it does rather than its filename. These want a saved tool, not a skill: write it once, then call it.Nothing written. `npm run propose -- --write 1,3` drafts those rows into .kompact/proposals/.
Real output, captured from eval/propose.ts when this page was
built, over the same committed corpus as the table above. Not a mock-up.
npm run propose is the same report over what the recorder saw on your machine,
which is nothing until the plugin has been installed a while.
Watch 3,228 commands land on the shapes they repeat
The classifier works and the counts are real. What is not established is the premise underneath: these sessions are an assistant’s tool use, and what an assistant repeats is reading files, editing them, and reaching for the shell. What that cost to find out, and what this table is not, is on the evidence page.
06Cost
281 ms to score 3,591 tool calls.
Compaction stops being a model call and becomes arithmetic. These are whole session logs, so the totals span more than one context window, so the ratio is what transfers.
| Session | Calls | Tokens before | After | Freed | Scoring |
|---|---|---|---|---|---|
| 20628921 | 3,003 | 2,111,814 | 1,435,062 | 32.0% | 247 ms |
| 25e65eab | 473 | 273,749 | 224,484 | 18.0% | 26 ms |
| agent-ab4cc5 | 49 | 43,400 | 39,451 | 9.1% | 3 ms |
| agent-afa3fd | 66 | 83,022 | 75,768 | 8.7% | 5 ms |
| Total | 3,591 | 2,511,985 | 1,774,765 | 29.3% | 281 ms |
Modest on purpose. The absolute cut this replaced frees 95.3% of the labelled corpus and keeps 8.9% of the outputs that were reused later: the first row of the policy table below, and the reason a third is the honest number. Measured on one machine’s own transcripts, which grow as work continues. The per-session column is the figure that carries.
What it replaces, timed. Over 11 real auto‑compactions on this machine,
Claude Code’s own summariser took a median 133 s to produce
(65 s to 224 s), each cutting the transcript by about
95%. kompact answers the same compaction with a pass measured in
milliseconds. This is the summary’s cost to produce, read from Claude Code’s
own durationMs — which times the whole compaction request, not only the pause
on screen, on one machine. What the assistant loses by not getting that narrative is a different
question and stays not verified.
And what each keeps alive. The built-in summary cut the window to a median 35k tokens; kompact’s own passes here (3, median 183 ms) left 262k — it fires earlier and compresses less, so more of the transcript stays verbatim. It does not compact less often; each pass is milliseconds, and the slow summary is left only for the residue kompact cannot drop.
Get this table for your own machine first
Clone the repository and run npm run dry-run. The commands are in
section 09, Install. It reads
~/.claude/projects, scores every tool call in every transcript it finds, and
prints the table above for your sessions instead of these. It writes nothing, installs no
plugin, and needs no key: it is the same compact the hook calls, run against
your own history. If the number it prints is not worth having, you have lost a minute.
Expect somewhere in 8.7% to 32.0%, the range four sessions covered, which is not a distribution. What moved it was tool mix and length: the 32.0% session is 3,003 calls of mostly commands whose output is never quoted afterwards, and the two at the bottom are 49 and 66 calls long, where there is almost nothing to free. At four sessions that is a reading of these four and not a rule. Nothing here has measured how much of the spread is noise. The dry run tells you which of those you are.
Why it frees a third and not four fifths
The scorer produces a ranking, and a ranking transfers across sessions better than an absolute probability cut: 0.789 held out one session at a time against 0.879 on the narrower corpus, so it degrades rather than collapses. A cut does not travel at all, because every session has its own mix of tools and so its own distribution: a Bash-heavy session scores low throughout, and a fixed cut sweeps all of it. Barely 11% of tool outputs are ever reused verbatim (247 of 2,239), so a well-calibrated scorer seldom passes 0.5 even for the ones that matter.
Calls are therefore dropped lowest-score-first only until the reduction target is met, with the keep threshold acting as a floor that nothing is dropped above. Simulated per session on the labelled corpus, scored leave-one-session-out:
| Policy | Output freed | Reused outputs kept | Reused characters kept |
|---|---|---|---|
| Absolute cut at 0.5 | 95.3% | 8.9% | 16.9% |
| Budget 0.5, floor 0.2 (shipped) | 23.4% | 77.3% | 92.2% |
| Budget 0.6, floor 0.2 | 28.3% | 72.1% | 91.4% |
| Budget 0.7, floor 0.2 | 33.9% | 68.8% | 90.2% |
What the cap costs, priced. The 33.6% above includes it; so does this. Of the 247 outputs that were genuinely reused later, 5 are longer than the cap, and shortening them removes 76,466 characters those later steps had quoted, 6.2% of every reused character. Counting the wrong drops beside it, 88.1% of reused characters survive the shipped defaults and 11.9% do not. The gain and the loss are measured the same way, because for a while only the gain was.
Freeing nearly everything is easy and nearly worthless. Run through the shipped code rather than this simulation, the same defaults give 33.6% freed and 79.4% kept. The sweep has no notion of two shipped rules: a call recording a change is never dropped, and a kept result is capped at 24,000 characters. The cap is most of the gap: 27 of the 2,239 calls are longer than it, and shortening those 27 alone frees 20.7% of every character in the corpus. Calibrating on your own sessions is an accuracy upgrade rather than a prerequisite, because the policy adapts to your distribution on its own.
What the neural sidecar would have cost to run
The neural model is not in this product any more. There is no option that turns it on,
and nothing shipped imports it. What it cost to run is a large part of why. The
eval/ scripts still measure all of this against a live sidecar, so every row below
is reproducible rather than remembered. One router process loads every checkpoint and routes
per request, so memory and start-up belong to the router rather than to any one checkpoint.
| Scoring 3,591 calls, the four sessions above | Time | Memory held | VRAM | Cold start |
|---|---|---|---|---|
| Built-in scorer | 281 ms | none | none | none |
| Neural sidecar, fastest checkpoint, on a GPU | 4.9 s | 3.2 GB | 5.8 GB | 6.1 s |
17× the time, for a lower AUC, on an RTX 4060 Laptop. Per request the fastest checkpoint answers eight questions in 653 ms and the slowest in 1,889 ms. The VRAM is the router with all three checkpoints resident, which is what the benchmark starts; one checkpoint alone is about 1.4 GB. Fine-tuning moves the AUC; it does not move this table.
This row has been wrong twice, and the second time is the more useful story. It first read
23.1 s and 122×, from a projection that
divided requests by the number of questions in a request, but a request carries one call's
two questions, so there is one request per call and the projection had halved itself. It then
read 90.7 s and 319× for weeks, because
the benchmark behind it was measuring a sidecar started with
LAYA_DEVICE=cpu while the card sat idle. The script printed the idle card two
lines above its own table and nobody read it. On the GPU the same work takes
4.9 s, so the magnitude was wrong by a factor of eighteen.
eval/sidecar-bench.ts now reads the sidecar's own /health, which
reports its device in one field, and refuses to print publishable rows from a CPU one.
One number was never the cost of running this model. It was the cost of one deployment of
it. laya-serve routes /health and /v1/systemone and
nothing else, so over HTTP every call is one state in one forward pass in one round trip,
which is the slowest thing the model can do; the library packs many states into shared passes.
Measured on the same card, the wire costs about 0.3 ms of that
round trip and batching roughly halves what is left. The evaluation
page prints the whole ladder. Even at its fastest the sidecar loses this comparison, which is the point of publishing the ladder rather than the one number that flattered our
own conclusion.
07Where it loses
The model this replaced fails quietly, which is worse than failing.
The English checkpoint reads 512 tokens and discards the rest with HTTP 200 and no
warning. eval/truncation.ts asks one question (the text says the
deployment was rolled back) of four states. Every one of them states the fact
verbatim. Only the position of the sentence changes.
| State | Characters | Tokens read | Answer |
|---|---|---|---|
| The sentence alone | 85 | 47 | 0.8472 |
| First, plus filler | 8,140 | 512 | 0.9152 |
| First, plus 5× filler | 40,360 | 512 | 0.9152 |
| Last, after 5× filler | 40,360 | 512 | 0.2191 |
Rows two and three are bit-identical: 32,220 characters were discarded and nothing in the response says so. Row four is the bill for that. It is the same 40,360 characters and the same sentence, moved to the end where the cut removes it, and the answer falls from 0.9152 to 0.2191 for a fact the text still states verbatim: from the right side of any threshold to the wrong one. A caller cannot tell rows three and four apart: same HTTP 200, same shape, same silence. Whether an answer is about your input or about the first 512 tokens of it is settled by where the evidence happened to fall. Both upstream projects send the whole conversation, up to 25,000 tokens, as one state.
This table used to be three rows that no script produced. The one claim here with no code
behind it, on a page arguing that every figure is generated. It is
npm run eval:truncation now, and the figures it prints are these. The
ones it replaced are retracted: they showed the cut but not what the cut costs.
What the other 20.6% costs
“79.4% of reused outputs kept” invites you to supply your own answer for the rest, so here it is, priced. (The sweep above says 77.3% and this says 20.6%, which sum to 97.9 rather than 100. The gap is the two rules the sweep has no notion of: a call recording a change is never dropped, and neither is a result nothing can produce again. Run through the shipped code the same defaults give 33.6% freed and 79.4% kept, which is the pair the masthead quotes.) Held-out scores, shipped policy, per session: 51 of the 247 reused outputs are dropped: 1.24 a session, and only one of them kept a head.
| How it would be recovered | Outputs | Characters | Produced by |
|---|---|---|---|
| Re-run the command, side effects unknown | 47 | 42,389 | Bash |
| Read the file again | 4 | 53,680 | Read |
| Not repeatable at all | 0 | 0 | — |
Only one of the 51 kept a head, because when a result scores below the floor the call
usually does too. Forty-seven are command output that a later tool input quoted verbatim, the shape this scorer is worst at, since most command output is never referred to again and
the minority that is gets swept along with it. The bottom row used to read
4 AskUserQuestion results, a human’s answer,
which no amount of re-running brings back. It reads zero because
a result nothing can produce again is no longer the scorer's to drop.
This is recovery cost, and it over-counts: it includes reuses that had
already happened when a real compaction fired. The next section measures what is actually
at risk.
Compacting where the engine would actually fire, 28 of 32 sessions lost nothing at all. Across the rest, 5 outputs were dropped that a later step went back to: 0.16 a session, and 6.9% of what was genuinely still needed. Nothing here replays an assistant against the compacted transcript, so whether that changed the work stays not verified.
How that was counted, and what the looser count says
Every figure above asks whether an output was reused somewhere later, which counts
reuses that had already happened by the time a compaction fired and could never have been
lost. eval/outcome.ts asks it in the right order: find where the engine would
actually compact (62% of a session's tool output), decide only the calls present at that
moment, then count the reuses that come afterwards whose source was dropped.
| Compacting at 62%, across 32 sessions | Count |
|---|---|
| Calls present when it fires | 1,071 |
| Reused only after that point | 72 |
| …and dropped anyway | 5 |
| Sessions that lose nothing at all | 28 of 32 |
0.16 lost outputs a session, and 6.9% of what was genuinely still needed. This is the necessary condition for the work to suffer, not proof that it did: nothing here replays an assistant against the compacted transcript, and that remains unverified.
Five more things it does not do
A summary compresses harder. Freeing a third of the tokens is less than replacing the conversation with a few thousand tokens of prose. What you buy with the difference is not re-running tools to recover what a summary threw away.
One feature does most of the work. Output size alone scores 0.878 with an eighth of the spread (sd 0.010 against 0.078) and a sixtieth of the variance. The other coefficients earn their place on characters freed at a given safety level, not on ranking.
It has never seen a Grep. The corpus contains no
Grep and no Glob call at all, so that feature never fired during
fitting and its weight is 0.0: zero by absence, not by
measurement. A search falls to the reference level, alongside WebFetch and
Agent, about which this scorer also knows nothing. Thirteen features were
defined; twelve were fitted.
The shipped coefficients are one person’s sessions. A different tool mix
scores differently and nothing detects that for you, so npm run calibrate
refits on your own transcripts and reports whether it helped.
It never shortens what a tool was asked. Every lever here acts on
output. eval/mass.ts says output is 46.8% of a
transcript and tool input is 43.6%. A
Write carries the whole file it wrote and a Bash its heredoc,
and nothing shortens any of it. eval/inputs.ts priced capping it and the trade
is poor: capping every input at 16,000 characters frees
1.9% of the transcript for
1.3% of the passages later steps quoted back, against
20.7% for 6.2% on the output side.
Reuse inside an input sits at median depth 0.50, so a head cap
samples it rather than trims it, which is the same reason head-and-tail truncation lost.
The full table
includes the one tool that would be worth capping and why a default built on 26 calls from
one machine would be a default fitted to this operator.
08Verification
Seven claims checked, five not.
The minimum Claude Code version — not verified
Everything here was run on 2.1.281, which has function hooks. Which release first shipped them has not been established, so “2.1.281 or newer” would be a claim from a sample of one. On a build without them the plugin installs and never runs.
An assistant still finishing the job — not verifiedblocks a 1.0
The measurement above shows how often the information would no longer be there. Whether that changes what the assistant does needs the same session replayed with and without compaction, which needs a live A/B.
That a proposed skill is worth having — not verified
The repetition in section 05, What you repeat is measured, and
the drafts validate as skills. Whether writing one saves anyone anything is not measured at
all: it needs the same work done twice, once with the skill and once without. Thirteen of the 17
repeated command shapes are generic shell verbs and the other four are
git incantations, which is evidence against rather than for. The ad-hoc
scripts it found are a better case, and get their own table above.
What deferring the summary costs — not verified
Six passes run before the model summary does, and how much room each one buys is measured. What the assistant loses by not getting a narrative summary for six compactions is not: this keeps verbatim tool calls and prose, which is a different thing from a story about them. maxPasses exists so the summary happens eventually, and its value is a judgement, not a result.
The engine calling session.compact live — not verifiedblocks a 1.0
The plugin loads and turn.complete fires in a real session, but the engine declines to compact below ~31% full and the model cannot invoke /compact itself. Covered by the real-session test; the engine-side handoff is not.
And the 7 claims that are verified
Beats the model it replaces
1,063 labelled calls, 18 sessions, 10 grouped splits, paired. eval/repeat.ts
No orphaned tool call survives
Run over a real session from disk. An orphan is rejected by the API and would break the session compaction was meant to save.
Prose is never touched
Same real-session test: only tool calls and results are dropped or truncated.
A dead scorer cannot break a session
Falls back to the built-in summary; a single failed request keeps its call.
Works for Codex CLI, against a fixture
jev-compact’s own parser, scorer and HTTP client driven over its recorded rollout fixture against our server. Not run against a live Codex: that plugin needs Codex ≥ 0.155 and this machine has 0.131.
What a compaction costs the work
Compaction simulated where the engine would fire it, then the reuses that come afterwards counted: 5 of 72, 0.16 a session, 28 of 32 sessions losing nothing. eval/outcome.ts
The sidecar truncates silently
One question asked of four states that all state the fact verbatim: two differing by 32,220 characters answer bit-identically, and moving the sentence past the 512-token cut takes the answer from 0.9152 to 0.2191. eval/truncation.ts
What repeats, and why the premise under it is not established
That is not the finding this was built to get. The classifier itself works: the counts are real, the signatures group work that belongs together, and three separate design ideas turned out to be wrong only because the data said so: 94% of command shapes were seen exactly once until a pipeline stopped keeping every stage's arguments, sequences never formed at all until the window slid across batches instead of inside one, and the 2,000-row cap cut the rarest kinds to nothing until each kind was pruned against itself.
What is not established is the premise underneath it. These sessions are an assistant's tool use, and what an assistant repeats is reading files, editing them, and reaching for the shell.
Two things this table is not. It is not from a live install: the recorder only sees sessions
from the moment it is switched on, so these counts come from replaying existing transcripts
through the recorder's own functions (eval/signals-fixture.ts) rather than from the
hook having run. And it carries no examples. The signatures are shapes and safe to publish, but
the samples kept beside them are raw text, and the first fixture built without stripping them
passed all eight credential checks while carrying another project's brief and a path with a
session id in it. On your own machine the samples are the point, because they are how a wrong
grouping becomes visible. In anything published they have no place.
Nothing recorded leaves the machine, and nothing is written into a skills directory: drafts land
in .kompact/proposals/, which Claude Code does not read. Set
recordSignals to true to record.
The gaps, in the order they matter
Everything above this line is measured. Everything below it is not yet, and is written in the future tense on purpose.
Open
The corpus is still one person's. Widening it from 18 sessions to 41 (all the same person) took leave-one-session-out AUC from 0.879 to 0.789, and calibration error from 0.035 to 0.051. Nine points and half again the error, inside one person's own work. A second operator's corpus would say how much of what is left is this person's habits.
Needs a machine
A live Codex CLI. Verified against a recorded rollout fixture; that plugin needs Codex ≥ 0.155 and this machine has 0.131.
Wanted
A fine-tuned neural checkpoint that wins. The bar is 0.905, not
chance. eval/ re-runs the whole comparison for anyone who tries, and the
section 06, Cost is the other half of that decision.
Structural
More than one person's sessions. The corpus is 41 sessions from one machine. A committed scrubbed copy ships with the repository so a second corpus can be compared against it rather than replacing it.
The questions a sceptic asks
Isn't this just a heuristic with extra steps?
It is a heuristic with error bars, which is the difference. Thirteen coefficients (twelve features and an intercept) fitted on labelled data, evaluated on sessions held out whole, and calibrated well enough that the threshold means a probability rather than a rank; section 01, Evidence carries the error itself. And most of the signal really is one feature: “big outputs get reused” alone reaches 0.878. That is three coefficients of the twelve, since size is banded. The other nine earn their place on characters freed at a fixed safety level, not on ranking, which is the answer to the question.
Why did the neural model lose?
It was asked to read facts as prose that the features encode directly, and it was never
trained on this task. The model’s own documentation says its base checkpoints sit near chance on
typed-decision workflows and should be treated as a fast base to specialise. That is what
was measured, and the wording mattered more than the checkpoint: on
typed-decisions, asking what the assistant should do next scores
0.719, and asking what is true of the text scores
0.416, below the coin flip, which is worse
than not scoring at all. That is the reverse of
what that documentation suggests, which is the finding rather than an
embarrassment: the checkpoint is better at a decision than at an entailment.
The checkpoint named for English is also not the best at English. On that same
wording it scores 0.480 (below the coin flip)
where the multilingual one scores 0.667, and the two ranges
do not overlap: English tops out at 0.502, multilingual bottoms out at 0.638. English wins
only on entailment, by 0.021, which is inside one standard deviation, and its best config
still loses to multilingual's best. It also reads half the context, 512 tokens against
1024, and runs about three times slower. None of that is why the sidecar loses:
typed‑decisions reads 1024 too and reaches 0.719, but it does mean
there is no configuration in which the English checkpoint is the one to pick.
Will it delete something I need?
Sometimes, and section 07, Where it loses prices it exactly. The asymmetry is what the whole design turns on: a wrong keep costs context, a wrong drop costs work, so a dropped result keeps its first 300 characters, a call that recorded a change is never dropped, and every failure mode falls back to Claude Code's own summary rather than to a half-compacted transcript.
Do I have to calibrate?
No: spending a budget lowest-score-first adapts on its own, where an absolute
probability cut could not survive a distribution it was not tuned on. So
npm run calibrate is an upgrade rather than a prerequisite, and it is still
worth running on your own sessions, for the reason in the ledger:
these coefficients come from one person's.
How is this different from the built-in /compact?
/compact?The built-in asks a model to summarise the session and replaces it with prose, which compresses far harder than a third. This keeps what survives verbatim, so a file you read is still the file, not a description of it. It is not a better compressor; it is a different trade, and it hands the session back to the built-in summary whenever it cannot do its job.