Evaluation results

Generated by npm run eval:results on 2026-09-27. Do not edit by hand — edit the scripts and re-run, so the numbers and the code that produced them stay together.

Everything in this file is reproducible. Every script here runs against the committed fixture in eval/fixtures/, so a runner that has never seen a private transcript regenerates it byte for byte — which is what lets CI gate it, and what lets test/published-figures.test.ts treat it as ground truth.

Figures measured on one machine's own transcripts live in SNAPSHOT.md instead. They move whenever their owner works, so nothing is bound to them: the page marks them as a dated snapshot. Keeping the two in one file is what let twelve figures go stale at once.

Ranking quality — eval/repeat.ts

Ten splits, each holding out 30% of sessions (not calls, so a session's own calls never sit on both sides). A single split flatters whatever it measures: the first run of this scorer reported 0.918, which is inside the range below but nowhere near its centre.

corpus (fixture, 1063 of 2239 rows scored by every config), laya answers (fixture): 1063 calls, 75 positives, 18 sessions
10 grouped splits, 30% of sessions held out each time

scorer                             AUC mean     sd    min    max    ECE  ECE(T)  drop@90%
-----------------------------------------------------------------------------------------
logistic (features)                   0.905  0.078  0.684  0.944  0.044   0.044     35.4%
output size only                      0.878  0.010  0.855  0.890  0.069   0.123     11.9%
laya typed-decisions/direct           0.719  0.026  0.666  0.754  0.413   0.423     12.1%
laya multilingual/direct              0.667  0.023  0.638  0.713  0.588   0.435      4.5%
laya english/entailment               0.625  0.016  0.602  0.658  0.417   0.424     18.0%
laya multilingual/reproducible        0.610  0.016  0.594  0.632  0.634   0.441      2.5%
laya multilingual/entailment          0.604  0.015  0.578  0.622  0.642   0.442      3.1%
laya english/reproducible             0.539  0.016  0.521  0.566  0.380   0.414      4.1%
laya english/direct                   0.480  0.012  0.462  0.502  0.445   0.425     10.1%
laya typed-decisions/reproducible     0.419  0.025  0.393  0.470  0.382   0.277      8.2%
laya typed-decisions/entailment       0.416  0.023  0.379  0.472  0.428   0.431      5.2%

ECE(T) is ECE after temperature scaling fitted on each split's own training
sessions — the calibration step the model's integration guide asks for and this
project never ran. AUC and drop@90% are unchanged by it, and cannot change: the
transform is monotone and both columns are rank-based.

paired vs best laya (typed-decisions/direct):
  logistic - laya AUC: mean +0.186 (sd 0.062, range 0.018 to 0.227)
  logistic won 10/10 splits
  -> the free scorer wins on every split; the sidecar is not justified for this task.

features: 13 (tool=Read, tool=Bash, tool=Edit|Write, ...)

The ECE column used to run without its companion, and it was not a fair measurement. The model's own integration guide says to refit a temperature on your own labels before trusting its probabilities, and warns that the multilingual checkpoint ships uncalibrated at 1.0. This project never ran that step, and then published a calibration error against the model as though it were a property of the model. ECE(T) is the same column with the step run — one temperature per scorer, fitted on each split's own training sessions, and every scorer gets one, including this one.

It matters most where the guide said it would: the multilingual rows fall from about 0.6 to about 0.44. It changes nothing for the scorer that ships, whose 0.044 is already the product of a maximum-likelihood fit on the same sessions — a temperature on top finds ~1 and moves it not at all, which is the check that the arithmetic is right.

And it cannot touch the result this table is actually about. Temperature scaling is strictly monotone in the probability, while AUC and drop@90% are rank-based — droppableAt sweeps the score's own values as candidate thresholds — so both columns are identical before and after, by construction rather than by luck. Correcting the unfair column leaves the ranking argument exactly where it was.

Does the neural model know anything the coefficients do not? — eval/teacher.ts

The table above asks which scorer ranks better and answers: this one. That is the right question for which one ships and the wrong one for was the sidecar worth building — a model can lose outright and still carry signal the winner lacks, and signal like that is worth having even when the model is not, because it can be distilled into the coefficients offline and shipped as floats.

So: the same logistic, the same ten splits, fitted once on the 13 features and once on those features plus the neural model's two probabilities for the same call. The gate was written down before the run — a mean paired gain above +0.018, the closest margin the comparison above already tolerates, on at least 8 of 10 splits.

corpus (fixture, 1063 of 2239 rows scored by every config), laya answers (fixture): 1063 calls, 75 positives, 18 sessions
10 grouped splits, 30% of sessions held out each time
the 13 features alone: AUC 0.905 (sd 0.078)

13 features + laya                     AUC    gain      sd   wins  verdict
--------------------------------------------------------------------------
english/reproducible                 0.906  +0.000   0.001   6/10        -
english/entailment                   0.906  +0.000   0.001   6/10        -
english/direct                       0.905  -0.000   0.000   4/10        -
typed-decisions/reproducible         0.905  -0.000   0.001   2/10        -
typed-decisions/direct               0.905  -0.001   0.001   1/10        -
typed-decisions/entailment           0.905  -0.001   0.001   2/10        -
multilingual/direct                  0.905  -0.001   0.009   7/10        -
multilingual/entailment              0.904  -0.002   0.004   2/10        -
multilingual/reproducible            0.904  -0.002   0.006   5/10        -

Gate A: gain > +0.018 and at least 8 of 10 splits won.
  FAILED. Best is english/reproducible at +0.000 over 6/10 splits.
  -> Laya adds nothing the same state already gives the features. It is a slower
     way to compute what src/features.ts computes, and that is the finding.

The features are read back out of the same state prose the model is given (src/features.ts), deliberately, so that neither side sees anything the other does not. This is what that choice buys: the result is not "a small model lost to a big one", it is "an encoder reading this prose extracts nothing from it that thirteen regexes miss". A negative result about our own idea, and the reason the fine-tune behind it was not run.

What the shipped coefficients generalise to — eval/fit.ts

The figure above holds out 30% of the 18 sessions the neural model was scored against. These are the coefficients that actually ship, fitted on all 2239 calls and held out one session at a time. It is the lower number, and it is the one to plan around.

corpus (fixture): 2239 calls, 41 sessions

result_needed: LOSO AUC 0.789  ECE 0.051  positives 247/2239
call_needed  : LOSO AUC 0.909  ECE 0.047  positives 544/2239

export const KEEP_RESULT_WEIGHTS: readonly number[] = [
  -0.854504,   // bias
  -0.080158,   // tool=Read
  -0.103559,   // tool=Bash
  -1.961938,   // tool=Edit|Write
  0.000000,    // tool=Grep|Glob
  -0.203870,   // isError
  2.283794,    // targetTouchedAfter
  -0.585116,   // targetReadAgain
  -3.047582,   // size=very short
  -1.474709,   // size=short
  -0.259475,   // size=very long
  0.074074,    // age=long ago
  0.097573,    // age=just now
];

export const KEEP_CALL_WEIGHTS: readonly number[] = [
  -0.614181,   // bias
  -0.379293,   // tool=Read
  -0.405222,   // tool=Bash
  0.438643,    // tool=Edit|Write
  0.000000,    // tool=Grep|Glob
  -0.059718,   // isError
  3.600786,    // targetTouchedAfter
  4.605452,    // targetReadAgain
  -2.658946,   // size=very short
  -1.263787,   // size=short
  0.195068,    // size=very long
  0.062363,    // age=long ago
  0.059608,    // age=just now
];

Decision policy — eval/policy.ts

The scorer produces a ranking; this turns it into a decision. The shipped defaults are budget 0.5, floor 0.20.

fixture: 2239 calls, 247 genuinely needed, 41 sessions

policy                         freed   wrong  +head     kept chars kept
-----------------------------------------------------------------------
threshold only, 0.5 (old)      95.3%     225     34      8.9%      16.9%

budget 0.7, floor 0.50         70.5%     176     17     28.7%      44.1%
budget 0.7, floor 0.30         60.4%     148      8     40.1%      53.9%
budget 0.7, floor 0.20         33.9%      77      4     68.8%      90.2%
budget 0.7, floor 0.15         11.9%      65      4     73.7%      91.6%
budget 0.7, floor 0.10          9.6%      58      7     76.5%      96.1%
budget 0.7, floor 0.05          3.7%      31     28     87.4%      98.7%

budget 0.5, floor 0.20         23.4%      56      4     77.3%      92.2%
budget 0.6, floor 0.20         28.3%      69      4     72.1%      91.4%
budget 0.7, floor 0.20         33.9%      77      4     68.8%      90.2%
budget 0.8, floor 0.20         35.7%      89      4     64.0%      87.3%

shipped code path (src/compact.ts decideAll, the same defaults):
  33.6% freed, 79.4% of reused outputs kept, 0 of 220 mutating calls dropped
  88.1% of reused CHARACTERS kept (11.9% lost, almost all of it to the 24,000-char cap)
  without that guard 15 of them would lose their call or output, so the rule saves 15, not 220
  the 24,000-char cap alone: 27 of 2239 calls run past it, and shortening those 27 frees 20.7% of every character in the corpus
  it costs 5 of the 247 reused outputs 76,466 characters later steps had quoted, 6.2% of every reused character
  a result nothing can produce again is also never dropped: 17 of 23 AskUserQuestion/ExitPlanMode outputs scored below the floor, and they are 1.67% of the corpus

wrong  = needed outputs that were dropped anyway
+head  = of those, how many kept their first 300 characters
kept   = share of needed outputs not dropped at all
chars kept = share of needed CHARACTERS still present afterwards,
             counting the heads that survived a partial drop.

What a wrong drop costs — eval/recovery.ts

The policy above keeps 77.3% of the outputs that were reused later. This prices the rest. Read the script's own caveats first: it measures recovery cost, not task outcome, and it over-counts, because a reuse that had already happened by the time a real compaction fired costs nothing when the output is dropped now.

fixture: 2239 calls in 41 sessions, 247 reused later

dropped anyway: 51 (20.6% of the reused outputs)
per session:    1.24 outputs
of those, kept their first 300 characters: 1, lost entirely: 50

how the assistant would get it back:
  command, side effects unknown    47      42,389 chars
  re-read                           4      53,680 chars

which tool produced it:
  Bash                             47      42,389 chars
  Read                              4      53,680 chars

why it was needed:
  reused in a later tool input     41      85,774 chars
  quoted in later text             10      10,295 chars

Not recoverable by re-running, and no head left: 47 of 2239 calls (2.1%), 1.15 per session.
NOT measured here: whether an assistant given the compacted transcript still
finishes the task. That needs a live A/B and is listed as unverified.

Whether the work survives — eval/outcome.ts

The measurement every other figure here stands in for. Every other one asks whether an output was reused somewhere later, which counts reuses that had already happened when a real compaction fired and could not therefore be lost. This finds where the engine would actually compact, decides only the calls that exist at that point, and counts the reuses that come afterwards whose source was dropped. It is still not a replay: it measures how often the information would no longer be there, which is the necessary condition for the work to suffer.

fixture: 2239 calls, 32 sessions with a compaction point
compacting at 62% of a session's tool output

calls present when it fires:        1071
of those, reused only afterwards:   72
and dropped anyway:                 5 (6.9% of them)
                                    54,576 characters
per session:                        0.16
sessions that lose nothing:         28 of 32

This is the necessary condition for the work to suffer, not proof that it
did: nothing here replays an assistant against the compacted transcript.

What the loop costs — eval/outcome.ts --passes

The premise of a ladder is that passes 2..N take cheap context rather than compounding loss. Asked properly — the loop replayed over the labelled corpus, each pass charged only for outputs whose first reuse comes after the point it fired at — that premise does not hold.

fixture: 2239 calls, 32 sessions with a compaction point
compacting at 62% of a session's tool output

calls present when it fires:        1071
of those, reused only afterwards:   72
and dropped anyway:                 5 (6.9% of them)
                                    54,576 characters
per session:                        0.16
sessions that lose nothing:         28 of 32

This is the necessary condition for the work to suffer, not proof that it
did: nothing here replays an assistant against the compacted transcript.


the loop, up to 6 passes a session
 pass  calls live  at risk  lost       freed   lost per 10k freed
-----------------------------------------------------------------
    1        1071       72     5   1,034,605                0.048
    2        1022       65     6     157,502                0.381
    3        1105       59     4     124,535                0.321
    4         786       53     1      22,525                0.444
    5         791       50     1      15,778                0.634
    6         794       44     0      12,751                0.000

whole loop: 17 outputs lost, 1,367,696 characters freed
            0.53 a session, against 0.16 for one pass
            1.32x the characters, 3.40x the loss

The ladder's premise is that the rate column stays flat: a later pass takes
cheap context rather than compounding loss. It does not. The first pass is the
efficient one, because it compacts the whole backlog at once; every pass after
it works on fresh material only and costs several times as much per character.

The reason is structural, not a defect in the later passes: the first pass compacts the whole accumulated backlog at once, which is where the cheap bulk is, and every pass after it works on fresh material only. It is still not an argument for handing over after one pass — the alternative to pass two is not "keep everything", it is the engine's model summary, which keeps no tool output verbatim at all. What the table settles is that the extra passes are not free, and anyone who wants the cheap pass and nothing else can set maxPasses: 1.

Not verified: whether deferring the engine's summary costs the assistant anything. applyDecisions never touches prose, so what kompact leaves behind is verbatim tool calls and the user's and assistant's own words — not a narrative. maxPasses is the backstop for that, and its value is a judgement: the floor would allow more.

Whether prose can be compacted fast — eval/prose-narration.ts

kompact's passes are mechanical and cost milliseconds; the model summary that runs when it hands over costs a median of two minutes (eval/summary-cost.ts). The obvious question is whether the prose — the assistant's and user's own words, which applyDecisions never touches — could be compacted the same cheap way.

The safest possible prose drop needs no scorer: when kompact drops every tool call in a message, the text that introduced them ("Let me read X") is orphaned, and dropping it recovers nothing that a re-run cannot. Measured over 35 real sessions, its yield is zero:

fixture: 35 sessions, window tail ~700k tokens each
co-located orphaned narration:  median 0%  (max 0%) of window
standalone prose pool:          median 3.92%  (mean 4.53%, max 15%)
tool share of window:           median 93.5%
pooled orphan:                  0 / 3,546,015 tokens

The zero is structural, not a null measurement. In Claude Code's transcripts a tool call sits in its own message with no text of its own — narration lives in separate, text-only rows. So "clear the text on a message whose calls were all dropped" has nothing to clear: those messages are already textless (msgsWithCalls === calls, asstWithText === 0 on every session checked). The Phase-1 drop was specified against a shape the data does not have; it is not shipped.

The only real prose lever is the standalone pool — the text-only assistant rows, a median 3.92% of the window. Reclaiming it is a different problem from dropping a tool result: a result is droppable because it can be re-run and keeps its first 300 characters, and a sentence has neither property. Attributing a standalone message to a call that was dropped is a judgement, not a fact, so this is the extractive-scorer problem, not the mechanical one.

Not built — research. An extractive prose scorer (the same logistic machinery as eval/logistic.ts, keeping user instructions, decisions, file paths and final answers; dropping acknowledgements and restated context) would need a labelled prose corpus built the way the 1,063 labelled tool calls were, and its wrong drops are unrecoverable. At a 3.92% median ceiling against a 93.5% tool share, the tool path is where the tokens are; the prose scorer is filed, not funded.

Can prose be compacted without a model? — eval/prose-extractive.ts

The ship target for prose is model-free and verbatim-safe. This asks whether cheap model-free features can rank an assistant prose sentence by whether it is referenced again later — the same verbatim-reuse proxy the tool-call scorer is judged on — well enough to drop the rest. Never a user message; assistant prose only. Features: length, LexRank-lite centrality, recency, has-path/code, average self-information from corpus unigram frequencies (a no-LM stand-in for Selective Context), is-question.

35 sessions, 4274 assistant prose sentences
reuse base rate:           17%  (share referenced later)
AUC, model-free logistic:  0.594   (out-of-fold, by session)
AUC, length baseline:      0.586
AUC, self-information:     0.418  (Selective-Context-lite, no LM)
droppable @ 95% retained:  8.1% of prose chars (36 reused sentences lost)
featurise + score:         6.3 ms/window

The model-free ranking is at chance: 0.594 against a 0.586 length baseline, so the logistic learns essentially nothing beyond "longer sentences are kept a little more often", and self-information does worse than chance — rare wording is not what gets reused. At a safe 95% retention it frees 8.1% of prose characters, and prose is 3.92% of the window, so this is about 0.3% of the window.

The ceiling is low for any method, not just this one. A perfect classifier that dropped every non-reused sentence would free 80.8% of prose characters — still only ~3.17% of the window, because prose is a small slice of it. A learned compressor (LLMLingua-2, Selective Context) cannot exceed that bound, so it cannot change the decision: on this corpus, the tokens are in tool output, and compacting prose — model-free or not — is not worth adding. A direct benchmark confirms the bound: LLMLingua-2 and a GPT-2 self-information compressor both rank at chance on this label (AUC 0.51 and 0.49), below model-free's 0.594 and below length, at 335 ms and 41 ms per sentence on CPU with a multi-GB dependency — no accuracy gained, large cost added. kompact stays mechanical and hands the prose-only residual to the engine's summary.

Can we reliably propose skills for full flows? — eval/flow-proposals.ts

A "full flow" is a re-usable multi-step tool sequence (Read → Edit → test), plus the session-start orient reads and the pre-handback verify checks. The ledger marks whether a proposed skill is worth having as not verified; this measures the answerable part: of the flows kompact would surface, how many are reliable — a habit across ≥ 2 sessions, and actionable rather than pure inspection. Flows are mined exactly as the recorder mines them (hooks/kompact-signals.ts).

195 sessions, 4818 distinct flow shapes
recur across ≥ 2 sessions:   90  (1.9%)   — a habit, not one afternoon
reliable (≥ 2 sess + actionable): 43  (0.9%)
current ranking (by saved), top 20:  20% cross-session, 10% reliable

Two things are true at once. The recurrence signal is sparse: only 1.9% of flow shapes are seen in more than one session, so most are one-off and cannot be proposed as habits at all. And the current ranking — modelled saved, what propose.ts shows — is unreliable: 10% of its top 20 are reliable, the rest being single-session or inspection noise, exactly the failure the ledger warned of.

A gate on (≥ 2 sessions AND actionable) fixes the precision — it isolates the 43 flows that are habits and do something, e.g.:

  [sequence] Read → Bash(python3 -c) → Read                        2 sess, 27x
  [sequence] Read → Edit → Edit                                    3 sess, 23x
  [sequence] Bash(python3 -c) → Read → Read                        2 sess, 20x
  [sequence] Read → Read → Edit                                    2 sess, 18x
  [sequence] Bash(python3 -c) → Read → Bash(python3 -c)            2 sess, 17x
  [sequence] Bash(python3 -) → Bash(grep -n) → Bash(grep -n)       2 sess, 15x
  [sequence] Bash(python3 -c) → Bash(sed -n) → Bash(sed -n)        2 sess, 11x
  [sequence] Write → Write → Read                                  2 sess, 11x

So proposals can be made reliable, but only with the gate and only modestly. Scanning the full corpus (recursively — including the ~80 subagent transcripts an earlier 2-level walk missed — and keying by parent session so one session's many subagents do not fake recurrence) lifts the reliable set from 25 to 43, across 90 cross-session flows: more data does raise the count. The character does not change, though. The reliable flows are generic edit / read / debug loops (Read → Edit → Edit, Read → python3 → Read), not distinctive procedures, and the saved-ranking still surfaces only 10% reliable in its top 20. More sessions buy more generic flows, not more skill-worthy ones.

Still not verified (per the ledger): that encoding any of these as a skill saves time — recurrence is measured, payoff is not. The shippable part is the gate; the honest next tests are a second operator's corpus (whether the same flows recur for someone else) and a SkillOpt-style held-out payoff check, neither of which one machine's history can answer.

A better flow-discovery method than n-grams? — eval/flow-discovery.ts

Study 4's near-zero recurrence might have been the fixed 3-gram's fault, so two model-free upgrades were tried: PrefixSpan (frequent gapped, variable-length subsequences, so Edit … test … commit survives interleaved noise) and intent-anchored flows (the tool run following each recurring intentSignature, keyed by goal rather than tool syntax).

distinct parent sessions: 12  (substantial, ≥ 10 calls: 4; 195 transcripts, mostly subagents)
PrefixSpan actionable patterns:         1369
  deduped to maximal (fair vs baseline):1280  (vs 25 for the 3-gram)
recurring intents (≥ 2 sessions):        2
  with a stable actionable flow:         0
top patterns by support:
  4x  AskUserQuestion → AskUserQuestion → AskUserQuestion
  3x  Bash(echo ;) → AskUserQuestion → AskUserQuestion → AskUserQuestion
  3x  AskUserQuestion → Write → Bash(ls -la) → Write
  3x  AskUserQuestion → Write → Bash(ls -la) → ToolSearch
  3x  AskUserQuestion → Write → Bash(ls -la) → Skill

The method is not the binding constraint — the data is. The 195 transcripts group into only 12 parent sessions (most are subagents of a few), and just 4 carry more than ten tool calls, dominated by one project and by the very sessions that built these studies. And the count comparison is itself unreliable: PrefixSpan returns 1369 actionable patterns, but deduping to maximal only trims that to 1280 — the bulk is combinatorial branching over a few sessions (AskUserQuestion → Write → Bash(ls -la) → {Write, Edit, Skill}), this session's own tooling, not an engineer's organic flows. Intent-anchoring finds 0 stable goal-flows. No heavier miner (PAM, Local Process Models, or an LLM auto-skill inducer) can conjure cross-operator regularity that 4 same-context sessions do not contain. The honest next step for skill proposals is a second operator's corpus, not a cleverer algorithm — and then a SkillOpt-style held-out check for whether a proposed skill actually saves time, which no amount of mining answers.

Do proposed flows recur on held-out sessions? — eval/skill-payoff.ts

SkillOpt keeps a skill only if it improves held-out performance. A live A/B is not possible here (nothing can replay a past session with a skill injected), so this tests the prerequisite a real payoff needs: split the sessions by time, propose the reliable flows from the earlier half, and see whether they recur in the later, unseen half.

114 parent sessions split by time (propose 57 → held-out 57)
sessions carrying flows:       propose 2, held-out 10  (activity is time-skewed to recent days)
gated proposals (≥ 2 sess + actionable): 8 → recurred in held-out: 3   (38% predictive)
ungated (any actionable flow):           745 → 5% predictive
held-out flow occurrences covered:       10 / 6,574  (0%)

Two honest signals and one honest limit. The gate helps: gated proposals recur on held-out data far more than ungated ones (38% vs 5%), so "recurs across ≥ 2 sessions and does something" is directionally the right filter. But the propose half carries only 2 sessions with flows — the corpus is time-skewed, with almost all substantial work in the last three days — so 38% over 8 proposals is not trustworthy, and the proposals cover about 0% of held-out activity.

The held-out test is sound; the data is too thin to run it. And even fully powered it would measure only predictive validity — the floor under any payoff. Realized time saved needs a live A/B and is bounded small: a skill does not stop you running the tools. That A/B and a second operator's corpus are the real next steps; no retrospective mining substitutes for either.

The prompt-cache cost of compaction — eval/cache-cost.ts

ContextPipe (arXiv:2609.00749) notes that editing history mid-prefix busts the prompt cache from the edit point on: the next turn re-creates cache (1.25x the base input rate) for content it had been reading at 0.1x. kompact drops stale tool outputs mid-history, so it pays this. Measured from the transcripts' own usage fields:

steady-state cache_creation per turn: ~990 tok
                    n   re-cache (empirical / modeled)   freed (median)
built-in (auto):   11   58k / 35k                     698k
kompact (manual):   3   0k / 262k                   395k

The cost is real, and larger for kompact than for the built-in: kompact re-caches ~262k tokens per pass (the kept prefix it must re-create once) against the built-in's measured ~58k, precisely because kompact keeps roughly seven times more context alive. But it amortizes. A pass pays back its cache bust in ~7.6 turns (re-cache 262k x 1.15 once against 395k freed x 0.1 saved every later turn), far below the spacing between compactions, so a kompact pass is net-positive on cache in the ordinary case.

Two honest limits keep this from being a shipping decision. Only 3 kompact passes exist on this machine, so the empirical spike for kompact is unreliable (the built-in's n=11, ~58k, is the trustworthy anchor); and this counts cache tokens only, not the primary benefit kompact exists for (staying under the window cap and avoiding the built-in's two-minute summary). A cache-aware gate would help only a compaction with very few turns left in the session. Measured, real, but not proven worth a gate — recorded, not shipped.

Could a local archive recover the dropped outputs that matter? — eval/archive-recall.ts

kompact drops a stale output betting it can be re-run; eval/recovery.ts showed that bet fails for outputs nothing reproduces. The fitting alternative (local, no model, no network): park the verbatim output in a BM25-indexed archive and re-inject on demand. This tests the retrieval mechanism — of the 1098 outputs needed verbatim later, does a BM25 query built from the reusing message return the correct original from the whole archive?

reused outputs: 1098 across 69 sessions;  archive 8,721 docs
BM25 recall@1: 23.2%    recall@5: 52.5%

The mechanism is not good enough on its own: over an archive of thousands of similar outputs (many near-duplicate reads and diffs), a naive keyword query returns the right original only 52.5% of the time in the top five. The idea is not dead — kompact knows each output's tool and target, so scoping the archive by target before BM25 should lift this sharply — but plain BM25 does not by itself beat re-running. And this is a ceiling: the query overlaps the output verbatim only because the output was present when reused; whether the model would issue such a query with the content absent needs a live A/B.

How far back is an output when it is reused? — eval/reuse-distance.ts

If reuse were mostly near-range, kompact's preserve-recent window would already protect what matters and dropping old outputs would be safe. It is not. Over 1098 verbatim reuses:

distance production -> first reuse: median 19 messages
within preserve-recent (<= 6):     37.6%
long-range (> 50 messages):        36.7%
recency predicts reuse:            AUC 0.407  (0.5 = no signal)

36.7% of reuses are long-range — an output produced fifty or more messages ago, reused verbatim. Dropping old outputs is exactly where kompact's risk sits, which is why the archive above matters and why the re-run guarantee is doing real work. Recency is a weak-to-negative predictor of reuse (AUC 0.407, partly a censoring artifact: late outputs have little session left in which to be reused), so "keep it because it is recent" is not the signal it feels like — the scorer's other features carry the weight.

What settles the payoff — the live A/B (specified, not run)

Every study here ends at the same wall: recurrence, retrieval recall, cache cost, reuse distance are all measurable, but whether kompact (or an archive, or a skill) changes what the assistant actually accomplishes is not, retrospectively. The test that would settle it: replay a set of real tasks twice — once with kompact, once without — through the model, and compare task success, tokens, and wall-clock. It is not run here, for two reasons this machine cannot fix: it needs a second operator's corpus (this one has four substantial sessions, one project, three days) so the result is not one person's habits, and it needs live session replay through the model, which is expensive and not reproducible in CI. This is the standing credibility ceiling; naming it precisely is the honest outcome. The harness belongs beside eval/outcome.ts, which already measures the necessary condition (whether the dropped information was later needed) without the sufficient one (whether its absence changed the work).

How far is the scorer from the offline optimum? — eval/offline-optimal.ts

kompact's keep/drop is an eviction policy; the reuse labels give perfect foresight, so we can compute the best-possible decision (keep an output iff it is needed later) and measure the gap. Over 2,239 labelled calls (11% needed later):

freed at 100% needed-retention (never drop a needed output):
  offline optimum (oracle):         75.1%   <- ceiling: drop exactly the not-needed
  kompact policy (force-keeps on):  0%
  raw logistic ranking:             0.2%
  size-only baseline:               25.7%
  (10.9% of calls force-kept: mutating / unrepeatable)

Read it carefully. freed@100% is a stringent, outlier-dominated metric — one needed output ranked below everything blocks all safe freeing, which is what drives the ~0% here. It does NOT mean kompact frees nothing in practice: at the shipped floor (0.2) it frees ~23% while accepting ~23% needed-loss (see the decision-policy section), leaning on re-run and the local archive to recover the rest. What the oracle (75.1%) against the size baseline (25.7%) shows is that most characters sit in not-needed outputs, so a policy that never dropped a needed one could free most of them — the scorer's tail just ranks some needed outputs too low to get there.

The honest catch: a better scorer was already tried and lost. The neural sidecars (Laya, and Study 2's LLMLingua-2 / Selective-Context) all scored worse than this logistic on the same task. So the headroom is real but unclaimed, and the lever is better features, not a heavier model — the next section tests exactly that. But no feature change can be trusted off this machine until it is refit and re-scored on a second operator's corpus. That corpus is the one thing every study here waits on: eval/calibrate.ts --contribute writes an aggregates-only file (counts and AUC/ECE, no session content, no weights) that a second operator can share to grow the evidence base — the scorer-corpus path, distinct from recordSignals, which counts repeated command shapes for skill proposals. A trustworthy result still needs more, and more diverse, real sessions than one machine holds.

Can better features close that gap? — eval/feature-search.ts

Study 11 said the lever is features, not a heavier model. The shipped scorer buckets output size into three levels because a model cannot read digits — but a logistic can, and three buckets throw away a long right tail. So this adds one continuous log-size feature and measures leave-one-session-out AUC and calibration (ECE) for both heads, over 2,239 calls / 41 sessions, against the shipped thirteen (continuous columns standardised with train-fold statistics only, so a fold never sees its test rows):

| feature set | result_needed AUC (Δ) | ECE | call_needed AUC (Δ) | ECE | | --- | --- | --- | --- | --- | | shipped-13 | 0.789 (—) | 0.051 | 0.909 (—) | 0.047 | | size-only (log chars) | 0.812 (+0.023) | 0.025 | 0.588 (-0.321) | 0.076 | | shipped + log chars | 0.807 (+0.018) | 0.033 | 0.916 (+0.007) | 0.027 | | shipped + log chars + sq | 0.807 (+0.018) | 0.033 | 0.917 (+0.008) | 0.023 | | shipped + log chars + tools | 0.808 (+0.019) | 0.032 | 0.916 (+0.007) | 0.030 |

Reading it: one continuous log-size feature helps both heads — result_needed 0.789→0.807 and call_needed 0.909→0.916, and it roughly halves calibration error (ECE 0.051→0.033 and 0.047→0.027). A squared term and extra tool indicators add nothing beyond that. The two heads disagree about size, which is the telling part: size alone already beats the full shipped set at deciding whether an OUTPUT is needed (0.812 vs 0.789), but for whether the CALL still matters it collapses to 0.588 — that decision rides on whether the target was touched or read again, not on how big the output was.

Modest, and free at scoring time — so it was built and wired all the way in to ship. It was then reverted, because the operating metric moved the wrong way. Refitting the 14-feature scorer and replaying the keep/drop policy (decideAll) over the corpus, at the shipped 0.2 threshold and the retention it holds (≈85% of needed kept), the log-size scorer freed only 8.3% of characters against the shipped scorer's 20.7% (the real decideAll freed, Study 15) — freed cut by more than half for no gain in retention. The shipped 13-feature scorer dominates its freed-vs-retention curve everywhere useful.

The reason is the mismatch AUC hides: AUC weights every call equally, but freed% is weighted by size, and log(chars) earns a positive weight (bigger output → likelier needed), so it protects large outputs wholesale — and large outputs are where the characters are. A per-call ranking gain (+0.02 AUC) became a character-weighted loss. This is the whole case for judging the scorer on Study 11's freed-at-retention curve rather than on AUC: the feature is a clean win on the metric that does not decide anything and a clear regression on the one that does. Not shipped.

Cost-aware eviction: drop big uncertain outputs first — eval/cost-aware.ts

Study 12's trap was that freed% is size-weighted while the score is per call. The size-aware decision (the CDN "beyond-Belady byte-miss-ratio" line, and cost-aware replacement like GDSF) is the principled response: keep small outputs cheaply, spend drops on the big uncertain ones. Tested realistically — the shipped scorer's out-of-fold P(needed), two keep-policies swept over the threshold, both read on the true labels (realized retention, no oracle): keep iff P ≥ t, versus keep iff P ≥ t OR the output is at or below the corpus-median 363 characters.

| realized retention | freed (P ≥ t) | freed (+ keep small) | Δ | | --- | --- | --- | --- | | 95% | 0.9% | 0.0% | -0.9 | | 90% | 0.9% | 2.0% | +1.1 | | 85% | 3.2% | 4.5% | +1.3 | | 80% | 4.4% | 10.1% | +5.7 |

At kompact's operating retention (85–90%) cost-aware is +1.3 pts — within noise; it only pulls ahead at 80%, below where kompact runs. No gain at the operating point. An earlier cut of this study used an oracle retention budget and reported +10.8 pts — but that was the oracle choosing which needed outputs to sacrifice, not the scorer; corrected here to realized retention. Speed: scoring all 2,239 calls takes ~1 ms (≈0.5 µs/call), three orders of magnitude inside the 500 ms budget.

Does imitating the Belady oracle's target help? — eval/reuse-target.ts

The imitation-learning line for cache replacement (Liu et al., ICML 2020) and reuse prediction (Faldu, 2020) train on the future reuse pattern, not a bare needed/not bit. kompact trains on binary result_needed (reused ever). This retrains on near-term targets — "reused within N messages" — and scores each on freed at the true result_needed retention:

| trained-on target | freed@90% | freed@85% | | --- | --- | --- | | any reuse (shipped) | 1.9% | 4.0% | | reused <= 10 msgs | 2.8% | 4.7% | | reused <= 25 msgs | 2.1% | 4.9% | | reused <= 50 msgs | 2.7% | 4.8% |

No near-term target beats the binary one by more than ~1 pt. The reuse-distance refinement that matters for CPU and CDN caches — where a line is evicted and re-fetched repeatedly over time — has no purchase on a one-shot drop at compaction: here "needed at all after this point" already is the Belady target. No gain.

Is the shipped keepThreshold on the frontier? — eval/operating-point.ts

Study 12 made the operating point a first-class lever, so before touching the model: sweep keepThreshold through the real policy (decideAll, with its force-keeps and drop_call/truncation) and read needed-retention and freed characters at each.

| keepThreshold | kept/needed | freed | | --- | --- | --- | | 0.10 | 86.6% | 6.6% | | 0.15 | 86.6% | 13.0% | | 0.20 (shipped) | 85.4% | 20.7% | | 0.25 | 83.0% | 26.0% | | 0.30 | 61.9% | 50.0% | | 0.35 | 61.5% | 51.9% | | 0.40 | 61.5% | 51.9% | | 0.50 | 61.5% | 51.9% |

No swept threshold beats the shipped 0.20 on both axes — 0.25 frees more but retains less, and 0.30 collapses retention to 62%. The default is Pareto-optimal on this corpus, and its 20.7% freed at 85.4% retention is the true decideAll figure the Study 12 regression is read against.

Taken together, Studies 13–15 test the three most-cited leads from the cache-eviction and reuse-prediction literature — cost-aware ordering, an imitation-learning target, and threshold re-tuning — and none beats the shipped design on freed-at-retention. With Study 12, that is four ranking-level ideas that looked promising and did not survive the operating metric. It is the same conclusion Study 11 reached from the other side: the scorer is near its ceiling for this corpus, and the remaining lever is a second operator's data, not a cleverer method or a heavier model.

Does the scorer transfer to a different harness? — eval/codex-transfer.ts

Every study above runs on one operator's Claude Code sessions, and Studies 11–15 keep concluding the lever is a second corpus, not method. The closest genuinely-different distribution on hand is Codex CLI (~/.codex/sessions) — a different harness, different tools (exec_command, apply_patch), the client jev-compact targets. Its outputs are mapped onto kompact's tool categories, reuse-labelled with the same 8-word-shingle method, and scored with the shipped weights, no refit.

Codex: 108 calls, 7 sessions, 28 needed (25.9% — higher than Claude Code's 11%)
shipped weights, no refit:  result_needed AUC 0.469   (in-distribution: 0.789)
                            freed@85% retention 1.3%   (Claude Code decideAll: 20.7%)

The scorer does not transfer. On Codex it scores at chance (0.469), and freed collapses to 1.3%. The corpus is labellable — outputs are reused more often than in Claude Code — so this is the scorer failing to generalise, not a data problem. Read it with its limits: 28 positives is a small sample (AUC 95% CI ≈ ±0.12, so "chance", not proven anti-correlation), and the Codex→kompact adapter is approximate. But the direction is unambiguous and it is the first cross-distribution evidence here: it is the concrete case for local calibration — eval/calibrate.ts refits on an operator's own sessions, and --contribute shares the aggregate so the "does it transfer" question can be answered with more than one machine. This aggregate is committed; the raw Codex corpus is not (it is another tool's private transcripts) — run eval/codex-transfer.ts on your own ~/.codex to reproduce it.

And the recovery makes the case concrete: refit the same features on Codex itself, out-of-fold by Codex session, and the scorer comes all the way back — AUC 0.904 against 0.469 for the shipped weights, freed@85% 67.5% against 1.3%. So the failure above is not that these features are wrong for Codex; it is that the weights are Claude Code's. Local calibration — which eval/calibrate.ts already does, and --contribute already shares — fully closes the gap. (Both Codex figures rest on 28 positives across 7 sessions, so the exact numbers are noisy; the 0.44-point swing is not.) That is the whole architecture in one experiment: ship a reasonable default, and refit locally where the distribution differs.

Snapshot — one machine's own transcripts

Generated by npm run eval:results on 2026-09-26.

Nothing in this file is reproducible, and no test is bound to it. Every script here reads whatever ~/.claude/projects happens to hold, and those transcripts grow every time their owner works — so these figures move for reasons that have nothing to do with any change to the code. The reproducible ones are in RESULTS.md.

Quote these as a dated snapshot or not at all. They are the range a reader should expect on their own machine, not a constant.

What it frees in practice — eval/sessions.ts

Not reproducible off this machine: it reads whatever ~/.claude/projects holds, and those transcripts grow as you work. Treat the total as a snapshot of one corpus, and the per-session column as the range that matters.

session              calls  tok before  tok after    freed  scoring
-------------------------------------------------------------------
20628921-046d-4fc0    3003   2,111,814  1,435,062    32.0%    247ms
25e65eab-4941-41ea     473     273,749    224,484    18.0%     26ms
agent-ab4cc58e7542      49      43,400     39,451     9.1%      3ms
agent-afa3fde141a2      66      83,022     75,768     8.7%      5ms
-------------------------------------------------------------------
total                 3591   2,511,985  1,774,765    29.3%    281ms

How many passes before the summary is due — eval/passes.ts

The other scripts measure one compaction. Production is a loop: the engine asks at compactAtPercent, kompact answers, the session keeps growing, the engine asks again. This simulates that loop on real transcripts — each pass gets the compacted prefix plus the real continuation, not its own output, because feeding a pass its own output measures exhaustion and reports a decay that production never sees.

Same caveat as sessions.ts: this machine's transcripts, which grow as you work, so it is a snapshot rather than a constant.

window 200,000 tokens, compacting at 60%, floor 5 pp (10,000 tokens), ceiling 6 passes

session              msgs  passes         pp reclaimed per pass           ms per pass
-------------------------------------------------------------------------------------
20628921             8663       6  8.7 9.9 8.4 8.3 7.2 6.6 5.8†       14 10 9 9 8 8 8
25e65eab             1455       2                 12.2 7.2 3.4*              18 14 12
78d7d176             1140       1                      5.7 3.4*                   8 6
agent-a2              294       0                          5.0*                    11
agent-ad              201       0                          5.0*                     7

* = refused by the floor, † = refused by the ceiling. Neither is taken, but only
the first says the loop had run out of things worth freeing.

9 engine summaries avoided across 3 sessions (3.0 per session that loops at all).
slowest single pass 18 ms.
  pass 1: median 5.7 pp over 5 sessions
  pass 2: median 7.2 pp over 3 sessions
  pass 3: median 8.4 pp over 2 sessions
  pass 4: median 8.3 pp over 1 sessions
  pass 5: median 7.2 pp over 1 sessions
  pass 6: median 6.6 pp over 1 sessions
  pass 7: median 5.8 pp over 1 sessions

of the tool results that survive, how many are now a truncation note:
  after pass 1: 2.9% (worst 7.6%)
  after pass 2: 6.3% (worst 8.3%)
  after pass 3: 8.3% (worst 8.3%)
  after pass 4: 7.3% (worst 7.3%)
  after pass 5: 7.1% (worst 7.1%)
  after pass 6: 6.7% (worst 6.7%)
  after pass 7: 6.8% (worst 6.8%)
  Nothing acts on this yet. It is the quality signal a single
  compaction does not have, and it is here to be watched.

what each bar takes, over the 14 passes measured:
  minReductionRatio  0.25  takes   0 of 14
  minReductionRatio  0.15  takes   2 of 14
  minReductionRatio  0.10  takes   8 of 14
  minReductionRatio  0.08  takes  11 of 14
  minReductionRatio  0.05  takes  14 of 14
  minFreedPercent      3  takes  14 of 14
  minFreedPercent      4  takes  12 of 14
  minFreedPercent      5  takes  10 of 14
  minFreedPercent      7  takes   7 of 14
  minFreedPercent     10  takes   1 of 14

ratio seen per pass: 0.14 0.16 0.14 0.14 0.12 0.11 0.10 0.20 0.12 0.06 0.09 0.06 0.08 0.08

wrote /home/oa/workspace/tools/laya-compact/eval/fixtures/passes.json: 7 rows, 6 taken
`npm run docs` renders it into the page via eval/passes-page.ts.

minReductionRatio: 0.25 — the bar that shipped before minFreedPercent — takes 0 of the 14 passes measured of them. It asks whether a pass was a large fraction of the transcript, and the passes above are 0.06 to 0.20 of theirs, so it was never cleared: kompact handed every one of these compactions to the model summary while it could still free 5.7 points of window in under 18 ms.

The unit is the fix rather than the value. Percentage points of the context window are what runs out, are comparable between a large session and a small one, and are the same unit as compactAtPercent — which makes the yield rule a hysteresis band for free: a taken pass leaves the fill at least minFreedPercent below the trigger, so the session has to grow back through it before another compaction can be requested.

The stub column is the one quality signal a repeated loop has and a single compaction does not. It rises for two passes and then stops, because fresh un-truncated output arrives between passes at about the rate the loop creates stubs — which is why no maxStubShare dial exists: there is nothing for it to catch.

What the cap costs, and one idea that did not work — eval/cap.ts

The cap (maxKeptChars) shortens outputs the ranking kept. It is a separate lever from the floor and it composes with it. Retention here is the share of reused shingles still present — the only measure both levers share, since dropping loses whole outputs and capping loses the far end of one.

The graded rows are a dead end, recorded rather than hidden: the cap cannot tell a 40,000-character output the scorer was confident about from one it merely did not drop, and the ranking has that number already, so the obvious move is to spend the tight cap only on the lukewarm outputs. Every grading frees more than the flat cap that matches it and costs more in the tail. keepResult is the probability an output is needed at all; it says nothing about where inside the output the reuse sits, and among outputs that survived the floor it has spent its information.

corpus: 81 sessions, 6685 calls, 12,604,050 characters
reused shingles to protect: 18,197
floor 0.2, dropped outputs keep their first 300 characters

policy                        freed    reuse kept
-------------------------------------------------
keep everything                0.0%         100.0%
ranking only (shipped)        22.0%          80.8%
cap 32,000 only                7.6%          99.9%
cap 24,000 only               10.0%          99.8%
cap 16,000 only               15.4%          99.1%
cap 8,000 only                29.7%          96.3%
cap 4,000 only                45.9%          86.1%
ranking + cap 32,000          23.4%          80.8%
ranking + cap 24,000          24.5%          80.7%
ranking + cap 16,000          28.6%          80.1%
ranking + cap 8,000           41.4%          77.2%
ranking + cap 4,000           56.7%          67.1%
ranking + graded 0.6/32k/12k    32.8%          79.3%
ranking + graded 0.6/24k/8k    40.8%          77.8%
ranking + graded 0.8/48k/8k    41.1%          77.6%

freed = share of all output characters removed.
reuse kept = share of shingles later text quoted that are still present.

per session, over the 53 sessions with 10+ reused shingles
ruinous = sessions keeping under half of what was quoted from them
policy                      freed med  kept med kept p10 kept worst  ruinous
----------------------------------------------------------------------------
ranking only (shipped)           7.2%     96.5%     66.7%       37.5%        2
cap 8,000 only                  16.1%     94.3%     61.7%        0.0%        3
ranking + cap 8,000             26.0%     83.7%     53.3%        0.0%        5
cap 16,000 only                  1.4%    100.0%     81.8%       37.5%        1
cap 32,000 only                  0.0%    100.0%    100.0%       43.8%        1
ranking + cap 32,000             7.3%     96.3%     66.7%       37.5%        2
ranking + cap 24,000             8.7%     96.3%     66.7%       37.5%        2
ranking + cap 16,000            14.3%     93.2%     66.7%       37.5%        3
ranking + graded 0.6/32k/12k      16.8%     91.0%     56.3%        0.0%        4
ranking + graded 0.6/24k/8k      24.9%     85.1%     53.3%        0.0%        5
ranking + graded 0.8/48k/8k      26.0%     84.7%     53.3%        0.0%        5

The half nothing touches — eval/mass.ts, eval/inputs.ts

Every lever in this repository acts on tool output. The first script asks whether that is where the characters are; the second prices capping the other half the way cap.ts prices the output cap.

An input cap is not shipped. Against the output cap — 20.7% of the corpus for 6.2% of reused characters — it is a far worse exchange rate at every setting, and the knee arrives early. The reason is the one that killed head-and-tail truncation: reuse inside an input is spread through it rather than gathered at the front, so a head cap samples it and loses reuse in proportion to what it frees.

The per-tool table names the one exception, and it is recorded rather than shipped for two reasons worth stating. It is a couple of dozen inputs on one machine, which is not a sample. And the tool it names is not one a stock Claude Code install has, so a default built on it would be a default fitted to this operator — the thing every other number here is arranged to avoid.

24 sessions, 12,875,280 characters

prose (never touched)         9.6%     1,241,042
tool input                   43.6%     5,614,810
tool output                  46.8%     6,019,428

tool                calls        input       output  in/call   share
--------------------------------------------------------------------
Bash                 4090    3,568,078    5,184,707      872   68.0%
Write                 116      806,837       32,056    6,955    6.5%
Read                  657       67,360      465,696      103    4.1%
SubagentHandback       20      439,487        1,220   21,974    3.4%
Agent                  70      289,864       79,348    4,141    2.9%
Edit                  214      242,197       42,941    1,132    2.2%
ExitPlanMode           10      135,611      134,853   13,561    2.1%
AskUserQuestion        24       46,729       13,345    1,947    0.5%
mcp__lightpanda__       2          130       41,840       65    0.3%
WebFetch                7        2,820        9,203      403    0.1%
Skill                  10        7,243          407      724    0.1%
TaskStop                9          207        4,799       23    0.0%
ToolSearch             30        2,420        1,460       81    0.0%
SendMessage             2        3,493          339    1,747    0.0%
69 sessions, 6,651 tool inputs, 6,265,370 characters
reused shingles to protect: 120,308

policy                                freed     reuse kept
----------------------------------------------------------
cap every input at 16,000              4.1%          98.6%
cap every input at 8,000              11.0%          89.3%
cap every input at 4,000              21.8%          70.5%
cap every input at 2,000              37.5%          48.8%
cap every input at 1,000              55.0%          29.2%

cap non-mutating inputs at 8,000       7.4%          98.6%
cap non-mutating inputs at 4,000      14.2%          93.2%
cap non-mutating inputs at 2,000      25.8%          82.7%

cap mutating inputs at 8,000           3.6%          90.7%
cap mutating inputs at 4,000           7.7%          77.4%
cap mutating inputs at 2,000          11.7%          66.1%

the same rows, as a share of the whole transcript:
policy                                freed     reuse kept
----------------------------------------------------------
cap every input at 32,000              0.1%          99.8%
cap every input at 16,000              1.8%          98.6%
cap every input at 12,000              2.7%          96.9%
cap every input at 8,000               4.7%          89.3%
cap every input at 6,000               6.4%          81.8%

how deep into an input the reuse falls:
tool                   reuses  median depth  in first 10%
---------------------------------------------------------
Write                   53915         0.498         10.1%
Bash                    53475         0.471          7.4%
Edit                     8444         0.614          4.7%
Agent                    3269         0.484         11.3%
Skill                     815         0.502          6.4%
ExitPlanMode              293         0.477          2.7%
AskUserQuestion            48         0.423         16.7%
SendMessage                20         0.676          0.0%

where the reuse is, by tool:
tool                  inputs       chars   reused  per input
------------------------------------------------------------
Bash                    5112   3,871,499    53475      10.46
Write                    131     944,079    53915     411.56
SubagentHandback          26     542,790       19       0.73
Edit                     258     318,503     8444      32.73
Agent                     70     289,864     3269      46.70
ExitPlanMode              10     135,611      293      29.30
Read                     911      89,747        0       0.00
AskUserQuestion           24      46,729       48       2.00
Skill                     11       7,370      815      74.09
WebFetch                  16       6,405        8       0.50

Where the sidecar fails quietly — eval/truncation.ts

Needs a live laya-serve, so this section is empty on a machine without one. The claim it backs is the page's, and until this script existed the rows behind it were prose nobody could re-run.

sidecar: http://127.0.0.1:8000/v1/systemone
checkpoint: english
question: "The text says the deployment was rolled back"
stated verbatim in every state: "The deployment was rolled back at 14:02 because the migration locked the users table."

state                    characters  tokens read   answer
---------------------------------------------------------
the sentence alone               85           47   0.8472
first, plus filler            8,140          512   0.9152
first, plus 5× filler        40,360          512   0.9152
last, after 5× filler        40,360          512   0.2191

The cut: 8,140 characters and 40,360 characters both read as 512 tokens and both answer 0.9152, bit-identical. 32,220 characters were discarded, with HTTP 200 and no warning.
The cost: move that same sentence to the end of that same filler and the answer falls from 0.9152 to 0.2191 — for a fact the text still states verbatim. Nothing in the response distinguishes the two.

Both upstream projects send the whole conversation, up to 25,000 tokens, as a single state.

And what it costs, which is less than it sounds

The corpus was built at one state budget, 700 tokens, and handed to every checkpoint. The English checkpoint's own budget is 320, so the server had been cutting 700 of 2,239 English rows — nearly a third — at HTTP 200 with no warning, and three published AUCs carried no footnote saying so.

The obvious conclusion was that those three numbers were unfair to the checkpoint. They are not. Giving it a state it can read whole makes it worse, at every wording, monotonically:

rows 2239  positives 247  corpus private  states as built (700 tokens)  sidecar http://127.0.0.1:8003/v1/systemone

checkpoint      phrasing         AUC    ECE  trunc    thr   drop%  wrongDrops
-----------------------------------------------------------------------------
english         reproducible   0.568  0.371    700  0.379     0.5           4
english         direct         0.525  0.411    681  0.463     2.2           4
english         entailment     0.578  0.394    659  0.419     1.4           4

the same rows at three state budgets, english:
state budget          phrasing         AUC   cut
------------------------------------------------
as built (700)        reproducible   0.568   700
as built (700)        direct         0.525   681
as built (700)        entailment     0.578   659
trimmed to 450        reproducible   0.528   169
trimmed to 450        direct         0.540   142
trimmed to 450        entailment     0.570   110
its own budget (320)  reproducible   0.406     2
its own budget (320)  direct         0.501     1
its own budget (320)  entailment     0.531     0

More state wins at every wording, even when a third of it is being cut.
The cut falls on the output excerpt, which buildCallState already puts last.

best by AUC: english/entailment (AUC 0.578, drops 1.4% of output chars at 98% safety)
AUC 0.5 = coin flip. Below ~0.65 zero-shot, fine-tuning is the only route to production.

buildCallState front-loads on purpose — task first, derived facts second, raw output excerpt last — so the server's cut lands on the part that was already the most expendable, while an honest trim to a smaller budget removes that part and then some. The silent cut was the problem, not the damage; the fix is the trunc column travelling as far as the AUC does, which it now does.

It also shows STATE_BUDGET.english is pessimistic. It reserves 192 tokens for the option head, and two noul questions with short criteria are nothing like that: at a 450-token state only about 150 rows of 2,239 are cut at all.

What the sidecar costs to run — eval/sidecar-bench.ts

Needs a live laya-serve, so this section is empty on a machine without one. One process serves all three checkpoints, so memory and start-up are properties of the router; only latency is per checkpoint.

This section was wrong once, and how it was wrong is worth keeping. Every figure in it was measured against a sidecar started with LAYA_DEVICE=cpu, while the card sat idle — and the script printed that idle card two lines above its own table, as gpu: ... 148 MiB. A ratio of "319x the time" reached the landing page from it. The script now reads /health, which reports the device in one field, and refuses to produce publishable rows from a CPU sidecar unless --allow-cpu labels them.

The ladder at the end is the other half of the same lesson. laya-serve routes /health and /v1/systemone and nothing else, so over HTTP every call is one state, one forward pass, one round trip — the slowest thing Laya can do. The library's predict_batch packs states into shared passes. A single number was never the cost of running Laya; it was the cost of this deployment of it.

sidecar: http://127.0.0.1:8003/v1/systemone
device:  cuda (loaded: english, typed-decisions, multilingual)
memory:  3170 MB resident (pid 3577571)
gpu:     NVIDIA GeForce RTX 4060 Laptop GPU, 5779 MiB, 8188 MiB

All three checkpoints are loaded by one process, so the memory above is the
whole router. Latency is per checkpoint, 5 runs, median and worst:

checkpoint           questions    median    worst  per question  input tok
--------------------------------------------------------------------------
english                      1     19 ms    31 ms       19.1 ms        104
english                      2     24 ms    25 ms       12.1 ms        208
english                      4     31 ms    31 ms        7.7 ms        416
english                      8     46 ms    47 ms        5.8 ms        832
multilingual                 1      9 ms    10 ms        9.3 ms         99
multilingual                 2     11 ms    11 ms        5.4 ms        198
multilingual                 4     14 ms    15 ms        3.6 ms        396
multilingual                 8     21 ms    21 ms        2.6 ms        792
typed-decisions              1     18 ms    19 ms       18.1 ms        104
typed-decisions              2     23 ms    24 ms       11.6 ms        208
typed-decisions              4     31 ms    31 ms        7.7 ms        416
typed-decisions              8     46 ms    46 ms        5.8 ms        832

input tok is usage.input_tokens, which is the per-question row count times the
number of questions — not the size of the state.

Scoring 3591 calls, the sessions above: 2 questions a request, 8 in flight.
  fastest checkpoint: 4.9 s and 3170 MB resident held for the session
  built-in scorer:    0.28 s and no process at all
  ratio:              17x the time
(3591 calls and 281 ms come from the eval/sessions.ts run in
 this same report, so the two sides are the same work.)

the same work off the wire, 2 questions a state, multilingual on cuda:
path                                           per call  vs slowest
-------------------------------------------------------------------
over HTTP on cuda, 3 resident                   10.9 ms        1.0x
over HTTP on cuda, 1 resident                   10.2 ms        1.1x
in process on cuda, one call at a time           9.4 ms        1.2x
in process on cuda, predict_batch(32)            4.5 ms        2.4x
(model resident in 4.4 s from a warm HF cache)

cold start:   6.1 s to first answer
  memory:     3915 MB resident (pid 3584247)
  gpu after:  NVIDIA GeForce RTX 4060 Laptop GPU, 7594 MiB, 8188 MiB
  gpu before: NVIDIA GeForce RTX 4060 Laptop GPU, 7424 MiB, 8188 MiB