A Measured Headroom That Nothing Could Reach

pycomplete is a code completer of mine that runs on a laptop CPU. It predicts the next token by mixing three things: an n-gram count model over your repository, a cache of the buffer you are currently editing, and a small transformer. One number, α, decides how much of the mixture is transformer, and it is a constant. The shipped value is 0.45, and a comment in the source says it is a constant on purpose.

The obvious idea is to stop it being a constant. Some positions in a file surely want more transformer and some want less; pick α per position and win. Three rounds later α is still 0.45, nothing shipped, and the reason is not the one I expected.

Everything below is scored in top-1: the fraction of positions where the token you were about to type is the single highest-ranked suggestion. And headroom means the gap between what a perfect chooser would score and what the best single fixed choice scores, which is the most any router could ever win.

A per-position oracle, allowed to pick the best α for every position after seeing the answer, scores +0.0705 top-1 over the best fixed α, with a paired 95% confidence interval of [+0.0623, +0.0786]. Three families of features captured 0.0%, −3.1% and −0.9% of that.

The completer lives in a private research repo, so there is no code to link. Every number below is traced to a results file or a re-run command.

The ceiling, and why it is not an artifact

6,000 positions sampled at a fixed stride through 42 held-out files, against a 116-file index of the Python standard library. The split is by file, fixed before any code existed, because positions inside one file share a buffer and an identifier vocabulary, and splitting by position leaks the answer across the boundary in exactly the direction being tested.

test top-1
shipped, α = 0.45 0.4948
best fixed α chosen on the training half (also 0.45) 0.4948
per-position oracle 0.5653

An oracle sweeping a grid can only look better as the grid gets denser; +0.0705 could therefore have been an artifact of trying seven values. On a 2,500-position subset, doubling the grid to thirteen values moves the oracle by +0.0012, which is 1.7% of the headroom. The headroom is structure, not resolution.

One piece of care matters more than it sounds. The mixture is replicated inside the measuring harness so that every α is scored on identical arm outputs and the transformer’s forward pass is paid once instead of seven times. That replication is a second implementation of the thing being measured, which is how harnesses quietly drift from tools; so the probe re-scores a sample of positions through the shipped completer and refuses to write its output if any disagree. The results file exists because 400 of 400 agreed.

Family one and two: nothing, and less than nothing

Every policy here is tabular. Sort each position into a cell, work out the best α for each cell on the training half, and use it. With no state transition and no discounting that is exactly a greedy contextual-bandit rule, which is the most legible thing to fail.

policy keyed on test gain over best fixed paired 95% CI
P-BUFFER how long the buffer is 0.4948 +0.0000 [−0.0033, +0.0033]
P-CONF the three arms’ confidence and agreement 0.4927 −0.0022 [−0.0065, +0.0022]
oracle the answer 0.5653 +0.0705 [+0.0623, +0.0786]

Buffer length capturing exactly nothing reproduces what the tool’s own source comments already claimed, on a different corpus from the one that settled it; confidence does slightly worse than doing nothing.

A flat result has two very different causes, and the round separated them in advance: a policy that overfits (big training gain, no test gain) is a different animal from one whose features cannot express anything at all. Permute the hit vectors across positions, which destroys any relation to the features while keeping every margin identical, and refit:

train-fit gain from pure noise : mean +0.0054   95th pct +0.0095
P-CONF's real train-fit gain   :      +0.0039   -> p = 0.740

Fitting noise beats fitting the features. With eight cells and seven values of α, picking the best-scoring α per cell is upward-biased by construction, and the confidence policy does not even reach that bias. So the finding is these features carry no routing signal, which is a narrower claim than routing does not work, and it is the only one that was measured.

Family three: the optimum really does move, and it is still worth nothing

The first round closed by naming the untried family: syntactic context, which is a different measurement rather than a better learner. The next round built it, importing the mixture, the α grid, the verification contract, the splitting rule and the placebo so that the features would be the substantive change.

The corpus is not held fixed, and that matters for reading the two rounds together. This one runs 12,000 positions over 72 held-out files rather than 6,000 over 42, and the best fixed α moves from 0.45 to 0.6 with it. So the tables below are internally consistent and are not comparable cell by cell with the tables above. The ship bar is the same in both: a router ships only if it clears +0.020 with a confidence interval excluding zero.

I predicted the payoff would come after a ., because the next token is an attribute and the repository index holds the project’s attribute names, so α should fall there. Measured, . wants α = 0.6, which is this round’s global best, with a lift of exactly 0.0000. The classes that move are , and =, which want less transformer at the start of a fresh expression. The direction of the mechanism I argued for was wrong, and I am leaving that on the page.

The gate held, though, in the sense of the pre-registered criterion: best α per previous-token class spans 0.30 across nine classes, where buffer length was flat and confidence was below the noise floor. Syntax genuinely moves the optimum, and it is the first family in three to do so.

Then the question the criterion does not ask: does moving pay? For each class, how much does its own best α beat the global best α on its own rows?

class n best α lift over the global best
, 261 0.3 +0.0421
: 202 0.45 +0.0198
= 204 0.3 +0.0098
id (identifier) 2137 0.45 +0.0066
kw (keyword) 616 0.45 +0.0016
( ) . op 1639 0.6 +0.0000

Size-weighted, the ceiling is +0.0063. Against this round’s own test-half oracle of +0.0662 that is 9.5%, and it is the most a per-class router could win, before anyone asks it to generalise. Four of nine classes want the global optimum exactly, and the two classes that actually pull away from it, , and =, hold 465 of 5,325 training rows between them.

The tool’s own source said why before any of this existed: the values from 0.45 to 0.6 are a plateau, and two of them differ on eighteen positions. Picking the best α on a plateau moves for free and is worth nothing.

Fitted, it behaves exactly as that predicts:

policy test gain paired 95% CI cells
best fixed α 0.4900 n/a n/a n/a
P-SYN, previous-token class by depth 0.4894 −0.0006 [−0.0049, +0.0036] 22, smallest 22 rows
P-SYN+, plus the token before that 0.4800 −0.0100 [−0.0153, −0.0048] 182, smallest 1 row

Finer cells bought variance, not reach; and the placebo gives the round in one line:

train-fit gain from pure noise : mean +0.0051   95th pct +0.0075
P-SYN's real train-fit gain    :      +0.0073   -> p = 0.065
the per-class ceiling          :      +0.0063

Syntax carries a little real signal, sitting above the noise mean where the confidence features sat below it. But the ceiling (+0.0063) and the noise floor (+0.0051) are the same size. The entire signal available to a per-class router is no larger than the bias from fitting one. That is a statement about per-class routing on this tool, not about every router anyone could build.

The same shape, one switch over

There is a second thing in this tool applied unconditionally. A suffix layer re-ranks candidates by how the token being typed ends, which helps when you are halfway through a word, and it is weighted the same everywhere. A per-position oracle over whether to apply it is worth +0.0235, CI [+0.0208, +0.0262], which clears the +0.020 bar. A perfect switch would be worth shipping, which is why it was worth trying.

The best fitted switch, given free choice of threshold and direction over the 5,325 training positions, returns the identity. It vetoes nothing.

The reason is arithmetic. The layer changes the top-ranked candidate on 2,670 positions: it breaks 282 of them and fixes 1,015. So a veto that fires blindly would trade away three and a half fixes for every break it prevents, and one that fires selectively has to find breaks at better than one per fix to come out ahead. Since breaks are only 0.28 of fixes to begin with, a useful veto needs to be about 3.6 times as selective as chance. The best ratio reachable by any threshold on any of the four available features is 0.48 breaks per fix, which is 1.7 times the base rate. Not enough, by a factor of two, so the optimiser correctly declines to fire at all.

My stated mechanism was wrong there too, in every feature. I had reasoned that the layer overrides a confident transformer, so a veto should fire on high confidence. Having the right answer ranked first and being confident are different quantities, and the breaks sit on the low-confidence side for all four features.

What this leaves

The ceiling is the contribution. +0.0705 with a confidence interval, stable under grid density, is a number this project did not have, and it bounds every future attempt. The next person does not have to find out whether there is anything there, only whether their features can reach it.

44.6% of positions are wrong at every α, which is what caps the oracle in the first round. Routing cannot touch those. Only a better arm can, and that is a different project.

What is untried is features, not method. The tabular policy was never the limitation: it beat nothing because its inputs carried nothing, and when its inputs finally carried something, the objective was flat.

The online version is now known to be unpromising rather than unmeasured. Measured offline over logged positions, every arm can be scored at every α, so the answer is fully available; a deployed system sees only whether the completion it showed was accepted. With full information these features found nothing, and less information cannot help.

What generalises

Measure the ceiling before you build the router. It costs one pass over logged data, it is the same measurement whatever policy you would have built, and it tells you the largest number any of them could return. Mine came back large, which is why I kept going, and three feature families later the interesting discovery is that a large ceiling and a flat objective are perfectly compatible.

A best choice that moves is not a best choice that pays. I wrote a criterion asking whether the optimum shifts across classes, it held at a span of 0.30, and I nearly reported that as the good news. The lift table is the one that matters and it was one extra column of the same data.

Print the bias from fitting noise next to every fitted gain. Two of the three families here produced a positive training-set gain. In both cases permuting the labels produced a larger one. That comparison is cheap, it fits in three lines of output, and without it both rounds would have had a number that looked like a small win.