Words Fitted to a Rule's Own Mistakes Do Not Travel
appgen is a tool of mine that turns a sentence into a running application with no language model in the generating path. Something has to read the sentence, and the thing that does it is a hand-written keyword rule that picks out the entity: the noun the generated app calls a row, which becomes its table name and its URL path. Given “Build a recipe manager with search” the rule has to answer recipe rather than manager, search or build.
It does that largely by elimination. Two closed word classes, a stop list and a determiner list, 130 words between them, mark the candidates that can never be the answer, and the rule takes the best of what is left. A language model reads this field better than my rule does, which raises the obvious cheap question. Can you just add the right words?
You can, in sample, and by a margin that is not luck. The same words are worth nothing anywhere else.
The measurement runs over two corpora of project titles with hand-written gold answers and no rows in common. One is 87 rows drawn from three third-party lists (build-your-own-x, project-based-learning and Project-Ideas-And-Resources); the other is 47 rows drawn from two more (karan/Projects and florinpop17/app-ideas). The lists are other people’s; only the gold answers are mine. The rule scores 57 of 87 on the first and 36 of 47 on the second. Those two scores are never added together, because the corpora are different registers and a pooled rate would hide exactly the effect this post is about.
There is no repo link here; the code is private research. Each figure is attributed to the results file or command that produced it.
Deletion is already at its best
Before adding anything, take things away: remove one word at a time from the 130 and rescore.
| words whose removal changes any row | 6 of 130 |
| of those, removals that gain a row | 0 |
| rows that deletion changes on the 87-row corpus | none at all |
So there is no row to be recovered by taking a word out of either class. Both lists are already at a local optimum in the deletion direction, and whatever headroom exists is in the other one.
Addition is real, and it is not selection noise
The candidates are drawn from the rule’s own wrong picks, which is the natural place to look and also the reason this needs a control. A word selected from the errors removes a wrong pick by construction; the question is whether it removes more of them than an arbitrary word from the same pool would.
Each corpus is used twice: in sample, meaning scored on the same rows the words were chosen from, and held out, meaning scored on the other corpus, which had no part in choosing them.
| selected on | words | candidates available | in sample | held out |
|---|---|---|---|---|
| the 87-row corpus | 7: business, checker, flask, golang, library, react, scheduler |
27 | 57 → 65 (+8) | on the 47-row corpus: +0 |
| the 47-row corpus | 4: browse, give, online, review |
11 | 36 → 40 (+4) | on the 87-row corpus: +0 |
The control is 200 seeded draws of the same number of words from the same pool, made without looking at the rows. That distribution is what “not luck” has to beat.
| selected on | fitted | random mean | random 95th percentile | random max |
|---|---|---|---|---|
| the 87-row corpus | +8 | +1.98 | +4 | +5 |
| the 47-row corpus | +4 | +1.07 | +3 | +3 |
Both fits exceed the maximum of 200 random draws. I had expected a random draw to land within two rows of the fit, and it missed by six. The selection is finding something real: these are not arbitrary words, they are words that carry weight on the corpus they were picked from.
The zero that means something, and the zero that does not
A held-out result of +0 is only informative if a different set of words could have scored otherwise. So the same 200 random draws were read on the held-out corpus too. The two directions turn out not to be the same measurement, and only one of them carries the finding.
| direction | fitted, held out | random draws, held out | draws that moved it |
|---|---|---|---|
| chosen on 87, tested on 47 | +0 | mean +0.00, min 0, max 0 | 0 of 200 |
| chosen on 47, tested on 87 | +0 | mean +0.34, max +1 | 67 of 200 positive, 0 negative |
In the first direction nothing at all can move the held-out corpus. No random draw does, so that zero is uninterpretable and I am not leaning on it. In the second direction the channel is demonstrably live: about a third of random draws help the held-out corpus by a row, none hurt it, and the four fitted words help it none.
That is the result. The fit beats chance in sample and comes in below the random mean out of it, and the prediction it refutes is my own: I had written down that the held-out figure would stay at zero for every draw, and a third of them moved it.
Why: the words belong to their corpus, not to the task
Look at what got selected. react, flask and golang are framework names. business, library, scheduler and checker are the vocabulary of a tutorial-title list. browse, give, online and review are imperatives, which is how the other corpus phrases things. Every one of them is a genuine wrong pick in its own corpus and dead weight in the other.
So the hypothesis space is not the constraint here, and it is not empty either. It is register-bound: tied to the house style of the particular collection of people who wrote those titles. A word list fitted to how one group writes project names is a word list about that group.
What this does not say
It does not say the rule cannot be improved. It tests exactly one hypothesis space, membership of two closed word classes. The part of the rule that decides whether a title is a noun phrase or a verb phrase, and routes accordingly, is untouched.
It does not say these words do not belong in a stop list. react and flask are framework names and a person may put them there on grounds this measurement cannot see. What is measured is that selecting them from the errors buys nothing that travels.
And it is not pooled. Two corpora, two registers, reported side by side.
The labels do not exist for the obvious next step
The obvious next step, once the hypothesis space is ruled out, is more labels. I had written down a curve at 75, then 150, then 300 rows. Those labels do not exist. The two gold sets are 87 rows and 47 rows with no overlap, so the entire hand-labelled budget is 134, against 861 titles extracted from the larger corpus alone. A 300-row curve needs about 170 new labels written by hand, and adding the two sets together to reach 134 is the pooling the whole design exists to avoid.
What is runnable instead is a curve inside the 87 rows, at 22 then 44 then 87, with the other corpus’s 47 held out as the transfer check. That is a smaller experiment than the one I wrote down, and saying so is cheaper than running the small one and describing it as the big one.
What generalises
A held-out zero is not evidence until you show the held-out corpus can move. I had two of them here, they looked identical in the results file, and one is a finding while the other is a dead channel that could not have said anything else. The only thing separating them is reading the control in the held-out direction, which costs one more loop and which I would not have run if the two zeros had not been sitting next to each other.
The other half is the more ordinary lesson and the more expensive one. Selection from your own errors beats chance in sample by construction, and a margin that clears the maximum of two hundred random draws feels like the opposite of overfitting. It was still overfitting. The in-sample control tells you the selection is finding structure; only the out-of-sample one tells you what the structure is about.