Reading a Reinforcement Learning Taxonomy Against My Own Tools

Somebody handed me a list of reinforcement learning algorithms and asked whether any of them could help my tools. It is a ChatGPT session printed to PDF: roughly ninety algorithm names across fifteen sections, with no descriptions, no citations and no results. I want to be precise about that, not to dismiss it. It is a menu of mechanism names, it was a good thing to be handed, and it is the reason I went and looked at one of my own tools properly. Where a claim below is about how an algorithm behaves, the authority is the third-party documentation I read alongside it, mostly OpenAI’s Spinning Up and a couple of tutorial pages, and never the list.

The tools are mine and none of them is a policy, which is what reinforcement learning optimises: a rule for choosing an action, tuned by whether past actions turned out well. appgen turns a sentence into a running application with no language model in the generating path. pycomplete is a code completer mixing an n-gram model, a buffer cache and a small transformer. growone is an online learner that adds hidden units as it goes.

Four rounds later, nothing shipped from any of it. The four reasons are different, and that is the interesting part.

The tools are in a private research repo, so none of them is linkable. The numbers are, in the sense that each one is attributed to where it came from.

Most of the list cannot apply, and that is the first result

Sections two through nine and eleven cover every value-based, policy-gradient, actor-critic, distributional, model-based, offline, multi-agent, hierarchical and large-scale family on the list. All of them need the same input: a number saying how well that went. Policy-gradient methods need to know whether an action beat the average; Q-learning needs a reward plus its own estimate of what the next state is worth. Without that number there is nothing to descend.

appgen has not got one, and the reason is not that nobody has built it yet. The obvious candidate is the tool’s own verification verdict, ALL CHECKS PASSED, and it is not a reward, because of what it means: the program matches its own published contract. A correct implementation of a misread request passes it. An earlier round drove 138 externally-written requests through the tool and got 34 wrong builds; all 34 printed ALL CHECKS PASSED. Optimising that signal trains the tool toward wrong programs that verify.

That is a structural fact about the tool rather than a gap in the literature, and it removes nine of the list’s fifteen sections in one paragraph. The remaining sections were not all assessed against that requirement, so this is nine of fifteen and not “all but one”.

The one branch that does apply

Section ten is exploration bonuses: pseudo-counts and Random Network Distillation. Both are reward-free. In reinforcement learning they manufacture a reward out of novelty, to push an agent toward states it has not seen; here you can keep the novelty estimator and throw the use away. The offline family motivates the same thing from the other side, which is that off the data’s support your estimates are untrustworthy, so stay on it.

Both reduce to one mechanism: estimate whether a request is inside the distribution the rules were built for. What the tool would do with that is refuse earlier and for a better reason. It already refuses most requests it cannot handle, by failing to find a grammar rule that matches; a novelty score would let it decline a request that parses fine and is simply unlike anything the rules were written for, which is the case that currently produces a confidently wrong app.

The population is 52 requests the tool actually built. 18 of them are apt, meaning the program it produced is the one the request asked for, judged against the request rather than against the tool’s own contract; 34 are not. The in-distribution reference is not curated: it is the repository’s own corpus generator, so the generator is the definition of a request the tool was built for, and no judgment call picks the fit set.

The scoring is AUC: take one apt request and one inapt one at random, and AUC is the probability the score ranks them the right way round. 0.5 is a coin flip and 1.0 is perfect. A permutation test at this sample size puts the floor for a resolvable result at 0.641, written down before any arm ran.

C1_length        AUC 0.717   p=0.0035
C2_coverage      AUC 0.739   p=0.0020
A1_pseudocount   AUC 0.784   p=0.0003
A2_rnd           AUC 0.676   p=0.0194

All four clear the floor, and the two that are not the imported machinery are the reason nothing shipped.

The control that ate the result

C1_length is token count. I put it in as the confound control, and wrote down in advance that it clearing the floor would matter more than the arms clearing it. It clears at 0.717: short requests are apt about 47% of the time, long ones about 11%.

So the arms have to survive length, and the round strips it two ways, both specified before either was run. Residual regresses the arm on length and scores what is left, which assumes the relationship is linear. Stratified takes the AUC separately inside each half of a median split, which assumes nothing but works on small groups.

arm raw residual stratified
C2_coverage 0.739 0.682 (p=0.016) 0.720 survives both
A1_pseudocount 0.784 0.567 (p=0.22) 0.716 the two controls disagree
A2_rnd 0.676 0.655 (p=0.033) 0.644 survives both, barely

The best raw arm is the only one the controls disagree about. Its 0.784 collapses to a null once the linear part of length is removed, and holds at 0.716 under the assumption-free read. I do not know which is right, and reporting the 0.716 alone would be choosing the favourable control after seeing both.

Then the part that decided it. The corpus was scraped from two websites, and they are not interchangeable. One contributed 40 requests at a 45% apt rate. The other contributed 12, all of them tutorial titles about reimplementing a technology (Chess Engine In C, mal - Make a Lisp, Write your own Operating System), and none of the twelve is apt. So knowing which website a request came from scores AUC 0.676 on this population, identical to the Random Network Distillation arm to three decimals.

Hold the corpus fixed and read inside the larger source only, with the floor recomputed for that smaller sample, and the RND arm drops to 0.644 and stops clearing. It was reading provenance. The pseudo-count arm survives that read; it is the one that does not survive the length control.

What actually decided it, which was not any of that

The confound is a reason to distrust an arm. It is not what stopped the tool shipping, and I want to be careful about that, because deciding afterwards is exactly how a round like this goes wrong.

The pre-registration said: a tool ships only if the separation holds and the operating point is usable. At a threshold keeping 90% or more of the apt builds, the best arm rejects 26% of the inapt ones, against the 40% named in advance. So the second condition failed and nothing ships, and that was settled before the arms ran.

A round that ran four arms, found p = 0.0003 on one of them, and then asked what counts as success would have shipped a gate on a 26% rejection rate.

The second tool has a return, and a plateau

pycomplete does have a usable number, because you can ask whether its top suggestion was right, so the machinery has the signal it needs. That turned into its own round with its own answer, and a third reason for shipping nothing.

One framing point is worth carrying here. Offline, over logged positions, every option can be scored at every setting, because the answer is recorded: the feedback is full, which makes the offline problem ordinary supervised learning and not reinforcement learning at all. The bandit framing only bites in deployment, where you see whether the completion you showed was accepted and nothing about the ones you did not. The offline version was measured first because if the offline ceiling is empty the online one cannot be better.

The third tool maximises a proxy, which is the one real hit

growone has a decision, a candidate set and a signal. When it adds a hidden unit it trains four candidates from random starting points and installs the one whose activation has the largest absolute covariance with what the model is currently getting wrong.

That is a proxy. The thing wanted is lower error going forward; the thing maximised is correlation with the current error right now. This is the exact split the literature names between directly optimising what you want and optimising a surrogate, and the list is what sent me to look.

Measured, the proxy is directionally right and already takes most of what is there. Perfect selection among its four candidates is worth 2.6% of loss, and the shipped rule captures 71.8% of that. No change ships, and that is the correct outcome.

The finding worth keeping is that the alignment is stratified. The measure is regret: how much of the available gap between the best and worst candidate the rule gives away, where 0 is perfect selection and 0.5 is choosing at random. On a smooth nonlinear stream regret is 0.228. On the open-key stream, where feature names arrive mid-run, it is 0.453, which is nearly blind. That stream is the one the tool is named for and is its distinguishing claim.

Two readings of that make opposite predictions. Either a unit installed for a feature that is still arriving looks worthless over a short window and pays later, so regret falls as the window lengthens; or the score simply does not rank candidates on that stream, so regret stays. Running each candidate set forward 600 points once and cutting the same run at three lengths, so the windows share the same installs:

window regret on the open-key stream
150 points 0.440
300 points 0.401
600 points 0.442

Paired change from four times the window: +0.002, CI [−0.052, +0.057]. The horizon explanation is rejected. The proxy is near-blind there rather than short-sighted.

(That 0.440 and the 0.453 above are the same measurement on different populations. The window study can only use installs with at least 600 points of stream left after them, which is 149 of the 224 the earlier round scored on that stream, so the two numbers should not be read as a change.)

And there is almost nothing to be blind about: the spread between the best and worst candidate on that stream is 2.5% to 3.1% of mean loss at any window. Ranking them perfectly is worth a rounding error, which is the real reason this closes rather than continues.

What I could not check

Seven of the tools the question was asked about were uninitialised submodules in the worktree that round ran in, so they were not examined. That was a fact about my environment rather than a finding about them, and two of them were checked out and audited in a later round, which found nothing across nineteen runners. The rest are still unexamined.

What generalises

A taxonomy is a shopping list, and the first question is which aisle you are allowed in. Nine of fifteen sections went away on one property of my tool: there is no reward, and its nearest candidate is a verdict saying this program matches its own contract, which is exactly what is true of a correct implementation of the wrong request. Working that out cost an afternoon and saved building any of them.

Put the dumb control in before the clever arm, and write down in advance that it matters more. Token count at 0.717 is the whole result here. Reporting the pseudo-count arm’s 0.784 next to a null, and not next to the length control, would have read as a success.

The corpus can be the classifier. When an arm ties with the name of the website a row was scraped from, and then falls below the bar as soon as you hold the website fixed, the arm was measuring provenance. That was visible in a two-row table of source against outcome, and I did not look at that table until after the arms had run.