How a Pluralisation Rule Broke a Correct Answer

appgen is a tool of mine that turns a sentence into a running application with no language model in the generating path. Give it “a support desk system with priorities and comments” and it writes a dependency-free Python file, starts it, and drives every feature you asked for over real HTTP before showing it to you. One of the things it has to decide is the entity: the noun the generated app calls a row, which becomes a table name, a form field and a URL path.

The checker that judges a finished artifact has two halves. The first starts the program and proves it stores and returns a record. The second, which everything below turns on, asks whether the entity is a word of the request, and it exists because an earlier round paid for an app whose nouns came from a domain table rather than from the sentence in front of it. A verdict is one of works, broken, or could not be driven.

In a previous post I measured what happens when you put a local language model in front of appgen to do the planning, and found that the model’s plan rescued zero requests that an empty seat lost. Five requests went the other way. This post is about one of them, because the model got that one right and the pipeline destroyed the answer afterwards.

The request was “Browse, Find Ratings, Check Actors and Find you next movie to watch”. The model answered movies. Here is what the tool did with it:

singular_of("movies")        -> 'movy'
_stem_forms("movy")          -> {'movy'}
'movy' in _stems(request)    -> False
verdict                      -> broken: "entity movy appears nowhere in the request"

appgen singularised a correct answer into a non-word, and its own checker then rejected the result as evidence that the reader had hallucinated.

The repo is private, so take the code on trust and the numbers on evidence: each one below is traced to the file or command it came from.

The guard that fixes the other two, and correctly declines this one

Three of the five lost requests are the entity field, and the obvious repair contains no model, no learner and no reading: if the supplied entity is not a word of the request, do not pass it, and let appgen derive its own. That is a one-line lexical constraint, and it calls the checker’s own stemming functions, so the arm agrees with the checker and not with a paraphrase of it.

The population is the 18 requests from that run where a built artifact reached the entity check at all. Replaying them with only the entity token changed, and with no model contacted:

built outcome
against the model’s own plans 2 recovered, 0 newly broken
against the plain derivation 17 of 18 identical; the 1 that differs, the derivation wins
on its own terms 15 of 18 working

That 17 of 18 is an agreement count, not a success count. It says the two arms land on the same verdict, including on rows where both of them lose.

The two it recovers are A clone of Facebook’s Instagram app (posts → instagram) and Click list item to display item details (plants → detail). Neither posts nor plants appears anywhere in its request, dropping the field puts a word of the request in the slot, and the artifact then builds and works. That guard shipped, as a withheld field rather than a substitution, which comes to the same thing because appgen then does its own derivation.

The constraint does not fire on movies, and it should not. movies stems to movie, and movie is in that sentence, so by the rule the checker itself applies the entity is a word of the request. The guard passes it, and the pipeline breaks it three functions later. The two changes are independent and they fix different halves of the same three rows.

Two things about that guard are worth stating because they cut against it. It is not free: against the frozen gold labels it costs the model three rows of raw read accuracy, 55 down to 52 out of 75, and it fires on 39 of those 75, over half. And against the plain derivation it is 7 : 2 discordant (seven rows it wins that the derivation loses, two the other way) at p = 0.18, which is unresolved, not better.

Those 75 rows are a different and larger scoring set from the 18 above: a held-out corpus of project titles with hand-written gold answers, used for grading the entity read on its own rather than the artifact it produces. Here it only says the constraint is not free.

The derivation does not always take a word from the request either

While checking that, I found a claim in my own tree that is not true. entity_from_request is described as taking a word out of the request by construction. It has a generic fallback:

entity_from_request("Your first Database app!")  ->  'record'
entity_from_request("Voting App")                ->  'record'

Neither request contains that word, so the same lexical check refuses both, and both arms lose those rows for the same reason. Anyone pricing that check as a free training label should price it at the 2-in-18 the fallback costs here rather than at zero.

Why movy happens: two candidates round-trip and the order decides

The singulariser works by generating candidates and keeping the ones that pluralise back to the word it started with. For movies there are three: movy, movie and movies itself. plural_of("movy") is movies. plural_of("movie") is also movies. Both round-trip, both are admissible, and the iteration order picks the loser. The -ies → y rule that produces movy is right for companies → company and wrong here, and the round-trip check cannot tell them apart because it is satisfied either way.

The table name is not what breaks. Both candidates pluralise to movies, so the generated app calls its table the same thing whichever one wins. What moves is the singular, which is the value the checker reads.

entity supplied singular table verdict
movies movy movies broken: entity movy appears nowhere in the request
movie movie movies works

Both obvious repairs break on the same two words

The first obvious repair is to prefer any non-identity candidate. An earlier commit priced exactly that and refused it, because over the closure of the tool’s irregular-plural table it turns series into sery and species into specy. I re-ran that measurement before touching anything, and it reproduces: the same two regressions, and only those two.

The second obvious repair is to break the tie with a part-of-speech lexicon, asking which candidate is a word anyone lists. The natural hook is the project’s existing out-of-vocabulary check, and it recurses until the stack ends, because that check normalises its input through the very function being repaired. It recurses on series and species. It does not recurse on movies, because the cycle needs the candidate’s own singularisation to hit a tie too, and movy does not.

That is the part I would not have predicted. A test written against this round’s headline word passes while the defect is fully present, so the tests assert the recursion on series and species, and assert that movies does not trigger it.

Those are the same two words in both failures, for one reason: series and species are themselves listed words with more than one round-tripping candidate. The narrower question, is this candidate a word anyone lists, is a raw table-membership test, it is cycle-free, and it answers the tie.

Across three populations, that narrower test makes four moves and all four are repairs, with no regressions:

population n ambiguous what moves
closure of the irregular-plural table 72 38 mouses: mous → mouse, gooses: goos → goose
every word appearing in either request corpus 433 55 s: '' → 's'
entities supplied to appgen (149 in total) 94 unique 82 movies: movy → movie

movies never appears in a request. It arrives as a supplied entity, 4 times in those 149, and it is a model’s read of the word movie in that sentence, pluralised. A round that measured reach over request vocabulary alone would have reported zero, correctly, for entirely the wrong reason.

The fix ships switched off

The lexicon is 200,878 entries in a 717,252-byte gzipped file, and the first lookup has to load all of it. The round measured that at 0.097 s and 40.7 MB of resident memory; re-running it while writing this gave 0.157 s and 47.7 MB, so treat it as a fraction of a second and tens of megabytes rather than as a constant. Warm lookups are free either way. It repairs one supplied entity in ninety-four. No import enters the generator at module scope, because a module-scope import would turn an optional check into a startup requirement for every caller.

The earlier commit refused a repair whose cost fell on every caller and whose benefit was one noun. This is the same shape with a better ratio and a smaller cost, and I did not turn it on either. What changed is that the price is now written down, so someone can decide to spend it.

One thing about the hook is not established and I am not going to round it up. It is in-process, and the test suite it was run against spawns appgen as a subprocess 24 times; those builds ran unhooked. What passed is the in-process half, and reporting that as “the suite is fine” would have left the denominator out.

The mutation that was wrong, not the suite

A way to find out whether a test suite can actually fail is to break the code on purpose, one defect at a time, and check that something goes red. Four such defects here, four caught. One of them showed up green on the first pass, and the defect was in the defect rather than in the suite.

Every tied word in every population has exactly three round-tripping candidates, movies being ['movy', 'movie', 'movies']. So a break that only consults the lexicon when there are more than two candidates leaves it consulted on every real input and changes nothing. Replacing it with one that offers only the first candidate tests the claim that actually matters, which is that every candidate reaches the tie-breaker rather than just the first one.

What generalises

An oracle that is satisfied by more than one answer is not an oracle. The round-trip check here is genuinely a good test, it rejects real garbage, and it accepts movy and movie with equal enthusiasm; everything after that is iteration order, which is to say it is luck that happened to be stable. Wherever a check admits a set rather than a value, something downstream is choosing from that set, and it is worth finding out what.

The second thing is smaller and I keep relearning it. I picked the repair, then wrote a test against the word that motivated the repair, and that test would have passed against a completely broken implementation. The word that shows a bug and the word that exercises the mechanism are often not the same word.