A Ceiling That Was Measuring My Test Harness

appgen is a tool of mine that turns a sentence into a running application without a language model in the generating path, and then proves the application works by starting it and driving it. It emits three kinds of program (an HTML web app, a JSON API, a command-line tool) in three languages (Python, Node and C). That is nine configurations, and anything planning a build in front of it has to pick one. Before measuring any planner I wanted the bound: what would a perfect choice of kind and language have won, against always picking the same one?

So I built every request under every configuration. 88 requests taken complete from florinpop17/app-ideas, 9 configurations, 792 cells (one cell is one request built under one configuration), all scored by the same function the real runs use.

requests where the configuration can change the outcome 20 of 88
requests constant across all nine configurations 68
ceiling: some configuration produces a working app 29
best fixed configuration (web/python, no input at all) 29
headroom over a constant 0

The headroom is the ceiling minus the best fixed choice: what a perfect chooser wins over a chooser that never looks at the request. Read straight down, that table says the choice is worth nothing. It does not say that. The reason is one row further in.

One number needs explaining before anything else, because it looks alarming. Only 29 of 88 requests produce a working app, and the other 59 are not failures: 57 of them appgen refuses outright, exiting with a message rather than guessing, because they ask for things outside its grammar. Two more build and come back broken. Refusing is the behaviour I want, and this post is about the 29 and the 2, not about the 57.

This lives in a private research repo, so there is nothing to link. Every figure below is read out of a results file or a command I re-ran, and I say which one.

Eight of the nine were invisible

Break the twenty movable requests out per configuration:

configuration works broken could not be driven
web/python 18 2 0
each of the other eight 0 0 20

The driver is the part of the harness that takes a built artifact and decides whether it works, and it had exactly one surface. It started the artifact, fetched /, looked for an HTML form and proved that a write round-tripped. That is the contract for a web app; it is not the contract for anything else. So a JSON API, which publishes no index page, a command-line tool, which has no port, and every non-Python language came back as could not be driven. My own harness records that as a defect in the harness, not as a verdict on the program. That is the right call, and it is exactly what makes the zero useless: a configuration that cannot lose also cannot win.

So the honest statement was not “the choice is worth nothing”. It was “the headroom over kind and language cannot be measured with this driver at all”. The zero was the instrument’s reach, not the tool’s. Four separate rounds shared that driver; every claim any of them made about kind or language rested on a scorer that could see one of nine.

A second casualty in the same round is the cleaner illustration. I had put a small online learner in that seat to choose the configuration per request, with its own controls beside it: the same learner with its search switched off, a fixed best guess that ignores the request, and an oracle that is allowed to see the answer.

arm working apps
the learner 29
the learner with its search switched off 29
always pick the best fixed configuration 29
an oracle over all nine 29

The learner picked web/python on 84 of 88 requests and reached the oracle. That is not the learner succeeding, and it is not the learner failing. Eight of nine options are unscoreable by construction, so the target is constant. The oracle equals a constant, and every arm from a random one upward hits it. A write-up reporting “the learner matches the oracle” would have been reporting the driver.

Building the driver the zero was asking for

A leg is one surface of the driver: a procedure for exercising an artifact of a given kind and deciding whether it did its job. There was one. Two more, dispatched on the kind the artifact publishes about itself:

  • The API leg reads the route the artifact names in its own printed curl line and fetches it. It learns the record shape from the service rather than guessing field names, posts a random sentinel value, reads it back, then asks for a record that was never created, which must answer 4xx. That last clause is the check an HTML form cannot make: a service that answers every path with the same document round-trips a write and discriminates nothing.
  • The CLI leg runs the published command chain as two separate subprocesses and requires the sentinel in the second one’s output. A program cannot pass by echoing its own arguments, because the process that echoed them has exited.

The existing web leg was not edited. I checked that function by function on the parsed syntax tree, not by reading the diff. “I only added a branch” is true right up until it is not.

The gate the whole round stands on is that no web row moves. Every cell was built twice and scored once by each dispatch, instead of scored twice from one build. The leg writes to the artifact’s store, so a second scorer would see the first one’s row.

cells scored 792
cells whose verdict moved 120
of those, whose artifact is neither API nor CLI 0
cells where the two builds’ output differed 0

The 120 are exactly the 60 API and 60 CLI record-keeping cells. Not one web cell moved, which is the claim every comparison this corpus has been used for depends on.

That took the reachable set from one configuration of nine to seven. The two left over were web/node and web/c, for a reason almost too small to write down: the web leg invokes an artifact as python3 <file>, and the run-line parser required that literal string. A follow-up round reached them by rebuilding the whole table again, 1,584 builds.

That rebuild was not one continuous run, and the reason the round gives is worth more than a clean number would be. The job was killed three times without a traceback, and during one resume two processes wrote to the same checkpoint for roughly 27 rows. Each checkpoint rewrites the table whole, so the published file is one process’s complete list rather than a merge, and it was checked for duplicates and corpus order before being read. Several hundred further builds were run and discarded in that incident, and the round cannot attest that none of them is in the file. Which process wrote the final checkpoint is a fact no artifact records.

The same zero, now saying something

the original driver the new one
configurations that can score a record-keeping artifact 1 of 9 9 of 9
ceiling 29 29
best fixed configuration web/python, 29 web/python, 29
headroom over a constant 0 0
requests the nine configurations disagree on 20 0
requests constant across all nine 68 88

The top four rows are identical. The bottom two are why the zero changed meaning entirely. Under the old driver the nine disagreed on all twenty movable requests, and every one of those disagreements was could not be driven against a real verdict. Under the new one they disagree on nothing. Row for row, all nine return the same verdict on all 88 requests, and the two that come back broken are broken in all nine.

So the answer to what would a perfect choice of kind and language have won is: nothing, and now for a reason about the tool rather than about the harness. appgen composes the same schema from the same sentence whatever the kind flag says; the kind decides which idiom those records are served in, and every idiom works.

What the new legs cannot see

A leg that returns “works” 120 times out of 120 has told you nothing until you show it can say otherwise. Breaking the leg’s own source does not do that; it shows only that the leg’s tests notice. What nobody had shown is that the legs fail when the artifact is broken, which is the only thing they exist for.

So: build appgen’s own artifacts, inject one named defect into the emitted source by a string replacement that must match exactly once, and score both arms. An anchor matching zero or twice is reported as not producible and counted in neither arm. An artifact that was never injured is one of appgen’s own, it works, and counting it would score the leg as lenient for my own failure to patch it.

cells verdict working verdict not working
appgen’s own artifact, untouched (30) 30, as they should be 0 false positives
one named defect injected (54) 15 false negatives 39 caught

Fifteen misses, and all fifteen were written down before the run; zero unpredicted.

Two things scope that 39 of 54, and both cut against reading it as a discrimination rate. All 54 injured cells are one request, Your first Database app!, carrying fifteen defects across two kinds and three languages, so the catch rate is measured on one artifact shape rather than on a corpus. And 12 of the 15 misses are refused by appgen’s own verifier before the leg ever runs, which turns them into could not be driven rather than into a passing broken app. Only 3 get past both.

The sharpest miss is worth naming. The status clause requires a 4xx for a record that was never created. Driven at three points on real artifacts in all three languages: a service answering 200 is caught, one answering 500 is caught, and one answering 400 passes. That last one is correct, because 400 is inside the accepted class. Then delete the detail route entirely, so every id falls through to the service’s no such route 404. It passes in all three languages, while appgen’s own verifier fails the same source on two separate checks. The leg’s “works” therefore means the collection endpoint is not a catch-all. It does not mean the record can be fetched back by its id, and I would have said it did.

Not every catch is an accusation either. Six of the 39 come back as could not be driven rather than broken: a service that never binds a port, and a service that answers 500 to its own published POST. Both are broken artifacts recorded as the harness’s fault, which is the conservative direction and conservative in appgen’s favour.

What this does not say

The zero is about this corpus. Of the 88 requests, appgen answers 20 with a record-keeping app at all, and those 20 are the ones the configuration can move. A corpus whose requests exercised something other than storing and listing records could separate the nine; this one does not.

“Works” is not “correct”. Everything above is a verdict that the artifact does the thing its own published invocation says it does. Whether it is the program the person asking wanted is a different question, measured elsewhere, and it comes out worse.

And the obvious objection to the headline is the right one to raise yourself: if every request in the corpus is record-shaped, the finding might be that CRUD is CRUD in any idiom rather than anything about appgen. That is exactly what it might be. What the round can say is that the choice buys nothing here, measured rather than assumed, which is more than it could say before.

What generalises

A ceiling computed with an instrument that cannot see most of the options is a measurement of the instrument. I got the same number twice and it meant two different things. The only way to tell them apart was to widen the instrument and watch whether the number moved. It did not, which is what makes the second zero worth having and the first one worth retracting.

The tell was available before any of this. It is the row saying eight of nine configurations returned could not be driven on every movable request. A result where one arm is unanimous and the rest are structurally silent is not a finding about the arms. The giveaway is that the oracle, the learner, the learner’s control and a constant all land on the same integer. When every arm from random upward hits the ceiling, the ceiling is the harness.