The Table Named After the Programming Language
appgen is a tool of mine that builds a small working application from one sentence, with no language model in the generating path. Before it can build anything it has to find the entity: the noun the app is about, which becomes its table name, its form fields and its URL paths. Given “Build a recipe manager with search” the answer should be recipe.
Given “Build your own Blockchain in JavaScript” the answer was javascript, and the tool built a table of javascripts.
It was not alone. “Let’s implement a cryptocurrency in Kotlin” produced kotlin. “Create a Twitter Bot in Python Using Tweepy” produced tweepy. “Build a Microblog with Flask” produced flask.
One more rule matters for everything below, because it decides whether a wrong answer is cosmetic or fatal. The tool’s checker rejects any artifact whose entity is not a word of the request. So an entity taken from the wrong part of the sentence still builds; an entity invented from nowhere does not.
Private repo, so no code link. What is checkable is the provenance of each figure, which I give as I go.
Why the language wins
The rule scans the request for candidate nouns. On an imperative sentence it prefers the rightmost content word, because English noun phrases put the head last: in “a recipe manager” the head is manager, and the informative word sits just to its left, so reading rightwards gets you to the right neighbourhood before scanning back.
A trailing in JavaScript, using Tweepy or with Flask is an adjunct. It says what the thing is built with, not what it is about. But it is also, almost always, where the rightmost content word lives. So the part of the sentence that is explicitly not the subject sits exactly where the rule looks for the subject.
The fix names no languages
Strip the tail, try the head, and keep today’s answer if the head names nothing. Fourteen lines, and it shipped.
Build your own Blockchain in JavaScript javascript -> blockchain
Let's implement a cryptocurrency in Kotlin kotlin -> cryptocurrency
Create a Twitter Bot in Python Using Tweepy tweepy -> bot
Code a 2D Game Engine using Java - Full Course beginner -> game
Build a Microblog with Flask flask -> microblog
The thing I like about this one is what is not in it. There is no list of programming languages. javascript, kotlin and tweepy are not named anywhere in the change, and the rule would work the same way on a framework invented next year. A hand-written list of languages would have scored the same on this corpus and would have started rotting the day it shipped.
And the title of this post is from the half with no evidence behind it
Here is the part that cuts against everything above.
The scored corpus is 75 project titles with hand-written gold answers. They come from three third-party lists, taken complete, none written with any knowledge of this tool: 861 linked items, of which 87 have gold answers, of which 75 reach the entity reader at all. On those 75 the fix moved the score from 31 to 37.
Break the six gained rows down by which preposition fired:
| preposition | rows changed | rows gained |
|---|---|---|
in |
1 | 0 |
using |
2 | 0 |
with |
12 | 6 |
Every one of the six gains is a with row. The in branch, which is the one that produced javascripts and the one this post is named after, changes a single scored row and gains nothing. It changes 57 rows on a larger corpus that has no gold answers, so I can see it doing something and cannot show that the something is right.
So javascripts is a real defect, really fixed, and “31 to 37” is not evidence that fixing it helped. The reach is carried by in and using; the measured accuracy is carried entirely by with. Those are two different claims and only one of them has gold labels underneath it.
Both arms in that measurement are the real functions rather than reimplementations. The before arm is the shipped rule with this change lifted back off; the after arm is the shipped entry point itself. An earlier round reimplemented an arm, watched it drift from the thing it claimed to measure, and took a whole round to notice.
One word was worth more than the grammar
X Clone means a clone of X, and the tool was answering clone.
The rule already knows about generic heads: nouns that say nothing about the domain, so that when it finds one it scans leftwards for the real subject. That set already held engine, box, book, collection, catalog, app, platform, system, service and program, which is why “Trello Platform” already answered trello. clone was in neither that set nor the stop list of words that can never be an answer.
Adding it is a one-word change, no new code path runs, and it shipped too. On the same 75 rows it moved the score from 37 to 49: fourteen rows moved, twelve gained and none lost. With clone in place the fourteen-line grammar fix is worth +8 rather than +6, since both arms move together, so the word is worth about one and a half times the grammar.
One of the twelve deserves a discount. Google Docs Clone answers docs where the gold set wants document or doc, and it scores as a gain on a technicality; a reader who thinks both answers are wrong in the same way should read eleven.
The mechanism was already written down in the tool’s own README, one paragraph from the problem the grammar fix solved: on noun phrases the leftmost scan takes a leading modifier, naming clone as four rows of it. The fix was sitting in the documentation waiting for me to read my own writing.
Where those twelve rows came from, and why that is a problem
27 of the 75 rows contain clone. The corpus is a list of project ideas, and “build a clone of X” is one of the two or three things such a list is mostly made of.
And the word was not chosen independently of these rows. The README paragraph that named it was itself written from an error analysis of this corpus, which lists four X Clone Application titles by name. So clone is a word selected from the rule’s own mistakes on the corpus it is then scored against, which is the setup an earlier post in this series measured directly: words picked that way gain in sample by construction, and the question is whether they travel.
Here they do not. The controls say so by being inert, which is the more useful way to read them:
- On a second corpus of 47 rows, nothing moves: 36 before, 36 after, no row gained or lost. One of those 47 does contain the word, and a different rule fires on it first, so the channel is blocked rather than empty.
- On the 1,440 requests the repository’s own generator can produce, nothing moves either: 772 before, 772 after. Those 1,440 requests contain the word zero times.
That is exactly the pattern the earlier post measured with a random-draw control: a real gain on the corpus the word came from, and precisely nothing anywhere else. Two rounds, two methods, same answer.
Where it does regress
It regresses, and the round went looking for where. Five hand-written requests a user could plausibly type come out worse:
a clone tracker clone -> record
a clone log with search clone -> record
a VM clone tracker clone -> record
a plant clone log for my nursery clone -> plant
a clone inventory with creating… clone -> inventory
If you genuinely want to track clones, clone is now a generic head and the tool scans past it. Two of those five land on a reasonable alternative. Three land on the tool’s fallback noun, record, which is not a word of the request, so the checker rejects the artifact outright.
None of these appears in any corpus this tool is scored on. They were written by hand precisely because no corpus contains them, which makes the regression real and unmeasured rather than real and small.
One more number, so two posts do not look like they disagree
49 of 75 is the entity reader scored on its own. Elsewhere in this series the same rule appears at 47 of 75, and that is the whole shipped path with a routing layer in front of the reader. Same corpus, different functions, and the two should not be subtracted. The rounds here measure the reader because the reader is what the change touches, and the source write-up is explicit that the delta is +6 either way and that you have to name the harness when quoting one.
What generalises
A rule that scans for the most important word will find whatever is where the important words usually are. The language name did not fool a heuristic about languages; it occupied a position. Fixing the position generalises to frameworks that do not exist yet. Fixing the names would have been a list, and a list is a promise to keep updating it.
The vivid example and the measured gain were in different halves of the corpus. I fixed javascripts, wrote the post about javascripts, and the gold labels only ever confirmed the with rows. The breakdown that shows this is one extra column of the same table, and I did not produce it until the round was almost written up.
Say which of your controls could have fired. Two controls here report no change. One of them covers a population where the word under test never appears, and the other has a single row holding it, blocked by a different rule. Listed as green checks, those read as two independent confirmations; they are one measurement and two abstentions.