Tagged llm
20 write-ups.
Searching for a Program Instead of Generating One
A primitive algebra searched from examples solved five of twenty-eight exercises, and every one already had a hand-written implementation in the same repo. A behavioural index recognised one function in forty-eight, and zero of eleven when handed other people's correct solutions to the same exercises.
When a Test Suite Rejects a Correct Program
Eleven of twenty-one operations never passed my frozen checks, and I wrote that up twice as a limit of the generator. A wrapper that only re-expressed each program's own stdout repaired 90 of 102 failures, changing no algorithm.
The Working Set That Never Saturated
A count table is random access, so in principle only the pages a document touches need be resident. Across eight real files, seven were still pulling in new pages after 2,700 completions. Then I measured the latency constant the verdict rested on and found it optimistic by 1.8x.
Just Train More: Measuring the Exchange Rate
The obvious objection to a zero-parameter cache beating a small transformer is that the transformer is undertrained. Sixteen times the training data moved the crossover 6.6x, which is document length times 1.60 per doubling: about as much as letting the cache read the rest of the file it was already sitting in.
When a Zero-Parameter Cache Overtakes a Transformer
A count table over the current document has no parameters and no training. Somewhere between 60 and 250 tokens of document it passes a 1.43M-parameter transformer, and by 1000 tokens it wins top-1 by 0.064, a 43% relative margin. Adding the transformer on top then buys 0.002.
The Last Non-Neural Candidate, and It Did Not Clear the Bar
The bar was a slope: keep converting extra data into accuracy after exact-context statistics saturate. A full hierarchical Pitman-Yor model with Gibbs sweeps and inferred discounts moved the intercept and left the slope alone, halving with every doubling exactly as cruder count models did.
The Embedding Table Was 72% of the Model
Nearly three quarters of my small transformer was a lookup table, so the dial that mattered was embedding precision, not network precision. int4 with one scale per row costs 16 KB more than one scale per tensor and recovers 2.2 of the 2.5 points per-tensor int4 loses.
Was TinyStories the Domain or the Vocabulary?
TinyStories is famous for a ~1,500-word vocabulary, so I capped mine at 1,500 and scored every arm on identical positions. It came out 4 points worse at top-1 than a 4,000-word cap, and more than a third of its output was the unknown-word token. The transferable lever was the domain, not the vocabulary.
Choosing Video Frames by Content Instead of by a Clock
At an equal frame budget, uniform sampling missed 9 of 28 shots across five real videos and spent 12 of its 40 frames re-photographing things it had already seen. Content-aware selection missed none and wasted two; here is the measure where it still comes out level, and the one video where uniform sampling wins.
Ranking Language Models by How Well They Spot Liars
I built a benchmark to rank language models on spotting liars, then ran 522 games of Mafia through it. The same model scored -0.400 and +0.078 in one run, 52 games apart. Day 1 accusations came in at -0.005 against chance on n=2,103, and the whole thing reproduces in 11 seconds with no API keys.
An AI Capture-the-Flag Tournament: What the Scoreboard Counted
A preliminary report concluded that models under 3B parameters cannot chain exploits. A later tournament with much larger models put an 8B model last: the engine credited 2 of its 279 logged captures, because a per-turn marker had counted each of the 232 times it read bonus flags off its own machine.
Checking a Cost Model Against a Stranger's Config File
I bounded an ambiguity in a config file at 2.4% and called it immaterial. The model's own header settled it months later, and correcting the assumption pushed a gate from 1.1% off to 3.4% off, outside the tolerance it had been passing.
Why a Small Transformer Can't Copy a Word It Hasn't Seen
An 11.9M-parameter transformer scores 83% strict on the domains it trained on and 0/8 on a single unseen noun. Three experiments to find out why, one of which corrected the diagnosis I had already written down, and a zero-parameter mechanism that does the copying 20 times out of 20.
How a Dedup Pass Deleted My Training Curriculum
A fine-tuning pipeline that weighted its best examples 6x by duplicating them to 94,022 lines, followed by a dedup pass that handed the trainer 50,145, each exactly once. Two retrains were rolled back for regressing before I found that the graded curriculum had never reached the trainer at all.
Error Feedback, Gradient Compression, and Why Adam Breaks It
Error feedback makes a biased gradient compressor unbiased over time, and under SGD it restored the full-precision trajectory to three digits. Under Adam it landed 1.9 times further from the optimum than no correction at all, and the fix I published helps just as much with no quantization in the run.
A Better FP4 Gradient Quantizer That Training Couldn't Notice
A per-block scale that cuts FP4 gradient-quantization error on 45 of 45 tensors, 14% against the published state of the art, two rented-GPU training runs that landed within noise anyway, and the measurement that explains both: the gap only shows at million-token batches, 35 to 643 times larger than anything I ran.
What a Code Completer's Eval Never Measures: Ghost Text
A hybrid n-gram and 2.48M-parameter code completer that gets 54.6% of tokens right, tuned dial by dial against its own held-out eval. Then I scored the one feature that emits more than a single token and it came back at 0.070, because the eval hands the model a true prefix and the feature hands it its own output.
Comparing INT4 and NVFP4 Palettes on Real Gradient Tensors
A palette study on synthetic blocks said evenly-spaced INT4 beats NVIDIA's FP4 grid once a Hadamard rotation gaussianizes the data, and loses by 2.3x without it. On real gradient tensors it wins either way, 41 of 45 without the rotation, because my stand-in for unrotated data put the outliers in the wrong place.
An App Generator That Verifies Everything Except Its Parser
A natural-language app generator whose verification sweep is exhaustive over everything except the part that reads English. The composer scores 100% on 810 cells built from feature sets; the parser that has to produce those feature sets scores 0.70, and nothing prints that number next to the other one.
Predicting the Speed of a 276B Model Streamed From an SSD
A from-config cost model for streaming a mixture-of-experts model off an SSD, validated byte-for-byte against a runtime I didn't write. Then the SSD benchmark turned out to be measuring RAM, and the real run missed the speed prediction by 23×, for the one reason the model had declared it couldn't account for.