Tagged measurement
37 write-ups.
Pinning What a Function Answers Across Rewrites and Time
A pin on a nondeterministic function fails at random on the next machine, and the blame lands on the pinning tool. Every candidate goes through nondet first; the committed control turns the gate off, changes not one byte of the tree, and asserts the check goes red anyway.
Measuring How Cost Scales by Counting Instead of Timing
Insertion sort on seed 17 at n=64 performs reads 3812 and writes 1848, the same integers in Python and JavaScript, because the generator, the iteration rule and the subscript rule are part of the contract. The same parity suite caught its own dependency formatting one number two ways above a million.
A Curve Fitter That Refuses to Answer
A tool that always produced a constant would be useless and would still pass every test that checks it produces one. 0.2.0 fixed a defect the suite had pinned: with error bars from scatter alone, a quantity measured exactly was reported as one that could not be determined.
Breaking CI Guards on Purpose to Prove They Can Fail
A fix I believed in was proved a no-op by the tool auditing itself: mtime invalidation has one-second granularity and the edit cycle is milliseconds, so stale bytecode looks fresh either way. PYTHONDONTWRITEBYTECODE is the guard actually holding; the surviving mutation corrected the README.
A Determinism Check Has to Leave the Process
Set iteration order is stable inside one Python interpreter and different in every new one, so the repeat-it-twice check reports deterministic every time. nondet probes in fresh processes instead: 9 of 9 nondeterministic fixtures caught, 0 of 10 deterministic ones falsely flagged.
Ten Packages, One Rule: A Check Must Be Able to Fail
Eight of the ten went up in one day, extracted from a year of measurement work in which the thing lying was usually the instrument, not the code. Each package encodes one way a green check can mean nothing, and each commits a named test that would report its own premise wrong.
What Mutation Testing Frameworks Do When a Timeout Kills Them
I expected a kill to leave all four with a mutated file on disk. Three of them never write a mutant into your tree at all, and the fourth is left dirty by an ordinary SIGTERM, which is a gentler signal than the one I built the suite to test for.
Searching for a Program Instead of Generating One
A primitive algebra searched from examples solved five of twenty-eight exercises, and every one already had a hand-written implementation in the same repo. A behavioural index recognised one function in forty-eight, and zero of eleven when handed other people's correct solutions to the same exercises.
An Open Redirect Guard That Rebuilt What It Rejected
isSafeReturnTo("//bad.example.com") returns false, and the fallback for that rejection handed the browser the same string back, rebuilt from the request path. 0.3.0 closed one open redirect and shipped a second one for eight more days; both had been there since the first publish in 2019.
When a Test Suite Rejects a Correct Program
Eleven of twenty-one operations never passed my frozen checks, and I wrote that up twice as a limit of the generator. A wrapper that only re-expressed each program's own stdout repaired 90 of 102 failures, changing no algorithm.
What a Rate Limiter Reports, and What the Server Receives
My crawler's per-process limiter reported a perfect interval while the host received every request at once. Pointing the same server-side test at Bottleneck, the most-downloaded rate limiter on npm: six cluster workers configured for 2 requests/second delivered 14.04, and every limiter reported 495.8 ms spacing.
The Working Set That Never Saturated
A count table is random access, so in principle only the pages a document touches need be resident. Across eight real files, seven were still pulling in new pages after 2,700 completions. Then I measured the latency constant the verdict rested on and found it optimistic by 1.8x.
Just Train More: Measuring the Exchange Rate
The obvious objection to a zero-parameter cache beating a small transformer is that the transformer is undertrained. Sixteen times the training data moved the crossover 6.6x, which is document length times 1.60 per doubling: about as much as letting the cache read the rest of the file it was already sitting in.
When a Zero-Parameter Cache Overtakes a Transformer
A count table over the current document has no parameters and no training. Somewhere between 60 and 250 tokens of document it passes a 1.43M-parameter transformer, and by 1000 tokens it wins top-1 by 0.064, a 43% relative margin. Adding the transformer on top then buys 0.002.
The Last Non-Neural Candidate, and It Did Not Clear the Bar
The bar was a slope: keep converting extra data into accuracy after exact-context statistics saturate. A full hierarchical Pitman-Yor model with Gibbs sweeps and inferred discounts moved the intercept and left the slope alone, halving with every doubling exactly as cruder count models did.
The Embedding Table Was 72% of the Model
Nearly three quarters of my small transformer was a lookup table, so the dial that mattered was embedding precision, not network precision. int4 with one scale per row costs 16 KB more than one scale per tensor and recovers 2.2 of the 2.5 points per-tensor int4 loses.
Was TinyStories the Domain or the Vocabulary?
TinyStories is famous for a ~1,500-word vocabulary, so I capped mine at 1,500 and scored every arm on identical positions. It came out 4 points worse at top-1 than a 4,000-word cap, and more than a third of its output was the unknown-word token. The transferable lever was the domain, not the vocabulary.
What a Browser Extension's Test Suite Cannot Reach
130 Node assertions passed while printing could not open a dialog at all, and choosing PDF silently broke the editor handoff. Both bugs sat at a boundary: my code meeting a browser API, and one internal stage meeting another. Pure-function suites are structurally blind to those.
Where Four Backends Stop Agreeing
My language compiles to C, WebAssembly, ARM64 and a bytecode VM, and correctness is defined as all four producing identical output. There are exactly four places where they are allowed not to, and casting an out-of-range float produces a different result in every one of them.
The Fixpoint No Test Suite Can Check
My language has 94 tests that run every program through four backends, compared byte for byte. All 94 run a program the JavaScript compiler built, so none can detect a compiler that disagrees with its own output. Self-hosting found three bugs the suite could not reach.
What a Language Needs Before It Can Compile Itself
Porting a 129-line JavaScript lexer to my own low-level language produced 355 lines. Almost none of the difference is verbosity: six specific absences account for it, one of them caused the port's only correctness bug, and one I filed as merely annoying turned out to be a correctness problem two stages later.
No Add-on Can Block Ads in This Browser
WebKit grants an extension webRequestBlocking without honouring it, and accepts declarativeNetRequest rulesets without applying them. uBO Lite with EasyList enabled blocked 0 of 5 tracker probes, while a content rule list the browser compiled itself blocked the same probe every run.
Measuring the Wrong Process for Eight Months
My benchmark credited each tab with one process and counted every other process as nothing. Re-measured against every process the browser actually owns, the light-pages result inverts: the ladder I published as 2.63x better than doing nothing is 1.35x worse.
What Maintaining a Forked npm Package Actually Buys
I would have said "eighteen releases of upkeep". What npm audit can see is one digit in a dependency range: a caret below 1.0.0 walls off the published fix, so it reports "No fix available". What it could not see is the open redirect my fork shipped for seven years while auditing clean.
Choosing Video Frames by Content Instead of by a Clock
At an equal frame budget, uniform sampling missed 9 of 28 shots across five real videos and spent 12 of its 40 frames re-photographing things it had already seen. Content-aware selection missed none and wasted two; here is the measure where it still comes out level, and the one video where uniform sampling wins.
Ranking Language Models by How Well They Spot Liars
I built a benchmark to rank language models on spotting liars, then ran 522 games of Mafia through it. The same model scored -0.400 and +0.078 in one run, 52 games apart. Day 1 accusations came in at -0.005 against chance on n=2,103, and the whole thing reproduces in 11 seconds with no API keys.
An AI Capture-the-Flag Tournament: What the Scoreboard Counted
A preliminary report concluded that models under 3B parameters cannot chain exploits. A later tournament with much larger models put an 8B model last: the engine credited 2 of its 279 logged captures, because a per-turn marker had counted each of the 232 times it read bonus flags off its own machine.
Checking a Cost Model Against a Stranger's Config File
I bounded an ambiguity in a config file at 2.4% and called it immaterial. The model's own header settled it months later, and correcting the assumption pushed a gate from 1.1% off to 3.4% off, outside the tolerance it had been passing.
Why a Small Transformer Can't Copy a Word It Hasn't Seen
An 11.9M-parameter transformer scores 83% strict on the domains it trained on and 0/8 on a single unseen noun. Three experiments to find out why, one of which corrected the diagnosis I had already written down, and a zero-parameter mechanism that does the copying 20 times out of 20.
How a Dedup Pass Deleted My Training Curriculum
A fine-tuning pipeline that weighted its best examples 6x by duplicating them to 94,022 lines, followed by a dedup pass that handed the trainer 50,145, each exactly once. Two retrains were rolled back for regressing before I found that the graded curriculum had never reached the trainer at all.
Error Feedback, Gradient Compression, and Why Adam Breaks It
Error feedback makes a biased gradient compressor unbiased over time, and under SGD it restored the full-precision trajectory to three digits. Under Adam it landed 1.9 times further from the optimum than no correction at all, and the fix I published helps just as much with no quantization in the run.
A Better FP4 Gradient Quantizer That Training Couldn't Notice
A per-block scale that cuts FP4 gradient-quantization error on 45 of 45 tensors, 14% against the published state of the art, two rented-GPU training runs that landed within noise anyway, and the measurement that explains both: the gap only shows at million-token batches, 35 to 643 times larger than anything I ran.
What a Code Completer's Eval Never Measures: Ghost Text
A hybrid n-gram and 2.48M-parameter code completer that gets 54.6% of tokens right, tuned dial by dial against its own held-out eval. Then I scored the one feature that emits more than a single token and it came back at 0.070, because the eval hands the model a true prefix and the feature hands it its own output.
Comparing INT4 and NVFP4 Palettes on Real Gradient Tensors
A palette study on synthetic blocks said evenly-spaced INT4 beats NVIDIA's FP4 grid once a Hadamard rotation gaussianizes the data, and loses by 2.3x without it. On real gradient tensors it wins either way, 41 of 45 without the rotation, because my stand-in for unrotated data put the outliers in the wrong place.
An App Generator That Verifies Everything Except Its Parser
A natural-language app generator whose verification sweep is exhaustive over everything except the part that reads English. The composer scores 100% on 810 cells built from feature sets; the parser that has to produce those feature sets scores 0.70, and nothing prints that number next to the other one.
What a Hibernated Browser Tab Actually Costs
My simulated tab scheduler priced a hibernated tab at 32 KB; real WebKit charges 39 MB. It claimed 7.9–11.5x less memory than an unmanaged browser and delivered 1.4–2.6x. That one number, wrong by three orders of magnitude, explains the entire gap.
Predicting the Speed of a 276B Model Streamed From an SSD
A from-config cost model for streaming a mixture-of-experts model off an SSD, validated byte-for-byte against a runtime I didn't write. Then the SSD benchmark turned out to be measuring RAM, and the real run missed the speed prediction by 23×, for the one reason the model had declared it couldn't account for.