Writing
Mostly post-mortems: a design claim, the measurement that contradicted it, and what the gap turned out to be.
Pinning What a Function Answers Across Rewrites and Time
· verification, python, measurement · 5 min read
A pin on a nondeterministic function fails at random on the next machine, and the blame lands on the pinning tool. Every candidate goes through nondet first; the committed control turns the gate off, changes not one byte of the tree, and asserts the check goes red anyway.
Finding Duplicate Functions by Executing Them, Not Reading Them
· verification, python, node · 6 min read
Two functions agreed on every rung of the input ladder and were not the same function; one character, ½, turned same into differs with a witness. The guard exists because agreement is cheap: constants, projections and copies agree with everything, and each mistake was made before it was guarded.
Measuring How Cost Scales by Counting Instead of Timing
· measurement, statistics, python, node · 6 min read
Insertion sort on seed 17 at n=64 performs reads 3812 and writes 1848, the same integers in Python and JavaScript, because the generator, the iteration rule and the subscript rule are part of the contract. The same parity suite caught its own dependency formatting one number two ways above a million.
A Curve Fitter That Refuses to Answer
· statistics, measurement, python, node · 6 min read
A tool that always produced a constant would be useless and would still pass every test that checks it produces one. 0.2.0 fixed a defect the suite had pinned: with error bars from scatter alone, a quantity measured exactly was reported as one that could not be determined.
A Check With a Zero Denominator Reports Clean
· verification, python, node · 5 min read
A JUnit file reading tests="50" skipped="50" is a green run of nothing wearing a total, so every parser here returns total and executed separately and the floor applies to executed. Thirteen mutations were applied to the source; the two that survived their first run were worth more than the eleven that were caught.
Breaking CI Guards on Purpose to Prove They Can Fail
· verification, python, measurement · 5 min read
A fix I believed in was proved a no-op by the tool auditing itself: mtime invalidation has one-second granularity and the edit cycle is milliseconds, so stale bytecode looks fresh either way. PYTHONDONTWRITEBYTECODE is the guard actually holding; the surviving mutation corrected the README.
An Exit Code Cannot Say Whether Anything Happened
· verification, node, python · 5 min read
go test on a tree with no test files prints [no test files] and exits 0, measured rather than assumed. didrun wraps a command, demands evidence it did something, and returns four states instead of one number; the committed control asserts that 0 passed satisfying the expected pattern still scores as did-not-run.
Proving the Tree Came Back After Breaking It on Purpose
· python, node, verification · 5 min read
finally does not run on SIGTERM, and a restore that ran is not a restore that worked. The suite kills real children: a try/finally harness leaves the mutated file on disk, the guarded one comes back byte for byte, and SIGKILL defeats both, so the last check lives one process out.
A Determinism Check Has to Leave the Process
· python, verification, measurement · 5 min read
Set iteration order is stable inside one Python interpreter and different in every new one, so the repeat-it-twice check reports deterministic every time. nondet probes in fresh processes instead: 9 of 9 nondeterministic fixtures caught, 0 of 10 deterministic ones falsely flagged.
Ten Packages, One Rule: A Check Must Be Able to Fail
· verification, measurement, python, node · 9 min read
Eight of the ten went up in one day, extracted from a year of measurement work in which the thing lying was usually the instrument, not the code. Each package encodes one way a green check can mean nothing, and each commits a named test that would report its own premise wrong.
What Mutation Testing Frameworks Do When a Timeout Kills Them
· mutation-testing, verification, measurement · 8 min read
I expected a kill to leave all four with a mutated file on disk. Three of them never write a mutant into your tree at all, and the fourth is left dirty by an ordinary SIGTERM, which is a gentler signal than the one I built the suite to test for.
Searching for a Program Instead of Generating One
· code-generation, measurement, llm, python · 11 min read
A primitive algebra searched from examples solved five of twenty-eight exercises, and every one already had a hand-written implementation in the same repo. A behavioural index recognised one function in forty-eight, and zero of eleven when handed other people's correct solutions to the same exercises.
An Open Redirect Guard That Rebuilt What It Rejected
· node, security, measurement · 7 min read
isSafeReturnTo("//bad.example.com") returns false, and the fallback for that rejection handed the browser the same string back, rebuilt from the request path. 0.3.0 closed one open redirect and shipped a second one for eight more days; both had been there since the first publish in 2019.
When a Test Suite Rejects a Correct Program
· verification, code-generation, measurement, llm · 14 min read
Eleven of twenty-one operations never passed my frozen checks, and I wrote that up twice as a limit of the generator. A wrapper that only re-expressed each program's own stdout repaired 90 of 102 failures, changing no algorithm.
What a Rate Limiter Reports, and What the Server Receives
· node, concurrency, measurement · 6 min read
My crawler's per-process limiter reported a perfect interval while the host received every request at once. Pointing the same server-side test at Bottleneck, the most-downloaded rate limiter on npm: six cluster workers configured for 2 requests/second delivered 14.04, and every limiter reported 495.8 ms spacing.
The Working Set That Never Saturated
· measurement, ssd, llm · 8 min read
A count table is random access, so in principle only the pages a document touches need be resident. Across eight real files, seven were still pulling in new pages after 2,700 completions. Then I measured the latency constant the verdict rested on and found it optimistic by 1.8x.
Just Train More: Measuring the Exchange Rate
· llm, training, measurement · 6 min read
The obvious objection to a zero-parameter cache beating a small transformer is that the transformer is undertrained. Sixteen times the training data moved the crossover 6.6x, which is document length times 1.60 per doubling: about as much as letting the cache read the rest of the file it was already sitting in.
When a Zero-Parameter Cache Overtakes a Transformer
· llm, measurement, statistics · 6 min read
A count table over the current document has no parameters and no training. Somewhere between 60 and 250 tokens of document it passes a 1.43M-parameter transformer, and by 1000 tokens it wins top-1 by 0.064, a 43% relative margin. Adding the transformer on top then buys 0.002.
The Last Non-Neural Candidate, and It Did Not Clear the Bar
· llm, measurement, statistics · 9 min read
The bar was a slope: keep converting extra data into accuracy after exact-context statistics saturate. A full hierarchical Pitman-Yor model with Gibbs sweeps and inferred discounts moved the intercept and left the slope alone, halving with every doubling exactly as cruder count models did.
The Embedding Table Was 72% of the Model
· llm, quantization, measurement · 6 min read
Nearly three quarters of my small transformer was a lookup table, so the dial that mattered was embedding precision, not network precision. int4 with one scale per row costs 16 KB more than one scale per tensor and recovers 2.2 of the 2.5 points per-tensor int4 loses.
Was TinyStories the Domain or the Vocabulary?
· llm, training, measurement · 6 min read
TinyStories is famous for a ~1,500-word vocabulary, so I capped mine at 1,500 and scored every arm on identical positions. It came out 4 points worse at top-1 than a 4,000-word cap, and more than a third of its output was the unknown-word token. The transferable lever was the domain, not the vocabulary.
A PDF Exporter With No PDF Library
· browser-extensions, node, verification · 6 min read
PDF wants JPEG bitstreams verbatim and zlib-wrapped deflate for everything else, and a browser already emits both. That makes the exporter 336 lines with no dependency, and moves the only place it can quietly break into the cross-reference table.
What a Browser Extension's Test Suite Cannot Reach
· browser-extensions, verification, measurement · 6 min read
130 Node assertions passed while printing could not open a dialog at all, and choosing PDF silently broke the editor handoff. Both bugs sat at a boundary: my code meeting a browser API, and one internal stage meeting another. Pure-function suites are structurally blind to those.
Where Four Backends Stop Agreeing
· compilers, verification, measurement · 5 min read
My language compiles to C, WebAssembly, ARM64 and a bytecode VM, and correctness is defined as all four producing identical output. There are exactly four places where they are allowed not to, and casting an out-of-range float produces a different result in every one of them.
The Fixpoint No Test Suite Can Check
· compilers, verification, measurement · 7 min read
My language has 94 tests that run every program through four backends, compared byte for byte. All 94 run a program the JavaScript compiler built, so none can detect a compiler that disagrees with its own output. Self-hosting found three bugs the suite could not reach.
What a Language Needs Before It Can Compile Itself
· compilers, node, measurement · 7 min read
Porting a 129-line JavaScript lexer to my own low-level language produced 355 lines. Almost none of the difference is verbosity: six specific absences account for it, one of them caused the port's only correctness bug, and one I filed as merely annoying turned out to be a correctness problem two stages later.
No Add-on Can Block Ads in This Browser
· webkit, browser-extensions, measurement, macos · 7 min read
WebKit grants an extension webRequestBlocking without honouring it, and accepts declarativeNetRequest rulesets without applying them. uBO Lite with EasyList enabled blocked 0 of 5 tracker probes, while a content rule list the browser compiled itself blocked the same probe every run.
Measuring the Wrong Process for Eight Months
· memory, swift, webkit, measurement · 10 min read
My benchmark credited each tab with one process and counted every other process as nothing. Re-measured against every process the browser actually owns, the light-pages result inverts: the ladder I published as 2.63x better than doing nothing is 1.35x worse.
When a SQL Engine Records Column Types but Never Reads Them
· sql, node, parsing · 4 min read
sql-nodejs accepts `id INT` and stores every value as a string. `WHERE age=25` matches a row and `WHERE age=25.0` matches nothing, because the comparison is string equality between two fragments of the query text.
Choosing an Unreleased API Over the One Already There
· node, api-compatibility, express · 5 min read
I removed a dependency on an unreleased router method in two lines and watched it pass CI on three Node versions, then put it back, because a public method is better code than the private field I had swapped in. I priced that choice at a few weeks and never priced it again.
What Maintaining a Forked npm Package Actually Buys
· node, security, measurement · 8 min read
I would have said "eighteen releases of upkeep". What npm audit can see is one digit in a dependency range: a caret below 1.0.0 walls off the published fix, so it reports "No fix available". What it could not see is the open redirect my fork shipped for seven years while auditing clean.
Choosing Video Frames by Content Instead of by a Clock
· measurement, llm, computer-vision, python · 5 min read
At an equal frame budget, uniform sampling missed 9 of 28 shots across five real videos and spent 12 of its 40 frames re-photographing things it had already seen. Content-aware selection missed none and wasted two; here is the measure where it still comes out level, and the one video where uniform sampling wins.
Ranking Language Models by How Well They Spot Liars
· llm, measurement, benchmarks, statistics · 11 min read
I built a benchmark to rank language models on spotting liars, then ran 522 games of Mafia through it. The same model scored -0.400 and +0.078 in one run, 52 games apart. Day 1 accusations came in at -0.005 against chance on n=2,103, and the whole thing reproduces in 11 seconds with no API keys.
An AI Capture-the-Flag Tournament: What the Scoreboard Counted
· llm, measurement, security, benchmarks · 8 min read
A preliminary report concluded that models under 3B parameters cannot chain exploits. A later tournament with much larger models put an 8B model last: the engine credited 2 of its 279 logged captures, because a per-turn marker had counted each of the 232 times it read bonus flags off its own machine.
Checking a Cost Model Against a Stranger's Config File
· llm, measurement, mixture-of-experts · 7 min read
I bounded an ambiguity in a config file at 2.4% and called it immaterial. The model's own header settled it months later, and correcting the assumption pushed a gate from 1.1% off to 3.4% off, outside the tolerance it had been passing.
Why a Small Transformer Can't Copy a Word It Hasn't Seen
· measurement, code-generation, llm, tokenization · 13 min read
An 11.9M-parameter transformer scores 83% strict on the domains it trained on and 0/8 on a single unseen noun. Three experiments to find out why, one of which corrected the diagnosis I had already written down, and a zero-parameter mechanism that does the copying 20 times out of 20.
How a Dedup Pass Deleted My Training Curriculum
· llm, measurement, fine-tuning, security · 9 min read
A fine-tuning pipeline that weighted its best examples 6x by duplicating them to 94,022 lines, followed by a dedup pass that handed the trainer 50,145, each exactly once. Two retrains were rolled back for regressing before I found that the graded curriculum had never reached the trainer at all.
Error Feedback, Gradient Compression, and Why Adam Breaks It
· llm, measurement, quantization, training, optimizers · 10 min read
Error feedback makes a biased gradient compressor unbiased over time, and under SGD it restored the full-precision trajectory to three digits. Under Adam it landed 1.9 times further from the optimum than no correction at all, and the fix I published helps just as much with no quantization in the run.
A Better FP4 Gradient Quantizer That Training Couldn't Notice
· llm, measurement, quantization, training · 9 min read
A per-block scale that cuts FP4 gradient-quantization error on 45 of 45 tensors, 14% against the published state of the art, two rented-GPU training runs that landed within noise anyway, and the measurement that explains both: the gap only shows at million-token batches, 35 to 643 times larger than anything I ran.
What a Code Completer's Eval Never Measures: Ghost Text
· llm, measurement, code-completion, python · 11 min read
A hybrid n-gram and 2.48M-parameter code completer that gets 54.6% of tokens right, tuned dial by dial against its own held-out eval. Then I scored the one feature that emits more than a single token and it came back at 0.070, because the eval hands the model a true prefix and the feature hands it its own output.
Comparing INT4 and NVFP4 Palettes on Real Gradient Tensors
· llm, quantization, training, measurement · 9 min read
A palette study on synthetic blocks said evenly-spaced INT4 beats NVIDIA's FP4 grid once a Hadamard rotation gaussianizes the data, and loses by 2.3x without it. On real gradient tensors it wins either way, 41 of 45 without the rotation, because my stand-in for unrotated data put the outliers in the wrong place.
An App Generator That Verifies Everything Except Its Parser
· measurement, code-generation, verification, llm · 14 min read
A natural-language app generator whose verification sweep is exhaustive over everything except the part that reads English. The composer scores 100% on 810 cells built from feature sets; the parser that has to produce those feature sets scores 0.70, and nothing prints that number next to the other one.
Reimplementing Enough of Kubernetes to Fool kubectl
· kubernetes, node, protocols, api-compatibility · 8 min read
Reimplementing enough of the Kubernetes API that the real kubectl can't tell the difference. The surprise wasn't the resource schemas; it was that kubectl delegates its own output formatting to the server, and that a Quantity sent as a plain string decodes as zero on the client rather than erroring.
What a Hibernated Browser Tab Actually Costs
· memory, swift, webkit, measurement · 17 min read
My simulated tab scheduler priced a hibernated tab at 32 KB; real WebKit charges 39 MB. It claimed 7.9–11.5x less memory than an unmanaged browser and delivered 1.4–2.6x. That one number, wrong by three orders of magnitude, explains the entire gap.
Predicting the Speed of a 276B Model Streamed From an SSD
· llm, measurement, macos, ssd, mixture-of-experts · 10 min read
A from-config cost model for streaming a mixture-of-experts model off an SSD, validated byte-for-byte against a runtime I didn't write. Then the SSD benchmark turned out to be measuring RAM, and the real run missed the speed prediction by 23×, for the one reason the model had declared it couldn't account for.
Rate Limiting a Crawler Across Node Cluster Workers
· concurrency, node, mongodb, crawlers · 7 min read
A per-process rate limiter under Node's cluster module enforced a perfect five-second delay while the host on the other end received a hundred requests at once. Fixing it took an atomic claim, and two more bugs on the way.