Projects
Each of these was built to find out how something works. Where a claim is measurable I've put the measurement in, including the ones that came out worse than I predicted.
nodejs-k8s125 stars · JavaScript · public
Kubernetes' core APIs reimplemented in Node. Point your real kubectl at it and it answers.
A single-process, single-node stand-in for kube-apiserver: state in MongoDB instead of etcd, "pods" as sibling Docker containers on the host. Around 55 resource kinds are routed; the ten that matter (Pod, Deployment, ReplicationController, Service, Endpoints, ConfigMap, Secret, Node, Namespace, Event ) have real lifecycle behaviour, and the rest round-trip as API objects. Init containers, exec/httpGet/tcpSocket probes, label and field selectors, watch streams, and all three patch types work.
The hard part wasn't the schemas. It was aggregated discovery, server-side table printing, the k8s\0-prefixed protobuf envelope, and the scalar coercions where a resource limit sent as a plain string decodes as zero on the client instead of erroring.
Write-up: reimplementing enough of Kubernetes to fool kubectl →
KestrelSwift · WebKit · public
A macOS browser with a memory budget for tabs. Above the budget's feasibility floor it holds 788 MB where discarding holds 802 MB, and destroys 6 tabs to its 8.
About 5,500 lines of Swift on WKWebView. Background tabs are demoted down a ladder of measured states rather than being killed outright, so the browser stays under a budget you set the way a game engine stays inside a frame budget. The rungs cost what the design said they cost, and re-measuring them changed nothing: 128 MB live, 106 MB warm, 39 MB cold, and 87 ms to bring a cold tab back to a live page.
It falsified its own design twice over. Simulation predicted 7.9–11.5× less memory than an unmanaged browser; the real engine gave about 1.43×, because the simulation priced a hibernated tab at 32 KB and WebKit charges 39 MB, since the renderer process survives and can't be terminated on request. Then the measurement turned out to be wrong as well: every per-tab figure was an attributed sum, and a process orphaned by demotion belongs to no tab and was counted as nothing. Against real process footprint the reduction is 1.24×, and "0% over budget" becomes 48% of samples over budget, because the scheduler was demoting until its own accounting said it was under the bound.
What survives is the comparison, on firmer ground than the figures it replaces: at an 800 MB budget over 10 heavy pages Kestrel holds 788 MB against discard-LRU's 802 MB, destroys 6 tabs to its 8, and is over budget 48% of the time against 62%. What does not survive is the light-pages run. By this design's own feasibility rule (39 MB per parked tab) eleven tabs need 429 MB before anything is displayed, and that run was given 150 MB; below the floor the ladder is 1.35× worse than managing nothing, since process churn costs more than the parked pages save.
Write-up: what a hibernated browser tab actually costs →
A small language with four independent backends (C, WebAssembly, ARM64, a bytecode VM) that have to agree byte for byte, and a compiler for it written in itself.
lm is statically typed with explicit memory, no closures and no garbage collector, and four backends compile the same typed AST: C, a WebAssembly binary emitted directly (no WAT and no wasm toolchain), Darwin arm64 assembly with no register allocator, and a stack VM whose arithmetic is BigInt with an explicit sign-extend after every operation. Twenty test programs and ten examples run through all four and are diffed against each other, which is what agreement on 64-bit wraparound, truncating division and float formatting is evidence of; a golden file matching one backend is not.
The disagreements it found are the point: an inverted || in the bytecode compiler, a C crash on a trailing else-less if, an arm64 shift computed from a non-power-of-two element size, and a u64 literal wasm encoded with a bit too many. Benchmarking found two more that the tests could not, because both produced correct output: the C backend was silently emitting x86_64 (so the "native" backend ran under translation), and clang was contracting a*b + c into a fused multiply-add, rounding once where lm rounds twice.
Then the compiler is rewritten in its own language, stage by stage against the JavaScript it replaces: the same token stream, the same syntax tree, the same verdict on 95 programs, and the same C byte for byte. bootstrap.sh compiles the emitter with the JavaScript compiler, again with the result, and once more with that; the last two agree across 6,883 lines. That fixpoint is the property a test suite cannot supply, because every test runs a program the JavaScript compiler built. It also found bugs the tests could not reach, including a parser that decided Name { ... } was a struct literal by whether the name began with a capital, so while R >= S { parsed one way in lm and another in JavaScript.
Where it stops is written down rather than left to be discovered: I/O is stdin and stdout and nothing else, there is no address-of operator, alloc never fails, and the README lists the four places the backends genuinely diverge (division by zero, out-of-range float casts, running off the end of memory, and printing infinities).
reckonerResearch · not a public repo
Fifty-plus experiments asking whether language-model capability really needs billions of parameters, and what today's largest models cost to run on modest hardware.
The strand with the sharpest result is a cost model for streaming a mixture-of-experts model's weights off an SSD. It was gated twice against systems I had nothing to do with: it reproduces a third-party runtime's on-disk container byte-for-byte from config alone (expert stride 1,769,472, per-layer blob 452,984,832, exact integer equality) and predicts an unrelated model's on-disk size, converted by different people with a different tool, to 1.1%.
Then it met the machine. Running a 276B-parameter model from SSD, the bytes-per-token prediction held under an independent method (kernel page-in counters against config arithmetic); the tokens-per-second prediction missed by 23×, because the model had declared that it did not account for compute and compute was what bound. On the way there, the SSD benchmark turned out to be measuring RAM.
Not under version control and holding ~119 GB of model weights, so there's no repo to link. The write-up is the artifact.
Write-up: predicting the speed of a 276B model streamed from an SSD →
Social deduction as an LLM benchmark, scored per turn and corrected for chance rather than ranked by win rate.
Language models play Mafia against each other, but nothing here ranks them by whether they won; a seven-player game is mostly variance, so win rate over any affordable batch mostly measures luck. Instead every town player makes a public accusation per statement, and because the engine knows both the ground truth and the exact probability of hitting a Mafia by chance for that turn's roster, accuracy minus that baseline gives detection lift: a signed number with a meaningful zero and 15–25 samples per game instead of one bit.
522 games across 19 models and 3,368 individual accusation records. Detection lift rises with game day (about zero on day one, then +0.13, +0.23, +0.26) and the rule-based control has not been beaten. The most useful finding is a correction to itself: the control scored +0.293 over twelve games and +0.071 pooled over 434, which is the project's own evidence that twelve games can't pin a policy down.
Concept, research questions and system design are mine; much of the implementation was AI-assisted. Findings are exploratory.
webCrawlerJavaScript · MongoDB · public
A search engine (crawler, inverted index, BM25 ranker) with no search library.
A cluster-based crawler pool scaled to available memory pulls URLs from a MongoDB frontier, parses with jsdom, and updates a hand-built inverted index ({ docId, tf, len } postings plus document frequency per term). search.js scores with a hand-implemented BM25. The tokenizer is shared between indexing and search, because those two have to agree exactly.
It also has a hand-written robots.txt parser and a politeness limiter that survives a hundred concurrent worker processes, which took three attempts, since the first two enforced a perfectly correct delay while the crawled host received a hundred simultaneous requests. The tests assert on what the crawled server received rather than what the limiter returned.
Write-up: rate limiting a crawler across Node cluster workers →
Hands a video to a language model as a measured timeline, not a pile of screenshots.
Sampling N frames from a video loses the thing that matters: a screenshot says nothing about when it happened, how long after the last one, or whether the camera moved or the subject did, so the gap gets filled with a plausible story. This turns an MP4 or MOV into timestamped intervals with shot boundaries, camera-motion vectors, on-screen text and audio activity attached, so the model reads intervals instead of guessing between samples. Python over ffmpeg and Tesseract.
No sample video ships with it, which is deliberate: the fixtures are generated, so the harness builds a clip whose contents (title cards, hard cuts, a fade from black, a pan, frozen frames, tone and silence) are known in advance and the detector's output can be checked against them rather than eyeballed.
Audits mutation harnesses for the six ways a green run can be a lie, and finds functions that already answer the question by running them rather than by reading their names.
Two questions ordinary CI does not answer, about work that already passes its tests. A green suite tells you the code did what the suite asked; it tells you nothing about whether the suite could have objected, and nothing about whether the code needed writing at all.
The first half audits mutation harnesses, the scripts that break your code on purpose to check that something notices. Six named properties, each one a way a run reports success without a test having executed. The sharpest: SIGTERM does not run finally, so a cancelled harness leaves the tree mutated and says nothing. It also checks the other half of a mutation table: every anchor string must match its target exactly once, because an anchor matching nothing is a guard that has quietly stopped being tested.
The second half decides duplication by running it. Every comparable function is probed once against one deterministic ladder of inputs, and two functions are candidates for being the same function exactly when their outcome vectors match, which makes discovery a hash bucket rather than a quadratic sweep and the decider execution rather than a name. In the tree it grew out of it paired is_wordy with _word. Both halves are standard library only in both languages, which is checkable from the two manifests, and the package is pointed at itself: 78 mutations across the Python and JavaScript halves, several of them defects it actually shipped.
An in-memory SQL engine with a hand-written parser and genuinely zero dependencies.
Parses SQL strings and answers SELECT with real column projection over its own table and row storage. "Zero dependencies" is checkable rather than aspirational: both dependencies and devDependencies are empty, and the tests run on Node's own built-in runner in CI.
Deliberately small, and the README says where it stops rather than leaving it to be discovered: one column=value equality per WHERE, one row per INSERT, no joins, updates, deletes or aggregates.
A mailroom tracker that reads a shipping label off a photo and emails the recipient.
Tesseract.js OCR over label images, wired into a Node/Express and MongoDB backend. Built for the problem of a package arriving and nobody knowing whose it is.
Contributions
Eleven merged pull requests into four projects I don't maintain.
Everything above is work I chose and assessed myself. These are the ones where someone else read the code and decided it was worth taking.
wesleytodd/express-openapi8 merged · @wesleytodd/openapi · ~25k installs/week · 146 stars
A year of feature work on a library that generates OpenAPI documents by reading an Express app's own routes: nested routers, routes passed as arrays, regex wildcard groups, configurable plugins, and SwaggerUI/Redoc configuration. Each one surfaced the next limitation.
The maintainer granted push access after these landed.
oracle/graaljs1 merged · 2,024 stars
Accept V8 --opt/--no-opt flags in the Node.js launcher, so tooling that passes them doesn't fall over on GraalVM's JavaScript engine. Two lines added, none removed: the smallest diff on this page, into the largest codebase on it.
aishek/axios-rate-limit1 merged · ~161k installs/week · 247 stars
Exposed the internal queue, so callers can see what's waiting instead of inferring it from timing.
guidesmiths/whoosh1 merged · 15 stars
Pass options through to sftp.createWriteStream, which the wrapper was dropping.
One more is written and deliberately unmerged: Express 5 support for the same library. Express 5 changed how routes can be introspected, so the fix isn't in this library at all: it needs getRoutes() from pillarjs/router#174, which hasn't shipped. It waits until that does.
Everything else on GitHub →