I rebuild real systems, then publish where the design broke.
Kubernetes' API server, reimplemented far enough that the real kubectl can't tell. A browser tab scheduler that destroys 6 tabs where discarding destroys 8. A search engine with its own crawler, index and ranker.
I'm Seth Wheeler, a software engineer and graduate student. Day to day I work on platform and integration engineering (Node microservices, Kubernetes, CI/CD) and, increasingly, MCP servers and retrieval-augmented tooling. What keeps pulling me in is systems that have to hold a resource contract under concurrency: schedulers, protocol implementations, anything whose correct behaviour is only observable from outside the process that's supposed to guarantee it.
Start here
- Evaluating my work
- The write-ups are the best evidence of how I work: a modelling error I found in my own design and quantified, and a bug that was only observable from outside my process. Background on the about page.
- Here to read code
- nodejs-k8s is the most substantial public codebase, and the compatibility claim is testable in about a minute with a
kubectlyou already have. The projects page has the measured detail on each one. - Just browsing
- All writing, or the feed.
Projects
nodejs-k8s
Kubernetes' core APIs, reimplemented in Node. Point your real kubectl at it.
Pods, deployments, replica sets, services and jobs served over the same REST surface the real API server exposes, so the unmodified kubectl CLI and standard YAML work against it unchanged.
Kestrel
A macOS browser with a memory budget for tabs: 788 MB against discard-LRU's 802 MB, destroying 6 tabs to its 8.
~5,500 lines of Swift on WKWebView. Background tabs are demoted down a ladder of measured states to stay under budget. Re-measured against real process footprint rather than per-tab attribution, the reduction on heavy pages is 1.24× rather than the 1.43× first reported, and on a budget 2.9× below its own feasibility floor the ladder is 1.35× worse than managing nothing at all.
lambda-language
A small language whose four independent backends (C, WebAssembly, ARM64, a bytecode VM) have to agree byte for byte.
Twenty test programs and ten examples run through all four code generators and diffed against each other, which caught an inverted || in the bytecode compiler and a u64 literal wasm encoded with a bit too many. The compiler is then rewritten in the language itself, stage by stage against the JavaScript it replaces, and bootstraps to a fixpoint: the last two stages are byte-identical across 6,883 lines.
social-deduction-bench
Social deduction as an LLM benchmark: scored per turn and corrected for chance, not ranked by win rate.
522 games across 19 models, measuring deception and deception-detection separately, with a rule-based control that nothing has beaten yet.
reckoner (research)
Fifty-plus experiments on whether language-model capability really needs billions of parameters.
A mixture-of-experts streaming cost model validated byte-for-byte against a runtime I didn't write, and to 1.1% against an unrelated model; then a 276B model run from SSD whose bytes-per-token prediction held while its speed prediction missed by 23×, for a limitation declared in advance. Not a public repo; the write-up is the artifact.
video-timeline
Hands a video to a language model as a measured timeline, not a pile of screenshots.
Turns an MP4 or MOV into timestamped intervals (shot boundaries, camera-motion vectors, on-screen text, audio activity) so the model reasons about what happened between frames instead of inventing it. Python over ffmpeg and Tesseract; no sample video ships, because the test fixtures are generated with known contents.
assay-checks
Finds duplicate functions by executing them, not by reading their names.
Two questions ordinary CI does not ask about code that already passes: could those tests have failed, and does the tree already answer this? Candidate pairs are functions whose outcome vectors match across one deterministic ladder of inputs, so the decider is execution rather than text: it paired is_wordy with _word, which no name-based detector puts together. The other half audits mutation harnesses against six named ways a run can report success without a test having run. Python and JavaScript, zero dependencies in both, and pointed at itself with 78 mutations.
sql-nodejs
An in-memory SQL engine on npm with a hand-written parser and genuinely zero dependencies.
Parses SQL strings and answers SELECT with real column projection over its own table and row storage. Deliberately small (one equality per WHERE, no joins or aggregates) and the limits are listed in the README rather than discovered.
webCrawler
A search engine (crawler, inverted index, BM25 ranker) with no Elasticsearch.
A hand-written robots.txt parser and a politeness limiter that survives 100 concurrent worker processes, asserted against what the crawled server actually received rather than what the limiter reported.
Postmastr-Backend
A mailroom tracker that reads shipping labels off a photo and emails the recipient.
Tesseract.js OCR over label images, wired into a Node/Express and MongoDB backend.
Writing
- Pinning What a Function Answers Across Rewrites and Time
A pin on a nondeterministic function fails at random on the next machine, and the blame lands on the pinning tool. Every candidate goes through nondet first; the committed control turns the gate off, changes not one byte of the tree, and asserts the check goes red anyway.
- Finding Duplicate Functions by Executing Them, Not Reading Them
Two functions agreed on every rung of the input ladder and were not the same function; one character, ½, turned same into differs with a witness. The guard exists because agreement is cheap: constants, projections and copies agree with everything, and each mistake was made before it was guarded.
- Measuring How Cost Scales by Counting Instead of Timing
Insertion sort on seed 17 at n=64 performs reads 3812 and writes 1848, the same integers in Python and JavaScript, because the generator, the iteration rule and the subscript rule are part of the contract. The same parity suite caught its own dependency formatting one number two ways above a million.
- A Curve Fitter That Refuses to Answer
A tool that always produced a constant would be useless and would still pass every test that checks it produces one. 0.2.0 fixed a defect the suite had pinned: with error bars from scatter alone, a quantity measured exactly was reported as one that could not be determined.
- A Check With a Zero Denominator Reports Clean
A JUnit file reading tests="50" skipped="50" is a green run of nothing wearing a total, so every parser here returns total and executed separately and the floor applies to executed. Thirteen mutations were applied to the source; the two that survived their first run were worth more than the eleven that were caught.