Tagged python
13 write-ups.
Pinning What a Function Answers Across Rewrites and Time
A pin on a nondeterministic function fails at random on the next machine, and the blame lands on the pinning tool. Every candidate goes through nondet first; the committed control turns the gate off, changes not one byte of the tree, and asserts the check goes red anyway.
Finding Duplicate Functions by Executing Them, Not Reading Them
Two functions agreed on every rung of the input ladder and were not the same function; one character, ½, turned same into differs with a witness. The guard exists because agreement is cheap: constants, projections and copies agree with everything, and each mistake was made before it was guarded.
Measuring How Cost Scales by Counting Instead of Timing
Insertion sort on seed 17 at n=64 performs reads 3812 and writes 1848, the same integers in Python and JavaScript, because the generator, the iteration rule and the subscript rule are part of the contract. The same parity suite caught its own dependency formatting one number two ways above a million.
A Curve Fitter That Refuses to Answer
A tool that always produced a constant would be useless and would still pass every test that checks it produces one. 0.2.0 fixed a defect the suite had pinned: with error bars from scatter alone, a quantity measured exactly was reported as one that could not be determined.
A Check With a Zero Denominator Reports Clean
A JUnit file reading tests="50" skipped="50" is a green run of nothing wearing a total, so every parser here returns total and executed separately and the floor applies to executed. Thirteen mutations were applied to the source; the two that survived their first run were worth more than the eleven that were caught.
Breaking CI Guards on Purpose to Prove They Can Fail
A fix I believed in was proved a no-op by the tool auditing itself: mtime invalidation has one-second granularity and the edit cycle is milliseconds, so stale bytecode looks fresh either way. PYTHONDONTWRITEBYTECODE is the guard actually holding; the surviving mutation corrected the README.
An Exit Code Cannot Say Whether Anything Happened
go test on a tree with no test files prints [no test files] and exits 0, measured rather than assumed. didrun wraps a command, demands evidence it did something, and returns four states instead of one number; the committed control asserts that 0 passed satisfying the expected pattern still scores as did-not-run.
Proving the Tree Came Back After Breaking It on Purpose
finally does not run on SIGTERM, and a restore that ran is not a restore that worked. The suite kills real children: a try/finally harness leaves the mutated file on disk, the guarded one comes back byte for byte, and SIGKILL defeats both, so the last check lives one process out.
A Determinism Check Has to Leave the Process
Set iteration order is stable inside one Python interpreter and different in every new one, so the repeat-it-twice check reports deterministic every time. nondet probes in fresh processes instead: 9 of 9 nondeterministic fixtures caught, 0 of 10 deterministic ones falsely flagged.
Ten Packages, One Rule: A Check Must Be Able to Fail
Eight of the ten went up in one day, extracted from a year of measurement work in which the thing lying was usually the instrument, not the code. Each package encodes one way a green check can mean nothing, and each commits a named test that would report its own premise wrong.
Searching for a Program Instead of Generating One
A primitive algebra searched from examples solved five of twenty-eight exercises, and every one already had a hand-written implementation in the same repo. A behavioural index recognised one function in forty-eight, and zero of eleven when handed other people's correct solutions to the same exercises.
Choosing Video Frames by Content Instead of by a Clock
At an equal frame budget, uniform sampling missed 9 of 28 shots across five real videos and spent 12 of its 40 frames re-photographing things it had already seen. Content-aware selection missed none and wasted two; here is the measure where it still comes out level, and the one video where uniform sampling wins.
What a Code Completer's Eval Never Measures: Ghost Text
A hybrid n-gram and 2.48M-parameter code completer that gets 54.6% of tokens right, tuned dial by dial against its own held-out eval. Then I scored the one feature that emits more than a single token and it came back at 0.070, because the eval hands the model a true prefix and the feature hands it its own output.