One rise per rung is a sample
Worth reading first: The survey this site cannot do · Counting the spirals · A head is a set of points.
Every census on this site is built the same way. Choose a set of lattices that spans the pairs and the branches worth covering, find a rise that produces each pair, grow the stem, do the experiment, record the row. It is a sensible design and it has answered the question it was designed for many times.
It also has a property nobody stated: within such a census, the counted pair and everything correlated with the rise move together. One rise per pair means one settled divergence per pair, one step ordering per pair, one front depth per pair. Any rule later scored across those rows inherits all of it.
This essay is about what that costs, using two results from this collection where the cost has now been measured.
It is worth being clear at the outset that this is not a general complaint about sampling. Every experiment holds something fixed, and holding things fixed is what makes a comparison a comparison. The specific problem here is narrower and has a name: the quantity held fixed is not independent of the quantity varied. Pair and rise are chosen together, so a census that varies the pair across its rows is also varying the rise across them, in a pattern nobody selected and nobody recorded. Any reading sensitive to the rise is then being scored against a confound rather than against a control.
The mechanism
A rung is a range of rises over which a counter returns one pair. Inside it, the settled divergence slides, the two contact steps can change places, and the front deepens. None of that is visible to a counter, which is what makes it a rung.
So when a census picks one rise per rung, it is not picking a representative point on a plateau. It is picking one point from a range across which several quantities vary, and fixing all of them at once by that single choice.
The word plateau is doing the damage, and it is worth replacing. A rung is a plateau in the counted pair and in nothing else — the geometry underneath slides continuously through it, and the pair is a step function laid over a smooth one. Reading “rung” as “a set of equivalent stems” is the mistake, and the vocabulary invites it. The finest sweep this collection has run puts twenty-three rises inside one rung and finds the divergence moving through all of them without a single flat stretch.
The consequence is precise. Any reading whose truth depends on one of those quantities will be scored, across the whole census, at one value of it per pair — and its score will be a joint statement about the reading and about which rises were chosen.
The size of the effect is not small. The 5/8 rung alone spans rises from eighteen thousandths to seven, over which the settled divergence moves by more than a degree, the step ordering reverses, and the number of wrecking offsets grows five-fold. A census samples one point from that. Two censuses built by different people, both entirely defensible, could pick 0.016 and 0.008 and disagree about a reading’s score without either being wrong about anything.
Two results that moved
The shortest hop. Scored across the census, the reading that a wrecked stem keeps its shortest hop is right at twelve of twenty-nine — a refutation. Scored along one rung, where the pair is held and the ordering reverses, it is right at sixteen of thirty-one, which is a coin flip. The reading is dead either way; the number was reporting the sampling.
The offset rule. Scored across the census it is right at twenty-five of thirty, and its two clauses look like halves of one statement. That reading also had a single anomaly — one pair of runs agreeing on pair and offset and disagreeing on the answer — which is exactly what a confound produces when it is almost but not quite hidden. Sweeping the confounded variable turned the anomaly into a reproducible effect, which is the usual fate of a lone outlier in a census with a variable nobody moved. Swept along a rung they come apart: the first clause applies at every rise, the second has nothing to apply to until the front is deep enough, and one offset changes its answer between the coarse end of a rung and the fine one.
What this is not
It is not a claim that the censuses were badly designed, and the distinction matters because the fix is different in the two cases.
A badly designed census would have chosen rises that produced the answer. These were chosen to produce the pairs, before any of the readings existed, and the rises attached were whatever the pair required. Nothing about that is careless, and the ordering matters: the census predates every reading scored on it, so there is no possibility of the rises having been selected, consciously or otherwise, to favour one. What is wrong is not the choosing but the reuse — a table built to answer one question, later asked a different one it was never arranged to answer.
The general form is that a sample assembled to vary one quantity holds others fixed by construction, and a rule fitted afterwards inherits the constancy as an untested assumption. The census is not wrong; it is answering a different question from the one later asked of it.
What to do instead
Not build bigger censuses. A census twice the size, still at one rise per rung, has the same defect and twice the confidence in it.
The fix is to notice which quantity a reading depends on and vary that on purpose. For a reading about step lengths, sweep the rung until the ordering reverses. For a reading about offsets, sweep until the front changes depth. For a reading about contact scales, sweep the rise until the contact numbers move.
There is a cheaper half-measure that is worth naming because it is often enough. A census does not have to become a sweep to stop being confounded: it only has to sample two rises per rung rather than one, at opposite ends. That doubles the cost, breaks the correlation between pair and rise, and turns any reading whose score differs between the two into a reading that has announced its own dependence. It would not have found the row that changes hands — that took twelve rises — but it would have found that something was wrong.
That is more expensive per reading and much cheaper per conclusion, because a reading scored over its own variable either survives or does not, and either way the number means something.
The cost is real and should not be understated. A rung sweep at a thousandth is twelve settled stems plus a cut stem per offset per rise, each grown to settle again and compared against a control — where a census row is one settled stem and its cuts. Sweeping one rung costs roughly what a fifth of a census costs, and there are more rungs than the census has rows. That is why the answer is not “sweep everything” but “sweep the quantity the reading is about”, which is one rung per reading rather than every rung for every reading.
How to tell whether a reading is exposed
There is a test, and it takes a minute rather than a sweep.
Write the reading down and ask what it would have to be given to be evaluated on a single stem. If the answer mentions only the counted pair, the offset, or which family is which, the reading is safe: those are all things a census row records and the rise cannot alter. If the answer mentions a length, a distance, a divergence, a depth, or anything measured in organs or degrees, the reading is exposed, because every one of those moves inside a rung.
By that test the shortest-hop reading was exposed from the day it was written — it is stated in step lengths — and the offset rule is exposed through its second clause, which is stated over a front. The reading about which family lost a member is safe, because divisibility of an offset by a counted number involves nothing the rise can touch.
The test is crude and it will occasionally flag something safe. It costs nothing and it would have flagged both of the results this essay is about.
The crude test says which readings could be exposed. There is a cheaper way to find out which ones actually are, and it needs no new stems.
Every row of every census here was grown at a stated rise, and a swept rung has stated ends. So each row has a number nobody has computed: where inside its own rung that rise sat, as a fraction — nought at the coarse end, one at the fine end. The 5/8 rows in the ablation census sit at 0.013 on a rung running from 0.018 to 0.007, a little under half way along. Nothing there is new information. It is arithmetic on numbers already recorded, and it turns a census that varies one quantity into a census that varies two.
With that column in hand, every reading already scored can be re-scored against it. If a reading’s failures cluster at one end of the rungs, it is exposed, and the column says in which direction. If its failures are scattered across the fraction, then the sampling is not what is wrong with it. Neither answer needs a single new stem, and the second is worth the more of the two, because it turns a caveat that currently applies to four censuses into one that applies to whichever of them the existing rows say it applies to.
The reason it was not available before is that a rung had no ends. A rung was a label attached to a pair, and a label has no inside; only sweeping one at a thousandth produced the two rises that bound it, and only then did where in the rung become a quantity at all. That is the ordinary way a confound becomes measurable — not by collecting more rows, but by finding the axis the rows already collected were spread along.
Two rungs have ends so far — the 5/8 rung swept at a thousandth, and the 3/5 rung at a tenth of that — so the column can be computed today for every row that fell inside either of them, and is a promise for the rest.
What is still scored the old way
Being specific is the point of writing this down, so: the ablation census, the two-organ table, the wreck destination list and the jugacy census are all one rise per rung, and every reading scored over any of them carries this caveat until it is swept.
The jugacy census deserves a particular mention, because the reading scored on it is about the arithmetic of the counted pair rather than about the geometry, and that is exactly the kind of reading this caveat does not reach. Whether a pair’s numbers share a factor is a property of the pair, and a census that fixes one rise per pair fixes nothing relevant to it. Sorting readings into those the sampling can touch and those it cannot is most of the work of applying any of this.
Two results are not affected and it is worth saying which, because a caveat that applies to everything is a caveat nobody applies. That the survivor is one of the two contact families is a statement about which lags exist, and the pair fixes those. That a wrecked stem keeps exactly one hop rigid is a statement about the motif it settles into, measured against its own control.
Where this leaves the collection
With a rule about its own method, which is the kind of result that is worth more than a figure and is harder to notice.
Every future census here should record the rise beside the pair, and every reading scored over one should say which of the two it depends on. Neither is expensive. Both would have caught this years of essays ago, and the reason neither was done is that the question the censuses were built for genuinely did not need them.
The uncomfortable part is that this was findable at any point. Nothing in the sweep that exposed it needed a new instrument, a new library or a new idea — only the decision to move a parameter that every table already recorded and no reading had ever varied. What made it invisible was that the parameter was settled when the method was chosen rather than argued about anywhere in the results, and nobody reads a method twice. The same shape has already cost this collection a result in a different thread, where a window was mistaken for a neighbourhood for the same reason: a quantity was named once, early, and then quoted rather than measured.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A stem too fine to settle — both name counting blind, claim testing, control, honest limits, lattice, measurement, negative result, parastichy pair, rise, rung, underdetermination
- One offset, two answers — both name claim testing, control, falsifiability, honest limits, lattice, measurement, negative result, parastichy pair, rise, rung, underdetermination
- The organ that was nobody's neighbour — both name claim testing, control, falsifiability, honest limits, lattice, measurement, negative result, parastichy pair, rung, underdetermination
- Two accounts of one number — both name counting blind, claim testing, honest limits, lattice, measurement, negative result, parastichy pair, rise, rung, underdetermination
- The shallower front turns over — both name claim testing, honest limits, lattice, measurement, negative result, parastichy pair, rise, rung, sample size
- Three organs and no mirror — both name counting blind, control, falsifiability, honest limits, lattice, measurement, parastichy pair, rise, rung
Named objects
A flat tag is an object no other essay names yet.
Counting blindClaim testingControlFalsifiabilityHonest limitsLatticeMeasurementNegative resultParastichy pairRiseRungSample sizeSelectionSummary statisticUnderdetermination