Matching instead of correcting
Worth reading first: What a summary throws away.
A question was open here for two rounds and could not be closed by measuring more carefully, because the difficulty was not in the care. The width of a disorder dip is measured either by a level crossing or by an integral; the first needs a level and the second needs a limit, and both are set by how close the next rational sits. The question was whether the width depends on how close the next rational sits.
An instrument calibrated by the thing under test returns an answer that cannot be read. That is not a precision problem and no amount of resolution fixes it.
What did fix it was not an instrument at all. It was a set — seven fractions whose neighbourhoods are the same to three per cent and whose denominators differ by a factor of two and a half — so that the confound had nothing to vary over. The measurement then took an afternoon and gave a spread of one per cent where two instruments had produced contradictory orderings.
This essay is about that move, because it is available elsewhere here and because naming it makes it easier to reach for.
The two ways of dealing with a confound
Correcting means measuring the confound, modelling its effect, and subtracting. It is the default, it is often the only option, and it has a specific failure mode: the model of the effect has to come from somewhere, and if it comes from the same machinery whose behaviour is in question then the correction cannot rescue that machinery’s evidence.
That was exactly the situation. A model of how the window’s reach biases the equivalent width would have had to be built out of the equivalent width’s own behaviour, and the quantity in dispute is what that behaviour depends on.
Matching means choosing the comparison so that the confound does not vary. It buys the same thing without a model, and it costs a search rather than an assumption. Its own failure mode is availability: sometimes no matched set exists, and sometimes matching one variable forces another to vary.
Here it was cheap. Twenty-four matched sets exist among the fractions with denominators from thirteen to sixty, and the head size — free, and previously thought of as a nuisance — could be spent to match two further parameters at the same time.
Three places this collection already did it without saying so
The move is not new here. It is what three earlier comparisons are, and seeing them together is what makes it a method rather than a trick.
Two stems matched on scatter. The question was whether the colour of a disturbance changes what a readout reports. Disturbances of equal nominal size produce different scatters, so a comparison at equal size is a comparison at unequal disturbance-as-measured. The repair was to match the stems on the quantity a botanist would report — the divergence scatter — and vary the colour.
Two branches matched on rise. The question was whether a wrecked stem’s block depends on the counts of the lattice it was cut from. Comparing lattices at different rises would confound the counts with the organ spacing, the depth of the neighbourhood and the width of the front. A Lucas-seeded stem and a golden-seeded one at the same rise differ in eight seed organs and nothing else.
A disturbance matched on structure and stripped of history. The question was whether a comb needs errors to be re-transmitted or merely shared. The repair was a disturbance built to have the same correlation at the same contact offsets, out of fresh deviates rather than out of previous values — matched on the structure, differing in the history.
Each of those was arrived at locally, as the obvious thing to do about a particular objection. What the fraction result adds is the observation that they are one move, and that it is worth asking for deliberately when an instrument’s own parameter is implicated.
When to reach for it
The trigger is specific and it is not “there is a confound”. It is:
A free parameter of the instrument is set by the quantity under test.
If the confound merely correlates with the answer, correcting is fine and often better, because matching throws away data. If the confound decides how the measurement is made — where a level lands, how far an integral reaches, which window is admissible, how large a head has to be — then a correction has to model the instrument with the instrument, and the evidence it produces is worth very little.
The second trigger is that a matched set has to be findable. That is an arithmetic question and it is usually cheap to answer: enumerate the objects, compute the confound for each, sort, and look for runs. Five hundred and twenty-six fractions took under a second.
Where it is not available
Three places here where the move would help and cannot be made, because being honest about the limits is what stops it becoming a slogan.
The dip’s own scale against the head size. The scale goes as q/n², so matching the scale across denominators forces the head sizes apart — which is exactly what was done, and it means the members are measured at heads from 487 to 766 organs. That is a second confound created by removing the first, and the only answer to it is the ordinary one: repeat at another common scale and check that each member’s own answer does not move.
The rise against the pair. A stem’s parastichy pair is set by its rise, so comparing pairs at one rise is only possible where two branches supply different pairs at the same rise. That is why the Lucas branch is useful and why there is no matched comparison between, say, 5/8 and 8/13 — nothing supplies both at one rise.
The organ count against everything. Almost every quantity here depends on how many organs a run has, and runs of different lengths cannot be matched on length while varying what the length is being used to measure. The answer there is the ordinary one too: hold the length fixed and vary something else.
The cost, honestly
Matching is not free and the bill is worth itemising, because “choose what to compare” sounds like it costs nothing.
It throws away data. Five hundred and twenty-six fractions were enumerated and seven were used. Everything outside the matched set is unavailable to the comparison, and a design that could use all of them — a regression with the confound as a covariate, say — would have far more statistical power if its model of the confound could be trusted. Here it could not, which is the whole argument.
It needs the confound computed for every candidate. That is usually the easy part and it is not always. Matching stems on the depth of their neighbourhood would need the neighbourhood computed for every candidate rise, which is a run rather than a formula.
It can force a second confound. Matching the dip’s scale across denominators forces the head sizes apart by a factor of 1.57, and that has to be handled separately.
And it constrains the range. The seven fractions span a factor of 2.47 in denominator, which is what the arithmetic allows at a fixed neighbourhood distance and a denominator ceiling of sixty. A wider span needs a higher ceiling, which changes every crowding in the search, so the range is not free to be extended.
What is bought for all that is a comparison in which one specific reading is unavailable to the sceptic. That is worth a great deal when the sceptic is correct — and the sceptic here was this collection, which had already withdrawn its own claim on exactly those grounds.
The relation to a null result
A matched design changes what a null is worth, and that is most of its value.
An unmatched comparison finding no difference is ambiguous: the effect may be absent, or it may be present and cancelled by a bias running the other way. A matched comparison finding no difference has no bias left to cancel against, so the null means what it says.
That is why the fraction measurement is worth an essay of its own rather than a sentence. Seven widths agreeing to a per cent would be a weak observation in an unmatched design and a strong one here, and the difference is entirely in how the seven were chosen.
The same design also supplies its own sensitivity check for free, which an unmatched null usually cannot. Widening the window makes the same seven fractions disagree by fifteen per cent, so the instrument demonstrably distinguishes them when it is allowed to — and reports nothing where the hypothesis predicted something.
The fourth place, which has not been done
The dek promises four places and three of them have been used. The fourth is available and outstanding, and it is worth naming so that it does not quietly become a fifth thing nobody noticed.
The ablation intervention reports which offsets are felt, and the boundary is the larger parastichy number. It has been measured at four rises on one branch and two on another, and the comparison between them confounds the pair with the rise: a coarse arrangement has a shallow neighbourhood and a wide organ spacing as well as a small pair, and every one of those could set a boundary.
The matched comparison exists. A golden-seeded stem and a Lucas-seeded one at one rise carry 5/8 and 4/7 — different pairs, one rise, one rule — and their boundaries are eight and seven. That is the matched version of the boundary result and it has been made once, in passing, in an essay about something else. What has not been done is the same comparison at two or three further rises, which is what would turn one matched pair into a matched design.
The reason it has not been done is cost rather than availability: each row is a few hundred continued placements, and the branch pairs available at a common rise are limited by where the two ladders happen to line up. It is the obvious next thing and it is written down here so that it stays obvious.
A note on what it does not fix
Matching removes one confound from one comparison. It does not make an instrument correct, it does not detect an instrument that is wrong in a way unrelated to the matched variable, and it does nothing at all about a quantity that is an artefact of the instrument’s resolution.
That last one has its own essay here and its own repair, which is to compare against a second resolution rather than to choose a better set. Different failure, different move.
Why this belongs in a collection about plants
A method essay in the middle of a collection about phyllotaxis needs a reason, and the reason is that the subject makes the failure unusually easy to fall into.
Almost everything measured here is measured on a neighbourhood. A spiral count is a statement about which organs are nearest each other. A disorder statistic is a statement about cells and their neighbours. An ablation’s boundary is the depth of a neighbourhood. And the quantity a great many of the interesting questions are about is also the neighbourhood: how crowded, how deep, how far the rule reaches.
So the instrument and the subject share a variable more often here than they would in most fields, and an instrument whose parameter is set by its subject is the exact condition under which correcting fails. A collection that keeps measuring neighbourhoods with neighbourhood-calibrated instruments should expect to need this move repeatedly, and it has needed it four times already.
What this does not say
It does not say matching is better than correcting. It is better in one situation, which is narrow and identifiable. In most situations a correction uses more of the data and is the right answer.
It does not say the three earlier comparisons were designed with this in mind. They were not; each was arrived at as the obvious response to a particular objection. What is new is noticing that they are the same response.
It does not say a matched set removes all confounds. It removes the one it matches. The fraction sets create a head-size spread by removing a scale spread, and that second confound is handled the ordinary way.
And it does not say the move is generally available. Two of the three places here where it would help most cannot use it, for arithmetic reasons stated above. A method that works when it works is worth naming and is not worth overselling.
The check that would refuse it
The claims in this essay are about method, so the checks are on the instances rather than on the principle.
A matched set has to be matched — every member’s confound within a few per cent — and has to span the thing under test by a factor of more than two. A set failing either is refused rather than measured, including a deliberately mismatched pair whose crowdings differ by a factor of fourteen.
The instrument’s other parameters have to be common as well: the same scale, the same reach towards the neighbour, the same floor on head size. Matching one parameter and letting two others vary would be a worse design than the unmatched one it replaces, because it would look controlled.
And the sensitivity half has to hold. A matched comparison reporting agreement must be shown to report disagreement somewhere — here, at a wider window — or the null is a statement about the instrument’s reach rather than about the objects.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A difference forgets a drift — both name artefact, claim testing, ensemble, evidence, honest limits, identifiability, measurement, negative result, null model, summary statistic
- A disturbance that is not passed on — both name artefact, discrimination, ensemble, evidence, honest limits, measurement, null model, summary statistic
- The order belonged to the method — both name artefact, discrimination, evidence, honest limits, identifiability, measurement, null model, summary statistic
- A disturbance with a memory — both name artefact, discrimination, ensemble, evidence, honest limits, measurement, null model
- A periodicity is not a lattice — both name artefact, discrimination, ensemble, evidence, identifiability, measurement, null model
- The band decides the answer — both name artefact, discrimination, evidence, honest limits, identifiability, measurement, residual
Named objects
A flat tag is an object no other essay names yet.
ArtefactBiasClaim testingCross validationDiscriminationEnsembleEvidenceHonest limitsIdentifiabilityMeasurementNegative resultNull modelResidualSelection effectSummary statistic