The claims, measured

Matching instead of correcting

Two rounds of work failed on one question because every instrument's free parameter was set by the thing under test. The repair was not a better instrument or a model of the bias: it was choosing what to compare so that the confound could not vary. That move is available in four other places here, and three of them have already used it without anybody naming it.

Worth reading first: What a summary throws away.

A question was open here for two rounds and could not be closed by measuring more carefully, because the difficulty was not in the care. The width of a disorder dip is measured either by a level crossing or by an integral; the first needs a level and the second needs a limit, and both are set by how close the next rational sits. The question was whether the width depends on how close the next rational sits.

An instrument calibrated by the thing under test returns an answer that cannot be read. That is not a precision problem and no amount of resolution fixes it.

What did fix it was not an instrument at all. It was a set — seven fractions whose neighbourhoods are the same to three per cent and whose denominators differ by a factor of two and a half — so that the confound had nothing to vary over. The measurement then took an afternoon and gave a spread of one per cent where two instruments had produced contradictory orderings.

This essay is about that move, because it is available elsewhere here and because naming it makes it easier to reach for.

Five fractions with one neighbour distance and every denominator. Each member of a matched set drawn on its own stretch of the divergence axis, 0.5° either side of itself, with the nearest other rational marked. The distances are 0.3651°, 0.3529°, 0.3692°, 0.3640°, 0.3557° — a spread of 4.6% — while the denominators run 17, 20, 25, 43, 44, a factor of 2.59. That is the construction this thread needed. Every instrument for the width of a disorder dip has a free parameter set by how close the neighbour is, so a hypothesis about the neighbourhood cannot be tested by varying the neighbourhood; on this set the neighbourhood is held fixed and the arithmetic of the fraction is what varies.
Fig. 1 The construction in its second instance: five fractions with one neighbour distance and denominators from seventeen to forty-four. Nothing about the instrument changed to make this measurable; what changed is which objects were put beside each other.

The two ways of dealing with a confound

Correcting means measuring the confound, modelling its effect, and subtracting. It is the default, it is often the only option, and it has a specific failure mode: the model of the effect has to come from somewhere, and if it comes from the same machinery whose behaviour is in question then the correction cannot rescue that machinery’s evidence.

That was exactly the situation. A model of how the window’s reach biases the equivalent width would have had to be built out of the equivalent width’s own behaviour, and the quantity in dispute is what that behaviour depends on.

Matching means choosing the comparison so that the confound does not vary. It buys the same thing without a model, and it costs a search rather than an assumption. Its own failure mode is availability: sometimes no matched set exists, and sometimes matching one variable forces another to vary.

Here it was cheap. Twenty-four matched sets exist among the fractions with denominators from thirteen to sixty, and the head size — free, and previously thought of as a nuisance — could be spent to match two further parameters at the same time.

Hold the neighbourhood and the denominator stops mattering. The equivalent width of the disorder dip — the area of the deficit divided by its own depth — for seven fractions whose nearest neighbours sit at the same distance and whose denominators run from 19 to 47. Each is measured at a head size chosen so that all of them share one scaled unit, which makes a window in scaled units the same window in degrees and the same fraction of the way to the neighbour for every member. At a window of 25 the seven widths are 38.8, 38.8, 39.0, 39.0, 38.7, 39.0, 39.0 — a spread of ×1.010 across a factor of 2.47 in denominator. The lines separate as the window widens, to ×1.147 at 200, and when they do they order by denominator rather than by crowding. So the residual this thread carried was the window: hold it and there is nothing left that belongs to the fraction.
Fig. 2 What the matching bought: seven equivalent widths agreeing to a per cent at a narrow window, where an unmatched comparison had produced three different orderings at three windows.

Three places this collection already did it without saying so

The move is not new here. It is what three earlier comparisons are, and seeing them together is what makes it a method rather than a trick.

Two stems matched on scatter. The question was whether the colour of a disturbance changes what a readout reports. Disturbances of equal nominal size produce different scatters, so a comparison at equal size is a comparison at unequal disturbance-as-measured. The repair was to match the stems on the quantity a botanist would report — the divergence scatter — and vary the colour.

Seven fractions with one neighbour distance and every denominator. Each member of a matched set drawn on its own stretch of the divergence axis, 0.5° either side of itself, with the nearest other rational marked. The distances are 0.3158°, 0.3158°, 0.3117°, 0.3069°, 0.3117°, 0.3077°, 0.3064° — a spread of 3.1% — while the denominators run 19, 20, 21, 23, 33, 45, 47, a factor of 2.47. That is the construction this thread needed. Every instrument for the width of a disorder dip has a free parameter set by how close the neighbour is, so a hypothesis about the neighbourhood cannot be tested by varying the neighbourhood; on this set the neighbourhood is held fixed and the arithmetic of the fraction is what varies.
Fig. 3 The first instance of the same construction. Matching removes a confound by choosing what to compare rather than by modelling it away.

Two branches matched on rise. The question was whether a wrecked stem’s block depends on the counts of the lattice it was cut from. Comparing lattices at different rises would confound the counts with the organ spacing, the depth of the neighbourhood and the width of the front. A Lucas-seeded stem and a golden-seeded one at the same rise differ in eight seed organs and nothing else.

Seven fractions with one neighbour distance and every denominator. Each member of a matched set drawn on its own stretch of the divergence axis, 0.5° either side of itself, with the nearest other rational marked. The distances are 0.3158°, 0.3158°, 0.3117°, 0.3069°, 0.3117°, 0.3077°, 0.3064° — a spread of 3.1% — while the denominators run 19, 20, 21, 23, 33, 45, 47, a factor of 2.47. That is the construction this thread needed. Every instrument for the width of a disorder dip has a free parameter set by how close the neighbour is, so a hypothesis about the neighbourhood cannot be tested by varying the neighbourhood; on this set the neighbourhood is held fixed and the arithmetic of the fraction is what varies.
Fig. 4 The same set read over a wider span. What the span decides is how much of each member’s neighbourhood is included, and the matching is on the neighbourhood itself.

A disturbance matched on structure and stripped of history. The question was whether a comb needs errors to be re-transmitted or merely shared. The repair was a disturbance built to have the same correlation at the same contact offsets, out of fresh deviates rather than out of previous values — matched on the structure, differing in the history.

Five fractions with one neighbour distance and every denominator. Each member of a matched set drawn on its own stretch of the divergence axis, 0.45° either side of itself, with the nearest other rational marked. The distances are 0.3651°, 0.3529°, 0.3692°, 0.3640°, 0.3557° — a spread of 4.6% — while the denominators run 17, 20, 25, 43, 44, a factor of 2.59. That is the construction this thread needed. Every instrument for the width of a disorder dip has a free parameter set by how close the neighbour is, so a hypothesis about the neighbourhood cannot be tested by varying the neighbourhood; on this set the neighbourhood is held fixed and the arithmetic of the fraction is what varies.
Fig. 5 The second set at its own span. Two constructions at two neighbour distances, each internally matched, is what makes the comparison a comparison.

Each of those was arrived at locally, as the obvious thing to do about a particular objection. What the fraction result adds is the observation that they are one move, and that it is worth asking for deliberately when an instrument’s own parameter is implicated.

When to reach for it

The trigger is specific and it is not “there is a confound”. It is:

A free parameter of the instrument is set by the quantity under test.

If the confound merely correlates with the answer, correcting is fine and often better, because matching throws away data. If the confound decides how the measurement is made — where a level lands, how far an integral reaches, which window is admissible, how large a head has to be — then a correction has to model the instrument with the instrument, and the evidence it produces is worth very little.

The second trigger is that a matched set has to be findable. That is an arithmetic question and it is usually cheap to answer: enumerate the objects, compute the confound for each, sort, and look for runs. Five hundred and twenty-six fractions took under a second.

Seven fractions with one neighbour distance and every denominator. Each member of a matched set drawn on its own stretch of the divergence axis, 0.4° either side of itself, with the nearest other rational marked. The distances are 0.3158°, 0.3158°, 0.3117°, 0.3069°, 0.3117°, 0.3077°, 0.3064° — a spread of 3.1% — while the denominators run 19, 20, 21, 23, 33, 45, 47, a factor of 2.47. That is the construction this thread needed. Every instrument for the width of a disorder dip has a free parameter set by how close the neighbour is, so a hypothesis about the neighbourhood cannot be tested by varying the neighbourhood; on this set the neighbourhood is held fixed and the arithmetic of the fraction is what varies.
Fig. 6 The first set over a narrower span. Narrowing it removes the crowding at the edges, which is the confound the enumeration was searched to avoid.

Where it is not available

Three places here where the move would help and cannot be made, because being honest about the limits is what stops it becoming a slogan.

The dip’s own scale against the head size. The scale goes as q/n², so matching the scale across denominators forces the head sizes apart — which is exactly what was done, and it means the members are measured at heads from 487 to 766 organs. That is a second confound created by removing the first, and the only answer to it is the ordinary one: repeat at another common scale and check that each member’s own answer does not move.

The rise against the pair. A stem’s parastichy pair is set by its rise, so comparing pairs at one rise is only possible where two branches supply different pairs at the same rise. That is why the Lucas branch is useful and why there is no matched comparison between, say, 5/8 and 8/13 — nothing supplies both at one rise.

The organ count against everything. Almost every quantity here depends on how many organs a run has, and runs of different lengths cannot be matched on length while varying what the length is being used to measure. The answer there is the ordinary one too: hold the length fixed and vary something else.

Five fractions with one neighbour distance and every denominator. Each member of a matched set drawn on its own stretch of the divergence axis, 0.55° either side of itself, with the nearest other rational marked. The distances are 0.3651°, 0.3529°, 0.3692°, 0.3640°, 0.3557° — a spread of 4.6% — while the denominators run 17, 20, 25, 43, 44, a factor of 2.59. That is the construction this thread needed. Every instrument for the width of a disorder dip has a free parameter set by how close the neighbour is, so a hypothesis about the neighbourhood cannot be tested by varying the neighbourhood; on this set the neighbourhood is held fixed and the arithmetic of the fraction is what varies.
Fig. 7 And the second over a wider one. Six readings across two matched sets is the whole of what matching instead of correcting buys.

The cost, honestly

Matching is not free and the bill is worth itemising, because “choose what to compare” sounds like it costs nothing.

It throws away data. Five hundred and twenty-six fractions were enumerated and seven were used. Everything outside the matched set is unavailable to the comparison, and a design that could use all of them — a regression with the confound as a covariate, say — would have far more statistical power if its model of the confound could be trusted. Here it could not, which is the whole argument.

It needs the confound computed for every candidate. That is usually the easy part and it is not always. Matching stems on the depth of their neighbourhood would need the neighbourhood computed for every candidate rise, which is a run rather than a formula.

It can force a second confound. Matching the dip’s scale across denominators forces the head sizes apart by a factor of 1.57, and that has to be handled separately.

And it constrains the range. The seven fractions span a factor of 2.47 in denominator, which is what the arithmetic allows at a fixed neighbourhood distance and a denominator ceiling of sixty. A wider span needs a higher ceiling, which changes every crowding in the search, so the range is not free to be extended.

What is bought for all that is a comparison in which one specific reading is unavailable to the sceptic. That is worth a great deal when the sceptic is correct — and the sceptic here was this collection, which had already withdrawn its own claim on exactly those grounds.

The exchange rate between the two, which is calculable

The itemised bill above makes matching look like a trade — safety bought with data — and the trade has a rate. Working it out shows that in this case it was not a trade at all.

Seven fractions were used out of five hundred and twenty-six, so the matched comparison has about one seventy-fifth of the sample. Precision goes as the square root of the sample, so the matched design’s standard error is roughly eight and a half times the one a regression over all of them would have had. On the face of it that is a heavy price.

It is not a price, because the two costs are of different kinds and only one of them falls with sample size. More data shrinks variance and leaves bias exactly where it is. A regression over five hundred and twenty-six fractions with a mis-specified confound model returns a very precise wrong number, and its precision is the part that makes it dangerous: the tighter the interval, the more confidently it excludes the truth.

So the criterion for reaching for a matched set is a comparison of two quantities rather than a matter of taste. Match when the bias exceeds the standard error the full sample would have given. Below that line the confound is a nuisance and correcting for it imperfectly still leaves the answer inside its interval; above it, extra data buys confidence in the wrong place, and every additional member of the sample makes the error harder to notice rather than easier.

Applied here the line is not close. The full-sample standard error would have been small, and the bias was large enough that two instruments produced contradictory orderings — a disagreement about sign, not about size. A bias that flips an ordering is not within any plausible standard error, so the seventy-five-fold reduction in data was the cheaper of the two options by a wide margin, and the one-per-cent spread the matched set returned is the evidence that it was.

The general form is worth stating because it is the opposite of the usual instinct. A large sample is a reason to match rather than a reason not to: it drives the standard error down past biases that a small sample would have hidden inside its own noise.

The relation to a null result

A matched design changes what a null is worth, and that is most of its value.

An unmatched comparison finding no difference is ambiguous: the effect may be absent, or it may be present and cancelled by a bias running the other way. A matched comparison finding no difference has no bias left to cancel against, so the null means what it says.

That is why the fraction measurement is worth an essay of its own rather than a sentence. Seven widths agreeing to a per cent would be a weak observation in an unmatched design and a strong one here, and the difference is entirely in how the seven were chosen.

The same design also supplies its own sensitivity check for free, which an unmatched null usually cannot. Widening the window makes the same seven fractions disagree by fifteen per cent, so the instrument demonstrably distinguishes them when it is allowed to — and reports nothing where the hypothesis predicted something.

The order follows the window, so it was never the fractions'. The four fractions with a denominator of 34, ordered four ways. The left column puts them in order of how close the nearest other rational is — the crowding — with the most crowded at the top. The other three order them by the area of their dip, at windows of 50, 100, 200 scaled units. The residual claim this thread carried was that the most crowded fraction gives the widest dip, which would make all four columns the same order. They are not: the order changes between the first two windows and settles, from a window of 100 outwards, into 11/34 > 15/34 > 13/34 > 9/34 — which is not the crowding order either. A quantity that reverses when the measurement is stopped somewhere else is a property of the stopping.
Fig. 8 The unmatched version of the same measurement, for contrast. Three windows, three orderings, and no way to tell which if any is the fractions’ own.

The fourth place, which has not been done

The dek promises four places and three of them have been used. The fourth is available and outstanding, and it is worth naming so that it does not quietly become a fifth thing nobody noticed.

The ablation intervention reports which offsets are felt, and the boundary is the larger parastichy number. It has been measured at four rises on one branch and two on another, and the comparison between them confounds the pair with the rise: a coarse arrangement has a shallow neighbourhood and a wide organ spacing as well as a small pair, and every one of those could set a boundary.

The matched comparison exists. A golden-seeded stem and a Lucas-seeded one at one rise carry 5/8 and 4/7 — different pairs, one rise, one rule — and their boundaries are eight and seven. That is the matched version of the boundary result and it has been made once, in passing, in an essay about something else. What has not been done is the same comparison at two or three further rises, which is what would turn one matched pair into a matched design.

The reason it has not been done is cost rather than availability: each row is a few hundred continued placements, and the branch pairs available at a common rise are limited by where the two ladders happen to line up. It is the obvious next thing and it is written down here so that it stays obvious.

A note on what it does not fix

Matching removes one confound from one comparison. It does not make an instrument correct, it does not detect an instrument that is wrong in a way unrelated to the matched variable, and it does nothing at all about a quantity that is an artefact of the instrument’s resolution.

That last one has its own essay here and its own repair, which is to compare against a second resolution rather than to choose a better set. Different failure, different move.

Why this belongs in a collection about plants

A method essay in the middle of a collection about phyllotaxis needs a reason, and the reason is that the subject makes the failure unusually easy to fall into.

Almost everything measured here is measured on a neighbourhood. A spiral count is a statement about which organs are nearest each other. A disorder statistic is a statement about cells and their neighbours. An ablation’s boundary is the depth of a neighbourhood. And the quantity a great many of the interesting questions are about is also the neighbourhood: how crowded, how deep, how far the rule reaches.

So the instrument and the subject share a variable more often here than they would in most fields, and an instrument whose parameter is set by its subject is the exact condition under which correcting fails. A collection that keeps measuring neighbourhoods with neighbourhood-calibrated instruments should expect to need this move repeatedly, and it has needed it four times already.

What this does not say

It does not say matching is better than correcting. It is better in one situation, which is narrow and identifiable. In most situations a correction uses more of the data and is the right answer.

It does not say the three earlier comparisons were designed with this in mind. They were not; each was arrived at as the obvious response to a particular objection. What is new is noticing that they are the same response.

It does not say a matched set removes all confounds. It removes the one it matches. The fraction sets create a head-size spread by removing a scale spread, and that second confound is handled the ordinary way.

And it does not say the move is generally available. Two of the three places here where it would help most cannot use it, for arithmetic reasons stated above. A method that works when it works is worth naming and is not worth overselling.

The check that would refuse it

The claims in this essay are about method, so the checks are on the instances rather than on the principle.

A matched set has to be matched — every member’s confound within a few per cent — and has to span the thing under test by a factor of more than two. A set failing either is refused rather than measured, including a deliberately mismatched pair whose crowdings differ by a factor of fourteen.

The instrument’s other parameters have to be common as well: the same scale, the same reach towards the neighbour, the same floor on head size. Matching one parameter and letting two others vary would be a worse design than the unmatched one it replaces, because it would look controlled.

And the sensitivity half has to hold. A matched comparison reporting agreement must be shown to report disagreement somewhere — here, at a wider window — or the null is a statement about the instrument’s reach rather than about the objects.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

  • A difference forgets a drift — both name artefact, claim testing, ensemble, evidence, honest limits, identifiability, measurement, negative result, null model, summary statistic
  • A disturbance that is not passed on — both name artefact, discrimination, ensemble, evidence, honest limits, measurement, null model, summary statistic
  • The order belonged to the method — both name artefact, discrimination, evidence, honest limits, identifiability, measurement, null model, summary statistic
  • A disturbance with a memory — both name artefact, discrimination, ensemble, evidence, honest limits, measurement, null model
  • A fifth of the hop — both name claim testing, honest limits, identifiability, measurement, negative result, residual, summary statistic
  • A periodicity is not a lattice — both name artefact, discrimination, ensemble, evidence, identifiability, measurement, null model

Named objects

A flat tag is an object no other essay names yet.

ArtefactBiasClaim testingCross validationDiscriminationEnsembleEvidenceHonest limitsIdentifiabilityMeasurementNegative resultNull modelResidualSelection effectSummary statistic