What a summary throws away
Worth reading first: The survey this site cannot do · What a count is worth · The sequence has a memory.
This phase measured four unrelated things: the order of a stem’s divergence angles, the side counts of a tessellated head, the size of a placement rule’s neighbourhood, and the shape of a branching junction. Two fields apart at the extremes, four separate libraries, no shared arithmetic.
They came out the same shape, and the shape is worth stating on its own because it is the part that transfers.
In each case a quantity was being reported that could not vary with the thing it was supposed to describe. Not “varied little”, not “was noisy” — was structurally incapable of carrying the information it was being read for. And in each case there was a second quantity, computable from the same data at no extra cost, that carried it.
The four
A spread is invariant to shuffling. Every statistic this collection had used on a divergence sequence — a mean, a scatter, a classified pair — gives the same answer if the angles are dealt out at random. Four phases of work on a sequence had thrown the order away. The second statistic is the autocorrelation, and it turns out to contain the parastichy number.
A mean side count is fixed by a theorem. Euler’s relation forces a planar tessellation to average six sides, so the number is the same on a golden-angle head and on a set of random points. The second statistic is the mean squared departure from six, which varies by a factor of eighty across arrangements whose means agree to a per cent.
A parameter that stopped mattering. The rule’s neighbourhood was swept sixfold to test a prediction about dilution, and nothing moved — because past four spacings the rule builds the identical lattice, internode for internode. The second statistic is a count of the placements that actually differ, which turns a flat curve from an insensitive measurement into an explanation.
A fitted exponent selected into being wrong. Fifty lopsided branching junctions from a tree built at 3 return 1.7 with a tight interval, because half the measurements come out physically impossible and dropping them keeps the biased half. The second statistic is the count of discarded junctions, which is zero on a good sample and 47% on a bad one.
What they have in common
Each of the four reported quantities is real, correctly computed, and standard. None of them is a mistake in the ordinary sense — there is no arithmetic to fix, and in three of the four cases there is no step in the analysis that could be pointed at.
What they share is that the reported number’s insensitivity was being read as a result. A scatter that agrees across three kinds of noise reads as “the kinds are alike”. A mean side count that agrees across species reads as “cellular tissue is universal”. A parameter sweep that changes nothing reads as “the model is robust”. A tight interval reads as “the data is good”.
In all four, the agreement was a property of the statistic rather than of the world.
The diagnostic
The useful output of noticing this four times is a question that can be asked before any data is collected, and it takes a minute.
What would have to be true of the sample for this number to come out differently?
If the honest answer is “nothing that could happen in this system”, the number is a check on the apparatus rather than a result. That is not useless — a mean side count that is not six means the tessellation was extracted badly, which is exactly what it is good for — but it is a different thing from evidence, and reporting it as evidence is the error.
The four cases answer the question differently and each answer is informative:
- shuffle the data — if the statistic is unchanged, the order is free information
- apply the theorem — if a theorem fixes the value, no arrangement can move it
- check the parameter binds — if the sweep changes no output at all, sweep something else
- count what was discarded — if the discard rate is high and correlated with the answer, the survivors are a selection
The one that is not like the others
Three of the four are cases where a second statistic was available and free. The branching one is different and the difference matters.
There, the second statistic — the count of impossible junctions — does not add information about the exponent. It adds information about whether the first number can be believed, which is a different service. Reporting it does not improve the estimate; it says the estimate is worthless and that different junctions should have been measured.
That distinction is worth keeping. Some second statistics are measurements that were being left on the table. Others are diagnostics that say the measurement did not happen. A survey needs both and they belong in different columns.
What this adds to the survey specification
This site has a standing specification for a study it cannot do, built up over three phases, and the phase adds four lines to it. All four are cheap and none of them is a new measurement.
Report the divergence angles in order, not only their spread. The order carries the parastichy number and, at lag one, which side of the placement decision the noise arrived on. A published table of angles is worth more than a published standard deviation and costs the same to collect.
Report the angular precision per organ, as a number obtained by remeasurement. Because reading error and one of the two candidate mechanisms are the same operation on the same numbers, a study that does not report it cannot distinguish its result from its instrument.
Report the second moment of the side distribution, and the whole distribution if there is room. The mean is a check on the tessellation, not a result.
Report attempted and retained, not retained. For any measurement with a domain — a junction that must have a thick enough parent, a run that must still have a lattice, a count that must be coprime — the discard rate is the only warning a selection bias will give.
Why nobody had taken the second statistic
The four cases were not hidden. The autocorrelation of a sequence is standard, the second moment of a side distribution is the cellular-tissue literature’s own measure, counting discards is elementary, and checking that a swept parameter binds is something anyone would agree to in the abstract.
So the interesting question is why all four were left, and the answer is the same in each case: the first statistic had an explanation attached to it.
The divergence scatter had a story — it is what a botanist can measure — and once a statistic is the one the field can obtain, asking what else is in the data stops feeling like an omission. The mean side count had a story, and it is a good one: a theorem explains it. That the theorem also drains it of evidential content is one step further than the explanation goes, and the step is not obvious while the explanation is satisfying. The neighbourhood parameter had a story about dilution among thirty neighbours. The tight interval had the best story of all, which is that it was correctly computed.
An explained number stops being interrogated. That is the mechanism, it is consistent across four unrelated cases, and it is the reason the diagnostic above is worth applying to quantities that are not puzzling. The puzzling ones get looked at anyway.
This collection’s own record bears it out. The +0.54 correlation was explained as a lattice correcting itself, and the explanation was good enough that nobody tested it where it made its strongest claim. The explanation was wrong and the number belonged to the rise.
What it costs to have found this
It is worth being clear that the four results are not equally weighty, because a synthesis essay can make a pattern look tidier than it is.
The sequence result is a genuine finding: a new instrument, a round trip against an existing counter, a control on a non-Fibonacci lattice, and a price in internodes. The second-moment result is a correction to a reporting convention, and the convention was already better in the cellular-tissue literature than in the phyllotaxis one. The neighbourhood result is a refutation of a prediction this site made itself. The junction result is a methodological warning about a study nobody has run.
Only the first produces a number about plants. The other three produce ways of not being wrong, which is a smaller thing and is the ordinary output of a phase spent checking.
The fifth case, which is this site’s own
Four cases were found by looking outward. A fifth was found by accident and belongs here because it is the same shape in the collection’s own machinery.
Nine calls in this site’s figure code drew point markers by passing a coordinate pair in the wrong form. The helper takes a point as an array and was being handed an object, so it returned nothing — and every one of those calls had been silently drawing no dots since the phase that introduced them.
Every gate stayed green. The site runs twenty-four shared checks and three of its own, and they ask whether a label fits, whether it contrasts, whether it stays inside the viewBox, whether the caption’s numbers match the figure’s. Not one of them asks whether an element that was asked for exists, because the symptom is absence and absence has nothing to inspect.
That is the same failure the site’s own instructions already record about a tick function returning an empty array for a descending axis — two figures in the fleet carried no gridlines at all for months, at full marks. The pattern recurring inside the checking apparatus, in the phase about statistics that cannot vary, is either tidy or embarrassing depending on how it is written up.
The repair is the one the diagnostic recommends: a check that can fail. The site’s gate now refuses any figure source calling that helper with an object, which is a one-line lint that would have caught all nine on the day they were written.
And what it does not license
The temptation after four cases is to generalise: that summary statistics are suspect, that fields report the wrong things, that the second moment is always where the information is.
None of that follows and the counter-examples are in this collection. The divergence scatter — a summary, invariant to shuffling, thrown away by nothing — is the quantity that establishes the site’s most robust result, that a lattice fails at about a degree and a half whichever way the noise got in. The counted pair, two integers summarising an entire arrangement, turns out to determine the whole family structure. Both are summaries and both are exactly right for what they are used for.
The difference is not that one kind of statistic is better. It is whether the quantity can vary with what it is being read for, and that is a question with a specific answer in each case rather than a disposition to hold.
What a reader should do with this
The four cases are about this subject and the shape is not, so it is worth saying what transfers and how far.
It is not a claim that fields report the wrong things. Three of the four conventions criticised here are perfectly sensible for what they were adopted for. The divergence scatter is what a botanist can measure; the mean side count is a tessellation check; a confidence interval prices what it prices. The failure in each case is a reuse — the quantity being asked a question it was not built to answer, usually years after it was adopted.
And the remedy is not more statistics. Adding a second number to every table would be its own kind of noise. The remedy is one question asked once per quantity, at the point where a number starts being used as evidence for something rather than as a description: what would have to differ for this to come out differently?
That question has a definite answer in every case here. Shuffle the sequence. Move the points. Widen the neighbourhood. Discard a different half. Where the answer is “the number would not move”, the quantity is not carrying the claim, and the second statistic is usually sitting in the same data already collected.
The habit, restated
Every figure here is generated from a stated rule and every claim is given a test it could fail. This phase adds a clause to that, and it is about the tests rather than the claims:
A test that cannot fail is not a test, and a statistic that cannot vary is not a measurement. The way to tell, in both cases, is to ask what would have to be different for the answer to change — and to keep asking it of the quantities that have been agreeing for a long time, because those are the ones where the question stopped being asked.
What is now open
The phase’s own leavings are three, and they follow from the four cases rather than from any one of them.
The second statistic on the sequence has a second statistic of its own. The spectrum contains the larger parastichy number and their sum as well as the period the readout names, and nothing in it distinguishes a family from a harmonic. A method that separated those would return the whole pair from angles alone, which is what the position counters give and this one does not.
The disorder measure should be swept against the continued fraction. The second moment of the side distribution is 0.253 at the golden angle and 0.291 half a degree away, and the reason appears to be the partial quotients rather than any distance from 137.5°. Whether that holds across many angles is one sweep and would say what the statistic is actually measuring.
And the mixture problem is still open with two statistics instead of one. The sequence now supplies a period as well as a lag-one correlation, and they have opposite requirements — one wants a disturbed plant, the other a quiet one — so they are not two readings of the same specimen. Whether a single stem can be found that yields both, and whether two numbers are enough to separate a small amount of one noise from a large amount of both, is the question this thread has now failed to close twice.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A sample that is confidently wrong — both name bias, branching exponent, measurement, sampling, selection, specimen, summary statistic, survey
- What the protractor has to be — both name autocorrelation, divergence angle, measurement, parastichy, sampling, specimen, summary statistic, survey
- A counter that sees no positions — both name autocorrelation, divergence angle, measurement, parastichy, sampling, summary statistic
- The test a plant could settle — both name autocorrelation, divergence angle, measurement, specimen, summary statistic, survey
- Which junctions say anything — both name branching exponent, measurement, sampling, specimen, summary statistic, survey
- What a quiet plant is worth — both name autocorrelation, divergence angle, measurement, specimen, survey
Named objects
A flat tag is an object no other essay names yet.
AutocorrelationBiasBranching exponentDisorderDivergence angleEuler's formulaThe range of the interactionMeasurementParastichySamplingSelectionSpecimenSummary statisticSurvey