A sample that is confidently wrong
Worth reading first: Fitting the exponent · How many plants would it take · The survey this site cannot do.
The arithmetic of what a branching junction can say about the exponent gives a factor of two thousand between an even fork and a twig. The natural expectation from that is a familiar one: measure the uninformative junctions and the answer comes back uncertain — a wide interval, honestly reported, saying that the data cannot choose.
That is not what happens, and this essay is what does.
The experiment
Build a synthetic tree at an exponent of exactly 3. Put two per cent of measurement error on every radius — parent and both daughters, independently, which is what measuring three branches with the same calipers does.
Take fifty junctions from the informative end, with daughter ratios between 0.6 and
- Fit the exponent. Bootstrap for an interval.
Then take fifty from the other end, ratios between 0.05 and 0.15, and do exactly the same.
Even forks: 3.01, interval [2.91, 3.10]. Correct, tight, excludes Da Vinci’s rule.
Twigs: 1.69, interval [1.56, 1.79]. Wrong by more than the entire distance between the two hypotheses being tested — and the interval excludes 3, excludes 2, and is only a fifth wider than the good sample’s.
Reported without the arithmetic in front of it, that is a striking result about branching networks. It is an artefact, and the tree it came from obeys Murray’s law exactly.
Where the bias comes from
Not from the fitter. A per-junction solve — asking, of each junction alone, what exponent satisfies — gives a median of 1.64 on the same data. Two estimators with nothing in common give the same wrong answer, so the cause is in the sample.
It is this. A parent is genuinely thicker than its larger daughter by a margin that collapses as the junction becomes lopsided.
| daughter ratio | true margin | noise on the ratio |
|---|---|---|
| 1.0 | 26% | 2.8% |
| 0.6 | 6.7% | 2.8% |
| 0.3 | 0.89% | 2.8% |
| 0.15 | 0.11% | 2.8% |
| 0.05 | 0.004% | 2.8% |
At an even fork the parent is a quarter thicker than its daughter and the noise is under three per cent — a nine-to-one margin, and no measurement ever comes out the wrong way round. At a daughter ratio of 0.15 the margin is a twenty-fifth of the noise.
So close to half the measured twig junctions come out with the parent thinner than a branch it carries. Measured: 112 of 240, which is 47%. At the informative end it is 0 of 240.
Why an impossible junction biases what is left
A junction whose parent is thinner than its larger daughter has no exponent at all. The equation has no solution in — the left side is smaller than the first term on the right at every exponent, and no power fixes that.
Every estimator therefore discards those junctions. The per-junction solve returns nothing for them. The least-squares fit gives them residuals that grow with , so they push the fitted exponent down and are effectively excluded at any plausible value.
Nobody chose to discard them. There is no line in the analysis that says drop the impossible ones; it is what the arithmetic does when handed an observation outside its domain.
And the half that survives is not a random half. It is precisely the half where the noise happened to inflate the parent, because that is what made them admissible. A sample selected on “the parent came out thick enough” is a sample of junctions whose parents are overstated, and an overstated parent implies a smaller exponent.
The bias is therefore in a definite direction and it is large: 3 measured as 1.7.
The margin, and why it collapses so fast
The table’s middle column deserves its own derivation, because the speed of the collapse is the whole mechanism and it is not obvious from the rule.
The parent’s radius relative to its larger daughter is under Murray’s law. For small that is approximately , so the margin above 1 goes as the cube of the daughter ratio.
Halve the asymmetry and the margin falls eightfold. From to it falls from 0.89% to 0.11%; from there to 0.075 it would fall to 0.014%. Meanwhile the noise on the measured ratio does not move at all — it is whatever the junction looks like.
So the ratio of signal to noise falls as , and the crossing — where the margin equals the noise — happens at about for a two-per-cent instrument. Below that, impossible junctions start appearing; well below it, half of them are.
That gives the threshold a physical meaning rather than an empirical one. The junctions worth measuring and the junctions that produce impossible readings are the same set, seen from two sides, because both are governed by how much thicker the parent genuinely is. The previous essay’s advice to measure comparable forks and this essay’s warning about lopsided ones are one statement.
Under Da Vinci’s rule the margin goes as rather than , so the collapse is slower and the crossing is at a smaller — which means the bias is milder for a network that really obeys the square law. That asymmetry is itself a small confound and it runs in the direction of making squares easier to confirm.
Why nothing in the output says so
This is the part that would do the damage in a real study.
The surviving sample looks clean. Thirty-two junctions instead of sixty, all of them physically sensible, all of them fitting a consistent exponent. The bootstrap interval is computed from those thirty-two and is honest about them — it correctly reports that this sample determines 1.69 to within about ±0.12.
What it cannot report is that the sample is not the population. The interval measures sampling variation and the error here is not sampling variation; it is a systematic displacement introduced before the interval was computed.
A confidence interval prices precision and is silent about selection. That is not a subtlety of this case; it is what a confidence interval is, and it is why the number of discarded junctions is the diagnostic rather than the width of the interval.
The diagnostic
It is one number and it should be reported.
The share of measured junctions that were physically impossible. A parent thinner than its larger daughter is not a borderline case or a judgement call — it is an observation no tree can produce, so every one of them is a measurement error, and the count of them is a direct measure of how much the noise exceeds the signal.
At the informative end it is zero, and a zero says the sample is safe. At the twig end it is 47%, and 47% says half the data has been silently thrown away on a criterion correlated with the answer.
Nothing else in the analysis carries that information. The fitted value does not; the interval does not; the number of junctions retained does not, unless the number attempted is reported beside it.
So the recommendation is small and specific: report attempted and retained, not retained. It is one extra integer and it converts an invisible bias into a visible one.
What a real study would look like from outside
It is worth imagining the paper this produces, because the point of the essay is that it would not look wrong.
A researcher measures a hundred and fifty junctions across several specimens, taking them as they come — which means mostly lopsided ones, since that is what a tree offers. Sixty or seventy are discarded during analysis for being unphysical, noted in passing as measurement failures. The remainder fit an exponent of about 1.7 with a tight interval.
That is a publishable result and it is an interesting one: it refutes Murray’s law, refutes Da Vinci’s rule, and suggests the network is doing something neither derivation anticipates. A reader would reasonably ask what mechanism gives 1.7, and several plausible ones could be constructed.
Every step of it is defensible in isolation. The measurements are careful, the discards are genuine — those junctions really were unphysical as measured — the fit is correct, the interval is honestly computed. The conclusion is wrong because of a selection that no individual step performed.
That combination is what makes this worth an essay rather than a footnote: there is no incorrect step to find. The correction is not to do any part of it better but to report one more number and to choose the junctions differently.
The retention rate is a measurement, not only a flag
Reporting attempted and retained is the minimum, and the number is worth more than a warning, because it is a measurement of how far the truth sits from the boundary.
Take the difference between a parent’s radius and its larger daughter’s. At the twig end the true value of that difference is 0.11% of a radius and the noise on it is 2.8%, so the difference is a scatter of width 2.8% centred 0.04 of its own width above zero. The share landing above zero is therefore just over a half — 52% — which is what the 47% impossible share says from the other side.
Two numbers already in the essay agreeing to a point is not decoration. It says the truncated-Gaussian picture is the right description of what happened rather than a story about it, and a right description is a thing that can be inverted.
What a corrected fit would return
Which suggests the repair, and following it through is what makes the essay’s opening expectation correct after all.
The bias comes from conditioning on retention: the analysis computes the probability of the surviving observations as though nothing had been discarded, when what it should compute is the probability of the whole sample including the discards, with the impossible region carrying its proper share. That is an ordinary truncated-likelihood correction, it needs the attempted count and nothing else, and it removes the displacement by construction.
What it does not do is recover information. Corrected properly, the twig sample would return an exponent near three with an interval so wide it excludes nothing — because a band whose junctions separate the candidate rules by 0.0012 of a radius against a noise of 0.028 genuinely cannot choose between them, and no amount of correct arithmetic invents what the geometry withheld.
So the essay’s opening expectation was right about the data and wrong about the analysis. Fifty twigs do say the exponent is undetermined. What was novel is that the standard procedure converts that honest ignorance into a confident wrong number, and the conversion happens in a step nobody performs — the silent drop.
Which is why the small recommendation is the right one. A truncated likelihood is the principled fix and it needs a model of the noise; reporting the attempted count needs nothing, and it is enough for a reader to see that the answer cannot mean what it says.
The general trap
Strip out the trees and the structure is one that recurs.
An observable has a domain — a region where it is defined, or physically possible. Measurement error scatters observations across the boundary of that domain. The analysis silently drops the ones that land outside. The survivors are selected on having been scattered inward, which correlates with the quantity being estimated, and the estimate is displaced.
The severity is governed by one ratio: how far the true value sits from the boundary, compared with the noise. Far from the boundary, nothing crosses and there is no selection. Near it, half the sample crosses and the bias is of the order of the effect being measured.
Written that way, the warning sign is available before any data is collected: is the quantity being measured close to a boundary of its own domain, relative to the instrument’s error? Here the answer at is that it is twenty-five times closer, and that alone predicts the whole result.
What it costs the previous essay’s conclusion
The leverage arithmetic recommended choosing junctions rather than accumulating them, on the grounds that lopsided ones carry almost no information. This essay strengthens that from an efficiency argument into a correctness one.
Measuring uninformative junctions is not merely wasteful. It is actively harmful, because the estimator does not degrade gracefully — it returns a precise wrong answer rather than an imprecise right one, and the failure is invisible in every quantity a study would normally report.
So the protocol gains a clause:
Measure the forks where the two branches are comparable. Skip the rest — and if they are measured anyway, report how many came out impossible, because that number is the only warning the analysis will give.
A note on what was tried first
The result above was not the expected one, and the route to it is worth recording because the first version of this essay would have been wrong in an interesting way.
The prediction, written from the previous essay’s arithmetic, was that the twig sample would return a wide interval containing both hypotheses — uninformative in the ordinary sense. The measurement gave 1.69 with a narrow interval, which read at first like a defect in the site’s own fitter.
Replacing the fitter with a per-junction solve was the obvious diagnostic, and it gave 1.64. Two unrelated estimators agreeing on a wrong answer is what moved the search from the estimator to the data, and counting the impossible junctions took a minute once the question was asked.
The lesson is the one this collection keeps arriving at from different directions: when a result is surprising, check whether the surprise is in the machinery or in the sample, and the cheapest way to tell is a second estimator that shares no code.
Where else this site is exposed to it
Having found the trap once, the honest thing is to look for it in this collection’s own measurements rather than to leave it as a warning about somebody else’s.
Two places are exposed and both turn out to be handled, which is worth stating with the reason rather than as reassurance.
The noise sweeps discard incoherent runs. Every ensemble here reports its scatter
over the runs that still have a lattice, and drops the ones that do not — which is
exactly a domain cut with selection on the outcome. It is safe only where the discard
rate is zero or one: at amplitudes where every run survives, nothing is selected; at
amplitudes where none does, there is no estimate. The site’s own machinery reports
coherentShare beside every scatter for this reason, and the essays that use a mean
scatter take it only from rows at 100%. The one place a partial share is used is the
tolerance boundary itself, where the quantity of interest is the share.
The angle recovery refuses non-coprime pairs. A whorled head’s counts share a factor and no divergence angle is recoverable, so those heads are excluded. That is a domain cut too — but it is not correlated with the recovered angle, because whether a pair is coprime is decided by the counts rather than by any measurement error on them. A cut on an exactly-known quantity cannot select on noise.
The distinction between those two cases is the useful one. A domain cut is dangerous when the thing that pushes an observation across the boundary is the same noise that would have biased the estimate. Where the boundary is crossed for a reason unrelated to the measurement error, discarding costs sample size and nothing else.
That gives the check a form that can be applied quickly to any of this collection’s results: for every discard rule, ask what would have to vary for an observation to cross it, and whether that thing is in the estimator.
The one-sentence version
An observation that cannot happen is a measurement error, dropping it is automatic, and the ones that survive are the ones the error pushed the right way — so a sample taken where the signal is smaller than the noise does not report uncertainty, it reports a different answer with confidence.
The number that would have caught it was available the whole time and is one integer: how many junctions were attempted and how many were used.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- What a summary throws away — both name bias, branching exponent, measurement, sampling, selection, specimen, summary statistic, survey
- What the protractor has to be — both name measurement, noise, sampling, specimen, summary statistic, survey
- A correction that keeps the overlap — both name bias, branching exponent, da vinci's rule, fitting, murray's law
- The test a plant could settle — both name measurement, noise, specimen, summary statistic, survey
- Two readings from one stem — both name measurement, noise, sampling, specimen, summary statistic
- A counter that sees no positions — both name measurement, noise, sampling, summary statistic
Named objects
A flat tag is an object no other essay names yet.
AllometryBiasBranching exponentDa Vinci's ruleFittingMeasurementMurray's lawNoiseSamplingSelectionSpecimenSummary statisticSurveyTransport