The fragile junctions are the informative ones
Worth reading first: The exponent an error moves · Fitting the exponent · How many plants would it take.
This collection has established two things about a branching junction and they point the same way. An even fork settles which of the two candidate rules a tree obeys and a twig settles nothing, by a factor of two thousand in separation; and a least-squares fit already weights by that separation squared, so a mixed sample is not dragged towards the twigs.
Once the radii carry error, the natural next worry is that the leverage cuts both ways — that the junctions doing the work are also the ones a measurement error damages most, and that the whole arrangement is therefore precarious.
It is not, and the replacement is worse than the worry. Damage and information are two almost unrelated quantities, and separating them is the whole of this essay.
The two quantities, per junction
Both come out of the same expansion that prices the displacement, and both are properties of a junction’s shape alone — neither knows the error level, and neither knows how many junctions are in the sample.
The leverage is , where is the derivative of the junction’s residual in the exponent. It is what a least-squares fit weights the junction by, so it is the information the junction carries about the answer.
The bias contribution is the junction’s expected cross-term between its residual and that derivative, per unit of . It is what the junction adds to the numerator of the step the fitter takes away from the truth.
The spans
At a true exponent of three, across daughter ratios from an even fork down to a twentieth:
| bias contributed (÷σ²) | leverage | ratio | |
|---|---|---|---|
| 1 | 4.500 | 5.338 × 10⁻² | 84.3 |
| 0.9 | 4.573 | 5.150 × 10⁻² | 88.8 |
| 0.8 | 4.801 | 4.553 × 10⁻² | 105 |
| 0.6 | 5.557 | 2.431 × 10⁻² | 229 |
| 0.5 | 5.887 | 1.352 × 10⁻² | 435 |
| 0.3 | 6.109 | 1.643 × 10⁻³ | 3,720 |
| 0.2 | 6.065 | 2.381 × 10⁻⁴ | 25,500 |
| 0.15 | 6.037 | 5.632 × 10⁻⁵ | 107,000 |
| 0.1 | 6.015 | 6.935 × 10⁻⁶ | 867,000 |
| 0.05 | 6.003 | 1.731 × 10⁻⁷ | 3.47 × 10⁷ |
The bias column runs from 4.500 to 6.109 — a factor of 1.36. The leverage column runs from 5.338 × 10⁻² to 1.731 × 10⁻⁷ — a factor of 308,352.
Both ends of the bias column are exact
The flatness is not an empirical observation about ten sampled ratios. It is what the expression does at its two limits, and both limits can be written down.
At an even fork, and , so the second bracket vanishes with and the first leaves , which at is 4.500 — the first row of the table, exactly.
At the other end tends to zero and tends to one, so the same expression tends to , which at is 6.000 — against 6.003 measured at a twentieth. So the whole column is bracketed between and , and the ratio of those is 4/3 whatever the exponent. A quantity confined between two multiples of that differ by a third is a quantity that cannot carry a sample design.
And the leverage’s upper end is a single logarithm
The same substitution does the other column. At an even fork reduces to , which is , and its square is the 5.338 × 10⁻² in the table.
That number is the most information a junction of any shape can carry about the exponent under this construction, and it is not large in absolute terms — a tenth of a nat per junction. What makes an even fork valuable is not that it is informative; it is that everything else is so much less so.
Which way round it goes
Not merely different by five orders of magnitude. Different in sign.
The junction that carries the least information contributes the most bias. A twig at a daughter ratio of a twentieth contributes 6.00 σ² against an even fork’s 4.50 σ² — thirty-three per cent more damage for one three-hundred-thousandth of the information.
So the worry that opened this essay is not merely unfounded; it is upside down. The informative junctions are the robust ones, and their robustness is not what saves a sample.
Why, in one line
A least-squares fit weights information by leverage and bias by counting.
The step the fitter takes away from the truth is a ratio: the summed bias contributions over the summed leverages. The denominator is dominated by whichever junctions have leverage, because leverage varies by five orders of magnitude. The numerator is not dominated by anything, because every junction contributes between 4.5 and 6.1 of it.
Add a junction with no leverage and the denominator barely moves while the numerator gains its full six. That is the mechanism, and it is arithmetic rather than a property of trees.
Damage per unit of information
Dividing one column by the other gives the quantity that matters for sample design: how much bias a junction buys per unit of information it supplies.
It runs from 84 at an even fork to 3.47 × 10⁷ at a twentieth — a factor of four hundred thousand. A twig is not a weak measurement of the exponent. It is a nearly pure delivery mechanism for bias.
Where the bias peaks
The bias column is not monotone, and the detail is worth a sentence because it is the only structure in an otherwise flat quantity.
It rises from 4.500 at an even fork to 6.109 at a daughter ratio of 0.3, then eases very slightly to 6.003 at a twentieth. So the maximum damage per junction sits at the same place the truncation mechanism switches on — about 0.3, where a parent’s genuine margin over its larger daughter falls through the noise.
The two mechanisms therefore peak together, which is unhelpful and is not a coincidence: both are about how small the signal at a junction has become relative to the perturbation.
The same arithmetic at the rival exponent
None of this is special to Murray’s law. Run the same expansion about an exponent of two and the shape is the same: a bias that barely moves across the range and a leverage that collapses.
The spans are narrower at two, which is the same fact that makes a tree built at two harder to displace, and the conclusion is unchanged. The pattern belongs to the form of the relation being fitted rather than to the particular exponent in it.
What that costs a chosen sample
The per-junction arithmetic predicts something specific about a sample somebody would actually gather, and the prediction is testable.
Take twenty even forks — daughter ratios 0.9 to 1, which is a good morning’s work on a dichotomous tree and about the best sample anyone can hope for. Then add twigs, which is what the same tree offers by the hundred.
Each mixture is measured with three hundred replicate samples, so the reported means and intervals are about the mixture rather than about one draw of the noise.
Why three hundred replicates
Each row of the dilution is measured over three hundred independent draws of the noise, which puts the standard error on a reported mean at about a seventeenth of the reported spread.
That matters here more than it usually does, because the quantity being reported is a displacement of a few hundredths at the top of the ladder. A displacement a tenth of the size of the spread has to be visible for the ladder’s first rungs to say anything at all, and at three hundred replicates it is. The mixtures are not one sample’s luck being read as a trend.
Twenty twigs double the displacement
At two per cent of error the twenty even forks alone return 2.971, a displacement of −0.029.
Adding twenty twigs takes it to 2.928, a displacement of −0.072. Twenty twigs therefore more than double the damage, while adding about one ten-thousandth of the information the twenty forks already carry.
That is the counting rule made concrete. Nothing was measured badly; twenty more junctions were measured, each of them contributing its six.
A thousand twigs
The end of the ladder is the composition a tree actually presents. Twenty even forks and a thousand twigs, at two per cent on every radius, return 1.915 with an interval of [1.88, 1.96].
The tree was built at three. The answer is wrong by more than the entire distance between the two rules being told apart, and the interval excludes both of them.
It does not go all the way to the twigs’ own answer
The twig band measured on its own at two per cent returns 1.648. The mixture of twenty forks and a thousand twigs returns 1.915, which sits above that and well below three.
So the leverage argument is still doing something: twenty even forks, outnumbered fifty to one, still hold the answer a quarter of a point above where the twigs alone would put it. Their information is not being ignored. It is simply being outvoted on the numerator by a crowd whose contribution to the denominator is negligible.
That is the sharpest way to see what has gone wrong. The estimator is behaving exactly as designed on both counts at once — weighting the forks properly and counting the twigs properly — and the composition of the sample is what turns those two correct behaviours into a wrong number.
The interval narrows as the answer worsens
That last clause is the part that would do the damage in a published study, and it is worth its own numbers.
The spread across replicates falls from 0.073 at twenty junctions to 0.025 at a thousand and twenty. The interval at the end is three times narrower than the interval the twenty even forks give on their own.
So every diagnostic that is actually reported improves as the answer deteriorates. More junctions, a tighter interval, a smaller standard error — and a number wrong by 1.085.
Nothing in the output says so
There is no line in the analysis that flags it, because nothing was done wrong. The junctions are real, the fitter is the same fitter that returns whatever exponent it is given, and the interval honestly prices the sampling variation in the sample it was computed from.
What an interval cannot price is a displacement introduced before it was computed. That is not a subtlety of this case; it is what an interval is, and it is why the composition of a sample has to be reported alongside its size.
At half a per cent it is still wrong
The obvious response is that two per cent is a poor instrument and a better one fixes it. It does not.
At half a per cent — a machined section under a microscope, about as good as a radius measurement gets — the same twenty forks return 2.998 on their own and 2.865 with a thousand twigs added, interval [2.83, 2.89]. That interval excludes the truth.
Quartering the error moved the displacement from 1.085 to 0.135, which is the square law working exactly as it should. It did not move the shape of the failure at all.
And the expansion predicts all of it
At half a per cent the derived displacements are −0.0021, −0.0050, −0.0164, −0.0730 and −0.1429 at 0, 20, 100, 500 and 1,000 added twigs, against measured −0.0023, −0.0051, −0.0164, −0.0708 and −0.1355.
Five per cent or better at every mixture out to 1,020 junctions, with nothing fitted. So the dilution is not an empirical curiosity of one seed: it is the per-junction arithmetic summed, and it can be computed for any proposed sample before the sample is gathered.
What this does to the earlier finding
It qualifies it rather than contradicting it, and the distinction is exact.
The earlier result was that daughter ratios spread evenly from 0.05 to 1 fit 2.906 at two per cent — safe, on the true value, dragged by almost nothing. That is measured again here and it holds.
But an even spread is not a tree. A tree has one trunk fork and hundreds of twigs, and at that composition the same fitter on the same trees returns 1.915. What was shown to be safe was the even spread, and nobody ever gathers one.
The sampling rule that follows
Stated plainly, because the whole essay is for it.
Do not add junctions to a branching sample. Choose them, and report the daughter ratio of every one. A junction below a daughter ratio of about 0.6 carries so little information that including it is very close to a pure addition of bias, and the arithmetic above says how much: six units of σ² apiece, whatever else is in the sample.
The corollary is uncomfortable and it is the honest one. A sample of twenty good forks is better than the same twenty forks plus everything else the tree offers, and it is better by more than a factor of thirty in displacement.
Why “measure more” is exactly wrong here
Adding data is the default remedy for a disappointing measurement and it is the wrong instinct in this family, for a reason that is not about branching at all.
More observations reduce variance. They do nothing to a bias, and if the added observations carry bias without carrying information they increase it. What a larger sample buys here is therefore confidence in a wrong number, which is the worst thing a sample can buy.
The general form is a warning about any estimator with weights in it: the weights protect the answer’s precision from uninformative data and do not protect its accuracy at all. The essay on what a sample size can and cannot fix is the same statement made against rather than against composition.
Which relocates the danger, again
The earlier work relocated the danger from the mixture to the selection: a sample taken by walking through a wood with callipers is a twig sample, because twigs are in the hand and the trunk fork is twenty metres up.
This moves it once more. A conscientious measurer who climbs for the trunk fork and takes every twig on the way down has produced a worse sample than one who climbed and took nothing else. Diligence about quantity is the failure mode, and it is the one no reviewer would query.
That is the same shape as a sample grid whose spacing decides its own answer: the defect is in the design rather than in the execution, and the execution can be flawless.
What a survey should report
Three things, and none costs anything.
The daughter ratios, as a distribution rather than a range — because the displacement is a sum over them and cannot be reconstructed from a minimum and a maximum. The error on a radius, since everything here scales as its square. And attempted against retained, which is the diagnostic the earlier work argued for and which reports the truncation mechanism that this arithmetic does not cover.
With those three a reader can compute the expected displacement of a published exponent to within a few per cent. Without the first, no reader can compute it at all.
What this does not support
It does not say the leverage argument was wrong. Least squares really does weight information correctly, and a mixed sample really is undragged in the sense that essay meant — its point estimate is not pulled towards the exponent the twigs would report on their own.
The dilution here is a different quantity. It is the bias each junction contributes under measurement error, which the leverage argument was not about and which is zero when the radii are exact. Both statements are true at once and the second only bites when there is an error.
Where the arithmetic stops
Everything above is the leading-order expansion, and it stops applying below a daughter ratio of about 0.3, where junctions start being made physically impossible by the noise and are silently dropped.
The mixtures here are measured rather than derived, so they include that second mechanism whether or not the expansion describes it — which is why the agreement at half a per cent is the interesting one. At half a per cent the truncation is mild, the expansion applies, and the two routes to the same number agree to five per cent. At two per cent the measurement is still right and the derivation is doing less of the work.
The two numbers to carry
1.36 and 308,352. A junction’s contribution to the damage barely depends on its shape; what it can tell anybody depends on its shape enormously. Every consequence in this essay follows from those two spans being so far apart.
And the consequence that matters outside this collection is the one about counting. An estimator that weights its observations is protecting the wrong thing — it is protecting the answer from noise in uninformative data, and the uninformative data is arriving with a systematic displacement attached. That is a property of weighted estimation rather than of trees, and it is why a sample is spoiled by counting rather than by weight.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- An optimum too flat to reach — both name bias, branching exponent, daughter ratio, honest limits, measurement error, murray's law, sample size, summary statistic
- Forty angles, and a limit — both name bias, honest limits, measurement error, sample size, sampling
- What a summary throws away — both name bias, branching exponent, sampling, selection, summary statistic
- A basin has a width — both name honest limits, measurement error, sample size, sampling
- A dip with no outer edge — both name honest limits, measurement error, sampling, summary statistic
- A wall that stopped moving — both name honest limits, measurement error, sample size, sampling
Named objects
A flat tag is an object no other essay names yet.
BiasBranching exponentDa Vinci's ruleDaughter ratioHonest limitsInterval estimateLeast-squares fitLeverageMeasurement errorMurray's lawSample sizeSamplingSelectionSummary statistic