Branching and transport

The fragile junctions are the informative ones

That is the obvious worry once the radii are uncertain, and it is false. Across the whole range of asymmetry a junction's contribution to the bias moves by a factor of 1.36 while its leverage moves by a factor of 308,352, so the junction that says nothing damages the answer as badly as the one that says everything — and a sample is spoiled by counting rather than by weight.

Worth reading first: The exponent an error moves · Fitting the exponent · How many plants would it take.

This collection has established two things about a branching junction and they point the same way. An even fork settles which of the two candidate rules a tree obeys and a twig settles nothing, by a factor of two thousand in separation; and a least-squares fit already weights by that separation squared, so a mixed sample is not dragged towards the twigs.

Once the radii carry error, the natural next worry is that the leverage cuts both ways — that the junctions doing the work are also the ones a measurement error damages most, and that the whole arrangement is therefore precarious.

It is not, and the replacement is worse than the worry. Damage and information are two almost unrelated quantities, and separating them is the whole of this essay.

One junction's bias moves 1.36-fold across the range its leverage moves 308,352-fold, at an exponent of 3. The bias a single junction of daughter ratio γ contributes to a least-squares fit, as a multiple of the squared measurement error, drawn against the leverage that junction carries — both computed from the expansion about a true exponent of 3 rather than fitted to anything. Across the whole range from an even fork to a twentieth the bias moves by a factor of 1.36 and the leverage by a factor of 308,352, and the uninformative junction contributes the larger share: 6.00σ² at γ = 0.05 against 4.50σ² at an even fork.
Fig. 1 Two quantities per junction against its daughter ratio: the bias it contributes to a fitted exponent, and the leverage it carries. One of them is flat and the other falls off the bottom of the picture.

The two quantities, per junction

Both come out of the same expansion that prices the displacement, and both are properties of a junction’s shape alone — neither knows the error level, and neither knows how many junctions are in the sample.

The leverage is c02c_0^2, where c0c_0 is the derivative of the junction’s residual in the exponent. It is what a least-squares fit weights the junction by, so it is the information the junction carries about the answer.

The bias contribution is the junction’s expected cross-term between its residual and that derivative, per unit of σ2\sigma^2. It is what the junction adds to the numerator of the step the fitter takes away from the truth.

The spans

At a true exponent of three, across daughter ratios from an even fork down to a twentieth:

γ\gamma bias contributed (÷σ²) leverage c02c_0^2 ratio
1 4.500 5.338 × 10⁻² 84.3
0.9 4.573 5.150 × 10⁻² 88.8
0.8 4.801 4.553 × 10⁻² 105
0.6 5.557 2.431 × 10⁻² 229
0.5 5.887 1.352 × 10⁻² 435
0.3 6.109 1.643 × 10⁻³ 3,720
0.2 6.065 2.381 × 10⁻⁴ 25,500
0.15 6.037 5.632 × 10⁻⁵ 107,000
0.1 6.015 6.935 × 10⁻⁶ 867,000
0.05 6.003 1.731 × 10⁻⁷ 3.47 × 10⁷

The bias column runs from 4.500 to 6.109 — a factor of 1.36. The leverage column runs from 5.338 × 10⁻² to 1.731 × 10⁻⁷ — a factor of 308,352.

Both ends of the bias column are exact

The flatness is not an empirical observation about ten sampled ratios. It is what the expression does at its two limits, and both limits can be written down.

At an even fork, γ=1\gamma = 1 and S=2S = 2, so the second bracket vanishes with logγ\log \gamma and the first leaves p[(1+1)/4+1]p\left[(1+1)/4 + 1\right], which at p=3p = 3 is 4.500 — the first row of the table, exactly.

At the other end γplogγ\gamma^{\,p} \log \gamma tends to zero and SS tends to one, so the same expression tends to 2p2p, which at p=3p = 3 is 6.000 — against 6.003 measured at a twentieth. So the whole column is bracketed between p[(1+1)/4+1]p[(1+1)/4 + 1] and 2p2p, and the ratio of those is 4/3 whatever the exponent. A quantity confined between two multiples of pp that differ by a third is a quantity that cannot carry a sample design.

And the leverage’s upper end is a single logarithm

The same substitution does the other column. At an even fork c0c_0 reduces to logR-\log R, which is log21/3=0.2310-\log 2^{1/3} = -0.2310, and its square is the 5.338 × 10⁻² in the table.

That number is the most information a junction of any shape can carry about the exponent under this construction, and it is not large in absolute terms — a tenth of a nat per junction. What makes an even fork valuable is not that it is informative; it is that everything else is so much less so.

Which way round it goes

Not merely different by five orders of magnitude. Different in sign.

The junction that carries the least information contributes the most bias. A twig at a daughter ratio of a twentieth contributes 6.00 σ² against an even fork’s 4.50 σ² — thirty-three per cent more damage for one three-hundred-thousandth of the information.

So the worry that opened this essay is not merely unfounded; it is upside down. The informative junctions are the robust ones, and their robustness is not what saves a sample.

The junction that can say nothing contributes 6.00σ² against an even fork's 4.50σ², at an exponent of 3. The bias a single junction of daughter ratio γ contributes to a least-squares fit, as a multiple of the squared measurement error, drawn against the leverage that junction carries — both computed from the expansion about a true exponent of 3 rather than fitted to anything. Across the whole range from an even fork to a twentieth the bias moves by a factor of 1.36 and the leverage by a factor of 308,352, and the uninformative junction contributes the larger share: 6.00σ² at γ = 0.05 against 4.50σ² at an even fork.
Fig. 2 The same two curves with the twig band shaded. The junction that says nothing about the exponent contributes more bias than the one that settles it, which is the opposite of what the leverage argument leads a reader to expect.

Why, in one line

A least-squares fit weights information by leverage and bias by counting.

The step the fitter takes away from the truth is a ratio: the summed bias contributions over the summed leverages. The denominator is dominated by whichever junctions have leverage, because leverage varies by five orders of magnitude. The numerator is not dominated by anything, because every junction contributes between 4.5 and 6.1 of it.

Add a junction with no leverage and the denominator barely moves while the numerator gains its full six. That is the mechanism, and it is arithmetic rather than a property of trees.

Damage per unit of information

Dividing one column by the other gives the quantity that matters for sample design: how much bias a junction buys per unit of information it supplies.

It runs from 84 at an even fork to 3.47 × 10⁷ at a twentieth — a factor of four hundred thousand. A twig is not a weak measurement of the exponent. It is a nearly pure delivery mechanism for bias.

The damage one junction does per unit of information, from 84 at an even fork to 34.7 million at a twentieth, at an exponent of 3. Bias divided by leverage, junction by junction, at a true exponent of 3. An even fork costs 84 units of bias for its information; a junction whose small daughter is a twentieth of the large one costs 34.7 million — a factor of 411,315 across a range over which the bias itself barely moves. A least-squares fit weights information by leverage and bias by counting, which is why the ratio and not the bias is what a sample has to be chosen against.
Fig. 3 Bias divided by leverage, on a logarithmic scale because no linear one holds both ends. This is the curve a sample design should be read against: it is the cost of including a junction, per unit of what including it buys.

Where the bias peaks

The bias column is not monotone, and the detail is worth a sentence because it is the only structure in an otherwise flat quantity.

It rises from 4.500 at an even fork to 6.109 at a daughter ratio of 0.3, then eases very slightly to 6.003 at a twentieth. So the maximum damage per junction sits at the same place the truncation mechanism switches on — about 0.3, where a parent’s genuine margin over its larger daughter falls through the noise.

The two mechanisms therefore peak together, which is unhelpful and is not a coincidence: both are about how small the signal at a junction has become relative to the perturbation.

The same arithmetic at the rival exponent

None of this is special to Murray’s law. Run the same expansion about an exponent of two and the shape is the same: a bias that barely moves across the range and a leverage that collapses.

The spans are narrower at two, which is the same fact that makes a tree built at two harder to displace, and the conclusion is unchanged. The pattern belongs to the form of the relation being fitted rather than to the particular exponent in it.

One junction's bias moves 1.36-fold across the range its leverage moves 1,580-fold, at an exponent of 2. The bias a single junction of daughter ratio γ contributes to a least-squares fit, as a multiple of the squared measurement error, drawn against the leverage that junction carries — both computed from the expansion about a true exponent of 2 rather than fitted to anything. Across the whole range from an even fork to a twentieth the bias moves by a factor of 1.36 and the leverage by a factor of 1,580, and the uninformative junction contributes the larger share: 4.02σ² at γ = 0.05 against 3.00σ² at an even fork.
Fig. 4 The per-junction arithmetic expanded about an exponent of two rather than three. The bias stays flat, the leverage still collapses, and the qualitative statement does not depend on which of the two rules is true.

What that costs a chosen sample

The per-junction arithmetic predicts something specific about a sample somebody would actually gather, and the prediction is testable.

Take twenty even forks — daughter ratios 0.9 to 1, which is a good morning’s work on a dichotomous tree and about the best sample anyone can hope for. Then add twigs, which is what the same tree offers by the hundred.

Each mixture is measured with three hundred replicate samples, so the reported means and intervals are about the mixture rather than about one draw of the noise.

Why three hundred replicates

Each row of the dilution is measured over three hundred independent draws of the noise, which puts the standard error on a reported mean at about a seventeenth of the reported spread.

That matters here more than it usually does, because the quantity being reported is a displacement of a few hundredths at the top of the ladder. A displacement a tenth of the size of the spread has to be visible for the ladder’s first rungs to say anything at all, and at three hundred replicates it is. The mixtures are not one sample’s luck being read as a trend.

Twenty twigs double the displacement

At two per cent of error the twenty even forks alone return 2.971, a displacement of −0.029.

Adding twenty twigs takes it to 2.928, a displacement of −0.072. Twenty twigs therefore more than double the damage, while adding about one ten-thousandth of the information the twenty forks already carry.

That is the counting rule made concrete. Nothing was measured badly; twenty more junctions were measured, each of them contributing its six.

A thousand twigs

The end of the ladder is the composition a tree actually presents. Twenty even forks and a thousand twigs, at two per cent on every radius, return 1.915 with an interval of [1.88, 1.96].

The tree was built at three. The answer is wrong by more than the entire distance between the two rules being told apart, and the interval excludes both of them.

20 even forks diluted with twigs at 2%: 2.971 down to 1.915. Twenty junctions of daughter ratio 0.9–1 — a good morning's work with a pair of callipers — with twigs of ratio 0.05–0.15 added to them, from a tree built at exactly 3. The dashed line is the displacement the expansion predicts from the daughter ratios alone; the shaded band is the central 90% of 300 replicate samples. At 1000 twigs the answer is 1.915 with an interval of [1.88, 1.96], and the interval is 2.9 times narrower than the one twenty even forks give on their own.
Fig. 5 Twenty even forks diluted with twigs at two per cent, from none to a thousand. The mean walks from 2.971 down to 1.915 and the interval narrows the whole way, so the report becomes more confident as it becomes more wrong.

It does not go all the way to the twigs’ own answer

The twig band measured on its own at two per cent returns 1.648. The mixture of twenty forks and a thousand twigs returns 1.915, which sits above that and well below three.

So the leverage argument is still doing something: twenty even forks, outnumbered fifty to one, still hold the answer a quarter of a point above where the twigs alone would put it. Their information is not being ignored. It is simply being outvoted on the numerator by a crowd whose contribution to the denominator is negligible.

That is the sharpest way to see what has gone wrong. The estimator is behaving exactly as designed on both counts at once — weighting the forks properly and counting the twigs properly — and the composition of the sample is what turns those two correct behaviours into a wrong number.

The interval narrows as the answer worsens

That last clause is the part that would do the damage in a published study, and it is worth its own numbers.

The spread across replicates falls from 0.073 at twenty junctions to 0.025 at a thousand and twenty. The interval at the end is three times narrower than the interval the twenty even forks give on their own.

So every diagnostic that is actually reported improves as the answer deteriorates. More junctions, a tighter interval, a smaller standard error — and a number wrong by 1.085.

Nothing in the output says so

There is no line in the analysis that flags it, because nothing was done wrong. The junctions are real, the fitter is the same fitter that returns whatever exponent it is given, and the interval honestly prices the sampling variation in the sample it was computed from.

What an interval cannot price is a displacement introduced before it was computed. That is not a subtlety of this case; it is what an interval is, and it is why the composition of a sample has to be reported alongside its size.

At half a per cent it is still wrong

The obvious response is that two per cent is a poor instrument and a better one fixes it. It does not.

At half a per cent — a machined section under a microscope, about as good as a radius measurement gets — the same twenty forks return 2.998 on their own and 2.865 with a thousand twigs added, interval [2.83, 2.89]. That interval excludes the truth.

Quartering the error moved the displacement from 1.085 to 0.135, which is the square law working exactly as it should. It did not move the shape of the failure at all.

20 even forks diluted with twigs at 0.5%: 2.998 down to 2.865. Twenty junctions of daughter ratio 0.9–1 — a good morning's work with a pair of callipers — with twigs of ratio 0.05–0.15 added to them, from a tree built at exactly 3. The dashed line is the displacement the expansion predicts from the daughter ratios alone; the shaded band is the central 90% of 300 replicate samples. At 1000 twigs the answer is 2.865 with an interval of [2.83, 2.89], and the interval is 1.1 times narrower than the one twenty even forks give on their own.
Fig. 6 The same dilution at half a per cent, with the expansion’s prediction drawn over it. A thousand twigs still take a correct answer to 2.865, and the interval around it still excludes three.

And the expansion predicts all of it

At half a per cent the derived displacements are −0.0021, −0.0050, −0.0164, −0.0730 and −0.1429 at 0, 20, 100, 500 and 1,000 added twigs, against measured −0.0023, −0.0051, −0.0164, −0.0708 and −0.1355.

Five per cent or better at every mixture out to 1,020 junctions, with nothing fitted. So the dilution is not an empirical curiosity of one seed: it is the per-junction arithmetic summed, and it can be computed for any proposed sample before the sample is gathered.

What this does to the earlier finding

It qualifies it rather than contradicting it, and the distinction is exact.

The earlier result was that daughter ratios spread evenly from 0.05 to 1 fit 2.906 at two per cent — safe, on the true value, dragged by almost nothing. That is measured again here and it holds.

But an even spread is not a tree. A tree has one trunk fork and hundreds of twigs, and at that composition the same fitter on the same trees returns 1.915. What was shown to be safe was the even spread, and nobody ever gathers one.

A junction's bias moves 1.36-fold where its leverage moves 308,352-fold, and 1000 twigs at 2% move the answer 1.085. Above: what one junction of daughter ratio γ contributes to the bias of a fitted exponent, against the leverage it carries, both from the expansion about 3 and neither fitted. The bias spans a factor of 1.36 and the leverage a factor of 308,352, so the junction that can say nothing does 33% more damage than the one that settles the question. Below: what that means for a sample somebody actually takes — 20 chosen even forks at 2% return 2.971, and the same forks with 1000 twigs beside them return 1.915, with an interval that has narrowed rather than widened.
Fig. 7 The per-junction arithmetic and what it does to a chosen sample, in one reading. The flat bias and the collapsing leverage above; the twenty even forks being diluted below, with their starting point marked.

The sampling rule that follows

Stated plainly, because the whole essay is for it.

Do not add junctions to a branching sample. Choose them, and report the daughter ratio of every one. A junction below a daughter ratio of about 0.6 carries so little information that including it is very close to a pure addition of bias, and the arithmetic above says how much: six units of σ² apiece, whatever else is in the sample.

The corollary is uncomfortable and it is the honest one. A sample of twenty good forks is better than the same twenty forks plus everything else the tree offers, and it is better by more than a factor of thirty in displacement.

Why “measure more” is exactly wrong here

Adding data is the default remedy for a disappointing measurement and it is the wrong instinct in this family, for a reason that is not about branching at all.

More observations reduce variance. They do nothing to a bias, and if the added observations carry bias without carrying information they increase it. What a larger sample buys here is therefore confidence in a wrong number, which is the worst thing a sample can buy.

The general form is a warning about any estimator with weights in it: the weights protect the answer’s precision from uninformative data and do not protect its accuracy at all. The essay on what a sample size can and cannot fix is the same statement made against nn rather than against composition.

Which relocates the danger, again

The earlier work relocated the danger from the mixture to the selection: a sample taken by walking through a wood with callipers is a twig sample, because twigs are in the hand and the trunk fork is twenty metres up.

This moves it once more. A conscientious measurer who climbs for the trunk fork and takes every twig on the way down has produced a worse sample than one who climbed and took nothing else. Diligence about quantity is the failure mode, and it is the one no reviewer would query.

That is the same shape as a sample grid whose spacing decides its own answer: the defect is in the design rather than in the execution, and the execution can be flawless.

What a survey should report

Three things, and none costs anything.

The daughter ratios, as a distribution rather than a range — because the displacement is a sum over them and cannot be reconstructed from a minimum and a maximum. The error on a radius, since everything here scales as its square. And attempted against retained, which is the diagnostic the earlier work argued for and which reports the truncation mechanism that this arithmetic does not cover.

With those three a reader can compute the expected displacement of a published exponent to within a few per cent. Without the first, no reader can compute it at all.

What this does not support

It does not say the leverage argument was wrong. Least squares really does weight information correctly, and a mixed sample really is undragged in the sense that essay meant — its point estimate is not pulled towards the exponent the twigs would report on their own.

The dilution here is a different quantity. It is the bias each junction contributes under measurement error, which the leverage argument was not about and which is zero when the radii are exact. Both statements are true at once and the second only bites when there is an error.

Where the arithmetic stops

Everything above is the leading-order expansion, and it stops applying below a daughter ratio of about 0.3, where junctions start being made physically impossible by the noise and are silently dropped.

The mixtures here are measured rather than derived, so they include that second mechanism whether or not the expansion describes it — which is why the agreement at half a per cent is the interesting one. At half a per cent the truncation is mild, the expansion applies, and the two routes to the same number agree to five per cent. At two per cent the measurement is still right and the derivation is doing less of the work.

The two numbers to carry

1.36 and 308,352. A junction’s contribution to the damage barely depends on its shape; what it can tell anybody depends on its shape enormously. Every consequence in this essay follows from those two spans being so far apart.

And the consequence that matters outside this collection is the one about counting. An estimator that weights its observations is protecting the wrong thing — it is protecting the answer from noise in uninformative data, and the uninformative data is arriving with a systematic displacement attached. That is a property of weighted estimation rather than of trees, and it is why a sample is spoiled by counting rather than by weight.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

BiasBranching exponentDa Vinci's ruleDaughter ratioHonest limitsInterval estimateLeast-squares fitLeverageMeasurement errorMurray's lawSample sizeSamplingSelectionSummary statistic