Branching and transport

A correction that keeps the overlap

The duel between a tree built at Murray's exponent and one built at Da Vinci's ended by saying the displacement is the geometry, and that no better estimator removes it. Correcting every replicate by simulation-extrapolation removes 92 per cent of the tree at three's displacement at five per cent of error and 68 per cent at twelve, and the error at which the two means cross leaves the measured range altogether. It pays in spread — the corrected readings are twice as wide at twelve per cent — so the error at which the two trees' intervals overlap does not move. Of the duel's two numbers, the inversion was the estimator's and the overlap is the question's.

Worth reading first: The exponent an error moves · Fitting the exponent · The cube law.

Two trees measured through the same fifty junctions — one built at Murray’s exponent of three, one at Da Vinci’s two — overlap at twelve per cent of relative error on every radius, and above twenty and a half per cent the tree built at three reads lower than the tree built at two. That essay drew a conclusion from the inversion: an estimator that did not invert would need a displacement the same at every truth, “and the displacement is the geometry”.

That is a claim about every estimator, and one estimator is enough to test it. The displacement itself is a known function of the error’s size, which is exactly the situation a standard correction exists for.

Adding error to carry a fitted exponent back to none, at 12%. At 12% of error on every radius each replicate is refitted with more error added at four levels, and the dots are the means over 300 replicates, the tree at 3 above and the tree at 2 below. Uncorrected they read 2.035 and 1.676. Carried back to no error through the mean points, the tree at 3 reads 2.695 by a quadratic curve, 2.380 by a linear curve, 2.914 by a rational curve; the tree at 2 reads 1.961 by a quadratic one, 1.877 by a linear one, 1.982 by a rational one. The curve carried back is a choice the method does not make.
Fig. 1 At twelve per cent of error, the mean fitted exponent of each tree as more error is deliberately added, and three curves carrying the trend back to no error at all. Where the curve is carried back is a choice, and the three choices land in different places.

Adding error to remove it

Simulation-extrapolation is the standard correction for measurement error of a known size, and its logic fits in a sentence. If adding error moves the answer in a predictable way, then measure how it moves by adding more, and carry the trend back past the error already present to the point where there would be none.

It needs nothing but the sample in hand and the size of the error. It is not a model of branching, and it knows nothing about Murray or Da Vinci; it treats the fitter as a black box whose answer drifts with the noise, and extrapolates the drift.

On the duel’s own replicates

The correction is run on exactly the replicates the duel used — three hundred at each error, fifty junctions of daughter ratio 0.6 to 1 — so its uncorrected rows are the duel’s rows, and every difference below is the correction’s.

Each replicate is refitted with added error at four levels, eight times at each, so every error level costs 2.88 million refits. A faster golden-section fitter lands within the original fitter’s own resolution of where the original fitter lands, checked on 120 samples at three errors and both trees, so what is corrected is the duel’s own estimator and not a substitute for it.

At twelve per cent

Uncorrected, the two trees read 2.035 and 1.676. As error is added the means fall steadily. Carried back to no error through the mean points, the tree at three reads 2.695 by a quadratic curve, 2.380 by a straight line and 2.914 by a rational curve; the tree at two reads 1.961, 1.877 and 1.982.

Every one of those is closer to its truth than the uncorrected reading. The tree at three recovers between 0.35 and 0.88 of an exponent depending on the curve, and the tree at two between 0.20 and 0.31. The size of the recovery is not in doubt; which curve to believe is.

Adding error to carry a fitted exponent back to none, at 5%. At 5% of error on every radius each replicate is refitted with more error added at four levels, and the dots are the means over 300 replicates, the tree at 3 above and the tree at 2 below. Uncorrected they read 2.763 and 1.935. Carried back to no error through the mean points, the tree at 3 reads 2.981 by a quadratic curve, 2.934 by a linear curve, 2.989 by a rational curve; the tree at 2 reads 1.993 by a quadratic one, 1.990 by a linear one, 1.993 by a rational one. The curve carried back is a choice the method does not make.
Fig. 2 The same construction at five per cent of error. The refitted points barely bend, and the three curves land within six hundredths of one another.

At five per cent, the choice barely matters

At five per cent of error the uncorrected readings are 2.763 and 1.935. Carried back, the tree at three reads 2.981 by the quadratic, 2.934 by the line and 2.989 by the rational; the tree at two reads 1.993, 1.990 and 1.993.

Where the refitted points lie close to a straight line, every reasonable curve through them agrees about where they came from. That is the regime in which the correction is a measurement rather than a choice, and it is also the regime in which the uncorrected duel was never in trouble.

The displacement is not the geometry

How much of a fitted exponent's displacement a quadratic correction removes, at every error. The share of each tree's displacement from its own truth that a quadratic correction removes. At 5% it removes 92% from the tree at 3, which reads 2.981 against 2.763 uncorrected, and 90% from the tree at 2; at 12% it removes 68% from the tree at 3, which reads 2.695 against 2.035 uncorrected, and 88% from the tree at 2; at 20% it removes 40% from the tree at 3, which reads 1.968 against 1.281 uncorrected, and 65% from the tree at 2; at 25% it removes 29% from the tree at 3, which reads 1.541 against 0.939 uncorrected, and 51% from the tree at 2.
Fig. 3 The share of each tree’s displacement from its own truth that a quadratic correction removes, at every error the duel was run at.

Measured as a share of the displacement from each tree’s own truth, the quadratic correction removes 92 per cent of the tree at three’s at five per cent of error, bringing it to 2.981 from 2.763. At twelve per cent it removes 68 per cent, bringing it to 2.695 from 2.035. At twenty it removes 40 and at twenty-five 29. For the tree at two the shares are 90, 88, 65 and 51.

So most of what the duel read as geometry at moderate error was the estimator’s. A fit of a noisy exponent junction by junction is pulled down by the noise, and a procedure that measures the pull can undo most of it without knowing anything about the tree.

What remains is not nothing

The correction weakens as the error grows, and the reason is in its construction. It extrapolates a trend measured between the error present and twice as much, and the trend it measures at large error has already bent — the fit has started to saturate — so a curve carried back along it under-corrects.

At twenty-five per cent, where the uncorrected tree at three read 0.939, the corrected one reads 1.541. That is a large recovery and it is still nowhere near three. The correction is a partial remedy whose partiality grows exactly where the duel’s inversion lived.

How much wider a corrected exponent's spread is than the uncorrected one, at every error. The standard deviation of the corrected readings over that of the uncorrected ones, over 300 replicates. At 2% it is ×1.57 for the tree at 3 and ×1.57 for the tree at 2; at 5% it is ×1.61 for the tree at 3 and ×1.56 for the tree at 2; at 12% it is ×2.04 for the tree at 3 and ×1.78 for the tree at 2; at 20% it is ×2.04 for the tree at 3 and ×2.05 for the tree at 2. The correction pays for what it removes in spread.
Fig. 4 The spread of the corrected readings over the spread of the uncorrected ones, for both trees, at every error. The correction’s price is paid here.

What a correction costs

Nothing is free in this. The standard deviation of the corrected readings over that of the uncorrected ones is ×1.57 at two per cent of error for both trees, ×1.61 and ×1.56 at five, ×2.04 and ×1.78 at twelve, and ×2.04 and ×2.05 at twenty.

The reason is general. A trend measured on noisy refits and extrapolated beyond its own range multiplies their scatter, because a small wobble in a slope at the measured points becomes a large wobble at a point outside them. Removing a bias by extrapolation is always bought with variance.

Why the tree at three widens more

At twelve per cent the tree at three’s spread doubles and the tree at two’s grows by three quarters. The difference follows from how far each is carried. The tree at three starts 0.97 below its truth and is carried back 0.66; the tree at two starts 0.32 below and is carried back 0.29.

An extrapolation magnifies the scatter of its refits by roughly how far past them it reaches, and the tree at three’s refitted curve is steeper — it falls from 2.035 to 1.245 across the added error where the tree at two falls from 1.676 to 1.250 — so the same four refits leave more uncertainty in where a steep line meets no error than a shallow one.

The overlap stays

The duel with every replicate corrected by a quadratic extrapolant. A tree at 3 and a tree at 2, 50 junctions of daughter ratio 0.6–1, 300 replicates at each error, every replicate corrected by a quadratic extrapolant. Shaded: the central 90% of the corrected readings; lines: the uncorrected means. The corrected intervals first overlap at 12%, the same error at which the uncorrected ones do. The uncorrected means cross at 20.5% and the corrected means do not cross anywhere on it; at 25% the corrected trees read 1.541 and 1.523.
Fig. 5 The duel with every replicate corrected by a quadratic curve: shaded, the central ninety per cent of the corrected readings; thin lines, the uncorrected means. The corrected bands close on each other at the same error, and their means never cross.

The corrected intervals first overlap at twelve per cent, the same error at which the uncorrected ones did. Each tree has been moved most of the way back towards its truth, and each interval has been widened by as much as the move was worth, and the two effects land the overlap at the same error.

That is the duel’s first number surviving a better estimator. Twelve per cent was not a property of a biased fitter; it is where fifty junctions measured to that error stop containing enough information to tell the two rules apart, however the fit is done.

The inversion goes

The second number does not survive. Uncorrected, the two means cross at 20.5 per cent, and above it the tree built at three reads lower than the tree built at two. Corrected, they do not cross at any error measured: at twenty-five per cent they read 1.541 and 1.523, the tree at three still — just — the higher.

So the inversion was the estimator’s. An estimator that does not invert does not need a displacement the same at every truth; it needs a displacement it can measure, and this one can.

At 12% the correction moves each tree towards its truth and widens it. Both trees at 12% of error, 300 replicates, before and after a quadratic correction. Uncorrected, the tree at 3 reads 2.035 [1.80, 2.30] and the tree at 2 1.676 [1.53, 1.84], overlapping. Corrected, they read 2.695 [2.21, 3.19] and 1.961 [1.70, 2.22], still overlapping: nearer their truths, and wider.
Fig. 6 Both trees at twelve per cent of error, before and after the correction. Each moves most of the way towards its own truth and widens as it goes, and the two corrected intervals still touch.

Twelve per cent in one frame

Uncorrected, the tree at three reads 2.035 with an interval of [1.80, 2.30] and the tree at two 1.676 with [1.53, 1.84]. Corrected, they read 2.695 [2.21, 3.19] and 1.961 [1.70, 2.22]. The means have moved apart by 0.37 and the intervals have widened by more, and the corrected intervals overlap by a hundredth.

A reader shown only the corrected numbers would conclude the tree at three is near three and the tree at two near two, which is right, and that a single study’s corrected exponent could still have come from either tree, which is also right.

The overlap as a ratio

Whether two intervals meet is one comparison: the distance between the means against the two half-widths that face each other, the lower half of the tree at three’s interval and the upper half of the tree at two’s. At twelve per cent, uncorrected, the means are 0.359 apart and the facing half-widths are 0.235 and 0.164, which sum to 0.399. The ratio is 0.90, and the intervals meet by 0.04.

Corrected, the means are 0.734 apart and the facing half-widths are 0.485 and 0.259, summing to 0.744. The ratio is 0.99, and the intervals meet by a hundredth. The correction doubled the distance between the means and very nearly doubled the widths facing each other, so the ratio rose by a tenth and stayed below one.

That is why the overlap did not move, and also how close it came to moving. A correction that widened its readings slightly less would have pulled the two trees apart at twelve per cent. Thirty-two refits a level is exactly such a setting, which is why the overlap’s error under it is given only as thirteen per cent or later.

Which of the duel’s numbers was which

The duel produced two thresholds and called the first a fact about a study design and the second a fact about the estimator. The correction confirms the second half exactly and qualifies the first: the overlap is a fact about the question — fitting a junction relation to measured radii — and not about any particular fitter.

That distinction decides what could improve it. A better fitter moves the inversion. Nothing that keeps asking the same junction-by-junction question moves the overlap, and the window that closes already showed that more junctions trade accuracy for confidence rather than buying both.

Three extrapolants carrying the tree at 3 back to no error, at every error measured. The mean corrected reading of the tree built at 3, by each of the three curves the correction can be carried back along. At 5% the quadratic gives 2.981, the linear 2.934 and the rational 3.511 with a spread of 2.75; at 12% the quadratic gives 2.695, the linear 2.380 and the rational 3.162 with a spread of 1.08; at 20% the quadratic gives 1.968, the linear 1.575 and the rational 5.217 with a spread of 10.69. The rational curve's pole lands at the edge of its search in 429 of the 600 corrections at 2% and in 39 at 25%; its readings that leave the axis are not drawn.
Fig. 7 The corrected reading of the tree built at three by each of the three curves the correction can be carried back along, at every error measured. The rational curve leaves the axis at most errors.

Correcting the mean, or averaging the corrections

There are two ways to state a corrected answer over three hundred replicates — correct each replicate and average the results, or average the refitted curves and correct the average — and for two of the three curves they are the same number exactly. A least-squares line or quadratic read at a fixed point is a weighted sum of the points it is fitted to, and a weighted sum of averages is the average of weighted sums.

So the quadratic reads 2.695 at twelve per cent either way, and the line reads 2.380 either way. The rational curve is not a weighted sum of its points: carried through the mean curve it reads 2.914, and averaged over its per-replicate corrections it reads 3.162. A method whose answer depends on the order of averaging and correcting has a second free choice hidden inside the first.

Three curves, three answers

The extrapolant is the method’s free choice and it is carried rather than made. The quadratic is the usual default and the one quoted above. A straight line under-corrects wherever the trend bends: 2.934 at five per cent and 2.380 at twelve, and its corrected means do cross, at 21.2 per cent.

The rational curve, exact for some error models, is unusable here. It reads 3.511 at five per cent with a spread of 2.75, 3.162 at twelve with a spread of 1.08, and 5.217 at twenty with a spread of 10.7, and its pole lands at the edge of its search in 429 of 600 corrections at two per cent of error. A correction built on it would report a tree at three reading five with a straight face.

And the grid decides as much

The corrected exponent at 12%, by how far the added error runs and which curve carries it back. The tree built at 3 at 12% of error, corrected on a grid of added error running to one, two or three times the error present, by each extrapolant: to 1 by quadratic, 2.777 with a spread of 0.423; to 1 by linear, 2.499 with a spread of 0.234; to 1 by rational, 7.300 with a spread of 12.114; to 2 by quadratic, 2.670 with a spread of 0.308; to 2 by linear, 2.369 with a spread of 0.192; to 2 by rational, 2.993 with a spread of 0.673; to 3 by quadratic, 2.600 with a spread of 0.254; to 3 by linear, 2.257 with a spread of 0.167; to 3 by rational, 2.910 with a spread of 0.463. Readings past 3.8 are drawn at the edge with their value beside them.
Fig. 8 At twelve per cent, the corrected reading of the tree built at three for grids of added error running to one, two and three times the error present, by each extrapolant, with a whisker of one standard deviation.

How far the added error runs is the same choice seen again. At twelve per cent the quadratic correction reads 2.777 on a grid running to once the error present, 2.670 to twice and 2.600 to three times; the straight line 2.499, 2.369 and 2.257. The rational curve reads 7.300 with a spread of 12.1 on the shortest grid and 2.910 on the longest.

Refitting more times at each level changes much less: two refits per level give 2.6945, eight give 2.6951 and thirty-two give 2.688. But thirty-two narrow the spread to 0.284 from 0.306, and at that setting the corrected intervals at twelve per cent no longer quite overlap — so the corrected overlap sits at twelve per cent with eight refits and at thirteen or later with thirty-two, and how much later is not measured.

The size of the error has to be known

The correction at 12% with the size of the error stated wrongly. The quadratic correction needs the size of the error it is correcting. Stated as a multiple of the true 12%: ×0.8 gives 2.466 for the tree at 3 and 1.857 for the tree at 2, overlapping; ×0.9 gives 2.577 for the tree at 3 and 1.906 for the tree at 2, overlapping; ×1 gives 2.695 for the tree at 3 and 1.961 for the tree at 2, overlapping; ×1.1 gives 2.820 for the tree at 3 and 2.022 for the tree at 2, overlapping; ×1.2 gives 2.949 for the tree at 3 and 2.087 for the tree at 2. Understating the error leaves the correction short and overstating it carries it past the truth, each with the spread of a correction that was done right.
Fig. 9 The quadratic correction at a true twelve per cent, with the error it is told to correct stated as eight tenths to twelve tenths of the truth.

The correction needs the error’s size, and a wrong size is a wrong answer delivered with a correction’s confidence. Told that a true twelve per cent is 9.6 per cent, the correction reads the tree at three as 2.466; told 10.8, 2.577; told 13.2, 2.820; told 14.4, 2.949. The tree at two moves from 1.857 to 2.087 across the same range.

A tenth of the error misstated moves the corrected answer by about 0.12 of an exponent — nearly as far as the difference between the quadratic and the linear curve. So the precision the correction needs in the error’s size is a precision the duel’s own closing sentence said nobody measures to: the instrument’s error, reported beside the exponent.

Reading a published exponent after this

Take an exponent of 2.0 reported from fifty good junctions measured to twelve per cent. Uncorrected, the duel says it is consistent with a tree built at three — whose readings at that error run from 1.80 to 2.30 — and with one built at two. Corrected by a quadratic curve with the error stated correctly, the same kind of reading would be carried towards whichever truth produced it, but the corrected intervals of the two trees still meet.

So the correction changes what a reported number means without changing whether it can decide between the rules. That is worth saying plainly, because a corrected number near three reads as a confirmation of Murray’s law in a way an uncorrected 2.0 never did, and at twelve per cent of error on fifty junctions it is no more of one.

What the correction cannot touch

Simulation-extrapolation removes the displacement that adding random error produces, because that is the displacement it measures. A systematic error does not move when random error is added to it, so the correction leaves it exactly where it was. A parent read fat at the fork lowers a corrected exponent by the same first-order amount it lowers an uncorrected one.

And it cannot invent information. The overlap it leaves is the part of the duel that belongs to the question, and a correction of the answer is not a change of question.

What this corrects

The duel said no better estimator removes the displacement by being better, because the displacement is the geometry. That sentence has been repaired where it was written. The displacement at moderate error is mostly the estimator’s, a standard correction removes most of it, and the inversion above twenty and a half per cent goes with it. What survives is the overlap, and a claim that survives its own best attempted refutation is a stronger claim than it was before.

What this does not establish

That the correction is usable on a real tree. It needs the error’s size to better than a tenth, and it assumes the error is of the kind it adds — Gaussian, relative and independent on each radius — which the control a survey would need would have to establish first. At the largest errors the added error sometimes makes a radius non-positive: fourteen of 2.88 million draws were redrawn at eighteen per cent and 2,304 at twenty-five, and one replicate was refused.

The refits also start to reach the edge of the range the fitter searches. Nine refits of 2.88 million were pinned at an end of it at sixteen per cent, 431 at twenty and 3,473 at twenty-five. They are counted rather than hidden, and they are part of why the correction’s grip weakens exactly where the inversion was: a refit pinned at the edge of the search reports the search, and a trend drawn through such refits is flatter than the fit’s real response to added error. Fitting the exponent with a wider search is one more setting the corrected numbers at the largest errors depend on.

It also tests one family of corrections. A different estimator — a fit that models the error explicitly, or one that weights junctions by their leverage — is not tested, and nothing here says what it would do to the overlap.

What would withdraw it

A quadratic correction that leaves most of the displacement in place at five per cent. Corrected intervals whose overlap moves by more than a couple of steps in the error. Corrected means that still cross below twenty-five per cent. A faster fitter that does not land where the original fitter lands. Each is checked whenever the measurement runs.

A different question

If the overlap belongs to the question, the remedy is to ask a different one of the same radii. Every branch in a tree carries a number of tips, and under either rule with tips of one size a branch’s radius is that count to the power one over the exponent. A count has no measurement error in it, so the noise sits on one side of the regression only. The next measurement fits the same measured radii against the count, and its test is whether a count that carries no error keeps the two trees apart past the twelve per cent at which every junction-by-junction estimator, corrected or not, cannot.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

BiasBranching exponentDa Vinci's ruleDiscriminationEvidenceFittingHonest limitsIdentifiabilityInterval estimateMeasurement errorMurray's lawSelf-correction