A correction that keeps the overlap
Worth reading first: The exponent an error moves · Fitting the exponent · The cube law.
Two trees measured through the same fifty junctions — one built at Murray’s exponent of three, one at Da Vinci’s two — overlap at twelve per cent of relative error on every radius, and above twenty and a half per cent the tree built at three reads lower than the tree built at two. That essay drew a conclusion from the inversion: an estimator that did not invert would need a displacement the same at every truth, “and the displacement is the geometry”.
That is a claim about every estimator, and one estimator is enough to test it. The displacement itself is a known function of the error’s size, which is exactly the situation a standard correction exists for.
Adding error to remove it
Simulation-extrapolation is the standard correction for measurement error of a known size, and its logic fits in a sentence. If adding error moves the answer in a predictable way, then measure how it moves by adding more, and carry the trend back past the error already present to the point where there would be none.
It needs nothing but the sample in hand and the size of the error. It is not a model of branching, and it knows nothing about Murray or Da Vinci; it treats the fitter as a black box whose answer drifts with the noise, and extrapolates the drift.
On the duel’s own replicates
The correction is run on exactly the replicates the duel used — three hundred at each error, fifty junctions of daughter ratio 0.6 to 1 — so its uncorrected rows are the duel’s rows, and every difference below is the correction’s.
Each replicate is refitted with added error at four levels, eight times at each, so every error level costs 2.88 million refits. A faster golden-section fitter lands within the original fitter’s own resolution of where the original fitter lands, checked on 120 samples at three errors and both trees, so what is corrected is the duel’s own estimator and not a substitute for it.
At twelve per cent
Uncorrected, the two trees read 2.035 and 1.676. As error is added the means fall steadily. Carried back to no error through the mean points, the tree at three reads 2.695 by a quadratic curve, 2.380 by a straight line and 2.914 by a rational curve; the tree at two reads 1.961, 1.877 and 1.982.
Every one of those is closer to its truth than the uncorrected reading. The tree at three recovers between 0.35 and 0.88 of an exponent depending on the curve, and the tree at two between 0.20 and 0.31. The size of the recovery is not in doubt; which curve to believe is.
At five per cent, the choice barely matters
At five per cent of error the uncorrected readings are 2.763 and 1.935. Carried back, the tree at three reads 2.981 by the quadratic, 2.934 by the line and 2.989 by the rational; the tree at two reads 1.993, 1.990 and 1.993.
Where the refitted points lie close to a straight line, every reasonable curve through them agrees about where they came from. That is the regime in which the correction is a measurement rather than a choice, and it is also the regime in which the uncorrected duel was never in trouble.
The displacement is not the geometry
Measured as a share of the displacement from each tree’s own truth, the quadratic correction removes 92 per cent of the tree at three’s at five per cent of error, bringing it to 2.981 from 2.763. At twelve per cent it removes 68 per cent, bringing it to 2.695 from 2.035. At twenty it removes 40 and at twenty-five 29. For the tree at two the shares are 90, 88, 65 and 51.
So most of what the duel read as geometry at moderate error was the estimator’s. A fit of a noisy exponent junction by junction is pulled down by the noise, and a procedure that measures the pull can undo most of it without knowing anything about the tree.
What remains is not nothing
The correction weakens as the error grows, and the reason is in its construction. It extrapolates a trend measured between the error present and twice as much, and the trend it measures at large error has already bent — the fit has started to saturate — so a curve carried back along it under-corrects.
At twenty-five per cent, where the uncorrected tree at three read 0.939, the corrected one reads 1.541. That is a large recovery and it is still nowhere near three. The correction is a partial remedy whose partiality grows exactly where the duel’s inversion lived.
What a correction costs
Nothing is free in this. The standard deviation of the corrected readings over that of the uncorrected ones is ×1.57 at two per cent of error for both trees, ×1.61 and ×1.56 at five, ×2.04 and ×1.78 at twelve, and ×2.04 and ×2.05 at twenty.
The reason is general. A trend measured on noisy refits and extrapolated beyond its own range multiplies their scatter, because a small wobble in a slope at the measured points becomes a large wobble at a point outside them. Removing a bias by extrapolation is always bought with variance.
Why the tree at three widens more
At twelve per cent the tree at three’s spread doubles and the tree at two’s grows by three quarters. The difference follows from how far each is carried. The tree at three starts 0.97 below its truth and is carried back 0.66; the tree at two starts 0.32 below and is carried back 0.29.
An extrapolation magnifies the scatter of its refits by roughly how far past them it reaches, and the tree at three’s refitted curve is steeper — it falls from 2.035 to 1.245 across the added error where the tree at two falls from 1.676 to 1.250 — so the same four refits leave more uncertainty in where a steep line meets no error than a shallow one.
The overlap stays
The corrected intervals first overlap at twelve per cent, the same error at which the uncorrected ones did. Each tree has been moved most of the way back towards its truth, and each interval has been widened by as much as the move was worth, and the two effects land the overlap at the same error.
That is the duel’s first number surviving a better estimator. Twelve per cent was not a property of a biased fitter; it is where fifty junctions measured to that error stop containing enough information to tell the two rules apart, however the fit is done.
The inversion goes
The second number does not survive. Uncorrected, the two means cross at 20.5 per cent, and above it the tree built at three reads lower than the tree built at two. Corrected, they do not cross at any error measured: at twenty-five per cent they read 1.541 and 1.523, the tree at three still — just — the higher.
So the inversion was the estimator’s. An estimator that does not invert does not need a displacement the same at every truth; it needs a displacement it can measure, and this one can.
Twelve per cent in one frame
Uncorrected, the tree at three reads 2.035 with an interval of [1.80, 2.30] and the tree at two 1.676 with [1.53, 1.84]. Corrected, they read 2.695 [2.21, 3.19] and 1.961 [1.70, 2.22]. The means have moved apart by 0.37 and the intervals have widened by more, and the corrected intervals overlap by a hundredth.
A reader shown only the corrected numbers would conclude the tree at three is near three and the tree at two near two, which is right, and that a single study’s corrected exponent could still have come from either tree, which is also right.
The overlap as a ratio
Whether two intervals meet is one comparison: the distance between the means against the two half-widths that face each other, the lower half of the tree at three’s interval and the upper half of the tree at two’s. At twelve per cent, uncorrected, the means are 0.359 apart and the facing half-widths are 0.235 and 0.164, which sum to 0.399. The ratio is 0.90, and the intervals meet by 0.04.
Corrected, the means are 0.734 apart and the facing half-widths are 0.485 and 0.259, summing to 0.744. The ratio is 0.99, and the intervals meet by a hundredth. The correction doubled the distance between the means and very nearly doubled the widths facing each other, so the ratio rose by a tenth and stayed below one.
That is why the overlap did not move, and also how close it came to moving. A correction that widened its readings slightly less would have pulled the two trees apart at twelve per cent. Thirty-two refits a level is exactly such a setting, which is why the overlap’s error under it is given only as thirteen per cent or later.
Which of the duel’s numbers was which
The duel produced two thresholds and called the first a fact about a study design and the second a fact about the estimator. The correction confirms the second half exactly and qualifies the first: the overlap is a fact about the question — fitting a junction relation to measured radii — and not about any particular fitter.
That distinction decides what could improve it. A better fitter moves the inversion. Nothing that keeps asking the same junction-by-junction question moves the overlap, and the window that closes already showed that more junctions trade accuracy for confidence rather than buying both.
Correcting the mean, or averaging the corrections
There are two ways to state a corrected answer over three hundred replicates — correct each replicate and average the results, or average the refitted curves and correct the average — and for two of the three curves they are the same number exactly. A least-squares line or quadratic read at a fixed point is a weighted sum of the points it is fitted to, and a weighted sum of averages is the average of weighted sums.
So the quadratic reads 2.695 at twelve per cent either way, and the line reads 2.380 either way. The rational curve is not a weighted sum of its points: carried through the mean curve it reads 2.914, and averaged over its per-replicate corrections it reads 3.162. A method whose answer depends on the order of averaging and correcting has a second free choice hidden inside the first.
Three curves, three answers
The extrapolant is the method’s free choice and it is carried rather than made. The quadratic is the usual default and the one quoted above. A straight line under-corrects wherever the trend bends: 2.934 at five per cent and 2.380 at twelve, and its corrected means do cross, at 21.2 per cent.
The rational curve, exact for some error models, is unusable here. It reads 3.511 at five per cent with a spread of 2.75, 3.162 at twelve with a spread of 1.08, and 5.217 at twenty with a spread of 10.7, and its pole lands at the edge of its search in 429 of 600 corrections at two per cent of error. A correction built on it would report a tree at three reading five with a straight face.
And the grid decides as much
How far the added error runs is the same choice seen again. At twelve per cent the quadratic correction reads 2.777 on a grid running to once the error present, 2.670 to twice and 2.600 to three times; the straight line 2.499, 2.369 and 2.257. The rational curve reads 7.300 with a spread of 12.1 on the shortest grid and 2.910 on the longest.
Refitting more times at each level changes much less: two refits per level give 2.6945, eight give 2.6951 and thirty-two give 2.688. But thirty-two narrow the spread to 0.284 from 0.306, and at that setting the corrected intervals at twelve per cent no longer quite overlap — so the corrected overlap sits at twelve per cent with eight refits and at thirteen or later with thirty-two, and how much later is not measured.
The size of the error has to be known
The correction needs the error’s size, and a wrong size is a wrong answer delivered with a correction’s confidence. Told that a true twelve per cent is 9.6 per cent, the correction reads the tree at three as 2.466; told 10.8, 2.577; told 13.2, 2.820; told 14.4, 2.949. The tree at two moves from 1.857 to 2.087 across the same range.
A tenth of the error misstated moves the corrected answer by about 0.12 of an exponent — nearly as far as the difference between the quadratic and the linear curve. So the precision the correction needs in the error’s size is a precision the duel’s own closing sentence said nobody measures to: the instrument’s error, reported beside the exponent.
Reading a published exponent after this
Take an exponent of 2.0 reported from fifty good junctions measured to twelve per cent. Uncorrected, the duel says it is consistent with a tree built at three — whose readings at that error run from 1.80 to 2.30 — and with one built at two. Corrected by a quadratic curve with the error stated correctly, the same kind of reading would be carried towards whichever truth produced it, but the corrected intervals of the two trees still meet.
So the correction changes what a reported number means without changing whether it can decide between the rules. That is worth saying plainly, because a corrected number near three reads as a confirmation of Murray’s law in a way an uncorrected 2.0 never did, and at twelve per cent of error on fifty junctions it is no more of one.
What the correction cannot touch
Simulation-extrapolation removes the displacement that adding random error produces, because that is the displacement it measures. A systematic error does not move when random error is added to it, so the correction leaves it exactly where it was. A parent read fat at the fork lowers a corrected exponent by the same first-order amount it lowers an uncorrected one.
And it cannot invent information. The overlap it leaves is the part of the duel that belongs to the question, and a correction of the answer is not a change of question.
What this corrects
The duel said no better estimator removes the displacement by being better, because the displacement is the geometry. That sentence has been repaired where it was written. The displacement at moderate error is mostly the estimator’s, a standard correction removes most of it, and the inversion above twenty and a half per cent goes with it. What survives is the overlap, and a claim that survives its own best attempted refutation is a stronger claim than it was before.
What this does not establish
That the correction is usable on a real tree. It needs the error’s size to better than a tenth, and it assumes the error is of the kind it adds — Gaussian, relative and independent on each radius — which the control a survey would need would have to establish first. At the largest errors the added error sometimes makes a radius non-positive: fourteen of 2.88 million draws were redrawn at eighteen per cent and 2,304 at twenty-five, and one replicate was refused.
The refits also start to reach the edge of the range the fitter searches. Nine refits of 2.88 million were pinned at an end of it at sixteen per cent, 431 at twenty and 3,473 at twenty-five. They are counted rather than hidden, and they are part of why the correction’s grip weakens exactly where the inversion was: a refit pinned at the edge of the search reports the search, and a trend drawn through such refits is flatter than the fit’s real response to added error. Fitting the exponent with a wider search is one more setting the corrected numbers at the largest errors depend on.
It also tests one family of corrections. A different estimator — a fit that models the error explicitly, or one that weights junctions by their leverage — is not tested, and nothing here says what it would do to the overlap.
What would withdraw it
A quadratic correction that leaves most of the displacement in place at five per cent. Corrected intervals whose overlap moves by more than a couple of steps in the error. Corrected means that still cross below twenty-five per cent. A faster fitter that does not land where the original fitter lands. Each is checked whenever the measurement runs.
A different question
If the overlap belongs to the question, the remedy is to ask a different one of the same radii. Every branch in a tree carries a number of tips, and under either rule with tips of one size a branch’s radius is that count to the power one over the exponent. A count has no measurement error in it, so the noise sits on one side of the regression only. The next measurement fits the same measured radii against the count, and its test is whether a count that carries no error keeps the two trees apart past the twelve per cent at which every junction-by-junction estimator, corrected or not, cannot.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- The band decides the answer — both name branching exponent, da vinci's rule, discrimination, evidence, fitting, honest limits, identifiability, measurement error, murray's law
- The fragile junctions are the informative ones — both name bias, branching exponent, da vinci's rule, honest limits, interval estimate, measurement error, murray's law
- An optimum too flat to reach — both name bias, branching exponent, discrimination, honest limits, measurement error, murray's law
- A disturbance with a memory — both name discrimination, evidence, honest limits, measurement error, self-correction
- A sample that is confidently wrong — both name bias, branching exponent, da vinci's rule, fitting, murray's law
- Matching instead of correcting — both name bias, discrimination, evidence, honest limits, identifiability
Named objects
A flat tag is an object no other essay names yet.
BiasBranching exponentDa Vinci's ruleDiscriminationEvidenceFittingHonest limitsIdentifiabilityInterval estimateMeasurement errorMurray's lawSelf-correction