Concept

Bias — where it appears

A systematic displacement of an estimate, which more data does not reduce because it is not sampling variation. Several of the quantities here carry one that reads as a result: a variance of block means is biased low in the number of blocks, and the bias points the same way at every setting.

Named by 16 essays across 5 fields — each of them below, with the objects they name alongside it.

The uninformative sample does not say so — it says something else. Above: how many junctions of a given asymmetry it takes to distinguish an exponent of 3 from an exponent of 2, at 2% measurement error on each radius. An even fork needs one; a junction whose small daughter is a twentieth of the large one needs 2195. Below: 50 junctions from each end of the range, on a synthetic tree built at exactly 3. The even forks return 3.01; the twigs return 1.69, with an interval no wider — because 47% of them measure as a parent thinner than its own larger daughter, and dropping those keeps only the half where the noise ran the right way.

A sample that is confidently wrong

Fifty lopsided junctions from a tree built at an exponent of exactly 3 return 1.7, with an interval that excludes 3 and excludes 2 as well. The sample carrying almost no information does not give a wide answer — it gives a narrow wrong one, and the cause is a selection nobody applies on purpose.

branching · Exponent
A tree built at 3, measured to 2%, reads 2.957 on the informative band. The exponent recovered from 100 junctions of a tree built at exactly 3, against the relative error placed independently on the parent and on both daughters — half a per cent is a machined section under a microscope, one to two per cent is callipers on a clean branch, five is a branch that is not round, ten is a radius read off a photograph. Each line is a band of daughter ratio and the whiskers are the central 90% of 300 replicate samples. Every band is displaced downward at every error and never upward: at 2% the informative band read 2.957 and the informative band read 2.957. Below, the same rows with the displacement and the spread drawn as separate bars, because only one of them falls when more junctions are measured.

The exponent an error moves

Every real measurement of a branch radius carries error and no synthetic tree does, so the question is what a symmetric error does to a fitted exponent. It does two things — a bias and a spread — and the bias runs downward at every error level and in every band, by an amount derivable from the daughter ratios alone.

branching · Exponent error
One junction's bias moves 1.36-fold across the range its leverage moves 308,352-fold, at an exponent of 3. The bias a single junction of daughter ratio γ contributes to a least-squares fit, as a multiple of the squared measurement error, drawn against the leverage that junction carries — both computed from the expansion about a true exponent of 3 rather than fitted to anything. Across the whole range from an even fork to a twentieth the bias moves by a factor of 1.36 and the leverage by a factor of 308,352, and the uninformative junction contributes the larger share: 6.00σ² at γ = 0.05 against 4.50σ² at an even fork.

The fragile junctions are the informative ones

That is the obvious worry once the radii are uncertain, and it is false. Across the whole range of asymmetry a junction's contribution to the bias moves by a factor of 1.36 while its leverage moves by a factor of 308,352, so the junction that says nothing damages the answer as badly as the one that says everything — and a sample is spoiled by counting rather than by weight.

branching · Exponent error
Between 3% and 5% of radius error, no sample size answers — 50 junctions among them. One row per error level. The pale bar is the sample sizes whose interval is narrow enough to state a claim from — half-width under ±0.25 and excluding 2 — and it starts where precision arrives. The second bar is the sample sizes whose interval still contains the 3 the tree was built at, and it ends where the displacement overtakes the width. Where the two overlap there is a usable window; at 5%, 7%, 10% they do not overlap at all, so below 50 junctions the answer is too wide to state and above 30 it no longer contains the truth.

The window that closes

The spread of a fitted branching exponent falls as the reciprocal root of the sample and its displacement does not fall at all, so there is a count past which every further junction buys confidence and no accuracy. Between three and five per cent of radius error the count arrives before the answer does, and no sample size both states a claim and contains the truth.

branching · Exponent error
Murray's law and Da Vinci's become one measurement at 12% of radius error on the informative band. Two trees measured the same way: 50 junctions of daughter ratio 0.6–1, 300 replicate samples at each of 15 error levels, one tree built at exactly 3 and one at exactly 2. Each shaded band is the central 90% of the recovered exponents. Both run downward, but the tree at 3 runs down faster — -113σ² against -29σ² — so the two close on each other.  At 11% they are still apart; at 12% the bands overlap and one study's answer could have come from either tree; at 20.5% the means cross, and above it a tree built at 3 measures lower than a tree built at 2.

Where three and two become one

Two trees, one built to obey Murray's law and one to obey Da Vinci's, are measured through the same fifty junctions with the same instrument. At twelve per cent of error on each radius the two answers overlap, and above twenty and a half the tree built at three measures lower than the tree built at two.

branching · Exponent error
What insisting on a fork angle costs: 43.3 degrees for one per cent of the network. The three ends are held where the optimum wants them, the branch point is moved everywhere inside them, and the cheapest network at each total angle is kept. The optimum sits at 74.93°. Everything within one per cent of the least cost runs 55.9° to 99.2° — a span of 43.3°, which read back as exponents covers 2.44 to 5.34 — and within a tenth of a per cent it still runs 13.6°. The prediction is steep in the exponent and the cost is nearly flat in the angle; they are the same curve read along its two axes.

An optimum too flat to reach

One per cent of a branching network's cost buys forty-three degrees of fork angle, covering exponents from 2.44 to 5.34, while the angle the theory predicts moves only fourteen and a half degrees across every daughter ratio there is. The prediction is steep and the cost is flat, and those are the same curve read along its two axes.

branching · Fork angle
The statistic everybody reports is the one that cannot vary. Six arrangements of 900 points, from a whorled lattice to a set with no rule in it. The mean number of sides per cell is 5.97–6.04 on all six, because Euler's formula forces it. The mean squared departure from six runs from 0.023 to 1.83 — a factor of 79 — and the most hexagonal tissue in the set is the whorled one, at a rational angle.

What a summary throws away

Four statistics this collection has relied on turn out to be incapable of varying with the thing they describe — one is invariant to shuffling, one is fixed by a theorem, one is a parameter that stopped mattering, one is a fitted number selected into being wrong. In each case the second statistic was free and nobody had taken it.

wrong · Second statistic
What a swelling does to the exponent read from the informative band, by where the swelling is. 50 junctions of daughter ratio 0.6–1, built at exactly 3 and read with no random error, but with one or more radii measured fat. A parent read fat lowers the reading: 1% gives 2.864, 3% gives 2.631, 10% gives 2.075, and at 11.3% the tree reads Da Vinci's 2. Daughters read fat raise it: 1% gives 3.150, 3% gives 3.499; from 8% some junctions have a daughter measured wider than their parent, which no exponent balances, and the line stops. All three radii read fat by one factor return 3.000 at every swelling — the unswollen reading exactly.

A swelling at the fork

A branch thickens where it forks, so a parent measured just below a junction and daughters measured just above it carry three different amounts of the same swelling. A swelling that fattens all three alike moves a fitted exponent by exactly nothing. A parent read one per cent fat moves it by as much as 3.6 per cent of random error on every radius, in a sign known in advance, and a tenth of a radius turns a tree built at Murray's three into one that reads Da Vinci's two with no noise at all. Added to the noise, it does not bring the two rules together any sooner: the two errors do not add.

branching · Exponent error
Seven fractions with one neighbour distance and every denominator. Each member of a matched set drawn on its own stretch of the divergence axis, 0.5° either side of itself, with the nearest other rational marked. The distances are 0.3158°, 0.3158°, 0.3117°, 0.3069°, 0.3117°, 0.3077°, 0.3064° — a spread of 3.1% — while the denominators run 19, 20, 21, 23, 33, 45, 47, a factor of 2.47. That is the construction this thread needed. Every instrument for the width of a disorder dip has a free parameter set by how close the neighbour is, so a hypothesis about the neighbourhood cannot be tested by varying the neighbourhood; on this set the neighbourhood is held fixed and the arithmetic of the fraction is what varies.

Fractions with the same neighbours

Every instrument this collection has for the width of a disorder dip has a free parameter set by how close the next rational sits — which makes a hypothesis about the neighbourhood untestable with any of them. The repair is not a better instrument. It is a set of fractions whose neighbourhoods are identical and whose denominators are not, and the arithmetic supplies twenty-four of them.

tissue · Second statistic
Adding error to carry a fitted exponent back to none, at 12%. At 12% of error on every radius each replicate is refitted with more error added at four levels, and the dots are the means over 300 replicates, the tree at 3 above and the tree at 2 below. Uncorrected they read 2.035 and 1.676. Carried back to no error through the mean points, the tree at 3 reads 2.695 by a quadratic curve, 2.380 by a linear curve, 2.914 by a rational curve; the tree at 2 reads 1.961 by a quadratic one, 1.877 by a linear one, 1.982 by a rational one. The curve carried back is a choice the method does not make.

A correction that keeps the overlap

The duel between a tree built at Murray's exponent and one built at Da Vinci's ended by saying the displacement is the geometry, and that no better estimator removes it. Correcting every replicate by simulation-extrapolation removes 92 per cent of the tree at three's displacement at five per cent of error and 68 per cent at twelve, and the error at which the two means cross leaves the measured range altogether. It pays in spread — the corrected readings are twice as wide at twelve per cent — so the error at which the two trees' intervals overlap does not move. Of the duel's two numbers, the inversion was the estimator's and the overlap is the question's.

branching · Exponent error
Both edges of the front heal; the middle of it does not. The same removals, followed for 300 organs each. A cut one to three places back is undone within fifty organs and a cut ten to thirteen places back within sixty. A cut in between is never undone: the divergence sequence settles into an exactly repeating cycle of  angles and holds it for the rest of the run. The rule corrects a displacement and cannot correct a deletion.

A period the grid invented

A wrecked stem was reported as settling into a repeating block of three angles — 219.84°, 220.31°, 220.78° — which is the smaller of its two spiral counts and would have confirmed a standing prediction. Those three numbers are three consecutive samples of the azimuth grid. There is no block; there is a constant the grid cannot write down, and the routine that found the block was working perfectly.

wrong · Instrument ceiling
Hold the neighbourhood and the denominator stops mattering. The equivalent width of the disorder dip — the area of the deficit divided by its own depth — for seven fractions whose nearest neighbours sit at the same distance and whose denominators run from 19 to 47. Each is measured at a head size chosen so that all of them share one scaled unit, which makes a window in scaled units the same window in degrees and the same fraction of the way to the neighbour for every member. At a window of 25 the seven widths are 38.8, 38.8, 39.0, 39.0, 38.7, 39.0, 39.0 — a spread of ×1.010 across a factor of 2.47 in denominator. The lines separate as the window widens, to ×1.147 at 200, and when they do they order by denominator rather than by crowding. So the residual this thread carried was the window: hold it and there is nothing left that belongs to the fraction.

The residual was the window

After the depth and the q over n squared scale are taken out of a disorder dip, something looked left over and looked ordered by how crowded the fraction's neighbourhood is. Measured on fractions whose neighbourhoods are identical by construction, seven widths across a factor of two and a half in denominator agree to one per cent. There is no residual; there was a comparison made at different effective windows.

wrong · Second statistic
Five fractions with one neighbour distance and every denominator. Each member of a matched set drawn on its own stretch of the divergence axis, 0.5° either side of itself, with the nearest other rational marked. The distances are 0.3651°, 0.3529°, 0.3692°, 0.3640°, 0.3557° — a spread of 4.6% — while the denominators run 17, 20, 25, 43, 44, a factor of 2.59. That is the construction this thread needed. Every instrument for the width of a disorder dip has a free parameter set by how close the neighbour is, so a hypothesis about the neighbourhood cannot be tested by varying the neighbourhood; on this set the neighbourhood is held fixed and the arithmetic of the fraction is what varies.

Matching instead of correcting

Two rounds of work failed on one question because every instrument's free parameter was set by the thing under test. The repair was not a better instrument or a model of the bias: it was choosing what to compare so that the confound could not vary. That move is available in four other places here, and three of them have already used it without anybody naming it.

wrong · Instrument ceiling
What one recount reads after 34/55 was first read as 34/54, for each belief the counter aims from. A head whose true pair is 34/55, first read as 34/54 with closing errors spread over 7.2°, recounted once; the two closings of the head are correlated by 0.9. Each bar is one counter, split into the recount reading 34/55, announcing itself again, and reading another pair silently: unaimed, aiming +0.0°, 41.8% right, 55.2% announced, 3.0% silent; expects Fibonacci, aiming +3.8°, 69.6% right, 30.2% announced, 0.2% silent; expects Fibonacci, capped at 2σ, aiming +3.8°, 69.6% right, 30.2% announced, 0.2% silent; expects any spiral, aiming −4.5°, 5.5% right, 62.7% announced, 31.8% silent; expects nothing, aiming +0.3°, 44.7% right, 52.8% announced, 2.5% silent.

The recount aims where the counter expects

A counter who recounts an announced reading knows it went wrong, and if the same habit spoils both counts of a head, the first error says where to aim the second. But the reading alone does not say which way the first erred: a reading of 34/54 is as well explained by a whorled 34/54 read right, or by 34/53 read long, as by 34/55 read short. The direction comes from what the counter expects. Expecting Fibonacci, an aimed recount at 7.2° and a habit correlated at 0.9 reads 34/55 69.6 per cent of the time where an unaimed one reads it 41.8, and the census needs seven specimens rather than fifteen. Expecting only a spiral, it aims the wrong way and reads 34/55 5.5 per cent of the time. The belief that helps is the hypothesis the census is testing: uncapped, it reads the geometry's whorled 3/6 heads as 3/5 and a census of a hundred rejects a true null 40 per cent of the time; capped, it still reads a silent 33/53 as Fibonacci twice as often. And no aimed recount spends fewer counts than 13/21 counted once.

wrong · Sample size
How often a stem settles, at three samplings of the starting angle. The share of runs that reach a lattice, over the whole table of four falloff exponents by eight rises. The upper three bars are the pooled share at nine, twenty and forty starting angles; the lower three are the share for each group of angles on its own. The pooled figure falls 40.6 to 32.3 to 30.0 per cent, by 8.3 points and then 2.3, so it is converging. The eleven angles added at twenty and the twenty added at forty settle at 25.6 and 27.7 per cent, which differ by less than their own error.

Forty angles, and a limit

Nine starting angles turned out to be a biased sample of the circle, and doubling to twenty said by how much. Doubling again says the estimate is converging — to a smaller correction than one doubling extrapolated to.

emergence · Settling
Six handovers relocated, in steps of the sweep that recorded them. The ladder finds a handover by stepping at a ratio of one per cent, which never lands on the grid the rises are named on, so a recorded handover is the nearest rise the sweep visited to a crossing nobody had located. This is the difference, in units of the sweep's own step at that rise: 0.082 to 0.835, mean 0.374. All six are positive and all six are inside a single step of the sweep, and both of those are predictions rather than summaries: the ladder reports the first rise it visits at which the ordering has already changed, so the rise it records must sit on the fine side of a crossing and within one of its own steps. The dashed rule is one step of the sweep, which is the bound the sampling predicts.

The handovers corrected

Six recorded handovers, relocated to the grid against where a one-per-cent sweep put them: all six sit on the fine side of a crossing and all six inside a single sweep step. Nothing about the rung explains the size of the discrepancy, which is what a sampling artefact is supposed to look like.

cylinder · Handover grid

Named alongside it

The objects these essays reach for when they reach for this one.

Honest limitsBranching exponentMeasurement errorMurray's lawSummary statisticDa Vinci's ruleMeasurementClaim testingSamplingDiscriminationNoiseSample size

All concepts