The error budget for a nautilus
Worth reading first: The nautilus question · What the centre costs.
The nautilus question settled that a nautilus does not grow as a golden spiral: 6.854 per turn against a measured 3.2, a factor of 2.142, and not a rounding error. Every essay on shells since has priced one way of getting that measurement wrong — an assumed centre, a short span of arc, a pair of dividers, an uneven clock, an oblique photograph, a miscounted septum, a shell that changed as it grew.
Each was a separate result and each returned a number. Nobody has added them up. Adding them up is the only thing that says what a section can decide, and it does not say what it was expected to.
Seven numbers, none of them typed in
The entries are computed rather than quoted, which matters more here than anywhere else in the collection: a figure repeated in prose can drift from the measurement it came from, and this essay is entirely about what happens when seven such figures are put beside each other. Each is imported from the file that produced it and recomputed every time the page is built.
An assumed centre displaced by a quarter of the innermost whorl’s radius, over every direction, costs 6.9 per cent — what the centre costs. The same displaced centre read over spans of one turn and more leaves a root-mean-square error of 22.4 per cent. A pair of dividers walked along the curve at the step counts a section allows reads high by up to 278.2 per cent — a measurement in steps. A curve sampled at equal intervals of time, where the marks pass half a turn and the fit’s unwrapping fails, gives 4.1 per cent — a spiral with no clock. A camera ten degrees off the normal, read over half a turn, gives 3.0 per cent. One septum miscounted at thirteen a whorl gives 9.1 — what the septa count. And a shell whose expansion rose from 2.8 to 3.6 a turn, read as one number, is 9.0 per cent away from what its last whorl is doing.
The sum is larger than the gap
Added without cancellation those seven come to 332.6 per cent. The gap the golden claim asks the measurement to resolve is 114.2.
That is not an arithmetic slip and it does not go away under a friendlier way of adding. In quadrature — the optimistic reading, which assumes the sources are independent and mostly cancel — the budget is 279.6 per cent, still more than twice the gap.
So the honest first conclusion is the one nobody wanted: the errors measured here do not, together, refuse the golden-spiral claim. A person reading a section badly enough, in ways every one of which has been documented here, could get 6.854 out of a shell grown at 3.2.
One entry decides it
Setting aside the centre leaves 325.7 per cent. The span, 310.3. The septum, 323.5. The changing shell, 323.6. The clock, 328.6. The oblique view, 329.7. None of them changes anything.
Setting aside the dividers leaves 54.4, which is less than half the gap.
So the whole verdict rests on one entry, and it is worth saying exactly which one. A measurement in steps found that walking a pair of dividers along a shell’s spiral is the oldest way to measure it, that it is the best route available once there are enough steps, and that under a count which follows exactly from the geometry it inflates the answer instead. It also found that it is the only route measured in this collection that pushes a nautilus towards a golden spiral rather than away from it.
The budget has now put a number on how far it pushes. It is not a contribution to the error; it is larger than the quantity in dispute.
What the dividers entry actually is
Two hundred and seventy-eight per cent is a large enough number to be worth opening rather than accepting, and opening it makes the entry sharper rather than softer.
The dividers reading picks a step length and walks it along the curve, counting how many steps fall in a turn. At a span of a turn and a half the best reading is 3.2000 at twenty-four steps, exact to fifteen decimal places. At two turns the best is 4.1176 at nine steps. At three turns it is 4.5029; at four, 9.3818; at five, 12.1032 — nearly four times the shell’s own growth factor, on a shell of five whorls, from a method a person would describe as careful.
The pattern is the step count. The best reading is the one with the smallest residual, and at a long span the smallest residual belongs to a coarse walk that skips most of the curve. Given enough steps the method is exact — the spanned column shows 3.2000 at thirty-six steps over two turns and at fifty over two and a half — and the failure is entirely in choosing the step by how well it fits.
That is why the entry is directional. A coarse walk cuts corners, a cut corner is shorter than the arc it replaces, and a shorter step per turn reads as a faster expansion. Every error it makes is upward, and 6.854 is upward of 3.2.
What that changes about the claim
Not the verdict, but its shape. The claim is refused by every way of reading a section except one, and the exception is directional rather than random: under-stepped dividers do not scatter the answer, they raise it, and they raise it towards exactly the number the claim names.
With the dividers set aside the six remaining sources together are 54.4 per cent without cancellation and 27.2 in quadrature, against a gap of 114.2. The margin is a factor of 2.1 on the pessimistic reading and 4.2 on the optimistic one. That is a comfortable refusal, and it is the refusal that should have been claimed all along — not “the gap is enormous and no error could reach it”, but “the gap exceeds every error except the one that points straight at it, and that one is avoidable by taking enough steps”.
The difference between those two statements is not rhetorical. The first is false and this essay is the measurement that shows it.
What the same budget cannot settle
This is the half that is not already known, and it is more useful than the half that is.
A claim that two growth factors differ by a stated fraction is decidable when the difference exceeds the budget. With the dividers set aside, the budget is 54.4 per cent — so a claim about a hundred-per-cent difference is decidable and a claim about anything up to fifty is not. The golden claim, at 114.2 per cent, sits just inside the decidable side.
A claim that a nautilus grows at 3.2 rather than 3.4 is a six-per-cent claim. It is not decidable at all by a section read the way these essays read one, and neither is a twenty-per-cent claim, and neither is a fifty. Every claim this subject might want to make about a difference between shells — between species, between juveniles and adults, between populations — is in that region.
So the position is sharper than “the famous number is wrong”. It is: the famous number is wrong by an amount large enough that a badly read section still refuses it, and almost every more interesting question about the same quantity is beyond what a badly read section can answer.
What a tighter claim would require
The budget also answers the question a person would actually ask next, which is what it would take.
A claim of 114 per cent needs the dividers controlled and nothing else. A claim of sixty per cent needs the same. A claim of thirty needs the dividers, the span and the septum count. A claim of ten needs those three plus the changing shell and the centre. A claim of five needs six of the seven, leaving only the oblique view, and even then the residual budget is 3.0 per cent — so a claim of three per cent is not reachable by controlling anything on this list.
“Controlled” has a different meaning for each entry and the list is short enough to say them. The dividers: take enough steps, which a measurement in steps quantifies. The span: measure over at least a turn, or use a span where a tilt cannot bite. The septum count: use a whorl-apart pair, which what the septa count shows needs no count at all. The changing shell: fit two arcs, with the caution a centre that invents a life history attaches. The centre: there is no way to control it on a real shell, which is why it sits near the bottom of what can be reached.
Why worst cases and not error bars
Every entry is a worst case under a stated condition, and that choice deserves defending because it is what makes the total so large.
The alternative is a standard deviation — what the reading typically does. That is the right unit when the errors are draws from a distribution nobody controls, and it is the wrong unit here, because most of these are not random at all. A centre is displaced by a definite amount in a definite direction. A camera is at a definite angle. A step length is chosen once. A shell either changed as it grew or it did not. Those are conditions, not draws, and a reading taken under one of them is wrong by its full amount every time.
So the budget answers “how wrong could a reading be, given that nobody recorded which of these applied” rather than “how wrong is a reading on average”. The first is the question a person reading somebody else’s published growth factor has to ask. The second is the question the original measurer could have answered and did not.
The one entry that genuinely is a distribution is reading noise, and it is not in the list — three points on a diameter measured it separately and found it below a per cent at a thousandth of the outer radius, which is inside the rounding on every other line here.
Which instrument the budget recommends
Three of the seven entries — the span, the oblique view and part of the centre — are properties of the fit rather than of the shell, and the calipers are immune or nearly immune to all three. Three points on a diameter found the caliper aim error entering as its square, and a section seen from the wrong angle found the caliper reading exactly unmoved by an oblique view.
That is the practical recommendation the budget supports, and it is not the one the earlier essays implied. The fit uses every point and is exposed to every geometric error in the picture. The calipers use three points on a line and are exposed to almost none of them. Where reading noise dominates the fit wins, which three points on a diameter measured; where the geometry of the picture is in doubt, which is the usual case with a photographed specimen, the calipers win by a margin the budget now quantifies.
Why a budget is worth assembling at all
Each of these errors was reported when it was measured, and each was reported as small. Six point nine per cent is small. Twenty-two per cent is uncomfortable but survivable against a claim that is out by a factor of two. Nine per cent is a footnote.
Put together they are 332.6 per cent, and the conclusion they were each individually consistent with is one none of them supports. That is the specific thing a budget is for, and it is the reason this one is a measurement rather than a summary: the sum was not predictable from the parts, because one part turned out to be twenty times the size of the rest put together and nobody had looked at them on one axis.
A budget can only grow
There is a structural point in this that is worth stating because it will bind on every future essay in this thread.
Each of the seven entries arrived as a result — a thing measured and reported. None of them was measured in order to be added to anything. But once the budget exists, every new measurement of a way the reading can go wrong is a new entry, and an entry can only make the total larger. A collection that keeps finding error sources is a collection whose own conclusions get progressively harder to support.
That is not a reason to stop looking, and it is not a paradox. It is the ordinary situation of any measurement whose uncertainty is being taken seriously for the first time, and the right response is the one the second half of this essay takes: stop asking whether the budget clears one famous claim, and start asking which claims it clears. The first question has a yes-or-no answer that a single new entry can overturn. The second has an answer that moves smoothly and stays useful.
It also puts a specific obligation on the essays after this one. A section seen from the wrong angle arrived at 3.0 per cent here, which is small — but it arrived at 23.7 per cent for a half-turn span at twenty degrees, and the entry above uses ten degrees because ten is what a person would fail to notice. A budget is only as honest as the conditions its worst cases are stated under, and every one of the seven has a condition attached that somebody chose.
What is claimed, in one line
The seven measured ways of getting a shell’s growth factor wrong come to 332.6 per cent added without cancellation and 279.6 in quadrature, both larger than the 114.2 per cent gap the golden-spiral claim asks a section to resolve; exactly one of them — the dividers, at 278.2 per cent alone — is responsible, and with it set aside the remaining six come to 54.4 per cent and refuse the claim by a factor of 2.1, while leaving every claim about a difference of half or less undecidable.
What none of this establishes
That these are all the ways a reading can be wrong. They are the ways measured so far, and a source nobody has thought of is not in the sum — which is the standing weakness of any error budget and is worse here, because the budget’s conclusion turned on one entry and a new entry of similar size would turn it back.
Nor that any real specimen carries any of them at these sizes. Each is a worst case under a stated condition, and a careful worker avoids most of them. The budget is what an uncontrolled reading could produce, not what a controlled one does.
And nothing here measures a nautilus. Every entry is computed on curves built to stated parameters, which is what makes the numbers comparable and what stops them from being a statement about animals.
What would withdraw it
A source whose worst case, recomputed, differs from what its own library returns. A second entry that alone brings the budget inside the gap, which would mean the conclusion does not rest on the dividers. A budget that clears the gap in quadrature but not without cancellation, or the reverse, which would make the verdict depend on how the sources are combined. A claim of five per cent shown decidable by controlling fewer sources than the list above names. Each is checked every time the page is built.
Still open: whether the entries are independent
The budget adds the sources two ways and neither is right. Adding without cancellation assumes they conspire; adding in quadrature assumes they are independent; and at least two of them are known not to be. A displaced centre and a short span are the same failure seen twice — how far a centre must move found that the centres which make a nautilus golden exist at every span up to 1.15 turns and at none from 1.2 upward, which is a correlation rather than two errors. An oblique view and a short span are also coupled, by the result that a tilt costs a half-turn span twenty times what it costs three turns. The next measurement is the joint distribution rather than the seven margins: a shell given a displaced centre, a tilt and a short span together, swept over all three, and the question of whether the combined worst case is the sum, the quadrature, or something neither.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- One number for a shell that changes — both name claim testing, growth factor, honest limits, measurement error, model scope, summary statistic, untested claim
- A difference forgets a drift — both name claim testing, honest limits, identifiability, negative result, summary statistic, untested claim
- A law that never stopped changing — both name claim testing, growth factor, honest limits, identifiability, model scope, negative result
- A wall that was never measured — both name claim testing, honest limits, identifiability, measurement error, negative result, summary statistic
- Four walls closer than they looked — both name claim testing, error propagation, honest limits, identifiability, measurement error, negative result
- The second statistic was the first — both name claim testing, honest limits, measurement error, negative result, summary statistic, untested claim
Named objects
A flat tag is an object no other essay names yet.
Claim testingError propagationGolden spiralGrowth factorHonest limitsIdentifiabilityMeasurement errorModel scopeNegative resultSelf-correctionSummary statisticUntested claim