The order belonged to the method
Worth reading first: What a summary throws away · Two laws that want opposite tissue.
The width of the disorder dip at a rational divergence has two laws attached to it. It falls as the inverse square of the head size, and its coefficient carries the denominator rather than how well the fraction approximates. Both survive.
After the two of them are taken out, something is left over, and the leftovers looked ordered. Inside each family of four fractions sharing a denominator, the one with the closest neighbour gave the widest dip and the one with the most space gave the narrowest. It was reported carefully — as a rank comparison rather than a fit, since four points do not support a fit — and with the observation that the order runs in the direction a measurement artefact would take.
This essay is what happened when it was measured again with a different instrument.
What the claim was, and why it was suspicious from the start
A crowded fraction has a neighbour whose own dip pulls the profile down nearby. Pulling the profile down near the shoulder makes the shoulder look lower, and a half-width read as a level crossing is read on the shoulder. So a crowded fraction would come out wider whether or not anything about the fraction is different.
That was said at the time, in those words. What it needed was an instrument whose free parameter is not on the shoulder — which is what the area was supposed to be, since an integral has no level in it.
The re-measurement
The area is measured on a scaled axis, out to a stated window, and divided by the dip’s depth: an equivalent width, in scaled units. Four fractions per family, two head sizes each, at three windows.
Take the denominator-34 family, which is where the residual was measured. Order the four by crowding, most crowded first: 15/34, 13/34, 9/34, 11/34, at 0.1795°, 0.1925°, 0.1998° and 0.2862°.
Now order them by the area of their dips:
| window | widest to narrowest | spread |
|---|---|---|
| 50 | 9/34 > 13/34 > 15/34 > 11/34 | 2% |
| 100 | 11/34 > 13/34 > 9/34 > 15/34 | 1% |
| 200 | 11/34 > 15/34 > 13/34 > 9/34 | 12% |
| 400 | 11/34 > 15/34 > 13/34 > 9/34 | 25% |
Three things at once.
At the narrow windows there is nothing to order. At a window of 50 the four equivalent widths are 100, 100, 98 and 98; at 100 they are 152, 152, 151 and 153. The integral out there is still measuring the floor of each dip, which is the same depth for everybody by construction, so the ranking is a ranking of rounding.
Where the numbers do separate, the order is not the crowding order. At windows of 200 and 400 the spread is 12% and 25%, which is enough to order — and the order is 11/34 first and 9/34 last. 11/34 is the loneliest of the four. The level crossing had the most crowded one widest; the area has the loneliest one widest.
And the order moves before it settles. The window that first produces a real spread is also the window at which the ranking stops changing, so there is no setting at which one can say the ranking has converged and separately say the numbers have.
The families do not agree with each other either
One family reversing could be one family being odd. Three families give three different answers, which is worse and more informative.
Denominator 21 never settles. The widest fraction is 8/21 at a window of 50, 10/21 at 100, and 8/21 again at 200 — and 10/21 is the loneliest of its family, so the ordering moves between the two ends of the crowding scale as the window moves. Its two head sizes stop agreeing about the ranking as well: at a window of 200 the smaller head puts 8/21 first and the larger puts 5/21 first, on the same four fractions.
Denominator 34 settles, backwards, as above.
Denominator 55 settles into an order that differs from the crowding order only in its last two members: 23/55 and 17/55 are the two most crowded and the two widest, which is what the residual claim predicts, and the remaining two are swapped. So the one family that nearly agrees with the claim is the family with the tightest neighbourhoods — 0.15° and 0.16° for its two leaders — which is also where the instrument’s own bias is largest.
So of three families, one agrees with the claim except in its tail, one comes out backwards, and one has no stable order at all. That is not a weak effect. It is what a quantity with no relationship to the ordering looks like when it is measured four times.
And the instrument has a bias of its own, in the same direction
There is a reason the area could not have settled this even if it had agreed.
A window fixed in scaled units is not a fixed distance into each fraction’s own neighbourhood. At a window of 200 the integral reaches 6.1% of the way to the neighbour for 15/34, the most crowded of its family, and 3.8% for 11/34, the loneliest. The crowded fraction’s window goes half again as far into its clear space, so it integrates more of its neighbour’s shoulder — which reduces its deficit and makes its area come out smaller.
That bias runs opposite to the level-crossing bias. The level crossing makes a crowded fraction look wider; the area makes it look narrower. The two instruments disagree about the ordering and each of them has a mechanism that predicts the disagreement.
Which is the strongest statement available here: the two instruments’ errors have opposite signs and the measured orders have opposite signs, and neither order is evidence about the fractions.
What a residual is, and why they are dangerous
It is worth generalising for a moment, because this is the fourth time on this site that a residual has turned out to belong to the instrument, and the pattern is recognisable in advance.
A residual is what is left after the effects one understands are subtracted. That makes it, by construction, the part of the measurement where the systematic errors live: every effect the model does not contain is in there, and so is every bias of the instrument, and there is nothing else in there except the thing being looked for. The signal-to-noise ratio of a residual is the worst in any measurement, and it is where people look for new effects because it is the only place a new effect could be.
The specific trap here has a shape worth naming. The residual was ordered by the quantity that sets the instrument’s own limit. Crowding decides how close a background sample may be taken, where a level crossing lands, and how far a window may be integrated. Any of those three, done imperfectly, produces an effect ordered by crowding. So the one hypothesis that could not be tested with these instruments is precisely the one that was tested.
That is not hindsight. It is a rule that can be applied before the measurement: list what an instrument’s free parameters depend on, and do not test a hypothesis about those quantities. Had that list been made here it would have had one entry — crowding — and the residual claim would have been recognised as untestable before the numbers were produced.
The verdict, stated as a refusal
The residual ordering by crowding is withdrawn. Not refuted — withdrawn, which is the weaker and more accurate word. What can be said is:
- the ordering measured with a level crossing is what that instrument’s bias predicts;
- the ordering measured with an area is roughly what that instrument’s bias predicts, in the other direction;
- no instrument here has a free parameter that is independent of the crowding, so nothing here can measure a residual ordered by crowding.
The two laws underneath are unaffected, and it is worth being clear about that, because a withdrawal in a thread makes everything in the thread look shaky. The 1/n² scaling is a statement about the shape of the profile and is visible directly: on the scaled axis the profiles at two head sizes lie on top of each other. The denominator coefficient is a comparison between families measured the same way, and comparisons of that kind are exactly what survives an instrument with a consistent bias.
The one number that did move, and where it went
There is a positive result inside all of this, and it belongs to the essay rather than to the withdrawal.
The area measures the two fractions the level crossing had to refuse. 9/34 gives 318 and 338 scaled units at two head sizes; 24/55 gives 348 and 362. Both are as reproducible as anybody else in their families — a disagreement of 6% and 4%, against a level-crossing method that gave one of them 2,665 at one head and 6,810 at the other and was right to refuse it.
That matters for the denominator result rather than for the crowding one. The claim that the width coefficient carries q was tested on ten of twelve fractions, with two missing because the instrument could not see them, and a missing case is always the one somebody suspects. With the area the family is complete, and the two recovered fractions sit inside their families’ spread rather than outside it: 9/34 at 318 against a family running 318 to 358, and 24/55 at 348 against a family running 350 to 373.
So the repair did what it was built for — it just did not do the extra thing the plan hoped it would do on the way.
What this costs, in claims
Three sentences elsewhere on this site have to be read differently now.
“The residual is ordered by crowding, in the direction a measurement artefact would take.” The second half of that sentence was the useful half and it should have been the whole of it.
“Inside a family the most crowded fraction gives the widest dip.” True of the numbers the level crossing produced, and not a statement about dips.
“Four points do not support a fit, so the claim is only that the order is not random.” The order is not random. It is systematic, and what it is systematic in is the instrument.
Is there any reading under which both instruments are right?
One is available and it is worth taking seriously for a paragraph, because dismissing it would be too convenient.
Suppose the dip really is wider for a crowded fraction near its floor and narrower far out on its shoulder — a genuinely different shape rather than a different width. A level crossing read at half the background is sensitive to the shoulder; an area at a wide window is sensitive to everything. The two could then disagree honestly, with each reporting a real property of a different part of the profile.
If that were so, the area at a narrow window should agree with the level crossing, since a narrow window weights the floor. It does not. At a window of 50 the four fractions of denominator 34 come out at 100, 100, 98 and 98 — no ordering at all, because down there every dip is at its own full depth and the integral is measuring the window.
So the shape reading predicts a place where the two instruments meet, and there is no such place. The disagreement is not between two views of a structured profile; it is between two instruments whose free parameters both key on the same quantity.
What would settle it
A measurement whose free parameter does not touch the crowding. Two shapes are available and both are expensive.
Hold the window at a fixed share of each fraction’s own crowding, and adjust the head size so that the dip occupies the same share of the window. That requires a different head for each fraction, chosen as a function of its neighbourhood, and a background band re-chosen for each.
Or build fractions with matched neighbourhoods. Rather than taking the fractions a denominator happens to supply, choose sets whose nearest-neighbour distances agree to within a few per cent and whose denominators differ — then crowding is held fixed by construction and whatever is left is the residual. The arithmetic here allows it; the sets are smaller and the heads are larger.
Neither has been run. What has been done is to say plainly that the earlier claim cannot be supported by the measurements that produced it, which is worth more than a third instrument with a third bias.
The check
The withdrawal is asserted rather than described, which is unusual and is the point: the claim that a claim cannot be made is itself a claim that can fail.
Two families are ordered by area at three windows each, and the check requires that at no window does the area’s order match the crowding order. If the area had agreed with the level crossing everywhere, that assertion would fail and this essay would be wrong — which is the shape a withdrawal has to have to be worth anything.
A second assertion requires the order to change with the window in at least one family, so that “the order belongs to the method” is not merely an interpretation of two instruments disagreeing once. And the fixed-degree comparison is asserted separately: at a fifth of the way to the neighbour, every fraction in two families gives a dip at the smaller head and a negative one at the larger, which is what says the window is not a tolerance.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A dip with no outer edge — both name artefact, convergents, disorder, falsifiability, honest limits, measurement, measurement error, null model, rational angle, rational approximation, sampling, summary statistic
- A width read off a staircase — both name artefact, convergents, falsifiability, honest limits, identifiability, measurement, measurement error, rational angle, sampling, summary statistic
- The background is not one sample — both name artefact, convergents, disorder, honest limits, measurement, measurement error, rational angle, rational approximation, sampling, summary statistic
- The width carries the denominator — both name artefact, convergents, disorder, honest limits, measurement, rational angle, rational approximation, sampling, summary statistic
- A dip belongs to the head — both name artefact, disorder, honest limits, measurement, rational angle, rational approximation, sampling, summary statistic
- A disturbance that is not passed on — both name artefact, discrimination, evidence, falsifiability, honest limits, measurement, null model, summary statistic
Named objects
A flat tag is an object no other essay names yet.
ArtefactConvergentsDiscriminationDisorderEvidenceFalsifiabilityHonest limitsIdentifiabilityMeasurementMeasurement errorNull modelRational angleRational approximationSamplingSummary statistic