No majority and no pair
Worth reading first: The damage has a period.
Two properties of a wrecked stem’s displacement profile, computed by different machinery for different reasons.
No balanced pair. The exchange is quantified over rows whose exceptional chains are exactly two, displaced in opposite directions by the same amount to within a twentieth. Rows that are not like that are set aside, and there are sixteen of them in the extended census.
No majority. The level the other chains sit at is the value the largest cluster of them agrees on, and on some rows that cluster is not more than half the profile. There are sixteen of those too.
Fifteen rows are in both
Which is the finding. Two quantities read off different things — one a fact about the largest agreeing cluster, one a fact about the two largest deviations — agree on fifteen of sixteen rows each.
Nothing arranged that. The exchange’s exclusion was written when the level was still a median, and it never looked at cluster sizes. The cluster reading was written a round later to test an objection to something else.
Two independent classifications of thirty-six rows agreeing on all but one each way is the kind of coincidence that means the two are measuring the same thing.
It is the same shape this site asks of every number it reports: computed twice by machinery that shares nothing, and worth reporting only because the two routes are separate. Here the object is a partition of a census rather than a single value, and the agreement is a count rather than a coincidence of digits.
Why they should agree
Because a balanced pair is a small number of exceptions by construction. Two exceptions on a lag of five leaves three at the level, which is a majority; on a lag of eight it leaves six.
So a row with a balanced pair almost always has a majority, and a row with many exceptions almost never does. The two readings are two ways of asking whether most of the profile is doing one thing.
What makes the agreement worth reporting is not that it is surprising but that it is exact: fifteen of sixteen each way, with the two disagreements identifiable and explicable.
The first disagreement: a majority and no pair
g008 cut at offset 6. Eight classes, five of them within a degree of each other at about 2
degrees, and three at −87.4, −133.3 and −132.8.
Five of eight is a majority, so the level is not in doubt and the median and the cluster agree to a tenth of a degree. But three exceptions is not a balanced pair, so the exchange sets it aside.
That row is the census’s one clean three-cycle: three adjacent chains rotating into one another’s places, whose displacements sum to a whole turn to four parts in ten thousand. It is the single most structured row in the excluded set and it has a comfortable majority.
The second: a pair and no majority
l013 cut at offset 4. Four classes, at −50.7, 35.7, 39.0 and 125.3 degrees.
Two of those — 35.7 and 39.0 — are within the tolerance of each other and are the level. The other two are exceptions, they are adjacent, and they are equal and opposite about that level at −89.8 and +86.3, which is a balanced pair. So the exchange keeps it.
And two of four is not a majority. It cannot be: a lag of four with a balanced pair has two exceptions and two chains left, and half is not more than half.
So this row parts from the other reading by arithmetic rather than by anything about the stem. Every row at a lag of four with a balanced pair is in the same position.
Which bounds the majority reading
At small lags it has no room. A lag of four gives at most three chains at the level; a lag of five gives four; and a balanced pair takes two of them away.
The census holds five rows at a lag of four and one of them has a balanced pair. Only that one is affected, but the limitation is general: a majority is not available to a short lag with an exchange in it, and the reading has to be read with the lag beside it.
That is not a reason to prefer the median. It is a reason to report the cluster size and the lag together, which the reading does.
What the cluster sizes look like
Twenty rows of thirty-six have a majority and sixteen do not. Among those that do, the smallest majority is five of eight and the largest is nine of eleven.
Among those that do not, the largest cluster runs from two of eight to three of seven. Two of eight is a quarter of the profile, which is the weakest level anything in this census is measured against.
Four rows have a tie — two or three distinct clusters of the same size — and three of the four are at two of eight or two of five. So the rows with the weakest levels are also the rows whose levels are least determined.
The weakest rows in the census
g005 cut at offsets 7 and 8. Eight classes each, and the class means fall into four groups
of two: about −132, about −43, about +1, about +93, and one more near +135.
Three clusters of two are tied for largest. Any of them could be the level and the reading says so, taking the tighter and reporting the tie.
Read against the median, which lands at about +2, six of the eight classes are exceptions. Read against any of the three clusters, six of eight are still exceptions. So this row is awkward whatever is done to it, and the recomputation does not rescue it.
What such a row means
A profile whose chains sit at four or five distinct values is not a pattern with exceptions. It is a stem in which most chains have moved and moved by different amounts, which is a different object from the one the reading was built for.
The exchange excluded them for the right reason and by the right criterion. What the cluster reading adds is a description of why they are excluded — not merely not a balanced pair but no value that most of the profile agrees on.
That is a more useful sentence, because it says the reading has nothing to measure rather than that the row failed a test.
The two disagreements are not symmetric
g008 at offset 6 has a majority and is excluded because it has three exceptions rather than
two. That is the exchange being narrow: three chains rotating is a real structure and the
exchange is about pairs.
l013 at offset 4 has a pair and no majority because its lag is four. That is the majority
reading running out of room.
So one disagreement is a property of the exchange’s definition and the other is a property of arithmetic, and neither is a defect. Both are the kind of edge two independent classifications produce when they nearly agree.
What this says about the median
That its condition — exceptions in a minority — fails on exactly the rows the exchange had already set aside, and holds on every row it keeps.
So the median was never being asked to do anything it could not do, in any file that quoted it. The eight rows where the two levels disagree are all in the excluded set, and the excluded set is where the majority fails.
That is a tidier state of affairs than it first looked. There is no correction to any published number here; there is a reading that was undefined on sixteen rows and returned a value anyway.
The count that was wrong before
The excluded set has been miscounted once already, in the other direction. It was described as rows with three or more exceptional chains, which is eleven rows, and the exchange’s actual criterion is rows with no balanced pair, which is sixteen.
The five in the difference carry exactly two exceptions displaced the same way, and a pair that is not equal and opposite is not a pair.
So this is the second time two descriptions of one set have been found to differ. Both times the difference was a handful of rows and both times the rows were the interesting ones.
The pattern is worth naming because it is cheap to check and keeps being productive. A set described two ways in two files is a set whose two descriptions can be compared, and the rows in the difference are where the descriptions stop meaning the same thing — which is how the one clean three-cycle was found and how the row that changes side in this round was.
What is not claimed
That the majority reading should replace the exchange’s criterion. It should not: the exchange is about balanced pairs and no balanced pair is the exact statement of what it cannot measure.
The majority reading is a description of the profile rather than a rule for inclusion, and its value is that it says something about a row that the exclusion does not. A row excluded with a majority is a row with structure; a row excluded without one is a row with nothing to measure against.
Fifteen of sixteen agreeing is a reason to trust both readings, not a reason to merge them.
What the tie count is worth
Four rows of thirty-six have a tied densest cluster, and every one of the four is in the excluded set. So a tie is another symptom of the same condition rather than a separate defect.
It matters because it is the one thing a median can never report. A median returns a value on a profile with four equal groups just as confidently as on a profile with six chains at one value, and the difference between those two situations is the whole of what this reading adds.
Reporting the cluster size and the tie count beside every level is therefore not decoration. It is the part of the reading that says how much the level is worth.
How the cluster is found
For each class mean in turn, count how many class means sit within the exception tolerance of it, measured the short way round the circle. Keep the candidate with the largest count. Where several candidates tie, take the tightest and record how many tied.
That is deliberately the simplest thing that could work, and it has no free parameter: the cluster’s width is the same ten degrees that decides which classes are exceptions, so there is nothing to tune.
It also has an honest failure. On a profile whose classes are spread evenly round the circle the largest cluster holds one, and the reading is refused rather than returning the mean of a single class — which is exactly what the median does without saying so.
Why not just count exceptions
Because the count of exceptions is measured against the level, and the level is what is in question. A row read against a bad level has too many exceptions by construction, so counting them cannot tell a genuinely scattered profile from a well-behaved one measured wrongly.
The cluster size is measured against nothing. It is a fact about how the class means sit relative to each other, and it is the same whatever value is chosen as the level.
That is why it is the right quantity for this comparison. Six rows lose an exception when the level is recomputed and not one gains one, and the cluster size does not move at all — because it never depended on the level.
The distribution, in full
Twenty rows with a majority: five of eight on one row, six of eight on five rows, three of five on six, five of seven on four, and nine of eleven on three. The three at nine of eleven are the fifth-lag lattice’s own rows, which are the cleanest profiles in the census.
Sixteen without: two of four on four rows, two of five on three, three of seven on eight, and two of eight on two.
So the split is not marginal. The weakest majority is five of eight at 62 per cent and the strongest non-majority is three of seven at 43, with nothing between them.
The gap is not a threshold
Between 43 per cent and 62 per cent there is nothing, and the line at a half sits in the middle of it. So majority is a separation on this census rather than a decision, in the same way the exception tolerance is and the closure threshold is not quite.
That is worth checking rather than assuming, because a classification whose line runs through a continuum is a classification that will move when the census grows.
It will move eventually. A row at exactly half — four of eight, or three of six — is arithmetically possible and none has appeared. When one does, the line will have to be stated as a convention rather than as a gap.
What the two readings are each good for
The exchange’s exclusion is a criterion: it says which rows that file can measure, and it is exact, because a balanced pair either exists or does not. Nothing here proposes replacing it.
The cluster size is a description: it says how much of a row’s profile agrees on anything, and it is the more informative of the two on a row that is excluded. No balanced pair is a statement about two chains; two of eight chains at the level is a statement about the whole profile.
Having both is what makes it possible to say that g008 at offset 6 is excluded for a reason
that has nothing to do with its profile being messy — five of its eight chains agree — and that
the two rows on g005 at offsets 7 and 8 are excluded for a reason that has everything to do
with it.
Why fifteen of sixteen is not a coincidence worth explaining away
Because the two readings are measuring the same underlying thing from opposite ends, and the surprise would have been disagreement.
A profile with two chains displaced and the rest at one value has a majority by arithmetic whenever the lag exceeds four. A profile with five or six chains displaced has no majority for the same reason. So the correlation is nearly forced, and what the count establishes is that the census contains no row of the awkward middle kind — three exceptions on a lag of eight, say, where the two readings would part.
There is exactly one such row and it is the three-cycle. That is the one the whole set-aside thread turned on, and finding that it is also the one row the two classifications disagree about is the part worth reporting.
What is claimed
That sixteen of the extended census’s thirty-six rows have no cluster of chains amounting to a majority, that sixteen have no balanced pair, and that fifteen rows are in both sets.
That the two rows they part on part for opposite reasons: g008 at offset 6 is the census’s
one clean three-cycle and has five of eight chains at its level, and l013 at offset 4 has a
balanced pair at a lag where a majority is arithmetically unavailable.
And that the median’s own condition therefore fails on exactly the rows the exchange sets aside, so no published number in that thread was ever measured against a level standing on an exception.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- The column nobody read — both name census design, claim testing, honest limits, measurement, negative result, summary statistic
- Three rows change sides — both name claim testing, classification, honest limits, measurement, negative result, summary statistic
- A difference forgets a drift — both name claim testing, honest limits, measurement, negative result, summary statistic
- A fifth of the hop — both name claim testing, honest limits, measurement, negative result, summary statistic
- A list that can only shrink — both name claim testing, classification, honest limits, measurement, negative result
- A period the grid invented — both name claim testing, honest limits, measurement, negative result, summary statistic
Named objects
A flat tag is an object no other essay names yet.
Census designClaim testingClassificationDamage profileExceptional chainExchangeHonest limitsMeasurementNegative resultResidue classRobust statisticSummary statistic