Is any of it true?
What is wrong with the numbers on the other five pages.
The calibrating null moves too fast
Its median per-frame velocity is 2.2× the corpus's. The excess measured against it is an upper bound.
α = 1e-4 is a choice
Picked so the null rate lands under MoSeq's artifact mass. Any other point on the curve is equally defensible.
Seeds are single exemplars
A state is a ball around one observed stretch, so variation within a behaviour falls outside it.
Human validation
The naive rater beats chance, unevenly
A naive rater's blind forced choice pools to 43.1%[32.0%, 54.6%] against 33.3% chance. Per-label accuracy spans 100.0% to 0.0% — the pooled number is the wrong object.
p = 0.0246, n = 102 trials over twenty labels; some are recognizable to a naive observer and most are not, which is exactly what the per-label range says and the pooled number cannot. Batch 2 — two experts, twenty different labels — is unscored. That half of “no human has validated any label” still holds.
The checks, both arms
nonparam: every check returns PASS, for every row below.
| check | moseq |
|---|---|
| dwell vs geometric, per animal | PASS |
| tracking-artifact enrichment | FAIL |
| animal identity leak | PASS |
| session identity leak | PASS |
What came before
An earlier instrument, retired
Before recurrence: 0.397 and 0.458 nats over 88 animals — a characterized negative, not a discreteness feature.
Tried before the recurrence measurement on this site: a rate-limited predictive-information channel between past and future, on the theory that the curve saturates for discrete behaviour and grows without bound for continuous behaviour. It did not produce the interior feature that would need. Superseded readings for both points are on the ledger — this is the corrected pair, not the numbers first reported.
The retraction
The MoSeq dwell claim was withdrawn: every surrogate reproduced it through the same fitted model, so the evidence was for the labeller's stickiness prior rather than for the corpus.
Q4's noise control, run
Run twice, on white noise and the OU continuum: zero labels, so nothing for a grammar test to run on — a step upstream of why MoSeq's noise passes the same test.
| labeller | applied to | coverage | grammar |
|---|---|---|---|
| MoSeq | white noise | 100% | PASS 21 / 69 |
| nonparam2 | white noise | 0.0000% | INCONCLUSIVE 0 / 0 |
| nonparam2 | OU continuum | 0.0000% | INCONCLUSIVE 0 / 0 |
0.0000% alone is not evidence of a tight threshold — the old, uncalibrated arm returned this same number because its surrogates sat 24× from the corpus (ledger 55), not because it was well-set. What makes v2's zero informative is that the same per-seed thresholds fire on 2.16% of the microstate null — the one that lands on the data manifold — while firing on neither of these. A labeller that declines to describe noise does not, by itself, make its eight surviving motifs behaviour.
Ledger 61: this exact result was first written into NONPARAMETRIC.md before either run had happened — the same failure the site retracted the MoSeq dwell claim for, caught here before it reached a page.