Experimental Branch · Current phase complete

Intermission.

What does the current representation preserve—and what did the spelling controls leave unexplained?

The current description

A symbolic trajectory of written form.

PCDVR deterministically recodes written characters as categorical states and retains their order as a trajectory. The current tests show measurable written-form information in state composition, transitions, boundaries, and complete paths. They have not established separate word-family information beyond the tested ordinary-spelling controls.

This is not a semantic embedding, an etymology detector, a validated phonetic or articulatory geometry, or a demonstrated improvement over spelling-based classification.

What the earlier tests measured

Order, then boundaries.

In the independent 48-word, four-family sequence challenge, ordered trajectories scored 68.8%; 2,000 randomized-order controls averaged 43.9%. One control matched or exceeded the ordered score (plus-one empirical p ≈ 0.001). This tested family classification, not prediction of the word itself.

In the separate 47-word boundary challenge, the raw final character scored 66.0%, the PCDVR final state 66.1%, and the full path 70.2%. Similar final-character and final-state scores motivated a direct test against ordinary spelling controls.

Inspect the sequence tests → · Inspect the boundary tests →

Experiment 22C · 96 reference / 48 challenge words · Four families

Testing what remains after spelling controls.

The declared model used edit distance, character unigram/bigram/trigram measurements, written-vowel sequence, length difference, a two-character prefix, and one-, two-, and three-character suffix features to predict PCDVR pairwise trajectory distance.

Mean R² across five word-blocked cross-validation folds was 0.606—about 60.6% of distance variance explained on average across folds. This is a distance-prediction score, not classification accuracy.

After subtracting the model-predicted component, residual-based family classification produced:

Observed residual accuracy22.9%22.9%
Label-permutation mean24.8%24.8%
Four-family chance25.0%25.0%
2,000 label-permutation controlsResidual accuracy 22.9%, chance 25%. 1,307 controls matched or exceeded the observed result. Bars count controls at each number correct out of 48.2,000 label-permutation controls3/48 correct: 4 controls4/48 correct: 10 controls5/48 correct: 23 controls6/48 correct: 47 controls7/48 correct: 80 controls8/48 correct: 135 controls9/48 correct: 176 controls10/48 correct: 218 controls11/48 correct: 255 controls12/48 correct: 211 controls13/48 correct: 231 controls14/48 correct: 205 controls15/48 correct: 143 controls16/48 correct: 99 controls17/48 correct: 66 controls18/48 correct: 35 controls19/48 correct: 28 controls20/48 correct: 15 controls21/48 correct: 11 controls22/48 correct: 4 controls23/48 correct: 1 controls24/48 correct: 2 controls25/48 correct: 1 controls0%50%100%Residual family-classification accuracy
Gold: observed residual (22.9%). Dashed blue: four-family chance (25%). Each bar represents an exact number correct out of 48, not a smoothed distribution.

These are label-permutation controls, not the within-word sequence shuffles used in Experiment 21P. Of 2,000 controls, 1,307 matched or exceeded the observed residual accuracy. Plus-one empirical p = (1,307 + 1) / (2,000 + 1) = 0.6537.

Experiment 22 did not detect residual word-family signal under the declared spelling controls in this fixed dataset.

This does not prove that every possible spelling model explains the representation, or that residual signal is absent in every dataset. The central 95% permutation interval was 12.5%–39.6%; it describes the controls, not a confidence interval for generalization.

Companion checks · 22A / 22B

Other views of the same challenge.

22A · Removing boundary characters

PCDVR accuracy was 68.8% for full words, 58.3% after removing the last character, 50.0% after removing the last two, and 45.8% after removing the first and last. These comparisons do not isolate independent linguistic information.

22B · Spelling-matched candidates

In the closest spelling-matched quartet condition, tie-aware candidate-family accuracy was:

Raw edit20.0%20.0%
Raw 1–3gram profile16.7%16.7%
Raw orthographic composite41.1%41.1%
Vowel-only trajectory30.0%30.0%
PCDVR full trajectory19.8%19.8%

This candidate-ranking task differs from the population and residual classifications above; its scores should not be treated as interchangeable.

Intermission, not erasure

Same nodes. Different paths.

LISTEN, SILENT, ENLIST, TINSEL, and INLETS still share a state inventory while following different ordered trajectories. That representational distinction remains; it does not by itself establish extra linguistic information.

The current classification phase is paused. Existing results, including negative findings, remain available. Future questions about resolution, diacritics, compression, or deliberately phonetic mappings would be new investigations.

Explore the anagram illustration →

Next · Star Branch · No results yet

A different geometry to explore.

Can a PCDVR constellation or trajectory be matched quantitatively to real astronomical geometry—and how do those matches compare with randomized controls?

This begins a new question, not an astronomical claim about names or an extension of the word-family results.

Support the project →Contact Bryan →
Source and scope

Values transcribed from pcdvr_orthographic_residual_challenge.json, schema pcdvr.exp22.orthographic_residual.v1. The 2,000 exported null outcomes were checked against their reported mean, exceedance count, and plus-one p-value. This website check is not an independent rerun of the experiment.

Reference SHA-256: 8bcca38473ee357440e7d24a88d43df7c26c609bb8f0db3ad0bec5d24ae1a189
Challenge SHA-256: c1aa2b43982a13709d60e86b4d8553a5afe2689d7aa68913a2c67fe11f744038