For now, I may have found a close independent result:
Yes — I think there are now at least a few reasonably close precedents for the pattern you are describing: reproducible structure across model depth without requiring a single universal generation-time trajectory.
The closest one I found is Bhattacharya & Kolli’s recent paper, An Analysis of Residual-Stream Geometry Across Transformer Depth. They analyze consecutive residual-stream transitions in six instruction-tuned models using relative displacement and Procrustes-based geometry. They report reproducible depth regularities — typically larger changes early and late, with a quieter middle region, and a particularly strong final transition — while describing the detailed curves as model-dependent but largely condition-stable.
That is not a replication of your setup: the tasks, measurements, model panel, and functional labels are different. But conceptually it seems very close to your narrower result: depth can have reproducible organization without implying one model-independent latent trajectory.
There is also a somewhat different but complementary result in Sun et al., LLM Reasoning as Trajectories: Step-Specific Representation Geometry and Correctness Signals. In mathematical reasoning, they find that representations associated with different reasoning steps become increasingly separable with layer depth. Again, this is not the same measurement as your progressive-structuring scaffold, but it is independent evidence that task-relevant representational organization can become more differentiated as depth increases.
On the time side, I found almost the opposite caution in Gjølbye et al., Reasoning Models Don’t Just Think Longer, They Move Differently: generation-time trajectory geometry is strongly affected by response length, and the corrected geometry remains domain-dependent. So your result that the previously plausible common temporal mode disappears when the panel is expanded does not strike me as anomalous. It may be exactly the kind of conditionality one should expect once depth and generation time are treated as genuinely separate axes.
In that sense, I actually like the mixed result:
ordered depth structure survives; the stronger temporal universality hypothesis does not.
Preserving that 0/8 temporal falsification makes the scope of the surviving claim much clearer.
If I were adding only one or two cheap follow-up controls, I would probably spend them on separating two questions that are currently close to each other but not quite identical:
- Is layer order non-exchangeable because adjacent layers form a structured local path?
- Does that structure also have a meaningful early → late orientation / absolute-depth organization?
Your full layer permutation is a perfectly sensible first null for the first question. For the second question, a secondary null or statistic that preserves more local adjacency could be informative.
Similarly, for the cross-model mean depth-profile correlation of r = 0.789, I would be curious whether the broad coherence is still visible under one very cheap robust view — for example Spearman correlation, first differences, detrending, or a version that reports terminal transitions separately.
Those would not be attempts to make the current positive results disappear. They would help say what kind of depth structure survived.
Why I think the depth/time split is a useful result in its own right
Your explicit separation
H \in \mathbb{R}^{T \times L \times d}
and
\Delta_L H_{t,l} = H_{t,l+1} - H_{t,l}
versus
\Delta_T H_{t,l} = H_{t+1,l} - H_{t,l}
seems more important after looking at the related work, not less.
A number of nearby literatures have independently arrived at the idea that depth should not be treated simply as another sample of generation time.
Layer-wise representational development
The older line of work already contains many examples where different information becomes accessible or reorganized at different depths:
So I think your statement that depth contains reproducible structure is connected to a fairly substantial existing family of observations.
What seems less established is the stronger possibility that one detailed normalized depth curve should be universal across architectures.
That distinction may matter because Csordás, Manning & Potts, Do Language Models Use Their Depth Efficiently?, find something particularly relevant: when mapping residual-stream representations between shallower and deeper models, layers at approximately the same relative depth align best. Their interpretation is that larger models may often spread related computations over more layers rather than simply appending qualitatively new computations.
That makes normalized relative depth itself an interesting baseline.
So mean r = 0.789 could potentially contain several components at once:
- coarse early/middle/late organization shared across Transformers;
- scaling-related relative-depth correspondence;
- genuinely finer shared topology specific to the phenomenon you are measuring.
A cheap robust comparison could help distinguish them without changing the main experiment.
Generation time seems freer to vary
The related evidence I found for temporal trajectories is much less universal.
In Gjølbye et al., raw path statistics during reasoning depend strongly on generated sequence length. After correcting for length, meaningful geometry remains, but its relationship to difficulty differs substantially by domain.
Other recent work also finds task- or context-dependent trajectory organization rather than one single time evolution shared across conditions.
That makes your negative temporal result quite plausible as a substantive result:
depth may be constrained by the stacked computational architecture, while token-time evolution has additional freedom from prompt structure, task, generated length, local decisions, and model-specific decoding dynamics.
I would still avoid promoting that sentence to a universal principle, but it seems like a useful working hypothesis for the next panel expansion.
A small sanity check that changed how I would interpret the permutation result
I tried a deliberately simple sanity check because I was curious what a full layer permutation can distinguish by itself.
I generated smooth high-dimensional layer trajectories with no functional stages encoded at all — just autocorrelated local updates — and evaluated a source-norm-normalized adjacent-path statistic.
As expected, the natural ordering was dramatically smoother than full random permutations.
More interestingly, increasingly conservative disruptions behaved differently:
- full random layer order produced a very large difference;
- small local swaps still produced a detectable difference;
- block-order permutation still produced a detectable difference;
- circular depth shifts also changed the score;
- but completely reversing the layer sequence gave almost the same path statistic as the natural order.
That last case is the useful part.
A symmetric adjacent-distance/path statistic can know that
A → B → C → D
forms a locally coherent path,
while being almost unable to distinguish it from
D → C → B → A.
So there are at least two scientifically different notions of “ordered depth”:
Local-topology / adjacency claim
These intermediate states are not an exchangeable bag; their natural neighborhood relationships matter.
Your existing full permutation directly addresses this.
Directional progression claim
There is something specifically early → middle → late about the organization.
This needs some component of the statistic or null that is sensitive to orientation or absolute normalized depth, rather than only to who is adjacent to whom.
This is why I would keep the full permutation rather than replace it, and optionally add a second test.
For example, depending on the exact topology statistic, possible secondary views could include:
natural depth order
vs.
reversed depth order
or:
natural
vs.
small-block reorderings that preserve most local structure
or an explicitly depth-sensitive statistic such as:
correlation(transition magnitude, normalized depth)
late-third mean - early-third mean
event-depth ordering statistic
The exact choice depends on what your current local-topology statistic actually contains.
There is a general statistical analogy here: with dependent observations, permutation schemes are often restricted so that the null destroys the property of interest while retaining other dependence as much as possible. Winkler et al.'s work on multi-level block permutation is one example of that general principle.
I would not import that method literally into Transformer layers — layers are not repeated-subject neuroimaging observations — but the design principle seems useful:
define which structure the null is supposed to destroy, and preserve unrelated structure where practical.
For your paper, that could make the existing result more interpretable rather than merely producing another p-value.
If the natural sequence beats both a destructive full permutation and a locality-preserving/orientation-sensitive control, the surviving claim becomes considerably sharper.
A second sanity check: raw cross-model correlation can hide very different kinds of agreement
I also tried two small checks on cross-model normalized-depth correlation because mean r = 0.789 is probably the other result where one inexpensive diagnostic could add a lot of information.
Synthetic coarse-depth check
It is easy to construct eight independent noisy profiles that share only a broad low-frequency early/middle/late trend and obtain a mean pairwise Pearson correlation around your reported value.
In one such deliberately generic construction:
raw mean pairwise Pearson ≈ 0.789
after coarse detrending ≈ 0.24
first-difference correlation ≈ 0.03
That obviously says nothing about what causes your 0.789.
It only shows that these alternative summaries answer different questions:
- raw Pearson: “do the overall profile shapes move together?”
- detrended correlation: “do they still agree beyond a coarse depth trend?”
- first differences: “do local increases/decreases line up?”
- Spearman: “is there similar rank/order structure even if amplitudes differ?”
Tiny real-model instrumentation check
I then tried a very small two-model hidden-state measurement, using generic prompts and a normalized adjacent-layer displacement profile.
This produced an initially striking result:
raw Pearson ≈ 0.953
Spearman ≈ 0.121
Looking at the profile made the reason obvious: both models had a very large terminal transition.
Removing only the terminal region changed the correlation dramatically; the final few normalized-depth points accounted for most of the Pearson covariance.
Again, this is not evidence that your eight-model r = 0.789 is terminal-transition driven. It is a different task, two models, and a different measurement.
What it did demonstrate is that an endpoint-robust view has useful discriminating power.
This seems especially relevant because the independent depth-geometry paper by Bhattacharya & Kolli also reports strong late-depth structure, including Procrustes residual peaking at the final transition.
So a lightweight reporting pattern might be something like:
raw profile correlation
+
rank-based correlation
+
one endpoint-robust or detrended view
rather than replacing the existing Pearson result.
For example:
mean pairwise Pearson, full profile
mean pairwise Spearman, full profile
mean pairwise Pearson, excluding final transition
would already tell a future reader whether the coherence is:
- broadly distributed across depth,
- primarily monotonic/rank-like,
- or concentrated in a few common high-amplitude transitions.
If the 0.789 survives those views, that would make the cross-model result substantially more informative.
If it changes, that is also useful, because it tells you what the shared structure actually consists of.
Temporal common mode: one conditional check rather than a blanket requirement
I would be a little more careful here than simply saying “control for response length.”
The right check depends on the definition of your local instability c_t.
For an additive/cumulative trajectory quantity, generated length can mechanically dominate the statistic. In a simple random-walk sanity check, for example, total path length and total accumulated turning angle become almost perfectly correlated with sequence length even when every local step is sampled from exactly the same distribution.
But local means or normalized local quantities need not have that dependence.
That distinction matches the result in Gjølbye et al.: their raw generation-time trajectory geometry required explicit length adjustment, after which meaningful domain-dependent structure remained.
So I would make this conditional:
If c_t or the common-mode summary accumulates across tokens:
compare raw and length-adjusted / length-matched results.
If c_t is already a strictly local normalized quantity:
length may be much less important;
event alignment or normalized generation progress may be the more relevant control.
This could also help interpret the 0/8 LOMO temporal result.
If the common mode remains absent after a reasonable alignment/length check, that is stronger evidence that the temporal trajectory is genuinely model/task conditional.
If it partially returns after alignment, the conclusion becomes more specific:
there is no universal trajectory in raw token time, but some lower-dimensional temporal organization may exist after accounting for generation geometry.
Either outcome would fit the broader progressive-structuring framework.
Where I would *not* push the interpretation yet
I think the manuscript is already appropriately careful about several boundaries, especially:
Representation ≠ Function ≠ Behavior.
That is consistent with the broader probing literature. For example, Hewitt & Liang’s control-task work is a useful general reminder that decodability alone does not establish that a represented feature is functionally used, while the Tuned Lens gives a more task-linked view of how predictions evolve across layers.
Given the claims you actually make, I would not treat causal intervention as the immediate missing requirement.
Activation patching, steering, or cross-model stitching could become interesting later if the goal shifts from:
“is there reproducible representational organization?”
to:
“does this particular depth region causally implement this function?”
But that is a different research stage.
Similarly, I would not infer from the current results that:
- one universal sequence of cognitive stages has been identified;
r = 0.789 proves equivalent computations at equivalent relative depths;
- the event-depth association localizes a unique function;
- the temporal null result means temporal organization does not exist;
- architecture family is irrelevant in general.
Your manuscript already avoids most of those jumps.
The current narrower statement seems much easier to defend:
some aspects of depth-resolved organization replicate across a heterogeneous small-model panel, while temporal and categorical similarities are substantially more conditional.
The fact that one earlier positive temporal hypothesis disappeared under panel expansion is part of the evidence for that narrower statement, rather than an embarrassment to hide.
If I had to choose only two next checks
I would probably choose these because they seem very low-cost relative to the information they add:
1. Separate non-exchangeability from depth orientation.
Keep the existing full permutation result, but add one secondary comparison that preserves much more local structure or explicitly reverses/tests absolute depth.
This answers:
Is the result mainly “neighboring layers form a coherent path,” or does it also carry an early → late organization?
2. Add one robust view of the r = 0.789 depth-profile coherence.
For example:
raw Pearson
+ Spearman
+ Pearson without the terminal transition
or:
raw Pearson
+ detrended Pearson
+ first-difference correlation
No need to run every variant. One small robustness panel would already clarify whether the shared structure is coarse, local, endpoint-driven, rank-like, or genuinely distributed.
Everything heavier — larger causal interventions, steering, model stitching, more elaborate probes — could come later if these descriptive distinctions remain stable.
So my current read would be:
Yes, there are independent results that look meaningfully related to the depth side of your finding. The interesting part of your eight-model result may be less “there is one universal trajectory” and more “depth organization survives heterogeneity while temporal organization does not.” The next high-information step is probably not a larger claim, but separating local adjacency, depth orientation, and broad cross-model profile coherence into slightly cleaner tests.