Role Tokens Are Representationally Invisible

A Decomposition of Chat-Format Effects in Cogito 671B

authorsAntra Tessera, Luxia, Janus
publishedMay 2026
Working draft, written with Claude Opus 4.7. Numbers, figures, and claims are subject to revision and author review.

Chat-formatted prompts wrap conversation content in structural scaffolding—special role tokens (<|User|>, <|Assistant|>), end-of-turn markers (<|EOS|>), system prompts, the alternation of speakers. A post-trained LLM behaves as “the Assistant” inside this scaffolding and differently without it, but it is not obvious which elements the model actually reads. We measure this directly. Using Cogito v2.1 671B (a DeepSeek V3 / MLA derivative), we capture pre-RoPE key vectors (K_nope) at the generation point across 54 multi-speaker conversation transcripts, ~15 turn-aligned depths up to 26k tokens, and 61 layers, comparing fifteen format variants that selectively ablate individual structural elements of the chat scaffold—including substitutions of the role-marker substrate for Anthropic-style text (\nHuman:/\nClaude:) and for random nonsense (\nZyx:/\nQwp:). We find: (i) the initial <|Assistant|> token, role-token identity, role-marker substrate (special vs. text vs. nonsense), role assignment, and system prompts are all representationally invisible (cos ≥ 0.99); replacing <|Assistant|> with <|tool_calls_begin|> produces literal 1.0 cosine at shallow depths, and substituting text or nonsense substrates yields cos = 0.999. This equivalence is behavioral, not merely representational: across the text- and nonsense-substrate swaps the model’s next-token prediction is identical in 100% of n = 811 comparisons each, with zero disagreements—the role marker leaves no trace in the output, not just in the key it writes. (ii) Only the repeated <|EOS|>-marked turn-segmentation pattern produces persistent representational divergence (cosine near zero, well below the 0.28 cross-conversation noise floor, flat from 2k to 26k tokens); and the signal is keyed to the <|EOS|> token specifically—replacing it with a text \n\n delimiter while keeping role markers reproduces the full divergence (cos ≈ 0). The substrate picture is thus sharply asymmetric: the role marker’s substrate is invisible, the terminator’s substrate is the entire effect. What chat formatting installs at the generation point is temporal segmentation structure, not persona-marking labels. We close with implications for accounts of how post-training conditions the Assistant persona, and an extrapolation: single-block CLI-simulation prefill should be representationally transparent for Anthropic-style models, where the format’s internal \n\nName: rhythm matches the model’s native turn markers, even at long context.

Introduction

Post-trained LLMs behave very differently inside chat formatting than outside it: wrapped in the chat scaffold the model acts as “the Assistant,” while the same content as a raw continuation elicits something else. The scaffold is a bundle of distinct elements—special role tokens (<|User|>, <|Assistant|>) marking whose turn it is, end-of-turn markers, system prompts, the alternation of speakers—and it is not obvious which of them the model actually reads. Which elements of the chat format carry the signal?

Two loci are plausible: (a) the special role tokens that explicitly mark turns, and (b) the structural pattern of the format itself—repeated turn boundaries, end-of-turn tokens, speaker alternation. The first is the simpler hypothesis: the <|Assistant|> token is the trigger that surfaces assistant behavior, so prompts that lack it (raw narrative continuations, or content packed into a single assistant block as in CLI-simulation framings) should look representationally distinct from the same content wrapped in chat scaffolding. The second locates the effect not in any label but in the temporal structure.

We test this directly on Cogito v2.1 671B, a DeepSeek V3 derivative using Multi-Head Latent Attention (MLA). For each of 54 multi-speaker conversation transcripts and ~15 turn-aligned depth points up to 26k tokens, we tokenize the same content under fifteen format variants that selectively retain or remove individual structural elements (role tokens, EOS delimiters, system prompts, role assignment, role token identity, and the substrate of both the role markers and the turn terminator), feed each through the model, and capture K_nope—the pre-RoPE key component from MLA’s latent decompression path—at all 61 layers at the last token position. K_nope is positionally independent and represents the part of the key vector that future attention reads back through; if format scaffolding installs a persistent processing mode, K_nope should record it.

Our finding is that role tokens are not the locus. The initial <|Assistant|> token, the identity of the role token, the assignment of roles to speakers, and the presence of an elaborate CLI-simulation system prompt are all invisible: cosine similarities of 0.99–1.000. And this is not merely a similarity of internal vectors that might wash out elsewhere in the computation: when we swap the role marker for arbitrary text or even random nonsense, the model’s next-token prediction is unchanged in 100% of 811 comparisons (zero disagreements)—the invisibility is behavioral. The only chat-format element that produces persistent representational divergence is the repeated <|EOS|>-marked turn-segmentation pattern itself. This places the action of chat format at the level of temporal segmentation structure, not at the level of persona-marking labels. We return in the Discussion to what this constrains about how post-training conditions the Assistant persona.

Method

Model and capture target. We use Cogito v2.1 671B (DeepseekV3ForCausalLM, 61 transformer layers, MLA attention) running under vLLM 0.16.0 on 8× NVIDIA B200 (TP=8), enforce_eager=True for hook compatibility, fp8 KV cache. We capture K_nope: the pre-RoPE key component obtained from MLA’s latent decompression path (fused_qkv_a_proj → split → kv_a_layernorm → kv_b_proj → split → k_nope), concatenated across TP ranks. This is the positionally-independent part of the key vector, captured at the last token position during a single-token-budget prefill. Capture is implemented as a vLLM entry-point plugin that hooks attention layers in every TP worker subprocess.

Data. 54 multi-speaker transcripts (~27k tokens each), generated by Cogito as the primary speaker with Opus 3 / Opus 4.5 / Kimi-2 acting as auditors in a backrooms-style setup. 34 transcripts were generated under chat formatting; 20 under raw continuation. Both formats are reconstructed from each transcript at capture time, so the content is identical between Format A and Format B for every measurement.

Format variants. We compare each of fifteen Format A variants against an appropriately matched Format B anchor. Format B for the six different-structure variants (NoRoles, Prefix-EOT, Roles-NoEOS, NoRoles-EOS, Prefix-EOS, TextDelim) is a full chat reference built by the same direct token construction as Format A—each turn rendered as role marker + content + <|EOS|> and concatenated—rather than the tokenizer’s apply_chat_template. We use this manual construction because the chat template inserts extra scaffolding (most consequentially a trailing generation prompt) that itself differs from a plain manually-built full chat at cos K_nope ≈ −0.03; anchoring on the template would fold that construction artifact into every structural comparison. (We verified this directly: a manually-built full chat with no structural change scores ≈ −0.03 against the template, identical to the apparent “divergence” it would impute to the ablated variants.) The remaining nine variants pair Format A against a matched Format B that already shares the same repeated turn-boundary structure (Table 1). Of these nine, two—TextRoles and NonsenseRoles—specifically probe the substrate of the role markers (Anthropic-style text or random labels in place of Cogito’s special tokens).

Table 1. Fifteen format variants. Format A varies; Format B is full chat unless noted. Notation: <U> = <|User|>, <A> = <|Assistant|>, EOT = <|EOT|> (token 128805), EOS = <|end_of_sentence|> (token 1).
Label Format A Isolates
Different-structure variants (Format B = full chat):
NoRoles <BOS>{text<EOT>text...} no role tokens at all
Prefix-EOT <BOS><U>.<A>{text<EOT>...} adds initial <U>.<A> only
Roles-NoEOS <BOS><U>text<A>text<U>... (no EOS) role alternation, no EOS
NoRoles-EOS <BOS>{text<EOS>text...} EOS as delimiter, no roles
Prefix-EOS <BOS><U>.<A>{text<EOS>...} single <A> block, EOS
TextDelim <U>text\n\n<A>text... (roles, \n\n delim) terminator: \n\n vs EOS
Same-structure variants (Format A vs. matched Format B):
Prefix-NL <BOS><U>.<A>{text\n\ntext...} vs raw \n\n prefix only, \n\n delim
Prefix-Chat full chat with vs. without <U>.<A> prefix prefix only, full chat
Prefix-RolesNL roles + \n\n, ± <U>.<A> prefix prefix only, \n\n roles
CliSys CLI-sim system prompt vs. none (chat template) system prompt content
CliSys-Raw same, raw token construction system prompt (raw)
RoleSwap normal vs. swapped <U>/<A> assignment role-to-speaker assignment
TokenSwap <A> replaced by <|tool_calls_begin|> role token identity
TextRoles <U>/<A> replaced by \nHuman:/\nClaude: role-marker substrate (text)
NonsenseRoles <U>/<A> replaced by \nZyx:/\nQwp: role-marker substrate (nonsense)

Measurement. For each transcript, at ~15 turn-aligned depth points spaced ~1.5k tokens apart, we (1) truncate to that turn, (2) tokenize under Format A and Format B independently, (3) prefill each through the model with max_tokens=1, (4) capture K_nope at the last position from all 61 layers. We compute per-layer cosine similarity cos(KAℓ, KBℓ) and aggregate across layers and seeds at each depth bin. The total dataset is 54 seeds × ~15 depths × 15 variants × 2 formats, ~24k forward passes, ~826 measurement points per variant.

Behavioral check (TextRoles, NonsenseRoles). For the two role-marker-substrate variants, we additionally record the top-1 next-token prediction under Format A and Format B at each generation point and compare them directly. This converts the representational comparison (cos K_nope) into a behavioral one: when the cosine is at 0.999, does the model also predict the same continuation token? We collect n = 811 comparisons per variant.

Controls. The cross-conversation noise floor (Format A of seed i vs Format A of seed j, i ≠ j, at matched depths) establishes the baseline for “how different are K_nope vectors for unrelated content under the same format.” We measure cos = 0.282 ± 0.302 (n = 12,200 layer comparisons). Format-comparison cosines below this floor indicate that the two formats produce more representational divergence than two unrelated conversations under the same format.

Results

The fifteen variants partition cleanly into two well-separated clusters—shared-structure variants at cos ≈ 1.0 and broken-structure variants at cos ≈ 0, with nothing in between (Table 2, Figure 1; per-depth detail in Table 3). TextRoles, NonsenseRoles, and TextDelim are discussed in the dedicated paragraphs below.

Table 2. All fifteen format variants and their overall mean K_nope cosine (Format A vs. matched Format B, averaged over depths and layers). For the different-structure variants, Format B is a manually-constructed full chat (see Method); same-structure variants use their matched manual anchor. The split is sharply bimodal, flat across depth (Figure 1); the cross-conversation noise floor is 0.282.
Variant Isolates cos Knope
Same-structure (turn-rhythm shared with Format B):
Prefix-Chat initial <U>.<A> prefix 1.000
TokenSwap role-token identity 1.000
CliSys CLI system-prompt content 1.000
RoleSwap role-to-speaker assignment 0.999
TextRoles role-marker substrate (text) 0.999
NonsenseRoles role-marker substrate (nonsense) 0.999
Prefix-RolesNL prefix (roles + \n\n) 0.994
CliSys-Raw system prompt (raw construction) 0.992
Prefix-NL prefix (\n\n delimiters) 0.992
Different-structure (Format A lacks the per-turn <EOS> rhythm):
Roles-NoEOS EOS removed, roles kept 0.006
TextDelim terminator \n\n vs <EOS> 0.006
NoRoles no role tokens 0.005
Prefix-EOT initial prefix only 0.005
Prefix-EOS single <A> block + EOS 0.001
NoRoles-EOS EOS delimiter, no role tokens 0.000
Figure 1
Figure 1Format-effect decay curves for the twelve original variants. Top cluster (right axis, same-structure variants): Prefix-NL, Prefix-Chat, Prefix-RolesNL, RoleSwap, TokenSwap, and the two CliSys variants—all cases in which Format A and Format B share the repeated turn-boundary structure and differ only in prefix, role-token identity, role assignment, or system prompt. Cosine similarities are 0.99–1.000, flat across depth. Bottom cluster (left axis, different-structure variants): NoRoles, Roles-NoEOS, NoRoles-EOS—variants in which Format A lacks the chat format’s repeated <EOS>-marked segmentation. Cosine similarities are  ≈ 0 (anti-correlated at the output layer), far below the 0.282 cross-conversation noise floor, through 26k tokens. The two role-marker-substrate variants TextRoles and NonsenseRoles (not pictured) join the top cluster at cos  ≈ 0.999.

Table 3. Mean K_nope cosine similarity (Format A vs Format B), across depth bins. Format B is a manually-constructed full chat (not the chat template; see Method). Bold values are within 0.01 of 1.0. The cross-conversation noise floor is 0.282; the different-structure variants sit near zero, well below it.

Depth NoRoles Pref-EOT Rol-NoEOS NoRol-EOS Pref-EOS Pref-NL Pref-Chat TokSwap
~0k −0.001 0.001 0.001 −0.004 −0.002 0.987 0.999 1.000
~2k −0.001 0.001 0.001 −0.004 −0.002 0.992 1.000 1.000
~8k 0.006 0.007 0.008 0.003 0.003 0.993 1.000 1.000
~16k 0.006 0.007 0.007 0.001 0.002 0.994 1.000 1.000
~24k 0.003 0.003 0.004 −0.001 −0.002 0.992 1.000 1.000
mean 0.005 0.005 0.006 0.000 0.001 0.992 1.000 1.000

Same-structure cluster — effectively identical. When both formats share the repeated turn-boundary structure, every other element of chat scaffolding is invisible. Specifically:

Figure 2 contrasts the three prefix tests (Prefix-NL, Prefix-Chat, Prefix-RolesNL) with Roles-NoEOS (role tokens without the EOS pattern) to show that the prefix-invisibility result holds whether or not the surrounding format also uses EOS.

Role-marker substrate (TextRoles, NonsenseRoles). TokenSwap establishes that the <|Assistant|> token has no special status within the set of trained special tokens. Two further variants replace the role tokens with non-special-token substrate while preserving the <|EOS|>-marked turn-segmentation pattern: TextRoles (\nHuman:/\nClaude:, Anthropic-style text markers, matched against a hand-built <|U|>/<|A|>+EOS Format B with identical layout) and NonsenseRoles (\nZyx:/\nQwp:, random labels with the same structure). Both sit in the same-structure cluster: cos K_nope = 0.998–0.999 across depth, climbing slightly with depth, well above the cross-conversation noise floor. In addition, a per-position next-token comparison (Format A vs Format B at the same generation point, n = 811 per variant) finds 100% top-1 token agreement with zero disagreements, together with vanishing distributional divergence (mean Jensen–Shannon divergence 4×10−4 nats over the top-100 logprobs; top-5 token-set overlap ≈ 70%). The greedy next token is identical—and the full distribution extremely close—whether the speaker is marked by a trained special token (<|Assistant|>), by semantically meaningful text (Claude:), or by random nonsense (Qwp:). The role-marker substrate is thus invisible behaviorally, not only representationally. This test is decisive precisely because it is discriminating: run on the different-structure variants the very same measurement collapses to ≈ 0% top-1 agreement with near-maximal divergence (Table 5), so the perfect agreement here reflects genuine substrate-invariance rather than a probe that cannot tell formats apart. A high K_nope cosine alone leaves open that the difference surfaces somewhere the probe does not look; an identical greedy prediction with vanishing JS divergence at the generation point closes that loophole.

Terminator substrate (TextDelim) — the asymmetry. The substrate-independence of role markers does not extend to the terminator. TextDelim holds the role markers fixed (special <U>/<A> on both sides) and replaces only the per-turn terminator: \n\n in Format A vs. <EOS> in Format B (full chat). This is the same roles+\n\n construction that is representationally identical to itself-plus-prefix (Prefix-RolesNL, cos = 0.994); yet compared against the EOS-terminated chat it collapses to cos K_nope ≈ 0 — squarely in the different-structure cluster, far below the 0.282 noise floor. Its per-layer profile is superimposable on the other different-structure variants: matched at the input layer (cos ≈ +0.99 at L00, where the final token is shared) but diverging through the stack to anti-correlation at the output (L60 ≈ −0.3). So the substrate picture is sharply asymmetric: swapping the role marker’s substrate (special → text → nonsense) is invisible, but swapping the terminator’s substrate (<EOS> → \n\n) is the entire effect. The segmentation signal is keyed to the trained <EOS> token specifically — not to “a repeated boundary,” which \n\n supplies but which carries none of it. This is consistent with <EOS> functioning as a trained attention-relevant marker whose identity, unlike the role labels’, cannot be repainted.

Figure 2
Figure 2Left: the prefix <U>.<A> is invisible across three delimiter regimes: Prefix-NL (\n\n), Prefix-Chat (full chat with <EOS>), and Prefix-RolesNL (roles + \n\n). All three sit at  ≥ 0.99 across depth. Right: the Prefix-Chat anchor (orange,  ≥ 0.999) compared against Roles-NoEOS (role alternation without <EOS>, purple,  ≈ 0). When the only difference is whether the format contains the repeated EOS-segmentation pattern, divergence is large; when EOS-segmentation is shared, all other elements wash out.

Different-structure cluster — divergent and persistent. Variants whose Format A lacks the repeated <EOS>-marked turn-boundary pattern diverge from full chat all the way to cos ≈ 0—in fact anti-correlated at the output layer (L60 ≈ −0.32, L00 ≈ +0.99)—far below the 0.282 noise floor, and they do not converge with depth through 26k tokens. The variants within this cluster are statistically indistinguishable from one another (all means in [0.000, 0.006]): neither an initial <U>.<A> prefix (NoRoles ≈ Prefix-EOT), role-token alternation without EOS (Prefix-EOT ≈ Roles-NoEOS), nor the <EOT>-vs-<EOS> delimiter identity (NoRoles ≈ NoRoles-EOS) moves the result. The contribution decomposition (Table 4) makes the structure explicit: only the presence of the full per-turn <EOS> rhythm separates this cluster from full chat.

Table 4. Contribution to format divergence by structural element.

Element Contribution to 1 - cos
Initial <U>.<A> prefix none (NoRoles≈Prefix-EOT, Prefix-Chat = 0.999)
Role-token alternation, no EOS none (Prefix-EOT ≈ Roles-NoEOS)
Role-token identity (<A> vs <|tool_calls_begin|>) none (TokenSwap ≈ 1.000)
Role-marker substrate (text / nonsense) none (TextRoles, NonsenseRoles = 0.999)
Role-to-speaker assignment none (RoleSwap = 0.999)
System prompt content (CLI sim) none (CliSys = 0.999)
Delimiter token identity (EOT vs EOS) negligible (NoRoles ≈ NoRoles-EOS, both ≈ 0)
Terminator substrate (<EOS> vs \n\n) primary (TextDelim ≈ 0)
Repeated <EOS>-marked turn segmentation primary driver

No depth-dependent washout. Across all five Cluster 2 variants, the format-comparison cosine is essentially flat from 2k to 26k tokens (Figure 1, Table 3). The hypothesis that format effects diminish with accumulating context is falsified for Cogito 671B at the depths and structural ablations tested here.

The terminator effect is dose-dependent. Substituting the per-turn <EOS> with a single newline on a growing fraction of turns (length preserved, compared against a manually-constructed full chat) yields a smooth, monotonic decline in cos K_nope: ≈ 1.0 for a single substitution, ≈ 0.94 at ~7% of turns, and ≈ 0 once all turns are substituted (the TextDelim limit). The two-cluster dichotomy is thus the continuous limit of a graded effect rather than a discrete switch; we found no evidence of a threshold or non-monotonicity along this axis.

Behavioral mirror: next-token agreement tracks the clusters. We complement the representational decomposition with a direct behavioral readout. At each aligned generation point we greedily decode one token under Format A and Format B and compare the next-token outputs by top-1 argmax agreement, top-5 token-set overlap, and Jensen–Shannon (JS) divergence over the top-100 logprobs (n ≈ 810–830 comparisons per variant). The behavioral signal mirrors the representational two-cluster structure exactly (Table 5). Every same-structure variant agrees with full chat on the next token ≥ 93% of the time—and exactly 100% for the substrate manipulations (TokenSwap, TextRoles, NonsenseRoles, CliSys, Prefix-Chat)—with JS divergence at or below 10−3 nats. Every different-structure variant disagrees on essentially every token (≤ 1% top-1 agreement) with near-maximal divergence (JS 0.63–0.67 nats, against the ln 2 ≈ 0.693 ceiling for disjoint distributions). There is no middle ground: behaviorally the formats are either interchangeable at the token level or almost completely disjoint, partitioned by exactly the same <EOS>-segmentation criterion that governs K_nope.

Table 5. Behavioral next-token agreement between Format A and full chat, by variant. Top-1: greedy argmax agreement; Top-5: mean top-5 token-set overlap; JS: mean Jensen–Shannon divergence (nats) over the top-100 logprobs. The two clusters partition exactly as in the representational analysis.
Variant Top-1 Top-5 JS
Same-structure cluster (shares <EOS> turn rhythm)
TokenSwap (V12) 100.0% 83.8% 0.0003
Prefix-Chat (V8) 100.0% 81.8% 0.0002
CliSys (V10a) 100.0% 92.6% 0.0000
TextRoles (V13/V15) 100.0% 71.6% 0.0004
NonsenseRoles (V14) 100.0% 69.1% 0.0004
RoleSwap (V11) 99.9% 77.2% 0.0008
Prefix-RolesNL (V9) 95.5% 93.6% 0.0031
CliSys-Raw (V10b) 95.5% 90.8% 0.0044
Prefix-NL (V7) 93.2% 92.3% 0.0057
Different-structure cluster (no repeated <EOS> rhythm)
NoRoles (V2) 0.4% 0.9% 0.666
Prefix-EOT (V3) 0.4% 0.9% 0.666
Roles-NoEOS (V4) 0.4% 0.7% 0.667
NoRoles-EOS (V5) 1.0% 9.0% 0.626
single-block + EOS (V6) 1.0% 8.6% 0.627

Discussion

Implications for the Persona Selection Model. The Persona Selection Model [1]—which treats the post-trained Assistant as a learned persona that runtime context conditions the model into enacting—is silent on which prompt elements perform persona conditioning—it states that runtime context selects from a learned distribution over Assistant variants, without specifying the mechanism. Our results constrain that mechanism: persona selection is not keyed to the role-token labels. The fact that swapping the <|Assistant|> token for <|tool_calls_begin|> produces literal 1.0 K_nope cosine—meaning the model’s pre-RoPE key vectors are bit-identical at the generation point—rules out any account in which the <|Assistant|> token serves as a persona-selection trigger that acts on the model’s internal state at the point where future attention will read it. Whatever the conditioning mechanism is, it is not a label-keyed circuit.

This is consistent with the more sophisticated reading of PSM in which persona conditioning is distributed across the entire prompt context. Our results are a positive constraint: that distribution lives in content + segmentation structure, not in the special-token labels that mark turns.

The EOS-segmentation pattern. The single chat-format element that produces persistent representational divergence is the repeated <EOS> token at every turn boundary—closer to a syntactic property (the temporal rhythm of segmenting the stream into bounded chunks) than to a semantic or persona-marking one. The effect is dose-dependent: substituting the terminator on a growing fraction of turns moves K_nope smoothly from full-chat-identical toward the no-EOS regime (see Results), so the two clusters are the limiting cases of a graded continuum rather than a discrete trigger. Whether the model’s reliance on <EOS> is a training artifact (EOS learned as a strong attention-relevant marker during instruction tuning) or reflects a deeper property of how the architecture processes segmented vs. continuous streams remains open.

Behavioral consistency. The next-token agreement test (Table 5) makes this concrete. Where K_nope says “same” (same-structure cluster), the model’s greedy next-token prediction agrees with full chat 93–100% of the time—including 99.9% for RoleSwap and 100% for TokenSwap—so the representational identity carries through to behavior by measurement, not merely by assumption. Where K_nope says “different” (different-structure cluster), the predictions disagree on essentially every token (≤ 1% agreement); and under raw-continuation generation (Format A of these variants) Cogito cannot sustain coherent multi-speaker conversation at depth: outputs degrade to short, fragmented turns, system-shutdown sequences, and emoji gibberish. The same structural absence that drives cos ≈ 0 in K_nope thus also produces both token-level divergence and unstable generation. We see no evidence in this experiment for a persona-state-machine view in which role tokens silently install a behavior-shaping mode invisible to local probes.

Format effects are model-relative; an extrapolation to Anthropic-style prefill. The “turn marker” that matters is whatever the target model was trained to read. For Anthropic-trained models, turn markers are text patterns (\n\nHuman:, \n\nAssistant:); for Cogito and most modern open-weight instruct models, they are special tokens (<|User|>, <|Assistant|>, <|EOS|>). Two practitioner formats are often conflated:

For Cogito, the second format is in our different-structure cluster, and we now measure this directly rather than infer it. TextDelim — role-marked turns delimited by \n\n instead of <|EOS|> — yields cos K_nope ≈ 0, the same as a prompt with no turn structure at all. The \n\nName: rhythm is not Cogito’s trained turn-marker pattern, so Cogito sees the backrooms transcript as a prompt without its native <|EOS|> boundaries. CliSys/CliSys-Raw separately rule out system-prompt content as the driver. Crucially, TextDelim contrasts with TextRoles: text role markers are invisible (0.999) but a text terminator is not (≈ 0) — the segmentation signal is keyed to the <|EOS|> token specifically, so substituting a text delimiter breaks it.

For an Anthropic-style model, however, the same content carries Anthropic’s native turn markers internally. The same-structure logic of our results then predicts that single-block CLI-simulation prefill should be representationally indistinguishable from running the same conversation as a sequence of API turns: both prompts share the trained \n\nName: segmentation rhythm, and our experiment finds that when this rhythm is shared, all other elements (prefix, role-token identity, role assignment, system prompt) wash out. Our same-structure plateau is flat from 2k to 26k tokens with no drift toward divergence, so this transparency would be expected to persist at long context.

TextRoles strengthens this extrapolation from the Cogito side. When we substitute Anthropic-style text markers (\nHuman:/\nClaude:) for Cogito’s native <|User|>/<|Assistant|> tokens while holding the <|EOS|>-marked segmentation rhythm fixed, Cogito’s K_nope cosines remain at 0.999 with 100% top-1 logprob agreement (see Results). So the role-marker substrate—special token vs text vs nonsense—is invariant on Cogito; what matters is the segmentation rhythm. The remaining extrapolation step is whether an Anthropic-trained model treats its native \n\nName: rhythm with the same transparency that Cogito treats its native <|EOS|> rhythm. We did not test this directly, but it is now the only step left in the chain.

Limitations

We measured K_nope at the last token position of the prefill. This captures what future attention will read back through, but does not capture (a) V projections at non-final positions, (b) the post-RoPE key component, (c) MLP outputs, or (d) attention-pattern differences that integrate over the prefix. Our claim is specifically about the pre-RoPE key vector at the generation point.

We tested one model (Cogito 671B). The same experiment on standard MHA architectures, on smaller models, on base (non-instruct) models, and on instruct models with very different training mixtures would all be informative. We particularly note that the EOS-driven divergence may be specific to instruction-tuned models that learned EOS as an attention-relevant marker.

The behavioral observation about generation degradation under raw continuation is a separate experiment that we have not run with controls (e.g., generating under RoleSwap or TokenSwap to confirm role-token identity is also behaviorally invisible). It is a qualitative observation derived from the transcript-generation pipeline that fed this experiment, included for consistency with the representational finding rather than as a primary result.

Acknowledgments

Compute on the experimental cluster. Capture infrastructure built on vLLM 0.16.0 with a custom entry-point plugin for cross-TP-worker hook propagation; see the project repository for the full plugin.

References

  1. Anthropic Alignment Team (2026). The Persona Selection Model: Why AI Assistants might Behave like Humans. Online. alignment.anthropic.com/2026/psm.