Representation changes across varying clinical input conditions: A dual-metric validation study of eight transformer architectures with length controls | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Representation changes across varying clinical input conditions: A dual-metric validation study of eight transformer architectures with length controls Yngve Mikkelsen This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-9237602/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract Background: Large language models are increasingly deployed in clinical decision support, yet the stability of their internal representations across diverse clinical input conditions remains poorly characterised. It is unclear whether changes in representation reflect geometric reorganisation (magnitude and directional shifts) or simple scaling artefacts. Methods: We used a four-group validation design across eight transformer models (2018–2023), including two modern architectures (Llama-2-7B, Mistral-7B). Group 1 comprised 450 MT samples of clinical notes, stratified into simple/moderate/complex length-proxied strata (n = 150 each). Group 2 comprised 450 matched synthetic texts. Groups 3–4 comprised 600 length-controlled texts isolating pure length effects. From the final hidden layer (excluding special tokens), we computed (i) per-token embedding magnitude (mean L2 norm) and (ii) mean pairwise cosine similarity between token embeddings. Analyses used one-way ANOVA with Bonferroni correction across eight models (α = 0.00625), Welch t-tests, and bootstrap 95% confidence intervals (10,000 iterations). Results: In Group 1, seven of eight models showed significant differences in magnitude across strata (p < 0.00625). Six of these seven also showed significant directional changes (cosine similarity changes of 3–26%), indicating geometric changes rather than scaling alone. BioBERT and ClinicalBERT showed the largest dual-metric effects (magnitude +8.5% and +7.3%; cosine −25.9% and −25.0%). Llama-2-7B showed no significant magnitude change (−0.6%, p = 0.062) and a non-significant simple-to-complex cosine change (+3.6%, p = 0.126). Mistral-7B showed a small but significant magnitude increase (+1.9%, p < 0.001) and significant directional convergence (cosine +14.4%, p < 0.001). Length-controlled analyses confirmed substantial length effects on both metrics. Conclusions: In older models, representation changes across length-proxied strata of clinical complexity are predominantly geometric. Modern architectures exhibit smaller magnitude shifts and a convergence trend in cosine similarity, in contrast to directional divergence in older models. Whether these representation-level changes translate into differences in downstream clinical task performance remains to be established. Integrative & Complementary Medicine Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Figure 6 Figure 7 Introduction Large language models are increasingly deployed in clinical settings for tasks such as documentation, clinical decision support, and information retrieval( 1 – 4 ). The reliability of these systems depends in part on the stability of their internal representations—the high-dimensional embeddings that encode semantic meaning. Understanding how these representations change under varying input conditions is relevant to characterising model behaviour in clinical applications( 5 , 6 ). Task-level performance degradation Performance limitations in transformer language models have been documented at the task level. Liu et al. ( 7 ) demonstrated the ‘lost in the middle’ phenomenon: language model performance is highest when relevant information occurs at the beginning or end of the input context, and degrades when it is in the middle of long contexts. This U-shaped performance curve persists even in explicitly long-context models. However, task-level performance metrics do not directly characterise the quality of internal representations. A model might show degraded task performance while maintaining stable representations—or conversely, show representation changes that do not immediately manifest as task failures. This distinction motivates representation-level analysis as a complementary approach. Geometric properties of embedding spaces Research has examined the geometric properties of transformer representations. Ethayarajh ( 8 ) found that contextualised word representations are anisotropic—concentrated in a narrow cone rather than uniformly distributed. Gao et al. ( 9 )identified this as the ‘representation degeneration problem.’ Timkey and van Schijndel ( 10 ) discovered that ‘rogue dimensions’ dominate similarity measures, with a mismatch between dimensions important for similarity and those important for model behaviour. Cai et al. ( 11 )offered a more nuanced view, identifying isolated isotropic clusters within globally anisotropic spaces. Length-induced embedding changes Zhou et al. ( 12 ) characterised ‘length collapse’ in transformer-based embedding models: longer text embeddings cluster in a narrower space, producing distributional differences across embeddings of varying text lengths. They theoretically demonstrated that the self-attention mechanism acts as a low-pass filter, with longer sequences increasing the rate of attenuation. Clinical documentation has distinctive characteristics, including specialised terminology and a correlation between document length and clinical complexity. Whether length-associated changes result from sequence length itself or from the clinical information density that accompanies longer texts has not been systematically examined. Domain-specific clinical language models Domain-specific language models for healthcare include BioBERT( 13 ), pre-trained on PubMed abstracts, and ClinicalBERT( 14 ), pre-trained on MIMIC-III clinical notes. More recently, Llama-2 ( 15 ) and Mistral ( 16 ) represent modern transformer architectures with improved positional encoding schemes. Comparative evaluations have shown that domain-specific models outperform general models on clinical NLP benchmarks( 17 , 18 ), but whether domain-specific training or modern architectural improvements affect representation stability remains unexamined. Study objectives This study characterises complexity-associated representation changes in clinical language models using two complementary metrics—embedding magnitude (L2 norm) and directional alignment (cosine similarity)—to distinguish geometric reorganisation from simple scaling artefacts. We address five questions: ( 1 ) Do transformer models exhibit significant representation changes across varying clinical complexity? ( 2 ) Are these changes geometric (affecting both magnitude and direction) or simple scaling artefacts? ( 3 ) Do findings generalise across real and synthetic clinical text? ( 4 ) Are changes driven by complexity or sequence length? ( 5 ) Do modern architectures (2023) show different patterns from older models (2018–2020)? We analyse 1,500 clinical texts (12,000 model–text observations) across eight transformer models spanning 2018–2023. Materials and methods Study design We conducted an observational study of changes in representation across input conditions in eight transformer models, using four experimental groups. This study was not pre-registered; all analyses should be considered exploratory. All analyses were carried out in February 2026. Experimental groups Group 1: Real clinical text. We extracted 450 clinical notes from the MT Samples corpus, a publicly available collection of medical transcriptions. The texts were stratified by complexity into three levels (n = 150 per level): Simple (mean 6.7 ± 1.8 words), Moderate (mean 21.9 ± 3.2 words), and Complex (mean 65.8 ± 7.6 words). We acknowledge that this operationalisation conflates complexity with length; Groups 3–4 address this confound. Group 2: Synthetic matched text. 450 synthetic clinical texts, matched to Group 1 for complexity and length (n = 150 per level), were constructed by the first author (a physician). Group 3: Length-controlled synthetic text. 300 synthetic texts with matched lengths but identical simple content at two lengths (~ 22 and ~ 63 words; n = 150 per condition), isolating pure length effects. Group 4: Length-controlled real text. 300 texts from MT Samples, padded with normal clinical findings to match moderate and complex lengths (n = 150 per condition). Models We evaluated eight transformer models across four architectural families. General-purpose: GPT-2 Small (117M parameters, 2019( 19 ); gpt2), GPT-2 Medium (345M, 2019; gpt2-medium), GPT-2 XL (1.5B, 2019; gpt2-xl), and BERT Base (110M, 2018; bert-base-uncased). Domain-specific: BioBERT v1.1 (110M, 2020; dmis-lab/biobert-v1.1) and ClinicalBERT (110M, 2019; emilyalsentzer/Bio_ClinicalBERT). Modern: Llama-2-7B (7B, 2023; meta-llama/Llama-2-7b-hf) and Mistral-7B (7B, 2023; mistralai/Mistral-7B-v0.1). All models were run at full precision (float32) using HuggingFace Transformers v4.36.0 with PyTorch v2.1.0 on an NVIDIA H100 GPU (80GB HBM3). Metrics We assessed two complementary representation metrics from the final hidden layer, excluding special tokens: ( 1 ) embedding magnitude (mean per-token L2 norm) to capture scale, and ( 2 ) directional alignment (mean pairwise cosine similarity between token embeddings) to capture geometry. If inputs induce only scaling changes, magnitude will change while cosine similarity remains stable; geometric reorganisation is indicated by changes in both metrics. For interpretation, we classified simple→complex effects per model as geometric, scaling-only, directional-only, or none, based on Bonferroni-corrected significance for each metric. Let token embeddings from the final hidden layer be hₜ ∈ ℝᵈ for tokens t = 1,…,T. The embedding magnitude for a text was M = (1/T)∑ₜ‖hₜ‖₂. Directional alignment was C = (2/(T(T − 1)))∑_{i < j} (h i ·hⱼ)/(‖h i ‖₂‖hⱼ‖₂). Statistical analysis For each model, a one-way ANOVA tested whether mean embedding magnitude differed across three complexity levels in Group 1. A Bonferroni correction was applied to eight comparisons (α = 0.05/8 = 0.00625). Post hoc comparisons used Welch’s t-test (unequal variances confirmed by Levene’s test for four of eight models). Effect sizes: Cohen’s d and eta-squared (η²). We note that Cohen’s d values for domain-specific models are very large (d > 3.5) because within-group standard deviations are small; this reflects the deterministic nature of the computation rather than an unusually strong behavioural effect. Bootstrap 95% CIs (10,000 iterations, bias-corrected percentile method). Non-parametric Kruskal-Wallis tests confirmed the parametric findings. Cross-corpus validation: Pearson correlation on three data points per model (interpreted as directional indicators). Cosine similarity changes were computed from mean values across conditions. All analyses: Python 3.11, SciPy v1.11, NumPy v1.25. Sample sizes (n = 150 per condition) were determined by corpus construction. Ethics statement This study used only publicly available text data (the MT Samples medical transcription corpus) and pre-trained language models. No human participants were recruited, no biological samples were collected, and no interventions were performed. The MT Samples corpus comprises de-identified medical transcription samples; no personal health identifiers are present in the texts used. The texts were used in accordance with the corpus’s publicly available terms. Institutional review board approval was not required. Data and code availability All data and analysis code are publicly available at: https://github.com/yngvemikkelsen/clinical-representation-stability . The repository contains: all derived datasets in CSV format (per-text-embedding magnitudes and cosine-similarity values for all eight models and ten conditions); Python analysis scripts that reproduce all tables, statistical tests, and figures; and model configuration files with the exact software versions. Raw MT Samples texts are not redistributed; instead, we provide text identifiers and the exact selection procedure to enable reproduction from the original source. Raw data are also provided as Supporting Information (S1–S2 Datasets). Results Primary finding: Magnitude changes across complexity levels Seven of eight models showed statistically significant magnitude changes across complexity levels in real clinical text (Group 1; ANOVA p < 0.001, surviving Bonferroni correction at α = 0.00625). Mistral-7B showed a small but significant magnitude increase (+ 1.9% [95% CI: +1.1%, + 2.8%], d = 0.54, p < 0.001). Llama-2-7B was the only model with a non-significant magnitude change (ANOVA: F(2,447) = 2.22, p = 0.109, η² = 0.010; Welch’s t-test, simple vs complex: t = − 1.87, p = 0.062). Table 1 presents the full results for both metrics; Fig. 1 shows the magnitude distributions. Table 1 Representation magnitude changes across complexity levels in real clinical text (Group 1) Model Simple (SD) Complex (SD) Δ% 95% CI p d η² Cos Δ% Type GPT-2 Small 182.0 (14.4) 189.7 (11.9) + 4.2 [+ 2.6, + 5.9] < .001 0.58 .078 −3.5 G GPT-2 Med 225.9 (29.2) 207.9 (18.8) −8.0 [− 10.2, − 5.6] < .001 −0.73 .113 −17.2 G GPT-2 XL 31.7 (2.8) 33.0 (1.6) + 3.9 [+ 2.3, + 5.6] < .001 0.55 .055 −1.1 S BERT Base 15.1 (0.8) 15.7 (0.4) + 3.8 [+ 2.9, + 4.8] < .001 0.91 .130 −11.1 G BioBERT 13.9 (0.3) 15.0 (0.3) + 8.5 [+ 8.0, + 9.0] < .001 3.85 .697 −25.9 G ClinicalBERT 14.9 (0.3) 16.0 (0.3) + 7.3 [+ 6.8, + 7.8] < .001 3.57 .650 −25.0 G Llama-2-7B 111.0 (4.0) 110.4 (1.8) −0.6 [− 1.2, + 0.0] .062 −0.22 .010 + 3.6 N Mistral-7B 340.6 (15.2) 347.2 (8.5) + 1.9 [+ 1.1, + 2.8] < .001 0.54 .060 + 14.4 G p = Welch’s t-test (simple vs complex) for magnitude. η² = eta-squared from a one-way ANOVA across all three strata (simple/moderate/complex). Cos Δ% = percentage change in mean pairwise cosine similarity (simple→complex); cosine p-values are reported in Table 2 . Type: G = geometric (both magnitude and cosine differences significant), S = scaling-only (magnitude significant, cosine not), D = directional-only (cosine significant, magnitude not), N = no significant simple→complex difference in either metric. Significance threshold: Bonferroni-corrected α = 0.00625 (0.05/8 models). Very large d values for BioBERT/ClinicalBERT reflect low within-group variance. Figure 1. Embedding magnitude distributions across complexity levels (Group 1: real clinical text). Violin plots with individual observations (jittered). Black lines indicate the means. n = 150 per complexity level per model. *** p < 0.001 (Bonferroni-corrected). Dual-metric analysis: Geometric vs scaling effects Table 2 Cosine similarity changes across complexity levels in real clinical text (Group 1). Model Simple Moderate Complex Δ% 95% CI p d KW p GPT-2 Small 0.957 0.946 0.924 −3.5 [− 5.3, − 1.7] < .001 −0.44 < .001 GPT-2 Med 0.941 0.843 0.780 −17.2 [− 20.0, − 14.3] < .001 −1.26 < .001 GPT-2 XL 0.360 0.339 0.356 −1.1 [− 4.8, + 2.8] .582 −0.06 .001 BERT Base 0.372 0.357 0.331 −11.1 [− 13.5, − 8.8] < .001 −1.00 < .001 BioBERT 0.640 0.562 0.474 −25.9 [− 27.2, − 24.6] < .001 −4.02 < .001 ClinicalBERT 0.637 0.552 0.477 −25.0 [− 26.0, − 24.0] < .001 −4.85 < .001 Llama-2-7B 0.316 0.313 0.328 + 3.6 [− 0.9, + 8.6] .126 + 0.18 .011 Mistral-7B 0.301 0.321 0.344 + 14.4 [+ 9.7, + 19.4] < .001 + 0.75 < .001 Mean pairwise cosine similarity between token embeddings from final hidden layer. Δ% = (complex–simple)/simple × 100. p = Welch’s t-test (simple vs complex). d = Cohen’s d. KW p = Kruskal-Wallis non-parametric test across three levels. 95% CI = bootstrap (10,000 iterations). Bonferroni-corrected α = 0.00625. n = 150 per level. The dual-metric analysis (Fig. 2, Table 2 ) indicates that representation changes across the simple→complex contrast are predominantly geometric—encompassing both magnitude and directional shifts—rather than simple scaling artefacts. Five of six older models with significant magnitude changes also showed significant directional changes (cosine similarity decreased by 3–26%). The exception was GPT-2 XL, which showed a significant magnitude increase (+ 3.9%) but no significant simple→complex cosine change (− 1.1%, p = .582), consistent with a scaling-dominant effect. Domain-specific models showed the largest directional reorganisation: BioBERT’s cosine similarity decreased by 25.9% (0.64→0.47, d = − 4.02) and ClinicalBERT by 25.0% (0.64→0.48, d = − 4.85). Among modern architectures, Mistral-7B showed significant directional convergence (cosine + 14.4%, p < .001), whereas Llama-2-7B showed a small, non-significant increase in cosine similarity (+ 3.6%, p = .126). Figure 2. Dual-metric analysis: Embedding magnitude and directional changes across complexity levels. Panel A: Side-by-side comparison of changes in the L2 norm (blue) and cosine similarity (red) across all eight models. Panel B: Scatter plot showing model-specific patterns. Older models (red circles) cluster in the lower-right quadrant (magnitude increase, directional divergence); modern models (blue squares) cluster near the origin or in the upper region (stability or convergence). Figure 3. Mean pairwise cosine similarity across complexity levels (Group 1: real clinical text, 8 models). Higher cosine similarity indicates more directionally aligned token embeddings. All six older models show decreasing cosine similarity with complexity (tokens become more directionally diverse). Modern models show increasing cosine similarity (directional convergence), with statistical significance for Mistral-7B and non-significance for Llama-2-7B in the simple→complex contrast. Figure 4. Effect sizes with 95% bootstrap confidence intervals (Group 1: real clinical text). Forest plot of L2-norm percentage changes. Bootstrap CIs from 10,000 iterations. Red = p < 0.001; grey = not significant. Bonferroni-corrected α = 0.00625. Within-group variability Within-group magnitude variability decreased with complexity across all eight models. Variance ratios (complex/simple) ranged from 0.19 (Llama-2-7B) to 0.76 (BioBERT). The coefficient of variation decreased consistently: e.g., GPT-2 Medium 12.9%→9.1%, BERT Base 5.4%→2.4%, Llama-2-7B 3.6%→1.6% (Fig. 5). This convergence aligns with Zhou et al.’s ( 12 ) theoretical prediction that longer sequences produce more homogeneous representations. Figure 5. Within-group magnitude variability across complexity levels (Group 1: Real clinical text). CV = SD/mean × 100. All models show decreasing CV from simple to complex, indicating representational convergence. Cross-corpus validation Five of the eight models showed high cross-corpus magnitude correlation (r > 0.80): GPT-2 Small (0.98), GPT-2 XL (1.00), BioBERT (0.98), ClinicalBERT (0.89), and Llama-2-7B (0.82). GPT-2 Medium showed a weak correlation (r = 0.34). BERT Base and Mistral-7B showed a reversal (r = − 0.93 and r = − 0.93). These correlations are computed from three data points and should be interpreted as directional indicators (Fig. 6). Figure 6. Cross-corpus validation: Real vs synthetic mean embedding magnitude (8 models). Each point represents one complexity level. Dashed lines show the linear fit. Five of eight models show consistent cross-corpus patterns. Modern architecture performance Both modern architectures (2023) showed patterns that differed from those of older models. Llama-2-7B showed no significant change in magnitude (− 0.6%, p = 0.062) and no significant change in simple→complex cosine (+ 3.6%, p = 0.126). Mistral-7B showed a small but significant magnitude increase (+ 1.9%, p < 0.001, d = 0.54) and significant directional convergence (cosine + 14.4%, p < 0.001, d = + 0.75). Overall, the modern models showed smaller magnitude shifts than older models and a convergence trend in cosine similarity, contrasting with the directional divergence observed in older models. Length-controlled analysis Length-controlled analyses (Fig. 7) confirmed that sequence length alone produces changes in both metrics. In Group 3 (synthetic, constant simple content), all eight models showed magnitude changes with length, and seven of eight showed cosine similarity changes exceeding 3%. Cosine similarity changes with length (− 18.6% to + 13.3%) were often larger than those with complexity, confirming that length is a major driver of representational geometry changes. The direction of length effects varied by architecture: older models generally showed decreasing cosine similarity with length (directional divergence), while modern models showed mixed patterns. Figure 7. Length effects on both magnitude and directional metrics (complexity held constant). Panel A: Synthetic (Group 3). Panel B: Real (Group 4). Blue = L2 norm change; red = cosine similarity change. All eight models tested. Discussion Principal findings This study characterised changes in representation across eight transformer models using dual metrics (magnitude and directional alignment), four experimental groups, 1,500 clinical texts, and 12,000 model–text observations. Five principal findings emerged. First, seven of eight models showed significant changes in magnitude with clinical complexity, with effect sizes ranging from medium to very large. Second, the dual-metric analysis showed that these changes are predominantly geometric—six of seven affected models exhibited both magnitude and directional shifts. Third, domain-specific models (BioBERT, ClinicalBERT) showed the largest effects on both metrics, with cosine similarity decreasing by approximately 25%. Fourth, modern architectures exhibited qualitatively different directional patterns: convergence rather than divergence, with Llama-2-7B showing a non-significant change in magnitude and Mistral-7B showing a small but significant effect. Fifth, length-controlled analyses confirmed that both length and complexity contribute to changes in representation across both metrics. The nature of representation changes: Geometric reorganisation A key contribution of this study is the establishment that complexity-associated changes reflect genuine geometric reorganisation of the embedding space rather than simple scaling artefacts. If embedding magnitudes grew proportionally with complexity, cosine similarity would remain constant. Instead, we observed dramatic directional changes—BioBERT and ClinicalBERT’s cosine similarity dropped by approximately 25%, indicating that token embeddings become substantially more directionally diverse when encoding complex clinical text. This has implications for downstream applications: directional changes could affect nearest-neighbour retrieval, clustering, and any application that depends on cosine similarity-based comparisons. Comparison with prior work Zhou et al. ( 12 ) theoretically demonstrated that self-attention acts as a low-pass filter, producing representational convergence with length. Our findings partially support and extend this framework. The magnitude convergence (decreasing within-group variance) aligns with the low-pass filter prediction. However, the directional divergence we observe (decreasing cosine similarity with complexity) suggests a more nuanced picture: while magnitude variance decreases (convergence), directional diversity increases (divergence). This suggests that complex clinical text activates more diverse representational pathways even as magnitude variance narrows. Scope and nature of claims It is important to distinguish characterisation studies from applied evaluation studies. This work characterises what happens to representations—it documents geometric reorganisation as a function of input complexity. It does not claim that these changes cause downstream task failures. Analogously, Ethayarajh ( 8 ) characterised anisotropy in contextualised embeddings, Gao et al. ( 9 ) characterised representation degeneration, and Zhou et al. ( 12 ) characterised length collapse—none included downstream task evaluation, because the research question was ‘what happens’ rather than ‘does it matter’. Whether geometric changes in representations translate into differences in clinical task performance is a distinct research question that warrants its own rigorous investigation (see Future Directions), and combining both questions in a single study would risk doing neither justice. Clinical implications With this scope in mind, three tentative implications emerge. First, systems using older models (2018–2020) should account for the fact that internal representations change geometrically—in magnitude and direction—with input complexity. These changes could affect cosine-similarity-based retrieval, nearest-neighbour classification, or embedding clustering in clinical pipelines, though empirical verification is needed. Second, domain-specific models showed the largest changes on both metrics, which is noteworthy given their widespread use in clinical NLP. This does not necessarily indicate inferior task performance; larger geometric changes could reflect more differentiated encoding of clinical complexity. Third, both modern architectures show greater stability across both metrics, supporting their consideration for applications requiring consistent representational behaviour. Limitations Several limitations merit consideration. First, although we employed two complementary metrics (L2 norm and cosine similarity), additional measures such as centred kernel alignment (CKA) and representational similarity analysis (RSA) would provide further characterisation. Second, we did not assess downstream task performance. As discussed above, this reflects the study’s scope as a characterisation study rather than a methodological gap; linking geometric changes to task performance is the highest-priority follow-up study. Third, our operationalisation of complexity is confounded with length in Groups 1–2; Groups 3–4 control for the length direction but not the reverse. Fourth, sample sizes were determined by corpus construction rather than by a priori power analysis. Fifth, we analysed only final-layer representations. Sixth, synthetic texts were generated by a single author. Seventh, MT Samples contains template transcriptions, not authentic EHR data. Eighth, cross-corpus correlations are based on three data points. Ninth, this study was not pre-registered, and all analyses are exploratory. Future directions Future work should prioritise linking changes in dual-metric representations to downstream clinical task performance using standardised benchmarks. Layer-wise analysis could reveal where geometric reorganisation occurs. Additional metrics (CKA, RSA) would further characterise the nature of directional changes. Testing additional modern architectures would establish generalisability. Finally, the divergent directional patterns between older and modern architectures warrant mechanistic investigation. Conclusions Representation changes across length-proxied clinical complexity strata are often geometric, involving both magnitude shifts and directional reorganisation of the embedding space. Domain-specific models show the largest dual-metric effects (BioBERT magnitude +8.5%, cosine similarity −25.9%; ClinicalBERT magnitude +7.3%, cosine similarity −25.0%). Modern architectures show comparatively small magnitude shifts (Llama-2-7B: non-significant; Mistral-7B: small but significant) and a convergence trend in cosine similarity, in contrast to the directional divergence observed in older models. Both length and complexity contribute to representation changes. Whether these geometric changes translate into differences in downstream clinical task performance remains an important open question. Declarations Data availability statement All data underpinning the findings are available without restriction. Individual-level measurements for all 12,000 observations (S2 Dataset) and summary statistics, including cosine similarity (S1 Dataset), are provided as Supporting Information. All data and analysis code are available at: https://github.com/yngvemikkelsen/clinical-representation-stability. Author contributions Conceptualisation: YM. Methodology: YM. Software: YM. Formal analysis: YM. Investigation: YM. Data curation: YM. Writing – original draft: YM. Writing – review & editing: YM. Visualisation: YM. Acknowledgements The author thanks the MT Samples community for making clinical transcription samples publicly available. References Lee P, Bubeck S, Petro J. Benefits, limits, and risks of GPT-4 as an AI chatbot for medicine. N Engl J Med. 2023;388(13):1233–9. Singhal K, Azizi S, Tu T, Mahdavi SS, Wei J, Chung HW. Large language models encode clinical knowledge. Nature. 2023;620(7972):172–80. Thirunavukarasu AJ, Ting DSJ, Elangovan K, Gutierrez L, Tan TF, Ting DSW. Large language models in medicine. Nat Med. 2023;29(8):1930–40. Moor M, Banerjee O, Abad ZSH. Foundation models for generalist medical artificial intelligence. Nature. 2023;616:259–65. Rogers A, Kovaleva O, Rumshisky A. A primer in BERTology: What we know about how BERT works. Trans Assoc Comput Linguist. 2020;8:842–66. Belinkov Y, Glass J. Analysis methods in neural language processing: A survey. Trans Assoc Comput Linguist. 2019;7:49–72. Liu NF, Lin K, Hewitt J. Lost in the middle: How language models use long contexts. Trans Assoc Comput Linguist. 2024;12:157–73. Ethayarajh K, editor How contextual are contextualized word representations? Proceedings of EMNLP-IJCNLP; 2019. Gao J, He D, Tan X, Qin T, Wang L, Liu TY, editors. Representation degeneration problem in training natural language generation models. Proceedings of ICLR; 2019. Timkey W, van Schijndel M, editors. All bark and no bite: Rogue dimensions in transformer language models. Proceedings of EMNLP; 2021. Cai X, Huang J, Bian Y, Church K, editors. Isotropy in the contextual embedding space: Clusters and manifolds. Proceedings of ICLR; 2021. Zhou Y, Wu J, Cui G, Zhang S, Zhang Z, Liu Z, editors. Length-induced embedding collapse in transformer-based models. Proceedings of ICLR; 2025. Lee J, Yoon W, Kim S. BioBERT: A pre-trained biomedical language representation model. Bioinformatics. 2020;36(4):1234–40. Alsentzer E, Murphy J, Boag W, editors. Publicly available clinical BERT embeddings. Proceedings of the 2nd Clinical NLP Workshop; 2019. Touvron H, Martin L, Stone K. Llama 2: Open foundation and fine-tuned chat models. 2023. Jiang AQ, Sablayrolles A, Mensch A. Mistral 7B. 2023. Rasmy L, Xiang Y, Xie Z, Tao C, Zhi D. Med-BERT: Pretrained contextualized embeddings on large-scale structured electronic health records. NPJ Digit Med. 2021;4(86). Gu Y, Tinn R, Cheng H. Domain-specific language model pretraining for biomedical NLP. ACM Trans Comput Healthcare. 2022;3(1):1–23. MT Samples: Medical transcription samples. Additional Declarations The authors declare no competing interests. Supplementary Files S1TableNormalityShapiroWilk.csv S1 Table. Normality tests. Shapiro-Wilk test results for all model-condition combinations. S2TableVarianceLevene.csv S2 Table. Variance homogeneity. Levene’s test results for all models across complexity levels. S3TableNonparametricKruskalWallis.csv S3 Table. Non-parametric analysis. Kruskal-Wallis test results confirming parametric findings. Supportinginformation.docx Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-9237602","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":612874447,"identity":"2ec9583c-7f8a-4a5c-a4c8-63cff8d60a98","order_by":0,"name":"Yngve Mikkelsen","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAABHElEQVRIie3RMWuDQBTA8SdCu7zE9aQl9iNccG3JVzE4uJipUBw6CAVdhK4HGfoVhELnk4NMNxdBB6XQOdkslBINhEI5ScYO958O5Xe+hwA63T8OCYLB+8MMyO9T8xSBgbhnEziSZXyK0Cp9J1t4u7YzwYsuqoOXddrALgHHitGlKlLLB5tBhVeTxBMoP1d5LanBJMwZR9dTkTL0XezJzEIqjESschKCiREYOaDLx8j3gVjb4utHBA4LGhMpLMZJUHzAYbAM+CQWHpQeHb6yHIhqMLsMjTajFdrZhgrciPmwS8Ek8Zm4uFetPy2DhndRtSDSb9vuUTjOOm2bXXJ795w+vRIFueH92PD3sn4FMvojnfiyUb/R6XQ63bE9Ti1njHT9LXQAAAAASUVORK5CYII=","orcid":"https://orcid.org/0000-0003-1543-3805","institution":"Saïd Business School, University of Oxford, Oxford, United Kingdom","correspondingAuthor":true,"prefix":"","firstName":"Yngve","middleName":"","lastName":"Mikkelsen","suffix":""}],"badges":[],"createdAt":"2026-03-26 19:27:57","currentVersionCode":1,"declarations":{"humanSubjects":false,"vertebrateSubjects":false,"conflictsOfInterestStatement":false,"humanSubjectEthicalGuidelines":false,"humanSubjectConsent":false,"humanSubjectClinicalTrial":false,"humanSubjectCaseReport":false,"vertebrateSubjectEthicalGuidelines":false},"doi":"10.21203/rs.3.rs-9237602/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-9237602/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":105699555,"identity":"ddbfe074-f1d2-440c-8889-1a805e0dba4b","added_by":"auto","created_at":"2026-03-30 05:25:45","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":9200760,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eEmbedding magnitude distributions across complexity levels (Group 1: real clinical text). \u003c/strong\u003e\u003cem\u003eViolin plots with individual observations (jittered). Black lines indicate the means. n = 150 per complexity level per model. *** p \u0026lt; 0.001 (Bonferroni-corrected).\u003c/em\u003e\u003c/p\u003e","description":"","filename":"Fig1MagnitudeviolinGroup1.png","url":"https://assets-eu.researchsquare.com/files/rs-9237602/v1/f08fe3269889de5be15b9a14.png"},{"id":105699567,"identity":"ba0d8c95-d9eb-42d8-b1cd-de3e0df58345","added_by":"auto","created_at":"2026-03-30 05:25:50","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":2272432,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eDual-metric analysis: Embedding magnitude and directional changes across complexity levels. \u003c/strong\u003e\u003cem\u003ePanel A: Side-by-side comparison of changes in the L2 norm (blue) and cosine similarity (red) across all eight models. Panel B: Scatter plot showing model-specific patterns. Older models (red circles) cluster in the lower-right quadrant (magnitude increase, directional divergence); modern models (blue squares) cluster near the origin or in the upper region (stability or convergence).\u003c/em\u003e\u003c/p\u003e","description":"","filename":"Fig2DualMetricGroup1.png","url":"https://assets-eu.researchsquare.com/files/rs-9237602/v1/7db82218af46b56b8602ffc8.png"},{"id":105699562,"identity":"346584e4-e12a-49c6-84a2-f93835fdba7e","added_by":"auto","created_at":"2026-03-30 05:25:49","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":1478416,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eMean pairwise cosine similarity across complexity levels (Group 1: real clinical text, 8 models).\u003c/strong\u003e Higher cosine similarity indicates more directionally aligned token embeddings. All six older models show decreasing cosine similarity with complexity (tokens become more directionally diverse). Modern models show increasing cosine similarity (directional convergence), with statistical significance for Mistral-7B and non-significance for Llama-2-7B in the simple→complex contrast.\u003c/p\u003e","description":"","filename":"Fig3CosinebycomplexityGroup1.png","url":"https://assets-eu.researchsquare.com/files/rs-9237602/v1/12154e1569267bd101d3b869.png"},{"id":105699552,"identity":"5b5ec24f-b328-4950-b59e-9a55123cc5a5","added_by":"auto","created_at":"2026-03-30 05:25:45","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":435240,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eEffect sizes with 95% bootstrap confidence intervals (Group 1: real clinical text). \u003c/strong\u003e\u003cem\u003eForest plot of L2-norm percentage changes. Bootstrap CIs from 10,000 iterations. Red = p \u0026lt; 0.001; grey = not significant. Bonferroni-corrected α = 0.00625.\u003c/em\u003e\u003c/p\u003e","description":"","filename":"Fig4ForestMagnitudeChangeGroup1.png","url":"https://assets-eu.researchsquare.com/files/rs-9237602/v1/0fb61509902b087d1db0d0e2.png"},{"id":105699563,"identity":"88dfacd1-e94b-4268-9d2d-bad4c1bc3121","added_by":"auto","created_at":"2026-03-30 05:25:49","extension":"png","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":1594897,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eWithin-group magnitude variability across complexity levels (Group 1: Real clinical text). \u003c/strong\u003e\u003cem\u003eCV = SD/mean × 100. All models show decreasing CV from simple to complex, indicating representational convergence.\u003c/em\u003e\u003c/p\u003e","description":"","filename":"Fig5CVbycomplexityGroup1.png","url":"https://assets-eu.researchsquare.com/files/rs-9237602/v1/fb418aef2dd5d927f6fe8788.png"},{"id":105699572,"identity":"99da4bc0-a93f-45d7-be78-2f52e0018e79","added_by":"auto","created_at":"2026-03-30 05:25:53","extension":"png","order_by":6,"title":"Figure 6","display":"","copyAsset":false,"role":"figure","size":1672516,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eCross-corpus validation: Real vs synthetic mean embedding magnitude (8 models). \u003c/strong\u003e\u003cem\u003eEach point represents one complexity level. Dashed lines show the linear fit. Five of eight models show consistent cross-corpus patterns.\u003c/em\u003e\u003c/p\u003e","description":"","filename":"Fig6CrossCorpusRealvsSynth.png","url":"https://assets-eu.researchsquare.com/files/rs-9237602/v1/e90ec1721978e7d3c008315c.png"},{"id":105699558,"identity":"1c9e3b8d-8e97-417f-bb64-667acf5c21cf","added_by":"auto","created_at":"2026-03-30 05:25:45","extension":"png","order_by":7,"title":"Figure 7","display":"","copyAsset":false,"role":"figure","size":2387496,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eLength effects on both magnitude and directional metrics (complexity held constant). \u003c/strong\u003e\u003cem\u003ePanel A: Synthetic (Group 3). Panel B: Real (Group 4). Blue = L2 norm change; red = cosine similarity change. All eight models tested.\u003c/em\u003e\u003c/p\u003e","description":"","filename":"Fig7LengthEffectsGroup3Group4.png","url":"https://assets-eu.researchsquare.com/files/rs-9237602/v1/2de7738d2a9fd8995a8724a2.png"},{"id":105729293,"identity":"2e8ee59e-3fda-4865-9d25-4e7941ffdb08","added_by":"auto","created_at":"2026-03-30 11:14:16","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":19228543,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-9237602/v1/11159246-8af5-4300-b4f2-ab896ead12d2.pdf"},{"id":105699548,"identity":"e7476906-95c8-40ad-9e98-c1717a9cc26b","added_by":"auto","created_at":"2026-03-30 05:25:44","extension":"csv","order_by":1,"title":"","display":"","copyAsset":false,"role":"supplement","size":7698,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eS1 Table. Normality tests. \u003c/strong\u003eShapiro-Wilk test results for all model-condition combinations.\u003c/p\u003e","description":"","filename":"S1TableNormalityShapiroWilk.csv","url":"https://assets-eu.researchsquare.com/files/rs-9237602/v1/c5f47c19bcffda9621a70d40.csv"},{"id":105699553,"identity":"014d7796-dffb-435e-b6ed-2c33a55dd9bd","added_by":"auto","created_at":"2026-03-30 05:25:45","extension":"csv","order_by":2,"title":"","display":"","copyAsset":false,"role":"supplement","size":913,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eS2 Table. Variance homogeneity. \u003c/strong\u003eLevene’s test results for all models across complexity levels.\u003c/p\u003e","description":"","filename":"S2TableVarianceLevene.csv","url":"https://assets-eu.researchsquare.com/files/rs-9237602/v1/925d3a3c08b54056dfef2ff9.csv"},{"id":105699507,"identity":"ca7cec73-3cd3-44fc-9e9d-b59d3f7f5e4e","added_by":"auto","created_at":"2026-03-30 05:25:38","extension":"csv","order_by":3,"title":"","display":"","copyAsset":false,"role":"supplement","size":919,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eS3 Table. Non-parametric analysis. \u003c/strong\u003eKruskal-Wallis test results confirming parametric findings.\u003c/p\u003e","description":"","filename":"S3TableNonparametricKruskalWallis.csv","url":"https://assets-eu.researchsquare.com/files/rs-9237602/v1/9a531f5d44cb2f8d3be7d45b.csv"},{"id":105699587,"identity":"cc5143e2-9c23-4f0f-8120-810e031968d3","added_by":"auto","created_at":"2026-03-30 05:26:00","extension":"docx","order_by":4,"title":"","display":"","copyAsset":false,"role":"supplement","size":13903,"visible":true,"origin":"","legend":"","description":"","filename":"Supportinginformation.docx","url":"https://assets-eu.researchsquare.com/files/rs-9237602/v1/abd5a7d09ea17a8becc6fdc4.docx"}],"financialInterests":"The authors declare no competing interests.","formattedTitle":"\u003cp\u003e\u003cstrong\u003eRepresentation changes across varying clinical input conditions: A dual-metric validation study of eight transformer architectures with length controls\u003c/strong\u003e\u003c/p\u003e","fulltext":[{"header":"Introduction","content":"\u003cp\u003eLarge language models are increasingly deployed in clinical settings for tasks such as documentation, clinical decision support, and information retrieval(\u003cspan additionalcitationids=\"CR2 CR3\" citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e). The reliability of these systems depends in part on the stability of their internal representations\u0026mdash;the high-dimensional embeddings that encode semantic meaning. Understanding how these representations change under varying input conditions is relevant to characterising model behaviour in clinical applications(\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e, \u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e).\u003c/p\u003e\n\u003ch3\u003eTask-level performance degradation\u003c/h3\u003e\n\u003cp\u003ePerformance limitations in transformer language models have been documented at the task level. Liu et al. (\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e) demonstrated the \u0026lsquo;lost in the middle\u0026rsquo; phenomenon: language model performance is highest when relevant information occurs at the beginning or end of the input context, and degrades when it is in the middle of long contexts. This U-shaped performance curve persists even in explicitly long-context models. However, task-level performance metrics do not directly characterise the quality of internal representations. A model might show degraded task performance while maintaining stable representations\u0026mdash;or conversely, show representation changes that do not immediately manifest as task failures. This distinction motivates representation-level analysis as a complementary approach.\u003c/p\u003e \u003cdiv id=\"Sec3\" class=\"Section2\"\u003e \u003ch2\u003eGeometric properties of embedding spaces\u003c/h2\u003e \u003cp\u003eResearch has examined the geometric properties of transformer representations. Ethayarajh (\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e) found that contextualised word representations are anisotropic\u0026mdash;concentrated in a narrow cone rather than uniformly distributed. Gao et al. (\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e)identified this as the \u0026lsquo;representation degeneration problem.\u0026rsquo; Timkey and van Schijndel (\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e) discovered that \u0026lsquo;rogue dimensions\u0026rsquo; dominate similarity measures, with a mismatch between dimensions important for similarity and those important for model behaviour. Cai et al. (\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e)offered a more nuanced view, identifying isolated isotropic clusters within globally anisotropic spaces.\u003c/p\u003e \u003c/div\u003e\n\u003ch3\u003eLength-induced embedding changes\u003c/h3\u003e\n\u003cp\u003eZhou et al. (\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e) characterised \u0026lsquo;length collapse\u0026rsquo; in transformer-based embedding models: longer text embeddings cluster in a narrower space, producing distributional differences across embeddings of varying text lengths. They theoretically demonstrated that the self-attention mechanism acts as a low-pass filter, with longer sequences increasing the rate of attenuation. Clinical documentation has distinctive characteristics, including specialised terminology and a correlation between document length and clinical complexity. Whether length-associated changes result from sequence length itself or from the clinical information density that accompanies longer texts has not been systematically examined.\u003c/p\u003e\n\u003ch3\u003eDomain-specific clinical language models\u003c/h3\u003e\n\u003cp\u003eDomain-specific language models for healthcare include BioBERT(\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e), pre-trained on PubMed abstracts, and ClinicalBERT(\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e), pre-trained on MIMIC-III clinical notes. More recently, Llama-2 (\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e) and Mistral (\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e) represent modern transformer architectures with improved positional encoding schemes. Comparative evaluations have shown that domain-specific models outperform general models on clinical NLP benchmarks(\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e, \u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e), but whether domain-specific training or modern architectural improvements affect representation stability remains unexamined.\u003c/p\u003e\n\u003ch3\u003eStudy objectives\u003c/h3\u003e\n\u003cp\u003eThis study characterises complexity-associated representation changes in clinical language models using two complementary metrics\u0026mdash;embedding magnitude (L2 norm) and directional alignment (cosine similarity)\u0026mdash;to distinguish geometric reorganisation from simple scaling artefacts. We address five questions: (\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e) Do transformer models exhibit significant representation changes across varying clinical complexity? (\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e) Are these changes geometric (affecting both magnitude and direction) or simple scaling artefacts? (\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e) Do findings generalise across real and synthetic clinical text? (\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e) Are changes driven by complexity or sequence length? (\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e) Do modern architectures (2023) show different patterns from older models (2018\u0026ndash;2020)? We analyse 1,500 clinical texts (12,000 model\u0026ndash;text observations) across eight transformer models spanning 2018\u0026ndash;2023.\u003c/p\u003e"},{"header":"Materials and methods","content":"\u003cdiv id=\"Sec8\" class=\"Section2\"\u003e \u003ch2\u003eStudy design\u003c/h2\u003e \u003cp\u003eWe conducted an observational study of changes in representation across input conditions in eight transformer models, using four experimental groups. This study was not pre-registered; all analyses should be considered exploratory. All analyses were carried out in February 2026.\u003c/p\u003e \u003c/div\u003e\n\u003ch3\u003eExperimental groups\u003c/h3\u003e\n\u003cp\u003e \u003cb\u003eGroup 1: Real clinical text.\u003c/b\u003e We extracted 450 clinical notes from the MT Samples corpus, a publicly available collection of medical transcriptions. The texts were stratified by complexity into three levels (n\u0026thinsp;=\u0026thinsp;150 per level): Simple (mean 6.7\u0026thinsp;\u0026plusmn;\u0026thinsp;1.8 words), Moderate (mean 21.9\u0026thinsp;\u0026plusmn;\u0026thinsp;3.2 words), and Complex (mean 65.8\u0026thinsp;\u0026plusmn;\u0026thinsp;7.6 words). We acknowledge that this operationalisation conflates complexity with length; Groups 3\u0026ndash;4 address this confound.\u003c/p\u003e \u003cp\u003e \u003cb\u003eGroup 2: Synthetic matched text.\u003c/b\u003e 450 synthetic clinical texts, matched to Group 1 for complexity and length (n\u0026thinsp;=\u0026thinsp;150 per level), were constructed by the first author (a physician).\u003c/p\u003e \u003cp\u003e \u003cb\u003eGroup 3: Length-controlled synthetic text.\u003c/b\u003e 300 synthetic texts with matched lengths but identical simple content at two lengths (~\u0026thinsp;22 and ~\u0026thinsp;63 words; n\u0026thinsp;=\u0026thinsp;150 per condition), isolating pure length effects.\u003c/p\u003e \u003cp\u003e \u003cb\u003eGroup 4: Length-controlled real text.\u003c/b\u003e 300 texts from MT Samples, padded with normal clinical findings to match moderate and complex lengths (n\u0026thinsp;=\u0026thinsp;150 per condition).\u003c/p\u003e\n\u003ch3\u003eModels\u003c/h3\u003e\n\u003cp\u003eWe evaluated eight transformer models across four architectural families. General-purpose: GPT-2 Small (117M parameters, 2019(\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e); gpt2), GPT-2 Medium (345M, 2019; gpt2-medium), GPT-2 XL (1.5B, 2019; gpt2-xl), and BERT Base (110M, 2018; bert-base-uncased). Domain-specific: BioBERT v1.1 (110M, 2020; dmis-lab/biobert-v1.1) and ClinicalBERT (110M, 2019; emilyalsentzer/Bio_ClinicalBERT). Modern: Llama-2-7B (7B, 2023; meta-llama/Llama-2-7b-hf) and Mistral-7B (7B, 2023; mistralai/Mistral-7B-v0.1). All models were run at full precision (float32) using HuggingFace Transformers v4.36.0 with PyTorch v2.1.0 on an NVIDIA H100 GPU (80GB HBM3).\u003c/p\u003e \u003cdiv id=\"Sec11\" class=\"Section2\"\u003e \u003ch2\u003eMetrics\u003c/h2\u003e \u003cp\u003eWe assessed two complementary representation metrics from the final hidden layer, excluding special tokens: (\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e) embedding magnitude (mean per-token L2 norm) to capture scale, and (\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e) directional alignment (mean pairwise cosine similarity between token embeddings) to capture geometry. If inputs induce only scaling changes, magnitude will change while cosine similarity remains stable; geometric reorganisation is indicated by changes in both metrics. For interpretation, we classified simple\u0026rarr;complex effects per model as geometric, scaling-only, directional-only, or none, based on Bonferroni-corrected significance for each metric.\u003c/p\u003e \u003cp\u003eLet token embeddings from the final hidden layer be hₜ \u0026isin; ℝᵈ for tokens t\u0026thinsp;=\u0026thinsp;1,\u0026hellip;,T. The embedding magnitude for a text was M = (1/T)\u0026sum;ₜ‖hₜ‖₂. Directional alignment was C = (2/(T(T\u0026thinsp;\u0026minus;\u0026thinsp;1)))\u0026sum;_{i\u0026thinsp;\u0026lt;\u0026thinsp;j} (h\u003csub\u003ei\u003c/sub\u003e\u0026middot;hⱼ)/(‖h\u003csub\u003ei\u003c/sub\u003e‖₂‖hⱼ‖₂).\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec12\" class=\"Section2\"\u003e \u003ch2\u003eStatistical analysis\u003c/h2\u003e \u003cp\u003eFor each model, a one-way ANOVA tested whether mean embedding magnitude differed across three complexity levels in Group 1. A Bonferroni correction was applied to eight comparisons (α\u0026thinsp;=\u0026thinsp;0.05/8\u0026thinsp;=\u0026thinsp;0.00625). Post hoc comparisons used Welch\u0026rsquo;s t-test (unequal variances confirmed by Levene\u0026rsquo;s test for four of eight models). Effect sizes: Cohen\u0026rsquo;s d and eta-squared (η\u0026sup2;). We note that Cohen\u0026rsquo;s d values for domain-specific models are very large (d\u0026thinsp;\u0026gt;\u0026thinsp;3.5) because within-group standard deviations are small; this reflects the deterministic nature of the computation rather than an unusually strong behavioural effect. Bootstrap 95% CIs (10,000 iterations, bias-corrected percentile method). Non-parametric Kruskal-Wallis tests confirmed the parametric findings. Cross-corpus validation: Pearson correlation on three data points per model (interpreted as directional indicators). Cosine similarity changes were computed from mean values across conditions. All analyses: Python 3.11, SciPy v1.11, NumPy v1.25. Sample sizes (n\u0026thinsp;=\u0026thinsp;150 per condition) were determined by corpus construction.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec13\" class=\"Section2\"\u003e \u003ch2\u003eEthics statement\u003c/h2\u003e \u003cp\u003eThis study used only publicly available text data (the MT Samples medical transcription corpus) and pre-trained language models. No human participants were recruited, no biological samples were collected, and no interventions were performed. The MT Samples corpus comprises de-identified medical transcription samples; no personal health identifiers are present in the texts used. The texts were used in accordance with the corpus\u0026rsquo;s publicly available terms. Institutional review board approval was not required.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec14\" class=\"Section2\"\u003e \u003ch2\u003eData and code availability\u003c/h2\u003e \u003cp\u003eAll data and analysis code are publicly available at: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://github.com/yngvemikkelsen/clinical-representation-stability\u003c/span\u003e\u003cspan address=\"https://github.com/yngvemikkelsen/clinical-representation-stability\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. The repository contains: all derived datasets in CSV format (per-text-embedding magnitudes and cosine-similarity values for all eight models and ten conditions); Python analysis scripts that reproduce all tables, statistical tests, and figures; and model configuration files with the exact software versions. Raw MT Samples texts are not redistributed; instead, we provide text identifiers and the exact selection procedure to enable reproduction from the original source. Raw data are also provided as Supporting Information (S1\u0026ndash;S2 Datasets).\u003c/p\u003e \u003c/div\u003e"},{"header":"Results","content":"\u003cdiv id=\"Sec16\" class=\"Section2\"\u003e \u003ch2\u003ePrimary finding: Magnitude changes across complexity levels\u003c/h2\u003e \u003cp\u003eSeven of eight models showed statistically significant magnitude changes across complexity levels in real clinical text (Group 1; ANOVA p\u0026thinsp;\u0026lt;\u0026thinsp;0.001, surviving Bonferroni correction at α\u0026thinsp;=\u0026thinsp;0.00625). Mistral-7B showed a small but significant magnitude increase (+\u0026thinsp;1.9% [95% CI: +1.1%, +\u0026thinsp;2.8%], d\u0026thinsp;=\u0026thinsp;0.54, p\u0026thinsp;\u0026lt;\u0026thinsp;0.001). Llama-2-7B was the only model with a non-significant magnitude change (ANOVA: F(2,447)\u0026thinsp;=\u0026thinsp;2.22, p\u0026thinsp;=\u0026thinsp;0.109, η\u0026sup2; = 0.010; Welch\u0026rsquo;s t-test, simple vs complex: t\u0026thinsp;=\u0026thinsp;\u0026minus;\u0026thinsp;1.87, p\u0026thinsp;=\u0026thinsp;0.062). Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e presents the full results for both metrics; Fig.\u0026nbsp;1 shows the magnitude distributions.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab1\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eRepresentation magnitude changes across complexity levels in real clinical text (Group 1)\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"10\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c7\" colnum=\"7\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c8\" colnum=\"8\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c9\" colnum=\"9\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c10\" colnum=\"10\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eModel\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eSimple (SD)\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eComplex (SD)\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eΔ%\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003e95% CI\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c6\"\u003e \u003cp\u003ep\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c7\"\u003e \u003cp\u003ed\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c8\"\u003e \u003cp\u003eη\u0026sup2;\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c9\"\u003e \u003cp\u003eCos Δ%\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c10\"\u003e \u003cp\u003eType\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eGPT-2 Small\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e182.0 (14.4)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e189.7 (11.9)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e+\u0026thinsp;4.2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e[+\u0026thinsp;2.6, +\u0026thinsp;5.9]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e\u0026lt;\u0026thinsp;.001\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.58\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003e.078\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c9\"\u003e \u003cp\u003e\u0026minus;3.5\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c10\"\u003e \u003cp\u003eG\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eGPT-2 Med\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e225.9 (29.2)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e207.9 (18.8)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e\u0026minus;8.0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e[\u0026minus;\u0026thinsp;10.2, \u0026minus;\u0026thinsp;5.6]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e\u0026lt;\u0026thinsp;.001\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e\u0026minus;0.73\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003e.113\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c9\"\u003e \u003cp\u003e\u0026minus;17.2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c10\"\u003e \u003cp\u003eG\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eGPT-2 XL\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e31.7 (2.8)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e33.0 (1.6)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e+\u0026thinsp;3.9\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e[+\u0026thinsp;2.3, +\u0026thinsp;5.6]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e\u0026lt;\u0026thinsp;.001\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.55\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003e.055\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c9\"\u003e \u003cp\u003e\u0026minus;1.1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c10\"\u003e \u003cp\u003eS\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eBERT Base\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e15.1 (0.8)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e15.7 (0.4)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e+\u0026thinsp;3.8\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e[+\u0026thinsp;2.9, +\u0026thinsp;4.8]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e\u0026lt;\u0026thinsp;.001\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.91\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003e.130\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c9\"\u003e \u003cp\u003e\u0026minus;11.1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c10\"\u003e \u003cp\u003eG\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eBioBERT\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e13.9 (0.3)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e15.0 (0.3)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e+\u0026thinsp;8.5\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e[+\u0026thinsp;8.0, +\u0026thinsp;9.0]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e\u0026lt;\u0026thinsp;.001\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e3.85\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003e.697\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c9\"\u003e \u003cp\u003e\u0026minus;25.9\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c10\"\u003e \u003cp\u003eG\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eClinicalBERT\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e14.9 (0.3)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e16.0 (0.3)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e+\u0026thinsp;7.3\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e[+\u0026thinsp;6.8, +\u0026thinsp;7.8]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e\u0026lt;\u0026thinsp;.001\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e3.57\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003e.650\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c9\"\u003e \u003cp\u003e\u0026minus;25.0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c10\"\u003e \u003cp\u003eG\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eLlama-2-7B\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e111.0 (4.0)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e110.4 (1.8)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e\u0026minus;0.6\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e[\u0026minus;\u0026thinsp;1.2, +\u0026thinsp;0.0]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e.062\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e\u0026minus;0.22\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003e.010\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c9\"\u003e \u003cp\u003e+\u0026thinsp;3.6\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c10\"\u003e \u003cp\u003eN\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eMistral-7B\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e340.6 (15.2)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e347.2 (8.5)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e+\u0026thinsp;1.9\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e[+\u0026thinsp;1.1, +\u0026thinsp;2.8]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e\u0026lt;\u0026thinsp;.001\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.54\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003e.060\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c9\"\u003e \u003cp\u003e+\u0026thinsp;14.4\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c10\"\u003e \u003cp\u003eG\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003ep\u0026thinsp;=\u0026thinsp;Welch\u0026rsquo;s t-test (simple vs complex) for magnitude. η\u0026sup2; = eta-squared from a one-way ANOVA across all three strata (simple/moderate/complex). Cos Δ% = percentage change in mean pairwise cosine similarity (simple\u0026rarr;complex); cosine p-values are reported in Table\u0026nbsp;\u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e. Type: G\u0026thinsp;=\u0026thinsp;geometric (both magnitude and cosine differences significant), S\u0026thinsp;=\u0026thinsp;scaling-only (magnitude significant, cosine not), D\u0026thinsp;=\u0026thinsp;directional-only (cosine significant, magnitude not), N\u0026thinsp;=\u0026thinsp;no significant simple\u0026rarr;complex difference in either metric. Significance threshold: Bonferroni-corrected α\u0026thinsp;=\u0026thinsp;0.00625 (0.05/8 models). Very large d values for BioBERT/ClinicalBERT reflect low within-group variance.\u003c/p\u003e \u003cp\u003e \u003cb\u003eFigure\u0026nbsp;1. Embedding magnitude distributions across complexity levels (Group 1: real clinical text).\u003c/b\u003e \u003cem\u003eViolin plots with individual observations (jittered). Black lines indicate the means. n\u0026thinsp;=\u0026thinsp;150 per complexity level per model. *** p\u0026thinsp;\u0026lt;\u0026thinsp;0.001 (Bonferroni-corrected).\u003c/em\u003e\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec17\" class=\"Section2\"\u003e \u003ch2\u003eDual-metric analysis: Geometric vs scaling effects\u003c/h2\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab2\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 2\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eCosine similarity changes across complexity levels in real clinical text (Group 1).\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"9\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c7\" colnum=\"7\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c8\" colnum=\"8\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c9\" colnum=\"9\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eModel\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eSimple\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eModerate\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eComplex\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eΔ%\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c6\"\u003e \u003cp\u003e95% CI\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c7\"\u003e \u003cp\u003ep\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c8\"\u003e \u003cp\u003ed\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c9\"\u003e \u003cp\u003eKW p\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eGPT-2 Small\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.957\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.946\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.924\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e\u0026minus;3.5\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e[\u0026minus;\u0026thinsp;5.3, \u0026minus;\u0026thinsp;1.7]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e\u0026lt;\u0026thinsp;.001\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e\u0026minus;0.44\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c9\"\u003e \u003cp\u003e\u0026lt;\u0026thinsp;.001\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eGPT-2 Med\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.941\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.843\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.780\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e\u0026minus;17.2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e[\u0026minus;\u0026thinsp;20.0, \u0026minus;\u0026thinsp;14.3]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e\u0026lt;\u0026thinsp;.001\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e\u0026minus;1.26\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c9\"\u003e \u003cp\u003e\u0026lt;\u0026thinsp;.001\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eGPT-2 XL\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.360\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.339\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.356\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e\u0026minus;1.1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e[\u0026minus;\u0026thinsp;4.8, +\u0026thinsp;2.8]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e.582\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e\u0026minus;0.06\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c9\"\u003e \u003cp\u003e.001\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eBERT Base\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.372\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.357\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.331\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e\u0026minus;11.1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e[\u0026minus;\u0026thinsp;13.5, \u0026minus;\u0026thinsp;8.8]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e\u0026lt;\u0026thinsp;.001\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e\u0026minus;1.00\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c9\"\u003e \u003cp\u003e\u0026lt;\u0026thinsp;.001\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eBioBERT\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.640\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.562\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.474\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e\u0026minus;25.9\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e[\u0026minus;\u0026thinsp;27.2, \u0026minus;\u0026thinsp;24.6]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e\u0026lt;\u0026thinsp;.001\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e\u0026minus;4.02\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c9\"\u003e \u003cp\u003e\u0026lt;\u0026thinsp;.001\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eClinicalBERT\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.637\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.552\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.477\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e\u0026minus;25.0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e[\u0026minus;\u0026thinsp;26.0, \u0026minus;\u0026thinsp;24.0]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e\u0026lt;\u0026thinsp;.001\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e\u0026minus;4.85\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c9\"\u003e \u003cp\u003e\u0026lt;\u0026thinsp;.001\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eLlama-2-7B\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.316\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.313\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.328\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e+\u0026thinsp;3.6\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e[\u0026minus;\u0026thinsp;0.9, +\u0026thinsp;8.6]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e.126\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e+\u0026thinsp;0.18\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c9\"\u003e \u003cp\u003e.011\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eMistral-7B\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.301\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.321\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.344\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e+\u0026thinsp;14.4\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e[+\u0026thinsp;9.7, +\u0026thinsp;19.4]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e\u0026lt;\u0026thinsp;.001\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e+\u0026thinsp;0.75\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c9\"\u003e \u003cp\u003e\u0026lt;\u0026thinsp;.001\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003e \u003cem\u003eMean pairwise cosine similarity between token embeddings from final hidden layer. Δ% = (complex\u0026ndash;simple)/simple \u0026times; 100. p\u0026thinsp;=\u0026thinsp;Welch\u0026rsquo;s t-test (simple vs complex). d\u0026thinsp;=\u0026thinsp;Cohen\u0026rsquo;s d. KW p\u0026thinsp;=\u0026thinsp;Kruskal-Wallis non-parametric test across three levels. 95% CI\u0026thinsp;=\u0026thinsp;bootstrap (10,000 iterations). Bonferroni-corrected α\u0026thinsp;=\u0026thinsp;0.00625. n\u0026thinsp;=\u0026thinsp;150 per level.\u003c/em\u003e \u003c/p\u003e \u003cp\u003eThe dual-metric analysis (Fig.\u0026nbsp;2, Table\u0026nbsp;\u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e) indicates that representation changes across the simple\u0026rarr;complex contrast are predominantly geometric\u0026mdash;encompassing both magnitude and directional shifts\u0026mdash;rather than simple scaling artefacts. Five of six older models with significant magnitude changes also showed significant directional changes (cosine similarity decreased by 3\u0026ndash;26%). The exception was GPT-2 XL, which showed a significant magnitude increase (+\u0026thinsp;3.9%) but no significant simple\u0026rarr;complex cosine change (\u0026minus;\u0026thinsp;1.1%, p = .582), consistent with a scaling-dominant effect. Domain-specific models showed the largest directional reorganisation: BioBERT\u0026rsquo;s cosine similarity decreased by 25.9% (0.64\u0026rarr;0.47, d\u0026thinsp;=\u0026thinsp;\u0026minus;\u0026thinsp;4.02) and ClinicalBERT by 25.0% (0.64\u0026rarr;0.48, d\u0026thinsp;=\u0026thinsp;\u0026minus;\u0026thinsp;4.85). Among modern architectures, Mistral-7B showed significant directional convergence (cosine\u0026thinsp;+\u0026thinsp;14.4%, p \u0026lt; .001), whereas Llama-2-7B showed a small, non-significant increase in cosine similarity (+\u0026thinsp;3.6%, p = .126).\u003c/p\u003e \u003cp\u003e \u003cb\u003eFigure\u0026nbsp;2. Dual-metric analysis: Embedding magnitude and directional changes across complexity levels.\u003c/b\u003e \u003cem\u003ePanel A: Side-by-side comparison of changes in the L2 norm (blue) and cosine similarity (red) across all eight models. Panel B: Scatter plot showing model-specific patterns. Older models (red circles) cluster in the lower-right quadrant (magnitude increase, directional divergence); modern models (blue squares) cluster near the origin or in the upper region (stability or convergence).\u003c/em\u003e\u003c/p\u003e \u003cp\u003e \u003cb\u003eFigure\u0026nbsp;3. Mean pairwise cosine similarity across complexity levels (Group 1: real clinical text, 8 models).\u003c/b\u003e Higher cosine similarity indicates more directionally aligned token embeddings. All six older models show decreasing cosine similarity with complexity (tokens become more directionally diverse). Modern models show increasing cosine similarity (directional convergence), with statistical significance for Mistral-7B and non-significance for Llama-2-7B in the simple\u0026rarr;complex contrast.\u003c/p\u003e \u003cp\u003e \u003cb\u003eFigure\u0026nbsp;4. Effect sizes with 95% bootstrap confidence intervals (Group 1: real clinical text).\u003c/b\u003e \u003cem\u003eForest plot of L2-norm percentage changes. Bootstrap CIs from 10,000 iterations. Red\u0026thinsp;=\u0026thinsp;p \u0026lt; 0.001; grey\u0026thinsp;=\u0026thinsp;not significant. Bonferroni-corrected α\u0026thinsp;=\u0026thinsp;0.00625.\u003c/em\u003e\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec18\" class=\"Section2\"\u003e \u003ch2\u003eWithin-group variability\u003c/h2\u003e \u003cp\u003eWithin-group magnitude variability decreased with complexity across all eight models. Variance ratios (complex/simple) ranged from 0.19 (Llama-2-7B) to 0.76 (BioBERT). The coefficient of variation decreased consistently: e.g., GPT-2 Medium 12.9%\u0026rarr;9.1%, BERT Base 5.4%\u0026rarr;2.4%, Llama-2-7B 3.6%\u0026rarr;1.6% (Fig.\u0026nbsp;5). This convergence aligns with Zhou et al.\u0026rsquo;s (\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e) theoretical prediction that longer sequences produce more homogeneous representations.\u003c/p\u003e \u003cp\u003e \u003cb\u003eFigure\u0026nbsp;5. Within-group magnitude variability across complexity levels (Group 1: Real clinical text).\u003c/b\u003e \u003cem\u003eCV\u0026thinsp;=\u0026thinsp;SD/mean \u0026times; 100. All models show decreasing CV from simple to complex, indicating representational convergence.\u003c/em\u003e\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec19\" class=\"Section2\"\u003e \u003ch2\u003eCross-corpus validation\u003c/h2\u003e \u003cp\u003eFive of the eight models showed high cross-corpus magnitude correlation (r\u0026thinsp;\u0026gt;\u0026thinsp;0.80): GPT-2 Small (0.98), GPT-2 XL (1.00), BioBERT (0.98), ClinicalBERT (0.89), and Llama-2-7B (0.82). GPT-2 Medium showed a weak correlation (r\u0026thinsp;=\u0026thinsp;0.34). BERT Base and Mistral-7B showed a reversal (r\u0026thinsp;=\u0026thinsp;\u0026minus;\u0026thinsp;0.93 and r\u0026thinsp;=\u0026thinsp;\u0026minus;\u0026thinsp;0.93). These correlations are computed from three data points and should be interpreted as directional indicators (Fig.\u0026nbsp;6).\u003c/p\u003e \u003cp\u003e \u003cb\u003eFigure\u0026nbsp;6. Cross-corpus validation: Real vs synthetic mean embedding magnitude (8 models).\u003c/b\u003e \u003cem\u003eEach point represents one complexity level. Dashed lines show the linear fit. Five of eight models show consistent cross-corpus patterns.\u003c/em\u003e\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec20\" class=\"Section2\"\u003e \u003ch2\u003eModern architecture performance\u003c/h2\u003e \u003cp\u003eBoth modern architectures (2023) showed patterns that differed from those of older models. Llama-2-7B showed no significant change in magnitude (\u0026minus;\u0026thinsp;0.6%, p\u0026thinsp;=\u0026thinsp;0.062) and no significant change in simple\u0026rarr;complex cosine (+\u0026thinsp;3.6%, p\u0026thinsp;=\u0026thinsp;0.126). Mistral-7B showed a small but significant magnitude increase (+\u0026thinsp;1.9%, p\u0026thinsp;\u0026lt;\u0026thinsp;0.001, d\u0026thinsp;=\u0026thinsp;0.54) and significant directional convergence (cosine\u0026thinsp;+\u0026thinsp;14.4%, p\u0026thinsp;\u0026lt;\u0026thinsp;0.001, d\u0026thinsp;=\u0026thinsp;+\u0026thinsp;0.75). Overall, the modern models showed smaller magnitude shifts than older models and a convergence trend in cosine similarity, contrasting with the directional divergence observed in older models.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec21\" class=\"Section2\"\u003e \u003ch2\u003eLength-controlled analysis\u003c/h2\u003e \u003cp\u003eLength-controlled analyses (Fig.\u0026nbsp;7) confirmed that sequence length alone produces changes in both metrics. In Group 3 (synthetic, constant simple content), all eight models showed magnitude changes with length, and seven of eight showed cosine similarity changes exceeding 3%. Cosine similarity changes with length (\u0026minus;\u0026thinsp;18.6% to +\u0026thinsp;13.3%) were often larger than those with complexity, confirming that length is a major driver of representational geometry changes. The direction of length effects varied by architecture: older models generally showed decreasing cosine similarity with length (directional divergence), while modern models showed mixed patterns.\u003c/p\u003e \u003cp\u003e \u003cb\u003eFigure\u0026nbsp;7. Length effects on both magnitude and directional metrics (complexity held constant).\u003c/b\u003e \u003cem\u003ePanel A: Synthetic (Group 3). Panel B: Real (Group 4). Blue\u0026thinsp;=\u0026thinsp;L2 norm change; red\u0026thinsp;=\u0026thinsp;cosine similarity change. All eight models tested.\u003c/em\u003e\u003c/p\u003e \u003c/div\u003e"},{"header":"Discussion","content":"\u003cdiv id=\"Sec23\" class=\"Section2\"\u003e \u003ch2\u003ePrincipal findings\u003c/h2\u003e \u003cp\u003eThis study characterised changes in representation across eight transformer models using dual metrics (magnitude and directional alignment), four experimental groups, 1,500 clinical texts, and 12,000 model\u0026ndash;text observations. Five principal findings emerged. First, seven of eight models showed significant changes in magnitude with clinical complexity, with effect sizes ranging from medium to very large. Second, the dual-metric analysis showed that these changes are predominantly geometric\u0026mdash;six of seven affected models exhibited both magnitude and directional shifts. Third, domain-specific models (BioBERT, ClinicalBERT) showed the largest effects on both metrics, with cosine similarity decreasing by approximately 25%. Fourth, modern architectures exhibited qualitatively different directional patterns: convergence rather than divergence, with Llama-2-7B showing a non-significant change in magnitude and Mistral-7B showing a small but significant effect. Fifth, length-controlled analyses confirmed that both length and complexity contribute to changes in representation across both metrics.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec24\" class=\"Section2\"\u003e \u003ch2\u003eThe nature of representation changes: Geometric reorganisation\u003c/h2\u003e \u003cp\u003eA key contribution of this study is the establishment that complexity-associated changes reflect genuine geometric reorganisation of the embedding space rather than simple scaling artefacts. If embedding magnitudes grew proportionally with complexity, cosine similarity would remain constant. Instead, we observed dramatic directional changes\u0026mdash;BioBERT and ClinicalBERT\u0026rsquo;s cosine similarity dropped by approximately 25%, indicating that token embeddings become substantially more directionally diverse when encoding complex clinical text. This has implications for downstream applications: directional changes could affect nearest-neighbour retrieval, clustering, and any application that depends on cosine similarity-based comparisons.\u003c/p\u003e \u003cdiv id=\"Sec25\" class=\"Section3\"\u003e \u003ch2\u003eComparison with prior work\u003c/h2\u003e \u003cp\u003eZhou et al. (\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e) theoretically demonstrated that self-attention acts as a low-pass filter, producing representational convergence with length. Our findings partially support and extend this framework. The magnitude convergence (decreasing within-group variance) aligns with the low-pass filter prediction. However, the directional divergence we observe (decreasing cosine similarity with complexity) suggests a more nuanced picture: while magnitude variance decreases (convergence), directional diversity increases (divergence). This suggests that complex clinical text activates more diverse representational pathways even as magnitude variance narrows.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec26\" class=\"Section3\"\u003e \u003ch2\u003eScope and nature of claims\u003c/h2\u003e \u003cp\u003eIt is important to distinguish characterisation studies from applied evaluation studies. This work characterises what happens to representations\u0026mdash;it documents geometric reorganisation as a function of input complexity. It does not claim that these changes cause downstream task failures. Analogously, Ethayarajh (\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e) characterised anisotropy in contextualised embeddings, Gao et al. (\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e) characterised representation degeneration, and Zhou et al. (\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e) characterised length collapse\u0026mdash;none included downstream task evaluation, because the research question was \u0026lsquo;what happens\u0026rsquo; rather than \u0026lsquo;does it matter\u0026rsquo;. Whether geometric changes in representations translate into differences in clinical task performance is a distinct research question that warrants its own rigorous investigation (see Future Directions), and combining both questions in a single study would risk doing neither justice.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec27\" class=\"Section3\"\u003e \u003ch2\u003eClinical implications\u003c/h2\u003e \u003cp\u003eWith this scope in mind, three tentative implications emerge. First, systems using older models (2018\u0026ndash;2020) should account for the fact that internal representations change geometrically\u0026mdash;in magnitude and direction\u0026mdash;with input complexity. These changes could affect cosine-similarity-based retrieval, nearest-neighbour classification, or embedding clustering in clinical pipelines, though empirical verification is needed. Second, domain-specific models showed the largest changes on both metrics, which is noteworthy given their widespread use in clinical NLP. This does not necessarily indicate inferior task performance; larger geometric changes could reflect more differentiated encoding of clinical complexity. Third, both modern architectures show greater stability across both metrics, supporting their consideration for applications requiring consistent representational behaviour.\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv id=\"Sec28\" class=\"Section2\"\u003e \u003ch2\u003eLimitations\u003c/h2\u003e \u003cp\u003eSeveral limitations merit consideration. First, although we employed two complementary metrics (L2 norm and cosine similarity), additional measures such as centred kernel alignment (CKA) and representational similarity analysis (RSA) would provide further characterisation. Second, we did not assess downstream task performance. As discussed above, this reflects the study\u0026rsquo;s scope as a characterisation study rather than a methodological gap; linking geometric changes to task performance is the highest-priority follow-up study. Third, our operationalisation of complexity is confounded with length in Groups 1\u0026ndash;2; Groups 3\u0026ndash;4 control for the length direction but not the reverse. Fourth, sample sizes were determined by corpus construction rather than by a priori power analysis. Fifth, we analysed only final-layer representations. Sixth, synthetic texts were generated by a single author. Seventh, MT Samples contains template transcriptions, not authentic EHR data. Eighth, cross-corpus correlations are based on three data points. Ninth, this study was not pre-registered, and all analyses are exploratory.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec29\" class=\"Section2\"\u003e \u003ch2\u003eFuture directions\u003c/h2\u003e \u003cp\u003eFuture work should prioritise linking changes in dual-metric representations to downstream clinical task performance using standardised benchmarks. Layer-wise analysis could reveal where geometric reorganisation occurs. Additional metrics (CKA, RSA) would further characterise the nature of directional changes. Testing additional modern architectures would establish generalisability. Finally, the divergent directional patterns between older and modern architectures warrant mechanistic investigation.\u003c/p\u003e \u003c/div\u003e"},{"header":"Conclusions","content":"\u003cp\u003eRepresentation changes across length-proxied clinical complexity strata are often geometric, involving both magnitude shifts and directional reorganisation of the embedding space. Domain-specific models show the largest dual-metric effects (BioBERT magnitude +8.5%, cosine similarity −25.9%; ClinicalBERT magnitude +7.3%, cosine similarity −25.0%). Modern architectures show comparatively small magnitude shifts (Llama-2-7B: non-significant; Mistral-7B: small but significant) and a convergence trend in cosine similarity, in contrast to the directional divergence observed in older models. Both length and complexity contribute to representation changes. Whether these geometric changes translate into differences in downstream clinical task performance remains an important open question.\u003c/p\u003e\n"},{"header":"Declarations","content":"\u003cp\u003eData availability statement\u003c/p\u003e\n\u003cp\u003eAll data underpinning the findings are available without restriction. Individual-level measurements for all 12,000 observations (S2 Dataset) and summary statistics, including cosine similarity (S1 Dataset), are provided as Supporting Information. All data and analysis code are available at: https://github.com/yngvemikkelsen/clinical-representation-stability.\u003c/p\u003e\n\u003cp\u003eAuthor contributions\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eConceptualisation:\u0026nbsp;\u003c/strong\u003eYM. \u003cstrong\u003eMethodology:\u0026nbsp;\u003c/strong\u003eYM. \u003cstrong\u003eSoftware:\u0026nbsp;\u003c/strong\u003eYM. \u003cstrong\u003eFormal analysis:\u0026nbsp;\u003c/strong\u003eYM. \u003cstrong\u003eInvestigation:\u0026nbsp;\u003c/strong\u003eYM. \u003cstrong\u003eData curation:\u0026nbsp;\u003c/strong\u003eYM. \u003cstrong\u003eWriting – original draft:\u0026nbsp;\u003c/strong\u003eYM. \u003cstrong\u003eWriting – review \u0026amp; editing:\u0026nbsp;\u003c/strong\u003eYM. \u003cstrong\u003eVisualisation:\u0026nbsp;\u003c/strong\u003eYM.\u003c/p\u003e\n\u003cp\u003eAcknowledgements\u003c/p\u003e\n\u003cp\u003eThe author thanks the MT Samples community for making clinical transcription samples publicly available.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\n \u003cli\u003eLee P, Bubeck S, Petro J. Benefits, limits, and risks of GPT-4 as an AI chatbot for medicine. N Engl J Med. 2023;388(13):1233\u0026ndash;9.\u003c/li\u003e\n \u003cli\u003eSinghal K, Azizi S, Tu T, Mahdavi SS, Wei J, Chung HW. Large language models encode clinical knowledge. Nature. 2023;620(7972):172\u0026ndash;80.\u003c/li\u003e\n \u003cli\u003eThirunavukarasu AJ, Ting DSJ, Elangovan K, Gutierrez L, Tan TF, Ting DSW. Large language models in medicine. Nat Med. 2023;29(8):1930\u0026ndash;40.\u003c/li\u003e\n \u003cli\u003eMoor M, Banerjee O, Abad ZSH. Foundation models for generalist medical artificial intelligence. Nature. 2023;616:259\u0026ndash;65.\u003c/li\u003e\n \u003cli\u003eRogers A, Kovaleva O, Rumshisky A. A primer in BERTology: What we know about how BERT works. Trans Assoc Comput Linguist. 2020;8:842\u0026ndash;66.\u003c/li\u003e\n \u003cli\u003eBelinkov Y, Glass J. Analysis methods in neural language processing: A survey. Trans Assoc Comput Linguist. 2019;7:49\u0026ndash;72.\u003c/li\u003e\n \u003cli\u003eLiu NF, Lin K, Hewitt J. Lost in the middle: How language models use long contexts. Trans Assoc Comput Linguist. 2024;12:157\u0026ndash;73.\u003c/li\u003e\n \u003cli\u003eEthayarajh K, editor How contextual are contextualized word representations? Proceedings of EMNLP-IJCNLP; 2019.\u003c/li\u003e\n \u003cli\u003eGao J, He D, Tan X, Qin T, Wang L, Liu TY, editors. Representation degeneration problem in training natural language generation models. Proceedings of ICLR; 2019.\u003c/li\u003e\n \u003cli\u003eTimkey W, van Schijndel M, editors. All bark and no bite: Rogue dimensions in transformer language models. Proceedings of EMNLP; 2021.\u003c/li\u003e\n \u003cli\u003eCai X, Huang J, Bian Y, Church K, editors. Isotropy in the contextual embedding space: Clusters and manifolds. Proceedings of ICLR; 2021.\u003c/li\u003e\n \u003cli\u003eZhou Y, Wu J, Cui G, Zhang S, Zhang Z, Liu Z, editors. Length-induced embedding collapse in transformer-based models. Proceedings of ICLR; 2025.\u003c/li\u003e\n \u003cli\u003eLee J, Yoon W, Kim S. BioBERT: A pre-trained biomedical language representation model. Bioinformatics. 2020;36(4):1234\u0026ndash;40.\u003c/li\u003e\n \u003cli\u003eAlsentzer E, Murphy J, Boag W, editors. Publicly available clinical BERT embeddings. Proceedings of the 2nd Clinical NLP Workshop; 2019.\u003c/li\u003e\n \u003cli\u003eTouvron H, Martin L, Stone K. Llama 2: Open foundation and fine-tuned chat models. 2023.\u003c/li\u003e\n \u003cli\u003eJiang AQ, Sablayrolles A, Mensch A. Mistral 7B. 2023.\u003c/li\u003e\n \u003cli\u003eRasmy L, Xiang Y, Xie Z, Tao C, Zhi D. Med-BERT: Pretrained contextualized embeddings on large-scale structured electronic health records. NPJ Digit Med. 2021;4(86).\u003c/li\u003e\n \u003cli\u003eGu Y, Tinn R, Cheng H. Domain-specific language model pretraining for biomedical NLP. ACM Trans Comput Healthcare. 2022;3(1):1\u0026ndash;23.\u003c/li\u003e\n \u003cli\u003eMT Samples: Medical transcription samples.\u003c/li\u003e\n\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":true,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"","lastPublishedDoi":"10.21203/rs.3.rs-9237602/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-9237602/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eBackground: Large language models are increasingly deployed in clinical decision support, yet the stability of their internal representations across diverse clinical input conditions remains poorly characterised. It is unclear whether changes in representation reflect geometric reorganisation (magnitude and directional shifts) or simple scaling artefacts.\u003c/p\u003e\n\u003cp\u003eMethods: We used a four-group validation design across eight transformer models (2018–2023), including two modern architectures (Llama-2-7B, Mistral-7B). Group 1 comprised 450 MT samples of clinical notes, stratified into simple/moderate/complex length-proxied strata (n = 150 each). Group 2 comprised 450 matched synthetic texts. Groups 3–4 comprised 600 length-controlled texts isolating pure length effects. From the final hidden layer (excluding special tokens), we computed (i) per-token embedding magnitude (mean L2 norm) and (ii) mean pairwise cosine similarity between token embeddings. Analyses used one-way ANOVA with Bonferroni correction across eight models (α = 0.00625), Welch t-tests, and bootstrap 95% confidence intervals (10,000 iterations).\u003c/p\u003e\n\u003cp\u003eResults: In Group 1, seven of eight models showed significant differences in magnitude across strata (p \u0026lt; 0.00625). Six of these seven also showed significant directional changes (cosine similarity changes of 3–26%), indicating geometric changes rather than scaling alone. BioBERT and ClinicalBERT showed the largest dual-metric effects (magnitude +8.5% and +7.3%; cosine −25.9% and −25.0%). Llama-2-7B showed no significant magnitude change (−0.6%, p = 0.062) and a non-significant simple-to-complex cosine change (+3.6%, p = 0.126). Mistral-7B showed a small but significant magnitude increase (+1.9%, p \u0026lt; 0.001) and significant directional convergence (cosine +14.4%, p \u0026lt; 0.001). Length-controlled analyses confirmed substantial length effects on both metrics.\u003c/p\u003e\n\u003cp\u003eConclusions: In older models, representation changes across length-proxied strata of clinical complexity are predominantly geometric. Modern architectures exhibit smaller magnitude shifts and a convergence trend in cosine similarity, in contrast to directional divergence in older models. Whether these representation-level changes translate into differences in downstream clinical task performance remains to be established.\u003c/p\u003e","manuscriptTitle":"Representation changes across varying clinical input conditions: A dual-metric validation study of eight transformer architectures with length controls","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2026-03-30 05:25:19","doi":"10.21203/rs.3.rs-9237602/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"4f7a5dae-8cbc-4f74-a8bc-c19183f27cbd","owner":[],"postedDate":"March 30th, 2026","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[{"id":65216710,"name":"Integrative \u0026 Complementary Medicine"}],"tags":[],"updatedAt":"2026-03-30T05:25:19+00:00","versionOfRecord":[],"versionCreatedAt":"2026-03-30 05:25:19","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-9237602","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-9237602","identity":"rs-9237602","version":["v1"]},"buildId":"XKTyCvWXoU3ODBz1xrDgd","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.