Reproducibility of AVM grading in clinical practice: A study of interobserver variability | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Reproducibility of AVM grading in clinical practice: A study of interobserver variability Francesco M.C. Lioi, Alessandro Benedictis, Davide Ferlito, Davide Luglietto, and 2 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-7525175/v1 This work is licensed under a CC BY 4.0 License Status: Published Journal Publication published 23 Jan, 2026 Read the published version in Neuroradiology → Version 1 posted You are reading this latest preprint version Abstract Purpose To quantify interobserver agreement in Spetzler–Martin (SM) and Spetzler–Ponce (SP) grading of pediatric brain arteriovenous malformations (AVMs), locate sources of variability, and test a composite Disagreement Index (DI). Methods Forty-five consecutive pediatric AVMs were independently graded by three neurosurgical residents without prior calibration. SM components (eloquence, venous drainage, nidus size) and SP class were assigned; nidus morphology (compact vs diffuse) was scored by two raters. Agreement was estimated with Fleiss’/Cohen’s κ and ICC(2,1); dispersion with across-rater standard deviation and Shannon entropy. Borderline cases were prespecified (SM 2–3; SP class transitions). DI combined entropy, SM-score dispersion, and component mismatch. Results SM scores showed moderate numerical agreement (ICC 0.72) but minimal categorical concordance (Fleiss’ κ 0.04). SP improved overall agreement (Fleiss’ κ 0.49), yet 17/45 (37.8%) cases crossed SP boundaries. Component reliability differed: eloquence κ 0.29, nidus size κ 0.41, venous drainage κ 0.58. Nidus morphology showed low reproducibility between two raters (≈ 49% agreement; Cohen’s κ − 0.15). Ten cases spanned the SM 2–3 threshold. DI ranged 0.15–1.00 (median 0.46) and isolated a small subset of highly discordant cases; eloquence was the primary driver in 8/10. Conclusions Interobserver variability concentrates at decision thresholds and is driven chiefly by how eloquence is interpreted. Standardized definitions, reporting measured nidus dimensions with SM bins, and routine lesion-to-eloquence distance may stabilize grading. DI can flag “teaching” cases and support calibration over time. Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Introduction In neurovascular practice, grading of brain arteriovenous malformations (AVMs) is central to therapeutic decision-making [ 11 , 13 , 12 , 10 , 8 , 5 , 1 ]. The Spetzler–Martin (SM) scale—based on nidus size, eloquence of adjacent cortex, and venous drainage—remains the most widely adopted risk stratification tool, while the Spetzler–Ponce (SP) classification provides a simplified derivative for treatment selection [ 11 , 12 ]. Despite their central role, the real-world application of these systems is often inconsistent [ 4 ]. Explicit AVM grades are rarely documented in multidisciplinary discussions and are frequently reconstructed retrospectively for research, often with unclear attribution or methodology. This raises doubts as to whether treatment strategies are genuinely informed by objective grading, and whether grading itself can serve as a reliable anchor for research comparability. Although the original SM study reported high reproducibility, subsequent investigations have demonstrated considerable interobserver variability, particularly in borderline or anatomically complex AVMs [ 6 , 4 , 2 ]. Lawton et al. noted disagreement rates approaching 25%, observing that “some unusual AVMs expose the system’s imprecision and subjectivity” [ 3 ]. The problem is especially evident in posterior fossa AVMs, where conventional definitions of “deep drainage” do not align with cerebellar anatomy, potentially distorting both classification and risk estimates [ 9 ]. Beyond anatomical ambiguities, subjective interpretation of parameters such as eloquence and nidus morphology further undermines consistency [ 7 ]. These limitations underscore the need for systematic and quantitative analyses of reproducibility. In this study, we evaluate interobserver agreement in SM and SP grading within a consecutive cohort of pediatric AVMs. Using blinded assessments by three neurosurgical residents, we quantify agreement for each component, characterize patterns of discordance, and examine the contribution of borderline cases and morphology. By applying entropy analysis, unsupervised clustering, and a composite Disagreement Index, we propose a reproducible framework to identify sources of variability and inform strategies for calibration, standardization, and more consistent application of AVM grading in both clinical and research contexts. Materials and Methods Study Design and Observers We assessed interobserver variability in AVM grading with the SM and SP systems. Forty-five consecutive pediatric AVMs were retrieved from a prospectively maintained, single-center neurovascular database. Digital subtraction angiography (DSA), MRI, and CT were independently reviewed by three neurosurgical residents familiar with AVM gradings. To reflect real-world practice, observers were not calibrated beforehand and were blinded to one another’s ratings and to treatment data. Cases were reviewed in fixed chronological order. Grading Protocol For each AVM, SM components were recorded: (1) eloquence of adjacent cortex, (2) venous drainage (superficial vs deep), and (3) nidus size ( 6 cm). The summed SM score (range 1–5) was converted to SP class (1–2 = A; 3 = B; 4–5 = C). Nidus morphology (compact vs diffuse) was also assessed without reference criteria to replicate real-world subjectivity. Agreement Metrics Categorical agreement was assessed with Fleiss’ κ (overall) and pairwise Cohen’s κ (A–B, A–C, B–C). Numerical agreement for SM scores was evaluated with the ICC (two-way random-effects, single measure). Case-level variability was quantified by the standard deviation (SD) of SM scores across raters and by Shannon entropy of the categorical distributions (base 2 unless otherwise specified; 0 = perfect agreement; higher values = greater disagreement). Borderline cases were defined a priori as AVMs graded across SM 2 vs 3 or across SP class boundaries (A↔B or B↔C). Temporal Trends and Clustering To test for learning or fatigue, agreement metrics were compared between early (cases 1–23) and late (cases 24–45) groups. Patterns of disagreement were explored with k-means clustering (k = 3) using entropy, SD, and component-level mismatch indicators as features. Cluster stability was checked across 100 random seeds. Composite Disagreement Index (operational definition) To aid calibration, we defined a case-level index combining dispersion and categorical discordance: Let H_SP be the Shannon entropy of the SP class distribution across raters; normalize by the maximum log base (H′_SP = H_SP / log_2 3). Let SD_SM be the across-rater SD of the SM score; normalize as SD′_SM = min(SD_SM / 2, 1) (bounded by the theoretical maximum for three raters on a 1–5 scale). Let M_comp be the proportion of SM components (0–3) with any disagreement (M′_comp = M_comp / 3). The Disagreement Index is then: DI = 0.4 × H′_SP + 0.4 × SD′_SM + 0.2 × M′_comp, rescaled to 0–1 for display. Higher values indicate greater disagreement. H′_SP captures threshold instability with direct clinical implications; SD′_SM reflects numeric dispersion; M′_comp localizes which domains disagree. We pre-specified weights to privilege clinical thresholds while maintaining balance. Sensitivity analyses with equal weights and alternative normalizations are reported in the Supplement. Software and Reporting Analyses were performed in Python (v3.10) and R (v4.3.1) with NumPy, SciPy, Pandas, and scikit-learn. Visualizations were generated with Matplotlib/Seaborn. Significance was set at p < 0.05. We followed GRRAS recommendations for reliability and agreement studies. Results Overall Interobserver Agreement Agreement on the Spetzler–Martin (SM) score was limited (Table 1 , Fig. 1 ). Numerical concordance was moderate (ICC = 0.72), whereas categorical agreement across SM categories was minimal (Fleiss’ κ = 0.04). For the Spetzler–Ponce (SP) classification, pairwise Cohen’s κ were A–B = 0.529, A–C = 0.583, and B–C = 0.435; the overall Fleiss’ κ was 0.49. At the case level, full concordance across all three raters occurred in 26/45 (57.8%), partial agreement in 14/45 (31.1%), and complete discordance in 3/45 (6.7%) (Fig. 2 ). Component-Level Variability and Borderline Cases Agreement differed across SM components: eloquence was least reproducible (κ = 0.29), nidus size showed intermediate consistency (κ = 0.41), and venous drainage was most reproducible (κ = 0.58). Numerical spread paralleled these patterns: the mean SD of SM scores was 0.32 (vs 0.21 for SP), and the mean Shannon entropy was 0.59 (vs 0.46 for SP). Collapsing SM to SP exposed threshold instability: 17/45 (37.8%) cases crossed SP class boundaries (A↔B = 6, B↔C = 7, A↔C = 4), while 28/45 (62.2%) showed no SP change (Fig. 3 ). Ten AVMs spanned the SM 2–3 boundary across observers. Consistent with these metrics, the primary driver of SM-component mismatch was eloquence in 24/45 (53%), with less frequent contributions from nidus size (4/45; 9%) and venous drainage (3/45; 7%); 14/45 (31%) had no SM-component mismatch (Fig. 3 ). In parallel, nidus morphology (compact vs diffuse) showed low agreement between two raters (≈ 49%; Cohen’s κ = −0.15), reinforcing the contribution of subjective elements to grading variability. Temporal Trends and Clustering Phenotypes Chronological analysis revealed no evidence of calibration or learning; variability increased over time (SM entropy 0.36 → 0.83; SP 0.30 → 0.63). Unsupervised clustering consistently identified three phenotypes: (1) complex AVMs with multi-component disagreement, (2) eloquence-centric discordance, and (3) straightforward cases with near-complete concordance (Fig. 2 ). Composite Disagreement Index The composite Disagreement Index (DI) ranged 0.15–1.00 (median 0.46). The ten most discordant cases (DI > 0.75) combined multi-component variability with frequent SP transitions; eloquence was the primary driver in 8/10, followed by venous drainage and nidus size (Fig. 4 ). These high-DI exemplars serve as practical benchmarks for targeted calibration and for tracking improvements after training. Discussion Overview of Main Findings Grading systems for AVMs appear most vulnerable precisely where they are intended to guide decisions. The Spetzler–Martin (SM) scale showed moderate numerical agreement (ICC = 0.72), yet categorical concordance was poor (Fleiss’ κ = 0.04), particularly near decision thresholds. The simplified Spetzler–Ponce (SP) scheme reduced dispersion (κ = 0.49), but 17/45 cases (37.8%) crossed at least one class boundary, revealing substantial threshold instability. Component-level agreement followed a consistent hierarchy: eloquence (κ = 0.29) was the most inconsistent, followed by nidus size (κ = 0.41), while venous drainage was relatively stable (κ = 0.58). Temporal analysis showed increasing entropy over time, without evidence of spontaneous calibration. Clustering analyses confirmed a structured pattern of disagreement, and our composite Disagreement Index (DI), ranging from 0.15 to 1.00 (median 0.46), isolated a small subset of highly unstable cases, 8/10 of which were driven by variability in eloquence assignment. These cases represent valuable opportunities for targeted calibration. Nidus: Size and Morphology Nidus size demonstrated only moderate reproducibility, consistent with the limitations of coarse binning in the SM system. Lesions at opposite ends of the same category (e.g., 0.9 cm vs 2.9 cm) may differ significantly in surgical risk despite being classified identically. Morphology proved even more problematic, with poor interobserver agreement (κ = − 0.15), highlighting the subjective nature of “compact” versus “diffuse” designations. These labels are sensitive to imaging quality, projection angle, and interpretive bias—particularly in deep or irregular lesions [ 2 , 5 ]. Our findings support incorporating measured diameter or volume alongside SM bins to mitigate intra-bin heterogeneity. Objective proxies for morphology—such as compactness indices or surface-to-volume ratios—may provide more reproducible alternatives. Until such tools are validated, consensus review of diffuse cases may improve consistency. This is particularly relevant for SM grade III AVMs, where heterogeneity in S/V/E composition can substantially influence treatment decisions [ 3 ]. Venous Drainage Venous drainage was the most reproducible SM component, yet classification remains imperfect, especially in infratentorial regions. In cerebellar AVMs, for example, veins may appear anatomically deep but drain superficially—or vice versa—complicating the application of supratentorial definitions [14]. These mismatches can distort grading and risk estimates, particularly near the SP class B/C threshold. Moreover, deep drainage retains independent prognostic value in multivariate models that include LED and diffuseness [ 9 ]. Our data reinforce the need for a more anatomically explicit approach: naming principal outflows (e.g., vein of Galen, thalamostriate vein) and including a posterior fossa supplement may reduce misclassification [14,15]. Distinguishing between true deep drainage and secondary rerouting from cortical venous thrombosis is also critical, as it may reflect lesion instability and influence surgical planning. Eloquent Location Eloquence was the least reproducible component and the principal driver of classification instability (Fig. 5 ). The binary SM label does not account for cortical plasticity, which can shift language or motor function to adjacent or even contralateral regions in patients with long-standing AVMs [13–16]. Nor does it consider proximity to eloquent subcortical tracts, such as the corticospinal tract or optic radiations, which may be more predictive of functional outcome than cortical landmarks alone. Models that incorporate lesion-to-eloquence distance (LED), such as the HDVL grading system, have shown superior prognostic performance (AUC 0.82 vs 0.71 for SM) [ 12 ]. In our series, eloquence drove 8/10 high-DI cases, highlighting its central role in interobserver variability. To improve grading reliability, we advocate integrating LED into routine documentation, specifying whether the at-risk substrate is cortical or tract-based, and using task-based imaging when available. Moving from a binary label to a graded, anatomically grounded measure of functional risk will improve both reproducibility and clinical relevance. Disagreement Index and Implications The Disagreement Index (DI) proved effective in stratifying cases by grading instability. High-DI cases concentrated near decision thresholds, where divergent treatment recommendations are most likely to arise. The increase in entropy over time suggests that ambiguity may accumulate, rather than resolve, in the absence of structured calibration. Operationally, the DI enables targeted quality improvement: calibration efforts can focus on high-DI cases using brief, reproducible strategies. Three habits—recording actual nidus dimensions, naming venous outflows (with posterior fossa notation), and documenting LED with the specific eloquent substrate—translate the main sources of variability into standardized, auditable fields. A curated atlas of high-DI cases can further support training and harmonization. These steps preserve the usability of SM/SP while aligning grading more closely with functional risk. Strengths and Limitations This study leveraged a consecutive pediatric cohort, blinded independent assessments, and a deliberately uncalibrated design to reflect real-world practice. The use of complementary metrics—κ, ICC, entropy, clustering, and the Disagreement Index—allowed for a multidimensional analysis of grading variability. Limitations include the single-center design, use of resident raters, fixed case order, lack of intra-rater testing, and restricted morphology assessment (available for only two observers). The generalizability of our findings and the weighting of the DI warrant external validation. Conclusions Both Spetzler–Martin and Spetzler–Ponce grading systems exhibit significant interobserver variability, particularly near decision thresholds. Eloquence and morphology are the least reproducible components, underscoring the need for structured definitions and objective augmentation. The Disagreement Index offers a practical tool for identifying unstable cases and guiding calibration. To preserve the clinical and research value of AVM grading, targeted standardization and functional integration are urgently needed. Declarations Ethics Approval and Consent to Participate All patients (or their legal guardians) provided written informed consent for the use of their anonymized imaging data. The study was conducted in accordance with the ethical standards of the Declaration of Helsinki. Since the study did not alter therapeutic strategies or involve interventional procedures, formal approval by an Institutional Review Board or Ethics Committee was not required. Human Ethics and Consent to Participate Declarations Human Ethics and Consent to Participate declarations: applicable. Patients (or parents/guardians) provided informed consent. Competing Interests The authors declare that they have no competing interests. Funding This research did not receive any specific grant from funding agencies in the public, commercial, or not-for-profit sectors. Author Contribution Author ContributionsF.M.C. Lioi: Conceptualization, methodology, data collection, analysis, visualization, writing—original draft, and project administration.D.F.: Support with figure preparation, data visualization, and writing—review and editing.C. E. M.: Supervision, critical review of the methodology, and writing—review and editing.A.D., D. L., and A. P. G.: Review of the final manuscript and contribution to critical revision of the content.All authors reviewed and approved the final version of the manuscript and agree to be accountable for all aspects of the work. References Davies JM, Kim H, Young WL, Lawton MT (2012) Classification schemes for arteriovenous malformations. Neurosurg Clin N Am 23:43–53. 10.1016/j.nec.2011.09.002 Du R, Dowd CF, Johnston SC, Young WL, Lawton MT (2005) Interobserver variability in grading of brain arteriovenous malformations using the Spetzler-Martin system. Neurosurgery 57:668–675 discussion 668–675 Du R, Keyoung HM, Dowd CF, Young WL, Lawton MT (2007) The effects of diffuseness and deep perforating artery supply on outcomes after microsurgical resection of brain arteriovenous malformations. Neurosurgery 60:638–646 discussion 646 – 638. 10.1227/01.Neu.0000255401.46151.8a Griessenauer CJ, Miller JH, Agee BS, Fisher WS 3rd, Curé JK, Chapman PR, Foreman PM, Fisher WA, Witcher AC, Walters BC (2014) Observer reliability of arteriovenous malformations grading scales using current imaging modalities. J Neurosurg 120:1179–1187. 10.3171/2014.2.Jns131262 Grüter BE, Sun W, Fierstra J, Regli L, Germans MR (2021) Systematic review of brain arteriovenous malformation grading systems evaluating microsurgical treatment recommendation. Neurosurg Rev 44:2571–2582. 10.1007/s10143-020-01464-3 Halim AX, Young WL, Johnston SC (2002) Reliability of angiographic assessment of brain arteriovenous malformations. Stroke 33:1508–1509 Jiao Y, Lin F, Wu J, Li H, Wang L, Jin Z, Wang S, Cao Y (2018) A supplementary grading scale combining lesion-to-eloquence distance for predicting surgical outcomes of patients with brain arteriovenous malformations. J Neurosurg 128:530–540. 10.3171/2016.10.Jns161415 Moon K, Levitt MR, Almefty RO, Nakaji P, Albuquerque FC, Zabramski JM, Wanebo JE, McDougall CG, Spetzler RF (2015) Safety and Efficacy of Surgical Resection of Unruptured Low-grade Arteriovenous Malformations From the Modern Decade. Neurosurgery 77:948–952 discussion 952 – 943. 10.1227/neu.0000000000000968 Nisson PL, Fard SA, Walter CM, Johnstone CM, Mooney MA, Tayebi Meybodi A, Lang M, Kim H, Jahnke H, Roe DJ, Dumont TM, Lemole GM, Spetzler RF, Lawton MT (2020) A novel proposed grading system for cerebellar arteriovenous malformations. J Neurosurg 132:1105–1115. 10.3171/2018.12.Jns181677 Rahme RJ, Singh R, De La Pena N, Turcotte EL, Bendok BR (2022) Arteriovenous Malformations: Treatment and Management. In: Mascitelli JR, Binning MJ (eds) Introduction to Vascular Neurosurgery. Springer International Publishing, Cham, pp 389–410. doi: 10.1007/978-3-030-88196-2_20 Spetzler RF, Martin NA (1986) A proposed grading system for arteriovenous malformations. J Neurosurg 65:476–483. 10.3171/jns.1986.65.4.0476 Spetzler RF, Ponce FA (2011) A 3-tier classification of cerebral arteriovenous malformations. Clinical article. J Neurosurg 114:842–849. 10.3171/2010.8.Jns10663 van Beijnum J, van der Worp HB, Buis DR, Al-Shahi Salman R, Kappelle LJ, Rinkel GJ, van der Sprenkel JW, Vandertop WP, Algra A, Klijn CJ (2011) Treatment of brain arteriovenous malformations: a systematic review and meta-analysis. JAMA 306:2011–2019. 10.1001/jama.2011.1632 Tables Table 1. Summary of interobserver agreement in AVM grading. Intraclass correlation (ICC) and Fleiss’ κ are reported for overall Spetzler–Martin (SM) and Spetzler–Ponce (SP) classifications. Component-level agreement is shown for eloquence, venous drainage, and nidus size (κ values), while nidus morphology is expressed as the percentage of cases with complete concordance. The data highlight moderate numerical agreement for SM scores but only fair categorical agreement, with eloquence and nidus morphology (compact vs diffuse) emerging as the least reproducible parameters. Domain Metric / Agreement measure Result Spetzler–Martin (SM) Intraclass correlation (ICC) 0.72 Fleiss’ κ (categorical) 0.04 Spetzler–Ponce (SP) Fleiss’ κ 0.49 Component-level Eloquence (κ) 0.29 Venous drainage (κ) 0.58 Nidus size (κ) 0.41 Nidus morphology (κ) –0.15 Additional Declarations No competing interests reported. Cite Share Download PDF Status: Published Journal Publication published 23 Jan, 2026 Read the published version in Neuroradiology → Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-7525175","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":533937657,"identity":"bf2e45b4-fcb2-4319-bd03-9faa52bdbcda","order_by":0,"name":"Francesco M.C. Lioi","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA60lEQVRIiWNgGAWjYBACPmYGhgMVcC6IxczcgFcLG0jLGTgXxGJmJKAFphAMGNvAJAEt7MwHDxzcY5fH397+8MPHebXR/O1ALT8qtuFxGFvCgQPPkoslzpwxlpy57XjujMOMDYw9Z27j0cJjcPjDAebEDRI5bMy8247lNgC1MDO24ddy4MCB+sQN8s+fMfPOOZY7n0gth4G2MJgx8zbU5G4grAXklwPHE2ecyTGWnHHsQO5GoJaD+PzCz38Y6JUD1Yn97ccffvhQU5c77/zhgw9+VODWgg4Og8kDRKsHgjpSFI+CUTAKRsEIAQDRyF4ftWcZvAAAAABJRU5ErkJggg==","orcid":"","institution":"Sapienza University of Rome","correspondingAuthor":true,"prefix":"","firstName":"Francesco","middleName":"M.C.","lastName":"Lioi","suffix":""},{"id":533937658,"identity":"6991bf11-ff02-4a6b-88ad-4f3e3c758c51","order_by":1,"name":"Alessandro Benedictis","email":"","orcid":"","institution":"Bambino Gesù Children's Hospital","correspondingAuthor":false,"prefix":"","firstName":"Alessandro","middleName":"","lastName":"Benedictis","suffix":""},{"id":533937659,"identity":"59644a8f-fbd9-4821-b058-4cc483527792","order_by":2,"name":"Davide Ferlito","email":"","orcid":"","institution":"Fondazione IRCCS San Gerardo dei Tintori","correspondingAuthor":false,"prefix":"","firstName":"Davide","middleName":"","lastName":"Ferlito","suffix":""},{"id":533937660,"identity":"c2843b1b-43c6-44bb-96cf-b56cdf3d1eeb","order_by":3,"name":"Davide Luglietto","email":"","orcid":"","institution":"University of Naples Federico II","correspondingAuthor":false,"prefix":"","firstName":"Davide","middleName":"","lastName":"Luglietto","suffix":""},{"id":533937661,"identity":"c0d252b0-e7dc-49f0-985e-2d9419140b32","order_by":4,"name":"Alberto P. Giraldo","email":"","orcid":"","institution":"Hospital del Mar","correspondingAuthor":false,"prefix":"","firstName":"Alberto","middleName":"P.","lastName":"Giraldo","suffix":""},{"id":533937662,"identity":"b62da753-bf75-4669-8ec1-cb27ddf2bc10","order_by":5,"name":"Carlo E. Marras","email":"","orcid":"","institution":"Bambino Gesù Children's Hospital","correspondingAuthor":false,"prefix":"","firstName":"Carlo","middleName":"E.","lastName":"Marras","suffix":""}],"badges":[],"createdAt":"2025-09-03 09:23:31","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-7525175/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-7525175/v1","draftVersion":[],"editorialEvents":[{"content":"https://doi.org/10.1007/s00234-025-03902-9","type":"published","date":"2026-01-23T15:57:17+00:00"}],"editorialNote":"","failedWorkflow":false,"files":[{"id":94480678,"identity":"2a76a3ad-b789-478b-88af-5a68290934b0","added_by":"auto","created_at":"2025-10-27 16:11:41","extension":"docx","order_by":0,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":45277,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript07.09.2025.docx","url":"https://assets-eu.researchsquare.com/files/rs-7525175/v1/b65027f7c22b7527421822ca.docx"},{"id":94480769,"identity":"3826e49a-9a06-4472-909d-c961b86536c5","added_by":"auto","created_at":"2025-10-27 16:11:56","extension":"docx","order_by":1,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":879139,"visible":true,"origin":"","legend":"","description":"","filename":"figures.docx","url":"https://assets-eu.researchsquare.com/files/rs-7525175/v1/2610b637e60bbbf7c0f3eb8f.docx"},{"id":94481037,"identity":"ea9dee80-6f32-4780-8a11-4dd3287b71e8","added_by":"auto","created_at":"2025-10-27 16:12:27","extension":"docx","order_by":2,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":14803,"visible":true,"origin":"","legend":"","description":"","filename":"table.docx","url":"https://assets-eu.researchsquare.com/files/rs-7525175/v1/67ce077169b9609d119dba94.docx"},{"id":94480170,"identity":"0578e47b-b67f-42a9-bf70-e988fde655d6","added_by":"auto","created_at":"2025-10-27 16:10:01","extension":"json","order_by":3,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":10124,"visible":true,"origin":"","legend":"","description":"","filename":"72c60f2bcaac473dadfbf151b27d9cf8.json","url":"https://assets-eu.researchsquare.com/files/rs-7525175/v1/00dfd19d58a87995fce7d659.json"},{"id":94480952,"identity":"02ac9d63-684c-4358-ab13-f5b769db0f16","added_by":"auto","created_at":"2025-10-27 16:12:17","extension":"xml","order_by":4,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":62331,"visible":true,"origin":"","legend":"","description":"","filename":"72c60f2bcaac473dadfbf151b27d9cf81enriched.xml","url":"https://assets-eu.researchsquare.com/files/rs-7525175/v1/2e685883ec0a06190b71181a.xml"},{"id":94480508,"identity":"228af982-e74d-4611-920c-e3de88f99599","added_by":"auto","created_at":"2025-10-27 16:11:20","extension":"jpeg","order_by":5,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":152931,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage1.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-7525175/v1/148f74ae081bcdd6fbaac279.jpeg"},{"id":94480771,"identity":"b840919c-a943-48e7-a231-370c568ca43c","added_by":"auto","created_at":"2025-10-27 16:11:57","extension":"jpeg","order_by":6,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":228226,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage2.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-7525175/v1/a4ca6be8de4ff4fb2982e0f2.jpeg"},{"id":94480657,"identity":"d47803a0-b6d1-45f4-a952-7962a4d6c2e4","added_by":"auto","created_at":"2025-10-27 16:11:34","extension":"jpeg","order_by":7,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":380249,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage3.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-7525175/v1/81267754e8429ca7716a0ad4.jpeg"},{"id":94480391,"identity":"43f59ad5-8dc5-46e4-8211-94b2b2b9f73a","added_by":"auto","created_at":"2025-10-27 16:10:56","extension":"png","order_by":8,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":74642,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage4.png","url":"https://assets-eu.researchsquare.com/files/rs-7525175/v1/124ebed62a5686cfe41cf967.png"},{"id":94480453,"identity":"0d7f8d5e-2da0-4e98-89de-97b9543ddf9a","added_by":"auto","created_at":"2025-10-27 16:11:14","extension":"jpeg","order_by":9,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":1074,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage5.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-7525175/v1/6d29e4eef576d8008b54e2dd.jpeg"},{"id":94480961,"identity":"bf3f6ffd-fd5a-4371-a160-a3423ebe0825","added_by":"auto","created_at":"2025-10-27 16:12:18","extension":"jpeg","order_by":10,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":54413,"visible":true,"origin":"","legend":"","description":"","filename":"groupimage1.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-7525175/v1/c0feda5f04f5f43fed99cb65.jpeg"},{"id":94480768,"identity":"e010a062-41e1-40db-b214-6e01d9814c96","added_by":"auto","created_at":"2025-10-27 16:11:56","extension":"png","order_by":11,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":28020,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage1.png","url":"https://assets-eu.researchsquare.com/files/rs-7525175/v1/6435930fa507c4c0f752c27f.png"},{"id":94480591,"identity":"4a3f86ab-b441-4c99-86d2-1cf60141bc78","added_by":"auto","created_at":"2025-10-27 16:11:24","extension":"png","order_by":12,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":44417,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage2.png","url":"https://assets-eu.researchsquare.com/files/rs-7525175/v1/e5f5c324d4bb04fef85c6d48.png"},{"id":94480925,"identity":"486bcaca-9872-4c35-bfca-81ec171e27ed","added_by":"auto","created_at":"2025-10-27 16:12:15","extension":"png","order_by":13,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":92386,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage3.png","url":"https://assets-eu.researchsquare.com/files/rs-7525175/v1/eb689b76aa145dd4fc692fe5.png"},{"id":94480619,"identity":"3f39ccef-fad9-4d75-bcea-72938310b66d","added_by":"auto","created_at":"2025-10-27 16:11:26","extension":"png","order_by":14,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":22735,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage4.png","url":"https://assets-eu.researchsquare.com/files/rs-7525175/v1/4ec0391790a67c406ec84539.png"},{"id":94480762,"identity":"2096554e-5906-4f7c-a8d7-953ccce0b7fc","added_by":"auto","created_at":"2025-10-27 16:11:55","extension":"png","order_by":15,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":935,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage5.png","url":"https://assets-eu.researchsquare.com/files/rs-7525175/v1/ea378bf05c3dc9c71aa437c2.png"},{"id":94480661,"identity":"c3b4bbf4-312a-4660-9d5a-8d92a45475e9","added_by":"auto","created_at":"2025-10-27 16:11:35","extension":"png","order_by":16,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":110046,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinegroupimage1.png","url":"https://assets-eu.researchsquare.com/files/rs-7525175/v1/686814636a6ab1997e585a27.png"},{"id":94480523,"identity":"af531f82-e0f1-4314-8953-e760b1908f91","added_by":"auto","created_at":"2025-10-27 16:11:22","extension":"xml","order_by":17,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":57153,"visible":true,"origin":"","legend":"","description":"","filename":"72c60f2bcaac473dadfbf151b27d9cf81structuring.xml","url":"https://assets-eu.researchsquare.com/files/rs-7525175/v1/5bca74c0d0f8fa18c6224889.xml"},{"id":94480943,"identity":"3cdf1edf-1288-45c2-91db-d1ca0d4d2303","added_by":"auto","created_at":"2025-10-27 16:12:16","extension":"html","order_by":18,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":68840,"visible":true,"origin":"","legend":"","description":"","filename":"earlyproof.html","url":"https://assets-eu.researchsquare.com/files/rs-7525175/v1/18363759a8a5efa8910659e5.html"},{"id":94481025,"identity":"b7ffb0a6-32ea-4131-998b-ed65642c2b49","added_by":"auto","created_at":"2025-10-27 16:12:25","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":36058,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eAgreement on SM components.\u003cbr\u003e\n \u003c/strong\u003eInterobserver reproducibility differed markedly across SM components. Agreement was lowest for eloquence (κ = 0.29), intermediate for nidus size (κ = 0.41), and highest for venous drainage (κ = 0.58). Vertical reference lines mark commonly used benchmarks for fair (κ ≈ 0.40) and moderate (κ ≈ 0.60) agreement.\u003c/p\u003e","description":"","filename":"image1.png","url":"https://assets-eu.researchsquare.com/files/rs-7525175/v1/263a01a831af76850bface99.png"},{"id":94481170,"identity":"208bcf1d-cc79-4ea9-99ac-486055ef58cb","added_by":"auto","created_at":"2025-10-27 16:12:51","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":64483,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eAgreement on SP classification.\u003cbr\u003e\n \u003c/strong\u003eHorizontal bars show the proportion of cases with full, partial, or complete disagreement across three independent raters. Labels report percentages and counts. Full concordance occurred in 26/45 (57.8%), partial agreement in 14/45 (31.1%), and complete discordance in 3/45 (6.7%). \u003cem\u003e(Two cases had missing SP ratings and are not displayed but are retained in the denominator.)\u003c/em\u003e Collapsing SM into SP yielded moderate overall agreement (Fleiss’ κ = 0.49), yet 17/45 (37.8%) cases still crossed SP class boundaries (A↔B, B↔C, or A↔C), underscoring residual threshold-related variability.\u003c/p\u003e","description":"","filename":"image2.png","url":"https://assets-eu.researchsquare.com/files/rs-7525175/v1/f7ec9915b7bc76619d50825b.png"},{"id":94480776,"identity":"2ad0d91a-4850-47a7-8588-ba6f9f1e70c9","added_by":"auto","created_at":"2025-10-27 16:11:57","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":133046,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eFrom SM component discordance to SP class transitions.\u003cbr\u003e\n \u003c/strong\u003eParallel-sets diagram for 45 pediatric AVMs showing how the Spetzler–Martin (SM) component selected as the primary driver of interobserver mismatch (left) relates to the resulting Spetzler–Ponce (SP) outcome across raters (right). Ribbon width is proportional to the number of cases; flows to No change are drawn with reduced opacity. Side bars report n (%) in each category. When multiple SM components disagreed, the primary driver was assigned by a prespecified hierarchy (Eloquence \u0026gt; Size \u0026gt; Venous); No SM mismatch denotes complete agreement on all three SM components. SP outcomes aggregate case-level dispersion into A↔B, B↔C, A↔C, or No change. In this cohort, 17/45 (37.8%) cases crossed an SP boundary (A↔B=6, B↔C=7, A↔C=4), with eloquence accounting for the majority of transitions.\u003c/p\u003e","description":"","filename":"image3.png","url":"https://assets-eu.researchsquare.com/files/rs-7525175/v1/89031128308cde1909a7a948.png"},{"id":94480772,"identity":"855d8d6c-766c-4f9a-b655-ecce63c8f7fb","added_by":"auto","created_at":"2025-10-27 16:11:57","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":74642,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eComposite disagreement summary (Top-5 cases, complete SP ratings).\u003cbr\u003e\n \u003c/strong\u003eHorizontal stacked bars show, for each anonymized patient (Pt #), the Disagreement Index (DI) decomposed into its components: H′SP (normalized Shannon entropy of the SP-class distribution across raters), SD′SM (normalized across-rater SD of Spetzler–Martin scores), and M′comp (proportion of SM-component mismatches). By definition, DI = 0.4·H′SP + 0.4·SD′SM + 0.2·M′comp (range 0–1); the numeric value of DI is reported at the end of each bar. Dashed vertical lines mark the cohort median DI and the Top-5 threshold (DI of the 5th-ranked case shown). To the right, the SP by rater triplet displays each rater’s SP class (A/B/C). The Driver panel flags which SM components disagreed (E = eloquence, V = venous drainage, S = size; filled square = mismatch, outlined = agreement). Thin horizontal rules separate the column headers (A/B/C; E/V/S) from the data. Cases are ordered by DI (highest to lowest).\u003c/p\u003e","description":"","filename":"image4.png","url":"https://assets-eu.researchsquare.com/files/rs-7525175/v1/f3d21494e9864e2b65fe2920.png"},{"id":94480437,"identity":"ee0da1d8-5b5b-4978-809a-f222ff3880ce","added_by":"auto","created_at":"2025-10-27 16:11:08","extension":"png","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":637800,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eEloquent-location disagreement illustrated.\u003cbr\u003e\nA\u003c/strong\u003e. Intraventricular AVM. On SM, “intraventricular” location is not eloquent per se, yet raters diverge because the nidus abuts periventricular pathways (fornix/optic radiations/capsular fibers). Some score this as non-eloquent while others as eloquent due to those reasons. \u003cstrong\u003eB\u003c/strong\u003e. Callosal AVM. This region itself is not classically included as SM-eloquent, but the nidus contacts the corpus callosum/cingulum, raising concern for interhemispheric and premotor–SMA networks. Here, some raters label eloquent on a tract basis, others do not.\u003c/p\u003e","description":"","filename":"image5.png","url":"https://assets-eu.researchsquare.com/files/rs-7525175/v1/9b9ff6d857bf2446c5b74bbd.png"},{"id":101151783,"identity":"a3f3b1b5-d891-40b9-9bef-6cd7a0bbc6ca","added_by":"auto","created_at":"2026-01-26 16:05:28","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":1875380,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-7525175/v1/d423765c-bd6a-4008-aa91-67d3c5755071.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"Reproducibility of AVM grading in clinical practice: A study of interobserver variability","fulltext":[{"header":"Introduction","content":"\u003cp\u003eIn neurovascular practice, grading of brain arteriovenous malformations (AVMs) is central to therapeutic decision-making [\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e, \u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e, \u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e, \u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e, \u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e, \u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e, \u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e]. The Spetzler\u0026ndash;Martin (SM) scale\u0026mdash;based on nidus size, eloquence of adjacent cortex, and venous drainage\u0026mdash;remains the most widely adopted risk stratification tool, while the Spetzler\u0026ndash;Ponce (SP) classification provides a simplified derivative for treatment selection [\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e, \u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e]. Despite their central role, the real-world application of these systems is often inconsistent [\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e]. Explicit AVM grades are rarely documented in multidisciplinary discussions and are frequently reconstructed retrospectively for research, often with unclear attribution or methodology. This raises doubts as to whether treatment strategies are genuinely informed by objective grading, and whether grading itself can serve as a reliable anchor for research comparability. Although the original SM study reported high reproducibility, subsequent investigations have demonstrated considerable interobserver variability, particularly in borderline or anatomically complex AVMs [\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e, \u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e, \u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e]. Lawton et al. noted disagreement rates approaching 25%, observing that \u0026ldquo;some unusual AVMs expose the system\u0026rsquo;s imprecision and subjectivity\u0026rdquo; [\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e]. The problem is especially evident in posterior fossa AVMs, where conventional definitions of \u0026ldquo;deep drainage\u0026rdquo; do not align with cerebellar anatomy, potentially distorting both classification and risk estimates [\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e]. Beyond anatomical ambiguities, subjective interpretation of parameters such as eloquence and nidus morphology further undermines consistency [\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e]. These limitations underscore the need for systematic and quantitative analyses of reproducibility. In this study, we evaluate interobserver agreement in SM and SP grading within a consecutive cohort of pediatric AVMs. Using blinded assessments by three neurosurgical residents, we quantify agreement for each component, characterize patterns of discordance, and examine the contribution of borderline cases and morphology. By applying entropy analysis, unsupervised clustering, and a composite Disagreement Index, we propose a reproducible framework to identify sources of variability and inform strategies for calibration, standardization, and more consistent application of AVM grading in both clinical and research contexts.\u003c/p\u003e"},{"header":"Materials and Methods","content":"\u003cdiv id=\"Sec3\" class=\"Section2\"\u003e\u003ch2\u003eStudy Design and Observers\u003c/h2\u003e\u003cp\u003eWe assessed interobserver variability in AVM grading with the SM and SP systems. Forty-five consecutive pediatric AVMs were retrieved from a prospectively maintained, single-center neurovascular database. Digital subtraction angiography (DSA), MRI, and CT were independently reviewed by three neurosurgical residents familiar with AVM gradings. To reflect real-world practice, observers were not calibrated beforehand and were blinded to one another\u0026rsquo;s ratings and to treatment data. Cases were reviewed in fixed chronological order.\u003c/p\u003e\u003c/div\u003e\n\u003ch3\u003eGrading Protocol\u003c/h3\u003e\n\u003cp\u003eFor each AVM, SM components were recorded: (1) eloquence of adjacent cortex, (2) venous drainage (superficial vs deep), and (3) nidus size (\u0026lt;\u0026thinsp;3 cm, 3\u0026ndash;6 cm, \u0026gt;\u0026thinsp;6 cm). The summed SM score (range 1\u0026ndash;5) was converted to SP class (1\u0026ndash;2\u0026thinsp;=\u0026thinsp;A; 3\u0026thinsp;=\u0026thinsp;B; 4\u0026ndash;5\u0026thinsp;=\u0026thinsp;C). Nidus morphology (compact vs diffuse) was also assessed without reference criteria to replicate real-world subjectivity.\u003c/p\u003e\n\u003ch3\u003eAgreement Metrics\u003c/h3\u003e\n\u003cp\u003eCategorical agreement was assessed with Fleiss\u0026rsquo; κ (overall) and pairwise Cohen\u0026rsquo;s κ (A\u0026ndash;B, A\u0026ndash;C, B\u0026ndash;C). Numerical agreement for SM scores was evaluated with the ICC (two-way random-effects, single measure). Case-level variability was quantified by the standard deviation (SD) of SM scores across raters and by Shannon entropy of the categorical distributions (base 2 unless otherwise specified; 0\u0026thinsp;=\u0026thinsp;perfect agreement; higher values\u0026thinsp;=\u0026thinsp;greater disagreement). Borderline cases were defined a priori as AVMs graded across SM 2 vs 3 or across SP class boundaries (A\u0026harr;B or B\u0026harr;C).\u003c/p\u003e\n\u003ch3\u003eTemporal Trends and Clustering\u003c/h3\u003e\n\u003cp\u003eTo test for learning or fatigue, agreement metrics were compared between early (cases 1\u0026ndash;23) and late (cases 24\u0026ndash;45) groups. Patterns of disagreement were explored with k-means clustering (k\u0026thinsp;=\u0026thinsp;3) using entropy, SD, and component-level mismatch indicators as features. Cluster stability was checked across 100 random seeds.\u003c/p\u003e\n\u003ch3\u003eComposite Disagreement Index (operational definition)\u003c/h3\u003e\n\u003cp\u003eTo aid calibration, we defined a case-level index combining dispersion and categorical discordance:\u003c/p\u003e\u003cp\u003e\u003cul\u003e\u003cli\u003e\u003cp\u003eLet H_SP be the Shannon entropy of the SP class distribution across raters; normalize by the maximum log base (H\u0026prime;_SP\u0026thinsp;=\u0026thinsp;H_SP / log_2 3).\u003c/p\u003e\u003c/li\u003e\u003cli\u003e\u003cp\u003eLet SD_SM be the across-rater SD of the SM score; normalize as SD\u0026prime;_SM\u0026thinsp;=\u0026thinsp;min(SD_SM / 2, 1) (bounded by the theoretical maximum for three raters on a 1\u0026ndash;5 scale).\u003c/p\u003e\u003c/li\u003e\u003cli\u003e\u003cp\u003eLet M_comp be the proportion of SM components (0\u0026ndash;3) with any disagreement (M\u0026prime;_comp\u0026thinsp;=\u0026thinsp;M_comp / 3).\u003c/p\u003e\u003c/li\u003e\u003c/ul\u003e\u003c/p\u003e\u003cp\u003eThe Disagreement Index is then: DI\u0026thinsp;=\u0026thinsp;0.4 \u0026times; H\u0026prime;_SP\u0026thinsp;+\u0026thinsp;0.4 \u0026times; SD\u0026prime;_SM\u0026thinsp;+\u0026thinsp;0.2 \u0026times; M\u0026prime;_comp, rescaled to 0\u0026ndash;1 for display. Higher values indicate greater disagreement. H\u0026prime;_SP captures threshold instability with direct clinical implications; SD\u0026prime;_SM reflects numeric dispersion; M\u0026prime;_comp localizes which domains disagree. We pre-specified weights to privilege clinical thresholds while maintaining balance. Sensitivity analyses with equal weights and alternative normalizations are reported in the Supplement.\u003c/p\u003e\u003cdiv id=\"Sec8\" class=\"Section2\"\u003e\u003ch2\u003eSoftware and Reporting\u003c/h2\u003e\u003cp\u003eAnalyses were performed in Python (v3.10) and R (v4.3.1) with NumPy, SciPy, Pandas, and scikit-learn. Visualizations were generated with Matplotlib/Seaborn. Significance was set at p\u0026thinsp;\u0026lt;\u0026thinsp;0.05. We followed GRRAS recommendations for reliability and agreement studies.\u003c/p\u003e\u003c/div\u003e"},{"header":"Results","content":"\u003cdiv id=\"Sec10\" class=\"Section2\"\u003e\n \u003ch2\u003eOverall Interobserver Agreement\u003c/h2\u003e\n \u003cp\u003eAgreement on the Spetzler\u0026ndash;Martin (SM) score was limited (Table \u003cspan class=\"InternalRef\"\u003e1\u003c/span\u003e, Fig. \u003cspan class=\"InternalRef\"\u003e1\u003c/span\u003e). Numerical concordance was moderate (ICC\u0026thinsp;=\u0026thinsp;0.72), whereas categorical agreement across SM categories was minimal (Fleiss\u0026rsquo; \u0026kappa;\u0026thinsp;=\u0026thinsp;0.04). For the Spetzler\u0026ndash;Ponce (SP) classification, pairwise Cohen\u0026rsquo;s \u0026kappa; were A\u0026ndash;B\u0026thinsp;=\u0026thinsp;0.529, A\u0026ndash;C\u0026thinsp;=\u0026thinsp;0.583, and B\u0026ndash;C\u0026thinsp;=\u0026thinsp;0.435; the overall Fleiss\u0026rsquo; \u0026kappa; was 0.49. At the case level, full concordance across all three raters occurred in 26/45 (57.8%), partial agreement in 14/45 (31.1%), and complete discordance in 3/45 (6.7%) (Fig. \u003cspan class=\"InternalRef\"\u003e2\u003c/span\u003e).\u003c/p\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec11\" class=\"Section2\"\u003e\n \u003ch2\u003eComponent-Level Variability and Borderline Cases\u003c/h2\u003e\n \u003cp\u003eAgreement differed across SM components: eloquence was least reproducible (\u0026kappa;\u0026thinsp;=\u0026thinsp;0.29), nidus size showed intermediate consistency (\u0026kappa;\u0026thinsp;=\u0026thinsp;0.41), and venous drainage was most reproducible (\u0026kappa;\u0026thinsp;=\u0026thinsp;0.58). Numerical spread paralleled these patterns: the mean SD of SM scores was 0.32 (vs 0.21 for SP), and the mean Shannon entropy was 0.59 (vs 0.46 for SP). Collapsing SM to SP exposed threshold instability: 17/45 (37.8%) cases crossed SP class boundaries (A\u0026harr;B\u0026thinsp;=\u0026thinsp;6, B\u0026harr;C\u0026thinsp;=\u0026thinsp;7, A\u0026harr;C\u0026thinsp;=\u0026thinsp;4), while 28/45 (62.2%) showed no SP change (Fig. \u003cspan class=\"InternalRef\"\u003e3\u003c/span\u003e). Ten AVMs spanned the SM 2\u0026ndash;3 boundary across observers. Consistent with these metrics, the primary driver of SM-component mismatch was eloquence in 24/45 (53%), with less frequent contributions from nidus size (4/45; 9%) and venous drainage (3/45; 7%); 14/45 (31%) had no SM-component mismatch (Fig. \u003cspan class=\"InternalRef\"\u003e3\u003c/span\u003e). In parallel, nidus morphology (compact vs diffuse) showed low agreement between two raters (\u0026asymp;\u0026thinsp;49%; Cohen\u0026rsquo;s \u0026kappa; = \u0026minus;0.15), reinforcing the contribution of subjective elements to grading variability.\u003c/p\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec12\" class=\"Section2\"\u003e\n \u003ch2\u003eTemporal Trends and Clustering Phenotypes\u003c/h2\u003e\n \u003cp\u003eChronological analysis revealed no evidence of calibration or learning; variability increased over time (SM entropy 0.36 \u0026rarr; 0.83; SP 0.30 \u0026rarr; 0.63). Unsupervised clustering consistently identified three phenotypes: (1) complex AVMs with multi-component disagreement, (2) eloquence-centric discordance, and (3) straightforward cases with near-complete concordance (Fig. \u003cspan class=\"InternalRef\"\u003e2\u003c/span\u003e).\u003c/p\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec13\" class=\"Section2\"\u003e\n \u003ch2\u003eComposite Disagreement Index\u003c/h2\u003e\n \u003cp\u003eThe composite Disagreement Index (DI) ranged 0.15\u0026ndash;1.00 (median 0.46). The ten most discordant cases (DI\u0026thinsp;\u0026gt;\u0026thinsp;0.75) combined multi-component variability with frequent SP transitions; eloquence was the primary driver in 8/10, followed by venous drainage and nidus size (Fig. \u003cspan class=\"InternalRef\"\u003e4\u003c/span\u003e). These high-DI exemplars serve as practical benchmarks for targeted calibration and for tracking improvements after training.\u003c/p\u003e\n\u003c/div\u003e"},{"header":"Discussion","content":"\u003cdiv id=\"Sec15\" class=\"Section2\"\u003e\u003ch2\u003eOverview of Main Findings\u003c/h2\u003e\u003cp\u003eGrading systems for AVMs appear most vulnerable precisely where they are intended to guide decisions. The Spetzler\u0026ndash;Martin (SM) scale showed moderate numerical agreement (ICC\u0026thinsp;=\u0026thinsp;0.72), yet categorical concordance was poor (Fleiss\u0026rsquo; κ\u0026thinsp;=\u0026thinsp;0.04), particularly near decision thresholds. The simplified Spetzler\u0026ndash;Ponce (SP) scheme reduced dispersion (κ\u0026thinsp;=\u0026thinsp;0.49), but 17/45 cases (37.8%) crossed at least one class boundary, revealing substantial threshold instability. Component-level agreement followed a consistent hierarchy: eloquence (κ\u0026thinsp;=\u0026thinsp;0.29) was the most inconsistent, followed by nidus size (κ\u0026thinsp;=\u0026thinsp;0.41), while venous drainage was relatively stable (κ\u0026thinsp;=\u0026thinsp;0.58). Temporal analysis showed increasing entropy over time, without evidence of spontaneous calibration. Clustering analyses confirmed a structured pattern of disagreement, and our composite Disagreement Index (DI), ranging from 0.15 to 1.00 (median 0.46), isolated a small subset of highly unstable cases, 8/10 of which were driven by variability in eloquence assignment. These cases represent valuable opportunities for targeted calibration.\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec16\" class=\"Section2\"\u003e\u003ch2\u003eNidus: Size and Morphology\u003c/h2\u003e\u003cp\u003eNidus size demonstrated only moderate reproducibility, consistent with the limitations of coarse binning in the SM system. Lesions at opposite ends of the same category (e.g., 0.9 cm vs 2.9 cm) may differ significantly in surgical risk despite being classified identically. Morphology proved even more problematic, with poor interobserver agreement (κ = \u0026minus;\u0026thinsp;0.15), highlighting the subjective nature of \u0026ldquo;compact\u0026rdquo; versus \u0026ldquo;diffuse\u0026rdquo; designations. These labels are sensitive to imaging quality, projection angle, and interpretive bias\u0026mdash;particularly in deep or irregular lesions [\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e, \u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e]. Our findings support incorporating measured diameter or volume alongside SM bins to mitigate intra-bin heterogeneity. Objective proxies for morphology\u0026mdash;such as compactness indices or surface-to-volume ratios\u0026mdash;may provide more reproducible alternatives. Until such tools are validated, consensus review of diffuse cases may improve consistency. This is particularly relevant for SM grade III AVMs, where heterogeneity in S/V/E composition can substantially influence treatment decisions [\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e].\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec17\" class=\"Section2\"\u003e\u003ch2\u003eVenous Drainage\u003c/h2\u003e\u003cp\u003eVenous drainage was the most reproducible SM component, yet classification remains imperfect, especially in infratentorial regions. In cerebellar AVMs, for example, veins may appear anatomically deep but drain superficially\u0026mdash;or vice versa\u0026mdash;complicating the application of supratentorial definitions [14]. These mismatches can distort grading and risk estimates, particularly near the SP class B/C threshold. Moreover, deep drainage retains independent prognostic value in multivariate models that include LED and diffuseness [\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e]. Our data reinforce the need for a more anatomically explicit approach: naming principal outflows (e.g., vein of Galen, thalamostriate vein) and including a posterior fossa supplement may reduce misclassification [14,15]. Distinguishing between true deep drainage and secondary rerouting from cortical venous thrombosis is also critical, as it may reflect lesion instability and influence surgical planning.\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec18\" class=\"Section2\"\u003e\u003ch2\u003eEloquent Location\u003c/h2\u003e\u003cp\u003eEloquence was the least reproducible component and the principal driver of classification instability (Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003e). The binary SM label does not account for cortical plasticity, which can shift language or motor function to adjacent or even contralateral regions in patients with long-standing AVMs [13\u0026ndash;16]. Nor does it consider proximity to eloquent subcortical tracts, such as the corticospinal tract or optic radiations, which may be more predictive of functional outcome than cortical landmarks alone. Models that incorporate lesion-to-eloquence distance (LED), such as the HDVL grading system, have shown superior prognostic performance (AUC 0.82 vs 0.71 for SM) [\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e]. In our series, eloquence drove 8/10 high-DI cases, highlighting its central role in interobserver variability. To improve grading reliability, we advocate integrating LED into routine documentation, specifying whether the at-risk substrate is cortical or tract-based, and using task-based imaging when available. Moving from a binary label to a graded, anatomically grounded measure of functional risk will improve both reproducibility and clinical relevance.\u003c/p\u003e\u003cp\u003e\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec19\" class=\"Section2\"\u003e\u003ch2\u003eDisagreement Index and Implications\u003c/h2\u003e\u003cp\u003eThe Disagreement Index (DI) proved effective in stratifying cases by grading instability. High-DI cases concentrated near decision thresholds, where divergent treatment recommendations are most likely to arise. The increase in entropy over time suggests that ambiguity may accumulate, rather than resolve, in the absence of structured calibration. Operationally, the DI enables targeted quality improvement: calibration efforts can focus on high-DI cases using brief, reproducible strategies. Three habits\u0026mdash;recording actual nidus dimensions, naming venous outflows (with posterior fossa notation), and documenting LED with the specific eloquent substrate\u0026mdash;translate the main sources of variability into standardized, auditable fields. A curated atlas of high-DI cases can further support training and harmonization. These steps preserve the usability of SM/SP while aligning grading more closely with functional risk.\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec20\" class=\"Section2\"\u003e\u003ch2\u003eStrengths and Limitations\u003c/h2\u003e\u003cp\u003eThis study leveraged a consecutive pediatric cohort, blinded independent assessments, and a deliberately uncalibrated design to reflect real-world practice. The use of complementary metrics\u0026mdash;κ, ICC, entropy, clustering, and the Disagreement Index\u0026mdash;allowed for a multidimensional analysis of grading variability. Limitations include the single-center design, use of resident raters, fixed case order, lack of intra-rater testing, and restricted morphology assessment (available for only two observers). The generalizability of our findings and the weighting of the DI warrant external validation.\u003c/p\u003e\u003c/div\u003e"},{"header":"Conclusions","content":"\u003cp\u003eBoth Spetzler\u0026ndash;Martin and Spetzler\u0026ndash;Ponce grading systems exhibit significant interobserver variability, particularly near decision thresholds. Eloquence and morphology are the least reproducible components, underscoring the need for structured definitions and objective augmentation. The Disagreement Index offers a practical tool for identifying unstable cases and guiding calibration. To preserve the clinical and research value of AVM grading, targeted standardization and functional integration are urgently needed.\u003c/p\u003e"},{"header":"Declarations","content":"\u003ch2\u003eEthics Approval and Consent to Participate\u003c/h2\u003e\u003cp\u003eAll patients (or their legal guardians) provided written informed consent for the use of their anonymized imaging data. The study was conducted in accordance with the ethical standards of the Declaration of Helsinki. Since the study did not alter therapeutic strategies or involve interventional procedures, formal approval by an Institutional Review Board or Ethics Committee was not required.\u003c/p\u003e\u003ch2\u003eHuman Ethics and Consent to Participate Declarations\u003c/h2\u003e\u003cp\u003eHuman Ethics and Consent to Participate declarations: applicable. Patients (or parents/guardians) provided informed consent.\u003c/p\u003e\u003ch2\u003eCompeting Interests\u003c/h2\u003e\u003cp\u003eThe authors declare that they have no competing interests.\u003c/p\u003e\u003ch2\u003eFunding\u003c/h2\u003e\u003cp\u003eThis research did not receive any specific grant from funding agencies in the public, commercial, or not-for-profit sectors.\u003c/p\u003e\u003ch2\u003eAuthor Contribution\u003c/h2\u003e\u003cp\u003eAuthor ContributionsF.M.C. Lioi: Conceptualization, methodology, data collection, analysis, visualization, writing\u0026mdash;original draft, and project administration.D.F.: Support with figure preparation, data visualization, and writing\u0026mdash;review and editing.C. E. M.: Supervision, critical review of the methodology, and writing\u0026mdash;review and editing.A.D., D. L., and A. P. G.: Review of the final manuscript and contribution to critical revision of the content.All authors reviewed and approved the final version of the manuscript and agree to be accountable for all aspects of the work.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eDavies JM, Kim H, Young WL, Lawton MT (2012) Classification schemes for arteriovenous malformations. Neurosurg Clin N Am 23:43\u0026ndash;53. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1016/j.nec.2011.09.002\u003c/span\u003e\u003cspan address=\"10.1016/j.nec.2011.09.002\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eDu R, Dowd CF, Johnston SC, Young WL, Lawton MT (2005) Interobserver variability in grading of brain arteriovenous malformations using the Spetzler-Martin system. Neurosurgery 57:668\u0026ndash;675 discussion 668\u0026ndash;675\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eDu R, Keyoung HM, Dowd CF, Young WL, Lawton MT (2007) The effects of diffuseness and deep perforating artery supply on outcomes after microsurgical resection of brain arteriovenous malformations. Neurosurgery 60:638\u0026ndash;646 discussion 646\u0026thinsp;\u0026ndash;\u0026thinsp;638. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1227/01.Neu.0000255401.46151.8a\u003c/span\u003e\u003cspan address=\"10.1227/01.Neu.0000255401.46151.8a\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eGriessenauer CJ, Miller JH, Agee BS, Fisher WS 3rd, Cur\u0026eacute; JK, Chapman PR, Foreman PM, Fisher WA, Witcher AC, Walters BC (2014) Observer reliability of arteriovenous malformations grading scales using current imaging modalities. J Neurosurg 120:1179\u0026ndash;1187. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.3171/2014.2.Jns131262\u003c/span\u003e\u003cspan address=\"10.3171/2014.2.Jns131262\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eGr\u0026uuml;ter BE, Sun W, Fierstra J, Regli L, Germans MR (2021) Systematic review of brain arteriovenous malformation grading systems evaluating microsurgical treatment recommendation. Neurosurg Rev 44:2571\u0026ndash;2582. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1007/s10143-020-01464-3\u003c/span\u003e\u003cspan address=\"10.1007/s10143-020-01464-3\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eHalim AX, Young WL, Johnston SC (2002) Reliability of angiographic assessment of brain arteriovenous malformations. Stroke 33:1508\u0026ndash;1509\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eJiao Y, Lin F, Wu J, Li H, Wang L, Jin Z, Wang S, Cao Y (2018) A supplementary grading scale combining lesion-to-eloquence distance for predicting surgical outcomes of patients with brain arteriovenous malformations. J Neurosurg 128:530\u0026ndash;540. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.3171/2016.10.Jns161415\u003c/span\u003e\u003cspan address=\"10.3171/2016.10.Jns161415\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eMoon K, Levitt MR, Almefty RO, Nakaji P, Albuquerque FC, Zabramski JM, Wanebo JE, McDougall CG, Spetzler RF (2015) Safety and Efficacy of Surgical Resection of Unruptured Low-grade Arteriovenous Malformations From the Modern Decade. Neurosurgery 77:948\u0026ndash;952 discussion 952\u0026thinsp;\u0026ndash;\u0026thinsp;943. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1227/neu.0000000000000968\u003c/span\u003e\u003cspan address=\"10.1227/neu.0000000000000968\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eNisson PL, Fard SA, Walter CM, Johnstone CM, Mooney MA, Tayebi Meybodi A, Lang M, Kim H, Jahnke H, Roe DJ, Dumont TM, Lemole GM, Spetzler RF, Lawton MT (2020) A novel proposed grading system for cerebellar arteriovenous malformations. J Neurosurg 132:1105\u0026ndash;1115. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.3171/2018.12.Jns181677\u003c/span\u003e\u003cspan address=\"10.3171/2018.12.Jns181677\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eRahme RJ, Singh R, De La Pena N, Turcotte EL, Bendok BR (2022) Arteriovenous Malformations: Treatment and Management. In: Mascitelli JR, Binning MJ (eds) Introduction to Vascular Neurosurgery. Springer International Publishing, Cham, pp 389\u0026ndash;410. doi:\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1007/978-3-030-88196-2_20\u003c/span\u003e\u003cspan address=\"10.1007/978-3-030-88196-2_20\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eSpetzler RF, Martin NA (1986) A proposed grading system for arteriovenous malformations. J Neurosurg 65:476\u0026ndash;483. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.3171/jns.1986.65.4.0476\u003c/span\u003e\u003cspan address=\"10.3171/jns.1986.65.4.0476\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eSpetzler RF, Ponce FA (2011) A 3-tier classification of cerebral arteriovenous malformations. Clinical article. J Neurosurg 114:842\u0026ndash;849. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.3171/2010.8.Jns10663\u003c/span\u003e\u003cspan address=\"10.3171/2010.8.Jns10663\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003evan Beijnum J, van der Worp HB, Buis DR, Al-Shahi Salman R, Kappelle LJ, Rinkel GJ, van der Sprenkel JW, Vandertop WP, Algra A, Klijn CJ (2011) Treatment of brain arteriovenous malformations: a systematic review and meta-analysis. JAMA 306:2011\u0026ndash;2019. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1001/jama.2011.1632\u003c/span\u003e\u003cspan address=\"10.1001/jama.2011.1632\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"},{"header":"Tables","content":"\u003cp\u003e\u003cstrong\u003eTable 1.\u003c/strong\u003e \u003cem\u003eSummary of interobserver agreement in AVM grading.\u003c/em\u003e\u003cbr\u003e\u0026nbsp;Intraclass correlation (ICC) and Fleiss\u0026rsquo; \u0026kappa; are reported for overall Spetzler\u0026ndash;Martin (SM) and Spetzler\u0026ndash;Ponce (SP) classifications. Component-level agreement is shown for eloquence, venous drainage, and nidus size (\u0026kappa; values), while nidus morphology is expressed as the percentage of cases with complete concordance. The data highlight moderate numerical agreement for SM scores but only fair categorical agreement, with eloquence and nidus morphology (compact vs diffuse) emerging as the least reproducible parameters.\u003c/p\u003e\n\u003ctable border=\"0\" cellspacing=\"3\" cellpadding=\"0\"\u003e\n \u003cthead\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003eDomain\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003eMetric / Agreement measure\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003eResult\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/thead\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003eSpetzler\u0026ndash;Martin (SM)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eIntraclass correlation (ICC)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.72\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eFleiss\u0026rsquo; \u0026kappa; (categorical)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003e0.04\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003eSpetzler\u0026ndash;Ponce (SP)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eFleiss\u0026rsquo; \u0026kappa;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.49\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003eComponent-level\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eEloquence (\u0026kappa;)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.29\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eVenous drainage (\u0026kappa;)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.58\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eNidus size (\u0026kappa;)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.41\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eNidus morphology (\u0026kappa;)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u0026ndash;0.15\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n\u003c/table\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":true,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"","lastPublishedDoi":"10.21203/rs.3.rs-7525175/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-7525175/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003ch2\u003ePurpose\u003c/h2\u003e\u003cp\u003eTo quantify interobserver agreement in Spetzler\u0026ndash;Martin (SM) and Spetzler\u0026ndash;Ponce (SP) grading of pediatric brain arteriovenous malformations (AVMs), locate sources of variability, and test a composite Disagreement Index (DI).\u003c/p\u003e\u003ch2\u003eMethods\u003c/h2\u003e\u003cp\u003eForty-five consecutive pediatric AVMs were independently graded by three neurosurgical residents without prior calibration. SM components (eloquence, venous drainage, nidus size) and SP class were assigned; nidus morphology (compact vs diffuse) was scored by two raters. Agreement was estimated with Fleiss\u0026rsquo;/Cohen\u0026rsquo;s κ and ICC(2,1); dispersion with across-rater standard deviation and Shannon entropy. Borderline cases were prespecified (SM 2\u0026ndash;3; SP class transitions). DI combined entropy, SM-score dispersion, and component mismatch.\u003c/p\u003e\u003ch2\u003eResults\u003c/h2\u003e\u003cp\u003e SM scores showed moderate numerical agreement (ICC 0.72) but minimal categorical concordance (Fleiss\u0026rsquo; κ 0.04). SP improved overall agreement (Fleiss\u0026rsquo; κ 0.49), yet 17/45 (37.8%) cases crossed SP boundaries. Component reliability differed: eloquence κ 0.29, nidus size κ 0.41, venous drainage κ 0.58. Nidus morphology showed low reproducibility between two raters (\u0026asymp;\u0026thinsp;49% agreement; Cohen\u0026rsquo;s κ\u0026thinsp;\u0026minus;\u0026thinsp;0.15). Ten cases spanned the SM 2\u0026ndash;3 threshold. DI ranged 0.15\u0026ndash;1.00 (median 0.46) and isolated a small subset of highly discordant cases; eloquence was the primary driver in 8/10.\u003c/p\u003e\u003ch2\u003eConclusions\u003c/h2\u003e\u003cp\u003eInterobserver variability concentrates at decision thresholds and is driven chiefly by how eloquence is interpreted. Standardized definitions, reporting measured nidus dimensions with SM bins, and routine lesion-to-eloquence distance may stabilize grading. DI can flag \u0026ldquo;teaching\u0026rdquo; cases and support calibration over time.\u003c/p\u003e","manuscriptTitle":"Reproducibility of AVM grading in clinical practice: A study of interobserver variability","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2025-10-27 15:22:49","doi":"10.21203/rs.3.rs-7525175/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"dcb66c61-4881-4507-9e93-6356bd2cba85","owner":[],"postedDate":"October 27th, 2025","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[],"tags":[],"updatedAt":"2026-01-26T16:01:50+00:00","versionOfRecord":{"articleIdentity":"rs-7525175","link":"https://doi.org/10.1007/s00234-025-03902-9","journal":{"identity":"neuroradiology","isVorOnly":false,"title":"Neuroradiology"},"publishedOn":"2026-01-23 15:57:17","publishedOnDateReadable":"January 23rd, 2026"},"versionCreatedAt":"2025-10-27 15:22:49","video":"","vorDoi":"10.1007/s00234-025-03902-9","vorDoiUrl":"https://doi.org/10.1007/s00234-025-03902-9","workflowStages":[]},"version":"v1","identity":"rs-7525175","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-7525175","identity":"rs-7525175","version":["v1"]},"buildId":"8U1c8b4HqxoKbykW_rLl7","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.