Artificial Intelligence Model’s assessment of intra-sample heterogeneity of HER2 IHC in breast cancer is related to interobserver and intraobserver agreement among pathologists

preprint OA: gold CC-BY-4.0

Abstract

Abstract Background Accurate assessment of HER2 status in breast cancer has been critical for guiding therapy and has become even more important with the emergence of antibody-drug conjugates, now also indicated in HER2-low tumors. However, inter- and intraobserver variability limits the reproducibility of HER2 IHC scoring among pathologists, particularly at the lower end of the expression spectrum. Artificial intelligence (AI) models offer potential to standardize and improve diagnostic accuracy and bring new insights into current practices shortcomings. Methods We conducted a study recruiting generalist and specialist pathologists from Rede D’Or centers across Brazil to assess digitized HER2 IHC whole slide images. The same images were presented for the pathologists with the interval of one month and to the AIM-HER2 (PathAI ®, Boston, MA) AI model. Intra- and interobserver agreement, as well as concordance with AI, were measured across 126 breast cancer samples. The association between sample features and agreement metrics was also analyzed using AI spatial breakdown data. Results Among 34 pathologists, median intraobserver concordance was 67.68% and median concordance with AI was 60.8%. Median interobserver agreement was 67.65%, with high agreement (> 85%) in 25.4% of samples. Significant positive correlations were observed among all agreement metrics. Samples with lower intra-sample heterogeneity, as determined by AI spatial breakdown scores, were associated with higher agreement levels. Conclusions Our findings highlight significant variability in HER2 IHC scoring among pathologists. AI models such as AIM-HER2 can be used to increase reproducibility of analysis and also to indicate the most difficult samples that can result in diagnostic discordance. We found that AI assessed intra-sample heterogeneity is correlated with a lower agreement rate among pathologists.
Full text 81,745 characters · extracted from preprint-html · click to expand
Artificial Intelligence Model’s assessment of intra-sample heterogeneity of HER2 IHC in breast cancer is related to interobserver and intraobserver agreement among pathologists | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Artificial Intelligence Model’s assessment of intra-sample heterogeneity of HER2 IHC in breast cancer is related to interobserver and intraobserver agreement among pathologists Pedro Simonis Seabra Martins Ferrari, Mariana Pettacia de Macedo, and 29 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-9023658/v1 This work is licensed under a CC BY 4.0 License Status: Under Revision Version 1 posted 17 You are reading this latest preprint version Abstract Background Accurate assessment of HER2 status in breast cancer has been critical for guiding therapy and has become even more important with the emergence of antibody-drug conjugates, now also indicated in HER2-low tumors. However, inter- and intraobserver variability limits the reproducibility of HER2 IHC scoring among pathologists, particularly at the lower end of the expression spectrum. Artificial intelligence (AI) models offer potential to standardize and improve diagnostic accuracy and bring new insights into current practices shortcomings. Methods We conducted a study recruiting generalist and specialist pathologists from Rede D’Or centers across Brazil to assess digitized HER2 IHC whole slide images. The same images were presented for the pathologists with the interval of one month and to the AIM-HER2 (PathAI ®, Boston, MA) AI model. Intra- and interobserver agreement, as well as concordance with AI, were measured across 126 breast cancer samples. The association between sample features and agreement metrics was also analyzed using AI spatial breakdown data. Results Among 34 pathologists, median intraobserver concordance was 67.68% and median concordance with AI was 60.8%. Median interobserver agreement was 67.65%, with high agreement (> 85%) in 25.4% of samples. Significant positive correlations were observed among all agreement metrics. Samples with lower intra-sample heterogeneity, as determined by AI spatial breakdown scores, were associated with higher agreement levels. Conclusions Our findings highlight significant variability in HER2 IHC scoring among pathologists. AI models such as AIM-HER2 can be used to increase reproducibility of analysis and also to indicate the most difficult samples that can result in diagnostic discordance. We found that AI assessed intra-sample heterogeneity is correlated with a lower agreement rate among pathologists. Digital Pathology Artificial Intelligence (AI) Computer Vision HER2 Immunohistochemistry interobserver agreement Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Figure 6 BACKGROUND Breast cancer is the most common cancer in women, corresponding to over 2.3 million cases and 670,000 deaths worldwide in 2022[ 1 ]. In the United States, approximately 1 in every 8 women will be diagnosed with breast cancer during their lifetime[ 2 ]. Breast cancer is a heterogeneous group of diseases when it comes to responses to treatment and clinical outcomes. To address this variability, treatment decision-making relies on phenotypic classification. Current mainstay methods of classification use the expression of four biomarkers by immunohistochemistry (IHC): the proliferation marker Ki-67, estrogen and progesterone receptors (ER and PR receptors, or HR for both), and human epidermal growth factor receptor 2 (HER2). This classification differentiates tumors in 4 major groups: Luminal A tumors are HR+, Ki-67-low and HER2 negative, Luminal B tumors are ER+, PR+/PR-, and Ki-67-high and/or HER2 positive, HER2-enriched tumors are HR- and HER2 positive, and triple negative tumors are negative to both hormone receptors and to HER2[ 3 ]. Therapeutic implications of these subtypes include the use of endocrine therapy for luminal tumors and the use of HER2-targeted therapy for HER2-enriched and HER2 positive luminal B tumors[ 4 ]. It was in the context of treating HER2 positive tumors that current IHC essays for this marker were designed: the main objective of pathologists when assessing a sample was to differentiate between HER2 positive (3+) tumors, very likely to be amplified and therefore responsive to treatment, and HER2 negative tumors (0 or 1+), very unlikely to be amplified[ 5 ]. For equivocal (2+) cases, there is the possibility of confirmation of amplification through in situ hybridization[ 6 ]. Scoring practices rely on guidelines from the American Society of Clinical Oncology (ASCO) and the College of American Pathologists (CAP) for standardization. In the most recent guidelines, HER2 0 is defined as a tumor without any HER2 staining or with faint or incomplete staining in 10% of cells or less, HER2 1 + is defined as a tumor with faint to incomplete staining in over 10% of cells, HER2 2 + is defined as a tumor with weak to moderate complete staining in over 10% of cells and HER2 3 + is defined as a tumor with complete and intense circumferential staining in over 10% of cells[ 6 ]. Current first-line treatment for HER2 positive advanced breast cancer patients relies on dual blockade of HER2, with the addition of Pertuzumab to the previous standard of care therapy with Trastuzumab, which has shown remarkable improvement of the effectiveness of treating this disease[ 3 , 7 ]. However, demand for new treatments still exists, especially for patients refractory to present HER2 blockers, and it has driven the development of new therapeutic options that offer a different approach to HER2-targeted treatment. Antibody-drug conjugates (ADCs) targeting HER2, such as Trastuzumab emtansine (TDM1) and Trastuzumab deruxtecan (T-DXd), were initially developed as second-line treatment for HER2 positive tumors that became resistant to first-line dual blockade[ 8 ]. These therapies differ from traditional receptor blockade therapies because they do not rely solely on inhibition of receptor activity for their therapeutic effects; instead, the antibody is also used as a “guide” to deliver the payload directly to tumor cells, allowing the delivery of more intensive chemotherapy with higher tolerability[ 9 ]. Due to this mechanism, second generation ADCs such as T-DXd were found to also be effective in HER2 expressing tumors of other origins, such as gastric tumors[ 10 ], and in breast tumors in which HER2 is not a driving oncogene, including HER2 2 + tumors that were negative for amplification of HER2 in ISH; HER2 1 + tumors, and even HER2 0 tumors that had visible staining that did not reach 10% of cells[ 11 , 12 ]. These new developments are driving a change in the nomenclature of IHC scoring of breast cancer. While originally all HER2 1 + and 0 tumors were deemed negative and put in the same category, now that these distinctions are clinically relevant, there’s a push to redefine the lower end of the HER2 scoring spectrum. In the newest nomenclatures, HER2 1 + is now called HER2 low, and HER2 0 is divided between HER2 ultra low when there’s any visible expression of HER2, and HER2 null when the sample is devoid of any HER2 staining on IHC[ 13 ]. To follow these new therapeutic developments, there is a growing demand for more precision and granularity in pathologists' scoring of HER2. This has raised concerns among pathologists and clinicians about the limitations of our current assay assessments. Multiple recent studies show low consistency and reproducibility of current scoring practices amongst pathologists, most notably in the lower end of HER2 expression. Fernandez et al.[ 14 ] found that for their cohort of 18 pathologists, only 26% of HER2 0 samples and 58% of HER2 + 3 samples reached over 90% agreement. Similarly, in the Baez Navarro et al.[ 15 ] study’s 16 pathologist cohort, agreement of 87.5% agreement was found in only 30.4% of cases. Another large cohort from Sakatani et al.[ 16 ] showed an average agreement among 144 pathologists from different skill levels of 63.4% with specialist consensus (which was not unanimous in several cases). In this context, there’s an evident need for new approaches to the scoring of HER2 IHC in breast cancer. A promising field is the application of AI models to whole-slide images to help pathologists in their decision-making process. This approach is currently being explored and applied to several diagnostic and prognostic challenges in pathology, with uses such as PDL-1 IHC scoring of melanoma and classification of non-small cell lung cancer[ 17 , 18 ], and in other areas of medicine, such as the classification of metabolic-associated steatohepatitis[ 19 ]. The advantages of applying AI-based methods can potentially be two-fold. First, there’s the possibility of improved reproducibility and standardization of assays that can be gained by the implementation of such models in clinical practice, second is the possibility of extracting data from these models that can provide insight into shortcomings of our current practices, providing actionable data for improvement, both during their use and even prior to the definite implementation of these tools in clinic. OBJECTIVES Measure AI, intra- and interobserver agreement metrics in a cohort of Brazilian pathologists’ scoring of HER2 IHC of breast cancer samples; Using the AI model, assess sample features that correlate with agreement metrics METHODS Study subject and materials To explore the issue of inconsistent scoring in HER2 IHC assessment of breast cancer samples, we invited pathologists from Rede D’Or centers in all regions of Brazil to take online surveys in the assessment of digitized HER2 IHC whole slide images, reassessed these same digitized images with the AIM-HER2 (PathAI ®, Boston, MA) AI model, and compared all results. The surveys were developed specifically for the purposes of this study. They consisted of on-line forms with links to whole slide images, each link being followed by a multiple choice question with all HER2 IHC scoring options (0 through 3+) to be selected for each sample. Samples included surgical and core biopsy specimens from our service’s diagnostic routine, anonymized and digitized for the purposes of this study. Both generalist pathologists and breast specialists were included. Study design Our study design included an initial 10-slide survey, which we named a pre-test, followed by an on-line standardization workshop and two 100-slide surveys separated by a month (test and post-test) for a total of 210 questions. All post-test questions were also assessed by the AIM-HER2 model (Fig. 1 ). Slides were scanned in the Aperio AT2 (Leica ® Biosystems, Nussloch, Germany) and made available through the Aperio ImageScope visualization tool. Each survey contained repeated rescanned samples with new identifications from the previous assessments to also measure intraobserver variability within our pathologist group. Two samples were present in all three assessments, eight samples were present in the pre-test and test assessments, and 72 samples were present in the test and post-test assessments. There were 18 unique samples for the test and 26 unique samples for the post-test. This totals of 126 samples (Fig. 2 ). Ethical considerations As an observational retrospective study with no direct patient impact, there were no significant ethical challenges for the execution of this project. All samples were anonymized for the protection of patient privacy. The project was assessed and approved by our institution’s assigned ethics committee. AI-powered HER2 scoring AIM-HER2 is an AI model designed for scoring HER2 IHC samples in breast cancer[ 20 ]. At this moment, the use of AIM-HER2 is restricted for research and should not be used in diagnostic procedures. The model uses an additive multiple instance learning (AMIL) architecture[ 21 ]. Training data for this model is comprised of whole-slide images with overall HER2 scoring; however, the scoring is performed through image segmentation and classification of individual regions. Each of these regions is classified independently, and the overall score is then defined by the composition of all subsets[ 21 ]. AIM-HER2 outputs include a slide-level score prediction, as well as the relative tissue areas corresponding to each ASCO/CAP HER2 scoring level. This model relies on the maxim that for an overall sample to be classified as a determined score, at least one of the subset images must represent this score. Due to this architecture, the model provides spatial breakdown analysis (or “scoring heatmaps”), predicting the likelihood of regions of the image to correspond to each score. This scoring system is not cell based, as the ASCO/CAP guidelines suggest, but instead uses overarching characteristics of all cells in each subset area to give a score prediction. Statistical analysis Statistical analysis and data visualization was performed using the R programming language (version 4.4.2), with the RStudio software (Version 2025.05.0 + 496). RESULTS Samples: A total of 34 pathologists participated in the test and pre-test surveys and four of them dropped out prior to completing the final survey, resulting in a total of 6,740 sample responses. Data from pathologists who did not complete the post-test survey were excluded from the intraobserver analysis but included in the interobserver and AI comparisons. Scoring: Out of the 126 samples included in the assessment, the majority of the 34 pathologists scored 35 samples as 0, 49 as 1+, 28 as 2 + and 14 as 3+. Samples scored in the extremes of the spectrum (0 and 3+) were more likely to reach a higher percentage of agreement among pathologists, whereas the intermediate scores (1+, and 2+), were more likely to have less agreement (Fig. 3 ). Out of 100 samples assessed by AI, 22 were scored as 0, 45 as 1+, 25 as 2+, and 8 as 3+. Agreement metrics: Measuring concordance using all answers from all assessments, median intraobserver concordance among pathologists was 67.68%, and median pathologist concordance with AI was 60.80%. Median interobserver agreement, assessed using one answer per sample for each pathologist, was 67.65% per sample. Agreement of over 85% was reached in 25.4% of samples (32/126). Median agreement with AI, using all responses, was 59.38% per sample. Figures 4 and 5 show a graphical representation of scoring for all samples, on aggregate for all pathologists (Fig. 4 ). Agreement metrics correlation: Significant positive correlations as measured by Pearson’s correlation coefficients were found between interobserver and AI agreement (r = 0.86, p < 2.2×10 − 16 ), intra- and interobserver agreement (r = 0.74, p = 1.19×10 − 15 ), and intraobserver and AI agreement (r = 0.68, p = 1.47×10 − 11 ) (Fig. 5 ). This corroborates the intuitive assumption that the difficulty in assessment is not equal among samples, with some samples being more difficult to assess than others, as measured by all correlation metrics. Intra-sample heterogeneity To assess the features in these samples that could explain this variability, we sought to look into the AIM-HER2 model’s spatial breakdown data. Samples in which the AI identified a higher proportion of the tumor area as matching the overall HER2 score (e.g., 90% of the tumor was 1 + and the final score was 1+) showed significantly higher interobserver (r = 0.66, p = 1.37x10 − 13 ) and AI agreement (r = 0.58, p = 2.57×10 − 10 )(Fig. 6 ). In contrast, neither the total tumor area nor the absolute area matching the overall score showed correlation with variability. These findings suggest that lower tumoral heterogeneity within a sample, regardless of the tumor area assessed, is associated with more reproducible scoring. DISCUSSION Our results in the evaluation of generalist and specialist pathologists’ assessments of HER2 IHC assessment in breast cancer showed very similar results to those found in the literature. We did find a lower percentage of high agreement samples, only 25.4% when using a cutoff of 85%, when compared to other studies over 30% of samples with high agreement even though they used higher cutoffs[ 14 , 15 ]. We believe this may be related to the larger overall pathologist cohort we gathered, which included mostly generalist pathologists. The attempt to find AI-based features that can explain the variability of results between samples makes our research different from previous work in this area. We were successful in finding that intrasample heterogeneity is a significant factor in describing the variability of pathologist responses, and that higher heterogeneity assessed by AI, defined for our study as a lower percentage of the total tumor area corresponding to the overall result for the sample, correlates to less agreement among pathologists and for a same pathologist in subsequent assessments. Some limitations of the study include the lack of standardization of pathologists’ workstation during assessments. We tried to partially address this issue by promoting a standardization workshop prior to the bulk of assessments, but pathologists were free to use whichever device they would like to assess the digitized samples. CONCLUSION As medicine becomes more precise and personalized, pathology increases its protagonism in the multidisciplinary scenario – especially in oncology. While new therapeutic targets are discovered and therapies evolve, there’s an equal demand for new ways to analyze tumor biology or new approaches to assess the currently used biomarkers in light of new clinical demands. To provide more precise and complete information faster is essential. The challenge of HER2 assessment is emblematic of this transformation and highlights the necessity of adaptability that the ever-evolving field of oncology requires. One of the ways to meet such a demand is through the use of novel computational based models of data assessment, especially for imaging analysis. For the specific issue of HER2 IHC scoring for breast cancer samples, our study shows that one of the potential pitfalls of scoring is in heterogeneous HER2 IHC slides. We believe this highlights the limitations of our current practices in scoring fringe cases, due to the challenges of establishing robust parameters for intermediate levels of staining and of visually defining “over 10% of cells” for each staining level. The latter of these criteria is overlooked by the AIM-HER2 model, as it scores areas of the slide, rather than individual cells. As an overall proof of concept, our study also demonstrates that AI can have an important role in the future of pathology – both as a mediator of diagnostic discussions as well as an indicator of critical points of analysis, helping pathologists focus on the diagnostic-defining regions of a slide. Declarations Funding and Disclosures: The project was entirely self-funded.PathAI provided us with the AIM-HER2 model and ongoing technical support to produce this research. They did not interfere with the design or execution of this project. Clinical trial number: not applicable. Author Contributions Statement: Conceptualization: All authors contributed to conceptualization and development of study methodology. All members of the Biomarkers in Breast Cancer Cooperative Study Group. contributed with sample analysis. Statistical analysis and original draft preparation was done primarily by PSSMF, writing, reviewing, and editing was done primarily by PSSMF, MPM, IWC and FAS. MPM, IWC and FAS facilitated sample digitalization and acquired technical support from Path AI. The project was supervised by FAS. ETHICS APPROVAL AND CONSENT TO PARTICIPATE Our project was approved by the Hospital e Maternidade São Luiz – HMSL ethics committee. The need for informed consent terms from participants was waived. Our project was done in accordance with the Declaration of Helsinki. The Brazilian unified identifying code for the ethics committee assessment (CAAE) for the project is: 94672525.2.0000.0087 Consent for publication : Not applicable. The project does not contain any identifying images or information for any patients. DATA AVAILABILITY Raw data generated from the results of the sample assessment surveys and survey images themselves will be made available upon reasonable request to the corresponding author MPM. References Kim J, Harper A, McCormack V, et al. Global patterns and trends in breast cancer incidence and mortality across 185 countries. Nat Med. 2025;31:1154–62. https://doi.org/10.1038/s41591-025-03502-3 . Sung H, Ferlay J, Siegel RL, et al. Global Cancer Statistics 2020: GLOBOCAN Estimates of Incidence and Mortality Worldwide for 36 Cancers in 185 Countries. CA Cancer J Clin. 2021;71:209–49. https://doi.org/10.3322/caac.21660 . Carvalho E, Canberk S, Schmitt F, Vale N. Molecular Subtypes and Mechanisms of Breast Cancer: Precision Medicine Approaches for Targeted Therapies. Cancers (Basel). 2025;17:1102. https://doi.org/10.3390/cancers17071102 . Corti C, Batra-Sharma H, Kelsten M, et al. Systemic Therapy in Breast Cancer. Am Soc Clin Oncol Educational Book. 2024;44. https://doi.org/10.1200/EDBK_432442 . Lehmann-Che J, Amira-Bouhidel F, Turpin E, et al. Immunohistochemical and molecular analyses of HER2 status in breast cancers are highly concordant and complementary approaches. Br J Cancer. 2011;104:1739–46. https://doi.org/10.1038/bjc.2011.135 . Wolff AC, Somerfield MR, Dowsett M, et al. Human Epidermal Growth Factor Receptor 2 Testing in Breast Cancer. Arch Pathol Lab Med. 2023;147:993–1000. https://doi.org/10.5858/arpa.2023-0950-SA . Swain SM, Baselga J, Kim S-B, et al. Pertuzumab, Trastuzumab, and Docetaxel in HER2-Positive Metastatic Breast Cancer. N Engl J Med. 2015;372:724–34. https://doi.org/10.1056/NEJMoa1413513 . Cortés J, Hurvitz SA, Im S-A, et al. Trastuzumab deruxtecan versus trastuzumab emtansine in HER2-positive metastatic breast cancer: long-term survival analysis of the DESTINY-Breast03 trial. Nat Med. 2024;30:2208–15. https://doi.org/10.1038/s41591-024-03021-7 . Fu Z, Li S, Han S, et al. Antibody drug conjugate: the biological missile for targeted cancer therapy. Signal Transduct Target Ther. 2022;7:93. https://doi.org/10.1038/s41392-022-00947-7 . Shitara K, Bang Y-J, Iwasa S, et al. Trastuzumab Deruxtecan in Previously Treated HER2-Positive Gastric Cancer. N Engl J Med. 2020;382:2419–30. https://doi.org/10.1056/NEJMoa2004413 . Modi S, Jacot W, Yamashita T, et al. Trastuzumab Deruxtecan in Previously Treated HER2-Low Advanced Breast Cancer. N Engl J Med. 2022;387:9–20. https://doi.org/10.1056/NEJMoa2203690 . Tang J, Li N, Yang L, et al. Trastuzumab deruxtecan (T-DXd) monotherapy in advanced HER2-positive or HER2-low/ultralow breast cancer: An updated pooled analysis of prospective clinical trials. J Clin Oncol. 2025. https://doi.org/10.1200/JCO.2025.43.16_suppl.e13103 . 43:. Franchina M, Pizzimenti C, Fiorentino V, et al. Low and Ultra-Low HER2 in Human Breast Cancer: An Effort to Define New Neoplastic Subtypes. Int J Mol Sci. 2023;24:12795. https://doi.org/10.3390/ijms241612795 . Fernandez AI, Liu M, Bellizzi A, et al. Examination of Low ERBB2 Protein Expression in Breast Cancer Tissue. JAMA Oncol. 2022;8:607. https://doi.org/10.1001/jamaoncol.2021.7239 . Baez-Navarro X, van Bockstal MR, Nawawi D, et al. Interobserver Variation in the Assessment of Immunohistochemistry Expression Levels in HER2-Negative Breast Cancer: Can We Improve the Identification of Low Levels of HER2 Expression by Adjusting the Criteria? An International Interobserver Study. Mod Pathol. 2023;36:100009. https://doi.org/10.1016/j.modpat.2022.100009 . Sakatani T, Tsuda H, Yoshida M, et al. Current status and challenges in HER2 IHC assessment: scoring survey results in Japan. Breast Cancer Res Treat. 2025;210:27–36. https://doi.org/10.1007/s10549-024-07532-2 . Koelzer VH, Gisler A, Hanhart JC, et al. Digital image analysis improves precision of PD-L1 scoring in cutaneous melanoma. Histopathology. 2018;73:397–406. https://doi.org/10.1111/his.13528 . Kludt C, Wang Y, Ahmad W, et al. Next-generation lung cancer pathology: Development and validation of diagnostic and prognostic algorithms. Cell Rep Med. 2024. https://doi.org/10.1016/j.xcrm.2024.101697 . Pulaski H, Harrison SA, Mehta SS, et al. Clinical validation of an AI-based pathology tool for scoring of metabolic dysfunction-associated steatohepatitis. Nat Med. 2024. https://doi.org/10.1038/s41591-024-03301-2 . Shamshoian J, Shanis Z, Cabeen R et al. (2025) Impact of a machine learning-powered algorithm on pathologist HER2 IHC scoring in breast cancer. Javed SA, Juyal D, Padigela H, et al. Additive MIL. Intrinsically Interpretable Multiple Instance Learning for Pathology; 2022. Additional Declarations No competing interests reported. Cite Share Download PDF Status: Under Revision Version 1 posted Editorial decision: Revision requested 19 Apr, 2026 Reviewers agreed at journal 19 Apr, 2026 Reviewers agreed at journal 19 Apr, 2026 Reviewers agreed at journal 18 Apr, 2026 Reviews received at journal 17 Apr, 2026 Reviewers agreed at journal 17 Apr, 2026 Reviewers agreed at journal 12 Apr, 2026 Reviewers agreed at journal 09 Apr, 2026 Reviews received at journal 07 Apr, 2026 Reviewers agreed at journal 04 Apr, 2026 Reviews received at journal 03 Apr, 2026 Reviewers agreed at journal 02 Apr, 2026 Reviewers invited by journal 02 Apr, 2026 Editor assigned by journal 26 Mar, 2026 Editor invited by journal 11 Mar, 2026 Submission checks completed at journal 11 Mar, 2026 First submitted to journal 11 Mar, 2026 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-9023658","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":619232061,"identity":"a65ae273-a76a-410c-93fe-08c1ff67487f","order_by":0,"name":"Pedro Simonis Seabra Martins Ferrari","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA6ElEQVRIiWNgGAWjYHACAyCWgzATKiQYG4jUYgzVcoZkLYxtDIS1mLM3b/vwgcFAznz24WcfHs6zkO2XPsD4uOIXbi2WPceKZ85gMDCWOZdmPCNxm4TxzL4EZsOzfXhcdSPHmJmH4U/iDB6g64BaEjecYWCTbOzBo+X+G5AWg/oZPOyfGRLnEKPlBg9YS4IEDw/QlgaoloYf+PySVsw4w8DAcAYPTzFDwjGgX3oYmw0bG3BrMWc/vJnhQ4WBvAQP+2bGHzV1sv08zAcfNvzB4zAkEgaAUQOMIAJaMAEeW0bBKBgFo2DEAQAu4Uh8XQfyvgAAAABJRU5ErkJggg==","orcid":"","institution":"D’Or Institute of Research \u0026 Education","correspondingAuthor":true,"prefix":"","firstName":"Pedro","middleName":"Simonis Seabra Martins","lastName":"Ferrari","suffix":""},{"id":619232062,"identity":"8ddeec2f-b1b6-4229-933d-b6fb8e58c991","order_by":1,"name":"Mariana Pettacia de Macedo","email":"","orcid":"","institution":"D’Or Institute of Research \u0026 Education","correspondingAuthor":false,"prefix":"","firstName":"Mariana","middleName":"Pettacia","lastName":"de Macedo","suffix":""},{"id":619232063,"identity":"f313e4b7-b0d9-43f4-8d73-8d39d59830a9","order_by":2,"name":"Isabela Werneck da Cunha","email":"","orcid":"","institution":"D’Or Institute of Research \u0026 Education","correspondingAuthor":false,"prefix":"","firstName":"Isabela","middleName":"Werneck da","lastName":"Cunha","suffix":""},{"id":619232064,"identity":"4ef2aa76-3946-426b-ab40-567d2a7219a1","order_by":3,"name":"Aline Alencar Giongo","email":"","orcid":"","institution":"Rede D'Or Hospital Network","correspondingAuthor":false,"prefix":"","firstName":"Aline","middleName":"Alencar","lastName":"Giongo","suffix":""},{"id":619232065,"identity":"2c37daf7-4b0a-4130-ad55-d7cda7e2c717","order_by":4,"name":"Aline Helen da Silva Camacho","email":"","orcid":"","institution":"Rede D'Or Hospital Network","correspondingAuthor":false,"prefix":"","firstName":"Aline","middleName":"Helen da Silva","lastName":"Camacho","suffix":""},{"id":619232066,"identity":"e1d66f48-1f36-4007-be5a-803ecc96803c","order_by":5,"name":"Amanda Goncalves Souza","email":"","orcid":"","institution":"Rede D'Or Hospital Network","correspondingAuthor":false,"prefix":"","firstName":"Amanda","middleName":"Goncalves","lastName":"Souza","suffix":""},{"id":619232067,"identity":"adffbfac-5bf7-4456-bcbf-f897753ab08b","order_by":6,"name":"Ana Lúcia de Brito Rodrigues","email":"","orcid":"","institution":"Rede D'Or Hospital Network","correspondingAuthor":false,"prefix":"","firstName":"Ana","middleName":"Lúcia de Brito","lastName":"Rodrigues","suffix":""},{"id":619232068,"identity":"f62a6c81-7dde-470f-8d61-ae8c0010177b","order_by":7,"name":"Ângela Gonçalves Karlinski","email":"","orcid":"","institution":"Rede D'Or Hospital Network","correspondingAuthor":false,"prefix":"","firstName":"Ângela","middleName":"Gonçalves","lastName":"Karlinski","suffix":""},{"id":619232069,"identity":"6ece42ac-14da-402d-b9d4-22d3dfbfbaf6","order_by":8,"name":"Ana Maria da Cunha Mercante","email":"","orcid":"","institution":"Rede D'Or Hospital Network","correspondingAuthor":false,"prefix":"","firstName":"Ana","middleName":"Maria da Cunha","lastName":"Mercante","suffix":""},{"id":619232070,"identity":"8124f8db-946e-4cb5-b637-6c710ce0fad0","order_by":9,"name":"Andrea Lima Cruz Monnerat","email":"","orcid":"","institution":"Rede D'Or Hospital Network","correspondingAuthor":false,"prefix":"","firstName":"Andrea","middleName":"Lima Cruz","lastName":"Monnerat","suffix":""},{"id":619232071,"identity":"8634adf6-93f5-4de8-9ce5-274e439b41da","order_by":10,"name":"Annelise Almeida Verdolin","email":"","orcid":"","institution":"Rede D'Or Hospital Network","correspondingAuthor":false,"prefix":"","firstName":"Annelise","middleName":"Almeida","lastName":"Verdolin","suffix":""},{"id":619232072,"identity":"eef39dc1-f957-4c9d-941c-ef7179c4731e","order_by":11,"name":"Antonio Alexandre Oliveira Lima Castro","email":"","orcid":"","institution":"Rede D'Or Hospital Network","correspondingAuthor":false,"prefix":"","firstName":"Antonio","middleName":"Alexandre Oliveira Lima","lastName":"Castro","suffix":""},{"id":619232073,"identity":"4f99c813-eb12-453f-aadc-42496a426526","order_by":12,"name":"Ariadne Callegario Nicolli","email":"","orcid":"","institution":"Rede D'Or Hospital Network","correspondingAuthor":false,"prefix":"","firstName":"Ariadne","middleName":"Callegario","lastName":"Nicolli","suffix":""},{"id":619232074,"identity":"ed9d840d-7d73-4237-98bc-29acb6cb6d94","order_by":13,"name":"Arthur Henrique Cunha Volpato","email":"","orcid":"","institution":"Rede D'Or Hospital Network","correspondingAuthor":false,"prefix":"","firstName":"Arthur","middleName":"Henrique Cunha","lastName":"Volpato","suffix":""},{"id":619232075,"identity":"035928c4-694a-4d4b-b610-b0bcd10e87d3","order_by":14,"name":"Bianca Nascimento de Aquino Garibaldi","email":"","orcid":"","institution":"Rede D'Or Hospital Network","correspondingAuthor":false,"prefix":"","firstName":"Bianca","middleName":"Nascimento de Aquino","lastName":"Garibaldi","suffix":""},{"id":619232076,"identity":"58ce470e-0ab7-44f4-a8ed-e106c461869a","order_by":15,"name":"Caio de Carvalho Santos","email":"","orcid":"","institution":"Rede D'Or Hospital Network","correspondingAuthor":false,"prefix":"","firstName":"Caio","middleName":"de Carvalho","lastName":"Santos","suffix":""},{"id":619232077,"identity":"617a6593-b344-4198-b97d-e9ff073907d0","order_by":16,"name":"Clovis Klock","email":"","orcid":"","institution":"D’Or Institute of Research \u0026 Education","correspondingAuthor":false,"prefix":"","firstName":"Clovis","middleName":"","lastName":"Klock","suffix":""},{"id":619232078,"identity":"5208377e-6254-4e7f-a571-7c9e1be9c701","order_by":17,"name":"Fabiano Pinto Saggioro","email":"","orcid":"","institution":"Rede D'Or Hospital Network","correspondingAuthor":false,"prefix":"","firstName":"Fabiano","middleName":"Pinto","lastName":"Saggioro","suffix":""},{"id":619232079,"identity":"cb854f98-87d4-498f-895b-70f8cfc2079d","order_by":18,"name":"Gerusa Tiburzio","email":"","orcid":"","institution":"Rede D'Or Hospital Network","correspondingAuthor":false,"prefix":"","firstName":"Gerusa","middleName":"","lastName":"Tiburzio","suffix":""},{"id":619232080,"identity":"1c1022d3-8897-4a7c-a35e-80314b8158a4","order_by":19,"name":"Geysa Bigi Maya Monteiro","email":"","orcid":"","institution":"Rede D'Or Hospital Network","correspondingAuthor":false,"prefix":"","firstName":"Geysa","middleName":"Bigi Maya","lastName":"Monteiro","suffix":""},{"id":619232081,"identity":"83eb0e6f-efa6-4a66-90dc-3ef5fbe9d13a","order_by":20,"name":"Giovanna Naves Teixeira Maximiano","email":"","orcid":"","institution":"Rede D'Or Hospital Network","correspondingAuthor":false,"prefix":"","firstName":"Giovanna","middleName":"Naves Teixeira","lastName":"Maximiano","suffix":""},{"id":619232082,"identity":"e0eb49e7-20cc-4d13-9cf9-13a79e309683","order_by":21,"name":"Giselle Maria Vignal","email":"","orcid":"","institution":"Rede D'Or Hospital Network","correspondingAuthor":false,"prefix":"","firstName":"Giselle","middleName":"Maria","lastName":"Vignal","suffix":""},{"id":619232083,"identity":"5e68f42e-4251-44fb-a2e8-67c460fa3915","order_by":22,"name":"Igor Campos da Silva","email":"","orcid":"","institution":"Rede D'Or Hospital Network","correspondingAuthor":false,"prefix":"","firstName":"Igor","middleName":"Campos da","lastName":"Silva","suffix":""},{"id":619232084,"identity":"7e0eb524-6b6c-49da-b963-becae46a9076","order_by":23,"name":"Janaina Nagel","email":"","orcid":"","institution":"Rede D'Or Hospital Network","correspondingAuthor":false,"prefix":"","firstName":"Janaina","middleName":"","lastName":"Nagel","suffix":""},{"id":619232085,"identity":"724a2401-d3f8-443b-aac4-1db6e80cb34f","order_by":24,"name":"Lidia M Magalhães C de Rezende","email":"","orcid":"","institution":"Rede D'Or Hospital Network","correspondingAuthor":false,"prefix":"","firstName":"Lidia","middleName":"M Magalhães C","lastName":"de Rezende","suffix":""},{"id":619232086,"identity":"524a8c48-0826-42a1-9e2c-247b3c69491a","order_by":25,"name":"Luciano Neder","email":"","orcid":"","institution":"Rede D'Or Hospital Network","correspondingAuthor":false,"prefix":"","firstName":"Luciano","middleName":"","lastName":"Neder","suffix":""},{"id":619232087,"identity":"0ab969c0-8dc6-470f-8f7f-4ae1bd35c752","order_by":26,"name":"Mariana Andozia Morini Matushita","email":"","orcid":"","institution":"Rede D'Or Hospital Network","correspondingAuthor":false,"prefix":"","firstName":"Mariana","middleName":"Andozia Morini","lastName":"Matushita","suffix":""},{"id":619232088,"identity":"da07dd0d-7f4e-4527-b135-76a65b61ac22","order_by":27,"name":"Nicolle Cavalcante Gaglionone","email":"","orcid":"","institution":"Rede D'Or Hospital Network","correspondingAuthor":false,"prefix":"","firstName":"Nicolle","middleName":"Cavalcante","lastName":"Gaglionone","suffix":""},{"id":619232091,"identity":"50211c8b-211b-40e8-9e39-7aaaaec28dcd","order_by":28,"name":"Patricia Sodre Dias Barreto","email":"","orcid":"","institution":"Rede D'Or Hospital Network","correspondingAuthor":false,"prefix":"","firstName":"Patricia","middleName":"Sodre Dias","lastName":"Barreto","suffix":""},{"id":619232093,"identity":"96ffb9d5-9d94-4bc6-ba9e-b00597251372","order_by":29,"name":"Roberto Bonfim Pimenta Peixoto","email":"","orcid":"","institution":"Rede D'Or Hospital Network","correspondingAuthor":false,"prefix":"","firstName":"Roberto","middleName":"Bonfim Pimenta","lastName":"Peixoto","suffix":""},{"id":619232094,"identity":"0f467e63-c14e-40d7-8a45-7e4c09dd4394","order_by":30,"name":"Fernando Augusto Soares","email":"","orcid":"","institution":"D’Or Institute of Research \u0026 Education","correspondingAuthor":false,"prefix":"","firstName":"Fernando","middleName":"Augusto","lastName":"Soares","suffix":""}],"badges":[],"createdAt":"2026-03-03 20:53:15","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-9023658/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-9023658/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":106536208,"identity":"20f34efa-a1e7-4d1d-8e7b-38be238c0c25","added_by":"auto","created_at":"2026-04-09 15:11:30","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":52406,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eSurveys structure and timeline\u003c/strong\u003e\u003c/p\u003e","description":"","filename":"image1.png","url":"https://assets-eu.researchsquare.com/files/rs-9023658/v1/b20560f783754a039d4843e6.png"},{"id":106536262,"identity":"830b2eef-75ef-4aca-8b0a-5094cf0cf8bc","added_by":"auto","created_at":"2026-04-09 15:11:50","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":85162,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eSample distribution\u003c/strong\u003e\u003c/p\u003e","description":"","filename":"image2.png","url":"https://assets-eu.researchsquare.com/files/rs-9023658/v1/676cac0a11df80f77ebce9b9.png"},{"id":106536254,"identity":"aedd8a2a-6caa-4287-b0c5-690ad08d24b7","added_by":"auto","created_at":"2026-04-09 15:11:44","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":72761,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003einterobserver agreement\u003c/strong\u003e\u003c/p\u003e","description":"","filename":"image3.png","url":"https://assets-eu.researchsquare.com/files/rs-9023658/v1/9a36ef3a678ae1700129c128.png"},{"id":106536260,"identity":"c86b89db-45fe-4b76-a1f3-fce9bcd73e4d","added_by":"auto","created_at":"2026-04-09 15:11:50","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":131295,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eOverall sample and survey-wise answers\u003c/strong\u003e\u003c/p\u003e","description":"","filename":"image4.png","url":"https://assets-eu.researchsquare.com/files/rs-9023658/v1/b63b90633edcdc81fef9cacb.png"},{"id":106536242,"identity":"4ad92f74-40a4-4cd4-aed0-b5919cfb1dee","added_by":"auto","created_at":"2026-04-09 15:11:41","extension":"png","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":103296,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eCorrelation between agreement metrics\u003c/strong\u003e\u003c/p\u003e","description":"","filename":"image5.png","url":"https://assets-eu.researchsquare.com/files/rs-9023658/v1/98a28fd30432ebb4af3d98c4.png"},{"id":106536317,"identity":"48e2a20a-4a24-41dd-961f-018e12e78ad3","added_by":"auto","created_at":"2026-04-09 15:12:02","extension":"png","order_by":6,"title":"Figure 6","display":"","copyAsset":false,"role":"figure","size":113558,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eCorrelation between heterogeneity and interobserver agreement\u003c/strong\u003e\u003c/p\u003e","description":"","filename":"image6.png","url":"https://assets-eu.researchsquare.com/files/rs-9023658/v1/4724aab33486b9fa79fdc06a.png"},{"id":106994002,"identity":"c0876096-5cf4-4062-a044-f46d0658b356","added_by":"auto","created_at":"2026-04-15 15:01:58","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":1467884,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-9023658/v1/ef8b8d80-75d9-4b55-9978-d90287ae854f.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"Artificial Intelligence Model’s assessment of intra-sample heterogeneity of HER2 IHC in breast cancer is related to interobserver and intraobserver agreement among pathologists","fulltext":[{"header":"BACKGROUND","content":"\u003cp\u003eBreast cancer is the most common cancer in women, corresponding to over 2.3\u0026nbsp;million cases and 670,000 deaths worldwide in 2022[\u003cspan class=\"CitationRef\"\u003e1\u003c/span\u003e]. In the United States, approximately 1 in every 8 women will be diagnosed with breast cancer during their lifetime[\u003cspan class=\"CitationRef\"\u003e2\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eBreast cancer is a heterogeneous group of diseases when it comes to responses to treatment and clinical outcomes. To address this variability, treatment decision-making relies on phenotypic classification. Current mainstay methods of classification use the expression of four biomarkers by immunohistochemistry (IHC): the proliferation marker Ki-67, estrogen and progesterone receptors (ER and PR receptors, or HR for both), and human epidermal growth factor receptor 2 (HER2). This classification differentiates tumors in 4 major groups: Luminal A tumors are HR+, Ki-67-low and HER2 negative, Luminal B tumors are ER+, PR+/PR-, and Ki-67-high and/or HER2 positive, HER2-enriched tumors are HR- and HER2 positive, and triple negative tumors are negative to both hormone receptors and to HER2[\u003cspan class=\"CitationRef\"\u003e3\u003c/span\u003e]. Therapeutic implications of these subtypes include the use of endocrine therapy for luminal tumors and the use of HER2-targeted therapy for HER2-enriched and HER2 positive luminal B tumors[\u003cspan class=\"CitationRef\"\u003e4\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eIt was in the context of treating HER2 positive tumors that current IHC essays for this marker were designed: the main objective of pathologists when assessing a sample was to differentiate between HER2 positive (3+) tumors, very likely to be amplified and therefore responsive to treatment, and HER2 negative tumors (0 or 1+), very unlikely to be amplified[\u003cspan class=\"CitationRef\"\u003e5\u003c/span\u003e]. For equivocal (2+) cases, there is the possibility of confirmation of amplification through in situ hybridization[\u003cspan class=\"CitationRef\"\u003e6\u003c/span\u003e]. Scoring practices rely on guidelines from the American Society of Clinical Oncology (ASCO) and the College of American Pathologists (CAP) for standardization. In the most recent guidelines, HER2 0 is defined as a tumor without any HER2 staining or with faint or incomplete staining in 10% of cells or less, HER2 1 + is defined as a tumor with faint to incomplete staining in over 10% of cells, HER2 2 + is defined as a tumor with weak to moderate complete staining in over 10% of cells and HER2 3 + is defined as a tumor with complete and intense circumferential staining in over 10% of cells[\u003cspan class=\"CitationRef\"\u003e6\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eCurrent first-line treatment for HER2 positive advanced breast cancer patients relies on dual blockade of HER2, with the addition of Pertuzumab to the previous standard of care therapy with Trastuzumab, which has shown remarkable improvement of the effectiveness of treating this disease[\u003cspan class=\"CitationRef\"\u003e3\u003c/span\u003e, \u003cspan class=\"CitationRef\"\u003e7\u003c/span\u003e]. However, demand for new treatments still exists, especially for patients refractory to present HER2 blockers, and it has driven the development of new therapeutic options that offer a different approach to HER2-targeted treatment.\u003c/p\u003e \u003cp\u003eAntibody-drug conjugates (ADCs) targeting HER2, such as Trastuzumab emtansine (TDM1) and Trastuzumab deruxtecan (T-DXd), were initially developed as second-line treatment for HER2 positive tumors that became resistant to first-line dual blockade[\u003cspan class=\"CitationRef\"\u003e8\u003c/span\u003e]. These therapies differ from traditional receptor blockade therapies because they do not rely solely on inhibition of receptor activity for their therapeutic effects; instead, the antibody is also used as a “guide” to deliver the payload directly to tumor cells, allowing the delivery of more intensive chemotherapy with higher tolerability[\u003cspan class=\"CitationRef\"\u003e9\u003c/span\u003e]. Due to this mechanism, second generation ADCs such as T-DXd were found to also be effective in HER2 expressing tumors of other origins, such as gastric tumors[\u003cspan class=\"CitationRef\"\u003e10\u003c/span\u003e], and in breast tumors in which HER2 is not a driving oncogene, including HER2 2 + tumors that were negative for amplification of HER2 in ISH; HER2 1 + tumors, and even HER2 0 tumors that had visible staining that did not reach 10% of cells[\u003cspan class=\"CitationRef\"\u003e11\u003c/span\u003e, \u003cspan class=\"CitationRef\"\u003e12\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eThese new developments are driving a change in the nomenclature of IHC scoring of breast cancer. While originally all HER2 1 + and 0 tumors were deemed negative and put in the same category, now that these distinctions are clinically relevant, there’s a push to redefine the lower end of the HER2 scoring spectrum. In the newest nomenclatures, HER2 1 + is now called HER2 low, and HER2 0 is divided between HER2 ultra low when there’s any visible expression of HER2, and HER2 null when the sample is devoid of any HER2 staining on IHC[\u003cspan class=\"CitationRef\"\u003e13\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eTo follow these new therapeutic developments, there is a growing demand for more precision and granularity in pathologists' scoring of HER2. This has raised concerns among pathologists and clinicians about the limitations of our current assay assessments.\u003c/p\u003e \u003cp\u003eMultiple recent studies show low consistency and reproducibility of current scoring practices amongst pathologists, most notably in the lower end of HER2 expression.\u003c/p\u003e \u003cp\u003eFernandez et al.[\u003cspan class=\"CitationRef\"\u003e14\u003c/span\u003e] found that for their cohort of 18 pathologists, only 26% of HER2 0 samples and 58% of HER2 + 3 samples reached over 90% agreement. Similarly, in the Baez Navarro et al.[\u003cspan class=\"CitationRef\"\u003e15\u003c/span\u003e] study’s 16 pathologist cohort, agreement of 87.5% agreement was found in only 30.4% of cases. Another large cohort from Sakatani et al.[\u003cspan class=\"CitationRef\"\u003e16\u003c/span\u003e] showed an average agreement among 144 pathologists from different skill levels of 63.4% with specialist consensus (which was not unanimous in several cases).\u003c/p\u003e \u003cp\u003eIn this context, there’s an evident need for new approaches to the scoring of HER2 IHC in breast cancer. A promising field is the application of AI models to whole-slide images to help pathologists in their decision-making process. This approach is currently being explored and applied to several diagnostic and prognostic challenges in pathology, with uses such as PDL-1 IHC scoring of melanoma and classification of non-small cell lung cancer[\u003cspan class=\"CitationRef\"\u003e17\u003c/span\u003e, \u003cspan class=\"CitationRef\"\u003e18\u003c/span\u003e], and in other areas of medicine, such as the classification of metabolic-associated steatohepatitis[\u003cspan class=\"CitationRef\"\u003e19\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eThe advantages of applying AI-based methods can potentially be two-fold. First, there’s the possibility of improved reproducibility and standardization of assays that can be gained by the implementation of such models in clinical practice, second is the possibility of extracting data from these models that can provide insight into shortcomings of our current practices, providing actionable data for improvement, both during their use and even prior to the definite implementation of these tools in clinic.\u003c/p\u003e\n\u003ch3\u003eOBJECTIVES\u003c/h3\u003e\n\u003cp\u003e \u003c/p\u003e\u003cul\u003e \u003cli\u003e \u003cp\u003eMeasure AI, intra- and interobserver agreement metrics in a cohort of Brazilian pathologists’ scoring of HER2 IHC of breast cancer samples;\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eUsing the AI model, assess sample features that correlate with agreement metrics\u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003cp\u003e\u003c/p\u003e \n\n "},{"header":"METHODS","content":"\u003ch2\u003eStudy subject and materials\u003c/h2\u003e\u003cp\u003eTo explore the issue of inconsistent scoring in HER2 IHC assessment of breast cancer samples, we invited pathologists from Rede D’Or centers in all regions of Brazil to take online surveys in the assessment of digitized HER2 IHC whole slide images, reassessed these same digitized images with the AIM-HER2 (PathAI ®, Boston, MA) AI model, and compared all results.\u003c/p\u003e\u003cp\u003eThe surveys were developed specifically for the purposes of this study. They consisted of on-line forms with links to whole slide images, each link being followed by a multiple choice question with all HER2 IHC scoring options (0 through 3+) to be selected for each sample. Samples included surgical and core biopsy specimens from our service’s diagnostic routine, anonymized and digitized for the purposes of this study.\u003c/p\u003e\u003cp\u003eBoth generalist pathologists and breast specialists were included.\u003c/p\u003e\u003ch3\u003eStudy design\u003c/h3\u003e\u003cp\u003eOur study design included an initial 10-slide survey, which we named a pre-test, followed by an on-line standardization workshop and two 100-slide surveys separated by a month (test and post-test) for a total of 210 questions. All post-test questions were also assessed by the AIM-HER2 model (Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e1\u003c/span\u003e). Slides were scanned in the Aperio AT2 (Leica ® Biosystems, Nussloch, Germany) and made available through the Aperio ImageScope visualization tool.\u003c/p\u003e\u003cp\u003eEach survey contained repeated rescanned samples with new identifications from the previous assessments to also measure intraobserver variability within our pathologist group. Two samples were present in all three assessments, eight samples were present in the pre-test and test assessments, and 72 samples were present in the test and post-test assessments. There were 18 unique samples for the test and 26 unique samples for the post-test. This totals of 126 samples (Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e2\u003c/span\u003e).\u003c/p\u003e\u003ch3\u003eEthical considerations\u003c/h3\u003e\u003cp\u003eAs an observational retrospective study with no direct patient impact, there were no significant ethical challenges for the execution of this project. All samples were anonymized for the protection of patient privacy. The project was assessed and approved by our institution’s assigned ethics committee.\u003c/p\u003e\u003ch3\u003eAI-powered HER2 scoring\u003c/h3\u003e\u003cp\u003eAIM-HER2 is an AI model designed for scoring HER2 IHC samples in breast cancer[\u003cspan class=\"CitationRef\"\u003e20\u003c/span\u003e]. At this moment, the use of AIM-HER2 is restricted for research and should not be used in diagnostic procedures.\u003c/p\u003e\u003cp\u003eThe model uses an additive multiple instance learning (AMIL) architecture[\u003cspan class=\"CitationRef\"\u003e21\u003c/span\u003e]. Training data for this model is comprised of whole-slide images with overall HER2 scoring; however, the scoring is performed through image segmentation and classification of individual regions. Each of these regions is classified independently, and the overall score is then defined by the composition of all subsets[\u003cspan class=\"CitationRef\"\u003e21\u003c/span\u003e]. AIM-HER2 outputs include a slide-level score prediction, as well as the relative tissue areas corresponding to each ASCO/CAP HER2 scoring level.\u003c/p\u003e\u003cp\u003eThis model relies on the maxim that for an overall sample to be classified as a determined score, at least one of the subset images must represent this score. Due to this architecture, the model provides spatial breakdown analysis (or “scoring heatmaps”), predicting the likelihood of regions of the image to correspond to each score.\u003c/p\u003e\u003cp\u003e This scoring system is not cell based, as the ASCO/CAP guidelines suggest, but instead uses overarching characteristics of all cells in each subset area to give a score prediction.\u003c/p\u003e\u003ch2\u003eStatistical analysis\u003c/h2\u003e\u003cp\u003eStatistical analysis and data visualization was performed using the R programming language (version 4.4.2), with the RStudio software (Version 2025.05.0 + 496).\u003c/p\u003e"},{"header":"RESULTS","content":"\u003cdiv id=\"Sec10\" class=\"Section2\"\u003e \u003ch2\u003eSamples:\u003c/h2\u003e \u003cp\u003eA total of 34 pathologists participated in the test and pre-test surveys and four of them dropped out prior to completing the final survey, resulting in a total of 6,740 sample responses.\u003c/p\u003e \u003cp\u003eData from pathologists who did not complete the post-test survey were excluded from the intraobserver analysis but included in the interobserver and AI comparisons.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec11\" class=\"Section2\"\u003e \u003ch2\u003eScoring:\u003c/h2\u003e \u003cp\u003eOut of the 126 samples included in the assessment, the majority of the 34 pathologists scored 35 samples as 0, 49 as 1+, 28 as 2\u0026thinsp;+\u0026thinsp;and 14 as 3+. Samples scored in the extremes of the spectrum (0 and 3+) were more likely to reach a higher percentage of agreement among pathologists, whereas the intermediate scores (1+, and 2+), were more likely to have less agreement (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003e). Out of 100 samples assessed by AI, 22 were scored as 0, 45 as 1+, 25 as 2+, and 8 as 3+.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec12\" class=\"Section2\"\u003e \u003ch2\u003eAgreement metrics:\u003c/h2\u003e \u003cp\u003eMeasuring concordance using all answers from all assessments, median intraobserver concordance among pathologists was 67.68%, and median pathologist concordance with AI was 60.80%.\u003c/p\u003e \u003cp\u003eMedian interobserver agreement, assessed using one answer per sample for each pathologist, was 67.65% per sample. Agreement of over 85% was reached in 25.4% of samples (32/126).\u003c/p\u003e \u003cp\u003eMedian agreement with AI, using all responses, was 59.38% per sample.\u003c/p\u003e \u003cp\u003eFigures \u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003e and \u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003e show a graphical representation of scoring for all samples, on aggregate for all pathologists (Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003e).\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec13\" class=\"Section2\"\u003e \u003ch2\u003eAgreement metrics correlation:\u003c/h2\u003e \u003cp\u003eSignificant positive correlations as measured by Pearson\u0026rsquo;s correlation coefficients were found between interobserver and AI agreement (r\u0026thinsp;=\u0026thinsp;0.86, p\u0026thinsp;\u0026lt;\u0026thinsp;2.2\u0026times;10\u003csup\u003e\u0026minus;\u0026thinsp;16\u003c/sup\u003e), intra- and interobserver agreement (r\u0026thinsp;=\u0026thinsp;0.74, p\u0026thinsp;=\u0026thinsp;1.19\u0026times;10\u003csup\u003e\u0026minus;\u0026thinsp;15\u003c/sup\u003e), and intraobserver and AI agreement (r\u0026thinsp;=\u0026thinsp;0.68, p\u0026thinsp;=\u0026thinsp;1.47\u0026times;10\u003csup\u003e\u0026minus;\u0026thinsp;11\u003c/sup\u003e) (Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003e). This corroborates the intuitive assumption that the difficulty in assessment is not equal among samples, with some samples being more difficult to assess than others, as measured by all correlation metrics.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec14\" class=\"Section2\"\u003e \u003ch2\u003eIntra-sample heterogeneity\u003c/h2\u003e \u003cp\u003eTo assess the features in these samples that could explain this variability, we sought to look into the AIM-HER2 model\u0026rsquo;s spatial breakdown data.\u003c/p\u003e \u003cp\u003eSamples in which the AI identified a higher proportion of the tumor area as matching the overall HER2 score (e.g., 90% of the tumor was 1\u0026thinsp;+\u0026thinsp;and the final score was 1+) showed significantly higher interobserver (r\u0026thinsp;=\u0026thinsp;0.66, p\u0026thinsp;=\u0026thinsp;1.37x10\u003csup\u003e\u0026minus;\u0026thinsp;13\u003c/sup\u003e) and AI agreement (r\u0026thinsp;=\u0026thinsp;0.58, p\u0026thinsp;=\u0026thinsp;2.57\u0026times;10\u003csup\u003e\u0026minus;\u0026thinsp;10\u003c/sup\u003e)(Fig.\u0026nbsp;\u003cspan refid=\"Fig6\" class=\"InternalRef\"\u003e6\u003c/span\u003e). In contrast, neither the total tumor area nor the absolute area matching the overall score showed correlation with variability. These findings suggest that lower tumoral heterogeneity within a sample, regardless of the tumor area assessed, is associated with more reproducible scoring.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e"},{"header":"DISCUSSION","content":"\u003cp\u003eOur results in the evaluation of generalist and specialist pathologists\u0026rsquo; assessments of HER2 IHC assessment in breast cancer showed very similar results to those found in the literature. We did find a lower percentage of high agreement samples, only 25.4% when using a cutoff of 85%, when compared to other studies over 30% of samples with high agreement even though they used higher cutoffs[\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e, \u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e]. We believe this may be related to the larger overall pathologist cohort we gathered, which included mostly generalist pathologists.\u003c/p\u003e \u003cp\u003eThe attempt to find AI-based features that can explain the variability of results between samples makes our research different from previous work in this area. We were successful in finding that intrasample heterogeneity is a significant factor in describing the variability of pathologist responses, and that higher heterogeneity assessed by AI, defined for our study as a lower percentage of the total tumor area corresponding to the overall result for the sample, correlates to less agreement among pathologists and for a same pathologist in subsequent assessments.\u003c/p\u003e \u003cp\u003eSome limitations of the study include the lack of standardization of pathologists\u0026rsquo; workstation during assessments. We tried to partially address this issue by promoting a standardization workshop prior to the bulk of assessments, but pathologists were free to use whichever device they would like to assess the digitized samples.\u003c/p\u003e"},{"header":"CONCLUSION","content":"\u003cp\u003eAs medicine becomes more precise and personalized, pathology increases its protagonism in the multidisciplinary scenario \u0026ndash; especially in oncology. While new therapeutic targets are discovered and therapies evolve, there\u0026rsquo;s an equal demand for new ways to analyze tumor biology or new approaches to assess the currently used biomarkers in light of new clinical demands. To provide more precise and complete information faster is essential. The challenge of HER2 assessment is emblematic of this transformation and highlights the necessity of adaptability that the ever-evolving field of oncology requires.\u003c/p\u003e \u003cp\u003eOne of the ways to meet such a demand is through the use of novel computational based models of data assessment, especially for imaging analysis.\u003c/p\u003e \u003cp\u003eFor the specific issue of HER2 IHC scoring for breast cancer samples, our study shows that one of the potential pitfalls of scoring is in heterogeneous HER2 IHC slides. We believe this highlights the limitations of our current practices in scoring fringe cases, due to the challenges of establishing robust parameters for intermediate levels of staining and of visually defining \u0026ldquo;over 10% of cells\u0026rdquo; for each staining level. The latter of these criteria is overlooked by the AIM-HER2 model, as it scores areas of the slide, rather than individual cells.\u003c/p\u003e \u003cp\u003eAs an overall proof of concept, our study also demonstrates that AI can have an important role in the future of pathology \u0026ndash; both as a mediator of diagnostic discussions as well as an indicator of critical points of analysis, helping pathologists focus on the diagnostic-defining regions of a slide.\u003c/p\u003e "},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eFunding and Disclosures: \u003c/strong\u003eThe project was entirely self-funded.PathAI provided us with the AIM-HER2 model and ongoing technical support to produce this research. They did not interfere with the design or execution of this project.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eClinical trial number: \u003c/strong\u003enot applicable.\u003c/p\u003e\n\n\u003cp\u003e\u003cstrong\u003eAuthor Contributions Statement: \u003c/strong\u003eConceptualization: All authors contributed to conceptualization and development of study methodology. All members of the Biomarkers in Breast Cancer Cooperative Study Group. contributed with sample analysis. Statistical analysis and original draft preparation was done primarily by PSSMF, writing, reviewing, and editing was done primarily by PSSMF, MPM, IWC and FAS. MPM, IWC and FAS facilitated sample digitalization and acquired technical support from Path AI. The project was supervised by FAS.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eETHICS APPROVAL AND CONSENT TO PARTICIPATE\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eOur project was approved by the \u003cem\u003eHospital e Maternidade S\u0026atilde;o Luiz \u0026ndash; HMSL\u003c/em\u003e ethics committee. The need for informed consent terms from participants was waived. Our project was done in accordance with the Declaration of Helsinki.\u003c/p\u003e\n\u003cp\u003eThe Brazilian unified identifying code for the ethics committee assessment (CAAE) for the project is: \u003cstrong\u003e94672525.2.0000.0087\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eConsent for publication\u003c/strong\u003e: Not applicable. The project does not contain any identifying images or information for any patients.\u003c/p\u003e\n\n\u003cp\u003e\u003cstrong\u003eDATA AVAILABILITY\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eRaw data generated from the results of the sample assessment surveys and survey images themselves will be made available upon reasonable request to the corresponding author MPM.\u003c/p\u003e\n"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eKim J, Harper A, McCormack V, et al. Global patterns and trends in breast cancer incidence and mortality across 185 countries. Nat Med. 2025;31:1154\u0026ndash;62. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1038/s41591-025-03502-3\u003c/span\u003e\u003cspan address=\"10.1038/s41591-025-03502-3\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSung H, Ferlay J, Siegel RL, et al. Global Cancer Statistics 2020: GLOBOCAN Estimates of Incidence and Mortality Worldwide for 36 Cancers in 185 Countries. CA Cancer J Clin. 2021;71:209\u0026ndash;49. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.3322/caac.21660\u003c/span\u003e\u003cspan address=\"10.3322/caac.21660\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCarvalho E, Canberk S, Schmitt F, Vale N. Molecular Subtypes and Mechanisms of Breast Cancer: Precision Medicine Approaches for Targeted Therapies. Cancers (Basel). 2025;17:1102. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.3390/cancers17071102\u003c/span\u003e\u003cspan address=\"10.3390/cancers17071102\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCorti C, Batra-Sharma H, Kelsten M, et al. Systemic Therapy in Breast Cancer. Am Soc Clin Oncol Educational Book. 2024;44. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1200/EDBK_432442\u003c/span\u003e\u003cspan address=\"10.1200/EDBK_432442\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLehmann-Che J, Amira-Bouhidel F, Turpin E, et al. Immunohistochemical and molecular analyses of HER2 status in breast cancers are highly concordant and complementary approaches. Br J Cancer. 2011;104:1739\u0026ndash;46. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1038/bjc.2011.135\u003c/span\u003e\u003cspan address=\"10.1038/bjc.2011.135\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWolff AC, Somerfield MR, Dowsett M, et al. Human Epidermal Growth Factor Receptor 2 Testing in Breast Cancer. Arch Pathol Lab Med. 2023;147:993\u0026ndash;1000. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.5858/arpa.2023-0950-SA\u003c/span\u003e\u003cspan address=\"10.5858/arpa.2023-0950-SA\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSwain SM, Baselga J, Kim S-B, et al. Pertuzumab, Trastuzumab, and Docetaxel in HER2-Positive Metastatic Breast Cancer. N Engl J Med. 2015;372:724\u0026ndash;34. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1056/NEJMoa1413513\u003c/span\u003e\u003cspan address=\"10.1056/NEJMoa1413513\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCort\u0026eacute;s J, Hurvitz SA, Im S-A, et al. Trastuzumab deruxtecan versus trastuzumab emtansine in HER2-positive metastatic breast cancer: long-term survival analysis of the DESTINY-Breast03 trial. Nat Med. 2024;30:2208\u0026ndash;15. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1038/s41591-024-03021-7\u003c/span\u003e\u003cspan address=\"10.1038/s41591-024-03021-7\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eFu Z, Li S, Han S, et al. Antibody drug conjugate: the biological missile for targeted cancer therapy. Signal Transduct Target Ther. 2022;7:93. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1038/s41392-022-00947-7\u003c/span\u003e\u003cspan address=\"10.1038/s41392-022-00947-7\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eShitara K, Bang Y-J, Iwasa S, et al. Trastuzumab Deruxtecan in Previously Treated HER2-Positive Gastric Cancer. N Engl J Med. 2020;382:2419\u0026ndash;30. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1056/NEJMoa2004413\u003c/span\u003e\u003cspan address=\"10.1056/NEJMoa2004413\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eModi S, Jacot W, Yamashita T, et al. Trastuzumab Deruxtecan in Previously Treated HER2-Low Advanced Breast Cancer. N Engl J Med. 2022;387:9\u0026ndash;20. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1056/NEJMoa2203690\u003c/span\u003e\u003cspan address=\"10.1056/NEJMoa2203690\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTang J, Li N, Yang L, et al. Trastuzumab deruxtecan (T-DXd) monotherapy in advanced HER2-positive or HER2-low/ultralow breast cancer: An updated pooled analysis of prospective clinical trials. J Clin Oncol. 2025. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1200/JCO.2025.43.16_suppl.e13103\u003c/span\u003e\u003cspan address=\"10.1200/JCO.2025.43.16_suppl.e13103\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. 43:.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eFranchina M, Pizzimenti C, Fiorentino V, et al. Low and Ultra-Low HER2 in Human Breast Cancer: An Effort to Define New Neoplastic Subtypes. Int J Mol Sci. 2023;24:12795. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.3390/ijms241612795\u003c/span\u003e\u003cspan address=\"10.3390/ijms241612795\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eFernandez AI, Liu M, Bellizzi A, et al. Examination of Low ERBB2 Protein Expression in Breast Cancer Tissue. JAMA Oncol. 2022;8:607. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1001/jamaoncol.2021.7239\u003c/span\u003e\u003cspan address=\"10.1001/jamaoncol.2021.7239\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBaez-Navarro X, van Bockstal MR, Nawawi D, et al. Interobserver Variation in the Assessment of Immunohistochemistry Expression Levels in HER2-Negative Breast Cancer: Can We Improve the Identification of Low Levels of HER2 Expression by Adjusting the Criteria? An International Interobserver Study. Mod Pathol. 2023;36:100009. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1016/j.modpat.2022.100009\u003c/span\u003e\u003cspan address=\"10.1016/j.modpat.2022.100009\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSakatani T, Tsuda H, Yoshida M, et al. Current status and challenges in HER2 IHC assessment: scoring survey results in Japan. Breast Cancer Res Treat. 2025;210:27\u0026ndash;36. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1007/s10549-024-07532-2\u003c/span\u003e\u003cspan address=\"10.1007/s10549-024-07532-2\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKoelzer VH, Gisler A, Hanhart JC, et al. Digital image analysis improves precision of PD-L1 scoring in cutaneous melanoma. Histopathology. 2018;73:397\u0026ndash;406. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1111/his.13528\u003c/span\u003e\u003cspan address=\"10.1111/his.13528\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKludt C, Wang Y, Ahmad W, et al. Next-generation lung cancer pathology: Development and validation of diagnostic and prognostic algorithms. Cell Rep Med. 2024. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1016/j.xcrm.2024.101697\u003c/span\u003e\u003cspan address=\"10.1016/j.xcrm.2024.101697\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePulaski H, Harrison SA, Mehta SS, et al. Clinical validation of an AI-based pathology tool for scoring of metabolic dysfunction-associated steatohepatitis. Nat Med. 2024. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1038/s41591-024-03301-2\u003c/span\u003e\u003cspan address=\"10.1038/s41591-024-03301-2\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eShamshoian J, Shanis Z, Cabeen R et al. (2025) Impact of a machine learning-powered algorithm on pathologist HER2 IHC scoring in breast cancer.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eJaved SA, Juyal D, Padigela H, et al. Additive MIL. Intrinsically Interpretable Multiple Instance Learning for Pathology; 2022.\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"bmc-cancer","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"bcan","sideBox":"Learn more about [BMC Cancer](http://bmccancer.biomedcentral.com/)","snPcode":"","submissionUrl":"https://www.editorialmanager.com/bcan/default.aspx","title":"BMC Cancer","twitterHandle":"BMC_series","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"em","reportingPortfolio":"BMC Series","inReviewEnabled":true,"inReviewRevisionsEnabled":true},"keywords":"Digital Pathology, Artificial Intelligence (AI), Computer Vision, HER2, Immunohistochemistry, interobserver agreement","lastPublishedDoi":"10.21203/rs.3.rs-9023658/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-9023658/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003ch2\u003eBackground\u003c/h2\u003e \u003cp\u003eAccurate assessment of HER2 status in breast cancer has been critical for guiding therapy and has become even more important with the emergence of antibody-drug conjugates, now also indicated in HER2-low tumors. However, inter- and intraobserver variability limits the reproducibility of HER2 IHC scoring among pathologists, particularly at the lower end of the expression spectrum. Artificial intelligence (AI) models offer potential to standardize and improve diagnostic accuracy and bring new insights into current practices shortcomings.\u003c/p\u003e\u003ch2\u003eMethods\u003c/h2\u003e \u003cp\u003eWe conducted a study recruiting generalist and specialist pathologists from Rede D\u0026rsquo;Or centers across Brazil to assess digitized HER2 IHC whole slide images. The same images were presented for the pathologists with the interval of one month and to the AIM-HER2 (PathAI \u0026reg;, Boston, MA) AI model. Intra- and interobserver agreement, as well as concordance with AI, were measured across 126 breast cancer samples. The association between sample features and agreement metrics was also analyzed using AI spatial breakdown data.\u003c/p\u003e\u003ch2\u003eResults\u003c/h2\u003e \u003cp\u003eAmong 34 pathologists, median intraobserver concordance was 67.68% and median concordance with AI was 60.8%. Median interobserver agreement was 67.65%, with high agreement (\u0026gt;\u0026thinsp;85%) in 25.4% of samples. Significant positive correlations were observed among all agreement metrics. Samples with lower intra-sample heterogeneity, as determined by AI spatial breakdown scores, were associated with higher agreement levels.\u003c/p\u003e\u003ch2\u003eConclusions\u003c/h2\u003e \u003cp\u003eOur findings highlight significant variability in HER2 IHC scoring among pathologists. AI models such as AIM-HER2 can be used to increase reproducibility of analysis and also to indicate the most difficult samples that can result in diagnostic discordance. We found that AI assessed intra-sample heterogeneity is correlated with a lower agreement rate among pathologists.\u003c/p\u003e","manuscriptTitle":"Artificial Intelligence Model’s assessment of intra-sample heterogeneity of HER2 IHC in breast cancer is related to interobserver and intraobserver agreement among pathologists","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2026-04-09 15:09:28","doi":"10.21203/rs.3.rs-9023658/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"decision","content":"Revision requested","date":"2026-04-20T01:51:38+00:00","index":"","fulltext":""},{"type":"reviewerAgreed","content":"159576951999960432651599279920745757858","date":"2026-04-19T14:55:23+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"91743561878171953405148292183732859137","date":"2026-04-19T13:49:27+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"83005802483937805596540541602680504601","date":"2026-04-18T07:53:41+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2026-04-17T11:21:54+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"336701762190546839491833087899080866842","date":"2026-04-17T09:01:09+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"37629091589457528050372909757359994060","date":"2026-04-12T12:20:19+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"15869449409998084070845436583447438466","date":"2026-04-09T13:35:53+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2026-04-08T03:43:26+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"240748430307780625754942061598726252841","date":"2026-04-04T23:41:29+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2026-04-03T05:56:28+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"101029555190022829450628286930816239575","date":"2026-04-02T17:39:33+00:00","index":"hide","fulltext":""},{"type":"reviewersInvited","content":"","date":"2026-04-02T17:18:59+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2026-03-26T08:21:05+00:00","index":"","fulltext":""},{"type":"editorInvited","content":"","date":"2026-03-12T03:33:19+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2026-03-12T02:39:51+00:00","index":"","fulltext":""},{"type":"submitted","content":"BMC Cancer","date":"2026-03-11T19:04:11+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"bmc-cancer","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"bcan","sideBox":"Learn more about [BMC Cancer](http://bmccancer.biomedcentral.com/)","snPcode":"","submissionUrl":"https://www.editorialmanager.com/bcan/default.aspx","title":"BMC Cancer","twitterHandle":"BMC_series","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"em","reportingPortfolio":"BMC Series","inReviewEnabled":true,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"17f73726-c1dc-4daf-aaed-ff4a0fd2dc8f","owner":[],"postedDate":"April 9th, 2026","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"in-revision","subjectAreas":[],"tags":[],"updatedAt":"2026-04-20T01:53:58+00:00","versionOfRecord":[],"versionCreatedAt":"2026-04-09 15:09:28","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-9023658","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-9023658","identity":"rs-9023658","version":["v1"]},"buildId":"XKTyCvWXoU3ODBz1xrDgd","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2026) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00
unpaywall
last seen: 2026-05-21T05:10:58.409756+00:00
License: CC-BY-4.0