Large Language Models for Automated Icd-10 Coding of Obstetric Clinical Notes in Portuguese: Comparison With Human Coders | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Article Large Language Models for Automated Icd-10 Coding of Obstetric Clinical Notes in Portuguese: Comparison With Human Coders Ricardo da Silva Santos, Murilo Gleyson Gazzola, Paulo Marcelino Figueira, and 3 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-8712058/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract Background Despite rapid advances in large language models (LLMs), automated ICD-10 coding of real-world clinical narratives remains unreliable. A key challenge lies not only in model limitations, but in the intrinsic ambiguity of clinical documentation and the substantial variability among human coders. Methods We benchmarked six general-purpose LLMs for hierarchical ICD-10 coding of 1,117 obstetric discharge summaries written in Brazilian Portuguese. Model performance was evaluated at both category and leaf levels and contextualized against blinded clinician validation to quantify realistic human agreement. We further assessed whether Portuguese-to-English translation or lightweight supervised fine-tuning improved performance. Results Even the best-performing model (GPT-4o) achieved only modest agreement with the human reference, with micro-F1 scores of 0.36 at the three-character level and 0.15 at the leaf level. Translation into English did not yield consistent gains, and direct fine-tuning on code–description pairs failed to improve accuracy. In clinician validation, human-coded references achieved a precision of 0.77, compared to 0.59 for the strongest model, revealing a substantial gap between automated predictions and clinically accepted codes. Conclusions Our findings indicate that current LLMs remain below human-level reliability for autonomous ICD-10 coding in Portuguese. More importantly, they show that conventional evaluation metrics such as F1-score substantially misrepresent clinical usefulness by conflating model error with human disagreement. The primary bottleneck for automated medical coding is therefore not model capacity, but the imperfect and variable human gold standard itself. LLMs should be positioned as decision-support tools that assist, rather than replace, expert clinical coders. Biological sciences/Computational biology and bioinformatics Health sciences/Diseases Health sciences/Health care Physical sciences/Mathematics and computing Health sciences/Medical research Figures Figure 1 Introduction Medical coding is fundamental to health systems, as it converts clinical documentation into standardized codes that underpin patient care, billing, research, and public health reporting 8 . Automating ICD coding has long been a central goal in natural language processing (NLP) 9 – 10 . The recent progress of large language models (LLMs) has fueled expectations that clinical coding might be solved off the shelf 11 – 12 . However, published evaluations consistently show that LLMs attain low accuracy when predicting codes. 13 – 14 Most prior studies on automated ICD coding implicitly assume that higher F1-scores reflect better clinical performance. However, medical coding is inherently subjective, with substantial inter-annotator disagreement even among trained professionals. As a result, traditional set-based metrics such as F1-score conflate model error with human disagreement, obscuring the true clinical value of automated systems. Against this backdrop, we benchmark multiple LLMs for hierarchical ICD-10 coding of Portuguese obstetric notes, compare them with a fine-tuned specialist, test whether Portuguese-to-English translation improves performance, and contextualize results against human inter-annotator agreement. In Brazilian Portuguese, systematic investigations of automated ICD coding remain scarce. Foundational Portuguese-language resources such as BRATECA and SemClinBr have advanced the local clinical NLP ecosystem, yet they have not been used to benchmark automated ICD-10 assignment on discharge summaries. 15 – 16 Prior Portuguese-language efforts either relied on pre-Transformer, classical methods to predict multiple ICD-10 codes from Brazilian clinical notes, or pursued different tasks (e.g., clinical assistants and note generation) rather than code assignment. Meanwhile, transformer-based ICD-10 coding has been examined for death certificates in European Portuguese, not for hospital discharge notes. 17 – 20 It is also essential to situate automation against realistic human performance. Disagreement among professional coders is common, reflecting ambiguities in clinical documentation, overlapping categories, and local conventions; automated systems should therefore be evaluated against this inherent variability. 21 – 22 Specifically, we address four gaps that are rarely examined together: (i) the realistic performance ceiling imposed by human disagreement; (ii) the behavior of general-purpose LLMs under hierarchical ICD-10 constraints; (iii) the limited transferability of lightweight fine-tuning strategies; and (iv) the questionable validity of conventional evaluation metrics for clinical coding. Methods Data Source and Reference Standard The corpus comprised 1,117 discharge summaries from a public, university-affiliated tertiary hospital specializing in women’s health in Brazil. These documents were retrospectively obtained and re-labeled by trained clinical coders specifically for this study, following a standardized protocol to ensure consistency and to establish the reference standard. Additional corpus statistics and note-length distributions are reported in the Supplementary Materials (S1.1; Table S1 and Figures S1 –S2). For secondary analyses requiring blinded human agreement, we selected a stratified subset of 315 summaries. This subset was constructed to maintain proportional representation of the major ICD-10 obstetric categories and to ensure sufficient variability in case complexity, enabling reliable estimation of inter-annotator disagreement. Task Definition and Label Space We framed coding as multi-label assignment of ICD-10 codes from full, free-text notes. Performance was assessed at two levels: Leaf level (full code with subcategory, e.g., O24.4 ), and Three-character category (e.g., O24 ), used to quantify hierarchical specificity. All codes were normalized to uppercase, whitespace was stripped, and duplicates removed. Out-of-scope/invalid tokens were discarded prior to scoring. We did not cap the number of codes a model could output per note. Models and Inference We evaluated six general-purpose LLMs plus one specialist variant: GPT-4o 1 , GPT-4o-mini 2 – 3 , Sabiá-3.1 4 , DeepSeek-V3 5 , Gemini 1.5 Flash 6 , and a fine-tuned GPT-4o-mini. All models were accessed via vendor APIs with temperature = 0 (deterministic decoding); remaining parameters followed vendor defaults unless stated otherwise. Inference was executed during June 10–11, 2025 (UTC); model names and API versions are listed in the Supplement to mitigate model-drift concerns. Exact model identifiers, endpoints, and inference dates are provided in the Supplementary Materials (S1.2; Table S2 ). Specialist fine-tuned variant We fine-tuned gpt-4o-mini-2024-07-18 to emit JSON-only ICD-10 outputs; the provider snapshot lists 2 epochs, batch 16, LR-multiplier 1.8, seed 1, and 2,325,050 trained tokens (job ftjob-wqj9laPJ3wHdBc75fVmpGEgU). Full metadata and the step log (1,557 steps; training loss 3.303→0.025) appear in Supplement Materials (S1.3; Figure S3 ) Importantly, the objective of this fine-tuning experiment was not to maximize performance through extensive hyperparameter optimization, data augmentation, or validation-based early stopping. Instead, it was designed to test a commonly proposed lightweight specialization strategy—training on isolated code–description pairs—and to evaluate whether such an approach transfers to multi-label coding of free-text clinical narratives. Prompting and Output Format For each model, we requested JSON-only outputs listing ICD-10 codes (no explanations). Where supported, we enforced structured output and validated the response schema with Pydantic (v2). For DeepSeek-V3, which does not expose structured output in the chat-completions route we used, we instructed JSON in the prompt and then parsed the returned text with conservative tokenization and our ICD-10 regex normalization. After validation/normalization, outputs were stored as Python lists of codes in the analysis DataFrame (one column per model). Full prompts are provided in the Supplement S1.4. Portuguese-to-English Translation Condition We translated each Portuguese discharge note to English using Google Translate accessed via the googletrans 7 4.0.0-rcl Python client, with default settings (one-pass, no stochastic sampling). Each translated note was then paired with an English version of the coding prompt—a faithful translation of the Portuguese prompt preserving the same JSON schema and decoding parameters—and submitted to the model. We applied the same post-processing pipeline as in the Portuguese condition (parsing, normalization, ICD-10 regex validation, de-duplication). No human post-editing of translations was performed. Human Agreement To contextualize model performance against a realistic human benchmark, we conducted a blinded clinician validation on the full test set of 315 obstetric discharge notes. For each note we built a candidate pool equal to the union of (i) the gold-standard ICD-10 codes and (ii) all codes proposed by the six model arms. A custom web interface (Streamlit) displayed each candidate code with its human-readable description; the reviewer was blinded to source (gold vs. model and model identity) and instructed to mark each code as clinically appropriate (TRUE) or not appropriate (FALSE) for that note. Multiple TRUE selections were allowed per note; notes were de-identified. Further details on the human validation workflow, including annotation steps and a representative screenshot of the annotation interface, are provided in the Supplementary Materials (Section S1.5; Figure S4 ). Statistical Analysis Model performance was primarily evaluated using precision, recall, and the F1-score, calculated at both the three-character category and the exact leaf levels. We report both micro- and macro-averaged metrics. Micro-averaged metrics were calculated globally by aggregating the counts of true positives, false positives, and false negatives across the entire dataset. Macro-averaged metrics were calculated by first computing the metric for each individual note, then taking the unweighted average of these per-note scores. To quantify the uncertainty in our performance estimates for Tables 2 , 3 , and 4 , we computed 95% confidence intervals (CIs) using non-parametric bootstrap resampling. For each metric, we generated 10,000 bootstrap samples by resampling the 1,117 notes with replacement. The 95% CI was then derived from the 2.5th and 97.5th percentiles of the resulting bootstrap distribution. Bootstrap distributions for all reported metrics are shown in the Supplementary Materials (S2.2; Figures S5 –S17). For the clinician-validated precision reported in Table 5 , which represents a binomial proportion, the 95% CI was calculated using the Wilson score interval, a method robust for proportions, including those near the boundaries of 0 or 1. Results Cohort and label characteristics We analyzed 1,117 obstetric discharge notes. The gold standard contained a median of 2 ICD-10 codes per note (IQR [2–3] ), spanning 261 three-character categories and 452 leaf codes. Across model outputs, distinct code diversity ranged from 490 to 978 at the leaf level and 242 to 481 at the category level. Additional per-model descriptive statistics are provided in the Supplementary Materials (S2.1; Table S4 ). Bars show, for GOLD (human annotations) and the six models, the number of unique codes observed in the corpus (n = 1,117 notes). In “Exact (leaf level)”, we count all leaf-level codes—including three-character categories and dotted subcategories—after normalization and within-note deduplication. In “Category (3 chars)”, we count only three-character categories (A00–Z99). Values above bars are the distinct counts. Greater diversity reflects broader code coverage, not necessarily higher accuracy. GOLD = human reference. All models were run with the same prompt and temperature = 0. Figure 1 — Diversity of distinct ICD-10 codes by set. Table 1 Descriptive statistics for the number of ICD-10 codes per note by model and GOLD.Values are reported as median (IQR) and mean ± SD over 1,117 discharge notes. Counts are computed at the leaf level (three-character categories and dotted subcategories), after normalization and within-note deduplication. All models were run deterministically (temperature = 0) with the same prompt. GOLD = human reference. Q1 = 25th percentile; Q3 = 75th percentile. Model n_notes median q1 q3 mean sd Sabiá-3.1 4 1117 6.0 5.0 7.0 6.24 1.90 DeepSeek-V3 5 1117 5.0 5.0 6.0 5.54 1.35 GPT-4o-Mini 2 – 3 1117 5.0 5.0 6.0 5.27 0.99 GPT-4o 1 1117 4.0 3.0 5.0 4.42 1.36 Gemini-1.5 Flash 6 1117 5.0 4.0 6.0 4.93 1.34 Gold 1117 2.0 2.0 3.0 2.58 1.39 Primary Endpoint: Three-Character Category On the primary endpoint (micro-F1 at the three-character category level), GPT-4o achieved the best performance with a score of 0.37 (95% CI 0.35–0.38). It was followed by GPT-4o-Mini (0.34, 95% CI 0.33–0.36) and DeepSeek-V3 (0.34, 95% CI 0.33–0.35); the substantial overlap in their confidence intervals suggests their performance was statistically comparable. The remaining models followed in descending order: Sabiá-3.1 (0.31, 95% CI 0.30–0.33), Gemini 1.5 Flash (0.23, 95% CI 0.22–0.25), and the fine-tuned GPT-4o-Mini (0.21, 95% CI 0.20–0.22). Full micro- and macro-averaged metrics for all models, including precision and recall, are detailed in Table 2 . Table 2 ICD-10 Coding Performance by Model — Portuguese, Three-Character Category Level. Modelo Micro Macro Precision Recall F1 Precision Recall F1 Sabiá-3.1 0.2255 (0.2159–0.2354) 0.5192 (0.5005–0.5378) 0.3145 (0.3025–0.3267) 0.2323 (0.2221–0.2425) 0.5467 (0.5259–0.5673) 0.3095 (0.2974–0.3216) DeepSeek-V3 0.2514 (0.2408–0.2619) 0.5330 (0.5146–0.5512) 0.3417 (0.3288–0.3544) 0.2496 (0.2390–0.2601) 0.5431 (0.5224–0.5638) 0.3285 (0.3157–0.3411) Fine-Tuned GPT-4o-Mini 0.1593 (0.1505–0.1681) 0.3110 (0.2940–0.3283) 0.2107 (0.1994–0.2218) 0.1689 (0.1599–0.1780) 0.3648 (0.3440–0.3852) 0.2180 (0.2068–0.2292) GPT-4o-mini 0.2585 (0.2478–0.2689) 0.5199 (0.5022–0.5380) 0.3453 (0.3326–0.3576) 0.2597 (0.2490–0.2705) 0.5462 (0.5258–0.5663) 0.3353 (0.3228–0.3474) GPT-4o 0.2933 (0.2805–0.3062) 0.4918 (0.4739–0.5096) 0.3675 (0.3531–0.3818) 0.3000 (0.2867–0.3131) 0.5036 (0.4830–0.5241) 0.3596 (0.3450–0.3742) Gemini-1.5 Flash 0.1799 (0.1702–0.1897) 0.3349 (0.3183–0.3525) 0.2341 (0.2222–0.2462) 0.1811 (0.1713–0.1910) 0.3567 (0.3371–0.3768) 0.2297 (0.2178–0.2416) Leaf Level Requiring leaf-level specificity substantially reduced performance for all systems. As detailed in Table 3 , the micro-F1 score for the top-performing model, GPT-4o , dropped from 0.37 to 0.15 (95% CI 0.14–0.16) . This degradation was consistent across all models, with relative performance drops (ΔF₁) ranging from − 52.5% for DeepSeek-V3 to a stark − 77.7% for the fine-tuned model. These results indicate that precise subcategory resolution (e.g., O24.4 vs. O24) remains a significant challenge. Table 3 ICD-10 Coding Performance by Model — Portuguese, Leaf Level Layout identical to Table 2 , but here each full code (including subcategories) is counted separately.F₁ = percentage change in Micro-F₁ from the three-character category level to the four-character leaf level. Δ Model Micro Macro ΔF₁ Precision Recall F1-Score Precision Recall F1-Score Sabiá-3.1 4 0.0798 (0.0737–0.0861) 0.1916 (0.1775–0.2064) 0.1126 (0.1042–0.1213) 0.0822 (0.0756–0.0890) 0.2060 (0.1891–0.2234) 0.1114 (0.1028–0.1202) -64.20% DeepSeek-V3 5 0.1193 (0.1115–0.1275) 0.2543 (0.2394–0.2702) 0.1624 (0.1524–0.1730) 0.1171 (0.1091–0.1253) 0.2433 (0.2265–0.2608) 0.1523 (0.1422–0.1628) -52.47% Fine-Tuned GPT-4o-mini 0.0341 (0.0297–0.0387) 0.0748 (0.0652–0.0849) 0.0469 (0.0409–0.0531) 0.0379 (0.0328–0.0432) 0.0890 (0.0766–0.1020) 0.0503 (0.0438–0.0571) -77.74% GPT-4o-mini 2 – 3 0.1043 (0.0964–0.1123) 0.2116 (0.1968–0.2269) 0.1398 (0.1296–0.1501) 0.1046 (0.0964–0.1128) 0.2100 (0.1936–0.2272) 0.1334 (0.1235–0.1435) -59.51% GPT-4o 1 0.1169 (0.1077–0.1265) 0.1992 (0.1845–0.2142) 0.1474 (0.1361–0.1589) 0.1186 (0.1091–0.1284) 0.1942 (0.1782–0.2110) 0.1410 (0.1299–0.1525) -59.89% Gemini-1.5 Flash 6 0.0653 (0.0589–0.0719) 0.1241 (0.1121–0.1363) 0.0855 (0.0773–0.0941) 0.0642 (0.0576–0.0709) 0.1225 (0.1089–0.1364) 0.0811 (0.0727–0.0895) -63.48% Portuguese-to-English translation To assess whether translating notes into English could improve ICD-10 coding, we ran a translation arm in which each Portuguese discharge summary was translated via Google Translate (deterministic configuration) and then submitted—using an English-language version of our standard JSON‐only prompt—to the same three models: Sabiá-3.1, DeepSeek-V3, and GPT-4o. Before selecting Google Translate, we benchmarked three PT→EN engines—Google Translate, the GPT-4o-mini API, and the Sabiá API—while holding the coder fixed (Sabiá-3.1); Google Translate achieved the highest scores and was therefore adopted for the English-input arm (see Supplementary Table S5 ). Table 4 reports micro‐precision, micro‐recall, and micro–F₁ at both leaf (exact code) and three-character category levels, along with ΔF₁ (defined as F₁English − F₁Portuguese). Sabiá-3.1 saw a small improvement on English notes (leaf ΔF₁ = +0.01410; category ΔF₁ = +0.0006). DeepSeek-V3 performance declined after translation (leaf ΔF₁ = − 0.0380; category ΔF₁ = − 0.0278). GPT-4o also dropped (leaf ΔF₁ = − 0.0251; category ΔF₁ = − 0.0436). Overall, translating into English did not yield consistent gains—and in two of three systems degraded coding accuracy—underscoring that native-language prompts remain preferable (Table 4 ). Exploratory role-specific precision for the instructed primary/secondary code roles is reported in the Supplementary Materials (S2.3; Table S6 ). Table 4 Comparative performance of Sabiá-3.1, DeepSeek-V3, and GPT-4o on English-translated vs Native-portuguese Notes For each condition (Leaf level and Three-character category) we report micro-precision, micro-recall, and micro-F₁. F₁ indicates the change in micro-F₁ resulting from translation and is defined as the micro-F₁ score on the English-translated notes minus the micro-F₁ score on the original Portuguese notes. Δ Level Model Micro Macro ΔF₁ Precision Recall F1 Precision Recall F1 Three Character GPT-4o 0.2553 (0.2444–0.2663) 0.4428 (0.4243–0.4616) 0.3239 (0.3109–0.3371) 0.2571 (0.2458–0.2685) 0.4828 (0.4620–0.5042) 0.3199 (0.3066–0.3332) -0.0436 Sabiá-3.1 0.2337 (0.2239–0.2438) 0.4836 (0.4648–0.5030) 0.3151 (0.3031–0.3276) 0.2323 (0.2223–0.2424) 0.5192 (0.4986–0.5402) 0.3052 (0.2933–0.3175) 0.0006 DeepSeek-V3 0.2357 (0.2263–0.2456) 0.4698 (0.4512–0.4891) 0.3139 (0.3021–0.3262) 0.2327 (0.2230–0.2426) 0.5080 (0.4875–0.5291) 0.3047 (0.2929–0.3169) -0.0278 Leaf Level GPT-4o 0.0857 (0.0781–0.0932) 0.1510 (0.1376–0.1644) 0.1094 (0.0997–0.1188) 0.0873 (0.0793–0.0951) 0.1626 (0.1470–0.1784) 0.1089 (0.0992–0.1185) -0.0251 Sabiá-3.1 0.0935 (0.0865–0.1006) 0.1969 (0.1825–0.2113) 0.1267 (0.1176–0.1361) 0.0930 (0.0860–0.1002) 0.2096 (0.1925–0.2267) 0.1227 (0.1136–0.1318) 0.0141 DeepSeek-V3 0.1021 (0.0949–0.1094) 0.2093 (0.1949–0.2239) 0.1373 (0.1278–0.1467) 0.1005 (0.0931–0.1079) 0.2140 (0.1974–0.2308) 0.1311 (0.1218–0.1406) -0.0380 Clinician validation In a blinded clinician validation of 315 obstetric discharge notes, the gold standard achieved a human-validated precision (PPV_H) of 0.77 (95% CI 0.74–0.79). Among the AI systems, GPT-4o led with PPV_H = 0.59 (95% CI 0.56–0.62), followed by DeepSeek-V3 at 0.54 (95% CI 0.51–0.56), GPT-4o-mini at 0.43 (95% CI 0.40–0.45), Sabiá-3.1 at 0.35 (95% CI 0.33–0.37), Gemini 1.5 Flash at 0.32 (95% CI 0.29–0.34), and the fine-tuned GPT-4o-mini at 0.30 (95% CI 0.28–0.33) (Table 5 ). The “gold gap” (proportion of gold-standard codes rejected by the clinician) was 23%, highlighting substantial intra-human variability. Table 5 Clinician-Validated Precision of ICD-10 Code Predictions by Model Model Trues False Total Codes Precision 95% CI LOWER 95% CI UPPER Sabiá-3.1 4 688 1300 1988 0.3461 0.3255 0.3673 DeepSeek-V3 5 937 806 1743 0.5376 0.5141 0.5609 Fine-Tuned GPT-4o-mini 532 1217 1749 0.3042 0.2831 0.3261 GPT-4o-mini 2 – 3 703 949 1652 0.4255 0.4019 0.4495 GPT-4o 1 819 569 1388 0.5901 0.5640 0.6156 Gemini-1.5 Flash 6 481 1042 1523 0.3158 0.2930 0.3396 Gold 624 190 814 0.7666 0.7363 0.7944 Fine-tuned specialist We evaluated whether light-weight supervised fine-tuning of GPT-4o-mini on single-label code–description pairs would improve multi-label coding of free-text discharge notes. The fine-tuned model generated a distribution of codes per note (median 5.0, IQR 5.0–6.0; mean 5.69 ± 2.06) that closely mirrored other systems (Table 1 ), but did not match the more conservative human reference (median 2.0, IQR 2.0–3.0). Despite this apparent alignment in output volume, fine-tuning failed to yield any accuracy gains. At the three-character category level, the specialist achieved micro-F₁ = 0.19 versus 0.36 for the best zero-shot model (GPT-4o), with similarly low macro-F₁ and no improvement in precision or recall (Table 2 ). Clinician-validated precision for the fine-tuned model was only 0.30 (95% CI 0.28–0.33), well below both the GPT-4o zero-shot system (0.59, 95% CI 0.56–0.62) and the gold standard (0.77, 95% CI 0.74–0.79) (Table 5 ). These results indicate that direct supervised fine-tuning on isolated code–description pairs does not transfer effectively to narrative clinical notes and offers no advantage in hierarchical specificity. Discussion In this study, we systematically evaluated six large language models (LLMs) for automated ICD-10 coding of Portuguese obstetric discharge notes, comparing them against a human reference standard and a fine-tuned specialist variant. Even the best performer, GPT-4o, achieved only a micro-F₁ of 0.36 at the three-character category level and 0.15 at the full (leaf) level. Translating notes into English conferred no performance benefit, and lightweight fine-tuning of GPT-4o-mini on description–code pairs likewise failed to improve outcomes. Our findings align with English-language benchmarks reporting modest off-the-shelf LLM performance on clinical coding tasks. For example, Dong et al. observed micro-F₁ scores below 0.50 for general discharge-summary coding 9 , and Soroush et al. characterized standard LLMs as “poor medical coders” in code-query settings 13 . Portuguese‐language studies to date have been limited to pre-Transformer techniques 17 or as ICD-10 assignment on death certificates 20 , which involve shorter, more standardized text, leaving a gap for Brazilian clinical notes that our work now fills. We observed a gold gap of 23%—the proportion of official codes the clinician rejected as non-applicable—a figure below the 30–40% disagreement typically reported in inter-coder reliability studies, in which two independent coders assign codes de novo and are then compared 21 – 22 . Although our protocol relied on validation against a pre-labeled set, the observed discrepancy reaffirms inherent human variability and sets a practical ceiling for automated systems. Nevertheless, GPT-4o achieved a clinician-validated precision of 59% (vs. 77% for the gold standard), demonstrating competitive zero-shot performance on Portuguese text. These results reinforce that LLMs should be regarded as assistance tools—proposing codes for expert review—rather than direct replacements for human coders. Moreover, the lack of improvement with English translation highlights the importance of developing and evaluating models directly in the native clinical language. This single-center obstetrics study may not generalize to other specialties or settings. Our fine-tuning protocol omitted data augmentation, validation splits, and early stopping—strategies shown to boost micro-F₁ to ~ 0.69 on English clinical notes through multi-phase, objective-aligned training 23 . We also did not explore few-shot prompting, chain-of-thought techniques, or retrieval-augmented architectures, all of which warrant investigation. In summary, while LLMs exhibit promise in medical natural language processing, their application to ICD-10 coding in Brazilian Portuguese should be positioned as coders’ assistants, supported by robust specialization pipelines and rigorous human oversight to ensure safety and reliability in real-world clinical practice. Declarations Ethics approval and consent to participate This study was approved by the Research Ethics Committee (Comitê de Ética em Pesquisa – CEP) of the University of Campinas (UNICAMP), Brazil, under approval number 7.446.157 (CAAE: 86136424.6.0000.5404). The study was conducted in accordance with national and institutional ethical standards and with the Declaration of Helsinki. It is a retrospective study based on anonymized electronic medical records, with no direct patient contact or intervention. The requirement for informed consent was waived by the Ethics Committee due to the use of fully anonymized retrospective data. Conflict of Interest The authors declare no competing interests. Funding This research did not receive any specific grant from funding agencies in the public, commercial, or not-for-profit sectors. Author Contribution R.S.S., M.G.G., and C.T. conceived the study and designed the experimental framework. R.S.S. implemented the computational pipeline, conducted the statistical analyses, and drafted the initial version of the manuscript. M.G.G. and C.T. contributed to the methodological design, interpretation of the results, and critical revision of the manuscript. P.M.F. and A.G.L. contributed to data curation, clinical interpretation, and validation of the reference annotations. R.C.P. contributed to clinical oversight, interpretation of obstetric outcomes, and refinement of the clinical evaluation protocol. All authors reviewed and approved the final manuscript. Acknowledgement The authors thank CAISM – Hospital da Mulher Prof. Dr. José Aristodemo Pinotti (UNICAMP) for providing access to the clinical data used in this study. Data Availability The code and scripts used to perform data processing, model evaluation, and statistical analysis in this study are publicly available in the following GitHub repository: https://github.com/ricardosantoss/llm-icd10-pt-obstetrics. This repository includes all code necessary to reproduce the experiments reported in the manuscript, including preprocessing, model inference, evaluation metrics, and visualizations. Dependencies and execution instructions are provided in the repository README. De-identified clinical data used in the study cannot be shared publicly due to ethical constraints; see the Data Availability statement for details on accessing de-identified data upon reasonable request. References OpenAI. GPT-4o System Card. OpenAI; 2024. Available at: https://openai.com/index/gpt-4o-system-card/ . OpenAI. GPT-4o mini: Advancing cost-efficient intelligence. OpenAI; July 18, 2024. https://openai.com/index/gpt-4o-mini-advancing-cost-efficient-intelligence/ . OpenAI. gpt-4o-mini — Model reference. OpenAI Platform Docs; 2024–2025. https://platform.openai.com/docs/models/gpt-4o-mini . Abonizio H, Almeida TS, Laitz T, et al. Sabiá-3 Technical Report. arXiv. 2024; DOI: 10.48550/arXiv.2410.12049 . Liu A, Feng B, Xue B, et al. DeepSeek-V3 Technical Report. arXiv. 2024; DOI: 10.48550/arXiv.2412.19437 . Georgiev P, Lei VI, Burnell R, et al. Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context. arXiv. 2024; DOI: 10.48550/arXiv.2403.05530 . Teo. SS. Googletrans: Free and Unlimited Google translate API for Python. PyPI; 2020. ( https://pypi.org/project/googletrans/) . World Health Organization. International Statistical Classification of Diseases and Related Health Problems, 10th Revision (ICD-10). Geneva; WHO; ( https://icd.who.int/browse10 ) Dong H, Falis M, Whiteley W, et al. Automated clinical coding: what, why, and where we are? Npj Digital Medicine. 2022. DOI: 10.1038/s41746-022-00705-7 . Yan C, Fu X, Liu X, et al. A survey of automated International Classification of Diseases coding: development, challenges, and applications. Intelligent Medicine. 2022; DOI: 10.1016/j.imed.2022.03.003 Nori H, King N, McKinney SM, Carignan D, Horvitz E. Capabilities of GPT-4 on medical challenge problems. arXiv. 2023; DOI: 10.48550/arXiv.2303.13375 . Singhal K, Tu T, Gottweis J, et al. Towards expert-level medical question answering with large language models. Nature Medicine. 2025. DOI: 10.1038/s41591-024-03423-7 . Soroush A, Glicksberg BS, Zimlichman E, et al. Large Language Models Are Poor Medical Coders — Benchmarking of Medical Code Querying. NEJM AI. 2024;. DOI: 10.1056/AIdbp2300040 . Kaur R, Ginige JA, Obst O. AI-based ICD coding and classification approaches using discharge summaries: a systematic literature review. Expert Systems with Applications. 2023; DOI: 10.1016/j.eswa.2022.118997 . Consoli B, Santos HDP, Ulbrich AHDPS, Vieira R, Bordini RH. BRATECA (Brazilian Tertiary Care Dataset): a Clinical Information Dataset for the Portuguese Language. In: Proceedings of LREC 2022. European Language Resources Association; 2022. ( https://aclanthology.org/2022.lrec-1.602/ ). Oliveira LES, Peters AC, Silva AMP, et al. SemClinBr—A multi-institutional and multi-specialty semantically annotated corpus for Portuguese clinical NLP tasks. Journal of Biomedical Semantics. 2022; DOI: 10.1186/s13326-022-00269-1 . Reys AD, Silva D, Severo D, Pedro S, Sá MMS, Salgado GAC. Predicting Multiple ICD-10 Codes from Brazilian-Portuguese Clinical Notes. In: Proceedings of the 30th Brazilian Conference on Intelligent Systems (BRACIS). 2020. DOI: 10.1007/978-3-030-61377-8_39 . Paiola PH, Garcia GL, Manesco JRR, Roder M, Rodrigues D, Papa JP. Adapting LLMs for the Medical Domain in Portuguese: A Study on Fine-Tuning and Model Evaluation. In: Proceedings of the WSCG Conference. 2025. ( http://wscg.zcu.cz/WSCG2025/papers/A13.pdf ) Pinto JGS, Freitas AR, Martins AC, Sawazaki CM, Vidal C, Oliveira LE. Developing resource-efficient clinical LLMs for Brazilian Portuguese. In: Proceedings of the 34th Brazilian Conference on Intelligent Systems (BRACIS). 2024. DOI: 10.1007/978-3-031-79038-6_4 . Coutinho I, Martins B. Transformer-based models for ICD-10 coding of death certificates with Portuguese text. Journal of Biomedical Informatics. 2022. DOI: 10.1016/j.jbi.2022.104232 . Wockenfuss R, Frese T, Herrmann K, Claussnitzer M, Sandholzer H. Three- and four-digit ICD-10 is not a reliable classification system in primary care. Scand J Prim Health Care. 2009. DOI: 10.1080/02813430903072215 . Stausberg J, Lehmann N, Kaczmarek D, Stein M. Reliability of diagnoses coding with ICD-10. Int J Med Inform. 2008. DOI: 10.1016/j.ijmedinf.2006.11.005 . Hou Z, Liu H, Bian J, He X, Yan Z. Enhancing medical coding efficiency through domain-specific fine-tuned large language models. npj Health Systems. 2025. DOI: 10.1038/s44401-025-00018-3 . Additional Declarations No competing interests reported. Supplementary Files SupplementalEvaluatingLargeLanguageModelsforAutomatedICD10CodingofObstetricClinicalNotesinPortugueseAComparativeStudy.docx Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-8712058","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Article","associatedPublications":[],"authors":[{"id":585999584,"identity":"9e6fa76e-2b33-4701-9919-b49c5ea97d35","order_by":0,"name":"Ricardo da Silva Santos","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAABHUlEQVRIiWNgGAWjYDACZhDBBufa8EBoAwZUcRxaGBsYGNKI0IIkBdJyGIs4GjBvZ374uaLMhkG+/ezxBx/+nJfRbW9gky4osJPXnd3A9rgCU4vMYTZjyTPn0hgMzuQlNs5su81jduYAm/QMg2TDbXcOsBuewdQiwczDINnYdhjo+BzDZt4GoJYbCWzSPAYHGLcBGZINWLUw/wRpke9/Y9jM8+ccj9n9B2At9ni0sIFtYbgBtIWH7QDQFgawlkTcWtjMLBvOpfEY3HhjOHNmWzLQL4nN1jwGycnb7hxsN8Smhf/w45sNZTZy8v05Bh8+/LGzNzt++OBtnj92tttuNx97iEULDPAgsRmh6iRgDOKBBKkaRsEoGAWjYJgCAHVIYSnEbZxpAAAAAElFTkSuQmCC","orcid":"","institution":"State University of Campinas","correspondingAuthor":true,"prefix":"","firstName":"Ricardo","middleName":"da Silva","lastName":"Santos","suffix":""},{"id":585999587,"identity":"7ad9d75a-ac7d-4a43-b094-352ef55bbcae","order_by":1,"name":"Murilo Gleyson Gazzola","email":"","orcid":"","institution":"Universidade Presbiteriana Mackenzie","correspondingAuthor":false,"prefix":"","firstName":"Murilo","middleName":"Gleyson","lastName":"Gazzola","suffix":""},{"id":585999589,"identity":"2ed19bf3-f3f9-430c-80be-580222a8081a","order_by":2,"name":"Paulo Marcelino Figueira","email":"","orcid":"","institution":"State University of Campinas","correspondingAuthor":false,"prefix":"","firstName":"Paulo","middleName":"Marcelino","lastName":"Figueira","suffix":""},{"id":585999609,"identity":"02115884-cf87-4c96-bc33-371d1fdece88","order_by":3,"name":"Adriana Gomes Luz","email":"","orcid":"","institution":"State University of Campinas","correspondingAuthor":false,"prefix":"","firstName":"Adriana","middleName":"Gomes","lastName":"Luz","suffix":""},{"id":585999610,"identity":"82f7405a-d800-4e3d-9c75-24441f41d397","order_by":4,"name":"Rodolfo de Carvalho Pacagnella","email":"","orcid":"","institution":"State University of Campinas","correspondingAuthor":false,"prefix":"","firstName":"Rodolfo","middleName":"de Carvalho","lastName":"Pacagnella","suffix":""},{"id":585999621,"identity":"bc131d00-f71c-4780-9453-a8ed18e07c72","order_by":5,"name":"Cristiano Torezzan","email":"","orcid":"","institution":"State University of Campinas","correspondingAuthor":false,"prefix":"","firstName":"Cristiano","middleName":"","lastName":"Torezzan","suffix":""}],"badges":[],"createdAt":"2026-01-27 15:09:46","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-8712058/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-8712058/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":101966431,"identity":"b9359391-734f-4284-9ed9-3373284c6d09","added_by":"auto","created_at":"2026-02-05 13:49:00","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":50915,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eDiversity of distinct ICD-10 codes by set.\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eBars show, for GOLD (human annotations) and the six models, the number of unique codes observed in the corpus (n = 1,117 notes). In “Exact (leaf level)”, we count all leaf-level codes—including three-character categories and dotted subcategories—after normalization and within-note deduplication. In “Category (3 chars)”, we count only three-character categories (A00–Z99). Values above bars are the distinct counts. Greater diversity reflects broader code coverage, not necessarily higher accuracy. GOLD = human reference. All models were run with the same prompt and temperature = 0.\u003c/p\u003e","description":"","filename":"1.png","url":"https://assets-eu.researchsquare.com/files/rs-8712058/v1/ae514fc519ce59f5bbf8c907.png"},{"id":105070231,"identity":"402028ad-69ee-4c63-8d2b-a50f7ae4a4a2","added_by":"auto","created_at":"2026-03-20 14:56:29","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":968729,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-8712058/v1/3c61e230-aba3-4282-9abb-1440bbbec158.pdf"},{"id":101966432,"identity":"202c9a4e-2b8f-4df4-b68c-1e97724efa68","added_by":"auto","created_at":"2026-02-05 13:49:00","extension":"docx","order_by":1,"title":"","display":"","copyAsset":false,"role":"supplement","size":4164465,"visible":true,"origin":"","legend":"","description":"","filename":"SupplementalEvaluatingLargeLanguageModelsforAutomatedICD10CodingofObstetricClinicalNotesinPortugueseAComparativeStudy.docx","url":"https://assets-eu.researchsquare.com/files/rs-8712058/v1/b27f6e4fb471ce5885e9f3fd.docx"}],"financialInterests":"No competing interests reported.","formattedTitle":"\u003cp\u003eLarge Language Models for Automated Icd-10 Coding of Obstetric Clinical Notes in Portuguese: Comparison With Human Coders\u003c/p\u003e","fulltext":[{"header":"Introduction","content":"\u003cp\u003eMedical coding is fundamental to health systems, as it converts clinical documentation into standardized codes that underpin patient care, billing, research, and public health reporting\u003csup\u003e\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e\u003c/sup\u003e. Automating ICD coding has long been a central goal in natural language processing (NLP)\u003csup\u003e\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e\u003c/sup\u003e. The recent progress of large language models (LLMs) has fueled expectations that clinical coding might be solved off the shelf \u003csup\u003e\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e\u003c/sup\u003e. However, published evaluations consistently show that LLMs attain low accuracy when predicting codes.\u003csup\u003e\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e\u003c/sup\u003e\u003c/p\u003e \u003cp\u003eMost prior studies on automated ICD coding implicitly assume that higher F1-scores reflect better clinical performance. However, medical coding is inherently subjective, with substantial inter-annotator disagreement even among trained professionals. As a result, traditional set-based metrics such as F1-score conflate model error with human disagreement, obscuring the true clinical value of automated systems.\u003c/p\u003e \u003cp\u003eAgainst this backdrop, we benchmark multiple LLMs for hierarchical ICD-10 coding of Portuguese obstetric notes, compare them with a fine-tuned specialist, test whether Portuguese-to-English translation improves performance, and contextualize results against human inter-annotator agreement.\u003c/p\u003e \u003cp\u003eIn Brazilian Portuguese, systematic investigations of automated ICD coding remain scarce. Foundational Portuguese-language resources such as BRATECA and SemClinBr have advanced the local clinical NLP ecosystem, yet they have not been used to benchmark automated ICD-10 assignment on discharge summaries.\u003csup\u003e\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e\u003c/sup\u003e\u003c/p\u003e \u003cp\u003ePrior Portuguese-language efforts either relied on pre-Transformer, classical methods to predict multiple ICD-10 codes from Brazilian clinical notes, or pursued different tasks (e.g., clinical assistants and note generation) rather than code assignment. Meanwhile, transformer-based ICD-10 coding has been examined for death certificates in European Portuguese, not for hospital discharge notes.\u003csup\u003e\u003cspan additionalcitationids=\"CR18 CR19\" citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e\u003c/sup\u003e\u003c/p\u003e \u003cp\u003eIt is also essential to situate automation against realistic human performance. Disagreement among professional coders is common, reflecting ambiguities in clinical documentation, overlapping categories, and local conventions; automated systems should therefore be evaluated against this inherent variability.\u003csup\u003e\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e22\u003c/span\u003e\u003c/sup\u003e\u003c/p\u003e \u003cp\u003eSpecifically, we address four gaps that are rarely examined together: (i) the realistic performance ceiling imposed by human disagreement; (ii) the behavior of general-purpose LLMs under hierarchical ICD-10 constraints; (iii) the limited transferability of lightweight fine-tuning strategies; and (iv) the questionable validity of conventional evaluation metrics for clinical coding.\u003c/p\u003e"},{"header":"Methods","content":"\u003cdiv id=\"Sec3\" class=\"Section2\"\u003e \u003ch2\u003eData Source and Reference Standard\u003c/h2\u003e \u003cp\u003eThe corpus comprised 1,117 discharge summaries from a public, university-affiliated tertiary hospital specializing in women\u0026rsquo;s health in Brazil. These documents were retrospectively obtained and re-labeled by trained clinical coders specifically for this study, following a standardized protocol to ensure consistency and to establish the reference standard. Additional corpus statistics and note-length distributions are reported in the Supplementary Materials (S1.1; Table \u003cspan refid=\"MOESM1\" class=\"InternalRef\"\u003eS1\u003c/span\u003e and Figures \u003cspan refid=\"MOESM1\" class=\"InternalRef\"\u003eS1\u003c/span\u003e\u0026ndash;S2).\u003c/p\u003e \u003cp\u003eFor secondary analyses requiring blinded human agreement, we selected a stratified subset of 315 summaries. This subset was constructed to maintain proportional representation of the major ICD-10 obstetric categories and to ensure sufficient variability in case complexity, enabling reliable estimation of inter-annotator disagreement.\u003c/p\u003e \u003c/div\u003e\n\u003ch3\u003eTask Definition and Label Space\u003c/h3\u003e\n\u003cp\u003eWe framed coding as multi-label assignment of ICD-10 codes from full, free-text notes. Performance was assessed at two levels:\u003c/p\u003e \u003cp\u003e \u003col\u003e \u003cspan\u003e \u003cli\u003e \u003cp\u003eLeaf level (full code with subcategory, e.g., \u003cem\u003eO24.4\u003c/em\u003e), and\u003c/p\u003e \u003c/li\u003e \u003c/span\u003e \u003cspan\u003e \u003cli\u003e \u003cp\u003eThree-character category (e.g., \u003cem\u003eO24\u003c/em\u003e), used to quantify hierarchical specificity.\u003c/p\u003e \u003c/li\u003e \u003c/span\u003e \u003c/ol\u003e \u003c/p\u003e \u003cp\u003eAll codes were normalized to uppercase, whitespace was stripped, and duplicates removed. Out-of-scope/invalid tokens were discarded prior to scoring. We did not cap the number of codes a model could output per note.\u003c/p\u003e\n\u003ch3\u003eModels and Inference\u003c/h3\u003e\n\u003cp\u003eWe evaluated six general-purpose LLMs plus one specialist variant: GPT-4o\u003csup\u003e1\u003c/sup\u003e, GPT-4o-mini\u003csup\u003e\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e\u003c/sup\u003e, Sabi\u0026aacute;-3.1\u003csup\u003e4\u003c/sup\u003e, DeepSeek-V3\u003csup\u003e5\u003c/sup\u003e, Gemini 1.5 Flash\u003csup\u003e\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e\u003c/sup\u003e, and a fine-tuned GPT-4o-mini. All models were accessed via vendor APIs with temperature\u0026thinsp;=\u0026thinsp;0 (deterministic decoding); remaining parameters followed vendor defaults unless stated otherwise. Inference was executed during June 10\u0026ndash;11, 2025 (UTC); model names and API versions are listed in the Supplement to mitigate model-drift concerns. Exact model identifiers, endpoints, and inference dates are provided in the Supplementary Materials (S1.2; Table \u003cspan refid=\"MOESM2\" class=\"InternalRef\"\u003eS2\u003c/span\u003e).\u003c/p\u003e\n\u003ch3\u003eSpecialist fine-tuned variant\u003c/h3\u003e\n\u003cp\u003eWe fine-tuned gpt-4o-mini-2024-07-18 to emit JSON-only ICD-10 outputs; the provider snapshot lists 2 epochs, batch 16, LR-multiplier 1.8, seed 1, and 2,325,050 trained tokens (job ftjob-wqj9laPJ3wHdBc75fVmpGEgU). Full metadata and the step log (1,557 steps; training loss 3.303\u0026rarr;0.025) appear in Supplement Materials (S1.3; Figure \u003cspan refid=\"MOESM3\" class=\"InternalRef\"\u003eS3\u003c/span\u003e)\u003c/p\u003e \u003cp\u003eImportantly, the objective of this fine-tuning experiment was not to maximize performance through extensive hyperparameter optimization, data augmentation, or validation-based early stopping. Instead, it was designed to test a commonly proposed lightweight specialization strategy\u0026mdash;training on isolated code\u0026ndash;description pairs\u0026mdash;and to evaluate whether such an approach transfers to multi-label coding of free-text clinical narratives.\u003c/p\u003e\n\u003ch3\u003ePrompting and Output Format\u003c/h3\u003e\n\u003cp\u003eFor each model, we requested JSON-only outputs listing ICD-10 codes (no explanations). Where supported, we enforced structured output and validated the response schema with Pydantic (v2). For DeepSeek-V3, which does not expose structured output in the \u003cem\u003echat-completions\u003c/em\u003e route we used, we instructed JSON in the prompt and then parsed the returned text with conservative tokenization and our ICD-10 regex normalization. After validation/normalization, outputs were stored as Python lists of codes in the analysis DataFrame (one column per model). Full prompts are provided in the Supplement S1.4.\u003c/p\u003e \u003cdiv id=\"Sec8\" class=\"Section2\"\u003e \u003ch2\u003ePortuguese-to-English Translation Condition\u003c/h2\u003e \u003cp\u003eWe translated each Portuguese discharge note to English using Google Translate accessed via the googletrans\u003csup\u003e\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e\u003c/sup\u003e 4.0.0-rcl Python client, with default settings (one-pass, no stochastic sampling). Each translated note was then paired with an English version of the coding prompt\u0026mdash;a faithful translation of the Portuguese prompt preserving the same JSON schema and decoding parameters\u0026mdash;and submitted to the model. We applied the same post-processing pipeline as in the Portuguese condition (parsing, normalization, ICD-10 regex validation, de-duplication). No human post-editing of translations was performed.\u003c/p\u003e \u003c/div\u003e\n\u003ch3\u003eHuman Agreement\u003c/h3\u003e\n\u003cp\u003eTo contextualize model performance against a realistic human benchmark, we conducted a blinded clinician validation on the full test set of 315 obstetric discharge notes. For each note we built a candidate pool equal to the union of (i) the gold-standard ICD-10 codes and (ii) all codes proposed by the six model arms. A custom web interface (Streamlit) displayed each candidate code with its human-readable description; the reviewer was blinded to source (gold vs. model and model identity) and instructed to mark each code as clinically appropriate (TRUE) or not appropriate (FALSE) for that note. Multiple TRUE selections were allowed per note; notes were de-identified. Further details on the human validation workflow, including annotation steps and a representative screenshot of the annotation interface, are provided in the Supplementary Materials (Section S1.5; Figure \u003cspan refid=\"MOESM4\" class=\"InternalRef\"\u003eS4\u003c/span\u003e).\u003c/p\u003e \u003cdiv id=\"Sec10\" class=\"Section2\"\u003e \u003ch2\u003eStatistical Analysis\u003c/h2\u003e \u003cp\u003eModel performance was primarily evaluated using precision, recall, and the F1-score, calculated at both the three-character category and the exact leaf levels. We report both micro- and macro-averaged metrics. Micro-averaged metrics were calculated globally by aggregating the counts of true positives, false positives, and false negatives across the entire dataset. Macro-averaged metrics were calculated by first computing the metric for each individual note, then taking the unweighted average of these per-note scores.\u003c/p\u003e \u003cp\u003eTo quantify the uncertainty in our performance estimates for Tables\u0026nbsp;\u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e, \u003cspan refid=\"Tab3\" class=\"InternalRef\"\u003e3\u003c/span\u003e, and \u003cspan refid=\"Tab4\" class=\"InternalRef\"\u003e4\u003c/span\u003e, we computed 95% confidence intervals (CIs) using non-parametric bootstrap resampling. For each metric, we generated 10,000 bootstrap samples by resampling the 1,117 notes with replacement. The 95% CI was then derived from the 2.5th and 97.5th percentiles of the resulting bootstrap distribution. Bootstrap distributions for all reported metrics are shown in the Supplementary Materials (S2.2; Figures \u003cspan refid=\"MOESM5\" class=\"InternalRef\"\u003eS5\u003c/span\u003e\u0026ndash;S17).\u003c/p\u003e \u003cp\u003eFor the clinician-validated precision reported in Table\u0026nbsp;\u003cspan refid=\"Tab5\" class=\"InternalRef\"\u003e5\u003c/span\u003e, which represents a binomial proportion, the 95% CI was calculated using the Wilson score interval, a method robust for proportions, including those near the boundaries of 0 or 1.\u003c/p\u003e \u003c/div\u003e"},{"header":"Results","content":"\u003cdiv id=\"Sec12\" class=\"Section2\"\u003e \u003ch2\u003eCohort and label characteristics\u003c/h2\u003e \u003cp\u003eWe analyzed \u003cb\u003e1,117\u003c/b\u003e obstetric discharge notes. The gold standard contained a median of \u003cb\u003e2\u003c/b\u003e ICD-10 codes per note (IQR \u003cb\u003e[2\u0026ndash;3]\u003c/b\u003e), spanning \u003cb\u003e261\u003c/b\u003e three-character categories and \u003cb\u003e452\u003c/b\u003e leaf codes. Across model outputs, distinct code diversity ranged from \u003cb\u003e490\u003c/b\u003e to \u003cb\u003e978\u003c/b\u003e at the leaf level and \u003cb\u003e242\u003c/b\u003e to \u003cb\u003e481\u003c/b\u003e at the category level. Additional per-model descriptive statistics are provided in the Supplementary Materials (S2.1; Table \u003cspan refid=\"MOESM4\" class=\"InternalRef\"\u003eS4\u003c/span\u003e).\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eBars show, for GOLD (human annotations) and the six models, the number of unique codes observed in the corpus (n\u0026thinsp;=\u0026thinsp;1,117 notes). In \u0026ldquo;Exact (leaf level)\u0026rdquo;, we count all leaf-level codes\u0026mdash;including three-character categories and dotted subcategories\u0026mdash;after normalization and within-note deduplication. In \u0026ldquo;Category (3 chars)\u0026rdquo;, we count only three-character categories (A00\u0026ndash;Z99). Values above bars are the distinct counts. Greater diversity reflects broader code coverage, not necessarily higher accuracy. GOLD\u0026thinsp;=\u0026thinsp;human reference. All models were run with the same prompt and temperature\u0026thinsp;=\u0026thinsp;0.\u003c/p\u003e \u003cp\u003e \u003cb\u003eFigure 1 \u0026mdash; Diversity of distinct ICD-10 codes by set.\u003c/b\u003e \u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab1\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eDescriptive statistics for the number of ICD-10 codes per note by model and GOLD.Values are reported as median (IQR) and mean\u0026thinsp;\u0026plusmn;\u0026thinsp;SD over 1,117 discharge notes. Counts are computed at the leaf level (three-character categories and dotted subcategories), after normalization and within-note deduplication. All models were run deterministically (temperature\u0026thinsp;=\u0026thinsp;0) with the same prompt. GOLD\u0026thinsp;=\u0026thinsp;human reference. Q1\u0026thinsp;=\u0026thinsp;25th percentile; Q3\u0026thinsp;=\u0026thinsp;75th percentile.\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"7\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c7\" colnum=\"7\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eModel\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003en_notes\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003emedian\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eq1\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eq3\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c6\"\u003e \u003cp\u003emean\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c7\"\u003e \u003cp\u003esd\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSabi\u0026aacute;-3.1\u003csup\u003e4\u003c/sup\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e1117\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e6.0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e5.0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e7.0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e6.24\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e1.90\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eDeepSeek-V3\u003csup\u003e5\u003c/sup\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e1117\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e5.0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e5.0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e6.0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e5.54\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e1.35\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eGPT-4o-Mini\u003csup\u003e\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e\u003c/sup\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e1117\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e5.0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e5.0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e6.0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e5.27\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.99\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eGPT-4o\u003csup\u003e1\u003c/sup\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e1117\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e4.0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e3.0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e5.0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e4.42\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e1.36\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eGemini-1.5 Flash\u003csup\u003e\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e\u003c/sup\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e1117\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e5.0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e4.0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e6.0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e4.93\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e1.34\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eGold\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e1117\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e2.0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e2.0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e3.0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e2.58\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e1.39\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec13\" class=\"Section2\"\u003e \u003ch2\u003ePrimary Endpoint: Three-Character Category\u003c/h2\u003e \u003cp\u003eOn the primary endpoint (micro-F1 at the three-character category level), GPT-4o achieved the best performance with a score of 0.37 (95% CI 0.35\u0026ndash;0.38). It was followed by GPT-4o-Mini (0.34, 95% CI 0.33\u0026ndash;0.36) and DeepSeek-V3 (0.34, 95% CI 0.33\u0026ndash;0.35); the substantial overlap in their confidence intervals suggests their performance was statistically comparable. The remaining models followed in descending order: Sabi\u0026aacute;-3.1 (0.31, 95% CI 0.30\u0026ndash;0.33), Gemini 1.5 Flash (0.23, 95% CI 0.22\u0026ndash;0.25), and the fine-tuned GPT-4o-Mini (0.21, 95% CI 0.20\u0026ndash;0.22). Full micro- and macro-averaged metrics for all models, including precision and recall, are detailed in Table\u0026nbsp;\u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab2\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 2\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eICD-10 Coding Performance by Model \u0026mdash; Portuguese, Three-Character Category Level.\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"7\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c7\" colnum=\"7\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003eModelo\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colspan=\"3\" nameend=\"c4\" namest=\"c2\"\u003e \u003cp\u003eMicro\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colspan=\"3\" nameend=\"c7\" namest=\"c5\"\u003e \u003cp\u003eMacro\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003ePrecision\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eRecall\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eF1\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003ePrecision\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c6\"\u003e \u003cp\u003eRecall\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c7\"\u003e \u003cp\u003eF1\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSabi\u0026aacute;-3.1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.2255 (0.2159\u0026ndash;0.2354)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.5192 (0.5005\u0026ndash;0.5378)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.3145 (0.3025\u0026ndash;0.3267)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.2323 (0.2221\u0026ndash;0.2425)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.5467 (0.5259\u0026ndash;0.5673)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.3095 (0.2974\u0026ndash;0.3216)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eDeepSeek-V3\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.2514 (0.2408\u0026ndash;0.2619)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.5330 (0.5146\u0026ndash;0.5512)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.3417 (0.3288\u0026ndash;0.3544)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.2496 (0.2390\u0026ndash;0.2601)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.5431 (0.5224\u0026ndash;0.5638)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.3285 (0.3157\u0026ndash;0.3411)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eFine-Tuned GPT-4o-Mini\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.1593 (0.1505\u0026ndash;0.1681)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.3110 (0.2940\u0026ndash;0.3283)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.2107 (0.1994\u0026ndash;0.2218)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.1689 (0.1599\u0026ndash;0.1780)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.3648 (0.3440\u0026ndash;0.3852)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.2180 (0.2068\u0026ndash;0.2292)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eGPT-4o-mini\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.2585 (0.2478\u0026ndash;0.2689)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.5199 (0.5022\u0026ndash;0.5380)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.3453 (0.3326\u0026ndash;0.3576)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.2597 (0.2490\u0026ndash;0.2705)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.5462 (0.5258\u0026ndash;0.5663)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.3353 (0.3228\u0026ndash;0.3474)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eGPT-4o\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.2933 (0.2805\u0026ndash;0.3062)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.4918 (0.4739\u0026ndash;0.5096)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.3675 (0.3531\u0026ndash;0.3818)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.3000 (0.2867\u0026ndash;0.3131)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.5036 (0.4830\u0026ndash;0.5241)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.3596 (0.3450\u0026ndash;0.3742)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eGemini-1.5 Flash\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.1799 (0.1702\u0026ndash;0.1897)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.3349 (0.3183\u0026ndash;0.3525)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.2341 (0.2222\u0026ndash;0.2462)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.1811 (0.1713\u0026ndash;0.1910)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.3567 (0.3371\u0026ndash;0.3768)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.2297 (0.2178\u0026ndash;0.2416)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec14\" class=\"Section2\"\u003e \u003ch2\u003eLeaf Level\u003c/h2\u003e \u003cp\u003eRequiring leaf-level specificity substantially reduced performance for all systems. As detailed in Table\u0026nbsp;\u003cspan refid=\"Tab3\" class=\"InternalRef\"\u003e3\u003c/span\u003e, the micro-F1 score for the top-performing model, \u003cb\u003eGPT-4o\u003c/b\u003e, dropped from 0.37 to \u003cb\u003e0.15 (95% CI 0.14\u0026ndash;0.16)\u003c/b\u003e. This degradation was consistent across all models, with relative performance drops (ΔF₁) ranging from \u003cb\u003e\u0026minus;\u0026thinsp;52.5%\u003c/b\u003e for DeepSeek-V3 to a stark\u0026thinsp;\u003cb\u003e\u0026minus;\u0026thinsp;77.7%\u003c/b\u003e for the fine-tuned model. These results indicate that precise subcategory resolution (e.g., O24.4 vs. O24) remains a significant challenge.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab3\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 3\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eICD-10 Coding Performance by Model \u0026mdash; Portuguese, Leaf Level Layout identical to Table\u0026nbsp;\u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e, but here each full code (including subcategories) is counted separately.F₁ = percentage change in Micro-F₁ from the three-character category level to the four-character leaf level.\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"8\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c7\" colnum=\"7\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c8\" colnum=\"8\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colspan=\"8\" nameend=\"c8\" namest=\"c1\"\u003e \u003cp\u003eΔ\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003e\u003cb\u003eModel\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colspan=\"3\" nameend=\"c4\" namest=\"c2\"\u003e \u003cp\u003e\u003cb\u003eMicro\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colspan=\"3\" nameend=\"c7\" namest=\"c5\"\u003e \u003cp\u003e\u003cb\u003eMacro\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003eΔF₁\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ePrecision\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eRecall\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eF1-Score\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003ePrecision\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eRecall\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003eF1-Score\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSabi\u0026aacute;-3.1\u003csup\u003e4\u003c/sup\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e0.0798 (0.0737\u0026ndash;0.0861)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.1916 (0.1775\u0026ndash;0.2064)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.1126 (0.1042\u0026ndash;0.1213)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.0822 (0.0756\u0026ndash;0.0890)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e0.2060 (0.1891\u0026ndash;0.2234)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e0.1114 (0.1028\u0026ndash;0.1202)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e-64.20%\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eDeepSeek-V3\u003csup\u003e5\u003c/sup\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e0.1193 (0.1115\u0026ndash;0.1275)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.2543 (0.2394\u0026ndash;0.2702)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.1624 (0.1524\u0026ndash;0.1730)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.1171 (0.1091\u0026ndash;0.1253)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e0.2433 (0.2265\u0026ndash;0.2608)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e0.1523 (0.1422\u0026ndash;0.1628)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e-52.47%\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eFine-Tuned GPT-4o-mini\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e0.0341 (0.0297\u0026ndash;0.0387)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.0748 (0.0652\u0026ndash;0.0849)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.0469 (0.0409\u0026ndash;0.0531)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.0379 (0.0328\u0026ndash;0.0432)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e0.0890 (0.0766\u0026ndash;0.1020)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e0.0503 (0.0438\u0026ndash;0.0571)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e-77.74%\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eGPT-4o-mini\u003csup\u003e\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e\u003c/sup\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e0.1043 (0.0964\u0026ndash;0.1123)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.2116 (0.1968\u0026ndash;0.2269)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.1398 (0.1296\u0026ndash;0.1501)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.1046 (0.0964\u0026ndash;0.1128)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e0.2100 (0.1936\u0026ndash;0.2272)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e0.1334 (0.1235\u0026ndash;0.1435)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e-59.51%\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eGPT-4o\u003csup\u003e1\u003c/sup\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e0.1169 (0.1077\u0026ndash;0.1265)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.1992 (0.1845\u0026ndash;0.2142)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.1474 (0.1361\u0026ndash;0.1589)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.1186 (0.1091\u0026ndash;0.1284)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e0.1942 (0.1782\u0026ndash;0.2110)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e0.1410 (0.1299\u0026ndash;0.1525)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e-59.89%\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eGemini-1.5 Flash\u003csup\u003e\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e\u003c/sup\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e0.0653 (0.0589\u0026ndash;0.0719)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.1241 (0.1121\u0026ndash;0.1363)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.0855 (0.0773\u0026ndash;0.0941)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.0642 (0.0576\u0026ndash;0.0709)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e0.1225 (0.1089\u0026ndash;0.1364)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e0.0811 (0.0727\u0026ndash;0.0895)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e-63.48%\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec15\" class=\"Section2\"\u003e \u003ch2\u003ePortuguese-to-English translation\u003c/h2\u003e \u003cp\u003eTo assess whether translating notes into English could improve ICD-10 coding, we ran a translation arm in which each Portuguese discharge summary was translated via Google Translate (deterministic configuration) and then submitted\u0026mdash;using an English-language version of our standard JSON‐only prompt\u0026mdash;to the same three models: Sabi\u0026aacute;-3.1, DeepSeek-V3, and GPT-4o. Before selecting Google Translate, we benchmarked three PT\u0026rarr;EN engines\u0026mdash;Google Translate, the GPT-4o-mini API, and the Sabi\u0026aacute; API\u0026mdash;while holding the coder fixed (Sabi\u0026aacute;-3.1); Google Translate achieved the highest scores and was therefore adopted for the English-input arm (see Supplementary Table \u003cspan refid=\"MOESM5\" class=\"InternalRef\"\u003eS5\u003c/span\u003e). Table\u0026nbsp;\u003cspan refid=\"Tab4\" class=\"InternalRef\"\u003e4\u003c/span\u003e reports micro‐precision, micro‐recall, and micro\u0026ndash;F₁ at both leaf (exact code) and three-character category levels, along with ΔF₁ (defined as F₁English\u0026thinsp;\u0026minus;\u0026thinsp;F₁Portuguese). Sabi\u0026aacute;-3.1 saw a small improvement on English notes (leaf ΔF₁ = +0.01410; category ΔF₁ = +0.0006). DeepSeek-V3 performance declined after translation (leaf ΔF₁ = \u0026minus;\u0026thinsp;0.0380; category ΔF₁ = \u0026minus;\u0026thinsp;0.0278). GPT-4o also dropped (leaf ΔF₁ = \u0026minus;\u0026thinsp;0.0251; category ΔF₁ = \u0026minus;\u0026thinsp;0.0436).\u003c/p\u003e \u003cp\u003eOverall, translating into English did not yield consistent gains\u0026mdash;and in two of three systems degraded coding accuracy\u0026mdash;underscoring that native-language prompts remain preferable (Table\u0026nbsp;\u003cspan refid=\"Tab4\" class=\"InternalRef\"\u003e4\u003c/span\u003e). Exploratory role-specific precision for the instructed primary/secondary code roles is reported in the Supplementary Materials (S2.3; Table \u003cspan refid=\"MOESM6\" class=\"InternalRef\"\u003eS6\u003c/span\u003e).\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab4\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 4\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eComparative performance of Sabi\u0026aacute;-3.1, DeepSeek-V3, and GPT-4o on English-translated vs Native-portuguese Notes For each condition (Leaf level and Three-character category) we report micro-precision, micro-recall, and micro-F₁. F₁ indicates the change in micro-F₁ resulting from translation and is defined as the micro-F₁ score on the English-translated notes minus the micro-F₁ score on the original Portuguese notes.\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"9\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c7\" colnum=\"7\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c8\" colnum=\"8\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c9\" colnum=\"9\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colspan=\"9\" nameend=\"c9\" namest=\"c1\"\u003e \u003cp\u003eΔ\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003e\u003cb\u003eLevel\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003e\u003cb\u003eModel\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colspan=\"3\" nameend=\"c5\" namest=\"c3\"\u003e \u003cp\u003e\u003cb\u003eMicro\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colspan=\"3\" nameend=\"c8\" namest=\"c6\"\u003e \u003cp\u003e\u003cb\u003eMacro\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c9\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003eΔF₁\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e\u003cb\u003ePrecision\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e\u003cb\u003eRecall\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e\u003cb\u003eF1\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e\u003cb\u003ePrecision\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e\u003cb\u003eRecall\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003e\u003cb\u003eF1\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"2\" rowspan=\"3\"\u003e \u003cp\u003eThree Character\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eGPT-4o\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.2553 (0.2444\u0026ndash;0.2663)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.4428 (0.4243\u0026ndash;0.4616)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.3239 (0.3109\u0026ndash;0.3371)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e0.2571 (0.2458\u0026ndash;0.2685)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e0.4828 (0.4620\u0026ndash;0.5042)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003e0.3199 (0.3066\u0026ndash;0.3332)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c9\"\u003e \u003cp\u003e-0.0436\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eSabi\u0026aacute;-3.1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.2337 (0.2239\u0026ndash;0.2438)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.4836 (0.4648\u0026ndash;0.5030)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.3151 (0.3031\u0026ndash;0.3276)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e0.2323 (0.2223\u0026ndash;0.2424)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e0.5192 (0.4986\u0026ndash;0.5402)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003e0.3052 (0.2933\u0026ndash;0.3175)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c9\"\u003e \u003cp\u003e0.0006\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eDeepSeek-V3\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.2357 (0.2263\u0026ndash;0.2456)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.4698 (0.4512\u0026ndash;0.4891)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.3139 (0.3021\u0026ndash;0.3262)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e0.2327 (0.2230\u0026ndash;0.2426)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e0.5080 (0.4875\u0026ndash;0.5291)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003e0.3047 (0.2929\u0026ndash;0.3169)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c9\"\u003e \u003cp\u003e-0.0278\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"2\" rowspan=\"3\"\u003e \u003cp\u003eLeaf Level\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eGPT-4o\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.0857 (0.0781\u0026ndash;0.0932)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.1510 (0.1376\u0026ndash;0.1644)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.1094 (0.0997\u0026ndash;0.1188)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e0.0873 (0.0793\u0026ndash;0.0951)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e0.1626 (0.1470\u0026ndash;0.1784)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003e0.1089 (0.0992\u0026ndash;0.1185)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c9\"\u003e \u003cp\u003e-0.0251\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eSabi\u0026aacute;-3.1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.0935 (0.0865\u0026ndash;0.1006)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.1969 (0.1825\u0026ndash;0.2113)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.1267 (0.1176\u0026ndash;0.1361)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e0.0930 (0.0860\u0026ndash;0.1002)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e0.2096 (0.1925\u0026ndash;0.2267)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003e0.1227 (0.1136\u0026ndash;0.1318)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c9\"\u003e \u003cp\u003e0.0141\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eDeepSeek-V3\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.1021 (0.0949\u0026ndash;0.1094)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.2093 (0.1949\u0026ndash;0.2239)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.1373 (0.1278\u0026ndash;0.1467)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e0.1005 (0.0931\u0026ndash;0.1079)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e0.2140 (0.1974\u0026ndash;0.2308)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003e0.1311 (0.1218\u0026ndash;0.1406)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c9\"\u003e \u003cp\u003e-0.0380\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec16\" class=\"Section2\"\u003e \u003ch2\u003eClinician validation\u003c/h2\u003e \u003cp\u003eIn a blinded clinician validation of 315 obstetric discharge notes, the gold standard achieved a human-validated precision (PPV_H) of 0.77 (95% CI 0.74\u0026ndash;0.79). Among the AI systems, GPT-4o led with PPV_H\u0026thinsp;=\u0026thinsp;0.59 (95% CI 0.56\u0026ndash;0.62), followed by DeepSeek-V3 at 0.54 (95% CI 0.51\u0026ndash;0.56), GPT-4o-mini at 0.43 (95% CI 0.40\u0026ndash;0.45), Sabi\u0026aacute;-3.1 at 0.35 (95% CI 0.33\u0026ndash;0.37), Gemini 1.5 Flash at 0.32 (95% CI 0.29\u0026ndash;0.34), and the fine-tuned GPT-4o-mini at 0.30 (95% CI 0.28\u0026ndash;0.33) (Table\u0026nbsp;\u003cspan refid=\"Tab5\" class=\"InternalRef\"\u003e5\u003c/span\u003e). The \u0026ldquo;gold gap\u0026rdquo; (proportion of gold-standard codes rejected by the clinician) was 23%, highlighting substantial intra-human variability.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab5\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 5\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eClinician-Validated Precision of ICD-10 Code Predictions by Model\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"7\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c7\" colnum=\"7\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eModel\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eTrues\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eFalse\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eTotal Codes\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003ePrecision\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c6\"\u003e \u003cp\u003e95% CI LOWER\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c7\"\u003e \u003cp\u003e95% CI UPPER\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSabi\u0026aacute;-3.1\u003csup\u003e4\u003c/sup\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e688\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e1300\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e1988\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.3461\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.3255\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.3673\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eDeepSeek-V3\u003csup\u003e5\u003c/sup\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e937\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e806\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e1743\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.5376\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.5141\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.5609\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eFine-Tuned GPT-4o-mini\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e532\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e1217\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e1749\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.3042\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.2831\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.3261\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eGPT-4o-mini\u003csup\u003e\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e\u003c/sup\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e703\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e949\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e1652\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.4255\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.4019\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.4495\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eGPT-4o\u003csup\u003e1\u003c/sup\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e819\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e569\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e1388\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.5901\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.5640\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.6156\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eGemini-1.5 Flash\u003csup\u003e\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e\u003c/sup\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e481\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e1042\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e1523\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.3158\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.2930\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.3396\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eGold\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e624\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e190\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e814\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.7666\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.7363\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.7944\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec17\" class=\"Section2\"\u003e \u003ch2\u003eFine-tuned specialist\u003c/h2\u003e \u003cp\u003eWe evaluated whether light-weight supervised fine-tuning of GPT-4o-mini on single-label code\u0026ndash;description pairs would improve multi-label coding of free-text discharge notes. The fine-tuned model generated a distribution of codes per note (median 5.0, IQR 5.0\u0026ndash;6.0; mean 5.69\u0026thinsp;\u0026plusmn;\u0026thinsp;2.06) that closely mirrored other systems (Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e), but did not match the more conservative human reference (median 2.0, IQR 2.0\u0026ndash;3.0).\u003c/p\u003e \u003cp\u003eDespite this apparent alignment in output volume, fine-tuning failed to yield any accuracy gains. At the three-character category level, the specialist achieved micro-F₁ = 0.19 versus 0.36 for the best zero-shot model (GPT-4o), with similarly low macro-F₁ and no improvement in precision or recall (Table\u0026nbsp;\u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e). Clinician-validated precision for the fine-tuned model was only 0.30 (95% CI 0.28\u0026ndash;0.33), well below both the GPT-4o zero-shot system (0.59, 95% CI 0.56\u0026ndash;0.62) and the gold standard (0.77, 95% CI 0.74\u0026ndash;0.79) (Table\u0026nbsp;\u003cspan refid=\"Tab5\" class=\"InternalRef\"\u003e5\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eThese results indicate that direct supervised fine-tuning on isolated code\u0026ndash;description pairs does not transfer effectively to narrative clinical notes and offers no advantage in hierarchical specificity.\u003c/p\u003e \u003c/div\u003e"},{"header":"Discussion","content":"\u003cp\u003eIn this study, we systematically evaluated six large language models (LLMs) for automated ICD-10 coding of Portuguese obstetric discharge notes, comparing them against a human reference standard and a fine-tuned specialist variant. Even the best performer, GPT-4o, achieved only a micro-F₁ of 0.36 at the three-character category level and 0.15 at the full (leaf) level. Translating notes into English conferred no performance benefit, and lightweight fine-tuning of GPT-4o-mini on description\u0026ndash;code pairs likewise failed to improve outcomes.\u003c/p\u003e \u003cp\u003eOur findings align with English-language benchmarks reporting modest off-the-shelf LLM performance on clinical coding tasks. For example, Dong et al. observed micro-F₁ scores below 0.50 for general discharge-summary coding\u003csup\u003e\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e\u003c/sup\u003e, and Soroush et al. characterized standard LLMs as \u0026ldquo;poor medical coders\u0026rdquo; in code-query settings\u003csup\u003e\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e\u003c/sup\u003e. Portuguese‐language studies to date have been limited to pre-Transformer techniques\u003csup\u003e\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e\u003c/sup\u003e or as ICD-10 assignment on death certificates\u003csup\u003e\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e\u003c/sup\u003e, which involve shorter, more standardized text, leaving a gap for Brazilian clinical notes that our work now fills.\u003c/p\u003e \u003cp\u003eWe observed a gold gap of 23%\u0026mdash;the proportion of official codes the clinician rejected as non-applicable\u0026mdash;a figure below the 30\u0026ndash;40% disagreement typically reported in inter-coder reliability studies, in which two independent coders assign codes de novo and are then compared\u003csup\u003e\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e22\u003c/span\u003e\u003c/sup\u003e. Although our protocol relied on validation against a pre-labeled set, the observed discrepancy reaffirms inherent human variability and sets a practical ceiling for automated systems. Nevertheless, GPT-4o achieved a clinician-validated precision of 59% (vs. 77% for the gold standard), demonstrating competitive zero-shot performance on Portuguese text. These results reinforce that LLMs should be regarded as assistance tools\u0026mdash;proposing codes for expert review\u0026mdash;rather than direct replacements for human coders. Moreover, the lack of improvement with English translation highlights the importance of developing and evaluating models directly in the native clinical language.\u003c/p\u003e \u003cp\u003eThis single-center obstetrics study may not generalize to other specialties or settings. Our fine-tuning protocol omitted data augmentation, validation splits, and early stopping\u0026mdash;strategies shown to boost micro-F₁ to ~\u0026thinsp;0.69 on English clinical notes through multi-phase, objective-aligned training\u003csup\u003e\u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e\u003c/sup\u003e. We also did not explore few-shot prompting, chain-of-thought techniques, or retrieval-augmented architectures, all of which warrant investigation.\u003c/p\u003e \u003cp\u003eIn summary, while LLMs exhibit promise in medical natural language processing, their application to ICD-10 coding in Brazilian Portuguese should be positioned as coders\u0026rsquo; assistants, supported by robust specialization pipelines and rigorous human oversight to ensure safety and reliability in real-world clinical practice.\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003eEthics approval and consent to participate This study was approved by the Research Ethics Committee (Comitê de Ética em Pesquisa – CEP) of the University of Campinas (UNICAMP), Brazil, under approval number 7.446.157 (CAAE: 86136424.6.0000.5404). The study was conducted in accordance with national and institutional ethical standards and with the Declaration of Helsinki. It is a retrospective study based on anonymized electronic medical records, with no direct patient contact or intervention. The requirement for informed consent was waived by the Ethics Committee due to the use of fully anonymized retrospective data.\u003c/p\u003e\u003cp\u003e \u003ch2\u003eConflict of Interest\u003c/h2\u003e \u003cp\u003eThe authors declare no competing interests.\u003c/p\u003e \u003c/p\u003e\u003ch2\u003eFunding\u003c/h2\u003e \u003cp\u003eThis research did not receive any specific grant from funding agencies in the public, commercial, or not-for-profit sectors.\u003c/p\u003e\u003ch2\u003eAuthor Contribution\u003c/h2\u003e\u003cp\u003eR.S.S., M.G.G., and C.T. conceived the study and designed the experimental framework. R.S.S. implemented the computational pipeline, conducted the statistical analyses, and drafted the initial version of the manuscript. M.G.G. and C.T. contributed to the methodological design, interpretation of the results, and critical revision of the manuscript. P.M.F. and A.G.L. contributed to data curation, clinical interpretation, and validation of the reference annotations. R.C.P. contributed to clinical oversight, interpretation of obstetric outcomes, and refinement of the clinical evaluation protocol. All authors reviewed and approved the final manuscript.\u003c/p\u003e\u003ch2\u003eAcknowledgement\u003c/h2\u003e\u003cp\u003eThe authors thank CAISM \u0026ndash; Hospital da Mulher Prof. Dr. Jos\u0026eacute; Aristodemo Pinotti (UNICAMP) for providing access to the clinical data used in this study.\u003c/p\u003e\u003ch2\u003eData Availability\u003c/h2\u003e\u003cp\u003eThe code and scripts used to perform data processing, model evaluation, and statistical analysis in this study are publicly available in the following GitHub repository: https://github.com/ricardosantoss/llm-icd10-pt-obstetrics. This repository includes all code necessary to reproduce the experiments reported in the manuscript, including preprocessing, model inference, evaluation metrics, and visualizations. Dependencies and execution instructions are provided in the repository README. De-identified clinical data used in the study cannot be shared publicly due to ethical constraints; see the Data Availability statement for details on accessing de-identified data upon reasonable request.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eOpenAI. GPT-4o System Card. OpenAI; 2024. Available at: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://openai.com/index/gpt-4o-system-card/\u003c/span\u003e\u003cspan address=\"https://openai.com/index/gpt-4o-system-card/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eOpenAI. GPT-4o mini: Advancing cost-efficient intelligence. OpenAI; July 18, 2024. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://openai.com/index/gpt-4o-mini-advancing-cost-efficient-intelligence/\u003c/span\u003e\u003cspan address=\"https://openai.com/index/gpt-4o-mini-advancing-cost-efficient-intelligence/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eOpenAI. gpt-4o-mini \u0026mdash; Model reference. OpenAI Platform Docs; 2024\u0026ndash;2025. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://platform.openai.com/docs/models/gpt-4o-mini\u003c/span\u003e\u003cspan address=\"https://platform.openai.com/docs/models/gpt-4o-mini\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAbonizio H, Almeida TS, Laitz T, et al. Sabi\u0026aacute;-3 Technical Report. arXiv. 2024; DOI:\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.48550/arXiv.2410.12049\u003c/span\u003e\u003cspan address=\"10.48550/arXiv.2410.12049\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLiu A, Feng B, Xue B, et al. DeepSeek-V3 Technical Report. arXiv. 2024; DOI:\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.48550/arXiv.2412.19437\u003c/span\u003e\u003cspan address=\"10.48550/arXiv.2412.19437\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGeorgiev P, Lei VI, Burnell R, et al. Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context. arXiv. 2024; DOI:\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.48550/arXiv.2403.05530\u003c/span\u003e\u003cspan address=\"10.48550/arXiv.2403.05530\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTeo. SS. Googletrans: Free and Unlimited Google translate API for Python. PyPI; 2020. (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://pypi.org/project/googletrans/)\u003c/span\u003e\u003cspan address=\"https://pypi.org/project/googletrans/)\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWorld Health Organization. International Statistical Classification of Diseases and Related Health Problems, 10th Revision (ICD-10). Geneva; WHO; (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://icd.who.int/browse10\u003c/span\u003e\u003cspan address=\"https://icd.who.int/browse10\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDong H, Falis M, Whiteley W, et al. Automated clinical coding: what, why, and where we are? Npj Digital Medicine. 2022. DOI:\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1038/s41746-022-00705-7\u003c/span\u003e\u003cspan address=\"10.1038/s41746-022-00705-7\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eYan C, Fu X, Liu X, et al. A survey of automated International Classification of Diseases coding: development, challenges, and applications. Intelligent Medicine. 2022; DOI: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1016/j.imed.2022.03.003\u003c/span\u003e\u003cspan address=\"10.1016/j.imed.2022.03.003\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eNori H, King N, McKinney SM, Carignan D, Horvitz E. Capabilities of GPT-4 on medical challenge problems. arXiv. 2023; DOI: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.48550/arXiv.2303.13375\u003c/span\u003e\u003cspan address=\"10.48550/arXiv.2303.13375\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSinghal K, Tu T, Gottweis J, et al. Towards expert-level medical question answering with large language models. Nature Medicine. 2025. DOI:\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1038/s41591-024-03423-7\u003c/span\u003e\u003cspan address=\"10.1038/s41591-024-03423-7\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSoroush A, Glicksberg BS, Zimlichman E, et al. Large Language Models Are Poor Medical Coders \u0026mdash; Benchmarking of Medical Code Querying. NEJM AI. 2024;. DOI:\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1056/AIdbp2300040\u003c/span\u003e\u003cspan address=\"10.1056/AIdbp2300040\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKaur R, Ginige JA, Obst O. AI-based ICD coding and classification approaches using discharge summaries: a systematic literature review. Expert Systems with Applications. 2023; DOI:\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1016/j.eswa.2022.118997\u003c/span\u003e\u003cspan address=\"10.1016/j.eswa.2022.118997\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eConsoli B, Santos HDP, Ulbrich AHDPS, Vieira R, Bordini RH. BRATECA (Brazilian Tertiary Care Dataset): a Clinical Information Dataset for the Portuguese Language. In: Proceedings of LREC 2022. European Language Resources Association; 2022. (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://aclanthology.org/2022.lrec-1.602/\u003c/span\u003e\u003cspan address=\"https://aclanthology.org/2022.lrec-1.602/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eOliveira LES, Peters AC, Silva AMP, et al. SemClinBr\u0026mdash;A multi-institutional and multi-specialty semantically annotated corpus for Portuguese clinical NLP tasks. Journal of Biomedical Semantics. 2022; DOI:\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1186/s13326-022-00269-1\u003c/span\u003e\u003cspan address=\"10.1186/s13326-022-00269-1\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eReys AD, Silva D, Severo D, Pedro S, S\u0026aacute; MMS, Salgado GAC. Predicting Multiple ICD-10 Codes from Brazilian-Portuguese Clinical Notes. In: Proceedings of the 30th Brazilian Conference on Intelligent Systems (BRACIS). 2020. DOI: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1007/978-3-030-61377-8_39\u003c/span\u003e\u003cspan address=\"10.1007/978-3-030-61377-8_39\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePaiola PH, Garcia GL, Manesco JRR, Roder M, Rodrigues D, Papa JP. Adapting LLMs for the Medical Domain in Portuguese: A Study on Fine-Tuning and Model Evaluation. In: Proceedings of the WSCG Conference. 2025. (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://wscg.zcu.cz/WSCG2025/papers/A13.pdf\u003c/span\u003e\u003cspan address=\"http://wscg.zcu.cz/WSCG2025/papers/A13.pdf\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePinto JGS, Freitas AR, Martins AC, Sawazaki CM, Vidal C, Oliveira LE. Developing resource-efficient clinical LLMs for Brazilian Portuguese. In: Proceedings of the 34th Brazilian Conference on Intelligent Systems (BRACIS). 2024. DOI: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1007/978-3-031-79038-6_4\u003c/span\u003e\u003cspan address=\"10.1007/978-3-031-79038-6_4\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCoutinho I, Martins B. Transformer-based models for ICD-10 coding of death certificates with Portuguese text. Journal of Biomedical Informatics. 2022. DOI:\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1016/j.jbi.2022.104232\u003c/span\u003e\u003cspan address=\"10.1016/j.jbi.2022.104232\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWockenfuss R, Frese T, Herrmann K, Claussnitzer M, Sandholzer H. Three- and four-digit ICD-10 is not a reliable classification system in primary care. Scand J Prim Health Care. 2009. DOI:\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1080/02813430903072215\u003c/span\u003e\u003cspan address=\"10.1080/02813430903072215\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eStausberg J, Lehmann N, Kaczmarek D, Stein M. Reliability of diagnoses coding with ICD-10. Int J Med Inform. 2008. DOI:\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1016/j.ijmedinf.2006.11.005\u003c/span\u003e\u003cspan address=\"10.1016/j.ijmedinf.2006.11.005\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHou Z, Liu H, Bian J, He X, Yan Z. Enhancing medical coding efficiency through domain-specific fine-tuned large language models. npj Health Systems. 2025. DOI:\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1038/s44401-025-00018-3\u003c/span\u003e\u003cspan address=\"10.1038/s44401-025-00018-3\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"","lastPublishedDoi":"10.21203/rs.3.rs-8712058/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-8712058/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003ch2\u003eBackground\u003c/h2\u003e \u003cp\u003eDespite rapid advances in large language models (LLMs), automated ICD-10 coding of real-world clinical narratives remains unreliable. A key challenge lies not only in model limitations, but in the intrinsic ambiguity of clinical documentation and the substantial variability among human coders.\u003c/p\u003e\u003ch2\u003eMethods\u003c/h2\u003e \u003cp\u003eWe benchmarked six general-purpose LLMs for hierarchical ICD-10 coding of 1,117 obstetric discharge summaries written in Brazilian Portuguese. Model performance was evaluated at both category and leaf levels and contextualized against blinded clinician validation to quantify realistic human agreement. We further assessed whether Portuguese-to-English translation or lightweight supervised fine-tuning improved performance.\u003c/p\u003e\u003ch2\u003eResults\u003c/h2\u003e \u003cp\u003eEven the best-performing model (GPT-4o) achieved only modest agreement with the human reference, with micro-F1 scores of 0.36 at the three-character level and 0.15 at the leaf level. Translation into English did not yield consistent gains, and direct fine-tuning on code\u0026ndash;description pairs failed to improve accuracy. In clinician validation, human-coded references achieved a precision of 0.77, compared to 0.59 for the strongest model, revealing a substantial gap between automated predictions and clinically accepted codes.\u003c/p\u003e\u003ch2\u003eConclusions\u003c/h2\u003e \u003cp\u003eOur findings indicate that current LLMs remain below human-level reliability for autonomous ICD-10 coding in Portuguese. More importantly, they show that conventional evaluation metrics such as F1-score substantially misrepresent clinical usefulness by conflating model error with human disagreement. The primary bottleneck for automated medical coding is therefore not model capacity, but the imperfect and variable human gold standard itself. LLMs should be positioned as decision-support tools that assist, rather than replace, expert clinical coders.\u003c/p\u003e","manuscriptTitle":"Large Language Models for Automated Icd-10 Coding of Obstetric Clinical Notes in Portuguese: Comparison With Human Coders","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2026-02-05 13:48:54","doi":"10.21203/rs.3.rs-8712058/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"6640ae45-7760-469c-a5fc-973aba37eaa5","owner":[],"postedDate":"February 5th, 2026","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[{"id":62388436,"name":"Biological sciences/Computational biology and bioinformatics"},{"id":62388437,"name":"Health sciences/Diseases"},{"id":62388438,"name":"Health sciences/Health care"},{"id":62388439,"name":"Physical sciences/Mathematics and computing"},{"id":62388440,"name":"Health sciences/Medical research"}],"tags":[],"updatedAt":"2026-03-20T14:55:44+00:00","versionOfRecord":[],"versionCreatedAt":"2026-02-05 13:48:54","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-8712058","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-8712058","identity":"rs-8712058","version":["v1"]},"buildId":"XKTyCvWXoU3ODBz1xrDgd","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.