EndoExtract: Co-Designing Structured Text Extraction from Endometriosis Ultrasound Reports

preprint OA: green CC0
AI-generated summary by claude@2026-06, 2026-06-14

EndoExtract is an LLM system co-designed with research assistants to extract structured data from unstructured endometriosis ultrasound reports, addressing workflow pain points by prioritizing human review of interpretive fields and highlighting source evidence.

One-sentence paraphrase of the abstract; not a substitute for reading it. No clinical advice. How this works

AI-generated deep summary by claude@2026-06, 2026-06-14 · read from full text

The paper studies how to extract structured, interpretable data from unstructured free-text endometriosis ultrasound reports to support downstream analytics, machine learning training, and clinical auditing. Using contextual inquiry with research assistants, the authors identify workflow pain points including asymmetrical trust between numerical and interpretive fields, repetitive manual highlighting and fatigue from sustained comparison, and terminology inconsistency across radiologists; these findings inform EndoExtract, an on-premise LLM-powered system that automatically highlights evidence in PDFs and separates batch extraction from mandatory human-paced verification of interpretive fields. In a formative workshop, participants reported that the tool shifts work from field-by-field entry to supervisory validation, while noting risks of over-skimming and challenges handling missing data. This paper is centrally about endometriosis — it proposes and evaluates EndoExtract for structuring data from endometriosis ultrasound reports.

Read from the paper's body, not the abstract. Not a substitute for reading the paper. No clinical advice. How this works

Abstract

Endometriosis ultrasound reports are often unstructured free-text documents that require manual abstraction for downstream tasks such as analytics, machine learning model training, and clinical auditing. We present \textbf{EndoExtract}, an on-premise LLM-powered system that extracts structured data from these reports and surfaces interpretive fields for human review. Through contextual inquiry with research assistants, we identified key workflow pain points: asymmetric trust between numerical and interpretive fields, repetitive manual highlighting, fatigue from sustained comparison, and terminology inconsistency across radiologists. These findings informed an interface that surfaces only interpretive fields for mandatory review, automatically highlights source evidence within PDFs, and separates batch extraction from human-paced verification. A formative workshop revealed that \textbf{EndoExtract} supports a shift from field-by-field data entry to supervisory validation, though participants noted risks of over-skimming and challenges in managing missing data.
Full text 2,647 characters · extracted from oa-doi-fallback · click to expand
Computer Science > Human-Computer Interaction [Submitted on 26 Jan 2026 (v1), last revised 14 Feb 2026 (this version, v3)] Title:EndoExtract: Co-Designing Structured Text Extraction from Endometriosis Ultrasound Reports View PDF HTML (experimental)Abstract:Endometriosis ultrasound reports are often unstructured free-text documents that require manual abstraction for downstream tasks such as analytics, machine learning model training, and clinical auditing. We present \textbf{EndoExtract}, an on-premise LLM-powered system that extracts structured data from these reports and surfaces interpretive fields for human review. Through contextual inquiry with research assistants, we identified key workflow pain points: asymmetric trust between numerical and interpretive fields, repetitive manual highlighting, fatigue from sustained comparison, and terminology inconsistency across radiologists. These findings informed an interface that surfaces only interpretive fields for mandatory review, automatically highlights source evidence within PDFs, and separates batch extraction from human-paced verification. A formative workshop revealed that \textbf{EndoExtract} supports a shift from field-by-field data entry to supervisory validation, though participants noted risks of over-skimming and challenges in managing missing data. Submission history From: Haiyi Li [view email][v1] Mon, 26 Jan 2026 05:17:32 UTC (7,483 KB) [v2] Thu, 29 Jan 2026 06:14:41 UTC (7,483 KB) [v3] Sat, 14 Feb 2026 10:38:52 UTC (7,483 KB) References & Citations Loading... Bibliographic and Citation Tools Bibliographic Explorer (What is the Explorer?) Connected Papers (What is Connected Papers?) Litmaps (What is Litmaps?) scite Smart Citations (What are Smart Citations?) Code, Data and Media Associated with this Article alphaXiv (What is alphaXiv?) CatalyzeX Code Finder for Papers (What is CatalyzeX?) DagsHub (What is DagsHub?) Gotit.pub (What is GotitPub?) Hugging Face (What is Huggingface?) ScienceCast (What is ScienceCast?) Demos Recommenders and Search Tools Influence Flower (What are Influence Flowers?) CORE Recommender (What is CORE?) arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: oa-doi-fallback

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Condition tags

endometriosis

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2026) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

openalex
last seen: 2026-06-04T00:00:01.174412+00:00
License: CC0 · commercial use OK