A systematic review of trial-matching pipelines using large language models

preprint OA: closed
Full text JSON View at publisher

Abstract

Abstract Matching patients to clinical trial options is critical for identifying novel treatments, especially in oncology. However, manual matching is labor-intensive and error-prone, leading to recruitment delays. Pipelines incorporating large language models (LLMs) offer a promising solution. We conducted a systematic review of studies published between 2020 and 2025 from three academic databases and one preprint server, identifying LLM-based approaches to clinical trial matching. Of 126 unique articles, 31 met inclusion criteria. Reviewed studies focused on matching patient-to-criterion only (n = 4), patient-to-trial only (n = 10), trial-to-patient only (n = 2), binary eligibility classification only (n = 1) or combined tasks (n = 14). Sixteen used synthetic data; fourteen used real patient data; one used both. Variability in datasets and evaluation metrics limited cross-study comparability. In studies with direct comparisons, the GPT-4 model consistently outperformed other models—even finely-tuned ones—in matching and eligibility extraction, albeit at higher cost. Promising strategies included zero-shot prompting with proprietary LLMs like the GPT-4o model, advanced retrieval methods, and fine-tuning smaller, open-source models for data privacy when incorporation of large models into hospital infrastructure is infeasible. Key challenges include accessing sufficiently large real-world data sets, and deployment-associated challenges such as reducing cost, mitigating risk of hallucinations, data leakage, and bias. This review synthesizes progress in applying LLMs to clinical trial matching, highlighting promising directions and key limitations. Standardized metrics, more realistic test sets, and attention to cost-efficiency and fairness will be critical for broader deployment.
Full text 217,319 characters · extracted from preprint-html · click to expand
A systematic review of trial-matching pipelines using large language models | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Article A systematic review of trial-matching pipelines using large language models Braxton A. Morrison, Madhumita Sushil, Jacob S. Young This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-8036235/v1 This work is licensed under a CC BY 4.0 License Status: Under Review Version 1 posted 11 You are reading this latest preprint version Abstract Matching patients to clinical trial options is critical for identifying novel treatments, especially in oncology. However, manual matching is labor-intensive and error-prone, leading to recruitment delays. Pipelines incorporating large language models (LLMs) offer a promising solution. We conducted a systematic review of studies published between 2020 and 2025 from three academic databases and one preprint server, identifying LLM-based approaches to clinical trial matching. Of 126 unique articles, 31 met inclusion criteria. Reviewed studies focused on matching patient-to-criterion only (n = 4), patient-to-trial only (n = 10), trial-to-patient only (n = 2), binary eligibility classification only (n = 1) or combined tasks (n = 14). Sixteen used synthetic data; fourteen used real patient data; one used both. Variability in datasets and evaluation metrics limited cross-study comparability. In studies with direct comparisons, the GPT-4 model consistently outperformed other models—even finely-tuned ones—in matching and eligibility extraction, albeit at higher cost. Promising strategies included zero-shot prompting with proprietary LLMs like the GPT-4o model, advanced retrieval methods, and fine-tuning smaller, open-source models for data privacy when incorporation of large models into hospital infrastructure is infeasible. Key challenges include accessing sufficiently large real-world data sets, and deployment-associated challenges such as reducing cost, mitigating risk of hallucinations, data leakage, and bias. This review synthesizes progress in applying LLMs to clinical trial matching, highlighting promising directions and key limitations. Standardized metrics, more realistic test sets, and attention to cost-efficiency and fairness will be critical for broader deployment. Biological sciences/Cancer Biological sciences/Computational biology and bioinformatics Health sciences/Health care Physical sciences/Mathematics and computing Health sciences/Medical research Clinical trial matching large language models LLMs systematic review GPT-4 patient recruitment automated matching clinical trials generative artificial intelligence patient-trial matching trial-patient matching patient-criterion matching Figures Figure 1 Figure 2 Figure 3 Introduction Clinical trials are essential for identifying novel treatment options and offering patients access to potentially life-saving therapies. This is particularly important in fields like oncology, where alternative therapeutic options may be limited. 1 – 3 One key stage of trial-matching includes patient recruitment, which represents ~ 32% of clinical trial cost and is the reason most frequently cited for the discontinuation of randomized controlled trials. 4 , 5 Another important stage is manual trial matching, which requires intensive review of complex eligibility criteria across extensive patient records to assess patient suitability for a given trial, a time-consuming and labor-intensive process. 6 – 8 This stage is also costly; one study reported a cost between $ 129.15 to $ 336.48 per enrolled patient. 8 Equity is another key issue with trial enrollment; minorities, elderly people and rural groups are often underrepresented in cancer trials, a problem which could potentially be addressed by screening all patients for a large number of trials. 9 Thus, screening more patients for eligibility across more trials could improve patient recruitment and equity, but cost, labor and time are core limiting factors to its implementation. Automated matching systems aim to address these inefficiencies, with early approaches relying on rigid, rule-based methods. One approach involved generating queries from trial eligibility criteria that could be applied to identify potentially eligible patients in clinical databases. 10 , 11 Another involved extracting key details from patient records and converting them to a structured format to filter lists of trials. 12 Earlier rule-based automation systems—such as the Watson for Clinical Trial Matching—demonstrated potential, as seen in a 2016 Mayo Clinic pilot showing an 80% increase in enrollment for breast cancer trials. 13 These systems often lacked generalizability across institutions, patient populations, or trial protocols. Recently, large language models (LLMs) have emerged as a flexible alternative capable of interpreting both structured and unstructured data, extracting eligibility criteria, and matching patients to trials with greater scalability. Synthesizing recent work from 2020 to 2025, we characterize datasets, common pipeline structures, comparative model performance, and evaluation strategies. We aim to provide a practical roadmap for researchers and implementers to understand the capabilities, limitations, and emerging best practices for deploying LLMs in this domain. Results Model Categories This review covered four main types of trial matching – “patient-criterion matching”, in which a patient is assessed for one or more criteria individually; “patient-to-trial,” through which a list of trials is determined for a given patient; “trial-to-patient,” through which a list of potentially eligible patients is provided for a given trial; and “binary classification”, in which it is determined whether a patient-trial pair is a match. Of the reviewed articles, four focused on matching patient-to-criterion only, ten on patient-to-trial only, two on trial-to-patient only, one on binary eligibility classification only and fourteen on combined tasks (Table 1 ). While in theory, a perfect match between a patient and all trial criteria should imply eligibility, real-world applications are more complex. Strict 100% patient-criterion match thresholds often exclude many patients who may still be eligible under fewer, more practical constraints. In a study by Gupta et al., they found that weighting more important criteria more heavily resulted in better performance in translating criterion-level assessments into patient-trial matches. 14 Table 1 Model Performance Assessment Study Study Data Set Main Model(s) Patient-to-Criterion Metrics Patient-to-Trial Metrics Trial-to-Patient Metrics Model Comparisons† Cost/Efficiency Beattie et al., 2024 26 − 2018 n2c2 GPT-3.5 Turbo; GPT-4 GPT-3.5 Turbo: - Accuracy: 0.81 - Sensitivity: 0.80 - Specificity: 0.82 - Micro F 1 : 0.79 GPT-4: - Accuracy: 0.87 - Sensitivity: 0.85 - Specificity: 0.89 - Micro F 1 : 0.86 N/A N/A N/A Model-processing time: 1–5 minutes per patient for all included criteria Beattie et al., 2024 27 - Article-specific patient and trial data GPT-3.5; GPT-4 GPT-3.5: - Accuracy: 0.761 - Sensitivity: 0.776 - Specificity: 0.732 - Youden Index: 0.357 GPT-4: - Accuracy: 0.838 - Sensitivity: 0.839 - Specificity: 0.830 - Youden Index: 0.668 GPT-3.5: - Median AUC: 0.64 - Median accuracy: 0.65 -Median sensitivity: 0.68 - Median specificity: 0.69 - Median Youden index: 0.27 GPT-4 - Median AUC: 0.74 - Median accuracy: 0.70 - Median sensitivity: 0.66 - Median specificity: 0.74 N/A Proprietary: GPT-3.5; GPT-4 Screening cost per patient: - GPT-3.5: $ 0.02- $ 0.03 - GPT-4: $ 0.15- $ 0.27. Screening time per patient: - GPT-3.5: 1.4-3 minutes - GPT-4: 7.9–12.4 minutes Cerami et al., 2024 28 - Article-specific patient and trial data TrialSpace text embedding + TrialChecker classification model (fine-tuned RoBERTa-Large; open-access) N/A DFCI trial enrollment dataset: - Precision@10: 0.89 - MAP@10: 0.93 DFCI standard-of-care dataset: - Precision@10: 0.87 - MAP@10: 0.91 DFCI trial enrollment dataset: - Precision@20: 0.91 - MAP@20: 0.93 DFCI standard-of-care dataset: - Precision@20: 0.87 - MAP@20: 0.90 N/A N/A Chowdhury et al., 2024 29 - Article-specific patient and trial data Siamese-Patient-Trial-Matching with LLaMA 2 to initialize embeddings of inputs (open-access) N/A Siamese-PTM (fine-grained): F 1 score of 0.92 *Note: article describes binary classification task of eligible/not eligible so could be labeled either “patient-trial” or “trial-patient” matching N/A N/A Datta et al., 2025 30 - TREC 2023 CT GPT-4 N/A - Precision@10: 0.7351 - NDCG@10: 0.8109 N/A N/A N/A Devi et al., 2024 32 - Article-specific patient and trial data GPT-3.5 Turbo N/A N/A Accuracy of GPT-3.5: No tuning tested on 20 data points: 100% Tuned and tested on 20 data points: 95% No tuning tested on 100 data points: 79% Tuned and tested on 50 data points: 82% N/A N/A Devi et al., 2024 31 31,32 - Article-specific patient and trial data GPT-4 N/A N/A Accuracy: No tuning, tested on 20 data points: 95% Tuned and tested on 20 data points: 100% No tuning, tested on 100 data points: 86% N/A N/A Ferber et al., 2024 33 - Article-specific patient and trial data GPT-4o Accuracy: 88.0% For 14/15 patients, their “target trial” was listed within the top 15 trial matches identified by the algorithm N/A N/A N/A Gueguen et al, 2025 34 TrialGPT to re-rank outputs from DigitalECMT TrialGPT with Qwen2.5-7B-Instruct as the LLM N/A Use of LLM-based re-ranking on results of DigitalECMT increased NDCG@3 from 0.61 to 0.64 N/A N/A N/A Gui et al, 2025 35 - Article-specific patient and trial data - Anthropo-morphized Experts’ Chain of Thought - LLMs converted eligibility criteria into questions: GPT-4o, Google Gemini Advanced, Anthropic Claude 3.5 Sonnet Pathway A, Majority Vote - Precision: ~0.921 - Recall: ~0.82 Pathway A, Majority Vote - Precision: ~0.922 - Recall: ~0.819 *Note: article describes binary classification task of eligible/not eligible so could be labeled either “patient-trial” or “trial-patient” matching - QWEN1.5, BAICHUAN, and GLM Pathway A, Majority Vote - Efficiency: ~0.44 seconds/task Gupta et al., 2024 14 - Article-specific patient and trial data OncoLLM (fine-tuned Qwen-1.5 14B model; open-access) Accuracy: - OncoLLM: 63% - GPT-3.5 Turbo: 53% - GPT-4: 68% - Qwen14B-Chat: 43% - Mitral-7B-Instruct: 41% - Mixtral-8 X 7B-Instruct: 49% - Meditron: 51% - MedLlama: 55% - TrialLlama: 57% Percentage of time model ranks ground truth trials in the top-3 positions among 10 considered trials - OncoLLM: 65% - GPT-3.5 Turbo (iterative): 61% Normalized Discounted Cumulative Gain - OncoLLM: 68% - GPT-3.5 Turbo: 62% - Proprietary: GPT-3.5 Turbo and GPT-4 - Open-source: Qwen14B-Chat, Mistral-7B-Instruct, Mixtral-8 × 7B-Instruct, Meditron, MedLlama, and TrialLlama Operating cost: - OncoLLM: ~ $ 170 - GPT-4: ~ $ 6055 Cost of a single patient-trial match: - OncoLLM: ~ $ 0.17 per patient-trial pair - GPT-4: ~ $ 6.18 per patient-trial pair Jin et al., 2024 22 - SIGIR 2016 - TREC 2021 CT - TREC 2022 CT TrialGPT (GPT-4) (zero-shot) TrialGPT-Matching: - Accuracy: 87.3% TrialGPT-Ranking: - NDCG@10: 0.7252 - Precision@10: 0.6724 N/A - Proprietary: GPT-3.5 - Open-source (for trial ranking): SciFive, BioBERT, PubMedBERT, SapBERT, and BioLinkBERT Reduction in screening time for patient recruitment: 42.6% Jullien et al., 2024 36 - TREC 2022 CT GPT-4 Turbo (proprietary) N/A - NDCG@10: 0.679 - Precision@10: 0.730 - Precision@25: 0.630 - MRR: 0.860 N/A - GPT-3.5 (proprietary) - TREC SOTA (open-source) *BM-25 for initial ranking N/A Kusa et al., 2023 37 - TREC 2021 CT - TREC 2022 CT TCRR with BioBERT (open-access) N/A - NDCG@5: 0.627 - NDCG@10: 0.604 - Precision@10: 0.482 - Reciprocal Rank: 0.672 N/A - Open-source: MonoBERT, TraditionalRR, TCRR initialized with bert-base-uncased , BioBERT, ClinicalBERT N/A Kusa et al., 2023 38 - TREC 2023 CT with structured data converted to unstructured text using GPT-3.5 TCRR neural re-ranking model with BlueBERT (open-access) and GPT-3.5 (zero-shot) to refine results N/A DoSSIER_3: - nDCG@5: 0.6653 - nDCG@10: 0.6837 - Precision@10: 0.5838 - Reciprocal Rank: 0.6421 N/A N/A N/A Lai et al., 2024 39 - Article-specific patient and trial data GPT-4o (zero-shot) GPT-4o: - Accuracy: 96.7% in agreeing with human raters on binary eligibility criteria GPT-4o: - Accuracy: 90.7% of eligible patient-trial matches - Sensitivity: 87.5–100% for 8/9 trials - Specificity: 73.3–100% for all 9 trials N/A N/A - Median cost to screen a single patient: $ 0.67 (range: $ 0.63- $ 0.74) - Median time elapsed per patient: 138 seconds (range: 130–146) - Median total token usage: 112,266.5 tokens (range: 102982.0-122174.2) Lin et al., 2024 40 - TREC 2021 CT - SIGIR 2016 - TrialAlign Panacea (finely-tuned Mistral-7B-Base model27; open-access) Yes; no metrics provided F1, precision, and recall N/A - Open-source: BioMistral, Mistral, Zephyr, LLAMA-2, MedAlpaca, Meditron N/A Nievas et al., 2024 23 - SIGIR 2016 - TREC CT 2021 - TREC CT 2022 Trial-LLAMA 70B (open-access) Trial-LLAMA 70B: - Implicit CLA: 68.77 - Explicit CLA: 59.9 GPT-4: - Implicit CLA: 75.31 - Explicit CLA: 58.8 Trial-LLAMA 70B: - NDCG@10: 0.6636 - Precision@10: 0.5886 - AUROC: 0.6528 - AURPC: 0.6515 GPT-4: - NDCG@10: 0.7728 - Precision@10: 0.7005 - AUROC: 0.7390 - AURPC: 0.7038 N/A - Proprietary: GPT-3.5 and GPT-4 - Open-source: LLAMA 7B, LLAMA 13B, and LLAMA 70B N/A Peikos et al., 2023 41 - TREC 2021 - TREC 2022 GPT-3.5 Turbo N/A TREC 2021 - R-precision: 0.212 - Binary preference: 0.275 - Precision@10: 0.323 - P- recision @25: 0.261 - MRR: 0.541 - nDCG@10: 0.512 TREC 2022 - R-precision: 0.276 - Binary preference: 0.298 - Precision @10: 0.372 - Precision @25: 0.338 - MRR: 0.576 - nDCG@10: 0.517 N/A N/A N/A Peikos et al, 2024 42 - TREC 2021 - TREC 2022 GPT-3.5 Turbo N/A TREC 2021** GPT-3.5: - NDCG@10: 0.486 - Binary preference: 0.213 - Precision@10: 0.276 - Recall@25: 0.115 - MRR: 0.440 Qwen2-7B-Instruct: - NDCG@10: 0.476 - Binary preference: 0.216 - Precision@10: 0.285 - Recall@25: 0.107 - MRR: 0.465 N/A -Proprietary: GPT-4 -Open access: Qwen2; Phi3-medium-4k-Instruct; Phi3-mini-4k-Instruct; Medical-Llama3-8B N/A Rahmanian et al, 2024 43 - n2c2 2018 GPT-3.5 Turbo Overall (micro): - F1: 0.9061 - AUC: 0.9035 Overall (macro): - F1: 0.8060 - AUC: 0.7949 N/A N/A N/A N/A Ruan et al., 2024 44 - Article-specific patient and trial data GPT-4 with a knowledge graph N/A GPT-4: - MSE: 27.27 - RMSE: 5.22 *Measures similarity to criteria of another trial identified by the patient N/A - Open-source: 11 models from the SBERT family N/A Rybinski et al., 2024 45 - TREC 2021 CT - TREC 2022 CT - TREC 2023 CT Fine-tuned GPT-3.5 Turbo with BM25 retrieval (zero-shot for final eligibility determination) N/A BM25 with TCRR and GPT-3.5 Turbo re-ranking: - nDCG@1000: 0.375 - nDCG@10: 0.777 - Precision@10: 0.697 - Reciprocal Rank: 0.783 BM25 with chain-of-thought GPT-4o: - nDCG@1000: 0.504 - nDCG@10: 0.785 - Precision@10: 0.603 - Reciprocal Rank: 0.844 N/A - Proprietary: GPT4-o - Cost of re-ranking with GPT-3.5 Turbo: ~ $ 0.25 per 100 API calls (per patient) Shi et al., 2024 15 − 2018 n2c2 - Synthesized a trial for which 28 patients from the 2018 n2c2 cohort would meet all criteria MAKA (proprietary) MAKA: - Accuracy: 0.909 - Precision: 0.822 - Recall: 0.846 - F 1 -score: 0.828 Strategy by Wornow et al.: - Accuracy: 0.884 - Precision: 0.727 - Recall: 0.0.894 - F 1 -score: 0.785 N/A MAKA: - Accuracy: 0.9306 - Precision: 0.6333 - Recall: 0.6786 - F 1 -score: 0.6552 - Strategies employed by Wornow et al and Beattie et al N/A Unlu et al., 2024 46 - Article-specific patient and trial data RAG-Enabled Clinical Trial Infrastructure for Inclusion Exclusion Review (RECTIFIER) – using GPT-4 Vision RECTIFIER: - Sensitivity: 75–100% - Specificity: 92.1–100% - PPV: 75–100% - MCC: 97.9–100% Study Staff: - Sensitivity: 66.7–100% - Specificity: 82.1–100% - PPV: 50–100% - MCC: 91.7–100% RECTIFIER - Sensitivity: 92.3% - Specificity: 93.9% - PPV: 98.1% - NPV: 78.6% - Accuracy: 92.7% - MCC: 81.3% Study staff - Sensitivity: 90.8% - Specificity: 83.6% - PPV: 94.9% - NPV: 73.0% - Accuracy: 89.1% - MCC: 71.1% N/A - Proprietary: GPT-3.5 RECTIFIER: - Individual-question approach: average of 11 cents/patient - Combined-question approach: average of 2 cents/patient Cost without RAG: - GPT-4: $ 15.88 per patient - GPT-3.5: $ 1.59 per patient Wong et al, 2023 47 - Article-specific patient and trial data GPT-4 (3-shot) Performed but metrics not provided GPT-3.5 (zero-shot) Precision: 88.5 Recall: 11.6 F1: 20.6 GPT-4 (zero-shot) Precision: 86.7 Recall: 46.8 F1: 60.8 GPT-4 (3-shot) Precision: 87.6 Recall: 67.3 F1: 76.1 *Note: article describes binary classification task of eligible/not eligible so could be labeled either “patient-trial” or “trial-patient” matching - Proprietary: GPT-3.5 N/A Woo, 2024 48 − 2018 n2c2 - Article-specific patient and trial data Llama-3.1-8B-All Llama-3.1-8B-All on MIMIC-IV-based data set: Balanced accuracy: 0.93 Micro-F1: 0.94 N/A N/A - Open-access: Llama-3.1-8B, Llama3.1-70B Cost of evaluating the Apixaban criteria for 10,000.patients: - Llama-3.1-8B: $ 929 - Llama3.1-70B: $ 4066 Wornow et al., 2024 24 − 2018 n2c2 − 1 novel exclusion criterion from a pulmonary arterial hypertension clinical trial - SIGIR 2016 Zero-shot GPT-4 with ACIN prompting strategy - Precision: 0.91 - Recall: 0.92 - Macro-F 1 : 0.81 - Micro-F 1 : 0.93 N/A N/A - Proprietary: GPT-3.5 - Open-source: Llama-2-70b, Mixtral-8x7B, Qwen2-72b, and Llama-3-70b - Cost to screen a single patient: ~ $ 1.55 - Evaluation time per patient: ~1 minute Yuan et al., 2024 49 - Article-specific patient and trial data LLM-based patient-trial matching with GPT-4 (LLM-PTM) - Precision: 0.964 - Recall: 0.862 - F 1 Score: 0.910 - Precision: 0.801 - Recall: 0.830 - F 1 Score: 0.815 N/A N/A N/A Zhuang et al., 2024 50 - TREC 2022 CT - TREC 2023 CT Hybrid PubmedBERT-based retriever with GPT-4 re-ranking N/A Bi-encoder dense retriever: NDCG@10: 0.5768 P@10: 0.3243 Recall@1000: 0.3670 PLADEv2 sparse retriever: NDCG@10: 0.5971 P@10: 0.3243 Recall@1000: 0.3482 Hybrid: - NDCG@10: 0.5763 - P@10: 0.2946 - Recall@1000: 0.3878 CE_weighted: - NDCG@10: 0.6716 - P@10: 0.4432 - Recall@1000: 0.3878 GPT-4: - NDCG@10: 0.7363 - P@10: 0.5108 - Recall@1000: 0.3878 N/A GPT-3.5 Turbo (proprietary) to generate extra training data N/A Zihang et al., 2025 51 - Article-specific patient and trial data llama3-70b-instruct Performed but metrics not provided Accuracy: - GLM-3-Turbo: 0.8973 - GLM-4: 0.9139 - llama3-70b-instruct: 0.9285 - Qwen-Turbo: 0.9166 N/A GLM-3-Turbo, GLM-4, Qwen-Turbo N/A * 2018 n2c2 denotes 2018 National Natural Language Processing Clinical Challenges cohort; ACIN, All criteria, individual notes: All notes are merged into a single prompt, but the model assesses one criterion at a time, requiring reprompting for each criterion; CLA: criterion-level accuracy; MAKA, Multi-Agents for Knowledge Augmentation; MCC = Mathew correlation coefficient; MRR = mean reciprocal rank; NCDG@10: Normalized Cumulative Discounted Gain; NDCG@10: Normalized Discounted Cumulative Gain; NPV: negative predictive value; PPV: positive predictive value; Program-rather-than-prompt: Ensures responses adhere to required format using structured programming objects instead of free-text prompts; QGMT, Query Generation, Medical Role & Task Description: Provides contextual information to ChatGPT and instructs it to generate a single keyword-based query; RAG, Retrieval-Augmented Generation; SIGIR 2016, Special Interest Group on Information Retrieval 2016 cohort; TCRR, Topical and Criteria Re-Ranking involves a two-step training schema focusing on topical relevance and eligibility classification; TREC 2021/2022/2023 CT, Text Retrieval Conference Clinical Trials Track. † Models compared within a given study include any models that were assessed by the authors alongside their best-performing LLM. **Additional metrics were calculated for TREC 2022 due to spacing concerns. Source: Data compiled from the studies included in this systematic review. Model Pipeline Design Most LLM-based trial-matching algorithms involve four stages: (1) acquisition of patient and trial data; (2) data pre-processing; (3) retrieval of relevant information from patient records; and (4) matching patients to trials (Fig. 2 A). Data Sets Pipeline inputs varied by study; patient data included case reports/vignettes, longitudinal prescription and medical claims data, medical records, and synthetic admission notes (Fig. 2 B; Table 2 ). Trial data were primarily drawn from ClinicalTrials.gov, pre-processed trial datasets, and hospital or international databases (Fig. 2 B; Table 3 ). Most data provided to model pipelines was unstructured text (Table 2 ). For articles where LLMs were fine-tuned, ground truth was generally derived from historical enrollments, researcher-determined matches or other LLMs (Table 2 – 3 ). Table 2 Patient Data Used for Training and Testing of LLMs Study Dataset Source of Data Data Description Data Size Patient Characteristics Data Type Publicly Available? 2018 n2c2 21 2014 i2b2/UTHealth shared tasks 52 , 53 Real, de-identified clinical notes; synthetic clinical trial 288 patients (2–5 records/patient); 13 predefined inclusion criteria Diabetes Unstructured text Yes Beattie et al., 2024 54 Enrollment list of phase II trial investigating hypo fractionated radiation therapy for head and neck cancer; head and neck radiation oncology team Real patient notes from surgical oncology, radiation oncology and medical oncology from the last 6 months; relevant pathology reports, imaging reports and lab results going back up to 1 year 35 patients enrolled in trial; 40 randomly identified patients seen by head and neck radiation oncology team Head and neck cancer Unstructured text Not Reported Cerami et al., 2024 28 Dana-Farber Cancer Institute (DFCI) Oncology Data Retrieval System - Retrospective trial enrollment data set: Real EHR data (all unstructured clinical notes, imaging reports, and pathology reports) for all adults who enrolled in cancer treatment clinical trials at DFCI from January 2016 to April 2024 - Standard of care (SOC) treatment dataset: Real data for patients who started SOC systemic therapies at DFCI from 2016 to 2024 - Retrospective trial enrollment data set: 16,139 enrollments for 13,425 patients onto 1,534 clinical trials - SOC treatment dataset: 86,042 treatment plans for 50,799 patients Cancer, including breast, lung, lymphoma and leukemia among others - Unstructured and structured text No Chowdhury et al., 2024 29 Mayo Clinic’s United Data Platform Real patient EHR data Structured EHR: Six types of clinical events (diagnosis, medication, allergy, family history of medical condition, lab tests, and admission (e.g., reason for visit) Unstructured EHR: radiology reports Demographics: age and gender. 180 patients 50 out of 180 patients were eligible for the clinical trial Unstructured and structured text No Devi et al., 2024 Synthetic case reports/patient descriptions 100 patients Non-small lung cancer Patient descriptions Unknown Ferber et al., 2024 32,33 Synthetic; based on fictional patient vignettes Synthetic oncology-focused patient EHRs (describe patient diagnoses, comorbidities, molecular data, imaging descriptions from staging CT or MRI scans, patient history) 51 records Cancer, including lung adenocarcinoma Unstructured text Yes Gueguen et al, 2025 34 Local Molecular Tumor Board at Centre Léon Bérard, France Real tumor board information on sequential patients; clinical data from full EHR and molecular data from molecular programmes 34 patients in the LLM subset Mixed adult solid tumors Unstructured text Yes; anonymized data available at: https://github.com/crcl-tm2/ trialmatch-tool-evaluation/blob/master/artifacts/data_raw/formatted_ data.csv Gui et al, 2025 35 The Hepatology Department of the First Affiliated Hospital of Guangxi University of Chinese Medicine Real, de-identified narrative admission notes 16,000 notes from over the last 10 years Hepatopathy Unstructured text Yes; can be obtained from corresponding author upon reasonable request Gupta et al., 2024 14 Institutional clinical research data warehouse from a single cancer center Real, de-identified EHR and clinical trial enrollment information, including: assessment & plan note, brief op note, consults, discharge instructions/ summary, h&p, op note, OR surgeon, procedures, progress notes, rad onc simulation, rad onc weekly review - Q&A Data: Notes from 50 patients - Clinical Data: Notes from over 5790 patients. - To evaluate patient-trial matching: 98 cancer patients - To evaluate trial-patient matching: ~ 1–3 patients who enrolled in each trial, 5–21 who did not Cancer, including breast and lung Unstructured text No Lai et al., 2024 39 The Pancreas Center Real de-identified medical oncology note, and surgical oncology note if available, from each patient who was screened for clinical trials at the Pancreas Center between January and May 2024 32 patients Pancreatic cancer; 19 out of 24 patients in the test set were eligible for at least one trial Unstructured text Not Reported SIGIR 2016 20 Adopted from TREC CDS 55 Real case reports 60 case reports Diagnosis varied Unstructured text Yes - Available at https://data.csiro.au/collection/csiro:17152 TREC 2021CT, 17 TREC 2022 CT 18 "cases created by individuals with medical training" Synthetic admissions notes - TREC 2021 CT: 75 topics - TREC 2022 CT: 50 topics Diagnosis varied Unstructured text Yes - TREC 2021 CT: Available at http://www.trec-cds.org/2021.html - TREC 2022 CT: Available at http://www.trec-cds.org/2022.html TREC 2023 CT 19 Synthetic Synthetic questionnaire data 451,538 documents for 50 patients Glaucoma, anxiety, COPD, breast cancer, COVID-19, rheumatoid arthritis, sickle cell anemia, type 2 diabetes Structured text Yes Available at https://www.trec-cds.org/2023.html Unlu et al., 2024 46 The ongoing Co-Operative Program for Implementation of Optimal Therapy in Heart Failure Trial Real EHR data over 2 years from ongoing trial including: progress notes, discharge summaries, history and physical, telephone encounters, notes to patients sent through the portal 1891 patients High rates of symptomatic heart failure Unstructured text No Wong, 2023 47 EHR from a collaborating health system NLP-based extraction on structured data from EHR 523 patient-trial enrollment pairs Patient-trial enrollment pairs pulled from historical enrollment data Structured text No Woo, 2024 (MIMIC-III based) 48 LLM-generated based on MIMIC-III 56 Llama-3.1-70B-Instruct was prompted to generate eligibility-criteria-based Q&A pairs based on discharge summaries from MIMIC-III 1,000 questions Diagnosis varied Structured text Yes https://physionet.org/content/mimic-ext-synth-trial-question/1.0.0/ Woo, 2024 (MIMIC-IV based) 48 Human-generated based on MIMIC-IV 57 Researchers wrote 23 questions similar to eligibility criteria modeled after the ARISTOTLE Apixaban vs. Warfarin trial, then manually annotated answers from each MIMIC-IV discharge summary 23 questions resembling eligibility criteria with a random sample of 100 patient notes from MIMIC-IV Diagnosis varied Structured text Yes https://physionet.org/content/mimic-iv-ext-apixaban-trial/1.0.0/ Yuan et al., 2024 49 Stroke patient database Real, longitudinal prescription and medical claims data (diagnoses, procedures, medications) Claims data for 825 patients Each patient enrolled in 1/6 stroke trials Unstructured text Not Reported Zhuang et al., 2024 50 ChatGPT TREC 2022 CT TREC 2023 CT - ChatGPT data: Patient descriptions generated for randomly selected clinical trials ~ 20,000 synthetic patient-trial pairs (25,000 total once they combined it with 5,000 patient description-trial pairs from TREC 2022 and 2023) Diagnosis varied Unstructured Text Not Reported Zihang et al., 2025 51 LLM-generated Simulated answers to eligibility questionnaires 85 patients Diagnosis varied Structured Available upon request *2018 n2c2 denotes 2018 National Natural Language Professing Clinical Challenges cohort; COPILOT-HF, Cooperative Program for ImpLementation of Optimal Therapy in Heart Failure; MGB, Mass General Brigham; MSKCC, Memorial Sloan Kettering Cancer Center; PAH, pulmonary arterial hypertension; RAG, retrieval-augmented generation; SIGIR 2016, Special Interest Group on Information Retrieval 2016 cohort; TCGA, The Cancer Genome Atlas; TREC 2021/2022/2023 CT, Text REtrieval Conference Clinical Trials Track Source: Data compiled from the studies included in this systematic review. Table 3 Clinical Trial Data Used for Training and Testing of LLMs Study Dataset Data Source Data Description Diagnoses Covered in Data Set Beattie et al., 2024 27 Phase II trial investigating hypofractionated radiation therapy for head and neck cancer 14 criteria identified based on trial inclusion/exclusion criteria for evaluation using the LLMs Head and neck cancer Cerami et al., 2024 28 Dana-Farber Harvard Cancer Center clinical trial database; ClinicalTrials.gov Therapeutic clinical trials open at Dana-Farber Harvard Cancer Center from January 2012 to June 2024; 500 trials open for cancer diagnoses on October 22, 2024 Cancer Chowdhury et al., 2024 29 ClinicalTrials.gov Five clinical trials (NCT02008357, NCT04468659, NCT02669433, NCT01767909, NCT02565511) Cognitive disorders, such as Alzheimer's disease and Dementia with Lewey Bodies Devi et al., 2024 32 Clinicaltrials.gov 10 U.S. drug-only interventional clinical trials Non Small Cell Lung Cancer Ferber et al., 2024 33 ClinicalTrials.gov 105,600 real clinical trials filtered for cancer collected on May 13, 2024 Cancer Gui et al, 2025 35 Six real-world clinical trials on hepatopathy: ChiCTR2100044187, NCT04353193, NCT04850534, NCT03911037, NCT01311167, NCT04021056 “Selected 58 criteria assessable using information typically documented in admission notes” chronic liver failure, cirrhosis, hepatocellular carcinoma Gupta et al., 2024 14 ClinicalTrials.gov For patient-trial matching: set of 10 trials for each patient, each for the same cancer type that were actively recruiting when patient enrolled in clinical trial (1/10 of those trials was one in which the patient actually enrolled) For trial-patient matching: Set of 36 clinical trials that all recruited patients from the same institution Cancer Lai et al., 2024 39 ClinicalTrials.gov Real trial data from nine ongoing clinical trials at the Pancreas Center Pancreatic cancer Lin et al., 2024 40 - ClinicalTrials.gov - PubMed Central - TrialAlign ClinicalTrials.gov: 467,944 trials, real data - ChiCTR (China): 76,186 trials, real data - TrialAlign: 14 sources of trial documents Diagnoses varied PROTECTOR1 dataset 58 ClinicalTrials.gov 764 phase 3 cancer trials Cancer Ruan et al., 2024 44 ClinicalTrials.gov 301 trials Fatty liver disease SIGIR 2016 20 ClinicalTrials.gov 204,855 trials Diagnosis varied TREC 2021CT, 17 TREC 2022 CT 18 ClinicalTrials.gov 375,581 clinical trial descriptions Diagnosis varied TrialAlign 40 14 sources, including ClinicalTrials.gov and ChiCTR (China) 793,279 trial documents 1,113,207 scientific papers related to clinical trials Diagnosis varied Unlu et al., 2024 46 Data from the ongoing Co-Operative Program for Implementation of Optimal Therapy in Heart Failure Trial in the Microsoft Dynamics 365 used by staff for patient screening 13 trial criteria Heart failure Wong, 2023 47 ClinicalTrials.gov 53 treatment-oriented, interventional trials Cancer Yuan et al., 2024 49 ClinicalTrials.gov 6 trials Stroke Zihang et al., 2025 51 ClinicalTrials.gov 579 clinical trial registration entries with Fudan University as the sponsor; 562 pre-recruitment questionnaires were generated based on 562 trials with complete data Diagnosis varied *MSKCC denotes Memorial Sloan Kettering Cancer Center. All articles that used TREC CT patient data sets pulled trial protocol data from clinicaltrials.gov unless otherwise specified. Source: Data compiled from the studies included in this systematic review. Diseases represented in these patient data sets include cancer, stroke, diabetes, cognitive disorders, liver disease, heart failure, or a mix of conditions (Table 2 ). No articles used non-text data like images, and only Unlu et al and Lai et al. used data from ongoing randomized clinical trials. Eleven studies utilized a selection of two or more clinical note types, which could be real or synthetic (Tables 1 – 2 ). For studies that used real EHR, selecting a subset of note types or from a set of note dates allowed them to select for only the most relevant notes while reducing the inputs to the LLM, increasing efficiency. Sixteen studies used publicly available synthetic or de-identified data (TREC CT 2021-2023 17–19 , SIGIR 2016 20 , 2018 n2c2 21 ), eleven used real patient data, and three used case reports/vignettes (Table 1 – 2 ). Notably, Woo et al used Llama-3.1-70B-Instruct to generate eligibility-criteria-based question/answer pairs based on discharge summaries from the publicly-available MIMIC-III data set, while Zihang et al used an LLM to generate simulated answers to patient questionnaires as their patient data inputs. Rybinski et al used the TREC CT 2021 and 2022 for training and validation, and TREC CT 2023 for testing. Since the latter includes data from earlier years, model performance metrics might be influenced by data leakage. Data Pre-Processing Data pre-processing pipelines, especially those used for pipelines run on clinical notes, generally included some combination of four stages: enrichment, query generation, retrieval, and restructuring (Fig. 2 C). Enrichment refers to the process of augmenting raw text with structured information that enhances semantic interpretability and downstream utility for computational models. In the context of trial matching, enrichment transforms unstructured patient or trial data into a machine-readable representation by extracting key medical concepts, standardizing terminology, and/or clarifying contextual meaning—thereby enabling accurate and efficient alignment between patient profiles and trial eligibility criteria. While enrichment implementation varied, it generally involved one or more of the following techniques: (1) Entity extraction : Identifies clinically meaningful phrases like "Stage IV non-small cell lung cancer" as discrete concepts to enable structured matching. (2) Concept normalization : Maps colloquial or alternate terms to standardized medical terminology—for example, recognizing “high blood pressure” as equivalent to “hypertension”—to prevent missed matches due to lexical variation. (3) Negation detection : Ensures that the model correctly interprets statements containing negation, such as interpreting “no history of diabetes" as the patient not having diabetes, so the patient is not matched to a diabetes-related trial. Once enriched, this structured data could be used to generate search queries that retrieve relevant patient or trial segments. Two key query optimization methods were common: (1) Query expansion : Adds related terms or synonyms to broaden the search scope (e.g., expanding “lung cancer” to include “NSCLC” or “pulmonary neoplasms”). (2) Query synthesis : Repackages the extracted and enriched data into coherent, task-specific prompts or structured queries suitable for the retrieval model. Retrieval involves extracting relevant portions of patient records based on the generated queries. By reducing the data ultimately input into the LLM, retrieval serves two purposes; (1) it keeps the data size within a given LLM’s context window and (2) it reduces computational costs by using a smaller, less computationally intensive model to reduce the data the larger, more computationally-intensive LLM ultimately needs to process. Retrieval-augmented generation (RAG) is commonly used to identify portions of the queried text most semantically similar to the query; however, this approach can lose chronological context, which is essential to properly analyze patient data. A variety of approaches could be used to circumvent this issue; for instance, one pipeline combined lexical (BM25) and semantic (MedCPT) retrieval strategies to better preserve clinical chronology. 22 Lexical retrieval served to identify relevant clinical trials by matching exact or near-exact keywords from a set of synthetic data from a given patient to the text in trial eligibility criteria, thereby capturing surface-level correspondences in terminology and phrasing. Semantic retrieval, on the other hand, encoded both patient data and trial descriptions into dense vector representations using pretrained language models, enabling the retrieval of conceptually aligned information even when lexical overlap is limited. For prompt-based trial matching pipelines, the last data pre-processing step is restructuring. This restructuring generally involves taking the retrieved chunks—usually consisting of one or multiple trial criteria and the corresponding extracted patient information—and reformatting them into a prompt provided to a model to request the model perform patient-criterion, patient-trial, or trial-patient eligibility. Pre-processed data is then usually fed into one of two matching pipeline types (Fig. 2 D): (1) Prompt-based : Prompts are provided to an LLM, the responses from which are subsequently integrated into article-specific scoring methods to generate ranked outputs. (2) Embedding-based : Patient and trial data were embedded in a shared vector space and directly ranked using similarity scores. LLMs are also applied to tasks beyond eligibility assessment or mapping into a shared vector space—they have been used to convert free text patient or trial data into structured formats like JSONs for use for downstream models; generate synthetic datasets for training smaller models; enrich data sets; and enhance retrieval through query generation, expansion, or synthesis (Fig. 2 E). Some models provided rationale for matching decisions or cited specific sentences to support their determination (Fig. 2 D). 22 – 24 This feature not only improved model performance, but increased transparency and the ability of end users to assess accuracy. Depending on the matching type, these pipelines are designed for different end-users — while patient-trial matching systems can generate a list of trials for a patient or their care team, trial-patient matching can help review patient data to extract a list of potentially eligible patients for coordinators of a particular trial (Fig. 2 E). Model Performance & Cost Analysis Although comparisons across studies were hindered by heterogeneous tasks and data, intra-study comparisons showed GPT-4 consistently outperformed traditional models for inclusion/exclusion extraction and patient-to-trial matching compared to open-access models, including ones that had been fine-tuned (Table 1 ). Only a few studies reported on cost and efficiency. In a study by Jin et al, their TrialGPT pipeline reduced clinician screening time by 42.6%. 22 Several studies reported costs associated with LLM-assisted trial matching lower than traditional human screening methods (Table 1 ). In a study by Gupta et al, a Qwen-1.5-14B-based pipeline, OncoLLM, incurred a per-patient cost of only $ 0.17 per patient-trial pair, as compared to the cost of $ 6.18 per pair incurred by GPT-4 (Table 1 ). For pipelines using either GPT-3.5 or GPT-4, screening cost per-patient ranged between $ 0.02 and $ 15.88 per patient-trial pair (Table 1 ). Of note—in work performed by Unlu et al, use of RAG methods and other strategies dropped that $ 15.88 value to only 2 cents/patient for GPT-4 (Table 1 ). For finely-tuned models, training cost is also a consideration; Gupta et al put the cost of training their smaller, fine-tuned model OncoLLM at $ 2,688. Overall, per-patient processing tended to be fast, with studies reporting between 1 and 12.4 minutes to evaluate a given patient for a trial (Table 1 ). Discussion There is a pressing need to improve patient screening for patient trials, both to accelerate scientific discovery and to expand access to potentially life-saving treatments. However, the expensive, time-consuming nature of patient matching limits the speed of this process and contributes to limited diversity of patients who ultimately enroll. LLM-based pipelines have the potential to offer a more flexible and scalable alternative to previous rule-based matching systems, reducing the workload for trial staff and facilitating clinical trial success. One barrier to implementation is that some models struggle with the nuanced and intricate nature of EHR data. Most trial-matching models are trained on data that is synthetic, simplified, abridged, or in structured data formats that, while more easily read by LLMs, do not fully represent the complexities of real-world, long-context EHR data. Their reliance on narrowly defined variables from trial criteria or patient records can limit their generalizability, and real-world applicability. Ideally, trial matching data sets should closely parallel the same real-world data that models will be used with when implemented in the clinic. A key barrier to achieving this goal is the high cost required to collect large, diverse datasets annotated by experts, especially since PHI concerns restrict data sharing. Despite these limitations, the success of models applied to diverse unstructured datasets, such as medical claims data and admission records, demonstrates the versatility of LLMs more broadly—a quality that is likely to increase with ongoing advancements in model architecture and capabilities. The effectiveness of trial-matching systems hinges not only on model architecture but also on the integration of stakeholder expertise into their design and evaluation. Developing successful pipelines requires collaboration with key stakeholders, such as trial coordinators, who can provide insights into how eligibility criteria are weighted in practice and how patients are evaluated for enrollment. This input can be used both to refine model performance and to ensure that trial-matching tools integrate seamlessly into existing workflows. Currently, GPT-4 offers the strongest performance in trial-matching tasks, with high zero-shot capabilities reducing the need for annotated datasets and advanced retrieval strategies helping to offset computational costs (Fig. 1 ). For institutions unable to deploy GPT-4 within HIPAA-compliant infrastructure, smaller LLMs present a viable alternative when fine-tuned for specific tasks, offering a balance between accuracy, cost, and data privacy ( Fig. 3 ). Using smaller models can also avoid issues with shifts in model performance and behavior over time, which can be seen with proprietary models like GPT-4. Standardization of evaluation metrics and evaluation on high-quality public benchmarks is crucial to facilitate comparison of model pipeline performance. For patient-to-criterion matching, metrics such as accuracy, sensitivity, specificity, recall and F 1 scores are appropriate. Similar metrics can be used for assessment of patient-to-trial or trial-to-patient matching capabilities, in addition to metrics like NDCG@10, Precision@10, AUROC and AURPC to assess ranking performance. Accuracy of LLM-generated explanations for eligibility determinations compared to qualified staff should also be determined. Assessing the performance of different pipelines on the same public benchmarks would also facilitate accurate comparisons. However, data memorization by LLMs is also a potential issue when using public data sets, so it is important to also use private data sets for model validation. Several steps can be taken to mitigate concerns of model errors and bias. To safeguard against model hallucinations, LLMs can be instructed to provide rationales for their matching decisions, along with explicit citations from the input. These explanations both improve model accuracy and facilitate the correction of model mistakes by human users. Trial staff should also remain a key component of patient assessment, ensuring critical clinical details are not overlooked. Moreover, human-in-the-loop validation studies are necessary to assess whether LLM-based matching pipelines improve human performance or change human behavior. Model performance across diverse patient groups should be assessed, and safeguards should be put in place to prevent models amplifying bias present in training data. Early trial matching systems show promise in ultimately reducing the time spent by staff in assessing eligibility, as well as the overall cost of trial recruitment. Future work should also assess the cost of data collection and annotation, infrastructure to host the model, personnel to design and maintain it, and the choice of cloud service. Recent advances in LLM-based systems offer improved generalizability and accuracy, with models being employed across diverse tasks. Implementation of data pre-processing strategies—such as enrichment, query generation, and retrieval—shows promise for improving performance and reducing computational costs. Intra-study comparisons suggest that GPT-4o achieves the strongest performance in trial-matching tasks, with high zero-shot capabilities reducing reliance on annotated datasets. For institutions unable to deploy GPT-4 within HIPAA-compliant infrastructure, smaller LLMs offer a viable alternative when fine-tuned for specific tasks. Future work should prioritize the use of standardized assessment metrics and test sets to enable cross-study comparison, as well as evaluate cost and bias compared to existing trial matching strategies. To ensure safe and effective implementation, models should be evaluated on real-world data and designed with transparency measures, such as LLM-generated eligibility justifications with sentences cited directly from records. Collaborative development with clinical stakeholders will be essential to ensure these tools augment, rather than replace, human decision-making and integrate effectively into real-world workflows. Methods Article Identification We identified 126 unique citations across three academic databases and one preprint server using a structured keyword search combining trial-related and LLM-related terms between January 1 st , 2020 and March 1 st , 2025 ( Fig. 1 ). Three articles were published as pre-prints within this time frame, then published in peer-reviewed journals after March 1 st , 2025, reflecting a common trend of first-step publishing in pre-print servers. The search string included: ("match[tiab]" OR "screen*[tiab]") AND ("clinical trial"[tiab] OR "clinical trials"[tiab]), combined with ("large language models"[tiab] OR "LLMs"[tiab] OR "large language model"[tiab] OR "LLM"[tiab] OR "ChatGPT"[tiab] OR "GPT-4"[tiab] OR "LLAMA"[tiab] OR "GPT-3.5"[tiab] OR "GPT-4"[tiab]). Covidence, a web-based review management tool, was employed to organize references, identify duplicate records, and document study inclusion/exclusion decisions. 25 After excluding 51 duplicates and 91 ineligible records, 31 full-text articles were included ( Fig. 1 ). All remaining records were accessible via full-text retrieval. Four were published in 2023, 23 in 2024, and four in 2025, reflecting rapid developments in the field. Declarations Data availability No new data were generated or analyzed in support of this review. Data sharing is not applicable. Code availability All data supporting this review are available within the cited articles. No new data were created. Author contributions B.A.M.: Conceptualization; Methodology; Search strategy; Data curation; Screening; Extraction; Formal analysis; Visualization; Writing–original draft; Writing–review & editing. M.S.: Methodology; Conceptualization; Supervision; Writing–review & editing. J.S.Y.: Conceptualization; Supervision; Writing–review & editing; Correspondence. Competing interests B.A.M.: No competing interests. M.S.: No competing interests. J.S.Y.: No competing interests. Funding This research received no specific grant from any funding agency in the public, commercial, or not-for-profit sectors. Acknowledgement Original art is the work of Mr. Kenneth Probst, Department of Neurosurgery, UCSF. The authors used the Elicit application (Ought, https://elicit.org) during the early scoping stage to roughly identify which articles contained certain types of information and to help organize thematic areas for review. UCSF Versa was used to assist with conciseness. No text or analysis generated by the tool appears in the submitted manuscript; all synthesis, interpretation, and writing were performed by the authors. References Wen PY, Weller M, Lee EQ, Alexander BM, Barnholtz-Sloan JS, Barthel FP, et al. Glioblastoma in adults: a Society for Neuro-Oncology (SNO) and European Society of Neuro-Oncology (EANO) consensus review on current management and future directions. Neuro-Oncol. 2020;22(8):1073–113. Yu W, Zhou D, Meng F, Wang J, Wang B, Qiang J, et al. The global, regional burden of pancreatic cancer and its attributable risk factors from 1990 to 2021. BMC Cancer. 2025;25(1):186. Mani K, Deng D, Lin C, Wang M, Hsu ML, Zaorsky NG. Causes of death among people living with metastatic cancer. Nat Commun. 2024;15(1):1519. Taylor K, Francesca P, Cru MJ, Ronte H, Haughey J. Intelligent clinical trials [Internet]. [cited 2025 Jan 26]. Available from: https://www2.deloitte.com/content/dam/insights/us/articles/22934_intelligent-clinical-trials/DI_Intelligent-clinical-trials.pdf Kasenda B, von Elm E, You J, Blümle A, Tomonaga Y, Saccilotto R, et al. Prevalence, Characteristics, and Publication of Discontinued Randomized Trials. JAMA. 2014;311(10):1045–52. Wong AR, Sun V, George K, Liu J, Padam S, Chen BA, et al. Barriers to Participation in Therapeutic Clinical Trials as Perceived by Community Oncologists. JCO Oncol Pract. 2020 Sept;16(9):e849–58. Durden K, Hurley P, Butler DL, Farner A, Shriver SP, Fleury ME. Provider motivations and barriers to cancer clinical trial screening, referral, and operations: Findings from a survey. Cancer. 2024;130(1):68–76. Penberthy LT, Dahman BA, Petkov VI, DeShazo JP. Effort Required in Eligibility Screening for Clinical Trials. J Oncol Pract. 2012;8(6):365–70. Guerra CE, Fleury ME, Byatt LP, Lian T, Pierce L. Strategies to Advance Equity in Cancer Clinical Trials. Am Soc Clin Oncol Educ Book Am Soc Clin Oncol Annu Meet. 2022;42:1–11. Yuan C, Ryan PB, Ta C, Guo Y, Li Z, Hardin J, et al. Criteria2Query: a natural language interface to clinical databases for cohort definition. J Am Med Inform Assoc. 2019;26(4):294–305. Thadani SR, Weng C, Bigger JT, Ennever JF, Wajngurt D. Electronic Screening Improves Efficiency in Clinical Trial Recruitment. J Am Med Inform Assoc JAMIA. 2009;16(6):869–73. Shriver SP, Arafat W, Potteiger C, Butler DL, Beg MS, Hullings M, et al. Feasibility of institution-agnostic, EHR-integrated regional clinical trial matching. Cancer. 2024;130(1):60–7. Haddad TC, Helgeson J, Pomerleau K, Makey M, Lombardo P, Coverdill S, et al. Impact of a cognitive computing clinical trial matching system in an ambulatory oncology practice. J Clin Oncol. 2018;36(15_suppl):6550–6550. Gupta S, Basu A, Nievas M, Thomas J, Wolfrath N, Ramamurthi A, et al. PRISM: Patient Records Interpretation for Semantic clinical trial Matching system using large language models. Npj Digit Med. 2024;7(1):1–12. Shi H, Zhang J, Zhang K. Enhancing Clinical Trial Patient Matching through Knowledge Augmentation with Multi-Agents [Internet]. arXiv; 2024 [cited 2025 Jan 19]. Available from: http://arxiv.org/abs/2411.14637 Unlu O, Shin J, Mailly CJ, Oates MF, Tucci MR, Varugheese M, et al. Retrieval-Augmented Generation–Enabled GPT-4 for Clinical Trial Screening. NEJM AI. 2024 June 27;1(7):AIoa2400181. 2021 TREC Clinical Trials Track [Internet]. [cited 2025 Jan 15]. Available from: http://www.trec-cds.org/2021.html#documents 2022 TREC Clinical Trials Track [Internet]. [cited 2025 Jan 15]. Available from: https://www.trec-cds.org/2023.html 2023 TREC Clinical Trials Track [Internet]. [cited 2025 Jan 15]. Available from: https://www.trec-cds.org/2023.html Koopman B, Zuccon G. A Test Collection for Matching Patients to Clinical Trials. In: Proceedings of the 39th International ACM SIGIR conference on Research and Development in Information Retrieval [Internet]. Pisa Italy: ACM; 2016 [cited 2025 Jan 4]. p. 669–72. Available from: https://dl.acm.org/doi/10.1145/2911451.2914672 Stubbs A, Filannino M, Soysal E, Henry S, Uzuner Ö. Cohort selection for clinical trials: n2c2 2018 shared task track 1. J Am Med Inform Assoc. 2019;26(11):1163–71. Jin Q, Wang Z, Floudas CS, Chen F, Gong C, Bracken-Clarke D, et al. Matching patients to clinical trials with large language models. Nat Commun. 2024;15(1):9074. Nievas M, Basu A, Wang Y, Singh H. Distilling large language models for matching patients to clinical trials. J Am Med Inform Assoc JAMIA. 2024 Sept 1;31(9):1953–63. Wornow M, Lozano A, Dash D, Jindal J, Mahaffey KW, Shah NH. Zero-Shot Clinical Trial Patient Matching with LLMs. NEJM AI. 2024;0(0):AIcs2400360. Covidence - Better systematic review management [Internet]. Covidence. [cited 2025 Sept 4]. Available from: https://www.covidence.org/ Beattie J, Neufeld S, Yang D, Chukwuma C, Gul A, Desai N, et al. Utilizing Large Language Models for Enhanced Clinical Trial Matching: A Study on Automation in Patient Screening. Cureus. 2024;16(5):e60044. Beattie J, Owens D, Navar AM, Giuliani Schmitt L, Taing K, Neufeld S, et al. ChatGPT augmented clinical trial screening. Mach Learn Health. 2025 July;1(1):015005. Cerami E, Trukhanov P, Paul MA, Hassett MJ, Riaz IB, Lindsay J, et al. MatchMiner-AI: An Open-Source Solution for Cancer Clinical Trial Matching [Internet]. arXiv; 2024 [cited 2025 Jan 11]. Available from: http://arxiv.org/abs/2412.17228 Chowdhury S, Rajaganapathy S, Yu Y, Tao C, Vassilaki M, Zong N. Matching Patients to Clinical Trials using LLaMA 2 Embeddings and Siamese Neural Network. medRxiv. 2024 June 30;2024.06.28.24309677. Datta S, Lee K, Huang LC, Paek H, Gildersleeve R, Gold J, et al. Patient2Trial: From Patient to Participant in Clinical Trials Using Large Language Models. Inform Med Unlocked. 2025;101615. Devi A, Uttrani S, Singla A, Jha S, Dasgupta N, Natarajan S, et al. Quantitative Analysis of GPT-4 model: Optimizing Patient Eligibility Classification for Clinical Trials and Reducing Expert Judgment Dependency. In: Proceedings of the 2024 8th International Conference on Medical and Health Informatics [Internet]. New York, NY, USA: Association for Computing Machinery; 2024 [cited 2024 Dec 23]. p. 230–7. (ICMHI ’24). Available from: https://dl.acm.org/doi/10.1145/3673971.3674014 Devi A, Uttrani S, Singla A, Jha S, Dasgupta N, Natarajan S, et al. Automating Clinical Trial Eligibility Screening: Quantitative Analysis of GPT Models versus Human Expertise. In: Proceedings of the 17th International Conference on PErvasive Technologies Related to Assistive Environments [Internet]. New York, NY, USA: Association for Computing Machinery; 2024 [cited 2024 Dec 22]. p. 626–32. (PETRA ’24). Available from: https://dl.acm.org/doi/10.1145/3652037.3663922 Ferber D, Hilgers L, Wiest IC, Leßmann ME, Clusmann J, Neidlinger P, et al. End-To-End Clinical Trial Matching with Large Language Models [Internet]. arXiv; 2024 [cited 2025 Jan 19]. Available from: http://arxiv.org/abs/2407.13463 Gueguen L, Olgiati L, Brutti-Mairesse C, Sans A, Le Texier V, Verlingue L. A prospective pragmatic evaluation of automatic trial matching tools in a molecular tumor board. Npj Precis Oncol. 2025;9(1):28. Gui X, Lv H, Wang X, Lv L, Xiao Y, Wang L. Enhancing hepatopathy clinical trial efficiency: a secure, large language model-powered pre-screening pipeline. BioData Min. 2025 June 14;18:42. Jullien M, Bogatu A, Unsworth H, Freitas A. Controlled LLM-based Reasoning for Clinical Trial Retrieval [Internet]. arXiv; 2024 [cited 2025 Jan 19]. Available from: http://arxiv.org/abs/2409.18998 Kusa W, Mendoza ÓE, Knoth P, Pasi G, Hanbury A. Effective matching of patients to clinical trials using entity extraction and neural re-ranking. J Biomed Inform. 2023;144:104444. Kusa W, Styll P, Seeliger M, Espitia Mendoza Ó, Hanbury A. DoSSIER at TREC 2023 Clinical Trials Track. In: # PLACEHOLDER_PARENT_METADATA_VALUE# [Internet]. NIST; 2023 [cited 2025 Jan 19]. Available from: https://repositum.tuwien.at/handle/20.500. 12708/203878 Lai SM, Malik AM, Sathe TS, Silvestri CJ, Manji GA, Kluger MD. A Proof-of-Concept Large Language Model Application to Support Clinical Trial Screening in Surgical Oncology [Internet]. medRxiv; 2024 [cited 2025 Feb 21]. p. 2024.09.20.24314053. Available from: https://www.medrxiv.org/content/ 10.1101/2024.09.20.24314053v2 Lin J, Xu H, Wang Z, Wang S, Sun J. Panacea: A foundation model for clinical trial search, summarization, design, and recruitment [Internet]. medRxiv; 2024 [cited 2024 Dec 23]. p. 2024.06.26.24309548. Available from: https://www.medrxiv.org/content/ 10.1101/2024.06.26.24309548v1 Peikos G, Symeonidis S, Kasela P, Pasi G. Utilizing ChatGPT to Enhance Clinical Trial Enrollment [Internet]. arXiv; 2023 [cited 2025 Jan 18]. Available from: http://arxiv.org/abs/2306.02077 Peikos G, Kasela P, Pasi G. Leveraging Large Language Models for Medical Information Extraction and Query Generation [Internet]. arXiv; 2024 [cited 2025 Aug 27]. Available from: http://arxiv.org/abs/2410.23851 Rahmanian M, Fakhrahmad SM, Mousavi SZ. Towards Efficient Patient Recruitment for Clinical Trials: Application of a Prompt-Based Learning Model [Internet]. arXiv.org. 2024 [cited 2025 Feb 1]. Available from: https://arxiv.org/abs/2404.16198v1 Ruan J, Su Q, Chen Z, Huang J, Li Y. CPRS: a clinical protocol recommendation system based on LLMs. Int J Med Inf. 2024;195:105746. Rybinski M, Kusa W, Karimi S, Hanbury A. Learning to match patients to clinical trials using large language models. J Biomed Inform. 2024;159:104734. Unlu O, Shin J, Mailly CJ, Oates MF, Tucci MR, Varugheese M, et al. Retrieval Augmented Generation Enabled Generative Pre-Trained Transformer 4 (GPT-4) Performance for Clinical Trial Screening. MedRxiv Prepr Serv Health Sci. 2024;2024.02.08.24302376. Wong C, Zhang S, Gu Y, Moung C, Abel J, Usuyama N, et al. Scaling Clinical Trial Matching Using Large Language Models: A Case Study in Oncology [Internet]. arXiv; 2023 [cited 2025 Jan 11]. Available from: http://arxiv.org/abs/2308.02180 Woo EG, Burkhart MC, Alsentzer E, Beaulieu-Jones BK. Synthetic data distillation enables the extraction of clinical information at scale. Npj Digit Med. 2025;8(1):267. Yuan J, Tang R, Jiang X, Hu X. Large Language Models for Healthcare Data Augmentation: An Example on Patient-Trial Matching. AMIA Annu Symp Proc. 2024;2023:1324–33. Zhuang S, Koopman B, Zuccon G. Team IELAB at TREC Clinical Trial Track 2023: Enhancing Clinical Trial Retrieval with Neural Rankers and Large Language Models [Internet]. arXiv; 2024 [cited 2024 Dec 21]. Available from: http://arxiv.org/abs/2401.01566 ZiHang C, QianMin S, GaoYi C, JiHan H, Ying L. Enhanced Pre-Recruitment Framework for Clinical Trial Questionnaires Through the Integration of Large Language Models and Knowledge Graphs [Internet]. Rochester, NY: Social Science Research Network; 2024 [cited 2025 Jan 19]. Available from: https://papers.ssrn.com/abstract=4713177 A S, Ö U. Annotating longitudinal clinical narratives for de-identification: The 2014 i2b2/UTHealth corpus. J Biomed Inform [Internet]. 2015 Dec [cited 2025 Aug 27];58 Suppl(Suppl). Available from: https://pubmed.ncbi.nlm.nih.gov/26319540/ A S, Ö U. Annotating risk factors for heart disease in clinical narratives for diabetic patients. J Biomed Inform [Internet]. 2015 Dec [cited 2025 Aug 27];58 Suppl(Suppl). Available from: https://pubmed.ncbi.nlm.nih.gov/26004790/ Beattie J, Owens D, Navar AM, Schmitt LG, Taing K, Neufeld S, et al. Large Language Model Augmented Clinical Trial Screening [Internet]. medRxiv; 2024 [cited 2024 Dec 22]. p. 2024.08.27.24312646. Available from: https://www.medrxiv.org/content/ 10.1101/2024.08.27.24312646v1 Voorhees EM, Hersh W. Overview of the TREC 2012 Medical Records Track. NIST [Internet]. 2013 June 28 [cited 2025 Aug 27]; Available from: https://www.nist.gov/publications/overview-trec-2012-medical-records-track Woo E, Burkhart MC, Alsentzer E, Beaulieu-Jones B. MIMIC-III-Ext-Synthetic-Clinical-Trial-Questions [Internet]. PhysioNet; [cited 2025 Sept 2]. Available from: https://physionet.org/content/mimic-ext-synth-trial-question/1.0.0/ Woo E, Burkhart MC, Alsentzer E, Beaulieu-Jones B. MIMIC-IV-Ext-Apixaban-Trial-Criteria-Questions [Internet]. PhysioNet; [cited 2025 Sept 2]. Available from: https://physionet.org/content/mimic-iv-ext-apixaban-trial/1.0.0/ Yang Y, Jayaraj S, Ludmir E, Roberts K. Text Classification of Cancer Clinical Trial Eligibility Criteria. AMIA Annu Symp Proc. 2024;2023:1304–13. Additional Declarations No competing interests reported. Cite Share Download PDF Status: Under Review Version 1 posted Editorial decision: Revision requested 07 Jan, 2026 Reviews received at journal 06 Jan, 2026 Reviewers agreed at journal 11 Dec, 2025 Reviewers agreed at journal 10 Dec, 2025 Reviews received at journal 17 Nov, 2025 Reviewers agreed at journal 17 Nov, 2025 Reviewers agreed at journal 13 Nov, 2025 Reviewers invited by journal 10 Nov, 2025 Editor assigned by journal 10 Nov, 2025 Submission checks completed at journal 10 Nov, 2025 First submitted to journal 05 Nov, 2025 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-8036235","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Article","associatedPublications":[],"authors":[{"id":546496149,"identity":"feac7330-9e9f-4c2a-a6cc-9c069747fda4","order_by":0,"name":"Braxton A. Morrison","email":"","orcid":"","institution":"University of California San Francisco","correspondingAuthor":false,"prefix":"","firstName":"Braxton","middleName":"A.","lastName":"Morrison","suffix":""},{"id":546496150,"identity":"4731aa68-ad54-400e-9696-7380ab7c3d24","order_by":1,"name":"Madhumita Sushil","email":"","orcid":"","institution":"University of California San Francisco","correspondingAuthor":false,"prefix":"","firstName":"Madhumita","middleName":"","lastName":"Sushil","suffix":""},{"id":546496151,"identity":"aae88c45-741e-4222-8941-8707bd90be5c","order_by":2,"name":"Jacob S. Young","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA5UlEQVRIiWNgGAWjYBACxgYGBmYY+wGQ4AdiA6K1MIOUSjYQ0gJWCqXZJIjSwtx++OHjAobDcga3269V89TUSTCwN2+TwOuwnjRj4xkMh40N7pwpu81z7LAEA8+xMvxaGnLYpHkYDiduuJGTdpu34UAdg0SOGX4t/W8QWop5G4AOk39DQMsMuC3px5h5G5glGCR4CGl5BvSLQbqx5I0cZsk5QL+w8aQVW+DTYtifDAyxCms5vhvpDz+8AYYYP/vhjTfwamkAkQbNQIIHEh1s+JSDgDyEqgNi9geEFI+CUTAKRsEIBQCYGUQlJI5WEQAAAABJRU5ErkJggg==","orcid":"","institution":"University of California San Francisco","correspondingAuthor":true,"prefix":"","firstName":"Jacob","middleName":"S.","lastName":"Young","suffix":""}],"badges":[],"createdAt":"2025-11-05 08:53:14","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-8036235/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-8036235/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":96454242,"identity":"e3076241-2222-498f-88df-665f892e4eba","added_by":"auto","created_at":"2025-11-21 10:02:29","extension":"docx","order_by":0,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":7311854,"visible":true,"origin":"","legend":"","description":"","filename":"LLMsforTrialMatching11.9.25Nature.docx","url":"https://assets-eu.researchsquare.com/files/rs-8036235/v1/55216a723f487a55137855fe.docx"},{"id":96454444,"identity":"24d0621e-2d8e-406f-a00e-46b609c43e96","added_by":"auto","created_at":"2025-11-21 10:02:45","extension":"json","order_by":1,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":5998,"visible":true,"origin":"","legend":"","description":"","filename":"74bff427124045ba80253068cca5bd75.json","url":"https://assets-eu.researchsquare.com/files/rs-8036235/v1/6facda43a9eb6b5a61bac3f8.json"},{"id":96400065,"identity":"f5b630dc-b8d9-4d49-a2b1-a61782735e49","added_by":"auto","created_at":"2025-11-20 16:03:30","extension":"xml","order_by":2,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":180864,"visible":true,"origin":"","legend":"","description":"","filename":"74bff427124045ba80253068cca5bd751enriched.xml","url":"https://assets-eu.researchsquare.com/files/rs-8036235/v1/959af22471fb66795511670b.xml"},{"id":96400062,"identity":"162a9633-4b41-44f5-8d61-b2fbb956b56c","added_by":"auto","created_at":"2025-11-20 16:03:30","extension":"png","order_by":6,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":40668,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage1.png","url":"https://assets-eu.researchsquare.com/files/rs-8036235/v1/4e06e8ea240897d804e659d2.png"},{"id":96400066,"identity":"7e08dc88-7aab-4c5f-a434-66c8ae06f2be","added_by":"auto","created_at":"2025-11-20 16:03:30","extension":"png","order_by":7,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":240942,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage2.png","url":"https://assets-eu.researchsquare.com/files/rs-8036235/v1/773760d2371874d92d59540f.png"},{"id":96400068,"identity":"60c413ab-78e8-4d3d-a527-0c3232c80138","added_by":"auto","created_at":"2025-11-20 16:03:30","extension":"png","order_by":8,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":122359,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage3.png","url":"https://assets-eu.researchsquare.com/files/rs-8036235/v1/af1e2c3d460c1dd891825f33.png"},{"id":96400069,"identity":"c4260dd8-0330-4b00-ab2f-a91f6941f036","added_by":"auto","created_at":"2025-11-20 16:03:30","extension":"xml","order_by":9,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":181318,"visible":true,"origin":"","legend":"","description":"","filename":"74bff427124045ba80253068cca5bd751structuring.xml","url":"https://assets-eu.researchsquare.com/files/rs-8036235/v1/4df18d31905d03bca08d3b90.xml"},{"id":96400071,"identity":"4db73a75-3c66-4665-899f-52560328ffa6","added_by":"auto","created_at":"2025-11-20 16:03:30","extension":"html","order_by":10,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":189474,"visible":true,"origin":"","legend":"","description":"","filename":"earlyproof.html","url":"https://assets-eu.researchsquare.com/files/rs-8036235/v1/dcd14abeaa356b8301e4a5dc.html"},{"id":96400063,"identity":"ee3a8625-90c8-477f-8da5-3073e2e55234","added_by":"auto","created_at":"2025-11-20 16:03:30","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":156684,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cem\u003eFlowchart depicting the systematic review search strategy and study selection according to PRISMA guidelines.\u003c/em\u003e Searches across three academic databases and one preprint server yielded 126 unique journal articles. Records were screened for eligibility, with exclusions primarily due to duplication or failure to meet predefined inclusion criteria. Source: Covidence systematic review software, Veritas Health Innovation, Melbourne, Australia. Available at www.covidence.org.\u003c/p\u003e","description":"","filename":"floatimage1.png","url":"https://assets-eu.researchsquare.com/files/rs-8036235/v1/84df568ecb26e29c70a4df93.png"},{"id":96454349,"identity":"3515cdf5-f179-4e1e-b80d-7365c5b15f2a","added_by":"auto","created_at":"2025-11-21 10:02:39","extension":"jpeg","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":1202957,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cem\u003eSchematic Overview of Common Trial Matching Pipeline Designs.\u003c/em\u003e (a) Key components of the reviewed trial matching pipelines include data input, data pre-processing, matching algorithms, and the output, which takes the form of patient-criterion, patient-trial, and/or trial-patient matches. (b) Each article used its own combination of patient and trial data. (c) Various data enrichment methods are employed on trial or patient data to improve the generation of queries, which are then used to retrieve the patient information most relevant to a given criterion or trial. In most pipelines, this patient and trial information is then restructured and inserted into a prompt to be fed into an LLM. (d) The first and most common matching pipeline involves feeding prompts with patient and trial data to an LLM, the output of which is inputted into article-specific scoring algorithms to generate ranked outputs. The second pipeline involves embedding specific sets of patient and trial data (e.g. patient summaries and eligibility criteria) into a shared vector space, then using similarity measures like cosine similarity to rank patients or trials. (e) While patient-trial matching can be used by patients and providers to identify trials for potential enrollment, trial-patient matching would be used by trial coordinators to identify patients to enroll in a given trial.\u003c/p\u003e","description":"","filename":"floatimage2.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-8036235/v1/cf756f21c5a356d137a56d62.jpeg"},{"id":96454228,"identity":"92ac5985-826a-4d49-b24f-109eca6bf1d6","added_by":"auto","created_at":"2025-11-21 10:02:29","extension":"jpeg","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":632767,"visible":true,"origin":"","legend":"\u003cp\u003eSee image above for figure legend\u003c/p\u003e","description":"","filename":"floatimage3.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-8036235/v1/c609bca45a978badbc410d63.jpeg"},{"id":96708066,"identity":"67d0d4a7-3d29-4ea6-98a4-350ad427dea9","added_by":"auto","created_at":"2025-11-25 09:54:49","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":3668214,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-8036235/v1/07799933-2d58-400d-bac4-cfe95f41c596.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"A systematic review of trial-matching pipelines using large language models","fulltext":[{"header":"Introduction","content":"\u003cp\u003eClinical trials are essential for identifying novel treatment options and offering patients access to potentially life-saving therapies. This is particularly important in fields like oncology, where alternative therapeutic options may be limited.\u003csup\u003e\u003cspan additionalcitationids=\"CR2\" citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e–\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e\u003c/sup\u003e One key stage of trial-matching includes patient recruitment, which represents ~ 32% of clinical trial cost and is the reason most frequently cited for the discontinuation of randomized controlled trials.\u003csup\u003e\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e,\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e\u003c/sup\u003e Another important stage is manual trial matching, which requires intensive review of complex eligibility criteria across extensive patient records to assess patient suitability for a given trial, a time-consuming and labor-intensive process.\u003csup\u003e\u003cspan additionalcitationids=\"CR7\" citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e–\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e\u003c/sup\u003e This stage is also costly; one study reported a cost between \u003cspan\u003e$\u003c/span\u003e129.15 to \u003cspan\u003e$\u003c/span\u003e336.48 per enrolled patient.\u003csup\u003e\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e\u003c/sup\u003e Equity is another key issue with trial enrollment; minorities, elderly people and rural groups are often underrepresented in cancer trials, a problem which could potentially be addressed by screening all patients for a large number of trials.\u003csup\u003e\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e\u003c/sup\u003e Thus, screening more patients for eligibility across more trials could improve patient recruitment and equity, but cost, labor and time are core limiting factors to its implementation.\u003c/p\u003e\u003cp\u003eAutomated matching systems aim to address these inefficiencies, with early approaches relying on rigid, rule-based methods. One approach involved generating queries from trial eligibility criteria that could be applied to identify potentially eligible patients in clinical databases.\u003csup\u003e\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e,\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e\u003c/sup\u003e Another involved extracting key details from patient records and converting them to a structured format to filter lists of trials.\u003csup\u003e\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e\u003c/sup\u003e Earlier rule-based automation systems—such as the Watson for Clinical Trial Matching—demonstrated potential, as seen in a 2016 Mayo Clinic pilot showing an 80% increase in enrollment for breast cancer trials.\u003csup\u003e\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e\u003c/sup\u003e These systems often lacked generalizability across institutions, patient populations, or trial protocols. Recently, large language models (LLMs) have emerged as a flexible alternative capable of interpreting both structured and unstructured data, extracting eligibility criteria, and matching patients to trials with greater scalability.\u003c/p\u003e\u003cp\u003eSynthesizing recent work from 2020 to 2025, we characterize datasets, common pipeline structures, comparative model performance, and evaluation strategies. We aim to provide a practical roadmap for researchers and implementers to understand the capabilities, limitations, and emerging best practices for deploying LLMs in this domain.\u003c/p\u003e"},{"header":"Results","content":"\u003cdiv id=\"Sec3\" class=\"Section2\"\u003e\u003ch2\u003eModel Categories\u003c/h2\u003e\u003cp\u003eThis review covered four main types of trial matching \u0026ndash; \u0026ldquo;patient-criterion matching\u0026rdquo;, in which a patient is assessed for one or more criteria individually; \u0026ldquo;patient-to-trial,\u0026rdquo; through which a list of trials is determined for a given patient; \u0026ldquo;trial-to-patient,\u0026rdquo; through which a list of potentially eligible patients is provided for a given trial; and \u0026ldquo;binary classification\u0026rdquo;, in which it is determined whether a patient-trial pair is a match. Of the reviewed articles, four focused on matching patient-to-criterion only, ten on patient-to-trial only, two on trial-to-patient only, one on binary eligibility classification only and fourteen on combined tasks (Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e). While in theory, a perfect match between a patient and all trial criteria should imply eligibility, real-world applications are more complex. Strict 100% patient-criterion match thresholds often exclude many patients who may still be eligible under fewer, more practical constraints. In a study by Gupta et al., they found that weighting more important criteria more heavily resulted in better performance in translating criterion-level assessments into patient-trial matches.\u003csup\u003e\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e\u003c/sup\u003e\u003c/p\u003e\u003cp\u003e\u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab1\" border=\"1\"\u003e\u003ccaption language=\"En\"\u003e\u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e\u003cdiv class=\"CaptionContent\"\u003e\u003cp\u003eModel Performance Assessment\u003c/p\u003e\u003c/div\u003e\u003c/caption\u003e\u003ccolgroup cols=\"9\"\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c7\" colnum=\"7\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c8\" colnum=\"8\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c9\" colnum=\"9\"\u003e\u003c/div\u003e\u003cthead\u003e\u003ctr\u003e\u003cth align=\"left\" colname=\"c1\"\u003e\u003cp\u003eStudy\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c2\"\u003e\u003cp\u003eStudy Data Set\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c3\"\u003e\u003cp\u003eMain Model(s)\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colspan=\"2\" nameend=\"c5\" namest=\"c4\"\u003e\u003cp\u003ePatient-to-Criterion Metrics\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c6\"\u003e\u003cp\u003ePatient-to-Trial Metrics\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c7\"\u003e\u003cp\u003eTrial-to-Patient Metrics\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c8\"\u003e\u003cp\u003eModel Comparisons\u0026dagger;\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c9\"\u003e\u003cp\u003eCost/Efficiency\u003c/p\u003e\u003c/th\u003e\u003c/tr\u003e\u003c/thead\u003e\u003ctbody\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eBeattie et al., 2024\u003csup\u003e26\u003c/sup\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e\u0026minus;\u0026thinsp;2018 n2c2\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eGPT-3.5 Turbo; GPT-4\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colspan=\"2\" nameend=\"c5\" namest=\"c4\"\u003e\u003cp\u003eGPT-3.5 Turbo:\u003c/p\u003e\u003cp\u003e- Accuracy: 0.81\u003c/p\u003e\u003cp\u003e- Sensitivity: 0.80\u003c/p\u003e\u003cp\u003e- Specificity: 0.82\u003c/p\u003e\u003cp\u003e- Micro F\u003csub\u003e1\u003c/sub\u003e: 0.79\u003c/p\u003e\u003cp\u003eGPT-4:\u003c/p\u003e\u003cp\u003e- Accuracy: 0.87\u003c/p\u003e\u003cp\u003e- Sensitivity: 0.85\u003c/p\u003e\u003cp\u003e- Specificity: 0.89\u003c/p\u003e\u003cp\u003e- Micro F\u003csub\u003e1\u003c/sub\u003e: 0.86\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003eN/A\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003eN/A\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003eN/A\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c9\"\u003e\u003cp\u003eModel-processing time: 1\u0026ndash;5 minutes per patient for all included criteria\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eBeattie et al., 2024\u003csup\u003e27\u003c/sup\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e- Article-specific patient and trial data\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colspan=\"2\" nameend=\"c4\" namest=\"c3\"\u003e\u003cp\u003eGPT-3.5; GPT-4\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003eGPT-3.5:\u003c/p\u003e\u003cp\u003e\u0026nbsp;\u0026nbsp;\u0026nbsp; - Accuracy: 0.761\u003c/p\u003e\u003cp\u003e\u0026nbsp;\u0026nbsp;\u0026nbsp; - Sensitivity: 0.776\u003c/p\u003e\u003cp\u003e\u0026nbsp;\u0026nbsp;\u0026nbsp; - Specificity: 0.732\u003c/p\u003e\u003cp\u003e\u0026nbsp;\u0026nbsp;\u0026nbsp; - Youden Index: 0.357\u003c/p\u003e\u003cp\u003eGPT-4:\u003c/p\u003e\u003cp\u003e\u0026nbsp;\u0026nbsp;\u0026nbsp; - Accuracy: 0.838\u003c/p\u003e\u003cp\u003e\u0026nbsp;\u0026nbsp;\u0026nbsp; - Sensitivity: 0.839\u003c/p\u003e\u003cp\u003e\u0026nbsp;\u0026nbsp;\u0026nbsp; - Specificity: 0.830\u003c/p\u003e\u003cp\u003e\u0026nbsp;\u0026nbsp;\u0026nbsp; - Youden Index: 0.668\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003eGPT-3.5:\u003c/p\u003e\u003cp\u003e- Median AUC: 0.64\u003c/p\u003e\u003cp\u003e- Median accuracy: 0.65\u003c/p\u003e\u003cp\u003e-Median sensitivity: 0.68\u003c/p\u003e\u003cp\u003e- Median specificity: 0.69\u003c/p\u003e\u003cp\u003e- Median Youden index: 0.27\u003c/p\u003e\u003cp\u003eGPT-4\u003c/p\u003e\u003cp\u003e- Median AUC: 0.74\u003c/p\u003e\u003cp\u003e- Median accuracy: 0.70\u003c/p\u003e\u003cp\u003e- Median sensitivity: 0.66\u003c/p\u003e\u003cp\u003e- Median specificity: 0.74\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003eN/A\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003eProprietary: GPT-3.5; GPT-4\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c9\"\u003e\u003cp\u003eScreening cost per patient:\u003c/p\u003e\u003cp\u003e- GPT-3.5: \u003cspan\u003e$\u003c/span\u003e0.02-\u003cspan\u003e$\u003c/span\u003e0.03\u003c/p\u003e\u003cp\u003e- GPT-4: \u003cspan\u003e$\u003c/span\u003e0.15-\u003cspan\u003e$\u003c/span\u003e0.27.\u003c/p\u003e\u003cp\u003eScreening time per patient:\u003c/p\u003e\u003cp\u003e- GPT-3.5: 1.4-3 minutes\u003c/p\u003e\u003cp\u003e- GPT-4: 7.9\u0026ndash;12.4 minutes\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eCerami et al., 2024\u003csup\u003e28\u003c/sup\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e- Article-specific patient and trial data\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colspan=\"2\" nameend=\"c4\" namest=\"c3\"\u003e\u003cp\u003eTrialSpace text embedding\u0026thinsp;+\u0026thinsp;TrialChecker classification model (fine-tuned RoBERTa-Large; open-access)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003eN/A\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003eDFCI trial enrollment dataset:\u003c/p\u003e\u003cp\u003e- Precision@10: 0.89\u003c/p\u003e\u003cp\u003e- MAP@10: 0.93\u003c/p\u003e\u003cp\u003eDFCI standard-of-care dataset:\u003c/p\u003e\u003cp\u003e- Precision@10: 0.87\u003c/p\u003e\u003cp\u003e- MAP@10: 0.91\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003eDFCI trial enrollment dataset:\u003c/p\u003e\u003cp\u003e- Precision@20: 0.91\u003c/p\u003e\u003cp\u003e- MAP@20: 0.93\u003c/p\u003e\u003cp\u003eDFCI standard-of-care dataset:\u003c/p\u003e\u003cp\u003e- Precision@20: 0.87\u003c/p\u003e\u003cp\u003e- MAP@20: 0.90\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003eN/A\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c9\"\u003e\u003cp\u003eN/A\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eChowdhury et al., 2024\u003csup\u003e29\u003c/sup\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e- Article-specific patient and trial data\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eSiamese-Patient-Trial-Matching with LLaMA 2 to initialize embeddings of inputs (open-access)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colspan=\"2\" nameend=\"c5\" namest=\"c4\"\u003e\u003cp\u003eN/A\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colspan=\"2\" nameend=\"c7\" namest=\"c6\"\u003e\u003cp\u003eSiamese-PTM (fine-grained): F\u003csub\u003e1\u003c/sub\u003e score of 0.92\u003c/p\u003e\u003cp\u003e*Note: article describes binary classification task of eligible/not eligible so could be labeled either \u0026ldquo;patient-trial\u0026rdquo; or \u0026ldquo;trial-patient\u0026rdquo; matching\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003eN/A\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c9\"\u003e\u003cp\u003eN/A\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eDatta et al., 2025\u003csup\u003e30\u003c/sup\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e- TREC 2023 CT\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eGPT-4\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colspan=\"2\" nameend=\"c5\" namest=\"c4\"\u003e\u003cp\u003eN/A\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e- Precision@10: 0.7351\u003c/p\u003e\u003cp\u003e- NDCG@10: 0.8109\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003eN/A\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003eN/A\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c9\"\u003e\u003cp\u003eN/A\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eDevi et al., 2024\u003csup\u003e32\u003c/sup\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e- Article-specific patient and trial data\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eGPT-3.5 Turbo\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colspan=\"2\" nameend=\"c5\" namest=\"c4\"\u003e\u003cp\u003eN/A\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003eN/A\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003eAccuracy of GPT-3.5:\u003c/p\u003e\u003cp\u003eNo tuning tested on 20 data points: 100%\u003c/p\u003e\u003cp\u003eTuned and tested on 20 data points: 95%\u003c/p\u003e\u003cp\u003eNo tuning tested on 100 data points: 79%\u003c/p\u003e\u003cp\u003eTuned and tested on 50 data points: 82%\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003eN/A\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c9\"\u003e\u003cp\u003eN/A\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eDevi et al., 2024\u003csup\u003e31\u003c/sup\u003e\u003c/p\u003e\u003cp\u003e\u003csup\u003e31,32\u003c/sup\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e- Article-specific patient and trial data\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eGPT-4\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colspan=\"2\" nameend=\"c5\" namest=\"c4\"\u003e\u003cp\u003eN/A\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003eN/A\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003eAccuracy:\u003c/p\u003e\u003cp\u003eNo tuning, tested on 20 data points: 95%\u003c/p\u003e\u003cp\u003eTuned and tested on 20 data points: 100%\u003c/p\u003e\u003cp\u003eNo tuning, tested on 100 data points: 86%\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003eN/A\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c9\"\u003e\u003cp\u003eN/A\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eFerber et al., 2024\u003csup\u003e33\u003c/sup\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e- Article-specific patient and trial data\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colspan=\"2\" nameend=\"c4\" namest=\"c3\"\u003e\u003cp\u003eGPT-4o\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003eAccuracy: 88.0%\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003eFor 14/15 patients, their \u0026ldquo;target trial\u0026rdquo; was listed within the top 15 trial matches identified by the algorithm\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003eN/A\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003eN/A\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c9\"\u003e\u003cp\u003eN/A\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eGueguen et al, 2025\u003csup\u003e34\u003c/sup\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eTrialGPT to re-rank outputs from DigitalECMT\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colspan=\"2\" nameend=\"c4\" namest=\"c3\"\u003e\u003cp\u003eTrialGPT with Qwen2.5-7B-Instruct as the LLM\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003eN/A\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003eUse of LLM-based re-ranking on results of DigitalECMT increased NDCG@3 from 0.61 to 0.64\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003eN/A\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003eN/A\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c9\"\u003e\u003cp\u003eN/A\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eGui et al, 2025\u003csup\u003e35\u003c/sup\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e- Article-specific patient and trial data\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colspan=\"2\" nameend=\"c4\" namest=\"c3\"\u003e\u003cp\u003e- Anthropo-morphized Experts\u0026rsquo; Chain of Thought\u003c/p\u003e\u003cp\u003e- LLMs converted eligibility criteria into questions: GPT-4o, Google Gemini Advanced, Anthropic Claude 3.5 Sonnet\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003ePathway A, Majority Vote\u003c/p\u003e\u003cp\u003e- Precision: ~0.921\u003c/p\u003e\u003cp\u003e- Recall: ~0.82\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colspan=\"2\" nameend=\"c7\" namest=\"c6\"\u003e\u003cp\u003ePathway A, Majority Vote\u003c/p\u003e\u003cp\u003e- Precision: ~0.922\u003c/p\u003e\u003cp\u003e- Recall: ~0.819\u003c/p\u003e\u003cp\u003e*Note: article describes binary classification task of eligible/not eligible so could be labeled either \u0026ldquo;patient-trial\u0026rdquo; or \u0026ldquo;trial-patient\u0026rdquo; matching\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003e- QWEN1.5, BAICHUAN, and GLM\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c9\"\u003e\u003cp\u003ePathway A, Majority Vote\u003c/p\u003e\u003cp\u003e- Efficiency: ~0.44 seconds/task\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eGupta et al., 2024\u003csup\u003e14\u003c/sup\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e- Article-specific patient and trial data\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eOncoLLM (fine-tuned Qwen-1.5 14B model; open-access)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colspan=\"2\" nameend=\"c5\" namest=\"c4\"\u003e\u003cp\u003eAccuracy:\u003c/p\u003e\u003cp\u003e- OncoLLM: 63%\u003c/p\u003e\u003cp\u003e- GPT-3.5 Turbo: 53%\u003c/p\u003e\u003cp\u003e- GPT-4: 68%\u003c/p\u003e\u003cp\u003e- Qwen14B-Chat: 43%\u003c/p\u003e\u003cp\u003e- Mitral-7B-Instruct: 41%\u003c/p\u003e\u003cp\u003e- Mixtral-8 X 7B-Instruct: 49%\u003c/p\u003e\u003cp\u003e- Meditron: 51%\u003c/p\u003e\u003cp\u003e- MedLlama: 55%\u003c/p\u003e\u003cp\u003e- TrialLlama: 57%\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003ePercentage of time model ranks ground truth trials in the top-3 positions among 10 considered trials\u003c/p\u003e\u003cp\u003e- OncoLLM: 65%\u003c/p\u003e\u003cp\u003e- GPT-3.5 Turbo (iterative): 61%\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003eNormalized Discounted Cumulative Gain\u003c/p\u003e\u003cp\u003e- OncoLLM: 68%\u003c/p\u003e\u003cp\u003e- GPT-3.5 Turbo: 62%\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003e- Proprietary: GPT-3.5 Turbo and GPT-4\u003c/p\u003e\u003cp\u003e- Open-source: Qwen14B-Chat, Mistral-7B-Instruct, Mixtral-8 \u0026times; 7B-Instruct, Meditron, MedLlama, and TrialLlama\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c9\"\u003e\u003cp\u003eOperating cost:\u003c/p\u003e\u003cp\u003e- OncoLLM: ~\u003cspan\u003e$\u003c/span\u003e170\u003c/p\u003e\u003cp\u003e- GPT-4: ~ \u003cspan\u003e$\u003c/span\u003e6055\u003c/p\u003e\u003cp\u003eCost of a single patient-trial match:\u003c/p\u003e\u003cp\u003e- OncoLLM: ~\u003cspan\u003e$\u003c/span\u003e0.17 per patient-trial pair\u003c/p\u003e\u003cp\u003e- GPT-4: ~ \u003cspan\u003e$\u003c/span\u003e6.18 per patient-trial pair\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eJin et al., 2024\u003csup\u003e22\u003c/sup\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e- SIGIR 2016\u003c/p\u003e\u003cp\u003e- TREC 2021 CT\u003c/p\u003e\u003cp\u003e- TREC 2022 CT\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eTrialGPT (GPT-4) (zero-shot)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colspan=\"2\" nameend=\"c5\" namest=\"c4\"\u003e\u003cp\u003eTrialGPT-Matching:\u003c/p\u003e\u003cp\u003e- Accuracy: 87.3%\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003eTrialGPT-Ranking:\u003c/p\u003e\u003cp\u003e- NDCG@10: 0.7252\u003c/p\u003e\u003cp\u003e- Precision@10: 0.6724\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003eN/A\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003e- Proprietary: GPT-3.5\u003c/p\u003e\u003cp\u003e- Open-source (for trial ranking): SciFive, BioBERT, PubMedBERT, SapBERT, and BioLinkBERT\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c9\"\u003e\u003cp\u003eReduction in screening time for patient recruitment:\u003c/p\u003e\u003cp\u003e42.6%\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eJullien et al., 2024\u003csup\u003e36\u003c/sup\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e- TREC 2022 CT\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eGPT-4 Turbo (proprietary)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colspan=\"2\" nameend=\"c5\" namest=\"c4\"\u003e\u003cp\u003eN/A\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e- NDCG@10: 0.679\u003c/p\u003e\u003cp\u003e- Precision@10: 0.730\u003c/p\u003e\u003cp\u003e- Precision@25: 0.630\u003c/p\u003e\u003cp\u003e- MRR: 0.860\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003eN/A\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003e- GPT-3.5 (proprietary)\u003c/p\u003e\u003cp\u003e- TREC SOTA (open-source)\u003c/p\u003e\u003cp\u003e*BM-25 for initial ranking\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c9\"\u003e\u003cp\u003eN/A\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eKusa et al., 2023\u003csup\u003e37\u003c/sup\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e- TREC 2021 CT\u003c/p\u003e\u003cp\u003e- TREC 2022 CT\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eTCRR with BioBERT (open-access)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colspan=\"2\" nameend=\"c5\" namest=\"c4\"\u003e\u003cp\u003eN/A\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e- NDCG@5: 0.627\u003c/p\u003e\u003cp\u003e- NDCG@10: 0.604\u003c/p\u003e\u003cp\u003e- Precision@10: 0.482\u003c/p\u003e\u003cp\u003e- Reciprocal Rank: 0.672\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003eN/A\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003e- Open-source: MonoBERT, TraditionalRR, TCRR initialized with \u003cem\u003ebert-base-uncased\u003c/em\u003e, BioBERT, ClinicalBERT\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c9\"\u003e\u003cp\u003eN/A\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eKusa et al., 2023\u003csup\u003e38\u003c/sup\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e- TREC 2023 CT with structured data converted to unstructured text using GPT-3.5\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eTCRR neural re-ranking model with BlueBERT (open-access) and GPT-3.5 (zero-shot) to refine results\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colspan=\"2\" nameend=\"c5\" namest=\"c4\"\u003e\u003cp\u003eN/A\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003eDoSSIER_3:\u003c/p\u003e\u003cp\u003e- nDCG@5: 0.6653\u003c/p\u003e\u003cp\u003e- nDCG@10: 0.6837\u003c/p\u003e\u003cp\u003e- Precision@10: 0.5838\u003c/p\u003e\u003cp\u003e- Reciprocal Rank: 0.6421\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003eN/A\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003eN/A\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c9\"\u003e\u003cp\u003eN/A\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eLai et al., 2024\u003csup\u003e39\u003c/sup\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e- Article-specific patient and trial data\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colspan=\"2\" nameend=\"c4\" namest=\"c3\"\u003e\u003cp\u003eGPT-4o (zero-shot)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003eGPT-4o:\u003c/p\u003e\u003cp\u003e\u0026nbsp;\u0026nbsp;\u0026nbsp; - Accuracy: 96.7% in agreeing with human raters on binary eligibility criteria\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003eGPT-4o:\u003c/p\u003e\u003cp\u003e\u0026nbsp;\u0026nbsp;\u0026nbsp; - Accuracy: 90.7% of eligible patient-trial matches\u003c/p\u003e\u003cp\u003e\u0026nbsp;\u0026nbsp;\u0026nbsp; - Sensitivity: 87.5\u0026ndash;100% for 8/9 trials\u003c/p\u003e\u003cp\u003e- Specificity: 73.3\u0026ndash;100% for all 9 trials\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003eN/A\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003eN/A\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c9\"\u003e\u003cp\u003e- Median cost to screen a single patient: \u003cspan\u003e$\u003c/span\u003e0.67 (range: \u003cspan\u003e$\u003c/span\u003e0.63-\u003cspan\u003e$\u003c/span\u003e0.74)\u003c/p\u003e\u003cp\u003e- Median time elapsed per patient: 138 seconds (range: 130\u0026ndash;146)\u003c/p\u003e\u003cp\u003e- Median total token usage: 112,266.5 tokens (range: 102982.0-122174.2)\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eLin et al., 2024\u003csup\u003e40\u003c/sup\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e- TREC 2021 CT\u003c/p\u003e\u003cp\u003e- SIGIR 2016\u003c/p\u003e\u003cp\u003e- TrialAlign\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003ePanacea (finely-tuned Mistral-7B-Base model27; open-access)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colspan=\"2\" nameend=\"c5\" namest=\"c4\"\u003e\u003cp\u003eYes; no metrics provided\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003eF1, precision, and recall\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003eN/A\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003e- Open-source: BioMistral, Mistral, Zephyr, LLAMA-2, MedAlpaca, Meditron\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c9\"\u003e\u003cp\u003eN/A\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eNievas et al., 2024\u003csup\u003e23\u003c/sup\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e- SIGIR 2016\u003c/p\u003e\u003cp\u003e- TREC CT 2021\u003c/p\u003e\u003cp\u003e- TREC CT 2022\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eTrial-LLAMA 70B (open-access)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colspan=\"2\" nameend=\"c5\" namest=\"c4\"\u003e\u003cp\u003eTrial-LLAMA 70B:\u003c/p\u003e\u003cp\u003e- Implicit CLA: 68.77\u003c/p\u003e\u003cp\u003e- Explicit CLA: 59.9\u003c/p\u003e\u003cp\u003eGPT-4:\u003c/p\u003e\u003cp\u003e- Implicit CLA: 75.31\u003c/p\u003e\u003cp\u003e- Explicit CLA: 58.8\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003eTrial-LLAMA 70B:\u003c/p\u003e\u003cp\u003e- NDCG@10: 0.6636\u003c/p\u003e\u003cp\u003e- Precision@10: 0.5886\u003c/p\u003e\u003cp\u003e- AUROC: 0.6528\u003c/p\u003e\u003cp\u003e- AURPC: 0.6515\u003c/p\u003e\u003cp\u003eGPT-4:\u003c/p\u003e\u003cp\u003e- NDCG@10: 0.7728\u003c/p\u003e\u003cp\u003e- Precision@10: 0.7005\u003c/p\u003e\u003cp\u003e- AUROC: 0.7390\u003c/p\u003e\u003cp\u003e- AURPC: 0.7038\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003eN/A\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003e- Proprietary: GPT-3.5 and GPT-4\u003c/p\u003e\u003cp\u003e- Open-source: LLAMA 7B, LLAMA 13B, and LLAMA 70B\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c9\"\u003e\u003cp\u003eN/A\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003ePeikos et al., 2023\u003csup\u003e41\u003c/sup\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e- TREC 2021 \u003c/p\u003e\u003cp\u003e- TREC 2022\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eGPT-3.5 Turbo\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colspan=\"2\" nameend=\"c5\" namest=\"c4\"\u003e\u003cp\u003eN/A\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003eTREC 2021 \u003c/p\u003e\u003cp\u003e- R-precision: 0.212\u003c/p\u003e\u003cp\u003e- Binary preference: 0.275\u003c/p\u003e\u003cp\u003e- Precision@10: 0.323\u003c/p\u003e\u003cp\u003e- P- recision @25: 0.261\u003c/p\u003e\u003cp\u003e- MRR: 0.541\u003c/p\u003e\u003cp\u003e- nDCG@10: 0.512\u003c/p\u003e\u003cp\u003eTREC 2022 \u003c/p\u003e\u003cp\u003e- R-precision: 0.276\u003c/p\u003e\u003cp\u003e- Binary preference: 0.298\u003c/p\u003e\u003cp\u003e- Precision @10: 0.372\u003c/p\u003e\u003cp\u003e- Precision @25: 0.338\u003c/p\u003e\u003cp\u003e- MRR: 0.576\u003c/p\u003e\u003cp\u003e- nDCG@10: 0.517\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003eN/A\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003eN/A\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c9\"\u003e\u003cp\u003eN/A\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003ePeikos et al, 2024\u003csup\u003e42\u003c/sup\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e- TREC 2021 \u003c/p\u003e\u003cp\u003e- TREC 2022\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eGPT-3.5 Turbo\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colspan=\"2\" nameend=\"c5\" namest=\"c4\"\u003e\u003cp\u003eN/A\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003eTREC 2021**\u003c/p\u003e\u003cp\u003eGPT-3.5:\u003c/p\u003e\u003cp\u003e- NDCG@10: 0.486\u003c/p\u003e\u003cp\u003e- Binary preference: 0.213\u003c/p\u003e\u003cp\u003e- Precision@10: 0.276\u003c/p\u003e\u003cp\u003e- Recall@25: 0.115\u003c/p\u003e\u003cp\u003e- MRR: 0.440\u003c/p\u003e\u003cp\u003eQwen2-7B-Instruct:\u003c/p\u003e\u003cp\u003e- NDCG@10: 0.476\u003c/p\u003e\u003cp\u003e- Binary preference: 0.216\u003c/p\u003e\u003cp\u003e- Precision@10: 0.285\u003c/p\u003e\u003cp\u003e- Recall@25: 0.107\u003c/p\u003e\u003cp\u003e- MRR: 0.465\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003eN/A\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003e-Proprietary: GPT-4\u003c/p\u003e\u003cp\u003e-Open access: Qwen2; Phi3-medium-4k-Instruct; Phi3-mini-4k-Instruct; Medical-Llama3-8B\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c9\"\u003e\u003cp\u003eN/A\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eRahmanian et al, 2024\u003csup\u003e43\u003c/sup\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e- n2c2 2018\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eGPT-3.5 Turbo\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colspan=\"2\" nameend=\"c5\" namest=\"c4\"\u003e\u003cp\u003eOverall (micro):\u003c/p\u003e\u003cp\u003e- F1: 0.9061\u003c/p\u003e\u003cp\u003e- AUC: 0.9035\u003c/p\u003e\u003cp\u003eOverall (macro):\u003c/p\u003e\u003cp\u003e- F1: 0.8060\u003c/p\u003e\u003cp\u003e- AUC: 0.7949\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003eN/A\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003eN/A\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003eN/A\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c9\"\u003e\u003cp\u003eN/A\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eRuan et al., 2024\u003csup\u003e44\u003c/sup\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e- Article-specific patient and trial data\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eGPT-4 with a knowledge graph\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colspan=\"2\" nameend=\"c5\" namest=\"c4\"\u003e\u003cp\u003eN/A\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003eGPT-4:\u003c/p\u003e\u003cp\u003e- MSE: 27.27\u003c/p\u003e\u003cp\u003e- RMSE: 5.22\u003c/p\u003e\u003cp\u003e*Measures similarity to criteria of another trial identified by the patient\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003eN/A\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003e- Open-source: 11 models from the SBERT family\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c9\"\u003e\u003cp\u003eN/A\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eRybinski et al., 2024\u003csup\u003e45\u003c/sup\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e- TREC 2021 CT\u003c/p\u003e\u003cp\u003e- TREC 2022 CT\u003c/p\u003e\u003cp\u003e- TREC 2023 CT\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eFine-tuned GPT-3.5 Turbo with BM25 retrieval\u003c/p\u003e\u003cp\u003e(zero-shot for final eligibility determination)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colspan=\"2\" nameend=\"c5\" namest=\"c4\"\u003e\u003cp\u003eN/A\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003eBM25 with TCRR and GPT-3.5 Turbo re-ranking:\u003c/p\u003e\u003cp\u003e- nDCG@1000: 0.375\u003c/p\u003e\u003cp\u003e- nDCG@10: 0.777\u003c/p\u003e\u003cp\u003e- Precision@10: 0.697\u003c/p\u003e\u003cp\u003e- Reciprocal Rank: 0.783\u003c/p\u003e\u003cp\u003eBM25 with chain-of-thought GPT-4o:\u003c/p\u003e\u003cp\u003e- nDCG@1000: 0.504\u003c/p\u003e\u003cp\u003e- nDCG@10: 0.785\u003c/p\u003e\u003cp\u003e- Precision@10: 0.603\u003c/p\u003e\u003cp\u003e- Reciprocal Rank: 0.844\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003eN/A\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003e- Proprietary: GPT4-o\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c9\"\u003e\u003cp\u003e- Cost of re-ranking with GPT-3.5 Turbo: ~\u003cspan\u003e$\u003c/span\u003e0.25 per 100 API calls (per patient)\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eShi et al., 2024\u003csup\u003e15\u003c/sup\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e\u0026minus;\u0026thinsp;2018 n2c2\u003c/p\u003e\u003cp\u003e- Synthesized a trial for which 28 patients from the 2018 n2c2 cohort would meet all criteria\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colspan=\"2\" nameend=\"c4\" namest=\"c3\"\u003e\u003cp\u003eMAKA (proprietary)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003eMAKA:\u003c/p\u003e\u003cp\u003e\u0026nbsp;\u0026nbsp;\u0026nbsp; - Accuracy: 0.909\u003c/p\u003e\u003cp\u003e\u0026nbsp;\u0026nbsp;\u0026nbsp; - Precision: 0.822\u003c/p\u003e\u003cp\u003e\u0026nbsp;\u0026nbsp;\u0026nbsp; - Recall: 0.846\u003c/p\u003e\u003cp\u003e\u0026nbsp;\u0026nbsp;\u0026nbsp; - F\u003csub\u003e1\u003c/sub\u003e-score: 0.828\u003c/p\u003e\u003cp\u003eStrategy by Wornow et al.:\u003c/p\u003e\u003cp\u003e\u0026nbsp;\u0026nbsp;\u0026nbsp; - Accuracy: 0.884\u003c/p\u003e\u003cp\u003e\u0026nbsp;\u0026nbsp;\u0026nbsp; - Precision: 0.727\u003c/p\u003e\u003cp\u003e\u0026nbsp;\u0026nbsp;\u0026nbsp; - Recall: 0.0.894\u003c/p\u003e\u003cp\u003e\u0026nbsp;\u0026nbsp;\u0026nbsp; - F\u003csub\u003e1\u003c/sub\u003e-score: 0.785\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003eN/A\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003eMAKA:\u003c/p\u003e\u003cp\u003e\u0026nbsp;\u0026nbsp;\u0026nbsp; - Accuracy: 0.9306\u003c/p\u003e\u003cp\u003e- Precision: 0.6333\u003c/p\u003e\u003cp\u003e- Recall: 0.6786\u003c/p\u003e\u003cp\u003e\u0026nbsp;\u0026nbsp;\u0026nbsp; - F\u003csub\u003e1\u003c/sub\u003e-score: 0.6552\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003e- Strategies employed by Wornow et al and Beattie et al\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c9\"\u003e\u003cp\u003eN/A\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eUnlu et al., 2024\u003csup\u003e46\u003c/sup\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e- Article-specific patient and trial data\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eRAG-Enabled Clinical Trial Infrastructure for Inclusion Exclusion Review (RECTIFIER) \u0026ndash; using GPT-4 Vision\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colspan=\"2\" nameend=\"c5\" namest=\"c4\"\u003e\u003cp\u003eRECTIFIER:\u003c/p\u003e\u003cp\u003e- Sensitivity: 75\u0026ndash;100%\u003c/p\u003e\u003cp\u003e- Specificity: 92.1\u0026ndash;100%\u003c/p\u003e\u003cp\u003e- PPV: 75\u0026ndash;100%\u003c/p\u003e\u003cp\u003e- MCC: 97.9\u0026ndash;100%\u003c/p\u003e\u003cp\u003eStudy Staff:\u003c/p\u003e\u003cp\u003e- Sensitivity: 66.7\u0026ndash;100%\u003c/p\u003e\u003cp\u003e- Specificity: 82.1\u0026ndash;100%\u003c/p\u003e\u003cp\u003e- PPV: 50\u0026ndash;100%\u003c/p\u003e\u003cp\u003e- MCC: 91.7\u0026ndash;100%\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003eRECTIFIER\u003c/p\u003e\u003cp\u003e- Sensitivity: 92.3%\u003c/p\u003e\u003cp\u003e- Specificity: 93.9%\u003c/p\u003e\u003cp\u003e- PPV: 98.1%\u003c/p\u003e\u003cp\u003e- NPV: 78.6%\u003c/p\u003e\u003cp\u003e- Accuracy: 92.7%\u003c/p\u003e\u003cp\u003e- MCC: 81.3%\u003c/p\u003e\u003cp\u003eStudy staff\u003c/p\u003e\u003cp\u003e- Sensitivity: 90.8%\u003c/p\u003e\u003cp\u003e- Specificity: 83.6%\u003c/p\u003e\u003cp\u003e- PPV: 94.9%\u003c/p\u003e\u003cp\u003e- NPV: 73.0%\u003c/p\u003e\u003cp\u003e- Accuracy: 89.1%\u003c/p\u003e\u003cp\u003e- MCC: 71.1%\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003eN/A\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003e- Proprietary: GPT-3.5\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c9\"\u003e\u003cp\u003eRECTIFIER:\u003c/p\u003e\u003cp\u003e- Individual-question approach: average of 11 cents/patient\u003c/p\u003e\u003cp\u003e- Combined-question approach: average of 2 cents/patient\u003c/p\u003e\u003cp\u003eCost without RAG:\u003c/p\u003e\u003cp\u003e- GPT-4: \u003cspan\u003e$\u003c/span\u003e15.88 per patient\u003c/p\u003e\u003cp\u003e- GPT-3.5: \u003cspan\u003e$\u003c/span\u003e1.59 per patient\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eWong et al, 2023\u003csup\u003e47\u003c/sup\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e- Article-specific patient and trial data\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eGPT-4 (3-shot)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colspan=\"2\" nameend=\"c5\" namest=\"c4\"\u003e\u003cp\u003ePerformed but metrics not provided\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colspan=\"2\" nameend=\"c7\" namest=\"c6\"\u003e\u003cp\u003eGPT-3.5 (zero-shot)\u003c/p\u003e\u003cp\u003ePrecision: 88.5\u003c/p\u003e\u003cp\u003eRecall: 11.6\u003c/p\u003e\u003cp\u003eF1: 20.6\u003c/p\u003e\u003cp\u003eGPT-4 (zero-shot)\u003c/p\u003e\u003cp\u003ePrecision: 86.7\u003c/p\u003e\u003cp\u003eRecall: 46.8\u003c/p\u003e\u003cp\u003eF1: 60.8\u003c/p\u003e\u003cp\u003eGPT-4 (3-shot)\u003c/p\u003e\u003cp\u003ePrecision: 87.6\u003c/p\u003e\u003cp\u003eRecall: 67.3\u003c/p\u003e\u003cp\u003eF1: 76.1\u003c/p\u003e\u003cp\u003e*Note: article describes binary classification task of eligible/not eligible so could be labeled either \u0026ldquo;patient-trial\u0026rdquo; or \u0026ldquo;trial-patient\u0026rdquo; matching\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003e- Proprietary: GPT-3.5\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c9\"\u003e\u003cp\u003eN/A\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eWoo, 2024\u003csup\u003e48\u003c/sup\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e\u0026minus;\u0026thinsp;2018 n2c2\u003c/p\u003e\u003cp\u003e- Article-specific patient and trial data\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eLlama-3.1-8B-All\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colspan=\"2\" nameend=\"c5\" namest=\"c4\"\u003e\u003cp\u003eLlama-3.1-8B-All on MIMIC-IV-based data set:\u003c/p\u003e\u003cp\u003eBalanced accuracy: 0.93\u003c/p\u003e\u003cp\u003eMicro-F1: 0.94\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003eN/A\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003eN/A\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003e- Open-access: Llama-3.1-8B, Llama3.1-70B\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c9\"\u003e\u003cp\u003eCost of evaluating the Apixaban criteria for 10,000.patients:\u003c/p\u003e\u003cp\u003e- Llama-3.1-8B: \u003cspan\u003e$\u003c/span\u003e929\u003c/p\u003e\u003cp\u003e- Llama3.1-70B: \u003cspan\u003e$\u003c/span\u003e4066\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eWornow et al., 2024\u003csup\u003e24\u003c/sup\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e\u0026minus;\u0026thinsp;2018 n2c2\u003c/p\u003e\u003cp\u003e\u0026minus;\u0026thinsp;1 novel exclusion criterion from a pulmonary arterial hypertension clinical trial\u003c/p\u003e\u003cp\u003e- SIGIR 2016\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eZero-shot GPT-4 with ACIN prompting strategy\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colspan=\"2\" nameend=\"c5\" namest=\"c4\"\u003e\u003cp\u003e- Precision: 0.91\u003c/p\u003e\u003cp\u003e- Recall: 0.92\u003c/p\u003e\u003cp\u003e- Macro-F\u003csub\u003e1\u003c/sub\u003e: 0.81\u003c/p\u003e\u003cp\u003e- Micro-F\u003csub\u003e1\u003c/sub\u003e: 0.93\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003eN/A\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003eN/A\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003e- Proprietary: GPT-3.5\u003c/p\u003e\u003cp\u003e- Open-source: Llama-2-70b, Mixtral-8x7B, Qwen2-72b, and Llama-3-70b\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c9\"\u003e\u003cp\u003e- Cost to screen a single patient: ~\u003cspan\u003e$\u003c/span\u003e1.55\u003c/p\u003e\u003cp\u003e- Evaluation time per patient: ~1 minute\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eYuan et al., 2024\u003csup\u003e49\u003c/sup\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e- Article-specific patient and trial data\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eLLM-based patient-trial matching with GPT-4 (LLM-PTM)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colspan=\"2\" nameend=\"c5\" namest=\"c4\"\u003e\u003cp\u003e- Precision: 0.964\u003c/p\u003e\u003cp\u003e- Recall: 0.862\u003c/p\u003e\u003cp\u003e- F\u003csub\u003e1\u003c/sub\u003e Score: 0.910\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e- Precision: 0.801\u003c/p\u003e\u003cp\u003e- Recall: 0.830\u003c/p\u003e\u003cp\u003e- F\u003csub\u003e1\u003c/sub\u003e Score: 0.815\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003eN/A\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003eN/A\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c9\"\u003e\u003cp\u003eN/A\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eZhuang et al., 2024\u003csup\u003e50\u003c/sup\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e- TREC 2022 CT\u003c/p\u003e\u003cp\u003e- TREC 2023 CT\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colspan=\"2\" nameend=\"c4\" namest=\"c3\"\u003e\u003cp\u003eHybrid PubmedBERT-based retriever with GPT-4 re-ranking\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003eN/A\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003eBi-encoder dense retriever:\u003c/p\u003e\u003cp\u003eNDCG@10: 0.5768\u003c/p\u003e\u003cp\u003eP@10: 0.3243\u003c/p\u003e\u003cp\u003eRecall@1000: 0.3670\u003c/p\u003e\u003cp\u003ePLADEv2 sparse retriever:\u003c/p\u003e\u003cp\u003eNDCG@10: 0.5971\u003c/p\u003e\u003cp\u003eP@10: 0.3243\u003c/p\u003e\u003cp\u003eRecall@1000: 0.3482\u003c/p\u003e\u003cp\u003eHybrid:\u003c/p\u003e\u003cp\u003e- NDCG@10: 0.5763\u003c/p\u003e\u003cp\u003e- P@10: 0.2946\u003c/p\u003e\u003cp\u003e- Recall@1000: 0.3878\u003c/p\u003e\u003cp\u003eCE_weighted:\u003c/p\u003e\u003cp\u003e- NDCG@10: 0.6716\u003c/p\u003e\u003cp\u003e- P@10: 0.4432\u003c/p\u003e\u003cp\u003e- Recall@1000: 0.3878\u003c/p\u003e\u003cp\u003eGPT-4:\u003c/p\u003e\u003cp\u003e\u0026nbsp;\u0026nbsp;\u0026nbsp; - NDCG@10: 0.7363\u003c/p\u003e\u003cp\u003e\u0026nbsp;\u0026nbsp;\u0026nbsp; - P@10: 0.5108\u003c/p\u003e\u003cp\u003e\u0026nbsp;\u0026nbsp;\u0026nbsp; - Recall@1000: 0.3878\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003eN/A\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003eGPT-3.5 Turbo (proprietary) to generate extra training data\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c9\"\u003e\u003cp\u003eN/A\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eZihang et al., 2025\u003csup\u003e51\u003c/sup\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e- Article-specific patient and trial data\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colspan=\"2\" nameend=\"c4\" namest=\"c3\"\u003e\u003cp\u003ellama3-70b-instruct\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003ePerformed but metrics not provided\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003eAccuracy:\u003c/p\u003e\u003cp\u003e- GLM-3-Turbo: 0.8973\u003c/p\u003e\u003cp\u003e- GLM-4: 0.9139\u003c/p\u003e\u003cp\u003e- llama3-70b-instruct: 0.9285\u003c/p\u003e\u003cp\u003e- Qwen-Turbo: 0.9166\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003eN/A\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003eGLM-3-Turbo, GLM-4, Qwen-Turbo\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c9\"\u003e\u003cp\u003eN/A\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003c/tbody\u003e\u003c/colgroup\u003e\u003ctfoot\u003e\u003ctr\u003e\u003ctd colspan=\"9\"\u003e* 2018 n2c2 denotes 2018 National Natural Language Processing Clinical Challenges cohort; ACIN, All criteria, individual notes: All notes are merged into a single prompt, but the model assesses one criterion at a time, requiring reprompting for each criterion; CLA: criterion-level accuracy; MAKA, Multi-Agents for Knowledge Augmentation; MCC\u0026thinsp;=\u0026thinsp;Mathew correlation coefficient; MRR\u0026thinsp;=\u0026thinsp;mean reciprocal rank; NCDG@10: Normalized Cumulative Discounted Gain; NDCG@10: Normalized Discounted Cumulative Gain; NPV: negative predictive value; PPV: positive predictive value; Program-rather-than-prompt: Ensures responses adhere to required format using structured programming objects instead of free-text prompts; QGMT, Query Generation, Medical Role \u0026amp; Task Description: Provides contextual information to ChatGPT and instructs it to generate a single keyword-based query; RAG, Retrieval-Augmented Generation; SIGIR 2016, Special Interest Group on Information Retrieval 2016 cohort; TCRR, Topical and Criteria Re-Ranking involves a two-step training schema focusing on topical relevance and eligibility classification; TREC 2021/2022/2023 CT, Text Retrieval Conference Clinical Trials Track.\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd colspan=\"9\"\u003e\u003cb\u003e\u0026dagger;\u003c/b\u003e Models compared within a given study include any models that were assessed by the authors alongside their best-performing LLM.\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd colspan=\"9\"\u003e**Additional metrics were calculated for TREC 2022 due to spacing concerns.\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd colspan=\"9\"\u003eSource: Data compiled from the studies included in this systematic review.\u003c/td\u003e\u003c/tr\u003e\u003c/tfoot\u003e\u003c/table\u003e\u003c/div\u003e\u003c/p\u003e\u003c/div\u003e\n\u003ch3\u003eModel Pipeline Design\u003c/h3\u003e\n\u003cp\u003eMost LLM-based trial-matching algorithms involve four stages: (1) acquisition of patient and trial data; (2) data pre-processing; (3) retrieval of relevant information from patient records; and (4) matching patients to trials (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e2\u003c/span\u003eA).\u003c/p\u003e\u003cp\u003e\u003c/p\u003e\n\u003ch3\u003eData Sets\u003c/h3\u003e\n\u003cp\u003ePipeline inputs varied by study; patient data included case reports/vignettes, longitudinal prescription and medical claims data, medical records, and synthetic admission notes (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e2\u003c/span\u003eB; Table\u0026nbsp;\u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e). Trial data were primarily drawn from ClinicalTrials.gov, pre-processed trial datasets, and hospital or international databases (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e2\u003c/span\u003eB; Table\u0026nbsp;\u003cspan refid=\"Tab3\" class=\"InternalRef\"\u003e3\u003c/span\u003e). Most data provided to model pipelines was unstructured text (Table\u0026nbsp;\u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e). For articles where LLMs were fine-tuned, ground truth was generally derived from historical enrollments, researcher-determined matches or other LLMs (Table\u0026nbsp;\u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e\u0026ndash;\u003cspan refid=\"Tab3\" class=\"InternalRef\"\u003e3\u003c/span\u003e).\u003c/p\u003e\u003cp\u003e\u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab2\" border=\"1\"\u003e\u003ccaption language=\"En\"\u003e\u003cdiv class=\"CaptionNumber\"\u003eTable 2\u003c/div\u003e\u003cdiv class=\"CaptionContent\"\u003e\u003cp\u003ePatient Data Used for Training and Testing of LLMs\u003c/p\u003e\u003c/div\u003e\u003c/caption\u003e\u003ccolgroup cols=\"7\"\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c7\" colnum=\"7\"\u003e\u003c/div\u003e\u003cthead\u003e\u003ctr\u003e\u003cth align=\"left\" colname=\"c1\"\u003e\u003cp\u003eStudy Dataset\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c2\"\u003e\u003cp\u003eSource of Data\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c3\"\u003e\u003cp\u003eData Description\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c4\"\u003e\u003cp\u003eData Size\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c5\"\u003e\u003cp\u003ePatient Characteristics\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c6\"\u003e\u003cp\u003eData Type\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c7\"\u003e\u003cp\u003ePublicly Available?\u003c/p\u003e\u003c/th\u003e\u003c/tr\u003e\u003c/thead\u003e\u003ctbody\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e2018 n2c2\u003csup\u003e21\u003c/sup\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e2014 i2b2/UTHealth shared tasks\u003csup\u003e\u003cspan citationid=\"CR52\" class=\"CitationRef\"\u003e52\u003c/span\u003e,\u003cspan citationid=\"CR53\" class=\"CitationRef\"\u003e53\u003c/span\u003e\u003c/sup\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eReal, de-identified clinical notes; synthetic clinical trial\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e288 patients (2\u0026ndash;5 records/patient); 13 predefined inclusion criteria\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003eDiabetes\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003eUnstructured text\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003eYes\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eBeattie et al., 2024\u003csup\u003e54\u003c/sup\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eEnrollment list of phase II trial investigating hypo fractionated radiation therapy for head and neck cancer; head and neck radiation oncology team\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eReal patient notes from surgical oncology, radiation oncology and medical oncology from the last 6 months; relevant pathology reports, imaging reports and lab results going back up to 1 year\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e35 patients enrolled in trial; 40 randomly identified patients seen by head and neck radiation oncology team\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003eHead and neck cancer\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003eUnstructured text\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003eNot Reported\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eCerami et al., 2024\u003csup\u003e28\u003c/sup\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eDana-Farber Cancer Institute (DFCI) Oncology Data Retrieval System\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e- Retrospective trial enrollment data set: Real EHR data (all unstructured clinical notes, imaging reports, and pathology reports) for all adults who enrolled in cancer treatment clinical trials at DFCI from January 2016 to April 2024\u003c/p\u003e\u003cp\u003e- Standard of care (SOC) treatment dataset: Real data for patients who started SOC systemic therapies at DFCI from 2016 to 2024\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e- Retrospective trial enrollment data set: 16,139 enrollments for 13,425 patients onto 1,534 clinical trials\u003c/p\u003e\u003cp\u003e- SOC treatment dataset: 86,042 treatment plans for 50,799 patients\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003eCancer, including breast, lung, lymphoma and leukemia among others\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e- Unstructured and structured text\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003eNo\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eChowdhury et al., 2024\u003csup\u003e29\u003c/sup\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eMayo Clinic\u0026rsquo;s United Data Platform\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eReal patient EHR data\u003c/p\u003e\u003cp\u003eStructured EHR: Six types of clinical events (diagnosis, medication, allergy, family history of medical condition, lab tests, and admission (e.g., reason for visit)\u003c/p\u003e\u003cp\u003eUnstructured EHR: radiology reports\u003c/p\u003e\u003cp\u003eDemographics: age and gender.\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e180 patients\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e50 out of 180 patients were eligible for the clinical trial\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003eUnstructured and structured text\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003eNo\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eDevi et al., 2024\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eSynthetic\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003ecase reports/patient descriptions\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e100 patients\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003eNon-small lung cancer\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003ePatient descriptions\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003eUnknown\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eFerber et al., 2024\u003csup\u003e32,33\u003c/sup\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eSynthetic; based on fictional patient vignettes\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eSynthetic oncology-focused patient EHRs (describe patient diagnoses, comorbidities, molecular data, imaging descriptions from staging CT or MRI scans, patient history)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e51 records\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003eCancer, including lung adenocarcinoma\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003eUnstructured text\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003eYes\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eGueguen et al, 2025\u003csup\u003e34\u003c/sup\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eLocal Molecular Tumor Board at Centre L\u0026eacute;on B\u0026eacute;rard, France\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eReal tumor board information on sequential patients; clinical data from full EHR and molecular data from molecular programmes\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e34 patients in the LLM subset\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003eMixed adult solid tumors\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003eUnstructured text\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003eYes; anonymized data available at:\u003c/p\u003e\u003cp\u003e\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://github.com/crcl-tm2/\u003c/span\u003e\u003cspan address=\"https://github.com/crcl-tm2/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e trialmatch-tool-evaluation/blob/master/artifacts/data_raw/formatted_ data.csv\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eGui et al, 2025\u003csup\u003e35\u003c/sup\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eThe Hepatology Department of the First Affiliated Hospital of Guangxi University of Chinese Medicine\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eReal, de-identified narrative admission notes\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e16,000 notes from over the last 10 years\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003eHepatopathy\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003eUnstructured text\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003eYes; can be obtained from corresponding author upon reasonable request\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eGupta et al., 2024\u003csup\u003e14\u003c/sup\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eInstitutional clinical research data warehouse from a single cancer center\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eReal, de-identified EHR and clinical trial enrollment information, including: assessment \u0026amp; plan note, brief op note, consults, discharge instructions/ summary, h\u0026amp;p, op note, OR surgeon, procedures, progress notes, rad onc simulation, rad onc weekly review\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e- Q\u0026amp;A Data: Notes from 50 patients\u003c/p\u003e\u003cp\u003e- Clinical Data: Notes from over 5790 patients.\u003c/p\u003e\u003cp\u003e- To evaluate patient-trial matching: 98 cancer patients\u003c/p\u003e\u003cp\u003e- To evaluate trial-patient matching: ~ 1\u0026ndash;3 patients who enrolled in each trial, 5\u0026ndash;21 who did not\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003eCancer, including breast and lung\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003eUnstructured text\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003eNo\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eLai et al., 2024\u003csup\u003e39\u003c/sup\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eThe Pancreas Center\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eReal de-identified medical oncology note, and surgical oncology note if available, from each patient who was screened for clinical trials at the Pancreas Center between January and May 2024\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e32 patients\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003ePancreatic cancer; 19 out of 24 patients in the test set were eligible for at least one trial\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003eUnstructured text\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003eNot Reported\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eSIGIR 2016\u003csup\u003e20\u003c/sup\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eAdopted from TREC CDS\u003csup\u003e\u003cspan citationid=\"CR55\" class=\"CitationRef\"\u003e55\u003c/span\u003e\u003c/sup\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eReal case reports\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e60 case reports\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003eDiagnosis varied\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003eUnstructured text\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003eYes\u003c/p\u003e\u003cp\u003e- Available at \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://data.csiro.au/collection/csiro:17152\u003c/span\u003e\u003cspan address=\"https://data.csiro.au/collection/csiro:17152\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eTREC 2021CT,\u003csup\u003e17\u003c/sup\u003e\u003c/p\u003e\u003cp\u003eTREC 2022 CT\u003csup\u003e18\u003c/sup\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e\"cases created by individuals with medical training\"\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eSynthetic admissions notes\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e- TREC 2021 CT: 75 topics\u003c/p\u003e\u003cp\u003e- TREC 2022 CT: 50 topics\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003eDiagnosis varied\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003eUnstructured text\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003eYes\u003c/p\u003e\u003cp\u003e- TREC 2021 CT:\u003c/p\u003e\u003cp\u003eAvailable at \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://www.trec-cds.org/2021.html\u003c/span\u003e\u003cspan address=\"http://www.trec-cds.org/2021.html\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/p\u003e\u003cp\u003e- TREC 2022 CT:\u003c/p\u003e\u003cp\u003eAvailable at \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://www.trec-cds.org/2022.html\u003c/span\u003e\u003cspan address=\"http://www.trec-cds.org/2022.html\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eTREC 2023 CT\u003csup\u003e19\u003c/sup\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eSynthetic\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eSynthetic questionnaire data\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e451,538 documents for 50 patients\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003eGlaucoma, anxiety, COPD, breast cancer, COVID-19, rheumatoid arthritis, sickle cell anemia, type 2 diabetes\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003eStructured text\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003eYes\u003c/p\u003e\u003cp\u003eAvailable at \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.trec-cds.org/2023.html\u003c/span\u003e\u003cspan address=\"https://www.trec-cds.org/2023.html\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eUnlu et al., 2024\u003csup\u003e46\u003c/sup\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eThe ongoing Co-Operative Program for Implementation of Optimal Therapy in Heart Failure Trial\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eReal EHR data over 2 years from ongoing trial including: progress notes, discharge summaries, history and physical, telephone encounters, notes to patients sent through the portal\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e1891 patients\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003eHigh rates of symptomatic heart failure\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003eUnstructured text\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003eNo\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eWong, 2023\u003csup\u003e47\u003c/sup\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eEHR from a collaborating health system\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eNLP-based extraction on structured data from EHR\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e523 patient-trial enrollment pairs\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003ePatient-trial enrollment pairs pulled from historical enrollment data\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003eStructured text\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003eNo\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eWoo, 2024 (MIMIC-III based)\u003csup\u003e\u003cspan citationid=\"CR48\" class=\"CitationRef\"\u003e48\u003c/span\u003e\u003c/sup\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eLLM-generated based on MIMIC-III\u003csup\u003e\u003cspan citationid=\"CR56\" class=\"CitationRef\"\u003e56\u003c/span\u003e\u003c/sup\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eLlama-3.1-70B-Instruct was prompted to generate eligibility-criteria-based Q\u0026amp;A pairs based on discharge summaries from MIMIC-III\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e1,000 questions\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003eDiagnosis varied\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003eStructured text\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003eYes\u003c/p\u003e\u003cp\u003e\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://physionet.org/content/mimic-ext-synth-trial-question/1.0.0/\u003c/span\u003e\u003cspan address=\"https://physionet.org/content/mimic-ext-synth-trial-question/1.0.0/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eWoo, 2024 (MIMIC-IV based)\u003csup\u003e\u003cspan citationid=\"CR48\" class=\"CitationRef\"\u003e48\u003c/span\u003e\u003c/sup\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eHuman-generated based on MIMIC-IV\u003csup\u003e57\u003c/sup\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eResearchers wrote 23 questions similar to eligibility criteria modeled after the ARISTOTLE Apixaban vs. Warfarin trial, then manually annotated answers from each MIMIC-IV discharge summary\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e23 questions resembling eligibility criteria with a random sample of 100 patient notes from MIMIC-IV\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003eDiagnosis varied\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003eStructured text\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003eYes\u003c/p\u003e\u003cp\u003e\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://physionet.org/content/mimic-iv-ext-apixaban-trial/1.0.0/\u003c/span\u003e\u003cspan address=\"https://physionet.org/content/mimic-iv-ext-apixaban-trial/1.0.0/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eYuan et al., 2024\u003csup\u003e49\u003c/sup\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eStroke patient database\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eReal, longitudinal prescription and medical claims data (diagnoses, procedures, medications)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003eClaims data for 825 patients\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003eEach patient enrolled in 1/6 stroke trials\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003eUnstructured text\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003eNot Reported\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eZhuang et al., 2024\u003csup\u003e50\u003c/sup\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eChatGPT\u003c/p\u003e\u003cp\u003eTREC 2022 CT\u003c/p\u003e\u003cp\u003eTREC 2023 CT\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e- ChatGPT data: Patient descriptions generated for randomly selected clinical trials\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e~\u0026thinsp;20,000 synthetic patient-trial pairs (25,000 total once they combined it with 5,000 patient description-trial pairs from TREC 2022 and 2023)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003eDiagnosis varied\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003eUnstructured Text\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003eNot Reported\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eZihang et al., 2025\u003csup\u003e51\u003c/sup\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eLLM-generated\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eSimulated answers to eligibility questionnaires\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e85 patients\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003eDiagnosis varied\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003eStructured\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003eAvailable upon request\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003c/tbody\u003e\u003c/colgroup\u003e\u003ctfoot\u003e\u003ctr\u003e\u003ctd colspan=\"7\"\u003e*2018 n2c2 denotes 2018 National Natural Language Professing Clinical Challenges cohort; COPILOT-HF, Cooperative Program for ImpLementation of Optimal Therapy in Heart Failure; MGB, Mass General Brigham; MSKCC, Memorial Sloan Kettering Cancer Center; PAH, pulmonary arterial hypertension; RAG, retrieval-augmented generation; SIGIR 2016, Special Interest Group on Information Retrieval 2016 cohort; TCGA, The Cancer Genome Atlas; TREC 2021/2022/2023 CT, Text REtrieval Conference Clinical Trials Track\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd colspan=\"7\"\u003eSource: Data compiled from the studies included in this systematic review.\u003c/td\u003e\u003c/tr\u003e\u003c/tfoot\u003e\u003c/table\u003e\u003c/div\u003e\u003c/p\u003e\u003cp\u003e\u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab3\" border=\"1\"\u003e\u003ccaption language=\"En\"\u003e\u003cdiv class=\"CaptionNumber\"\u003eTable 3\u003c/div\u003e\u003cdiv class=\"CaptionContent\"\u003e\u003cp\u003eClinical Trial Data Used for Training and Testing of LLMs\u003c/p\u003e\u003c/div\u003e\u003c/caption\u003e\u003ccolgroup cols=\"4\"\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e\u003cthead\u003e\u003ctr\u003e\u003cth align=\"left\" colname=\"c1\"\u003e\u003cp\u003eStudy Dataset\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c2\"\u003e\u003cp\u003eData Source\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c3\"\u003e\u003cp\u003eData Description\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c4\"\u003e\u003cp\u003eDiagnoses Covered in Data Set\u003c/p\u003e\u003c/th\u003e\u003c/tr\u003e\u003c/thead\u003e\u003ctbody\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eBeattie et al., 2024\u003csup\u003e27\u003c/sup\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003ePhase II trial investigating hypofractionated radiation therapy for head and neck cancer\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e14 criteria identified based on trial inclusion/exclusion criteria for evaluation using the LLMs\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003eHead and neck cancer\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eCerami et al., 2024\u003csup\u003e28\u003c/sup\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eDana-Farber Harvard Cancer Center clinical trial database;\u003c/p\u003e\u003cp\u003eClinicalTrials.gov\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eTherapeutic clinical trials open at Dana-Farber Harvard Cancer Center from January 2012 to June 2024;\u003c/p\u003e\u003cp\u003e500 trials open for cancer diagnoses on October 22, 2024\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003eCancer\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eChowdhury et al., 2024\u003csup\u003e29\u003c/sup\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eClinicalTrials.gov\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eFive clinical trials (NCT02008357, NCT04468659, NCT02669433, NCT01767909, NCT02565511)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003eCognitive disorders, such as Alzheimer's disease and Dementia with Lewey Bodies\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eDevi et al., 2024\u003csup\u003e32\u003c/sup\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eClinicaltrials.gov\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e10 U.S. drug-only interventional clinical trials\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003eNon Small Cell Lung Cancer\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eFerber et al., 2024\u003csup\u003e33\u003c/sup\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eClinicalTrials.gov\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e105,600 real clinical trials filtered for cancer collected on May 13, 2024\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003eCancer\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eGui et al, 2025\u003csup\u003e35\u003c/sup\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eSix real-world clinical trials on hepatopathy: ChiCTR2100044187, NCT04353193, NCT04850534, NCT03911037, NCT01311167, NCT04021056\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e\u0026ldquo;Selected 58 criteria assessable using information typically documented in admission notes\u0026rdquo;\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003echronic liver failure, cirrhosis, hepatocellular carcinoma\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eGupta et al., 2024\u003csup\u003e14\u003c/sup\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eClinicalTrials.gov\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eFor patient-trial matching: set of 10 trials for each patient, each for the same cancer type that were actively recruiting when patient enrolled in clinical trial (1/10 of those trials was one in which the patient actually enrolled)\u003c/p\u003e\u003cp\u003eFor trial-patient matching: Set of 36 clinical trials that all recruited patients from the same institution\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003eCancer\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eLai et al., 2024\u003csup\u003e39\u003c/sup\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eClinicalTrials.gov\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eReal trial data from nine ongoing clinical trials at the Pancreas Center\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003ePancreatic cancer\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eLin et al., 2024\u003csup\u003e40\u003c/sup\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e- ClinicalTrials.gov\u003c/p\u003e\u003cp\u003e- PubMed Central\u003c/p\u003e\u003cp\u003e- TrialAlign\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eClinicalTrials.gov: 467,944 trials, real data\u003c/p\u003e\u003cp\u003e- ChiCTR (China): 76,186 trials, real data\u003c/p\u003e\u003cp\u003e- TrialAlign: 14 sources of trial documents\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003eDiagnoses varied\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003ePROTECTOR1 dataset\u003csup\u003e\u003cspan citationid=\"CR58\" class=\"CitationRef\"\u003e58\u003c/span\u003e\u003c/sup\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eClinicalTrials.gov\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e764 phase 3 cancer trials\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003eCancer\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eRuan et al., 2024\u003csup\u003e44\u003c/sup\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eClinicalTrials.gov\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e301 trials\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003eFatty liver disease\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eSIGIR 2016\u003csup\u003e20\u003c/sup\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eClinicalTrials.gov\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e204,855 trials\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003eDiagnosis varied\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eTREC 2021CT,\u003csup\u003e17\u003c/sup\u003e\u003c/p\u003e\u003cp\u003eTREC 2022 CT\u003csup\u003e18\u003c/sup\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eClinicalTrials.gov\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e375,581 clinical trial descriptions\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003eDiagnosis varied\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eTrialAlign\u003csup\u003e\u003cspan citationid=\"CR40\" class=\"CitationRef\"\u003e40\u003c/span\u003e\u003c/sup\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e14 sources, including ClinicalTrials.gov and ChiCTR (China)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e793,279 trial documents\u003c/p\u003e\u003cp\u003e1,113,207 scientific papers related to clinical trials\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003eDiagnosis varied\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eUnlu et al., 2024\u003csup\u003e46\u003c/sup\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eData from the ongoing Co-Operative Program for Implementation of Optimal Therapy in Heart Failure Trial in the Microsoft Dynamics 365 used by staff for patient screening\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e13 trial criteria\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003eHeart failure\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eWong, 2023\u003csup\u003e47\u003c/sup\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eClinicalTrials.gov\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e53 treatment-oriented, interventional trials\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003eCancer\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eYuan et al., 2024\u003csup\u003e49\u003c/sup\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eClinicalTrials.gov\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e6 trials\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003eStroke\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eZihang et al., 2025\u003csup\u003e51\u003c/sup\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eClinicalTrials.gov\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e579 clinical trial registration entries with Fudan University as the sponsor; 562 pre-recruitment questionnaires were generated based on 562 trials with complete data\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003eDiagnosis varied\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003c/tbody\u003e\u003c/colgroup\u003e\u003ctfoot\u003e\u003ctr\u003e\u003ctd colspan=\"4\"\u003e*MSKCC denotes Memorial Sloan Kettering Cancer Center. All articles that used TREC CT patient data sets pulled trial protocol data from clinicaltrials.gov unless otherwise specified.\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd colspan=\"4\"\u003eSource: Data compiled from the studies included in this systematic review.\u003c/td\u003e\u003c/tr\u003e\u003c/tfoot\u003e\u003c/table\u003e\u003c/div\u003e\u003c/p\u003e\u003cp\u003eDiseases represented in these patient data sets include cancer, stroke, diabetes, cognitive disorders, liver disease, heart failure, or a mix of conditions (Table\u0026nbsp;\u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e). No articles used non-text data like images, and only Unlu et al and Lai et al. used data from ongoing randomized clinical trials. Eleven studies utilized a selection of two or more clinical note types, which could be real or synthetic (Tables\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e\u0026ndash;\u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e). For studies that used real EHR, selecting a subset of note types or from a set of note dates allowed them to select for only the most relevant notes while reducing the inputs to the LLM, increasing efficiency. Sixteen studies used publicly available synthetic or de-identified data (TREC CT 2021-2023\u003csup\u003e17\u0026ndash;19\u003c/sup\u003e, SIGIR 2016\u003csup\u003e20\u003c/sup\u003e, 2018 n2c2\u003csup\u003e21\u003c/sup\u003e), eleven used real patient data, and three used case reports/vignettes (Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e\u0026ndash;\u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e). Notably, Woo et al used Llama-3.1-70B-Instruct to generate eligibility-criteria-based question/answer pairs based on discharge summaries from the publicly-available MIMIC-III data set, while Zihang et al used an LLM to generate simulated answers to patient questionnaires as their patient data inputs.\u003c/p\u003e\u003cp\u003eRybinski et al used the TREC CT 2021 and 2022 for training and validation, and TREC CT 2023 for testing. Since the latter includes data from earlier years, model performance metrics might be influenced by data leakage.\u003c/p\u003e\n\u003ch3\u003eData Pre-Processing\u003c/h3\u003e\n\u003cp\u003eData pre-processing pipelines, especially those used for pipelines run on clinical notes, generally included some combination of four stages: enrichment, query generation, retrieval, and restructuring (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e2\u003c/span\u003eC).\u003c/p\u003e\u003cp\u003eEnrichment refers to the process of augmenting raw text with structured information that enhances semantic interpretability and downstream utility for computational models. In the context of trial matching, enrichment transforms unstructured patient or trial data into a machine-readable representation by extracting key medical concepts, standardizing terminology, and/or clarifying contextual meaning\u0026mdash;thereby enabling accurate and efficient alignment between patient profiles and trial eligibility criteria.\u003c/p\u003e\u003cp\u003eWhile enrichment implementation varied, it generally involved one or more of the following techniques: (1) \u003cb\u003eEntity extraction\u003c/b\u003e: Identifies clinically meaningful phrases like \"Stage IV non-small cell lung cancer\" as discrete concepts to enable structured matching. (2) \u003cb\u003eConcept normalization\u003c/b\u003e: Maps colloquial or alternate terms to standardized medical terminology\u0026mdash;for example, recognizing \u0026ldquo;high blood pressure\u0026rdquo; as equivalent to \u0026ldquo;hypertension\u0026rdquo;\u0026mdash;to prevent missed matches due to lexical variation. (3) \u003cb\u003eNegation detection\u003c/b\u003e: Ensures that the model correctly interprets statements containing negation, such as interpreting \u0026ldquo;no history of diabetes\" as the patient not having diabetes, so the patient is not matched to a diabetes-related trial.\u003c/p\u003e\u003cp\u003eOnce enriched, this structured data could be used to generate search queries that retrieve relevant patient or trial segments. Two key query optimization methods were common: (1) \u003cb\u003eQuery expansion\u003c/b\u003e: Adds related terms or synonyms to broaden the search scope (e.g., expanding \u0026ldquo;lung cancer\u0026rdquo; to include \u0026ldquo;NSCLC\u0026rdquo; or \u0026ldquo;pulmonary neoplasms\u0026rdquo;). (2) \u003cb\u003eQuery synthesis\u003c/b\u003e: Repackages the extracted and enriched data into coherent, task-specific prompts or structured queries suitable for the retrieval model.\u003c/p\u003e\u003cp\u003eRetrieval involves extracting relevant portions of patient records based on the generated queries. By reducing the data ultimately input into the LLM, retrieval serves two purposes; (1) it keeps the data size within a given LLM\u0026rsquo;s context window and (2) it reduces computational costs by using a smaller, less computationally intensive model to reduce the data the larger, more computationally-intensive LLM ultimately needs to process. Retrieval-augmented generation (RAG) is commonly used to identify portions of the queried text most semantically similar to the query; however, this approach can lose chronological context, which is essential to properly analyze patient data. A variety of approaches could be used to circumvent this issue; for instance, one pipeline combined lexical (BM25) and semantic (MedCPT) retrieval strategies to better preserve clinical chronology.\u003csup\u003e\u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e22\u003c/span\u003e\u003c/sup\u003e Lexical retrieval served to identify relevant clinical trials by matching exact or near-exact keywords from a set of synthetic data from a given patient to the text in trial eligibility criteria, thereby capturing surface-level correspondences in terminology and phrasing. Semantic retrieval, on the other hand, encoded both patient data and trial descriptions into dense vector representations using pretrained language models, enabling the retrieval of conceptually aligned information even when lexical overlap is limited.\u003c/p\u003e\u003cp\u003eFor prompt-based trial matching pipelines, the last data pre-processing step is restructuring. This restructuring generally involves taking the retrieved chunks\u0026mdash;usually consisting of one or multiple trial criteria and the corresponding extracted patient information\u0026mdash;and reformatting them into a prompt provided to a model to request the model perform patient-criterion, patient-trial, or trial-patient eligibility.\u003c/p\u003e\u003cp\u003ePre-processed data is then usually fed into one of two matching pipeline types (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e2\u003c/span\u003eD): (1) \u003cb\u003ePrompt-based\u003c/b\u003e: Prompts are provided to an LLM, the responses from which are subsequently integrated into article-specific scoring methods to generate ranked outputs. (2) \u003cb\u003eEmbedding-based\u003c/b\u003e: Patient and trial data were embedded in a shared vector space and directly ranked using similarity scores.\u003c/p\u003e\u003cp\u003eLLMs are also applied to tasks beyond eligibility assessment or mapping into a shared vector space\u0026mdash;they have been used to convert free text patient or trial data into structured formats like JSONs for use for downstream models; generate synthetic datasets for training smaller models; enrich data sets; and enhance retrieval through query generation, expansion, or synthesis (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e2\u003c/span\u003eE). Some models provided rationale for matching decisions or cited specific sentences to support their determination (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e2\u003c/span\u003eD).\u003csup\u003e\u003cspan additionalcitationids=\"CR23\" citationid=\"CR22\" class=\"CitationRef\"\u003e22\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e24\u003c/span\u003e\u003c/sup\u003e This feature not only improved model performance, but increased transparency and the ability of end users to assess accuracy. Depending on the matching type, these pipelines are designed for different end-users \u0026mdash; while patient-trial matching systems can generate a list of trials for a patient or their care team, trial-patient matching can help review patient data to extract a list of potentially eligible patients for coordinators of a particular trial (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e2\u003c/span\u003eE).\u003c/p\u003e\n\u003ch3\u003eModel Performance \u0026 Cost Analysis\u003c/h3\u003e\n\u003cp\u003eAlthough comparisons across studies were hindered by heterogeneous tasks and data, intra-study comparisons showed GPT-4 consistently outperformed traditional models for inclusion/exclusion extraction and patient-to-trial matching compared to open-access models, including ones that had been fine-tuned (Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e). Only a few studies reported on cost and efficiency. In a study by Jin et al, their TrialGPT pipeline reduced clinician screening time by 42.6%.\u003csup\u003e22\u003c/sup\u003e Several studies reported costs associated with LLM-assisted trial matching lower than traditional human screening methods (Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e). In a study by Gupta et al, a Qwen-1.5-14B-based pipeline, OncoLLM, incurred a per-patient cost of only \u003cspan\u003e$\u003c/span\u003e0.17 per patient-trial pair, as compared to the cost of \u003cspan\u003e$\u003c/span\u003e6.18 per pair incurred by GPT-4 (Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e). For pipelines using either GPT-3.5 or GPT-4, screening cost per-patient ranged between \u003cspan\u003e$\u003c/span\u003e0.02 and \u003cspan\u003e$\u003c/span\u003e15.88 per patient-trial pair (Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e). Of note\u0026mdash;in work performed by Unlu et al, use of RAG methods and other strategies dropped that \u003cspan\u003e$\u003c/span\u003e15.88 value to only 2 cents/patient for GPT-4 (Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e). For finely-tuned models, training cost is also a consideration; Gupta et al put the cost of training their smaller, fine-tuned model OncoLLM at \u003cspan\u003e$\u003c/span\u003e2,688. Overall, per-patient processing tended to be fast, with studies reporting between 1 and 12.4 minutes to evaluate a given patient for a trial (Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e).\u003c/p\u003e"},{"header":"Discussion","content":"\u003cp\u003eThere is a pressing need to improve patient screening for patient trials, both to accelerate scientific discovery and to expand access to potentially life-saving treatments. However, the expensive, time-consuming nature of patient matching limits the speed of this process and contributes to limited diversity of patients who ultimately enroll. LLM-based pipelines have the potential to offer a more flexible and scalable alternative to previous rule-based matching systems, reducing the workload for trial staff and facilitating clinical trial success.\u003c/p\u003e\u003cp\u003eOne barrier to implementation is that some models struggle with the nuanced and intricate nature of EHR data. Most trial-matching models are trained on data that is synthetic, simplified, abridged, or in structured data formats that, while more easily read by LLMs, do not fully represent the complexities of real-world, long-context EHR data. Their reliance on narrowly defined variables from trial criteria or patient records can limit their generalizability, and real-world applicability. Ideally, trial matching data sets should closely parallel the same real-world data that models will be used with when implemented in the clinic. A key barrier to achieving this goal is the high cost required to collect large, diverse datasets annotated by experts, especially since PHI concerns restrict data sharing. Despite these limitations, the success of models applied to diverse unstructured datasets, such as medical claims data and admission records, demonstrates the versatility of LLMs more broadly—a quality that is likely to increase with ongoing advancements in model architecture and capabilities.\u003c/p\u003e\u003cp\u003eThe effectiveness of trial-matching systems hinges not only on model architecture but also on the integration of stakeholder expertise into their design and evaluation. Developing successful pipelines requires collaboration with key stakeholders, such as trial coordinators, who can provide insights into how eligibility criteria are weighted in practice and how patients are evaluated for enrollment. This input can be used both to refine model performance and to ensure that trial-matching tools integrate seamlessly into existing workflows. Currently, GPT-4 offers the strongest performance in trial-matching tasks, with high zero-shot capabilities reducing the need for annotated datasets and advanced retrieval strategies helping to offset computational costs (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e1\u003c/span\u003e). For institutions unable to deploy GPT-4 within HIPAA-compliant infrastructure, smaller LLMs present a viable alternative when fine-tuned for specific tasks, offering a balance between accuracy, cost, and data privacy (\u003cb\u003eFig.\u0026nbsp;3\u003c/b\u003e). Using smaller models can also avoid issues with shifts in model performance and behavior over time, which can be seen with proprietary models like GPT-4.\u003c/p\u003e\u003cp\u003e\u003c/p\u003e\u003cp\u003eStandardization of evaluation metrics and evaluation on high-quality public benchmarks is crucial to facilitate comparison of model pipeline performance. For patient-to-criterion matching, metrics such as accuracy, sensitivity, specificity, recall and F\u003csub\u003e1\u003c/sub\u003e scores are appropriate. Similar metrics can be used for assessment of patient-to-trial or trial-to-patient matching capabilities, in addition to metrics like NDCG@10, Precision@10, AUROC and AURPC to assess ranking performance. Accuracy of LLM-generated explanations for eligibility determinations compared to qualified staff should also be determined. Assessing the performance of different pipelines on the same public benchmarks would also facilitate accurate comparisons. However, data memorization by LLMs is also a potential issue when using public data sets, so it is important to also use private data sets for model validation.\u003c/p\u003e\u003cp\u003eSeveral steps can be taken to mitigate concerns of model errors and bias. To safeguard against model hallucinations, LLMs can be instructed to provide rationales for their matching decisions, along with explicit citations from the input. These explanations both improve model accuracy and facilitate the correction of model mistakes by human users. Trial staff should also remain a key component of patient assessment, ensuring critical clinical details are not overlooked. Moreover, human-in-the-loop validation studies are necessary to assess whether LLM-based matching pipelines improve human performance or change human behavior. Model performance across diverse patient groups should be assessed, and safeguards should be put in place to prevent models amplifying bias present in training data.\u003c/p\u003e\u003cp\u003eEarly trial matching systems show promise in ultimately reducing the time spent by staff in assessing eligibility, as well as the overall cost of trial recruitment. Future work should also assess the cost of data collection and annotation, infrastructure to host the model, personnel to design and maintain it, and the choice of cloud service.\u003c/p\u003e\u003cp\u003eRecent advances in LLM-based systems offer improved generalizability and accuracy, with models being employed across diverse tasks. Implementation of data pre-processing strategies—such as enrichment, query generation, and retrieval—shows promise for improving performance and reducing computational costs. Intra-study comparisons suggest that GPT-4o achieves the strongest performance in trial-matching tasks, with high zero-shot capabilities reducing reliance on annotated datasets. For institutions unable to deploy GPT-4 within HIPAA-compliant infrastructure, smaller LLMs offer a viable alternative when fine-tuned for specific tasks. Future work should prioritize the use of standardized assessment metrics and test sets to enable cross-study comparison, as well as evaluate cost and bias compared to existing trial matching strategies. To ensure safe and effective implementation, models should be evaluated on real-world data and designed with transparency measures, such as LLM-generated eligibility justifications with sentences cited directly from records. Collaborative development with clinical stakeholders will be essential to ensure these tools augment, rather than replace, human decision-making and integrate effectively into real-world workflows.\u003c/p\u003e"},{"header":"Methods","content":"\u003ch2\u003eArticle Identification\u003c/h2\u003e\n\u003cp\u003eWe identified 126 unique citations across three academic databases and one preprint server using a structured keyword search combining trial-related and LLM-related terms between January 1\u003csup\u003est\u003c/sup\u003e , 2020 and March 1\u003csup\u003est\u003c/sup\u003e, 2025 (\u003cstrong\u003eFig. 1\u003c/strong\u003e). Three articles were published as pre-prints within this time frame, then published in peer-reviewed journals after March 1\u003csup\u003est\u003c/sup\u003e, 2025, reflecting a common trend of first-step publishing in pre-print servers. The search string included: (\u0026quot;match[tiab]\u0026quot; OR \u0026quot;screen*[tiab]\u0026quot;) AND (\u0026quot;clinical trial\u0026quot;[tiab] OR \u0026quot;clinical trials\u0026quot;[tiab]), combined with (\u0026quot;large language models\u0026quot;[tiab] OR \u0026quot;LLMs\u0026quot;[tiab] OR \u0026quot;large language model\u0026quot;[tiab] OR \u0026quot;LLM\u0026quot;[tiab] OR \u0026quot;ChatGPT\u0026quot;[tiab] OR \u0026quot;GPT-4\u0026quot;[tiab] OR \u0026quot;LLAMA\u0026quot;[tiab] OR \u0026quot;GPT-3.5\u0026quot;[tiab] OR \u0026quot;GPT-4\u0026quot;[tiab]). Covidence, a web-based review management tool, was employed to organize references, identify duplicate records, and document study inclusion/exclusion decisions.\u003csup\u003e25\u003c/sup\u003e After excluding 51 duplicates and 91 ineligible records, 31 full-text articles were included (\u003cstrong\u003eFig. 1\u003c/strong\u003e). All remaining records were accessible via full-text retrieval. Four were published in 2023, 23 in 2024, and four in 2025, reflecting rapid developments in the field.\u0026nbsp;\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eData availability\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eNo new data were generated or analyzed in support of this review. Data sharing is not applicable.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eCode availability\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eAll data supporting this review are available within the cited articles. No new data were created.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAuthor contributions\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eB.A.M.: Conceptualization; Methodology; Search strategy; Data curation; Screening; Extraction; Formal analysis; Visualization; Writing\u0026ndash;original draft; Writing\u0026ndash;review \u0026amp; editing.\u003c/p\u003e\n\u003cp\u003eM.S.: Methodology; Conceptualization; Supervision; Writing\u0026ndash;review \u0026amp; editing.\u003c/p\u003e\n\u003cp\u003eJ.S.Y.: Conceptualization; Supervision; Writing\u0026ndash;review \u0026amp; editing; Correspondence.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eCompeting interests\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eB.A.M.: No competing interests.\u003c/p\u003e\n\u003cp\u003eM.S.: No competing interests.\u003c/p\u003e\n\u003cp\u003eJ.S.Y.: No competing interests.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eFunding\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThis research received no specific grant from any funding agency in the public, commercial, or not-for-profit sectors.\u003c/p\u003e\u003ch2\u003eAcknowledgement\u003c/h2\u003e\u003cp\u003eOriginal art is the work of Mr. Kenneth Probst, Department of Neurosurgery, UCSF. The authors used the Elicit application (Ought, https://elicit.org) during the early scoping stage to roughly identify which articles contained certain types of information and to help organize thematic areas for review. UCSF Versa was used to assist with conciseness. No text or analysis generated by the tool appears in the submitted manuscript; all synthesis, interpretation, and writing were performed by the authors.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eWen PY, Weller M, Lee EQ, Alexander BM, Barnholtz-Sloan JS, Barthel FP, et al. Glioblastoma in adults: a Society for Neuro-Oncology (SNO) and European Society of Neuro-Oncology (EANO) consensus review on current management and future directions. Neuro-Oncol. 2020;22(8):1073\u0026ndash;113.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eYu W, Zhou D, Meng F, Wang J, Wang B, Qiang J, et al. The global, regional burden of pancreatic cancer and its attributable risk factors from 1990 to 2021. BMC Cancer. 2025;25(1):186.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eMani K, Deng D, Lin C, Wang M, Hsu ML, Zaorsky NG. Causes of death among people living with metastatic cancer. Nat Commun. 2024;15(1):1519.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eTaylor K, Francesca P, Cru MJ, Ronte H, Haughey J. Intelligent clinical trials [Internet]. [cited 2025 Jan 26]. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www2.deloitte.com/content/dam/insights/us/articles/22934_intelligent-clinical-trials/DI_Intelligent-clinical-trials.pdf\u003c/span\u003e\u003cspan address=\"https://www2.deloitte.com/content/dam/insights/us/articles/22934_intelligent-clinical-trials/DI_Intelligent-clinical-trials.pdf\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eKasenda B, von Elm E, You J, Bl\u0026uuml;mle A, Tomonaga Y, Saccilotto R, et al. Prevalence, Characteristics, and Publication of Discontinued Randomized Trials. JAMA. 2014;311(10):1045\u0026ndash;52.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eWong AR, Sun V, George K, Liu J, Padam S, Chen BA, et al. Barriers to Participation in Therapeutic Clinical Trials as Perceived by Community Oncologists. JCO Oncol Pract. 2020 Sept;16(9):e849\u0026ndash;58.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eDurden K, Hurley P, Butler DL, Farner A, Shriver SP, Fleury ME. Provider motivations and barriers to cancer clinical trial screening, referral, and operations: Findings from a survey. Cancer. 2024;130(1):68\u0026ndash;76.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003ePenberthy LT, Dahman BA, Petkov VI, DeShazo JP. Effort Required in Eligibility Screening for Clinical Trials. J Oncol Pract. 2012;8(6):365\u0026ndash;70.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eGuerra CE, Fleury ME, Byatt LP, Lian T, Pierce L. Strategies to Advance Equity in Cancer Clinical Trials. Am Soc Clin Oncol Educ Book Am Soc Clin Oncol Annu Meet. 2022;42:1\u0026ndash;11.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eYuan C, Ryan PB, Ta C, Guo Y, Li Z, Hardin J, et al. Criteria2Query: a natural language interface to clinical databases for cohort definition. J Am Med Inform Assoc. 2019;26(4):294\u0026ndash;305.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eThadani SR, Weng C, Bigger JT, Ennever JF, Wajngurt D. Electronic Screening Improves Efficiency in Clinical Trial Recruitment. J Am Med Inform Assoc JAMIA. 2009;16(6):869\u0026ndash;73.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eShriver SP, Arafat W, Potteiger C, Butler DL, Beg MS, Hullings M, et al. Feasibility of institution-agnostic, EHR-integrated regional clinical trial matching. Cancer. 2024;130(1):60\u0026ndash;7.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eHaddad TC, Helgeson J, Pomerleau K, Makey M, Lombardo P, Coverdill S, et al. Impact of a cognitive computing clinical trial matching system in an ambulatory oncology practice. J Clin Oncol. 2018;36(15_suppl):6550\u0026ndash;6550.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eGupta S, Basu A, Nievas M, Thomas J, Wolfrath N, Ramamurthi A, et al. PRISM: Patient Records Interpretation for Semantic clinical trial Matching system using large language models. Npj Digit Med. 2024;7(1):1\u0026ndash;12.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eShi H, Zhang J, Zhang K. Enhancing Clinical Trial Patient Matching through Knowledge Augmentation with Multi-Agents [Internet]. arXiv; 2024 [cited 2025 Jan 19]. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://arxiv.org/abs/2411.14637\u003c/span\u003e\u003cspan address=\"http://arxiv.org/abs/2411.14637\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eUnlu O, Shin J, Mailly CJ, Oates MF, Tucci MR, Varugheese M, et al. Retrieval-Augmented Generation\u0026ndash;Enabled GPT-4 for Clinical Trial Screening. NEJM AI. 2024 June 27;1(7):AIoa2400181.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003e2021 TREC Clinical Trials Track [Internet]. [cited 2025 Jan 15]. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://www.trec-cds.org/2021.html#documents\u003c/span\u003e\u003cspan address=\"http://www.trec-cds.org/2021.html#documents\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003e2022 TREC Clinical Trials Track [Internet]. [cited 2025 Jan 15]. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.trec-cds.org/2023.html\u003c/span\u003e\u003cspan address=\"https://www.trec-cds.org/2023.html\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003e2023 TREC Clinical Trials Track [Internet]. [cited 2025 Jan 15]. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.trec-cds.org/2023.html\u003c/span\u003e\u003cspan address=\"https://www.trec-cds.org/2023.html\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eKoopman B, Zuccon G. A Test Collection for Matching Patients to Clinical Trials. In: Proceedings of the 39th International ACM SIGIR conference on Research and Development in Information Retrieval [Internet]. Pisa Italy: ACM; 2016 [cited 2025 Jan 4]. p. 669\u0026ndash;72. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://dl.acm.org/doi/10.1145/2911451.2914672\u003c/span\u003e\u003cspan address=\"https://dl.acm.doi/10.1145/2911451.2914672\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eStubbs A, Filannino M, Soysal E, Henry S, Uzuner \u0026Ouml;. Cohort selection for clinical trials: n2c2 2018 shared task track 1. J Am Med Inform Assoc. 2019;26(11):1163\u0026ndash;71.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eJin Q, Wang Z, Floudas CS, Chen F, Gong C, Bracken-Clarke D, et al. Matching patients to clinical trials with large language models. Nat Commun. 2024;15(1):9074.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eNievas M, Basu A, Wang Y, Singh H. Distilling large language models for matching patients to clinical trials. J Am Med Inform Assoc JAMIA. 2024 Sept 1;31(9):1953\u0026ndash;63.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eWornow M, Lozano A, Dash D, Jindal J, Mahaffey KW, Shah NH. Zero-Shot Clinical Trial Patient Matching with LLMs. NEJM AI. 2024;0(0):AIcs2400360.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eCovidence - Better systematic review management [Internet]. Covidence. [cited 2025 Sept 4]. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.covidence.org/\u003c/span\u003e\u003cspan address=\"https://www.covidence.org/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eBeattie J, Neufeld S, Yang D, Chukwuma C, Gul A, Desai N, et al. Utilizing Large Language Models for Enhanced Clinical Trial Matching: A Study on Automation in Patient Screening. Cureus. 2024;16(5):e60044.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eBeattie J, Owens D, Navar AM, Giuliani Schmitt L, Taing K, Neufeld S, et al. ChatGPT augmented clinical trial screening. Mach Learn Health. 2025 July;1(1):015005.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eCerami E, Trukhanov P, Paul MA, Hassett MJ, Riaz IB, Lindsay J, et al. MatchMiner-AI: An Open-Source Solution for Cancer Clinical Trial Matching [Internet]. arXiv; 2024 [cited 2025 Jan 11]. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://arxiv.org/abs/2412.17228\u003c/span\u003e\u003cspan address=\"http://arxiv.org/abs/2412.17228\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eChowdhury S, Rajaganapathy S, Yu Y, Tao C, Vassilaki M, Zong N. Matching Patients to Clinical Trials using LLaMA 2 Embeddings and Siamese Neural Network. medRxiv. 2024 June 30;2024.06.28.24309677.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eDatta S, Lee K, Huang LC, Paek H, Gildersleeve R, Gold J, et al. Patient2Trial: From Patient to Participant in Clinical Trials Using Large Language Models. Inform Med Unlocked. 2025;101615.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eDevi A, Uttrani S, Singla A, Jha S, Dasgupta N, Natarajan S, et al. Quantitative Analysis of GPT-4 model: Optimizing Patient Eligibility Classification for Clinical Trials and Reducing Expert Judgment Dependency. In: Proceedings of the 2024 8th International Conference on Medical and Health Informatics [Internet]. New York, NY, USA: Association for Computing Machinery; 2024 [cited 2024 Dec 23]. p. 230\u0026ndash;7. (ICMHI \u0026rsquo;24). Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://dl.acm.org/doi/10.1145/3673971.3674014\u003c/span\u003e\u003cspan address=\"https://dl.acm.doi/10.1145/3673971.3674014\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eDevi A, Uttrani S, Singla A, Jha S, Dasgupta N, Natarajan S, et al. Automating Clinical Trial Eligibility Screening: Quantitative Analysis of GPT Models versus Human Expertise. In: Proceedings of the 17th International Conference on PErvasive Technologies Related to Assistive Environments [Internet]. New York, NY, USA: Association for Computing Machinery; 2024 [cited 2024 Dec 22]. p. 626\u0026ndash;32. (PETRA \u0026rsquo;24). Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://dl.acm.org/doi/10.1145/3652037.3663922\u003c/span\u003e\u003cspan address=\"https://dl.acm.doi/10.1145/3652037.3663922\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eFerber D, Hilgers L, Wiest IC, Le\u0026szlig;mann ME, Clusmann J, Neidlinger P, et al. End-To-End Clinical Trial Matching with Large Language Models [Internet]. arXiv; 2024 [cited 2025 Jan 19]. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://arxiv.org/abs/2407.13463\u003c/span\u003e\u003cspan address=\"http://arxiv.org/abs/2407.13463\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eGueguen L, Olgiati L, Brutti-Mairesse C, Sans A, Le Texier V, Verlingue L. A prospective pragmatic evaluation of automatic trial matching tools in a molecular tumor board. Npj Precis Oncol. 2025;9(1):28.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eGui X, Lv H, Wang X, Lv L, Xiao Y, Wang L. Enhancing hepatopathy clinical trial efficiency: a secure, large language model-powered pre-screening pipeline. BioData Min. 2025 June 14;18:42.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eJullien M, Bogatu A, Unsworth H, Freitas A. Controlled LLM-based Reasoning for Clinical Trial Retrieval [Internet]. arXiv; 2024 [cited 2025 Jan 19]. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://arxiv.org/abs/2409.18998\u003c/span\u003e\u003cspan address=\"http://arxiv.org/abs/2409.18998\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eKusa W, Mendoza \u0026Oacute;E, Knoth P, Pasi G, Hanbury A. Effective matching of patients to clinical trials using entity extraction and neural re-ranking. J Biomed Inform. 2023;144:104444.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eKusa W, Styll P, Seeliger M, Espitia Mendoza \u0026Oacute;, Hanbury A. DoSSIER at TREC 2023 Clinical Trials Track. In: # PLACEHOLDER_PARENT_METADATA_VALUE# [Internet]. NIST; 2023 [cited 2025 Jan 19]. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://repositum.tuwien.at/handle/20.500.\u003c/span\u003e\u003cspan address=\"https://repositum.tuwien.at/handle/20.500.\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e12708/203878\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eLai SM, Malik AM, Sathe TS, Silvestri CJ, Manji GA, Kluger MD. A Proof-of-Concept Large Language Model Application to Support Clinical Trial Screening in Surgical Oncology [Internet]. medRxiv; 2024 [cited 2025 Feb 21]. p. 2024.09.20.24314053. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.medrxiv.org/content/\u003c/span\u003e\u003cspan address=\"https://www.medrxiv.org/content/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1101/2024.09.20.24314053v2\u003c/span\u003e\u003cspan address=\"10.1101/2024.09.20.24314053v2\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eLin J, Xu H, Wang Z, Wang S, Sun J. Panacea: A foundation model for clinical trial search, summarization, design, and recruitment [Internet]. medRxiv; 2024 [cited 2024 Dec 23]. p. 2024.06.26.24309548. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.medrxiv.org/content/\u003c/span\u003e\u003cspan address=\"https://www.medrxiv.org/content/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1101/2024.06.26.24309548v1\u003c/span\u003e\u003cspan address=\"10.1101/2024.06.26.24309548v1\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003ePeikos G, Symeonidis S, Kasela P, Pasi G. Utilizing ChatGPT to Enhance Clinical Trial Enrollment [Internet]. arXiv; 2023 [cited 2025 Jan 18]. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://arxiv.org/abs/2306.02077\u003c/span\u003e\u003cspan address=\"http://arxiv.org/abs/2306.02077\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003ePeikos G, Kasela P, Pasi G. Leveraging Large Language Models for Medical Information Extraction and Query Generation [Internet]. arXiv; 2024 [cited 2025 Aug 27]. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://arxiv.org/abs/2410.23851\u003c/span\u003e\u003cspan address=\"http://arxiv.org/abs/2410.23851\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eRahmanian M, Fakhrahmad SM, Mousavi SZ. Towards Efficient Patient Recruitment for Clinical Trials: Application of a Prompt-Based Learning Model [Internet]. arXiv.org. 2024 [cited 2025 Feb 1]. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://arxiv.org/abs/2404.16198v1\u003c/span\u003e\u003cspan address=\"https://arxiv.org/abs/2404.16198v1\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eRuan J, Su Q, Chen Z, Huang J, Li Y. CPRS: a clinical protocol recommendation system based on LLMs. Int J Med Inf. 2024;195:105746.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eRybinski M, Kusa W, Karimi S, Hanbury A. Learning to match patients to clinical trials using large language models. J Biomed Inform. 2024;159:104734.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eUnlu O, Shin J, Mailly CJ, Oates MF, Tucci MR, Varugheese M, et al. Retrieval Augmented Generation Enabled Generative Pre-Trained Transformer 4 (GPT-4) Performance for Clinical Trial Screening. MedRxiv Prepr Serv Health Sci. 2024;2024.02.08.24302376.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eWong C, Zhang S, Gu Y, Moung C, Abel J, Usuyama N, et al. Scaling Clinical Trial Matching Using Large Language Models: A Case Study in Oncology [Internet]. arXiv; 2023 [cited 2025 Jan 11]. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://arxiv.org/abs/2308.02180\u003c/span\u003e\u003cspan address=\"http://arxiv.org/abs/2308.02180\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eWoo EG, Burkhart MC, Alsentzer E, Beaulieu-Jones BK. Synthetic data distillation enables the extraction of clinical information at scale. Npj Digit Med. 2025;8(1):267.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eYuan J, Tang R, Jiang X, Hu X. Large Language Models for Healthcare Data Augmentation: An Example on Patient-Trial Matching. AMIA Annu Symp Proc. 2024;2023:1324\u0026ndash;33.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eZhuang S, Koopman B, Zuccon G. Team IELAB at TREC Clinical Trial Track 2023: Enhancing Clinical Trial Retrieval with Neural Rankers and Large Language Models [Internet]. arXiv; 2024 [cited 2024 Dec 21]. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://arxiv.org/abs/2401.01566\u003c/span\u003e\u003cspan address=\"http://arxiv.org/abs/2401.01566\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eZiHang C, QianMin S, GaoYi C, JiHan H, Ying L. Enhanced Pre-Recruitment Framework for Clinical Trial Questionnaires Through the Integration of Large Language Models and Knowledge Graphs [Internet]. Rochester, NY: Social Science Research Network; 2024 [cited 2025 Jan 19]. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://papers.ssrn.com/abstract=4713177\u003c/span\u003e\u003cspan address=\"https://papers.ssrn.com/abstract=4713177\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eA S, \u0026Ouml; U. Annotating longitudinal clinical narratives for de-identification: The 2014 i2b2/UTHealth corpus. J Biomed Inform [Internet]. 2015 Dec [cited 2025 Aug 27];58 Suppl(Suppl). Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://pubmed.ncbi.nlm.nih.gov/26319540/\u003c/span\u003e\u003cspan address=\"https://pubmed.ncbi.nlm.nih.gov/26319540/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eA S, \u0026Ouml; U. Annotating risk factors for heart disease in clinical narratives for diabetic patients. J Biomed Inform [Internet]. 2015 Dec [cited 2025 Aug 27];58 Suppl(Suppl). Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://pubmed.ncbi.nlm.nih.gov/26004790/\u003c/span\u003e\u003cspan address=\"https://pubmed.ncbi.nlm.nih.gov/26004790/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eBeattie J, Owens D, Navar AM, Schmitt LG, Taing K, Neufeld S, et al. Large Language Model Augmented Clinical Trial Screening [Internet]. medRxiv; 2024 [cited 2024 Dec 22]. p. 2024.08.27.24312646. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.medrxiv.org/content/\u003c/span\u003e\u003cspan address=\"https://www.medrxiv.org/content/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1101/2024.08.27.24312646v1\u003c/span\u003e\u003cspan address=\"10.1101/2024.08.27.24312646v1\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eVoorhees EM, Hersh W. Overview of the TREC 2012 Medical Records Track. NIST [Internet]. 2013 June 28 [cited 2025 Aug 27]; Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.nist.gov/publications/overview-trec-2012-medical-records-track\u003c/span\u003e\u003cspan address=\"https://www.nist.gov/publications/overview-trec-2012-medical-records-track\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eWoo E, Burkhart MC, Alsentzer E, Beaulieu-Jones B. MIMIC-III-Ext-Synthetic-Clinical-Trial-Questions [Internet]. PhysioNet; [cited 2025 Sept 2]. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://physionet.org/content/mimic-ext-synth-trial-question/1.0.0/\u003c/span\u003e\u003cspan address=\"https://physionet.org/content/mimic-ext-synth-trial-question/1.0.0/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eWoo E, Burkhart MC, Alsentzer E, Beaulieu-Jones B. MIMIC-IV-Ext-Apixaban-Trial-Criteria-Questions [Internet]. PhysioNet; [cited 2025 Sept 2]. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://physionet.org/content/mimic-iv-ext-apixaban-trial/1.0.0/\u003c/span\u003e\u003cspan address=\"https://physionet.org/content/mimic-iv-ext-apixaban-trial/1.0.0/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eYang Y, Jayaraj S, Ludmir E, Roberts K. Text Classification of Cancer Clinical Trial Eligibility Criteria. AMIA Annu Symp Proc. 2024;2023:1304\u0026ndash;13.\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"npj-digital-medicine","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"npjdigitalmed","sideBox":"Learn more about [npj Digital Medicine](http://www.nature.com/npjdigitalmed/)","snPcode":"41746","submissionUrl":"https://submission.springernature.com/new-submission/41746/3","title":"npj Digital Medicine","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"stoa","reportingPortfolio":"NPJ","inReviewEnabled":true,"inReviewRevisionsEnabled":true},"keywords":"Clinical trial matching, large language models, LLMs, systematic review, GPT-4, patient recruitment, automated matching, clinical trials, generative artificial intelligence, patient-trial matching, trial-patient matching, patient-criterion matching","lastPublishedDoi":"10.21203/rs.3.rs-8036235/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-8036235/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eMatching patients to clinical trial options is critical for identifying novel treatments, especially in oncology. However, manual matching is labor-intensive and error-prone, leading to recruitment delays. Pipelines incorporating large language models (LLMs) offer a promising solution. We conducted a systematic review of studies published between 2020 and 2025 from three academic databases and one preprint server, identifying LLM-based approaches to clinical trial matching. Of 126 unique articles, 31 met inclusion criteria. Reviewed studies focused on matching patient-to-criterion only (n\u0026thinsp;=\u0026thinsp;4), patient-to-trial only (n\u0026thinsp;=\u0026thinsp;10), trial-to-patient only (n\u0026thinsp;=\u0026thinsp;2), binary eligibility classification only (n\u0026thinsp;=\u0026thinsp;1) or combined tasks (n\u0026thinsp;=\u0026thinsp;14). Sixteen used synthetic data; fourteen used real patient data; one used both. Variability in datasets and evaluation metrics limited cross-study comparability. In studies with direct comparisons, the GPT-4 model consistently outperformed other models\u0026mdash;even finely-tuned ones\u0026mdash;in matching and eligibility extraction, albeit at higher cost. Promising strategies included zero-shot prompting with proprietary LLMs like the GPT-4o model, advanced retrieval methods, and fine-tuning smaller, open-source models for data privacy when incorporation of large models into hospital infrastructure is infeasible. Key challenges include accessing sufficiently large real-world data sets, and deployment-associated challenges such as reducing cost, mitigating risk of hallucinations, data leakage, and bias. This review synthesizes progress in applying LLMs to clinical trial matching, highlighting promising directions and key limitations. Standardized metrics, more realistic test sets, and attention to cost-efficiency and fairness will be critical for broader deployment.\u003c/p\u003e","manuscriptTitle":"A systematic review of trial-matching pipelines using large language models","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2025-11-20 16:03:25","doi":"10.21203/rs.3.rs-8036235/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"decision","content":"Revision requested","date":"2026-01-07T08:40:54+00:00","index":"","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2026-01-06T16:50:32+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"124124068038278379633166520329462814775","date":"2025-12-11T17:30:21+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"75205166754374005887812180332067815517","date":"2025-12-10T07:43:42+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2025-11-17T18:22:32+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"233079740227408476420756976310839180024","date":"2025-11-17T10:56:40+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"7258983503161405894890344004540318786","date":"2025-11-13T15:30:50+00:00","index":"hide","fulltext":""},{"type":"reviewersInvited","content":"","date":"2025-11-11T02:56:37+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2025-11-10T21:39:40+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2025-11-10T05:39:16+00:00","index":"","fulltext":""},{"type":"submitted","content":"npj Digital Medicine","date":"2025-11-05T08:38:37+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"npj-digital-medicine","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"npjdigitalmed","sideBox":"Learn more about [npj Digital Medicine](http://www.nature.com/npjdigitalmed/)","snPcode":"41746","submissionUrl":"https://submission.springernature.com/new-submission/41746/3","title":"npj Digital Medicine","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"stoa","reportingPortfolio":"NPJ","inReviewEnabled":true,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"ead7b476-0dff-413a-b4b2-85efc9112c15","owner":[],"postedDate":"November 20th, 2025","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"under-review","subjectAreas":[{"id":58137584,"name":"Biological sciences/Cancer"},{"id":58137585,"name":"Biological sciences/Computational biology and bioinformatics"},{"id":58137586,"name":"Health sciences/Health care"},{"id":58137587,"name":"Physical sciences/Mathematics and computing"},{"id":58137588,"name":"Health sciences/Medical research"}],"tags":[],"updatedAt":"2026-05-05T18:53:18+00:00","versionOfRecord":[],"versionCreatedAt":"2025-11-20 16:03:25","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-8036235","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-8036235","identity":"rs-8036235","version":["v1"]},"buildId":"8U1c8b4HqxoKbykW_rLl7","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00