An exploration of available methods and tools to improve the efficiency of systematic review production - a scoping review | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article An exploration of available methods and tools to improve the efficiency of systematic review production - a scoping review Lisa Affengruber, Miriam M. van der Maten, Isa Spiero, Barbara Nussbaumer-Streit, and 18 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-4595777/v1 This work is licensed under a CC BY 4.0 License Status: Published Journal Publication published 18 Sep, 2024 Read the published version in BMC Medical Research Methodology → Version 1 posted 11 You are reading this latest preprint version Abstract Background Systematic reviews (SRs) are time-consuming and labor-intensive to perform. With the growing number of scientific publications, the SR development process becomes even more laborious. This is problematic because timely SR evidence is essential for decision-making in evidence-based healthcare and policymaking. Numerous methods and tools that accelerate SR development have recently emerged. To date, no scoping review has been conducted to provide a comprehensive summary of methods and ready-to-use tools to improve efficiency in SR production. Objective To present an overview of primary studies that evaluated the use of ready-to-use applications of tools or review methods to improve efficiency in the review process. Methods We conducted a scoping review. An information specialist performed a systematic literature search in four databases, supplemented with citation-based and grey literature searching. We included studies reporting the performance of methods and ready-to-use tools for improving efficiency when producing or updating a SR in the health field. We performed dual, independent title and abstract screening, full-text selection, and data extraction. The results were analyzed descriptively and presented narratively. Results We included 103 studies: 51 studies reported on methods, 54 studies on tools, and 2 studies reported on both methods and tools to make SR production more efficient. A total of 72 studies evaluated the validity (n = 69) or usability (n = 3) of one method (n = 33) or tool (n = 39), and 31 studies performed comparative analyses of different methods (n = 15) or tools (n = 16). 20 studies conducted prospective evaluations in real-time workflows. Most studies evaluated methods or tools that aimed at screening titles and abstracts (n = 42) and literature searching (n = 24), while for other steps of the SR process, only a few studies were found. Regarding the outcomes included, most studies reported on validity outcomes (n = 84), while outcomes such as impact on results (n = 23), time-saving (n = 26), usability (n = 13), and cost-saving (n = 3) were less often evaluated. Conclusion For title and abstract screening and literature searching, various evaluated methods and tools are available that aim at improving the efficiency of SR production. However, only few studies have addressed the influence of these methods and tools in real-world workflows. Few studies exist that evaluate methods or tools supporting the remaining tasks. Additionally, while validity outcomes are frequently reported, there is a lack of evaluation regarding other outcomes. rapid review systematic review evidence synthesis scoping review method automation tools Figures Figure 1 Figure 2 Figure 3 Figure 4 Introduction Systematic reviews (SRs) provide the most valid way of synthesizing evidence as they follow a structured, rigorous, and transparent research process. Because of their thoroughness, SRs have a long history in informing health policy decision-making, clinical guidelines, and primary research [3]. However, the standard method employed in high-quality SRs involves many steps that are predominantly conducted manually, resulting in a laborious and time-intensive process lasting an average of fifteen months [4]. With the exponential growth of scientific literature, the challenge of SR development is exacerbated. Especially, this is the case in contexts where timely evidence is imperative and decision-makers need urgent answers, as demonstrated during the coronavirus pandemic [5]. Additionally, researchers’ aspirations to undertake SRs prior to initiating new primary studies is hindered by the complex and resource-intensive nature of the SR development process [6]. In response to these challenges, a surge of interest in methods to accelerate SR development occurred in recent years, leading to the emergence of “rapid reviews” (RRs). Methods used in RR development include searching a limited number of bibliographic databases, single-reviewer literature screening, or abbreviated quality assessment [7]. However, depending on various factors, the trade-off between the time saved and the potential reduction in quality and comprehensiveness is a critical issue that must be carefully weighed and discussed with stakeholders. Concurrently, efforts are underway to leverage technological innovations to expedite the research process, involving machine learning, natural language processing, active learning, or text mining to mimic human activities in SR tasks [8]. Supportive tools can offer varying levels of automation and decision-making, ranging from basic file management to fully automated decision-making, as outlined by O’Connor et al. [9]. These levels include semiautomated tools for workflow prioritization to fully automated decision-making processes [9]. Some tools offer researchers ready-to-use applications, while other algorithms are not yet developed into user-friendly tools [10]. Table 1 provides an overview of commonly used terms in this scoping review and their definitions. Table 1: Commonly used terms Term Definition Text mining Text mining is the process of extracting useful information from unstructured text data. It involves analyzing large amounts of text data to discover patterns, relationships, and insights that can be used to make informed decisions. Text mining techniques include natural language processing, machine learning, active learning, and statistical analysis [9, 11]. Machine learning Machine learning is a type of artificial intelligence that allows computer systems to automatically improve their performance at a specific task by learning from data without being explicitly programmed. It involves developing algorithms that can learn patterns and relationships in data and then use this knowledge to make predictions or decisions about new data [9]. Active learning Active learning is a machine learning technique that selects which data to learn from, rather than being trained on a prelabeled dataset. In active learning, the algorithm iteratively selects the most informative data points for labeling by a human expert and then uses these labeled data points to update its model. This process continues until the model achieves a desired level of accuracy or until labeling additional data points no longer improves the model’s performance [9]. Fully automated Full automation means that the tool makes decisions without human input, for example, the tool automatically conducts a risk of bias assessment of a randomized controlled trial without requiring input from a human reviewer [9, 12]. Semiautomated Semiautomated tools recommend decisions, but manual confirmation is required. For instance, the tool suggest whether a record is probably eligible for inclusion, but additional human judgment is necessary to make the final decision for in- or exclusion, which is useful in, for example, the prioritization of relevant abstracts and duplicate detection [9, 12]. Supportive tools without automation Tools that improve the file management process, for example, citation databases, reference management tools, and SR management tools, but do not include any form of automation [9]. Supportive tools Tools that use text mining, machine and active learning, and/or fully or semiautomated approaches, but also tools without automation to support RR production. Ready-to-use tools Tools that offer researchers ready-to-use, fully developed applications, while other algorithms are not yet developed into user-friendly tools. Crowdsourcing Utilizing a large group of human volunteers to perform a task online. Workload saving Workload saving refers to the reduction in the amount of work required to complete a task or process through abbreviated methods or tools in comparison to SR methodology. In this context, workload saving refers to reducing the number of references or documents that must be screened, reviewed, or analyzed during the review process. Time-saving Time-saving in this context refers to the reduction in the amount of time measured in hours or minutes required to complete the review process through abbreviated methods or supportive tools. Cost-saving Cost-saving in this context refers to the reduction in financial expenditures associated with conducting the review process. It mainly encompasses personnel costs required to complete the review. Impact on results Impact on results in this context refers to changes in statistical significance, effect estimates, or conclusions as a result of abbreviated methods or supporting tools during the review process. Validity outcomes Validity outcomes in this context describe the degree to which the method or tool is aligned with the SR methodology. Validity is measured as the tool’s precision, sensitivity, specificity, accuracy, errors made, etc., compared to SR methods done by humans. Usability Usability refers to the extent to which end users can use the method/tool in a specific context to achieve specific goals effectively, efficiently, and satisfactorily, and how the end users experienced it [13]. In the context of SR tools, usability is often reported as evaluation results (e.g., usability scores) of the method or tool or end-user experiences including joy of use, interface design, etc. Abbreviations: RR: rapid review, SR: systematic review While these methods and tools hold promise for enhancing the efficiency of SR production, their widespread adoption faces challenges, including limited awareness among review teams, concerns about validity, and usability issues [14]. Addressing these barriers requires evaluations to determine the validity and usability of various methods and tools across different stages of the review process [10, 15]. To bridge this gap, we conducted a scoping review to comprehensively map the landscape of methods and tools aimed at improving the efficiency of SR development, assessing their validity, resource utilization (workload/time/costs), and impact on results, as well as exploring usability for all steps of the review process. This review complements another scoping review that identified the most resource-intensive areas when conducting or updating a SR [2]. We mapped the efficiency outcomes of each method and tool against the steps of the SR process. Specifically, our scoping review aimed to answer the following key questions (KQs): Which methods and tools are used to improve the efficiency of SR production? How efficient are these methods or tools regarding validity, resource use, and impact on results? How was the user experience when using these methods and tools? Methods We conducted this review as part of working group 3 of EVBRES (EVidence-Bases RESearch) COST Action CA17117 (www.ebvres.eu). We published the protocol for this scoping review on June 18, 2020 (Open Science Framework: https://osf.io/9423z). Study design We conducted a scoping review following the guidance of Arksey and O’Malley [16], Levac et al. [17], and Peters et al. [7]. Within EVBRES, we adopted the definition of a scoping review as “a form of knowledge synthesis that addresses an exploratory research question aimed at mapping key concepts, types of evidence, and gaps in research related to a defined area or field by systematically searching, selecting, and synthesising existing knowledge” [18]. We report our review in accordance with the PRISMA Extension for Scoping Reviews (PRISMA-ScR) [19]. Information sources and search The search for this scoping review followed an iterative three-step process recommended by the Joanna Briggs Institute [20]: 1) First, an information specialist (RS) conducted a preliminary limited, focused search in Scopus in March 2020. We screened the search results and analyzed relevant studies to discover additional relevant keywords and sources. 2) Second, based on identified search terms from the included studies, the information specialist performed a comprehensive search (November 2021) in MEDLINE and Embase, both via Ovid. The comprehensive MEDLINE strategy was reviewed by another information specialist (IK) in accordance with the Peer Review of Electronic Search Strategies (PRESS) guideline [21]. 3) Third, we checked the reference lists of the identified studies and background articles, conducted grey literature searches (e.g., organizations that produce SRs and RRs), and contacted experts in the field. In addition, to identify grey literature, we searched for conference proceedings covered in Embase and checked if the associated full text was available. We limited the database searches to articles on methodological adaptations published since 1997, as this was the first year of mention in the published literature of methods to make the review process more efficient [22]. For tools, we limited the search to articles published since 2005, as this was the first year of mention of a text mining model in the published literature, according to Jonnalagadda and Petitti (2013) [23]. The search strategies are provided in Appendix 1. Our search was updated on December 14, 2023 to include evidence published since our initial search. Eligibility criteria The eligibility criteria are outlined in Table 2. Our focus was on incorporating primary studies that assessed the efficacy of automated, ready-to-use applications of tools, or RR methods within the SR process. Specifically, we sought tools that demand no programming expertise, relying instead on user-friendly interfaces devoid of complex codes, syntaxes, or algorithms. We were interested in studies assessing their use within one or more of the fifteen steps of the SR process as defined by Tsafnat et al. (2014), further supplemented by the steps “critical appraisal,” “grading the certainty of evidence,” and “administration/project management” [24] [2] (Figure 1). Table 2: Eligibility criteria for study inclusion in the scoping review Inclusion criteria Exclusion criteria Topic Studies reporting the performance of methods/ready-to-use tools with regard to improving efficiency when producing or updating a SR in the health field Studies not reporting on improving efficiency Studies assessing algorithms, classifiers, or models that are not usable as tools Concept Addressing improvement in efficiency within one or more steps of a SR or RR (as depicted in Figure 1) of health intervention or diagnostic or prognostic studies Studies assessing other steps such as the dissemination of results Outcomes Workload saved Time saved Personnel effort saved Costs saved Impact on results, conclusions Usability Validity outcomes: for example, recall (sensitivity), precision, accuracy, specificity, false positives, false negatives, reliability, proportion of relevant studies missed, prediction performance, ordering performance Any other outcomes Study design/ publication type Evaluation study of method/ready-to-use tool (e.g., method studies, RCTs, nRCTs, cohort studies,..) Development studies of ready-to-use tools Validation studies of of ready-to-use tools Evaluation studies of tools that are no longer available Full text availability Full text retrievable via our libraries Full text not retrievable via our libraries Timing Studies assessing methods: 1997 Studies assessing tools: 2005 Language All languages Abbreviations: nRCT: nonrandomized-controlled trial, RCT: randomized controlled trial, RR: rapid review, SR: systematic review Study selection We piloted the abstract screening with 50 records and the full-text screening with five records. Following the piloting, the results were discussed with all reviewers, and the screening guidance was updated to include clarifications wherever necessary. The review team used Covidence (www.covidence.org) to dually screen titles/abstracts and full texts. We resolved conflicts throughout the screening process through re-examination of the study and subsequent discussion and, if necessary, by consulting a third reviewer. Data charting We developed a data extraction form and pilot-tested it before implementation using Google Forms. The data abstraction was done by one author and checked by a second author to ensure consistency and correctness in the extracted data. A third author made final decisions in cases of discrepancies. We extracted relevant study characteristics and outcomes per review step. Data mapping We mapped the identified methods and tools by each review step and summarized the outcomes of individual studies. As the objective of this scoping review was to descriptively map efficiency outcomes and the usability of methods and tools against SR production steps, we did not apply a formal certainty of evidence or risk of bias (RoB) assessment. Additionally, we used data mapping to identify research gaps. Results We included 103 studies [12, 25-126] evaluating 21 methods (n=51) [25, 28, 29, 31, 33, 37, 39-42, 50, 52, 53, 55, 60, 61, 64-66, 72, 74, 76, 78, 79, 81, 82, 84-92, 99-101, 104, 106, 108-111, 113, 114, 116, 121, 122, 125, 126] and 35 tools (n=54) [12, 26, 27, 30, 32, 34-36, 38, 43-49, 51, 54, 56-59, 62, 63, 67-71, 73-75, 77, 80, 83, 93-98, 102, 103, 105, 107, 108, 112, 115, 117-120, 124, 127] (Figure 2: PRISMA study flowchart). Table 3 provides an overview of the identified methods and tools. A total of 73 studies were validity studies (n=70) [25-34, 37, 39, 43-46, 48, 49, 51, 52, 56-58, 60, 62-66, 69, 70, 75, 77-79, 81-85, 87-91, 94-97, 99, 101-103, 105-112, 114-117, 119-122, 124, 125] or usability studies (n=3) [59, 67, 68] assessing a single method or tool, and 30 studies performed comparative analyses of different methods or tools [12, 35, 36, 38, 40-42, 47, 50, 53-55, 61, 71-74, 76, 80, 86, 92, 93, 98, 100, 104, 108, 113, 118, 126, 127]. Few studies prospectively evaluated methods or tools in a real-world workflow (n=20) [12, 27, 32, 35, 46, 50, 67, 68, 77, 78, 88, 90, 94, 98, 105, 108, 112, 114, 125, 127], 7 studies of those used independent testing (by a different reviewer team) with external data [12, 35, 46, 94, 98, 112, 127]. The majority of studies evaluated methods or tools for supporting the tasks of title and abstract screening (n=42) [32, 35, 36, 38, 43-45, 47, 48, 51, 55, 58, 59, 63, 73, 79, 82-84, 86-90, 94-96, 100, 102, 105-108, 112-114, 117-120, 125, 127] or devising the search strategy and performing the search (n=24) [28, 34, 37, 39, 42, 52, 53, 56, 60, 64-66, 72, 81, 91-93, 98, 99, 104, 110, 121, 122, 126] (see Figure 2). For several steps of the SR process, only a few studies that evaluated methods or tools were identified: deduplication: n=6 [30, 36, 54, 57, 71, 80], additional search: n=2 [33, 97], update search: n=6 [36, 50, 61, 77, 109, 111], full-text selection: n=4 [85, 113, 114, 125], data extraction: n=11 [31, 36, 46, 67, 69, 70, 74, 103, 112, 124, 125]; critical appraisal: n=9, [26, 27, 36, 49, 62, 68, 75, 101, 115], and combination of abbreviated methods/tools: n=6 [12, 25, 76, 78, 100, 116] (see Figure 2). No studies were found for some steps of the SR process, such as administration/project management, formulating the review question, searching for existing reviews, writing the protocol, full-text retrieval, synthesis/meta-analysis, certainty of evidence assessment, and report preparation. In Appendix 2, we summarize the characteristics of all the included studies. Most studies reported on validity outcomes (n=84, 46%) [12, 25-34, 38, 41-58, 61-63, 65, 66, 68-75, 77, 79, 80, 82-90, 93, 95, 96, 98, 100-113, 115-124, 126], while outcomes such as workload saving (n=35, 19%) [12, 27, 28, 31, 32, 38, 43-45, 47, 48, 51, 58, 63, 66, 83-86, 90, 94, 96, 102, 104-108, 110, 113, 114, 116, 118, 120, 126], time-saving (n=24, 13%) [32-34, 38, 44, 46, 47, 50, 51, 69-71, 73, 74, 86, 90, 95-98, 103, 108, 124, 125]; impact on results (n=23, 13%) [25, 31, 37, 39-42, 58, 60, 61, 64, 76, 78, 81, 91, 92, 99, 100, 110, 114, 116, 121, 126], usability (n=13, 7%) [35, 36, 38, 59, 67, 68, 77, 93, 94, 96, 97, 120, 124] and cost-saving (n=3, 2%) [32, 82, 113] were less evaluated (Figure 4: Outcomes reported in the included studies). In Appendix 2, we map the efficiency and usability outcomes per tool and method against the review steps of the SR process.The included studies reported various validity outcomes (i.e., specificity, precision, accuracy) and time, costs, or workload savings to undertake the review. None of the studies reported the personnel effort saved. Figure 3 presents the frequency of reported outcomes in the included studies. Table 3: Identified methods and tools per review step Review step Methods Tools Administration/project management - - Review question formulation - - Search for existing reviews - - Protocol preparation - - Literature search Search strategy and database search Citation-based searching [28, 65, 66] Search restrictions: database [53, 92, 104, 121], language [37, 39, 60, 64, 81, 91, 99] Abbreviated search strategies for study type [52, 72, 110], topic [122], or search date [42, 126] MeSH on Demand [93, 98] PubReMiner [93, 98] Polyglot Search Translator [34] Risklick search platform [56] Yale MeSH Analyzer [93, 98] Deduplication - ASySD [57] EBSCO [71] EndNote [54, 57, 71, 80] Covidence [80] Deduklick [30] Mendely [54, 71, 80] OVID [71, 80] Rayyan [54, 80] Refworks [71] SRA- Deduplicator [36, 54, 57] Zotero [54, 80] Additional search Scopus approach [33] Paperfetcher [97] Update search Clinical Query search combined with PubMed-related articles search [109, 111] Clinical Query search in MEDLINE and Embase [61] · McMaster Premium LiteratUre Service [61] PubMed Similar Articles Search [50] Scopus Citation Tracking [50] RobotReviewer LIVE [36, 77] Study selection Title and abstract selection Crowdsourcing using different (automation) tools [84, 86-90] Dual computer monitors [125] Single-reviewer screening [55, 100, 113, 114] Limited screening: PICO-based title-only screening [106] Review of reviews [108] Title-first screening [79] AbstrackR [36, 38, 44, 45, 47, 48, 51, 59, 73, 107, 108, 118] ASReview [83, 95, 102, 120] ChatGPT [112] Colandr [38] Covidence [35, 36, 59] DistillerSR [36, 43, 47, 58] EPPI-reviewer [36, 59, 118, 127] Rayyan [35, 36, 38, 59, 73, 94, 96, 119, 127] RCT classifier [117] Research screener [32] RobotAnalyst [35, 36, 47, 105, 108] SRA-helper for EndNote [35] SWIFT-active screener [36, 63] SWIFT-Review [73] Full-text retrieval - - Full-text selection Crowdsourcing using different (automation-) tools [85] Dual computer monitors [125] Single-reviewer screening [113, 114] - Data extraction Dual computer monitors [125] Single-reviewer data extraction [31, 74] ChatGPT [112] Data Abstraction Assistant [36, 67, 74] Dextr [124] ExaCT [46, 70, 103] Plot Digitizer [69] Critical appraisal Crowdsourcing with CrowdCARE [101] RobotReviewer [26, 27, 36, 49, 62, 68, 75, 115] Synthesis/meta-analysis - - Certainty of evidence assessment - - Report preparation - - Combination of abbreviated methods/tools Rapid review methods with multiple steps combined: abbreviated search/search limits [25, 76, 100, 116], single-reviewer screening/data extraction [25, 116], review of reviews [78], abbreviated critical appraisal [116] Rapid review tools with multiple steps combined: SRA-Deduplicator, EndNote, Polyglot Search Translator, RobotReviewer, SRA-Helper [12] Abbreviations: ASReview: Automated Systematic Review, ASySd: Automated Systematic Synthesis and Data extraction, ChatGPT: Chat Generative Pre-trained Transformer (by OpenAI), EPPI: Evidence for Policy and Practice Information and Co-ordinating Centre, ExaCT: Extraction of Citations Tool, MeSH: Medical Subject Headings, PICO: Population, Intervention, Comparison, Outcome, PST: Polyglot Search Translator, PLUS: McMaster Premium LiteratUre Service, RCT: randomized controlled trial, SR: systematic review, SRA: Systematic Review Assistant, SWIFT: Sciome Workbench for Interactive Computer-Facilitated Text-mining Methods or tools for literature search Search strategy and database search Five tools (MeSH on Demand [93, 98], PubReMiner [93, 98], Polyglot Search Translator [34], Risklick search platform [56], and Yale MeSH Analyzer [93, 98]) and three methods (abbreviated search strategies study type [52, 72, 110], topic [122], or search date [42, 126]; citation-based searching [28, 65, 66]; search restrictions for database [53, 92, 104, 121] and language (e.g. only English articles) [37, 39, 60, 64, 81, 91, 99]) were evaluated in 24 studies [28, 34, 37, 39, 42, 52, 53, 56, 60, 64-66, 72, 81, 91-93, 98, 99, 104, 110, 121, 122, 126] to support devising search strategies and/or performing literature searches. Tools for search strategies Using text mining tools for search strategy development (MeSH on Demand, PubReMiner, and YaleMeSH Analyzer) reduced time expenditure compared to manual searches, with tools saving over half the time required for manual searches (5 hours, standard deviation [SD=2] vs. 12 hours [SD=8]) [98]. Using supportive tools such as Polyglot Search Translator [34], MeSH on Demand, PubReMiner, and YaleMeSH Analyzer [93, 98] was less sensitive [98] and showed a slightly reduced precision compared to manual searches (11% to 14% vs. 15%) [93]. The Risklick search platform demonstrated a high precision for identifying clinical trials (96%) and COVID-19–related publications (94%) [56]. User ratings by the study authors indicated that PubReMiner and YaleMeSH Analyzer were considered “useful” or “extremely useful,” while MeSH on Demand received the rating “not very useful” on a 5-point Likert scale (from extremely useful to least useful) [93]. Abbreviated search strategies for study type, topic, or search date Two studies evaluated an abbreviated search strategy (i.e., Cochrane highly sensitive search strategy) [52] and a brief RCT strategy [110] for identifying RCTs. Both achieved high sensitivity rates of 99.5% [52] and 94% [110] while reducing the number of records requiring screening by 16% [110]. Although some RCTs were missed using abbreviated search strategies, there were no significant differences in the conclusions [110]. One study [122] assessed an abbreviated search strategy using only the generic drug name to identify drug-related RCTs, achieving high sensitivities in both MEDLINE (99%) and Embase (99.6%) [122]. Lee et al. (2012) evaluated 31 search filters for SRs and meta-analyses, with the health-evidence.ca Systematic Review search filter performing best, maintaining a high sensitivity while reducing the number of articles needing for screening (90% in MEDLINE, 88% in Embase, 90% in CINAHL) [72]. Furuya-Kanamori et al. [42] and Xu et al. [126] investigated the impact of restricting search timeframes on effect estimates and found that limiting searches to the most recent 10 to 15 years resulted in minimal changes in effect estimates (<5%) while reducing workload by up to 45% [42, 126]. Nevertheless, this approach missed 21% to 35% of the relevant studies [42]. Citation-based searching Three studies [28, 65, 66] assessed whether citation-based searching can improve efficiency in systematic reviewing. Citation-based searching achieved a reduction in the number of retrieved articles (50% to 89% fewer articles) compared to the original searches while still capturing a substantial proportion of 75% to 82% of the included articles [28, 66]. Restricted database searching Seven studies assessed the validity of restricted database searching and suggested that searching at least two topic-related databases yielded high recall and precision for various types of studies [29, 40, 41, 53, 92, 104, 121]. Preston et al. (2015) demonstrated that searching only MEDLINE and Embase plus reference list checking identified 93% of the relevant references while saving 24% of the workload [104]. Beyer et al. (2013) emphasized the necessity of searching at least two databases along with reference list checking to retrieve all included studies [29]. Goossen et al. (2018) highlighted that combining MEDLINE with CENTRAL and hand searching was the most effective for RCTs (Recall: 99%), while for nonrandomized studies, combining MEDLINE with Web of Science yielded the highest recall (99.5%) [53]. Ewald et al. (2022) showed that searching two or more databases (MEDLINE/CENTRAL/Embase) reached a recall of ≥87.9% for identifying mainly RCTs [41]. Additionally, Van Enst et al. (2014) indicated that restricting searches to MEDLINE alone might slightly overestimate the results compared to broader database searches in diagnostic accuracy SRs (relative diagnostic odds ratio: 1.04; 95% confidence interval [CI], 0.95 to 1.15) [121]. Nussbaumer-Streit et al. (2018) and Ewald et al. (2020) found that combining one database with another or with searches of reference lists was noninferior to comprehensive searches (2%; 95% CI, 0% to 9%; if opposite concusion was of concern) [92] as the effect estimates were similar (ratio of odds ratios [ROR] median: 1.0 interquartile range [IQR]: 1.0–1.01) [40]. Restricted language searching Seven studies found that excluding non-English articles to reduce workload would minimally alter the conclusions or effect estimates of the meta-analyses. Two studies found no change in the overall conclusions [60, 91], and five studies [37, 60, 64, 91, 99] reported changes in the effect estimates or statistical significance of the meta-analyses. Specifically, the statistical significance of the effect estimates changed in 3% to 12% of the meta-analyses[37, 60, 64, 91, 99]. Deduplication Six studies [30, 36, 54, 57, 71, 80] compared eleven supportive software tools (ASySD, EBSCO, EndNote, Covidence, Deduklick, Mendeley, OVID, Rayyan, RefWorks, Systematic Review Accelerator, Zotero). Manual deduplication took approximately 4 hours 45 minutes, whereas using the tools reduced the time by 4 hours 42 minutes to only 3 minutes [71]. False negative duplicates varied from 36 (Mendeley) to 258 (EndNote), while false positives ranged from 0 (OVID) to 43 (EBSCO) [71]. The precision was high with 99% to 100% for Deduklick and ASySD, and the sensitivity was highest for Rayyan (ranging from 99% to 100%) [54, 57, 80], followed by Covidence, OVID, Systematic Review Accelerator, Mendeley, EndNote, and Zotero [54, 57, 80]. However, Cowie et al. reported that the Systematic Review Accelerator received a low rating of 9/30 for its features and usability [36]. Additional literature search Paperfetcher, identified as an application to automate additional searches as handsearching and citation searching, saved up to 92.0% of the time compared to manual handsearching and reference list checking, though validity outcomes for Paperfetcher were not reported [97]. Additionally, the Scopus approach, in which reviewers electronically downloaded the reference lists of relevant articles and screened only new references dually, saved approximately 62.5% of the time compared to manual checking [33]. Update literature search We identified one tool (RobotReviewer LIVE) [36, 77] and five methods (Clinical Query search combined with PubMed-related articles search, Clinical Query search in MEDLINE and Embase, searching the McMaster Premium LiteratUre Service [PLUS], PubMed similar articles search, and Scopus citation tracking) [50, 61, 109, 111] for improving the efficiency of updating literature searches. RobotReviewer LIVE showed a precision of 55% and a high recall of 100% [77] with limitations including search restricted to MEDLINE, consideration of only RCTs, and low usability scores for features [36, 77]. The Clinical Query (CQ) search, combined with the PubMed-related articles search and the CQ search in MEDLINE and Embase, exhibited high recall rates ranging from 84% to 91% [61, 109, 111], while the PLUS database had a lower recall rate of 23% [61]. The PubMed similar articles search and Scopus citation tracking had a low sensitivity of 25% each, with time-saving percentages of 24% and 58%, respectively [50]. However, the omission of studies from searching the PLUS database only did not significantly change the effect estimates in most reviews (ROR: 0.99; 95% CI, 0.87 to 1.14) [61]. Methods or tools for study selection Title and abstract selection We identified 42 studies evaluating 14 supportive software tools (AbstrackR, ASReview, ChatGPT, Colandr, Covidence, DistillerSR, EPPI-reviewer, Rayyan, RCT classifier, Research screener, RobotAnalyst, SRA-Helper for EndNote, SWIFT-active screener, SWIFT-review) [32, 35, 36, 38, 43-45, 47, 48, 51, 58, 59, 63, 73, 83, 94-96, 102, 105, 107, 108, 112, 117-120, 127] using advanced text mining and machine and active learning techniques, and five methods (crowdsourcing using different [automation] tools, dual computer monitors, single-reviewer screening, PICO-based title-only screening, limited screening [review of reviews]) [55, 79, 82, 84, 86-90, 100, 106, 108, 113, 114, 125] for improving the title and abstract screening efficiency. The tested datasets ranged from 1 to 60 SRs and 148 to 190,555 records. Tools for title and abstract selection Various tools (e.g., EPPI-Reviewer, Covidence, DistillerSR, and Rayyan) offer collaborative online platforms for SRs, enhancing efficiency by managing and distributing screening tasks, facilitating multiuser screening, and tracking records throughout the review process [128]. In a semiautomated tool, the tool provide suggestions or probabilities regarding the eligibility of a reference for inclusion in the review, but human judgment is still required to make the final decision [9, 12]. In contrast, in a fully automated system, the tool makes the final decision without human intervention based on predetermined criteria or algorithms. Some tools provide fully automated screening options (e.g., DistillerSR), semiautomated (e.g., RobotAnalyst), or both (e.g., AbstrackR, DistillerAI) using machine learning or natural language processing methods [9, 12] (see Table 1). Among the eleven semi- and fully -automated tools (AbstrackR [36, 38, 44, 45, 47, 48, 51, 59, 73, 107, 108, 118], ASReview [83, 95, 102, 120], ChatGPT [112], Colandr [38], DistillerSR [36, 43, 47, 58], EPPI-reviewer [36, 59, 118, 127], Rayyan [35, 36, 38, 73, 94, 96, 119, 127], RCT classifier [117], Research screener [32], RobotAnalyst [35, 36, 47, 105, 108], SRA-helper for EndNote [35], SWIFT-active screener [36, 63], SWIFT-review [73]). ASReview [83, 95, 102, 120] and Research Screener [32] demonstrated a robust performance, identifying 95% of the relevant studies while saving 37% to 92% of the workload. SWIFT-active screener[36, 63], RobotAnalyst [35, 36, 47, 105, 108], and Rayyan [35, 36, 94, 96, 119] also performed well, identifying 95% of the relevant studies with workload savings from 34% to 49%. EPPI-Reviewer identified all the relevant abstracts after screening 40% to 99% of the references across reviews [118]. DistillerAI showed substantial workload savings of 53% while identifying 95% of the relevant studies [58], with varying degrees of validity [43, 47, 58]. Colandr and SWIFT-Review exhibited sensitivity rates of 65% and 91%, respectively, with 97% workload savings and around 400 minutes of time saved [38, 73]. ChatGPT’s sensitivity was 100% [112] and RCT classifiers recall was 99% [117]; workload or time savings were not reported [112, 117]. AbstrackR showed a moderate performance with a potential workload savings from 4% up to 97% [38, 44, 45, 47, 48, 51, 107, 108, 118] while missing up to 44% of the relevant studies [38, 45, 48, 51, 108]. Covidence and SRA-Helper for EndNote did not report validity outcomes. Most of the supportive software tools were easy to use or learn, suitable for collaboration, and straightforward for inexperienced users (ASReview, AbstarckR, Covidence, SRA-Helper, Rayyan, RobotAnalyst) [35, 36, 44, 59, 94, 96, 120]. Other tools were more complex regarding their usability but were useful for large and complex projects (DistillerSR, EPPI-Reviewer) [36, 47, 59]. Poor interface quality (AbstrackR) [59], issues with help section/response time (RobotAnalyst, Covidence, EPPI) [35, 59], and overloaded side panel (Rayyan) [35] were weaknesses reported in the studies. Methods for title and abstract selection Among the four methods identified (dual computer monitors, single-reviewer screening, crowdsourcing using different [automation] tools, and limited screening [review of reviews, PICO-based title-only screening, title-first screening]) [55, 79, 82, 84, 86-90, 100, 113, 114, 125] for supporting title and abstract screening, crowdsourcing in combination with screening platforms or machine learning tools demonstrated the most promising performance in improving efficiency. Studies by Noel-Storr et al. [82, 87-90] found that the Cochrane Crowd plus Screen4Me/RCT classifier achieved a high sensitivity ranging from 84% to 100% in abstract screening and reduced screening time. Crowdsourcing via Amazon Mechanical Turk yielded correct inclusions of 95% to 99% with a substantial cost reduction of 82% [82]. However, the sensitivity was moderate when the screening was conducted manually by medical students or on web-based platforms (47% to 67%) [86]. Single-reviewer screening missed 0% to 19% of the relevant studies [55, 100, 113, 114] while saving 50% to 58% of the time and costs, respectively [113]. The findings indicate that single-reviewer screening by less-experienced reviewers could substantially alter the results, whereas experienced reviewers had a negligible impact [100]. Limited screening methods, such as reviews of reviews (also known as umbrella reviews), exhibited a moderate sensitivity (56%) and significantly reduced the number of citations needed to screen [108]. Title-first screening and PICO-based title screening demonstrated an accurate validity, with a recall of 100% [79, 107] and a reduction in screening effort ranging from 11% to 78% [107]. However, screening with dual computer monitors did not notably improve the time saved [125]. Full-text selection For full-text screening, we identified three methods: crowdsourcing [85], using dual computer monitors [125], and single-reviewer screening [113, 114]. Using crowdsourcing in combination with the CrowdScreenSR saved 68% of the workload [85]. With dual computer monitors, no significant difference of time taken for the full-text screening was reported [125]. Single-reviewer screening missed 7% to 12% of the relevant studies [114] while saving only 4% of the time and costs [113]. Methods or tools for data extraction We identified 11 studies evaluating five tools (ChatGPT [112], Data Abstraction Assistant (DAA) [36, 67, 74], Dextr [124], ExaCT [46, 70, 103] and Plot Digitizer [69]) and two methods (dual computer monitors [125] and single data extraction [31, 74]) to expedite the data extraction process. ExaCT [46, 70, 103], DAA [36, 67, 74], Dextr [124], and Plot Digitizer [69] achieved a time reduction of up to 60% [103, 124], with precision rates of 93% for ExaCT [70, 103] and 96% for Dextr and an error rate of 17% for DAA [67, 74]. Manual extraction by two reviewers and with the assistance of Plot Digitizer showed a similar agreement with the original data, showing a slightly higher agreement with the assistance of Plot Digitizer (Plot Digitizer: 73% and 75%, manual extraction: 66% and 69%) [69]. A total of 87% of manually extracted data elements matched with ExaCT, resulting in qualitatively altered meta-analysis results [103]. ChatGPT demonstrated consistent agreement with human researchers across various parameters (κ=0.79-1) extracted from studies, such as language, targeted disease, natural language processing model, sample size, and performance parameters, and moderate to fair agreement for clinical task (κ=0.58) and clinical implementation (κ=0.34) [112]. Usability was assessed only for DAA and Dextr, with both tools deemed very easy to use [67, 124], although DAA scored lower on feature scores, while Dextr was noted for its flexible interface [124]. Single data extraction and dual monitors reduced the time for extracting data by 24 to 65 minutes per article [31, 74, 125], with similar error rates between single and dual data extraction methods (single: 16% [74], 18% [31], dual: 15% [31, 74]) and comparable pooled estimates [31, 74]. Methods or tools for critical appraisal We identified nine studies reporting on one software tool (RobotReviewer) [26, 27, 36, 49, 62, 68, 75, 115] and one method (crowdsourcing via CrowdCARE) [101] for improving critical appraisal efficiency. Collectively, the study authors suggested that RobotReviewer can support but not replace RoB assessments by humans [26, 27, 49, 62, 68, 75] as performance varied per RoB domain [26, 49, 62, 115]. The authors reported similar completion times for RoB appraisal with and without RobotReviewer assistance [27]. Reviewers were equally as likely to accept RobotReviewer’s judgments as one another’s during consensus (83% for RobotReviewer, 81% for humans) [68], showing similar accuracy (RobotReviewer assisted RoB appraisal: 71%, RoB appraisal by two reviewers: 78%) [27, 75]. The reviewers generally described the tool as acceptable and useful [68], whereby collaboration with other users is not possible [36]. Combination of abbreviated methods/tools Five studies evaluated RR methods [25, 76, 78, 100, 116], and one study evaluated various tools [12] combining multiple review steps. While two case studies found no differences in findings between RR and SR approaches [78, 116], another author/paper/study found in two of three RRs that no conclusion could be drawn due to insufficient information [25]. Additionally, in a study including three RRs, RR methods affected almost one-third of the meta-analyses with less precise pooled estimates [100]. Marshall et al. (2019) included 2,512 SRs and reported a loss of all data in 4% to 45% of the meta-analyses and changes of 7% to 39% in the statistical significance due to RR methods [76]. Automation tools (SRA- Deduplicator, EndNote, Polyglot Search Translator, RobotReviewer, SRA-Helper) reduced the person-time spent on SR tasks (42 hours versus 12 hours) [12]. However, error rates, falsely excluded studies, and sensitivity varied immensely across studies [25, 116]. Discussion To the best of our knowledge, this is the first scoping review to map evaluated methods and tools for improving the efficiency of SR production across the various review steps. We describe which review steps methods and ready-to-use tools are available and have been evaluated. Additionally, we provide an overview of the contexts in which these methods and tools were evaluated, such as real-time workflow testing and the use of internal or external data. Across all the SR review steps, most studies evaluated study selection, followed by literature searching and data extraction. Around half of the studies evaluated tools and half of the studies methods. For study selection, most of the tools offered semiautomated-assisted screening by classifying or ranking references. The methods focused mainly on limiting the review team’s human resources, for example, through single-reviewer screening or by distributing the tasks to a crowd or students. Two scoping reviews, one on tools [10] and one on methods [15] to support the SR process, are in line with this result, as the authors also identified these tasks as the most frequently evaluated in the literature [10, 15]. As shown by our scoping review and others [10], a major focus on (semi) automation tools for study selection occurred in recent years. This is important, as a recent study on resource use found that study selection and data extraction are the most resource-intensive tasks besides administration and project management as well as critical appraisal [2]. For the following tasks, we could not identify a single study evaluating a tool or method: administration/project management, formulating the review question, writing the protocol, searching for existing reviews, full-text retrieval, synthesis/meta-analysis, certainty of evidence assessment, and report preparation. However, all of these tasks are also time-consuming, especially according to the study by Nussbaumer-Streit et al [2]. Project management requires the largest proportion of SR production time [2]. To our knowledge, tools supporting project management are already available, such as Covidence, DistillerSR, EPPI-Reviewer, or simple online platforms such as Google Forms, which can also support managing and coordinating projects. However, no evaluations of these support platforms were found. Similarly, while innovative software tools such as large language models (e.g., ChatGPT) show promise in supporting tasks such as report preparation, there is a lack of formal evaluation in this context as well. This is relevant for future research that aims to improve SR production since these tasks are extremely resource-intensive. Our scoping review identified several research gaps. There is a lack of studies evaluating the usability of tools and methods. No study evaluated the usability of any single method. Only for the study selection task did we identify multiple studies evaluating the usability of tools. However, an important factor in adopting tools and methods is their user-friendliness [14, 129] and their fit with standard SR workflows [9, 47]. Furthermore, if usability was considered, this was often evaluated in a nonformal or standardized way employing various scales, questions, and feedback mechanisms. To enable meaningful comparisons between different methods, there is a clear need for a formal analysis of user experience and usability. Therefore, authors and review teams would benefit from comparable usability studies on methods and tools that aim to improve the SR process’s efficiency. Few studies exist that evaluate the impact on results when using accelerated methods or tools. We identified 25 studies (13%) where the odds ratio changed by 0% to 63% depending on the method or tool [61, 91, 121, 130-133]. Marshall et al. (2019) stated that there has not been a large-scale evaluation of the effects of RR methods on the number of falsely excluded studies and the consequent changes in meta-analysis results [134]. Indeed, understanding the potential impact of different methods and tools on the results is fundamental, as emphasized by Wagner et al. (2016) [135]. Wagner et al. conducted an online survey of guideline developers and policymakers, aiming to discover how much incremental uncertainty about the correctness of an answer would be acceptable in RRs. They found that participants demanded very high levels of accuracy and would tolerate a median 10% risk of wrong answers [135]. Therefore, studies focusing on the impact on results and conclusions through RR methods and tools are warranted. The majority of studies retrospectively evaluated only a single tool or method using existing internal data, offering limited insights into real-world adoption. Prospective studies within a real-time workflow study (n=20) [12, 27, 32, 35, 46, 50, 67, 68, 77, 78, 88, 90, 94, 98, 105, 108, 112, 114, 125, 127], comparing several tools and methods (n=5) [12, 35, 50, 98, 108, 127] or from independent reviewer teams using their own dataset (n=7) [12, 35, 46, 94, 98, 112, 127], are scarce. However, such studies are crucial for providing valid comparative evidence on validity, workload savings, usability, impact on results, and real-world benefits. Particularly for automated title and abstract screening, where most tools function similarly, key information such as the stopping rule (indicating when screening can cease) is essential. Notably, there is limited research (n=8) [51, 63, 94, 95, 107, 108, 119, 127] exploring the combined effects of algorithmic re-ranking and stopping criterion determination in automated title and abstract screening. Furthermore, studies assessing the influence of automated tools (e.g., re-ranking algorithms) on human decision-making are lacking. Therefore, simulation and prospective real-time studies evaluating the workflow between manual procedures and tools with stopping criterion are warranted. The balance between the time saved and the potential reduction in quality and comprehensiveness is influenced by various factors, including the decision-making urgency, resource availability, and the decision-makers’ specific needs. However, accelerated approaches are not universally appropriate. In situations where the thoroughness and rigor of the evidence are paramount—such as in developing clinical guidelines, conducting health technology assessments, or addressing areas with significant scientific uncertainty—the risk of missing critical evidence or drawing inaccurate conclusions outweighs the benefits of speed. Given the heterogeneity in study designs and contexts, there is a pressing need for standardized frameworks for the evaluation and reporting of tools and methods for SR production. Furthermore, researchers should be aware of the importance of testing methods and tools with their own datasets and contextual factors. Pretraining both the tools and the crowd before implementation is essential for optimizing efficiency and ensuring reliable outcomes. However, we believe our findings are generalizable to other types of evidence syntheses, such as scoping reviews, reviews of reviews, or RRs. Additionally, we highlight that a part of the Rapid Review Methods Series [128] offers guidance on the use of supportive tools, aiming to assist researchers in effectively navigating the complexities of SR and RR production. Our scoping review has several limitations. First, our inclusion criteria focused solely on studies mentioning efficiency improvements. While this criterion aimed to strike a balance between screening workload and sensitivity, it may have inadvertently excluded relevant studies that did not explicitly highlight efficiency gains. We might have missed relevant studies. Second, the heterogeneity among the included studies poses challenges in generalizing the findings to other review teams and contexts. The narrow focus of many studies, along with their publication primarily in English and focus on specific study types, further limits the generalizability of our findings. Moreover, the limited proportion of studies (29%, 30/103) comparing different tools and methods on the same dataset within the same study accentuates the need for caution while interpreting the findings. However, we think the validity and usability outcomes reported in this scoping review provide a good orientation. Conclusion Based on the identified evidence, various methods and tools for literature searching and title and abstract screening are available with the aim of improving efficiency. However, only few studies have addressed the influence of these methods and tools in real-world workflows. Fewer studies evaluated methods or tools supporting the other tasks of SR production. Moreover, the reporting of the outcomes of existing evaluations varies considerably, revealing significant research gaps, especially in assessing usability and impact on results. Future research should prioritize addressing these gaps, evaluating real-world adoption, and establishing standardized frameworks for the evaluation of methods and tools to enhance the overall effectiveness of SR development processes. Declarations Ethics approval and consent to participate: not applicable. Consent for publication: not applicable. Availability of data and materials: No/Not applicable (this manuscript does not report data generation or analysis). Competing interests: The authors declare that they have no competing interests. Funding: This work was supported by funds from the EU funded COST Action EVBRES (CA17117), two COST STSM fundings (LA, MM), and by a scholarship for the first author (LA) of the Gesellschaft für Forschungsförderung Niederösterreich m.b.H. Authors’ contributions: LA, BNS, RS, and LH developed the study concept. LA wrote the protocol. RS conducted the literature searches and similar article searches. LA, MMM, IS, BNS, MMK, MEM, EB, ME, RS, PNL, GP, NR, KG, LK, MS, AMP, and PM screened the references and extracted the data. LA and MMM conducted grey literature searches, reference list checking, and hand searches. RS provided methodological input throughout the study. LA and MMM created the first draft of the manuscript, which all other authors critically revised. The authors read and approved the final version of the submitted manuscript. Acknowledgments: We would like to thank Hans Lund and Jos Kleijnen for their input on the manuscript. References Tsafnat G, Glasziou P, Choong MK, Dunn A, Galgani F, Coiera E. Systematic review automation technologies. Syst Rev. 2014;3:74. Nussbaumer-Streit B, Ellen M, Klerings I, Sfetcu R, Riva N, Mahmić-Kaknjo M, Poulentzas G, Martinez P, Baladia E, Ziganshina LE et al. Resource use during systematic review production varies widely: a scoping review. J Clin Epidemiol 2021. Oliver S, Dickson K, Bangpan M. Systematic reviews: making them policy relevant. A briefing for policy makers and systematic reviewers. EPPI-Centre, Social Science Research Unit, UCL Institute of Education, University College London, London 2015. Borah R, Brown AW, Capers PL, Kaiser KA. Analysis of the time and workers needed to conduct systematic reviews of medical interventions using data from the PROSPERO registry. BMJ Open. 2017;7(2):e012545–012545. Donnelly CA, Boyd I, Campbell P, Craig C, Vallance P, Walport M, Whitty CJM, Woods E, Wormald C. Four principles to make evidence synthesis more useful for policy. Nature. 2018;558(7710):361–4. Clayton GL, Smith IL, Higgins JPT, Mihaylova B, Thorpe B, Cicero R, Lokuge K, Forman JR, Tierney JF, White IR, et al. The INVEST project: investigating the use of evidence synthesis in the design and analysis of clinical trials. Trials. 2017;18(1):219. Peters MDJ, Marnie C, Tricco AC, Pollock D, Munn Z, Alexander L, McInerney P, Godfrey CM, Khalil H. Updated methodological guidance for the conduct of scoping reviews. JBI Evid Synthesis 2020, 18(10). Beller E, Clark J, Tsafnat G, Adams C, Diehl H, Lund H, Ouzzani M, Thayer K, Thomas J, Turner T. Making progress with the automation of systematic reviews: principles of the International Collaboration for the Automation of Systematic Reviews (ICASR). Syst reviews. 2018;7(1):77. O’Connor AM, Tsafnat G, Thomas J, Glasziou P, Gilbert SB, Hutton B. A question of trust: can we build an evidence base to gain trust in systematic review automation technologies? Syst reviews. 2019;8(1):1–8. Khalil H, Ameen D, Zarnegar A. Tools to support the automation of systematic reviews: a scoping review. J Clin Epidemiol. 2022;144:22–42. Khurana D, Koli A, Khatter K, Singh S. Natural language processing: state of the art, current trends and challenges. Multimedia Tools Appl. 2023;82(3):3713–44. Clark J, McFarlane C, Cleo G, Ishikawa Ramos C, Marshall S. The Impact of Systematic Review Automation Tools on Methodological Quality and Time Taken to Complete Systematic Review Tasks: Case Study. JMIR Med Educ. 2021;7(2):e24418. Ergonomics of human-system interaction - Part 11. Usability: Definitions and concepts, (ISO 9241-11: 2018). ISO 9241-11:2018 - Ergonomics of human-system interaction — Part 11: Usability: Definitions and concepts. Scott AM, Forbes C, Clark J, Carter M, Glasziou P, Munn Z. Systematic review automation tools improve efficiency but lack of knowledge impedes their adoption: a survey. J Clin Epidemiol. 2021;138:80–94. Hamel C, Michaud A, Thuku M, Affengruber L, Skidmore B, Nussbaumer-Streit B, Stevens A, Garritty C. Few evaluative studies exist examining rapid review methodology across stages of conduct: a systematic scoping review. J Clin Epidemiol. 2020;126:131–40. Arksey H, O'Malley L. Scoping studies: towards a methodological framework. Int J Soc Res Methodol. 2005;8(1):19–32. Levac D, Colquhoun H, O'Brien KK. Scoping studies: advancing the methodology. Implement Sci. 2010;5(1):69. Colquhoun HL, Levac D, O'Brien KK, Straus S, Tricco AC, Perrier L, Kastner M, Moher D. Scoping reviews: time for clarity in definition, methods, and reporting. J Clin Epidemiol. 2014;67(12):1291–4. Tricco AC, Lillie E, Zarin W, O'Brien KK, Colquhoun H, Levac D, Moher D, Peters MD, Horsley T, Weeks L. PRISMA extension for scoping reviews (PRISMA-ScR): checklist and explanation. Ann Intern Med. 2018;169(7):467–73. JBI Reviewer's Manual. Chapter 11.2.5. Search Strategy. McGowan J, Sampson M, Salzwedel DM, Cogo E, Foerster V, Lefebvre C. PRESS peer review of electronic search strategies: 2015 guideline statement. J Clin Epidemiol. 2016;75:40–6. Best L, Stevens A, Colin-Jones D. Rapid and responsive health technology assessment: the development and evaluation process in the South and West region of England. J Clin Eff 1997. Jonnalagadda S, Petitti D. A new iterative method to reduce workload in systematic review process. Int J Comput Biol Drug Des. 2013;6(1–2):5–17. Higgins JPTTJ, Chandler J, Cumpston M, Li T, Page MJ, Welch VA, editors. Cochrane Handbook for Systematic Reviews of Interventions version 6.2 (updated February 2021); 2021. Affengruber L, Wagner G, Waffenschmidt S, Lhachimi SK, Nussbaumer-Streit B, Thaler K, Griebler U, Klerings I, Gartlehner G. Combining abbreviated literature searches with single-reviewer screening: three case studies of rapid reviews. Syst Rev. 2020;9(1):162. Armijo-Olivo S, Craig R, Campbell S. Comparing machine and human reviewers to evaluate the risk of bias in randomized controlled trials. Res. 2020;11(3):484–93. Arno A, Thomas J, Wallace B, Marshall IJ, McKenzie JE, Elliott JH. Accuracy and Efficiency of Machine Learning-Assisted Risk-of-Bias Assessments in Real-World Systematic Reviews: A Noninferiority Randomized Controlled Trial. Ann Intern Med. 2022;175(7):1001–9. Belter CW. Citation analysis as a literature search method for systematic reviews. J Assoc Soc Inf Sci Technol. 2016;67(11):2766–77. Beyer FR, Wright K. Can we prioritise which databases to search? A case study using a systematic review of frozen shoulder management. Health Info Libr J. 2013;30(1):49–58. Borissov N, Haas Q, Minder B, Kopp-Heim D, von Gernler M, Janka H, Teodoro D, Amini P. Reducing systematic review burden using Deduklick: a novel, automated, reliable, and explainable deduplication algorithm to foster medical research. Syst. 2022;11(1):172. Buscemi N, Hartling L, Vandermeer B, Tjosvold L, Klassen TP. Single data extraction generated more errors than double data extraction in systematic reviews. J Clin Epidemiol. 2006;59(7):697–703. Chai KEK, Lines RLJ, Gucciardi DF, Ng L. Research Screener: a machine learning tool to semi-automate abstract screening for systematic reviews. Syst Reviews 2021, 10(1). Chapman AL, Morgan LC, Gartlehner G. Semi-automating the manual literature search for systematic reviews increases efficiency. Health Info Libr J. 2010;27(1):22–7. Clark JM, Sanders S, Carter M, Honeyman D, Cleo G, Auld Y, Booth D, Condron P, Dalais C, Bateup S, et al. Improving the translation of search strategies using the Polyglot Search Translator: a randomized controlled trial. J Med Libr Assoc. 2020;108(2):195–207. Cleo G, Scott AM, Islam F, Julien B, Beller E. Usability and acceptability of four systematic review automation software packages: a mixed method design. Syst. 2019;8(1):145. Cowie K, Rahmatullah A, Hardy N, Holub K, Kallmes K. Web-Based Software Tools for Systematic Literature Review in Medicine: Systematic Search and Feature Analysis. JMIR Med Inf. 2022;10(5):e33219. Dechartres A, Atal I, Riveros C, Meerpohl J, Ravaud P. Association between publication characteristics and treatment effect estimates: a meta-epidemiologic study. Ann Intern Med. 2018;169(6):385–93. dos Reis AHS, de Oliveira ALM, Fritsch C, Zouch J, Ferreira P, Polese JC. Usefulness of machine learning softwares to screen titles of systematic reviews: a methodological study. Syst Reviews. 2023;12(1):68. Egger M, Juni P, Bartlett C, Holenstein F, Sterne J. How important are comprehensive literature searches and the assessment of trial quality in systematic reviews? Empirical study. Health Technol Assess. 2003;7(1):1–76. Ewald H, Klerings I, Wagner G, Heise TL, Dobrescu AI, Armijo-Olivo S, Stratil JM, Lhachimi SK, Mittermayr T, Gartlehner G, et al. Abbreviated and comprehensive literature searches led to identical or very similar effect estimates: a meta-epidemiological study. J Clin Epidemiol. 2020;128:1–12. Ewald H, Klerings I, Wagner G, Heise TL, Stratil JM, Lhachimi SK, Hemkens LG, Gartlehner G, Armijo-Olivo S, Nussbaumer-Streit B. Searching two or more databases decreased the risk of missing relevant studies: a metaresearch study. J Clin Epidemiol. 2022;149:154–64. Furuya-Kanamori L, Lin L, Kostoulas P, Clark J, Xu C. Limits in the search date for rapid reviews of diagnostic test accuracy studies. Res. 2022;13:13. Gartlehner G, Wagner G, Lux L, Affengruber L, Dobrescu A, Kaminski-Hartenthaler A, Viswanathan M. Assessing the accuracy of machine-assisted abstract screening with DistillerAI: a user study. Syst. 2019;8(1):277. Gates A, Gates M, DaRosa D, Elliott SA, Pillay J, Rahman S, Vandermeer B, Hartling L. Decoding semi-automated title-abstract screening: findings from a convenience sample of reviews. Syst Reviews 2020, 9(1). Gates A, Gates M, Sebastianski M, Guitard S, Elliott SA, Hartling L. The semi-automation of title and abstract screening: a retrospective exploration of ways to leverage Abstrackr's relevance predictions in systematic and rapid reviews. BMC Med Res Methodol. 2020;20(1):139. Gates A, Gates M, Sim S, Elliott SA, Pillay J, Hartling L. Creating efficiencies in the extraction of data from randomized trials: a prospective evaluation of a machine learning and text mining tool. BMC Med Res Methodol. 2021;21(1):169. Gates A, Guitard S, Pillay J, Elliott SA, Dyson MP, Newton AS, Hartling L. Performance and usability of machine learning for screening in systematic reviews: a comparative evaluation of three tools. Syst. 2019;8(1):278. Gates A, Johnson C, Hartling L. Technology-assisted title and abstract screening for systematic reviews: a retrospective evaluation of the Abstrackr machine learning tool. Syst. 2018;7(1):45. Gates A, Vandermeer B, Hartling L. Technology-assisted risk of bias assessment in systematic reviews: a prospective cross-sectional evaluation of the RobotReviewer machine learning tool. J Clin Epidemiol. 2018;96:54–62. Gates M, Elliott SA, Gates A, Sebastianski M, Pillay J, Bialy L, Hartling L. LOCATE: a prospective evaluation of the value of Leveraging Ongoing Citation Acquisition Techniques for living Evidence syntheses. Syst. 2021;10(1):116. Giummarra MJ, Lau G, Gabbe BJ. Evaluation of text mining to reduce screening workload for injury-focused systematic reviews. Inj Prev. 2020;26(1):55–60. Glanville JM, Lefebvre C, Miles JN, Camosso-Stefinovic J. How to identify randomized controlled trials in MEDLINE: ten years on. J Med Libr Assoc. 2006;94(2):130–6. Goossen K, Tenckhoff S, Probst P, Grummich K, Mihaljevic AL, Büchler MW, Diener MK. Optimal literature search for systematic reviews in surgery. Langenbecks Arch Surg. 2018;403(1):119–29. Guimaraes NS, Ferreira AJF, Ribeiro Silva RC, de Paula AA, Lisboa CS, Magno L, Ichiara MY, Barreto ML. Deduplicating records in systematic reviews: there are free, accurate automated ways to do so. J Clin Epidemiol. 2022;152:110–5. Gartlehner G, Affengruber L, Titscher V, Noel-Storr A, Dooley G, Ballarini N, König F. Single-reviewer abstract screening missed 13 percent of relevant studies: a crowd-based, randomized controlled trial. J Clin Epidemiol. 2020;121:20–8. Haas Q, Alvarez DV, Borissov N, Ferdowsi S, von Meyenn L, Trelle S, Teodoro D, Amini P. Utilizing Artificial Intelligence to Manage COVID-19 Scientific Evidence Torrent with Risklick AI: A Critical Tool for Pharmacology and Therapy Development. Pharmacology. 2021;106(5–6):244–53. Hair K, Bahor Z, Macleod M, Liao J, Sena ES. The Automated Systematic Search Deduplicator (ASySD): a rapid, open-source, interoperable tool to remove duplicate citations in biomedical systematic reviews. BMC Biol. 2023;21(1):189. Hamel C, Kelly SE, Thavorn K, Rice DB, Wells GA, Hutton B. An evaluation of DistillerSR's machine learning-based prioritization tool for title/abstract screening - impact on reviewer-relevant outcomes. BMC Med Res Methodol. 2020;20(1):256. Harrison H, Griffin SJ, Kuhn I, Usher-Smith JA. Software tools to support title and abstract screening for systematic reviews in healthcare: An evaluation. BMC Med Res Methodol. 2020;20(1):7. Hartling L, Featherstone R, Nuspl M, Shave K, Dryden DM, Vandermeer B. Grey literature in systematic reviews: a cross-sectional study of the contribution of non-English reports, unpublished studies and dissertations to the results of meta-analyses in child-relevant reviews. BMC Med Res Methodol 2017, 17(1):64. Hemens BJ, Haynes RB. McMaster Premium LiteratUre Service (PLUS) performed well for identifying new studies for updated Cochrane reviews. J Clin Epidemiol. 2012;65(1):62–e7261. Hirt J, Meichlinger J, Schumacher P, Mueller G. Agreement in Risk of Bias Assessment Between RobotReviewer and Human Reviewers: An Evaluation Study on Randomised Controlled Trials in Nursing-Related Cochrane Reviews. J Nurs Scholarsh. 2021;53(2):246–54. Howard BE, Phillips J, Tandon A, Maharana A, Elmore R, Mav D, Sedykh A, Thayer K, Merrick BA, Walker V et al. SWIFT-Active Screener: Accelerated document screening through active learning and integrated recall estimation. Environ Int 2020, 138 (no pagination). Hugues A, Di Marco J, Bonan I, Rode G, Cucherat M, Gueyffier F. Publication language and the estimate of treatment effects of physical therapy on balance and postural control after stroke in meta-analyses of randomised controlled trials. PLoS ONE. 2020;15(3):e0229822. Janssens A, Gwinn M, Brockman JE, Powell K, Goodman M. Novel citation-based search method for scientific literature: a validation study. BMC Med Res Methodol. 2020;20(1):25. Janssens AC, Gwinn M. Novel citation-based search method for scientific literature: application to meta-analyses. BMC Med Res Methodol. 2015;15:84. Jap J, Saldanha IJ, Smith BT, Lau J, Schmid CH, Li T, Investigators obotDAA. Features and functioning of Data Abstraction Assistant, a software application for data abstraction during systematic reviews. Res Synthesis Methods. 2019;10(1):2–14. Jardim PSJ, Rose CJ, Ames HM, Echavez JFM, Van de Velde S, Muller AE. Automating risk of bias assessment in systematic reviews: a real-time mixed methods comparison of human researchers to a machine learning system. BMC Med Res Methodol. 2022;22(1):167. Jelicic Kadic A, Vucic K, Dosenovic S, Sapunar D, Puljak L. Extracting data from figures with software was faster, with higher interrater reliability than manual extraction. J Clin Epidemiol. 2016;74:119–23. Kiritchenko S, de Bruijn B, Carini S, Martin J, Sim I. ExaCT: automatic extraction of clinical trial characteristics from journal publications. BMC Med Inf Decis Mak. 2010;10:56. Kwon Y, Lemieux M, McTavish J, Wathen N. Identifying and removing duplicate records from systematic review searches. J Med Libr Association: JMLA. 2015;103(4):184–8. Lee E, Dobbins M, Decorby K, McRae L, Tirilis D, Husson H. An optimal search filter for retrieving systematic reviews and meta-analyses. BMC Med Res Methodol. 2012;12:51. Li J, Kabouji J, Bouhadoun S, Tanveer S, Filion KB, Gore G, Josephson CB, Kwon CS, Jette N, Bauer PR, et al. Sensitivity and specificity of alternative screening methods for systematic reviews using text mining tools. J Clin Epidemiol. 2023;162:72–80. Li T, Saldanha IJ, Jap J, Smith BT, Canner J, Hutfless SM, Branch V, Carini S, Chan W, de Bruijn B, et al. A randomized trial provided new evidence on the accuracy and efficiency of traditional vs. electronically annotated abstraction approaches in systematic reviews. J Clin Epidemiol. 2019;115:77–89. Marshall IJ, Kuiper J, Wallace BC. RobotReviewer: evaluation of a system for automatically assessing bias in clinical trials. J Am Med Inf Assoc. 2016;23(1):193–201. Marshall IJ, Marshall R, Wallace BC, Brassey J, Thomas J. Rapid reviews may produce different results to systematic reviews: a meta-epidemiological study. J Clin Epidemiol. 2019;109:30–41. Marshall IJ, Trikalinos TA, Soboczenski F, Yun HS, Kell G, Marshall R, Wallace BC. In a pilot study, automated real-time systematic review updates were feasible, accurate, and work-saving. J Clin Epidemiol. 2023;153:26–33. Martyn-St James M, Cooper K, Kaltenthaler E. Methods for a rapid systematic review and metaanalysis in evaluating selective serotonin reuptake inhibitors for premature ejaculation. Evid Policy: J Res Debate Pract. 2017;13(3):517–38. Mateen FJ, Oh J, Tergas AI, Bhayani NH, Kamdar BB. Titles versus titles and abstracts for initial screening of articles for systematic reviews. Clin Epidemiol. 2013;5:89–95. McKeown S, Mir ZM. Considerations for conducting systematic reviews: evaluating the performance of different methods for de-duplicating references. Syst. 2021;10(1):38. Moher D, Klassen TP, Schulz KF, Berlin JA, Jadad AR, Liberati A. What contributions do languages other than English make on the results of meta-analyses? J Clin Epidemiol. 2000;53(9):964–72. Mortensen ML, Adam GP, Trikalinos TA, Kraska T, Wallace BC. An exploration of crowdsourcing citation screening for systematic reviews. Res. 2017;8(3):366–86. Muthu S. The efficiency of machine learning-assisted platform for article screening in systematic reviews in orthopaedics. Int Orthop. 2023;47(2):551–6. Nama N, Iliriani K, Xia MY, Chen BP, Zhou LL, Pojsupap S, Kappel C, O'Hearn K, Sampson M, Menon K, et al. A pilot validation study of crowdsourcing systematic reviews: update of a searchable database of pediatric clinical trials of high-dose vitamin D. Transl. 2017;6(1):18–26. Nama N, Sampson M, Barrowman N, Sandarage R, Menon K, Macartney G, Murto K, Vaccani JP, Katz S, Zemek R, et al. Crowdsourcing the Citation Screening Process for Systematic Reviews: Validation Study. J Med Internet Res. 2019;21(4):e12953. Ng L, Pitt V, Huckvale K, Clavisi O, Turner T, Gruen R, Elliott JH. Title and Abstract Screening and Evaluation in Systematic Reviews (TASER): a pilot randomised controlled trial of title and abstract screening by medical students. Syst Rev. 2014;3:121. Noel-Storr A, Dooley G, Affengruber L, Gartlehner G. Citation screening using crowdsourcing and machine learning produced accurate results: evaluation of Cochrane's modified Screen4Me service. J Clin Epidemiol. 2020;29:29. Noel-Storr A, Dooley G, Elliott J, Steele E, Shemilt I, Mavergames C, Wisniewski S, McDonald S, Murano M, Glanville J, et al. An evaluation of Cochrane Crowd found that crowdsourcing produced accurate results in identifying randomized trials. J Clin Epidemiol. 2021;133:130–9. Noel-Storr A, Gartlehner G, Dooley G, Persad E, Nussbaumer-Streit B. Crowdsourcing the identification of studies for COVID-19-related Cochrane Rapid Reviews. Res Synth Methods. 2022;13(5):585–94. Noel-Storr AH, Redmond P, Lamé G, Liberati E, Kelly S, Miller L, Dooley G, Paterson A, Burt J. Crowdsourcing citation-screening in a mixed-studies systematic review: a feasibility study. BMC Med Res Methodol. 2021;21(1):88. Nussbaumer-Streit B, Klerings I, Dobrescu AI, Persad E, Stevens A, Garritty C, Kamel C, Affengruber L, King VJ, Gartlehner G. Excluding non-English publications from evidence-syntheses did not change conclusions: a meta-epidemiological study. J Clin Epidemiol. 2020;118:42–54. Nussbaumer-Streit B, Klerings I, Wagner G, Heise TL, Dobrescu AI, Armijo-Olivo S, Stratil JM, Persad E, Lhachimi SK, Van Noord MG, et al. Abbreviated literature searches were viable alternatives to comprehensive searches: a meta-epidemiological study. J Clin Epidemiol. 2018;102:1–11. O'Keefe H, Rankin J, Wallace SA, Beyer F. Investigation of text-mining methodologies to aid the construction of search strategies in systematic reviews of diagnostic test accuracy-a case study. Res. 2023;14(1):79–98. Olofsson H, Brolund A, Hellberg C, Silverstein R, Stenstrom K, Osterberg M, Dagerhamn J. Can abstract screening workload be reduced using text mining? User experiences of the tool Rayyan. Res. 2017;8(3):275–80. Oude Wolcherink MJ, Pouwels X, van Dijk SHB, Doggen CJM, Koffijberg H. Can artificial intelligence separate the wheat from the chaff in systematic reviews of health economic articles? Expert Rev Pharmacoecon Outcomes Res. 2023;23(9):1049–56. Ouzzani M, Hammady H, Fedorowicz Z, Elmagarmid A. Rayyan-a web and mobile app for systematic reviews. Syst. 2016;5(1):210. Pallath A, Zhang Q. Paperfetcher: A tool to automate handsearching and citation searching for systematic reviews. Res. 2022;19:19. Paynter RA, Featherstone R, Stoeger E, Fiordalisi C, Voisin C, Adam GP. A prospective comparison of evidence synthesis search strategies developed with and without text-mining tools. J Clin Epidemiol. 2021;139:350–60. Pham B, Klassen TP, Lawson ML, Moher D. Language of publication restrictions in systematic reviews gave different results depending on whether the intervention was conventional or complementary. J Clin Epidemiol. 2005;58(8):769–76. Pham MT, Waddell L, Rajic A, Sargeant JM, Papadopoulos A, McEwen SA. Implications of applying methodological shortcuts to expedite systematic reviews: three case studies using systematic reviews from agri-food public health. Res. 2016;7(4):433–46. Pianta MJ, Makrai E, Verspoor KM, Cohn TA, Downie LE. Crowdsourcing critical appraisal of research evidence (CrowdCARE) was found to be a valid approach to assessing clinical research quality. J Clin Epidemiol. 2018;104:8–14. Pijls BG. Machine Learning assisted systematic reviewing in orthopaedics. J Orthop. 2024;48:103–6. Pradhan R, Hoaglin DC, Cornell M, Liu W, Wang V, Yu H. Automatic extraction of quantitative data from ClinicalTrials.gov to conduct meta-analyses. J Clin Epidemiol. 2019;105:92–100. Preston L, Carroll C, Gardois P, Paisley S, Kaltenthaler E. Improving search efficiency for systematic reviews of diagnostic test accuracy: An exploratory study to assess the viability of limiting to MEDLINE, EMBASE and reference checking. Syst Reviews 2015, 4(1). Przybyla P, Brockmeier AJ, Kontonatsios G, Le Pogam MA, McNaught J, von Elm E, Nolan K, Ananiadou S. Prioritising references for systematic reviews with RobotAnalyst: A user study. Res. 2018;9(3):470–88. Rathbone J, Albarqouni L, Bakhit M, Beller E, Byambasuren O, Hoffmann T, Scott AM, Glasziou P. Expediting citation screening using PICo-based title-only screening for identifying studies in scoping searches and rapid reviews. Syst. 2017;6(1):233. Rathbone J, Hoffmann T, Glasziou P. Faster title and abstract screening? Evaluating Abstrackr, a semi-automated online screening program for systematic reviewers. Syst. 2015;4:80. Reddy SM, Patel S, Weyrich M, Fenton J, Viswanathan M. Comparison of a traditional systematic review approach with review-of-reviews and semi-automation as strategies to update the evidence. Syst. 2020;9(1):243. Rice M, Ali MU, Fitzpatrick-Lewis D, Kenny M, Raina P, Sherifali D. Testing the effectiveness of simplified search strategies for updating systematic reviews. J Clin Epidemiol. 2017;88:148–53. Royle P, Waugh N. A simplified search strategy for identifying randomised controlled trials for systematic reviews of health care interventions: a comparison with more exhaustive strategies. BMC Med Res Methodol. 2005;5:23. Sampson M, de Bruijn B, Urquhart C, Shojania K. Complementary approaches to searching MEDLINE may be sufficient for updating systematic reviews. J Clin Epidemiol. 2016;78:108–15. Schopow N, Osterhoff G, Baur D. Applications of the Natural Language Processing Tool ChatGPT in Clinical Practice: Comparative Study and Augmented Systematic Review. JMIR Med Inf. 2023;11:e48933. Shemilt I, Khan N, Park S, Thomas J. Use of cost-effectiveness analysis to compare the efficiency of study identification methods in systematic reviews. Syst. 2016;5(1):140. Stoll CRT, Izadi S, Fowler S, Green P, Suls J, Colditz GA. The value of a second reviewer for study selection in systematic reviews. Res. 2019;10(4):539–45. Šuster S, Baldwin T, Verspoor K. Analysis of predictive performance and reliability of classifiers for quality assessment of medical evidence revealed important variation by medical area. J Clin Epidemiol. 2023;159:58–69. Taylor-Phillips S, Geppert J, Stinton C, Freeman K, Johnson S, Fraser H, Sutcliffe P, Clarke A. Comparison of a full systematic review versus rapid review approaches to assess a newborn screening test for tyrosinemia type 1. Res. 2017;8(4):475–84. Thomas J, McDonald S, Noel-Storr A, Shemilt I, Elliott J, Mavergames C, Marshall IJ. Machine learning reduced workload with minimal risk of missing studies: development and evaluation of a randomized controlled trial classifier for Cochrane Reviews. J Clin Epidemiol. 2021;133:140–51. Tsou AY, Treadwell JR, Erinoff E, Schoelles K. Machine learning for screening prioritization in systematic reviews: comparative performance of Abstrackr and EPPI-Reviewer. Syst. 2020;9(1):73. Valizadeh A, Moassefi M, Nakhostin-Ansari A, Hosseini Asl SH, Saghab Torbati M, Aghajani R, Maleki Ghorbani Z, Faghani S. Abstract screening using the automated tool Rayyan: results of effectiveness in three diagnostic test accuracy systematic reviews. BMC Med Res Methodol. 2022;22(1):160. van de Schoot R, de Bruin J, Schram R, Zahedi P, de Boer J, Weijdema F, Kramer B, Huijts M, Hoogerwerf M, Ferdinands G, et al. An open source machine learning framework for efficient and transparent systematic reviews. Nat Mach Intell. 2021;3(2):125–33. Van Enst WA, Scholten RJPM, Whiting P, Zwinderman AH, Hooft L. Meta-epidemiologic analysis indicates that MEDLINE searches are sufficient for diagnostic test accuracy systematic reviews. J Clin Epidemiol. 2014;67(11):1192–9. Waffenschmidt S, Guddat C. Searches for randomized controlled trials of drugs in MEDLINE and EMBASE using only generic drug names compared with searches applied in current practice in systematic reviews. Res. 2015;6(2):188–94. Waffenschmidt S, Knelangen M, Sieben W, Bühn S, Pieper D. Single screening versus conventional double screening for study selection in systematic reviews: a methodological systematic review. BMC Med Res Methodol. 2019;19(1):132. Walker VR, Rooney AA. CEC02-02 Automated and Semi-Automated Approaches for Literature Searching, Screening, and Data Extraction for Systematic Reviews in Environmental Health. Toxicol Lett. 2021;350(Supplement):S4. Wang Z, Asi N, Elraiyah TA, Abu Dabrh AM, Undavalli C, Glasziou P, Montori V, Murad MH. Dual computer monitors to increase efficiency of conducting systematic reviews. J Clin Epidemiol. 2014;67(12):1353–7. Xu C, Ju K, Lin L, Jia P, Kwong JSW, Syed A, Furuya-Kanamori L. Rapid evidence synthesis approach for limits on the search date: How rapid could it be? Res. 2022;13(1):68–76. Waffenschmidt S, Sieben W, Jakubeit T, Knelangen M, Overesch I, Bühn S, Pieper D, Skoetz N, Hausner E. Increasing the efficiency of study selection for systematic reviews using prioritization tools and a single-screening approach. Syst Reviews. 2023;12(1):161. Affengruber L, Nussbaumer-Streit B, Hamel C, Maten MVd, Thomas J, Mavergames C, Spijker R, Gartlehner G. Rapid review methods series: Guidance on the use of supportive software. BMJ Evidence-Based Med 2024:bmjebm–2023. van Altena AJ, Spijker R, Olabarriaga SD. Usage of automation tools in systematic reviews. Res Synthesis Methods. 2019;10(1):72–82. Ewald H, Klerings I, Wagner G, Heise TL, Dobrescu AI, Armijo-Olivo S, Stratil JM, Lhachimi SK, Mittermayr T, Gartlehner G, et al. Abbreviated and comprehensive literature searches led to identical or very similar effect estimates: a meta-epidemiological study. J Clin Epidemiol. 2020;128:1–12. Hartling L, Featherstone R, Nuspl M, Shave K, Dryden DM, Vandermeer B. Grey literature in systematic reviews: a cross-sectional study of the contribution of non-English reports, unpublished studies and dissertations to the results of meta-analyses in child-relevant reviews. BMC Medical Research Methodology 2017, 17(1):64. Egger M, J�ni P, Bartlett C, Holenstein F, Sterne J. How important are comprehensive literature searches and the assessment of trial quality in systematic reviews? Empir Study. 2003;7(1):1–76. Pham B, Klassen TP, Lawson ML, Moher D. Language of publication restrictions in systematic reviews gave different results depending on whether the intervention was conventional or complementary. J Clin Epidemiol. 2005;58(8):769–e776762. Marshall IJ, Wallace BC. Toward systematic review automation: a practical guide to using machine learning tools in research synthesis. Syst Reviews. 2019;8(1):163. Wagner G, Nussbaumer-Streit B, Greimel J, Ciapponi A, Gartlehner G. Trading certainty for speed - how much uncertainty are decisionmakers and guideline developers willing to accept when using rapid reviews: an international survey. BMC Med Res Methodol. 2017;17:121. Additional Declarations No competing interests reported. Supplementary Files Appendix1.docx Appendix2.docx Appendix3.docx Cite Share Download PDF Status: Published Journal Publication published 18 Sep, 2024 Read the published version in BMC Medical Research Methodology → Version 1 posted Editorial decision: Revision requested 22 Jul, 2024 Reviews received at journal 20 Jul, 2024 Reviewers agreed at journal 11 Jul, 2024 Reviewers agreed at journal 11 Jul, 2024 Reviews received at journal 11 Jul, 2024 Reviewers agreed at journal 11 Jul, 2024 Reviewers invited by journal 05 Jul, 2024 Editor invited by journal 21 Jun, 2024 Editor assigned by journal 21 Jun, 2024 Submission checks completed at journal 21 Jun, 2024 First submitted to journal 17 Jun, 2024 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-4595777","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":325892505,"identity":"e4a002f3-bdc8-4fd3-80e3-1a95595c64d0","order_by":0,"name":"Lisa Affengruber","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAABRElEQVRIie2RP0vDUBDALzxolsS3Xoj6GU4CwYLUr9ISaJeAgyCCUgtCulTngvodOmndCg/SJegmGSuBTgoBQSL+qY/EIibBWTC/5eDu/bh7dwAVFX+QFZaGdZ4GpQfAoflVY+VKLctbRm+pGL2cgnkF8gpNmj+fFBRVvY/hkJCb/bsnfby1Y00df/biQetKBSVOoNEtDKZZCD6hcR7snulBu34ZzDsbJ1K5PgZmDMApdGGa7Fx761LoOkz3BNmha6PuLVojAWDKHxUVNUrgg3A7UxZkDV3beJddpMJeAY6KCtioeISEHSGVCRG6tqlnSk12ESWD2ZutU0IMXaZcBA5hMG+bqzdgjYTi1Qc0zZb5DefTKIyf5caGnQgexw3ifcc3HvZgbXQrRJjsH3AoITuERvl8eqYyYYk6+61aUVFR8Y/5BI0oYr4ya2wnAAAAAElFTkSuQmCC","orcid":"","institution":"Cochrane Austria, Department for Evidence-based Medicine and Clinical Epidemiology, University for Continuing Education Krems, Krems","correspondingAuthor":true,"prefix":"","firstName":"Lisa","middleName":"","lastName":"Affengruber","suffix":""},{"id":325892506,"identity":"81dde6ef-0e95-441f-bc77-509cb48b93fa","order_by":1,"name":"Miriam M. van der Maten","email":"","orcid":"","institution":"Knowledge Institute of Federation of Medical Specialists, Utrecht","correspondingAuthor":false,"prefix":"","firstName":"Miriam","middleName":"M. van der","lastName":"Maten","suffix":""},{"id":325892507,"identity":"5860f63e-9a70-4030-8abe-4f49a7738ffc","order_by":2,"name":"Isa Spiero","email":"","orcid":"","institution":"Julius Center for Health Sciences and Primary Care, University Medical Center Utrecht, Utrecht University, Utrecht","correspondingAuthor":false,"prefix":"","firstName":"Isa","middleName":"","lastName":"Spiero","suffix":""},{"id":325892508,"identity":"49c087c7-a62c-49ca-b45a-6f00e44fc9b2","order_by":3,"name":"Barbara Nussbaumer-Streit","email":"","orcid":"","institution":"Cochrane Austria, Department for Evidence-based Medicine and Clinical Epidemiology, University for Continuing Education Krems, Krems","correspondingAuthor":false,"prefix":"","firstName":"Barbara","middleName":"","lastName":"Nussbaumer-Streit","suffix":""},{"id":325892509,"identity":"3700821c-1905-4948-95e1-fa8fba1c2750","order_by":4,"name":"Mersiha Mahmić-Kaknjo","email":"","orcid":"","institution":"Zenica Cantonal Hospital, Department for Clinical Pharmacology, Zenica","correspondingAuthor":false,"prefix":"","firstName":"Mersiha","middleName":"","lastName":"Mahmić-Kaknjo","suffix":""},{"id":325892510,"identity":"3573e5dd-54d2-4696-94d8-7b10df1e4635","order_by":5,"name":"Moriah E. Ellen","email":"","orcid":"","institution":"Department of Health Policy and Management, Guilford Glazer Faculty of Business and Management and Faculty of Health Sciences, Ben-Gurion University of the Negev, Beer-Sheva","correspondingAuthor":false,"prefix":"","firstName":"Moriah","middleName":"E.","lastName":"Ellen","suffix":""},{"id":325892511,"identity":"61c9b664-3afc-4a8e-98da-04eb1b8f7107","order_by":6,"name":"Käthe Goossen","email":"","orcid":"","institution":"Witten/Herdecke University, Institute for Research in Operative Medicine (IFOM), Cologne","correspondingAuthor":false,"prefix":"","firstName":"Käthe","middleName":"","lastName":"Goossen","suffix":""},{"id":325892512,"identity":"370a8e82-bc12-4908-8fe7-e24ea024b77e","order_by":7,"name":"Lucia Kantorova","email":"","orcid":"","institution":"Czech National Centre for Evidence-Based Healthcare and Knowledge Translation, Institute of Biostatistics and Analyses, Faculty of Medicine, Masaryk University, Brno","correspondingAuthor":false,"prefix":"","firstName":"Lucia","middleName":"","lastName":"Kantorova","suffix":""},{"id":325892513,"identity":"605f400f-f112-44d6-902c-af45c2969a1f","order_by":8,"name":"Lotty Hooft","email":"","orcid":"","institution":"Cochrane Netherlands, Julius Center for Health Sciences and Primary Care, University Medical Center Utrecht, Utrecht University, Utrecht","correspondingAuthor":false,"prefix":"","firstName":"Lotty","middleName":"","lastName":"Hooft","suffix":""},{"id":325892514,"identity":"2df9f481-a949-4c9f-8c7d-7ff0d13a32ca","order_by":9,"name":"Nicoletta Riva","email":"","orcid":"","institution":"Department of Pathology, Faculty of Medicine and Surgery, University of Malta, Msida","correspondingAuthor":false,"prefix":"","firstName":"Nicoletta","middleName":"","lastName":"Riva","suffix":""},{"id":325892515,"identity":"01a91a91-165c-46aa-afa9-d3d64f649199","order_by":10,"name":"Georgios Poulentzas","email":"","orcid":"","institution":"Laboratory of Hygiene and Environmental Protection, Department of Medicine, Democritus University of Thrace, Alexandroupolis","correspondingAuthor":false,"prefix":"","firstName":"Georgios","middleName":"","lastName":"Poulentzas","suffix":""},{"id":325892516,"identity":"249fca0f-b6da-4cd9-a907-6f7ab9ead262","order_by":11,"name":"Panagiotis Nikolaos Lalagkas","email":"","orcid":"","institution":"Laboratory of Hygiene and Environmental Protection, Department of Medicine, Democritus University of Thrace, Alexandroupolis","correspondingAuthor":false,"prefix":"","firstName":"Panagiotis","middleName":"Nikolaos","lastName":"Lalagkas","suffix":""},{"id":325892517,"identity":"2f5af0e3-55a2-44d8-ad8a-7ee76da664cd","order_by":12,"name":"Anabela G. Silva","email":"","orcid":"","institution":"CINTESIS.RISE@UA, University of Aveiro, Campus Universitário de Santiago, Aveiro","correspondingAuthor":false,"prefix":"","firstName":"Anabela","middleName":"G.","lastName":"Silva","suffix":""},{"id":325892518,"identity":"33450218-d23f-4f08-aba1-96ce425f9ac1","order_by":13,"name":"Michele Sassano","email":"","orcid":"","institution":"Section of Hygiene, University Department of Life Sciences and Public Health, Università Cattolica del Sacro Cuore, Rome","correspondingAuthor":false,"prefix":"","firstName":"Michele","middleName":"","lastName":"Sassano","suffix":""},{"id":325892519,"identity":"167198de-fda1-4c02-bc26-ca8de997e10c","order_by":14,"name":"Raluca Sfetcu","email":"","orcid":"","institution":"National Institute for Health Services Management, Bucharest","correspondingAuthor":false,"prefix":"","firstName":"Raluca","middleName":"","lastName":"Sfetcu","suffix":""},{"id":325892520,"identity":"e0a1f094-32ca-42ca-9f94-210ec0979c7d","order_by":15,"name":"María E. Marqués","email":"","orcid":"","institution":"Red de Nutrición Basada en la Evidencia, Academia Española de Nutrición y Dietética, Pamplona","correspondingAuthor":false,"prefix":"","firstName":"María","middleName":"E.","lastName":"Marqués","suffix":""},{"id":325892521,"identity":"326075a9-fe8f-4cc0-8435-5a381d537bb1","order_by":16,"name":"Tereza Friessova","email":"","orcid":"","institution":"Department of Health Sciences, Faculty of Medicine, Masaryk University, Brno","correspondingAuthor":false,"prefix":"","firstName":"Tereza","middleName":"","lastName":"Friessova","suffix":""},{"id":325892522,"identity":"0c593b55-37b7-4710-bf1c-8be28eb93934","order_by":17,"name":"Eduard Baladia","email":"","orcid":"","institution":"Red de Nutrición Basada en la Evidencia, Academia Española de Nutrición y Dietética, Pamplona","correspondingAuthor":false,"prefix":"","firstName":"Eduard","middleName":"","lastName":"Baladia","suffix":""},{"id":325892523,"identity":"3e7fb8ab-8e04-4c6f-a395-eb6da4c75a4e","order_by":18,"name":"Angelo Maria Pezzullo","email":"","orcid":"","institution":"Section of Hygiene, University Department of Life Sciences and Public Health, Università Cattolica del Sacro Cuore, Rome","correspondingAuthor":false,"prefix":"","firstName":"Angelo","middleName":"Maria","lastName":"Pezzullo","suffix":""},{"id":325892524,"identity":"0937fb56-3d44-4a44-92f4-50ea168dd368","order_by":19,"name":"Patricia Martinez","email":"","orcid":"","institution":"Red de Nutrición Basada en la Evidencia, Academia Española de Nutrición y Dietética, Pamplona","correspondingAuthor":false,"prefix":"","firstName":"Patricia","middleName":"","lastName":"Martinez","suffix":""},{"id":325892525,"identity":"21966958-c74e-4927-b7cb-f8631bc48a38","order_by":20,"name":"Gerald Gartlehner","email":"","orcid":"","institution":"Cochrane Austria, Department for Evidence-based Medicine and Clinical Epidemiology, University for Continuing Education Krems, Krems","correspondingAuthor":false,"prefix":"","firstName":"Gerald","middleName":"","lastName":"Gartlehner","suffix":""},{"id":325892526,"identity":"2133f2cf-ab96-4015-b891-27b4c59d2fa7","order_by":21,"name":"René Spijker","email":"","orcid":"","institution":"Cochrane Netherlands, Julius Center for Health Sciences and Primary Care, University Medical Center Utrecht, Utrecht University, Utrecht","correspondingAuthor":false,"prefix":"","firstName":"René","middleName":"","lastName":"Spijker","suffix":""}],"badges":[],"createdAt":"2024-06-17 18:54:05","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-4595777/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-4595777/v1","draftVersion":[],"editorialEvents":[{"content":"https://doi.org/10.1186/s12874-024-02320-4","type":"published","date":"2024-09-18T15:56:53+00:00"}],"editorialNote":"","failedWorkflow":false,"files":[{"id":60123167,"identity":"e09bef0e-25a0-40d2-8f1c-ac3053ea3fa6","added_by":"auto","created_at":"2024-07-12 05:19:29","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":64298,"visible":true,"origin":"","legend":"\u003cp\u003eSteps of the SR process. *added to Tsafnat et al.’s list [1]. ** added based on Nussbaumer-Streit et al. [2]\u003c/p\u003e","description":"","filename":"1.png","url":"https://assets-eu.researchsquare.com/files/rs-4595777/v1/6c7d11cb56c8a5336d7a1882.png"},{"id":60123930,"identity":"d449a9c1-fa86-4410-82a6-e0c1301310d2","added_by":"auto","created_at":"2024-07-12 05:27:29","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":68229,"visible":true,"origin":"","legend":"\u003cp\u003ePRISMA flowchart\u003c/p\u003e","description":"","filename":"2.png","url":"https://assets-eu.researchsquare.com/files/rs-4595777/v1/170cffaee4660c177261c198.png"},{"id":60123931,"identity":"df3a6fcf-7d46-4968-8ef3-0007d2893d4b","added_by":"auto","created_at":"2024-07-12 05:27:29","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":20909,"visible":true,"origin":"","legend":"\u003cp\u003eThe number of identified evaluation studies per review step\u003c/p\u003e","description":"","filename":"3.png","url":"https://assets-eu.researchsquare.com/files/rs-4595777/v1/76bc21a031f215fadd8016ea.png"},{"id":60123173,"identity":"f16e1b7e-c619-4adf-87db-59535fc24b7c","added_by":"auto","created_at":"2024-07-12 05:19:29","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":20315,"visible":true,"origin":"","legend":"\u003cp\u003eOutcomes reported in the included studies\u003c/p\u003e","description":"","filename":"4.png","url":"https://assets-eu.researchsquare.com/files/rs-4595777/v1/eb8bb89ae77cae97fee5488e.png"},{"id":65104071,"identity":"c78ee393-bf36-4c92-b7ef-eed8396419de","added_by":"auto","created_at":"2024-09-23 16:11:27","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":1149380,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-4595777/v1/cbd2d967-e73d-4916-ad56-c033b5943ec6.pdf"},{"id":60123171,"identity":"642c1006-47b6-45ba-b48a-c2faf238b163","added_by":"auto","created_at":"2024-07-12 05:19:29","extension":"docx","order_by":1,"title":"","display":"","copyAsset":false,"role":"supplement","size":26495,"visible":true,"origin":"","legend":"","description":"","filename":"Appendix1.docx","url":"https://assets-eu.researchsquare.com/files/rs-4595777/v1/663e42e4706d710b1df54511.docx"},{"id":60123168,"identity":"b22e5846-7251-40e0-92da-8f80434b51e7","added_by":"auto","created_at":"2024-07-12 05:19:29","extension":"docx","order_by":2,"title":"","display":"","copyAsset":false,"role":"supplement","size":147559,"visible":true,"origin":"","legend":"","description":"","filename":"Appendix2.docx","url":"https://assets-eu.researchsquare.com/files/rs-4595777/v1/2a70e7dc5f77dffa93cb87d2.docx"},{"id":60123166,"identity":"444bd780-be35-4dbc-a0bc-bdc5f7514bf3","added_by":"auto","created_at":"2024-07-12 05:19:28","extension":"docx","order_by":3,"title":"","display":"","copyAsset":false,"role":"supplement","size":271367,"visible":true,"origin":"","legend":"","description":"","filename":"Appendix3.docx","url":"https://assets-eu.researchsquare.com/files/rs-4595777/v1/258331f402480f15655b2077.docx"}],"financialInterests":"No competing interests reported.","formattedTitle":"An exploration of available methods and tools to improve the efficiency of systematic review production - a scoping review ","fulltext":[{"header":"Introduction","content":"\u003cp\u003eSystematic reviews (SRs) provide the most valid way of synthesizing evidence as they follow a structured, rigorous, and transparent research process. Because of their thoroughness, SRs have a long history in informing health policy decision-making, clinical guidelines, and primary research\u0026nbsp;[3].\u0026nbsp;However, the standard method employed in high-quality SRs involves many steps that are predominantly conducted manually, resulting in a laborious and time-intensive process lasting an average of fifteen months\u0026nbsp;[4]. With the exponential growth of scientific literature, the challenge of SR development is exacerbated. Especially, this is the case in contexts where timely evidence is imperative and decision-makers need urgent answers, as demonstrated during the coronavirus pandemic\u0026nbsp;[5]. Additionally, researchers\u0026rsquo; aspirations to undertake SRs prior to initiating new primary studies is hindered by the complex and resource-intensive nature of the SR development process\u0026nbsp;[6].\u003c/p\u003e\n\u003cp\u003eIn response to these challenges, a surge of interest in methods to accelerate SR development occurred in recent years, leading to the emergence of \u0026ldquo;rapid reviews\u0026rdquo; (RRs). Methods used in RR development include searching a limited number of bibliographic databases, single-reviewer literature screening, or abbreviated quality assessment\u0026nbsp;[7]. However, depending on various factors, the trade-off between the time saved and the potential reduction in quality and comprehensiveness is a critical issue that must be carefully weighed and discussed with stakeholders. Concurrently, efforts are underway to leverage technological innovations to expedite the research process, involving machine learning, natural language processing, active learning, or text mining to mimic human activities in SR tasks\u0026nbsp;[8]. Supportive tools can offer varying levels of automation and decision-making, ranging from basic file management to fully automated decision-making, as outlined by O\u0026rsquo;Connor et al.\u0026nbsp;[9]. These levels include semiautomated tools for workflow prioritization to fully automated decision-making processes\u0026nbsp;[9]. Some tools offer researchers ready-to-use applications, while other algorithms are not yet developed into user-friendly tools\u0026nbsp;[10].\u0026nbsp;Table 1\u0026nbsp;provides an overview of commonly used terms in this scoping review and their definitions.\u003c/p\u003e\n\u003cp\u003eTable\u0026nbsp;1:\u0026nbsp;Commonly used terms\u0026nbsp;\u003c/p\u003e\n\u003ctable border=\"1\" cellspacing=\"0\" cellpadding=\"0\" width=\"606\"\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd width=\"24.628099173553718%\" valign=\"top\"\u003e\n \u003cp\u003e\u003cstrong\u003eTerm\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"75.37190082644628%\" valign=\"top\"\u003e\n \u003cp\u003e\u003cstrong\u003eDefinition\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"24.628099173553718%\" valign=\"top\"\u003e\n \u003cp\u003eText mining\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"75.37190082644628%\" valign=\"top\"\u003e\n \u003cp\u003eText mining is the process of extracting useful information from unstructured text data. It involves analyzing large amounts of text data to discover patterns, relationships, and insights that can be used to make informed decisions. Text mining techniques include natural language processing, machine learning, active learning, and statistical analysis\u0026nbsp;[9, 11].\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"24.628099173553718%\" valign=\"top\"\u003e\n \u003cp\u003eMachine learning\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"75.37190082644628%\" valign=\"top\"\u003e\n \u003cp\u003eMachine learning is a type of artificial intelligence that allows computer systems to automatically improve their performance at a specific task by learning from data without being explicitly programmed. It involves developing algorithms that can learn patterns and relationships in data and then use this knowledge to make predictions or decisions about new data\u0026nbsp;[9].\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"24.628099173553718%\" valign=\"top\"\u003e\n \u003cp\u003eActive learning\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"75.37190082644628%\" valign=\"top\"\u003e\n \u003cp\u003eActive learning is a machine learning technique that selects which data to learn from, rather than being trained on a prelabeled dataset. In active learning, the algorithm iteratively selects the most informative data points for labeling by a human expert and then uses these labeled data points to update its model. This process continues until the model achieves a desired level of accuracy or until labeling additional data points no longer improves the model\u0026rsquo;s performance\u0026nbsp;[9].\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"24.628099173553718%\" valign=\"top\"\u003e\n \u003cp\u003eFully automated\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"75.37190082644628%\" valign=\"top\"\u003e\n \u003cp\u003eFull automation means that the tool makes decisions without human input, for example, the tool automatically conducts a risk of bias assessment of a randomized controlled trial without requiring input from a human reviewer\u0026nbsp;[9, 12].\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"24.628099173553718%\" valign=\"top\"\u003e\n \u003cp\u003eSemiautomated\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"75.37190082644628%\" valign=\"top\"\u003e\n \u003cp\u003eSemiautomated tools recommend decisions, but manual confirmation is required. For instance, the tool suggest whether a record is probably eligible for inclusion, but additional human judgment is necessary to make the final decision for in- or exclusion, which is useful in, for example, the prioritization of relevant abstracts and duplicate detection\u0026nbsp;[9, 12].\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"24.628099173553718%\" valign=\"top\"\u003e\n \u003cp\u003eSupportive tools without automation\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"75.37190082644628%\" valign=\"top\"\u003e\n \u003cp\u003eTools that improve the file management process, for example, citation databases, reference management tools, and SR management tools, but do not include any form of automation\u0026nbsp;[9].\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"24.628099173553718%\" valign=\"top\"\u003e\n \u003cp\u003eSupportive tools\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"75.37190082644628%\" valign=\"top\"\u003e\n \u003cp\u003eTools that use text mining, machine and active learning, and/or fully or semiautomated approaches, but also tools without automation to support RR production.\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"24.628099173553718%\" valign=\"top\"\u003e\n \u003cp\u003eReady-to-use tools\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"75.37190082644628%\" valign=\"top\"\u003e\n \u003cp\u003eTools that offer researchers ready-to-use, fully developed applications, while other algorithms are not yet developed into user-friendly tools.\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"24.628099173553718%\" valign=\"top\"\u003e\n \u003cp\u003eCrowdsourcing\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"75.37190082644628%\" valign=\"top\"\u003e\n \u003cp\u003eUtilizing a large group of human volunteers to perform a task online.\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"24.628099173553718%\" valign=\"top\"\u003e\n \u003cp\u003eWorkload saving\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"75.37190082644628%\" valign=\"top\"\u003e\n \u003cp\u003eWorkload saving refers to the reduction in the amount of work required to complete a task or process through abbreviated methods or tools in comparison to SR methodology. In this context, workload saving refers to reducing the number of references or documents that must be screened, reviewed, or analyzed during the review process.\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"24.628099173553718%\" valign=\"top\"\u003e\n \u003cp\u003eTime-saving\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"75.37190082644628%\" valign=\"top\"\u003e\n \u003cp\u003eTime-saving in this context refers to the reduction in the amount of time measured in hours or minutes required to complete the review process through abbreviated methods or supportive tools.\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"24.628099173553718%\" valign=\"top\"\u003e\n \u003cp\u003eCost-saving\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"75.37190082644628%\" valign=\"top\"\u003e\n \u003cp\u003eCost-saving in this context refers to the reduction in financial expenditures associated with conducting the review process. It mainly encompasses personnel costs required to complete the review.\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"24.628099173553718%\" valign=\"top\"\u003e\n \u003cp\u003eImpact on results\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"75.37190082644628%\" valign=\"top\"\u003e\n \u003cp\u003eImpact on results in this context refers to changes in statistical significance, effect estimates, or conclusions as a result of abbreviated methods or supporting tools during the review process.\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"24.628099173553718%\" valign=\"top\"\u003e\n \u003cp\u003eValidity outcomes\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"75.37190082644628%\" valign=\"top\"\u003e\n \u003cp\u003eValidity outcomes in this context describe the degree to which the method or tool is aligned with the SR methodology. Validity is measured as the tool\u0026rsquo;s precision, sensitivity, specificity, accuracy, errors made, etc., compared to SR methods done by humans.\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"24.628099173553718%\" valign=\"top\"\u003e\n \u003cp\u003eUsability\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"75.37190082644628%\" valign=\"top\"\u003e\n \u003cp\u003eUsability refers to the extent to which end users can use the method/tool in a specific context to achieve specific goals effectively, efficiently, and satisfactorily, and how the end users experienced it\u0026nbsp;[13]. In the context of SR tools, usability is often reported as evaluation results (e.g., usability scores) of the method or tool or end-user experiences including joy of use, interface design, etc.\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n\u003c/table\u003e\n\u003cp\u003e\u003cem\u003eAbbreviations: RR: rapid review, SR: systematic review\u003c/em\u003e\u003c/p\u003e\n\u003cp\u003eWhile these methods and tools hold promise for enhancing the efficiency of SR production, their widespread adoption faces challenges, including limited awareness among review teams, concerns about validity, and usability issues\u0026nbsp;[14]. Addressing these barriers requires evaluations to determine the validity and usability of various methods and tools across different stages of the review process\u0026nbsp;[10, 15].\u003c/p\u003e\n\u003cp\u003eTo bridge this gap, we conducted a scoping review to comprehensively map the landscape of methods and tools aimed at improving the efficiency of SR development, assessing their validity, resource utilization (workload/time/costs), and impact on results, as well as exploring usability for all steps of the review process. This review complements another scoping review that identified the most resource-intensive areas when conducting or updating a SR\u0026nbsp;[2]. We mapped the efficiency outcomes of each method and tool against the steps of the SR process. Specifically, our scoping review aimed to answer the following key questions (KQs):\u003c/p\u003e\n\u003col start=\"1\" type=\"1\"\u003e\n \u003cli\u003eWhich methods and tools are used to improve the efficiency of SR production?\u003c/li\u003e\n \u003cli\u003eHow efficient are these methods or tools regarding validity, resource use, and impact on results?\u003c/li\u003e\n \u003cli\u003eHow was the user experience when using these methods and tools?\u003c/li\u003e\n\u003c/ol\u003e"},{"header":"Methods","content":"\u003cp\u003eWe conducted this review as part of working group 3 of EVBRES (EVidence-Bases RESearch) COST Action CA17117 (www.ebvres.eu). We published the protocol for this scoping review on June 18, 2020 (Open Science Framework:\u0026nbsp;https://osf.io/9423z).\u003c/p\u003e\n\u003ch2\u003eStudy design\u003c/h2\u003e\n\u003cp\u003eWe conducted a scoping review following the guidance of Arksey and O’Malley\u0026nbsp;[16], Levac et al.\u0026nbsp;[17], and Peters et al.\u0026nbsp;[7]. Within EVBRES, we adopted the definition of a scoping review as “a form of knowledge synthesis that addresses an exploratory research question aimed at mapping key concepts, types of evidence, and gaps in research related to a defined area or field by systematically searching, selecting, and synthesising existing knowledge”\u0026nbsp;[18]. We report our review in accordance with the PRISMA Extension for Scoping Reviews (PRISMA-ScR)\u0026nbsp;[19].\u003c/p\u003e\n\u003ch2\u003eInformation sources and search\u003c/h2\u003e\n\u003cp\u003eThe search for this scoping review followed an iterative three-step process recommended by the Joanna Briggs Institute\u0026nbsp;[20]:\u003c/p\u003e\n\u003cp\u003e1) First, an information specialist (RS) conducted a preliminary limited, focused search in Scopus in March 2020. We screened the search results and analyzed relevant studies to discover additional relevant keywords and sources.\u003c/p\u003e\n\u003cp\u003e\u003cbr\u003e\u0026nbsp;2) Second, based on identified search terms from the included studies, the information specialist performed a comprehensive search (November 2021) in MEDLINE and Embase,\u0026nbsp;both via Ovid. The comprehensive MEDLINE strategy was reviewed by another information specialist (IK) in accordance with the Peer Review of Electronic Search Strategies (PRESS) guideline\u0026nbsp;[21].\u003c/p\u003e\n\u003cp\u003e3) Third, we checked the reference lists of the identified studies and background articles, conducted grey literature searches (e.g., organizations that produce SRs and RRs), and contacted experts in the field. In addition, to identify grey literature, we searched for conference proceedings covered in Embase and checked if the associated full text was available.\u003c/p\u003e\n\u003cp\u003eWe limited the database searches to articles on methodological adaptations published since 1997, as this was the first year of mention in the published literature of methods to make the review process more efficient\u0026nbsp;[22]. For tools, we limited the search to articles published since 2005, as this was the first year of mention of a text mining model in the published literature, according to Jonnalagadda and Petitti (2013)\u0026nbsp;[23]. The search strategies are provided in Appendix 1. Our search was updated on December 14, 2023 to include evidence published since our initial search.\u003c/p\u003e\n\u003cp\u003e\u003cem\u003eEligibility criteria\u003c/em\u003e\u003c/p\u003e\n\u003cp\u003eThe eligibility criteria are outlined in Table 2. Our focus was on incorporating primary studies that assessed the efficacy of automated, ready-to-use applications of tools, or RR methods within the SR process. Specifically, we sought tools that demand no programming expertise, relying instead on user-friendly interfaces devoid of complex codes, syntaxes, or algorithms. We were interested in studies assessing their use within one or more of the fifteen steps of the SR process as defined by Tsafnat et al. (2014), further supplemented by the steps “critical appraisal,” “grading the certainty of evidence,” and “administration/project management”\u0026nbsp;[24]\u0026nbsp;[2]\u0026nbsp;(Figure 1).\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eTable\u0026nbsp;2: Eligibility criteria for study inclusion in the scoping review\u003c/p\u003e\n\u003ctable border=\"1\" cellspacing=\"0\" cellpadding=\"0\" width=\"601\"\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd width=\"12.645590682196339%\" valign=\"top\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd width=\"42.42928452579035%\" valign=\"top\"\u003e\n \u003cp\u003e\u003cstrong\u003eInclusion criteria\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"44.925124792013314%\" valign=\"top\"\u003e\n \u003cp\u003e\u003cstrong\u003eExclusion criteria\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"12.645590682196339%\" valign=\"top\"\u003e\n \u003cp\u003e\u003cstrong\u003eTopic\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"42.42928452579035%\" valign=\"top\"\u003e\n \u003cp\u003eStudies reporting the performance of methods/ready-to-use tools with regard to improving efficiency when producing or updating a SR in the health field\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"44.925124792013314%\" valign=\"top\"\u003e\n \u003cp\u003eStudies not reporting on improving efficiency\u003c/p\u003e\n \u003cp\u003eStudies assessing algorithms, classifiers, or models that are not usable as tools\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"12.645590682196339%\" valign=\"top\"\u003e\n \u003cp\u003e\u003cstrong\u003eConcept\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"42.42928452579035%\" valign=\"top\"\u003e\n \u003cp\u003eAddressing improvement in efficiency within one or more steps of a SR or RR (as depicted in Figure 1) of health intervention or diagnostic or prognostic studies\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"44.925124792013314%\" valign=\"top\"\u003e\n \u003cp\u003eStudies assessing other steps such as the dissemination of results\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"12.645590682196339%\" valign=\"top\"\u003e\n \u003cp\u003e\u003cstrong\u003eOutcomes\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"42.42928452579035%\" valign=\"top\"\u003e\n \u003cp\u003eWorkload saved\u003c/p\u003e\n \u003cp\u003eTime saved\u003c/p\u003e\n \u003cp\u003ePersonnel effort saved\u003c/p\u003e\n \u003cp\u003eCosts saved\u003c/p\u003e\n \u003cp\u003eImpact on results, conclusions\u003c/p\u003e\n \u003cp\u003eUsability\u003c/p\u003e\n \u003cp\u003eValidity outcomes: for example, recall (sensitivity), precision, accuracy, specificity, false positives, false negatives, reliability, proportion of relevant studies missed, prediction performance, ordering performance\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"44.925124792013314%\" valign=\"top\"\u003e\n \u003cp\u003eAny other outcomes\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"12.645590682196339%\" valign=\"top\"\u003e\n \u003cp\u003e\u003cstrong\u003eStudy design/ publication type\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"42.42928452579035%\" valign=\"top\"\u003e\n \u003cp\u003eEvaluation study of method/ready-to-use\u0026nbsp;tool (e.g., method studies, RCTs, nRCTs, cohort studies,..)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"44.925124792013314%\" valign=\"top\"\u003e\n \u003cp\u003eDevelopment studies of ready-to-use tools\u003c/p\u003e\n \u003cp\u003eValidation studies of of ready-to-use tools\u003c/p\u003e\n \u003cp\u003eEvaluation studies of tools that are no longer available\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"12.645590682196339%\" valign=\"top\"\u003e\n \u003cp\u003e\u003cstrong\u003eFull text availability\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"42.42928452579035%\" valign=\"top\"\u003e\n \u003cp\u003eFull text retrievable via our libraries\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"44.925124792013314%\" valign=\"top\"\u003e\n \u003cp\u003eFull text not retrievable via our libraries\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"12.645590682196339%\" valign=\"top\"\u003e\n \u003cp\u003e\u003cstrong\u003eTiming\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"87.35440931780366%\" colspan=\"2\" valign=\"top\"\u003e\n \u003cp\u003eStudies assessing methods: 1997\u003c/p\u003e\n \u003cp\u003eStudies assessing tools: 2005\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"12.645590682196339%\" valign=\"top\"\u003e\n \u003cp\u003e\u003cstrong\u003eLanguage\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"87.35440931780366%\" colspan=\"2\" valign=\"top\"\u003e\n \u003cp\u003eAll languages\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n\u003c/table\u003e\n\u003cp\u003e\u003cem\u003eAbbreviations: nRCT: nonrandomized-controlled trial, RCT: randomized controlled trial, RR: rapid review, SR: systematic review\u003c/em\u003e\u003c/p\u003e\n\u003ch2\u003eStudy selection\u003c/h2\u003e\n\u003cp\u003eWe piloted the abstract screening with 50 records and the full-text screening with five records. Following the piloting, the results were discussed with all reviewers, and the screening guidance was updated to include clarifications wherever necessary. The review team used Covidence (www.covidence.org) to dually screen titles/abstracts and full texts. We resolved conflicts throughout the screening process through re-examination of the study and subsequent discussion and, if necessary, by consulting a third reviewer.\u003c/p\u003e\n\u003ch2\u003eData charting\u0026nbsp;\u003c/h2\u003e\n\u003cp\u003eWe developed a data extraction form and pilot-tested it before implementation using Google Forms. The data abstraction was done by one author and checked by a second author to ensure consistency and correctness in the extracted data. A third author made final decisions in cases of discrepancies. We extracted relevant study characteristics and outcomes per review step.\u003c/p\u003e\n\u003ch2\u003eData mapping\u0026nbsp;\u003c/h2\u003e\n\u003cp\u003eWe mapped the identified methods and tools by each review step and summarized the outcomes of individual studies. As the objective of this scoping review was to descriptively map efficiency outcomes and the usability of methods and tools against SR production steps, we did not apply a formal certainty of evidence or risk of bias (RoB) assessment. Additionally, we used data mapping to identify research gaps.\u003c/p\u003e"},{"header":"Results","content":"\u003cp\u003eWe included 103 studies\u0026nbsp;[12, 25-126]\u0026nbsp;evaluating 21 methods (n=51)\u0026nbsp;[25, 28, 29, 31, 33, 37, 39-42, 50, 52, 53, 55, 60, 61, 64-66, 72, 74, 76, 78, 79, 81, 82, 84-92, 99-101, 104, 106, 108-111, 113, 114, 116, 121, 122, 125, 126]\u0026nbsp;and 35 tools (n=54)\u0026nbsp;[12, 26, 27, 30, 32, 34-36, 38, 43-49, 51, 54, 56-59, 62, 63, 67-71, 73-75, 77, 80, 83, 93-98, 102, 103, 105, 107, 108, 112, 115, 117-120, 124, 127]\u0026nbsp;(Figure 2: PRISMA study flowchart).\u0026nbsp;Table 3\u0026nbsp;provides an overview of the identified methods and tools. A total of 73 studies were validity studies (n=70)\u0026nbsp;[25-34, 37, 39, 43-46, 48, 49, 51, 52, 56-58, 60, 62-66, 69, 70, 75, 77-79, 81-85, 87-91, 94-97, 99, 101-103, 105-112, 114-117, 119-122, 124, 125]\u0026nbsp;or usability studies (n=3)\u0026nbsp;[59, 67, 68]\u0026nbsp;assessing a single method or tool, and 30 studies performed comparative analyses of different methods or tools\u0026nbsp;[12, 35, 36, 38, 40-42, 47, 50, 53-55, 61, 71-74, 76, 80, 86, 92, 93, 98, 100, 104, 108, 113, 118, 126, 127]. Few studies prospectively evaluated methods or tools in a real-world workflow (n=20)\u0026nbsp;[12, 27, 32, 35, 46, 50, 67, 68, 77, 78, 88, 90, 94, 98, 105, 108, 112, 114, 125, 127],\u0026nbsp;7 studies of those used independent testing (by a different reviewer team) with external data\u0026nbsp;[12, 35, 46, 94, 98, 112, 127].\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eThe majority of studies evaluated methods or tools for supporting the tasks of title and abstract screening (n=42)\u0026nbsp;[32, 35, 36, 38, 43-45, 47, 48, 51, 55, 58, 59, 63, 73, 79, 82-84, 86-90, 94-96, 100, 102, 105-108, 112-114, 117-120, 125, 127]\u0026nbsp;or devising the search strategy and performing the search (n=24)\u0026nbsp;[28, 34, 37, 39, 42, 52, 53, 56, 60, 64-66, 72, 81, 91-93, 98, 99, 104, 110, 121, 122, 126]\u0026nbsp;(see\u0026nbsp;Figure 2). For several steps of the SR process, only a few studies that evaluated methods or tools were identified: deduplication: n=6\u0026nbsp;[30, 36, 54, 57, 71, 80], additional search: n=2\u0026nbsp;[33, 97], update search: n=6\u0026nbsp;[36, 50, 61, 77, 109, 111], full-text selection: n=4\u0026nbsp;[85, 113, 114, 125], data extraction: n=11\u0026nbsp;[31, 36, 46, 67, 69, 70, 74, 103, 112, 124, 125]; critical appraisal: n=9,\u0026nbsp;[26, 27, 36, 49, 62, 68, 75, 101, 115], and\u0026nbsp;combination of abbreviated methods/tools: n=6\u0026nbsp;[12, 25, 76, 78, 100, 116]\u0026nbsp;(see\u0026nbsp;Figure 2). No studies were found for some steps of the SR process, such as administration/project management, formulating the review question, searching for existing reviews, writing the protocol, full-text retrieval, synthesis/meta-analysis, certainty of evidence assessment, and report preparation. In Appendix 2, we summarize the characteristics of all the included studies.\u003c/p\u003e\n\u003cp\u003eMost studies reported on validity outcomes (n=84, 46%) [12, 25-34, 38, 41-58, 61-63, 65, 66, 68-75, 77, 79, 80, 82-90, 93, 95, 96, 98, 100-113, 115-124, 126], while outcomes such as workload saving (n=35, 19%) [12, 27, 28, 31, 32, 38, 43-45, 47, 48, 51, 58, 63, 66, 83-86, 90, 94, 96, 102, 104-108, 110, 113, 114, 116, 118, 120, 126], time-saving (n=24, 13%) [32-34, 38, 44, 46, 47, 50, 51, 69-71, 73, 74, 86, 90, 95-98, 103, 108, 124, 125]; impact on results (n=23, 13%) [25, 31, 37, 39-42, 58, 60, 61, 64, 76, 78, 81, 91, 92, 99, 100, 110, 114, 116, 121, 126], usability (n=13, 7%) [35, 36, 38, 59, 67, 68, 77, 93, 94, 96, 97, 120, 124] and cost-saving (n=3, 2%) [32, 82, 113] were less evaluated (Figure 4: Outcomes reported in the included studies). In Appendix 2, we map the efficiency and usability outcomes per tool and method against the review steps of the SR process.The included studies reported various validity outcomes (i.e., specificity, precision, accuracy) and time, costs, or workload savings to undertake the review. None of the studies reported the personnel effort saved. Figure 3 presents the frequency of reported outcomes in the included studies.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eTable\u0026nbsp;3: Identified methods and tools per review step\u003c/p\u003e\n\u003ctable border=\"1\" cellspacing=\"0\" cellpadding=\"0\" width=\"604\"\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd width=\"34.437086092715234%\" valign=\"top\"\u003e\n \u003cp\u003e\u003cstrong\u003eReview step\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"32.78145695364238%\" valign=\"top\"\u003e\n \u003cp\u003e\u003cstrong\u003eMethods\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"32.78145695364238%\" valign=\"top\"\u003e\n \u003cp\u003e\u003cstrong\u003eTools\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"34.437086092715234%\" valign=\"top\"\u003e\n \u003cp\u003e\u003cstrong\u003eAdministration/project management\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"32.78145695364238%\" valign=\"top\"\u003e\n \u003cp\u003e-\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"32.78145695364238%\" valign=\"top\"\u003e\n \u003cp\u003e-\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"34.437086092715234%\" valign=\"top\"\u003e\n \u003cp\u003e\u003cstrong\u003eReview question formulation\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"32.78145695364238%\" valign=\"top\"\u003e\n \u003cp\u003e-\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"32.78145695364238%\" valign=\"top\"\u003e\n \u003cp\u003e-\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"34.437086092715234%\" valign=\"top\"\u003e\n \u003cp\u003e\u003cstrong\u003eSearch for existing reviews\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"32.78145695364238%\" valign=\"top\"\u003e\n \u003cp\u003e-\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"32.78145695364238%\" valign=\"top\"\u003e\n \u003cp\u003e-\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"34.437086092715234%\" valign=\"top\"\u003e\n \u003cp\u003e\u003cstrong\u003eProtocol preparation\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"32.78145695364238%\" valign=\"top\"\u003e\n \u003cp\u003e-\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"32.78145695364238%\" valign=\"top\"\u003e\n \u003cp\u003e-\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"34.437086092715234%\" valign=\"top\"\u003e\n \u003cp\u003e\u003cstrong\u003eLiterature search\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"32.78145695364238%\" valign=\"top\"\u003e\n \u003c/td\u003e\n \u003ctd width=\"32.78145695364238%\" valign=\"top\"\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"34.437086092715234%\" valign=\"top\"\u003e\n \u003cp\u003e\u003cstrong\u003e\u003cem\u003eSearch strategy and database search\u003c/em\u003e\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"32.78145695364238%\" valign=\"top\"\u003e\n \u003cul\u003e\n \u003cli\u003eCitation-based searching\u0026nbsp;[28, 65, 66]\u003c/li\u003e\n \u003cli\u003eSearch restrictions: database\u0026nbsp;[53, 92, 104, 121], language\u0026nbsp;[37, 39, 60, 64, 81, 91, 99]\u003c/li\u003e\n \u003cli\u003eAbbreviated search strategies for study type\u0026nbsp;[52, 72, 110], topic\u0026nbsp;[122], or search date\u0026nbsp;[42, 126]\u003c/li\u003e\n \u003c/ul\u003e\n \u003c/td\u003e\n \u003ctd width=\"32.78145695364238%\" valign=\"top\"\u003e\n \u003cul\u003e\n \u003cli\u003eMeSH on Demand\u0026nbsp;[93, 98]\u003c/li\u003e\n \u003cli\u003ePubReMiner\u0026nbsp;[93, 98]\u003c/li\u003e\n \u003cli\u003ePolyglot Search Translator\u0026nbsp;[34]\u003c/li\u003e\n \u003cli\u003eRisklick search platform\u0026nbsp;[56]\u003c/li\u003e\n \u003cli\u003eYale MeSH Analyzer\u0026nbsp;[93, 98]\u003c/li\u003e\n \u003c/ul\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"34.437086092715234%\" valign=\"top\"\u003e\n \u003cp\u003e\u003cstrong\u003e\u003cem\u003eDeduplication\u003c/em\u003e\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"32.78145695364238%\" valign=\"top\"\u003e\n \u003cp\u003e-\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"32.78145695364238%\" valign=\"top\"\u003e\n \u003cul\u003e\n \u003cli\u003eASySD\u0026nbsp;[57]\u003c/li\u003e\n \u003cli\u003eEBSCO\u0026nbsp;[71]\u003c/li\u003e\n \u003cli\u003eEndNote\u0026nbsp;[54, 57, 71, 80]\u003c/li\u003e\n \u003cli\u003eCovidence\u0026nbsp;[80]\u003c/li\u003e\n \u003cli\u003eDeduklick\u0026nbsp;[30]\u003c/li\u003e\n \u003cli\u003eMendely\u0026nbsp;[54, 71, 80]\u003c/li\u003e\n \u003cli\u003eOVID\u0026nbsp;[71, 80]\u003c/li\u003e\n \u003cli\u003eRayyan\u0026nbsp;[54, 80]\u003c/li\u003e\n \u003cli\u003eRefworks\u0026nbsp;[71]\u003c/li\u003e\n \u003cli\u003eSRA- Deduplicator\u0026nbsp;[36, 54, 57]\u003c/li\u003e\n \u003cli\u003eZotero\u0026nbsp;[54, 80]\u003c/li\u003e\n \u003c/ul\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"34.437086092715234%\" valign=\"top\"\u003e\n \u003cp\u003e\u003cstrong\u003e\u003cem\u003eAdditional search\u003c/em\u003e\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"32.78145695364238%\" valign=\"top\"\u003e\n \u003cul\u003e\n \u003cli\u003eScopus approach\u0026nbsp;[33]\u003c/li\u003e\n \u003c/ul\u003e\n \u003c/td\u003e\n \u003ctd width=\"32.78145695364238%\" valign=\"top\"\u003e\n \u003cul\u003e\n \u003cli\u003ePaperfetcher\u0026nbsp;[97]\u003c/li\u003e\n \u003c/ul\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"34.437086092715234%\" valign=\"top\"\u003e\n \u003cp\u003e\u003cstrong\u003e\u003cem\u003eUpdate search\u003c/em\u003e\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"32.78145695364238%\" valign=\"top\"\u003e\n \u003cul\u003e\n \u003cli\u003eClinical Query search combined with PubMed-related articles search\u0026nbsp;[109, 111]\u003c/li\u003e\n \u003cli\u003eClinical Query search in MEDLINE and Embase\u0026nbsp;[61]\u003c/li\u003e\n \u003cli\u003e·\u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;McMaster Premium LiteratUre Service\u0026nbsp;[61]\u003c/li\u003e\n \u003cli\u003ePubMed Similar Articles Search\u0026nbsp;[50]\u003c/li\u003e\n \u003cli\u003eScopus Citation Tracking\u0026nbsp;[50]\u003c/li\u003e\n \u003c/ul\u003e\n \u003c/td\u003e\n \u003ctd width=\"32.78145695364238%\" valign=\"top\"\u003e\n \u003cul\u003e\n \u003cli\u003eRobotReviewer LIVE\u0026nbsp;[36, 77]\u003c/li\u003e\n \u003c/ul\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"34.437086092715234%\" valign=\"top\"\u003e\n \u003cp\u003e\u003cstrong\u003eStudy selection\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"32.78145695364238%\" valign=\"top\"\u003e\n \u003c/td\u003e\n \u003ctd width=\"32.78145695364238%\" valign=\"top\"\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"34.437086092715234%\" valign=\"top\"\u003e\n \u003cp\u003e\u003cstrong\u003e\u003cem\u003eTitle and abstract selection\u003c/em\u003e\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"32.78145695364238%\" valign=\"top\"\u003e\n \u003cul\u003e\n \u003cli\u003eCrowdsourcing using different (automation) tools\u0026nbsp;[84, 86-90]\u003c/li\u003e\n \u003cli\u003eDual computer monitors\u0026nbsp;[125]\u003c/li\u003e\n \u003cli\u003eSingle-reviewer screening\u0026nbsp;[55, 100, 113, 114]\u003c/li\u003e\n \u003cli\u003eLimited screening:\u003col\u003e\n \u003cli\u003ePICO-based title-only screening\u0026nbsp;[106]\u003c/li\u003e\n \u003cli\u003eReview of reviews\u0026nbsp;[108]\u003c/li\u003e\n \u003c/ol\u003e\n \u003c/li\u003e\n \u003cli\u003eTitle-first screening\u0026nbsp;[79]\u003c/li\u003e\n \u003c/ul\u003e\n \u003c/td\u003e\n \u003ctd width=\"32.78145695364238%\" valign=\"top\"\u003e\n \u003cul\u003e\n \u003cli\u003eAbstrackR\u0026nbsp;[36, 38, 44, 45, 47, 48, 51, 59, 73, 107, 108, 118]\u003c/li\u003e\n \u003cli\u003eASReview\u0026nbsp;[83, 95, 102, 120]\u003c/li\u003e\n \u003cli\u003eChatGPT\u0026nbsp;[112]\u003c/li\u003e\n \u003cli\u003eColandr\u0026nbsp;[38]\u003c/li\u003e\n \u003cli\u003eCovidence\u0026nbsp;[35, 36, 59]\u003c/li\u003e\n \u003cli\u003eDistillerSR\u0026nbsp;[36, 43, 47, 58]\u003c/li\u003e\n \u003cli\u003eEPPI-reviewer\u0026nbsp;[36, 59, 118, 127]\u003c/li\u003e\n \u003cli\u003eRayyan\u0026nbsp;[35, 36, 38, 59, 73, 94, 96, 119, 127]\u003c/li\u003e\n \u003cli\u003eRCT classifier\u0026nbsp;[117]\u003c/li\u003e\n \u003cli\u003eResearch screener\u0026nbsp;[32]\u003c/li\u003e\n \u003cli\u003eRobotAnalyst\u0026nbsp;[35, 36, 47, 105, 108]\u003c/li\u003e\n \u003cli\u003eSRA-helper for EndNote\u0026nbsp;[35]\u003c/li\u003e\n \u003cli\u003eSWIFT-active screener\u0026nbsp;[36, 63]\u003c/li\u003e\n \u003cli\u003eSWIFT-Review\u0026nbsp;[73]\u003c/li\u003e\n \u003c/ul\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"34.437086092715234%\" valign=\"top\"\u003e\n \u003cp\u003e\u003cstrong\u003e\u003cem\u003eFull-text retrieval\u003c/em\u003e\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"32.78145695364238%\" valign=\"top\"\u003e\n \u003cp\u003e-\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"32.78145695364238%\" valign=\"top\"\u003e\n \u003cp\u003e-\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"34.437086092715234%\" valign=\"top\"\u003e\n \u003cp\u003e\u003cstrong\u003e\u003cem\u003eFull-text selection\u003c/em\u003e\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"32.78145695364238%\" valign=\"top\"\u003e\n \u003cul\u003e\n \u003cli\u003eCrowdsourcing using different (automation-) tools\u0026nbsp;[85]\u003c/li\u003e\n \u003cli\u003eDual computer monitors\u0026nbsp;[125]\u003c/li\u003e\n \u003cli\u003eSingle-reviewer screening\u0026nbsp;[113, 114]\u003c/li\u003e\n \u003c/ul\u003e\n \u003c/td\u003e\n \u003ctd width=\"32.78145695364238%\" valign=\"top\"\u003e\n \u003cp\u003e-\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"34.437086092715234%\" valign=\"top\"\u003e\n \u003cp\u003e\u003cstrong\u003eData extraction\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"32.78145695364238%\" valign=\"top\"\u003e\n \u003cul\u003e\n \u003cli\u003eDual computer monitors\u0026nbsp;[125]\u003c/li\u003e\n \u003cli\u003eSingle-reviewer data extraction\u0026nbsp;[31, 74]\u003c/li\u003e\n \u003c/ul\u003e\n \u003c/td\u003e\n \u003ctd width=\"32.78145695364238%\" valign=\"top\"\u003e\n \u003cul\u003e\n \u003cli\u003eChatGPT\u0026nbsp;[112]\u003c/li\u003e\n \u003cli\u003eData Abstraction Assistant\u0026nbsp;[36, 67, 74]\u003c/li\u003e\n \u003cli\u003eDextr\u0026nbsp;[124]\u003c/li\u003e\n \u003cli\u003eExaCT\u0026nbsp;[46, 70, 103]\u003c/li\u003e\n \u003cli\u003ePlot Digitizer\u0026nbsp;[69]\u003c/li\u003e\n \u003c/ul\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"34.437086092715234%\" valign=\"top\"\u003e\n \u003cp\u003e\u003cstrong\u003eCritical appraisal\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"32.78145695364238%\" valign=\"top\"\u003e\n \u003cul\u003e\n \u003cli\u003eCrowdsourcing with CrowdCARE\u0026nbsp;[101]\u003c/li\u003e\n \u003c/ul\u003e\n \u003c/td\u003e\n \u003ctd width=\"32.78145695364238%\" valign=\"top\"\u003e\n \u003cul\u003e\n \u003cli\u003eRobotReviewer\u0026nbsp;[26, 27, 36, 49, 62, 68, 75, 115]\u003c/li\u003e\n \u003c/ul\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"34.437086092715234%\" valign=\"top\"\u003e\n \u003cp\u003e\u003cstrong\u003eSynthesis/meta-analysis\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"32.78145695364238%\" valign=\"top\"\u003e\n \u003cp\u003e-\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"32.78145695364238%\" valign=\"top\"\u003e\n \u003cp\u003e-\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"34.437086092715234%\" valign=\"top\"\u003e\n \u003cp\u003e\u003cstrong\u003eCertainty of evidence assessment\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"32.78145695364238%\" valign=\"top\"\u003e\n \u003cp\u003e-\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"32.78145695364238%\" valign=\"top\"\u003e\n \u003cp\u003e-\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"34.437086092715234%\" valign=\"top\"\u003e\n \u003cp\u003e\u003cstrong\u003eReport preparation\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"32.78145695364238%\" valign=\"top\"\u003e\n \u003cp\u003e-\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"32.78145695364238%\" valign=\"top\"\u003e\n \u003cp\u003e-\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"34.437086092715234%\" valign=\"top\"\u003e\n \u003cp\u003e\u003cstrong\u003eCombination of abbreviated methods/tools\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"32.78145695364238%\" valign=\"top\"\u003e\n \u003cul\u003e\n \u003cli\u003eRapid review methods with multiple steps combined: abbreviated search/search limits\u0026nbsp;[25, 76, 100, 116], single-reviewer screening/data extraction\u0026nbsp;[25, 116], review of reviews\u0026nbsp;[78], abbreviated critical appraisal\u0026nbsp;[116]\u0026nbsp;\u003c/li\u003e\n \u003c/ul\u003e\n \u003c/td\u003e\n \u003ctd width=\"32.78145695364238%\" valign=\"top\"\u003e\n \u003cp\u003eRapid review tools with multiple steps combined: SRA-Deduplicator, EndNote, Polyglot Search Translator, RobotReviewer, SRA-Helper\u0026nbsp;[12]\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n\u003c/table\u003e\n\u003cp\u003e\u003cem\u003eAbbreviations:\u0026nbsp;\u003c/em\u003e\u003cem\u003eASReview: Automated Systematic Review, ASySd:\u003c/em\u003e\u003cem\u003eAutomated Systematic Synthesis and Data extraction, ChatGPT: Chat Generative Pre-trained Transformer (by OpenAI), EPPI: Evidence for Policy and Practice Information and Co-ordinating Centre, ExaCT: Extraction of Citations Tool, MeSH: Medical Subject Headings, PICO: Population, Intervention, Comparison, Outcome, PST: Polyglot Search Translator, PLUS: McMaster Premium LiteratUre Service, RCT: randomized controlled trial, SR: systematic review, SRA: Systematic Review Assistant, SWIFT: Sciome Workbench for Interactive Computer-Facilitated Text-mining\u003c/em\u003e\u003c/p\u003e\n\u003ch2\u003eMethods or tools for literature search\u0026nbsp;\u003c/h2\u003e\n\u003ch3\u003eSearch strategy and database search\u0026nbsp;\u003c/h3\u003e\n\u003cp\u003eFive tools (MeSH on Demand\u0026nbsp;[93, 98], PubReMiner\u0026nbsp;[93, 98], Polyglot Search Translator\u0026nbsp;[34], Risklick search platform\u0026nbsp;[56], and Yale MeSH Analyzer\u0026nbsp;[93, 98]) and three methods (abbreviated search strategies study type\u0026nbsp;[52, 72, 110], topic\u0026nbsp;[122], or search date\u0026nbsp;[42, 126]; citation-based searching\u0026nbsp;[28, 65, 66]; search restrictions for database\u0026nbsp;[53, 92, 104, 121]\u0026nbsp;and language (e.g. only English articles)\u0026nbsp;[37, 39, 60, 64, 81, 91, 99]) were evaluated in 24 studies\u0026nbsp;[28, 34, 37, 39, 42, 52, 53, 56, 60, 64-66, 72, 81, 91-93, 98, 99, 104, 110, 121, 122, 126]\u0026nbsp;to support devising search strategies and/or performing literature searches.\u003c/p\u003e\n\u003ch4\u003eTools for search strategies\u003c/h4\u003e\n\u003cp\u003eUsing text mining tools for search strategy development (MeSH on Demand,\u0026nbsp;PubReMiner, and YaleMeSH Analyzer) reduced time expenditure compared to manual searches, with tools saving over half the time required for manual searches (5 hours, standard deviation [SD=2] vs. 12 hours [SD=8])\u0026nbsp;[98].\u0026nbsp;Using supportive tools such as Polyglot Search Translator\u0026nbsp;[34],\u0026nbsp;MeSH on Demand,\u0026nbsp;PubReMiner, and YaleMeSH Analyzer\u0026nbsp;[93, 98]\u0026nbsp;was less sensitive\u0026nbsp;[98]\u0026nbsp;and showed a slightly reduced precision compared to manual searches (11% to 14% vs. 15%)\u0026nbsp;[93]. The Risklick search platform demonstrated a high precision for identifying clinical trials (96%) and COVID-19–related publications (94%)\u0026nbsp;[56].\u003c/p\u003e\n\u003cp\u003eUser ratings by the study authors indicated that PubReMiner and YaleMeSH Analyzer were considered “useful” or “extremely useful,” while MeSH on Demand received the rating “not very useful” on a 5-point Likert scale (from extremely useful to least useful)\u0026nbsp;[93].\u003c/p\u003e\n\u003ch4\u003eAbbreviated search strategies for study type, topic, or search date\u003c/h4\u003e\n\u003cp\u003eTwo studies evaluated an abbreviated search strategy (i.e., Cochrane highly sensitive search strategy)\u0026nbsp;[52]\u0026nbsp;and a brief RCT strategy\u0026nbsp;[110]\u0026nbsp;for identifying RCTs. Both achieved high sensitivity rates of 99.5%\u0026nbsp;[52]\u0026nbsp;and\u0026nbsp;94%\u0026nbsp;[110]\u0026nbsp;while reducing the number of records requiring screening by 16%\u0026nbsp;[110]. Although some RCTs were missed using abbreviated search strategies, there were no significant differences in the conclusions\u0026nbsp;[110].\u003c/p\u003e\n\u003cp\u003eOne study\u0026nbsp;[122]\u0026nbsp;assessed an abbreviated search strategy using only the generic drug name to identify drug-related RCTs, achieving high sensitivities in both MEDLINE (99%) and Embase (99.6%)\u0026nbsp;[122].\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eLee et al. (2012) evaluated 31 search filters for SRs and meta-analyses, with the health-evidence.ca Systematic Review search filter performing best, maintaining a high sensitivity while reducing the number of articles needing for screening (90% in MEDLINE, 88% in Embase, 90% in CINAHL)\u0026nbsp;[72].\u003c/p\u003e\n\u003cp\u003eFuruya-Kanamori et al.\u0026nbsp;[42]\u0026nbsp;and Xu et al.\u0026nbsp;[126]\u0026nbsp;investigated the impact of restricting search timeframes on effect estimates and found that limiting searches to the most recent 10 to 15 years resulted in minimal changes in effect estimates (\u0026lt;5%) while reducing workload by up to 45%\u0026nbsp;[42, 126]. Nevertheless, this approach missed 21% to 35% of the relevant studies\u0026nbsp;[42].\u003c/p\u003e\n\u003ch4\u003eCitation-based searching\u003c/h4\u003e\n\u003cp\u003eThree studies\u0026nbsp;[28, 65, 66]\u0026nbsp;assessed whether citation-based searching can improve efficiency in systematic reviewing. Citation-based searching achieved a reduction in the number of retrieved articles (50% to 89% fewer articles) compared to the original searches while still capturing a substantial proportion of 75% to 82% of the included articles\u0026nbsp;[28, 66].\u003c/p\u003e\n\u003ch4\u003eRestricted database searching\u003c/h4\u003e\n\u003cp\u003eSeven studies assessed the validity of restricted database searching and suggested that searching at least two topic-related databases yielded high recall and precision for various types of studies\u0026nbsp;[29, 40, 41, 53, 92, 104, 121].\u003c/p\u003e\n\u003cp\u003ePreston et al. (2015) demonstrated that searching only MEDLINE and Embase plus reference list checking identified 93% of the relevant references while saving 24% of the workload\u0026nbsp;[104]. Beyer et al. (2013) emphasized the necessity of searching at least two databases along with reference list checking to retrieve all included studies\u0026nbsp;[29]. Goossen et al. (2018) highlighted that combining MEDLINE with CENTRAL and hand searching was the most effective for RCTs (Recall: 99%), while for nonrandomized studies, combining MEDLINE with Web of Science yielded the highest recall (99.5%)\u0026nbsp;[53]. Ewald et al. (2022) showed that searching two or more databases (MEDLINE/CENTRAL/Embase) reached a recall of ≥87.9% for identifying mainly RCTs\u0026nbsp;[41]. Additionally, Van Enst et al. (2014) indicated that restricting searches to MEDLINE alone might slightly overestimate the results compared to broader database searches in diagnostic accuracy SRs (relative diagnostic odds ratio: 1.04; 95% confidence interval [CI], 0.95 to 1.15)\u0026nbsp;[121]. Nussbaumer-Streit et al. (2018) and Ewald et al. (2020) found that combining one database with another or with searches of reference lists was noninferior to comprehensive searches (2%; 95% CI, 0% to 9%; if opposite concusion was of concern)\u0026nbsp;[92]\u0026nbsp;as the effect estimates were similar (ratio of odds ratios [ROR] median: 1.0 interquartile range [IQR]: 1.0–1.01)\u0026nbsp;[40].\u003c/p\u003e\n\u003ch4\u003eRestricted language searching\u003c/h4\u003e\n\u003cp\u003eSeven studies found that excluding non-English articles to reduce workload would minimally alter the conclusions or effect estimates of the meta-analyses. Two studies found no change in the overall conclusions\u0026nbsp;[60, 91], and five studies\u0026nbsp;[37, 60, 64, 91, 99]\u0026nbsp;reported changes in the effect estimates or statistical significance of the meta-analyses. Specifically, the statistical significance of the effect estimates changed in 3% to 12% of the meta-analyses[37, 60, 64, 91, 99].\u0026nbsp;\u003c/p\u003e\n\u003ch3\u003eDeduplication\u003c/h3\u003e\n\u003cp\u003eSix studies\u0026nbsp;[30, 36, 54, 57, 71, 80]\u0026nbsp;compared eleven supportive software tools (ASySD, EBSCO, EndNote, Covidence, Deduklick, Mendeley, OVID, Rayyan, RefWorks, Systematic Review Accelerator, Zotero). Manual deduplication took approximately 4 hours 45 minutes, whereas using the tools reduced the time by 4 hours 42 minutes to only 3 minutes\u0026nbsp;[71]. False negative duplicates varied from 36 (Mendeley) to 258 (EndNote), while false positives ranged from 0 (OVID) to 43 (EBSCO)\u0026nbsp;[71]. The precision was high with 99% to 100% for Deduklick and ASySD, and the sensitivity was highest for Rayyan (ranging from 99% to 100%)\u0026nbsp;[54, 57, 80], followed by Covidence, OVID, Systematic Review Accelerator, Mendeley, EndNote, and Zotero\u0026nbsp;[54, 57, 80]. However, Cowie et al. reported that the Systematic Review Accelerator received a low rating of 9/30 for its features and usability\u0026nbsp;[36].\u003c/p\u003e\n\u003ch3\u003eAdditional literature search\u003c/h3\u003e\n\u003cp\u003ePaperfetcher, identified as an application to automate additional searches as handsearching and citation searching, saved up to 92.0% of the time compared to manual handsearching and reference list checking, though validity outcomes for Paperfetcher were not reported\u0026nbsp;[97]. Additionally, the Scopus approach, in which reviewers electronically downloaded the reference lists of relevant articles and screened only new references dually, saved approximately 62.5% of the time compared to manual checking\u0026nbsp;[33].\u003c/p\u003e\n\u003ch3\u003eUpdate literature search\u003c/h3\u003e\n\u003cp\u003eWe identified one tool (RobotReviewer LIVE)\u0026nbsp;[36, 77]\u0026nbsp;and five methods (Clinical Query search combined with PubMed-related articles search, Clinical Query search in MEDLINE and Embase, searching the McMaster Premium LiteratUre Service [PLUS], PubMed similar articles search, and Scopus citation tracking)\u0026nbsp;[50, 61, 109, 111]\u0026nbsp;for improving the efficiency of updating literature searches. RobotReviewer LIVE showed a precision of 55% and a high recall of 100%\u0026nbsp;[77]\u0026nbsp;with limitations including search restricted to MEDLINE, consideration of only RCTs, and low usability scores for features\u0026nbsp;[36, 77].\u003c/p\u003e\n\u003cp\u003eThe Clinical Query (CQ) search, combined with the PubMed-related articles search and the CQ search in MEDLINE and Embase, exhibited high recall rates ranging from 84% to 91%\u0026nbsp;[61, 109, 111], while the PLUS database had a lower recall rate of 23%\u0026nbsp;[61]. The PubMed similar articles search and Scopus citation tracking had a low sensitivity of 25% each, with time-saving percentages of 24% and 58%, respectively\u0026nbsp;[50]. However, the omission of studies from searching the PLUS database only did not significantly change the effect estimates in most reviews (ROR: 0.99; 95% CI, 0.87 to 1.14)\u0026nbsp;[61].\u003c/p\u003e\n\u003ch2\u003eMethods or tools for study selection\u003c/h2\u003e\n\u003ch3\u003eTitle and abstract selection\u003c/h3\u003e\n\u003cp\u003eWe identified 42 studies evaluating 14 supportive software tools (AbstrackR, ASReview, ChatGPT, Colandr, Covidence, DistillerSR, EPPI-reviewer, Rayyan, RCT classifier, Research screener, RobotAnalyst, SRA-Helper for EndNote, SWIFT-active screener, SWIFT-review)\u0026nbsp;[32, 35, 36, 38, 43-45, 47, 48, 51, 58, 59, 63, 73, 83, 94-96, 102, 105, 107, 108, 112, 117-120, 127]\u0026nbsp;using advanced text mining and machine and active learning techniques, and five methods (crowdsourcing using different [automation] tools, dual computer monitors, single-reviewer screening, PICO-based title-only screening, limited screening [review of reviews])\u0026nbsp;[55, 79, 82, 84, 86-90, 100, 106, 108, 113, 114, 125]\u0026nbsp;for improving the title and abstract screening efficiency. The tested datasets ranged from 1 to 60 SRs and 148 to 190,555 records.\u003c/p\u003e\n\u003ch4\u003eTools for title and abstract selection\u003c/h4\u003e\n\u003cp\u003eVarious tools (e.g., EPPI-Reviewer, Covidence, DistillerSR, and Rayyan) offer collaborative online platforms for SRs, enhancing efficiency by managing and distributing screening tasks, facilitating multiuser screening, and tracking records throughout the review process\u0026nbsp;[128].\u003c/p\u003e\n\u003cp\u003eIn a semiautomated tool, the tool provide suggestions or probabilities regarding the eligibility of a reference for inclusion in the review, but human judgment is still required to make the final decision\u0026nbsp;[9, 12]. In contrast, in a fully automated system, the tool makes the final decision without human intervention based on predetermined criteria or algorithms. Some tools provide fully automated screening options (e.g., DistillerSR), semiautomated (e.g., RobotAnalyst), or both (e.g., AbstrackR, DistillerAI) using machine learning or natural language processing methods\u0026nbsp;[9, 12]\u0026nbsp;(see\u0026nbsp;Table 1).\u003c/p\u003e\n\u003cp\u003eAmong the eleven semi- and fully -automated tools (AbstrackR\u0026nbsp;[36, 38, 44, 45, 47, 48, 51, 59, 73, 107, 108, 118], ASReview\u0026nbsp;[83, 95, 102, 120], ChatGPT\u0026nbsp;[112], Colandr\u0026nbsp;[38], DistillerSR\u0026nbsp;[36, 43, 47, 58], EPPI-reviewer\u0026nbsp;[36, 59, 118, 127], Rayyan\u0026nbsp;[35, 36, 38, 73, 94, 96, 119, 127], RCT classifier\u0026nbsp;[117], Research screener\u0026nbsp;[32], RobotAnalyst\u0026nbsp;[35, 36, 47, 105, 108], SRA-helper for EndNote\u0026nbsp;[35], SWIFT-active screener\u0026nbsp;[36, 63], SWIFT-review\u0026nbsp;[73]). ASReview\u0026nbsp;[83, 95, 102, 120]\u0026nbsp;and Research Screener\u0026nbsp;[32]\u0026nbsp;demonstrated a robust performance, identifying 95% of the relevant studies while saving 37% to 92% of the workload. SWIFT-active screener[36, 63], RobotAnalyst\u0026nbsp;[35, 36, 47, 105, 108], and Rayyan\u0026nbsp;[35, 36, 94, 96, 119]\u0026nbsp;also performed well, identifying 95% of the relevant studies with workload savings from 34% to 49%. EPPI-Reviewer identified all the relevant abstracts after screening 40% to 99% of the references across reviews\u0026nbsp;[118]. DistillerAI showed substantial workload savings of 53% while identifying 95% of the relevant studies\u0026nbsp;[58], with varying degrees of validity\u0026nbsp;[43, 47, 58]. Colandr and SWIFT-Review exhibited sensitivity rates of 65% and 91%, respectively, with 97% workload savings and around 400 minutes of time saved\u0026nbsp;[38, 73]. ChatGPT’s sensitivity was 100%\u0026nbsp;[112]\u0026nbsp;and RCT classifiers recall was 99%\u0026nbsp;[117]; workload or time savings were not reported\u0026nbsp;[112, 117]. AbstrackR showed a moderate performance with a potential workload savings from 4% up to 97%\u0026nbsp;[38, 44, 45, 47, 48, 51, 107, 108, 118]\u0026nbsp;while missing up to 44% of the relevant studies\u0026nbsp;[38, 45, 48, 51, 108]. Covidence and SRA-Helper for EndNote did not report validity outcomes.\u003c/p\u003e\n\u003cp\u003eMost of the supportive software tools were easy to use or learn, suitable for collaboration, and straightforward for inexperienced users (ASReview, AbstarckR, Covidence, SRA-Helper, Rayyan, RobotAnalyst)\u0026nbsp;[35, 36, 44, 59, 94, 96, 120]. Other tools were more complex regarding their usability but were useful for large and complex projects (DistillerSR, EPPI-Reviewer)\u0026nbsp;[36, 47, 59]. Poor interface quality (AbstrackR)\u0026nbsp;[59], issues with help section/response time (RobotAnalyst, Covidence, EPPI)\u0026nbsp;[35, 59], and overloaded side panel (Rayyan)\u0026nbsp;[35]\u0026nbsp;were weaknesses reported in the studies.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cem\u003eMethods for title and abstract selection\u003c/em\u003e\u003c/p\u003e\n\u003cp\u003eAmong the four methods identified (dual computer monitors, single-reviewer screening, crowdsourcing using different [automation] tools, and limited screening [review of reviews, PICO-based title-only screening, title-first screening])\u0026nbsp;[55, 79, 82, 84, 86-90, 100, 113, 114, 125]\u0026nbsp;for supporting title and abstract screening, crowdsourcing in combination with screening platforms or machine learning tools demonstrated the most promising performance in improving efficiency. Studies by Noel-Storr et al.\u0026nbsp;[82, 87-90]\u0026nbsp;found that the Cochrane Crowd plus Screen4Me/RCT classifier achieved a high sensitivity ranging from 84% to 100% in abstract screening and reduced screening time. Crowdsourcing via Amazon Mechanical Turk yielded correct inclusions of 95% to 99% with a substantial cost reduction of 82%\u0026nbsp;[82]. However, the sensitivity was moderate when the screening was conducted manually by medical students or on web-based platforms (47% to 67%)\u0026nbsp;[86].\u003c/p\u003e\n\u003cp\u003eSingle-reviewer screening missed 0% to 19% of the relevant studies\u0026nbsp;[55, 100, 113, 114]\u0026nbsp;while saving 50% to 58% of the time and costs, respectively\u0026nbsp;[113]. The findings indicate that single-reviewer screening by less-experienced reviewers could substantially alter the results, whereas experienced reviewers had a negligible impact\u0026nbsp;[100].\u003c/p\u003e\n\u003cp\u003eLimited screening methods, such as reviews of reviews (also known as umbrella reviews), exhibited a moderate sensitivity (56%) and significantly reduced the number of citations needed to screen\u0026nbsp;[108]. Title-first screening and PICO-based title screening demonstrated an accurate validity, with a recall of 100%\u0026nbsp;[79, 107]\u0026nbsp;and a reduction in screening effort ranging from 11% to 78%\u0026nbsp;[107]. However, screening with dual computer monitors did not notably improve the time saved\u0026nbsp;[125].\u003c/p\u003e\n\u003ch3\u003eFull-text selection\u003c/h3\u003e\n\u003cp\u003eFor full-text screening, we identified three methods: crowdsourcing\u0026nbsp;[85], using dual computer monitors\u0026nbsp;[125], and single-reviewer screening\u0026nbsp;[113, 114]. Using crowdsourcing in combination with the CrowdScreenSR saved 68% of the workload\u0026nbsp;[85]. With dual computer monitors, no significant difference of time taken for the full-text screening was reported\u0026nbsp;[125]. Single-reviewer screening missed 7% to 12% of the relevant studies\u0026nbsp;[114]\u0026nbsp;while saving only 4% of the time and costs\u0026nbsp;[113].\u0026nbsp;\u003c/p\u003e\n\u003ch2\u003eMethods\u0026nbsp;or tools for data extraction\u0026nbsp;\u003c/h2\u003e\n\u003cp\u003eWe identified 11 studies evaluating five tools (ChatGPT\u0026nbsp;[112], Data Abstraction Assistant (DAA)\u0026nbsp;[36, 67, 74], Dextr\u0026nbsp;[124], ExaCT\u0026nbsp;[46, 70, 103]\u0026nbsp;and Plot Digitizer\u0026nbsp;[69]) and two methods (dual computer monitors\u0026nbsp;[125]\u0026nbsp;and single data extraction\u0026nbsp;[31, 74]) to expedite the data extraction process.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eExaCT\u0026nbsp;[46, 70, 103], DAA\u0026nbsp;[36, 67, 74], Dextr\u0026nbsp;[124], and Plot Digitizer\u0026nbsp;[69]\u0026nbsp;achieved a time reduction of up to 60%\u0026nbsp;[103, 124], with precision rates of 93% for ExaCT\u0026nbsp;[70, 103]\u0026nbsp;and 96% for Dextr and an error rate of 17% for DAA\u0026nbsp;[67, 74]. Manual extraction by two reviewers and with the assistance of Plot Digitizer showed a similar agreement with the original data, showing a slightly higher agreement with the assistance of Plot Digitizer (Plot Digitizer: 73% and 75%, manual extraction: 66% and 69%)\u0026nbsp;[69]. A total of 87% of manually extracted data elements matched with ExaCT, resulting in qualitatively altered meta-analysis results\u0026nbsp;[103]. ChatGPT demonstrated consistent agreement with human researchers across various parameters (κ=0.79-1) extracted from studies, such as language, targeted disease, natural language processing model, sample size, and performance parameters, and moderate to fair agreement for clinical task (κ=0.58) and clinical implementation (κ=0.34)\u0026nbsp;[112]. Usability was assessed only for DAA and Dextr, with both tools deemed very easy to use\u0026nbsp;[67, 124], although DAA scored lower on feature scores, while Dextr was noted for its flexible interface\u0026nbsp;[124].\u003c/p\u003e\n\u003cp\u003eSingle data extraction and dual monitors reduced the time for extracting data by 24 to 65 minutes per article\u0026nbsp;[31, 74, 125], with similar error rates between single and dual data extraction methods (single: 16%\u0026nbsp;[74], 18%\u0026nbsp;[31], dual: 15%\u0026nbsp;[31, 74]) and comparable pooled estimates\u0026nbsp;[31, 74].\u0026nbsp;\u003c/p\u003e\n\u003ch2\u003eMethods or tools for critical appraisal\u0026nbsp;\u003c/h2\u003e\n\u003cp\u003eWe identified nine studies reporting on one software tool (RobotReviewer)\u0026nbsp;[26, 27, 36, 49, 62, 68, 75, 115]\u0026nbsp;and one method (crowdsourcing via CrowdCARE)\u0026nbsp;[101]\u0026nbsp;for improving critical appraisal efficiency. Collectively, the study authors suggested that RobotReviewer can support but not replace RoB assessments by humans\u0026nbsp;[26, 27, 49, 62, 68, 75]\u0026nbsp;as performance varied per RoB domain\u0026nbsp;[26, 49, 62, 115]. The authors reported similar completion times for RoB appraisal with and without RobotReviewer assistance\u0026nbsp;[27]. Reviewers were equally as likely to accept RobotReviewer’s judgments as one another’s during consensus (83% for RobotReviewer, 81% for humans)\u0026nbsp;[68], showing similar accuracy (RobotReviewer assisted RoB appraisal: 71%, RoB appraisal by two reviewers: 78%)\u0026nbsp;[27, 75]. The reviewers generally described the tool as acceptable and useful\u0026nbsp;[68], whereby collaboration with other users is not possible\u0026nbsp;[36].\u0026nbsp;\u003c/p\u003e\n\u003ch2\u003eCombination of abbreviated methods/tools\u0026nbsp;\u003c/h2\u003e\n\u003cp\u003eFive studies evaluated RR methods [25, 76, 78, 100, 116], and one study evaluated various tools [12] combining multiple review steps. While two case studies found no differences in findings between RR and SR approaches [78, 116], another author/paper/study found in two of three RRs that no conclusion could be drawn due to insufficient information [25]. Additionally, in a study including three RRs, RR methods affected almost one-third of the meta-analyses with less precise pooled estimates [100]. Marshall et al. (2019) included 2,512 SRs and reported a loss of all data in 4% to 45% of the meta-analyses and changes of 7% to 39% in the statistical significance due to RR methods [76]. Automation tools (SRA- Deduplicator, EndNote, Polyglot Search Translator, RobotReviewer, SRA-Helper) reduced the person-time spent on SR tasks (42 hours versus 12 hours) [12]. However, error rates, falsely excluded studies, and sensitivity varied immensely across studies [25, 116].\u0026nbsp;\u003c/p\u003e"},{"header":"Discussion","content":"\u003cp\u003eTo the best of our knowledge, this is the first scoping review to map evaluated methods and tools for improving the efficiency of SR production across the various review steps. We describe which review steps methods and\u0026nbsp;ready-to-use\u0026nbsp;tools are available and have been evaluated.\u0026nbsp;Additionally, we provide an overview of the contexts in which these methods and tools were evaluated, such as real-time workflow testing and the use of internal or external data.\u003c/p\u003e\n\u003cp\u003eAcross all the SR review steps, most studies evaluated study selection, followed by literature searching and data extraction. Around half of the studies evaluated tools and half of the studies methods. For study selection, most of the tools offered semiautomated-assisted screening by classifying or ranking references. The methods focused mainly on limiting the review team’s human resources, for example, through single-reviewer screening or by distributing the tasks to a crowd or students.\u003c/p\u003e\n\u003cp\u003eTwo scoping reviews, one on tools\u0026nbsp;[10]\u0026nbsp;and one on methods\u0026nbsp;[15]\u0026nbsp;to support the SR process, are in line with this result, as the authors also identified these tasks as the most frequently evaluated in the literature\u0026nbsp;[10, 15]. As shown by our scoping review and others\u0026nbsp;[10], a major focus on (semi) automation tools for study selection occurred in recent years. This is important, as a recent study on resource use found that study selection and data extraction are the most resource-intensive tasks besides administration and project management as well as critical appraisal\u0026nbsp;[2].\u003c/p\u003e\n\u003cp\u003eFor the following tasks, we could not identify a single study evaluating a tool or method:\u0026nbsp;administration/project management, formulating the review question, writing the protocol, searching for existing reviews, full-text retrieval, synthesis/meta-analysis, certainty of evidence assessment, and report preparation. However, all of these tasks are also time-consuming, especially according to the study by Nussbaumer-Streit et al\u0026nbsp;[2]. Project management requires the largest proportion of SR production time\u0026nbsp;[2]. To our knowledge, tools supporting project management are already available, such as Covidence, DistillerSR, EPPI-Reviewer, or simple online platforms such as Google Forms, which can also support managing and coordinating projects. However, no evaluations of these support platforms were found. Similarly, while innovative software tools such as large language models (e.g., ChatGPT) show promise in supporting tasks such as report preparation, there is a lack of formal evaluation in this context as well. This is relevant for future research that aims to improve SR production since these tasks are extremely resource-intensive.\u003c/p\u003e\n\u003cp\u003eOur scoping review identified several research gaps. There is a lack of studies evaluating the usability of tools and methods. No study evaluated the usability of any single method. Only for the study selection task did we identify multiple studies evaluating the usability of tools. However, an important factor in adopting tools and methods is their user-friendliness\u0026nbsp;[14, 129]\u0026nbsp;and their fit with standard SR workflows\u0026nbsp;[9, 47]. Furthermore, if usability was considered, this was often evaluated in a nonformal or standardized way employing various scales, questions, and feedback mechanisms. To enable meaningful comparisons between different methods, there is a clear need for a formal analysis of user experience and usability. Therefore, authors and review teams would benefit from comparable usability studies on methods and tools that aim to improve the SR process’s efficiency.\u003c/p\u003e\n\u003cp\u003eFew studies exist that evaluate the impact on results when using accelerated methods or tools. We identified 25 studies (13%) where the odds ratio changed by 0% to 63% depending on the method or tool\u0026nbsp;[61, 91, 121, 130-133]. Marshall et al. (2019) stated that there has not been a large-scale evaluation of the effects of RR methods on the number of falsely excluded studies and the consequent changes in meta-analysis results\u0026nbsp;[134]. Indeed, understanding the potential impact of different methods and tools on the results is fundamental, as emphasized by Wagner et al. (2016)\u0026nbsp;[135]. Wagner et al. conducted an online survey of guideline developers and policymakers, aiming to discover how much incremental uncertainty about the correctness of an answer would be acceptable in RRs. They found that participants demanded very high levels of accuracy and would tolerate a median 10% risk of wrong answers\u0026nbsp;[135]. Therefore, studies focusing on the impact on results and conclusions through RR methods and tools are warranted.\u003c/p\u003e\n\u003cp\u003eThe majority of studies retrospectively evaluated only a single tool or method using existing internal data, offering limited insights into real-world adoption. Prospective studies within a real-time workflow study (n=20)\u0026nbsp;[12, 27, 32, 35, 46, 50, 67, 68, 77, 78, 88, 90, 94, 98, 105, 108, 112, 114, 125, 127], comparing several tools and methods (n=5)\u0026nbsp;[12, 35, 50, 98, 108, 127]\u0026nbsp;or from independent reviewer teams using their own dataset (n=7)\u0026nbsp;[12, 35, 46, 94, 98, 112, 127], are scarce. However, such studies are crucial for providing valid comparative evidence on validity, workload savings, usability, impact on results, and real-world benefits. Particularly for automated title and abstract screening, where most tools function similarly, key information such as the stopping rule (indicating when screening can cease) is essential. Notably, there is limited research (n=8)\u0026nbsp;[51, 63, 94, 95, 107, 108, 119, 127]\u0026nbsp;exploring the combined effects of algorithmic re-ranking and stopping criterion determination in automated title and abstract screening. Furthermore, studies assessing the influence of automated tools (e.g., re-ranking algorithms) on human decision-making are lacking. Therefore, simulation and prospective real-time studies evaluating the workflow between manual procedures and tools with stopping criterion are warranted.\u003c/p\u003e\n\u003cp\u003eThe balance between the time saved and the potential reduction in quality and comprehensiveness is influenced by various factors, including the decision-making urgency, resource availability, and the decision-makers’ specific needs. However, accelerated approaches are not universally appropriate. In situations where the thoroughness and rigor of the evidence are paramount—such as in developing clinical guidelines, conducting health technology assessments, or addressing areas with significant scientific uncertainty—the risk of missing critical evidence or drawing inaccurate conclusions outweighs the benefits of speed.\u003c/p\u003e\n\u003cp\u003eGiven the heterogeneity in study designs and contexts, there is a pressing need for standardized frameworks for the evaluation and reporting of tools and methods for SR production. Furthermore, researchers should be aware of the importance of testing methods and tools with their own datasets and contextual factors. Pretraining both the tools and the crowd before implementation is essential for optimizing efficiency and ensuring reliable outcomes. However, we believe our findings are generalizable to other types of evidence syntheses, such as scoping reviews, reviews of reviews, or RRs. Additionally, we highlight that a part of the Rapid Review Methods Series\u0026nbsp;[128]\u0026nbsp;offers guidance on the use of supportive tools, aiming to assist researchers in effectively navigating the complexities of SR and RR production.\u003c/p\u003e\n\u003cp\u003eOur scoping review has several limitations. First, our inclusion criteria focused solely on studies mentioning efficiency improvements. While this criterion aimed to strike a balance between screening workload and sensitivity, it may have inadvertently excluded relevant studies that did not explicitly highlight efficiency gains. We might have missed relevant studies. Second, the heterogeneity among the included studies poses challenges in generalizing the findings to other review teams and contexts. The narrow focus of many studies, along with their publication primarily in English and focus on specific study types, further limits the generalizability of our findings. Moreover, the limited proportion of studies (29%, 30/103) comparing different tools and methods on the same dataset within the same study accentuates the need for caution while interpreting the findings. However, we think the validity and usability outcomes reported in this scoping review provide a good orientation.\u003c/p\u003e"},{"header":"Conclusion","content":"\u003cp\u003eBased on the identified evidence, various methods and tools for literature searching and title and abstract screening are available with the aim of improving efficiency. However, only few studies have addressed the influence of these methods and tools in real-world workflows. Fewer studies evaluated methods or tools supporting the other tasks of SR production. Moreover, the reporting of the outcomes of existing evaluations varies considerably, revealing significant research gaps, especially in assessing usability and impact on results. Future research should prioritize addressing these gaps, evaluating real-world adoption, and establishing standardized frameworks for the evaluation of methods and tools to enhance the overall effectiveness of SR development processes.\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u003cem\u003eEthics approval and consent to participate:\u003c/em\u003e not applicable.\u003c/p\u003e\n\u003cp\u003e\u003cem\u003eConsent for publication:\u003c/em\u003e not applicable.\u003c/p\u003e\n\u003cp\u003e\u003cem\u003eAvailability of data and materials:\u0026nbsp;\u003c/em\u003eNo/Not applicable (this manuscript does not report data generation or analysis).\u003c/p\u003e\n\u003cp\u003e\u003cem\u003eCompeting interests:\u003c/em\u003e The authors declare that they have no competing interests.\u003c/p\u003e\n\u003cp\u003e\u003cem\u003eFunding:\u003c/em\u003e This work was supported by funds from the EU funded COST Action EVBRES (CA17117), two COST STSM fundings (LA, MM), and by a scholarship for the first author (LA) of the Gesellschaft f\u0026uuml;r Forschungsf\u0026ouml;rderung Nieder\u0026ouml;sterreich m.b.H.\u003c/p\u003e\n\u003cp\u003e\u003cem\u003eAuthors\u0026rsquo; contributions:\u003c/em\u003e LA, BNS, RS, and LH developed the study concept. LA wrote the protocol. RS conducted the literature searches and similar article searches. LA, MMM, IS, BNS, MMK, MEM, EB, ME, RS, PNL, GP, NR, KG, LK, MS, AMP, and PM screened the references and extracted the data. LA and MMM conducted grey literature searches, reference list checking, and hand searches. RS provided methodological input throughout the study. LA and MMM created the first draft of the manuscript, which all other authors critically revised. The authors read and approved the final version of the submitted manuscript.\u003c/p\u003e\n\u003cp\u003e\u003cem\u003eAcknowledgments:\u003c/em\u003e We would like to thank Hans Lund and Jos Kleijnen for their input on the manuscript.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eTsafnat G, Glasziou P, Choong MK, Dunn A, Galgani F, Coiera E. Systematic review automation technologies. Syst Rev. 2014;3:74.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eNussbaumer-Streit B, Ellen M, Klerings I, Sfetcu R, Riva N, Mahmić-Kaknjo M, Poulentzas G, Martinez P, Baladia E, Ziganshina LE et al. Resource use during systematic review production varies widely: a scoping review. J Clin Epidemiol 2021.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eOliver S, Dickson K, Bangpan M. Systematic reviews: making them policy relevant. A briefing for policy makers and systematic reviewers. \u003cem\u003eEPPI-Centre, Social Science Research Unit, UCL Institute of Education, University College London, London\u003c/em\u003e 2015.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBorah R, Brown AW, Capers PL, Kaiser KA. Analysis of the time and workers needed to conduct systematic reviews of medical interventions using data from the PROSPERO registry. BMJ Open. 2017;7(2):e012545\u0026ndash;012545.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDonnelly CA, Boyd I, Campbell P, Craig C, Vallance P, Walport M, Whitty CJM, Woods E, Wormald C. Four principles to make evidence synthesis more useful for policy. Nature. 2018;558(7710):361\u0026ndash;4.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eClayton GL, Smith IL, Higgins JPT, Mihaylova B, Thorpe B, Cicero R, Lokuge K, Forman JR, Tierney JF, White IR, et al. The INVEST project: investigating the use of evidence synthesis in the design and analysis of clinical trials. Trials. 2017;18(1):219.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePeters MDJ, Marnie C, Tricco AC, Pollock D, Munn Z, Alexander L, McInerney P, Godfrey CM, Khalil H. Updated methodological guidance for the conduct of scoping reviews. JBI Evid Synthesis 2020, 18(10).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBeller E, Clark J, Tsafnat G, Adams C, Diehl H, Lund H, Ouzzani M, Thayer K, Thomas J, Turner T. Making progress with the automation of systematic reviews: principles of the International Collaboration for the Automation of Systematic Reviews (ICASR). Syst reviews. 2018;7(1):77.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eO\u0026rsquo;Connor AM, Tsafnat G, Thomas J, Glasziou P, Gilbert SB, Hutton B. A question of trust: can we build an evidence base to gain trust in systematic review automation technologies? Syst reviews. 2019;8(1):1\u0026ndash;8.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKhalil H, Ameen D, Zarnegar A. Tools to support the automation of systematic reviews: a scoping review. J Clin Epidemiol. 2022;144:22\u0026ndash;42.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKhurana D, Koli A, Khatter K, Singh S. Natural language processing: state of the art, current trends and challenges. Multimedia Tools Appl. 2023;82(3):3713\u0026ndash;44.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eClark J, McFarlane C, Cleo G, Ishikawa Ramos C, Marshall S. The Impact of Systematic Review Automation Tools on Methodological Quality and Time Taken to Complete Systematic Review Tasks: Case Study. JMIR Med Educ. 2021;7(2):e24418.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eErgonomics of human-system interaction - Part 11. Usability: Definitions and concepts, (ISO 9241-11: 2018). ISO 9241-11:2018 - Ergonomics of human-system interaction \u0026mdash; Part 11: Usability: Definitions and concepts.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eScott AM, Forbes C, Clark J, Carter M, Glasziou P, Munn Z. Systematic review automation tools improve efficiency but lack of knowledge impedes their adoption: a survey. J Clin Epidemiol. 2021;138:80\u0026ndash;94.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHamel C, Michaud A, Thuku M, Affengruber L, Skidmore B, Nussbaumer-Streit B, Stevens A, Garritty C. Few evaluative studies exist examining rapid review methodology across stages of conduct: a systematic scoping review. J Clin Epidemiol. 2020;126:131\u0026ndash;40.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eArksey H, O'Malley L. Scoping studies: towards a methodological framework. Int J Soc Res Methodol. 2005;8(1):19\u0026ndash;32.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLevac D, Colquhoun H, O'Brien KK. Scoping studies: advancing the methodology. Implement Sci. 2010;5(1):69.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eColquhoun HL, Levac D, O'Brien KK, Straus S, Tricco AC, Perrier L, Kastner M, Moher D. Scoping reviews: time for clarity in definition, methods, and reporting. J Clin Epidemiol. 2014;67(12):1291\u0026ndash;4.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTricco AC, Lillie E, Zarin W, O'Brien KK, Colquhoun H, Levac D, Moher D, Peters MD, Horsley T, Weeks L. PRISMA extension for scoping reviews (PRISMA-ScR): checklist and explanation. Ann Intern Med. 2018;169(7):467\u0026ndash;73.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eJBI Reviewer's Manual. Chapter 11.2.5. Search Strategy.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMcGowan J, Sampson M, Salzwedel DM, Cogo E, Foerster V, Lefebvre C. PRESS peer review of electronic search strategies: 2015 guideline statement. J Clin Epidemiol. 2016;75:40\u0026ndash;6.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBest L, Stevens A, Colin-Jones D. Rapid and responsive health technology assessment: the development and evaluation process in the South and West region of England. J Clin Eff 1997.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eJonnalagadda S, Petitti D. A new iterative method to reduce workload in systematic review process. Int J Comput Biol Drug Des. 2013;6(1\u0026ndash;2):5\u0026ndash;17.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHiggins JPTTJ, Chandler J, Cumpston M, Li T, Page MJ, Welch VA, editors. Cochrane Handbook for Systematic Reviews of Interventions version 6.2 (updated February 2021); 2021.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAffengruber L, Wagner G, Waffenschmidt S, Lhachimi SK, Nussbaumer-Streit B, Thaler K, Griebler U, Klerings I, Gartlehner G. Combining abbreviated literature searches with single-reviewer screening: three case studies of rapid reviews. Syst Rev. 2020;9(1):162.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eArmijo-Olivo S, Craig R, Campbell S. Comparing machine and human reviewers to evaluate the risk of bias in randomized controlled trials. Res. 2020;11(3):484\u0026ndash;93.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eArno A, Thomas J, Wallace B, Marshall IJ, McKenzie JE, Elliott JH. Accuracy and Efficiency of Machine Learning-Assisted Risk-of-Bias Assessments in Real-World Systematic Reviews: A Noninferiority Randomized Controlled Trial. Ann Intern Med. 2022;175(7):1001\u0026ndash;9.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBelter CW. Citation analysis as a literature search method for systematic reviews. J Assoc Soc Inf Sci Technol. 2016;67(11):2766\u0026ndash;77.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBeyer FR, Wright K. Can we prioritise which databases to search? A case study using a systematic review of frozen shoulder management. Health Info Libr J. 2013;30(1):49\u0026ndash;58.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBorissov N, Haas Q, Minder B, Kopp-Heim D, von Gernler M, Janka H, Teodoro D, Amini P. Reducing systematic review burden using Deduklick: a novel, automated, reliable, and explainable deduplication algorithm to foster medical research. Syst. 2022;11(1):172.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBuscemi N, Hartling L, Vandermeer B, Tjosvold L, Klassen TP. Single data extraction generated more errors than double data extraction in systematic reviews. J Clin Epidemiol. 2006;59(7):697\u0026ndash;703.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChai KEK, Lines RLJ, Gucciardi DF, Ng L. Research Screener: a machine learning tool to semi-automate abstract screening for systematic reviews. Syst Reviews 2021, 10(1).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChapman AL, Morgan LC, Gartlehner G. Semi-automating the manual literature search for systematic reviews increases efficiency. Health Info Libr J. 2010;27(1):22\u0026ndash;7.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eClark JM, Sanders S, Carter M, Honeyman D, Cleo G, Auld Y, Booth D, Condron P, Dalais C, Bateup S, et al. Improving the translation of search strategies using the Polyglot Search Translator: a randomized controlled trial. J Med Libr Assoc. 2020;108(2):195\u0026ndash;207.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCleo G, Scott AM, Islam F, Julien B, Beller E. Usability and acceptability of four systematic review automation software packages: a mixed method design. Syst. 2019;8(1):145.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCowie K, Rahmatullah A, Hardy N, Holub K, Kallmes K. Web-Based Software Tools for Systematic Literature Review in Medicine: Systematic Search and Feature Analysis. JMIR Med Inf. 2022;10(5):e33219.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDechartres A, Atal I, Riveros C, Meerpohl J, Ravaud P. Association between publication characteristics and treatment effect estimates: a meta-epidemiologic study. Ann Intern Med. 2018;169(6):385\u0026ndash;93.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003edos Reis AHS, de Oliveira ALM, Fritsch C, Zouch J, Ferreira P, Polese JC. Usefulness of machine learning softwares to screen titles of systematic reviews: a methodological study. Syst Reviews. 2023;12(1):68.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eEgger M, Juni P, Bartlett C, Holenstein F, Sterne J. How important are comprehensive literature searches and the assessment of trial quality in systematic reviews? Empirical study. Health Technol Assess. 2003;7(1):1\u0026ndash;76.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eEwald H, Klerings I, Wagner G, Heise TL, Dobrescu AI, Armijo-Olivo S, Stratil JM, Lhachimi SK, Mittermayr T, Gartlehner G, et al. Abbreviated and comprehensive literature searches led to identical or very similar effect estimates: a meta-epidemiological study. J Clin Epidemiol. 2020;128:1\u0026ndash;12.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eEwald H, Klerings I, Wagner G, Heise TL, Stratil JM, Lhachimi SK, Hemkens LG, Gartlehner G, Armijo-Olivo S, Nussbaumer-Streit B. Searching two or more databases decreased the risk of missing relevant studies: a metaresearch study. J Clin Epidemiol. 2022;149:154\u0026ndash;64.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eFuruya-Kanamori L, Lin L, Kostoulas P, Clark J, Xu C. Limits in the search date for rapid reviews of diagnostic test accuracy studies. Res. 2022;13:13.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGartlehner G, Wagner G, Lux L, Affengruber L, Dobrescu A, Kaminski-Hartenthaler A, Viswanathan M. Assessing the accuracy of machine-assisted abstract screening with DistillerAI: a user study. Syst. 2019;8(1):277.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGates A, Gates M, DaRosa D, Elliott SA, Pillay J, Rahman S, Vandermeer B, Hartling L. Decoding semi-automated title-abstract screening: findings from a convenience sample of reviews. Syst Reviews 2020, 9(1).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGates A, Gates M, Sebastianski M, Guitard S, Elliott SA, Hartling L. The semi-automation of title and abstract screening: a retrospective exploration of ways to leverage Abstrackr's relevance predictions in systematic and rapid reviews. BMC Med Res Methodol. 2020;20(1):139.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGates A, Gates M, Sim S, Elliott SA, Pillay J, Hartling L. Creating efficiencies in the extraction of data from randomized trials: a prospective evaluation of a machine learning and text mining tool. BMC Med Res Methodol. 2021;21(1):169.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGates A, Guitard S, Pillay J, Elliott SA, Dyson MP, Newton AS, Hartling L. Performance and usability of machine learning for screening in systematic reviews: a comparative evaluation of three tools. Syst. 2019;8(1):278.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGates A, Johnson C, Hartling L. Technology-assisted title and abstract screening for systematic reviews: a retrospective evaluation of the Abstrackr machine learning tool. Syst. 2018;7(1):45.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGates A, Vandermeer B, Hartling L. Technology-assisted risk of bias assessment in systematic reviews: a prospective cross-sectional evaluation of the RobotReviewer machine learning tool. J Clin Epidemiol. 2018;96:54\u0026ndash;62.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGates M, Elliott SA, Gates A, Sebastianski M, Pillay J, Bialy L, Hartling L. LOCATE: a prospective evaluation of the value of Leveraging Ongoing Citation Acquisition Techniques for living Evidence syntheses. Syst. 2021;10(1):116.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGiummarra MJ, Lau G, Gabbe BJ. Evaluation of text mining to reduce screening workload for injury-focused systematic reviews. Inj Prev. 2020;26(1):55\u0026ndash;60.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGlanville JM, Lefebvre C, Miles JN, Camosso-Stefinovic J. How to identify randomized controlled trials in MEDLINE: ten years on. J Med Libr Assoc. 2006;94(2):130\u0026ndash;6.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGoossen K, Tenckhoff S, Probst P, Grummich K, Mihaljevic AL, B\u0026uuml;chler MW, Diener MK. Optimal literature search for systematic reviews in surgery. Langenbecks Arch Surg. 2018;403(1):119\u0026ndash;29.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGuimaraes NS, Ferreira AJF, Ribeiro Silva RC, de Paula AA, Lisboa CS, Magno L, Ichiara MY, Barreto ML. Deduplicating records in systematic reviews: there are free, accurate automated ways to do so. J Clin Epidemiol. 2022;152:110\u0026ndash;5.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGartlehner G, Affengruber L, Titscher V, Noel-Storr A, Dooley G, Ballarini N, K\u0026ouml;nig F. Single-reviewer abstract screening missed 13 percent of relevant studies: a crowd-based, randomized controlled trial. J Clin Epidemiol. 2020;121:20\u0026ndash;8.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHaas Q, Alvarez DV, Borissov N, Ferdowsi S, von Meyenn L, Trelle S, Teodoro D, Amini P. Utilizing Artificial Intelligence to Manage COVID-19 Scientific Evidence Torrent with Risklick AI: A Critical Tool for Pharmacology and Therapy Development. Pharmacology. 2021;106(5\u0026ndash;6):244\u0026ndash;53.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHair K, Bahor Z, Macleod M, Liao J, Sena ES. The Automated Systematic Search Deduplicator (ASySD): a rapid, open-source, interoperable tool to remove duplicate citations in biomedical systematic reviews. BMC Biol. 2023;21(1):189.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHamel C, Kelly SE, Thavorn K, Rice DB, Wells GA, Hutton B. An evaluation of DistillerSR's machine learning-based prioritization tool for title/abstract screening - impact on reviewer-relevant outcomes. BMC Med Res Methodol. 2020;20(1):256.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHarrison H, Griffin SJ, Kuhn I, Usher-Smith JA. Software tools to support title and abstract screening for systematic reviews in healthcare: An evaluation. BMC Med Res Methodol. 2020;20(1):7.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHartling L, Featherstone R, Nuspl M, Shave K, Dryden DM, Vandermeer B. Grey literature in systematic reviews: a cross-sectional study of the contribution of non-English reports, unpublished studies and dissertations to the results of meta-analyses in child-relevant reviews. \u003cem\u003eBMC Med Res Methodol\u003c/em\u003e 2017, 17(1):64.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHemens BJ, Haynes RB. McMaster Premium LiteratUre Service (PLUS) performed well for identifying new studies for updated Cochrane reviews. J Clin Epidemiol. 2012;65(1):62\u0026ndash;e7261.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHirt J, Meichlinger J, Schumacher P, Mueller G. Agreement in Risk of Bias Assessment Between RobotReviewer and Human Reviewers: An Evaluation Study on Randomised Controlled Trials in Nursing-Related Cochrane Reviews. J Nurs Scholarsh. 2021;53(2):246\u0026ndash;54.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHoward BE, Phillips J, Tandon A, Maharana A, Elmore R, Mav D, Sedykh A, Thayer K, Merrick BA, Walker V et al. SWIFT-Active Screener: Accelerated document screening through active learning and integrated recall estimation. Environ Int 2020, 138 (no pagination).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHugues A, Di Marco J, Bonan I, Rode G, Cucherat M, Gueyffier F. Publication language and the estimate of treatment effects of physical therapy on balance and postural control after stroke in meta-analyses of randomised controlled trials. PLoS ONE. 2020;15(3):e0229822.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eJanssens A, Gwinn M, Brockman JE, Powell K, Goodman M. Novel citation-based search method for scientific literature: a validation study. BMC Med Res Methodol. 2020;20(1):25.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eJanssens AC, Gwinn M. Novel citation-based search method for scientific literature: application to meta-analyses. BMC Med Res Methodol. 2015;15:84.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eJap J, Saldanha IJ, Smith BT, Lau J, Schmid CH, Li T, Investigators obotDAA. Features and functioning of Data Abstraction Assistant, a software application for data abstraction during systematic reviews. Res Synthesis Methods. 2019;10(1):2\u0026ndash;14.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eJardim PSJ, Rose CJ, Ames HM, Echavez JFM, Van de Velde S, Muller AE. Automating risk of bias assessment in systematic reviews: a real-time mixed methods comparison of human researchers to a machine learning system. BMC Med Res Methodol. 2022;22(1):167.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eJelicic Kadic A, Vucic K, Dosenovic S, Sapunar D, Puljak L. Extracting data from figures with software was faster, with higher interrater reliability than manual extraction. J Clin Epidemiol. 2016;74:119\u0026ndash;23.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKiritchenko S, de Bruijn B, Carini S, Martin J, Sim I. ExaCT: automatic extraction of clinical trial characteristics from journal publications. BMC Med Inf Decis Mak. 2010;10:56.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKwon Y, Lemieux M, McTavish J, Wathen N. Identifying and removing duplicate records from systematic review searches. J Med Libr Association: JMLA. 2015;103(4):184\u0026ndash;8.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLee E, Dobbins M, Decorby K, McRae L, Tirilis D, Husson H. An optimal search filter for retrieving systematic reviews and meta-analyses. BMC Med Res Methodol. 2012;12:51.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLi J, Kabouji J, Bouhadoun S, Tanveer S, Filion KB, Gore G, Josephson CB, Kwon CS, Jette N, Bauer PR, et al. Sensitivity and specificity of alternative screening methods for systematic reviews using text mining tools. J Clin Epidemiol. 2023;162:72\u0026ndash;80.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLi T, Saldanha IJ, Jap J, Smith BT, Canner J, Hutfless SM, Branch V, Carini S, Chan W, de Bruijn B, et al. A randomized trial provided new evidence on the accuracy and efficiency of traditional vs. electronically annotated abstraction approaches in systematic reviews. J Clin Epidemiol. 2019;115:77\u0026ndash;89.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMarshall IJ, Kuiper J, Wallace BC. RobotReviewer: evaluation of a system for automatically assessing bias in clinical trials. J Am Med Inf Assoc. 2016;23(1):193\u0026ndash;201.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMarshall IJ, Marshall R, Wallace BC, Brassey J, Thomas J. Rapid reviews may produce different results to systematic reviews: a meta-epidemiological study. J Clin Epidemiol. 2019;109:30\u0026ndash;41.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMarshall IJ, Trikalinos TA, Soboczenski F, Yun HS, Kell G, Marshall R, Wallace BC. In a pilot study, automated real-time systematic review updates were feasible, accurate, and work-saving. J Clin Epidemiol. 2023;153:26\u0026ndash;33.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMartyn-St James M, Cooper K, Kaltenthaler E. Methods for a rapid systematic review and metaanalysis in evaluating selective serotonin reuptake inhibitors for premature ejaculation. Evid Policy: J Res Debate Pract. 2017;13(3):517\u0026ndash;38.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMateen FJ, Oh J, Tergas AI, Bhayani NH, Kamdar BB. Titles versus titles and abstracts for initial screening of articles for systematic reviews. Clin Epidemiol. 2013;5:89\u0026ndash;95.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMcKeown S, Mir ZM. Considerations for conducting systematic reviews: evaluating the performance of different methods for de-duplicating references. Syst. 2021;10(1):38.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMoher D, Klassen TP, Schulz KF, Berlin JA, Jadad AR, Liberati A. What contributions do languages other than English make on the results of meta-analyses? J Clin Epidemiol. 2000;53(9):964\u0026ndash;72.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMortensen ML, Adam GP, Trikalinos TA, Kraska T, Wallace BC. An exploration of crowdsourcing citation screening for systematic reviews. Res. 2017;8(3):366\u0026ndash;86.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMuthu S. The efficiency of machine learning-assisted platform for article screening in systematic reviews in orthopaedics. Int Orthop. 2023;47(2):551\u0026ndash;6.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eNama N, Iliriani K, Xia MY, Chen BP, Zhou LL, Pojsupap S, Kappel C, O'Hearn K, Sampson M, Menon K, et al. A pilot validation study of crowdsourcing systematic reviews: update of a searchable database of pediatric clinical trials of high-dose vitamin D. Transl. 2017;6(1):18\u0026ndash;26.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eNama N, Sampson M, Barrowman N, Sandarage R, Menon K, Macartney G, Murto K, Vaccani JP, Katz S, Zemek R, et al. Crowdsourcing the Citation Screening Process for Systematic Reviews: Validation Study. J Med Internet Res. 2019;21(4):e12953.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eNg L, Pitt V, Huckvale K, Clavisi O, Turner T, Gruen R, Elliott JH. Title and Abstract Screening and Evaluation in Systematic Reviews (TASER): a pilot randomised controlled trial of title and abstract screening by medical students. Syst Rev. 2014;3:121.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eNoel-Storr A, Dooley G, Affengruber L, Gartlehner G. Citation screening using crowdsourcing and machine learning produced accurate results: evaluation of Cochrane's modified Screen4Me service. J Clin Epidemiol. 2020;29:29.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eNoel-Storr A, Dooley G, Elliott J, Steele E, Shemilt I, Mavergames C, Wisniewski S, McDonald S, Murano M, Glanville J, et al. An evaluation of Cochrane Crowd found that crowdsourcing produced accurate results in identifying randomized trials. J Clin Epidemiol. 2021;133:130\u0026ndash;9.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eNoel-Storr A, Gartlehner G, Dooley G, Persad E, Nussbaumer-Streit B. Crowdsourcing the identification of studies for COVID-19-related Cochrane Rapid Reviews. Res Synth Methods. 2022;13(5):585\u0026ndash;94.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eNoel-Storr AH, Redmond P, Lam\u0026eacute; G, Liberati E, Kelly S, Miller L, Dooley G, Paterson A, Burt J. Crowdsourcing citation-screening in a mixed-studies systematic review: a feasibility study. BMC Med Res Methodol. 2021;21(1):88.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eNussbaumer-Streit B, Klerings I, Dobrescu AI, Persad E, Stevens A, Garritty C, Kamel C, Affengruber L, King VJ, Gartlehner G. Excluding non-English publications from evidence-syntheses did not change conclusions: a meta-epidemiological study. J Clin Epidemiol. 2020;118:42\u0026ndash;54.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eNussbaumer-Streit B, Klerings I, Wagner G, Heise TL, Dobrescu AI, Armijo-Olivo S, Stratil JM, Persad E, Lhachimi SK, Van Noord MG, et al. Abbreviated literature searches were viable alternatives to comprehensive searches: a meta-epidemiological study. J Clin Epidemiol. 2018;102:1\u0026ndash;11.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eO'Keefe H, Rankin J, Wallace SA, Beyer F. Investigation of text-mining methodologies to aid the construction of search strategies in systematic reviews of diagnostic test accuracy-a case study. Res. 2023;14(1):79\u0026ndash;98.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eOlofsson H, Brolund A, Hellberg C, Silverstein R, Stenstrom K, Osterberg M, Dagerhamn J. Can abstract screening workload be reduced using text mining? User experiences of the tool Rayyan. Res. 2017;8(3):275\u0026ndash;80.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eOude Wolcherink MJ, Pouwels X, van Dijk SHB, Doggen CJM, Koffijberg H. Can artificial intelligence separate the wheat from the chaff in systematic reviews of health economic articles? Expert Rev Pharmacoecon Outcomes Res. 2023;23(9):1049\u0026ndash;56.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eOuzzani M, Hammady H, Fedorowicz Z, Elmagarmid A. Rayyan-a web and mobile app for systematic reviews. Syst. 2016;5(1):210.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePallath A, Zhang Q. Paperfetcher: A tool to automate handsearching and citation searching for systematic reviews. Res. 2022;19:19.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePaynter RA, Featherstone R, Stoeger E, Fiordalisi C, Voisin C, Adam GP. A prospective comparison of evidence synthesis search strategies developed with and without text-mining tools. J Clin Epidemiol. 2021;139:350\u0026ndash;60.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePham B, Klassen TP, Lawson ML, Moher D. Language of publication restrictions in systematic reviews gave different results depending on whether the intervention was conventional or complementary. J Clin Epidemiol. 2005;58(8):769\u0026ndash;76.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePham MT, Waddell L, Rajic A, Sargeant JM, Papadopoulos A, McEwen SA. Implications of applying methodological shortcuts to expedite systematic reviews: three case studies using systematic reviews from agri-food public health. Res. 2016;7(4):433\u0026ndash;46.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePianta MJ, Makrai E, Verspoor KM, Cohn TA, Downie LE. Crowdsourcing critical appraisal of research evidence (CrowdCARE) was found to be a valid approach to assessing clinical research quality. J Clin Epidemiol. 2018;104:8\u0026ndash;14.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePijls BG. Machine Learning assisted systematic reviewing in orthopaedics. J Orthop. 2024;48:103\u0026ndash;6.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePradhan R, Hoaglin DC, Cornell M, Liu W, Wang V, Yu H. Automatic extraction of quantitative data from ClinicalTrials.gov to conduct meta-analyses. J Clin Epidemiol. 2019;105:92\u0026ndash;100.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePreston L, Carroll C, Gardois P, Paisley S, Kaltenthaler E. Improving search efficiency for systematic reviews of diagnostic test accuracy: An exploratory study to assess the viability of limiting to MEDLINE, EMBASE and reference checking. Syst Reviews 2015, 4(1).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePrzybyla P, Brockmeier AJ, Kontonatsios G, Le Pogam MA, McNaught J, von Elm E, Nolan K, Ananiadou S. Prioritising references for systematic reviews with RobotAnalyst: A user study. Res. 2018;9(3):470\u0026ndash;88.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRathbone J, Albarqouni L, Bakhit M, Beller E, Byambasuren O, Hoffmann T, Scott AM, Glasziou P. Expediting citation screening using PICo-based title-only screening for identifying studies in scoping searches and rapid reviews. Syst. 2017;6(1):233.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRathbone J, Hoffmann T, Glasziou P. Faster title and abstract screening? Evaluating Abstrackr, a semi-automated online screening program for systematic reviewers. Syst. 2015;4:80.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eReddy SM, Patel S, Weyrich M, Fenton J, Viswanathan M. Comparison of a traditional systematic review approach with review-of-reviews and semi-automation as strategies to update the evidence. Syst. 2020;9(1):243.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRice M, Ali MU, Fitzpatrick-Lewis D, Kenny M, Raina P, Sherifali D. Testing the effectiveness of simplified search strategies for updating systematic reviews. J Clin Epidemiol. 2017;88:148\u0026ndash;53.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRoyle P, Waugh N. A simplified search strategy for identifying randomised controlled trials for systematic reviews of health care interventions: a comparison with more exhaustive strategies. BMC Med Res Methodol. 2005;5:23.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSampson M, de Bruijn B, Urquhart C, Shojania K. Complementary approaches to searching MEDLINE may be sufficient for updating systematic reviews. J Clin Epidemiol. 2016;78:108\u0026ndash;15.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSchopow N, Osterhoff G, Baur D. Applications of the Natural Language Processing Tool ChatGPT in Clinical Practice: Comparative Study and Augmented Systematic Review. JMIR Med Inf. 2023;11:e48933.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eShemilt I, Khan N, Park S, Thomas J. Use of cost-effectiveness analysis to compare the efficiency of study identification methods in systematic reviews. Syst. 2016;5(1):140.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eStoll CRT, Izadi S, Fowler S, Green P, Suls J, Colditz GA. The value of a second reviewer for study selection in systematic reviews. Res. 2019;10(4):539\u0026ndash;45.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eŠuster S, Baldwin T, Verspoor K. Analysis of predictive performance and reliability of classifiers for quality assessment of medical evidence revealed important variation by medical area. J Clin Epidemiol. 2023;159:58\u0026ndash;69.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTaylor-Phillips S, Geppert J, Stinton C, Freeman K, Johnson S, Fraser H, Sutcliffe P, Clarke A. Comparison of a full systematic review versus rapid review approaches to assess a newborn screening test for tyrosinemia type 1. Res. 2017;8(4):475\u0026ndash;84.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eThomas J, McDonald S, Noel-Storr A, Shemilt I, Elliott J, Mavergames C, Marshall IJ. Machine learning reduced workload with minimal risk of missing studies: development and evaluation of a randomized controlled trial classifier for Cochrane Reviews. J Clin Epidemiol. 2021;133:140\u0026ndash;51.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTsou AY, Treadwell JR, Erinoff E, Schoelles K. Machine learning for screening prioritization in systematic reviews: comparative performance of Abstrackr and EPPI-Reviewer. Syst. 2020;9(1):73.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eValizadeh A, Moassefi M, Nakhostin-Ansari A, Hosseini Asl SH, Saghab Torbati M, Aghajani R, Maleki Ghorbani Z, Faghani S. Abstract screening using the automated tool Rayyan: results of effectiveness in three diagnostic test accuracy systematic reviews. BMC Med Res Methodol. 2022;22(1):160.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003evan de Schoot R, de Bruin J, Schram R, Zahedi P, de Boer J, Weijdema F, Kramer B, Huijts M, Hoogerwerf M, Ferdinands G, et al. An open source machine learning framework for efficient and transparent systematic reviews. Nat Mach Intell. 2021;3(2):125\u0026ndash;33.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eVan Enst WA, Scholten RJPM, Whiting P, Zwinderman AH, Hooft L. Meta-epidemiologic analysis indicates that MEDLINE searches are sufficient for diagnostic test accuracy systematic reviews. J Clin Epidemiol. 2014;67(11):1192\u0026ndash;9.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWaffenschmidt S, Guddat C. Searches for randomized controlled trials of drugs in MEDLINE and EMBASE using only generic drug names compared with searches applied in current practice in systematic reviews. Res. 2015;6(2):188\u0026ndash;94.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWaffenschmidt S, Knelangen M, Sieben W, B\u0026uuml;hn S, Pieper D. Single screening versus conventional double screening for study selection in systematic reviews: a methodological systematic review. BMC Med Res Methodol. 2019;19(1):132.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWalker VR, Rooney AA. CEC02-02 Automated and Semi-Automated Approaches for Literature Searching, Screening, and Data Extraction for Systematic Reviews in Environmental Health. Toxicol Lett. 2021;350(Supplement):S4.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWang Z, Asi N, Elraiyah TA, Abu Dabrh AM, Undavalli C, Glasziou P, Montori V, Murad MH. Dual computer monitors to increase efficiency of conducting systematic reviews. J Clin Epidemiol. 2014;67(12):1353\u0026ndash;7.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eXu C, Ju K, Lin L, Jia P, Kwong JSW, Syed A, Furuya-Kanamori L. Rapid evidence synthesis approach for limits on the search date: How rapid could it be? Res. 2022;13(1):68\u0026ndash;76.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWaffenschmidt S, Sieben W, Jakubeit T, Knelangen M, Overesch I, B\u0026uuml;hn S, Pieper D, Skoetz N, Hausner E. Increasing the efficiency of study selection for systematic reviews using prioritization tools and a single-screening approach. Syst Reviews. 2023;12(1):161.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAffengruber L, Nussbaumer-Streit B, Hamel C, Maten MVd, Thomas J, Mavergames C, Spijker R, Gartlehner G. Rapid review methods series: Guidance on the use of supportive software. BMJ Evidence-Based Med 2024:bmjebm\u0026ndash;2023.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003evan Altena AJ, Spijker R, Olabarriaga SD. Usage of automation tools in systematic reviews. Res Synthesis Methods. 2019;10(1):72\u0026ndash;82.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eEwald H, Klerings I, Wagner G, Heise TL, Dobrescu AI, Armijo-Olivo S, Stratil JM, Lhachimi SK, Mittermayr T, Gartlehner G, et al. Abbreviated and comprehensive literature searches led to identical or very similar effect estimates: a meta-epidemiological study. J Clin Epidemiol. 2020;128:1\u0026ndash;12.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHartling L, Featherstone R, Nuspl M, Shave K, Dryden DM, Vandermeer B. Grey literature in systematic reviews: a cross-sectional study of the contribution of non-English reports, unpublished studies and dissertations to the results of meta-analyses in child-relevant reviews. \u003cem\u003eBMC Medical Research Methodology\u003c/em\u003e 2017, 17(1):64.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eEgger M, J\u0026iuml;\u0026iquest;\u0026frac12;ni P, Bartlett C, Holenstein F, Sterne J. How important are comprehensive literature searches and the assessment of trial quality in systematic reviews? Empir Study. 2003;7(1):1\u0026ndash;76.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePham B, Klassen TP, Lawson ML, Moher D. Language of publication restrictions in systematic reviews gave different results depending on whether the intervention was conventional or complementary. J Clin Epidemiol. 2005;58(8):769\u0026ndash;e776762.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMarshall IJ, Wallace BC. Toward systematic review automation: a practical guide to using machine learning tools in research synthesis. Syst Reviews. 2019;8(1):163.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWagner G, Nussbaumer-Streit B, Greimel J, Ciapponi A, Gartlehner G. Trading certainty for speed - how much uncertainty are decisionmakers and guideline developers willing to accept when using rapid reviews: an international survey. BMC Med Res Methodol. 2017;17:121.\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":true,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"bmc-medical-research-methodology","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"bmrm","sideBox":"Learn more about [BMC Medical Research Methodology](http://bmcmedresmethodol.biomedcentral.com/)","snPcode":"","submissionUrl":"https://www.editorialmanager.com/bmrm/default.aspx","title":"BMC Medical Research Methodology","twitterHandle":"BMC_series","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"em","reportingPortfolio":"BMC Series","inReviewEnabled":true,"inReviewRevisionsEnabled":true},"keywords":"rapid review, systematic review, evidence synthesis, scoping review, method, automation tools","lastPublishedDoi":"10.21203/rs.3.rs-4595777/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-4595777/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003ch2\u003eBackground\u003c/h2\u003e \u003cp\u003eSystematic reviews (SRs) are time-consuming and labor-intensive to perform. With the growing number of scientific publications, the SR development process becomes even more laborious. This is problematic because timely SR evidence is essential for decision-making in evidence-based healthcare and policymaking. Numerous methods and tools that accelerate SR development have recently emerged. To date, no scoping review has been conducted to provide a comprehensive summary of methods and ready-to-use tools to improve efficiency in SR production.\u003c/p\u003e\u003ch2\u003eObjective\u003c/h2\u003e \u003cp\u003e To present an overview of primary studies that evaluated the use of ready-to-use applications of tools or review methods to improve efficiency in the review process.\u003c/p\u003e\u003ch2\u003eMethods\u003c/h2\u003e \u003cp\u003e We conducted a scoping review. An information specialist performed a systematic literature search in four databases, supplemented with citation-based and grey literature searching. We included studies reporting the performance of methods and ready-to-use tools for improving efficiency when producing or updating a SR in the health field. We performed dual, independent title and abstract screening, full-text selection, and data extraction. The results were analyzed descriptively and presented narratively.\u003c/p\u003e\u003ch2\u003eResults\u003c/h2\u003e \u003cp\u003eWe included 103 studies: 51 studies reported on methods, 54 studies on tools, and 2 studies reported on both methods and tools to make SR production more efficient. A total of 72 studies evaluated the validity (n\u0026thinsp;=\u0026thinsp;69) or usability (n\u0026thinsp;=\u0026thinsp;3) of one method (n\u0026thinsp;=\u0026thinsp;33) or tool (n\u0026thinsp;=\u0026thinsp;39), and 31 studies performed comparative analyses of different methods (n\u0026thinsp;=\u0026thinsp;15) or tools (n\u0026thinsp;=\u0026thinsp;16). 20 studies conducted prospective evaluations in real-time workflows. Most studies evaluated methods or tools that aimed at screening titles and abstracts (n\u0026thinsp;=\u0026thinsp;42) and literature searching (n\u0026thinsp;=\u0026thinsp;24), while for other steps of the SR process, only a few studies were found. Regarding the outcomes included, most studies reported on validity outcomes (n\u0026thinsp;=\u0026thinsp;84), while outcomes such as impact on results (n\u0026thinsp;=\u0026thinsp;23), time-saving (n\u0026thinsp;=\u0026thinsp;26), usability (n\u0026thinsp;=\u0026thinsp;13), and cost-saving (n\u0026thinsp;=\u0026thinsp;3) were less often evaluated.\u003c/p\u003e\u003ch2\u003eConclusion\u003c/h2\u003e \u003cp\u003eFor title and abstract screening and literature searching, various evaluated methods and tools are available that aim at improving the efficiency of SR production. However, only few studies have addressed the influence of these methods and tools in real-world workflows. Few studies exist that evaluate methods or tools supporting the remaining tasks. Additionally, while validity outcomes are frequently reported, there is a lack of evaluation regarding other outcomes.\u003c/p\u003e","manuscriptTitle":"An exploration of available methods and tools to improve the efficiency of systematic review production - a scoping review ","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2024-07-12 05:19:24","doi":"10.21203/rs.3.rs-4595777/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"decision","content":"Revision requested","date":"2024-07-22T05:49:17+00:00","index":"","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2024-07-20T08:58:13+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"128183586044054529317826204822171557827","date":"2024-07-12T00:25:56+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"33723699812702970427020308811304869644","date":"2024-07-11T17:22:31+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2024-07-11T11:35:03+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"40702906633502621628861451348434258339","date":"2024-07-11T10:14:18+00:00","index":"hide","fulltext":""},{"type":"reviewersInvited","content":"","date":"2024-07-05T17:17:10+00:00","index":"","fulltext":""},{"type":"editorInvited","content":"","date":"2024-06-21T16:06:39+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2024-06-21T16:03:55+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2024-06-21T16:02:59+00:00","index":"","fulltext":""},{"type":"submitted","content":"BMC Medical Research Methodology","date":"2024-06-17T18:52:47+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"bmc-medical-research-methodology","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"bmrm","sideBox":"Learn more about [BMC Medical Research Methodology](http://bmcmedresmethodol.biomedcentral.com/)","snPcode":"","submissionUrl":"https://www.editorialmanager.com/bmrm/default.aspx","title":"BMC Medical Research Methodology","twitterHandle":"BMC_series","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"em","reportingPortfolio":"BMC Series","inReviewEnabled":true,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"dc9667b7-8c81-430f-859f-c78f447616db","owner":[],"postedDate":"July 12th, 2024","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"published-in-journal","subjectAreas":[],"tags":[],"updatedAt":"2024-09-23T16:04:52+00:00","versionOfRecord":{"articleIdentity":"rs-4595777","link":"https://doi.org/10.1186/s12874-024-02320-4","journal":{"identity":"bmc-medical-research-methodology","isVorOnly":false,"title":"BMC Medical Research Methodology"},"publishedOn":"2024-09-18 15:56:53","publishedOnDateReadable":"September 18th, 2024"},"versionCreatedAt":"2024-07-12 05:19:24","video":"","vorDoi":"10.1186/s12874-024-02320-4","vorDoiUrl":"https://doi.org/10.1186/s12874-024-02320-4","workflowStages":[]},"version":"v1","identity":"rs-4595777","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-4595777","identity":"rs-4595777","version":["v1"]},"buildId":"qtupq5eGEP_6zYnWcrvyt","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.