Managing Pregnancies of Unknown Location With the M4 Prediction Model or the NICE Algorithm: A Randomised Controlled Trial With Cross-Sectional Diagnostic Accuracy Data.

OA: gold CC-BY-NC-4.0

Abstract

ObjectiveTo determine the diagnostic performance and clinical utility of the M4 prediction model and the NICE algorithm managing women with pregnancy of unknown location (PUL).DesignThe study has a superiority design regarding specificity for non-ectopic pregnancy for M4, given that the primary outcome of sensitivity for ectopic pregnancy (EP) is non-inferior in comparison with the NICE algorithm.SettingEmergency gynaecology units in Sweden.Population595 women with PUL.MethodsParticipants were randomised (1:1) to M4 or the NICE algorithm after two serum human chorionic (hCG) levels and were categorised as high or low risk of having an EP. The diagnostic performance was evaluated on cross-sectional data and utility by parallel groups.Main outcome measuresThe proportion of EP categorised as high risk (sensitivity) and non-ectopic pregnancies categorised as low risk (specificity). Clinical outcomes were assessed.ResultsThe sensitivity for EP was 79% (115 of 146) for M4 versus 85% (124 of 146) for the NICE algorithm, p = 0.1496 and the specificity for non-ectopic pregnancies was 67% (300 of 449) for M4 and 74% (334 of 449) for the NICE algorithm, p = 0.0003. Clinical outcomes were similar between groups.ConclusionsThe sensitivity for EP by M4 was non-inferior to NICE, but specificity was better for the NICE algorithm. No between group differences were observed for clinical outcomes.Trial registrationNCT03461835, https://www.Clinicaltrialsgov.
Full text 26,629 characters · extracted from pmc-nxml · 8 sections · click to expand

Author

All persons listed as authors fulfil the International Committee of Medical Journal Editors (ICMJE) criteria for authorship: J.F. and A.S. contributed substantially to conception and design of the study, interpreted the data, and approved the final version after revision. J.F. was responsible for site monitoring and drafted the article. H.G. and K.B. contributed to the collection of data, site monitoring, commented on the drafts, and approved the final version. C.B. commented on the drafts, revised the article, and approved the final version. All authors agree to be accountable for all aspects of the work in ensuring that questions related to the accuracy or integrity of any part of the work are appropriately investigated and resolved.

Methods

This was a multicentre randomised trial of two protocols for PUL management with a cross‐sectional design for the diagnostic performance. The inference on the clinical utility of the protocols was based on their diagnostic performance in combination with subsequent safety and effectiveness, delineated by secondary outcomes. The trial was approved by the Regional Ethical Review Board, Gothenburg, Sweden (approval #382–17) and all patients provided written informed consent. The reporting is based on the CONSORT guidelines. The patients and/or the public were not involved in the design, conduct, reporting, or dissemination plans of this research. A statistical analysis plan (SAP) was constructed by a senior statistician not involved in the trial and can be found along with the study protocol at clinicaltrials.gov . Women with PUL that were managed as outpatients at participating sites were eligible. We defined four types of PUL based on TVS, published in a consensus statement, but used slightly different terminology to align with new recommendations and recent research (Data  S1 ) [ 1 , 20 , 21 ]: Women were excluded from participation if (1) confirmation of pregnancy site at the first visit (2) hospitalisation at the first visit (3) decline to participate or (4) not understanding the oral or written (Swedish) study information. Randomisation occurred after the second hCG and required that (1) the first hCG level was < 10.000 IU/L (2) no hospitalisation before the second hCG and (3) the second hCG was measured within 44–56 h after the first hCG. Upon arrival at the gynaecology emergency unit at Sahlgrenska University hospital, Södra Älvsborg Hospital or Skaraborg Hospital Skövde, a triage nurse assessed whether women with complaints or worries in early pregnancy needed to see a physician or not. After examination eligible women were enrolled by the physician who filled out a case report form (CRF), and two hCG samplings were scheduled by the nurse, details of serum and urine hCG assays used are found in Data  S1 . The principal investigator entered hCG levels into an electronic CRF (eCRF) and randomisation module. A nurse contacted women by telephone and scheduled an appointment for follow‐up. Data was collected from the interlinked electronic medical system between hospitals and the eCRF. Participants were assigned in a 1:1 ratio to M4 or the NICE algorithm group using a computerised randomisation module. A concealed randomisation sequence was generated using t ‐test minimisation to ensure an even allocation ratio between groups. Participants, clinicians, and investigators were blinded to randomised groups. Participants in both groups were categorised by the respective protocols as either high or low risk of having an EP. Both protocols provided the same recommendation for follow‐up, depending on the categorisation to high or low risk: (1) A high‐risk PUL had a re‐examination within 24–48 h (2) a low‐risk PUL predicted to be a spontaneously resolving PUL had a home urine pregnancy test after 2 weeks and if tested positive a re‐examination was performed after 24–48 h (3) a low‐risk PUL predicted to be a normally sited pregnancy was re‐examined after 1 week. Details of the M4 prediction model, including covariates (Table  S1 ), and the NICE algorithm can be viewed in Supplement 1. The primary outcomes were the sensitivity for EP and the specificity for non‐EP, reported along with AUC and negative positive predictive values (NPV) and positive predictive values (PPV). Secondary outcomes were time from randomisation to diagnosis, successful first line treatment of EP, time from diagnosis until resolution (hCG < 5.3 IU/L) of EP and persistent PUL, length of follow‐up, number of hCG and TVS, interventions and adverse events (as defined in the SAP). If a repeat dose of methotrexate was needed to achieve resolution of pregnancy it was still considered a successful treatment. We defined emergency surgery as when a patient made an unscheduled visit to the emergency unit due to deteriorating symptoms and surgery was commenced shortly thereafter. No core outcome set exists for PUL but there is a high degree of convergence with those developed for EP, including treatment success, resolution time, number of additional interventions and adverse events [ 22 ]. Reasons for nonadherence to protocol were recorded and were either due to nurse planning, physician decision or unplanned visit by patient. The observed outcome of PUL served as reference standard and included normally sited pregnancy, EP and spontaneously resolving PUL. Definitions of outcomes are described in Data  S1 . Based on previous data from our clinic the NICE algorithm has a 70% specificity for non‐EP and an 86% sensitivity for EP [ 19 ]. The present study was originally designed to have 80% power to detect an 8% points higher specificity (superiority) for non‐EP and a non‐inferior sensitivity for EP by M4 compared with the NICE algorithm with a non‐inferiority margin for the difference between groups of 7% points and would require 600 participants per group to account for dropouts. Due to slow recruitment and recommendation after external review 30 March 2022 it was decided to analyse the primary outcome for both protocols on the whole sample data as a cross‐sectional design instead of a parallel design, when half of the first estimated sample size was reached, i.e., 300 participants per group. For evaluating diagnostic performance, the doubled sample size could be utilised to increase power. The parallel design was kept for secondary outcomes. Due to these considerations the study was prematurely terminated in the end of 2022 when approximately 600 participants were enrolled. An important aspect making it possible to transfer the design from a solely RCT design to a partly cross‐sectional design was that diagnostic performance of the two protocols were not dependent on the clinical management. All randomised subjects with a known outcome of PUL were included in the full analysis set (FAS). Subjects included in FAS, that adhered to protocol according to categorisation in high or low risk constituted the per‐protocol population. Nonadherent subjects were not included in the per‐protocol population (Figure  1 ). Participant trial flow. FAS, full analysis set; PP, per‐protocol. The primary analysis was performed for both M4 and NICE on cross‐sectional data for all subjects in the FAS (Figure  1 ) for the primary outcome. The sensitivity represented EP and persistent PUL correctly categorised as high risk and the specificity represented non‐EP (spontaneously resolving PUL and normally sited pregnancies) correctly categorised as low risk. The AUC described the protocols' ability to discriminate between EP and non‐EP. Positive predictive value represented EP among PUL categorised as high risk and NPV represented non‐EP among PUL categorised as low risk. A sensitivity analysis was conducted to test differences in diagnostic performance by randomised groups as originally intended for which the per‐protocol population was used (non‐inferiority design). An exploratory analysis for the FAS was also performed. An exact 95% confidence interval (CI) was calculated for the difference in sensitivity and specificity. Two main analyses were performed on the FAS for secondary outcomes corresponding to the clinical utility between randomised groups and between subjects categorised as high or low risk. Complementary analyses were also done for per‐protocol populations. The risk ratio and 95% CI were calculated. Descriptive statistics were presented by mean, SD, median, minimum and maximum (continuous variables) and number and percentage (categorical variables). Kaplan–Meier curves were used to report the time to event. The applied tests are presented in the respective tables and figures. All tests were two‐tailed and conducted at the 0.05 significance level. All analyses were performed by using SAS v9.4 (Cary, NC).

Results

There were 883 women screened between 20 February 2018 and 31 December 2022, of which 16 declined to participate and 260 did not fulfil eligibility criteria as shown in the trial flow (Figure  1 ). Out of 607 women randomised, 595 (98%) had data available for the primary outcome and were included in the primary cross‐sectional analysis. Adherence to the protocol was seen for 256 of 302 (84.8%) women in the M4 group and 257 of 293 (87.7%) in the NICE group ( p  = 0.74), details including reasons for non‐adherence can be found in Table  S2 . Baseline characteristics of participants were similar between groups (Table  1 ) as were characteristics of the first TVS and hCG (Table  2 ). HCG data was comparable between the M4 and the NICE group (Tables  S3–S7 ). In all, 2136 hCG blood test were obtained and 1350 TVSs performed. Of 113 EP (19.3%) in the study, 53 underwent laparoscopy of which 92.4% (49 of 53) were detected on TVS before surgery. A surgical procedure was carried out on 78 of 595 (13.1%) women. After randomisation, 175 (29.4%) women with spontaneously resolving PUL had a two‐week telephone call in line with protocol. Six had a positive urine pregnancy test, two were re‐examined and had additional hCG levels analysed that declined spontaneously. Two had a negative home urine pregnancy test after another week and two had a negative test when re‐attending the clinic. Baseline and clinical characteristics of patients by randomised group in the full analysis set. Note : Data presented as n /total (%) if not stated otherwise. Percentages may not add up to 100% due to rounding. Abbreviation: BMI, body mass index. Participants can be part of multiple previous pregnancy categories. Worries due to previous ectopic pregnancy. Self‐reported. Characteristics of transvaginal sonography and hCG at the first visit. Note : Data presented as n /total (%) if not stated otherwise. Percentages may not add up to 100% due to rounding. Abbreviation: hCG, human chorionic gonadotrophin. Descriptions used in these women's medical record was either blighted ovum or retained products of conception. M4 had a 79% sensitivity for EP versus 85% for the NICE algorithm, p  = 0.15 and lower specificity for non‐EP, 67% versus 74% (Table  3 ). The NICE algorithm had a higher AUC for predicting EP than M4 0.84 versus 0.80, p  = 0.0374, (Table  3 and Figure  S1 ), but lower AUC for normally sited pregnancies and spontaneously resolving PUL (Table  S8 ). In the sensitivity analysis the diagnostic performance was analysed by randomised groups and the sensitivity for EP by M4 was non‐inferior to NICE by not exceeding the non‐inferiority margin of seven percentage points while no statistically significant difference was demonstrated in the specificity for non‐EP risk (Table  S9A ). The sensitivity was inflated in both groups since primarily EP categorised as low risk (21 vs. 3) did not adhere to protocol (M4, n  = 14, NICE, n  = 7) which reduced the number of false negative results whilst not as much the number of true positive results in the per‐protocol population. The diagnostic performance analysed by randomised groups in the FAS was similar to the cross‐sectional data in the primary analysis (Table  S9B ). Primary outcome analysed on cross‐sectional data ( n  = 595). a 67 (62–71) 300/449 74 (70–78) 334/449 Note : Data presented as % (95% CI) and n/ total if not stated otherwise. Abbreviations: AUC, area under the curve; hCG, human chorionic gonadotrophin; NA, not applicable; PUL, pregnancy of unknown location. The sign test was used for testing the difference in sensitivity and specificity and DeLong's test was used for the AUC. Persistent PUL ( n  = 33) included among ectopic pregnancies. Normally sited pregnancies and spontaneously resolving PUL. Time to diagnosis of EP was similar between groups (Figure  S2 ). The overall rate of successful first line treatment of EP was not statistically significantly different in the M4 versus the NICE group p  = 0.35, neither was the type of treatment nor was the hCG resolution time (Table  4 ). In the M4 group 63.4% (33/52) of EP were managed without laparoscopic surgery versus 44.3% (27/61) in the NICE group ( p  = 0.064). No differences were observed in the rate of medical and surgical interventions for any PUL outcome (Table  S10 ). No between groups differences were seen for time to diagnosis, length of follow‐up and number of hCG and TVS (Table  4 and Table S11 ). Secondary outcomes by randomised groups. a Note : Data presented as n/ total (%) if not stated otherwise. Differences are reported as M4‐NICE and risk ratio as M4/NICE. Differences between percentages are given in percentage points; differences between other values are given in the unit for that value. Abbreviations: AUC, area under the curve; hCG, human chorionic gonadotrophin; NA, not applicable; sPUL, pregnancy of unknown location. The Fisher's Exact test (lowest 1‐sided p‐value multiplied by 2) was used for testing difference between groups for dichotomous variables and the t‐test for continuous variables. Persistent PUL not included among ectopic pregnancies. Calculated from the time of diagnosis until hCG < 5.3 IU/L, data were missing for one woman in the M4 group. Includes all PUL outcomes. Calculated from randomisation, data were missing for one woman in the M4 group. The risk ratio for adverse events between M4 and NICE was 1.28; 95% CI, 0.85 to 1.95 (Table  S12 ), meaning that there might be an almost doubled adverse event rate in the M4 group without this study being able to detect that. Six women were diagnosed with a ruptured EP, all were discharged from the hospital the following day after laparoscopy, one received blood‐transfusion (Table  S13 ). Unplanned visit was the most common adverse event in both groups. Seven women re‐presented to the clinic and had emergency laparoscopy, an EP was diagnosed in five cases and two women were diagnosed with a miscarriage. Three out of five EPs requiring emergency surgery were labelled low risk by M4 ( n  = 1) or NICE ( n  = 2). However, two had high risk assessment due to either physician concern or suspicion of EP on the first TVS. A higher percentage of women categorised as high risk experienced any adverse event than women categorised as low risk (18.3% vs. 9.7%), p  = 0.0039 (Table  S14 ). Ectopic pregnancies categorised as high risk had a longer time to resolution than EP categorised as low risk (17.2 vs. 10.7 days, p  = 0.022) and length of follow‐up (Table  S15 ). The time to diagnosis of EP, successful first line treatment and the number of hCG and TVS were not statistically significantly different (Table  S15 ). There was a higher rate of medical and surgical procedures among normally sited pregnancies with false positive results than those with true negative results (Table  S16 ). False positive normally sited pregnancies and spontaneously resolving PUL had a higher number of hCG samples and TVS than true negative cases (Table  S17 ). Spontaneously resolving PUL with false positive results had longer time to diagnosis and length of follow‐up. HCG characteristics are described in Table  S18 . Results of secondary outcomes for the per‐protocol population were generally consistent with those in the FAS population regarding the direction of relative risk, but the effect size was slightly higher for some results (Tables  S19–S25 ).

Discussion

In this randomised controlled trial, M4 had non‐inferior sensitivity for EP but lower specificity for non‐EP than the NICE algorithm when assessed on cross‐sectional data. These results were supported by analyses according to randomised groups, although the sensitivity for EP was higher in the per‐protocol population in both groups. The NICE algorithm was better than M4 at predicting EP. No differences were observed for clinical outcomes between randomised groups but for high and low risk categories. The rate of adverse events was almost twice as high among women categorised as high risk compared with women categorised as low risk. Non‐EP with false positive results had prolonged time to diagnosis, length of follow and had more clinic examinations and hCG samplings. The time to diagnosis of EP categorised as high risk was not shorter than for EP categorised as low risk (false negative results) while having longer hCG resolution time and length of follow‐up. The main strengths of this study are the partly randomised design across multiple centres, large sample size and minor losses to follow‐up. Further, this is the first study evaluating both diagnostic performance and clinical impact on women with PUL to determine utility between protocols in a real‐world setting. We included different types of PUL and accepted a range in time span between hCG measurement, as commonly done, to ensure a representative cohort of participating clinics and making our results generalisable to other settings [ 21 ]. The main limitation is the premature ending of the study, thus only half of the calculated sample size was reached due to slow recruitment. Second, the trial was not powered to detect differences in clinical outcomes, which was further aggravated by the reduced sample size. Third, a successor to M4, named M6, has been developed, but was not used because no external validation study had yet been carried out [ 23 ]. Fourth, almost 15% of participants were not managed according to protocol, mostly due to a physician decision, and should be considered when interpreting the results. However, clinical judgement is an important safety aspect of the initial PUL management as shown in the present study, where EPs categorised as low risk in many cases were assessed as high risk. The rate of non‐adherence in the present study was comparable with that reported from a study evaluating the implementation of M6 in the UK [ 24 ]. Fifth, a large proportion of patients with early pregnancy complications was managed by physicians before being specialist. Although a senior physician was available upon request, study findings may not be generalizable to clinical settings with a different seniority level. The M4 had similar sensitivity for EP as in an implementation study in the UK [ 2 ]. In a retrospective study from our unit M4 had an 8% point higher sensitivity while specificity was identical [ 19 ]. The absolute hCG concentrations is a covariate in M4 that increase the estimated probability for EP and in PUL populations with lower average hCG levels the sensitivity is impaired as seen in an ART setting [ 25 ]. The average hCG levels for EP was 1471 IU/L in our previous study versus 862 IU/L in the present study contributing to lower estimates. There is a substantial assay‐related variation for hCG analysis why the same assay should be used when monitoring early pregnancies [ 26 , 27 ]. The hCG levels between our study populations were similar for other PUL outcomes and we do not attribute assay variations to have caused the dissimilarities. These aspects are important to recollect when adopting any decision rule that rely on absolute hCG levels. Contrary, the diagnostic performance of the NICE algorithm was highly consistent. The average number of blood tests and TVS performed was similar as in the M4 implementation study, at least for correctly categorised non‐EPs adhering to protocol [ 2 ]. Ectopic pregnancies in our previous cohort of PUL were mainly treated surgically without a need for further hCG monitoring, which doubled the number of hCG samplings in the present study [ 19 ]. Approximately 60% of EPs were managed conservatively in the present study as in the M4 and a M6 implementation study [ 2 , 24 ]. They reported fewer hCG samplings and scans for EP, but recordings stopped once diagnosed, making comparisons faulty. The overall surgical intervention rate in our study was 13.1%, in accordance with another recent study [ 28 ]. The rate of medical or surgical procedures among non‐EP was in level with the M6 study [ 24 ]. In the present study 55.7% of EPs in the NICE group were treated laparoscopically and 35.6% in the M4 group. We found that the median hCG level for EP in the NICE group was 610 IU/L versus 405 IU/L in the M4 group, p  = 0.11. Salpingectomy might have been necessary to a greater extent in the NICE group since EPs with higher hCG levels more often fail to resolve spontaneously [ 29 ]. This could be associated with a higher economic cost for EP treatment in the NICE group [ 30 ]. The difference in hCG levels for EP between groups is likely to be at random, as this was not seen for other PUL outcomes. The rupture rate of EP was similar between groups and there was one case in each group with disruption of a possibly live pregnancy, two well recognised complications [ 2 , 24 ]. By carefully selecting women with EP for expectant management and identifying EP at increased risk of failing methotrexate and prepare for surgical intervention are means for limiting rupture rates [ 29 , 31 ]. A majority of ruptured EPs in the present study were already under close surveillance or received methotrexate and seemed not to be associated with a delay in TVS confirmation. It is also important to be familiar with ultrasonographic criteria to safely rule out a live pregnancy and not rely on too slow increment in hCG levels [ 32 , 33 ]. Although an expectant approach is often feasible, an exception might be persistent PUL that according to one study had better chance of resolution from active management, with vacuum aspiration or methotrexate [ 34 ]. The number of clinic visits and length of follow‐up of women with spontaneously resolving PUL seems to have a casual relationship with the application of a PUL protocol. Whether this is true for safety outcomes is questionable given our results. Within the framework of a standardised management of PUL and the oversight of patient flow that it brings, together with our findings and current body of evidence, it may be preferable to prioritise a lower rate of false positive results over sensitivity. This could reduce the workload and burden for patients without compromising safety. By accepting a 21% hCG level decline as previously proposed or utilise a progesterone cut‐off < 10 nmol/L for allowing no further clinic visits could aid this objective [ 35 , 36 , 37 ]. A third option is the M6 model that has a better overall predictive performance than M4, potentially exceeding NICE [ 21 , 38 ].

Conclusions

The NICE algorithm had better diagnostic and predictive ability than M4. Resemblances between the protocols and study size hampered detection of differences in clinical outcomes. Women with non‐EP benefit from being correctly categorised as low risk, while EP being correctly categorised as high risk did not ensure a faster diagnosis. The study was not powered to detect differences for secondary outcomes requiring careful interpretation of these findings. Still, a protocol with higher sensitivity for EP may enhance patient safety. There is a need for well‐designed studies to compare protocols with substantial dissimilarities, such as single‐visit approaches using progesterone versus serial hCG levels. Development of a core outcome set for PUL could serve as a support for future research.

Introduction

A main objective of transvaginal ultrasonography (TVS) in early pregnancy is to determine if the implanted embryo is normally sited, ectopic (EP) or of unknown location (PUL) [ 1 ]. The distinction is not always clear but will induce specific clinical pathways wherein the risk with diagnostic errors lies. The unintended interruption of a live pregnancy or fatal intraabdominal bleeding from a ruptured EP has been reported among women with PUL [ 2 , 3 ]. PUL protocols are designed to prevent such events, stemming from hasty or delayed clinical interventions [ 4 ]. The management of PUL is associated with a significant workload and therefore needs to be efficient to alleviate patient and societal burden [ 5 ]. No global PUL strategy exists but many guidelines recommend repeated serum human chorionic gonadotrophin (hCG) levels to guide subsequent management [ 6 , 7 , 8 , 9 , 10 , 11 ]. Comparisons between PUL protocols have been done with traditional statistical measures such as sensitivity, specificity, and area under the receiver operating characteristic curves (AUC) [ 12 ]. One promising protocol is the M4 prediction model, which in a meta‐analysis had the best ability to predict an EP among PUL [ 13 ]. It is first and foremost recognised that prediction models can excel in one clinical setting but yield unacceptable performance when externally validated, with the risk of undermining patients' safety if implemented carelessly [ 14 , 15 ]. Second, the impact on relevant clinical outcomes has scarcely been studied for available prediction models [ 16 ]. Third, as a decision support software, M4 qualifies as a medical device and is required under international regulations, to provide positive patients outcomes compared with standard of care, which in the UK would be a clinical management algorithm (NICE guideline 126) [ 7 , 17 , 18 ]. The M4 had better specificity for non‐ectopic pregnancy (non‐EP) than the NICE algorithm and similar sensitivity for EP in a retrospective study [ 19 ]. Despite the general use of the NICE algorithm, its diagnostic performance has not been evaluated prospectively and to even a lesser extent has the clinical utility of any PUL protocols been compared head‐to‐head. In this pragmatic randomised clinical trial, we aimed to determine if M4 has superior specificity for non‐EP and noninferior sensitivity for EP compared with the NICE algorithm. The clinical utility and aspects of PUL with false‐positive or false‐negative results were also assessed.

Coi Statement

J.F., A.S. and C.B. have received research grants for this work.

Supplementary Material

Data S1: Data S2:

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: pmc-nxml

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-08-30T09:23:35.175841+00:00
unpaywall
last seen: 2026-08-13T06:47:16.638238+00:00
License: CC-BY-NC-4.0