Question Which risk score for lower gastrointestinal bleeding best discriminates safe discharge, major bleeding, need for transfusion, and need for hemostasis?
Findings In this systematic review and meta-analysis of 9 studies of 4 risk scores, the Oakland score was the most discriminative for predicting safe discharge, major bleeding, and need for transfusion, whereas the Strate score was the best at predicting need for hemostasis.
Meaning This study suggests that the Oakland and Strate scores can be used to predict lower gastrointestinal bleeding outcomes with a high degree of certainty.
Importance Clinical prediction models, or risk scores, can be used to risk stratify patients with lower gastrointestinal bleeding (LGIB), although the most discriminative score is unknown.
Objective To identify all LGIB risk scores available and compare their prognostic performance.
Data Sources A systematic search of Ovid MEDLINE, Embase, and the Cochrane Central Register of Controlled Trials from January 1, 1990, through August 31, 2021, was conducted. Non–English-language articles were excluded.
Study Selection Observational and interventional studies deriving or validating an LGIB risk score for the prediction of a clinical outcome were included. Studies including patients younger than 16 years or limited to a specific patient population or a specific cause of bleeding were excluded. Two investigators independently screened the studies, and disagreements were resolved by consensus.
Data Extraction and Synthesis Data were abstracted according to the Preferred Reporting Items for Systematic Reviews and Meta-analyses (PRISMA) guideline independently by 2 investigators and pooled using random-effects models.
Main Outcomes and Measures Summary diagnostic performance measures (sensitivity, specificity, and area under the receiver operating characteristic curve [AUROC]) determined a priori were calculated for each risk score and outcome combination.
Results A total of 3268 citations were identified, of which 9 studies encompassing 12 independent cohorts and 4 risk scores (Oakland, Strate, NOBLADS [nonsteroidal anti-inflammatory drug use, no diarrhea, no abdominal tenderness, blood pressure ≤100 mm Hg, antiplatelet drug use (nonaspirin), albumin <3.0 g/dL, disease score ≥2 (according to the Charlson Comorbidity Index), and syncope], and BLEED [ongoing bleeding, low systolic blood pressure, elevated prothrombin time, erratic mental status, and unstable comorbid disease]) were included in the meta-analysis. For the prediction of safe discharge, the AUROC for the Oakland score was 0.86 (95% CI, 0.82-0.88). For major bleeding, the AUROC was 0.93 (95% CI, 0.90-0.95) for the Oakland score, 0.73 (95% CI, 0.69-0.77) for the Strate score, 0.58 (95% CI, 0.53-0.62) for the NOBLADS score, and 0.65 (95% CI, 0.61-0.69) for the BLEED score. For transfusion, the AUROC was 0.99 (95% CI, 0.98-1.00) for the Oakland score and 0.88 (95% CI, 0.85-0.90) for the NOBLADS score. For hemostasis, the AUROC was 0.36 (95% CI, 0.32-0.40) for the Oakland score, 0.82 (95% CI, 0.79-0.85) for the Strate score, and 0.24 (95% CI, 0.20-0.28) for the NOBLADS score.
Conclusions and Relevance The Oakland score was the most discriminative LGIB risk score for predicting safe discharge, major bleeding, and need for transfusion, whereas the Strate score was best for predicting need for hemostasis. This study suggests that these scores can be used to predict outcomes from LGIB and guide clinical care accordingly.
Lower gastrointestinal bleeding (LGIB) is a common reason for emergency hospitalization, with an annual incidence rate upward of 87 cases per 100 000 individuals.1–4 Most often, LGIB is a self-limiting condition, with most cases resolving spontaneously.1,5 However, major bleeding resulting in blood transfusion, surgery, and even death can occur in some cases. In a large, prospective cohort study involving 2528 cases of LGIB across 143 hospitals, 26.3% of patients required blood transfusion, 1% required embolization or surgery, and 3.4% died.5 On a population level, LGIB is expensive and imposes considerable costs on the health care system, largely owing to hospitalizations. In a registry study of 30 498 hospitalized patients with gastrointestinal bleeding, those with LGIB had longer lengths of stay (mean [SD], 13.9 [8.8] days vs 11.6 [7.9] days) and higher resource use than patients with upper gastrointestinal bleeding.2
Of critical importance to the management of these patients is differentiating the majority of people who can be safely discharged for outpatient care from those who are at risk for serious adverse events and require hospitalization. Doing so would avoid the cost and burden of unnecessary hospitalization for patients at low risk of adverse outcomes while reliably identifying patients at risk for hospital admission. Clinical prediction models, or risk scores, are well suited for this purpose.6 A highly discriminative LGIB risk score would allow for the dichotomization of patients into high-risk and low-risk groups for adverse outcomes and can be used to guide clinical care.
To date, numerous LGIB risk scores have been developed.7 However, the quality of these risk scores, based in part on the representativeness of the derivation cohort, the degree of external validation, and the risk score’s accuracy, are largely unknown. Thus, the objective of this study was to conduct the first meta-analysis of LGIB risk scores, to our knowledge, based on prognostic performance.
This systematic review and meta-analysis followed the Preferred Reporting Items for Systematic Reviews and Meta-analyses (PRISMA) reporting guideline8 and the Transparent Reporting of a Multivariable Prediction Model for Individual Prognosis or Diagnosis (TRIPOD) reporting guideline, and systematic searches were conducted in Ovid MEDLINE, Embase, and the Cochrane Central Register of Controlled Trials databases from January 1, 1990, through August 31, 2021. All non–English-language articles were excluded. The search queries were developed using a combination of subject headings and alternative free-text terms (eAppendix 1 in the Supplement). Optimized methodological search filters and text words were used to refine search results. The search strategies were modified for each database to include database-specific index terms. We also searched ClinicalTrials.gov to identify unpublished trials and abstracts from scientific meetings for the past 5 years from Digestive Disease Week, the American College of Gastroenterology Annual Scientific Meeting, and United European Gastroenterology Week. Reference lists of relevant articles and reviews were examined.
Observational studies and clinical trials with the objective of deriving or validating an LGIB risk score to predict an outcome, such as safe discharge, mortality, rebleeding, need for hemostatic intervention, and need for blood transfusions, were included. Studies including patients younger than 16 years or limited to a restricted patient population, such as older patients, or cause of bleeding, such as diverticular bleeding, were excluded. Because most studies did not directly report the number of true positives, true negatives, false positives, and false negatives, these values were calculated from extracted data; studies with insufficient information to calculate these metrics were excluded because they could not be meta-analyzed. In these situations, corresponding authors were contacted twice, 1 month apart, in an attempt to garner the missing information before the study was excluded.
Titles and abstracts were screened followed by full-text review independently by 2 of us (M.A. and M.G.). Discrepancies between the 2 reviewers were resolved by consensus, and failing that, they were resolved by a third reviewer (M.S.) who made the final determination. The study was registered in PROSPERO (CRD42018110347).
Data were extracted from eligible studies independently by 2 of us (M.A. and M.G.) and preceded by trialing of the data collection document. Disagreements were resolved by consensus, and failing that, they were resolved by a third reviewer (M.S.) who made the final determination. Data describing the study population, the risk scores used, and their diagnostic performances (true positive, true negative, false positive, or false negative for each risk score, cutoff, and outcome combination) were extracted. Studies containing more than 1 independent cohort underwent data extraction in which each cohort was denoted by a lowercase letter following the first author’s name and year of publication (eg, Strate [2005A], Strate [2005B]). Study quality was measured using the modified Quality Assessment on Diagnostic Accuracy Studies tool (QUADAS-2; eAppendix 2 in the Supplement).9
Outcomes of interest for the meta-analysis included the prediction of safe discharge, major bleeding, transfusion, need for hemostasis, and mortality, although there were insufficient publications to meta-analyze the last outcome. Safe discharge was defined as the absence of major bleeding, transfusion, need for hemostasis, readmission for LGIB within 28 days, and death. Major bleeding was defined as having either recurrent bleeding or severe bleeding, as these terms were not standardized between studies but ultimately represented some form of heightened bleeding (eAppendix 3 in the Supplement). Need for hemostasis was defined as the requirement for endoscopic hemostasis, radiologic embolization, or surgery to control bleeding; radiology for diagnostic purposes and surgery for nonhemostatic purposes did not qualify. Need for blood transfusion was defined as receiving at least 1 unit of red blood cells.
The ability of LGIB risk scores to predict each outcome was measured by its discrimination, calculated based on the area under the receiver operating characteristic curve (AUROC).8 Discrimination is a measure of a risk score’s ability to differentiate between patients who have developed and patients who have not developed an outcome event.6 Thus, a highly discriminative score allows for the identification of patients who are unlikely to develop an outcome event and separate them from those who are at higher risk for the outcome event.
Binary study outcomes were analyzed using the random-effects regression model called the hierarchical summary receiver operating characteristic curve (HSROC) model as described by Rutter and Gatsonis.10 This model is useful for its connection to the bivariate-normal model with random effects11,12 and for producing summary receiver operating characteristic curves (ROCs) for the diagnostic test.10 For each outcome (safe discharge, major bleeding, transfusions, and hemostasis) and for each risk score, the HSROC model was used to explicitly model the reported risk score thresholds as fixed cofactors, allowing sensitivity and specificity to vary with threshold. These models typically require at least 3 studies for convergence. If the summary ROC models could not be fit owing to insufficient data, we attempted to use a continuity correction of one that was added to zero cells to try to improve model stability. Using the HSROC model, predicted sensitivity, specificity, and derived AUROC for a target threshold were derived. This threshold, or risk score cutoff value, was chosen based on the original publication of the risk score to predict specific LGIB-related outcomes. This approach was used because the additional information allows for more precise estimates and, in some cases, fitting a bivariate-normal regression model to the subset of studies reporting the threshold of interest resulted in unstable model estimation. There are no consensus measures or methods for quantifying or testing heterogeneity for the accuracy of diagnostic tests in meta-analyses. For selected scores and specific cutoff values, we therefore conducted a meta-analysis for each diagnostic performance measure (eg, sensitivity and specificity) individually. We used a random-effects meta-analysis model with an empirical Bayes estimator of between-study heterogeneity for the purposes of summarizing between-study heterogeneity. Although there are limitations with this approach, this model did provide familiar parameters of heterogeneity from the perspective of a meta-analysis of clinical trials. One limitation is that correlation between paired measures, such as sensitivity, specificity, and likelihood ratios, are not modeled, and, as such, these heterogeneity estimates should be interpreted cautiously. Because positive and negative predictive values depend on the prevalence of the outcome, which varied between studies, heterogeneity was not further characterized for these metrics.
The HSROC models were estimated using the PROC NLMIXED procedure in SAS, version 9.4 (SAS Institute Inc) with the MetaDAS macro published by the Cochrane Collaboration.13 Model parameters and summary diagnostic classification data were input to RevMan, version 5.4.1 (The Cochrane Collaboration) to produce summary ROC and forest plots.
LEGGI TUTTO




