Meta-Analysis of Observational Studies: MOOSE, Bias and Pooling

How to run a meta-analysis of observational studies: MOOSE and PRISMA 2020 reporting, adjusted odds ratios, NOS/ROBINS-I appraisal, random-effects pooling.

A meta-analysis of observational studies pools cohort, case-control and cross-sectional evidence when randomised trials are unethical, impractical or absent. The statistics look familiar, but the threats do not: every included estimate carries its own confounding and bias, and pooling can make a biased answer look remarkably precise. This guide covers what changes when the evidence is observational.

Why a meta-analysis of observational studies is different

In a trial, randomisation balances known and unknown confounders; in observational research nothing does. Three problems travel straight into the pooled estimate:

  • Confounding: exposed and unexposed groups differ in age, socioeconomic status, smoking or comorbidity, and each study adjusts for a different subset.
  • Selection bias: who enters the cohort, who is lost to follow-up and how controls are recruited can distort the association.
  • Information bias: self-reported exposures, recall bias in case-control studies and misclassified outcomes shift estimates, often unpredictably.

Designs also answer different questions: a cohort follows incidence over time, a case-control study works backwards from outcome to exposure, and a cross-sectional survey captures prevalence at one moment, leaving temporality uncertain. Large registry studies then add very narrow confidence intervals around estimates that may still be confounded. Precision is not validity.

Protocol and reporting: PROSPERO, PRISMA 2020 and MOOSE

Register the protocol on PROSPERO before screening starts. Frame the question with PECO (population, exposure, comparator, outcome) and pre-specify what would otherwise be decided post hoc: the primary effect measure, the minimum confounder set, the design subgroups and the sensitivity analyses.

Report with PRISMA 2020 for the flow diagram and checklist, plus the MOOSE checklist (Meta-analysis Of Observational Studies in Epidemiology, 2000), which asks for observational-specific detail such as how confounding was assessed and how study quality entered the analysis. Many epidemiology journals expect both. For the core review steps, see our meta-analysis and PRISMA guide.

Study designs, effect measures and adjusted estimates

Never pool cohort, case-control and cross-sectional studies blindly. The defensible default is to analyse each design as a subgroup, report the test for subgroup differences and give an overall estimate only if the designs agree. Cross-sectional studies, which cannot establish temporality, often belong in a separate or sensitivity analysis.

Studies usually report odds ratios (OR), risk ratios (RR) or hazard ratios (HR). When the outcome is rare (roughly below 10% in the unexposed group, as a rule of thumb), the OR approximates the RR and the HR approximates both, so they are often analysed together with a sensitivity analysis by measure. When the outcome is common, the OR exaggerates the RR: convert using the baseline risk or analyse the measures separately.

Extract the adjusted estimate with its 95% confidence interval, preferably from the model that meets your pre-specified confounder set; the most adjusted model is not automatically best if it adjusts for mediators. Ratios are pooled on the log scale. Worked example: adjusted OR = 1.45, 95% CI 1.12 to 1.88, so ln(OR) = 0.372 and SE = (ln(UL) − ln(LL)) / 3.92 = (0.631 − 0.113) / 3.92 = 0.132. Include each cohort once, even if it produced several papers. For studies reporting mean differences or correlations, our effect size guide covers the d and r families.

Pooling model and heterogeneity

A random-effects model is the default, because populations, exposure definitions and adjustment sets genuinely differ. The classic DerSimonian–Laird estimator tends to underestimate between-study variance when studies are few, giving overly narrow intervals. REML is a better default estimator of τ², and the Hartung–Knapp (Hartung–Knapp–Sidik–Jonkman) adjustment, based on a t distribution, gives more honest intervals with few studies.

Cochran's Q tests whether variation exceeds chance but has low power with few studies. I² is the share of variability due to between-study differences; 25%, 50% and 75% are rough benchmarks for low, moderate and high, but I² rises with study precision, so large registry studies inflate it even when absolute differences are modest. τ² gives the absolute between-study variance, and the 95% prediction interval shows where the effect in a new study is likely to fall; if it crosses 1, say so plainly.

7858.53919.5078All studies34Cohort52Case-control69Cross-sectional
I² (%) for all 24 studies and within the cohort (k = 11), case-control (k = 8) and cross-sectional (k = 5) subgroups; illustrative example

Design often explains part of the heterogeneity. Explore the rest with meta-regression on pre-specified moderators, such as smoking adjustment or length of follow-up, allowing roughly ten studies per moderator.

Risk of bias, small-study effects and GRADE certainty

Risk-of-bias tools for observational studies (report judgements by domain, not as a total score)
ToolDesigned forOutputWatch out for
Newcastle–Ottawa Scale (NOS)Cohort and case-control studiesStars for selection, comparability and outcome/exposure (up to 9)A summed score hides which domain failed
ROBINS-INon-randomised studies of interventionsLow, moderate, serious or critical riskJudged against a hypothetical target trial; confounding is decisive
ROBINS-EObservational studies of exposures (environmental, occupational, lifestyle)Domain judgements from low to very high riskMore demanding; budget time for reviewer calibration
JBI checklistsAnalytical cross-sectional, cohort, case-control and prevalence studiesYes / no / unclear / not applicable per itemUse the checklist that matches each design

For small-study effects, inspect a funnel plot and run Egger's test only with about 10 or more studies; with fewer, it lacks power. In observational reviews, asymmetry can reflect confounding or design differences as well as unpublished null results, so interpret it cautiously and use trim-and-fill only as a sensitivity analysis. Pre-specify these sensitivity analyses:

  • Leave-one-out: re-pool omitting each study in turn to find influential estimates.
  • Excluding high risk of bias: keep only studies at low or moderate risk.
  • Adjusted versus crude: a large gap between pooled adjusted and unadjusted estimates signals strong confounding.
  • By measure and design: ORs separately from RRs and HRs, plus a cohort-only analysis.

Finally, rate certainty with GRADE. Observational evidence starts at low certainty and is rated down for risk of bias, inconsistency, indirectness, imprecision or publication bias; it can be rated up for a large effect, a dose–response gradient, or plausible residual confounding that would have weakened the observed effect.

A Stata workflow for observational meta-analysis

  1. Build one row per study: label, design, the log of the adjusted estimate (logor) and its standard error (se).
  2. meta set logor se, studylabel(study) random(reml) declares the data with a REML random-effects model (Stata 16 or later).
  3. meta summarize, eform predinterval gives the pooled OR with Q, I², τ² and the prediction interval; add random(dlaird) for DerSimonian–Laird or se(khartung) for Hartung–Knapp.
  4. meta forestplot, subgroup(design) eform plots the results by design with the test of group differences.
  5. meta funnelplot, meta bias, egger and meta trimfill address small-study effects.
  6. meta regress fits moderators, and in recent releases meta summarize, leaveoneout runs the leave-one-out analysis.

Comprehensive Meta-Analysis (CMA) accepts adjusted ratios with confidence limits directly, recent versions of SPSS include meta-analysis procedures, and we use Python scripts for reproducible extraction checks. To hand over pooling and appraisal, see our systematic review and meta-analysis service or send us your protocol.

In observational meta-analysis, a narrow confidence interval measures precision; it is never a certificate of validity.

Frequently Asked Questions

Is a meta-analysis of observational studies reliable?

It can be, if confounding and bias are handled explicitly through adjusted estimates, design subgroups, formal risk-of-bias appraisal and sensitivity analyses. Even then, GRADE starts observational evidence at low certainty, so conclusions should be worded as associations rather than proven causal effects.

Can cohort and case-control studies be combined in one meta-analysis?

They can, but not blindly. Analyse each design as a subgroup, test for subgroup differences, and report an overall estimate only when the designs give consistent results and address the same question.

Should I use adjusted or unadjusted odds ratios in a meta-analysis?

Use adjusted estimates for the primary analysis, ideally those meeting the confounder set in your protocol, and convert them to log ORs with SE = (ln(UL) − ln(LL)) / 3.92. Crude estimates belong in a sensitivity analysis showing how far confounding shifts the result.

What meta-analysis services does Celsus offer for observational studies?

Our meta-analysis service covers observational evidence end to end: PECO question and PROSPERO protocol, search and dual screening, NOS, ROBINS-I or ROBINS-E appraisal, adjusted-estimate extraction, and random-effects pooling in Stata or CMA with subgroup, sensitivity and publication-bias analyses. Delivery includes GRADE ratings, forest and funnel plots, and PRISMA 2020 and MOOSE-compliant reporting with reproducible files.

← All posts