METHODS OVERVIEW
How to conduct a systematic review
Last reviewed September 14, 2026
A systematic review answers a focused research question by searching, screening, and synthesizing every relevant study using a documented, repeatable method - as opposed to a narrative review, which summarizes literature at the author's discretion. This is the full sequence, stage by stage, with a free tool and a deeper guide linked at each step.
What makes a review "systematic"
Four things distinguish a systematic review from an ordinary literature summary: a question specified in advance, a documented search strategy across named databases, explicit inclusion/exclusion criteria applied by more than one reviewer, and a transparent accounting of how many records moved through each stage. Remove any of these and a review stops being reproducible - a second team following your methods section should arrive at close to the same set of included studies.
Meta-analysis (statistically pooling results) is common in systematic reviews but not required - a review can be systematic in method and still report its findings narratively if the included studies are too different to pool.
Step 1: Formulate your question and register a protocol
Frame the question with PICO - Population, Intervention, Comparison, Outcome - before searching anything. A vague question ("does exercise help depression?") produces an unmanageable search; a PICO question ("in adults with major depressive disorder, does supervised aerobic exercise, compared with usual care, reduce depressive symptom scores?") defines your search terms and your inclusion criteria in the same sentence.
Writing this down as a protocol - and registering it publicly (PROSPERO is the standard registry for health-related reviews) before screening begins - is what prevents outcome-switching after the fact. See Guide 1: Protocol and PICO and the free protocol template.
Step 2: Search and import the literature
Translate your PICO question into a search string for each relevant database (commonly PubMed/MEDLINE, Embase, Cochrane CENTRAL, plus subject-specific databases), combining synonyms with OR and concepts with AND. Export results as RIS, BibTeX, or CSV and import them into whatever tool you're screening in.
Duplicate records across databases are normal and expected - dedup before screening starts, not after, so reviewers aren't wasting time re-deciding on the same record twice.
Database searching alone typically misses some relevant studies, which is why most systematic review methods also call for supplementary searching: checking the reference lists of included studies (backward citation chasing), checking who has cited them since (forward citation chasing), and searching grey literature - conference abstracts, dissertations, trial registries, and preprints - where negative or unpublished results are more likely to surface. Skipping this step is one of the more common, and more consequential, shortcuts in a rushed review, since it systematically biases the included evidence toward published, positive findings.
Step 3: Screen title/abstract, then full text
Screening happens in two passes. First, title/abstract: quick exclusions against your PICO criteria get most records out. Second, full text: the remaining records are read in full against the same criteria, and this is where most detailed exclusion reasons come from (wrong population, wrong comparator, no usable outcome data, wrong study design).
Both passes should use at least two independent reviewers, blind to each other's decisions, with disagreements resolved by discussion or a third reviewer - not by whichever reviewer screened first. See Guide 2: Screening efficiency.
Step 4: Extract data with a structured form
Design your extraction fields before extraction starts - study design, sample size, population characteristics, intervention/comparator details, outcome measures and effect estimates, funding source - and have two reviewers extract independently where feasible, reconciling discrepancies against the source text. A shared spreadsheet with inconsistent column usage across reviewers is the most common way extraction data becomes unusable at the analysis stage.
See Guide 3: Precision extraction and the free data extraction template.
Step 5: Assess risk of bias
Every included study should be judged for methodological quality - randomization, blinding, attrition, and other domain-specific concerns - so readers can weigh how much confidence the pooled result deserves. Established structured tools exist for this (Cochrane's RoB 2 for randomized trials, ROBINS-I for non-randomized studies); at minimum, record a judgment against each relevant domain rather than a single unexplained "quality: good/bad" label.
The reason these tools break judgment into separate domains, rather than one overall score, is that a study can be strong on one axis and weak on another - well-randomized but with high dropout, say - and collapsing that into a single number hides exactly the information a reader needs to judge whether a specific result is trustworthy. Two reviewers assessing risk of bias independently, the same way they screen independently, catches disagreements about how a study actually behaved versus how it was described in the abstract.
Step 6: Report with a PRISMA 2020 flow diagram
PRISMA 2020 is the reporting standard nearly every journal expects: a flow diagram showing records identified, duplicates removed, records screened and excluded, reports sought and not retrieved, reports assessed and excluded (with reasons), and studies finally included. The numbers in this diagram must reconcile exactly with the numbers in your methods and results text - a mismatch is a reporting defect reviewers will flag.
See Guide 4: PRISMA 2020 and the free PRISMA flow diagram generator.
Step 7: Synthesize the findings
If included studies measure comparable outcomes with comparable populations, a meta-analysis pools their effect sizes into one estimate - fixed-effect if you believe all studies share one true effect, random-effects if you expect genuine variation between them (check this with Cochran's Q and I² first). If studies are too heterogeneous in population, intervention, or outcome definition to pool meaningfully, synthesize narratively instead and say so explicitly - forcing a pooled number onto incompatible studies is worse than not pooling at all.
See Guide 5: Statistical synthesis and the free I² calculator.
Team size and realistic timelines
A minimum viable team is two reviewers plus someone to arbitrate disagreements - often a third reviewer or a senior author, though on small teams this is frequently the same person who wrote the protocol. Larger reviews (thousands of records, multiple outcomes) benefit from more screeners working in parallel, but extraction and synthesis are harder to parallelize without a shared, structured template, since inconsistent field usage between extractors creates work that has to be redone at analysis time.
Timelines vary widely, but a rough shape holds across most reviews: protocol and search strategy development takes 2-6 weeks; title/abstract screening of a few thousand records takes 3-8 weeks with two reviewers working part-time; full-text screening and extraction together often take as long as everything before them combined; risk-of-bias assessment and synthesis add another few weeks; write-up and internal review before submission typically add 4-8 more. Reviews with a narrow, well-defined question and a small literature move faster; reviews on a broad or contested topic with tens of thousands of initial records can take a year or more regardless of team size.
Common pitfalls
Changing inclusion criteria after seeing results. This is exactly what protocol registration exists to prevent - decide your criteria before you know which studies they'll include or exclude.
Single-reviewer screening. Faster, but it removes the main safeguard against inconsistent or careless inclusion decisions - disclose it as a limitation if you do it.
PRISMA numbers that don't add up. Usually caused by exclusion reasons that don't sum to the eligibility-exclude box, or duplicates subtracted twice. Fix the source counts; don't adjust the diagram to look balanced.
Pooling studies that shouldn't be pooled. A low I² supports pooling; a high I² without exploring why (subgroup differences, outlier studies) is a sign to slow down, not a number to explain away.
Ignoring publication bias in the synthesis. Studies with null or negative results are less likely to be published at all, which can skew a pooled effect toward showing more benefit than actually exists. A funnel plot (or a formal test like Egger's) on your pooled studies is the standard check, and it matters more the fewer studies you have.
FAQ
How long does a systematic review take?
Most take 6-18 months from protocol to submission, depending on team size, the volume of literature, and how many databases are searched. Screening and data extraction - not writing - are usually the largest time sinks, which is why dual-reviewer throughput and structured extraction forms matter more than they might seem to at the outset.
What's the difference between a systematic review and a literature review?
A systematic review follows a pre-specified, documented protocol - a defined search strategy, explicit inclusion/exclusion criteria, dual independent screening, and a reported flow of records through each stage (typically as a PRISMA 2020 diagram) - so another team could in principle reproduce it. A narrative literature review has none of these requirements; it's a synthesis shaped by the author's judgment about what to include.
Do I need to register a protocol before starting?
It's considered best practice and is required by some journals and funders, particularly for health-related reviews (registries like PROSPERO are the common venue). Registering before screening begins protects against outcome-switching - quietly changing your inclusion criteria or outcomes after seeing which results support a preferred conclusion.
How many reviewers do I need for screening?
At least two, screening independently, is the standard for both title/abstract and full-text stages - the point is to catch each reviewer's individual blind spots and inconsistencies. A single reviewer is faster but sacrifices the main protection against selective or careless inclusion decisions.
Can one person do a systematic review alone?
It's done, particularly for scoping or rapid reviews, but it forgoes independent dual screening entirely, which is a real limitation to disclose in the methods section - not a technicality to omit. If you're working solo, look for a workflow that at least makes your screening decisions and extraction data auditable after the fact.
Do I need statistical software for meta-analysis?
Only if you're pooling numeric outcomes across studies (fixed/random-effects models, forest plots). A review that only summarizes findings narratively doesn't require it. Many teams still run screening and extraction in one tool and export data to a separate stats package (R, RevMan) for pooling; some, including EvidenceFlow, keep both steps in one workspace.
EvidenceFlow runs import, dedup, dual screening, extraction, PRISMA reporting, and meta-analysis in one workspace, free with no review-count limit. AI-assisted screening, extraction auto-fill, and manuscript drafting are an optional paid upgrade.