GUIDE 3 OF 5
Precision extraction: forms, omics, and risk of bias
After inclusion, the unit of work is a field, not a yes/no. Structured forms, RoB judgements, and a link back to the PDF are what make a pooled estimate defensible.
Binary screening decisions become granular capture: who was studied, what was compared, which effect can be pooled, and whether the trial’s design undermines that effect. For omics reviews, the same stage is where GWAS variants or gene-level RNA-seq summaries enter a table instead of a slide deck.
Designing structured extraction forms for clinical and omics data
Start from fields you will actually analyse. Clinical reviews typically need bibliographic identity, PICO, dichotomous 2×2 counts or continuous mean/SD/N, a reported effect with SE or CI, and an overall RoB call. Omics reviews add gene, variant, phenotype, and study-level summary statistics you can map across papers.
The free data-extraction template is that column set in the browser, with CSV download. EvidenceFlow projects use the same idea as an extraction table you can dual-review.
Implementing Cochrane RoB 2 frameworks
RoB 2 is domain-based, not a single “good/bad” stamp. Two domains repay extra attention because they most often move a trial from “some concerns” to “high”:
| Domain | What you are asking |
|---|---|
| 1 — Randomization process | Was allocation sequence and concealment adequate? |
| 2 — Deviations from intended interventions | Did blinding and adherence hold? |
| 3 — Missing outcome data | Is loss to follow-up likely related to the outcome? |
| 4 — Measurement of the outcome | Could outcome assessment be influenced by knowledge of assignment? |
| 5 — Selection of the reported result | Were outcomes and analyses pre-specified? |
A senior methodologist will often start with Domain 1 and Domain 3. Weak randomization and informative missing data both bias a pooled estimate even when sample sizes look large.
Ensuring data provenance
Every extracted value that will enter a forest plot should be recoverable from the PDF: page, table number, or a short quote. Without that, disagreements at analysis time become archaeology.
EvidenceFlow AI+ can propose auto-fill from full text; you still edit before it feeds analysis. Continue with Guide 4: PRISMA 2020.
FAQ
Raw counts or published OR/RR?
Prefer events and totals, or mean/SD/N. Pooling can derive OR, RR, RD, MD, or SMD instead of copying an effect that used a different model.
Is AI extraction the final number?
No. Auto-fill proposes values from the PDF. A reviewer must accept or edit them. Auto-fill is paid; manual extraction is free.
What about missing outcomes?
Record missingness on the form. That feeds RoB 2 Domain 3 and stops you from quietly dropping studies at pooling.
Core import, screening, extraction, meta-analysis, and PRISMA reporting are free. AI relevance scoring and extraction auto-fill are the paid upgrade.