What is a meta-analysis?
This guide assumes you have never run a systematic review or a meta-analysis before. It starts with the most basic question, namely what these things even are, and builds up one layer at a time, until you can run a complete review in Synthesis with a team and submit the result to a journal.
It is written in two interleaved tracks. The first track teaches the underlying research method, so you understand what you are doing and why. The second track shows you how Synthesis performs each step, so you are never just clicking buttons you do not understand. Look for the two callout styles throughout.
Start with a single study and its problem
Imagine one clinical study tests whether a new drug lowers blood pressure. It enrolls 120 patients, half on the drug and half on placebo, and finds the drug group did slightly better. Is that enough to change how doctors treat patients?
Usually, no. A single study of 120 people can be wrong for ordinary reasons: the patients happened to be unusually healthy, the effect was a fluke of chance, the clinic was unusually good, or the result was simply too imprecise to trust. Any one study, however well run, is a single data point in a noisy world.
The idea behind a meta-analysis
Now imagine fifteen separate teams around the world each ran a similar study of that same drug. Individually each is limited. But if you could carefully gather all fifteen and combine their results into one pooled answer, you would effectively have a study of thousands of patients, far more reliable than any single one.
A meta-analysis is exactly that: a statistical method for combining the numerical results of many separate studies into one pooled estimate. The whole point is to turn many small, uncertain answers into one larger, more trustworthy answer.
Researchers do this because better-pooled evidence leads to better decisions for clinical guidelines, for policy, and for patients. Meta-analyses sit at the very top of the evidence hierarchy for precisely this reason.
Systematic review vs meta-analysis: not the same thing
These two terms are constantly confused, so fix the distinction now:
- A systematic review is the whole rigorous process of finding, screening, appraising, and summarizing all the evidence on a question, following a pre-declared method so it can be reproduced and audited.
- A meta-analysis is the optional statistical step inside a systematic review where you actually pool the numbers. Not every systematic review contains one. If the studies are too different to combine, you report them narratively instead.
Put simply: the systematic review is the entire factory, from raw literature to finished paper. The meta-analysis is one machine inside that factory, the statistics engine. This guide teaches the whole factory, because that is what you will actually run.
Synthesis runs the entire factory in a single tool. The traditional path uses one app to screen, a separate program to run statistics, and a word processor to write, forcing you to re-key data between them. Synthesis connects all of it so your data flows through untouched.
The seven-stage pipeline
Every systematic review with meta-analysis follows the same sequence of stages. Learn this map first; the rest of the guide is just each stage in detail.
| Stage | What you do | What comes out |
|---|---|---|
| 1. Frame the question | Define a precise question and write a protocol | A registered, locked plan |
| 2. Search the literature | Search databases to find every relevant study | Hundreds to thousands of records |
| 3. Screen and select | Filter records down to the studies that qualify | A handful of included studies |
| 4. Extract the data | Pull the raw numbers out of each study | A structured data table |
| 5. Assess risk of bias | Judge how trustworthy each study is | A quality rating per study |
| 6. Pool statistically | Combine the numbers into one estimate | A pooled result and forest plot |
| 7. Rate certainty and report | Grade the evidence and write it up | A submission-ready manuscript |
Two checks run alongside the whole pipeline. Heterogeneity: are the studies consistent enough to combine at all? Publication bias: are missing or unpublished studies quietly skewing the result? Both are covered in The two checks below.
Stage 1: Frame the question and write the protocol
PICO: turning a vague idea into an answerable question
A research question like “is this drug any good?” cannot be answered systematically because it is not specific. You make it answerable with the PICO framework, which forces you to name four things:
- P, Population: who exactly? (e.g. adults with type 2 diabetes)
- I, Intervention: what treatment or exposure? (e.g. drug X, 10 mg daily)
- C, Comparison: compared against what? (e.g. placebo)
- O, Outcome: measuring what, and when? (e.g. HbA1c reduction at 12 weeks)
Strung together: “In adults with type 2 diabetes (P), does drug X (I) compared with placebo (C) reduce HbA1c at 12 weeks (O)?” That is a question a review can actually answer.
The protocol and why you write it first
A protocol is a written plan that states your question, your search strategy, your inclusion and exclusion rules, and exactly how you will analyze the data, all decided and written down before you look at any results. You then register it publicly, most commonly on a registry called PROSPERO.
This pre-commitment is the single most important safeguard for honesty in a review. If you decide your methods after seeing the data, you can, even unconsciously, pick the rules that produce the answer you were hoping for. Locking the plan first makes that impossible and is what reviewers and journals look for.
Synthesis gives you a protocol-writing stage with a structured PICO template, so your question, inclusion criteria, and planned analysis are captured in one place and carried forward automatically into screening and extraction. You define the rules once, and the tool enforces them downstream.
Changing your inclusion criteria after you have started screening, without documenting why, is one of the fastest ways to get a review rejected. If you must change the plan, record the change and the reason. Reviewers expect an audit trail, not a silent edit.
Stage 2: Search the literature
Your goal is to find every relevant study, not just the convenient ones. A biased search produces a biased review no matter how good the later steps are.
How a real search is built
You translate your PICO terms into a structured search query using Boolean logic (AND, OR, NOT) and controlled vocabulary. In medicine, these standardized tags are called MeSH terms. You then run that query across multiple databases, because no single database contains everything. The usual set includes PubMed, Embase, the Cochrane Library, and Web of Science, plus clinical-trial registries to catch studies that were run but never published.
You also hand-search the reference lists of the studies you find, and sometimes email authors for unpublished data. A typical search returns several hundred to several thousand records.
Synthesis lets you import search results directly from databases like PubMed and de-duplicates them automatically, so the same study found in three databases collapses into one record. You start screening from a clean, merged list instead of a messy spreadsheet.
Stage 3: Screen and select studies
Screening is a funnel. You start with everything the search returned and progressively remove what does not qualify, in two rounds.
The two screening rounds
- Title and abstract screening: you read only the title and abstract of each record and decide whether it could plausibly meet your criteria. This is fast and removes the large majority of records.
- Full-text screening: you read the complete paper of every survivor and apply your inclusion and exclusion criteria strictly. What remains is your set of included studies.
A worked example of the funnel:
| Step | Records remaining | Removed |
|---|---|---|
| Identified by search | 2,400 | n/a |
| After de-duplication | 1,600 | 800 duplicates |
| After title/abstract | 120 | 1,480 irrelevant |
| After full-text review | 18 | 102 failed criteria |
Why two people screen independently
Each record is screened by two reviewers working separately, neither able to see the other’s decisions. Where they disagree, a third reviewer breaks the tie. This is called dual or double screening, and the blinding matters: if reviewer two can see that reviewer one already said “include,” they tend to go along with it, and the second check becomes worthless.
Synthesis enforces true dual-blinded screening with proper reviewer roles, so the two screeners genuinely cannot see each other’s calls, and disagreements are routed to a resolver. Every decision is logged to an audit trail. It also offers optional priority screening: a built-in classifier learns from your include/exclude decisions and reorders the queue so the most likely-relevant records surface first. This is an aid to ordering your work; you still make every decision yourself.
Sharing one login between two reviewers completely breaks dual-blinding, because there is no real second opinion and the audit trail is meaningless. Each reviewer must have their own account. This is a known weakness of some older tools; do not replicate it.
Stage 4: Extract the data (the heart of the review)
Extraction is where you pull the raw numbers out of each included study so they can be pooled. The single most important principle to internalize:
You record the raw numbers reported in each paper, not the paper’s conclusions. The meta-analysis recomputes everything from scratch. Your job is to faithfully transcribe data and prove where it came from, not to interpret it.
Two passes: study characteristics, then outcomes
Extraction has two parts. The first describes the study once; the second captures its results, of which there may be several.
Pass 1: study characteristics (one set per study). This is context: who was studied and how. Typical fields are study design, country or region, setting, enrollment period, total participants, mean age, percentage male, the diagnostic criteria or definition used, and free-text notes. It is also good practice to record the study’s funding source and any conflicts of interest, because these feed into your bias assessment later.
Pass 2: outcomes (one or more per study). A single study can report several outcomes, for example mortality, response rate, and an adverse event, and each can be stratified by its own criteria. For each outcome you record what was measured and the numbers behind it. What numbers you need depends entirely on the type of outcome.
The data you extract, by outcome type
This is the part new reviewers most often get wrong, so go slowly. The kind of outcome decides which numbers you must capture.
Binary (dichotomous) outcomes: did the event happen or not? Examples: died vs survived, cured vs not, relapsed vs not. For a two-group comparison you need four numbers, the events and the group total, for each of the two arms:
| Study | Treatment events | Treatment total | Control events | Control total |
|---|---|---|---|---|
| Smith 2019 | 12 | 100 | 20 | 98 |
| Lee 2021 | 8 | 75 | 15 | 80 |
| Patel 2022 | 22 | 150 | 30 | 145 |
Continuous outcomes: a measured value. Examples: blood pressure, weight, a pain score. For a two-group comparison you need six numbers, the mean, the standard deviation (SD), and the sample size, for each arm.
Other outcome types:
- Proportion: events divided by total in a single group (e.g. prevalence).
- Incidence rate: events over person-time observed.
- Single mean: one mean and SD with its sample size.
- Generic effect: used when a paper only reports a pre-computed effect (an odds ratio, risk ratio, or hazard ratio) with its 95% confidence interval, rather than the raw counts.
The details that quietly decide whether your pooling is valid
Beyond the headline numbers, capture these or your later analysis can be silently wrong:
- Timepoint: when was it measured? Pooling a 12-week result with a 1-year result as if they were the same is an error.
- Endpoint vs change-from-baseline: a value at the end is not the same as the change from the start. Never mix them in one analysis.
- Dispersion type: a reported spread might be an SD, a standard error (SE), or a confidence interval, and they are not interchangeable. Record which one it is, and let the engine convert, rather than converting by hand.
- Effect measure: if you record a pre-computed effect, note whether it is an OR, RR, or HR.
- Unit: mmHg is not mmol/L. A unit mismatch ruins a pooled continuous result.
- Derived flag and source quote: if you calculated a number rather than reading it directly, mark it as derived; and always paste the verbatim sentence you took the number from.
Synthesis structures extraction into the same two passes: study characteristics, then criteria-stratified outcomes. When you pick an outcome’s measure type, the form adapts to show exactly the fields that measure needs (both arms for a comparison, mean/SD/N per arm for continuous, and so on), so you cannot leave out the control arm or fill in irrelevant boxes. Every field has a verbatim source-quote box for provenance, and a derived flag for any number you calculated. Two AI agents can pre-extract independently and the tool only pre-fills fields where they agree; disagreements are flagged for your judgement. The AI drafts; you decide and submit. The final numbers are always yours.
The most common fatal extraction error is recording only one arm of a comparison. A binary comparison needs four numbers and a continuous one needs six. If you capture only the treatment arm, the statistics engine has nothing to compare against and the outcome is unusable.
Stage 5: Assess risk of bias
Before combining studies, you judge how trustworthy each one is. A flawless 2,000-patient randomized trial deserves more confidence than a small, poorly controlled one. This judgement is separate from how big the study is; it is about how well it was conducted.
You use a standardized tool so the judgement is consistent and reproducible. For randomized trials the common one is the Cochrane Risk of Bias tool (RoB 2); for observational studies, the Newcastle-Ottawa Scale. These rate each study across domains, such as whether the randomization was proper, whether participants and assessors were blinded, whether outcome data was complete, and whether there was selective reporting, and each domain is rated as low risk, some concerns, or high risk.
Note the distinction from statistical weighting: risk-of-bias is your structured judgement of study quality, whereas the weighting in Stage 6 is an automatic mathematical function of each study’s precision. Both matter, and they are not the same thing.
Synthesis provides the standard appraisal instruments as structured forms, records a rating and a justification for each domain, and produces the risk-of-bias summary figures journals expect, all linked to the same studies you screened and extracted.
Stage 6: Pool the studies statistically
This is the meta-analysis proper. It happens in three conceptual moves.
Move one: convert each study to a common effect size
Studies report results in different ways, so you first translate each into a single comparable number called an effect size:
- For binary outcomes: an odds ratio, risk ratio, or risk difference.
- For continuous outcomes: a mean difference, or a standardized mean difference when studies used different measurement scales.
An effect size of “no difference” is 1.0 for ratio measures (odds and risk ratios) and 0 for difference measures. That reference point is the line everything is judged against.
Move two: combine them as a weighted average
You do not take a simple average. Larger, more precise studies are given more weight. Formally, each study’s weight is proportional to the inverse of its variance, which means precise studies with tight confidence intervals count for more. A 2,000-patient study moves the pooled result far more than a 50-patient one.
There are two weighting models, and choosing between them matters:
- Fixed-effect model: assumes every study is estimating the exact same single true effect. Appropriate only when studies are very similar.
- Random-effects model: assumes the true effect varies a little between studies (different populations, doses, settings) and accounts for that extra variation. Most real-world medical meta-analyses use this, because identical studies are rare.
Move three: read the forest plot
The result is displayed as a forest plot, the signature image of meta-analysis. Here is how to read one:
- Each horizontal row is one study.
- The square sits at that study’s effect size; the size of the square shows its weight, so a bigger square means more influence.
- The horizontal line through the square is its 95% confidence interval; a wide line means an imprecise study.
- The vertical line down the middle is “no effect.”
- The diamond at the bottom is the pooled result. Its center is the combined estimate and its width is the combined confidence interval. If the whole diamond sits to one side of the no-effect line and does not touch it, the pooled effect is statistically significant.
Synthesis runs the pooling on a deterministic statistics engine and produces forest plots directly from your extracted data, with no exporting to a separate statistics program. It supports both fixed-effect and random-effects models, the standard effect measures, and methods (such as GLMM for proportion and prevalence data) that some traditional tools lack. Critically, the AI assistance never touches the computed estimate. Every pooled number is produced by transparent, cross-verified statistics on the numbers you extracted, so the result is reproducible and defensible.
The two checks: heterogeneity and publication bias
Heterogeneity: should these studies be combined at all?
If one study shows the drug helps and another shows it harms, averaging them produces a meaningless middle number. Heterogeneity is the degree to which study results disagree beyond what chance alone would explain.
You measure it most commonly with the I-squared statistic, a percentage where 0% means the studies are perfectly consistent and 75% or more signals high inconsistency. High heterogeneity is not necessarily fatal, but it is a signal to stop and investigate why the studies disagree, often by running subgroup analyses (for example, splitting by dose or by population) to find the source.
Publication bias: are the missing studies skewing you?
Studies with exciting positive results get published more readily than studies that found nothing. If half the “no effect” studies are sitting unpublished in researchers’ file drawers, the studies you can find will overstate the true effect.
You probe for this with a funnel plot, which is a scatter of studies by size and effect that should look like a symmetric inverted funnel if nothing is missing (asymmetry suggests absent studies), and with formal tests such as Egger’s test.
Synthesis computes heterogeneity statistics automatically alongside every pooled estimate, and generates funnel plots and the standard publication-bias tests from the same dataset, so these checks are part of the workflow rather than an afterthought.
Stage 7: Rate the certainty and write the manuscript
GRADE: how confident should a reader be?
Producing a pooled number is not the end. You then rate the overall certainty of the whole body of evidence using a system called GRADE, which assigns one of four levels: High, Moderate, Low, or Very Low. You start high for randomized trials and downgrade for problems you already measured: serious risk of bias, high heterogeneity, imprecision, or evidence of publication bias. The GRADE rating is what tells a guideline committee how much to trust your conclusion.
PRISMA: reporting it so others can reproduce it
Finally you write up the review following PRISMA, the reporting standard for systematic reviews. PRISMA is essentially a checklist plus a flow diagram (the screening funnel from Stage 3) that ensures you have transparently reported what you searched, what you included and excluded, and how you analyzed it. Journals expect PRISMA compliance; a review that cannot be reproduced from its own methods section will struggle to pass review.
Because every stage lived in one tool, Synthesis can assemble the manuscript for you: it generates the PRISMA flow diagram from your actual screening counts, the GRADE summary, the forest and funnel plots, and the evidence tables, then drafts the manuscript text around them. You export to Word, PDF, or other formats ready for journal submission.
Working as a team in Synthesis
Real reviews are team efforts, and the integrity of the method depends on people doing distinct jobs. Synthesis supports multiple collaboration roles so that responsibilities, and the audit trail, stay clean.
Why roles matter
The dual-blinded screening above only works if the two screeners are genuinely different people with their own accounts, and if a third person resolves conflicts. The same logic applies to extraction, where two independent extractions catch transcription errors. Roles enforce this separation rather than relying on good intentions.
A practical team workflow
- The lead defines the protocol and PICO, and invites the team.
- Two reviewers are assigned to screen each record independently and blind.
- A senior reviewer resolves any screening disagreements.
- Two extractors independently extract data from included studies; disagreements are reconciled into an agreed consensus record.
- A statistician or the lead runs the pooled analysis and the two checks.
- The team appraises risk of bias and agrees the GRADE rating.
- The lead generates and edits the manuscript, with the team reviewing before submission.
Synthesis provides distinct collaboration roles with appropriate permissions, a consensus view that shows your own extraction against the agreed team version, and a complete audit trail of who did what and when. Because everyone works in the same platform, there are no emailed spreadsheets to reconcile and no broken provenance.
Your end-to-end checklist
Use this as a single reminder of the whole journey from question to submission.
| # | Step | Done when… |
|---|---|---|
| 1 | Frame PICO and write protocol | Question is specific and the plan is registered |
| 2 | Build and run the search | All target databases searched, results imported and de-duplicated |
| 3 | Title/abstract screen | Two blind reviewers done, conflicts resolved |
| 4 | Full-text screen | Inclusion criteria applied; included set finalized |
| 5 | Extract data | Both arms captured per outcome, provenance recorded |
| 6 | Assess risk of bias | Every study rated with justification |
| 7 | Run the meta-analysis | Model chosen, effect sizes pooled, forest plot produced |
| 8 | Run the two checks | Heterogeneity quantified, publication bias assessed |
| 9 | Rate with GRADE | Certainty level assigned and justified |
| 10 | Generate and submit manuscript | PRISMA diagram, plots, tables, and text assembled and exported |
A systematic review is judged on two things: whether its method was rigorous, and whether that method is fully transparent. Synthesis exists to make the rigorous path also the easy path, but the judgement at every decision point is, and must remain, yours. Use the tool to remove the busywork, not to remove your thinking.
Glossary
| Term | Meaning |
|---|---|
| Confidence interval (CI) | The range within which the true value is likely to lie; a wider interval means more uncertainty. |
| Dual / double screening | Two reviewers screening the same records independently and blind to each other. |
| Effect size | A single number expressing a study’s result on a comparable scale (e.g. odds ratio, mean difference). |
| Fixed-effect model | A pooling model assuming all studies share one identical true effect. |
| Forest plot | The standard chart showing each study’s effect and the pooled result as a diamond. |
| Funnel plot | A scatter plot used to detect publication bias through its symmetry. |
| GRADE | A system rating overall certainty of evidence as High, Moderate, Low, or Very Low. |
| Heterogeneity | The degree to which study results disagree beyond chance; measured by I-squared. |
| I-squared (I²) | A percentage quantifying heterogeneity; higher means more inconsistency. |
| MeSH terms | Standardized medical vocabulary tags used to build precise database searches. |
| Meta-analysis | The statistical step of pooling results from multiple studies into one estimate. |
| Odds ratio / Risk ratio | Ratio effect measures for binary outcomes; 1.0 means no difference. |
| PICO | Population, Intervention, Comparison, Outcome: the framework for a focused question. |
| PRISMA | The reporting standard (checklist and flow diagram) for systematic reviews. |
| PROSPERO | The public registry where review protocols are pre-registered. |
| Protocol | The pre-declared written plan for the whole review, locked before seeing results. |
| Publication bias | Distortion caused by positive studies being published more than null ones. |
| Random-effects model | A pooling model allowing the true effect to vary between studies. |
| Risk of bias | A structured judgement of how trustworthy an individual study’s conduct was. |
| Standard deviation (SD) | A measure of spread of values around the mean within a study. |
| Standardized mean difference | A continuous effect size used when studies measured on different scales. |
| Systematic review | The full rigorous process of finding, appraising, and summarizing all evidence on a question. |