How to Choose Outcomes for an Evidence Table: Quantitative vs Qualitative Reviews

Choosing outcomes for an evidence table is a synthesis-design choice, not a list of everything a paper reports. A quantitative review needs outcomes that are comparable across studies: the same construct, a compatible measure, a compatible timepoint, and the same direction of effect. A qualitative review needs outcomes that are conceptually coherent: findings that describe the same phenomenon, so they can be brought into one synthesis. In either design, keep an outcome only when it can move the decision the review is there to support. The failure mode is a full table that cannot be synthesised.
Cochrane Handbook Chapter 3 (section 3.2.4) requires review authors to define critical and important outcomes in advance, including the outcome domain and the measure, and to prespecify time points. Outcomes that are not critical or important to users of the review can be left out. Where a core outcome set exists, COMET is the place to look it up. COMET defines a core outcome set as an agreed minimum to measure and report in a health area, and says researchers still study other outcomes as well.
The grid those outcomes sit in is an evidence table. Column design is covered in What is an evidence table. For how the decision drives the columns, see Analysis-driven design of evidence tables.
Same question, two ways
Question (synthetic). In adults with type 2 diabetes starting a weekly injectable, what is the effect of adding a smartphone dose reminder, compared with usual care, at about six months — and why do people keep or drop the reminder?
Studies A–D below are illustrations. They are not published trials.
Quantitative: same measure, timepoint, and direction
The critical benefit outcome is change in HbA1c from baseline, in percentage points, with lower values as benefit. The window is 24 weeks, accepting 20–28 weeks, and only one result per study from that window. A harm outcome (severe hypoglycaemia) belongs in the set as well: Chapter 3 says critical outcomes should normally include a benefit and a harm. It is a separate row, not a second number stuffed into the HbA1c cell.
| Study | Measure | Timepoint | Direction | Result | Synthesis decision |
|---|---|---|---|---|---|
| A | HbA1c change from baseline, percentage points | 24 weeks | Lower is benefit | Mean difference −0.4 (95% CI −0.6 to −0.2) | Pool. Same construct, measure, window, and direction. |
| B | Fasting plasma glucose, mmol/L | 4 weeks | Lower is benefit | Mean difference −0.3 (95% CI −0.8 to 0.2) | Do not pool with A. Different measure and different timepoint. |
Chapter 3, section 3.2.4.3, says to specify the domain by the construct ("what" is measured) before specifying the tool and the timing. Study B can sit in its own domain or in a narrative paragraph. It cannot share a meta-analysis with Study A.
The full table that cannot be synthesised. One column headed "glycaemia" containing both −0.4 percentage points and −0.3 mmol/L. Every study has a number. No mean difference can be calculated, because the units and the timepoints are not the same. That column fails the review even though extraction looks complete.
Qualitative: one phenomenon
Cochrane Handbook Chapter 21 treats the phenomenon of interest as part of the qualitative question (section 21.5). Here the phenomenon is why adults keep or drop the reminder, not "anything patients said".
| Study | Finding as reported | Same phenomenon? | Synthesis decision |
|---|---|---|---|
| C | People muted the prompt because it arrived during work meetings. | Yes. Barrier to continued use. | Eligible for a finding about stopping the reminder. |
| C | One person stopped after a hypoglycaemia episode they linked to an extra prompted dose. | Yes. Same phenomenon, different reason. | Eligible. A second theme under the same phenomenon, not a new outcome. |
| D | Participants said clinic parking was expensive. | No. Access cost, not reminder use. | Out of this synthesis. Park it under access, or drop it. |
A coherent qualitative outcome is one phenomenon, described so that each contributing finding is about that phenomenon. Study D is a real report from participants. Forcing it into "why people stop the reminder" produces a theme that no longer answers the question. The column is full, and the synthesis finding is not.
A filled-in selection worksheet
Two candidates for the question above. The first stays. The second is the column teams add because a paper mentioned it.
| Worksheet question | HbA1c change at about 6 months | "Clinical improvement" |
|---|---|---|
| What decision does this support? | Whether adding the reminder changes glycaemic control enough to matter for the care pathway. | None. The label does not name a decision. |
| Quantitative, qualitative, or both? | Quantitative. | Unspecified. Papers use the phrase for symptoms, labs, and clinician impression. |
| Exact definition and measurement rule? | Change from baseline in HbA1c, percentage points, laboratory assay. Fasting glucose is not a substitute. | No shared definition. |
| Required timepoint? | 24 weeks. Accept 20–28 weeks. One result per study in that window. | None stated. |
| Expected synthesis method? | Mean-difference meta-analysis if at least two studies share measure, window, and direction. Otherwise narrative, labelled as not pooled. | Cannot be prespecified. Pooling would mix constructs. |
| Risk if missing? | The review cannot answer the glycaemic decision. Do not backfill with Study B. | Low. Leaving it out avoids a column that cannot be synthesised. |
| Keep? | Yes. Critical benefit outcome. Check COMET for a diabetes core outcome set and treat HbA1c as a minimum if that set lists it, not as the only outcome allowed. | No. |
Qualitative candidate, same question:
| Worksheet question | Barriers to continued reminder use | Clinic parking cost |
|---|---|---|
| What decision does this support? | Whether non-use of the reminder explains a small or null HbA1c effect, and what would have to change in the prompt. | Whether parking affects clinic attendance. That is a different decision. |
| Quantitative, qualitative, or both? | Qualitative. | Qualitative, but a different phenomenon. |
| Exact definition? | Participant-reported reasons for muting, ignoring, or stopping the reminder. | Money or time spent parking at the clinic. |
| Required timepoint? | During the period the reminder was offered. | Not tied to the reminder. |
| Expected synthesis method? | Thematic synthesis of findings about keeping or dropping the reminder. Chapter 21 also lists framework synthesis and meta-ethnography when the finding will sit next to an intervention review. | Not this synthesis. |
| Risk if missing? | The effect estimate has no account of non-use. | Low for this question. |
| Keep? | Yes. | No, unless the review question is about access costs. |
If two extractors both see a 12-week and a 24-week HbA1c in Study A, and each picks a different week, both numbers are in the paper. That is an interpretation mismatch. Resolve it with the ladder in Resolving dual-extraction disagreements: the prespecified window wins, and both source pointers stay on the row.
Quantitative checks
For a quantitative outcome, four checks come before extraction starts:
- Comparability. Can values be compared across studies, or only listed?
- Definition. Is the endpoint the same construct, not the same English word?
- Timepoint. Are the windows compatible, or did Chapter 3's prespecified interval already decide which result to take?
- Effect estimate. Can the reported numbers be turned into the effect measure the synthesis will use? Chapter 6 is the conversion reference (for example, a standard error from a confidence interval). It does not rescue a column that mixed different constructs.
Strong candidates: mortality or event outcomes with a stated definition and denominator; a validated scale with a declared transformation; resource use with one unit and one time base.
High-risk candidates: many incompatible instruments under one heading; "clinical improvement" with no definition; baseline and follow-up values with no fixed window.
If an outcome fails comparability, report it narratively. Do not pool it to keep the forest plot full.
Qualitative checks
Chapter 21 describes qualitative evidence synthesis as a way to address stakeholder preferences, experiences, implementation, and context. The outcome is a phenomenon, specified tightly enough that someone else can tell whether a finding belongs.
Useful domains, when they are the decision:
- patient experience and burden
- barriers and facilitators to implementation
- clinician workflow
- system constraints
Fix a small set of phenomena before coding. Subthemes can appear inside a phenomenon. A new phenomenon is a new outcome, with its own decision, or it is out of scope.
Mixed methods
Use shared domains, and keep the data types in their own columns.
- Glycaemic control: HbA1c mean difference. Not a theme score.
- Continued use: a quantitative adherence measure, if one was prespecified, plus the qualitative phenomenon "why people stop the reminder".
- Harm: severe hypoglycaemia counts, plus any qualitative account of hypoglycaemia linked to the prompt.
The qualitative finding can explain a small quantitative effect. It is not a second effect size. Scoring themes as 1–5 so they can be "pooled" with HbA1c drops both the meaning and the analysis.
Common mistakes
Including every reported endpoint
Extraction sprawls, and the synthesis has no hierarchy. Chapter 3 allows more outcomes than the summary table, and still says outcomes that are not critical or important can be left out. Summary-of-findings tables are limited to seven outcomes chosen in advance, not the seven with the largest effects.
No timepoint policy
Late arguments about week 12 versus week 24 are a missing protocol line, not a difficult judgement at extraction.
Mixing constructs in one column
This is the glycaemia failure above. The table looks finished. The synthesis cannot start.
Forcing qualitative themes into pseudo-quantitative scores
A count of papers that "mentioned parking" is not an effect estimate, and it is not a coherent finding about the reminder.
What to keep
A useful evidence table stores the outcomes the synthesis will actually use. Comparable numbers for a quantitative review. One phenomenon at a time for a qualitative review. Anything else waits, or stays out.
Related reading
Tags:

About the Author
Connect on LinkedInGeorge Burchell
George Burchell is a specialist in systematic literature reviews and scientific evidence synthesis with significant expertise in integrating advanced AI technologies and automation tools into the research process. With over four years of consulting and practical experience, he has developed and led multiple projects focused on accelerating and refining the workflow for systematic reviews within medical and scientific research.