Study Design Clinical Research Evidence-Based Medicine 35 min read August 16, 2026
BY: Statistics Fundamentals Team
Reviewed By: Minsa A (Senior Statistics Editor)

Randomized Controlled Trials

A new blood pressure medication is tested against a placebo in 400 patients. Half are randomly assigned to each group without anyone knowing who got what until the data are collected. That structure — intervention, control, random assignment, concealment — is the randomized controlled trial. It is among the most carefully reasoned study designs in clinical and public health research, though not the right tool for every question.

This guide covers what an RCT is, how randomization and blinding work, what the different trial types are, how results are measured, what biases can arise, and how to critically appraise a published RCT. Interactive tools for sample size, effect size, risk, and critical appraisal are included throughout.

What You'll Learn
  • ✓ The definition of a randomized controlled trial and how it differs from observational research
  • ✓ How randomization, allocation concealment, and blinding each serve a different purpose
  • ✓ All major RCT designs: parallel, crossover, cluster, factorial, adaptive, pragmatic
  • ✓ Sample size planning, statistical power, and effect size calculation
  • ✓ Risk ratio, odds ratio, ARR, RRR, NNT, NNH with worked examples
  • ✓ Intention-to-treat vs per-protocol analysis, and when each applies
  • ✓ Bias types: selection, performance, detection, attrition, reporting
  • ✓ CONSORT reporting standards and a full critical appraisal checklist

What Is a Randomized Controlled Trial?

Definition — Randomized Controlled Trial
A randomized controlled trial (RCT) is an experimental study in which eligible participants are randomly allocated to two or more groups — typically an intervention group and a control group — so that the groups can be compared on pre-specified outcomes.
Randomization helps distribute both known and unknown confounders across groups, which can strengthen the basis for causal inference. It does not guarantee that groups are perfectly balanced in any given trial, and it does not address all sources of bias on its own.

The basic logic is comparative: if participants are allocated by chance, then differences in outcomes between groups are more likely to reflect the effect of the intervention rather than pre-existing differences between participants. That is why RCTs are often described as strong designs for evaluating whether an intervention causes an outcome — though the quality of any specific trial depends on how well it was designed, conducted, and analyzed.

Three structural features distinguish an RCT from other study designs. First, there is a researcher-assigned intervention rather than naturally occurring exposure. Second, there is a comparison group. Third, allocation to groups is made by a random process. Additional design choices — blinding, allocation concealment, stratification, and prespecified analysis plans — strengthen the trial's ability to produce unbiased estimates.

3
Core RCT features
α = 0.05
Common significance threshold
80%
Typical power target
CONSORT
Reporting standard
📌
Quick Answer — Featured Snippet

A randomized controlled trial is an experimental study design where participants are randomly assigned to an intervention group or a control group to compare outcomes. Random allocation distributes confounders across groups, supporting causal inference when the trial is well designed and conducted.

Basic Structure of an RCT

A typical parallel-group RCT follows a sequence from eligibility assessment through analysis. Understanding each stage helps with both designing a trial and critically reading one.

1

Define the Research Question

Specify the population, intervention, comparator, and outcome (PICO). The research question should be answerable by randomized comparison and should correspond to a state of genuine uncertainty — clinical equipoise — about the intervention.

2

Set Eligibility Criteria

Inclusion and exclusion criteria define who can participate. They balance internal validity (recruiting a relatively homogeneous population) against external validity (making results generalizable to a broader clinical or policy context). Unnecessarily restrictive criteria limit generalizability without always improving validity.

3

Recruit and Obtain Informed Consent

Participants are screened, and those who meet eligibility criteria are invited to participate. Informed consent must be obtained before randomization. Consent covers the trial's purpose, what participation involves, potential risks and benefits, and the right to withdraw.

4

Randomly Allocate Participants

A pre-generated random sequence determines group assignment. The assignment should be concealed from those enrolling participants until the participant is formally entered into the trial. This is allocation concealment — a separate concept from randomization itself.

5

Administer the Intervention

Participants receive the assigned intervention or control condition. Where feasible, blinding prevents participants, care providers, or outcome assessors from knowing which group a participant is in, reducing performance and detection bias.

6

Follow Up and Measure Outcomes

Participants are followed for the pre-specified time period. The primary outcome, secondary outcomes, and safety outcomes are measured according to the protocol. Minimizing loss to follow-up is important because systematic loss can reintroduce bias.

7

Analyze and Interpret

Data are analyzed according to the pre-specified statistical analysis plan. The primary analysis typically uses intention-to-treat. Effect sizes, confidence intervals, and p-values are reported, along with a consideration of clinical significance alongside statistical significance.

Randomization

Randomization is the process of using a chance mechanism to determine which group each participant is assigned to. Its main purpose is to distribute both known and unknown confounders across the comparison groups without the investigator selecting who goes where.

⚠️
Common Misconception

Randomization does not guarantee that groups are identical at baseline. In smaller trials, chance imbalances can occur even with proper randomization. Stratification, block randomization, and covariate adjustment are tools that can reduce this risk.

Methods of Randomization

MethodDescriptionStrengthLimitation
Simple randomizationEach participant assigned independently (e.g., coin flip)Unpredictable sequenceGroups may differ in size, especially in small trials
Block randomizationParticipants allocated in blocks to ensure balanced group sizes at intervalsMaintains balance throughout recruitmentIf block size is known, the final assignment in a block may be predictable
Stratified randomizationRandomization performed separately within predefined strata (e.g., age, sex, disease severity)Ensures balance on important prognostic factorsComplexity increases with many stratification variables
MinimizationAdaptive procedure that assigns each new participant to the group that minimizes imbalance on specified factorsAchieves balance on multiple factors simultaneouslyPartly deterministic; some debate on whether it constitutes strict randomization
Cluster randomizationGroups (clinics, schools, villages) rather than individuals are randomizedPractical when individual randomization is not feasibleRequires adjustment for clustering in analysis; reduces effective sample size

Randomization vs Allocation Concealment

These two concepts are distinct and both matter. Confusing them is a recognized methodological error in clinical research.

ConceptWhat it doesWhen it actsBias it addresses
RandomizationDetermines which group a participant will enter, using a chance mechanismWhen generating the allocation sequence (before recruitment)Confounding at baseline
Allocation concealmentPrevents those enrolling participants from knowing the upcoming assignment until a participant is committed to the trialDuring enrollmentSelection bias from selective enrollment

Inadequate allocation concealment allows an investigator who knows the next assignment is "drug" to delay enrolling a sicker patient until a "placebo" assignment comes up. This can introduce systematic differences between groups that randomization alone would not prevent. Sealed opaque envelopes, central telephone or web-based allocation systems, and pharmacy-controlled assignment are common concealment methods.

Control Groups and Blinding

What Is a Control Group?

The control group provides the comparison against which the intervention group is evaluated. The choice of control is a scientific and ethical decision that depends on the research question. Common control conditions include usual care (what participants would receive in standard clinical practice), placebo (an inert or sham treatment that mirrors the intervention without the active component), active comparator (a different treatment rather than no treatment), no intervention, and wait-list control.

Placebo-controlled trials are useful when there is no established effective treatment, when the placebo effect is particularly relevant to the outcome, or when withholding treatment does not expose participants to unacceptable risk. They are not automatically superior to active-comparator trials, and in conditions where effective treatments already exist, a placebo control may raise ethical concerns.

Blinding and Masking

Blinding — also called masking — refers to withholding knowledge of group assignment from specified parties in the trial. The goal is to prevent that knowledge from influencing outcomes or their measurement.

Who is blindedBias it addressesPractical example
ParticipantsReduces placebo effects, co-intervention differences, and differential attritionIdentical capsules for drug and placebo groups
Care providersReduces performance bias — differential management of participants based on groupPharmacist dispenses coded medication without label identifying drug or placebo
Outcome assessorsReduces detection bias — differential measurement or adjudication of outcomesRadiologist reading CT scans does not know group assignment
Data analystsReduces reporting bias during analysis stageAnalysis completed before unblinding the group codes
⚠️
On the Labels "Single-Blind" and "Double-Blind"

The terms single-blind, double-blind, and triple-blind are not uniformly defined across the literature. "Double-blind" sometimes means participants and investigators are blinded, sometimes participants and outcome assessors. Trial reports should specify exactly which parties were blinded rather than relying on these labels alone.

Types of Randomized Controlled Trials

Different RCT designs address different research questions. Choosing the appropriate design depends on the intervention, the target population, ethical constraints, feasibility, and what question the trial is meant to answer.

Most Common
↔️

Parallel-Group RCT

Each participant is randomized to one group and remains in that group throughout the study. This is the simplest and most frequently used design. Treatment groups are compared at the end of the follow-up period.

Within-Subject
🔄

Crossover RCT

Participants receive multiple interventions in sequence, separated by a washout period. Each participant serves as their own control. Requires stable conditions, an adequate washout to avoid carryover effects, and careful handling of period effects. Not suitable when the intervention has lasting effects.

Group-Level
🏥

Cluster Randomized Trial

Groups — clinics, schools, practices, communities — rather than individuals are randomized. Used when individual randomization is not feasible or when contamination between participants is likely. Analysis must account for intracluster correlation and the resulting design effect.

Multi-Arm

Factorial RCT

Participants are randomized to combinations of two or more interventions. For a two-by-two design, groups would receive A alone, B alone, A and B together, or neither. Efficient when interactions between interventions are of interest or when additive effects can be assumed.

Real-World
🌍

Pragmatic RCT

Designed to evaluate effectiveness — how well an intervention performs in routine clinical or public health practice. Broad eligibility criteria, flexible delivery, and practical outcomes reflect what would happen in real-world application.

Controlled Conditions
🔬

Explanatory RCT

Designed to evaluate efficacy — does the intervention work under ideal conditions? Narrower eligibility criteria, strict protocol adherence, and careful outcome assessment answer the mechanistic question of whether the intervention can work.

Flexible Design
📊

Adaptive Trial

Allows pre-specified modifications to the trial based on accumulating data. Adaptations can include sample size re-estimation, dropping treatment arms, or adjusting randomization ratios. Requires rigorous statistical planning to maintain Type I error control.

Non-Inferiority
↕️

Non-Inferiority Trial

Tests whether a new treatment is not worse than a comparator by more than a pre-specified margin. The margin must be clinically justified. Analysis populations and confidence interval placement around the margin are critical for correct interpretation. "No significant difference" in a superiority trial does not demonstrate non-inferiority.

RCT design typology adapted from: Friedman LM, Furberg CD, DeMets DL, Reboussin DM, Granger CB. Fundamentals of Clinical Trials (5th ed). Springer, 2015. Also see NCBI Bookshelf — Clinical Trials.

Sample Size and Statistical Power

The sample size determines whether a trial can reliably detect an effect if one exists. An underpowered trial may fail to find a real difference — not because the intervention does not work, but because too few participants were enrolled. Pre-study sample size planning is required for any ethically and scientifically sound RCT.

Statistical Power

Power is the probability that the test will correctly reject the null hypothesis when the null hypothesis is actually false. Power = 1 − β, where β is the Type II error rate. Most trials target 80% power (β = 0.20) or 90% power (β = 0.10).

Power and Sample Size Relationships
Power = 1 − β  |  n ↑ → Power ↑
α = significance level (e.g., 0.05) β = Type II error rate δ = expected effect size σ = standard deviation

Sample size depends on the primary outcome type, the minimum clinically meaningful effect size, variability in the outcome, the chosen significance level (commonly α = 0.05, though not universally required), the desired power, the allocation ratio between groups, and expected attrition. These are common choices, not universal requirements — specific contexts may call for different values.

RCT Sample Size Calculator

Sample Size Calculator

For educational use only. Verify results with validated statistical software for actual research planning. Calculations run locally in your browser — no data is sent anywhere.

n per group (adjusted)
Total n (adjusted)
Cohen's d
n per group (adjusted)
Total n (adjusted)
ARR
NNT

How RCT Results Are Measured

The choice of effect measure depends on the outcome type and the research question. Each measure answers a slightly different question about the magnitude and direction of the intervention's effect.

Continuous Outcomes

Mean Difference
MD = x̄₁ − x̄₀
x̄₁ = intervention group mean x̄₀ = control group mean

The mean difference (MD) is expressed in the original units of measurement and is directly interpretable when groups used the same scale. When combining outcomes across studies using different scales, the standardized mean difference (SMD, Cohen's d) can be used, though it is harder to interpret clinically.

Binary Outcomes and Risk Measures

Key Risk and Ratio Measures
Risk = events / total  |  RR = Risk₁ / Risk₀  |  ARR = Risk₀ − Risk₁
RR = Risk Ratio (Relative Risk) ARR = Absolute Risk Reduction RRR = ARR / Risk₀ NNT = 1 / ARR
MeasureFormulaNull valueInterpretation
Risk Ratio (RR)Risk₁ / Risk₀RR = 1How many times more likely the event is in the intervention vs. control group
Odds Ratio (OR)(a/b) / (c/d) from 2×2 tableOR = 1Ratio of odds of event; approximates RR when events are rare; not the same as RR when events are common
ARRRisk₀ − Risk₁ARR = 0Absolute reduction in event rate; reflects absolute benefit regardless of baseline risk
RRRARR / Risk₀RRR = 0Proportional reduction in risk; can appear large even when ARR is small
NNT1 / ARRNumber of patients needing treatment for one additional good outcome; requires time horizon specification
NNH1 / ARI (absolute risk increase)Number treated for one additional adverse event; must be reported alongside NNT
⚠️
RR vs OR — a common confusion

The odds ratio is not the same as the risk ratio. When the outcome is common (event rate above roughly 10%), the OR can substantially overestimate the magnitude of the RR. Always check which measure is being reported and whether the OR is being interpreted as if it were an RR.

RCT Risk and Effect Size Calculator

Risk, NNT, and Effect Measures

Enter event counts or rates for each group. All calculations are performed locally in your browser.

Risk Ratio
Odds Ratio
ARR
RRR
NNT
Risk (Intervention)
Risk (Control)

Confidence Intervals and P-Values

Confidence Intervals in RCTs

A confidence interval quantifies the precision of an estimate. A 95% confidence interval means that if the same study were conducted many times under the same conditions, approximately 95% of the intervals constructed would contain the true population parameter. This is a property of the procedure, not a statement that there is a 95% probability that this specific interval contains the true value — the parameter is fixed, not random.

For difference measures (mean difference, ARR), the null value is zero. For ratio measures (RR, OR), the null value is one. When a 95% CI excludes the null value, the result is statistically significant at α = 0.05. When a CI is wide, the estimate is imprecise regardless of the p-value. A narrow CI around a small effect is still a small effect.

P-Values

The p-value is the probability, under the null hypothesis, of obtaining a test statistic as extreme as or more extreme than what was observed. A small p-value means the observed result is unlikely if H₀ were true — not that H₀ is false or that the effect is large.

🚫
What "statistically significant" does not mean

Statistical significance (p < 0.05) does not mean the effect is clinically important, that the treatment works in all patients, that the null hypothesis is false, or that the result will replicate. A large trial can find a statistically significant but clinically trivial effect. Effect size and confidence intervals are essential context for any p-value.

Statistical vs Clinical Significance

A result can be statistically significant without being clinically meaningful, and clinically meaningful without reaching statistical significance. The distinction matters enormously in practice.

SituationStatistical significanceClinical significanceWhat to conclude
Large trial, tiny effectp < 0.001Low — effect smaller than MCIDReal but trivial effect — unlikely to matter to patients
Small trial, meaningful effectp = 0.18High — effect exceeds MCIDTrial underpowered; effect plausible but unconfirmed
Adequate trial, meaningful effectp = 0.02High — effect exceeds MCIDBoth statistically and clinically significant — strong evidence
Large trial, null resultp = 0.62Low — narrow CI excludes MCIDEvidence of absence of meaningful effect

MCID = minimal clinically important difference. For p-value interpretation and confidence interval construction, see the dedicated guides at Statistics Fundamentals.

Intention-to-Treat and Per-Protocol Analysis

Intention-to-Treat Analysis

Intention-to-treat (ITT) analysis includes all randomized participants in the group to which they were originally assigned, regardless of whether they completed the intervention, deviated from the protocol, or were found post-randomization to be ineligible. ITT preserves the benefits of randomization: if allocation was successful, any confounders that randomization distributed across groups remain distributed across groups in the analysis, even if participants switched groups.

ITT gives an estimate of the effect of being assigned to the intervention — the policy-relevant question in most settings. It tends to be conservative (attenuating the effect toward null when non-adherence is present) rather than anticonservative. Missing data handling is a genuine challenge in ITT analyses and should be addressed through pre-specified imputation or sensitivity analysis strategies.

Per-Protocol Analysis

Per-protocol analysis restricts the analysis to participants who completed the trial in accordance with the protocol. It can help answer whether the intervention works in those who actually received it, but it reintroduces potential confounding because the subgroup who adheres may differ systematically from those who do not.

💡
ITT and Per-Protocol in Non-Inferiority Trials

In non-inferiority trials, both ITT and per-protocol analyses are relevant. An ITT analysis that dilutes the effect toward null could make an inferior treatment appear non-inferior. Both analyses are typically expected, and the non-inferiority conclusion should be consistent across both.

Bias in Randomized Controlled Trials

Randomization reduces but does not eliminate all sources of bias. Understanding each bias type helps both in designing trials to minimize them and in appraising published trials.

Selection Bias

Arises from how participants are recruited or allocated to groups. If high-risk participants are systematically directed to one arm because the upcoming assignment is known, baseline differences emerge.

✓ Addressed by: allocation concealment

Performance Bias

Occurs when participants or care providers change their behavior based on knowledge of group assignment — for example, the intervention group receiving more attention or co-interventions.

✓ Addressed by: blinding of participants and providers

Detection Bias

Results from differences in how outcomes are measured or assessed between groups. An unblinded assessor may record outcomes differently depending on which group a participant is in.

✓ Addressed by: blinding of outcome assessors

Attrition Bias

Systematic differences between groups in who withdraws or is lost to follow-up. If drop-outs occur for reasons related to the treatment or its side effects, and this differs between arms, the remaining sample no longer reflects the randomized population.

✓ Addressed by: ITT analysis, minimizing loss to follow-up

Reporting Bias

Selective reporting of outcomes — publishing only those that are statistically significant while leaving others unreported — distorts the evidence base. Occurs at the trial level and is distinct from publication bias.

✓ Addressed by: trial registration, pre-specified outcomes

Does Randomization Eliminate Confounding?

Randomization distributes confounders — both measured and unmeasured — across groups by a chance process. In expectation, the groups will be similar on all baseline characteristics. In any finite sample, however, particularly a small one, chance imbalances can occur. This is not a failure of the randomization process but a consequence of small sample variability.

Strategies to reduce the probability of chance imbalance include stratified randomization (randomizing within subgroups of important prognostic variables), block randomization (ensuring approximately equal group sizes throughout recruitment), and minimization. Even after randomization, covariate adjustment in the analysis — recommended in the statistical analysis plan — can improve precision and further reduce any residual imbalance.

Internal and External Validity

Internal Validity

Internal validity refers to the degree to which the trial's results reflect a true causal effect of the intervention, rather than bias or confounding. A well-conducted RCT with proper randomization, allocation concealment, blinding, complete follow-up, and intention-to-treat analysis can have strong internal validity. Each element addresses a different threat to valid causal inference.

External Validity and Generalizability

External validity — also called generalizability or applicability — concerns whether the results apply to patients or settings beyond those in the trial. Even a methodologically rigorous trial may not generalize if its participants differ substantially from those in clinical practice, if the intervention was delivered under unusually controlled conditions, or if the comparator does not reflect current standard care.

Pragmatic trials are designed to maximize external validity; explanatory trials prioritize internal validity. The difference between efficacy (does it work under ideal conditions?) and effectiveness (does it work in real-world practice?) is partly a question of where on this spectrum the trial sits.

RCTs vs Other Study Designs

Feature RCT Cohort Study Case-Control Study
Intervention assignmentResearcher-assigned, randomizedSelf-selected or naturally occurringDetermined after outcome
Direction in timeProspectiveProspective (or retrospective)Retrospective
Confounding controlRandomization distributes confoundersStatistical adjustment onlyMatching and statistical adjustment
Causal inferenceStrongest basis when well designedModerate; residual confounding possibleWeaker; selection and recall bias risks
Rare outcomesRequires large sample; costlyPossible with long follow-upWell-suited
Ethical limitsCannot randomize to harmful exposuresCan study harmful exposuresCan study harmful exposures
Typical cost and timeHighModerate to highModerate
Effect measureRR, ARR, MDRR, HROR (approximates RR for rare outcomes)

For a full treatment of observational study designs, see the study design section. For further reading on correlation vs causation, see correlation vs causation.

Advantages and Limitations

Advantages of RCTs

When well designed and conducted, RCTs offer several methodological strengths relative to observational research. Randomization distributes confounders, including unmeasured ones, in a way that statistical adjustment alone cannot guarantee. Prospective data collection reduces recall bias and allows outcomes to be measured consistently. Pre-specification of the primary outcome and analysis plan reduces the opportunity for selective reporting. The controlled setting allows the effect of a single well-defined intervention to be evaluated in isolation.

Limitations of RCTs

RCTs are not the right tool for every question. Several categories of limitation deserve consideration.

LimitationWhy it matters
Ethical constraintsRandomization cannot be applied when it would expose participants to known harm — researchers cannot randomize people to smoking to study lung cancer
Cost and timeLarge multicentre trials can take years and cost tens of millions of dollars, limiting feasibility for many interventions
GeneralizabilityRestrictive eligibility criteria produce homogeneous study populations that may differ from real-world patients
Hawthorne effectParticipants may change their behavior because they are being observed, affecting both arms
Rare or delayed outcomesEvents that occur in fewer than 1% of patients or take decades to appear require very large or very long trials
Complex interventionsBehavioral, educational, or multi-component interventions are difficult to standardize and blind
External validityTightly controlled efficacy trials may not predict what happens when an intervention is delivered in routine care
Attrition and non-adherenceDropout and protocol deviations are common in long trials and must be handled carefully in analysis

Ethical Considerations

Ethical standards in RCTs exist to protect participants. Several principles deserve attention.

Clinical equipoise means that there must be genuine uncertainty in the expert clinical community about which treatment is better. Without equipoise, randomizing participants would mean exposing some to an inferior treatment. Establishing equipoise is a prerequisite for ethical randomization, though it does not make all specific design choices automatically acceptable.

Informed consent requires that participants understand the purpose of the study, what participation involves, the risks and potential benefits, alternatives to participation, and their right to withdraw at any time without penalty. Obtaining consent must occur before randomization.

Independent ethics committee or IRB review is required before a trial begins. The committee evaluates whether the scientific rationale, design, consent procedures, and participant protections meet accepted ethical standards.

Data monitoring committees review accumulating data during a trial to identify safety signals or evidence of clear efficacy or futility. Pre-specified stopping rules govern whether and when the trial should be terminated early.

Trial registration in a public registry before participant enrollment is now expected by major journals and required by many ethics bodies. Registration creates a public record of the trial's hypotheses, design, and primary outcomes, which reduces selective reporting.

CONSORT and Trial Reporting

CONSORT (Consolidated Standards of Reporting Trials) is a reporting guideline that specifies what information should be included in a published RCT report. It exists because incomplete and unclear reporting makes it difficult for readers to assess whether a trial was conducted properly and what its results actually mean.

CONSORT includes a checklist of reporting items covering the title and abstract, background, methods (randomization, blinding, outcomes, sample size), results (participant flow, recruitment, baseline data, outcomes, harms), discussion, and other information (registration, funding). The CONSORT flow diagram is a particularly useful element: it shows how many participants were assessed, excluded, randomized, lost to follow-up, and included in the analysis.

CONSORT Participant Flow

CONSORT-Style Flow Diagram (Parallel-Group RCT)
Assessed for Eligibility (n = X)
Excluded (n = X)
Did not meet criteria / declined / other
Randomized (n = X)
↓                        ↓
Allocated to Intervention (n = X)
Received intervention (n = X)
Did not receive (n = X, reason)
Allocated to Control (n = X)
Received control (n = X)
Did not receive (n = X, reason)
↓                        ↓
Follow-Up (n = X)
Lost to follow-up (n = X, reason)
Discontinued (n = X, reason)
Follow-Up (n = X)
Lost to follow-up (n = X, reason)
Discontinued (n = X, reason)
↓                        ↓
Analysed — Intervention (n = X)
Excluded from analysis (n = X, reason)
Analysed — Control (n = X)
Excluded from analysis (n = X, reason)
Exact reporting structure should follow current CONSORT guidance. The diagram above illustrates the general flow; actual numbers and reasons must be specified from the trial data.

Worked Examples

All examples below use hypothetical educational data. No real trial results are fabricated or implied.

Example 1 — Calculating Risk Ratio and ARR

Worked Example 1 — Risk Ratio and ARR

Problem: A parallel-group RCT tests a new antibiotic (n = 150 per group) for post-operative infection. In the intervention group, 18 of 150 develop infection. In the control group, 36 of 150 do.

1

Calculate risks: Risk₁ = 18/150 = 0.12 (12%). Risk₀ = 36/150 = 0.24 (24%).

2

Risk Ratio: RR = 0.12 / 0.24 = 0.50. The intervention group has half the risk of infection compared to control.

3

ARR: ARR = 0.24 − 0.12 = 0.12 (12%). Twelve fewer infections per 100 patients treated.

4

RRR: RRR = 0.12 / 0.24 = 0.50 (50%). A 50% relative reduction in infection risk.

5

NNT: NNT = 1 / 0.12 = 8.3. Approximately 9 patients need to receive the antibiotic to prevent one infection over the follow-up period. (NNT always requires a specified time horizon and outcome.)

✅ The new antibiotic reduced absolute infection risk by 12 percentage points (RR = 0.50, NNT ≈ 9). The relative risk reduction of 50% is more dramatic-sounding than the absolute reduction; reporting both alongside confidence intervals gives the clearest picture.

Example 2 — Sample Size for a Continuous Outcome

Worked Example 2 — Sample Size, Continuous Outcome

Problem: Researchers plan a trial of a dietary intervention on systolic blood pressure. They expect the intervention to reduce SBP by 8 mmHg (SD = 15 mmHg). Using α = 0.05 two-tailed and 80% power, with equal group sizes and 10% expected attrition — how many participants are needed?

1

Effect size: Cohen's d = δ/σ = 8/15 = 0.533. A medium effect.

2

Z-values: For α = 0.05 two-tailed: z_α/2 = 1.96. For 80% power: z_β = 0.842.

3

Base n per group: n = 2 × (z_α/2 + z_β)² / d² = 2 × (1.96 + 0.842)² / (0.533)² ≈ 2 × 7.85 / 0.284 ≈ 55 per group.

4

Attrition adjustment: n_adjusted = 55 / (1 − 0.10) = 62 per group → Total = 124 participants.

✅ The trial needs approximately 62 participants per group (124 total) to detect an 8 mmHg difference with 80% power at α = 0.05, accounting for 10% attrition. Verify with validated statistical software before finalizing any protocol.

Example 3 — Interpreting a Crossover Trial Result

Worked Example 3 — Crossover RCT

Problem: A crossover trial tests two analgesics for chronic back pain in 40 participants. Each participant receives Drug A first, then Drug B (or vice versa), with a 2-week washout. Mean pain score (0–10) on Drug A = 4.1, on Drug B = 5.6. Paired t-test: t = 3.2, p = 0.003, 95% CI for difference = [0.56, 2.44].

1

Design check: Crossover is appropriate only if the condition is stable. Chronic back pain qualifies. Washout of 2 weeks should be adequate if the analgesic's effects do not persist beyond that.

2

Effect: Drug A achieves a mean score 1.5 points lower than Drug B (4.1 vs 5.6). The 95% CI [0.56, 2.44] excludes zero, confirming the difference is statistically significant.

3

Clinical significance: A 1.5-point difference on a 0–10 pain scale is typically considered clinically meaningful (commonly cited MCID ≈ 2 points, though this varies). The lower bound of the CI (0.56) falls below a common MCID threshold — the uncertainty range includes effects that may not be clinically important.

4

Carryover concern: If participants in one sequence (A then B) showed a different treatment effect than those in the other sequence (B then A), a carryover effect from the first period may be distorting results. This should be tested statistically.

✅ Drug A produced a statistically significant pain reduction versus Drug B. The effect size (1.5 points) is plausibly clinically meaningful, though the confidence interval does not rule out effects below a typical MCID. Carryover effects and period effects should be assessed before drawing firm conclusions.

Critical Appraisal of an RCT

Critical appraisal asks: was this trial designed, conducted, and reported in a way that allows me to trust its results? The following questions address the main methodological concerns.

DomainQuestion to askWhy it matters
RandomizationWas the random sequence generated by an appropriate method?Inadequate sequence generation leaves allocation predictable
Allocation concealmentWere upcoming assignments concealed from those enrolling participants?Unconcealed allocation enables selective enrollment
BlindingWho was blinded? Were methods adequate?Unblinded parties may behave differently or assess outcomes with bias
AttritionHow many dropped out? Were reasons reported? Were they balanced across groups?Differential dropout reintroduces confounding
ITTWere all randomized participants analyzed in their assigned group?Departures from ITT can bias estimates
OutcomesWas the primary outcome pre-specified? Were all outcomes reported?Post hoc selection inflates false positive rates
Effect sizeWhat is the magnitude and direction of the effect? Is it clinically meaningful?Statistical significance is not sufficient; clinical relevance matters
Confidence intervalsAre confidence intervals reported? Are they narrow enough to be informative?Wide CIs indicate imprecision regardless of p-value
HarmsWere adverse events reported systematically?Benefits without harms data give an incomplete risk-benefit picture
GeneralizabilityAre the participants and setting similar to the target population?Efficacy results may not translate to effectiveness in different settings

Interactive RCT Quality Checklist

RCT Critical Appraisal Checklist

items met out of 14

This checklist is an educational aid, not a validated risk-of-bias tool. For formal appraisal, use the Cochrane RoB 2 tool or equivalent validated instrument.

Is an RCT Right for Your Research Question?

RCT Suitability Decision Tool

Is there a specific intervention you want to evaluate?

✅ An RCT appears appropriate for this research question. Proceed with careful protocol development, ethics review, trial registration, and consultation with a statistician and trial methodologist before finalizing the design.

⚠️ An RCT may face practical or ethical constraints. Consider whether a modified design (cluster RCT, pragmatic trial, stepped-wedge) or a high-quality observational design might address the question more feasibly. Consult a methodologist.

❌ An RCT is likely not the right design for this question. Observational designs (cohort study, case-control study) or qualitative research may be more appropriate. This tool is not a substitute for ethics review or formal study-design consultation.

Practice Questions

Practice 1

A trial reports RR = 0.70 (95% CI: 0.45 to 1.09). What does this tell you?

The intervention is associated with a 30% relative reduction in risk compared to control. However, the 95% CI includes 1.0 (the null value for RR), meaning the result is not statistically significant at α = 0.05. The wide interval reflects imprecision — the true effect could range from a 55% reduction to a 9% increase. The trial may be underpowered.
Practice 2

What is the difference between randomization and allocation concealment? Why does confusing them matter?

Randomization generates a chance-based sequence that determines which group participants enter. Allocation concealment prevents those enrolling participants from knowing the next upcoming assignment before a participant is formally enrolled. A trial can have a valid random sequence but poor concealment (e.g., a visible list) allowing selective enrollment. Confusing the two leads to underestimating selection bias risk in trials with poor concealment but described randomization.
Practice 3

In a crossover trial, what is a carryover effect and why is it a problem?

A carryover effect occurs when the effect of the first intervention persists into the second period, contaminating the measurement of the second intervention. This violates the assumption that observations in each period are independent of prior treatment. It is addressed by an adequate washout period between interventions, though determining the required washout length can be difficult, especially for interventions with delayed or uncertain duration of action.
Practice 4

Control event rate = 0.40, intervention event rate = 0.28. Calculate ARR, RRR, and NNT.

ARR = 0.40 − 0.28 = 0.12 (12%). RRR = 0.12 / 0.40 = 0.30 (30%). NNT = 1 / 0.12 = 8.3 ≈ 9. Approximately 9 patients would need to receive the intervention for one additional patient to avoid the event, at the time horizon of the study.
Practice 5

A trial of 400 participants found a statistically significant result (p = 0.03) for a difference of 0.3 kg weight loss (95% CI: 0.1 to 0.5 kg). Is this clinically significant?

Statistically yes — the CI excludes zero and p < 0.05. Clinically almost certainly not: a 0.3 kg difference in weight is below any accepted MCID for meaningful weight loss interventions and would be imperceptible to patients. This example shows how large samples can produce statistically significant but practically irrelevant findings. Effect size and clinical context must accompany any p-value.
Practice 6

Why should the primary outcome be prespecified rather than selected after seeing the data?

If investigators choose the primary outcome after observing the data, they can select the measure that happens to show the most favorable or significant result. This inflates the false positive rate — the probability that a statistically significant result is actually a chance finding. Pre-specification, combined with trial registration, creates a public record of the original hypothesis that limits outcome switching.

Real-World RCT Applications

💊

Pharmacology

Drug efficacy and safety evaluation before regulatory approval. Typically Phase III superiority or non-inferiority trials with pre-specified primary endpoints.

💉

Vaccines

Efficacy against infection, severe disease, or transmission. Often large cluster or individual randomized trials with blinding of participants and assessors.

🧠

Psychology

Evaluation of behavioral therapies, cognitive interventions, and psychosocial programs. Blinding is often limited; outcome assessors may be blinded even when participants cannot be.

🏥

Surgery

Surgical vs non-surgical comparisons or comparisons of surgical techniques. Full blinding is often not possible; sham-controlled trials exist but raise their own ethical questions.

🥦

Nutrition

Dietary interventions, supplement trials, and weight management programs. Adherence and contamination between groups are common practical challenges.

🏫

Education

Cluster randomized trials of teaching methods, curricula, or school-based programs. Randomization at the school or classroom level; outcomes include academic achievement or behavior.

📱

Digital Health

Apps, wearable-based interventions, and telehealth programs. Recruitment, randomization, and outcome collection can occur remotely; blinding is typically limited.

🌍

Public Health

Community health programs, sanitation interventions, and policy evaluations. Often pragmatic cluster randomized designs with population-level outcomes.

Frequently Asked Questions

An RCT is a study where participants are randomly sorted into groups, typically one that receives a treatment and one that does not, so that any difference in outcomes can more reliably be attributed to the treatment rather than to pre-existing differences between people.

Randomization distributes both measured and unmeasured characteristics across groups by chance rather than by choice. This means the groups start out similar on average, and any difference in outcomes is more likely due to the intervention than to background differences between participants.

The control group provides the comparison against which the intervention is evaluated. Without it, you cannot determine whether any change observed in the intervention group is due to the treatment, natural disease progression, the placebo effect, or other factors occurring simultaneously.

In an RCT, the researcher assigns participants to groups using randomization. In a cohort study, participants self-select their exposure in ordinary life and the researcher observes what happens. RCTs have a stronger basis for causal inference because randomization distributes confounders; cohort studies rely on statistical adjustment for measured confounders only.

A well-conducted RCT provides strong evidence for a causal effect of the assigned intervention, but causation is a conclusion that requires judgment about the entire body of evidence rather than a single trial. Factors like biological plausibility, consistency across studies, and dose-response relationships also inform causal conclusions.

Clinical equipoise means there is genuine uncertainty in the expert community about which of the treatments being compared is better. It is an ethical prerequisite for randomizing participants: if one treatment is clearly superior, assigning participants to the inferior treatment would be unjustified.

Loss to follow-up can introduce bias if participants drop out for reasons related to their treatment or outcome. Intention-to-treat analysis keeps all participants in their assigned groups to preserve the benefits of randomization. Missing outcome data should be handled through pre-specified methods such as multiple imputation, and sensitivity analyses should test whether results change under different missing-data assumptions.

No. RCTs are well suited for evaluating the effect of interventions under conditions of uncertainty, but they are not appropriate for studying harmful exposures for ethical reasons, rare outcomes because of feasibility, questions about natural history or prognosis, or experiences and perspectives best studied qualitatively. The best design depends on the specific research question.

A superiority trial tests whether one treatment is better than another. A non-inferiority trial tests whether a new treatment is not worse than a comparator by more than a pre-specified margin. An equivalence trial tests whether two treatments differ by no more than a pre-specified margin in either direction. These are distinct hypotheses with different analysis approaches; a non-significant superiority result does not establish non-inferiority or equivalence.

CONSORT is a reporting guideline that specifies what information should appear in a published RCT report. It exists because incomplete reporting prevents readers from assessing whether a trial was conducted properly. A CONSORT flow diagram traces participants from eligibility assessment through to final analysis, making losses and exclusions transparent.

Consort Statement: consort-statement.org — current CONSORT checklist and flow diagram template.
Cochrane Handbook: training.cochrane.org/handbook — guidance on randomization, bias assessment, and evidence synthesis.
ClinicalTrials.gov: clinicaltrials.gov — publicly accessible trial registration database (NIH/National Library of Medicine).
WHO ICTRP: who.int/clinical-trials-registry-platform — international registry network.
For hypothesis testing foundations underpinning RCT analysis, see: Hypothesis Testing, Confidence Intervals, and Effect Size at Statistics Fundamentals.