100 High-Yield Cards for Paper Critique Mastery
Companion to FRCS Part 2 Paper Critique Handbook
Pro Tip: Cover the answer side with your hand or a piece of paper. Speak your answer out loud before checking!
The absolute difference in event rates between control and experimental groups.
Example: If control death rate = 20% and treatment death rate = 15%:
ARR = 20% - 15% = 5%
The proportional reduction in risk compared to the control group.
Example: ARR = 5%, CER = 20%:
RRR = 5/20 = 25%
How many patients need treatment to prevent one adverse event.
Example: ARR = 5% = 0.05:
NNT = 1/0.05 = 20
Meaning: Treat 20 patients to prevent 1 death
p-value = the probability of observing a result at least as extreme as this one, if the null hypothesis were true.
It is not "the probability the result is due to chance," nor the probability the null hypothesis is true — that is the classic viva trap.
CI = Range within which the true effect likely lies (95% confident)
For differences:
For ratios (OR, RR, HR):
Ability of a test to correctly identify people WHO HAVE the disease.
Think: SnNout - High Sensitivity rules Negative result out
Ability of a test to correctly identify people WHO DON'T HAVE the disease.
Think: SpPin - High Specificity rules positive result in
Type I Error (α): False positive
→ Saying there's an effect when there isn't (rejecting true null hypothesis)
→ Set by p-value threshold (usually 0.05 = 5% risk)
Type II Error (β): False negative
→ Missing a real effect (failing to reject false null hypothesis)
→ Related to statistical power (Power = 1 - β)
Power = Ability to detect a real effect if it exists
Power = 1 - β (Type II error)
Standard target: 80% power
Increased by:
Interpretation:
Example: RR = 0.75 means 25% reduction in risk
From 2×2 table: Odds of exposure in cases vs controls
Interpretation (same direction as RR):
HR = the ratio of the hazard (event rate at a given instant) in one group vs the other, across the whole follow-up.
Used in survival analysis (time-to-event data)
Interpretation:
Comparing 2 independent groups, continuous data:
Independent t-test:
Mann-Whitney U test:
Comparing 2 categorical variables:
Chi-square test:
Fisher's exact test:
Analysis of Variance (ANOVA)
Use: Comparing means of 3+ groups
Requirements:
If significant → Post-hoc tests to see which specific pairs differ
Linear Regression:
Logistic Regression:
Cox Proportional Hazards Regression
Use: Time-to-event data (survival analysis)
Outcome: Hazard Ratios (HR) for each predictor
Example: Identifying independent predictors of recurrence-free survival while adjusting for age, stage, grade
A result can have p<0.05 but be too small to matter in practice.
Example: New drug reduces operative time by 2 minutes
ITT: Analyze patients in the groups they were randomized to, regardless of what actually happened
Why?
Example: Patient randomized to surgery but refused → still counted in surgery group
I² = Measure of how different the studies are
Interpretation:
Prevalence: Proportion with disease AT ONE TIME POINT
Incidence: NEW cases over a TIME PERIOD
Correlation: Two variables are associated/related
Causation: One variable CAUSES the other
Classic example: Ice cream sales correlate with drowning deaths
The diamond at the bottom of a forest plot shows the pooled result from all studies combined.
Interpretation:
How many patients need to be exposed to cause one additional harm
Example: New drug increases bleeding by 3%
NNH = 1/0.03 = 33
Meaning: Treat 33 patients to cause 1 additional bleed
Standard Deviation (SD):
Standard Error (SE):
When: Sample doesn't represent target population
Example: Trial only enrolls fit, healthy patients but results applied to elderly with comorbidities
Prevention:
When: Groups treated differently beyond the intervention
Example: Intervention group gets more clinic visits, closer monitoring, extra support
Prevention:
When: Outcomes measured differently between groups
Example: Assessor knows treatment allocation and "looks harder" for complications in one group
Prevention:
When: Systematic differences in dropouts/withdrawals
Example: Sicker patients preferentially drop out from treatment group because of side effects
Prevention:
When: Participants remember past exposures differently
Example: Patients with cancer remember risk factors (smoking, diet) better than healthy controls
Most common in: Case-control studies
Prevention:
When: Positive studies more likely to be published than negative ones
Result: Meta-analyses overestimate treatment effects
Detection: Funnel plot asymmetry
Prevention:
Definition: Variable associated with BOTH exposure AND outcome that distorts the apparent relationship
Classic Example: Coffee → Lung cancer?
At Design Stage:
Example of restriction: Only include non-smokers to eliminate smoking as confounder
At Analysis Stage:
When: Early detection appears to improve survival without actually changing death
Example: Screening detects cancer 2 years earlier
Solution: Use mortality, not survival from diagnosis
When: Screening preferentially detects slow-growing cancers
Why: Aggressive cancers progress quickly and present between screens
Result: Screened cancers have better prognosis (but only because we caught the slow ones)
Solution: Randomized screening trials
When: Assessor's expectations influence measurements
Example: Surgeon who performed novel technique assesses outcomes
Prevention:
When: People change behavior because they know they're being observed
Example: Patients in trial comply better with medication than real-world patients
Result: Trial overestimates real-world effectiveness
Solution: Pragmatic trial design, minimal monitoring
When: Selective reporting of favorable outcomes
Example: Trial pre-specifies primary outcome A, but reports secondary outcome B because it was significant
Prevention:
Randomization:
Blinding:
Single-blind: Patient doesn't know (doctor does)
Double-blind: Patient AND doctor don't know
Triple-blind: Patient, doctor, AND outcome assessor don't know
Impossible to blind:
CAN still blind:
When: Treatment decision is based on prognosis
Example: Sicker patients get more aggressive treatment
Internal Validity: Are results true for THIS study population?
External Validity (Generalizability): Do results apply to OTHER populations?
When: Errors in measuring/classifying variables
Examples:
Types: Recall bias, observer bias, measurement error
Direction: Forward (Exposure → Outcome)
Setup:
Measure: Relative Risk
Example: Follow smokers vs non-smokers → lung cancer
Direction: Backward (Outcome → Exposure)
Setup:
Measure: Odds Ratio (NOT relative risk!)
Example: Lung cancer patients vs controls → smoking history
Best for RARE OUTCOMES
Why: You start with cases that already have the rare outcome
Other advantages:
Main problem: Recall bias
Other limitations:
Gold standard because:
Prospective Cohort:
Retrospective Cohort:
Design: Snapshot at one time point
Measures: Prevalence (not incidence)
Example: Survey of current smoking rates
Advantage: Quick, cheap
Limitation: Cannot establish temporal relationship or causation
Systematic Review:
Meta-Analysis:
Don't pool if:
P = Population/Patient
I = Intervention
C = Comparison/Control
O = Outcome
Example: In patients with gastric cancer (P), does perioperative chemotherapy (I) vs surgery alone (C) improve survival (O)?
Method: Randomize within subgroups to ensure balance of important prognostic factors
Example: Stratify by cancer stage
Result: Equal stage distribution in both groups
Method: Randomize in blocks to ensure equal group sizes
Example: Block of 4
Result: Balanced allocation throughout trial
Method: Randomize groups (hospitals, wards) not individuals
When used: Intervention is at group level
Example: Surgical safety checklist implementation
Analysis: Need to account for clustering
Definition: Keeping randomization sequence hidden until patient enrolled
Prevents: Selection bias
How:
Intention-to-Treat (ITT):
Per-Protocol:
Level 1: Systematic review of RCTs (1a) or a single high-quality RCT (1b)
Level 2: Cohort study, or a lower-quality / underpowered RCT (2a/2b)
Level 3: Case-control study
Level 4: Case series
Level 5: Expert opinion
Goal: Show new treatment is "not worse than" standard (within acceptable margin)
Example: New operation is less invasive - want to show outcomes not significantly worse
Pre-specify: Non-inferiority margin (acceptable difference)
Definition: Combining multiple outcomes into one endpoint
Example: MACE (Major Adverse Cardiac Events)
Advantage: Increases power, fewer patients needed
Problem: Mixing important (death) with less important outcomes
Design: Real-world conditions, broad inclusion, minimal intervention
vs Explanatory trial: Ideal conditions, strict criteria, close monitoring
Pragmatic advantages:
Definition: Planned analysis before trial completion
Reasons:
Important: Must be pre-planned, use adjusted significance levels
The 4-Sentence Opening:
❌ "Age causes complications"
✅ Say: "Age correlates with complications"
❌ "p=0.05 so it's clinically important"
✅ Say: "While statistically significant, we must consider clinical significance"
❌ "The paper says..."
✅ Say: "The authors found/reported..."
❌ "I don't know"
✅ Say: "Based on what I can see in the methods..."
S = Situation (what you see)
"This is a... The main finding is..."
B = Background (context)
"This is important because current practice is..."
A = Assessment (your critique)
"The strengths are... However, I'm concerned about..."
R = Recommendation (conclusion)
"This suggests we should/shouldn't..."
Good viva response:
"I would check if randomization was appropriate by examining Table 1 for baseline balance. The groups should be similar for age, sex, comorbidities, and disease severity. If there are significant differences, randomization may have failed or the sample was too small to achieve balance. I'd also check if allocation was concealed and if the randomization method is described."
Professional response:
"True double-blinding is challenging in surgical trials because the operating surgeon necessarily knows the procedure performed. However, we can still minimize bias by blinding the outcome assessors, data analysts, and potentially the patients if incisions are similar and they're under general anaesthetic. The key is documented, standardized outcome assessment by independent, blinded assessors."
Where to look:
Viva response structure:
"While the result is statistically significant (p=[value]), I need to consider clinical significance. The absolute difference is [ARR], meaning [interpretation]. The NNT is [value], which means treating [X] patients prevents one event. Given the [risks/costs/burden] of treatment, I would consider this [clinically meaningful / clinically trivial]."
Good answer:
"I would check if a sample size calculation was performed. The authors should report the assumed effect size, significance level (usually 0.05), power (usually 80-90%), and expected dropout rate. If the calculated sample wasn't achieved, the study may be underpowered to detect a real effect, increasing risk of type II error."
Template response:
"The generalizability depends on how representative this sample is. I note the patients were [describe population]. In my practice, patients tend to be [older/younger/sicker/healthier]. The [strict inclusion criteria / single center / specialized setting] may limit external validity. However, the internal validity appears sound, so results likely apply to similar populations."
Always prepared to give:
2 Strengths:
2 Weaknesses:
Professional way to mention:
"I note the study was funded by [pharmaceutical company / device manufacturer]. While this doesn't automatically invalidate results, it does warrant careful scrutiny of the methods and consideration of whether outcomes favor the sponsor. I would look for independent replication and check if protocol was pre-registered."
6-Step Description:
Good response:
"The wide confidence interval indicates imprecision in the estimate. This usually reflects a small sample size, high variability, or few events. While the point estimate suggests [effect], the CI ranges from [lower] to [upper], meaning the true effect could be anywhere in this range. I would interpret this result cautiously and consider whether additional data is needed."
Structured answer:
"The loss to follow-up was [X]%. Generally, <5% is acceptable, 5-20% introduces potential bias, and >20% threatens validity. I'm concerned because: (1) dropout was differential between groups [[%] vs [%]], and (2) reasons suggest sicker patients may have dropped out. This could introduce attrition bias. Ideally, I'd want sensitivity analysis assuming worst-case scenarios."
0-2 min: Quick scan - title, abstract, figures
2-7 min: Deep dive - methods (PICO, design, randomization), Table 1 (baseline), main results table
7-10 min: Prepare critique - identify 2 strengths, 2 weaknesses, calculate NNT if possible, note discussion points
Never say "I don't know" - instead:
Critical points to make:
"While the subgroup analysis suggests [finding], I'm cautious because: (1) it appears post-hoc rather than pre-specified, (2) multiple subgroups increase risk of false positives, and (3) each subgroup is underpowered. Subgroup findings should be hypothesis-generating only and require validation in dedicated trials."
The complete answer:
"The control event rate was [X]% and the experimental rate was [Y]%, giving an absolute risk reduction of [ARR]%. This represents a relative risk reduction of [RRR]%. The number needed to treat is [NNT], meaning we need to treat [N] patients to prevent one [outcome]. Given the [burden/cost/risks] of treatment, I consider this [clinically significant/marginal]."
The Perfect 60-Second Pitch:
"This [design] of [N] patients with [condition] compared [intervention] to [control] and found [primary outcome with numbers]. The [ARR/RRR/NNT] was [values]. Main strengths are [2 strengths]. However, I'm concerned about [2 limitations]. This provides Level [X] evidence that in [population], [intervention] [does/doesn't] improve [outcome], and I [would/wouldn't] change practice based on this."
Balanced response template:
"This is Level [X] evidence showing [finding]. Before changing practice I would consider: (1) Are these results internally valid? (2) Do they apply to my patients? (3) Are benefits clinically meaningful given risks/costs? (4) Has this been replicated? Based on this single study, I would [await further evidence / discuss with MDT / consider in selected patients]."
Trap 1: Confusing correlation with causation
Trap 2: Saying "relative risk" for case-control
Trap 3: Ignoring clinical vs statistical significance
Trap 4: Not checking baseline balance (Table 1)
Trap 5: Assuming subgroup analysis is valid
Professional stalling phrases:
Professional uncertainty phrases:
The 8-Point Critique:
The bridge phrase:
"So statistically, [summarize numbers]. Clinically, this means for my patients, [translate to practice]. I would [counseling approach / treatment modification / shared decision-making]. The key message for the patient is [simple language explanation]."
Complete forest plot answer:
"This forest plot meta-analysis of [N] studies shows [intervention] compared to [control] for [outcome]. Each square represents a study, with size reflecting weight. The diamond shows the pooled estimate of [value] with 95% CI [range]. This [does/doesn't] cross the line of no effect (1.0 for ratio measures, 0 for mean differences), indicating [significance]. The I² of [X]% suggests [low/moderate/high] heterogeneity."
Professional responses:
Before entering the room:
If you realize you're wrong:
"Actually, let me correct that. On reflection..." OR "I apologize, I misspoke. What I meant to say was..."
If examiner corrects you:
"Thank you for clarifying. So if [restate correctly]..."
In the last 90 seconds:
Final 30 seconds template:
"In summary, this [study design] of [N] patients shows [main finding]. The [ARR/NNT] is [value], which I consider [clinically significant/not significant] given [rationale]. The main limitation is [key weakness]. This provides Level [X] evidence, and based on this, I would [clinical action]."
Honest fallback:
"I'm not familiar with this specific statistical method. However, I can see the p-value is [X] and confidence interval is [Y], suggesting [interpretation]. The clinical finding appears to be [describe]. I would discuss this with a statistician to fully understand the analysis."
Comparison response:
"Multi-center trials have better external validity because they include diverse populations, surgical techniques, and settings. However, single-center studies often have more consistent protocols and data quality. This [multi/single]-center design means results are [more/less] generalizable but [more/less] internally consistent."
Professional transitions:
"I'm PREPARED. I'm CALM. I'm a SURGEON."
You know PICO.
You know 2+2 (strengths + weaknesses).
You know ARR, RRR, NNT.
You've practiced.
You've got this. 💪
These flashcards distill the essential knowledge from the FRCS Part 2 Paper Critique Handbook. Review them daily, focus on your weak areas, and use them for last-minute revision. Remember: you don't need to know everything perfectly - you need to demonstrate structured thinking and professional communication.
Good luck in your viva! 🎓