Causal inference is the scientific and statistical discipline concerned with drawing conclusions about cause-and-effect relationships from data, distinguishing genuine causal mechanisms from mere statistical association. It employs frameworks such as potential outcomes (Rubin causal model), structural causal models (Pearl’s do-calculus), and graphical models (directed acyclic graphs) to formalise interventions and reason about counterfactuals. Applications span medicine, economics, social science, and AI alignment, wherever understanding the effect of an action — not merely its correlation with outcomes — is required.
Semantic Classification
Content
Compositional Relationships (Components)
SubClassOf(ai:CausalInference
ObjectSomeValuesFrom(ai:hasPart ai:DirectedAcyclicGraph))
SubClassOf(ai:CausalInference
ObjectSomeValuesFrom(ai:hasPart ai:InstrumentalVariables))
SubClassOf(ai:CausalInference
ObjectSomeValuesFrom(ai:hasPart ai:PotentialOutcomesFramework))
SubClassOf(ai:CausalInference
ObjectSomeValuesFrom(ai:hasPart ai:StructuralCausalModel))
SubClassOf(ai:CausalInference
ObjectSomeValuesFrom(ai:hasPart ai:DoCalculus))
SubClassOf(ai:CausalInference
ObjectSomeValuesFrom(ai:hasPart ai:PropensityScoreMatching))
SubClassOf(ai:CausalInference
ObjectSomeValuesFrom(ai:hasPart ai:DifferenceInDifferences))
SubClassOf(ai:CausalInference
ObjectSomeValuesFrom(ai:hasPart ai:CausalForest))
SubClassOf(ai:CausalInference
ObjectSomeValuesFrom(ai:hasPart ai:CounterfactualReasoning))
SubClassOf(ai:CausalInference
ObjectSomeValuesFrom(ai:hasPart ai:RegressionDiscontinuity))
Dependency Relationships
SubClassOf(ai:CausalInference
ObjectSomeValuesFrom(ai:requires ai:ObservationalData))
SubClassOf(ai:CausalInference
ObjectSomeValuesFrom(ai:requires ai:ConfoundingVariable))
SubClassOf(ai:CausalInference
ObjectSomeValuesFrom(ai:requires ai:ProbabilityTheory))
SubClassOf(ai:CausalInference
ObjectSomeValuesFrom(ai:requires ai:GraphTheory))
SubClassOf(ai:CausalInference
ObjectSomeValuesFrom(ai:requires ai:StatisticalModel))
SubClassOf(ai:CausalInference
ObjectSomeValuesFrom(ai:requires ai:IdentificationAssumption))
SubClassOf(ai:CausalInference
ObjectSomeValuesFrom(ai:requires ai:IgnorabilityCondition))
Capability Relationships
SubClassOf(ai:CausalInference
ObjectSomeValuesFrom(ai:enables ai:CausalLanguageModelling))
SubClassOf(ai:CausalInference
ObjectSomeValuesFrom(ai:enables ai:CounterfactualReasoning))
SubClassOf(ai:CausalInference
ObjectSomeValuesFrom(ai:enables ai:AlgorithmicFairness))
SubClassOf(ai:CausalInference
ObjectSomeValuesFrom(ai:enables ai:PolicyEvaluation))
SubClassOf(ai:CausalInference
ObjectSomeValuesFrom(ai:enables ai:HeterogeneousTreatmentEffect))
SubClassOf(ai:CausalInference
ObjectSomeValuesFrom(ai:enables ai:ChainOfThoughtReasoning))
SubClassOf(ai:CausalInference
ObjectSomeValuesFrom(ai:enables ai:CausalReinforcementLearning))
SubClassOf(ai:CausalInference
ObjectSomeValuesFrom(ai:enables ai:CausalRepresentationLearning))
Implementation Relationships
SubClassOf(ai:CausalInference
ObjectSomeValuesFrom(ai:implements ai:BayesianInference))
SubClassOf(ai:CausalInference
ObjectSomeValuesFrom(ai:implements ai:AutomatedReasoning))
SubClassOf(ai:CausalInference
ObjectSomeValuesFrom(ai:implements ai:MathematicalReasoning))
SubClassOf(ai:CausalInference
ObjectSomeValuesFrom(ai:implements ai:InferenceAlgorithm))
SubClassOf(ai:CausalInference
ObjectSomeValuesFrom(ai:implements ai:StatisticalIdentification))
SubClassOf(ai:CausalInference
ObjectSomeValuesFrom(ai:implements ai:EvidenceSynthesis))
Reduction Relationships
SubClassOf(ai:CausalInference
ObjectSomeValuesFrom(ai:reducesTo ai:AssociationAnalysis))
SubClassOf(ai:CausalInference
ObjectSomeValuesFrom(ai:reducesTo ai:ProbabilisticInference))
SubClassOf(ai:CausalInference
ObjectSomeValuesFrom(ai:reducesTo ai:StatisticalEstimation))
SubClassOf(ai:CausalInference
ObjectSomeValuesFrom(ai:reducesTo ai:CounterfactualQuery))
About
Causal inference addresses one of the most fundamental challenges in empirical science: how to move from observing patterns in data to understanding the mechanisms that produced them. The field rests on a simple but profound distinction — the difference between asking “what is associated with what?” (a statistical question answerable by correlation) and “what would happen if we intervened?” (a causal question that requires additional assumptions about mechanism). This distinction, formalised mathematically in the late twentieth century, revealed that conventional statistical methods were systematically ill-equipped to answer causal questions from observational data without additional structure.
The intellectual lineage of causal inference spans over a century. Sewall Wright’s path analysis (1918–1920) introduced graphical representations of causal structure in genetics, allowing researchers to distinguish direct and indirect effects in systems of interrelated variables. In economics, structural equation modelling (Haavelmo, 1944; Cowles Commission) represented causal mechanisms as algebraic equations, though the causal interpretation of these equations remained philosophically contested for decades. The field’s modern foundations were laid in two partially independent traditions: Jerzy Neyman’s (1923) and Donald Rubin’s (1974, 1978) potential outcomes framework, which defines causal effects as contrasts between counterfactual outcomes under different treatment assignments, and Judea Pearl’s graphical causal modelling programme (1988–2000), which provided a formal language — directed acyclic graphs, the do-operator, and do-calculus — for representing and reasoning about causal structure.
Pearl’s Causality (2000, second edition 2009) unified the graphical and potential-outcomes traditions and established the do-calculus as a complete algorithm for determining whether a causal quantity is identifiable from observational data given a causal graph. Pearl’s hierarchy of causal reasoning — association (seeing), intervention (doing), and counterfactual (imagining) — provides a principled taxonomy of the questions that data analysis can and cannot answer without additional causal assumptions. This hierarchy has proved influential in AI alignment and evaluation research, where distinguishing associative pattern matching from genuine causal understanding is a central concern in assessing the reliability of large language models.
The methodological history of causal inference in the second half of the twentieth century was characterised by disciplinary divergence: epidemiologists developed confounding adjustment methods for observational cohort studies; economists developed instrumental variable and natural experiment approaches; social scientists applied path analysis and structural equation modelling; biostatisticians developed survival analysis with time-varying confounders; and machine learning researchers largely ignored causation in favour of prediction. Pearl’s structural causal model framework (1988–2009) provided for the first time a unified mathematical language that subsumed all these approaches, proved their theoretical equivalence in key cases, and identified precisely when and why each is applicable. The subsequent convergence of the field around DAGs, do-calculus, potential outcomes, and identification theory is the most significant methodological development in applied statistics since the development of maximum likelihood estimation.
The 2021 Nobel Prize in Economic Sciences crystallised causal inference’s status as a mainstream scientific paradigm. The Nobel Committee’s citation for Joshua Angrist, David Card, and Guido Imbens described the “natural experiments” methodology as a genuinely new tool for answering causal questions in settings where randomised experiments are infeasible — applicable equally to labour economics, health economics, development economics, and public policy evaluation. The prize attracted attention from researchers across the quantitative sciences to causal identification as a distinct methodological challenge, accelerating adoption of causal methods in fields ranging from medicine to computer science.
Components and Architecture
Causal inference consists of several interacting methodological components:
Frameworks:
-
Rubin Causal Model (Potential Outcomes): Defines the individual treatment effect as ITE = Y(1) - Y(0) for each unit, where Y(1) and Y(0) are potential outcomes under treatment and control. The fundamental problem of causal inference — only one potential outcome is observable — means that individual effects are never directly observed. Population-level estimands (ATE, ATT, ATU) are defined as expectations over these contrasts and are identifiable under assumptions of ignorability (unconfoundedness) and positivity (overlap).
-
Pearl’s Structural Causal Model: Represents causal structure as a DAG with nodes (variables) and directed edges (direct causal effects), plus structural equations specifying how each variable is determined by its parents and exogenous noise. The do-operator formalises intervention: do(X=x) removes all incoming edges to X and sets X to x, modelling a surgical external intervention. Three rules of do-calculus suffice to derive all identifiable interventional and counterfactual distributions from observational data under the graph.
Identification Strategies:
-
Randomised Controlled Trials (RCTs): Random assignment ensures treatment is independent of potential outcomes, making naive difference in means a valid estimator of ATE. The gold standard but often ethically or practically infeasible.
-
Instrumental Variables (IV): Uses variables that affect treatment but affect outcomes only through treatment (exclusion restriction) to identify causal effects in the presence of unmeasured confounding. Local Average Treatment Effect (LATE) theorem (Imbens & Angrist, 1994) provides the interpretable estimand for heterogeneous-treatment-effect settings.
-
Regression Discontinuity Design: Exploits discontinuities in treatment assignment rules (e.g. policy thresholds) to identify causal effects for units near the threshold.
-
Difference-in-Differences (DiD): Compares pre-post changes in treated and control groups, removing time-invariant confounders under the parallel trends assumption.
-
Back-door and Front-door Adjustment: Graphical criteria specifying which sets of observed variables suffice to adjust for confounding, enabling observational identification without instruments.
Estimation Methods:
-
Propensity Score Methods: Model the probability of treatment given observed covariates; matching, inverse probability weighting (IPW), and doubly robust estimators use this to adjust for observed confounding.
-
Meta-learners: T-learner, S-learner, X-learner, DR-learner — modular approaches that combine off-the-shelf machine learning predictions to estimate heterogeneous treatment effects (HTEs).
-
Causal Forests: Non-parametric, tree-based estimators (Wager & Athey, 2018) that estimate HTEs adaptively, building on generalised random forests (grf) with honest sample splitting, asymptotic normality guarantees, and double machine learning nuisance model removal.
-
Double Machine Learning (DML): Uses cross-fitting and Neyman orthogonality to debias causal estimates when high-dimensional nuisance models (outcome prediction, propensity score) are estimated with machine learning methods (Chernozhukov et al., 2018).
Sensitivity Analysis:
-
E-values (VanderWeele & Ding, 2017) quantify the minimum confounding strength required to explain away an observed causal estimate, enabling calibrated robustness claims without measuring unmeasured confounders directly.
-
Rosenbaum bounds for matched observational studies.
-
Simulation-based sensitivity analysis using partial identification bounds.
Use Cases and Major Application Families
Clinical Medicine and Pharmacoepidemiology: Randomised trials establish efficacy; causal inference extends this to effectiveness in real-world populations using electronic health record data. Target trial emulation (Hernán & Robins, 2016) provides a principled framework for designing observational studies that mimic the structure of hypothetical RCTs. UK Biobank studies (involving hundreds of thousands of participants) routinely apply Mendelian randomisation — a form of IV analysis using genetic variants as instruments — to estimate causal effects of exposures such as BMI and smoking on disease incidence across the landscape of conditions. Applications include estimating drug efficacy in subgroups, rare adverse event detection, and personalised medicine through HTE estimation.
Economics and Labour Market Policy: The Nobel Prize-winning work of Angrist, Card, and Imbens (2021) established natural experiments as the dominant design for empirical labour economics. Card’s analysis of the 1980 Mariel boatlift (using Miami as treated, comparison cities as control) estimated the labour market impact of immigration; Angrist and Krueger’s (1991) quarter-of-birth instrument identified the returns to education; Imbens and Angrist’s (1994) LATE theorem provided the formal framework for interpreting IV estimates in heterogeneous populations. These methods now underpin policy evaluation across minimum wage, welfare reform, education interventions, and active labour market programmes worldwide.
Technology Industry and Online Experimentation: Large technology platforms (Alphabet, Meta, Amazon, Netflix) run millions of randomised A/B experiments annually, but increasingly complement these with observational causal methods for settings where experimentation is infeasible or expensive. Causal forests and meta-learners are used for HTE estimation to personalise product interventions. Interrupted time-series and synthetic control methods (Abadie et al.) estimate platform-wide policy impacts. Causal recommender systems model the causal effect of recommendations on long-term engagement rather than naively optimising predicted click rates.
Algorithmic Fairness: Causal inference provides the theoretical foundations for distinguishing permissible from impermissible algorithmic discrimination. Counterfactual fairness (Kusner et al., 2017) defines a decision as fair if it would be the same in a counterfactual world where the protected attribute were different but all causally downstream variables were unchanged. Mediation analysis decomposes observed disparities into direct and indirect causal pathways, enabling targeted remediation. Causal RL fairness research (2025) extends these ideas to sequential decision-making, where fairness constraints must account for dynamic feedback between algorithmic decisions and population outcomes.
AI Safety and Alignment: Causal methods are being applied to evaluate whether large language model outputs reflect genuine causal understanding or spurious statistical regularities. Causal language model evaluation benchmarks (EconCausal, CausalVLBench, 2025) test models on counterfactual and interventional reasoning tasks. Structural causal models are used to audit whether RLHF-trained models have learned causally appropriate reward responses or are susceptible to reward hacking via spurious correlates. Causal representation learning aims to produce model internals that respect causal structure, improving robustness to distributional shift — directly relevant to AI Safety and Catastrophic Risk Reduction.
Causal Reinforcement Learning: Causal RL embeds causal knowledge into Reinforcement Learning to address four challenges: spurious correlations in reward attribution, sample inefficiency, poor generalisation across environments, and fairness. By modelling the causal structure of the environment, causal RL agents learn policies that are robust to distributional shifts and can transfer knowledge across tasks more efficiently than purely associative agents. The key insight is that the reward function in standard RL is an associative model: it predicts reward given observed state-action pairs without representing the causal mechanism by which actions produce outcomes. A causal RL agent instead models the interventional distribution P(reward | do(action), state), enabling it to correctly attribute credit for outcomes even in environments with confounded state-action distributions (e.g. when exploration policy induces selection bias in the collected experience). A 2025 IEEE TNNLS survey on causal RL documents rapid growth in this application area. Practical implementations combine structural causal world models (learned from interaction data using causal discovery algorithms) with model-based RL policy optimisation, achieving substantially better sample efficiency and out-of-distribution generalisation than model-free or model-based RL without causal structure.
Personalised Medicine and Precision Oncology: Causal inference for personalised medicine estimates individual-level treatment effects from heterogeneous patient populations, moving beyond population-average effects to identify which subgroups benefit from which treatments. This application is particularly important in oncology, where treatment response varies enormously across molecular subtypes of cancer and the same tumour type may respond to targeted therapy or immunotherapy depending on specific genomic features. Meta-learner approaches (T-learner, X-learner, DR-learner) combined with genomic feature selection enable heterogeneous treatment effect estimation across large cohorts. The fundamental challenge is that no patient receives both treatments (the fundamental problem of causal inference) — valid HTE estimation requires either large randomised trials with adequate subgroup power or causal machine learning methods that can estimate HTEs from observational data under unconfoundedness. UK NHS data infrastructure (NHS Digital, CPRD, OpenSAFELY) provides large-scale observational data for UK-specific causal pharmacoepidemiology research.
Education Policy and Programme Evaluation: Natural experiment designs have transformed education research by enabling causal evaluation of teaching interventions, curriculum changes, and school accountability policies without requiring randomised assignment. Regression discontinuity designs exploit school admission cutoffs, test score thresholds, and age-at-entry cutoffs to estimate the causal effects of education interventions on outcomes including academic achievement, earnings, and health. DiD designs estimate the effects of policy changes (teacher pay, school funding formulae, free school meal eligibility) using variation across local authorities or across time. The Education Endowment Foundation (EEF) in the UK has funded hundreds of randomised trials of educational interventions, and causal meta-analysis methods are used to synthesise effect sizes across diverse contexts and populations.
Academic Context
The modern field was shaped by the convergence of several intellectual traditions:
- Foundational Theoretical Contributions:
- Jerzy Neyman (1923): potential outcomes notation; counterfactual comparison as formal language for treatment effects.
- Ronald Fisher (1935): randomisation theory; hypothesis testing in designed experiments; exact tests for treatment effects.
- Sewall Wright (1918–1921): path coefficients and path analysis; graphical representation of causal structure in genetics.
- Trygve Haavelmo (1944): structural econometrics; simultaneous equation models with probabilistic causal interpretation; Cowles Commission methodology.
- Donald Rubin (1974, 1978): Rubin Causal Model; formal potential outcomes framework for observational studies; propensity score introduction.
- Paul Holland (1986): “Statistics and Causal Inference” — classic exposition of potential outcomes framework in statistics literature.
- Judea Pearl (1988–2009): graphical causal models, do-calculus, structural causal models; unification of all major causal frameworks; Causality (2000, 2009).
- Applied Methodology Milestones:
- Paul Rosenbaum & Donald Rubin (1983): propensity score theorem; balancing score for observational study covariate adjustment.
- James Heckman (1979): selection model; two-stage least squares; Heckman correction for sample selection bias. Nobel Prize 2000.
- Joshua Angrist & Alan Krueger (1991): quarter-of-birth instrumental variable; returns to education; popularised IV in labour economics.
- Guido Imbens & Joshua Angrist (1994): LATE theorem; interpreted IV estimand as average treatment effect for compliers.
- Miguel Hernán & James Robins (2000–2020): marginal structural models; g-computation; target trial emulation; Causal Inference: What If (2020).
- Susan Athey & Stefan Wager (2018): causal forests; honest splitting; asymptotic theory for non-parametric HTEs.
- Victor Chernozhukov et al. (2018): double machine learning; Neyman orthogonality; debiased ML-based causal estimates.
- Institutional and Research Group Landscape:
- Judea Pearl’s group (UCLA): do-calculus, counterfactual theory, causal hierarchy, causal representation learning.
- Guido Imbens and Susan Athey (Stanford): causal machine learning, causal forests, grf package, policy tree.
- James Robins (Harvard): semiparametric efficiency theory, marginal structural models, targeted learning.
- Miguel Hernán (Harvard): target trial emulation, pharmaco-epidemiology, causal inference in biomedical research.
- Victor Chernozhukov (MIT): double machine learning, high-dimensional econometrics, quantile treatment effects.
- London School of Hygiene and Tropical Medicine (LSHTM): Mendelian randomisation methodology, mediation analysis, time-varying confounders.
- MRC Integrative Epidemiology Unit (Bristol): world leader in Mendelian randomisation applied to biobank data; George Davey Smith, Kate Tilling.
- Edinburgh Causal AI Lab: causal discovery algorithms, causal representation learning, LLM causal reasoning evaluation.
- Nobel Prize Recognition (2021):
- The 2021 Nobel Memorial Prize in Economic Sciences awarded to David Card (UC Berkeley), Joshua Angrist (MIT), and Guido Imbens (Stanford).
- Card: “for his empirical contributions to labour economics” using natural experiments (Mariel boatlift, minimum wage studies).
- Angrist and Imbens: “for their methodological contributions to the analysis of causal relationships” including the LATE theorem and IV interpretation.
- Landmark recognition of causal inference as a central scientific paradigm in quantitative social science.
- Prize attracted researchers across many quantitative disciplines to causal identification as a distinct methodological challenge.
- Key Academic Venues and Resources:
- NeurIPS, ICML, ICLR workshops on causal machine learning and causal representation learning.
- American Causal Inference Conference (ACIC): primary US meeting for methodology researchers.
- Journal of Causal Inference (De Gruyter): dedicated journal for the field since 2013.
- Causal Inference in Statistics — A Primer (Pearl, Glymour & Jewell, 2016): accessible introduction to SCM-based causal analysis.
- Evidence Synthesis International and Cochrane Collaboration: causal synthesis methods in systematic reviews.
Current Landscape (2026)
By mid-2026, causal inference has achieved mainstream status across empirical science and is undergoing rapid integration with large-scale machine learning:
- LLM Causal Evaluation:
- New benchmarks systematically evaluate whether frontier language models exhibit genuine causal reasoning or associative mimicry.
- EconCausal (2025): tests LLMs on economics-style causal questions requiring confounder identification and natural experiment reasoning.
- CausalVLBench (2025): visual causal reasoning benchmark for large vision-language models.
- Causal Methods for LLM Development framework (2025, arXiv:2605.25998): debiased evaluation of LLM design decisions using causal inference on annotation processes.
- Results mixed: strong on verbal tasks addressable via pattern completion; poor on genuine do-calculus manipulation or counterfactual consistency.
- Finding (arXiv:2506.00844, 2025): “LLMs cannot discover causality” — reflects fundamental associative architecture limitation.
- Chain-of-thought reasoning often fails to be causally responsible for LLM answers: the visible reasoning process may not drive the actual computation.
- Causal Representation Learning:
- Project of learning representations that respect causal structure — disentangling latent causal factors, identifying causal variables from raw sensory data.
- Linear causal representation learning by topological ordering, pruning, and disentanglement (arXiv:2509.22553, 2025): progress toward identifiable causal latent structure.
- Integration with foundation models aims to produce more robust and interpretable learned representations.
- Key theoretical result: causal representations are identifiable under weaker assumptions than ICA-based independent component analysis.
- Applications to domain generalisation and out-of-distribution robustness: causally-structured representations should generalise better across environment shifts.
- Causal Methods for LLM Development and Evaluation:
- Causal inference applied to engineering decisions during LLM development: training data composition, architecture choices, RLHF implementation.
- Counterfactual questions: what would this model’s capabilities be had we used different training data? What caused the observed capability difference between model versions?
- Methods combining large-scale LLM annotations with gold-standard human labels (Imai & Li, 2025) produce debiased estimates with formal statistical guarantees.
- Causal mediation analysis of model internals: which intermediate representations causally mediate safety-relevant input-output relationships?
- Heterogeneous Treatment Effects in Medicine (2025):
- Causal forests and double machine learning now standard in clinical research for HTE estimation.
- 2025 systematic review (International Statistical Review, Rehill et al.): causal forests among most widely adopted HTE methods; grf R package the dominant implementation.
- Psychiatry: causal forests for personalised treatment effect estimation in antidepressant trials (PMC, 2025).
- HIV care: causal forest DML for TB preventive therapy impact on ART adherence (Scientific Reports, 2025).
- Oncology: HTE estimation for targeted cancer treatment selection using electronic health record data.
- Public health: synthetic control and DiD methods for UK-wide COVID-19 intervention evaluation using NHS administrative data.
- Causal Reinforcement Learning Integration (2025):
- IEEE TNNLS 2025 survey: causal RL integrating structural causal models with policy gradient, model-based RL, and offline RL algorithms.
- Causally-Enhanced Reinforcement Policy Optimisation (arXiv:2509.23095, 2025): substantial sample efficiency gains over standard RL by incorporating causal world model structure.
- Causal RL addresses: spurious correlations in reward attribution; poor generalisation across environments; sample inefficiency; fairness constraints.
- Online platforms (Netflix, Spotify, LinkedIn) deploying causal recommendation systems that model long-term engagement effects rather than naive click prediction.
- Causal off-policy evaluation for healthcare decision support: estimating outcomes of counterfactual treatment policies using observational hospital data.
- Econometric Policy Evaluation:
- Staggered DiD designs with heterogeneous timing (Callaway & Sant’Anna, 2021; Sun & Abraham, 2021) address problems with conventional TWFE estimators.
- Synthetic difference-in-differences (Arkhangelsky et al., 2021): matrix completion for improved synthetic control estimation.
- High-dimensional IV estimation using LASSO-based first stages (Chernozhukov et al.) handles settings with many potential instruments.
- Real-time policy evaluation using online causal methods: estimating effects of platform algorithmic changes within days rather than months.
UK Context
The United Kingdom has strong academic traditions in causal inference across multiple disciplines:
- Medical Research Council and Biobank Infrastructure:
- MRC has been a leading funder of causal epidemiology research for over four decades.
- MRC Integrative Epidemiology Unit (IEU), University of Bristol: world-leading centre for Mendelian randomisation.
- George Davey Smith and colleagues developed Mendelian randomisation as a formal causal inference method in epidemiology.
- Kate Tilling contributes to causal longitudinal methods and life-course epidemiology.
- UK Biobank: 500,000+ participants with linked health records; foundational infrastructure for large-scale causal epidemiology.
- Landmark causal attribution analyses of smoking and BMI effects across disease landscape (PMC, 2022) exemplify UK biobank causal inference capacity.
- London School of Hygiene and Tropical Medicine (LSHTM):
- Contributed extensively to time-varying treatment causal methodology, marginal structural models, and g-estimation.
- Collaborative work with Harvard’s Hernán group on target trial emulation frameworks.
- Target trial emulation now shapes randomised trial emulation across UK health data infrastructure.
- Applied to CPRD (Clinical Practice Research Datalink), SAIL Databank, NHS Digital, and OpenSAFELY.
- Economic Policy Research:
- Centre for Economic Performance (CEP, LSE): applies natural experiment designs to UK labour market, trade, and education policy questions.
- Institute for Fiscal Studies (IFS): universal credit evaluation, early years education interventions, NHS resource allocation — all using causal identification strategies.
- NBER-style natural experiment tradition in UK economics research: UK as a natural laboratory for policy variation (devolution, Brexit, differential minimum wage implementation).
- AI and Causal Inference Intersection:
- Edinburgh Causal AI Lab: causal discovery algorithms, LLM causal reasoning evaluation, causal representation learning.
- Alan Turing Institute: causal inference designated as strategic research priority; Data-Centric Engineering programme.
- UKRI AI programme: explicitly includes causal reasoning and out-of-distribution robustness as funded themes.
- Recognition that associative machine learning alone is insufficient for high-stakes AI applications driving policy-level investment.
- Northern England Contributions:
- University of Leeds: computational epidemiology applying DiD and synthetic control methods to public health interventions; regional health inequality causal analysis.
- University of Sheffield (ScHARR — School of Health and Related Research): causal methods for health technology assessment; NICE guidance methodology.
- University of Manchester (Alliance Manchester Business School): causal econometrics applied to Northern Powerhouse labour market questions; regional economic impact evaluation.
- Newcastle University: causal analysis of health determinants in deprived populations; Fuse Centre for Translational Research in Public Health.
- UKRI and National Strategy:
- UKRI Strategic Priorities Fund includes causal inference methods for large-scale administrative data.
- Connected Health Cities and Trusted Research Environment programmes provide causal analysis infrastructure.
- Health Data Research UK (HDRUK): federated causal inference across NHS data systems; privacy-preserving causal estimation.
Causal Discovery Methods
Causal discovery — learning causal graph structure from data, rather than estimating effects given a known graph — is a distinct methodological tradition within causal inference:
- Constraint-Based Methods:
- PC Algorithm (Spirtes, Glymour & Scheines, 1991): uses conditional independence tests to eliminate edges and orient directed edges in a DAG. Complexity: O(p^d) independence tests for p variables with maximum degree d.
- FCI Algorithm (Fast Causal Inference): extends PC to settings with latent variables and selection bias; produces Partial Ancestral Graphs (PAGs) encoding uncertainty about causal directions.
- RFCI (Really Fast Causal Inference): computationally efficient approximation for large variable sets.
- Key assumption: Faithfulness — observed conditional independencies reflect graph structure, with no “accidental” cancellations.
- Score-Based Methods:
- GES (Greedy Equivalence Search, Chickering 2002): searches over equivalence classes of DAGs using a penalised likelihood score (BIC, AIC, MDL); polynomial time under faithfulness.
- NOTEARS (Zheng et al., 2018): reformulates DAG structure learning as a continuous optimisation problem with an algebraic acyclicity constraint; gradient descent over the space of weighted adjacency matrices.
- DAG-GNN, GraN-DAG, GOLEM: neural network-based score-based methods enabling non-linear causal structure learning.
- Functional Causal Model Methods:
- LiNGAM (Linear Non-Gaussian Acyclic Model, Shimizu et al. 2006): exploits non-Gaussianity of noise to identify full causal ordering (not just Markov equivalence class) from observational data; unique identifiability under non-Gaussian noise.
- ANM (Additive Noise Models): non-linear generalisations of LiNGAM where each variable is a non-linear function of its parents plus independent noise.
- Bivariate causal direction testing (Mooij et al.): tests which of X → Y or Y → X is more plausible using asymmetries in residual distributions.
- LLM-Assisted Causal Discovery (2025):
- Evidence Triangulator framework: uses LLMs to extract and synthesise causal evidence across study designs (RCTs, observational studies, mechanism descriptions).
- Causal-LLM: unified one-shot framework combining prompt-driven and data-driven causal graph discovery.
- Key limitation: LLMs cannot discover novel causality from data (arXiv:2506.00844, 2025); they retrieve and synthesise prior knowledge, not causal patterns.
- Hybrid approaches combining LLM knowledge elicitation (graph initialisation) with data-driven refinement show promise.
- LLM-elicited causal graphs as priors for Bayesian causal discovery is an emerging research direction.
- Time-Series Causal Discovery:
- Granger Causality: X Granger-causes Y if past values of X improve prediction of Y beyond what past Y alone provides. A statistical (not causal) criterion without additional assumptions.
- PCMCI (Runge et al., 2019): PC algorithm adapted for time series; discovers lagged and contemporaneous causal links using conditional independence on lagged values.
- NeurIPS 2025 benchmark on time-series causal discovery highlights progress and remaining challenges in financial and climate data.
Pearl’s Causal Hierarchy and AI Implications
Judea Pearl’s ladder of causation provides a principled taxonomy for distinguishing the types of reasoning that different AI architectures can and cannot perform:
Rung 1 — Association (Seeing): Queries of the form P(Y | X=x) — what is the probability of Y given that we observe X=x? Standard statistical learning, including all current deep learning, operates at this level. Large language models, trained on observational data, are intrinsically Rung 1 systems: they learn associations between tokens and concepts as they appear in training corpora. However rich and nuanced these associations, they do not without further structure support valid answers to causal or counterfactual questions.
Rung 2 — Intervention (Doing): Queries of the form P(Y | do(X=x)) — what is the probability of Y if I intervene and set X to x? Answering such questions requires identifying the causal mechanism relating X to Y, not merely their co-occurrence. This level is achievable by combining Rung 1 data with a causal graph (Pearl’s approach) or by running a randomised experiment. Reinforcement learning with environment interaction operates at this level, since the agent intervenes in the world and observes outcomes.
Rung 3 — Counterfactual (Imagining): Queries of the form P(Y(x’) | X=x, Y=y) — given that we observed X=x and Y=y, what would Y have been had X been x’ instead? Answering such questions requires a full structural causal model, not just observational data plus a causal graph. Counterfactual reasoning is required for explanations (“why did Y happen?”), blame attribution (“would the harm have occurred had the action been different?”), and certain fairness analyses.
Pearl (2025 reprint) argues that all human cognitive tasks relevant to intelligence require at minimum Rung 2 and often Rung 3 reasoning, and that current AI systems operating at Rung 1 are therefore fundamentally limited in their ability to generalise beyond training distribution, understand causation, or reason about hypothetical scenarios. This theoretical limitation motivates causal AI research: the integration of causal structure into neural architectures to enable genuine Rung 2 and 3 capabilities.
The implication for AI safety is direct: if LLMs operate only at Rung 1, they cannot reliably reason about the consequences of interventions (Rung 2) or counterfactuals (Rung 3). This limits their ability to engage in robust safety-relevant reasoning — e.g. correctly predicting the consequences of a novel action in a high-stakes setting. Causal AI safety research aims to address this limitation.
Future Directions (2026–2030)
Several developments will shape causal inference over the next four years:
Causal Foundation Models: Integrating causal inductive biases into large pre-trained models — through architectures that explicitly represent interventional vs. observational distributions, or through causal regularisation objectives during pre-training — is an active research direction. Approaches include: incorporating causal graph structure as a relational bias in transformer attention; pre-training on synthetic datasets with known causal structure; and fine-tuning on tasks requiring do-calculus manipulation. Success would produce models capable of reliable Rung 2 and Rung 3 reasoning on novel scenarios without task-specific fine-tuning, addressing a fundamental limitation of current associative architectures.
Automated Causal Discovery at Scale: Learning causal graph structure (not just causal effects given a graph) from data remains computationally hard in the general case (NP-hard for exact structure learning). Recent progress on constraint-based (PC algorithm, FCI algorithm for latent variables), score-based (GES, NOTEARS differentiable DAG learning), and functional (LiNGAM, ANM) approaches, combined with LLM-based causal graph elicitation and structure initialisation (Evidence Triangulator, 2025), may produce practical automated causal discovery for moderate-dimensional systems. Extensions to time series (Granger causality, PCMCI) and high-dimensional imaging data are active frontiers.
Causal AI Safety: The intersection of causal inference and AI Safety will deepen, with causal frameworks being used to define and audit alignment properties — whether a model’s reward function corresponds to the intended causal objective, whether safety-relevant behaviours have the right causal structure, and how causal reasoning about model internals can support Mechanistic Interpretability. Causal mediation analysis applied to neural network internals promises to identify which intermediate representations causally mediate safety-relevant input-output relationships, supporting more targeted interpretability and intervention.
Policy-Relevant Heterogeneity at National Scale: Governments and international organisations are increasingly demanding causal evaluations of large-scale policies — climate adaptation investments, education reforms, NHS treatment protocols, labour market interventions — with subgroup-specific estimates. Methods for causal inference from administrative data at national scale (UK-APB, UKRI Data Infrastructure, NHS federated analytics) will be critical enablers. The integration of privacy-preserving methods (differentially private causal estimation, federated causal inference) with HTE estimation is an important methodological challenge.
Causal RL at Scale: As reinforcement learning is applied to progressively longer-horizon, higher-stakes problems (drug discovery, autonomous systems, scientific research automation), causal world models will become essential for robust policy learning. Theoretical foundations for causal RL are being actively developed, and 2026–2030 will likely see the first large-scale deployments of causal RL in pharmaceutical discovery, precision agriculture, and energy grid management. The integration of offline causal RL — learning causal policies from observational data without environment interaction — is particularly important for safety-critical domains where online exploration is dangerous.
Instrumental Causal AI: The “instrumental variables” of language — finding parts of the training data or prompt that causally influence model outputs through specific pathways without confounding — may enable new forms of model steering, alignment verification, and safety auditing. Causal inference on the data-generating process for model training can in principle identify which training examples causally drove which capabilities, enabling targeted data interventions for safety improvement.
Causal Fairness in Dynamic Systems: Static causal fairness criteria (counterfactual fairness, path-specific effects) are being extended to dynamic settings where algorithmic decisions affect population outcomes, which in turn affect future data. Causal long-term fairness frameworks must reason about the feedback loops created by algorithmic interventions, integrating causal RL, population dynamics modelling, and heterogeneous treatment effect estimation into a unified policy design framework. This is particularly important for Algorithmic Fairness in domains like credit, hiring, and criminal justice where repeated decisions shape population distributions.
Formal Methods and Algorithms
The Do-Calculus — Formal Statement: Pearl’s do-calculus consists of three rules that govern when intervention distributions can be simplified by removing do-operators:
-
Rule 1 (Insertion/deletion of observations): P(y | do(x), z, w) = P(y | do(x), w) if (Y ⊥ Z | X, W) in G_X (the graph with all incoming edges to X deleted).
-
Rule 2 (Action/observation exchange): P(y | do(x), do(z), w) = P(y | do(x), z, w) if (Y ⊥ Z | X, W) in G_{X,overbar{Z}} (the graph with incoming edges to X deleted and outgoing edges from Z deleted).
-
Rule 3 (Insertion/deletion of actions): P(y | do(x), do(z), w) = P(y | do(x), w) if (Y ⊥ Z | X, W) in G_{X, Z(W)} where Z(W) are Z-nodes not ancestors of any W-node in G_X.
These rules are sound and complete for identifying interventional distributions from observational data: if a query is identifiable, do-calculus can derive it; if it is not, do-calculus will fail to simplify it.
Back-Door Criterion: A set of variables Z satisfies the back-door criterion relative to (X, Y) in a DAG G if: (i) no node in Z is a descendant of X, and (ii) Z blocks every path between X and Y that contains an arrow into X (a “back-door path”). If Z satisfies the back-door criterion, then P(Y | do(X)) = Σ_z P(Y | X, Z=z) P(Z=z). This is the most widely applied identification result in practical causal analysis.
Propensity Score Theorem (Rosenbaum & Rubin, 1983): Under the ignorability assumption (Y(0), Y(1) ⊥ T | X), adjustment for the propensity score e(X) = P(T=1|X) suffices to remove confounding: (Y(0), Y(1)) ⊥ T | e(X). This allows covariate-dimensional reduction from many confounders X to a single scalar e(X), enabling matching and weighting on the propensity score without controlling for all individual covariates.
Doubly Robust Estimator: The augmented inverse probability weighted (AIPW) estimator of ATE is: ATE_DR = (1/n) Σ_i [ μ_1(X_i) - μ_0(X_i) + T_i(Y_i - μ_1(X_i))/e(X_i) - (1-T_i)(Y_i - μ_0(X_i))/(1-e(X_i)) ], where μ_t(X) = E[Y|T=t, X]. The estimator is doubly robust: it is consistent if either the outcome model μ or the propensity score e is correctly specified, but not necessarily both. This property is particularly valuable in high-dimensional settings.
Local Average Treatment Effect (LATE) Theorem: Under monotonicity (no defiers), the IV estimator with instrument Z equals LATE = E[Y(1) - Y(0) | T(1) > T(0)], the ATE for compliers — units who take treatment if and only if assigned to treatment by the instrument. This interpretational result (Imbens & Angrist, 1994) clarified what IV estimates identify in heterogeneous-treatment-effect settings, resolving decades of ambiguity.
Causal Forest Algorithm: Wager & Athey’s (2018) causal forest builds on random forests with two modifications: (i) honest splitting, where the estimation sample is split into a training half (used to determine splits) and an estimation half (used to compute leaf-level estimates), preventing overfitting; (ii) causal criterion for split quality, which optimises for heterogeneity in treatment effects rather than outcome prediction. Under regularity conditions, the resulting estimator achieves asymptotic normality and pointwise confidence intervals for the conditional average treatment effect τ(x) = E[Y(1) - Y(0) | X=x].
Cross-Disciplinary Connections
Causal inference bridges multiple academic disciplines, and each connection enriches the methodological toolkit:
- Epidemiology:
- Earliest large-scale application domain for causal methods in health research.
- Developed propensity score matching, regression discontinuity, DiD, and Mendelian randomisation methods in population health context.
- Contributed target trial emulation framework for rigorous observational study design.
- Hill’s criteria for causal inference (1965) remain influential: strength, consistency, temporality, biological gradient, plausibility, coherence, experiment, analogy.
- Modern epidemiology increasingly uses DAG-based confounding analysis rather than checklist-based approaches.
- Econometrics:
- Developed IV, LATE theorem, natural experiment designs, regression discontinuity, and high-dimensional methods.
- Nobel Prize 2021 validates causal econometrics as central to the empirical economics discipline.
- Structural econometrics tradition (Haavelmo, Cowles Commission) provides the foundational SCM-like framework for macroeconomic causal analysis.
- Policy evaluation is the primary application: education, healthcare, labour market, and environmental economics.
- Statistics:
- Formal potential outcomes framework; doubly robust estimation; semiparametric efficiency theory.
- Exact randomisation tests; Bayesian causal inference; sensitivity analysis for unmeasured confounders.
- Targeted Maximum Likelihood Estimation (TMLE): semiparametrically efficient doubly robust estimator with influence function-based variance estimation.
- Computer Science / AI:
- Causal discovery algorithms (PC, FCI, GES, LiNGAM, NOTEARS); causal representation learning; causal RL.
- Integration of causal methods with LLMs; neural causal models; causal fairness theory.
- Causal program synthesis: generating programs that perform specified causal computations from data.
- Philosophy of Science:
- Conceptual foundations of causation: manipulationist theories (Woodward), counterfactual theories (Lewis), mechanistic theories.
- Causal pluralism: different causal concepts may be appropriate for different scientific domains.
- Causal explanation and mechanism: how do we explain why something happened in causal terms?
- The interventionist criterion (Woodward, 2003): X causes Y if and only if an ideal intervention on X would change Y.
- Political Science:
- Regression discontinuity designs for electoral threshold effects; synthetic control for case study policy evaluation.
- Natural experiments in institutional design: federalism as a natural experiment; constitutional change as treatment.
- Causal inference for international relations: estimating effects of diplomatic interventions, sanctions, and treaties.
- Medicine / Clinical Trials:
- Adaptive clinical trial designs; personalised medicine HTEs; target trial emulation.
- Subgroup analysis with multiple testing correction; adaptive enrichment designs based on biomarker response.
- Pharmacovigilance signal detection: causal inference methods for post-marketing safety surveillance.
Benchmark Datasets and Evaluation Resources
Causal inference methods are evaluated against a combination of semi-synthetic datasets (observational covariates from real populations, synthetic outcome-generation mechanisms with known causal truth) and natural experimental archives:
Jobs (LaLonde, 1986 / Dehejia-Wahba, 1999): The canonical benchmark for propensity score methods. The National Supported Work (NSW) job training programme dataset consists of randomised trial participants augmented with observational comparison groups from the Current Population Survey and PSID. ATE on earnings at 18 months is used to compare estimators’ ability to recover the RCT benchmark from the observational subsample.
IHDP (Hill, 2011): The Infant Health and Development Programme dataset with semi-synthetic outcomes. Surface B covariates and 1000 synthetic realisations of outcome-generation mechanisms allow systematic comparison of heterogeneous treatment effect estimators including causal forests, Bayesian additive regression trees (BART), and meta-learners.
ACIC Data Challenge Benchmarks: The American Causal Inference Conference (ACIC) has run annual benchmarks (2016, 2017, 2019, 2022) with semi-synthetic data at varying levels of confounding, treatment effect heterogeneity, and overlap violation. These provide comprehensive comparisons of competing estimators under controlled conditions.
401(k) Eligibility (Poterba, Venti & Wise, 1994; Chernozhukov & Hansen, 2004): A standard IV benchmark. Eligibility for 401(k) pension plans (plausibly random conditional on income) is used as an instrument for 401(k) participation, with financial wealth as the outcome. Used to evaluate DML, TSLS, and related IV estimators.
Twins Dataset: A semi-synthetic benchmark based on US twin births (1989–1991) where the binary “treatment” is birth weight category and outcomes are infant mortality. The fact that twins share the same womb environment enables within-pair comparisons that isolate treatment from genetic and environmental confounders.
EconCausal Benchmark (2025): A new 2025 benchmark specifically evaluating LLMs’ causal reasoning capabilities on economics-style causal questions, testing the ability to identify confounders, reason about natural experiments, and perform counterfactual analysis in structured social science settings.
CausalVLBench (2025): A visual causal reasoning benchmark assessing whether large vision-language models can correctly identify causal relationships from visual scenarios, counterfactual images, and causal chain diagrams.
Key Terminology
- Causal Estimand: The precise mathematical quantity a causal study aims to estimate — e.g. ATE (average treatment effect over the whole population), ATT (average treatment effect on the treated), LATE (local ATE for compliers), CATE (conditional ATE for a subgroup with characteristics x).
- Confounding Variable / Confounder: A variable that causes both the treatment and the outcome, creating spurious associations between them. Uncontrolled confounders bias naive treatment effect estimates toward or away from zero.
- Counterfactual: A statement about what would have happened under a different assignment — e.g. “what would this patient’s outcome have been had they not received treatment?” Counterfactuals are the definitional objects in the Rubin Causal Model: individual treatment effects are defined as differences between factual and counterfactual outcomes.
- Directed Acyclic Graph (DAG): A graph with directed edges (arrows) and no directed cycles, used in Pearl’s framework to represent causal relationships. Nodes are variables; directed edges represent direct causal influence from parent to child. DAGs encode conditional independence restrictions via d-separation.
- D-Separation: A graphical criterion for determining conditional independence in a DAG. Variables X and Y are d-separated by a set Z if every path between them is blocked by Z (either via a chain or fork with a member in Z, or via a collider not in Z and with no descendant in Z).
- Do-Operator: Pearl’s formal notation for an external intervention. P(Y | do(X=x)) denotes the probability distribution of Y when X is set to x by external manipulation, as opposed to P(Y | X=x) which conditions on observing X=x (potentially a selection effect).
- Exchangeability / Ignorability / Unconfoundedness: The key identification assumption for observational causal inference in the potential outcomes framework: (Y(0), Y(1)) ⊥ T | X — potential outcomes are independent of treatment assignment conditional on observed covariates X. This means that conditional on X, treatment assignment is “as good as random.”
- Heterogeneous Treatment Effect (HTE): The variation in treatment effects across individuals or subgroups, characterised by the conditional average treatment effect CATE(x) = E[Y(1) - Y(0) | X=x]. Understanding HTE is essential for personalised medicine, targeted policy, and precision marketing.
- Instrumental Variable (IV): A variable Z that (i) causes the treatment T (relevance), (ii) affects the outcome Y only through its effect on T (exclusion restriction), and (iii) is not associated with unmeasured confounders of T and Y (exogeneity). Valid instruments allow identification of causal effects even in the presence of unmeasured confounders.
- Mediation Analysis: Decomposing a total causal effect into direct effects (X → Y pathways that do not pass through mediator M) and indirect effects (X → M → Y pathways), enabling understanding of causal mechanisms rather than just total effects.
- Natural Experiment: A setting where treatment assignment is determined by a natural process (geographic variation, policy cutoffs, lottery allocation) that is plausibly exogenous, enabling quasi-experimental causal identification without researcher-controlled randomisation.
- Potential Outcomes: The two hypothetical outcome values Y(1) and Y(0) that would occur for each unit under treatment and control respectively. Only one is observed for any given unit — the “fundamental problem of causal inference.” The individual treatment effect ITE = Y(1) - Y(0) is therefore never directly observable.
- Structural Causal Model (SCM): A tuple (U, V, F, P_U) where V are endogenous variables, U are exogenous noise variables, F is a set of structural equations specifying each V_i as a function of its parents and noise, and P_U is a distribution over U. SCMs encode both observational and interventional distributions.
Research and Literature
- Pearl, J. (2009). Causality: Models, Reasoning, and Inference (2nd ed.). Cambridge University Press.
- Pearl, J. & Mackenzie, D. (2018). The Book of Why: The New Science of Cause and Effect. Basic Books.
- Rubin, D.B. (1974). “Estimating Causal Effects of Treatments in Randomized and Nonrandomized Studies.” Journal of Educational Psychology, 66(5), 688–701.
- Neyman, J. (1923). “On the Application of Probability Theory to Agricultural Experiments.” Statistical Science (transl. 1990), 5(4), 465–472.
- Imbens, G.W. & Angrist, J.D. (1994). “Identification and Estimation of Local Average Treatment Effects.” Econometrica, 62(2), 467–475.
- Angrist, J.D. & Krueger, A.B. (1991). “Does Compulsory School Attendance Affect Schooling and Earnings?” Quarterly Journal of Economics, 106(4), 979–1014.
- Card, D. (1990). “The Impact of the Mariel Boatlift on the Miami Labor Market.” Industrial and Labor Relations Review, 43(2), 245–257.
- Wager, S. & Athey, S. (2018). “Estimation and Inference of Heterogeneous Treatment Effects Using Random Forests.” Journal of the American Statistical Association, 113(523), 1228–1242.
- Chernozhukov, V. et al. (2018). “Double/Debiased Machine Learning for Treatment and Structural Parameters.” Econometrics Journal, 21(1), C1–C68.
- Hernán, M.A. & Robins, J.M. (2020). Causal Inference: What If. Chapman & Hall/CRC.
- Imbens, G.W. & Rubin, D.B. (2015). Causal Inference for Statistics, Social, and Biomedical Sciences. Cambridge University Press.
- Wright, S. (1921). “Correlation and Causation.” Journal of Agricultural Research, 20(7), 557–585.
- Haavelmo, T. (1944). “The Probability Approach in Econometrics.” Econometrica, 12(Suppl.), 1–115.
- Holland, P.W. (1986). “Statistics and Causal Inference.” Journal of the American Statistical Association, 81(396), 945–960.
- VanderWeele, T.J. & Ding, P. (2017). “Sensitivity Analysis in Observational Research: Introducing the E-Value.” Annals of Internal Medicine, 167(4), 268–274.
- Kusner, M.J., Loftus, J., Russell, C., & Silva, R. (2017). “Counterfactual Fairness.” NeurIPS 2017.
- Abadie, A., Diamond, A., & Hainmueller, J. (2010). “Synthetic Control Methods for Comparative Case Studies.” Journal of the American Statistical Association, 105(490), 493–505.
- Davey Smith, G. & Hemani, G. (2014). “Mendelian Randomisation: Genetic Anchors for Causal Inference in Epidemiological Studies.” Human Molecular Genetics, 23(R1), R89–R98.
- Huang, L. et al. (2025). “A Survey on Causal Reinforcement Learning.” IEEE Transactions on Neural Networks and Learning Systems (TNNLS).
- Rehill, P.M. et al. (2025). “How Do Applied Researchers Use the Causal Forest? A Methodological Review.” International Statistical Review, doi:10.1111/insr.12610.
- Imai, K. & Li, M. (2025). “Causal Representation Learning with Generative Artificial Intelligence: Application to Texts as Treatments.” Harvard Faculty, imai.fas.harvard.edu.
- Zhou, Y. et al. (2025). “Causal Methods for LLM Development and Evaluation.” arXiv:2605.25998.
- Pearl, J. (2009). “Causal Inference in Statistics: An Overview.” Statistics Surveys, 3, 96–146. (cs.columbia.edu reprint)
- Nobel Committee for Economic Sciences. (2021). “Scientific Background: Natural Experiments Help Answer Important Questions for Society.” The Royal Swedish Academy of Sciences.
- Williamson, E.J., Forbes, A., & White, I.R. (2018). “Variance Reduction in Randomised Trials by Inverse Probability Weighting Using the Propensity Score.” Statistics in Medicine, 33(5), 721–737.
- Oxera. (2021). “Causality and Natural Experiments: The 2021 Nobel Prize in Economic Sciences.” oxera.com.
- Causal AI Research Group, University of Edinburgh. (2025). “Causal Discovery with LLMs: Unified Frameworks for Evidence Triangulation.” Preprint.
- MRC Integrative Epidemiology Unit, University of Bristol. (2022). “Causal Attribution Fractions, and the Attribution of Smoking and BMI to the Landscape of Disease Incidence in UK Biobank.” medRxiv / PMC9667855.
Identification Assumptions and Their Violation
Every causal inference method rests on identification assumptions that cannot be directly verified from data — they require domain knowledge, theoretical reasoning, or institutional design. Understanding when these assumptions are violated and how violations affect conclusions is a core competency of the field:
- Ignorability / Unconfoundedness:
- Violated when unmeasured variables cause both treatment and outcome.
- Mendelian randomisation addresses this by using genetic variants (assumed exogenous) as instruments.
- Sensitivity analysis using E-values (VanderWeele & Ding, 2017): minimum unmeasured confounding needed to explain away the observed effect.
- Diagnostic checks: testing balance of observed covariates between treated and control after adjustment; testing whether propensity score model predicts known non-confounders.
- Stable Unit Treatment Value Assumption (SUTVA):
- Violated when treatment of one unit affects outcomes of other units (spillovers, general equilibrium effects).
- Example violations: vaccination campaigns (herd immunity creates spillovers); price interventions (market equilibrium effects).
- Methods for interference: partial population designs; network causal inference; weighted network estimators.
- Exclusion Restriction (IV):
- Violated when the instrument affects the outcome through pathways other than the treatment.
- Mendelian randomisation violations: horizontal pleiotropy (genetic variant affects multiple biological pathways).
- Tests: MR-Egger test for systematic pleiotropy; weighted median estimator robust to minority IV violations.
- Monotonicity (IV):
- Violated when some units are “defiers” — they take treatment if and only if the instrument assigns them to control.
- Empirically hard to test; domain knowledge about mechanisms is the primary tool.
- Overlap / Positivity:
- Violated when some covariate patterns are observed only in treated or only in control units.
- Practical trimming strategies: restrict inference to regions of common support; use doubly robust estimators that are less sensitive to poor propensity score overlap.
- Faithfulness:
- In causal discovery: violated when causal effects cancel exactly, creating conditional independencies that do not correspond to graph structure.
- Rare in practice but creates theoretical identification failures; addressed by robust PC algorithm variants.