@Zen_with_AI: ๐—” ๐—™๐—ถ๐—ฟ๐˜€๐˜ ๐—–๐—ผ๐˜‚๐—ฟ๐˜€๐—ฒ ๐—ถ๐—ป ๐—–๐—ฎ๐˜‚๐˜€๐—ฎ๐—น ๐—œ๐—ป๐—ณ๐—ฒ๐—ฟ๐—ฒ๐—ป๐—ฐ๐—ฒ (Free 490-Page Textbook) Author Peng Ding, Professor inโ€ฆ

X AI KOLs Timeline Papers

Summary

A free 490-page textbook on causal inference by Peng Ding, covering topics from correlation to longitudinal data with accompanying R code and datasets.

๐—” ๐—™๐—ถ๐—ฟ๐˜€๐˜ ๐—–๐—ผ๐˜‚๐—ฟ๐˜€๐—ฒ ๐—ถ๐—ป ๐—–๐—ฎ๐˜‚๐˜€๐—ฎ๐—น ๐—œ๐—ป๐—ณ๐—ฒ๐—ฟ๐—ฒ๐—ป๐—ฐ๐—ฒ (Free 490-Page Textbook) Author Peng Ding, Professor in the Department of Statistics at UC Berkeley, this book is a compilation of his lecture notes from seven years of teaching causal inference courses, corresponding to the undergraduate course Stat 156 and the graduate course Stat 256, so the content only requires a foundation in probability theory, statistical inference, linear and logistic regressionโ€”undergrads can tackle it too. It kicks off with the Yule-Simpson paradox, first drawing the line between correlation and causation, then unfolding into topics like identification, estimation, and longitudinal data. The accompanying R code and datasets are hosted on Harvard Dataverse, available for free download under the CC BY 4.0 license. If you want another reference book to round things out, Kirk Borne also recommends Judea Pearl's *Causal Inference in Statistics: A Primer*โ€”read them together: one teaches you how to do it, the other explains why you think that way. Link:
Original Article
View Cached Full Text

Cached at: 09/11/26, 12:27 AM

A First Course in Causal Inference (Free 490-Page Textbook) Author Peng Ding, Professor in the Department of Statistics at UC Berkeley, this book is a compilation of his lecture notes from seven years of teaching causal inference courses, corresponding to the undergraduate course Stat 156 and the graduate course Stat 256, so the content only requires a foundation in probability theory, statistical inference, linear and logistic regressionโ€”undergrads can tackle it too. It kicks off with the Yule-Simpson paradox, first drawing the line between correlation and causation, then unfolding into topics like identification, estimation, and longitudinal data. The accompanying R code and datasets are hosted on Harvard Dataverse, available for free download under the CC BY 4.0 license. If you want another reference book to round things out, Kirk Borne also recommends Judea Pearlโ€™s Causal Inference in Statistics: A Primerโ€”read them together: one teaches you how to do it, the other explains why you think that way. Link:


A First Course in Causal Inference

Source: https://arxiv.org/html/2305.18793

To students and readers who are interested in causal inference

Contents
  1. I. Introduction
    1. 1. Correlation, Association, and the Yuleโ€“Simpson Paradox
      1. 1.1 Traditional view of statistics
      2. 1.2 Some commonly-used measures of association
        1. 1.2.1 Correlation and regression
        2. 1.2.2 Contingency tables
      3. 1.3 An example of the Yuleโ€“Simpson Paradox
        1. 1.3.1 Data
        2. 1.3.2 Explanation
        3. 1.3.3 Geometry of the Yuleโ€“Simpson Paradox
      4. 1.4 The Berkeley graduate school admission data
      5. 1.5 Homework Problems
    2. 2. Potential Outcomes
      1. 2.6 Experimentalistsโ€™ view of causal inference
      2. 2.7 Formal notation of potential outcomes
        1. 2.7.1 Causal effects, subgroups, and the non-existence of Yuleโ€“Simpson Paradox
        2. 2.7.2 Subtlety of the definition of the experimental unit
      3. 2.8 Treatment assignment mechanism
      4. 2.9 Homework Problems
  2. II. Randomized experiments
    1. 3. The Completely Randomized Experiment and the Fisher Randomization Test
      1. 3.10 CRE
      2. 3.11 FRT
      3. 3.12 Canonical choices of the test statistic
      4. 3.13 A case study of the LaLonde experimental data
      5. 3.14 Some history of randomized experiments and FRT
        1. 3.14.1 James Lindโ€™s experiment
        2. 3.14.2 Lady tasting tea
        3. 3.14.3 Two Fisherian principles for experiments
      6. 3.15 Discussion
        1. 3.15.1 Other sharp null hypotheses and confidence intervals
        2. 3.15.2 Other test statistics
        3. 3.15.3 Final remarks
      7. 3.16 Homework Problems
    2. 4. Neymanian Repeated Sampling Inference in Completely Randomized Experiments
      1. 4.17 Finite population quantities
      2. 4.18 Neyman, (1923)โ€™s theorem
      3. 4.19 Proofs
      4. 4.20 Regression analysis of the CRE
      5. 4.21 Examples
        1. 4.21.1 Simulation
        2. 4.21.2 Heavy-tailed outcome and failure of Normal approximations
        3. 4.21.3 Application
      6. 4.22 Homework Problems
    3. 5. Stratification and Post-Stratification in Randomized Experiments
      1. 5.23 Stratification
      2. 5.24 FRT
        1. 5.24.1 Theory
        2. 5.24.2 An application
      3. 5.25 Neymanian inference
        1. 5.25.1 Point and interval estimation
        2. 5.25.2 Numerical examples
        3. 5.25.3 Comparing the SRE and the CRE
      4. 5.26 Post-stratification in a CRE
        1. 5.26.1 Meinert et al., (1970)โ€™s Example
        2. 5.26.2 Chong et al., (2016)โ€™s Example
      5. 5.27 Practical questions
      6. 5.28 Homework Problems
    4. 6. Rerandomization and Regression Adjustment
      1. 6.29 Rerandomization
        1. 6.29.1 Experimental design
        2. 6.29.2 Statistical inference
      2. 6.30 Regression adjustment
        1. 6.30.1 Covariate-adjusted FRT
        2. 6.30.2 Analysis of covariance and extensions
          1. 6.30.2.1 Some heuristics for Lin, (2013)โ€™s results
          2. 6.30.2.2 Understanding Lin, (2013)โ€™s estimator via predicting the potential outcomes
          3. 6.30.2.3 Understanding Lin, (2013)โ€™s estimator via adjusting for covariate imbalance
        3. 6.30.3 Some additional remarks on regression adjustment
          1. 6.30.3.1 Duality between ReM and regression adjustment
          2. 6.30.3.2 Equivalence of regression adjustment and post-stratification
          3. 6.30.3.3 Difference-in-difference as a special case of covariate adjustment ฯ„ฬ‚(ฮฒโ‚,ฮฒโ‚€)
        4. 6.30.4 Extension to the SRE
      3. 6.31 Unification, combination, and comparison
      4. 6.32 Simulation
      5. 6.33 Final remarks
      6. 6.34 Homework Problems
    5. 7. Matched-Pairs Experiment
      1. 7.35 Design of the experiment and potential outcomes
      2. 7.36 FRT
      3. 7.37 Neymanian inference
      4. 7.38 Covariate adjustment
        1. 7.38.1 FRT
        2. 7.38.2 Regression adjustment
      5. 7.39 Examples
        1. 7.39.1 Darwinโ€™s data comparing cross-fertilizing and self-fertilizing on the height of corns
        2. 7.39.2 Childrenโ€™s television workshop experiment data
      6. 7.40 Comparing the MPE and CRE
      7. 7.41 Extension to the general matched experiment
        1. 7.41.1 FRT
        2. 7.41.2 Estimating the average of the within-strata effects
        3. 7.41.3 A more general causal estimand
      8. 7.42 Homework Problems
    6. 8. Unification of the Fisherian and Neymanian Inferences in Randomized Experiments
      1. 8.43 Testing strong and weak null hypotheses in the CRE
      2. 8.44 Covariate-adjusted FRTs in the CRE
      3. 8.45 A simulation study
      4. 8.46 General recommendations
      5. 8.47 A case study
      6. 8.48 Homework Problems
    7. 9. Bridging Finite and Super Population Causal Inference
      1. 9.49 CRE
      2. 9.50 Simulation under the CRE: the super population perspective
      3. 9.51 Extension to the SRE
      4. 9.52 Homework Problems
  3. III. Observational studies
    1. 10. Observational Studies, Selection Bias, and Nonparametric Identification of Causal Effects
      1. 10.53 Motivating Examples
      2. 10.54 Causal effects and selection bias under the potential outcomes framework
      3. 10.55 Sufficient conditions for nonparametric identification of causal effects
        1. 10.55.1 Identification
        2. 10.55.2 Plausibility of the ignorability assumption
      4. 10.56 Two simple estimation strategies and their limitations
        1. 10.56.1 Stratification or standardization based on discrete covariates
        2. 10.56.2 Outcome regression
      5. 10.57 Homework Problems
    2. 11. The Central Role of the Propensity Score in Observational Studies for Causal Effects
      1. 11.58 The propensity score as a dimension reduction tool
        1. 11.58.1 Theory
        2. 11.58.2 Propensity score stratification
        3. 11.58.3 Application
      2. 11.59 Propensity score weighting
        1. 11.59.1 Theory
        2. 11.59.2 Inverse propensity score weighting estimators
        3. 11.59.3 A problem of IPW and a fundamental problem of causal inference
        4. 11.59.4 Application
      3. 11.60 The balancing property of the propensity score
        1. 11.60.1 Theory
        2. 11.60.2 Covariate balance check
      4. 11.61 Homework Problems
    3. 12. The Doubly Robust or the Augmented Inverse Propensity Score Weighting Estimator for the Average Causal Effect
      1. 12.62 The doubly robust estimator
        1. 12.62.1 Population version
        2. 12.62.2 Sample version
      2. 12.63 More intuition and theory for the doubly robust estimator
        1. 12.63.1 Reducing the variance of the IPW estimator
        2. 12.63.2 Reducing the bias of the outcome regression estimator
      3. 12.64 Examples
        1. 12.64.1 Summary of some canonical estimators for ฯ„
        2. 12.64.2 Simulation
        3. 12.64.3 Applications
      4. 12.65 Some further discussion
      5. 12.66 Homework problems
    4. 13. The Average Causal Effect on the Treated Units and Other Estimands
      1. 13.67 Nonparametric identification of ฯ„_T
      2. 13.68 Inverse propensity score weighting and doubly robust estimation of ฯ„_T
      3. 13.69 An example
      4. 13.70 Other estimands
      5. 13.71 Homework Problems
    5. 14. Using the Propensity Score in Regressions for Causal Effects
      1. 14.72 Regressions with the propensity score as a covariate
      2. 14.73 Regressions weighted by the inverse of the propensity score
        1. 14.73.1 Average causal effect
        2. 14.73.2 Average causal effect on the treated units
      3. 14.74 Homework problems
    6. 15. Matching in Observational Studies
      1. 15.75 A simple starting point: many more control units
      2. 15.76 A more complicated but realistic scenario
      3. 15.77 Matching estimator for the average causal effect
        1. 15.77.1 Point estimation and bias correction
        2. 15.77.2 Connection with the doubly robust estimators
      4. 15.78 Matching estimator for the average causal effect on the treated
      5. 15.79 A case study
        1. 15.79.1 Experimental data
        2. 15.79.2 Observational data
        3. 15.79.3 Covariate balance checks
      6. 15.80 Discussion
      7. 15.81 Homework Problems
  4. IV. Difficulties and challenges of observational studies
    1. 16. Difficulties of Unconfoundedness in Observational Studies for Causal Effects
      1. 16.82 Some basics of the causal diagram
      2. 16.83 Assessing the unconfoundedness assumption
        1. 16.83.1 Using negative outcomes
        2. 16.83.2 Using negative exposures
        3. 16.83.3 Summary
      3. 16.84 Problems of over-adjustment
        1. 16.84.1 M-bias
        2. 16.84.2 Z-bias
        3. 16.84.3 What covariates should we adjust for in observational studies?
      4. 16.85 Homework Problems
    2. 17. E-Value: Evidence for Causation in Observational Studies with Unmeasured Confounding

Similar Articles