BINF 4008:004 · Fall 2026

Generating Real-World Evidence in Medicine

A course on thoughtful observational study design.

Course description

Healthcare systems generate enormous volumes of data in various forms, with electronic health records and administrative claims in particular serving as rich sources of information. These real-world health data sources are potential reservoirs of insights for biomedical discovery. Real-world data, however, arises from a data-generating process, shaped by clinical workflow and billing practices, that is distinct from the true biomedical processes we aim to study. This can make distilling reliable insights from these observational data sources challenging. This course introduces the theoretical foundations and practical methods required for addressing these challenges through rigorous observational study design using real-world health data that can be used to generate reliable evidence to answer a range of clinical questions.

Each course topic is developed through three layers of theory, application in informatics, and implementation through tools developed by the Observational Health and Data Sciences and Informatics (OHDSI) consortium, which has developed analytic tools based on best practices for high-quality observational studies. The theory for this course is developed around the frameworks of probabilistic and causal inference, to facilitate understanding of study-design principles and analytic approaches that are thoughtfully informed by domain knowledge. We use theory to understand the challenges and biases of real-world health data, and how observational study design choices and methods can address them. Each stage of the observational study pipeline is covered including formulating a precise research question, identifying the right data, defining study populations and phenotypes, and determining and executing the appropriate analyses.

Tools and conventions developed by the OHDSI community will serve as a practical framework for implementing these concepts. We will examine how OHDSI tools can be used to quickly translate a well-designed research question into an executable study and to evaluate whether the resulting estimates are credible. Each class meeting includes both lecture and hands-on exercises so that concepts are translated into practice as they are introduced.

Learning objectives

Theory

Probabilistic & causal inference

  • Distinguish between associational, interventional, and counterfactual queries
  • Construct causal diagrams representing assumptions about a data-generating process and sources of bias
  • Analyze a Bayesian network to infer independencies and dependencies
  • Evaluate causal effect identifiability from a diagram using key identification criteria
Application

Informatics for observational studies

  • Articulate potential sources of bias in EHR and claims databases
  • Formulate precise research questions with clearly specified study components
  • Define phenotypes and study cohorts that minimize measurement and other biases
  • Determine the appropriate analysis for a specified question
Tools

The OHDSI workflow

  • Understand the OMOP common data model and standardized vocabularies
  • Construct cohort definitions in ATLAS and evaluate phenotypes
  • Execute characterization, prediction, and estimation studies with Strategus
  • Evaluate results using diagnostics

Schedule

Twelve weeks, each moving from theory to application to tool. Week titles to come.

Theory
Application
Tool
01
9/8/26
Introduction
  • Motivation and introduction to observational studies in informatics
  • Pearl's causal hierarchy
  • Introduction to OHDSI
02
9/15/26
Understanding Data-Generating Processes
  • Introduction to structural causal models (SCMs)
  • Domain context: EMR vs. claims data, and the OMOP common data model / vocabularies
  • ATLAS overview
03
9/22/26
Encoding Knowledge
  • From structural causal models to causal diagrams
  • Phenotyping: from diagrams to data
  • ATLAS: cohort construction, part 1
04
9/29/26
  • d-Separation with causal diagrams, part 1 — determining probabilities from diagrams
  • Formulating associational queries (characterization)
  • Strategus Characterization package
05
10/6/26
  • d-Separation with causal diagrams, part 2 — reading independence structure
  • Formulating associational queries (prediction)
  • Strategus Patient-level prediction package
06
10/13/26
  • Identification, part 1 — backdoor criterion and adjustment sets
  • Graphically encoding biases in medical data, part 1
  • ATLAS: cohort construction, part 2
07
10/20/26
  • Identification, part 2 — conditional backdoor, frontdoor criterion
  • Graphically encoding biases in medical data, part 2
  • Strategus CohortDiagnostics
08
10/27/26
  • From identification to estimation — propensity-score analysis
  • Cohort study design
  • Strategus Population-level estimation (CohortMethod) package
—
11/2/26
Election day — no class
—
11/10/26
AMIA — no class. Sign up for a 30-minute 1:1 final project consult.
09
11/17/26
  • Bias evaluation and diagnostics
  • Self-controlled case series study design
  • Strategus SCCS package
10
11/24/26
  • Transportability and data fusion
  • Meta-analyses and network studies
  • Strategus EvidenceSynthesis tool
11
12/1/26
  • Project presentation
12
12/8/26
  • Special topics in observational research

Format & grading

Weekly assignments

25% of grade

Putting theory and application into practice with OHDSI tools. Assignments not finished in class are due before the next session.

Final project

50% of grade

An independent or paired OHDSI-style observational study, developed in consultation with the instructor, answering a clinically relevant research question of your choosing.

Participation & attendance

25% of grade

Attendance is expected at every session (two weeks' notice for planned absences; 24 hours or ASAP for illness). Assessed through active engagement in lecture and lab.

Final project

Work independently or in pairs to design and execute one or more OHDSI-style observational studies. The goal is to leave the course with a manuscript draft ready to submit to a clinical journal.

A

Journal manuscript — 3,000 to 4,000 words

A journal-style research report formatted for a clinical journal, presenting your research question, analytical approach, results, and interpretation of findings.

  1. Introduction / Background
  2. Methods
  3. Results
  4. Discussion
  5. Conclusion
B

Study design & execution reflection — 1 to 3 pages

A reflection piece describing a more thorough account of your research process, including intermediate steps that won't be reflected in the manuscript itself.

  1. What informed your phenotype and study design choices? What alternative choices could you have made, and why did you settle on the choices you did?
  2. What decisions did you make initially that may not have worked? Why did they require refinement?
  3. How did you evaluate your choices along the way?

Policies

Data Use & Security

You will access to and analyze real patient health data for educational purposes. This data is protected under HIPAA, and every student is bound by the corresponding standards of appropriate use.

  1. Data is accessed and analyzed through direct database access only. Do not download, copy, print, screenshot, or store any data on personal devices, USB drives, cloud storage, or anywhere beyond the approved database.
  2. Patient data or your database access credentials may not be shared with another individual or GenAI tool. You may not upload patient data directly to a GenAI tool or share your database access credentials. This includes copying code to GenAI where your credentials are hard-coded and using agentic AI that can read and access files where your credentials are stored.
  3. When you need to review individual records to understand data or verify an analysis, do so to the minimum extent necessary.
  4. If you suspect data may have been compromised, report it to the instructor immediately. There is no penalty for promptly reported incidents.

Generative AI

You're generally encouraged to use GenAI tools as you see fit throughout the course, with the data-security exceptions above. For any submitted work, you'll be asked to describe with some specificity how you used GenAI. If your use appears to have hindered your understanding of key concepts or prevented you from submitting accurate, high-quality work, you'll meet with the instructor to discuss it and revise and resubmit with that guidance.

Lectures

Recording or transcribing lectures with any tool isn't permitted, except with written approval from the accommodations office. Slides are posted after each class.

Weekly assignments

Use GenAI to debug code or otherwise support your work, as long as it doesn't hinder your comprehension or compromise data security, including tools with full file access, such as Claude Code, that could read your database connection details directly.

Final project

Same allowances and exceptions as weekly assignments. You're encouraged to develop your research ideas, write-ups, and slides on your own as each has requirements of accuracy and specificity you're unlikely to meet without your own critical thinking.

Textbook & resources

The Book of OHDSIrequired
Causal Artificial Intelligencerequired
The Book of Whyoptional

Logistics

Class sessions
Tuesdays, 2:10–4:00pm
Location
PH20-200
Instructor
Tara Anand — tara.v.anand@columbia.edu
Teaching assistant
Muying Li — ml5218@cumc.columbia.edu
Office hours
By appointment