Update repo
This commit is contained in:
@@ -0,0 +1,202 @@
|
||||
---
|
||||
title: "Full Emulation Pipeline"
|
||||
format:
|
||||
html:
|
||||
embed-resources: true
|
||||
docx: default
|
||||
execute:
|
||||
echo: true
|
||||
warning: false
|
||||
message: false
|
||||
---
|
||||
|
||||
## Goal
|
||||
|
||||
This notebook puts the entire target trial emulation pipeline together in one place:
|
||||
|
||||
1. Simulate the observational data.
|
||||
2. Define eligibility at time zero.
|
||||
3. Clone eligible patients into treatment arms.
|
||||
4. Apply censoring rules at the 2-hour mark.
|
||||
5. Compute IPTW + IPCW weights.
|
||||
6. Estimate the per-protocol effect.
|
||||
|
||||
Then we compare all four methods we have learned so far.
|
||||
|
||||
## Setup
|
||||
|
||||
```{r}
|
||||
suppressPackageStartupMessages({
|
||||
library(dplyr)
|
||||
library(gt)
|
||||
library(tibble)
|
||||
library(tidyr)
|
||||
})
|
||||
|
||||
source("../R/simulate_icu_cohort.R")
|
||||
source("../R/simulate_icu_cohort_longitudinal.R")
|
||||
source("../R/estimate_naive_vasopressor_mortality_effect.R")
|
||||
source("../R/estimate_standardized_vasopressor_mortality_effect.R")
|
||||
source("../R/estimate_iptw_vasopressor_mortality_effect.R")
|
||||
source("../R/clone_trial_arms.R")
|
||||
source("../R/estimate_iptw_and_ipcw_effect.R")
|
||||
```
|
||||
|
||||
## Step 1: Simulate Data
|
||||
|
||||
We generate both the cross-sectional and longitudinal datasets using the same seed.
|
||||
|
||||
```{r}
|
||||
icu_data <- simulate_icu_cohort(n_patients = 1000, seed = 20260531)
|
||||
longitudinal_data <- simulate_icu_cohort_longitudinal(n_patients = 1000, seed = 20260531)
|
||||
```
|
||||
|
||||
## Step 2: Cross-Sectional Estimates
|
||||
|
||||
### Naive comparison
|
||||
|
||||
```{r}
|
||||
naive_analysis <- estimate_naive_vasopressor_mortality_effect(icu_data)
|
||||
```
|
||||
|
||||
### Outcome regression standardization
|
||||
|
||||
```{r}
|
||||
standardized_analysis <- estimate_standardized_vasopressor_mortality_effect(icu_data)
|
||||
```
|
||||
|
||||
### Cross-sectional IPTW
|
||||
|
||||
```{r}
|
||||
iptw_analysis <- estimate_iptw_vasopressor_mortality_effect(icu_data)
|
||||
```
|
||||
|
||||
## Step 3: Longitudinal Emulation
|
||||
|
||||
### Clone and censor
|
||||
|
||||
```{r}
|
||||
clones <- clone_trial_arms(longitudinal_data)
|
||||
```
|
||||
|
||||
### Compute IPTW + IPCW weights and estimate effect
|
||||
|
||||
```{r}
|
||||
longitudinal_analysis <- estimate_iptw_and_ipcw_effect(clones$clone_baseline)
|
||||
```
|
||||
|
||||
## Four-Way Comparison
|
||||
|
||||
We now extract the risk difference and risk ratio from each method and place them in one table.
|
||||
|
||||
```{r}
|
||||
# Naive
|
||||
naive_rd <- naive_analysis$mortality_effect_estimates |>
|
||||
filter(estimate == "Naive risk difference") |>
|
||||
pull(value)
|
||||
|
||||
naive_rr <- naive_analysis$mortality_effect_estimates |>
|
||||
filter(estimate == "Naive risk ratio") |>
|
||||
pull(value)
|
||||
|
||||
# Standardized
|
||||
std_rd <- standardized_analysis$standardized_mortality_effect_estimates |>
|
||||
filter(estimate == "Standardized risk difference") |>
|
||||
pull(value)
|
||||
|
||||
std_rr <- standardized_analysis$standardized_mortality_effect_estimates |>
|
||||
filter(estimate == "Standardized risk ratio") |>
|
||||
pull(value)
|
||||
|
||||
# Cross-sectional IPTW
|
||||
iptw_rd <- iptw_analysis$iptw_mortality_effect_estimates |>
|
||||
filter(estimate == "IPTW risk difference") |>
|
||||
pull(value)
|
||||
|
||||
iptw_rr <- iptw_analysis$iptw_mortality_effect_estimates |>
|
||||
filter(estimate == "IPTW risk ratio") |>
|
||||
pull(value)
|
||||
|
||||
# Longitudinal IPTW + IPCW
|
||||
long_rd <- longitudinal_analysis$weighted_effect_estimates |>
|
||||
filter(estimate == "IPTW + IPCW risk difference") |>
|
||||
pull(value)
|
||||
|
||||
long_rr <- longitudinal_analysis$weighted_effect_estimates |>
|
||||
filter(estimate == "IPTW + IPCW risk ratio") |>
|
||||
pull(value)
|
||||
|
||||
four_way <- tibble(
|
||||
method = c(
|
||||
"1. Naive observed comparison",
|
||||
"2. Outcome regression standardization",
|
||||
"3. Cross-sectional IPTW",
|
||||
"4. Longitudinal IPTW + IPCW (per-protocol)"
|
||||
),
|
||||
risk_difference = c(naive_rd, std_rd, iptw_rd, long_rd),
|
||||
risk_ratio = c(naive_rr, std_rr, iptw_rr, long_rr)
|
||||
)
|
||||
|
||||
four_way |>
|
||||
gt() |>
|
||||
tab_header(title = "Four-Way Comparison of Effect Estimation Methods") |>
|
||||
cols_label(
|
||||
method = "Method",
|
||||
risk_difference = "Risk difference",
|
||||
risk_ratio = "Risk ratio"
|
||||
) |>
|
||||
fmt_number(columns = c(risk_difference, risk_ratio), decimals = 3)
|
||||
```
|
||||
|
||||
## Interpretation
|
||||
|
||||
### What each method answers
|
||||
|
||||
| Method | Question it answers |
|
||||
|--------|-------------------|
|
||||
| **Naive** | "What was the mortality difference between patients who did and did not receive early vasopressors?" |
|
||||
| **Standardization** | "What would the mortality risk be if everyone followed each strategy, holding the eligible population fixed?" |
|
||||
| **Cross-sectional IPTW** | "What is the effect of receiving early vasopressors versus not, reweighted to balance confounders?" |
|
||||
| **Longitudinal IPTW + IPCW** | "What is the per-protocol effect of the early vasopressor strategy versus the no early strategy, respecting the grace period and accounting for censoring?" |
|
||||
|
||||
### What to notice in the simulated data
|
||||
|
||||
In the data-generating process, early vasopressors have a modest **protective** effect.
|
||||
|
||||
1. The **naive** estimate shows **harm** (positive risk difference, risk ratio > 1). This is because sicker patients are more likely to receive early vasopressors and more likely to die — confounding by indication.
|
||||
|
||||
2. **Standardization** and **cross-sectional IPTW** both move the estimate toward a **protective** direction, correcting for baseline confounding.
|
||||
|
||||
3. The **longitudinal IPTW + IPCW** estimate may differ slightly from cross-sectional IPTW because it:
|
||||
- Restricts to the per-protocol population (patients who could adhere through the 2-hour window)
|
||||
- Re-weights for informative censoring (patients censored due to protocol deviation or early death)
|
||||
- Respects the grace period as a longitudinal treatment assignment window
|
||||
|
||||
None of the methods perfectly recover the true effect in a single finite sample, but the longitudinal method most closely emulates the target trial design.
|
||||
|
||||
## Assumption Checklist
|
||||
|
||||
Every method above relies on assumptions. A quick checklist:
|
||||
|
||||
| Assumption | Naive | Standardized | IPTW | IPTW + IPCW |
|
||||
|------------|-------|-------------|------|-------------|
|
||||
| No unmeasured confounding | ❌ | ✅ | ✅ | ✅ |
|
||||
| Positivity (every patient has some probability of each treatment) | ❌ | ✅ | ✅ | ✅ |
|
||||
| Correct model specification | N/A | ✅ | ✅ | ✅ |
|
||||
| No informative censoring | N/A | N/A | N/A | ✅ |
|
||||
|
||||
- **No unmeasured confounding**: We must have measured all variables that affect both treatment assignment and mortality. In a real study, this is never fully testable.
|
||||
- **Positivity**: Every patient must have a non-zero probability of receiving each treatment. Extreme propensity scores cause unstable weights.
|
||||
- **Correct model specification**: The logistic regression models must capture the true relationships between confounders, treatment, and outcome.
|
||||
- **No informative censoring (for IPCW)**: Censoring must be independent of the outcome given measured confounders. In practice, we use baseline-only IPCW here; a full analysis would include time-varying covariates.
|
||||
|
||||
## Next Steps
|
||||
|
||||
From here, the tutorial can extend into several directions:
|
||||
|
||||
1. **Time-varying confounders**: Include `map_current` and `lactate_current` in the censoring model to handle post-baseline severity changes.
|
||||
2. **Survival analysis**: Model time-to-death rather than 28-day binary mortality, using Kaplan-Meier or Cox models with IPCW.
|
||||
3. **Sensitivity analysis**: Test how sensitive the results are to unmeasured confounding (e.g., E-values).
|
||||
4. **`targets` pipeline**: Convert the analysis into a reproducible pipeline with dependency tracking.
|
||||
|
||||
These topics move the project from a teaching scaffold toward a production-ready target trial emulation workflow.
|
||||
Reference in New Issue
Block a user