--- title: "Rotating panels: net change, chaining and gross flows" output: rmarkdown::html_vignette vignette: > %\VignetteIndexEntry{Rotating panels: net change, chaining and gross flows} %\VignetteEngine{knitr::rmarkdown} %\VignetteEncoding{UTF-8} --- ```{r setup, include = FALSE} knitr::opts_chunk$set(collapse = TRUE, comment = "#>") library(weightflow) set.seed(20260910) ``` A continuous household survey with a rotating design measures the same units more than once. That is what makes the **change** between two periods far more precise than the levels themselves -- the shared sample cancels part of the sampling error -- and it is also what makes the change **harder** to estimate. The two samples are not independent, so the variance of the change is not the sum of the two variances: $$V\!\left(\hat\theta_t - \hat\theta_{t-1}\right) = V\!\left(\hat\theta_t\right) + V\!\left(\hat\theta_{t-1}\right) - 2\,\mathrm{Cov}\!\left(\hat\theta_t, \hat\theta_{t-1}\right)$$ and that covariance term is the whole problem: it is not a design constant you can look up, it has to be *produced* by the way the replicates are drawn. This vignette covers the three things a rotating panel makes possible, in the order an office needs them: 1. the **net change** between two periods, with its honest variance; 2. **chaining** period after period the way production actually runs, one month at a time, with only a small object traveling between runs; 3. the **gross flows** -- who moved between states -- which need a longitudinal weight. The package's panel layer follows the methodology of Statistics Canada's Labour Force Survey (cat. 71-526-X, sec. 7.2.2) for the coordinated replication, and ECLAC's household-survey manual (chapters XVI-XVII) for the longitudinal weight and the flows. ## The structure comes first Before any weighting, `panel_design()` reads the unit x wave crossing and describes what is actually there. It computes nothing about weights. ```{r design} pd <- panel_design(panel_ine, unit = c("household_id", "person_no"), wave = "wave", rotation_group = "rotation_group", pattern = "6") pd ``` Two things to read in that output. The **overlap matrix** is observed, from the data. The **profile implied by the pattern** is theoretical, derived from the rotation calendar. The declared `pattern` is therefore *verification*, not configuration: when the observed overlap falls below what the design implies, the linkage key is suspect, and `PN-01` says so. Ordinary attrition pulls the observed overlap down a little; a broken key pulls it down a lot. The pattern is parsed into the whole profile, not just the adjacent lag, because the informative lag is not always lag 1: ```{r profiles} prof <- function(p) round(weightflow:::.wf_pattern_overlap(p)$profile[1:5], 3) rbind(`6 (Canada LFS, ECH Uruguay)` = prof("6"), `4(0)1 (ECLAC ch. XVI example)` = prof("4(0)1"), `2-(2)-2 (Chile ENE, Italy)` = prof("2-(2)-2"), `4-8-4 (US CPS)` = prof("4-8-4"), `1(2)5 (PNAD Continua)` = prof("1(2)5")) ``` `2-(2)-2` shares **no** sample at lag 2 and half of it at lag 4 -- as much as at lag 1 -- so a year-on-year change there needs as much coordination as a quarter-on-quarter one. The PNAD Continua design `1(2)5` shares nothing with the adjacent quarter at all: a design built around "the previous period" would be exactly backwards for it. Both notations are accepted and describe the same calendar: `"2-(2)-2"` and `"2(2)2"` are the same design. ## Net change between two periods The coordinated bootstrap resamples **PSUs**, and a PSU present in both waves is resampled the same way in both. That is what lets the overlap show up as covariance. ```{r change} w1 <- subset(panel_ine, wave == 1 & disposition == "R") w2 <- subset(panel_ine, wave == 2 & disposition == "R") rec <- function(d) weighting_spec(d, base_weights = pw) |> step_nonresponse(respondent = disposition == "R", by = "sex") wb <- wave_bootstrap(list(T1 = rec(w1), T2 = rec(w2)), replicates = 100, strata = "stratum", psu = "psu", seed = 1, progress = FALSE) change_mean(wb, "unemployed") ``` `rho` is the correlation $\rho$ the overlap induces, and `deff_change` is $$\mathrm{deff}_\Delta = \frac{V\!\left(\hat\theta_t - \hat\theta_{t-1}\right)}{V\!\left(\hat\theta_t\right) + V\!\left(\hat\theta_{t-1}\right)}$$ the ratio between the variance reported here and what an office would report if it treated the two periods as independent samples. Ignoring the overlap does not give a conservative answer -- it gives a wrong one, in either direction depending on the sign of the covariance. For a combination over more than two waves -- a rolling quarter, an annual average -- `panel_estimate()` takes an arbitrary contrast: ```{r contrast, eval = FALSE} panel_estimate(wb, mean_of("unemployed"), contrast = c(-1, 1)) # the net change panel_estimate(wb, mean_of("unemployed")) # the average ``` ## Chaining: how production actually runs `wave_bootstrap()` wants every wave at once. An office does not have them: it publishes month `t` weeks after month `t-1`, and the twelfth month of the year cannot wait for the first eleven to be reprocessed. `wave_step()` runs **one period at a time** and hands the next period a small object -- the *carry* -- that is all it needs. ```{r chain, eval = FALSE} # period 1: nothing to coordinate with yet s1 <- wave_step(rec(w1), estimands = EST, replicates = 500, strata = "stratum", psu = "psu", period = "2026-01", seed = 1) saveRDS(wave_carry(s1), "carry/2026-01.rds") # period 2, weeks later, in a fresh session prev <- readRDS("carry/2026-01.rds") s2 <- wave_step(rec(w2), previous = prev, estimands = EST, replicates = 500, strata = "stratum", psu = "psu", period = "2026-02", seed = 2) s2$weights # the cross-sectional weights the office publishes s2$change # the net change against 2026-01, with rho and deff_change s2$strata # the coordination diagnostic, stratum by stratum saveRDS(wave_carry(s2), "carry/2026-02.rds") ``` Three properties worth stating plainly. **The cross-sectional weights are untouched.** `s2$weights` is identical to `prep(spec)$final_weight`. Coordination adds the change and the carry; it never moves the point estimate the office publishes. **`previous` is a list, not a file.** Which earlier periods share sample with this one is decided by the rotation calendar, not by proximity: with `2-(2)-2` the useful lags are 1, 3, 4 and 5, and lag 2 is empty. Each PSU inherits its multiplicity from the most recent carry that contains it, so gaps and returning cohorts resolve themselves. **The coordination is exact when the stratum keeps its size.** Coordination transfers the PSU *multiplicities*, and when `n_h` is unchanged the transfer is a permutation -- case (ii) of the LFS methodology -- so no replicate needs adjusting. `$strata` reports the case and the share of replicates that closed without adjustment, per stratum. That column is the quality indicator that decides whether a figure is publishable. ### Combinations over the chain `wave_step()` reports pairwise changes. A rolling quarter is not pairwise, and `wave_contrast()` estimates any linear combination straight from the saved carries: ```{r wavecontrast, eval = FALSE} tr <- lapply(c("2026-01", "2026-02", "2026-03"), \(m) readRDS(sprintf("carry/%s.rds", m))) wave_contrast(tr, "unemployment_rate") # rolling quarter wave_contrast(tr, "unemployment_rate", contrast = c(-1, 0, 1)) # T3 - T1 ``` This works without the waves being in memory because each carry stores the `R` replicate values of every declared estimand, and those replicates are **paired across periods** -- the coordination transferred the multiplicities. Stacking them recovers the full covariance matrix, and any contrast follows from it. ### The composite estimator When the recipe ends in `step_cre()`, the calibration of period $t$ targets two blocks at once: the known demographic totals $\mathbf{X}$, and composite totals $\widehat{\mathbf{Z}}$ **estimated with the previous wave**. Treating the second block as if it were known makes the variance anticonservative, so replicate $b$ of period $t$ rebuilds $\widehat{\mathbf{Z}}$ from replicate $b$ of period $t-1$. That is why a chain with `step_cre()` needs the "fat" carry: it must bring the previous period's replicate weights, not just its replicate estimates. `$n_cre_injected` and `$n_cre_skipped` audit that the injection happened. The estimator, its two imputations for the birth rotation group and the tuning constant $\alpha$ are in `vignette("composite-estimation")`. ## Gross flows: who moved A net change is a difference of aggregates. It cannot distinguish an immobile population from one where equal numbers enter and leave employment. For that you need the **longitudinal** file and a longitudinal weight. ```{r flows} wide <- panel_merge(list(T1 = subset(panel_ine, wave == 1), T2 = subset(panel_ine, wave == 2)), by = c("household_id", "person_no"), require = "all") lw <- weighting_spec(wide, base_weights = pw_T1) |> step_drop_ineligible(disposition_T2 == "OS", reason = "left the target population") |> step_attrition(respondent = disposition_T2 == "R", method = "propensity", formula = ~ age_T1 + sex_T1) |> prep() transition_matrix(lw, from = "lf_status_T1", to = "lf_status_T2", format = "row") ``` `boot_transition()` adds a standard error per cell, and `boot_flows()` gives the same information as population totals plus the net flows `i->j` minus `j->i` and the margins -- how many started in each state, ended in each, stayed, and moved. A caveat about the illustration, not the method: the four shipped panel datasets are synthetic, and in all of them **inactivity is an absorbing state** -- nobody enters or leaves it. So the matrix above shows movement only between employment and unemployment, and the `inact` row and column are degenerate. The mechanics are the point here; on real microdata the transitions in and out of inactivity are usually the interesting ones. Note the order of the two steps, which is not cosmetic. **Leaving the target population is not nonresponse.** Someone who died or emigrated is removed from the universe with `step_drop_ineligible()`; their weight is not redistributed to anyone. Someone who is still in the universe but did not answer is attrition, and their weight *is* redistributed, among the units that remain. Collapsing the two inflates the population. Three further conventions govern a longitudinal weight -- who belongs to the longitudinal population, which period's totals to calibrate to, and why cross-sectional estimates from a longitudinal file are reference only. They are decisions the analyst makes and the package cannot check, and they are set out in `vignette("panel-longitudinal")` and in `?step_cross_sectional`. ## Where to look next This vignette is the entry point; three companions go deeper into the pieces it uses. * `vignette("coordinated-replication")` -- what actually travels between waves, the four coordination cases, and how to read the `$strata` diagnostic that decides whether a change is publishable. * `vignette("composite-estimation")` -- `step_cre()` in full: what the composite estimator buys, and why its control totals make its variance a special case. * `vignette("panel-longitudinal")` -- the **pure** panel: attrition over many waves, the longitudinal weight, and gross flows. For the reference pages: `?wave_step` and `?wave_carry` for the chaining engine, `?wave_contrast` for combinations over a chain, `?panel_design` for the structure layer, `?transition_matrix` for the flows, and `?report_panel` for a quality report of a panel run. `vignette("variance-estimation")` covers the single-period bootstrap this one builds on, and `vignette("validation")` checks the change variance against an analytic estimator from a different family.