Creates a cross-sectional person spine with demographics, education, occupation (real 6-digit ANZSCO codes with 6D task scores), and income parameters. All datasets (Census, ITR, STP, etc.) are later projected from this spine.
Usage
generate_spine(
n = 200000L,
seed = 42L,
output_dir = NULL,
format = c("parquet", "csv"),
suffix = NULL,
return_data = TRUE,
use_template = TRUE,
template_n = .SPINE_TEMPLATE_DEFAULT_N
)Arguments
- n
Integer. Number of persons to generate.
- seed
Integer. Random seed for reproducibility.
- output_dir
Character or NULL. Base output directory. If NULL, uses
get_data_path(). One of the two must be set.- format
Character. Output format: "parquet" (default) or "csv".
- suffix
Character or NULL. Optional suffix appended to the run directory name to avoid overwriting (e.g. "seed3" creates
fplida_5m_seed3/). By default, runs with the same N (rounded) overwrite each other.- return_data
Logical. If TRUE (default), the function returns an R data.frame of the spine for the caller's use. At very large N this conversion can take many minutes; orchestration paths that discard the return value should pass
return_data = FALSE.- use_template
Logical. If TRUE (default), use the spine template cache: at the first call for a given seed, a ~100k template is built once; subsequent calls sample from it with replacement. This is dramatically faster at large N. Set to FALSE to build each person directly without template sampling.
- template_n
Integer. Template size when building (default 100000).