Skip to contents

Splits the spine into k_slices disjoint person ranges and runs the full product pipeline in parallel across fresh R worker processes. For any non-trivial N this is substantially faster than a sequential build because it breaks the per-generator chunk-level serialisation and reduces per-process allocator pressure.

Usage

build_fplida(
  n = 1000000L,
  seed = 42L,
  years = 2015:2025,
  k_slices = NULL,
  rayon_threads = NULL,
  products = "all",
  exclude_products = NULL,
  export_format = c("parquet", "csv"),
  output_dir = NULL,
  suffix = NULL,
  slice_parent_dir = NULL,
  keep_slice_dirs = FALSE,
  keep_parquet = TRUE,
  export_base_file = FALSE,
  complete_dil_schema = identical(products, "all") && is.null(exclude_products),
  complete_dil_rows = 100L,
  messy_files = TRUE,
  messy_names = TRUE,
  years_by_product = NULL,
  n_workers = NULL,
  stp_zstd_level = NULL
)

Arguments

stp_zstd_level

Integer or NULL. Optional ZSTD compression level from 1 to 22 for STP Parquet files. Each slice compresses its STP files after generation and verifies their schemas and values before replacement. NULL preserves the native compression. Other products are unchanged.

n

Integer. Total number of persons (default 1,000,000).

seed

Integer. Base random seed.

years

Integer vector. Panel years for time-varying datasets. Each product narrows this to the reference period its dataset publishes, so a build never writes a year PLIDA does not have. A product whose dataset covers none of these years is left out of the build and reported. See plida_dataset_years().

k_slices

Integer. Number of parallel slice workers. Defaults to detectCores() at n < 15M and cores/2 at larger N (memory headroom per worker).

rayon_threads

Integer or NULL. Rayon threads per worker. NULL auto-computes from cores / k_slices.

products

Character. Either "all" or a character vector of product names.

exclude_products

Character or NULL. Products to exclude when products = "all". Cannot include "spine".

export_format

"parquet" (default) or "csv".

output_dir

Character or NULL. Base output directory.

suffix

Character or NULL. Optional canonical run_dir suffix.

slice_parent_dir

Character or NULL. Parent directory for the ephemeral slice run_dirs. NULL uses tempdir().

keep_slice_dirs

Logical. If TRUE, slice run_dirs are not removed after merge (useful for debugging). Default FALSE.

keep_parquet

Logical. When export_format = "csv", should the source parquet files be kept alongside the converted CSVs? Defaults to TRUE. If FALSE, the parquet run directory is deleted after a successful CSV conversion. Ignored when export_format = "parquet".

export_base_file

Logical. If TRUE, include the internal base-spine file in the final output. A Parquet build retains _system/base-spine.parquet. A CSV build also exports base-spine-v6/base-spine-v6.csv. The build always creates the Parquet file for internal processing. The default is FALSE.

complete_dil_schema

Logical. If TRUE, write one canonical table for every product-table structure in the bundled PLIDA Data Item List for the selected build targets. The default is TRUE for products = "all" and FALSE for a partial build. Survey structures are included, but their unvalidated values remain typed missing where a bespoke generator does not supply them.

complete_dil_rows

Positive integer. Maximum rows in each canonical DIL schema companion. The default is 100. Richer bespoke product outputs retain their normal row counts.

messy_files

Logical. When export_format = "csv", write each PLIDA data product as a top-level folder, keep BLADE product folders under abs-blade/, keep STP parquet product folders under ato-stp/stp-standard/ or ato-stp/stp-extended/, omit other agency grouping folders, and write agency spines as top-level <agency>-spine-v6/ folders. Defaults to TRUE.

messy_names

Logical. When export_format = "csv", vary a small subset of variables across related PIT/PAYG, MBS, and PBS year products. Defaults to TRUE.

years_by_product

Named list of integer year vectors. Overrides years for selected year-aware products, for example list(mbs = 2024L, pbs = 2024L). Published coverage still applies. Requires complete_dil_schema = FALSE: schema companions cover the full bundled registry rather than a restricted observation window.

n_workers

Integer or NULL. Maximum concurrent slice workers. NULL uses one worker per slice. Use fewer workers than slices to bound memory.

Value

Invisibly, a list with build metadata and per-slice stats.

Details

The canonical output run directory has one directory per product: per-product outputs are stored as part files (part-NNN.parquet), naturally readable via arrow::open_dataset().

Cross-person- or household-dependent generators (spine, core, blade, lfs) run once centrally on the full population before workers start. All other products run per slice.

CSV output: when export_format = "csv", the build runs to parquet internally (all Rust fast paths are parquet-only), then the canonical run directory is converted to CSV via DuckDB. The CSV run directory is written to <run_dir>_csv/. STP is preserved as parquet rather than converted to CSV. With messy_files = TRUE, preserved STP product directories are written under ato-stp/stp-standard/ or ato-stp/stp-extended/ in the CSV run directory. With keep_parquet = FALSE the source parquet run directory is deleted after conversion.

Examples

if (FALSE) { # \dontrun{
res <- build_fplida(n = 1e6)
res <- build_fplida(n = 25e6, k_slices = 5, rayon_threads = 2)
res <- build_fplida(n = 1e6, export_format = "csv")
} # }