Build a complete fplida dataset using parallel slice workers
Source:R/build_fplida.R
build_fplida.RdSplits the spine into k_slices disjoint person ranges and
runs the full product pipeline in parallel across fresh R worker
processes. For any non-trivial N this is substantially faster than
a sequential build because it breaks the per-generator chunk-level
serialisation and reduces per-process allocator pressure.
Usage
build_fplida(
n = 1000000L,
seed = 42L,
years = 2015:2025,
k_slices = NULL,
rayon_threads = NULL,
products = "all",
exclude_products = NULL,
export_format = c("parquet", "csv"),
output_dir = NULL,
suffix = NULL,
slice_parent_dir = NULL,
keep_slice_dirs = FALSE,
keep_parquet = TRUE,
export_base_file = FALSE,
complete_dil_schema = identical(products, "all") && is.null(exclude_products),
complete_dil_rows = 100L,
messy_files = TRUE,
messy_names = TRUE,
years_by_product = NULL,
n_workers = NULL,
stp_zstd_level = NULL
)Arguments
- stp_zstd_level
Integer or NULL. Optional ZSTD compression level from 1 to 22 for STP Parquet files. Each slice compresses its STP files after generation and verifies their schemas and values before replacement. NULL preserves the native compression. Other products are unchanged.
- n
Integer. Total number of persons (default 1,000,000).
- seed
Integer. Base random seed.
- years
Integer vector. Panel years for time-varying datasets. Each product narrows this to the reference period its dataset publishes, so a build never writes a year PLIDA does not have. A product whose dataset covers none of these years is left out of the build and reported. See
plida_dataset_years().- k_slices
Integer. Number of parallel slice workers. Defaults to
detectCores()atn < 15Mandcores/2at larger N (memory headroom per worker).- rayon_threads
Integer or NULL. Rayon threads per worker. NULL auto-computes from
cores / k_slices.- products
Character. Either
"all"or a character vector of product names.- exclude_products
Character or NULL. Products to exclude when
products = "all". Cannot include"spine".- export_format
"parquet"(default) or"csv".- output_dir
Character or NULL. Base output directory.
- suffix
Character or NULL. Optional canonical run_dir suffix.
- slice_parent_dir
Character or NULL. Parent directory for the ephemeral slice run_dirs. NULL uses
tempdir().- keep_slice_dirs
Logical. If TRUE, slice run_dirs are not removed after merge (useful for debugging). Default FALSE.
- keep_parquet
Logical. When
export_format = "csv", should the source parquet files be kept alongside the converted CSVs? Defaults toTRUE. IfFALSE, the parquet run directory is deleted after a successful CSV conversion. Ignored whenexport_format = "parquet".- export_base_file
Logical. If TRUE, include the internal base-spine file in the final output. A Parquet build retains
_system/base-spine.parquet. A CSV build also exportsbase-spine-v6/base-spine-v6.csv. The build always creates the Parquet file for internal processing. The default is FALSE.- complete_dil_schema
Logical. If TRUE, write one canonical table for every product-table structure in the bundled PLIDA Data Item List for the selected build targets. The default is TRUE for
products = "all"and FALSE for a partial build. Survey structures are included, but their unvalidated values remain typed missing where a bespoke generator does not supply them.- complete_dil_rows
Positive integer. Maximum rows in each canonical DIL schema companion. The default is 100. Richer bespoke product outputs retain their normal row counts.
- messy_files
Logical. When
export_format = "csv", write each PLIDA data product as a top-level folder, keep BLADE product folders underabs-blade/, keep STP parquet product folders underato-stp/stp-standard/orato-stp/stp-extended/, omit other agency grouping folders, and write agency spines as top-level<agency>-spine-v6/folders. Defaults to TRUE.- messy_names
Logical. When
export_format = "csv", vary a small subset of variables across related PIT/PAYG, MBS, and PBS year products. Defaults to TRUE.- years_by_product
Named list of integer year vectors. Overrides
yearsfor selected year-aware products, for examplelist(mbs = 2024L, pbs = 2024L). Published coverage still applies. Requirescomplete_dil_schema = FALSE: schema companions cover the full bundled registry rather than a restricted observation window.- n_workers
Integer or NULL. Maximum concurrent slice workers. NULL uses one worker per slice. Use fewer workers than slices to bound memory.
Details
The canonical output run directory has one directory per product:
per-product outputs are stored as part files
(part-NNN.parquet), naturally readable via
arrow::open_dataset().
Cross-person- or household-dependent generators (spine,
core, blade, lfs) run once centrally on the full
population before workers start. All other products run per slice.
CSV output: when export_format = "csv", the build runs to
parquet internally (all Rust fast paths are parquet-only), then the
canonical run directory is converted to CSV via DuckDB. The CSV run
directory is written to <run_dir>_csv/. STP is preserved as
parquet rather than converted to CSV. With messy_files = TRUE,
preserved STP product directories are written under
ato-stp/stp-standard/ or ato-stp/stp-extended/ in the CSV
run directory. With keep_parquet = FALSE the source parquet run
directory is deleted after conversion.