Convert a parquet fplida run directory to CSV via DuckDB
Source:R/convert_to_csv.R
convert_parquet_dir_to_csv.RdWalks src_dir recursively. Every .parquet file is rewritten as a
sibling .csv under dst_dir, preserving the agency-dataset folder
structure, except datasets listed in preserve_parquet_datasets.
Preserved datasets are linked/copied as parquet. When
messy_files = TRUE, each PLIDA data product is written to its own
top-level <product>/ folder, agency spine files are written once under
top-level <agency>-spine-v6/ folders, and agency-dataset folders are
omitted. BLADE products keep the same product-folder structure under
abs-blade/, and STP parquet products keep it under
ato-stp/stp-standard/ or ato-stp/stp-extended/.
Partitioned product directories (containing part-NNN.parquet files) are
rewritten as a single unified CSV at <product>/<product>.csv, or under
the dataset exception folder for BLADE.
Usage
convert_parquet_dir_to_csv(
src_dir,
dst_dir = NULL,
threads = NULL,
memory_limit = "4GB",
log_file = NULL,
verbose = interactive(),
preserve_parquet_datasets = character(),
messy_files = TRUE,
messy_names = TRUE
)Arguments
- src_dir
Character. Parquet run directory (e.g. fplida_30m/).
- dst_dir
Character. Output root directory. Created if missing. If NULL, uses
paste0(src_dir, "_csv").- threads
Integer.
PRAGMA threads— DuckDB parallelism. Default isNULL, which usesparallel::detectCores()capped at 10. On a host where core detection fails, this falls back to 1.- memory_limit
Character.
PRAGMA memory_limit. Default "4GB".- log_file
Character or NULL. Path to a plain-text log. NULL skips logging (default).
- verbose
Logical. If TRUE, print per-file progress to the console (default TRUE when interactive, FALSE otherwise).
- preserve_parquet_datasets
Character vector of top-level dataset directory names to keep as parquet under
dst_dir, e.g."ato-stp".- messy_files
Logical. If TRUE, emit each PLIDA data product as a top-level
<product>/<product>.csvfolder, preserve STP parquet products underato-stp/stp-standard/<product>/orato-stp/stp-extended/<product>/, emit BLADE products underabs-blade/<product>/, and emit agency spines as top-level<agency>-spine-v6/<agency>-spine-v6.csvfolders instead of keeping other agency-dataset folders.- messy_names
Logical. If TRUE, vary a small subset of variable identifiers across related PIT/PAYG, MBS, and PBS year products by using case changes and close aliases.
Value
Invisibly, a list with fields:
- total_files
Number of parquet inputs converted
- total_rows
Total row count across all files
- total_bytes
Total CSV output bytes
- total_elapsed_sec
Wall-clock time in seconds
- per_file
data.frame of per-file timings
Details
DuckDB's COPY (SELECT * FROM read_parquet(...)) TO 'x.csv' (FORMAT CSV, HEADER) pipeline was benchmarked on 1 GB 30m fplida census data at
roughly 6x arrow::write_csv_arrow and 85x arrow::read_parquet +
data.table::fwrite. It streams end-to-end in C++ without any R-side
data.frame materialisation. Row totals are read from Parquet row-group
metadata, avoiding a second full data scan after each CSV write.