Skip to contents

Walks src_dir recursively. Every .parquet file is rewritten as a sibling .csv under dst_dir, preserving the agency-dataset folder structure, except datasets listed in preserve_parquet_datasets. Preserved datasets are linked/copied as parquet. When messy_files = TRUE, each PLIDA data product is written to its own top-level <product>/ folder, agency spine files are written once under top-level <agency>-spine-v6/ folders, and agency-dataset folders are omitted. BLADE products keep the same product-folder structure under abs-blade/, and STP parquet products keep it under ato-stp/stp-standard/ or ato-stp/stp-extended/. Partitioned product directories (containing part-NNN.parquet files) are rewritten as a single unified CSV at <product>/<product>.csv, or under the dataset exception folder for BLADE.

Usage

convert_parquet_dir_to_csv(
  src_dir,
  dst_dir = NULL,
  threads = NULL,
  memory_limit = "4GB",
  log_file = NULL,
  verbose = interactive(),
  preserve_parquet_datasets = character(),
  messy_files = TRUE,
  messy_names = TRUE
)

Arguments

src_dir

Character. Parquet run directory (e.g. fplida_30m/).

dst_dir

Character. Output root directory. Created if missing. If NULL, uses paste0(src_dir, "_csv").

threads

Integer. PRAGMA threads — DuckDB parallelism. Default is NULL, which uses parallel::detectCores() capped at 10. On a host where core detection fails, this falls back to 1.

memory_limit

Character. PRAGMA memory_limit. Default "4GB".

log_file

Character or NULL. Path to a plain-text log. NULL skips logging (default).

verbose

Logical. If TRUE, print per-file progress to the console (default TRUE when interactive, FALSE otherwise).

preserve_parquet_datasets

Character vector of top-level dataset directory names to keep as parquet under dst_dir, e.g. "ato-stp".

messy_files

Logical. If TRUE, emit each PLIDA data product as a top-level <product>/<product>.csv folder, preserve STP parquet products under ato-stp/stp-standard/<product>/ or ato-stp/stp-extended/<product>/, emit BLADE products under abs-blade/<product>/, and emit agency spines as top-level <agency>-spine-v6/<agency>-spine-v6.csv folders instead of keeping other agency-dataset folders.

messy_names

Logical. If TRUE, vary a small subset of variable identifiers across related PIT/PAYG, MBS, and PBS year products by using case changes and close aliases.

Value

Invisibly, a list with fields:

total_files

Number of parquet inputs converted

total_rows

Total row count across all files

total_bytes

Total CSV output bytes

total_elapsed_sec

Wall-clock time in seconds

per_file

data.frame of per-file timings

Details

DuckDB's COPY (SELECT * FROM read_parquet(...)) TO 'x.csv' (FORMAT CSV, HEADER) pipeline was benchmarked on 1 GB 30m fplida census data at roughly 6x arrow::write_csv_arrow and 85x arrow::read_parquet + data.table::fwrite. It streams end-to-end in C++ without any R-side data.frame materialisation. Row totals are read from Parquet row-group metadata, avoiding a second full data scan after each CSV write.