Skip to contents

generate_pit_ps() writes synthetic Payment Summary (PIT_PS) product tables. The covered population contains people with an employer payment summary. Each record represents one person, one payer, and one financial year.

Usage

generate_pit_ps(
  spine = NULL,
  seed = 42L,
  years = 2002:2023,
  output_dir = NULL,
  format = c("parquet", "csv"),
  return_data = FALSE
)

Arguments

spine

A data frame from generate_spine(), or NULL. If the value is NULL, the function reads the spine from the run directory.

seed

An integer random seed.

years

A vector of financial-year end years. The delivery covers 2002 to 2023; any other year is dropped and no file is written for it.

output_dir

The base output directory, or NULL.

format

The output format. This function supports "parquet" only.

return_data

This argument has no effect. The function always writes the data to disk.

Value

An invisible metadata list with n_records, n_rows, years, and path.

Details

The Rust pipeline uses the shared employment panel. It also writes the occupation-panel files that generate_pit_itr() uses.

What each product contains

A product holds one to three tables. Each is written as its own file, named <product>--<table>, which is the convention the package uses for every product the data item list splits into tables. The schema changes by year, from four variables in 2001-02 to thirty-six in 2022-23, and every table carries exactly the variables the registry declares for it. From 2019-20 several tables report one year and differ only in how long after 30 June the extract was cut; the six-month cut misses the payment summaries lodged after it. The four products from 2015-16 to 2018-19 also carry a separate geography table, one row per person with a payment summary that year.

Which identifier names the employer

The tables delivered up to 2021-22 name the employer in ABN_HASH_TRUNC and the tables from 2021-22 on name it in BN. Which one a table carries comes from the data item list, not from the financial year: 2021-22 is mixed, with ABN_HASH_TRUNC on the six-month extract and BN on the sixteen-month re-extract. The two are different hashings of the same ABN and never share a value, so joining across the change needs the blade-key-abn-hash-trunc-to-bn-key product in abs-blade.

Dataset and variable information

The ABS administrative income sources website gives information about this dataset. Use dataset_info("PIT_PS") for dataset information. Use variable_info("PIT_PS") for variables, sources, value support, and topic tags.

Occupation codes are not ANZSCO

The occupation panel this function writes for generate_pit_itr() holds ATO salary and wage occupation codes rather than ANZSCO. See ?generate_pit_itr for the mapping, and the ATO's published list.