Medallion

Survey dataset storage tiers

Survey dataset storage tiers One household survey dataset held at four ascending quality tiers on MinIO. Raw lands in raw-bucket, written by Apache NiFi as CSV, Parquet and JSON by a data engineer, and still carries respondent names and national identifiers. A PII REMOVE promotion lifts it to De-identified in anon-bucket, written by a Trino INSERT into partitioned Iceberg. A CLEAN and WEIGHT promotion lifts that to Curated in staging-bucket, typed and deduplicated with design weights applied by a data scientist in Trino and JupyterHub. An accent AGGREGATE promotion lands in the focal Aggregated tier in aggregated-bucket, holding Iceberg indicator tables that reporting and statistics consumers query. Promotions are adjacent-tier only and never flow backwards. PII REMOVE CLEAN + WEIGHT AGGREGATE Raw raw-bucket Tool Apache NiFi raw write Format CSV · Parquet · JSON Writer Data engineer Example payload survey_2026q1.csv — 412 cols, respondent name and national ID De-identified anon-bucket Tool Trino INSERT Format Iceberg · partitioned Writer Data engineer Example payload hashed respondent key; name and ID columns dropped at write time Curated staging-bucket Tool Trino · JupyterHub Format Iceberg · cleaned Writer Data scientist Example payload typed and deduplicated rows with design weights applied per stratum Aggregated aggregated-bucket Tool Trino INSERT · SAS JDBC Format Iceberg · indicators Writer Data scientist Example payload employment_rate by region and quarter, with confidence bounds