← Gallery
Diagram
medallion
← Prev
Next →
Medallion
Survey dataset storage tiers
Survey dataset storage tiers
One household survey dataset held at four ascending quality tiers on MinIO. Raw lands in raw-bucket, written by Apache NiFi as CSV, Parquet and JSON by a data engineer, and still carries respondent names and national identifiers. A PII REMOVE promotion lifts it to De-identified in anon-bucket, written by a Trino INSERT into partitioned Iceberg. A CLEAN and WEIGHT promotion lifts that to Curated in staging-bucket, typed and deduplicated with design weights applied by a data scientist in Trino and JupyterHub. An accent AGGREGATE promotion lands in the focal Aggregated tier in aggregated-bucket, holding Iceberg indicator tables that reporting and statistics consumers query. Promotions are adjacent-tier only and never flow backwards.
PII REMOVE
CLEAN + WEIGHT
AGGREGATE
Raw
raw-bucket
Tool
Apache NiFi raw write
Format
CSV · Parquet · JSON
Writer
Data engineer
Example payload
survey_2026q1.csv — 412 cols,
respondent name and national ID
De-identified
anon-bucket
Tool
Trino INSERT
Format
Iceberg · partitioned
Writer
Data engineer
Example payload
hashed respondent key; name and
ID columns dropped at write time
Curated
staging-bucket
Tool
Trino · JupyterHub
Format
Iceberg · cleaned
Writer
Data scientist
Example payload
typed and deduplicated rows with
design weights applied per stratum
Aggregated
aggregated-bucket
Tool
Trino INSERT · SAS JDBC
Format
Iceberg · indicators
Writer
Data scientist
Example payload
employment_rate by region and
quarter, with confidence bounds