Data EngineeringActiveOpen source
PECOS Data Extraction Pipeline
A production-ready data pipeline demonstrating AWS Step Functions orchestrating four parallel PySpark ETL jobs (Clinicians, Practices, Canonical Providers, Enrollment Detail Tables) that process CMS PECOS healthcare provider data. Each job implements the Template Method pattern via a BaseETL class, applies domain-specific transformations, and outputs Snappy-compressed Parquet partitioned by state. The pipeline includes retry with exponential backoff, SNS notifications on success/failure, a local Python simulator mirroring the ASL state machine, and a bootstrap script that provisions S3, IAM, Glue, Step Functions, and SNS in one command.
Highlights
- ▸Step Functions state machine with parallel execution, retry, and choice routing
- ▸Four PySpark ETL jobs using Template Method pattern with healthcare domain logic
- ▸One-command bootstrap/teardown for complete AWS deployment
Tech Stack
AWS Step FunctionsPySparkAWS GlueS3SNSPythonDockerpytest