Pavan Kalyan Kambala
Pavan Kalyan Kambala

Data Engineer

Open to offers · Member since 17 Sep 2026
Location
Atlanta, United States
Desired salary
Unspecified
Work preference
Remote Only / Full Time
Experience level
Mid

About

Professional summary

I am a Data Engineer with over three years of experience building cloud-native ETL and ELT pipelines, data lakes, and analytics platforms for healthcare and financial-services organizations.

I specialize in Python, SQL, PySpark, Apache Spark, Apache Airflow, Snowflake, and AWS. I design scalable ingestion and transformation workflows that process high-volume data reliably while meeting operational SLAs.

I have experience architecting multi-layer AWS data lake environments using S3, Glue, Redshift, Lambda, IAM, and related services. My work includes dimensional modeling, Data Vault modeling, incremental loading, CDC patterns, and data warehouse development.

I focus on improving pipeline performance, data quality, and deployment reliability. I have implemented automated validation frameworks with Great Expectations, CI/CD workflows with Jenkins and GitHub Actions, and governance controls for sensitive healthcare data.

I have delivered data solutions that improved pipeline runtimes, query performance, reporting accuracy, and release processes. I am AWS, Snowflake, Databricks, and Microsoft Power BI certified, and I am interested in data engineering opportunities involving modern cloud data platforms.

Skills

19 capabilities

Tech stack & tools

Working toolkit

Application Hosting

Collaboration

Data Stores

Development

Languages & Frameworks

Experience

Career history

Data Engineer CVS Health

Architected a multi-layer AWS data lake using S3, Glue, Redshift, and Snowflake to integrate EHR, claims, and flat-file sources for enterprise healthcare analytics and machine learning use cases. Supported more than 20 Apache Airflow-orchestrated workflows processing over 300 GB daily at a 99% SLA.

Built and optimized Python-based ETL pipelines on Apache Airflow, reducing end-to-end runtime by 30%. Redesigned data models and tuned Snowflake and Redshift SQL through indexing, clustering keys, and partition pruning, improving query and Tableau dashboard performance by 35%.

Developed an automated Great Expectations data-quality framework covering null checks, schema enforcement, and referential integrity, reducing reporting defects by 25%. Automated CI/CD deployments with Jenkins and GitHub Actions, reducing manual deployment effort by 50%.

Implemented HIPAA-compliant governance using IAM access controls, field-level encryption, audit logging, and AWS Glue Data Catalog lineage and metadata management for PHI datasets, reducing unauthorized-access incidents by 30%.

Fin Data Engineer Cyient

Designed a PySpark ingestion framework on AWS using S3, Glue, and Lambda to process more than 500 GB of nightly financial transaction data. Applied dynamic partitioning and optimized shuffle configuration for efficient large-scale processing.

Implemented a watermark-based incremental loading strategy with timestamp tracking and change detection, reducing batch processing time by 40%. Built modular and parameterized Apache Airflow DAGs to orchestrate more than 15 production ingestion workflows with SLA monitoring.

Developed Star Schema dimensional models and curated data marts in Snowflake and PostgreSQL to support Tableau and Power BI dashboards. Implemented dbt- and PySpark-based data-quality checks for regulatory KPI accuracy.

Containerized ETL workloads with Docker, reducing environment-related defects by 30%, and mentored three junior engineers on ETL and Airflow best practices.

Education

Learning history

University of Central Missouri

Master of Science, Computer Science

VVIT

Bachelor of Technology, Computer Science & Engineering

This professional hasn’t added portfolio projects yet.

This professional hasn’t listed any services yet.

People also viewed

All talent ›
Jobs Talent AI Tools Salaries
Menu