| Prabhu Jay - Senior Data Engineer |
| [email protected] |
| Location: Dallas, Texas, USA |
| Relocation: Yes |
| Visa: H1B |
| Resume file: Prabhu_Senior Data Engineer_1783520226706.docx Please check the file(s) for viruses. Files are checked manually and then made available for download. |
|
Prabhu Jay
Senior Databricks Engineer +1 224-252-5162 [email protected] www.linkedin.com/in/prabhu-jay-25a937250 -------------------------------------------------------------------------------------------------------------------------------------------------- Professional Summary: 13 years Hands-on experience working with data across analytics, big data, and data engineering roles, supporting and healthcare projects. Strong background in building and maintaining large-scale data pipelines using Spark, Hadoop, and Kafka for both batch and real-time processing. Practical experience developing Spark applications using PySpark and Scala for data cleaning, transformations, joins, and aggregations on high-volume datasets. Comfortable working with distributed systems and parallel processing concepts to handle structured, semi-structured, and streaming data efficiently. Experience in using AWS services such as S3, EMR, Redshift, Glue, Athena, Lambda, SNS, and CloudWatch for data storage, processing, analytics, and monitoring. Solid hands-on experience with Azure services including Azure Data Factory, Azure Data Lake Storage Gen2, Azure Synapse Analytics, Azure SQL DB, and Event Hubs, primarily for data migration and analytics use cases. Participates in the development improvement and maintenance of snowflake database applications Strong experience building end-to-end data workflows, starting from data ingestion through transformation and ending with reporting and analytics. Regularly worked with Python for data processing and automation, along with SQL for querying, validation, and reporting across multiple databases. Experience working with relational and NoSQL databases such as PostgreSQL, MySQL, SQL Server, MongoDB, Cassandra, DynamoDB, and HBase. Good understanding of data warehousing concepts, including schema design, partitioning strategies, and performance considerations for analytical workloads. Experience supporting reporting and visualization needs using tools like Tableau and Power BI to deliver clear, actionable insights. Familiar with creating and managing cloud-based data lakes and organizing data for analytics, reporting, and downstream applications. Hands-on experience scheduling and managing data pipelines using tools like Apache Airflow and Azure Data Factory. In-depth knowledge of Snowflake Database, Schema and Table structures. Practical exposure to handling pipeline failures, reruns, alerts, and monitoring to keep production data systems reliable. Experience using Git for version control and working in team-based development environments. Comfortable working in Agile teams, with exposure to SDLC and Waterfall models depending on project requirements. Known for a problem-solving mindset, attention to data quality, and the ability to work independently while collaborating effectively with cross-functional teams. Technical Skills Cloud Platforms AWS (S3, EMR, Redshift, Glue, Athena, Lambda, SNS, CloudWatch, DynamoDB, Step Functions, IAM), Azure (Azure Data Factory, Azure Data Lake Gen2, Azure Synapse Analytics, Azure SQL DB) Programming Languages Python, SQL, Java, Scala, PySpark, Shell Scripting Big Data & Distributed Processing Apache Spark, Hadoop (HDFS, YARN), Hive, Kafka, Sqoop, NiFi, Oozie, Cloudera (CDH/HDP), Dataproc Data Engineering & Orchestration Apache Airflow, Cloud Composer, Azure Data Factory, AWS Step Functions Databases PostgreSQL, MySQL, Oracle, Microsoft SQL Server, DB2, Azure SQL DB NoSQL Databases Amazon Redshift, Snowflake, Azure Synapse Analytics, Teradata Data Warehousing & Analytics Looker, Tableau, Power BI, SSRS, Crystal Reports ETL & Data Integration Tools AWS Glue,Informatica, SSIS, Matillion, Alteryx, Talend Data Modeling &Governance ERwin, ER/Studio, Data Catalogs, Data Lineage, Data Quality Checks, Great Expectations, Deequ BI & Visualization Tools Tableau, Power BI, Looker, SSRS CI/CD & DevOps Jenkins, GitHub Actions, Git, Terraform, CloudFormation Security & Compliance IAM, Role-Based Access Control, Encryption (at rest & in transit), HIPAA Compliance, PII Handling Methodologies Agile (Scrum), SDLC, Waterfall Work Experience: Client: Humana June 2024 Present Role: Sr. Data/Databricks Engineer Designed and implemented scalable data processing solutions using Azure Databricks and Spark for large-scale analytics workloads. Built and optimized data pipelines using PySpark to process structured and semi-structured data efficiently. Collaborated with cross-functional teams to gather requirements and translate business needs into technical solutions. Built scalable end-to-end ETL pipelin es using PySpark and Python to ingest, transform, and load large-scale healthcare data into analytics-ready formats. Developed automated data validation and exception handling frameworks to detect, log, and resolve inconsistencies across pipeline stages. Designed MDM-style data integration workflows to synchronize master data (customer, provider, claims entities) across distributed systems using REST APIs, aligning with Reltio-like architecture patterns. Built API-driven data exchange frameworks enabling near real-time synchronization between upstream systems and centralized master datasets, similar to Reltio API integrations Built low-latency streaming pipelines using Kafka and Spark Structured Streaming for real-time data ingestion and processing. Optimized complex queries using indexing, execution plan analysis, and stored procedures, improving performance for large-scale transactional datasets. Integrated data pipelines with REST-based microservices and backend systems, enabling seamless communication across distributed application layers. Implemented event-driven pipelines using Kafka with producer-consumer architecture, ensuring reliable, asynchronous message processing comparable to enterprise MQ systems. Developed schema-driven ingestion frameworks for structured message data (XML/JSON/Avro), standardizing payloads into canonical models for downstream processing. Built high-volume transaction processing and reconciliation workflows with validation and matching logic similar to post-trade settlement systems. Designed data pipelines aligned with trade lifecycle stages (ingestion, validation, enrichment, and reconciliation) to ensure accurate end-to-end transaction processing. Designed and supported scalable ELT pipelines using Snowflake and Python to transform and load large-scale structured and semi-structured datasets Designed and optimized Amazon Aurora PostgreSQL databases for high-volume analytics workloads, implementing efficient schema design and query tuning to improve data retrieval performance. Built event-driven data pipelines using Amazon EventBridge to trigger AWS Glue jobs and Lambda functions based on S3 data events, enabling scalable and loosely coupled workflow orchestration. Developed end-to-end AWS-native data pipelines using Glue (PySpark), S3, Step Functions, and Athena, delivering scalable batch and near real-time data processing solutions. Developed complex SQL transformations in Snowflake using CTEs, window functions, and query optimization techniques to build high-performance analytics-ready datasets Built and orchestrated end-to-end data workflows integrating Apache Airflow and Azure Data Factory with Snowflake for automated, reliable pipeline execution Optimized Snowflake performance and cost efficiency through efficient schema design, query tuning, and effective compute utilization strategies Built automated data workflows using Power Automate integrated with Azure Data Factory and APIs, and developed lightweight Power Apps dashboards to enable business users to trigger pipelines and monitor data processing status in real time. Developed Python-based AI prototypes using machine learning libraries to automate data quality checks and anomaly detection, integrating models into data pipelines for intelligent validation and improved data reliability.Developed and optimized Kusto (KQL) queries in Azure Data Explorer to analyze large-scale telemetry and pipeline logs, enabling real-time monitoring, anomaly detection, and faster root cause analysis of data pipeline failures. Designed and implemented a Microsoft Fabric Lakehouse architecture leveraging OneLake, Dataflows Gen2, and Fabric Warehouses to ingest, transform, and serve analytics-ready datasets, improving data accessibility and reducing pipeline latency across business domains. Contributed to designing lakehouse-based data architecture using Azure Databricks, ADLS Gen2, and Synapse for structured and semi-structured data processing. Worked with data modeling teams to define and implement structured data models supporting reporting, analytics, and downstream data consumption. Developed and optimized Spark jobs in Azure Databricks for large-scale distributed processing of batch and streaming healthcare datasets. Developed reusable Python modules for data extraction, transformation, validation, and automation of ETL workflows. Wrote complex SQL queries for data transformation, validation, and performance optimization across relational and analytical databases. Designed and maintained Airflow DAGs to orchestrate and automate complex multi-stage data pipelines with dependency management. Built cloud-native data solutions using Azure Data Lake, Databricks, Synapse, and AWS services for scalable data processing and analytics. Implemented real-time data pipelines using Kafka and Spark Structured Streaming to process event-driven healthcare data with low latency. Designed and optimized analytical datasets in Synapse and SQL-based systems to support high-performance reporting and BI workloads. Applied modular transformation principles inspired by dbt methodology to structure layered data transformations (staging, intermediate, and marts). Implemented automated data validation checks and monitoring frameworks using Azure Monitor to ensure pipeline reliability and accuracy. Built unified data consumption layers by integrating multiple domain datasets into reusable structures for enterprise analytics. Designed API-based data integration workflows using REST and event-driven architecture patterns aligned with enterprise middleware concepts. Supported ingestion and processing of enterprise application data using standardized data models and API-based integration patterns. Implemented CI/CD pipelines using Azure DevOps and GitHub Actions to automate deployment of data pipelines and Spark applications. Deployed and managed containerized Spark and data processing applications on Kubernetes clusters to support scalable and highly available analytics workloads. Containerized PySpark, Python, and Scala-based data engineering applications using Docker to ensure consistent execution across development, testing, and production environments. Assisted in deploying data engineering workloads on Google Kubernetes Engine (GKE) for scalable execution of containerized analytics applications. Designed and implemented CI/CD pipelines using AWS CodePipeline to automate build, test, and deployment processes for Spark and Databricks workloads across development, QA, and production environments. Designed and implemented scalable end-to-end cloud data pipelines using Azure Data Factory, Azure Databricks, and Azure Synapse Analytics to build Lakehouse-style architectures aligned with Microsoft Fabric concepts, including centralized data storage, ELT processing, reusable transformation layers, and Power BI-ready analytical datasets for enterprise reporting and near real-time analytics. Owned end-to-end delivery of data projects including design, development, deployment, and ongoing support. Conducted code reviews and enforced best practices to improve code quality, performance, and maintainability. Designed data lake solutions using Azure Data Lake Storage (ADLS Gen2) to support scalable and cost-efficient data storage. Facilitated communication between technical and non-technical stakeholders to align on project goals and deliverables. Designed and implemented master data consolidation pipelines to create unified views of enterprise datasets, supporting golden record generation across multiple data sources. Applied data standardization, deduplication, and survivorship logic to ensure consistency and accuracy of critical business entities. Configured schema-driven data ingestion frameworks using JSON/Avro to standardize entity models, similar to Reltio entity configuration. Leveraged AI-assisted coding tools like GitHub Copilot to accelerate development of PySpark scripts, SQL queries, and data validation logic. Defined data models, relationships, and validation rules for master datasets within Snowflake and Databricks environments. Identified and mitigated risks related to project timelines, system performance, and data quality. Stayed current with emerging technologies in Azure Databricks and modern data engineering practices to drive innovation. Proficient in Python, SQL, and Scala for building scalable data pipelines and transformati Experience with Azure Data Lake Storage (ADLS Gen2) and Azure SQL Database for cloud-based data storage solutions. Strong understanding of ETL/ELT processes, data modeling, and data warehousing concepts. Experience building and optimizing data pipelines for performance, scalability, and reliability. Proven ability to lead technical teams and manage multiple projects in Agile environments. Strong problem-solving skills with experience in troubleshooting complex distributed data systems. Environment: Azure (Azure Data Lake Storage Gen2, Azure SQL Database, Azure Synapse Analytics, Azure Databricks, Azure Data Factory, Azure Functions, Azure Event Hubs, Azure Cosmos DB, Azure Monitor, Azure Logic Apps, Azure Resource Manager), Spark, PySpark, Python, Apache Kafka, Apache Airflow, Delta Lake, Terraform, Tableau, Git, Azure DevOps, GitHub Actions, CI/CD Pipelines Client: Blue cross Blue shield. Sep 2022 May 2024 Role: Sr. Databricks Engineer Responsibilities: Set up a centralized healthcare Data Lake on Azure to manage diverse clinical datasets including electronic health records (EHR), patient demographics, insurance claims, lab results, imaging metadata, and IoT patient monitoring data. Collected and ingested data from multiple sources such as Azure SQL Database, Azure Data Lake Storage Gen2, Azure Cosmos DB, on-premise clinical systems, and HL7/FHIR-based healthcare APIs via Azure API Management. Used Azure Data Factory to discover, clean, and structure raw healthcare data, creating ETL pipelines to transform data for analytics and reporting. Implemented data reconciliation and survivorship rules to prioritize and retain the most accurate records across multiple systems. Designed workflows to maintain consistent master data hierarchies across patient and provider domains. Integrated FHIR-based REST APIs to ingest and synchronize healthcare master data across systems. Developed API-based services to expose curated master datasets for downstream applications. Built real-time streaming pipelines using Azure Event Hubs / Apache Kafka on HDInsight and Azure Databricks (PySpark Structured Streaming) to process patient vital signs and clinical event streams. Configured Azure Functions for serverless processing of incoming healthcare events and triggered downstream workflows automatically. Utilized AI-powered development assistants to improve productivity in ETL development and pipeline debugging. Applied AI tools to generate reusable code templates and optimize data transformation logic. Designed and maintained logical and physical data models representing master entities and hierarchies, aligning with MDM configuration best practices. Implemented metadata-driven configurations for data ingestion and transformation pipelines. Implemented centralized master data pipelines for healthcare entities (patients, providers) using Azure Data Lake and Databricks, following MDM principles comparable to Reltio-based solutions. Integrated healthcare APIs (FHIR/HL7) to standardize and consolidate master data across multiple clinical systems, replicating Reltio-style data ingestion patterns. Used Azure Synapse Analytics (Serverless SQL Pool) to allow teams to query structured and semi-structured healthcare datasets directly from Azure Data Lake Storage without provisioning servers. Created data catalogs with Azure Purview to maintain metadata, track schema changes, and ensure discoverability of all datasets. Scheduled and managed workflows using Apache Airflow (on Azure Managed Airflow / AKS) and Azure Logic Apps, ensuring that ETL jobs and dependent processes ran reliably and in the correct sequence. Deployed Azure infrastructure using Terraform and ARM/Bicep templates, including Data Lake Storage accounts, Synapse dedicated pools, Databricks workspaces, Azure Functions, and Azure AD RBAC to ensure reproducibility and security. Optimized Azure Synapse Analytics dedicated SQL pools for high-performance analytical queries on claims and EHR datasets, including partitioning, distribution strategies, and workload management. Processed large-scale clinical data using Azure Databricks with Spark and PySpark, performing transformations, aggregations, and joining multiple healthcare datasets efficiently. Developed Java-based microservices and REST APIs hosted on Azure App Service / AKS to expose processed healthcare data to applications while enforcing HIPAA-compliant access controls. Set up monitoring and alerting using Azure Monitor, Log Analytics, and Azure Alerts to track pipeline health, data anomalies, and system performance. Built interactive dashboards and visualizations in Power BI to support clinical analytics, patient risk scoring, treatment outcome evaluation, and operational efficiency metrics. Migrated legacy healthcare databases to Azure Data Lake and Azure Synapse Analytics while validating data quality and ensuring no loss of critical clinical information. Implemented CI/CD pipelines with Azure DevOps and GitHub Actions to automate deployment of ETL workflows, Spark jobs, dashboards, and infrastructure changes. Maintained detailed documentation in Confluence and Git to support knowledge sharing and onboarding of new team members. Implemented data governance and lineage tracking using Azure Purview to maintain compliance with HIPAA and healthcare data standards, ensuring proper handling of sensitive patient information. Integrated Delta Lake on Azure Databricks to unify structured and semi-structured datasets, enabling advanced analytics, predictive modeling, and machine learning on patient care and operational data. Delivered end-to-end solutions from raw data ingestion to analytics dashboards, supporting data-driven clinical decisions, population health management, and hospital operations. Environment: Azure (Azure Data Lake Storage Gen2, Azure SQL Database, Azure Synapse Analytics, Azure Databricks, Azure Functions, Azure Data Factory, Azure Event Hubs, Azure Cosmos DB, Azure Monitor, Azure Logic Apps, Azure Purview, ARM/Bicep), Spark, PySpark, Python, Apache Kafka, Apache Airflow, Delta Lake, Terraform, Power BI, Git, Azure DevOps, GitHub Actions, CI/CD Pipelines Client: Accenture Feb 2018 Mar 2022 Role: Big Data Engineer Responsibilities: Worked on setting up and maintaining multi-cluster Hadoop environments using Cloudera and Hortonworks to support large-scale retail data processing, including deployments running on AWS EC2 infrastructure. Handled cluster upgrades, patching, and production issues for HDP and CDH clusters to keep the platform stable and available across on-prem and AWS environments. Ingested retail data from multiple sources such as POS systems, transaction feeds, invoices, and customer activity logs into the Hadoop ecosystem, including data landing in Amazon S3. Set up Kafka brokers and managed schemas to stream structured retail events reliably into the data lake, with Kafka running on AWS instances. Built Spark jobs using Spark Core, Spark SQL, and DataFrames to process structured sales and transaction data at scale on Hadoop and AWS-based clusters. Reworked existing Hive and SQL queries into Spark DataFrame logic to improve processing speed and handle growing data volumes. Implemented real-time processing using Spark Structured Streaming to handle live sales and customer events. Stored raw and processed data in HDFS and Amazon S3 and organized it for downstream processing and reporting. Used Spark SQL to prepare curated datasets and loaded the results into Hive tables for analytics and reporting use cases. Designed Spark and Hive-based ETL flows to clean, transform, and aggregate retail data for business teams. Used Hive features such as partitioning, dynamic partitions, and bucketing to manage large transaction tables and improve query performance. Created and maintained both managed and external Hive tables based on access and data retention needs. Loaded processed data into HBase for fast access where near-real-time lookups were required. Wrote SQL, Hive, and Python checks to validate data loads and confirm record counts, duplicates, and data consistency. Built Python scripts to automate sampling and verification of retail datasets to ensure data accuracy. Worked with Avro data formats for Kafka producers and consumers and processed Avro files using Spark and Python. Improved Spark job performance by moving older RDD-based logic to DataFrames and Spark SQL. Built small PySpark proof-of-concepts to test new processing ideas before moving them into production. Created dashboards and reports in Tableau and Python to show sales trends, customer behavior, and operational metrics to business users. Supported and configured tools such as Elastic search, Logstash, Kibana, Cassandra, NiFi, and Kafka as part of the overall data platform deployed on AWS. Monitored and supported Kafka topics and clusters using Kafka Manager to ensure smooth data flow across environments. Environment: Hadoop, Hive, Spark, HDFS, HBase, Kafka, Sqoop, Flume, NiFi, MapReduce, Oozie, Pig, SQL, Python, Tableau, AWS (EC2, S3), Cloudera Manager, ETL Client: Rentomojo, Hyderabad, India Nov 2014 Dec 2017 Role: Data Engineer Responsibilities: Responsible for creating Data ingestion pipelines from different data sources. Developed data pipelines using Azure Databricks to automate the data preparation for modelling. Developed Spark applications using Spark - SQL in Azure Databricks for data extraction, transformation and aggregation from multiple file formats for analyzing & transforming the data to uncover insights into the customer usage patterns Created delta tables and scheduled jobs in Databricks for automating the data extraction process from various sources. Implemented the project using ASP.NET, Visual C# and back-end database as Oracle Developer. Implemented Spark using Python (pySpark) and SparkSQL for faster testing and processing of data. Imported real time weblogs using Kafka as a messaging system and ingested the data to Spark Streaming. Performed Data Analysis, data quality, designs and implemented best practices. Built and managed ETL processes to extract, transform, and load data into Big Query from various sources, including cloud storage, Hadoop, and relational databases. Support the private cloud environment which consists of Confidential as the cloud product, KVM as the hypervisor, Confidential Nova dashboard for orchestration. Involved in the development of DOM parsing, SQL procedures and in development of IVR in VXML,CCXML by using Java and JSP. Migrating SQL database to Azure data Lake, Azure data lake Analytics, Azure SQL Database, Data Bricks and Azure SQL Data warehouse and Controlling and granting database access and Migrating On premise databases to Azure Data lake store using Azure Data factory (ADF). Created Pipelines in ADF using Linked Services/Datasets/Pipeline/ to Extract, Transform and load data from different sources like Azure SQL, Blob storage, Azure SQL Data warehouse, write-back tool and backwards. Environment: Azure Data Factory, ADLS, Azure Blob Storage, ADF, SQL, python, Pyspark, Git Azure DevOps, MongoDB, Snowflake, Confluence, JIRA, Service Now EDUCATION: Bachelor s in computer science engineering, from CR Reddy College, India. Masters in IT Management, from Concordia university St Paul CERTIFICATIONS: AZ-305: Designing Microsoft Azure Infrastructure Solutions | ID H708-1085 Keywords: csharp continuous integration continuous deployment quality analyst artificial intelligence message queue business intelligence sthree database active directory information technology Arizona Idaho |