Home

MD Mubassir - Data Engineer
[email protected]
Location: Chicago, Illinois, USA
Relocation: Yes
Visa: GC
MD MUBASSIR
Data Engineer
[email protected] | (872) 212-4111 | Linkedin
PROFESSIONAL SUMMARY:
A results-driven Data Engineer with 10+ years of progressive experience in data quality analysis, SQL-driven data investigation, and cross-functional collaboration between business and engineering teams to ensure accurate, complete, and operationally reliable data across enterprise environments.
Proven expertise delivering 25-35% gains in operational efficiency through predictive analytics, machine learning, and cloud-native data architecture on AWS, Azure, and Snowflake platforms.
Expert-level SQL skills including complex query writing, ad-hoc data investigation, advanced indexing, and partitioning across Snowflake, Redshift, SQL Server, and PostgreSQL, consistently achieving 40%+ query performance improvements in production environments.
Proficient in Python with Pandas for large-scale data manipulation, analysis of 10M+ record datasets, and scripted data quality checks, data completeness validation, and anomaly detection on modern platforms including Snowflake and AWS Redshift.
Specialization in creating executive dashboards in Power BI and Tableau that accelerated leadership decision-making by more than 40% through real-time data visualization.
Skilled in data governance, data quality monitoring, and regulatory compliance frameworks meeting HIPAA and GDPR standards, with extensive experience in forecasting and demand planning models improving financial planning accuracy by 30%.
Experienced acting as a liaison between business and factory/operations teams and technical data engineering teams; providing structured feedback to external partners and vendors to resolve data quality issues, drive continuous data improvement, and establish enterprise-grade data governance standards across analytics pipelines.
Demonstrated track record of root-cause analysis on data pipeline failures, investigating data completeness gaps and structural issues, and implementing continuous improvement processes that reduced data errors by 30-35% across enterprise reporting systems.
Led cloud migration initiatives transitioning legacy on-premises and Hadoop-based workloads to scalable GCP (BigQuery, Dataproc),Azure and AWS cloud-native platforms, achieving 25-35% infrastructure cost savings while enabling CI/CD-driven ETL pipeline automation, Apache Airflow orchestration, and Snowflake-based data warehousing at enterprise scale.

TECHNICAL SKILLS:
BI & Visualization Tools: Tableau, Power BI, Looker, QlikView
Programming Languages: Python, R, SQL, SAS
Libraries & ML Frameworks: Pandas, NumPy, Scikit-learn, XGBoost, TensorFlow, LangChain, Matplotlib, Seaborn
Databases: MySQL, PostgreSQL, SQL Server, Oracle, Snowflake, Redshift
Cloud Platforms: AWS (Redshift, S3, Glue, SageMaker), Azure (Synapse, Data Factory, Blob Storage), GCP (BigQuery, Dataflow, Dataproc, GCS, Cloud Composer, Pub/Sub, Dataplex)
ETL & Orchestration: Informatica, Talend, SSIS, Apache Airflow, Cloud Composer, Dataflow, dbt
Big Data Technologies: Spark, Hadoop, PySpark, Databricks, Hive, BigQuery, Dataproc
AI / GenAI: LLMs (GPT-4, Claude), RAG Systems, AI Agents, NLP, Predictive Modelling, Time-Series Forecasting
APIs & Data Formats: REST APIs, JSON, XML
DevOps & Version Control: Git, Jenkins, Docker, CI/CD Pipelines
Governance & Compliance: HIPAA, GDPR, Data Quality Frameworks, Data Governance

PROFESSIONAL EXPERIENCE:
CIGNA, Wilmington, DE, Aug 2025 Present
Senior Data Engineer
Responsibilities:
Architected a GenAI-powered clinical decision-support platform using RAG architecture and Azure OpenAI GPT-4, enabling clinicians to query 5M+ patient records in natural language and reducing time-to-insight by 65% for care management teams.
Deployed an LLM-based NLP pipeline using LangChain to extract structured insights from unstructured physician notes and medical transcripts, improving patient risk stratification accuracy by 32% and driving an 18% reduction in 30-day readmission rates.
Built an AI agent framework that automates end-to-end claims anomaly detection workflows, integrating predictive models with real-time Azure Synapse data feeds to flag fraudulent claims reducing insurance processing errors by 25% and saving $3M+ annually.
Designed and scaled HIPAA-compliant data pipelines on Azure processing 10TB+ of daily patient data, improving pipeline reliability to 99.9% uptime and reducing data latency by 55%; applied equivalent patterns using GCP BigQuery, GCS, and Dataflow for cloud-native batch and streaming pipelines.
Built and deployed XGBoost and Scikit-learn models to forecast healthcare utilization across multi-regional business units, improving financial planning accuracy by 30% and reducing budgetary risk through data-driven projections.
Built executive Tableau and Power BI dashboards unifying performance data across 20+ healthcare facilities, delivering real-time insights that accelerated leadership decision-making by 40% and informed strategic initiatives across clinical operations, finance, and quality.
Optimized SQL Server databases with advanced indexing and partitioning strategies, improving dashboard query performance by 45% and reducing reporting SLA breaches.
Designed scalable ETL processes using Talend, SSIS, and Apache Airflow, decreasing data latency and improving reporting efficiency across enterprise healthcare systems.
Established enterprise-wide data governance, KPI standardization, and regulatory compliance frameworks ensuring HIPAA and GDPR audit readiness across all analytics outputs.
Led migration of legacy reporting infrastructure to cloud-native Power BI platform, eliminating 60% of manual reporting and improving data consistency for 500+ clinical users.
Environment: Python, SQL, R, LangChain, Azure OpenAI, GPT-4, RAG Systems, NLP, Scikit-learn, XGBoost, Pandas, NumPy, Azure (Synapse Analytics, Data Factory, Blob Storage), GCP (BigQuery, Dataflow, Dataproc, GCS, Cloud Composer, Pub/Sub, Dataplex), Snowflake, Power BI, Tableau, Apache Airflow, REST APIs, JSON, SQL Server, PostgreSQL, Git, Jenkins, Docker, Agile/Scrum, JIRA, HIPAA, GDPR, Data Governance Frameworks, CI/CD Pipelines.
FIFTH THIRD BANK, Cincinnati, OH, Mar 2023 Jul 2025
Senior Data Engineer
Responsibilities:
Spearheaded enterprise analytics transformation for a revenue portfolio in excess of $500M by integrating predictive analytics and automated business intelligence solutions that improved financial decision-making accuracy by 30% year-over-year.
Designed and engineered scalable cloud data architecture solutions using AWS Redshift, S3, and Glue to handle large-scale transactional and customer data exceeding 5TB per day, reducing batch reporting cycle from 5 days to near-real-time.
Built and trained sophisticated machine learning forecasting models using Python, Scikit-learn, and XGBoost, resulting in a 30% year-over-year improvement in revenue forecasting accuracy directly informing CFO-level strategic planning.
Developed customer segmentation and lifetime value analysis models driving a 22% improvement in retail banking product cross-sell rates across consumer and commercial portfolios.
Implemented enterprise data governance practices that drove a 35% improvement in data accuracy while improving regulatory compliance and audit preparedness across all reporting pipelines.
Led the cloud migration of legacy on-premises and Hadoop-based reporting to AWS and GCP (BigQuery) cloud-native BI platform, reducing infrastructure costs by 25% and enabling self-service analytics for 300+ business users.
Created executive KPI dashboards to track financial risk, operational, and compliance metrics in support of board-level reporting.
Developed A/B testing infrastructure and statistical experimentation models to support loan product pricing optimization and measurement of deposit growth ROI.
Optimized complex SQL query performance, indexing, and partitioning designs resulting in over 40% improvements in query execution times in production environments.
Developed automated anomaly detection models to pinpoint financial irregularities and minimize reporting errors across enterprise data pipelines.
Environment: Python, R, SQL Server, AWS (Redshift, S3, Glue), GCP (BigQuery, GCS, Dataflow, Cloud Composer), Databricks, Power BI, Tableau, Talend, SSIS, Scikit-learn, Pandas, NumPy, REST APIs, Git, JIRA, Agile Methodology, Data Warehouse Architecture, ETL Pipelines.
DEPOSITORY TRUST & CLEARING CORPORATION, Coppell, TX, Oct 2020 Feb 2023
Data Analytics Consultant
Responsibilities:
Analysed broker-dealer trade submission and settlement cycle patterns, reducing participant fail rates by 20% across NSCC/DTC clearing operations.
Built predictive models forecasting trade settlement volumes and net debit cap utilization, reducing settlement failures by 15% and improving real-time risk decisions.
Designed Tableau dashboards for real-time visibility into trade flow, clearing fund requirements, and NSCC/DTC settlement obligations for executive and regulatory reporting.
Developed SQL data marts consolidating trade, position, and collateral data for risk analytics, regulatory capital reporting, and fee reconciliation.
Built participant risk-scoring models to flag margin breach and settlement default risks, reducing exposure incidents by 17% and improving systemic risk monitoring.
Automated SEC, FINRA, and Federal Reserve regulatory reporting pipelines, cutting manual effort by 50% and eliminating recurring compliance errors.
Used Python and Pandas and PySpark on Databricks to process 10M+ daily trade records for intraday settlement monitoring and end-of-day position reconciliation at market infrastructure scale.
Applied time-series modelling to detect trade volume and margin volatility cycles, supporting liquidity planning at quarter-end and options expiry events.
Built AWS SQL-based reporting infrastructure ensuring high availability and scalability for clearing analytics, with data governance, automated data completeness checks, and validation controls for auditability.
Optimized Talend and SSIS ETL pipelines for real-time trade feed ingestion, improving data refresh latency for intraday risk decisions.
Environment: Python, SQL, Pandas, PySpark, Databricks, NumPy, Tableau, Power BI, Azure SQL Database, MySQL, PostgreSQL, Talend, SSIS, Excel (Advanced Macros & VBA), REST APIs, Time-Series Modelling, Git, Agile Framework, Data Marts, Data Modelling, ETL Development.
TAILORED BRANDS, Fremont, CA, Jan 2019 Sep 2020
Business Intelligence Analyst
Responsibilities:
Created enterprise dashboards to assist product and engineering teams with performance analysis and key performance indicator (KPI) monitoring.
Designed relational data models to improve the performance of reporting and optimize queries in SQL environments.
Developed automated stored procedures and reporting scripts, leading to a 40% decrease in the time spent manually creating reports.
Developed KPI scorecards to improve accountability and transparency in business units.
Performed root cause analysis to determine operational inefficiencies and recommend process improvements.
Developed extract, transform, and load (ETL) processes using Informatica to improve data consolidation and transformation accuracy.
Improved data warehouse schema design to accommodate complex analytical scenarios and advanced reporting requirements.
Built merchandise demand forecasting models to inform seasonal product assortment decisions and store-level inventory planning, reducing stockout events by 18%.
Designed visual analytics for store performance, sales trends, and markdown effectiveness to improve executive-level retail reporting clarity.
Environment: SQL Server, Oracle, Informatica, Tableau, Power BI, Excel, Stored Procedures, T-SQL, Data Warehouse Schema Design, ETL Processes, Data Modelling (Star & Snowflake Schema), Git, JIRA, Agile/Scrum, KPI Scorecards, Reporting Automation Tools.
ADVITHRI TECHNOLOGIES, Hyderabad, India, Jun 2015 Aug 2018
Data Analyst
Responsibilities:
Created complex SQL queries and reporting solutions for operational decision-making purposes in various departments, reducing ad-hoc reporting turnaround from 3 days to same-day delivery.
Performed exploratory data analysis to uncover trends and areas for performance improvement across client delivery and product usage datasets.
Built Excel dashboard solutions that automated financial and operational reporting processes, reducing reporting errors by 30% through automated validation rules.
Pre-processed large datasets, leading to enhanced data quality and integrity across reporting systems.
Contributed to the development of the company's data warehouse and schema optimization projects, establishing a foundational BI infrastructure across 5+ departments.
Analysed trends and variances to improve forecasting models and support planning decisions.
Assisted in CRM reporting and customer analytics projects.
Created ad-hoc analytical reports on product usage, client delivery metrics, and operational performance to guide executive-level decision-making.
Environment: SQL, MySQL, SQL Server, Excel (Pivot Tables, Macros, VBA), Crystal Reports, Basic Python Scripting, Data Cleaning Tools, ETL Support Processes, Relational Database Design, Reporting Automation, Version Control Systems, Business Intelligence Foundations.

EDUCATION: Jun 2011 May 2015
Bachelor's in Computer Science Engineering, Veermata Jijabai Technological Institute, Mumbai, India.
Keywords: continuous integration continuous deployment artificial intelligence machine learning business intelligence sthree active directory rlang California Delaware Maryland Ohio Texas

To remove this resume please click here or send an email from [email protected] to [email protected] with subject as "delete" (without inverted commas)
[email protected];7493
Enter the captcha code and we will send and email at [email protected]
with a link to edit / delete this resume
Captcha Image: