| Mohammed Numair Ahmed - Data Scientist/ AIML Engineer/ AI Engineer |
| [email protected] |
| Location: Remote, Remote, USA |
| Relocation: Yes |
| Visa: Green Card |
|
Senior Data Scientist and Gen AI/ML Engineer with 12+ years of experience building and deploying enterprise machine learning systems across healthcare, energy, and financial services. Started out writing SQL pipelines and Tableau dashboards at Paychex, moved into Python-based ML at Resido and Shell, then spent four years at CVS Health building predictive healthcare analytics and RAG-based document intelligence platforms. Currently at Community Health System leading clinical AI systems that automate physician documentation, patient risk scoring, and prior authorization workflows using LangChain, fine-tuned LLMs, and AWS SageMaker. Experienced across the full ML lifecycle from raw EHR data ingestion and feature engineering through model training, deployment, drift monitoring, and presenting findings to clinical and executive stakeholders.
TECHNICAL SKILLS ML / Modeling XGBoost, LightGBM, scikit-learn, Random Forest, SVM, KNN, Logistic Regression, Ridge/Lasso Regression, GBM, Decision Trees, Ensemble Methods, Hyperparameter Tuning, Cross-Validation, SHAP, LIME Deep Learning PyTorch, TensorFlow, Keras, CNN, ANN, Autoencoders, Transfer Learning (VGG, ResNet, InceptionNet, MobileNet), Object Detection (YOLO, Faster R-CNN), Image Segmentation (U-Net, Mask R-CNN), Optimization Algorithms (SGD, Adam, RMSprop) Generative AI LangChain, LangGraph, LangSmith, Agentic AI, RAG pipeline design, Fine-tuning (LoRA, QLoRA, PEFT), LLMs (GPT, Claude, LLaMA, BERT, T5), Ollama, Vector Databases (Pinecone, FAISS, Neo4j), Prompt Engineering, Hugging Face Transformers, OpenAI API AWS SageMaker (Pipelines, Training Jobs, Batch Transform, Feature Store, Model Monitor, Model Registry), Bedrock, Lambda, Step Functions, Glue, Redshift, EMR, Kinesis, DynamoDB, S3, EC2, API Gateway, Athena, QuickSight, Textract, Comprehend, CloudWatch, IAM, CloudFormation Data Engineering PySpark, Apache Spark, Hadoop, Hive, HBase, Sqoop, Pig, Kafka, Apache Airflow, AWS Glue ETL, Snowflake, ETL pipeline design, HL7/FHIR data integration MLOps MLflow, Weights & Biases, SageMaker Model Monitor, DVC, Docker, Kubernetes, GitHub Actions, Jenkins, CI/CD pipelines, Model drift detection, A/B testing, Experiment tracking, Feature Stores, Kubeflow, Terraform NLP BioBERT, ClinicalBERT, scispaCy, NLTK, spaCy, Named Entity Recognition, Sentiment Analysis, Sequence-to-Sequence, LSTM, Bi-LSTM, Transformer models (BERT, GPT, T5), Machine Translation, Hugging Face Analytics & BI Power BI, Tableau, Amazon QuickSight, Python (pandas, numpy, scipy, matplotlib, seaborn), R, SQL, Jupyter Notebook, Google Colab Languages Python, SQL, PySpark, R, Java, JavaScript, Bash, Go, Scala, HTML, CSS Databases MySQL, PostgreSQL, MongoDB, MS SQL Server, SQLite, Redshift, RDS, Teradata, Oracle, HBase Compliance & Security HIPAA, ICD-10, SNOMED CT, CPT, LOINC, FHIR, HL7, AWS IAM, Model Explainability, Bias Detection, AI Governance Frameworks, Ethical AI PROFESSIONAL EXPERIENCE Community Health System Franklin, TN Senior Data Scientist / Gen AI Engineer August 2023 Present Community Health System (CHS) operates 80+ hospitals across 16 states and is one of the largest publicly traded hospital companies in the US. Joined the Enterprise Analytics and AI team at a time when the organization was investing heavily in reducing clinician documentation burden and improving care consistency across its hospital network. All work runs within a HIPAA-compliant architecture under formal model governance review before anything reaches production. The core problem was that clinical reviewers were spending 60 to 70 percent of their working day manually cross-referencing physician notes, EHR records, and policy documents before they could make any care decision. The team needed something that could surface the right clinical context automatically, without putting PHI at risk. Designed and built a RAG-based clinical knowledge retrieval system using LangChain, vector embeddings, and semantic search across EHRs, clinical notes, and medical literature. The retrieval layer combined structured SQL-based reasoning with unstructured LangChain-powered question answering, delivering highly accurate contextual search results across clinical databases and document repositories. Implemented enterprise-scale semantic search using vector embeddings and vector databases enabling real-time retrieval across millions of structured and unstructured healthcare records. Fine-tuned large language models including Claude and GPT-4 using LoRA and PEFT parameter-efficient techniques to handle discharge summary generation, call transcript summarization, clinical documentation assistance, and regulatory compliance validation. Tuned prompts and grounding thresholds carefully the first pass was suppressing valid outputs because paraphrased policy text was being flagged as hallucination. After calibration, documentation accuracy across care teams improved substantially. Engineered clinical NLP pipelines using BioBERT, ClinicalBERT, and scispaCy to extract medical entities diagnoses, medications, symptoms, procedures, and clinical events from unstructured physician notes, discharge summaries, and patient records. Applied ICD-10, SNOMED CT, CPT, and LOINC terminology standards to normalize clinical data across analytics, predictive modeling, and interoperability systems. PHI handling was the non-negotiable constraint throughout. Every piece of clinical text ran through HIPAA-compliant de-identification protocols before reaching model endpoints. Enabled interoperability between multiple healthcare platforms by integrating FHIR and HL7-based data exchange frameworks, supporting standardized ingestion and real-time data sharing across hospital systems. AWS IAM security protocols governed all access to sensitive analytics services. Built and maintained comprehensive MLOps frameworks using AWS SageMaker Model Monitor, CloudWatch, MLflow, and automated CI/CD pipelines to track model performance, detect drift, enforce governance, and enable continuous retraining. Implemented SHAP and LIME explainability dashboards to satisfy regulatory compliance requirements for clinical AI models in decision support environments. Designed scalable ETL pipelines using Python, PySpark, AWS Glue, and Airflow to ingest, transform, and integrate large-scale healthcare datasets into Snowflake-based data warehouses. Built optimized healthcare data models including star schema and snowflake schema to support business intelligence reporting, clinical analytics, and financial performance monitoring. Developed and deployed an automated prior authorization workflow using LangGraph multi-agent orchestration and AWS Step Functions, reducing prior auth turnaround times from three to five business days down to under four hours for routine cases. The system parsed payer policy documents in real time, cross-referenced them against patient clinical records, and generated structured justification letters that compliance teams reviewed rather than authored from scratch, cutting per-case processing costs substantially. Architected a patient risk stratification platform using ensemble models (XGBoost, LightGBM, and logistic regression) trained on structured EHR data covering diagnoses, labs, vitals, and medication history. Models predicted 30-day readmission, sepsis onset, and deterioration risk scores that surfaced directly in the clinical workflow dashboard. Integrated SHAP-based explainability output into the clinician-facing UI so physicians could see the top contributing factors behind every risk prediction, improving trust and adoption across the hospital network. Led the design of a real-time clinical alerting system integrated with the hospital s EHR using HL7 FHIR APIs and AWS Kinesis data streams. The pipeline ingested continuously updating patient vitals, lab results, and nursing notes, ran them through trained deterioration models on a sliding window basis, and pushed risk alerts to care team dashboards within 90 seconds of a triggering event. Worked closely with clinical informatics and nursing leadership to calibrate alert thresholds and reduce alert fatigue, increasing alert actionability rates across pilot units. Impact: Reviewer document review time dropped significantly across pilot markets. Deployed AI-powered clinical copilots that assist physicians with real-time summarization of patient histories, automated chart review, and contextual recommendations based on clinical guidelines. Mentored junior data scientists and ML engineers on generative AI development, data engineering pipelines, experiment tracking, and production deployment strategies in regulated healthcare environments. Environment: Python, Scikit-learn, NumPy, SciPy, Matplotlib, Pandas, AWS S3, DynamoDB, Lambda, EC2, SageMaker, EMR, Redshift, Snowflake, LangChain, BioBERT, ClinicalBERT, SHAP, LIME, MLflow, Airflow, Docker, Kubernetes, Power BI CVS Health Irving, TX AI Engineer February 2021 August 2023 CVS Health operates one of the largest pharmacy networks in the US and carries a significant managed care business through Aetna. Joined the Enterprise AI team when the organization was actively scaling its machine learning investment to improve clinical decision support and reduce operational friction across its pharmacy and care management workflows. The clinical analytics team was working off disconnected data sources with no unified ML platform patient engagement signals, claims activity, and operational metrics sat in separate systems and nobody had a reliable way to combine them for predictive modeling. The ask was to build a production-ready analytics layer that could support both batch prediction and real-time inference. Developed and deployed real-time predictive models using TensorFlow, Keras, and scikit-learn to analyze patient engagement patterns, claims activity, and operational signals. Built advanced models including Ridge Regression, Lasso, XGBoost, and K-Means clustering to predict patient adherence, identify high-risk populations, and improve care management strategies with feature engineering, data preprocessing, imputation, and feature selection at each stage to ensure high-quality inputs. Fine-tuned transformer-based LLMs using Hugging Face frameworks and LLaMA architectures with LoRA and PEFT techniques to automate clinical document classification, healthcare policy analysis, and regulatory compliance monitoring. Designed and implemented vector embeddings, and semantic search to enable contextual retrieval of clinical documentation, medical policies, and patient support knowledge bases. Integrated AWS cloud services including EC2, S3, Redshift, RDS, API Gateway, ELB, SNS, and EBS with Google Cloud Vertex AI platforms to support scalable ML training workflows, model experimentation, and enterprise AI deployments. Built containerized AI model deployment pipelines using Docker and Flask-based microservices, enabling seamless integration of predictive models and inference APIs within enterprise healthcare applications. Designed secure private cloud networking to safely expose ML inference endpoints under HIPAA-compliant access controls. Implemented enterprise monitoring and observability using AWS CloudWatch, GCP Monitoring, Amazon QuickSight, and Power BI to track model performance, latency, prediction accuracy, and data drift across production AI systems. Optimized Snowflake-based healthcare data warehouses by designing shared dimension schemas enabling efficient ad hoc querying and cross-domain healthcare analytics. Designed and delivered a pharmacy adherence intelligence platform using gradient boosting models trained on historical prescription fill patterns, chronic disease indicators, and member demographic features. The platform identified members at high risk of medication non-adherence 60 to 90 days in advance, enabling targeted pharmacist outreach. Deployed as a batch inference pipeline on AWS SageMaker with daily scoring refreshed against the pharmacy data warehouse in Snowflake, directly supporting the Aetna care management team s population health program. Built a healthcare policy document QA system using a multi-stage RAG architecture over a corpus of 4,000+ internal payer policies, clinical guidelines, and regulatory documentation. Implemented hybrid retrieval combining BM25 keyword ranking with dense vector search using Pinecone, with a re-ranking layer to maximize relevance for clinical policy questions. The system reduced the time compliance reviewers spent manually searching policy documentation, enabling them to handle higher caseloads without additional headcount. Impact: Delivered scalable predictive analytics that reduced patient risk prediction time and improved care management efficiency across CVS Health platforms. Collaborated with cross-functional teams including healthcare analysts, data engineers, product managers, and compliance teams to design AI-driven solutions improving patient outcomes and regulatory compliance. Environment: Python, Scikit-learn, TensorFlow, Keras, PySpark, AWS (S3, DynamoDB, Lambda, EC2, SageMaker, EMR, Redshift), GCP Vertex AI, Snowflake, LangChain, Docker, XGBoost, MLflow, Power BI, QuickSight Shell Houston, TX Data Scientist / ML Engineer April 2017 February 2021 Shell is one of the world's largest energy companies, operating across upstream, downstream, and integrated gas segments globally. Joined the Enterprise Analytics team in an Agile delivery environment to build data science solutions supporting workforce performance, demand forecasting, and operational intelligence across Shell's US operations. Performed exploratory data analysis, univariate and multivariate statistical analysis, and time series forecasting to identify workforce trends, demand patterns, and operational performance indicators across Shell's US business units. Developed machine learning models using XGBoost, Random Forest, and SVM in Python with full feature engineering, cross-validation, and SHAP-based model interpretability to predict performance outcomes and detect operational risks before they escalated. Built scalable ETL pipelines integrating enterprise datasets from AWS S3, RDS, and Snowflake using Python, Pandas, NumPy, and Scikit-learn to support ML model training and analytics reporting. Designed star and snowflake schema data models supporting both OLTP and OLAP systems, and developed RESTful APIs and Python backend services to expose ML model outputs for enterprise web applications. Implemented NLP solutions using NLTK and scikit-learn for sentiment analysis and text classification of workforce feedback and operational documentation. Built containerized ML workflows using Docker and AWS SageMaker, with Kubernetes orchestrating distributed compute to support scalable training and inference pipelines. Designed hybrid data retrieval solutions combining advanced SQL with semantic search techniques to improve enterprise knowledge discovery. Developed interactive dashboards using Power BI and Power Query to visualize workforce performance metrics, operational KPIs, and model insights for business stakeholders. Built automated data ingestion and transformation pipelines to ensure high data quality, integrity, and timely availability across all analytics datasets. Collaborated with data governance and compliance teams to implement a centralized ML experiment tracking and model versioning system using MLflow on AWS, establishing standardized conventions for model registration, artifact storage, and performance metadata tracking across the analytics team. This formalized the path from experimentation to production deployment and significantly reduced the time to promote models from development into scheduled batch inference workflows on SageMaker, establishing ML reproducibility practices that were later adopted across other Shell analytics teams globally. Impact: Delivered a predictive analytics platform actively used by workforce and operations teams across four years. ETL pipelines ensured reliable, timely availability of analytics datasets across distributed enterprise systems. Dashboard adoption spread across multiple business units as the primary tool for operational performance monitoring. Environment: Python, Pandas, NumPy, Scikit-learn, SciPy, NLTK, XGBoost, Random Forest, SVM, SQL, Oracle, Teradata, Snowflake, AWS (S3, RDS, SageMaker), Docker, Kubernetes, ETL, OLAP/OLTP, Power BI, Tableau Resido Melville, NY Junior Data Scientist / ML Engineer September 2016 March 2017 Resido was a real estate analytics firm in Melville, NY. Joined the data science team to support customer segmentation and market analytics initiatives, working across Python-based ETL pipelines, ML model development, and stakeholder reporting. Performed exploratory data analysis and data preparation using Python, Pandas, NumPy, and SciPy to support customer segmentation and market analytics. Built ETL pipelines using Python, Alteryx, and Snowflake to integrate and standardize enterprise datasets from 40+ distributed internal and external data sources, including web-scraped data and third-party feeds. Implemented automated data processing workflows using Hadoop, Mahout, and MongoDB to support scalable ingestion and transformation pipelines. Developed machine learning models using scikit-learn for customer segmentation, behavioral analysis, and predictive scoring to support targeted business strategies. Performed statistical analysis and multivariate data validation to identify data quality issues and ensure reliability of enterprise datasets. Designed and deployed reporting solutions using Python APIs and Tableau dashboards to provide business stakeholders with actionable insights, and monitored data pipelines in collaboration with engineering and QA teams to ensure system reliability and regulatory compliance. Conducted real estate market price modeling using regression and tree-based models to forecast property valuation trends across target markets. Integrated web-scraped listing data, zip code economic indicators, and historical transaction records into a unified Snowflake analytics layer to support property acquisition and pricing strategy decisions. Delivered model output summaries and visual trend analysis via Tableau dashboards consumed by the business development and acquisitions teams, sharpening their ability to identify undervalued markets ahead of competitors. Environment: Python, Pandas, NumPy, SciPy, Seaborn, Matplotlib, Scikit-learn, NLTK, SQL, Snowflake, Alteryx, Hadoop, MongoDB, ETL, Tableau, OLTP/OLAP, Oracle, SQL Server Paychex Rochester, NY Data Analyst January 2014 August 2016 Paychex is one of the largest payroll and HR services companies in the US, serving over 700,000 clients. Joined the analytics and reporting team during a period when the business was moving away from static Excel-based reporting toward self-service BI and automated data workflows. Developed and optimized 200+ stored procedures, database views, and SQL queries in MS SQL Server to support payroll, tax, compliance, and financial reporting systems. Collected, validated, and integrated high-volume sales and operational data from 40+ internal and external sources to support analytics and reporting initiatives across departments. Designed and delivered interactive Tableau dashboards and KPI reports that enabled business stakeholders to monitor sales performance, compliance metrics, and operational trends without submitting reporting tickets. Automated data processing workflows using Python and SQL to extract, clean, and standardize large datasets reducing manual reporting effort materially. Implemented web scraping solutions using Python and built API-based reporting integrations for real-time dashboard updates, and collaborated with database administrators and business analysts to deliver scalable data solutions across the organization. Partnered with payroll operations and compliance leadership to redesign the client compliance reporting framework, migrating 60+ static Excel-based reports into parameterized Tableau dashboards backed by optimized MS SQL Server stored procedures. The new reporting layer reduced report generation cycle times from two to three hours per report down to on-demand self-service, freeing the analytics team from routine report production and allowing them to redirect capacity toward higher-value analytical projects. Partnered with the database administration team to optimize underlying query execution plans, cutting average report load time significantly across the most-used compliance reports. Environment: MS SQL Server, T-SQL, Python, Tableau, Advanced Excel, Oracle, Web Scraping, ETL Automation, APIs, Basic JavaScript EDUCATION Master's in Information Technology and Project Management Cumberland University, TN, USA (Aug 2012 Dec 2013) Bachelor of Technology in Computer Science Engineering MRCET (Aug 2008 July 2012) CERTIFICATIONS AWS Certified Machine Learning Specialty Databricks Associate ML Engineer TensorFlow Developer Certificate Keywords: continuous integration continuous deployment quality analyst artificial intelligence machine learning user interface business intelligence sthree active directory rlang golang trade national microsoft mississippi Connecticut Delaware New York Tennessee Texas |