Home

Pavan Kalyan Kondle - Gen AI/DS/AI ML
[email protected]
Location: Dallas, Texas, USA
Relocation: YES
Visa: GC
Resume file: Pavan Kalyan Kondle_1783526598807.docx
Please check the file(s) for viruses. Files are checked manually and then made available for download.
Pavan Kalyan Kondle
AI/ML Engineer / Generative AI Engineer / Data Scientist
[email protected] / Phone: +18049932915 / LinkedIn
OBJECTIVE:
Senior AI/ML Engineer and Data Scientist with 10+ years of experience delivering end-to-end AI, machine learning, and Generative AI solutions for financial services, healthcare, insurance, telecom, and banking organizations. Skilled in Python, SQL, Claude, AWS, Azure, GCP, LLMs, RAG, LangChain, vector databases, and FastAPI, with hands-on experience building scalable, high-impact systems that improve automation, analytics, and business decision-making.

PROFESSIONAL SUMMARY:
Strong hands-on experience building scalable machine learning and statistical modeling solutions using Python, Scikit-learn, TensorFlow, PyTorch, SQL, and PySpark for fraud detection, forecasting, customer segmentation, operational analytics, and anomaly detection use cases. Built reusable data pipelines and model workflows capable of supporting both project-based initiatives and fast-moving executive analytical requests.
Developed data visualization dashboards and executive reporting solutions using Power BI, Tableau, Plotly, and Python visualizations to communicate trends, operational KPIs, model outcomes, and business insights to leadership teams. Translated technical findings into actionable recommendations that supported business planning, operational prioritization, and strategic decision-making.
Hands-on experience with classical Data Science algorithms including Regression, Random Forest, XGBoost, Clustering, PCA, Forecasting, and Anomaly Detection models.
Built and supported Generative AI and Agentic AI workflows using OpenAI, Claude, LangChain, LangGraph, and Prompt Engineering techniques for enterprise search, conversational analytics, and intelligent workflow automation. Worked with GitHub Copilot and AI-assisted development workflows to accelerate development tasks, improve code quality, and simplify repetitive implementation activities.
Experienced working across distributed enterprise data environments including Snowflake, Databricks, SQL Server, PostgreSQL, MongoDB, and cloud-native storage systems. Collaborated with cross-functional business, operations, analytics, and engineering teams to independently deliver scalable AI and data science solutions in fast-paced enterprise environments.
CERTIFICATIONS:

Applied Data Science and Machine Learning, IIT Madras, 2023
Azure AI Engineer Associate (AI-102)

TECHNICAL SKILLS:

Programming Python, R, SQL, Java, Next.js, JavaScript, React, .NET, C#, Shell Scripting
ML/AI Frameworks Scikit-learn, TensorFlow, PyTorch, Keras, XGBoost, LightGBM, CatBoost, Prophet, ARIMA, RAPIDS, MLlib
GenAI & LLMs GPT-3/3.5/4/4o, Claude 2/3, Gemini, Mistral, LLaMA 2/3, Falcon, Mixtral, Cohere Command R, Hugging Face Transformers, Anthropic Claude, OpenAI API, Azure OpenAI, AWS Bedrock, AutoGen, MCP
Prompt Engineering Chain-of-Thought (CoT), ReAct, Role Prompting, Retrieval-Aware Prompting, Toolformer, System Messages
RAG & Memory LangChain, LangGraph, Semantic Kernel, DSPy, PromptLayer, AutoGen, CrewAI, Haystack, FAISS, Pinecone, Weaviate, Qdrant, Elasticsearch, RAG Pipelines, RAG Fusion, Self-RAG
NLP / CV / Time-Series SpaCy, NLTK, Hugging Face Datasets, OpenCV, Tesseract OCR, DeepStream, Whisper, CLIP, BLIP, LLaVA, Gemini Vision, Sentiment Analysis, NER, Semantic Search, Document Summarization, Embedding Models, Anomaly Detection, Time-Series Forecasting
MLOps & Pipelines MLflow, TFX, Kubeflow, Airflow, DVC, Feast, Model Registry, CI/CD, Docker, Kubernetes (EKS, GKE, AKS, OpenShift), Jenkins, Terraform, Helm, ArgoCD, GitHub Actions, Serverless (AWS Lambda, Cloud Run)
Model Tuning LoRA, QLoRA, PEFT, SFT, RLHF, vLLM, Hugging Face Transformers
Data Engineering BigQuery, Apache Iceberg, Apache Spark, Databricks, Snowflake, ETL/ELT Pipelines, Data Modeling, Streaming Analytics, Snowpark, Snowflake cortex
Cloud Platforms AWS (S3, SageMaker, Amazon Q, Quicksight, EKS, EMR, EC2, Fargate, IAM, Lambda, Redshift, DynamoDB, Bedrock, CloudWatch, Secrets Manager, Step Functions, Webhooks), Azure (ML Studio, Synapse, Cosmos DB, Azure OpenAI), GCP (Vertex AI, BigQuery, Cloud Run, ADK, A2A, Pub/Sub, Spanner, Dataflow, Cloud IAM)
Backend & API Java, Spring Boot, Spring Security, Python, Express.js, Next.js, REST APIs, BFF Proxy, JWT Forwarding, GraphQL, SSE Streaming, JPA/Hibernate
Monitoring & Security Prometheus, Grafana, Evidently AI, AWS CloudWatch Logs Insights, ELK Stack (Elasticsearch, Logstash, Kibana), Datadog, Sentry, IAM, HIPAA Compliance, Audit Logging, Secure APIs, RBAC
Databases Qdrant, Milvus, PostgreSQL, MySQL, MongoDB, Redshift, BigQuery, Cosmos DB, Spanner, DynamoDB
Visualization & BI Tableau, Power BI, R Shiny, Jupyter Notebook, Matplotlib, Seaborn, Plotly, Looker
DevOps & Automation Docker, Kubernetes (EKS, GKE, AKS, OpenShift), Terraform, Jenkins, GitHub Actions, Git, GitOps, Infrastructure as Code (IaC), Event Hubs, Step Functions, Pub/Sub
Other AI Capabilities Agentic AI Systems, Multi-Agent Coordination, Generative AI Automation, Chatbot Design, Knowledge Extraction, Personalized Recommendations, Clinical Risk Modeling, Real-time Inference, A/B Testing

WORK EXPERIENCE:
Client: Capital One, Dallas, TX Nov 2023 Till Date
Role: AI/Gen AI Engineer
Responsibilities:
Executed end-to-end data science workflows using Python, SQL, PySpark, and Databricks to process fraud transactions, customer interactions, operational logs, and analyst investigation data from multiple enterprise systems. Built scalable preprocessing and feature engineering pipelines that improved downstream model performance and reduced manual data preparation effort.
Developed supervised machine learning models using Scikit-learn, TensorFlow, and PyTorch for fraud detection, transaction risk scoring, anomaly detection, and customer behavior analysis. Improved fraud detection precision by approximately 18% through iterative feature engineering and model tuning workflows.
Built forecasting and workload prediction models using statistical analysis and machine learning techniques to estimate investigation volumes, analyst workload, and customer support trends. Helped operations teams improve workforce planning and reduce investigation backlog during high-volume periods.
Designed and implemented reusable Python APIs and FastAPI-based microservices to expose fraud insights, risk scores, semantic retrieval outputs, and GenAI-driven summaries to downstream enterprise applications. Reduced dependency on manual analyst reviews by integrating model outputs directly into operational workflows.
Developed RAG-based enterprise search workflows using LangChain, OpenAI, Pinecone, and FAISS to support conversational fraud investigation and policy lookup capabilities. Improved contextual retrieval quality by tuning chunking logic, embedding selection, and semantic reranking strategies.
Built semantic search pipelines combining vector retrieval and keyword search across fraud investigation records, customer transcripts, and policy documents. Enabled analysts to query enterprise datasets using natural language instead of relying only on traditional SQL-based lookup processes.
Implemented Agentic AI workflows using LangGraph and Prompt Engineering techniques where AI agents handled retrieval, summarization, validation, and recommendation generation tasks across fraud investigation workflows. Improved investigation efficiency by separating complex reasoning tasks into modular execution steps.
Developed feature engineering pipelines using Python, PySpark, Pandas, NumPy, AWS Glue, and S3 to process transaction records, customer profiles, operational logs, and fraud case data.
Deployed ML models using SageMaker Endpoints, Lambda, Docker containers, and API Gateway to support real-time and batch inference workloads.
Developed Prompt Engineering strategies using Chain-of-Thought prompting, few-shot prompting, and structured response templates for fraud analytics and customer interaction use cases. Refined prompts using analyst feedback to improve response consistency and reduce unsupported outputs.
Worked with GitHub Copilot and AI-assisted development workflows to accelerate Python API development, SQL generation, test case creation, and repetitive code implementation tasks. Improved developer productivity during rapid iteration cycles across AI and analytics projects.
Built executive dashboards using Power BI, Tableau, and Plotly to visualize fraud trends, operational KPIs, model accuracy, investigation backlog, and customer behavior insights. Presented analytical findings and ML-driven recommendations to business stakeholders and operational leadership teams.
Performed exploratory data analysis and statistical modeling on large transaction datasets using Python, Pandas, NumPy, and SQL. Identified seasonal trends, high-risk transaction behaviors, and operational bottlenecks that influenced downstream ML and business decisions.
Developed scalable data pipelines using PySpark, SQL, and Snowflake to process structured and semi-structured enterprise datasets for analytics and machine learning workflows. Reduced batch processing time by approximately 30% through query optimization and distributed execution improvements.
Built anomaly detection models using unsupervised learning techniques including clustering and Isolation Forest approaches to identify emerging fraud patterns not captured through labeled datasets. Increased visibility into previously undetected transaction anomalies and high-risk behaviors.
Implemented model monitoring and drift detection workflows using MLflow, logging frameworks, and scheduled evaluation pipelines to track prediction quality and feature distribution changes over time. Automated retraining triggers when model performance dropped below operational thresholds.
Created reusable feature engineering pipelines using Python and SQL to standardize fraud-related features such as transaction velocity, geolocation variance, device changes, merchant risk patterns, and historical fraud linkage. Reduced duplicate feature creation effort across multiple ML initiatives.
Worked with Snowflake and Databricks to build enterprise analytical datasets supporting fraud analytics, customer segmentation, and executive reporting workflows. Improved access to consistent and reliable datasets for analytics and machine learning teams.
Built NoSQL-backed metadata and session tracking workflows using MongoDB and DynamoDB for conversational AI and investigation support systems. Improved retrieval consistency and state management across multi-step AI-driven workflows.
Designed optimized PostgreSQL queries and stored procedures for model-serving datasets and analytical reporting workflows.
Developed production-grade AI and analytics systems within regulated financial environments using RBAC, audit logging, encrypted APIs, and controlled data access workflows. Supported enterprise compliance requirements while maintaining scalable analytics operations.
Implemented multimodal AI workflows involving OCR extraction, document parsing, and text analysis to process customer-uploaded evidence and fraud-related documents. Reduced manual review effort by automating extraction and classification tasks for unstructured inputs.
Collaborated with fraud analysts, operations teams, product owners, and data engineers to convert business requirements into scalable analytics and AI solutions. Delivered both scheduled project work and ad hoc executive-driven analysis requests under tight timelines.
Participated in CI/CD workflows using GitHub Actions, Docker, Jenkins, and Kubernetes to automate testing, deployment, monitoring, and rollback support for AI applications and APIs. Reduced deployment effort and improved release stability across analytics environments.
Environment: Python, PyTorch, TensorFlow, Scikit-learn, AWS SageMaker, AWS Bedrock, AWS Lambda, AWS S3, AWS Glue, AWS EMR, AWS CloudWatch, API Gateway, Amazon EKS, Docker, Kubernetes, MLflow, LangChain, LlamaIndex, CrewAI, AutoGen, OpenAI, Anthropic Claude, Llama, Mistral, Hugging Face Transformers, Pinecone, FAISS, Chroma, FastAPI, REST APIs, PySpark, Pandas, NumPy, Databricks, Snowflake, GitHub Actions, Jenkins, Terraform, Git, Prometheus, Grafana, PostgreSQL, Redis, Linux.
Client: Parallon Healthcare, Nashville, TN Dec 2020 to Oct 2023
Role: AI/ML Engineer / Data Scientist
Responsibilities:
Built healthcare-focused NLP and sentiment classification workflows using Python, SpaCy, NLTK, Scikit-learn, TensorFlow, PyTorch, and Hugging Face Transformers to analyze clinical documentation, claims notes, denial comments, payer responses, and operational review notes.
Developed text classification models to categorize denial reasons, documentation gaps, claim review outcomes, payer response patterns, and compliance risk indicators, helping healthcare operations teams identify high-risk claims earlier and reduce manual review bottlenecks.
Used Azure Databricks for scalable feature engineering, model training, and distributed healthcare analytics workloads.
Built preprocessing pipelines using Python, SQL, Pandas, NumPy, and Spark to clean healthcare text, normalize claim terminology, standardize denial codes, remove noisy fields, and prepare model-ready datasets for NLP and forecasting use cases.
Built distributed NLP preprocessing and feature engineering pipelines using PySpark and Databricks to process large-scale healthcare claims and clinical datasets.
Developed transformer-based NLP workflows using BERT/RoBERTa models and Hugging Face Transformers to improve classification accuracy on clinical notes, payer correspondence, and denial explanations compared with earlier keyword-based review approaches.
Built feature engineering pipelines from claims history, denial codes, payer behavior, service dates, clinical documentation fields, and operational workflow events to support denial prediction, workload forecasting, and compliance risk scoring models.
Built API-driven integrations to expose claim insights, NLP classifications, risk scores, denial predictions, and forecasting outputs into internal healthcare applications, allowing operational users to consume AI outputs within existing workflows.
Developed automated Airflow workflows to schedule data ingestion, text preprocessing, feature generation, model scoring, evaluation reporting, and dashboard refreshes, reducing recurring manual analytics preparation effort.
Partnered with compliance teams, revenue cycle analysts, and business stakeholders to validate model outputs, review false positives, refine denial categories, and ensure AI recommendations aligned with healthcare operational rules.
Implemented data quality checks for missing claim fields, inconsistent denial codes, duplicate records, invalid service dates, and incomplete clinical documentation, improving reliability of downstream NLP and forecasting outputs.
Created dashboards and reports using Power BI, Plotly, and Python visualizations to communicate denial trends, documentation patterns, workload forecasts, model accuracy, and operational KPIs to healthcare business stakeholders.
Supported production improvement efforts by documenting data preparation logic, model evaluation steps, deployment notes, known limitations, and troubleshooting steps so support teams could maintain recurring AI/ML workflows more efficiently.
Used PostgreSQL for feature storage, model metadata management, and downstream analytics integration.
Environment: Python, SQL, Azure, Scikit-learn, TensorFlow, PyTorch, Hugging Face Transformers, SpaCy, NLTK, Pandas, NumPy, Spark, Airflow, Snowflake, PostgreSQL, Power BI, Plotly, Jupyter Notebook, NLP, Text Classification, Sentiment Analysis, Forecasting, Model Evaluation, Data Pipelines
Client: Geico Insurance, Getzville, NY Oct 2018 to Nov 2020
Role: Data Scientist
Responsibilities:
Built end-to-end credit risk prediction pipelines using Python, SQL, integrating policy, claims, and customer datasets to develop machine learning models for risk scoring and fraud detection.
Enabled insurance operations teams to improve decision-making through AI-powered analytics dashboards and chatbot interfaces, reducing claim processing turnaround time and improving risk assessment accuracy.
Performed large-scale exploratory data analysis (EDA) using Python (Pandas, NumPy, Matplotlib) to uncover claim anomalies, customer behavior patterns, and fraud indicators.
Engineered predictive features from structured policy and claims data using Python + SQL, improving model signal quality and boosting fraud detection precision by 18%.
Built ETL and data integration workflows using Snowflake and cloud-native services to support analytics and model training.
Developed supervised machine learning models using Scikit-learn (Random Forest, Gradient Boosting, Logistic Regression) to classify high-risk claims and potential fraud cases.
Implemented credit risk scoring models for insurance customers, enabling operational teams to prioritize investigations and reduce fraud losses.
Designed insurance data models integrating underwriting, claims, billing, and policy datasets, enabling unified analytics workflows across multiple insurance business functions.
Designed customer segmentation models using K-Means clustering, identifying high-risk behavioral groups and improving targeted fraud monitoring strategies.
Built time-series forecasting models using ARIMA and statistical techniques, predicting claim volumes and operational risk trends across insurance portfolios.
Performed model evaluation using cross-validation, ROC-AUC, precision-recall metrics, and statistical testing, ensuring robustness of risk prediction models.
Conducted feature importance analysis using SHAP and model explainability techniques, enabling regulatory-friendly interpretation of machine learning decisions.
Developed Python-based data preprocessing and transformation pipelines, cleaning and standardizing multi-source insurance datasets for model training.
Implemented automated model retraining workflows using Python scripts and scheduled batch pipelines, ensuring models stayed accurate as new claim data arrived.
Built fraud anomaly detection models using unsupervised learning techniques, identifying suspicious claims that deviated from historical behavior patterns.
Deployed machine learning models into production using AWS SageMaker for model hosting and batch inference, enabling real-time fraud scoring services.
Built scalable data ingestion pipelines using AWS S3 and EMR with PySpark, enabling distributed processing of large claims and policy datasets.
Environment: Python 3.7, Python, SQL, Pandas, NumPy, Matplotlib, Scikit-learn, Random Forest, Gradient Boosting, Logistic Regression, K-Means Clustering, ARIMA Forecasting, SHAP Explainability, PySpark, AWS SageMaker (Model Hosting), AWS S3, AWS EMR, Jupyter Notebook, Git, Linux.
Client: AT&T, Dallas, TX Mar 2016 to Sept 2018
Role: Python Data Developer
Responsibilities:
Supported analytics and machine-learning workflows on telecom operational data, helping identify service issues, usage anomalies, and customer-impacting patterns across large enterprise datasets.
Worked with operations-focused datasets tied to network performance, service utilization, and customer behavior, giving strong exposure to telecom-scale distributed systems and operational troubleshooting.
Built ETL and data validation workflows that improved reliability of reporting and analytical pipelines used by telecom operations teams.
Performed root-cause analysis on anomalous operational patterns using Python and SQL, helping teams diagnose service issues faster in complex data environments.
Created and optimized analytics solutions in Power BI, including Star Schema / Snowflake Schema data models, semantic models, and advanced DAX (Time Intelligence, YTD/MTD, Dynamic Ranking), supporting self-service analytics for 500+ users; managed Power BI Deployment Pipelines (Dev/Test/Prod) with Git-based versioning.
Created KPI dashboards for telecom operations that improved visibility into performance trends, usage behavior, and emerging service risks.
Collaborated with business and technical stakeholders to translate telecom operational pain points into scalable analytics and automation solutions.
Built, trained, and tuned ML models across classification, regression, clustering, and time-series forecasting using Scikit-learn, XGBoost, TensorFlow, and PyTorch, with structured evaluation frameworks to improve robustness and scalability.
Developed high-volume ETL/data prep workflows using Alteryx, blending and cleansing 10TB+ data from Oracle and SQL Server, cutting monthly preparation time by 40%.
Environment: Python (Pandas, NumPy, Scikit-Learn, Matplotlib), SQL, Power BI, Alteryx, Snowflake, Oracle DB, Microsoft SQL Server, Hadoop (Hive), Excel (VBA/Macros), Git, Linux, JIRA, ETL Pipelines, Telecom KPIs, KPI dashboards
Client: Federal Home Loan Bank of New York, NY April 2014 to Feb 2016
Role: Data Warehouse Developer
Responsibilities
Translated complex banking regulatory requirements into technical specifications, delivering Power BI Dashboards that
became the standard for senior leadership s weekly operational reviews.
Automated recurring liquidity and risk reports using SQL and Python (Pandas), replacing manual Excel macros and reducing monthly reporting cycle time by 40%.
Developed robust data validation and reconciliation scripts using Python and NumPy to detect discrepancies between the General Ledger and Risk systems, reducing data quality incidents by 30%.
Executed complex SQL queries to extract and aggregate operational metrics from Hadoop (Hive) and Oracle databases, creating optimized datasets for downstream analytics.
Designed and maintained the "Daily Liquidity" dashboard in Tableau, visualizing cash flow variances and intraday positions to support the Treasury desk s funding decisions.
Facilitated User Acceptance Testing (UAT) for new data warehouse releases, creating test cases to verify data accuracy across financial reporting systems.
Environment: Python (Pandas, NumPy), SQL, Power BI, SAS, Hadoop (Hive), Oracle DB, Excel (Advanced), Linux, JIRA.

EDUCATION:
JNTU Hyderabad, Bachelor of Technology (Computer Science)
Keywords: csharp continuous integration continuous deployment artificial intelligence machine learning javascript business intelligence sthree database active directory rlang trade national New York Tennessee Texas

To remove this resume please click here or send an email from [email protected] to [email protected] with subject as "delete" (without inverted commas)
[email protected];7533
Enter the captcha code and we will send and email at [email protected]
with a link to edit / delete this resume
Captcha Image: