Home

Yashwanth - Data Scientist & AI Engineer
[email protected]
Location: Iselin, New Jersey, USA
Relocation: Open
Visa:
Resume file: Yashwanth_Kotha_Resume_1782748586724.docx
Please check the file(s) for viruses. Files are checked manually and then made available for download.
PROFESSIONAL SUMMARY

Data Scientist and AI Engineer with 8 plus years of progressive experience across data analytics, data engineering, machine learning, deep learning, and generative AI, with production deployments across healthcare, insurance, and banking sectors.
Skilled in Advanced Regression Modeling, Time Series Analysis, Statistical Testing, Correlation, Multivariate Analysis, Forecasting, and application of Statistical Concepts to solve complex business problems.
Proficient in Data Acquisition, Storage, Analysis, Integration, Predictive Modeling, Logistic Regression, Decision Trees, Factor Analysis, Cluster Analysis, Neural Networks, and advanced statistical econometric techniques.
Adept at writing code in Python, R, SAS, and SQL to manipulate, transform, and analyze large datasets for machine learning model development, data loads, and extracts.
Strong experience in Data Visualization, Exploratory Data Analysis (EDA), Heatmaps, and communicating analytical findings to both technical and non-technical stakeholders using Tableau and Power BI.
Good knowledge and understanding of data mining techniques including classification, clustering, regression, random forests, and deep learning methods such as CNN, RNN, ANN, reinforcement learning, and transfer learning.
Experience implementing machine learning algorithms including Naive Bayes, Random Forests, Decision Trees, Linear and Logistic Regression, SVM, K-Means Clustering, PCA, and Neural Networks using scikit-learn, XGBoost, TensorFlow, and PyTorch.
Strong experience in Big Data technologies including Apache Spark, Hadoop, HDFS, Hive, and Kafka, with hands-on experience processing large structured and unstructured datasets using PySpark and Python.
Experience in all phases of data warehouse development from requirements, analysis, design, development, testing, and post-production support using Oracle, SQL Server, Snowflake, and Redshift.
Skilled in MLOps and cloud deployment using MLflow, Docker, Kubernetes, AWS SageMaker, and AWS Lambda to build scalable inference services, automate retraining workflows, and monitor model performance in production.
Experience in full stack development using React.js, Angular, Spring Boot, and FastAPI to build data-driven web applications and REST API services that expose ML model predictions to end users.
Experienced in Agile and Scrum methodologies, with strong knowledge in all phases of the SDLC including analysis, design, development, testing, and deployment, working closely with cross-functional teams to deliver data science solutions on schedule.
Experienced working in regulated and compliance-driven environments including healthcare, insurance, and government sectors, ensuring data handling, model documentation, and analytical outputs meet HIPAA, GDPR, and state data governance standards.

TECHNICAL SKILLS

Languages and Tools: Python, R, SAS, SQL, Scala, Java, Jupyter Notebook, Git, Agile, SDLC
Machine Learning: Scikit-learn, XGBoost, LightGBM, Random Forest, Logistic Regression, SVM, Naive Bayes, K-Means, PCA, A/B Testing, SHAP, SPSS, Time Series (ARIMA, Prophet)
Deep Learning and NLP: PyTorch, TensorFlow, Keras, CNN, RNN, LSTM, HuggingFace Transformers, BERT, BioBERT, spaCy, NLTK, Transfer Learning
Generative AI: OpenAI GPT-4o, Claude 3, Llama 3, LangChain, RAG Pipelines, Pinecone, ChromaDB, Prompt Engineering, LoRA Fine-tuning, AWS Bedrock, Azure OpenAI
Data and Cloud: Apache Spark, Kafka, Hadoop, Hive, Airflow, dbt, Snowflake, Redshift, AWS (S3, EC2, SageMaker, Lambda), Azure ML, MLflow, Docker, Kubernetes
Databases and BI: Oracle, PostgreSQL, MySQL, SQL Server, MongoDB, Snowflake, Tableau, Power BI, Excel (PivotTables, VBA, Power Query), SSRS


CERTIFICATIONS

AWS Certified Solutions Architect - Associate
AWS Certified Machine Learning Engineer - Associate
Microsoft Certified: Azure Data Scientist Associate
IBM Data Science Professional Certificate
Oracle Certified Associate, Java SE 8 Programmer
Oracle Cloud Infrastructure 2023 AI Certified Foundations Associate

WORK EXPERIENCE
Regeneron October 2024 - Present
Senior Data Scientist Rensselaer, NY
Analyzed clinical trial data for rare disease and immunology programs using SAS and R, performing statistical testing, survival analysis, and correlation analysis on patient outcome datasets to support regulatory submissions and trial reporting.
Built drug efficacy prediction models using logistic regression and Random Forest on patient biomarker and genomics data, helping research scientists identify patient subgroups most likely to respond to investigational therapies.
Applied time series analysis and regression modeling on longitudinal patient lab values and vital signs to track disease progression over time and flag early signals of adverse drug reactions across active trials.
Used K-Means clustering and PCA on genomics and clinical feature sets to stratify patients into meaningful subgroups for trial enrollment criteria and post-hoc efficacy analysis, supporting precision medicine research goals.
Performed exploratory data analysis and correlation analysis on large clinical and biomarker datasets using Python and Pandas, surfacing key drivers of patient outcomes for scientific and regulatory review.
Built Tableau and Power BI dashboards for clinical operations teams to monitor trial enrollment progress, lab result trends, and patient safety metrics across ongoing Regeneron studies in real time.
Wrote Python and SQL scripts to clean, transform, and integrate clinical trial data from multiple source systems into a unified analysis-ready dataset, reducing manual preparation time for the biostatistics team.
Collaborated with research scientists, bioinformaticians, and clinical teams in Agile working sessions to frame scientific questions as data science problems and deliver findings through clear visualizations and written reports

Innova Solutions Jul 2023 Sep 2024
Lead Data Scientist / AI Engineer Atlanta, GA
Built and deployed Retrieval-Augmented Generation (RAG) pipelines using LangChain, GPT-4o, and Pinecone vector database for client-facing document search and Q&A applications, enabling business users to retrieve accurate answers from large internal knowledge bases and document repositories.
Developed NLP solutions using HuggingFace Transformers and fine-tuned BERT and Llama 3 models to classify, summarize, and extract information from unstructured client documents across multiple engagements in healthcare, retail, and financial services.
Built machine learning models using XGBoost, Random Forest, and scikit-learn for client predictive analytics use cases including churn prediction, demand forecasting, and risk scoring, delivering end-to-end pipelines from data preparation to production deployment.
Implemented prompt engineering strategies and evaluation frameworks for GPT-4o and Claude 3 based applications, iterating on system prompts and few-shot examples to improve output quality and accuracy for client-specific use cases.
Fine-tuned LLMs using LoRA and QLoRA on proprietary client datasets via AWS Bedrock and Azure OpenAI, improving domain-specific answer quality for enterprise Gen AI applications in insurance and healthcare verticals.
Managed the full ML and Gen AI model lifecycle using MLflow for experiment tracking, Docker for containerization, and AWS SageMaker for deployment and monitoring, keeping client production models versioned and auditable.
Built FastAPI REST services to expose Gen AI and ML model outputs to client application platforms, enabling both real-time inference and batch scoring modes with clean, versioned API contracts.
Led a team of data scientists and AI engineers in Agile sprints, conducting sprint planning, code reviews, and client demonstrations to deliver Gen AI solutions on schedule across multiple concurrent client engagements.

New York State Office of Mental Health Sep 2021 - Jun 2023
Senior Data Scientist Albany, NY
Built machine learning models using XGBoost and Random Forest on patient admission records, diagnosis history, and treatment data to predict risk of inpatient psychiatric hospitalization, helping care coordinators prioritize outreach for high-risk individuals in community mental health programs.
Developed NLP pipelines using BERT and spaCy to extract structured clinical information from unstructured psychiatrist notes and mental health assessment forms in the NYS OMH electronic health record system, reducing manual documentation burden for clinical staff.
Applied time series analysis and regression modeling on longitudinal patient treatment records to track mental health episode trends and forecast medication adherence patterns, supporting clinical team planning for long-term care programs.
Used K-Means clustering on patient demographic and clinical data to group populations by mental health risk profile, helping NYS OMH program teams identify high-risk communities for targeted outreach and intervention.
Performed exploratory data analysis and statistical testing on mental health program data across NYS psychiatric centers using Python, identifying outcome trends, geographic disparities, and risk factors to inform program evaluation and policy decisions.
Used MLflow to track and version all model experiments across mental health analytics projects, maintaining reproducible and audit-ready documentation of model performance in compliance with state data governance standards.
Built Power BI dashboards for NYS OMH program managers to monitor patient admission volumes, discharge outcomes, and community program participation rates, replacing static monthly reports with interactive self-service analytics.
Worked with clinical analysts and program coordinators in Agile working sessions to validate model outputs and translate data science findings into actionable recommendations for mental health service delivery.

LS Power Sep 2020 - Aug 2021
Senior Data Analyst Austin, TX
Wrote SQL queries and stored procedures in SQL Server to extract and aggregate generation output, fuel consumption, and equipment performance data from LS Power plant operations databases for daily and weekly operations reporting.
Used Python with Pandas, NumPy, and Matplotlib to analyze historical electricity generation and load data, identifying trends in plant capacity utilization and supporting operations teams in understanding performance patterns.
Applied regression analysis and time series analysis on energy market pricing and generation cost data to support the asset management team in evaluating plant economics and short-term dispatch planning decisions.
Built Tableau and Power BI dashboards to monitor key generation metrics including plant availability, output versus forecast, heat rates, and forced outage rates, replacing manual spreadsheet reports with live self-service analytics.
Performed statistical analysis and hypothesis testing on equipment sensor data to identify correlations between operating conditions and maintenance events, supporting the reliability engineering team in scheduling preventive maintenance.
Created Excel reporting tools using PivotTables, Power Query, and VBA macros to automate weekly fuel cost and generation performance reports delivered to LS Power management and asset teams.
Worked with plant operations engineers and asset managers to understand data needs, validate analytical findings, and present energy performance insights in clear, business-friendly formats.

Chubb Jul 2019 - Aug 2020
Data Analyst Philadelphia, PA
Wrote SQL queries, stored procedures, and functions in SQL Server and Oracle to extract and aggregate policy, premium, and claims data for operational reports and ad hoc analysis requests from underwriting and finance teams.
Used Python with Pandas, NumPy, and Matplotlib to clean, analyze, and visualize large insurance datasets, automating data preparation tasks that were previously handled manually in Excel.
Performed exploratory data analysis on commercial lines claims data, building correlation matrices and distribution plots to identify patterns in loss frequency and severity, sharing findings with the actuarial pricing team.
Applied regression analysis and hypothesis testing on policy and premium data to identify key drivers of profitability and loss ratio trends, supporting data-driven decisions in underwriting strategy.
Built Tableau and Power BI dashboards to track premium growth, claims frequency, loss ratios, and renewal retention rates across commercial and personal lines, giving management real-time visibility into portfolio performance.
Developed a logistic regression churn model using scikit-learn on commercial lines renewal data to flag accounts at risk of non-renewal, supporting the underwriting team in prioritizing retention efforts.
Created Excel reporting tools using PivotTables, Power Query, and VBA macros to automate weekly and monthly regulatory and performance reports, reducing manual reporting effort for the finance and operations teams.

RealPage Jul 2018 - Jun 2019
Software Engineer Richardson, TX
Developed full stack web application features for RealPage property management and leasing platforms using Java, Spring Boot, Hibernate, and Angular, building backend APIs and frontend UI modules used by property managers and leasing teams.
Built and maintained REST API services using Spring Boot to integrate RealPage platform modules with third-party payment processors, background screening vendors, and utility billing systems, handling JSON and XML data formats.
Designed Angular frontend components for property analytics and lease management dashboards, consuming Spring Boot APIs to display real-time occupancy rates, rent roll data, and maintenance request status to property managers.
Wrote SQL Server stored procedures, triggers, and queries for property operations reporting, tenant transaction audit trails, and nightly data aggregation jobs feeding RealPage reporting modules.
Built Java batch processing jobs to aggregate daily rent payment and lease renewal records from multiple property databases into a consolidated reporting schema, supporting financial reporting for property owners.
Wrote unit and integration tests using JUnit and Mockito, integrating the test suite into a Maven CI pipeline to maintain code quality and catch regressions early before QA handoff.
Participated in full Agile SDLC including sprint planning, daily standups, peer code reviews, and retrospectives, working with product managers to deliver platform features on schedule.
Keywords: continuous integration quality analyst artificial intelligence machine learning user interface javascript business intelligence sthree active directory rlang Georgia New York Pennsylvania Texas

To remove this resume please click here or send an email from [email protected] to [email protected] with subject as "delete" (without inverted commas)
[email protected];7501
Enter the captcha code and we will send and email at [email protected]
with a link to edit / delete this resume
Captcha Image: