About Me

Data rarely tells a clean story. It’s buried in noise, nulls, and inconsistent formatting. I use Python to clean it up and SQL to gig into it, until the “what happened” turns into a clear “why” and an actual next step.

Outside of queries, I care about the same thing in my day-to-day: efficient, low-waste ways of doing things, whether that’s a workflow or a trail I’m hiking or a wave I’m out chasing.


Technical Skills

  • Languages: Python, SQL
  • Libraries & Frameworks: Pandas, NumPy, Scikit-Learn, Matplotlib, Seaborn, SQLAlchemy
  • Data Tools: SQLite, DBeaver, Jupyter, Tableau, Git
  • Core Competencies: Data Cleaning, Exploratory Data Analysis (EDA), Predictive Modeling, Feature Engineering, Root Cause Analysis

Medicare Part D Fraud Detection

Tech Stack: Python (Pandas, NumPy, Matplotlib, Seaborn, Folium, Scikit-Learn), Statistical Anomaly Detection

Overview: An end-to-end fraud detection pipeline analyzing 1.8 million Medicare Part D prescription records across 62,301 Texas providers to identify statistically anomalous billing patterns consistent with fraud, waste, and abuse.

Key Highlights:

  • Engineered three fraud signal features and applied Z-score peer benchmarking to compare each provider against their specialty group, flagging 100 high-risk providers representing $274.5 million in potentially suspicious Medicare spending.
  • Identified 16 ophthalmologists prescribing opioids, a clear specialty mismatch with no legitimate clinical justification, using targeted drug-specialty cross-analysis.
  • Built an interactive geographic map of flagged providers across Texas using Folium, with color-coded risk tiers and clickable provider detail popups.

Apple Global Sales Analysis

Tech Stack: SQL, Python (Pandas, Seaborn), Exploratory Data Analysis

Overview: An end-to-end exploratory data analysis of synthetic data regarding historical sales data to identify core revenue drivers and regional market trends.

Key Highlights:

  • Analyzed complex financial datasets to isolate product lifecycle plateaus and growth markets.
  • Translated raw sales metrics into actionable business intelligence and clear visual dashboards.

Steam Hidden Gems: Machine Learning Classifier

Tech Stack: Python (Pandas, Scikit-Learn, Matplotlib, Seaborn), Feature Engineering

Overview: A predictive machine learning pipeline analyzing 10,000 Steam games to identify “hidden gems” by engineering a custom engagement metric and building a binary classifier.

Key Highlights:

  • Engineered a custom quality metric from raw review data to isolate a target class of undiscovered games, utilizing balanced class weights to handle a strict 90/10 dataset imbalance.
  • Identified and successfully mitigated critical data leakage during the feature selection phase, ensuring the model relied purely on independent commercial variables.
  • Evaluated both Random Forest and Logistic Regression models, ultimately proving that linear models generalized better and that commercial strategy (price, discounts) is a statistically weak predictor of community-driven game quality.

Credentials

  • B.S. Data Analytics — Western Governors University, Dec 2025
  • AWS Cloud Practitioner
  • COMPTIA Data+
  • COMPTIA Project+