About Me
Data rarely tells a clean story. It’s buried in noise, nulls, and inconsistent formatting. I use Python to clean it up and SQL to gig into it, until the “what happened” turns into a clear “why” and an actual next step.
Outside of queries, I care about the same thing in my day-to-day: efficient, low-waste ways of doing things, whether that’s a workflow or a trail I’m hiking or a wave I’m out chasing.
Technical Skills
- Languages: Python, SQL
- Libraries & Frameworks: Pandas, NumPy, Scikit-Learn, Matplotlib, Seaborn, SQLAlchemy
- Data Tools: SQLite, DBeaver, Jupyter, Tableau, Git
- Core Competencies: Data Cleaning, Exploratory Data Analysis (EDA), Predictive Modeling, Feature Engineering, Root Cause Analysis
Featured Projects
Medicare Part D Fraud Detection
Tech Stack: Python (Pandas, NumPy, Matplotlib, Seaborn, Folium, Scikit-Learn), Statistical Anomaly Detection
Overview: An end-to-end fraud detection pipeline analyzing 1.8 million Medicare Part D prescription records across 62,301 Texas providers to identify statistically anomalous billing patterns consistent with fraud, waste, and abuse.
Key Highlights:
- Engineered three fraud signal features and applied Z-score peer benchmarking to compare each provider against their specialty group, flagging 100 high-risk providers representing $274.5 million in potentially suspicious Medicare spending.
- Identified 16 ophthalmologists prescribing opioids, a clear specialty mismatch with no legitimate clinical justification, using targeted drug-specialty cross-analysis.
- Built an interactive geographic map of flagged providers across Texas using Folium, with color-coded risk tiers and clickable provider detail popups.
Apple Global Sales Analysis
Tech Stack: SQL, Python (Pandas, Seaborn), Exploratory Data Analysis
Overview: An end-to-end exploratory data analysis of synthetic data regarding historical sales data to identify core revenue drivers and regional market trends.
Key Highlights:
- Analyzed complex financial datasets to isolate product lifecycle plateaus and growth markets.
- Translated raw sales metrics into actionable business intelligence and clear visual dashboards.
Steam Hidden Gems: Machine Learning Classifier
Tech Stack: Python (Pandas, Scikit-Learn, Matplotlib, Seaborn), Feature Engineering
Overview: A predictive machine learning pipeline analyzing 10,000 Steam games to identify “hidden gems” by engineering a custom engagement metric and building a binary classifier.
Key Highlights:
- Engineered a custom quality metric from raw review data to isolate a target class of undiscovered games, utilizing balanced class weights to handle a strict 90/10 dataset imbalance.
- Identified and successfully mitigated critical data leakage during the feature selection phase, ensuring the model relied purely on independent commercial variables.
- Evaluated both Random Forest and Logistic Regression models, ultimately proving that linear models generalized better and that commercial strategy (price, discounts) is a statistically weak predictor of community-driven game quality.
Credentials
- B.S. Data Analytics — Western Governors University, Dec 2025
- AWS Cloud Practitioner
- COMPTIA Data+
- COMPTIA Project+