Portfolio
Predicting Loan Defaults with R
Can income and loan amount predict default? I built and compared two classification models in R, kNN and Random Forest, on real lending data from Kaggle. Both hit around 85% accuracy, and the project digs into why that number alone doesn't tell the whole story.
Yelp III: Are cities really different, or do they just have more reviews?
Las Vegas restaurants dominated Montreal's on check-ins. Then we controlled for review count and the gap shrank by two thirds. Much of what looked like city culture was review volume in disguise. A lesson in why the right covariate changes everything.
Yelp II: Same model, three cities, three different answers
Do restaurants behave differently in different cities? We compared 1,000 restaurants each in Las Vegas, Montreal, and Charlotte. All three differed significantly. But ANOVA can only confirm a difference exists. It can't explain it. That took one more model.
Yelp I: What actually predicts restaurant check-ins (hint: not stars)
We modeled check-ins for 20,000+ restaurants in the Yelp dataset. Four variables explained 82 percent of the variation. Star ratings weren't one of them. Review volume mattered most: 10 percent more reviews predicted 12 percent more check-ins.
Experience: Merging and Data Cleaning - When Inbuilt Functions are Impossible
Spreadsheet duplicate-delete functions are blunt instruments. Run them on two closely related sheets and you lose real data along with the duplicates. We needed to merge two 10,000-row spreadsheets without that risk. Carefully using find-and-replace with conditional formatting, we merged both sheets and cleared every true duplicate in a single workday.
Experience: Government Reports and Dissertations - When Brilliant Researchers Can’t Explain What They Do
A technical team hired us to write the commercialization plan for their government proposal. Their first draft was so dense with jargon we couldn’t follow it. If we couldn't follow it, neither could the non-expert reviewers scoring it. We translated the science into plain language while meeting the agency's strict formatting rules. The fully compliant proposal went out on time, with compliments from the team.
Data Analysis: New Listings by Area
A real estate client wanted to forecast new listings by area, using reports from an independent housing statistics agency. We built the dataset in Excel with standardized region names, so no listing got miscounted from a typo or naming mismatch. Then we visualized listings by region and month. The charts revealed a market potentially more profitable than San Francisco, one our client hadn't been targeting.