Yelp I: What actually predicts restaurant check-ins (hint: not stars)

Star ratings are the most visible number on Yelp. So when our six-person research team set out to model restaurant check-ins using Yelp's public dataset of 20,000+ restaurants, we expected stars to matter.

They didn't.

First, some honesty about the setup. We originally wanted check-ins to measure how many people visit a restaurant. Then we found out Yelp runs marketing incentives for checking in, which contaminates that measure. So we adjusted the goal: check-ins became a proxy for customers willing to publicly announce where they eat. When the data undermines your plan, you change the plan. Not the data.

With around 40 variables to sort through, we used stepwise and hierarchical methods to find the ones that mattered. Several variables were heavily skewed and had to be log-transformed first to meet the assumptions of regression. What survived was a four-variable model that explained 82 percent of the variation in check-ins:

Number of reviews. Whether the restaurant was still open. Price range. Delivery.

Reviews dominated. A 10 percent increase in reviews predicted a 12 percent increase in check-ins. Cheaper restaurants got more check-ins than pricier ones. And stars never made the cut, even when we tested interactions between stars and reviews, and a Bayesian rating we built to account for review volume. None were significant.

The takeaway: on Yelp, the crowd size matters more than the crowd's verdict. People check in where other people already are.

This was the first of three models. Next we asked whether the answer changes depending on the city. That's Yelp II.

Previous
Previous

Yelp II: Same model, three cities, three different answers

Next
Next

Experience: Merging and Data Cleaning - When Inbuilt Functions are Impossible