What a Drunken Dwarf & Spring Breakers Can Teach Us About Game Jam Ratings

Hank the Drunken Dwarf

game-of-thrones-real-wine-inside-tyrion.jpg ...Not this one.

When we’re rating an entry, we’re giving other people a sense of how it fares—is it visually, sonically, or mechanically better than some other game?

But flawed methods can lead to skewed results. Take this example from the American Association for Public Opinion Research:

People Magazine asked visitors to its website to vote in an online poll for the Most Beautiful Person of 1998. The ballot included Julia Roberts, Leonardo DiCaprio, Madonna and the other usual suspects. It also allowed write-in nominations. The temptation proved too great for radio bad boy Howard Stern, who advised his listeners to email votes for Hank the Angry, Drunken Dwarf, a Stern sidekick who died in 2001. A small army of online pranksters quickly took up the campaign. Hank swamped the competition, finishing with 230,169 votes, or about 16 times the number who supported DiCaprio, the pretty face whom People declared the fairest of the fair and put on its cover.

For a variety of reasons, these comparisons can fail us—even with when hundreds of thousands of people join in. And the problems may get worse when you’re drawing from a small population.

Polls that gauge the feelings of ten million people can do so confidently with a random sample of around ten thousand people—less than 1%. But, if we’re talking about a small, finite population—say, a few thousand developers participating in a game jam—you’ll need a larger percentage of a small group to gauge those feelings with the same accuracy.

A much larger percentage, in fact. That 1% can become 10% or even 50%. Is Horror Date Simulator 6 really a five-star game, according to the 1,000 developers participating in Lucky’s Dreamjam 42?

Ask a few hundred people.

So what? What’s at stake?

So, yeah, that’s too rigorous. Just a year ago, Ludum Dare’s ratings period was extended because fewer than half of all games met the minimum threshold for scoring—about 20 votes.

I’m going to hazard a guess and say that it’s unlikely we’d see the kind of participation needed to make a higher minimum threshold or a two-tier ranking system viable (if we’re trying to provide as many participants as possible with a rating, anyway).

It’s worth thinking about why we’d want a rating system in the first place. Ludum Dare’s smart balance filter encourages people to play, praise, critique, and promote each other’s games—and pretty successfully, I’d say. If that’s all we’re here for, reforms to the ratings system are really solutions in search of a problem.

What’s at stake? Besides a bit of prestige (a nice salve when mom asks why you don’t just put those skills to use for Big Company), a really high score encourages people to try your game after the jam ends. And that’s a big deal for those of us who’d like to see games we’re passionate about reach an audience beyond our basement.

Random samples -> reliable(ish) results

If we’re looking to improve the accuracy of a small sample, it helps if that sample is randomly-selected. Here’s another example of a failed poll from the American Association for Public Opinion Research:

In March of 2006, the American Medical Association reported disturbing rates of binge drinking and unprotected sex among college women during spring break. The report was based on what the researchers claimed was a survey of “a random sample” of 644 women.The survey results were breathlessly reported on the Today Show, the CBS Early Show, and hundreds of reports followed on local television and radio newscasts… One problem: The sample was not random. The results were based on only women who volunteered to answer the question as part of an online survey panel. Only about a quarter of these women had ever gone on a spring break trip.

Anyone who’s participated in these game jams can probably identify a few reasons why these samples aren’t quite random.

1. Self-Promotion/Self-Selection

Developers post on ldjam.com’s front page in search of more ratings. Besides side-stepping the smart balance filter, posts like this are a problem for ratings in that they’re most likely to recruit from a group of people who are excited about or interested in that kind of game.

2. Tit-For-Tat/Quid-Pro-Quo

“I’ve played your game, now play mine”—or “I’ve rated your game highly! Here’s mine”—might give that game a slight bump.

3. Voting Blocs

If you’re on friendly terms with other, competing participants, you’re likely to give their game a look. Even if this results in only five or ten additional ratings, the ratings for a game that’s just over the minimum threshold can experience substantial skewing.

Restricted Ratings: A Modest Proposal

I’m wondering how everyone would feel about something like this: everything—or almost everything—stays as it is. Anyone can play, share, comment on, and promote any game they’d like.

But, you can’t rate just any game—you’ll have a page with 50-100 randomly-selected games, ordered by their smart balance score. That list won’t change until you rate a game from that list. At that point, another randomly-chosen game will be available to rate… and so on.

There's a discussion in progress on randomness and group voting at GitHub, where Ludum Dare's source code is contained, and another on a "homework" feature, which rewards users for playing "less desirable" (less trendy, not web-based) games (thank you, @samusoidal).

I like this solution because it makes the most of these small samples and still allows developers to look for feedback from the kinds of players they’d like to make games for.

I’ve had fun thinking about this—and I hope you’ve had fun reading it!