{"author_link":"\/users\/corc0","author_name":"corc0","author_uid":"corc0","comments":[],"epoch":1526769570,"event":"LD41","format":"md","ldjam_node_id":97178,"likes":21,"metadata":{"p_key":"115536","p_author":"corc0","p_authorkey":"1053449","p_urlkey":"331496","p_title":"What a Drunken Dwarf & Spring Breakers Can Teach Us About Game Jam Ratings","p_cat":"LDJam ","p_event":"LD41","p_time":"1526769570","p_likes":"21","p_comments":"0","p_status":"WAYBACK","us_key":"1053449","us_name":"corc0","us_username":"corc0","event_start":"1524182400","event_key":"72","event_name":"LD 41"},"node":{"_collation":{"body_sanitizer":"TextUtils::SanitizeHTML via existing importer","event":"LD41","removed_author":false},"_superparent":73256,"_trust":3,"author":53449,"body":"## Hank the Drunken Dwarf\n\n![game-of-thrones-real-wine-inside-tyrion.jpg](\/\/\/raw\/9c0\/d\/z\/14ee8.jpg)\n...*Not this one.*\n\nWhen we\u2019re rating an entry, we\u2019re giving other people a sense of how it fares\u2014is it visually, sonically, or mechanically better than some other game? \n\nBut flawed methods can lead to skewed results. Take this example from [the American Association for Public Opinion Research](https:\/\/www.aapor.org\/Education-Resources\/For-Researchers\/Poll-Survey-FAQ\/Bad-Samples.aspx):\n\n> People Magazine asked visitors to its website to vote in an online poll for the Most Beautiful Person of 1998. The ballot included Julia Roberts, Leonardo DiCaprio, Madonna and the other usual suspects. It also allowed write-in nominations. The temptation proved too great for radio bad boy Howard Stern, who advised his listeners to email votes for Hank the Angry, Drunken Dwarf, a Stern sidekick who died in 2001. A small army of online pranksters quickly took up the campaign. Hank swamped the competition, finishing with 230,169 votes, or about 16 times the number who supported DiCaprio, the pretty face whom People declared the fairest of the fair and put on its cover.\n\nFor a variety of reasons, these comparisons can fail us\u2014even with when hundreds of thousands of people join in. And the problems may get worse when you\u2019re drawing from a small population. \n\nPolls that gauge the feelings of ten million people can do so confidently with a random sample of around ten thousand people\u2014less than 1%. But, if we\u2019re talking about a [small, finite population](https:\/\/newonlinecourses.science.psu.edu\/stat414\/node\/264\/)\u2014say, a few thousand developers participating in a game jam\u2014you\u2019ll need a larger percentage of a small group to gauge those feelings with the same accuracy. \n\nA **much** larger percentage, in fact. That 1% can become 10% or even 50%. Is *Horror Date Simulator 6* **really** a five-star game, according to the 1,000 developers participating in *Lucky\u2019s Dreamjam 42*? \n\nAsk a few hundred people.\n\n## So what? What\u2019s at stake? \n\nSo, yeah, that\u2019s too rigorous. Just a year ago, Ludum Dare\u2019s ratings period [was extended](https:\/\/ldjam.com\/events\/ludum-dare\/38\/ludum-dare-dot-com\/short-on-ratings-friday-new-features) because fewer than half of all games met the minimum threshold for scoring\u2014about 20 votes. \n\nI\u2019m going to hazard a guess and say that it\u2019s unlikely we\u2019d see the kind of participation needed to make a higher minimum threshold or a [two-tier ranking system](https:\/\/ldjam.com\/events\/ludum-dare\/41\/super-dashball\/ranking-overhaul-idea-tier-based-2-round-ranking) viable (if we\u2019re trying to provide as many participants as possible with a rating, anyway).\n\nIt\u2019s worth thinking about why we\u2019d want a rating system in the first place. Ludum Dare\u2019s smart balance filter encourages people to play, praise, critique, and promote each other\u2019s games\u2014and pretty successfully, I\u2019d say. If that\u2019s all we\u2019re here for, reforms to the ratings system are really solutions in search of a problem. \n\nWhat\u2019s at stake? Besides a bit of **prestige** (a nice salve when mom asks why you don\u2019t just put those skills to use for Big Company), a really high score encourages people to try your game after the jam ends. And that\u2019s a big deal for those of us who\u2019d like to see games we\u2019re passionate about reach an audience beyond our basement.\n\n## Random samples -> reliable(ish) results\nIf we\u2019re looking to improve the accuracy of a small sample, it helps if that sample is randomly-selected. Here\u2019s another example of a failed poll from [the American Association for Public Opinion Research](https:\/\/www.aapor.org\/Education-Resources\/For-Researchers\/Poll-Survey-FAQ\/Bad-Samples.aspx):\n\n> In March of 2006, the American Medical Association reported disturbing rates of binge drinking and unprotected sex among college women during spring break. The report was based on what the researchers claimed was a survey of \u201ca random sample\u201d of 644 women.The survey results were breathlessly reported on the Today Show, the CBS Early Show, and hundreds of reports followed on local television and radio newscasts\u2026 One problem: The sample was not random. The results were based on only women who volunteered to answer the question as part of an online survey panel. Only about a quarter of these women had ever gone on a spring break trip. \n\nAnyone who\u2019s participated in these game jams can probably identify a few reasons why these samples aren\u2019t quite random. \n\n\n### 1. Self-Promotion\/Self-Selection\n\n\nDevelopers post on [ldjam.com](ldjam.com)\u2019s front page in search of more ratings. Besides side-stepping the smart balance filter, posts like this are a problem for ratings in that they\u2019re most likely to recruit from a group of people who are excited about or interested in that kind of game.\n\n\n### 2. Tit-For-Tat\/Quid-Pro-Quo\n\n\n\u201cI\u2019ve played your game, now play mine\u201d\u2014or \u201cI\u2019ve rated your game highly! Here\u2019s mine\u201d\u2014might give that game a slight bump.\n\n\n### 3. Voting Blocs\n\n\nIf you\u2019re on friendly terms with other, competing participants, you\u2019re likely to give their game a look. Even if this results in only five or ten additional ratings, the ratings for a game that\u2019s just over the minimum threshold can experience substantial skewing.   \n\n## Restricted Ratings: A Modest Proposal\n\nI\u2019m wondering how everyone would feel about something like this: everything\u2014or almost everything\u2014stays as it is. Anyone can play, share, comment on, and promote any game they\u2019d like. \n\n\nBut, you can\u2019t rate just any game\u2014you\u2019ll have a page with 50-100 randomly-selected games, ordered by their smart balance score. That list won\u2019t change until you rate a game from that list. At that point, another randomly-chosen game will be available to rate\u2026 and so on.\n\nThere's a discussion in progress [on randomness and group voting](https:\/\/github.com\/ludumdare\/ludumdare\/issues\/1653) at GitHub, where Ludum Dare's source code is contained, and another on a [\"homework\"](https:\/\/github.com\/ludumdare\/ludumdare\/issues\/954) feature, which rewards users for playing \"less desirable\" (less trendy, not web-based) games  (thank you, @samusoidal). \n\nI like this solution because it makes the most of these small samples and still allows developers to look for feedback from the kinds of players they\u2019d like to make games for.\n\n\nI\u2019ve had fun thinking about this\u2014and I hope you\u2019ve had fun reading it!  \n","comments":11,"comments-timestamp":"2018-05-25T23:42:07Z","created":"2018-05-19T01:34:23Z","files":[],"files-timestamp":0,"id":97178,"love":21,"love-timestamp":"2018-05-23T00:53:16Z","meta":[],"modified":"2018-05-25T23:42:07Z","name":"What a Drunken Dwarf & Spring Breakers Can Teach Us About Game Jam Ratings","node-timestamp":"2018-05-20T00:47:01Z","parent":85741,"parents":[1,5,9,73256,85741],"path":"\/events\/ludum-dare\/41\/retrograde\/what-a-drunken-dwarf-spring-breakers-can-teach-us-about-game-jam-ratings","published":"2018-05-19T22:39:30Z","scope":"public","slug":"what-a-drunken-dwarf-spring-breakers-can-teach-us-about-game-jam-ratings","subsubtype":"","subtype":"","type":"post","version":288106},"node_metadata":{"n_key":"97178","n_urlkey":"331496","n_parent":"85741","n_path":"\/events\/ludum-dare\/41\/retrograde\/what-a-drunken-dwarf-spring-breakers-can-teach-us-about-game-jam-ratings","n_slug":"what-a-drunken-dwarf-spring-brea","n_type":"post","n_subtype":"","n_subsubtype":"","n_author":"53449","n_created":"1526693663","n_modified":"1527291727","n_version":"288106","n_status":"WAYBACK"},"source_url":"https:\/\/ldjam.com\/events\/ludum-dare\/41\/retrograde\/what-a-drunken-dwarf-spring-breakers-can-teach-us-about-game-jam-ratings","text":"## Hank the Drunken Dwarf\n\n![game-of-thrones-real-wine-inside-tyrion.jpg](\/\/\/raw\/9c0\/d\/z\/14ee8.jpg)\n...*Not this one.*\n\nWhen we\u2019re rating an entry, we\u2019re giving other people a sense of how it fares\u2014is it visually, sonically, or mechanically better than some other game? \n\nBut flawed methods can lead to skewed results. Take this example from [the American Association for Public Opinion Research](https:\/\/www.aapor.org\/Education-Resources\/For-Researchers\/Poll-Survey-FAQ\/Bad-Samples.aspx):\n\n> People Magazine asked visitors to its website to vote in an online poll for the Most Beautiful Person of 1998. The ballot included Julia Roberts, Leonardo DiCaprio, Madonna and the other usual suspects. It also allowed write-in nominations. The temptation proved too great for radio bad boy Howard Stern, who advised his listeners to email votes for Hank the Angry, Drunken Dwarf, a Stern sidekick who died in 2001. A small army of online pranksters quickly took up the campaign. Hank swamped the competition, finishing with 230,169 votes, or about 16 times the number who supported DiCaprio, the pretty face whom People declared the fairest of the fair and put on its cover.\n\nFor a variety of reasons, these comparisons can fail us\u2014even with when hundreds of thousands of people join in. And the problems may get worse when you\u2019re drawing from a small population. \n\nPolls that gauge the feelings of ten million people can do so confidently with a random sample of around ten thousand people\u2014less than 1%. But, if we\u2019re talking about a [small, finite population](https:\/\/newonlinecourses.science.psu.edu\/stat414\/node\/264\/)\u2014say, a few thousand developers participating in a game jam\u2014you\u2019ll need a larger percentage of a small group to gauge those feelings with the same accuracy. \n\nA **much** larger percentage, in fact. That 1% can become 10% or even 50%. Is *Horror Date Simulator 6* **really** a five-star game, according to the 1,000 developers participating in *Lucky\u2019s Dreamjam 42*? \n\nAsk a few hundred people.\n\n## So what? What\u2019s at stake? \n\nSo, yeah, that\u2019s too rigorous. Just a year ago, Ludum Dare\u2019s ratings period [was extended](https:\/\/ldjam.com\/events\/ludum-dare\/38\/ludum-dare-dot-com\/short-on-ratings-friday-new-features) because fewer than half of all games met the minimum threshold for scoring\u2014about 20 votes. \n\nI\u2019m going to hazard a guess and say that it\u2019s unlikely we\u2019d see the kind of participation needed to make a higher minimum threshold or a [two-tier ranking system](https:\/\/ldjam.com\/events\/ludum-dare\/41\/super-dashball\/ranking-overhaul-idea-tier-based-2-round-ranking) viable (if we\u2019re trying to provide as many participants as possible with a rating, anyway).\n\nIt\u2019s worth thinking about why we\u2019d want a rating system in the first place. Ludum Dare\u2019s smart balance filter encourages people to play, praise, critique, and promote each other\u2019s games\u2014and pretty successfully, I\u2019d say. If that\u2019s all we\u2019re here for, reforms to the ratings system are really solutions in search of a problem. \n\nWhat\u2019s at stake? Besides a bit of **prestige** (a nice salve when mom asks why you don\u2019t just put those skills to use for Big Company), a really high score encourages people to try your game after the jam ends. And that\u2019s a big deal for those of us who\u2019d like to see games we\u2019re passionate about reach an audience beyond our basement.\n\n## Random samples -> reliable(ish) results\nIf we\u2019re looking to improve the accuracy of a small sample, it helps if that sample is randomly-selected. Here\u2019s another example of a failed poll from [the American Association for Public Opinion Research](https:\/\/www.aapor.org\/Education-Resources\/For-Researchers\/Poll-Survey-FAQ\/Bad-Samples.aspx):\n\n> In March of 2006, the American Medical Association reported disturbing rates of binge drinking and unprotected sex among college women during spring break. The report was based on what the researchers claimed was a survey of \u201ca random sample\u201d of 644 women.The survey results were breathlessly reported on the Today Show, the CBS Early Show, and hundreds of reports followed on local television and radio newscasts\u2026 One problem: The sample was not random. The results were based on only women who volunteered to answer the question as part of an online survey panel. Only about a quarter of these women had ever gone on a spring break trip. \n\nAnyone who\u2019s participated in these game jams can probably identify a few reasons why these samples aren\u2019t quite random. \n\n\n### 1. Self-Promotion\/Self-Selection\n\n\nDevelopers post on [ldjam.com](ldjam.com)\u2019s front page in search of more ratings. Besides side-stepping the smart balance filter, posts like this are a problem for ratings in that they\u2019re most likely to recruit from a group of people who are excited about or interested in that kind of game.\n\n\n### 2. Tit-For-Tat\/Quid-Pro-Quo\n\n\n\u201cI\u2019ve played your game, now play mine\u201d\u2014or \u201cI\u2019ve rated your game highly! Here\u2019s mine\u201d\u2014might give that game a slight bump.\n\n\n### 3. Voting Blocs\n\n\nIf you\u2019re on friendly terms with other, competing participants, you\u2019re likely to give their game a look. Even if this results in only five or ten additional ratings, the ratings for a game that\u2019s just over the minimum threshold can experience substantial skewing.   \n\n## Restricted Ratings: A Modest Proposal\n\nI\u2019m wondering how everyone would feel about something like this: everything\u2014or almost everything\u2014stays as it is. Anyone can play, share, comment on, and promote any game they\u2019d like. \n\n\nBut, you can\u2019t rate just any game\u2014you\u2019ll have a page with 50-100 randomly-selected games, ordered by their smart balance score. That list won\u2019t change until you rate a game from that list. At that point, another randomly-chosen game will be available to rate\u2026 and so on.\n\nThere's a discussion in progress [on randomness and group voting](https:\/\/github.com\/ludumdare\/ludumdare\/issues\/1653) at GitHub, where Ludum Dare's source code is contained, and another on a [\"homework\"](https:\/\/github.com\/ludumdare\/ludumdare\/issues\/954) feature, which rewards users for playing \"less desirable\" (less trendy, not web-based) games  (thank you, @samusoidal). \n\nI like this solution because it makes the most of these small samples and still allows developers to look for feedback from the kinds of players they\u2019d like to make games for.\n\n\nI\u2019ve had fun thinking about this\u2014and I hope you\u2019ve had fun reading it!  \n","title":"What a Drunken Dwarf & Spring Breakers Can Teach Us About Game Jam Ratings","wayback_source":[]}