Half-star rankings? Why not make 10 louder?

Ludum Dare 40 is this weekend.

For this jam, they've made a minor change to the ratings system: to allow half-star increments. I don't see the point of this. Apparently people were torn between giving a 4 and 5, so we're giving them 4.5. I'm sure there was popular demand for it, but I still don't see what benefit this will bring.

Simple is better

These days, I like simple, and I like quick.

If I asked you to pick your favorite color crayon out of a box of 8 crayons, you're going to have an easier time doing that than if I give you a box of 64 crayons. 

With the 8-color box, you pick whichever color is closest to your favorite, and you don't think about in-between shades all that much, because they're not there, and because every color is very distinct from the others.

8crayons.jpeg It's blue, isn't it?

With the 64-color box, you still pick whichever color is closest to your favorite, and maybe now there is a color that is closer to your actual favorite, but the trade-off is that you now have to do a lot more thinking, because there are a few shades that are really close, and it's harder to choose between them all.

64crayons.jpg But which blue?

I'm willing to bet you could rank an 8-pack of crayons in order of favorite to least favorite color in about a minute, maybe less.  With the 64 pack, if you can do it at all, it'll probably take you way more than 8 times as long to sort 8 times as many colors.

Numbers and perception

For theme voting, we get a 3-position scale:

-1 | 0 | 1

And it works. It's fine. You either like it, or you don't. Make a decision and move on.

You wouldn't "improve" this system by adding a "kindof like it" and "kindof don't like it". There would be no point, other than to muddle the theme voting and make it more complicated.

For game rankings, LD has always (as far as I know) used a 5-star system. Which has just been changed to allow half-star ratings. So, if you count the number of actual steps, you don't get 6 anymore, but now 11:

Then - 6 gradations: n/a | ★ | ★★ | ★★★ | ★★★★ | ★★★★★

Now - 11 gradations: n/a | ½ | ★ | ★½ | ★★ | ★★½ | ★★★ | ★★★½ | ★★★★ | ★★★★½ | ★★★★★

My thinking is that from a mathematical standpoint, all we've done is switch from a 5-point scale to a 10-point scale.  Calling them "half stars" is just semantics.

But, from a psychological standpoint, there's more to it than that.

With the three-position scale, -1 implies that you do not like it, 0 implies that you are neutral, and 1 implies that you like it.

With a 5-star scale, many people assume that 3-stars is "neutral" or "OK", while 1- and 2- star rankings are "bad" and 4- and 5- star ratings are "good".  People who think this way tend to use 3 stars as their default rating, and "reward" good things with a 4- or 5- star rating, and "punish" bad things with a 1- or 2- star rating.  The result of this is that the 1-2 ratings are seldom used.  Some raters will also take the position that 5-stars should be  ultra-rare, as it is a signifier of the highest standards, of perfection, and so they're really constrained to using 3 and 4 star ratings for almost everything, the occasional rare 5-star for something that is truly outsanding, and withholding 1-2 star ratings most of the time.

The implication is that there is a spectrum of quality, and that the scale represents the entire spectrum, but that the vast majority of things are somewhere in the middle of the spectrum and thus the middle of the scale.  The ends of the scale don't get used, other than to show the range of the scale. And this, I conjecture, is why some people feel like they need half-stars: having sacrificed the bottom end and sometimes the very top of the scale, they need to slice up what's left in order to have "enough" meaningful gradations between 3 and 4.5 stars.

This is similar to the loudness wars in audio formats, where dynamic range is sacrificed in order to make every song as loud as all the others, because being loud is what gets you noticed.

So, to restore the lost dynamic range, people want to cram in extra numbers somehow, e.g. to add in half-stars.

I would prefer to see a ratings system with a wide dynamic range, so I recommend using the entire scale.  Hand out those 1- and 2- star ratings.  Don't feel bad about it.  1 > 0, so 1 star is better than nothing.  It's like a score: more is better, but you can still win with just one point.

The thing is, it's difficult to remove the psychological stigma associated with 1-2 stars.  Like fighting against the loudness wars, advocating that people change the way we think about the meaning of <3 stars, so that we can use the entire available spectrum is an uphill battle.  It's easier to just slice up the stars and let people have their 6-point range (5 stars plus n/a), and label the points n/a, 3, 3.5, 4, 4.5, 5.

But it also feels dumb.  I keep thinking about the scene from Spinal Tap, "Why not make 10 louder?"

Instead of using fractions, why not switch to a 10-point scale?  We could even eliminate the stigma scores, and have still have a 5-point scale, but labeled so that they're values from the 10-point scale that people would actually use:

6 (3 stars) | 7 (3.5 stars) | 8 (4 stars) | 9 (4.5 stars) | |10 (5 stars)

I look at it like a way to fool people into having the 5 points they really need without having to think that they're handing out "negative" ratings when they rate something with only 1-2 stars.

5-point system in 5-point scale.png 5-point system in 10-point inflated scale.png

Two graphs: in one, the same scores are inflated with 5 "vanity" or "participation trophy" points, in order to signify the reluctance of judges to give "negative" ratings by rating below 5, which is very common in 10-point ratings systems. We feel better about handing out those scores than we do the lower scores, but the difference between a 10 and a 7 looks like less than the difference between a 5 and a 2, even though the difference in both cases is 3.

The only problem with that approach, as I see it, is that you can keep slicing stars arbitrarily thin, and eventually you end up with a big box of crayons problem.  We could use a 5-star scale, a 1-10 scale, or a 100 point scale.  But it's harder to decide if something is closer to 80% or 82% than it is to decide if something is closer to 80% or 90%. And you don't actually gain anything, just more complexity and more muddle.

I can tell quickly if I want to give something a 4-star or a 5-star rating, but if I start asking myself does it deserve 4 stars, or 4.5 stars, I have to slow down and think about it too much.  As a result, I think I would spend more time on fewer games.

What was the point of all this, again? Oh yeah, to rate games

Ranking things like games is subjective, full of gut instinct and arbitrary taste. But with quantifying a gut reaction, there's a point at which additional granularity isn't really useful, and actually becomes counter-productive, because it makes people more indecisive.

It's already somewhat laughable to think that these are things that can be quantified, but we do it, because you can't really compare and rank things otherwise.  Picking a number is way quicker, and easier, than writing a paragraph or an essay.  And it's way more feasible compare numbers than it is to convey subjective full-text reviews.  That's the reason we use number ratings.  Like, that's literally the only reason to do it.

I want to rate quickly and decisively and move on, so I can play and rate the most games. Precision and accuracy aren't the most important thing, here.

Consistency is far more important.  If I can rate a lot of games, and apply my personal, subjective approach to rating consistently across all of them, it's possible to compare those scores against each other.  This is true even if "my three" is very different from "your three" -- as long as we're both consistent in how we rate.  We get precision and accuracy from the weight of averages.  If you can get a few dozen reviews of your game, average the ratings numbers, you get pretty good confidence that the numbers are about right, even if you don't really know what a "3" or a "4" actually means, and even if any N individual judges can't agree on what "3" or "4" means.  The math of averages end up giving us decimals anyway.

Conclusion

To get better ratings(1), we need more ratings, not more gradations. This is true even when the quality of individual reviews is low, (so long as the quality of reviews is consistent from game to game by any particular reviewer.) 

Thus, simplifying and streamlining the process of rating games, so that we can rate more games, would be better than giving individuals the means to assign ratings with finer granularity.

__

(1) I'm talking about better quantitative ratings, obviously; better qualitative ratings would be nice, too -- but that's another topic.