{"author_link":"\/users\/david-boeger","author_name":"David Boeger","author_uid":"david-boeger","comments":[],"epoch":1526678319,"event":"LD41","format":"md","ldjam_node_id":97175,"likes":5,"metadata":{"p_key":"115535","p_author":"David Boeger","p_authorkey":"1079774","p_urlkey":"331495","p_title":"Another Suggestion for \"Improved\" Rankings","p_cat":"LDJam ","p_event":"LD41","p_time":"1526678319","p_likes":"5","p_comments":"0","p_status":"WAYBACK","us_key":"1079774","us_name":"David Boeger","us_username":"david-boeger","event_start":"1524182400","event_key":"72","event_name":"LD 41"},"node":{"_collation":{"body_sanitizer":"TextUtils::SanitizeHTML via existing importer","event":"LD41","removed_author":false},"_superparent":73256,"_trust":1,"author":79774,"body":"Disclaimer: a friend has respectfully asked me to no longer comment on a set of related issues, and so out of respect for him, this is purely a technical discussion on how to improve the ranking system.\n\nHere's the problem I see: despite people going into the competition fully understanding that varying sample sizes are not good for the objectivity of the relative rankings, despite the fact that these people are fully aware in their logical minds that the current rankings are an imperfect science which says little to nothing about one game versus another, despite the fact that people recognize that by setting the rating threshold required for ranking too high, it would exclude many busy people who simply don't have enough time to generate that much buzz around their game, there's something deeply unsatisfying about seeing a game with very few ratings beat a game with very many ratings. And this is for good reason. It's not only an emotional reaction, but there is a logical justification: because the system is very easy to abuse, and nobody wants to feel cheated after dedicating an entire weekend developing free entertainment for others. So here's my take on a solution.\n\nI think in order to create a more objective ranking system, you absolutely must create an objective target goal. The most trivial example would be having a single authoritative user be responsible for all ratings. Obviously, that's impractical with 3k+ submissions, but it illustrates the point. The objective target goal is now to GET THE HIGHEST RATING WITH THIS PERSON. It is directly measurable by the ratings that person gives. And believe it or not, despite shifting the focus from some nebulous concept like being the best game overall to something tangential like BEING THE GAME THIS PERSON RATES THE HIGHEST, the fact that it is DIRECTLY MEASURABLE actually allows you to draw MORE INFORMATIVE CONCLUSIONS over time. If your submissions win multiple events in a row, you can actually now make a solid argument in favor of your games being the best AT ACHIEVING THE STATED GOAL. You now have something that is directly comparable between submissions and across events.\n\nUnfortunately, solutions like that all have one major flaw in the context of LD, which is that they directly contradict some of the stated purposes of LD. Namely, that it's a community event, developers want a randomized sample of feedback from the general population, and nobody should be mandated to participate in any capacity. We'll call these kinds of solutions CLASS A solutions. CLASS A solutions compromise on those stated goals in some way or another. Most of the solutions I've seen proposed following this event so far are CLASS A solutions, because they involve some sort of authoritative voter round, where trusted voters, streamers, etc. can, to varying degrees, OVERRIDE ratings given by the general community in the initial voting round. I get it, they're not just completely throwing away the early results. But they are being granted more power to some degree.\n\nOn the other hand, we have what we will call the CLASS B solutions. Class B solutions are more desirable because they do not compromise on the goals of LD whatsoever (or at least do so very minimally, but that's up to you guys to decide). CLASS B solutions cannot weight votes, and they cannot mandate participation in any capacity. They can SUGGEST or ENCOURAGE participation in order to generate more ratings and hopefully reduce controversy or feelings of suspicion. An example of a CLASS B solution I've seen proposed is one that was brought up a few events back, where there would be something like a homework queue that rewards participants with extra Coolness in exchange for rating a suggested subset of submissions. This is certainly a cool idea, and does encourage participants to rate more games in general, but it DOES NOT DIRECTLY ADDRESS THE CORE COMPLAINT, which is that CATEGORY WINNERS SHOULD HAVE RELATIVELY HIGH COUNTS OF RATINGS.\n\nSo to summarize what I've said up until this point, what we really want is a CLASS B solution which stands a chance to address the core issue of winners with few ratings, without going as far as a CLASS A solution would to reduce community voting power or mandate participation. This is what I've come up with:\n\nYou can think of it as 2 voting rounds, similar to the CLASS A proposals so far, but technically, there is only 1 voting round. They're more like 2 voting phases. Participants only get 1 vote, as they do now, and votes are not weighted, so everyone's vote is equal, as it is now. Voter power is never once reduced in the slightest. The only difference is that the intermediate results are made public prior to the 2nd phase. Then, the final results are tallied after both phases are complete. Again, I want to make very clear, this is almost exactly the same as the way it is now. Each participant gets to vote on each entry exactly once, and that 1 vote is equal. The ONLY difference is that for the 2nd portion of the voting period, the intermediate results are made public. By doing this, you allow people who haven't voted on a game near the top of the rankings to vote on it, thereby increasing the sample size of the winners.\n\nThink of it like getting dressed in the morning. Sometimes, you're in a rush, and you quickly slip into your pants, but your underwear doesn't feel quite right, so you adjust. The 2nd phase is the adjustment. If the intermediate results are released and people feel there is a game which doesn't belong at the top, they can add their vote, thereby increasing the confidence in the overall ranking. If they choose not to vote, and the game stays where it's at, that's okay too. The game isn't automatically disqualified, its fans don't have their votes questioned, and the game wins, because people had every chance to vote against it but decided not to. If you wanted, you could even have a confidence threshold, where you bump a game up or down a slot for having a massive gap in number of ratings compared to those around it, but honestly, at that point, we're just splitting split ends of hair. You could also prohibit the top 100 in a category from voting on each other's games as well, just to eliminate spiteful voting between them, but it's not really worth doing considering they each get only 1 vote and could just as easily make a bunch of fake accounts.\n\nSo far, this is the only CLASS B system I've been able to come up with in my head that comes anywhere close to protecting the values of LD while potentially addressing this recurring issue. Unfortunately, it is not without its flaws. It does absolutely nothing to prevent intentional abuse, just like the current system. Short of implementing a strictly managed CLASS A solution, I don't think there's any reasonable way to prevent intentional cheating. We're better off trying to find ways to utilize the resources we have, namely thousands of participants, to generate more honest ratings to combat the cheaters. It's also kind of sad to think that a submission could be a category leader going into phase 2, and then be edged out by a competitor generating late buzz. I mean, sure, the results are not locked in until after phase 2, but still, that's a bummer.\n\nIf you've made it this far, I apologize for the long description. The system itself is actually quite simple, but understanding the reasons behind it was quite the brain exercise. Please leave your comments below.\n\nADDENDUM: And for what it's worth, the reason I put \"improved\" in quotes in the post's subject line is that I personally do not believe this system will improve the objectivity of the rankings overall. I believe the system will be just as flawed, maybe even more so, because of the fundamental principle that I mentioned earlier about not having an objective target goal.\n\nThink of it this way. Imagine attending an NBA Finals match to determine the winning team of the basketball season. There is a very objective goal. Whichever team scores the most points in the allotted time wins. That may not determine the best team overall, but it does determine the winner of that particular contest according to a strict set of rules. Now, imagine if instead of tracking points scored, the winner was determined by the volume of cheers from the stands. At the end of the game, the referee asks the fans of each team to take turns cheering as loudly as possible for who they thought won (I know there are instruments to objectively measure noise volume, but for the purposes of this example, pretend the ref is doing it by ear, and he's deaf on one side and spinning in a circle). Whichever one he thinks sounds the loudest wins. This system I've proposed is a little bit like fitting more people in the stadium to improve the confidence in this win-by-sound system. I'm just not convinced it's actually better in such a noisy (no pun intended) system.\n\nYou have to consider the very real possibility that someone in the middle tiers has a strong entry and wants honest feedback, but they can't generate buzz because everyone is too busy rating top tier games in phase 2 which don't even necessarily appeal to them. You've actually shifted the balance of exposure negatively, all for the sake of reducing suspicion, and the value of the change in objectivity is questionable at best. This \"fix\" is nothing more than a placebo to ease the suspicions of people who don't actually understand why community votes are not objective. In fact, that person who can't generate buzz may then be encouraged to seek out friendly votes from people they've curried favor with, introducing even more bias and muddying the middle tiers. Butterfly effect. Everything affects everything else.\n\nIf my suggestion were implemented and I was offered to participate in it, I'd have to politely decline. Because again, I recognize a CLASS B system will never have anywhere near the level of objectivity that a proper CLASS A system will. You need a CLASS A system for this to be a true competition. A CLASS B system is something you use to judge a silly informal game of hop-scotch on a kindergarten playground. A CLASS A system is needed if you really care about the competitive aspect of something. It's up to you, the LD community, to decide where your priorities are.","comments":10,"comments-timestamp":"2018-05-23T21:07:28Z","created":"2018-05-18T20:14:04Z","files":[],"files-timestamp":0,"id":97175,"love":5,"love-timestamp":"2018-05-19T22:16:24Z","meta":[],"modified":"2018-05-23T21:07:28Z","name":"Another Suggestion for \"Improved\" Rankings","node-timestamp":"2018-05-18T22:42:31Z","parent":79778,"parents":[1,5,9,73256,79778],"path":"\/events\/ludum-dare\/41\/war-maker\/another-suggestion-for-improved-rankings","published":"2018-05-18T21:18:39Z","scope":"public","slug":"another-suggestion-for-improved-rankings","subsubtype":"","subtype":"","type":"post","version":288084},"node_metadata":{"n_key":"97175","n_urlkey":"331495","n_parent":"79778","n_path":"\/events\/ludum-dare\/41\/war-maker\/another-suggestion-for-improved-rankings","n_slug":"another-suggestion-for-improved-","n_type":"post","n_subtype":"","n_subsubtype":"","n_author":"79774","n_created":"1526674444","n_modified":"1527109648","n_version":"288084","n_status":"WAYBACK"},"source_url":"https:\/\/ldjam.com\/events\/ludum-dare\/41\/war-maker\/another-suggestion-for-improved-rankings","text":"Disclaimer: a friend has respectfully asked me to no longer comment on a set of related issues, and so out of respect for him, this is purely a technical discussion on how to improve the ranking system.\n\nHere's the problem I see: despite people going into the competition fully understanding that varying sample sizes are not good for the objectivity of the relative rankings, despite the fact that these people are fully aware in their logical minds that the current rankings are an imperfect science which says little to nothing about one game versus another, despite the fact that people recognize that by setting the rating threshold required for ranking too high, it would exclude many busy people who simply don't have enough time to generate that much buzz around their game, there's something deeply unsatisfying about seeing a game with very few ratings beat a game with very many ratings. And this is for good reason. It's not only an emotional reaction, but there is a logical justification: because the system is very easy to abuse, and nobody wants to feel cheated after dedicating an entire weekend developing free entertainment for others. So here's my take on a solution.\n\nI think in order to create a more objective ranking system, you absolutely must create an objective target goal. The most trivial example would be having a single authoritative user be responsible for all ratings. Obviously, that's impractical with 3k+ submissions, but it illustrates the point. The objective target goal is now to GET THE HIGHEST RATING WITH THIS PERSON. It is directly measurable by the ratings that person gives. And believe it or not, despite shifting the focus from some nebulous concept like being the best game overall to something tangential like BEING THE GAME THIS PERSON RATES THE HIGHEST, the fact that it is DIRECTLY MEASURABLE actually allows you to draw MORE INFORMATIVE CONCLUSIONS over time. If your submissions win multiple events in a row, you can actually now make a solid argument in favor of your games being the best AT ACHIEVING THE STATED GOAL. You now have something that is directly comparable between submissions and across events.\n\nUnfortunately, solutions like that all have one major flaw in the context of LD, which is that they directly contradict some of the stated purposes of LD. Namely, that it's a community event, developers want a randomized sample of feedback from the general population, and nobody should be mandated to participate in any capacity. We'll call these kinds of solutions CLASS A solutions. CLASS A solutions compromise on those stated goals in some way or another. Most of the solutions I've seen proposed following this event so far are CLASS A solutions, because they involve some sort of authoritative voter round, where trusted voters, streamers, etc. can, to varying degrees, OVERRIDE ratings given by the general community in the initial voting round. I get it, they're not just completely throwing away the early results. But they are being granted more power to some degree.\n\nOn the other hand, we have what we will call the CLASS B solutions. Class B solutions are more desirable because they do not compromise on the goals of LD whatsoever (or at least do so very minimally, but that's up to you guys to decide). CLASS B solutions cannot weight votes, and they cannot mandate participation in any capacity. They can SUGGEST or ENCOURAGE participation in order to generate more ratings and hopefully reduce controversy or feelings of suspicion. An example of a CLASS B solution I've seen proposed is one that was brought up a few events back, where there would be something like a homework queue that rewards participants with extra Coolness in exchange for rating a suggested subset of submissions. This is certainly a cool idea, and does encourage participants to rate more games in general, but it DOES NOT DIRECTLY ADDRESS THE CORE COMPLAINT, which is that CATEGORY WINNERS SHOULD HAVE RELATIVELY HIGH COUNTS OF RATINGS.\n\nSo to summarize what I've said up until this point, what we really want is a CLASS B solution which stands a chance to address the core issue of winners with few ratings, without going as far as a CLASS A solution would to reduce community voting power or mandate participation. This is what I've come up with:\n\nYou can think of it as 2 voting rounds, similar to the CLASS A proposals so far, but technically, there is only 1 voting round. They're more like 2 voting phases. Participants only get 1 vote, as they do now, and votes are not weighted, so everyone's vote is equal, as it is now. Voter power is never once reduced in the slightest. The only difference is that the intermediate results are made public prior to the 2nd phase. Then, the final results are tallied after both phases are complete. Again, I want to make very clear, this is almost exactly the same as the way it is now. Each participant gets to vote on each entry exactly once, and that 1 vote is equal. The ONLY difference is that for the 2nd portion of the voting period, the intermediate results are made public. By doing this, you allow people who haven't voted on a game near the top of the rankings to vote on it, thereby increasing the sample size of the winners.\n\nThink of it like getting dressed in the morning. Sometimes, you're in a rush, and you quickly slip into your pants, but your underwear doesn't feel quite right, so you adjust. The 2nd phase is the adjustment. If the intermediate results are released and people feel there is a game which doesn't belong at the top, they can add their vote, thereby increasing the confidence in the overall ranking. If they choose not to vote, and the game stays where it's at, that's okay too. The game isn't automatically disqualified, its fans don't have their votes questioned, and the game wins, because people had every chance to vote against it but decided not to. If you wanted, you could even have a confidence threshold, where you bump a game up or down a slot for having a massive gap in number of ratings compared to those around it, but honestly, at that point, we're just splitting split ends of hair. You could also prohibit the top 100 in a category from voting on each other's games as well, just to eliminate spiteful voting between them, but it's not really worth doing considering they each get only 1 vote and could just as easily make a bunch of fake accounts.\n\nSo far, this is the only CLASS B system I've been able to come up with in my head that comes anywhere close to protecting the values of LD while potentially addressing this recurring issue. Unfortunately, it is not without its flaws. It does absolutely nothing to prevent intentional abuse, just like the current system. Short of implementing a strictly managed CLASS A solution, I don't think there's any reasonable way to prevent intentional cheating. We're better off trying to find ways to utilize the resources we have, namely thousands of participants, to generate more honest ratings to combat the cheaters. It's also kind of sad to think that a submission could be a category leader going into phase 2, and then be edged out by a competitor generating late buzz. I mean, sure, the results are not locked in until after phase 2, but still, that's a bummer.\n\nIf you've made it this far, I apologize for the long description. The system itself is actually quite simple, but understanding the reasons behind it was quite the brain exercise. Please leave your comments below.\n\nADDENDUM: And for what it's worth, the reason I put \"improved\" in quotes in the post's subject line is that I personally do not believe this system will improve the objectivity of the rankings overall. I believe the system will be just as flawed, maybe even more so, because of the fundamental principle that I mentioned earlier about not having an objective target goal.\n\nThink of it this way. Imagine attending an NBA Finals match to determine the winning team of the basketball season. There is a very objective goal. Whichever team scores the most points in the allotted time wins. That may not determine the best team overall, but it does determine the winner of that particular contest according to a strict set of rules. Now, imagine if instead of tracking points scored, the winner was determined by the volume of cheers from the stands. At the end of the game, the referee asks the fans of each team to take turns cheering as loudly as possible for who they thought won (I know there are instruments to objectively measure noise volume, but for the purposes of this example, pretend the ref is doing it by ear, and he's deaf on one side and spinning in a circle). Whichever one he thinks sounds the loudest wins. This system I've proposed is a little bit like fitting more people in the stadium to improve the confidence in this win-by-sound system. I'm just not convinced it's actually better in such a noisy (no pun intended) system.\n\nYou have to consider the very real possibility that someone in the middle tiers has a strong entry and wants honest feedback, but they can't generate buzz because everyone is too busy rating top tier games in phase 2 which don't even necessarily appeal to them. You've actually shifted the balance of exposure negatively, all for the sake of reducing suspicion, and the value of the change in objectivity is questionable at best. This \"fix\" is nothing more than a placebo to ease the suspicions of people who don't actually understand why community votes are not objective. In fact, that person who can't generate buzz may then be encouraged to seek out friendly votes from people they've curried favor with, introducing even more bias and muddying the middle tiers. Butterfly effect. Everything affects everything else.\n\nIf my suggestion were implemented and I was offered to participate in it, I'd have to politely decline. Because again, I recognize a CLASS B system will never have anywhere near the level of objectivity that a proper CLASS A system will. You need a CLASS A system for this to be a true competition. A CLASS B system is something you use to judge a silly informal game of hop-scotch on a kindergarten playground. A CLASS A system is needed if you really care about the competitive aspect of something. It's up to you, the LD community, to decide where your priorities are.","title":"Another Suggestion for \"Improved\" Rankings","wayback_source":[]}