Jozef

Ludum Dare 53

The best jokes of Ludum Dare 53 (and other interesting stats)

Can You Outwit AI with Your Jokes? Try Our Game and Find Out!

In our game, Impasta, you are given one of 300+ joke setups to which you need to come up with a punchline. There's a little twist though! You need to first play a mini-game of Peggle in which you unlock letters and then you can only use these letters to write the punchline. The punchlines are rated on a scale 0-10 by our backend using GPT-4.

blog0.png

At the moment of writing this post, our game has received a 35.5 ratings on Ludum Dare, and there have been 321 jokes submitted and rated. Let's dive a little bit more into these stats!

First, the number of players that passed through each of the 10 rounds:

blog1.png

While sadly we can see about a 30% drop-off after the first round, for the other rounds there's a rather consistent drop-off of around 5 players per round. In the ideal world, we would, of course, like to have every player play the game until the end, but on the positive side, we can at least see that there isn't any single "stumper" level that would make players more likely to quit the game.

Let's have a look at how the players actually did in Peggle. This chart shows how many letters out of 26 players unlocked on average in each level. Keep in mind the vertical axis is on scale 20-26.

blog2.png

We've planned for players to be able to get all 26 letters most of the time, so it's nice to see that on average they've managed to get at least 20 letters in each round. We also planned to have every level except for the last one to be increasingly harder, so it's nice to see how the player performance steadily declined over the first 5 levels. After that, it jumped sharply back up but that's possibly because most of the players who were not as good at Peggle quit the game out of frustration?

blog3.png

Image showing levels 1, 4, 5 and 9

But how did the players fare at coming up with punchlines? Our first version of GPT-4 prompt we've used for rating gave us only 4 different number ratings even though we told it to rate the punchlines on a scale of 0-10: 0, 2, 4, and 7. This was obviously not ideal as we wanted the ratings to be spread throughout the entire scale and also be able to score higher than 7. We've tried many approaches with modifying the prompt and other parameters. This process was too long to describe here, so I might write another post where I go more in-depth about it. In the end, we've come up with a prompt and parameters that seemed to perform quite well, so let's see how it did with the actual playerbase:

blog4.png

We can see that with our current prompt, GPT-4 still has a tendency to avoid certain number ratings like 5 or 9, but overall it wasn't that bad. There have even been some 10/10 jokes! I'm sure you must be curious what the best jokes were. I want to note that the majority of the 10/10 came up from the first round, where we used popular jokes setups that most people already know, for example:

What do you call a fake noodle? An impasta!

(Yes, that's where the name of the game comes from).

Now without any further ado, here are the "original jokes" our players came up with that scored 10/10:

Why are robots terrible at picking up on social cues? They only get ones and zeroes

Why do rabbits make great comedians? They don't carrot about critics

What do you call an alligator that knows the alphabet? Alliterator

How does a lobster feel when it's about to perform stand-up comedy? It feels a little shell-doubt!

What kind of music does a tree listen to? Oakestral

And my favorite:

What did the sushi say to the bento box? Wasa-b

I hope you found this post at least a little funny and maybe educational. If there's interest, I might write another one about how we went about implementing our GPT-4 powered joke rating system.

Our game is available to play and rate at https://ldjam.com/events/ludum-dare/53/impasta