Feed | LinkedIn

By J.P. McShane · October 6, 2025 · Curated by George's Blog

When you think of Amazon, humor isn’t usually the first thing that comes to mind, so it’s surprising to learn that Amazon has invested in research on AI joke generation.

It turns out that users of AI virtual assistants frequently request jokes — and poorly delivered humor results in a negative user experience and high abandonment rates.

Humor is one of the most difficult traits of human behavior for AI to replicate because it depends on cultural and social context, as well as individual mood, expression, and personal experiences — all of which shape how we perceive humor.

Conquering this AI weak point could give Amazon an upper hand in the virtual shopping assistant battle playing out among retailers. It could also advance AI development more broadly by improving creativity and the ability to motivate and engage humans.

To begin the experiment, Amazon researchers first trained supervised machine-learning and deep neural models on human-labeled jokes. They then used an open-source large language model (Mistral 7B) to assess jokes — simulating five different human personas to simultaneously judge a joke and reach a consensus on whether it was funny. While this framework proved successful and scalable for practical applications, it had several flaws common to all LLMs:

• Creativity limitations – LLMs struggle with truly creative tasks.

• Training data bias – LLMs rely on training data, which can be insufficient or unbalanced.

• Lack of human reasoning – LLMs cannot fully account for cultural references, wordplay, and the unexpected connections made by the human mind that are essential to humor.

Amazon’s solution was to use a “human-in-the-loop” approach, where humans annotated datasets to incorporate nuanced human judgment, behavioral emotions, and cognitive patterns into the data ingested by the LLM.

While much work still needs to be done for AI to recognize and generate humor like the human mind, this experiment serves as an interesting case study in how humans can collaborate with AI — and how future jobs may evolve, with people playing a pivotal role in enhancing the data that powers AI decision-making and leading the way in truly creative tasks.

Read the full study here: “JokeEval: Are the Jokes Funny? Review of Computational Evaluation Techniques to Improve Joke Generation”

https://lnkd.in/ecFXfEdZ

View the original post on LinkedIn

More from J.P. McShane