Humor Arena: Which LLM is the funniest?
Summary
The article presents a comparison of 20 LLM versions on humor generation using a fine-tuned judge, with Fable 5 scoring highest in a joke benchmark.
Similar Articles
HumorRank: A Tournament-Based Leaderboard for Evaluating Humor Generation in Large Language Models
HumorRank introduces a tournament-based leaderboard using pairwise evaluations and Bradley-Terry MLE to rank LLMs on humor generation, showing humor quality depends on comedic mastery rather than scale.
i built a benchmark to test whether LLMs can understand and create jokes
A developer built a benchmark called lolbench to evaluate LLMs on understanding and creating jokes, revealing models are strong at explaining real jokes but struggle with failed ones, and incorporating human voting for preference testing.
What taste in humor do LLMs have?
A research report reveals that LLMs have varying levels of agreement in humor taste, with a general preference for absurd humor and roast joke formats, and Gemini being the most consistent judge.
Computational Humor with Multimodal LLMs: Methods, Datasets, Evaluation, and Challenges
This survey examines computational humor understanding in multimodal LLMs, covering methods, datasets, evaluation protocols, and challenges such as shortcut-prone evaluation and weak evidence grounding.
Computational Humor with Multimodal LLMs: Methods, Datasets, Evaluation, and Challenges
A comprehensive survey on multimodal humor understanding using large language models, covering methods, datasets, evaluation protocols, and challenges in interpreting humor in memes, cartoons, and comics.