humor-evaluation

Tag

Cards List
#humor-evaluation

i built a benchmark to test whether LLMs can understand and create jokes

Reddit r/artificial · 3d ago

A developer built a benchmark called lolbench to evaluate LLMs on understanding and creating jokes, revealing models are strong at explaining real jokes but struggle with failed ones, and incorporating human voting for preference testing.

0 favorites 0 likes
← Back to home

Submit Feedback