Tag
The article discusses how public benchmarking of hallucination in LLMs has driven rapid improvements and explores implications for alignment, emphasizing the need for benchmarks on transparency and honesty.
A tweet highlights PoolsideAI's unusual openness, praising their release of a small coding model, publication of papers, and full evaluation datasets, setting a standard for transparency in AI.