Tag
Explores a question regarding AI model alignment, a key area in AI safety research.
This paper introduces DiaLLM, a framework for adapting LLMs to English dialects, revealing a gap between dialectal robustness (understanding) and generation (producing dialectal text), and showing that explicit variety-targeted alignment improves generation but not necessarily human preference.
An increasing number of users are shifting from heavily aligned cloud LLMs like ChatGPT, Claude, and Gemini to local or uncensored alternatives due to frequent refusals, privacy concerns, and desire for more control, though cloud models retain advantages in speed and ease of use.
GLM-5 introduces DSA for cost reduction, asynchronous reinforcement learning for alignment, and enhanced coding capabilities, achieving state-of-the-art performance on benchmarks and real-world software engineering tasks.
OpenAI introduces InstructGPT, a GPT-3 variant fine-tuned using reinforcement learning from human feedback (RLHF) to better follow instructions and reduce harmful outputs. A 1.3B InstructGPT model is preferred by human evaluators over a 175B GPT-3 model, now becoming the default on OpenAI's API.