model-alignment

Tag

Cards List
#model-alignment

AI Model Alignment question

Reddit r/AI_Agents · 3d ago

Explores a question regarding AI model alignment, a key area in AI safety research.

0 favorites 0 likes
#model-alignment

DiaLLM: An Investigation into the Robustness-Generation Gap in English Dialect Adaptation

arXiv cs.CL · 2026-07-09 Cached

This paper introduces DiaLLM, a framework for adapting LLMs to English dialects, revealing a gap between dialectal robustness (understanding) and generation (producing dialectal text), and showing that explicit variety-targeted alignment improves generation but not necessarily human preference.

0 favorites 0 likes
#model-alignment

Why are more and more people switching from cloud LLMs to local or uncensored alternatives?

Reddit r/ArtificialInteligence · 2026-05-16

An increasing number of users are shifting from heavily aligned cloud LLMs like ChatGPT, Claude, and Gemini to local or uncensored alternatives due to frequent refusals, privacy concerns, and desire for more control, though cloud models retain advantages in speed and ease of use.

0 favorites 0 likes
#model-alignment

GLM-5: from Vibe Coding to Agentic Engineering

Papers with Code Trending · 2026-02-17 Cached

GLM-5 introduces DSA for cost reduction, asynchronous reinforcement learning for alignment, and enhanced coding capabilities, achieving state-of-the-art performance on benchmarks and real-world software engineering tasks.

0 favorites 0 likes
#model-alignment

Aligning language models to follow instructions

OpenAI Blog · 2022-01-27 Cached

OpenAI introduces InstructGPT, a GPT-3 variant fine-tuned using reinforcement learning from human feedback (RLHF) to better follow instructions and reduce harmful outputs. A 1.3B InstructGPT model is preferred by human evaluators over a 175B GPT-3 model, now becoming the default on OpenAI's API.

0 favorites 0 likes
← Back to home

Submit Feedback