Tag
GPT-5.6 achieves first place on Eq-Bench's Creative Writing benchmark, demonstrating significant advancement in AI-generated creative text.
An opinion piece discusses how recent updates to Claude have made the AI chatbot more prone to refusing requests and lecturing users, particularly on sensitive creative topics, frustrating long-time users.
Jerry Liu shares that creating a skill to reflect his writing style and hillclimbing based on real samples solves 70% of the style problem, but notes that articulating clear insights remains a post-training challenge.
Akwin123 fine-tuned Gemma-4-31B-it specifically for copywriting and creative writing tasks, achieving a +290 Elo improvement over the base model on a custom benchmark using QLoRA. The model is open-source and available on Hugging Face.
This paper introduces Loom, an assisted writing framework that leverages a three-layer pipeline based on narratological story/discourse distinction to control narrative intent and rendering density, achieving improved factual integrity and descriptive intensity compared to baselines.
This article presents a technique to improve LLM creative writing by modifying the sampling process using entropy, aiming to reduce the generic 'LLM feel' in generated text.
PageStorm is an AI model specifically designed for creative book writing, aimed at helping authors generate narratives and plots.
A site ran longform speculative-fiction prompts through AI models including Claude Fable 5, publishing the resulting story 'Headwaters' with process notes, raising questions about language becoming training material that people might need to hide.
OPERA proposes a reinforcement learning method for open-ended tasks using intrinsic rewards based on perplexity dynamics, replacing unreliable LLM-as-a-judge reward models. It achieves state-of-the-art results on Qwen3-8B, matching proprietary models in creative writing and other open-ended tasks.
A user expresses concern that current AI models have become less creative and more corporate-sounding due to safety guardrails, contrasting them with earlier open models that were more imaginative.
An analysis exploring why Gemma 4, despite advantages like QAT and vision support, lacks community finetunes compared to Mistral, and whether community inertia will eventually shift.
GLM-5.2 is an open weight AI model optimized for creative writing tasks, claimed to be the best in its category.
Researchers from CUHK-Shenzhen introduce a jailbreak method using fanfiction subgenres from Archive of Our Own as attack carriers, embedding harmful content within creative writing scenes. Their method achieves a mean attack success rate of 0.731 on eight aligned LLMs, with a multi-turn extension (Saga-A4) reaching 0.924 ASR, outperforming existing methods.
POLARIS is a training recipe using GRPO with LLM-as-judge rewards and human-reference injection to improve long-form story generation in small models. Applied to Qwen3.5-9B, the resulting POLARIS-9B model matches Qwen3.5-27B performance on creative writing benchmarks while better adhering to length instructions.
An opinion piece argues that AI-generated fiction is like fast food, lacking the depth and originality of human-written stories, emphasizing the continued need for human authors.
Gemini 3.5 Flash outperforms Gemini 3.1 Pro on a short story creative writing benchmark, improving from -2.3 to -1.8 in head-to-head comparisons.
This paper introduces a dataset and training framework that transforms human-authored novels into multi-resolution planning scaffolds, enabling long-context language models to generate book-scale fiction with more human-like prose and narrative dynamics.
A writing-focused fine-tune of Google's Gemma 4 31B model, aiming for more natural English and better prose, with reduced refusals suitable for creative writing, translations, and roleplay.
OpenAI has published content about GPT-5's creative writing capabilities, highlighting the model's performance in generating creative text. This follows the release of GPT-5, OpenAI's latest and most advanced language model.