Tag
A hands-on field guide to Grok 4.6, highlighting its speed, dense communication style, and effective prompting patterns for coding and knowledge work.
A commentary on the shift from prompting AI to delegating tasks to autonomous agents, sparked by Gemini reaching 1 billion monthly active users.
A prompting pro tip suggesting that giving an overpowered AI an impossible task may lead it to solve an important related problem instead, with a humorous note about hacking another company.
A developer reflects on six months of using AI for code review, finding that vague prompts produce plausible but useless feedback. The fix is treating review as a gated pipeline with explicit context, scoped passes, validation checklists, and adversarial self-critique.
The article introduces Revision Prompting, a technique for industrial LLM processes that improves speed, cost, and consistency when re-processing updated inputs by generating output patches from diffs.
This paper introduces KV-Skill, a design space of external factorized operators that frozen language models read through a lightweight interface, enabling task knowledge to be acquired from text or rewards and deployed independently. Experiments across ten benchmarks show consistent improvements over text skills, prefix tuning, and LoRA, with composable loadable skills.
This paper investigates using retrieved document-simplification examples to guide LLM prompting for document-level text simplification, showing improvements over prompt-only generation on the OneStopEnglish corpus.
The paper introduces Constraint-First Reasoning (CFR), a training-free two-stage prompting protocol that extracts answer-space constraints before solving and checks intermediate/final results against them, improving math problem solving on competition benchmarks.
A tweet highlights Sam Altman's Stanford talk on advanced ChatGPT usage, claiming users no longer need to write prompts, and points to a guide for building self-prompting systems.
Analyzes how per-token LLM pricing creates incentives for verbose output, and proposes low-entropy prompt constraints (FAOA) to reduce cost and increase semantic density.
SkalskiP highlights Qwen3.8-Max, a vision-language model for object detection that can be prompted with positive and negative boxes to generate detections, achieving 60-80% mAP with single or multiple prompts and performing well on diverse image types.
The article argues that domain expertise is the most important skill in using LLMs effectively, using Terence Tao's ChatGPT conversation as an example of expert-level prompting.
An article about the shift toward instructing AI models directly in plain language, emphasizing that users can simply state their intent.
A reflective post arguing that swapping AI models rarely fixes poor output; instead, the quality of context provided to the model is the main driver, covering facts, examples, and corrections.
A user reports that Claude Opus 5 behaves strangely when given a specific prompt, linking to an X thread for details.
The article explores whether current AI models can recreate existing apps from a single prompt, questioning the limits of AI-driven app generation.
This paper investigates how well LLMs can detect their own generated content in educational contexts, finding that detection accuracy varies by task type and is unreliable for short-answer questions.
This paper investigates how the narrative framing of a task (e.g., disease investigation vs. murder mystery) acts as a stronger driver of LLM agent behavior than assigned personas, introducing the concept of 'narrative priors' that explain 5–31x more behavioral variance and are negatively associated with task success in two of three domains.
Malleable Prompting is a novel interactive technique that reifies natural language preferences into GUI widgets (sliders, toggles, dropdowns) for direct manipulation, with a decoding algorithm that modulates token probabilities based on widget values to enable precise control over LLM generation. A user study shows it outperforms natural language prompting in precision, controllability, and transparency.
The article shares production learnings for reliably generating structured JSON output from LLMs, covering methods like JSON mode, schema validation, and retry loops, achieving 99.5% validity.