Tag
Box tested Claude Opus 5.5 and found it delivers significant performance improvements over Opus 5 for complex enterprise knowledge tasks, with major gains in token efficiency, speed, and cost.
Anthropic has released Opus 5.5, a new AI model with lower prices and performance matching or exceeding larger models like Fable, featuring improved communication and alignment with safety pacing efforts.
Grok 4.7 is SpaceXAI's most capable AI model for coding and knowledge work, offering improved task handling and safeguards while maintaining the same price and speed as Grok 4.6.
The tweet argues that knowledge work is harder to automate with AI agents compared to software development due to advantages like automated feedback loops and organized documentation in coding.
Claude Fable 5.1 is Anthropic's release of its most advanced AI models for coding and knowledge work, with research capabilities that offer an early glimpse into AI contributions to scientific progress.
Claude introduces Claude Fable 5.1 and Claude Mythos 5.1, claiming they are the world's most advanced models for coding and knowledge work.
Anthropic introduces Claude Fable 5.1 and Claude Mythos 5.1, described as the world's most advanced models for coding and knowledge work.
The paper proposes a harness paradigm for AI agents in large enterprises, focusing on governance and standardization to make AI tools more manageable and compliant.
The tweet compares the historical promotion of smoking as harmless to the current uncertainty about AI's impact on knowledge work, suggesting it may take decades to understand AI's effects.
The author built an AI agent for market research that automates report assembly but still requires human judgment for contextual decisions, highlighting the balance between AI automation and human expertise in knowledge work.
WANDR is a benchmark for evaluating AI agents on wide and deep research tasks, focusing on high-volume data collection with verifiable accuracy. It includes 500 tasks and an evaluation harness to stress-test current systems.
The post highlights the success of Grok Bot as a transformative tool for knowledge work, akin to Claude Code, and criticizes AI labs like OpenAI, Anthropic, and Google for delaying similar releases, which could lead to lost market share.
Garry Tan's insight that startups will become markdown files plus AI agents, turning expertise into durable assets, with commentary on the era of personal AGI.
An essay exploring the existential disillusionment among knowledge workers, who increasingly question the meaning of their careers amid AI disruption and seek analog hobbies or escapes from tech work.
OpenAI released ChatGPT Work, an AI agent for knowledge work that integrates with Slack, email, Drive, and other tools to bring cloud agents to a billion users. The article unpacks its design, position in OpenAI's lineup, and its planned merge with ChatGPT by year's end.
Levie and Max Spero discuss how verifiability determines which 'hard' knowledge work gets automated first, arguing that objectively testable domains like math, cyber, and code will see faster AI automation than fields with subjective outcomes.
APEX-Accounting is a benchmark created by Mercor and Ramp to assess frontier AI models on real accounting tasks. The best model, Claude-Fable-5 (Max), achieved 56.4% mean criteria.
This paper presents a reusable template (llm-wiki-memory-template) that implements the llm-wiki pattern to preserve session memory and failure paths in collaborative knowledge work involving multiple humans, AI agents, and domains. It argues for the substrate's generalizability across three axes and includes case studies demonstrating improved evidence tracking.
The article argues that AI agents will replace human workers not because they are geniuses, but because much of knowledge work consists of repetitive, menial tasks that are already 'agent-shaped'. It questions what parts of jobs truly require human input.
Kimi Work is a desktop AI agent that connects to local files, automates tasks via a cron engine, and offers autonomous web browsing and multi-agent swarm capabilities, designed for knowledge workers and finance professionals.