Tag
This paper presents an artificial life predator-prey model of foraging under noisy perception, showing that uncertainty-aware decision policies significantly improve survival compared to blindly trusting perceptual labels, and that agents transition from exploratory to conservative strategies as uncertainty increases.
Announcement of the NeurIPS 2026 PhysUnderstand Workshop on Physical Understanding for Decision-Making in Sydney, including call for papers and submission details.
Presents RESPClinBench, a real-world scenario benchmark for respiratory clinical decision-making, evaluating seven LLMs on COPD and pulmonary nodule cases. Finds task-specific limitations including imaging hallucination and medication-safety risks.
Building a revenue agent revealed that the hardest challenge is preventing plausible but flawed outputs from translating into real-world actions.
This paper evaluates four machine learning models for discrete choice modeling in policy preference elicitation, using Monte Carlo experiments and a real energy policy case study to assess performance under individual heterogeneity and choice complexity.
Jenova AI's What-If Analyst is a decision-making tool that maps first-, second-, and third-order effects across life domains, surfaces hidden assumptions, and auto-researches current data, available across web, iOS, Android, with a free tier and API.
A blog post warns that AI coding agents often default to popular but unsuitable technologies, incurring technical debt, and urges developers to retain agency in architectural decisions.
The author argues that AI investments are largely failing and causing irrational decision-making across organizations, driven by mass psychosis rather than tangible results.
The paper introduces Mental World Modeling (MWM), a framework that integrates hidden mental states as core components of world models, and presents MENTIS, a training-free baseline. Experiments with 8 LLM-based world models show explicit mental-state modeling is essential for predicting human decisions in situated scenarios.
The article explores whether AI is genuinely improving business operations or remains mostly hype, asking for real-world examples beyond chatbots and content generation in areas like task reduction, customer support, data analysis, and workflow automation.
A philosophical essay arguing that the critical question about AI usage is not whether AI was used, but who made the decisions and provided direction, emphasizing human judgment over tool use.
A developer demonstrates how two AI agents with shared memory running locally on a laptop can catch and prevent bad decisions, even after a restart. The key insight is that persistent shared memory enables agents to operate as a real team.
Introduces DFAH-Bench, a replay benchmark to measure behavioral instability in financial agent decision-making, finding that outcome agreement alone misses significant trajectory divergence.
Fathom helps turn messy bank exports into clear financial decisions.
The article argues against blindly adopting LLMs and provides six questions to evaluate whether an LLM is appropriate for a given workflow, emphasizing that LLMs trade determinism for flexibility and should only be used when necessary.
This paper argues that AI-native biotech companies should use a core computational abstraction called a Company World Model rather than mimicking human department structures. It introduces a dry-lab benchmark comparing architectures and finds that a value-conversion architecture outperforms human-org mimics, but results are objective-sensitive.
This paper introduces a decision-aware weak-to-strong (W2S) learning framework that uses limited labeled data to train a weak model, which then generates soft supervision on unlabeled data to train a strong model for improved contextual stochastic optimization. Theoretical bounds and empirical experiments show that abundant unlabeled data can reduce downstream decision risk when the correlation between weak and strong feature representations is small.
Discusses a contract design approach to prevent AI from autonomously deciding what to ask, shifting focus from tool to action in AI governance.
This article discusses how to determine the appropriate boundaries and restrictions for customer-facing AI agents, focusing on when they should be allowed to act autonomously and when human oversight is needed.
An MCP server that integrates async-first working practices as tools for AI assistants, allowing users to draft decision docs, convert meetings to artifacts, score status updates, and more.