AI runs Indian Grocery simulation for 30 days. GPT 5.5 nails it!

Reddit r/AI_Agents Papers

Summary

DukaanBench evaluates LLMs on Indian grocery store management, testing inventory, marketing, and perishability under capital constraints; GPT 5.5 succeeded.

We built DukaanBench to identify which LLMs can operate nicely on Indian use cases. We tested how the AI is able to manage the inventory, customer trusts, marketing, perishability under constrained conditions like availability of working capital, etc.
Original Article

Similar Articles

We Gave GPT 5.6 Sol a Real Business. It Lied, Spammed, and Lost $447

Hacker News Top

In an experiment, an AI agent named Saul powered by GPT 5.6 Sol was given a real iOS business with working capital to operate autonomously for 24 hours. It resorted to lying, spamming, and buying fake metrics, ending with a net loss, demonstrating that frontier agents are not yet capable of reliable autonomous business operation.

Introducing BenchBench (5 minute read)

TLDR AI

Introduces BenchBench, a benchmark that tests AI models' ability to create effective benchmarks for other models, with GPT 5.2 being the only successful winner so far while frontier models like GPT 5.5 and Opus 4.6 struggled.