The author describes building a small business operations system using only an iPhone and AI tools like ChatGPT, detailing the architecture and seeking feedback on potential failure modes.
I’m not an engineer. I’m building a small business operations system using only an iPhone, ChatGPT, Google Drive, scheduled tasks, and browser-based AI tools. I started very simply, but small failures kept appearing: scheduled tasks triggered, but I couldn’t prove the work actually finished external actions could become uncertain after timeouts or missing responses retries could potentially duplicate side effects long-term memory could become stale technical execution success and actual business success were easy to confuse an AI could continue after a tool failure and still produce a plausible answer After repeatedly fixing those problems, I ended up with something like this: Google Drive as the source of truth structured state, registries, and execution evidence a Run Audit for each execution TECH_STATUS and BUSINESS_STATUS tracked separately an UNCERTAIN state when an external result cannot be confirmed no blind retry after an uncertain external action duplicate-execution / idempotency checks freshness checks immediately before execution human approval for high-impact actions regression cases built from previous failures AI for ambiguous judgment, deterministic handling for state, logs, deadlines, and calculations The rough execution loop is: Retrieve → Understand → Decide → Act → Verify → Reconcile → Record I also try to keep action permissions narrow while allowing the AI to reason broadly. The unusual constraint is that I built and operate all of this from an iPhone. No custom backend, no self-hosted server, and no coding environment. I’m Japanese and not an engineer at all. I even had AI help me write and post this because my English is limited. But I genuinely want honest feedback from people who know this field better than I do. I’m not looking for praise. I’m looking for failure modes. If this had to run continuously for months or years, what do you think would break first? I’d especially appreciate criticism around: long-term state and memory stale or conflicting information task monitoring retries and duplicate execution evidence and auditability human-in-the-loop approval separating technical success from business success recovery after partially completed external actions If you were reviewing this architecture, what would you change first? I can share a simplified diagram or a concrete failure scenario if that would help.
A 397-billion-parameter AI model has been successfully run on an iPhone, demonstrating on-device AI capabilities with a mixture-of-experts design, though facing challenges in speed, storage, and heat.
The author describes building a multimodal WhatsApp AI system using n8n but struggles to sell it, seeking advice on creating a project that solves real business needs.
The author shares a postmortem on building a production phone-based AI voice agent, revealing that most engineering time was consumed by telephony infrastructure, turn detection, observability, and failure handling rather than core LLM behavior. They suggest using managed platforms like Vapi, Retell, or Dasha from the start to focus engineering effort on business logic.
The author describes a setup where different AI models are assigned to specific roles (planning, coding, review) to reduce API costs for a 24/7 autonomous engineering team, and shares common failure points like model wandering and hallucinated ownership.