I reached 1,000+ telemetry events across 38 deployments — what should I be thinking about now?
Summary
A solo developer of an open-source AI workflow automation platform discusses scaling to 1,000+ telemetry events and seeks community advice on maintenance, backward compatibility, and observability for growing usage.
Similar Articles
Which platform is your company using for ai agent observability and reliability needs?
A developer building multi-agent financial workflows seeks community advice on observability and reliability tooling for AI agents in production, sharing frustration with fragmented landscape and cascading failures.
Anyone actually running AI agents in production with real users - not demos, not 10 beta testers. What's your stack? And has anyone moved back to traditional code after trying agents in prod - why?
A discussion prompt asking about real-world AI agent deployments with 100+ users, covering tech stacks and scaling issues, plus experiences of moving back to traditional code.
What are you using for observability?
A developer discusses the lack of suitable observability tools for AI agents, expressing disappointment with existing solutions like Opik and hoping for a service that supports OpenTelemetry for analyzing agent sessions and failure modes.
How are you structuring production-ready development with AI coding agents?
A web developer shares experiences building systems around AI coding agents for reliable development and seeks advice on workflows and orchestration tools.
Experience sharing: building an AI Agent to Triage GitHub, Discourse, and Email (A Real-World Use Case for OSS Maintenance)
The author shares a case study on building an AI agent for Seafile that triages support requests across GitHub, Discourse, and Email by synchronizing knowledge and providing actionable suggestions to maintainers.