Show HN: Product analytics (and evals) for agent sessions on your MCP

Hacker News Top Products

Summary

Armature is a new product analytics tool for AI agent sessions, capturing user intent, agent thinking, and success scores via an SDK for MCP servers, with automatic PII redaction.

Hi HN! We’re Theodore and Louis, founders of Armature (YC P26). We reconstruct the entire session behind the MCP tool calls you receive, including what the user asked their agent to do and what the agent thought.<p>You wrap your MCP in 3 lines of code (our SDK is available in Typescript, Python and Go) and start seeing in your dashboard: - All sessions reconstructed: it’s like reading the real conversation the user had inside Claude or ChatGPT! - A ranking of your MCP most popular use cases, built from sessions clustering - The most frequent issues your users’ agents encounter so you can fix them.<p>Here is a quick demo: <a href="https:&#x2F;&#x2F;youtu.be&#x2F;ZFlvquhyNMQ" rel="nofollow">https:&#x2F;&#x2F;youtu.be&#x2F;ZFlvquhyNMQ</a><p>The story behind this is that we initially launched Armature as a standalone testing tool (<a href="https:&#x2F;&#x2F;www.ycombinator.com&#x2F;launches&#x2F;QQc-armature-making-your-app-finally-usable-by-ai-agents">https:&#x2F;&#x2F;www.ycombinator.com&#x2F;launches&#x2F;QQc-armature-making-you...</a>) that could naturally be used through an MCP itself. We quickly realized we had no idea how our users were using Armature MCP and if they were satisfied with it or frustrated. It’s something we had also experienced in our previous companies: Louis built MCPs exposed to millions of users and Theo was a Forward Deployed Engineer at Palantir before joining a Datadog spin-off as Founding Engineer. Both testing and product analytics had always been real pains when exposing a product to agents but we always thought there wasn’t much we could do about analytics because the conversation lived in our users’ AI client.<p>Then it struck us: what if we asked the agents why they were making this or that tool call? And what’s the user&#x27;s intent or potential frustration? So we started experimenting with MCP instrumentation and the use-cases actually surprised us! Many of our first customers had implemented workarounds for their CI to trigger new tests or for their coding agents to fetch the results efficiently. Even though we talked to our first users regularly, they had never shared this feedback with us. We then built automations to automatically cluster use-cases, identify issues frequently encountered and let our own coding agents fix them. When our CTO friends heard about this, they wanted to try it for themselves so we gave them access to a cloned version of our internal product and they started sharing feedback like they never did on our “real” product!<p>That’s when we decided to start working seriously on MCP Analytics as a product. At first we were afraid of degrading MCP performance so we iterated until we reached the exact same success rate as without our instrumentation (89.17 % vs 89.15 % pass rate out of 870 runs). Then privacy was an obvious constraint so we applied the same methods we had learned from working with banking data or building sensitive data scanning in logs. Today, redaction runs client-side before reaching our servers. There are still a lot of things we haven’t fully figured out: not all fields are equally filled by all models, session fingerprinting for serverless &#x2F; stateless MCPs isn’t perfect, and use-case clustering remains to be optimized.<p>But we are finally launching our analytics product to everyone, self-serve at <a href="https:&#x2F;&#x2F;armature.tech">https:&#x2F;&#x2F;armature.tech</a> with a set-up that takes less than 5 minutes and a generous free tier.<p>And now we are working on fully closing the loop, bringing evals back in our product so we can: identify top workflows and issues -&gt; recommend fixes and improvements -&gt; test fixes at scale on the same workflows run by users, across all harnesses and models -&gt; open PRs to ship fixes directly. The evals can be generated automatically from the session analytics so you can catch every regression and can test every improvement’s real impact across all models and harnesses before shipping it.<p>Here’s an example to make it more concrete: 10 days ago, a marketing automation platform which has had early access to what we built for weeks identified thanks to MCP Analytics that users were frustrated not being able to change their target audience after campaign creation. So they shipped the feature and tested it successfully locally with Claude Code on Fable 5. Then a few days later when preparing their new MCP public release, they ran a suite of evals on Armature and realized that small models could hallucinate audience_ids which would lead their MCP to send the campaign to ALL their contacts by default (which could obviously lead to disasters in prod). This is the kind of story that makes what we are building feel so helpful!<p>Now, the most useful feedback for us would be to know what’s still missing in our product so you can feel you are now in full control of the “Agent Experience”. And if you run an MCP in production we’d also love to know: what do you do today to know if agents succeed and if the users behind them are happy?
Original Article
View Cached Full Text

Cached at: 08/03/26, 07:34 PM

# Armature · Product analytics for agent sessions Source: [https://armature.tech/](https://armature.tech/) Product analytics for agent sessions User sessions on your![](https://armature.tech/assets/logos/claude.svg)Claude Connector![](https://armature.tech/assets/logos/chatgpt.png)ChatGPT ApporMCPlive inside their AI client, not your UI\. Armature captures and shows you what users ask, what agents think and how they use your product\. list\_customers◂ create\_subscription◂ send\_invoice402◂ send\_invoice◂ search\_customers◂ get\_plans◂ list\_invoices◂ update\_plan◂ usage\_report◂ ![Claude](https://armature.tech/assets/logos/claude-wordmark.png) User intent · 14:16Set up a paid plan for ACME and invoice them **✦**Agent thinking · 14:16Invoice needs a billing contact — retrying with the owner email\. ✓ Agent completed user task successfully**92***score* With Armature - Session rebuilt - User intent captured - Agent thinking captured - Session evaluated Use cases ## Learn your users' topuse cases Armature's models read every session and identify what the user came to do\. Sessions are grouped into use cases and ranked by volume and success rate,**including the use cases you don't support yet**\. Create & send invoices38% Bulk refundsnot supported yet14% Issues ## Spot the topissuesand fix them Models scan every session for failures, loops and dead ends, group them by root cause and rank them by how many users encountered them,**even when every API response was 200 OK**\. 127Agent loops on missing auth scope↑ 43 this weekfix → 64Search misses "refund" phrasing 31Export truncated by pagination 12Rate limit hit on bulk updates Session replay ## Replay and investigate anysession Every session is scored on whether the user got what they asked for\. Open any of them and replay the full trace: the ask, the agent's thinking and every call\.**When something breaks, you see exactly where\.** User intentSet up a paid plan for ACME and invoice them list\_customers◂ **✦**Agent thinkingInvoice needs a billing contact — retrying with the owner email\. send\_invoice402◂ send\_invoice◂ ✓ Agent completed user task successfully**92***score* ## Set up in a few minutes ### Create your key Sign up, generate an API key and drop it in your deployment secrets\. ### Prompt your coding agent Paste one prompt and let your agent wire the SDK in\. ### Deploy Ship it and watch sessions flow in within minutes\. ## Your users' data stays safe Sessions can carry personal information and secrets, so we treat every one of them as sensitive\. Detection models scan sessions and redact PII and secrets by default, before anything reaches storage\. ## Start for free Or reach out to us for custom needs\. Free$0 - 1,000 sessions / month includedthen $50 / 1,000 sessions - Unlimited projectslaunch offer - Unlimited users - 7\-day retention - Dashboard access - Session replay - Use\-case groupinglaunch offer - Issue identificationlaunch offer [Start for free→](https://app.armature.tech/)No credit card required CustomLet's talk - Everything in Free - Priority support & SLA - SSO / SAML - Audit logs - Custom retention - Dedicated onboarding [Contact us](mailto:[email protected]?subject=Armature%20custom%20plan) ## Common questions 01How does Armature capture sessions?You add our SDK to your MCP server, Claude Connector or ChatGPT App backend\. It takes a few lines and one deploy\. From there Armature rebuilds every session, classifies it and scores it automatically\. 02Do I have to change my MCP server?No\. You keep the server you already run\. The SDK wraps it to capture sessions and changes nothing about how it behaves for your users\. 03Which agents and clients do you support?Any client your users bring: Claude, ChatGPT, Cursor, Codex, Gemini CLI and the rest\. If it can reach your MCP server, Armature can capture the session\. 04How is this different from PostHog, Amplitude or Mixpanel?Those tools track humans clicking your UI\. When a user delegates to an agent, the session happens inside Claude or ChatGPT and your UI never sees it\. Armature captures exactly those sessions\. 05How is this different from LangSmith or Langfuse?They observe the agents you build yourself, and they serve the engineers who build them\. Armature shows product teams how your users' agents experience what you ship\. 06Is my users' data safe?Yes\. Detection models scan every session and redact PII and secrets by default, before anything reaches storage\. You control retention and can delete your data at any time\. Ask us for the full security details, we are happy to share them\. 07What does it cost?It is free up to 1,000 sessions a month, with unlimited projects and users\. Past that you pay $50 per 1,000 sessions\. Custom plans cover teams with serious traffic\. ## Start improving your agent experience Capture the sessions users have with your product through any AI agent\. See what they ask, watch how agents deliver and fix the issues they encounter\.

Similar Articles

Show HN: Ax-check.com – Can agents use your product?

Hacker News Top

Ax-check.com tests how well AI agents can use your product by having three agents attempt end-to-end onboarding and provides concrete guidance for fixing documentation, marketing sites, CLI, MCP, and skills.

Paid Agent MCP servers

Reddit r/openclaw

A developer is building 402 MCP servers for AI agents and seeks community feedback on which paid microservices would be most valuable for making agents more useful.