Update: parked HITL writes. Cheap Flash for a read-only analytics agent — dates, bilingual routing, and Gemini cache at 0%

Reddit r/AI_Agents News

Summary

The author updates on a read-only analytics agent project using Gemini Flash Lite, discusses caching inefficiencies and bilingual routing challenges, and seeks advice on date handling.

Thanks for the replies on the last post (read-only POS copilot, route → fetch → narrate → ground, Node/TS + Vercel AI SDK). The write-path comments were consistent: preview ≠ confirm in the same loop, idempotency + optimistic lock, fail closed near mutations. That’s enough for me to hold HITL. I’m not adding mutate tools until the read loop is tighter. What I did take: honor pins (no more “computed but not registered”), keep CORE-first schemas, eval on traces / fetch / step count / $/% ⊆ DTO. I did not split into a two-call “CORE then pins” router. Model. Still Gemini 2.5/3.5 Flash Lite in dev. It’s cheap, and this job is “pick a tool, read a small JSON, say it in English or French.” I don’t want a high-thinking model in the loop. Dev is on their free tier. If that’s a trap (cache, tool-calling, bilingual), I’m open to another LLM — same constraint: low cost, good at tools + short narration, not a reasoner. Cost, typical ask (50 intent turns, one store, Redis miss, Flash Lite): Metrics mean median input tokens 4,840 4,718 Gemini cached input 0 0 output tokens 123 107 LLM steps (round trips) 2.0 2 tools called 1.0 1 wall time 2.3s 2.0s registered tools 12.2 12 Shape is almost always: step 0 pick one tool (~2.3k in) → fetch slim DTO → step 1 narrate (~2.6k in). Wider log (n=229): Gemini implicit cache 2.2% overall, 0/96 on current CORE-12, 0/51 on pinned. The only hits were an older CORE-10 prefix (~40% of that request’s input). So the prefix can cache; it isn’t right now. Pins aren’t proven as the breaker — CORE-12 alone is cold too. That’s why I didn’t add a second model call. Caveat: local eval + one store, short sessions. Redis exact-match is only for first-turn how-to, not analytics. Different cache. I’m optimizing the read agent next (Canada EN+FR). Two asks: Dates. Portal always sends a visible dateRange. A phrase table in our code may override it. The model never sets tool from/to — generateObject invented windows last time. When the table misses (hier, “this month vs last month”), do you pause with chips (Claude/OpenClaw-style, but as a UI interrupt, not a Gemini tool), or keep the bar default and name the window in the first sentence? I don’t want to ask on every “what’s our revenue?” Bilingual routing. Gemini can write French. Keyword pins / how-to terms are English, so « préparation cuisine » never loads get_fulfillment. I’m leaning EN+FR synonym tables on the same regexes, not translate-for-routing (mangles $/%) and not htt p 400 for other languages. Anyone run a bilingual keyword router that didn’t rot? Is request locale enough for answer language? Not selling anything. Happy to go into comments.
Original Article

Similar Articles

Gemini 3.5 Flash (Low) (1 minute read)

TLDR AI

Google introduces Gemini 3.5 Flash (Low), a new model variant that uses about 45% fewer tokens than the Medium version while outperforming the older Gemini 3 Flash (High) on SWE tasks. They have also reset quotas for all paid plans.

Gemini API Managed Agents: 3.6 Flash, hooks, and more

Google AI Blog

Google announced updates to Managed Agents in the Gemini API, including a default to the Gemini 3.6 Flash model, environment hooks for tool call auditing, budget controls, scheduled triggers, and free tier access.

Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber

Hacker News Top

Google introduces Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber models, offering improved token efficiency, lower latency, and better performance for agentic workflows, with 3.6 Flash reducing output token usage by 17% and showing gains in coding and knowledge tasks.

Gemini 2.5 Flash-Lite is now ready for scaled production use

Google DeepMind Blog

Google releases Gemini 2.5 Flash-Lite as stable and generally available, the fastest and lowest-cost model in the Gemini 2.5 family at $0.10 input/$0.40 output per 1M tokens, featuring native reasoning capabilities and full feature parity with native tools.