@_philschmid: Gemini 3.7 Flash just took #1 on @ArtificialAnlys new AA-AnalystAgent. AA-AnalystAgent evaluates against 80 real-world …
Summary
Gemini 3.7 Flash takes first place on the new AA-AnalystAgent benchmark, excelling in accuracy (60% pass^5), speed (1.32s per task), and cost efficiency across 80 real-world quantitative analysis tasks in multiple domains.
View Cached Full Text
Cached at: 08/19/26, 02:52 PM
Gemini 3.7 Flash just took #1 on @ArtificialAnlys new AA-AnalystAgent.
AA-AnalystAgent evaluates against 80 real-world quantitative analysis tasks across 14 business and scientific domains (finance, healthcare, hydrology, government appropriations).
The Agent run inside an isolated Python 3.12 sandbox using AA’s open-source Stirrup harness. They receive reference spreadsheets (.xlsx) and documents (.docx) alongside standard data libraries (pandas, polars, openpyxl, scipy, PyMuPDF) to inspect schemas, write scripts, handle edge cases, and calculate final figures.
Gemini 3.7 Flash: • Accuracy: #1 with 60.0% pass^5 (70.5% pass@1, 77.5% pass@5) • Speed: 1.32s per task (fastest) • Cost: $0.54 avg per task (middle)
Similar Articles
@narens: Benchmaxxed
Gemini 3.7 flash outperforms Fable 5, Opus 5, and GPT-5.6 on the Analyst Agent benchmark by Artificial Analysis.
Gemini 3.5 Flash looks worse than it seems on Artificial Analysis
Comparison showing that Gemini 3.5 Flash scores slightly lower than Gemini 3.1 Pro in Artificial Analysis benchmarks and has a higher total benchmark cost despite lower per-token API pricing.
Gemini 3.5 Flash Benchmarks
Benchmark results for the Gemini 3.5 Flash model are discussed, likely showcasing its performance across various AI tasks.
Gemini 3 Flash: frontier intelligence built for speed
Google has released Gemini 3 Flash, a fast, cost-effective AI model that combines Pro-grade reasoning with Flash-level speed for tasks like coding, complex analysis, and agentic workflows.
Gemini 3.5 Flash Looks Good For How Fast It Is (8 minute read)
Google released Gemini 3.5 Flash, a hybrid speed model that rivals Opus 4.7 and GPT-5.5 in speed and cost while performing well on agentic and coding benchmarks.