llm-benchmarking

Tag

Cards List
#llm-benchmarking

PICasso: An AI-Enabled Design Framework for Autonomous Optimization of Silicon Photonic Devices

arXiv cs.AI · 19h ago Cached

PICasso is an AI-enabled framework for autonomous optimization of silicon photonic devices from natural-language specifications, demonstrating significant improvements in design satisfaction and loss reduction through LLMs and simulation feedback.

0 favorites 0 likes
#llm-benchmarking

From Global Benchmarks to Local Evaluations: Benchmarking LLMs for the German Public Sector

arXiv cs.CL · 2026-08-19 Cached

This paper introduces MÖVE, a holistic evaluation framework for LLMs in the German public sector, examining governance dimensions like energy consumption, provider transparency, and knowledge of German-party positions, revealing trade-offs that necessitate context-specific model selection.

0 favorites 0 likes
#llm-benchmarking

When Stories Evolve: Benchmarking LLM Storytelling Across Agent Architectures in Open-Ended World Simulations

arXiv cs.CL · 2026-08-18 Cached

This paper introduces WSE-bench, a process benchmark for evaluating LLMs in open-ended world simulations, separately assessing sustained generation, canonical coherence, and meaningful development.

0 favorites 0 likes
#llm-benchmarking

SemPlan: Benchmarking Structured Semantic Planning for LLM-Based Queries over Enterprise Data

arXiv cs.AI · 2026-08-17 Cached

SemPlan is a benchmark for evaluating structured semantic planning in LLM-based queries over enterprise data, comparing four architectures using a synthetic bilingual dataset of 1,800 cases.

0 favorites 0 likes
#llm-benchmarking

Format Sensitivity Index: Token-Controlled Prompt Wrapper Robustness and Schema Compliance in LLM Benchmarking

arXiv cs.AI · 2026-07-14 Cached

This paper introduces the Format Sensitivity Index (FSI) and Parseability Sensitivity Index (PSI) to quantify how much LLM accuracy varies under different prompt wrappers. Through 140,000 generations across models and tasks, it shows that wrapper choice can drastically affect scores, with parseability failures being a key driver.

0 favorites 0 likes
#llm-benchmarking

Added direct model downloads right from the UI in Anubis OSS - if anyone would help test that would be great

Reddit r/LocalLLaMA · 2026-05-26

Anubis OSS, an Apple Silicon Mac app for benchmarking local LLMs, now supports direct model downloads from the UI via a 'Browse Models' button that pulls from ollama.com library. The developer is seeking testers to confirm installation and functionality.

0 favorites 0 likes
#llm-benchmarking

Energy per Successful Goal: Goal-Level Energy Accounting for Agentic AI Systems

arXiv cs.AI · 2026-05-25 Cached

This paper proposes A-LEMS, a framework that redefines AI energy accounting from per-inference to Energy per Successful Goal (EpG), and introduces the Orchestration Overhead Index (OOI) to measure energy costs of multi-step orchestration in agentic systems. Empirical results show agentic workflows consume 4.33× higher mean energy per goal than linear baselines, but OOI can invert for tool-augmented tasks, demonstrating goal-level accounting is necessary.

0 favorites 0 likes
#llm-benchmarking

Built a 10-agent pipeline for portfolio construction — macro, screener, 6 analysts, orchestrator, constructor — runs across 6 LLM providers

Reddit r/AI_Agents · 2026-05-11

1rok is a TypeScript framework that enables running multi-agent portfolio construction pipelines across multiple LLM providers to benchmark their performance on financial tasks like stock selection and position sizing.

0 favorites 0 likes
#llm-benchmarking

@no_stp_on_snek: mrcr v2 8-needle at 1m, open weights stack, single rented mi300x. longctx directional 0.688 (n=30, mass-val rerun pendi…

X AI KOLs Following · 2026-05-08 Cached

Shares early benchmark scores and evaluation metrics for an open-weight model stack run on a single AMD MI300X, noting competitive performance against closed-source alternatives.

0 favorites 0 likes
#llm-benchmarking

Qwen3.5-27B, Qwen3.5-122B, and Qwen3.6-35B on 4x RTX 3090 — MoEs struggle with strict global rules

Reddit r/LocalLLaMA · 2026-04-20

A user benchmarks three Qwen models (Qwen3.5-27B dense, Qwen3.5-122B-A10B MoE, Qwen3.6-35B-A3B MoE) on 4x RTX 3090 GPUs under real agentic workloads, finding that MoE models consistently underperform the dense 27B at following strict global rules despite speed advantages, with the Qwen3.6-35B leading in generation throughput.

0 favorites 0 likes
← Back to home

Submit Feedback