Tag
PICasso is an AI-enabled framework for autonomous optimization of silicon photonic devices from natural-language specifications, demonstrating significant improvements in design satisfaction and loss reduction through LLMs and simulation feedback.
This paper introduces MÖVE, a holistic evaluation framework for LLMs in the German public sector, examining governance dimensions like energy consumption, provider transparency, and knowledge of German-party positions, revealing trade-offs that necessitate context-specific model selection.
This paper introduces WSE-bench, a process benchmark for evaluating LLMs in open-ended world simulations, separately assessing sustained generation, canonical coherence, and meaningful development.
SemPlan is a benchmark for evaluating structured semantic planning in LLM-based queries over enterprise data, comparing four architectures using a synthetic bilingual dataset of 1,800 cases.
This paper introduces the Format Sensitivity Index (FSI) and Parseability Sensitivity Index (PSI) to quantify how much LLM accuracy varies under different prompt wrappers. Through 140,000 generations across models and tasks, it shows that wrapper choice can drastically affect scores, with parseability failures being a key driver.
Anubis OSS, an Apple Silicon Mac app for benchmarking local LLMs, now supports direct model downloads from the UI via a 'Browse Models' button that pulls from ollama.com library. The developer is seeking testers to confirm installation and functionality.
This paper proposes A-LEMS, a framework that redefines AI energy accounting from per-inference to Energy per Successful Goal (EpG), and introduces the Orchestration Overhead Index (OOI) to measure energy costs of multi-step orchestration in agentic systems. Empirical results show agentic workflows consume 4.33× higher mean energy per goal than linear baselines, but OOI can invert for tool-augmented tasks, demonstrating goal-level accounting is necessary.
1rok is a TypeScript framework that enables running multi-agent portfolio construction pipelines across multiple LLM providers to benchmark their performance on financial tasks like stock selection and position sizing.
Shares early benchmark scores and evaluation metrics for an open-weight model stack run on a single AMD MI300X, noting competitive performance against closed-source alternatives.
A user benchmarks three Qwen models (Qwen3.5-27B dense, Qwen3.5-122B-A10B MoE, Qwen3.6-35B-A3B MoE) on 4x RTX 3090 GPUs under real agentic workloads, finding that MoE models consistently underperform the dense 27B at following strict global rules despite speed advantages, with the Qwen3.6-35B leading in generation throughput.