Tag
This paper introduces a taxonomy of 9 model selection algorithms for multi-LLM collaboration, showing that capability-aware selection strategies outperform random or heuristic team assembly by up to 36.1% across math, coding, QA, and reasoning tasks.
This paper presents EVIL-Detect, a multi-signal ensemble framework with conflict-aware fusion for detecting LLM-generated, refined, and human-written Chinese text in the NLPCC 2026 Shared Task 6, achieving first place with a macro-F1 of 0.8888.
This paper builds a multi-scale stacking ensemble for credit risk scoring and audits LLM-generated explanations, finding that ranking gains are real but small while the narrative explanations are often unfaithful, with SHAP and LIME agreeing on important features but not their order or sign.
LymphoSAT, an ensemble of 126 specialized solvers generated with LLM assistance, won the SAT Competition 2026, demonstrating domain-specific hyperspecialization as a new approach to SAT solving.
This paper proposes a domain-knowledge-free metacognitive layer for fusing multiple pre-trained ViT-based perception models, using label vector pools and consistency-based abduction. It matches majority-vote baselines on clean data and is particularly robust against coordinated label-flipping attacks.
This paper proposes a framework that ensembles the reasoning structures of multiple LLMs by weighted merging of extracted Directed Acyclic Graphs (DAGs), enabling consensus reasoning with improved accuracy and interpretability across several benchmarks.
This paper introduces an uncertainty-aware trust estimation method for aggregating predictions from multiple LLMs, adapting structured expert judgment with Cooke-style log weighting to penalize overconfident incorrect predictions. Evaluations on MMLU and MMLU-Pro show that this approach achieves superior accuracy-reliability balance under heterogeneous and contaminated expert panels.
A user highlights a finding from Hugh Madden's writeup on multi-agent systems: even a strong arbiter (GPT-5.5) can be biased by seeing weaker agents' outputs first, collapsing from ~98% solo accuracy to 7/9.
This paper presents LV-ROVER, a multi-stream Tesseract ensemble for Maltese OCR, achieving a 70% reduction in character error rate through synthetic data training and post-processing, addressing the challenges of low-resource OCR for Maltese.
This paper applies ensemble machine learning models (Random Forest, Gradient Boosting, XGBoost, Extra Trees) to detect cirrhosis in hepatitis C patients using 28 features from 2038 Egyptian patients. The Extra Trees model achieved 96.92% accuracy with only 16 features, outperforming other models.
This paper evaluates deep learning models (LSTM, TCN, Transformer) on the WESAD dataset for multimodal emotion recognition from physiological signals, showing that an ensemble achieves 98.91% accuracy.
A reflection arguing that in multi-model setups, the consensus output is less valuable than the disagreements, which reveal genuinely contested parts of a problem. The post questions whether consensus should be the goal and how to distinguish productive disagreement from noise.
This paper compares multiple machine learning and transformer models for sentiment classification on movie reviews, finding RoBERTa achieves 93.02% accuracy, and a soft voting ensemble improves performance.
This paper presents the winning system for SemEval-2026 Task 8's generation subtask, using a heterogeneous ensemble of seven LLMs with dual prompting strategies and a GPT-4o-mini judge to select the best response. The system achieved first place with a conditioned harmonic mean of 0.7827, outperforming all baselines and demonstrating the value of model diversity.