ensemble

Tag

Cards List
#ensemble

You're Hired: Strategic Model Selection for LLM Collaboration

arXiv cs.CL ↗ · 2d ago Cached

This paper introduces a taxonomy of 9 model selection algorithms for multi-LLM collaboration, showing that capability-aware selection strategies outperform random or heuristic team assembly by up to 36.1% across math, coding, QA, and reasoning tasks.

0 favorites 0 likes
#ensemble

EVIL-Detect for NLPCC 2026 Shared Task 6: LLM-Generated Text Detection

arXiv cs.CL ↗ · 2026-08-12 Cached

This paper presents EVIL-Detect, a multi-signal ensemble framework with conflict-aware fusion for detecting LLM-generated, refined, and human-written Chinese text in the NLPCC 2026 Shared Task 6, achieving first place with a macro-F1 of 0.8888.

0 favorites 0 likes
#ensemble

Accurate Ensembles, Fragile Narratives: Multi-Scale Stacking and a Fidelity Audit of LLM-Generated Explanations for Credit Risk

arXiv cs.LG ↗ · 2026-08-11 Cached

This paper builds a multi-scale stacking ensemble for credit risk scoring and audits LLM-generated explanations, finding that ranking gains are real but small while the narrative explanations are often unfaithful, with SHAP and LIME agreeing on important features but not their order or sign.

0 favorites 0 likes
#ensemble

Domain-specific hyperspecialization (for SAT)

Lobsters Hottest ↗ · 2026-08-07 Cached

LymphoSAT, an ensemble of 126 specialized solvers generated with LLM assistance, won the SAT Competition 2026, demonstrating domain-specific hyperspecialization as a new approach to SAT solving.

0 favorites 0 likes
#ensemble

Adversarially Robust Abductive Fusion of Pre-trained Transformer-based Perception Models

arXiv cs.AI ↗ · 2026-08-06 Cached

This paper proposes a domain-knowledge-free metacognitive layer for fusing multiple pre-trained ViT-based perception models, using label vector pools and consistency-based abduction. It matches majority-vote baselines on clean data and is particularly robust against coordinated label-flipping attacks.

0 favorites 0 likes
#ensemble

Reasoning Consensus: Structural Ensembling of LLM Reasoning via Weighted DAG Aggregation

arXiv cs.CL ↗ · 2026-07-31 Cached

This paper proposes a framework that ensembles the reasoning structures of multiple LLMs by weighted merging of extracted Directed Acyclic Graphs (DAGs), enabling consensus reasoning with improved accuracy and interpretability across several benchmarks.

0 favorites 0 likes
#ensemble

Uncertainty-Aware Trust Estimation for Multi-LLM Systems via Structured Expert Judgement

arXiv cs.LG ↗ · 2026-07-24 Cached

This paper introduces an uncertainty-aware trust estimation method for aggregating predictions from multiple LLMs, adapting structured expert judgment with Cooke-style log weighting to penalize overconfident incorrect predictions. Evaluations on MMLU and MMLU-Pro show that this approach achieves superior accuracy-reliability balance under heterogeneous and contaminated expert panels.

0 favorites 0 likes
#ensemble

@no_stp_on_snek: if you build multi-agent or mixture-of-agents systems, read @dangerm00se's writeup. the finding that stuck with me: eve…

X AI KOLs Timeline ↗ · 2026-07-06 Cached

A user highlights a finding from Hugh Madden's writeup on multi-agent systems: even a strong arbiter (GPT-5.5) can be biased by seeing weaker agents' outputs first, collapsing from ~98% solo accuracy to 7/9.

0 favorites 0 likes
#ensemble

LV-ROVER: Multi-Stream Tesseract Voting for Maltese Paragraph OCR

arXiv cs.CL ↗ · 2026-07-02 Cached

This paper presents LV-ROVER, a multi-stream Tesseract ensemble for Maltese OCR, achieving a 70% reduction in character error rate through synthetic data training and post-processing, addressing the challenges of low-resource OCR for Maltese.

0 favorites 0 likes
#ensemble

Explainable Ensemble-Based Machine Learning Models for Detecting the Presence of Cirrhosis in Hepatitis C Patients

arXiv cs.AI ↗ · 2026-06-26 Cached

This paper applies ensemble machine learning models (Random Forest, Gradient Boosting, XGBoost, Extra Trees) to detect cirrhosis in hepatitis C patients using 28 features from 2038 Egyptian patients. The Extra Trees model achieved 96.92% accuracy with only 16 features, outperforming other models.

0 favorites 0 likes
#ensemble

Deep Temporal Modeling and Ensemble Fusion for Multimodal Emotion Recognition from Physiological Signals

arXiv cs.CL ↗ · 2026-06-16 Cached

This paper evaluates deep learning models (LSTM, TCN, Transformer) on the WESAD dataset for multimodal emotion recognition from physiological signals, showing that an ensemble achieves 98.91% accuracy.

0 favorites 0 likes
#ensemble

the more i use multiple models, the more i think "AI consensus" is a trap — the disagreement is the only part worth paying attention to

Reddit r/artificial ↗ · 2026-06-06

A reflection arguing that in multi-model setups, the consensus output is less valuable than the disagreements, which reveal genuinely contested parts of a problem. The post questions whether consensus should be the goal and how to distinguish productive disagreement from noise.

0 favorites 0 likes
#ensemble

From TF-IDF to Transformers: A Comparative and Ensemble Approach to Sentiment Classification

arXiv cs.CL ↗ · 2026-05-22 Cached

This paper compares multiple machine learning and transformer models for sentiment classification on movie reviews, finding RoBERTa achieves 93.02% accuracy, and a soft voting ensemble improves performance.

0 favorites 0 likes
#ensemble

RaguTeam at SemEval-2026 Task 8: Meno and Friends in a Judge-Orchestrated LLM Ensemble for Faithful Multi-Turn Response Generation

Hugging Face Daily Papers ↗ · 2026-05-06 Cached

This paper presents the winning system for SemEval-2026 Task 8's generation subtask, using a heterogeneous ensemble of seven LLMs with dual prompting strategies and a GPT-4o-mini judge to select the best response. The system achieved first place with a conditioned harmonic mean of 0.7827, outperforming all baselines and demonstrating the value of model diversity.

0 favorites 0 likes
← Back to home

Submit Feedback