Tag
EnSiTa is a trilingual multi-domain parallel dataset and benchmark for English, Sinhala, and Tamil, featuring human post-edited training data and extensive experiments on domain-specific machine translation to address low-resource language challenges.
JEPA-Anything presents a domain-agnostic framework based on orthogonal predictive factorization for learning predictive models across diverse systems like vision, biology, and control, with demonstrated improvements and experimental validation.
Odyssey-3 is a new foundation world model that can control robots, cars, drones, and virtual worlds by reusing world knowledge across machines with minimal experience.
TurnBench introduces a multi-domain benchmark for assessing turn-taking dynamics in spoken dialogue, featuring a hand-labeled corpus and standardized evaluation protocols for end-of-turn and interruption detection.
This paper diagnoses intra-adapter contention in MoE+LoRA fine-tuning and introduces SpawnLoRA to dynamically add sub-adapters, reducing negative transfer across domains.
FirstPass is a large-scale peer review dataset from Nature Communications, covering multiple scientific domains and multi-round dialogues to improve AI models for scientific judgment.
Flower Hub is a reproducible benchmarking platform for federated learning that enables execution and evaluation across both simulation and deployment runtimes.
This paper proposes LT-MKT, a method for multi-domain knowledge tracing that incorporates cognitive load and knowledge transfer using large language models to construct a hierarchical graph, achieving state-of-the-art performance on real-world datasets.
This paper proposes PAMT, a process-aligned reinforcement learning framework for multi-domain machine translation that combines domain-aware long chain-of-thought supervision with step-level process rewards to improve domain-sensitive translation decisions.
This paper introduces Translation with Thought (TwT), a resource-rational framework for multi-domain machine translation that adaptively modulates reasoning effort based on input difficulty, trained via supervised fine-tuning on difficulty-aware reasoning traces and reinforcement learning. TwT-7B and TwT-14B outperform larger SOTA reasoning models while reducing token usage by 32–60%.
This paper introduces Tencent WorkBuddy Bench, a multi-domain evaluation suite for coding agents designed to resist data contamination by reverse-engineering tasks from real commits and business scenarios, covering Code, Web, Office, and Security domains.
Relay-Bench is a new benchmark designed to evaluate large language models on reasoning chains that require knowledge across multiple domains.
The paper presents a scalable framework for multi-domain dialogue state tracking using BERT, achieving zero-shot generalization and improving performance on the SGD dataset.
Qwen releases Qwen-AgentWorld-35B-A3B, a native language world model that simulates agentic environments across seven domains via long chain-of-thought reasoning. The model is trained with a three-stage pipeline and supports MCP, Search, Terminal, SWE, Android, Web, and OS interactions.
LOGOS is a scientific generative language model that encodes diverse scientific objects and spatial interactions as token sequences, enabling a unified autoregressive framework for tasks across natural sciences. Models at 1B, 3B, and 8B parameters show consistent performance scaling and are released to facilitate research.
Count Anything is a generalist model for text-guided object counting that unifies multiple domains, supported by the new CLOC dataset with 220K images across six visual domains. It achieves strong accuracy and multi-domain generalization.
This paper introduces DoRA-RBAC, a framework for composing LLM adapters, and tests whether geometry-aware merging improves multi-domain performance. Results show no consistent advantage over standard averaging, suggesting adapter interference is not primarily driven by parameter-space geometry.
Arbor is an AI framework for autonomous scientific research that uses a coordinator, executors, and a persistent hypothesis tree to iteratively improve research outcomes across multiple domains, achieving strong results on six real research tasks.
SoCRATES introduces a realistic multi-domain benchmark for evaluating proactive LLM mediators, showing that top models resolve only about one-third of the consensus gap in conflict resolution.
This paper proposes a local perturbation theory to explain cross-domain interference in multi-domain RL for LLMs, showing that interference is driven by a second-order damage term in a low-dimensional conflict subspace, and demonstrates that brief domain refresh or training-free rollback can selectively recover lost capabilities.