hybrid-models

Tag

Cards List
#hybrid-models

A hybrid quantum-classical neural network for learning to route

arXiv cs.LG · 5d ago Cached

This paper investigates hybrid quantum-classical neural networks for the vehicle routing problem, finding that encoder feed-forward replacement can reduce model parameters by 56.6% while maintaining near-baseline performance for small to medium instances.

0 favorites 0 likes
#hybrid-models

Conservative Hybrid Graph Networks for Process Systems with Learned Routing

arXiv cs.LG · 6d ago Cached

The paper presents Conservative Hybrid Graph Network (CHGN), a method for modeling dynamic process systems that allows extrapolation to unseen larger graphs without retraining and improves fault prediction in industrial applications.

0 favorites 0 likes
#hybrid-models

Technical Comparative Benchmarking Study: Advanced AI Hybrid Methods for Renewable Energy Farm Optimization and Forecasting

arXiv cs.LG · 2026-08-28 Cached

A comparative benchmarking study evaluates various AI methods for renewable energy farm optimization and forecasting, showing that ensemble and hybrid approaches excel in different data scenarios.

0 favorites 0 likes
#hybrid-models

Empirical Characterization of Learning Geometry in Hybrid Quantum Forecasting Models

arXiv cs.LG · 2026-08-21 Cached

The paper empirically characterizes the learning geometry of hybrid quantum forecasting models, comparing them to classical baselines using Neural Tangent Kernel dynamics and other metrics, showing that similar generalization can emerge from different optimization trajectories.

0 favorites 0 likes
#hybrid-models

Post-training Quantization for Hybrid Iterative Generative Models

arXiv cs.LG · 2026-08-17 Cached

HyGenQ is a post-training quantization framework for hybrid iterative generative models that addresses challenges like excessive outliers and amplified anomalies, enabling 8-bit precision quantization while maintaining generation quality.

0 favorites 0 likes
#hybrid-models

Massive Activations in Hybrid Linear Attention Large Language Models: Pre-Attention Spikes and Inter-Spike Plateaus

Hugging Face Daily Papers · 2026-08-12 Cached

This paper presents the first systematic study of massive activations in hybrid linear-attention LLMs, uncovering pre-attention spikes and inter-spike plateaus governed by cancellation timing, and showing how their morphology recovers at full-attention limits.

0 favorites 0 likes
#hybrid-models

Multistage Defer Trees for Hybrid Interpretability: If at First You Can't Succeed, Tree Again

arXiv cs.LG · 2026-07-01 Cached

Introduces Multistage Defer Trees, a sequence of sparse decision trees that defer hard samples to later trees or a black box, aiming to match ensemble accuracy while keeping most predictions interpretable.

0 favorites 0 likes
#hybrid-models

Comparing Transformers and Hybrid Models at the Token Level

Lobsters Hottest · 2026-06-27 Cached

This paper analyzes token-level prediction differences between transformers and hybrid attention-recurrent models using Olmo 3 and Olmo Hybrid, finding that hybrids improve on semantic state tracking while transformers excel at n-gram copying and syntactic bracket matching.

0 favorites 0 likes
#hybrid-models

@_albertgu: Transformers are better at copying, while RNNs are better at modeling "meaning-bearing words—the nouns, verbs, & adject…

X AI KOLs Following · 2026-06-26 Cached

A thread from Ai2 compares transformer (Olmo 3) and hybrid (Olmo Hybrid) models, finding that transformers excel at copying while RNNs better model meaning-bearing words, highlighting the growing viability of hybrid architectures.

0 favorites 0 likes
#hybrid-models

@ZhihuFrontier: Half a year ago, a Zhihu contributor predicted that the next Transformer would absorb loops, recurrent state, sparse ro…

X AI KOLs Timeline · 2026-06-26 Cached

A Zhihu contributor's half-year-old prediction that the next Transformer would absorb loops, recurrent state, sparse routing, and latent reasoning is gaining relevance as Loop Engineering advances. The article explores how future Transformer architectures may evolve into hybrid models blending linear-complexity layers for background context with attention for precise reasoning, plus finer-grained sparsity and native System 2 reasoning.

0 favorites 0 likes
#hybrid-models

Which tokens does a hybrid model predict better?

Hugging Face Blog · 2026-06-25 Cached

A study comparing Olmo Hybrid and Olmo 3 transformers at the token level shows hybrid models better predict meaningful tokens like nouns/verbs, while transformers excel at copying tokens from input.

0 favorites 0 likes
#hybrid-models

PE-MHL: Physics-Encoded Modular Hybrid Layers for Scalable Learning of Complex Systems

arXiv cs.LG · 2026-06-04

This paper proposes PE-MHL, a Physics-Encoded Modular Hybrid Layer framework that incrementally refines a physics-based model with data-driven sub-models, providing theoretical convergence guarantees and outperforming monolithic networks on control benchmarks.

0 favorites 0 likes
#hybrid-models

A Systematic Evaluation of Current Architectures in Wind Power Forecasting

arXiv cs.LG · 2026-06-03 Cached

This paper presents a systematic literature review of hybrid approaches for interval wind speed forecasting, combining deep learning, modal decomposition, and statistical methods to enhance prediction accuracy and reliability.

0 favorites 0 likes
#hybrid-models

Reciprocal Co-Training (RCT): Coupling Gradient-Based and Non-Differentiable Models via Reinforcement Learning

arXiv cs.CL · 2026-04-21 Cached

Researchers from Fordham University introduce Reciprocal Co-Training (RCT), a framework that couples LLMs and Random Forest classifiers via reinforcement learning, creating an iterative feedback loop where each model improves using signals from the other. Experiments on three medical datasets show consistent performance gains for both models, demonstrating a general mechanism for integrating incompatible model families.

0 favorites 0 likes
#hybrid-models

Olmo Hybrid: From Theory to Practice and Back

arXiv cs.CL · 2026-04-20 Cached

This paper presents Olmo Hybrid, a 7B-parameter language model that combines attention and Gated DeltaNet recurrent layers, demonstrating both theoretical and empirical advantages over pure transformers. The work shows that hybrid models have greater expressivity, scale more efficiently during pretraining, and outperform comparable transformer baselines.

0 favorites 0 likes
← Back to home

Submit Feedback