frozen-backbone

Tag

Cards List
#frozen-backbone

Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio

arXiv cs.CL · 2026-07-22 Cached

Fusion Embedding introduces a family of models that add audio to a frozen vision-language embedding backbone, enabling a unified space for text, image, video, and audio retrieval. The models train only lightweight adapters and achieve audio-image retrieval without paired audio-visual data.

0 favorites 0 likes
#frozen-backbone

baseten/GLM-5.2-Vision-NVFP4

Hugging Face Models Trending · 2026-07-20 Cached

Baseten releases GLM-5.2-Vision, a vision-language model that adds MoonViT vision encoder to GLM-5.2 via a trained PatchMerger projector, keeping the text backbone and vision tower frozen. The model is quantized to NVFP4 for efficient inference on Blackwell hardware.

0 favorites 0 likes
#frozen-backbone

Is Domain Adaptation Always Helpful? A Frozen-Backbone Study of Cross-Domain Sentiment Transfer

arXiv cs.CL · 2026-07-08 Cached

This paper investigates whether explicit domain adaptation methods are beneficial for sentiment transfer when using frozen pre-trained language model backbones, finding that effectiveness depends on whether the backbone already possesses target-domain knowledge.

0 favorites 0 likes
#frozen-backbone

Time-Series Foundation Model Embeddings for Remaining Useful Life Estimation

Hugging Face Daily Papers · 2026-06-10 Cached

This paper introduces a lightweight approach for remaining useful life estimation using frozen embeddings from the Chronos-2 time-series foundation model combined with a simple regression head, achieving superior performance on industrial sensor data compared to baseline methods.

0 favorites 0 likes
#frozen-backbone

Orthrus-Qwen3-8B : up to 7.8×tokens/forward on Qwen3-8B, frozen backbone, provably identical output distribution

Reddit r/LocalLLaMA · 2026-05-15

Introduces Orthrus, a method that injects a trainable diffusion attention module into a frozen autoregressive transformer to achieve up to 7.8× tokens per forward pass and ~6× wall-clock speedup on MATH-500, with provably identical output distribution to the base Qwen3-8B model. The approach requires minimal additional parameters and training, and avoids the TTFT penalty of external drafters.

0 favorites 0 likes
← Back to home

Submit Feedback