HuggingFace

Articles from HuggingFace

Cards List

Altworld/Hemmingway-1

Hugging Face Models Trending · 3d ago Cached

Hemmingway-1 is a 27B-parameter open-source AI model specialized for everyday writing tasks, outperforming leading models on benchmarks for human-like communication.

0 favorites 0 likes

Grounded Action Model: 3D Grounding as a Foundation for Robotics

Hugging Face Daily Papers · 3d ago Cached

This paper proposes Grounded Action Models (GAMs), a new paradigm for robot foundation models that integrates 3D grounding, achieving state-of-the-art performance on manipulation tasks.

0 favorites 0 likes

Mira-Scene: Pixel-Aligned Layouts for Generative 3D Scene

Hugging Face Daily Papers · 3d ago Cached

Mira-Scene introduces a compositional 3D scene reconstruction framework using pixel-aligned canonical coordinate maps for accurate object layouts, achieving significant improvements in layout accuracy over existing methods.

0 favorites 0 likes

One to More, More to One: Category-Aware Iterative Expert Training for Software Engineering Agents

Hugging Face Daily Papers · 3d ago Cached

This paper introduces a category-aware expert training framework for software engineering agents to mitigate uneven progress across task categories, using iterative training and multi-teacher distillation, with significant performance gains on Pro-618 and SWE-bench Multilingual benchmarks.

0 favorites 0 likes

Why Do Video Diffusion Models Violate Physics? Unveiling the Flaws in Attention Mechanisms

Hugging Face Daily Papers · 3d ago Cached

This paper presents an interpretability study on video diffusion models, revealing that Rotary Position Embedding (RoPE) induces excessive spatial attention decay, causing physics violations, and proposes a lightweight architectural modification to enhance physical coherence in generated videos.

0 favorites 0 likes

convaiinnovations/laya-multilingual

Hugging Face Models Trending · 4d ago Cached

Laya Multilingual is a multilingual AI decision model that provides typed answers with probabilities across 100+ languages in a single forward pass, improving accuracy over English-only models.

0 favorites 0 likes

Towards Full Pipeline FP8 Reinforcement Learning for LLMs

Hugging Face Daily Papers · 4d ago Cached

The paper proposes Calibrated Clipping to stabilize FP8 quantization in reinforcement learning for LLMs by aligning clipping bounds with high-precision distributions, eliminating entropy surges and restoring performance.

0 favorites 0 likes

UltraTex: Unleashing 2K Multi-View Diffusion for 3D Texturing

Hugging Face Daily Papers · 4d ago Cached

UltraTex is an efficient framework for high-resolution multi-view diffusion-based 3D texturing, introducing techniques to reduce redundancy and achieve significant speedups in training and inference.

0 favorites 0 likes

OmniEdu: Open Foundation Models for Learning and Teaching

Hugging Face Daily Papers · 4d ago Cached

OmniEdu introduces an open family of foundation models for K-12 education, trained on a curated corpus to improve problem-solving, curriculum grounding, and pedagogical tutoring capabilities.

0 favorites 0 likes

Transferring the Intelligence of VLMs to Robotic Control

Hugging Face Daily Papers · 4d ago Cached

This paper presents RoboDawn, a method to transfer Vision-Language Model intelligence to robotic control, achieving state-of-the-art results on benchmarks with zero-shot and one-shot learning and successful real-world applications.

0 favorites 0 likes

convaiinnovations/laya

Hugging Face Models Trending · 5d ago Cached

Laya is an open-source non-autoregressive decision model that provides typed answers with calibrated probabilities, designed for tasks like email triage and conversational AI, showing significant performance improvements over existing models.

0 favorites 0 likes

HuRo: Robotizing Human Videos for Scalable VLA Pretraining

Hugging Face Daily Papers · 5d ago Cached

This paper presents HuRo, a pipeline for robotizing human videos to create scalable VLA pretraining data, showing significant improvements in task completion and robustness on real-world manipulation tasks.

0 favorites 0 likes

From Pretraining to Proficiency: Real-World Subtask RL for Long-Horizon Manipulation with Minimal Human Intervention

Hugging Face Daily Papers · 5d ago Cached

The paper presents PARTS, a real-world subtask reinforcement learning framework that improves long-horizon manipulation tasks by focusing on bottleneck subtasks with minimal human intervention, achieving higher success rates in experiments.

0 favorites 0 likes

APort Vault: Benchmarking AI Agent Payment Authorization with the Open Agent Passport

Hugging Face Daily Papers · 5d ago Cached

This paper presents APort Vault, a benchmark for evaluating payment authorization in AI agents, featuring over 225,000 evaluations across 14 models to test security policies and the Open Agent Passport specification.

0 favorites 0 likes

Gricea: An Open Science Platform for Conversational AI Research

Hugging Face Daily Papers · 5d ago Cached

Gricea is an open-science platform for conversational AI research that enables configurable and deployable research artifacts to support replication, extension, and cumulative knowledge building.

0 favorites 0 likes

OmniVChat: Synthesizing, Benchmarking, and Training for Native Audio-Visual Dialogue

Hugging Face Daily Papers · 5d ago Cached

This paper introduces OmniVChat, a task for native audio-visual dialogue, and presents a data engine, benchmark, and reinforcement learning method to train and evaluate omni models, demonstrating improvements on synthesized and human-recorded data.

0 favorites 0 likes

MintAct: A Unified Visual Agent for Digital Environments

Hugging Face Daily Papers · 5d ago Cached

MintAct is a family of vision-language models that unifies UI grounding, multi-step navigation across mobile, desktop, and web, and visual tool use, achieving state-of-the-art performance on various benchmarks.

0 favorites 0 likes

OmniVBench: A Benchmark and Large-Scale Dataset for Omni Reference-to-Video Generation

Hugging Face Daily Papers · 5d ago Cached

The paper introduces OmniVBench, a comprehensive benchmark for omni reference-to-video generation, and the Omni-R2VDataset, a large-scale training dataset, to evaluate and improve R2V models.

0 favorites 0 likes

RecreationWorld: Scalable and Verifiable Environments for Hybrid Computer-Use Agents

Hugging Face Daily Papers · 5d ago Cached

RecreationWorld introduces a scalable and verifiable framework for hybrid computer-use agents, providing environments across five platforms and a benchmark for evaluation.

0 favorites 0 likes

GraphSkillEvo: Evolutionary Optimization of Graph-Structured Agent Skills

Hugging Face Daily Papers · 5d ago Cached

GraphSkillEvo is an evolutionary optimization framework that represents agent skills as graph-structured artifacts to improve LLM performance, outperforming baselines on multiple benchmarks.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback