Tag
ViDiHand leverages pretrained video diffusion model representations to reconstruct 4D hand motion directly from egocentric video frames, outperforming existing methods on ARCTIC, HOT3D, and HOI4D without detectors or optimization.
This paper introduces Internal Coherence Maximization (ICM) to generate persona-specific examples for aligning AI with diverse human values without human supervision, demonstrating that coherent examples generalize better across benchmarks.
This paper demonstrates that fine-tuned AI text detectors amplify a pretrained typicality axis rather than learning an AI-vs-human boundary, with raw encoder projections often matching or exceeding fine-tuned performance.
MC-RFM proposes a novel Riemannian flow-matching framework for few-shot adaptation that models feature displacement on a mixed-curvature manifold combining hyperbolic and Euclidean spaces, outperforming existing methods across multiple visual recognition benchmarks.
This paper presents the Transformers library, an open-source library that provides state-of-the-art Transformer architectures and pretrained models under a unified API, designed to make advanced natural language processing accessible to the broader machine learning community.