TAPNext: Tracking Any Point (TAP) as Next Token Prediction
Summary
TAPNext reframes Tracking Any Point (TAP) as sequential masked token decoding, achieving state-of-the-art online and offline tracking performance without tracking-specific inductive biases, with tracking heuristics emerging naturally from end-to-end training.
View Cached Full Text
Cached at: 10/03/26, 09:55 AM
Paper page - TAPNext: Tracking Any Point (TAP) as Next Token Prediction
Source: https://huggingface.co/papers/2504.05579
Abstract
TAPNext addresses video tracking by framing it as sequential masked token decoding, achieving state-of-the-art performance while eliminating tracking-specific biases.
Tracking Any Point (TAP) in a video is a challenging computer vision problem with many demonstrated applications in robotics, video editing, and 3D reconstruction. Existing methods for TAP rely heavily on complex tracking-specific inductive biases and heuristics, limiting their generality and potential for scaling. To address these challenges, we present TAPNext, a new approach that casts TAP assequential masked token decoding. Our model is causal, tracks in a purely online fashion, and removes tracking-specific inductive biases. This enables TAPNext to run with minimal latency, and removes the temporal windowing required by many existing state of art trackers. Despite its simplicity, TAPNext achieves a new state-of-the-art tracking performance among both online and offline trackers. Finally, we present evidence that many widely used tracking heuristics emerge naturally in TAPNext through end-to-end training.
View arXiv pageView PDFProject pageGitHub2.09kAdd to collection
Get this paper in your agent:
hf papers read 2504\.05579
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2504.05579 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2504.05579 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2504.05579 in a Space README.md to link it from this page.
Collections including this paper2
Similar Articles
Next-Latent Prediction Transformers [R]
Microsoft Research introduces Next-Latent Prediction (NextLat), a self-supervised method that trains transformers to predict their own next latent state, enabling compact world models for reasoning and planning and achieving up to 3.3x faster inference via self-speculative decoding.
AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction
This paper introduces AdaMTP, an adaptive training paradigm for multi-token prediction that dynamically aligns prediction horizons with sequence predictability using entropy-based segmentation, consistently outperforming standard MTP on math, code, and general benchmarks across three LLM backbones.
@FinanceYF5: Next token prediction is short-sighted. What if the Transformer learns to predict its own next hidden state? Jayden Teoh proposes Next-Latent Prediction (NextLat): a self-supervised learning method that teaches the Transformer to form...
Jayden Teoh proposes Next-Latent Prediction (NextLat), a self-supervised learning method that teaches the Transformer to learn to predict the next hidden state, thereby forming a compact world model for reasoning and planning, and achieves up to 3.3x inference speedup through self-speculative decoding.
PTP: Previous-Token Prediction based LLM Inversion for Near-Exact Prompt Reconstruction
The paper presents PTP, a functional approach to LLM inversion that trains an inverse language model from scratch using previous-token prediction on synthetic data from a target black-box LLM, enabling near-exact prompt reconstruction from responses and outperforming prior work.
NITP: Next Implicit Token Prediction for LLM Pre-training
Next Implicit Token Prediction (NITP) enhances language model pre-training by adding dense continuous supervision in representation space, improving generalization and performance across model sizes with minimal computational overhead.