Google DeepMind Unifies Seven Tracking Models in TapNet
Google DeepMind has consolidated its complete Tracking Any Point ecosystem into a single open-source repository, offering unified code and weights for seven video tracking architectures.

Google DeepMind has released TapNet, a consolidated GitHub repository that unifies its entire suite of Tracking Any Point research under an open-source Apache 2.0 license. The trending project, which has accumulated roughly 2,200 GitHub stars, replaces multiple standalone codebases with a single maintained repository. It incorporates seven distinct architectures: TAP-Net, TAPIR, BootsTAPIR, TAPNext, TAPNext++, RoboTAP, and TRAJAN. Pretrained weights for these systems are available in both JAX and PyTorch via Hugging Face.
To help developers quickly deploy and evaluate these systems, the repository ships with eleven Colab notebooks supporting offline, online, and clustering workflows, alongside evaluation setups for the TAP-Vid, RoboTAP, and TAPVid-3D benchmarks. Given any input video and specified query points, these models compute frame-by-frame 2D coordinates while predicting visibility states to track points across occlusions, deformable surfaces, fabric, masonry, skin, and robotic grippers.
The updated library includes notable performance milestones across different hardware footprints. Among them, TAPNext++ reaches a 67.0% Average Jaccard score on the DAVIS benchmark at 512x512 resolution, maintaining stable tracking 40x longer than prior iterations. For real-time mobile deployment, Causal TAPIR processes incoming frames at approximately 17 frames per second on a 2018 mobile GPU, accessible via a live webcam demonstration provided in the repo.
This unified repository streamlines development for practitioners building solutions in robot imitation, video editing, dynamic 3D reconstruction, motion analysis, and generative-video evaluation. Rather than navigating fragmented implementations, engineers can now directly select between offline tracking over whole clips or causal processing for streaming inputs based on their target framework, latency constraints, and operational horizons.
This is our own summary of reporting by AlphaSignal



