Encode any input
Images, videos, and multi-view streams share one context, without an imposed input order.
Capabilities
Photo collections, moving-camera videos, and synchronized camera streams provide complementary evidence of geometry and motion. ARROW reconstructs and tracks points across these inputs without requiring a temporal ordering of the observations.
Paper
Dynamic scenes may be captured by a moving camera, multiple video streams, or images taken at different times. These observations reveal complementary aspects of scene geometry and motion, yet bringing them together requires establishing correspondence across viewpoints, capture times, and visibility changes. We introduce ARROW, a feed-forward model that unifies 3D reconstruction and 3D point tracking from arbitrary image sets. At its core is a novel order-invariant querying approach, which allows the association of queries with observations across arbitrary inputs. We show that exposing the model to more diverse sets of inputs during training results in improved task performance. Moreover, the resulting model is capable of generalization to a wider range of tasks including multi-view tracking. Trained with this strategy, ARROW establishes a new state of the art in 3D tracking on WorldTrack and TAPVid-3D and outperforms dedicated multi-view trackers on an adapted RGB-only MVTracker benchmark, while remaining competitive across 3D reconstruction tasks.
At a glance
Reconstruct and track across image sets, videos, multi-view streams, or any combination through the same rich scene representation.
Predicted observation identities replace learned timestep embeddings, enabling meaningful cross-attention without explicit order and tracking beyond fixed temporal horizons.
Dropping temporal inductive bias broadens supervision across viewpoints and distant times, improving tracking performance.
Approach
A query identifies a pixel in a source image, a target observation's capture time, and a reference camera's coordinate system. ARROW predicts that physical point's 3D position using the shared context of all input observations.
Images, videos, and multi-view streams share one context, without an imposed input order.
Predicted identities associate each query with the intended observations in the global context.
Query any source point, target observation, and reference view for sparse tracking or dense reconstruction.
Examples
Explore predictions from the same observations, with ARROW as the reference.
ARROW prediction not exported for this scene yet.
No prediction exported for this method yet.
Settings
Evaluation
ARROW sets a new state of the art in 3D tracking on WorldTrack and TAPVid-3D, and outperforms dedicated multi-view trackers on an adapted RGB-only MVTracker benchmark. Reconstruction and camera pose estimation remain competitive. Longer bars are always better; tracking and video-depth plots use an adjusted range to show differences more clearly. Labels show the original metric values. See the paper for full comparisons and training ablations.
Compare ARROW, OmniX, 4RC, V-DPM and OpenD4RT across tracking, camera pose, reconstruction and depth. Hover or tap a spoke or dataset label for the original numbers. Select a method in the legend to highlight it; select it again to restore all methods.
One metric per task and dataset (APD All, ATE, Chamfer and δ₁). Each spoke is normalized to the best displayed method: score / best for ↑, best / error for ↓. A reversed-log scale highlights differences near the best score; missing values remain gaps. Source: paper evaluation tables.
Average percentage of points within distance thresholds (APD), across all tracks. Higher is better.
The mean weights all four datasets equally; ARROW also has 21% lower mean EPE than UniQuery4R. Full results: Table 1 in the paper.
Fraction of pixels within a 1.25× depth error threshold. Higher is better.
Values from the video-depth evaluation table. All reported baselines are shown.
Mean bidirectional Chamfer distance in metres. Lower is better.
Values from the 3D reconstruction evaluation table. All reported baselines are shown.
Absolute trajectory error after similarity alignment. Lower is better.
Values from the camera-pose evaluation table. All reported baselines are shown.
Reference
@article{fradlin2026arrow,
title = {{ARROW}: Arbitrary Reconstruction and Tracking of 4D Observations in the Wild},
author = {Fradlin, Ilya and Schmidt, Christian and Piekenbrinck, Jens and Knaebel, Karim and Martin Garcia, Gonzalo and Leibe, Bastian},
year = 2026,
journal = {arXiv preprint arXiv:2610.01314},
}