ARROW: Arbitrary Reconstruction and Tracking of 4D Observations in the Wild

Ilya Fradlin Christian Schmidt Jens Piekenbrinck Karim Knaebel Gonzalo Martin Garcia Bastian Leibe

RWTH Aachen University

Capabilities

Reconstruction and tracking from any RGB observations

Photo collections, moving-camera videos, and synchronized camera streams provide complementary evidence of geometry and motion. ARROW reconstructs and tracks points across these inputs without requiring a temporal ordering of the observations.

Paper

Abstract

Dynamic scenes may be captured by a moving camera, multiple video streams, or images taken at different times. These observations reveal complementary aspects of scene geometry and motion, yet bringing them together requires establishing correspondence across viewpoints, capture times, and visibility changes. We introduce ARROW, a feed-forward model that unifies 3D reconstruction and 3D point tracking from arbitrary image sets. At its core is a novel order-invariant querying approach, which allows the association of queries with observations across arbitrary inputs. We show that exposing the model to more diverse sets of inputs during training results in improved task performance. Moreover, the resulting model is capable of generalization to a wider range of tasks including multi-view tracking. Trained with this strategy, ARROW establishes a new state of the art in 3D tracking on WorldTrack and TAPVid-3D and outperforms dedicated multi-view trackers on an adapted RGB-only MVTracker benchmark, while remaining competitive across 3D reconstruction tasks.

At a glance

Key-Takeaways

  1. 01

    One representation.
    Any observation mix.

    Reconstruct and track across image sets, videos, multi-view streams, or any combination through the same rich scene representation.

  2. 02

    Predict identities.
    Query in any order.

    Predicted observation identities replace learned timestep embeddings, enabling meaningful cross-attention without explicit order and tracking beyond fixed temporal horizons.

  3. 03

    Less temporal bias.
    Better tracking.

    Dropping temporal inductive bias broadens supervision across viewpoints and distant times, improving tracking performance.

Approach

Method

A query identifies a pixel in a source image, a target observation's capture time, and a reference camera's coordinate system. ARROW predicts that physical point's 3D position using the shared context of all input observations.

Three-panel comparison: order-dependent queries match ordered video features; fixed query embeddings fail with unordered image features; ARROW uses content-derived identity tokens to associate queries with observations.
From indexed to input-dependent correspondence. (a) Order-aware global context enables position-based associations. (b) Order-invariant global context exposes no fixed order to exploit. (c) Content-derived identity embeddings enable order-invariant associations.
ARROW architecture: a multi-view encoder produces shared context and observation identity tokens; a query encoder combines a source RGB patch, pixel coordinates, and source, target, and reference identities; independent queries decode 3D predictions.
ARROW architecture. A multi-view encoder produces shared image context and one content-derived ID token per observation, without temporal encodings. Each query independently attends to this context, so one encoding serves both sparse tracking and dense reconstruction.
01

Encode any input

Images, videos, and multi-view streams share one context, without an imposed input order.

02

Compose queries with identities

Predicted identities associate each query with the intended observations in the global context.

03

Flexible inputs, flexible predictions

Query any source point, target observation, and reference view for sparse tracking or dense reconstruction.

Examples

Compare geometry and motion

Explore predictions from the same observations, with ARROW as the reference.

Evaluation

Quantitative results

ARROW sets a new state of the art in 3D tracking on WorldTrack and TAPVid-3D, and outperforms dedicated multi-view trackers on an adapted RGB-only MVTracker benchmark. Reconstruction and camera pose estimation remain competitive. Longer bars are always better; tracking and video-depth plots use an adjusted range to show differences more clearly. Labels show the original metric values. See the paper for full comparisons and training ablations.

Four tasks

Performance across tasks

Compare ARROW, OmniX, 4RC, V-DPM and OpenD4RT across tracking, camera pose, reconstruction and depth. Hover or tap a spoke or dataset label for the original numbers. Select a method in the legend to highlight it; select it again to restore all methods.

25 50 75 90 100 3D tracking · APD (All) ↑ Camera pose · ATE ↓ 3D reconstruction · Chamfer ↓ Video depth · δ₁ ↑ Aria Digital Twin Dynamic Replica PointOdyssey Panoptic Studio Bonn Sintel PStudio NRGBD 7-Scenes Bonn KITTI 7-Scenes

One metric per task and dataset (APD All, ATE, Chamfer and δ₁). Each spoke is normalized to the best displayed method: score / best for ↑, best / error for ↓. A reversed-log scale highlights differences near the best score; missing values remain gaps. Source: paper evaluation tables.

WorldTrack · APD (All) ↑

3D tracking across four datasets

+4.38 pp

Average percentage of points within distance thresholds (APD), across all tracks. Higher is better.

Four-dataset mean1 / 5
ARROW88.14%
UniQuery4R83.76%
SM4RT82.36%
OmniX81.72%
V-DPM80.05%
4RC79.73%
SpatialTracker-V275.61%
OpenD4RT73.78%
St4RTrack71.36%
Point4D70.81%

The mean weights all four datasets equally; ARROW also has 21% lower mean EPE than UniQuery4R. Full results: Table 1 in the paper.

Video depth · δ₁ ↑

Video depth across three datasets

+0.6 pp

Fraction of pixels within a 1.25× depth error threshold. Higher is better.

Three-dataset mean1 / 4
ARROW96.5%
DA395.9%
SM4RT95.3%
V-DPM95.0%
Pi394.9%
4RC94.6%
VGGT-Ω94.5%
VGGT94.4%
OmniX92.8%
OpenD4RT90.6%

Values from the video-depth evaluation table. All reported baselines are shown.

3D reconstruction · Chamfer ↓

Reconstruction quality across datasets

2.0% lower

Mean bidirectional Chamfer distance in metres. Lower is better.

Two-dataset mean1 / 3
ARROW0.0375 m
VGGT-Ω0.0382 m
OmniX0.0387 m
DA30.0442 m
VGGT0.0484 m
4RC0.0495 m
Pi30.0510 m
V-DPM0.0568 m
SM4RT0.0956 m
OpenD4RT0.3290 m

Values from the 3D reconstruction evaluation table. All reported baselines are shown.

Camera pose · ATE ↓

Camera pose across three datasets

14.0% lower

Absolute trajectory error after similarity alignment. Lower is better.

Three-dataset mean1 / 4
ARROW0.0470 m
OmniX0.0547 m
Pi30.0785 m
SM4RT0.0924 m
4RC0.0929 m
DA30.1153 m
V-DPM0.1183 m
VGGT-Ω0.1351 m
VGGT0.2039 m
OpenD4RT0.6064 m

Values from the camera-pose evaluation table. All reported baselines are shown.

Reference

Citation

BibTeX
@article{fradlin2026arrow,
    title         = {{ARROW}: Arbitrary Reconstruction and Tracking of 4D Observations in the Wild},
    author        = {Fradlin, Ilya and Schmidt, Christian and Piekenbrinck, Jens and Knaebel, Karim and Martin Garcia, Gonzalo and Leibe, Bastian},
    year          = 2026,
    journal       = {arXiv preprint arXiv:2610.01314},
}