HA-SLAM / Human Archive
Ego500 · HA-HAND · July 2026

HA-SLAM tracking natural head motion.  Watch the full three-minute segment →

Technical Report

HA-SLAM

The most robust SLAM system for egocentric data.

Problem

Open-source SLAM systems are built on the foundations of drones and autonomous vehicles. Such bodies move through space in controlled patterns characterized by smooth, planned, and predictable motion. Head-worn motion from egocentric rigs breaks all of these priors, as humans move in non-traditional patterns with fast head movements and abrupt viewpoint changes.

As a result, traditional SLAM systems lack the robust, continuous tracking required for large-scale egocentric data collection and robot learning1. Egocentric human video is a primary scaling axis for robot learning, and an accurate, continuous camera pose is what transforms raw video into useful training data by separating camera motion from motion in the world.

Today, Human Archive is announcing HA-SLAM, a state-of-the-art SLAM system built specifically for egocentric motion. Across our benchmark, HA-SLAM significantly outperforms the leading open-source alternatives.

HA-SLAM

HA-SLAM is built for one thing: high-accuracy, complete tracking under natural egocentric motion.

As one of the largest collectors of multi-camera egocentric data, Human Archive has deployed capture systems at scale across commercial and household environments worldwide. That experience exposed a fundamental limitation of traditional SLAM systems: their feature detectors collapse under natural human motion.

Across a significant body of internal evaluation and research, we found that existing frontends often generate anywhere from 0 to 50 noisy features frame to frame — far too few, and far too unstable, for the classical backend to work with. Starved of correspondences, it is forced to produce heavily biased poses or none at all. The degradation consistently stemmed from three conditions:

To solve this, we built our own custom frontend feature extraction and matching system, purpose-built for egocentric environments. Against the hand-crafted baseline, it produces an 8× increase in both feature counts and cross-frame match counts. Under the conditions that matter most — high rotational velocity and heavily blurred frames — that margin widens to a 17× increase in matches, providing the stable correspondences needed for accurate pose estimation.

We also optimized our backend and Global Bundle Adjustment pipeline, tuning keyframe selection and tracking window length specifically for egocentric motion rather than relying on parameters designed for drones and autonomous vehicles.

Setup

Hardware

HA-SLAM runs on our own head-worn capture rig: a global-shutter stereo pair with a hardware-synchronized IMU. Every system in this evaluation ran on the same recordings from that rig — the same cameras, the same inertial stream, and the same sequences. The visual and inertial streams share a hardware clock, leaving only a negligible residual offset. This synchronization is critical because, during rapid head motion (up to 236°/s), even a 1 ms timing error introduces nearly a quarter-degree of rotational misalignment. That is enough to degrade the IMU preintegration the backend depends on exactly when the visual signal is weakest.

Baselines

We evaluated HA-SLAM against five of the most robust open-source SLAM systems available, each run in its stock, best published configuration.

SystemTypeArchitecture
HA-SLAMStereo-inertialLearned feature frontend + stereo-inertial SLAM
ORB-SLAM3Stereo-inertialHand-crafted ORB features, keyframe SLAM
OKVIS2Stereo-inertialSliding-window optimization with marginalization
BasaltStereo-inertialVisual-inertial mapping, non-linear factor recovery
DPVOMonocular VOLearned patch-based tracking (no IMU)
DROID-SLAMMonocular VODense differentiable SLAM, recurrent flow (no IMU)

Eval Segments

We tested on 8 segments of natural head motion, captured indoors with millimeter-accurate motion capture ground truth. Ground truth was used only for post-hoc evaluation, never inside the estimation pipeline. The segments are not interchangeable, so we group them into three buckets by angular velocity and translational speed.

DifficultySegmentsAvg RotationPeak RotationAvg Speed
Easyseg03, seg06, seg0726–36 deg/s70–106 deg/s0.15–0.34 m/s
Mediumseg00, seg01, seg0441–50 deg/s107–153 deg/s0.18–0.49 m/s
Hardseg02, seg0572–89 deg/s192–264 deg/s0.42–0.72 m/s

Evaluations

No single number describes whether a SLAM system works. We report three, each measuring something the others cannot see.

Raw ATE. Absolute Trajectory Error measures global consistency: estimated positions compared to ground truth after rigid SE(3) alignment, capturing accumulated drift across the full session. It is computed only on frames where a system returned a pose.

Completeness. The proportion of frames where the system returned a pose at all, regardless of how accurate that pose was. A system tracking 40% of a session at excellent accuracy and one tracking 100% at the same accuracy report identical raw ATE. Completeness is what separates them.

Gap-adjusted ATE. Every frame in the session is scored. Where no pose exists, the last known pose is held frozen — exactly as a runtime would render it — and error accumulates while the world moves and the estimate does not. This folds accuracy and coverage into one number, and is the closest measure of what a downstream consumer of the pose stream actually receives.

Reading the tables An X marks a run we treat as unusable: the system returned poses for under 85% of the frames. No error figure is reported in that case, because an ATE computed over a self-selected fraction of the run describes which frames the system accepted rather than how accurately it tracked. Coverage below this line means the pose stream has gaps too large and too frequent for a downstream consumer to bridge.

Overall results

Mean values across all eight segments, for every system tested. These are averages over the complete run set, not a selected subset.

SystemCompletenessRaw ATEGap-Adjusted ATE
HA-SLAM98.8%1.43 cm1.57 cm
ORB-SLAM366.8%X21.55 cm
OKVIS280.7%X13.57 cm
Basalt99.8%7.83 cm7.85 cm
DPVO99.8%73.89 cm73.85 cm
DROID-SLAM99.8%84.57 cm84.57 cm

Segment breakdown

Every segment, every system, grouped into the three difficulty buckets. Raw ATE is marked X where coverage falls below 85%.

Raw ATE

SegmentHA-SLAMORB-SLAM3OKVIS2BasaltDPVODROID-SLAM
Easy
seg030.87 cm0.94 cmX6.91 cm17.84 cm103.28 cm
seg060.54 cmX1.44 cm8.88 cm10.40 cm104.87 cm
seg070.46 cmXX6.39 cm18.63 cm68.06 cm
Medium
seg000.96 cm0.94 cm1.73 cm4.88 cm36.57 cm114.19 cm
seg011.93 cm2.06 cm2.73 cm10.55 cm41.28 cm109.64 cm
seg040.18 cmX1.05 cm6.86 cm363.06 cm52.85 cm
Hard
seg021.21 cmX2.56 cm6.90 cm16.44 cm46.43 cm
seg055.27 cm5.37 cm6.16 cm11.26 cm86.87 cm77.22 cm

Completeness

SegmentHA-SLAMORB-SLAM3OKVIS2BasaltDPVODROID-SLAM
Easy
seg0399.4%99.5%38.0%99.8%99.8%99.8%
seg0699.8%82.3%99.7%99.8%99.8%99.8%
seg0799.1%71.1%21.8%99.8%99.8%99.8%
Medium
seg0097.2%97.4%99.8%99.8%99.8%99.8%
seg0199.0%99.3%99.4%99.8%99.8%99.8%
seg0499.2%3.9%87.3%99.8%99.8%99.8%
Hard
seg0299.1%3.6%99.7%99.8%99.8%99.8%
seg0597.8%85.7%98.3%99.8%99.8%99.8%

Gap-Adjusted ATE

SegmentHA-SLAMORB-SLAM3OKVIS2BasaltDPVODROID-SLAM
Easy
seg030.95 cm1.04 cm61.72 cm6.92 cm17.84 cm103.28 cm
seg060.61 cm22.87 cm1.47 cm8.88 cm10.40 cm104.87 cm
seg070.58 cm37.47 cm24.25 cm6.40 cm18.63 cm68.06 cm
Medium
seg001.03 cm1.05 cm1.79 cm4.90 cm36.57 cm114.19 cm
seg011.98 cm2.12 cm2.79 cm10.55 cm41.28 cm109.64 cm
seg040.72 cm26.79 cm7.67 cm6.89 cm363.06 cm52.85 cm
Hard
seg021.32 cm64.58 cm2.61 cm6.94 cm16.44 cm46.43 cm
seg055.36 cm16.45 cm6.25 cm11.30 cm86.87 cm77.22 cm

Conclusion

HA-SLAM posts the best completeness and the best trajectory error of any system in this evaluation, and to our knowledge the strongest egocentric SLAM results reported publicly. It averages 98.8% completeness and never drops below 97.2% on any segment. It records the lowest gap-adjusted error in the field at 1.57 cm, 5.0× better than the next system of any architecture and 13.7× better than stock ORB-SLAM3. And its raw ATE of 1.43 cm is measured over effectively the entire run set rather than a self-selected slice of easy frames.

Holding both at once is what separates it. Every baseline here trades one for the other: ORB-SLAM3 and OKVIS2 reach competitive accuracy but surrender a third and a fifth of the run set to tracking loss, while Basalt, DPVO and DROID-SLAM stay continuously available and give up an order of magnitude or more on error. HA-SLAM concedes neither, and the plot below shows it sitting alone in the corner where both hold.

Model Performance
Gap-adjusted ATE against completeness · upper right is better
low error, high coverage 1 2 5 10 20 50 100 200 60% 70% 80% 90% 100% Completeness → ← Gap-adjusted ATE (cm, log) HA-SLAM Basalt OKVIS2 ORB-SLAM3 DPVO DROID-SLAM
Gap-adjusted ATE plotted against completeness, log error axis. The upper-right region is the target: accurate and continuously available. HA-SLAM is the only system in it. Basalt, DPVO and DROID-SLAM hold coverage but sit one to two orders of magnitude behind on error; ORB-SLAM3 and OKVIS2 reach usable accuracy only on the frames they accept, and fall left as a result. DPVO and DROID-SLAM share the same completeness and are labelled with leader lines.

That makes HA-SLAM the only system in this evaluation usable for generating robotic training data. Egocentric human video is a primary scaling axis for robot learning today, and continuous pose is what converts it into training data at all, since it separates the camera's own motion from motion in the world. A pose stream with gaps is not a slightly worse dataset — frames without pose cannot be used, and the frames around them cannot be trusted. We believe HA-SLAM takes us one step closer to large-scale training of robots on egocentric data, and we are excited to share these results with the community. If any researcher, lab, or company is interested in collaborating or using our SLAM system, we would be glad to hear from you — raj@humanarchive.ai.


References

  1. Krishnan, Liu, Sarlin et al. Benchmarking Egocentric Visual-Inertial SLAM at City Scale. ETH Zurich, Google, Meta Reality Labs, Microsoft, 2025. arXiv:2509.26639.
  2. Campos, Elvira, Gómez Rodríguez, Montiel, Tardós. ORB-SLAM3: An Accurate Open-Source Library for Visual, Visual-Inertial and Multi-Map SLAM. IEEE Transactions on Robotics 37(6):1874–1890, 2021. arXiv:2007.11898.
  3. Leutenegger. OKVIS2: Realtime Scalable Visual-Inertial SLAM with Loop Closure. 2022. arXiv:2202.09199.
  4. Usenko, Demmel, Schubert, Stückler, Cremers. Visual-Inertial Mapping with Non-Linear Factor Recovery. IEEE Robotics and Automation Letters 5(2):422–429, 2020. arXiv:1904.06504.
  5. Teed, Lipson, Deng. Deep Patch Visual Odometry. Advances in Neural Information Processing Systems 36:39033–39051, 2023. arXiv:2208.04726.
  6. Teed, Deng. DROID-SLAM: Deep Visual SLAM for Monocular, Stereo, and RGB-D Cameras. Advances in Neural Information Processing Systems 34:16558–16569, 2021. arXiv:2108.10869.