overthereality.ai

Research note · reconstruction accuracy

How accurate is our automatic 360°-camera reconstruction pipeline?

We benchmarked our 360° photogrammetry pipeline against a LIDAR survey of the same site, captured simultaneously on a shared rig. The short answer: the reconstruction agrees with the LIDAR to a median of 3.9 cm — the reference scanner's own declared accuracy — with no accumulated drift over a 118 m path. The longer answer is the one disagreement we found, a uniform 2.9% scale offset, and the chain of measurements showing the reconstruction did not create it: the model copies its metric reference to 0.03%, the disagreement lies between that reference and the LIDAR, and the only tape measure on record sides with the model. Either way, it is one scalar away from disappearing.

OVR 360° pipeline — dual-fisheye SfM [1] + dense reconstruction · metric scale from mobile-AR poses [4] (iPhone 15, no LiDAR — a median-user device)
Reference: LIDAR survey (3DMakerpro Eagle, declared accuracy ~4 cm [6]), same session, shared rig · 1,414 poses · 8.4 M points
Eval: cloud-to-cloud distance, both clouds @ 5 cm voxel (1.07 M pts) · Umeyama [2] + scaled ICP [3] · 40 × 25 m site, 118 m path

01The question

«How accurate is it?» deserves a number, not an adjective

Our pipeline turns a ten-minute walk with a 360° camera into a dense, metric 3D model of a site. We say «survey-grade» in slides; this note is about earning that phrase. We put a LIDAR scanner on the same pole as the 360° camera, walked the site once, and built the two reconstructions by fully independent means — ours from imagery through structure-from-motion [1], the reference from LIDAR range measurements. Same walk, same light, two instruments that share no assumptions.

The capture is genuinely the thing we sell: one operator, one pass, no tripods, no targets. Below, the input stream and what the pipeline makes of it.

input: 360° stream
output: reconstructed scene, free camera
Left: the raw dual-fisheye walkthrough (one of two lenses shown; the black patches are privacy masking of people and plates, applied at source). Right: a free-camera flythrough rendered from the radiance field trained on this capture — the same underlying geometry this note validates. The flythrough is a render of the model, not video from the walk: that camera path was never flown.

Three numbers hold the story together; the rest of the note reconstructs them one at a time.

3.9 cm
median distance of the reconstruction from the LIDAR surface (1.07 M points, 5 cm voxel, after scale correction)
2.9%
uniform scale offset vs the LIDAR — the one real disagreement, measured and traced to its source
0.08%
scale agreement between two independent halves of the scene — the reconstruction does not drift along the 118 m path

02The result

Half the surface within 3.9 cm, three quarters within 10 cm

The headline protocol is deliberately blunt: downsample both clouds to a 5 cm voxel grid and measure, for each of the 1.07 million reconstructed points, the distance to the nearest LIDAR surface point. After the scale correction discussed below, the median is 3.9 cm; 59.5% of the surface is within 5 cm and 77.6% within 10 cm.

One calibration before the plot: the reference scanner itself has a manufacturer-declared accuracy of about 4 cm [6]. A median agreement of 3.9 cm therefore sits at the floor of what this reference can certify — at this level we are measuring the pair of instruments, and the reconstruction is not distinguishable from the reference's own error budget.

Cumulative distribution of point-to-LIDAR distances
Cumulative distribution of point→LIDAR-surface distance, recomputed for this note from the archived clouds with the report protocol (both clouds at 5 cm voxel; reproduces the published medians to the millimetre and the within-5/10 cm shares to half a point). The long tail beyond ~15 cm is dominated by coverage the two surveys do not share — the LIDAR reaches ~190 m while the usable photogrammetric reconstruction spans ~40 m — rather than by reconstruction error (see §05).

Numbers first, but geometry is easier to trust when you can look at it. Here is the same site as both instruments see it, from the same viewpoint:

ours — 12.8 M pointsPhotogrammetric dense point cloud, true colors
LIDAR — 8.4 M pointsLIDAR point cloud, true colors
The two point clouds in true color, rendered from the same camera after alignment: left our dense reconstruction, right the LIDAR reference. Every comparison in this note is computed between these two datasets.

And because a claim like «the walls coincide» should be checkable, a 1 m horizontal slice of both clouds, seen from above — the classic floor-plan test. Where the alignment is good, green (ours) draws directly on top of gray (LIDAR):

Top-down 1m slice of both clouds overlaid
A horizontal slice 0.6–1.6 m above the floor, top view: LIDAR in gray, our reconstruction in green. Walls, columns and street furniture land on the same lines; green fringes appear where vegetation moved between the two passes of the same walk and where the fisheye cameras saw over the low wall the LIDAR could not.

The camera path itself tells the same story. The pipeline registered 385 rig frames along 115 m; the LIDAR tracked 1,414 poses along 118 m. Aligned with the same transform as the clouds, the two paths agree to a median of 6.1 cm (p90 11.2 cm) over the whole loop [2]:

Top view of both camera trajectories overlaid
Both camera trajectories, top view, from the archived reconstructions: LIDAR in near-black, our 360° rig in green. The 6.1 cm median is measured with the report's trajectory protocol; the short black stub bottom-left is a corridor the LIDAR entered after the 360° recording had stopped.

03A risk we had to defuse

The obvious alignment method quietly converges to a lie

Before any accuracy number exists, the two reconstructions must be aligned — and the alignment itself is the first place this comparison could have silently gone wrong. The obvious tool is free-scale ICP between the two camera trajectories. It is also degenerate: two smooth curves can be brought closer by shrinking one onto the other, so the optimizer happily trades scale for proximity and returns a spurious optimum. A curve, unlike a surface, does not constrain scale.

So the protocol splits the problem in two. First, a closed-form similarity transform (Umeyama [2]) over arc-length correspondences between the trajectories — searching across travel direction and trim, immune to the shrinking failure because it never iterates. Then the scale is refined where it is actually observable: scaled ICP on the dense surfaces [3], at four decreasing correspondence thresholds (50, 25, 12, 6 cm). The scale estimate moved by less than ±0.0003 across all four thresholds — the sign of a well-conditioned optimum, not a lucky one.

The trajectory-only estimate, for the record, landed at 2.49% — half a point off the surface-based consensus. That gap is the degeneracy talking, and it is why every scale number quoted in this note comes from surfaces.

04The cost

One honest disagreement: 2.9% in scale, uniform everywhere

After alignment, one systematic discrepancy survives: our reconstruction is uniformly 2.9% larger than the LIDAR (scale factor 0.9717, pipeline→LIDAR). Where the LIDAR measures 10 m, the model reads 10.29 m — which of the two is right is a question §06 returns to. It is not a bug in the comparison and not noise: four independent estimators — trajectories, full-scene surface ICP, and surface ICP restricted to each half of the scene separately — all land on the same number.

Four independent scale estimates
Scale offset vs LIDAR from four independent estimators (published analysis of 2026-07-28). The trajectory estimate (slate) underestimates for the structural reason of §03 and is shown for the record; the three surface-based estimates agree within 0.08%. Being uniform, the offset is removable with a single scalar.

Ignoring the offset is not an option, and not only for the tape-measure test: locking scale to 1 degrades the whole accuracy table — the geometry gets blamed for what is really a gauge error.

Alignmentmedianp90within 5 cm
rigid only — scale locked to 15.9 cm41.9 cm44.9%
with the measured scale correction3.9 cm27.3 cm59.5%

A 2.9% scale error costing 2 cm of median accuracy is the whole story of this section in one row: the residual error budget of this reconstruction is so small that a scale gauge most surveys would shrug at is the dominant term. So — where does the 2.9% come from? Two suspects: the reconstruction accumulated it along the walk, or it inherited it from somewhere. The next two sections interrogate them in turn.

05Ruling out drift

It is one global gauge factor, not error accumulating along the path

The tempting story is the classic one: image-based surveys drift, error compounds with distance, and 118 m is a long walk. We nearly wrote that sentence. It is wrong, and the scene itself proves it. If error accumulated along the path, two independent halves of the scene would disagree about scale, and the residual would grow toward the periphery. Neither happens: the two halves return the same scale to within 0.08%, and the median residual moves only from 3.2 cm at the scene centre to 5.5 cm at the periphery.

Median residual vs distance from scene centre
Median point→LIDAR distance by distance from the scene centre, recomputed from the archived clouds. Within the coverage the two surveys share (green) the profile is essentially flat — drift would draw a rising line. The orange region is dominated by structure the cameras only saw from tens of metres away while the LIDAR scanned it directly; the growth there measures coverage, not drift.

The same effect explains the long tail of the headline distribution, and it is worth seeing where those «bad» points actually live. Painting every reconstructed point by its measured distance to the LIDAR puts the red exactly where the geometry argument says it should be — distant façades and vegetation at the edge of coverage, not the surveyed core:

Reconstruction colored by measured distance to LIDAR
The reconstruction with each point colored by its measured distance to the LIDAR surface: green < 5 cm, orange < 10 cm, red ≥ 10 cm. The walked area — ground, walkway, near structures — is solidly green; red concentrates on the building façade tens of metres outside the walked loop and on vegetation. The same layer is toggleable on the interactive model in §07.

06The verdict on the 2.9%

The reconstruction did not create the 2.9% — and it can prove it

With drift ruled out, the remaining suspect is the pipeline's source of absolute scale. Our reconstructions are made metric by the AR poses of mobile-device photos captured alongside the 360° walk [4] — 95 of them in this scene. That gives us a decisive experiment: compare the finished model against its own reference. If the reconstruction distorted scale, it would disagree with the very poses it was scaled from.

It does not. The model reproduces its AR reference to 0.03% in scale and 1.2 cm median in position (p90 2.7 cm). The reconstruction is a faithful copy of the gauge it was handed; the 2.9% is a disagreement between the gauge and the LIDAR, passed through intact.

Scale disagreement: model vs its reference, model vs LIDAR
The blame test, from the published analysis: the reconstruction agrees with its own metric reference two orders of magnitude more tightly than reference and LIDAR agree with each other. Whatever created the 2.9% happened upstream of the reconstruction.

Two structural observations turn this from an excuse into a roadmap. First, the device: the reference photos were taken with an iPhone 15 without a LiDAR sensor, on purpose — that is the median device of our users, and we wanted the validation to reflect a typical capture, not a best-case one. Its ARKit scale comes from visual-inertial odometry alone, exactly the quantity a phone LiDAR helps pin down, so a gauge error of a few percent is unsurprising — and repeating the capture with a LiDAR-equipped phone is the obvious next measurement. Second, the coverage: the 95 reference photos span roughly 6 × 5.7 m of a 40 × 25 m site, so global scale is extrapolated from under 4% of the surveyed area, and any local bias in the AR poses there propagates to the whole model. Both levers — a better reference device and wider reference coverage — are capture-time fixes that require no change to the reconstruction itself.

And one last twist: which side of the 2.9% disagreement is actually right? The LIDAR is a consumer instrument with its own ~4 cm declared error budget [6], so «LIDAR = truth» is not granted. We ran the crudest possible third gauge: a tape measure over a row of six floor tiles on site — 201 cm in the real world, 200.5 cm over the same span in the delivered model, a −0.25% difference well within tape-and-tile tolerance. On that span the model's absolute scale is essentially right, which — given that the model copies its AR reference to 0.03% — hints that a good share of the 2.9% may sit on the LIDAR's side of the disagreement, not ours. One hand-taped span at one location decides nothing at percent level across a 40 m scene; read it as directional. But it is the difference between «our model is 2.9% too large» and the more honest «our model and the LIDAR disagree by 2.9%, and the only tape measure on record votes for the model».

07A blind check

The alignment recovers rig geometry it was never told

Every number so far rests on one alignment, so the alignment deserves its own audit — ideally against a quantity it could not have optimized for. The rig provides one: the LIDAR rode about 70 cm below the 360° camera on the same pole. The cloud-to-cloud alignment was computed from surface correspondences alone; it knows nothing about hardware. Yet measuring both camera heights above the LIDAR's floor plane in the aligned frame returns 2.15 m and 1.45 m — a recovered offset of 69.8 cm.

The capture rig in use: dual-fisheye 360-degree camera on top of the pole, Eagle LIDAR scanner mounted on the crossbar roughly 70 cm below
The shared rig during this capture: the dual-fisheye 360° camera rides at the top of the pole (~2.15 m), the Eagle LIDAR sits on the crossbar below (~1.45 m). The ~70 cm between them is the quantity the alignment recovers blind in the table that follows — from surface geometry alone, never having been told the rig exists.
Quantitymeasured from the aligned clouds
360° camera height above LIDAR floor2.15 m
LIDAR sensor height above LIDAR floor1.45 m
recovered vertical rig offset (physical: ~70 cm)69.8 cm

Two honest footnotes. First, «~70 cm» is a tape-measure figure for the physical rig, so read the last millimetres of that agreement as luck; the point is that a blind geometric estimate lands on the hardware to well under the error budget it is auditing. Second, the trajectory figure of §02 uses the report's trajectory-alignment protocol; our recomputation for this note, which reuses the cloud alignment instead, gives an even smaller 3.4 cm median — we quote the larger published number and treat the difference as protocol sensitivity, not accuracy gained.

Below, the full evidence in one place: both clouds, both trajectories — the two paths ride visibly ~70 cm apart, as the rig dictates — and the measured per-point error as a toggleable layer. The claim of this note is that green and gray are the same site to about four centimetres; here you can check it.

preview of the interactive 3D viewer: both aligned point clouds and the camera trajectories
The two aligned reconstructions, live: LIDAR reference and our model (320 k points each, decimated from 8.4 M / 12.8 M), the two camera trajectories (gray tube: LIDAR; green: 360° rig — the constant vertical gap is the physical rig offset), and the measured error toggle, which repaints our cloud by its true per-point distance to the LIDAR with the same color scale as the figure above. If this note is right, turning it on should show a green core with red only at the coverage fringe.
Surface agreement (after scale correction)medianp68p90within 5 cmwithin 10 cm
published table (2026-07-28)3.9 cm6.6 cm27.3 cm59.5%77.6%
recomputed for this note from the archived clouds3.9 cm6.7 cm28.1 cm59.2%77.1%