Spatial and Physical Computation · Systems Study

The Geometry of Operational State

Reconstructing physically coherent reality from sensor streams, digital twins, and incomplete observations.

Constructed benchmark Coordinate frames treated as evidence Physical feasibility separated from correctness Uncertainty regions replace false precision

Empirical status. This study uses a controlled, programmatically constructed spatial-state benchmark. Its three domain labels identify typed operational geometry families, not real clinical, warehouse, or industrial environments. Each channel is reduced to at most one observation per object at a reconstruction boundary; sequence-level stream processing is not evaluated. The study reports no customer records, production sensor logs, field failure frequencies, domain-performance estimates, or measured Cortonex deployment performance. The fitted trace result is conditional on the explicitly parameterized diagnostic-response model in Section 5.3. The exact geometry result is a deterministic oracle ceiling whose value equals benchmark release validity by construction.

Abstract

A reconstructed operational scene can be composed entirely from individually plausible observations and still describe the wrong world. Sensor coordinates may be expressed in an unverified frame. Two objects may exchange identities while retaining plausible positions. A digital twin may be geometrically clean but one revision behind the physical environment. A collision-removal step may produce a feasible configuration at the wrong location. An occluded object may be represented as a precise point even though the observations support only a broad region. These are not ordinary data-quality defects. They are failures of spatial meaning.

We study such failures with a constructed benchmark containing 12,000 operational scenes, 48,000 typed objects, and 144,000 potential source-object observation slots. Scenes are divided equally among clinical spaces, warehouse operations, and industrial equipment bays. Each scene contains a bounded floor region, fixed obstacles, six operational zones, four typed objects, three observation channels, source-frame metadata, four fixed fiducials, digital-twin revision state, and an authoritative reference geometry at the reconstruction boundary.

The benchmark contains nine conditions: clean execution, rigid frame mismatch, identity collision, occlusion overcommitment, stale digital twin, impossible occupancy, topology breach, support gap, and compound geometry fault. Nonclean scenes receive a material fault with probability 0.86; the remaining cases are condition-stratified safe controls with ambiguous telemetry but a valid final state. Each source is represented by a boundary-time observation rather than a temporal sequence. We compare reported-coordinate fusion, correspondence-based fiducial normalization, hard physical constraint projection, uncertainty-aware state representation, scalar-confidence release, a geometry-trace release policy, and an exact benchmark oracle.

The held-out test population contains 2,400 scenes: 1,363 release-valid scenes, 1,037 invalid scenes, and 993 invalid scenes that the deliberately minimal base release rule would pass silently. All 44 held-out base holds arise in support-gap scenes with a missing required field. Reported-coordinate fusion is exactly valid in 45.50% of scenes and physically feasible in 78.29%. Fiducial frame normalization raises exact validity to 55.38% and feasibility to 85.08%. Hard constraint projection makes all 2,400 test scenes physically feasible, but exact validity rises by only two scenes, from 1,329 to 1,331. An uncertainty-aware representation raises release validity to 56.79% by replacing unsupported point certainty with bounded state regions whose widths depend only on available-source count and post-registration inter-source spread.

A scalar confidence gate holds 114 of 1,037 invalid scenes, or 10.99%, including 70 of 993 base-silent failures, or 7.05%. Under the diagnostic-response model specified in Section 5.3, a geometry-trace gate receiving noisy frame-registration, identity, revision, observation-support, zone, lineage, association, and source-integrity channels holds 946 invalid scenes, or 91.22%, including 902 base-silent failures, or 90.84%. It false-holds 25 of 1,363 valid scenes, or 1.83%. The scalar gate releases 93.83% of the stress suite with 40.99% invalidity among releases. The trace gate releases 59.54% with 6.37% invalidity among releases.

The result is an observability measurement under an explicitly parameterized constructed diagnostic-response model. It is not production performance or an independent estimate of raw-sensor detection accuracy. The architectural conclusion is narrower: local geometric plausibility does not establish operational spatial truth. A release system must treat coordinate frames, object identity, revision state, physical occupancy, operational topology, observation support, and uncertainty as separate evidence-bearing properties.

Primary held-out results
PolicyInvalid scenes heldBase-silent scenes heldValid false holdsRelease coverageInvalidity among released scenes
Scalar confidence gate114 of 1,037, 10.99%70 of 993, 7.05%34 of 1,363, 2.49%93.83%40.99%
Geometry trace gate946 of 1,037, 91.22%902 of 993, 90.84%25 of 1,363, 1.83%59.54%6.37%
Exact geometry oracle1,037 of 1,037, by construction993 of 993, by construction0 of 1,363, by construction56.79%0%, by construction

A scene can be locally coherent and globally wrong

Rigid transformations preserve local geometry. If every object in a sensor frame is rotated and translated by the same transformation, pairwise distances remain unchanged:

because a rigid rotation satisfies:

A local consistency check can therefore report a coherent scene while the complete scene is displaced or rotated relative to the world frame. The problem is not that the observations contradict one another. The problem is that they agree in the wrong coordinate system.

Spatial plate A Rigid-frame error Constructed scene

Interactive replacement for this plate. Plan view of a published constructed scene, with positions, zones, and footprints read from the study payload. Any of the six held-out scenes can be selected; the extruded view is an optional illustration.

Spatial plate A. Rigid-frame error preserves local structure while corrupting world-state meaning.
Spatial plate A. Rigid-frame error preserves local structure while corrupting world-state meaning.

Spatial plate A. Rigid-frame error. A held-out constructed scene is shown in the authoritative world frame, after reported-coordinate fusion under a shared source-frame error, and after fiducial normalization. The example is illustrative of the benchmark mechanism. It is not a customer scene.

This distinction is central to operational state. A scene can satisfy all of the following and still be wrong:

  • object-to-object distances look plausible;
  • every source record is well formed;
  • sensor confidence is high;
  • no object overlaps another;
  • the scene fits inside the mapped boundary;
  • the assembled state is internally consistent.

Spatial truth requires more. The state must be expressed in the correct frame, attached to the correct identities, current to the controlling revision, supported at the claimed precision, physically feasible, and topologically consistent with the operational environment.

1Operational state as a constrained geometric object

An operational scene is not merely a set of coordinates. At reconstruction boundary τ\tau, define:

where:

  • ii is object identity;
  • gig_i is the typed axis-aligned floor footprint;
  • piR2p_i\in\mathbb R^2 is floor position in the authoritative world frame;
  • ziz_i is operational zone or topological state;
  • νi\nu_i is the controlling revision;

The benchmark is deliberately planar. State estimation occurs in the floor plane. The static plates extrude the evaluated footprints using four fixed illustrative heights solely to make object identity and overlap legible; vertical height is not a generated, fitted, or evaluated benchmark variable. Object orientation is fixed by type and is not an evaluated degree of freedom. The construction is sufficient to study frame registration, identity, floor occupancy, zone topology, stale twins, and planar uncertainty without claiming arbitrary three-dimensional surface or pose reconstruction.

A source ss observes object ii through a source-specific measurement operator:

where xix_i denotes the relevant state variables and εis\varepsilon_{is} represents observation error. If the source reports coordinates in its own frame Fs\mathcal F_s, the world-frame relation is:

with:

A reconstruction engine does not know RWsR_{Ws} and tWst_{Ws} merely because a source supplies a coordinate. It must inherit, estimate, or verify the transform. Related localization and mapping systems make the same separation explicit: observations acquire spatial meaning through an estimated relation between sensor pose and map state [8,9]. In the benchmark, fixed fiducials provide independent frame anchors. The normalization stage estimates:

where aksa_k^s and akWa_k^W are corresponding fiducials. This is the classical absolute-orientation problem [1,2]. Iterative closest-point methods generalize registration when correspondences are not fixed in advance [3], while robust estimators such as random sample consensus address contaminated correspondence sets [4].

The benchmark intentionally keeps four fiducial correspondences known. Their reported coordinates are transformed by the same live-source frame error and independent anchor noise. A two-dimensional orthogonal Procrustes fit estimates the inverse rigid map from the fiducials alone, then normalizes the camera-like and ranging observations before they are fused again with the already world-referenced digital twin. Authoritative object positions are never used for registration.

1.1A scene is more than a point cloud

For each object, define its axis-aligned floor footprint under reconstructed pose:

Hard physical feasibility requires boundary inclusion:

obstacle exclusion:

and object clearance. Let Bi+0.01(p^i)B_i^{+0.01}(\widehat p_i) denote the footprint expanded by 0.010.01 m on every side. The benchmark requires the interiors of the expanded footprints to be disjoint:

Equation (11) is an axis-aligned two-centimetre object-to-object clearance rule. Interior disjointness permits the two expanded boundaries to touch, which corresponds to exactly 0.02 m of separation between the original footprints and matches the benchmark implementation. These conditions are necessary, not sufficient. An object can be collision-free and still occupy the wrong ward, aisle, bay, compartment, or control zone. Let:

map spatial position to operational zone. Topological consistency requires:

where ziz_i^{*} is the authorized zone claim at boundary τ\tau.

Identity consistency is separate again. Let πs\pi_s map source detections to canonical object identities. A correct association requires:

If two objects exchange identities, the set of positions can remain plausible while the operational meaning is wrong. This is the classical data-association problem in another form. Multiple-hypothesis tracking and joint probabilistic data association were developed precisely because assignment uncertainty cannot be reduced to coordinate noise [5,6].

Revision consistency requires:

where νi\nu_i^{*} is the controlling digital-twin or source revision at τ\tau. A stale twin can be geometrically smooth, complete, and self-consistent while representing an earlier physical world.

1.2Contributions and scope

We make seven contributions.

First, we define operational spatial validity as a conjunction of world-frame positional support, including frame registration and bounded-uncertainty conditions, together with identity continuity, revision currency, physical feasibility, and zone topology.

Second, we construct a controlled benchmark in which each failure mechanism alters concrete sensor, transform, identity, revision, occupancy, topology, or support state rather than merely flipping an abstract error label.

Third, we separate four reconstruction stages: reported-coordinate fusion, fiducial frame normalization, hard constraint projection, and uncertainty-aware representation.

Fourth, we show experimentally that physical constraint satisfaction can increase sharply without a comparable increase in operational truth. In the held-out set, hard projection raises physical feasibility from 85.08% to 100%, while exact point-state validity increases from 55.38% to 55.46%.

Fifth, we compare confidence-only release with a geometry-trace policy under disjoint fitting, calibration, threshold-selection, and test populations.

Sixth, we provide interactive fault, calibration, risk-coverage, diagnostic-ladder, exact-recheck, and scene-trace views together with six static geometric plates.

Seventh, we state the boundary of the result. The benchmark does not establish real sensor performance, general three-dimensional reconstruction, field fault prevalence, digital-twin correctness, or measured Cortonex deployment performance.

2Benchmark construction

2.1Scene families

The benchmark contains:

constructed scenes, divided equally among:

Typed scene familyScenesConstructed objectsPotential source-object observation slots
Clinical spaces4,00016,00048,000
Warehouse operations4,00016,00048,000
Industrial equipment bays4,00016,00048,000
Total12,00048,000144,000

Each scene contains:

  • a bounded 24 m×16 m24\text{ m}\times16\text{ m} floor region;
  • six labeled operational zones;
  • two fixed obstacle bands;
  • four typed axis-aligned floor footprints, rendered as extruded cuboids in the static plates;
  • camera-like, ranging, and digital-twin observation channels;
  • source-frame and fiducial state;
  • current and prior digital-twin revisions;
  • an authoritative reference geometry at boundary τ\tau.

The domain labels alter object types and dimensions. They do not reproduce clinical workflows, warehouse physics, industrial safety rules, real sensor characteristics, or domain-specific operating distributions. Each source contributes at most one boundary-time observation per object. The study does not evaluate stream latency, asynchronous delivery, temporal filtering, motion dynamics, or sequence-level tracking.

The six zone rectangles, expressed as [x0,x1]×[y0,y1][x_0,x_1]\times[y_0,y_1] in metres, are:

ZoneBounds
Staging[0.5,6.2]×[0.5,5.0][0.5,6.2]\times[0.5,5.0]
Dock[7.0,17.0]×[0.5,5.0][7.0,17.0]\times[0.5,5.0]
Quality hold[18.0,23.5]×[0.5,5.0][18.0,23.5]\times[0.5,5.0]
Aisle west[0.5,6.2]×[6.0,15.5][0.5,6.2]\times[6.0,15.5]
Aisle central[8.0,15.0]×[6.0,15.5][8.0,15.0]\times[6.0,15.5]
Aisle east[16.8,23.5]×[6.0,15.5][16.8,23.5]\times[6.0,15.5]

The fixed obstacle bands are [6.45,7.65]×[5.4,15.8][6.45,7.65]\times[5.4,15.8] and [15.15,16.35]×[5.4,15.8][15.15,16.35]\times[5.4,15.8]. Objects 1 through 4 are assigned, respectively, to aisle east, staging, aisle west, and quality hold. Their authoritative centers are sampled uniformly inside the assigned zone after accounting for half-footprint extent and an additional 0.15 m inward margin.

Typed scene familyObject 1 footprintObject 2 footprintObject 3 footprintObject 4 footprint
Clinical spacesBed, 2.15×0.952.15\times0.95 mSupply cart, 1.05×0.721.05\times0.72 mInfusion stand, 0.55×0.550.55\times0.55 mIsolation cabinet, 0.82×0.820.82\times0.82 m
Warehouse operationsPallet, 1.22×1.021.22\times1.02 mForklift, 2.45×1.252.45\times1.25 mAutonomous cart, 0.76×0.700.76\times0.70 mQuality cage, 0.92×0.820.92\times0.82 m
Industrial equipment baysMachine skid, 2.30×1.552.30\times1.55 mService cart, 1.25×0.851.25\times0.85 mInspection unit, 0.88×0.880.88\times0.88 mTool cabinet, 1.02×0.681.02\times0.68 m

Source and registration model

Before intervention, camera-like, ranging, and digital-twin positions are generated from the authoritative object centers with independent isotropic Gaussian coordinate noise:

ChannelCoordinate-noise standard deviationBaseline fusion weightFrame state
Camera-like observation0.18 m per axis0.38Live-source frame
Ranging observation0.10 m per axis0.42Live-source frame
Digital-twin observation0.04 m per axis0.20World frame

The baseline weights are fixed across all conditions and renormalized only over the channels available for each object. No injected condition name or materiality flag changes the reconstruction weights. The four fixed fiducials receive independent Gaussian coordinate noise with standard deviation 0.028 m per axis before rigid registration. The controlling twin revision is sampled uniformly from the integers 3 through 9, and the prior revision is one less.

For nonmaterial controls, a shared perturbation with coordinate standard deviation 0.18 m is applied at 0.25 scale to the camera-like channel and at negative 0.15 scale to the ranging channel. Nonmaterial occlusion and support-gap controls also remove one camera-like observation while preserving final validity. These values are benchmark design parameters, not estimates of deployed sensor accuracy.

2.2Condition allocation

Each scene receives one condition. Per scene family, the fixed allocation is:

ConditionPer familyComplete benchmark
Clean execution1,2803,840
Rigid frame mismatch4801,440
Identity collision4001,200
Occlusion overcommitment4401,320
Stale digital twin4001,200
Impossible occupancy3601,080
Topology breach320960
Support gap200600
Compound geometry fault120360
Total4,00012,000

For a nonclean condition, materiality is sampled as:

The frozen benchmark realized 7,015 material nonclean scenes. The 0.86 value and the condition allocation are stress-design choices. They are not estimates of field failure prevalence.

Within this benchmark, material is an intervention-strength flag, not a synonym for final invalidity. It selects the higher-amplitude branch of the named mechanism before normalization, projection, uncertainty representation, and release evaluation. A material frame mismatch can be corrected by fiducial registration, and a material occlusion intervention can become release-valid through a bounded uncertainty region. Materiality therefore describes the injected challenge, while validity describes the final reconstructed state.

Nonmaterial controls preserve the final valid state while introducing ambiguous telemetry, such as partial source absence or mild inter-source disagreement. They are assigned within the same condition strata, but they are not mechanism-matched counterfactuals for every material intervention. Their purpose is to prevent the fitted trace policy from treating every anomaly as a fault.

2.3Failure mechanisms

Rigid frame mismatch

The camera-like and ranging channels receive the same rotation and translation in their reported coordinate frame:

The pairwise geometry can remain plausible even when the world placement is wrong.

Identity collision

Two object assignments are exchanged through an incorrect association map:

The geometry may still contain the correct number of boxes, but the state belongs to the wrong objects.

Occlusion overcommitment

Live support is removed for one object, and the remaining estimate is biased while initially represented with unjustified precision. The fault is not merely positional error. It is the mismatch between evidential support and reported uncertainty.

Stale digital twin

A prior revision contributes state as if it were current:

Live observations for the affected object are partially absent, allowing the stale twin to dominate fusion.

Impossible occupancy

All sources agree on a placement intersecting a fixed obstacle or another object. This creates a physically impossible but observationally coherent state.

Topology breach

An object is assigned to a region outside its authorized operational zone:

The injected topology mechanism does not enforce hard feasibility before fusion. Some topology-labelled scenes therefore also activate boundary, obstacle, or object-clearance checks. The condition name identifies the intervention family, not an exclusive validity component.

Support gap

Current live observations are absent, and the system falls back to an older twin state. Some cases are overt because a required field is missing; others remain silent.

Compound geometry fault

A shared frame error is combined with either identity collision or stale-twin wrong-zone state.

The nine condition labels are intervention families rather than mutually exclusive failure predicates. An identity exchange can also create a footprint conflict because object dimensions differ. An impossible-occupancy placement can also violate zone topology. A topology intervention can be infeasible before projection. Fault-level tables must therefore be read as behavior under the named intervention, not as a decomposition into orthogonal causes.

The material intervention parameters are fixed as follows:

Intervention familyConstructed parameterization
Rigid frame mismatchRotation magnitude uniform from 12 to 28 degrees with random sign; each translation coordinate uniform from -1.45 m to 1.45 m
Identity collisionTwo distinct object assignments selected uniformly and exchanged in both live channels
Occlusion overcommitmentCamera absent for one object; ranging available with probability 0.35; twin displacement magnitude uniform from 0.95 m to 1.85 m in a random direction
Stale digital twinPrior revision used; twin displacement magnitude uniform from 1.10 m to 2.50 m; camera and ranging each available with probability 0.45
Impossible occupancyWith probability 0.55, the object is placed inside a fixed obstacle band; otherwise it is placed near another object
Topology breachThe object is moved near the center of a nonauthorized zone with isotropic coordinate perturbation of standard deviation 0.35 m
Support gapBoth live channels are absent; the prior twin is displaced by 0.90 m to 2.10 m; a required field is marked missing with probability 0.35
Compound geometry faultRotation magnitude uniform from 10 to 23 degrees and translation coordinates uniform from -1.20 m to 1.20 m, followed with probability 0.55 by identity exchange and otherwise by stale wrong-zone twin state

These choices define the stress suite. They are not empirical distributions of field faults.

2.4Data separation

The 12,000 scenes are partitioned before final evaluation:

PopulationScenesPurpose
Training7,200Fit policy coefficients
Probability calibration1,200Fit sigmoid probability calibration
Threshold selection1,200Select release thresholds under a valid false-hold budget
Held-out testing2,400Final measurement only

The split is stratified by typed scene family and condition. Held-out aggregate validity labels do not estimate policy coefficients, calibration parameters, or release thresholds. Diagnostic channels in every split are generated by the same frozen response formulas, several of which use constructed latent predicates disclosed in Section 5.3. Let T\mathcal T denote the held-out test population and define:

3Reconstruction stages

3.1Reported-coordinate fusion

For object ii, let Ai\mathcal A_i be the available sources and wisw_{is} their normalized weights. The baseline point estimate is:

This stage trusts each source's reported world coordinates and canonical identity.

3.2Fiducial frame normalization

The second stage estimates an inverse rigid transform from four fixed anchors using Equation (7), maps the camera-like and ranging observations into the world frame, and then repeats fusion while leaving the already world-referenced digital twin unchanged:

The benchmark does not use authoritative object positions to estimate this correction. Only fixed fiducials are available.

3.3Hard physical constraint projection

Let Sˉ\bar S denote the normalized point reconstruction, with normalized fused object positions pˉi\bar p_i, and let Chard\mathcal C_{\mathrm{hard}} be the set satisfying Equations (9) through (11). The ideal Euclidean projection is:

Projection methods are standard tools in convex feasibility [11]. The benchmark hard set is nonconvex because of pairwise clearance, and the study does not claim to solve Equation (24) globally. Its implementation is an order-dependent deterministic approximation. Objects are processed sequentially. Candidate displacements use radii from 0.25 m to 5.00 m in 0.25 m increments, with 24 equally spaced directions at each radius, plus the zero-displacement candidate. The nearest feasible candidate is selected. This is not a general optimizer for a nonconvex three-dimensional scene.

Even an exact selection from the minimizers in Equation (24) would choose a closest feasible state to the input, not necessarily the authoritative state. If the input is wrong because of identity, topology, or stale revision, a geometrically minimal repair can remain operationally wrong.

Spatial plate B Feasible but wrong Constructed scene

Interactive replacement for this plate. Plan view of a published constructed scene, with positions, zones, and footprints read from the study payload. Any of the six held-out scenes can be selected; the extruded view is an optional illustration.

Spatial plate B. Hard physical constraint projection restores feasibility without proving correctness.
Spatial plate B. Hard physical constraint projection restores feasibility without proving correctness.

Spatial plate B. Feasible but wrong. The hard projection removes the constructed occupancy violation. It cannot infer the authoritative position merely from the fact that a nearby feasible position exists.

3.4Uncertainty-aware representation

A point estimate is not always the correct output. For object ii, the benchmark uses an isotropic Gaussian-form support model:

Let ci{1,2,3}c_i\in\{1,2,3\} be the number of available source channels in the frozen benchmark. The digital-twin channel is always retained, so the zero-source state is not evaluated. Let pˉi\bar p_i be the normalized weighted fusion from Section 3.2. The spread used by the benchmark is the unweighted root mean squared distance of the available normalized source points from that weighted fusion, si=Ai1sAip^isWpˉi22s_i=\sqrt{|\mathcal A_i|^{-1}\sum_{s\in\mathcal A_i}\lVert\widehat p_{is}^{W}-\bar p_i\rVert_2^2}. In Equation (25), μi\mu_i is the projected point center from Section 3.3. The nominal scale is the explicit design rule:

No intervention label, condition name, materiality flag, authoritative reference error, or final validity label enters Equation (25a). The same width rule is applied to every object in every scene. The support model remains constructed and nominal, but it is label-independent at inference time.

The corresponding nominal 1α1-\alpha support contour is:

For the planar benchmark:

For the isotropic planar model, contour area is:

The benchmark accepts uncertain positional support only when the authoritative reference is inside the nominal 95% contour and the contour area does not exceed a fixed ceiling. The word nominal is important: the construction does not claim frequentist 95% coverage in a field population.

The ceiling prevents an arbitrarily broad region from making every point trivially supported.

Spatial plate C Partial observation Constructed scene

Interactive replacement for this plate. Plan view of a published constructed scene, with positions, zones, and footprints read from the study payload. Any of the six held-out scenes can be selected; the extruded view is an optional illustration.

Spatial plate C. Occlusion should widen the state instead of inventing a precise point.
Spatial plate C. Occlusion should widen the state instead of inventing a precise point.

Spatial plate C. Partial observation. The uncertainty region exposes what the observations do not determine. It does not convert an unsupported point into a verified location.

4Validity and release

4.1Positional support

For point-valued stages, object error is:

Point position is accepted when:

Scene positional root mean squared error is:

The stage table reports the arithmetic mean of scene-level RMSE, so each scene receives equal weight:

This is not the pooled object-level RMSE. For reference, pooled object-level RMSE is 4.0097 m for reported-coordinate fusion, 3.8879 m after frame normalization, and 3.9061 m after projection. The article uses mean scene-level RMSE because the scene is the evaluation unit.

For the uncertainty-aware stage, an object is position-supported when either Equation (31) holds or Equations (28) and (29) hold. Hard physical feasibility is evaluated at the projected point centers. The benchmark does not require every point inside an uncertainty contour to satisfy the hard constraints.

4.2Release validity

From this section onward, index ii denotes a scene rather than an object. Define binary scene conditions:

  • PiP_i: every object is supported in the authoritative world frame, either by the point tolerance or by the bounded uncertainty rule;
  • IiI_i: object identities remain correct;
  • RiR_i: the controlling revision is current;
  • OiO_i: hard occupancy constraints hold;
  • TiT_i: operational zone topology is correct.

Frame error is evaluated through world-frame positional and zone support, while uncertainty honesty is part of PiP_i. The benchmark release-validity label is therefore:

The multiplication denotes Boolean conjunction. It does not assume statistical independence.

A physically feasible scene satisfies only:

Therefore:

but:

Equation (36) is the central distinction in the study.

4.3Base hold and silent invalidity

Let Bi=1B_i=1 when the base assembler visibly holds a scene because observation confidence is below 0.49 or a required object field is missing.

A base-silent invalid scene is:

An overt invalid scene is:

The held-out identities are:

and:

Thus, 993 invalid held-out scenes pass the ordinary base rule without a visible hold.

The base comparator is deliberately minimal. All 183 base holds, including all 44 held-out base holds, occur in material support-gap scenes with at least one required object field missing. Seventeen of those complete-benchmark scenes, including two held-out scenes, also fall below the scalar confidence threshold, but the missing required field already holds them. The base-silent count therefore means silent relative to this limited comparator, not silent relative to a mature production release policy.

5Release policies

5.1Scalar observation-confidence gate

The scalar policy receives only the one-world observation-confidence value. It does not see the condition label, materiality, authoritative scene, or benchmark validity bit.

5.2Geometry trace gate

The geometry trace policy receives eight noisy one-world diagnostics:

  1. observation confidence;
  2. constructed identity-consistency response;
  3. digital-twin revision freshness;
  4. observation support completeness;
  5. zone-claim consistency;
  6. spatial lineage completeness;
  7. constructed association-consistency response;
  8. frame, revision, and projection integrity.

5.3Constructed diagnostic-response model

The trace channels are generated by an explicit response model. They are not measurements from deployed sensors or a production reference monitor. For a latent channel value [0,1]\ell\in[0,1] and declared reliability rr, the observed diagnostic is:

The exact latent definitions and response reliabilities are:

DiagnosticLatent value before response noiseReliability rr
Constructed identity-consistency response0.950.95 when identity is correct, otherwise 0.080.080.86
Revision freshness0.950.95 when the controlling revision is current, otherwise 0.080.080.87
Observation supportMean available-source fraction across the four objects0.78
Zone-claim consistency0.940.94 when every projected object is in its authorized zone, otherwise 0.130.130.82
Spatial-lineage completeness0.42cˉ+0.28uνline+0.30uf0.42\bar c+0.28u_\nu^{\mathrm{line}}+0.30u_f, clipped to [0,1][0,1]0.75
Constructed association-consistency response0.88esˉ/(0.9 m)+uI0.88e^{-\bar s/(0.9\text{ m})}+u_I, clipped to [0,1][0,1]0.76
Frame, revision, and projection integrityerfpre/(0.55 m)uνintuΦedproj/(1.8 m)e^{-r_f^{\mathrm{pre}}/(0.55\text{ m})}u_\nu^{\mathrm{int}}u_\Phi e^{-d_{\mathrm{proj}}/(1.8\text{ m})}0.84

Here cˉ\bar c is mean source availability, sˉ\bar s is mean normalized inter-source spread, and rfprer_f^{\mathrm{pre}} is the root mean squared discrepancy between the reported and known world-frame fiducial coordinates before rigid correction. In the lineage channel, uνline=1u_\nu^{\mathrm{line}}=1 for a current revision and 0.150.15 when stale. The frame term is uf=1u_f=1 when post-registration fiducial residual is below 0.180.18 m and 0.350.35 otherwise. The association offset is uI=0.10u_I=0.10 for a correct identity and 0.080.08 otherwise; when identity is incorrect, the complete association latent is capped at 0.220.22. In the integrity channel, uνint=1u_\nu^{\mathrm{int}}=1 for a current revision and 0.420.42 otherwise, uΦ=1u_\Phi=1 when the normalized point scene is physically feasible and 0.320.32 otherwise, and dprojd_{\mathrm{proj}} is the sum of object-center displacement introduced by hard projection. No injected condition name or materiality flag enters this channel.

Observation confidence is generated separately. Let cˉ\bar c again denote mean source availability and let ai=esˉi/(0.9 m)a_i=e^{-\bar s_i/(0.9\text{ m})} denote position agreement for scene ii. Its pre-noise center is:

where δimiss=1\delta_i^{\mathrm{miss}}=1 when a required object field is missing and is zero otherwise. The reported value is clip[0,1](N(μC,i,0.0652))\operatorname{clip}_{[0,1]}(\mathcal N(\mu_{C,i},0.065^2)). Coherent faults can therefore remain confident through source agreement itself; the confidence generator does not receive the intervention class.

Several latent channels are noisy responses to benchmark validity components rather than outputs of independently evaluated raw-sensor algorithms. The trace result is therefore a conditional observability result under this diagnostic model. It is not an independent estimate of how accurately a deployed system would infer identity, revision, topology, or integrity from field data. On the held-out set, the individual channel AUROCs for validity range from 0.607 to 0.795; the combined policy gains its separation by composing several partially informative channels.

5.4Fitting and calibration

For policy kk with standardized feature vector xi(k)\mathbf x_i^{(k)}, the fitted validity score is:

where:

The feature standardizer and an L2L_2-regularized logistic model are fitted only on the training population. The implementation uses balanced class weights, regularization parameter C=1C=1, and the LBFGS solver. A separate sigmoid calibration map is then fitted on the calibration population using a near-unregularized one-dimensional logistic fit with C=106C=10^6:

Probability calibration and ranking answer different questions [12]. The study therefore reports Brier score, expected calibration error, ranking metrics, detection, false holds, coverage, and selective risk together.

5.5Threshold selection

A scene is held when the base rule already holds it or fitted validity falls below threshold:

The numerical tolerance prevents machine-level equality changes from altering a threshold tie.

Thresholds are chosen on the threshold-selection population to maximize invalid-scene detection subject to a 2% valid false-hold budget. When two feasible observed thresholds hold the same number of invalid scenes, the larger threshold is selected:

subject to:

The selected thresholds are:

The threshold-selection population contains 680 valid scenes, so the integer budget is:

Both selected policies hold exactly 13 valid threshold-selection scenes. Their separate held-out false-hold rates are allowed to differ.

6Reconstruction results

6.1Stage comparison

Reconstruction stageScene validityPhysical feasibilityMean scene-level positional RMSE
Reported-coordinate fusion1,092 of 2,400, 45.50%1,879 of 2,400, 78.29%2.2763 m
Fiducial frame normalization1,329 of 2,400, 55.38%2,042 of 2,400, 85.08%1.9840 m
Hard physical constraint projection1,331 of 2,400, 55.46%2,400 of 2,400, 100.00%1.9898 m
Uncertainty-aware representation1,363 of 2,400, 56.79% release-valid2,400 of 2,400, 100.00%1.9898 m point RMSE

Frame normalization raises exact scene validity by:

percentage points, while reducing mean positional RMSE by:

Hard projection raises physical feasibility by:

percentage points relative to frame normalization. Yet exact scene validity increases by only:

percentage points.

This is not a paradox. Projection solves a feasibility problem. It does not recover lost identity, establish the current twin revision, restore missing observations, or infer the correct operational zone.

Uncertainty-aware representation adds 32 release-valid scenes beyond the exact projected point state:

percentage points. All 32 additional release-valid scenes are occlusion-overcommitment cases that satisfy the common source-count and inter-source-spread rule in Equation (25a). No condition-specific width is applied. The gain remains benchmark-specific because the width formula and area ceiling are design choices, not field-calibrated uncertainty estimates.

6.2Interactive figure 1: Spatial observability matrix

The first interactive figure maps each fault class against nine diagnostic channels. It should be read as an observability matrix, not a causal attribution chart.

The matrix compares observation confidence, identity consistency, revision freshness, observation support, zone consistency, lineage completeness, association consistency, frame and projection integrity, and exact benchmark validity.

Figure 1 Spatial observability by failure mechanism Complete constructed benchmark

Mean diagnostic values for clean scenes and each material spatial fault, including the definition-bound exact benchmark validity column.

Heat map of nine geometry diagnostics across clean scenes and the material spatial fault classes. High signal is not universally good: a scene can be locally plausible and physically feasible while its reconstructed state describes the wrong world.

Figure 1. Spatial observability matrix. Mean diagnostic values for clean scenes and material fault scenes in the complete constructed benchmark. No single channel exposes every failure class.
Figure values

6.3Hard feasibility can hide wrong identity

Spatial plate D Identity collision Constructed scene

Interactive replacement for this plate. Plan view of a published constructed scene, with positions, zones, and footprints read from the study payload. Any of the six held-out scenes can be selected; the extruded view is an optional illustration.

Spatial plate D. Plausible positions can belong to the wrong objects.
Spatial plate D. Plausible positions can belong to the wrong objects.

Spatial plate D. Identity collision. The spatial set can remain plausible while the correspondence map is wrong. A collision checker cannot recover identity.

6.4Projection can erase the visible symptom

Spatial plate E Projection displacement Constructed scene

Interactive replacement for this plate. Plan view of a published constructed scene, with positions, zones, and footprints read from the study payload. Any of the six held-out scenes can be selected; the extruded view is an optional illustration.

Spatial plate E. Constraint projection repairs overlap but not provenance.
Spatial plate E. Constraint projection repairs overlap but not provenance.

Spatial plate E. Projection displacement. The scene becomes collision-free after projection. The repair may remove the most obvious symptom while leaving the reconstructed location unsupported by the authoritative scene.

6.5Topology is not reducible to occupancy

Spatial plate F Topology breach Constructed scene

Interactive replacement for this plate. Plan view of a published constructed scene, with positions, zones, and footprints read from the study payload. Any of the six held-out scenes can be selected; the extruded view is an optional illustration.

Spatial plate F. A collision-free state can violate operational topology.
Spatial plate F. A collision-free state can violate operational topology.

Spatial plate F. Topology breach. This selected held-out example is physically feasible after reconstruction but violates its authorized operational zone. Across the intervention family, pre-projection hard feasibility is not guaranteed.

7Release-policy results

7.1Operating-point measurements

For policy kk, invalid detection is:

Base-silent detection is:

Valid false-hold rate is:

Release coverage is:

Invalidity among released scenes, the selective risk, is defined when at least one test scene is released:

The scalar gate holds:

of invalid scenes and:

of base-silent invalid scenes.

The geometry trace gate holds:

of invalid scenes and:

of base-silent invalid scenes.

It false-holds:

of valid scenes.

The scalar policy releases 2,252 scenes, of which 923 are invalid:

The trace policy releases 1,429 scenes, of which 91 are invalid:

The reduction in released invalidity is therefore:

percentage points.

The corresponding reduction in release coverage is:

percentage points.

The stronger gate is not free. It routes substantially more scenes to review.

7.2Ranking and calibration

The Brier score is:

Ten-bin expected calibration error is:

MetricScalar confidence gateGeometry trace gate
AUROC for release validity0.608580.98231
Average precision for validity0.639460.98220
Brier score0.234310.03883
Ten-bin expected calibration error0.029790.01077

Here acc(Bm)\operatorname{acc}(B_m) is the empirical valid-scene fraction in bin BmB_m, and conf(Bm)\operatorname{conf}(B_m) is its mean fitted validity probability. The bins are the ten equal-width intervals [0,0.1),[0.1,0.2),,[0.8,0.9),[0.9,1][0,0.1),[0.1,0.2),\ldots,[0.8,0.9),[0.9,1]; empty bins contribute zero and are omitted from the displayed reliability points.

The scalar score provides weak discrimination for ranking validity. This is intentional. Several faults are constructed to remain locally coherent and highly confident. Confidence is not a substitute for frame, identity, revision, occupancy, topology, and support diagnostics.

7.3Nominal Wilson intervals

For an observed rate p^=x/n\widehat p=x/n, the two-sided nominal 95% Wilson interval is:

For trace invalid detection, the interval is approximately:

For trace base-silent detection:

For trace valid false holds:

These are finite-count Bernoulli summaries. They are not deployment confidence intervals, design-based intervals over all possible constructed benchmarks, or field-performance guarantees.

7.4Interactive figure 2: Fault-level geometry atlas

Figure 2 Invalid-scene detection by spatial condition Held-out test set

Base-silent detection for the two fitted policies, with the exact geometry oracle shown as a definition-bound ceiling.

Grouped comparison of detection by fault class for each release policy, with the exact oracle shown as a definition-bound ceiling rather than fitted detector performance. Every value is listed in the figure values table.

Figure 2. Invalid-scene detection by spatial intervention family. The grouped horizontal view compares scalar confidence, the geometry trace, and the exact oracle ceiling across the nine conditions. The families are not orthogonal validity components, so one intervention can activate several checks. Under the simplified four-fiducial construction, normalization corrects every held-out rigid-frame scene to release validity. That class therefore contributes no invalid denominator after the complete reconstruction sequence, although the trace policy false-holds 2 of its 288 valid scenes, or 0.69%.
Figure values

The two fitted residual classes are occlusion overcommitment and impossible occupancy. The trace policy holds 48 of 121 invalid held-out occlusion scenes, or 39.67%, leaving 73 invalid releases. It holds 167 of 185 invalid impossible-occupancy scenes, or 90.27%, leaving 18 invalid releases. It detects every invalid held-out case in identity collision, stale twin, topology breach, support gap, and compound-fault classes. Such complete separation is a property of the instrumented construction. It is not evidence that field systems expose equally clean diagnostics.

7.5Interactive figure 3: Progressive geometry-control ladder

The progressive figure independently fits five policies:

Evidence availableInvalid detectionBase-silent detectionValid false holdsRelease coverageResidual silent rateAUROCBrier
Observation confidence10.99%7.05%2.49%93.83%38.46%0.608580.23431
+ frame, revision, and projection integrity8.39%4.33%1.83%95.33%39.58%0.740000.20460
+ identity and association30.38%27.29%1.47%86.04%30.08%0.788100.17756
+ revision and zone diagnostics87.85%87.31%1.10%61.42%5.25%0.945430.05104
Full geometry trace91.22%90.84%1.83%59.54%3.79%0.982310.03883
Exact geometry oracle100%, by construction100%, by construction0%, by construction56.79%0%not applicablenot applicable
Figure 3 Progressive geometry-control ladder Refit and recalibrated per step

Invalid and base-silent detection as geometry evidence channels are added, each row fitted independently under the same split protocol.

Ladder of independently fitted rows showing invalid and base-silent detection as diagnostic channels are added. Each row is trained, calibrated, thresholded, and evaluated separately, so operating-point behaviour need not improve monotonically. Every value is listed in the figure values table and in the results tables in the article body.

Figure 3. Progressive geometry-control ladder. The largest gain appears only after revision and topology become explicit. Every fitted row is trained, calibrated, thresholded, and evaluated independently.
Figure values

7.6Interactive figure 4: Spatial risk versus release coverage

Selective classification trades coverage for lower released risk [13]. The fixed threshold grid preserves the irreversible base hold.

Figure 4 Spatial risk versus release coverage Held-out threshold grid

Exact step curves over the fixed threshold grid. The fitted policies cannot reverse the irreversible base hold.

Risk-coverage step curves over the disclosed threshold grid, with the irreversible base hold active at every point. Lower invalidity among released items requires a larger review population. The exact oracle is excluded because it is not a fitted threshold curve. Every value is listed in the figure values table.

Figure 4. Invalidity among released scenes versus release coverage. The geometry trace creates a materially different threshold curve from scalar confidence, but lower risk requires a larger review population. The exact oracle is excluded because it is not a fitted threshold curve.
Figure values

7.7Interactive figure 5: Validity-probability reliability

Figure 5 Reliability diagrams Held-out reliability bins

Mean predicted validity against empirical validity in ten equal-width bins for the two fitted policies. The oracle is excluded.

Reliability diagram comparing mean predicted validity with empirical validity in equal-width held-out probability bins for the two fitted policies. The oracle is excluded because it is not a calibrated probability model. Every value is listed in the figure values table.

Figure 5. Reliability of fitted scene-validity probabilities. Predicted release validity is compared with empirical validity in held-out probability bins. Global calibration is not proof that every spatial failure class is observable.
Figure values

8Fault interpretation

8.1Rigid frame mismatch

Frame mismatch is conceptually important because local geometry can remain intact. In the held-out class, reported-coordinate fusion is valid in 51 of 288 scenes, or 17.71%, while correspondence-based fiducial normalization makes all 288 scenes release-valid under the simplified anchor model. This is a recovery result under known, well-spread correspondences, not a claim about unconstrained field registration. The trace policy still false-holds 2 corrected scenes because it treats evidence of a large pre-normalization frame discrepancy conservatively.

A translation-only correction is insufficient when the source frame is rotated. If:

then subtracting only tt leaves:

This equals the world-frame point for every possible pp only when:

8.2Identity collision

Identity is not recoverable from non-overlap alone. A feasible assignment of boxes can correspond to an incorrect permutation. Because the benchmark objects have different footprints and authorized zones, an identity exchange can also create occupancy or topology failure. The intervention label names the assignment fault, not an isolated geometric predicate. The state space is therefore partly discrete:

where Sn\mathfrak S_n is the permutation group over object identities.

8.3Occlusion overcommitment

Occupancy-grid and probabilistic-robotics methods treat unknown space as a distinct state rather than automatically declaring it free or occupied [7,10]. The same principle applies at object level. Under partial observation, uncertainty should expand. A narrow point estimate is a stronger claim than the evidence supports.

The benchmark's uncertainty-aware stage rescues many occlusion cases because it asks whether the authoritative state lies inside a bounded region, not whether an unsupported point happens to be exact.

8.4Stale digital twin

A digital twin is not current merely because its geometry is internally consistent. The ISO 23247 series provides a structured digital-twin framework for observable manufacturing elements [14]. Operational use still requires an explicit rule identifying which revision controls at the decision boundary.

8.5Impossible occupancy

Hard projection is highly effective at removing constructed overlap and boundary violations under the explicit two-centimetre object-clearance rule. It is deliberately weak at identifying truth. An impossible-occupancy intervention can also move an object into the wrong operational zone. In the held-out benchmark, all projected scenes are physically feasible, yet 1,069 remain invalid as exact point states and 1,037 remain invalid after the bounded uncertainty rule is applied.

8.6Topology breach

Operational topology adds semantics to geometry. A room, aisle, hold zone, sterile area, restricted bay, or safety envelope is not equivalent to free Euclidean space. The injected topology intervention selects a wrong zone but does not guarantee pre-projection physical feasibility, so some cases also activate hard constraints. After projection, a state can be collision-free while remaining operationally forbidden.

8.7Support gap

A missing live observation should not be replaced silently by an older twin without a current-revision check. The base rule catches some support gaps as overt. The remainder require trace-level revision and support diagnostics.

9Exact recheck and conditional transfer

9.1Exact geometry oracle

The exact oracle is defined as:

Its hold rule is:

Perfect separation follows by construction. The oracle has no AUROC, average precision, Brier score, expected calibration error, or reliability curve. It is not a fitted detector and not measured Cortonex performance.

9.2Risk-prioritized exact recheck

The idealized recheck overlay applies exact benchmark validity to increasing fractions of trace-released scenes, ordered from lowest to highest fitted trace probability. It can add holds but cannot reverse an existing trace hold.

Share of trace releases recheckedAll-scene recheck coverageBase-silent detectionResidual silent rate over all scenesStress-suite release coverage
0%0.00%90.84%3.79%59.54%
10%5.96%94.56%2.25%58.00%
25%14.88%97.28%1.13%56.88%
50%29.75%99.19%0.33%56.08%
75%44.67%99.70%0.13%55.88%
100%59.54%100.00%0.00%55.75%
Figure 6 Exact geometry recheck budget Risk-prioritized; coverage budget

Base-silent detection and residual risk as a function of the fraction of trace-released scenes granted exact geometry recheck.

Dual-axis budget curve over increasing exact-recheck coverage. Greater idealized recheck coverage removes residual base-silent failures while preserving holds already imposed by the trace policy. This is an architectural budget ceiling, not measured verifier performance. Every value is listed in the figure values table.

Figure 6. Risk-prioritized exact geometry recheck. The curve is an architectural budget ceiling. It assumes exact access to benchmark validity for selected scenes and is not measured verifier performance.
Figure values

Full-overlay release remains below the standalone oracle's 56.79% because the overlay preserves the trace policy's 25 valid false holds.

9.3Conditional prevalence scenarios

Let:

  • ϱ\varrho be an assumed operating invalid-scene prevalence;
  • dkd_k be held-out invalid detection;
  • fkf_k be held-out valid false-hold rate.

Projected hold rate is:

Projected invalidity among released scenes is:

At an assumed invalid prevalence of 5%, the conditional calculations are:

PolicyEstimated holdEstimated releaseEstimated invalidity among releases
Scalar confidence gate2.92%97.08%4.58%
Geometry trace gate6.30%93.70%0.47%
Exact geometry oracle5.00%95.00%0%, by construction

These are sensitivity calculations, not deployment forecasts. They assume the benchmark's conditional detection and false-hold rates remain unchanged when prevalence changes.

10Interactive scene traces

The scene explorer contains six explicitly constructed held-out examples:

  1. rigid frame mismatch;
  2. identity collision;
  3. occlusion overcommitment;
  4. stale digital twin;
  5. impossible occupancy;
  6. topology breach.

For each scene, the explorer exposes:

  • typed scene family;
  • authoritative object-zone facts;
  • emitted object-zone facts;
  • current revision;
  • attached twin revision;
  • source-bundle description;
  • fault mechanism;
  • scalar confidence;
  • fitted trace probability and hold state;
  • exact oracle status.
Interactive Prevalence scenario calculator Scenario, not a field estimate

Projection of hold, release, and residual rates from the measured held-out conditional rates at a reader-chosen invalid prevalence.

Scenario calculator over an assumed invalid prevalence chosen by the reader, applying the measured held-out conditional rates of each policy. Scenario projection only: the prevalence is not estimated from the constructed stress suite, and the calculation assumes those rates transfer unchanged.

Conditional operating-prevalence scenarios. Estimated hold rate, release rate, and invalidity among released scenes under an assumed operational invalid prevalence. This is conditional algebra applied to the held-out operating points, not a forecast.
Figure values
Figure 7 Operational geometry trace explorer Held-out examples

One constructed scene per spatial condition: query, expected state, reconstructed state, diagnostics, and each gate decision.

Interactive trace explorer over six held-out constructed scenes, one per spatial condition. Each trace shows the authoritative geometry, reported observation, twin revision, operational zone, reconstructed state, injected mechanism, diagnostics, and the release decision under each policy.

Figure 7. Operational geometry trace explorer. The examples are constructed benchmark scenes, not customer records. Their purpose is to show how the same aggregate metric can arise from different geometric mechanisms. When a point estimate exceeds the 0.75 m tolerance but satisfies the bounded uncertainty rule, the explorer reports a nominal 95% support region rather than presenting the center point as an exact accepted location.

Print view shows the rigid frame mismatch trace. The remaining traces are available in the online version.

Figure values

11Study specification and verification

This study uses a programmatically constructed benchmark developed for controlled evaluation. The frozen seed is 70,629,43170{,}629{,}431. Exploratory development informed the benchmark and diagnostic-response design before the final protocol was frozen; the study was not preregistered. The final protocol nevertheless keeps fitting, calibration, threshold selection, and held-out testing disjoint. The public article reports:

  • scene-family composition;
  • facility geometry and object abstraction;
  • condition allocation;
  • materiality process;
  • source and reconstruction stages;
  • frame-normalization objective and fixed-fiducial construction;
  • physical and topological constraints;
  • uncertainty-support rule;
  • release-validity definition;
  • four-way data separation;
  • fitted-policy, diagnostic-response, and calibration method;
  • threshold-selection and numerical tie rule;
  • aggregate held-out results;
  • confidence intervals;
  • calibration and selective-risk measurements;
  • principal limitations.

The Cortonex Lab retains the versioned benchmark, scene-level outputs, fitted-policy state, generator, and separate verification materials used to produce the reported results. These internal research materials are not part of the public release.

All published operating-point values were regenerated from the frozen scene-level test output and checked through a separate verification path before packaging. This is internal technical verification, not outside certification or peer review.

Data and code availability. The experimental design, scene composition, fault mechanisms, metrics, aggregate results, geometric plates, and limitations are documented in this publication. Scene-level benchmark data, fitted-policy state, construction code, and private verification materials are retained by The Cortonex Lab and are not publicly distributed.

12Limitations

12.1Constructed scenes

The benchmark uses generated geometry, not field data. Its scene distributions, dimensions, fault allocation, materiality probability, noise scales, source weights, thresholds, and review budget are design choices.

12.2Planar state with illustrative extrusion

The evaluated state contains axis-aligned floor footprints and center positions only. The static geometric plates extrude those footprints using fixed illustrative heights that do not enter any benchmark label or metric. The study does not evaluate object height, arbitrary meshes, deformable objects, articulated systems, full six-degree-of-freedom pose, surface reconstruction, photometric consistency, or dynamic contact mechanics.

12.3Boundary-time observations rather than streams

Although the publication title refers to sensor streams, the benchmark reduces each channel to a constructed observation at one reconstruction boundary. It does not evaluate streaming latency, asynchronous arrival, dynamic filtering, temporal association, motion continuity, or sequence-level state estimation.

12.4Simplified registration

Four fiducial correspondences are known, well spread across the constructed floor, and corrupted only by modest independent coordinate noise. Real registration can fail through poor anchor geometry, outliers, moving anchors, repeated structure, partial overlap, nonrigid distortion, and degenerate viewpoints. ICP and related methods can converge to local solutions [3].

12.5Simplified data association

Identity faults are constructed as assignment collisions. The benchmark does not model long-horizon track birth and death, dense crowding, appearance drift, adversarial identity mimicry, or full multiple-hypothesis tracking.

12.6Hard projection is deliberately narrow

The projection stage enforces boundary, obstacle, and the explicit two-centimetre object-clearance rule through deterministic local search. It is not a general solver for nonconvex three-dimensional feasibility. It does not infer hidden semantics or recover authoritative state from physical constraints alone.

12.7Uncertainty regions are isotropic and constructed

The benchmark uses isotropic planar Gaussian-form support regions derived only from available-source count and post-registration inter-source spread. The same width rule is applied across all conditions, without intervention labels or materiality flags. Even so, the formula and the 22 square metre area ceiling are benchmark design choices rather than field-calibrated uncertainty estimates. The 95% contour is nominal and is not calibrated as a population coverage guarantee. Projection displacement does not widen the support region, so the construction can understate uncertainty after a large geometric repair. Hard feasibility is checked at the projected point centers; the full uncertainty contours are not truncated by obstacles, boundaries, other objects, or zone topology. Real sensor uncertainty can be anisotropic, multimodal, non-Gaussian, correlated across objects, and conditional on visibility, calibration, motion, or the repair operation itself.

12.8Operational topology is typed but simple

Zones are fixed rectangles, and zone validity is evaluated from each object center rather than full-footprint containment. Real operational topology can include access direction, reachability, containment hierarchies, temporal occupancy rules, safety envelopes, route constraints, and policy-dependent zone semantics.

12.9Digital-twin revision is explicit

The benchmark supplies current and prior revision identifiers. Field systems can have missing revisions, branching histories, inconsistent clocks, partial updates, and unclear authority.

12.10The base comparator is deliberately minimal

Every base hold occurs in a material support-gap scene with a missing required field, and the confidence threshold adds no unique hold beyond that missing-field condition. The 993 base-silent held-out scenes are therefore silent relative to this deliberately limited comparator. They are not an estimate of what would pass a mature Cortonex release policy or a well-instrumented production geometry system.

12.11The diagnostic response model is assumed

The geometry trace receives parameterized noisy responses to identity, revision, support, zone, lineage, association, and integrity state. The identity and association channels are constructed responses to latent benchmark state, not outputs of an independently evaluated signature or multi-object association algorithm. The integrity channel uses only fiducial residual, revision metadata, pre-projection feasibility, and projection displacement, but its response formula and reliability remain designed rather than field measured. These channels are cleaner and more directly instrumented than many deployed systems can provide. The 91.22% result is conditional on the disclosed response formulas and reliability settings. Field performance would depend on whether the same properties can be inferred from raw sources without access to benchmark truth.

12.12Fault families are not orthogonal

Condition names identify the injected mechanism. They do not partition final invalidity into exclusive causes. Identity exchange, impossible occupancy, topology, stale revision, and compound interventions can activate multiple positional, occupancy, topology, identity, or revision predicates at once.

12.13One primary frozen seed and no preregistration

The reported benchmark uses one primary frozen seed, 70,629,43170{,}629{,}431. Exploratory work informed the final design before freezing, and no preregistration was filed. Deterministic regeneration verifies the retained experiment. A limited post hoc internal sensitivity check over additional seeds is retained with the private verification record, but it was not preregistered, does not alter the primary tables, and does not establish robustness to structural benchmark choices.

12.14No operational cost model

The study reports release coverage and false holds but does not measure human review time, correction cost, latency, compute cost, sensor cost, or consequences of delay.

12.15No field prevalence

The held-out invalid prevalence is:

This high prevalence is created by the stress design. It is not an estimate of real operational invalidity.

12.16No production claim

The reported 91.22% invalid-scene detection is a held-out measurement on the disclosed constructed benchmark. It is not Cortonex production performance, a customer outcome, or a guarantee.

13Conclusion

Operational state is not established by collecting more coordinates. A defensible reconstruction must answer several independent questions:

  • Are observations expressed in the correct frame?
  • Do observations belong to the correct objects?
  • Is the digital-twin revision current?
  • Does the state satisfy hard physical constraints?
  • Does it satisfy operational topology?
  • Is the claimed precision supported by the available observations?
  • Can the release system distinguish a feasible repair from an authoritative state?

The benchmark exposes the consequences of collapsing these questions into one confidence value. Reported-coordinate fusion is often locally plausible. Frame normalization repairs a specific global failure class. Hard projection eliminates constructed physical impossibility. Uncertainty-aware representation prevents partial observation from becoming false precision. None of these alone establishes operational truth.

The strongest result is not the aggregate detection rate. It is the separation between physical feasibility and release validity. On the held-out stress suite, hard projection makes every scene physically feasible, yet only 55.46% are exactly valid as point states. A physically possible world can still be the wrong world.

The Cortonex Lab therefore treats geometry as evidence-bearing state. Coordinate frames, identities, revisions, constraints, topology, uncertainty, and release status must remain attached to the reconstruction. A world model should not be released because it looks coherent. It should be released only when the system can state what makes that coherence defensible.

Study specification and verification reference benchmark implementation, not production Cortonex software
Version1.7
Benchmark seed70629431
Benchmark size12,000 constructed scenes
Splits7,200 train / 1,200 calibration / 1,200 threshold / 2,400 test

Empirical status. A controlled spatial observability study on a programmatically constructed planar benchmark with an authoritative reference geometry at the reconstruction boundary, chosen so that the validity of every scene is known and silent spatial failure can be measured rather than estimated. Each channel contributes at most one boundary-time observation per object; sequence-level stream processing is not evaluated. The static plates use fixed illustrative extrusion heights that enter no label or metric. The fitted trace result is conditional on the explicitly parameterized diagnostic-response model disclosed in the article, and is not raw-sensor detection accuracy. Every quantity reported here is a benchmark measurement, not a customer record, production sensor log, field failure frequency, or measured Cortonex deployment result.

The base comparator is deliberately minimal: all 44 held-out base holds arise in support-gap scenes with a missing required field, and the confidence threshold adds no unique held-out hold. Base-silent scenes should therefore not be read as failures that a mature production policy would necessarily pass. The protocol keeps fitting, calibration, threshold selection, and held-out testing disjoint, and every published operating point was regenerated from the frozen scene-level output and checked through a separate verification path. Internal freezing is the assurance used here; the study was not preregistered or externally peer reviewed. Scene-level data, fitted-policy state, and the generator are retained by Cortonex and are not publicly distributed.

Cite this study

The Cortonex Lab. The Geometry of Operational State: Reconstructing Physically Coherent Reality from Sensor Streams, Digital Twins, and Incomplete Observations. Version 1.7. Cortonex Technologies Inc. https://cortonex.com/lab/geometry-of-operational-state/

@techreport{cortonexlab-geometry-of-operational-state,
  author      = {{The Cortonex Lab}},
  title       = {The Geometry of Operational State: Reconstructing Physically
                 Coherent Reality from Sensor Streams, Digital Twins, and
                 Incomplete Observations},
  institution = {The Cortonex Lab, Cortonex Technologies Inc.},
  version     = {1.7},
  url         = {https://cortonex.com/lab/geometry-of-operational-state/},
  note        = {Controlled constructed planar spatial benchmark;
                 no production-performance claim.}
}

References

  1. Horn, B. K. P. Closed-form solution of absolute orientation using unit quaternions. Journal of the Optical Society of America A, 4(4), 629-642, 1987.
  2. Arun, K. S., Huang, T. S., and Blostein, S. D. Least-squares fitting of two 3-D point sets. IEEE Transactions on Pattern Analysis and Machine Intelligence, 9(5), 698-700, 1987.
  3. Besl, P. J., and McKay, N. D. A method for registration of 3-D shapes. IEEE Transactions on Pattern Analysis and Machine Intelligence, 14(2), 239-256, 1992.
  4. Fischler, M. A., and Bolles, R. C. Random sample consensus: a paradigm for model fitting with applications to image analysis and automated cartography. Communications of the ACM, 24(6), 381-395, 1981.
  5. Reid, D. B. An algorithm for tracking multiple targets. IEEE Transactions on Automatic Control, 24(6), 843-854, 1979.
  6. Fortmann, T., Bar-Shalom, Y., and Scheffe, M. Sonar tracking of multiple targets using joint probabilistic data association. IEEE Journal of Oceanic Engineering, 8(3), 173-184, 1983.
  7. Elfes, A. Using occupancy grids for mobile robot perception and navigation. Computer, 22(6), 46-57, 1989.
  8. Durrant-Whyte, H., and Bailey, T. Simultaneous localization and mapping: part I. IEEE Robotics and Automation Magazine, 13(2), 99-110, 2006.
  9. Bailey, T., and Durrant-Whyte, H. Simultaneous localization and mapping: part II. IEEE Robotics and Automation Magazine, 13(3), 108-117, 2006.
  10. Thrun, S., Burgard, W., and Fox, D. Probabilistic Robotics. MIT Press, 2005.
  11. Bauschke, H. H., and Borwein, J. M. On projection algorithms for solving convex feasibility problems. SIAM Review, 38(3), 367-426, 1996.
  12. Guo, C., Pleiss, G., Sun, Y., and Weinberger, K. Q. On calibration of modern neural networks. Proceedings of the 34th International Conference on Machine Learning, 70, 1321-1330, 2017.
  13. Geifman, Y., and El-Yaniv, R. Selective classification for deep neural networks. Advances in Neural Information Processing Systems, 30, 2017.
  14. International Organization for Standardization. ISO 23247-1:2021, Automation systems and integration, digital twin framework for manufacturing, part 1: overview and general principles. 2021.