The Geometry of Operational State
Reconstructing physically coherent reality from sensor streams, digital twins, and incomplete observations.
Empirical status. This study uses a controlled, programmatically constructed spatial-state benchmark. Its three domain labels identify typed operational geometry families, not real clinical, warehouse, or industrial environments. Each channel is reduced to at most one observation per object at a reconstruction boundary; sequence-level stream processing is not evaluated. The study reports no customer records, production sensor logs, field failure frequencies, domain-performance estimates, or measured Cortonex deployment performance. The fitted trace result is conditional on the explicitly parameterized diagnostic-response model in Section 5.3. The exact geometry result is a deterministic oracle ceiling whose value equals benchmark release validity by construction.
A reconstructed operational scene can be composed entirely from individually plausible observations and still describe the wrong world. Sensor coordinates may be expressed in an unverified frame. Two objects may exchange identities while retaining plausible positions. A digital twin may be geometrically clean but one revision behind the physical environment. A collision-removal step may produce a feasible configuration at the wrong location. An occluded object may be represented as a precise point even though the observations support only a broad region. These are not ordinary data-quality defects. They are failures of spatial meaning.
We study such failures with a constructed benchmark containing 12,000 operational scenes, 48,000 typed objects, and 144,000 potential source-object observation slots. Scenes are divided equally among clinical spaces, warehouse operations, and industrial equipment bays. Each scene contains a bounded floor region, fixed obstacles, six operational zones, four typed objects, three observation channels, source-frame metadata, four fixed fiducials, digital-twin revision state, and an authoritative reference geometry at the reconstruction boundary.
The benchmark contains nine conditions: clean execution, rigid frame mismatch, identity collision, occlusion overcommitment, stale digital twin, impossible occupancy, topology breach, support gap, and compound geometry fault. Nonclean scenes receive a material fault with probability 0.86; the remaining cases are condition-stratified safe controls with ambiguous telemetry but a valid final state. Each source is represented by a boundary-time observation rather than a temporal sequence. We compare reported-coordinate fusion, correspondence-based fiducial normalization, hard physical constraint projection, uncertainty-aware state representation, scalar-confidence release, a geometry-trace release policy, and an exact benchmark oracle.
The held-out test population contains 2,400 scenes: 1,363 release-valid scenes, 1,037 invalid scenes, and 993 invalid scenes that the deliberately minimal base release rule would pass silently. All 44 held-out base holds arise in support-gap scenes with a missing required field. Reported-coordinate fusion is exactly valid in 45.50% of scenes and physically feasible in 78.29%. Fiducial frame normalization raises exact validity to 55.38% and feasibility to 85.08%. Hard constraint projection makes all 2,400 test scenes physically feasible, but exact validity rises by only two scenes, from 1,329 to 1,331. An uncertainty-aware representation raises release validity to 56.79% by replacing unsupported point certainty with bounded state regions whose widths depend only on available-source count and post-registration inter-source spread.
A scalar confidence gate holds 114 of 1,037 invalid scenes, or 10.99%, including 70 of 993 base-silent failures, or 7.05%. Under the diagnostic-response model specified in Section 5.3, a geometry-trace gate receiving noisy frame-registration, identity, revision, observation-support, zone, lineage, association, and source-integrity channels holds 946 invalid scenes, or 91.22%, including 902 base-silent failures, or 90.84%. It false-holds 25 of 1,363 valid scenes, or 1.83%. The scalar gate releases 93.83% of the stress suite with 40.99% invalidity among releases. The trace gate releases 59.54% with 6.37% invalidity among releases.
The result is an observability measurement under an explicitly parameterized constructed diagnostic-response model. It is not production performance or an independent estimate of raw-sensor detection accuracy. The architectural conclusion is narrower: local geometric plausibility does not establish operational spatial truth. A release system must treat coordinate frames, object identity, revision state, physical occupancy, operational topology, observation support, and uncertainty as separate evidence-bearing properties.
| Policy | Invalid scenes held | Base-silent scenes held | Valid false holds | Release coverage | Invalidity among released scenes |
|---|---|---|---|---|---|
| Scalar confidence gate | 114 of 1,037, 10.99% | 70 of 993, 7.05% | 34 of 1,363, 2.49% | 93.83% | 40.99% |
| Geometry trace gate | 946 of 1,037, 91.22% | 902 of 993, 90.84% | 25 of 1,363, 1.83% | 59.54% | 6.37% |
| Exact geometry oracle | 1,037 of 1,037, by construction | 993 of 993, by construction | 0 of 1,363, by construction | 56.79% | 0%, by construction |
A scene can be locally coherent and globally wrong
Rigid transformations preserve local geometry. If every object in a sensor frame is rotated and translated by the same transformation, pairwise distances remain unchanged:
because a rigid rotation satisfies:
A local consistency check can therefore report a coherent scene while the complete scene is displaced or rotated relative to the world frame. The problem is not that the observations contradict one another. The problem is that they agree in the wrong coordinate system.
Spatial plate A
Rigid-frame error
Constructed scene
Interactive replacement for this plate. Plan view of a published constructed scene, with positions, zones, and footprints read from the study payload. Any of the six held-out scenes can be selected; the extruded view is an optional illustration.
Spatial plate A. Rigid-frame error. A held-out constructed scene is shown in the authoritative world frame, after reported-coordinate fusion under a shared source-frame error, and after fiducial normalization. The example is illustrative of the benchmark mechanism. It is not a customer scene.
This distinction is central to operational state. A scene can satisfy all of the following and still be wrong:
- object-to-object distances look plausible;
- every source record is well formed;
- sensor confidence is high;
- no object overlaps another;
- the scene fits inside the mapped boundary;
- the assembled state is internally consistent.
Spatial truth requires more. The state must be expressed in the correct frame, attached to the correct identities, current to the controlling revision, supported at the claimed precision, physically feasible, and topologically consistent with the operational environment.
1Operational state as a constrained geometric object
An operational scene is not merely a set of coordinates. At reconstruction boundary , define:
where:
- is object identity;
- is the typed axis-aligned floor footprint;
- is floor position in the authoritative world frame;
- is operational zone or topological state;
- is the controlling revision;
The benchmark is deliberately planar. State estimation occurs in the floor plane. The static plates extrude the evaluated footprints using four fixed illustrative heights solely to make object identity and overlap legible; vertical height is not a generated, fitted, or evaluated benchmark variable. Object orientation is fixed by type and is not an evaluated degree of freedom. The construction is sufficient to study frame registration, identity, floor occupancy, zone topology, stale twins, and planar uncertainty without claiming arbitrary three-dimensional surface or pose reconstruction.
A source observes object through a source-specific measurement operator:
where denotes the relevant state variables and represents observation error. If the source reports coordinates in its own frame , the world-frame relation is:
with:
A reconstruction engine does not know and merely because a source supplies a coordinate. It must inherit, estimate, or verify the transform. Related localization and mapping systems make the same separation explicit: observations acquire spatial meaning through an estimated relation between sensor pose and map state [8,9]. In the benchmark, fixed fiducials provide independent frame anchors. The normalization stage estimates:
where and are corresponding fiducials. This is the classical absolute-orientation problem [1,2]. Iterative closest-point methods generalize registration when correspondences are not fixed in advance [3], while robust estimators such as random sample consensus address contaminated correspondence sets [4].
The benchmark intentionally keeps four fiducial correspondences known. Their reported coordinates are transformed by the same live-source frame error and independent anchor noise. A two-dimensional orthogonal Procrustes fit estimates the inverse rigid map from the fiducials alone, then normalizes the camera-like and ranging observations before they are fused again with the already world-referenced digital twin. Authoritative object positions are never used for registration.
1.1A scene is more than a point cloud
For each object, define its axis-aligned floor footprint under reconstructed pose:
Hard physical feasibility requires boundary inclusion:
obstacle exclusion:
and object clearance. Let denote the footprint expanded by m on every side. The benchmark requires the interiors of the expanded footprints to be disjoint:
Equation (11) is an axis-aligned two-centimetre object-to-object clearance rule. Interior disjointness permits the two expanded boundaries to touch, which corresponds to exactly 0.02 m of separation between the original footprints and matches the benchmark implementation. These conditions are necessary, not sufficient. An object can be collision-free and still occupy the wrong ward, aisle, bay, compartment, or control zone. Let:
map spatial position to operational zone. Topological consistency requires:
where is the authorized zone claim at boundary .
Identity consistency is separate again. Let map source detections to canonical object identities. A correct association requires:
If two objects exchange identities, the set of positions can remain plausible while the operational meaning is wrong. This is the classical data-association problem in another form. Multiple-hypothesis tracking and joint probabilistic data association were developed precisely because assignment uncertainty cannot be reduced to coordinate noise [5,6].
Revision consistency requires:
where is the controlling digital-twin or source revision at . A stale twin can be geometrically smooth, complete, and self-consistent while representing an earlier physical world.
1.2Contributions and scope
We make seven contributions.
First, we define operational spatial validity as a conjunction of world-frame positional support, including frame registration and bounded-uncertainty conditions, together with identity continuity, revision currency, physical feasibility, and zone topology.
Second, we construct a controlled benchmark in which each failure mechanism alters concrete sensor, transform, identity, revision, occupancy, topology, or support state rather than merely flipping an abstract error label.
Third, we separate four reconstruction stages: reported-coordinate fusion, fiducial frame normalization, hard constraint projection, and uncertainty-aware representation.
Fourth, we show experimentally that physical constraint satisfaction can increase sharply without a comparable increase in operational truth. In the held-out set, hard projection raises physical feasibility from 85.08% to 100%, while exact point-state validity increases from 55.38% to 55.46%.
Fifth, we compare confidence-only release with a geometry-trace policy under disjoint fitting, calibration, threshold-selection, and test populations.
Sixth, we provide interactive fault, calibration, risk-coverage, diagnostic-ladder, exact-recheck, and scene-trace views together with six static geometric plates.
Seventh, we state the boundary of the result. The benchmark does not establish real sensor performance, general three-dimensional reconstruction, field fault prevalence, digital-twin correctness, or measured Cortonex deployment performance.
2Benchmark construction
2.1Scene families
The benchmark contains:
constructed scenes, divided equally among:
| Typed scene family | Scenes | Constructed objects | Potential source-object observation slots |
|---|---|---|---|
| Clinical spaces | 4,000 | 16,000 | 48,000 |
| Warehouse operations | 4,000 | 16,000 | 48,000 |
| Industrial equipment bays | 4,000 | 16,000 | 48,000 |
| Total | 12,000 | 48,000 | 144,000 |
Each scene contains:
- a bounded floor region;
- six labeled operational zones;
- two fixed obstacle bands;
- four typed axis-aligned floor footprints, rendered as extruded cuboids in the static plates;
- camera-like, ranging, and digital-twin observation channels;
- source-frame and fiducial state;
- current and prior digital-twin revisions;
- an authoritative reference geometry at boundary .
The domain labels alter object types and dimensions. They do not reproduce clinical workflows, warehouse physics, industrial safety rules, real sensor characteristics, or domain-specific operating distributions. Each source contributes at most one boundary-time observation per object. The study does not evaluate stream latency, asynchronous delivery, temporal filtering, motion dynamics, or sequence-level tracking.
The six zone rectangles, expressed as in metres, are:
| Zone | Bounds |
|---|---|
| Staging | |
| Dock | |
| Quality hold | |
| Aisle west | |
| Aisle central | |
| Aisle east |
The fixed obstacle bands are and . Objects 1 through 4 are assigned, respectively, to aisle east, staging, aisle west, and quality hold. Their authoritative centers are sampled uniformly inside the assigned zone after accounting for half-footprint extent and an additional 0.15 m inward margin.
| Typed scene family | Object 1 footprint | Object 2 footprint | Object 3 footprint | Object 4 footprint |
|---|---|---|---|---|
| Clinical spaces | Bed, m | Supply cart, m | Infusion stand, m | Isolation cabinet, m |
| Warehouse operations | Pallet, m | Forklift, m | Autonomous cart, m | Quality cage, m |
| Industrial equipment bays | Machine skid, m | Service cart, m | Inspection unit, m | Tool cabinet, m |
Source and registration model
Before intervention, camera-like, ranging, and digital-twin positions are generated from the authoritative object centers with independent isotropic Gaussian coordinate noise:
| Channel | Coordinate-noise standard deviation | Baseline fusion weight | Frame state |
|---|---|---|---|
| Camera-like observation | 0.18 m per axis | 0.38 | Live-source frame |
| Ranging observation | 0.10 m per axis | 0.42 | Live-source frame |
| Digital-twin observation | 0.04 m per axis | 0.20 | World frame |
The baseline weights are fixed across all conditions and renormalized only over the channels available for each object. No injected condition name or materiality flag changes the reconstruction weights. The four fixed fiducials receive independent Gaussian coordinate noise with standard deviation 0.028 m per axis before rigid registration. The controlling twin revision is sampled uniformly from the integers 3 through 9, and the prior revision is one less.
For nonmaterial controls, a shared perturbation with coordinate standard deviation 0.18 m is applied at 0.25 scale to the camera-like channel and at negative 0.15 scale to the ranging channel. Nonmaterial occlusion and support-gap controls also remove one camera-like observation while preserving final validity. These values are benchmark design parameters, not estimates of deployed sensor accuracy.
2.2Condition allocation
Each scene receives one condition. Per scene family, the fixed allocation is:
| Condition | Per family | Complete benchmark |
|---|---|---|
| Clean execution | 1,280 | 3,840 |
| Rigid frame mismatch | 480 | 1,440 |
| Identity collision | 400 | 1,200 |
| Occlusion overcommitment | 440 | 1,320 |
| Stale digital twin | 400 | 1,200 |
| Impossible occupancy | 360 | 1,080 |
| Topology breach | 320 | 960 |
| Support gap | 200 | 600 |
| Compound geometry fault | 120 | 360 |
| Total | 4,000 | 12,000 |
For a nonclean condition, materiality is sampled as:
The frozen benchmark realized 7,015 material nonclean scenes. The 0.86 value and the condition allocation are stress-design choices. They are not estimates of field failure prevalence.
Within this benchmark, material is an intervention-strength flag, not a synonym for final invalidity. It selects the higher-amplitude branch of the named mechanism before normalization, projection, uncertainty representation, and release evaluation. A material frame mismatch can be corrected by fiducial registration, and a material occlusion intervention can become release-valid through a bounded uncertainty region. Materiality therefore describes the injected challenge, while validity describes the final reconstructed state.
Nonmaterial controls preserve the final valid state while introducing ambiguous telemetry, such as partial source absence or mild inter-source disagreement. They are assigned within the same condition strata, but they are not mechanism-matched counterfactuals for every material intervention. Their purpose is to prevent the fitted trace policy from treating every anomaly as a fault.
2.3Failure mechanisms
Rigid frame mismatch
The camera-like and ranging channels receive the same rotation and translation in their reported coordinate frame:
The pairwise geometry can remain plausible even when the world placement is wrong.
Identity collision
Two object assignments are exchanged through an incorrect association map:
The geometry may still contain the correct number of boxes, but the state belongs to the wrong objects.
Occlusion overcommitment
Live support is removed for one object, and the remaining estimate is biased while initially represented with unjustified precision. The fault is not merely positional error. It is the mismatch between evidential support and reported uncertainty.
Stale digital twin
A prior revision contributes state as if it were current:
Live observations for the affected object are partially absent, allowing the stale twin to dominate fusion.
Impossible occupancy
All sources agree on a placement intersecting a fixed obstacle or another object. This creates a physically impossible but observationally coherent state.
Topology breach
An object is assigned to a region outside its authorized operational zone:
The injected topology mechanism does not enforce hard feasibility before fusion. Some topology-labelled scenes therefore also activate boundary, obstacle, or object-clearance checks. The condition name identifies the intervention family, not an exclusive validity component.
Support gap
Current live observations are absent, and the system falls back to an older twin state. Some cases are overt because a required field is missing; others remain silent.
Compound geometry fault
A shared frame error is combined with either identity collision or stale-twin wrong-zone state.
The nine condition labels are intervention families rather than mutually exclusive failure predicates. An identity exchange can also create a footprint conflict because object dimensions differ. An impossible-occupancy placement can also violate zone topology. A topology intervention can be infeasible before projection. Fault-level tables must therefore be read as behavior under the named intervention, not as a decomposition into orthogonal causes.
The material intervention parameters are fixed as follows:
| Intervention family | Constructed parameterization |
|---|---|
| Rigid frame mismatch | Rotation magnitude uniform from 12 to 28 degrees with random sign; each translation coordinate uniform from -1.45 m to 1.45 m |
| Identity collision | Two distinct object assignments selected uniformly and exchanged in both live channels |
| Occlusion overcommitment | Camera absent for one object; ranging available with probability 0.35; twin displacement magnitude uniform from 0.95 m to 1.85 m in a random direction |
| Stale digital twin | Prior revision used; twin displacement magnitude uniform from 1.10 m to 2.50 m; camera and ranging each available with probability 0.45 |
| Impossible occupancy | With probability 0.55, the object is placed inside a fixed obstacle band; otherwise it is placed near another object |
| Topology breach | The object is moved near the center of a nonauthorized zone with isotropic coordinate perturbation of standard deviation 0.35 m |
| Support gap | Both live channels are absent; the prior twin is displaced by 0.90 m to 2.10 m; a required field is marked missing with probability 0.35 |
| Compound geometry fault | Rotation magnitude uniform from 10 to 23 degrees and translation coordinates uniform from -1.20 m to 1.20 m, followed with probability 0.55 by identity exchange and otherwise by stale wrong-zone twin state |
These choices define the stress suite. They are not empirical distributions of field faults.
2.4Data separation
The 12,000 scenes are partitioned before final evaluation:
| Population | Scenes | Purpose |
|---|---|---|
| Training | 7,200 | Fit policy coefficients |
| Probability calibration | 1,200 | Fit sigmoid probability calibration |
| Threshold selection | 1,200 | Select release thresholds under a valid false-hold budget |
| Held-out testing | 2,400 | Final measurement only |
The split is stratified by typed scene family and condition. Held-out aggregate validity labels do not estimate policy coefficients, calibration parameters, or release thresholds. Diagnostic channels in every split are generated by the same frozen response formulas, several of which use constructed latent predicates disclosed in Section 5.3. Let denote the held-out test population and define:
3Reconstruction stages
3.1Reported-coordinate fusion
For object , let be the available sources and their normalized weights. The baseline point estimate is:
This stage trusts each source's reported world coordinates and canonical identity.
3.2Fiducial frame normalization
The second stage estimates an inverse rigid transform from four fixed anchors using Equation (7), maps the camera-like and ranging observations into the world frame, and then repeats fusion while leaving the already world-referenced digital twin unchanged:
The benchmark does not use authoritative object positions to estimate this correction. Only fixed fiducials are available.
3.3Hard physical constraint projection
Let denote the normalized point reconstruction, with normalized fused object positions , and let be the set satisfying Equations (9) through (11). The ideal Euclidean projection is:
Projection methods are standard tools in convex feasibility [11]. The benchmark hard set is nonconvex because of pairwise clearance, and the study does not claim to solve Equation (24) globally. Its implementation is an order-dependent deterministic approximation. Objects are processed sequentially. Candidate displacements use radii from 0.25 m to 5.00 m in 0.25 m increments, with 24 equally spaced directions at each radius, plus the zero-displacement candidate. The nearest feasible candidate is selected. This is not a general optimizer for a nonconvex three-dimensional scene.
Even an exact selection from the minimizers in Equation (24) would choose a closest feasible state to the input, not necessarily the authoritative state. If the input is wrong because of identity, topology, or stale revision, a geometrically minimal repair can remain operationally wrong.
Spatial plate B
Feasible but wrong
Constructed scene
Interactive replacement for this plate. Plan view of a published constructed scene, with positions, zones, and footprints read from the study payload. Any of the six held-out scenes can be selected; the extruded view is an optional illustration.
Spatial plate B. Feasible but wrong. The hard projection removes the constructed occupancy violation. It cannot infer the authoritative position merely from the fact that a nearby feasible position exists.
3.4Uncertainty-aware representation
A point estimate is not always the correct output. For object , the benchmark uses an isotropic Gaussian-form support model:
Let be the number of available source channels in the frozen benchmark. The digital-twin channel is always retained, so the zero-source state is not evaluated. Let be the normalized weighted fusion from Section 3.2. The spread used by the benchmark is the unweighted root mean squared distance of the available normalized source points from that weighted fusion, . In Equation (25), is the projected point center from Section 3.3. The nominal scale is the explicit design rule:
No intervention label, condition name, materiality flag, authoritative reference error, or final validity label enters Equation (25a). The same width rule is applied to every object in every scene. The support model remains constructed and nominal, but it is label-independent at inference time.
The corresponding nominal support contour is:
For the planar benchmark:
For the isotropic planar model, contour area is:
The benchmark accepts uncertain positional support only when the authoritative reference is inside the nominal 95% contour and the contour area does not exceed a fixed ceiling. The word nominal is important: the construction does not claim frequentist 95% coverage in a field population.
The ceiling prevents an arbitrarily broad region from making every point trivially supported.
Spatial plate C
Partial observation
Constructed scene
Interactive replacement for this plate. Plan view of a published constructed scene, with positions, zones, and footprints read from the study payload. Any of the six held-out scenes can be selected; the extruded view is an optional illustration.
Spatial plate C. Partial observation. The uncertainty region exposes what the observations do not determine. It does not convert an unsupported point into a verified location.
4Validity and release
4.1Positional support
For point-valued stages, object error is:
Point position is accepted when:
Scene positional root mean squared error is:
The stage table reports the arithmetic mean of scene-level RMSE, so each scene receives equal weight:
This is not the pooled object-level RMSE. For reference, pooled object-level RMSE is 4.0097 m for reported-coordinate fusion, 3.8879 m after frame normalization, and 3.9061 m after projection. The article uses mean scene-level RMSE because the scene is the evaluation unit.
For the uncertainty-aware stage, an object is position-supported when either Equation (31) holds or Equations (28) and (29) hold. Hard physical feasibility is evaluated at the projected point centers. The benchmark does not require every point inside an uncertainty contour to satisfy the hard constraints.
4.2Release validity
From this section onward, index denotes a scene rather than an object. Define binary scene conditions:
- : every object is supported in the authoritative world frame, either by the point tolerance or by the bounded uncertainty rule;
- : object identities remain correct;
- : the controlling revision is current;
- : hard occupancy constraints hold;
- : operational zone topology is correct.
Frame error is evaluated through world-frame positional and zone support, while uncertainty honesty is part of . The benchmark release-validity label is therefore:
The multiplication denotes Boolean conjunction. It does not assume statistical independence.
A physically feasible scene satisfies only:
Therefore:
but:
Equation (36) is the central distinction in the study.
4.3Base hold and silent invalidity
Let when the base assembler visibly holds a scene because observation confidence is below 0.49 or a required object field is missing.
A base-silent invalid scene is:
An overt invalid scene is:
The held-out identities are:
and:
Thus, 993 invalid held-out scenes pass the ordinary base rule without a visible hold.
The base comparator is deliberately minimal. All 183 base holds, including all 44 held-out base holds, occur in material support-gap scenes with at least one required object field missing. Seventeen of those complete-benchmark scenes, including two held-out scenes, also fall below the scalar confidence threshold, but the missing required field already holds them. The base-silent count therefore means silent relative to this limited comparator, not silent relative to a mature production release policy.
5Release policies
5.1Scalar observation-confidence gate
The scalar policy receives only the one-world observation-confidence value. It does not see the condition label, materiality, authoritative scene, or benchmark validity bit.
5.2Geometry trace gate
The geometry trace policy receives eight noisy one-world diagnostics:
- observation confidence;
- constructed identity-consistency response;
- digital-twin revision freshness;
- observation support completeness;
- zone-claim consistency;
- spatial lineage completeness;
- constructed association-consistency response;
- frame, revision, and projection integrity.
5.3Constructed diagnostic-response model
The trace channels are generated by an explicit response model. They are not measurements from deployed sensors or a production reference monitor. For a latent channel value and declared reliability , the observed diagnostic is:
The exact latent definitions and response reliabilities are:
| Diagnostic | Latent value before response noise | Reliability |
|---|---|---|
| Constructed identity-consistency response | when identity is correct, otherwise | 0.86 |
| Revision freshness | when the controlling revision is current, otherwise | 0.87 |
| Observation support | Mean available-source fraction across the four objects | 0.78 |
| Zone-claim consistency | when every projected object is in its authorized zone, otherwise | 0.82 |
| Spatial-lineage completeness | , clipped to | 0.75 |
| Constructed association-consistency response | , clipped to | 0.76 |
| Frame, revision, and projection integrity | 0.84 |
Here is mean source availability, is mean normalized inter-source spread, and is the root mean squared discrepancy between the reported and known world-frame fiducial coordinates before rigid correction. In the lineage channel, for a current revision and when stale. The frame term is when post-registration fiducial residual is below m and otherwise. The association offset is for a correct identity and otherwise; when identity is incorrect, the complete association latent is capped at . In the integrity channel, for a current revision and otherwise, when the normalized point scene is physically feasible and otherwise, and is the sum of object-center displacement introduced by hard projection. No injected condition name or materiality flag enters this channel.
Observation confidence is generated separately. Let again denote mean source availability and let denote position agreement for scene . Its pre-noise center is:
where when a required object field is missing and is zero otherwise. The reported value is . Coherent faults can therefore remain confident through source agreement itself; the confidence generator does not receive the intervention class.
Several latent channels are noisy responses to benchmark validity components rather than outputs of independently evaluated raw-sensor algorithms. The trace result is therefore a conditional observability result under this diagnostic model. It is not an independent estimate of how accurately a deployed system would infer identity, revision, topology, or integrity from field data. On the held-out set, the individual channel AUROCs for validity range from 0.607 to 0.795; the combined policy gains its separation by composing several partially informative channels.
5.4Fitting and calibration
For policy with standardized feature vector , the fitted validity score is:
where:
The feature standardizer and an -regularized logistic model are fitted only on the training population. The implementation uses balanced class weights, regularization parameter , and the LBFGS solver. A separate sigmoid calibration map is then fitted on the calibration population using a near-unregularized one-dimensional logistic fit with :
Probability calibration and ranking answer different questions [12]. The study therefore reports Brier score, expected calibration error, ranking metrics, detection, false holds, coverage, and selective risk together.
5.5Threshold selection
A scene is held when the base rule already holds it or fitted validity falls below threshold:
The numerical tolerance prevents machine-level equality changes from altering a threshold tie.
Thresholds are chosen on the threshold-selection population to maximize invalid-scene detection subject to a 2% valid false-hold budget. When two feasible observed thresholds hold the same number of invalid scenes, the larger threshold is selected:
subject to:
The selected thresholds are:
The threshold-selection population contains 680 valid scenes, so the integer budget is:
Both selected policies hold exactly 13 valid threshold-selection scenes. Their separate held-out false-hold rates are allowed to differ.
6Reconstruction results
6.1Stage comparison
| Reconstruction stage | Scene validity | Physical feasibility | Mean scene-level positional RMSE |
|---|---|---|---|
| Reported-coordinate fusion | 1,092 of 2,400, 45.50% | 1,879 of 2,400, 78.29% | 2.2763 m |
| Fiducial frame normalization | 1,329 of 2,400, 55.38% | 2,042 of 2,400, 85.08% | 1.9840 m |
| Hard physical constraint projection | 1,331 of 2,400, 55.46% | 2,400 of 2,400, 100.00% | 1.9898 m |
| Uncertainty-aware representation | 1,363 of 2,400, 56.79% release-valid | 2,400 of 2,400, 100.00% | 1.9898 m point RMSE |
Frame normalization raises exact scene validity by:
percentage points, while reducing mean positional RMSE by:
Hard projection raises physical feasibility by:
percentage points relative to frame normalization. Yet exact scene validity increases by only:
percentage points.
This is not a paradox. Projection solves a feasibility problem. It does not recover lost identity, establish the current twin revision, restore missing observations, or infer the correct operational zone.
Uncertainty-aware representation adds 32 release-valid scenes beyond the exact projected point state:
percentage points. All 32 additional release-valid scenes are occlusion-overcommitment cases that satisfy the common source-count and inter-source-spread rule in Equation (25a). No condition-specific width is applied. The gain remains benchmark-specific because the width formula and area ceiling are design choices, not field-calibrated uncertainty estimates.
6.2Interactive figure 1: Spatial observability matrix
The first interactive figure maps each fault class against nine diagnostic channels. It should be read as an observability matrix, not a causal attribution chart.
The matrix compares observation confidence, identity consistency, revision freshness, observation support, zone consistency, lineage completeness, association consistency, frame and projection integrity, and exact benchmark validity.
Figure 1
Spatial observability by failure mechanism
Complete constructed benchmark
Mean diagnostic values for clean scenes and each material spatial fault, including the definition-bound exact benchmark validity column.
Heat map of nine geometry diagnostics across clean scenes and the material spatial fault classes. High signal is not universally good: a scene can be locally plausible and physically feasible while its reconstructed state describes the wrong world.
Figure values
6.3Hard feasibility can hide wrong identity
Spatial plate D
Identity collision
Constructed scene
Interactive replacement for this plate. Plan view of a published constructed scene, with positions, zones, and footprints read from the study payload. Any of the six held-out scenes can be selected; the extruded view is an optional illustration.
Spatial plate D. Identity collision. The spatial set can remain plausible while the correspondence map is wrong. A collision checker cannot recover identity.
6.4Projection can erase the visible symptom
Spatial plate E
Projection displacement
Constructed scene
Interactive replacement for this plate. Plan view of a published constructed scene, with positions, zones, and footprints read from the study payload. Any of the six held-out scenes can be selected; the extruded view is an optional illustration.
Spatial plate E. Projection displacement. The scene becomes collision-free after projection. The repair may remove the most obvious symptom while leaving the reconstructed location unsupported by the authoritative scene.
6.5Topology is not reducible to occupancy
Spatial plate F
Topology breach
Constructed scene
Interactive replacement for this plate. Plan view of a published constructed scene, with positions, zones, and footprints read from the study payload. Any of the six held-out scenes can be selected; the extruded view is an optional illustration.
Spatial plate F. Topology breach. This selected held-out example is physically feasible after reconstruction but violates its authorized operational zone. Across the intervention family, pre-projection hard feasibility is not guaranteed.
7Release-policy results
7.1Operating-point measurements
For policy , invalid detection is:
Base-silent detection is:
Valid false-hold rate is:
Release coverage is:
Invalidity among released scenes, the selective risk, is defined when at least one test scene is released:
The scalar gate holds:
of invalid scenes and:
of base-silent invalid scenes.
The geometry trace gate holds:
of invalid scenes and:
of base-silent invalid scenes.
It false-holds:
of valid scenes.
The scalar policy releases 2,252 scenes, of which 923 are invalid:
The trace policy releases 1,429 scenes, of which 91 are invalid:
The reduction in released invalidity is therefore:
percentage points.
The corresponding reduction in release coverage is:
percentage points.
The stronger gate is not free. It routes substantially more scenes to review.
7.2Ranking and calibration
The Brier score is:
Ten-bin expected calibration error is:
| Metric | Scalar confidence gate | Geometry trace gate |
|---|---|---|
| AUROC for release validity | 0.60858 | 0.98231 |
| Average precision for validity | 0.63946 | 0.98220 |
| Brier score | 0.23431 | 0.03883 |
| Ten-bin expected calibration error | 0.02979 | 0.01077 |
Here is the empirical valid-scene fraction in bin , and is its mean fitted validity probability. The bins are the ten equal-width intervals ; empty bins contribute zero and are omitted from the displayed reliability points.
The scalar score provides weak discrimination for ranking validity. This is intentional. Several faults are constructed to remain locally coherent and highly confident. Confidence is not a substitute for frame, identity, revision, occupancy, topology, and support diagnostics.
7.3Nominal Wilson intervals
For an observed rate , the two-sided nominal 95% Wilson interval is:
For trace invalid detection, the interval is approximately:
For trace base-silent detection:
For trace valid false holds:
These are finite-count Bernoulli summaries. They are not deployment confidence intervals, design-based intervals over all possible constructed benchmarks, or field-performance guarantees.
7.4Interactive figure 2: Fault-level geometry atlas
Figure 2
Invalid-scene detection by spatial condition
Held-out test set
Base-silent detection for the two fitted policies, with the exact geometry oracle shown as a definition-bound ceiling.
Grouped comparison of detection by fault class for each release policy, with the exact oracle shown as a definition-bound ceiling rather than fitted detector performance. Every value is listed in the figure values table.
Figure values
The two fitted residual classes are occlusion overcommitment and impossible occupancy. The trace policy holds 48 of 121 invalid held-out occlusion scenes, or 39.67%, leaving 73 invalid releases. It holds 167 of 185 invalid impossible-occupancy scenes, or 90.27%, leaving 18 invalid releases. It detects every invalid held-out case in identity collision, stale twin, topology breach, support gap, and compound-fault classes. Such complete separation is a property of the instrumented construction. It is not evidence that field systems expose equally clean diagnostics.
7.5Interactive figure 3: Progressive geometry-control ladder
The progressive figure independently fits five policies:
| Evidence available | Invalid detection | Base-silent detection | Valid false holds | Release coverage | Residual silent rate | AUROC | Brier |
|---|---|---|---|---|---|---|---|
| Observation confidence | 10.99% | 7.05% | 2.49% | 93.83% | 38.46% | 0.60858 | 0.23431 |
| + frame, revision, and projection integrity | 8.39% | 4.33% | 1.83% | 95.33% | 39.58% | 0.74000 | 0.20460 |
| + identity and association | 30.38% | 27.29% | 1.47% | 86.04% | 30.08% | 0.78810 | 0.17756 |
| + revision and zone diagnostics | 87.85% | 87.31% | 1.10% | 61.42% | 5.25% | 0.94543 | 0.05104 |
| Full geometry trace | 91.22% | 90.84% | 1.83% | 59.54% | 3.79% | 0.98231 | 0.03883 |
| Exact geometry oracle | 100%, by construction | 100%, by construction | 0%, by construction | 56.79% | 0% | not applicable | not applicable |
Figure 3
Progressive geometry-control ladder
Refit and recalibrated per step
Invalid and base-silent detection as geometry evidence channels are added, each row fitted independently under the same split protocol.
Ladder of independently fitted rows showing invalid and base-silent detection as diagnostic channels are added. Each row is trained, calibrated, thresholded, and evaluated separately, so operating-point behaviour need not improve monotonically. Every value is listed in the figure values table and in the results tables in the article body.
Figure values
7.6Interactive figure 4: Spatial risk versus release coverage
Selective classification trades coverage for lower released risk [13]. The fixed threshold grid preserves the irreversible base hold.
Figure 4
Spatial risk versus release coverage
Held-out threshold grid
Exact step curves over the fixed threshold grid. The fitted policies cannot reverse the irreversible base hold.
Risk-coverage step curves over the disclosed threshold grid, with the irreversible base hold active at every point. Lower invalidity among released items requires a larger review population. The exact oracle is excluded because it is not a fitted threshold curve. Every value is listed in the figure values table.
Figure values
7.7Interactive figure 5: Validity-probability reliability
Figure 5
Reliability diagrams
Held-out reliability bins
Mean predicted validity against empirical validity in ten equal-width bins for the two fitted policies. The oracle is excluded.
Reliability diagram comparing mean predicted validity with empirical validity in equal-width held-out probability bins for the two fitted policies. The oracle is excluded because it is not a calibrated probability model. Every value is listed in the figure values table.
Figure values
8Fault interpretation
8.1Rigid frame mismatch
Frame mismatch is conceptually important because local geometry can remain intact. In the held-out class, reported-coordinate fusion is valid in 51 of 288 scenes, or 17.71%, while correspondence-based fiducial normalization makes all 288 scenes release-valid under the simplified anchor model. This is a recovery result under known, well-spread correspondences, not a claim about unconstrained field registration. The trace policy still false-holds 2 corrected scenes because it treats evidence of a large pre-normalization frame discrepancy conservatively.
A translation-only correction is insufficient when the source frame is rotated. If:
then subtracting only leaves:
This equals the world-frame point for every possible only when:
8.2Identity collision
Identity is not recoverable from non-overlap alone. A feasible assignment of boxes can correspond to an incorrect permutation. Because the benchmark objects have different footprints and authorized zones, an identity exchange can also create occupancy or topology failure. The intervention label names the assignment fault, not an isolated geometric predicate. The state space is therefore partly discrete:
where is the permutation group over object identities.
8.3Occlusion overcommitment
Occupancy-grid and probabilistic-robotics methods treat unknown space as a distinct state rather than automatically declaring it free or occupied [7,10]. The same principle applies at object level. Under partial observation, uncertainty should expand. A narrow point estimate is a stronger claim than the evidence supports.
The benchmark's uncertainty-aware stage rescues many occlusion cases because it asks whether the authoritative state lies inside a bounded region, not whether an unsupported point happens to be exact.
8.4Stale digital twin
A digital twin is not current merely because its geometry is internally consistent. The ISO 23247 series provides a structured digital-twin framework for observable manufacturing elements [14]. Operational use still requires an explicit rule identifying which revision controls at the decision boundary.
8.5Impossible occupancy
Hard projection is highly effective at removing constructed overlap and boundary violations under the explicit two-centimetre object-clearance rule. It is deliberately weak at identifying truth. An impossible-occupancy intervention can also move an object into the wrong operational zone. In the held-out benchmark, all projected scenes are physically feasible, yet 1,069 remain invalid as exact point states and 1,037 remain invalid after the bounded uncertainty rule is applied.
8.6Topology breach
Operational topology adds semantics to geometry. A room, aisle, hold zone, sterile area, restricted bay, or safety envelope is not equivalent to free Euclidean space. The injected topology intervention selects a wrong zone but does not guarantee pre-projection physical feasibility, so some cases also activate hard constraints. After projection, a state can be collision-free while remaining operationally forbidden.
8.7Support gap
A missing live observation should not be replaced silently by an older twin without a current-revision check. The base rule catches some support gaps as overt. The remainder require trace-level revision and support diagnostics.
9Exact recheck and conditional transfer
9.1Exact geometry oracle
The exact oracle is defined as:
Its hold rule is:
Perfect separation follows by construction. The oracle has no AUROC, average precision, Brier score, expected calibration error, or reliability curve. It is not a fitted detector and not measured Cortonex performance.
9.2Risk-prioritized exact recheck
The idealized recheck overlay applies exact benchmark validity to increasing fractions of trace-released scenes, ordered from lowest to highest fitted trace probability. It can add holds but cannot reverse an existing trace hold.
| Share of trace releases rechecked | All-scene recheck coverage | Base-silent detection | Residual silent rate over all scenes | Stress-suite release coverage |
|---|---|---|---|---|
| 0% | 0.00% | 90.84% | 3.79% | 59.54% |
| 10% | 5.96% | 94.56% | 2.25% | 58.00% |
| 25% | 14.88% | 97.28% | 1.13% | 56.88% |
| 50% | 29.75% | 99.19% | 0.33% | 56.08% |
| 75% | 44.67% | 99.70% | 0.13% | 55.88% |
| 100% | 59.54% | 100.00% | 0.00% | 55.75% |
Figure 6
Exact geometry recheck budget
Risk-prioritized; coverage budget
Base-silent detection and residual risk as a function of the fraction of trace-released scenes granted exact geometry recheck.
Dual-axis budget curve over increasing exact-recheck coverage. Greater idealized recheck coverage removes residual base-silent failures while preserving holds already imposed by the trace policy. This is an architectural budget ceiling, not measured verifier performance. Every value is listed in the figure values table.
Figure values
Full-overlay release remains below the standalone oracle's 56.79% because the overlay preserves the trace policy's 25 valid false holds.
9.3Conditional prevalence scenarios
Let:
- be an assumed operating invalid-scene prevalence;
- be held-out invalid detection;
- be held-out valid false-hold rate.
Projected hold rate is:
Projected invalidity among released scenes is:
At an assumed invalid prevalence of 5%, the conditional calculations are:
| Policy | Estimated hold | Estimated release | Estimated invalidity among releases |
|---|---|---|---|
| Scalar confidence gate | 2.92% | 97.08% | 4.58% |
| Geometry trace gate | 6.30% | 93.70% | 0.47% |
| Exact geometry oracle | 5.00% | 95.00% | 0%, by construction |
These are sensitivity calculations, not deployment forecasts. They assume the benchmark's conditional detection and false-hold rates remain unchanged when prevalence changes.
10Interactive scene traces
The scene explorer contains six explicitly constructed held-out examples:
- rigid frame mismatch;
- identity collision;
- occlusion overcommitment;
- stale digital twin;
- impossible occupancy;
- topology breach.
For each scene, the explorer exposes:
- typed scene family;
- authoritative object-zone facts;
- emitted object-zone facts;
- current revision;
- attached twin revision;
- source-bundle description;
- fault mechanism;
- scalar confidence;
- fitted trace probability and hold state;
- exact oracle status.
Interactive
Prevalence scenario calculator
Scenario, not a field estimate
Projection of hold, release, and residual rates from the measured held-out conditional rates at a reader-chosen invalid prevalence.
Scenario calculator over an assumed invalid prevalence chosen by the reader, applying the measured held-out conditional rates of each policy. Scenario projection only: the prevalence is not estimated from the constructed stress suite, and the calculation assumes those rates transfer unchanged.
Figure values
Figure 7
Operational geometry trace explorer
Held-out examples
One constructed scene per spatial condition: query, expected state, reconstructed state, diagnostics, and each gate decision.
Interactive trace explorer over six held-out constructed scenes, one per spatial condition. Each trace shows the authoritative geometry, reported observation, twin revision, operational zone, reconstructed state, injected mechanism, diagnostics, and the release decision under each policy.
Print view shows the rigid frame mismatch trace. The remaining traces are available in the online version.
Figure values
11Study specification and verification
This study uses a programmatically constructed benchmark developed for controlled evaluation. The frozen seed is . Exploratory development informed the benchmark and diagnostic-response design before the final protocol was frozen; the study was not preregistered. The final protocol nevertheless keeps fitting, calibration, threshold selection, and held-out testing disjoint. The public article reports:
- scene-family composition;
- facility geometry and object abstraction;
- condition allocation;
- materiality process;
- source and reconstruction stages;
- frame-normalization objective and fixed-fiducial construction;
- physical and topological constraints;
- uncertainty-support rule;
- release-validity definition;
- four-way data separation;
- fitted-policy, diagnostic-response, and calibration method;
- threshold-selection and numerical tie rule;
- aggregate held-out results;
- confidence intervals;
- calibration and selective-risk measurements;
- principal limitations.
The Cortonex Lab retains the versioned benchmark, scene-level outputs, fitted-policy state, generator, and separate verification materials used to produce the reported results. These internal research materials are not part of the public release.
All published operating-point values were regenerated from the frozen scene-level test output and checked through a separate verification path before packaging. This is internal technical verification, not outside certification or peer review.
Data and code availability. The experimental design, scene composition, fault mechanisms, metrics, aggregate results, geometric plates, and limitations are documented in this publication. Scene-level benchmark data, fitted-policy state, construction code, and private verification materials are retained by The Cortonex Lab and are not publicly distributed.
12Limitations
12.1Constructed scenes
The benchmark uses generated geometry, not field data. Its scene distributions, dimensions, fault allocation, materiality probability, noise scales, source weights, thresholds, and review budget are design choices.
12.2Planar state with illustrative extrusion
The evaluated state contains axis-aligned floor footprints and center positions only. The static geometric plates extrude those footprints using fixed illustrative heights that do not enter any benchmark label or metric. The study does not evaluate object height, arbitrary meshes, deformable objects, articulated systems, full six-degree-of-freedom pose, surface reconstruction, photometric consistency, or dynamic contact mechanics.
12.3Boundary-time observations rather than streams
Although the publication title refers to sensor streams, the benchmark reduces each channel to a constructed observation at one reconstruction boundary. It does not evaluate streaming latency, asynchronous arrival, dynamic filtering, temporal association, motion continuity, or sequence-level state estimation.
12.4Simplified registration
Four fiducial correspondences are known, well spread across the constructed floor, and corrupted only by modest independent coordinate noise. Real registration can fail through poor anchor geometry, outliers, moving anchors, repeated structure, partial overlap, nonrigid distortion, and degenerate viewpoints. ICP and related methods can converge to local solutions [3].
12.5Simplified data association
Identity faults are constructed as assignment collisions. The benchmark does not model long-horizon track birth and death, dense crowding, appearance drift, adversarial identity mimicry, or full multiple-hypothesis tracking.
12.6Hard projection is deliberately narrow
The projection stage enforces boundary, obstacle, and the explicit two-centimetre object-clearance rule through deterministic local search. It is not a general solver for nonconvex three-dimensional feasibility. It does not infer hidden semantics or recover authoritative state from physical constraints alone.
12.7Uncertainty regions are isotropic and constructed
The benchmark uses isotropic planar Gaussian-form support regions derived only from available-source count and post-registration inter-source spread. The same width rule is applied across all conditions, without intervention labels or materiality flags. Even so, the formula and the 22 square metre area ceiling are benchmark design choices rather than field-calibrated uncertainty estimates. The 95% contour is nominal and is not calibrated as a population coverage guarantee. Projection displacement does not widen the support region, so the construction can understate uncertainty after a large geometric repair. Hard feasibility is checked at the projected point centers; the full uncertainty contours are not truncated by obstacles, boundaries, other objects, or zone topology. Real sensor uncertainty can be anisotropic, multimodal, non-Gaussian, correlated across objects, and conditional on visibility, calibration, motion, or the repair operation itself.
12.8Operational topology is typed but simple
Zones are fixed rectangles, and zone validity is evaluated from each object center rather than full-footprint containment. Real operational topology can include access direction, reachability, containment hierarchies, temporal occupancy rules, safety envelopes, route constraints, and policy-dependent zone semantics.
12.9Digital-twin revision is explicit
The benchmark supplies current and prior revision identifiers. Field systems can have missing revisions, branching histories, inconsistent clocks, partial updates, and unclear authority.
12.10The base comparator is deliberately minimal
Every base hold occurs in a material support-gap scene with a missing required field, and the confidence threshold adds no unique hold beyond that missing-field condition. The 993 base-silent held-out scenes are therefore silent relative to this deliberately limited comparator. They are not an estimate of what would pass a mature Cortonex release policy or a well-instrumented production geometry system.
12.11The diagnostic response model is assumed
The geometry trace receives parameterized noisy responses to identity, revision, support, zone, lineage, association, and integrity state. The identity and association channels are constructed responses to latent benchmark state, not outputs of an independently evaluated signature or multi-object association algorithm. The integrity channel uses only fiducial residual, revision metadata, pre-projection feasibility, and projection displacement, but its response formula and reliability remain designed rather than field measured. These channels are cleaner and more directly instrumented than many deployed systems can provide. The 91.22% result is conditional on the disclosed response formulas and reliability settings. Field performance would depend on whether the same properties can be inferred from raw sources without access to benchmark truth.
12.12Fault families are not orthogonal
Condition names identify the injected mechanism. They do not partition final invalidity into exclusive causes. Identity exchange, impossible occupancy, topology, stale revision, and compound interventions can activate multiple positional, occupancy, topology, identity, or revision predicates at once.
12.13One primary frozen seed and no preregistration
The reported benchmark uses one primary frozen seed, . Exploratory work informed the final design before freezing, and no preregistration was filed. Deterministic regeneration verifies the retained experiment. A limited post hoc internal sensitivity check over additional seeds is retained with the private verification record, but it was not preregistered, does not alter the primary tables, and does not establish robustness to structural benchmark choices.
12.14No operational cost model
The study reports release coverage and false holds but does not measure human review time, correction cost, latency, compute cost, sensor cost, or consequences of delay.
12.15No field prevalence
The held-out invalid prevalence is:
This high prevalence is created by the stress design. It is not an estimate of real operational invalidity.
12.16No production claim
The reported 91.22% invalid-scene detection is a held-out measurement on the disclosed constructed benchmark. It is not Cortonex production performance, a customer outcome, or a guarantee.
13Conclusion
Operational state is not established by collecting more coordinates. A defensible reconstruction must answer several independent questions:
- Are observations expressed in the correct frame?
- Do observations belong to the correct objects?
- Is the digital-twin revision current?
- Does the state satisfy hard physical constraints?
- Does it satisfy operational topology?
- Is the claimed precision supported by the available observations?
- Can the release system distinguish a feasible repair from an authoritative state?
The benchmark exposes the consequences of collapsing these questions into one confidence value. Reported-coordinate fusion is often locally plausible. Frame normalization repairs a specific global failure class. Hard projection eliminates constructed physical impossibility. Uncertainty-aware representation prevents partial observation from becoming false precision. None of these alone establishes operational truth.
The strongest result is not the aggregate detection rate. It is the separation between physical feasibility and release validity. On the held-out stress suite, hard projection makes every scene physically feasible, yet only 55.46% are exactly valid as point states. A physically possible world can still be the wrong world.
The Cortonex Lab therefore treats geometry as evidence-bearing state. Coordinate frames, identities, revisions, constraints, topology, uncertainty, and release status must remain attached to the reconstruction. A world model should not be released because it looks coherent. It should be released only when the system can state what makes that coherence defensible.
reference benchmark implementation, not production Cortonex software
Empirical status. A controlled spatial observability study on a programmatically constructed planar benchmark with an authoritative reference geometry at the reconstruction boundary, chosen so that the validity of every scene is known and silent spatial failure can be measured rather than estimated. Each channel contributes at most one boundary-time observation per object; sequence-level stream processing is not evaluated. The static plates use fixed illustrative extrusion heights that enter no label or metric. The fitted trace result is conditional on the explicitly parameterized diagnostic-response model disclosed in the article, and is not raw-sensor detection accuracy. Every quantity reported here is a benchmark measurement, not a customer record, production sensor log, field failure frequency, or measured Cortonex deployment result.
The base comparator is deliberately minimal: all 44 held-out base holds arise in support-gap scenes with a missing required field, and the confidence threshold adds no unique held-out hold. Base-silent scenes should therefore not be read as failures that a mature production policy would necessarily pass. The protocol keeps fitting, calibration, threshold selection, and held-out testing disjoint, and every published operating point was regenerated from the frozen scene-level output and checked through a separate verification path. Internal freezing is the assurance used here; the study was not preregistered or externally peer reviewed. Scene-level data, fitted-policy state, and the generator are retained by Cortonex and are not publicly distributed.
The Cortonex Lab. The Geometry of Operational State: Reconstructing Physically Coherent Reality from Sensor Streams, Digital Twins, and Incomplete Observations. Version 1.7. Cortonex Technologies Inc. https://cortonex.com/lab/geometry-of-operational-state/
@techreport{cortonexlab-geometry-of-operational-state,
author = {{The Cortonex Lab}},
title = {The Geometry of Operational State: Reconstructing Physically
Coherent Reality from Sensor Streams, Digital Twins, and
Incomplete Observations},
institution = {The Cortonex Lab, Cortonex Technologies Inc.},
version = {1.7},
url = {https://cortonex.com/lab/geometry-of-operational-state/},
note = {Controlled constructed planar spatial benchmark;
no production-performance claim.}
}
References
- Horn, B. K. P. Closed-form solution of absolute orientation using unit quaternions. Journal of the Optical Society of America A, 4(4), 629-642, 1987.
- Arun, K. S., Huang, T. S., and Blostein, S. D. Least-squares fitting of two 3-D point sets. IEEE Transactions on Pattern Analysis and Machine Intelligence, 9(5), 698-700, 1987.
- Besl, P. J., and McKay, N. D. A method for registration of 3-D shapes. IEEE Transactions on Pattern Analysis and Machine Intelligence, 14(2), 239-256, 1992.
- Fischler, M. A., and Bolles, R. C. Random sample consensus: a paradigm for model fitting with applications to image analysis and automated cartography. Communications of the ACM, 24(6), 381-395, 1981.
- Reid, D. B. An algorithm for tracking multiple targets. IEEE Transactions on Automatic Control, 24(6), 843-854, 1979.
- Fortmann, T., Bar-Shalom, Y., and Scheffe, M. Sonar tracking of multiple targets using joint probabilistic data association. IEEE Journal of Oceanic Engineering, 8(3), 173-184, 1983.
- Elfes, A. Using occupancy grids for mobile robot perception and navigation. Computer, 22(6), 46-57, 1989.
- Durrant-Whyte, H., and Bailey, T. Simultaneous localization and mapping: part I. IEEE Robotics and Automation Magazine, 13(2), 99-110, 2006.
- Bailey, T., and Durrant-Whyte, H. Simultaneous localization and mapping: part II. IEEE Robotics and Automation Magazine, 13(3), 108-117, 2006.
- Thrun, S., Burgard, W., and Fox, D. Probabilistic Robotics. MIT Press, 2005.
- Bauschke, H. H., and Borwein, J. M. On projection algorithms for solving convex feasibility problems. SIAM Review, 38(3), 367-426, 1996.
- Guo, C., Pleiss, G., Sun, Y., and Weinberger, K. Q. On calibration of modern neural networks. Proceedings of the 34th International Conference on Machine Learning, 70, 1321-1330, 2017.
- Geifman, Y., and El-Yaniv, R. Selective classification for deep neural networks. Advances in Neural Information Processing Systems, 30, 2017.
- International Organization for Standardization. ISO 23247-1:2021, Automation systems and integration, digital twin framework for manufacturing, part 1: overview and general principles. 2021.