Failure Modes · Benchmark

Where Evidence Pipelines Fail

A controlled fault-injection study of source fidelity, citation lineage, and governed release.

Constructed benchmark Six-stage evidence pipeline Six-document bundle per case Three release policies compared

Empirical status. This is a controlled fault-injection study on a programmatically constructed benchmark in which the ground truth of every record is known exactly. Exact reference state is the point of the design: real corpora cannot supply ground-truth validity labels, so silent failure can only be estimated on them, never measured. The three domain labels denote constructed schema families, not clinical, legal, or supply-chain evidence, and no figure here is a customer record, a production log, an industry failure rate, or a measured Cortonex deployment result. The exact current-source oracle equals benchmark validity by construction; it is a logical ceiling, not a detector.

Abstract

Source-grounded systems can fail while their transformed evidence paths remain internally consistent and their scalar confidence signals remain high. We study this problem with an instrumented reference pipeline spanning capture, segmentation, indexing, retrieval, synthesis, citation attachment, and governed release. The benchmark contains 24,000 constructed, versioned records divided equally among three domain-labelled schema families. Each case is evaluated against its own six-document candidate bundle; the bundles are not merged into a shared retrieval corpus. After the experimental protocol was frozen, a new seed generated the reported benchmark. The resulting split contains 14,401 training records, a 4,802-record validation pool divided into disjoint calibration and threshold-selection subsets, and 4,797 final test records.

A released record is valid only if its two emitted facts are correct, its cited source belongs to the requested entity, the citation identifies the controlling revision, and the canonical cited source supports both facts. The held-out test set contains 2,133 valid records, 2,664 invalid records, and 2,242 invalid records that the base pipeline would release silently.

At validation-selected operating points under a 2 percent false-hold budget, a fitted scalar-confidence gate holds 422 of 2,664 invalid records, or 15.8%, and none of the 2,242 base-silent failures. A fitted trace gate holds 2,266 invalid records, or 85.1%, including 1,844 base-silent failures, or 82.2%. The scalar gate false-holds 0 of 2,133 valid test records. The trace gate false-holds 1 of 2,133, or 0.047%. The trace gate reduces invalidity among released stress-suite records from 51.2% under the scalar gate to 15.7%, while reducing release coverage from 91.2% to 52.7%.

An exact current-source oracle holds every invalid benchmark record and no valid record because its indicator is definitionally equivalent to the validity label. It is not a learned detector, a calibrated probability, or a Cortonex product result. The central finding is narrower and more durable: internal agreement within a transformed evidence path does not establish fidelity to authoritative current state. Identity, revision, support, transformation integrity, and release state must be measured separately.

Benchmark design, not product metrics
24,000constructed records
3operational domains
7 + cleanfault conditions plus clean state
2primary fitted policies

1The operational problem

An evidence-bearing conclusion is not produced by one operation. A source is captured, transformed into retrievable units, indexed, selected, composed into a claim, attached to a citation, and routed through a release policy. Every transition can alter the source identity, source revision, qualifier, value, or relation between the claim and its evidence.

Some failures announce themselves. A retrieval returns no complete evidence span. A source is missing. A parser cannot recover a required field. These failures can be held by a simple completeness or confidence rule.

Other failures preserve the appearance of a valid record. A stale revision can support an answer perfectly while no longer controlling the decision. A correct answer can be attached to evidence from another entity. A current citation can fail to support one emitted value. Capture corruption can alter a decisive token before retrieval begins, leaving every downstream component internally consistent with the corrupted representation.

The relevant systems question is therefore not only whether the final text is plausible or factually correct in isolation. It is whether the record remains faithful to the authoritative current source and whether the release layer has enough independent evidence to detect a broken chain.

Prior work distinguishes retrieval quality, answer quality, atomic factuality, citation quality, calibration, and selective release [2-8]. Provenance standards provide a vocabulary for entities, activities, derivation, revision, and responsible agents [1]. Recent failure taxonomies also show that retrieval-grounded systems fail at multiple pipeline stages rather than through one homogeneous error process [9,10]. This study focuses on a narrower operational problem: which failures become observable when different evidence channels are available to the release gate.

1.1Contributions and scope

We make five contributions.

First, we define a record-level validity condition that separates answer correctness, source identity, revision currentness, and canonical claim support.

Second, we construct an instrumented reference pipeline and inject faults into concrete stage artifacts rather than assigning output labels without a mechanism.

Third, we fit two primary, deliberately small release policies: one using a single scalar confidence signal, and one using a seven-signal trace. Their purpose is to expose observability, not to establish state-of-the-art classification performance. Figure 3 adds three independently fitted diagnostic ablations under the same split protocol; those ablations are analytical decompositions, not additional principal policy claims.

Fourth, we separate model fitting, probability calibration, threshold selection, and final evaluation. The 4,802-record validation pool is divided into two disjoint 2,401-record subsets. The reported test benchmark was generated only after the protocol was internally frozen. The protocol was not externally preregistered, and this publication is not peer reviewed.

Fifth, we report an exact current-source oracle only as a definition-bound ceiling. It is explicitly excluded from calibration and ranking claims.

Cortonex retains the versioned benchmark, record-level outputs, frozen generator, and separate numerical verification materials used to produce the study. Those internal materials are not part of the public release.

1.2What is assumed, constructed, fitted, measured, and derived

The study contains five epistemically different kinds of quantities.

Assumed design choices include the schema families, per-case six-document candidate bundles, fault-assignment probabilities, 85 percent material-fault probability, 2 percent validation false-hold budget, scalar-confidence formula, deterministic extraction rules, and benign signal perturbations. These are not estimates of real institutions.

Constructed observations include the 24,000 source records, their current and superseded revisions, the injected stage artifacts, emitted answers, citations, and trace signals.

Fitted quantities include logistic-regression coefficients, isotonic calibration maps, and release thresholds. These are learned from the constructed training and validation records.

Held-out measurements include the 422 of 2,664, 2,266 of 2,664, 1,844 of 2,242, release-coverage, selective-risk, calibration, and fault-level results reported on the 4,797-record test set.

Derived scenarios include the operating-prevalence table and source-recheck budget curve. They transform held-out conditional rates under additional assumptions. They are not field measurements or forecasts.


2Formal system definition

2.1Pipeline state

For record ii, let Xi(0)X_i^{(0)} denote the authoritative source state. The reference pipeline applies six transformations before release:

The transformations represent capture, segmentation, indexing, retrieval, synthesis, and citation attachment. A release gate then observes a restricted signal vector and either releases or holds the candidate record.

The pre-release candidate record is

where qiq_i is the query, y^i\widehat{\mathbf y}_i contains the two emitted facts, d^i\widehat d_i is the cited source identifier, ν^i\widehat \nu_i is the cited revision, x^icit\widehat x_i^{\mathrm{cit}} is the observed citation text, and zi\mathbf z_i is the release-policy signal vector. The gate state is not included in Ri\mathcal{R}_i^- because the gate has not yet acted.

2.2Validity

Let yi\mathbf y_i^ denote the two authoritative current facts, did_i^ the controlling source identifier, and νi\nu_i^* its revision. Define four ground-truth conditions:

and

The benchmark validity label is

The multiplication is Boolean conjunction. It does not assume statistical independence among the four conditions. In this normalized benchmark, the exact controlling-source condition makes some terms logically redundant: FiF_i implies EiE_i, and FiCiF_iC_i implies AiA_i because each controlling source contains one labelled value for each requested field. We retain the four terms to expose distinct audit properties and to keep the definition usable for less regular records. The conjunction is not claimed to be logically minimal.

This definition is intentionally strict. A factually correct answer with a stale citation is invalid. A current citation that supports only one of two emitted facts is invalid. A correct answer attached to another entity is invalid. The definition measures integrity of the complete intelligence record, not only truth of the surface answer.

2.3Base hold, fitted release, and silent failure

The reference pipeline first applies an overt base hold:

where cic_i is the scalar pipeline-confidence signal and sis_i is span completeness.

A base-silent failure is an invalid record that the base rule would release:

For fitted policy kk, let p^i(k)\widehat p_i^{(k)} be the calibrated probability that the record is valid, and let τk\tau_k be the selected threshold. Final release is

Final hold is

The fitted gate can add a hold but cannot reverse an overt base hold. Residual silent failure is

At threshold τ\tau, release coverage and selective risk are

when at least one record is released.

2.4Exact current-source oracle

The constructed benchmark contains one exact controlling source and deterministic reference facts. Define

Under the benchmark construction,

The corresponding oracle hold is hioracle=1Oih_i^{\mathrm{oracle}}=1-O_i. Equation (16) is a logical identity established by the benchmark definition. It is not an empirical detector result.

Schematic The reference evidence pipeline System schematic, not a measured figure

Interactive stage map of the six transformations, the fault class injected at each boundary, and the three nested release-gate levels.

Canonical source -> Capture -> Segmentation -> Index -> Retrieval -> Synthesis -> Citation -> Release gate. Each stage button reveals the artifact produced at that stage, the material fault injected at its boundary, why that fault can remain silent, and the evidence channel that detects it. The capture stage makes the core architectural point: the downstream trace can remain internally consistent even when the canonical source comparison fails. This schematic is conceptual and carries no measured values.

Interactive schematic of the evidence pipeline. Selecting a stage shows its artifact, material fault, silent mechanism, and detecting evidence channel. The release gate contains three nested levels: answer confidence, trace signals, and exact current-source comparison.


3Benchmark construction

Why a constructed benchmark. Silent failure is defined by the absence of any signal: an invalid output released as if valid. Measuring it therefore requires certain ground truth about which outputs are invalid, for every record. Real corpora do not carry ground-truth validity labels, so on real data silent failure can only be estimated by another imperfect system, which turns measurement into circularity. A programmatically constructed benchmark with exact reference state is the instrument that makes the quantity measurable at all. Controlled fault injection against known ground truth is the standard method of systems reliability research, and this study follows it.

3.1Domain-labelled schema families

The 24,000 records are divided equally among three schema families. The labels supply different field types and terminology; they do not reproduce real domain complexity.

Domain labelFirst fieldSecond field
Healthcare operationsReview state from READY, HELD, ESCALATED, or PENDINGCapacity index sampled from 0.45 to 0.95
Legal operationsApproval state from APPROVED, HELD, COUNSEL_REVIEW, or REJECTEDNotice window from 7, 10, 14, 21, 30, or 45 days
Supply chain operationsRelease state from RELEASED, HELD, CONDITIONAL, or ESCALATEDCommitted quantity from 80 through 1,799 units

Each record has:

  • one unique entity;
  • one controlling current revision sampled from revisions 2 through 9;
  • the immediately preceding superseded revision;
  • two atomic current facts and two stale facts;
  • an accountable owner;
  • a same-entity scenario hard negative;
  • a current review note;
  • a control policy record;
  • a cross-entity distractor;
  • a query requesting both current facts and the supporting revision.

The regular schema permits exact evaluation without a subjective grader.

Each case is evaluated only against its own six-document candidate bundle. The 24,000 bundles are not merged into one shared corpus. The experiment therefore isolates within-case source fidelity and release behavior; it does not measure corpus-scale retrieval, index construction, latency, throughput, or cross-case interference.

The reference operations are intentionally minimal. Capture copies canonical text, segmentation uses fixed token windows, indexing operates over the six local documents, retrieval uses lexical scoring, synthesis extracts two labelled fields, and citation attachment retains selected-source metadata. The value of the experiment lies in controlled observability, not in claiming realistic end-to-end model complexity.

3.2Local retrieval and scalar confidence

Sources are segmented into chunks and scored by a small lexical retrieval function with explicit currentness bonuses and stale-source penalties. Let uiu_i and viv_i be the top and runner-up extractable retrieval scores. The normalized retrieval margin is

The scalar pipeline-confidence signal is

This scalar is a designed benchmark signal. It is not a measured probability from a commercial or neural model. It incorporates retrieval separation and selected-source currentness metadata. Span completeness is a separate signal applied by the overt base hold in Equation (8).

3.3Fault assignment and materiality

Each record receives one condition:

ConditionAssignment probability
Clean execution0.34
Capture mutation0.10
Segment context mismatch0.10
Stale index pointer0.10
Retrieval substitution0.10
Synthesis substitution0.10
Citation swap0.10
Compound fault0.06

For a nonclean condition,

where Mi=1M_i=1 means the fault changes record validity. The remaining 15 percent are nonmaterial perturbations that exercise review burden without changing the current facts, source identity, controlling revision, or canonical support.

The realized benchmark contains 8,298 clean records and 15,702 nonclean records. Of the nonclean records, 2,371 are nonmaterial, or 15.10 percent. These values are realized draws from the design, not target industry frequencies.

3.4Benign signal perturbations

Perfect detector features would make review burden uninformative. The benchmark therefore includes deterministic benign irregularities on a small subset of valid records:

  • a harmless trailing newline can break byte-level source equality while preserving semantic validity;
  • a trace event can be omitted while the required evidence remains complete;
  • a registry-propagation signal can report partial freshness while the cited source is in fact current;
  • nonmaterial capture faults can rename an owner label without altering controlling facts.

These perturbations are generated from stable record-identifier hashes. They are design choices, not measured infrastructure rates. They also establish an important distinction: byte identity, observed trace completeness, and semantic validity are related but not identical.

3.5Split and post-freeze evaluation protocol

Records are stratified by domain label, fault class, and validity into:

SplitRecordsPurpose
Training14,401Fit logistic-regression coefficients
Validation pool4,802Reserved for calibration and threshold selection
Calibration subset2,401Fit isotonic calibration maps
Threshold subset2,401Select release thresholds under the false-hold budget
Final test4,797Report held-out results

The calibration and threshold subsets are disjoint. Test labels do not fit coefficients, calibration maps, or thresholds.

The method was developed against a separate development seed. The generator, validity rule, feature definitions, split logic, calibration procedure, threshold rule, metrics, and figure algorithms were then frozen. Seed 20260727 generated the reported benchmark. No method change was made after the final-seed outputs were produced. This is an internal post-freeze evaluation protocol, not an external preregistration.

3.6Realized final test set

The final test set contains:

Record stateCountShare of test set
Valid2,13344.5%
Invalid2,66455.5%
Base-silent invalid2,24246.7%
Overt invalid4228.8%

The invalid prevalence of 55.5% is intentionally high. The benchmark is a fault-rich stress suite, not an estimate of how often deployed systems fail.


4Fault operators

FaultInjected mechanismPrincipal observability
Capture mutationOne current captured value is replaced by the same entity's superseded value while the current source identifier and revision remain stableInternal trace stays consistent; byte integrity and authoritative current-source comparison fail
Segment context mismatchThe captured source remains intact, but the segmented artifact pairs current context with one superseded valueCurrent identity and revision remain correct; exact-token citation support falls to one of two fields
Stale index pointerThe current index entry is removed and the same entity's immediately preceding revision is promotedCitation can fully support the emitted answer; revision freshness fails
Retrieval substitutionA stale or scenario record is forced above the current sourceThe designed scalar confidence falls below the base threshold, making the failure overt in this benchmark
Synthesis substitutionOne emitted value is replaced after retrieval while the selected source remains currentCitation support falls to one of two fields
Citation swapThe answer remains unchanged but the citation is attached to another recordAnswer correctness can remain true while identity, currentness, or support fails
Compound faultA synthesis substitution and citation substitution occur together, with an incomplete traceMultiple signals fail
Figure 1 Stage observability under each fault class Complete constructed benchmark

Mean verifier-signal values by fault class. Toggle between the mean signal and its delta from the clean baseline.

Heat map of nine verifier signals under eight fault classes. High signal is not universally good: a high citation-support value can coexist with an invalid stale source, and capture mutation leaves every internal signal high while source integrity and current-source match collapse.

Figure 1. Mean observed signal values by fault class. The heat map uses clean and material-fault records across the full 24,000-record constructed benchmark. It is a design audit of the injected observability channels, not a held-out policy-performance estimate or a real-world prevalence estimate. The exact current-source column is the definition-bound oracle and must be visually separated from the deployable trace signals.
Figure values

Heat map summary: high signal is not universally good. Citation support remains near one under stale-index faults even though the record is invalid, and capture mutation leaves confidence, entity, revision, span, support, trace, and margin nearly unchanged while source integrity and current-source match collapse.

The capture fault is the central blind spot. Entity match, version freshness, span completeness, observed citation support, trace completeness, and retrieval margin can all look normal because every downstream operation consumes the same corrupted captured representation. Only an independent representation of authoritative source state exposes the error.


5Release policies

The principal comparison uses two fitted policies: scalar confidence and full trace. Three additional policies in Figure 3 are independently fitted diagnostic ablations used to identify what each evidence channel contributes. The exact current-source row is a deterministic ceiling, not a fitted policy.

5.1Scalar confidence gate

The first fitted policy observes only cic_i, the scalar signal in Equation (18). A balanced logistic regression maps the scalar to a validity score. Isotonic calibration is fitted on the calibration subset, and the threshold is selected on the disjoint threshold subset.

This policy is intentionally weak. It tests whether a single aggregate confidence signal can substitute for explicit provenance checks.

5.2Trace gate

The trace gate observes seven features:

where:

  • cic_i is scalar pipeline confidence;
  • eie_i is observed entity match;
  • fif_i is observed revision freshness;
  • sis_i is span completeness;
  • i\ell_i is exact contiguous-token support for the two emitted values in the observed citation text;
  • tit_i is trace completeness;
  • mim_i is retrieval margin.

The exact-token rule prevents character-substring errors such as treating 5 days as supported by 45 days, or 11 units as supported by 111 units. It remains a weak support signal because it tests value presence, not full contextual entailment or authority.

5.3Fitting, calibration, and threshold selection

For fitted policy kk,

with balanced logistic regression and inverse-regularization parameter CLR=1C_{\mathrm{LR}}=1. Isotonic calibration produces

On the threshold-selection subset, let nkvaln_k^{\mathrm{val}} be the number of valid records and sort their calibrated scores in ascending order. The integer budget is bk=0.02nkvalb_k=\lfloor0.02n_k^{\mathrm{val}}\rfloor. The implementation sets τk\tau_k to the score at zero-based index bkb_k and holds a record only when p^i(k)<τk\widehat p_i^{(k)}<\tau_k. This strict inequality does not split an isotonic score group, and the score-based component holds at most bkb_k valid records. When no valid record is already base-held, as in the final threshold subset, the resulting final-hold rate satisfies

In the final threshold subset, nkval=1,065n_k^{\mathrm{val}}=1{,}065 and bk=21b_k=21 for each primary policy. No valid threshold-subset record was already base-held, and no larger observed calibrated score satisfied the 2 percent final-hold budget.

5.4Exact current-source oracle

The third row in the results is not a fitted policy. It uses Oi=ViO_i=V_i from Equation (16). It is reported to show the logical ceiling when authoritative current state is exact and directly comparable. It receives no AUROC, average precision, Brier score, ECE, or calibration curve.


6Results

6.1Held-out operating points

Policy or ceilingInvalid records heldBase-silent failures heldValid records false-heldRelease coverageInvalidity among released recordsBrier score
Scalar confidence gate422 / 2,664 (15.8%)0 / 2,242 (0.0%)0 / 2,133 (0.0% observed)4,375 / 4,797 (91.2%)2,242 / 4,375 (51.2%)0.224
Trace gate2,266 / 2,664 (85.1%)1,844 / 2,242 (82.2%)1 / 2,133 (0.047%)2,530 / 4,797 (52.7%)398 / 2,530 (15.7%)0.069
Exact current-source oracle2,664 / 2,664 (100.0% by definition)2,242 / 2,242 (100.0% by definition)0 / 2,133 (0.0% by definition)2,133 / 4,797 (44.5%)0 / 2,133 (0.0% by definition)Not applicable

The two-sided nominal 95 percent Wilson interval for scalar invalid detection is 14.5% to 17.3%. The corresponding trace interval is 83.7% to 86.4%. For trace detection of base-silent failures, the interval is 80.6% to 83.8%.

Zero observed false holds under the scalar gate does not establish a zero underlying probability. For 0 of 2,133, the upper endpoint of the two-sided nominal 95 percent Wilson interval is 0.18%. For the trace gate's 1 of 2,133, the interval is 0.008% to 0.265%.

These are nominal finite-count summaries under an iid Bernoulli approximation. The test set itself was constructed through stratification by domain label, fault class, and validity, so the intervals are not design-based uncertainty statements for this benchmark and are not confidence intervals for real deployment rates.

Figure 2 Fault detection by release policy Held-out test set

Per-fault detection under the three policies. Switch between silent-fault detection, all invalid-output detection, and residual invalidity after the gate.

Grouped horizontal bars of detection by fault class for the three release policies. The capture-mutation row shows the trace policy at zero silent detection and the exact current-source oracle at one hundred percent. Retrieval substitution is overt in this reference implementation.

Figure 2. Detection by fault class. The default view shows base-silent-fault detection for the two fitted policies. A secondary view shows invalid detection and valid-record false holds. The exact-source oracle must be visually labelled as definition-bound, not as a third trained model.
Figure values

Bar chart summary: on capture mutation the trace policy detects zero percent of silent faults while the exact current-source oracle detects one hundred percent. Selecting a fault row opens the matching record trace in Figure 7.

6.2A scalar score does not recover provenance

The scalar gate holds all 422 overt retrieval substitutions and no base-silent failure. This is not surprising: the benchmark deliberately lowers scalar confidence when a noncontrolling retrieval is forced. Capture, segment, stale-index, synthesis, citation, and compound faults can preserve a high scalar value.

The result does not prove that every scalar confidence score will fail in the same way. It shows that a one-dimensional summary cannot identify failure properties it does not encode.

6.3The trace exposes most designed silent faults, but not capture corruption

The final trace policy holds every material test fault in the segment-context, stale-index, retrieval, synthesis, citation, and compound classes. The 422 retrieval substitutions are already held by the overt base rule; the fitted trace score adds the holds for the other listed silent-fault classes. The policy holds none of the 398 material capture mutations in the test set.

This sharp separation is a property of the constructed signals. The faults and observability channels were co-designed. It should be read as an architectural demonstration: a trace can expose faults that alter identity, revision, support, completeness, or retrieval state, but it cannot expose a corruption inherited by every downstream component unless an independent source representation is available.

The trace gate also false-holds one valid legal-schema record. The record is semantically valid, but a deterministic benign registry-lag signal sets observed freshness to 0.5. That single case is useful precisely because it prevents the trace gate from appearing infallible.

6.4Progressive evidence channels

Figure 3 Progressive evidence-channel ladder Refit and recalibrated per step

Six policies from confidence only to the exact current-source oracle. Select a step to see its feature set and operating metrics.

Ladder of independently fitted rows showing invalid and base-silent detection as diagnostic channels are added. Each row is trained, calibrated, thresholded, and evaluated separately, so operating-point behaviour need not improve monotonically. Every value is listed in the figure values table and in the results tables in the article body.

Figure 3. Detection as evidence channels are added. The scalar and full-trace rows are the two primary fitted policies. The three intermediate fitted rows are diagnostic ablations, each trained, calibrated, and thresholded independently under the same split protocol. The oracle row is a logical ceiling and must not be joined to the fitted-policy sequence as though it were another model.
Figure values

Ladder summary: every value in this figure is listed in the figure values table below it and in the results tables in the article body.

At the selected operating points:

Evidence available to the fitted policyInvalid detectionBase-silent detectionValid false holdsRelease coverageHeld-out Brier score
Scalar confidence15.8%0.0%0.000%91.2%0.224
Confidence plus identity and version55.2%46.8%0.047%69.3%0.158
Trace without citation support55.2%46.8%0.047%69.3%0.156
Full trace85.1%82.2%0.047%52.7%0.069
Trace plus source hash85.1%82.2%0.000%52.8%0.015

The source-integrity feature compares the delivered citation bytes with the canonical selected document checksum. It greatly improves ranking and Brier score and separates capture mutation in score space, but it adds no invalid detections at its independently selected operating threshold because benign byte mismatches constrain a stricter threshold. It also cannot establish current authority: an intact stale document can pass a checksum comparison while remaining noncontrolling.

6.5Risk and coverage

Figure 4 Selective risk as release coverage changes Held-out threshold grid

Threshold sweep of released-record invalidity against release coverage for each policy, with the validation-selected operating points.

Risk-coverage step curves over the disclosed threshold grid, with the irreversible base hold active at every point. Lower invalidity among released items requires a larger review population. The exact oracle is excluded because it is not a fitted threshold curve. Every value is listed in the figure values table.

Figure 4. Exact step curves for selective risk against release coverage. The fitted policy can add holds but cannot reverse the base hold. Maximum displayed coverage is therefore 91.2%, not 100 percent. Render the curves as steps through exact score groups, not as smooth interpolations. The oracle is a separate point at 44.5% coverage and zero benchmark risk by definition.
Figure values

Risk-coverage summary: every value in this figure is listed in the figure values table below it and in the results tables in the article body.

At the selected thresholds, the scalar gate releases 4,375 records, including 2,242 invalid records. The trace gate releases 2,530, including 398 invalid records. The reduction in selective risk comes with a reduction of 1,845 released records.

The scalar curve has few distinct steps because isotonic calibration maps the constructed scalar into a small number of score groups. A visually smooth curve would create information that does not exist in the data.

6.6Calibration

For fitted policy kk, the Brier score is

For ten equal-width bins B1,,B10B_1,\ldots,B_{10}, expected calibration error is

Figure 5 Held-out reliability diagrams Held-out reliability bins

Ten-bin reliability of the calibrated validity probabilities against the ideal diagonal. Marker size encodes bin count.

Reliability diagram comparing mean predicted validity with empirical validity in equal-width held-out probability bins for the two fitted policies. The oracle is excluded because it is not a calibrated probability model. Every value is listed in the figure values table.

Figure 5. Held-out reliability diagrams for the two fitted policies. The oracle is excluded. The scalar gate has Brier score 0.224 and 10-bin ECE 0.000126. The trace gate has Brier score 0.069 and ECE 0.000157.
Figure values

Reliability summary: the near-binary calibration of the independent policy is caused by exact structured ground truth and direct source comparison. It is a ceiling, not evidence of universal calibration.

The very small ECE values should not be overinterpreted. Isotonic calibration produces a few large score groups whose mean predictions closely match their group frequencies. ECE is bin-dependent and does not measure ranking or selective risk. The reliability diagrams, Brier scores, AUROC values, operating-point counts, and risk-coverage curves must be read together.

6.7Recheck coverage as a logical budget calculation

The study overlays the trace hold with exact-current-source checks applied to nested seeded subsets of the test records.

Figure 6 Source-recheck coverage and silent-fault detection Random assignment; coverage budget

Mixed policy: every record receives the trace gate while a controlled fraction also receives direct current-source recheck.

Dual-axis budget curve over increasing exact-recheck coverage. Greater idealized recheck coverage removes residual base-silent failures while preserving holds already imposed by the trace policy. This is an architectural budget ceiling, not measured verifier performance. Every value is listed in the figure values table.

Figure 6. Exact-source recheck coverage against residual stress-suite risk. At zero recheck coverage, the mixed policy equals the trace gate: 82.2% base-silent detection and 8.3% residual base-silent failures across all test records. At full recheck coverage, all invalid benchmark records are held and residual benchmark failure is zero.
Figure values

Coverage summary: every value in this figure is listed in the figure values table below it and in the results tables in the article body.

The full-coverage mixed policy still retains the trace gate's one benign false hold because the overlay only adds holds; it does not reverse an existing trace hold. It therefore releases 2,132 records, while the standalone oracle releases 2,133.

This is not a measured routing strategy. A deployable verifier would be imperfect and costly, and the choice of which records to recheck should depend on consequence, novelty, conflict, source volatility, and permission sensitivity.

6.8Conditional operating-prevalence scenarios

Let π\pi be an assumed operating invalid prevalence, dkd_k the held-out invalid-detection rate, and fkf_k the held-out false-hold rate. The projected hold rate is

and projected invalidity among released records is

when Hk(π)<1H_k(\pi)<1.

Interactive Prevalence scenario calculator Scenario, not a field estimate

Projection of hold, release, and residual rates from the measured held-out conditional rates at a reader-chosen invalid prevalence.

Scenario calculator over an assumed invalid prevalence chosen by the reader, applying the measured held-out conditional rates of each policy. Scenario projection only: the prevalence is not estimated from the constructed stress suite, and the calculation assumes those rates transfer unchanged.

Interactive prevalence scenario. Estimated hold rate, release rate, and residual invalidity for each policy at the selected prevalence, computed from the published conditional rates using equations 26 and 27.
Figure values

Prevalence scenario summary: the projection applies measured conditional rates at a reader-chosen prevalence and assumes those rates transfer unchanged, which may not hold outside the constructed benchmark.

At an assumed invalid prevalence of 5 percent:

Policy or ceilingConditional hold rateConditional release rateConditional invalidity among released records
Scalar confidence gate0.79%99.21%4.24%
Trace gate4.30%95.70%0.78%
Exact current-source oracle5.00%95.00%0.00% by definition

These values assume that the benchmark conditional detection and false-hold rates transfer unchanged into another environment. That assumption has not been tested. The table is a mathematical transformation, not a deployment forecast.

6.9Domain-label strata

The three test strata each contain 1,599 records. Trace invalid detection is 84.9 percent in the healthcare-labelled schema, 85.4 percent in the legal-labelled schema, and 84.9 percent in the supply-chain-labelled schema. The similarity is expected because the same pipeline, fault process, and feature definitions operate across all three simple schemas. It does not establish performance in healthcare, law, or supply-chain operations.


7Record-level failure traces

Figure 7 Record-level fault trace explorer Held-out examples

One held-out record per silent fault condition: source state, emitted record, verifier signals, and the decision of each gate.

Interactive trace explorer over six held-out records: capture mutation, segmentation context loss, stale index pointer, synthesis substitution, citation swap, and compound fault. Each trace shows the current source, expected and observed answers, attached citation, injected fault, verifier signals, and the release decision under each policy.

Figure 7. Interactive constructed trace explorer. The six public examples separate expected facts from emitted facts and expected source revision from attached source revision. For each fault class with a base-silent test case, the displayed record is the case returned by the frozen implementation after sorting candidates by descending scalar confidence; ties are resolved by that fixed implementation. Retrieval substitution is absent because every material retrieval substitution is overtly held. The explorer also shows the query, authoritative current source, attached citation, injected mechanism, trace probability, and gate decisions. These are constructed test records, not customer records.

Print view shows the capture mutation trace. The remaining traces are available in the online version.

Figure values

Trace explorer summary: each tab is one held-out record. The complete selected trace is readable without color; every stage carries a text state label, and changed value tokens are marked in the cited evidence text.

7.1Capture mutation

The current captured payload changes one value to the same entity's superseded value while retaining the current identifier and revision. Retrieval, answer, and citation agree with the corrupted captured payload. The trace gate releases the example because every observed trace signal remains strong. The exact-source oracle holds it.

7.2Segment context mismatch

The captured source remains correct, but a chunk artifact pairs current context with one stale value. The current citation identity and revision remain intact. Exact-token support falls to one of two emitted fields, and the trace gate holds the record.

7.3Stale index pointer

The current index entry is absent and the same entity's immediately previous revision is promoted. The answer is fully supported by that stale revision. Freshness, not citation text, exposes the failure.

7.4Synthesis substitution

Retrieval selects the controlling source, but one value changes after retrieval. Entity identity and revision stay correct. The attached source supports only one of the two emitted values.

7.5Citation swap

The two emitted answer facts remain correct while the citation is attached to another record. Answer-only evaluation would report success. Record-level validity fails because the cited evidence no longer establishes the complete claim.

7.6Compound fault

A value substitution and citation substitution occur in the same record, with a degraded trace-completeness signal. The example illustrates why a generic confidence value is insufficient for diagnosis: the reviewer needs the violated invariants and affected sources.


8Systems implications

8.1Source identity and revision are part of the conclusion

Citation text alone does not establish that the evidence belongs to the requested entity or that it is the controlling revision. Source keys and revision state should remain first-class fields through release.

8.2Verification must cross a failure boundary

A verifier that consumes the same corrupted capture can confirm the corruption. Independent verification requires another source path, immutable bytes, an authoritative registry, or another representation whose failures are not identical to those of the primary path.

8.3Support must be claim-level and context-sensitive

Exact token presence is stronger than character-substring matching but weaker than entailment. A production evaluator should test labelled field context, negation, qualifiers, scope, authority, and whether the citation supports each atomic claim.

8.4Base holds should not be reversed by a downstream score

If the primary pipeline lacks a complete evidence span or fails an explicit confidence floor, a later fitted score should not silently convert that overt failure into a release. Equations (10) and (11) enforce monotonic hold composition.

8.5Calibration is not governance

A calibrated probability can still conceal the failure mechanism. Governed release requires typed reasons, affected sources, accountable owners, and the evidence needed to resolve the hold.

8.6Review strength should follow consequence

Exact source comparison may be costly. The recheck curve provides a logical budget surface, not an optimal policy. Real allocation should be evaluated against consequence-weighted loss, source volatility, review capacity, and permission boundaries.


9Limitations

  1. Constructed records. The benchmark uses regular, deterministic schemas. It does not reproduce ambiguity, drafting variation, missing authority, multilingual content, tables, images, handwriting, temporal overlap, or source-system complexity.
  1. Domain labels are not domain evidence. Healthcare, legal, and supply-chain labels alter field vocabularies and value types only. No domain-performance claim follows from the strata.
  1. Faults and signals are co-designed. Most material faults directly alter one of the trace features. High trace detection therefore demonstrates observability under the design, not general detector performance.
  1. Per-case retrieval scope. Each query is evaluated against one independent six-document candidate bundle. The study does not test shared-corpus retrieval, approximate indexing, corpus-scale latency, throughput, or cross-case interference.
  1. No neural generation system is evaluated. Retrieval and synthesis are deterministic reference operations with injected substitutions. The study does not benchmark a commercial model or a deployed Cortonex system.
  1. Designed prevalence. Fault probabilities and materiality are experimental choices. The 55.5 percent invalid test prevalence is a stress condition, not an industry estimate.
  1. Definition-bound oracle. The exact current-source indicator equals validity by construction. Real authoritative state may be incomplete, disputed, inaccessible, multiply valid, or itself wrong.
  1. Weak observed support signal. Exact contiguous-token matching avoids substring errors but does not establish contextual entailment, authority, or absence of contradiction.
  1. Internal post-freeze protocol. The final seed was generated after an internal protocol freeze, but the study was not externally preregistered, independently replicated by another institution, or peer reviewed.
  1. Nominal intervals. Wilson intervals are finite-count summaries under an iid Bernoulli approximation. Because the test set was stratified by designed variables, they are not design-based intervals for the benchmark and do not quantify uncertainty about real-world deployment rates.
  1. No human review experiment. The study does not measure reviewer time, correction rate, disagreement, fatigue, or cost.
  1. Limited common-cause structure. Capture mutation demonstrates one common inherited failure, but the benchmark does not model correlated infrastructure incidents, adversarial sources, permission failures, or cascading organizational error.
  1. Security and privacy are outside scope. No claim is made about access control, data leakage, prompt injection, model extraction, or deployment security.

10Study specification and verification

This publication reports the benchmark design, exact validity rule, fault operators, split structure, release-policy definitions, calibration and threshold procedure, aggregate results, nominal intervals, and limitations.

Cortonex retains the frozen generator, versioned benchmark, record-level outputs, run logs, and separate verification implementation. Those internal research materials are not publicly distributed.

The final-seed study was regenerated from a frozen protocol and seed 20260727. Thirteen of the fourteen generated data outputs were reproduced byte-for-byte in a clean second run. The remaining file, study_metadata.json, matched semantically except for runtime duration. A separate arithmetic verifier recomputed count identities, fault semantics, final release composition, Wilson intervals, Brier scores, ECE, exact risk-coverage steps, reliability bins, source-recheck overlays, and prevalence transformations directly from record-level outputs. The frozen-study audit passed 316 record-level and aggregate checks. A separate publication audit then checked the article equations, displayed ratios, metadata, public graph payload, trace semantics, bibliography, formula rendering, disclosure boundary, and package checksums.

Data and code availability. This study uses a constructed reference benchmark. The experimental design, benchmark composition, fault conditions, metrics, aggregate results, and limitations are documented in this publication. Record-level benchmark data and the reference implementation are retained by Cortonex and are not publicly distributed.

Study specification and verification reference benchmark implementation, not production Cortonex software
Version1.4
Benchmark seed20260727
Benchmark size24,000 constructed records
Splits14,401 train / 4,802 validation / 4,797 test

Empirical status. A controlled fault-injection study on a programmatically constructed benchmark with exact reference state, chosen so that the validity of every record is known and silent failure can be measured rather than estimated. Every quantity reported here is a benchmark measurement, not a customer record, production log, commercial-system measurement, industry failure rate, or measured Cortonex deployment result.

The protocol was frozen before the final seed was drawn, and the published operating-point metrics were regenerated from that frozen benchmark version and cross-checked through a separate numerical verification implementation. Internal freezing is the assurance used here; the study was not externally preregistered or peer reviewed. Record-level data and the reference implementation are retained by Cortonex and are not publicly distributed.

Cite this study

The Cortonex Lab. Where Evidence Pipelines Fail: A Controlled Fault-Injection Study of Source Fidelity, Citation Lineage, and Governed Release. Version 1.4. Cortonex Technologies Inc. https://cortonex.com/lab/where-evidence-pipelines-fail/

@techreport{cortonexlab-where-evidence-pipelines-fail,
  author      = {{The Cortonex Lab}},
  title       = {Where Evidence Pipelines Fail: A Controlled Fault-Injection
                 Study of Source Fidelity, Citation Lineage, and Governed Release},
  institution = {The Cortonex Lab, Cortonex Technologies Inc.},
  version     = {1.4},
  url         = {https://cortonex.com/lab/where-evidence-pipelines-fail/},
  note        = {Controlled constructed benchmark;
                 no production-performance claim.}
}

References

  1. Lebo, T., Sahoo, S., and McGuinness, D. (eds.). PROV-O: The PROV Ontology. W3C Recommendation, April 30, 2013.
  2. Petroni, F., Piktus, A., Fan, A., et al. KILT: a Benchmark for Knowledge Intensive Language Tasks. NAACL-HLT, 2021.
  3. Gao, T., Yen, H., Yu, J., and Chen, D. Enabling Large Language Models to Generate Text with Citations. EMNLP, 2023.
  4. Min, S., Krishna, K., Lyu, X., et al. FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text Generation. EMNLP, 2023.
  5. Es, S., James, J., Espinosa Anke, L., and Schockaert, S. RAGAs: Automated Evaluation of Retrieval Augmented Generation. EACL System Demonstrations, 2024.
  6. Saad-Falcon, J., Khattab, O., Potts, C., and Zaharia, M. ARES: An Automated Evaluation Framework for Retrieval-Augmented Generation Systems. NAACL-HLT, 2024.
  7. Guo, C., Pleiss, G., Sun, Y., and Weinberger, K. Q. On Calibration of Modern Neural Networks. ICML, 2017.
  8. Geifman, Y. and El-Yaniv, R. Selective Classification for Deep Neural Networks. NeurIPS, 2017.
  9. Garani, A. A Systematic Taxonomy of Failure Modes in Retrieval-Augmented Generation Systems. TrustNLP, 2026.
  10. Leung, K. K., Belbahri, M., Sui, Y., et al. Classifying and Addressing the Diversity of Errors in Retrieval-Augmented Generation Systems. EACL, 2026.

Appendix A. Exact held-out count identities

The final test-set identities are

For the scalar gate,

For the trace gate,

Appendix B. Wilson score interval

For xx observed events in nn trials, with p^=x/n\widehat p=x/n and z=1.959963984540054z=1.959963984540054, the two-sided nominal 95 percent Wilson interval is

For x=0x=0 and n=2,133n=2{,}133, the upper endpoint is 0.0017977276, or 0.18%. For x=1x=1 and n=2,133n=2{,}133, the interval is 0.0000827636 to 0.0026509248.

Appendix C. Publication interpretation boundary

The article supports these statements:

  • internally consistent evidence paths can remain wrong relative to authoritative source state;
  • explicit identity, revision, support, and trace signals expose different designed fault classes;
  • a downstream fitted gate should not reverse an overt base hold;
  • exact source comparison is a logical ceiling in this benchmark;
  • held-out stress-suite measurements and prevalence-transfer scenarios are different kinds of quantities.

The article does not support these statements:

  • Cortonex detects 85.1 percent of real customer evidence failures;
  • enterprise evidence pipelines have a 55.5 percent invalid-output rate;
  • exact current-source verification is perfect in deployment;
  • the three domain-labelled strata establish healthcare, legal, or supply-chain performance;
  • the reported rates generalize without a field evaluation.