correction governance results

When analogue correction should be withheld

The ARC evidence layer is not presented as a universal accuracy booster. It is evaluated as deploy-time governance: measuring correction headroom, testing when correction should abstain, and exposing why reliable action remains limited under year and spatial shift.

Reading guide

headroom here, governance details in dedicated pages

This page now follows the manuscript argument: analogue residuals have real headroom, rolling-history deployment is harder, and the deployable contribution is asymmetric governance rather than automatic correction.

Forced per-field routing is summarized in Router audit. Selective abstention, selector validation, and sanity controls are summarized in Controls. Spatial residual structure and paired evidence cards are summarized in Reliability cases.

Scope note. This public demo covers ARC's agronomic representation. The AEF representation audit and Track-H / Track-P information contracts remain manuscript analyses and are intentionally outside this release.

Asymmetric governance

main result in one figure
Asymmetry between knowing when to correct and knowing when to abstain
Deployment value is asymmetric. The evidence is weak for choosing active correction policies, but stronger for identifying when correction should be withheld. This is the site-wide reading frame for the ARC layer.

Background

YieldSAT dataset coverage map and table
Dataset coverage. Miranda et al. (2026). YieldSAT field coverage across countries, crops, years, and Sentinel-2 images.
YieldSAT data modalities for crop yield modeling
Data modalities. Miranda et al. (2026). Satellite, weather, soil, and terrain inputs used for field-level yield modeling.
Najjar-style multimodal yield prediction pipeline
Base multimodal model. Najjar et al. (2025). Transformer-based multimodal crop-yield model used as the base encoder.

Calibration vs analogue correction

Calibration

global shift
Same correction for all fields.
+β_tField A
+β_tField B
+β_tField C
+β_tQuery
Uses alllabeled target calibration fields estimate one shared bias shift.

Analogue correction

local analogue vote
Query-specific analogue correction.
w .52Top-1
w .31Top-2
w .17Top-3
local ΔQuery
Uses Nkanalogues are ranked in the agronomic feature space used by ARC; closer fields get larger weights.

Why scenes differ

ARC retrieval geometry across LOYO and LORO domain structure in the 349-feature agronomic representation
LOYO/LORO structure. Independent projections show heterogeneous organization under temporal and regional shift. This is a descriptive view of the explicit 349-feature agronomic representation, not a causal explanation of correction performance.

Primary result: label efficiency

matched target-label budgets · F1.4
Label efficiency of ARC against frozen base, scalar calibration, and matched-budget retraining across crop and protocol settings
ARC is the primary label-efficiency result. Across 30 crop × protocol × sample-ratio settings, ARC has lower mean held-out-domain RMSE than frozen base in 29/30, scalar calibration in 30/30, and matched-budget retraining in 22/30. Paired bootstrap favors ARC in 25/30, 26/30, and 14/30 settings respectively; the complete paired summary is shown below.

Current showcase: one Contract-A query

wheat · LORO · SR 0.2 · r0
Observed yield and frozen, calibrated, and ARC prediction maps for the current wheat LORO showcase field
Current F1.4 showcase. Germany Wheat/LORO field `wheat_loro_sr0.2_r0_b9bc2db689b4`: ARC reduces field-level absolute error relative to scalar calibration in this illustrative Contract-A case. Open the full retrieval walkthrough.

Rolling-history deployment

real similarity signal, year-drift limit
Rolling-history ARC signal compared with randomized null controls
Rolling-history sanity control. Real rolling-history correction beats randomized nulls in 19 of 30 settings, but aggregate all-ARC correction is still harmful. Similarity signal exists; deploying it as an always-correct policy does not follow.

Appendix baseline comparison

target-calibration reference, not the main governance claim
Label-efficiency of ARC analogue correction against target-calibration training
Analogue correction versus target calibration. This figure is retained as a baseline comparison. It is not evidence that ARC should be applied blindly; the governance pages test that separate deployment question.

Decision workflow

leakage-safe deploy-time evidence
Base model
Start from the Najjar-style multimodal Transformer prediction.
Target calibration
Estimate target-domain bias from a small labeled calibration bank.
Analogue retrieval
Find similar fields in target-domain agro-feature space.
Evidence scores
Score support, distance, anomaly strength, view consistency, and historical risk without target ground truth.
Apply or abstain
Use correction only when the evidence passes the selected policy; otherwise keep calibration-only.
Disclose limits
Report when evidence cannot distinguish helpful from harmful correction, rather than hiding uncertainty.

Paired held-out-domain summary

current primary result · 30 settings collapsed by crop × protocol

Each row aggregates five sample ratios (0.1–0.5). “Favoured” counts paired bootstrap comparisons in which ARC is preferred; “not separated” is the remainder after ARC-favoured and matched-budget-retraining-favoured settings.

cropprotocolsettingsARC bootstrap favoured vs frozen baseARC bootstrap favoured vs scalar calibrationARC bootstrap favoured vs retrainingretraining favoured vs ARCnot separated
CornLORO511023
CornLOYO555302
SoybeanLORO545203
SoybeanLOYO555203
WheatLORO555302
WheatLOYO555401