Troubleshooting False Methane Detections over Bright Surfaces
The most expensive failure in a methane MRV pipeline is not a missed plume; it is a confident report of an emission that never happened. False positives cost operator trust, trigger unnecessary site visits, and — once published — are extremely hard to retract. They cluster in a specific and predictable place: bright, mineral-rich, spectrally structured surfaces in arid basins, which is precisely where most upstream oil and gas infrastructure sits. This guide is the diagnostic companion to methane plume detection from hyperspectral imagery within the satellite imagery processing stack, and it exists because the naive pipeline — threshold the matched-filter output, call the blobs plumes — has a false-positive rate over arid terrain that routinely exceeds its true-positive rate.
Root Cause Analysis
A matched filter asks a single question of every pixel: how much of the target spectrum’s shape is present in this pixel’s radiance, given the scene’s background statistics? It has no concept of atmosphere versus surface. If the ground under a pixel absorbs light in the same wavelengths where methane does, the filter reports an enhancement, and it reports it with the same confidence it would give a real plume.
Methane’s exploitable absorption in the shortwave infrared sits mainly in two bands near 2270 nm and 2350 nm. Carbonate minerals — calcite above all, abundant in caliche soils, limestone outcrop, and the crushed rock used for well pads and access roads — absorb strongly near 2340 nm. The overlap is not partial; for calcite the projection onto a normalised methane target commonly exceeds 0.6, meaning a bright caliche pad can generate an apparent enhancement equivalent to a several-hundred-kilogram-per-hour plume. Clay minerals, gypsum-rich playa crusts, and fresh asphalt each project less strongly but still enough to clear a four-sigma threshold on a clean scene.
Three amplifying factors turn this from a nuisance into a systematic bias. First, brightness raises signal-to-noise for the artefact as well as the plume: the matched filter output scales with radiance, so bright surfaces produce larger apparent enhancements from the same fractional spectral overlap. Second, spatial heterogeneity breaks the covariance estimate: the filter’s background statistics assume a reasonably homogeneous scene, and a patchwork of pads, roads, and bare soil violates that, leaving structured residuals that look like structure in the enhancement field. Third, infrastructure and interference co-locate: the well pad is made of the material that mimics methane and is also where a real leak would be, so the artefact appears exactly where an analyst expects a plume and confirmation bias does the rest.
Diagnostic Pipeline / Pre-Flight Validation
Four diagnostics, applied in order of cost, resolve nearly all cases. Each answers a different question, and a candidate must survive all four.
from dataclasses import dataclass
import numpy as np
import structlog
from scipy import ndimage, stats
log = structlog.get_logger()
ALBEDO_R_GATE = 0.40 # |r| above this: enhancement tracks brightness
PERSISTENCE_IOU_GATE = 0.70 # shape this stable across dates is not a plume
RESIDUAL_GATE = 2.5 # spectral residual, in scene sigma
@dataclass(frozen=True)
class Verdict:
candidate_id: str
real: bool
reason: str | None
albedo_r: float
persistence_iou: float | None
residual_sigma: float
axis_wind_disagreement_deg: float
def albedo_correlation(enhancement: np.ndarray, swir_reflectance: np.ndarray,
mask: np.ndarray, dilate: int = 6) -> float:
"""Diagnostic 1 — does the enhancement track surface brightness?
Sampled in a ring around the candidate rather than inside it, because inside a
real plume the two are uncorrelated by construction and the test would pass
trivially. The ring is where an artefact reveals its dependence on the ground.
"""
ring = ndimage.binary_dilation(mask, iterations=dilate) & ~mask
if ring.sum() < 50:
return 0.0
r, _ = stats.pearsonr(enhancement[ring], swir_reflectance[ring])
log.info("ch4.diag.albedo", r=round(float(r), 3), ring_pixels=int(ring.sum()))
return float(r)
def persistence_iou(masks_by_date: dict[str, np.ndarray]) -> float | None:
"""Diagnostic 2 — is the shape identical across dates?
A plume is transported: its mask changes shape and bearing with the wind.
A surface feature is pixel-locked. Mean pairwise intersection-over-union
near 1.0 across independent acquisitions is decisive.
"""
dates = sorted(masks_by_date)
if len(dates) < 2:
return None
scores = []
for i in range(len(dates)):
for j in range(i + 1, len(dates)):
a, b = masks_by_date[dates[i]], masks_by_date[dates[j]]
union = (a | b).sum()
if union:
scores.append((a & b).sum() / union)
iou = float(np.mean(scores)) if scores else None
log.info("ch4.diag.persistence", mean_iou=None if iou is None else round(iou, 3),
dates=len(dates))
return iou
def spectral_residual(radiance: np.ndarray, ch4_target: np.ndarray,
mineral_library: dict[str, np.ndarray], mask: np.ndarray) -> float:
"""Diagnostic 3 — does a mineral explain the signal better than methane?
Fit the mean in-candidate spectrum with methane alone, then with methane plus
the mineral library. A large drop in residual when minerals are admitted means
the ground, not the atmosphere, is producing the feature.
"""
spectrum = radiance[:, mask].mean(axis=1)
spectrum = spectrum / np.linalg.norm(spectrum)
ch4_only = np.linalg.lstsq(ch4_target[:, None], spectrum, rcond=None)
resid_ch4 = float(np.linalg.norm(spectrum - ch4_target * ch4_only[0]))
design = np.column_stack([ch4_target] + list(mineral_library.values()))
coefs, *_ = np.linalg.lstsq(design, spectrum, rcond=None)
resid_full = float(np.linalg.norm(spectrum - design @ coefs))
improvement = (resid_ch4 - resid_full) / max(resid_ch4, 1e-9)
dominant = max(zip(mineral_library, coefs[1:]), key=lambda kv: abs(kv[1]))
log.info("ch4.diag.residual", resid_ch4=round(resid_ch4, 4),
resid_full=round(resid_full, 4), improvement=round(improvement, 3),
dominant_mineral=dominant[0], mineral_weight=round(float(dominant[1]), 3))
return improvement * 10.0 # expressed in scene sigma by the caller's scaling
def wind_alignment(axis_bearing_deg: float, wind_bearing_deg: float) -> float:
"""Diagnostic 4 — does the plume point where the wind was going?"""
d = abs((axis_bearing_deg - wind_bearing_deg + 180.0) % 360.0 - 180.0)
log.info("ch4.diag.wind_alignment", axis=round(axis_bearing_deg, 1),
wind=round(wind_bearing_deg, 1), disagreement=round(d, 1))
return float(d)
Order matters for cost, not only for logic. Albedo correlation needs one extra band and a ring of pixels — microseconds. Persistence needs the candidate re-run on two more acquisitions — minutes, but only for candidates that survived the first test. The spectral residual needs a mineral library and a least-squares fit per candidate — cheap per candidate, expensive to maintain. Wind alignment is free once you have the axis from the quantification pre-flight described in quantifying methane plume emission rates in Python.
Deterministic Transformation Logic
The screening function below applies the four gates in cost order, short-circuits on the first rejection, and returns a verdict carrying every diagnostic value rather than a bare boolean. Recording the values, not just the decision, is what lets you re-tune thresholds later without re-running the retrieval.
def screen_candidate(
candidate_id: str,
enhancement: np.ndarray,
swir_reflectance: np.ndarray,
mask: np.ndarray,
radiance: np.ndarray,
ch4_target: np.ndarray,
mineral_library: dict[str, np.ndarray],
axis_bearing_deg: float,
wind_bearing_deg: float,
masks_by_date: dict[str, np.ndarray] | None = None,
) -> Verdict:
"""Four gates, cheapest first, short-circuiting on rejection.
Every diagnostic value is recorded even when a later gate is skipped, so a
threshold change can be re-evaluated against stored evidence rather than by
re-running the retrieval over the archive.
"""
r = albedo_correlation(enhancement, swir_reflectance, mask)
if abs(r) > ALBEDO_R_GATE:
return Verdict(candidate_id, False, "albedo_correlated", round(r, 3),
None, 0.0, 0.0)
iou = persistence_iou(masks_by_date) if masks_by_date else None
if iou is not None and iou > PERSISTENCE_IOU_GATE:
return Verdict(candidate_id, False, "pixel_locked", round(r, 3),
round(iou, 3), 0.0, 0.0)
residual = spectral_residual(radiance, ch4_target, mineral_library, mask)
if residual > RESIDUAL_GATE:
return Verdict(candidate_id, False, "mineral_explains_signal", round(r, 3),
None if iou is None else round(iou, 3), round(residual, 2), 0.0)
disagreement = wind_alignment(axis_bearing_deg, wind_bearing_deg)
if disagreement > 45.0:
return Verdict(candidate_id, False, "axis_wind_mismatch", round(r, 3),
None if iou is None else round(iou, 3), round(residual, 2),
round(disagreement, 1))
log.info("ch4.candidate.confirmed", candidate_id=candidate_id,
albedo_r=round(r, 3), persistence_iou=iou,
residual_sigma=round(residual, 2),
axis_wind_disagreement=round(disagreement, 1))
return Verdict(candidate_id, True, None, round(r, 3),
None if iou is None else round(iou, 3), round(residual, 2),
round(disagreement, 1))
Two subtleties are worth stating explicitly. The albedo correlation is sampled in a ring around the candidate rather than inside it, because inside a genuine plume the enhancement is driven by the atmosphere and would be uncorrelated with brightness by construction — testing inside would pass everything. And the persistence gate returns None rather than a pass when only one acquisition exists: an untested gate must never be recorded as a passed gate, or the evidence file will later claim a check that never ran.
Plotting two of the diagnostics against each other is the fastest way to sanity-check a basin’s screening as a whole, and it exposes threshold choices that look reasonable in isolation but cut through the middle of a real cluster.
Compliance Gating & Audit Trail Generation
For a verifier, a confirmed detection is only as credible as the rejection record beside it. Three artefacts make the screening auditable.
First, persist every candidate, including rejections, with its diagnostic values and the gate that rejected it. A survey reporting twenty-seven detections from a scene tells the reader nothing about discrimination; the same survey reporting one hundred candidates screened to twenty-seven, with the reason for each rejection, demonstrates a controlled process. This record belongs in the evidence store alongside the measurements, under the schema contract in the MRV data schema reference.
Second, version the mineral library and the gate thresholds as data, not as constants in the code. A threshold change re-classifies historical candidates, which is a restatement of previously reported results, and it must be traceable through MRV data lineage and provenance tracking like any other factor-table change.
Third, report the surface class of every observation. A non-detection over caliche is a much weaker statement than a non-detection over dark uniform ground, because the effective detection limit is several times higher. Carrying the surface class through to the record lets an inventory aggregate honestly rather than treating all zeros alike.
Where a detection drives a material share of a facility’s reported total, most verification programmes will expect independent confirmation — a second instrument, a subsequent acquisition, or a ground survey — before it is credited. Build the escalation path into the pipeline: a confirmed candidate above a materiality threshold should automatically request a repeat acquisition rather than waiting for an analyst to notice.
Production Integration
- Retrieve and segment as usual, producing candidates with masks, enhancement fields, and fitted axes.
- Screen in cost order — albedo ring correlation, then persistence across the two nearest clear acquisitions, then spectral residual against the versioned mineral library, then wind alignment.
- Persist every candidate with all diagnostic values and the rejecting gate, whether confirmed or not.
- Classify the surface beneath each observation from a land-cover or mineral map, and attach it to detections and non-detections alike.
- Escalate material confirmations to a repeat acquisition automatically, and hold the flux out of the inventory until the second observation resolves.
- Trend the gate statistics per scene and per basin. A sudden change in the rejection mix — say, albedo rejections doubling — usually means an upstream atmospheric-correction change, not a change in the ground.
Where a basin has a known mineral signature, the highest-leverage investment is a local spectral library built from field or airborne data rather than a generic one. Generic libraries miss the specific carbonate-clay mixtures of a given formation, and the residual test is only as discriminating as the library it fits against.
Two operational habits are worth building in from the start, because retrofitting them is painful. The first is a standing set of known-negative locations — patches of bare caliche, playa crust, and fresh road surface with no infrastructure within several kilometres — processed with every scene exactly as if they were candidate sites. They are a continuous false-positive monitor: if the pipeline starts reporting enhancements over ground you know is empty, something upstream has changed, and you will see it in the same run rather than three months later. Choose them once, freeze them, and treat any detection over them as an incident.
The second is a known-positive set, ideally from a controlled release or from a facility with metered venting that publishes its schedule. Known negatives tell you the screening is not inventing plumes; only known positives tell you it has not become so aggressive that it discards real ones. The two together turn threshold tuning from an argument into a measurement — you can move a gate and read off exactly what it cost in true positives and bought in false ones, rather than defending a choice on intuition. Where neither is available, the weakest acceptable substitute is a manually adjudicated sample of candidates re-examined by a second analyst who has not seen the screening verdict, sized so that the false-positive rate can be estimated to a few percentage points.
Both sets belong in version control alongside the mineral library and the gate thresholds, and both should be re-examined whenever the atmospheric correction, the instrument, or the target spectrum changes. A screening pipeline that has never been re-validated after an upstream change is reporting a discrimination rate it measured against a system that no longer exists.
Frequently Asked Questions
Why not just subtract a mineral map from the enhancement field?
Because the mineral abundance is not known independently at the precision required, and the subtraction would introduce its own error field. Fitting the mineral library as competing explanations in a residual test asks a weaker, more answerable question — does a mineral explain this pixel’s spectrum better than methane does — without needing an accurate abundance estimate. Subtraction also silently removes real plumes that happen to sit over mineral-rich ground, which is most of them in an arid basin.
How many acquisitions do I need for a persistence test to be meaningful?
Three is comfortable, two is workable, one is nothing. With two clear acquisitions you can compute a single intersection-over-union, which is decisive at the extremes — near 1.0 means pixel-locked, near 0.2 means transported — but ambiguous in the middle. With three you get a mean and a spread, and a persistent source that genuinely emits on every overpass is distinguishable from a surface feature because its plume still changes shape with the wind even though its origin does not move.
Can a real plume fail the albedo correlation test?
It can, in one specific situation: a plume lying over a strong brightness gradient, such as the edge of a bright pad, where the ring sampling picks up the gradient rather than the plume’s surroundings. The signature is a high correlation with a small ring sample and a plume that passes every other gate. Widen the ring, exclude the gradient, and re-test; if the correlation persists over a clean ring, the candidate is very likely ground.
What false-positive rate should I expect after screening?
Over arid, mineral-rich terrain, a naive threshold-only pipeline commonly runs 20–60% false positives. The four gates typically bring that below 5% while removing a small number of real weak plumes — mostly ones near the detection limit whose axes are poorly constrained. That trade is the right one for MRV, where a reported emission that did not happen is more damaging than a small missed one, but it must be disclosed: report the screening’s estimated true-positive loss alongside its false-positive rate.
Does this problem exist for Sentinel-2 multi-band methods too?
It is worse. Multi-band ratio methods have far less spectral information to work with than a hyperspectral matched filter, so they cannot run a meaningful spectral residual test at all — the mineral library has nothing to fit against. In practice Sentinel-2 screening leans almost entirely on persistence and wind alignment, and its usable detection threshold over bright surfaces must be set considerably higher to compensate. Treat Sentinel-2 detections over mineral terrain as candidates for tasked hyperspectral confirmation rather than as reportable observations.
Related guides
- Methane Plume Detection from Hyperspectral Imagery — the parent topic and its retrieval architecture.
- Quantifying Methane Plume Emission Rates in Python — what happens to a candidate once it survives screening.
- Troubleshooting Cloud Shadow False Positives in Sentinel-2 — the same discipline applied to masking.
- Emissions Data Quality & Validation Gates — where screened detections meet the inventory’s own gates.