GeoParquet and STAC Metadata Conventions for MRV
A carbon MRV dataset that outlives its authors is a metadata problem before it is a storage problem. The pixels and the rows are easy to keep; what disappears is the knowledge of which projection they are in, what version of which emission factor table produced them, whether a null means unobserved or zero, and which of three similarly named files was the one that fed the submitted report. This guide sets out a working convention for recording that, within the MRV data schema reference in the pipeline orchestration and compliance reference stack.
The division of labour between the two formats is the thing to get right first. GeoParquet metadata travels inside the file and answers questions about the data itself — geometry column, CRS, units, null semantics. STAC metadata lives outside the file and answers questions about the file’s place in the world — what period it covers, what produced it, what it supersedes, where its siblings are. Trying to make either do the other’s job produces either files nobody can search or catalogues that go stale the moment a file is copied.
Root Cause Analysis
Three recurring defects account for most of the metadata problems in carbon datasets, and each has a specific structural cause.
Units live in column names or nowhere. A column called emissions in a table produced by three different teams has been kilograms, tonnes, and tonnes of CO₂ equivalent in the same organisation. Encoding the unit in the name — emissions_tco2e — is a substantial improvement and still fails when a unit changes and the name does not, or when a consumer parses the name for something else. Recording the unit as structured column metadata makes it machine-checkable, which is what allows a validation gate to reject a mismatch rather than a human to notice one.
Null carries three different meanings in one column. Not measured, measured as zero, and not applicable are semantically distinct and are all stored as null by default. In carbon accounting the distinction is material: a parcel with null emissions because it was not surveyed is not a parcel with zero emissions, and summing a column that conflates them understates the total by exactly the unsurveyed portion. The convention that fixes this is to record the null semantics per column and, where more than one meaning is genuinely needed, to add an explicit status column rather than overloading null.
Catalogue and file drift apart. A STAC item is a copy of facts about a file, and copies go stale. A file reprojected, a column added, a period corrected — each changes the file without changing the catalogue, and the catalogue then describes something that no longer exists. This is why the small overlap set matters: CRS, temporal extent, and checksum are cheap to verify programmatically, and a periodic reconciliation over those three catches nearly every drift.
The pattern behind all three is that metadata which is not mechanically checked is metadata that is eventually wrong. The convention below is designed around what a validator can assert, not around what is nice to document.
Diagnostic Pipeline / Pre-Flight Validation
Validate metadata at write time rather than at read time. A file written without its required metadata is cheap to fix in the minute after it was produced and expensive to fix once it has been copied into three downstream systems.
import json
from dataclasses import dataclass, field
import pyarrow.parquet as pq
import structlog
log = structlog.get_logger()
REQUIRED_FILE_KEYS = frozenset({
"mrv:schema_version",
"mrv:producer_run_id",
"mrv:emission_factor_version",
"mrv:methodology",
"mrv:period_start",
"mrv:period_end",
})
NULL_SEMANTICS = frozenset({"not_measured", "not_applicable", "no_nulls"})
@dataclass(frozen=True)
class ColumnConvention:
"""The metadata every measure column must carry.
`null_means` is mandatory and has no default. Requiring the author to
state it is the whole point — a default of 'not_measured' would be
guessed correctly most of the time and wrongly exactly where it matters.
"""
name: str
unit: str
null_means: str
valid_min: float | None = None
valid_max: float | None = None
definition: str = ""
def to_metadata(self) -> dict[bytes, bytes]:
payload = {
"unit": self.unit,
"null_means": self.null_means,
"valid_min": self.valid_min,
"valid_max": self.valid_max,
"definition": self.definition,
}
return {b"mrv:column": json.dumps(payload, sort_keys=True).encode()}
class MetadataError(ValueError):
"""Raised when a file cannot be published as it stands."""
def validate_geoparquet(path: str, conventions: dict[str, ColumnConvention]) -> None:
"""Assert file- and column-level metadata before publication."""
meta = pq.read_metadata(path)
kv = {k.decode(): v.decode() for k, v in (meta.metadata or {}).items()}
missing = REQUIRED_FILE_KEYS - kv.keys()
if missing:
raise MetadataError(
f"{path}: missing required file metadata {sorted(missing)}"
)
if "geo" not in kv:
raise MetadataError(
f"{path}: no 'geo' key — this is not a valid GeoParquet file, "
"whatever its extension says"
)
geo = json.loads(kv["geo"])
primary = geo.get("primary_column")
col_meta = geo.get("columns", {}).get(primary, {})
if not col_meta.get("crs"):
raise MetadataError(
f"{path}: geometry column '{primary}' has no CRS. A missing CRS "
"is not an assertion of WGS 84 — it is an assertion of nothing, "
"and every consumer will guess differently"
)
schema = pq.read_schema(path)
for name, conv in conventions.items():
if name not in schema.names:
raise MetadataError(f"{path}: expected column '{name}' is absent")
field_meta = schema.field(name).metadata or {}
raw = field_meta.get(b"mrv:column")
if raw is None:
raise MetadataError(
f"{path}: column '{name}' carries no mrv:column metadata; "
f"a bare column is a column whose unit is folklore"
)
declared = json.loads(raw.decode())
if declared["unit"] != conv.unit:
raise MetadataError(
f"{path}: column '{name}' declares unit '{declared['unit']}' "
f"but the convention requires '{conv.unit}'"
)
if declared["null_means"] not in NULL_SEMANTICS:
raise MetadataError(
f"{path}: column '{name}' declares null_means "
f"'{declared['null_means']}', not one of {sorted(NULL_SEMANTICS)}"
)
log.info(
"geoparquet.validated", path=path,
columns=len(conventions), crs=col_meta["crs"].get("id", {}),
schema_version=kv["mrv:schema_version"],
)
The refusal to treat a missing CRS as WGS 84 is the assertion that saves the most trouble. Defaulting is convenient exactly once and then produces a dataset whose coordinates are silently in a local grid, which the consuming pipeline dutifully plots in the Gulf of Guinea.
Deterministic Transformation Logic
The STAC side of the convention is a small set of required fields plus two extensions worth adopting wholesale. Building the item from the file rather than alongside it is what keeps the two consistent.
import hashlib
from dataclasses import dataclass
from datetime import date, datetime, timezone
from pathlib import Path
@dataclass(frozen=True)
class MrvStacItem:
"""A STAC item for one MRV output, with the fields a verifier looks for.
The mrv: namespace fields are the project-specific additions. Everything
else is standard STAC plus the processing and version extensions, which
already model exactly what an audit needs — who produced it, with what
software, and what it replaces.
"""
item_id: str
collection: str
geometry: dict
bbox: tuple[float, float, float, float]
datetime_utc: datetime | None
start_datetime: datetime
end_datetime: datetime
asset_href: str
asset_checksum: str
asset_bytes: int
crs_code: str
processing_run_id: str
processing_software: dict[str, str]
methodology: str
emission_factor_version: str
schema_version: str
superseded_item: str | None
deprecated: bool = False
def to_dict(self) -> dict:
return {
"type": "Feature",
"stac_version": "1.0.0",
"stac_extensions": [
"https://stac-extensions.github.io/processing/v1.1.0/schema.json",
"https://stac-extensions.github.io/version/v1.2.0/schema.json",
"https://stac-extensions.github.io/file/v2.1.0/schema.json",
"https://stac-extensions.github.io/projection/v1.1.0/schema.json",
],
"id": self.item_id,
"collection": self.collection,
"geometry": self.geometry,
"bbox": list(self.bbox),
"properties": {
# A period product has no single instant. Null datetime with
# a start/end pair is the correct STAC encoding, and picking
# an arbitrary midpoint — the common shortcut — makes every
# temporal search subtly wrong.
"datetime": None,
"start_datetime": self.start_datetime.isoformat(),
"end_datetime": self.end_datetime.isoformat(),
"proj:code": self.crs_code,
"processing:lineage": (
f"run {self.processing_run_id} under {self.methodology}"
),
"processing:software": self.processing_software,
"version": self.schema_version,
"deprecated": self.deprecated,
"mrv:methodology": self.methodology,
"mrv:emission_factor_version": self.emission_factor_version,
"mrv:schema_version": self.schema_version,
"mrv:run_id": self.processing_run_id,
},
"assets": {
"data": {
"href": self.asset_href,
"type": "application/vnd.apache.parquet",
"roles": ["data"],
"file:checksum": self.asset_checksum,
"file:size": self.asset_bytes,
}
},
"links": (
[{"rel": "predecessor-version", "href": self.superseded_item}]
if self.superseded_item else []
),
}
def reconcile(item: MrvStacItem, parquet_path: Path) -> list[str]:
"""Check the catalogue against the file it describes.
Runs periodically over the whole catalogue. Three comparisons catch
almost every drift, and all three are cheap: checksum, CRS, and period.
"""
problems: list[str] = []
digest = hashlib.sha256()
with parquet_path.open("rb") as fh:
for block in iter(lambda: fh.read(1 << 20), b""):
digest.update(block)
actual = f"sha256:{digest.hexdigest()}"
if actual != item.asset_checksum:
problems.append(
f"checksum drift: catalogue {item.asset_checksum[:20]}… "
f"vs file {actual[:20]}…"
)
kv = {k.decode(): v.decode()
for k, v in (pq.read_metadata(parquet_path).metadata or {}).items()}
if kv.get("mrv:period_start") != item.start_datetime.date().isoformat():
problems.append(
f"period drift: catalogue {item.start_datetime.date()} "
f"vs file {kv.get('mrv:period_start')}"
)
if kv.get("mrv:schema_version") != item.schema_version:
problems.append(
f"schema version drift: catalogue {item.schema_version} "
f"vs file {kv.get('mrv:schema_version')}"
)
return problems
The datetime: None with a start and end pair is a small point that matters disproportionately. Almost every MRV output covers an interval rather than an instant, and setting datetime to the interval’s midpoint — which many tools do by default — makes a search for items covering a given day return the wrong set, in a way that is very hard to notice.
Compliance Gating & Audit Trail Generation
The gate is straightforward: no file publishes without passing metadata validation, and no catalogue entry publishes without reconciling against its file. Beyond that, three things belong in the record.
The schema version alongside every output, in both layers. A dataset spanning a schema change is normal; one where the version is inferable only from the column list is not, and inference goes wrong at exactly the transition.
The supersession link rather than an overwrite. STAC’s version extension already models this — deprecated on the old item and a predecessor-version link on the new one — and using it means a verifier who has a reference to an older item finds it, marked, rather than finding nothing.
Reconciliation results with their date. A catalogue that has been reconciled monthly and has a clean record is materially more credible than one asserted to be consistent, and the record costs a log line.
Production Integration
Write the metadata in the same function that writes the file, and build the STAC item from the file’s own metadata rather than from the variables that produced both. This sounds like a small distinction and it is the one that keeps the two layers in agreement: deriving the item from the file means an error in the file surfaces as an obviously wrong item, whereas deriving both from the same in-memory variables means an error appears consistently in both and looks correct.
The column conventions themselves should live in one place shared by every producer, not repeated per pipeline. That shared definition is the natural companion to the canonical Parquet schema data dictionary for MRV, and the emission factor version referenced in the file metadata should be a key into the versioned table described in versioning emission factor databases for reproducible MRV rather than a free-text string.
Frequently Asked Questions
Which STAC extensions are worth adopting for MRV?
Four carry their weight. Projection supplies the CRS and transform fields, which every raster consumer needs. Processing records lineage and software versions, which is what an audit asks for first. Version models supersession, which is unavoidable once figures get restated. File carries checksums and sizes, which makes integrity checkable. Beyond those, extensions tend to describe sensor characteristics that matter for imagery and not for derived carbon products, so adopting them adds surface without adding answers.
Should the geometry be stored as WKB or as GeoArrow encoding?
GeoArrow encoding is preferable where the toolchain supports it, because it avoids a serialisation round trip on read and enables columnar operations on coordinates directly. WKB remains the safer default for datasets that will be read by unknown tools, since support for it is universal. The important thing is that whichever is used is declared in the geo metadata — a file whose encoding must be inferred from the bytes is a file that will eventually be inferred wrongly.
How should partitioning interact with the metadata?
Partition columns should be excluded from the file’s own data but declared in the collection-level STAC metadata, so a consumer knows a partition key exists without reading a file to find out. A common and painful mistake is partitioning by a column that is then dropped from the files, which makes each individual file ambiguous about which partition it belongs to once it is moved. Keep the partition value in the file metadata even when the column itself is elided.
Does a small project need a STAC catalogue at all?
The catalogue’s value scales with the number of files and the number of years, and both grow faster than expected. A project with a dozen outputs can manage with directory conventions and a README, and the same project five years later has several hundred outputs across three schema versions and no way to find which one fed the 2027 submission. Adopting STAC early is cheap; retrofitting it across an archive whose provenance was never recorded is mostly archaeology.
What belongs in the collection rather than the item?
Anything true of every item: the licence, the providers, the summary of available fields, the extents that bound all items, and the column conventions themselves. Repeating those per item makes them expensive to correct and lets them drift between items. The rule of thumb is that an item carries what distinguishes it from its siblings, and everything else belongs one level up.
How should a schema change be handled in the catalogue?
As a new collection when the change is breaking, and as a version bump within the collection when it is additive. Adding an optional column is additive; changing a column’s unit, its null semantics, or its meaning is breaking even if the column name is unchanged, and treating that as additive is how a consumer ends up summing two incompatible vintages. The version extension makes the distinction visible, and a validation gate keyed on schema version makes it enforceable.
Is a checksum in the catalogue worth the cost of computing it?
Yes, and it is cheaper than it appears because it is computed once at write time when the bytes are already in memory. It buys two things: detection of silent corruption, which does happen in long-lived archives, and detection of an unrecorded rewrite, which happens rather more often. The second is the real value — a file that changed without its catalogue entry changing is a provenance break, and the checksum is the only cheap way to find it.
Where should the convention itself be published?
Somewhere a consumer can reach without asking anyone — a schema document alongside the catalogue, versioned with the same identifier the files carry. The convention is part of the data’s meaning, and a dataset whose column semantics are defined only in a wiki behind an organisation’s login has the same long-term problem as one with no definitions at all. Publishing the convention with the collection costs a file and removes an entire class of question.
The convention document should also state which of its rules are enforced by a validator and which are advisory, because the two age differently: enforced rules stay true, and advisory ones quietly stop being followed.
The same logic applies to the vocabulary a column draws on. A land use class column whose permitted values are enumerated in the convention is checkable; one whose values are whatever the producing pipeline happened to emit will accumulate near-duplicates — “forest”, “Forest”, “forest_land” — that no aggregation will notice and every join will silently drop.
Related guides
- MRV Data Schema Reference — the parent topic and the wider schema conventions.
- Canonical Parquet Schema Data Dictionary for MRV — the shared column definitions this convention validates against.
- Versioning Emission Factor Databases for Reproducible MRV — what the emission factor version field points at.
- COG vs Zarr vs GeoParquet for MRV Workloads — choosing the format this metadata describes.