DEMETER

A world model for the living field.

Etherion / Messis / DEMETER

Agriculture is the oldest human technology. DEMETER is a small step in AI’s return to it: a 963M-parameter world model that learns how a crop’s state evolves under weather, and projects the field forward from its observed history. It is maintained as a forcing-conditioned predictive state, not a score.

One disclosed baseline: each field’s own continuity floor, computed on the same windows as the model error.

Held-out field-parents

Model error vs each parent’s own persistence floor

Measured result

Model error versus each held-out field-parent's persistence floor
Research infrastructureNVIDIA Inception Program

963,270,656

parameters

30

climate anchors · every Köppen–Geiger group

16

supported crops

53

held-out field-parents

00 / Primary measured result

Below each field's own continuity floor, on 96% of parents.

Persistence is the honest reference for a next-state task: the next state equals the latest observed state, and the floor is the error that the naive forecast makes on the same windows. The model must beat each field's own floor — not a global average.

51 of 53 field-parents beat their own floor.

A parent is one independent held-out field series. The median next-state error is 0.0739 against a persistence floor of 0.0807; the median parent-level error is 0.1501 against 0.1869. Skill versus persistence at the median is a measured 8.4%.

Because MSE is quadratic, a mean over a few windows is a rainfall-spike detector. This release leads with the median, the floor, the parent win-rate, and the tail — never a lone mean.

51 / 53

field-parents beating their own floor

0.0739

median next-state error (normalized)

0.0807

median persistence floor (same windows)

8.4%

median skill vs persistence

0.3961 / 1.442

model p90 / p99 per-window error

0.5328 / 2.199

floor p90 / p99 per-window error

01 / Projected versus observed

Predicted versus observed, in native units.

On the same held-out fields, the projected state is compared with the observed state. Every value here is native — nothing is normalized.

Feature / unitObservedPredictedRMSECorrelation
NDVI / index0.069550.069700.021840.973
Soil moisture / m3/m30.20510.20480.14540.560
Temperature / °C21.0420.981.7910.977
Rainfall / mm/day2.3982.3125.5470.393
ET0 / mm/day4.4474.4440.85450.912
Humidity / %59.2959.288.0620.931
Wind speed / m/s2.5352.5800.87280.721
Solar radiation / MJ/m2/day19.2719.163.7060.818
GDD / °C-day1882.021881.851.9841.000
ET deficit / mm/day2.0492.1325.9420.503

FIGURE 02

Predicted vs observed

Predicted vs observed

Predicted against observed field state in native units, one point per scored field, y = x drawn. Tight parity along the diagonal; the observed range on each axis keeps the error honest.

Source: Held-out field-parents · native units

Correlation is the Pearson correlation between predicted and observed field state across the scored fields (1.000 is perfect). Canopy condition and temperature are tracked closely; soil moisture and rainfall are the honest hard cases, reported exactly as measured — in native units, never normalized.

02 / The tail, disclosed

No mean without spike accounting.

A handful of unpredictable target-day forcing jumps dominate any mean over a sliding-window panel — and the persistence floor is equally large on those same rows. So the tail is stated, not hidden.

FIGURE 03

Tail robustness

Tail robustness

The tail is stated, not hidden. Per-window median, p90, and p99 are shown for the model and the floor, with the same-day forcing jumps that dominate any mean.

Source: p90/p99 · spike accounting disclosed

How to read a mean

The model’s median per-window error is 0.0739; its p90 is 0.3961 and its p99 is 1.442. The persistence floor on the same windows is 0.0807 at the median, 0.5328 at p90, and 2.199 at p99 — the tail is shared, not a model failure.

Every reported row in this release carries the median, the same-window floor, the parent win-rate, the per-feature physical errors, p90/p99, and skill versus persistence together.

03 / Task contract

One task, fully disclosed.

What the model receives, what it predicts, how it is scored, and what is held out.

  1. 01

    The model processes a 60-day causal window of one field and emits 59 next-day predictions — one per position. The released result scores the window's final position.

  2. 02

    The forcing is the ERA5 weather block; the state covers vegetation condition and soil moisture.

  3. 03

    Fitting covers 2017–2023; the 2024–2026 streaming holdout is locked out of all fitting.

  4. 04

    Overlapping daily windows are reduced to their distinct field-parents before any aggregate is reported.

  5. 05

    The predicted latent is decoded with a per-checkpoint linear readout fit on training and validation, and scored once on the holdout.

Persistence floor

Persistence predicts zero residual: the next state equals the latest observed state. It is computed on the same windows as the model error, so each field carries its own difficulty. It is the only baseline in this release.

53

held-out field-parents

10,713

streaming windows

Representative rule

The published benchmark is one crop specialist, trained across 30 climate anchors that span every Köppen–Geiger group. The product is organized by crop, and all 16 ship.

04 / World operations

The result is a trajectory, not a scalar.

A tabular model returns a point estimate. A maintained predictive state supports a family of operations over the same trajectory — and keeps every step inspectable.

FIGURE 04

World operations surface

World operations surface

One forcing-conditioned state supports prediction, simulation, observation, speculation, windback, intervention, and multi-head inference. Operational utility requires separate policy evaluation.

Source: One maintained predictive state · roadmap surface

Prediction

The trajectory that follows an observed state and forcing.

Simulation

What changes under different rainfall, heat, or management.

Observation

Which dimensions move, and where surprise concentrates.

Speculation

Which counterfactual futures remain consistent with changed forcing.

Windback

Which past disturbance best explains present divergence.

Intervention

Which action sequence satisfies a future state constraint.

Multi-head inference

Stress, phenology, harvest, risk, and confidence on one trajectory.

Operational utility — water saved, stress avoided, intervention acceptance, and yield or quality outcomes — requires separate policy evaluation and is not folded into the predictive benchmark.

05 / World indexes

A prediction that arrives with its own reading.

One error number says how the model did on average. It cannot say what this projection rests on, how steady it is, or where it is fragile. Every DEMETER projection ships with two indexes, reported in the field's own physical units, so a number becomes a statement about one field.

Stability / Instability

What this projection is sensitive to, how steady it stays, which inputs disturb it only in combination, and how far the tail can reach. A stable field reads stable; a field sitting in an unstable regime says so — before anyone acts on it.

Attention / Influence

What the projection actually depends on, measured rather than assumed — kept deliberately separate from what the model draws upon. Where the two disagree is the finding: inputs that are read without moving the answer, and inputs that stay quiet while carrying it.

Why a maintained state, not a fit

A fitted curve returns an answer and forgets the question. Asking what a projection rests on, what would change it, and how far to lean on it requires a maintained state of the field under its weather — something that can be perturbed, rolled forward, and constrained. The indexes are not a report bolted onto a score. They are what the object is.

How to read them

These are statements about one field’s physics, not a league table. Confidence is a band for the model on that field; it does not rank one field above another.

06 / Figure set

Four figures. Each one evidence, not decoration.

Every figure is rendered from the hash-preserved holdout artifact and carries its floor and unit footer.

FIGURE 01

Parent win-rate

Parent win-rate

Every independent field-parent, plotted as model error against its own persistence floor. Points below the diagonal beat the floor: 51 of 53 parents do.

Source: 53 held-out field-parents · ERA5 holdout · persistence floor

FIGURE 02

Predicted vs observed

Predicted vs observed

Predicted against observed field state in native units, one point per scored field, y = x drawn. Tight parity along the diagonal; the observed range on each axis keeps the error honest.

Source: Held-out field-parents · native units

FIGURE 03

Tail robustness

Tail robustness

The tail is stated, not hidden. Per-window median, p90, and p99 are shown for the model and the floor, with the same-day forcing jumps that dominate any mean.

Source: p90/p99 · spike accounting disclosed

FIGURE 04

World operations surface

World operations surface

One forcing-conditioned state supports prediction, simulation, observation, speculation, windback, intervention, and multi-head inference. Operational utility requires separate policy evaluation.

Source: One maintained predictive state · roadmap surface

07 / Public provenance

Three public source families. One causal frame.

The benchmark uses real environmental observations. The source families are public.

NASA SMAP

Soil evidence

Public soil-moisture observations contribute the ground-water side of a crop system's evolving condition, in native m3/m3.

Sentinel-2

Vegetation evidence

Public satellite observations contribute canopy and vegetation evidence across the trajectory.

ERA5

Meteorological evidence

The Copernicus ERA5 reanalysis supplies the physical forcing context for the next state.

Sixteen supported crops

Maize, sorghum, pearl millet, cassava, groundnut, cowpea, rice, sweet potato, yam, sesame, soybean, coffee, wheat, sugarcane, tobacco, cocoa.

08 / Research program

Measured now. Roadmap next.

The release states one measured task contract. The research program widens the same world object without overclaiming it.

Richer temporal coverage

Longer and denser held-out history to measure how the floor moves with more distinct field-parents.

Broader operating scope

More climate anchors and more crop profiles extend the same object, each measured before it is claimed.

Spatial sensing for delicate crops

For high-value crops where geometry matters, Etherion intends to integrate spatial LiDAR insight with temporal prediction.

Measured versus future

Measured now: the next-state result on 53 held-out field-parents, with the physical table and the tail. Roadmap: richer temporal coverage, broader scope, and spatial sensing.

09 / The operating model

Messis is the path from model to observability layer.

Messis is the agricultural decision engine being built on Etherion's own infrastructure — the next stage of this work, not a capability of this release.

Daily observability by reader level

The same evidence becomes practical instructions for field workers, inspectable context for agronomists, and regional briefs for cooperatives.

Operations over one state

Prediction, simulation, speculation, windback, and intervention run over one maintained predictive state, with the evidence path preserved.

Organizational data sovereignty

The organization is to retain full sovereignty over the raw data Messis produces for it. Insight can travel; ownership of observations remains with the organization.

Measured versus in development

This section describes the delivery layer under construction. The only measured result on this page is the next-state benchmark above.