AI · mapping

Finding similar features without training data

Point at a few example cells on a Sentinel-2 image; the foundation model ranks the other cells of the scene by similarity. The models are used as is, without labelling or training.

Expertise France · 2026

River meanders found by similarity search on a Sentinel-2 image
Tana valley, Sentinel-2 from January 2024, Clay model: four example cells on the river (green); the 40 best results (yellow) follow its meanders. Right: their thumbnails, to accept or reject.

Who it is for

For a first screening of an area (irrigated plots, settlements, quarries, water bodies), before finer mapping or to gather training examples.

What the application does

Search by example

A click on the map adds a positive or negative cell. Once embeddings are computed, the search updates on each click, validation or setting change.

Several models

Clay (80 m patch), TerraMind (160 m) and a no-AI baseline. Two models can be combined, for example optical and radar.

Validation and export

Accept or reject each result on its thumbnail, automatic threshold, linear classifier from 5 positives and 5 negatives. GeoTIFF and GeoJSON export.

How it works

  1. 01

    Scene

    Sentinel-2 L2A image (10 bands) over the drawn area, and optionally the Sentinel-1 image closest in time or a radar series.

  2. 02

    Embeddings

    One vector per patch, frozen weights: a few minutes per model on GPU for a scene.

  3. 03

    Score

    Cosine between each cell and the mean of the positives minus that of the negatives. A cell is the mean of the patches it contains.

Screenshots

Continuous score map of the similarity search
The same search as a continuous score map: how much each cell resembles the examples, across the scene.
Irrigated plots found by similarity search
Irrigated plots: eight example cells (green), 21 results above the threshold (yellow).

Known limitations

  • Cells of 80 m minimum (160 m for TerraMind): the tool finds areas, not small objects.
  • No quantified comparison yet between Clay, TerraMind and the no-AI baseline. For water or bare soil, the baseline may do as well.
  • Cloudy cells are excluded in optical imagery. Radar avoids this but measures something else: roughness and the structure of buildings and vegetation.

Going further

  • Count true positives on a few queries to compare the models
  • Export validated results as training examples

Role

  • Application design
  • Model comparison
  • Development and GPU testing

Technologies

PyTorch · TerraTorch · Clay · TerraMind · Sentinel-2 · Sentinel-1 · STAC · ipyleaflet · Voilà

Context

Demonstrator designed and built in 2026 by Guillaume Rieu, author of the Earth Innovation Labs specifications, as a volunteer contribution to the acceptance testing of these platforms in Kenya for Expertise France. It runs in the platform's JupyterHub environment.

A similar need?

These demonstrators can serve as the basis for a tool adapted to your data, your area or your question.

Contact us