Embeddings are one way we make use of the rich representations learned by geospatial foundation models, supporting everything from search and retrieval to natural language interfaces. Yet for many practitioners, the concept remains opaque. This primer explores how embeddings are generated, the kinds of questions they help answer, and why they’re useful across the geospatial ecosystem.
Illustrations by Gus Becker
Every day, satellites collect enormous amounts of data about the planet: optical imagery, radar, elevation, weather observations, and more.
Traditionally, geospatial analysis has depended on turning that raw data into insights via remote sensing methods, statistical analysis, or computer vision/machine learning techniques. Band indices like NDVI or NDWI, land cover classifications, and statistical summaries help transform imagery into something we can analyze and reason about.
Those approaches still matter. Many are scientifically rigorous, interpretable, and extremely effective for targeted problems.
But a growing class of geospatial models are starting from a different premise. Instead of relying only on human-defined categories or measurements, foundation models learn representations directly from Earth observation data. The outputs of those systems — embeddings — are becoming an increasingly important building block for retrieval, similarity search, clustering, and other emerging geospatial AI workflows.
If you’ve spent the last year hearing terms like “GeoFM,” “foundation model,” or “embedding” without a clear mental model for how they connect together, this post is for you.
The illustrations throughout this post—and the accompanying downloadable poster and mini-zine— use a crocheting metaphor to visualize a core idea behind geospatial foundation models and embeddings: transforming large volumes of raw Earth observation data into reusable representations that can support many downstream tasks and workflows.
From Raw Data to Foundation Models
Most modern geospatial embeddings are generated using geospatial foundation models (GeoFMs): large machine learning models trained on massive collections of Earth observation data.
Unlike models designed for a single task, foundation models are trained to learn general representations that can be reused across many different applications, including classification, retrieval, similarity search, and change detection.
Training data for foundation models may include:
- Sentinel-1 radar imagery
- Sentinel-2/Landsat optical imagery
- Aerial imagery
- Elevation models
- Climate and environmental datasets
There are two separate stages involved here that are easy to conflate: training a foundation model and generating embeddings.
High level workflow diagram showing geospatial foundation model training and inference.
First, a geospatial foundation model is trained on large amounts of Earth observation data. This is the expensive, large-scale process usually performed by research labs, institutions, or other model providers. During this stage, the model learns patterns and relationships within the training data.
Later, that already-trained model can be used to process new imagery and generate embeddings. The model applies what it learned during training to infer new representations of the new data. No new learning happens during inference.
A growing ecosystem of geospatial foundation models has emerged in recent years, including Clay, AlphaEarth, OlmoEarth, Tessera, and Prithvi.
Different models generate different kinds of embeddings depending on a variety of factors, including the data they were trained on, geographic coverage, spatial resolution, temporal frequency, training objectives, and how time and space are represented within the model itself. This is why two foundation models trained on similar imagery archives may still produce different embeddings and perform differently across tasks.
What are embeddings
A satellite image contains enormous amounts of information. While imagery records values across pixels and bands, embeddings represent patterns and relationships within that imagery in a new form.
Embeddings are not the imagery itself, but rather a compressed representation that preserves patterns and relationships the model learned are important during training.
An embedding represents this information as numbers grouped into dimensions of a “latent space.” We are used to describing our world with three spatial dimensions, but dimensions in latent space are mathematical rather than physical. Together, they encode patterns and relationships the model has learned from the data.
As humans living in a three-dimensional world, it’s hard to imagine a space with dozens, hundreds, or even thousands of dimensions. This makes embeddings difficult for humans to interpret directly, but computers have no problem working with these mathematical representations. In practice, an embedding is usually stored as a list of numbers.
Each dot is a place on Earth. As the dots move from their geographic positions into embedding space, deserts, forests, and ice sheets group with similar places, even places on different continents. Visualization by Earth Genome.
Unlike a classification, an embedding does not assign a label such as “forest,” “urban,” or “water.” Instead, it describes data in a way that allows models to compare one piece of data with another.
Embeddings sit between raw Earth observation data and downstream applications. Rather than building directly from raw imagery every time, many modern AI systems first generate embeddings, then use those embeddings to support search, retrieval, similarity, and other tasks.
What embeddings can do
Embeddings provide another way to represent and search geospatial information. For some use cases, embeddings offer a way to answer questions about our planet faster, more efficiently, and more cost effectively than with traditional remote sensing techniques.
A few emerging use cases for embeddings include:
Similarity Search and Retrieval
Finding places or objects that resemble another place.
Embeddings make it easier to search large collections of imagery and spatial data based on similarity and context rather than metadata alone. Instead of filtering only by coordinates, timestamps, or keywords, systems can retrieve embeddings representing pixels or patches in imagery because they resemble other embeddings representing pixels or patches at a different point in space or time.
Similarity search results. The reference image is the first one on the left, the others are selected by similarity score. This is useful to find features such as pools or solar panels in no-code applications.
For example, imagery of agricultural regions with similar growing patterns may produce similar embeddings, while imagery of dense urban areas may produce a different embedding pattern. The model is not necessarily assigning labels to those places. Instead, it is learning how they relate to other examples it has seen.
Clustering and Exploration
Embeddings can help organize large geospatial datasets into groups with shared characteristics, making exploratory analysis easier at a planetary scale (for example, an entire cluster can be assigned a label by manually inspecting only one or two examples.)
3 of 64 embeddings dimensions visualized in RGB near urban and agricultural areas near Longmont, Colorado.
Change Detection and Anomaly Identification
Embedding-based approaches can help surface unusual patterns or regions that diverge from expected conditions without requiring every possible change type to be predefined ahead of time.
Post Cameron Peak Fire Burn Scar: Sentinel-2 2021 (left) and AlphaEarth Embedding Change Intensity based on Cosine Similarity (right).
Median pixel level embedding change calculated with cosine similarity indicates that embeddings are picking up on changes in the landscape the fire caused. Pixels outside the fire were randomly sampled forest and rangeland classes to match the land use land cover types within the burn scar.
This changes the kinds of questions geospatial systems can ask. Instead of only asking, “Does this pixel belong in class X?” systems can ask, “What else looks like this? What changed? What patterns are similar across regions or time?”
New Kinds of Workflows
That shift is also enabling new ways of interacting with Earth observation data. With the rise of satellite-specific vision-language models, a growing number of platforms and research systems are exploring natural language interfaces that let users ask questions such as, “Show me examples of burned areas in the western United States.”
Behind the scenes, embeddings help connect those human-language questions to relevant imagery, locations, and environmental data. While the user interacts through language, the underlying system relies on embeddings to identify relationships and retrieve meaningful results.
Interpreting Embeddings
By now, you may have noticed that embeddings feel different from many geospatial data products practitioners are used to working with.
Imagery, maps, and classification layers are designed to be interpreted by people. We can look at them, assign meaning to them, and use them to answer specific questions about the world. Embeddings are different. They are primarily designed for computational tasks such as similarity search, retrieval, clustering, and pattern discovery.
This is one reason embeddings can feel abstract at first. The individual values within an embedding are usually not meaningful on their own. What matters is how embeddings relate to one another and the patterns those relationships reveal.
Visualizing embeddings can still be useful. While embeddings are not intended to be viewed in the same way as imagery or maps, visualization can help build intuition and make these representations more approachable. New cloud-native geospatial tools make it possible to explore embeddings as datasets and identify patterns that may not be obvious in traditional imagery alone.
One open-source example is deck.gl-raster, which can be used to visualize embeddings stored in formats such as Zarr directly in the browser. Tools like these help bridge the gap between machine-readable representations and human understanding.
It is also important to remember that an embedding is not a universal representation of “truth.” Embeddings reflect the data, assumptions, and objectives of the foundation model that produced them. Different models may produce different embeddings from the same imagery because they were trained with different goals, datasets, or architectures. Most foundation models are associated with scientific papers, and some models even have open source code and model weights to increase transparency and provide users with insights on choices made during training.
The same farmland on Colorado’s Western Slope, embedded by Tessera (left) and AlphaEarth Foundations (right).
Many geospatial workflows still rely on predefined categories and carefully designed measurements, classifications, and traditional remote sensing or machine learning methods. Embeddings provide another way to represent geospatial information and explore relationships within Earth observation data.
Accessing Embeddings
Once a foundation model has been trained, there are generally two ways to work with embeddings.
Some providers release pre-generated embeddings datasets that users can immediately explore and analyze. Google’s AlphaEarth Foundations, for example, includes embeddings generated across large portions of the globe for 2017-2025 and is made available as a dataset on Google Earth Engine. These embeddings are also available via Source Cooperative through an open-source community effort that converted, packaged, and documented access to the dataset. These releases make it possible to experiment with embeddings without needing to run a foundation model yourself.
Other providers publish foundation models, open-source code, and model weights that allow users to generate embeddings of their own areas of interest and time ranges. Tessera, OlmoEarth, and Prithvi are examples of projects that make models and tooling available to the broader community.
As more organizations begin working with embeddings, a growing ecosystem of datasets, tools, and cloud-native workflows is emerging to help store, share, visualize, and analyze them at scale.
What Comes Next
Embeddings are still an emerging area of geospatial technology. New foundation models, datasets, tools, and workflows are appearing quickly, and many aspects of the ecosystem are still evolving.
But the core idea is simple. Embeddings provide a way to represent complex Earth observation data so that machines can more easily search, compare, and identify relationships within it. They don’t replace traditional geospatial analysis; instead, they offer another way of working with Earth data, particularly when similarity, retrieval, discovery, and large-scale exploration are useful.
Understanding what embeddings are, where they come from, and the kinds of questions they can help answer is a useful first step toward understanding the broader shift taking place across geospatial AI.
Explore Further
Ready to dig deeper? Here are a few resources, datasets, and tools that can help you explore geospatial foundation models and embeddings in practice.
View the Full Visual Guide
From Earth Data to Embeddings Illustrated Poster & Mini-Zine
The illustrations in this post were designed as a visual companion to the concepts covered throughout this article. View the complete infographic or download the mini-zine to explore the relationship between Earth observation data, foundation models, embeddings, and their application in a single visual reference.
Learn More about Geospatial Foundation Models and Embeddings
-
A beautiful visual explainer of embeddings from Ode.
-
Tutorials and notebook examples from Klemmer, Konstantin, et al.
-
Visualize Embeddings at a Planetary Scale from Earth Genome.
Explore Embeddings Through Real Applications
Access Open Embeddings Datasets
- Source Coop: AlphaEarth, Clay v 1.5
- AWS Open Data: Tessera
Open Earth Foundation Models
Related content
More for you
What we're doing.
Latest


















