Tag
This paper presents a systematic comparison of two geospatial foundation models, TerraMind and THOR, developed under ESA's φ-lab, analyzing how architectural choices like patch size and decoder type affect performance across ten use cases in Earth observation tasks.
The first three operational FireSat satellites, backed by Google and Bezos Earth Fund, launched to detect wildfires as small as 5x5 meters, aiming to provide hourly global coverage by 2029.
This paper surveys the emerging paradigm of Geospatial Foundation Models (GeoFMs), which are pre-trained on massive geospatial datasets to enable rapid fine-tuning and zero-shot analysis of satellite and aerial imagery. It covers the paradigm shift, model adaptation strategies, and a forward-looking vision of Agentic Geospatial Reasoning using LLMs as orchestrators.
This chapter reviews design principles and current landscape of foundation models for Earth observation, highlighting the need for domain-specific adaptation, physically plausible representations, and consistent evaluation benchmarks. It includes case studies on harmful algal bloom prediction and adaptive monitoring station selection.
The paper presents the largest controlled scaling study for Earth-observation foundation models, showing that pretraining loss poorly predicts downstream performance and providing an optimal compute allocation rule. It trains scaled pixel-wise models (0.5B and 1B parameters) and distills them into compact student models that outperform larger open and proprietary models.
EO-Agents presents a three-agent LLM pipeline for generating Earth observation hypotheses, leveraging a NASA knowledge graph and graph neural network to rank candidate dataset pairings, with LLM agents filtering, generating, and evaluating structured research hypotheses.
EO-WM proposes a video diffusion transformer for probabilistic Earth observation forecasting that incorporates physically informed conditioning to capture weather-driven uncertainties, achieving improved prediction of vegetation indices under extreme weather.
UniverSat introduces a Universal Patch Encoder for Vision Transformers that enables robust, sensor-agnostic spatial feature extraction across diverse Earth Observation data types, achieving strong results on classification and segmentation benchmarks.
NAVI-Orbital demonstrates the first in-orbit deployment of a zero-shot vision-language model (Gemma 3) on a LEO satellite, enabling autonomous scene classification and semantic compression of Earth observation data without fine-tuning.
A satellite called Yam-9 used Google DeepMind's Gemma 3 vision-language model in orbit to autonomously identify areas of interest based on natural language queries, marking the first reported use of a VLM in space and signaling a shift toward more autonomous satellite operations.
This paper proposes HADT, a transformer-based architecture for autonomous resource management in heterogeneous satellite clusters for Earth observation, using differential attention and relational tokenization. Experiments show significant improvements over baselines and strong adaptability to varying cluster sizes.
This paper introduces a novel uncertainty-aware PINN framework for flood inference from SAR data, addressing 'physics shock' by dynamically relaxing physical constraints in noisy regions. Evaluated on Sen1Floods11, the method achieves a 25% improvement in IoU and provides calibrated uncertainty bounds for operational disaster response.
This paper presents a unified benchmark for composed image retrieval in Earth observation, evaluating vision-language backbones and introducing a change-centric dataset (xView2-CIR) for disaster monitoring, highlighting distinct challenges compared to attribute-based retrieval.
OlmoEarth v1.1 is a new family of satellite imagery analysis models from Allen AI that reduces compute costs by up to 3x while maintaining performance, achieved by decreasing token sequence lengths in transformer-based models.
This paper audits 152 papers on geospatial foundation models and finds severe lack of standardization, making it impossible to determine state-of-the-art. The authors propose six concrete expectations to improve reproducibility and comparability.
Analyzes the 64-D embedding manifold of Google AlphaEarth across 12.1M U.S. samples, shows non-Euclidean structure and poor vector arithmetic, then builds an agentic system with geometry-aware tools that outperforms parametric baselines on environmental queries.
Google DeepMind introduces AlphaEarth Foundations, an AI model that integrates petabytes of Earth observation data into unified embeddings to map and monitor the planet at 10x10 meter resolution. The model's compact representations enable efficient planetary-scale analysis for applications in food security, deforestation tracking, and environmental monitoring.