Announced 8 Oct 2026 · Sources checked
What did ESA launch on 8 October?
The European Space Agency announced Earth Virtual Expert, or EVE, on 8 October 2026. ESA describes it as an AI-powered platform that answers questions in a conversational format so that Earth observation knowledge, scattered across scientific papers, technical reports, mission documentation, models and repositories, becomes usable by researchers, policymakers, journalists and the general public. Example questions in ESA's announcement include how vegetation in southern Europe has changed over the past decade and how satellites help monitor greenhouse-gas emissions.
EVE was developed by ESA Φ-lab, the agency's innovation unit at the ESRIN Centre for Earth Observation in Frascati, Italy, together with Pi School, an AI research and training company, Mistral AI, and Imperative Space, a company that works on innovation and communication for the space sector. ESA adds that after the collaboration Mistral has signed a letter of intent with the agency, though the announcement does not say what that letter covers.
Two dates matter. The announcement and public opening are dated 8 October 2026. The underlying model is older: Hugging Face lists the eve-esa/EVE-Instruct repository as created on 16 February 2026, and the model's technical documentation gives a release date of 13 February 2026. The training and benchmark datasets in the eve-esa organisation date from late November 2025. So the news is ESA formally opening the full system to everyone, not a brand-new set of weights.
Europe has been pushing for AI it can run on its own terms; for comparison, see our coverage of Aleph Alpha's Kolibri open-weight model and of Mistral Large 4, Mistral's latest flagship preview.

How does EVE work under the hood?
EVE combines three layers. The first is EVE-Instruct, a language model adapted to what the team calls Earth Intelligence. According to the model card, it starts from Mistral Small 3.2 (24B parameters, 128k-token context) and is trained with a mix of instruction data and long-form text, each blending general-domain replay data with Earth-observation and Earth-science material. The final training mixture is about 33.5 billion tokens; because of licensing conditions on some sources, the team publicly releases a curated 10.7-billion-token subset.
Much of the domain data is synthetic. The card says about 21 billion tokens were generated with models including Mistral Large 3, Mistral Medium 3.1, GPT-4o Mini, Qwen3-235B, DeepSeek-R1, DeepSeek V3.1 and Qwen2.5-72B, then filtered by LLM judges for relevance, factual quality and grounding. The team merged checkpoints from ten training runs with different data mixtures and finished with online direct preference optimization using the same recipe as Mistral's Ministral models.
The second layer is retrieval. ESA says EVE does not rely only on what it learned in training: a retrieval-augmented generation system searches additional sources, including recent articles published by Wiley under a collaboration, and folds them into answers. Training text came from ESA, NASA and Copernicus websites, manuscripts and peer-reviewed research. If you want the basics of the technique, our explainer on retrieval-augmented generation covers how it works and where it fails.
The third layer is a guardrail. ESA describes an 'LLM-as-a-judge' feature that reviews generated responses and flags content to revise or filter before users see it, an attempt to reduce statements not supported by scientific evidence. It narrows rather than removes the problem described in our guide to why AI models hallucinate.
How well does EVE-Instruct score?
The EVE team built what it calls the first manually created Earth-observation and Earth-science benchmarks, 5,693 samples covering multiple-choice questions, open-ended answers, answers with supplied context and hallucination detection, according to the project's technology page. On those tests the model card reports clear gains over the base model at zero-shot: 77.73% versus 70.30% accuracy on multiple-answer multiple-choice questions, 96.35% versus 83.51% on single-answer questions, an F1 of 84.70 versus 82.19 on hallucination detection, and an LLM-judge score of 96.40 versus 91.78 on open-ended questions.
The card also compares EVE-Instruct with Llama 4 Scout, Qwen3 30B-A3B and Gemma 3 27B. EVE-Instruct posts the top score in most columns, but not all: on open-ended questions with supplied context, Qwen3 scores 81.81 against EVE-Instruct's 78.28. On general benchmarks the team reports an overall category average of 74.0 against 72.2 for Mistral Small 3.2, with small gains in reasoning, coding, knowledge, tool calling, instruction following and chat quality. Every one of these numbers comes from the developers, measured on benchmarks they created.
| Test | Mistral Small 3.2 | EVE-Instruct |
|---|---|---|
| MCQA multiple-answer accuracy | 70.30 | 77.73 |
| MCQA single-answer accuracy | 83.51 | 96.35 |
| Hallucination detection F1 | 82.19 | 84.70 |
| Open-ended (LLM judge) | 91.78 | 96.40 |
| General-capability average | 72.2 | 74.0 |

How can you use EVE today?
There are two routes. Most people will use the hosted chat at eve.philab.esa.int, which ESA says is open to all interested users. The model's technical documentation describes it as an AI chatbot system, discloses that users are talking to a machine rather than ESA staff, and asks people to verify critical information against primary sources. The project's technology page lists streaming answers, source citations, a choice of knowledge bases and feedback tools in the interface, and a FastAPI-based API for chat, retrieval queries, document ingestion and evaluation.
Developers can run the model themselves. EVE-Instruct is on Hugging Face under Apache 2.0, with quantized variants (GGUF Q4_K_M and Q3_K_M, AWQ, AutoRound W4A16 and W8A8) published between April and July 2026. The team recommends vLLM 0.9.1 or newer with the mistral-common tokenizer, and warns that the full-precision model needs about 55 GB of GPU memory and is not compatible with Ollama. The supporting code, including the backend, frontend, data-scraping and processing pipelines, retrieval experiments and an evaluation kit, sits in Apache-2.0 repositories under the eve-esa GitHub organisation, alongside newer repositories for agents and an MCP tool registry.
That last detail fits ESA's roadmap. The agency says EVE entered a second development phase in September 2026, aimed at retrieving and analysing Earth-observation data and carrying out multi-step tasks, and adding environmental data, weather reports and numerical models. Searching and presenting satellite imagery is planned for 2027.
Why does a domain model from a space agency matter?
General chatbots can already explain remote sensing, but they are weakest exactly where Earth-observation users need precision: mission-specific details, processing levels, indices, the provenance of a dataset and what a recent paper actually found. EVE's bet is that a mid-sized model trained on curated domain text, paired with retrieval from vetted publishers and an answer-checking step, can be more reliable on those questions than a larger general model, and cheap enough to host publicly.
Openness is the other point. Because weights, data subsets, benchmarks and pipeline code are released under permissive licences, universities, national agencies or companies can rebuild, audit or extend the system rather than depend on one hosted service. The licence is genuinely open source in the sense we described in open weights versus open source, though the full training mixture is not public because some source material is licensed. For European institutions that want sovereign tools, a public agency shipping a reproducible domain model is a template others can copy.
What are the limits and open questions?
The evidence is still self-reported. The domain benchmarks were written by the EVE team, some scoring uses LLM judges, and we found no independent evaluation. A 13-point gain on single-answer multiple-choice questions says little about how often EVE gives a subtly wrong answer to an open policy question. ESA's own documentation tells users to check critical information against primary sources.
There are also small inconsistencies to note. The model card text says EVE-Instruct is fine-tuned from Mistral Small 3.2, while its metadata lists the 3.1 checkpoint as the base model. The training data includes private datasets obtained from third parties, and the hosted retrieval layer depends on a publisher collaboration whose scope is not detailed. Finally, EVE is text-only today: it can discuss satellite missions and indices but cannot yet look at imagery or run analyses, capabilities ESA places in its second phase and in 2027. We did not test the hosted service or run the model for this article.
Common questions
Is EVE free to use?
ESA says the hosted EVE system is live for all interested users, and the EVE-Instruct model is free to download under the Apache 2.0 licence. Running it yourself needs roughly 55 GB of GPU memory at full precision, or less with the published quantized versions.
Which model is EVE based on?
EVE-Instruct is a 24-billion-parameter fine-tune of Mistral AI's Mistral Small 3.2 with a 128k-token context window, adapted on Earth-observation and Earth-science text by ESA Φ-lab and Pi School with Mistral AI and Imperative Space.
Can EVE analyse satellite images?
Not yet. Today EVE answers questions from text sources. ESA says its second phase adds retrieval and analysis of Earth-observation data and multi-step tasks, with satellite-imagery search and presentation planned for 2027.
What to remember
ESA's EVE turns a decade of Earth-observation literature into an openly licensed question-answering assistant with retrieval and an answer check; it is worth trying and building on, with the caution that its accuracy claims are the developers' own and imagery support is still a 2027 plan.
Sources & further reading
How this story was made
Written by Kristian Kostov with AI assistance and checked against the linked sources. Company performance claims are attributed to the company. Analysis reflects AiLookout’s interpretation; we have not independently tested the products discussed. Cover photography is illustrative and does not depict the specific announcement or product.
Our editorial standards





