Breakpoint

How Google automates geospatial outbreak prediction

Building a geospatial prediction model for a crisis takes specialist teams weeks: finding covariates across portals, cleaning them, guarding against…

google research··PT2M15S

video loads only when you press play

Building a geospatial prediction model for a crisis takes specialist teams weeks: finding covariates across portals, cleaning them, guarding against…

Building a geospatial prediction model for a crisis takes specialist teams weeks: finding covariates across portals, cleaning them, guarding against target leakage, training and validating. Google Research's Planetary Prediction Engine automates the whole workflow from a natural-language query, in minutes.

  • The system turns a natural-language question into data discovery, curation and model selection.
  • Feature gates remove columns that leak the answer or use information unavailable at prediction time.
  • Held-out places test whether a model generalizes geographically instead of memorizing known regions.

Right now, an outbreak is moving faster than the map of it. The model that predicts where a virus goes next takes specialist teams weeks to build, and the virus doesn't wait. Most of those weeks aren't even spent on the model. They go on hunting the data: case counts in one portal, roads in another, rainfall in a third, each in its own format, each needing to be cleaned, lined up and checked before anything can train. Google just built a system that does that whole job itself. You ask it a question in plain English, and it hands back a trained model and a full report, in minutes instead of weeks, working through 3 stages. First, it goes looking for the data. It reads the question, decides which signals might matter, and pulls them from the big public catalogues, and when a signal isn't in any catalogue, it searches the open web while it works. Then it adds something no spreadsheet has. Google has two models that compress the world into a fingerprint per place: one learned from how people live somewhere, their searches, their businesses, how busy the streets are, and one learned from satellites, describing every 10 metre square of land on Earth. Both get stitched onto the table. But before anything trains, every column has to get past a gate, because the easiest way for a model to score well is to cheat. A column that quietly contains the answer, or data leaking in from the future, gets thrown out. The third stage trains a whole family of models, from simple lines to decision trees to small neural networks, watches itself for memorising instead of learning, and keeps the one that holds up on places it has never seen. Across dozens of public health indicators in the United States, it beat the pipelines experts had built by hand. In Nigeria, it took food security estimates that only existed for whole states and predicted them district by district, twice as accurately as the standard approach. Then came the real test. This year, Ebola broke out in the Democratic Republic of the Congo. For 5 weeks in a row, the system forecast which health zones the virus would reach next, and it called 15 of the 18, ten points ahead of the published expert model. The people who used to spend weeks building the map can now just ask for it. And for the first time, the map moves faster than the disease.

This explainer is based on Planetary Prediction Engine: Automating global models via Earth AI by Google Research ↗. The original reporting and technical work belong to its publisher.