Skip to content
TrackPodcasts
scienceSep 9, 202644:24

Can AI help us better predict the weather?

About this episode

Can artificial intelligence solve one of science’s most chaotic physics puzzles? In this episode, Professor Hannah Fry sits down with Peter Battaglia, Senior Director of Research at Google DeepMind, to explore how machine learning is transforming global weather forecasting. From providing critical early warnings for Category 5 storms like Hurricane Melissa to predicting renewable energy supply and agricultural impacts, see how models like GraphCast and WeatherNext 3 are building upon decades of numerical physics to reshape our understanding of the atmosphere.

Learn more about WeatherNext 3, our most advanced global weather AI model yet: https://deepmind.google/science/weathernext/

Timestamps: 

00:00 Introduction

00:38 Hurricane Melissa

11:50 Why weather forecasting is hard

14:13 Traditional models vs AI models 

21:55 Probabilistic forecasting 

26:00 WeatherNext 3 

33:00 Future outlook 

43:44 Hannah's reflections 
 

Please leave us a review on Spotify or Apple Podcasts if you enjoyed this episode. We always want to hear from our audience whether that's in the form of feedback, new idea or a guest recommendation!


Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.

Get every episode summarized

Each time Google DeepMind: The Podcast publishes, we email you a written briefing from the transcript — the topics, who appeared, and any specific claims, with the ad reads skipped.

Email me new episodes

Free for 3 shows. No card needed.

Hosts & guests

Transcript ready

475 searchable segments. Every word is indexed and playable.

Can AI help us better predict the weather?

Google DeepMind: The Podcast

0:00
44:24

Full transcript

Google DeepMind: The PodcastCan AI help us better predict the weather?. Machine-transcribed; use the interactive transcript above to jump the player to any line.

Welcome to Google Deep Mind The Podcast, I'm Hannah Frey. When a hurricane is heading towards land, time becomes absolutely critical. An extra day of warning can clear the roads, empty the hospitals, open up shelters and get people to safety. Even a few hours can be the difference between a crisis and a catastrophe. But while we can calculate the exact second the sun will rise a thousand years from now, predicting the weather in the future with certainty remains one of the trickiest physics problems on the planet. How do you handle something that is so sensitive to tiny fluctuations, where a small change in today's air pressure can turn into a major storm a week later, the famous butterfly effect. Well, the traditional way is solving fluid dynamics equations with supercomputers. But now there is another way that is doing something completely different. The team at Google Deep Mind have been thinking about the challenge of AI weather prediction for years. They started back in 2020 with models that use satellite images to predict short-term rainfall.

Today, the latest system, Weather Next, provides detailed hour by hour predictions for the entire planet. Joining me to discuss how AI models the weather is Peter Bataglia, Senior Director of Research at Google Deep Mind. Welcome to the Podcast, Peter. Thank you. I want to start with a really concrete example of how AI has made a difference. Tell me the story about Hurricane Melissa. Yeah, so Hurricane Melissa was a very powerful hurricane that struck Jamaica and several Caribbean islands in October 2025. And we had been working on specific models for hurricane forecasting for several years. And earlier in 2025, about a year ago, we released our best weather model and we created a website called Weather Lab that had live hurricane forecasts. We had also established these partnerships with a number of agencies,

including the National Hurricane Center, which is in the US under NOAA. So throughout the season, we had been seeing that the model was fairly accurate at predicting hurricanes. And when Hurricane Melissa, what became Hurricane Melissa, started to form in the Atlantic, about a week before it eventually made landfall, our forecast started showing that it was going to become a very intense hurricane. Because unusually, it started really small, didn't it? And stayed small for a while. Yeah, and this is how they work. They started as a disturbance. It's just some pattern that could eventually, given the warm sea temperature, drop enough energy to then become a very intense and ferocious storm. So about a week before, a little less and a week before, we started to see the model becoming quite confident that it was going to turn into a category five hurricane, which is the strongest category of hurricane. What was at this point? I don't think it was even a tropical storm. You might have been a tropical depression at that point. So actually, we watch the folks on Twitter or X. There's actually a pretty active meteorological

community and they follow a lot of the models. So they started to notice and say, this is the Google model is forecasting a very intense storm. They're saying it's going to be category five and people started to notice and we were really watching intently. Were the other models, I mean, because there's numerous of the simulations that are running, looking exactly that error, at the same time, were they also predicting at this point that it was going to be catching up? Some were predicting that it was going to rise in intensity, but none to my knowledge had the confidence and the confidence around that it was going to become that intense and the specific trajectory and the specific level of intensity. They, in the National Hurricane Center, their forecasters are making, they're constantly observing and making guesses and determinations and trying to refine their own official projection about what's going to happen. They're drawing on all sorts of different information, like traditional models that are well understood, newer models, and they're designed, as a center, they're set up to take lots of different pieces of information

and then stitch that all together to form their forecast. On Saturday, they made an official determination that it would become a category five storm and hit Jamaica. It still wasn't even, I don't think it was even a category one storm yet. They had predicted that it was going to undergo rapid intensification in a two-day, two-and-a-half day period. So this is Saturday and it was predicted to hit land on Monday. I think it was going to reach category five, I think, on Monday and maybe make landfall on Tuesday, if I remember correctly. So this is still actually a fair amount of warning there. Yeah, I think it had about three-day lead time and when they could issue that category five warning, that also prompts triggers a set of preparatory responses and then the local officials know, okay, there's a lot of confidence and information suggests and this is going to be very dangerous. We're going to start evacuating when we do all the different things that wouldn't have taken place had it been at the less intense storm. And this was, we were told later, this was the lowest intensity storm that the National Hurricane Center had ever forecast to become

category five. They had never made a category five forecast from such a low intensity. And we learned later that they said they were heavily informed by watching our model, that our model confidence in the forecast was really increasing their own confidence. Because you were saying it with 80% confidence, right? Yeah, I think it varied by which forecast, because we made forecasts every six hours. But about 80, as the week of unfolded, it went from lower probability to a pretty high probability. So they issued the category five warning, it did strike Jamaica. I think 95 people lost their lives, a huge amount of damage. I think on record, it was the highest wind ever occurred. Yeah, 190 mile an hour winds. Yeah. I mean, the thing is, it's devastating as it was, it could have been worse after less warning. Yeah, we can never really know what would happen. But we do know that these types of operational centers are crucially important for protecting people and protecting property and then also getting the services back online

when to storm passes. So yeah, and we just in general through that whole experience that was, we've had this wonderful partnership with the National Hurricane Center and our researchers talked them all the time. And we learn a lot about how they make their determination. So then we can feed that back into how we operate. So even though it's sort of a case study in a single instance, we actually, that partnership and seeing that whole thing through lets us then learn about how to improve our models and the types of information that the decision makers need to see in order to inform their decisions. We're going to get into the model a little bit later. But I just wonder on this topic of Melissa, do you think that the AI model has a sort of kind of advantage when it comes to predicting hurricanes that far out in the future? Or is it just is it just that you're doing something different to the other models and that's maybe why you get a different zone? The way I think about it, it's pretty simple. What's happened in the past is the same physical process that happens in the future. So when you get enough evidence of the past

and you have models that are capable of really making sense of the fine sort of subtle structure among these input and output and past and future relationships, you just get more accurate prediction. So the better you do your modeling and your machine learning and your evaluation and your data, all the data handling, the better your models are going to be. Because it's still according to the same rules, it's still the same time. Same rules. You just have to extract as much information about that data from the past as you can and then use it to forecast the future. So tell me a bit about this cyclone forecasting that you're releasing to the world. Tell me a bit about that. What was the motivation behind that? We wanted to make a special case of the model that was designed to predict tropical cyclones. These are like some of the most intense, strongest, rarest, really weather phenomena. And we felt if we could do this well, it will have impact. It'll also be a pretty significant scientific challenge that will force us to improve our methods. This sort of phase of our project began in earnest about four years ago and we were working on global models, sort of

forecasting all the weather on the surface and the atmosphere for 10, 15 day forecasts. As we made progress on that, we wanted to have like downstream impact forecasting the actual weather events. So we looked at things like tropical cyclones have been in all of our papers. We've done little analyses. Okay, so this predicted the temperature in seven days, but can it predict the track of a tropical cyclone? So that had been going on for a while, but then we began a sort of focused line of work, work stream with a dedicated sub team about two years ago or so. And we put a lot of investment in this and we also built partnerships right from the get go academics, operational, meteorological centers. We had trusted testers for many months before we released it, but then we released it last year and we made improvements to it over time. When we run the experimental evaluations over the last several years and across the globe, we see about an extra day of accuracy in the sense that a forecast, which would have had an accuracy of some level at a two-day

forecast, we achieved that same accuracy as at a three-day forecast in many cases. And this allows you to make, again, preparations for the storm and plan and issue warnings earlier. And in general, this is going to potentially be a pretty impactful tool for people to use. Do people have access to these forecasts? Okay, if I wanted to know whether it was going to rain, next week, could I look at your forecast? Yeah, so there's different ways to access the forecasts. One way is we actually have a live feed of our forecast and you can access. So you can Google for weather next and you'll find it in various forms. If you want the tropical cyclone forecast, you can go on weather lab. We actually just released beyond the tropical cyclones, like temperature and wind and pressure as well. And if you want to know what the weather is going to be, like if you want to just use your weather app, you can use pixel or search and the weather

next forecast are also contributing to those forecasts as well. So there's a lot of ways, you know, whether you're an enterprise customer and you want to get the data or you just want to kind of see the icon, you can see the forecast. And we were talking about all of these extreme weather events, right? Of hurricanes and cyclones and so on. Is weather prediction becoming more urgent as time goes on? Weather is becoming more extreme as time goes on. That's pretty clear. The real reason is because weather is energy. I mean, you have like, you know, the sun is electromagnetic radiation, the wind is mechanical energy. Temperature is kinetic energy. When you have a warmer earth, you have more energy. And so the weather will be more intense and that intensity also manifests is less predictable at least by historical standards because it's sort of changing weather patterns. And we're seeing that there's, you know, warmer temperatures globally, I think last year, the year before was the hottest year on record. And if you look back over the last 10, 15 years,

all the hottest years on record have basically been occurring in the last few decades. And so I think, you know, you see much more wildfires, other types of flooding and things like this. So I think as we see AI advance, it's calling is there to try to apply it to some of the challenges that are also becoming more difficult and important. I can imagine there's some people listening to there too, like, look, why is it so hot? We've done amazing, phenomenal things of science. Why can't you tell me whether it's going to rain next Wednesday, you know? What makes weather forecasting hard is that little things can have big effects. So this is like the butterfly effect. If a butterfly flaps its wings or it doesn't, you know, this can have actually an impact on what large scale if it's a storm or it's clear or what. Well, we can't observe every butterfly. So there's a fundamental limit to what we can observe and that then translates to a fundamental limit and what we can predict. But that doesn't mean that again, on the basis of

all the evidence we've ever seen, we can't detect subtle little patterns that maybe have gone unnoticed historically with previous methods and as technology advances and gets better, it's always going to see an increase in the ability, our ability to predict things more accurately. That's the reason it's hard. It's hard because we don't observe important features of that drive it. But they do lead little hints and breadcrumbs that we can capture with statistical learning. Because when people talk about the butterfly effect, I mean, they don't, it's sort of not meant metaphorically, right? It's like, it's sort of quite literally true, right? The in theory about a fly could, I'm not sure how many butterflies have actually caused hurricanes, like I couldn't say. But the point is that the atmosphere is a fluid and the fluid has a particular type of physical characteristics that mean that little things can make big impact. So again, even if you drop like a small rock into a pond, you might see a little ring or a little wave or a little splash of

drop of water. And if you can capture that and understand that that indicates that there's probably something happened that will cause a sort of concentric circle to grow, you can take advantage of that. So this is where I think AI and machine learning are really how they work. They're able to realize that I've seen droplets before that got splashed and little ring over here. And then what happens is you get a bigger ring of, you know, little waves that hit the other shore of the pond. When you put it like that, it sort of does make sense that this might be an AI problem because it's pattern recognition in a lot of ways. But why wasn't AI lab like Deep Mind? Why did you guys get involved in weather forecasting in the first place? What was the motivation? I think there was two main reasons. So one is because weather is very important. It's like probably the oldest prediction problem, one of the oldest problems in science. And, you know, part of our mission at Deep Mind is to try to solve major scientific challenges. The humanity faces, you know, one of the goals of AI is to be able to, you know, increase our ability to solve major

scientific challenges that we face. Whether itself is a challenge is not, but it's not a, you know, kind of abstract problem. It's actually touch it literally. Everybody knows a lot about weather. And it impacts everyone's lives every single day. For me, more personally, I had been working on machine learning for simulation for many, many years, even before this. And we had even been working on using machine learning to simulate fluids. So several teams at Deep Mind came together. Folks were interested in weather. Folks were interested in simulation. And we realized that the methods that we had probably could apply to some of the methods or some of the problems that in weather. And then one other thing that actually helped us, helped push us over to work on weather was that there was amazing data sets on weather that were already available. So the ECMWF, the European Center for Media Range, weather forecasting, have been building these records of weather on Earth that like span decades. So we see this and we're like, this is perfect for machine learning. Like, this is, we're so fortunate to be able to build on this. And it's no surprise that they also have

very powerful good AI weather models as well. So when you have, when the time is right, you have the right sort of raw materials and then you have like the teams that can put this together and it's aligned with your missions, something that's a good choice to work on it. We should probably describe actually what the more traditional numerical method is. What was happening before? How would we solve the weather before? Yeah, so one of the dominant methods of forecasting the weather is called numerical weather prediction. So the way this works is you have a supercomputer and this, it runs an algorithm that takes in the current estimate of the weather on the Earth. And then it makes a prediction about the weather in the next, you know, hour or six hours. And it takes its own prediction and feeds it back in and it makes another prediction for a few hours later. And it's all based on essentially physics equations. So the, yeah, the algorithm itself, it's an approximation to the solution to the equations of fluid motion. So again, the atmosphere

is a fluid. And, but if you say, I only know, you know, the variables of the present and I want to solve for the variables of the future, I need to solve that somehow. Now this, because again, because of the butterfly effect and because there's coupling across scales, it's very difficult to solve the equations exactly because you'd have to know very, very detailed information to be able to solve for the whole thing. It's a very big problem. So instead, they approximate this and they approximate it by simulating the weather at a sort of course, blurry scale. And they find approximate solutions to those equations. And I mean, really in practice, this has been a sort of triumph of science and engineering for decades, right? The fact that you can like know what's going to happen in a week and a half or two weeks, like what's the sky, you know, what there's rain, there's a fall, there's something amazing. Really? Like this was never historically, it was never possible, maybe a day or two in advance, but it just sort of shows that through decades of, you know, physics and engineering and computing and data collection, this has already been a huge

triumph for, you know, humanity and an amazing demonstration of like, you know, human ingenuity intelligence. But then how does Google DeepMinds approach differ from from everything that went before it? So these AI weather forecasts that we're building are based on looking at a history of weather and how it unfolds in time and then learning the statistical patterns. And on the basis of the past patterns, making predictions about the future in the same way. So in the same way that you might like fit a line to a trend of numbers that are increasing, here we fit a very complicated line to a very complicated set of weather from the past and then extrapolated out into the future. Talked me through the lineage then of AI and weather forecasting, like how is it progressed over time? So the earliest use of machine learning and AI and weather forecasting was usually to help traditional numerical models fill in the blanks, especially in the finer scale structure that they

weren't explicitly simulating. Then we saw as the image modeling work in machine learning evolved in the last maybe eight, ten years, a new type of weather, AI weather model that was predicting, say, the rain over a country for 90 minutes or directly trying to predict the wind velocities in a small region. They weren't using numerical weather prediction, but they also weren't simulating all of the weather over the earth. They were just trying to basically do image or video modeling, but with weather images and video. And this third phase we're in now is simulating all of the weather over the earth in the same way that the numerical weather prediction methods are also simulating them, so they're trying to predict what's happening in the surface and in the atmosphere at all parts of the globe. And in principle, these are able to capture everything about

the weather, at least at the resolution they're modeling the same way the numerical weather prediction does. That thing you're describing, that's graphcast essentially, is it? Yeah, graphcast was one of the first models in phase three where they were taking all the weather on the earth and then simulating a forecast out to 10 days. So graphcast, the way it works is it takes the full state of the weather over the globe. And similar to other AI machine learning method, it runs the same local operation or learned function everywhere. And this roughly reflects that physics is the same everywhere. And then it processes this into a larger set of representations that cover the whole globe. And then it takes that and then predicts back down to what will happen locally again. What will happen in London? Does this make sense across the entire global at once and back down to what's going to happen in London and you carry on going up and down and up and down? That's right.

And the reason that's helpful, so numerical weather prediction doesn't do that. It just makes predictions at a very local scale. But statistical learning methods are able to learn large-scale structure of weather. So for example, a very large hurricane can span 100 or more kilometers. And there's a lot of structure there. So even the west side of a hurricane can tell you a lot about the east side of the hurricane. So by having the model able to represent the whole hurricane in one representation, it can better inform the local properties of what's going to happen next. And this is one big difference between AI weather models and traditional numerical weather prediction that is probably why the AI models are so effective. Because they can see the whole global ones. See large-scale patterns as well and learn that structure and exploit that as well as the local structure. That makes sense. At what point did these models become probabilistic? Probabilistic means probability or probably. So a deterministic model always makes exactly one

guess based on its inputs. A probabilistic model makes many guesses because this could happen or that could happen. And this allows a model to express the range of scenarios most likely. And even some of the more unlikely but still possible events. So in the field of weather forecasting, they've been moving to a probabilistic forecast for a while. So even traditional numerical weather forecasting has been using probabilistic models. And they just they roll out different scenarios that represent different possible futures. We started working on probabilistic weather models just after graphcast. So this was around 2023. And the first model we had in that was called Gencast. So it was in many ways a follow on to graphcast, it changed a bunch of things. But the biggest difference was instead of trying to predict the average weather that was expected to happen, it predicted many different scenarios that were likely to happen. And this is much more useful for extreme events.

And for any kind of decision making scenario where if you have extreme things or rare things that matter a lot, you really want to kind of know. Even if it's only a five or 10% chance of happening, you probably want to prepare for it. I mean, these are I guess in a way the sort of cones of probability that you see when you're looking at the trajectory of a hurricane. Exactly. It will probably land somewhere within these. Exactly. If you've seen the like all the little lines they call them spaghetti plots, the different scenarios. Yeah, absolutely. So how do you how do you do it then within the air models? How do you get them to be probabilistic? Because by nature, they're making a prediction at a time, aren't they? Yeah. So we have now we have two different ways of doing this. So one way of doing it is again, like many other AI techniques for video generation, we use diffusion models. And the idea there is the model it learns to take a very noisy image and turn it into something that looks like weather in this case. By starting with lots

of different noisy images, when you when you refine it to look like weather, you wind up at a slightly different scenario. And that is how you get this diversity in these different scenarios. So it's almost like we were talking earlier about the butterfly and these like little perturbations. So it's almost like you're injecting a bit of statistical perturbation and then running it forward so that you can see if anything changes. Yeah, so I think there's again, we have two techniques. So diffusion models are traditionally used. We have another technique we use, which we invented called functional generative networks. And what this does is injects different input scenarios. And it changes the actual weights in the neural network, the parameters of the neural network slightly. And this gives a diversity of outputs. So it's almost like you're saying, what would the world look like if there was a butterfly here or a butterfly here or a butterfly here, run them over and over and over and over again. And then when you've done it, hundreds, maybe thousands of times, you sort of get a sense of what the future is most likely to look like.

Yeah, the best way to just, you know, it's just these models just give a range of scenarios, each of which could have happened. And that allows people to look at the thing, they can, again, go and actually the Hurray came a list example. When you see them all sort of making the same prediction about the intensity, you can start to have a lot of confidence. That's what's going to happen. If they're spaghetti plots and they're all over the place, then the model is saying that's very difficult for us to understand what could happen. And that's a very important thing because it's a fact of how we forecast whether we cannot know certain things. So it's important for our models to recognize that and to essentially be sort of humble in a sense of not trying to over, be overly confident about that. And I guess the flip side of that is that is how you are coming up with your confidence scores for the predictions that you do make. That's right. That's right. So when they all are making the same prediction, you know that the model is, that is a very high chance that will happen. How's weather next three different to all of this then? So weather next three

differs because instead of strictly taking in the estimate of the state of the weather, this produced by a weather bureau, it also takes in raw satellite imagery. And instead of simply predicting the estimate of the global weather, it also predicts what's measured at high quality weather stations like airports. So this is one big, big feature where weather next three is moving beyond traditional numerical weather prediction to take raw data as input and predict raw data as output. So traditionally in the field of weather forecasting, there's actually a pretty mature and robust sequence of steps. So I like to use this analogy of a tree. I call it like the weather tree. The roots represent the data. So you get data from all sorts of places. And then where the roots meet the trunk, that's where we make the guess about the state of the weather at the current time. The trunk represents the operational models that predict all the weather on the earth.

And then as you go up into the branches, you start using the weather forecast for many, many different things. So you can maybe have a regional model or a model that's focused on energy forecasting or extremes. You have all sorts of different types of weather models that start to be developed by different groups for specific use cases. And at the leaves of the tree, that's like your applications. I want to know what the weather is. I pull out my phone. I want to know what's the temperature right where I am or right now. We want to change the whole value chain of weather. And we want to reshape this weather tree. So a single model takes the information from the roots and makes predictions about the leaves all at once. And so weather next is the first attempt for us to do that where we're taking in raw satellite observations. And we're also predicting raw station observations on the other side, all in a single model. Traditionally, this has been done in at least three or four different stages. Many groups are thinking about this. It's a pretty obvious idea to take more raw data and try to do the thing end to end. But weather next three is

probably the first that is doing this at a level that's significantly more accurate than any previous AI or traditional model. Is this sort of the AI version of looking out the window? In some ways, yes, it's very sophisticated. It's sophisticated. It's sophisticated. It's sophisticated. It uses more computers. But yeah, I mean, exactly, exactly, right. When you look at the window, you're looking at the sky. You kind of go, oh, well, it's like cloudy. It's probably not going to be as warm. Or I can see the trees moving. It's probably windy or it's rain. So I'm going to get wet. And so we are collapsing the many, many stages of processing into a more singular architecture and algorithm. Some other features are that it's operating at a higher resolution. So this means instead of seeing a sort of blurry or pixelated image, it's a much more refined image. In fact, it has different resolutions that you can make predictions at. And another feature is instead of making a prediction about the weather

every six hours, it makes a new prediction forecast every one hour. And the reason we can do that is because it's taking in satellite imagery that changes on an hourly basis. So by having data that's changing more frequently, we can then make forecasts that are different more frequently. And those forecasts themselves also have a higher temporal resolution. The model act natively predicts a range of different weather states across a six hour window. So it says, and we're going to one o'clock, two o'clock, three o'clock. Yeah, we're going to get this weather instead of just saying it's six o'clock, at 12 o'clock, we're going to have this weather. Which I can imagine ends up being quite useful for, I don't know, renewable energy prediction as to what you're actually going to get. Yeah. So I think a lot of the decisions that went into weather next three were made on the basis of what people have been telling us they would like to see from the next generation of weather models. So for renewable energy forecasting or load forecasting. What do we mean load forecasting? So load is, yeah, that's a technical term for electricity demand. So, you know, renewable wind power

is mostly dependent on the wind speed. The solar, you know, solar power is dependent on how sunny it is. It turns out that the amount of electricity that the power grid needs to provide tends to be highly influenced by the temperature. Because people turn on the heat when it's cold, they turn on their air conditioners when it's hot. So you get peaks in the hottest parts of the summer and the coldest parts of the winter in electrical demand. So if we can predict the wet the temperature and also the humidity more accurately, we can make better predictions of the grids, the demand for electricity. And then we also, if we can make better wind and solar predictions, we can make better predictions about the supply of electricity. So weather next three also offers new wind and solar variables that we hadn't been providing before. And that should help with the generation and also with the higher resolution surface temperature, we should be able to help with the electrical demand forecasting. I'm thinking about the impact on agriculture here as well. I mean, this is,

they care a lot about weather. I think, I'm not sure this is true, but I think farmers are probably the oldest category of weather forecasters because they, like timing when to plant and when to harvest, is like a completely determined based on the temperature and the rain. So if you could start to forecast even like weekly rainfall or average temperature over a month, things like this, that can inform agricultural decision making. You can also plant different types of crops, drought resistant seeds, rain resistant seeds. As AI weather models start to bridge out into these impact areas, we can think about making much longer term forecast with people starting to study and also trying to model more directly these agricultural decisions. What how much rainfall is too much rainfall that we'd want to then prepare for, how much is sort of normal and it's okay to just sort of ignore. Has this required a complete overall of the architecture that went before or are you able to just sort of plug in the data into what was there already? Some parts of

the innovations in weather next three were relatively straightforward. So adding in solar variables and wind variables was relatively straightforward. It wasn't trivial because these variables have different characteristics that have to be accounted for, but making predictions about finer temporal resolutions or taking in satellite imagery or predicting stations were much more profound dramatic changes to the architecture of the model and the training process that we use to actually train it with data. So let me ask about their performance then. How do they compare to the sort of more traditional numerical processes? Yeah, I mean these days the AI weather models are considerably more accurate. For us, graphcast I believe was the first AI weather model to be more accurate than the traditional deterministic forecast. Gencast was the first that was significantly more accurate than the best probabilistic forecast from numerical weather prediction. Our new models are building on the performances of our old models. There's other, again there's many other

groups who are doing great work. I think the models we're building tend to be more accurate than those, but it's also hard to say because there's different variables, there's temperature, there's wind, there's different forecast horizons in the far future, the near term, there's lots of regions. So you can't always say one model is uniformly better, even in tropical cyclones. Historically they usually relied on two different types of models to predict the track and the intensity. It turned out that predicting the track was better for models that were sort of coarse grain and in large scale. The intensity was better for models that had very fine grain that kind of highly resolved to local features. Our model was the best at track and intensity and it's a single model. Even then though you say, well you have to look at, is it better at most intense storms, is it better at the Pacific, is it better at the Atlantic? It's very important to not make a single claim about this because it actually matters for decision makers. If you say this is the best model

and then well in this region it's not, you don't want people to rely on it as much. So there's a lot, there's a lot, I just the way I would say is we're really on a journey. It's hard to say that we're better than anyone uniformly than models, but when we evaluate we tend to outperform the competitors at the, you know, the ones we're comparing. What about extreme weather events? The air model is as good, just the traditional ones. Our models appear to be best on tropical cyclones. That's some of the most extreme weather on earth. There's been less work on things like extreme heat, extreme cold, even fine grain like tornadoes and things like this. There's less work that I'm aware of, but I think all of this stuff is going to, when you apply these same methods and using high quality data sets, you're going to, the next few years, you're going to see mostly these phenomena be best predicted by AI models. Here's one question about the fact that you are using the current time step to predict the next time step. If you end up getting a prediction slightly wrong, if there's like a little error in there, can those errors not compound over time?

So that that has to be more a problem of traditional methods. There's nothing to kind of bring them back to what's most plausible. So it's just using the equations. Yeah, once it just starts to go off into something that's rare, there's nothing about the equation that says it must be like normal weather. With AI models, their failure mode tends to be, when they start to make errors, they tend to more just predict average weather. Like predicting the average weather is actually a pretty good way to predict weather. And it's actually the oldest way to predict weather. So if you're like a farmer and you're trying to figure out like, you know, when do I plant my seeds, you use like the season in the spring. Before it gets warm, when we're going to have a lot of, you know, sunlight. So we see that statistical models tend to, when they start to generate errors, they just tend to predict, you know, they regress to the mean. It's called they regress to the average weather. So they don't really blow up or explode very much like we see in traditional numerical methods. But then I also wonder about unusual weather events, you know, that we're getting more and more up with climate change. It's sort of statistical outliers as it were. So those two

things feel almost contradictory that it tends towards the average weather, but can also do even when events you haven't seen before. It's exactly. It's a strange phenomena that these statistical learning methods can also predict the strongest hurricanes on that have ever been recorded. Right. The reason we think this is possible is because the weather is, it's really a mosaic of lots of local weather. So while the model might not have seen this specific instance of this storm at this location, this trajectory, it has seen intense weather in other parts of the globe, in other, you know, smaller scale things like storms and things like this. So the models learn the statistics, the local statistics, and then they can form a sort of novel new extreme forecast because they've seen the parts. And that's the best that we are understanding of why this works. But I think also, this is an area that probably needs to be explored better as well. So as we try to get into regimes where we want to predict like the monsoon or things that are just, they only happen once

a year. We just don't really actually have many examples of this. We want to make sure that we're also innovating on the efficiency of the learning that it can take not that many examples and still make use of them in a way that's sort of unbiased and accurate. Because even if you've got a few decades weather data, it's still not that many of them. For the point. Right. Right. What's next for weather next then? What improvements would you like to see in the future? So there's so many applications of weather forecasting. Like whether the estimate weather itself affects like a third of the economy. Now, it's not that a better weather forecast can help a third of the economy necessarily, but it makes contact with all sorts of things. Energy, agriculture, insurance, transportation, you know, I've again disasters. All of those things need very specific information about what's going to happen. And also there's a lot of applications, whether like an energy, especially where they're also collecting a lot of weather observations themselves. And we could potentially incorporate these into our weather models. And we

haven't done any of this yet. So the way I think about it is we've kind of made a significant advance on just forecasting the global weather. But now as we go into the many, many different branches of whether there's just zillions of problems that can be tackled. And you can bring in data from the problem itself. Much higher resolution. I mean, you often want to you see weather patterns that are like happening on a less than a kilometer scale. And there's all sorts of data sources that can be brought in as well. I mean, most weather observations are actually not even used by weather agencies. So there's, I actually think there's a ton of data. There's a ton of problems that haven't even touched yet. This is really just kind of sort of demonstrated that we can make a very accurate forecast. And now it's time to really like get to work, I think. Some people sort of, I think imagine that there will be a point where the AI stuff just replaces the existing numerical methods, the existing supercomputers and so on and they've got running. Do you think that will be the case? Do you think we'll move away from physics-based equations and those kind of

simulations? Yes. So I think we will move away from approximating the solutions to equations. But we're not moving away from physics. The data was also created by physics. We're just going to approximate the physics in a different way. And I actually think that we're probably going to open a lot more applications of weather forecasting. My hope is that this will actually create a lot more opportunities and value and sort of cottage industries and things like this as it's possible to make forecasts that are more accurate and also make forecasts that are more directly tied to the impact of weather. So instead of forecasting the wind, maybe you forecast the damage that the wind will cause to the power lines. If we can get to a regime where we're starting to forecast the impact, I think the sky's the limit. It's just so interesting that this idea that you're still not quite sure why it's able to do it so well. Are you working on interpretability too?

To some extent, the lesson that I've learned about AI and machine learning is that ultimately it's the evaluation that determines the truth and the quality of the model. So while it's very helpful, I think to understand how the model is making predictions, why it makes errors and things like this. It can also help folks who want to understand why the weather is, they maybe are meteorologists and they actually have a lot of knowledge about how the weather works. So if the model can help, can speak to them in that language, then they can make more use of it. But at the end of the day, I think if we want more accurate models, we need to be very diligent about keeping our focus on the evaluation. We just expand our evaluation beyond the temperature or the wind to, again, these many more variables and many more observations. That's what's most important. My final question. I think the public is really focused on large language models at the moment, right? And to some extent, the forgotten that there are all these other phenomenal research projects that are going on all over the place. Do you think there's things that the work on weather prediction

has learned that can be applied to other applications? Yeah, I mean, our weather forecasts are much bigger than the videos that we generate. The scale that we're operating at is enormous. The data is enormous. These are gigabytes of data, many gigabytes of data that are going in and out of the models. Even the technique we've used, we've invented for the probabilistic forecasting, the functional generative network is a new technique of it's much more efficient to train and turn out to be much more effective for us than diffusion models. So we're also looking at that, how to incorporate some of the technical advances back into video modeling, probabilistic modeling that can span into other areas. By the same token, I want for us to also think about how to take LLMs and draw in insights from them. And you can start to think about maybe the weather model takes in text data. Why not? People are talking about the weather. Why doesn't operate on that? Maybe you can interrogate the model with an LLM the same way that you can interrogate a video model with an LLM and ask a question about what's happening in the video. Maybe you can say

what's going to, is it going to rain on that mountain top or in that valley? So we haven't even really scratched the surface of this, but I think there's a nice two-way interaction between the innovations we're making and feeding those back into video modeling and large-scale models and also taking the insights from language and these models take very diverse information. So much exciting stuff going on. Peter, thank you so much. That's fun to see you. Thank you. Weather prediction is one of the oldest puzzles we have. For most of human history, all we had to go on were signs like whether the sky was red at dusk or if you got a bad knee before the rain. And about a hundred years ago, we started calculating the future using the best equations we had at our disposal. And now, this is one place where AI is out there in the real world making a gigantic difference already, buying ourselves a bit more warning and extra hour and extra day. But on the right day, that makes all the difference. You have been listening to Google DeepMind, the podcast,

with me Professor Hannah Frye. We've got lots more coming up in this series. So make sure you subscribe and we'll see you next time.

More episodes

More from Google DeepMind: The Podcast

View all episodes →