May 19, 2008
Do IPCC Temperature Forecasts Have Skill?
[UPDATE] Roger Pielke, Sr. tells us that we are barking up the wrong tree looking at surface temperatures anyway. He says that the real action is in looking at oceanic heat content, for which predictions have far less variability over short terms than do surface temperatures. And he says that observations of accumulated heat content over the past 4 years “are not even close” to the model predictions. For the details, please see for your self at his site.]
“Skill” is a technical term in the forecast verification literature that means the ability to beat a naïve baseline when making forecasts. If your forecasting methodology can’t beat some simple heuristic, then it will likely be of little use.
What are examples of such naïve baselines? In weather forecasting historical climatology is often used. So if the average temperature in Boulder for May 20 is 75 degrees, and my prediction is for 85 degree, then any observed temperature below 80 degrees will mean that my forecast had no skill. In the mutual fund industry stock indexes are examples of naive baselines used to evaluate performance of fund managers. Of course, no forecasting method can always show skill in every forecast, so the appropriate metric is the degree of skill present in your forecasts. Like many other aspects of forecast verification, skill is a matter of degree, and is not black or white.
Skill is preferred to “consistency” if only because the addition of bad forecasts to a forecasting ensemble does not improve skill unless it improves forecast accuracy, which is not the case with certain measures of “consistency,” as we have seen. Skill also provides a clear metric of success for forecasts, once a naïve baseline is agreed upon. As time goes on, forecasts such as those issued by the IPCC should tend toward increasing skill, as the gap between a naive forecast and a prediction grows. If a forecasting methodology shows no skill then it would be appropriate to question the usefulness and/or accuracy of the forecasting methodology.
In this post I use the IPCC forecasts of 1990, 2001, and 2007 to illustrate the concept of skill, and to explain why it is a much better metric that “consistency” to evaluate forecasts of the IPCC.
The first task is to choose a naïve baseline. This choice is subjective and people often argue over it. People making forecasts usually want a baseline that is easy to beat, people using or paying for forecasts often want a more rigorous baseline. For this exercise I will use the observed temperature trend over the 100 years ending in 2005, as reported by the 2007 IPCC, which is 0.076 degrees per decade. So in this exercise the baseline that the IPCC forecasts have to beat is a naïve assumption that future temperature increases will increase by the same rate as has been observed over the past 100 years. Obviously, one could argue for a different naïve baseline, but this is the one I’ve chosen to use.
I will also use the ensemble average “best guess” from the IPCC for the most appropriate emissions scenario as the prediction. And for observations I will use the average value from the four main group tracking global temperature trends. These choices could be made differently, and a more comprehensive analysis would explore different ways to do the analysis.
So then, using these metrics how does the IPCC 1990 best estimate forecast for future increases in temperature compare for 1990-2007? The figure below shows that the IPCC forecast, while over-predicting the observed trend, outperformed this naïve baseline. So the forecast can be claimed to be skillful, but not by very much.

A more definitive example of a skillful forecast is the 2001 IPCC prediction, which the following figure shows demonstrated a high degree of skill.

Similarly, the 2000-2007 forecast of the IPCC 2007 also shows a high degree of skill, as seen in the next figure.

But in 2008 things get interesting. With data from 2008 included, rather than ending in 2007, then the 2007 IPCC forecast is no longer skillful, as shown below.

If one starts the IPCC predictions in 2001, then the lack of skill is even greater, as seen below.

What does all of this mean for the ability of the IPCC to predict longer-term climate change? Perhaps nothing, as many scientists would claim that it makes no sense to discuss IPCC predictions on time scales less than 20 or 30 years. If so, then it would also be inappropriate to claim that IPCC forecasts on the shorter scales are skillful or accurate. One way to interpret the recent Keenlyside et al. paper in Nature is that their analysis suggests that the IPCC predictions of future temperature evolution won’t be skillful unless they account for various factors not included in the IPCC predictions.
The point of this exercise is to show that there are simple, unambiguous alternatives to using the notion of “consistency” as the basis for comparing IPCC forecasts with observations. “Consistency” between models and observations is a misleading, and I would say fairly useless way to talk about climate forecasts. Measures of skill provide an unambiguous way to evaluate how the IPCC is doing over time.
But make no mistake, the longer the IPCC forecasts lie in a zone of “no skill” — which the most recent ones (2007) currently do (for the time of the forecast to present) — the more interest they will receive. This time period may be for only one more month, or perhaps many years. I don’t know. This situation creates interesting incentives for forecasters who want their predictions to show skill.