Showing posts with label regression towards the mean. Show all posts
Showing posts with label regression towards the mean. Show all posts

Friday, December 30, 2022

149: Portugal, Spain and France - a comparison

In Post 138 I compared the temperature trends of the Scandinavian countries to see it there were any similarities. There were. In fact there was almost perfect agreement between the 5-year average trends of Norway, Sweden and Finland as far back as 1900 (see Fig. 138.3). As both Norway and Sweden had about twenty sets of station data that went back to 1900 and Finland had about ten, this demonstrated that averaging over a large number of independent data sets eliminates most measurements errors: a consequence of regression towards the mean.

In Post 144 I repeated this procedure for trends from Ireland, Scotland and England and obtained a similar result (see Fig. 144.3), although the trend for England differed slightly from the other two due to its greater urbanization. This demonstrates that neighbouring regions should have similar climates, or at least they should experience similar changes to their climates. So is this also the case for Portugal, Spain and France? I ask this because the results from Post 146 suggest that before 1980 the climate of Spain was cooling while Post 145 suggests that that of Portugal was warming. Well, the results in Fig. 149.1 below show that in fact the temperature trends of Spain and Portugal are very well correlated as far back as 1940, then they diverge. France, on the other hand, is only very weakly correlated to both Spain and Portugal.


Fig. 149.1: A comparison of the 5-year average temperature trends since 1800 for Portugal (green), Spain (red) and France (blue). The two upper trends are offset by +2°C for clarity and the bottom two trends are offset by -2°C.


This discrepancy can be explained in part by the number of stations contributing to the mean temperature anomaly (MTA) of each country per month (see Fig. 149.2 below). Before 1940 there are only two stations contributing to the Portugal MTA, which is probably why it diverges from the Spain MTA which consistently has over ten contributing stations. However, this cannot fully explain the poor correlation of the French data to that of either Spain or Portugal, even though France also has a low number of stations before 1940. The issue here is that the France MTA has a high number of contributing stations after 1960, as do Spain and Portugal, and yet its correlation to both of their MTAs is still poor after 1960. That said, its overall trend since 1860 does follow that of Portugal quite closely.


Fig. 149.2: The number of station records included each month in the averaging for the mean temperature trends of each country in Fig. 149.1.


It should be remembered, though, that the MTAs of both Portugal and France before 1940 are strongly dependent on only two or three sets of station data, and in both cases most of these stations are located in the biggest cities: Paris, Marseille, Lisbon and Porto. These four stations also all appear to exhibit severe continuous warming since 1900 consistent with the effect of urban heat islands. In which case the similarity between the MTA trends of Portugal and France before 1940 may simply be a consequence of parallel economic development in their largest cities.

Instead these comparisons suggest that France may actually have a completely different climate to the Iberian Peninsula even though it is its closest neighbour. The reason for this may be down to geography and the influence of the Pyrenees mountain range at the border that effectively insulates one region from the other.


Tuesday, December 13, 2022

144: Evidence against temperature adjustments #4 (British Isles)

In the previous four posts I examined the temperature changes for Ireland (see Post 140), Scotland (see Post 142), England (see Post 143) and Great Britain (see Post 141). While all four sets of temperature data appeared similar from 1900 onwards, there were some differences, and these differences were most apparent in a comparison of the earlier data for Ireland and Great Britain. When the Great Britain data was separated into different trends for Scotland and England a similar degree of difference was observed with the Scotland data appearing to correlate more closely with Ireland, and England with Great Britain. In this post I will look to show this pictorially by comparing the various trends directly.

First, if we compare the data for Ireland, Scotland and England with Great Britain we see that England shows the closest agreement after 1900 but Scotland shows the better agreement before 1840 (see Fig. 144.1 below). The data depicted here are the 5-year moving averages of the mean temperature anomalies (MTAs) for each country as shown by the yellow curves in Fig. 140.2, Fig. 141.2, Fig. 142.2 and Fig. 143.2 in previous posts.


Fig. 144.1: The 5-year average temperature trends since 1760 for Ireland, Scotland and England each compared to that of Great Britain. For clarity the trends for Ireland and England are offset by +2°C and -1.5°C respectively.


What is striking about the trends in Fig. 144.1 is how similar they all are after 1860, while the greatest disparities occur before 1860. The reason for this is evident from Fig. 144.2 below which shows that the number of stations used to calculate each of the MTA for Ireland, Scotland and England drops below five before 1870. From this we can conclude two things. First, this suggests that if there are too few stations used in determining the MTA the accuracy decreases. Secondly we see that when there are sufficient stations used to determine the MTA the accuracy is so good that there is little difference between the MTA for different neighbouring countries. 

This is not the first time such conclusions have been drawn. The same effects were seen in Post 138 (Evidence against temperature adjustments #3) comparing trends in the different Scandinavian countries and Post 57 (The case against temperature data adjustments #1) comparing them in various central European countries. In all cases the conclusion is the same. If trends for neighbouring countries agree, then they are likely to all be correct, not all equally incorrect. Therefore no adjustments to the temperature data are needed or justified. A similar result is also encountered when comparing random samples of stations from the same region as was shown for the USA in Post 67 (More evidence against temperature data adjustments #2). The reason for this is that averaging a sufficiently large number of independent data sets results in a reduction in the size of the errors imported from each. This is known as regression towards the mean.


Fig. 144.2: The number of station records included each month in the averaging for the mean temperature trends in Fig. 144.1.


The second comparison I have performed is to compare data for Ireland, Scotland and England with each other. This is shown in Fig. 144.3 below. Now we see that the two countries that agree most closely are Scotland and Ireland while the data for England appears to exhibit more warming after 1980 and before 1900. This additional warming could be in excess of 0.5°C since 1840.


Fig. 144.3: Comparisons of the 5-year average temperature trends since 1760 for England and Scotland (two top curves, both offset by +2°C), Scotland and Ireland (two middle curves), and Ireland and England (two bottom curves, both offset by -2°C).


Conclusions

Once again a comparison of temperature data for neighbouring countries indicates that most adjustments to the data are unnecessary as the averaging process will correct for most errors via regression towards the mean.

The data for Scotland and Ireland are in closest agreement, probably because both have similar population densities and are more rural.

The data for England is in closest agreement with that of Great Britain, probably because England is the largest country in Great Britain and so its stations will always make the dominant contribution compared to other countries such as Scotland or Wales. 

The greater warming seen in England (of over 0.5°C) is further evidence that warming within countries is driven not just by carbon dioxide levels in the atmosphere and the greenhouse effect, but by local energy consumption as well. So net-zero will not be a panacea.


Saturday, September 24, 2022

138: Evidence against temperature adjustments #3 (Scandinavia)

One of the main aims of this blog has been to investigate the extent to which the various datasets in the global temperature record have been adjusted and to ascertain both the impact of these adjustments and their validity. Most of the blog posts for individual countries or territories have sought to quantify the magnitude of these adjustments by calculating two versions of the mean temperature anomaly (MTA) for each region; one based on its raw unadjusted data and a second using Berkeley Earth adjusted data. Then the two are compared and the difference calculated. This difference is often considerable and often shows that the adjustments have increased the amount of reported warming. But I have also investigated the second issue, that of validity. One way to do this is to compare the MTA for neighbouring regions or different data samples from the same region. 

The rationale is as follows. If there are errors in the data that are sufficient to affect the MTA, then comparing MTAs from different samples from the same region, or samples from adjacent regions that would be expected to be almost identical, could highlight the errors. Of course any difference between MTAs from different regions does not prove that the data is wrong; it may be that the regions aren't as similar as one supposed. But if the data is virtually identical then that does suggest both that the temperature trends for the two samples or regions are behaving the same, and that any data errors in the temperature datasets (which are likely to be numerous) are not significant and so are not in need of correction or adjustment.

In Post 57 I used this approach to compare the temperature trends of neighbouring countries in central Europe (Germany, Czechoslovakia, Austria and Hungary). The results showed that if the MTA for a country was determined using data from more than about fifteen different station records then there was little difference between MTAs for different countries, and thus very little error in the MTA of each country. This is because of a property of statistics called regression towards the mean. This basically states that if any dataset contains errors in its measurements (which most data does), and those errors are random in their size and distribution (which they often are), then the errors will tend to cancel each other when you average the data. Moreover, the more data you average, the greater the cancellation of errors and so the more accurate will be the result. If errors don't cancel, then that is because the errors are systematic not random, so the process also helps to identify these as well.

In Post 67 I repeated this process for temperature data from the USA. In this case instead of comparing data from adjacent regions I compared different samples of one hundred stations from the same region: the entire contiguous United States. The result was the same as in Post 57 with each sample exhibiting an identical temperature trend over time with identical fluctuations in the 5-year moving average of the trend.

In this post I will repeat the country comparison of Post 57 but using the 5-year moving average of the temperature trend data from the four neighbouring Scandinavian countries of Norway, Sweden, Finland and Denmark. These trends were determined in Post 135, Post 136, Post 137 and Post 48 respectively. The results are shown in Fig. 139.1 below.


Fig. 138.1: A comparison of the 5-year average temperature trends since 1700 for Norway, Finland and Denmark compared to that of Sweden. The trends for Finland and Norway are offset by ±3°C for clarity.


In Fig. 139.1 I have compared the trends of Norway, Finland and Denmark with that of Sweden. The reasons for choosing Sweden as the comparator were both geographic and practical. It sits between the other three countries and so is a near neighbour for each (Finland and Denmark are not near neighbours so would not be good comparators). But it also has the most stations of the four countries and so should have the most reliable trend.

The data in Fig. 139.1 clearly shows that the trends for all four countries are very similar after 1900 but diverge as one looks further back in time towards 1800. The reason for this is the reduction in station numbers seen in each country as one moves back in time from 1950 (see Fig. 138.2 below). Given that it seems that somewhere between ten and thirty stations are needed in the MTA average in order for the errors to be minimized, we can see from Fig. 138.2 that this condition is satisfied for all four countries after 1890. That is why the MTAs diverge before 1890 but are very similar after that date.


Fig. 138.2: The number of station records included each month in the averaging for the mean temperature trends in Fig. 138.1.


If we just consider the data after 1850 we see that the agreement between trends for the different countries is remarkably good after 1890 (see Fig. 138.3 below). The agreement between Norway and Sweden, and Finland and Sweden are both particularly good to the point of their three trends being almost identical. There is also excellent agreement between Denmark's trend and that of Sweden after 1980 but less so before. This is probably the result of Denmark not only having much fewer stations than the other three countries, but also having fewer than ten stations before 1975.


Fig. 138.3: A comparison of the 5-year average temperature trends since 1850 for Norway, Finland and Denmark compared to that of Sweden. The trends for Finland and Norway are offset by ±3°C for clarity.


Summary

The data in Fig. 138.3 once again demonstrates the futility of temperature adjustments. The fact that the mean temperature anomalies (MTAs) of Norway, Sweden and Finland agree so well for over 120 years from 1890 onwards without data adjustments indicates that the averaging process alone can eliminate most errors.

The Denmark data also adds weight to the conjecture that between ten and thirty stations are needed in the average in order to eliminate most of the data errors. As the error size decreases with the square root of the sample size, an average of 25 datasets should decrease the error size by 80% (reducing each error to a fifth of its nominal value). 

Comparing the data of these four countries in this way also gives us more confidence in the determining the true nature of the regional temperature trend. All the data after 1900 pretty much agree so we can conclude that temperatures from 1900 to 1980 rose marginally by less than 0.3°C and then jumped by about 1°C in the 1980s. But this jump is still only comparable to the size of the fluctuations in the 5-year average.

From 1850 to 1900 both Denmark and Norway diverge from Sweden slightly but in different directions. But this is based on a comparison of only one or two stations in each case and so is not unexpected.


Friday, May 14, 2021

67. More evidence against temperature data adjustments (USA)

In Post 57 (The case against temperature data adjustments) I presented evidence that seemed to cast doubt on the need to adjust temperature data. The main argument from climate scientists in favour of these adjustments is their belief that the raw data cannot always be trusted. Over time, changes to the data collection process may occur. These changes may be due to changes in the location of the weather station, changes to the environment around the original site, or changes to the instrumentation or data collection methods. 

It is certainly true that these issues affect many, if not most temperature records, and the longer the temperature record, the more likely such issues will probably occur. The important questions, though, are: how large are these data errors, and what is the best way to eliminate them from historical data?


Fig. 67.1: The temperature anomalies for Baker City Municipal Airport as calculated by Berkeley Earth.


The approach most climate groups use is to adjust data within each individual temperature record, an example of which can be seen by comparing the data in Fig. 67.1 above and Fig. 67.2 below. Both graphs show temperature data from Baker City Municipal Airport (Berkeley Earth ID: 164703) in the state of Oregon in the USA, which has then been adjusted by Berkeley Earth. The original data in Fig. 67.1 has no discernible temperature trend (green line), but after the data has been chopped into multiple autonomous segments, and those segments each subjected to its own separate corrective bias, the overall trend becomes strongly positive with a gradient of +0.67°C per century. Thus warming appears where before there was none.


Fig. 67.2: The adjustments made by Berkeley Earth to the temperature anomalies for Baker City Municipal Airport.


The justification for using these adjustments is that climate scientists believe they can identify points in the data where errors have been introduced, and also that they can determine what the correction factor needs to be in order to eradicate the error. The size of the adjustments is usually determined by comparing the station temperature time series with that of its neighbours, but as I showed in Post 43, even identical neighbours can display temperature differences of up to ±0.25°C just due to measurement uncertainties.

Adjustments are most commonly made at positions in the time series corresponding to known or documented station moves (red diamonds in Fig. 67.2), or at points where there is a gap in the temperature record (green diamonds). But, groups such as NOAA and Berkeley Earth have also developed algorithms that they claim can identify other points in the time series where undocumented changes have occurred. These positions in the data are referred to by NOAA and Berkeley Earth as changepoints and breakpoints respectively. To many climate sceptics, however, these techniques remain controversial. But I would argue that in many cases they are also unnecessary because of a statistical phenomenon called regression towards the mean.

 

Fig. 67.3: The average temperature trend for the 100 longest temperature records in the USA. The best fit is applied to the monthly mean data from 1921 to 2010 and has a positive gradient of +0.25 ± 0.15 °C per century. The monthly temperature changes are defined relative to the 1951-1980 monthly averages.

 

Basically, if the errors are randomly distributed between different records, and also at different times in those records, and if they are of comparable size, then any averaging process will cause the errors to partially cancel. The bigger the number of records in the average, the more precisely they will cancel.

In Post 57 I demonstrated that these errors can be eradicated using a simple averaging process. I did this by averaging unadjusted temperature data from stations located in neighbouring European countries (Germany, Austria, Hungary and Czechoslovakia), and showing that the averaging process gave the same result for the 5-year average trend for each country, provided there were more than about twenty stations in the average for each country. This was despite the fact that Berkeley Earth had applied over three adjustments on average to each temperature record during its own analysis process for those same stations.

 

Fig. 67.4: The average temperature trend for the 101st to the 200th longest temperature records in the USA. The best fit is applied to the monthly mean data from 1921 to 2010 and has a negative gradient of -0.11 ± 0.15 °C per century. The monthly temperature changes are defined relative to the 1951-1980 monthly averages.

 

The key to this is having sufficient data. In the case of the USA we have more than sufficient data. In Post 66 I analysed the 400 longest temperature records for the USA and determined the temperature trend since 1750. This in turn showed no evidence of any global warming in the USA over the last 100 years. But suppose we split those 400 records into four sets of 100 records, and compare the four results for the different mean temperature trends. What would we expect to see?


Fig. 67.5: The average temperature trend for the 201st to the 300th longest temperature records in the USA. The best fit is applied to the monthly mean data from 1921 to 2010 and has a slight negative gradient of -0.003 ± 0.135 °C per century. The monthly temperature changes are defined relative to the 1951-1980 monthly averages.

 

Well, the answer is shown in Fig. 67.3-Fig. 67.6. The result is that the four temperature trends look very similar (it is probably easiest to compare the 5-year moving average curves). But judging by the adjustments made by Berkeley Earth to the data in Fig. 67.2, it would not be unreasonable to expect Berkeley Earth to have made over 1000 adjustments in total to the 100 station records used to generate each of these four temperature trends. So not making these 1000 adjustments should result in large discrepancies between the four different trends, assuming the adjustments are needed. But they aren't needed, and there are no large discrepancies.

 

Fig. 67.6: The average temperature trend for the 301st to the 400th longest temperature records in the USA. The best fit is applied to the monthly mean data from 1921 to 2010 and has a negative gradient of -0.22 ± 0.14 °C per century. The monthly temperature changes are defined relative to the 1951-1980 monthly averages.

 

The four trends are compared in more detail in Fig. 67.7 below. The average of the anomalies from the 100 longest records in the USA are shown in yellow and are offset by -1°C for clarity. The mean of next longest 100 records is shown in blue. The mean of the third longest set is shown in red and offset by +1°C, with the fourth longest set shown in black, but not offset. To aid the analysis process, the blue curve is plotted three times, with three different offsets so that it can be compared with the other three trends.

 

Fig. 67.7:  A comparison of the 5-year averaged temperature trends for four sets of 100 temperature records in the USA. The trends are offset for clarity with the trend for stations 101-200 used as a comparator for each of the other three trends.

 

What is clear is that for all the data from 1890 onwards the four trends are virtually identical. This implies that the averaging process has eliminated almost all the data errors. Before 1890 the number of stations in each average decreases dramatically as shown in Fig. 67.8 below, which is why the level of agreement between the curves is much lower. From 1900 onwards, however, there is almost total agreement. The only significant differences are for stations 001-100 in the 1930s, and stations 301-400 post-1995, but in both cases the discrepancy is generally less than 0.25°C. These differences also largely account for the differences in the best fit lines in Fig. 67.3-Fig. 67.6. As for their causes, well the lower temperature anomaly for stations 301-400 after 1995 could be due to these stations being newer than the rest. That would imply they are located in smaller towns with less waste heat production. For stations 001-100 the opposite is probably true as these station time series are the longest. That in turn suggests they are more likely to be located near the largest cities.


Fig. 67.8: The number of station records included each month in the mean temperature trends.


What this demonstrates unequivocally is that data adjustments are unnecessary when determining global or regional mean trends. This is because the errors in the individual station records will cancel when averaged. If they did not, then the four trends in Fig. 67.7 would not be so alike. Instead there would be significant differences. And remember, the same result was demonstrated in Post 57.

But what this also implies is that, if the errors in the individual station records cancel, then so too should the adjustments that are applied by Berkeley Earth and others to correct these errors. Except they don't.

 

Fig. 67.9: The average temperature trend for the 301st to the 400th longest temperature records in the USA after adjustments made by Berkeley Earth. The best fit is applied to the monthly mean data from 1921 to 2010 and has a positive gradient of +0.54 ± 0.05 °C per century.

 

The graph in Fig. 67.9 above shows the mean temperature trend for stations 301 to 400 with their Berkeley Earth adjustments included. If the adjustments cancelled, then the graph should resemble the data in Fig. 67.6, but it doesn't. Instead of a significant negative trend, there is a sizeable positive trend. In fact the adjustments have added a net warming of over +0.7°C to the data since 1920. The same positive trend is also seen for the means of the adjusted data for the other three sets of 100 stations, so at least they are consistent, but that does not mean they are correct. In fact all they are doing here is adding warming where none existed previously.

 

Summary

What I have presented here is yet more compelling evidence against the statistical validity of temperature adjustments.

I have shown that the true temperature trend can be determined simply by averaging the anomalies from the raw data. This confirms the similar result for Central European data that I presented in Post 57.

This adds further weight to my claim in Post 66 that there has been no global warming in the USA since 1900.