Showing posts with label measurement error. Show all posts
Showing posts with label measurement error. Show all posts

Saturday, September 24, 2022

138: Evidence against temperature adjustments #3 (Scandinavia)

One of the main aims of this blog has been to investigate the extent to which the various datasets in the global temperature record have been adjusted and to ascertain both the impact of these adjustments and their validity. Most of the blog posts for individual countries or territories have sought to quantify the magnitude of these adjustments by calculating two versions of the mean temperature anomaly (MTA) for each region; one based on its raw unadjusted data and a second using Berkeley Earth adjusted data. Then the two are compared and the difference calculated. This difference is often considerable and often shows that the adjustments have increased the amount of reported warming. But I have also investigated the second issue, that of validity. One way to do this is to compare the MTA for neighbouring regions or different data samples from the same region. 

The rationale is as follows. If there are errors in the data that are sufficient to affect the MTA, then comparing MTAs from different samples from the same region, or samples from adjacent regions that would be expected to be almost identical, could highlight the errors. Of course any difference between MTAs from different regions does not prove that the data is wrong; it may be that the regions aren't as similar as one supposed. But if the data is virtually identical then that does suggest both that the temperature trends for the two samples or regions are behaving the same, and that any data errors in the temperature datasets (which are likely to be numerous) are not significant and so are not in need of correction or adjustment.

In Post 57 I used this approach to compare the temperature trends of neighbouring countries in central Europe (Germany, Czechoslovakia, Austria and Hungary). The results showed that if the MTA for a country was determined using data from more than about fifteen different station records then there was little difference between MTAs for different countries, and thus very little error in the MTA of each country. This is because of a property of statistics called regression towards the mean. This basically states that if any dataset contains errors in its measurements (which most data does), and those errors are random in their size and distribution (which they often are), then the errors will tend to cancel each other when you average the data. Moreover, the more data you average, the greater the cancellation of errors and so the more accurate will be the result. If errors don't cancel, then that is because the errors are systematic not random, so the process also helps to identify these as well.

In Post 67 I repeated this process for temperature data from the USA. In this case instead of comparing data from adjacent regions I compared different samples of one hundred stations from the same region: the entire contiguous United States. The result was the same as in Post 57 with each sample exhibiting an identical temperature trend over time with identical fluctuations in the 5-year moving average of the trend.

In this post I will repeat the country comparison of Post 57 but using the 5-year moving average of the temperature trend data from the four neighbouring Scandinavian countries of Norway, Sweden, Finland and Denmark. These trends were determined in Post 135, Post 136, Post 137 and Post 48 respectively. The results are shown in Fig. 139.1 below.


Fig. 138.1: A comparison of the 5-year average temperature trends since 1700 for Norway, Finland and Denmark compared to that of Sweden. The trends for Finland and Norway are offset by ±3°C for clarity.


In Fig. 139.1 I have compared the trends of Norway, Finland and Denmark with that of Sweden. The reasons for choosing Sweden as the comparator were both geographic and practical. It sits between the other three countries and so is a near neighbour for each (Finland and Denmark are not near neighbours so would not be good comparators). But it also has the most stations of the four countries and so should have the most reliable trend.

The data in Fig. 139.1 clearly shows that the trends for all four countries are very similar after 1900 but diverge as one looks further back in time towards 1800. The reason for this is the reduction in station numbers seen in each country as one moves back in time from 1950 (see Fig. 138.2 below). Given that it seems that somewhere between ten and thirty stations are needed in the MTA average in order for the errors to be minimized, we can see from Fig. 138.2 that this condition is satisfied for all four countries after 1890. That is why the MTAs diverge before 1890 but are very similar after that date.


Fig. 138.2: The number of station records included each month in the averaging for the mean temperature trends in Fig. 138.1.


If we just consider the data after 1850 we see that the agreement between trends for the different countries is remarkably good after 1890 (see Fig. 138.3 below). The agreement between Norway and Sweden, and Finland and Sweden are both particularly good to the point of their three trends being almost identical. There is also excellent agreement between Denmark's trend and that of Sweden after 1980 but less so before. This is probably the result of Denmark not only having much fewer stations than the other three countries, but also having fewer than ten stations before 1975.


Fig. 138.3: A comparison of the 5-year average temperature trends since 1850 for Norway, Finland and Denmark compared to that of Sweden. The trends for Finland and Norway are offset by ±3°C for clarity.


Summary

The data in Fig. 138.3 once again demonstrates the futility of temperature adjustments. The fact that the mean temperature anomalies (MTAs) of Norway, Sweden and Finland agree so well for over 120 years from 1890 onwards without data adjustments indicates that the averaging process alone can eliminate most errors.

The Denmark data also adds weight to the conjecture that between ten and thirty stations are needed in the average in order to eliminate most of the data errors. As the error size decreases with the square root of the sample size, an average of 25 datasets should decrease the error size by 80% (reducing each error to a fifth of its nominal value). 

Comparing the data of these four countries in this way also gives us more confidence in the determining the true nature of the regional temperature trend. All the data after 1900 pretty much agree so we can conclude that temperatures from 1900 to 1980 rose marginally by less than 0.3°C and then jumped by about 1°C in the 1980s. But this jump is still only comparable to the size of the fluctuations in the 5-year average.

From 1850 to 1900 both Denmark and Norway diverge from Sweden slightly but in different directions. But this is based on a comparison of only one or two stations in each case and so is not unexpected.


Friday, May 14, 2021

67. More evidence against temperature data adjustments (USA)

In Post 57 (The case against temperature data adjustments) I presented evidence that seemed to cast doubt on the need to adjust temperature data. The main argument from climate scientists in favour of these adjustments is their belief that the raw data cannot always be trusted. Over time, changes to the data collection process may occur. These changes may be due to changes in the location of the weather station, changes to the environment around the original site, or changes to the instrumentation or data collection methods. 

It is certainly true that these issues affect many, if not most temperature records, and the longer the temperature record, the more likely such issues will probably occur. The important questions, though, are: how large are these data errors, and what is the best way to eliminate them from historical data?


Fig. 67.1: The temperature anomalies for Baker City Municipal Airport as calculated by Berkeley Earth.


The approach most climate groups use is to adjust data within each individual temperature record, an example of which can be seen by comparing the data in Fig. 67.1 above and Fig. 67.2 below. Both graphs show temperature data from Baker City Municipal Airport (Berkeley Earth ID: 164703) in the state of Oregon in the USA, which has then been adjusted by Berkeley Earth. The original data in Fig. 67.1 has no discernible temperature trend (green line), but after the data has been chopped into multiple autonomous segments, and those segments each subjected to its own separate corrective bias, the overall trend becomes strongly positive with a gradient of +0.67°C per century. Thus warming appears where before there was none.


Fig. 67.2: The adjustments made by Berkeley Earth to the temperature anomalies for Baker City Municipal Airport.


The justification for using these adjustments is that climate scientists believe they can identify points in the data where errors have been introduced, and also that they can determine what the correction factor needs to be in order to eradicate the error. The size of the adjustments is usually determined by comparing the station temperature time series with that of its neighbours, but as I showed in Post 43, even identical neighbours can display temperature differences of up to ±0.25°C just due to measurement uncertainties.

Adjustments are most commonly made at positions in the time series corresponding to known or documented station moves (red diamonds in Fig. 67.2), or at points where there is a gap in the temperature record (green diamonds). But, groups such as NOAA and Berkeley Earth have also developed algorithms that they claim can identify other points in the time series where undocumented changes have occurred. These positions in the data are referred to by NOAA and Berkeley Earth as changepoints and breakpoints respectively. To many climate sceptics, however, these techniques remain controversial. But I would argue that in many cases they are also unnecessary because of a statistical phenomenon called regression towards the mean.

 

Fig. 67.3: The average temperature trend for the 100 longest temperature records in the USA. The best fit is applied to the monthly mean data from 1921 to 2010 and has a positive gradient of +0.25 ± 0.15 °C per century. The monthly temperature changes are defined relative to the 1951-1980 monthly averages.

 

Basically, if the errors are randomly distributed between different records, and also at different times in those records, and if they are of comparable size, then any averaging process will cause the errors to partially cancel. The bigger the number of records in the average, the more precisely they will cancel.

In Post 57 I demonstrated that these errors can be eradicated using a simple averaging process. I did this by averaging unadjusted temperature data from stations located in neighbouring European countries (Germany, Austria, Hungary and Czechoslovakia), and showing that the averaging process gave the same result for the 5-year average trend for each country, provided there were more than about twenty stations in the average for each country. This was despite the fact that Berkeley Earth had applied over three adjustments on average to each temperature record during its own analysis process for those same stations.

 

Fig. 67.4: The average temperature trend for the 101st to the 200th longest temperature records in the USA. The best fit is applied to the monthly mean data from 1921 to 2010 and has a negative gradient of -0.11 ± 0.15 °C per century. The monthly temperature changes are defined relative to the 1951-1980 monthly averages.

 

The key to this is having sufficient data. In the case of the USA we have more than sufficient data. In Post 66 I analysed the 400 longest temperature records for the USA and determined the temperature trend since 1750. This in turn showed no evidence of any global warming in the USA over the last 100 years. But suppose we split those 400 records into four sets of 100 records, and compare the four results for the different mean temperature trends. What would we expect to see?


Fig. 67.5: The average temperature trend for the 201st to the 300th longest temperature records in the USA. The best fit is applied to the monthly mean data from 1921 to 2010 and has a slight negative gradient of -0.003 ± 0.135 °C per century. The monthly temperature changes are defined relative to the 1951-1980 monthly averages.

 

Well, the answer is shown in Fig. 67.3-Fig. 67.6. The result is that the four temperature trends look very similar (it is probably easiest to compare the 5-year moving average curves). But judging by the adjustments made by Berkeley Earth to the data in Fig. 67.2, it would not be unreasonable to expect Berkeley Earth to have made over 1000 adjustments in total to the 100 station records used to generate each of these four temperature trends. So not making these 1000 adjustments should result in large discrepancies between the four different trends, assuming the adjustments are needed. But they aren't needed, and there are no large discrepancies.

 

Fig. 67.6: The average temperature trend for the 301st to the 400th longest temperature records in the USA. The best fit is applied to the monthly mean data from 1921 to 2010 and has a negative gradient of -0.22 ± 0.14 °C per century. The monthly temperature changes are defined relative to the 1951-1980 monthly averages.

 

The four trends are compared in more detail in Fig. 67.7 below. The average of the anomalies from the 100 longest records in the USA are shown in yellow and are offset by -1°C for clarity. The mean of next longest 100 records is shown in blue. The mean of the third longest set is shown in red and offset by +1°C, with the fourth longest set shown in black, but not offset. To aid the analysis process, the blue curve is plotted three times, with three different offsets so that it can be compared with the other three trends.

 

Fig. 67.7:  A comparison of the 5-year averaged temperature trends for four sets of 100 temperature records in the USA. The trends are offset for clarity with the trend for stations 101-200 used as a comparator for each of the other three trends.

 

What is clear is that for all the data from 1890 onwards the four trends are virtually identical. This implies that the averaging process has eliminated almost all the data errors. Before 1890 the number of stations in each average decreases dramatically as shown in Fig. 67.8 below, which is why the level of agreement between the curves is much lower. From 1900 onwards, however, there is almost total agreement. The only significant differences are for stations 001-100 in the 1930s, and stations 301-400 post-1995, but in both cases the discrepancy is generally less than 0.25°C. These differences also largely account for the differences in the best fit lines in Fig. 67.3-Fig. 67.6. As for their causes, well the lower temperature anomaly for stations 301-400 after 1995 could be due to these stations being newer than the rest. That would imply they are located in smaller towns with less waste heat production. For stations 001-100 the opposite is probably true as these station time series are the longest. That in turn suggests they are more likely to be located near the largest cities.


Fig. 67.8: The number of station records included each month in the mean temperature trends.


What this demonstrates unequivocally is that data adjustments are unnecessary when determining global or regional mean trends. This is because the errors in the individual station records will cancel when averaged. If they did not, then the four trends in Fig. 67.7 would not be so alike. Instead there would be significant differences. And remember, the same result was demonstrated in Post 57.

But what this also implies is that, if the errors in the individual station records cancel, then so too should the adjustments that are applied by Berkeley Earth and others to correct these errors. Except they don't.

 

Fig. 67.9: The average temperature trend for the 301st to the 400th longest temperature records in the USA after adjustments made by Berkeley Earth. The best fit is applied to the monthly mean data from 1921 to 2010 and has a positive gradient of +0.54 ± 0.05 °C per century.

 

The graph in Fig. 67.9 above shows the mean temperature trend for stations 301 to 400 with their Berkeley Earth adjustments included. If the adjustments cancelled, then the graph should resemble the data in Fig. 67.6, but it doesn't. Instead of a significant negative trend, there is a sizeable positive trend. In fact the adjustments have added a net warming of over +0.7°C to the data since 1920. The same positive trend is also seen for the means of the adjusted data for the other three sets of 100 stations, so at least they are consistent, but that does not mean they are correct. In fact all they are doing here is adding warming where none existed previously.

 

Summary

What I have presented here is yet more compelling evidence against the statistical validity of temperature adjustments.

I have shown that the true temperature trend can be determined simply by averaging the anomalies from the raw data. This confirms the similar result for Central European data that I presented in Post 57.

This adds further weight to my claim in Post 66 that there has been no global warming in the USA since 1900.

 

Wednesday, March 31, 2021

57. The case against temperature data adjustments (EU)

Fig. 57.1: The number of weather stations with temperature data in the Northern Hemisphere since 1700 according to Berkeley Earth.

 

There are four major problems with global temperature data.

1) It is not spread evenly

Only about 10% of all available data covers the Southern Hemisphere (compare Fig. 57.2 below with Fig. 57.1 above), while in the Northern Hemisphere over half the data is from the USA alone (as shown in Fig. 57.3). In addition, there is no reliable temperature data covering the oceans from before 1998 when the Argo programme for a global array of 3000 autonomous profiling floats was proposed. The Argo programme has since been used to measure ocean temperatures and salinity on a continuous basis down to depths of 1000m across most of the oceans between the polar regions, but that means we only have reliable data for the last 20 years. 

The result is that only land-based data is available before 1998, and this tends to cluster around urban areas. The solution to this clustering employed by climate scientists is to resort to techniques such as gridding, weighting and homogenization. 

Gridding involves creating a virtual grid of points across the Earth's surface, usually 1° of longitude or latitude apart. This is limited by two factors: computing power and data coverage. As there are unlikely to be any weather stations at these grid points, unless by coincidence, virtual station records are created at these points by averaging the temperatures from the nearest real stations. This averaging of stations is not equal. Instead the average usually weights the different stations according to their closeness in distance (although even stations 1000 km away can be included) and their correlation to the mean of all those datasets. This process of weighting based on correlation is often called homogenization. 

 

Fig. 57.2: The number of weather stations with temperature data in the Southern Hemisphere since 1750 according to Berkeley Earth.

 

2) It does not go back far enough in time

As I have shown previously, the earliest temperature records are from Germany (see Post 49) and the Netherlands (see Post 41) and go back to the early 18th century. However, there is no Southern Hemisphere temperature data before 1830, and only two datasets in the USA from before 1810. The principal reason is that the amount of available data is positively correlated with economic development. As more countries have industrialized, the number of weather stations has increased. Unfortunately, climate change involves measuring the change in temperature since a previous epoch or reference period (over say 100 or 200 years), and in those times the availability of data is much, much, worse. So increasing the quality of current data cannot increase the quality of the measured temperature change. This will always be constrained by how much data we had in the distant past.


Fig. 57.3: The number of weather stations with temperature data in the United States since 1700 according to Berkeley Earth.


3) The data is often subject to measurement errors

Over time weather stations are often moved, instruments are ungraded, and the local environment changes as well. The conventional wisdom is that all these changes have profound impacts on the temperature records that need to be compensated for. This is the rationale behind data adjustments. The problem is, none of it is really justified, as I will demonstrate in this post.

If there are problems with the temperature data at different times and locations, these issues should be randomly distributed. That means any adjustments to correct these errors should be randomly distributed as well. This in turn means that averaging a sufficiently large number of stations for a regional or global trend should result in the cancellation of both the errors and the adjustments. As I have shown in many previous posts here, this does not happen. In fact in many cases the adjustments can add (or subtract) as much, or even more, warming (or cooling) to the mean trend than is present in the original data, particularly in the Southern Hemisphere. For examples see my posts for Texas, Indonesia, PNG, the South Pacific (East and West), NSW, Victoria, South Australia, Northern Territory and New Zealand among others.

One contentious issue is the problem of station moves or changes to the local environment. The conventional wisdom is that these will both strongly affect the temperature record. Frankly, I disagree. In my view those who say they will are failing to understand what is being measured. One example is, what would happen if the weather station was to be moved from open ground to an area under a large tree? Does the increased shade reduce the temperature? The answer is no because the thermometer is already in the shade inside its Stevenson screen. Moreover, the thermometer is measuring air temperature, not the temperature on the ground, and the air is continuously circulating. So the air under the tree is at virtually the same temperature as the air above open ground. The one adjustment that does affect temperature is altitude. Air (almost) always gets colder as you ascend in height.

4) There just isn't enough data

There are currently about 40,000 weather stations across the globe. This sounds like a lot, but it is only about one for every 13,000 square kilometres of area. That means that on average, these stations are over 110 km apart, or more than 1° of longitude or latitude. Even today, that is probably the bare minimum of what is required to measure a global temperature. Unfortunately, in previous times, the availability of data was much, much, worse.

Of course, now there are alternatives. One is to use satellites, but again this only provides data back to about 1980. The other problem with satellites is that their orbits generally no not cover the polar regions. And finally, they can only see what is emitted at the top of the atmosphere (TOA). So they can measure temperatures at the TOA, but measuring surface temperatures can be problematic as the infra-red radiation emitted by the surface is largely absorbed by carbon dioxide and water vapour in the lower atmosphere.

Over the course of the last eleven months I have posted 56 articles to this blog. Over half of these have analysed the surface temperature trends in various countries, states and regions. In virtually every case, the trend I have determined by averaging station anomalies has differed from the conventional widely publicized versions. These differences are largely due to homogenization and data adjustments. 

Homogenization

There are two potential issues with homogenization. Firstly, there are more urban stations than rural ones. This is because stations tend to be located near to where people live. Secondly, urban stations tend to be closer together. So they are more likely to be strongly correlated. As homogenization uses correlation for weighting the influence of each station's data in the mean temperature for the local region, this means that the influence of urban stations will be stronger. 

So both potential issues are likely to favour urban stations over rural ones. Yet it is the urban ones that are more likely to be biased due to the urban heat island (UHI) effect. The result is that that bias is often transmitted to the less contaminated rural stations, thereby biasing the whole regional trend upwards. This is why I do not use homogenization in my analysis. The other problematic intervention is data adjustment.

Data adjustments

The rationale for data adjustments is that they are needed to compensate for measurement errors that may occur from changes of station site, instrument or method. The justification for using them is that climate scientists believe they can identify weak points in the data. Some might call that hubris. The alternative viewpoint is that these adjustments are unnecessary and that averaging a sufficiently large sample will erase the errors automatically via regression to the mean. I will now demonstrate that with real data.


Fig. 57.4: The 5-year average temperature trends for Austria, Hungary and Czechoslovakia together with best fit lines for the interval 1791-1980 (m is the gradient in °C per century). The Austria and Czechoslovakia data are offset by +2°C and -2°C respectively to aid clarity.


In three recent posts I calculated and examined the temperature trends for Czechoslovakia (Post 53), Hungary (Post 54) and Austria (Post 55). The five-year moving averages of the temperature trends in these three countries are shown in Fig. 57.1 above. What is immediately apparent is the high degree of similarity that these trends display, particularly after 1940. This is indicated by the red and black arrows which mark the positions of coincident peaks and troughs respectively in the three datasets.

It turns out that all three datasets are also very similar to that of Germany (see Post 49). This is shown in Fig. 57.5 below. This is not surprising as the four countries are all close neighbours. What is surprising is that there are not greater differences between the four datasets, particularly given the number of adjustments that Berkeley Earth felt needed to be made to the individual station records for these countries when undertaking their analysis.


Fig. 57.5: The 5-year average temperature trends for Austria, Hungary and Czechoslovakia compared to that of Germany.


To understand the potential impact of these adjustments, consider this. The temperature trend for Austria in Fig. 55.1 of Post 55 was determined by averaging up to 26 individual temperature records. Yet the total number of adjustments made to those records by Berkeley Earth in the time interval 1940-2013 was more than 90. That is more than three adjustments per temperature record, or at least one for every 21 years of data. Yet if the adjustments are ignored, and the data for each country is just averaged normally, the results for each country, Austria, Czechoslovakia, Germany and Hungary, are virtually identical. This leads to the following conclusions and implications.


Conclusions

1) The data in Fig. 57.5 indicates that the temperature trends for Austria, Czechoslovakia, Germany and Hungary are virtually identical after 1940. The probability that this is due to random chance is minimal. It therefore implies that the temperature trends for these countries from 1940 onwards are indeed virtually identical. This is not a total surprise as they are all close neighbours.

2) As all the individual temperature anomaly time series used to generate these trends are not identical, and all are likely to have data irregularities from time to time, this also means that those data irregularities are highly likely to be random in both their size and distribution across the various time series. This means that when averaged to create the regional trend, their irregularities will partially cancel. If the number of sites is large enough, the cancellation will be almost total. This is what is seen in Fig. 57.5, and it is why all the trends shown are virtually identical post-1940.


Implications

1) If the temperature trends for Austria, Czechoslovakia, Germany and Hungary are virtually identical after 1940, as conclusion #1 suggests, then it is reasonable to suppose that they should be virtually identical before 1940 as well. But they aren't, as the data in Fig. 57.5 illustrates. This is because the trends in each case are based on the average of too few individual anomaly time-series for the irregularities from each station time-series to be fully cancelled by the irregularities from the remainder. Before 1940 there are only sixteen valid temperature records in Austria, three in Hungary and three in Czechoslovakia. Germany, on the other hand has about thirty.

2) However, if it is true that all the temperature trends for Austria, Czechoslovakia, Germany and Hungary before 1940 should be the same, then there is no reason why we cannot combine them all into a single trend. This would dramatically increase the number of individual time-series being averaged, and so reduce the discrepancy between the calculated value for the trend and the true value. This has been done in Fig. 57.6 below.


Fig. 57.6: The temperature trend for Central Europe since 1700. The best fit is applied to the interval 1791-1980 and has a negative gradient of -0.05 ± 0.07 °C per century. The monthly temperature changes are defined relative to the 1981-2010 monthly averages.

 

The data in Fig. 57.6 represents the temperature trend for the combined region of Austria, Czechoslovakia, Germany and Hungary. The trend after 1940 is the same as that seen in those individual countries and the gradient of the best fit line for 1791-1980 more closely resembles the equivalent lines for Germany and Hungary than it does those of Austria and Czechoslovakia. But now we also have a more accurate trend before 1940. The question is, how much more accurate?

 

Fig. 57.7: The number of station time-series included in the average each month for the temperature trend in Fig. 57.6

 

The data from Austria, Czechoslovakia, Germany and Hungary suggest that approximately 20 different time-series are required in the average for the irregularities in the different station time-series to almost fully cancel. The graph in Fig. 57.7 suggests that this threshold is surpassed for almost every month of every year after 1830.

 

Fig. 57.8: The temperature trend for Central Europe since 1700. The best fit is applied to the interval 1831-2010 and has a positive gradient of 0.62 ± 0.07 °C per century. The monthly temperature changes are defined relative to the 1981-2010 monthly averages.

 

If we now calculate the best fit to the data in Fig. 57.8, but only use data after 1830, we get a gradient for the trend line of 0.62 °C per century. This equates to a temperature rise since 1830 of over 1.1 °C.

 

 
Fig. 57.9: The temperature trend for Central Europe since 1700. The best fit is applied to the interval 1781-2010 and has a positive gradient of 0.21 ± 0.05 °C per century. The monthly temperature changes are defined relative to the 1981-2010 monthly averages.

 

However, you could argue that the regional monthly average data in Fig. 57.6 is still reasonably accurate all the way back to 1780 as it continues to have over a dozen temperature records incorporated into the average every month of every year after this time. In which case the temperature rise since 1780, as indicated by the best fit line in Fig. 57.9, is actually less than 0.5 °C. This suggests that we can be reasonably confident that temperatures in central Europe between 1750 and 1830 were fairly similar to those of today.


Summary

What I have demonstrated here is that adjustments to the raw temperature data are unnecessary and can be avoided simply by averaging sufficient datasets (i.e. more than about 20).

I have also shown that it is highly likely that the mean temperature in central Europe is not much higher now than it was at the start of the Industrial Revolution (1750-1830). 


Disclaimer: No data were harmed or mistreated during the writing of this post. This blog believes that all data deserve to be respected and to have their values protected.


Thursday, December 31, 2020

45. Review of the year 2020

I started this blog in May, in part to occupy my time during the Covid-19 lockdown. But I was also motivated by a growing dissatisfaction with the quality of data analysis I was witnessing in climate science, and in particular the lack of any objectivity in the way much of the data was being presented and reported. My concerns were twofold. 

The first was the drip-drip of selective alarmism with an overt confirmation bias that kept appearing in the media with no comparable reporting of events that contradicted that narrative. The worry here is that extreme events that are just part of the natural variation of the climate were being portrayed as the new normal, while events of the opposite extreme were being ignored. It appeared that balance was being sacrificed for publicity.

The second was the over-reliance of much of the climate analysis on complex statistical analysis techniques of doubtful accuracy or veracity. To paraphrase Lord Rutherford: if you need to use complex statistics to see any trends in your data, then you would be better off using better data. Or to put it more simply, if you can't see a trend with simple regression analysis, then the odds are there is no trend to see.

The purpose of this blog has not been to repeat the methods of climate scientists, nor to improve on them. It has merely been to set a benchmark against which their claims can be measured and tested.

My first aim has been to go back to basics, to examine the original temperature data, look for trends in that data, and to apply some basic error analysis to determine how significant those trends really are. Then I have sought to compare what I see in the original data with what climate scientists claim is happening. In most cases I have found that the temperature trends in the real data are significantly less than those reported by climate scientists. In other words, much of the reported temperature rises, particularly in Southern Hemisphere data, result from the data manipulations performed by the climate scientists on the data. This implies that many of the reported temperature rises are an exaggeration.

In addition, I have tried to look at the physics and mathematics underpinning the data in order to test other possible hypotheses that could explain the observed temperature trends that I could detect. Below I have set out a summary of my conclusions so far.


1) The physics and mathematics

There are two alternative theories that I have considered as explanations of the temperature changes. The first is natural variation. The problem here is that in order to conclusively prove this to be the case you need temperature data that extends back in time for dozens of centuries, and we simply do not have that data. Climate scientists have tried to solve this by using proxy data from tree rings and sediments and other biological or geological sources, but in my opinion these are wholly inadequate as they are badly calibrated. The idea that you can measure the average annual temperature of an entire region to an accuracy of better than 0.1 °C simply by measuring the width of a few tree rings, when you have no idea of the degree of linearity of your proxy, or the influence of numerous external variables (e.g. rainfall, soil quality, disease, access to sunlight), is preposterous. But there is another way.

i) Fractals and self-similarity

If you can show that the fluctuations in temperature over different timescales follow a clear pattern, then you can extrapolate back in time. One such pattern is that resulting from fractal behaviour and self-similarity in the temperature record. By self-similarity I mean that every time you average the data you end up with a pattern of fluctuations that looks similar to the one you started with, but with amplitudes and periods that change according to a precise mathematical scaling function.

In Post 9 I applied this analysis to various sets of temperature data from New Zealand. I then repeated it for data from Australia and then again in Post 42 for data from De Bilt in the Netherlands. In virtually all these cases I found a consistent power law for the scaling parameter indicative of a fractal dimension of between 0.20 and 0.30, with most values clustered close to 0.25. The low magnitude of this scaling term suggests that the fluctuations in long term temperatures are much greater in amplitude than conventional statistical analysis would predict. 

For example, in the case of De Bilt it suggests that the standard deviation in the average 100-year temperature is more than 0.2 °C. This means that there is a 16% probability of the mean temperature for any century being more than 0.3°C more (or less) than the mean temperature for the previous century, and therefore a one in six possibility of a 0.6 °C temperature rise in any given century. So a 0.6 °C temperature rise over a century could occur once every 600 years purely because of natural variations in temperature. It also suggests that similar temperature variations that we have seen in temperature data in the last 50 or 100 years might have been repeated frequently in the not so distant past.

ii) Direct anthropogenic surface heating (DASH) and the urban heat island (UHI)

Another possible explanation for any observed rise in temperature is the heating of the environment that occurs due to human industrial activity. All energy use produces waste heat. Not only that, but all energy must end up as heat and entropy in the end. The Second Law of Thermodynamics tells us that. It is therefore inevitable that human activity must heat the local environment. The only question is by how much.

Most discussions in this area focus on what is known as the urban heat island (UHI). This is a phenomenon whereby urban areas either absorb extra solar radiation because of changes made to the surface albedo by urban development (e.g. concrete, tarmac, etc), or tall buildings trap the absorbed heat and reduce the circulation of warm air, thereby concentrating the heat. But there is another contribution that continually gets overlooked - direct anthropogenic surface heating (DASH). 

When humans generate and consume energy they liberate heat or thermal energy. This energy heats up the ground, and the air just above it, in much the same way that radiation from the Sun does. In so doing DASH adds to the heat that is re-emitted from the Earth's surface, and therefore increases the Earth's surface temperature at that location.

In Post 14 I showed that this heating can be significant - up to 1 °C in countries such as Belgium and the Netherlands with high levels of economic output and high population densities. In Post 29 I extended this idea to look at suburban energy usage and found a similar result. 

What this shows is that you don't need to invoke the Greenhouse Effect to find a plausible mechanism via which humans are heating the planet. Simple thermodynamics will suffice. Of course climate scientists dismiss this because they assume that this heat is dissipated uniformly across the Earth's surface - but it isn't. And just as significant is the fact that the majority of weather stations are in places where most people live, and therefore they also tend to be in regions where the direct anthropogenic surface heating (DASH) is most pronounced. So this direct heating effect is magnified in the temperature data.

iii) The data reliability

It is taken as read that the temperature data used to determine the magnitude of the observed global warming is accurate. But is it? Every measurement has an error. In the case of temperature data it appears that these errors are comparable in magnitude to many of the effects climate scientists are trying to measure.

In Post 43 I looked at pairs of stations in the Netherlands that were less than 1.6 km apart. One might expect that most such pairs would exhibit identical datasets for the two stations in the pair, but they don't. In virtually every case the fluctuations in the difference in their monthly average temperatures was about 0.2 °C. While this was consistent with the values one would expect based on error analysis, it does highlight the limits to the accuracy of this data. It also raises questions about how valid techniques such as breakpoint adjustment are, given that these techniques depend on detecting relatively small differences in temperature for data from neighbouring stations.

iv) Temperature correlations between stations

In Post 11 I looked at the product moment correlation coefficients (PMCC) between temperature data from different stations, and compared the correlation coefficients with the station separation. What became apparent was evidence for a strong negative linear relationship between the maximum correlation coefficient for temperature anomalies between pairs of station and their separation. For station separations of less than 500 km positive correlations of better than 0.9 were possible, but this dropped to a maximum correlation of about 0.7 for separations of 1000 km and 0.3 at 2000 km.

There were also clear differences between the behaviour of the raw anomaly data and the Berkeley Earth adjusted data. The Berkeley Earth adjustments appear to reduce the scatter in the correlations for the 12-month averaged data, but do so at the expense of the quality of the monthly data. This suggests that these adjustments may be making the data less reliable not more so. The improvement in the scatter of the Berkeley Earth 12-month averaged data is also curious. Is it because it is this data that is used to determine the adjustments and not the monthly data, or is this not the case and instead there is some other reason? And what of the scatter in the data? Can we use this to measure the quality and reliability of the original data? This clearly warrants further study.


Fig. 45.1: Correlations (PMCC) for the period 1971-2010 between temperature anomalies for all stations in New Zealand with a minimum overlap of 200 months. Three datasets were studied: a) the monthly anomalies; b) the 12-month average of the monthly anomalies; c) the 5-year average of the monthly anomalies. Also studied were the equivalent for the Berkeley Earth adjusted data.



2) The data

Over the last eight months I have analysed most of the temperature data in the Southern Hemisphere as well as all the data in Europe that predates 1850. The results are summarized below.

i) Antarctica

In Post 4 I showed that the temperature at the South Pole has been stable since the 1950s. There is no instrumental temperature data before 1956 and there are only two stations of note near the South Pole (Amundsen-Scott and Vostok). Both show stable or negative trends.

Then in Post 30 I looked at the temperature data from the periphery of the continent. This I divided into three geographical regions: the Atlantic coast, the Pacific coast and the Peninsula. The first two only have data from about 1950 onwards. In both cases the temperature data is also stable with no statistically significant trend either upwards or downwards. Only the Peninsula exhibited a strong and statistically significant upward trend of about 2 °C since 1945.


ii) New Zealand

Fig. 45.2: Average warming trend of for long and medium stations in New Zealand. The best fit to the data has a gradient of +0.27 ± 0.04 °C per century.

In Posts 6-9 I looked at the temperature data from New Zealand. Although the country only has about 27 long or medium length temperature records, with only ten having data before 1880, there is sufficient data before 1930 to suggest temperatures in this period were almost comparable to those of today. The difference is less than 0.3 °C.


iii) Australia

Fig. 45.3: The temperature trend for Australia since 1853. The best fit is applied to the interval 1871-2010 and has a gradient of 0.24 ± 0.04 °C per century.

The temperature trend for Australia (see Post 26) is very similar to that of New Zealand. Most states and territories exhibited high temperatures in the latter part of the 19th century that then declined before increasing in the latter quarter of the 20th century. The exceptions were Queensland (see Post 24) and Western Australia (see Post 22), but this was largely due to an absence of data before 1900. While there is much less temperature data for Australia before 1900 compared to the latter part of the 20th century, there is sufficient to indicate that, as in New Zealand, temperatures in the late 19th century were similar to those of the present day.


iv) Indonesia

Fig. 45.4: The temperature trend for Indonesia since 1840. The best fit is applied to the interval 1908-2002 and has a negative gradient of -0.03 ± 0.04 °C per century.

The temperature data for Indonesia is complicated by the lack of quality data before 1960 (see Post 31). The temperature trend after 1960 is the average of between 33 and 53 different datasets, but between 1910 and 1960 it generally comprises less than ten. Nevertheless, this is sufficient data to suggest that temperatures in the first half of the 20th century were greater than those in the latter half. This is despite the data from Jakarta Observatorium which exhibits an overall warming trend of nearly 3 °C from 1870 to 2010 (see Fig. 31.1 in Post 31).

It is also worth noting that the temperature data from Papua New Guinea (see Post 32) is similar to that for Indonesia for the period from 1940 onwards. Unfortunately Papua New Guinea only has one significant dataset that predates 1940, so conclusions regarding the temperature trend in this earlier time period are difficult to ascertain.


v) South Pacific

Most of the temperature data from the South Pacific comes from the various islands in the western half of the ocean. This data exhibits little if any warming, but does exhibit large fluctuations in temperature over the course of the 20th century (see Post 33). The eastern half of the South Pacific, on the other hand, exhibits a small but discernible negative temperature trend of between -0.1 and -0.2 °C per century (see Post 34).


vi) South America

Fig. 45.5: The temperature trend for South America since 1832. The best fit is applied to the interval 1900-1999 and has a gradient of +0.54 ± 0.05 °C per century.

In Post 35 I analysed over 300 of the longest temperature records from South America, including over 20 with more than 100 years of data. The overall trend suggests that temperatures fluctuated significantly before 1900 and have risen by about 0.5 °C since. The high temperatures seen before 1850 are exclusively due to the data from Rio de Janeiro and so may not be representative of the region as a whole.


vii) Southern Africa

Fig. 45.6: The temperature trend for South Africa since 1840. The best fit is applied to the interval 1857-1976 and has a gradient of +0.017 ± 0.056 °C per century.

In Posts 37-39 I looked at the temperature trends for South Africa, Botswana and Namibia. Botswana and Namibia were both found to have less than four usable sets of station data before 1960 and only about 10-12 afterwards. South Africa had much more data, but the general trends were the same. Before 1980 the temperature trends were stable or perhaps slightly negative, but after 1980 there was a sudden rise of between 0.5 °C and 2 °C in all three trends, with the largest being found in Botswana. This does not correlate with accepted theories on global warming (the rises in temperature are too large and too sudden, and do not correlate with rises in atmospheric carbon dioxide), and so the exact origin of these rises appears to be unexplained.

 

viii) Europe

Fig. 45.7: The temperature trend for Europe since 1700. The best fit is applied to the interval 1731-1980 and has a positive gradient of +0.10 ± 0.04 °C per century.

In Post 44 I used the 109 longest temperature records to determine the temperature trend in Europe since 1700. The resulting data suggests that temperatures were stable from 1700 to 1980 (they rose by less than 0.25 °C), and then rose suddenly by about 0.8 °C after 1986. The reason for this change is unclear, but one possibility is that it has occurred due to a significant improvement in air quality that reduced the amount of particulates in the atmosphere. These particulates, that may have been present in earlier years, could have induced a cooling that compensated for the underlying warming trend. Once removed, the temperature then rebounded. Even if this is true, it suggests a maximum warming of about 1 °C since 1700, much of which could be the result of direct anthropogenic surface heating (DASH) as discussed in Post 14. In countries such as Belgium and the Netherlands the temperature rise is even less than that expected from such surface heating. It is also much less than that expected from an enhanced Greenhouse Effect due to increasing carbon dioxide levels in the atmosphere (i.e. about 1.5 °C in the Northern Hemisphere since 1910). In fact the total temperature rise should exceed 2.5 °C. So here is the BIG question? Where has all that missing temperature rise gone?


Tuesday, December 8, 2020

43. The reliability of individual temperature records

One of my many criticisms of climate scientists is their use of adjustments to temperature data to supposedly correct for errors in the measurements, corrections which in my opinion are probably not needed for errors that are not real. These corrections come in two main types: homogenization and breakpoint adjustments.

In the case of homogenization, records from neighbouring stations (and the definition of what constitutes a neighbouring station can be somewhat variable) are used to create an average expected temperature for that location, with differences in latitude and elevation compensated for during the process. This homogenization is used to infill missing monthly data points in each record. But it is also used to define the monthly reference temperatures (MRTs) that then define the monthly anomalies.

Breakpoint adjustments, or changepoint adjustments (see this PDF from NOAA) as they are alternatively called, are supposedly used to correct for false trends in the data. This generally means adjusting the slope of all sets of station data so that they look more or less the same, and more importantly have the same general trends as those quoted by the IPCC. So a station like Jakarta Observatorium in Indonesia (Berkeley Earth ID:155660) which actually has a very large warming trend of 1.84 °C per century, and has had since 1870 (and therefore has a total warming since 1870 of over 2.6 °C), gets adjusted down so that its trend is only 0.95 °C. This is because its warming is too high to fit with the IPCC narrative of only 1.0 °C of warming in the Southern Hemisphere since 1900. 

On the other hand, Dubbo (Darling Street) in New South Wales (Berkeley Earth ID:152082) which also has temperature data dating from about 1870, but instead has a negative warming (or cooling) trend of -0.32 °C per century, gets adjusted up so that its trend becomes +0.56 °C per century, and thus closer to accepted "real" value of 1.0 °C per century. 

If this all sounds a bit fishy, then welcome to the wonderful and wacky world of climate science, where nothing is quite as it seems. Central to all these data corrections is the assumption that most of the underlying data is reliable, but more importantly, that it is possible to detect the bad data from the good data. The questions, is any of this true? Is most of the data good? Can we really detect the small amount of bad data? And can the good data actually be so unreliable, or subject to so many unknown hidden variables, that it looks like bad data? One way to test this is by comparing data from stations that are very close neighbours.

As I pointed out in Post 41, the Netherlands has a number of stations that are located very close to a neighbouring station. In fact I have identified nine pairs of stations in the Netherlands where both stations have over 480 months of data, where there is significant temporal overlap of their data (i.e. they have a lot of months where both stations have active data), and where their spatial separation is less than 1.6 km (or about one mile for those dinosaurs from the USA who can't do metric). This allows direct comparisons of data to be made for stations that are, or should be, virtually identical. It is worth noting here that for this purpose the Netherlands has another unique advantage: it is very flat. That means that we do not need to worry about temperature differences occurring between stations due to differences in altitude.

In order to test the reliability of these temperature records I will apply three tests to their data. The first will look at the difference in the mean temperature of each set of station data in the pair. Ideally this should be zero, but there may be a systematic offset between stations due to local geography that could be significant. Such a difference would not necessarily raise question-marks over the validity of the data.

The second test will look at the difference in monthly temperatures between the two stations over time. The issue here is how much randomness is there in the temperature difference, and how significant is it. This will be measured by calculating the standard deviation of the temperature difference. Again, I would expect to see a low value here with noise levels in this data being at least at least a factor of √30 less than the accuracy of the daily mean temperature of each station (which I would estimate conservatively at 1 °C). Overall, this suggests that the standard deviation of this dataset should be less than 0.2 °C, and probably less than 0.1 °C.

Finally, I will look at the trend of the difference in temperature over time. If this is significantly large and comparable to the trends seen in the anomaly data for either station, that would indicate significant reliability problems with this type of data.

The results of these three test are summarized below for each of the nine pairs of stations.


Case 1: Soesterberg

Fig 43.1: The difference is monthly mean temperatures for two stations at Soesterberg. The mean of the monthly differences is 0.17 °C, the standard deviation of the differences is 0.27 °C, and the trend in the differences is -0.29 ± 0.10 °C per century.


The two stations at Soesterberg are BE-92835 (trend of +2.55 °C per century) and BE-139138 (trend of +2.29 °C per century). According to Berkeley Earth they are 1.06 km apart.



Case 2: Schiphol

Fig 43.2: The difference is monthly mean temperatures for two stations at Schiphol. The mean of the monthly differences is 0.09 °C, the standard deviation of the differences is 0.17 °C, and the trend in the differences is 0.017 ± 0.056 °C per century.


The two stations at Schiphol are BE-18517 (trend of +2.53 °C per century) and BE-157005 (trend of +2.12 °C per century). According to Berkeley Earth they are 1.2 km apart.



Case 3: Valkenberg

Fig 43.3: The difference is monthly mean temperatures for two stations at Valkenberg. The mean of the monthly differences is 0.18 °C, the standard deviation of the differences is 0.20 °C, and the trend in the differences is -0.07 ± 0.09 °C per century.


The two stations at Valkenberg are BE-174609 (trend of +2.29 °C per century) and BE-157004 (trend of +1.65 °C per century). According the Berkeley Earth they are 0.25 km apart.



Case 4: Eindhoven

Fig 43.4: The difference is monthly mean temperatures for two stations at Eindhoven. The mean of the monthly differences is 0.10 °C, the standard deviation of the differences is 0.20 °C, and the trend in the differences is 0.20 ± 0.06 °C per century.


The two stations at Eindhoven are BE-18478 (trend of +2.31 °C per century) and BE-156991 (trend of +2.06 °C per century). According to Berkeley Earth they are 1.42 km apart.



Case 5: Volkel

Fig 43.5: The difference is monthly mean temperatures for two stations at Volkel. The mean of the monthly differences is 0.10 °C, the standard deviation of the differences is 0.23 °C, and the trend in the differences is 0.20 ± 0.07 °C per century.


The two stations at Volkel are BE-92832 (trend of +2.31 °C per century) and BE-156995 (trend of +2.10 °C per century). According to Berkeley Earth they are 0.81 km apart.



Case 6: Gilze Rijen

Fig 43.6: The difference is monthly mean temperatures for two stations at Gilze Rijen. The mean of the monthly differences is 0.11 °C, the standard deviation of the differences is 0.30 °C, and the trend in the differences is -0.01 ± 0.09 °C per century.


The two stations at Gilze Rijen are BE-18485 (trend of +2.41 °C per century) and BE-156994 (trend of +1.93 °C per century). According to Berkeley Earth they are 0.16 km apart.



Case 7: Deelen

Fig 43.7: The difference is monthly mean temperatures for two stations at Deelen. The mean of the monthly differences is 0.11 °C, the standard deviation of the differences is 0.25 °C, and the trend in the differences is -0.13 ± 0.09 °C per century.


The two stations at Deelen are BE-18506 (trend of +2.50 °C per century) and BE-157001 (trend of +1.78 °C per century). According to Berkeley Earth they are 1.62 km apart.



Case 8: Rotterdam

Fig 43.8: The difference is monthly mean temperatures for two stations at Rotterdam. The mean of the monthly differences is 0.21 °C, the standard deviation of the differences is 0.21 °C, and the trend in the differences is -0.26 ± 0.14 °C per century.


The two stations at Rotterdam are BE-18497 (trend of +2.17 °C per century) and BE-18496 (trend of +1.80 °C per century). According to Berkeley Earth they are 0.89 km apart.



Case 9: Hoek Van Holland

Fig 43.9: The difference is monthly mean temperatures for two stations at Hoek Van Holland. The mean of the monthly differences is 0.07 °C, the standard deviation of the differences is 0.29 °C, and the trend in the differences is 0.50 ± 0.18 °C per century.


The two stations at Hoek Van Holland and BE-156999 (trend of +1.95 °C per century) and BE-18500 (trend of +1.62 °C per century). According to Berkeley Earth they are 0.87 km apart.


Summary

The three measures I have used to assess the reliability of the temperature records are the difference in the mean temperatures of various pairs of stations, the standard deviation of that difference in monthly temperatures between the two stations, and the magnitude of the trend difference in monthly temperatures. It is important to point out that the data used in the analysis shown in the figures above was the raw monthly temperature data, and not the monthly anomaly data. Overall, the results can be summarized as follows.

1) The difference in mean temperatures

The data shown above for nine pairs of stations indicates that in each case the mean temperature of the two stations can differ by up to 0.2 °C. In fact the mean difference is about 0.13 °C. The question we then need to answer is, is this difference in line with expectations based on known measurement accuracies for the actual data? Or is it determined by other factors such as random variations in the local climate or systematic differences due to differing local environments?

The expected error in the difference in mean temperatures comes from two main sources. One arises from the error in calculating the mean temperature of each station, while the second comes from the expected temperature difference due to their spatial separation.

In order to estimate the first error we start with the original measurement error in the mean daily temperature. This should be less than 1 °C. Then, as each station has over 480 months of data, and each month is itself the average of approximately 30 daily readings, the total number of daily readings being averaged for each station will be N ≥ 30x480. This implies that N ≥ 14400. Now statistical theory states that the error in measuring the mean temperature of a particular station over N readings should be a factor of √N less than the error in a single daily mean temperature measurement. So, this component of the error should be less than 1/120 of 1 °C, in other words less than about 0.008 °C. Combining the error from second station will increase this error by a factor of √2 to give 0.012 °C

The second error component can be estimated by looking at how the global mean temperature changes with latitude. At the equator mean temperatures are about 25 °C, while around the Arctic Circle they drop to near zero. this implies that mean temperatures drop by about 1 °C for every 300 km of latitude. As the two stations in each station pairs are never more than about 1.5 km apart, this implies a maximum difference in temperature due to location of about 0.005 °C. 

Combining the two errors above (by summing their squares) give a combined maximum expected error of 0.013 °C. This is an order of magnitude less than what we observe. This suggests the difference in the mean temperatures is too high to be solely due to measurement uncertainties, even if we allow for differences in local geographical location. It seems likely that local environment differences are the dominant factor here, but these will probably be in the form of fixed temperature offsets that should not impact significantly on the anomaly data over time. If they do, then there will be evidence for this in the form of excessive differences in the trends.


2) The standard deviation

The mean standard deviation of the monthly temperature differences for the nine pairs of stations shown in the figures above is 0.24 °C. While this is much less than the standard deviation of the monthly anomalies of individual stations (typically about 1 °C), it is still significant.

At the start of this post I suggested that 0.2 °C should be a more likely upper limit for the standard deviation, based on the measurement accuracy of the daily mean temperatures, and the number of daily readings that combine to form the monthly mean temperature. This will be heavily dependent on the accuracy of the mean daily temperature, though. 

If the daily mean temperature measurements have an error or uncertainty of 1 °C, then combining 30 of them into a monthly mean will decrease the error or uncertainty for the monthly mean by a factor of √30. However, then comparing the monthly means of two different stations will increase the error in the temperature difference by √2, so overall, the error in the difference in monthly temperatures should be a factor of √15 less than the error in a single mean daily temperature. This is approximately what we see.


3) The long term trend of the temperature difference

Of the three test results, this is probably the most surprising. While one might expect adjacent stations to experience a relative offset in their local temperatures, or differences due to statistical fluctuations over time, generally one would expect their temperature trends to follow each other. Yet the data shown above suggests otherwise.

Overall, the various station pairs exhibited a wide range of trends for their difference in monthly temperatures over time, as illustrated in the figures above. The mean trend seen for the first five pairs of stations (ignoring sign) is approximately 0.15 °C per century. This seems much higher than I would intuitively expect, but is it?

The difference in the trends is likely to be related to the uncertainty in the trends for the anomalies of each station dataset. These depend to the standard deviation of the residuals and inversely with the length of the dataset. For any best fit or trend line the error in the gradient can by estimated by dividing the standard deviation of the residuals by the standard deviation of the x-values multiplied by the square root of the number of x-values. 

In this case the residual is effectively the difference in monthly mean temperatures between stations, and the x-values are the time axis in the graphs above. The standard deviation of the x-values is roughly 12 years and there are roughly 400 points, while the standard deviation of the residuals is effectively 0.24 °C. This suggests that the trend seen in the temperature difference data is likely to be in the range ±0.001 °C per year, or ±0.1 °C per century. Again this is roughly what we see, although the actual trends in the graphs shown above are about double this value, so maybe there is some additional (but relatively small) influence here due to differences in the local environment for the two stations over time. 


Conclusions

The analysis above indicates that even weather stations that are located close together can yield significantly different results from each other for their temperature trends, mean temperatures and temperature distributions over time, just through the presence of known measurement errors. These differences between nearby stations are much greater than I expected to see before I performed this analysis, but are generally consistent with the measurement data and known sources of error. What it does indicate, though, is that even the best data is not that accurate, reproducible or reliable. Given the lack of long term temperature data for many parts of the world, this raises questions over the accuracy of any climate analysis that relies on this imperfect data.