<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing with OASIS Tables v3.0 20080202//EN" "journalpub-oasis3.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:oasis="http://docs.oasis-open.org/ns/oasis-exchange/table" xml:lang="en" dtd-version="3.0" article-type="research-article">
  <front>
    <journal-meta><journal-id journal-id-type="publisher">NPG</journal-id><journal-title-group>
    <journal-title>Nonlinear Processes in Geophysics</journal-title>
    <abbrev-journal-title abbrev-type="publisher">NPG</abbrev-journal-title><abbrev-journal-title abbrev-type="nlm-ta">Nonlin. Processes Geophys.</abbrev-journal-title>
  </journal-title-group><issn pub-type="epub">1607-7946</issn><publisher>
    <publisher-name>Copernicus Publications</publisher-name>
    <publisher-loc>Göttingen, Germany</publisher-loc>
  </publisher></journal-meta>
    <article-meta>
      <article-id pub-id-type="doi">10.5194/npg-28-409-2021</article-id><title-group><article-title>The blessing of dimensionality for the analysis of climate data</article-title><alt-title>Blessing of dimensionality in climate</alt-title>
      </title-group><?xmltex \runningtitle{Blessing of dimensionality in climate}?><?xmltex \runningauthor{B. Christiansen}?>
      <contrib-group>
        <contrib contrib-type="author" corresp="yes">
          <name><surname>Christiansen</surname><given-names>Bo</given-names></name>
          <email>boc@dmi.dk</email>
        <ext-link>https://orcid.org/0000-0003-2792-4724</ext-link></contrib>
        <aff id="aff1"><institution>Danish Meteorological Institute, Copenhagen, Denmark</institution>
        </aff>
      </contrib-group>
      <author-notes><corresp id="corr1">Bo Christiansen (boc@dmi.dk)</corresp></author-notes><pub-date><day>3</day><month>September</month><year>2021</year></pub-date>
      
      <volume>28</volume>
      <issue>3</issue>
      <fpage>409</fpage><lpage>422</lpage>
      <history>
        <date date-type="received"><day>11</day><month>January</month><year>2021</year></date>
           <date date-type="rev-request"><day>19</day><month>January</month><year>2021</year></date>
           <date date-type="rev-recd"><day>20</day><month>July</month><year>2021</year></date>
           <date date-type="accepted"><day>29</day><month>July</month><year>2021</year></date>
      </history>
      <permissions>
        <copyright-statement>Copyright: © 2021 Bo Christiansen</copyright-statement>
        <copyright-year>2021</copyright-year>
      <license license-type="open-access"><license-p>This work is licensed under the Creative Commons Attribution 4.0 International License. To view a copy of this licence, visit <ext-link ext-link-type="uri" xlink:href="https://creativecommons.org/licenses/by/4.0/">https://creativecommons.org/licenses/by/4.0/</ext-link></license-p></license></permissions><self-uri xlink:href="https://npg.copernicus.org/articles/28/409/2021/npg-28-409-2021.html">This article is available from https://npg.copernicus.org/articles/28/409/2021/npg-28-409-2021.html</self-uri><self-uri xlink:href="https://npg.copernicus.org/articles/28/409/2021/npg-28-409-2021.pdf">The full text article is available as a PDF file from https://npg.copernicus.org/articles/28/409/2021/npg-28-409-2021.pdf</self-uri>
      <abstract><title>Abstract</title>
    <p id="d1e78">We give a simple description of the blessing of dimensionality with
the main focus on the concentration phenomena. These phenomena imply that in
high dimensions the lengths of independent random vectors from the same distribution have almost the same length and that independent vectors
are almost orthogonal. In the climate and atmospheric sciences we rely increasingly on ensemble modelling and face the challenge of analysing
large samples of long time series and spatially extended fields. We show how the properties of high dimensions allow us to obtain analytical
results for e.g. correlations between sample members and the behaviour of the sample mean when the size of the sample grows. We find
that the properties of high dimensionality with reasonable success can be
applied to climate data. This is the case although most climate
data show strong anisotropy and both spatial and temporal dependence, resulting in effective dimensions around 25–100.</p>
  </abstract>
    </article-meta>
  </front>
<body>
      

<sec id="Ch1.S1" sec-type="intro">
  <label>1</label><title>Introduction</title>
      <p id="d1e90">In many areas of geophysics we operate in high-dimensional spaces.
Examples from the atmospheric and climate sciences include extended
spatial fields, such as precipitation or near-surface temperature,
and long time series of atmospheric variables, such as the global mean temperature.  These fields and time series may be either observed or
modelled.  Over the last decades ensemble modelling has been generally
accepted as a valuable tool to gauge the unpredictability and error
originating from uncertain initial conditions or deficiencies in model
physics. There is also an increased tendency for gridded observational
products and reanalyses to apply ensemble techniques to represent the
different uncertainties.  We are therefore often in a situation where we
need to analyse large samples of high-dimensional fields. These samples could consist of the individual ensemble members or just individual
years of a spatial field.</p>
      <p id="d1e93">This might seem a daunting challenge as the properties of high-dimensional space often appear counterintuitive to minds experienced
only in the low-dimensional world.  However, the properties of high-dimensional space may sometimes simplify the analysis and allow us to
obtain rather general analytical results. A central result in this
respect is the concentration of measures which, with a quote from
<xref ref-type="bibr" rid="bib1.bibx11" id="text.1"/>, loosely states that “A random variable that smoothly depends on the influence of many weakly dependent random
variables is, on an appropriate scale, very close to a constant”. The importance of this is described with another quote: “The idea of
concentration of measures is arguably one of the great ideas of analysis
in our time” <xref ref-type="bibr" rid="bib1.bibx45" id="paren.2"/>.  We will see how in high dimensions such concentration properties often allow us to substitute the length
of a random vector with its expectation value and to treat independent
vectors as orthogonal (i.e. having a zero dot product).</p>
      <p id="d1e102">These advantageous properties of high dimensionality – often referred
to as the blessing of dimensionality – have rarely been applied to
the atmospheric and climate sciences. Exceptions are our previous
papers on the subject.  In <xref ref-type="bibr" rid="bib1.bibx13" id="text.3"/> we described how
the blessing of dimensionality explains why the ensemble mean often
outperforms the individual ensemble members and why the ensemble
mean often has an error that is 30 % smaller than the median error
of the individual ensemble members.  In <xref ref-type="bibr" rid="bib1.bibx14" id="text.4"/> we
used the properties of high dimensions to analyse a global ensemble
reforecast. We described how the behaviour of the ensemble mean forecast
can be described by a simple model in which variances and bias depend on
lead time. In <xref ref-type="bibr" rid="bib1.bibx15" id="text.5"/> we analysed a multi-model climate ensemble using the properties of high dimensions to separate two<?pagebreak page410?> competing understandings of the ensemble – the indistinguishable interpretation and
the truth-centred interpretation.  In this paper we aim to give a more comprehensive and coherent discussion of the blessing of dimensionality
and to which extent it applies to the situation in atmospheric science.</p>
      <p id="d1e114">In Sect. <xref ref-type="sec" rid="Ch1.S2"/> we describe the properties of
high-dimensional spaces, focusing first on what is often called the curse/blessing of dimensionality (Sect. <xref ref-type="sec" rid="Ch1.S2"/>.1) and then more specifically on the concentration of measures (Sect. <xref ref-type="sec" rid="Ch1.S2"/>.2). The mathematical results are often only proved
for independent and identically distributed (iid) random variables. In
Sect. <xref ref-type="sec" rid="Ch1.S3"/> we discuss how this requirement can be
loosened and how it relates to geophysical fields which often contain strong temporal and spatial dependence. In Sect. <xref ref-type="sec" rid="Ch1.S4"/>
we focus on the application to atmospheric and climate science. First,
in Sect. <xref ref-type="sec" rid="Ch1.S4"/>.1 we directly investigate to
which extent the climatic fields fulfill the requirements of high
dimensionality. We then (Sects. <xref ref-type="sec" rid="Ch1.S4"/>.2 and
<xref ref-type="sec" rid="Ch1.S4"/>.3) discuss analytical results for distances
and correlations between samples and how well these hold for climate
fields. In Sect. <xref ref-type="sec" rid="Ch1.S4"/>.4 we likewise explore
analytical results for how the ensemble mean depends on ensemble size. The
paper is closed with the conclusions in Sect. <xref ref-type="sec" rid="Ch1.S5"/>.</p>

<?xmltex \floatpos{t}?><table-wrap id="Ch1.T1" specific-use="star"><?xmltex \currentcnt{1}?><label>Table 1</label><caption><p id="d1e142">Results for a unit cube in <inline-formula><mml:math id="M1" display="inline"><mml:mi>N</mml:mi></mml:math></inline-formula> dimensions.
The vertices of a unit cube  <inline-formula><mml:math id="M2" display="inline"><mml:mrow><mml:mo>[</mml:mo><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>/</mml:mo><mml:mn mathvariant="normal">2</mml:mn><mml:mo>,</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>/</mml:mo><mml:mn mathvariant="normal">2</mml:mn><mml:msup><mml:mo>]</mml:mo><mml:mi>N</mml:mi></mml:msup></mml:mrow></mml:math></inline-formula>
are <inline-formula><mml:math id="M3" display="inline"><mml:mrow><mml:mo>[</mml:mo><mml:mo>±</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>/</mml:mo><mml:mn mathvariant="normal">2</mml:mn><mml:mo>,</mml:mo><mml:mo>±</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>/</mml:mo><mml:mn mathvariant="normal">2</mml:mn><mml:mo>,</mml:mo><mml:mi mathvariant="normal">…</mml:mi><mml:mo>±</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>/</mml:mo><mml:mn mathvariant="normal">2</mml:mn><mml:mo>]</mml:mo></mml:mrow></mml:math></inline-formula>.  The number of vertices is <inline-formula><mml:math id="M4" display="inline"><mml:mrow><mml:msup><mml:mn mathvariant="normal">2</mml:mn><mml:mi>N</mml:mi></mml:msup></mml:mrow></mml:math></inline-formula>
and the length of the vertices <inline-formula><mml:math id="M5" display="inline"><mml:mrow><mml:msqrt><mml:mi>N</mml:mi></mml:msqrt><mml:mo>/</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:mrow></mml:math></inline-formula>.  The fraction of volume
within <inline-formula><mml:math id="M6" display="inline"><mml:mi mathvariant="italic">ϵ</mml:mi></mml:math></inline-formula> of the edge is <inline-formula><mml:math id="M7" display="inline"><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>-</mml:mo><mml:mo>(</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>-</mml:mo><mml:mi mathvariant="italic">ϵ</mml:mi><mml:msup><mml:mo>)</mml:mo><mml:mi>N</mml:mi></mml:msup></mml:mrow></mml:math></inline-formula>. The volume of the inscribed sphere is <inline-formula><mml:math id="M8" display="inline"><mml:mrow><mml:msup><mml:mi mathvariant="italic">π</mml:mi><mml:mrow><mml:mi>N</mml:mi><mml:mo>/</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:mrow></mml:msup><mml:mo>(</mml:mo><mml:mi>d</mml:mi><mml:mo>/</mml:mo><mml:mn mathvariant="normal">4</mml:mn><mml:msup><mml:mo>)</mml:mo><mml:mi>N</mml:mi></mml:msup><mml:mo>/</mml:mo><mml:mi mathvariant="normal">Γ</mml:mi><mml:mo>(</mml:mo><mml:mi>N</mml:mi><mml:mo>/</mml:mo><mml:mn mathvariant="normal">2</mml:mn><mml:mo>+</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> with <inline-formula><mml:math id="M9" display="inline"><mml:mrow><mml:mi>d</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:math></inline-formula>.</p></caption><oasis:table frame="topbot"><oasis:tgroup cols="6">
     <oasis:colspec colnum="1" colname="col1" align="left"/>
     <oasis:colspec colnum="2" colname="col2" align="center"/>
     <oasis:colspec colnum="3" colname="col3" align="right"/>
     <oasis:colspec colnum="4" colname="col4" align="center"/>
     <oasis:colspec colnum="5" colname="col5" align="right"/>
     <oasis:colspec colnum="6" colname="col6" align="center"/>
     <oasis:thead>
       <oasis:row>
         <oasis:entry colname="col1"><inline-formula><mml:math id="M10" display="inline"><mml:mi>N</mml:mi></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col2">Volume</oasis:entry>
         <oasis:entry colname="col3">No. of vertices</oasis:entry>
         <oasis:entry colname="col4">Length of</oasis:entry>
         <oasis:entry colname="col5">Volume of</oasis:entry>
         <oasis:entry colname="col6">Fraction of volume</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2"/>
         <oasis:entry colname="col3"/>
         <oasis:entry colname="col4">vertices</oasis:entry>
         <oasis:entry colname="col5">inscribed sphere</oasis:entry>
         <oasis:entry colname="col6">within 0.05 of edge</oasis:entry>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row>
         <oasis:entry colname="col1">2</oasis:entry>
         <oasis:entry colname="col2">1</oasis:entry>
         <oasis:entry colname="col3">4</oasis:entry>
         <oasis:entry colname="col4">0.707</oasis:entry>
         <oasis:entry colname="col5">0.785</oasis:entry>
         <oasis:entry colname="col6">0.0975</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">3</oasis:entry>
         <oasis:entry colname="col2">1</oasis:entry>
         <oasis:entry colname="col3">8</oasis:entry>
         <oasis:entry colname="col4">0.866</oasis:entry>
         <oasis:entry colname="col5">0.524</oasis:entry>
         <oasis:entry colname="col6">0.1426</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">5</oasis:entry>
         <oasis:entry colname="col2">1</oasis:entry>
         <oasis:entry colname="col3">32</oasis:entry>
         <oasis:entry colname="col4">1.118</oasis:entry>
         <oasis:entry colname="col5">0.164</oasis:entry>
         <oasis:entry colname="col6">0.2262</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">10</oasis:entry>
         <oasis:entry colname="col2">1</oasis:entry>
         <oasis:entry colname="col3">1024</oasis:entry>
         <oasis:entry colname="col4">1.581</oasis:entry>
         <oasis:entry colname="col5">0.00249</oasis:entry>
         <oasis:entry colname="col6">0.4013</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">25</oasis:entry>
         <oasis:entry colname="col2">1</oasis:entry>
         <oasis:entry colname="col3">3.35 <inline-formula><mml:math id="M11" display="inline"><mml:mrow><mml:mo>×</mml:mo><mml:msup><mml:mn mathvariant="normal">10</mml:mn><mml:mn mathvariant="normal">7</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col4">2.500</oasis:entry>
         <oasis:entry colname="col5">2.85 <inline-formula><mml:math id="M12" display="inline"><mml:mrow><mml:mo>×</mml:mo><mml:msup><mml:mn mathvariant="normal">10</mml:mn><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">11</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col6">0.7226</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">50</oasis:entry>
         <oasis:entry colname="col2">1</oasis:entry>
         <oasis:entry colname="col3">1.13 <inline-formula><mml:math id="M13" display="inline"><mml:mrow><mml:mo>×</mml:mo><mml:msup><mml:mn mathvariant="normal">10</mml:mn><mml:mn mathvariant="normal">15</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col4">3.535</oasis:entry>
         <oasis:entry colname="col5">1.54 <inline-formula><mml:math id="M14" display="inline"><mml:mrow><mml:mo>×</mml:mo><mml:msup><mml:mn mathvariant="normal">10</mml:mn><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">28</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col6">0.9231</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">100</oasis:entry>
         <oasis:entry colname="col2">1</oasis:entry>
         <oasis:entry colname="col3">1.27 <inline-formula><mml:math id="M15" display="inline"><mml:mrow><mml:mo>×</mml:mo><mml:msup><mml:mn mathvariant="normal">10</mml:mn><mml:mn mathvariant="normal">30</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col4">5.000</oasis:entry>
         <oasis:entry colname="col5">1.87 <inline-formula><mml:math id="M16" display="inline"><mml:mrow><mml:mo>×</mml:mo><mml:msup><mml:mn mathvariant="normal">10</mml:mn><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">70</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col6">0.9941</oasis:entry>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup></oasis:table></table-wrap>

</sec>
<sec id="Ch1.S2">
  <label>2</label><title>Properties of high-dimensional spaces</title>
      <p id="d1e651">Here we give a brief overview of the properties of high-dimensional spaces.  We begin in Sect. <xref ref-type="sec" rid="Ch1.S2"/>.1 with some
general considerations about high-dimensional spaces, while we in Sect. <xref ref-type="sec" rid="Ch1.S2"/>.2 focus more on the concentration of
measures. Some of the simple examples were also, but more briefly,
described in <xref ref-type="bibr" rid="bib1.bibx13" id="text.6"/>.</p>
<sec id="Ch1.S2.SS1">
  <label>2.1</label><title>Curse of dimensionality</title>
      <p id="d1e668">The properties of high-dimensional spaces often defy our intuition based on two and three dimensions <xref ref-type="bibr" rid="bib1.bibx12 bib1.bibx5 bib1.bibx7" id="paren.7"/>.
Apart from the well-known fact – sometimes called the empty space
phenomenon – that the number of samples needed to obtain a given
coverage grows exponentially with dimension, there are other less
appreciated features of high-dimensional spaces <xref ref-type="bibr" rid="bib1.bibx7" id="paren.8"/>. For example, almost every point is an outlier in its own projection, and independent vectors are almost always orthogonal. The latter property
is called waist concentration and, more precisely, states that when
the dimension increases, the angles between independent vectors become narrowly distributed around the mean <inline-formula><mml:math id="M17" display="inline"><mml:mrow><mml:mi mathvariant="italic">π</mml:mi><mml:mo>/</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:mrow></mml:math></inline-formula>, with a variance that
converges towards zero.</p>
      <p id="d1e689">The properties of high-dimensional spaces are sometimes called the curse and sometimes the blessing of dimensionality, depending on the considered problem.  In the present context these properties
turn out to be a blessing as they strongly simplify the analysis
and make analytical results possible.</p>
      <p id="d1e692">As a simple example we consider a cube in <inline-formula><mml:math id="M18" display="inline"><mml:mi>N</mml:mi></mml:math></inline-formula> dimensions with side
<inline-formula><mml:math id="M19" display="inline"><mml:mi>d</mml:mi></mml:math></inline-formula> and centred around <inline-formula><mml:math id="M20" display="inline"><mml:mn mathvariant="bold">0</mml:mn></mml:math></inline-formula>.  The cube has <inline-formula><mml:math id="M21" display="inline"><mml:mrow><mml:msup><mml:mn mathvariant="normal">2</mml:mn><mml:mi>N</mml:mi></mml:msup></mml:mrow></mml:math></inline-formula> vertices with the positions <inline-formula><mml:math id="M22" display="inline"><mml:mrow><mml:mi>d</mml:mi><mml:mo>(</mml:mo><mml:mo>±</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>/</mml:mo><mml:mn mathvariant="normal">2</mml:mn><mml:mo>,</mml:mo><mml:mo>±</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>/</mml:mo><mml:mn mathvariant="normal">2</mml:mn><mml:mo>,</mml:mo><mml:mi mathvariant="normal">…</mml:mi><mml:mo>±</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>/</mml:mo><mml:mn mathvariant="normal">2</mml:mn><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>.  The distance
between each vertex and the centre is <inline-formula><mml:math id="M23" display="inline"><mml:mrow><mml:mi>d</mml:mi><mml:msqrt><mml:mi>N</mml:mi></mml:msqrt><mml:mo>/</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:mrow></mml:math></inline-formula>. The volume of the cube within a distance <inline-formula><mml:math id="M24" display="inline"><mml:mrow><mml:mi mathvariant="italic">ϵ</mml:mi><mml:mi>d</mml:mi></mml:mrow></mml:math></inline-formula> of the edge is <inline-formula><mml:math id="M25" display="inline"><mml:mrow><mml:mo>(</mml:mo><mml:msup><mml:mi>d</mml:mi><mml:mi>N</mml:mi></mml:msup><mml:mo>-</mml:mo><mml:mo>(</mml:mo><mml:mi>d</mml:mi><mml:mo>-</mml:mo><mml:mi mathvariant="italic">ϵ</mml:mi><mml:mi>d</mml:mi><mml:msup><mml:mo>)</mml:mo><mml:mi>N</mml:mi></mml:msup><mml:mo>)</mml:mo><mml:mo>/</mml:mo><mml:msup><mml:mi>d</mml:mi><mml:mi>N</mml:mi></mml:msup><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>-</mml:mo><mml:mo>(</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>-</mml:mo><mml:mi mathvariant="italic">ϵ</mml:mi><mml:msup><mml:mo>)</mml:mo><mml:mi>N</mml:mi></mml:msup></mml:mrow></mml:math></inline-formula> and the volume of the
inscribed sphere is <inline-formula><mml:math id="M26" display="inline"><mml:mrow><mml:msup><mml:mi mathvariant="italic">π</mml:mi><mml:mrow><mml:mi>N</mml:mi><mml:mo>/</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:mrow></mml:msup><mml:mo>(</mml:mo><mml:mi>d</mml:mi><mml:mo>/</mml:mo><mml:mn mathvariant="normal">4</mml:mn><mml:msup><mml:mo>)</mml:mo><mml:mi>N</mml:mi></mml:msup><mml:mo>/</mml:mo><mml:mi mathvariant="normal">Γ</mml:mi><mml:mo>(</mml:mo><mml:mi>N</mml:mi><mml:mo>/</mml:mo><mml:mn mathvariant="normal">2</mml:mn><mml:mo>+</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>.  The situation
is shown in Table <xref ref-type="table" rid="Ch1.T1"/> for a unit cube (<inline-formula><mml:math id="M27" display="inline"><mml:mrow><mml:mi>d</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:math></inline-formula>) for
different values of <inline-formula><mml:math id="M28" display="inline"><mml:mi>N</mml:mi></mml:math></inline-formula>.  For <inline-formula><mml:math id="M29" display="inline"><mml:mrow><mml:mi>N</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">100</mml:mn></mml:mrow></mml:math></inline-formula> there are more than <inline-formula><mml:math id="M30" display="inline"><mml:mrow><mml:msup><mml:mn mathvariant="normal">10</mml:mn><mml:mn mathvariant="normal">30</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula>
vertices<fn id="Ch1.Footn1"><p id="d1e948">Comparable to the number of atoms in 30 <inline-formula><mml:math id="M31" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">t</mml:mi></mml:mrow></mml:math></inline-formula> of water
or the number of bacteria on the Earth: a factor of 1 million larger than the estimated number of stars in the universe.</p></fn> and more that 99 %
of the volume is within a distance 0.05 of the edge. The volume of the
inscribed sphere – which for two dimensions contains the bulk of the cube – is virtually zero. Thus, the volume increasingly concentrates near
the surface when the dimension increases. The form of the <inline-formula><mml:math id="M32" display="inline"><mml:mi>N</mml:mi></mml:math></inline-formula>-dimensional
cube has been compared to that of a sea urchin <xref ref-type="bibr" rid="bib1.bibx29" id="paren.9"/>.</p>
      <p id="d1e970">Consider now a sample of points drawn independently from the high-dimensional cube.  For moderate sample size (<inline-formula><mml:math id="M33" display="inline"><mml:mrow><mml:mo>≪</mml:mo><mml:msup><mml:mn mathvariant="normal">2</mml:mn><mml:mi>N</mml:mi></mml:msup></mml:mrow></mml:math></inline-formula>, which already for
<inline-formula><mml:math id="M34" display="inline"><mml:mrow><mml:mi>N</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">25</mml:mn></mml:mrow></mml:math></inline-formula> is larger than <inline-formula><mml:math id="M35" display="inline"><mml:mrow><mml:msup><mml:mn mathvariant="normal">10</mml:mn><mml:mn mathvariant="normal">7</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula>, so moderate is probably not the right word), all samples will be located in different vertices. This means that all
samples will have almost the same distance from the centre and that all pairs of samples will be almost perpendicular. The distances between pairs
of samples will also be almost identical, making concepts such as nearest neighbours problematic.  However, the ensemble mean will be different
as it will be located near the otherwise vacant centre of the cube. These properties are not particular for the cube but are quite general
also for unbounded distributions, as we will see in the next subsection.</p>
      <p id="d1e1010">The beneficial properties of high dimensionality are recognized in many
areas of machine learning <xref ref-type="bibr" rid="bib1.bibx33 bib1.bibx26" id="paren.10"/>, but the lack
of contrast between distances may also pose problems for algorithms, such as clustering <xref ref-type="bibr" rid="bib1.bibx46 bib1.bibx32" id="paren.11"/>.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F1" specific-use="star"><?xmltex \currentcnt{1}?><?xmltex \def\figurename{Figure}?><label>Figure 1</label><caption><p id="d1e1021"><bold>(a)</bold> <inline-formula><mml:math id="M36" display="inline"><mml:mi>N</mml:mi></mml:math></inline-formula>-dimensional Gaussian distributions with unit variances and
zero means as a function of <inline-formula><mml:math id="M37" display="inline"><mml:mrow><mml:mi>r</mml:mi><mml:mo>=</mml:mo><mml:mo>|</mml:mo><mml:mo>|</mml:mo><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>|</mml:mo><mml:mo>|</mml:mo></mml:mrow></mml:math></inline-formula> for different values of <inline-formula><mml:math id="M38" display="inline"><mml:mi>N</mml:mi></mml:math></inline-formula> (Eq. <xref ref-type="disp-formula" rid="Ch1.E3"/>). The position of the mode goes like <inline-formula><mml:math id="M39" display="inline"><mml:msqrt><mml:mi>N</mml:mi></mml:msqrt></mml:math></inline-formula>
and the width is approximately the constant <inline-formula><mml:math id="M40" display="inline"><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>/</mml:mo><mml:msqrt><mml:mn mathvariant="normal">2</mml:mn></mml:msqrt></mml:mrow></mml:math></inline-formula>. <bold>(b)</bold> The distribution of angles
between pairs of independent <inline-formula><mml:math id="M41" display="inline"><mml:mi>N</mml:mi></mml:math></inline-formula>-dimensional Gaussian vectors for the same values of <inline-formula><mml:math id="M42" display="inline"><mml:mi>N</mml:mi></mml:math></inline-formula>.
Thick curves are calculated in the large ensemble limit. For both <bold>(a)</bold>
and <bold>(b)</bold> the thin dashed curves illustrate the distributions for a sample of size 50.
</p></caption>
          <?xmltex \igopts{width=497.923228pt}?><graphic xlink:href="https://npg.copernicus.org/articles/28/409/2021/npg-28-409-2021-f01.png"/>

        </fig>

</sec>
<sec id="Ch1.S2.SS2">
  <label>2.2</label><title>Concentration of measures</title>
      <?pagebreak page411?><p id="d1e1121">We first look at a very simple example to describe the general idea of
concentration of measures.  Consider <inline-formula><mml:math id="M43" display="inline"><mml:mi>N</mml:mi></mml:math></inline-formula> iid random variables <inline-formula><mml:math id="M44" display="inline"><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>,
<inline-formula><mml:math id="M45" display="inline"><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>,</mml:mo><mml:mn mathvariant="normal">2</mml:mn><mml:mi mathvariant="normal">…</mml:mi><mml:mi>N</mml:mi></mml:mrow></mml:math></inline-formula>, each with mean <inline-formula><mml:math id="M46" display="inline"><mml:mi mathvariant="italic">μ</mml:mi></mml:math></inline-formula> and variance <inline-formula><mml:math id="M47" display="inline"><mml:mrow><mml:msup><mml:mi mathvariant="italic">σ</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula>.  For the
sum of the variables, <inline-formula><mml:math id="M48" display="inline"><mml:mrow><mml:msub><mml:mo>∑</mml:mo><mml:mi>i</mml:mi></mml:msub><mml:msub><mml:mi>x</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>, we have for the expectation and
variance

                <disp-formula id="Ch1.E1" content-type="numbered"><label>1</label><mml:math id="M49" display="block"><mml:mstyle displaystyle="true" class="stylechange"/><mml:mrow><mml:mstyle displaystyle="true" class="stylechange"/><mml:mi>E</mml:mi><mml:mfenced close=")" open="("><mml:mrow><mml:munder><mml:mo movablelimits="false">∑</mml:mo><mml:mi>i</mml:mi></mml:munder><mml:msub><mml:mi>x</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:mfenced><mml:mo>=</mml:mo><mml:munder><mml:mo movablelimits="false">∑</mml:mo><mml:mi>i</mml:mi></mml:munder><mml:mi>E</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:mi>N</mml:mi><mml:mi mathvariant="italic">μ</mml:mi></mml:mrow></mml:math></disp-formula>

          and

                <disp-formula id="Ch1.E2" content-type="numbered"><label>2</label><mml:math id="M50" display="block"><mml:mstyle displaystyle="true" class="stylechange"/><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><mml:mtext>Var</mml:mtext><mml:mfenced open="(" close=")"><mml:mrow><mml:munder><mml:mo movablelimits="false">∑</mml:mo><mml:mi>i</mml:mi></mml:munder><mml:msub><mml:mi>x</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:mfenced><mml:mo>=</mml:mo><mml:munder><mml:mo movablelimits="false">∑</mml:mo><mml:mi>i</mml:mi></mml:munder><mml:mtext>Var</mml:mtext><mml:mo>(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:mi>N</mml:mi><mml:msup><mml:mi mathvariant="italic">σ</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo>.</mml:mo></mml:mrow></mml:math></disp-formula>

          Thus, <inline-formula><mml:math id="M51" display="inline"><mml:mrow><mml:msqrt><mml:mrow><mml:mtext>Var</mml:mtext><mml:mo>(</mml:mo><mml:msub><mml:mo>∑</mml:mo><mml:mi>i</mml:mi></mml:msub><mml:msub><mml:mi>x</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:msqrt><mml:mo>/</mml:mo><mml:mi>E</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mo>∑</mml:mo><mml:mi>i</mml:mi></mml:msub><mml:msub><mml:mi>x</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:msqrt><mml:mstyle displaystyle="false"><mml:mfrac style="text"><mml:mn mathvariant="normal">1</mml:mn><mml:mi>N</mml:mi></mml:mfrac></mml:mstyle></mml:msqrt><mml:mstyle displaystyle="false"><mml:mfrac style="text"><mml:mi mathvariant="italic">μ</mml:mi><mml:mi mathvariant="italic">σ</mml:mi></mml:mfrac></mml:mstyle></mml:mrow></mml:math></inline-formula>.
Therefore, when <inline-formula><mml:math id="M52" display="inline"><mml:mi>N</mml:mi></mml:math></inline-formula> grows, both the expectation of the sum and the width of its distribution will grow, but the relative width will decrease. We
can therefore, with some reason, say that the distribution of the sum
becomes more and more sharply defined around its mean. If we normalize
the sum with <inline-formula><mml:math id="M53" display="inline"><mml:mi>N</mml:mi></mml:math></inline-formula> to get the mean, <inline-formula><mml:math id="M54" display="inline"><mml:mrow><mml:mover accent="true"><mml:mi>x</mml:mi><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:mo>=</mml:mo><mml:msub><mml:mo>∑</mml:mo><mml:mi>i</mml:mi></mml:msub><mml:msub><mml:mi>x</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>/</mml:mo><mml:mi>N</mml:mi></mml:mrow></mml:math></inline-formula>, we have <inline-formula><mml:math id="M55" display="inline"><mml:mrow><mml:mi>E</mml:mi><mml:mo>(</mml:mo><mml:mover accent="true"><mml:mi>x</mml:mi><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:mi mathvariant="italic">μ</mml:mi></mml:mrow></mml:math></inline-formula> and  <inline-formula><mml:math id="M56" display="inline"><mml:mrow><mml:mtext>Var</mml:mtext><mml:mo>(</mml:mo><mml:mover accent="true"><mml:mi>x</mml:mi><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:msup><mml:mi mathvariant="italic">σ</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo>/</mml:mo><mml:mi>N</mml:mi></mml:mrow></mml:math></inline-formula>.
Therefore, the mean becomes increasingly narrowly distributed
around the constant <inline-formula><mml:math id="M57" display="inline"><mml:mi mathvariant="italic">μ</mml:mi></mml:math></inline-formula>. Thus, for large <inline-formula><mml:math id="M58" display="inline"><mml:mi>N</mml:mi></mml:math></inline-formula> we can in many situations
treat the mean, <inline-formula><mml:math id="M59" display="inline"><mml:mover accent="true"><mml:mi>x</mml:mi><mml:mo mathvariant="normal">‾</mml:mo></mml:mover></mml:math></inline-formula>, as a constant.</p>
      <p id="d1e1468">The considerations above are basically the rationale behind the law
of large numbers and are also closely related to the central limit
theorem which states that <inline-formula><mml:math id="M60" display="inline"><mml:mrow><mml:mo>(</mml:mo><mml:mover accent="true"><mml:mi>x</mml:mi><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:mo>-</mml:mo><mml:mi mathvariant="italic">μ</mml:mi><mml:mo>)</mml:mo><mml:mo>/</mml:mo><mml:msqrt><mml:mi>N</mml:mi></mml:msqrt><mml:mo>/</mml:mo><mml:mi mathvariant="italic">σ</mml:mi></mml:mrow></mml:math></inline-formula>
converges towards a standard Gaussian distribution, <inline-formula><mml:math id="M61" display="inline"><mml:mrow><mml:mi mathvariant="script">N</mml:mi><mml:mo>(</mml:mo><mml:mn mathvariant="normal">0</mml:mn><mml:mo>,</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>.
The concentration of measures can be extended beyond the iid situation
(see Sect. <xref ref-type="sec" rid="Ch1.S3"/>) as indicated by the quotation from
<xref ref-type="bibr" rid="bib1.bibx11" id="text.12"/> in the introduction.</p>
      <p id="d1e1522">Let us organize the random variables into an <inline-formula><mml:math id="M62" display="inline"><mml:mi>N</mml:mi></mml:math></inline-formula> vector <inline-formula><mml:math id="M63" display="inline"><mml:mrow><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>=</mml:mo><mml:mo>(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msub><mml:mi mathvariant="normal">…</mml:mi><mml:msub><mml:mi>x</mml:mi><mml:mi>N</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>. Now <inline-formula><mml:math id="M64" display="inline"><mml:mrow><mml:mo>|</mml:mo><mml:mo>|</mml:mo><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>|</mml:mo><mml:msup><mml:mo>|</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo>=</mml:mo><mml:msubsup><mml:mi>x</mml:mi><mml:mn mathvariant="normal">1</mml:mn><mml:mn mathvariant="normal">2</mml:mn></mml:msubsup><mml:mo>+</mml:mo><mml:msubsup><mml:mi>x</mml:mi><mml:mn mathvariant="normal">2</mml:mn><mml:mn mathvariant="normal">2</mml:mn></mml:msubsup><mml:mi mathvariant="normal">…</mml:mi><mml:msubsup><mml:mi>x</mml:mi><mml:mi>N</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msubsup></mml:mrow></mml:math></inline-formula> is
a sum of independent variables and will therefore – according to the
arguments above – for large <inline-formula><mml:math id="M65" display="inline"><mml:mi>N</mml:mi></mml:math></inline-formula> be approximately a constant.</p>
      <p id="d1e1618">Let us consider a multi-variate standard Gaussian distribution
<inline-formula><mml:math id="M66" display="inline"><mml:mrow><mml:mi>P</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:mo>(</mml:mo><mml:mn mathvariant="normal">2</mml:mn><mml:mi mathvariant="italic">π</mml:mi><mml:msup><mml:mo>)</mml:mo><mml:mrow><mml:mo>-</mml:mo><mml:mi>N</mml:mi><mml:mo>/</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:mrow></mml:msup><mml:mi>exp⁡</mml:mi><mml:mo>(</mml:mo><mml:mo>-</mml:mo><mml:msubsup><mml:mo>∑</mml:mo><mml:mrow><mml:mi>n</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mi>N</mml:mi></mml:msubsup><mml:msubsup><mml:mi>x</mml:mi><mml:mi>n</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msubsup><mml:mo>/</mml:mo><mml:mn mathvariant="normal">2</mml:mn><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>.
The surface area of a hyper-sphere with radius <inline-formula><mml:math id="M67" display="inline"><mml:mi>r</mml:mi></mml:math></inline-formula> in <inline-formula><mml:math id="M68" display="inline"><mml:mi>N</mml:mi></mml:math></inline-formula> dimensions is
<inline-formula><mml:math id="M69" display="inline"><mml:mrow><mml:msub><mml:mi>S</mml:mi><mml:mrow><mml:mi>N</mml:mi><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mn mathvariant="normal">2</mml:mn><mml:msup><mml:mi mathvariant="italic">π</mml:mi><mml:mrow><mml:mi>N</mml:mi><mml:mo>/</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:mrow></mml:msup><mml:msup><mml:mi>r</mml:mi><mml:mrow><mml:mi>N</mml:mi><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msup><mml:mo>/</mml:mo><mml:mi mathvariant="normal">Γ</mml:mi><mml:mfenced close=")" open="("><mml:mstyle displaystyle="false"><mml:mfrac style="text"><mml:mi>N</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:mfrac></mml:mstyle></mml:mfenced></mml:mrow></mml:math></inline-formula>.
So, as a function of <inline-formula><mml:math id="M70" display="inline"><mml:mrow><mml:mi>r</mml:mi><mml:mo>=</mml:mo><mml:mo>|</mml:mo><mml:mo>|</mml:mo><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>|</mml:mo><mml:mo>|</mml:mo></mml:mrow></mml:math></inline-formula>, we get the <inline-formula><mml:math id="M71" display="inline"><mml:mi mathvariant="italic">χ</mml:mi></mml:math></inline-formula> distribution

                <disp-formula id="Ch1.E3" content-type="numbered"><label>3</label><mml:math id="M72" display="block"><mml:mstyle class="stylechange" displaystyle="true"/><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><mml:mi>P</mml:mi><mml:mo>(</mml:mo><mml:mi>r</mml:mi><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:msub><mml:mi>S</mml:mi><mml:mrow><mml:mi>N</mml:mi><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msub><mml:mi>P</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:msup><mml:mn mathvariant="normal">2</mml:mn><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>-</mml:mo><mml:mi>N</mml:mi><mml:mo>/</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:mrow></mml:msup></mml:mrow><mml:mrow><mml:mi mathvariant="normal">Γ</mml:mi><mml:mfenced open="(" close=")"><mml:mstyle displaystyle="false"><mml:mfrac style="text"><mml:mi>N</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:mfrac></mml:mstyle></mml:mfenced></mml:mrow></mml:mfrac></mml:mstyle><mml:msup><mml:mi>r</mml:mi><mml:mrow><mml:mi>N</mml:mi><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msup><mml:mi>exp⁡</mml:mi><mml:mo>(</mml:mo><mml:mo>-</mml:mo><mml:msup><mml:mi>r</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo>/</mml:mo><mml:mn mathvariant="normal">2</mml:mn><mml:mo>)</mml:mo><mml:mo>.</mml:mo></mml:mrow></mml:math></disp-formula>

          The maximum of <inline-formula><mml:math id="M73" display="inline"><mml:mrow><mml:mi>P</mml:mi><mml:mo>(</mml:mo><mml:mi>r</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> is reached for <inline-formula><mml:math id="M74" display="inline"><mml:mrow><mml:mi>r</mml:mi><mml:mo>=</mml:mo><mml:msqrt><mml:mrow><mml:mi>N</mml:mi><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msqrt></mml:mrow></mml:math></inline-formula> and the width
(standard deviation) of the peak converges quickly with <inline-formula><mml:math id="M75" display="inline"><mml:mi>N</mml:mi></mml:math></inline-formula> towards <inline-formula><mml:math id="M76" display="inline"><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>/</mml:mo><mml:msqrt><mml:mn mathvariant="normal">2</mml:mn></mml:msqrt></mml:mrow></mml:math></inline-formula>. This is illustrated in Fig. <xref ref-type="fig" rid="Ch1.F1"/>.</p>
      <?pagebreak page412?><p id="d1e1931">The concentration of measures is the backbone of statistical mechanics.
As a simple example, we consider the canonical ensemble of weakly
interacting identical particles. This ensemble describes a system with a
constant number of particles, <inline-formula><mml:math id="M77" display="inline"><mml:mi>N</mml:mi></mml:math></inline-formula>, in a heat bath. All particles have the same spectrum of energy states, <inline-formula><mml:math id="M78" display="inline"><mml:mrow><mml:msub><mml:mi>E</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>, and the probability of a particle
being in the <italic>i</italic>th state is proportional to <inline-formula><mml:math id="M79" display="inline"><mml:mrow><mml:mi>exp⁡</mml:mi><mml:mo>(</mml:mo><mml:mo>-</mml:mo><mml:mi mathvariant="italic">β</mml:mi><mml:msub><mml:mi>E</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>. The total energy grows like <inline-formula><mml:math id="M80" display="inline"><mml:mi>N</mml:mi></mml:math></inline-formula>, while the fluctuations (standard deviation)
in the total energy grow like <inline-formula><mml:math id="M81" display="inline"><mml:msqrt><mml:mi>N</mml:mi></mml:msqrt></mml:math></inline-formula>. Thus, the relative fluctuations in the total energy go as <inline-formula><mml:math id="M82" display="inline"><mml:mrow><mml:mo>∼</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>/</mml:mo><mml:msqrt><mml:mi>N</mml:mi></mml:msqrt></mml:mrow></mml:math></inline-formula>, and in the thermodynamic
limit, <inline-formula><mml:math id="M83" display="inline"><mml:mrow><mml:mi>N</mml:mi><mml:mo>→</mml:mo><mml:mi mathvariant="normal">∞</mml:mi></mml:mrow></mml:math></inline-formula>, these fluctuations and fluctuations
in other macroscopic quantities can be neglected.  This holds also
for non-identical and interacting particles <xref ref-type="bibr" rid="bib1.bibx26" id="paren.13"><named-content content-type="pre">see</named-content><named-content content-type="post">for a recent
discussion</named-content></xref>, just as the concentration of measures can be extended beyond the iid situation.</p>
      <p id="d1e2026">Let us take a brief look at waist concentration. Consider two independent
unit vectors <inline-formula><mml:math id="M84" display="inline"><mml:mi mathvariant="bold-italic">a</mml:mi></mml:math></inline-formula> and <inline-formula><mml:math id="M85" display="inline"><mml:mi mathvariant="bold-italic">b</mml:mi></mml:math></inline-formula>. Without lack of generality we
can set <inline-formula><mml:math id="M86" display="inline"><mml:mrow><mml:mi mathvariant="bold-italic">a</mml:mi><mml:mo>=</mml:mo><mml:mo>(</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>,</mml:mo><mml:mn mathvariant="normal">0</mml:mn><mml:mo>,</mml:mo><mml:mn mathvariant="normal">0</mml:mn><mml:mi mathvariant="normal">…</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>. The dot product then becomes <inline-formula><mml:math id="M87" display="inline"><mml:mrow><mml:msub><mml:mi>b</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula>. It is therefore easy to see that <inline-formula><mml:math id="M88" display="inline"><mml:mrow><mml:mi mathvariant="bold-italic">a</mml:mi><mml:mo>⋅</mml:mo><mml:mi mathvariant="bold-italic">b</mml:mi></mml:mrow></mml:math></inline-formula> has zero
mean and that its spread converges to zero as <inline-formula><mml:math id="M89" display="inline"><mml:mrow><mml:mo>∼</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>/</mml:mo><mml:msqrt><mml:mi>N</mml:mi></mml:msqrt></mml:mrow></mml:math></inline-formula>. This
result does not require Gaussianity; see e.g. <xref ref-type="bibr" rid="bib1.bibx36" id="text.14"/> for a general derivation. The angle <inline-formula><mml:math id="M90" display="inline"><mml:mi mathvariant="italic">ϕ</mml:mi></mml:math></inline-formula> between <inline-formula><mml:math id="M91" display="inline"><mml:mi mathvariant="bold-italic">a</mml:mi></mml:math></inline-formula> and <inline-formula><mml:math id="M92" display="inline"><mml:mi mathvariant="bold-italic">b</mml:mi></mml:math></inline-formula> will therefore converge towards <inline-formula><mml:math id="M93" display="inline"><mml:mrow><mml:mi mathvariant="italic">π</mml:mi><mml:mo>/</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:mrow></mml:math></inline-formula> as <inline-formula><mml:math id="M94" display="inline"><mml:mrow><mml:mi>cos⁡</mml:mi><mml:mi mathvariant="italic">ϕ</mml:mi><mml:mo>=</mml:mo><mml:mi mathvariant="bold-italic">a</mml:mi><mml:mo>⋅</mml:mo><mml:mi mathvariant="bold-italic">b</mml:mi></mml:mrow></mml:math></inline-formula>. This is illustrated in Fig. <xref ref-type="fig" rid="Ch1.F1"/> for
Gaussian-distributed vectors for different values of <inline-formula><mml:math id="M95" display="inline"><mml:mi>N</mml:mi></mml:math></inline-formula>: for <inline-formula><mml:math id="M96" display="inline"><mml:mrow><mml:mi>N</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:mrow></mml:math></inline-formula> the distribution of angles is flat, but for larger values of <inline-formula><mml:math id="M97" display="inline"><mml:mi>N</mml:mi></mml:math></inline-formula> it
becomes increasingly peaked around <inline-formula><mml:math id="M98" display="inline"><mml:mrow><mml:mi mathvariant="italic">π</mml:mi><mml:mo>/</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:mrow></mml:math></inline-formula>.</p>
      <p id="d1e2203">The topic of concentration properties is an active mathematical
field with focus on probabilistic bounds on how quickly empirical means converge to the ensemble means for different classes of random variables,
including non-iid variables <xref ref-type="bibr" rid="bib1.bibx49 bib1.bibx51" id="paren.15"/>. Such
bounds include Bernstein's and Hoeffding's inequalities and give strict
mathematical meaning to the looser considerations above.  As an example, the Hoeffding bound states that for all <inline-formula><mml:math id="M99" display="inline"><mml:mrow><mml:mi>t</mml:mi><mml:mo>≥</mml:mo><mml:mn mathvariant="normal">0</mml:mn></mml:mrow></mml:math></inline-formula>,

                <disp-formula id="Ch1.E4" content-type="numbered"><label>4</label><mml:math id="M100" display="block"><mml:mstyle displaystyle="true" class="stylechange"/><mml:mrow><mml:mstyle displaystyle="true" class="stylechange"/><mml:mi>P</mml:mi><mml:mfenced open="(" close=")"><mml:mrow><mml:mfenced open="|" close="|"><mml:mrow><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mn mathvariant="normal">1</mml:mn><mml:mi>N</mml:mi></mml:mfrac></mml:mstyle><mml:mo movablelimits="false">∑</mml:mo><mml:mo>(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>-</mml:mo><mml:msub><mml:mi mathvariant="italic">μ</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:mfenced><mml:mo>≥</mml:mo><mml:mi>t</mml:mi></mml:mrow></mml:mfenced><mml:mo>≤</mml:mo><mml:mi>exp⁡</mml:mi><mml:mo>(</mml:mo><mml:mo>-</mml:mo><mml:mi>N</mml:mi><mml:msup><mml:mi>t</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo>/</mml:mo><mml:mo>(</mml:mo><mml:mn mathvariant="normal">2</mml:mn><mml:msup><mml:mi mathvariant="italic">σ</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo>)</mml:mo><mml:mo>)</mml:mo><mml:mo>.</mml:mo></mml:mrow></mml:math></disp-formula>

          This holds for independent random variables drawn from a distribution which tails decay at
least as quickly as the tails of a Gaussian distribution <xref ref-type="bibr" rid="bib1.bibx51" id="paren.16"><named-content content-type="pre">sub-Gaussian,</named-content></xref>. Here <inline-formula><mml:math id="M101" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">μ</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> is the mean of <inline-formula><mml:math id="M102" display="inline"><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>
and <inline-formula><mml:math id="M103" display="inline"><mml:mi mathvariant="italic">σ</mml:mi></mml:math></inline-formula> is a constant.  For the angle <inline-formula><mml:math id="M104" display="inline"><mml:mi mathvariant="italic">ϕ</mml:mi></mml:math></inline-formula> between two independent
vectors, we have similarly for all <inline-formula><mml:math id="M105" display="inline"><mml:mrow><mml:mi>t</mml:mi><mml:mo>≥</mml:mo><mml:mn mathvariant="normal">0</mml:mn></mml:mrow></mml:math></inline-formula> <xref ref-type="bibr" rid="bib1.bibx26" id="paren.17"><named-content content-type="pre">from</named-content></xref>

                <disp-formula id="Ch1.E5" content-type="numbered"><label>5</label><mml:math id="M106" display="block"><mml:mstyle class="stylechange" displaystyle="true"/><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><mml:mi>P</mml:mi><mml:mfenced open="(" close=")"><mml:mrow><mml:mfenced open="|" close="|"><mml:mrow><mml:mi>cos⁡</mml:mi><mml:mi mathvariant="italic">ϕ</mml:mi></mml:mrow></mml:mfenced><mml:mo>≥</mml:mo><mml:mi>t</mml:mi></mml:mrow></mml:mfenced><mml:mo>≤</mml:mo><mml:mn mathvariant="normal">2</mml:mn><mml:mi>exp⁡</mml:mi><mml:mo>(</mml:mo><mml:mo>-</mml:mo><mml:mi>N</mml:mi><mml:msup><mml:mi>t</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo>/</mml:mo><mml:mn mathvariant="normal">2</mml:mn><mml:mo>)</mml:mo><mml:mo>.</mml:mo></mml:mrow></mml:math></disp-formula>

          Such relations are also central to the related field of large-deviation theory, which specifically studies the exponential decay of probabilities of large
fluctuations. See <xref ref-type="bibr" rid="bib1.bibx47" id="text.18"/> for a general review and
<xref ref-type="bibr" rid="bib1.bibx24" id="text.19"/> for an application to weather extremes.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F2" specific-use="star"><?xmltex \currentcnt{2}?><?xmltex \def\figurename{Figure}?><label>Figure 2</label><caption><p id="d1e2412">The AgERA data set. Daily means with 5 d separation for June 1980–1990. <bold>(a, c)</bold> Distribution of the lengths. <bold>(b, d)</bold> Distribution of the angles.
<bold>(a, b)</bold> Near-surface temperature (K). <bold>(c, d)</bold> Precipitation (mm <inline-formula><mml:math id="M107" display="inline"><mml:mrow class="unit"><mml:msup><mml:mi mathvariant="normal">d</mml:mi><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula>).
</p></caption>
          <?xmltex \igopts{width=497.923228pt}?><graphic xlink:href="https://npg.copernicus.org/articles/28/409/2021/npg-28-409-2021-f02.png"/>

        </fig>

</sec>
</sec>
<sec id="Ch1.S3">
  <label>3</label><title>Extension to situations with dependent and non-identical variables</title>
      <p id="d1e2456">Like the central limit theorem (CLT), the concentration properties are
originally developed for iid variables. However, also as the central limit
theorem, they can be extended to classes of dependent variables. Although
no general condition exists for the CLT <xref ref-type="bibr" rid="bib1.bibx17" id="paren.20"/>, an important
factor for both the CLT and the concentration properties is the strength
of the dependence <xref ref-type="bibr" rid="bib1.bibx35 bib1.bibx11" id="paren.21"/>.  Many properties
of iid processes can be extended to processes where the rate of mixing
is strong enough <xref ref-type="bibr" rid="bib1.bibx11" id="paren.22"/>.  Here, mixing processes are
defined by a decay of correlations towards zero; i.e. <inline-formula><mml:math id="M108" display="inline"><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M109" display="inline"><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> should become independent when  <inline-formula><mml:math id="M110" display="inline"><mml:mrow><mml:mo>|</mml:mo><mml:mi>i</mml:mi><mml:mo>-</mml:mo><mml:mi>j</mml:mi><mml:mo>|</mml:mo></mml:mrow></mml:math></inline-formula> increases.</p>
      <p id="d1e2507">Here correlations generally refer to measures of the dependence, e.g. the distance between the joint distribution and the product of the marginal
distributions. Note that the decay of Pearson's correlation coefficient is not necessarily sufficient as a zero correlation coefficient does
not guarantee independence as it only gauges linear dependence. The
auto-regressive moving-average (ARMA) models which are often used in
geophysics are examples of mixing processes <xref ref-type="bibr" rid="bib1.bibx40" id="paren.23"/>. More
generally, <xref ref-type="bibr" rid="bib1.bibx11" id="text.24"/> finds that the concentration of measures holds for a random variable that smoothly depends on the influence of
many weakly dependent random variables.</p>
      <p id="d1e2516">The mixing and decay of correlations are closely related
to the concept of effective degrees of freedom also known
as the effective dimension <xref ref-type="bibr" rid="bib1.bibx17" id="paren.25"><named-content content-type="pre">e.g.</named-content></xref>.  Shalizi (2006)<fn id="Ch1.Footn2"><p id="d1e2524"><uri>https://www.stat.cmu.edu/~cshalizi/754/2006/notes/lecture-27.pdf</uri> (last access: 23 August 2021)</p></fn>
shows an example of the CLT for dependent variables, “only with the true sample size replaced by an effective sample size.”  The basic idea
is that dependent variables of effective dimension <inline-formula><mml:math id="M111" display="inline"><mml:mrow><mml:msup><mml:mi>N</mml:mi><mml:mo>*</mml:mo></mml:msup></mml:mrow></mml:math></inline-formula> gives the same
information as independent variables of dimension <inline-formula><mml:math id="M112" display="inline"><mml:mi>N</mml:mi></mml:math></inline-formula>.  Heuristically,
consider a function in a two-dimensional square region with each side of length <inline-formula><mml:math id="M113" display="inline"><mml:mi>L</mml:mi></mml:math></inline-formula>. If correlations decay exponentially with characteristic length <inline-formula><mml:math id="M114" display="inline"><mml:mi mathvariant="italic">ξ</mml:mi></mml:math></inline-formula>, then we can to a first approximation describe the function
by <inline-formula><mml:math id="M115" display="inline"><mml:mrow><mml:msup><mml:mi>N</mml:mi><mml:mo>*</mml:mo></mml:msup><mml:mo>=</mml:mo><mml:mo>(</mml:mo><mml:mi>L</mml:mi><mml:mo>/</mml:mo><mml:mi mathvariant="italic">ξ</mml:mi><mml:msup><mml:mo>)</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula> independent variables.  For fixed <inline-formula><mml:math id="M116" display="inline"><mml:mi mathvariant="italic">ξ</mml:mi></mml:math></inline-formula> the number
of independent variables go to infinity with increasing <inline-formula><mml:math id="M117" display="inline"><mml:mi>L</mml:mi></mml:math></inline-formula>, and in this situation we may assume that the limit theorems hold.  Note that some methods to calculate the number of effective dimensions of e.g. surface temperature are directly based on these arguments using an average <inline-formula><mml:math id="M118" display="inline"><mml:mi mathvariant="italic">ξ</mml:mi></mml:math></inline-formula>
<xref ref-type="bibr" rid="bib1.bibx16" id="paren.26"><named-content content-type="pre">see the summary in</named-content></xref>.</p>
      <?pagebreak page413?><p id="d1e2615">The situation is well known in the study of one-dimensional time series <xref ref-type="bibr" rid="bib1.bibx50" id="paren.27"><named-content content-type="pre">see e.g.</named-content><named-content content-type="post">Sect. 17</named-content></xref>.  As a simple example we consider a time series, <inline-formula><mml:math id="M119" display="inline"><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>, of length <inline-formula><mml:math id="M120" display="inline"><mml:mi>N</mml:mi></mml:math></inline-formula>  generated with a first-order auto-regressive, AR(1), process with coefficient <inline-formula><mml:math id="M121" display="inline"><mml:mi mathvariant="italic">ρ</mml:mi></mml:math></inline-formula>. The
auto-correlations behave as <inline-formula><mml:math id="M122" display="inline"><mml:mrow><mml:mo>∼</mml:mo><mml:msup><mml:mi mathvariant="italic">ρ</mml:mi><mml:mi>r</mml:mi></mml:msup></mml:mrow></mml:math></inline-formula>, where <inline-formula><mml:math id="M123" display="inline"><mml:mi>r</mml:mi></mml:math></inline-formula> is the lag. A
decorrelation time, <inline-formula><mml:math id="M124" display="inline"><mml:mi mathvariant="italic">τ</mml:mi></mml:math></inline-formula>, can be found where the auto-correlations
have decayed to <inline-formula><mml:math id="M125" display="inline"><mml:mrow><mml:msup><mml:mi>e</mml:mi><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula>: <inline-formula><mml:math id="M126" display="inline"><mml:mrow><mml:mi mathvariant="italic">τ</mml:mi><mml:mo>=</mml:mo><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>/</mml:mo><mml:mi>ln⁡</mml:mi><mml:mi mathvariant="italic">ρ</mml:mi></mml:mrow></mml:math></inline-formula>.  The effective degrees
of freedom would therefore be <inline-formula><mml:math id="M127" display="inline"><mml:mrow><mml:mo>-</mml:mo><mml:mi>N</mml:mi><mml:mi>ln⁡</mml:mi><mml:mi mathvariant="italic">ρ</mml:mi></mml:mrow></mml:math></inline-formula>.  Less heuristically, we
have for the ensemble mean, <inline-formula><mml:math id="M128" display="inline"><mml:mrow><mml:mover accent="true"><mml:mi>x</mml:mi><mml:mo mathvariant="normal">¯</mml:mo></mml:mover><mml:mo>=</mml:mo><mml:msub><mml:mo>∑</mml:mo><mml:mi>i</mml:mi></mml:msub><mml:msub><mml:mi>x</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>/</mml:mo><mml:mi>N</mml:mi></mml:mrow></mml:math></inline-formula>, that <inline-formula><mml:math id="M129" display="inline"><mml:mrow><mml:msqrt><mml:mi>N</mml:mi></mml:msqrt><mml:mover accent="true"><mml:mi>x</mml:mi><mml:mo mathvariant="normal">¯</mml:mo></mml:mover><mml:mo>∼</mml:mo><mml:mi mathvariant="script">N</mml:mi><mml:mo>(</mml:mo><mml:mn mathvariant="normal">0</mml:mn><mml:mo>,</mml:mo><mml:msup><mml:mi mathvariant="italic">σ</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo>(</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>+</mml:mo><mml:mi mathvariant="italic">ρ</mml:mi><mml:mo>)</mml:mo><mml:mo>/</mml:mo><mml:mo>(</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>-</mml:mo><mml:mi mathvariant="italic">ρ</mml:mi><mml:mo>)</mml:mo><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>. Compared with the similar results for iid Gaussian variables, <inline-formula><mml:math id="M130" display="inline"><mml:mrow><mml:msqrt><mml:mi>N</mml:mi></mml:msqrt><mml:mover accent="true"><mml:mi>x</mml:mi><mml:mo mathvariant="normal">¯</mml:mo></mml:mover><mml:mo>∼</mml:mo><mml:mi mathvariant="script">N</mml:mi><mml:mo>(</mml:mo><mml:mn mathvariant="normal">0</mml:mn><mml:mo>,</mml:mo><mml:msup><mml:mi mathvariant="italic">σ</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>, this suggests an effective dimension
of <inline-formula><mml:math id="M131" display="inline"><mml:mrow><mml:mi>N</mml:mi><mml:mo>(</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>-</mml:mo><mml:mi mathvariant="italic">ρ</mml:mi><mml:mo>)</mml:mo><mml:mo>/</mml:mo><mml:mo>(</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>+</mml:mo><mml:mi mathvariant="italic">ρ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> <xref ref-type="bibr" rid="bib1.bibx18 bib1.bibx50" id="paren.28"/>.  For the
sum of squares we have <inline-formula><mml:math id="M132" display="inline"><mml:mrow><mml:msqrt><mml:mi>N</mml:mi></mml:msqrt><mml:mover accent="true"><mml:mrow><mml:msup><mml:mi>x</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow><mml:mo mathvariant="normal">¯</mml:mo></mml:mover><mml:mo>∼</mml:mo><mml:mi mathvariant="script">N</mml:mi><mml:mo>(</mml:mo><mml:msup><mml:mi mathvariant="italic">σ</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo>,</mml:mo><mml:mn mathvariant="normal">2</mml:mn><mml:msup><mml:mi mathvariant="italic">σ</mml:mi><mml:mn mathvariant="normal">4</mml:mn></mml:msup><mml:mo>(</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>+</mml:mo><mml:msup><mml:mi mathvariant="italic">ρ</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo>)</mml:mo><mml:mo>/</mml:mo><mml:mo>(</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>-</mml:mo><mml:msup><mml:mi mathvariant="italic">ρ</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo>)</mml:mo><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>, now suggesting an effective dimension
of <inline-formula><mml:math id="M133" display="inline"><mml:mrow><mml:mi>N</mml:mi><mml:mo>(</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>-</mml:mo><mml:msup><mml:mi mathvariant="italic">ρ</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo>)</mml:mo><mml:mo>/</mml:mo><mml:mo>(</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>+</mml:mo><mml:msup><mml:mi mathvariant="italic">ρ</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> <xref ref-type="bibr" rid="bib1.bibx3" id="paren.29"/>.  We note that the effective dimension <inline-formula><mml:math id="M134" display="inline"><mml:mrow><mml:msup><mml:mi>N</mml:mi><mml:mo>*</mml:mo></mml:msup></mml:mrow></mml:math></inline-formula> depends on the measure of interest.</p>
      <p id="d1e2995">In the case of two-dimensional fields, different methods exist to estimate the number of effective dimensions <inline-formula><mml:math id="M135" display="inline"><mml:mrow><mml:msup><mml:mi>N</mml:mi><mml:mo>*</mml:mo></mml:msup></mml:mrow></mml:math></inline-formula> <xref ref-type="bibr" rid="bib1.bibx53 bib1.bibx9" id="paren.30"/>. Some methods are directly based on the characteristic length, <inline-formula><mml:math id="M136" display="inline"><mml:mi mathvariant="italic">ξ</mml:mi></mml:math></inline-formula>,
using an average over the different directions <xref ref-type="bibr" rid="bib1.bibx16" id="paren.31"/>.
The estimated number depends both on the method used and on the field, the timescale, and the geographical region.  For the annual mean surface temperature values of <inline-formula><mml:math id="M137" display="inline"><mml:mrow><mml:msup><mml:mi>N</mml:mi><mml:mo>*</mml:mo></mml:msup></mml:mrow></mml:math></inline-formula> vary between 50 and 100 depending
on the method <xref ref-type="bibr" rid="bib1.bibx10 bib1.bibx28 bib1.bibx44" id="paren.32"/> when the whole globe is considered. Values in the same range have been found for monthly surface
temperatures in the Northern Hemisphere <xref ref-type="bibr" rid="bib1.bibx53 bib1.bibx9" id="paren.33"/>.</p>
      <p id="d1e3040">These numbers are of course small compared to Avogadro's number relevant for statistical mechanics, but they are still comparable to the dimensions
in Fig. <xref ref-type="fig" rid="Ch1.F1"/> where the concentration properties hold
to a reasonable degree.  In Sect. <xref ref-type="sec" rid="Ch1.S4.SS1"/> we directly investigate to which extent the concentration properties hold for atmospheric fields.</p>
</sec>
<sec id="Ch1.S4">
  <label>4</label><title>Atmospheric and climate science</title>
      <p id="d1e3055">As we saw in Sect. <xref ref-type="sec" rid="Ch1.S2"/>, concentration of measures
and waist concentration allow us in high dimensions to set dot products
of independent vectors to zero and substitute the length of a random
vector with its expectation value. In Sect. <xref ref-type="sec" rid="Ch1.S3"/>
we argued that when the components of the fields or time series are dependent, the concentration phenomena hold when the effective dimension is large.  However, to test the concentration properties, we also need
independent samples.</p>
      <p id="d1e3062">For initial condition ensembles consisting of experiments with the same
model but with different initial conditions, the different ensemble
members can be considered independent (considering anomalies with
respect to the ensemble centre as explained in the next subsection). For multi-model ensembles where experiments are performed with models
with different physical parameterizations (but the same external
forcings), the situation is more complicated <xref ref-type="bibr" rid="bib1.bibx34 bib1.bibx8 bib1.bibx15" id="paren.34"><named-content content-type="pre">e.g.</named-content><named-content content-type="post">and references therein</named-content></xref>. The annual or monthly climatologies are obvious measures for comparing models or for validating
the models against observations <xref ref-type="bibr" rid="bib1.bibx25" id="paren.35"/>. Another used measure
is the forced response in e.g. time series of global means.</p>
      <p id="d1e3075">Another way to obtain independent samples from the same distribution is to
consider a given variable at different times. For example, we could look
at the spatial field of precipitation or temperature at different days or
months. To ensure that the fields are drawn from the same distribution, we need to avoid or remove the annual cycle and – if longer periods
are considered – to make sure that there is no external forcing.
The sample times should also be sufficiently separated.</p>
      <p id="d1e3078">In the next sections we will consider the following geophysical
data sets. (1)  Daily means of near-surface temperature and precipitation from AgERA for June in the period 1980–1990. The
AgERA provides daily<?pagebreak page414?> surface meteorological data for agro-ecological
studies (doi: 10.24381/cds.6c68c9bb) based on ECMWF's ERA5 reanalysis
<xref ref-type="bibr" rid="bib1.bibx31" id="paren.36"/>.  AgERA is land-only and of high resolution, with more than 2 million (<inline-formula><mml:math id="M138" display="inline"><mml:mrow><mml:mi>N</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">2</mml:mn><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mn mathvariant="normal">353</mml:mn><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mn mathvariant="normal">526</mml:mn></mml:mrow></mml:math></inline-formula>) grid points. Using daily means taken every fifth day for June in 11 <inline-formula><mml:math id="M139" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">years</mml:mi></mml:mrow></mml:math></inline-formula>, we have a sample size
of 66.  (2) Monthly near-surface temperature from the multi-model
CMIP5 ensemble <xref ref-type="bibr" rid="bib1.bibx22" id="paren.37"/> consisting of 45 (the sample
size) historical experiments. The models are identified in Table 1 of <xref ref-type="bibr" rid="bib1.bibx15" id="text.38"/>. (3) The Max Planck Institute Grand
Ensemble <xref ref-type="bibr" rid="bib1.bibx38" id="paren.39"><named-content content-type="pre">MPI-GE,</named-content></xref> consisting of 100 (the sample size)
members differing only in initial conditions.  From MPI-GE we consider the
monthly mean near-surface temperature and precipitation.  For both model
ensembles we consider the monthly climatology in the period 1980–2005 and
the annual Northern Hemisphere (NH) mean values in the period 1961–2005.</p>
      <p id="d1e3123">In addition to the geophysical data, we also include two simple samples of independent vectors. The first sample consists of independent vectors
drawn from an <inline-formula><mml:math id="M140" display="inline"><mml:mi>N</mml:mi></mml:math></inline-formula>-dimensional spherical (all components have zero mean and
unit variance) Gaussian distribution as in Fig. <xref ref-type="fig" rid="Ch1.F1"/>. The
second sample is drawn from a standard Gamma distribution with shape
parameter <inline-formula><mml:math id="M141" display="inline"><mml:mn mathvariant="normal">3</mml:mn></mml:math></inline-formula> (location and scale parameters 0 and 1).  In the latter
case we include anisotropy (not identically distributed components) by
multiplying the <italic>n</italic>th component by <inline-formula><mml:math id="M142" display="inline"><mml:mrow><mml:mn mathvariant="normal">5</mml:mn><mml:mi>n</mml:mi><mml:mo>/</mml:mo><mml:mi>N</mml:mi></mml:mrow></mml:math></inline-formula>, so the mean and variance of the <italic>n</italic>th component become <inline-formula><mml:math id="M143" display="inline"><mml:mrow><mml:mn mathvariant="normal">15</mml:mn><mml:mi>n</mml:mi><mml:mo>/</mml:mo><mml:mi>N</mml:mi></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M144" display="inline"><mml:mrow><mml:mn mathvariant="normal">75</mml:mn><mml:mo>(</mml:mo><mml:mi>n</mml:mi><mml:mo>/</mml:mo><mml:mi>N</mml:mi><mml:msup><mml:mo>)</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula>, respectively.
For the simple random variables we let the dimension, <inline-formula><mml:math id="M145" display="inline"><mml:mi>N</mml:mi></mml:math></inline-formula>, vary from 1
to 100.  The sample size is chosen to 50.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F3" specific-use="star"><?xmltex \currentcnt{3}?><?xmltex \def\figurename{Figure}?><label>Figure 3</label><caption><p id="d1e3207">The multi-model 45-member CMIP5 ensemble. Monthly climatology in TAS (K). <bold>(a)</bold> Distribution of the lengths. <bold>(b)</bold> Distribution of the angles.
</p></caption>
        <?xmltex \igopts{width=469.470472pt}?><graphic xlink:href="https://npg.copernicus.org/articles/28/409/2021/npg-28-409-2021-f03.png"/>

      </fig>

      <?xmltex \floatpos{t}?><fig id="Ch1.F4" specific-use="star"><?xmltex \currentcnt{4}?><?xmltex \def\figurename{Figure}?><label>Figure 4</label><caption><p id="d1e3224">The normalized distances as a  function of dimension <inline-formula><mml:math id="M146" display="inline"><mml:mi>N</mml:mi></mml:math></inline-formula>. For each <inline-formula><mml:math id="M147" display="inline"><mml:mi>N</mml:mi></mml:math></inline-formula> we draw 50 <inline-formula><mml:math id="M148" display="inline"><mml:mi>N</mml:mi></mml:math></inline-formula>-dimensional random vectors and calculate the
pairwise distances <inline-formula><mml:math id="M149" display="inline"><mml:msqrt><mml:mrow><mml:mo>|</mml:mo><mml:mo>|</mml:mo><mml:msup><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>k</mml:mi></mml:msup><mml:mo>-</mml:mo><mml:msup><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>l</mml:mi></mml:msup><mml:mo>|</mml:mo><mml:msup><mml:mo>|</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo>/</mml:mo><mml:mi>N</mml:mi></mml:mrow></mml:msqrt></mml:math></inline-formula> (blue) and the distances to the sample mean <inline-formula><mml:math id="M150" display="inline"><mml:msqrt><mml:mrow><mml:mo>|</mml:mo><mml:mo>|</mml:mo><mml:msup><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>k</mml:mi></mml:msup><mml:mo>-</mml:mo><mml:mover accent="true"><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:mo>|</mml:mo><mml:msup><mml:mo>|</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo>/</mml:mo><mml:mi>N</mml:mi></mml:mrow></mml:msqrt></mml:math></inline-formula> (black). The thick curves show the mean of the distances and the broken curves the mean <inline-formula><mml:math id="M151" display="inline"><mml:mo>±</mml:mo></mml:math></inline-formula>2 standard deviations. In <bold>(a)</bold> each
component of the vectors is drawn from a standard Gaussian; in <bold>(b)</bold> each component is drawn from a standard Gamma distribution with shape parameter <inline-formula><mml:math id="M152" display="inline"><mml:mn mathvariant="normal">3</mml:mn></mml:math></inline-formula> (location and scale parameters 0 and 1). In the
latter case we include  anisotropy by multiplying the <italic>n</italic>th component by <inline-formula><mml:math id="M153" display="inline"><mml:mrow><mml:mn mathvariant="normal">5</mml:mn><mml:mi>n</mml:mi><mml:mo>/</mml:mo><mml:mi>N</mml:mi></mml:mrow></mml:math></inline-formula>, so the mean and variance of the <italic>n</italic>th component become <inline-formula><mml:math id="M154" display="inline"><mml:mrow><mml:mn mathvariant="normal">15</mml:mn><mml:mi>n</mml:mi><mml:mo>/</mml:mo><mml:mi>N</mml:mi></mml:mrow></mml:math></inline-formula>
and <inline-formula><mml:math id="M155" display="inline"><mml:mrow><mml:mn mathvariant="normal">75</mml:mn><mml:mo>(</mml:mo><mml:mi>n</mml:mi><mml:mo>/</mml:mo><mml:mi>N</mml:mi><mml:msup><mml:mo>)</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula>.  Note the factor of <inline-formula><mml:math id="M156" display="inline"><mml:msqrt><mml:mn mathvariant="normal">2</mml:mn></mml:msqrt></mml:math></inline-formula> between the mean distances.
</p></caption>
        <?xmltex \igopts{width=469.470472pt}?><graphic xlink:href="https://npg.copernicus.org/articles/28/409/2021/npg-28-409-2021-f04.png"/>

      </fig>

<?xmltex \floatpos{t}?><table-wrap id="Ch1.T2" specific-use="star"><?xmltex \currentcnt{2}?><label>Table 2</label><caption><p id="d1e3410">Summary of the different measures. Entries show mean/standard deviation.
Units are K for temperature and <inline-formula><mml:math id="M157" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">mm</mml:mi><mml:mspace width="0.125em" linebreak="nobreak"/><mml:msup><mml:mi mathvariant="normal">d</mml:mi><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula> for precipitation.</p></caption><oasis:table frame="topbot"><?xmltex \begin{scaleboxenv}{.9}[.9]?><oasis:tgroup cols="9">
     <oasis:colspec colnum="1" colname="col1" align="left"/>
     <oasis:colspec colnum="2" colname="col2" align="right"/>
     <oasis:colspec colnum="3" colname="col3" align="center"/>
     <oasis:colspec colnum="4" colname="col4" align="center"/>
     <oasis:colspec colnum="5" colname="col5" align="center"/>
     <oasis:colspec colnum="6" colname="col6" align="center"/>
     <oasis:colspec colnum="7" colname="col7" align="center"/>
     <oasis:colspec colnum="8" colname="col8" align="center"/>
     <oasis:colspec colnum="9" colname="col9" align="left"/>
     <oasis:thead>
       <oasis:row>
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2">No. of</oasis:entry>
         <oasis:entry colname="col3">Lengths</oasis:entry>
         <oasis:entry colname="col4">Angles</oasis:entry>
         <oasis:entry colname="col5">Distances between</oasis:entry>
         <oasis:entry colname="col6">Distances between</oasis:entry>
         <oasis:entry colname="col7">Correlations</oasis:entry>
         <oasis:entry colname="col8">Correlations between</oasis:entry>
         <oasis:entry colname="col9"/>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2">samples</oasis:entry>
         <oasis:entry colname="col3"/>
         <oasis:entry colname="col4"/>
         <oasis:entry colname="col5">pairs of ensemble</oasis:entry>
         <oasis:entry colname="col6">ensemble members</oasis:entry>
         <oasis:entry colname="col7">between pairs</oasis:entry>
         <oasis:entry colname="col8">ensemble members</oasis:entry>
         <oasis:entry colname="col9"/>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2"/>
         <oasis:entry colname="col3"/>
         <oasis:entry colname="col4"/>
         <oasis:entry colname="col5">members</oasis:entry>
         <oasis:entry colname="col6">and ensemble</oasis:entry>
         <oasis:entry colname="col7">of ensemble</oasis:entry>
         <oasis:entry colname="col8">and ensemble</oasis:entry>
         <oasis:entry colname="col9"/>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2"/>
         <oasis:entry colname="col3"/>
         <oasis:entry colname="col4"/>
         <oasis:entry colname="col5"/>
         <oasis:entry colname="col6">mean</oasis:entry>
         <oasis:entry colname="col7">members</oasis:entry>
         <oasis:entry colname="col8">mean</oasis:entry>
         <oasis:entry colname="col9"/>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row>
         <oasis:entry colname="col1">AgERA, temperature,</oasis:entry>
         <oasis:entry colname="col2">66</oasis:entry>
         <oasis:entry colname="col3"><inline-formula><mml:math id="M158" display="inline"><mml:mrow><mml:mn mathvariant="normal">4.36</mml:mn><mml:mo>/</mml:mo><mml:mn mathvariant="normal">0.58</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col4"><inline-formula><mml:math id="M159" display="inline"><mml:mrow><mml:mn mathvariant="normal">1.59</mml:mn><mml:mo>/</mml:mo><mml:mn mathvariant="normal">0.21</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col5"><inline-formula><mml:math id="M160" display="inline"><mml:mrow><mml:mn mathvariant="normal">6.21</mml:mn><mml:mo>/</mml:mo><mml:mn mathvariant="normal">0.88</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col6"><inline-formula><mml:math id="M161" display="inline"><mml:mrow><mml:mn mathvariant="normal">4.36</mml:mn><mml:mo>/</mml:mo><mml:mn mathvariant="normal">0.58</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col7"><inline-formula><mml:math id="M162" display="inline"><mml:mrow><mml:mn mathvariant="normal">0.41</mml:mn><mml:mo>/</mml:mo><mml:mn mathvariant="normal">0.13</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col8"><inline-formula><mml:math id="M163" display="inline"><mml:mrow><mml:mn mathvariant="normal">0.64</mml:mn><mml:mo>/</mml:mo><mml:mn mathvariant="normal">0.07</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col9"/>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">daily, June</oasis:entry>
         <oasis:entry colname="col2"/>
         <oasis:entry colname="col3"/>
         <oasis:entry colname="col4"/>
         <oasis:entry colname="col5"/>
         <oasis:entry colname="col6"/>
         <oasis:entry colname="col7"/>
         <oasis:entry colname="col8"/>
         <oasis:entry colname="col9"/>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">AgERA, precipitation,</oasis:entry>
         <oasis:entry colname="col2">66</oasis:entry>
         <oasis:entry colname="col3"><inline-formula><mml:math id="M164" display="inline"><mml:mrow><mml:mn mathvariant="normal">4.34</mml:mn><mml:mo>/</mml:mo><mml:mn mathvariant="normal">0.44</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col4"><inline-formula><mml:math id="M165" display="inline"><mml:mrow><mml:mn mathvariant="normal">1.59</mml:mn><mml:mo>/</mml:mo><mml:mn mathvariant="normal">0.05</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col5"><inline-formula><mml:math id="M166" display="inline"><mml:mrow><mml:mn mathvariant="normal">6.20</mml:mn><mml:mo>/</mml:mo><mml:mn mathvariant="normal">0.47</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col6"><inline-formula><mml:math id="M167" display="inline"><mml:mrow><mml:mn mathvariant="normal">4.34</mml:mn><mml:mo>/</mml:mo><mml:mn mathvariant="normal">0.44</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col7"><inline-formula><mml:math id="M168" display="inline"><mml:mrow><mml:mn mathvariant="normal">0.46</mml:mn><mml:mo>/</mml:mo><mml:mn mathvariant="normal">0.05</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col8"><inline-formula><mml:math id="M169" display="inline"><mml:mrow><mml:mn mathvariant="normal">0.69</mml:mn><mml:mo>/</mml:mo><mml:mn mathvariant="normal">0.04</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col9"/>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">daily, June</oasis:entry>
         <oasis:entry colname="col2"/>
         <oasis:entry colname="col3"/>
         <oasis:entry colname="col4"/>
         <oasis:entry colname="col5"/>
         <oasis:entry colname="col6"/>
         <oasis:entry colname="col7"/>
         <oasis:entry colname="col8"/>
         <oasis:entry colname="col9"/>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">AgERA, temperature,</oasis:entry>
         <oasis:entry colname="col2">66</oasis:entry>
         <oasis:entry colname="col3"><inline-formula><mml:math id="M170" display="inline"><mml:mrow><mml:mn mathvariant="normal">3.25</mml:mn><mml:mo>/</mml:mo><mml:mn mathvariant="normal">1.06</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col4"><inline-formula><mml:math id="M171" display="inline"><mml:mrow><mml:mn mathvariant="normal">1.58</mml:mn><mml:mo>/</mml:mo><mml:mn mathvariant="normal">0.50</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col5"><inline-formula><mml:math id="M172" display="inline"><mml:mrow><mml:mn mathvariant="normal">4.61</mml:mn><mml:mo>/</mml:mo><mml:mn mathvariant="normal">1.58</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col6"><inline-formula><mml:math id="M173" display="inline"><mml:mrow><mml:mn mathvariant="normal">3.25</mml:mn><mml:mo>/</mml:mo><mml:mn mathvariant="normal">1.06</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col7"><inline-formula><mml:math id="M174" display="inline"><mml:mrow><mml:mn mathvariant="normal">0.35</mml:mn><mml:mo>/</mml:mo><mml:mn mathvariant="normal">0.31</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col8"><inline-formula><mml:math id="M175" display="inline"><mml:mrow><mml:mn mathvariant="normal">0.57</mml:mn><mml:mo>/</mml:mo><mml:mn mathvariant="normal">0.25</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col9"/>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">daily, June, N. Europe</oasis:entry>
         <oasis:entry colname="col2"/>
         <oasis:entry colname="col3"/>
         <oasis:entry colname="col4"/>
         <oasis:entry colname="col5"/>
         <oasis:entry colname="col6"/>
         <oasis:entry colname="col7"/>
         <oasis:entry colname="col8"/>
         <oasis:entry colname="col9"/>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">AgERA, precipitation,</oasis:entry>
         <oasis:entry colname="col2">66</oasis:entry>
         <oasis:entry colname="col3"><inline-formula><mml:math id="M176" display="inline"><mml:mrow><mml:mn mathvariant="normal">4.05</mml:mn><mml:mo>/</mml:mo><mml:mn mathvariant="normal">1.88</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col4"><inline-formula><mml:math id="M177" display="inline"><mml:mrow><mml:mn mathvariant="normal">1.56</mml:mn><mml:mo>/</mml:mo><mml:mn mathvariant="normal">0.28</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col5"><inline-formula><mml:math id="M178" display="inline"><mml:mrow><mml:mn mathvariant="normal">5.92</mml:mn><mml:mo>/</mml:mo><mml:mn mathvariant="normal">2.32</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col6"><inline-formula><mml:math id="M179" display="inline"><mml:mrow><mml:mn mathvariant="normal">4.95</mml:mn><mml:mo>/</mml:mo><mml:mn mathvariant="normal">1.88</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col7"><inline-formula><mml:math id="M180" display="inline"><mml:mrow><mml:mn mathvariant="normal">0.66</mml:mn><mml:mo>/</mml:mo><mml:mn mathvariant="normal">0.17</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col8"><inline-formula><mml:math id="M181" display="inline"><mml:mrow><mml:mn mathvariant="normal">0.82</mml:mn><mml:mo>/</mml:mo><mml:mn mathvariant="normal">0.13</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col9"/>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">daily, June, N. Europe</oasis:entry>
         <oasis:entry colname="col2"/>
         <oasis:entry colname="col3"/>
         <oasis:entry colname="col4"/>
         <oasis:entry colname="col5"/>
         <oasis:entry colname="col6"/>
         <oasis:entry colname="col7"/>
         <oasis:entry colname="col8"/>
         <oasis:entry colname="col9"/>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">CMIP5,  monthly</oasis:entry>
         <oasis:entry colname="col2">45</oasis:entry>
         <oasis:entry colname="col3"><inline-formula><mml:math id="M182" display="inline"><mml:mrow><mml:mn mathvariant="normal">2.57</mml:mn><mml:mo>/</mml:mo><mml:mn mathvariant="normal">0.46</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col4"><inline-formula><mml:math id="M183" display="inline"><mml:mrow><mml:mn mathvariant="normal">1.59</mml:mn><mml:mo>/</mml:mo><mml:mn mathvariant="normal">0.28</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col5"><inline-formula><mml:math id="M184" display="inline"><mml:mrow><mml:mn mathvariant="normal">3.66</mml:mn><mml:mo>/</mml:mo><mml:mn mathvariant="normal">0.73</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col6"><inline-formula><mml:math id="M185" display="inline"><mml:mrow><mml:mn mathvariant="normal">2.57</mml:mn><mml:mo>/</mml:mo><mml:mn mathvariant="normal">0.46</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col7"><inline-formula><mml:math id="M186" display="inline"><mml:mrow><mml:mn mathvariant="normal">0.44</mml:mn><mml:mo>/</mml:mo><mml:mn mathvariant="normal">0.18</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col8"><inline-formula><mml:math id="M187" display="inline"><mml:mrow><mml:mn mathvariant="normal">0.64</mml:mn><mml:mo>/</mml:mo><mml:mn mathvariant="normal">0.12</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col9"/>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">climatology,</oasis:entry>
         <oasis:entry colname="col2"/>
         <oasis:entry colname="col3"/>
         <oasis:entry colname="col4"/>
         <oasis:entry colname="col5"/>
         <oasis:entry colname="col6"/>
         <oasis:entry colname="col7"/>
         <oasis:entry colname="col8"/>
         <oasis:entry colname="col9"/>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">precipitation</oasis:entry>
         <oasis:entry colname="col2"/>
         <oasis:entry colname="col3"/>
         <oasis:entry colname="col4"/>
         <oasis:entry colname="col5"/>
         <oasis:entry colname="col6"/>
         <oasis:entry colname="col7"/>
         <oasis:entry colname="col8"/>
         <oasis:entry colname="col9"/>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">MPI-GE, monthly</oasis:entry>
         <oasis:entry colname="col2">100</oasis:entry>
         <oasis:entry colname="col3"><inline-formula><mml:math id="M188" display="inline"><mml:mrow><mml:mn mathvariant="normal">0.34</mml:mn><mml:mo>/</mml:mo><mml:mn mathvariant="normal">0.03</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col4"><inline-formula><mml:math id="M189" display="inline"><mml:mrow><mml:mn mathvariant="normal">1.58</mml:mn><mml:mo>/</mml:mo><mml:mn mathvariant="normal">0.12</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col5"><inline-formula><mml:math id="M190" display="inline"><mml:mrow><mml:mn mathvariant="normal">0.49</mml:mn><mml:mo>/</mml:mo><mml:mn mathvariant="normal">0.04</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col6"><inline-formula><mml:math id="M191" display="inline"><mml:mrow><mml:mn mathvariant="normal">0.34</mml:mn><mml:mo>/</mml:mo><mml:mn mathvariant="normal">0.03</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col7"><inline-formula><mml:math id="M192" display="inline"><mml:mrow><mml:mn mathvariant="normal">0.49</mml:mn><mml:mo>/</mml:mo><mml:mn mathvariant="normal">0.07</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col8"><inline-formula><mml:math id="M193" display="inline"><mml:mrow><mml:mn mathvariant="normal">0.70</mml:mn><mml:mo>/</mml:mo><mml:mn mathvariant="normal">0.04</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col9"/>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">climatology,</oasis:entry>
         <oasis:entry colname="col2"/>
         <oasis:entry colname="col3"/>
         <oasis:entry colname="col4"/>
         <oasis:entry colname="col5"/>
         <oasis:entry colname="col6"/>
         <oasis:entry colname="col7"/>
         <oasis:entry colname="col8"/>
         <oasis:entry colname="col9"/>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">precipitation</oasis:entry>
         <oasis:entry colname="col2"/>
         <oasis:entry colname="col3"/>
         <oasis:entry colname="col4"/>
         <oasis:entry colname="col5"/>
         <oasis:entry colname="col6"/>
         <oasis:entry colname="col7"/>
         <oasis:entry colname="col8"/>
         <oasis:entry colname="col9"/>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">MPI-GE, monthly</oasis:entry>
         <oasis:entry colname="col2">100</oasis:entry>
         <oasis:entry colname="col3"><inline-formula><mml:math id="M194" display="inline"><mml:mrow><mml:mn mathvariant="normal">0.27</mml:mn><mml:mo>/</mml:mo><mml:mn mathvariant="normal">0.01</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col4"><inline-formula><mml:math id="M195" display="inline"><mml:mrow><mml:mn mathvariant="normal">1.58</mml:mn><mml:mo>/</mml:mo><mml:mn mathvariant="normal">0.06</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col5"><inline-formula><mml:math id="M196" display="inline"><mml:mrow><mml:mn mathvariant="normal">0.39</mml:mn><mml:mo>/</mml:mo><mml:mn mathvariant="normal">0.02</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col6"><inline-formula><mml:math id="M197" display="inline"><mml:mrow><mml:mn mathvariant="normal">0.27</mml:mn><mml:mo>/</mml:mo><mml:mn mathvariant="normal">0.01</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col7"><inline-formula><mml:math id="M198" display="inline"><mml:mrow><mml:mn mathvariant="normal">0.51</mml:mn><mml:mo>/</mml:mo><mml:mn mathvariant="normal">0.04</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col8"><inline-formula><mml:math id="M199" display="inline"><mml:mrow><mml:mn mathvariant="normal">0.72</mml:mn><mml:mo>/</mml:mo><mml:mn mathvariant="normal">0.03</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col9"/>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">climatology,</oasis:entry>
         <oasis:entry colname="col2"/>
         <oasis:entry colname="col3"/>
         <oasis:entry colname="col4"/>
         <oasis:entry colname="col5"/>
         <oasis:entry colname="col6"/>
         <oasis:entry colname="col7"/>
         <oasis:entry colname="col8"/>
         <oasis:entry colname="col9"/>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">precipitation</oasis:entry>
         <oasis:entry colname="col2"/>
         <oasis:entry colname="col3"/>
         <oasis:entry colname="col4"/>
         <oasis:entry colname="col5"/>
         <oasis:entry colname="col6"/>
         <oasis:entry colname="col7"/>
         <oasis:entry colname="col8"/>
         <oasis:entry colname="col9"/>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">CMIP5, NH annual</oasis:entry>
         <oasis:entry colname="col2">45</oasis:entry>
         <oasis:entry colname="col3"><inline-formula><mml:math id="M200" display="inline"><mml:mrow><mml:mn mathvariant="normal">0.67</mml:mn><mml:mo>/</mml:mo><mml:mn mathvariant="normal">0.35</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col4"><inline-formula><mml:math id="M201" display="inline"><mml:mrow><mml:mn mathvariant="normal">1.60</mml:mn><mml:mo>/</mml:mo><mml:mn mathvariant="normal">1.14</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col5"><inline-formula><mml:math id="M202" display="inline"><mml:mrow><mml:mn mathvariant="normal">0.90</mml:mn><mml:mo>/</mml:mo><mml:mn mathvariant="normal">0.59</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col6"><inline-formula><mml:math id="M203" display="inline"><mml:mrow><mml:mn mathvariant="normal">0.67</mml:mn><mml:mo>/</mml:mo><mml:mn mathvariant="normal">0.35</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col7"><inline-formula><mml:math id="M204" display="inline"><mml:mrow><mml:mn mathvariant="normal">0.39</mml:mn><mml:mo>/</mml:mo><mml:mn mathvariant="normal">0.19</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col8"><inline-formula><mml:math id="M205" display="inline"><mml:mrow><mml:mn mathvariant="normal">0.63</mml:mn><mml:mo>/</mml:mo><mml:mn mathvariant="normal">0.10</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col9"/>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">means, temperature</oasis:entry>
         <oasis:entry colname="col2"/>
         <oasis:entry colname="col3"/>
         <oasis:entry colname="col4"/>
         <oasis:entry colname="col5"/>
         <oasis:entry colname="col6"/>
         <oasis:entry colname="col7"/>
         <oasis:entry colname="col8"/>
         <oasis:entry colname="col9"/>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">MPI-GE, NH annual</oasis:entry>
         <oasis:entry colname="col2">100</oasis:entry>
         <oasis:entry colname="col3"><inline-formula><mml:math id="M206" display="inline"><mml:mrow><mml:mn mathvariant="normal">0.16</mml:mn><mml:mo>/</mml:mo><mml:mn mathvariant="normal">0.02</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col4"><inline-formula><mml:math id="M207" display="inline"><mml:mrow><mml:mn mathvariant="normal">1.58</mml:mn><mml:mo>/</mml:mo><mml:mn mathvariant="normal">0.20</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col5"><inline-formula><mml:math id="M208" display="inline"><mml:mrow><mml:mn mathvariant="normal">0.22</mml:mn><mml:mo>/</mml:mo><mml:mn mathvariant="normal">0.03</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col6"><inline-formula><mml:math id="M209" display="inline"><mml:mrow><mml:mn mathvariant="normal">0.16</mml:mn><mml:mo>/</mml:mo><mml:mn mathvariant="normal">0.02</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col7"><inline-formula><mml:math id="M210" display="inline"><mml:mrow><mml:mn mathvariant="normal">0.54</mml:mn><mml:mo>/</mml:mo><mml:mn mathvariant="normal">0.12</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col8"><inline-formula><mml:math id="M211" display="inline"><mml:mrow><mml:mn mathvariant="normal">0.74</mml:mn><mml:mo>/</mml:mo><mml:mn mathvariant="normal">0.07</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col9"/>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">means, temperature</oasis:entry>
         <oasis:entry colname="col2"/>
         <oasis:entry colname="col3"/>
         <oasis:entry colname="col4"/>
         <oasis:entry colname="col5"/>
         <oasis:entry colname="col6"/>
         <oasis:entry colname="col7"/>
         <oasis:entry colname="col8"/>
         <oasis:entry colname="col9"/>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup><?xmltex \end{scaleboxenv}?></oasis:table></table-wrap>

<sec id="Ch1.S4.SS1">
  <label>4.1</label><title>Concentration of measures in atmospheric fields</title>
      <p id="d1e4665">In this subsection we directly investigate the distributions of the
lengths of the sample members and the distributions of the angles between
them. The results from this and the following subsections are summarized
in Table <xref ref-type="table" rid="Ch1.T2"/>.</p>
      <p id="d1e4670">We centre the sample, <inline-formula><mml:math id="M212" display="inline"><mml:mrow><mml:msup><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>k</mml:mi></mml:msup></mml:mrow></mml:math></inline-formula>, <inline-formula><mml:math id="M213" display="inline"><mml:mrow><mml:mi>k</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>,</mml:mo><mml:mi mathvariant="normal">…</mml:mi><mml:mi>K</mml:mi></mml:mrow></mml:math></inline-formula>, to the sample mean, <inline-formula><mml:math id="M214" display="inline"><mml:mrow><mml:mover accent="true"><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:mo>=</mml:mo><mml:msub><mml:mo>∑</mml:mo><mml:mi>k</mml:mi></mml:msub><mml:msup><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>k</mml:mi></mml:msup><mml:mo>/</mml:mo><mml:mi>K</mml:mi></mml:mrow></mml:math></inline-formula>, and calculate the lengths as the
square root of <inline-formula><mml:math id="M215" display="inline"><mml:mrow><mml:mo>|</mml:mo><mml:mo>|</mml:mo><mml:msup><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>k</mml:mi></mml:msup><mml:mo>-</mml:mo><mml:mover accent="true"><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:mo>|</mml:mo><mml:msup><mml:mo>|</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo>/</mml:mo><mml:mi>N</mml:mi></mml:mrow></mml:math></inline-formula> for each sample
member. The angle <inline-formula><mml:math id="M216" display="inline"><mml:mi mathvariant="italic">ϕ</mml:mi></mml:math></inline-formula> between two sample members, <inline-formula><mml:math id="M217" display="inline"><mml:mi>k</mml:mi></mml:math></inline-formula> and <inline-formula><mml:math id="M218" display="inline"><mml:mi>l</mml:mi></mml:math></inline-formula>, is given
by <inline-formula><mml:math id="M219" display="inline"><mml:mrow><mml:mo>(</mml:mo><mml:msup><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>k</mml:mi></mml:msup><mml:mo>-</mml:mo><mml:mover accent="true"><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:mo>)</mml:mo><mml:mo>⋅</mml:mo><mml:mo>(</mml:mo><mml:msup><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>l</mml:mi></mml:msup><mml:mo>-</mml:mo><mml:mover accent="true"><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:mo>|</mml:mo><mml:mo>|</mml:mo><mml:msup><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>k</mml:mi></mml:msup><mml:mo>-</mml:mo><mml:mover accent="true"><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:mo>|</mml:mo><mml:mo>|</mml:mo><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mo>|</mml:mo><mml:mo>|</mml:mo><mml:msup><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>l</mml:mi></mml:msup><mml:mo>-</mml:mo><mml:mover accent="true"><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:mo>|</mml:mo><mml:mo>|</mml:mo><mml:mi>cos⁡</mml:mi><mml:mi mathvariant="italic">ϕ</mml:mi></mml:mrow></mml:math></inline-formula>. This gives us <inline-formula><mml:math id="M220" display="inline"><mml:mi>K</mml:mi></mml:math></inline-formula> lengths and <inline-formula><mml:math id="M221" display="inline"><mml:mrow><mml:mi>K</mml:mi><mml:mo>(</mml:mo><mml:mi>K</mml:mi><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>)</mml:mo><mml:mo>/</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:mrow></mml:math></inline-formula> angles. This centering – the subtraction of the sample mean – is not important for the calculation
of the lengths, as we explain at the end of this subsection.</p>
      <p id="d1e4902">We first consider the near-surface temperature and precipitation fields from the AgERA data set.  Figure <xref ref-type="fig" rid="Ch1.F2"/> shows the lengths and angles for daily means taken every fifth day for June
in the period 1980–1990. The 11 <inline-formula><mml:math id="M222" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">years</mml:mi></mml:mrow></mml:math></inline-formula> give us 66 samples.  We see that for temperature the lengths are relatively tightly distributed around
4.36 <inline-formula><mml:math id="M223" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">K</mml:mi></mml:mrow></mml:math></inline-formula> (<inline-formula><mml:math id="M224" display="inline"><mml:mi mathvariant="italic">σ</mml:mi></mml:math></inline-formula> in Eq. <xref ref-type="disp-formula" rid="Ch1.E6"/>) with a standard deviation of 0.58 <inline-formula><mml:math id="M225" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">K</mml:mi></mml:mrow></mml:math></inline-formula>. The angles are likewise distributed around <inline-formula><mml:math id="M226" display="inline"><mml:mrow><mml:mi mathvariant="italic">π</mml:mi><mml:mo>/</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:mrow></mml:math></inline-formula> with a standard
deviation of 0.21.  For precipitation the distributions are somewhat
narrower, in particular for the angles. This is what we would expect due to the larger number of effective degrees of freedom compared to
temperature. However, this effect is reduced as we include both dry and
wet days in the analysis. While the precipitation amount on wet days has
a short decorrelation length, this does not hold for the spatial field indicating wet/dry days. Note also that the distribution of precipitation
is extremely non-Gaussian.  These results indicate that the concentration
of measures and the waist concentration hold at least to some extent
for these fields.</p>
      <p id="d1e4953">Figure <xref ref-type="fig" rid="Ch1.F3"/> shows the lengths and angles for
the monthly seasonal cycle in near-surface temperature, 1980–2015, for
the multi-model CMIP5 ensemble. The models have been regridded to a common
<inline-formula><mml:math id="M227" display="inline"><mml:mrow><mml:mn mathvariant="normal">144</mml:mn><mml:mo>×</mml:mo><mml:mn mathvariant="normal">73</mml:mn></mml:mrow></mml:math></inline-formula> grid, so <inline-formula><mml:math id="M228" display="inline"><mml:mrow><mml:mi>N</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">144</mml:mn><mml:mo>×</mml:mo><mml:mn mathvariant="normal">73</mml:mn><mml:mo>×</mml:mo><mml:mn mathvariant="normal">12</mml:mn></mml:mrow></mml:math></inline-formula>. The sample has a size of 45 and consists of one ensemble member from each of the models.  The lengths are distributed
around 2.57 <inline-formula><mml:math id="M229" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">K</mml:mi></mml:mrow></mml:math></inline-formula> with a standard deviation of 0.46 <inline-formula><mml:math id="M230" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">K</mml:mi></mml:mrow></mml:math></inline-formula> and the angles around
<inline-formula><mml:math id="M231" display="inline"><mml:mrow><mml:mi mathvariant="italic">π</mml:mi><mml:mo>/</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:mrow></mml:math></inline-formula> with a standard deviation of 0.28.  Thus, compared to the example
in Fig. <xref ref-type="fig" rid="Ch1.F2"/>, the distributions are less tightly distributed. The main explanation is probably that the effective degrees
of freedom in the monthly climatology is smaller than that of the daily
fields. However, there are also reasons to believe that the multi-model
ensemble is not totally independent <xref ref-type="bibr" rid="bib1.bibx34 bib1.bibx8" id="paren.40"/>.  Note the negative skewness in the distribution of the angles. Angles close
to zero indicate pairs of models that are almost parallel and therefore
strongly dependent. These pairs correspond to variants of the same model,
such as MIROC-ESM and MIROC-ESM-CHEM, which are well known to be close
in the model genealogy <xref ref-type="bibr" rid="bib1.bibx34" id="paren.41"/>. A simple comparison between
the distributions of <inline-formula><mml:math id="M232" display="inline"><mml:mi mathvariant="italic">ϕ</mml:mi></mml:math></inline-formula> in Figs. <xref ref-type="fig" rid="Ch1.F2"/>
and <xref ref-type="fig" rid="Ch1.F3"/> with the distributions in
Fig. <xref ref-type="fig" rid="Ch1.F1"/> (from Gaussians) shows that the effective
dimension is between 25 and 50 for temperature and several hundreds for precipitation.</p>
      <p id="d1e5042">Results for the MPI-GE 100-member initial condition ensemble are shown in Table <xref ref-type="table" rid="Ch1.T2"/>. Here we have <inline-formula><mml:math id="M233" display="inline"><mml:mrow><mml:mn mathvariant="normal">192</mml:mn><mml:mo>×</mml:mo><mml:mn mathvariant="normal">96</mml:mn></mml:mrow></mml:math></inline-formula> grid points, so <inline-formula><mml:math id="M234" display="inline"><mml:mrow><mml:mi>N</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">192</mml:mn><mml:mo>×</mml:mo><mml:mn mathvariant="normal">96</mml:mn><mml:mo>×</mml:mo><mml:mn mathvariant="normal">12</mml:mn></mml:mrow></mml:math></inline-formula> and the sample size is 100. The distributions of lengths and
angles are now narrower compared to the multi-model CMIP5 ensemble.
This corresponds to a larger effective dimension in the monthly
climatology which now reflects only different initial conditions and
not model differences. Also in this example are the distributions for
precipitation narrower than those for temperature.</p>
      <p id="d1e5079">Reducing the spatial area decreases the effective dimension.  As an
example we have included in Table <xref ref-type="table" rid="Ch1.T2"/> the results for the
AgERA when applied to northern Europe (50–65<inline-formula><mml:math id="M235" display="inline"><mml:msup><mml:mi/><mml:mo>∘</mml:mo></mml:msup></mml:math></inline-formula> N, 0–25<inline-formula><mml:math id="M236" display="inline"><mml:msup><mml:mi/><mml:mo>∘</mml:mo></mml:msup></mml:math></inline-formula> E). As expected, we see an increase in the width of the distributions for
both precipitation and near-surface temperature.</p>
      <?pagebreak page415?><p id="d1e5102">In the analysis above we centred the sample to the sample mean before calculating the lengths; i.e. we used <inline-formula><mml:math id="M237" display="inline"><mml:mrow><mml:mo>|</mml:mo><mml:mo>|</mml:mo><mml:msup><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>k</mml:mi></mml:msup><mml:mo>-</mml:mo><mml:mover accent="true"><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:mo>|</mml:mo><mml:msup><mml:mo>|</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo>/</mml:mo><mml:mi>N</mml:mi></mml:mrow></mml:math></inline-formula> instead of  <inline-formula><mml:math id="M238" display="inline"><mml:mrow><mml:mo>|</mml:mo><mml:mo>|</mml:mo><mml:msup><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>k</mml:mi></mml:msup><mml:mo>|</mml:mo><mml:msup><mml:mo>|</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo>/</mml:mo><mml:mi>N</mml:mi></mml:mrow></mml:math></inline-formula>. However, these expressions only differ by the length of the mean, <inline-formula><mml:math id="M239" display="inline"><mml:mrow><mml:mo>|</mml:mo><mml:mo>|</mml:mo><mml:msup><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>k</mml:mi></mml:msup><mml:mo>|</mml:mo><mml:msup><mml:mo>|</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo>/</mml:mo><mml:mi>N</mml:mi><mml:mo>=</mml:mo><mml:msup><mml:mover accent="true"><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo>/</mml:mo><mml:mi>N</mml:mi><mml:mo>+</mml:mo><mml:mo>|</mml:mo><mml:mo>|</mml:mo><mml:msup><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>k</mml:mi></mml:msup><mml:mo>-</mml:mo><mml:mover accent="true"><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:mo>|</mml:mo><mml:msup><mml:mo>|</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo>/</mml:mo><mml:mi>N</mml:mi></mml:mrow></mml:math></inline-formula>, as <inline-formula><mml:math id="M240" display="inline"><mml:mrow><mml:msup><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>k</mml:mi></mml:msup><mml:mo>-</mml:mo><mml:mover accent="true"><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo mathvariant="normal">‾</mml:mo></mml:mover></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M241" display="inline"><mml:mover accent="true"><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo mathvariant="normal">‾</mml:mo></mml:mover></mml:math></inline-formula> are orthogonal in high
dimensions (due to waist concentration).
The absence of centering makes most sense for precipitation
that has a natural zero point.  For AgERA precipitation <inline-formula><mml:math id="M242" display="inline"><mml:mrow><mml:mo>|</mml:mo><mml:mo>|</mml:mo><mml:mover accent="true"><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:mo>|</mml:mo><mml:mo>|</mml:mo><mml:mo>/</mml:mo><mml:msqrt><mml:mi>N</mml:mi></mml:msqrt></mml:mrow></mml:math></inline-formula> is <inline-formula><mml:math id="M243" display="inline"><mml:mn mathvariant="normal">3.35</mml:mn></mml:math></inline-formula>, the mean of <inline-formula><mml:math id="M244" display="inline"><mml:mrow><mml:mo>|</mml:mo><mml:mo>|</mml:mo><mml:msup><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>k</mml:mi></mml:msup><mml:mo>-</mml:mo><mml:mover accent="true"><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:mo>|</mml:mo><mml:mo>|</mml:mo><mml:mo>/</mml:mo><mml:msqrt><mml:mi>N</mml:mi></mml:msqrt></mml:mrow></mml:math></inline-formula> is <inline-formula><mml:math id="M245" display="inline"><mml:mn mathvariant="normal">4.34</mml:mn></mml:math></inline-formula>, and the mean of <inline-formula><mml:math id="M246" display="inline"><mml:mrow><mml:mo>|</mml:mo><mml:mo>|</mml:mo><mml:msup><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>k</mml:mi></mml:msup><mml:mo>|</mml:mo><mml:mo>|</mml:mo><mml:msqrt><mml:mi>N</mml:mi></mml:msqrt></mml:mrow></mml:math></inline-formula> is
<inline-formula><mml:math id="M247" display="inline"><mml:mn mathvariant="normal">5.48</mml:mn></mml:math></inline-formula>  (all <inline-formula><mml:math id="M248" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">mm</mml:mi><mml:mspace width="0.125em" linebreak="nobreak"/><mml:msup><mml:mi mathvariant="normal">d</mml:mi><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula>), fulfilling the Pythagorean relationship (<inline-formula><mml:math id="M249" display="inline"><mml:mrow><mml:msup><mml:mn mathvariant="normal">5.48</mml:mn><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo>=</mml:mo><mml:msup><mml:mn mathvariant="normal">3.35</mml:mn><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo>+</mml:mo><mml:msup><mml:mn mathvariant="normal">4.34</mml:mn><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula>).</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F5" specific-use="star"><?xmltex \currentcnt{5}?><?xmltex \def\figurename{Figure}?><label>Figure 5</label><caption><p id="d1e5406">Distances between samples (cyan) and between sample and sample  mean (blue).  <bold>(a)</bold>
AgERA, daily mean precipitation for June. <bold>(b)</bold> CMIP5, monthly climatology of near-surface
temperature.  Note the factor of <inline-formula><mml:math id="M250" display="inline"><mml:msqrt><mml:mn mathvariant="normal">2</mml:mn></mml:msqrt></mml:math></inline-formula> between mean distances.
</p></caption>
          <?xmltex \igopts{width=426.791339pt}?><graphic xlink:href="https://npg.copernicus.org/articles/28/409/2021/npg-28-409-2021-f05.png"/>

        </fig>

      <?xmltex \floatpos{t}?><fig id="Ch1.F6" specific-use="star"><?xmltex \currentcnt{6}?><?xmltex \def\figurename{Figure}?><label>Figure 6</label><caption><p id="d1e5431">Correlations as a function of dimension <inline-formula><mml:math id="M251" display="inline"><mml:mi>N</mml:mi></mml:math></inline-formula>. For each <inline-formula><mml:math id="M252" display="inline"><mml:mi>N</mml:mi></mml:math></inline-formula> we draw 50 <inline-formula><mml:math id="M253" display="inline"><mml:mi>N</mml:mi></mml:math></inline-formula>-dimensional random vectors and calculate correlations of pairs of
sample differences (blue, <inline-formula><mml:math id="M254" display="inline"><mml:mrow><mml:mtext>corr</mml:mtext><mml:mo>(</mml:mo><mml:msup><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>k</mml:mi></mml:msup><mml:mo>-</mml:mo><mml:msup><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>m</mml:mi></mml:msup><mml:mo>,</mml:mo><mml:msup><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>l</mml:mi></mml:msup><mml:mo>-</mml:mo><mml:msup><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>m</mml:mi></mml:msup><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>) and correlations of sample differences and differences between
ensemble mean and individual ensemble members (black, <inline-formula><mml:math id="M255" display="inline"><mml:mrow><mml:mtext>corr</mml:mtext><mml:mo>(</mml:mo><mml:msup><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>k</mml:mi></mml:msup><mml:mo>-</mml:mo><mml:msup><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>m</mml:mi></mml:msup><mml:mo>,</mml:mo><mml:mover accent="true"><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:mo>-</mml:mo><mml:msup><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>m</mml:mi></mml:msup><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>).  The thick curves show the mean
of the correlations and the broken curves the mean <inline-formula><mml:math id="M256" display="inline"><mml:mo>±</mml:mo></mml:math></inline-formula>2 standard deviations. In <bold>(a)</bold> each component of the vectors is drawn from a
standard Gaussian; in <bold>(b)</bold> each component is drawn from a Gamma distribution (see caption to Fig. <xref ref-type="fig" rid="Ch1.F4"/>).  The horizontal black lines indicate <inline-formula><mml:math id="M257" display="inline"><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>/</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M258" display="inline"><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>/</mml:mo><mml:msqrt><mml:mn mathvariant="normal">2</mml:mn></mml:msqrt></mml:mrow></mml:math></inline-formula>.
</p></caption>
          <?xmltex \igopts{width=426.791339pt}?><graphic xlink:href="https://npg.copernicus.org/articles/28/409/2021/npg-28-409-2021-f06.png"/>

        </fig>

</sec>
<sec id="Ch1.S4.SS2">
  <label>4.2</label><title>Distances between samples and between samples and ensemble mean</title>
      <p id="d1e5586">If the sample members are drawn independently from the same distribution
in high dimensions, they have approximately the same length, and we can write

                <disp-formula id="Ch1.E6" content-type="numbered"><label>6</label><mml:math id="M259" display="block"><mml:mstyle class="stylechange" displaystyle="true"/><mml:mrow><mml:mstyle displaystyle="true" class="stylechange"/><mml:mo>|</mml:mo><mml:mo>|</mml:mo><mml:msup><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>k</mml:mi></mml:msup><mml:mo>-</mml:mo><mml:mover accent="true"><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:mo>|</mml:mo><mml:msup><mml:mo>|</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo>/</mml:mo><mml:mi>N</mml:mi><mml:mo>=</mml:mo><mml:msup><mml:mi mathvariant="italic">σ</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo>.</mml:mo></mml:mrow></mml:math></disp-formula>

          For the distance between two different sample members, we get

                <disp-formula id="Ch1.E7" content-type="numbered"><label>7</label><mml:math id="M260" display="block"><mml:mstyle displaystyle="true" class="stylechange"/><mml:mrow><mml:mstyle displaystyle="true" class="stylechange"/><mml:mo>|</mml:mo><mml:mo>|</mml:mo><mml:msup><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>k</mml:mi></mml:msup><mml:mo>-</mml:mo><mml:msup><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>l</mml:mi></mml:msup><mml:mo>|</mml:mo><mml:msup><mml:mo>|</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo>/</mml:mo><mml:mi>N</mml:mi><mml:mo>=</mml:mo><mml:mo>|</mml:mo><mml:mo>|</mml:mo><mml:mo>(</mml:mo><mml:msup><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>k</mml:mi></mml:msup><mml:mo>-</mml:mo><mml:mover accent="true"><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:mo>)</mml:mo><mml:mo>-</mml:mo><mml:mo>(</mml:mo><mml:msup><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>l</mml:mi></mml:msup><mml:mo>-</mml:mo><mml:mover accent="true"><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:mo>)</mml:mo><mml:mo>|</mml:mo><mml:msup><mml:mo>|</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo>/</mml:mo><mml:mi>N</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">2</mml:mn><mml:msup><mml:mi mathvariant="italic">σ</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>

          where we have used the fact that <inline-formula><mml:math id="M261" display="inline"><mml:mrow><mml:msup><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>k</mml:mi></mml:msup><mml:mo>-</mml:mo><mml:mover accent="true"><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo mathvariant="normal">‾</mml:mo></mml:mover></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M262" display="inline"><mml:mrow><mml:msup><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>l</mml:mi></mml:msup><mml:mo>-</mml:mo><mml:mover accent="true"><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo mathvariant="normal">‾</mml:mo></mml:mover></mml:mrow></mml:math></inline-formula> are orthogonal.</p>
      <p id="d1e5769">Therefore, the distance between two sample members is a square root of 2 larger than the distance between a sample member and the sample mean. The geometric interpretation is that the sample mean and any two
sample members form an isosceles right triangle with the right angle at
the sample mean <xref ref-type="bibr" rid="bib1.bibx27 bib1.bibx41" id="paren.42"/>. The factor of <inline-formula><mml:math id="M263" display="inline"><mml:msqrt><mml:mn mathvariant="normal">2</mml:mn></mml:msqrt></mml:math></inline-formula>
then comes from Pythagoras' equation. It is worth noting that the sample mean is special and is not drawn from the same distribution as
the sample members. As mentioned when discussing the example of the
high-dimensional unit cube from Sect. <xref ref-type="sec" rid="Ch1.S2"/>a, the
sample members would be located in the spikes, while the sample mean would be close to the centre.</p>
      <p id="d1e5785">Figure <xref ref-type="fig" rid="Ch1.F4"/> demonstrates this in the simple situation
where 50 <inline-formula><mml:math id="M264" display="inline"><mml:mi>N</mml:mi></mml:math></inline-formula>-dimensional vectors are drawn from prescribed distributions.
The results are shown as a function of <inline-formula><mml:math id="M265" display="inline"><mml:mi>N</mml:mi></mml:math></inline-formula>. When <inline-formula><mml:math id="M266" display="inline"><mml:mi>N</mml:mi></mml:math></inline-formula> increases, the spread of the distances decreases, and for large <inline-formula><mml:math id="M267" display="inline"><mml:mi>N</mml:mi></mml:math></inline-formula> the factor of <inline-formula><mml:math id="M268" display="inline"><mml:msqrt><mml:mn mathvariant="normal">2</mml:mn></mml:msqrt></mml:math></inline-formula> is
clearly seen. This holds both for simple spherical Gaussian-distributed vectors (left panel) and for Gamma-distributed vectors with strong
anisotropy in the components (right panel), although the convergence is faster in the Gaussian case.</p>
      <p id="d1e5827">Figure <xref ref-type="fig" rid="Ch1.F5"/> shows the distances for AgERA daily mean
precipitation for June (left panel) and for the monthly climatology of
near-surface temperature for the CMIP5 ensemble<?pagebreak page416?> (right panel). The
distribution of the distances between sample members is shown together
with the distribution of the distances between the sample members and the
sample mean. The mean and width of these distributions are also shown
in Table <xref ref-type="table" rid="Ch1.T2"/> for both these and the other data sets. In all cases the factor of <inline-formula><mml:math id="M269" display="inline"><mml:msqrt><mml:mn mathvariant="normal">2</mml:mn></mml:msqrt></mml:math></inline-formula> is clearly seen for the mean values, although the widths of the distributions are substantial in all cases.
For the AgERA daily precipitation (Fig. <xref ref-type="fig" rid="Ch1.F5"/> left), the two distributions are almost separated, while this is not the case for the CMIP5 ensemble.</p>
      <p id="d1e5845">The indistinguishable interpretation claims that observations are
drawn from the same distribution as the ensemble members.  With this
assumption and the considerations above, <xref ref-type="bibr" rid="bib1.bibx13" id="text.43"/>
explained the ubiquitous observation that the error (compared
to observations) of the ensemble mean often is 30 % smaller
(<inline-formula><mml:math id="M270" display="inline"><mml:mrow><mml:mo>(</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>-</mml:mo><mml:msqrt><mml:mn mathvariant="normal">2</mml:mn></mml:msqrt><mml:mo>)</mml:mo><mml:mo>/</mml:mo><mml:msqrt><mml:mn mathvariant="normal">2</mml:mn></mml:msqrt><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> than the typical error of the individual
ensemble members <xref ref-type="bibr" rid="bib1.bibx25" id="paren.44"><named-content content-type="pre">e.g.</named-content></xref>. We also explained why the ensemble mean very often has a smaller error than all individual
ensemble members <xref ref-type="bibr" rid="bib1.bibx14" id="paren.45"/>.</p>
      <p id="d1e5883">The results in this subsection and Sects. <xref ref-type="sec" rid="Ch1.S4"/>.2
and 4 not only hold for the Euclidean (square) norm distance, but also for e.g. the maximum norm distance and the correlation distance (<inline-formula><mml:math id="M271" display="inline"><mml:msqrt><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>-</mml:mo><mml:msup><mml:mi>r</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:msqrt></mml:math></inline-formula>, where <inline-formula><mml:math id="M272" display="inline"><mml:mi>r</mml:mi></mml:math></inline-formula> is correlation).</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F7" specific-use="star"><?xmltex \currentcnt{7}?><?xmltex \def\figurename{Figure}?><label>Figure 7</label><caption><p id="d1e5913">Correlations of pairs of sample differences <inline-formula><mml:math id="M273" display="inline"><mml:mrow><mml:mtext>corr</mml:mtext><mml:mo>(</mml:mo><mml:msup><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>k</mml:mi></mml:msup><mml:mo>-</mml:mo><mml:msup><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>m</mml:mi></mml:msup><mml:mo>,</mml:mo><mml:msup><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>l</mml:mi></mml:msup><mml:mo>-</mml:mo><mml:msup><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>m</mml:mi></mml:msup><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> (cyan) and correlations of sample differences
and differences between sample mean and individual sample members
<inline-formula><mml:math id="M274" display="inline"><mml:mrow><mml:mtext>corr</mml:mtext><mml:mo>(</mml:mo><mml:msup><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>k</mml:mi></mml:msup><mml:mo>-</mml:mo><mml:msup><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>m</mml:mi></mml:msup><mml:mo>,</mml:mo><mml:mover accent="true"><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:mo>-</mml:mo><mml:msup><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>m</mml:mi></mml:msup><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> (blue).  <bold>(a)</bold>
Daily mean precipitation June from AgERA. <bold>(b)</bold> Monthly climatology of
near-surface temperature in CMIP5.
</p></caption>
          <?xmltex \igopts{width=426.791339pt}?><graphic xlink:href="https://npg.copernicus.org/articles/28/409/2021/npg-28-409-2021-f07.png"/>

        </fig>

      <?xmltex \floatpos{t}?><fig id="Ch1.F8" specific-use="star"><?xmltex \currentcnt{8}?><?xmltex \def\figurename{Figure}?><label>Figure 8</label><caption><p id="d1e6006">The length of the sample mean <inline-formula><mml:math id="M275" display="inline"><mml:mrow><mml:mo>|</mml:mo><mml:mo>|</mml:mo><mml:mover accent="true"><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:mo>|</mml:mo><mml:msup><mml:mo>|</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo>/</mml:mo><mml:mi>N</mml:mi></mml:mrow></mml:math></inline-formula> as a function of sample size <inline-formula><mml:math id="M276" display="inline"><mml:mi>K</mml:mi></mml:math></inline-formula>. Black curves show results from <inline-formula><mml:math id="M277" display="inline"><mml:mrow><mml:mi>N</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">100</mml:mn></mml:mrow></mml:math></inline-formula>
and blue curves for <inline-formula><mml:math id="M278" display="inline"><mml:mrow><mml:mi>N</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">10</mml:mn></mml:mrow></mml:math></inline-formula>. For each <inline-formula><mml:math id="M279" display="inline"><mml:mi>K</mml:mi></mml:math></inline-formula> results are based on 200 draws.
The solid curves show the mean over these draws and the broken curves
the mean <inline-formula><mml:math id="M280" display="inline"><mml:mo>±</mml:mo></mml:math></inline-formula>2 standard deviations. The red curve is an analytic result (Eq. <xref ref-type="disp-formula" rid="Ch1.E12"/>) with the theoretical values for <inline-formula><mml:math id="M281" display="inline"><mml:mi mathvariant="italic">μ</mml:mi></mml:math></inline-formula> and
<inline-formula><mml:math id="M282" display="inline"><mml:mi mathvariant="italic">σ</mml:mi></mml:math></inline-formula>. In <bold>(a)</bold> each component of the vectors is drawn from a
standard Gaussian <inline-formula><mml:math id="M283" display="inline"><mml:mrow><mml:msup><mml:mi mathvariant="italic">μ</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0</mml:mn></mml:mrow></mml:math></inline-formula>, <inline-formula><mml:math id="M284" display="inline"><mml:mrow><mml:msup><mml:mi mathvariant="italic">σ</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:math></inline-formula>; in <bold>(b)</bold> each component is drawn from a Gamma distribution, <inline-formula><mml:math id="M285" display="inline"><mml:mrow><mml:msup><mml:mi mathvariant="italic">μ</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo>=</mml:mo><mml:mn mathvariant="normal">75.00</mml:mn></mml:mrow></mml:math></inline-formula>, <inline-formula><mml:math id="M286" display="inline"><mml:mrow><mml:msup><mml:mi mathvariant="italic">σ</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo>=</mml:mo><mml:mn mathvariant="normal">25.00</mml:mn></mml:mrow></mml:math></inline-formula>
(see caption to Fig. <xref ref-type="fig" rid="Ch1.F4"/>).
</p></caption>
          <?xmltex \igopts{width=497.923228pt}?><graphic xlink:href="https://npg.copernicus.org/articles/28/409/2021/npg-28-409-2021-f08.png"/>

        </fig>

</sec>
<?pagebreak page417?><sec id="Ch1.S4.SS3">
  <label>4.3</label><title>Correlations between sample differences</title>
      <p id="d1e6180">Error correlations and correlations between model differences are
important when studying the structure of a model ensemble and when comparing
an ensemble to observations <xref ref-type="bibr" rid="bib1.bibx2 bib1.bibx42 bib1.bibx6" id="paren.46"/>.</p>
      <p id="d1e6186">We have in general
<inline-formula><mml:math id="M287" display="inline"><mml:mrow><mml:mtext>corr</mml:mtext><mml:mo>(</mml:mo><mml:msup><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>k</mml:mi></mml:msup><mml:mo>,</mml:mo><mml:msup><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>l</mml:mi></mml:msup><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>-</mml:mo><mml:mstyle displaystyle="false"><mml:mfrac style="text"><mml:mn mathvariant="normal">1</mml:mn><mml:mn mathvariant="normal">2</mml:mn></mml:mfrac></mml:mstyle><mml:mo>|</mml:mo><mml:mo>|</mml:mo><mml:msup><mml:mover accent="true"><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo mathvariant="normal" stretchy="false">^</mml:mo></mml:mover><mml:mi>k</mml:mi></mml:msup><mml:mo>-</mml:mo><mml:msup><mml:mover accent="true"><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo mathvariant="normal" stretchy="false">^</mml:mo></mml:mover><mml:mi>l</mml:mi></mml:msup><mml:mo>|</mml:mo><mml:msup><mml:mo>|</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo>/</mml:mo><mml:mi>N</mml:mi></mml:mrow></mml:math></inline-formula>,
where <inline-formula><mml:math id="M288" display="inline"><mml:mover accent="true"><mml:mspace width="0.25em" linebreak="nobreak"/><mml:mo stretchy="false" mathvariant="normal">^</mml:mo></mml:mover></mml:math></inline-formula> indicates variables standardized to zero mean and unit variance.
Therefore, with <inline-formula><mml:math id="M289" display="inline"><mml:mrow><mml:msup><mml:mi mathvariant="italic">ϵ</mml:mi><mml:mi>k</mml:mi></mml:msup><mml:mo>=</mml:mo><mml:msup><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>k</mml:mi></mml:msup><mml:mo>-</mml:mo><mml:msup><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>m</mml:mi></mml:msup></mml:mrow></mml:math></inline-formula> we have
<inline-formula><mml:math id="M290" display="inline"><mml:mrow><mml:msup><mml:mover accent="true"><mml:mi mathvariant="bold-italic">ϵ</mml:mi><mml:mo mathvariant="normal" stretchy="false">^</mml:mo></mml:mover><mml:mi>k</mml:mi></mml:msup><mml:mo>=</mml:mo><mml:mo>(</mml:mo><mml:msup><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>k</mml:mi></mml:msup><mml:mo>-</mml:mo><mml:msup><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>m</mml:mi></mml:msup><mml:mo>)</mml:mo><mml:mo>/</mml:mo><mml:msqrt><mml:mn mathvariant="normal">2</mml:mn></mml:msqrt><mml:mo>/</mml:mo><mml:mi mathvariant="italic">σ</mml:mi></mml:mrow></mml:math></inline-formula>. We now get

                <disp-formula specific-use="align" content-type="numbered"><mml:math id="M291" display="block"><mml:mtable displaystyle="true"><mml:mtr><mml:mtd><mml:mrow><mml:mstyle displaystyle="true" class="stylechange"/><mml:mtext>corr</mml:mtext><mml:mo>(</mml:mo><mml:msup><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>k</mml:mi></mml:msup><mml:mo>-</mml:mo><mml:msup><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>m</mml:mi></mml:msup><mml:mo>,</mml:mo><mml:msup><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>l</mml:mi></mml:msup><mml:mo>-</mml:mo><mml:msup><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>m</mml:mi></mml:msup><mml:mo>)</mml:mo></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><mml:mo>=</mml:mo><mml:mtext>corr</mml:mtext><mml:mo>(</mml:mo><mml:msup><mml:mi mathvariant="bold-italic">ϵ</mml:mi><mml:mi>k</mml:mi></mml:msup><mml:mo>,</mml:mo><mml:msup><mml:mi mathvariant="bold-italic">ϵ</mml:mi><mml:mi>l</mml:mi></mml:msup><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mstyle displaystyle="true" class="stylechange"/></mml:mtd><mml:mtd><mml:mrow><mml:mstyle displaystyle="true" class="stylechange"/><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>-</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mn mathvariant="normal">1</mml:mn><mml:mn mathvariant="normal">2</mml:mn></mml:mfrac></mml:mstyle><mml:msup><mml:mfenced open="∥" close="∥"><mml:mrow><mml:msup><mml:mover accent="true"><mml:mi mathvariant="bold-italic">ϵ</mml:mi><mml:mo mathvariant="normal" stretchy="false">^</mml:mo></mml:mover><mml:mi>k</mml:mi></mml:msup><mml:mo>-</mml:mo><mml:msup><mml:mover accent="true"><mml:mi mathvariant="bold-italic">ϵ</mml:mi><mml:mo mathvariant="normal" stretchy="false">^</mml:mo></mml:mover><mml:mi>l</mml:mi></mml:msup></mml:mrow></mml:mfenced><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo>/</mml:mo><mml:mi>N</mml:mi></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mstyle class="stylechange" displaystyle="true"/></mml:mtd><mml:mtd><mml:mrow><mml:mstyle displaystyle="true" class="stylechange"/><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>-</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mn mathvariant="normal">1</mml:mn><mml:mn mathvariant="normal">2</mml:mn></mml:mfrac></mml:mstyle><mml:msup><mml:mfenced close="∥" open="∥"><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:msup><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>k</mml:mi></mml:msup><mml:mo>-</mml:mo><mml:msup><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>l</mml:mi></mml:msup></mml:mrow><mml:mrow><mml:msqrt><mml:mn mathvariant="normal">2</mml:mn></mml:msqrt><mml:mi mathvariant="italic">σ</mml:mi></mml:mrow></mml:mfrac></mml:mstyle></mml:mfenced><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo>/</mml:mo><mml:mi>N</mml:mi></mml:mrow></mml:mtd></mml:mtr><mml:mlabeledtr id="Ch1.E8"><mml:mtd><mml:mtext>8</mml:mtext></mml:mtd><mml:mtd><mml:mstyle displaystyle="true" class="stylechange"/></mml:mtd><mml:mtd><mml:mrow><mml:mstyle displaystyle="true" class="stylechange"/><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>-</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mn mathvariant="normal">1</mml:mn><mml:mn mathvariant="normal">2</mml:mn></mml:mfrac></mml:mstyle><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mn mathvariant="normal">1</mml:mn><mml:mn mathvariant="normal">2</mml:mn></mml:mfrac></mml:mstyle><mml:mo>,</mml:mo></mml:mrow></mml:mtd></mml:mlabeledtr></mml:mtable></mml:math></disp-formula>

            where in the last step we have used Eq. (<xref ref-type="disp-formula" rid="Ch1.E7"/>).  Thus, in high
dimensions the correlation between sample differences is <inline-formula><mml:math id="M292" display="inline"><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>/</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:mrow></mml:math></inline-formula>.</p>
      <p id="d1e6537"><?xmltex \hack{\newpage}?>Replacing <inline-formula><mml:math id="M293" display="inline"><mml:mrow><mml:msup><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>l</mml:mi></mml:msup></mml:mrow></mml:math></inline-formula> with the sample mean, we get

                <disp-formula specific-use="align" content-type="numbered"><mml:math id="M294" display="block"><mml:mtable displaystyle="true"><mml:mlabeledtr id="Ch1.E9"><mml:mtd><mml:mtext>9</mml:mtext></mml:mtd><mml:mtd><mml:mstyle displaystyle="true" class="stylechange"/></mml:mtd><mml:mtd><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><?xmltex \hack{\hbox\bgroup\fontsize{9.3}{9.3}\selectfont$\displaystyle}?><mml:mtext mathvariant="normal">corr</mml:mtext><mml:mo>(</mml:mo><mml:msup><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>k</mml:mi></mml:msup><mml:mo>-</mml:mo><mml:msup><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>m</mml:mi></mml:msup><mml:mo>,</mml:mo><mml:mover accent="true"><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:mo>-</mml:mo><mml:msup><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>m</mml:mi></mml:msup><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>-</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mn mathvariant="normal">1</mml:mn><mml:mn mathvariant="normal">2</mml:mn></mml:mfrac></mml:mstyle><mml:msup><mml:mfenced open="∥" close="∥"><mml:mrow><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:msup><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>k</mml:mi></mml:msup><mml:mo>-</mml:mo><mml:msup><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>m</mml:mi></mml:msup></mml:mrow><mml:mrow><mml:msqrt><mml:mn mathvariant="normal">2</mml:mn></mml:msqrt><mml:mi mathvariant="italic">σ</mml:mi></mml:mrow></mml:mfrac></mml:mstyle><mml:mo>-</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:mover accent="true"><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:mo>-</mml:mo><mml:msup><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>m</mml:mi></mml:msup></mml:mrow><mml:mi mathvariant="italic">σ</mml:mi></mml:mfrac></mml:mstyle></mml:mrow></mml:mfenced><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo>/</mml:mo><mml:mi>N</mml:mi><?xmltex \hack{$\egroup}?></mml:mrow></mml:mtd></mml:mlabeledtr><mml:mlabeledtr id="Ch1.E10"><mml:mtd><mml:mtext>10</mml:mtext></mml:mtd><mml:mtd><mml:mstyle displaystyle="true" class="stylechange"/></mml:mtd><mml:mtd><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>-</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mn mathvariant="normal">1</mml:mn><mml:mn mathvariant="normal">2</mml:mn></mml:mfrac></mml:mstyle><mml:msup><mml:mfenced close="∥" open="∥"><mml:mrow><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:msup><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>k</mml:mi></mml:msup><mml:mo>-</mml:mo><mml:mover accent="true"><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo mathvariant="normal">‾</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:msqrt><mml:mn mathvariant="normal">2</mml:mn></mml:msqrt><mml:mi mathvariant="italic">σ</mml:mi></mml:mrow></mml:mfrac></mml:mstyle><mml:mo>+</mml:mo><mml:mo>(</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>/</mml:mo><mml:msqrt><mml:mn mathvariant="normal">2</mml:mn></mml:msqrt><mml:mo>)</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:msup><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>m</mml:mi></mml:msup><mml:mo>-</mml:mo><mml:mover accent="true"><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo mathvariant="normal">‾</mml:mo></mml:mover></mml:mrow><mml:mi mathvariant="italic">σ</mml:mi></mml:mfrac></mml:mstyle></mml:mrow></mml:mfenced><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo>/</mml:mo><mml:mi>N</mml:mi></mml:mrow></mml:mtd></mml:mlabeledtr><mml:mlabeledtr id="Ch1.E11"><mml:mtd><mml:mtext>11</mml:mtext></mml:mtd><mml:mtd><mml:mstyle displaystyle="true" class="stylechange"/></mml:mtd><mml:mtd><mml:mrow><mml:mstyle displaystyle="true" class="stylechange"/><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>-</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mn mathvariant="normal">1</mml:mn><mml:mn mathvariant="normal">2</mml:mn></mml:mfrac></mml:mstyle><mml:mo>(</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>/</mml:mo><mml:mn mathvariant="normal">2</mml:mn><mml:mo>+</mml:mo><mml:mo>(</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>/</mml:mo><mml:msqrt><mml:mn mathvariant="normal">2</mml:mn></mml:msqrt><mml:msup><mml:mo>)</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>/</mml:mo><mml:msqrt><mml:mn mathvariant="normal">2</mml:mn></mml:msqrt><mml:mo>.</mml:mo></mml:mrow></mml:mtd></mml:mlabeledtr></mml:mtable></mml:math></disp-formula>

            In the last step we used the independence of the two terms and applied
Eq. (<xref ref-type="disp-formula" rid="Ch1.E6"/>) to each.</p>
      <p id="d1e6814">Figure <xref ref-type="fig" rid="Ch1.F6"/> shows the correlations for the simple
random vectors as also used in Fig. <xref ref-type="fig" rid="Ch1.F4"/>. For all <inline-formula><mml:math id="M295" display="inline"><mml:mi>N</mml:mi></mml:math></inline-formula>
the correlations are distributed around <inline-formula><mml:math id="M296" display="inline"><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>/</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M297" display="inline"><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>/</mml:mo><mml:msqrt><mml:mn mathvariant="normal">2</mml:mn></mml:msqrt></mml:mrow></mml:math></inline-formula>. For
small <inline-formula><mml:math id="M298" display="inline"><mml:mi>N</mml:mi></mml:math></inline-formula> the spread is large, but it decreases when <inline-formula><mml:math id="M299" display="inline"><mml:mi>N</mml:mi></mml:math></inline-formula> increases, and for large <inline-formula><mml:math id="M300" display="inline"><mml:mi>N</mml:mi></mml:math></inline-formula> the correlations are very narrowly distributed around <inline-formula><mml:math id="M301" display="inline"><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>/</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:mrow></mml:math></inline-formula>
and <inline-formula><mml:math id="M302" display="inline"><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>/</mml:mo><mml:msqrt><mml:mn mathvariant="normal">2</mml:mn></mml:msqrt><mml:mo>≈</mml:mo><mml:mn mathvariant="normal">0.71</mml:mn></mml:mrow></mml:math></inline-formula>.</p>
      <p id="d1e6905">The correlations for AgERA daily mean precipitation for June  and for
the CMIP5 monthly climatology of near-surface temperate are shown in
Fig. <xref ref-type="fig" rid="Ch1.F7"/>. The mean values are close to the high-dimensional
values from Eqs. (<xref ref-type="disp-formula" rid="Ch1.E8"/>) and (<xref ref-type="disp-formula" rid="Ch1.E9"/>), although the spread is rather high.  This is also the case for the other fields as reported
in Table <xref ref-type="table" rid="Ch1.T2"/>.</p>
      <?pagebreak page418?><p id="d1e6916">If we again assume that the observations are drawn from the same
distribution as the ensemble members – the indistinguishable
interpretation – the error correlation is <inline-formula><mml:math id="M303" display="inline"><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>/</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:mrow></mml:math></inline-formula> (Eq. <xref ref-type="disp-formula" rid="Ch1.E8"/>).
On the other hand, if observations are near the ensemble mean
– the truth-centred interpretation – the error correlations will be zero as <inline-formula><mml:math id="M304" display="inline"><mml:mrow><mml:msup><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>k</mml:mi></mml:msup><mml:mo>-</mml:mo><mml:mover accent="true"><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo mathvariant="normal">‾</mml:mo></mml:mover></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M305" display="inline"><mml:mrow><mml:msup><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>l</mml:mi></mml:msup><mml:mo>-</mml:mo><mml:mover accent="true"><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo mathvariant="normal">‾</mml:mo></mml:mover></mml:mrow></mml:math></inline-formula> are orthogonal. Error correlations around <inline-formula><mml:math id="M306" display="inline"><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>/</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:mrow></mml:math></inline-formula> have been observed in many studies of climate
models <xref ref-type="bibr" rid="bib1.bibx42 bib1.bibx30 bib1.bibx1" id="paren.47"><named-content content-type="pre">e.g.</named-content></xref>, providing evidence for the indistinguishable interpretation.</p>
      <p id="d1e6987">With <inline-formula><mml:math id="M307" display="inline"><mml:mrow><mml:msup><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>m</mml:mi></mml:msup></mml:mrow></mml:math></inline-formula> replaced by observations, Eq. (<xref ref-type="disp-formula" rid="Ch1.E9"/>) gives the
correlation between individual model errors and the model mean error. This
quantity is shown in Fig. 2 of <xref ref-type="bibr" rid="bib1.bibx42" id="text.48"/> for the climatology
of different variables in the CMIP3 multi-model ensemble, and it is always close to <inline-formula><mml:math id="M308" display="inline"><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>/</mml:mo><mml:msqrt><mml:mn mathvariant="normal">2</mml:mn></mml:msqrt><mml:mo>≈</mml:mo><mml:mn mathvariant="normal">0.71</mml:mn></mml:mrow></mml:math></inline-formula>, as predicted by Eq. (<xref ref-type="disp-formula" rid="Ch1.E9"/>).</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F9" specific-use="star"><?xmltex \currentcnt{9}?><?xmltex \def\figurename{Figure}?><label>Figure 9</label><caption><p id="d1e7027"><bold>(a)</bold> Time series of annual NH mean temperature from MPI-GE (black) and CMIP5 (cyan). Thick solid curves are ensemble means, a dashed curves ensemble means <inline-formula><mml:math id="M309" display="inline"><mml:mo>±</mml:mo></mml:math></inline-formula>2 standard deviations, and thin curves are individual models.  Each ensemble has been centred to its ensemble mean in the first 10 <inline-formula><mml:math id="M310" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">years</mml:mi></mml:mrow></mml:math></inline-formula>.  <bold>(b)</bold> The length of the ensemble mean <inline-formula><mml:math id="M311" display="inline"><mml:mrow><mml:mo>|</mml:mo><mml:mo>|</mml:mo><mml:mover accent="true"><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:mo>|</mml:mo><mml:msup><mml:mo>|</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo>/</mml:mo><mml:mi>N</mml:mi></mml:mrow></mml:math></inline-formula> as a function of ensemble size <inline-formula><mml:math id="M312" display="inline"><mml:mi>K</mml:mi></mml:math></inline-formula> for MPI-GE (black) and CMIP5 (cyan). The ensemble means <inline-formula><mml:math id="M313" display="inline"><mml:mo>±</mml:mo></mml:math></inline-formula>2 standard deviations are also shown. Theoretical results from Eq. (<xref ref-type="disp-formula" rid="Ch1.E12"/>) with  <inline-formula><mml:math id="M314" display="inline"><mml:mrow><mml:msup><mml:mi mathvariant="italic">σ</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.094</mml:mn></mml:mrow></mml:math></inline-formula>, <inline-formula><mml:math id="M315" display="inline"><mml:mrow><mml:msup><mml:mi mathvariant="italic">μ</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.164</mml:mn><mml:msup><mml:mi>K</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula> for MPI-GE and <inline-formula><mml:math id="M316" display="inline"><mml:mrow><mml:msup><mml:mi mathvariant="italic">σ</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.654</mml:mn></mml:mrow></mml:math></inline-formula>, <inline-formula><mml:math id="M317" display="inline"><mml:mrow><mml:msup><mml:mi mathvariant="italic">μ</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.193</mml:mn><mml:msup><mml:mi>K</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula> for CMIP5 are shown in red.
</p></caption>
          <?xmltex \igopts{width=497.923228pt}?><graphic xlink:href="https://npg.copernicus.org/articles/28/409/2021/npg-28-409-2021-f09.png"/>

        </fig>

</sec>
<sec id="Ch1.S4.SS4">
  <label>4.4</label><title>Effect of sample size</title>
      <p id="d1e7177">We now consider how the sample mean depends on the
sample size. The ensemble mean is often used to estimate
the forced response from initial condition and multi-model
ensembles <xref ref-type="bibr" rid="bib1.bibx23 bib1.bibx4 bib1.bibx37" id="paren.49"/>, and it is of interest to know how large an ensemble is needed for the estimation
to be saturated <xref ref-type="bibr" rid="bib1.bibx39" id="paren.50"/>.</p>
      <p id="d1e7186">Letting <inline-formula><mml:math id="M318" display="inline"><mml:mover accent="true"><mml:mrow><mml:msup><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi mathvariant="normal">∞</mml:mi></mml:msup></mml:mrow><mml:mo mathvariant="normal">‾</mml:mo></mml:mover></mml:math></inline-formula> represent the true (i.e. the distribution) mean of the sample, we get in the high-dimensional case (reformulating Eq. <xref ref-type="disp-formula" rid="Ch1.E6"/>)

                <disp-formula id="Ch1.E12" content-type="numbered"><label>12</label><mml:math id="M319" display="block"><mml:mstyle class="stylechange" displaystyle="true"/><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><mml:mo>|</mml:mo><mml:mo>|</mml:mo><mml:mover accent="true"><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:mo>|</mml:mo><mml:msup><mml:mo>|</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo>/</mml:mo><mml:mi>N</mml:mi><mml:mo>=</mml:mo><mml:msup><mml:mi mathvariant="italic">μ</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo>+</mml:mo><mml:msup><mml:mi mathvariant="italic">σ</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo>/</mml:mo><mml:mi>K</mml:mi><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>

          where <inline-formula><mml:math id="M320" display="inline"><mml:mrow><mml:msup><mml:mi mathvariant="italic">μ</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo>=</mml:mo><mml:mo>|</mml:mo><mml:mo>|</mml:mo><mml:mover accent="true"><mml:mrow><mml:msup><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi mathvariant="normal">∞</mml:mi></mml:msup></mml:mrow><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:mo>|</mml:mo><mml:msup><mml:mo>|</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo>/</mml:mo><mml:mi>N</mml:mi></mml:mrow></mml:math></inline-formula>. Thus, <inline-formula><mml:math id="M321" display="inline"><mml:mrow><mml:mo>|</mml:mo><mml:mo>|</mml:mo><mml:mover accent="true"><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:mo>|</mml:mo><mml:msup><mml:mo>|</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula> converges like <inline-formula><mml:math id="M322" display="inline"><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>/</mml:mo><mml:mi>K</mml:mi></mml:mrow></mml:math></inline-formula>, and the convergence is slowest where the sample spread is largest. Similar results have been
presented by <xref ref-type="bibr" rid="bib1.bibx48" id="text.51"/> and <xref ref-type="bibr" rid="bib1.bibx43" id="text.52"/> based on other
arguments. See also <xref ref-type="bibr" rid="bib1.bibx15" id="text.53"/> for the decay of the error
of the ensemble mean when compared to observations.</p>
      <p id="d1e7335">The practical way to estimate the effect of sample size is to apply a
bootstrap procedure to a large sample of size <inline-formula><mml:math id="M323" display="inline"><mml:mrow><mml:msup><mml:mi>K</mml:mi><mml:mn mathvariant="normal">0</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula>. From this sample we draw (with replacement) a number of sub-samples of size <inline-formula><mml:math id="M324" display="inline"><mml:mi>K</mml:mi></mml:math></inline-formula>, <inline-formula><mml:math id="M325" display="inline"><mml:mrow><mml:mi>K</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>,</mml:mo><mml:mi mathvariant="normal">…</mml:mi><mml:msup><mml:mi>K</mml:mi><mml:mn mathvariant="normal">0</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula>.  From these sub-samples we calculate the mean and
spread of <inline-formula><mml:math id="M326" display="inline"><mml:mrow><mml:mo>|</mml:mo><mml:mo>|</mml:mo><mml:mover accent="true"><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:mo>|</mml:mo><mml:msup><mml:mo>|</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo>/</mml:mo><mml:mi>N</mml:mi></mml:mrow></mml:math></inline-formula> for each <inline-formula><mml:math id="M327" display="inline"><mml:mi>K</mml:mi></mml:math></inline-formula>.</p>
      <p id="d1e7410">The mean is shown as a function of <inline-formula><mml:math id="M328" display="inline"><mml:mi>K</mml:mi></mml:math></inline-formula> – using the bootstrap procedure – in Fig. <xref ref-type="fig" rid="Ch1.F8"/> for the simple examples with <inline-formula><mml:math id="M329" display="inline"><mml:mrow><mml:mi>N</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">100</mml:mn></mml:mrow></mml:math></inline-formula>
and <inline-formula><mml:math id="M330" display="inline"><mml:mrow><mml:mi>N</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">10</mml:mn></mml:mrow></mml:math></inline-formula>. For <inline-formula><mml:math id="M331" display="inline"><mml:mrow><mml:mi>N</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">100</mml:mn></mml:mrow></mml:math></inline-formula> (black curves) <inline-formula><mml:math id="M332" display="inline"><mml:mrow><mml:mo>|</mml:mo><mml:mo>|</mml:mo><mml:mover accent="true"><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:mo>|</mml:mo><mml:msup><mml:mo>|</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo>/</mml:mo><mml:mi>N</mml:mi></mml:mrow></mml:math></inline-formula> is
narrowly distributed around the theoretical mean (Eq. <xref ref-type="disp-formula" rid="Ch1.E12"/>)
for both the Gaussian- (left) and Gamma-distributed samples (right). For <inline-formula><mml:math id="M333" display="inline"><mml:mrow><mml:mi>N</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">10</mml:mn></mml:mrow></mml:math></inline-formula> (cyan curves) <inline-formula><mml:math id="M334" display="inline"><mml:mrow><mml:mo>|</mml:mo><mml:mo>|</mml:mo><mml:mover accent="true"><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:mo>|</mml:mo><mml:msup><mml:mo>|</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo>/</mml:mo><mml:mi>N</mml:mi></mml:mrow></mml:math></inline-formula> is also
distributed around the theoretical mean, but with larger spread.</p>
      <p id="d1e7526">In the three previous subsections we studied the samples of daily
June temperatures and of monthly climatologies. In the former the
<inline-formula><mml:math id="M335" display="inline"><mml:mi>N</mml:mi></mml:math></inline-formula> vectors consisted of spatial maps and in the latter of combined spatial climatologies for all 12 months.  However, we also work in
high dimensionality when considering a single long time series.  The left panel in Fig. <xref ref-type="fig" rid="Ch1.F9"/> shows time series of the annual NH mean near-surface temperature for the MPI-GE 100 member initial condition
ensemble and the CMIP5 45-member multi-model ensemble for the period 1961–2005. Both ensembles have been centred to their ensemble means in the first 10 <inline-formula><mml:math id="M336" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">years</mml:mi></mml:mrow></mml:math></inline-formula>. Both ensemble means agree on a forced response
consisting of an overall trend with some signals of volcanic eruptions
after 1982 (El Chrichón) and 1991 (Mount Pinatubo). The spread of
the multi-model ensemble is much larger than the spread of the initial
condition ensemble.</p>
      <?pagebreak page419?><p id="d1e7546">The right panel shows <inline-formula><mml:math id="M337" display="inline"><mml:mrow><mml:mo>|</mml:mo><mml:mo>|</mml:mo><mml:mover accent="true"><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:mo>|</mml:mo><mml:msup><mml:mo>|</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo>/</mml:mo><mml:mi>N</mml:mi></mml:mrow></mml:math></inline-formula> as a function of <inline-formula><mml:math id="M338" display="inline"><mml:mi>K</mml:mi></mml:math></inline-formula>. As expected from Eq. (<xref ref-type="disp-formula" rid="Ch1.E12"/>), the initial condition ensemble
converges more quickly than the multi-model ensemble due to its smaller variance.  Note the excellent agreement with Eq. (<xref ref-type="disp-formula" rid="Ch1.E12"/>) (red
curves), where <inline-formula><mml:math id="M339" display="inline"><mml:mrow><mml:msup><mml:mi mathvariant="italic">σ</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula> has been estimated as the variance over time and all ensemble members and <inline-formula><mml:math id="M340" display="inline"><mml:mrow><mml:msup><mml:mi mathvariant="italic">μ</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula> likewise estimated from the ensemble
mean over all ensemble members.  The large spread for the CMIP5 ensemble
is due to the well-known fact that the bias in global mean temperature is different for different models <xref ref-type="bibr" rid="bib1.bibx52" id="paren.54"/>, which led to a
breakdown of the condition of independence. This is not the case for the initial condition ensemble (see also Table <xref ref-type="table" rid="Ch1.T2"/>).
Smaller spread is obtained for the CMIP5 ensemble if each model is
centred to its own (and not the ensemble) mean in the first 10 <inline-formula><mml:math id="M341" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">years</mml:mi></mml:mrow></mml:math></inline-formula>.</p>
</sec>
</sec>
<sec id="Ch1.S5" sec-type="conclusions">
  <label>5</label><title>Conclusions</title>
      <p id="d1e7632">It is well known that the number of samples necessary for a given
coverage increases exponentially with the dimension.  In this paper we
have described other more non-intuitive properties of high-dimensional
space such as the concentration of measures and waist concentration. In
loose terms these properties state that independent sample members from
the same distribution have the same lengths and that pairs of independent
sample members are orthogonal.  While most results are derived for iid
random variables, we discussed the extension to the non-iid situation and how the strength of the dependence is related to the effective dimension.</p>
      <p id="d1e7635">We directly investigated to which extent these properties hold for
typical climate fields and time series.  Ensemble modelling provides an obvious source of samples, but samples can also be obtained by
considering e.g. different days or years.  We investigated the monthly climatology of both an initial condition ensemble and a multi-model
ensemble.  We also investigated fields of daily means from a reanalysis.
While the nominal dimensions of such fields are high, the effective
dimensions are typically of the order 25–100, and it is not obvious
to which degree the properties of high-dimensional dimension apply to
such fields.</p>
      <p id="d1e7638">We found that for the global-scale fields of near-surface temperature and precipitation, both the concentration of measures and the  waist concentration hold to a reasonable degree. The lengths of the sample
members are rather narrowly distributed around the mean length, with widths (standard deviation) around <inline-formula><mml:math id="M342" display="inline"><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>/</mml:mo><mml:mn mathvariant="normal">5</mml:mn></mml:mrow></mml:math></inline-formula>–<inline-formula><mml:math id="M343" display="inline"><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>/</mml:mo><mml:mn mathvariant="normal">10</mml:mn></mml:mrow></mml:math></inline-formula> of the mean value. The
angles between pairs of sample members are also rather narrowly
distributed around <inline-formula><mml:math id="M344" display="inline"><mml:mrow><mml:mi mathvariant="italic">π</mml:mi><mml:mo>/</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:mrow></mml:math></inline-formula>. This holds both when the samples consist of
the climatology of different ensemble members from a model and when the
samples consist of different daily means from a reanalysis.</p>
      <p id="d1e7677">Regarding the model ensembles, the concentration properties are better
fulfilled for the initial condition ensemble (MPI-GE) than for the
multi-model ensemble (CMIP5). In the latter case the dependence of
related models will result in these models being far from  orthogonal.</p>
      <p id="d1e7681">Based on the concentration properties, we derived simple analytical results that hold for large dimensions.  These analytical results
include (1) the distances between two sample members are a factor of <inline-formula><mml:math id="M345" display="inline"><mml:msqrt><mml:mn mathvariant="normal">2</mml:mn></mml:msqrt></mml:math></inline-formula> larger than the distance between sample members and the sample mean.
(2) The correlations between differences of pair of sample members are <inline-formula><mml:math id="M346" display="inline"><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>/</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:mrow></mml:math></inline-formula>, while the correlations between differences of sample members and the sample
mean are <inline-formula><mml:math id="M347" display="inline"><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>/</mml:mo><mml:msqrt><mml:mn mathvariant="normal">2</mml:mn></mml:msqrt></mml:mrow></mml:math></inline-formula>. (3) An expression for how the sample mean depends
on the sample size and on the sample spread. We found that these results
describe the behaviour of the climate fields reasonably well.</p>
      <p id="d1e7717">We conclude that in many cases the concentration properties allow
us a deeper understanding the behaviour of samples of climate
fields. However, in each case it is important to investigate whether the conditions of high dimensionality and independence are fulfilled. Even for
global fields there is a substantial spread around the values predicted
for the high-dimensional limit.</p>
      <?pagebreak page420?><p id="d1e7720">We have only briefly mentioned the relation between observations and
models. The relation depends on whether we assume that observations
are drawn from the same distribution as the model ensemble
(the indistinguishable interpretation) or whether we assume
that the ensemble members are centred around the observations (truth-centred interpretation). In the former case the results for
individual model members also hold for observations, as we discussed in Sect. <xref ref-type="sec" rid="Ch1.S4"/>.2, while in the latter case results
may be different.  Many of the simple analytical results can be
extended to situations where e.g. the models are biased as explored in <xref ref-type="bibr" rid="bib1.bibx15" id="text.55"/> using a simple statistical model that included
both interpretations as limits.</p>
</sec>

      
      </body>
    <back><notes notes-type="codeavailability"><title>Code availability</title>

      <p id="d1e7732">The Interactive Data Language (IDL) code used in the analysis can be requested from the author.</p>
  </notes><notes notes-type="dataavailability"><title>Data availability</title>

      <p id="d1e7738">The AgERA reanalysis was downloaded from <ext-link xlink:href="https://doi.org/10.24381/cds.6c68c9bb" ext-link-type="DOI">10.24381/cds.6c68c9bb</ext-link> <xref ref-type="bibr" rid="bib1.bibx19" id="paren.56"/>.</p>

      <p id="d1e7747">The MPI Grand Ensemble Project
<uri>https://www.mpimet.mpg.de/en/grand-ensemble/</uri> (last access: 23 August 2021, <xref ref-type="bibr" rid="bib1.bibx38" id="altparen.57"/>, <ext-link xlink:href="https://doi.org/10.1029/2019MS001639" ext-link-type="DOI">10.1029/2019MS001639</ext-link>) was downloaded via <xref ref-type="bibr" rid="bib1.bibx21" id="text.58"/> from
<uri>https://esgf-data.dkrz.de/projects/esgf-dkrz/</uri> (last access: 23 August 2021).</p>

      <p id="d1e7765">The CMIP5 data were downloaded from <uri>https://esgf-node.llnl.gov/projects/esgf-llnl/</uri> (last access: 23 August 2021, <xref ref-type="bibr" rid="bib1.bibx20" id="altparen.59"/>).</p>
  </notes><notes notes-type="competinginterests"><title>Competing interests</title>

      <p id="d1e7777">The author declares that there is no conflict of interest.</p>
  </notes><notes notes-type="disclaimer"><title>Disclaimer</title>

      <p id="d1e7783">Publisher's note: Copernicus Publications remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.</p>
  </notes><ack><title>Acknowledgements</title><p id="d1e7789">The author acknowledges the support from the NordForsk-funded Nordic Centre of Excellence and the European Union.</p><p id="d1e7791">We acknowledge the World Climate Research Programme's Working Group
on Coupled Modelling, which is responsible for CMIP, and we thank the
climate modeling groups for producing and making available their model
output. For CMIP the U.S. Department of Energy's Program for Climate
Model Diagnosis and Intercomparison provided coordinating support and led development of software infrastructure in partnership with the Global
Organization for Earth System Science Portals.</p></ack><notes notes-type="financialsupport"><title>Financial support</title>

      <p id="d1e7796">This research has been supported by the NordForsk (award no. 76654) and the Horizon 2020 (EUCP (grant no. 776613)).</p>
  </notes><notes notes-type="reviewstatement"><title>Review statement</title>

      <p id="d1e7802">This paper was edited by Stéphane Vannitsem and reviewed by Maarten Ambaum and one anonymous referee.</p>
  </notes><ref-list>
    <title>References</title>

      <ref id="bib1.bibx1"><?xmltex \def\ref@label{Abramowitz et~al.(2019)}?><label>Abramowitz et al.(2019)</label><?label Abramowitz2019?><mixed-citation>Abramowitz, G., Herger, N., Gutmann, E., Hammerling, D., Knutti, R., Leduc, M., Lorenz, R., Pincus, R., and Schmidt, G. A.: ESD Reviews: Model dependence in multi-model climate ensembles: weighting, sub-selection and out-of-sample testing, Earth Syst. Dynam., 10, 91–105, <ext-link xlink:href="https://doi.org/10.5194/esd-10-91-2019" ext-link-type="DOI">10.5194/esd-10-91-2019</ext-link>, 2019.</mixed-citation></ref>
      <ref id="bib1.bibx2"><?xmltex \def\ref@label{Annan and Hargreaves(2010)}?><label>Annan and Hargreaves(2010)</label><?label Annan2010?><mixed-citation>Annan, J. D. and Hargreaves, J. C.: Reliability of the CMIP3 ensemble, Geophys. Res. Lett., 37, L02703, <ext-link xlink:href="https://doi.org/10.1029/2009GL041994" ext-link-type="DOI">10.1029/2009GL041994</ext-link>, 2010.</mixed-citation></ref>
      <ref id="bib1.bibx3"><?xmltex \def\ref@label{Bartlett(1935)}?><label>Bartlett(1935)</label><?label Bartlett1935?><mixed-citation>Bartlett, M. S.: Some aspects of the time-correlation problem in regard to tests of significance, J. R. Stat. Soc., 98, 536–543, <ext-link xlink:href="https://doi.org/10.2307/2342284" ext-link-type="DOI">10.2307/2342284</ext-link>, 1935.</mixed-citation></ref>
      <ref id="bib1.bibx4"><?xmltex \def\ref@label{Bengtsson and Hodges(2019)}?><label>Bengtsson and Hodges(2019)</label><?label Bengtsson2019?><mixed-citation>Bengtsson, L. and Hodges, K. I.: Can an ensemble climate simulation be used to separate climate change signals from internal unforced variability?, Clim. Dynam., 52, 3553–3573, <ext-link xlink:href="https://doi.org/10.1007/s00382-018-4343-8" ext-link-type="DOI">10.1007/s00382-018-4343-8</ext-link>, 2019.</mixed-citation></ref>
      <ref id="bib1.bibx5"><?xmltex \def\ref@label{Bishop(2007)}?><label>Bishop(2007)</label><?label Bishop2007?><mixed-citation>
Bishop, C.: Pattern recognition and machine learning (Information science and statistics), Springer-Verlag New York, Inc., Secaucus, NJ, USA, 2nd edn., 2007.</mixed-citation></ref>
      <ref id="bib1.bibx6"><?xmltex \def\ref@label{Bishop and Abramowitz(2013)}?><label>Bishop and Abramowitz(2013)</label><?label Bishop2013?><mixed-citation>Bishop, C. H. and Abramowitz, G.: Climate model dependence and the replicate Earth paradigm, Clim. Dynam., 41, 885–900, <ext-link xlink:href="https://doi.org/10.1007/s00382-012-1610-y" ext-link-type="DOI">10.1007/s00382-012-1610-y</ext-link>, 2013.</mixed-citation></ref>
      <ref id="bib1.bibx7"><?xmltex \def\ref@label{Blum et~al.(2020)}?><label>Blum et al.(2020)</label><?label Blum2017?><mixed-citation>Blum, A., Hopcroft, J., and Kannan, R.: Foundations of data science, Cambridge University Press, Cambridge, UK, available at: <uri>https://www.cs.cornell.edu/jeh/book.pdf</uri> (last access: 23 August 2021),  2020.</mixed-citation></ref>
      <ref id="bib1.bibx8"><?xmltex \def\ref@label{Bo\'{e}(2018)}?><label>Boé(2018)</label><?label Boe2018?><mixed-citation>Boé, J.: Interdependency in multimodel climate projections: Component replication and result similarity, Geophys. Res. Lett., 45, 2771–2779, <ext-link xlink:href="https://doi.org/10.1002/2017GL076829" ext-link-type="DOI">10.1002/2017GL076829</ext-link>, 2018.</mixed-citation></ref>
      <ref id="bib1.bibx9"><?xmltex \def\ref@label{Bretherton et~al.(1999)}?><label>Bretherton et al.(1999)</label><?label Bretherton1999?><mixed-citation>Bretherton, C. S., Widmann, M., Dymnikov, V. P., Wallace, J. M., and Bladé, I.: The effective number of spatial degrees of freedom of a time-varying field, J. Climate, 12, 1990–2009, <ext-link xlink:href="https://doi.org/10.1175/1520-0442(1999)012&lt;1990:TENOSD&gt;2.0.CO;2" ext-link-type="DOI">10.1175/1520-0442(1999)012&lt;1990:TENOSD&gt;2.0.CO;2</ext-link>, 1999.</mixed-citation></ref>
      <ref id="bib1.bibx10"><?xmltex \def\ref@label{Briffa and Jones(1993)}?><label>Briffa and Jones(1993)</label><?label Briffa1993?><mixed-citation> Briffa, K. R. and Jones, P. D.: Global surface air temperature variations during the twentieth century: Part 2, implications for large-scale high-frequency palaeoclimatic studies, Holocene, 3, 77–88, 1993.</mixed-citation></ref>
      <ref id="bib1.bibx11"><?xmltex \def\ref@label{Chazottes(2015)}?><label>Chazottes(2015)</label><?label Chazottes2015?><mixed-citation> Chazottes, J.-R.: Fluctuations of observables in dynamical systems: from limit theorems to concentration inequalities, in: Nonlinear Dynamics New Directions,  Springer, Cham, Switzerland, 47–85, 2015.</mixed-citation></ref>
      <ref id="bib1.bibx12"><?xmltex \def\ref@label{Cherkassky and Mulier(2007)}?><label>Cherkassky and Mulier(2007)</label><?label Cherkassky2007?><mixed-citation> Cherkassky, V. S. and Mulier, F.: Learning from data: concepts, theory, and methods, John Wiley and Sons, Hoboken, N.J, 2nd edn., 2007.</mixed-citation></ref>
      <ref id="bib1.bibx13"><?xmltex \def\ref@label{Christiansen(2018)}?><label>Christiansen(2018)</label><?label Christiansen2018?><mixed-citation>Christiansen, B.: Ensemble averaging and the curse of dimensionality, J. Climate, 31, 1587–1596, <ext-link xlink:href="https://doi.org/10.1175/JCLI-D-17-0197.1" ext-link-type="DOI">10.1175/JCLI-D-17-0197.1</ext-link>, 2018.</mixed-citation></ref>
      <ref id="bib1.bibx14"><?xmltex \def\ref@label{Christiansen(2019)}?><label>Christiansen(2019)</label><?label Christiansen2019?><mixed-citation>Christiansen, B.: Analysis of ensemble mean forecasts: The blessings of high dimensionality, Mon. Weather Rev., 147, 1699–1712, <ext-link xlink:href="https://doi.org/10.1175/MWR-D-18-0211.1" ext-link-type="DOI">10.1175/MWR-D-18-0211.1</ext-link>, 2019.</mixed-citation></ref>
      <?pagebreak page421?><ref id="bib1.bibx15"><?xmltex \def\ref@label{Christiansen(2020)}?><label>Christiansen(2020)</label><?label Christiansen2020?><mixed-citation>Christiansen, B.: Understanding the distribution of multi-model ensembles, J. Climate, 33, 9447–9465, <ext-link xlink:href="https://doi.org/10.1175/JCLI-D-20-0186.1" ext-link-type="DOI">10.1175/JCLI-D-20-0186.1</ext-link>, 2020.</mixed-citation></ref>
      <ref id="bib1.bibx16"><?xmltex \def\ref@label{Christiansen and Ljungqvist(2017)}?><label>Christiansen and Ljungqvist(2017)</label><?label Christiansen2017?><mixed-citation>Christiansen, B. and Ljungqvist, F. C.: Challenges and perspectives for large-scale temperature reconstructions of the past two millennia, Rev. Geophys., 2016RG000521, <ext-link xlink:href="https://doi.org/10.1002/2016RG000521" ext-link-type="DOI">10.1002/2016RG000521</ext-link>, 2017.</mixed-citation></ref>
      <ref id="bib1.bibx17"><?xmltex \def\ref@label{Clusel and Bertin(2008)}?><label>Clusel and Bertin(2008)</label><?label Clusel2008?><mixed-citation>Clusel, M. and Bertin, E.: Global fluctuations in physical systems: a subtle interplay between sum and extreme value statistics, Int. J. Mod. Phys. B, 22, 3311–3368, <ext-link xlink:href="https://doi.org/10.1142/S021797920804853X" ext-link-type="DOI">10.1142/S021797920804853X</ext-link>, 2008.</mixed-citation></ref>
      <ref id="bib1.bibx18"><?xmltex \def\ref@label{Crack and Ledoit(2010)}?><label>Crack and Ledoit(2010)</label><?label Crack2009?><mixed-citation>
Crack, T. F. and Ledoit, O.: Central limit theorems when data are dependent: Addressing the pedagogical gaps, Journal of Financial Education, 36, 38–60, 2010.</mixed-citation></ref>
      <ref id="bib1.bibx19"><?xmltex \def\ref@label{ECMWF(2021)}?><label>ECMWF(2021)</label><?label ECMWF?><mixed-citation>ECMWF: Daily surface meteorological data set for agronomic use, based on ERA5, ECMWF [dat set], <ext-link xlink:href="https://doi.org/10.24381/cds.6c68c9bb" ext-link-type="DOI">10.24381/cds.6c68c9bb</ext-link>, 2021.</mixed-citation></ref>
      <ref id="bib1.bibx20"><?xmltex \def\ref@label{ESGF(2021a)}?><label>ESGF(2021a)</label><?label ESGFa?><mixed-citation>ESGF: Coupled Model Intercomparison Project – Phase 5, World Climate Research Programme (WCRP), ESGF [dat set], available at: <uri>https://esgf-node.llnl.gov/projects/esgf-llnl/</uri>, last access: 23 August 2021a.</mixed-citation></ref>
      <ref id="bib1.bibx21"><?xmltex \def\ref@label{ESGF(2021b)}?><label>ESGF(2021b)</label><?label ESGFb?><mixed-citation>ESGF (Earth System Grid Federation): ESGF-CoG Node, DKRZ (German Climate Computing Centre), available at: <uri>https://esgf-data.dkrz.de/projects/esgf-dkrz/</uri>, last access: 23 August 2021b.</mixed-citation></ref>
      <ref id="bib1.bibx22"><?xmltex \def\ref@label{Flato et~al.(2013)}?><label>Flato et al.(2013)</label><?label Flato2013?><mixed-citation>Flato, G., Marotzke, J., Abiodun, B., Braconnot, P., Chou, S. C., Collins, W. J., Cox, P., Driouech, F., Emori, S., Eyring, V., Forest, C., Gleckler, P., Guilyardi, E., Jakob, C., Kattsov, V., Reason, C., and Rummukainen, M.: Evaluation of Climate Models, in: Climate Change 2013. Contribution of Working Group I to the Fifth Assessment Report of the Intergovernmental Panel on Climate Change, edited by: Stocker, T. F., Qin, D., Plattner, G.-K., Tignor, M., Allen, S. K., Boschung, J., Nauels, A., Xia, Y., Bex, V., and Midgley, P. M., Cambridge University Press, Cambridge, UK and New York, NY, USA, chap. 9, 741–866, <ext-link xlink:href="https://doi.org/10.1017/CBO9781107415324.020" ext-link-type="DOI">10.1017/CBO9781107415324.020</ext-link>, 2013.</mixed-citation></ref>
      <ref id="bib1.bibx23"><?xmltex \def\ref@label{Frankcombe et~al.(2018)}?><label>Frankcombe et al.(2018)</label><?label Frankcombe2018?><mixed-citation>Frankcombe, L. M., England, M. H., Kajtar, J. B., Mann, M. E., and Steinman, B. A.: On the choice of ensemble mean for estimating the forced signal in the presence of internal variability, J. Climate, 31, 5681–5693, <ext-link xlink:href="https://doi.org/10.1175/JCLI-D-17-0662.1" ext-link-type="DOI">10.1175/JCLI-D-17-0662.1</ext-link>, 2018.</mixed-citation></ref>
      <ref id="bib1.bibx24"><?xmltex \def\ref@label{G\'{a}lfi et~al.(2019)}?><label>Gálfi et al.(2019)</label><?label Galfi2019?><mixed-citation>Gálfi, V. M., Lucarini, V., and Wouters, J.: A large deviation theory-based analysis of heat waves and cold spells in a simplified model of the general circulation of the atmosphere, J. Stat. Mech.-Theory E., 2019, 033404, <ext-link xlink:href="https://doi.org/10.1088/1742-5468/ab02e8" ext-link-type="DOI">10.1088/1742-5468/ab02e8</ext-link>, 2019.</mixed-citation></ref>
      <ref id="bib1.bibx25"><?xmltex \def\ref@label{Gleckler et~al.(2008)}?><label>Gleckler et al.(2008)</label><?label Gleckler2008?><mixed-citation>Gleckler, P., Taylor, K., and Doutriaux, C.: Performance metrics for climate models, J. Geophys. Res., 113, D06104, <ext-link xlink:href="https://doi.org/10.1029/2007JD008972" ext-link-type="DOI">10.1029/2007JD008972</ext-link>, 2008.</mixed-citation></ref>
      <ref id="bib1.bibx26"><?xmltex \def\ref@label{Gorban and Tyukin(2018)}?><label>Gorban and Tyukin(2018)</label><?label Gorban2018?><mixed-citation>Gorban, A. N. and Tyukin, I. Y.: Blessing of dimensionality: mathematical foundations of the statistical physics of data, Philos. T. Roy. Soc. A, 376, 20170237, <ext-link xlink:href="https://doi.org/10.1098/rsta.2017.0237" ext-link-type="DOI">10.1098/rsta.2017.0237</ext-link>, 2018.</mixed-citation></ref>
      <ref id="bib1.bibx27"><?xmltex \def\ref@label{Hall et~al.(2005)}?><label>Hall et al.(2005)</label><?label Hall2005?><mixed-citation>Hall, P., Marron, J. S., and Neeman, A.: Geometric representation of high dimension, low sample size data, J. R. Stat. Soc. B, 67, 427–444, <ext-link xlink:href="https://doi.org/10.1111/j.1467-9868.2005.00510.x" ext-link-type="DOI">10.1111/j.1467-9868.2005.00510.x</ext-link>, 2005.</mixed-citation></ref>
      <ref id="bib1.bibx28"><?xmltex \def\ref@label{Hansen and Lebedeff(1987)}?><label>Hansen and Lebedeff(1987)</label><?label Hansen1987?><mixed-citation> Hansen, J. and Lebedeff, S.: Global trends of measured surface air temperature, J. Geophys. Res., 92, 13345–13372, 1987.</mixed-citation></ref>
      <ref id="bib1.bibx29"><?xmltex \def\ref@label{Hecht-Nielsen(1990)}?><label>Hecht-Nielsen(1990)</label><?label Hecht-nielsen1990?><mixed-citation>
Hecht-Nielsen, R.:
Neurocomputing, Addison-Wesley, Reading, Massachusetts, 1990.</mixed-citation></ref>
      <ref id="bib1.bibx30"><?xmltex \def\ref@label{Herger et~al.(2018)}?><label>Herger et al.(2018)</label><?label Herger2018?><mixed-citation>Herger, N., Abramowitz, G., Knutti, R., Angélil, O., Lehmann, K., and Sanderson, B. M.: Selecting a climate model subset to optimise key ensemble properties, Earth Syst. Dynam., 9, 135–151, <ext-link xlink:href="https://doi.org/10.5194/esd-9-135-2018" ext-link-type="DOI">10.5194/esd-9-135-2018</ext-link>, 2018.</mixed-citation></ref>
      <ref id="bib1.bibx31"><?xmltex \def\ref@label{Hersbach et~al.(2019)}?><label>Hersbach et al.(2019)</label><?label Hersbach2019?><mixed-citation>Hersbach, H., Bell, W.,
Berrisford, P., Horányi, A., J., M.-S., Nicolas, J., Radu, R., Schepers, D., Simmons, A., Soci, C., and Dee, D.: Global reanalysis: goodbye ERA-Interim, hello ERA5, ECMWF Newsletter, 159, 17–24,
<ext-link xlink:href="https://doi.org/10.21957/vf291hehd7" ext-link-type="DOI">10.21957/vf291hehd7</ext-link>, 2019.</mixed-citation></ref>
      <ref id="bib1.bibx32"><?xmltex \def\ref@label{Kab\'{a}n(2012)}?><label>Kabán(2012)</label><?label Kaban2012?><mixed-citation>Kabán, A.: Non-parametric detection of meaningless distances in high dimensional data, Stat. Comput., 22, 375–385, <ext-link xlink:href="https://doi.org/10.1007/s11222-011-9229-0" ext-link-type="DOI">10.1007/s11222-011-9229-0</ext-link>, 2012.</mixed-citation></ref>
      <ref id="bib1.bibx33"><?xmltex \def\ref@label{Kainen(1997)}?><label>Kainen(1997)</label><?label Kainen1997?><mixed-citation>Kainen, P. C.: Utilizing geometric anomalies of high dimension: When complexity makes computation easier, in: Computer intensive methods in control and signal processing, pp. 283–294, Birkhäuser, Boston, MA, <ext-link xlink:href="https://doi.org/10.1007/978-1-4612-1996-5_18" ext-link-type="DOI">10.1007/978-1-4612-1996-5_18</ext-link>, 1997.</mixed-citation></ref>
      <ref id="bib1.bibx34"><?xmltex \def\ref@label{Knutti et~al.(2013)}?><label>Knutti et al.(2013)</label><?label Knutti2013?><mixed-citation>Knutti, R., Masson, D., and Gettelman, A.: Climate model genealogy: Generation CMIP5 and how we got there, Geophys. Res. Lett., 40, 1194–1199, <ext-link xlink:href="https://doi.org/10.1002/grl.50256" ext-link-type="DOI">10.1002/grl.50256</ext-link>, 2013.</mixed-citation></ref>
      <ref id="bib1.bibx35"><?xmltex \def\ref@label{Kontorovich and Ramanan(2008)}?><label>Kontorovich and Ramanan(2008)</label><?label Kontorovich2008?><mixed-citation>Kontorovich, L. and Ramanan, K.: Concentration inequalities for dependent random variables via the martingale method, Ann. Probab., 36, 2126–2158, <ext-link xlink:href="https://doi.org/10.1214/07-AOP384" ext-link-type="DOI">10.1214/07-AOP384</ext-link>, 2008.</mixed-citation></ref>
      <ref id="bib1.bibx36"><?xmltex \def\ref@label{Lehmann and Romano(2005)}?><label>Lehmann and Romano(2005)</label><?label Lehmann2005?><mixed-citation> Lehmann, E. L. and Romano, J. P.: Testing statistical hypotheses, Springer texts in statistics, Springer, New York, 3rd edn., 2005.</mixed-citation></ref>
      <ref id="bib1.bibx37"><?xmltex \def\ref@label{Liang et~al.(2020)}?><label>Liang et al.(2020)</label><?label Liang2020?><mixed-citation>Liang, Y.-C., Kwon, Y.-O., Frankignoul, C., Danabasoglu, G., Yeager, S., Cherchi, A., Gao, Y., Gastineau, G., Ghosh, R., Matei, D., Mecking, J. V., Peano, D., Suo, L., and Tian, T.: Quantification of the Arctic sea ice-driven atmospheric circulation variability in coordinated large ensemble simulations, Geophys. Res. Lett., 47, e2019GL085397, <ext-link xlink:href="https://doi.org/10.1029/2019GL085397" ext-link-type="DOI">10.1029/2019GL085397</ext-link>, 2020.</mixed-citation></ref>
      <ref id="bib1.bibx38"><?xmltex \def\ref@label{Maher et~al.(2019)}?><label>Maher et al.(2019)</label><?label Maher2019?><mixed-citation>Maher, N., Milinski, S., Suarez-Gutierrez, L., Botzet, M., Dobrynin, M., Kornblueh, L., Kröger, J., Takano, Y., Ghosh, R., Hedemann, C., Li, C., Li, H., Manzini, E., Notz, D., Putrasahan, D., Boysen, L., Claussen, M., Ilyina, T., Olonscheck, D., Raddatz, T., Stevens, B., and Marotzke, J.: The Max Planck Institute Grand Ensemble: Enabling the exploration of climate system variability, J. Adv. Model. Earth Sy., 11, 2050–2069, <ext-link xlink:href="https://doi.org/10.1029/2019MS001639" ext-link-type="DOI">10.1029/2019MS001639</ext-link>, 2019 (available at: <uri>https://www.mpimet.mpg.de/en/grand-ensemble/</uri>, last access: 23 August 2021).</mixed-citation></ref>
      <ref id="bib1.bibx39"><?xmltex \def\ref@label{Milinski et~al.(2019)}?><label>Milinski et al.(2019)</label><?label Milinski2020?><mixed-citation>Milinski, S., Maher, N., and Olonscheck, D.: How large does a large ensemble need to be?, Earth Syst. Dynam., 11, 885–901, <ext-link xlink:href="https://doi.org/10.5194/esd-11-885-2020" ext-link-type="DOI">10.5194/esd-11-885-2020</ext-link>, 2020.</mixed-citation></ref>
      <ref id="bib1.bibx40"><?xmltex \def\ref@label{Mokkadem(1988)}?><label>Mokkadem(1988)</label><?label Mokkadem1988?><mixed-citation>Mokkadem, A.: Mixing properties of ARMA processes, Stoch. Proc. Appl., 29, 309–315, <ext-link xlink:href="https://doi.org/10.1016/0304-4149(88)90045-2" ext-link-type="DOI">10.1016/0304-4149(88)90045-2</ext-link>, 1988.</mixed-citation></ref>
      <ref id="bib1.bibx41"><?xmltex \def\ref@label{Palmer et~al.(2006)}?><label>Palmer et al.(2006)</label><?label Palmer2005?><mixed-citation>
Palmer, T., Buizza, R., Hagedorn, R., Lorenze, A., Leutbecher, M., and Lenny, S.: Ensemble prediction: A pedagogical perspective, ECMWF Newsletter, 106, 10–17, 2006.</mixed-citation></ref>
      <ref id="bib1.bibx42"><?xmltex \def\ref@label{Pennell and Reichler(2011)}?><label>Pennell and Reichler(2011)</label><?label Pennell2011?><mixed-citation>Pennell, C. and Reichler, T.: On the effective number of climate models, J. Climate, 24, 2358–2367, <ext-link xlink:href="https://doi.org/10.1175/2010JCLI3814.1" ext-link-type="DOI">10.1175/2010JCLI3814.1</ext-link>, 2011.</mixed-citation></ref>
      <ref id="bib1.bibx43"><?xmltex \def\ref@label{Potempski and Galmarini(2009)}?><label>Potempski and Galmarini(2009)</label><?label Potempski2009?><mixed-citation>Potempski, S. and Galmarini, S.: <italic>Est modus in rebus</italic>: analytical properties of multi-model ensembles, Atmos. Chem. Phys., 9, 9471–9489, <ext-link xlink:href="https://doi.org/10.5194/acp-9-9471-2009" ext-link-type="DOI">10.5194/acp-9-9471-2009</ext-link>, 2009.</mixed-citation></ref>
      <?pagebreak page422?><ref id="bib1.bibx44"><?xmltex \def\ref@label{Shen et~al.(1994)}?><label>Shen et al.(1994)</label><?label Shen1994?><mixed-citation> Shen, S. S. P., North, G. R., and Kim, K.-Y.: Spectral approach to optimal estimation of the global average temperature, J. Climate, 7, 1999–2007, 1994.</mixed-citation></ref>
      <ref id="bib1.bibx45"><?xmltex \def\ref@label{Talagrand(1996)}?><label>Talagrand(1996)</label><?label Talagrand1996?><mixed-citation> Talagrand, M.: A new look at independence, Ann. Probab., 24, 1–34, 1996.</mixed-citation></ref>
      <ref id="bib1.bibx46"><?xmltex \def\ref@label{Toma\v{s}ev and Radovanovi\'{c}(2016)}?><label>Tomašev and Radovanović(2016)</label><?label Tomasev2016?><mixed-citation>Tomašev, N. and Radovanović, M.: Clustering Evaluation in High-Dimensional Data, in: Unsupervised Learning Algorithms, edited by: Celebi, M. E. and Aydin, K., pp. 71–107,  Springer, Cham,  <ext-link xlink:href="https://doi.org/10.1007/978-3-319-24211-8_4" ext-link-type="DOI">10.1007/978-3-319-24211-8_4</ext-link>, 2016.</mixed-citation></ref>
      <ref id="bib1.bibx47"><?xmltex \def\ref@label{Touchette(2009)}?><label>Touchette(2009)</label><?label Touchette2009?><mixed-citation>Touchette, H.: The large deviation approach to statistical mechanics, Phys. Rep., 478, 1–69, <ext-link xlink:href="https://doi.org/10.1016/j.physrep.2009.05.002" ext-link-type="DOI">10.1016/j.physrep.2009.05.002</ext-link>, 2009.</mixed-citation></ref>
      <ref id="bib1.bibx48"><?xmltex \def\ref@label{van Loon et~al.(2007)}?><label>van Loon et al.(2007)</label><?label vanLoon2007?><mixed-citation>van Loon, M., Vautard, R., Schaap, M., Bergström, R., Bessagnet, B., Brandt, J., Builtjes, P., Christensen, J., Cuvelier, C., Graff, A., Jonson, J., Krol, M., Langner, J., Roberts, P., Rouil, L., Stern, R., Tarrasón, L., Thunis, P., Vignati, E., White, L., and Wind, P.: Evaluation of long-term ozone simulations from seven regional air quality models and their ensemble, Atmos. Environ., 41, 2083–2097, <ext-link xlink:href="https://doi.org/10.1016/j.atmosenv.2006.10.073" ext-link-type="DOI">10.1016/j.atmosenv.2006.10.073</ext-link>, 2007.
</mixed-citation></ref><?xmltex \hack{\newpage}?>
      <ref id="bib1.bibx49"><?xmltex \def\ref@label{Vershynin(2018)}?><label>Vershynin(2018)</label><?label Vershynin2018?><mixed-citation>Vershynin, R.: High-dimensional probability – Probability theory and stochastic processes, Cambridge University Press, Cambridge, <ext-link xlink:href="https://doi.org/10.1017/9781108231596" ext-link-type="DOI">10.1017/9781108231596</ext-link>, 2018.</mixed-citation></ref>
      <ref id="bib1.bibx50"><?xmltex \def\ref@label{von Storch and Zwiers(1999)}?><label>von Storch and Zwiers(1999)</label><?label vonStorch1999?><mixed-citation>
von Storch, H. and Zwiers, F. W.: Statistical analysis in climate research, Cambridge University Press, Cambridge, ISBN 0 521 45071 3, 1999.</mixed-citation></ref>
      <ref id="bib1.bibx51"><?xmltex \def\ref@label{Wainwright(2019)}?><label>Wainwright(2019)</label><?label Wainwright2019?><mixed-citation>Wainwright, M. J.: High-dimensional statistics: A non-asymptotic viewpoint, Cambridge Series in Statistical and Probabilistic Mathematics, Cambridge University Press, Cambridge, <ext-link xlink:href="https://doi.org/10.1017/9781108627771" ext-link-type="DOI">10.1017/9781108627771</ext-link>, 2019.</mixed-citation></ref>
      <ref id="bib1.bibx52"><?xmltex \def\ref@label{Wang et~al.(2014)}?><label>Wang et al.(2014)</label><?label Wang2014?><mixed-citation>Wang, C., Zhang, L., Lee, S.-K., Wu, L., and Mechoso, C. R.: A global perspective on CMIP5 climate model biases, Nat. Clim. Change, 4, 201–205, <ext-link xlink:href="https://doi.org/10.1038/nclimate2118" ext-link-type="DOI">10.1038/nclimate2118</ext-link>, 2014.</mixed-citation></ref>
      <ref id="bib1.bibx53"><?xmltex \def\ref@label{Wang and Shen(1999)}?><label>Wang and Shen(1999)</label><?label Wang1999?><mixed-citation>Wang, X. and Shen, S. S.: Estimation of spatial degrees of freedom of a climate field, J. Climate, 12, 1280–1291, <ext-link xlink:href="https://doi.org/10.1175/1520-0442(1999)012&lt;1280:EOSDOF&gt;2.0.CO;2" ext-link-type="DOI">10.1175/1520-0442(1999)012&lt;1280:EOSDOF&gt;2.0.CO;2</ext-link>, 1999.</mixed-citation></ref>

  </ref-list></back>
    <!--<article-title-html>The blessing of dimensionality for the analysis of climate data</article-title-html>
<abstract-html><p>We give a simple description of the blessing of dimensionality with
the main focus on the concentration phenomena. These phenomena imply that in
high dimensions the lengths of independent random vectors from the same distribution have almost the same length and that independent vectors
are almost orthogonal. In the climate and atmospheric sciences we rely increasingly on ensemble modelling and face the challenge of analysing
large samples of long time series and spatially extended fields. We show how the properties of high dimensions allow us to obtain analytical
results for e.g. correlations between sample members and the behaviour of the sample mean when the size of the sample grows. We find
that the properties of high dimensionality with reasonable success can be
applied to climate data. This is the case although most climate
data show strong anisotropy and both spatial and temporal dependence, resulting in effective dimensions around 25–100.</p></abstract-html>
<ref-html id="bib1.bib1"><label>Abramowitz et al.(2019)</label><mixed-citation>
Abramowitz, G., Herger, N., Gutmann, E., Hammerling, D., Knutti, R., Leduc, M., Lorenz, R., Pincus, R., and Schmidt, G. A.: ESD Reviews: Model dependence in multi-model climate ensembles: weighting, sub-selection and out-of-sample testing, Earth Syst. Dynam., 10, 91–105, <a href="https://doi.org/10.5194/esd-10-91-2019" target="_blank">https://doi.org/10.5194/esd-10-91-2019</a>, 2019.
</mixed-citation></ref-html>
<ref-html id="bib1.bib2"><label>Annan and Hargreaves(2010)</label><mixed-citation>
Annan, J. D. and Hargreaves, J. C.: Reliability of the CMIP3 ensemble, Geophys. Res. Lett., 37, L02703, <a href="https://doi.org/10.1029/2009GL041994" target="_blank">https://doi.org/10.1029/2009GL041994</a>, 2010.
</mixed-citation></ref-html>
<ref-html id="bib1.bib3"><label>Bartlett(1935)</label><mixed-citation>
Bartlett, M. S.: Some aspects of the time-correlation problem in regard to tests of significance, J. R. Stat. Soc., 98, 536–543, <a href="https://doi.org/10.2307/2342284" target="_blank">https://doi.org/10.2307/2342284</a>, 1935.
</mixed-citation></ref-html>
<ref-html id="bib1.bib4"><label>Bengtsson and Hodges(2019)</label><mixed-citation>
Bengtsson, L. and Hodges, K. I.: Can an ensemble climate simulation be used to separate climate change signals from internal unforced variability?, Clim. Dynam., 52, 3553–3573, <a href="https://doi.org/10.1007/s00382-018-4343-8" target="_blank">https://doi.org/10.1007/s00382-018-4343-8</a>, 2019.
</mixed-citation></ref-html>
<ref-html id="bib1.bib5"><label>Bishop(2007)</label><mixed-citation>
Bishop, C.: Pattern recognition and machine learning (Information science and statistics), Springer-Verlag New York, Inc., Secaucus, NJ, USA, 2nd edn., 2007.
</mixed-citation></ref-html>
<ref-html id="bib1.bib6"><label>Bishop and Abramowitz(2013)</label><mixed-citation>
Bishop, C. H. and Abramowitz, G.: Climate model dependence and the replicate Earth paradigm, Clim. Dynam., 41, 885–900, <a href="https://doi.org/10.1007/s00382-012-1610-y" target="_blank">https://doi.org/10.1007/s00382-012-1610-y</a>, 2013.
</mixed-citation></ref-html>
<ref-html id="bib1.bib7"><label>Blum et al.(2020)</label><mixed-citation>
Blum, A., Hopcroft, J., and Kannan, R.: Foundations of data science, Cambridge University Press, Cambridge, UK, available at: <a href="https://www.cs.cornell.edu/jeh/book.pdf" target="_blank"/> (last access: 23 August 2021),  2020.
</mixed-citation></ref-html>
<ref-html id="bib1.bib8"><label>Boé(2018)</label><mixed-citation> Boé, J.: Interdependency in multimodel climate projections: Component replication and result similarity, Geophys. Res. Lett., 45, 2771–2779, <a href="https://doi.org/10.1002/2017GL076829" target="_blank">https://doi.org/10.1002/2017GL076829</a>, 2018.
</mixed-citation></ref-html>
<ref-html id="bib1.bib9"><label>Bretherton et al.(1999)</label><mixed-citation> Bretherton, C. S., Widmann, M., Dymnikov, V. P., Wallace, J. M., and Bladé, I.: The effective number of spatial degrees of freedom of a time-varying field, J. Climate, 12, 1990–2009, <a href="https://doi.org/10.1175/1520-0442(1999)012&lt;1990:TENOSD&gt;2.0.CO;2" target="_blank">https://doi.org/10.1175/1520-0442(1999)012&lt;1990:TENOSD&gt;2.0.CO;2</a>, 1999.
</mixed-citation></ref-html>
<ref-html id="bib1.bib10"><label>Briffa and Jones(1993)</label><mixed-citation> Briffa, K. R. and Jones, P. D.: Global surface air temperature variations during the twentieth century: Part 2, implications for large-scale high-frequency palaeoclimatic studies, Holocene, 3, 77–88, 1993.
</mixed-citation></ref-html>
<ref-html id="bib1.bib11"><label>Chazottes(2015)</label><mixed-citation> Chazottes, J.-R.: Fluctuations of observables in dynamical systems: from limit theorems to concentration inequalities, in: Nonlinear Dynamics New Directions,  Springer, Cham, Switzerland, 47–85, 2015.
</mixed-citation></ref-html>
<ref-html id="bib1.bib12"><label>Cherkassky and Mulier(2007)</label><mixed-citation> Cherkassky, V. S. and Mulier, F.: Learning from data: concepts, theory, and methods, John Wiley and Sons, Hoboken, N.J, 2nd edn., 2007.
</mixed-citation></ref-html>
<ref-html id="bib1.bib13"><label>Christiansen(2018)</label><mixed-citation> Christiansen, B.: Ensemble averaging and the curse of dimensionality, J. Climate, 31, 1587–1596, <a href="https://doi.org/10.1175/JCLI-D-17-0197.1" target="_blank">https://doi.org/10.1175/JCLI-D-17-0197.1</a>, 2018.
</mixed-citation></ref-html>
<ref-html id="bib1.bib14"><label>Christiansen(2019)</label><mixed-citation> Christiansen, B.: Analysis of ensemble mean forecasts: The blessings of high dimensionality, Mon. Weather Rev., 147, 1699–1712, <a href="https://doi.org/10.1175/MWR-D-18-0211.1" target="_blank">https://doi.org/10.1175/MWR-D-18-0211.1</a>, 2019.
</mixed-citation></ref-html>
<ref-html id="bib1.bib15"><label>Christiansen(2020)</label><mixed-citation> Christiansen, B.: Understanding the distribution of multi-model ensembles, J. Climate, 33, 9447–9465, <a href="https://doi.org/10.1175/JCLI-D-20-0186.1" target="_blank">https://doi.org/10.1175/JCLI-D-20-0186.1</a>, 2020.
</mixed-citation></ref-html>
<ref-html id="bib1.bib16"><label>Christiansen and Ljungqvist(2017)</label><mixed-citation> Christiansen, B. and Ljungqvist, F. C.: Challenges and perspectives for large-scale temperature reconstructions of the past two millennia, Rev. Geophys., 2016RG000521, <a href="https://doi.org/10.1002/2016RG000521" target="_blank">https://doi.org/10.1002/2016RG000521</a>, 2017.
</mixed-citation></ref-html>
<ref-html id="bib1.bib17"><label>Clusel and Bertin(2008)</label><mixed-citation> Clusel, M. and Bertin, E.: Global fluctuations in physical systems: a subtle interplay between sum and extreme value statistics, Int. J. Mod. Phys. B, 22, 3311–3368, <a href="https://doi.org/10.1142/S021797920804853X" target="_blank">https://doi.org/10.1142/S021797920804853X</a>, 2008.
</mixed-citation></ref-html>
<ref-html id="bib1.bib18"><label>Crack and Ledoit(2010)</label><mixed-citation>
Crack, T. F. and Ledoit, O.: Central limit theorems when data are dependent: Addressing the pedagogical gaps, Journal of Financial Education, 36, 38–60, 2010.
</mixed-citation></ref-html>
<ref-html id="bib1.bib19"><label>ECMWF(2021)</label><mixed-citation>
ECMWF: Daily surface meteorological data set for agronomic use, based on ERA5, ECMWF [dat set], <a href="https://doi.org/10.24381/cds.6c68c9bb" target="_blank">https://doi.org/10.24381/cds.6c68c9bb</a>, 2021.
</mixed-citation></ref-html>
<ref-html id="bib1.bib20"><label>ESGF(2021a)</label><mixed-citation>
ESGF: Coupled Model Intercomparison Project – Phase 5, World Climate Research Programme (WCRP), ESGF [dat set], available at: <a href="https://esgf-node.llnl.gov/projects/esgf-llnl/" target="_blank"/>, last access: 23 August 2021a.
</mixed-citation></ref-html>
<ref-html id="bib1.bib21"><label>ESGF(2021b)</label><mixed-citation>
ESGF (Earth System Grid Federation): ESGF-CoG Node, DKRZ (German Climate Computing Centre), available at: <a href="https://esgf-data.dkrz.de/projects/esgf-dkrz/" target="_blank"/>, last access: 23 August 2021b.
</mixed-citation></ref-html>
<ref-html id="bib1.bib22"><label>Flato et al.(2013)</label><mixed-citation>
Flato, G., Marotzke, J., Abiodun, B., Braconnot, P., Chou, S. C., Collins, W. J., Cox, P., Driouech, F., Emori, S., Eyring, V., Forest, C., Gleckler, P., Guilyardi, E., Jakob, C., Kattsov, V., Reason, C., and Rummukainen, M.: Evaluation of Climate Models, in: Climate Change 2013. Contribution of Working Group I to the Fifth Assessment Report of the Intergovernmental Panel on Climate Change, edited by: Stocker, T. F., Qin, D., Plattner, G.-K., Tignor, M., Allen, S. K., Boschung, J., Nauels, A., Xia, Y., Bex, V., and Midgley, P. M., Cambridge University Press, Cambridge, UK and New York, NY, USA, chap. 9, 741–866, <a href="https://doi.org/10.1017/CBO9781107415324.020" target="_blank">https://doi.org/10.1017/CBO9781107415324.020</a>, 2013.
</mixed-citation></ref-html>
<ref-html id="bib1.bib23"><label>Frankcombe et al.(2018)</label><mixed-citation> Frankcombe, L. M., England, M. H., Kajtar, J. B., Mann, M. E., and Steinman, B. A.: On the choice of ensemble mean for estimating the forced signal in the presence of internal variability, J. Climate, 31, 5681–5693, <a href="https://doi.org/10.1175/JCLI-D-17-0662.1" target="_blank">https://doi.org/10.1175/JCLI-D-17-0662.1</a>, 2018.
</mixed-citation></ref-html>
<ref-html id="bib1.bib24"><label>Gálfi et al.(2019)</label><mixed-citation> Gálfi, V. M., Lucarini, V., and Wouters, J.: A large deviation theory-based analysis of heat waves and cold spells in a simplified model of the general circulation of the atmosphere, J. Stat. Mech.-Theory E., 2019, 033404, <a href="https://doi.org/10.1088/1742-5468/ab02e8" target="_blank">https://doi.org/10.1088/1742-5468/ab02e8</a>, 2019.
</mixed-citation></ref-html>
<ref-html id="bib1.bib25"><label>Gleckler et al.(2008)</label><mixed-citation> Gleckler, P., Taylor, K., and Doutriaux, C.: Performance metrics for climate models, J. Geophys. Res., 113, D06104, <a href="https://doi.org/10.1029/2007JD008972" target="_blank">https://doi.org/10.1029/2007JD008972</a>, 2008.
</mixed-citation></ref-html>
<ref-html id="bib1.bib26"><label>Gorban and Tyukin(2018)</label><mixed-citation> Gorban, A. N. and Tyukin, I. Y.: Blessing of dimensionality: mathematical foundations of the statistical physics of data, Philos. T. Roy. Soc. A, 376, 20170237, <a href="https://doi.org/10.1098/rsta.2017.0237" target="_blank">https://doi.org/10.1098/rsta.2017.0237</a>, 2018.
</mixed-citation></ref-html>
<ref-html id="bib1.bib27"><label>Hall et al.(2005)</label><mixed-citation> Hall, P., Marron, J. S., and Neeman, A.: Geometric representation of high dimension, low sample size data, J. R. Stat. Soc. B, 67, 427–444, <a href="https://doi.org/10.1111/j.1467-9868.2005.00510.x" target="_blank">https://doi.org/10.1111/j.1467-9868.2005.00510.x</a>, 2005.
</mixed-citation></ref-html>
<ref-html id="bib1.bib28"><label>Hansen and Lebedeff(1987)</label><mixed-citation> Hansen, J. and Lebedeff, S.: Global trends of measured surface air temperature, J. Geophys. Res., 92, 13345–13372, 1987.
</mixed-citation></ref-html>
<ref-html id="bib1.bib29"><label>Hecht-Nielsen(1990)</label><mixed-citation>
Hecht-Nielsen, R.:
Neurocomputing, Addison-Wesley, Reading, Massachusetts, 1990.
</mixed-citation></ref-html>
<ref-html id="bib1.bib30"><label>Herger et al.(2018)</label><mixed-citation>
Herger, N., Abramowitz, G., Knutti, R., Angélil, O., Lehmann, K., and Sanderson, B. M.: Selecting a climate model subset to optimise key ensemble properties, Earth Syst. Dynam., 9, 135–151, <a href="https://doi.org/10.5194/esd-9-135-2018" target="_blank">https://doi.org/10.5194/esd-9-135-2018</a>, 2018.
</mixed-citation></ref-html>
<ref-html id="bib1.bib31"><label>Hersbach et al.(2019)</label><mixed-citation> Hersbach, H., Bell, W.,
Berrisford, P., Horányi, A., J., M.-S., Nicolas, J., Radu, R., Schepers, D., Simmons, A., Soci, C., and Dee, D.: Global reanalysis: goodbye ERA-Interim, hello ERA5, ECMWF Newsletter, 159, 17–24,
<a href="https://doi.org/10.21957/vf291hehd7" target="_blank">https://doi.org/10.21957/vf291hehd7</a>, 2019.
</mixed-citation></ref-html>
<ref-html id="bib1.bib32"><label>Kabán(2012)</label><mixed-citation> Kabán, A.: Non-parametric detection of meaningless distances in high dimensional data, Stat. Comput., 22, 375–385, <a href="https://doi.org/10.1007/s11222-011-9229-0" target="_blank">https://doi.org/10.1007/s11222-011-9229-0</a>, 2012.
</mixed-citation></ref-html>
<ref-html id="bib1.bib33"><label>Kainen(1997)</label><mixed-citation> Kainen, P. C.: Utilizing geometric anomalies of high dimension: When complexity makes computation easier, in: Computer intensive methods in control and signal processing, pp. 283–294, Birkhäuser, Boston, MA, <a href="https://doi.org/10.1007/978-1-4612-1996-5_18" target="_blank">https://doi.org/10.1007/978-1-4612-1996-5_18</a>, 1997.
</mixed-citation></ref-html>
<ref-html id="bib1.bib34"><label>Knutti et al.(2013)</label><mixed-citation> Knutti, R., Masson, D., and Gettelman, A.: Climate model genealogy: Generation CMIP5 and how we got there, Geophys. Res. Lett., 40, 1194–1199, <a href="https://doi.org/10.1002/grl.50256" target="_blank">https://doi.org/10.1002/grl.50256</a>, 2013.
</mixed-citation></ref-html>
<ref-html id="bib1.bib35"><label>Kontorovich and Ramanan(2008)</label><mixed-citation> Kontorovich, L. and Ramanan, K.: Concentration inequalities for dependent random variables via the martingale method, Ann. Probab., 36, 2126–2158, <a href="https://doi.org/10.1214/07-AOP384" target="_blank">https://doi.org/10.1214/07-AOP384</a>, 2008.
</mixed-citation></ref-html>
<ref-html id="bib1.bib36"><label>Lehmann and Romano(2005)</label><mixed-citation> Lehmann, E. L. and Romano, J. P.: Testing statistical hypotheses, Springer texts in statistics, Springer, New York, 3rd edn., 2005.
</mixed-citation></ref-html>
<ref-html id="bib1.bib37"><label>Liang et al.(2020)</label><mixed-citation> Liang, Y.-C., Kwon, Y.-O., Frankignoul, C., Danabasoglu, G., Yeager, S., Cherchi, A., Gao, Y., Gastineau, G., Ghosh, R., Matei, D., Mecking, J. V., Peano, D., Suo, L., and Tian, T.: Quantification of the Arctic sea ice-driven atmospheric circulation variability in coordinated large ensemble simulations, Geophys. Res. Lett., 47, e2019GL085397, <a href="https://doi.org/10.1029/2019GL085397" target="_blank">https://doi.org/10.1029/2019GL085397</a>, 2020.
</mixed-citation></ref-html>
<ref-html id="bib1.bib38"><label>Maher et al.(2019)</label><mixed-citation>
Maher, N., Milinski, S., Suarez-Gutierrez, L., Botzet, M., Dobrynin, M., Kornblueh, L., Kröger, J., Takano, Y., Ghosh, R., Hedemann, C., Li, C., Li, H., Manzini, E., Notz, D., Putrasahan, D., Boysen, L., Claussen, M., Ilyina, T., Olonscheck, D., Raddatz, T., Stevens, B., and Marotzke, J.: The Max Planck Institute Grand Ensemble: Enabling the exploration of climate system variability, J. Adv. Model. Earth Sy., 11, 2050–2069, <a href="https://doi.org/10.1029/2019MS001639" target="_blank">https://doi.org/10.1029/2019MS001639</a>, 2019 (available at: <a href="https://www.mpimet.mpg.de/en/grand-ensemble/" target="_blank"/>, last access: 23 August 2021).
</mixed-citation></ref-html>
<ref-html id="bib1.bib39"><label>Milinski et al.(2019)</label><mixed-citation> Milinski, S., Maher, N., and Olonscheck, D.: How large does a large ensemble need to be?, Earth Syst. Dynam., 11, 885–901, <a href="https://doi.org/10.5194/esd-11-885-2020" target="_blank">https://doi.org/10.5194/esd-11-885-2020</a>, 2020.
</mixed-citation></ref-html>
<ref-html id="bib1.bib40"><label>Mokkadem(1988)</label><mixed-citation> Mokkadem, A.: Mixing properties of ARMA processes, Stoch. Proc. Appl., 29, 309–315, <a href="https://doi.org/10.1016/0304-4149(88)90045-2" target="_blank">https://doi.org/10.1016/0304-4149(88)90045-2</a>, 1988.
</mixed-citation></ref-html>
<ref-html id="bib1.bib41"><label>Palmer et al.(2006)</label><mixed-citation>
Palmer, T., Buizza, R., Hagedorn, R., Lorenze, A., Leutbecher, M., and Lenny, S.: Ensemble prediction: A pedagogical perspective, ECMWF Newsletter, 106, 10–17, 2006.
</mixed-citation></ref-html>
<ref-html id="bib1.bib42"><label>Pennell and Reichler(2011)</label><mixed-citation>
Pennell, C. and Reichler, T.: On the effective number of climate models, J. Climate, 24, 2358–2367, <a href="https://doi.org/10.1175/2010JCLI3814.1" target="_blank">https://doi.org/10.1175/2010JCLI3814.1</a>, 2011.
</mixed-citation></ref-html>
<ref-html id="bib1.bib43"><label>Potempski and Galmarini(2009)</label><mixed-citation>
Potempski, S. and Galmarini, S.: <i>Est modus in rebus</i>: analytical properties of multi-model ensembles, Atmos. Chem. Phys., 9, 9471–9489, <a href="https://doi.org/10.5194/acp-9-9471-2009" target="_blank">https://doi.org/10.5194/acp-9-9471-2009</a>, 2009.
</mixed-citation></ref-html>
<ref-html id="bib1.bib44"><label>Shen et al.(1994)</label><mixed-citation> Shen, S. S. P., North, G. R., and Kim, K.-Y.: Spectral approach to optimal estimation of the global average temperature, J. Climate, 7, 1999–2007, 1994.
</mixed-citation></ref-html>
<ref-html id="bib1.bib45"><label>Talagrand(1996)</label><mixed-citation> Talagrand, M.: A new look at independence, Ann. Probab., 24, 1–34, 1996.
</mixed-citation></ref-html>
<ref-html id="bib1.bib46"><label>Tomašev and Radovanović(2016)</label><mixed-citation>
Tomašev, N. and Radovanović, M.: Clustering Evaluation in High-Dimensional Data, in: Unsupervised Learning Algorithms, edited by: Celebi, M. E. and Aydin, K., pp. 71–107,  Springer, Cham,  <a href="https://doi.org/10.1007/978-3-319-24211-8_4" target="_blank">https://doi.org/10.1007/978-3-319-24211-8_4</a>, 2016.
</mixed-citation></ref-html>
<ref-html id="bib1.bib47"><label>Touchette(2009)</label><mixed-citation> Touchette, H.: The large deviation approach to statistical mechanics, Phys. Rep., 478, 1–69, <a href="https://doi.org/10.1016/j.physrep.2009.05.002" target="_blank">https://doi.org/10.1016/j.physrep.2009.05.002</a>, 2009.
</mixed-citation></ref-html>
<ref-html id="bib1.bib48"><label>van Loon et al.(2007)</label><mixed-citation> van Loon, M., Vautard, R., Schaap, M., Bergström, R., Bessagnet, B., Brandt, J., Builtjes, P., Christensen, J., Cuvelier, C., Graff, A., Jonson, J., Krol, M., Langner, J., Roberts, P., Rouil, L., Stern, R., Tarrasón, L., Thunis, P., Vignati, E., White, L., and Wind, P.: Evaluation of long-term ozone simulations from seven regional air quality models and their ensemble, Atmos. Environ., 41, 2083–2097, <a href="https://doi.org/10.1016/j.atmosenv.2006.10.073" target="_blank">https://doi.org/10.1016/j.atmosenv.2006.10.073</a>, 2007.

</mixed-citation></ref-html>
<ref-html id="bib1.bib49"><label>Vershynin(2018)</label><mixed-citation>
Vershynin, R.: High-dimensional probability – Probability theory and stochastic processes, Cambridge University Press, Cambridge, <a href="https://doi.org/10.1017/9781108231596" target="_blank">https://doi.org/10.1017/9781108231596</a>, 2018.
</mixed-citation></ref-html>
<ref-html id="bib1.bib50"><label>von Storch and Zwiers(1999)</label><mixed-citation>
von Storch, H. and Zwiers, F. W.: Statistical analysis in climate research, Cambridge University Press, Cambridge, ISBN&thinsp;0&thinsp;521&thinsp;45071&thinsp;3, 1999.
</mixed-citation></ref-html>
<ref-html id="bib1.bib51"><label>Wainwright(2019)</label><mixed-citation>
Wainwright, M. J.: High-dimensional statistics: A non-asymptotic viewpoint, Cambridge Series in Statistical and Probabilistic Mathematics, Cambridge University Press, Cambridge, <a href="https://doi.org/10.1017/9781108627771" target="_blank">https://doi.org/10.1017/9781108627771</a>, 2019.
</mixed-citation></ref-html>
<ref-html id="bib1.bib52"><label>Wang et al.(2014)</label><mixed-citation>
Wang, C., Zhang, L., Lee, S.-K., Wu, L., and Mechoso, C. R.: A global perspective on CMIP5 climate model biases, Nat. Clim. Change, 4, 201–205, <a href="https://doi.org/10.1038/nclimate2118" target="_blank">https://doi.org/10.1038/nclimate2118</a>, 2014.
</mixed-citation></ref-html>
<ref-html id="bib1.bib53"><label>Wang and Shen(1999)</label><mixed-citation>
Wang, X. and Shen, S. S.: Estimation of spatial degrees of freedom of a climate field, J. Climate, 12, 1280–1291, <a href="https://doi.org/10.1175/1520-0442(1999)012&lt;1280:EOSDOF&gt;2.0.CO;2" target="_blank">https://doi.org/10.1175/1520-0442(1999)012&lt;1280:EOSDOF&gt;2.0.CO;2</a>, 1999.
</mixed-citation></ref-html>--></article>
