<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing with OASIS Tables v3.0 20080202//EN" "journalpub-oasis3.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:oasis="http://docs.oasis-open.org/ns/oasis-exchange/table" xml:lang="en" dtd-version="3.0"><?xmltex \makeatother\@nolinetrue\makeatletter?>
  <front>
    <journal-meta><journal-id journal-id-type="publisher">NPG</journal-id><journal-title-group>
    <journal-title>Nonlinear Processes in Geophysics</journal-title>
    <abbrev-journal-title abbrev-type="publisher">NPG</abbrev-journal-title><abbrev-journal-title abbrev-type="nlm-ta">Nonlin. Processes Geophys.</abbrev-journal-title>
  </journal-title-group><issn pub-type="epub">1607-7946</issn><publisher>
    <publisher-name>Copernicus Publications</publisher-name>
    <publisher-loc>Göttingen, Germany</publisher-loc>
  </publisher></journal-meta>
    <article-meta>
      <article-id pub-id-type="doi">10.5194/npg-27-121-2020</article-id><title-group><article-title>Seasonal statistical–dynamical prediction of the North Atlantic Oscillation by probabilistic post-processing and its evaluation</article-title><alt-title>Statistical–dynamical prediction of the NAO</alt-title>
      </title-group><?xmltex \runningtitle{Statistical--dynamical prediction of the NAO}?><?xmltex \runningauthor{A. D\"{u}sterhus}?>
      <contrib-group>
        <contrib contrib-type="author" corresp="yes" rid="aff1 aff2">
          <name><surname>Düsterhus</surname><given-names>André</given-names></name>
          <email>andre.duesterhus@mu.ie</email>
        <ext-link>https://orcid.org/0000-0003-2192-175X</ext-link></contrib>
        <aff id="aff1"><label>1</label><institution>Institute of Oceanography, Center for Earth System Research and Sustainability (CEN), <?xmltex \hack{\break}?>Universität Hamburg, Hamburg, Germany</institution>
        </aff>
        <aff id="aff2"><label>2</label><institution>ICARUS, Department of Geography, Maynooth University, Maynooth, Ireland</institution>
        </aff>
      </contrib-group>
      <author-notes><corresp id="corr1">André Düsterhus (andre.duesterhus@mu.ie)</corresp></author-notes><pub-date><day>27</day><month>February</month><year>2020</year></pub-date>
      
      <volume>27</volume>
      <issue>1</issue>
      <fpage>121</fpage><lpage>131</lpage>
      <history>
        <date date-type="received"><day>30</day><month>September</month><year>2019</year></date>
           <date date-type="rev-request"><day>10</day><month>October</month><year>2019</year></date>
           <date date-type="rev-recd"><day>16</day><month>January</month><year>2020</year></date>
           <date date-type="accepted"><day>17</day><month>January</month><year>2020</year></date>
      </history>
      <permissions>
        <copyright-statement>Copyright: © 2020 André Düsterhus</copyright-statement>
        <copyright-year>2020</copyright-year>
      <license license-type="open-access"><license-p>This work is licensed under the Creative Commons Attribution 4.0 International License. To view a copy of this licence, visit <ext-link ext-link-type="uri" xlink:href="https://creativecommons.org/licenses/by/4.0/">https://creativecommons.org/licenses/by/4.0/</ext-link></license-p></license></permissions><self-uri xlink:href="https://npg.copernicus.org/articles/27/121/2020/npg-27-121-2020.html">This article is available from https://npg.copernicus.org/articles/27/121/2020/npg-27-121-2020.html</self-uri><self-uri xlink:href="https://npg.copernicus.org/articles/27/121/2020/npg-27-121-2020.pdf">The full text article is available as a PDF file from https://npg.copernicus.org/articles/27/121/2020/npg-27-121-2020.pdf</self-uri>
      <abstract><title>Abstract</title>
    <p id="d1e88">Dynamical models of various centres have shown in recent years seasonal prediction skill of the North Atlantic Oscillation (NAO). By filtering the ensemble members on the basis of statistical predictors, known as subsampling, it is possible to achieve even higher prediction skill. In this study the aim is to design a generalisation of the subsampling approach and establish it as a post-processing procedure.</p>
    <p id="d1e91">Instead of selecting discrete ensemble members for each year, as the subsampling approach does, the distributions of ensembles and statistical predictors are combined to create a probabilistic prediction of the winter NAO. By comparing the combined statistical–dynamical prediction with the predictions of its single components, it can be shown that it achieves similar results to the statistical prediction. At the same time it can be shown that, unlike the statistical prediction, the combined prediction has fewer years where it performs worse than the dynamical prediction.</p>
    <p id="d1e94">By applying the gained distributions to other meteorological variables, like geopotential height, precipitation and surface temperature, it can be shown that evaluating prediction skill depends highly on the chosen metric. Besides the common anomaly correlation (ACC) this study also presents scores based on the Earth mover's distance (EMD) and the integrated quadratic distance (IQD), which are designed to evaluate skills of probabilistic predictions. It shows that by evaluating the predictions for each year separately compared to applying a metric to all years at the same time, like correlation-based metrics, leads to different interpretations of the analysis.</p>
  </abstract>
    </article-meta>
  </front>
<body>
      

<sec id="Ch1.S1" sec-type="intro">
  <label>1</label><title>Introduction</title>
      <p id="d1e106">Seasonal prediction of the North Atlantic Oscillation (NAO) is a challenge. During the year the NAO describes a high portion of the explained variability of the pressure field over the North Atlantic region and with it has a high influence on European weather. While the winter NAO (WNAO) is a dominant factor in changes in the storm tracks over the North Atlantic <xref ref-type="bibr" rid="bib1.bibx11" id="paren.1"/>, the summer NAO (SNAO) is associated with precipitation and temperature differences between Scandinavia and the Mediterranean <xref ref-type="bibr" rid="bib1.bibx8" id="paren.2"/>.</p>
      <p id="d1e115">Predicting the WNAO on the seasonal scale is a long-standing aim of the community <xref ref-type="bibr" rid="bib1.bibx5 bib1.bibx12 bib1.bibx16" id="paren.3"/> and various current seasonal prediction systems have demonstrated limited significant correlation skill for the WNAO <xref ref-type="bibr" rid="bib1.bibx2" id="paren.4"/>.
<xref ref-type="bibr" rid="bib1.bibx6" id="text.5"/> have shown that by combining statistical and dynamical predictions, a much higher significant correlation skill is achievable.
This paper applies an ensemble subsampling algorithm, which bases selection of ensemble members on their closeness to statistical predictors. The selected ensembles are then used to create a new sub-selected ensemble mean, which has for the NAO index, but also for many other variables and regions, a better prediction skill than the ensemble mean of all ensemble members.</p>
      <p id="d1e127">Statistical–dynamical predictions based on different strategies are common in many fields in geoscience. <xref ref-type="bibr" rid="bib1.bibx9" id="text.6"/> developed a framework for the dynamical evolution of statistical distributions in phase space with applications to meteorological fields. <xref ref-type="bibr" rid="bib1.bibx19" id="text.7"/> apply a combined statistical–dynamical approach by using an emulator based<?pagebreak page122?> on dynamical forecasts to create seasonal hurricane prediction. <xref ref-type="bibr" rid="bib1.bibx14" id="text.8"/> developed a “best member” concept, which uses verification statistics to dress a dynamical ensemble prediction. Statistical post-processing procedures to enhance forecast skill by dynamical models are applied in various ways in atmospheric science <xref ref-type="bibr" rid="bib1.bibx21" id="paren.9"/>. Especially Bayesian model averaging <xref ref-type="bibr" rid="bib1.bibx13" id="paren.10"/>, which creates weights for ensemble members based on their performance in a training period, has been well established.</p>
      <p id="d1e145">The focus of this paper is to implement the subsampling algorithm as a probabilistic post-processing procedure, demonstrated for the seasonal prediction of the WNAO. In contrast to <xref ref-type="bibr" rid="bib1.bibx6" id="text.11"/>, which worked with deterministic ensemble members, it interprets ensemble members and the statistical predictors as values with uncertainties. The combination of statistical and dynamical models does not happen by selecting the ensemble members directly, but by combinations of probability density functions to create a new probabilistic forecast. This approach allows us to evaluate a prediction skill not only for a long time series, but also for each individual year. We use for this two newly developed skill scores, the 1D-continuous-EMD score and the 1D-continuous-IQD score, based on the Earth mover's distance (EMD) and the integrated quadratic distance (IQD). The WNAO has a severe influence on various meteorological fields over the European continent. Therefore, we also use the probabilistic information of the prediction to create a weighted mean of the ensemble members, which creates a better hindcast skill for important meteorological variables like surface temperature and precipitation.</p>
</sec>
<sec id="Ch1.S2">
  <label>2</label><title>Data and model</title>
      <p id="d1e159">To demonstrate the procedure we use the seasonal prediction system based on the MPI-ESM <xref ref-type="bibr" rid="bib1.bibx6" id="paren.12"/> with a model resolution of T63/L95 (200 km/1.875<inline-formula><mml:math id="M1" display="inline"><mml:msup><mml:mi/><mml:mo>∘</mml:mo></mml:msup></mml:math></inline-formula>, 95 vertical layers) in the atmosphere and T0.4/L40 (40 km/0.4<inline-formula><mml:math id="M2" display="inline"><mml:msup><mml:mi/><mml:mo>∘</mml:mo></mml:msup></mml:math></inline-formula>, 40 vertical layers) in the ocean (also known as mixed resolution, MR). As described by <xref ref-type="bibr" rid="bib1.bibx1" id="text.13"/>, we initialise in each November between 1982 and 2017 a 30 ensemble member hindcast from an assimilation run based on assimilated reanalysis/observations in the atmospheric, oceanic and sea-ice components. As an observational reference we use the ERA-Interim reanalysis <xref ref-type="bibr" rid="bib1.bibx3" id="paren.14"/>.
<?xmltex \hack{\break}?>For the observations and the hindcasts the NAO is calculated by an empirical orthogonal function (EOF) analysis <xref ref-type="bibr" rid="bib1.bibx10" id="paren.15"/>. For the WNAO we calculate the mean sea-level pressure field for December, January and February and calculate the EOF of the North Atlantic sector limited by 20–80<inline-formula><mml:math id="M3" display="inline"><mml:msup><mml:mi/><mml:mo>∘</mml:mo></mml:msup></mml:math></inline-formula> N and 70<inline-formula><mml:math id="M4" display="inline"><mml:msup><mml:mi/><mml:mo>∘</mml:mo></mml:msup></mml:math></inline-formula> W–40<inline-formula><mml:math id="M5" display="inline"><mml:msup><mml:mi/><mml:mo>∘</mml:mo></mml:msup></mml:math></inline-formula> E.</p>
</sec>
<sec id="Ch1.S3">
  <label>3</label><title>Methodology</title>
<sec id="Ch1.S3.SS1">
  <label>3.1</label><title>Seasonal prediction of the WNAO</title>
      <p id="d1e237">The seasonal prediction of the WNAO for the period of 1982 to 2017 is shown in Fig. <xref ref-type="fig" rid="Ch1.F1"/>. Every dot represents one WNAO value of one ensemble member, which has also available the full meteorological and oceanographical fields during the associated winter period. These hindcast predictions for the WNAO have a large spread, covering the range of the observations given by the reanalysis, but do not give indication of a specific NAO value 2 to 4 months ahead. As a general skill measure the community applies correlation skills. Those measures have indicated in recent years significant hindcast skill for several different prediction systems <xref ref-type="bibr" rid="bib1.bibx2" id="paren.16"/>.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F1"><?xmltex \currentcnt{1}?><label>Figure 1</label><caption><p id="d1e247">Seasonal prediction of the WNAO. Single dynamical models (black) initialised in November predicting the DJF-NAO (red).</p></caption>
          <?xmltex \igopts{width=241.848425pt}?><graphic xlink:href="https://npg.copernicus.org/articles/27/121/2020/npg-27-121-2020-f01.png"/>

        </fig>

</sec>
<sec id="Ch1.S3.SS2">
  <label>3.2</label><title>Statistical–dynamical prediction</title>
      <p id="d1e264">Our approach will be applied to every single year independently. As an example we choose the year 2010, which shows an extreme negative WNAO value. The first step is to generate one probability density function (pdf) for each ensemble member prediction (<inline-formula><mml:math id="M6" display="inline"><mml:mrow><mml:msub><mml:mi>e</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>) of the WNAO value, which is generated by a 2000-member bootstrap of the EOF fields <xref ref-type="bibr" rid="bib1.bibx20" id="paren.17"/>. In the bootstrap the first EOF field is recalculated by resampling the mean sea-level pressure fields from each year. To create from these predictions a pdf for all ensemble members (<inline-formula><mml:math id="M7" display="inline"><mml:mi mathvariant="script">E</mml:mi></mml:math></inline-formula>), mixture modelling <xref ref-type="bibr" rid="bib1.bibx17" id="paren.18"/> at discrete NAO index values is applied:
            <disp-formula id="Ch1.E1" content-type="numbered"><label>1</label><mml:math id="M8" display="block"><mml:mrow><mml:mi mathvariant="script">E</mml:mi><mml:mo>(</mml:mo><mml:mi>v</mml:mi><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:munder><mml:mo movablelimits="false">∑</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>∈</mml:mo><mml:mi mathvariant="script">I</mml:mi></mml:mrow></mml:munder><mml:msub><mml:mi>e</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>(</mml:mo><mml:mi>v</mml:mi><mml:mo>)</mml:mo><mml:mo>.</mml:mo></mml:mrow></mml:math></disp-formula>
          Here, <inline-formula><mml:math id="M9" display="inline"><mml:mi>v</mml:mi></mml:math></inline-formula> corresponds to each value of the discretised NAO values and <inline-formula><mml:math id="M10" display="inline"><mml:mi mathvariant="script">I</mml:mi></mml:math></inline-formula> to the indices of the ensemble members. The chosen resolution for the discretised NAO values is 0.01 and,<?pagebreak page123?> after creating the sum of all single-member pdfs, the overall pdf <inline-formula><mml:math id="M11" display="inline"><mml:mi mathvariant="script">E</mml:mi></mml:math></inline-formula> is normalised. For 2010 the results are shown in Fig. <xref ref-type="fig" rid="Ch1.F2"/>. As expected from Fig. <xref ref-type="fig" rid="Ch1.F1"/>, the dynamical model prediction has a very broad pdf equating to a low signal.</p>
      <p id="d1e356">To sharpen the prediction we introduce literature-backed physical statistical predictors. As predictors (<inline-formula><mml:math id="M12" display="inline"><mml:mrow><mml:msub><mml:mi>p</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>) we use those defined by <xref ref-type="bibr" rid="bib1.bibx6" id="text.19"/>: sea-surface temperature in the Northern Hemisphere, Arctic sea-ice volume, Siberian snow cover and stratospheric temperature at 100 hPa. All predictors and their influence on the WNAO have been discussed in the paper. For the physical validity of a prediction the selection of the correct predictors is essential and has to be adapted to any newly analysed phenomena individually. Each predictor makes a prediction from the climatic state taken from the ERA-Interim reanalysis <xref ref-type="bibr" rid="bib1.bibx3" id="paren.20"/> before the initialisation of the dynamical model for a WNAO value in the following winter. For the predictors a normalised index over the hindcast period is calculated by forming the mean over the significantly correlated areas between the physical field and the WNAO index. It has been shown by a real forecast test in <xref ref-type="bibr" rid="bib1.bibx6" id="text.21"/> that this approach is usable also in cases where the predictor is only formed with past information instead of the whole hindcast period.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F2" specific-use="star"><?xmltex \currentcnt{2}?><label>Figure 2</label><caption><p id="d1e381">Dynamical prediction of the WNAO for 2010. Single models (grey) as pdfs of their bootstrapped uncertainties. From this the overall model prediction (black) is created by empirical mixture modelling.</p></caption>
          <?xmltex \igopts{width=341.433071pt}?><graphic xlink:href="https://npg.copernicus.org/articles/27/121/2020/npg-27-121-2020-f02.png"/>

        </fig>

      <?xmltex \floatpos{t}?><fig id="Ch1.F3" specific-use="star"><?xmltex \currentcnt{3}?><label>Figure 3</label><caption><p id="d1e393">Statistical prediction of the WNAO for 2010. Single predictors (light blue) as pdfs of their bootstrapped uncertainties. From this the statistical prediction (pink) is created by empirical mixture modelling.</p></caption>
          <?xmltex \igopts{width=341.433071pt}?><graphic xlink:href="https://npg.copernicus.org/articles/27/121/2020/npg-27-121-2020-f03.png"/>

        </fig>

      <p id="d1e402">We treat the predictors <inline-formula><mml:math id="M13" display="inline"><mml:mrow><mml:msub><mml:mi>p</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> like the ensemble members before and apply an empirical mixture modelling. For the year 2010 the results are shown in Fig. <xref ref-type="fig" rid="Ch1.F3"/>. Due to the limited number of predictors compared to the ensemble members, and in the shown case also due to their alignment, the resulting statistical prediction pdf (<inline-formula><mml:math id="M14" display="inline"><mml:mi mathvariant="script">P</mml:mi></mml:math></inline-formula>) is much sharper than the dynamical model prediction.</p>
      <p id="d1e425">To create a combined prediction the two pdfs (<inline-formula><mml:math id="M15" display="inline"><mml:mi mathvariant="script">E</mml:mi></mml:math></inline-formula> and <inline-formula><mml:math id="M16" display="inline"><mml:mi mathvariant="script">P</mml:mi></mml:math></inline-formula>) are after normalisation multiplied at each of the discretised NAO values:
            <disp-formula id="Ch1.E2" content-type="numbered"><label>2</label><mml:math id="M17" display="block"><mml:mrow><mml:mi mathvariant="script">M</mml:mi><mml:mo>(</mml:mo><mml:mi>v</mml:mi><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:mi mathvariant="script">E</mml:mi><mml:mo>(</mml:mo><mml:mi>v</mml:mi><mml:mo>)</mml:mo><mml:mo>⋅</mml:mo><mml:mi mathvariant="script">P</mml:mi><mml:mo>(</mml:mo><mml:mi>v</mml:mi><mml:mo>)</mml:mo><mml:mo>.</mml:mo></mml:mrow></mml:math></disp-formula>
          After another normalisation the final combined prediction <inline-formula><mml:math id="M18" display="inline"><mml:mi mathvariant="script">M</mml:mi></mml:math></inline-formula> creates the statistical–dynamical prediction for the seasonal NAO prediction in the specific year. The pdfs of the observations (<inline-formula><mml:math id="M19" display="inline"><mml:mi mathvariant="script">O</mml:mi></mml:math></inline-formula>) are determined by the same bootstrapping mechanism as the one applied for the hindcasts.
The result for the year 2010 is shown in Fig. <xref ref-type="fig" rid="Ch1.F4"/>. The pdf of the combined prediction is close to the one of the statistical predictions, but shows differences where there is additional information from the dynamical model prediction. Therefore, the combined prediction shows a clearer signal than the dynamical model prediction, which does not give any indication of a specific NAO value at all.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F4" specific-use="star"><?xmltex \currentcnt{4}?><label>Figure 4</label><caption><p id="d1e498">Sequence of the post-processing procedure for the WNAO in 2010. Combining dynamical (black) and statistical (pink) predictions to a combined prediction (blue) and comparing it to the observations (red).</p></caption>
          <?xmltex \igopts{width=341.433071pt}?><graphic xlink:href="https://npg.copernicus.org/articles/27/121/2020/npg-27-121-2020-f04.png"/>

        </fig>

</sec>
<sec id="Ch1.S3.SS3">
  <label>3.3</label><title>NAO evaluation</title>
      <p id="d1e515">To evaluate the performance of the three different predictions (<inline-formula><mml:math id="M20" display="inline"><mml:mi mathvariant="script">E</mml:mi></mml:math></inline-formula>, <inline-formula><mml:math id="M21" display="inline"><mml:mi mathvariant="script">P</mml:mi></mml:math></inline-formula> and <inline-formula><mml:math id="M22" display="inline"><mml:mi mathvariant="script">M</mml:mi></mml:math></inline-formula>) and compare the predictions with the observation, we use two different scores based on the same formulation. The first is based on the Earth mover's distance <xref ref-type="bibr" rid="bib1.bibx15" id="paren.22"/>. The one-dimensional EMD <xref ref-type="bibr" rid="bib1.bibx7" id="paren.23"/> can be derived by
            <disp-formula id="Ch1.E3" content-type="numbered"><label>3</label><mml:math id="M23" display="block"><mml:mrow><mml:msub><mml:mi>D</mml:mi><mml:mi mathvariant="normal">EMD</mml:mi></mml:msub><mml:mo>(</mml:mo><mml:mi>f</mml:mi><mml:mo>,</mml:mo><mml:mi>g</mml:mi><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mn mathvariant="normal">1</mml:mn><mml:mrow><mml:msub><mml:mi>n</mml:mi><mml:mi>b</mml:mi></mml:msub></mml:mrow></mml:mfrac></mml:mstyle><mml:munderover><mml:mo movablelimits="false">∑</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mi>n</mml:mi><mml:mi>b</mml:mi></mml:msub></mml:mrow></mml:munderover><mml:mfenced open="|" close="|"><mml:mrow><mml:mi>F</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi>v</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>)</mml:mo><mml:mo>-</mml:mo><mml:mi>G</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi>v</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:mfenced><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>
          where <inline-formula><mml:math id="M24" display="inline"><mml:mi>f</mml:mi></mml:math></inline-formula> and <inline-formula><mml:math id="M25" display="inline"><mml:mi>g</mml:mi></mml:math></inline-formula> are two pdfs and <inline-formula><mml:math id="M26" display="inline"><mml:mi>F</mml:mi></mml:math></inline-formula> and <inline-formula><mml:math id="M27" display="inline"><mml:mi>G</mml:mi></mml:math></inline-formula> the associated cumulative distribution functions (cdfs). <inline-formula><mml:math id="M28" display="inline"><mml:mrow><mml:msub><mml:mi>n</mml:mi><mml:mi>b</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> describe in this case the number of discretised values <inline-formula><mml:math id="M29" display="inline"><mml:mrow><mml:msub><mml:mi>v</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> of the cdfs.</p>
      <p id="d1e674">The second is the IQD, which is defined in its discrete formulation as <xref ref-type="bibr" rid="bib1.bibx18" id="paren.24"/>
            <disp-formula id="Ch1.E4" content-type="numbered"><label>4</label><mml:math id="M30" display="block"><mml:mrow><mml:msub><mml:mi>D</mml:mi><mml:mi mathvariant="normal">IQD</mml:mi></mml:msub><mml:mo>(</mml:mo><mml:mi>f</mml:mi><mml:mo>,</mml:mo><mml:mi>g</mml:mi><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mn mathvariant="normal">1</mml:mn><mml:mrow><mml:msub><mml:mi>n</mml:mi><mml:mi>b</mml:mi></mml:msub></mml:mrow></mml:mfrac></mml:mstyle><mml:munderover><mml:mo movablelimits="false">∑</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mi>n</mml:mi><mml:mi>b</mml:mi></mml:msub></mml:mrow></mml:munderover><mml:mo>(</mml:mo><mml:mi>F</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi>v</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>)</mml:mo><mml:mo>-</mml:mo><mml:mi>G</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi>v</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>)</mml:mo><mml:msup><mml:mo>)</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo>.</mml:mo></mml:mrow></mml:math></disp-formula>
          It must be mentioned that the IQD is similar to the continuous ranked probability score (CRPS) but is defined for non-deterministic observations. As a consequence, while CRPS needs to have a point observation, the IQD can take into account the full uncertainty distribution of an observation.</p>
      <p id="d1e763">We define the scores for both metrics by comparing the pdfs of the model prediction (<inline-formula><mml:math id="M31" display="inline"><mml:mi mathvariant="script">M</mml:mi></mml:math></inline-formula>), the observations (<inline-formula><mml:math id="M32" display="inline"><mml:mi mathvariant="script">O</mml:mi></mml:math></inline-formula>) and the climatology (<inline-formula><mml:math id="M33" display="inline"><mml:mi mathvariant="script">C</mml:mi></mml:math></inline-formula>). It is calculated for any prediction <inline-formula><mml:math id="M34" display="inline"><mml:mi mathvariant="script">A</mml:mi></mml:math></inline-formula> by
            <disp-formula id="Ch1.E5" content-type="numbered"><label>5</label><mml:math id="M35" display="block"><mml:mrow><mml:mi>q</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="script">A</mml:mi><mml:mo>,</mml:mo><mml:mi mathvariant="script">O</mml:mi><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>-</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:mi>D</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="script">A</mml:mi><mml:mo>,</mml:mo><mml:mi mathvariant="script">O</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mi>D</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="script">C</mml:mi><mml:mo>,</mml:mo><mml:mi mathvariant="script">O</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mfrac></mml:mstyle><mml:mo>.</mml:mo></mml:mrow></mml:math></disp-formula>
          When <inline-formula><mml:math id="M36" display="inline"><mml:mi>D</mml:mi></mml:math></inline-formula> is <inline-formula><mml:math id="M37" display="inline"><mml:mrow><mml:msub><mml:mi>D</mml:mi><mml:mi mathvariant="normal">EMD</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> we call the score the 1D-continuous-EMD score, and when we apply <inline-formula><mml:math id="M38" display="inline"><mml:mrow><mml:msub><mml:mi>D</mml:mi><mml:mi mathvariant="normal">IQD</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> it is the 1D-continuous-IQD score.</p>
      <p id="d1e878">In the case of a perfect prediction the score becomes 1, a model prediction equal to a climatology 0 and negative for a worse prediction than the climatology. Since the NAO index is normalised for mean and standard deviation, we use as climatology a standard normal distribution <inline-formula><mml:math id="M39" display="inline"><mml:mrow><mml:mi mathvariant="script">N</mml:mi><mml:mo>(</mml:mo><mml:mn mathvariant="normal">0</mml:mn><mml:mo>,</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>.
It is important to note here that <xref ref-type="bibr" rid="bib1.bibx18" id="text.25"/> compared the two metrics (EMD as <italic>area validation metric</italic>). While the EMD is a metric measuring the distance between the pdfs, it is in contrast to the IQD not a proper divergence measure. As a consequence, the EMD prefers, unlike the IQD, underdispersed model simulations. In the following we will demonstrate the effect that the choice of the two different metrics has on the evaluation.</p>
      <p id="d1e906">To estimate uncertainties, we use 500 randomly selected uniformly distributed weightings of the ensemble members between 1 and 0 and create with those a pdf for the scores.</p>
</sec>
<sec id="Ch1.S3.SS4">
  <label>3.4</label><title>Variable field evaluation</title>
      <?pagebreak page124?><p id="d1e917">To estimate the post-processed variable field, we calculate a weighted mean of the meteorological variable fields, where the field of each individual member is weighted by a coefficient <inline-formula><mml:math id="M40" display="inline"><mml:mrow><mml:msub><mml:mi>c</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>. The weighting coefficients <inline-formula><mml:math id="M41" display="inline"><mml:mrow><mml:msub><mml:mi>c</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> are estimated by weighting the predictions <inline-formula><mml:math id="M42" display="inline"><mml:mi mathvariant="script">A</mml:mi></mml:math></inline-formula> (each of <inline-formula><mml:math id="M43" display="inline"><mml:mi mathvariant="script">E</mml:mi></mml:math></inline-formula>, <inline-formula><mml:math id="M44" display="inline"><mml:mi mathvariant="script">M</mml:mi></mml:math></inline-formula> and <inline-formula><mml:math id="M45" display="inline"><mml:mi mathvariant="script">P</mml:mi></mml:math></inline-formula>) with each of the pdfs of the ensemble members (<inline-formula><mml:math id="M46" display="inline"><mml:mrow><mml:msub><mml:mi>e</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>):
            <disp-formula id="Ch1.E6" content-type="numbered"><label>6</label><mml:math id="M47" display="block"><mml:mrow><mml:msub><mml:mi>c</mml:mi><mml:mrow><mml:mi>A</mml:mi><mml:mo>,</mml:mo><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:munder><mml:mo movablelimits="false">∑</mml:mo><mml:mi>v</mml:mi></mml:munder><mml:msub><mml:mi>e</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>(</mml:mo><mml:mi>v</mml:mi><mml:mo>)</mml:mo><mml:mo>⋅</mml:mo><mml:mi mathvariant="script">A</mml:mi><mml:mo>(</mml:mo><mml:mi>v</mml:mi><mml:mo>)</mml:mo><mml:mo>.</mml:mo></mml:mrow></mml:math></disp-formula>
          Weighting each ensemble member with its associated coefficient <inline-formula><mml:math id="M48" display="inline"><mml:mrow><mml:msub><mml:mi>c</mml:mi><mml:mrow><mml:mi>A</mml:mi><mml:mo>,</mml:mo><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula> and calculating the weighted mean of the atmospheric fields of the individual ensembles then generate the model prediction for the specified field and prediction.</p>
      <p id="d1e1045">For evaluation of the meteorological variable fields, we apply three different strategies. The first is the anomaly correlation coefficient (ACC), a common measure of skill in seasonal predictions. The second and third approaches are to use the 1D-continuous-EMD and 1D-continuous-IQD scores at every grid point. As a climatology all observational values for the investigated time frame are chosen. The observation in each year is a single value with 100 % as a weight. In the case of the weights for the ensemble member, each value of the variable at the grid point gets weighted with the relative weight <inline-formula><mml:math id="M49" display="inline"><mml:mrow><mml:msub><mml:mi>c</mml:mi><mml:mrow><mml:mi>A</mml:mi><mml:mo>,</mml:mo><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula> given by the three different predictions. With this approach it is possible to calculate the 1D-continuous-EMD and 1D-continuous-IQD scores for each of the three different predictions. In Sect. <xref ref-type="sec" rid="Ch1.S4.SS2.SSS2"/> the relative positioning between two predictions is shown. Significances are here determined by <xref ref-type="bibr" rid="bib1.bibx4" id="text.26"/>, which determines the skill significances by comparisons to random walks.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F5" specific-use="star"><?xmltex \currentcnt{5}?><label>Figure 5</label><caption><p id="d1e1071">Yearly comparison of the WNAO scores for dynamical (black), statistical (pink) and combined (blue) predictions. Each vertical bar represents the 5 % to 95 % bootstrapped 1D-continuous-EMD score (above) and 1D-continuous-IQD score (below). The filled parts of these bars are the 25th to 75th quartiles and the small vertical lines the associated median. The long vertical lines are the averaged yearly scores for the different predictions.</p></caption>
          <?xmltex \igopts{width=312.980315pt}?><graphic xlink:href="https://npg.copernicus.org/articles/27/121/2020/npg-27-121-2020-f05.png"/>

        </fig>

</sec>
</sec>
<sec id="Ch1.S4">
  <label>4</label><title>Results</title>
<sec id="Ch1.S4.SS1">
  <label>4.1</label><title>Evaluating the seasonal NAO prediction</title>
      <p id="d1e1096">In a next step we evaluate the yearly performance of the WNAO prediction of the three different predictions (<inline-formula><mml:math id="M50" display="inline"><mml:mi mathvariant="script">E</mml:mi></mml:math></inline-formula>, <inline-formula><mml:math id="M51" display="inline"><mml:mi mathvariant="script">P</mml:mi></mml:math></inline-formula> and <inline-formula><mml:math id="M52" display="inline"><mml:mi mathvariant="script">M</mml:mi></mml:math></inline-formula>) with the 1D-continuous-EMD and 1D-continuous-IQD scores. Figure <xref ref-type="fig" rid="Ch1.F5"/> shows that the results of the combined (<inline-formula><mml:math id="M53" display="inline"><mml:mi mathvariant="script">M</mml:mi></mml:math></inline-formula>) and statistical (<inline-formula><mml:math id="M54" display="inline"><mml:mi mathvariant="script">P</mml:mi></mml:math></inline-formula>) predictions are clearly better performing than the dynamical model results (<inline-formula><mml:math id="M55" display="inline"><mml:mi mathvariant="script">E</mml:mi></mml:math></inline-formula>). In most years, the combined and statistical predictions demonstrate skill for the 1D-continuous-EMD score compared to a climatological prediction over the whole uncertainty range. The dynamical model prediction has less variability over the years in skill than the other two predictions and in only a few years is able to reach the average skill of the combined prediction. The median and interquartile range of the summed-up prediction skill for all evaluated years for the combined prediction (<inline-formula><mml:math id="M56" display="inline"><mml:mrow><mml:mn mathvariant="normal">0.39</mml:mn><mml:mspace linebreak="nobreak" width="0.25em"/><mml:mo>[</mml:mo><mml:mn mathvariant="normal">0.21</mml:mn><mml:mo>;</mml:mo><mml:mn mathvariant="normal">0.60</mml:mn><mml:mo>]</mml:mo></mml:mrow></mml:math></inline-formula>) is higher compared to the dynamical (<inline-formula><mml:math id="M57" display="inline"><mml:mrow><mml:mn mathvariant="normal">0.12</mml:mn><mml:mspace linebreak="nobreak" width="0.25em"/><mml:mo>[</mml:mo><mml:mn mathvariant="normal">0.01</mml:mn><mml:mo>;</mml:mo><mml:mn mathvariant="normal">0.22</mml:mn><mml:mo>]</mml:mo></mml:mrow></mml:math></inline-formula>) and statistical (<inline-formula><mml:math id="M58" display="inline"><mml:mrow><mml:mn mathvariant="normal">0.37</mml:mn><mml:mspace width="0.25em" linebreak="nobreak"/><mml:mo>[</mml:mo><mml:mn mathvariant="normal">0.17</mml:mn><mml:mo>;</mml:mo><mml:mn mathvariant="normal">0.56</mml:mn><mml:mo>]</mml:mo></mml:mrow></mml:math></inline-formula>) predictions. There is only one year (2003) with a strong discrepancy of the combined and statistical predictions and a clearly negative score. For 1D-continuous-IQD the results are less clear. In this case the median of the combined prediction is much closer to the dynamical prediction than the statistical prediction. The uncertainty range for the combined and statistical predictions also increases relative to the dynamical prediction, which can be explained by their sharpness.
To better evaluate the performance of each prediction with respect to the other predictions, we determine the relative ranking of the median of each prediction in each year for both scores. The rankings are counted for the whole hindcast period and the results displayed in Table <xref ref-type="table" rid="Ch1.T1"/>. For the 1D-continuous-EMD score the dynamical prediction has in only a few years a better prediction skill than the other two predictions. In the majority of the years its prediction skill is lower<?pagebreak page125?> than both other predictions. Looking at the best prediction for each year, the statistical and combined predictions are on equal terms. Nevertheless, the combined prediction is much more unlikely to be the worst of the three predictions in a year, while the statistical prediction takes much more often the third rank. These results show that the combined prediction is closer to the statistical rather than dynamical prediction. In case the combined prediction is not the best one, it is in almost all cases better than one of the two. As such it offers a smoothing of the prediction skill, preventing many worse predictions. In the case of the 1D-continuous-IQD score, the result differs clearly. Here the dynamical prediction is much more competitive. It shares almost equally with the statistical prediction first place, while the statistical prediction hardly changes its statistics of positions. As a consequence the combined prediction is much more often in last place. Still, it is the prediction with the most middle places of the three predictions, stressing the argument that the combined prediction is a mixture of the other two.</p>

<?xmltex \floatpos{t}?><table-wrap id="Ch1.T1" specific-use="star"><?xmltex \currentcnt{1}?><label>Table 1</label><caption><p id="d1e1206">Count of years of relative positioning of the three different predictions using the median of the 1D-continuous-EMD score.</p></caption><oasis:table frame="topbot"><oasis:tgroup cols="7">
     <oasis:colspec colnum="1" colname="col1" align="left"/>
     <oasis:colspec colnum="2" colname="col2" align="right"/>
     <oasis:colspec colnum="3" colname="col3" align="right"/>
     <oasis:colspec colnum="4" colname="col4" align="right" colsep="1"/>
     <oasis:colspec colnum="5" colname="col5" align="right"/>
     <oasis:colspec colnum="6" colname="col6" align="right"/>
     <oasis:colspec colnum="7" colname="col7" align="right"/>
     <oasis:thead>
       <oasis:row>
         <oasis:entry colname="col1"/>
         <oasis:entry rowsep="1" namest="col2" nameend="col4" align="center" colsep="1">EMD </oasis:entry>
         <oasis:entry rowsep="1" namest="col5" nameend="col7" align="center">IQD </oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">rank</oasis:entry>
         <oasis:entry colname="col2">dynamical</oasis:entry>
         <oasis:entry colname="col3">statistical</oasis:entry>
         <oasis:entry colname="col4">combination</oasis:entry>
         <oasis:entry colname="col5">dynamical</oasis:entry>
         <oasis:entry colname="col6">statistical</oasis:entry>
         <oasis:entry colname="col7">combination</oasis:entry>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row>
         <oasis:entry colname="col1">1</oasis:entry>
         <oasis:entry colname="col2">5</oasis:entry>
         <oasis:entry colname="col3">17</oasis:entry>
         <oasis:entry colname="col4">14</oasis:entry>
         <oasis:entry colname="col5">13</oasis:entry>
         <oasis:entry colname="col6">16</oasis:entry>
         <oasis:entry colname="col7">7</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">2</oasis:entry>
         <oasis:entry colname="col2">5</oasis:entry>
         <oasis:entry colname="col3">10</oasis:entry>
         <oasis:entry colname="col4">21</oasis:entry>
         <oasis:entry colname="col5">7</oasis:entry>
         <oasis:entry colname="col6">12</oasis:entry>
         <oasis:entry colname="col7">17</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">3</oasis:entry>
         <oasis:entry colname="col2">26</oasis:entry>
         <oasis:entry colname="col3">9</oasis:entry>
         <oasis:entry colname="col4">1</oasis:entry>
         <oasis:entry colname="col5">16</oasis:entry>
         <oasis:entry colname="col6">8</oasis:entry>
         <oasis:entry colname="col7">12</oasis:entry>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup></oasis:table></table-wrap>

      <?xmltex \floatpos{t}?><fig id="Ch1.F6" specific-use="star"><?xmltex \currentcnt{6}?><label>Figure 6</label><caption><p id="d1e1348">ACC results for the WNAO for three different atmospheric variables: surface temperature <bold>(a, b, c)</bold>, total precipitation <bold>(d, e, f)</bold> and geopotential height <bold>(g, h, i)</bold>. Shown are the combined prediction <bold>(a, d, g)</bold>, the difference between the combined and dynamical predictions <bold>(b, e, h)</bold> and the difference between the combined and statistical predictions <bold>(c, f, i)</bold>. Black dots indicate significances estimated by a 500-sample bootstrap.</p></caption>
          <?xmltex \igopts{width=497.923228pt}?><graphic xlink:href="https://npg.copernicus.org/articles/27/121/2020/npg-27-121-2020-f06.png"/>

        </fig>

</sec>
<sec id="Ch1.S4.SS2">
  <label>4.2</label><title>Analysis of atmospheric variable fields</title>
<sec id="Ch1.S4.SS2.SSS1">
  <label>4.2.1</label><title>Climatological analysis</title>
      <p id="d1e1391">In the following we investigate three different atmospheric variable fields: surface temperature, total precipitation and 500 hPa geopotential height. In Fig. <xref ref-type="fig" rid="Ch1.F6"/> the results are shown for the winter (DJF) season with the ACC.
For the winter surface temperature, the main areas of significant hindcast skill of the combined prediction can be found over large parts of the North Atlantic and in a band reaching from northern France to eastern Europe, sparing northern Scandinavia and the Mediterranean. These results are comparable to those shown by <xref ref-type="bibr" rid="bib1.bibx6" id="text.27"/>. Comparing it to the dynamical prediction shows that the main significant increase in skill can be found over western Europe, with a general non-significant increase over the whole continent. Some significant increase in prediction skill can also be found in the Labrador Sea, while a significant decrease is located over Greenland. The comparison to the statistical prediction shows only small differences. The areas shown as significant have to be assumed to be random and an artefact of the bootstrapping approach.</p>
      <p id="d1e1399">The total precipitation has significant positive hindcast skill north of the British Isles, east of the Baltic Sea, in the Mediterranean and between the Canaries and the Azores. Compared to the dynamical prediction the area east of the Baltic Sea and the Mediterranean has significantly increased skill, while again compared to the statistical prediction not much change is detectable. Finally, for the geopotential height, the hindcast skill for the combined prediction is found in areas over the Iberian Peninsula, between the Canaries and the Azores and between the British Isles and Greenland. Compared to the dynamical prediction, some increase in hindcast skill can be found over southern Scandinavia and the east of Greenland. In the comparison to the statistical prediction the combined prediction shows significantly lower hindcast skill in areas over Greenland and the British Isles. This can be explained by the conditioning of the statistical prediction on the NAO directly, while the dynamic component of the statistical dynamical prediction decreases the skill in the main influence areas of the NAO.</p>
      <p id="d1e1402">The analysis shows that there exist changes between the dynamical and combined predictions. Generally, the hindcast skill of the combined prediction is very close to the one of the statistical prediction.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F7" specific-use="star"><?xmltex \currentcnt{7}?><label>Figure 7</label><caption><p id="d1e1408">Relative positioning of the predictions of the variable fields on the basis of the 1D-continuous-EMD score. Shown is the number of years in which the first named prediction has a better score than the second named prediction. Compared are the combined and model predictions <bold>(a, d, g)</bold>, statistical and dynamical predictions <bold>(b, e, h)</bold> and combined and statistical predictions <bold>(c, f, i)</bold>. Significances are determined by a comparison towards a random walk at a confidence level of 0.05. Variables are positioned as in Fig. <xref ref-type="fig" rid="Ch1.F6"/>.</p></caption>
            <?xmltex \igopts{width=497.923228pt}?><graphic xlink:href="https://npg.copernicus.org/articles/27/121/2020/npg-27-121-2020-f07.png"/>

          </fig>

</sec>
<sec id="Ch1.S4.SS2.SSS2">
  <label>4.2.2</label><title>Analysing single years</title>
      <?pagebreak page127?><p id="d1e1436">In a next step, the same atmospheric fields are compared with the 1D-continuous-EMD score (Fig. <xref ref-type="fig" rid="Ch1.F7"/>). To prevent influences of biases and trends, the data are grid-point-wise normalised and de-trended. Again the analysis shows the relative positioning of two predictions. For the surface temperature in winter the difference between the dynamical and combined prediction is only significant in small patches distributed over the North Atlantic. Generally, no clear patterns can be identified. Especially the large significant areas determined by the ACC before do not show any significance with this score. The significant area in the ACC over western Europe has some increased values in favour of the combined predictions, but is not significant. In the comparison between the statistical and dynamical predictions the increases and decreases are consistent with what has been seen for the dynamical prediction compared to the combined prediction. This consistency shows that the statistical model plays a dominating role in the combination. Better hindcast skill for the combined prediction compared to the statistical prediction can be identified in the west of the Mediterranean.</p>
      <p id="d1e1441">For the total precipitation the only significant change is a stretch north of Scandinavia. Also for the other comparisons for this variable the changes are small and do not show a consistent pattern. This is different for the geopotential height, where large areas in  the north-eastern Atlantic and north of Scandinavia are significantly better represented in the combined prediction rather than the dynamical prediction. Both areas are not identified in the equivalent comparison with the anomaly correlation. The comparison of the statistical prediction compared to the dynamical prediction shows very similar patterns. The last comparison shows that the combination has areas between the Canaries and the Azores, where it is significantly higher, while in large areas of western Europe it has consistently better skill but does not show significantly better skill.</p>
      <p id="d1e1444">This analysis shows that the three predictions do not have in all cases a clear relative ranking towards each other. Generally the results are very patchy, and apart from the north of Scandinavia, no consistency can be seen.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F8" specific-use="star"><?xmltex \currentcnt{8}?><label>Figure 8</label><caption><p id="d1e1450">Relative positioning of the predictions of the variable fields on the basis of the 1D-continuous-IQD score. As Fig. <xref ref-type="fig" rid="Ch1.F7"/> but calculated with the 1D-continuous-IQD.</p></caption>
            <?xmltex \igopts{width=497.923228pt}?><graphic xlink:href="https://npg.copernicus.org/articles/27/121/2020/npg-27-121-2020-f08.png"/>

          </fig>

      <p id="d1e1461">In the case of the analysis of the 1D-continuous-IQD score (Fig. <xref ref-type="fig" rid="Ch1.F8"/>), the comparison between the statistical and combined models shows, in terms of significant areas, comparable results to the one seen in Fig. <xref ref-type="fig" rid="Ch1.F7"/>. When the two predictions are compared to the dynamic prediction, the latter performs<?pagebreak page128?> much better with this score than with the 1D-continuous-EMD score. While the general pattern of the areas stays the same, the dynamic prediction is in most areas the best prediction. Comparing the combined and statistical models shows remarkably similar results to the 1D-continuous-EMD score. All these results are consistent with the results we have seen in Sect. <xref ref-type="sec" rid="Ch1.S4.SS1"/> for the single time series.</p>
</sec>
</sec>
</sec>
<sec id="Ch1.S5" sec-type="conclusions">
  <label>5</label><title>Discussion and conclusion</title>
      <p id="d1e1481">This paper shows a post-processing procedure, generalising the newly established subsampling procedure by <xref ref-type="bibr" rid="bib1.bibx6" id="text.28"/>. By not only selecting single ensemble members, but also utilising their uncertainty ranges, a much better understanding of the reason for its success is possible. As seen in Sect. <xref ref-type="sec" rid="Ch1.S3.SS2"/> the better prediction skill for the NAO by the combination of the statistical and dynamical model compared to the unprocessed dynamical prediction results from the sharper prediction of the statistical prediction. As by construction the different statistical predictions are highly connected towards the target value, in this case NAO, the predictor-driven predictions result in higher skill. Furthermore, advantages of using this post-processing approach compared to the pure subsampling are the availability of non-parametric uncertainties for the predictions and the possibility of weighting the different ensemble members for the analysis of variable fields with unequal weights. As such especially outliers can therefore be much better handled, without giving them too high of a weight within the analysis.</p>
      <?pagebreak page129?><p id="d1e1489">Compared to the statistical prediction, the combined prediction achieves similar results for the NAO prediction. In a three-way comparison together with the dynamical prediction we have shown that it generally does not show more skill than the statistical prediction, but it observes less negative outliers in skill. Nevertheless, in the case of the atmospheric variable predictions, the prediction based on predictors is not entirely a statistical prediction. The construction of weighting the ensemble members leads to a statistical–dynamical prediction as well, where the weight of the dynamical model is less pronounced. As such, the skill between the two dynamical–statistical predictions is more similar in this case than the NAO prediction itself.
We have seen that the two categories of scores show the hindcast skill of the different forecasts from a different perspective. The 1D-continuous-EMD and 1D-continuous-IQD scores allow us to effectively evaluate the skill of two probabilistic results, like observations and predictions. The scores have similar characteristics like the RMSE in cases of undetected trends, different variability of different forecasts or a bias. In the case of this study it is noted that the combined prediction is sharper than the dynamical prediction for each year's prediction, but also varies more from year to year. Also compared to the correlation, the two presented scores can decompose the skill in a consistent way for every single year.</p>
      <?pagebreak page130?><p id="d1e1492">As each year is compared to the climatology, a value close to the climatology can have a huge influence by creating substantive negative scores. To prevent this, the application of other references, like uniform distributions over the whole measurement range, can be an appropriate alternative. Comparing the results of the 1D-continuous-EMD and 1D-continuous-IQD scores shows that the latter infers a much harder penalty for mispredictions. While the EMD metric uses a linear distance measure, the IQD divergence increases the distance by the square in Eq. (<xref ref-type="disp-formula" rid="Ch1.E4"/>). Discussion and comparison of the properties of two measures have been done by <xref ref-type="bibr" rid="bib1.bibx18" id="text.29"/>. In the practical implementation done in this paper we have seen that the IQD tends to prefer a non-informative prediction over a wrong sharp prediction, while the EMD is more tolerant of wrong prediction in order to achieve a better score.</p>
      <p id="d1e1500">Evaluating the skill on a yearly basis and taking a look at the relative positioning the approach allow for a paradigm change as also described by <xref ref-type="bibr" rid="bib1.bibx4" id="text.30"/>. By counting the years in which one prediction is better than another, a single outlier cannot drive the whole verification result as it can do for correlations or RMSE. It also answers a typical question in forecast verification in a much more appropriate way: how sure can we be that a single prediction is better than another? The evaluation procedure presented here is able to quantify this answer for non-parametric predictions.</p>
      <p id="d1e1507">The ACC is well used in the literature, and its main disadvantages are parametric assumptions in the interpretation of its results. We have seen that there are considerable differences when all years are evaluated at the same time, as is done in a correlation-based score or the evaluation based on evaluating single years. Correlations can be misleading and show skill where there is not necessarily a good argument for it as it is prone to outliers. These discussions are well known when correlation-like measures are compared with distance-like measures, like the RMSE. Further progress in the creation of appropriate skill evaluation is therefore necessary.
It is noted that while we show in this analysis only the results for the winter season, the results for the summer season are comparable.</p>
      <p id="d1e1510">The methodology and verification techniques shown in this analysis are widely applicable within predictions of many different phenomena. This is especially valid in the case of non-parametric datasets like in the analysis of extremes. The statistical–dynamical approach as illustrated here delivers consistent improved results compared to one of its components. Seen as a post-processing step, it forms a useful step to condition predictions on a physical basis in order to reduce noise and intensify the signal. Using non-parametric approaches in the analysis offers a more appropriate path to verify predictions in general.</p>
</sec>

      
      </body>
    <back><notes notes-type="dataavailability"><title>Data availability</title>

      <p id="d1e1517">All data are stored in the DKRZ archive and can be made accessible upon request (<uri>https://www.dkrz.de/up</uri>, last access: 26 February 2020).</p>
  </notes><notes notes-type="competinginterests"><title>Competing interests</title>

      <p id="d1e1526">The author declares that there is no conflict of interest.</p>
  </notes><notes notes-type="sistatement"><title>Special issue statement</title>

      <p id="d1e1532">This article is part of the special issue “Advances in post-processing and blending of deterministic and ensemble forecasts”. It is not associated with a conference.</p>
  </notes><ack><title>Acknowledgements</title><p id="d1e1538">The author would like to thank Johanna Baehr and Mikhail Dobrynin for the fruitful discussions. The author would also like to thank three anonymous reviewers and editor Sebastian Lerch for very helpful comments on this paper. Model simulations were performed using the high-performance computer at the German Climate Computing Center (DKRZ).</p></ack><notes notes-type="financialsupport"><title>Financial support</title>

      <p id="d1e1543">This research has been supported the University of Hamburg's Cluster of Excellence Integrated Climate System Analysis and Prediction (CliSAP). It was also supported by A4 (Aigéin, Aeráid, agus athrú Atlantaigh), funded by the Marine Institute and the European Regional Development fund (grant: PBA/CC/18/01).</p>
  </notes><notes notes-type="reviewstatement"><title>Review statement</title>

      <p id="d1e1549">This paper was edited by Sebastian Lerch and reviewed by three anonymous referees.</p>
  </notes><ref-list>
    <title>References</title>

      <ref id="bib1.bibx1"><?xmltex \def\ref@label{{Baehr et~al.(2015)Baehr, Fr{\"{o}}hlich, Botzet, Domeisen, Kornblueh,
Notz, Piontek, Pohlmann, Tietsche, and M{\"{u}}ller}}?><label>Baehr et al.(2015)Baehr, Fröhlich, Botzet, Domeisen, Kornblueh,
Notz, Piontek, Pohlmann, Tietsche, and Müller</label><?label BaehrFrohlich2015?><mixed-citation>
Baehr, J., Fröhlich, K., Botzet, M., Domeisen, D. I. V., Kornblueh, L.,
Notz, D., Piontek, R., Pohlmann, H., Tietsche, S., and Müller, W. A.: The
prediction of surface temperature in the new seasonal prediction system based
on the MPI-ESM coupled climate model, Clim. Dynam., 44, 2723–2735, 2015.</mixed-citation></ref>
      <ref id="bib1.bibx2"><?xmltex \def\ref@label{{Butler et~al.(2016)Butler, Arribas, Athanassiadou, Baehr, Calvo,
Charlton-Perez, D{\'{e}}qu{\'{e}}, Domeisen, Fr{\"{o}}hlich, Hendon, Imada, Ishii,
Iza, Karpechko, Kumar, MacLachlan, Merryfield, M{\"{u}}ller, O'Neill, Scaife,
Scinocca, Sigmond, Stockdale, and Yasuda}}?><label>Butler et al.(2016)Butler, Arribas, Athanassiadou, Baehr, Calvo,
Charlton-Perez, Déqué, Domeisen, Fröhlich, Hendon, Imada, Ishii,
Iza, Karpechko, Kumar, MacLachlan, Merryfield, Müller, O'Neill, Scaife,
Scinocca, Sigmond, Stockdale, and Yasuda</label><?label ButlerArribas2016?><mixed-citation>
Butler, A. H., Arribas, A., Athanassiadou, M., Baehr, J., Calvo, N.,
Charlton-Perez, A., Déqué, M., Domeisen, D. I. V., Fröhlich, K.,
Hendon, H., Imada, Y., Ishii, M., Iza, M., Karpechko, A. Y., Kumar, A.,
MacLachlan, C., Merryfield, W. J., Müller, W. A., O'Neill, A., Scaife,
A. A., Scinocca, J., Sigmond, M., Stockdale, T. N., and Yasuda, T.: The
Climate-system Historical Forecast Project: do stratosphere-resolving models
make better seasonal climate predictions in boreal winter?, Q. J.
Roy. Meteor. Soc., 142, 1413–1427, 2016.</mixed-citation></ref>
      <ref id="bib1.bibx3"><?xmltex \def\ref@label{{Dee et~al.(2011)Dee, Uppala, Simmons, Berrisford, Poli, Kobayashi,
Andrae, Balmaseda, Balsamo, Bauer, Bechtold, Beljaars, van~de Berg, Bidlot,
Bormann, Delsol, Dragani, Fuentes, Geer, Haimberger, Healy, Hersbach,
H{\'{o}}lm, Isaksen, K{\aa}llberg, K{\"{o}}hler, Matricardi, McNally,
Monge‐Sanz, Morcrette, Park, Peubey, de~Rosnay, Tavolato, Th{\'{e}}paut, and
Vitart}}?><label>Dee et al.(2011)Dee, Uppala, Simmons, Berrisford, Poli, Kobayashi,
Andrae, Balmaseda, Balsamo, Bauer, Bechtold, Beljaars, van de Berg, Bidlot,
Bormann, Delsol, Dragani, Fuentes, Geer, Haimberger, Healy, Hersbach,
Hólm, Isaksen, Kållberg, Köhler, Matricardi, McNally,
Monge‐Sanz, Morcrette, Park, Peubey, de Rosnay, Tavolato, Thépaut, and
Vitart</label><?label DeeUppala2011?><mixed-citation>
Dee, D. P., Uppala, S. M., Simmons, A. J., Berrisford, P., Poli, P., Kobayashi,
S., Andrae, U., Balmaseda, M. A., Balsamo, G., Bauer, P., Bechtold, P.,
Beljaars, A. C. M., van de Berg, L., Bidlot, J., Bormann, N., Delsol, C.,
Dragani, R., Fuentes, M., Geer, A. J., Haimberger, L., Healy, S. B.,
Hersbach, H., Hólm, E. V., Isaksen, L., Kållberg, P., Köhler, M.,
Matricardi, M., McNally, A. P., Monge‐Sanz, B. M., Morcrette, J., Park, B.,
Peubey, C., de Rosnay, P., Tavolato, C., Thépaut, J., and Vitart, F.: The
ERA‐Interim reanalysis: configuration and performance of the data
assimilation system, Q. J. Roy. Meteor. Soc.,
137, 553–597, 2011.</mixed-citation></ref>
      <ref id="bib1.bibx4"><label>DelSole and Tippett(2016)</label><?label DelTip16?><mixed-citation>
DelSole, T. and Tippett, M. K.: Forecast Comparison Based on Random Walks,
Mon. Weather Rev., 144, 615–626, 2016.</mixed-citation></ref>
      <ref id="bib1.bibx5"><label>Doblas-Reyes et al.(2003)Doblas-Reyes, Pavan, and
Stephenson</label><?label DobPavSte03?><mixed-citation>
Doblas-Reyes, F. J., Pavan, V., and Stephenson, D. B.: The skill of multi-model
seasonal forecasts of the wintertime North Atlantic Oscillation, Clim.
Dynam., 21, 501–514, 2003.</mixed-citation></ref>
      <ref id="bib1.bibx6"><?xmltex \def\ref@label{{Dobrynin et~al.(2018)Dobrynin, Domeisen, M{\"{u}}ller, Bell, Brune,
Bunzel, D\"{u}sterhus, Fr{\"{o}}hlich, Pohlmann, and Baehr}}?><label>Dobrynin et al.(2018)Dobrynin, Domeisen, Müller, Bell, Brune,
Bunzel, Düsterhus, Fröhlich, Pohlmann, a<?pagebreak page131?>nd Baehr</label><?label DobDomMul1804?><mixed-citation>
Dobrynin, M., Domeisen, D. I. V., Müller, W. A., Bell, L., Brune, S.,
Bunzel, F., Düsterhus, A., Fröhlich, K., Pohlmann, H., and Baehr, J.:
Improved teleconnection-based seasonal predictions of boreal winter,
Geophys. Res. Lett., 45, 3605–3614, 2018.</mixed-citation></ref>
      <ref id="bib1.bibx7"><?xmltex \def\ref@label{{D\"{u}sterhus and Hense(2012)}}?><label>Düsterhus and Hense(2012)</label><?label DusHen12?><mixed-citation>
Düsterhus, A. and Hense, A.: Advanced information criterion for
environmental data quality assurance, Adv. Sci. Res., 8,
99–104, 2012.</mixed-citation></ref>
      <ref id="bib1.bibx8"><label>Folland et al.(2009)Folland, Knight, Linderholm, Fereday, Ineson, and
Hurrell</label><?label FollandKnight2009?><mixed-citation>
Folland, C. K., Knight, J., Linderholm, H. W., Fereday, D., Ineson, S., and
Hurrell, J. W.: The Summer North Atlantic Oscillation: Past, Present, and
Future, J. Climate, 22, 1082–1103, 2009.</mixed-citation></ref>
      <ref id="bib1.bibx9"><label>Gleeson(1970)</label><?label Gle70?><mixed-citation>
Gleeson, T. A.: Statistical-Dynamical Predictions, J. Appl.
Meteorol., 9, 333–334, 1970.</mixed-citation></ref>
      <ref id="bib1.bibx10"><label>Glowienka-Hense(1990)</label><?label GlowienkaHense1990?><mixed-citation>
Glowienka-Hense, R.: The North Atlantic Oscillation in the
Atlantic-European SLP, Tellus, 42, 497–507, 1990.</mixed-citation></ref>
      <ref id="bib1.bibx11"><label>Hurrell(1995)</label><?label Hur95?><mixed-citation>
Hurrell, J. W.: Decadal Trends in the North Atlantic Oscillation: Regional
Temperatures and Precipitation, Science, 269, 676–679, 1995.</mixed-citation></ref>
      <ref id="bib1.bibx12"><?xmltex \def\ref@label{{M{\"{u}}ller et~al.(2005)M{\"{u}}ller, Appenzeller, and
Sch{\"{a}}r}}?><label>Müller et al.(2005)Müller, Appenzeller, and
Schär</label><?label MullerAppenzeller2005?><mixed-citation>
Müller, W. A., Appenzeller, C., and Schär, C.: Probabilistic seasonal
prediction of the winter North Atlantic Oscillation and its impact on near
surface temperature, Clim. Dynam., 24, 213–226, 2005.</mixed-citation></ref>
      <ref id="bib1.bibx13"><label>Raftery et al.(2005)Raftery, Gneiting, Balabdaoui, and
Polakowski</label><?label RafGneBal0505?><mixed-citation>
Raftery, A. E., Gneiting, T., Balabdaoui, F., and Polakowski, M.: Using
Bayesian Model Averaging to Calibrate Forecast Ensembles, Mon. Weather
Rev., 133, 1155–1174, 2005.</mixed-citation></ref>
      <ref id="bib1.bibx14"><label>Roulston and Smith(2003)</label><?label RouSmi03?><mixed-citation>Roulston, M. S. and Smith, L. A.: Combining dynamical and statistical
ensembles, Tellus A, 55, 16–30,
<ext-link xlink:href="https://doi.org/10.3402/tellusa.v55i1.12082" ext-link-type="DOI">10.3402/tellusa.v55i1.12082</ext-link>, 2003.</mixed-citation></ref>
      <ref id="bib1.bibx15"><label>Rubner et al.(2001)Rubner, Puzicha, Tomasi, and
Buhmann</label><?label RubPuzTom01?><mixed-citation>Rubner, Y., Puzicha, J., Tomasi, C., and Buhmann, J. M.: Empirical Evaluation
of Dissimilarity Measures for Color and Texture, Comput. Vis. Image
Und., 84, 25–43, 2001.
 </mixed-citation></ref><?xmltex \hack{\newpage}?>
      <ref id="bib1.bibx16"><label>Scaife et al.(2014)Scaife, Arribas, Blockley, Brookshaw, Clark,
Dunstone, Eade, Fereday, Folland, Gordon, Hermanson, Knight, Lea, MacLachlan,
Maidens, Martin, Peterson, Smith, Vellinga, Wallace, Waters, and
Williams</label><?label ScaifeArribas2014?><mixed-citation>
Scaife, A. A., Arribas, A., Blockley, E., Brookshaw, A., Clark, R. T.,
Dunstone, N., Eade, R., Fereday, D., Folland, C. K., Gordon, M., Hermanson,
L., Knight, J. R., Lea, D. J., MacLachlan, C., Maidens, A., Martin, M.,
Peterson, A. K., Smith, D., Vellinga, M., Wallace, E., Waters, J., and
Williams, A.: Skillful long-range prediction of European and North American
winters, Geophys. Res. Lett., 41, 2514–2519, 2014.</mixed-citation></ref>
      <ref id="bib1.bibx17"><?xmltex \def\ref@label{{Sch{\"{o}}lzel and Hense(2011)}}?><label>Schölzel and Hense(2011)</label><?label SchHen1105?><mixed-citation>
Schölzel, C. and Hense, A.: Probabilistic assessment of regional climate
change in Southwest Germany by ensemble dressing, Clim. Dynam., 36,
2003–2014, 2011.</mixed-citation></ref>
      <ref id="bib1.bibx18"><label>Thorarinsdottir et al.(2013)Thorarinsdottir, Gneiting, and
Gissibl</label><?label ThoGneGis1306?><mixed-citation>
Thorarinsdottir, T. L., Gneiting, T., and Gissibl, N.: Using Proper Divergence
Functions to Evaluate Climate Models, SIAM/ASA J. Uncertainty Quantification,
1, 522–534, 2013.</mixed-citation></ref>
      <ref id="bib1.bibx19"><label>Vecchi et al.(2011)Vecchi, Zhao, Wang, Villarini, Rosati, Kumar,
Held, and Gudgel</label><?label VecZhaWan11?><mixed-citation>Vecchi, G. A., Zhao, M., Wang, H., Villarini, G., Rosati, A., Kumar, A., Held,
I. M., and Gudgel, R.: Statistical–Dynamical Predictions of Seasonal North
Atlantic Hurricane Activity, Mon. Weather Rev., 139, 1070–1082,
<ext-link xlink:href="https://doi.org/10.1175/2010MWR3499.1" ext-link-type="DOI">10.1175/2010MWR3499.1</ext-link>, 2011.</mixed-citation></ref>
      <ref id="bib1.bibx20"><label>Wang et al.(2014)Wang, Magnusdottir, Stern, Tian, and
Yu</label><?label WanMagSte1402?><mixed-citation>
Wang, Y.-H., Magnusdottir, G., Stern, H., Tian, X., and Yu, Y.: Uncertainty
Estimates of the EOF-Derived North Atlantic Oscillation, J. Climate,
27, 1290–1301, 2014.</mixed-citation></ref>
      <ref id="bib1.bibx21"><label>Williams et al.(2014)Williams, Ferro, and Kwasniok</label><?label WilFerKwa1404?><mixed-citation>
Williams, R. M., Ferro, C. A. T., and Kwasniok, F.: A comparison of ensemble
post-processing methods for extreme events, Q. J. Roy.
Meteor. Soc., 140, 1112–1120, 2014.</mixed-citation></ref>

  </ref-list></back>
    <!--<article-title-html>Seasonal statistical–dynamical prediction of the North Atlantic Oscillation by probabilistic post-processing and its evaluation</article-title-html>
<abstract-html><p>Dynamical models of various centres have shown in recent years seasonal prediction skill of the North Atlantic Oscillation (NAO). By filtering the ensemble members on the basis of statistical predictors, known as subsampling, it is possible to achieve even higher prediction skill. In this study the aim is to design a generalisation of the subsampling approach and establish it as a post-processing procedure.</p><p>Instead of selecting discrete ensemble members for each year, as the subsampling approach does, the distributions of ensembles and statistical predictors are combined to create a probabilistic prediction of the winter NAO. By comparing the combined statistical–dynamical prediction with the predictions of its single components, it can be shown that it achieves similar results to the statistical prediction. At the same time it can be shown that, unlike the statistical prediction, the combined prediction has fewer years where it performs worse than the dynamical prediction.</p><p>By applying the gained distributions to other meteorological variables, like geopotential height, precipitation and surface temperature, it can be shown that evaluating prediction skill depends highly on the chosen metric. Besides the common anomaly correlation (ACC) this study also presents scores based on the Earth mover's distance (EMD) and the integrated quadratic distance (IQD), which are designed to evaluate skills of probabilistic predictions. It shows that by evaluating the predictions for each year separately compared to applying a metric to all years at the same time, like correlation-based metrics, leads to different interpretations of the analysis.</p></abstract-html>
<ref-html id="bib1.bib1"><label>Baehr et al.(2015)Baehr, Fröhlich, Botzet, Domeisen, Kornblueh,
Notz, Piontek, Pohlmann, Tietsche, and Müller</label><mixed-citation>
Baehr, J., Fröhlich, K., Botzet, M., Domeisen, D. I. V., Kornblueh, L.,
Notz, D., Piontek, R., Pohlmann, H., Tietsche, S., and Müller, W. A.: The
prediction of surface temperature in the new seasonal prediction system based
on the MPI-ESM coupled climate model, Clim. Dynam., 44, 2723–2735, 2015.
</mixed-citation></ref-html>
<ref-html id="bib1.bib2"><label>Butler et al.(2016)Butler, Arribas, Athanassiadou, Baehr, Calvo,
Charlton-Perez, Déqué, Domeisen, Fröhlich, Hendon, Imada, Ishii,
Iza, Karpechko, Kumar, MacLachlan, Merryfield, Müller, O'Neill, Scaife,
Scinocca, Sigmond, Stockdale, and Yasuda</label><mixed-citation>
Butler, A. H., Arribas, A., Athanassiadou, M., Baehr, J., Calvo, N.,
Charlton-Perez, A., Déqué, M., Domeisen, D. I. V., Fröhlich, K.,
Hendon, H., Imada, Y., Ishii, M., Iza, M., Karpechko, A. Y., Kumar, A.,
MacLachlan, C., Merryfield, W. J., Müller, W. A., O'Neill, A., Scaife,
A. A., Scinocca, J., Sigmond, M., Stockdale, T. N., and Yasuda, T.: The
Climate-system Historical Forecast Project: do stratosphere-resolving models
make better seasonal climate predictions in boreal winter?, Q. J.
Roy. Meteor. Soc., 142, 1413–1427, 2016.
</mixed-citation></ref-html>
<ref-html id="bib1.bib3"><label>Dee et al.(2011)Dee, Uppala, Simmons, Berrisford, Poli, Kobayashi,
Andrae, Balmaseda, Balsamo, Bauer, Bechtold, Beljaars, van de Berg, Bidlot,
Bormann, Delsol, Dragani, Fuentes, Geer, Haimberger, Healy, Hersbach,
Hólm, Isaksen, Kållberg, Köhler, Matricardi, McNally,
Monge‐Sanz, Morcrette, Park, Peubey, de Rosnay, Tavolato, Thépaut, and
Vitart</label><mixed-citation>
Dee, D. P., Uppala, S. M., Simmons, A. J., Berrisford, P., Poli, P., Kobayashi,
S., Andrae, U., Balmaseda, M. A., Balsamo, G., Bauer, P., Bechtold, P.,
Beljaars, A. C. M., van de Berg, L., Bidlot, J., Bormann, N., Delsol, C.,
Dragani, R., Fuentes, M., Geer, A. J., Haimberger, L., Healy, S. B.,
Hersbach, H., Hólm, E. V., Isaksen, L., Kållberg, P., Köhler, M.,
Matricardi, M., McNally, A. P., Monge‐Sanz, B. M., Morcrette, J., Park, B.,
Peubey, C., de Rosnay, P., Tavolato, C., Thépaut, J., and Vitart, F.: The
ERA‐Interim reanalysis: configuration and performance of the data
assimilation system, Q. J. Roy. Meteor. Soc.,
137, 553–597, 2011.
</mixed-citation></ref-html>
<ref-html id="bib1.bib4"><label>DelSole and Tippett(2016)</label><mixed-citation>
DelSole, T. and Tippett, M. K.: Forecast Comparison Based on Random Walks,
Mon. Weather Rev., 144, 615–626, 2016.
</mixed-citation></ref-html>
<ref-html id="bib1.bib5"><label>Doblas-Reyes et al.(2003)Doblas-Reyes, Pavan, and
Stephenson</label><mixed-citation>
Doblas-Reyes, F. J., Pavan, V., and Stephenson, D. B.: The skill of multi-model
seasonal forecasts of the wintertime North Atlantic Oscillation, Clim.
Dynam., 21, 501–514, 2003.
</mixed-citation></ref-html>
<ref-html id="bib1.bib6"><label>Dobrynin et al.(2018)Dobrynin, Domeisen, Müller, Bell, Brune,
Bunzel, Düsterhus, Fröhlich, Pohlmann, and Baehr</label><mixed-citation>
Dobrynin, M., Domeisen, D. I. V., Müller, W. A., Bell, L., Brune, S.,
Bunzel, F., Düsterhus, A., Fröhlich, K., Pohlmann, H., and Baehr, J.:
Improved teleconnection-based seasonal predictions of boreal winter,
Geophys. Res. Lett., 45, 3605–3614, 2018.
</mixed-citation></ref-html>
<ref-html id="bib1.bib7"><label>Düsterhus and Hense(2012)</label><mixed-citation>
Düsterhus, A. and Hense, A.: Advanced information criterion for
environmental data quality assurance, Adv. Sci. Res., 8,
99–104, 2012.
</mixed-citation></ref-html>
<ref-html id="bib1.bib8"><label>Folland et al.(2009)Folland, Knight, Linderholm, Fereday, Ineson, and
Hurrell</label><mixed-citation>
Folland, C. K., Knight, J., Linderholm, H. W., Fereday, D., Ineson, S., and
Hurrell, J. W.: The Summer North Atlantic Oscillation: Past, Present, and
Future, J. Climate, 22, 1082–1103, 2009.
</mixed-citation></ref-html>
<ref-html id="bib1.bib9"><label>Gleeson(1970)</label><mixed-citation>
Gleeson, T. A.: Statistical-Dynamical Predictions, J. Appl.
Meteorol., 9, 333–334, 1970.
</mixed-citation></ref-html>
<ref-html id="bib1.bib10"><label>Glowienka-Hense(1990)</label><mixed-citation>
Glowienka-Hense, R.: The North Atlantic Oscillation in the
Atlantic-European SLP, Tellus, 42, 497–507, 1990.
</mixed-citation></ref-html>
<ref-html id="bib1.bib11"><label>Hurrell(1995)</label><mixed-citation>
Hurrell, J. W.: Decadal Trends in the North Atlantic Oscillation: Regional
Temperatures and Precipitation, Science, 269, 676–679, 1995.
</mixed-citation></ref-html>
<ref-html id="bib1.bib12"><label>Müller et al.(2005)Müller, Appenzeller, and
Schär</label><mixed-citation>
Müller, W. A., Appenzeller, C., and Schär, C.: Probabilistic seasonal
prediction of the winter North Atlantic Oscillation and its impact on near
surface temperature, Clim. Dynam., 24, 213–226, 2005.
</mixed-citation></ref-html>
<ref-html id="bib1.bib13"><label>Raftery et al.(2005)Raftery, Gneiting, Balabdaoui, and
Polakowski</label><mixed-citation>
Raftery, A. E., Gneiting, T., Balabdaoui, F., and Polakowski, M.: Using
Bayesian Model Averaging to Calibrate Forecast Ensembles, Mon. Weather
Rev., 133, 1155–1174, 2005.
</mixed-citation></ref-html>
<ref-html id="bib1.bib14"><label>Roulston and Smith(2003)</label><mixed-citation>
Roulston, M. S. and Smith, L. A.: Combining dynamical and statistical
ensembles, Tellus A, 55, 16–30,
<a href="https://doi.org/10.3402/tellusa.v55i1.12082" target="_blank">https://doi.org/10.3402/tellusa.v55i1.12082</a>, 2003.
</mixed-citation></ref-html>
<ref-html id="bib1.bib15"><label>Rubner et al.(2001)Rubner, Puzicha, Tomasi, and
Buhmann</label><mixed-citation>
Rubner, Y., Puzicha, J., Tomasi, C., and Buhmann, J. M.: Empirical Evaluation
of Dissimilarity Measures for Color and Texture, Comput. Vis. Image
Und., 84, 25–43, 2001.

</mixed-citation></ref-html>
<ref-html id="bib1.bib16"><label>Scaife et al.(2014)Scaife, Arribas, Blockley, Brookshaw, Clark,
Dunstone, Eade, Fereday, Folland, Gordon, Hermanson, Knight, Lea, MacLachlan,
Maidens, Martin, Peterson, Smith, Vellinga, Wallace, Waters, and
Williams</label><mixed-citation>
Scaife, A. A., Arribas, A., Blockley, E., Brookshaw, A., Clark, R. T.,
Dunstone, N., Eade, R., Fereday, D., Folland, C. K., Gordon, M., Hermanson,
L., Knight, J. R., Lea, D. J., MacLachlan, C., Maidens, A., Martin, M.,
Peterson, A. K., Smith, D., Vellinga, M., Wallace, E., Waters, J., and
Williams, A.: Skillful long-range prediction of European and North American
winters, Geophys. Res. Lett., 41, 2514–2519, 2014.
</mixed-citation></ref-html>
<ref-html id="bib1.bib17"><label>Schölzel and Hense(2011)</label><mixed-citation>
Schölzel, C. and Hense, A.: Probabilistic assessment of regional climate
change in Southwest Germany by ensemble dressing, Clim. Dynam., 36,
2003–2014, 2011.
</mixed-citation></ref-html>
<ref-html id="bib1.bib18"><label>Thorarinsdottir et al.(2013)Thorarinsdottir, Gneiting, and
Gissibl</label><mixed-citation>
Thorarinsdottir, T. L., Gneiting, T., and Gissibl, N.: Using Proper Divergence
Functions to Evaluate Climate Models, SIAM/ASA J. Uncertainty Quantification,
1, 522–534, 2013.
</mixed-citation></ref-html>
<ref-html id="bib1.bib19"><label>Vecchi et al.(2011)Vecchi, Zhao, Wang, Villarini, Rosati, Kumar,
Held, and Gudgel</label><mixed-citation>
Vecchi, G. A., Zhao, M., Wang, H., Villarini, G., Rosati, A., Kumar, A., Held,
I. M., and Gudgel, R.: Statistical–Dynamical Predictions of Seasonal North
Atlantic Hurricane Activity, Mon. Weather Rev., 139, 1070–1082,
<a href="https://doi.org/10.1175/2010MWR3499.1" target="_blank">https://doi.org/10.1175/2010MWR3499.1</a>, 2011.
</mixed-citation></ref-html>
<ref-html id="bib1.bib20"><label>Wang et al.(2014)Wang, Magnusdottir, Stern, Tian, and
Yu</label><mixed-citation>
Wang, Y.-H., Magnusdottir, G., Stern, H., Tian, X., and Yu, Y.: Uncertainty
Estimates of the EOF-Derived North Atlantic Oscillation, J. Climate,
27, 1290–1301, 2014.
</mixed-citation></ref-html>
<ref-html id="bib1.bib21"><label>Williams et al.(2014)Williams, Ferro, and Kwasniok</label><mixed-citation>
Williams, R. M., Ferro, C. A. T., and Kwasniok, F.: A comparison of ensemble
post-processing methods for extreme events, Q. J. Roy.
Meteor. Soc., 140, 1112–1120, 2014.
</mixed-citation></ref-html>--></article>
