<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing with OASIS Tables v3.0 20080202//EN" "journalpub-oasis3.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:oasis="http://docs.oasis-open.org/ns/oasis-exchange/table" xml:lang="en" dtd-version="3.0"><?xmltex \makeatother\@nolinetrue\makeatletter?>
  <front>
    <journal-meta><journal-id journal-id-type="publisher">NPG</journal-id><journal-title-group>
    <journal-title>Nonlinear Processes in Geophysics</journal-title>
    <abbrev-journal-title abbrev-type="publisher">NPG</abbrev-journal-title><abbrev-journal-title abbrev-type="nlm-ta">Nonlin. Processes Geophys.</abbrev-journal-title>
  </journal-title-group><issn pub-type="epub">1607-7946</issn><publisher>
    <publisher-name>Copernicus Publications</publisher-name>
    <publisher-loc>Göttingen, Germany</publisher-loc>
  </publisher></journal-meta>
    <article-meta>
      <article-id pub-id-type="doi">10.5194/npg-27-411-2020</article-id><title-group><article-title>Beyond univariate calibration: verifying spatial structure in ensembles of forecast fields</article-title><alt-title>Verifying spatial structure in ensembles of forecast fields</alt-title>
      </title-group><?xmltex \runningtitle{Verifying spatial structure in ensembles of forecast fields}?><?xmltex \runningauthor{J.~Jacobson~et~al.}?>
      <contrib-group>
        <contrib contrib-type="author" corresp="yes" rid="aff1">
          <name><surname>Jacobson</surname><given-names>Josh</given-names></name>
          <email>josh.jacobson@colorado.edu</email>
        <ext-link>https://orcid.org/0000-0003-4418-2208</ext-link></contrib>
        <contrib contrib-type="author" corresp="no" rid="aff1">
          <name><surname>Kleiber</surname><given-names>William</given-names></name>
          
        </contrib>
        <contrib contrib-type="author" corresp="no" rid="aff2 aff3">
          <name><surname>Scheuerer</surname><given-names>Michael</given-names></name>
          
        <ext-link>https://orcid.org/0000-0003-4540-9478</ext-link></contrib>
        <contrib contrib-type="author" corresp="no" rid="aff2 aff3">
          <name><surname>Bellier</surname><given-names>Joseph</given-names></name>
          
        </contrib>
        <aff id="aff1"><label>1</label><institution>Department of Applied Mathematics, University of Colorado, Boulder, Colorado, USA</institution>
        </aff>
        <aff id="aff2"><label>2</label><institution>Cooperative Institute for Research in Environmental Sciences, University of Colorado, Boulder, Colorado, USA</institution>
        </aff>
        <aff id="aff3"><label>3</label><institution>Physical Sciences Laboratory, National Oceanic and Atmospheric Administration, Boulder, Colorado, USA</institution>
        </aff>
      </contrib-group>
      <author-notes><corresp id="corr1">Josh Jacobson (josh.jacobson@colorado.edu)</corresp></author-notes><pub-date><day>31</day><month>August</month><year>2020</year></pub-date>
      
      <volume>27</volume>
      <issue>3</issue>
      <fpage>411</fpage><lpage>427</lpage>
      <history>
        <date date-type="received"><day>24</day><month>December</month><year>2019</year></date>
           <date date-type="accepted"><day>21</day><month>July</month><year>2020</year></date>
           <date date-type="rev-recd"><day>15</day><month>July</month><year>2020</year></date>
           <date date-type="rev-request"><day>20</day><month>January</month><year>2020</year></date>
      </history>
      <permissions>
        <copyright-statement>Copyright: © 2020 </copyright-statement>
        <copyright-year>2020</copyright-year>
      <license license-type="open-access"><license-p>This work is licensed under the Creative Commons Attribution 4.0 International License. To view a copy of this licence, visit <ext-link ext-link-type="uri" xlink:href="https://creativecommons.org/licenses/by/4.0/">https://creativecommons.org/licenses/by/4.0/</ext-link></license-p></license></permissions><self-uri xlink:href="https://npg.copernicus.org/articles/.html">This article is available from https://npg.copernicus.org/articles/.html</self-uri><self-uri xlink:href="https://npg.copernicus.org/articles/.pdf">The full text article is available as a PDF file from https://npg.copernicus.org/articles/.pdf</self-uri>
      <abstract><title>Abstract</title>
    <p id="d1e121">Most available verification metrics for ensemble forecasts focus on univariate quantities. That is, they assess whether the ensemble provides an
adequate representation of the forecast uncertainty about the quantity of interest at a particular location and time. For spatially indexed ensemble  forecasts, however, it is also important that forecast fields reproduce the spatial structure of the observed field and represent the uncertainty  about spatial properties such as the size of the area for which heavy precipitation, high winds, critical fire weather conditions, etc., are
expected. In this article we study the properties of the fraction of threshold exceedance (FTE) histogram, a new diagnostic tool designed for
spatially indexed ensemble forecast fields. Defined as the fraction of grid points where a prescribed threshold is exceeded, the FTE is calculated  for the verification field and separately for each ensemble member. It yields a projection of a – possibly high-dimensional – multivariate
quantity onto a univariate quantity that can be studied with standard tools like verification rank histograms. This projection is appealing since it
reflects a spatial property that is intuitive and directly relevant in applications, though it is not obvious whether the FTE is sufficiently
sensitive to misrepresentation of spatial structure in the ensemble. In a comprehensive simulation study we find that departures from uniformity of
the FTE histograms can indeed be related to forecast ensembles with biased spatial variability and that these histograms detect shortcomings in the  spatial structure of ensemble forecast fields that are not obvious by eye. For demonstration, FTE histograms are applied in the context of spatially
downscaled ensemble precipitation forecast fields from NOAA's Global Ensemble Forecast System.</p>
  </abstract>
    </article-meta>
  </front>
<body>
      

<sec id="Ch1.S1" sec-type="intro">
  <label>1</label><title>Introduction</title>
      <p id="d1e133">Ensemble prediction systems like the ECMWF ensemble <xref ref-type="bibr" rid="bib1.bibx4" id="paren.1"/> or NOAA's Global Ensemble Forecast System (GEFS; <xref ref-type="bibr" rid="bib1.bibx33" id="altparen.2"/>) are now
state of the art in operational meteorological forecasting at weather prediction centers worldwide. One of the goals of ensemble forecasting is the
representation of uncertainty about the state of the atmosphere at a future time <xref ref-type="bibr" rid="bib1.bibx31 bib1.bibx23" id="paren.3"/>, and verification
metrics are required that can assess to what extent this goal is achieved. For univariate quantities, i.e., if forecasts are studied separately for each location and each forecast lead time, diagnostic tools like verification rank histograms <xref ref-type="bibr" rid="bib1.bibx1 bib1.bibx15" id="paren.4"/>, or reliability diagrams <xref ref-type="bibr" rid="bib1.bibx24" id="paren.5"/> can be used to check whether ensemble forecasts are calibrated, i.e., statistically consistent with the values that materialize.</p>
      <p id="d1e151">When entire forecast fields are considered, aspects beyond univariate calibration are important. For example, ensembles that yield reliable
probabilistic forecasts at each location may still over- or under-forecast regional minima/maxima if their members exhibit an inaccurate spatial
structure (e.g., <xref ref-type="bibr" rid="bib1.bibx9" id="altparen.6"/>, their Fig. 6). For weather variables like precipitation, which are used as inputs to hydrological
forecast models, it is crucial that accumulations over space and time (and the associated uncertainty) are predicted accurately, and this again
requires an adequate representation of spatial structure and temporal persistence of precipitation by the ensemble.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F1" specific-use="star"><?xmltex \currentcnt{1}?><label>Figure 1</label><caption><p id="d1e159">Simulated verification field and three associated forecast fields (arbitrary color scale) in which the spatial correlation length is either the same as for the verification, 10 % miscalibrated, or 50 % miscalibrated. Can you tell which is correct?</p></caption>
        <?xmltex \igopts{width=497.923228pt}?><graphic xlink:href="https://npg.copernicus.org/articles/27/411/2020/npg-27-411-2020-f01.png"/>

      </fig>

      <p id="d1e169">There is an added difficulty for forecasters in that misrepresentation of the spatial structure of weather variables by<?pagebreak page412?> ensemble forecast fields may not be discernible by eye. For example, consider the simulated fields in Fig. <xref ref-type="fig" rid="Ch1.F1"/>: perhaps one of these forecast fields has a
clearly different spatial correlation length than the verification, but we suspect that even the sharp-eyed reader cannot distinguish between the
remaining fields with confidence. Even if the differences are obvious, a quantitative verification metric is required to objectively compare different
forecast systems or methodologies.</p>
      <p id="d1e174">Several multivariate generalizations of verification rank histograms, such as minimum spanning tree histograms <xref ref-type="bibr" rid="bib1.bibx29 bib1.bibx32" id="paren.7"/>, multivariate rank histograms <xref ref-type="bibr" rid="bib1.bibx13" id="paren.8"/>, average-rank and band-depth rank histograms <xref ref-type="bibr" rid="bib1.bibx30" id="paren.9"/>, and copula probability
integral transform histograms <xref ref-type="bibr" rid="bib1.bibx34" id="paren.10"/>, have been proposed and allow one to assess different aspects of multivariate calibration. They are all based on different projections of the multivariate quantity of interest onto a univariate quantity that can then be studied
using standard verification rank histograms. Unfortunately, most of these projections do not allow an intuitive understanding of exactly what
multivariate aspect is being assessed, and none are tailored to the special case where the multivariate quantity of interest is a spatial field. Two
recent papers <xref ref-type="bibr" rid="bib1.bibx6 bib1.bibx5" id="paren.11"/> propose a wavelet-based verification approach in which wavelet transformations
of forecast and observed fields are performed to characterize and compare the fields' texture. The authors demonstrate that this approach is able to
detect differences in the spatial correlation length similar to those shown in Fig. <xref ref-type="fig" rid="Ch1.F1"/>. <xref ref-type="bibr" rid="bib1.bibx20" id="text.12"/> define a
skill score based on wavelet spectra and study the score differences between a randomly selected ensemble member and the verification field in order
to detect possible deficiencies in the texture of the forecast fields. Our goal is similar, but the approach studied here follows the idea of defining
a projection from the multivariate quantity (here: a spatial field) onto a univariate quantity that can be analyzed via verification rank histograms. Our main focus is on the probabilistic nature of the forecasts; that is, we want to test whether the ensemble adequately represents the <italic>uncertainty</italic> about spatial quantities.</p>
      <p id="d1e201">The projection underlying the verification metric studied here is based on threshold exceedances of the forecast and observation fields. This
binarization of continuous weather variables is common in spatial forecast verification (see <xref ref-type="bibr" rid="bib1.bibx12" id="altparen.13"/>) as it allows one to study, for
example, low, intermediate, and high precipitation amounts separately. In the context of deterministic forecast verification, <xref ref-type="bibr" rid="bib1.bibx25" id="text.14"/>
define the fractions skill score (FSS) based on the fraction of threshold exceedances (FTEs) within a certain neighborhood of every grid point and use it to examine at which spatial scale the forecast FTEs become skillful. <xref ref-type="bibr" rid="bib1.bibx28" id="text.15"/> use a similar concept to study whether
an ensemble of forecast fields adequately represents spatial forecast uncertainty. They calculate the FTE for all ensemble members and the verifying
observation field and study verification rank histograms of the resulting univariate quantity in order to diagnose the advantages and limitations of different statistical methods to generate high-resolution ensemble precipitation forecast fields based on lower-resolution NWP model output. The FTE is an interpretable quantity that is highly relevant in applications where the fraction of the forecast domain for which severe weather conditions are
expected (e.g., heavy rain, extreme wind speeds) may be of interest. However, it is not obvious whether FTE histograms are sufficiently sensitive to misrepresentation of the spatial structure by the ensemble, and the goal of the present paper is to investigate this discrimination
ability in detail.</p>
      <p id="d1e213">In Sect. 2, we describe the calculation of the FTE and the construction of the FTE histogram in detail. In Sect. 3, a simulation study is designed and
implemented that allows us to analyze the discrimination capability of the FTE histograms with regard to spatial structures. In Sect. 4, we
demonstrate the utility of FTE histograms in the context of spatially downscaled ensemble precipitation forecast fields from NOAA's Global Ensemble
Forecast System. A discussion and concluding remarks are given in Sect. 5.</p>
</sec>
<?pagebreak page413?><sec id="Ch1.S2">
  <label>2</label><title>The fraction of threshold exceedance metric</title>
      <p id="d1e224">Let <inline-formula><mml:math id="M1" display="inline"><mml:mrow><mml:mi>Z</mml:mi><mml:mo>(</mml:mo><mml:mi>s</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> be a scalar field on a domain <inline-formula><mml:math id="M2" display="inline"><mml:mrow><mml:mi>s</mml:mi><mml:mo>∈</mml:mo><mml:mi>D</mml:mi></mml:mrow></mml:math></inline-formula>. Here, we describe a strategy of studying exceedances of <inline-formula><mml:math id="M3" display="inline"><mml:mi>Z</mml:mi></mml:math></inline-formula> at various thresholds. That is, we
focus interest in statistics based on <inline-formula><mml:math id="M4" display="inline"><mml:mrow><mml:msub><mml:mn mathvariant="normal">1</mml:mn><mml:mrow><mml:mo mathvariant="italic">{</mml:mo><mml:mi>Z</mml:mi><mml:mo>(</mml:mo><mml:mi>s</mml:mi><mml:mo>)</mml:mo><mml:mo>&gt;</mml:mo><mml:mi mathvariant="italic">τ</mml:mi><mml:mo mathvariant="italic">}</mml:mo></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula> for a given threshold <inline-formula><mml:math id="M5" display="inline"><mml:mrow><mml:mi mathvariant="italic">τ</mml:mi><mml:mo>∈</mml:mo><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow></mml:math></inline-formula>. In the domain <inline-formula><mml:math id="M6" display="inline"><mml:mi>D</mml:mi></mml:math></inline-formula>, we define the
FTE as the fraction of all points at which <inline-formula><mml:math id="M7" display="inline"><mml:mi mathvariant="italic">τ</mml:mi></mml:math></inline-formula> is exceeded. Specifically, let

              <disp-formula specific-use="align" content-type="numbered"><mml:math id="M8" display="block"><mml:mtable displaystyle="true"><mml:mtr><mml:mtd><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><mml:mtext>FTE</mml:mtext><mml:mo>(</mml:mo><mml:mi>Z</mml:mi><mml:mo>,</mml:mo><mml:mi mathvariant="italic">τ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:mstyle displaystyle="true" class="stylechange"/><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mn mathvariant="normal">1</mml:mn><mml:mrow><mml:mo>|</mml:mo><mml:mi>D</mml:mi><mml:mo>|</mml:mo></mml:mrow></mml:mfrac></mml:mstyle><mml:munder><mml:mo movablelimits="false">∫</mml:mo><mml:mi>D</mml:mi></mml:munder><mml:msub><mml:mn mathvariant="normal">1</mml:mn><mml:mrow><mml:mo mathvariant="italic">{</mml:mo><mml:mi>Z</mml:mi><mml:mo>(</mml:mo><mml:mi>s</mml:mi><mml:mo>)</mml:mo><mml:mo>&gt;</mml:mo><mml:mi mathvariant="italic">τ</mml:mi><mml:mo mathvariant="italic">}</mml:mo></mml:mrow></mml:msub><mml:mo>(</mml:mo><mml:mi>s</mml:mi><mml:mo>)</mml:mo><mml:mi mathvariant="normal">d</mml:mi><mml:mi>s</mml:mi></mml:mrow></mml:mtd></mml:mtr><mml:mlabeledtr id="Ch1.E1"><mml:mtd><mml:mtext>1</mml:mtext></mml:mtd><mml:mtd><mml:mstyle class="stylechange" displaystyle="true"/></mml:mtd><mml:mtd><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mn mathvariant="normal">1</mml:mn><mml:mi>n</mml:mi></mml:mfrac></mml:mstyle><mml:munderover><mml:mo movablelimits="false">∑</mml:mo><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mi>n</mml:mi></mml:munderover><mml:msub><mml:mn mathvariant="normal">1</mml:mn><mml:mrow><mml:mo mathvariant="italic">{</mml:mo><mml:mi>Z</mml:mi><mml:mo>(</mml:mo><mml:mi>s</mml:mi><mml:mo>)</mml:mo><mml:mo>&gt;</mml:mo><mml:mi mathvariant="italic">τ</mml:mi><mml:mo mathvariant="italic">}</mml:mo></mml:mrow></mml:msub><mml:mo>(</mml:mo><mml:msub><mml:mi>s</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>)</mml:mo><mml:mo>,</mml:mo></mml:mrow></mml:mtd></mml:mlabeledtr></mml:mtable></mml:math></disp-formula>

          where the first equality represents the idealized continuous spatial process definition, while the second reflects the discrete nature of spatial
sampling in an operational probabilistic forecasting context with <inline-formula><mml:math id="M9" display="inline"><mml:mrow><mml:mi>D</mml:mi><mml:mo>=</mml:mo><mml:mo mathvariant="italic">{</mml:mo><mml:msub><mml:mi>s</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:mi mathvariant="normal">…</mml:mi><mml:mo>,</mml:mo><mml:msub><mml:mi>s</mml:mi><mml:mi>n</mml:mi></mml:msub><mml:mo mathvariant="italic">}</mml:mo></mml:mrow></mml:math></inline-formula>. The resulting univariate quantity can be evaluated by common
univariate verification metrics <xref ref-type="bibr" rid="bib1.bibx28" id="paren.16"/>.</p>
      <p id="d1e479">Suppose we have a <inline-formula><mml:math id="M10" display="inline"><mml:mi>k</mml:mi></mml:math></inline-formula>-member ensemble <inline-formula><mml:math id="M11" display="inline"><mml:mrow><mml:msub><mml:mi>Z</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub><mml:mo>(</mml:mo><mml:mi>s</mml:mi><mml:mo>)</mml:mo><mml:mo>,</mml:mo><mml:mi mathvariant="normal">…</mml:mi><mml:mo>,</mml:mo><mml:msub><mml:mi>Z</mml:mi><mml:mi>k</mml:mi></mml:msub><mml:mo>(</mml:mo><mml:mi>s</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> and associated verification field <inline-formula><mml:math id="M12" display="inline"><mml:mrow><mml:msub><mml:mi>Z</mml:mi><mml:mn mathvariant="normal">0</mml:mn></mml:msub><mml:mo>(</mml:mo><mml:mi>s</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> (e.g., observation or analysis) all on <inline-formula><mml:math id="M13" display="inline"><mml:mi>D</mml:mi></mml:math></inline-formula>;
let <inline-formula><mml:math id="M14" display="inline"><mml:mrow><mml:mi mathvariant="italic">π</mml:mi><mml:mo>=</mml:mo><mml:mo mathvariant="italic">{</mml:mo><mml:mtext>FTE</mml:mtext><mml:mo>(</mml:mo><mml:msub><mml:mi>Z</mml:mi><mml:mn mathvariant="normal">0</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:mi mathvariant="italic">τ</mml:mi><mml:mo>)</mml:mo><mml:mo>,</mml:mo><mml:mi mathvariant="normal">…</mml:mi><mml:mo>,</mml:mo><mml:mtext>FTE</mml:mtext><mml:mo>(</mml:mo><mml:msub><mml:mi>Z</mml:mi><mml:mi>k</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:mi mathvariant="italic">τ</mml:mi><mml:mo>)</mml:mo><mml:mo mathvariant="italic">}</mml:mo></mml:mrow></mml:math></inline-formula>. Note that <inline-formula><mml:math id="M15" display="inline"><mml:mi mathvariant="italic">π</mml:mi></mml:math></inline-formula> depends on the threshold, but for ease of exposition we do not
include this dependence in notation. We call <inline-formula><mml:math id="M16" display="inline"><mml:mi>r</mml:mi></mml:math></inline-formula> the rank of the verification FTE relative to the set of verification and ensemble forecast FTEs or the rank of <inline-formula><mml:math id="M17" display="inline"><mml:mrow><mml:mtext>FTE</mml:mtext><mml:mo>(</mml:mo><mml:msub><mml:mi>Z</mml:mi><mml:mn mathvariant="normal">0</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:mi mathvariant="italic">τ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> in <inline-formula><mml:math id="M18" display="inline"><mml:mi mathvariant="italic">π</mml:mi></mml:math></inline-formula>. There are three cases of interest when computing <inline-formula><mml:math id="M19" display="inline"><mml:mi>r</mml:mi></mml:math></inline-formula>: (1) no ties exist in <inline-formula><mml:math id="M20" display="inline"><mml:mi mathvariant="italic">π</mml:mi></mml:math></inline-formula>, (2) ties exist among a
subset of <inline-formula><mml:math id="M21" display="inline"><mml:mi mathvariant="italic">π</mml:mi></mml:math></inline-formula> that includes <inline-formula><mml:math id="M22" display="inline"><mml:mrow><mml:mtext>FTE</mml:mtext><mml:mo>(</mml:mo><mml:msub><mml:mi>Z</mml:mi><mml:mn mathvariant="normal">0</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:mi mathvariant="italic">τ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>, or (3) there is only one unique value in <inline-formula><mml:math id="M23" display="inline"><mml:mi mathvariant="italic">π</mml:mi></mml:math></inline-formula>. In the first case no special action is required,
and in the second case ties in rank are simply broken uniformly at random. The third case arises when all ensemble members have the exact same FTE as
the verification, as may occur, for example, when the precipitation amount reported by the verification and predicted by all ensemble members is below
the threshold <inline-formula><mml:math id="M24" display="inline"><mml:mi mathvariant="italic">τ</mml:mi></mml:math></inline-formula> everywhere in <inline-formula><mml:math id="M25" display="inline"><mml:mi>D</mml:mi></mml:math></inline-formula>. Instances of this case are completely uninformative for the purpose of diagnosing miscalibration and can be
discarded.</p>
      <p id="d1e705">Gathering ranks over <inline-formula><mml:math id="M26" display="inline"><mml:mi>N</mml:mi></mml:math></inline-formula> instances of forecast–verification pairs, <inline-formula><mml:math id="M27" display="inline"><mml:mrow><mml:msub><mml:mi>r</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:mi mathvariant="normal">…</mml:mi><mml:mo>,</mml:mo><mml:msub><mml:mi>r</mml:mi><mml:mi>N</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>, a natural way to communicate the FTE rank behavior is through a histogram (termed <italic>FTE histogram</italic> by <xref ref-type="bibr" rid="bib1.bibx28" id="altparen.17"/>) over the <inline-formula><mml:math id="M28" display="inline"><mml:mrow><mml:mi>k</mml:mi><mml:mo>+</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:math></inline-formula> possible ranks. Its construction is akin to that of the
univariate verification rank histogram discussed in <xref ref-type="bibr" rid="bib1.bibx1" id="text.18"/> and <xref ref-type="bibr" rid="bib1.bibx15" id="text.19"/>, but the latter only evaluates
the marginal distribution of the ensemble. The FTE histogram behaves similarly to the univariate verification rank histogram under marginal miscalibration in that overpopulated low (high) bins are an indication of an over-forecast (under-forecast) bias, and a <inline-formula><mml:math id="M29" display="inline"><mml:mo>∪</mml:mo></mml:math></inline-formula>-shaped (<inline-formula><mml:math id="M30" display="inline"><mml:mo lspace="0mm">∩</mml:mo></mml:math></inline-formula>-shaped) histogram is an indication of an under-dispersed (over-dispersed) ensemble. However, it is also sensitive to misrepresentation of
spatial correlations by the ensemble forecast fields. To see this, consider first the extreme case where the forecast fields are spatially
uncorrelated (i.e., spatial white noise) while the verification fields have maximal spatial correlations. In this setup, if <inline-formula><mml:math id="M31" display="inline"><mml:mi mathvariant="italic">τ</mml:mi></mml:math></inline-formula> is equal to the
climatological median of the marginal distributions at each grid point, the FTE for each ensemble member is close to <inline-formula><mml:math id="M32" display="inline"><mml:mn mathvariant="normal">0.5</mml:mn></mml:math></inline-formula>, while the FTE for the verification field is either <inline-formula><mml:math id="M33" display="inline"><mml:mn mathvariant="normal">0</mml:mn></mml:math></inline-formula> or <inline-formula><mml:math id="M34" display="inline"><mml:mn mathvariant="normal">1</mml:mn></mml:math></inline-formula>, with equal probability. The associated FTE histogram is <inline-formula><mml:math id="M35" display="inline"><mml:mo>∪</mml:mo></mml:math></inline-formula>-shaped, with half of the cases in the lowest bin and the other half in the highest bin. If <inline-formula><mml:math id="M36" display="inline"><mml:mi mathvariant="italic">τ</mml:mi></mml:math></inline-formula> is equal to the 95th climatological percentile, the FTE of each ensemble member is close to
<inline-formula><mml:math id="M37" display="inline"><mml:mn mathvariant="normal">0.05</mml:mn></mml:math></inline-formula> and the FTE of the verification field is <inline-formula><mml:math id="M38" display="inline"><mml:mn mathvariant="normal">0</mml:mn></mml:math></inline-formula> with probability <inline-formula><mml:math id="M39" display="inline"><mml:mn mathvariant="normal">0.95</mml:mn></mml:math></inline-formula> and <inline-formula><mml:math id="M40" display="inline"><mml:mn mathvariant="normal">1</mml:mn></mml:math></inline-formula> with probability <inline-formula><mml:math id="M41" display="inline"><mml:mn mathvariant="normal">0.05</mml:mn></mml:math></inline-formula>. The associated FTE histogram is
<inline-formula><mml:math id="M42" display="inline"><mml:mo>∪</mml:mo></mml:math></inline-formula>-shaped <italic>and</italic> skewed, with 95 % of all cases in the lowest bin and 5 % of all cases in the highest bin. For <inline-formula><mml:math id="M43" display="inline"><mml:mi mathvariant="italic">τ</mml:mi></mml:math></inline-formula> equal to the 5th climatological percentile, the skewness is in the other direction, with 5 % (95 %) of all cases in the lowest (highest) bin. In a more
realistic situation, where both forecast end verification fields are spatially correlated but the spatial correlations of the forecast fields are too
weak (i.e., they exhibit too much spatial variability), we can still expect to see a somewhat <inline-formula><mml:math id="M44" display="inline"><mml:mo>∪</mml:mo></mml:math></inline-formula>-shaped FTE histogram since the verification FTE values are more likely to assume extreme ranks than the ensemble FTE values. For large values of <inline-formula><mml:math id="M45" display="inline"><mml:mi mathvariant="italic">τ</mml:mi></mml:math></inline-formula>, the lower bins will be more populated; for small values of <inline-formula><mml:math id="M46" display="inline"><mml:mi mathvariant="italic">τ</mml:mi></mml:math></inline-formula>, the higher bins will be more populated. Conversely, if the spatial correlations of the forecast fields are too strong, the
verification ranks will over-populate the central bins, slightly shifted upward or downward from the center depending on <inline-formula><mml:math id="M47" display="inline"><mml:mi mathvariant="italic">τ</mml:mi></mml:math></inline-formula>. An ensemble that is
marginally and spatially calibrated (i.e., the strength of spatial correlations within each forecast field matches that of the verification) will
result in a flat FTE histogram.</p>
      <p id="d1e901">If the marginal forecast distributions are miscalibrated, the resulting effects on the rank of the verification FTE are superimposed on those caused
by misrepresentation of spatial correlations. This complicates interpretation because it is often impossible to disentangle the different sources of
miscalibration (this loss of information is an inevitable consequence of projecting a multivariate quantity onto a univariate one), and it can even
happen that different effects cancel each other out. For example, ensemble forecast fields which are both under-dispersive and have too strong spatial
correlations may result in flat FTE histograms. This serves as a reminder that – as in the univariate case – a flat histogram is a necessary but not
sufficient condition for probabilistic calibration. It simply indicates that the verification and the ensemble are indistinguishable with regard to the particular aspect of the forecast fields (here: exceedance of a prespecified threshold) assessed by this metric. Systematic over- or
under-forecast biases can be accounted for by using different (depending on the respective climatology) threshold values <inline-formula><mml:math id="M48" display="inline"><mml:mi mathvariant="italic">τ</mml:mi></mml:math></inline-formula> for the forecast and
verification fields. We are not aware of an equally straightforward way to account for dispersion errors, so we encourage users to always check the
marginal forecast distributions first and then study FTE histograms for different thresholds, possibly<?pagebreak page414?> in conjunction with other multivariate verification metrics in order to obtain a comprehensive picture of the multivariate properties of the ensemble forecasts.</p>

<?xmltex \floatpos{t}?><table-wrap id="Ch1.T1" specific-use="star"><?xmltex \currentcnt{1}?><label>Table 1</label><caption><p id="d1e915">Characterization of FTE histogram shapes via <inline-formula><mml:math id="M49" display="inline"><mml:mi mathvariant="italic">β</mml:mi></mml:math></inline-formula>-score and <inline-formula><mml:math id="M50" display="inline"><mml:mi mathvariant="italic">β</mml:mi></mml:math></inline-formula>-bias and their interpretation with regard to potential deficiencies of the ensemble forecast fields.</p></caption><oasis:table frame="topbot"><oasis:tgroup cols="4">
     <oasis:colspec colnum="1" colname="col1" align="left"/>
     <oasis:colspec colnum="2" colname="col2" align="left"/>
     <oasis:colspec colnum="3" colname="col3" align="left"/>
     <oasis:colspec colnum="4" colname="col4" align="left"/>
     <oasis:thead>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">Histogram</oasis:entry>
         <oasis:entry colname="col2">Parameters</oasis:entry>
         <oasis:entry colname="col3">Score and bias</oasis:entry>
         <oasis:entry colname="col4">Interpretation</oasis:entry>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row>
         <oasis:entry colname="col1">Uniform</oasis:entry>
         <oasis:entry colname="col2"><inline-formula><mml:math id="M51" display="inline"><mml:mrow><mml:mi>a</mml:mi><mml:mo>=</mml:mo><mml:mi>b</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col3"><inline-formula><mml:math id="M52" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">β</mml:mi><mml:mi mathvariant="normal">S</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi mathvariant="italic">β</mml:mi><mml:mi mathvariant="normal">B</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col4">Ensemble FTEs consistent with verification FTE</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1"><inline-formula><mml:math id="M53" display="inline"><mml:mo>∪</mml:mo></mml:math></inline-formula>-shaped</oasis:entry>
         <oasis:entry colname="col2"><inline-formula><mml:math id="M54" display="inline"><mml:mrow><mml:mi>a</mml:mi><mml:mo>,</mml:mo><mml:mi>b</mml:mi><mml:mo>&lt;</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col3"><inline-formula><mml:math id="M55" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">β</mml:mi><mml:mi mathvariant="normal">S</mml:mi></mml:msub><mml:mo>&lt;</mml:mo><mml:mn mathvariant="normal">0</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col4">Under-dispersed marginal distributions OR excessive spatial variability</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1"><inline-formula><mml:math id="M56" display="inline"><mml:mo>∩</mml:mo></mml:math></inline-formula>-shaped</oasis:entry>
         <oasis:entry colname="col2"><inline-formula><mml:math id="M57" display="inline"><mml:mrow><mml:mi>a</mml:mi><mml:mo>,</mml:mo><mml:mi>b</mml:mi><mml:mo>&gt;</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col3"><inline-formula><mml:math id="M58" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">β</mml:mi><mml:mi mathvariant="normal">S</mml:mi></mml:msub><mml:mo>&gt;</mml:mo><mml:mn mathvariant="normal">0</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col4">Over-dispersed marginal distributions OR insufficient spatial variability</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Right-skewed</oasis:entry>
         <oasis:entry colname="col2"><inline-formula><mml:math id="M59" display="inline"><mml:mrow><mml:mi>a</mml:mi><mml:mo>&lt;</mml:mo><mml:mi>b</mml:mi></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col3"><inline-formula><mml:math id="M60" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">β</mml:mi><mml:mi mathvariant="normal">B</mml:mi></mml:msub><mml:mo>&gt;</mml:mo><mml:mn mathvariant="normal">0</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col4">Over-forecast bias OR excessive spatial variability at high thresholds</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Left-skewed</oasis:entry>
         <oasis:entry colname="col2"><inline-formula><mml:math id="M61" display="inline"><mml:mrow><mml:mi>a</mml:mi><mml:mo>&gt;</mml:mo><mml:mi>b</mml:mi></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col3"><inline-formula><mml:math id="M62" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">β</mml:mi><mml:mi mathvariant="normal">B</mml:mi></mml:msub><mml:mo>&lt;</mml:mo><mml:mn mathvariant="normal">0</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col4">Under-forecast bias OR insufficient spatial variability at high thresholds</oasis:entry>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup></oasis:table><table-wrap-foot><p id="d1e932">Skewness is exaggerated by high thresholds; see text for more detail.</p></table-wrap-foot></table-wrap>

      <p id="d1e1191">While the FTE histogram is a useful visual diagnostic tool, a quantitative measure for studying departures from uniformity is desirable. Akin to
<xref ref-type="bibr" rid="bib1.bibx21" id="text.20"/>, we fit a beta distribution to the histogram values (transformed to the unit interval) and characterize the histogram shape
based on the <inline-formula><mml:math id="M63" display="inline"><mml:mi mathvariant="italic">β</mml:mi></mml:math></inline-formula><italic>-score</italic> and <inline-formula><mml:math id="M64" display="inline"><mml:mi mathvariant="italic">β</mml:mi></mml:math></inline-formula><italic>-bias</italic>, respectively, defined as
          <disp-formula id="Ch1.E2" content-type="numbered"><label>2</label><mml:math id="M65" display="block"><mml:mrow><mml:msub><mml:mi mathvariant="italic">β</mml:mi><mml:mi mathvariant="normal">S</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>-</mml:mo><mml:msqrt><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mn mathvariant="normal">1</mml:mn><mml:mrow><mml:mi>a</mml:mi><mml:mo>⋅</mml:mo><mml:mi>b</mml:mi></mml:mrow></mml:mfrac></mml:mstyle></mml:msqrt><mml:mo>,</mml:mo><mml:mspace width="1em" linebreak="nobreak"/><mml:msub><mml:mi mathvariant="italic">β</mml:mi><mml:mi mathvariant="normal">B</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mi>b</mml:mi><mml:mo>-</mml:mo><mml:mi>a</mml:mi><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>
        where <inline-formula><mml:math id="M66" display="inline"><mml:mi>a</mml:mi></mml:math></inline-formula> and <inline-formula><mml:math id="M67" display="inline"><mml:mi>b</mml:mi></mml:math></inline-formula> are the two distribution parameters. Since histogram values only occur at discrete points in <inline-formula><mml:math id="M68" display="inline"><mml:mrow><mml:mo>[</mml:mo><mml:mn mathvariant="normal">0</mml:mn><mml:mo>,</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>]</mml:mo></mml:mrow></mml:math></inline-formula>, parameter estimation methods
will incur some bias due to the lack of data on the interior of adjacent ranks. Thus, we stochastically disaggregate the (transformed) ranks
<inline-formula><mml:math id="M69" display="inline"><mml:mrow><mml:msub><mml:mi>r</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:mi mathvariant="normal">…</mml:mi><mml:mo>,</mml:mo><mml:msub><mml:mi>r</mml:mi><mml:mi>N</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> to continuous values in <inline-formula><mml:math id="M70" display="inline"><mml:mrow><mml:mo>[</mml:mo><mml:mn mathvariant="normal">0</mml:mn><mml:mo>,</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>]</mml:mo></mml:mrow></mml:math></inline-formula> (see Appendix A for details) and fit a beta distribution via maximum likelihood. Together, the
<inline-formula><mml:math id="M71" display="inline"><mml:mi mathvariant="italic">β</mml:mi></mml:math></inline-formula>-score and <inline-formula><mml:math id="M72" display="inline"><mml:mi mathvariant="italic">β</mml:mi></mml:math></inline-formula>-bias provide a pair of succinct descriptive statistics which communicate the visual characteristics of the histogram and therefore the ensemble's calibration properties. In the ideal case, <inline-formula><mml:math id="M73" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">β</mml:mi><mml:mi mathvariant="normal">S</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M74" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">β</mml:mi><mml:mi mathvariant="normal">B</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> are both exactly zero, indicating that
the FTE histogram is perfectly uniform. In practice, these metrics are never exactly zero. The resulting set of possible deviations and broad
interpretations of the corresponding histogram shapes is outlined in Table <xref ref-type="table" rid="Ch1.T1"/>. With the <inline-formula><mml:math id="M75" display="inline"><mml:mi mathvariant="italic">β</mml:mi></mml:math></inline-formula>-score and <inline-formula><mml:math id="M76" display="inline"><mml:mi mathvariant="italic">β</mml:mi></mml:math></inline-formula>-bias, we have an easily interpreted measure of spatial forecast calibration.</p>
      <p id="d1e1385">In summary, the FTE metric is composed of three steps: (1) calculate the FTE of each verification and ensemble forecast field, (2) construct an FTE
histogram over available instances of forecast and verification times, and (3) derive the <inline-formula><mml:math id="M77" display="inline"><mml:mi mathvariant="italic">β</mml:mi></mml:math></inline-formula>-score and <inline-formula><mml:math id="M78" display="inline"><mml:mi mathvariant="italic">β</mml:mi></mml:math></inline-formula>-bias from the stochastically
disaggregated FTE histogram to characterize departure from uniformity.</p>
</sec>
<sec id="Ch1.S3">
  <label>3</label><title>Simulation study</title>
      <p id="d1e1410">In this section we consider an extensive simulation study to assess the ability of the proposed FTE histogram to diagnose deficiencies in the
representation of spatial variability by the ensemble forecast fields. Our simulations will be based on multivariate Gaussian processes where the
notion of “spatial variability” can be quantified in terms of a correlation length parameter. The various meteorological quantities of interest such
as precipitation and wind speeds can be quite heterogeneous and spatially nonstationary over the study domain. However, since we study the spatial
structure of threshold exceedances, a suitable choice of thresholds can mitigate these effects to a degree that multivariate, stationary Gaussian
processes can be viewed as a sufficiently flexible model for simulating realistic spatial fields. To see this, consider a strictly positive and
continuous variable <inline-formula><mml:math id="M79" display="inline"><mml:mrow><mml:mi>Z</mml:mi><mml:mo>(</mml:mo><mml:mi>s</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> at two spatial locations <inline-formula><mml:math id="M80" display="inline"><mml:mrow><mml:mi>s</mml:mi><mml:mo>=</mml:mo><mml:msub><mml:mi>s</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>s</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula> with possibly unequal continuous cumulative distribution functions <inline-formula><mml:math id="M81" display="inline"><mml:mrow><mml:msub><mml:mi>F</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M82" display="inline"><mml:mrow><mml:msub><mml:mi>F</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula>,
respectively. Rather than considering a spatially constant threshold such as 10 <inline-formula><mml:math id="M83" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">m</mml:mi><mml:mspace width="0.125em" linebreak="nobreak"/><mml:msup><mml:mi mathvariant="normal">s</mml:mi><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula> for wind gusts, we can use a location-dependent threshold, say the <inline-formula><mml:math id="M84" display="inline"><mml:mrow><mml:mn mathvariant="normal">90</mml:mn><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mi mathvariant="italic">%</mml:mi></mml:mrow></mml:math></inline-formula> climatological quantiles <inline-formula><mml:math id="M85" display="inline"><mml:mrow><mml:mi>q</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi>s</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M86" display="inline"><mml:mrow><mml:mi>q</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi>s</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> representing local characteristics. Then both quantities
<inline-formula><mml:math id="M87" display="inline"><mml:mrow><mml:msub><mml:mn mathvariant="normal">1</mml:mn><mml:mrow><mml:mo mathvariant="italic">{</mml:mo><mml:mi>Z</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi>s</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>)</mml:mo><mml:mo>&gt;</mml:mo><mml:mi>q</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi>s</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>)</mml:mo><mml:mo mathvariant="italic">}</mml:mo></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula>, <inline-formula><mml:math id="M88" display="inline"><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>,</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:mrow></mml:math></inline-formula> are identically distributed Bernoulli(0.1) random variables. Exploiting a standard Gaussian probability integral transformation method, we note that <inline-formula><mml:math id="M89" display="inline"><mml:mrow><mml:msup><mml:mi mathvariant="normal">Φ</mml:mi><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msup><mml:mo>(</mml:mo><mml:mi>F</mml:mi><mml:mo>(</mml:mo><mml:mi>Z</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi>s</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>)</mml:mo><mml:mo>)</mml:mo><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> is a standard normal random variable, where <inline-formula><mml:math id="M90" display="inline"><mml:mi mathvariant="normal">Φ</mml:mi></mml:math></inline-formula> is the cumulative distribution
function of a standard normal. Thus, the original probability of threshold exceedance can be written as

              <disp-formula specific-use="align" content-type="numbered"><mml:math id="M91" display="block"><mml:mtable displaystyle="true"><mml:mtr><mml:mtd><mml:mrow><mml:mstyle displaystyle="true" class="stylechange"/><mml:mi>P</mml:mi><mml:mo>(</mml:mo><mml:mi>Z</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi>s</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>)</mml:mo><mml:mo>&gt;</mml:mo><mml:mi>q</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi>s</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>)</mml:mo><mml:mo>)</mml:mo></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><mml:mo>=</mml:mo><mml:mi>P</mml:mi><mml:mo>(</mml:mo><mml:mi>F</mml:mi><mml:mo>(</mml:mo><mml:mi>Z</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi>s</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>)</mml:mo><mml:mo>)</mml:mo><mml:mo>&gt;</mml:mo><mml:mi>F</mml:mi><mml:mo>(</mml:mo><mml:mi>q</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi>s</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>)</mml:mo><mml:mo>)</mml:mo><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mstyle class="stylechange" displaystyle="true"/></mml:mtd><mml:mtd><mml:mrow><mml:mstyle displaystyle="true" class="stylechange"/><mml:mo>=</mml:mo><mml:mi>P</mml:mi><mml:mo>(</mml:mo><mml:msup><mml:mi mathvariant="normal">Φ</mml:mi><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msup><mml:mo>(</mml:mo><mml:mi>F</mml:mi><mml:mo>(</mml:mo><mml:mi>Z</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi>s</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>)</mml:mo><mml:mo>)</mml:mo><mml:mo>)</mml:mo><mml:mo>&gt;</mml:mo><mml:msup><mml:mi mathvariant="normal">Φ</mml:mi><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msup><mml:mo>(</mml:mo><mml:mi>F</mml:mi><mml:mo>(</mml:mo><mml:mi>q</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi>s</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>)</mml:mo><mml:mo>)</mml:mo><mml:mo>)</mml:mo><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mlabeledtr id="Ch1.E3"><mml:mtd><mml:mtext>3</mml:mtext></mml:mtd><mml:mtd><mml:mstyle displaystyle="true" class="stylechange"/></mml:mtd><mml:mtd><mml:mrow><mml:mstyle displaystyle="true" class="stylechange"/><mml:mo>=</mml:mo><mml:mi>P</mml:mi><mml:mo>(</mml:mo><mml:mi>X</mml:mi><mml:mo>&gt;</mml:mo><mml:msup><mml:mi mathvariant="normal">Φ</mml:mi><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msup><mml:mo>(</mml:mo><mml:mn mathvariant="normal">0.9</mml:mn><mml:mo>)</mml:mo><mml:mo>)</mml:mo><mml:mo>,</mml:mo></mml:mrow></mml:mtd></mml:mlabeledtr></mml:mtable></mml:math></disp-formula>

          where <inline-formula><mml:math id="M92" display="inline"><mml:mi>X</mml:mi></mml:math></inline-formula> is a standard normal. Thus, we have shown that a field of random variables with continuous, possibly distinct local probability
distributions can be transformed to standard Gaussian marginal distributions, and using local quantiles as the threshold is then equivalent to a
spatially constant threshold on the transformed variables. For weather variables with discrete–continuous marginal distributions (e.g., precipitation), this direction of the transformation is not quite as straightforward. Conversely, however, simulated fields from Gaussian processes
can always be transformed to any desired marginal distributions (including discrete–continuous ones). In our ensuing simulation studies we therefore consider stationary spatial Gaussian processes to represent forecast and verification fields.</p>
      <p id="d1e1841">The main technical difficulty in setting up the simulation study is in generating multiple, stationary Gaussian random fields that have different
correlation lengths while being correlated with each other. That is, we would like to generate <inline-formula><mml:math id="M93" display="inline"><mml:mrow><mml:msub><mml:mi>Z</mml:mi><mml:mn mathvariant="normal">0</mml:mn></mml:msub><mml:mo>(</mml:mo><mml:mi>s</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M94" display="inline"><mml:mrow><mml:msub><mml:mi>Z</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub><mml:mo>(</mml:mo><mml:mi>s</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> in such a way that
<inline-formula><mml:math id="M95" display="inline"><mml:mrow><mml:mtext>Cov</mml:mtext><mml:mo>(</mml:mo><mml:msub><mml:mi>Z</mml:mi><mml:mn mathvariant="normal">0</mml:mn></mml:msub><mml:mo>(</mml:mo><mml:mi>s</mml:mi><mml:mo>)</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mi>Z</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub><mml:mo>(</mml:mo><mml:mi>s</mml:mi><mml:mo>)</mml:mo><mml:mo>)</mml:mo><mml:mo>&gt;</mml:mo><mml:mn mathvariant="normal">0</mml:mn></mml:mrow></mml:math></inline-formula> (representing that the forecast field is correlated with the verification field) and where <inline-formula><mml:math id="M96" display="inline"><mml:mrow><mml:msub><mml:mi>Z</mml:mi><mml:mn mathvariant="normal">0</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M97" display="inline"><mml:mrow><mml:msub><mml:mi>Z</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula> have possibly
distinct correlation lengths (representing that the forecast field is spatially miscalibrated). A natural approach is to use multivariate random field
models.</p>
<sec id="Ch1.S3.SS1">
  <label>3.1</label><title>Multivariate Gaussian processes</title>
      <p id="d1e1948">We call a vector of processes <inline-formula><mml:math id="M98" display="inline"><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>Z</mml:mi><mml:mn mathvariant="normal">0</mml:mn></mml:msub><mml:mo>(</mml:mo><mml:mi>s</mml:mi><mml:mo>)</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mi>Z</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub><mml:mo>(</mml:mo><mml:mi>s</mml:mi><mml:mo>)</mml:mo><mml:mo>,</mml:mo><mml:mi mathvariant="normal">…</mml:mi><mml:mo>,</mml:mo><mml:msub><mml:mi>Z</mml:mi><mml:mi>k</mml:mi></mml:msub><mml:mo>(</mml:mo><mml:mi>s</mml:mi><mml:mo>)</mml:mo><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> a multivariate Gaussian process if its finite-dimensional distributions are multivariate
normal. We focus on second-order stationary mean zero multivariate Gaussian processes in that <inline-formula><mml:math id="M99" display="inline"><mml:mrow><mml:mi>E</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi>Z</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>(</mml:mo><mml:mi>s</mml:mi><mml:mo>)</mml:mo><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0</mml:mn></mml:mrow></mml:math></inline-formula> for all <inline-formula><mml:math id="M100" display="inline"><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0</mml:mn><mml:mo>,</mml:mo><mml:mi mathvariant="normal">…</mml:mi><mml:mo>,</mml:mo><mml:mi>k</mml:mi></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M101" display="inline"><mml:mrow><mml:mi>s</mml:mi><mml:mo>∈</mml:mo><mml:mi>D</mml:mi></mml:mrow></mml:math></inline-formula>. Stationarity implies that the stochastic process is characterized by

                <disp-formula specific-use="align" content-type="numbered"><mml:math id="M102" display="block"><mml:mtable displaystyle="true"><mml:mtr><mml:mtd><mml:mrow><mml:mstyle displaystyle="true" class="stylechange"/><mml:msub><mml:mi>C</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>(</mml:mo><mml:mi>h</mml:mi><mml:mo>)</mml:mo><mml:mo>=</mml:mo></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:mstyle displaystyle="true" class="stylechange"/><mml:mspace width="0.25em" linebreak="nobreak"/><mml:mtext>Cov</mml:mtext><mml:mo>(</mml:mo><mml:msub><mml:mi>Z</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>(</mml:mo><mml:mi>s</mml:mi><mml:mo>+</mml:mo><mml:mi>h</mml:mi><mml:mo>)</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mi>Z</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>(</mml:mo><mml:mi>s</mml:mi><mml:mo>)</mml:mo><mml:mo>)</mml:mo><mml:mo>,</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mlabeledtr id="Ch1.E4"><mml:mtd><mml:mtext>4</mml:mtext></mml:mtd><mml:mtd><mml:mstyle displaystyle="true" class="stylechange"/></mml:mtd><mml:mtd><mml:mrow><mml:mstyle displaystyle="true" class="stylechange"/><mml:mtext>for all</mml:mtext><mml:mspace width="0.25em" linebreak="nobreak"/><mml:mi>h</mml:mi><mml:mspace width="0.25em" linebreak="nobreak"/><mml:mtext>such that</mml:mtext><mml:mspace width="0.25em" linebreak="nobreak"/><mml:mi>s</mml:mi><mml:mo>+</mml:mo><mml:mi>h</mml:mi><mml:mo>∈</mml:mo><mml:mi>D</mml:mi><mml:mo>,</mml:mo></mml:mrow></mml:mtd></mml:mlabeledtr></mml:mtable></mml:math></disp-formula>

            which are called covariance functions for <inline-formula><mml:math id="M103" display="inline"><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:math></inline-formula> and cross-covariance functions for <inline-formula><mml:math id="M104" display="inline"><mml:mrow><mml:mi>i</mml:mi><mml:mo>≠</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:math></inline-formula>. Not all choices of functions <inline-formula><mml:math id="M105" display="inline"><mml:mrow><mml:msub><mml:mi>C</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula> will result in a
valid model; in particular, we require that the matrix of functions <inline-formula><mml:math id="M106" display="inline"><mml:mrow><mml:mi mathvariant="bold">C</mml:mi><mml:mo>(</mml:mo><mml:mi>h</mml:mi><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:mo>(</mml:mo><mml:msub><mml:mi>C</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>(</mml:mo><mml:mi>h</mml:mi><mml:mo>)</mml:mo><mml:msubsup><mml:mo>)</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0</mml:mn></mml:mrow><mml:mi>k</mml:mi></mml:msubsup></mml:mrow></mml:math></inline-formula> be a<?pagebreak page415?> nonnegative definite matrix function,
the technical definition of which can be found in <xref ref-type="bibr" rid="bib1.bibx11" id="text.21"/>.</p>
      <p id="d1e2250">There are many models for multivariate processes <xref ref-type="bibr" rid="bib1.bibx11" id="paren.22"/>, and here we exploit a particular class called the multivariate Matérn
<xref ref-type="bibr" rid="bib1.bibx14 bib1.bibx2" id="paren.23"/>. We rely on the popular Matérn correlation function
            <disp-formula id="Ch1.E5" content-type="numbered"><label>5</label><mml:math id="M107" display="block"><mml:mrow><mml:mi>M</mml:mi><mml:mo>(</mml:mo><mml:mi>d</mml:mi><mml:mo>|</mml:mo><mml:mi mathvariant="italic">ν</mml:mi><mml:mo>,</mml:mo><mml:mi>a</mml:mi><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:msup><mml:mn mathvariant="normal">2</mml:mn><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>-</mml:mo><mml:mi mathvariant="italic">ν</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mrow><mml:mi mathvariant="normal">Γ</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="italic">ν</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mfrac></mml:mstyle><mml:msup><mml:mfenced open="(" close=")"><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mi>d</mml:mi><mml:mi>a</mml:mi></mml:mfrac></mml:mstyle></mml:mfenced><mml:mi mathvariant="italic">ν</mml:mi></mml:msup><mml:msub><mml:mi>K</mml:mi><mml:mi mathvariant="italic">ν</mml:mi></mml:msub><mml:mfenced open="(" close=")"><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mi>d</mml:mi><mml:mi>a</mml:mi></mml:mfrac></mml:mstyle></mml:mfenced><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>
          where <inline-formula><mml:math id="M108" display="inline"><mml:mi mathvariant="normal">Γ</mml:mi></mml:math></inline-formula> is the gamma function, <inline-formula><mml:math id="M109" display="inline"><mml:mrow><mml:msub><mml:mi>K</mml:mi><mml:mi mathvariant="italic">ν</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> is the modified Bessel function of the second kind of order <inline-formula><mml:math id="M110" display="inline"><mml:mi mathvariant="italic">ν</mml:mi></mml:math></inline-formula>, and <inline-formula><mml:math id="M111" display="inline"><mml:mi>d</mml:mi></mml:math></inline-formula> is a nonnegative scalar. Parameters have interpretations as a smoothness (<inline-formula><mml:math id="M112" display="inline"><mml:mi mathvariant="italic">ν</mml:mi></mml:math></inline-formula>) and spatial range or correlation length (<inline-formula><mml:math id="M113" display="inline"><mml:mi>a</mml:mi></mml:math></inline-formula>). The multivariate Matérn correlation
function is defined as
            <disp-formula id="Ch1.E6" content-type="numbered"><label>6</label><mml:math id="M114" display="block"><mml:mrow><mml:msub><mml:mi>C</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>(</mml:mo><mml:mi>h</mml:mi><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:msubsup><mml:mi mathvariant="italic">σ</mml:mi><mml:mi>i</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msubsup><mml:mi>M</mml:mi><mml:mo>(</mml:mo><mml:mo>‖</mml:mo><mml:mi>h</mml:mi><mml:mo>‖</mml:mo><mml:mo>|</mml:mo><mml:msub><mml:mi mathvariant="italic">ν</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>a</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>)</mml:mo><mml:mo>,</mml:mo><mml:mspace linebreak="nobreak" width="1em"/><mml:mtext>for </mml:mtext><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0</mml:mn><mml:mo>,</mml:mo><mml:mi mathvariant="normal">…</mml:mi><mml:mo>,</mml:mo><mml:mi>k</mml:mi></mml:mrow></mml:math></disp-formula>
          and

                <disp-formula specific-use="align" content-type="numbered"><mml:math id="M115" display="block"><mml:mtable displaystyle="true"><mml:mtr><mml:mtd><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><mml:msub><mml:mi>C</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>(</mml:mo><mml:mi>h</mml:mi><mml:mo>)</mml:mo><mml:mo>=</mml:mo></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><mml:mspace width="0.25em" linebreak="nobreak"/><mml:msub><mml:mi>C</mml:mi><mml:mrow><mml:mi>j</mml:mi><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>(</mml:mo><mml:mi>h</mml:mi><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:msub><mml:mi mathvariant="italic">ρ</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mi mathvariant="italic">σ</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:msub><mml:mi mathvariant="italic">σ</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mi>M</mml:mi><mml:mo>(</mml:mo><mml:mo>‖</mml:mo><mml:mi>h</mml:mi><mml:mo>‖</mml:mo><mml:mo>|</mml:mo><mml:msub><mml:mi mathvariant="italic">ν</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>a</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo><mml:mo>,</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mlabeledtr id="Ch1.E7"><mml:mtd><mml:mtext>7</mml:mtext></mml:mtd><mml:mtd><mml:mstyle displaystyle="true" class="stylechange"/></mml:mtd><mml:mtd><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><mml:mtext>for </mml:mtext><mml:mn mathvariant="normal">0</mml:mn><mml:mo>≤</mml:mo><mml:mi>i</mml:mi><mml:mo>≠</mml:mo><mml:mi>j</mml:mi><mml:mo>≤</mml:mo><mml:mi>k</mml:mi><mml:mo>,</mml:mo></mml:mrow></mml:mtd></mml:mlabeledtr></mml:mtable></mml:math></disp-formula>

            where <inline-formula><mml:math id="M116" display="inline"><mml:mrow><mml:mo>‖</mml:mo><mml:mo>⋅</mml:mo><mml:mo>‖</mml:mo></mml:mrow></mml:math></inline-formula> is the Euclidean norm. In this latter equation, <inline-formula><mml:math id="M117" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">ρ</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>∈</mml:mo><mml:mo>[</mml:mo><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>,</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>]</mml:mo></mml:mrow></mml:math></inline-formula> is the co-located cross-correlation coefficient. Interpretation
of the cross-covariance parameters requires spectral techniques <xref ref-type="bibr" rid="bib1.bibx22" id="paren.24"/>.</p>
</sec>
<sec id="Ch1.S3.SS2">
  <label>3.2</label><title>Simulation setup</title>
      <p id="d1e2631">Simultaneously simulating the verification field <inline-formula><mml:math id="M118" display="inline"><mml:mrow><mml:msub><mml:mi>Z</mml:mi><mml:mn mathvariant="normal">0</mml:mn></mml:msub><mml:mo>(</mml:mo><mml:mi>s</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> and all forecast fields <inline-formula><mml:math id="M119" display="inline"><mml:mrow><mml:msub><mml:mi>Z</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub><mml:mo>(</mml:mo><mml:mi>s</mml:mi><mml:mo>)</mml:mo><mml:mo>,</mml:mo><mml:mi mathvariant="normal">…</mml:mi><mml:mo>,</mml:mo><mml:msub><mml:mi>Z</mml:mi><mml:mi>k</mml:mi></mml:msub><mml:mo>(</mml:mo><mml:mi>s</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> is difficult due to the high-dimensional
joint covariance matrix. Instead, we approach simulations by jointly simulating the verification fields <inline-formula><mml:math id="M120" display="inline"><mml:mrow><mml:msub><mml:mi>Z</mml:mi><mml:mn mathvariant="normal">0</mml:mn></mml:msub><mml:mo>(</mml:mo><mml:mi>s</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> and the (scaled) ensemble mean field <inline-formula><mml:math id="M121" display="inline"><mml:mrow><mml:msub><mml:mi>Z</mml:mi><mml:mi mathvariant="normal">M</mml:mi></mml:msub><mml:mo>(</mml:mo><mml:mi>s</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> from a bivariate Matérn model. We then perturb the mean field with independent univariate Gaussian random fields to generate an 11-member
ensemble, <inline-formula><mml:math id="M122" display="inline"><mml:mrow><mml:mi>k</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">11</mml:mn></mml:mrow></mml:math></inline-formula>.</p>
      <p id="d1e2731">The simulation setup follows a series of steps.
<list list-type="order"><list-item>
      <p id="d1e2736">Generate <inline-formula><mml:math id="M123" display="inline"><mml:mrow><mml:msub><mml:mi>Z</mml:mi><mml:mn mathvariant="normal">0</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M124" display="inline"><mml:mrow><mml:msub><mml:mi>Z</mml:mi><mml:mi mathvariant="normal">M</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>, the verification and (scaled) ensemble mean as a mean zero bivariate Gaussian random field with multivariate
Matérn correlation length parameters <inline-formula><mml:math id="M125" display="inline"><mml:mrow><mml:msub><mml:mi>a</mml:mi><mml:mn mathvariant="normal">0</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula>, <inline-formula><mml:math id="M126" display="inline"><mml:mrow><mml:msub><mml:mi>a</mml:mi><mml:mi mathvariant="normal">M</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>, and <inline-formula><mml:math id="M127" display="inline"><mml:mrow><mml:msub><mml:mi>a</mml:mi><mml:mrow><mml:mn mathvariant="normal">0</mml:mn><mml:mi mathvariant="normal">M</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msqrt><mml:mrow><mml:msub><mml:mi>a</mml:mi><mml:mn mathvariant="normal">0</mml:mn></mml:msub><mml:msub><mml:mi>a</mml:mi><mml:mi mathvariant="normal">M</mml:mi></mml:msub></mml:mrow></mml:msqrt></mml:mrow></mml:math></inline-formula>, smoothness parameters <inline-formula><mml:math id="M128" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">ν</mml:mi><mml:mn mathvariant="normal">0</mml:mn></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi mathvariant="italic">ν</mml:mi><mml:mrow><mml:mn mathvariant="normal">0</mml:mn><mml:mi mathvariant="normal">M</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi mathvariant="italic">ν</mml:mi><mml:mi mathvariant="normal">M</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1.5</mml:mn></mml:mrow></mml:math></inline-formula>, and co-located correlation coefficient <inline-formula><mml:math id="M129" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">ρ</mml:mi><mml:mrow><mml:mn mathvariant="normal">0</mml:mn><mml:mi mathvariant="normal">M</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi mathvariant="italic">ω</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.8</mml:mn></mml:mrow></mml:math></inline-formula>.</p></list-item><list-item>
      <p id="d1e2867">Generate 11 independent mean zero Gaussian random fields <inline-formula><mml:math id="M130" display="inline"><mml:mrow><mml:msub><mml:mi>W</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub><mml:mo>(</mml:mo><mml:mi>s</mml:mi><mml:mo>)</mml:mo><mml:mo>,</mml:mo><mml:mi mathvariant="normal">…</mml:mi><mml:mo>,</mml:mo><mml:msub><mml:mi>W</mml:mi><mml:mn mathvariant="normal">11</mml:mn></mml:msub><mml:mo>(</mml:mo><mml:mi>s</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> with Matérn covariance having correlation length
<inline-formula><mml:math id="M131" display="inline"><mml:mrow><mml:mi>a</mml:mi><mml:mo>=</mml:mo><mml:msub><mml:mi>a</mml:mi><mml:mi mathvariant="normal">M</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> and smoothness <inline-formula><mml:math id="M132" display="inline"><mml:mrow><mml:mi mathvariant="italic">ν</mml:mi><mml:mo>=</mml:mo><mml:msub><mml:mi mathvariant="italic">ν</mml:mi><mml:mi mathvariant="normal">M</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1.5</mml:mn></mml:mrow></mml:math></inline-formula>.</p></list-item><list-item>
      <p id="d1e2939">The ensemble member fields <inline-formula><mml:math id="M133" display="inline"><mml:mrow><mml:msub><mml:mi>Z</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub><mml:mo>(</mml:mo><mml:mi>s</mml:mi><mml:mo>)</mml:mo><mml:mo>,</mml:mo><mml:mi mathvariant="normal">…</mml:mi><mml:mo>,</mml:mo><mml:msub><mml:mi>Z</mml:mi><mml:mn mathvariant="normal">11</mml:mn></mml:msub><mml:mo>(</mml:mo><mml:mi>s</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> are constructed as<disp-formula id="Ch1.E8" content-type="numbered"><label>8</label><mml:math id="M134" display="block"><mml:mrow><mml:msub><mml:mi>Z</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>(</mml:mo><mml:mi>s</mml:mi><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:mi mathvariant="italic">ω</mml:mi><mml:msub><mml:mi>Z</mml:mi><mml:mi mathvariant="normal">M</mml:mi></mml:msub><mml:mo>(</mml:mo><mml:mi>s</mml:mi><mml:mo>)</mml:mo><mml:mo>+</mml:mo><mml:msqrt><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>-</mml:mo><mml:msup><mml:mi mathvariant="italic">ω</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:msqrt><mml:msub><mml:mi>W</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>(</mml:mo><mml:mi>s</mml:mi><mml:mo>)</mml:mo><mml:mo>,</mml:mo><mml:mspace width="1em" linebreak="nobreak"/><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>,</mml:mo><mml:mi mathvariant="normal">…</mml:mi><mml:mo>,</mml:mo><mml:mn mathvariant="normal">11</mml:mn><mml:mo>.</mml:mo></mml:mrow></mml:math></disp-formula></p></list-item></list>
The third step implies that each field in the ensemble is a Gaussian process with mean zero, variance one, correlation length <inline-formula><mml:math id="M135" display="inline"><mml:mrow><mml:msub><mml:mi>a</mml:mi><mml:mi mathvariant="normal">M</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>,
smoothness <inline-formula><mml:math id="M136" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">ν</mml:mi><mml:mi mathvariant="normal">M</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>, and univariate “forecast skill” controlled by the parameter <inline-formula><mml:math id="M137" display="inline"><mml:mi mathvariant="italic">ω</mml:mi></mml:math></inline-formula> (see Appendix B). Note that by choosing the
co-located correlation coefficient <inline-formula><mml:math id="M138" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">ρ</mml:mi><mml:mrow><mml:mn mathvariant="normal">0</mml:mn><mml:mi mathvariant="normal">M</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi mathvariant="italic">ω</mml:mi></mml:mrow></mml:math></inline-formula>, the correlation between the verification and each ensemble member is <inline-formula><mml:math id="M139" display="inline"><mml:mrow><mml:msup><mml:mi mathvariant="italic">ω</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula>,
the same as the correlation between ensemble members themselves. That is, <inline-formula><mml:math id="M140" display="inline"><mml:mrow><mml:mtext>Cov</mml:mtext><mml:mo>[</mml:mo><mml:msub><mml:mi>Z</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>Z</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>]</mml:mo><mml:mo>=</mml:mo><mml:msup><mml:mi mathvariant="italic">ω</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula> for <inline-formula><mml:math id="M141" display="inline"><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0</mml:mn><mml:mi mathvariant="normal">…</mml:mi><mml:mo>,</mml:mo><mml:mn mathvariant="normal">11</mml:mn></mml:mrow></mml:math></inline-formula> when <inline-formula><mml:math id="M142" display="inline"><mml:mrow><mml:mi>i</mml:mi><mml:mo>≠</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:math></inline-formula>
(derivation in Appendix C), and thus the ensemble forecasts are calibrated in the univariate sense.</p>
      <p id="d1e3177">In this study, fields were constructed on a square grid over the domain <inline-formula><mml:math id="M143" display="inline"><mml:mrow><mml:mo>[</mml:mo><mml:mo>-</mml:mo><mml:mn mathvariant="normal">20</mml:mn><mml:mo>,</mml:mo><mml:mn mathvariant="normal">20</mml:mn><mml:mo>]</mml:mo><mml:mo>×</mml:mo><mml:mo>[</mml:mo><mml:mo>-</mml:mo><mml:mn mathvariant="normal">20</mml:mn><mml:mo>,</mml:mo><mml:mn mathvariant="normal">20</mml:mn><mml:mo>]</mml:mo></mml:mrow></mml:math></inline-formula> with resolution 0.2. Verification-ensemble
samples were collected by repeating the simulation above 5000 times for each combination of

                <disp-formula specific-use="align"><mml:math id="M144" display="block"><mml:mtable displaystyle="true"><mml:mtr><mml:mtd><mml:mstyle displaystyle="true" class="stylechange"/></mml:mtd><mml:mtd><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><mml:msub><mml:mi>a</mml:mi><mml:mn mathvariant="normal">0</mml:mn></mml:msub><mml:mo>∈</mml:mo><mml:mo mathvariant="italic">{</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>,</mml:mo><mml:mn mathvariant="normal">1.5</mml:mn><mml:mo>,</mml:mo><mml:mi mathvariant="normal">…</mml:mi><mml:mo>,</mml:mo><mml:mn mathvariant="normal">3.5</mml:mn><mml:mo>,</mml:mo><mml:mn mathvariant="normal">4</mml:mn><mml:mo mathvariant="italic">}</mml:mo><mml:mo>,</mml:mo><mml:mspace width="1em" linebreak="nobreak"/></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mstyle displaystyle="true" class="stylechange"/></mml:mtd><mml:mtd><mml:mrow><mml:mstyle displaystyle="true" class="stylechange"/><mml:msub><mml:mi>a</mml:mi><mml:mi mathvariant="normal">M</mml:mi></mml:msub><mml:mo>∈</mml:mo><mml:mo mathvariant="italic">{</mml:mo><mml:mn mathvariant="normal">0.5</mml:mn><mml:msub><mml:mi>a</mml:mi><mml:mn mathvariant="normal">0</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:mn mathvariant="normal">0.6</mml:mn><mml:msub><mml:mi>a</mml:mi><mml:mn mathvariant="normal">0</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:mi mathvariant="normal">…</mml:mi><mml:mo>,</mml:mo><mml:mn mathvariant="normal">1.4</mml:mn><mml:msub><mml:mi>a</mml:mi><mml:mn mathvariant="normal">0</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:mn mathvariant="normal">1.5</mml:mn><mml:msub><mml:mi>a</mml:mi><mml:mn mathvariant="normal">0</mml:mn></mml:msub><mml:mo mathvariant="italic">}</mml:mo><mml:mo>,</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>

            resulting in a total of 77 experiments. Note that in practice, each sample corresponds to a date for which forecasts have been issued and verifying
observations are available, meaning the sample size is governed by the time period for which the verification is performed. For each experiment, FTE
histograms were constructed from the 5000 samples using each of <inline-formula><mml:math id="M145" display="inline"><mml:mrow><mml:mi mathvariant="italic">τ</mml:mi><mml:mo>∈</mml:mo><mml:mo mathvariant="italic">{</mml:mo><mml:mn mathvariant="normal">0</mml:mn><mml:mo>,</mml:mo><mml:mn mathvariant="normal">0.5</mml:mn><mml:mo>,</mml:mo><mml:mi mathvariant="normal">…</mml:mi><mml:mo>,</mml:mo><mml:mn mathvariant="normal">3.5</mml:mn><mml:mo>,</mml:mo><mml:mn mathvariant="normal">4</mml:mn><mml:mo mathvariant="italic">}</mml:mo></mml:mrow></mml:math></inline-formula>. That is, for a given <inline-formula><mml:math id="M146" display="inline"><mml:mrow><mml:msub><mml:mi>a</mml:mi><mml:mn mathvariant="normal">0</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M147" display="inline"><mml:mrow><mml:msub><mml:mi>a</mml:mi><mml:mi mathvariant="normal">M</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>
we analyzed nine FTE histograms, for a total of 693 histograms across all experiments.</p><?xmltex \hack{\newpage}?>
</sec>
<?pagebreak page416?><sec id="Ch1.S3.SS3">
  <label>3.3</label><title>Simulation analysis</title>
      <p id="d1e3381">The question of primary interest in this analysis is whether the FTE histogram accurately identifies miscalibration of ensemble correlation lengths.</p>
<sec id="Ch1.S3.SS3.SSS1">
  <label>3.3.1</label><title>Illustrative examples of FTE histograms</title>
      <p id="d1e3391">First, we study the discrimination ability of the FTE histogram in something of an exaggerated setting, where the miscalibration is obvious. We choose
the median of the marginal distribution as the threshold (i.e., <inline-formula><mml:math id="M148" display="inline"><mml:mrow><mml:mi mathvariant="italic">τ</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0</mml:mn></mml:mrow></mml:math></inline-formula>) and a verification correlation length of 2. On this grid, binary fields
produced in this way appear qualitatively similar to the binary precipitation fields analyzed later in this paper (see Fig. <xref ref-type="fig" rid="Ch1.F7"/>). The correlation length ratio is the ratio of the ensemble correlation length to that of the verification field. We study
ensembles with too small of a correlation length using ratio 0.5 (Fig. <xref ref-type="fig" rid="Ch1.F2"/>, row A), a correct correlation length using ratio 1.0 (Fig. <xref ref-type="fig" rid="Ch1.F2"/>, row B), and too large of a correlation length using ratio 1.5 (Fig. <xref ref-type="fig" rid="Ch1.F2"/>, row C). Corresponding FTE
histograms are then constructed with respect to these three ratios using 5000 verification-ensemble samples in each case. This revealing example is
depicted in Fig. <xref ref-type="fig" rid="Ch1.F2"/> and behaves as described in Table <xref ref-type="table" rid="Ch1.T1"/>, where the FTE histogram takes a
<inline-formula><mml:math id="M149" display="inline"><mml:mo>∪</mml:mo></mml:math></inline-formula> shape (<inline-formula><mml:math id="M150" display="inline"><mml:mo lspace="0mm">∩</mml:mo></mml:math></inline-formula> shape) when the ensemble correlation length is too small (large), indicating excessive (insufficient) spatial variability. As desired, the FTE histogram is approximately flat when the ensemble fields have the same correlation length as the verification field.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F2" specific-use="star"><?xmltex \currentcnt{2}?><label>Figure 2</label><caption><p id="d1e3435">Example binary exceedance verification field and a subset of ensemble fields with representative FTE histogram for threshold <inline-formula><mml:math id="M151" display="inline"><mml:mrow><mml:mi mathvariant="italic">τ</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0</mml:mn></mml:mrow></mml:math></inline-formula> found using 5000 samples. Dark blue regions indicate threshold exceedance. All verification fields have correlation length <inline-formula><mml:math id="M152" display="inline"><mml:mrow><mml:msub><mml:mi>a</mml:mi><mml:mn mathvariant="normal">0</mml:mn></mml:msub><mml:mo>=</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:mrow></mml:math></inline-formula> and ensemble fields have correlation length <inline-formula><mml:math id="M153" display="inline"><mml:mrow><mml:msub><mml:mi>a</mml:mi><mml:mi mathvariant="normal">M</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>,</mml:mo><mml:mn mathvariant="normal">2</mml:mn><mml:mo>,</mml:mo><mml:mn mathvariant="normal">3</mml:mn></mml:mrow></mml:math></inline-formula> in rows A, B, and C, respectively. FTE histograms are density histograms with dotted line <inline-formula><mml:math id="M154" display="inline"><mml:mrow><mml:mi>y</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:math></inline-formula> and corresponding <inline-formula><mml:math id="M155" display="inline"><mml:mi mathvariant="italic">β</mml:mi></mml:math></inline-formula>-score (left) and <inline-formula><mml:math id="M156" display="inline"><mml:mi mathvariant="italic">β</mml:mi></mml:math></inline-formula>-bias (right) annotated.</p></caption>
            <?xmltex \igopts{width=341.433071pt}?><graphic xlink:href="https://npg.copernicus.org/articles/27/411/2020/npg-27-411-2020-f02.png"/>

          </fig>

      <?xmltex \floatpos{t}?><fig id="Ch1.F3" specific-use="star"><?xmltex \currentcnt{3}?><label>Figure 3</label><caption><p id="d1e3523">As Fig. <xref ref-type="fig" rid="Ch1.F2"/> but ensemble fields have correlation length <inline-formula><mml:math id="M157" display="inline"><mml:mrow><mml:msub><mml:mi>a</mml:mi><mml:mi mathvariant="normal">M</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1.8</mml:mn><mml:mo>,</mml:mo><mml:mn mathvariant="normal">2</mml:mn><mml:mo>,</mml:mo><mml:mn mathvariant="normal">2.2</mml:mn></mml:mrow></mml:math></inline-formula> in rows A, B, and C, respectively.</p></caption>
            <?xmltex \igopts{width=341.433071pt}?><graphic xlink:href="https://npg.copernicus.org/articles/27/411/2020/npg-27-411-2020-f03.png"/>

          </fig>

      <p id="d1e3558">While the FTE histogram is able to correctly identify the obvious miscalibration of the ensemble for the scenario in Fig. <xref ref-type="fig" rid="Ch1.F2"/>, one
could likely draw the same conclusions by visual inspection and would not use the FTE histogram for these fields in practice. However, ensemble
forecast models are not generally so grossly miscalibrated; though a true correlation length ratio does not exist in reality, the theoretical ratio
will often be much closer to unity. Therefore, the true utility of the FTE histogram is realized when the miscalibration is not so visually
obvious. This more realistic example is illustrated in Fig. <xref ref-type="fig" rid="Ch1.F3"/>, where the above experiment is repeated using different correlation length ratios. In row A, the ensembles have ratio 0.9 and the resulting FTE histogram is still noticeably <inline-formula><mml:math id="M158" display="inline"><mml:mo>∪</mml:mo></mml:math></inline-formula>-shaped. The ratio in row B is 1.0, which yields a flat FTE histogram. In row C, the ratio is 1.1 and the FTE histogram is noticeably <inline-formula><mml:math id="M159" display="inline"><mml:mo>∩</mml:mo></mml:math></inline-formula>-shaped. Again, these results are consistent
with Table <xref ref-type="table" rid="Ch1.T1"/>, and we conclude that the FTE histogram maintains accurate discrimination ability even when ensemble
members are only slightly miscalibrated.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F4" specific-use="star"><?xmltex \currentcnt{4}?><label>Figure 4</label><caption><p id="d1e3583">As Fig. <xref ref-type="fig" rid="Ch1.F3"/> but with threshold <inline-formula><mml:math id="M160" display="inline"><mml:mrow><mml:mi mathvariant="italic">τ</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:mrow></mml:math></inline-formula>.</p></caption>
            <?xmltex \igopts{width=341.433071pt}?><graphic xlink:href="https://npg.copernicus.org/articles/27/411/2020/npg-27-411-2020-f04.png"/>

          </fig>

      <p id="d1e3606">Of course, one may often want to use a threshold parameter other than the median of the marginal distributions. The choice of <inline-formula><mml:math id="M161" display="inline"><mml:mi mathvariant="italic">τ</mml:mi></mml:math></inline-formula> is somewhat
application specific; for example, it can be chosen such as to focus on high precipitation amounts. Thus, it is important that the FTE histogram
maintains discrimination ability for different choices of <inline-formula><mml:math id="M162" display="inline"><mml:mi mathvariant="italic">τ</mml:mi></mml:math></inline-formula>. For a visual example, the same experiment depicted in Fig. <xref ref-type="fig" rid="Ch1.F3"/>
is repeated in Fig. <xref ref-type="fig" rid="Ch1.F4"/> but with FTE histograms constructed using <inline-formula><mml:math id="M163" display="inline"><mml:mrow><mml:mi mathvariant="italic">τ</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:mrow></mml:math></inline-formula> (equivalent to 2 SDs from the mean in this case). When the ensemble fields have a correlation length that is slightly too small (row A), the resulting FTE histogram is <inline-formula><mml:math id="M164" display="inline"><mml:mo>∪</mml:mo></mml:math></inline-formula>-shaped and
has a slight right skew due to the higher threshold but correctly indicates excessive spatial variability. When the ensemble exhibits insufficient spatial variability, i.e., correlation length is slightly too large (row C), the FTE histogram is <inline-formula><mml:math id="M165" display="inline"><mml:mo>∩</mml:mo></mml:math></inline-formula>-shaped and somewhat left-skewed. Reassuringly, the FTE histogram remains flat when the ensemble fields share the same correlation length as the verification fields (row B). While these results are
in agreement with Table <xref ref-type="table" rid="Ch1.T1"/>, the effect of the threshold can be studied more generally using the estimated <inline-formula><mml:math id="M166" display="inline"><mml:mi mathvariant="italic">β</mml:mi></mml:math></inline-formula>-score
and <inline-formula><mml:math id="M167" display="inline"><mml:mi mathvariant="italic">β</mml:mi></mml:math></inline-formula>-bias.</p>
</sec>
<sec id="Ch1.S3.SS3.SSS2">
  <label>3.3.2</label><title>Quantifying deviation from uniformity</title>
      <p id="d1e3678">Recall that we propose quantifying the shape of the FTE histogram with the <inline-formula><mml:math id="M168" display="inline"><mml:mi mathvariant="italic">β</mml:mi></mml:math></inline-formula>-score and <inline-formula><mml:math id="M169" display="inline"><mml:mi mathvariant="italic">β</mml:mi></mml:math></inline-formula>-bias. When <inline-formula><mml:math id="M170" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">β</mml:mi><mml:mi mathvariant="normal">S</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi mathvariant="italic">β</mml:mi><mml:mi mathvariant="normal">B</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0</mml:mn></mml:mrow></mml:math></inline-formula>, the FTE histogram is perfectly uniform. How do these metrics change as the threshold <inline-formula><mml:math id="M171" display="inline"><mml:mi mathvariant="italic">τ</mml:mi></mml:math></inline-formula> increases? Figure <xref ref-type="fig" rid="Ch1.F5"/> demonstrates how
the <inline-formula><mml:math id="M172" display="inline"><mml:mi mathvariant="italic">β</mml:mi></mml:math></inline-formula>-score and <inline-formula><mml:math id="M173" display="inline"><mml:mi mathvariant="italic">β</mml:mi></mml:math></inline-formula>-bias vary over increasing thresholds for different correlation length ratios; the estimated <inline-formula><mml:math id="M174" display="inline"><mml:mi mathvariant="italic">β</mml:mi></mml:math></inline-formula>-distribution
parameters are also depicted for comparison with Table <xref ref-type="table" rid="Ch1.T1"/>. Where provided, confidence intervals were estimated via the nonparametric bootstrap method (see <xref ref-type="bibr" rid="bib1.bibx8 bib1.bibx7" id="altparen.25"/>). When the correlation length ratio is 1.0, both the <inline-formula><mml:math id="M175" display="inline"><mml:mi mathvariant="italic">β</mml:mi></mml:math></inline-formula>-score
and <inline-formula><mml:math id="M176" display="inline"><mml:mi mathvariant="italic">β</mml:mi></mml:math></inline-formula>-bias are approximately zero for every choice of <inline-formula><mml:math id="M177" display="inline"><mml:mi mathvariant="italic">τ</mml:mi></mml:math></inline-formula>, correctly indicating a spatially calibrated ensemble. When the correlation length
ratio is less (greater) than 1.0, the <inline-formula><mml:math id="M178" display="inline"><mml:mi mathvariant="italic">β</mml:mi></mml:math></inline-formula>-scores are themselves generally less (greater) than zero, indicating excessive (insufficient) spatial
variability. As previously discussed, the <inline-formula><mml:math id="M179" display="inline"><mml:mi mathvariant="italic">β</mml:mi></mml:math></inline-formula>-bias becomes more pronounced at higher thresholds, thus highlighting the inextricable link between
threshold and skewness. Above a very high threshold of about 3 SDs (i.e., <inline-formula><mml:math id="M180" display="inline"><mml:mrow><mml:mi mathvariant="italic">τ</mml:mi><mml:mo>≈</mml:mo><mml:mn mathvariant="normal">3</mml:mn></mml:mrow></mml:math></inline-formula>), both metrics exhibit a tendency toward zero. This is partially due to the fact that the number of exceedances for high thresholds will often be zero, and since ranks are only discarded from the histogram if all ensemble members and the verification field have the same FTE, the FTE histogram for very high thresholds will be
composed largely of ranks resulting from ties broken uniformly at random. This results in a more deceptively uniform histogram which explains the
tendency toward zero. For the most extreme thresholds studied here, the confounding effect of resolving ties (which can exist between all but a single
ensemble member FTE) at random becomes very dominant, and histogram shapes get distorted to a degree where the interpretations provided in
Table <xref ref-type="table" rid="Ch1.T1"/> no longer hold. The associated FTE histograms still exhibit non-uniformity and thus indicate that the ensemble
forecasts are not perfectly calibrated, but it becomes impossible to diagnose the particular type of miscalibration from the histogram shape. We<?pagebreak page417?> can
also see that the sampling variability increases with increasing threshold since more and more uninformative cases with fully tied FTE values exist,
so a much larger total number of verification cases is required in order to have a comparable number of informative cases. In practice, if FTE
histograms are used as a diagnostic tool, we recommend focusing on moderate thresholds. If they are used to compare the calibration of different
forecast systems, they can still be effective at more extreme thresholds.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F5" specific-use="star"><?xmltex \currentcnt{5}?><label>Figure 5</label><caption><p id="d1e3805">Estimated beta distribution parameters (top) and corresponding <inline-formula><mml:math id="M181" display="inline"><mml:mi mathvariant="italic">β</mml:mi></mml:math></inline-formula>-score and <inline-formula><mml:math id="M182" display="inline"><mml:mi mathvariant="italic">β</mml:mi></mml:math></inline-formula>-bias (bottom) of FTE histograms calculated over different thresholds for forecasts with low, even, and high correlation length ratios about <inline-formula><mml:math id="M183" display="inline"><mml:mrow><mml:msub><mml:mi>a</mml:mi><mml:mn mathvariant="normal">0</mml:mn></mml:msub><mml:mo>=</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:mrow></mml:math></inline-formula>. Vertical lines denote the 95 % confidence interval found via the nonparametric bootstrap method.</p></caption>
            <?xmltex \igopts{width=369.885827pt}?><graphic xlink:href="https://npg.copernicus.org/articles/27/411/2020/npg-27-411-2020-f05.png"/>

          </fig>

      <?xmltex \floatpos{t}?><fig id="Ch1.F6" specific-use="star"><?xmltex \currentcnt{6}?><label>Figure 6</label><caption><p id="d1e3845">Estimated <inline-formula><mml:math id="M184" display="inline"><mml:mi mathvariant="italic">β</mml:mi></mml:math></inline-formula>-score and <inline-formula><mml:math id="M185" display="inline"><mml:mi mathvariant="italic">β</mml:mi></mml:math></inline-formula>-bias of FTE histograms constructed with <inline-formula><mml:math id="M186" display="inline"><mml:mrow><mml:mi mathvariant="italic">τ</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:math></inline-formula> for verification correlation length <inline-formula><mml:math id="M187" display="inline"><mml:mrow><mml:msub><mml:mi>a</mml:mi><mml:mn mathvariant="normal">0</mml:mn></mml:msub><mml:mo>∈</mml:mo><mml:mo mathvariant="italic">{</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>,</mml:mo><mml:mn mathvariant="normal">3</mml:mn><mml:mo mathvariant="italic">}</mml:mo></mml:mrow></mml:math></inline-formula> and ensemble range <inline-formula><mml:math id="M188" display="inline"><mml:mrow><mml:msub><mml:mi>a</mml:mi><mml:mi mathvariant="normal">M</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> varying from <inline-formula><mml:math id="M189" display="inline"><mml:mrow><mml:mn mathvariant="normal">0.5</mml:mn><mml:msub><mml:mi>a</mml:mi><mml:mn mathvariant="normal">0</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula> to <inline-formula><mml:math id="M190" display="inline"><mml:mrow><mml:mn mathvariant="normal">1.5</mml:mn><mml:msub><mml:mi>a</mml:mi><mml:mn mathvariant="normal">0</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula>.</p></caption>
            <?xmltex \igopts{width=312.980315pt}?><graphic xlink:href="https://npg.copernicus.org/articles/27/411/2020/npg-27-411-2020-f06.png"/>

          </fig>

      <p id="d1e3942">Another variable of interest in evaluating the FTE histograms is the size of the domain to which the metric is applied. In our simulation framework,
making the domain larger or smaller while keeping the correlation length constant is equivalent to keeping the domain size constant and varying the
correlation length of the verification field. That is, for a fixed domain size, a smaller correlation length mimics a “large domain” (with low
resolution) and a larger correlation length mimics a “small domain” (with high resolution). Analyzing the <inline-formula><mml:math id="M191" display="inline"><mml:mi mathvariant="italic">β</mml:mi></mml:math></inline-formula>-score and <inline-formula><mml:math id="M192" display="inline"><mml:mi mathvariant="italic">β</mml:mi></mml:math></inline-formula>-bias over a range of<?pagebreak page418?> correlation length ratios is then equivalent to studying the FTE histograms' utility for different domain sizes. For the domain used in this study,
a correlation length of 1 is considered small and 3 is considered large (see Fig. <xref ref-type="fig" rid="Ch1.F2"/>). In either case, Fig. <xref ref-type="fig" rid="Ch1.F6"/>
shows that the <inline-formula><mml:math id="M193" display="inline"><mml:mi mathvariant="italic">β</mml:mi></mml:math></inline-formula>-score quickly deviates from zero when the correlation length <italic>ratio</italic> is different from 1. The <inline-formula><mml:math id="M194" display="inline"><mml:mi mathvariant="italic">β</mml:mi></mml:math></inline-formula>-score's relatively steep slope around the correlation length of 1.0 in both cases indicates that the FTE histogram maintains good discrimination ability
regardless of domain size, provided that there are sufficiently many grid points within the domain to keep the (spatial) sampling variability
associated with the calculation of the FTE values low. Notably, the inverse relationship between <inline-formula><mml:math id="M195" display="inline"><mml:mi mathvariant="italic">β</mml:mi></mml:math></inline-formula>-bias and the correlation length ratio is also
in agreement with Table <xref ref-type="table" rid="Ch1.T1"/>.</p>
      <p id="d1e3990">We now turn attention back to our motivating figure (Fig. <xref ref-type="fig" rid="Ch1.F1"/>), which was created with verification correlation length
<inline-formula><mml:math id="M196" display="inline"><mml:mrow><mml:msub><mml:mi>a</mml:mi><mml:mn mathvariant="normal">0</mml:mn></mml:msub><mml:mo>=</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:mrow></mml:math></inline-formula>. Forecast 1 exhibited the correct spatial structure (i.e., <inline-formula><mml:math id="M197" display="inline"><mml:mrow><mml:msub><mml:mi>a</mml:mi><mml:mi mathvariant="normal">M</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:mrow></mml:math></inline-formula>), forecast 2 was incorrectly specified with correlation length
<inline-formula><mml:math id="M198" display="inline"><mml:mrow><mml:msub><mml:mi>a</mml:mi><mml:mi mathvariant="normal">M</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mn mathvariant="normal">3</mml:mn></mml:mrow></mml:math></inline-formula>, and forecast 3 was incorrectly specified with correlation length <inline-formula><mml:math id="M199" display="inline"><mml:mrow><mml:msub><mml:mi>a</mml:mi><mml:mi mathvariant="normal">M</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1.8</mml:mn></mml:mrow></mml:math></inline-formula>. While it may be obvious that forecast 2 has
incorrect spatial structure, the structural difference between forecasts 1 and 3 is not so apparent. However, as demonstrated by the analysis
above (Figs. <xref ref-type="fig" rid="Ch1.F2"/>–<xref ref-type="fig" rid="Ch1.F5"/>), these misspecifications are certainly identifiable using FTE histograms.</p>
</sec>
</sec>
</sec>
<?pagebreak page419?><sec id="Ch1.S4">
  <label>4</label><title>Application to downscaling of ensemble precipitation forecasts</title>
      <p id="d1e4070">Distributed hydrological models like NOAA's National Water Model (NWM) require meteorological inputs at a relatively high spatial resolution. At
shorter forecast lead times (typically up to 1 or 2 days ahead) limited-area NWP models provide such high-resolution forecasts, but for longer lead times only forecasts from global ensemble forecast systems like NOAA's GEFS are available. These come at a relatively coarse resolution and need
to be downscaled (statistically or dynamically) to the high-resolution output grid. Here, we use a combination of the statistical post-processing
algorithm proposed by <xref ref-type="bibr" rid="bib1.bibx27" id="text.26"/>, ensemble copula coupling (ECC; <xref ref-type="bibr" rid="bib1.bibx26" id="altparen.27"/>), and the spatial downscaling method
proposed by <xref ref-type="bibr" rid="bib1.bibx10" id="text.28"/> to obtain calibrated, high-resolution precipitation forecast fields based on GEFS ensemble forecasts. Does the
spatial disaggregation method produce precipitation fields with appropriate sub-grid-scale variability?  This question will be answered using the FTE-based verification metric discussed above.</p>
<sec id="Ch1.S4.SS1">
  <label>4.1</label><title>Data and downscaling methodology</title>
      <p id="d1e4089">We consider 6 h precipitation accumulations over a region in the southeastern US between <inline-formula><mml:math id="M200" display="inline"><mml:mo>-</mml:mo></mml:math></inline-formula>91 and <inline-formula><mml:math id="M201" display="inline"><mml:mo>-</mml:mo></mml:math></inline-formula>81<inline-formula><mml:math id="M202" display="inline"><mml:msup><mml:mi/><mml:mo>∘</mml:mo></mml:msup></mml:math></inline-formula> longitude and 30 and 40<inline-formula><mml:math id="M203" display="inline"><mml:msup><mml:mi/><mml:mo>∘</mml:mo></mml:msup></mml:math></inline-formula> latitude during the period from January 2002 to December 2016. Ensemble precipitation forecasts for lead time 66 to 72 h were obtained from NOAA's
second-generation GEFS reforecast dataset <xref ref-type="bibr" rid="bib1.bibx16" id="paren.29"/> at a horizontal resolution of <inline-formula><mml:math id="M204" display="inline"><mml:mo>∼</mml:mo></mml:math></inline-formula> 0.5<inline-formula><mml:math id="M205" display="inline"><mml:msup><mml:mi/><mml:mo>∘</mml:mo></mml:msup></mml:math></inline-formula>. Downscaling and verification is
performed against precipitation analyses from the <inline-formula><mml:math id="M206" display="inline"><mml:mo>∼</mml:mo></mml:math></inline-formula> 0.125<inline-formula><mml:math id="M207" display="inline"><mml:msup><mml:mi/><mml:mo>∘</mml:mo></mml:msup></mml:math></inline-formula> climatology-calibrated precipitation analysis (CCPA) dataset
<xref ref-type="bibr" rid="bib1.bibx17" id="paren.30"/>.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F7" specific-use="star"><?xmltex \currentcnt{7}?><label>Figure 7</label><caption><p id="d1e4165">Examples of different data fields for 6 h precipitation accumulation on 24 July 2004 <bold>(a–c)</bold> and corresponding 5 <inline-formula><mml:math id="M208" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">mm</mml:mi></mml:mrow></mml:math></inline-formula> binary exceedance fields (<bold>d–f</bold>; dark blue regions indicate threshold exceedance). From left to right: coarse-scale GEFS ensemble member, the same member downscaled to the analysis resolution, and the corresponding CCPA analysis.</p></caption>
          <?xmltex \igopts{width=369.885827pt}?><graphic xlink:href="https://npg.copernicus.org/articles/27/411/2020/npg-27-411-2020-f07.png"/>

        </fig>

      <p id="d1e4188">In order to obtain calibrated ensemble precipitation forecasts at the CCPA grid resolution, we proceed in three steps. First, we apply the post-processing algorithm by <xref ref-type="bibr" rid="bib1.bibx27" id="text.31"/> to the GEFS forecasts and upscaled (to the GEFS grid resolution) precipitation analyses in
order to remove systematic biases and ensure adequate representation of forecast uncertainty at this coarse grid scale. The resulting predictive
distributions are turned back into an 11-member ensemble using the ECC-mQ-SNP variation <xref ref-type="bibr" rid="bib1.bibx28" id="paren.32"/> of the ECC technique. This
variation removes discontinuities and avoids randomization that can occur when the standard ECC approach is applied to precipitation fields. Finally,
each ensemble member is downscaled from the GEFS to the CCPA grid resolution using a slightly simplified version of the Gibbs sampling disaggregation model (GSDM) proposed by <xref ref-type="bibr" rid="bib1.bibx10" id="text.33"/>. To generate downscaled fields with spatial properties that vary depending on the season, we rely here
on a monthly calibration of the GSDM, rather than on meteorological predictors as in the original model. The 15 years of data are cross-validated: 1 year at a time is left out for verification and the post-processing and downscaling models are fitted with data from the remaining 14 years. Repeating
this process for all years leaves us with 15 years of downscaled ensemble forecasts and verifying analyses. See Fig. <xref ref-type="fig" rid="Ch1.F7"/> for a
visual reference.</p>
</sec>
<sec id="Ch1.S4.SS2">
  <label>4.2</label><title>Univariate verification</title>
      <p id="d1e4210">Before applying the FTE histogram to investigate whether the spatial disaggregation used in the downscaling method produces precipitation fields with
appropriate sub-grid-scale variability, we check the calibration of the univariate ensemble forecasts across all fine-scale grid points. We study (separately) the months January, April, July, and October in order to represent winter, spring, summer, and fall, respectively. Daily analyses and
corresponding ensemble forecasts from each of these months are pooled over the entire verification period and all grid points within the study area and are used to construct the verification rank histograms in Fig. <xref ref-type="fig" rid="Ch1.F8"/>. Cases where all ensemble member forecasts and the analysis
are tied – for example, when there is zero accumulation at a grid point for all fields – are withheld from the histogram to avoid<?pagebreak page420?> artificial
uniformity introduced by breaking ties in rank at random.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F8" specific-use="star"><?xmltex \currentcnt{8}?><label>Figure 8</label><caption><p id="d1e4217">Verification rank histograms (density) for downscaled fields at representative months with cases of fully tied ranks removed. Estimated <inline-formula><mml:math id="M209" display="inline"><mml:mi mathvariant="italic">β</mml:mi></mml:math></inline-formula>-score (left) and <inline-formula><mml:math id="M210" display="inline"><mml:mi mathvariant="italic">β</mml:mi></mml:math></inline-formula>-bias (right) annotated.</p></caption>
          <?xmltex \igopts{width=341.433071pt}?><graphic xlink:href="https://npg.copernicus.org/articles/27/411/2020/npg-27-411-2020-f08.png"/>

        </fig>

      <p id="d1e4240">Ideally, the statistical post-processing and downscaling should yield calibrated ensemble forecasts and thus rank histograms that are approximately
uniform. Clearly, the rank histograms for the downscaled forecast fields shown in Fig. <xref ref-type="fig" rid="Ch1.F8"/> are not uniform; there is a consistent
peak in the higher ranks indicating that the downscaled ensemble forecasts tend to underestimate precipitation accumulations, especially in fall and
winter. This bias could be an indication that either the post-processing distribution (gamma) or the disaggregation distribution (log-normal) is not perfectly suited for representing the respective forecast uncertainties. It may also be a result of a superposition of biases in different sub-domains or for different weather situations. Univariate calibration in July – which happens to be a month with more frequent precipitation in this region of the
US – is relatively good, and while the histograms of other months show clear departures from uniformity, there is at least no strong <inline-formula><mml:math id="M211" display="inline"><mml:mo>∪</mml:mo></mml:math></inline-formula> or <inline-formula><mml:math id="M212" display="inline"><mml:mo>∩</mml:mo></mml:math></inline-formula> shape to indicate significant dispersion errors. We thus continue with our analysis of the spatial calibration of the downscaled ensemble
forecast fields, keeping in mind though that the under-forecast biases seen in Fig. <xref ref-type="fig" rid="Ch1.F8"/> will carry over to the FTE histograms and will superimpose any shape resulting from spatial miscalibration. In July, where only a mild under-forecast bias is observed, we will have the best chance of drawing unambiguous conclusions about the spatial structure from the shape of the FTE histogram.</p>
</sec>
<sec id="Ch1.S4.SS3">
  <label>4.3</label><title>Verification of spatial structure</title>
      <p id="d1e4269">In the remaining analysis, we employ FTE histograms to investigate the spatial properties of the ensemble forecast fields obtained by the downscaling
algorithm for the same representative months outlined above. Spatial variability of precipitation fields depends on whether precipitation is
stratiform or convective, and in the latter case also on the type of convection (local vs. synoptically forced). The frequency of occurrence of these categories has a seasonal cycle, and it is therefore interesting to study how well the downscaling methodology works in different seasons. The first
step in computing the FTE is deciding what value to use for the threshold. If the climatology varies strongly across the domain, it may be desirable
to use a variable threshold such as a climatology percentile. However, the southeastern US is a flat and relatively homogeneous region, meaning the precipitation accumulation patterns will not be affected as much by orography, and we therefore select a fixed threshold for constructing FTE
histograms. Another advantage of this approach is that a fixed threshold has a direct physical interpretation; here we use thresholds of 5, 10, and
20 <inline-formula><mml:math id="M213" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">mm</mml:mi></mml:mrow></mml:math></inline-formula> to study the spatial calibration of the ensemble for low, medium, and high accumulation levels over the 6 h window.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F9" specific-use="star"><?xmltex \currentcnt{9}?><label>Figure 9</label><caption><p id="d1e4282">FTE histograms for downscaled fields at different thresholds in representative months. Estimated <inline-formula><mml:math id="M214" display="inline"><mml:mi mathvariant="italic">β</mml:mi></mml:math></inline-formula>-score (left) and <inline-formula><mml:math id="M215" display="inline"><mml:mi mathvariant="italic">β</mml:mi></mml:math></inline-formula>-bias (right) annotated.</p></caption>
          <?xmltex \igopts{width=341.433071pt}?><graphic xlink:href="https://npg.copernicus.org/articles/27/411/2020/npg-27-411-2020-f09.png"/>

        </fig>

      <p id="d1e4305">In Fig. <xref ref-type="fig" rid="Ch1.F9"/>, it is clear by visual inspection that the FTE histograms are all <inline-formula><mml:math id="M216" display="inline"><mml:mo>∪</mml:mo></mml:math></inline-formula>-shaped to some extent, though the
corresponding <inline-formula><mml:math id="M217" display="inline"><mml:mi mathvariant="italic">β</mml:mi></mml:math></inline-formula>-scores highlight that the histograms are explicitly more uniform in the fall and winter months. In the spring and summer months
(i.e., April and July) the histograms reveal a clear under-dispersion in the ensemble FTEs at all<?pagebreak page421?> analyzed thresholds. This would suggest that the downscaled ensemble overestimates fine-scale variability during the seasons with more convective events. This could indicate that the calibration
procedure of the GSDM downscaling method in <xref ref-type="bibr" rid="bib1.bibx10" id="text.34"/> struggles with selecting good parameters that produce downscaled precipitation fields
with just the right amount of spatial variability during the summer season with mainly (but not exclusively) convective precipitation. The FTE
histogram can thus provide valuable diagnostic information that helps identify shortcomings of a forecast methodology. Indeed, in one of our current
projects we seek to improve the GSDM, with one objective being to calibrate the model such that the downscaled fields reproduce the correct amount of
spatial variability, in a flow-dependent fashion, using meteorological predictors such as instability indices and vertical wind shear
<xref ref-type="bibr" rid="bib1.bibx3" id="paren.35"/>. For the <inline-formula><mml:math id="M218" display="inline"><mml:mi mathvariant="italic">β</mml:mi></mml:math></inline-formula>-biases seen in Fig. <xref ref-type="fig" rid="Ch1.F9"/>, the interpretation is more difficult. Their value at higher
thresholds has the opposite sign to what we would expect from Table <xref ref-type="table" rid="Ch1.T1"/> in a situation where the forecast fields simulate excessive fine-scale variability. As noted above, however, this is likely due to an under-forecast bias in the marginal distributions and the
associated effect on the <inline-formula><mml:math id="M219" display="inline"><mml:mi mathvariant="italic">β</mml:mi></mml:math></inline-formula>-bias which is opposite to (and seems to be dominating) the effects of spatial miscalibration.</p>
</sec>
</sec>
<sec id="Ch1.S5" sec-type="conclusions">
  <label>5</label><title>Conclusions</title>
      <p id="d1e4359">When forecasting meteorological variables on a spatial domain, it is important for many applications that not only the marginal forecast distributions, but also the spatial (and/or temporal) correlation structure is represented adequately. In some instances, misrepresentation of spatial structure by
ensemble forecast fields may be visually obvious; otherwise, a quantitative verification metric is desired to objectively<?pagebreak page422?> evaluate the ensemble
calibration. The FTE metric studied here is a projection of a multivariate quantity (i.e., a spatial field) to a univariate quantity and can be combined with the concept of a (univariate) verification rank histogram to analyze the spatial structure of ensemble forecast fields. This idea was first applied by <xref ref-type="bibr" rid="bib1.bibx28" id="text.36"/> to study the properties of downscaled ensemble precipitation forecasts, but an understanding of
the general capability of the FTE metric to detect misrepresentation of the spatial structure by the ensemble has been lacking as yet.</p>
      <p id="d1e4365">In this paper, we performed a systematic study in which we simulated ensemble forecast and verification fields with different correlation lengths to
understand how well a misspecification of the correlation length can be detected by the FTE metric. To this end, the metric was slightly extended and
is composed of three steps: (1) calculate the FTE of each verification and ensemble forecast field, (2) construct an FTE histogram over available
instances of forecast and verification times, and (3) derive the <inline-formula><mml:math id="M220" display="inline"><mml:mi mathvariant="italic">β</mml:mi></mml:math></inline-formula>-score and <inline-formula><mml:math id="M221" display="inline"><mml:mi mathvariant="italic">β</mml:mi></mml:math></inline-formula>-bias from the stochastically disaggregated FTE histogram to
characterize departure from uniformity. We have found that the FTE metric is capable of detecting even minor issues with the correlation length (e.g.,
10 % miscalibration) in ensemble forecasts, and this conclusion was consistent across a range of thresholds and domain sizes. Applied in a data
example with downscaled precipitation forecast fields, the FTE metric pointed to some shortcomings of the underlying spatial disaggregation algorithm
during the seasons where precipitation is driven by local convection.</p>
      <p id="d1e4382"><?xmltex \hack{\newpage}?>The FTE metric is relatively simple and enjoys an easy and intuitive interpretation. In particular, the <inline-formula><mml:math id="M222" display="inline"><mml:mi mathvariant="italic">β</mml:mi></mml:math></inline-formula>-score and <inline-formula><mml:math id="M223" display="inline"><mml:mi mathvariant="italic">β</mml:mi></mml:math></inline-formula>-bias can be compared
according to Table <xref ref-type="table" rid="Ch1.T1"/> to diagnose shortcomings in the calibration of ensemble forecasts. If different types of
miscalibration occur together, additional diagnostic tools like univariate verification rank histograms have to be considered alongside the FTE
histograms to disentangle the different effects. While we have focused on histograms in the analysis of the verification FTE rank here, the same
projection could also be used in combination with proper scoring rules. We believe that FTE histograms are a useful addition to the set of spatial
verification metrics. They complement metrics like the wavelet-based verification approach proposed by <xref ref-type="bibr" rid="bib1.bibx6" id="text.37"/> which has
additional capabilities when it comes to analyzing aspects of the spatial texture of forecast fields but is not primarily targeted at proper
uncertainty quantification by an ensemble.</p><?xmltex \hack{\clearpage}?>
</sec>

      
      </body>
    <back><app-group>

<?pagebreak page423?><app id="App1.Ch1.S1">
  <?xmltex \currentcnt{A}?><label>Appendix A</label><title>Stochastic disaggregation of transformed ranks</title>
      <p id="d1e4417">The data vector of ranks <inline-formula><mml:math id="M224" display="inline"><mml:mi mathvariant="bold-italic">r</mml:mi></mml:math></inline-formula>
has discrete elements <inline-formula><mml:math id="M225" display="inline"><mml:mrow><mml:msub><mml:mi>r</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>∈</mml:mo><mml:mo mathvariant="italic">{</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>,</mml:mo><mml:mn mathvariant="normal">2</mml:mn><mml:mo>,</mml:mo><mml:mi mathvariant="normal">…</mml:mi><mml:mo>,</mml:mo><mml:mi>k</mml:mi><mml:mo>+</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo mathvariant="italic">}</mml:mo></mml:mrow></mml:math></inline-formula>, where <inline-formula><mml:math id="M226" display="inline"><mml:mi>k</mml:mi></mml:math></inline-formula> is the number of ensemble members. In order to disaggregate these elements to a continuous domain for use with maximum likelihood estimation, the following algorithm is applied to each element <inline-formula><mml:math id="M227" display="inline"><mml:mrow><mml:msub><mml:mi>r</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>.
<list list-type="order"><list-item>
      <p id="d1e4483">Let <inline-formula><mml:math id="M228" display="inline"><mml:mrow><mml:msub><mml:mi>d</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mo>(</mml:mo><mml:msub><mml:mi>r</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>-</mml:mo><mml:mstyle displaystyle="false"><mml:mfrac style="text"><mml:mn mathvariant="normal">1</mml:mn><mml:mn mathvariant="normal">2</mml:mn></mml:mfrac></mml:mstyle><mml:mo>)</mml:mo><mml:mo>/</mml:mo><mml:mo>(</mml:mo><mml:mi>k</mml:mi><mml:mo>+</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>.</p></list-item><list-item>
      <p id="d1e4530">Simulate a (continuous) uniform random variable:
<inline-formula><mml:math id="M229" display="inline"><mml:mrow><mml:msub><mml:mi>U</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>∼</mml:mo><mml:mtext>Uniform</mml:mtext><mml:mfenced close=")" open="("><mml:mrow><mml:msub><mml:mi>d</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>-</mml:mo><mml:mstyle displaystyle="false"><mml:mfrac style="text"><mml:mn mathvariant="normal">1</mml:mn><mml:mrow><mml:mn mathvariant="normal">2</mml:mn><mml:mo>(</mml:mo><mml:mi>k</mml:mi><mml:mo>+</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>)</mml:mo></mml:mrow></mml:mfrac></mml:mstyle><mml:mo>,</mml:mo><mml:msub><mml:mi>d</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>+</mml:mo><mml:mstyle displaystyle="false"><mml:mfrac style="text"><mml:mn mathvariant="normal">1</mml:mn><mml:mrow><mml:mn mathvariant="normal">2</mml:mn><mml:mo>(</mml:mo><mml:mi>k</mml:mi><mml:mo>+</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>)</mml:mo></mml:mrow></mml:mfrac></mml:mstyle></mml:mrow></mml:mfenced></mml:mrow></mml:math></inline-formula>.</p></list-item><list-item>
      <p id="d1e4601">Set <inline-formula><mml:math id="M230" display="inline"><mml:mrow><mml:msub><mml:mi>r</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>U</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>.</p></list-item></list>
The effect of Step 1 is a mapping into <inline-formula><mml:math id="M231" display="inline"><mml:mrow><mml:mo>[</mml:mo><mml:mn mathvariant="normal">0</mml:mn><mml:mo>,</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>]</mml:mo></mml:mrow></mml:math></inline-formula>, while Step 2 is the stochastic disaggregation to evenly spaced uniform intervals whose supports form a partition of unity of [0,1].</p>
</app>

<app id="App1.Ch1.S2">
  <?xmltex \currentcnt{B}?><label>Appendix B</label><title>Properties of simulated ensemble members</title>
      <p id="d1e4648">Let <inline-formula><mml:math id="M232" display="inline"><mml:mrow><mml:msub><mml:mi>Z</mml:mi><mml:mi mathvariant="normal">M</mml:mi></mml:msub><mml:mo>(</mml:mo><mml:mi>s</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M233" display="inline"><mml:mrow><mml:msub><mml:mi>W</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>(</mml:mo><mml:mi>s</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> be independent, mean-zero Gaussian processes, each with Matérn covariance function <inline-formula><mml:math id="M234" display="inline"><mml:mrow><mml:mi>M</mml:mi><mml:mo>(</mml:mo><mml:mi>d</mml:mi><mml:mo>|</mml:mo><mml:msub><mml:mi mathvariant="italic">ν</mml:mi><mml:mi mathvariant="normal">M</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>a</mml:mi><mml:mi mathvariant="normal">M</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>. Now suppose <inline-formula><mml:math id="M235" display="inline"><mml:mrow><mml:msub><mml:mi>Z</mml:mi><mml:mi mathvariant="normal">M</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M236" display="inline"><mml:mrow><mml:msub><mml:mi>W</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> are independent standard Gaussian random variables representing the marginal distribution
of processes <inline-formula><mml:math id="M237" display="inline"><mml:mrow><mml:msub><mml:mi>Z</mml:mi><mml:mi mathvariant="normal">M</mml:mi></mml:msub><mml:mo>(</mml:mo><mml:mi>s</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M238" display="inline"><mml:mrow><mml:msub><mml:mi>W</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>(</mml:mo><mml:mi>s</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>. Setting random variable
          <disp-formula id="App1.Ch1.S2.E9" content-type="numbered"><label>B1</label><mml:math id="M239" display="block"><mml:mrow><mml:msub><mml:mi>Z</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mi mathvariant="italic">ω</mml:mi><mml:msub><mml:mi>Z</mml:mi><mml:mi mathvariant="normal">M</mml:mi></mml:msub><mml:mo>+</mml:mo><mml:msqrt><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>-</mml:mo><mml:msup><mml:mi mathvariant="italic">ω</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:msqrt><mml:msub><mml:mi>W</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:mspace linebreak="nobreak" width="1em"/><mml:mi mathvariant="italic">ω</mml:mi><mml:mo>∈</mml:mo><mml:mo>[</mml:mo><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>,</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>]</mml:mo><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>
        we see <inline-formula><mml:math id="M240" display="inline"><mml:mrow><mml:mi mathvariant="double-struck">E</mml:mi><mml:mo>[</mml:mo><mml:msub><mml:mi>Z</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>]</mml:mo><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0</mml:mn></mml:mrow></mml:math></inline-formula> by linearity of the expectation operator and

              <disp-formula specific-use="align" content-type="numbered"><mml:math id="M241" display="block"><mml:mtable displaystyle="true"><mml:mtr><mml:mtd><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><mml:mtext>Var</mml:mtext><mml:mo>[</mml:mo><mml:msub><mml:mi>Z</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>]</mml:mo></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:mstyle displaystyle="true" class="stylechange"/><mml:mo>=</mml:mo><mml:mtext>Var</mml:mtext><mml:mo>[</mml:mo><mml:mi mathvariant="italic">ω</mml:mi><mml:msub><mml:mi>Z</mml:mi><mml:mi mathvariant="normal">M</mml:mi></mml:msub><mml:mo>+</mml:mo><mml:msqrt><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>-</mml:mo><mml:msup><mml:mi mathvariant="italic">ω</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:msqrt><mml:msub><mml:mi>W</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>]</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mstyle displaystyle="true" class="stylechange"/></mml:mtd><mml:mtd><mml:mrow><mml:mstyle displaystyle="true" class="stylechange"/><mml:mo>=</mml:mo><mml:msup><mml:mi mathvariant="italic">ω</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mtext>Var</mml:mtext><mml:mo>[</mml:mo><mml:msub><mml:mi>Z</mml:mi><mml:mi mathvariant="normal">M</mml:mi></mml:msub><mml:mo>]</mml:mo><mml:mo>+</mml:mo><mml:mo>(</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>-</mml:mo><mml:msup><mml:mi mathvariant="italic">ω</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo>)</mml:mo><mml:mtext>Var</mml:mtext><mml:mo>[</mml:mo><mml:msub><mml:mi>W</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>]</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mlabeledtr id="App1.Ch1.S2.E10"><mml:mtd><mml:mtext>B2</mml:mtext></mml:mtd><mml:mtd><mml:mstyle displaystyle="true" class="stylechange"/></mml:mtd><mml:mtd><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:mtd></mml:mlabeledtr></mml:mtable></mml:math></disp-formula>

          using independence of <inline-formula><mml:math id="M242" display="inline"><mml:mrow><mml:msub><mml:mi>Z</mml:mi><mml:mi mathvariant="normal">M</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M243" display="inline"><mml:mrow><mml:msub><mml:mi>W</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>. Then <inline-formula><mml:math id="M244" display="inline"><mml:mrow><mml:msub><mml:mi>Z</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> is a standard Gaussian random variable representing the marginal distribution of ensemble
member <inline-formula><mml:math id="M245" display="inline"><mml:mrow><mml:msub><mml:mi>Z</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>(</mml:mo><mml:mi>s</mml:mi><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:mi mathvariant="italic">ω</mml:mi><mml:msub><mml:mi>Z</mml:mi><mml:mi mathvariant="normal">M</mml:mi></mml:msub><mml:mo>(</mml:mo><mml:mi>s</mml:mi><mml:mo>)</mml:mo><mml:mo>+</mml:mo><mml:msqrt><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>-</mml:mo><mml:msup><mml:mi mathvariant="italic">ω</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:msqrt><mml:msub><mml:mi>W</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>(</mml:mo><mml:mi>s</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>. Further, observe that

              <disp-formula specific-use="align" content-type="numbered"><mml:math id="M246" display="block"><mml:mtable displaystyle="true"><mml:mtr><mml:mtd><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><mml:mtext>Cov</mml:mtext></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><mml:mo>[</mml:mo><mml:msub><mml:mi>Z</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>(</mml:mo><mml:mi>s</mml:mi><mml:mo>+</mml:mo><mml:mi>h</mml:mi><mml:mo>)</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mi>Z</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>(</mml:mo><mml:mi>s</mml:mi><mml:mo>)</mml:mo><mml:mo>]</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mstyle class="stylechange" displaystyle="true"/></mml:mtd><mml:mtd><mml:mrow><mml:mstyle displaystyle="true" class="stylechange"/><mml:mo>=</mml:mo><mml:mtext>Cov</mml:mtext><mml:mo>[</mml:mo><mml:mi mathvariant="italic">ω</mml:mi><mml:msub><mml:mi>Z</mml:mi><mml:mi mathvariant="normal">M</mml:mi></mml:msub><mml:mo>(</mml:mo><mml:mi>s</mml:mi><mml:mo>+</mml:mo><mml:mi>h</mml:mi><mml:mo>)</mml:mo><mml:mo>+</mml:mo><mml:msqrt><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>-</mml:mo><mml:msup><mml:mi mathvariant="italic">ω</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:msqrt><mml:msub><mml:mi>W</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>(</mml:mo><mml:mi>s</mml:mi><mml:mo>+</mml:mo><mml:mi>h</mml:mi><mml:mo>)</mml:mo><mml:mo>,</mml:mo><mml:mi mathvariant="italic">ω</mml:mi><mml:msub><mml:mi>Z</mml:mi><mml:mi mathvariant="normal">M</mml:mi></mml:msub><mml:mo>(</mml:mo><mml:mi>s</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mstyle class="stylechange" displaystyle="true"/></mml:mtd><mml:mtd><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><?xmltex \hack{\protect\hphantom{123}}?><mml:mo>+</mml:mo><mml:msqrt><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>-</mml:mo><mml:msup><mml:mi mathvariant="italic">ω</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:msqrt><mml:msub><mml:mi>W</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>(</mml:mo><mml:mi>s</mml:mi><mml:mo>)</mml:mo><mml:mo>]</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mstyle class="stylechange" displaystyle="true"/></mml:mtd><mml:mtd><mml:mrow><mml:mstyle displaystyle="true" class="stylechange"/><mml:mo>=</mml:mo><mml:msup><mml:mi mathvariant="italic">ω</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mtext>Cov</mml:mtext><mml:mo>[</mml:mo><mml:msub><mml:mi>Z</mml:mi><mml:mi mathvariant="normal">M</mml:mi></mml:msub><mml:mo>(</mml:mo><mml:mi>s</mml:mi><mml:mo>+</mml:mo><mml:mi>h</mml:mi><mml:mo>)</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mi>Z</mml:mi><mml:mi mathvariant="normal">M</mml:mi></mml:msub><mml:mo>(</mml:mo><mml:mi>s</mml:mi><mml:mo>)</mml:mo><mml:mo>]</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mstyle class="stylechange" displaystyle="true"/></mml:mtd><mml:mtd><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><?xmltex \hack{\protect\hphantom{123}}?><mml:mo>+</mml:mo><mml:mo>(</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>-</mml:mo><mml:msup><mml:mi mathvariant="italic">ω</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo>)</mml:mo><mml:mtext>Cov</mml:mtext><mml:mo>[</mml:mo><mml:msub><mml:mi>W</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>(</mml:mo><mml:mi>s</mml:mi><mml:mo>+</mml:mo><mml:mi>h</mml:mi><mml:mo>)</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mi>W</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>(</mml:mo><mml:mi>s</mml:mi><mml:mo>)</mml:mo><mml:mo>]</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mstyle class="stylechange" displaystyle="true"/></mml:mtd><mml:mtd><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><mml:mo>=</mml:mo><mml:msup><mml:mi mathvariant="italic">ω</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mi>M</mml:mi><mml:mo>(</mml:mo><mml:mo>|</mml:mo><mml:mo>|</mml:mo><mml:mi>h</mml:mi><mml:mo>|</mml:mo><mml:mo>|</mml:mo><mml:mo>|</mml:mo><mml:msub><mml:mi mathvariant="italic">ν</mml:mi><mml:mi mathvariant="normal">M</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>a</mml:mi><mml:mi mathvariant="normal">M</mml:mi></mml:msub><mml:mo>)</mml:mo><mml:mo>+</mml:mo><mml:mo>(</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>-</mml:mo><mml:msup><mml:mi mathvariant="italic">ω</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo>)</mml:mo><mml:mi>M</mml:mi><mml:mo>(</mml:mo><mml:mo>|</mml:mo><mml:mo>|</mml:mo><mml:mi>h</mml:mi><mml:mo>|</mml:mo><mml:mo>|</mml:mo><mml:mo>|</mml:mo><mml:msub><mml:mi mathvariant="italic">ν</mml:mi><mml:mi mathvariant="normal">M</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>a</mml:mi><mml:mi mathvariant="normal">M</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mlabeledtr id="App1.Ch1.S2.E11"><mml:mtd><mml:mtext>B3</mml:mtext></mml:mtd><mml:mtd><mml:mstyle class="stylechange" displaystyle="true"/></mml:mtd><mml:mtd><mml:mrow><mml:mstyle displaystyle="true" class="stylechange"/><mml:mo>=</mml:mo><mml:mi>M</mml:mi><mml:mo>(</mml:mo><mml:mo>|</mml:mo><mml:mo>|</mml:mo><mml:mi>h</mml:mi><mml:mo>|</mml:mo><mml:mo>|</mml:mo><mml:mo>|</mml:mo><mml:msub><mml:mi mathvariant="italic">ν</mml:mi><mml:mi mathvariant="normal">M</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>a</mml:mi><mml:mi mathvariant="normal">M</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mlabeledtr></mml:mtable></mml:math></disp-formula>

          by independence of <inline-formula><mml:math id="M247" display="inline"><mml:mrow><mml:msub><mml:mi>Z</mml:mi><mml:mi mathvariant="normal">M</mml:mi></mml:msub><mml:mo>(</mml:mo><mml:mi>s</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M248" display="inline"><mml:mrow><mml:msub><mml:mi>W</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>(</mml:mo><mml:mi>s</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>. That is, ensemble members <inline-formula><mml:math id="M249" display="inline"><mml:mrow><mml:msub><mml:mi>Z</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>(</mml:mo><mml:mi>s</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>, <inline-formula><mml:math id="M250" display="inline"><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>,</mml:mo><mml:mi mathvariant="normal">…</mml:mi><mml:mo>,</mml:mo><mml:mi>k</mml:mi></mml:mrow></mml:math></inline-formula> preserve the covariance structure of the  ensemble mean <inline-formula><mml:math id="M251" display="inline"><mml:mrow><mml:msub><mml:mi>Z</mml:mi><mml:mi mathvariant="normal">M</mml:mi></mml:msub><mml:mo>(</mml:mo><mml:mi>s</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>.</p><?xmltex \hack{\newpage}?>
</app>

<app id="App1.Ch1.S3">
  <?xmltex \currentcnt{C}?><label>Appendix C</label><title>Derivation of an appropriate co-located correlation coefficient</title>
      <p id="d1e5553">Suppose <inline-formula><mml:math id="M252" display="inline"><mml:mrow><mml:msub><mml:mi>Z</mml:mi><mml:mn mathvariant="normal">0</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M253" display="inline"><mml:mrow><mml:msub><mml:mi>Z</mml:mi><mml:mi mathvariant="normal">M</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> are standard Gaussian random variables with <inline-formula><mml:math id="M254" display="inline"><mml:mrow><mml:mtext>Corr</mml:mtext><mml:mo>[</mml:mo><mml:msub><mml:mi>Z</mml:mi><mml:mn mathvariant="normal">0</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>Z</mml:mi><mml:mi mathvariant="normal">M</mml:mi></mml:msub><mml:mo>]</mml:mo><mml:mo>=</mml:mo><mml:mi mathvariant="italic">ρ</mml:mi></mml:mrow></mml:math></inline-formula>, and <inline-formula><mml:math id="M255" display="inline"><mml:mrow><mml:mo mathvariant="italic">{</mml:mo><mml:msub><mml:mi>W</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:msubsup><mml:mo mathvariant="italic">}</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mi>k</mml:mi></mml:msubsup></mml:mrow></mml:math></inline-formula> is a set
of independent standard Gaussian random variables, each independent of <inline-formula><mml:math id="M256" display="inline"><mml:mrow><mml:msub><mml:mi>Z</mml:mi><mml:mn mathvariant="normal">0</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M257" display="inline"><mml:mrow><mml:msub><mml:mi>Z</mml:mi><mml:mi mathvariant="normal">M</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>. Define
          <disp-formula id="App1.Ch1.S3.E12" content-type="numbered"><label>C1</label><mml:math id="M258" display="block"><mml:mtable rowspacing="0.2ex" class="split" columnspacing="1em" displaystyle="true" columnalign="right left"><mml:mtr><mml:mtd><mml:mrow><mml:msub><mml:mi>Z</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>=</mml:mo></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:mi mathvariant="italic">ω</mml:mi><mml:msub><mml:mi>Z</mml:mi><mml:mi mathvariant="normal">M</mml:mi></mml:msub><mml:mo>+</mml:mo><mml:msqrt><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>-</mml:mo><mml:msup><mml:mi mathvariant="italic">ω</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:msqrt><mml:msub><mml:mi>W</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd/><mml:mtd><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>,</mml:mo><mml:mi mathvariant="normal">…</mml:mi><mml:mo>,</mml:mo><mml:mi>k</mml:mi><mml:mo>,</mml:mo><mml:mspace linebreak="nobreak" width="1em"/><mml:mi mathvariant="italic">ω</mml:mi><mml:mo>∈</mml:mo><mml:mo>[</mml:mo><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>,</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>]</mml:mo><mml:mo>.</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
        From Appendix B we have that each <inline-formula><mml:math id="M259" display="inline"><mml:mrow><mml:msub><mml:mi>Z</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> is again a standard Gaussian random variable. Then, for <inline-formula><mml:math id="M260" display="inline"><mml:mrow><mml:mi>i</mml:mi><mml:mo>≠</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:math></inline-formula>, we see that

              <disp-formula specific-use="align" content-type="numbered"><mml:math id="M261" display="block"><mml:mtable displaystyle="true"><mml:mtr><mml:mtd><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><mml:mtext>Cov</mml:mtext><mml:mo>[</mml:mo><mml:msub><mml:mi>Z</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>Z</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>]</mml:mo></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><mml:mo>=</mml:mo><mml:mtext>Cov</mml:mtext><mml:mo>(</mml:mo><mml:mi mathvariant="italic">ω</mml:mi><mml:msub><mml:mi>Z</mml:mi><mml:mi mathvariant="normal">M</mml:mi></mml:msub><mml:mo>+</mml:mo><mml:msqrt><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>-</mml:mo><mml:msup><mml:mi mathvariant="italic">ω</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:msqrt><mml:msub><mml:mi>W</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:mi mathvariant="italic">ω</mml:mi><mml:msub><mml:mi>Z</mml:mi><mml:mi mathvariant="normal">M</mml:mi></mml:msub></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mstyle class="stylechange" displaystyle="true"/></mml:mtd><mml:mtd><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><?xmltex \hack{\protect\hphantom{123}}?><mml:mo>+</mml:mo><mml:msqrt><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>-</mml:mo><mml:msup><mml:mi mathvariant="italic">ω</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:msqrt><mml:msub><mml:mi>W</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mstyle displaystyle="true" class="stylechange"/></mml:mtd><mml:mtd><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><mml:mo>=</mml:mo><mml:msup><mml:mi mathvariant="italic">ω</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mtext>Cov</mml:mtext><mml:mo>(</mml:mo><mml:msub><mml:mi>Z</mml:mi><mml:mi mathvariant="normal">M</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>Z</mml:mi><mml:mi mathvariant="normal">M</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mlabeledtr id="App1.Ch1.S3.E13"><mml:mtd><mml:mtext>C2</mml:mtext></mml:mtd><mml:mtd><mml:mstyle displaystyle="true" class="stylechange"/></mml:mtd><mml:mtd><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><mml:mo>=</mml:mo><mml:msup><mml:mi mathvariant="italic">ω</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:mtd></mml:mlabeledtr></mml:mtable></mml:math></disp-formula>

          using pairwise independence of <inline-formula><mml:math id="M262" display="inline"><mml:mrow><mml:msub><mml:mi>Z</mml:mi><mml:mi mathvariant="normal">M</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>W</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>, and <inline-formula><mml:math id="M263" display="inline"><mml:mrow><mml:msub><mml:mi>W</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>. Using a similar technique we see
          <disp-formula id="App1.Ch1.S3.E14" content-type="numbered"><label>C3</label><mml:math id="M264" display="block"><mml:mrow><mml:mtext>Cov</mml:mtext><mml:mo>[</mml:mo><mml:msub><mml:mi>Z</mml:mi><mml:mn mathvariant="normal">0</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>Z</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>]</mml:mo><mml:mo>=</mml:mo><mml:mi mathvariant="italic">ω</mml:mi><mml:mi mathvariant="italic">ρ</mml:mi><mml:mo>.</mml:mo></mml:mrow></mml:math></disp-formula>
        Now let <inline-formula><mml:math id="M265" display="inline"><mml:mrow><mml:mi mathvariant="bold-italic">Z</mml:mi><mml:mo>=</mml:mo><mml:mo>(</mml:mo><mml:msub><mml:mi>Z</mml:mi><mml:mn mathvariant="normal">0</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>Z</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:mi mathvariant="normal">…</mml:mi><mml:mo>,</mml:mo><mml:msub><mml:mi>Z</mml:mi><mml:mi>k</mml:mi></mml:msub><mml:msup><mml:mo>)</mml:mo><mml:mo>′</mml:mo></mml:msup></mml:mrow></mml:math></inline-formula>. Then,
          <disp-formula id="App1.Ch1.S3.E15" content-type="numbered"><label>C4</label><mml:math id="M266" display="block"><mml:mrow><mml:mtext>Cov</mml:mtext><mml:mo>[</mml:mo><mml:mi mathvariant="bold-italic">Z</mml:mi><mml:mo>]</mml:mo><mml:mo>=</mml:mo><mml:mfenced open="(" close=")"><mml:mtable class="matrix" columnalign="center center center center center" framespacing="0em"><mml:mtr><mml:mtd><mml:mn mathvariant="normal">1</mml:mn></mml:mtd><mml:mtd><mml:mrow><mml:mi mathvariant="italic">ω</mml:mi><mml:mi mathvariant="italic">ρ</mml:mi></mml:mrow></mml:mtd><mml:mtd><mml:mi mathvariant="normal">…</mml:mi></mml:mtd><mml:mtd><mml:mi mathvariant="normal">…</mml:mi></mml:mtd><mml:mtd><mml:mrow><mml:mi mathvariant="italic">ω</mml:mi><mml:mi mathvariant="italic">ρ</mml:mi></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mrow><mml:mi mathvariant="italic">ω</mml:mi><mml:mi mathvariant="italic">ρ</mml:mi></mml:mrow></mml:mtd><mml:mtd><mml:mi mathvariant="normal">⋱</mml:mi></mml:mtd><mml:mtd><mml:mrow><mml:msup><mml:mi mathvariant="italic">ω</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:mtd><mml:mtd><mml:mi mathvariant="normal">…</mml:mi></mml:mtd><mml:mtd><mml:mrow><mml:msup><mml:mi mathvariant="italic">ω</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mi mathvariant="normal">⋮</mml:mi></mml:mtd><mml:mtd><mml:mrow><mml:msup><mml:mi mathvariant="italic">ω</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:mtd><mml:mtd><mml:mi mathvariant="normal">⋱</mml:mi></mml:mtd><mml:mtd><mml:mi mathvariant="normal">⋱</mml:mi></mml:mtd><mml:mtd><mml:mi mathvariant="normal">⋮</mml:mi></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mi mathvariant="normal">⋮</mml:mi></mml:mtd><mml:mtd><mml:mi mathvariant="normal">⋮</mml:mi></mml:mtd><mml:mtd><mml:mi mathvariant="normal">⋱</mml:mi></mml:mtd><mml:mtd><mml:mi mathvariant="normal">⋱</mml:mi></mml:mtd><mml:mtd><mml:mrow><mml:msup><mml:mi mathvariant="italic">ω</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mrow><mml:mi mathvariant="italic">ω</mml:mi><mml:mi mathvariant="italic">ρ</mml:mi></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:msup><mml:mi mathvariant="italic">ω</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:mtd><mml:mtd><mml:mi mathvariant="normal">…</mml:mi></mml:mtd><mml:mtd><mml:mrow><mml:msup><mml:mi mathvariant="italic">ω</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:mtd><mml:mtd><mml:mn mathvariant="normal">1</mml:mn></mml:mtd></mml:mtr></mml:mtable></mml:mfenced><mml:mo>.</mml:mo></mml:mrow></mml:math></disp-formula>
        Setting <inline-formula><mml:math id="M267" display="inline"><mml:mrow><mml:mi mathvariant="italic">ρ</mml:mi><mml:mo>=</mml:mo><mml:mi mathvariant="italic">ω</mml:mi></mml:mrow></mml:math></inline-formula> is thus necessary (except for the trivial case where <inline-formula><mml:math id="M268" display="inline"><mml:mrow><mml:mi mathvariant="italic">ω</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0</mml:mn></mml:mrow></mml:math></inline-formula>) and sufficient for univariate probabilistic calibration of
the ensemble as this choice makes <inline-formula><mml:math id="M269" display="inline"><mml:mrow><mml:msub><mml:mi>Z</mml:mi><mml:mn mathvariant="normal">0</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula> indistinguishable from <inline-formula><mml:math id="M270" display="inline"><mml:mrow><mml:msub><mml:mi>Z</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:mi mathvariant="normal">…</mml:mi><mml:mo>,</mml:mo><mml:msub><mml:mi>Z</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> in distribution.</p>
</app>

<app id="App1.Ch1.S4">
  <?xmltex \currentcnt{D}?><label>Appendix D</label><title>Sensitivity of FTE histograms to forecast skill</title>
      <?pagebreak page424?><p id="d1e6212">The simulation setup introduced in Sect. <xref ref-type="sec" rid="Ch1.S3.SS2"/> allows us to control the skill of the synthetic ensemble forecasts through the
parameters <inline-formula><mml:math id="M271" display="inline"><mml:mi mathvariant="italic">ρ</mml:mi></mml:math></inline-formula> and <inline-formula><mml:math id="M272" display="inline"><mml:mi mathvariant="italic">ω</mml:mi></mml:math></inline-formula>, where (as shown above) the restriction <inline-formula><mml:math id="M273" display="inline"><mml:mrow><mml:mi mathvariant="italic">ρ</mml:mi><mml:mo>=</mml:mo><mml:mi mathvariant="italic">ω</mml:mi></mml:mrow></mml:math></inline-formula> is required to ensure calibration of the marginal
distributions. Does the sensitivity of the FTE histogram to misspecified correlation lengths change with changing forecast skill? To investigate this further, we show results for the extreme “no skill” case where <inline-formula><mml:math id="M274" display="inline"><mml:mrow><mml:mi mathvariant="italic">ρ</mml:mi><mml:mo>=</mml:mo><mml:mi mathvariant="italic">ω</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0</mml:mn></mml:mrow></mml:math></inline-formula>, to complement those shown above where we simulated forecasts
with a relatively high correlation (<inline-formula><mml:math id="M275" display="inline"><mml:mrow><mml:mi mathvariant="italic">ρ</mml:mi><mml:mo>=</mml:mo><mml:mi mathvariant="italic">ω</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.8</mml:mn><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> between ensemble mean and verification at each grid point. The other parameters remain
unchanged; i.e., simulation experiments are performed for <inline-formula><mml:math id="M276" display="inline"><mml:mrow><mml:msub><mml:mi>a</mml:mi><mml:mn mathvariant="normal">0</mml:mn></mml:msub><mml:mo>=</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M277" display="inline"><mml:mrow><mml:msub><mml:mi>a</mml:mi><mml:mi mathvariant="normal">M</mml:mi></mml:msub><mml:mo>∈</mml:mo><mml:mo mathvariant="italic">{</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>,</mml:mo><mml:mn mathvariant="normal">1.8</mml:mn><mml:mo>,</mml:mo><mml:mn mathvariant="normal">2</mml:mn><mml:mo>,</mml:mo><mml:mn mathvariant="normal">2.2</mml:mn><mml:mo>,</mml:mo><mml:mn mathvariant="normal">3</mml:mn><mml:mo mathvariant="italic">}</mml:mo></mml:mrow></mml:math></inline-formula>. By Eq. (<xref ref-type="disp-formula" rid="Ch1.E8"/>), the
ensemble members are then independent realizations of a mean zero Gaussian random field with Matérn covariance having correlation length <inline-formula><mml:math id="M278" display="inline"><mml:mrow><mml:mi>a</mml:mi><mml:mo>=</mml:mo><mml:msub><mml:mi>a</mml:mi><mml:mi mathvariant="normal">M</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>.</p>
      <p id="d1e6346">Figure <xref ref-type="fig" rid="App1.Ch1.S4.F10"/> gives an example of simulated ensemble member fields which are mutually independent and uncorrelated with the
verification field, while their spatial structure is 10 % miscalibrated in the case of rows A and C. In contrast to Fig. <xref ref-type="fig" rid="Ch1.F3"/>,
where the positive skill of the ensemble forecasts entails some degree of correspondence between the features in the forecast and verification fields,
there is no such correspondence in between the fields in Fig. <xref ref-type="fig" rid="App1.Ch1.S4.F10"/>. The corresponding FTE histograms and their associated
<inline-formula><mml:math id="M279" display="inline"><mml:mi mathvariant="italic">β</mml:mi></mml:math></inline-formula>-scores, however, are able to identify row B as spatially calibrated and row A (row C) as having a correlation length ratio that is too small
(large).</p>
      <p id="d1e6362">A similar story is provided by Fig. <xref ref-type="fig" rid="App1.Ch1.S4.F11"/>, which, like Fig. <xref ref-type="fig" rid="Ch1.F5"/> where <inline-formula><mml:math id="M280" display="inline"><mml:mrow><mml:mi mathvariant="italic">ρ</mml:mi><mml:mo>=</mml:mo><mml:mi mathvariant="italic">ω</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.8</mml:mn></mml:mrow></mml:math></inline-formula>, demonstrates how the estimated <inline-formula><mml:math id="M281" display="inline"><mml:mi mathvariant="italic">β</mml:mi></mml:math></inline-formula>-distribution parameters, <inline-formula><mml:math id="M282" display="inline"><mml:mi mathvariant="italic">β</mml:mi></mml:math></inline-formula>-score, and <inline-formula><mml:math id="M283" display="inline"><mml:mi mathvariant="italic">β</mml:mi></mml:math></inline-formula>-bias vary over increasing thresholds for different correlation length
ratios. While the associated experiments differ in that ensemble members have no univariate skill here (i.e., <inline-formula><mml:math id="M284" display="inline"><mml:mrow><mml:mi mathvariant="italic">ρ</mml:mi><mml:mo>=</mml:mo><mml:mi mathvariant="italic">ω</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0</mml:mn></mml:mrow></mml:math></inline-formula>), the behavior
witnessed in the two figures is nearly indistinguishable. Thus, we conclude that correlation between the ensemble and verification has a negligible
effect on the FTE histograms' ability to detect miscalibration in spatial structure. This was not obvious to us a priori, but perhaps one can think of
these “no skill” simulations as the residual fields that remain after the “predictable component” has been subtracted from both ensemble member
and verification fields. In any case, the insensitivity to forecast skill is good news for the practical application of FTE histograms, where forecast
skill is usually unknown and confounding effects are undesirable. It is a reminder though that they are a tool for assessing forecast calibration, not forecast skill.</p><?xmltex \hack{\clearpage}?><?xmltex \floatpos{h!}?><fig id="App1.Ch1.S4.F10"><?xmltex \currentcnt{D1}?><label>Figure D1</label><caption><p id="d1e6426">As Fig. <xref ref-type="fig" rid="Ch1.F3"/> but ensemble fields are constructed with skill parameter <inline-formula><mml:math id="M285" display="inline"><mml:mrow><mml:mi mathvariant="italic">ρ</mml:mi><mml:mo>=</mml:mo><mml:mi mathvariant="italic">ω</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0</mml:mn></mml:mrow></mml:math></inline-formula>.</p></caption>
        <?xmltex \hack{\hsize\textwidth}?>
        <?xmltex \igopts{width=341.433071pt}?><graphic xlink:href="https://npg.copernicus.org/articles/27/411/2020/npg-27-411-2020-f10.png"/>

      </fig>

      <?xmltex \floatpos{h!}?><fig id="App1.Ch1.S4.F11"><?xmltex \currentcnt{D2}?><label>Figure D2</label><caption><p id="d1e6457">As Fig. <xref ref-type="fig" rid="Ch1.F5"/> but ensemble fields are constructed with skill parameter <inline-formula><mml:math id="M286" display="inline"><mml:mrow><mml:mi mathvariant="italic">ρ</mml:mi><mml:mo>=</mml:mo><mml:mi mathvariant="italic">ω</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0</mml:mn></mml:mrow></mml:math></inline-formula>.</p></caption>
        <?xmltex \hack{\hsize\textwidth}?>
        <?xmltex \igopts{width=369.885827pt}?><graphic xlink:href="https://npg.copernicus.org/articles/27/411/2020/npg-27-411-2020-f11.png"/>

      </fig>

<?xmltex \hack{\clearpage}?>
</app>
  </app-group><notes notes-type="codedataavailability"><title>Code and data availability</title>

      <p id="d1e6492">The code and data used for this study are available in the accompanying Zenodo repositories (code:
<ext-link xlink:href="https://doi.org/10.5281/zenodo.3945515" ext-link-type="DOI">10.5281/zenodo.3945515</ext-link>, <xref ref-type="bibr" rid="bib1.bibx18" id="altparen.38"/>; data: <ext-link xlink:href="https://doi.org/10.5281/zenodo.3945512" ext-link-type="DOI">10.5281/zenodo.3945512</ext-link>, <xref ref-type="bibr" rid="bib1.bibx19" id="altparen.39"/>). The most recent version of the code can also be found in Josh Jacobson's Git repository (<uri>https://github.com/joshhjacobson/FTE</uri>, last access: 14 July 2020).</p>
  </notes><notes notes-type="authorcontribution"><title>Author contributions</title>

      <p id="d1e6513">This study is based on the Master's work of JJ under supervision of WK and MS. The concept of this study was developed by MS and extended upon by all involved. JJ implemented the study and performed the analysis with guidance from WK and MS. JB provided the downscaled GEFS
forecast fields. JJ, WK, and MS collaborated in discussing the results and composing the manuscript, with input from JB on Sect. 4.</p>
  </notes><notes notes-type="competinginterests"><title>Competing interests</title>

      <p id="d1e6519">The authors declare that they have no conflict of interest.</p>
  </notes><notes notes-type="sistatement"><title>Special issue statement</title>

      <p id="d1e6525">This article is part of the special issue “Advances in post-processing and blending of deterministic and ensemble forecasts”. It is not associated with a conference.</p>
  </notes><ack><title>Acknowledgements</title><p id="d1e6531">Josh Jacobson was supported by NSF DMS-1407340. William Kleiber was supported by NSF DMS-1811294 and DMS-1923062. Michael Scheuerer and Joseph Bellier were supported by funding from the
US NWS Office of Science &amp; Technology Integration through the Meteorological
Development Laboratory, project no. 720T8MWQML. This work utilized resources from the University of Colorado Boulder Research Computing Group, which is supported by the NSF (awards ACI-1532235 and ACI-1532236), the University of Colorado Boulder, and
Colorado State University.</p></ack><notes notes-type="financialsupport"><title>Financial support</title>

      <p id="d1e6536">This research has been supported by the National Science Foundation (grant nos. DMS-1407340, DMS-1811294, and DMS-1923062) and by the US National Weather Service Office of Science &amp; Technology
Integration (grant no. 720T8MWQML).</p>
  </notes><notes notes-type="reviewstatement"><title>Review statement</title>

      <p id="d1e6542">This paper was edited by Maxime Taillardat and reviewed by two anonymous referees.</p>
  </notes><ref-list>
    <title>References</title>

      <ref id="bib1.bibx1"><label>Anderson(1996)</label><?label anderson_method_1996?><mixed-citation>Anderson, J. L.:
A method for producing and evaluating probabilistic forecasts from ensemble model integrations,
J. Climate,
9, 1518–1530, <ext-link xlink:href="https://doi.org/10.1175/1520-0442(1996)009&lt;1518:AMFPAE&gt;2.0.CO;2" ext-link-type="DOI">10.1175/1520-0442(1996)009&lt;1518:AMFPAE&gt;2.0.CO;2</ext-link>, 1996.</mixed-citation></ref>
      <ref id="bib1.bibx2"><label>Apanasovich et al.(2012)Apanasovich, Genton, and Sun</label><?label apanasovich2012?><mixed-citation>Apanasovich, T. V., Genton, M. G., and Sun, Y.:
A valid Matérn class of cross-covariance functions for multivariate random fields with any number of components,
J. Am. Stat. Assoc.,
107, 180–193, <ext-link xlink:href="https://doi.org/10.1080/01621459.2011.643197" ext-link-type="DOI">10.1080/01621459.2011.643197</ext-link>, 2012.</mixed-citation></ref>
      <ref id="bib1.bibx3"><label>Bellier et al.(2020)Bellier, Scheuerer, and Hamill</label><?label Bellier2020?><mixed-citation>Bellier, J., Scheuerer, M., and Hamill, T. M.:
Precipitation downscaling with Gibbs sampling: An improved method for producing realistic, weather-dependent and anisotropic fields,
J. Hydrometeorol.,
<ext-link xlink:href="https://doi.org/10.1175/JHM-D-20-0069.1" ext-link-type="DOI">10.1175/JHM-D-20-0069.1</ext-link>, online first, 2020.</mixed-citation></ref>
      <ref id="bib1.bibx4"><label>Buizza et al.(2007)Buizza, Bidlot, Wedi, Fuentes, Hamrud, Holt, and Vitart</label><?label BuizzaEA2007?><mixed-citation>Buizza, R., Bidlot, J.-R., Wedi, N., Fuentes, M., Hamrud, M., Holt, G., and Vitart, F.:
The new ECMWF VAREPS (variable resolution ensemble prediction system),
Q. J. Roy. Meteor. Soc.,
133, 681–695, <ext-link xlink:href="https://doi.org/10.1002/qj.75" ext-link-type="DOI">10.1002/qj.75</ext-link>, 2007.</mixed-citation></ref>
      <ref id="bib1.bibx5"><label>Buschow and Friederichs(2020)</label><?label BuschowEA_wavelets_2020?><mixed-citation>Buschow, S. and Friederichs, P.: Using wavelets to verify the scale structure of precipitation forecasts, Adv. Stat. Clim. Meteorol. Oceanogr., 6, 13–30, <ext-link xlink:href="https://doi.org/10.5194/ascmo-6-13-2020" ext-link-type="DOI">10.5194/ascmo-6-13-2020</ext-link>, 2020.</mixed-citation></ref>
      <ref id="bib1.bibx6"><label>Buschow et al.(2019)Buschow, Pidstrigach, and Friederichs</label><?label BuschowEA_wavelets_2019?><mixed-citation>Buschow, S., Pidstrigach, J., and Friederichs, P.: Assessment of wavelet-based spatial verification by means of a stochastic precipitation model (wv_verif v0.1.0), Geosci. Model Dev., 12, 3401–3418, <ext-link xlink:href="https://doi.org/10.5194/gmd-12-3401-2019" ext-link-type="DOI">10.5194/gmd-12-3401-2019</ext-link>, 2019.</mixed-citation></ref>
      <ref id="bib1.bibx7"><label>Cullen and Frey(1999)</label><?label bootstrap_method?><mixed-citation>Cullen, A. C. and Frey, H. C.:
Probabilistic techniques in exposure assessment,
Plenum Press, New York, USA, London, UK, <ext-link xlink:href="https://doi.org/10.1002/sim.958" ext-link-type="DOI">10.1002/sim.958</ext-link>, 1999.</mixed-citation></ref>
      <ref id="bib1.bibx8"><label>Delignette-Muller and Dutang(2015)</label><?label fitdistrplus?><mixed-citation>Delignette-Muller, M. L. and Dutang, C.:
fitdistrplus: An R package for fitting distributions,
J. Stat. Softw.,
64, 1–34, <ext-link xlink:href="https://doi.org/10.18637/jss.v064.i04" ext-link-type="DOI">10.18637/jss.v064.i04</ext-link>, 2015.</mixed-citation></ref>
      <ref id="bib1.bibx9"><label>Feldmann et al.(2015)Feldmann, Scheuerer, and Thorarinsdottir</label><?label FeldmannEA2015?><mixed-citation>Feldmann, K., Scheuerer, M., and Thorarinsdottir, T. L.:
Spatial postprocessing of ensemble forecasts for temperature using nonhomogeneous Gaussian regression,
Mon. Weather Rev.,
143, 955–971, <ext-link xlink:href="https://doi.org/10.1175/MWR-D-14-00210.1" ext-link-type="DOI">10.1175/MWR-D-14-00210.1</ext-link>, 2015.</mixed-citation></ref>
      <ref id="bib1.bibx10"><label>Gagnon et al.(2012)Gagnon, Rousseau, Mailhot, and Caya</label><?label GagnonEA2012?><mixed-citation>Gagnon, P., Rousseau, A. N., Mailhot, A., and Caya, D.:
Spatial disaggregation of mean areal rainfall using Gibbs sampling,
J. Hydrometeorol.,
13, 324–337, <ext-link xlink:href="https://doi.org/10.1175/JHM-D-11-034.1" ext-link-type="DOI">10.1175/JHM-D-11-034.1</ext-link>, 2012.</mixed-citation></ref>
      <ref id="bib1.bibx11"><label>Genton and Kleiber(2015)</label><?label genton2015?><mixed-citation>Genton, M. G. and Kleiber, W.:
Cross-covariance functions for multivariate geostatistics,
Stat. Sci.,
30, 147–163, <ext-link xlink:href="https://doi.org/10.1214/14-STS487" ext-link-type="DOI">10.1214/14-STS487</ext-link>, 2015.</mixed-citation></ref>
      <ref id="bib1.bibx12"><label>Gilleland et al.(2009)Gilleland, Ahijevych, Brown, Casati, and Ebert</label><?label GillelandEA2009?><mixed-citation>Gilleland, E., Ahijevych, D., Brown, B. G., Casati, B., and Ebert, E. E.:
Intercomparison of spatial forecast verification methods,
Weather Forecast.,
24, 1416–1430, <ext-link xlink:href="https://doi.org/10.1175/2009WAF2222269.1" ext-link-type="DOI">10.1175/2009WAF2222269.1</ext-link>, 2009.</mixed-citation></ref>
      <ref id="bib1.bibx13"><label>Gneiting et al.(2008)Gneiting, Stanberry, Grimit, Held, and Johnson</label><?label GneitingEA2008?><mixed-citation>Gneiting, T., Stanberry, L. I., Grimit, E. P., Held, L., and Johnson, N. A.:
Assessing probabilistic forecasts of multivariate quantities, with an application to ensemble predictions of surface winds,
TEST,
17, 211, <ext-link xlink:href="https://doi.org/10.1007/s11749-008-0114-x" ext-link-type="DOI">10.1007/s11749-008-0114-x</ext-link>, 2008.</mixed-citation></ref>
      <ref id="bib1.bibx14"><label>Gneiting et al.(2010)Gneiting, Kleiber, and Schlather</label><?label gneiting_matern_2010?><mixed-citation>Gneiting, T., Kleiber, W., and Schlather, M.:
Matérn cross-covariance functions for multivariate random fields,
J. Am. Stat. Assoc.,
105, 1167–1177, <ext-link xlink:href="https://doi.org/10.1198/jasa.2010.tm09420" ext-link-type="DOI">10.1198/jasa.2010.tm09420</ext-link>, 2010.</mixed-citation></ref>
      <ref id="bib1.bibx15"><label>Hamill(2001)</label><?label hamill_interpretation_2001?><mixed-citation>Hamill, T. M.:
Interpretation of rank histograms for verifying ensemble forecasts,
Mon. Weather Rev.,
129, 550–560, <ext-link xlink:href="https://doi.org/10.1175/1520-0493(2001)129&lt;0550:IORHFV&gt;2.0.CO;2" ext-link-type="DOI">10.1175/1520-0493(2001)129&lt;0550:IORHFV&gt;2.0.CO;2</ext-link>, 2001.</mixed-citation></ref>
      <ref id="bib1.bibx16"><label>Hamill et al.(2013)Hamill, Bates, Whitaker, Murray, Fiorino, Galarneau, Zhu, and Lapenta</label><?label HamillEA2013?><mixed-citation>Hamill, T. M., Bates, G. T., Whitaker, J. S., Murray, D. R., Fiorino, M., Galarneau, T. J., Zhu, Y., and Lapenta, W.:
NOAA's second-generation global medium-range ensemble reforecast dataset,
B. Am. Meteorol. Soc.,
94, 1553–1565, <ext-link xlink:href="https://doi.org/10.1175/BAMS-D-12-00014.1" ext-link-type="DOI">10.1175/BAMS-D-12-00014.1</ext-link>, 2013.</mixed-citation></ref>
      <?pagebreak page427?><ref id="bib1.bibx17"><label>Hou et al.(2014)Hou, Charles, Luo, Toth, Zhu, Krzysztofowicz, Lin, Xie, Seo, Pena, and Cui</label><?label HouEA2014?><mixed-citation>Hou, D., Charles, M., Luo, Y., Toth, Z., Zhu, Y., Krzysztofowicz, R., Lin, Y., Xie, P., Seo, D.-J., Pena, M., and Cui, B.:
Climatology-calibrated precipitation analysis at fine scales: Statistical adjustment of stage IV toward CPC gauge-based analysis,
J. Hydrometeorol.,
15, 2542–2557, <ext-link xlink:href="https://doi.org/10.1175/JHM-D-11-0140.1" ext-link-type="DOI">10.1175/JHM-D-11-0140.1</ext-link>, 2014.</mixed-citation></ref>
      <ref id="bib1.bibx18"><label>Jacobson(2020)</label><?label Jacobson2020?><mixed-citation>Jacobson, J.: FTE: Final revised paper (Version v1.1.0), Zenodo, <ext-link xlink:href="https://doi.org/10.5281/zenodo.3945515" ext-link-type="DOI">10.5281/zenodo.3945515</ext-link>, 2020.</mixed-citation></ref>
      <ref id="bib1.bibx19"><label>Jacobson et al.(2020)</label><?label JacobsonForecast2020?><mixed-citation>Jacobson, J., Bellier, J., Scheuerer, M., and Kleiber, W.: Verifying Spatial Structure in Ensembles of Forecast Fields (Version v1.1) [Data set], Zenodo, <ext-link xlink:href="https://doi.org/10.5281/zenodo.3945512" ext-link-type="DOI">10.5281/zenodo.3945512</ext-link>, 2020.</mixed-citation></ref>
      <ref id="bib1.bibx20"><label>Kapp et al.(2018)Kapp, Friederichs, Brune, and Weniger</label><?label KappEA_wavelets_2018?><mixed-citation>Kapp, F., Friederichs, P., Brune, S., and Weniger, M.:
Spatial verification of high-resolution ensemble precipitation forecasts using local wavelet spectra,
Meteorol. Z.,
27, 467–480, <ext-link xlink:href="https://doi.org/10.1127/metz/2018/0903" ext-link-type="DOI">10.1127/metz/2018/0903</ext-link>, 2018.</mixed-citation></ref>
      <ref id="bib1.bibx21"><label>Keller and Hense(2011)</label><?label Keller+Hense2011?><mixed-citation>Keller, J. D. and Hense, A.:
A new non-Gaussian evaluation method for ensemble forecasts based on analysis rank histograms,
Meteorol. Z.,
20, 107–117, <ext-link xlink:href="https://doi.org/10.1127/0941-2948/2011/0217" ext-link-type="DOI">10.1127/0941-2948/2011/0217</ext-link>, 2011.</mixed-citation></ref>
      <ref id="bib1.bibx22"><label>Kleiber(2017)</label><?label kleiber2017?><mixed-citation>Kleiber, W.:
Coherence for multivariate random fields,
Stat. Sinica,
27, 1675–1697, <ext-link xlink:href="https://doi.org/10.5705/ss.202015.0309" ext-link-type="DOI">10.5705/ss.202015.0309</ext-link>, 2017.</mixed-citation></ref>
      <ref id="bib1.bibx23"><label>Leutbecher and Palmer(2008)</label><?label Leutbecher+Palmer2008?><mixed-citation>Leutbecher, M. and Palmer, T.:
Ensemble forecasting,
J. Comput. Phys.,
227, 3515–3539, <ext-link xlink:href="https://doi.org/10.1016/j.jcp.2007.02.014" ext-link-type="DOI">10.1016/j.jcp.2007.02.014</ext-link>, 2008.</mixed-citation></ref>
      <ref id="bib1.bibx24"><label>Murphy and Winkler(1977)</label><?label Murphy+Winkler1977?><mixed-citation>Murphy, A. H. and Winkler, R. L.:
Reliability of subjective probability forecasts of precipitation and temperature,
J. R. Stat. Soc. C-Appl.,
26, 41–47, <ext-link xlink:href="https://doi.org/10.2307/2346866" ext-link-type="DOI">10.2307/2346866</ext-link>, 1977.</mixed-citation></ref>
      <ref id="bib1.bibx25"><label>Roberts and Lean(2008)</label><?label RobertsLeanFSS?><mixed-citation>Roberts, N. M. and Lean, H. W.:
Scale-selective verification of rainfall accumulations from high-resolution forecasts of convective events,
Mon. Weather Rev.,
136, 78–97, <ext-link xlink:href="https://doi.org/10.1175/2007MWR2123.1" ext-link-type="DOI">10.1175/2007MWR2123.1</ext-link>, 2008.</mixed-citation></ref>
      <ref id="bib1.bibx26"><label>Schefzik et al.(2013)Schefzik, Thorarinsdottir, and Gneiting</label><?label SchefzikEA2013?><mixed-citation>Schefzik, R., Thorarinsdottir, T. L., and Gneiting, T.:
Uncertainty quantification in complex simulation models using ensemble copula coupling,
Stat. Sci.,
28, 616–640, <ext-link xlink:href="https://doi.org/10.1214/13-STS443" ext-link-type="DOI">10.1214/13-STS443</ext-link>, 2013.
</mixed-citation></ref><?xmltex \hack{\newpage}?>
      <ref id="bib1.bibx27"><label>Scheuerer and Hamill(2015)</label><?label Scheuerer+Hamill2015?><mixed-citation>Scheuerer, M. and Hamill, T. M.:
Statistical postprocessing of ensemble precipitation forecasts by fitting censored, shifted Gamma distributions,
Mon. Weather Rev.,
143, 4578–4596, <ext-link xlink:href="https://doi.org/10.1175/MWR-D-15-0061.1" ext-link-type="DOI">10.1175/MWR-D-15-0061.1</ext-link>, 2015.</mixed-citation></ref>
      <ref id="bib1.bibx28"><label>Scheuerer and Hamill(2018)</label><?label scheuerer_generating_2018?><mixed-citation>Scheuerer, M. and Hamill, T. M.:
Generating calibrated ensembles of physically realistic, high-resolution precipitation forecast fields based on GEFS model output,
J. Hydrometeorol.,
19, 1651–1670, <ext-link xlink:href="https://doi.org/10.1175/JHM-D-18-0067.1" ext-link-type="DOI">10.1175/JHM-D-18-0067.1</ext-link>, 2018.</mixed-citation></ref>
      <ref id="bib1.bibx29"><label>Smith and Hansen(2004)</label><?label Smith+Hansen2004?><mixed-citation>Smith, L. A. and Hansen, J. A.:
Extending the limits of ensemble forecast verification with the minimum spanning tree,
Mon. Weather Rev.,
132, 1522–1528, <ext-link xlink:href="https://doi.org/10.1175/1520-0493(2004)132&lt;1522:ETLOEF&gt;2.0.CO;2" ext-link-type="DOI">10.1175/1520-0493(2004)132&lt;1522:ETLOEF&gt;2.0.CO;2</ext-link>, 2004.</mixed-citation></ref>
      <ref id="bib1.bibx30"><label>Thorarinsdottir et al.(2016)Thorarinsdottir, Scheuerer, and Heinz</label><?label ThorarinsdottirEA2016?><mixed-citation>Thorarinsdottir, T. L., Scheuerer, M., and Heinz, C.:
Assessing the calibration of high-dimensional ensemble forecasts using rank histograms,
J. Comput. Graph. Stat.,
25, 105–122, <ext-link xlink:href="https://doi.org/10.1080/10618600.2014.977447" ext-link-type="DOI">10.1080/10618600.2014.977447</ext-link>, 2016.</mixed-citation></ref>
      <ref id="bib1.bibx31"><label>Toth and Kalnay(1993)</label><?label Toth+Kalnay1993?><mixed-citation>Toth, Z. and Kalnay, E.:
Ensemble forecasting at NMC: The generation of perturbations,
B. Am. Meteorol. Soc.,
74, 2317–2330, <ext-link xlink:href="https://doi.org/10.1175/1520-0477(1993)074&lt;2317:EFANTG&gt;2.0.CO;2" ext-link-type="DOI">10.1175/1520-0477(1993)074&lt;2317:EFANTG&gt;2.0.CO;2</ext-link>, 1993.</mixed-citation></ref>
      <ref id="bib1.bibx32"><label>Wilks(2004)</label><?label Wilks2004?><mixed-citation>Wilks, D. S.:
The minimum spanning tree histogram as a verification tool for multidimensional ensemble forecasts,
Mon. Weather Rev.,
132, 1329–1340, <ext-link xlink:href="https://doi.org/10.1175/1520-0493(2004)132&lt;1329:TMSTHA&gt;2.0.CO;2" ext-link-type="DOI">10.1175/1520-0493(2004)132&lt;1329:TMSTHA&gt;2.0.CO;2</ext-link>, 2004.</mixed-citation></ref>
      <ref id="bib1.bibx33"><label>Zhou et al.(2017)Zhou, Zhu, Hou, Luo, Peng, and Wobus</label><?label ZhuEA2017?><mixed-citation>Zhou, X., Zhu, Y., Hou, D., Luo, Y., Peng, J., and Wobus, R.:
Performance of the new NCEP Global Ensemble Forecast System in a parallel experiment,
Weather Forecast.,
32, 1989–2004, <ext-link xlink:href="https://doi.org/10.1175/WAF-D-17-0023.1" ext-link-type="DOI">10.1175/WAF-D-17-0023.1</ext-link>, 2017.</mixed-citation></ref>
      <ref id="bib1.bibx34"><label>Ziegel and Gneiting(2014)</label><?label Ziegel+Gneiting2014?><mixed-citation>Ziegel, J. F. and Gneiting, T.:
Copula calibration,
Electron. J. Stat.,
8, 2619–2638, <ext-link xlink:href="https://doi.org/10.1214/14-EJS964" ext-link-type="DOI">10.1214/14-EJS964</ext-link>, 2014.</mixed-citation></ref>

  </ref-list></back>
    <!--<article-title-html>Beyond univariate calibration: verifying spatial structure in ensembles of forecast fields</article-title-html>
<abstract-html><p>Most available verification metrics for ensemble forecasts focus on univariate quantities. That is, they assess whether the ensemble provides an
adequate representation of the forecast uncertainty about the quantity of interest at a particular location and time. For spatially indexed ensemble  forecasts, however, it is also important that forecast fields reproduce the spatial structure of the observed field and represent the uncertainty  about spatial properties such as the size of the area for which heavy precipitation, high winds, critical fire weather conditions, etc., are
expected. In this article we study the properties of the fraction of threshold exceedance (FTE) histogram, a new diagnostic tool designed for
spatially indexed ensemble forecast fields. Defined as the fraction of grid points where a prescribed threshold is exceeded, the FTE is calculated  for the verification field and separately for each ensemble member. It yields a projection of a – possibly high-dimensional – multivariate
quantity onto a univariate quantity that can be studied with standard tools like verification rank histograms. This projection is appealing since it
reflects a spatial property that is intuitive and directly relevant in applications, though it is not obvious whether the FTE is sufficiently
sensitive to misrepresentation of spatial structure in the ensemble. In a comprehensive simulation study we find that departures from uniformity of
the FTE histograms can indeed be related to forecast ensembles with biased spatial variability and that these histograms detect shortcomings in the  spatial structure of ensemble forecast fields that are not obvious by eye. For demonstration, FTE histograms are applied in the context of spatially
downscaled ensemble precipitation forecast fields from NOAA's Global Ensemble Forecast System.</p></abstract-html>
<ref-html id="bib1.bib1"><label>Anderson(1996)</label><mixed-citation>
Anderson, J. L.:
A method for producing and evaluating probabilistic forecasts from ensemble model integrations,
J. Climate,
9, 1518–1530, <a href="https://doi.org/10.1175/1520-0442(1996)009&lt;1518:AMFPAE&gt;2.0.CO;2" target="_blank">https://doi.org/10.1175/1520-0442(1996)009&lt;1518:AMFPAE&gt;2.0.CO;2</a>, 1996.
</mixed-citation></ref-html>
<ref-html id="bib1.bib2"><label>Apanasovich et al.(2012)Apanasovich, Genton, and Sun</label><mixed-citation>
Apanasovich, T. V., Genton, M. G., and Sun, Y.:
A valid Matérn class of cross-covariance functions for multivariate random fields with any number of components,
J. Am. Stat. Assoc.,
107, 180–193, <a href="https://doi.org/10.1080/01621459.2011.643197" target="_blank">https://doi.org/10.1080/01621459.2011.643197</a>, 2012.
</mixed-citation></ref-html>
<ref-html id="bib1.bib3"><label>Bellier et al.(2020)Bellier, Scheuerer, and Hamill</label><mixed-citation>
Bellier, J., Scheuerer, M., and Hamill, T. M.:
Precipitation downscaling with Gibbs sampling: An improved method for producing realistic, weather-dependent and anisotropic fields,
J. Hydrometeorol.,
<a href="https://doi.org/10.1175/JHM-D-20-0069.1" target="_blank">https://doi.org/10.1175/JHM-D-20-0069.1</a>, online first, 2020.
</mixed-citation></ref-html>
<ref-html id="bib1.bib4"><label>Buizza et al.(2007)Buizza, Bidlot, Wedi, Fuentes, Hamrud, Holt, and Vitart</label><mixed-citation>
Buizza, R., Bidlot, J.-R., Wedi, N., Fuentes, M., Hamrud, M., Holt, G., and Vitart, F.:
The new ECMWF VAREPS (variable resolution ensemble prediction system),
Q. J. Roy. Meteor. Soc.,
133, 681–695, <a href="https://doi.org/10.1002/qj.75" target="_blank">https://doi.org/10.1002/qj.75</a>, 2007.
</mixed-citation></ref-html>
<ref-html id="bib1.bib5"><label>Buschow and Friederichs(2020)</label><mixed-citation>
Buschow, S. and Friederichs, P.: Using wavelets to verify the scale structure of precipitation forecasts, Adv. Stat. Clim. Meteorol. Oceanogr., 6, 13–30, <a href="https://doi.org/10.5194/ascmo-6-13-2020" target="_blank">https://doi.org/10.5194/ascmo-6-13-2020</a>, 2020.
</mixed-citation></ref-html>
<ref-html id="bib1.bib6"><label>Buschow et al.(2019)Buschow, Pidstrigach, and Friederichs</label><mixed-citation>
Buschow, S., Pidstrigach, J., and Friederichs, P.: Assessment of wavelet-based spatial verification by means of a stochastic precipitation model (wv_verif v0.1.0), Geosci. Model Dev., 12, 3401–3418, <a href="https://doi.org/10.5194/gmd-12-3401-2019" target="_blank">https://doi.org/10.5194/gmd-12-3401-2019</a>, 2019.
</mixed-citation></ref-html>
<ref-html id="bib1.bib7"><label>Cullen and Frey(1999)</label><mixed-citation>
Cullen, A. C. and Frey, H. C.:
Probabilistic techniques in exposure assessment,
Plenum Press, New York, USA, London, UK, <a href="https://doi.org/10.1002/sim.958" target="_blank">https://doi.org/10.1002/sim.958</a>, 1999.
</mixed-citation></ref-html>
<ref-html id="bib1.bib8"><label>Delignette-Muller and Dutang(2015)</label><mixed-citation>
Delignette-Muller, M. L. and Dutang, C.:
fitdistrplus: An R package for fitting distributions,
J. Stat. Softw.,
64, 1–34, <a href="https://doi.org/10.18637/jss.v064.i04" target="_blank">https://doi.org/10.18637/jss.v064.i04</a>, 2015.
</mixed-citation></ref-html>
<ref-html id="bib1.bib9"><label>Feldmann et al.(2015)Feldmann, Scheuerer, and Thorarinsdottir</label><mixed-citation>
Feldmann, K., Scheuerer, M., and Thorarinsdottir, T. L.:
Spatial postprocessing of ensemble forecasts for temperature using nonhomogeneous Gaussian regression,
Mon. Weather Rev.,
143, 955–971, <a href="https://doi.org/10.1175/MWR-D-14-00210.1" target="_blank">https://doi.org/10.1175/MWR-D-14-00210.1</a>, 2015.
</mixed-citation></ref-html>
<ref-html id="bib1.bib10"><label>Gagnon et al.(2012)Gagnon, Rousseau, Mailhot, and Caya</label><mixed-citation>
Gagnon, P., Rousseau, A. N., Mailhot, A., and Caya, D.:
Spatial disaggregation of mean areal rainfall using Gibbs sampling,
J. Hydrometeorol.,
13, 324–337, <a href="https://doi.org/10.1175/JHM-D-11-034.1" target="_blank">https://doi.org/10.1175/JHM-D-11-034.1</a>, 2012.
</mixed-citation></ref-html>
<ref-html id="bib1.bib11"><label>Genton and Kleiber(2015)</label><mixed-citation>
Genton, M. G. and Kleiber, W.:
Cross-covariance functions for multivariate geostatistics,
Stat. Sci.,
30, 147–163, <a href="https://doi.org/10.1214/14-STS487" target="_blank">https://doi.org/10.1214/14-STS487</a>, 2015.
</mixed-citation></ref-html>
<ref-html id="bib1.bib12"><label>Gilleland et al.(2009)Gilleland, Ahijevych, Brown, Casati, and Ebert</label><mixed-citation>
Gilleland, E., Ahijevych, D., Brown, B. G., Casati, B., and Ebert, E. E.:
Intercomparison of spatial forecast verification methods,
Weather Forecast.,
24, 1416–1430, <a href="https://doi.org/10.1175/2009WAF2222269.1" target="_blank">https://doi.org/10.1175/2009WAF2222269.1</a>, 2009.
</mixed-citation></ref-html>
<ref-html id="bib1.bib13"><label>Gneiting et al.(2008)Gneiting, Stanberry, Grimit, Held, and Johnson</label><mixed-citation>
Gneiting, T., Stanberry, L. I., Grimit, E. P., Held, L., and Johnson, N. A.:
Assessing probabilistic forecasts of multivariate quantities, with an application to ensemble predictions of surface winds,
TEST,
17, 211, <a href="https://doi.org/10.1007/s11749-008-0114-x" target="_blank">https://doi.org/10.1007/s11749-008-0114-x</a>, 2008.
</mixed-citation></ref-html>
<ref-html id="bib1.bib14"><label>Gneiting et al.(2010)Gneiting, Kleiber, and Schlather</label><mixed-citation>
Gneiting, T., Kleiber, W., and Schlather, M.:
Matérn cross-covariance functions for multivariate random fields,
J. Am. Stat. Assoc.,
105, 1167–1177, <a href="https://doi.org/10.1198/jasa.2010.tm09420" target="_blank">https://doi.org/10.1198/jasa.2010.tm09420</a>, 2010.
</mixed-citation></ref-html>
<ref-html id="bib1.bib15"><label>Hamill(2001)</label><mixed-citation>
Hamill, T. M.:
Interpretation of rank histograms for verifying ensemble forecasts,
Mon. Weather Rev.,
129, 550–560, <a href="https://doi.org/10.1175/1520-0493(2001)129&lt;0550:IORHFV&gt;2.0.CO;2" target="_blank">https://doi.org/10.1175/1520-0493(2001)129&lt;0550:IORHFV&gt;2.0.CO;2</a>, 2001.
</mixed-citation></ref-html>
<ref-html id="bib1.bib16"><label>Hamill et al.(2013)Hamill, Bates, Whitaker, Murray, Fiorino, Galarneau, Zhu, and Lapenta</label><mixed-citation>
Hamill, T. M., Bates, G. T., Whitaker, J. S., Murray, D. R., Fiorino, M., Galarneau, T. J., Zhu, Y., and Lapenta, W.:
NOAA's second-generation global medium-range ensemble reforecast dataset,
B. Am. Meteorol. Soc.,
94, 1553–1565, <a href="https://doi.org/10.1175/BAMS-D-12-00014.1" target="_blank">https://doi.org/10.1175/BAMS-D-12-00014.1</a>, 2013.
</mixed-citation></ref-html>
<ref-html id="bib1.bib17"><label>Hou et al.(2014)Hou, Charles, Luo, Toth, Zhu, Krzysztofowicz, Lin, Xie, Seo, Pena, and Cui</label><mixed-citation>
Hou, D., Charles, M., Luo, Y., Toth, Z., Zhu, Y., Krzysztofowicz, R., Lin, Y., Xie, P., Seo, D.-J., Pena, M., and Cui, B.:
Climatology-calibrated precipitation analysis at fine scales: Statistical adjustment of stage IV toward CPC gauge-based analysis,
J. Hydrometeorol.,
15, 2542–2557, <a href="https://doi.org/10.1175/JHM-D-11-0140.1" target="_blank">https://doi.org/10.1175/JHM-D-11-0140.1</a>, 2014.
</mixed-citation></ref-html>
<ref-html id="bib1.bib18"><label>Jacobson(2020)</label><mixed-citation>
Jacobson, J.: FTE: Final revised paper (Version v1.1.0), Zenodo, <a href="https://doi.org/10.5281/zenodo.3945515" target="_blank">https://doi.org/10.5281/zenodo.3945515</a>, 2020.
</mixed-citation></ref-html>
<ref-html id="bib1.bib19"><label>Jacobson et al.(2020)</label><mixed-citation>
Jacobson, J., Bellier, J., Scheuerer, M., and Kleiber, W.: Verifying Spatial Structure in Ensembles of Forecast Fields (Version v1.1) [Data set], Zenodo, <a href="https://doi.org/10.5281/zenodo.3945512" target="_blank">https://doi.org/10.5281/zenodo.3945512</a>, 2020.
</mixed-citation></ref-html>
<ref-html id="bib1.bib20"><label>Kapp et al.(2018)Kapp, Friederichs, Brune, and Weniger</label><mixed-citation>
Kapp, F., Friederichs, P., Brune, S., and Weniger, M.:
Spatial verification of high-resolution ensemble precipitation forecasts using local wavelet spectra,
Meteorol. Z.,
27, 467–480, <a href="https://doi.org/10.1127/metz/2018/0903" target="_blank">https://doi.org/10.1127/metz/2018/0903</a>, 2018.
</mixed-citation></ref-html>
<ref-html id="bib1.bib21"><label>Keller and Hense(2011)</label><mixed-citation>
Keller, J. D. and Hense, A.:
A new non-Gaussian evaluation method for ensemble forecasts based on analysis rank histograms,
Meteorol. Z.,
20, 107–117, <a href="https://doi.org/10.1127/0941-2948/2011/0217" target="_blank">https://doi.org/10.1127/0941-2948/2011/0217</a>, 2011.
</mixed-citation></ref-html>
<ref-html id="bib1.bib22"><label>Kleiber(2017)</label><mixed-citation>
Kleiber, W.:
Coherence for multivariate random fields,
Stat. Sinica,
27, 1675–1697, <a href="https://doi.org/10.5705/ss.202015.0309" target="_blank">https://doi.org/10.5705/ss.202015.0309</a>, 2017.
</mixed-citation></ref-html>
<ref-html id="bib1.bib23"><label>Leutbecher and Palmer(2008)</label><mixed-citation>
Leutbecher, M. and Palmer, T.:
Ensemble forecasting,
J. Comput. Phys.,
227, 3515–3539, <a href="https://doi.org/10.1016/j.jcp.2007.02.014" target="_blank">https://doi.org/10.1016/j.jcp.2007.02.014</a>, 2008.
</mixed-citation></ref-html>
<ref-html id="bib1.bib24"><label>Murphy and Winkler(1977)</label><mixed-citation>
Murphy, A. H. and Winkler, R. L.:
Reliability of subjective probability forecasts of precipitation and temperature,
J. R. Stat. Soc. C-Appl.,
26, 41–47, <a href="https://doi.org/10.2307/2346866" target="_blank">https://doi.org/10.2307/2346866</a>, 1977.
</mixed-citation></ref-html>
<ref-html id="bib1.bib25"><label>Roberts and Lean(2008)</label><mixed-citation>
Roberts, N. M. and Lean, H. W.:
Scale-selective verification of rainfall accumulations from high-resolution forecasts of convective events,
Mon. Weather Rev.,
136, 78–97, <a href="https://doi.org/10.1175/2007MWR2123.1" target="_blank">https://doi.org/10.1175/2007MWR2123.1</a>, 2008.
</mixed-citation></ref-html>
<ref-html id="bib1.bib26"><label>Schefzik et al.(2013)Schefzik, Thorarinsdottir, and Gneiting</label><mixed-citation>
Schefzik, R., Thorarinsdottir, T. L., and Gneiting, T.:
Uncertainty quantification in complex simulation models using ensemble copula coupling,
Stat. Sci.,
28, 616–640, <a href="https://doi.org/10.1214/13-STS443" target="_blank">https://doi.org/10.1214/13-STS443</a>, 2013.

</mixed-citation></ref-html>
<ref-html id="bib1.bib27"><label>Scheuerer and Hamill(2015)</label><mixed-citation>
Scheuerer, M. and Hamill, T. M.:
Statistical postprocessing of ensemble precipitation forecasts by fitting censored, shifted Gamma distributions,
Mon. Weather Rev.,
143, 4578–4596, <a href="https://doi.org/10.1175/MWR-D-15-0061.1" target="_blank">https://doi.org/10.1175/MWR-D-15-0061.1</a>, 2015.
</mixed-citation></ref-html>
<ref-html id="bib1.bib28"><label>Scheuerer and Hamill(2018)</label><mixed-citation>
Scheuerer, M. and Hamill, T. M.:
Generating calibrated ensembles of physically realistic, high-resolution precipitation forecast fields based on GEFS model output,
J. Hydrometeorol.,
19, 1651–1670, <a href="https://doi.org/10.1175/JHM-D-18-0067.1" target="_blank">https://doi.org/10.1175/JHM-D-18-0067.1</a>, 2018.
</mixed-citation></ref-html>
<ref-html id="bib1.bib29"><label>Smith and Hansen(2004)</label><mixed-citation>
Smith, L. A. and Hansen, J. A.:
Extending the limits of ensemble forecast verification with the minimum spanning tree,
Mon. Weather Rev.,
132, 1522–1528, <a href="https://doi.org/10.1175/1520-0493(2004)132&lt;1522:ETLOEF&gt;2.0.CO;2" target="_blank">https://doi.org/10.1175/1520-0493(2004)132&lt;1522:ETLOEF&gt;2.0.CO;2</a>, 2004.
</mixed-citation></ref-html>
<ref-html id="bib1.bib30"><label>Thorarinsdottir et al.(2016)Thorarinsdottir, Scheuerer, and Heinz</label><mixed-citation>
Thorarinsdottir, T. L., Scheuerer, M., and Heinz, C.:
Assessing the calibration of high-dimensional ensemble forecasts using rank histograms,
J. Comput. Graph. Stat.,
25, 105–122, <a href="https://doi.org/10.1080/10618600.2014.977447" target="_blank">https://doi.org/10.1080/10618600.2014.977447</a>, 2016.
</mixed-citation></ref-html>
<ref-html id="bib1.bib31"><label>Toth and Kalnay(1993)</label><mixed-citation>
Toth, Z. and Kalnay, E.:
Ensemble forecasting at NMC: The generation of perturbations,
B. Am. Meteorol. Soc.,
74, 2317–2330, <a href="https://doi.org/10.1175/1520-0477(1993)074&lt;2317:EFANTG&gt;2.0.CO;2" target="_blank">https://doi.org/10.1175/1520-0477(1993)074&lt;2317:EFANTG&gt;2.0.CO;2</a>, 1993.
</mixed-citation></ref-html>
<ref-html id="bib1.bib32"><label>Wilks(2004)</label><mixed-citation>
Wilks, D. S.:
The minimum spanning tree histogram as a verification tool for multidimensional ensemble forecasts,
Mon. Weather Rev.,
132, 1329–1340, <a href="https://doi.org/10.1175/1520-0493(2004)132&lt;1329:TMSTHA&gt;2.0.CO;2" target="_blank">https://doi.org/10.1175/1520-0493(2004)132&lt;1329:TMSTHA&gt;2.0.CO;2</a>, 2004.
</mixed-citation></ref-html>
<ref-html id="bib1.bib33"><label>Zhou et al.(2017)Zhou, Zhu, Hou, Luo, Peng, and Wobus</label><mixed-citation>
Zhou, X., Zhu, Y., Hou, D., Luo, Y., Peng, J., and Wobus, R.:
Performance of the new NCEP Global Ensemble Forecast System in a parallel experiment,
Weather Forecast.,
32, 1989–2004, <a href="https://doi.org/10.1175/WAF-D-17-0023.1" target="_blank">https://doi.org/10.1175/WAF-D-17-0023.1</a>, 2017.
</mixed-citation></ref-html>
<ref-html id="bib1.bib34"><label>Ziegel and Gneiting(2014)</label><mixed-citation>
Ziegel, J. F. and Gneiting, T.:
Copula calibration,
Electron. J. Stat.,
8, 2619–2638, <a href="https://doi.org/10.1214/14-EJS964" target="_blank">https://doi.org/10.1214/14-EJS964</a>, 2014.
</mixed-citation></ref-html>--></article>
