<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing with OASIS Tables v3.0 20080202//EN" "journalpub-oasis3.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:oasis="http://docs.oasis-open.org/ns/oasis-exchange/table" xml:lang="en" dtd-version="3.0" article-type="research-article">
  <front>
    <journal-meta><journal-id journal-id-type="publisher">NPG</journal-id><journal-title-group>
    <journal-title>Nonlinear Processes in Geophysics</journal-title>
    <abbrev-journal-title abbrev-type="publisher">NPG</abbrev-journal-title><abbrev-journal-title abbrev-type="nlm-ta">Nonlin. Processes Geophys.</abbrev-journal-title>
  </journal-title-group><issn pub-type="epub">1607-7946</issn><publisher>
    <publisher-name>Copernicus Publications</publisher-name>
    <publisher-loc>Göttingen, Germany</publisher-loc>
  </publisher></journal-meta>
    <article-meta>
      <article-id pub-id-type="doi">10.5194/npg-29-171-2022</article-id><title-group><article-title>Using neural networks to improve simulations in the gray zone</article-title><alt-title>NNs to improve gray zone simulations</alt-title>
      </title-group><?xmltex \runningtitle{NNs to improve gray zone simulations}?><?xmltex \runningauthor{R. Kriegmair et al.}?>
      <contrib-group>
        <contrib contrib-type="author" corresp="no" rid="aff1">
          <name><surname>Kriegmair</surname><given-names>Raphael</given-names></name>
          
        </contrib>
        <contrib contrib-type="author" corresp="yes" rid="aff1">
          <name><surname>Ruckstuhl</surname><given-names>Yvonne</given-names></name>
          <email>yvonne.ruckstuhl@lmu.de</email>
        </contrib>
        <contrib contrib-type="author" corresp="no" rid="aff2">
          <name><surname>Rasp</surname><given-names>Stephan</given-names></name>
          
        </contrib>
        <contrib contrib-type="author" corresp="no" rid="aff1">
          <name><surname>Craig</surname><given-names>George</given-names></name>
          
        </contrib>
        <aff id="aff1"><label>1</label><institution>Meteorological Institute Munich, Ludwig-Maximilians-Universität München, Munich, Germany</institution>
        </aff>
        <aff id="aff2"><label>2</label><institution>ClimateAi, Inc., San Francisco, USA</institution>
        </aff>
      </contrib-group>
      <author-notes><corresp id="corr1">Yvonne Ruckstuhl (yvonne.ruckstuhl@lmu.de)</corresp></author-notes><pub-date><day>2</day><month>May</month><year>2022</year></pub-date>
      
      <volume>29</volume>
      <issue>2</issue>
      <fpage>171</fpage><lpage>181</lpage>
      <history>
        <date date-type="received"><day>7</day><month>May</month><year>2021</year></date>
           <date date-type="accepted"><day>27</day><month>February</month><year>2022</year></date>
           <date date-type="rev-recd"><day>6</day><month>September</month><year>2021</year></date>
           <date date-type="rev-request"><day>17</day><month>May</month><year>2021</year></date>
      </history>
      <permissions>
        <copyright-statement>Copyright: © 2022 Raphael Kriegmair et al.</copyright-statement>
        <copyright-year>2022</copyright-year>
      <license license-type="open-access"><license-p>This work is licensed under the Creative Commons Attribution 4.0 International License. To view a copy of this licence, visit <ext-link ext-link-type="uri" xlink:href="https://creativecommons.org/licenses/by/4.0/">https://creativecommons.org/licenses/by/4.0/</ext-link></license-p></license></permissions><self-uri xlink:href="https://npg.copernicus.org/articles/29/171/2022/npg-29-171-2022.html">This article is available from https://npg.copernicus.org/articles/29/171/2022/npg-29-171-2022.html</self-uri><self-uri xlink:href="https://npg.copernicus.org/articles/29/171/2022/npg-29-171-2022.pdf">The full text article is available as a PDF file from https://npg.copernicus.org/articles/29/171/2022/npg-29-171-2022.pdf</self-uri>
      <abstract><title>Abstract</title>

      <p id="d1e114">Machine learning represents a potential method to cope with the gray zone problem of representing motions in dynamical systems on scales comparable to the model resolution. Here we explore the possibility of using a neural network to directly learn the error caused by unresolved scales. We use a modified shallow water model which includes highly nonlinear processes mimicking atmospheric convection. To create the training dataset, we run the model in a high- and a low-resolution setup and compare the difference after one low-resolution time step, starting from the same initial conditions, thereby obtaining an exact target. The neural network is able to learn a large portion of the difference when evaluated on single time step predictions on a validation dataset. When coupled to the low-resolution model, we find large forecast improvements up to 1 d on average. After this, the accumulated error due to the mass conservation violation of the neural network starts to dominate and deteriorates the forecast. This deterioration can effectively be delayed by adding a penalty term to the loss function used to train the ANN to conserve mass in a weak sense. This study reinforces the need to include physical constraints in neural network parameterizations.</p>
  </abstract>
    </article-meta>
  </front>
<body>
      

<sec id="Ch1.S1" sec-type="intro">
  <label>1</label><title>Introduction</title>
      <p id="d1e126">Current limitations on computational power force weather and climate prediction to use relatively low-resolution simulations. Subgrid-scale processes, i.e., processes that are not resolved by the model grid, are typically represented using physical parameterizations <xref ref-type="bibr" rid="bib1.bibx32" id="paren.1"/>. Inaccuracies in these parameterizations are known to cause errors in weather forecasts and biases in climate projections. While parameterizations are becoming more sophisticated over time, there is evidence that key structural uncertainties remain <xref ref-type="bibr" rid="bib1.bibx24 bib1.bibx25 bib1.bibx17" id="paren.2"/>.</p>
      <p id="d1e135">A particularly difficult problem in the representation of unresolved processes is the so-called gray zone <xref ref-type="bibr" rid="bib1.bibx9 bib1.bibx15" id="paren.3"/>, where a certain physical phenomenon such as a cumulus cloud is similar in size to the model resolution and, hence, partially resolved. In the development of many classical parameterizations, features are assumed to be small in comparison to the model resolution. This scale separation provides a conceptual basis for specifying the average effects of the unresolved flow features on the resolved flow. In contrast, there is no theoretical basis for determining such a relationship in the gray zone. Instead, the truncation error of the numerical model is a significant factor. While we might still expect there to be some relationship between the resolved and unresolved parts of the flow, we have no way to define it.</p>
      <p id="d1e141">Viewing the atmosphere as a turbulent flow, with up- and downscale cascades, phenomena like synoptic cyclones and cumulus clouds emerge where geometric or physical constraints impose length scales on the flow <xref ref-type="bibr" rid="bib1.bibx20 bib1.bibx21 bib1.bibx12" id="paren.4"/>. If a numerical model is truncated near one of these scales, the corresponding phenomenon will be only partially resolved, and the simulation will be inaccurate. In particular, the properties of the phenomenon may be determined by the truncation length rather than by the physical scale. A thorough review of the gray zone problem from a turbulence perspective is provided by <xref ref-type="bibr" rid="bib1.bibx15" id="text.5"/>.</p>
      <p id="d1e150">An important example of the gray zone in practice is the simulation of deep convective clouds in kilometer-scale models used operationally for regional weather prediction. The models typically have a horizontal resolution of 2–4 km, which is not sufficient to fully resolve the cumulus clouds with sizes in the range from 1 to 10 km. In these models, the simulated cumulus clouds collapse to a scale proportional to the model grid length, unrealistically becoming smaller and more intense as the resolution is increased <xref ref-type="bibr" rid="bib1.bibx7 bib1.bibx33" id="paren.6"/>. In models with grid lengths over 10 km, the convective clouds are completely subgrid and should be parameterized, while models with resolution under 100 m will accurately reproduce the dynamics of cumulus clouds, provided that the turbulent mixing processes are well represented. In the gray zone in between, the performance of the models depends sensitively on the resolution and details of the parameterizations that are used <xref ref-type="bibr" rid="bib1.bibx16" id="paren.7"/>.</p>
      <p id="d1e160">Using machine learning (ML) methods such as artificial neural networks (ANNs) for alleviating the problems described above has received increasing attention over the past few years. One approach is to avoid the need for parameterizations altogether by emulating the entire model using observations  <xref ref-type="bibr" rid="bib1.bibx6 bib1.bibx23 bib1.bibx13 bib1.bibx11 bib1.bibx31 bib1.bibx10" id="paren.8"/>. In these studies, a dense and noise-free observation network is often assumed. <xref ref-type="bibr" rid="bib1.bibx3" id="text.9"/> and <xref ref-type="bibr" rid="bib1.bibx1" id="text.10"/> circumvent the requirement of this assumption by using data assimilation to form targets for ANNs from sparse and noisy observations.</p>
      <p id="d1e172">Though studies have shown that surrogate models produced by machine learning can be accurate for small dynamical systems, replacing an entire numerical weather prediction model for operational use is not yet within our reach. Therefore, a more practical approach is to use ANNs as replacement for uncertain parameterizations. This has been done either by learning from physics-based expensive parametrization schemes <xref ref-type="bibr" rid="bib1.bibx22 bib1.bibx27" id="paren.11"/> or high-resolution simulations <xref ref-type="bibr" rid="bib1.bibx19 bib1.bibx5 bib1.bibx2 bib1.bibx26 bib1.bibx35" id="paren.12"/>, which is the approach we take here.
Such data-driven techniques could be a way to reduce the structural uncertainty of traditional parameterizations, even at gray zone resolutions where the physical basis of the parameterization is no longer valid. The first challenge is to create the training data, i.e., to separate the resolved and unresolved scales from the high-resolution simulation. <xref ref-type="bibr" rid="bib1.bibx5" id="text.13"/> use a coarse-graining approach based on subtracting the coarse-grained advection term from the local tendencies. This approach can be used for any model and resolution but is sensitive to the choice of grid and time step. Furthermore, the resulting subgrid tendencies are only an approximation and may not represent the real difference between the low- and high-resolution model. <xref ref-type="bibr" rid="bib1.bibx35" id="text.14"/> use the same model version for low- and high-resolution simulations and compute exact differences after a single low-resolution time step by starting both model versions from the same initial conditions. They manage to obtain stable long-term simulations, using the low-resolution model with a machine learning correction, that come close the high-resolution ground truth.</p>
      <p id="d1e187">Here, we use the modified rotating shallow water (modRSW) model to explore the use of a machine learning subgrid representation in a highly nonlinear dynamical system. The modRSW is an idealized fluid model of convective-scale numerical weather prediction in which convection is triggered by orography. As such, the model mimics the gray zone problem of operational kilometer-scale models. Using a simplified model allows us to focus on some key conceptual questions surrounding machine learning parameterizations, such as how choices in the neural network training affect long-term physical consistency. In particular, we include weak physical constraints in the training procedure.</p>
      <p id="d1e190">The contents of this work are outlined in the following.
Section <xref ref-type="sec" rid="Ch1.S2"/> introduces the experiment setup used to obtain and analyze results. The modRSW model is briefly explained in Sect. <xref ref-type="sec" rid="Ch1.S2.SS1"/>, followed by a description of the training data generation in Sect. <xref ref-type="sec" rid="Ch1.S2.SS2"/>. The architecture and training process of the ANN used in this research are given in Sect. <xref ref-type="sec" rid="Ch1.S2.SS3"/>. Our verification metrics are defined in Sect. <xref ref-type="sec" rid="Ch1.S3"/>. The results are presented in Sect. <xref ref-type="sec" rid="Ch1.S4"/>, followed by concluding remarks in Sect. <xref ref-type="sec" rid="Ch1.S5"/>.</p>
</sec>
<sec id="Ch1.S2">
  <label>2</label><title>Experiment setup</title>
<sec id="Ch1.S2.SS1">
  <label>2.1</label><title>The modRSW model</title>
      <p id="d1e223">The modRSW model <xref ref-type="bibr" rid="bib1.bibx18" id="paren.15"/> used in this research represents an extended version of the 1D shallow water equations, i.e., 1D fluid flow over orography. Its prognostic variables are fluid height <inline-formula><mml:math id="M1" display="inline"><mml:mi>h</mml:mi></mml:math></inline-formula>, wind speed <inline-formula><mml:math id="M2" display="inline"><mml:mi>u</mml:mi></mml:math></inline-formula>, and a rain mass fraction <inline-formula><mml:math id="M3" display="inline"><mml:mi>r</mml:mi></mml:math></inline-formula>.
Based on the model by <xref ref-type="bibr" rid="bib1.bibx34" id="text.16"/>, it implements two threshold heights, <inline-formula><mml:math id="M4" display="inline"><mml:mrow><mml:msub><mml:mi>H</mml:mi><mml:mi mathvariant="normal">c</mml:mi></mml:msub><mml:mo>&lt;</mml:mo><mml:msub><mml:mi>H</mml:mi><mml:mi mathvariant="normal">r</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>, which initiate convection and rain production, respectively. Convection is stimulated by modifying the pressure term to remain constant where <inline-formula><mml:math id="M5" display="inline"><mml:mi>h</mml:mi></mml:math></inline-formula> rises above <inline-formula><mml:math id="M6" display="inline"><mml:mrow><mml:msub><mml:mi>H</mml:mi><mml:mi mathvariant="normal">c</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>. In contrast to <xref ref-type="bibr" rid="bib1.bibx34" id="text.17"/>, the modRSW model does not apply diffusion or stochastic forcing. The model is mass conserving, meaning that the domain mean of <inline-formula><mml:math id="M7" display="inline"><mml:mi>h</mml:mi></mml:math></inline-formula> is constant over time.
In this study, a small but significant model-intrinsic drift in the domain mean of <inline-formula><mml:math id="M8" display="inline"><mml:mi>u</mml:mi></mml:math></inline-formula> is removed by adding a relaxation term.
This term is defined by using a corresponding timescale <inline-formula><mml:math id="M9" display="inline"><mml:mrow><mml:msub><mml:mi>t</mml:mi><mml:mtext>relax</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula>, as <inline-formula><mml:math id="M10" display="inline"><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mover accent="true"><mml:mi>u</mml:mi><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:mn mathvariant="normal">0</mml:mn></mml:msub><mml:mo>-</mml:mo><mml:msub><mml:mover accent="true"><mml:mi>u</mml:mi><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:mi>t</mml:mi></mml:msub><mml:mo>)</mml:mo><mml:mo>⋅</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mtext>relax</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula>, where the overbar denotes the domain mean.
Depending on the orography used, this model yields a range of dynamical organization between regular and chaotic behaviors. Orography is defined as a superposition of cosines with the wavenumbers <inline-formula><mml:math id="M11" display="inline"><mml:mrow><mml:mi>k</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>/</mml:mo><mml:mi>L</mml:mi><mml:mo>,</mml:mo><mml:mi mathvariant="normal">…</mml:mi><mml:mo>,</mml:mo><mml:msub><mml:mi>k</mml:mi><mml:mtext>max</mml:mtext></mml:msub><mml:mo>/</mml:mo><mml:mi>L</mml:mi></mml:mrow></mml:math></inline-formula> (<inline-formula><mml:math id="M12" display="inline"><mml:mi>L</mml:mi></mml:math></inline-formula> domain length). Amplitudes are given as <inline-formula><mml:math id="M13" display="inline"><mml:mrow><mml:mi>A</mml:mi><mml:mo>(</mml:mo><mml:mi>k</mml:mi><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>/</mml:mo><mml:mi>k</mml:mi></mml:mrow></mml:math></inline-formula>, while phase shifts for each term are randomly chosen from <inline-formula><mml:math id="M14" display="inline"><mml:mrow><mml:mo>[</mml:mo><mml:mn mathvariant="normal">0</mml:mn><mml:mo>,</mml:mo><mml:mi>L</mml:mi><mml:mo>]</mml:mo></mml:mrow></mml:math></inline-formula>. In this work, two realizations of the orography are selected to represent the regular and more chaotic dynamical behaviors. Figure <xref ref-type="fig" rid="Ch1.F1"/> displays a 24 h segment of the simulation corresponding to each orography.</p><?xmltex \hack{\newpage}?><?xmltex \floatpos{t}?><fig id="Ch1.F1" specific-use="star"><?xmltex \currentcnt{1}?><?xmltex \def\figurename{Figure}?><label>Figure 1</label><caption><p id="d1e435">A 24 h segment of the HR simulation for the three model variables <inline-formula><mml:math id="M15" display="inline"><mml:mi>h</mml:mi></mml:math></inline-formula>, <inline-formula><mml:math id="M16" display="inline"><mml:mi>u</mml:mi></mml:math></inline-formula>, and <inline-formula><mml:math id="M17" display="inline"><mml:mi>r</mml:mi></mml:math></inline-formula> (from top to bottom), corresponding to the regular case <bold>(a)</bold> and the chaotic case <bold>(b)</bold>.</p></caption>
          <?xmltex \igopts{width=398.338583pt}?><graphic xlink:href="https://npg.copernicus.org/articles/29/171/2022/npg-29-171-2022-f01.png"/>

        </fig>

</sec>
<sec id="Ch1.S2.SS2">
  <label>2.2</label><title>Training data generation</title>
      <p id="d1e479">Conceptually, the ANN's task is to correct a low-resolution (LR) model forecast towards the model truth, which is a coarse-grained high-resolution (HR) model simulation. The coarse-graining factor in this study is set to 4, which is analogous to the range of scales found in the gray zone where deep cumulus convection is partially resolved (e.g., 2.5–10 km). <xref ref-type="bibr" rid="bib1.bibx13" id="text.18"/> show that the choice of coarse-graining factor can substantially affect the performance of ML methods. In our case, however, choosing a larger factor would correspond to a coarse model grid length that is larger than the typical cloud size, changing the nature of the problem from learning to improve poorly resolved existing features in the coarse simulation to parameterizing features that might not be seen at all. The dynamical time step of the model is determined at each iteration, based on the Courant–Friedrichs–Lewy (CFL) criterion. To achieve temporally equidistant output states for both resolutions, the time step is truncated accordingly when necessary.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F2"><?xmltex \currentcnt{2}?><?xmltex \def\figurename{Figure}?><label>Figure 2</label><caption><p id="d1e487">Schematic of the training data generation process. A HR run is coarse grained to LR to generate the model truth. Each model truth state is integrated forward for one time step using LR dynamics. The difference between the obtained states and corresponding model truth defines the desired network output (red arrows), while the preceding model truth defines the network input.</p></caption>
          <?xmltex \igopts{width=236.157874pt}?><graphic xlink:href="https://npg.copernicus.org/articles/29/171/2022/npg-29-171-2022-f02.png"/>

        </fig>

      <p id="d1e496">A training sample (input target pair) is defined by the model truth at some time <inline-formula><mml:math id="M18" display="inline"><mml:mrow><mml:msub><mml:mi>t</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> and the difference between the model truth and the corresponding LR forecast at <inline-formula><mml:math id="M19" display="inline"><mml:mrow><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>+</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>+</mml:mo><mml:mi mathvariant="normal">d</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:math></inline-formula>, respectively (see Fig. <xref ref-type="fig" rid="Ch1.F2"/>). To generate the model truth, HR data are obtained by integrating the modRSW model forward using the parameters shown in Table <xref ref-type="table" rid="Ch1.T1"/>. All states and the orography are subsequently coarse grained to LR, resulting in model truth (LR1).
Each LR1 state is integrated forward for a single time step, using the modRSW model on LR with the coarse grained orography, resulting in a single step prediction (LR2). The synchronized differences <inline-formula><mml:math id="M20" display="inline"><mml:mrow><mml:mtext>LR1</mml:mtext><mml:mo>(</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>)</mml:mo><mml:mo>-</mml:mo><mml:mtext>LR2</mml:mtext><mml:mo>(</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> then define the training targets corresponding to the input <inline-formula><mml:math id="M21" display="inline"><mml:mrow><mml:mtext>LR1</mml:mtext><mml:mo>(</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>, which includes the orography. A time series of <inline-formula><mml:math id="M22" display="inline"><mml:mrow><mml:mi>T</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">200</mml:mn><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mn mathvariant="normal">000</mml:mn></mml:mrow></mml:math></inline-formula> time steps, which is equivalent to approximately 57 d in real time, is generated for both orographies. The first day of the simulation is discarded as spin up, the subsequent 30 d are used for training, and the remaining 26 d are used for validation purposes. The decorrelation length scale of the model is roughly 4 h.</p>

<?xmltex \floatpos{t}?><table-wrap id="Ch1.T1"><?xmltex \currentcnt{1}?><label>Table 1</label><caption><p id="d1e615">Model setting parameters.</p></caption><oasis:table frame="topbot"><oasis:tgroup cols="3">
     <oasis:colspec colnum="1" colname="col1" align="left"/>
     <oasis:colspec colnum="2" colname="col2" align="left"/>
     <oasis:colspec colnum="3" colname="col3" align="right"/>
     <oasis:thead>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">Model parameter</oasis:entry>
         <oasis:entry colname="col2">Symbol</oasis:entry>
         <oasis:entry colname="col3">Value</oasis:entry>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row>
         <oasis:entry colname="col1">HR grid point number</oasis:entry>
         <oasis:entry colname="col2"><inline-formula><mml:math id="M23" display="inline"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mtext>HR</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col3"><inline-formula><mml:math id="M24" display="inline"><mml:mn mathvariant="normal">800</mml:mn></mml:math></inline-formula></oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">LR grid point number</oasis:entry>
         <oasis:entry colname="col2"><inline-formula><mml:math id="M25" display="inline"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mtext>LR</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col3"><inline-formula><mml:math id="M26" display="inline"><mml:mn mathvariant="normal">200</mml:mn></mml:math></inline-formula></oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Time step</oasis:entry>
         <oasis:entry colname="col2"><inline-formula><mml:math id="M27" display="inline"><mml:mrow><mml:mi mathvariant="normal">d</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col3">0.001</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Domain size (non-dim.)</oasis:entry>
         <oasis:entry colname="col2"><inline-formula><mml:math id="M28" display="inline"><mml:mi>L</mml:mi></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col3">1.0</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">CFL</oasis:entry>
         <oasis:entry colname="col2">–</oasis:entry>
         <oasis:entry colname="col3">0.5</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Convection threshold</oasis:entry>
         <oasis:entry colname="col2"><inline-formula><mml:math id="M29" display="inline"><mml:mrow><mml:msub><mml:mi>H</mml:mi><mml:mi mathvariant="normal">c</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col3">1.02</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Rain threshold</oasis:entry>
         <oasis:entry colname="col2"><inline-formula><mml:math id="M30" display="inline"><mml:mrow><mml:msub><mml:mi>H</mml:mi><mml:mi mathvariant="normal">r</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col3">1.05</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Initial total height</oasis:entry>
         <oasis:entry colname="col2"><inline-formula><mml:math id="M31" display="inline"><mml:mrow><mml:msub><mml:mi>H</mml:mi><mml:mn mathvariant="normal">0</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col3">1.0</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Rossby number</oasis:entry>
         <oasis:entry colname="col2"><inline-formula><mml:math id="M32" display="inline"><mml:mi mathvariant="italic">Ro</mml:mi></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col3"><inline-formula><mml:math id="M33" display="inline"><mml:mi mathvariant="normal">∞</mml:mi></mml:math></inline-formula></oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Froude number</oasis:entry>
         <oasis:entry colname="col2"><inline-formula><mml:math id="M34" display="inline"><mml:mi mathvariant="italic">Fr</mml:mi></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col3">1.1</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Effective gravity</oasis:entry>
         <oasis:entry colname="col2"><inline-formula><mml:math id="M35" display="inline"><mml:mi>g</mml:mi></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col3"><inline-formula><mml:math id="M36" display="inline"><mml:mrow><mml:msup><mml:mi mathvariant="italic">Fr</mml:mi><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula></oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Beta</oasis:entry>
         <oasis:entry colname="col2"><inline-formula><mml:math id="M37" display="inline"><mml:mi mathvariant="italic">β</mml:mi></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col3">0.2</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Alpha</oasis:entry>
         <oasis:entry colname="col2"><inline-formula><mml:math id="M38" display="inline"><mml:mrow><mml:msup><mml:mi mathvariant="italic">α</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col3">10</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Rain conversion factor</oasis:entry>
         <oasis:entry colname="col2"><inline-formula><mml:math id="M39" display="inline"><mml:mrow><mml:msup><mml:mi>c</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col3"><inline-formula><mml:math id="M40" display="inline"><mml:mrow><mml:mn mathvariant="normal">0.1</mml:mn><mml:mo>×</mml:mo><mml:mi>g</mml:mi><mml:mo>×</mml:mo><mml:msub><mml:mi>H</mml:mi><mml:mi mathvariant="normal">r</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula></oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">Wind relaxation timescale</oasis:entry>
         <oasis:entry colname="col2"><inline-formula><mml:math id="M41" display="inline"><mml:mrow><mml:msub><mml:mi>t</mml:mi><mml:mtext>relax</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col3"><inline-formula><mml:math id="M42" display="inline"><mml:mrow><mml:mi mathvariant="normal">d</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:math></inline-formula></oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">Orography generation</oasis:entry>
         <oasis:entry colname="col2"/>
         <oasis:entry colname="col3"/>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Maximum wave number</oasis:entry>
         <oasis:entry colname="col2"><inline-formula><mml:math id="M43" display="inline"><mml:mrow><mml:msub><mml:mi>k</mml:mi><mml:mtext>max</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col3">100</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Maximum amplitude</oasis:entry>
         <oasis:entry colname="col2"><inline-formula><mml:math id="M44" display="inline"><mml:mrow><mml:msub><mml:mi>B</mml:mi><mml:mtext>max</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col3">0.1</oasis:entry>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup></oasis:table></table-wrap>

</sec>
<sec id="Ch1.S2.SS3">
  <label>2.3</label><title>Convolutional ANN</title>
      <p id="d1e1047">A characteristic property of convolutional ANNs is that they reflect spatial invariance and localization. These two properties also apply to the dynamics of many physical systems, such as the one investigated here. They differ from, e.g., dense networks by the use of a so-called kernel. This vector is shifted across the domain grid, covering <inline-formula><mml:math id="M45" display="inline"><mml:mi>k</mml:mi></mml:math></inline-formula> grid points at each position. At each position, the dot product of the kernel and current grid values is computed, determining (along with an activation function) the corresponding output value. For more details on convolutional ANNs, we refer to <xref ref-type="bibr" rid="bib1.bibx14" id="text.19"/>.</p>
      <p id="d1e1060">The ANN structure used in this research is described in the following. There are five hidden layers applied, each using the ReLU activation function. The input layer uses ReLU as well, while the output layer uses a linear activation function. All hidden layers have 32 filters. The input and output layer shapes are defined by the input and target data. The kernel size is set uniformly to three grid points.</p>
      <p id="d1e1063">The loss is determined during training by comparing the ANN output to the corresponding target. A standard measure for loss is the mean squared error (MSE). However, any loss function can be used to tailor the application. For example, additional terms can be added to impose weak constraints on the training process as, for example, done in <xref ref-type="bibr" rid="bib1.bibx30" id="text.20"/>. This possibility is exploited here to impose mass conservation in a weak sense. The constraint is implemented by penalizing the deviation of the square of the domain mean corrections of  <italic>h</italic> from zero, resulting in the following loss function:
            <disp-formula id="Ch1.E1" content-type="numbered"><label>1</label><mml:math id="M46" display="block"><mml:mrow><mml:mtext>MSE</mml:mtext><mml:mo>(</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mtext>out</mml:mtext></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mtext>target</mml:mtext></mml:msub><mml:mo>)</mml:mo><mml:mo>+</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mtext>mass</mml:mtext></mml:msub><mml:mo>⋅</mml:mo><mml:msup><mml:mfenced close=")" open="("><mml:mover accent="true"><mml:mrow><mml:msubsup><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mtext>out</mml:mtext><mml:mi>h</mml:mi></mml:msubsup></mml:mrow><mml:mo mathvariant="normal">‾</mml:mo></mml:mover></mml:mfenced><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>
          where the second term represents a weighted mass conservation constraint. In this expression, <inline-formula><mml:math id="M47" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mtext>out</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M48" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mtext>target</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula> are the output and corresponding target of the ANN, respectively, MSE denotes the mean squared error, the tunable scalar <inline-formula><mml:math id="M49" display="inline"><mml:mrow><mml:msub><mml:mi>w</mml:mi><mml:mtext>mass</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula> is the mass conservation  constraint weighting, <inline-formula><mml:math id="M50" display="inline"><mml:mrow><mml:msubsup><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mtext>out</mml:mtext><mml:mi>h</mml:mi></mml:msubsup></mml:mrow></mml:math></inline-formula> is the ANN output for <inline-formula><mml:math id="M51" display="inline"><mml:mi>h</mml:mi></mml:math></inline-formula>,  and the overbar denotes the domain mean.</p>
      <p id="d1e1176">The Adam algorithm, with a learning rate of <inline-formula><mml:math id="M52" display="inline"><mml:mrow><mml:msup><mml:mn mathvariant="normal">10</mml:mn><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">3</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula>, is used to minimize the loss function over the ANN weights in batches of 256 samples. Since the loss function is typically not convex, the ANN likely converges to a local rather than the global minimum. To sample this error, we repeat the training of each ANN presented in this paper, with randomly chosen initial weights, five times. For all ANNs, a total of 1000 epochs is performed. The ANN architecture and hyperparameters were selected based on a loose tuning procedure where no strong sensitivities were detected. The training is done using Python library Keras <xref ref-type="bibr" rid="bib1.bibx8" id="paren.21"/>.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F3"><?xmltex \currentcnt{3}?><?xmltex \def\figurename{Figure}?><label>Figure 3</label><caption><p id="d1e1199">Loss function value for <inline-formula><mml:math id="M53" display="inline"><mml:mrow><mml:msub><mml:mi>w</mml:mi><mml:mtext>mass</mml:mtext></mml:msub><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0</mml:mn></mml:mrow></mml:math></inline-formula> (MSE) of the validation data corresponding to the last five epochs of the training process (blue) for each trained ANN (<inline-formula><mml:math id="M54" display="inline"><mml:mi>x</mml:mi></mml:math></inline-formula> axis). For each ANN, the mean loss function value over the last five epochs is depicted in orange.</p></caption>
          <?xmltex \igopts{width=199.169291pt}?><graphic xlink:href="https://npg.copernicus.org/articles/29/171/2022/npg-29-171-2022-f03.png"/>

        </fig>

      <?xmltex \floatpos{t}?><fig id="Ch1.F4" specific-use="star"><?xmltex \currentcnt{4}?><?xmltex \def\figurename{Figure}?><label>Figure 4</label><caption><p id="d1e1232">Mean (bars) and standard deviations
<inline-formula><mml:math id="M55" display="inline"><mml:mrow><mml:msub><mml:mtext>SD</mml:mtext><mml:mtext>total</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula>, <inline-formula><mml:math id="M56" display="inline"><mml:mrow><mml:msub><mml:mtext>SD</mml:mtext><mml:mtext>time</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula>, and
<inline-formula><mml:math id="M57" display="inline"><mml:mrow><mml:msub><mml:mtext>SD</mml:mtext><mml:mtext>ANN</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula> (error bars; from dark to light,
respectively) of the RMSE <bold>(a, d, g)</bold>, SME
<bold>(b, e, h)</bold>, and bias <bold>(c, f, i)</bold> of ANN corrected (blue) and uncorrected (orange) single time step predictions of the validation data with respect to the model truth for the regular and chaotic cases (<inline-formula><mml:math id="M58" display="inline"><mml:mi>y</mml:mi></mml:math></inline-formula> axis) and for the variables <inline-formula><mml:math id="M59" display="inline"><mml:mi>h</mml:mi></mml:math></inline-formula>, <inline-formula><mml:math id="M60" display="inline"><mml:mi>u</mml:mi></mml:math></inline-formula>, and <inline-formula><mml:math id="M61" display="inline"><mml:mi>r</mml:mi></mml:math></inline-formula> (from top to bottom).</p></caption>
          <?xmltex \igopts{width=341.433071pt}?><graphic xlink:href="https://npg.copernicus.org/articles/29/171/2022/npg-29-171-2022-f04.png"/>

        </fig>

</sec>
</sec>
<sec id="Ch1.S3">
  <label>3</label><title>Verification methods</title>
      <p id="d1e1321">As the initial training weights of the ANNs and the exact number of epochs performed are, to some extent, arbitrary, it is desirable to measure the sensitivity of our results to the realization of these quantities. Figure <xref ref-type="fig" rid="Ch1.F3"/> shows the MSE of the validation dataset of the last five epochs (<inline-formula><mml:math id="M62" display="inline"><mml:mi>y</mml:mi></mml:math></inline-formula> axis) for five ANNs with different realizations of initial training weights (<inline-formula><mml:math id="M63" display="inline"><mml:mi>x</mml:mi></mml:math></inline-formula> axis) for both orographies.  Since the MSE appears sensitive to both the initial weights and the epoch number, we use both to sample the total ANN variability, resulting in <inline-formula><mml:math id="M64" display="inline"><mml:mrow><mml:mn mathvariant="normal">5</mml:mn><mml:mo>×</mml:mo><mml:mn mathvariant="normal">5</mml:mn><mml:mo>=</mml:mo><mml:mn mathvariant="normal">25</mml:mn></mml:mrow></mml:math></inline-formula> samples for each ANN training setup that is presented in the remainder of this paper.</p>
      <p id="d1e1356">In the following, the main scores that are used to verify the efficacy of the ANNs are the root mean squared error (RMSE), spatial mean error (SME), and bias:

              <disp-formula specific-use="gather" content-type="numbered"><mml:math id="M65" display="block"><mml:mtable displaystyle="true"><mml:mlabeledtr id="Ch1.E2"><mml:mtd><mml:mtext>2</mml:mtext></mml:mtd><mml:mtd><mml:mrow><mml:mstyle displaystyle="true" class="stylechange"/><mml:mtext>RMSE</mml:mtext><mml:mo>(</mml:mo><mml:mtext mathvariant="bold">y</mml:mtext><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:msqrt><mml:mrow><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mn mathvariant="normal">1</mml:mn><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mtext>LR</mml:mtext></mml:msub></mml:mrow></mml:mfrac></mml:mstyle><mml:munderover><mml:mo movablelimits="false">∑</mml:mo><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mtext>LR</mml:mtext></mml:msub></mml:mrow></mml:munderover><mml:mo mathsize="1.1em">(</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>-</mml:mo><mml:msubsup><mml:mi>y</mml:mi><mml:mi>j</mml:mi><mml:mtext>true</mml:mtext></mml:msubsup><mml:msup><mml:mo mathsize="1.1em">)</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:msqrt></mml:mrow></mml:mtd></mml:mlabeledtr><mml:mlabeledtr id="Ch1.E3"><mml:mtd><mml:mtext>3</mml:mtext></mml:mtd><mml:mtd><mml:mrow><mml:mstyle displaystyle="true" class="stylechange"/><mml:mtext>SME</mml:mtext><mml:mo>(</mml:mo><mml:mtext mathvariant="bold">y</mml:mtext><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:mfenced open="|" close="|"><mml:mrow><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mn mathvariant="normal">1</mml:mn><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mtext>LR</mml:mtext></mml:msub></mml:mrow></mml:mfrac></mml:mstyle><mml:munderover><mml:mo movablelimits="false">∑</mml:mo><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mtext>LR</mml:mtext></mml:msub></mml:mrow></mml:munderover><mml:mo mathsize="1.1em">(</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>-</mml:mo><mml:msubsup><mml:mi>y</mml:mi><mml:mi>j</mml:mi><mml:mtext>true</mml:mtext></mml:msubsup><mml:mo mathsize="1.1em">)</mml:mo></mml:mrow></mml:mfenced></mml:mrow></mml:mtd></mml:mlabeledtr><mml:mlabeledtr id="Ch1.E4"><mml:mtd><mml:mtext>4</mml:mtext></mml:mtd><mml:mtd><mml:mrow><mml:mstyle displaystyle="true" class="stylechange"/><mml:mtext>bias</mml:mtext><mml:mo>(</mml:mo><mml:mtext mathvariant="bold">y</mml:mtext><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mn mathvariant="normal">1</mml:mn><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mtext>LR</mml:mtext></mml:msub></mml:mrow></mml:mfrac></mml:mstyle><mml:munderover><mml:mo movablelimits="false">∑</mml:mo><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mtext>LR</mml:mtext></mml:msub></mml:mrow></mml:munderover><mml:mo mathsize="1.1em">(</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>-</mml:mo><mml:msubsup><mml:mi>y</mml:mi><mml:mi>j</mml:mi><mml:mtext>true</mml:mtext></mml:msubsup><mml:mo mathsize="1.1em">)</mml:mo><mml:mo>,</mml:mo></mml:mrow></mml:mtd></mml:mlabeledtr></mml:mtable></mml:math></disp-formula>

          where <inline-formula><mml:math id="M66" display="inline"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mtext>LR</mml:mtext></mml:msub><mml:mo>=</mml:mo><mml:mn mathvariant="normal">200</mml:mn></mml:mrow></mml:math></inline-formula> is the number of grid points, and <inline-formula><mml:math id="M67" display="inline"><mml:mrow><mml:msup><mml:mtext mathvariant="bold">y</mml:mtext><mml:mtext>true</mml:mtext></mml:msup><mml:mo>∈</mml:mo><mml:msup><mml:mi mathvariant="double-struck">R</mml:mi><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mtext>LR</mml:mtext></mml:msub></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula> is a snapshot of the model truth. Multiple samples of these scores are obtained using both the 25 realizations of the ANN and a sequence of initial conditions provided by the time dimension. The final verification metrics are then the mean and standard deviation (SD) of the respective scores <inline-formula><mml:math id="M68" display="inline"><mml:mrow><mml:mi>X</mml:mi><mml:mo>∈</mml:mo><mml:mo mathvariant="italic">{</mml:mo><mml:mtext>RMSE, SME, bias</mml:mtext><mml:mo mathvariant="italic">}</mml:mo></mml:mrow></mml:math></inline-formula>, as follows:

              <disp-formula specific-use="gather" content-type="numbered"><mml:math id="M69" display="block"><mml:mtable displaystyle="true"><mml:mlabeledtr id="Ch1.E5"><mml:mtd><mml:mtext>5</mml:mtext></mml:mtd><mml:mtd><mml:mrow><mml:mstyle displaystyle="true" class="stylechange"/><mml:mover accent="true"><mml:mi>X</mml:mi><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:mo>(</mml:mo><mml:mtext mathvariant="bold">y</mml:mtext><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mn mathvariant="normal">1</mml:mn><mml:mrow><mml:msub><mml:mi>L</mml:mi><mml:mtext>ANN</mml:mtext></mml:msub><mml:msub><mml:mi>T</mml:mi><mml:mtext>veri</mml:mtext></mml:msub></mml:mrow></mml:mfrac></mml:mstyle><mml:munderover><mml:mo movablelimits="false">∑</mml:mo><mml:mrow><mml:mi>l</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mi>L</mml:mi><mml:mtext>ANN</mml:mtext></mml:msub></mml:mrow></mml:munderover><mml:munderover><mml:mo movablelimits="false">∑</mml:mo><mml:mrow><mml:mi>t</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mi>T</mml:mi><mml:mtext>veri</mml:mtext></mml:msub></mml:mrow></mml:munderover><mml:mi>X</mml:mi><mml:mo mathsize="1.1em">(</mml:mo><mml:msubsup><mml:mtext mathvariant="bold">y</mml:mtext><mml:mi>t</mml:mi><mml:mi>l</mml:mi></mml:msubsup><mml:mo mathsize="1.1em">)</mml:mo></mml:mrow></mml:mtd></mml:mlabeledtr><mml:mlabeledtr id="Ch1.E6"><mml:mtd><mml:mtext>6</mml:mtext></mml:mtd><mml:mtd><mml:mrow><mml:mstyle displaystyle="true" class="stylechange"/><?xmltex \hack{\hbox\bgroup\fontsize{9.5}{9.5}\selectfont$\displaystyle}?><mml:msub><mml:mtext mathvariant="normal">SD</mml:mtext><mml:mtext>total</mml:mtext></mml:msub><mml:mfenced close=")" open="("><mml:mrow><mml:mi>X</mml:mi><mml:mo>(</mml:mo><mml:mtext mathvariant="bold">y</mml:mtext><mml:mo>)</mml:mo></mml:mrow></mml:mfenced><mml:mo>=</mml:mo><mml:msqrt><mml:mrow><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mn mathvariant="normal">1</mml:mn><mml:mrow><mml:msub><mml:mi>L</mml:mi><mml:mtext>ANN</mml:mtext></mml:msub><mml:msub><mml:mi>T</mml:mi><mml:mtext>veri</mml:mtext></mml:msub></mml:mrow></mml:mfrac></mml:mstyle><mml:munderover><mml:mo movablelimits="false">∑</mml:mo><mml:mrow><mml:mi>l</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mi>L</mml:mi><mml:mtext>ANN</mml:mtext></mml:msub></mml:mrow></mml:munderover><mml:munderover><mml:mo movablelimits="false">∑</mml:mo><mml:mrow><mml:mi>t</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mi>T</mml:mi><mml:mtext>veri</mml:mtext></mml:msub></mml:mrow></mml:munderover><mml:mo mathsize="1.1em">(</mml:mo><mml:mi>X</mml:mi><mml:mo mathsize="1.1em">(</mml:mo><mml:msubsup><mml:mtext mathvariant="bold">y</mml:mtext><mml:mi>t</mml:mi><mml:mi>l</mml:mi></mml:msubsup><mml:mo mathsize="1.1em">)</mml:mo><mml:mo>-</mml:mo><mml:mover accent="true"><mml:mi>X</mml:mi><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:mo>(</mml:mo><mml:mtext mathvariant="bold">y</mml:mtext><mml:mo>)</mml:mo><mml:msup><mml:mo mathsize="1.1em">)</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:msqrt><?xmltex \hack{$\egroup}?></mml:mrow></mml:mtd></mml:mlabeledtr><mml:mlabeledtr id="Ch1.E7"><mml:mtd><mml:mtext>7</mml:mtext></mml:mtd><mml:mtd><mml:mrow><mml:mstyle displaystyle="true" class="stylechange"/><?xmltex \hack{\hbox\bgroup\fontsize{9.5}{9.5}\selectfont$\displaystyle}?><mml:msub><mml:mtext mathvariant="normal">SD</mml:mtext><mml:mtext>time</mml:mtext></mml:msub><mml:mfenced close=")" open="("><mml:mrow><mml:mi>X</mml:mi><mml:mo>(</mml:mo><mml:mtext mathvariant="bold">y</mml:mtext><mml:mo>)</mml:mo></mml:mrow></mml:mfenced><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mn mathvariant="normal">1</mml:mn><mml:mrow><mml:msub><mml:mi>L</mml:mi><mml:mtext>ANN</mml:mtext></mml:msub></mml:mrow></mml:mfrac></mml:mstyle><mml:munderover><mml:mo movablelimits="false">∑</mml:mo><mml:mrow><mml:mi>l</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mi>L</mml:mi><mml:mtext>ANN</mml:mtext></mml:msub></mml:mrow></mml:munderover><mml:msqrt><mml:mrow><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mn mathvariant="normal">1</mml:mn><mml:mrow><mml:msub><mml:mi>T</mml:mi><mml:mtext>veri</mml:mtext></mml:msub></mml:mrow></mml:mfrac></mml:mstyle><mml:munderover><mml:mo movablelimits="false">∑</mml:mo><mml:mrow><mml:mi>t</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mi>T</mml:mi><mml:mtext>veri</mml:mtext></mml:msub></mml:mrow></mml:munderover><mml:mo mathsize="1.1em">(</mml:mo><mml:mi>X</mml:mi><mml:mo mathsize="1.1em">(</mml:mo><mml:msubsup><mml:mtext mathvariant="bold">y</mml:mtext><mml:mi>t</mml:mi><mml:mi>l</mml:mi></mml:msubsup><mml:mo mathsize="1.1em">)</mml:mo><mml:mo>-</mml:mo><mml:mover accent="true"><mml:mi>X</mml:mi><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:mo>(</mml:mo><mml:msup><mml:mtext mathvariant="bold">y</mml:mtext><mml:mi>l</mml:mi></mml:msup><mml:mo>)</mml:mo><mml:msup><mml:mo mathsize="1.1em">)</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:msqrt><?xmltex \hack{$\egroup}?></mml:mrow></mml:mtd></mml:mlabeledtr><mml:mlabeledtr id="Ch1.E8"><mml:mtd><mml:mtext>8</mml:mtext></mml:mtd><mml:mtd><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><?xmltex \hack{\hbox\bgroup\fontsize{9.5}{9.5}\selectfont$\displaystyle}?><mml:msub><mml:mtext mathvariant="normal">SD</mml:mtext><mml:mtext>ANN</mml:mtext></mml:msub><mml:mfenced close=")" open="("><mml:mrow><mml:mi>X</mml:mi><mml:mo>(</mml:mo><mml:mtext mathvariant="bold">y</mml:mtext><mml:mo>)</mml:mo></mml:mrow></mml:mfenced><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mn mathvariant="normal">1</mml:mn><mml:mrow><mml:msub><mml:mi>T</mml:mi><mml:mtext>veri</mml:mtext></mml:msub></mml:mrow></mml:mfrac></mml:mstyle><mml:munderover><mml:mo movablelimits="false">∑</mml:mo><mml:mrow><mml:mi>l</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mi>T</mml:mi><mml:mtext>veri</mml:mtext></mml:msub></mml:mrow></mml:munderover><mml:msqrt><mml:mrow><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mn mathvariant="normal">1</mml:mn><mml:mrow><mml:msub><mml:mi>L</mml:mi><mml:mtext>ANN</mml:mtext></mml:msub></mml:mrow></mml:mfrac></mml:mstyle><mml:munderover><mml:mo movablelimits="false">∑</mml:mo><mml:mrow><mml:mi>l</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mi>L</mml:mi><mml:mtext>ANN</mml:mtext></mml:msub></mml:mrow></mml:munderover><mml:mo mathsize="1.1em">(</mml:mo><mml:mi>X</mml:mi><mml:mo mathsize="1.1em">(</mml:mo><mml:msubsup><mml:mtext mathvariant="bold">y</mml:mtext><mml:mi>t</mml:mi><mml:mi>l</mml:mi></mml:msubsup><mml:mo mathsize="1.1em">)</mml:mo><mml:mo>-</mml:mo><mml:mover accent="true"><mml:mi>X</mml:mi><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:mo>(</mml:mo><mml:msub><mml:mtext mathvariant="bold">y</mml:mtext><mml:mi>t</mml:mi></mml:msub><mml:mo>)</mml:mo><mml:msup><mml:mo mathsize="1.1em">)</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:msqrt><?xmltex \hack{$\egroup}?><mml:mo>,</mml:mo></mml:mrow></mml:mtd></mml:mlabeledtr></mml:mtable></mml:math></disp-formula>

          where <inline-formula><mml:math id="M70" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula> and <inline-formula><mml:math id="M71" display="inline"><mml:mi>l</mml:mi></mml:math></inline-formula> index time and ANN realizations, respectively, <inline-formula><mml:math id="M72" display="inline"><mml:mrow><mml:mover accent="true"><mml:mi>X</mml:mi><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:mo>(</mml:mo><mml:msub><mml:mtext mathvariant="bold">y</mml:mtext><mml:mi>t</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> indicates the mean over time steps, and <inline-formula><mml:math id="M73" display="inline"><mml:mrow><mml:mover accent="true"><mml:mi>X</mml:mi><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:mo>(</mml:mo><mml:msup><mml:mtext mathvariant="bold">y</mml:mtext><mml:mi>l</mml:mi></mml:msup><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> the mean over ANN realizations. Note that Eqs. (<xref ref-type="disp-formula" rid="Ch1.E7"/>) and (<xref ref-type="disp-formula" rid="Ch1.E8"/>) are meant to isolate the variability inherited from initial conditions and ANN realizations, respectively. We apply these verification metrics to both single time step predictions and 48 h forecasts.
<list list-type="bullet"><list-item>
      <p id="d1e2105"><italic>Single time step predictions.</italic> For each time step corresponding to the validation dataset (<inline-formula><mml:math id="M74" display="inline"><mml:mrow><mml:msub><mml:mi>T</mml:mi><mml:mtext>veri</mml:mtext></mml:msub><mml:mo>=</mml:mo><mml:mn mathvariant="normal">92</mml:mn><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mn mathvariant="normal">863</mml:mn></mml:mrow></mml:math></inline-formula>), the model truth is used as initial condition for a single time step prediction of the LR model, creating a LR prediction (LR). This LR prediction is subsequently corrected by the ANN, creating the corresponding ANN-corrected prediction (<inline-formula><mml:math id="M75" display="inline"><mml:mrow><mml:msub><mml:mtext>LR</mml:mtext><mml:mtext>ANN</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula>).</p></list-item><list-item>
      <p id="d1e2140"><italic>The 48 h forecasts.</italic> The 48 h forecasts are generated from a set of 50 initial conditions (<inline-formula><mml:math id="M76" display="inline"><mml:mrow><mml:msub><mml:mi>T</mml:mi><mml:mtext>veri</mml:mtext></mml:msub><mml:mo>=</mml:mo><mml:mn mathvariant="normal">50</mml:mn></mml:mrow></mml:math></inline-formula>) taken from the validation dataset. To ensure independence, the initial conditions are set 4 h apart, which is roughly the decorrelation length scale of the model. After each low-resolution single time step prediction, the ANN is applied to create initial conditions for the next LR single time step prediction, creating a 48 h LR ANN-corrected forecast (<inline-formula><mml:math id="M77" display="inline"><mml:mrow><mml:msub><mml:mtext>LR</mml:mtext><mml:mtext>ANN</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula>). As a reference, a LR simulation without ANN corrections (LR) is run in parallel.</p></list-item></list>
For the single time step predictions and the 48 h forecasts, our verification metrics are applied to both LR and <inline-formula><mml:math id="M78" display="inline"><mml:mrow><mml:msub><mml:mtext>LR</mml:mtext><mml:mtext>ANN</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula> and compared in Sect. <xref ref-type="sec" rid="Ch1.S4"/>. Note that <inline-formula><mml:math id="M79" display="inline"><mml:mrow><mml:msub><mml:mi>L</mml:mi><mml:mtext>ANN</mml:mtext></mml:msub><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:math></inline-formula> for LR, yielding <inline-formula><mml:math id="M80" display="inline"><mml:mrow><mml:mover accent="true"><mml:mi>X</mml:mi><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:mo>(</mml:mo><mml:mtext>LR</mml:mtext><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:msub><mml:mtext>SD</mml:mtext><mml:mtext>ANN</mml:mtext></mml:msub><mml:mo>(</mml:mo><mml:mi>X</mml:mi><mml:mo>(</mml:mo><mml:mtext>LR</mml:mtext><mml:mo>)</mml:mo><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M81" display="inline"><mml:mrow><mml:msub><mml:mtext>SD</mml:mtext><mml:mtext>total</mml:mtext></mml:msub><mml:mo>(</mml:mo><mml:mi>X</mml:mi><mml:mo>(</mml:mo><mml:mtext>LR</mml:mtext><mml:mo>)</mml:mo><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:msub><mml:mtext>SD</mml:mtext><mml:mtext>time</mml:mtext></mml:msub><mml:mo>(</mml:mo><mml:mi>X</mml:mi><mml:mo>(</mml:mo><mml:mtext>LR</mml:mtext><mml:mo>)</mml:mo><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>.</p>
</sec>
<sec id="Ch1.S4">
  <label>4</label><title>Results</title>
      <p id="d1e2288">We performed a series of experiments designed to investigate the feasibility of using an ANN to correct for model error due to unresolved scales. In Sect. <xref ref-type="sec" rid="Ch1.S4.SS1"/>, we first explore the performance of the ANNs trained with the standard MSE as the loss function (<inline-formula><mml:math id="M82" display="inline"><mml:mrow><mml:msub><mml:mi>w</mml:mi><mml:mtext>mass</mml:mtext></mml:msub><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0</mml:mn></mml:mrow></mml:math></inline-formula> in Eq. <xref ref-type="disp-formula" rid="Ch1.E1"/>). Next, the weak constraint is added to the loss function as in Eq. (<xref ref-type="disp-formula" rid="Ch1.E1"/>), and the benefits are examined in Sect. <xref ref-type="sec" rid="Ch1.S4.SS2"/>.</p>
<sec id="Ch1.S4.SS1">
  <label>4.1</label><title>ANN with standard loss function</title>
      <p id="d1e2321">Figure <xref ref-type="fig" rid="Ch1.F4"/> shows the results for the single time step
predictions. The improvements achieved by the ANN with respect to LR in terms of RMSE are substantial for <inline-formula><mml:math id="M83" display="inline"><mml:mi>h</mml:mi></mml:math></inline-formula>, <inline-formula><mml:math id="M84" display="inline"><mml:mi>u</mml:mi></mml:math></inline-formula>, and <inline-formula><mml:math id="M85" display="inline"><mml:mi>r</mml:mi></mml:math></inline-formula>, amounting to 97 %, 89 %, and 90 %  for the regular case and 96 %, 84 %, and 92 % for the chaotic case. Also, the negative biases present in <inline-formula><mml:math id="M86" display="inline"><mml:mi>u</mml:mi></mml:math></inline-formula> and <inline-formula><mml:math id="M87" display="inline"><mml:mi>r</mml:mi></mml:math></inline-formula> for LR virtually disappear when the ANN is applied. However, as the ANN is not explicitly instructed to conserve mass, a small SME is introduced in variable <inline-formula><mml:math id="M88" display="inline"><mml:mi>h</mml:mi></mml:math></inline-formula>. Its significance will become apparent when analyzing the 48 h forecasts. It is interesting to note that the ANN's architecture and learning process is unbiased (small <inline-formula><mml:math id="M89" display="inline"><mml:mover accent="true"><mml:mtext>bias</mml:mtext><mml:mo mathvariant="normal">‾</mml:mo></mml:mover></mml:math></inline-formula>), although single ANN realizations may be biased (large <inline-formula><mml:math id="M90" display="inline"><mml:mrow><mml:msub><mml:mtext>SD</mml:mtext><mml:mtext>ANN</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula>). Also, in contrast to the SME and bias, for the RMSE, the initial conditions are the main source of variability (<inline-formula><mml:math id="M91" display="inline"><mml:mrow><mml:msub><mml:mtext>SD</mml:mtext><mml:mtext>time</mml:mtext></mml:msub><mml:mo>≫</mml:mo><mml:msub><mml:mtext>SD</mml:mtext><mml:mtext>ANN</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula>). This is better visible in Figs. <xref ref-type="fig" rid="Ch1.F7"/> and <xref ref-type="fig" rid="Ch1.F8"/>, where the results for <inline-formula><mml:math id="M92" display="inline"><mml:mrow><mml:msub><mml:mi>w</mml:mi><mml:mtext>mass</mml:mtext></mml:msub><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0</mml:mn></mml:mrow></mml:math></inline-formula> are plotted again.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F5"><?xmltex \currentcnt{5}?><?xmltex \def\figurename{Figure}?><label>Figure 5</label><caption><p id="d1e2430">The <inline-formula><mml:math id="M93" display="inline"><mml:mover accent="true"><mml:mtext>RMSE</mml:mtext><mml:mo mathvariant="normal">‾</mml:mo></mml:mover></mml:math></inline-formula> evolution of 48 h forecasts for model variables <inline-formula><mml:math id="M94" display="inline"><mml:mi>h</mml:mi></mml:math></inline-formula>, <inline-formula><mml:math id="M95" display="inline"><mml:mi>u</mml:mi></mml:math></inline-formula>, and <inline-formula><mml:math id="M96" display="inline"><mml:mi>r</mml:mi></mml:math></inline-formula> (from top to bottom) of LR (black) and  <inline-formula><mml:math id="M97" display="inline"><mml:mrow><mml:msub><mml:mtext>LR</mml:mtext><mml:mtext>ANN</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula> (blue). The shaded region corresponds to <inline-formula><mml:math id="M98" display="inline"><mml:mrow><mml:msub><mml:mtext>SD</mml:mtext><mml:mtext>total</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula>.</p></caption>
          <?xmltex \igopts{width=236.157874pt}?><graphic xlink:href="https://npg.copernicus.org/articles/29/171/2022/npg-29-171-2022-f05.png"/>

        </fig>

      <?xmltex \floatpos{t}?><fig id="Ch1.F6"><?xmltex \currentcnt{6}?><?xmltex \def\figurename{Figure}?><label>Figure 6</label><caption><p id="d1e2495">Same as Fig. <xref ref-type="fig" rid="Ch1.F5"/> but for the MSE.</p></caption>
          <?xmltex \igopts{width=236.157874pt}?><graphic xlink:href="https://npg.copernicus.org/articles/29/171/2022/npg-29-171-2022-f06.png"/>

        </fig>

      <p id="d1e2507">Next we examine the effect of the ANN on a 48 h forecast. Here we compare the LR simulation with (<inline-formula><mml:math id="M99" display="inline"><mml:mrow><mml:msub><mml:mtext>LR</mml:mtext><mml:mtext>ANN</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula>) and without (LR) the use of the ANN. Both simulations start from the same initial conditions as the model truth. The evolution of <inline-formula><mml:math id="M100" display="inline"><mml:mover accent="true"><mml:mtext>RMSE</mml:mtext><mml:mo mathvariant="normal">‾</mml:mo></mml:mover></mml:math></inline-formula> is presented in Fig. <xref ref-type="fig" rid="Ch1.F5"/>. The RMSEs corresponding to the regular case are higher than for the chaotic case. This is because the regular case exhibits a repeating pattern of long-lived, high-amplitude convective events. In comparison, the chaotic case produces short-lived perturbations with very small amplitude, leading to smaller climatological variability.</p>
      <p id="d1e2533">For both orographies, the ANN has a clear positive effect on the forecast until the error of LR saturates, after which the error of <inline-formula><mml:math id="M101" display="inline"><mml:mrow><mml:msub><mml:mtext>LR</mml:mtext><mml:mtext>ANN</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula> continues to grow. For the chaotic case, this leads to a detrimental impact of the ANN after about 1 d. Also, the <inline-formula><mml:math id="M102" display="inline"><mml:mrow><mml:msub><mml:mtext>SD</mml:mtext><mml:mtext>total</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula> of <inline-formula><mml:math id="M103" display="inline"><mml:mrow><mml:msub><mml:mtext>LR</mml:mtext><mml:mtext>ANN</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula> is rapidly exceeding that of LR. This is because, in contrast to LR, the shaded region for <inline-formula><mml:math id="M104" display="inline"><mml:mrow><mml:msub><mml:mtext>LR</mml:mtext><mml:mtext>ANN</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula> includes the variability due to the ANN realizations which significantly contributes to the total variability. This is seen in Fig. <xref ref-type="fig" rid="Ch1.F12"/> and discussed further in the next section.</p>
      <p id="d1e2582">It is not surprising that <inline-formula><mml:math id="M105" display="inline"><mml:mrow><mml:msub><mml:mtext>LR</mml:mtext><mml:mtext>ANN</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula> deteriorates as the forecast lead time increases, since the ANNs are not perfect (as opposed to the data they were trained on) and the resulting errors accumulate over time, leading to biases. This is clearly visible in Fig. <xref ref-type="fig" rid="Ch1.F6"/>, where it is seen that the SME of <inline-formula><mml:math id="M106" display="inline"><mml:mi>h</mml:mi></mml:math></inline-formula> and <inline-formula><mml:math id="M107" display="inline"><mml:mi>u</mml:mi></mml:math></inline-formula> diverge, which is in contrast to LR. The SME of <inline-formula><mml:math id="M108" display="inline"><mml:mi>r</mml:mi></mml:math></inline-formula> for LR is the result of a negative bias in the amount of rain produced (see Fig. <xref ref-type="fig" rid="Ch1.F4"/>), which is caused by the coarse graining of the orography. This bias is  significantly reduced by the ANNs. The divergence of the SME of <inline-formula><mml:math id="M109" display="inline"><mml:mi>h</mml:mi></mml:math></inline-formula> is the result of applying ANNs that, in contrast to the model, do not conserve mass. This leads to accumulated mass errors, causing biases in the wind field due to momentum conservation and a change in probability for the fluid to rise above <inline-formula><mml:math id="M110" display="inline"><mml:mrow><mml:msub><mml:mi>H</mml:mi><mml:mi mathvariant="normal">c</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M111" display="inline"><mml:mrow><mml:msub><mml:mi>H</mml:mi><mml:mi mathvariant="normal">r</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>. We therefore investigate if reducing the mass error, by adding a penalty term to the loss function of the ANN, can increase the forecast skill further.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F7" specific-use="star"><?xmltex \currentcnt{7}?><?xmltex \def\figurename{Figure}?><label>Figure 7</label><caption><p id="d1e2653">The <inline-formula><mml:math id="M112" display="inline"><mml:mover accent="true"><mml:mtext>RMSE</mml:mtext><mml:mo mathvariant="normal">‾</mml:mo></mml:mover></mml:math></inline-formula> <bold>(a, d, g)</bold>, MSE <bold>(b, e, h)</bold>, and bias <bold>(c, f, i)</bold> of the validation data for the different weightings (<inline-formula><mml:math id="M113" display="inline"><mml:mi>x</mml:mi></mml:math></inline-formula> axis) and the respective model variables (rows) for the regular case. Error bars indicate (from dark to light) <inline-formula><mml:math id="M114" display="inline"><mml:mrow><mml:msub><mml:mtext>SD</mml:mtext><mml:mtext>total</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula>, <inline-formula><mml:math id="M115" display="inline"><mml:mrow><mml:msub><mml:mtext>SD</mml:mtext><mml:mtext>time</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula>,  and  <inline-formula><mml:math id="M116" display="inline"><mml:mrow><mml:msub><mml:mtext>SD</mml:mtext><mml:mtext>ANN</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula>, and the red lines in the right panel indicate the zero line.</p></caption>
          <?xmltex \igopts{width=341.433071pt}?><graphic xlink:href="https://npg.copernicus.org/articles/29/171/2022/npg-29-171-2022-f07.png"/>

        </fig>

      <?xmltex \floatpos{t}?><fig id="Ch1.F8" specific-use="star"><?xmltex \currentcnt{8}?><?xmltex \def\figurename{Figure}?><label>Figure 8</label><caption><p id="d1e2724">Same as Fig. <xref ref-type="fig" rid="Ch1.F7"/> but for the chaotic case.</p></caption>
          <?xmltex \igopts{width=341.433071pt}?><graphic xlink:href="https://npg.copernicus.org/articles/29/171/2022/npg-29-171-2022-f08.png"/>

        </fig>

</sec>
<sec id="Ch1.S4.SS2">
  <label>4.2</label><title>ANN with mass conservation in a weak sense</title>
      <p id="d1e2743">Instead of including mass conservation in the training process of the ANN, it is natural to first try to correct the mass violation by post processing the ANN corrections. We tested two approaches, i.e., homogeneously subtracting the spatial mean of the <inline-formula><mml:math id="M117" display="inline"><mml:mi>h</mml:mi></mml:math></inline-formula> corrections and multiplying the vector of positive (negative) <inline-formula><mml:math id="M118" display="inline"><mml:mi>h</mml:mi></mml:math></inline-formula> corrections with the appropriate scalar when the mass violation is positive (negative). Neither of these simple approaches led to improvements. We, therefore, included mass conservation in a weak sense in the training process of the ANN, as described in Eq. (<xref ref-type="disp-formula" rid="Ch1.E1"/>). We trained ANNs with mass conservation weightings of <inline-formula><mml:math id="M119" display="inline"><mml:mrow><mml:msub><mml:mi>w</mml:mi><mml:mtext>mass</mml:mtext></mml:msub><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>,</mml:mo><mml:mn mathvariant="normal">10</mml:mn><mml:mo>,</mml:mo><mml:mn mathvariant="normal">100</mml:mn><mml:mo>,</mml:mo><mml:mn mathvariant="normal">1000</mml:mn></mml:mrow></mml:math></inline-formula>. These weightings result in a contribution to the loss function of roughly 0.2 %, 0.7 %, 2 %, and 5 % throughout the training process, respectively (not shown). Note that the ANNs presented in the previous section correspond to <inline-formula><mml:math id="M120" display="inline"><mml:mrow><mml:msub><mml:mi>w</mml:mi><mml:mtext>mass</mml:mtext></mml:msub><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0</mml:mn></mml:mrow></mml:math></inline-formula>.</p>
      <p id="d1e2804">Figures <xref ref-type="fig" rid="Ch1.F7"/> and <xref ref-type="fig" rid="Ch1.F8"/> show the single time step predictions for the regular and chaotic cases, respectively. Clearly, the mass conservation penalty term in the loss function has the desired effect of reducing the mass error for both orographies. Also, the error bars of the mass bias go down. A clear, convincing correlation between the reduction in SME and bias for <inline-formula><mml:math id="M121" display="inline"><mml:mi>h</mml:mi></mml:math></inline-formula> and any other field and/or metric is not detected, with the possible exception of the SME for <inline-formula><mml:math id="M122" display="inline"><mml:mi>u</mml:mi></mml:math></inline-formula> in the chaotic case. A tradeoff between increasing RMSE and decreasing MSE for increasing <inline-formula><mml:math id="M123" display="inline"><mml:mrow><mml:msub><mml:mi>w</mml:mi><mml:mtext>mass</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula> was expected but is not observed. The RMSE even tends to decrease a minimal amount for the chaotic case.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F9"><?xmltex \currentcnt{9}?><?xmltex \def\figurename{Figure}?><label>Figure 9</label><caption><p id="d1e2838">The <inline-formula><mml:math id="M124" display="inline"><mml:mover accent="true"><mml:mtext>RMSE</mml:mtext><mml:mo mathvariant="normal">‾</mml:mo></mml:mover></mml:math></inline-formula> evolution of 48 h forecasts for model variables <inline-formula><mml:math id="M125" display="inline"><mml:mi>h</mml:mi></mml:math></inline-formula>, <inline-formula><mml:math id="M126" display="inline"><mml:mi>u</mml:mi></mml:math></inline-formula>, and <inline-formula><mml:math id="M127" display="inline"><mml:mi>r</mml:mi></mml:math></inline-formula> (from top to bottom) of LR (black) and  <inline-formula><mml:math id="M128" display="inline"><mml:mrow><mml:msub><mml:mtext>LR</mml:mtext><mml:mtext>ANN</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula> for the different weightings (blue colors). </p></caption>
          <?xmltex \igopts{width=236.157874pt}?><graphic xlink:href="https://npg.copernicus.org/articles/29/171/2022/npg-29-171-2022-f09.png"/>

        </fig>

      <?xmltex \floatpos{t}?><fig id="Ch1.F10"><?xmltex \currentcnt{10}?><?xmltex \def\figurename{Figure}?><label>Figure 10</label><caption><p id="d1e2892">Same as Fig. <xref ref-type="fig" rid="Ch1.F9"/> but for the MSE.</p></caption>
          <?xmltex \igopts{width=236.157874pt}?><graphic xlink:href="https://npg.copernicus.org/articles/29/171/2022/npg-29-171-2022-f10.png"/>

        </fig>

      <?xmltex \floatpos{t}?><fig id="Ch1.F11"><?xmltex \currentcnt{11}?><?xmltex \def\figurename{Figure}?><label>Figure 11</label><caption><p id="d1e2905">Correlation of the different weighting (blue colors) between the bias of <inline-formula><mml:math id="M129" display="inline"><mml:mi>h</mml:mi></mml:math></inline-formula> and the bias of <inline-formula><mml:math id="M130" display="inline"><mml:mi>u</mml:mi></mml:math></inline-formula> <bold>(a, b)</bold> and <inline-formula><mml:math id="M131" display="inline"><mml:mi>r</mml:mi></mml:math></inline-formula> <bold>(c, d)</bold> for the regular <bold>(a, c)</bold> and the chaotic <bold>(b, d)</bold> cases. </p></caption>
          <?xmltex \igopts{width=236.157874pt}?><graphic xlink:href="https://npg.copernicus.org/articles/29/171/2022/npg-29-171-2022-f11.png"/>

        </fig>

      <p id="d1e2948">Figure <xref ref-type="fig" rid="Ch1.F9"/> presents the mean RMSE of the 48 h forecasts for all weightings. The weak mass conservation constraint has the desired effect on the forecast skill. For the chaotic case, more than 15 h in forecast quality is gained. For the regular case,  the number is unclear since the RMSE is still lower than LR and has not yet saturated after 48 h. However, we can say that it is at least 30 h. As hypothesized, Fig. <xref ref-type="fig" rid="Ch1.F10"/> indicates that the divergence of the domain mean error of the wind <inline-formula><mml:math id="M132" display="inline"><mml:mi>u</mml:mi></mml:math></inline-formula> is delayed as the weighting <inline-formula><mml:math id="M133" display="inline"><mml:mrow><mml:msub><mml:mi>w</mml:mi><mml:mtext>mass</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula> is increased. This, in turn, positively affects the domain mean of the rain <inline-formula><mml:math id="M134" display="inline"><mml:mi>r</mml:mi></mml:math></inline-formula>. To support these claims, we look at Fig. <xref ref-type="fig" rid="Ch1.F11"/>, which shows the correlation between the bias in <inline-formula><mml:math id="M135" display="inline"><mml:mi>h</mml:mi></mml:math></inline-formula> and the bias in wind <inline-formula><mml:math id="M136" display="inline"><mml:mi>u</mml:mi></mml:math></inline-formula> and rain <inline-formula><mml:math id="M137" display="inline"><mml:mi>r</mml:mi></mml:math></inline-formula>, respectively. In the single step predictions, these correlations were not conclusively detected. However, as the forecast evolves, the wind bias becomes almost completely anticorrelated to the mass bias. A strong correlation between the mass bias and the rain bias is also established after a few time steps, likely when the change in the probability of crossing the rain threshold resulting from the mass bias has taken effect. We also note that the larger  <inline-formula><mml:math id="M138" display="inline"><mml:mrow><mml:msub><mml:mi>w</mml:mi><mml:mtext>mass</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula>, the weaker the correlations. We hypothesize that, as the mass bias weakens, other causes for introducing domain mean biases in the wind and rain field become more significant. Such other causes may, for example, depend on the orography or the state of <inline-formula><mml:math id="M139" display="inline"><mml:mi>u</mml:mi></mml:math></inline-formula> and <inline-formula><mml:math id="M140" display="inline"><mml:mi>r</mml:mi></mml:math></inline-formula>.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F12"><?xmltex \currentcnt{12}?><?xmltex \def\figurename{Figure}?><label>Figure 12</label><caption><p id="d1e3032">Evolution of <inline-formula><mml:math id="M141" display="inline"><mml:mrow><mml:msub><mml:mtext>SD</mml:mtext><mml:mtext>total</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula>, <inline-formula><mml:math id="M142" display="inline"><mml:mrow><mml:msub><mml:mtext>SD</mml:mtext><mml:mtext>time</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula>, and  <inline-formula><mml:math id="M143" display="inline"><mml:mrow><mml:msub><mml:mtext>SD</mml:mtext><mml:mtext>ANN</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula> (solid, dotted, and dashed lines, respectively) for the different weightings (blue colors) for the regular <bold>(a, c, e)</bold> and the chaotic <bold>(b, d, f)</bold> cases.</p></caption>
          <?xmltex \igopts{width=236.157874pt}?><graphic xlink:href="https://npg.copernicus.org/articles/29/171/2022/npg-29-171-2022-f12.png"/>

        </fig>

      <p id="d1e3080">Next we look at the variability in the forecast errors in terms of <inline-formula><mml:math id="M144" display="inline"><mml:mrow><mml:msub><mml:mtext>SD</mml:mtext><mml:mtext>total</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula>, <inline-formula><mml:math id="M145" display="inline"><mml:mrow><mml:msub><mml:mtext>SD</mml:mtext><mml:mtext>time</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula>, and  <inline-formula><mml:math id="M146" display="inline"><mml:mrow><mml:msub><mml:mtext>SD</mml:mtext><mml:mtext>ANN</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula> in Fig. <xref ref-type="fig" rid="Ch1.F12"/>. For small weightings, the variability caused by ANN realizations dominates the total variability. However, as the weighting increases, the variability due to the initial conditions takes over. This again confirms the benefits of adding the mass penalty term to the loss function, as it demonstrates a decrease in the sensitivity of the forecast to the training process of the ANN.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F13"><?xmltex \currentcnt{13}?><?xmltex \def\figurename{Figure}?><label>Figure 13</label><caption><p id="d1e3121">Snapshot of the state variables for the chaotic case of a 6 h forecast starting from initial conditions of the validation dataset for the truth (red), LR (black), and <inline-formula><mml:math id="M147" display="inline"><mml:mrow><mml:msub><mml:mtext>LR</mml:mtext><mml:mtext>ANN</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula> corresponding to weightings <inline-formula><mml:math id="M148" display="inline"><mml:mrow><mml:msub><mml:mi>w</mml:mi><mml:mtext>mass</mml:mtext></mml:msub><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0</mml:mn></mml:mrow></mml:math></inline-formula> (light blue) and <inline-formula><mml:math id="M149" display="inline"><mml:mrow><mml:msub><mml:mi>w</mml:mi><mml:mtext>mass</mml:mtext></mml:msub><mml:mo>=</mml:mo><mml:mn mathvariant="normal">100</mml:mn></mml:mrow></mml:math></inline-formula> (dark blue). The dotted red lines are the convection threshold <inline-formula><mml:math id="M150" display="inline"><mml:mrow><mml:msub><mml:mi>H</mml:mi><mml:mi mathvariant="normal">c</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> and rain threshold <inline-formula><mml:math id="M151" display="inline"><mml:mrow><mml:msub><mml:mi>H</mml:mi><mml:mi mathvariant="normal">r</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>, respectively. </p></caption>
          <?xmltex \igopts{width=236.157874pt}?><graphic xlink:href="https://npg.copernicus.org/articles/29/171/2022/npg-29-171-2022-f13.png"/>

        </fig>

      <p id="d1e3193">A visual examination of the animations of the forecast evolution suggests that convective events produced in the LR run are wider and shallower than in the coarse-grained HR run. This behavior mimics the collapse of convective clouds towards the grid length that is typical of kilometer-scale numerical weather prediction models, as noted in the introduction. This then leads to a lack of rain mass but also, via conservation of momentum, a drift in the wind field. The convective events in the LR simulations are therefore also increasingly misplaced as the forecast lead time increases. The ANNs are capable of sharpening the gradients of the convective events, leading to highly accurate forecasts of convective events up to 6–12 h. After this, spurious, missed, and misplaced events start to occur, although the forecast skill remains significant a while longer, which is in contrast to the LR simulations where the forecast skill dissolves after just a few hours. A snapshot of the state for the chaotic case is presented in Fig. <xref ref-type="fig" rid="Ch1.F13"/>. The main rain event is misplaced for LR due to the bias in the wind field. Also, LR misses the small neighboring events which the <inline-formula><mml:math id="M152" display="inline"><mml:mrow><mml:msub><mml:mtext>LR</mml:mtext><mml:mtext>ANN</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula>s do catch. Furthermore, it is also clear that <inline-formula><mml:math id="M153" display="inline"><mml:mrow><mml:msub><mml:mtext>LR</mml:mtext><mml:mtext>ANN</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula> for <inline-formula><mml:math id="M154" display="inline"><mml:mrow><mml:msub><mml:mi>w</mml:mi><mml:mtext>mass</mml:mtext></mml:msub><mml:mo>=</mml:mo><mml:mn mathvariant="normal">100</mml:mn></mml:mrow></mml:math></inline-formula> is closer to the truth for all variables than for <inline-formula><mml:math id="M155" display="inline"><mml:mrow><mml:msub><mml:mi>w</mml:mi><mml:mtext>mass</mml:mtext></mml:msub><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0</mml:mn></mml:mrow></mml:math></inline-formula>.</p>
</sec>
</sec>
<sec id="Ch1.S5" sec-type="conclusions">
  <label>5</label><title>Conclusions</title>
      <p id="d1e3260">In this paper, we evaluated the feasibility of using an ANN to correct for model error in the gray zone, where important features occur on scales comparable to the model resolution. The model that was used in our idealized setup mimics key aspects of convection, such as conditional instability triggered by orography, and resulting convective events, including rain. As such, this model is representative of fully complex convective-scale numerical weather prediction models and, in particular, the corresponding errors due to unresolved scales in the gray zone. We considered two cases, each with a different realization of the orography, leading to two different regimes. We considered one where the convective events are large and long-lived and one where the convective events are small and short-lived. We refer to the former case as regular and the latter case as chaotic. We showed that the ANNs are capable of accurately sharpening gradients where necessary in both cases to prevent the missing and flattening of convective events that are caused by the low-resolution model's inability to resolve fine scales. For the regular case, the RMSE is still significantly lower than the low-resolution simulation (LR) after 48 h. For the chaotic case, the RMSE surpasses LR after about 1 d. Since the ANNs are not perfect, their errors accumulate over time, deteriorating the forecast skill. In particular, the accumulated mass error causes biases which are not present in LR because the model conserves mass exactly. We, therefore, investigated the effects of adding a term to the loss function of the ANN's training process to penalize mass conservation violation. We found that reducing the mass error reduces the biases in the wind and rain field, leading to further forecasts improvements. For the chaotic case, an additional 15 h in forecast lead time is gained before the RMSE exceeds the LR control simulation and at least 30 h for regular case. Such a positive effect of mass conservation was also found in, for example, <xref ref-type="bibr" rid="bib1.bibx36 bib1.bibx29 bib1.bibx30" id="text.22"/>. Furthermore, we showed that including the penalty term in the loss function reduces the sensitivity of the model forecasts to the learning process of the ANN, rendering the approach more robust.</p>
      <p id="d1e3266">While these results are encouraging, there are some issues to consider when applying this method to operational configurations. On a technical level, the generation of the training data and the training of the ANN can be costly and time consuming due to the requirement of sufficient HR data and the cumbersome exercise of tuning the ANN. The latter is a known problem that can be minimized through the clever iteration of tested ANN settings, but it cannot be fully avoided. Depending on the costs of generating HR data, using observations could be considered instead, as done by <xref ref-type="bibr" rid="bib1.bibx4" id="text.23"/>. They use data assimilation to generate HR data from available sparse and noisy observations. Aside from saving computational costs by replacing HR simulations with data assimilation, it might offer an advantage on a different issue as well, i.e., the effect of other model errors. In contrast to what was assumed in this paper, in reality not all model error stems from unresolved scales. By using observations of the true state of the atmosphere, all model error is accounted for by the trained ANN. On the other hand, the training data contain the errors inherited from data assimilation. It is not clear which error source is more important, and therefore, both approaches are worthy of investigation – not only to improve model forecasts but also to gain more insight in the model error itself and its comparison to errors stemming from data assimilation.</p>
</sec>

      
      </body>
    <back><notes notes-type="codeavailability"><title>Code availability</title>

      <p id="d1e3276">The provided source code (<ext-link xlink:href="https://doi.org/10.5281/zenodo.4740252" ext-link-type="DOI">10.5281/zenodo.4740252</ext-link>; <xref ref-type="bibr" rid="bib1.bibx28" id="altparen.24"/>) includes the necessary scripts to produce the data.</p>
  </notes><notes notes-type="authorcontribution"><title>Author contributions</title>

      <p id="d1e3288">RK produced the source code. RK and YR ran the experiments and visualized the results. SR provided expertise on neural networks. GC provided expertise on convective-scale dynamics. All authors contributed to the scientific design of the study, the analysis of the numerical results, and the writing of the paper.</p>
  </notes><notes notes-type="competinginterests"><title>Competing interests</title>

      <p id="d1e3294">The contact author has declared that neither they nor their co-authors have any competing interests.</p>
  </notes><notes notes-type="disclaimer"><title>Disclaimer</title>

      <p id="d1e3300">Publisher's note: Copernicus Publications remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.</p>
  </notes><notes notes-type="financialsupport"><title>Financial support</title>

      <p id="d1e3306">This research has been supported by the German Research Foundation (DFG; subproject B6 of the Transregional Collaborative Research Project SFB/TRR 165, “Waves to
Weather” and grant no. JA 1077/4-1).</p>
  </notes><notes notes-type="reviewstatement"><title>Review statement</title>

      <p id="d1e3312">This paper was edited by Stéphane Vannitsem and reviewed by Julien Brajard and Davide Faranda.</p>
  </notes><ref-list>
    <title>References</title>

      <ref id="bib1.bibx1"><?xmltex \def\ref@label{{Bocquet et~al.(2020)Bocquet, Brajard, Carrassi, and Bertino}}?><label>Bocquet et al.(2020)Bocquet, Brajard, Carrassi, and Bertino</label><?label Bocquet2020?><mixed-citation>Bocquet, M., Brajard, J., Carrassi, A., and Bertino, L.: Bayesian inference of chaotic dynamics by merging data assimilation, machine learning and expectation-maximization, Foundations of Data Science, 2, 55–80, <ext-link xlink:href="https://doi.org/10.3934/fods.2020004" ext-link-type="DOI">10.3934/fods.2020004</ext-link>, 2020.</mixed-citation></ref>
      <ref id="bib1.bibx2"><?xmltex \def\ref@label{{Bolton and Zanna(2019)}}?><label>Bolton and Zanna(2019)</label><?label Bolton_etal_2019?><mixed-citation>Bolton, T. and Zanna, L.: Applications of Deep Learning to Ocean Data Inference and Subgrid Parameterization, J. Adv. Model. Earth Sy., 11, 376–399, <ext-link xlink:href="https://doi.org/10.1029/2018MS001472" ext-link-type="DOI">10.1029/2018MS001472</ext-link>, 2019.</mixed-citation></ref>
      <ref id="bib1.bibx3"><?xmltex \def\ref@label{{Brajard et~al.(2020)Brajard, Carrassi, Bocquet, and Bertino}}?><label>Brajard et al.(2020)Brajard, Carrassi, Bocquet, and Bertino</label><?label brajardetal_2019?><mixed-citation>Brajard, J., Carrassi, A., Bocquet, M., and Bertino, L.: Combining data assimilation and machine learning to emulate a dynamical model from sparse and noisy observations: A case study with the Lorenz 96 model, J. Comput. Sci., 44, 101171, <ext-link xlink:href="https://doi.org/10.1016/j.jocs.2020.101171" ext-link-type="DOI">10.1016/j.jocs.2020.101171</ext-link>, 2020.</mixed-citation></ref>
      <ref id="bib1.bibx4"><?xmltex \def\ref@label{{Brajard et~al.(2021)Brajard, Carrassi, Bocquet, and Bertino}}?><label>Brajard et al.(2021)Brajard, Carrassi, Bocquet, and Bertino</label><?label brajard_etal2021b?><mixed-citation>Brajard, J., Carrassi, A., Bocquet, M., and Bertino, L.: Combining data assimilation and machine learning to infer unresolved scale parametrization, Philos. T. Roy. Soc. A, 379, 20200086, <ext-link xlink:href="https://doi.org/10.1098/rsta.2020.0086" ext-link-type="DOI">10.1098/rsta.2020.0086</ext-link>, 2021.</mixed-citation></ref>
      <ref id="bib1.bibx5"><?xmltex \def\ref@label{{Brenowitz and Bretherton(2019)}}?><label>Brenowitz and Bretherton(2019)</label><?label brenowitz_2019?><mixed-citation>Brenowitz, N. D. and Bretherton, C. S.: Spatially Extended Tests of a Neural Network Parametrization Trained by Coarse-Graining, J. Adv. Model. Earth Sy., 11, 2728–2744, <ext-link xlink:href="https://doi.org/10.1029/2019MS001711" ext-link-type="DOI">10.1029/2019MS001711</ext-link>, 2019.</mixed-citation></ref>
      <ref id="bib1.bibx6"><?xmltex \def\ref@label{{Brunton et~al.(2016)Brunton, Proctor, and Kutz}}?><label>Brunton et al.(2016)Brunton, Proctor, and Kutz</label><?label Brunton2016?><mixed-citation>Brunton, S. L., Proctor, J. L., and Kutz, J. N.: Discovering governing equations from data by sparse identification of nonlinear dynamical systems, P. Natl. Acad. Sci. USA, 113, 3932–3937, <ext-link xlink:href="https://doi.org/10.1073/pnas.1517384113" ext-link-type="DOI">10.1073/pnas.1517384113</ext-link>, 2016.</mixed-citation></ref>
      <ref id="bib1.bibx7"><?xmltex \def\ref@label{{Bryan et~al.(2003)Bryan, Wyngaard, and Fritsch}}?><label>Bryan et al.(2003)Bryan, Wyngaard, and Fritsch</label><?label Bryan_etal_2003?><mixed-citation>Bryan, G. H., Wyngaard, J. C., and Fritsch, J. M.: Resolution Requirements for the Simulation of Deep Moist Convection, Mon. Weather Rev., 131, 2394–2416, <ext-link xlink:href="https://doi.org/10.1175/1520-0493(2003)131&lt;2394:RRFTSO&gt;2.0.CO;2" ext-link-type="DOI">10.1175/1520-0493(2003)131&lt;2394:RRFTSO&gt;2.0.CO;2</ext-link>, 2003.</mixed-citation></ref>
      <ref id="bib1.bibx8"><?xmltex \def\ref@label{{Chollet(2017)}}?><label>Chollet(2017)</label><?label Chollet2015?><mixed-citation>
Chollet, F.: Deep Learning with Python, Manning Publications
Company, Greenwich, CT, USA, ISBN 9781617294433, 2017.</mixed-citation></ref>
      <ref id="bib1.bibx9"><?xmltex \def\ref@label{{Chow et~al.(2019)Chow, Sch\"{a}r, Ban, Lundquist, Schlemmer, and Shi}}?><label>Chow et al.(2019)Chow, Schär, Ban, Lundquist, Schlemmer, and Shi</label><?label Chow_etal_2019?><mixed-citation>Chow, F. K., Schär, C., Ban, N., Lundquist, K. A., Schlemmer, L., and Shi, X.: Crossing Multiple Gray Zones in the Transition from Mesoscale to Microscale Simulation over Complex Terrain, Atmosphere, 10, 274, <ext-link xlink:href="https://doi.org/10.3390/atmos10050274" ext-link-type="DOI">10.3390/atmos10050274</ext-link>, 2019.</mixed-citation></ref>
      <ref id="bib1.bibx10"><?xmltex \def\ref@label{{Dueben and Bauer(2018)}}?><label>Dueben and Bauer(2018)</label><?label dueben_etal_2018?><mixed-citation>Dueben, P. D. and Bauer, P.: Challenges and design choices for global weather and climate models based on machine learning, Geosci. Model Dev., 11, 3999–4009, <ext-link xlink:href="https://doi.org/10.5194/gmd-11-3999-2018" ext-link-type="DOI">10.5194/gmd-11-3999-2018</ext-link>, 2018.</mixed-citation></ref>
      <ref id="bib1.bibx11"><?xmltex \def\ref@label{{Fablet et~al.(2018)Fablet, Ouala, and Herzet}}?><label>Fablet et al.(2018)Fablet, Ouala, and Herzet</label><?label Fablet_etal_2018?><mixed-citation>Fablet, R., Ouala, S., and Herzet, C.: Bilinear Residual Neural Network for the Identification and Forecasting of Geophysical Dynamics, in: 2018 26th European Signal Processing Conference (EUSIPCO), 3–7 September 2018, Rome, Italy, 1477–1481, <ext-link xlink:href="https://doi.org/10.23919/EUSIPCO.2018.8553492" ext-link-type="DOI">10.23919/EUSIPCO.2018.8553492</ext-link>, 2018.</mixed-citation></ref>
      <ref id="bib1.bibx12"><?xmltex \def\ref@label{{Faranda et~al.(2018)Faranda, Lembo, Iyer, Kuzzay, Chibbaro, Daviaud, and Dubrulle}}?><label>Faranda et al.(2018)Faranda, Lembo, Iyer, Kuzzay, Chibbaro, Daviaud, and Dubrulle</label><?label Faranda2018?><mixed-citation>Faranda, D., Lembo, V., Iyer, M., Kuzzay, D., Chibbaro, S., Daviaud, F., and Dubrulle, B.: Computation and Characterization of Local Subfilter-Scale Energy Transfers in Atmospheric Flows, J. Atmos. Sci., 75, 2175–2186, <ext-link xlink:href="https://doi.org/10.1175/JAS-D-17-0114.1" ext-link-type="DOI">10.1175/JAS-D-17-0114.1</ext-link>, 2018.</mixed-citation></ref>
      <ref id="bib1.bibx13"><?xmltex \def\ref@label{{Faranda et~al.(2021)Faranda, Vrac, Yiou, Pons, Hamid, Carella, Ngoungue~Langue, Thao, and Gautard}}?><label>Faranda et al.(2021)Faranda, Vrac, Yiou, Pons, Hamid, Carella, Ngoungue Langue, Thao, and Gautard</label><?label Faranda_etal_2020?><mixed-citation>Faranda, D., Vrac, M., Yiou, P., Pons, F. M. E., Hamid, A., Carella, G., Ngoungue Langue, C., Thao, S., and Gautard, V.: Enhancing geophysical flow machine learning performance via scale separation, Nonlin. Processes Geophys., 28, 423–443, <ext-link xlink:href="https://doi.org/10.5194/npg-28-423-2021" ext-link-type="DOI">10.5194/npg-28-423-2021</ext-link>, 2021.</mixed-citation></ref>
      <ref id="bib1.bibx14"><?xmltex \def\ref@label{{Goodfellow et~al.(2016)Goodfellow, Bengio, and Courville}}?><label>Goodfellow et al.(2016)Goodfellow, Bengio, and Courville</label><?label Goodfellow-et-al-2016?><mixed-citation>Goodfellow, I., Bengio, Y., and Courville, A.: Deep Learning, MIT Press, <uri>http://www.deeplearningbook.org</uri>, ISBN: 9780262035613, 2016.</mixed-citation></ref>
      <ref id="bib1.bibx15"><?xmltex \def\ref@label{{Honnert et~al.(2020)Honnert, Efstathiou, Beare, Ito, Lock, Neggers, Plant, Shin, Tomassini, and Zhou}}?><label>Honnert et al.(2020)Honnert, Efstathiou, Beare, Ito, Lock, Neggers, Plant, Shin, Tomassini, and Zhou</label><?label Honnert_etal_2020?><mixed-citation>Honnert, R., Efstathiou, G. A., Beare, R. J., Ito, J., Lock, A., Neggers, R., Plant, R. S., Shin, H. H., Tomassini, L., and Zhou, B.: The Atmospheric Boundary Layer and the “Gray Zone” of Turbulence: A Critical Review, J. Geophys. Res.-Atmos., 125, e2019JD030317, <ext-link xlink:href="https://doi.org/10.1029/2019JD030317" ext-link-type="DOI">10.1029/2019JD030317</ext-link>, 2020.</mixed-citation></ref>
      <ref id="bib1.bibx16"><?xmltex \def\ref@label{{Jeworrek et~al.(2019)Jeworrek, West, and Stull}}?><label>Jeworrek et al.(2019)Jeworrek, West, and Stull</label><?label Jeworrek_etal_2019?><mixed-citation>Jeworrek, J., West, G., and Stull, R.: Evaluation of Cumulus and Microphysics Parameterizations in WRF across the Convective Gray Zone, Weather Forecast., 34, 1097–1115, <ext-link xlink:href="https://doi.org/10.1175/WAF-D-18-0178.1" ext-link-type="DOI">10.1175/WAF-D-18-0178.1</ext-link>, 2019.</mixed-citation></ref>
      <ref id="bib1.bibx17"><?xmltex \def\ref@label{{Jones and Randall(2011)}}?><label>Jones and Randall(2011)</label><?label Jones_etal_2011?><mixed-citation>Jones, T. R. and Randall, D. A.: Quantifying the limits of convective parameterizations, J. Geophys. Res.-Atmos., 116, D08210, <ext-link xlink:href="https://doi.org/10.1029/2010JD014913" ext-link-type="DOI">10.1029/2010JD014913</ext-link>, 2011.</mixed-citation></ref>
      <ref id="bib1.bibx18"><?xmltex \def\ref@label{{Kent et~al.(2017)Kent, Bokhove, and Tobias}}?><label>Kent et al.(2017)Kent, Bokhove, and Tobias</label><?label Kent_etal_2017?><mixed-citation>Kent, T., Bokhove, O., and Tobias, S.: Dynamics of an idealized fluid model for investigating convective-scale data assimilation, Tellus A, 69, 1369332, <ext-link xlink:href="https://doi.org/10.1080/16000870.2017.1369332" ext-link-type="DOI">10.1080/16000870.2017.1369332</ext-link>, 2017.</mixed-citation></ref>
      <ref id="bib1.bibx19"><?xmltex \def\ref@label{{Krasnopolsky et~al.(2013)Krasnopolsky, Fox-Rabinovitz, and Belochitski}}?><label>Krasnopolsky et al.(2013)Krasnopolsky, Fox-Rabinovitz, and Belochitski</label><?label Krasnopolsky2013?><mixed-citation>Krasnopolsky, V. M., Fox-Rabinovitz, M. S., and Belochitski, A. A.: Using ensemble of neural networks to learn stochastic convection parameterizations for climate and numerical weather prediction models from data simulated by a cloud resolving model, Adv. Artif. Neural Syst., 2013, 485913, <ext-link xlink:href="https://doi.org/10.1155/2013/485913" ext-link-type="DOI">10.1155/2013/485913</ext-link>, 2013.</mixed-citation></ref>
      <ref id="bib1.bibx20"><?xmltex \def\ref@label{{Lovejoy and Schertzer(2010)}}?><label>Lovejoy and Schertzer(2010)</label><?label Lovejoy2010?><mixed-citation>Lovejoy, S. and Schertzer, D.: Towards a new synthesis for atmospheric dynamics: Space–time cascades, Atmos. Res., 96, 1–52, <ext-link xlink:href="https://doi.org/10.1016/j.atmosres.2010.01.004" ext-link-type="DOI">10.1016/j.atmosres.2010.01.004</ext-link>, 2010.</mixed-citation></ref>
      <ref id="bib1.bibx21"><?xmltex \def\ref@label{{Marino et~al.(2013)Marino, Mininni, Rosenberg, and Pouquet}}?><label>Marino et al.(2013)Marino, Mininni, Rosenberg, and Pouquet</label><?label Marino2013?><mixed-citation>Marino, R., Mininni, P. D., Rosenberg, D., and Pouquet, A.: Inverse cascades in rotating stratified turbulence: Fast growth of large scales, EPL-Europhys. Lett., 102, 44006, <ext-link xlink:href="https://doi.org/10.1209/0295-5075/102/44006" ext-link-type="DOI">10.1209/0295-5075/102/44006</ext-link>, 2013.</mixed-citation></ref>
      <ref id="bib1.bibx22"><?xmltex \def\ref@label{{O'Gorman and Dwyer(2018)}}?><label>O'Gorman and Dwyer(2018)</label><?label OGorman_et_al2018?><mixed-citation>O'Gorman, P. A. and Dwyer, J. G.: Using Machine Learning to Parameterize Moist Convection: Potential for Modeling of Climate, Climate Change, and Extreme Events, J. Adv. Model. Earth Sy., 10, 2548–2563, <ext-link xlink:href="https://doi.org/10.1029/2018MS001351" ext-link-type="DOI">10.1029/2018MS001351</ext-link>, 2018.
</mixed-citation></ref><?xmltex \hack{\newpage}?>
      <ref id="bib1.bibx23"><?xmltex \def\ref@label{{Pathak et~al.(2018)Pathak, Hunt, Girvan, Lu, and Ott}}?><label>Pathak et al.(2018)Pathak, Hunt, Girvan, Lu, and Ott</label><?label Pathak2018?><mixed-citation>Pathak, J., Hunt, B., Girvan, M., Lu, Z., and Ott, E.: Model-Free Prediction of Large Spatiotemporally Chaotic Systems from Data: A Reservoir Computing Approach, Phys. Rev. Lett., 120, 024102, <ext-link xlink:href="https://doi.org/10.1103/PhysRevLett.120.024102" ext-link-type="DOI">10.1103/PhysRevLett.120.024102</ext-link>, 2018.</mixed-citation></ref>
      <ref id="bib1.bibx24"><?xmltex \def\ref@label{{Randall et~al.(2003)Randall, Khairoutdinov, Arakawa, and Grabowski}}?><label>Randall et al.(2003)Randall, Khairoutdinov, Arakawa, and Grabowski</label><?label Randall_etal_2003?><mixed-citation>Randall, D., Khairoutdinov, M., Arakawa, A., and Grabowski, W.: Breaking the Cloud Parameterization Deadlock, B. Am. Meteorol. Soc., 84, 1547–1564, <ext-link xlink:href="https://doi.org/10.1175/BAMS-84-11-1547" ext-link-type="DOI">10.1175/BAMS-84-11-1547</ext-link>, 2003.</mixed-citation></ref>
      <ref id="bib1.bibx25"><?xmltex \def\ref@label{{Randall(2013)}}?><label>Randall(2013)</label><?label Randall_2013?><mixed-citation>Randall, D. A.: Beyond deadlock, Geophys. Res. Lett., 40, 5970–5976, <ext-link xlink:href="https://doi.org/10.1002/2013GL057998" ext-link-type="DOI">10.1002/2013GL057998</ext-link>, 2013.</mixed-citation></ref>
      <ref id="bib1.bibx26"><?xmltex \def\ref@label{{Rasp(2020)}}?><label>Rasp(2020)</label><?label Rasp_2019?><mixed-citation>Rasp, S.: Coupled online learning as a way to tackle instabilities and biases in neural network parameterizations: general algorithms and Lorenz 96 case study (v1.0), Geosci. Model Dev., 13, 2185–2196, <ext-link xlink:href="https://doi.org/10.5194/gmd-13-2185-2020" ext-link-type="DOI">10.5194/gmd-13-2185-2020</ext-link>, 2020.</mixed-citation></ref>
      <ref id="bib1.bibx27"><?xmltex \def\ref@label{{Rasp et~al.(2018)Rasp, Pritchard, and Gentine}}?><label>Rasp et al.(2018)Rasp, Pritchard, and Gentine</label><?label Rasp_2018?><mixed-citation>Rasp, S., Pritchard, M. S., and Gentine, P.: Deep learning to represent subgrid processes in climate models, P. Natl. Acad. Sci. USA, 115, 9684–9689, <ext-link xlink:href="https://doi.org/10.1073/pnas.1810286115" ext-link-type="DOI">10.1073/pnas.1810286115</ext-link>, 2018.</mixed-citation></ref>
      <ref id="bib1.bibx28"><?xmltex \def\ref@label{{Kriegmair(2021)}}?><label>Kriegmair(2021)</label><?label ruckstuhl2021?><mixed-citation>Kriegmair, R.:  wavestoweather/NNforModelError: NN to improve simulations in the grey zone (v1.0.0), Zenodo [code],  <ext-link xlink:href="https://doi.org/10.5281/zenodo.4740252" ext-link-type="DOI">10.5281/zenodo.4740252</ext-link>, 2021.</mixed-citation></ref>
      <ref id="bib1.bibx29"><?xmltex \def\ref@label{{Ruckstuhl and Janji{\'{c}}(2018)}}?><label>Ruckstuhl and Janjić(2018)</label><?label Ruckstuhl2018?><mixed-citation>Ruckstuhl, Y. and Janjić, T.: Parameter and state estimation with ensemble Kalman filter based algorithms for convective-scale applications, Q. J. Roy. Meteor. Soc, 144, 826–841, <ext-link xlink:href="https://doi.org/10.1002/qj.3257" ext-link-type="DOI">10.1002/qj.3257</ext-link>, 2018.</mixed-citation></ref>
      <ref id="bib1.bibx30"><?xmltex \def\ref@label{{Ruckstuhl et~al.(2021)Ruckstuhl, Janji\'{c}, and Rasp}}?><label>Ruckstuhl et al.(2021)Ruckstuhl, Janjić, and Rasp</label><?label ruckstuhl_etal_2021?><mixed-citation>Ruckstuhl, Y., Janjić, T., and Rasp, S.: Training a convolutional neural network to conserve mass in data assimilation, Nonlin. Processes Geophys., 28, 111–119, <ext-link xlink:href="https://doi.org/10.5194/npg-28-111-2021" ext-link-type="DOI">10.5194/npg-28-111-2021</ext-link>, 2021.</mixed-citation></ref>
      <ref id="bib1.bibx31"><?xmltex \def\ref@label{{Scher(2018)}}?><label>Scher(2018)</label><?label Sher2018?><mixed-citation>Scher, S.: Toward Data-Driven Weather and Climate Forecasting: Approximating a Simple General Circulation Model With Deep Learning, Geophys. Res. Lett., 45, 12616–12622, <ext-link xlink:href="https://doi.org/10.1029/2018GL080704" ext-link-type="DOI">10.1029/2018GL080704</ext-link>, 2018.</mixed-citation></ref>
      <ref id="bib1.bibx32"><?xmltex \def\ref@label{{Stensrud(2009)}}?><label>Stensrud(2009)</label><?label stensrud_2009?><mixed-citation>
Stensrud, D.: Parameterization Schemes: Keys to Understanding Numerical Weather Prediction Models, Cambridge University Press, ISBN 978-0-521-12676-2, 2009.</mixed-citation></ref>
      <ref id="bib1.bibx33"><?xmltex \def\ref@label{{Wagner et~al.(2018)Wagner, Heinzeller, Wagner, Rummler, and Kunstmann}}?><label>Wagner et al.(2018)Wagner, Heinzeller, Wagner, Rummler, and Kunstmann</label><?label Wagner_etal_2018?><mixed-citation>Wagner, A., Heinzeller, D., Wagner, S., Rummler, T., and Kunstmann, H.: Explicit Convection and Scale-Aware Cumulus Parameterizations: High-Resolution Simulations over Areas of Different Topography in Germany, Mon. Weather Rev., 146, 1925–1944, <ext-link xlink:href="https://doi.org/10.1175/MWR-D-17-0238.1" ext-link-type="DOI">10.1175/MWR-D-17-0238.1</ext-link>, 2018.</mixed-citation></ref>
      <ref id="bib1.bibx34"><?xmltex \def\ref@label{{W{\"{u}}rsch and Craig(2014)}}?><label>Würsch and Craig(2014)</label><?label wursch_etal_2014?><mixed-citation>Würsch, M. and Craig, G. C.: A simple dynamical model of cumulus convection for data assimilation research, Meteorol. Z., 23, 483–490, <ext-link xlink:href="https://doi.org/10.1127/0941-2948/2014/0492" ext-link-type="DOI">10.1127/0941-2948/2014/0492</ext-link>, 2014.</mixed-citation></ref>
      <ref id="bib1.bibx35"><?xmltex \def\ref@label{{Yuval and O'Gorman(2020)}}?><label>Yuval and O'Gorman(2020)</label><?label Yuval2020?><mixed-citation>
Yuval, J. and O'Gorman, P. A.: Stable machine-learning parameterization of subgrid processes for climate modeling at a range of resolutions, Nat. Commun., 11, 1–10, 2020.</mixed-citation></ref>
      <ref id="bib1.bibx36"><?xmltex \def\ref@label{{Zeng et~al.(2017)Zeng, Janji{\'{c}}, Ruckstuhl, and Verlaan}}?><label>Zeng et al.(2017)Zeng, Janjić, Ruckstuhl, and Verlaan</label><?label Zeng2017?><mixed-citation>Zeng, Y., Janjić, T., Ruckstuhl, Y., and Verlaan, M.: Ensemble-type Kalman filter algorithm conserving mass, total energy and enstrophy, Q. J. Roy. Meteor. Soc., 143, 2902–2914, <ext-link xlink:href="https://doi.org/10.1002/qj.3142" ext-link-type="DOI">10.1002/qj.3142</ext-link>, 2017.</mixed-citation></ref>

  </ref-list></back>
    <!--<article-title-html>Using neural networks to improve simulations in the gray zone</article-title-html>
<abstract-html/>
<ref-html id="bib1.bib1"><label>Bocquet et al.(2020)Bocquet, Brajard, Carrassi, and Bertino</label><mixed-citation>
Bocquet, M., Brajard, J., Carrassi, A., and Bertino, L.: Bayesian inference of chaotic dynamics by merging data assimilation, machine learning and expectation-maximization, Foundations of Data Science, 2, 55–80, <a href="https://doi.org/10.3934/fods.2020004" target="_blank">https://doi.org/10.3934/fods.2020004</a>, 2020.
</mixed-citation></ref-html>
<ref-html id="bib1.bib2"><label>Bolton and Zanna(2019)</label><mixed-citation>
Bolton, T. and Zanna, L.: Applications of Deep Learning to Ocean Data Inference and Subgrid Parameterization, J. Adv. Model. Earth Sy., 11, 376–399, <a href="https://doi.org/10.1029/2018MS001472" target="_blank">https://doi.org/10.1029/2018MS001472</a>, 2019.
</mixed-citation></ref-html>
<ref-html id="bib1.bib3"><label>Brajard et al.(2020)Brajard, Carrassi, Bocquet, and Bertino</label><mixed-citation>
Brajard, J., Carrassi, A., Bocquet, M., and Bertino, L.: Combining data assimilation and machine learning to emulate a dynamical model from sparse and noisy observations: A case study with the Lorenz 96 model, J. Comput. Sci., 44, 101171, <a href="https://doi.org/10.1016/j.jocs.2020.101171" target="_blank">https://doi.org/10.1016/j.jocs.2020.101171</a>, 2020.
</mixed-citation></ref-html>
<ref-html id="bib1.bib4"><label>Brajard et al.(2021)Brajard, Carrassi, Bocquet, and Bertino</label><mixed-citation>
Brajard, J., Carrassi, A., Bocquet, M., and Bertino, L.: Combining data assimilation and machine learning to infer unresolved scale parametrization, Philos. T. Roy. Soc. A, 379, 20200086, <a href="https://doi.org/10.1098/rsta.2020.0086" target="_blank">https://doi.org/10.1098/rsta.2020.0086</a>, 2021.
</mixed-citation></ref-html>
<ref-html id="bib1.bib5"><label>Brenowitz and Bretherton(2019)</label><mixed-citation>
Brenowitz, N. D. and Bretherton, C. S.: Spatially Extended Tests of a Neural Network Parametrization Trained by Coarse-Graining, J. Adv. Model. Earth Sy., 11, 2728–2744, <a href="https://doi.org/10.1029/2019MS001711" target="_blank">https://doi.org/10.1029/2019MS001711</a>, 2019.
</mixed-citation></ref-html>
<ref-html id="bib1.bib6"><label>Brunton et al.(2016)Brunton, Proctor, and Kutz</label><mixed-citation>
Brunton, S. L., Proctor, J. L., and Kutz, J. N.: Discovering governing equations from data by sparse identification of nonlinear dynamical systems, P. Natl. Acad. Sci. USA, 113, 3932–3937, <a href="https://doi.org/10.1073/pnas.1517384113" target="_blank">https://doi.org/10.1073/pnas.1517384113</a>, 2016.
</mixed-citation></ref-html>
<ref-html id="bib1.bib7"><label>Bryan et al.(2003)Bryan, Wyngaard, and Fritsch</label><mixed-citation>
Bryan, G. H., Wyngaard, J. C., and Fritsch, J. M.: Resolution Requirements for the Simulation of Deep Moist Convection, Mon. Weather Rev., 131, 2394–2416, <a href="https://doi.org/10.1175/1520-0493(2003)131&lt;2394:RRFTSO&gt;2.0.CO;2" target="_blank">https://doi.org/10.1175/1520-0493(2003)131&lt;2394:RRFTSO&gt;2.0.CO;2</a>, 2003.
</mixed-citation></ref-html>
<ref-html id="bib1.bib8"><label>Chollet(2017)</label><mixed-citation>
Chollet, F.: Deep Learning with Python, Manning Publications
Company, Greenwich, CT, USA, ISBN 9781617294433, 2017.
</mixed-citation></ref-html>
<ref-html id="bib1.bib9"><label>Chow et al.(2019)Chow, Schär, Ban, Lundquist, Schlemmer, and Shi</label><mixed-citation>
Chow, F. K., Schär, C., Ban, N., Lundquist, K. A., Schlemmer, L., and Shi, X.: Crossing Multiple Gray Zones in the Transition from Mesoscale to Microscale Simulation over Complex Terrain, Atmosphere, 10, 274, <a href="https://doi.org/10.3390/atmos10050274" target="_blank">https://doi.org/10.3390/atmos10050274</a>, 2019.
</mixed-citation></ref-html>
<ref-html id="bib1.bib10"><label>Dueben and Bauer(2018)</label><mixed-citation>
Dueben, P. D. and Bauer, P.: Challenges and design choices for global weather and climate models based on machine learning, Geosci. Model Dev., 11, 3999–4009, <a href="https://doi.org/10.5194/gmd-11-3999-2018" target="_blank">https://doi.org/10.5194/gmd-11-3999-2018</a>, 2018.
</mixed-citation></ref-html>
<ref-html id="bib1.bib11"><label>Fablet et al.(2018)Fablet, Ouala, and Herzet</label><mixed-citation>
Fablet, R., Ouala, S., and Herzet, C.: Bilinear Residual Neural Network for the Identification and Forecasting of Geophysical Dynamics, in: 2018 26th European Signal Processing Conference (EUSIPCO), 3–7 September 2018, Rome, Italy, 1477–1481, <a href="https://doi.org/10.23919/EUSIPCO.2018.8553492" target="_blank">https://doi.org/10.23919/EUSIPCO.2018.8553492</a>, 2018.
</mixed-citation></ref-html>
<ref-html id="bib1.bib12"><label>Faranda et al.(2018)Faranda, Lembo, Iyer, Kuzzay, Chibbaro, Daviaud, and Dubrulle</label><mixed-citation>
Faranda, D., Lembo, V., Iyer, M., Kuzzay, D., Chibbaro, S., Daviaud, F., and Dubrulle, B.: Computation and Characterization of Local Subfilter-Scale Energy Transfers in Atmospheric Flows, J. Atmos. Sci., 75, 2175–2186, <a href="https://doi.org/10.1175/JAS-D-17-0114.1" target="_blank">https://doi.org/10.1175/JAS-D-17-0114.1</a>, 2018.
</mixed-citation></ref-html>
<ref-html id="bib1.bib13"><label>Faranda et al.(2021)Faranda, Vrac, Yiou, Pons, Hamid, Carella, Ngoungue Langue, Thao, and Gautard</label><mixed-citation>
Faranda, D., Vrac, M., Yiou, P., Pons, F. M. E., Hamid, A., Carella, G., Ngoungue Langue, C., Thao, S., and Gautard, V.: Enhancing geophysical flow machine learning performance via scale separation, Nonlin. Processes Geophys., 28, 423–443, <a href="https://doi.org/10.5194/npg-28-423-2021" target="_blank">https://doi.org/10.5194/npg-28-423-2021</a>, 2021.
</mixed-citation></ref-html>
<ref-html id="bib1.bib14"><label>Goodfellow et al.(2016)Goodfellow, Bengio, and Courville</label><mixed-citation>
Goodfellow, I., Bengio, Y., and Courville, A.: Deep Learning, MIT Press, <a href="http://www.deeplearningbook.org" target="_blank"/>, ISBN: 9780262035613, 2016.
</mixed-citation></ref-html>
<ref-html id="bib1.bib15"><label>Honnert et al.(2020)Honnert, Efstathiou, Beare, Ito, Lock, Neggers, Plant, Shin, Tomassini, and Zhou</label><mixed-citation>
Honnert, R., Efstathiou, G. A., Beare, R. J., Ito, J., Lock, A., Neggers, R., Plant, R. S., Shin, H. H., Tomassini, L., and Zhou, B.: The Atmospheric Boundary Layer and the “Gray Zone” of Turbulence: A Critical Review, J. Geophys. Res.-Atmos., 125, e2019JD030317, <a href="https://doi.org/10.1029/2019JD030317" target="_blank">https://doi.org/10.1029/2019JD030317</a>, 2020.
</mixed-citation></ref-html>
<ref-html id="bib1.bib16"><label>Jeworrek et al.(2019)Jeworrek, West, and Stull</label><mixed-citation>
Jeworrek, J., West, G., and Stull, R.: Evaluation of Cumulus and Microphysics Parameterizations in WRF across the Convective Gray Zone, Weather Forecast., 34, 1097–1115, <a href="https://doi.org/10.1175/WAF-D-18-0178.1" target="_blank">https://doi.org/10.1175/WAF-D-18-0178.1</a>, 2019.
</mixed-citation></ref-html>
<ref-html id="bib1.bib17"><label>Jones and Randall(2011)</label><mixed-citation>
Jones, T. R. and Randall, D. A.: Quantifying the limits of convective parameterizations, J. Geophys. Res.-Atmos., 116, D08210, <a href="https://doi.org/10.1029/2010JD014913" target="_blank">https://doi.org/10.1029/2010JD014913</a>, 2011.
</mixed-citation></ref-html>
<ref-html id="bib1.bib18"><label>Kent et al.(2017)Kent, Bokhove, and Tobias</label><mixed-citation>
Kent, T., Bokhove, O., and Tobias, S.: Dynamics of an idealized fluid model for investigating convective-scale data assimilation, Tellus A, 69, 1369332, <a href="https://doi.org/10.1080/16000870.2017.1369332" target="_blank">https://doi.org/10.1080/16000870.2017.1369332</a>, 2017.
</mixed-citation></ref-html>
<ref-html id="bib1.bib19"><label>Krasnopolsky et al.(2013)Krasnopolsky, Fox-Rabinovitz, and Belochitski</label><mixed-citation>
Krasnopolsky, V. M., Fox-Rabinovitz, M. S., and Belochitski, A. A.: Using ensemble of neural networks to learn stochastic convection parameterizations for climate and numerical weather prediction models from data simulated by a cloud resolving model, Adv. Artif. Neural Syst., 2013, 485913, <a href="https://doi.org/10.1155/2013/485913" target="_blank">https://doi.org/10.1155/2013/485913</a>, 2013.
</mixed-citation></ref-html>
<ref-html id="bib1.bib20"><label>Lovejoy and Schertzer(2010)</label><mixed-citation>
Lovejoy, S. and Schertzer, D.: Towards a new synthesis for atmospheric dynamics: Space–time cascades, Atmos. Res., 96, 1–52, <a href="https://doi.org/10.1016/j.atmosres.2010.01.004" target="_blank">https://doi.org/10.1016/j.atmosres.2010.01.004</a>, 2010.
</mixed-citation></ref-html>
<ref-html id="bib1.bib21"><label>Marino et al.(2013)Marino, Mininni, Rosenberg, and Pouquet</label><mixed-citation>
Marino, R., Mininni, P. D., Rosenberg, D., and Pouquet, A.: Inverse cascades in rotating stratified turbulence: Fast growth of large scales, EPL-Europhys. Lett., 102, 44006, <a href="https://doi.org/10.1209/0295-5075/102/44006" target="_blank">https://doi.org/10.1209/0295-5075/102/44006</a>, 2013.
</mixed-citation></ref-html>
<ref-html id="bib1.bib22"><label>O'Gorman and Dwyer(2018)</label><mixed-citation>
O'Gorman, P. A. and Dwyer, J. G.: Using Machine Learning to Parameterize Moist Convection: Potential for Modeling of Climate, Climate Change, and Extreme Events, J. Adv. Model. Earth Sy., 10, 2548–2563, <a href="https://doi.org/10.1029/2018MS001351" target="_blank">https://doi.org/10.1029/2018MS001351</a>, 2018.

</mixed-citation></ref-html>
<ref-html id="bib1.bib23"><label>Pathak et al.(2018)Pathak, Hunt, Girvan, Lu, and Ott</label><mixed-citation>
Pathak, J., Hunt, B., Girvan, M., Lu, Z., and Ott, E.: Model-Free Prediction of Large Spatiotemporally Chaotic Systems from Data: A Reservoir Computing Approach, Phys. Rev. Lett., 120, 024102, <a href="https://doi.org/10.1103/PhysRevLett.120.024102" target="_blank">https://doi.org/10.1103/PhysRevLett.120.024102</a>, 2018.
</mixed-citation></ref-html>
<ref-html id="bib1.bib24"><label>Randall et al.(2003)Randall, Khairoutdinov, Arakawa, and Grabowski</label><mixed-citation>
Randall, D., Khairoutdinov, M., Arakawa, A., and Grabowski, W.: Breaking the Cloud Parameterization Deadlock, B. Am. Meteorol. Soc., 84, 1547–1564, <a href="https://doi.org/10.1175/BAMS-84-11-1547" target="_blank">https://doi.org/10.1175/BAMS-84-11-1547</a>, 2003.
</mixed-citation></ref-html>
<ref-html id="bib1.bib25"><label>Randall(2013)</label><mixed-citation>
Randall, D. A.: Beyond deadlock, Geophys. Res. Lett., 40, 5970–5976, <a href="https://doi.org/10.1002/2013GL057998" target="_blank">https://doi.org/10.1002/2013GL057998</a>, 2013.
</mixed-citation></ref-html>
<ref-html id="bib1.bib26"><label>Rasp(2020)</label><mixed-citation>
Rasp, S.: Coupled online learning as a way to tackle instabilities and biases in neural network parameterizations: general algorithms and Lorenz 96 case study (v1.0), Geosci. Model Dev., 13, 2185–2196, <a href="https://doi.org/10.5194/gmd-13-2185-2020" target="_blank">https://doi.org/10.5194/gmd-13-2185-2020</a>, 2020.
</mixed-citation></ref-html>
<ref-html id="bib1.bib27"><label>Rasp et al.(2018)Rasp, Pritchard, and Gentine</label><mixed-citation>
Rasp, S., Pritchard, M. S., and Gentine, P.: Deep learning to represent subgrid processes in climate models, P. Natl. Acad. Sci. USA, 115, 9684–9689, <a href="https://doi.org/10.1073/pnas.1810286115" target="_blank">https://doi.org/10.1073/pnas.1810286115</a>, 2018.
</mixed-citation></ref-html>
<ref-html id="bib1.bib28"><label>Kriegmair(2021)</label><mixed-citation>
Kriegmair, R.:  wavestoweather/NNforModelError: NN to improve simulations in the grey zone (v1.0.0), Zenodo [code],  <a href="https://doi.org/10.5281/zenodo.4740252" target="_blank">https://doi.org/10.5281/zenodo.4740252</a>, 2021.
</mixed-citation></ref-html>
<ref-html id="bib1.bib29"><label>Ruckstuhl and Janjić(2018)</label><mixed-citation>
Ruckstuhl, Y. and Janjić, T.: Parameter and state estimation with ensemble Kalman filter based algorithms for convective-scale applications, Q. J. Roy. Meteor. Soc, 144, 826–841, <a href="https://doi.org/10.1002/qj.3257" target="_blank">https://doi.org/10.1002/qj.3257</a>, 2018.
</mixed-citation></ref-html>
<ref-html id="bib1.bib30"><label>Ruckstuhl et al.(2021)Ruckstuhl, Janjić, and Rasp</label><mixed-citation>
Ruckstuhl, Y., Janjić, T., and Rasp, S.: Training a convolutional neural network to conserve mass in data assimilation, Nonlin. Processes Geophys., 28, 111–119, <a href="https://doi.org/10.5194/npg-28-111-2021" target="_blank">https://doi.org/10.5194/npg-28-111-2021</a>, 2021.
</mixed-citation></ref-html>
<ref-html id="bib1.bib31"><label>Scher(2018)</label><mixed-citation>
Scher, S.: Toward Data-Driven Weather and Climate Forecasting: Approximating a Simple General Circulation Model With Deep Learning, Geophys. Res. Lett., 45, 12616–12622, <a href="https://doi.org/10.1029/2018GL080704" target="_blank">https://doi.org/10.1029/2018GL080704</a>, 2018.
</mixed-citation></ref-html>
<ref-html id="bib1.bib32"><label>Stensrud(2009)</label><mixed-citation>
Stensrud, D.: Parameterization Schemes: Keys to Understanding Numerical Weather Prediction Models, Cambridge University Press, ISBN 978-0-521-12676-2, 2009.
</mixed-citation></ref-html>
<ref-html id="bib1.bib33"><label>Wagner et al.(2018)Wagner, Heinzeller, Wagner, Rummler, and Kunstmann</label><mixed-citation>
Wagner, A., Heinzeller, D., Wagner, S., Rummler, T., and Kunstmann, H.: Explicit Convection and Scale-Aware Cumulus Parameterizations: High-Resolution Simulations over Areas of Different Topography in Germany, Mon. Weather Rev., 146, 1925–1944, <a href="https://doi.org/10.1175/MWR-D-17-0238.1" target="_blank">https://doi.org/10.1175/MWR-D-17-0238.1</a>, 2018.
</mixed-citation></ref-html>
<ref-html id="bib1.bib34"><label>Würsch and Craig(2014)</label><mixed-citation>
Würsch, M. and Craig, G. C.: A simple dynamical model of cumulus convection for data assimilation research, Meteorol. Z., 23, 483–490, <a href="https://doi.org/10.1127/0941-2948/2014/0492" target="_blank">https://doi.org/10.1127/0941-2948/2014/0492</a>, 2014.
</mixed-citation></ref-html>
<ref-html id="bib1.bib35"><label>Yuval and O'Gorman(2020)</label><mixed-citation>
Yuval, J. and O'Gorman, P. A.: Stable machine-learning parameterization of subgrid processes for climate modeling at a range of resolutions, Nat. Commun., 11, 1–10, 2020.
</mixed-citation></ref-html>
<ref-html id="bib1.bib36"><label>Zeng et al.(2017)Zeng, Janjić, Ruckstuhl, and Verlaan</label><mixed-citation>
Zeng, Y., Janjić, T., Ruckstuhl, Y., and Verlaan, M.: Ensemble-type Kalman filter algorithm conserving mass, total energy and enstrophy, Q. J. Roy. Meteor. Soc., 143, 2902–2914, <a href="https://doi.org/10.1002/qj.3142" target="_blank">https://doi.org/10.1002/qj.3142</a>, 2017.
</mixed-citation></ref-html>--></article>
