<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing with OASIS Tables v3.0 20080202//EN" "journalpub-oasis3.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:oasis="http://docs.oasis-open.org/ns/oasis-exchange/table" dtd-version="3.0"><?xmltex \makeatother\@nolinetrue\makeatletter?>
  <front>
    <journal-meta>
<journal-id journal-id-type="publisher">NPG</journal-id>
<journal-title-group>
<journal-title>Nonlinear Processes in Geophysics</journal-title>
<abbrev-journal-title abbrev-type="publisher">NPG</abbrev-journal-title>
<abbrev-journal-title abbrev-type="nlm-ta">Nonlin. Processes Geophys.</abbrev-journal-title>
</journal-title-group>
<issn pub-type="epub">1607-7946</issn>
<publisher><publisher-name>Copernicus Publications</publisher-name>
<publisher-loc>Göttingen, Germany</publisher-loc>
</publisher>
</journal-meta>

    <article-meta>
      <article-id pub-id-type="doi">10.5194/npg-24-351-2017</article-id><title-group><article-title>Controllability, not chaos, key criterion for ocean state estimation</article-title>
      </title-group><?xmltex \runningtitle{Controllability, not chaos}?><?xmltex \runningauthor{G.~Gebbie and T.-L.~Hsieh}?>
      <contrib-group>
        <contrib contrib-type="author" corresp="yes" rid="aff1">
          <name><surname>Gebbie</surname><given-names>Geoffrey</given-names></name>
          <email>ggebbie@whoi.edu</email>
        </contrib>
        <contrib contrib-type="author" corresp="no" rid="aff1 aff2 aff3">
          <name><surname>Hsieh</surname><given-names>Tsung-Lin</given-names></name>
          
        </contrib>
        <aff id="aff1"><label>1</label><institution>Department of Physical Oceanography, Woods Hole Oceanographic Institution, Woods Hole, MA, USA</institution>
        </aff>
        <aff id="aff2"><label>2</label><institution>Summer Student Fellow, Woods Hole Oceanographic Institution, Woods Hole, MA, USA</institution>
        </aff>
        <aff id="aff3"><label>3</label><institution>Program in Atmospheric and Oceanic Sciences, Princeton University, Princeton, NJ, USA</institution>
        </aff>
      </contrib-group>
      <author-notes><corresp id="corr1">Geoffrey Gebbie (ggebbie@whoi.edu)</corresp></author-notes><pub-date><day>19</day><month>July</month><year>2017</year></pub-date>
      
      <volume>24</volume>
      <issue>3</issue>
      <fpage>351</fpage><lpage>366</lpage>
      <history>
        <date date-type="received"><day>28</day><month>September</month><year>2016</year></date>
           <date date-type="rev-request"><day>30</day><month>September</month><year>2016</year></date>
           <date date-type="rev-recd"><day>29</day><month>April</month><year>2017</year></date>
           <date date-type="accepted"><day>5</day><month>June</month><year>2017</year></date>
      </history>
      <permissions>
<license license-type="open-access">
<license-p>This work is licensed under the Creative Commons Attribution 3.0 Unported License. To view a copy of this licence, visit <ext-link ext-link-type="uri" xlink:href="https://creativecommons.org/licenses/by/3.0/">https://creativecommons.org/licenses/by/3.0/</ext-link></license-p>
</license>
</permissions><self-uri xlink:href="https://npg.copernicus.org/articles/24/351/2017/npg-24-351-2017.html">This article is available from https://npg.copernicus.org/articles/24/351/2017/npg-24-351-2017.html</self-uri>
<self-uri xlink:href="https://npg.copernicus.org/articles/24/351/2017/npg-24-351-2017.pdf">The full text article is available as a PDF file from https://npg.copernicus.org/articles/24/351/2017/npg-24-351-2017.pdf</self-uri>


      <abstract>
    <p>The Lagrange multiplier method for combining observations and
models (i.e., the adjoint method or “4D-VAR”) has been avoided or
approximated when the numerical model is highly nonlinear or chaotic. This
approach has been adopted primarily due to difficulties in the initialization
of low-dimensional chaotic models, where the search for optimal initial
conditions by gradient-descent algorithms is hampered by multiple local
minima. Although initialization is an important task for numerical weather
prediction, ocean state estimation usually demands an additional task – a
solution of the time-dependent surface boundary conditions that result from
atmosphere–ocean interaction. Here, we apply the Lagrange multiplier method
to an analogous boundary control problem, tracking the trajectory of the
forced chaotic pendulum. Contrary to previous assertions, it is demonstrated
that the Lagrange multiplier method can track multiple chaotic transitions
through time, so long as the boundary conditions render the system
controllable. Thus, the nonlinear timescale poses no limit to the time
interval for successful Lagrange multiplier-based estimation. That the key
criterion is controllability, not a pure measure of dynamical stability or
chaos, illustrates the similarities between the Lagrange multiplier method
and other state estimation methods. The results with the chaotic pendulum
suggest that nonlinearity should not be a fundamental obstacle to ocean state
estimation with eddy-resolving models, especially when using an improved
first-guess trajectory.</p>
  </abstract>
    </article-meta>
  </front>
<body>
      

      <?xmltex \hack{\newpage}?>
<sec id="Ch1.S1" sec-type="intro">
  <title>Introduction</title>
      <p>The most complicated, and probably most realistic, numerical models of the
ocean circulation are eddy-resolving ocean general circulation models
<xref ref-type="bibr" rid="bib1.bibx1 bib1.bibx39 bib1.bibx26" id="paren.1"><named-content content-type="pre">e.g.,</named-content></xref>.
Such models are a natural choice in ocean state estimation, the combination
of models and observations to reconstruct our best estimate of what the ocean
has actually done <xref ref-type="bibr" rid="bib1.bibx53" id="paren.2"><named-content content-type="pre">e.g.,</named-content></xref>. Here, we
restrict our focus to state estimation as the transient reconstruction of the
ocean state over a finite time interval where observations have been
collected, following the convention of <xref ref-type="bibr" rid="bib1.bibx64" id="text.3"/>.
In order to unambiguously diagnose physical mechanisms of interest, the ocean
state must be dynamically consistent: a solution to the dynamical equations
of motion without unphysical sources and sinks. The Lagrange multiplier
method <xref ref-type="bibr" rid="bib1.bibx58 bib1.bibx62" id="paren.4"><named-content content-type="pre">e.g.,</named-content></xref>,
sometimes called the adjoint method <xref ref-type="bibr" rid="bib1.bibx28 bib1.bibx59" id="paren.5"><named-content content-type="pre">e.g.,</named-content></xref>,
“4D-VAR” <xref ref-type="bibr" rid="bib1.bibx10 bib1.bibx16" id="paren.6"><named-content content-type="pre">e.g.,</named-content></xref>,
or variational data assimilation <xref ref-type="bibr" rid="bib1.bibx36 bib1.bibx5 bib1.bibx4" id="paren.7"><named-content content-type="pre">e.g.,</named-content></xref>,
is a method that satisfies these criteria, unlike the Kalman filter
<xref ref-type="bibr" rid="bib1.bibx18" id="paren.8"><named-content content-type="pre">e.g.,</named-content></xref> or nudging techniques
<xref ref-type="bibr" rid="bib1.bibx38" id="paren.9"><named-content content-type="pre">e.g.,</named-content></xref>.</p>
      <p>For the Lagrange multiplier method to be successful in state-of-the-art ocean
models, two major issues need to be addressed: (1) the high dimensionality of
the forward model and estimation problem, and (2) the nonlinearity of ocean
models at increasingly fine resolution. Research conducted by the ECCO
(Estimating the Circulation and Climate of the Ocean) Consortium
<xref ref-type="bibr" rid="bib1.bibx54 bib1.bibx55" id="paren.10"/> has
demonstrated that (1) the dimensionality of many million state variables
presents a challenge, but it can be overcome insofar as a solution can be
found that fits the ocean data <xref ref-type="bibr" rid="bib1.bibx63" id="paren.11"><named-content content-type="pre">e.g.,</named-content></xref>. One caveat is that the
convergence of the optimization process may be slower than hoped, but this is
primarily an issue of computational efficiency. Regarding nonlinearity (2),
the adjoint model has the same stability characteristics as the forward
model, as the eigenvalues of linearized state transition matrix are the same
as the transpose of the matrix <xref ref-type="bibr" rid="bib1.bibx49" id="paren.12"/>.
Therefore, nonlinearity in the forward model may be accompanied by an
unstable adjoint model and Lagrange multipliers that grow exponentially with
time. When the Lagrange multiplier method is used to enforce a nonlinear
constraint such as a chaotic model, the search for a solution becomes
iterative and the Lagrange multipliers provide gradient information that is
used to minimize an objective function that describes the model fit to
observations <xref ref-type="bibr" rid="bib1.bibx41" id="paren.13"><named-content content-type="pre">e.g.,</named-content></xref>. For a
bounded objective function with growing gradients, multiple local minima are
present that complicate the search for a global minimum
<xref ref-type="bibr" rid="bib1.bibx43" id="paren.14"><named-content content-type="pre">e.g.,</named-content></xref>. Even sophisticated gradient-descent
algorithms such as the variable-storage quasi-Newton method
<xref ref-type="bibr" rid="bib1.bibx47 bib1.bibx24" id="paren.15"/> can become
stalled in a local minimum and are not guaranteed to fit the observations
adequately. For example, <xref ref-type="bibr" rid="bib1.bibx34" id="text.16"/> used the
<xref ref-type="bibr" rid="bib1.bibx37" id="text.17"/> model to conclude that the “adjoint does
not tend to useful sensitivity values”, echoing previous concerns with
simple, chaotic models <xref ref-type="bibr" rid="bib1.bibx20 bib1.bibx45 bib1.bibx56" id="paren.18"><named-content content-type="pre">e.g.,</named-content></xref>.</p>
      <p>Due in part to the concerns raised about nonlinearity in simple models, the
method of Lagrange multipliers has rarely been applied to realistic models
over time windows longer than the eddy scale. For example, some studies
restricted the time windows to be short enough that unstable modes would not
grow too large <xref ref-type="bibr" rid="bib1.bibx51 bib1.bibx9" id="paren.19"><named-content content-type="pre">e.g.,</named-content></xref>.
The Southern Ocean State Estimate was produced with an approximate version of
the method of Lagrange multipliers, where the Lagrange multipliers are
calculated by an adjoint model with artificially large diffusivities that
stabilize the model <xref ref-type="bibr" rid="bib1.bibx42" id="paren.20"/>. Such an approach is
not guaranteed to work, as the Lagrange multipliers of the stabilized model
have no simple relation with those of the original eddy-resolving model. The
iterative search technique could then be led in the opposite direction as the
truth, as was shown to occur in a quasi-geostrophic ocean model
<xref ref-type="bibr" rid="bib1.bibx32" id="paren.21"/>. We are aware of only one case where the
unmodified method of Lagrange multipliers was applied to an eddy-permitting
ocean GCM over a timescale longer than the eddy scale of a few months
<xref ref-type="bibr" rid="bib1.bibx23" id="paren.22"/>. Contrary to expectation given by the
simple chaotic models, an acceptable fit was found to oceanographic
observations over a 1-year interval in the northeast Atlantic Ocean
<xref ref-type="bibr" rid="bib1.bibx22" id="paren.23"/>. No clear explanation for these disparate results
has been put forward.</p>
      <p>In this research, we wish to re-examine (2) the influence of nonlinear
models on the method of Lagrange multipliers and ocean state estimation. Is
the adjoint method useless with a highly nonlinear or chaotic system, as
studies with low-dimensional chaotic models suggest? Here we posit that the
initialization problem that has informed much of the current thinking about
the Lagrange multiplier method is not the relevant analogy for ocean state
estimation. As has been documented in textbooks
<xref ref-type="bibr" rid="bib1.bibx3 bib1.bibx61" id="paren.24"><named-content content-type="pre">e.g.,</named-content></xref>, the ocean state
estimation problem is better described as a time-variable boundary value
problem because synoptic atmospheric variability acts as an external forcing
on the ocean. Given our relatively uncertain knowledge regarding air–sea
fluxes, the ocean state estimation is rightfully considered a time-variable
boundary value problem where both the initial conditions and boundary
conditions must be found. For example, <xref ref-type="bibr" rid="bib1.bibx4" id="text.25"/>
described an estimation method for the external forcing, initial and boundary
conditions that solves the Euler–Lagrange equations for a linear model. In
the typical implementation of ocean state estimation with a general
circulation model <xref ref-type="bibr" rid="bib1.bibx31" id="paren.26"><named-content content-type="pre">e.g.,</named-content></xref>, the
surface forcing is defined to be part of the control vector. Because the
effect of nonlinearity is seen as the major roadblock for application of the
Lagrange multipler method, we isolate this effect by choosing a model that is
highly nonlinear but low-dimensional: the forced, chaotic pendulum (Sect. 2).
Toy models are worth revisiting because the dynamics is comparatively
simple to understand, the nonlinear coupling to periodic forcing has been
shown to be important in atmosphere–ocean dynamics
<xref ref-type="bibr" rid="bib1.bibx60" id="paren.27"><named-content content-type="pre">e.g.,</named-content></xref>, and these models have strongly
influenced when the Lagrange multiplier method has been deployed to realistic
ocean problems. We will show that previous toy models have sometimes been misinterpreted.</p>
      <p>Rather than developing a new state-of-the-art data assimilation technique, we
proceed by taking the existing Lagrange multipler method and developing
diagnostics regarding when and why it succeeds or fails, as evaluated by the
ability to fit observations. Relative to the initialization problem, the
prospects for a successful state estimate are shown to be improved in the
boundary control problem, even if one uses a highly nonlinear model such as
the forced, chaotic pendulum. If the chaotic nature of the model is not a
roadblock, what is the relevant criterion for success with the Lagrange
multiplier method? Our results with the chaotic pendulum suggest that
“controllability”, defined as the ability to move from one arbitrary state
to another by adjustments on the control variables (e.g., external forcing),
is the relevant diagnostic. The control variables of the pendulum are
analogous to adjustments of the atmospheric boundary forcing in an ocean
model. Therefore, there is a wide variety of situations where the Lagrange
multipiers of an ocean general circulation model (GCM) are useful, and that
previous GCM results can be explained in this context.</p>
</sec>
<sec id="Ch1.S2">
  <title>Lagrange multiplier method</title>
<sec id="Ch1.S2.SS1">
  <title>Pendulum model and synthetic data</title>
      <p>The fixed, single pendulum can be modeled as a nonlinear or linear set of
equations, and it can also be easily modified to be stable or unstable. In
many ways, the pendulum is a more flexible and easily interpreted physical
system than the often-used <xref ref-type="bibr" rid="bib1.bibx37" id="text.28"/> equations that
approximate atmospheric convection. The relevance of the pendulum to the
ocean is obviously indirect, but much of the community's knowledge of state
estimation has been formed through the intuition of simple models. The motion
of the forced pendulum is described by the deterministic equation
<xref ref-type="bibr" rid="bib1.bibx2" id="paren.29"/>:

                <disp-formula id="Ch1.E1" content-type="numbered"><mml:math id="M1" display="block"><mml:mstyle class="stylechange" displaystyle="true"/><mml:mrow><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:msup><mml:mi mathvariant="normal">d</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mi mathvariant="italic">θ</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="normal">d</mml:mi><mml:msup><mml:mi>t</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:mfrac></mml:mstyle><mml:mo>+</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mn mathvariant="normal">1</mml:mn><mml:mi>q</mml:mi></mml:mfrac></mml:mstyle><mml:mspace width="0.25em" linebreak="nobreak"/><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:mi mathvariant="normal">d</mml:mi><mml:mi mathvariant="italic">θ</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="normal">d</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:mfrac></mml:mstyle><mml:mo>+</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mi>g</mml:mi><mml:mi>l</mml:mi></mml:mfrac></mml:mstyle><mml:mi>sin⁡</mml:mi><mml:mi mathvariant="italic">θ</mml:mi><mml:mo>=</mml:mo><mml:mi>f</mml:mi><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>

          where <inline-formula><mml:math id="M2" display="inline"><mml:mi mathvariant="italic">θ</mml:mi></mml:math></inline-formula> is the displacement angle from vertical, <inline-formula><mml:math id="M3" display="inline"><mml:mi>q</mml:mi></mml:math></inline-formula> is a damping
coefficient, <inline-formula><mml:math id="M4" display="inline"><mml:mi>g</mml:mi></mml:math></inline-formula> is gravitational acceleration, <inline-formula><mml:math id="M5" display="inline"><mml:mi>l</mml:mi></mml:math></inline-formula> is the pendulum length,
and <inline-formula><mml:math id="M6" display="inline"><mml:mrow><mml:mi>f</mml:mi><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> is an external forcing term. Later, the external forcing will be
broken into a first guess and a perturbation, <inline-formula><mml:math id="M7" display="inline"><mml:mrow><mml:mi>f</mml:mi><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> <inline-formula><mml:math id="M8" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> <inline-formula><mml:math id="M9" display="inline"><mml:mrow><mml:msub><mml:mi>f</mml:mi><mml:mn mathvariant="normal">0</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula> <inline-formula><mml:math id="M10" display="inline"><mml:mo>+</mml:mo></mml:math></inline-formula> <inline-formula><mml:math id="M11" display="inline"><mml:mrow><mml:mi mathvariant="italic">δ</mml:mi><mml:mi>f</mml:mi><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>,
where the first guess is set to periodic forcing,
<inline-formula><mml:math id="M12" display="inline"><mml:mrow><mml:msub><mml:mi>f</mml:mi><mml:mn mathvariant="normal">0</mml:mn></mml:msub><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> <inline-formula><mml:math id="M13" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> <inline-formula><mml:math id="M14" display="inline"><mml:mrow><mml:mi>b</mml:mi><mml:mi>cos⁡</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi mathvariant="italic">ω</mml:mi><mml:mi>d</mml:mi></mml:msub><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>. With parameters <inline-formula><mml:math id="M15" display="inline"><mml:mi>q</mml:mi></mml:math></inline-formula> <inline-formula><mml:math id="M16" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> 100 s,  <inline-formula><mml:math id="M17" display="inline"><mml:mrow><mml:mi>g</mml:mi><mml:mo>/</mml:mo><mml:mi>l</mml:mi></mml:mrow></mml:math></inline-formula> <inline-formula><mml:math id="M18" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> 1.0 s<inline-formula><mml:math id="M19" display="inline"><mml:msup><mml:mi/><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula>, <inline-formula><mml:math id="M20" display="inline"><mml:mi>b</mml:mi></mml:math></inline-formula> <inline-formula><mml:math id="M21" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> 1.5 rad s<inline-formula><mml:math id="M22" display="inline"><mml:msup><mml:mi/><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula>, and
<inline-formula><mml:math id="M23" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">ω</mml:mi><mml:mi>d</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> <inline-formula><mml:math id="M24" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> <inline-formula><mml:math id="M25" display="inline"><mml:mrow><mml:mn mathvariant="normal">2</mml:mn><mml:mo>/</mml:mo><mml:mn mathvariant="normal">3</mml:mn></mml:mrow></mml:math></inline-formula> s<inline-formula><mml:math id="M26" display="inline"><mml:msup><mml:mi/><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula>, the pendulum is chaotic (here defined as extreme
sensitivity to initial conditions). Following the numerical implementation in
Appendix <xref ref-type="sec" rid="App1.Ch1.S1"/>, the state vector is defined, <inline-formula><mml:math id="M27" display="inline"><mml:mrow><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> <inline-formula><mml:math id="M28" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> <inline-formula><mml:math id="M29" display="inline"><mml:mrow><mml:mo>[</mml:mo><mml:mi mathvariant="italic">ω</mml:mi><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo><mml:mspace width="0.25em" linebreak="nobreak"/><mml:mi mathvariant="italic">θ</mml:mi><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo><mml:msup><mml:mo>]</mml:mo><mml:mi>T</mml:mi></mml:msup></mml:mrow></mml:math></inline-formula>,
where <inline-formula><mml:math id="M30" display="inline"><mml:msup><mml:mi/><mml:mi>T</mml:mi></mml:msup></mml:math></inline-formula> is the vector transpose and the state
variables are related by <inline-formula><mml:math id="M31" display="inline"><mml:mi mathvariant="italic">ω</mml:mi></mml:math></inline-formula> <inline-formula><mml:math id="M32" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> d<inline-formula><mml:math id="M33" display="inline"><mml:mrow><mml:mi mathvariant="italic">θ</mml:mi><mml:mo>/</mml:mo></mml:mrow></mml:math></inline-formula>d<inline-formula><mml:math id="M34" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula>. Matrices and vectors are
indicated in boldface. The state has dimension <inline-formula><mml:math id="M35" display="inline"><mml:mi>M</mml:mi></mml:math></inline-formula> <inline-formula><mml:math id="M36" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> 2 and the forcing vector
has dimension 1. The evolution of the state is succinctly written as

                <disp-formula id="Ch1.E2" content-type="numbered"><mml:math id="M37" display="block"><mml:mstyle class="stylechange" displaystyle="true"/><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mi mathvariant="normal">Δ</mml:mi><mml:mi>t</mml:mi><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:mi mathvariant="script">L</mml:mi><mml:mo>[</mml:mo><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo><mml:mo>,</mml:mo><mml:mi>f</mml:mi><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo><mml:mo>]</mml:mo><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>

          where the model state is stepped from time <inline-formula><mml:math id="M38" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula> to <inline-formula><mml:math id="M39" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula> <inline-formula><mml:math id="M40" display="inline"><mml:mo>+</mml:mo></mml:math></inline-formula> <inline-formula><mml:math id="M41" display="inline"><mml:mrow><mml:mi mathvariant="normal">Δ</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:math></inline-formula>, and
<inline-formula><mml:math id="M42" display="inline"><mml:mi mathvariant="script">L</mml:mi></mml:math></inline-formula> is the discretized, nonlinear operator that represents
Eq. (<xref ref-type="disp-formula" rid="Ch1.E1"/>). In the ocean model case, the state would
correspond to velocities and property fields, and the external forcing would
include air–sea momentum, heat, and freshwater fluxes.</p>
      <p>We consider an “identical twin” experiment where the true solution is known
(solid line, Fig. <xref ref-type="fig" rid="Ch1.F1"/>), and we observe the pendulum angle
episodically through time with normally distributed random errors of standard
deviation, <inline-formula><mml:math id="M43" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">σ</mml:mi><mml:mi mathvariant="italic">θ</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> <inline-formula><mml:math id="M44" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> 0.5 rad. In most oceanographically relevant
cases, observations have already been collected over some fixed time interval
(0 <inline-formula><mml:math id="M45" display="inline"><mml:mo>≤</mml:mo></mml:math></inline-formula> <inline-formula><mml:math id="M46" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula> <inline-formula><mml:math id="M47" display="inline"><mml:mo>≤</mml:mo></mml:math></inline-formula> <inline-formula><mml:math id="M48" display="inline"><mml:mi>T</mml:mi></mml:math></inline-formula>). Here, observations, <inline-formula><mml:math id="M49" display="inline"><mml:mrow><mml:mi>y</mml:mi><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>, are taken at a set of <inline-formula><mml:math id="M50" display="inline"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi>y</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>
evenly spaced times with an time interval of <inline-formula><mml:math id="M51" display="inline"><mml:mrow><mml:mi mathvariant="normal">Δ</mml:mi><mml:msub><mml:mi>t</mml:mi><mml:mi>y</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> <inline-formula><mml:math id="M52" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> <inline-formula><mml:math id="M53" display="inline"><mml:mrow><mml:mi>T</mml:mi><mml:mo>/</mml:mo><mml:mo>(</mml:mo><mml:msub><mml:mi>N</mml:mi><mml:mi>y</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> <inline-formula><mml:math id="M54" display="inline"><mml:mo>-</mml:mo></mml:math></inline-formula> 1).</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F1"><caption><p>The rapid divergence of pendulum trajectories is indicated by the
path density of trajectories (background shading), and the evolution of
three sample trajectories: the “truth” or reference trajectory (solid line), a
“first-guess” trajectory with incorrect initial angular velocity that
diverges within 5 s (dashed line), and a first-guess trajectory with
incorrect initial angle, <inline-formula><mml:math id="M55" display="inline"><mml:mi mathvariant="italic">θ</mml:mi></mml:math></inline-formula>, which diverges after 30 s (other dashed
line). The path density of trajectories is computed with 10 000 forward
integrations with normally distributed perturbations about the truth
(standard deviation: <inline-formula><mml:math id="M56" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">σ</mml:mi><mml:mi mathvariant="italic">ω</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> <inline-formula><mml:math id="M57" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> 1 rad s<inline-formula><mml:math id="M58" display="inline"><mml:msup><mml:mi/><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula>,
<inline-formula><mml:math id="M59" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">σ</mml:mi><mml:mi mathvariant="italic">θ</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> <inline-formula><mml:math id="M60" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> 0.5 rad).</p></caption>
          <?xmltex \igopts{width=199.169291pt}?><graphic xlink:href="https://npg.copernicus.org/articles/24/351/2017/npg-24-351-2017-f01.pdf"/>

        </fig>

</sec>
<sec id="Ch1.S2.SS2">
  <title>Cost function</title>
      <p>We proceed by defining a least-squares cost function to be minimized. The
data-based contribution to the cost function, <inline-formula><mml:math id="M61" display="inline"><mml:mrow><mml:msub><mml:mi>J</mml:mi><mml:mi mathvariant="normal">d</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>, measures the squared
misfit between the model and observations:

                <disp-formula id="Ch1.E3" content-type="numbered"><mml:math id="M62" display="block"><mml:mstyle displaystyle="true" class="stylechange"/><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><mml:msub><mml:mi>J</mml:mi><mml:mi mathvariant="normal">d</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mn mathvariant="normal">1</mml:mn><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi>y</mml:mi></mml:msub></mml:mrow></mml:mfrac></mml:mstyle><mml:munderover><mml:mo movablelimits="false">∑</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi>y</mml:mi></mml:msub><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:munderover><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:msup><mml:mfenced open="[" close="]"><mml:mi mathvariant="italic">θ</mml:mi><mml:mfenced close=")" open="("><mml:mi>i</mml:mi><mml:mi mathvariant="normal">Δ</mml:mi><mml:msub><mml:mi>t</mml:mi><mml:mi>y</mml:mi></mml:msub></mml:mfenced><mml:mo>-</mml:mo><mml:mi>y</mml:mi><mml:mfenced close=")" open="("><mml:mi>i</mml:mi><mml:mi mathvariant="normal">Δ</mml:mi><mml:msub><mml:mi>t</mml:mi><mml:mi>y</mml:mi></mml:msub></mml:mfenced></mml:mfenced><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow><mml:mrow><mml:msubsup><mml:mi mathvariant="italic">σ</mml:mi><mml:mi mathvariant="italic">θ</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msubsup></mml:mrow></mml:mfrac></mml:mstyle><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>

          where the penalty is weighted by the number of observations, <inline-formula><mml:math id="M63" display="inline"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi>y</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>, and their
standard error, <inline-formula><mml:math id="M64" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">σ</mml:mi><mml:mi mathvariant="italic">θ</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>, such that the expected value of <inline-formula><mml:math id="M65" display="inline"><mml:mrow><mml:msub><mml:mi>J</mml:mi><mml:mi mathvariant="normal">d</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> is
near 1. As we have imposed Gaussian error statistics, minimizing this
least-squares cost function also leads to the maximum likelihood solution
<xref ref-type="bibr" rid="bib1.bibx29" id="paren.30"><named-content content-type="pre">e.g.,</named-content></xref>. In matrix–vector notation,
Eq. (<xref ref-type="disp-formula" rid="Ch1.E3"/>) becomes

                <disp-formula specific-use="align" content-type="numbered"><mml:math id="M66" display="block"><mml:mtable displaystyle="true"><mml:mtr><mml:mtd><mml:mrow><mml:mstyle displaystyle="true" class="stylechange"/><mml:msub><mml:mi>J</mml:mi><mml:mi mathvariant="normal">d</mml:mi></mml:msub><mml:mo>=</mml:mo></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:mstyle displaystyle="true" class="stylechange"/><mml:munderover><mml:mo movablelimits="false">∑</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi>y</mml:mi></mml:msub><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:munderover><mml:mo>[</mml:mo><mml:mi mathvariant="bold">E</mml:mi><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mfenced close=")" open="("><mml:mi>i</mml:mi><mml:mi mathvariant="normal">Δ</mml:mi><mml:msub><mml:mi>t</mml:mi><mml:mi>y</mml:mi></mml:msub></mml:mfenced><mml:mo>-</mml:mo><mml:mi>y</mml:mi><mml:mfenced close=")" open="("><mml:mi>i</mml:mi><mml:mi mathvariant="normal">Δ</mml:mi><mml:msub><mml:mi>t</mml:mi><mml:mi>y</mml:mi></mml:msub></mml:mfenced><mml:msup><mml:mo>]</mml:mo><mml:mi>T</mml:mi></mml:msup><mml:msup><mml:mi>W</mml:mi><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:mtd></mml:mtr><mml:mlabeledtr id="Ch1.E4"><mml:mtd/><mml:mtd><mml:mstyle displaystyle="true" class="stylechange"/></mml:mtd><mml:mtd><mml:mrow><mml:mstyle displaystyle="true" class="stylechange"/><mml:mfenced close="]" open="["><mml:mi mathvariant="bold">E</mml:mi><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mfenced open="(" close=")"><mml:mi>i</mml:mi><mml:mi mathvariant="normal">Δ</mml:mi><mml:msub><mml:mi>t</mml:mi><mml:mi>y</mml:mi></mml:msub></mml:mfenced><mml:mo>-</mml:mo><mml:mi>y</mml:mi><mml:mfenced open="(" close=")"><mml:mi>i</mml:mi><mml:mi mathvariant="normal">Δ</mml:mi><mml:msub><mml:mi>t</mml:mi><mml:mi>y</mml:mi></mml:msub></mml:mfenced></mml:mfenced><mml:mo>,</mml:mo></mml:mrow></mml:mtd></mml:mlabeledtr></mml:mtable></mml:math></disp-formula>

            where <inline-formula><mml:math id="M67" display="inline"><mml:mrow><mml:mi>y</mml:mi><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> is a scalar, <inline-formula><mml:math id="M68" display="inline"><mml:mi mathvariant="bold">E</mml:mi></mml:math></inline-formula> is the observational matrix that samples
the observable part of the state and has dimension 1 <inline-formula><mml:math id="M69" display="inline"><mml:mo>×</mml:mo></mml:math></inline-formula> <inline-formula><mml:math id="M70" display="inline"><mml:mi>M</mml:mi></mml:math></inline-formula>, and <inline-formula><mml:math id="M71" display="inline"><mml:mi>W</mml:mi></mml:math></inline-formula> is a
weight. Comparison of the first term in Eqs. (<xref ref-type="disp-formula" rid="Ch1.E4"/>) to (<xref ref-type="disp-formula" rid="Ch1.E3"/>)
shows that <inline-formula><mml:math id="M72" display="inline"><mml:mi>W</mml:mi></mml:math></inline-formula> <inline-formula><mml:math id="M73" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> <inline-formula><mml:math id="M74" display="inline"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi>y</mml:mi></mml:msub><mml:msubsup><mml:mi mathvariant="italic">σ</mml:mi><mml:mi mathvariant="italic">θ</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msubsup></mml:mrow></mml:math></inline-formula>. While it is
unconventional to transpose the scalar data–model misfit in
Eq. (<xref ref-type="disp-formula" rid="Ch1.E4"/>), we retain this notation so that the equations are
applicable to cases where multiple observations are available at each time.</p>
      <p>A second contribution to the cost function includes two terms that constrain
the difference between our posterior and prior estimates of the initial
conditions and forcing,
<?xmltex \hack{\newpage}?><?xmltex \hack{\vspace*{-6mm}}?>

                <disp-formula specific-use="align" content-type="numbered"><mml:math id="M75" display="block"><mml:mtable displaystyle="true"><mml:mtr><mml:mtd><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><mml:msub><mml:mi>J</mml:mi><mml:mn mathvariant="normal">0</mml:mn></mml:msub></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><mml:mo>=</mml:mo><mml:msup><mml:mfenced open="[" close="]"><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>(</mml:mo><mml:mn mathvariant="normal">0</mml:mn><mml:mo>)</mml:mo><mml:mo>-</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mn mathvariant="normal">0</mml:mn></mml:msub><mml:mo>(</mml:mo><mml:mn mathvariant="normal">0</mml:mn><mml:mo>)</mml:mo></mml:mfenced><mml:mi>T</mml:mi></mml:msup><mml:msubsup><mml:mi mathvariant="bold">S</mml:mi><mml:mi>x</mml:mi><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msubsup><mml:mfenced close="]" open="["><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>(</mml:mo><mml:mn mathvariant="normal">0</mml:mn><mml:mo>)</mml:mo><mml:mo>-</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mn mathvariant="normal">0</mml:mn></mml:msub><mml:mo>(</mml:mo><mml:mn mathvariant="normal">0</mml:mn><mml:mo>)</mml:mo></mml:mfenced></mml:mrow></mml:mtd></mml:mtr><mml:mlabeledtr id="Ch1.E5"><mml:mtd/><mml:mtd><mml:mstyle displaystyle="true" class="stylechange"/></mml:mtd><mml:mtd><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><?xmltex \hack{\hbox\bgroup\fontsize{9.5}{9.5}\selectfont$\displaystyle}?><mml:mo>+</mml:mo><mml:munderover><mml:mo movablelimits="false">∑</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:munderover><mml:msup><mml:mfenced open="[" close="]"><mml:mi>f</mml:mi><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mi mathvariant="normal">Δ</mml:mi><mml:mi>t</mml:mi><mml:mo>)</mml:mo><mml:mo>-</mml:mo><mml:msub><mml:mi>f</mml:mi><mml:mn mathvariant="normal">0</mml:mn></mml:msub><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mi mathvariant="normal">Δ</mml:mi><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mfenced><mml:mi>T</mml:mi></mml:msup><mml:msubsup><mml:mi>S</mml:mi><mml:mi>f</mml:mi><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msubsup><mml:mfenced close="]" open="["><mml:mi>f</mml:mi><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mi mathvariant="normal">Δ</mml:mi><mml:mi>t</mml:mi><mml:mo>)</mml:mo><mml:mo>-</mml:mo><mml:msub><mml:mi>f</mml:mi><mml:mn mathvariant="normal">0</mml:mn></mml:msub><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mi mathvariant="normal">Δ</mml:mi><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mfenced><mml:mo>,</mml:mo><?xmltex \hack{$\egroup}?></mml:mrow></mml:mtd></mml:mlabeledtr></mml:mtable></mml:math></disp-formula>

            where <inline-formula><mml:math id="M76" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mn mathvariant="normal">0</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula>(0) is the first-guess initial conditions, there are
<inline-formula><mml:math id="M77" display="inline"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> model time steps, <inline-formula><mml:math id="M78" display="inline"><mml:mrow><mml:msub><mml:mi>f</mml:mi><mml:mn mathvariant="normal">0</mml:mn></mml:msub><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> is the first-guess forcing, and <inline-formula><mml:math id="M79" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold">S</mml:mi><mml:mi>x</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> and
<inline-formula><mml:math id="M80" display="inline"><mml:mrow><mml:msub><mml:mi>S</mml:mi><mml:mi>f</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> are weighted by 5 rad and 10 rad s<inline-formula><mml:math id="M81" display="inline"><mml:msup><mml:mi/><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula>, respectively, to
penalize deviations. Note that this cost function is also normalized by the
number of observations, <inline-formula><mml:math id="M82" display="inline"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi>y</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>, to be consistent in the posterior tests later
in this work. Here we seek values of <inline-formula><mml:math id="M83" display="inline"><mml:mrow><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M84" display="inline"><mml:mrow><mml:mi>f</mml:mi><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> that minimize
the sum, <inline-formula><mml:math id="M85" display="inline"><mml:mrow><mml:msup><mml:mi>J</mml:mi><mml:mo>′</mml:mo></mml:msup></mml:mrow></mml:math></inline-formula> <inline-formula><mml:math id="M86" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> <inline-formula><mml:math id="M87" display="inline"><mml:mrow><mml:msub><mml:mi>J</mml:mi><mml:mi mathvariant="normal">d</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> <inline-formula><mml:math id="M88" display="inline"><mml:mo>+</mml:mo></mml:math></inline-formula> <inline-formula><mml:math id="M89" display="inline"><mml:mrow><mml:msub><mml:mi>J</mml:mi><mml:mn mathvariant="normal">0</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula>, but the stationary point found by individually
minimizing the values d<inline-formula><mml:math id="M90" display="inline"><mml:mrow><mml:msup><mml:mi>J</mml:mi><mml:mo>′</mml:mo></mml:msup><mml:mo>/</mml:mo></mml:mrow></mml:math></inline-formula>d<inline-formula><mml:math id="M91" display="inline"><mml:mrow><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> and d<inline-formula><mml:math id="M92" display="inline"><mml:mrow><mml:msup><mml:mi>J</mml:mi><mml:mo>′</mml:mo></mml:msup><mml:mo>/</mml:mo></mml:mrow></mml:math></inline-formula>d<inline-formula><mml:math id="M93" display="inline"><mml:mrow><mml:mi>f</mml:mi><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> will
almost certainly violate the model constraint in Eq. (<xref ref-type="disp-formula" rid="Ch1.E2"/>).
We enforce the model constraint by appending a Lagrange multiplier term to
the combined cost function,

                <disp-formula specific-use="align" content-type="numbered"><mml:math id="M94" display="block"><mml:mtable displaystyle="true"><mml:mtr><mml:mtd><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><mml:mi>J</mml:mi></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><mml:mo>=</mml:mo><mml:msub><mml:mi>J</mml:mi><mml:mi mathvariant="normal">d</mml:mi></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>J</mml:mi><mml:mn mathvariant="normal">0</mml:mn></mml:msub><mml:mo>-</mml:mo><mml:mn mathvariant="normal">2</mml:mn><mml:munderover><mml:mo movablelimits="false">∑</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:munderover><mml:mi mathvariant="bold-italic">μ</mml:mi><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mi mathvariant="normal">Δ</mml:mi><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mi mathvariant="normal">Δ</mml:mi><mml:mi>t</mml:mi><mml:msup><mml:mo>)</mml:mo><mml:mi>T</mml:mi></mml:msup><mml:mo mathvariant="italic">{</mml:mo><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mi mathvariant="normal">Δ</mml:mi><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mi mathvariant="normal">Δ</mml:mi><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mlabeledtr id="Ch1.E6"><mml:mtd/><mml:mtd><mml:mstyle class="stylechange" displaystyle="true"/></mml:mtd><mml:mtd><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><mml:mo>-</mml:mo><mml:mi mathvariant="script">L</mml:mi><mml:mo>[</mml:mo><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mi mathvariant="normal">Δ</mml:mi><mml:mi>t</mml:mi><mml:mo>)</mml:mo><mml:mo>,</mml:mo><mml:mi>f</mml:mi><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mi mathvariant="normal">Δ</mml:mi><mml:mi>t</mml:mi><mml:mo>)</mml:mo><mml:mo>]</mml:mo><mml:mo mathvariant="italic">}</mml:mo><mml:mo>,</mml:mo></mml:mrow></mml:mtd></mml:mlabeledtr></mml:mtable></mml:math></disp-formula>

            where <inline-formula><mml:math id="M95" display="inline"><mml:mrow><mml:mi mathvariant="bold-italic">μ</mml:mi><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> is a Lagrange multiplier, and the scaling with “2” is
helpful in later derivations and does not change the numerical value of <inline-formula><mml:math id="M96" display="inline"><mml:mi>J</mml:mi></mml:math></inline-formula>
because the quantity inside curly brackets vanishes. Now the cost function
can be minimized by independently setting the partial derivatives of <inline-formula><mml:math id="M97" display="inline"><mml:mi>J</mml:mi></mml:math></inline-formula> with
respect to the state, the forcing, and the Lagrange multipliers to zero. This
problem will be solved using a gradient-descent method (detailed later) that
is excellent at finding the nearest minimum. If the first guess is good, then
the closest minimum may actually be the global minimum
<xref ref-type="bibr" rid="bib1.bibx50" id="paren.31"><named-content content-type="pre">e.g.,</named-content></xref>, and therefore we design an
improved first guess next.</p>
</sec>
<sec id="Ch1.S2.SS3">
  <title>First-guess trajectory</title>
      <p>Minimizing <inline-formula><mml:math id="M98" display="inline"><mml:mi>J</mml:mi></mml:math></inline-formula> requires a first-guess of the full model trajectory, <inline-formula><mml:math id="M99" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mn mathvariant="normal">0</mml:mn></mml:msub><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>.
A sensible and common approach is to use the observation at initial
time, <inline-formula><mml:math id="M100" display="inline"><mml:mi>y</mml:mi></mml:math></inline-formula>(0), to inform the initial conditions for the state, <inline-formula><mml:math id="M101" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mn mathvariant="normal">0</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula>(0).
Then, the first-guess forcing, <inline-formula><mml:math id="M102" display="inline"><mml:mrow><mml:msub><mml:mi>f</mml:mi><mml:mn mathvariant="normal">0</mml:mn></mml:msub><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> <inline-formula><mml:math id="M103" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> <inline-formula><mml:math id="M104" display="inline"><mml:mrow><mml:mi>b</mml:mi><mml:mi>cos⁡</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi mathvariant="italic">ω</mml:mi><mml:mi>d</mml:mi></mml:msub><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>, is used to drive
the model forward in time. In this case, the state at any time, <inline-formula><mml:math id="M105" display="inline"><mml:mi mathvariant="italic">τ</mml:mi></mml:math></inline-formula>, can
be computed directly from the initial state,

                <disp-formula specific-use="align" content-type="numbered"><mml:math id="M106" display="block"><mml:mtable displaystyle="true"><mml:mtr><mml:mtd><mml:mrow><mml:mstyle displaystyle="true" class="stylechange"/><?xmltex \hack{\hbox\bgroup\fontsize{8}{8}\selectfont$\displaystyle}?><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mn mathvariant="normal">0</mml:mn></mml:msub><mml:mo>(</mml:mo><mml:mi mathvariant="italic">τ</mml:mi><mml:mo>)</mml:mo><?xmltex \hack{$\egroup}?></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:mstyle displaystyle="true" class="stylechange"/><?xmltex \hack{\hbox\bgroup\fontsize{8}{8}\selectfont$\displaystyle}?><mml:mo>=</mml:mo><mml:msub><mml:mi mathvariant="script">L</mml:mi><mml:mrow><mml:mi>K</mml:mi><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msub><mml:mfenced open="[" close="]"><mml:mi mathvariant="normal">…</mml:mi><mml:mfenced close="]" open="["><mml:msub><mml:mi mathvariant="script">L</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub><mml:mfenced open="[" close="]"><mml:msub><mml:mi mathvariant="script">L</mml:mi><mml:mn mathvariant="normal">0</mml:mn></mml:msub><mml:mfenced close="]" open="["><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mn mathvariant="normal">0</mml:mn></mml:msub><mml:mo>(</mml:mo><mml:mn mathvariant="normal">0</mml:mn><mml:mo>)</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mi>f</mml:mi><mml:mn mathvariant="normal">0</mml:mn></mml:msub><mml:mo>(</mml:mo><mml:mn mathvariant="normal">0</mml:mn><mml:mo>)</mml:mo></mml:mfenced><mml:mo>,</mml:mo><mml:msub><mml:mi>f</mml:mi><mml:mn mathvariant="normal">0</mml:mn></mml:msub><mml:mo>(</mml:mo><mml:mi mathvariant="normal">Δ</mml:mi><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mfenced><mml:mi mathvariant="normal">…</mml:mi></mml:mfenced><mml:mo>,</mml:mo><mml:msub><mml:mi>f</mml:mi><mml:mn mathvariant="normal">0</mml:mn></mml:msub><mml:mo>(</mml:mo><mml:mi>K</mml:mi><mml:mi mathvariant="normal">Δ</mml:mi><mml:mi>t</mml:mi><mml:mo>-</mml:mo><mml:mi mathvariant="normal">Δ</mml:mi><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mfenced><?xmltex \hack{$\egroup}?></mml:mrow></mml:mtd></mml:mtr><mml:mlabeledtr id="Ch1.E7"><mml:mtd/><mml:mtd><mml:mstyle class="stylechange" displaystyle="true"/></mml:mtd><mml:mtd><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><?xmltex \hack{\hbox\bgroup\fontsize{8}{8}\selectfont$\displaystyle}?><mml:mo>=</mml:mo><mml:mi mathvariant="script">R</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="italic">τ</mml:mi><mml:mo>,</mml:mo><mml:mn mathvariant="normal">0</mml:mn><mml:mo>)</mml:mo><mml:mfenced open="[" close="]"><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mn mathvariant="normal">0</mml:mn></mml:msub><mml:mo>(</mml:mo><mml:mn mathvariant="normal">0</mml:mn><mml:mo>)</mml:mo></mml:mfenced><mml:mo>,</mml:mo><?xmltex \hack{$\egroup}?></mml:mrow></mml:mtd></mml:mlabeledtr></mml:mtable></mml:math></disp-formula>

            where <inline-formula><mml:math id="M107" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="script">L</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> indicates the nonlinear model operator at time step <inline-formula><mml:math id="M108" display="inline"><mml:mi>k</mml:mi></mml:math></inline-formula>,
<inline-formula><mml:math id="M109" display="inline"><mml:mi>K</mml:mi></mml:math></inline-formula> <inline-formula><mml:math id="M110" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> <inline-formula><mml:math id="M111" display="inline"><mml:mrow><mml:mi mathvariant="italic">τ</mml:mi><mml:mo>/</mml:mo><mml:mi mathvariant="normal">Δ</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:math></inline-formula> is the number of time steps between <inline-formula><mml:math id="M112" display="inline"><mml:mrow><mml:mi>t</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0</mml:mn></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M113" display="inline"><mml:mrow><mml:mi>t</mml:mi><mml:mo>=</mml:mo><mml:mi mathvariant="italic">τ</mml:mi></mml:mrow></mml:math></inline-formula>, and
the state transition matrix, <inline-formula><mml:math id="M114" display="inline"><mml:mrow><mml:mi mathvariant="script">R</mml:mi><mml:mo>(</mml:mo><mml:mi>m</mml:mi></mml:mrow></mml:math></inline-formula>, <inline-formula><mml:math id="M115" display="inline"><mml:mrow><mml:mi>n</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>, defines the aggregate,
nonlinear model step to time <inline-formula><mml:math id="M116" display="inline"><mml:mi>m</mml:mi></mml:math></inline-formula> from <inline-formula><mml:math id="M117" display="inline"><mml:mi>n</mml:mi></mml:math></inline-formula>. In the following, we refer to this
trajectory as the “standard” first-guess state.</p>
      <p>For a nonlinear system, and a chaotic system in particular, this first-guess
trajectory usually diverges from the already-collected observations at some
point, and thus can be ruled out as a possible solution <italic>a priori</italic>. When
the pendulum initial conditions are imperfectly known, the range of possible
pendulum trajectories expands greatly with time, even if the forcing
evolution is perfectly known (background shading, Fig. <xref ref-type="fig" rid="Ch1.F1"/>).
Normally distributed initial perturbations to the truth with standard
deviation of 1 rad s<inline-formula><mml:math id="M118" display="inline"><mml:msup><mml:mi/><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula> in angular velocity and 0.5 rad in the initial
angle lead to a divergence of roughly 200 rad between extreme trajectories
(background shading, Fig. <xref ref-type="fig" rid="Ch1.F1"/>). The angle is not renormalized
when the angle is greater or less than <inline-formula><mml:math id="M119" display="inline"><mml:mi mathvariant="italic">π</mml:mi></mml:math></inline-formula>, and thus the angle records a
history of how many times the pendulum has rotated. If no information about
the initial angular velocity is available, a reasonable assumption is that
<inline-formula><mml:math id="M120" display="inline"><mml:mi mathvariant="italic">ω</mml:mi></mml:math></inline-formula> <inline-formula><mml:math id="M121" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> 0 with some large error, but the pendulum trajectory with this
initial velocity and the correct initial angle diverges from truth in less
than 5 s (first dashed line, Fig. <xref ref-type="fig" rid="Ch1.F1"/>). In the case where
the initial velocity is known perfectly but the initial angle is observed
with an initial error of 0.5 rad (second dashed line,
Fig. <xref ref-type="fig" rid="Ch1.F1"/>), the trajectory follows truth for 30 s before
eventually diverging. As the time interval of interest increases, any
uncertainty in the initial conditions will ultimately lead to a divergence
between truth and this first-guess model trajectory. While these sample model
trajectories may seem overly naive, the first-guess trajectory used for ocean
state estimation usually has similar characteristics: usage of an observation
at the initial time, some prior knowledge of the forcing, and a freely running forward model.</p>
</sec>
<sec id="Ch1.S2.SS4">
  <title>An improved first guess</title>
      <p>The aforementioned standard approach does not use the observational
information already in hand that could inform the time evolution of the
forcing. There are many methods that are available to update the forcing,
such as the Kalman filter <xref ref-type="bibr" rid="bib1.bibx30" id="paren.32"><named-content content-type="pre">e.g.,</named-content></xref>,
but these methods rival or exceed the method of Lagrange multipliers in
computational cost because of the explicit representation of the solution
covariance matrix <xref ref-type="bibr" rid="bib1.bibx17" id="paren.33"/>. One remedy is to solve
the Kalman filter equation in a reduced space with the covariance represented
by an ensemble rather than being explicitly represented. Instead, we design a
whole-domain method that is computationally efficient and provides a good
first guess for the boundary control problem.</p>
      <p>Here we seek an update to the initial conditions and the forcing (i.e.,
<inline-formula><mml:math id="M122" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">ω</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula>(0) <inline-formula><mml:math id="M123" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> <inline-formula><mml:math id="M124" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">ω</mml:mi><mml:mn mathvariant="normal">0</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula>(0) <inline-formula><mml:math id="M125" display="inline"><mml:mo>+</mml:mo></mml:math></inline-formula> <inline-formula><mml:math id="M126" display="inline"><mml:mrow><mml:mi mathvariant="italic">δ</mml:mi><mml:mi mathvariant="italic">ω</mml:mi></mml:mrow></mml:math></inline-formula>, <inline-formula><mml:math id="M127" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">θ</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula>(0) <inline-formula><mml:math id="M128" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> <inline-formula><mml:math id="M129" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">θ</mml:mi><mml:mn mathvariant="normal">0</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula>(0) <inline-formula><mml:math id="M130" display="inline"><mml:mo>+</mml:mo></mml:math></inline-formula> <inline-formula><mml:math id="M131" display="inline"><mml:mrow><mml:mi mathvariant="italic">δ</mml:mi><mml:mi mathvariant="italic">θ</mml:mi></mml:mrow></mml:math></inline-formula>,
<inline-formula><mml:math id="M132" display="inline"><mml:mrow><mml:msub><mml:mi>f</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> <inline-formula><mml:math id="M133" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> <inline-formula><mml:math id="M134" display="inline"><mml:mrow><mml:msub><mml:mi>f</mml:mi><mml:mn mathvariant="normal">0</mml:mn></mml:msub><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> <inline-formula><mml:math id="M135" display="inline"><mml:mo>+</mml:mo></mml:math></inline-formula> <inline-formula><mml:math id="M136" display="inline"><mml:mrow><mml:mi mathvariant="italic">δ</mml:mi><mml:mi>f</mml:mi><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>), which takes the observations
into account. For small perturbations, we derive a linearized equation for
the change to the state at the time of the first observation,
<inline-formula><mml:math id="M137" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula> <inline-formula><mml:math id="M138" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> <inline-formula><mml:math id="M139" display="inline"><mml:mrow><mml:mi mathvariant="normal">Δ</mml:mi><mml:msub><mml:mi>t</mml:mi><mml:mi>y</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>
<?xmltex \hack{\newpage}?><?xmltex \hack{\vspace*{-6mm}}?>

                <disp-formula specific-use="align" content-type="numbered"><mml:math id="M140" display="block"><mml:mtable displaystyle="true"><mml:mtr><mml:mtd><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><?xmltex \hack{\hbox\bgroup\fontsize{8.9}{8.9}\selectfont$\displaystyle}?><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub><?xmltex \hack{$\egroup}?></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><?xmltex \hack{\hbox\bgroup\fontsize{8.9}{8.9}\selectfont$\displaystyle}?><mml:mfenced close=")" open="("><mml:mi mathvariant="normal">Δ</mml:mi><mml:msub><mml:mi>t</mml:mi><mml:mi>y</mml:mi></mml:msub></mml:mfenced><mml:mo>=</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mn mathvariant="normal">0</mml:mn></mml:msub><mml:mfenced open="(" close=")"><mml:mi mathvariant="normal">Δ</mml:mi><mml:msub><mml:mi>t</mml:mi><mml:mi>y</mml:mi></mml:msub></mml:mfenced><mml:mo>+</mml:mo><mml:mfenced close="" open="["><mml:msubsup><mml:mi mathvariant="normal">Π</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mi>K</mml:mi></mml:msubsup><mml:mi mathvariant="bold">A</mml:mi><mml:mo>(</mml:mo><mml:mo>(</mml:mo><mml:mi>K</mml:mi><mml:mo>-</mml:mo><mml:mi>i</mml:mi><mml:mo>)</mml:mo><mml:mi mathvariant="normal">Δ</mml:mi><mml:mi>t</mml:mi><mml:mo>)</mml:mo><mml:mo>|</mml:mo><mml:msubsup><mml:mi mathvariant="normal">Π</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mrow><mml:mi>K</mml:mi><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msubsup><mml:mi mathvariant="bold">A</mml:mi><mml:mo>(</mml:mo><mml:mo>(</mml:mo><mml:mi>K</mml:mi><mml:mo>-</mml:mo><mml:mi>i</mml:mi><mml:mo>)</mml:mo><mml:mi mathvariant="normal">Δ</mml:mi><mml:mi>t</mml:mi><mml:mo>)</mml:mo><mml:mi mathvariant="bold">B</mml:mi></mml:mfenced><?xmltex \hack{$\egroup}?></mml:mrow></mml:mtd></mml:mtr><mml:mlabeledtr id="Ch1.E8"><mml:mtd/><mml:mtd><mml:mstyle displaystyle="true" class="stylechange"/></mml:mtd><mml:mtd><mml:mrow><mml:mstyle displaystyle="true" class="stylechange"/><?xmltex \hack{\hbox\bgroup\fontsize{8.9}{8.9}\selectfont$\displaystyle}?><mml:mfenced open="." close="]"><mml:mo>|</mml:mo><mml:msubsup><mml:mi mathvariant="normal">Π</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mrow><mml:mi>K</mml:mi><mml:mo>-</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:mrow></mml:msubsup><mml:mi mathvariant="bold">A</mml:mi><mml:mo>(</mml:mo><mml:mo>(</mml:mo><mml:mi>K</mml:mi><mml:mo>-</mml:mo><mml:mi>i</mml:mi><mml:mo>)</mml:mo><mml:mi mathvariant="normal">Δ</mml:mi><mml:mi>t</mml:mi><mml:mo>)</mml:mo><mml:mi mathvariant="bold">B</mml:mi><mml:mo>|</mml:mo><mml:mi mathvariant="normal">…</mml:mi><mml:mo>|</mml:mo><mml:mi mathvariant="bold">B</mml:mi></mml:mfenced><mml:mfenced open="(" close=")"><mml:mtable class="array" columnalign="center"><mml:mtr><mml:mtd><mml:mrow><mml:mi mathvariant="italic">δ</mml:mi><mml:mi mathvariant="italic">ω</mml:mi><mml:mo>(</mml:mo><mml:mn mathvariant="normal">0</mml:mn><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mrow><mml:mi mathvariant="italic">δ</mml:mi><mml:mi mathvariant="italic">θ</mml:mi><mml:mo>(</mml:mo><mml:mn mathvariant="normal">0</mml:mn><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mrow><mml:mi mathvariant="italic">δ</mml:mi><mml:mi>f</mml:mi><mml:mo>(</mml:mo><mml:mn mathvariant="normal">0</mml:mn><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mrow><mml:mi mathvariant="italic">δ</mml:mi><mml:mi>f</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="normal">Δ</mml:mi><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mi mathvariant="normal">⋮</mml:mi></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mrow><mml:mi mathvariant="italic">δ</mml:mi><mml:mi>f</mml:mi><mml:mfenced close=")" open="("><mml:mi mathvariant="normal">Δ</mml:mi><mml:msub><mml:mi>t</mml:mi><mml:mi>y</mml:mi></mml:msub><mml:mo>-</mml:mo><mml:mi mathvariant="normal">Δ</mml:mi><mml:mi>t</mml:mi></mml:mfenced></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:mfenced><mml:mo>+</mml:mo><mml:mi mathvariant="italic">ϵ</mml:mi><mml:mo>,</mml:mo><?xmltex \hack{$\egroup}?></mml:mrow></mml:mtd></mml:mlabeledtr></mml:mtable></mml:math></disp-formula>

            where <inline-formula><mml:math id="M141" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula> is the “improved” first-guess, <inline-formula><mml:math id="M142" display="inline"><mml:mi>K</mml:mi></mml:math></inline-formula> <inline-formula><mml:math id="M143" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> <inline-formula><mml:math id="M144" display="inline"><mml:mrow><mml:mi mathvariant="normal">Δ</mml:mi><mml:msub><mml:mi>t</mml:mi><mml:mi>y</mml:mi></mml:msub><mml:mo>/</mml:mo><mml:mi mathvariant="normal">Δ</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:math></inline-formula>
is the number of model time steps from <inline-formula><mml:math id="M145" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula> <inline-formula><mml:math id="M146" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> 0 to <inline-formula><mml:math id="M147" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula> <inline-formula><mml:math id="M148" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> <inline-formula><mml:math id="M149" display="inline"><mml:mrow><mml:msub><mml:mi>t</mml:mi><mml:mi>y</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>,
<inline-formula><mml:math id="M150" display="inline"><mml:mrow><mml:mi mathvariant="bold">A</mml:mi><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> <inline-formula><mml:math id="M151" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> <inline-formula><mml:math id="M152" display="inline"><mml:mrow><mml:mo>∂</mml:mo><mml:mi mathvariant="script">L</mml:mi><mml:mo>/</mml:mo><mml:mo>∂</mml:mo><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> is the tangent-linear model,
<inline-formula><mml:math id="M153" display="inline"><mml:mi mathvariant="bold">B</mml:mi></mml:math></inline-formula> <inline-formula><mml:math id="M154" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> <inline-formula><mml:math id="M155" display="inline"><mml:mrow><mml:mo>∂</mml:mo><mml:mi mathvariant="script">L</mml:mi><mml:mo>/</mml:mo><mml:mo>∂</mml:mo><mml:mi>f</mml:mi><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> is constant in time, and <inline-formula><mml:math id="M156" display="inline"><mml:mi mathvariant="italic">ϵ</mml:mi></mml:math></inline-formula> is
the error due to linearization. We define the column vector of perturbations
in Eq. (<xref ref-type="disp-formula" rid="Ch1.E8"/>) to be the control vector, <inline-formula><mml:math id="M157" display="inline"><mml:mi mathvariant="bold-italic">u</mml:mi></mml:math></inline-formula>, so
that the equation becomes

                <disp-formula id="Ch1.E9" content-type="numbered"><mml:math id="M158" display="block"><mml:mstyle displaystyle="true" class="stylechange"/><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub><mml:mfenced close=")" open="("><mml:mi mathvariant="normal">Δ</mml:mi><mml:msub><mml:mi>t</mml:mi><mml:mi>y</mml:mi></mml:msub></mml:mfenced><mml:mo>=</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mn mathvariant="normal">0</mml:mn></mml:msub><mml:mfenced close=")" open="("><mml:mi mathvariant="normal">Δ</mml:mi><mml:msub><mml:mi>t</mml:mi><mml:mi>y</mml:mi></mml:msub></mml:mfenced><mml:mo>+</mml:mo><mml:mi mathvariant="bold">C</mml:mi><mml:mi mathvariant="bold-italic">u</mml:mi><mml:mo>+</mml:mo><mml:mi mathvariant="italic">ϵ</mml:mi><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>

          where <inline-formula><mml:math id="M159" display="inline"><mml:mi mathvariant="bold">C</mml:mi></mml:math></inline-formula> is the controllability (or reachability) matrix
<xref ref-type="bibr" rid="bib1.bibx11 bib1.bibx62" id="paren.34"><named-content content-type="pre">e.g.,</named-content></xref>.</p>
      <p>The observation, <inline-formula><mml:math id="M160" display="inline"><mml:mrow><mml:mi>y</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="normal">Δ</mml:mi><mml:msub><mml:mi>t</mml:mi><mml:mi>y</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>, and the combination of
Eqs. (<xref ref-type="disp-formula" rid="Ch1.E7"/>) and (<xref ref-type="disp-formula" rid="Ch1.E9"/>) provides one
constraint:

                <disp-formula id="Ch1.E10" content-type="numbered"><mml:math id="M161" display="block"><mml:mstyle class="stylechange" displaystyle="true"/><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><mml:mi>y</mml:mi><mml:mfenced close=")" open="("><mml:mi mathvariant="normal">Δ</mml:mi><mml:msub><mml:mi>t</mml:mi><mml:mi>y</mml:mi></mml:msub></mml:mfenced><mml:mo>=</mml:mo><mml:mi mathvariant="bold">E</mml:mi><mml:mi mathvariant="script">R</mml:mi><mml:mfenced close=")" open="("><mml:mi mathvariant="normal">Δ</mml:mi><mml:msub><mml:mi>t</mml:mi><mml:mi>y</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:mn mathvariant="normal">0</mml:mn></mml:mfenced><mml:mfenced close="]" open="["><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mn mathvariant="normal">0</mml:mn></mml:msub><mml:mo>(</mml:mo><mml:mn mathvariant="normal">0</mml:mn><mml:mo>)</mml:mo></mml:mfenced><mml:mo>+</mml:mo><mml:mi mathvariant="bold">EC</mml:mi><mml:mi mathvariant="bold-italic">u</mml:mi><mml:mo>+</mml:mo><mml:mi>n</mml:mi><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>

          where the controllability matrix can be calculated given the trajectory,
<inline-formula><mml:math id="M162" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mn mathvariant="normal">0</mml:mn></mml:msub><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>, and <inline-formula><mml:math id="M163" display="inline"><mml:mi>n</mml:mi></mml:math></inline-formula> is the misfit. Here we minimize the squared misfit,

                <disp-formula specific-use="align" content-type="numbered"><mml:math id="M164" display="block"><mml:mtable displaystyle="true"><mml:mtr><mml:mtd><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><?xmltex \hack{\hbox\bgroup\fontsize{9}{9}\selectfont$\displaystyle}?><mml:msub><mml:mi>J</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub><mml:mo>=</mml:mo><?xmltex \hack{$\egroup}?></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><?xmltex \hack{\hbox\bgroup\fontsize{9}{9}\selectfont$\displaystyle}?><mml:msup><mml:mfenced close="}" open="{"><mml:mi>y</mml:mi><mml:mfenced close=")" open="("><mml:mi mathvariant="normal">Δ</mml:mi><mml:msub><mml:mi>t</mml:mi><mml:mi>y</mml:mi></mml:msub></mml:mfenced><mml:mo>-</mml:mo><mml:mi mathvariant="bold">E</mml:mi><mml:mi mathvariant="script">R</mml:mi><mml:mfenced open="(" close=")"><mml:mi mathvariant="normal">Δ</mml:mi><mml:msub><mml:mi>t</mml:mi><mml:mi>y</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:mn mathvariant="normal">0</mml:mn></mml:mfenced><mml:mfenced close="]" open="["><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mn mathvariant="normal">0</mml:mn></mml:msub><mml:mo>(</mml:mo><mml:mn mathvariant="normal">0</mml:mn><mml:mo>)</mml:mo></mml:mfenced><mml:mo>-</mml:mo><mml:mi mathvariant="bold">EC</mml:mi><mml:mi mathvariant="bold-italic">u</mml:mi></mml:mfenced><mml:mi>T</mml:mi></mml:msup><mml:msup><mml:mi>W</mml:mi><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msup><?xmltex \hack{$\egroup}?></mml:mrow></mml:mtd></mml:mtr><mml:mlabeledtr id="Ch1.E11"><mml:mtd/><mml:mtd><mml:mstyle displaystyle="true" class="stylechange"/></mml:mtd><mml:mtd><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><?xmltex \hack{\hbox\bgroup\fontsize{9}{9}\selectfont$\displaystyle}?><mml:mfenced open="{" close="}"><mml:mi>y</mml:mi><mml:mfenced close=")" open="("><mml:mi mathvariant="normal">Δ</mml:mi><mml:msub><mml:mi>t</mml:mi><mml:mi>y</mml:mi></mml:msub></mml:mfenced><mml:mo>-</mml:mo><mml:mi mathvariant="bold">E</mml:mi><mml:mi mathvariant="script">R</mml:mi><mml:mfenced open="(" close=")"><mml:mi mathvariant="normal">Δ</mml:mi><mml:msub><mml:mi>t</mml:mi><mml:mi>y</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:mn mathvariant="normal">0</mml:mn></mml:mfenced><mml:mfenced open="[" close="]"><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mn mathvariant="normal">0</mml:mn></mml:msub><mml:mo>(</mml:mo><mml:mn mathvariant="normal">0</mml:mn><mml:mo>)</mml:mo></mml:mfenced><mml:mo>-</mml:mo><mml:mi mathvariant="bold">EC</mml:mi><mml:mi mathvariant="bold-italic">u</mml:mi></mml:mfenced><mml:mo>+</mml:mo><mml:msup><mml:mi mathvariant="bold-italic">u</mml:mi><mml:mi>T</mml:mi></mml:msup><mml:msup><mml:mi mathvariant="bold">Q</mml:mi><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msup><mml:mi mathvariant="bold-italic">u</mml:mi><mml:mo>,</mml:mo><?xmltex \hack{$\egroup}?></mml:mrow></mml:mtd></mml:mlabeledtr></mml:mtable></mml:math></disp-formula>

            where <inline-formula><mml:math id="M165" display="inline"><mml:mi mathvariant="bold">Q</mml:mi></mml:math></inline-formula> is a block diagonal matrix with <inline-formula><mml:math id="M166" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold">S</mml:mi><mml:mi>x</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M167" display="inline"><mml:mrow><mml:msub><mml:mi>S</mml:mi><mml:mi>f</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> on the
diagonal. In the case of an underdetermined problem, we solve for <inline-formula><mml:math id="M168" display="inline"><mml:mi mathvariant="bold-italic">u</mml:mi></mml:math></inline-formula>
with the least-squares formula,

                <disp-formula specific-use="align" content-type="numbered"><mml:math id="M169" display="block"><mml:mtable displaystyle="true"><mml:mtr><mml:mtd><mml:mrow><mml:mstyle displaystyle="true" class="stylechange"/><mml:mi mathvariant="bold-italic">u</mml:mi><mml:mo>=</mml:mo></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><mml:msup><mml:mi mathvariant="bold">QC</mml:mi><mml:mi>T</mml:mi></mml:msup><mml:msup><mml:mi mathvariant="bold">E</mml:mi><mml:mi>T</mml:mi></mml:msup><mml:msup><mml:mfenced close="]" open="["><mml:msup><mml:mi mathvariant="bold">ECQC</mml:mi><mml:mi>T</mml:mi></mml:msup><mml:msup><mml:mi mathvariant="bold">E</mml:mi><mml:mi>T</mml:mi></mml:msup><mml:mo>+</mml:mo><mml:mi>W</mml:mi></mml:mfenced><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:mtd></mml:mtr><mml:mlabeledtr id="Ch1.E12"><mml:mtd/><mml:mtd><mml:mstyle class="stylechange" displaystyle="true"/></mml:mtd><mml:mtd><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><mml:mfenced close="}" open="{"><mml:mi>y</mml:mi><mml:mfenced close=")" open="("><mml:mi mathvariant="normal">Δ</mml:mi><mml:msub><mml:mi>t</mml:mi><mml:mi>y</mml:mi></mml:msub></mml:mfenced><mml:mo>-</mml:mo><mml:mi mathvariant="bold">E</mml:mi><mml:mi mathvariant="script">R</mml:mi><mml:mfenced open="(" close=")"><mml:mi mathvariant="normal">Δ</mml:mi><mml:msub><mml:mi>t</mml:mi><mml:mi>y</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:mn mathvariant="normal">0</mml:mn></mml:mfenced><mml:mfenced open="[" close="]"><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mn mathvariant="normal">0</mml:mn></mml:msub><mml:mo>(</mml:mo><mml:mn mathvariant="normal">0</mml:mn><mml:mo>)</mml:mo></mml:mfenced></mml:mfenced><mml:mo>,</mml:mo></mml:mrow></mml:mtd></mml:mlabeledtr></mml:mtable></mml:math></disp-formula>

            but note the nonlinearity due to <inline-formula><mml:math id="M170" display="inline"><mml:mrow><mml:mi mathvariant="script">R</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="normal">Δ</mml:mi><mml:msub><mml:mi>t</mml:mi><mml:mi>y</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>, 0). To handle this
quantity, we update the state transition and controllability matrices
iteratively, which is identical to the method of total inversion
<xref ref-type="bibr" rid="bib1.bibx57" id="paren.35"/>. In cases where it saves
computations, we employ the overdetermined least-squares formula instead of
Eq. (<xref ref-type="disp-formula" rid="Ch1.E12"/>). The full nonlinear model is run with the updated
controls to produce the improved first-guess trajectory, <inline-formula><mml:math id="M171" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>, for
the first segment (0 <inline-formula><mml:math id="M172" display="inline"><mml:mo>≤</mml:mo></mml:math></inline-formula> <inline-formula><mml:math id="M173" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula> <inline-formula><mml:math id="M174" display="inline"><mml:mo>≤</mml:mo></mml:math></inline-formula> <inline-formula><mml:math id="M175" display="inline"><mml:mrow><mml:mi mathvariant="normal">Δ</mml:mi><mml:msub><mml:mi>t</mml:mi><mml:mi>y</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>). The algorithm proceeds
sequentially <inline-formula><mml:math id="M176" display="inline"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi>y</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> <inline-formula><mml:math id="M177" display="inline"><mml:mo>-</mml:mo></mml:math></inline-formula> 1 times, where the terminal state from one segment becomes
the initial condition for the next.</p><?xmltex \hack{\newpage}?>
</sec>
<sec id="Ch1.S2.SS5">
  <title>Solution for Lagrange multipliers</title>
      <p>We obtain the sensitivity of <inline-formula><mml:math id="M178" display="inline"><mml:mi>J</mml:mi></mml:math></inline-formula> to the initial conditions by taking the
partial derivative,

                <disp-formula specific-use="align" content-type="numbered"><mml:math id="M179" display="block"><mml:mtable displaystyle="true"><mml:mtr><mml:mtd><mml:mrow><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:mo>∂</mml:mo><mml:mi>J</mml:mi></mml:mrow><mml:mrow><mml:mo>∂</mml:mo><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>(</mml:mo><mml:mn mathvariant="normal">0</mml:mn><mml:mo>)</mml:mo></mml:mrow></mml:mfrac></mml:mstyle></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:mstyle displaystyle="true" class="stylechange"/><mml:mo>=</mml:mo><mml:mn mathvariant="normal">2</mml:mn><mml:mi mathvariant="bold-italic">μ</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="normal">Δ</mml:mi><mml:mi>t</mml:mi><mml:mo>)</mml:mo><mml:mo>+</mml:mo><mml:mn mathvariant="normal">2</mml:mn><mml:msubsup><mml:mi mathvariant="bold">S</mml:mi><mml:mi>x</mml:mi><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msubsup><mml:mfenced open="[" close="]"><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub><mml:mo>(</mml:mo><mml:mn mathvariant="normal">0</mml:mn><mml:mo>)</mml:mo><mml:mo>-</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mn mathvariant="normal">0</mml:mn></mml:msub><mml:mo>(</mml:mo><mml:mn mathvariant="normal">0</mml:mn><mml:mo>)</mml:mo></mml:mfenced></mml:mrow></mml:mtd></mml:mtr><mml:mlabeledtr id="Ch1.E13"><mml:mtd/><mml:mtd><mml:mstyle class="stylechange" displaystyle="true"/></mml:mtd><mml:mtd><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><mml:mo>+</mml:mo><mml:mn mathvariant="normal">2</mml:mn><mml:msup><mml:mi mathvariant="bold">E</mml:mi><mml:mi>T</mml:mi></mml:msup><mml:msup><mml:mi>W</mml:mi><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msup><mml:mfenced open="[" close="]"><mml:mi mathvariant="bold">E</mml:mi><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub><mml:mo>(</mml:mo><mml:mn mathvariant="normal">0</mml:mn><mml:mo>)</mml:mo><mml:mo>-</mml:mo><mml:mi>y</mml:mi><mml:mo>(</mml:mo><mml:mn mathvariant="normal">0</mml:mn><mml:mo>)</mml:mo></mml:mfenced><mml:mo>,</mml:mo></mml:mrow></mml:mtd></mml:mlabeledtr></mml:mtable></mml:math></disp-formula>

            where the improved first guess, <inline-formula><mml:math id="M180" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>, is used. Taking the
derivative with respect to the other set of unknowns, we find

                <disp-formula id="Ch1.E14" content-type="numbered"><mml:math id="M181" display="block"><mml:mstyle class="stylechange" displaystyle="true"/><mml:mrow><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:mo>∂</mml:mo><mml:mi>J</mml:mi></mml:mrow><mml:mrow><mml:mo>∂</mml:mo><mml:mi>f</mml:mi><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mfrac></mml:mstyle><mml:mo>=</mml:mo><mml:mn mathvariant="normal">2</mml:mn><mml:msup><mml:mi mathvariant="bold">B</mml:mi><mml:mi>T</mml:mi></mml:msup><mml:mi mathvariant="bold-italic">μ</mml:mi><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mi mathvariant="normal">Δ</mml:mi><mml:mi>t</mml:mi><mml:mo>)</mml:mo><mml:mo>+</mml:mo><mml:mn mathvariant="normal">2</mml:mn><mml:msubsup><mml:mi>S</mml:mi><mml:mi>f</mml:mi><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msubsup><mml:mfenced close="]" open="["><mml:msub><mml:mi>f</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo><mml:mo>-</mml:mo><mml:msub><mml:mi>f</mml:mi><mml:mn mathvariant="normal">0</mml:mn></mml:msub><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mfenced><mml:mo>.</mml:mo></mml:mrow></mml:math></disp-formula>

          With knowledge of these gradients, we could improve the initial conditions
and forcing, but both Eqs. (<xref ref-type="disp-formula" rid="Ch1.E13"/>) and (<xref ref-type="disp-formula" rid="Ch1.E14"/>) depend
upon the Lagrange multipliers, <inline-formula><mml:math id="M182" display="inline"><mml:mrow><mml:mi mathvariant="bold-italic">μ</mml:mi><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>, that we must solve for first.</p>
      <p>Extending the partial derivative of <inline-formula><mml:math id="M183" display="inline"><mml:mi>J</mml:mi></mml:math></inline-formula> with respect to <inline-formula><mml:math id="M184" display="inline"><mml:mrow><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M185" display="inline"><mml:mrow><mml:mi mathvariant="bold-italic">μ</mml:mi><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>
at all times, we recover the Euler–Lagrange (or “normal”) equations.
The Lagrange multipliers are determined by time stepping backward in time:

                <disp-formula id="Ch1.E15" content-type="numbered"><mml:math id="M186" display="block"><mml:mstyle displaystyle="true" class="stylechange"/><mml:mrow><mml:mstyle displaystyle="true" class="stylechange"/><mml:mi mathvariant="bold-italic">μ</mml:mi><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:mi mathvariant="bold">A</mml:mi><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:msup><mml:mo>)</mml:mo><mml:mi>T</mml:mi></mml:msup><mml:mi mathvariant="bold-italic">μ</mml:mi><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mi mathvariant="normal">Δ</mml:mi><mml:mi>t</mml:mi><mml:mo>)</mml:mo><mml:mo>+</mml:mo><mml:msup><mml:mi mathvariant="bold">E</mml:mi><mml:mi>T</mml:mi></mml:msup><mml:msup><mml:mi>W</mml:mi><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msup><mml:mfenced open="[" close="]"><mml:mi mathvariant="bold">E</mml:mi><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo><mml:mo>-</mml:mo><mml:mi>y</mml:mi><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mfenced><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>

          where the last term on the right-hand side only appears if an observation,
<inline-formula><mml:math id="M187" display="inline"><mml:mrow><mml:mi>y</mml:mi><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>, is available. Initial conditions are given by

                <disp-formula id="Ch1.E16" content-type="numbered"><mml:math id="M188" display="block"><mml:mstyle displaystyle="true" class="stylechange"/><mml:mrow><mml:mstyle displaystyle="true" class="stylechange"/><mml:mi mathvariant="bold-italic">μ</mml:mi><mml:mo>(</mml:mo><mml:mi>T</mml:mi><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:msup><mml:mi mathvariant="bold">E</mml:mi><mml:mi>T</mml:mi></mml:msup><mml:msup><mml:mi>W</mml:mi><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msup><mml:mfenced close="]" open="["><mml:mi mathvariant="bold">E</mml:mi><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub><mml:mo>(</mml:mo><mml:mi>T</mml:mi><mml:mo>)</mml:mo><mml:mo>-</mml:mo><mml:mi>y</mml:mi><mml:mo>(</mml:mo><mml:mi>T</mml:mi><mml:mo>)</mml:mo></mml:mfenced><mml:mo>.</mml:mo></mml:mrow></mml:math></disp-formula>

          Equations (<xref ref-type="disp-formula" rid="Ch1.E15"/>) and (<xref ref-type="disp-formula" rid="Ch1.E16"/>) are
collectively known as the <italic>adjoint model</italic>
<xref ref-type="bibr" rid="bib1.bibx6" id="paren.36"><named-content content-type="pre">e.g.,</named-content></xref>, where the model–observation
misfit is part of the adjoint model forcing, <inline-formula><mml:math id="M189" display="inline"><mml:mrow><mml:msup><mml:mi mathvariant="bold">E</mml:mi><mml:mi>T</mml:mi></mml:msup><mml:msup><mml:mi>W</mml:mi><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msup><mml:mo>[</mml:mo><mml:mi mathvariant="bold">E</mml:mi><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> <inline-formula><mml:math id="M190" display="inline"><mml:mo>-</mml:mo></mml:math></inline-formula> <inline-formula><mml:math id="M191" display="inline"><mml:mrow><mml:mi>y</mml:mi><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo><mml:mo>]</mml:mo></mml:mrow></mml:math></inline-formula>.
Now the result of Eq. (<xref ref-type="disp-formula" rid="Ch1.E15"/>) can be
substituted into Eqs. (<xref ref-type="disp-formula" rid="Ch1.E13"/>) and (<xref ref-type="disp-formula" rid="Ch1.E14"/>) to solve for the gradients.</p>
      <p>In summary, the method of Lagrange multipliers is implemented with the
following steps. Starting with a guess for the initial conditions and
forcing, <inline-formula><mml:math id="M192" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mn mathvariant="normal">0</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula>(0) and <inline-formula><mml:math id="M193" display="inline"><mml:mrow><mml:msub><mml:mi>f</mml:mi><mml:mn mathvariant="normal">0</mml:mn></mml:msub><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>, we improve the first guess and solve
for the full trajectory, <inline-formula><mml:math id="M194" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>, where the method of total inversion
attempts to account for nonlinearities by re-linearizing about each successive
update <xref ref-type="bibr" rid="bib1.bibx57" id="paren.37"/>. The adjoint model is then
run backward in time to solve for the sensitivity of <inline-formula><mml:math id="M195" display="inline"><mml:mi>J</mml:mi></mml:math></inline-formula> with respect to the
two types of unknowns, <inline-formula><mml:math id="M196" display="inline"><mml:mi mathvariant="bold-italic">x</mml:mi></mml:math></inline-formula>(0) and <inline-formula><mml:math id="M197" display="inline"><mml:mrow><mml:mi>f</mml:mi><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>. Here we use a quasi-Newton
gradient-descent optimization <xref ref-type="bibr" rid="bib1.bibx47" id="paren.38"/> to update these
uncertain control parameters. Because the model is nonlinear, the
tangent-linear model, <inline-formula><mml:math id="M198" display="inline"><mml:mrow><mml:mi mathvariant="bold">A</mml:mi><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>, will depend upon the nonlinear model
trajectory, and we re-run the full nonlinear model to get an updated
trajectory that will replace <inline-formula><mml:math id="M199" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>. Forward-adjoint model
integrations are repeated until <inline-formula><mml:math id="M200" display="inline"><mml:mi>J</mml:mi></mml:math></inline-formula> has an acceptable value by a
<inline-formula><mml:math id="M201" display="inline"><mml:mrow><mml:msup><mml:mi mathvariant="italic">χ</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula> statistical test. Here we consider the state estimation method successful if
any solution that acceptably fits the data is found. We acknowledge that many
of the cases presented here are underdetermined, and thus we expect those
solutions to not be unique. We emphasize, however, that finding any solution
would be a breakthrough, as this test has been difficult to satisfy with
chaotic models and the Lagrange multiplier method.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F2"><caption><p>Cost-function values as a function of initial angle, <inline-formula><mml:math id="M202" display="inline"><mml:mi mathvariant="italic">θ</mml:mi></mml:math></inline-formula>, and
angular velocity, <inline-formula><mml:math id="M203" display="inline"><mml:mi mathvariant="italic">ω</mml:mi></mml:math></inline-formula>, for an observational time window of 50 s. The
local minimum found by the optimization (yellow dot) is not the absolute
minimum value of the cost function.</p></caption>
          <?xmltex \igopts{width=199.169291pt}?><graphic xlink:href="https://npg.copernicus.org/articles/24/351/2017/npg-24-351-2017-f02.png"/>

        </fig>

</sec>
</sec>
<sec id="Ch1.S3">
  <title>Results</title>
<sec id="Ch1.S3.SS1">
  <title>Tracking chaotic transitions</title>
      <p>Synthetic observations of the pendulum angle are generated every 2.5 s
over a 50 s time interval, where a random error of 0.5 rad is added to
every observation. We first illustrate the futility of a brute force search
for the optimal initial conditions by re-running the forward model with
combinations of the initial angle and angular velocity in the neighborhood of
the truth. The data contribution to the cost function
(Eq. <xref ref-type="disp-formula" rid="Ch1.E3"/>) is then evaluated for each forward model
trajectory, giving rise to a complex topology where the global minimum is not
immediately visible (Fig. <xref ref-type="fig" rid="Ch1.F2"/>). That the topology of the
cost function is not conducive for gradient descent search was previously
documented in models of convection, quasi-geostrophic flow, and the oceanic
double gyre model <xref ref-type="bibr" rid="bib1.bibx46 bib1.bibx33 bib1.bibx35" id="paren.39"><named-content content-type="pre">e.g.,</named-content></xref>.
When the initial angle and angular velocity are slightly perturbed from the
truth, the cost-function values can become extremely large due to the
divergence of trajectories. Furthermore, the cost function varies irregularly
with many local extrema at locations other than the true solution. The basin
of attraction of the true solution, defined in analogy to a drainage basin on
a topographic map, is much smaller than the observational uncertainty, and
thus, it is likely that the iterations of the adjoint method will converge to
a local, non-global, minimum.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F3"><caption><p>Control of the chaotic pendulum with an improved-first-guess and the
Lagrange multiplier method. <bold>(a)</bold> Observations are taken every 2.5 s
with standard error of 0.5 rad (circles with 1<inline-formula><mml:math id="M204" display="inline"><mml:mi mathvariant="italic">σ</mml:mi></mml:math></inline-formula> error bars). The
trajectory of the pendulum angle (<inline-formula><mml:math id="M205" display="inline"><mml:mrow><mml:mi mathvariant="italic">θ</mml:mi><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>, <bold>a</bold>), angular velocity
(<inline-formula><mml:math id="M206" display="inline"><mml:mrow><mml:mi mathvariant="italic">ω</mml:mi><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>, <bold>b</bold>), and the forcing (<inline-formula><mml:math id="M207" display="inline"><mml:mrow><mml:mi>f</mml:mi><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>, <bold>c</bold>) are given for
a standard first-guess (dashed line), the improved first-guess (gray line),
and the final Lagrange multiplier-based estimate (black line). The standard
first-guess forcing is not shown due to its similarity to the final
estimate.</p></caption>
          <?xmltex \igopts{width=236.157874pt}?><graphic xlink:href="https://npg.copernicus.org/articles/24/351/2017/npg-24-351-2017-f03.png"/>

        </fig>

      <p>We start the application of the methods of Sect. <xref ref-type="sec" rid="Ch1.S2"/> by
determining the first-guess trajectory. Here we implement a model time step of
0.01 s over a 50 s integration time and thus the control vector has 5000 forcing
variables and two initial condition variables. The first-guess
initial conditions, forcing, and trajectory are calculated according to
Sect. <xref ref-type="sec" rid="Ch1.S2.SS3"/>. Despite the assumption of linearity in the
calculation of the first guess, the improved first guess has a trajectory
that is nearly consistent with the error bars of the observations (gray line,
top panel, Fig. <xref ref-type="fig" rid="Ch1.F3"/>). The first guess also tracks the rapid
transitions in the interval, 12 s <inline-formula><mml:math id="M208" display="inline"><mml:mo>&lt;</mml:mo></mml:math></inline-formula> <inline-formula><mml:math id="M209" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula> <inline-formula><mml:math id="M210" display="inline"><mml:mo>&lt;</mml:mo></mml:math></inline-formula> 20 s, where revolutions of the
pendulum occur due to chaotic dynamics. These results contrast with a
seemingly reasonable first-guess trajectory that is determined by a model
simulation initialized with the first observation (<inline-formula><mml:math id="M211" display="inline"><mml:mi mathvariant="italic">θ</mml:mi></mml:math></inline-formula>(0) <inline-formula><mml:math id="M212" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> <inline-formula><mml:math id="M213" display="inline"><mml:mi>y</mml:mi></mml:math></inline-formula>(0)) and zero
angular velocity that goes off track in less than 5 s (see “standard
first guess”, top panel, Fig. <xref ref-type="fig" rid="Ch1.F3"/>). Similar first-guesses are
common in ocean models, where the circulation field is started from rest with
the assumption that geostrophic balance will equilibrate the velocity field
rapidly. Our more sophisticated, but still linear, method of deriving an
improved first guess makes the Lagrange multiplier method more likely to succeed.</p>
      <p>The improved first-guess trajectory is a better fit to the data in large part
due to the updated initial angular velocity (middle panel,
Fig. <xref ref-type="fig" rid="Ch1.F3"/>). Starting with an angular velocity of about 2 rad s<inline-formula><mml:math id="M214" display="inline"><mml:msup><mml:mi/><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula>,
this trajectory has frequent changes in the sign of angular velocity
consistent with reversals in pendulum rotation. Conversely, the standard
trajectory has a long period of strictly positive angular velocity
(10 s <inline-formula><mml:math id="M215" display="inline"><mml:mo>&lt;</mml:mo></mml:math></inline-formula> <inline-formula><mml:math id="M216" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula> <inline-formula><mml:math id="M217" display="inline"><mml:mrow><mml:mo>&lt;</mml:mo><mml:mn mathvariant="normal">4</mml:mn></mml:mrow></mml:math></inline-formula> 0 s) that is inconsistent with the observations.
Another important factor is the position of the pendulum around 10 s after the
start of the integration, when small differences in the state become greatly
magnified. The true pendulum trajectory then enters a period where several
revolutions occur successively. The inaccurate initial velocity causes errors
at this critical time of instability and thus the trajectories diverge.</p>
      <p>For <inline-formula><mml:math id="M218" display="inline"><mml:mi mathvariant="italic">θ</mml:mi></mml:math></inline-formula> and <inline-formula><mml:math id="M219" display="inline"><mml:mi mathvariant="italic">ω</mml:mi></mml:math></inline-formula>, the difference between the improved-first-guess
and final (Lagrange multiplier method) trajectories is smaller than the
changes brought about by the first-guess improvement itself. The method of
Lagrange multipliers acts similarly to a combined filter–smoother that
simultaneously takes into account past and future observations, leading to an
angular velocity evolution with somewhat smaller range while still fitting
the observations. Consequently, the evolution of the pendulum angle is also
smoother, with fewer variations at the observational sampling frequency of
<inline-formula><mml:math id="M220" display="inline"><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>/</mml:mo><mml:mn mathvariant="normal">2.5</mml:mn></mml:mrow></mml:math></inline-formula> s. The full impact of the Lagrange multiplier method only becomes
clear when considering the external forcing in the following section.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F4"><caption><p>Comparison of the reconstructed pendulum and truth.
<bold>(a)</bold> Difference between the observed (circles), standard first guess
(dashed), improved first guess (solid gray line), and final estimate (solid,
black line) of pendulum angle relative to the truth. The standard first-guess
is off scale for much of the panel. <bold>(b)</bold> Similar but for angular
velocity, <inline-formula><mml:math id="M221" display="inline"><mml:mi mathvariant="italic">ω</mml:mi></mml:math></inline-formula>. <bold>(c)</bold> Same but for forcing, <inline-formula><mml:math id="M222" display="inline"><mml:mi>f</mml:mi></mml:math></inline-formula>. The standard
first-guess forcing is suppressed because it is identically
zero.</p></caption>
          <?xmltex \igopts{width=236.157874pt}?><graphic xlink:href="https://npg.copernicus.org/articles/24/351/2017/npg-24-351-2017-f04.png"/>

        </fig>

      <?xmltex \floatpos{t}?><fig id="Ch1.F5" specific-use="star"><caption><p>State estimation of the chaotic pendulum with a reduced set of
two observations with standard error 0.5 rad <bold>(a, c)</bold> and a set of
20 observations with standard error 5 rad. <bold>(a, b)</bold> Comparison of the
observations (circles with 1<inline-formula><mml:math id="M223" display="inline"><mml:mi mathvariant="italic">σ</mml:mi></mml:math></inline-formula> error bars), the standard first-guess
(dashed), the improved first-guess (gray solid line), and the final state
estimate (solid black line), as in Fig. <xref ref-type="fig" rid="Ch1.F3"/>.
<bold>(c, d)</bold> Similar to the top row, except all quantities are referenced
to the truth, <inline-formula><mml:math id="M224" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">θ</mml:mi><mml:mi mathvariant="normal">true</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>, as in Fig. <xref ref-type="fig" rid="Ch1.F4"/>. Again
the standard first guess is off scale for much of the time window. The
improved first guess is nearly identical (and obscured) to the final estimate
in all panels.</p></caption>
          <?xmltex \igopts{width=455.244094pt}?><graphic xlink:href="https://npg.copernicus.org/articles/24/351/2017/npg-24-351-2017-f05.png"/>

        </fig>

</sec>
<sec id="Ch1.S3.SS2">
  <title>Reconstruction of the forcing</title>
      <p>The improved first-guess trajectory better fits the observations than a
standard first-guess, but there are tradeoffs in the estimated forcing
(bottom panel, Fig. <xref ref-type="fig" rid="Ch1.F3"/>). In order to track the chaotic
transitions of the pendulum, the improved first-guess trajectory makes
forcing adjustments that are sometimes strong and abrupt in order to
compensate for previous errors in the trajectory. In other words, these
adjustments take the forcing evolution farther away from the truth that was
used in the standard first-guess trajectory. This is reminiscent of the
small-scale features that are added to the surface forcing of ocean models in
order to fit observations <xref ref-type="bibr" rid="bib1.bibx54" id="paren.40"><named-content content-type="pre">e.g.,</named-content></xref>,
although our case is not a compensation for inaccurate model dynamics because
we are operating under a perfect model assumption. This tradeoff is probably
unacceptable for those wishing to physically interpret the forcing field, and
indicates that the improved first-guess estimate is not a good final solution
despite fitting the observations. The power of the Lagrange multiplier method
is now clear; not only is the final estimate smoother than the first guess,
but also the final forcing estimate accurately reproduces the amplitude and frequency
of the true forcing: <inline-formula><mml:math id="M225" display="inline"><mml:mrow><mml:mi>f</mml:mi><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> <inline-formula><mml:math id="M226" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> 1.5<inline-formula><mml:math id="M227" display="inline"><mml:mrow><mml:mi>cos⁡</mml:mi><mml:mo>(</mml:mo><mml:mn mathvariant="normal">2</mml:mn><mml:mi>t</mml:mi><mml:mo>/</mml:mo><mml:mn mathvariant="normal">3</mml:mn><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>.</p>
      <p>Our identical twin experiment permits a comparison with the truth to diagnose
actual errors even at times without observations. While the improved first
guess appeared to fit the data well, the misfit to the truth displays
considerable structure, including a large deviation around <inline-formula><mml:math id="M228" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula> <inline-formula><mml:math id="M229" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> 45 s to values
that are inconsistent with the observations (top panel,
Fig. <xref ref-type="fig" rid="Ch1.F4"/>). Such a deviation reflects an inaccurate
interpolation between data points during the construction of the improved
first-guess trajectory. The first guess also appears to overfit the
observations, as this estimate deviates from the truth in the neighborhood of
observations with large error (e.g., <inline-formula><mml:math id="M230" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula> <inline-formula><mml:math id="M231" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> 32.5, 42.5 s). The final estimate,
on the other hand, hovers near 0 for the entire time interval, with a
standard deviation of 0.46 rad, very near to the actual observational
uncertainty of 0.50 rad. The final estimate reproduces 72 % of the variance
in the observational error, computed by comparison of the estimated to true
observational error. Visually, the final estimate is closer to the truth than
the observations over the majority of the time interval, indicating that the
Lagrange multiplier method filters out the observational noise even in this
chaotic system.</p>
      <p>For the angular velocity and forcing (middle and bottom panels,
Fig. <xref ref-type="fig" rid="Ch1.F4"/>), the Lagrange multiplier method reproduces the
truth despite the imperfect, sparse observations and the chaotic model
dynamics. The suppression of the abrupt and large changes in forcing is not
simply a smoothing or averaging of the forcing, but instead is seen to
reflect the true forcing, as evidenced by the deviation from true forcing
being small. Strictly speaking, a <inline-formula><mml:math id="M232" display="inline"><mml:mrow><mml:msup><mml:mi mathvariant="italic">χ</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula> posterior statistical test,
discussed later in Sect. 3.4, is needed to assess what is meant by
“small”. Under this condition, the method of Lagrange multipliers is
superior to the first-guess estimate because all components of the solution,
both the state and forcing, can be physically interpreted.</p>
</sec>
<sec id="Ch1.S3.SS3">
  <title>Influence of the data stream</title>
      <p>In the case where only two observations are available (<inline-formula><mml:math id="M233" display="inline"><mml:mi mathvariant="italic">θ</mml:mi></mml:math></inline-formula>(0) <inline-formula><mml:math id="M234" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> <inline-formula><mml:math id="M235" display="inline"><mml:mo>-</mml:mo></mml:math></inline-formula>2 <inline-formula><mml:math id="M236" display="inline"><mml:mo>±</mml:mo></mml:math></inline-formula> 0.5,
<inline-formula><mml:math id="M237" display="inline"><mml:mi mathvariant="italic">θ</mml:mi></mml:math></inline-formula>(50) <inline-formula><mml:math id="M238" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> 10 <inline-formula><mml:math id="M239" display="inline"><mml:mo>±</mml:mo></mml:math></inline-formula> 0.5), the time between observations is greater than
both the fundamental period and the nonlinear timescale of the pendulum.
The nonlinear timescale is defined in this work as the time interval that the
tangent-linear model well approximates the nonlinear dynamics, which depends
upon the size of initial perturbation. Despite the long time between
observations, the estimated trajectory fits both observations via the
Lagrange multiplier method (top left panel, Fig. <xref ref-type="fig" rid="Ch1.F5"/>).
Thus, there appears to be no lower limit on the number of observations
necessary in order to produce an acceptable state estimate with this model.
Of the two steps in our method, it is the first-guess calculation that is
responsible for fitting the data within their errors, and the optimization
with Lagrange multipliers does not substantially change or improve this
estimate. When comparing the estimate to the true trajectory that was
withheld from the reconstruction method, differences larger than 10 rad exist
(bottom left panel, Fig. <xref ref-type="fig" rid="Ch1.F5"/>). Thus, reconstruction of
the full, partially unobserved trajectory without any intervening
observations is a challenging task, as expected. Here, we emphasize that the
first goal in state estimation is to find any model trajectory that fits the
observations, and that in realistic cases we will not know whether the model
interpolates between the observations in the correct way. We recognize that a
chaotic model usually has many trajectories that satisfy the initial and
final times, and thus, any one trajectory is unlikely to reconstruct the
truth at all intervening times.</p>
      <p>In the case that many (<inline-formula><mml:math id="M240" display="inline"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi>y</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> <inline-formula><mml:math id="M241" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> 20) observations are taken but the standard
error is large (5 rad), all observations are again fit within their 1<inline-formula><mml:math id="M242" display="inline"><mml:mi mathvariant="italic">σ</mml:mi></mml:math></inline-formula>
error bars. Strictly speaking, the data appear to be overfit, as 32 % of
the points are expected to reside outside the one standard error level but
none do. This fit has a lower standard deviation, 3.3 rad, than that expected
by the observational error of 5.0 rad, suggesting that the numerical model is
adding information that is complementary to the observations. Only in short
time intervals does the estimate differ from the truth by more than 5 rad,
such as near the observation of the 2<inline-formula><mml:math id="M243" display="inline"><mml:mi mathvariant="italic">σ</mml:mi></mml:math></inline-formula> outlier at <inline-formula><mml:math id="M244" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula> <inline-formula><mml:math id="M245" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> 25 s. Unlike the
case where 20 observations were taken with a smaller standard error of
0.5 rad in the previous section, this estimate is only partially successful at
filtering noise out of the observations. The correlation coefficient between
the estimated and actual observational error is <inline-formula><mml:math id="M246" display="inline"><mml:mi>r</mml:mi></mml:math></inline-formula> <inline-formula><mml:math id="M247" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> 0.43, indicating that
some noise remains. Both case studies in this section indicate that neither
the quality nor the quantity of the data stream affect whether the Lagrange
multiplier method can be successful with a chaotic model.</p>
</sec>
<sec id="Ch1.S3.SS4">
  <title>Influence of the number of controls</title>
      <p>The previous section addresses cases where the forcing is adjusted at every
time step, leading to 5002 control variables. After application of the
Lagrange multiplier method, the resulting value of the cost function
(Eq. <xref ref-type="disp-formula" rid="Ch1.E3"/>) depends upon both the number of control parameters
and the number of observations (Fig. <xref ref-type="fig" rid="Ch1.F6"/>). The value
of <inline-formula><mml:math id="M248" display="inline"><mml:mi>J</mml:mi></mml:math></inline-formula> is always below 1 when 5002 control variables are defined, consistent
with the previously reported results. Here we test the null hypothesis that
the model is consistent with the observations, and we perform a
<inline-formula><mml:math id="M249" display="inline"><mml:mrow><mml:msup><mml:mi mathvariant="italic">χ</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula> statistical test with <inline-formula><mml:math id="M250" display="inline"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi>y</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> degrees of freedom as is appropriate for our cost
function. Our one-sided test statistic is the value of <inline-formula><mml:math id="M251" display="inline"><mml:mi>J</mml:mi></mml:math></inline-formula> where 5 % of
cases are expected to have larger values by chance. For 5002 controls, we
find that all values are small enough that the null hypothesis cannot be
rejected at the 5 % insignificance level (area above black line,
Fig. <xref ref-type="fig" rid="Ch1.F6"/>). In this one-sided test, the Lagrange
multiplier method is expected to acceptably fit the observations if enough
controls are available. In the ocean state estimation problem, all air–sea
fluxes are uncertain and temporally variable; therefore, a large number of controls
can usually be defined.</p>
      <p><?xmltex \hack{\newpage}?>For some cases where only 5 or 10 observations are available, the cost
function is small enough that overfitting may be occurring. In these cases,
we find that the control perturbations necessary to fit the data are very
small, and this impacts the size of the cost function through the <inline-formula><mml:math id="M252" display="inline"><mml:mrow><mml:msub><mml:mi>J</mml:mi><mml:mn mathvariant="normal">0</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula> term
in Eq. (<xref ref-type="disp-formula" rid="Ch1.E6"/>). This effect has been documented in chaotic systems
by the control engineering literature <xref ref-type="bibr" rid="bib1.bibx48" id="paren.41"><named-content content-type="pre">e.g.,</named-content></xref>.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F6"><caption><p>Influence of the number of observations and controls on the ability
to track the chaotic pendulum. The base-10 logarithm of the cost function,
<inline-formula><mml:math id="M253" display="inline"><mml:mrow><mml:msub><mml:mi>log⁡</mml:mi><mml:mn mathvariant="normal">10</mml:mn></mml:msub><mml:mo>(</mml:mo><mml:mi>J</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>, is calculated as a function of the number of evenly spaced
observations over a 50 s window (the abscissa), and the number of effective
degrees of freedom in the control perturbations to the forcing (ordinate). A
<inline-formula><mml:math id="M254" display="inline"><mml:mrow><mml:msup><mml:mi mathvariant="italic">χ</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula> statistical test determines the limit where 95 % of realizations
are expected to have smaller <inline-formula><mml:math id="M255" display="inline"><mml:mi>J</mml:mi></mml:math></inline-formula> values (thick, black line); therefore, cases
below this threshold represent an unacceptable fit to the data at the 5 %
insignificance level.</p></caption>
          <?xmltex \igopts{width=199.169291pt}?><graphic xlink:href="https://npg.copernicus.org/articles/24/351/2017/npg-24-351-2017-f06.png"/>

        </fig>

      <?xmltex \floatpos{t}?><fig id="Ch1.F7" specific-use="star"><caption><p>Escaping an apparent local minimum. <bold>(a)</bold> Cost-function
values as a function of <inline-formula><mml:math id="M256" display="inline"><mml:mi mathvariant="italic">θ</mml:mi></mml:math></inline-formula>(0) and <inline-formula><mml:math id="M257" display="inline"><mml:mi mathvariant="italic">ω</mml:mi></mml:math></inline-formula>(0), in a region of phase
space where a local minimum is present in this slice (yellow dot). The same
cost function, but oriented along a slice with constant <inline-formula><mml:math id="M258" display="inline"><mml:mi mathvariant="italic">θ</mml:mi></mml:math></inline-formula>(0) <inline-formula><mml:math id="M259" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> 0
in the dimensions of <inline-formula><mml:math id="M260" display="inline"><mml:mrow><mml:mi mathvariant="italic">δ</mml:mi><mml:mi>f</mml:mi></mml:mrow></mml:math></inline-formula>(0) and <inline-formula><mml:math id="M261" display="inline"><mml:mi mathvariant="italic">ω</mml:mi></mml:math></inline-formula>(0). The local minimum in the
first two dimensions is no longer an extremum in the other two dimensions or
the combined three-dimensional (3-D) space. This case used <inline-formula><mml:math id="M262" display="inline"><mml:mi>T</mml:mi></mml:math></inline-formula> <inline-formula><mml:math id="M263" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> 10 s for
illustration.</p></caption>
          <?xmltex \igopts{width=341.433071pt}?><graphic xlink:href="https://npg.copernicus.org/articles/24/351/2017/npg-24-351-2017-f07.png"/>

        </fig>

      <p>We investigate the effect of a decrease in the number of controls by redefining
the external forcing control perturbation. For <inline-formula><mml:math id="M264" display="inline"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi>u</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> forcing controls, we define

                <disp-formula id="Ch1.E17" content-type="numbered"><mml:math id="M265" display="block"><mml:mstyle displaystyle="true" class="stylechange"/><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><mml:mi>f</mml:mi><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:msub><mml:mi>f</mml:mi><mml:mn mathvariant="normal">0</mml:mn></mml:msub><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo><mml:mo>+</mml:mo><mml:mi mathvariant="bold">Γ</mml:mi><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo><mml:mfenced open="(" close=")"><mml:mtable class="array" columnalign="center"><mml:mtr><mml:mtd><mml:mrow><mml:mi mathvariant="italic">δ</mml:mi><mml:mi>f</mml:mi><mml:mo>(</mml:mo><mml:mn mathvariant="normal">0</mml:mn><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mrow><mml:mi mathvariant="italic">δ</mml:mi><mml:mi>f</mml:mi><mml:mfenced close=")" open="("><mml:mi>T</mml:mi><mml:mo>/</mml:mo><mml:mfenced close=")" open="("><mml:msub><mml:mi>N</mml:mi><mml:mi>u</mml:mi></mml:msub><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mfenced></mml:mfenced></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mrow><mml:mi mathvariant="italic">δ</mml:mi><mml:mi>f</mml:mi><mml:mfenced close=")" open="("><mml:mn mathvariant="normal">2</mml:mn><mml:mi>T</mml:mi><mml:mo>/</mml:mo><mml:mfenced open="(" close=")"><mml:msub><mml:mi>N</mml:mi><mml:mi>u</mml:mi></mml:msub><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mfenced></mml:mfenced></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mi mathvariant="normal">⋮</mml:mi></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mrow><mml:mi mathvariant="italic">δ</mml:mi><mml:mi>f</mml:mi><mml:mo>(</mml:mo><mml:mi>T</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:mfenced><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>

          where <inline-formula><mml:math id="M266" display="inline"><mml:mrow><mml:mi mathvariant="bold">Γ</mml:mi><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> is a row vector that performs linear interpolation in time,
and <inline-formula><mml:math id="M267" display="inline"><mml:mrow><mml:mi mathvariant="italic">δ</mml:mi><mml:mi>f</mml:mi><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> is only defined at <inline-formula><mml:math id="M268" display="inline"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi>u</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> control times. This formulation
enforces some temporal correlation in the external forcing. Alternatively,
this could be accomplished using a non-diagonal weighting matrix, <inline-formula><mml:math id="M269" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold">S</mml:mi><mml:mi>f</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>.
In this case, the number of degrees of freedom is reduced relative to the
total number of controls.</p>
      <p>For a given number of observations, a decrease in the number of controls
leads to a decrease in the likelihood of a successful fit to the data. The
initialization problem is equivalent to the case with two control variables,
and Fig. <xref ref-type="fig" rid="Ch1.F6"/> suggests that the Lagrange multiplier
method will not produce a good fit to data, as documented by previous works.
The criterion of a good fit also depends upon the number of observations,
where more observations decrease the likelihood of success. To understand why
the data can or cannot be fit, consider that each observation gives a
constraint of the type documented in Eq. (<xref ref-type="disp-formula" rid="Ch1.E10"/>). If all
of these constraints are enforced simultaneously, the problem is formally
underdetermined when the number of controls exceeds the number of
observations, and it is generally likely that a solution exists. The simple
interpretation that the number of controls must exceed the number of
observations does not strictly hold due to the logarithmic scale in
Fig. <xref ref-type="fig" rid="Ch1.F6"/>. Even when the problem is formally
underdetermined, a singular value decomposition analysis of the
controllability matrix, <inline-formula><mml:math id="M270" display="inline"><mml:mi mathvariant="bold">C</mml:mi></mml:math></inline-formula>, reveals that not all controls are
independent and that the data cannot always be fit perfectly. Here, we
identify formally underdetermined cases with a <inline-formula><mml:math id="M271" display="inline"><mml:mi>J</mml:mi></mml:math></inline-formula> value that is unacceptably
high (Fig. <xref ref-type="fig" rid="Ch1.F6"/>); thus, the solvability condition is
sometimes violated. Such a result indicates that the controllability matrix
has an effective rank less than the number of observations, showing the
importance of this quantity as a diagnostic measure of the conditioning of
the estimation problem. In practice, the singular values need not be strictly
zero, as a large discrepancy between the magnitudes of singular values can
give ill conditioning.</p>
      <p>We also find cases where the gradient-descent method is capable of navigating
the complex cost-function topology with Lagrange multiplier sensitivity
information. A slice of <inline-formula><mml:math id="M272" display="inline"><mml:mi>J</mml:mi></mml:math></inline-formula> along <inline-formula><mml:math id="M273" display="inline"><mml:mi mathvariant="italic">ω</mml:mi></mml:math></inline-formula>(0) and <inline-formula><mml:math id="M274" display="inline"><mml:mi mathvariant="italic">θ</mml:mi></mml:math></inline-formula>(0) is focused on a
region of phase space that appears to contain a local minimum, although this
is not the true solution (left panel, Fig. <xref ref-type="fig" rid="Ch1.F7"/>). In the
initial control problem, the optimization would proceed in these two
dimensions and be trapped by the local minimum. Taking a two-dimensional (2-D) slice of the cost
function in the dimension of the initial angular velocity and forcing,
<inline-formula><mml:math id="M275" display="inline"><mml:mrow><mml:mi mathvariant="italic">δ</mml:mi><mml:mi>f</mml:mi></mml:mrow></mml:math></inline-formula>(0), however, the same location may no longer be a local minimum
in the expanded phase space (right panel, Fig. <xref ref-type="fig" rid="Ch1.F7"/>). In our
example, the cost function can be further minimized by decreasing <inline-formula><mml:math id="M276" display="inline"><mml:mrow><mml:mi mathvariant="italic">δ</mml:mi><mml:mi>f</mml:mi></mml:mrow></mml:math></inline-formula>(0),
and the optimization process may eventually get out of the
trap in <inline-formula><mml:math id="M277" display="inline"><mml:mi mathvariant="italic">ω</mml:mi></mml:math></inline-formula>(0)<inline-formula><mml:math id="M278" display="inline"><mml:mrow><mml:mo>/</mml:mo><mml:mi mathvariant="italic">θ</mml:mi></mml:mrow></mml:math></inline-formula>(0) space. Thus, additional dimensions in the
optimization space can sometimes alleviate problems with the gradient-descent algorithm.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F8"><caption><p>Control of the chaotic pendulum with an inaccurate first guess of
the external forcing. <bold>(a)</bold> Observations are taken every 2.5 s with
standard error of 0.5 rad (circles with 1<inline-formula><mml:math id="M279" display="inline"><mml:mi mathvariant="italic">σ</mml:mi></mml:math></inline-formula> error bars). The
trajectory of the pendulum angle (<inline-formula><mml:math id="M280" display="inline"><mml:mrow><mml:mi mathvariant="italic">θ</mml:mi><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>, <bold>a</bold>), angular velocity
(<inline-formula><mml:math id="M281" display="inline"><mml:mrow><mml:mi mathvariant="italic">ω</mml:mi><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>, <bold>b</bold>), and the forcing (<inline-formula><mml:math id="M282" display="inline"><mml:mrow><mml:mi>f</mml:mi><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>, <bold>c</bold>) are given for
a standard first-guess (dashed line), the improved first-guess (gray line),
and the final Lagrange multiplier-based estimate (black line). In the forcing
panel, we also include the true forcing (dash-dot
line).</p></caption>
          <?xmltex \igopts{width=236.157874pt}?><graphic xlink:href="https://npg.copernicus.org/articles/24/351/2017/npg-24-351-2017-f08.png"/>

        </fig>

</sec>
<sec id="Ch1.S3.SS5">
  <title>Influence of prior forcing information</title>
      <p>The previous examples in Sect. 3 proceed with prior information that the
forcing is periodic with an accurate magnitude and phase. A good analogy is
the regular forcing of solar insolation on the ocean surface. Here, we test
the performance of the Lagrange multiplier method with inaccurate prior
information about the forcing, as is a more realistic analogy to the
uncertainty of air–sea fluxes. In particular, our first guess of the forcing,
<inline-formula><mml:math id="M283" display="inline"><mml:mrow><mml:msub><mml:mi>f</mml:mi><mml:mn mathvariant="normal">0</mml:mn></mml:msub><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>, is systematically biased by decreasing <inline-formula><mml:math id="M284" display="inline"><mml:mi>b</mml:mi></mml:math></inline-formula> from 1.5 to 0.75 rad s<inline-formula><mml:math id="M285" display="inline"><mml:msup><mml:mi/><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula>.
The trajectory driven by inaccurate forcing is no worse than the
previous cases with accurate forcing due to the dominance of the chaotic
dynamics of system (Fig. <xref ref-type="fig" rid="Ch1.F8"/>). Using the same observations
as shown in Fig. 3, we find that the chaotic pendulum trajectory is tracked
over multiple nonlinear timescales despite this more stringent test. In this
case, however, the forcing estimate still contains errors relative to the
true forcing calculated with <inline-formula><mml:math id="M286" display="inline"><mml:mi>b</mml:mi></mml:math></inline-formula> <inline-formula><mml:math id="M287" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> 1.5 rad s<inline-formula><mml:math id="M288" display="inline"><mml:msup><mml:mi/><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula>, and some high-frequency
structures remain in <inline-formula><mml:math id="M289" display="inline"><mml:mrow><mml:mi>f</mml:mi><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> (see “improved first guess” in bottom panel,
Fig. <xref ref-type="fig" rid="Ch1.F8"/>). If instead the Lagrange multiplier method is
started from the standard first guess, a smoother and more accurate estimate
of the forcing is obtained at the expense of not fitting the data as well
(see “final estimate” in bottom panel). Any remaining irregular structures
can be handled by imposing temporal correlations as was done in
Sect. <xref ref-type="sec" rid="Ch1.S3.SS4"/>. If such measures are not taken, the
investigator must take care to decide what elements of the forcing represent
true variability and which are compensating for model error. In our simple
system of equations, model errors and forcing errors are mathematically
equivalent. In state estimates with eddy-resolving GCMs, however, small-scale
forcing variability is found near oceanic fronts and the investigator must
determine on a case-by-case basis to what extent it reflects real variability.</p><?xmltex \hack{\newpage}?>
</sec>
</sec>
<sec id="Ch1.S4">
  <title>Discussion</title>
<sec id="Ch1.S4.SS1">
  <title>Relation to Kalman filter/smoother</title>
      <p>Our suggestion that controllability is a key criterion brings our
understanding of the Lagrange multiplier method into closer consistency with
the Kalman filter/smoother (i.e., the combined usage of the Kalman filter and
smoother). Both methods solve the same least-squares problem, and the
solution of a linear problem should not depend upon the chosen method
<xref ref-type="bibr" rid="bib1.bibx62" id="paren.42"><named-content content-type="pre">e.g.,</named-content></xref>. <xref ref-type="bibr" rid="bib1.bibx19" id="text.43"/> found that the problem must be
controllable for the Kalman filter/smoother to be successful. In addition,
the chaotic <xref ref-type="bibr" rid="bib1.bibx37" id="text.44"/> model was tracked with the
Kalman filter/smoother over time windows much longer than the nonlinear
timescale when the system was completely controllable (i.e., all estimated
quantities are uncertain and are treated as control variables)
<xref ref-type="bibr" rid="bib1.bibx13" id="paren.45"/>. Our results suggest that the equivalence of
the Kalman filter/smoother and Lagrange multiplier method may be extended to
nonlinear problems, thus explaining why the chaotic estimation problem may be
solved by the Lagrange multiplier method.</p>
      <p>To recover the true trajectory of a system, observability is also important,
as the estimation problem is the dual of the control problem
<xref ref-type="bibr" rid="bib1.bibx19 bib1.bibx40" id="paren.46"/>.
For the linear problem, <xref ref-type="bibr" rid="bib1.bibx8" id="text.47"/> showed that
complete observability implies asymptotic stability of the Kalman
filter/smoother. Defining observability and controllability conditions for
nonlinear state estimation problems is difficult
<xref ref-type="bibr" rid="bib1.bibx7" id="paren.48"/>. In practice, the important criterion is
ability to solve Eq. (<xref ref-type="disp-formula" rid="Ch1.E10"/>). Strictly speaking, the
solution criteria will therefore depend upon both the controllability matrix,
<inline-formula><mml:math id="M290" display="inline"><mml:mi mathvariant="bold">C</mml:mi></mml:math></inline-formula>, and the observational matrix, <inline-formula><mml:math id="M291" display="inline"><mml:mi mathvariant="bold">E</mml:mi></mml:math></inline-formula>, which combines the issues
of observability and controllability. Here, we suggest the operational
definition that a system is effectively controllable when the solution to
Eq. (<xref ref-type="disp-formula" rid="Ch1.E10"/>), generalized to multiple observations, exists.</p>
      <p>Related to the idea of observability, <xref ref-type="bibr" rid="bib1.bibx61" id="text.49"/> stated that
“problems owing to the multiple minima in the cost function can always be
overcome by having enough observations to keep the estimates close to the
true state.” To evaluate this statement, we emphasize that there are two
levels of successful reconstruction: (1) one that accurately fits the data,
and (2) one that accurately fits the truth at all times and locations.
Criterion (1) has been our metric for success in this work, as in real-world
problems, criterion (2) cannot be tested. Here we have shown that only
controllability is necessary for (1) even with a nonlinear system. In
addition, we show that the data can still be fit even if very few
observations are available, as an off-track estimate can be righted by
precise adjustments to the forcing (recall Fig. <xref ref-type="fig" rid="Ch1.F3"/>). That
short-lived forcing adjustments can put the estimate on track is likely a
consequence of the nonlinear dynamics of our particular problem, although we
believe that an eddy-resolving ocean general circulation model could behave
the same way. Our interpretation is consistent with work in the control of
chaotic systems. Engineers have described the control of a chaotic system as
being “easier” than control of other systems because the necessary control
adjustments are very small <xref ref-type="bibr" rid="bib1.bibx48" id="paren.50"><named-content content-type="pre">e.g.,</named-content></xref>.</p>
</sec>
<sec id="Ch1.S4.SS2">
  <title>Comparison of controllability and stability metrics</title>
      <p>In this section, we compare criteria for the success of the Lagrange
multiplier method. Previously suggested criteria include Lyapunov exponents
or other stability metrics of the tangent-linear model
<xref ref-type="bibr" rid="bib1.bibx34" id="paren.51"><named-content content-type="pre">e.g.,</named-content></xref>. The tangent-linear matrix has
eigenvalues with absolute value greater than one when linearized about a
state in the upper-half plane (<inline-formula><mml:math id="M292" display="inline"><mml:mrow><mml:mi mathvariant="italic">π</mml:mi><mml:mo>/</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:mrow></mml:math></inline-formula> <inline-formula><mml:math id="M293" display="inline"><mml:mo>&lt;</mml:mo></mml:math></inline-formula> <inline-formula><mml:math id="M294" display="inline"><mml:mi mathvariant="italic">θ</mml:mi></mml:math></inline-formula> <inline-formula><mml:math id="M295" display="inline"><mml:mo>&lt;</mml:mo></mml:math></inline-formula> 3<inline-formula><mml:math id="M296" display="inline"><mml:mrow><mml:mi mathvariant="italic">π</mml:mi><mml:mo>/</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:mrow></mml:math></inline-formula>), reflecting the
divergence of neighboring nonlinear trajectories when a pendulum perturbed
towards the horizontal is more rapidly accelerated downwards. Conversely, the
lower-half plane is locally linearly stable. The unforced pendulum with
initial conditions in the upper-half plane is episodically unstable, until
damping brings the pendulum permanently into a stable configuration in the
lower-half plane.</p>
      <p>Here we investigate the influence of stability versus that of nonlinearity.
The pendulum is a useful system because it is easily modified to have
four distinct dynamical states: (1) nonlinear, unstable, (2) nonlinear, stable,
(3) linear, stable, and (4) linear, unstable. Case (1) is the original dynamical
equation for the pendulum (Eq. <xref ref-type="disp-formula" rid="Ch1.E1"/>). By restricting the
phase space to the lower-half plane, the pendulum is locally linearly stable
at all times, although it is still nonlinear (Case 2). When the pendulum is
linearized by the small-angle approximation with the linear term of the
Taylor series expansion (<inline-formula><mml:math id="M297" display="inline"><mml:mrow><mml:mi>sin⁡</mml:mi><mml:mi mathvariant="italic">θ</mml:mi></mml:mrow></mml:math></inline-formula> <inline-formula><mml:math id="M298" display="inline"><mml:mo>≈</mml:mo></mml:math></inline-formula> <inline-formula><mml:math id="M299" display="inline"><mml:mi mathvariant="italic">θ</mml:mi></mml:math></inline-formula>, see
Appendix <xref ref-type="sec" rid="App1.Ch1.S1.SS2"/>), we obtain a linear, stable model (Case 3). If
instead the pendulum is linearized around its apex, the sign of <inline-formula><mml:math id="M300" display="inline"><mml:mi mathvariant="italic">θ</mml:mi></mml:math></inline-formula> in
the linearized equation is reversed, rendering the system linear but unstable (Case 4).</p>
      <p>We revisit the problem of estimating the initial angle when <inline-formula><mml:math id="M301" display="inline"><mml:mi mathvariant="italic">θ</mml:mi></mml:math></inline-formula> is
observed at the final time. A synthetic observation is generated by running
the model with initial displacement of <inline-formula><mml:math id="M302" display="inline"><mml:mrow><mml:mo>-</mml:mo><mml:mi mathvariant="italic">π</mml:mi><mml:mo>/</mml:mo><mml:mn mathvariant="normal">6</mml:mn></mml:mrow></mml:math></inline-formula> rad and zero velocity.
Assuming a perfect model and observation, the shape of the cost function is
generated by changing the initial conditions and evaluating <inline-formula><mml:math id="M303" display="inline"><mml:mi>J</mml:mi></mml:math></inline-formula>. In the two
nonlinear cases, a slice of the cost function contains many local minima that
emerge when the time window is extended from 5 to 50 s
(Fig. <xref ref-type="fig" rid="Ch1.F9"/>, lower panels). The cost function in the
nonlinear, stable case deviates from a parabola because the state transition
matrix is non-self-adjoint and non-normal growth occurs <xref ref-type="bibr" rid="bib1.bibx14 bib1.bibx15" id="paren.52"/>. Thus, even nonlinear
models that are stable are subject to local minima, and linear stability is
not always a good metric to determine whether a gradient-descent search will
successfully find the global minimum.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F9"><caption><p>Cost function with respect to the initial pendulum angle. A
synthetic observation was made from a model run with initial angle,
<inline-formula><mml:math id="M304" display="inline"><mml:mi mathvariant="italic">θ</mml:mi></mml:math></inline-formula> <inline-formula><mml:math id="M305" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> <inline-formula><mml:math id="M306" display="inline"><mml:mrow><mml:mo>-</mml:mo><mml:mi mathvariant="italic">π</mml:mi><mml:mo>/</mml:mo><mml:mn mathvariant="normal">6</mml:mn></mml:mrow></mml:math></inline-formula>. The time between the initial state and the cost-function evaluation is 0.5, 5, or 50 s. <bold>(a)</bold> Linear, stable
pendulum. <bold>(b)</bold> Linear, unstable pendulum. <bold>(c)</bold> Nonlinear,
stable pendulum. <bold>(d)</bold> Nonlinear, unstable pendulum. Notice the wider
scale for <inline-formula><mml:math id="M307" display="inline"><mml:mi mathvariant="italic">θ</mml:mi></mml:math></inline-formula> in the lower, right panel. The pendulum's dynamical
regimes are further explained in the text. Reproduced with permission from
<xref ref-type="bibr" rid="bib1.bibx21" id="text.53"/>.</p></caption>
          <?xmltex \igopts{width=236.157874pt}?><graphic xlink:href="https://npg.copernicus.org/articles/24/351/2017/npg-24-351-2017-f09.png"/>

        </fig>

      <p>Conversely, the linear, unstable case does yield a parabolic cost function
(Fig. <xref ref-type="fig" rid="Ch1.F9"/>, upper panels), implying that instability
does not impede the search for the minimum. Again, local linear stability
does not appear to be a good metric for determining the presence of local
minima, because an unstable system may yield a well-behaved function. This
example reinforces the counterintuitive relationship between stability and
local minima, where a linearly stable system does not have a paraboloidal
cost function but an unstable system does. While this reversed relationship
does not always hold, linear stability metrics are not reliable. We suggest
that controllability is a better metric, but note that controllability and
stability are not unrelated, as a system with a growing unstable mode could
lead to a controllability matrix that effectively drops rank.</p>
</sec>
<sec id="Ch1.S4.SS3">
  <title>Relevance to ocean state estimation</title>
      <p>In the Introduction, we remarked on the only ocean state estimate known to
the authors that successfully implemented the Lagrange multiplier method in
an eddy-permitting ocean GCM without any modification to the adjoint model
<xref ref-type="bibr" rid="bib1.bibx23" id="paren.54"/>. In light of the results of this
work, a combination of factors appears to have been responsible for that
success. While the ocean model had <inline-formula><mml:math id="M308" display="inline"><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>/</mml:mo><mml:mn mathvariant="normal">6</mml:mn></mml:mrow></mml:math></inline-formula><inline-formula><mml:math id="M309" display="inline"><mml:msup><mml:mi/><mml:mo>∘</mml:mo></mml:msup></mml:math></inline-formula> <inline-formula><mml:math id="M310" display="inline"><mml:mo>×</mml:mo></mml:math></inline-formula> <inline-formula><mml:math id="M311" display="inline"><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>/</mml:mo><mml:mn mathvariant="normal">6</mml:mn></mml:mrow></mml:math></inline-formula><inline-formula><mml:math id="M312" display="inline"><mml:msup><mml:mi/><mml:mo>∘</mml:mo></mml:msup></mml:math></inline-formula> horizontal
resolution and contained mesoscale eddies (Fig. <xref ref-type="fig" rid="Ch1.F10"/>), the
resolution was not adequate to fully resolve the eddies. Also, the model
domain of the northeast Atlantic Ocean was a relatively quiescent one. Both
factors likely led to the ocean model being more linear than other studies.
In addition, the adjoint model of a coarse-resolution twin was used to form
an improved first guess, which would improve the likelihood of success with
gradient descent much as our method did here. Perhaps most importantly, the
ocean state estimate included air–sea control fields that were updated every
10 days, leading to a total of 5.5 <inline-formula><mml:math id="M313" display="inline"><mml:mo>×</mml:mo></mml:math></inline-formula> 10<inline-formula><mml:math id="M314" display="inline"><mml:msup><mml:mi/><mml:mn mathvariant="normal">6</mml:mn></mml:msup></mml:math></inline-formula> control variables. Given the
rapid adjustment of the ocean to barotropic waves, it is likely that the
system passed the controllability criterion derived in this work.
Controllability could be numerically evaluated in a GCM by a series of
impulse functions: a dynamical equivalent to the passive response recorded by
transit time distributions <xref ref-type="bibr" rid="bib1.bibx12 bib1.bibx27" id="paren.55"><named-content content-type="pre">e.g.,</named-content></xref>. Open
questions include whether the deep ocean is completely controllable by
surface boundary conditions, and whether ocean data require variability at
timescales shorter than 10 days to be introduced through the surface forcing.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F10"><caption><p>Nested view of the <inline-formula><mml:math id="M315" display="inline"><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>/</mml:mo><mml:mn mathvariant="normal">6</mml:mn></mml:mrow></mml:math></inline-formula><inline-formula><mml:math id="M316" display="inline"><mml:msup><mml:mi/><mml:mo>∘</mml:mo></mml:msup></mml:math></inline-formula> regional state estimate of
<xref ref-type="bibr" rid="bib1.bibx23" id="text.56"/> inside the 2<inline-formula><mml:math id="M317" display="inline"><mml:msup><mml:mi/><mml:mo>∘</mml:mo></mml:msup></mml:math></inline-formula> state estimate
of <xref ref-type="bibr" rid="bib1.bibx54" id="text.57"/>. Potential temperature at 310 m
depth, with a contour interval of 1 <inline-formula><mml:math id="M318" display="inline"><mml:msup><mml:mi/><mml:mo>∘</mml:mo></mml:msup></mml:math></inline-formula>C, is shown. The boundary
between the two estimates (thick black line) is discontinuous in temperature
because of the open-boundary control adjustments. Reproduced with permission
from <xref ref-type="bibr" rid="bib1.bibx21" id="text.58"/>.</p></caption>
          <?xmltex \igopts{width=199.169291pt}?><graphic xlink:href="https://npg.copernicus.org/articles/24/351/2017/npg-24-351-2017-f10.pdf"/>

        </fig>

</sec>
</sec>
<sec id="Ch1.S5" sec-type="conclusions">
  <title>Conclusions</title>
      <p>Nonlinearity is not a fundamental obstacle to constraining a model to
observations using the Lagrange multiplier method. On the basis of research
primarily with toy models, chaotic systems were thought to represent such an
obstacle if the estimation time window was too long. Here we find that the
trajectories of the nonlinear pendulum can be tracked over multiple rapid
transitions that are due to chaotic dynamics. The Lagrange multiplier method
is successful under the condition that enough boundary controls are available
through time, and that the system passes a test of controllability. In the
case of the pendulum, the rank of the controllability matrix is a better
metric to predict a success of state estimation rather than a measure of
dynamical stability. The ocean state estimation problem is analogous to the
problem posed here; uncertain air–sea fluxes contain large errors that
require control adjustments through time.</p>
      <p><?xmltex \hack{\newpage}?>Our implementation of the Lagrange multiplier method includes a step to
construct a good first guess that helps the iterative gradient-descent
search. The first-guess method has been developed with implementation in an
ocean GCM in mind. Specifically, sub-problems are defined over the interval
between observations and thus require less memory than a whole-domain
approach. In addition, we suggest that the particular first-guess method of
this work is not the only way to produce a good first guess, and that other
methods would bring the first-guess state close enough to the truth to
increase the likelihood of success. A good example is the Green's function method
<xref ref-type="bibr" rid="bib1.bibx52 bib1.bibx44" id="paren.59"><named-content content-type="pre">e.g.,</named-content></xref>
that selects a subset of the full control variables and makes some linearity
assumptions. Following up the Green's function optimization with a gradient-descent search with the Lagrange multiplier method is therefore a worthwhile
research goal. The results of this work suggest that ocean state estimation
should continue with the Lagrange multiplier method and models that resolve
higher and higher resolution physics.</p>
</sec>

      
      </body>
    <back><notes notes-type="dataavailability">

      <p>No datasets were used or produced in this work.</p>
  </notes><?xmltex \hack{\clearpage}?><app-group>

<app id="App1.Ch1.S1">
  <title>Numerical implementation of pendulum</title>
<sec id="App1.Ch1.S1.SS1">
  <title>Nonlinear pendulum</title>
      <p>The forced, nonlinear pendulum is governed by the following equation,

                <disp-formula id="App1.Ch1.E1" content-type="numbered"><mml:math id="M319" display="block"><mml:mrow><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:msup><mml:mi mathvariant="normal">d</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mi mathvariant="italic">θ</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="normal">d</mml:mi><mml:msup><mml:mi>t</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:mfrac></mml:mstyle><mml:mo>+</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mn mathvariant="normal">1</mml:mn><mml:mi>q</mml:mi></mml:mfrac></mml:mstyle><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:mi mathvariant="normal">d</mml:mi><mml:mi mathvariant="italic">θ</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="normal">d</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:mfrac></mml:mstyle><mml:mo>+</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mi>g</mml:mi><mml:mi>l</mml:mi></mml:mfrac></mml:mstyle><mml:mi>sin⁡</mml:mi><mml:mi mathvariant="italic">θ</mml:mi><mml:mo>=</mml:mo><mml:mi>f</mml:mi><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>

          where the symbols were defined in Sect. <xref ref-type="sec" rid="Ch1.S2"/>. Discretizing in
time, we obtain the symbolic form of the model equation used in the main text:

                <disp-formula specific-use="align" content-type="numbered"><mml:math id="M320" display="block"><mml:mtable displaystyle="true"><mml:mtr><mml:mtd><mml:mrow><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:mi mathvariant="normal">d</mml:mi><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mi mathvariant="normal">d</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:mfrac></mml:mstyle></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mi mathvariant="normal">d</mml:mi><mml:mrow><mml:mi mathvariant="normal">d</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:mfrac></mml:mstyle><mml:mfenced open="(" close=")"><mml:mtable class="array" columnalign="center"><mml:mtr><mml:mtd><mml:mrow><mml:mi mathvariant="italic">ω</mml:mi><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mrow><mml:mi mathvariant="italic">θ</mml:mi><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:mfenced></mml:mrow></mml:mtd></mml:mtr><mml:mlabeledtr id="App1.Ch1.E2"><mml:mtd/><mml:mtd><mml:mstyle displaystyle="true" class="stylechange"/></mml:mtd><mml:mtd><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><mml:mo>=</mml:mo><mml:mfenced open="(" close=")"><mml:mtable class="array" columnalign="center"><mml:mtr><mml:mtd><mml:mrow><mml:mo>-</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mn mathvariant="normal">1</mml:mn><mml:mi>q</mml:mi></mml:mfrac></mml:mstyle><mml:mi mathvariant="italic">ω</mml:mi><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo><mml:mo>-</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mi>g</mml:mi><mml:mi>l</mml:mi></mml:mfrac></mml:mstyle><mml:mi>sin⁡</mml:mi><mml:mi mathvariant="italic">θ</mml:mi><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo><mml:mo>+</mml:mo><mml:mi>f</mml:mi><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mrow><mml:mi mathvariant="italic">ω</mml:mi><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:mfenced><mml:mo>.</mml:mo></mml:mrow></mml:mtd></mml:mlabeledtr></mml:mtable></mml:math></disp-formula>

            If the system is discretized with a forward Euler time step of time <inline-formula><mml:math id="M321" display="inline"><mml:mrow><mml:mi mathvariant="normal">Δ</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:math></inline-formula>,
the discrete-time state space realization is:

                <disp-formula specific-use="align" content-type="numbered"><mml:math id="M322" display="block"><mml:mtable displaystyle="true"><mml:mtr><mml:mtd><mml:mstyle displaystyle="true" class="stylechange"/></mml:mtd><mml:mtd><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><mml:mfenced open="(" close=")"><mml:mtable class="array" columnalign="center"><mml:mtr><mml:mtd><mml:mrow><mml:mi mathvariant="italic">ω</mml:mi><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mrow><mml:mi mathvariant="italic">θ</mml:mi><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:mfenced><mml:mo>=</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mlabeledtr id="App1.Ch1.E3"><mml:mtd/><mml:mtd><mml:mstyle displaystyle="true" class="stylechange"/></mml:mtd><mml:mtd><mml:mrow><mml:mstyle displaystyle="true" class="stylechange"/><mml:mfenced close=")" open="("><mml:mtable class="array" columnalign="center"><mml:mtr><mml:mtd><mml:mrow><mml:mo>(</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>-</mml:mo><mml:mi mathvariant="normal">Δ</mml:mi><mml:mi>t</mml:mi><mml:mo>/</mml:mo><mml:mi>q</mml:mi><mml:mo>)</mml:mo><mml:mi mathvariant="italic">ω</mml:mi><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo><mml:mo>-</mml:mo><mml:mo>(</mml:mo><mml:mi>g</mml:mi><mml:mi>sin⁡</mml:mi><mml:mi mathvariant="italic">θ</mml:mi><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo><mml:mo>/</mml:mo><mml:mi>l</mml:mi><mml:mo>+</mml:mo><mml:mi>f</mml:mi><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo><mml:mo>)</mml:mo><mml:mi mathvariant="normal">Δ</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mrow><mml:mi mathvariant="normal">Δ</mml:mi><mml:mi>t</mml:mi><mml:mi mathvariant="italic">ω</mml:mi><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo><mml:mo>+</mml:mo><mml:mi mathvariant="italic">θ</mml:mi><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:mfenced><mml:mo>,</mml:mo></mml:mrow></mml:mtd></mml:mlabeledtr></mml:mtable></mml:math></disp-formula>

            or simply,

                <disp-formula id="App1.Ch1.E4" content-type="numbered"><mml:math id="M323" display="block"><mml:mstyle class="stylechange" displaystyle="true"/><mml:mrow><mml:mstyle displaystyle="true" class="stylechange"/><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:mi mathvariant="script">L</mml:mi><mml:mo>[</mml:mo><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo><mml:mo>,</mml:mo><mml:mi>f</mml:mi><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo><mml:mo>]</mml:mo><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>

          <?xmltex \hack{\newpage}?><?xmltex \hack{\noindent}?>where <inline-formula><mml:math id="M324" display="inline"><mml:mi mathvariant="script">L</mml:mi></mml:math></inline-formula> is a nonlinear operator due to the sine
function. For use in the Euler–Lagrange equations, we also produce the following
linearized operators:

                <disp-formula specific-use="align" content-type="numbered"><mml:math id="M325" display="block"><mml:mtable displaystyle="true"><mml:mtr><mml:mtd><mml:mrow><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:mo>∂</mml:mo><mml:mi mathvariant="script">L</mml:mi></mml:mrow><mml:mrow><mml:mo>∂</mml:mo><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mfrac></mml:mstyle><mml:mo>≡</mml:mo></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:mstyle displaystyle="true" class="stylechange"/><mml:mi mathvariant="bold">A</mml:mi><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:mfenced close=")" open="("><mml:mtable class="array" columnalign="center center"><mml:mtr><mml:mtd><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>-</mml:mo><mml:mi mathvariant="normal">Δ</mml:mi><mml:mi>t</mml:mi><mml:mo>/</mml:mo><mml:mi>q</mml:mi></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:mo>-</mml:mo><mml:mi>g</mml:mi><mml:mi>cos⁡</mml:mi><mml:mi mathvariant="italic">θ</mml:mi><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo><mml:mi mathvariant="normal">Δ</mml:mi><mml:mi>t</mml:mi><mml:mo>/</mml:mo><mml:mi>l</mml:mi></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mrow><mml:mi mathvariant="normal">Δ</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:mtd><mml:mtd><mml:mn mathvariant="normal">1</mml:mn></mml:mtd></mml:mtr></mml:mtable></mml:mfenced><mml:mo>,</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mlabeledtr id="App1.Ch1.E5"><mml:mtd/><mml:mtd><mml:mstyle class="stylechange" displaystyle="true"/></mml:mtd><mml:mtd><mml:mrow><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:mo>∂</mml:mo><mml:mi mathvariant="script">L</mml:mi></mml:mrow><mml:mrow><mml:mo>∂</mml:mo><mml:mi>f</mml:mi><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mfrac></mml:mstyle><mml:mo>≡</mml:mo><mml:mi mathvariant="bold">B</mml:mi><mml:mo>=</mml:mo><mml:mfenced close=")" open="("><mml:mtable class="array" columnalign="center"><mml:mtr><mml:mtd><mml:mrow><mml:mi mathvariant="normal">Δ</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mn mathvariant="normal">0</mml:mn></mml:mtd></mml:mtr></mml:mtable></mml:mfenced><mml:mo>.</mml:mo></mml:mrow></mml:mtd></mml:mlabeledtr></mml:mtable></mml:math></disp-formula>

            Here we use a second-order Taylor time stepping (i.e., midpoint forward Euler
method) for increased accuracy. We code the tangent-linear model in
accordance with differentiation rules for numerical codes
<xref ref-type="bibr" rid="bib1.bibx25" id="paren.60"/>, and we run this linearized model with
perturbations to all elements of <inline-formula><mml:math id="M326" display="inline"><mml:mrow><mml:mi mathvariant="bold-italic">δ</mml:mi><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M327" display="inline"><mml:mrow><mml:mi mathvariant="italic">δ</mml:mi><mml:mi>f</mml:mi><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> to recover the values. Numerical
parameters include <inline-formula><mml:math id="M328" display="inline"><mml:mrow><mml:mi mathvariant="normal">Δ</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:math></inline-formula> <inline-formula><mml:math id="M329" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> 0.01 s, <inline-formula><mml:math id="M330" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">ω</mml:mi><mml:mi mathvariant="normal">true</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>(0) <inline-formula><mml:math id="M331" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> 1.2959 rad s<inline-formula><mml:math id="M332" display="inline"><mml:msup><mml:mi/><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula>,
<inline-formula><mml:math id="M333" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">θ</mml:mi><mml:mi mathvariant="normal">true</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>(0) <inline-formula><mml:math id="M334" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> <inline-formula><mml:math id="M335" display="inline"><mml:mo>-</mml:mo></mml:math></inline-formula>2.4667 rad, and the forcing phase,
<inline-formula><mml:math id="M336" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">ϕ</mml:mi><mml:mi mathvariant="normal">true</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>(0) <inline-formula><mml:math id="M337" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> 0.3412 rad.</p>
</sec>
<sec id="App1.Ch1.S1.SS2">
  <title>Linear, stable pendulum</title>
      <p>The linear, stable pendulum is derived with the small-angle approximation.
This approximation is a linearization around zero displacement

                <disp-formula id="App1.Ch1.E6" content-type="numbered"><mml:math id="M338" display="block"><mml:mstyle class="stylechange" displaystyle="true"/><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><?xmltex \hack{\hbox\bgroup\fontsize{9.5}{9.5}\selectfont$\displaystyle}?><mml:mfenced close=")" open="("><mml:mtable class="array" columnalign="center"><mml:mtr><mml:mtd><mml:mrow><mml:mi mathvariant="italic">ω</mml:mi><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mi mathvariant="normal">Δ</mml:mi><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mrow><mml:mi mathvariant="italic">θ</mml:mi><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mi mathvariant="normal">Δ</mml:mi><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:mfenced><mml:mo>=</mml:mo><mml:mfenced close=")" open="("><mml:mtable class="array" columnalign="center center"><mml:mtr><mml:mtd><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>-</mml:mo><mml:mi mathvariant="normal">Δ</mml:mi><mml:mi>t</mml:mi><mml:mo>/</mml:mo><mml:mi>q</mml:mi></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:mo>-</mml:mo><mml:mi>g</mml:mi><mml:mi mathvariant="normal">Δ</mml:mi><mml:mi>t</mml:mi><mml:mo>/</mml:mo><mml:mi>l</mml:mi></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mrow><mml:mi mathvariant="normal">Δ</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:mtd><mml:mtd><mml:mn mathvariant="normal">1</mml:mn></mml:mtd></mml:mtr></mml:mtable></mml:mfenced><mml:mfenced open="(" close=")"><mml:mtable class="array" columnalign="center"><mml:mtr><mml:mtd><mml:mrow><mml:mi mathvariant="italic">ω</mml:mi><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mrow><mml:mi mathvariant="italic">θ</mml:mi><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:mfenced><mml:mo>.</mml:mo><?xmltex \hack{$\egroup}?></mml:mrow></mml:math></disp-formula>

          The more general tangent-linear model is re-linearized around a changing
nonlinear model trajectory.</p><?xmltex \hack{\clearpage}?>
</sec>
</app>
  </app-group><notes notes-type="competinginterests">

      <p>The authors declare that they have no conflict of interest.</p>
  </notes><ack><title>Acknowledgements</title><p>We thank Geir Evensen, Armin Köhl, Olivier Marchal, and Eli Tziperman for
discussions on this topic over the last decade, and to Jacques Verron for his
note that has encouraged this work. Geoffrey Gebbie also acknowledges
Carl Wunsch, Patrick Heimbach, Detlef Stammer, and Julio Sheinbaum for
guidance getting started on this project. Geoffrey Gebbie was funded through
the Ocean and Climate Change Institute of the Woods Hole Oceanographic
Institution. Tsung-Lin Hsieh was funded by the Arthur Vining Davis
Foundations Fund for Summer Student Fellows through the Woods Hole
Oceanographic Institution. <?xmltex \hack{\newline}?><?xmltex \hack{\newline}?>
Edited by: Zoltan Toth <?xmltex \hack{\newline}?>
Reviewed by: three anonymous referees</p></ack><ref-list>
    <title>References</title>

      <ref id="bib1.bibx1"><label>Arbic et al.(2010)Arbic, Wallcraft, and Metzger</label><mixed-citation>
Arbic, B., Wallcraft, A., and Metzger, E.: Concurrent simulation of the
eddying general circulation and tides in a global ocean model, Ocean Model.,
32, 175–187, 2010.</mixed-citation></ref>
      <ref id="bib1.bibx2"><label>Baker and Gollub(1990)</label><mixed-citation>
Baker, G. L. and Gollub, J. P.: Chaotic Dynamics: An Introduction, Cambridge
University Press, Cambridge, 1990.</mixed-citation></ref>
      <ref id="bib1.bibx3"><label>Bennett(1992)</label><mixed-citation>
Bennett, A. F.: Inverse methods in physical oceanography, in: Cambridge Monographs,
Cambridge University Press, Cambridge, 1992.</mixed-citation></ref>
      <ref id="bib1.bibx4"><label>Bennett(2002)</label><mixed-citation>
Bennett, A. F.: Inverse Modeling of the Ocean and Atmosphere, Cambridge
University Press, Cambridge, 2002.</mixed-citation></ref>
      <ref id="bib1.bibx5"><label>Bonekamp et al.(2001)Bonekamp, Van oldenborgh, and Burgers</label><mixed-citation>
Bonekamp, H., Van Oldenborgh, G. J., and Burgers, G.: Variational Assimilation
of Tropical Atmosphere–Ocean and expendable bathythermograph data in the
Hamburg Ocean Primitive Equation ocean general circulation model, adjusting the
surface fluxes in the tropical ocean, J. Geophys. Res.-Oceans, 106, 16693–16709, 2001.</mixed-citation></ref>
      <ref id="bib1.bibx6"><label>Bugnion et al.(2006)Bugnion, Hill, and Stone</label><mixed-citation>
Bugnion, V., Hill, C., and Stone, P. H.: An adjoint analysis of the meridional
overturning circulation in a hybrid coupled model, J. Climate, 19, 3751–3767, 2006.</mixed-citation></ref>
      <ref id="bib1.bibx7"><label>Casti(1985)</label><mixed-citation>
Casti, J. L.: Nonlinear system theory, in: vol. 175, Academic Press, Cambridge, MA, USA, 1985.</mixed-citation></ref>
      <ref id="bib1.bibx8"><label>Cohn and Dee(1988)</label><mixed-citation>
Cohn, S. E. and Dee, D. P.: Observability of discretized partial differential
equations, SIAM J. Numer. Anal., 25, 586–617, 1988.</mixed-citation></ref>
      <ref id="bib1.bibx9"><label>Cong et al.(1998)Cong, Ikeda, and Hendry</label><mixed-citation>
Cong, L. Z., Ikeda, M., and Hendry, R. M.: Variational assimilation of Geosat
altimeter data into a two-layer quasi-geostrophic model over the Newfoundland
ridge and basin, J. Geophys. Res., 103, 7719–7734, 1998.</mixed-citation></ref>
      <ref id="bib1.bibx10"><label>Courtier et al.(1994)Courtier, Thepaut, and Hollingsworth</label><mixed-citation>
Courtier, P., Thepaut, J.-N., and Hollingsworth, A.: A Strategy for Operational
Implementation of 4D-Var, Using an Incremental Approach, Q. J. Roy. Meteorol.
Soc., 120, 1367–1387, 1994.</mixed-citation></ref>
      <ref id="bib1.bibx11"><label>Dahleh and Diaz-Bobillo(1999)</label><mixed-citation>
Dahleh, M. A. and Diaz-Bobillo, I.: Control of Uncertain Systems: A Linear
Programming Approach, Pergamon Press, Oxford, UK, 1999.</mixed-citation></ref>
      <ref id="bib1.bibx12"><label>Delhez et al.(1999)Delhez, Campin, Hirst, and Deleersnijder</label><mixed-citation>
Delhez, E., Campin, J., Hirst, A., and Deleersnijder, E.: Toward a general
theory of the age in ocean modelling, Ocean Model., 1, 17–27, 1999.</mixed-citation></ref>
      <ref id="bib1.bibx13"><label>Evensen(1997)</label><mixed-citation>
Evensen, G.: Advanced data assimilation for strongly nonlinear dynamics,
Mon. Weather Rev., 125, 1342–1354, 1997.</mixed-citation></ref>
      <ref id="bib1.bibx14"><label>Farrell(1989)</label><mixed-citation>
Farrell, B.: Optimal excitation of baroclinic waves, J. Atmos. Sci., 46, 1193–1206, 1989.</mixed-citation></ref>
      <ref id="bib1.bibx15"><label>Farrell and Ioannou(1993)</label><mixed-citation>
Farrell, B. and Ioannou, P.: Transient development of perturbations in stratified
shear-flow, J. Atmos. Sci., 50, 2201–2214, 1993.</mixed-citation></ref>
      <ref id="bib1.bibx16"><label>Ferron and Marotzke(2003)</label><mixed-citation>
Ferron, B. and Marotzke, J.: Impact of 4D-Variational Assimilation of WOCE
Hydrography on the Meridional Circulation of the Indian Ocean, Deep-Sea
Res. Pt. II, 50, 2005–2021, 2003.</mixed-citation></ref>
      <ref id="bib1.bibx17"><label>Fukumori(2002)</label><mixed-citation>
Fukumori, I.: A partitioned Kalman filter and smoother, Mon. Weather Rev., 130, 1370–1383, 2002.</mixed-citation></ref>
      <ref id="bib1.bibx18"><label>Fukumori and Malanotte-Rizzoli(1995)</label><mixed-citation>
Fukumori, I. and Malanotte-Rizzoli, P.: An approximate Kalman Filter for
ocean data assimilation: an example with an idealized Gulf-Stream model, J.
Geophys. Res., 100, 6777–6793, 1995.</mixed-citation></ref>
      <ref id="bib1.bibx19"><label>Fukumori et al.(1993)Fukumori, Benveniste, Wunsch, and Haidvogel</label><mixed-citation>
Fukumori, I., Benveniste, J., Wunsch, C., and Haidvogel, D. B.: Assimilation of
Sea Surface Topography into an Ocean Circulation Model Using a Steady-State
Smoother, J. Phys. Oceanogr., 23, 1831–1855, 1993.</mixed-citation></ref>
      <ref id="bib1.bibx20"><label>Gauthier(1992)</label><mixed-citation>
Gauthier, P.: Chaos and Quadri-dimensional Data Assimilation: A Study Based on
the Lorenz Model, Tellus A, 44, 2–17, 1992.</mixed-citation></ref>
      <ref id="bib1.bibx21"><label>Gebbie(2004)</label><mixed-citation>
Gebbie, G.: Subduction in an Eddy-resolving State Estimate of the Northeast
Atlantic Ocean, PhD thesis, Massachusetts Institute of Technology/Woods Hole
Oceanographic Institution Joint Program in Oceanography, Cambridge/Woods Hole, MA, USA, 2004.</mixed-citation></ref>
      <ref id="bib1.bibx22"><label>Gebbie(2007)</label><mixed-citation>Gebbie, G.: Does eddy subduction matter in the Northeast Atlantic Ocean?,
J. Geophys. Res., 112, C06007, <ext-link xlink:href="https://doi.org/10.1029/2006JC003568" ext-link-type="DOI">10.1029/2006JC003568</ext-link>, 2007.</mixed-citation></ref>
      <ref id="bib1.bibx23"><label>Gebbie et al.(2006)Gebbie, Heimbach, and Wunsch</label><mixed-citation>Gebbie, G., Heimbach, P., and Wunsch, C.: Strategies for Nested and Eddy-Permitting
State Estimation, J. Geophys. Res., 111, C10073, <ext-link xlink:href="https://doi.org/10.1029/2005JC003094" ext-link-type="DOI">10.1029/2005JC003094</ext-link>, 2006.</mixed-citation></ref>
      <ref id="bib1.bibx24"><label>Gilbert and Lemaréchal(1989)</label><mixed-citation>
Gilbert, J. C. and Lemaréchal, C.: Some Numerical Experiments with
Variable-storage Quasi-Newton Algorithms, Math. Program., 45, 407–435, 1989.</mixed-citation></ref>
      <ref id="bib1.bibx25"><label>Griewank(2000)</label><mixed-citation>
Griewank, A.: Evaluating Derivatives: Principles and Techniques of Algorithmic
Differentiation, Society for Industrial and Applied Mathematics, Philadelphia, 2000.</mixed-citation></ref>
      <ref id="bib1.bibx26"><label>Griffies et al.(2015)Griffies, Winton, Anderson, Benson, Delworth,
Dufour, Dunne, Goddard, Morrison, Rosati et al.</label><mixed-citation>
Griffies, S. M., Winton, M., Anderson, W. G., Benson, R., Delworth, T. L.,
Dufour, C. O., Dunne, J. P., Goddard, P., Morrison, A. K., Rosati, A., Wittenberg,
A. T., Yin, J., and Zhang, R.: Impacts on ocean heat from transient mesoscale
eddies in a hierarchy of climate models, J. Climate, 28, 952–977, 2015.</mixed-citation></ref>
      <ref id="bib1.bibx27"><label>Haine and Hall(2002)</label><mixed-citation>
Haine, T. W. N. and Hall, T. M.: A generalized transport theory: Water-mass
composition and age, J. Phys. Oceanogr., 32, 1932–1946, 2002.</mixed-citation></ref>
      <ref id="bib1.bibx28"><label>Hall et al.(1982)Hall, Cacuci, and Schlesinger</label><mixed-citation>
Hall, M. C. G., Cacuci, D. G., and Schlesinger, M. E.: Sensitivity Analysis of
a Radiative-Convective Model by the Adjoint Method, J. Atmos. Sci., 39, 2038–2050, 1982.</mixed-citation></ref>
      <ref id="bib1.bibx29"><label>Jazwinski(1970)</label><mixed-citation>
Jazwinski, A.: Stochastic Processes and Filtering Theory, Academic Press, Cambridge, MA, 1970.</mixed-citation></ref>
      <ref id="bib1.bibx30"><label>Keppenne et al.(2005)Keppenne, Rienecker, Kurkowski, Adamec et al.</label><mixed-citation>Keppenne, C. L., Rienecker, M. M., Kurkowski, N. P., and Adamec, D. A.: Ensemble
Kalman filter assimilation of temperature and altimeter data with bias correction
and application to seasonal prediction, Nonlin. Processes Geophys., 12, 491–503,
<ext-link xlink:href="https://doi.org/10.5194/npg-12-491-2005" ext-link-type="DOI">10.5194/npg-12-491-2005</ext-link>, 2005.</mixed-citation></ref>
      <ref id="bib1.bibx31"><label>Köhl and Stammer(2008)</label><mixed-citation>
Köhl, A. and Stammer, D.: Variability of the meridional overturning in the
North Atlantic from the 50-year GECCO state estimation, J. Phys. Oceanogr.,
38, 1913–1930, 2008.</mixed-citation></ref>
      <ref id="bib1.bibx32"><label>Köhl and Willebrand(2002)</label><mixed-citation>
Köhl, A. and Willebrand, J.: An Adjoint Method for the Assimilation of
Statistical Characteristics into Eddy-resolving Ocean Models, Tellus, 54, 406–425, 2002.</mixed-citation></ref>
      <ref id="bib1.bibx33"><label>Köhl and Willebrand(2003)</label><mixed-citation>Köhl, A. and Willebrand, J.: Variational Assimilation of SSH Variability
from TOPEX/POSEIDON and ERS1 into an Eddy-permitting Model of the North Atlantic,
J. Geophys. Res., 108, 3092 <ext-link xlink:href="https://doi.org/10.1029/2001JC000982" ext-link-type="DOI">10.1029/2001JC000982</ext-link>, 2003.</mixed-citation></ref>
      <ref id="bib1.bibx34"><label>Lea et al.(2000)Lea, Allen, and Haine</label><mixed-citation>
Lea, D. J., Allen, M. R., and Haine, T. W. N.: Sensitivity Analysis of the
Climate of a Chaotic System, Tellus A, 52, 523–532, 2000.</mixed-citation></ref>
      <ref id="bib1.bibx35"><label>Lea et al.(2006)Lea, Haine, and Gasparovic</label><mixed-citation>
Lea, D. J., Haine, T. W. N., and Gasparovic, R. F.: Observability of the
Irminger Sea Circulation using Variational Data Assimilation, Q. J. Roy.
Meteorol. Soc., 132, 1545–1576, 2006.</mixed-citation></ref>
      <ref id="bib1.bibx36"><label>LeDimet and Talagrand(1986)</label><mixed-citation>
LeDimet, F. and Talagrand, O.: Variational algorithm for analysis and
assimilation of meteorological observations: Theoretical aspects, Tellus A,
38, 97–110, 1986.</mixed-citation></ref>
      <ref id="bib1.bibx37"><label>Lorenz(1963)</label><mixed-citation>
Lorenz, E. N.: Deterministic, Nonperiodic Flow, J. Atmos. Sci., 20, 130–141, 1963.</mixed-citation></ref>
      <ref id="bib1.bibx38"><label>Malanotte-Rizzoli and Tziperman(1996)</label><mixed-citation>
Malanotte-Rizzoli, P. and Tziperman, E.: The oceanographic data assimilation
problem: overview, motivation and purposes, Elsevier Oceanography Series 61,
Amsterdam, the Netherlands, 3–17, 1996.</mixed-citation></ref>
      <ref id="bib1.bibx39"><label>Maltrud et al.(2010)Maltrud, Bryan, and Peacock</label><mixed-citation>
Maltrud, M., Bryan, F., and Peacock, S.: Boundary impulse response functions in
a century-long eddying global ocean simulation, Environ. Fluid Mech., 10, 275–295, 2010.</mixed-citation></ref>
      <ref id="bib1.bibx40"><label>Marchal(2014)</label><mixed-citation>
Marchal, O.: On the observability of oceanic gyres, J. Phys. Oceanogr., 44, 2498–2523, 2014.</mixed-citation></ref>
      <ref id="bib1.bibx41"><label>Marotzke et al.(1999)Marotzke, Giering, Zhang, Stammer, Hill, and Lee</label><mixed-citation>
Marotzke, J., Giering, R., Zhang, K. Q., Stammer, D., Hill, C., and Lee, T.:
Construction of the Adjoint MIT Ocean General Circulation Model and Application
to Atlantic Heat Transport Sensitivity, J. Geophys. Res., 104, 529–547, 1999.</mixed-citation></ref>
      <ref id="bib1.bibx42"><label>Mazloff et al.(2010)Mazloff, Heimbach, and Wunsch</label><mixed-citation>
Mazloff, M., Heimbach, P., and Wunsch, C.: An eddy-permitting Southern Ocean
state estimate, J. Phys. Oceanogr., 40, 880–899, 2010.</mixed-citation></ref>
      <ref id="bib1.bibx43"><label>McShane(1989)</label><mixed-citation>
McShane, E. J.: The Calculus of Variations from the Beginning Through to
Optimal Control Theory, SIAM J. Control Optim., 27, 916–939, 1989.</mixed-citation></ref>
      <ref id="bib1.bibx44"><label>Menemenlis et al.(2004)Menemenlis, Fukumori, and Lee</label><mixed-citation>
Menemenlis, D., Fukumori, I., and Lee, T.: Using Green's functions to calibrate
an ocean general circulation model, Mon. Weather Rev., 133, 1224–1240, 2004.</mixed-citation></ref>
      <ref id="bib1.bibx45"><label>Miller et al.(1994a)Miller, Ghil, and Gauthiez</label><mixed-citation>
Miller, R. N., Ghil, M., and Gauthiez, F.: Advanced Data Assimilation in
Strongly Nonlinear Systems, J. Atmos. Sci., 51, 1037–1056, 1994a.</mixed-citation></ref>
      <ref id="bib1.bibx46"><label>Miller et al.(1994b)Miller, Zaron, and Bennett</label><mixed-citation>
Miller, R. N., Zaron, E. D., and Bennett, A. F.: Data assimilation in models
with convective adjustment, Mon. Weath. Rev., 122, 2607–2613, 1994b.</mixed-citation></ref>
      <ref id="bib1.bibx47"><label>Nocedal(1980)</label><mixed-citation>
Nocedal, J.: Updating quasi-Newton matrices with limited storage, Math. Comput.,
35, 773–782, 1980.</mixed-citation></ref>
      <ref id="bib1.bibx48"><label>Ott et al.(1990)Ott, Grebogi, and Yorke</label><mixed-citation>
Ott, E., Grebogi, C., and Yorke, J.: Controlling chaos, Phys. Rev. Lett., 64, 1196–1199, 1990.</mixed-citation></ref>
      <ref id="bib1.bibx49"><label>Palmer(1996)</label><mixed-citation>Palmer, T. N.: Predictability of the Atmosphere and Oceans: From Days to
Decades, in: vol. 44 of NATO ASI Series, chap. 3, Decadal Climate Variability:
Dynamics and Predictability, Springer, New York, NY, USA, 83–156, 1996.
 </mixed-citation></ref><?xmltex \hack{\newpage}?>
      <ref id="bib1.bibx50"><label>Pires et al.(1996)Pires, Vautard, and Talagrand</label><mixed-citation>
Pires, C., Vautard, R., and Talagrand, O.: On extending the limits of variational
assimilation in nonlinear chaotic systems, Tellus A, 48, 96–121, 1996.</mixed-citation></ref>
      <ref id="bib1.bibx51"><label>Schröter et al.(1993)Schröter, Seiler, and Wenzel</label><mixed-citation>
Schröter, J., Seiler, U., and Wenzel, M.: Variational Assimilation of
GEOSAT Data into an Eddy-resolving Model of the Gulf Stream Extension Area, J.
Phys. Oceanogr., 23, 925–953, 1993.</mixed-citation></ref>
      <ref id="bib1.bibx52"><label>Stammer and Wunsch(1996)</label><mixed-citation>
Stammer, D. and Wunsch, C.: The determination of the large-scale circulation of
the Pacific Ocean from satellite altimetry using model Green's functions, J.
Geophys. Res., 101, 18409–18432, 1996.</mixed-citation></ref>
      <ref id="bib1.bibx53"><label>Stammer et al.(2002a)Stammer, Wunsch, Fukumori, and Marshall</label><mixed-citation>
Stammer, D., Wunsch, C., Fukumori, I., and Marshall, J.: State Estimation in
Modern Oceanographic Research, EOS, 83, 294–295, 2002a.</mixed-citation></ref>
      <ref id="bib1.bibx54"><label>Stammer et al.(2002b)Stammer, Wunsch, Giering, Eckert,
Heimbach, Marotzke, Adcroft, Hill, and Marshall</label><mixed-citation>Stammer, D., Wunsch, C., Giering, R., Eckert, C., Heimbach, P., Marotzke, J.,
Adcroft, A., Hill, C. N., and Marshall, J.: The global ocean circulation
during 1992–1997, estimated from ocean observations and a general circulation
model, J. Geophys. Res.-Oceans, 107, 3118 <ext-link xlink:href="https://doi.org/10.1029/2001JC000888" ext-link-type="DOI">10.1029/2001JC000888</ext-link>, 2002b.</mixed-citation></ref>
      <ref id="bib1.bibx55"><label>Stammer et al.(2004)Stammer, Ueyoshi, Kohl, Large, Josey, and
Wunsch</label><mixed-citation>Stammer, D., Ueyoshi, K., Kohl, A., Large, W. G., Josey, S. A., and Wunsch, C.:
Estimating air-sea fluxes of heat, freshwater, and momentum through global
ocean data assimilation, J. Geophys. Res., 109, C05023, <ext-link xlink:href="https://doi.org/10.1029/2003JC002082" ext-link-type="DOI">10.1029/2003JC002082</ext-link>, 2004.</mixed-citation></ref>
      <ref id="bib1.bibx56"><label>Tanguay et al.(1995)Tanguay, Bartello, and Gauthier</label><mixed-citation>
Tanguay, M., Bartello, P., and Gauthier, P.: Four-dimensional Data Assimilation
with a Wide Range of Scales, Tellus A, 47, 974–997, 1995.</mixed-citation></ref>
      <ref id="bib1.bibx57"><label>Tarantola and Valette(1982)</label><mixed-citation>
Tarantola, A. and Valette, B.: Generalized Nonlinear Inverse Problems Solved
using the Least Squares Criterion, Rev. Geophys. Space Phys., 20, 219–232, 1982.</mixed-citation></ref>
      <ref id="bib1.bibx58"><label>Thacker and Long(1988)</label><mixed-citation>
Thacker, W. C. and Long, R. B.: Fitting dynamics to data, J. Geophys. Res.,
93, 1227–1240, 1988.</mixed-citation></ref>
      <ref id="bib1.bibx59"><label>Tziperman and Thacker(1989)</label><mixed-citation>
Tziperman, E. and Thacker, W. C.: An Optimal-Control/Adjoint-Equations Approach
to Studying the Oceanic General Circulation, J. Phys. Oceanogr., 19, 1471–1485, 1989.</mixed-citation></ref>
      <ref id="bib1.bibx60"><label>Tziperman et al.(1994)Tziperman, Stone, Cane, and Jarosh</label><mixed-citation>
Tziperman, E., Stone, L., Cane, M. A., and Jarosh, H.: El Niño chaos:
overlapping of resonances between the seasonal cycle and the Pacific
ocean–atmosphere oscillator, Science, 264, 72–74, 1994.</mixed-citation></ref>
      <ref id="bib1.bibx61"><label>Wunsch(1996)</label><mixed-citation>
Wunsch, C.: The Ocean Circulation Inverse Problem, Cambridge University Press, New York, 1996.</mixed-citation></ref>
      <ref id="bib1.bibx62"><label>Wunsch(2010)</label><mixed-citation>
Wunsch, C.: Discrete Inverse and State Estimation Problems. With Geophysical
Fluid Applications, Cambridge University Press, Cambridge, UK, 2010.</mixed-citation></ref>
      <ref id="bib1.bibx63"><label>Wunsch and Heimbach(2007)</label><mixed-citation>
Wunsch, C. and Heimbach, P.: Practical global oceanic state estimation, Physica D,
230, 192–208, 2007.</mixed-citation></ref>
      <ref id="bib1.bibx64"><label>Wunsch et al.(2009)Wunsch, Heimbach, Ponte, and Fukumori</label><mixed-citation>
Wunsch, C., Heimbach, P., Ponte, R., and Fukumori, I.: The global general
circulation of the oceans estimated by the ECCO-Consortium, Oceanography,
20, 88–103, 2009.</mixed-citation></ref>

  </ref-list><app-group content-type="float"><app><title/>

    </app></app-group></back>
    <!--<article-title-html>Controllability, not chaos, key criterion for ocean state estimation</article-title-html>
<abstract-html><p class="p">The Lagrange multiplier method for combining observations and
models (i.e., the adjoint method or <q>4D-VAR</q>) has been avoided or
approximated when the numerical model is highly nonlinear or chaotic. This
approach has been adopted primarily due to difficulties in the initialization
of low-dimensional chaotic models, where the search for optimal initial
conditions by gradient-descent algorithms is hampered by multiple local
minima. Although initialization is an important task for numerical weather
prediction, ocean state estimation usually demands an additional task – a
solution of the time-dependent surface boundary conditions that result from
atmosphere–ocean interaction. Here, we apply the Lagrange multiplier method
to an analogous boundary control problem, tracking the trajectory of the
forced chaotic pendulum. Contrary to previous assertions, it is demonstrated
that the Lagrange multiplier method can track multiple chaotic transitions
through time, so long as the boundary conditions render the system
controllable. Thus, the nonlinear timescale poses no limit to the time
interval for successful Lagrange multiplier-based estimation. That the key
criterion is controllability, not a pure measure of dynamical stability or
chaos, illustrates the similarities between the Lagrange multiplier method
and other state estimation methods. The results with the chaotic pendulum
suggest that nonlinearity should not be a fundamental obstacle to ocean state
estimation with eddy-resolving models, especially when using an improved
first-guess trajectory.</p></abstract-html>
<ref-html id="bib1.bib1"><label>Arbic et al.(2010)Arbic, Wallcraft, and Metzger</label><mixed-citation>
Arbic, B., Wallcraft, A., and Metzger, E.: Concurrent simulation of the
eddying general circulation and tides in a global ocean model, Ocean Model.,
32, 175–187, 2010.
</mixed-citation></ref-html>
<ref-html id="bib1.bib2"><label>Baker and Gollub(1990)</label><mixed-citation>
Baker, G. L. and Gollub, J. P.: Chaotic Dynamics: An Introduction, Cambridge
University Press, Cambridge, 1990.
</mixed-citation></ref-html>
<ref-html id="bib1.bib3"><label>Bennett(1992)</label><mixed-citation>
Bennett, A. F.: Inverse methods in physical oceanography, in: Cambridge Monographs,
Cambridge University Press, Cambridge, 1992.
</mixed-citation></ref-html>
<ref-html id="bib1.bib4"><label>Bennett(2002)</label><mixed-citation>
Bennett, A. F.: Inverse Modeling of the Ocean and Atmosphere, Cambridge
University Press, Cambridge, 2002.
</mixed-citation></ref-html>
<ref-html id="bib1.bib5"><label>Bonekamp et al.(2001)Bonekamp, Van oldenborgh, and Burgers</label><mixed-citation>
Bonekamp, H., Van Oldenborgh, G. J., and Burgers, G.: Variational Assimilation
of Tropical Atmosphere–Ocean and expendable bathythermograph data in the
Hamburg Ocean Primitive Equation ocean general circulation model, adjusting the
surface fluxes in the tropical ocean, J. Geophys. Res.-Oceans, 106, 16693–16709, 2001.
</mixed-citation></ref-html>
<ref-html id="bib1.bib6"><label>Bugnion et al.(2006)Bugnion, Hill, and Stone</label><mixed-citation>
Bugnion, V., Hill, C., and Stone, P. H.: An adjoint analysis of the meridional
overturning circulation in a hybrid coupled model, J. Climate, 19, 3751–3767, 2006.
</mixed-citation></ref-html>
<ref-html id="bib1.bib7"><label>Casti(1985)</label><mixed-citation>
Casti, J. L.: Nonlinear system theory, in: vol. 175, Academic Press, Cambridge, MA, USA, 1985.
</mixed-citation></ref-html>
<ref-html id="bib1.bib8"><label>Cohn and Dee(1988)</label><mixed-citation>
Cohn, S. E. and Dee, D. P.: Observability of discretized partial differential
equations, SIAM J. Numer. Anal., 25, 586–617, 1988.
</mixed-citation></ref-html>
<ref-html id="bib1.bib9"><label>Cong et al.(1998)Cong, Ikeda, and Hendry</label><mixed-citation>
Cong, L. Z., Ikeda, M., and Hendry, R. M.: Variational assimilation of Geosat
altimeter data into a two-layer quasi-geostrophic model over the Newfoundland
ridge and basin, J. Geophys. Res., 103, 7719–7734, 1998.
</mixed-citation></ref-html>
<ref-html id="bib1.bib10"><label>Courtier et al.(1994)Courtier, Thepaut, and Hollingsworth</label><mixed-citation>
Courtier, P., Thepaut, J.-N., and Hollingsworth, A.: A Strategy for Operational
Implementation of 4D-Var, Using an Incremental Approach, Q. J. Roy. Meteorol.
Soc., 120, 1367–1387, 1994.
</mixed-citation></ref-html>
<ref-html id="bib1.bib11"><label>Dahleh and Diaz-Bobillo(1999)</label><mixed-citation>
Dahleh, M. A. and Diaz-Bobillo, I.: Control of Uncertain Systems: A Linear
Programming Approach, Pergamon Press, Oxford, UK, 1999.
</mixed-citation></ref-html>
<ref-html id="bib1.bib12"><label>Delhez et al.(1999)Delhez, Campin, Hirst, and Deleersnijder</label><mixed-citation>
Delhez, E., Campin, J., Hirst, A., and Deleersnijder, E.: Toward a general
theory of the age in ocean modelling, Ocean Model., 1, 17–27, 1999.
</mixed-citation></ref-html>
<ref-html id="bib1.bib13"><label>Evensen(1997)</label><mixed-citation>
Evensen, G.: Advanced data assimilation for strongly nonlinear dynamics,
Mon. Weather Rev., 125, 1342–1354, 1997.
</mixed-citation></ref-html>
<ref-html id="bib1.bib14"><label>Farrell(1989)</label><mixed-citation>
Farrell, B.: Optimal excitation of baroclinic waves, J. Atmos. Sci., 46, 1193–1206, 1989.
</mixed-citation></ref-html>
<ref-html id="bib1.bib15"><label>Farrell and Ioannou(1993)</label><mixed-citation>
Farrell, B. and Ioannou, P.: Transient development of perturbations in stratified
shear-flow, J. Atmos. Sci., 50, 2201–2214, 1993.
</mixed-citation></ref-html>
<ref-html id="bib1.bib16"><label>Ferron and Marotzke(2003)</label><mixed-citation>
Ferron, B. and Marotzke, J.: Impact of 4D-Variational Assimilation of WOCE
Hydrography on the Meridional Circulation of the Indian Ocean, Deep-Sea
Res. Pt. II, 50, 2005–2021, 2003.
</mixed-citation></ref-html>
<ref-html id="bib1.bib17"><label>Fukumori(2002)</label><mixed-citation>
Fukumori, I.: A partitioned Kalman filter and smoother, Mon. Weather Rev., 130, 1370–1383, 2002.
</mixed-citation></ref-html>
<ref-html id="bib1.bib18"><label>Fukumori and Malanotte-Rizzoli(1995)</label><mixed-citation>
Fukumori, I. and Malanotte-Rizzoli, P.: An approximate Kalman Filter for
ocean data assimilation: an example with an idealized Gulf-Stream model, J.
Geophys. Res., 100, 6777–6793, 1995.
</mixed-citation></ref-html>
<ref-html id="bib1.bib19"><label>Fukumori et al.(1993)Fukumori, Benveniste, Wunsch, and Haidvogel</label><mixed-citation>
Fukumori, I., Benveniste, J., Wunsch, C., and Haidvogel, D. B.: Assimilation of
Sea Surface Topography into an Ocean Circulation Model Using a Steady-State
Smoother, J. Phys. Oceanogr., 23, 1831–1855, 1993.
</mixed-citation></ref-html>
<ref-html id="bib1.bib20"><label>Gauthier(1992)</label><mixed-citation>
Gauthier, P.: Chaos and Quadri-dimensional Data Assimilation: A Study Based on
the Lorenz Model, Tellus A, 44, 2–17, 1992.
</mixed-citation></ref-html>
<ref-html id="bib1.bib21"><label>Gebbie(2004)</label><mixed-citation>
Gebbie, G.: Subduction in an Eddy-resolving State Estimate of the Northeast
Atlantic Ocean, PhD thesis, Massachusetts Institute of Technology/Woods Hole
Oceanographic Institution Joint Program in Oceanography, Cambridge/Woods Hole, MA, USA, 2004.
</mixed-citation></ref-html>
<ref-html id="bib1.bib22"><label>Gebbie(2007)</label><mixed-citation>
Gebbie, G.: Does eddy subduction matter in the Northeast Atlantic Ocean?,
J. Geophys. Res., 112, C06007, <a href="https://doi.org/10.1029/2006JC003568" target="_blank">https://doi.org/10.1029/2006JC003568</a>, 2007.
</mixed-citation></ref-html>
<ref-html id="bib1.bib23"><label>Gebbie et al.(2006)Gebbie, Heimbach, and Wunsch</label><mixed-citation>
Gebbie, G., Heimbach, P., and Wunsch, C.: Strategies for Nested and Eddy-Permitting
State Estimation, J. Geophys. Res., 111, C10073, <a href="https://doi.org/10.1029/2005JC003094" target="_blank">https://doi.org/10.1029/2005JC003094</a>, 2006.
</mixed-citation></ref-html>
<ref-html id="bib1.bib24"><label>Gilbert and Lemaréchal(1989)</label><mixed-citation>
Gilbert, J. C. and Lemaréchal, C.: Some Numerical Experiments with
Variable-storage Quasi-Newton Algorithms, Math. Program., 45, 407–435, 1989.
</mixed-citation></ref-html>
<ref-html id="bib1.bib25"><label>Griewank(2000)</label><mixed-citation>
Griewank, A.: Evaluating Derivatives: Principles and Techniques of Algorithmic
Differentiation, Society for Industrial and Applied Mathematics, Philadelphia, 2000.
</mixed-citation></ref-html>
<ref-html id="bib1.bib26"><label>Griffies et al.(2015)Griffies, Winton, Anderson, Benson, Delworth,
Dufour, Dunne, Goddard, Morrison, Rosati et al.</label><mixed-citation>
Griffies, S. M., Winton, M., Anderson, W. G., Benson, R., Delworth, T. L.,
Dufour, C. O., Dunne, J. P., Goddard, P., Morrison, A. K., Rosati, A., Wittenberg,
A. T., Yin, J., and Zhang, R.: Impacts on ocean heat from transient mesoscale
eddies in a hierarchy of climate models, J. Climate, 28, 952–977, 2015.
</mixed-citation></ref-html>
<ref-html id="bib1.bib27"><label>Haine and Hall(2002)</label><mixed-citation>
Haine, T. W. N. and Hall, T. M.: A generalized transport theory: Water-mass
composition and age, J. Phys. Oceanogr., 32, 1932–1946, 2002.
</mixed-citation></ref-html>
<ref-html id="bib1.bib28"><label>Hall et al.(1982)Hall, Cacuci, and Schlesinger</label><mixed-citation>
Hall, M. C. G., Cacuci, D. G., and Schlesinger, M. E.: Sensitivity Analysis of
a Radiative-Convective Model by the Adjoint Method, J. Atmos. Sci., 39, 2038–2050, 1982.
</mixed-citation></ref-html>
<ref-html id="bib1.bib29"><label>Jazwinski(1970)</label><mixed-citation>
Jazwinski, A.: Stochastic Processes and Filtering Theory, Academic Press, Cambridge, MA, 1970.
</mixed-citation></ref-html>
<ref-html id="bib1.bib30"><label>Keppenne et al.(2005)Keppenne, Rienecker, Kurkowski, Adamec et al.</label><mixed-citation>
Keppenne, C. L., Rienecker, M. M., Kurkowski, N. P., and Adamec, D. A.: Ensemble
Kalman filter assimilation of temperature and altimeter data with bias correction
and application to seasonal prediction, Nonlin. Processes Geophys., 12, 491–503,
<a href="https://doi.org/10.5194/npg-12-491-2005" target="_blank">https://doi.org/10.5194/npg-12-491-2005</a>, 2005.
</mixed-citation></ref-html>
<ref-html id="bib1.bib31"><label>Köhl and Stammer(2008)</label><mixed-citation>
Köhl, A. and Stammer, D.: Variability of the meridional overturning in the
North Atlantic from the 50-year GECCO state estimation, J. Phys. Oceanogr.,
38, 1913–1930, 2008.
</mixed-citation></ref-html>
<ref-html id="bib1.bib32"><label>Köhl and Willebrand(2002)</label><mixed-citation>
Köhl, A. and Willebrand, J.: An Adjoint Method for the Assimilation of
Statistical Characteristics into Eddy-resolving Ocean Models, Tellus, 54, 406–425, 2002.
</mixed-citation></ref-html>
<ref-html id="bib1.bib33"><label>Köhl and Willebrand(2003)</label><mixed-citation>
Köhl, A. and Willebrand, J.: Variational Assimilation of SSH Variability
from TOPEX/POSEIDON and ERS1 into an Eddy-permitting Model of the North Atlantic,
J. Geophys. Res., 108, 3092 <a href="https://doi.org/10.1029/2001JC000982" target="_blank">https://doi.org/10.1029/2001JC000982</a>, 2003.
</mixed-citation></ref-html>
<ref-html id="bib1.bib34"><label>Lea et al.(2000)Lea, Allen, and Haine</label><mixed-citation>
Lea, D. J., Allen, M. R., and Haine, T. W. N.: Sensitivity Analysis of the
Climate of a Chaotic System, Tellus A, 52, 523–532, 2000.
</mixed-citation></ref-html>
<ref-html id="bib1.bib35"><label>Lea et al.(2006)Lea, Haine, and Gasparovic</label><mixed-citation>
Lea, D. J., Haine, T. W. N., and Gasparovic, R. F.: Observability of the
Irminger Sea Circulation using Variational Data Assimilation, Q. J. Roy.
Meteorol. Soc., 132, 1545–1576, 2006.
</mixed-citation></ref-html>
<ref-html id="bib1.bib36"><label>LeDimet and Talagrand(1986)</label><mixed-citation>
LeDimet, F. and Talagrand, O.: Variational algorithm for analysis and
assimilation of meteorological observations: Theoretical aspects, Tellus A,
38, 97–110, 1986.
</mixed-citation></ref-html>
<ref-html id="bib1.bib37"><label>Lorenz(1963)</label><mixed-citation>
Lorenz, E. N.: Deterministic, Nonperiodic Flow, J. Atmos. Sci., 20, 130–141, 1963.
</mixed-citation></ref-html>
<ref-html id="bib1.bib38"><label>Malanotte-Rizzoli and Tziperman(1996)</label><mixed-citation>
Malanotte-Rizzoli, P. and Tziperman, E.: The oceanographic data assimilation
problem: overview, motivation and purposes, Elsevier Oceanography Series 61,
Amsterdam, the Netherlands, 3–17, 1996.
</mixed-citation></ref-html>
<ref-html id="bib1.bib39"><label>Maltrud et al.(2010)Maltrud, Bryan, and Peacock</label><mixed-citation>
Maltrud, M., Bryan, F., and Peacock, S.: Boundary impulse response functions in
a century-long eddying global ocean simulation, Environ. Fluid Mech., 10, 275–295, 2010.
</mixed-citation></ref-html>
<ref-html id="bib1.bib40"><label>Marchal(2014)</label><mixed-citation>
Marchal, O.: On the observability of oceanic gyres, J. Phys. Oceanogr., 44, 2498–2523, 2014.
</mixed-citation></ref-html>
<ref-html id="bib1.bib41"><label>Marotzke et al.(1999)Marotzke, Giering, Zhang, Stammer, Hill, and Lee</label><mixed-citation>
Marotzke, J., Giering, R., Zhang, K. Q., Stammer, D., Hill, C., and Lee, T.:
Construction of the Adjoint MIT Ocean General Circulation Model and Application
to Atlantic Heat Transport Sensitivity, J. Geophys. Res., 104, 529–547, 1999.
</mixed-citation></ref-html>
<ref-html id="bib1.bib42"><label>Mazloff et al.(2010)Mazloff, Heimbach, and Wunsch</label><mixed-citation>
Mazloff, M., Heimbach, P., and Wunsch, C.: An eddy-permitting Southern Ocean
state estimate, J. Phys. Oceanogr., 40, 880–899, 2010.
</mixed-citation></ref-html>
<ref-html id="bib1.bib43"><label>McShane(1989)</label><mixed-citation>
McShane, E. J.: The Calculus of Variations from the Beginning Through to
Optimal Control Theory, SIAM J. Control Optim., 27, 916–939, 1989.
</mixed-citation></ref-html>
<ref-html id="bib1.bib44"><label>Menemenlis et al.(2004)Menemenlis, Fukumori, and Lee</label><mixed-citation>
Menemenlis, D., Fukumori, I., and Lee, T.: Using Green's functions to calibrate
an ocean general circulation model, Mon. Weather Rev., 133, 1224–1240, 2004.
</mixed-citation></ref-html>
<ref-html id="bib1.bib45"><label>Miller et al.(1994a)Miller, Ghil, and Gauthiez</label><mixed-citation>
Miller, R. N., Ghil, M., and Gauthiez, F.: Advanced Data Assimilation in
Strongly Nonlinear Systems, J. Atmos. Sci., 51, 1037–1056, 1994a.
</mixed-citation></ref-html>
<ref-html id="bib1.bib46"><label>Miller et al.(1994b)Miller, Zaron, and Bennett</label><mixed-citation>
Miller, R. N., Zaron, E. D., and Bennett, A. F.: Data assimilation in models
with convective adjustment, Mon. Weath. Rev., 122, 2607–2613, 1994b.
</mixed-citation></ref-html>
<ref-html id="bib1.bib47"><label>Nocedal(1980)</label><mixed-citation>
Nocedal, J.: Updating quasi-Newton matrices with limited storage, Math. Comput.,
35, 773–782, 1980.
</mixed-citation></ref-html>
<ref-html id="bib1.bib48"><label>Ott et al.(1990)Ott, Grebogi, and Yorke</label><mixed-citation>
Ott, E., Grebogi, C., and Yorke, J.: Controlling chaos, Phys. Rev. Lett., 64, 1196–1199, 1990.
</mixed-citation></ref-html>
<ref-html id="bib1.bib49"><label>Palmer(1996)</label><mixed-citation>
Palmer, T. N.: Predictability of the Atmosphere and Oceans: From Days to
Decades, in: vol. 44 of NATO ASI Series, chap. 3, Decadal Climate Variability:
Dynamics and Predictability, Springer, New York, NY, USA, 83–156, 1996.

</mixed-citation></ref-html>
<ref-html id="bib1.bib50"><label>Pires et al.(1996)Pires, Vautard, and Talagrand</label><mixed-citation>
Pires, C., Vautard, R., and Talagrand, O.: On extending the limits of variational
assimilation in nonlinear chaotic systems, Tellus A, 48, 96–121, 1996.
</mixed-citation></ref-html>
<ref-html id="bib1.bib51"><label>Schröter et al.(1993)Schröter, Seiler, and Wenzel</label><mixed-citation>
Schröter, J., Seiler, U., and Wenzel, M.: Variational Assimilation of
GEOSAT Data into an Eddy-resolving Model of the Gulf Stream Extension Area, J.
Phys. Oceanogr., 23, 925–953, 1993.
</mixed-citation></ref-html>
<ref-html id="bib1.bib52"><label>Stammer and Wunsch(1996)</label><mixed-citation>
Stammer, D. and Wunsch, C.: The determination of the large-scale circulation of
the Pacific Ocean from satellite altimetry using model Green's functions, J.
Geophys. Res., 101, 18409–18432, 1996.
</mixed-citation></ref-html>
<ref-html id="bib1.bib53"><label>Stammer et al.(2002a)Stammer, Wunsch, Fukumori, and Marshall</label><mixed-citation>
Stammer, D., Wunsch, C., Fukumori, I., and Marshall, J.: State Estimation in
Modern Oceanographic Research, EOS, 83, 294–295, 2002a.
</mixed-citation></ref-html>
<ref-html id="bib1.bib54"><label>Stammer et al.(2002b)Stammer, Wunsch, Giering, Eckert,
Heimbach, Marotzke, Adcroft, Hill, and Marshall</label><mixed-citation>
Stammer, D., Wunsch, C., Giering, R., Eckert, C., Heimbach, P., Marotzke, J.,
Adcroft, A., Hill, C. N., and Marshall, J.: The global ocean circulation
during 1992–1997, estimated from ocean observations and a general circulation
model, J. Geophys. Res.-Oceans, 107, 3118 <a href="https://doi.org/10.1029/2001JC000888" target="_blank">https://doi.org/10.1029/2001JC000888</a>, 2002b.
</mixed-citation></ref-html>
<ref-html id="bib1.bib55"><label>Stammer et al.(2004)Stammer, Ueyoshi, Kohl, Large, Josey, and
Wunsch</label><mixed-citation>
Stammer, D., Ueyoshi, K., Kohl, A., Large, W. G., Josey, S. A., and Wunsch, C.:
Estimating air-sea fluxes of heat, freshwater, and momentum through global
ocean data assimilation, J. Geophys. Res., 109, C05023, <a href="https://doi.org/10.1029/2003JC002082" target="_blank">https://doi.org/10.1029/2003JC002082</a>, 2004.
</mixed-citation></ref-html>
<ref-html id="bib1.bib56"><label>Tanguay et al.(1995)Tanguay, Bartello, and Gauthier</label><mixed-citation>
Tanguay, M., Bartello, P., and Gauthier, P.: Four-dimensional Data Assimilation
with a Wide Range of Scales, Tellus A, 47, 974–997, 1995.
</mixed-citation></ref-html>
<ref-html id="bib1.bib57"><label>Tarantola and Valette(1982)</label><mixed-citation>
Tarantola, A. and Valette, B.: Generalized Nonlinear Inverse Problems Solved
using the Least Squares Criterion, Rev. Geophys. Space Phys., 20, 219–232, 1982.
</mixed-citation></ref-html>
<ref-html id="bib1.bib58"><label>Thacker and Long(1988)</label><mixed-citation>
Thacker, W. C. and Long, R. B.: Fitting dynamics to data, J. Geophys. Res.,
93, 1227–1240, 1988.
</mixed-citation></ref-html>
<ref-html id="bib1.bib59"><label>Tziperman and Thacker(1989)</label><mixed-citation>
Tziperman, E. and Thacker, W. C.: An Optimal-Control/Adjoint-Equations Approach
to Studying the Oceanic General Circulation, J. Phys. Oceanogr., 19, 1471–1485, 1989.
</mixed-citation></ref-html>
<ref-html id="bib1.bib60"><label>Tziperman et al.(1994)Tziperman, Stone, Cane, and Jarosh</label><mixed-citation>
Tziperman, E., Stone, L., Cane, M. A., and Jarosh, H.: El Niño chaos:
overlapping of resonances between the seasonal cycle and the Pacific
ocean–atmosphere oscillator, Science, 264, 72–74, 1994.
</mixed-citation></ref-html>
<ref-html id="bib1.bib61"><label>Wunsch(1996)</label><mixed-citation>
Wunsch, C.: The Ocean Circulation Inverse Problem, Cambridge University Press, New York, 1996.
</mixed-citation></ref-html>
<ref-html id="bib1.bib62"><label>Wunsch(2010)</label><mixed-citation>
Wunsch, C.: Discrete Inverse and State Estimation Problems. With Geophysical
Fluid Applications, Cambridge University Press, Cambridge, UK, 2010.
</mixed-citation></ref-html>
<ref-html id="bib1.bib63"><label>Wunsch and Heimbach(2007)</label><mixed-citation>
Wunsch, C. and Heimbach, P.: Practical global oceanic state estimation, Physica D,
230, 192–208, 2007.
</mixed-citation></ref-html>
<ref-html id="bib1.bib64"><label>Wunsch et al.(2009)Wunsch, Heimbach, Ponte, and Fukumori</label><mixed-citation>
Wunsch, C., Heimbach, P., Ponte, R., and Fukumori, I.: The global general
circulation of the oceans estimated by the ECCO-Consortium, Oceanography,
20, 88–103, 2009.
</mixed-citation></ref-html>--></article>
