<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing with OASIS Tables v3.0 20080202//EN" "https://jats.nlm.nih.gov/nlm-dtd/publishing/3.0/journalpub-oasis3.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:oasis="http://docs.oasis-open.org/ns/oasis-exchange/table" xml:lang="en" dtd-version="3.0" article-type="research-article">
  <front>
    <journal-meta><journal-id journal-id-type="publisher">NPG</journal-id><journal-title-group>
    <journal-title>Nonlinear Processes in Geophysics</journal-title>
    <abbrev-journal-title abbrev-type="publisher">NPG</abbrev-journal-title><abbrev-journal-title abbrev-type="nlm-ta">Nonlin. Processes Geophys.</abbrev-journal-title>
  </journal-title-group><issn pub-type="epub">1607-7946</issn><publisher>
    <publisher-name>Copernicus Publications</publisher-name>
    <publisher-loc>Göttingen, Germany</publisher-loc>
  </publisher></journal-meta>
    <article-meta>
      <article-id pub-id-type="doi">10.5194/npg-33-425-2026</article-id><title-group><article-title>Bayesian data selection to quantify the value of  data for landslide runout calibration</article-title><alt-title>Bayesian data selection for landslide runout calibration</alt-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author" corresp="yes" rid="aff1">
          <name><surname>Kumar</surname><given-names>V. Mithlesh</given-names></name>
          <email>kumar@mbd.rwth-aachen.de</email>
        <ext-link>https://orcid.org/0009-0000-9790-664X</ext-link></contrib>
        <contrib contrib-type="author" corresp="no" rid="aff1">
          <name><surname>Yildiz</surname><given-names>Anil</given-names></name>
          
        <ext-link>https://orcid.org/0000-0002-2257-7025</ext-link></contrib>
        <contrib contrib-type="author" corresp="yes" rid="aff1">
          <name><surname>Kowalski</surname><given-names>Julia</given-names></name>
          <email>kowalski@mbd.rwth-aachen.de</email>
        <ext-link>https://orcid.org/0000-0003-4123-5896</ext-link></contrib>
        <aff id="aff1"><label>1</label><institution>Chair of Methods for Model-based Development in Computational Engineering,  RWTH Aachen University, Aachen, 52062, Germany</institution>
        </aff>
      </contrib-group>
      <author-notes><corresp id="corr1">V. Mithlesh Kumar (kumar@mbd.rwth-aachen.de) and Julia Kowalski (kowalski@mbd.rwth-aachen.de)</corresp></author-notes><pub-date><day>25</day><month>August</month><year>2026</year></pub-date>
      
      <volume>33</volume>
      <issue>3</issue>
      <fpage>425</fpage><lpage>453</lpage>
      <history>
        <date date-type="received"><day>15</day><month>September</month><year>2025</year></date>
           <date date-type="rev-request"><day>2</day><month>October</month><year>2025</year></date>
           <date date-type="rev-recd"><day>3</day><month>February</month><year>2026</year></date>
           <date date-type="accepted"><day>16</day><month>June</month><year>2026</year></date>
      </history>
      <permissions>
        <copyright-statement>Copyright: © 2026 V. Mithlesh Kumar et al.</copyright-statement>
        <copyright-year>2026</copyright-year>
      <license license-type="open-access"><license-p>This work is licensed under the Creative Commons Attribution 4.0 International License. To view a copy of this licence, visit <ext-link ext-link-type="uri" xlink:href="https://creativecommons.org/licenses/by/4.0/">https://creativecommons.org/licenses/by/4.0/</ext-link></license-p></license></permissions><self-uri xlink:href="https://npg.copernicus.org/articles/33/425/2026/npg-33-425-2026.html">This article is available from https://npg.copernicus.org/articles/33/425/2026/npg-33-425-2026.html</self-uri><self-uri xlink:href="https://npg.copernicus.org/articles/33/425/2026/npg-33-425-2026.pdf">The full text article is available as a PDF file from https://npg.copernicus.org/articles/33/425/2026/npg-33-425-2026.pdf</self-uri>
      <abstract><title>Abstract</title>

      <p id="d2e102">The reliability of physics-based landslide runout models depends on the effective calibration of their parameters, which are often conceptual and cannot be physically measured. Bayesian methods offer a robust framework to incorporate uncertainties in both model and observations into the calibration process. Therefore, they are increasingly used to calibrate physics-based landslide runout models. However, the reliability of Bayesian calibration in its practical application to real-world landslide events depends on the availability and quality of observational data – from aggregated post-event measurements such as impact area to time-resolved data such as force time histories. Despite this, systematic investigation of the influence of observational data on the Bayesian calibration of landslide runout models has been limited.</p>

      <p id="d2e105">We propose quantifying the impact of observational data on calibration outcomes by measuring the information gained during the calibration process using an information-theoretic measure called Kullback-Leibler (KL) divergence. Building on this, we present a unified Bayesian data selection workflow to identify the most informative dataset for calibrating a given parameter. The workflow runs parallel calibration routines across available observation datasets. It then computes the information gained relative to the observations by calculating the KL divergence between prior and posterior distributions and selects the dataset that yields the highest KL divergence.</p>

      <p id="d2e108">We demonstrate our workflow using an elementary landslide runout model, calibrating friction parameters with a diverse set of synthetic observations to evaluate the impact of data selection on parameter calibration. Specifically, we compare and quantify the information gained from calibration routines using observations that differ in information content (runout distance versus maximum velocity), granularity (aggregated versus time series data), and temporal characteristics (resolution and length of the velocity time series). The insights from this study will optimize the use of available observations for calibration and guide the design of effective data acquisition strategies.</p>
  </abstract>
    
<funding-group>
<award-group id="gs1">
<funding-source>Deutsche Forschungsgemeinschaft</funding-source>
<award-id>333849990/GRK2379</award-id>
<award-id>441527981</award-id>
</award-group>
</funding-group>
</article-meta>
  </front>
<body>
      

<sec id="Ch1.S1" sec-type="intro">
  <label>1</label><title>Introduction</title>
      <p id="d2e120">Landslides are a significant natural hazard, with their frequency and intensity increasing as climate change makes extreme weather patterns more likely <xref ref-type="bibr" rid="bib1.bibx42 bib1.bibx41 bib1.bibx49" id="paren.1"/>. Understanding and predicting landslide runout behavior – how far and fast landslides travel <xref ref-type="bibr" rid="bib1.bibx53" id="paren.2"/> – is therefore crucial for impact-based risk and hazard assessment <xref ref-type="bibr" rid="bib1.bibx50 bib1.bibx14" id="paren.3"/> and for the development of effective mitigation strategies <xref ref-type="bibr" rid="bib1.bibx34 bib1.bibx25" id="paren.4"/>. To this end, we use physics-based runout models, which are adept at capturing the bulk behavior of landslides, an essential objective in runout forecasting <xref ref-type="bibr" rid="bib1.bibx35" id="paren.5"/>. A wide variety of physics-based computational runout models are available, and they can be broadly classified on the basis of the fidelity level and flow characteristics accounted for, the underlying rheological relationships, and the computational methods used to solve them. <xref ref-type="bibr" rid="bib1.bibx35" id="text.6"/> and <xref ref-type="bibr" rid="bib1.bibx45" id="text.7"/> compiled a collection of selected computational models. Most of these landslide runout models are semi-empirical due to the absence of universal constitutive laws governing the complex and diverse phenomena involved in landslides <xref ref-type="bibr" rid="bib1.bibx40" id="paren.8"/>. Consequently, these models rely on empirical constitutive relations rather than mechanistic subscale formulations representing microscale behavior. This implies the presence of conceptual parameters <xref ref-type="bibr" rid="bib1.bibx26" id="paren.9"/> that cannot be physically measured and thus must be calibrated based on the re-analyses of past landslide events. By calibration, we refer to the process of inferring model parameters from observational data. The applicability of such calibrated computational models is therefore limited to the physical regime covered by the available calibration data and offers little insight into the complex micromechanical properties of real landslides <xref ref-type="bibr" rid="bib1.bibx35" id="paren.10"/>. However, their ability to reproduce the bulk behavior of landslides makes them a pragmatic choice to analyze the landslide runout behavior <xref ref-type="bibr" rid="bib1.bibx23" id="paren.11"/> and allows their use as operational tools in hazard mitigation.</p>
      <p id="d2e157">Computational calibration methods play an essential role in the model-based prediction tool chain, particularly in landslide runout modeling. In the past, landslide runout models were predominantly calibrated using deterministic methods, traditionally based on subjective trial and error choices <xref ref-type="bibr" rid="bib1.bibx24" id="paren.12"/>. These deterministic methods are limited by equifinality and non-uniqueness issues <xref ref-type="bibr" rid="bib1.bibx36" id="paren.13"/> and do not offer a robust framework to handle multiple uncertainties, for instance due to model and data errors, involved in the calibration process <xref ref-type="bibr" rid="bib1.bibx5" id="paren.14"/>. In contrast, probabilistic methods effectively address these limitations by aiming for the probability distributions of the parameters rather than a deterministic estimate. Bayesian methods represent initial parameter uncertainty with a prior distribution, which is updated based on observational data to yield a posterior distribution reflecting the reduced uncertainty. This process thus offers a comprehensive framework to explicitly handle the various uncertainties involved in calibrating computational models that include conceptual parameters.</p>
      <p id="d2e169">Recently, Bayesian methods have attracted significant interest from the geohazard research community. For example, <xref ref-type="bibr" rid="bib1.bibx12" id="text.15"/> and <xref ref-type="bibr" rid="bib1.bibx20" id="text.16"/> calibrated the rheological parameters of a snow avalanche propagation model employing a Bayesian approach. <xref ref-type="bibr" rid="bib1.bibx2" id="text.17"/> applied this approach to calibrate a landslide runout model and compared its performance with a deterministic parameter identification algorithm. Furthermore, <xref ref-type="bibr" rid="bib1.bibx37" id="text.18"/> inverted landslide characteristics using seismic data within a Bayesian framework. One significant finding in these studies that is also known from other application fields is that applying Bayesian methods typically entails a significant computational burden because of the large number of so-called forward computational model evaluations required. This high computational cost is the primary limitation of Bayesian methods in their practical applications <xref ref-type="bibr" rid="bib1.bibx1 bib1.bibx6" id="paren.19"/>. A recent trend in simulation for computationally costly high-throughput tasks is, therefore, to train non-intrusive surrogate models that aim at substituting the original simulation model with a fast-to-evaluate alternative that, for example, can be utilized for uncertainty quantification <xref ref-type="bibr" rid="bib1.bibx55" id="paren.20"/>. <xref ref-type="bibr" rid="bib1.bibx58" id="text.21"/> leveraged surrogates based on Gaussian Process (GP) emulation to develop a computationally feasible Bayesian calibration workflow, while <xref ref-type="bibr" rid="bib1.bibx38" id="text.22"/> adopted a similar approach using the polynomial chaos expansion ansatz to build surrogates.</p>
      <p id="d2e197">Although integrating fast-evaluating surrogates into the calibration workflow alleviates the computational burden associated with Bayesian calibration methods, this addresses only one part of the challenge. Effective calibration essentially relies on reducing the discrepancy between model predictions and observational data <xref ref-type="bibr" rid="bib1.bibx8 bib1.bibx3" id="paren.23"/>. Two important factors contribute to this discrepancy: (i) the potential inadequacy of the model to capture the real-world process to be predicted, and (ii) the innate measurement uncertainties within the observational data. Model inadequacy and its impact on the efficacy of the calibration process have been addressed, for example, in the work of <xref ref-type="bibr" rid="bib1.bibx28" id="text.24"/>, <xref ref-type="bibr" rid="bib1.bibx19" id="text.25"/>, and <xref ref-type="bibr" rid="bib1.bibx52" id="text.26"/>.</p>
      <p id="d2e213">The measurement uncertainty in the observational data affects the uncertainty of the predicted landslide runout when those data are integrated into the computational calibration method underlying the model-based prediction tool chain. Methods for handling measurement-related uncertainty are well established. These typically involve incorporating a statistical noise model, such as a Gaussian distribution, into the calibration process. Hyperparameters such as mean and standard deviation are assumed heuristically or inferred along with the model parameters <xref ref-type="bibr" rid="bib1.bibx58 bib1.bibx20" id="paren.27"/> from the data. Several studies use this approach to assess the influence of uncertainty in the observational data on the calibration outcomes <xref ref-type="bibr" rid="bib1.bibx2" id="paren.28"/>. A typical outcome is, hence, insight into the acceptable measurement uncertainty to guarantee a specific quality of the calibration result, which in turn dictates the reliability of the computational model predictions.</p>
      <p id="d2e222">However, it is seldom considered that different types of observational data, such as an outline of the area affected by a landslide versus a localized measurement of its deposition height, result in markedly different calibration results. This observation holds even if identical Gaussian noise assumptions are being used. Such differences indicate that the choice of observational data not only plays a critical role in shaping calibration outcomes but also constitutes a lever to improve the quality of calibration outcomes. This insight is echoed in the work of <xref ref-type="bibr" rid="bib1.bibx58" id="text.29"/>, who found that remarkably localized spatial data, such as maximum velocities or deposit heights, provided better constraints for friction coefficients than aggregated data like deposit volume and impact area. <xref ref-type="bibr" rid="bib1.bibx37" id="text.30"/> made similar observations and found that force time history data was more adept at inverting the characteristics of the landslide than static data such as deposit area or runout distance. This is a surprising result since it strongly seems to indicate that even if we are ultimately interested in predicting the impact area, physics-based computational landslide models should not necessarily be calibrated only based on the impact area of past results, but should also take into consideration local information such as a measurement of the deposition height at a specific location.</p>
      <p id="d2e231">Any systematic investigation of this effect requires an approach that quantifies the value of concrete choices of observational data and lays the foundation for systematically assessing the value-add of specific choices of observational data on the calibration outcome and, eventually, the predictive quality of the computational prediction pipeline. Such methods are currently unavailable, yet would be highly relevant to the geohazard community since observational data are often sparse due to logistical and financial constraints. Gaining clearer insight into how different observations influence calibration outcomes could support more efficient use of available data and guide the design of smarter data acquisition strategies. Similar ideas have been followed in other fields. <xref ref-type="bibr" rid="bib1.bibx27" id="text.31"/> examined the impact of data resolution on the inference of hydrological model parameters using data from an experimental basin, while <xref ref-type="bibr" rid="bib1.bibx33" id="text.32"/> and <xref ref-type="bibr" rid="bib1.bibx9" id="text.33"/> investigated the effect of dataset length on the calibration of the hydrological model in data-limited catchments. In building energy modeling, <xref ref-type="bibr" rid="bib1.bibx19" id="text.34"/> studied the role of data quantity and quality in Bayesian calibration of the EnergyPlus model <xref ref-type="bibr" rid="bib1.bibx46" id="paren.35"/>. However, even these studies qualitatively compare the calibration outcomes and do not attempt to quantify the impact, which would help us to optimize data acquisition.</p>
      <p id="d2e249">We can assess the impact of observations on the Bayesian calibration outcome by examining the resulting posterior distributions since they represent the uncertainty reduced during calibration. While <xref ref-type="bibr" rid="bib1.bibx58" id="text.36"/> reported variation in calibration outcomes across observational datasets by qualitatively comparing the corresponding posterior distributions, <xref ref-type="bibr" rid="bib1.bibx37" id="text.37"/> analyzed this variation using the modes of the posterior distributions. However, neither of these approaches captures the inherent pathway by which observations influence Bayesian calibration: updating prior beliefs with information to obtain the posterior distribution. We therefore measure this information gained during calibration to quantify the impact of observational data on calibration outcomes. To this end, we employ the Kullback-Leibler (KL) divergence, an information-theoretic concept, to compare probability distributions. Specifically, it measures the information lost when approximating a probability distribution relative to the true distribution. Thus, we quantify the information gained during calibration by measuring the KL divergence between the posterior and prior distributions. In contrast to earlier studies that relied on qualitative and mode-based comparisons, our approach provides a quantitative metric for each observation's impact on the calibration outcome, which can be used to select the most informative dataset. However, computing the KL divergence poses a computational challenge because it requires calculating integrals that include the intractable posterior distributions. A widely adopted approach to tackle this challenge involves Monte Carlo-based estimators, which are unbiased but suffer from high variance due to random sampling. Consequently, we choose a novel universal divergence estimator proposed by <xref ref-type="bibr" rid="bib1.bibx48" id="text.38"/>. This estimator offers robust estimates of KL divergence leveraging <inline-formula><mml:math id="M1" display="inline"><mml:mi>k</mml:mi></mml:math></inline-formula>-nearest-neighbor (<inline-formula><mml:math id="M2" display="inline"><mml:mi>k</mml:mi></mml:math></inline-formula>-NN) distances.</p>
      <p id="d2e275">This study addresses the identified research gap: a lack of a method to select the observational data best suited for the Bayesian calibration of landslide runout models by (a) proposing a methodological approach and (b) introducing a computational framework that demonstrates its feasibility. The developed Bayesian data selection workflow, which we define as finding the most informative dataset to calibrate a given parameter, allows us to quantify the information gained during calibration. This workflow orchestrates multiple calibration routines across different observational datasets in parallel. It then quantifies the calibration performance using KL divergence, which can be systematically exploited to optimize the data acquisition. To our knowledge, it is the first time that such information-theoretic concepts have been incorporated to compare and quantify the calibration performance of landslide runout models.</p>
      <p id="d2e278">We will demonstrate the proficiency of our Bayesian data selection methodology based on the so-called lumped mass model representing an idealized landslide runout model. The primary objective of this work is to introduce and describe a novel methodology, namely a Bayesian data selection workflow for landslide runout models. Application to the lumped mass model will allow us to use this workflow to investigate the role of the type and scope of observational data in the calibration process. In order to further isolate the role of data selection from secondary effects, we will limit ourselves to synthetic data generated from simulating an idealized lumped mass model at preselected set of parameters. While we know the limitations of a lumped mass point model, it provides an ideal testbed to assess the value-add of optimized data selection, which constitutes an unused potential hidden in landslide runout prediction. The community can also use our results as a future reproducible benchmark case.</p>
</sec>
<sec id="Ch1.S2">
  <label>2</label><title>Methodology</title>
      <p id="d2e289">This section outlines a novel approach to quantifying the information gain in an automated Bayesian data selection workflow. As a first step, it is necessary to define the statistical model that underpins the embedded Bayesian calibration task, before introducing the KL divergence as a metric for measuring calibration performance across multiple datasets. Finally, we detail the computational workflow used to apply this methodology.</p>
<sec id="Ch1.S2.SS1">
  <label>2.1</label><title>Statistical model formulation considering measurement and model error</title>
      <p id="d2e299">The computational model, referred to as <inline-formula><mml:math id="M3" display="inline"><mml:mi mathvariant="script">M</mml:mi></mml:math></inline-formula>, comprises a physics-based theoretical model and a solution algorithm. Both in conjunction allow predicting the relevant aspects of the landslide hazard mitigation task, for example, the length of the runout. The computational landslide model can hence be written as a <italic>parameter-to-observable mapping</italic> given by

            <disp-formula id="Ch1.E1" content-type="numbered"><label>1</label><mml:math id="M4" display="block"><mml:mrow><mml:mi mathvariant="script">M</mml:mi><mml:mo>:</mml:mo><mml:mi>S</mml:mi><mml:mo>→</mml:mo><mml:mi>D</mml:mi></mml:mrow></mml:math></disp-formula>

          Here, <inline-formula><mml:math id="M5" display="inline"><mml:mrow><mml:mi>s</mml:mi><mml:mo>∈</mml:mo><mml:mi>S</mml:mi></mml:mrow></mml:math></inline-formula> contains all the parametrized information needed to initialize a specific landslide simulation scenario, for example, topographic information and initial mass distribution, while <inline-formula><mml:math id="M6" display="inline"><mml:mrow><mml:mi>d</mml:mi><mml:mo>∈</mml:mo><mml:mi>D</mml:mi></mml:mrow></mml:math></inline-formula> denotes the concrete prediction that is being made, such as the runout length or the impact area. Typically, space <inline-formula><mml:math id="M7" display="inline"><mml:mi>D</mml:mi></mml:math></inline-formula> also comprises the space of observables <inline-formula><mml:math id="M8" display="inline"><mml:mi>y</mml:mi></mml:math></inline-formula> that can be measured in the field, such that we assume <inline-formula><mml:math id="M9" display="inline"><mml:mrow><mml:mi>y</mml:mi><mml:mo>∈</mml:mo><mml:mi>D</mml:mi></mml:mrow></mml:math></inline-formula>. In our case, <inline-formula><mml:math id="M10" display="inline"><mml:mi mathvariant="script">M</mml:mi></mml:math></inline-formula> additionally depends on parameter <inline-formula><mml:math id="M11" display="inline"><mml:mrow><mml:mi mathvariant="italic">θ</mml:mi><mml:mo>∈</mml:mo><mml:mi mathvariant="normal">Θ</mml:mi></mml:mrow></mml:math></inline-formula> that cannot be determined independently and thus must be calibrated based on field observations <inline-formula><mml:math id="M12" display="inline"><mml:mi>y</mml:mi></mml:math></inline-formula>. Note that <inline-formula><mml:math id="M13" display="inline"><mml:mi mathvariant="italic">θ</mml:mi></mml:math></inline-formula> can either represent a scalar parameter or refer to a tuple of parameters, such as a single or several friction parameters.</p>
      <p id="d2e413">The primary objective of model calibration is to leverage observations <inline-formula><mml:math id="M14" display="inline"><mml:mi>y</mml:mi></mml:math></inline-formula> in order to infer on optimal parameters <inline-formula><mml:math id="M15" display="inline"><mml:mi mathvariant="italic">θ</mml:mi></mml:math></inline-formula>, such that for a given scenario <inline-formula><mml:math id="M16" display="inline"><mml:mi>s</mml:mi></mml:math></inline-formula> the discrepancy between observations <inline-formula><mml:math id="M17" display="inline"><mml:mi>y</mml:mi></mml:math></inline-formula> and model predictions <inline-formula><mml:math id="M18" display="inline"><mml:mrow><mml:mi mathvariant="script">M</mml:mi><mml:mo>(</mml:mo><mml:mi>s</mml:mi><mml:mo>;</mml:mo><mml:mi mathvariant="italic">θ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> is minimized:

            <disp-formula id="Ch1.E2" content-type="numbered"><label>2</label><mml:math id="M19" display="block"><mml:mtable rowspacing="0.2ex" class="split" displaystyle="true" columnalign="right left"><mml:mtr><mml:mtd/><mml:mtd><mml:mrow><mml:mtext>Calibration task:</mml:mtext></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd/><mml:mtd><mml:mrow><mml:mspace linebreak="nobreak" width="1em"/><mml:mtext>Find </mml:mtext><mml:mi mathvariant="italic">θ</mml:mi><mml:mtext> such that </mml:mtext><mml:mo>|</mml:mo><mml:mi mathvariant="script">M</mml:mi><mml:mo>(</mml:mo><mml:mi>s</mml:mi><mml:mo>;</mml:mo><mml:mi mathvariant="italic">θ</mml:mi><mml:mo>)</mml:mo><mml:mo>-</mml:mo><mml:mi>y</mml:mi><mml:mo>|</mml:mo><mml:mtext> is minimal</mml:mtext></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>

          Note, that we did not yet specify the metric, in which this deviation is measured. The discrepancy can be attributed to two significant sources of uncertainties, as discussed in the seminal work of <xref ref-type="bibr" rid="bib1.bibx28" id="text.39"/>. First, we have the uncertainty resulting from the noise in the observations, referred to as measurement noise, denoted by <inline-formula><mml:math id="M20" display="inline"><mml:mi mathvariant="italic">ε</mml:mi></mml:math></inline-formula>. Second, we have the model inadequacy, <inline-formula><mml:math id="M21" display="inline"><mml:mrow><mml:mi mathvariant="italic">δ</mml:mi><mml:mo>(</mml:mo><mml:mi>s</mml:mi><mml:mo>;</mml:mo><mml:mi mathvariant="italic">θ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>, which results from uncertainty and error in the model formulation itself.</p>
<sec id="Ch1.S2.SS1.SSS1">
  <label>2.1.1</label><title>Measurement noise</title>
      <p id="d2e544">Field observations and measurements are inevitably affected by errors and noise <xref ref-type="bibr" rid="bib1.bibx28 bib1.bibx39" id="paren.40"/>, which arise due to inherent limitations in measurement processes. For instance, measuring the runout distance of a landslide is often plagued with uncertainties resulting from sensor accuracy and topography resolution. Let us assume that our measurement <inline-formula><mml:math id="M22" display="inline"><mml:mi>y</mml:mi></mml:math></inline-formula> is subject to noise, such that we have to differentiate it from <inline-formula><mml:math id="M23" display="inline"><mml:mrow><mml:msup><mml:mi>y</mml:mi><mml:mi>r</mml:mi></mml:msup></mml:mrow></mml:math></inline-formula>, denoting the true, yet unknown, representation of the physical phenomenon we want to capture. There will always be a discrepancy between <inline-formula><mml:math id="M24" display="inline"><mml:mi>y</mml:mi></mml:math></inline-formula> and <inline-formula><mml:math id="M25" display="inline"><mml:mrow><mml:msup><mml:mi>y</mml:mi><mml:mi>r</mml:mi></mml:msup></mml:mrow></mml:math></inline-formula>, referred to as the measurement uncertainty and denoted by <inline-formula><mml:math id="M26" display="inline"><mml:mi mathvariant="italic">ε</mml:mi></mml:math></inline-formula>:

                  <disp-formula id="Ch1.E3" content-type="numbered"><label>3</label><mml:math id="M27" display="block"><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><mml:mi mathvariant="italic">ε</mml:mi><mml:mo>:</mml:mo><mml:mo>=</mml:mo><mml:mi>y</mml:mi><mml:mo>-</mml:mo><mml:msup><mml:mi>y</mml:mi><mml:mi>r</mml:mi></mml:msup></mml:mrow></mml:math></disp-formula></p>
      <p id="d2e615">Observations can hence be expressed as a function of <inline-formula><mml:math id="M28" display="inline"><mml:mrow><mml:msup><mml:mi>y</mml:mi><mml:mi>r</mml:mi></mml:msup></mml:mrow></mml:math></inline-formula> and the associated measurement noise, according to

              <disp-formula id="Ch1.E4" content-type="numbered"><label>4</label><mml:math id="M29" display="block"><mml:mrow><mml:mi>y</mml:mi><mml:mo>=</mml:mo><mml:msup><mml:mi>y</mml:mi><mml:mi>r</mml:mi></mml:msup><mml:mo>+</mml:mo><mml:mi mathvariant="italic">ε</mml:mi><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula></p>
      <p id="d2e649">Typically, the noise is not known, such that we need to assume an ansatz, which constitutes our noise model. The additive Gaussian noise model is one of the most common noise models, where <inline-formula><mml:math id="M30" display="inline"><mml:mi mathvariant="italic">ε</mml:mi></mml:math></inline-formula> is considered to be a realization of a Gaussian distribution, often with zero mean and a covariance matrix <inline-formula><mml:math id="M31" display="inline"><mml:mi mathvariant="bold">Σ</mml:mi></mml:math></inline-formula>, i.e. <inline-formula><mml:math id="M32" display="inline"><mml:mrow><mml:mi mathvariant="italic">ε</mml:mi><mml:mo>∼</mml:mo><mml:mi mathvariant="script">N</mml:mi><mml:mo>(</mml:mo><mml:mn mathvariant="normal">0</mml:mn><mml:mo>,</mml:mo><mml:mi mathvariant="bold">Σ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>. The structure of the covariance matrix depends on the observational dataset we are calibrating. For scalar observations, the covariance matrix reduces to the variance <inline-formula><mml:math id="M33" display="inline"><mml:mrow><mml:msup><mml:mi mathvariant="italic">σ</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula>. For multidimensional observations, we assume independent errors with constant standard deviation, leading to <inline-formula><mml:math id="M34" display="inline"><mml:mrow><mml:mi mathvariant="bold">Σ</mml:mi><mml:mo>=</mml:mo><mml:msup><mml:mi mathvariant="italic">σ</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mi mathvariant="bold">I</mml:mi></mml:mrow></mml:math></inline-formula>. We refer to <inline-formula><mml:math id="M35" display="inline"><mml:mi mathvariant="italic">σ</mml:mi></mml:math></inline-formula> as the discrepancy parameter and in this study, it is either fixed based on heuristic assumptions or calibrated alongside the model parameters.</p>
</sec>
<sec id="Ch1.S2.SS1.SSS2">
  <label>2.1.2</label><title>Model Inadequacy</title>
      <p id="d2e732">Model inadequacy denotes the inherent limitations of the computational model in replicating the true value <inline-formula><mml:math id="M36" display="inline"><mml:mrow><mml:msup><mml:mi>y</mml:mi><mml:mi>r</mml:mi></mml:msup></mml:mrow></mml:math></inline-formula>, due to the idealizing assumptions and theoretical simplifications that underlie the physics-based model formulation. A certain idealization is evident in the case of landslide modeling, where the process complexity cannot be fully resolved.</p>
      <p id="d2e746"><xref ref-type="bibr" rid="bib1.bibx28" id="text.41"/> addressed this discrepancy by explicitly incorporating a model inadequacy term into the statistical model. Following <xref ref-type="bibr" rid="bib1.bibx39" id="text.42"/>, we refer to the model inadequacy term as <inline-formula><mml:math id="M37" display="inline"><mml:mrow><mml:mi mathvariant="italic">δ</mml:mi><mml:mo>(</mml:mo><mml:mi>s</mml:mi><mml:mo>;</mml:mo><mml:mi mathvariant="italic">θ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> and we have the relation

              <disp-formula id="Ch1.E5" content-type="numbered"><label>5</label><mml:math id="M38" display="block"><mml:mrow><mml:msup><mml:mi>y</mml:mi><mml:mi>r</mml:mi></mml:msup><mml:mo>-</mml:mo><mml:mi mathvariant="script">M</mml:mi><mml:mo>(</mml:mo><mml:mi>s</mml:mi><mml:mo>;</mml:mo><mml:mi mathvariant="italic">θ</mml:mi><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:mi mathvariant="italic">δ</mml:mi><mml:mo>(</mml:mo><mml:mi>s</mml:mi><mml:mo>;</mml:mo><mml:mi mathvariant="italic">θ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula></p>
</sec>
<sec id="Ch1.S2.SS1.SSS3">
  <label>2.1.3</label><title>Statistical model and focus of this work</title>
      <p id="d2e820">Combining the noise model Eq. (<xref ref-type="disp-formula" rid="Ch1.E4"/>) and model inadequacy Eq. (<xref ref-type="disp-formula" rid="Ch1.E5"/>) yields the statistical model.</p>
      <p id="d2e827"><disp-formula id="Ch1.E6" content-type="numbered"><label>6</label><mml:math id="M39" display="block"><mml:mrow><mml:mi>y</mml:mi><mml:mo>-</mml:mo><mml:mi mathvariant="script">M</mml:mi><mml:mo>(</mml:mo><mml:mi>s</mml:mi><mml:mo>;</mml:mo><mml:mi mathvariant="italic">θ</mml:mi><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:mi mathvariant="italic">δ</mml:mi><mml:mo>(</mml:mo><mml:mi>s</mml:mi><mml:mo>;</mml:mo><mml:mi mathvariant="italic">θ</mml:mi><mml:mo>)</mml:mo><mml:mo>+</mml:mo><mml:mi mathvariant="italic">ε</mml:mi><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>
            which states that the deviation between the model predictions and the observations results from the superposition of measurement and model errors. As illustrated in Fig. <xref ref-type="fig" rid="F1"/>, this statistical model establishes the relation between model predictions, observations, and reality. Thus, it is used in the Bayesian calibration framework to find the <italic>most probable</italic> set of parameters for a given model <inline-formula><mml:math id="M40" display="inline"><mml:mi mathvariant="script">M</mml:mi></mml:math></inline-formula> and observation <inline-formula><mml:math id="M41" display="inline"><mml:mi>y</mml:mi></mml:math></inline-formula>.</p>

      <fig id="F1" specific-use="star"><label>Figure 1</label><caption><p id="d2e893">Statistical model connecting model predictions, reality, and observations. Adapted from <xref ref-type="bibr" rid="bib1.bibx51" id="text.43"/>.</p></caption>
            <graphic xlink:href="https://npg.copernicus.org/articles/33/425/2026/npg-33-425-2026-f01.png"/>

          </fig>

      <p id="d2e906">However, through this study, we aim to answer a different question: <italic>Given a model parameter</italic> <inline-formula><mml:math id="M42" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">θ</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>, <italic>which is the most informative observation</italic> <inline-formula><mml:math id="M43" display="inline"><mml:mrow><mml:msubsup><mml:mi>y</mml:mi><mml:mi>j</mml:mi><mml:mo>*</mml:mo></mml:msubsup></mml:mrow></mml:math></inline-formula> <italic>during calibration?</italic> Consequently, we isolate the effect of data selection through the use of synthetic data generated from the computational model that guarantees observations to be consistent with the model predictions, hence <inline-formula><mml:math id="M44" display="inline"><mml:mrow><mml:mi mathvariant="italic">δ</mml:mi><mml:mo>(</mml:mo><mml:mi>s</mml:mi><mml:mo>;</mml:mo><mml:mi mathvariant="italic">θ</mml:mi><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:msup><mml:mi>y</mml:mi><mml:mi>r</mml:mi></mml:msup><mml:mo>-</mml:mo><mml:mi mathvariant="script">M</mml:mi><mml:mo>(</mml:mo><mml:mi>s</mml:mi><mml:mo>;</mml:mo><mml:mi mathvariant="italic">θ</mml:mi><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0</mml:mn></mml:mrow></mml:math></inline-formula>.</p>
      <p id="d2e987">The resulting simplified statistical model can formally be re-written as

              <disp-formula id="Ch1.E7" content-type="numbered"><label>7</label><mml:math id="M45" display="block"><mml:mrow><mml:mi mathvariant="italic">θ</mml:mi><mml:mo>=</mml:mo><mml:msup><mml:mi mathvariant="script">M</mml:mi><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msup><mml:mo>(</mml:mo><mml:mi>s</mml:mi><mml:mo>;</mml:mo><mml:mi>y</mml:mi><mml:mo>,</mml:mo><mml:mi mathvariant="italic">ε</mml:mi><mml:mo>)</mml:mo><mml:mo>.</mml:mo></mml:mrow></mml:math></disp-formula></p>
      <p id="d2e1023">In reality, of course, the situation is much more involved in the sense that <inline-formula><mml:math id="M46" display="inline"><mml:mrow><mml:msup><mml:mi mathvariant="script">M</mml:mi><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula> denotes a computationally hard-to-solve calibration task. The formal write-up, however, indicates the aim of this study, which is not only to estimate parameter <inline-formula><mml:math id="M47" display="inline"><mml:mi mathvariant="italic">θ</mml:mi></mml:math></inline-formula> for a given observational dataset <inline-formula><mml:math id="M48" display="inline"><mml:mi>y</mml:mi></mml:math></inline-formula> but also to interpret <inline-formula><mml:math id="M49" display="inline"><mml:mi>y</mml:mi></mml:math></inline-formula> as a lever to improve the calibration result. The latter will be referred to as the <italic>data selection</italic> task. Developing a methodological approach to address this task is the major goal of this study. Although this goal differs from classical Bayesian calibration studies, it turns out that much of the existing work on Bayesian inference can be utilized.</p>
</sec>
</sec>
<sec id="Ch1.S2.SS2">
  <label>2.2</label><title>Bayesian Inference framework</title>
      <p id="d2e1073">We adopt the Bayesian approach to solve the statistical model formulated in Eq. (<xref ref-type="disp-formula" rid="Ch1.E7"/>) following <xref ref-type="bibr" rid="bib1.bibx39" id="text.44"/>. Underpinning everything that follows is Bayes' theorem, which reads

            <disp-formula id="Ch1.E8" content-type="numbered"><label>8</label><mml:math id="M50" display="block"><mml:mrow><mml:munder><mml:munder class="underbrace"><mml:mrow><mml:mi>P</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="italic">θ</mml:mi><mml:mo>∣</mml:mo><mml:mi>y</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo mathvariant="normal">︸</mml:mo></mml:munder><mml:mtext>Posterior</mml:mtext></mml:munder><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:munder><mml:munder class="underbrace"><mml:mrow><mml:mi>P</mml:mi><mml:mo>(</mml:mo><mml:mi>y</mml:mi><mml:mo>∣</mml:mo><mml:mi mathvariant="italic">θ</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo mathvariant="normal">︸</mml:mo></mml:munder><mml:mtext>Likelihood</mml:mtext></mml:munder><mml:mo>⋅</mml:mo><mml:munder><mml:munder class="underbrace"><mml:mrow><mml:mi>P</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="italic">θ</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo mathvariant="normal">︸</mml:mo></mml:munder><mml:mtext>Prior</mml:mtext></mml:munder></mml:mrow><mml:mrow><mml:munder><mml:munder class="underbrace"><mml:mrow><mml:msub><mml:mo>∫</mml:mo><mml:mi mathvariant="normal">Θ</mml:mi></mml:msub><mml:mi>P</mml:mi><mml:mo>(</mml:mo><mml:mi>y</mml:mi><mml:mo>∣</mml:mo><mml:mi mathvariant="italic">θ</mml:mi><mml:mo>)</mml:mo><mml:mo>⋅</mml:mo><mml:mi>P</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="italic">θ</mml:mi><mml:mo>)</mml:mo><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mi mathvariant="normal">d</mml:mi><mml:mi mathvariant="italic">θ</mml:mi></mml:mrow><mml:mo mathvariant="normal">︸</mml:mo></mml:munder><mml:mtext>Evidence</mml:mtext></mml:munder></mml:mrow></mml:mfrac></mml:mstyle></mml:mrow></mml:math></disp-formula>

          and captures the relation between prior knowledge of the parameter distribution for <inline-formula><mml:math id="M51" display="inline"><mml:mi mathvariant="italic">θ</mml:mi></mml:math></inline-formula>, observation <inline-formula><mml:math id="M52" display="inline"><mml:mi>y</mml:mi></mml:math></inline-formula>, and information return after combining both. Bayes' theorem enables us to determine the optimal parameter values consistent with the observed data by computing the probability distribution of the parameters conditioned on the observations. The resulting distribution is referred to as the posterior distribution – or simply, the posterior – and is denoted by <inline-formula><mml:math id="M53" display="inline"><mml:mrow><mml:mi>P</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="italic">θ</mml:mi><mml:mo>∣</mml:mo><mml:mi>y</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> in Eq. (<xref ref-type="disp-formula" rid="Ch1.E8"/>). The posterior represents the updated beliefs about the parameters after observing the data. To arrive at this, we encode our prior beliefs regarding the parameters into a probability distribution known as the prior distribution, also known as the prior, denoted in Eq. (<xref ref-type="disp-formula" rid="Ch1.E8"/>) as <inline-formula><mml:math id="M54" display="inline"><mml:mrow><mml:mi>P</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="italic">θ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>. We then update the prior distribution by multiplying it with the so-called likelihood function, <inline-formula><mml:math id="M55" display="inline"><mml:mrow><mml:mi>P</mml:mi><mml:mo>(</mml:mo><mml:mi>y</mml:mi><mml:mo>∣</mml:mo><mml:mi mathvariant="italic">θ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>, which, as the name implies, represents the likelihood of observing the data <inline-formula><mml:math id="M56" display="inline"><mml:mi>y</mml:mi></mml:math></inline-formula> for a given set of parameters <inline-formula><mml:math id="M57" display="inline"><mml:mi mathvariant="italic">θ</mml:mi></mml:math></inline-formula>. The prior distribution typically results from earlier Bayesian calibration steps or requires domain expertise or even empirical knowledge regarding the parameters. The likelihood function, on the other hand, follows from the noise model employed in the statistical model formulation (Eq. <xref ref-type="disp-formula" rid="Ch1.E7"/>). Assuming an additive Gaussian noise model with zero mean and covariance <inline-formula><mml:math id="M58" display="inline"><mml:mi mathvariant="bold">Σ</mml:mi></mml:math></inline-formula>, we formulate the likelihood function as shown below, where <inline-formula><mml:math id="M59" display="inline"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">d</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> denotes the dimension of the observation <inline-formula><mml:math id="M60" display="inline"><mml:mi>y</mml:mi></mml:math></inline-formula>.

            <disp-formula id="Ch1.E9" content-type="numbered"><label>9</label><mml:math id="M61" display="block"><mml:mtable rowspacing="0.2ex" class="split" displaystyle="true" columnalign="right left"><mml:mtr><mml:mtd><mml:mrow><mml:mi>P</mml:mi><mml:mo>(</mml:mo><mml:mi>y</mml:mi><mml:mo>∣</mml:mo><mml:mi mathvariant="italic">θ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mn mathvariant="normal">1</mml:mn><mml:msqrt><mml:mrow><mml:mo>(</mml:mo><mml:mn mathvariant="normal">2</mml:mn><mml:mi mathvariant="italic">π</mml:mi><mml:msup><mml:mo>)</mml:mo><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="normal">d</mml:mi></mml:msub></mml:mrow></mml:msup><mml:mi mathvariant="normal">det</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold">Σ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msqrt></mml:mfrac></mml:mstyle></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd/><mml:mtd><mml:mrow><mml:mo>⋅</mml:mo><mml:mi>exp⁡</mml:mi><mml:mfenced open="(" close=")"><mml:mrow><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mn mathvariant="normal">2</mml:mn></mml:mfrac></mml:mstyle><mml:mo>(</mml:mo><mml:mi>y</mml:mi><mml:mo>-</mml:mo><mml:mi mathvariant="script">M</mml:mi><mml:mo>(</mml:mo><mml:mi>s</mml:mi><mml:mo>;</mml:mo><mml:mi mathvariant="italic">θ</mml:mi><mml:mo>)</mml:mo><mml:msup><mml:mo>)</mml:mo><mml:mi>T</mml:mi></mml:msup><mml:msup><mml:mi mathvariant="bold">Σ</mml:mi><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msup><mml:mo>(</mml:mo><mml:mi>y</mml:mi><mml:mo>-</mml:mo><mml:mi mathvariant="script">M</mml:mi><mml:mo>(</mml:mo><mml:mi>s</mml:mi><mml:mo>;</mml:mo><mml:mi mathvariant="italic">θ</mml:mi><mml:mo>)</mml:mo><mml:mo>)</mml:mo></mml:mrow></mml:mfenced></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p>
      <p id="d2e1423">With the prior distribution and the likelihood function defined, we can now compute the posterior distribution using Bayes' theorem presented in Eq. (<xref ref-type="disp-formula" rid="Ch1.E8"/>). However, computing the posterior distribution involves integrating the product of likelihood and prior with respect to the parameter <inline-formula><mml:math id="M62" display="inline"><mml:mi mathvariant="italic">θ</mml:mi></mml:math></inline-formula>, an integral referred to as the evidence in Eq. (<xref ref-type="disp-formula" rid="Ch1.E8"/>). This integral is infeasible as soon as the computational model encapsulated in the likelihood function Eq. (<xref ref-type="disp-formula" rid="Ch1.E9"/>) gets costly to solve. Also, the dimension of the parameter space <inline-formula><mml:math id="M63" display="inline"><mml:mi mathvariant="normal">Θ</mml:mi></mml:math></inline-formula> impacts on computational feasibility. For a high-dimensional parameter distribution, the integral is also high-dimensional, making it one of the primary practical challenges of the Bayesian approach. Following <xref ref-type="bibr" rid="bib1.bibx15 bib1.bibx43" id="paren.45"/>, we address this challenge by approximating the posterior distribution using samples generated by Markov Chain Monte Carlo (MCMC) methods, averting the need to calculate the integral. These MCMC samples can be used to approximate the posterior distribution using methods such as kernel density estimation.</p>
      <p id="d2e1450">A schematic representation of the Bayesian approach is shown in Fig. <xref ref-type="fig" rid="F2"/>. We begin by defining the prior distributions for the parameters <inline-formula><mml:math id="M64" display="inline"><mml:mrow><mml:mi mathvariant="italic">θ</mml:mi><mml:mo>∈</mml:mo><mml:mo mathvariant="italic">{</mml:mo><mml:msub><mml:mi mathvariant="italic">θ</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi mathvariant="italic">θ</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:mi mathvariant="normal">…</mml:mi><mml:mo>,</mml:mo><mml:msub><mml:mi mathvariant="italic">θ</mml:mi><mml:mi>m</mml:mi></mml:msub><mml:mo mathvariant="italic">}</mml:mo></mml:mrow></mml:math></inline-formula>. Then, the likelihood function is used to update these priors, thereby obtaining the posterior distributions. As Fig. <xref ref-type="fig" rid="F2"/> illustrates, the likelihood function can be interpreted as the core driver of the process, with data resulting from observations guiding and refining each step. Consequently, the posterior distributions strongly depend on the observations used for calibration. The next section will be devoted to quantifying this dependence, ultimately leading to a better understanding of how to leverage observations <inline-formula><mml:math id="M65" display="inline"><mml:mi>y</mml:mi></mml:math></inline-formula> for more informative and reliable posterior distributions.</p>

      <fig id="F2"><label>Figure 2</label><caption><p id="d2e1504">Schematic illustration of the Bayesian inference process. The likelihood function acts as the core driver of updating the prior distribution to posterior distribution based on observational data. The resulting posterior depends strongly on type and quality of the observation.</p></caption>
          <graphic xlink:href="https://npg.copernicus.org/articles/33/425/2026/npg-33-425-2026-f02.png"/>

        </fig>

</sec>
<sec id="Ch1.S2.SS3">
  <label>2.3</label><title>Data selection using Information-theoretic measure</title>
      <p id="d2e1521">In the Bayesian framework, our prior beliefs about the parameters <inline-formula><mml:math id="M66" display="inline"><mml:mi mathvariant="italic">θ</mml:mi></mml:math></inline-formula> are updated by observations, but the extent of this update varies depending on which observations are used. To quantify how informative each candidate observation is, we measure the information gained during calibration with observation <inline-formula><mml:math id="M67" display="inline"><mml:mi>y</mml:mi></mml:math></inline-formula>. Observations reduce parameter uncertainty by updating the prior distribution <inline-formula><mml:math id="M68" display="inline"><mml:mrow><mml:mi>P</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="italic">θ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> to the posterior distribution <inline-formula><mml:math id="M69" display="inline"><mml:mrow><mml:mi>P</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="italic">θ</mml:mi><mml:mo>∣</mml:mo><mml:mi>y</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>. Thus, the information gained during this updating process corresponds to the reduced uncertainty. A fundamental information-theoretic measure of uncertainty is the entropy. For a random variable <inline-formula><mml:math id="M70" display="inline"><mml:mi mathvariant="italic">θ</mml:mi></mml:math></inline-formula> with probability distribution <inline-formula><mml:math id="M71" display="inline"><mml:mrow><mml:mi>P</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="italic">θ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>, entropy <inline-formula><mml:math id="M72" display="inline"><mml:mrow><mml:mi>H</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="italic">θ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> is given as:

            <disp-formula id="Ch1.E10" content-type="numbered"><label>10</label><mml:math id="M73" display="block"><mml:mrow><mml:mi>H</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="italic">θ</mml:mi><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:mo>-</mml:mo><mml:mo movablelimits="false">∫</mml:mo><mml:mi>P</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="italic">θ</mml:mi><mml:mo>)</mml:mo><mml:mi>log⁡</mml:mi><mml:mi>P</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="italic">θ</mml:mi><mml:mo>)</mml:mo><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mi mathvariant="normal">d</mml:mi><mml:mi mathvariant="italic">θ</mml:mi></mml:mrow></mml:math></disp-formula></p>
      <p id="d2e1649">Higher entropy indicates greater uncertainty about the parameter. Therefore, the change in entropy from prior to posterior quantifies the information gained from an observation. A standard measure to quantify information gain based on change in entropy is KL divergence <xref ref-type="bibr" rid="bib1.bibx30" id="paren.46"/>, also known as relative entropy. In the Bayesian setting, KL divergence between the posterior distribution (<inline-formula><mml:math id="M74" display="inline"><mml:mrow><mml:mi>P</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="italic">θ</mml:mi><mml:mo>∣</mml:mo><mml:mi>y</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>) and the prior distribution (<inline-formula><mml:math id="M75" display="inline"><mml:mrow><mml:mi>P</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="italic">θ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>), as defined in Eq. (<xref ref-type="disp-formula" rid="Ch1.E11"/>), admits a direct interpretation as information gain, since it quantifies the reduction in uncertainty about parameter <inline-formula><mml:math id="M76" display="inline"><mml:mi mathvariant="italic">θ</mml:mi></mml:math></inline-formula> achieved by conditioning on observation <inline-formula><mml:math id="M77" display="inline"><mml:mi>y</mml:mi></mml:math></inline-formula>. Higher KL divergence indicates greater reduction in uncertainty from prior to posterior, implying that the observation provides greater information about the parameters. Accordingly, we use this quantity as a criterion for data selection, consistent with its standard use in Bayesian optimal experimental design and inverse problems <xref ref-type="bibr" rid="bib1.bibx17 bib1.bibx7 bib1.bibx22 bib1.bibx4" id="paren.47"/>.</p>
      <p id="d2e1707"><disp-formula id="Ch1.E11" content-type="numbered"><label>11</label><mml:math id="M78" display="block"><mml:mrow><mml:msub><mml:mi>D</mml:mi><mml:mi mathvariant="normal">KL</mml:mi></mml:msub><mml:mo>(</mml:mo><mml:mi>P</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="italic">θ</mml:mi><mml:mo>∣</mml:mo><mml:mi>y</mml:mi><mml:mo>)</mml:mo><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mo>‖</mml:mo><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mi>P</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="italic">θ</mml:mi><mml:mo>)</mml:mo><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:mo movablelimits="false">∫</mml:mo><mml:mi>P</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="italic">θ</mml:mi><mml:mo>∣</mml:mo><mml:mi>y</mml:mi><mml:mo>)</mml:mo><mml:mi>log⁡</mml:mi><mml:mfenced open="(" close=")"><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:mi>P</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="italic">θ</mml:mi><mml:mo>∣</mml:mo><mml:mi>y</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mi>P</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="italic">θ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mfrac></mml:mstyle></mml:mfenced><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mi mathvariant="normal">d</mml:mi><mml:mi mathvariant="italic">θ</mml:mi></mml:mrow></mml:math></disp-formula></p>
      <p id="d2e1797">In the context of our work, we focus on calibrating landslide runout models with conceptual parameters that cannot be measured physically. For these parameters, we have limited prior information such as physically plausible bounds derived from literature and domain expertise. We therefore use uniform priors within these bounds, following standard practice in landslide runout calibration <xref ref-type="bibr" rid="bib1.bibx37 bib1.bibx2 bib1.bibx38" id="paren.48"/>. For uniform priors, the prior entropy <inline-formula><mml:math id="M79" display="inline"><mml:mrow><mml:mi>H</mml:mi><mml:mo>(</mml:mo><mml:mi>P</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="italic">θ</mml:mi><mml:mo>)</mml:mo><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> is constant, so maximizing KL divergence is equivalent to minimizing the posterior entropy <inline-formula><mml:math id="M80" display="inline"><mml:mrow><mml:mi>H</mml:mi><mml:mo>(</mml:mo><mml:mi>P</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="italic">θ</mml:mi><mml:mo>|</mml:mo><mml:mi>y</mml:mi><mml:mo>)</mml:mo><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> (see Appendix <xref ref-type="sec" rid="App1.Ch1.S1"/> for derivation). In our numerical experiments with uniform priors, maximizing KL divergence is therefore equivalent to selecting observations that minimize posterior entropy, yielding the sharpest and most informative posterior distributions.</p>
      <p id="d2e1850">A schematic representation of the data selection process is shown in Fig. <xref ref-type="fig" rid="F3"/>. We use the Bayesian calibration approach detailed in Sect. <xref ref-type="sec" rid="Ch1.S2.SS2"/> to calibrate the parameters <inline-formula><mml:math id="M81" display="inline"><mml:mi mathvariant="italic">θ</mml:mi></mml:math></inline-formula> using multiple observational datasets, where <inline-formula><mml:math id="M82" display="inline"><mml:mrow><mml:msub><mml:mi>y</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>∈</mml:mo><mml:mo mathvariant="italic">{</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:mi mathvariant="normal">…</mml:mi><mml:mo>,</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mi>n</mml:mi></mml:msub><mml:mo mathvariant="italic">}</mml:mo></mml:mrow></mml:math></inline-formula> denotes the <inline-formula><mml:math id="M83" display="inline"><mml:mi>i</mml:mi></mml:math></inline-formula>th dataset. Note that each individual dataset <inline-formula><mml:math id="M84" display="inline"><mml:mrow><mml:msub><mml:mi>y</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> can either constitute a single scalar observation or a sequence of observations such as time series of velocity or position. Consequently, we obtain <inline-formula><mml:math id="M85" display="inline"><mml:mi>n</mml:mi></mml:math></inline-formula> posterior distributions, one for each observational dataset, which we compare against the prior distribution using the KL divergence defined in Eq. (<xref ref-type="disp-formula" rid="Ch1.E11"/>). This yields a quantitative measure of the information that each observation contributes to the calibration process, allowing us to identify the most informative dataset <inline-formula><mml:math id="M86" display="inline"><mml:mrow><mml:msup><mml:mover accent="true"><mml:mi>y</mml:mi><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:mo>*</mml:mo></mml:msup></mml:mrow></mml:math></inline-formula> to infer the parameters <inline-formula><mml:math id="M87" display="inline"><mml:mi mathvariant="italic">θ</mml:mi></mml:math></inline-formula>. The bar above <inline-formula><mml:math id="M88" display="inline"><mml:mrow><mml:msup><mml:mover accent="true"><mml:mi>y</mml:mi><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:mo>*</mml:mo></mml:msup></mml:mrow></mml:math></inline-formula> indicates its aggregate nature. Although <inline-formula><mml:math id="M89" display="inline"><mml:mrow><mml:msup><mml:mover accent="true"><mml:mi>y</mml:mi><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:mo>*</mml:mo></mml:msup></mml:mrow></mml:math></inline-formula> maximizes the information for joint calibration, it may not be the most informative dataset to calibrate a specific parameter <inline-formula><mml:math id="M90" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">θ</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>.</p>

      <fig id="F3" specific-use="star"><label>Figure 3</label><caption><p id="d2e1995">Schematic illustration of the Bayesian data selection process using information-theoretic metrics. Multiple calibration routines are performed in parallel, each using the same likelihood function but leveraging different observations to update the prior distributions. By comparing each resulting posterior distribution against the prior, we can quantify the information gained during calibration relative to observations.</p></caption>
          <graphic xlink:href="https://npg.copernicus.org/articles/33/425/2026/npg-33-425-2026-f03.png"/>

        </fig>

      <p id="d2e2004">To identify the most informative dataset for a given parameter <inline-formula><mml:math id="M91" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">θ</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>, we marginalize the posteriors with respect to parameters and obtain <inline-formula><mml:math id="M92" display="inline"><mml:mi>n</mml:mi></mml:math></inline-formula> posteriors for each parameter <inline-formula><mml:math id="M93" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">θ</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>∈</mml:mo><mml:mo mathvariant="italic">{</mml:mo><mml:msub><mml:mi mathvariant="italic">θ</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi mathvariant="italic">θ</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:mi mathvariant="normal">…</mml:mi><mml:mo>,</mml:mo><mml:msub><mml:mi mathvariant="italic">θ</mml:mi><mml:mi>m</mml:mi></mml:msub><mml:mo mathvariant="italic">}</mml:mo></mml:mrow></mml:math></inline-formula>, given as <inline-formula><mml:math id="M94" display="inline"><mml:mrow><mml:mo mathvariant="italic">{</mml:mo><mml:mi>P</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi mathvariant="italic">θ</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>∣</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub><mml:mo>)</mml:mo><mml:mo>,</mml:mo><mml:mi>P</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi mathvariant="italic">θ</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>∣</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msub><mml:mo>)</mml:mo><mml:mo>,</mml:mo><mml:mi mathvariant="normal">…</mml:mi><mml:mo>,</mml:mo><mml:mi>P</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi mathvariant="italic">θ</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>∣</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mi>n</mml:mi></mml:msub><mml:mo>)</mml:mo><mml:mo mathvariant="italic">}</mml:mo></mml:mrow></mml:math></inline-formula>. We then compute the KL divergence between the individual prior distribution <inline-formula><mml:math id="M95" display="inline"><mml:mrow><mml:mi>P</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi mathvariant="italic">θ</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> and the posterior distributions of <inline-formula><mml:math id="M96" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">θ</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>  corresponding to each of the observations. The generalized formulation for computing the KL divergence between the prior and posterior distribution of the parameter <inline-formula><mml:math id="M97" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">θ</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> calibrated using observation <inline-formula><mml:math id="M98" display="inline"><mml:mrow><mml:msub><mml:mi>y</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> is given in Eq. (<xref ref-type="disp-formula" rid="Ch1.E12"/>).</p>
      <p id="d2e2192"><disp-formula id="Ch1.E12" content-type="numbered"><label>12</label><mml:math id="M99" display="block"><mml:mtable class="split" rowspacing="0.2ex" displaystyle="true" columnalign="right left"><mml:mtr><mml:mtd/><mml:mtd><mml:mrow><mml:msubsup><mml:mi>D</mml:mi><mml:mi mathvariant="normal">KL</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msubsup><mml:mo>(</mml:mo><mml:mi>P</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi mathvariant="italic">θ</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>∣</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>)</mml:mo><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mo>‖</mml:mo><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mi>P</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi mathvariant="italic">θ</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>)</mml:mo><mml:mo>)</mml:mo><mml:mo>=</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd/><mml:mtd><mml:mrow><mml:mspace width="1em" linebreak="nobreak"/><mml:mo movablelimits="false">∫</mml:mo><mml:mi>P</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi mathvariant="italic">θ</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>∣</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>)</mml:mo><mml:mi>log⁡</mml:mi><mml:mfenced open="(" close=")"><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:mi>P</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi mathvariant="italic">θ</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>∣</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mi>P</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi mathvariant="italic">θ</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:mfrac></mml:mstyle></mml:mfenced><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mi mathvariant="normal">d</mml:mi><mml:msub><mml:mi mathvariant="italic">θ</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p>
      <p id="d2e2325">Thus, for each parameter <inline-formula><mml:math id="M100" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">θ</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>, we have <inline-formula><mml:math id="M101" display="inline"><mml:mi>n</mml:mi></mml:math></inline-formula> KL divergence values denoted as <inline-formula><mml:math id="M102" display="inline"><mml:mrow><mml:msubsup><mml:mi>D</mml:mi><mml:mi mathvariant="normal">KL</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msubsup></mml:mrow></mml:math></inline-formula>. These values correspond to observations <inline-formula><mml:math id="M103" display="inline"><mml:mrow><mml:msub><mml:mi>y</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>, with <inline-formula><mml:math id="M104" display="inline"><mml:mi>i</mml:mi></mml:math></inline-formula> ranging from <inline-formula><mml:math id="M105" display="inline"><mml:mrow><mml:mo mathvariant="italic">{</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mi mathvariant="normal">…</mml:mi><mml:mi>n</mml:mi><mml:mo mathvariant="italic">}</mml:mo></mml:mrow></mml:math></inline-formula>. Since these values quantify the information gained through the Bayesian update, we can utilize them to select the dataset <inline-formula><mml:math id="M106" display="inline"><mml:mrow><mml:msubsup><mml:mi>y</mml:mi><mml:mi>j</mml:mi><mml:mo>*</mml:mo></mml:msubsup></mml:mrow></mml:math></inline-formula> that constitutes the most informative dataset for calibrating the parameter <inline-formula><mml:math id="M107" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">θ</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>.</p>
      <p id="d2e2424"><disp-formula id="Ch1.E13" content-type="numbered"><label>13</label><mml:math id="M108" display="block"><mml:mrow><mml:mtable rowspacing="0.2ex" class="split" displaystyle="true" columnalign="right left"><mml:mtr><mml:mtd/><mml:mtd><mml:mrow><mml:mtext>Data selection task:</mml:mtext></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd/><mml:mtd><mml:mrow><mml:mspace linebreak="nobreak" width="1em"/><mml:mtext>For a given parameter </mml:mtext><mml:msub><mml:mi mathvariant="italic">θ</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mspace width="0.25em" linebreak="nobreak"/><mml:mspace width="0.25em" linebreak="nobreak"/><mml:mtext>find</mml:mtext><mml:mspace linebreak="nobreak" width="0.25em"/><mml:mspace width="0.25em" linebreak="nobreak"/><mml:msubsup><mml:mi>y</mml:mi><mml:mi>j</mml:mi><mml:mo>*</mml:mo></mml:msubsup><mml:mo>=</mml:mo><mml:mi>arg⁡</mml:mi><mml:munder><mml:mo movablelimits="false">max⁡</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>∈</mml:mo><mml:mo mathvariant="italic">{</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>,</mml:mo><mml:mi mathvariant="normal">…</mml:mi><mml:mo>,</mml:mo><mml:mi>n</mml:mi><mml:mo mathvariant="italic">}</mml:mo></mml:mrow></mml:munder><mml:msubsup><mml:mi>D</mml:mi><mml:mi mathvariant="normal">KL</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msubsup></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:mrow></mml:math></disp-formula></p>
      <p id="d2e2504">Bayesian calibration combined with data selection based on KL divergence constitutes the methodology of Bayesian data selection. This allows us to understand and quantify how adept available observations are in constraining a given parameter.</p>
</sec>
<sec id="Ch1.S2.SS4">
  <label>2.4</label><title>Workflow</title>
      <p id="d2e2516">We implement our Bayesian data selection workflow, as illustrated in Fig. <xref ref-type="fig" rid="F4"/>, using <monospace>PSimPy</monospace>, a Python-based package for predictive and probabilistic simulations <xref ref-type="bibr" rid="bib1.bibx57" id="paren.49"/>. The workflow comprises three distinct phases: (i) Surrogate Modeling; (ii) Bayesian Parameter Calibration; (iii) Data Selection.</p>

      <fig id="F4" specific-use="star"><label>Figure 4</label><caption><p id="d2e2529">Overview of the Bayesian data selection workflow. The workflow consists of three key phases: (i) Surrogate Modeling, where a computationally efficient surrogate replaces the expensive forward model; (ii) Bayesian Parameter Calibration, where an MCMC sampler is used to infer the posterior distribution of the parameters; and (iii) Data Selection, where KL divergence quantifies the information gain from different observational data.  </p></caption>
          <graphic xlink:href="https://npg.copernicus.org/articles/33/425/2026/npg-33-425-2026-f04.png"/>

        </fig>

<sec id="Ch1.S2.SS4.SSS1">
  <label>2.4.1</label><title>Surrogate Modeling</title>
      <p id="d2e2545">Surrogate modeling serves as a computational enabler to overcome computational bottlenecks in calculating the KL divergence. As illustrated in Eq. (<xref ref-type="disp-formula" rid="Ch1.E11"/>), calculating KL divergence requires posterior distributions, which are approximated through MCMC sampling (see Sect. <xref ref-type="sec" rid="Ch1.S2.SS2"/>). The accuracy of this approximation is based on the ability of MCMC chains to effectively explore the posterior space, which typically requires a large number of samples. Each sample necessitates evaluating the likelihood function, entailing one complete execution of the computational model. Therefore, this approach becomes infeasible for computationally expensive models. To tackle this challenge, we employ GP emulators, a non-intrusive surrogate modelling technique that reduces computational costs in Bayesian calibration workflows <xref ref-type="bibr" rid="bib1.bibx58" id="paren.50"/>. The widespread adoption of GP emulators stems from their ability to provide probabilistic predictions, allowing for rigorous quantification of the uncertainty associated with predictions. Furthermore, they offer efficient performance with limited training datasets compared to other machine learning approaches. Mathematically, a GP is defined by a mean function <inline-formula><mml:math id="M109" display="inline"><mml:mrow><mml:mi>m</mml:mi><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> and a covariance function <inline-formula><mml:math id="M110" display="inline"><mml:mrow><mml:mi>k</mml:mi><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:msup><mml:mi>x</mml:mi><mml:mo>′</mml:mo></mml:msup><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> as shown in Eq. (<xref ref-type="disp-formula" rid="Ch1.E14"/>).</p>
      <p id="d2e2593"><disp-formula id="Ch1.E14" content-type="numbered"><label>14</label><mml:math id="M111" display="block"><mml:mrow><mml:mi>f</mml:mi><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mo>)</mml:mo><mml:mo>∼</mml:mo><mml:mi mathvariant="script">G</mml:mi><mml:mi mathvariant="script">P</mml:mi><mml:mo>(</mml:mo><mml:mi>m</mml:mi><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mo>)</mml:mo><mml:mo>,</mml:mo><mml:mi>k</mml:mi><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:msup><mml:mi>x</mml:mi><mml:mo>′</mml:mo></mml:msup><mml:mo>)</mml:mo><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula></p>
      <p id="d2e2644">Both the mean and covariance functions in Eq. (<xref ref-type="disp-formula" rid="Ch1.E14"/>) are parameterized by hyperparameters that are inferred from the training data. To implement the GP emulator, we utilize the <bold><monospace>emulator</monospace></bold> module of <monospace>PSimPy</monospace>, which harnesses RobustGaSP, an R package for Gaussian stochastic process emulation that provides robust estimates of the hyperparameters leading to enhanced predictive performance <xref ref-type="bibr" rid="bib1.bibx16" id="paren.51"/>.</p>
      <p id="d2e2659">Figure <xref ref-type="fig" rid="F4"/> illustrates the key steps involved in the surrogate modeling phase. We start by generating a set of input parameters using the <bold><monospace>sampler</monospace></bold> module of <monospace>PSimPy</monospace>, which leverages space-filling schemes like Latin hypercube sampling. Next, we employ the <bold><monospace>simulator</monospace></bold> module to evaluate our computational model <inline-formula><mml:math id="M112" display="inline"><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="script">M</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>, at these input points. The resulting model outputs are then post-processed according to the emulation strategy determined by the observation dimensionality. For scalar observations, we emulate the parameter-to-observable map, replacing the forward model <inline-formula><mml:math id="M113" display="inline"><mml:mi mathvariant="script">M</mml:mi></mml:math></inline-formula> in Eq. (<xref ref-type="disp-formula" rid="Ch1.E9"/>) with a cost-effective surrogate <inline-formula><mml:math id="M114" display="inline"><mml:mover accent="true"><mml:mi mathvariant="script">M</mml:mi><mml:mo stretchy="true" mathvariant="normal">^</mml:mo></mml:mover></mml:math></inline-formula>. Alternatively, for high-dimensional observations (e.g., velocity or position time series), we directly emulate the likelihood function, which quantifies the mismatch between model output and observation. This approach leverages the fact that the likelihood is scalar-valued regardless of observation dimensionality, thereby avoiding the computational challenges of constructing GP emulators with high-dimensional outputs. These post-processed outputs, together with the set of input points, constitute our training data used to build and train the GP emulator. We then validate the trained surrogate using <inline-formula><mml:math id="M115" display="inline"><mml:mi>k</mml:mi></mml:math></inline-formula>-fold cross-validation. After successful validation, the surrogate model is available for predictions.</p>
</sec>
<sec id="Ch1.S2.SS4.SSS2">
  <label>2.4.2</label><title>Bayesian Parameter calibration</title>
      <p id="d2e2723">The trained GP surrogate from the surrogate modeling phase replaces the expensive computational model in the likelihood function in Eq. (<xref ref-type="disp-formula" rid="Ch1.E9"/>), allowing us to sample the posterior distribution using MCMC methods. For this purpose, we utilize the <bold><monospace>mcmc sampler</monospace></bold> of <monospace>PSimPy</monospace>, which is based on Python's <monospace>emcee</monospace> package, an affine invariant MCMC ensemble sampler <xref ref-type="bibr" rid="bib1.bibx13" id="paren.52"/>. This implies that the sampler is unaffected by affine transformations of parameter space, allowing it to sample complex probability distributions without the need for problem-specific tuning. Furthermore, it employs multiple chains that evolve in parallel, resulting in an efficient exploration of the probability distribution. Additionally, we assess the convergence of the MCMC chains using the <bold><monospace>diagnostics</monospace></bold> module, leveraging Python's <monospace>Arviz</monospace> package <xref ref-type="bibr" rid="bib1.bibx31" id="paren.53"/>. Using this module, we qualitatively assess the convergence using trace plots and then quantify it using standard diagnostic metrics such as Gelman-Rubin statistic <inline-formula><mml:math id="M116" display="inline"><mml:mover accent="true"><mml:mi>R</mml:mi><mml:mo mathvariant="normal" stretchy="true">^</mml:mo></mml:mover></mml:math></inline-formula> (see Appendix <xref ref-type="sec" rid="App1.Ch1.S2"/>).</p>
</sec>
<sec id="Ch1.S2.SS4.SSS3">
  <label>2.4.3</label><title>Data selection</title>
      <p id="d2e2772">In this phase, we compute the KL divergence between the posterior and prior distributions obtained from the Bayesian calibration phase. As discussed in Sect. <xref ref-type="sec" rid="Ch1.S2.SS3"/>, this involves computing the intractable integral, which we approximate using the distance-based KL divergence estimator proposed by <xref ref-type="bibr" rid="bib1.bibx48" id="text.54"/>. To this end, we incorporate the code from <xref ref-type="bibr" rid="bib1.bibx18" id="text.55"/> into our workflow, as it implements the estimators described by <xref ref-type="bibr" rid="bib1.bibx48" id="text.56"/>. We marginalize the posterior distributions from the Bayesian calibration phase with respect to the parameters and compute KL divergence, as presented in Sect. <xref ref-type="sec" rid="Ch1.S2.SS3"/>. By calibrating <inline-formula><mml:math id="M117" display="inline"><mml:mi>m</mml:mi></mml:math></inline-formula> parameters using <inline-formula><mml:math id="M118" display="inline"><mml:mi>n</mml:mi></mml:math></inline-formula> observations, we end up with <inline-formula><mml:math id="M119" display="inline"><mml:mrow><mml:mi>n</mml:mi><mml:mo>×</mml:mo><mml:mi>m</mml:mi></mml:mrow></mml:math></inline-formula> KL divergence matrix, as shown in Fig. <xref ref-type="fig" rid="F5"/>, where each entry, <inline-formula><mml:math id="M120" display="inline"><mml:mrow><mml:msubsup><mml:mi>D</mml:mi><mml:mi mathvariant="normal">KL</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msubsup></mml:mrow></mml:math></inline-formula>, quantifies the information that <inline-formula><mml:math id="M121" display="inline"><mml:mi>i</mml:mi></mml:math></inline-formula>th observation <inline-formula><mml:math id="M122" display="inline"><mml:mrow><mml:msub><mml:mi>y</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> provides for calibrating <inline-formula><mml:math id="M123" display="inline"><mml:mi>j</mml:mi></mml:math></inline-formula>th parameter <inline-formula><mml:math id="M124" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">θ</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>. Using this KL divergence matrix we can identify the most informative observation <inline-formula><mml:math id="M125" display="inline"><mml:mrow><mml:msubsup><mml:mi>y</mml:mi><mml:mi>j</mml:mi><mml:mo>*</mml:mo></mml:msubsup></mml:mrow></mml:math></inline-formula> for a given parameter <inline-formula><mml:math id="M126" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">θ</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> by maximizing the KL divergence across available observations (as defined in Eq. <xref ref-type="disp-formula" rid="Ch1.E13"/>).</p>

      <fig id="F5"><label>Figure 5</label><caption><p id="d2e2900">KL divergence matrix generated in the calibration of <inline-formula><mml:math id="M127" display="inline"><mml:mi>m</mml:mi></mml:math></inline-formula> parameters with <inline-formula><mml:math id="M128" display="inline"><mml:mi>n</mml:mi></mml:math></inline-formula> observation. Each entry of the matrix quantifies the information provided by observation <inline-formula><mml:math id="M129" display="inline"><mml:mrow><mml:msub><mml:mi>y</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> for calibrating parameter <inline-formula><mml:math id="M130" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">θ</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>. </p></caption>
            <graphic xlink:href="https://npg.copernicus.org/articles/33/425/2026/npg-33-425-2026-f05.png"/>

          </fig>

</sec>
</sec>
</sec>
<sec id="Ch1.S3">
  <label>3</label><title>Case Study</title>
      <p id="d2e2955">This study investigates how the choice of observational data influences the calibration of model parameters. Specifically, we aim to identify the most informative dataset for calibrating a given parameter. To this end, we performed multiple calibration routines using the Bayesian calibration workflow outlined in Sect. <xref ref-type="sec" rid="Ch1.S2.SS4"/>, each with a different observational dataset. We then compute the KL divergence between the prior and posterior distributions to quantify the information gained from each dataset. As shown in Sect. <xref ref-type="sec" rid="Ch1.S2.SS1"/>, this process requires three key components: (1) a computational model with parameters to be calibrated, (2) observations to update those parameters, and (3) a noise model that captures uncertainty in observations. The noise model has already been described in Eq. (<xref ref-type="disp-formula" rid="Ch1.E4"/>). This section describes the computational model and multiple observational datasets used in the calibration.</p>
<sec id="Ch1.S3.SS1">
  <label>3.1</label><title>Computational Model</title>
      <p id="d2e2971"><xref ref-type="bibr" rid="bib1.bibx23" id="text.57"/> categorized the numerical landslide runout models into models based on continuum mechanics and lumped mass models. In both models, gravity primarily drives the motion, while a friction term, dependent on the chosen rheological model, resists it. However, due to the complexity of landslide dynamics, these rheological formulations often involve conceptual parameters that are not directly measurable and must, therefore, be inferred through model calibration. While lumped mass models idealize the sliding landslide mass as a single mass point, continuum-based models treat it as a spatially distributed deformable mass governed by conservation laws <xref ref-type="bibr" rid="bib1.bibx55" id="paren.58"/>. As a result, lumped mass models are limited in their ability to represent internal deformations, which are captured by continuum-based models, thereby enabling a more accurate simulation of flow dynamics and deposit morphology <xref ref-type="bibr" rid="bib1.bibx21" id="paren.59"/>. However, lumped mass models offer a conceptually straightforward framework for estimating bulk characteristics such as runout distances, velocities, and accelerations <xref ref-type="bibr" rid="bib1.bibx56" id="paren.60"/>. Furthermore, conceptual simplicity allows for a clear, tractable mapping between model parameters and landslide dynamics, which is often obscured in continuum models. Thus, we deliberately employ lumped mass models, given that the primary aim of this study is to assess how observational data influence parameter calibration.</p>
      <p id="d2e2985">Governing equation of a lumped mass model is mathematically described using Newton's second law of motion, as shown in Eq. (<xref ref-type="disp-formula" rid="Ch1.E15"/>).</p>
      <p id="d2e2990"><disp-formula id="Ch1.E15" content-type="numbered"><label>15</label><mml:math id="M131" display="block"><mml:mrow><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:mi mathvariant="normal">d</mml:mi><mml:mi>u</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="normal">d</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:mfrac></mml:mstyle><mml:mo>=</mml:mo><mml:mi>g</mml:mi><mml:mi>sin⁡</mml:mi><mml:mi mathvariant="italic">β</mml:mi><mml:mo>-</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:msub><mml:mi>F</mml:mi><mml:mi mathvariant="normal">res</mml:mi></mml:msub></mml:mrow><mml:mi>m</mml:mi></mml:mfrac></mml:mstyle></mml:mrow></mml:math></disp-formula>
          Here, <inline-formula><mml:math id="M132" display="inline"><mml:mi>u</mml:mi></mml:math></inline-formula> is the tangential velocity of the idealized mass point, <inline-formula><mml:math id="M133" display="inline"><mml:mi>g</mml:mi></mml:math></inline-formula> is the gravitational constant, <inline-formula><mml:math id="M134" display="inline"><mml:mi mathvariant="italic">β</mml:mi></mml:math></inline-formula> is the slope angle, and <inline-formula><mml:math id="M135" display="inline"><mml:mrow><mml:msub><mml:mi>F</mml:mi><mml:mi mathvariant="normal">res</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> is the resisting force term. We choose the Voellmy rheological model, which includes a classical dry Coulomb friction coefficient <inline-formula><mml:math id="M136" display="inline"><mml:mi mathvariant="italic">μ</mml:mi></mml:math></inline-formula> representing the basal resistance to landslide motion, along with a velocity-dependent friction term known as the turbulent friction coefficient <inline-formula><mml:math id="M137" display="inline"><mml:mi mathvariant="italic">ξ</mml:mi></mml:math></inline-formula>. The corresponding formulation of the resistance force is given as:

            <disp-formula id="Ch1.E16" content-type="numbered"><label>16</label><mml:math id="M138" display="block"><mml:mrow><mml:msub><mml:mi>F</mml:mi><mml:mi mathvariant="normal">res</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mi mathvariant="italic">μ</mml:mi><mml:mi>m</mml:mi><mml:mi>g</mml:mi><mml:mi>cos⁡</mml:mi><mml:mi mathvariant="italic">β</mml:mi><mml:mo>+</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:mi>m</mml:mi><mml:mi>g</mml:mi></mml:mrow><mml:mi mathvariant="italic">ξ</mml:mi></mml:mfrac></mml:mstyle><mml:msup><mml:mi>u</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:math></disp-formula></p>
      <p id="d2e3115">In addition to the rheological parameters described in Eq. (<xref ref-type="disp-formula" rid="Ch1.E16"/>) <inline-formula><mml:math id="M139" display="inline"><mml:mrow><mml:mo>(</mml:mo><mml:mo mathvariant="italic">{</mml:mo><mml:mi mathvariant="italic">μ</mml:mi><mml:mo>,</mml:mo><mml:mi mathvariant="italic">ξ</mml:mi><mml:mo mathvariant="italic">}</mml:mo><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>, we need to provide parameterized information <inline-formula><mml:math id="M140" display="inline"><mml:mrow><mml:mo>(</mml:mo><mml:mi>s</mml:mi><mml:mo>∈</mml:mo><mml:mo mathvariant="italic">{</mml:mo><mml:mi>T</mml:mi><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi><mml:mo>)</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mi>u</mml:mi><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mn mathvariant="normal">0</mml:mn></mml:msub></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>u</mml:mi><mml:mrow><mml:msub><mml:mi>y</mml:mi><mml:mn mathvariant="normal">0</mml:mn></mml:msub></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mn mathvariant="normal">0</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mn mathvariant="normal">0</mml:mn></mml:msub><mml:mo mathvariant="italic">}</mml:mo><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> specifying the landslide simulation scenario, where <inline-formula><mml:math id="M141" display="inline"><mml:mrow><mml:mi>T</mml:mi><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> denotes the topography and the other parameters represent the initial conditions, such as initial velocity and position. In this study, we employ a synthetic topography representing an idealized digital elevation model (DEM). As shown in Fig. <xref ref-type="fig" rid="F6"/>, the topography consists of a curved longitudinal profile that is uniform in the transverse direction.</p>

      <fig id="F6" specific-use="star"><label>Figure 6</label><caption><p id="d2e3230">Vertical cross-section of the synthetic topographic model used in this study, with elevation (<inline-formula><mml:math id="M142" display="inline"><mml:mi>z</mml:mi></mml:math></inline-formula>) plotted against the horizontal coordinate (<inline-formula><mml:math id="M143" display="inline"><mml:mi>x</mml:mi></mml:math></inline-formula>). The topography is constant along the transverse direction. The red marker denotes the position of the initial release point projected onto this cross-section.</p></caption>
          <graphic xlink:href="https://npg.copernicus.org/articles/33/425/2026/npg-33-425-2026-f06.png"/>

        </fig>

</sec>
<sec id="Ch1.S3.SS2">
  <label>3.2</label><title>Observational datasets</title>
      <p id="d2e3261">We curate a diverse set of observations, each capturing different aspects of landslide dynamics, to systematically assess the influence of data selection on Bayesian calibration outcome. However, obtaining such diverse observational datasets in the real world is often infeasible due to logistical and financial constraints <xref ref-type="bibr" rid="bib1.bibx44" id="paren.61"/>. To address this, the study uses synthetic data to calibrate the friction parameters of the lumped mass model described in Sect. <xref ref-type="sec" rid="Ch1.S3.SS1"/>. Synthetic data also provide the advantage of known ground-truth parameters that can be directly compared with the inferred values.</p>
      <p id="d2e3269">The observational data used in this study can be categorized as: (1) aggregate observations, which reflect bulk characteristics such as runout distance and maximum velocity, and (2) time series observations, which capture the velocity and position of the sliding mass over time. We generate these observations by evaluating the lumped mass model with a selected set of friction parameters. These parameters are arbitrarily selected within the bounds reported in the literature: the dry Coulomb friction coefficient <inline-formula><mml:math id="M144" display="inline"><mml:mrow><mml:mi mathvariant="italic">μ</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.23</mml:mn></mml:mrow></mml:math></inline-formula>, chosen from the range <inline-formula><mml:math id="M145" display="inline"><mml:mrow><mml:mo>[</mml:mo><mml:mn mathvariant="normal">0.02</mml:mn><mml:mo>,</mml:mo><mml:mn mathvariant="normal">0.3</mml:mn><mml:mo>]</mml:mo></mml:mrow></mml:math></inline-formula>,  and the turbulent friction coefficient <inline-formula><mml:math id="M146" display="inline"><mml:mrow><mml:mi mathvariant="italic">ξ</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1000</mml:mn></mml:mrow></mml:math></inline-formula>, chosen from the range <inline-formula><mml:math id="M147" display="inline"><mml:mrow><mml:mo>[</mml:mo><mml:mn mathvariant="normal">100</mml:mn><mml:mo>,</mml:mo><mml:mn mathvariant="normal">2200</mml:mn><mml:mo>]</mml:mo></mml:mrow></mml:math></inline-formula>. The lumped mass model is then evaluated with these parameters, generating a time history of velocity and position of the sliding mass. These outputs are post-processed to create both aggregate and time series observational datasets, with added noise to mimic real-world measurement uncertainty. For aggregate observations, which are scalar quantities (e.g., runout distance or maximum velocity), perturbations are drawn from a Gaussian distribution with zero mean and a standard deviation equal to one-tenth of the scalar's magnitude. For time series observations, we assume that errors in time steps are independent of each other. Then, each time step is perturbed using noise drawn from a Gaussian distribution with zero mean and unit standard deviation and then scaled by one-tenth of the maximum value in the respective time series (velocity or position). The resulting noisy values are clipped to zero wherever they become negative.</p>
</sec>
<sec id="Ch1.S3.SS3">
  <label>3.3</label><title>Design of Numerical Experiments</title>
      <p id="d2e3337">We conducted a series of eight numerical experiments using the synthetic datasets described in Sect. <xref ref-type="sec" rid="Ch1.S3.SS2"/> to calibrate the friction parameters of the lumped mass model. In experiments 1 through 4, we vary the information content and granularity of the observational data to examine the influence of data selection on the calibration outcome. Experiments 1 and 2 use aggregated observations – maximum velocity and runout distance, respectively – while Experiments 3 and 4 use time series data of velocity <inline-formula><mml:math id="M148" display="inline"><mml:mrow><mml:mi>u</mml:mi><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> and position <inline-formula><mml:math id="M149" display="inline"><mml:mrow><mml:mi>x</mml:mi><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>.</p>
      <p id="d2e3370">Experiments 5 and 6 investigate how the temporal characteristics of time series data, specifically length, and resolution, impact calibration outcomes. While experiment 5 involves calibration routines performed using time series data of varying lengths, experiment 6 uses time series data of varying resolution. Together, these experiments evaluate how the length and resolution of the time series data influence the accuracy of parameter estimation.</p>
      <p id="d2e3373">Experiments 7 and 8, extend the calibration to include the noise model’s discrepancy parameter (see Eq. <xref ref-type="disp-formula" rid="Ch1.E4"/>), which was previously fixed using heuristic assumptions. These experiments explore how jointly estimating observational uncertainty, along with <inline-formula><mml:math id="M150" display="inline"><mml:mi mathvariant="italic">μ</mml:mi></mml:math></inline-formula> and <inline-formula><mml:math id="M151" display="inline"><mml:mi mathvariant="italic">ξ</mml:mi></mml:math></inline-formula>, affects calibration outcomes. Velocity and position time series are again used as observational datasets in these cases. In all the 8 experiments we assume uniform priors for the friction parameters. The bounds for the priors are chosen from the literature as discussed in Sect. <xref ref-type="sec" rid="Ch1.S2.SS3"/>, for <inline-formula><mml:math id="M152" display="inline"><mml:mrow><mml:mi mathvariant="italic">μ</mml:mi><mml:mo>∈</mml:mo><mml:mo>[</mml:mo><mml:mn mathvariant="normal">0.02</mml:mn><mml:mo>,</mml:mo><mml:mn mathvariant="normal">0.3</mml:mn><mml:mo>]</mml:mo></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M153" display="inline"><mml:mrow><mml:mi mathvariant="italic">ξ</mml:mi><mml:mo>∈</mml:mo><mml:mo>[</mml:mo><mml:mn mathvariant="normal">100</mml:mn><mml:mo>,</mml:mo><mml:mn mathvariant="normal">2200</mml:mn><mml:mo>]</mml:mo></mml:mrow></mml:math></inline-formula> m s<sup>−2</sup>. Additionally in experiment 7 and 8, we assume an uniform prior for the discrepancy parameters <inline-formula><mml:math id="M155" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">σ</mml:mi><mml:mrow><mml:msub><mml:mi mathvariant="normal">vel</mml:mi><mml:mi mathvariant="normal">TS</mml:mi></mml:msub></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M156" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">σ</mml:mi><mml:mrow><mml:msub><mml:mi mathvariant="normal">pos</mml:mi><mml:mi mathvariant="normal">TS</mml:mi></mml:msub></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula> with bounds <inline-formula><mml:math id="M157" display="inline"><mml:mrow><mml:mo>[</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>,</mml:mo><mml:mn mathvariant="normal">5</mml:mn><mml:mo>]</mml:mo></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M158" display="inline"><mml:mrow><mml:mo>[</mml:mo><mml:mn mathvariant="normal">0</mml:mn><mml:mo>,</mml:mo><mml:mn mathvariant="normal">500</mml:mn><mml:mo>]</mml:mo></mml:mrow></mml:math></inline-formula> respectively.</p>

<table-wrap id="T1" specific-use="star"><label>Table 1</label><caption><p id="d2e3513">Summary of numerical experiments investigating the impact of observational data choice on Bayesian calibration outcomes. Experiments 1–6 calibrate friction coefficients (<inline-formula><mml:math id="M159" display="inline"><mml:mi mathvariant="italic">μ</mml:mi></mml:math></inline-formula>, <inline-formula><mml:math id="M160" display="inline"><mml:mi mathvariant="italic">ξ</mml:mi></mml:math></inline-formula>) using different observation types, while experiments 7–8 additionally calibrate discrepancy parameters (<inline-formula><mml:math id="M161" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">σ</mml:mi><mml:mrow><mml:msub><mml:mi mathvariant="normal">vel</mml:mi><mml:mi mathvariant="normal">TS</mml:mi></mml:msub></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula>, <inline-formula><mml:math id="M162" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">σ</mml:mi><mml:mrow><mml:msub><mml:mi mathvariant="normal">pos</mml:mi><mml:mi mathvariant="normal">TS</mml:mi></mml:msub></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula>). Apart from experiments 5–6, all other experiments involve single calibration routines. All experiments use synthetic observations generated by adding random noise drawn from a Gaussian distribution with zero mean and specified standard deviations.</p></caption><oasis:table frame="topbot"><oasis:tgroup cols="5">
     <oasis:colspec colnum="1" colname="col1" align="left"/>
     <oasis:colspec colnum="2" colname="col2" align="right"/>
     <oasis:colspec colnum="3" colname="col3" align="left"/>
     <oasis:colspec colnum="4" colname="col4" align="left"/>
     <oasis:colspec colnum="5" colname="col5" align="right"/>
     <oasis:thead>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">Exp. No.</oasis:entry>
         <oasis:entry colname="col2">No. of Calibrations</oasis:entry>
         <oasis:entry colname="col3">Calibrated Parameters</oasis:entry>
         <oasis:entry colname="col4">Observation</oasis:entry>
         <oasis:entry colname="col5">Std. Dev. of Applied Noise</oasis:entry>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row>
         <oasis:entry colname="col1">1</oasis:entry>
         <oasis:entry colname="col2">1</oasis:entry>
         <oasis:entry colname="col3"><inline-formula><mml:math id="M163" display="inline"><mml:mrow><mml:mi mathvariant="italic">μ</mml:mi><mml:mo>,</mml:mo><mml:mi mathvariant="italic">ξ</mml:mi></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col4">Maximum velocity <inline-formula><mml:math id="M164" display="inline"><mml:mrow><mml:msub><mml:mi>U</mml:mi><mml:mtext>max</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col5">2.08</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">2</oasis:entry>
         <oasis:entry colname="col2">1</oasis:entry>
         <oasis:entry colname="col3"><inline-formula><mml:math id="M165" display="inline"><mml:mrow><mml:mi mathvariant="italic">μ</mml:mi><mml:mo>,</mml:mo><mml:mi mathvariant="italic">ξ</mml:mi></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col4">runout distance <inline-formula><mml:math id="M166" display="inline"><mml:mrow><mml:msub><mml:mi>X</mml:mi><mml:mtext>end</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col5">225</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">3</oasis:entry>
         <oasis:entry colname="col2">1</oasis:entry>
         <oasis:entry colname="col3"><inline-formula><mml:math id="M167" display="inline"><mml:mrow><mml:mi mathvariant="italic">μ</mml:mi><mml:mo>,</mml:mo><mml:mi mathvariant="italic">ξ</mml:mi></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col4">Velocity time series <inline-formula><mml:math id="M168" display="inline"><mml:mrow><mml:mi>u</mml:mi><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col5">2.08</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">4</oasis:entry>
         <oasis:entry colname="col2">1</oasis:entry>
         <oasis:entry colname="col3"><inline-formula><mml:math id="M169" display="inline"><mml:mrow><mml:mi mathvariant="italic">μ</mml:mi><mml:mo>,</mml:mo><mml:mi mathvariant="italic">ξ</mml:mi></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col4">Position time series <inline-formula><mml:math id="M170" display="inline"><mml:mrow><mml:mi>x</mml:mi><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col5">225</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">5</oasis:entry>
         <oasis:entry colname="col2">100</oasis:entry>
         <oasis:entry colname="col3"><inline-formula><mml:math id="M171" display="inline"><mml:mrow><mml:mi mathvariant="italic">μ</mml:mi><mml:mo>,</mml:mo><mml:mi mathvariant="italic">ξ</mml:mi></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col4">Multiple velocity time series, varying lengths</oasis:entry>
         <oasis:entry colname="col5">2.08</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">6</oasis:entry>
         <oasis:entry colname="col2">10</oasis:entry>
         <oasis:entry colname="col3"><inline-formula><mml:math id="M172" display="inline"><mml:mrow><mml:mi mathvariant="italic">μ</mml:mi><mml:mo>,</mml:mo><mml:mi mathvariant="italic">ξ</mml:mi></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col4">Multiple velocity time series, varying frequencies</oasis:entry>
         <oasis:entry colname="col5">2.08</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">7</oasis:entry>
         <oasis:entry colname="col2">1</oasis:entry>
         <oasis:entry colname="col3"><inline-formula><mml:math id="M173" display="inline"><mml:mrow><mml:mi mathvariant="italic">μ</mml:mi><mml:mo>,</mml:mo><mml:mi mathvariant="italic">ξ</mml:mi><mml:mo>,</mml:mo><mml:msub><mml:mi mathvariant="italic">σ</mml:mi><mml:mrow><mml:msub><mml:mi mathvariant="normal">vel</mml:mi><mml:mi mathvariant="normal">TS</mml:mi></mml:msub></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col4">Velocity time series <inline-formula><mml:math id="M174" display="inline"><mml:mrow><mml:mi>u</mml:mi><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col5">2.08</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">8</oasis:entry>
         <oasis:entry colname="col2">1</oasis:entry>
         <oasis:entry colname="col3"><inline-formula><mml:math id="M175" display="inline"><mml:mrow><mml:mi mathvariant="italic">μ</mml:mi><mml:mo>,</mml:mo><mml:mi mathvariant="italic">ξ</mml:mi><mml:mo>,</mml:mo><mml:msub><mml:mi mathvariant="italic">σ</mml:mi><mml:mrow><mml:msub><mml:mi mathvariant="normal">pos</mml:mi><mml:mi mathvariant="normal">TS</mml:mi></mml:msub></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col4">Position time series <inline-formula><mml:math id="M176" display="inline"><mml:mrow><mml:mi>x</mml:mi><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col5">225</oasis:entry>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup></oasis:table></table-wrap>


</sec>
</sec>
<sec id="Ch1.S4">
  <label>4</label><title>Results</title>
      <p id="d2e3930">This section includes results for the curated set of numerical experiments discussed in Sect. <xref ref-type="sec" rid="Ch1.S3.SS3"/>. We adopt the workflow presented in Sect. <xref ref-type="sec" rid="Ch1.S2.SS4"/> and the associated data: (i) Training dataset including set of sampled parameters and the corresponding model outputs; (ii) Ground-truth data used to generate the synthetic observations detailed in Sect. <xref ref-type="sec" rid="Ch1.S3.SS2"/> are archived in this repository <ext-link xlink:href="https://doi.org/10.5281/zenodo.17120721" ext-link-type="DOI">10.5281/zenodo.17120721</ext-link> <xref ref-type="bibr" rid="bib1.bibx32" id="paren.62"/>.</p>
<sec id="Ch1.S4.SS1">
  <label>4.1</label><title>Experiment 1: Calibration with maximum velocity as observation</title>
      <p id="d2e3952">Figure <xref ref-type="fig" rid="F7"/> illustrates the prior distribution of the friction parameters and their corresponding posterior distributions obtained by performing the calibration using the maximum velocity (<inline-formula><mml:math id="M177" display="inline"><mml:mrow><mml:msub><mml:mi>U</mml:mi><mml:mi mathvariant="normal">max</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>) as observation. The maximum a posteriori (MAP) estimates, indicated by the black dashed lines in Fig. <xref ref-type="fig" rid="F7"/>a and b, provide the most probable values for the parameters. The MAP estimate for <inline-formula><mml:math id="M178" display="inline"><mml:mi mathvariant="italic">μ</mml:mi></mml:math></inline-formula> is 0.284, while for <inline-formula><mml:math id="M179" display="inline"><mml:mi mathvariant="italic">ξ</mml:mi></mml:math></inline-formula> it is 1187. The uncertainty associated with these estimates is summarized by highest density intervals (HDI), listed in the Table <xref ref-type="table" rid="T2"/>. The parameter values within 95 % HDI have a higher probability than those outside this interval. Thus, a narrower HDI, as observed for <inline-formula><mml:math id="M180" display="inline"><mml:mi mathvariant="italic">ξ</mml:mi></mml:math></inline-formula>, indicates a lower degree of uncertainty. This suggests that the maximum velocity provided greater information for <inline-formula><mml:math id="M181" display="inline"><mml:mi mathvariant="italic">ξ</mml:mi></mml:math></inline-formula> during calibration. This is also highlighted by the difference in KL divergence values (refer Table <xref ref-type="table" rid="T2"/>), which had a higher value for <inline-formula><mml:math id="M182" display="inline"><mml:mi mathvariant="italic">ξ</mml:mi></mml:math></inline-formula>.</p>

      <fig id="F7" specific-use="star"><label>Figure 7</label><caption><p id="d2e4012">Posterior and prior distributions of <bold>(a)</bold> Dry Coulomb friction coefficient (<inline-formula><mml:math id="M183" display="inline"><mml:mi mathvariant="italic">μ</mml:mi></mml:math></inline-formula>) and <bold>(b)</bold>  Turbulent friction coefficient (<inline-formula><mml:math id="M184" display="inline"><mml:mi mathvariant="italic">ξ</mml:mi></mml:math></inline-formula>) based on calibration with maximum velocity (<inline-formula><mml:math id="M185" display="inline"><mml:mrow><mml:msub><mml:mi>U</mml:mi><mml:mtext>max</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula>) as observation. Maximum velocity conveyed more information about <inline-formula><mml:math id="M186" display="inline"><mml:mi mathvariant="italic">ξ</mml:mi></mml:math></inline-formula> during calibration, as reflected in the narrower posterior distribution.</p></caption>
          <graphic xlink:href="https://npg.copernicus.org/articles/33/425/2026/npg-33-425-2026-f07.png"/>

        </fig>

<table-wrap id="T2" specific-use="star"><label>Table 2</label><caption><p id="d2e4063">Kullback-Leibler Divergence and 95 % Highest Density Intervals for friction coefficients calibrated using multiple datasets.</p></caption><oasis:table frame="topbot"><oasis:tgroup cols="5">
     <oasis:colspec colnum="1" colname="col1" align="left"/>
     <oasis:colspec colnum="2" colname="col2" align="right"/>
     <oasis:colspec colnum="3" colname="col3" align="right" colsep="1"/>
     <oasis:colspec colnum="4" colname="col4" align="right"/>
     <oasis:colspec colnum="5" colname="col5" align="right"/>
     <oasis:thead>
       <oasis:row>
         <oasis:entry colname="col1">Observations</oasis:entry>
         <oasis:entry namest="col2" nameend="col3" align="center" colsep="1">Dry Coulomb Friction </oasis:entry>
         <oasis:entry namest="col4" nameend="col5" align="center">Turbulent Friction </oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1"/>
         <oasis:entry rowsep="1" namest="col2" nameend="col3" align="center" colsep="1">coefficient (<inline-formula><mml:math id="M187" display="inline"><mml:mi mathvariant="italic">μ</mml:mi></mml:math></inline-formula>) </oasis:entry>
         <oasis:entry rowsep="1" namest="col4" nameend="col5" align="center">coefficient (<inline-formula><mml:math id="M188" display="inline"><mml:mi mathvariant="italic">ξ</mml:mi></mml:math></inline-formula>) </oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2">KL Divergence</oasis:entry>
         <oasis:entry colname="col3">95 % HDI</oasis:entry>
         <oasis:entry colname="col4">KL Divergence</oasis:entry>
         <oasis:entry colname="col5">95 % HDI</oasis:entry>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row>
         <oasis:entry colname="col1">Maximum Velocity</oasis:entry>
         <oasis:entry colname="col2">0.06</oasis:entry>
         <oasis:entry colname="col3"><inline-formula><mml:math id="M189" display="inline"><mml:mrow><mml:mo>[</mml:mo><mml:mn mathvariant="normal">0.04</mml:mn><mml:mo>,</mml:mo><mml:mn mathvariant="normal">0.3</mml:mn><mml:mo>]</mml:mo></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col4">0.63</oasis:entry>
         <oasis:entry colname="col5"><inline-formula><mml:math id="M190" display="inline"><mml:mrow><mml:mo>[</mml:mo><mml:mn mathvariant="normal">768</mml:mn><mml:mo>,</mml:mo><mml:mn mathvariant="normal">1876</mml:mn><mml:mo>]</mml:mo></mml:mrow></mml:math></inline-formula></oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">runout distance</oasis:entry>
         <oasis:entry colname="col2">0.43</oasis:entry>
         <oasis:entry colname="col3"><inline-formula><mml:math id="M191" display="inline"><mml:mrow><mml:mo>[</mml:mo><mml:mn mathvariant="normal">0.11</mml:mn><mml:mo>,</mml:mo><mml:mn mathvariant="normal">0.3</mml:mn><mml:mo>]</mml:mo></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col4">0.05</oasis:entry>
         <oasis:entry colname="col5"><inline-formula><mml:math id="M192" display="inline"><mml:mrow><mml:mo>[</mml:mo><mml:mn mathvariant="normal">100</mml:mn><mml:mo>,</mml:mo><mml:mn mathvariant="normal">2080</mml:mn><mml:mo>]</mml:mo></mml:mrow></mml:math></inline-formula></oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Velocity Time Series</oasis:entry>
         <oasis:entry colname="col2">2.69</oasis:entry>
         <oasis:entry colname="col3"><inline-formula><mml:math id="M193" display="inline"><mml:mrow><mml:mo>[</mml:mo><mml:mn mathvariant="normal">0.22</mml:mn><mml:mo>,</mml:mo><mml:mn mathvariant="normal">0.23</mml:mn><mml:mo>]</mml:mo></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col4">3.59</oasis:entry>
         <oasis:entry colname="col5"><inline-formula><mml:math id="M194" display="inline"><mml:mrow><mml:mo>[</mml:mo><mml:mn mathvariant="normal">971</mml:mn><mml:mo>,</mml:mo><mml:mn mathvariant="normal">1028</mml:mn><mml:mo>]</mml:mo></mml:mrow></mml:math></inline-formula></oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Position Time Series</oasis:entry>
         <oasis:entry colname="col2">0.84</oasis:entry>
         <oasis:entry colname="col3"><inline-formula><mml:math id="M195" display="inline"><mml:mrow><mml:mo>[</mml:mo><mml:mn mathvariant="normal">0.14</mml:mn><mml:mo>,</mml:mo><mml:mn mathvariant="normal">0.26</mml:mn><mml:mo>]</mml:mo></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col4">1.89</oasis:entry>
         <oasis:entry colname="col5"><inline-formula><mml:math id="M196" display="inline"><mml:mrow><mml:mo>[</mml:mo><mml:mn mathvariant="normal">792</mml:mn><mml:mo>,</mml:mo><mml:mn mathvariant="normal">1105</mml:mn><mml:mo>]</mml:mo></mml:mrow></mml:math></inline-formula></oasis:entry>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup></oasis:table></table-wrap>

</sec>
<sec id="Ch1.S4.SS2">
  <label>4.2</label><title>Experiment 2: Calibration with runout distance as observation</title>
      <p id="d2e4334">We repeated the calibration of the friction parameters using the runout distance as observation, and the results are depicted in Fig. <xref ref-type="fig" rid="F8"/>. This calibration yielded a MAP estimate for <inline-formula><mml:math id="M197" display="inline"><mml:mi mathvariant="italic">μ</mml:mi></mml:math></inline-formula>  of 0.277, nearly identical to the MAP value presented in Fig. <xref ref-type="fig" rid="F7"/>a. However, the 95 % HDI listed in Table <xref ref-type="table" rid="T2"/>, is considerably shorter, suggesting greater confidence in the estimate. In contrast, the 95 % HDI interval for <inline-formula><mml:math id="M198" display="inline"><mml:mi mathvariant="italic">ξ</mml:mi></mml:math></inline-formula> has widened, and even the MAP estimate (shown in Fig. <xref ref-type="fig" rid="F8"/>b) of 159 significantly differs from the true value of 1000. This indicates that more information was gained for <inline-formula><mml:math id="M199" display="inline"><mml:mi mathvariant="italic">μ</mml:mi></mml:math></inline-formula> when calibrated with the runout distance. A higher KL divergence value of 0.43 for <inline-formula><mml:math id="M200" display="inline"><mml:mi mathvariant="italic">μ</mml:mi></mml:math></inline-formula> compared to 0.05 for <inline-formula><mml:math id="M201" display="inline"><mml:mi mathvariant="italic">ξ</mml:mi></mml:math></inline-formula> further corroborates this assertion, refer Table <xref ref-type="table" rid="T2"/>.</p>

      <fig id="F8" specific-use="star"><label>Figure 8</label><caption><p id="d2e4385">Posterior and prior distributions of <bold>(a)</bold> Dry Coulomb friction coefficient (<inline-formula><mml:math id="M202" display="inline"><mml:mi mathvariant="italic">μ</mml:mi></mml:math></inline-formula>) and <bold>(b)</bold>  Turbulent friction coefficient (<inline-formula><mml:math id="M203" display="inline"><mml:mi mathvariant="italic">ξ</mml:mi></mml:math></inline-formula>) based on a calibration with runout distance (<inline-formula><mml:math id="M204" display="inline"><mml:mrow><mml:msub><mml:mi>X</mml:mi><mml:mtext>end</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula>) as observation. Greater contraction of the posterior of <inline-formula><mml:math id="M205" display="inline"><mml:mi mathvariant="italic">μ</mml:mi></mml:math></inline-formula> indicates higher information gain during calibration with the runout distance.</p></caption>
          <graphic xlink:href="https://npg.copernicus.org/articles/33/425/2026/npg-33-425-2026-f08.png"/>

        </fig>

</sec>
<sec id="Ch1.S4.SS3">
  <label>4.3</label><title>Experiment 3: Calibration with velocity time series as observation</title>
      <p id="d2e4441">A third calibration was conducted using the velocity time series <inline-formula><mml:math id="M206" display="inline"><mml:mrow><mml:mi>u</mml:mi><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>, and the parameter distributions are presented in Fig. <xref ref-type="fig" rid="F9"/>. Unlike the results of calibration with aggregated data (e.g., maximum velocity and runout distance), we see a significant information gain for both <inline-formula><mml:math id="M207" display="inline"><mml:mi mathvariant="italic">μ</mml:mi></mml:math></inline-formula> and <inline-formula><mml:math id="M208" display="inline"><mml:mi mathvariant="italic">ξ</mml:mi></mml:math></inline-formula>, reflected by the much smaller 95 % HDI listed in Table <xref ref-type="table" rid="T2"/>. Even the KL divergence values outlined in Table <xref ref-type="table" rid="T2"/> are significantly higher than those corresponding to the aggregated data. Furthermore, the MAP estimates for both parameters are nearly identical to the true values.</p>

      <fig id="F9" specific-use="star"><label>Figure 9</label><caption><p id="d2e4481">Posterior and prior distributions of <bold>(a)</bold> Dry Coulomb friction coefficient (<inline-formula><mml:math id="M209" display="inline"><mml:mi mathvariant="italic">μ</mml:mi></mml:math></inline-formula>) and <bold>(b)</bold>  Turbulent friction coefficient (<inline-formula><mml:math id="M210" display="inline"><mml:mi mathvariant="italic">ξ</mml:mi></mml:math></inline-formula>) based on a calibration with velocity time series as observation. Velocity time series provides considerably greater information about both <inline-formula><mml:math id="M211" display="inline"><mml:mi mathvariant="italic">μ</mml:mi></mml:math></inline-formula> and <inline-formula><mml:math id="M212" display="inline"><mml:mi mathvariant="italic">ξ</mml:mi></mml:math></inline-formula> during calibration, as evidenced by the narrower posterior distributions.</p></caption>
          <graphic xlink:href="https://npg.copernicus.org/articles/33/425/2026/npg-33-425-2026-f09.png"/>

        </fig>

</sec>
<sec id="Ch1.S4.SS4">
  <label>4.4</label><title>Experiment 4: Calibration with position time series as observation</title>
      <p id="d2e4534">Figure <xref ref-type="fig" rid="F10"/> shows the parameter distributions, calibrated with position time series, <inline-formula><mml:math id="M213" display="inline"><mml:mrow><mml:mi>x</mml:mi><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>. Again, we observed information gain for both parameters, reflected in the contracted posterior distributions,  illustrated in Fig. <xref ref-type="fig" rid="F10"/>a and b. MAP estimates for <inline-formula><mml:math id="M214" display="inline"><mml:mi mathvariant="italic">μ</mml:mi></mml:math></inline-formula> and <inline-formula><mml:math id="M215" display="inline"><mml:mi mathvariant="italic">ξ</mml:mi></mml:math></inline-formula> are also closer to the true values compared to the estimates in Figs. <xref ref-type="fig" rid="F7"/> and <xref ref-type="fig" rid="F8"/>. However, this information gain is lower than that observed in the calibration using velocity time series (<inline-formula><mml:math id="M216" display="inline"><mml:mrow><mml:mi>u</mml:mi><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>), as indicated by the lower KL divergence values of 0.84 for <inline-formula><mml:math id="M217" display="inline"><mml:mi mathvariant="italic">μ</mml:mi></mml:math></inline-formula> and 1.89 for <inline-formula><mml:math id="M218" display="inline"><mml:mi mathvariant="italic">ξ</mml:mi></mml:math></inline-formula>, compared to 2.69 and 3.59 (as shown in Table <xref ref-type="table" rid="T2"/>).</p>

      <fig id="F10" specific-use="star"><label>Figure 10</label><caption><p id="d2e4606">Posterior and prior distributions of <bold>(a)</bold> Dry Coulomb friction coefficient (<inline-formula><mml:math id="M219" display="inline"><mml:mi mathvariant="italic">μ</mml:mi></mml:math></inline-formula>) and <bold>(b)</bold> Turbulent friction coefficient (<inline-formula><mml:math id="M220" display="inline"><mml:mi mathvariant="italic">ξ</mml:mi></mml:math></inline-formula>) based on calibration with position time series as observation. Similar to velocity time series, position time series also provides information about both parameters, albeit to a lesser extent.</p></caption>
          <graphic xlink:href="https://npg.copernicus.org/articles/33/425/2026/npg-33-425-2026-f10.png"/>

        </fig>

</sec>
<sec id="Ch1.S4.SS5">
  <label>4.5</label><title>Experiment 5: Value of information in data: Length of velocity time series data</title>
      <p id="d2e4643">We calibrated the friction parameters with 100 different datasets of velocity time series, each containing a varying number of time steps. Seven selected posterior distribution of <inline-formula><mml:math id="M221" display="inline"><mml:mi mathvariant="italic">ξ</mml:mi></mml:math></inline-formula> from these calibrations is depicted in Fig. <xref ref-type="fig" rid="F11"/>, with the <inline-formula><mml:math id="M222" display="inline"><mml:mi>y</mml:mi></mml:math></inline-formula>-axis listing the number of velocity time steps used in the calibration. Calibrating with a higher number of time steps led to greater information gain, reflected in the contracting posteriors with increasing time step count. However, from the Fig. <xref ref-type="fig" rid="F11"/>, we can see that the difference between the posteriors in the initial time steps is considerably greater than the posteriors in the later time steps. This suggests that the rate of information gain decreases as we use more time steps for calibration.</p>

      <fig id="F11" specific-use="star"><label>Figure 11</label><caption><p id="d2e4666">Posterior distributions of turbulent friction coefficient <inline-formula><mml:math id="M223" display="inline"><mml:mi mathvariant="italic">ξ</mml:mi></mml:math></inline-formula> calibrated with a varying number of velocity time steps. The rate of information gain is more profound in the initial time steps than the later.</p></caption>
          <graphic xlink:href="https://npg.copernicus.org/articles/33/425/2026/npg-33-425-2026-f11.png"/>

        </fig>

      <fig id="F12" specific-use="star"><label>Figure 12</label><caption><p id="d2e4684"><bold>(a)</bold> Variation of Kullback-Leibler divergence of the posterior and prior distributions of turbulent friction coefficient <inline-formula><mml:math id="M224" display="inline"><mml:mi mathvariant="italic">ξ</mml:mi></mml:math></inline-formula> with velocity time steps. <bold>(b)</bold> Variation of velocity and acceleration with time steps. </p></caption>
          <graphic xlink:href="https://npg.copernicus.org/articles/33/425/2026/npg-33-425-2026-f12.png"/>

        </fig>

      <p id="d2e4706">Figure <xref ref-type="fig" rid="F12"/>a, plots the variation of the KL divergence of the posterior and prior distributions of <inline-formula><mml:math id="M225" display="inline"><mml:mi mathvariant="italic">ξ</mml:mi></mml:math></inline-formula> with the number of velocity time steps used for calibration. Increasing number of time steps resulted in higher KL divergence, indicating a positive correlation between the information gain and the number of time steps. However, the initial steep slope of the plot in Fig. <xref ref-type="fig" rid="F12"/>a and its subsequent plateauing point to a diminishing return effect, where additional time steps beyond a critical threshold yield progressively smaller information gain. Interestingly, this threshold corresponds to the time step at which velocity attains its peak, as seen in Fig. <xref ref-type="fig" rid="F12"/>b.</p>
</sec>
<sec id="Ch1.S4.SS6">
  <label>4.6</label><title>Experiment 6: Value of information in data: Temporal resolution of velocity time series data</title>
      <p id="d2e4730">Figure <xref ref-type="fig" rid="F13"/> illustrates the posterior distributions of <inline-formula><mml:math id="M226" display="inline"><mml:mi mathvariant="italic">ξ</mml:mi></mml:math></inline-formula> calibrated using velocity time series datasets of varying temporal resolutions. Each posterior distribution is associated with a specific time step size indicated on the <inline-formula><mml:math id="M227" display="inline"><mml:mi>y</mml:mi></mml:math></inline-formula>-axis. For example, a time step size of <inline-formula><mml:math id="M228" display="inline"><mml:mn mathvariant="normal">4</mml:mn></mml:math></inline-formula> on the <inline-formula><mml:math id="M229" display="inline"><mml:mi>y</mml:mi></mml:math></inline-formula>-axis indicates that the corresponding posterior was generated using a velocity time series of time step size <inline-formula><mml:math id="M230" display="inline"><mml:mn mathvariant="normal">4</mml:mn></mml:math></inline-formula>. As the resolution of the time series data increases (i.e., with a smaller time step size), the information gain increases, as evidenced by the contracting posteriors.</p>

      <fig id="F13" specific-use="star"><label>Figure 13</label><caption><p id="d2e4773">Posterior distributions of turbulent friction coefficient <inline-formula><mml:math id="M231" display="inline"><mml:mi mathvariant="italic">ξ</mml:mi></mml:math></inline-formula> calibrated with a varying temporal resolution of the velocity data. The information gained during calibration increases with temporal resolution.</p></caption>
          <graphic xlink:href="https://npg.copernicus.org/articles/33/425/2026/npg-33-425-2026-f13.png"/>

        </fig>


</sec>
<sec id="Ch1.S4.SS7">
  <label>4.7</label><title>Experiment 7 and 8: Calibration of the discrepancy parameters</title>
      <p id="d2e4799">Figures <xref ref-type="fig" rid="F7"/>, <xref ref-type="fig" rid="F8"/>, <xref ref-type="fig" rid="F9"/>, and <xref ref-type="fig" rid="F10"/>, presented posterior distributions of friction parameters calibrated using multiple datasets listed in Table <xref ref-type="table" rid="T2"/>. During these calibrations, we made heuristic assumptions for the discrepancy parameter <inline-formula><mml:math id="M232" display="inline"><mml:mi mathvariant="italic">σ</mml:mi></mml:math></inline-formula> of the noise model, refer to Sect. <xref ref-type="sec" rid="Ch1.S2.SS1"/> for detailed discussion. However, we can include this parameter in the calibration routine and calibrate it along with the friction parameters. Figure <xref ref-type="fig" rid="F14"/>, depicts the prior and posterior distributions of discrepancy parameters corresponding to velocity and position time series. In Fig. <xref ref-type="fig" rid="F14"/>a, we see that using velocity time series <inline-formula><mml:math id="M233" display="inline"><mml:mrow><mml:mi>u</mml:mi><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>, we can learn extensively for the velocity discrepancy parameter <inline-formula><mml:math id="M234" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">σ</mml:mi><mml:mrow><mml:msub><mml:mi mathvariant="normal">vel</mml:mi><mml:mi mathvariant="normal">TS</mml:mi></mml:msub></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula>, indicated by the highly contracted posterior. Similarly, the position time series provides substantial information regarding the position discrepancy parameter <inline-formula><mml:math id="M235" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">σ</mml:mi><mml:mrow><mml:msub><mml:mi mathvariant="normal">pos</mml:mi><mml:mi mathvariant="normal">TS</mml:mi></mml:msub></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula>, refer Fig. <xref ref-type="fig" rid="F14"/>b.</p>

      <fig id="F14" specific-use="star"><label>Figure 14</label><caption><p id="d2e4875">Posterior and prior distributions of <bold>(a)</bold> Velocity discrepancy parameter (<inline-formula><mml:math id="M236" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">σ</mml:mi><mml:mrow><mml:msub><mml:mi mathvariant="normal">vel</mml:mi><mml:mi mathvariant="normal">TS</mml:mi></mml:msub></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula>) and <bold>(b)</bold> Position discrepancy parameter (<inline-formula><mml:math id="M237" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">σ</mml:mi><mml:mrow><mml:msub><mml:mi mathvariant="normal">pos</mml:mi><mml:mi mathvariant="normal">TS</mml:mi></mml:msub></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula>) based on a calibration with velocity and position time series respectively.</p></caption>
          <graphic xlink:href="https://npg.copernicus.org/articles/33/425/2026/npg-33-425-2026-f14.png"/>

        </fig>

</sec>
</sec>
<sec id="Ch1.S5">
  <label>5</label><title>Discussions</title>
      <p id="d2e4930">In our numerical experiments, we investigated the impact of the scope and quantity of the data used as observations in the Bayesian parameter calibration of landslide runout models. To this end, we calibrated the friction parameters of a lumped mass model, namely the dry Coulomb friction coefficient (<inline-formula><mml:math id="M238" display="inline"><mml:mi mathvariant="italic">μ</mml:mi></mml:math></inline-formula>) and the turbulent friction coefficient (<inline-formula><mml:math id="M239" display="inline"><mml:mi mathvariant="italic">ξ</mml:mi></mml:math></inline-formula>), with a diverse set of observational data summarized in Table <xref ref-type="table" rid="T1"/>. Using our novel Bayesian data selection workflow, we quantified the information gained during the calibration by means of a statistical distance measure from information theory, referred to as KL divergence <xref ref-type="bibr" rid="bib1.bibx30" id="paren.63"/>. Comparison of KL divergence associated with posterior distributions of alternative observation data (Figs. <xref ref-type="fig" rid="F7"/>, <xref ref-type="fig" rid="F8"/>, <xref ref-type="fig" rid="F9"/>, and <xref ref-type="fig" rid="F10"/>) highlights the critical role of data selection in the calibration of these parameters. In our experiment, we found that calibration of a lumped mass model using maximum velocity and runout distance revealed contrasting trends: the maximum velocity provided substantial information for <inline-formula><mml:math id="M240" display="inline"><mml:mi mathvariant="italic">ξ</mml:mi></mml:math></inline-formula> while minimally contributing to <inline-formula><mml:math id="M241" display="inline"><mml:mi mathvariant="italic">μ</mml:mi></mml:math></inline-formula>. In contrast, calibration based on the runout distance provided a greater constraint, hence more information, for <inline-formula><mml:math id="M242" display="inline"><mml:mi mathvariant="italic">μ</mml:mi></mml:math></inline-formula>, but minimally contributed to gaining a better understanding of <inline-formula><mml:math id="M243" display="inline"><mml:mi mathvariant="italic">ξ</mml:mi></mml:math></inline-formula>. These contrasting trends can be explained by examining the variations of these parameters with maximum velocity and runout distance, respectively.  Figure <xref ref-type="fig" rid="F15"/> shows a clear dependence of the runout distance on <inline-formula><mml:math id="M244" display="inline"><mml:mi mathvariant="italic">μ</mml:mi></mml:math></inline-formula>, which decreases with increasing <inline-formula><mml:math id="M245" display="inline"><mml:mi mathvariant="italic">μ</mml:mi></mml:math></inline-formula> while remaining unaffected by changes in <inline-formula><mml:math id="M246" display="inline"><mml:mi mathvariant="italic">ξ</mml:mi></mml:math></inline-formula>. In contrast, the maximum velocity varies significantly with <inline-formula><mml:math id="M247" display="inline"><mml:mi mathvariant="italic">ξ</mml:mi></mml:math></inline-formula> but shows little to no change with <inline-formula><mml:math id="M248" display="inline"><mml:mi mathvariant="italic">μ</mml:mi></mml:math></inline-formula>. The narrower posterior distributions for <inline-formula><mml:math id="M249" display="inline"><mml:mi mathvariant="italic">ξ</mml:mi></mml:math></inline-formula> when calibrating with maximum velocity and for <inline-formula><mml:math id="M250" display="inline"><mml:mi mathvariant="italic">μ</mml:mi></mml:math></inline-formula> when calibrating with the runout distance highlight the complementary strengths of these datasets for Bayesian parameter inference. This behavior is consistent with earlier findings of <xref ref-type="bibr" rid="bib1.bibx35" id="text.64"/>, who noted that <inline-formula><mml:math id="M251" display="inline"><mml:mi mathvariant="italic">μ</mml:mi></mml:math></inline-formula> and <inline-formula><mml:math id="M252" display="inline"><mml:mi mathvariant="italic">ξ</mml:mi></mml:math></inline-formula> influence different aspects of flow behavior. Specifically, <inline-formula><mml:math id="M253" display="inline"><mml:mi mathvariant="italic">μ</mml:mi></mml:math></inline-formula> is associated with the runout distance, while <inline-formula><mml:math id="M254" display="inline"><mml:mi mathvariant="italic">ξ</mml:mi></mml:math></inline-formula> limits the flow velocities. Our results extend this qualitative understanding by providing a methodology to quantify these relationships using the information gained during Bayesian calibration. These outcomes point to the hidden potential of a systematic data selection regarding its anticipated value-add with the critical process governed by the parameter of interest.</p>

      <fig id="F15" specific-use="star"><label>Figure 15</label><caption><p id="d2e5076">Variation of the friction coefficients (<inline-formula><mml:math id="M255" display="inline"><mml:mi mathvariant="italic">μ</mml:mi></mml:math></inline-formula> and <inline-formula><mml:math id="M256" display="inline"><mml:mi mathvariant="italic">ξ</mml:mi></mml:math></inline-formula>) with maximum velocity (<inline-formula><mml:math id="M257" display="inline"><mml:mrow><mml:msub><mml:mi>U</mml:mi><mml:mi mathvariant="normal">max</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>) and run out distance (<inline-formula><mml:math id="M258" display="inline"><mml:mrow><mml:msub><mml:mi>X</mml:mi><mml:mtext>end</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula>). The dry Coulomb friction coefficient <inline-formula><mml:math id="M259" display="inline"><mml:mi mathvariant="italic">μ</mml:mi></mml:math></inline-formula> primarily varies with <inline-formula><mml:math id="M260" display="inline"><mml:mrow><mml:msub><mml:mi>X</mml:mi><mml:mtext>end</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula>, while the turbulent friction coefficient <inline-formula><mml:math id="M261" display="inline"><mml:mi mathvariant="italic">ξ</mml:mi></mml:math></inline-formula> is more strongly influenced by <inline-formula><mml:math id="M262" display="inline"><mml:mrow><mml:msub><mml:mi>U</mml:mi><mml:mi mathvariant="normal">max</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>.</p></caption>
        <graphic xlink:href="https://npg.copernicus.org/articles/33/425/2026/npg-33-425-2026-f15.png"/>

      </fig>

      <p id="d2e5158">In numerical experiments 1 and 2, aggregated data proved inadequate in the joint calibration of the parameters – they could infer only one parameter each – we, therefore, explored time series data as an alternative. We calibrated the parameters using velocity and position time series data, each depicting the time history of velocity and position of the sliding mass, respectively. Figures <xref ref-type="fig" rid="F9"/> and <xref ref-type="fig" rid="F10"/> illustrate that the time series data provided substantial information during the calibration of both parameters, as evidenced by the highly contracted posteriors. This higher information gain implies that time series data are better equipped to calibrate friction parameters than aggregated data. Our findings are consistent with the work of <xref ref-type="bibr" rid="bib1.bibx37" id="text.65"/>, who observed that the time history data offered a better calibration of the landslide parameters than the static data. Specifically, the force-time history derived from seismic records was more adept at constraining landslide parameters than runout distance and deposit area because it captures the temporal evolution of landslide dynamics. In the same way, the velocity and position time series data used in our study captured the evolving dynamics of landslides more efficiently than aggregated data. A similar observation was made by <xref ref-type="bibr" rid="bib1.bibx54" id="text.66"/>, who emphasized the limitations of static data in the adequate inversion of the landslide characteristics. Additionally, the friction parameters we want to calibrate govern the acceleration and deceleration phases of the landslide motion, which are better reflected in the time series data <xref ref-type="bibr" rid="bib1.bibx37" id="paren.67"/>. These findings further suggest that aligning observations with the critical process governed by the parameter of interest improves the calibration performance <xref ref-type="bibr" rid="bib1.bibx27" id="paren.68"/>.</p>
      <p id="d2e5179">The superior performance of time series data compared to aggregated data implies a positive correlation between data quantity and calibration performance, where data quantity refers to the length and resolution of the time series. From Fig. <xref ref-type="fig" rid="F11"/>, we can see that an increase in the length of time series leads to enhancement in calibration indicated by the contracting posteriors. Similarly, the calibration improves when we increase the temporal resolution, evidenced by posterior contraction with decreasing time step size, as shown in Fig. <xref ref-type="fig" rid="F13"/>. These observations further support the claim that as the data quantity used for calibration increases, the calibration performance enhances accordingly. These observations echo findings in the field of hydrology, which highlight that increasing the data series length enhances the reliability of the calibration of a hydrological model <xref ref-type="bibr" rid="bib1.bibx9 bib1.bibx33" id="paren.69"/>. <xref ref-type="bibr" rid="bib1.bibx27" id="text.70"/> reported similar findings; they compared the impact of the temporal resolution of the data on the calibration of the hydrological model parameters. They found that high-resolution data better captured parameters associated with fast hydrological processes, as these finer-scale data preserve the dynamics that are lost in coarser resolutions because of data averaging. However, more data does not necessarily lead to better calibration; for instance, <xref ref-type="bibr" rid="bib1.bibx10" id="text.71"/> found that increasing the length of calibration data series beyond 10 years did not improve the validation performance of the hydrological model. Similarly, <xref ref-type="bibr" rid="bib1.bibx11" id="text.72"/> observed that the ability of the data to capture critical processes related to the parameters we want to calibrate was more critical than the temporal resolution of the data.</p>
      <p id="d2e5199">As the practical availability of data is often limited due to logistical and financial constraints <xref ref-type="bibr" rid="bib1.bibx44" id="paren.73"/>, we analyzed the information gained relative to the data points in the time series data. For this study, we calibrated friction parameters using multiple time series datasets, each containing several time steps. Figure <xref ref-type="fig" rid="F11"/> depicts the posterior distributions of the parameter <inline-formula><mml:math id="M263" display="inline"><mml:mi mathvariant="italic">ξ</mml:mi></mml:math></inline-formula> calibrated using a selected set of time series data of varying length. From Fig. <xref ref-type="fig" rid="F11"/>, we can infer that as the number of time steps increases, the rate of information gain decreases, which implies that the ratio of information gain to the length of time series data is skewed after a certain threshold. This inference aligns with the study of <xref ref-type="bibr" rid="bib1.bibx33" id="text.74"/>, who found that the hydrological model calibration did not improve after a certain threshold, even with increased data series length. This inference is further reflected in Fig. <xref ref-type="fig" rid="F12"/>a, where we plotted the information gain (quantified by KL divergence) against time steps. We can observe that the slope of this plot flattens after a certain threshold, indicating again that the rate of information gain diminishes after a threshold. We further observed from Fig. <xref ref-type="fig" rid="F12"/>b that this threshold corresponds to an observation window in the velocity time series during which the sliding mass accelerates and attains its maximum velocity. Thus, this is the duration which marks the point at which the system's dynamics have evolved and stabilized, as critical processes affecting the dynamics have happened. Therefore, data capturing these changes is significantly more informative and relevant than the rest. This finding reinforces our earlier point: it is not solely the data quantity that matters, but rather its ability to capture the critical processes governing the parameters.</p>
      <p id="d2e5224">Beyond the insights from our synthetic case study, the proposed Bayesian data selection workflow provides a systematic framework for performing a priori assessment of data informativeness before costly field campaigns. To implement this framework, practitioners should first prepare a synthetic test case using site-specific topography and their computational model. By generating candidate observational datasets through forward model evaluation at known parameters, they establish a virtual testbed for scenario-based testing that quantifies how different candidate datasets inform specific parameters of interest and tests case-specific hypotheses about data value.</p>
      <p id="d2e5227">While we demonstrate the workflow using idealized topography and synthetic data, the model-agnostic approach can be transferred to complex models and real topographies. However, two practical limitations should be acknowledged. First, our statistical model assumes negligible model discrepancy – an assumption valid for synthetic data but potentially problematic when applying the framework to field observations where structural model inadequacies may prevent full explanation of the data. Second, when model misspecification or prior-data conflict occurs, posteriors can contract to incorrect regions of parameter space, where higher KL divergence does not necessarily indicate better calibration. This limitation is exacerbated when using non-uniform priors. For non-uniform priors, KL divergence between the posterior and the prior no longer reduces to posterior entropy alone, as it captures both the sharpness of the posterior and its shift from the prior distribution. This introduces a trade-off: when a posterior contracts to an incorrect region, KL divergence can assign high information gain to a misleading result because it rewards both concentration and shift. Conversely, entropy focuses solely on posterior sharpness and does not account for whether the posterior has shifted toward more plausible parameter regions. When applying this framework in such scenarios, practitioners should therefore complement these information-theoretic metrics with robust validation methods such as posterior predictive checks.</p>
      <p id="d2e5230">Figure <xref ref-type="fig" rid="F14"/>a and b indicate that the discrepancy parameter is successfully calibrated using velocity and position time series. In this study, since we used synthetic data for calibration, we had control over the noise in the data; refer Sect. <xref ref-type="sec" rid="Ch1.S2.SS1"/> and Table <xref ref-type="table" rid="T1"/> for details. Specifically, we know the standard deviation of the Gaussian distribution from which the noise was drawn; for velocity time series, it was <inline-formula><mml:math id="M264" display="inline"><mml:mn mathvariant="normal">2.08</mml:mn></mml:math></inline-formula>, and for position time series <inline-formula><mml:math id="M265" display="inline"><mml:mn mathvariant="normal">225</mml:mn></mml:math></inline-formula>, see Sect. <xref ref-type="sec" rid="Ch1.S3.SS2"/>. These values closely align with the MAP estimates of the discrepancy parameter from the calibration of the velocity and position time series of <inline-formula><mml:math id="M266" display="inline"><mml:mn mathvariant="normal">2.03</mml:mn></mml:math></inline-formula> and <inline-formula><mml:math id="M267" display="inline"><mml:mn mathvariant="normal">220</mml:mn></mml:math></inline-formula>, indicating our ability to infer this parameter and, thus, reflecting our ability to quantify the uncertainty associated with measurement noise. These findings further indicate that higher data quantity (like time series data) can help us quantify the uncertainty associated with the data quality. Our results are consistent with <xref ref-type="bibr" rid="bib1.bibx29" id="text.75"/>, who reported that higher data quantity can offset the impact of poor data quality.</p>
</sec>
<sec id="Ch1.S6" sec-type="conclusions">
  <label>6</label><title>Conclusions</title>
      <p id="d2e5281">It is a well-known fact that the outcome of Bayesian calibration is highly dependent on the available observational data. However, there is a lack of studies that systematically investigate the impact of the selection of observation on the calibration result. We propose a Bayesian data selection workflow to address this challenge and identify the most informative observation in calibrating a given parameter. This workflow quantifies the impact of data selection on calibration performance by assessing the information gained during calibration, utilizing KL divergence, an established information-theoretic metric. Computing the KL divergence based on posterior distributions resulting from Bayesian parameter calibration presents itself as an extremely computationally intense task. We addressed the latter challenge by integrating a surrogate modeling technique based on GP emulation. The complete Bayesian data selection workflow is being made available with this article.</p>
      <p id="d2e5284">We have demonstrated the feasibility of the workflow through numerical experiments in which we systematically investigated the influence of data selection in calibrating two friction parameters of an idealized landslide runout model. To achieve this, we designed rigorous experiments that quantitatively assess how observations with variations in information content, specifically velocity versus position – and granularity, such as aggregated data versus time series data – affect the calibration outcome. The experimental results indicate that the information content and the observation granularity significantly impacted the calibration outcome. We found that time series data considerably outperforms aggregated data in constraining parameters, owing to its superior ability to capture the landslide dynamics. However, this does not imply that calibration performance scales linearly with the data quantity. While increasing the length and frequency of time series data enhances calibration performance, these improvements yield diminishing returns if the observation window exceeds a specific duration. Remarkably, the optimal length of the observation window that yields the maximum rate of information gain corresponds to the time the sliding mass needs to attain its maximum velocity. Thus, the evolution of landslide dynamics has stabilized. These findings suggest that data capturing the specific dynamics for an observation window of that duration are better suited to calibrate landslide model parameters. The landslide community can use these insights to optimize calibration strategies based on available data and to design effective future data acquisition strategies.</p>
</sec>

      
      </body>
    <back><app-group>

<app id="App1.Ch1.S1">
  <label>Appendix A</label><title/>
      <p id="d2e5297">In this appendix section, we demonstrate that when using uniform priors, maximizing KL divergence between posterior and prior distributions is equivalent to minimizing the entropy of the posterior distribution.</p>
      <p id="d2e5300">Entropy of a random variable <inline-formula><mml:math id="M268" display="inline"><mml:mi mathvariant="italic">θ</mml:mi></mml:math></inline-formula> with probability distribution <inline-formula><mml:math id="M269" display="inline"><mml:mrow><mml:mi>P</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="italic">θ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> is given as:

          <disp-formula id="App1.Ch1.S1.E17" content-type="numbered"><label>A1</label><mml:math id="M270" display="block"><mml:mrow><mml:mi>H</mml:mi><mml:mo>(</mml:mo><mml:mi>P</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="italic">θ</mml:mi><mml:mo>)</mml:mo><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:mo>-</mml:mo><mml:mo movablelimits="false">∫</mml:mo><mml:mi>P</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="italic">θ</mml:mi><mml:mo>)</mml:mo><mml:mi>log⁡</mml:mi><mml:mi>P</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="italic">θ</mml:mi><mml:mo>)</mml:mo><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mi mathvariant="normal">d</mml:mi><mml:mi mathvariant="italic">θ</mml:mi></mml:mrow></mml:math></disp-formula></p>
      <p id="d2e5373">KL divergence between the posterior and prior of a parameter <inline-formula><mml:math id="M271" display="inline"><mml:mi mathvariant="italic">θ</mml:mi></mml:math></inline-formula> calibrated using observation <inline-formula><mml:math id="M272" display="inline"><mml:mi>y</mml:mi></mml:math></inline-formula> is given as:

          <disp-formula id="App1.Ch1.S1.E18" content-type="numbered"><label>A2</label><mml:math id="M273" display="block"><mml:mrow><mml:msub><mml:mi>D</mml:mi><mml:mi mathvariant="normal">KL</mml:mi></mml:msub><mml:mo>(</mml:mo><mml:mi>P</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="italic">θ</mml:mi><mml:mo>∣</mml:mo><mml:mi>y</mml:mi><mml:mo>)</mml:mo><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mo>‖</mml:mo><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mi>P</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="italic">θ</mml:mi><mml:mo>)</mml:mo><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:mo movablelimits="false">∫</mml:mo><mml:mi>P</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="italic">θ</mml:mi><mml:mo>∣</mml:mo><mml:mi>y</mml:mi><mml:mo>)</mml:mo><mml:mi>log⁡</mml:mi><mml:mfenced close=")" open="("><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:mi>P</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="italic">θ</mml:mi><mml:mo>∣</mml:mo><mml:mi>y</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mi>P</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="italic">θ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mfrac></mml:mstyle></mml:mfenced><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mi mathvariant="normal">d</mml:mi><mml:mi mathvariant="italic">θ</mml:mi></mml:mrow></mml:math></disp-formula></p>
      <p id="d2e5480">Expanding the logarithm, we obtain:

          <disp-formula id="App1.Ch1.S1.E19" content-type="numbered"><label>A3</label><mml:math id="M274" display="block"><mml:mrow><mml:mtable rowspacing="0.2ex" class="split" displaystyle="true" columnalign="right left"><mml:mtr><mml:mtd/><mml:mtd><mml:mrow><mml:msub><mml:mi>D</mml:mi><mml:mi mathvariant="normal">KL</mml:mi></mml:msub><mml:mo>(</mml:mo><mml:mi>P</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="italic">θ</mml:mi><mml:mo>∣</mml:mo><mml:mi>y</mml:mi><mml:mo>)</mml:mo><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mo>‖</mml:mo><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mi>P</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="italic">θ</mml:mi><mml:mo>)</mml:mo><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:mo movablelimits="false">∫</mml:mo><mml:mi>P</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="italic">θ</mml:mi><mml:mo>∣</mml:mo><mml:mi>y</mml:mi><mml:mo>)</mml:mo><mml:mfenced close="]" open="["><mml:mrow><mml:mi>log⁡</mml:mi><mml:mi>P</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="italic">θ</mml:mi><mml:mo>∣</mml:mo><mml:mi>y</mml:mi><mml:mo>)</mml:mo><mml:mo>-</mml:mo><mml:mi>log⁡</mml:mi><mml:mi>P</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="italic">θ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mfenced><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mi mathvariant="normal">d</mml:mi><mml:mi mathvariant="italic">θ</mml:mi></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd/><mml:mtd><mml:mrow><mml:mspace width="1em" linebreak="nobreak"/><mml:mo>=</mml:mo><mml:mo movablelimits="false">∫</mml:mo><mml:mi>P</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="italic">θ</mml:mi><mml:mo>∣</mml:mo><mml:mi>y</mml:mi><mml:mo>)</mml:mo><mml:mi>log⁡</mml:mi><mml:mi>P</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="italic">θ</mml:mi><mml:mo>∣</mml:mo><mml:mi>y</mml:mi><mml:mo>)</mml:mo><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mi mathvariant="normal">d</mml:mi><mml:mi mathvariant="italic">θ</mml:mi><mml:mo>-</mml:mo><mml:mo movablelimits="false">∫</mml:mo><mml:mi>P</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="italic">θ</mml:mi><mml:mo>∣</mml:mo><mml:mi>y</mml:mi><mml:mo>)</mml:mo><mml:mi>log⁡</mml:mi><mml:mi>P</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="italic">θ</mml:mi><mml:mo>)</mml:mo><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mi mathvariant="normal">d</mml:mi><mml:mi mathvariant="italic">θ</mml:mi></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd/><mml:mtd><mml:mrow><mml:mspace width="1em" linebreak="nobreak"/><mml:mo>=</mml:mo><mml:mo>-</mml:mo><mml:mi>H</mml:mi><mml:mo>(</mml:mo><mml:mi>P</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="italic">θ</mml:mi><mml:mo>∣</mml:mo><mml:mi>y</mml:mi><mml:mo>)</mml:mo><mml:mo>)</mml:mo><mml:mo>-</mml:mo><mml:mo movablelimits="false">∫</mml:mo><mml:mi>P</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="italic">θ</mml:mi><mml:mo>∣</mml:mo><mml:mi>y</mml:mi><mml:mo>)</mml:mo><mml:mi>log⁡</mml:mi><mml:mi>P</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="italic">θ</mml:mi><mml:mo>)</mml:mo><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mi mathvariant="normal">d</mml:mi><mml:mi mathvariant="italic">θ</mml:mi></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:mrow></mml:math></disp-formula></p>
      <p id="d2e5715">For a uniform prior over a bounded domain <inline-formula><mml:math id="M275" display="inline"><mml:mi mathvariant="normal">Θ</mml:mi></mml:math></inline-formula> with volume <inline-formula><mml:math id="M276" display="inline"><mml:mi>V</mml:mi></mml:math></inline-formula>, we have <inline-formula><mml:math id="M277" display="inline"><mml:mrow><mml:mi>P</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="italic">θ</mml:mi><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:mstyle displaystyle="false"><mml:mfrac style="text"><mml:mn mathvariant="normal">1</mml:mn><mml:mi>V</mml:mi></mml:mfrac></mml:mstyle></mml:mrow></mml:math></inline-formula> (constant). Substituting this:

          <disp-formula id="App1.Ch1.S1.E20" content-type="numbered"><label>A4</label><mml:math id="M278" display="block"><mml:mrow><mml:mtable class="split" rowspacing="0.2ex" displaystyle="true" columnalign="right left"><mml:mtr><mml:mtd/><mml:mtd><mml:mrow><mml:msub><mml:mi>D</mml:mi><mml:mi mathvariant="normal">KL</mml:mi></mml:msub><mml:mo>(</mml:mo><mml:mi>P</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="italic">θ</mml:mi><mml:mo>∣</mml:mo><mml:mi>y</mml:mi><mml:mo>)</mml:mo><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mo>‖</mml:mo><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mi>P</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="italic">θ</mml:mi><mml:mo>)</mml:mo><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:mo>-</mml:mo><mml:mi>H</mml:mi><mml:mo>(</mml:mo><mml:mi>P</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="italic">θ</mml:mi><mml:mo>∣</mml:mo><mml:mi>y</mml:mi><mml:mo>)</mml:mo><mml:mo>)</mml:mo><mml:mo>-</mml:mo><mml:mo movablelimits="false">∫</mml:mo><mml:mi>P</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="italic">θ</mml:mi><mml:mo>∣</mml:mo><mml:mi>y</mml:mi><mml:mo>)</mml:mo><mml:mi>log⁡</mml:mi><mml:mfenced open="(" close=")"><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mn mathvariant="normal">1</mml:mn><mml:mi>V</mml:mi></mml:mfrac></mml:mstyle></mml:mfenced><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mi mathvariant="normal">d</mml:mi><mml:mi mathvariant="italic">θ</mml:mi></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd/><mml:mtd><mml:mrow><mml:mspace width="1em" linebreak="nobreak"/><mml:mo>=</mml:mo><mml:mo>-</mml:mo><mml:mi>H</mml:mi><mml:mo>(</mml:mo><mml:mi>P</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="italic">θ</mml:mi><mml:mo>∣</mml:mo><mml:mi>y</mml:mi><mml:mo>)</mml:mo><mml:mo>)</mml:mo><mml:mo>+</mml:mo><mml:mi>log⁡</mml:mi><mml:mi>V</mml:mi><mml:mo movablelimits="false">∫</mml:mo><mml:mi>P</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="italic">θ</mml:mi><mml:mo>∣</mml:mo><mml:mi>y</mml:mi><mml:mo>)</mml:mo><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mi mathvariant="normal">d</mml:mi><mml:mi mathvariant="italic">θ</mml:mi></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd/><mml:mtd><mml:mrow><mml:mspace linebreak="nobreak" width="1em"/><mml:mo>=</mml:mo><mml:mo>-</mml:mo><mml:mi>H</mml:mi><mml:mo>(</mml:mo><mml:mi>P</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="italic">θ</mml:mi><mml:mo>∣</mml:mo><mml:mi>y</mml:mi><mml:mo>)</mml:mo><mml:mo>)</mml:mo><mml:mo>+</mml:mo><mml:mi>log⁡</mml:mi><mml:mi>V</mml:mi></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:mrow></mml:math></disp-formula></p>
      <p id="d2e5942">Thus KL divergence between the posterior and prior distributions reduces to the negative entropy of the posterior distribution plus a constant <inline-formula><mml:math id="M279" display="inline"><mml:mrow><mml:mi>log⁡</mml:mi><mml:mo>(</mml:mo><mml:mi>V</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>, when using uniform priors. Therefore, selecting observations to maximize the KL divergence is equivalent to selecting observations that yield the sharpest posterior as discussed in Sect. <xref ref-type="sec" rid="Ch1.S2.SS3"/>.</p>
</app>

<app id="App1.Ch1.S2">
  <label>Appendix B</label><title/>
      <p id="d2e5969">In this appendix, we present trace plots of the MCMC chains from selected experiments (specifically experiments 1, 2, 3, 4, 7, and 8). The corresponding <inline-formula><mml:math id="M280" display="inline"><mml:mover accent="true"><mml:mi>R</mml:mi><mml:mo mathvariant="normal" stretchy="false">^</mml:mo></mml:mover></mml:math></inline-formula> values are reported in Table <xref ref-type="table" rid="TB1"/>. While the trace plots provide qualitative assessment of the convergence of the MCMC chains, <inline-formula><mml:math id="M281" display="inline"><mml:mover accent="true"><mml:mi>R</mml:mi><mml:mo stretchy="false" mathvariant="normal">^</mml:mo></mml:mover></mml:math></inline-formula> values provides a quantitative measure of convergence <xref ref-type="bibr" rid="bib1.bibx47" id="paren.76"/>.</p>

      <fig id="FB1"><label>Figure B1</label><caption><p id="d2e5999">MCMC trace plots for <bold>(a)</bold> Dry Coulomb friction coefficient <inline-formula><mml:math id="M282" display="inline"><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="italic">μ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> and <bold>(b)</bold> Turbulent friction coefficient <inline-formula><mml:math id="M283" display="inline"><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="italic">ξ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> based on calibration with maximum velocity <inline-formula><mml:math id="M284" display="inline"><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>U</mml:mi><mml:mi mathvariant="normal">max</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> as observation. Each color represents an independent chain. Left panels show kernel density estimates of the marginal posterior distributions. Right panels show trace plots across iterations, demonstrating chain convergence and adequate mixing.</p></caption>
        
        <graphic xlink:href="https://npg.copernicus.org/articles/33/425/2026/npg-33-425-2026-f16.png"/>

      </fig>

<fig id="FB2"><label>Figure B2</label><caption><p id="d2e6059">MCMC trace plots for <bold>(a)</bold> Dry Coulomb friction coefficient <inline-formula><mml:math id="M285" display="inline"><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="italic">μ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> and <bold>(b)</bold> Turbulent friction coefficient <inline-formula><mml:math id="M286" display="inline"><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="italic">ξ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> based on calibration with runout distance <inline-formula><mml:math id="M287" display="inline"><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>X</mml:mi><mml:mi mathvariant="normal">end</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> as observation. Each color represents an independent chain. Left panels show kernel density estimates of the marginal posterior distributions. Right panels show trace plots across iterations, demonstrating chain convergence and adequate mixing.</p></caption>
        
        <graphic xlink:href="https://npg.copernicus.org/articles/33/425/2026/npg-33-425-2026-f17.png"/>

      </fig>

      <fig id="FB3"><label>Figure B3</label><caption><p id="d2e6117">MCMC trace plots for <bold>(a)</bold> Dry Coulomb friction coefficient <inline-formula><mml:math id="M288" display="inline"><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="italic">μ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> and <bold>(b)</bold> Turbulent friction coefficient <inline-formula><mml:math id="M289" display="inline"><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="italic">ξ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> based on calibration with velocity time series as observation. Each color represents an independent chain. Left panels show kernel density estimates of the marginal posterior distributions. Right panels show trace plots across iterations, demonstrating chain convergence and adequate mixing.</p></caption>
        
        <graphic xlink:href="https://npg.copernicus.org/articles/33/425/2026/npg-33-425-2026-f18.png"/>

      </fig>

<fig id="FB4"><label>Figure B4</label><caption><p id="d2e6161">MCMC trace plots for <bold>(a)</bold> Dry Coulomb friction coefficient <inline-formula><mml:math id="M290" display="inline"><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="italic">μ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> and <bold>(b)</bold> Turbulent friction coefficient <inline-formula><mml:math id="M291" display="inline"><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="italic">ξ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> based on calibration with position time series as observation. Each color represents an independent chain. Left panels show kernel density estimates of the marginal posterior distributions. Right panels show trace plots across iterations, demonstrating chain convergence and adequate mixing.</p></caption>
        
        <graphic xlink:href="https://npg.copernicus.org/articles/33/425/2026/npg-33-425-2026-f19.png"/>

      </fig>

<fig id="FB5"><label>Figure B5</label><caption><p id="d2e6206">MCMC trace plots for <bold>(a)</bold> Dry Coulomb friction coefficient <inline-formula><mml:math id="M292" display="inline"><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="italic">μ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> and <bold>(b)</bold> Turbulent friction coefficient <inline-formula><mml:math id="M293" display="inline"><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="italic">ξ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> <bold>(c)</bold> Velocity discrepancy parameter <inline-formula><mml:math id="M294" display="inline"><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi mathvariant="italic">σ</mml:mi><mml:mrow><mml:msub><mml:mi mathvariant="normal">vel</mml:mi><mml:mi mathvariant="normal">TS</mml:mi></mml:msub></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> based on calibration with velocity time series as observation. Each color represents an independent chain. Left panels show kernel density estimates of the marginal posterior distributions. Right panels show trace plots across iterations, demonstrating chain convergence and adequate mixing.</p></caption>
        
        <graphic xlink:href="https://npg.copernicus.org/articles/33/425/2026/npg-33-425-2026-f20.png"/>

      </fig>

<fig id="FB6"><label>Figure B6</label><caption><p id="d2e6273">MCMC trace plots for <bold>(a)</bold> Dry Coulomb friction coefficient <inline-formula><mml:math id="M295" display="inline"><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="italic">μ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> and <bold>(b)</bold> Turbulent friction coefficient <inline-formula><mml:math id="M296" display="inline"><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="italic">ξ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> <bold>(c)</bold> Position discrepancy parameter <inline-formula><mml:math id="M297" display="inline"><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi mathvariant="italic">σ</mml:mi><mml:mrow><mml:msub><mml:mi mathvariant="normal">pos</mml:mi><mml:mi mathvariant="normal">TS</mml:mi></mml:msub></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> based on calibration with position time series as observation. Each color represents an independent chain. Left panels show kernel density estimates of the marginal posterior distributions. Right panels show trace plots across iterations, demonstrating chain convergence and adequate mixing.</p></caption>
        
        <graphic xlink:href="https://npg.copernicus.org/articles/33/425/2026/npg-33-425-2026-f21.png"/>

      </fig>

<table-wrap id="TB1"><label>Table B1</label><caption><p id="d2e6342">Gelman-Rubin Statistic <inline-formula><mml:math id="M298" display="inline"><mml:mrow><mml:mo>(</mml:mo><mml:mover accent="true"><mml:mi>R</mml:mi><mml:mo mathvariant="normal" stretchy="false">^</mml:mo></mml:mover><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> for MCMC sampling of friction coefficients and the discrepancy parameter. <inline-formula><mml:math id="M299" display="inline"><mml:mover accent="true"><mml:mi>R</mml:mi><mml:mo mathvariant="normal" stretchy="false">^</mml:mo></mml:mover></mml:math></inline-formula> values indicate successful chain convergence for all parameters across different observations.</p></caption><oasis:table frame="topbot"><oasis:tgroup cols="4">
     <oasis:colspec colnum="1" colname="col1" align="left"/>
     <oasis:colspec colnum="2" colname="col2" align="right"/>
     <oasis:colspec colnum="3" colname="col3" align="right"/>
     <oasis:colspec colnum="4" colname="col4" align="right"/>
     <oasis:thead>
       <oasis:row>
         <oasis:entry colname="col1">Observations</oasis:entry>
         <oasis:entry rowsep="1" namest="col2" nameend="col4" align="center">Gelman-Rubin Statistic <inline-formula><mml:math id="M300" display="inline"><mml:mrow><mml:mo>(</mml:mo><mml:mover accent="true"><mml:mi>R</mml:mi><mml:mo stretchy="false" mathvariant="normal">^</mml:mo></mml:mover><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula></oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2">Dry Coulomb friction coefficient</oasis:entry>
         <oasis:entry colname="col3">Turbulent friction coefficient</oasis:entry>
         <oasis:entry colname="col4">Discrepancy parameter</oasis:entry>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row>
         <oasis:entry colname="col1">Maximum velocity <inline-formula><mml:math id="M301" display="inline"><mml:mrow><mml:msub><mml:mi>U</mml:mi><mml:mtext>max</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col2">1.007</oasis:entry>
         <oasis:entry colname="col3">1.007</oasis:entry>
         <oasis:entry colname="col4">–</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">runout distance <inline-formula><mml:math id="M302" display="inline"><mml:mrow><mml:msub><mml:mi>X</mml:mi><mml:mtext>end</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col2">1.008</oasis:entry>
         <oasis:entry colname="col3">1.009</oasis:entry>
         <oasis:entry colname="col4">–</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Velocity time series <inline-formula><mml:math id="M303" display="inline"><mml:mrow><mml:mi>u</mml:mi><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col2">1.006</oasis:entry>
         <oasis:entry colname="col3">1.005</oasis:entry>
         <oasis:entry colname="col4">–</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Position time series <inline-formula><mml:math id="M304" display="inline"><mml:mrow><mml:mi>x</mml:mi><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col2">1.010</oasis:entry>
         <oasis:entry colname="col3">1.010</oasis:entry>
         <oasis:entry colname="col4">–</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Velocity time series <inline-formula><mml:math id="M305" display="inline"><mml:mrow><mml:mi>u</mml:mi><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> (discrepancy)</oasis:entry>
         <oasis:entry colname="col2">1.007</oasis:entry>
         <oasis:entry colname="col3">1.008</oasis:entry>
         <oasis:entry colname="col4">1.011</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Position time series <inline-formula><mml:math id="M306" display="inline"><mml:mrow><mml:mi>x</mml:mi><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>  (discrepancy)</oasis:entry>
         <oasis:entry colname="col2">1.009</oasis:entry>
         <oasis:entry colname="col3">1.009</oasis:entry>
         <oasis:entry colname="col4">1.007</oasis:entry>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup></oasis:table></table-wrap>


</app>

<app id="App1.Ch1.S3">
  <label>Appendix C</label><title/>
      <p id="d2e6601">Figure <xref ref-type="fig" rid="F11"/> plots the posterior distributions of <inline-formula><mml:math id="M307" display="inline"><mml:mi mathvariant="italic">ξ</mml:mi></mml:math></inline-formula> calibrated with velocity time series datasets of increasing length. While information gain increases with time series length (reflected in progressively narrower posteriors), the differences between consecutive posteriors are much larger for initial time steps than for later ones, indicating a decreasing rate of information gain. To determine whether early time steps inherently contain more information, we performed calibration experiments using fixed-length observation time windows at different temporal positions. Figure <xref ref-type="fig" rid="FC1"/> shows the posterior distributions of <inline-formula><mml:math id="M308" display="inline"><mml:mi mathvariant="italic">ξ</mml:mi></mml:math></inline-formula> calibrated using 100 time steps from observation windows positioned at: (0–100), (100–200), (200–300), (300–400), (700–800), (1000–1100), (1100–1200), and (1200–1300). This experimental setup allows direct comparison of information content across temporal positions. The results clearly show that posteriors become progressively wider for later time windows, confirming that initial time steps contain significantly more information about <inline-formula><mml:math id="M309" display="inline"><mml:mi mathvariant="italic">ξ</mml:mi></mml:math></inline-formula> than later time steps.</p>

      <fig id="FC1"><label>Figure C1</label><caption><p id="d2e6631">Posterior distributions of turbulent friction coefficient <inline-formula><mml:math id="M310" display="inline"><mml:mi mathvariant="italic">ξ</mml:mi></mml:math></inline-formula> calibrated with a fixed observation window of <inline-formula><mml:math id="M311" display="inline"><mml:mn mathvariant="normal">100</mml:mn></mml:math></inline-formula> time steps, with varying onset time. Initial time steps provide the highest information gain.</p></caption>
        
        <graphic xlink:href="https://npg.copernicus.org/articles/33/425/2026/npg-33-425-2026-f22.png"/>

      </fig>


</app>
  </app-group><notes notes-type="codeavailability"><title>Code availability</title>

      <p id="d2e6662">Code required to perform the numerical experiments listed in Sect. <xref ref-type="sec" rid="Ch1.S4"/> is available at this repository <ext-link xlink:href="https://doi.org/10.5281/zenodo.17120721" ext-link-type="DOI">10.5281/zenodo.17120721</ext-link> <xref ref-type="bibr" rid="bib1.bibx32" id="paren.77"/>.</p>
  </notes><notes notes-type="dataavailability"><title>Data availability</title>

      <p id="d2e6676">Data required for the experiments: (i) Posterior samples corresponding to the calibration routines, (ii) Training data comprising of the design (sampled set of input parameters) and the corresponding model outputs (iii) Ground truth data used in the calibration routines is hosted in this repository <ext-link xlink:href="https://doi.org/10.5281/zenodo.17120721" ext-link-type="DOI">10.5281/zenodo.17120721</ext-link> <xref ref-type="bibr" rid="bib1.bibx32" id="paren.78"/>.</p>
  </notes><notes notes-type="ercavailability"><title>Interactive computing environment (ICE)</title>

      <p id="d2e6688">The Zenodo repository (<ext-link xlink:href="https://doi.org/10.5281/zenodo.17120721" ext-link-type="DOI">10.5281/zenodo.17120721</ext-link>, <xref ref-type="bibr" rid="bib1.bibx32" id="altparen.79"/>) contains the code base used to produce the results in this paper: a set of Jupyter notebooks (data selection, posterior analysis, and figure generation), the training/model-output data (HDF5 and joblib), MCMC results, and a conda environment specification (Yaml/BDS_environment.yaml) listing the required packages.</p>
  </notes><notes notes-type="authorcontribution"><title>Author contributions</title>

      <p id="d2e6700">Author contributions follow the CRediT taxonomy. VMK: Conceptualization, Methodology, Software, Formal analysis, Investigation, Writing – original draft, Writing – review and editing, Visualization. AY: Conceptualization, Methodology, Writing – review and editing, Supervision, Visualization. JK: Conceptualization, Methodology, Writing – review and editing, Supervision, Funding acquisition.</p>
  </notes><notes notes-type="competinginterests"><title>Competing interests</title>

      <p id="d2e6706">The contact author has declared that none of the authors has any competing interests.</p>
  </notes><notes notes-type="disclaimer"><title>Disclaimer</title>

      <p id="d2e6712">Publisher's note: Copernicus Publications remains neutral with regard to jurisdictional claims made in the text, published maps, institutional affiliations, or any other geographical representation in this paper. The authors bear the ultimate responsibility for providing appropriate place names. Views expressed in the text are those of the authors and do not necessarily reflect the views of the publisher.</p>
  </notes><notes notes-type="financialsupport"><title>Financial support</title>

      <p id="d2e6719">This research has been supported by financial support from DFG-Deutsche Forschungsgemeinschaft (German Research Foundation) under the grant 333849990/GRK2379 (International Research Training Group (IRTG-2379): Hierarchical and Hybrid Approaches in Modern Inverse Problems). Authors also acknowledge the funding by Deutsche Forschungsgemeinschaft (DFG) within the framework of the research project OptiData: Improving the Predictivity of Simulating Natural Hazards due to Mass Movements – Optimal Design and Model Selection (Project no. 441527981).This open-access publication was funded  by the RWTH Aachen University.</p>
  </notes><notes notes-type="reviewstatement"><title>Review statement</title>

      <p id="d2e6730">This paper was edited by Amit Apte and reviewed by Reyko Schachtschneider, Aki Vehtari, Flavia Pinheiro, and one anonymous referee.</p>
  </notes><ref-list>
    <title>References</title>

      <ref id="bib1.bibx1"><label>Aaron(2017)</label><mixed-citation>Aaron, J.: Advancement and calibration of a 3D numerical model for landslide runout analysis, PhD thesis, Univ. British Columbia, Vancouver, <ext-link xlink:href="https://doi.org/10.14288/1.0357191" ext-link-type="DOI">10.14288/1.0357191</ext-link>, 2017.</mixed-citation></ref>
      <ref id="bib1.bibx2"><label>Aaron et al.(2019)Aaron, McDougall, and Nolde</label><mixed-citation>Aaron, J., McDougall, S., and Nolde, N.: Two methodologies to calibrate landslide runout models, Landslides, 16, 907–920, <ext-link xlink:href="https://doi.org/10.1007/s10346-018-1116-8" ext-link-type="DOI">10.1007/s10346-018-1116-8</ext-link>, 2019.</mixed-citation></ref>
      <ref id="bib1.bibx3"><label>Aaron et al.(2022)Aaron, McDougall, Kowalski, Mitchell, and Nolde</label><mixed-citation>Aaron, J., McDougall, S., Kowalski, J., Mitchell, A., and Nolde, N.: Probabilistic prediction of rock avalanche runout using a numerical model, Landslides, 19, 2853–2869, <ext-link xlink:href="https://doi.org/10.1007/s10346-022-01939-y" ext-link-type="DOI">10.1007/s10346-022-01939-y</ext-link>, 2022.</mixed-citation></ref>
      <ref id="bib1.bibx4"><label>Baptista et al.(2022)Baptista, Cao, Chen, Ghattas, Li, Marzouk, and Oden</label><mixed-citation>Baptista, R., Cao, L., Chen, J., Ghattas, O., Li, F., Marzouk, Y. M., and Oden, J. T.: Bayesian model calibration for block copolymer self-assembly: Likelihood-free inference and expected information gain computation via measure transport, arXiv, <ext-link xlink:href="https://doi.org/10.48550/arXiv.2206.11343" ext-link-type="DOI">10.48550/arXiv.2206.11343</ext-link>, 2022.</mixed-citation></ref>
      <ref id="bib1.bibx5"><label>Barros et al.(2009)Barros, Kirby, and Mavris</label><mixed-citation>Barros, P. A., Kirby, M. R., and Mavris, D. N.: A Review of Calibration under Uncertainty within the Environmental Design Space, in: 47th AIAA Aerospace Sciences Meeting including The New Horizons Forum and Aerospace Exposition, American Institute of Aeronautics and Astronautics, Reston, Virginia, <ext-link xlink:href="https://doi.org/10.2514/6.2009-974" ext-link-type="DOI">10.2514/6.2009-974</ext-link>, 2009.</mixed-citation></ref>
      <ref id="bib1.bibx6"><label>Brezzi et al.(2016)Brezzi, Gabrieli, Marcato, Pastor, and Cola</label><mixed-citation>Brezzi, L., Gabrieli, F., Marcato, G., Pastor, M., and Cola, S.: A new data assimilation procedure to develop a debris flow run-out model, Landslides, 13, 1083–1096, <ext-link xlink:href="https://doi.org/10.1007/s10346-015-0625-y" ext-link-type="DOI">10.1007/s10346-015-0625-y</ext-link>, 2016.</mixed-citation></ref>
      <ref id="bib1.bibx7"><label>Chowdhary et al.(2024)Chowdhary, Tong, Stadler, and Alexanderian</label><mixed-citation>Chowdhary, A., Tong, S., Stadler, G., and Alexanderian, A.: Sensitivity Analysis of the information gain in infinite-dimensional Bayesian linear inverse problems, Int. J. Uncertain. Quan., 14, 17–35, <ext-link xlink:href="https://doi.org/10.1615/int.j.uncertaintyquantification.2024051416" ext-link-type="DOI">10.1615/int.j.uncertaintyquantification.2024051416</ext-link>, 2024.</mixed-citation></ref>
      <ref id="bib1.bibx8"><label>Cotter(2024)</label><mixed-citation>Cotter, S. L.: Hierarchical Bayesian Data Selection, ACM Trans. Probab. Mach. Learn., 1, 7, <ext-link xlink:href="https://doi.org/10.1145/3699721" ext-link-type="DOI">10.1145/3699721</ext-link>, 2024.</mixed-citation></ref>
      <ref id="bib1.bibx9"><label>Cui et al.(2015)Cui, Sun, Teng, Song, and Yao</label><mixed-citation>Cui, X., Sun, W., Teng, J., Song, H., and Yao, X.: Effect of length of the observed dataset on the calibration of a distributed hydrological model, Proc. IAHS, 368, 305–311, <ext-link xlink:href="https://doi.org/10.5194/piahs-368-305-2015" ext-link-type="DOI">10.5194/piahs-368-305-2015</ext-link>, 2015.</mixed-citation></ref>
      <ref id="bib1.bibx10"><label>Ekmekcioğlu et al.(2022)Ekmekcioğlu, Demirel, and Booij</label><mixed-citation>Ekmekcioğlu, Ö., Demirel, M. C., and Booij, M. J.: Effect of data length, spin-up period and spatial model resolution on fully distributed hydrological model calibration in the Moselle basin, Hydrol. Sci. J., 67, 759–772, <ext-link xlink:href="https://doi.org/10.1080/02626667.2022.2046754" ext-link-type="DOI">10.1080/02626667.2022.2046754</ext-link>, 2022.</mixed-citation></ref>
      <ref id="bib1.bibx11"><label>Etter et al.(2018)Etter, Strobl, Seibert, and van Meerveld</label><mixed-citation>Etter, S., Strobl, B., Seibert, J., and van Meerveld, H. J. I.: Value of uncertain streamflow observations for hydrological modelling, Hydrol. Earth Syst. Sci., 22, 5243–5257, <ext-link xlink:href="https://doi.org/10.5194/hess-22-5243-2018" ext-link-type="DOI">10.5194/hess-22-5243-2018</ext-link>, 2018.</mixed-citation></ref>
      <ref id="bib1.bibx12"><label>Fischer et al.(2020)Fischer, Kofler, Huber, Fellin, Mergili, and Oberguggenberger</label><mixed-citation>Fischer, J. T., Kofler, A., Huber, A., Fellin, W., Mergili, M., and Oberguggenberger, M.: Bayesian inference in snow avalanche simulation with r.Avaflow, Geosci., 10, 191, <ext-link xlink:href="https://doi.org/10.3390/geosciences10050191" ext-link-type="DOI">10.3390/geosciences10050191</ext-link>, 2020.</mixed-citation></ref>
      <ref id="bib1.bibx13"><label>Foreman-Mackey et al.(2013)Foreman-Mackey, Hogg, Lang, and Goodman</label><mixed-citation>Foreman-Mackey, D., Hogg, D. W., Lang, D., and Goodman, J.: emcee: The MCMC Hammer, Publ. Astron. Soc. Pac., 125, 306–312, <ext-link xlink:href="https://doi.org/10.1086/670067" ext-link-type="DOI">10.1086/670067</ext-link>, 2013.</mixed-citation></ref>
      <ref id="bib1.bibx14"><label>Froese et al.(2012)Froese, Charrière, Humair, Jaboyedoff, and Pedrazzini</label><mixed-citation> Froese, C. R., Charrière, M., Humair, F., Jaboyedoff, M., and Pedrazzini, A.: Characterization and management of rockslide hazard at Turtle Mountain, Alberta, Canada, 310–322, Cambridge Univ. Press, ISBN 9781107002067, 2012.</mixed-citation></ref>
      <ref id="bib1.bibx15"><label>Gelman et al.(2013)Gelman, Carlin, Stern, Dunson, Vehtari, and Rubin</label><mixed-citation>Gelman, A., Carlin, J. B., Stern, H. S., Dunson, D. B., Vehtari, A., and Rubin, D. B.: Bayesian Data Analysis, 3rd edn., Chapman and Hall/CRC, <ext-link xlink:href="https://doi.org/10.1201/b16018" ext-link-type="DOI">10.1201/b16018</ext-link>, 2013.</mixed-citation></ref>
      <ref id="bib1.bibx16"><label>Gu and Berger(2016)</label><mixed-citation>Gu, M. and Berger, J. O.: Parallel partial Gaussian process emulation for computer models with massive output, Ann. Appl. Stat., 10, 1317–1347, <ext-link xlink:href="https://doi.org/10.1214/16-AOAS934" ext-link-type="DOI">10.1214/16-AOAS934</ext-link>, 2016.</mixed-citation></ref>
      <ref id="bib1.bibx17"><label>Haeusel et al.(2026)Haeusel, Nitzler, Köglmeier, and Wall</label><mixed-citation>Haeusel, L. J., Nitzler, J., Köglmeier, L. J., and Wall, W. A.: Multi-physics-enhanced Bayesian inverse analysis: Information gain from additional fields, Comput. Method. Appl. M., 452, 118735, <ext-link xlink:href="https://doi.org/10.1016/j.cma.2026.118735" ext-link-type="DOI">10.1016/j.cma.2026.118735</ext-link>, 2026.</mixed-citation></ref>
      <ref id="bib1.bibx18"><label>Hartland(2020)</label><mixed-citation>Hartland, N.: KL-divergence-estimators, GitHub [code], <uri>https://github.com/nhartland/KL-divergence-estimators</uri> (last access: 15 September 2025), 2020.</mixed-citation></ref>
      <ref id="bib1.bibx19"><label>Heo et al.(2015)Heo, Graziano, Guzowski, and Muehleisen</label><mixed-citation>Heo, Y., Graziano, D. J., Guzowski, L., and Muehleisen, R. T.: Evaluation of calibration efficacy under different levels of uncertainty, J. Build. Perform. Simu., 8, 135–144, <ext-link xlink:href="https://doi.org/10.1080/19401493.2014.896947" ext-link-type="DOI">10.1080/19401493.2014.896947</ext-link>, 2015.</mixed-citation></ref>
      <ref id="bib1.bibx20"><label>Heredia et al.(2020)Heredia, Eckert, Prieur, and Thibert</label><mixed-citation>Heredia, M. B., Eckert, N., Prieur, C., and Thibert, E.: Bayesian calibration of an avalanche model from autocorrelated measurements along the flow: Application to velocities extracted from photogrammetric images, J. Glaciol., 66, 373–385, <ext-link xlink:href="https://doi.org/10.1017/jog.2020.11" ext-link-type="DOI">10.1017/jog.2020.11</ext-link>, 2020.</mixed-citation></ref>
      <ref id="bib1.bibx21"><label>Hergarten(2024)</label><mixed-citation>Hergarten, S.: Scaling between volume and runout of rock avalanches explained by a modified Voellmy rheology, Earth Surf. Dynam., 12, 219–229, <ext-link xlink:href="https://doi.org/10.5194/esurf-12-219-2024" ext-link-type="DOI">10.5194/esurf-12-219-2024</ext-link>, 2024.</mixed-citation></ref>
      <ref id="bib1.bibx22"><label>Huber et al.(2023)Huber, Georgia, and Finley</label><mixed-citation>Huber, H. A., Georgia, S. K., and Finley, S. D.: Systematic Bayesian posterior analysis guided by Kullback-Leibler divergence facilitates hypothesis formation, J. Theor. Biol., 558, 111341, <ext-link xlink:href="https://doi.org/10.1016/j.jtbi.2022.111341" ext-link-type="DOI">10.1016/j.jtbi.2022.111341</ext-link>, 2023.</mixed-citation></ref>
      <ref id="bib1.bibx23"><label>Hungr(1995)</label><mixed-citation>Hungr, O.: A model for the runout analysis of rapid flow slides, debris flows and avalanches, Can. Geotech. J., 32, 610–623, <ext-link xlink:href="https://doi.org/10.1139/t95-063" ext-link-type="DOI">10.1139/t95-063</ext-link>, 1995.</mixed-citation></ref>
      <ref id="bib1.bibx24"><label>Hungr and McDougall(2009)</label><mixed-citation>Hungr, O. and McDougall, S.: Two numerical models for landslide dynamic analysis, Comput. Geosci., 35, 978–992, <ext-link xlink:href="https://doi.org/10.1016/j.cageo.2007.12.003" ext-link-type="DOI">10.1016/j.cageo.2007.12.003</ext-link>, 2009.</mixed-citation></ref>
      <ref id="bib1.bibx25"><label>Hübl et al.(2009)Hübl, Suda, Proske, Kaitna, and Scheidl</label><mixed-citation> Hübl, J., Suda, J., Proske, D., Kaitna, R., and Scheidl, C.: Debris Flow Impact Estimation, in: Proc. 11th Int. Symp. Water Manag. Hydraul. Eng., Faculty of Civil Engineering, edited by: Popovska, C. and Jovanovski, M., University of Ss. Cyril and Methodius, Skopje, Macedonia, 137–148, ISBN 978-9989-2469-7-5,  2009.</mixed-citation></ref>
      <ref id="bib1.bibx26"><label>Iverson(2003)</label><mixed-citation>Iverson, R. M.: How should mathematical models of geomorphic processes be judged?, in: Prediction in Geomorphology, vol. 135 of Geophys. Monogr., Am. Geophys. Union, <ext-link xlink:href="https://doi.org/10.1029/135GM07" ext-link-type="DOI">10.1029/135GM07</ext-link>, 2003.</mixed-citation></ref>
      <ref id="bib1.bibx27"><label>Kavetski et al.(2011)Kavetski, Fenicia, and Clark</label><mixed-citation>Kavetski, D., Fenicia, F., and Clark, M. P.: Impact of temporal data resolution on parameter inference and model identification in conceptual hydrological modeling: Insights from an experimental catchment, Water Resour. Res., 47, W05501, <ext-link xlink:href="https://doi.org/10.1029/2010WR009525" ext-link-type="DOI">10.1029/2010WR009525</ext-link>, 2011.</mixed-citation></ref>
      <ref id="bib1.bibx28"><label>Kennedy and O'Hagan(2001)</label><mixed-citation>Kennedy, M. C. and O'Hagan, A.: Bayesian Calibration of Computer Models, J. R. Stat. Soc. B, 63, 425–464, <ext-link xlink:href="https://doi.org/10.1111/1467-9868.00294" ext-link-type="DOI">10.1111/1467-9868.00294</ext-link>, 2001.</mixed-citation></ref>
      <ref id="bib1.bibx29"><label>Khorashadi Zadeh et al.(2019)Khorashadi Zadeh, Nossent, Woldegiorgis, Bauwens, and van Griensven</label><mixed-citation>Khorashadi Zadeh, F., Nossent, J., Woldegiorgis, B. T., Bauwens, W., and van Griensven, A.: Impact of measurement error and limited data frequency on parameter estimation and uncertainty quantification, Environ. Model. Softw., 118, 35–47, <ext-link xlink:href="https://doi.org/10.1016/j.envsoft.2019.03.022" ext-link-type="DOI">10.1016/j.envsoft.2019.03.022</ext-link>, 2019.</mixed-citation></ref>
      <ref id="bib1.bibx30"><label>Kullback and Leibler(1951)</label><mixed-citation>Kullback, S. and Leibler, R. A.: On Information and Sufficiency, Ann. Math. Stat., 22, 79–86, <ext-link xlink:href="https://doi.org/10.1214/aoms/1177729694" ext-link-type="DOI">10.1214/aoms/1177729694</ext-link>, 1951.</mixed-citation></ref>
      <ref id="bib1.bibx31"><label>Kumar et al.(2019)Kumar, Carroll, Hartikainen, and Martin</label><mixed-citation>Kumar, R., Carroll, C., Hartikainen, A., and Martin, O.: ArviZ: a unified library for exploratory analysis of Bayesian models in Python, J. Open Source Softw., 4, 1143, <ext-link xlink:href="https://doi.org/10.21105/joss.01143" ext-link-type="DOI">10.21105/joss.01143</ext-link>, 2019.</mixed-citation></ref>
      <ref id="bib1.bibx32"><label>Kumar(2025)</label><mixed-citation>Kumar, V. M.: Bayesian data selection to quantify the value of data for landslide runout calibration, Zenodo [code, data set], <ext-link xlink:href="https://doi.org/10.5281/zenodo.17120721" ext-link-type="DOI">10.5281/zenodo.17120721</ext-link>, 2025.</mixed-citation></ref>
      <ref id="bib1.bibx33"><label>Li et al.(2010)Li, Wang, Liu, Yan, Yu, and Zhang</label><mixed-citation>Li, C., Wang, H., Liu, J., Yan, D., Yu, F., and Zhang, L.: Effect of calibration data series length on performance and optimal parameters of hydrological model, Water Sci. Eng., 3, 378–393, <ext-link xlink:href="https://doi.org/10.3882/j.issn.1674-2370.2010.04.002" ext-link-type="DOI">10.3882/j.issn.1674-2370.2010.04.002</ext-link>, 2010.</mixed-citation></ref>
      <ref id="bib1.bibx34"><label>Mancarella and Hungr(2010)</label><mixed-citation>Mancarella, D. and Hungr, O.: Analysis of run-up of granular avalanches against steep, adverse slopes and protective barriers, Can. Geotech. J., 47, 827–841, <ext-link xlink:href="https://doi.org/10.1139/t09-143" ext-link-type="DOI">10.1139/t09-143</ext-link>, 2010.</mixed-citation></ref>
      <ref id="bib1.bibx35"><label>McDougall(2017)</label><mixed-citation>McDougall, S.: 2014 Canadian Geotechnical Colloquium: Landslide runout analysis – current practice and challenges, Can. Geotech. J., 54, 605–620, <ext-link xlink:href="https://doi.org/10.1139/cgj-2016-0104" ext-link-type="DOI">10.1139/cgj-2016-0104</ext-link>, 2017.</mixed-citation></ref>
      <ref id="bib1.bibx36"><label>McMillan and Clark(2009)</label><mixed-citation>McMillan, H. and Clark, M.: Rainfall-runoff model calibration using informal likelihood measures within a Markov chain Monte Carlo sampling scheme, Water Resour. Res., 45, <ext-link xlink:href="https://doi.org/10.1029/2008WR007288" ext-link-type="DOI">10.1029/2008WR007288</ext-link>, 2009.</mixed-citation></ref>
      <ref id="bib1.bibx37"><label>Moretti et al.(2020)Moretti, Mangeney, Walter, Capdeville, Bodin, Stutzmann, and Le Friant</label><mixed-citation>Moretti, L., Mangeney, A., Walter, F., Capdeville, Y., Bodin, T., Stutzmann, E., and Le Friant, A.: Constraining landslide characteristics with Bayesian inversion of field and seismic data, Geophys. J. Int., 221, 1341–1348, <ext-link xlink:href="https://doi.org/10.1093/gji/ggaa056" ext-link-type="DOI">10.1093/gji/ggaa056</ext-link>, 2020.</mixed-citation></ref>
      <ref id="bib1.bibx38"><label>Navarro et al.(2018)Navarro, Le Maître, Hoteit, George, Mandli, and Knio</label><mixed-citation>Navarro, M., Le Maître, O. P., Hoteit, I., George, D. L., Mandli, K. T., and Knio, O. M.: Surrogate-based parameter inference in debris flow model, Comput. Geosci., 22, 1447–1463, <ext-link xlink:href="https://doi.org/10.1007/s10596-018-9765-1" ext-link-type="DOI">10.1007/s10596-018-9765-1</ext-link>, 2018.</mixed-citation></ref>
      <ref id="bib1.bibx39"><label>Oden(2016)</label><mixed-citation>Oden, J. T.: Foundations of Predictive Computational Science, Lecture Notes, CSE 397/EM 397: Special Topics in Computational Science, ICES, The University of Texas at Austin, <uri>https://www.oden.utexas.edu/media/reports/2017/1701.pdf</uri> (last access: 15 September 2025), 2016.</mixed-citation></ref>
      <ref id="bib1.bibx40"><label>Pastor et al.(2012)Pastor, Blanc, Manzanal, Drempetic, Pastor, Sanchez, Crosta, Imposimato, Roddeman et al.</label><mixed-citation>Pastor, M., Blanc, T., Manzanal, D., Drempetic, V., Pastor, M. J., Sánchez, M., Crosta, G., Imposimato, S., Roddeman, D., Foester, E., Kobayashi, H., Delattre, M., and Issler, D.:  Landslide Runout: Review of Analytical/Empirical Models for Subaerial Slides, Submarine Slides and Snow Avalanche. Numerical Modelling. Software Tools, Material Models, Validation and Benchmarking for Selected Case Studies, Deliverable D1.7, Revision 2, SafeLand, EU FP7 Grant Agreement No. 226479, <uri>https://www.ngi.no/globalassets/bilder/prosjekter/safeland/rapporter/d1.7_revised.pdf</uri> (last access: 15 September 2025), 2012.</mixed-citation></ref>
      <ref id="bib1.bibx41"><label>Perkins(2012)</label><mixed-citation>Perkins, S.: Death toll from landslides vastly underestimated, Nature, <ext-link xlink:href="https://doi.org/10.1038/nature.2012.11140" ext-link-type="DOI">10.1038/nature.2012.11140</ext-link>, 2012.</mixed-citation></ref>
      <ref id="bib1.bibx42"><label>Petley(2012)</label><mixed-citation>Petley, D.: Global patterns of loss of life from landslides, Geology, 40, 927–930, <ext-link xlink:href="https://doi.org/10.1130/G33217.1" ext-link-type="DOI">10.1130/G33217.1</ext-link>, 2012.</mixed-citation></ref>
      <ref id="bib1.bibx43"><label>Robert and Casella(2004)</label><mixed-citation> Robert, C. P. and Casella, G.: Monte Carlo Statistical Methods, Springer Texts in Statistics, Springer, New York, 2nd edn., ISBN 978-0387212395, 2004.</mixed-citation></ref>
      <ref id="bib1.bibx44"><label>Seibert et al.(2024)Seibert, Clerc-Schwarzenbach, and van Meerveld</label><mixed-citation>Seibert, J., Clerc-Schwarzenbach, F. M., and van Meerveld, H. J.: Getting your money's worth: Testing the value of data for hydrological model calibration, Hydrol. Process., 38, e15094, <ext-link xlink:href="https://doi.org/10.1002/hyp.15094" ext-link-type="DOI">10.1002/hyp.15094</ext-link>, 2024.</mixed-citation></ref>
      <ref id="bib1.bibx45"><label>Trujillo-Vela et al.(2022)Trujillo-Vela, Ramos-Cañón, Escobar-Vargas, and Galindo-Torres</label><mixed-citation>Trujillo-Vela, M. G., Ramos-Cañón, A. M., Escobar-Vargas, J. A., and Galindo-Torres, S. A.: An overview of debris-flow mathematical modelling, Earth-Sci. Rev., 232, 104135, <ext-link xlink:href="https://doi.org/10.1016/j.earscirev.2022.104135" ext-link-type="DOI">10.1016/j.earscirev.2022.104135</ext-link>, 2022.</mixed-citation></ref>
      <ref id="bib1.bibx46"><label>U.S. Department of Energy(2024)</label><mixed-citation>U.S. Department of Energy: EnergyPlus, <uri>https://energyplus.net/</uri> (last access: 7 June 2025), 2024.</mixed-citation></ref>
      <ref id="bib1.bibx47"><label>Vehtari et al.(2021)Vehtari, Gelman, Simpson, Carpenter, and Bürkner</label><mixed-citation>Vehtari, A., Gelman, A., Simpson, D., Carpenter, B., and Bürkner, P.-C.: Rank-Normalization, Folding, and Localization: An Improved <inline-formula><mml:math id="M312" display="inline"><mml:mover accent="true"><mml:mi>R</mml:mi><mml:mo mathvariant="normal" stretchy="true">^</mml:mo></mml:mover></mml:math></inline-formula> for Assessing Convergence of MCMC (with Discussion), Bayesian Anal., 16, 667–718, <ext-link xlink:href="https://doi.org/10.1214/20-BA1221" ext-link-type="DOI">10.1214/20-BA1221</ext-link>, 2021.</mixed-citation></ref>
      <ref id="bib1.bibx48"><label>Wang et al.(2009)Wang, Kulkarni, and Verdú</label><mixed-citation>Wang, Q., Kulkarni, S. R., and Verdú, S.: Divergence Estimation for Multidimensional Densities Via <inline-formula><mml:math id="M313" display="inline"><mml:mi>k</mml:mi></mml:math></inline-formula>-Nearest-Neighbor Distances, IEEE T. Inform. Theory, 55, 2392–2405, <ext-link xlink:href="https://doi.org/10.1109/TIT.2009.2016060" ext-link-type="DOI">10.1109/TIT.2009.2016060</ext-link>, 2009.</mixed-citation></ref>
      <ref id="bib1.bibx49"><label>Wang et al.(2023)Wang, Wang, Lin, and Yang</label><mixed-citation>Wang, X., Wang, Y., Lin, Q., and Yang, X.: Assessing global landslide casualty risk under moderate climate change based on multiple GCM projections, Int. J. Disast. Risk Sc., 14, 751–767, <ext-link xlink:href="https://doi.org/10.1007/s13753-023-00514-w" ext-link-type="DOI">10.1007/s13753-023-00514-w</ext-link>, 2023.</mixed-citation></ref>
      <ref id="bib1.bibx50"><label>Willenberg et al.(2009)Willenberg, Eberhardt, and Loew</label><mixed-citation>Willenberg, H., Eberhardt, E., and Loew, S.: Hazard assessment and runout analysis for an unstable rock slope above an industrial site in the Riviera valley, Switzerland, Landslides, 6, 111–119, <ext-link xlink:href="https://doi.org/10.1007/s10346-009-0146-7" ext-link-type="DOI">10.1007/s10346-009-0146-7</ext-link>, 2009. </mixed-citation></ref>
      <ref id="bib1.bibx51"><label>Wu et al.(2018)Wu, Kozlowski, Meidani, and Shirvan</label><mixed-citation>Wu, X., Kozlowski, T., Meidani, H., and Shirvan, K.: Inverse uncertainty quantification using the modular Bayesian approach based on Gaussian process, Part 1: Theory, Nucl. Eng. Des., 335, 339–355, <ext-link xlink:href="https://doi.org/10.1016/j.nucengdes.2018.06.004" ext-link-type="DOI">10.1016/j.nucengdes.2018.06.004</ext-link>, 2018.</mixed-citation></ref>
      <ref id="bib1.bibx52"><label>Xu and Valocchi(2015)</label><mixed-citation>Xu, T. and Valocchi, A. J.: A Bayesian approach to improved calibration and prediction of groundwater models with structural error, Water Resour. Res., 51, 9290–9311, <ext-link xlink:href="https://doi.org/10.1002/2015WR017912" ext-link-type="DOI">10.1002/2015WR017912</ext-link>, 2015.</mixed-citation></ref>
      <ref id="bib1.bibx53"><label>Xu et al.(2019)Xu, Jin, Sun, Soga, and Zhou</label><mixed-citation>Xu, X., Jin, F., Sun, Q., Soga, K., and Zhou, G. G.: Three-dimensional material point method modeling of runout behaviour of the Hongshiyan landslide, Can. Geotech. J., 56, 1318–1337, <ext-link xlink:href="https://doi.org/10.1139/cgj-2017-0638" ext-link-type="DOI">10.1139/cgj-2017-0638</ext-link>, 2019.</mixed-citation></ref>
      <ref id="bib1.bibx54"><label>Yan et al.(2022)Yan, Cui, Huang, Zhou, Zhang, Yin, Guo, and Hu</label><mixed-citation>Yan, Y., Cui, Y., Huang, X., Zhou, J., Zhang, W., Yin, S., Guo, J., and Hu, S.: Combining seismic signal dynamic inversion and numerical modeling improves landslide process reconstruction, Earth Surf. Dynam., 10, 1233–1252, <ext-link xlink:href="https://doi.org/10.5194/esurf-10-1233-2022" ext-link-type="DOI">10.5194/esurf-10-1233-2022</ext-link>, 2022.</mixed-citation></ref>
      <ref id="bib1.bibx55"><label>Yildiz et al.(2023)Yildiz, Zhao, and Kowalski</label><mixed-citation>Yildiz, A., Zhao, H., and Kowalski, J.: Computationally-feasible uncertainty quantification in model-based landslide risk assessment, Front. Earth Sci., 10, 1032438, <ext-link xlink:href="https://doi.org/10.3389/feart.2022.1032438" ext-link-type="DOI">10.3389/feart.2022.1032438</ext-link>, 2023.</mixed-citation></ref>
      <ref id="bib1.bibx56"><label>Zahra(2010)</label><mixed-citation>Zahra, T.: Quantifying uncertainties in Landslide Runout Modelling, PhD thesis, Univ. Twente, <ext-link xlink:href="https://doi.org/10.13140/RG.2.1.1585.0323" ext-link-type="DOI">10.13140/RG.2.1.1585.0323</ext-link>, 2010.</mixed-citation></ref>
      <ref id="bib1.bibx57"><label>Zhao(2022)</label><mixed-citation>Zhao, H.: PSimPy: Predictive and probabilistic simulation with Python, Git, <uri>https://git.rwth-aachen.de/mbd/psimpy</uri> (last access: 15 September 2025), 2022.</mixed-citation></ref>
      <ref id="bib1.bibx58"><label>Zhao and Kowalski(2022)</label><mixed-citation>Zhao, H. and Kowalski, J.: Bayesian active learning for parameter calibration of landslide run-out models, Landslides, 19, 2033–2045, <ext-link xlink:href="https://doi.org/10.1007/s10346-022-01857-z" ext-link-type="DOI">10.1007/s10346-022-01857-z</ext-link>, 2022.</mixed-citation></ref>

  </ref-list></back>
    <!--<article-title-html>Bayesian data selection to quantify the value of  data for landslide runout calibration</article-title-html>
<abstract-html/>
<ref-html id="bib1.bib1"><label>Aaron(2017)</label><mixed-citation>
      
Aaron, J.: Advancement and calibration of a 3D numerical model for landslide
runout analysis, PhD thesis, Univ. British Columbia, Vancouver, <a href="https://doi.org/10.14288/1.0357191" target="_blank">https://doi.org/10.14288/1.0357191</a>, 2017.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib2"><label>Aaron et al.(2019)Aaron, McDougall, and Nolde</label><mixed-citation>
      
Aaron, J., McDougall, S., and Nolde, N.: Two methodologies to calibrate
landslide runout models, Landslides, 16, 907–920,
<a href="https://doi.org/10.1007/s10346-018-1116-8" target="_blank">https://doi.org/10.1007/s10346-018-1116-8</a>, 2019.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib3"><label>Aaron et al.(2022)Aaron, McDougall, Kowalski, Mitchell, and
Nolde</label><mixed-citation>
      
Aaron, J., McDougall, S., Kowalski, J., Mitchell, A., and Nolde, N.:
Probabilistic prediction of rock avalanche runout using a numerical model,
Landslides, 19, 2853–2869, <a href="https://doi.org/10.1007/s10346-022-01939-y" target="_blank">https://doi.org/10.1007/s10346-022-01939-y</a>, 2022.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib4"><label>Baptista et al.(2022)Baptista, Cao, Chen, Ghattas, Li, Marzouk, and
Oden</label><mixed-citation>
      
Baptista, R., Cao, L., Chen, J., Ghattas, O., Li, F., Marzouk, Y. M., and Oden,
J. T.: Bayesian model calibration for block copolymer self-assembly:
Likelihood-free inference and expected information gain computation via
measure transport, arXiv, <a href="https://doi.org/10.48550/arXiv.2206.11343" target="_blank">https://doi.org/10.48550/arXiv.2206.11343</a>, 2022.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib5"><label>Barros et al.(2009)Barros, Kirby, and Mavris</label><mixed-citation>
      
Barros, P. A., Kirby, M. R., and Mavris, D. N.: A Review of Calibration under Uncertainty within the Environmental Design Space, in: 47th AIAA Aerospace Sciences Meeting including The New Horizons Forum and Aerospace Exposition, American Institute of Aeronautics and Astronautics, Reston, Virginia, <a href="https://doi.org/10.2514/6.2009-974" target="_blank">https://doi.org/10.2514/6.2009-974</a>, 2009.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib6"><label>Brezzi et al.(2016)Brezzi, Gabrieli, Marcato, Pastor, and
Cola</label><mixed-citation>
      
Brezzi, L., Gabrieli, F., Marcato, G., Pastor, M., and Cola, S.: A new data
assimilation procedure to develop a debris flow run-out model, Landslides,
13, 1083–1096, <a href="https://doi.org/10.1007/s10346-015-0625-y" target="_blank">https://doi.org/10.1007/s10346-015-0625-y</a>, 2016.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib7"><label>Chowdhary et al.(2024)Chowdhary, Tong, Stadler, and
Alexanderian</label><mixed-citation>
      
Chowdhary, A., Tong, S., Stadler, G., and Alexanderian, A.: Sensitivity
Analysis of the information gain in infinite-dimensional Bayesian linear
inverse problems, Int. J. Uncertain. Quan., 14,
17–35, <a href="https://doi.org/10.1615/int.j.uncertaintyquantification.2024051416" target="_blank">https://doi.org/10.1615/int.j.uncertaintyquantification.2024051416</a>, 2024.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib8"><label>Cotter(2024)</label><mixed-citation>
      
Cotter, S. L.: Hierarchical Bayesian Data Selection, ACM Trans. Probab. Mach.
Learn., 1, 7, <a href="https://doi.org/10.1145/3699721" target="_blank">https://doi.org/10.1145/3699721</a>, 2024.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib9"><label>Cui et al.(2015)Cui, Sun, Teng, Song, and Yao</label><mixed-citation>
      
Cui, X., Sun, W., Teng, J., Song, H., and Yao, X.: Effect of length of the observed dataset on the calibration of a distributed hydrological model, Proc. IAHS, 368, 305–311, <a href="https://doi.org/10.5194/piahs-368-305-2015" target="_blank">https://doi.org/10.5194/piahs-368-305-2015</a>, 2015.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib10"><label>Ekmekcioğlu et al.(2022)Ekmekcioğlu, Demirel, and
Booij</label><mixed-citation>
      
Ekmekcioğlu, Ö., Demirel, M. C., and Booij, M. J.: Effect of data
length, spin-up period and spatial model resolution on fully distributed
hydrological model calibration in the Moselle basin, Hydrol. Sci. J., 67,
759–772, <a href="https://doi.org/10.1080/02626667.2022.2046754" target="_blank">https://doi.org/10.1080/02626667.2022.2046754</a>, 2022.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib11"><label>Etter et al.(2018)Etter, Strobl, Seibert, and van
Meerveld</label><mixed-citation>
      
Etter, S., Strobl, B., Seibert, J., and van Meerveld, H. J. I.: Value of uncertain streamflow observations for hydrological modelling, Hydrol. Earth Syst. Sci., 22, 5243–5257, <a href="https://doi.org/10.5194/hess-22-5243-2018" target="_blank">https://doi.org/10.5194/hess-22-5243-2018</a>, 2018.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib12"><label>Fischer et al.(2020)Fischer, Kofler, Huber, Fellin, Mergili, and
Oberguggenberger</label><mixed-citation>
      
Fischer, J. T., Kofler, A., Huber, A., Fellin, W., Mergili, M., and
Oberguggenberger, M.: Bayesian inference in snow avalanche simulation with
r.Avaflow, Geosci., 10, 191, <a href="https://doi.org/10.3390/geosciences10050191" target="_blank">https://doi.org/10.3390/geosciences10050191</a>, 2020.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib13"><label>Foreman-Mackey et al.(2013)Foreman-Mackey, Hogg, Lang, and
Goodman</label><mixed-citation>
      
Foreman-Mackey, D., Hogg, D. W., Lang, D., and Goodman, J.: emcee: The MCMC
Hammer, Publ. Astron. Soc. Pac., 125, 306–312, <a href="https://doi.org/10.1086/670067" target="_blank">https://doi.org/10.1086/670067</a>, 2013.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib14"><label>Froese et al.(2012)Froese, Charrière, Humair, Jaboyedoff, and
Pedrazzini</label><mixed-citation>
      
Froese, C. R., Charrière, M., Humair, F., Jaboyedoff, M., and Pedrazzini, A.:
Characterization and management of rockslide hazard at Turtle Mountain,
Alberta, Canada, 310–322, Cambridge Univ. Press, ISBN 9781107002067,
2012.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib15"><label>Gelman et al.(2013)Gelman, Carlin, Stern, Dunson, Vehtari, and
Rubin</label><mixed-citation>
      
Gelman, A., Carlin, J. B., Stern, H. S., Dunson, D. B., Vehtari, A., and Rubin,
D. B.: Bayesian Data Analysis, 3rd edn., Chapman and Hall/CRC,
<a href="https://doi.org/10.1201/b16018" target="_blank">https://doi.org/10.1201/b16018</a>, 2013.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib16"><label>Gu and Berger(2016)</label><mixed-citation>
      
Gu, M. and Berger, J. O.: Parallel partial Gaussian process emulation for
computer models with massive output, Ann. Appl. Stat., 10, 1317–1347,
<a href="https://doi.org/10.1214/16-AOAS934" target="_blank">https://doi.org/10.1214/16-AOAS934</a>, 2016.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib17"><label>Haeusel et al.(2026)Haeusel, Nitzler, Köglmeier, and
Wall</label><mixed-citation>
      
Haeusel, L. J., Nitzler, J., Köglmeier, L. J., and Wall, W. A.:
Multi-physics-enhanced Bayesian inverse analysis: Information gain from
additional fields, Comput. Method. Appl. M.,
452, 118735, <a href="https://doi.org/10.1016/j.cma.2026.118735" target="_blank">https://doi.org/10.1016/j.cma.2026.118735</a>, 2026.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib18"><label>Hartland(2020)</label><mixed-citation>
      
Hartland, N.: KL-divergence-estimators, GitHub [code],
<a href="https://github.com/nhartland/KL-divergence-estimators" target="_blank"/> (last access: 15 September 2025), 2020.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib19"><label>Heo et al.(2015)Heo, Graziano, Guzowski, and Muehleisen</label><mixed-citation>
      
Heo, Y., Graziano, D. J., Guzowski, L., and Muehleisen, R. T.: Evaluation of
calibration efficacy under different levels of uncertainty, J. Build.
Perform. Simu., 8, 135–144, <a href="https://doi.org/10.1080/19401493.2014.896947" target="_blank">https://doi.org/10.1080/19401493.2014.896947</a>, 2015.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib20"><label>Heredia et al.(2020)Heredia, Eckert, Prieur, and
Thibert</label><mixed-citation>
      
Heredia, M. B., Eckert, N., Prieur, C., and Thibert, E.: Bayesian calibration
of an avalanche model from autocorrelated measurements along the flow:
Application to velocities extracted from photogrammetric images, J. Glaciol.,
66, 373–385, <a href="https://doi.org/10.1017/jog.2020.11" target="_blank">https://doi.org/10.1017/jog.2020.11</a>, 2020.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib21"><label>Hergarten(2024)</label><mixed-citation>
      
Hergarten, S.: Scaling between volume and runout of rock avalanches explained by a modified Voellmy rheology, Earth Surf. Dynam., 12, 219–229, <a href="https://doi.org/10.5194/esurf-12-219-2024" target="_blank">https://doi.org/10.5194/esurf-12-219-2024</a>, 2024.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib22"><label>Huber et al.(2023)Huber, Georgia, and Finley</label><mixed-citation>
      
Huber, H. A., Georgia, S. K., and Finley, S. D.: Systematic Bayesian posterior
analysis guided by Kullback-Leibler divergence facilitates hypothesis
formation, J. Theor. Biol., 558, 111341,
<a href="https://doi.org/10.1016/j.jtbi.2022.111341" target="_blank">https://doi.org/10.1016/j.jtbi.2022.111341</a>, 2023.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib23"><label>Hungr(1995)</label><mixed-citation>
      
Hungr, O.: A model for the runout analysis of rapid flow slides, debris flows
and avalanches, Can. Geotech. J., 32, 610–623, <a href="https://doi.org/10.1139/t95-063" target="_blank">https://doi.org/10.1139/t95-063</a>, 1995.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib24"><label>Hungr and McDougall(2009)</label><mixed-citation>
      
Hungr, O. and McDougall, S.: Two numerical models for landslide dynamic
analysis, Comput. Geosci., 35, 978–992, <a href="https://doi.org/10.1016/j.cageo.2007.12.003" target="_blank">https://doi.org/10.1016/j.cageo.2007.12.003</a>,
2009.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib25"><label>Hübl et al.(2009)Hübl, Suda, Proske, Kaitna, and
Scheidl</label><mixed-citation>
      
Hübl, J., Suda, J., Proske, D., Kaitna, R., and Scheidl, C.: Debris Flow
Impact Estimation, in: Proc. 11th Int. Symp. Water Manag. Hydraul. Eng., Faculty of Civil Engineering, edited by: Popovska, C. and Jovanovski, M., University of Ss. Cyril and Methodius, Skopje, Macedonia,
137–148, ISBN 978-9989-2469-7-5,  2009.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib26"><label>Iverson(2003)</label><mixed-citation>
      
Iverson, R. M.: How should mathematical models of geomorphic processes be
judged?, in: Prediction in Geomorphology, vol. 135 of Geophys.
Monogr., Am. Geophys. Union, <a href="https://doi.org/10.1029/135GM07" target="_blank">https://doi.org/10.1029/135GM07</a>, 2003.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib27"><label>Kavetski et al.(2011)Kavetski, Fenicia, and Clark</label><mixed-citation>
      
Kavetski, D., Fenicia, F., and Clark, M. P.: Impact of temporal data resolution
on parameter inference and model identification in conceptual hydrological
modeling: Insights from an experimental catchment, Water Resour. Res., 47,
W05501, <a href="https://doi.org/10.1029/2010WR009525" target="_blank">https://doi.org/10.1029/2010WR009525</a>, 2011.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib28"><label>Kennedy and O'Hagan(2001)</label><mixed-citation>
      
Kennedy, M. C. and O'Hagan, A.: Bayesian Calibration of Computer Models, J. R. Stat. Soc. B, 63, 425–464, <a href="https://doi.org/10.1111/1467-9868.00294" target="_blank">https://doi.org/10.1111/1467-9868.00294</a>, 2001.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib29"><label>Khorashadi Zadeh et al.(2019)Khorashadi Zadeh, Nossent, Woldegiorgis,
Bauwens, and van Griensven</label><mixed-citation>
      
Khorashadi Zadeh, F., Nossent, J., Woldegiorgis, B. T., Bauwens, W., and van
Griensven, A.: Impact of measurement error and limited data frequency on
parameter estimation and uncertainty quantification, Environ. Model. Softw.,
118, 35–47, <a href="https://doi.org/10.1016/j.envsoft.2019.03.022" target="_blank">https://doi.org/10.1016/j.envsoft.2019.03.022</a>, 2019.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib30"><label>Kullback and Leibler(1951)</label><mixed-citation>
      
Kullback, S. and Leibler, R. A.: On Information and Sufficiency, Ann. Math. Stat., 22, 79–86, <a href="https://doi.org/10.1214/aoms/1177729694" target="_blank">https://doi.org/10.1214/aoms/1177729694</a>, 1951.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib31"><label>Kumar et al.(2019)Kumar, Carroll, Hartikainen, and
Martin</label><mixed-citation>
      
Kumar, R., Carroll, C., Hartikainen, A., and Martin, O.: ArviZ: a unified
library for exploratory analysis of Bayesian models in Python, J. Open Source
Softw., 4, 1143, <a href="https://doi.org/10.21105/joss.01143" target="_blank">https://doi.org/10.21105/joss.01143</a>, 2019.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib32"><label>Kumar(2025)</label><mixed-citation>
      
Kumar, V. M.: Bayesian data selection to quantify the value of data for
landslide runout calibration, Zenodo [code, data set],
<a href="https://doi.org/10.5281/zenodo.17120721" target="_blank">https://doi.org/10.5281/zenodo.17120721</a>, 2025.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib33"><label>Li et al.(2010)Li, Wang, Liu, Yan, Yu, and Zhang</label><mixed-citation>
      
Li, C., Wang, H., Liu, J., Yan, D., Yu, F., and Zhang, L.: Effect of
calibration data series length on performance and optimal parameters of
hydrological model, Water Sci. Eng., 3, 378–393,
<a href="https://doi.org/10.3882/j.issn.1674-2370.2010.04.002" target="_blank">https://doi.org/10.3882/j.issn.1674-2370.2010.04.002</a>, 2010.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib34"><label>Mancarella and Hungr(2010)</label><mixed-citation>
      
Mancarella, D. and Hungr, O.: Analysis of run-up of granular avalanches against steep, adverse slopes and protective barriers, Can. Geotech. J., 47, 827–841, <a href="https://doi.org/10.1139/t09-143" target="_blank">https://doi.org/10.1139/t09-143</a>, 2010.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib35"><label>McDougall(2017)</label><mixed-citation>
      
McDougall, S.: 2014 Canadian Geotechnical Colloquium: Landslide runout analysis
– current practice and challenges, Can. Geotech. J., 54, 605–620,
<a href="https://doi.org/10.1139/cgj-2016-0104" target="_blank">https://doi.org/10.1139/cgj-2016-0104</a>, 2017.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib36"><label>McMillan and Clark(2009)</label><mixed-citation>
      
McMillan, H. and Clark, M.: Rainfall-runoff model calibration using informal
likelihood measures within a Markov chain Monte Carlo sampling scheme, Water
Resour. Res., 45, <a href="https://doi.org/10.1029/2008WR007288" target="_blank">https://doi.org/10.1029/2008WR007288</a>, 2009.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib37"><label>Moretti et al.(2020)Moretti, Mangeney, Walter, Capdeville, Bodin,
Stutzmann, and Le Friant</label><mixed-citation>
      
Moretti, L., Mangeney, A., Walter, F., Capdeville, Y., Bodin, T., Stutzmann, E., and Le Friant, A.: Constraining landslide characteristics with Bayesian inversion of field and seismic data, Geophys. J. Int., 221, 1341–1348, <a href="https://doi.org/10.1093/gji/ggaa056" target="_blank">https://doi.org/10.1093/gji/ggaa056</a>, 2020.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib38"><label>Navarro et al.(2018)Navarro, Le Maître, Hoteit, George, Mandli, and
Knio</label><mixed-citation>
      
Navarro, M., Le Maître, O. P., Hoteit, I., George, D. L., Mandli, K. T., and
Knio, O. M.: Surrogate-based parameter inference in debris flow model,
Comput. Geosci., 22, 1447–1463, <a href="https://doi.org/10.1007/s10596-018-9765-1" target="_blank">https://doi.org/10.1007/s10596-018-9765-1</a>, 2018.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib39"><label>Oden(2016)</label><mixed-citation>
      
Oden, J. T.: Foundations of Predictive Computational Science, Lecture Notes, CSE 397/EM 397: Special Topics in Computational Science, ICES, The University of Texas at Austin, <a href="https://www.oden.utexas.edu/media/reports/2017/1701.pdf" target="_blank"/> (last access: 15 September 2025), 2016.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib40"><label>Pastor et al.(2012)Pastor, Blanc, Manzanal, Drempetic, Pastor,
Sanchez, Crosta, Imposimato, Roddeman et al.</label><mixed-citation>
      
Pastor, M., Blanc, T., Manzanal, D., Drempetic, V., Pastor, M. J., Sánchez, M., Crosta, G., Imposimato, S., Roddeman, D., Foester, E., Kobayashi, H., Delattre, M., and Issler, D.:  Landslide Runout: Review of Analytical/Empirical Models for Subaerial Slides, Submarine Slides and Snow Avalanche. Numerical Modelling. Software Tools, Material Models, Validation and Benchmarking for Selected Case Studies, Deliverable D1.7, Revision 2, SafeLand, EU FP7 Grant Agreement No. 226479, <a href="https://www.ngi.no/globalassets/bilder/prosjekter/safeland/rapporter/d1.7_revised.pdf" target="_blank"/> (last access: 15 September 2025), 2012.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib41"><label>Perkins(2012)</label><mixed-citation>
      
Perkins, S.: Death toll from landslides vastly underestimated, Nature,
<a href="https://doi.org/10.1038/nature.2012.11140" target="_blank">https://doi.org/10.1038/nature.2012.11140</a>, 2012.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib42"><label>Petley(2012)</label><mixed-citation>
      
Petley, D.: Global patterns of loss of life from landslides, Geology, 40,
927–930, <a href="https://doi.org/10.1130/G33217.1" target="_blank">https://doi.org/10.1130/G33217.1</a>, 2012.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib43"><label>Robert and Casella(2004)</label><mixed-citation>
      
Robert, C. P. and Casella, G.: Monte Carlo Statistical Methods, Springer Texts
in Statistics, Springer, New York, 2nd edn., ISBN 978-0387212395, 2004.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib44"><label>Seibert et al.(2024)Seibert, Clerc-Schwarzenbach, and van
Meerveld</label><mixed-citation>
      
Seibert, J., Clerc-Schwarzenbach, F. M., and van Meerveld, H. J.: Getting your money's worth: Testing the value of data for hydrological model calibration, Hydrol. Process., 38, e15094, <a href="https://doi.org/10.1002/hyp.15094" target="_blank">https://doi.org/10.1002/hyp.15094</a>, 2024.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib45"><label>Trujillo-Vela et al.(2022)Trujillo-Vela, Ramos-Cañón,
Escobar-Vargas, and Galindo-Torres</label><mixed-citation>
      
Trujillo-Vela, M. G., Ramos-Cañón, A. M., Escobar-Vargas, J. A., and Galindo-Torres, S. A.: An overview of debris-flow mathematical modelling, Earth-Sci. Rev., 232, 104135, <a href="https://doi.org/10.1016/j.earscirev.2022.104135" target="_blank">https://doi.org/10.1016/j.earscirev.2022.104135</a>, 2022.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib46"><label>U.S. Department of Energy(2024)</label><mixed-citation>
      
U.S. Department of Energy: EnergyPlus, <a href="https://energyplus.net/" target="_blank"/> (last
access: 7 June 2025), 2024.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib47"><label>Vehtari et al.(2021)Vehtari, Gelman, Simpson, Carpenter, and
Bürkner</label><mixed-citation>
      
Vehtari, A., Gelman, A., Simpson, D., Carpenter, B., and Bürkner, P.-C.:
Rank-Normalization, Folding, and Localization: An Improved <mover accent="true"><i>R</i> <mo form="infix">^</mo> </mover> for
Assessing Convergence of MCMC (with Discussion), Bayesian Anal., 16, 667–718, <a href="https://doi.org/10.1214/20-BA1221" target="_blank">https://doi.org/10.1214/20-BA1221</a>, 2021.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib48"><label>Wang et al.(2009)Wang, Kulkarni, and Verdú</label><mixed-citation>
      
Wang, Q., Kulkarni, S. R., and Verdú, S.: Divergence Estimation for Multidimensional Densities Via <i>k</i>-Nearest-Neighbor Distances, IEEE T. Inform. Theory, 55, 2392–2405, <a href="https://doi.org/10.1109/TIT.2009.2016060" target="_blank">https://doi.org/10.1109/TIT.2009.2016060</a>, 2009.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib49"><label>Wang et al.(2023)Wang, Wang, Lin, and Yang</label><mixed-citation>
      
Wang, X., Wang, Y., Lin, Q., and Yang, X.: Assessing global landslide casualty
risk under moderate climate change based on multiple GCM projections, Int.
J. Disast. Risk Sc., 14, 751–767, <a href="https://doi.org/10.1007/s13753-023-00514-w" target="_blank">https://doi.org/10.1007/s13753-023-00514-w</a>, 2023.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib50"><label>Willenberg et al.(2009)Willenberg, Eberhardt, and
Loew</label><mixed-citation>
      
Willenberg, H., Eberhardt, E., and Loew, S.: Hazard assessment and runout
analysis for an unstable rock slope above an industrial site in the Riviera
valley, Switzerland, Landslides, 6, 111–119,
<a href="https://doi.org/10.1007/s10346-009-0146-7" target="_blank">https://doi.org/10.1007/s10346-009-0146-7</a>, 2009.


    </mixed-citation></ref-html>
<ref-html id="bib1.bib51"><label>Wu et al.(2018)Wu, Kozlowski, Meidani, and Shirvan</label><mixed-citation>
      
Wu, X., Kozlowski, T., Meidani, H., and Shirvan, K.: Inverse uncertainty
quantification using the modular Bayesian approach based on Gaussian
process, Part 1: Theory, Nucl. Eng. Des., 335, 339–355,
<a href="https://doi.org/10.1016/j.nucengdes.2018.06.004" target="_blank">https://doi.org/10.1016/j.nucengdes.2018.06.004</a>, 2018.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib52"><label>Xu and Valocchi(2015)</label><mixed-citation>
      
Xu, T. and Valocchi, A. J.: A Bayesian approach to improved calibration and
prediction of groundwater models with structural error, Water Resour. Res.,
51, 9290–9311, <a href="https://doi.org/10.1002/2015WR017912" target="_blank">https://doi.org/10.1002/2015WR017912</a>, 2015.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib53"><label>Xu et al.(2019)Xu, Jin, Sun, Soga, and Zhou</label><mixed-citation>
      
Xu, X., Jin, F., Sun, Q., Soga, K., and Zhou, G. G.: Three-dimensional material point method modeling of runout behaviour of the Hongshiyan landslide, Can. Geotech. J., 56, 1318–1337, <a href="https://doi.org/10.1139/cgj-2017-0638" target="_blank">https://doi.org/10.1139/cgj-2017-0638</a>, 2019.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib54"><label>Yan et al.(2022)Yan, Cui, Huang, Zhou, Zhang, Yin, Guo, and
Hu</label><mixed-citation>
      
Yan, Y., Cui, Y., Huang, X., Zhou, J., Zhang, W., Yin, S., Guo, J., and Hu, S.: Combining seismic signal dynamic inversion and numerical modeling improves landslide process reconstruction, Earth Surf. Dynam., 10, 1233–1252, <a href="https://doi.org/10.5194/esurf-10-1233-2022" target="_blank">https://doi.org/10.5194/esurf-10-1233-2022</a>, 2022.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib55"><label>Yildiz et al.(2023)Yildiz, Zhao, and Kowalski</label><mixed-citation>
      
Yildiz, A., Zhao, H., and Kowalski, J.: Computationally-feasible uncertainty
quantification in model-based landslide risk assessment, Front. Earth Sci.,
10, 1032438, <a href="https://doi.org/10.3389/feart.2022.1032438" target="_blank">https://doi.org/10.3389/feart.2022.1032438</a>, 2023.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib56"><label>Zahra(2010)</label><mixed-citation>
      
Zahra, T.: Quantifying uncertainties in Landslide Runout Modelling, PhD
thesis, Univ. Twente, <a href="https://doi.org/10.13140/RG.2.1.1585.0323" target="_blank">https://doi.org/10.13140/RG.2.1.1585.0323</a>, 2010.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib57"><label>Zhao(2022)</label><mixed-citation>
      
Zhao, H.: PSimPy: Predictive and probabilistic simulation with Python, Git,
<a href="https://git.rwth-aachen.de/mbd/psimpy" target="_blank"/> (last access: 15 September 2025), 2022.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib58"><label>Zhao and Kowalski(2022)</label><mixed-citation>
      
Zhao, H. and Kowalski, J.: Bayesian active learning for parameter calibration
of landslide run-out models, Landslides, 19, 2033–2045,
<a href="https://doi.org/10.1007/s10346-022-01857-z" target="_blank">https://doi.org/10.1007/s10346-022-01857-z</a>, 2022.

    </mixed-citation></ref-html>--></article>
