<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Pedestrian Dynamics – Notes</title><link>https://pedestriandynamics.org/notes/</link><description>Recent content in Notes on Pedestrian Dynamics</description><generator>Hugo -- gohugo.io</generator><language>en-us</language><atom:link href="https://pedestriandynamics.org/notes/index.xml" rel="self" type="application/rss+xml"/><item><title>Calibrating a pedestrian model against real data, part 2: where it stops</title><link>https://pedestriandynamics.org/notes/dakota-validation/</link><pubDate>Mon, 07 Sep 2026 00:00:00 +0000</pubDate><guid>https://pedestriandynamics.org/notes/dakota-validation/</guid><description>
&lt;div class="hx-overflow-x-auto hx-mt-6 hx-flex hx-rounded-lg hx-border hx-py-2 ltr:hx-pr-4 rtl:hx-pl-4 contrast-more:hx-border-current contrast-more:dark:hx-border-current hx-border-blue-200 hx-bg-blue-100 hx-text-blue-900 dark:hx-border-blue-200/30 dark:hx-bg-blue-900/30 dark:hx-text-blue-200">
&lt;div class="ltr:hx-pl-3 ltr:hx-pr-2 rtl:hx-pr-3 rtl:hx-pl-2">&lt;/div>
&lt;div class="hx-w-full hx-min-w-0 hx-leading-7">
&lt;div class="hx-mt-6 hx-leading-7 first:hx-mt-0">
&lt;strong>Abstract.&lt;/strong> &lt;a href="https://pedestriandynamics.org/notes/dakota-calibration/" >Part 1&lt;/a> calibrated JuPedSim&amp;rsquo;s Collision Free Speed model with Dakota on the Hermes bottleneck experiment and found two parameter sets with near-equal residual norms. This part takes them to two other open experiments without refitting: a 0.5 m gate with the motivation varied between runs, and an entrance without guiding barriers. At the gate one set stalls in most seeds and the other passes it 20 to 40 % too fast. Recalibrating on the gate makes passage reliable but not every observable accurate, and two optimizer starts again disagree. Separate timing parameters absorb part of the low-motivation condition, in-sample, but not the speed. At the unguided entrance every set drains the crowd two to three times faster than the people did. A joint calibration over both experiments is worse than either specialist. The note closes with what we would do differently.
&lt;/div>
&lt;/div>
&lt;/div>
&lt;p>A calibration is only validated within the domain it was tested in. Part 1 ended with two parameter sets, A and B, that reproduce flow, density and speed at five bottleneck widths of 2.4 to 5.0 m mostly within ten percent, though neither passes the stated tolerance everywhere, and that disagree with each other on what the model gets wrong at the widest opening. The desired speed is fixed at the 1.55 m/s measured on the Hermes participants in every set below; carrying that value to other experiments is an assumption. The sets found along the way:&lt;/p>
&lt;table>
&lt;thead>
&lt;tr>
&lt;th>parameter&lt;/th>
&lt;th>default&lt;/th>
&lt;th>set A&lt;/th>
&lt;th>set B&lt;/th>
&lt;th>set C&lt;/th>
&lt;th>set D&lt;/th>
&lt;th>joint&lt;/th>
&lt;/tr>
&lt;/thead>
&lt;tbody>
&lt;tr>
&lt;td>desired speed [m/s]&lt;/td>
&lt;td>1.2&lt;/td>
&lt;td>1.55&lt;/td>
&lt;td>1.55&lt;/td>
&lt;td>1.55&lt;/td>
&lt;td>1.55&lt;/td>
&lt;td>1.55&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>radius [m]&lt;/td>
&lt;td>0.20&lt;/td>
&lt;td>0.13&lt;/td>
&lt;td>0.15&lt;/td>
&lt;td>0.124&lt;/td>
&lt;td>0.101&lt;/td>
&lt;td>0.115&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>time gap [s]&lt;/td>
&lt;td>1.0&lt;/td>
&lt;td>0.81&lt;/td>
&lt;td>0.55&lt;/td>
&lt;td>1.038&lt;/td>
&lt;td>0.958&lt;/td>
&lt;td>0.962&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>neighbor repulsion strength&lt;/td>
&lt;td>8&lt;/td>
&lt;td>2.1&lt;/td>
&lt;td>9.4&lt;/td>
&lt;td>9.01&lt;/td>
&lt;td>1.70&lt;/td>
&lt;td>8.62&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>neighbor repulsion range [m]&lt;/td>
&lt;td>0.10&lt;/td>
&lt;td>0.25&lt;/td>
&lt;td>0.10&lt;/td>
&lt;td>0.062&lt;/td>
&lt;td>0.336&lt;/td>
&lt;td>0.189&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>wall repulsion strength&lt;/td>
&lt;td>5&lt;/td>
&lt;td>5 (fixed)&lt;/td>
&lt;td>5 (fixed)&lt;/td>
&lt;td>2.96&lt;/td>
&lt;td>1.39&lt;/td>
&lt;td>4.69&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>wall repulsion range [m]&lt;/td>
&lt;td>0.02&lt;/td>
&lt;td>0.02 (fixed)&lt;/td>
&lt;td>0.02 (fixed)&lt;/td>
&lt;td>0.072&lt;/td>
&lt;td>0.105&lt;/td>
&lt;td>0.031&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>fitted to&lt;/td>
&lt;td>&lt;/td>
&lt;td>Hermes 2.4, 3.6, 5.0 m&lt;/td>
&lt;td>Hermes 2.4, 3.6, 5.0 m&lt;/td>
&lt;td>CrowdQueue h0&lt;/td>
&lt;td>CrowdQueue h0&lt;/td>
&lt;td>both&lt;/td>
&lt;/tr>
&lt;/tbody>
&lt;/table>
&lt;p>Sets A and B, and sets C and D, are two optimizer starts on the same data. The tools, observables and tolerances are those of part 1: JuPedSim, PedPy on both experiment and simulation, Dakota&amp;rsquo;s surrogate-based optimizer, and a decision rule of 6 % on flow and 10 % on density and speed, chosen as calibration weights and adopted as tolerances after the fact.&lt;/p>
&lt;h2>Step 6 — a second experiment, and where the model stops&lt;span class="hx-absolute -hx-mt-20" id="step-6--a-second-experiment-and-where-the-model-stops">&lt;/span>
&lt;a href="#step-6--a-second-experiment-and-where-the-model-stops" class="subheading-anchor" aria-label="Permalink for this section">&lt;/a>&lt;/h2>&lt;figure>&lt;img src="https://pedestriandynamics.org/notes/dakota-validation/exp2.jpg"
alt="The CrowdQueue experiment, Wuppertal 2018: a run in the 5.6 m corridor seen from above, and the setup with the 0.5 m gate at the origin. Images: Pedestrian Dynamics Data Archive, Forschungszentrum Jülich.">&lt;figcaption>
&lt;p>The CrowdQueue experiment, Wuppertal 2018: a run in the 5.6 m corridor seen from above, and the setup with the 0.5 m gate at the origin. Images: Pedestrian Dynamics Data Archive, Forschungszentrum Jülich.&lt;/p>
&lt;/figcaption>
&lt;/figure>
&lt;p>The second open dataset is the &lt;a href="https://ped.fz-juelich.de/db/doku.php?id=crowdqueue" target="_blank" rel="noopener">CrowdQueue experiment&lt;/a> from Wuppertal 2018 (Adrian et al. 2020): a crowd in front of an entrance again, but through a 0.5 m gate instead of a 2.4 to 5 m opening, in corridors from 1.2 to 5.6 m wide, with 11 to 75 participants and the motivation varied between runs. The archive files carry the full geometry and every trajectory, and each simulation is seeded from the data: agents in view at the first frame start there, and agents that enter the tracked corridor later are injected at the time and place of their first observation. Every tracked participant is simulated. An injection that finds its spot occupied is retried for 10 s and then dropped; this happened to one to four people in a few seeds of the two densest 1.2 m runs, and such a seed cannot count as emptied. Density and speed use the same window in experiment and simulation, the frames between the tenth and ninetieth percent of that run&amp;rsquo;s own crossings. Density is the time average over every frame; speed is averaged only over frames in which the area is occupied. &lt;strong>Whether an observable is conditional on occupancy is part of its definition, and sparse runs expose the difference:&lt;/strong> averaging the zeros of empty frames had once reported 0.12 m/s instead of 0.56 m/s for the sparsest run. The flow is measured at a gate line and density and speed in a 2.2 m² upstream area, so the three values are not a flux identity and should not be forced to satisfy \(q=\rho v\).&lt;/p>
&lt;p>Each run is simulated three times and every seed is shown, because averages hide the thing that matters here. A seed counts as emptied if every expected participant was injected and passed the gate within 120 s, twice the longest experiment. Otherwise it is not emptied, and if it also had an interval of 20 s or more without a crossing while agents remained, it counts as stalled. Flow is the slope of N(t) between the tenth and ninetieth percent of crossings for emptied runs; for runs that did not empty it is the number who passed divided by the time from the first crossing to the end of the run, which is lower by construction.&lt;/p>
&lt;figure>&lt;img src="https://pedestriandynamics.org/notes/dakota-validation/fig10.png"
alt="N(t) at the gate for three baseline runs, experiment against the three seeds of CrowdQueue set C. The simulated slopes are steeper than the measurements. In run 110 each seed passes 62 of 63 tracked participants because one late entrant could not be placed; none is classified as stalled.">&lt;figcaption>
&lt;p>N(t) at the gate for three baseline runs, experiment against the three seeds of CrowdQueue set C. The simulated slopes are steeper than the measurements. In run 110 each seed passes 62 of 63 tracked participants because one late entrant could not be placed; none is classified as stalled.&lt;/p>
&lt;/figcaption>
&lt;/figure>
&lt;p>The transfer test comes first: the two Hermes parameter sets applied to the 21 runs without refitting. Set A, the long time gap with weak long-range repulsion, stalls at the funnel in 46 of 63 seeds and empties 16. Set B, the short gap with strong short-range repulsion, empties 59 of 63 and stalls in none, but where it flows it overshoots the measured gate flow by 22 % on average in the baseline runs and by 37 % in the low-motivation runs. &lt;strong>Of the two Hermes optima with near-equal residual norms on the wide bottleneck, one mostly stalls at a half-metre gate and the other passes it too fast.&lt;/strong> That is the practical meaning of the parameter compensation in part 1: sets with the same fit quality on one experiment behave differently on another.&lt;/p>
&lt;p>Recalibrating on the seven baseline runs at 1.2, 3.4 and 5.6 m, with the wall parameters free because the 1.2 m corridors exercise them, was done from two optimizer starts, and again they disagree. Set C has a three-seed norm of 11.3 on its calibration runs and empties 59 of 63 seeds over all conditions. Set D has a norm of 16.2 and stalls in 14 seeds. Set C&amp;rsquo;s baseline flows are 12 % high on average, with individual runs from 2 % low to 30 % high. Its radius is 0.124 m; set D puts the radius almost at its 0.10 m lower bound. &lt;strong>A second optimizer start changed both the inferred parameters and whether the narrow gate remained passable.&lt;/strong>&lt;/p>
&lt;figure>&lt;img src="https://pedestriandynamics.org/notes/dakota-validation/fig6.png"
alt="CrowdQueue: experiment (bars) against Hermes set B transferred without fitting (blue), the CrowdQueue calibration, set C (orange), and the joint calibration (green), one marker per seed, columns ordered by corridor width with the run number; asterisks mark the runs used in the calibration. Filled circle: the run emptied within 120 s; open circle: not emptied, discharge continuing; open square: stalled. Top: baseline motivation, bottom: low motivation.">&lt;figcaption>
&lt;p>CrowdQueue: experiment (bars) against Hermes set B transferred without fitting (blue), the CrowdQueue calibration, set C (orange), and the joint calibration (green), one marker per seed, columns ordered by corridor width with the run number; asterisks mark the runs used in the calibration. Filled circle: the run emptied within 120 s; open circle: not emptied, discharge continuing; open square: stalled. Top: baseline motivation, bottom: low motivation.&lt;/p>
&lt;/figcaption>
&lt;/figure>
&lt;p>The low-motivation runs show what a flow-only validation would miss. Between the baseline and the low-motivation condition the measured flow at the same gate is about 10 % lower; the model has no input for motivation, so with parameters fixed at the baseline calibration it predicts the same or a higher flow, and overshoots the low-motivation runs by 24 % on average. In the 1.2 m corridor with 24 people, set C gives a flow 29 % high, a front area 44 % denser than measured and people moving 21 % slower: throughput alone hides two compensating errors in the upstream state.&lt;/p>
&lt;h3>What static parameters can absorb&lt;span class="hx-absolute -hx-mt-20" id="what-static-parameters-can-absorb">&lt;/span>
&lt;a href="#what-static-parameters-can-absorb" class="subheading-anchor" aria-label="Permalink for this section">&lt;/a>&lt;/h3>&lt;p>Can a different parameter set stand in for motivation? Freeing all seven parameters is a poor diagnostic because every interaction can compensate for every other one. We therefore freeze the five interaction parameters at set C and profile only desired speed &lt;code>v0&lt;/code> and time gap &lt;code>T&lt;/code>, separately for the baseline condition &lt;code>h0&lt;/code> and the low-motivation condition &lt;code>h−&lt;/code>. These profiles are fitted and evaluated on the same runs, with one-seed grids and Dakota runs that stopped at the iteration cap with their convergence criteria unmet. They are in-sample diagnostics, not a validated motivation model.&lt;/p>
&lt;p>For &lt;code>h0&lt;/code>, Dakota selects (&lt;code>v0&lt;/code>, &lt;code>T&lt;/code>) = (1.17 m/s, 0.79 s) with a three-seed norm of 9.19 over 33 residuals. For &lt;code>h−&lt;/code> it selects (1.50, 1.73) with 11.74. Nearby low cells follow a diagonal valley in each grid, so the pair is constrained jointly more clearly than either coordinate is identified alone, the same degeneracy between desired speed and headway that Kretz et al. (2026) describe for the social force model. Neither profile point transfers to the other condition: the &lt;code>h0&lt;/code> point scores 15.65 on &lt;code>h−&lt;/code> data and the &lt;code>h−&lt;/code> point 13.95 on &lt;code>h0&lt;/code> data. The first profile had stopped at a bound of &lt;code>T = 1.2&lt;/code> s and appeared to show that no pair could represent &lt;code>h−&lt;/code>; widening the box to 2.0 s removed that conclusion. &lt;strong>A bound hit is a prompt for a sensitivity check, not a result.&lt;/strong>&lt;/p>
&lt;p>The larger time gap reduces the residuals in this conditional slice. At the &lt;code>h−&lt;/code> profile point, flow is within 13 % in every run and density within 11 % in seven of nine. It pays for that with speed: seven runs are 24 to 53 % too slow. Part of that residual is the observable rather than the model. The speed is a PedPy individual speed over a 0.4 s window, and people in the front area who move less than 10 cm in 2 s still register 0.09 to 0.13 m/s, a standing-frame floor that plausibly comes from head tracking and body sway. In the 4.5 and 5.6 m corridors the measured speeds of 0.12 to 0.21 m/s are within a factor of two of that floor, so a 10 % tolerance there is tighter than the observable resolves. We have not built an error model for it: the floor is an empirical warning, not a distribution, and we do not know how it varies with density or width or how it covaries with the density measurement. A post-hoc check with the radial speed towards the gate lowers the &lt;code>h−&lt;/code> norm from 11.7 to 9.8 and leaves five runs more than two assumed standard deviations too slow. &lt;strong>The static &lt;code>v0&lt;/code>–&lt;code>T&lt;/code> pair absorbs part of motivation, but it does not reproduce flow, density and speed together across corridor widths, and the speed residuals near the floor are poorly resolved.&lt;/strong>&lt;/p>
&lt;p>The densities become tangible as head counts because the front area is 2.2 m². In the two 1.2 m pairs, the experiment averages 6.1 people for &lt;code>h0&lt;/code> against 3.9 for &lt;code>h−&lt;/code> with 24 participants, and 12.2 against 5.8 with 63. With one common set C, simulation compresses those contrasts to 5.9 against 5.7 and 9.7 against 7.2. After the separate &lt;code>v0&lt;/code>–&lt;code>T&lt;/code> fits, the means become 5.6 against 5.0 and 12.3 against 5.6. The simulations replay each run&amp;rsquo;s observed initial positions and arrivals, so this is conditional evidence: it tests whether the model preserves the supplied positioning difference, not whether it generates that difference from motivation.&lt;/p>
&lt;figure>&lt;img src="https://pedestriandynamics.org/notes/dakota-validation/front_occupancy.png"
alt="People inside the 2.2 m² area immediately upstream of the gate, experiment against three simulation seeds at the condition-specific v0–T profile points. The paired runs have the same corridor width and nominal population. Each simulation is conditioned on that run&amp;rsquo;s observed initial positions and arrivals; the figure is not a generative test of motivation.">&lt;figcaption>
&lt;p>People inside the 2.2 m² area immediately upstream of the gate, experiment against three simulation seeds at the condition-specific v0–T profile points. The paired runs have the same corridor width and nominal population. Each simulation is conditioned on that run&amp;rsquo;s observed initial positions and arrivals; the figure is not a generative test of motivation.&lt;/p>
&lt;/figcaption>
&lt;/figure>
&lt;figure>&lt;img src="https://pedestriandynamics.org/notes/dakota-validation/motivation_pair.gif"
alt="The two 63-person runs in the 1.2 m corridor, tracked experiment beside one simulation seed at the condition-specific v0–T profile point. The dashed rectangle is the front measurement area; the counts below each panel are the people inside it and the people who have passed the gate. Illustration only: the simulation starts from the observed positions, so the sparser low-motivation queue is partly inherited.">&lt;figcaption>
&lt;p>The two 63-person runs in the 1.2 m corridor, tracked experiment beside one simulation seed at the condition-specific v0–T profile point. The dashed rectangle is the front measurement area; the counts below each panel are the people inside it and the people who have passed the gate. Illustration only: the simulation starts from the observed positions, so the sparser low-motivation queue is partly inherited.&lt;/p>
&lt;/figcaption>
&lt;/figure>
&lt;p>We also tested a spacing explanation by fixing &lt;code>v0&lt;/code> and &lt;code>T&lt;/code> at the &lt;code>h0&lt;/code> profile point and varying radius and neighbor range. Its best point, radius 0.150 m and range 0.083 m, has a three-seed norm of 14.4, worse than the timing profile, and most of the surface is not a fit landscape but a stall cliff: from a neighbor range of 0.337 m upward all nine runs stall or fail at every radius, because radius and range govern both upstream separation and access through the same 0.5 m gate. This negative slice does not by itself identify the missing state variable.&lt;/p>
&lt;figure>&lt;img src="https://pedestriandynamics.org/notes/dakota-validation/motivation_profile.png"
alt="One-seed residual profiles with all unshown parameters fixed. Left and centre: desired speed v0 against time gap T for h0 and h−; right: radius against neighbor range for h− with v0 and T fixed at the h0 profile point. Open circles mark grid minima and white crosses the set-C coordinates. On the spacing panel, black dots mark grid points with some stalled or failed runs and black crosses points where all nine were stalled or failed. The common colour scale is clipped at norm 30.">&lt;figcaption>
&lt;p>One-seed residual profiles with all unshown parameters fixed. Left and centre: desired speed v0 against time gap T for h0 and h−; right: radius against neighbor range for h− with v0 and T fixed at the h0 profile point. Open circles mark grid minima and white crosses the set-C coordinates. On the spacing panel, black dots mark grid points with some stalled or failed runs and black crosses points where all nine were stalled or failed. The common colour scale is clipped at norm 30.&lt;/p>
&lt;/figcaption>
&lt;/figure>
&lt;p>The one high-motivation run, 11 people through the 1.2 m corridor at 2.1 persons per second, gets its own sweep over the time gap. A time gap near 0.3 s reaches the measured flow; there the density is about three times the measurement and the speed 30 % low. This condition has one usable run, so the sweep is a diagnostic, not a calibration.&lt;/p>
&lt;figure>&lt;img src="https://pedestriandynamics.org/notes/dakota-validation/fig8.png"
alt="The high-motivation run: flow, density and occupied-frame speed in the corridor against the time gap, all other parameters at set C, three seeds per point shown individually. Filled markers and the line: runs that emptied. Grey band: the measured value with the uncertainty assumed in the calibration weights.">&lt;figcaption>
&lt;p>The high-motivation run: flow, density and occupied-frame speed in the corridor against the time gap, all other parameters at set C, three seeds per point shown individually. Filled markers and the line: runs that emptied. Grey band: the measured value with the uncertainty assumed in the calibration weights.&lt;/p>
&lt;/figcaption>
&lt;/figure>
&lt;p>&lt;strong>With the stated tolerance, recalibration on CrowdQueue makes passage reliable without bringing every observable within tolerance, and parameter sets that fit the wide bottleneck equally well behave very differently at the narrow gate.&lt;/strong> Why the remaining discrepancies are there, whether the round body, the interaction rules or something else, these tests do not say.&lt;/p>
&lt;h2>Step 7 — a third experiment, without barriers&lt;span class="hx-absolute -hx-mt-20" id="step-7--a-third-experiment-without-barriers">&lt;/span>
&lt;a href="#step-7--a-third-experiment-without-barriers" class="subheading-anchor" aria-label="Permalink for this section">&lt;/a>&lt;/h2>&lt;figure>&lt;img src="https://pedestriandynamics.org/notes/dakota-validation/exp3.jpg"
alt="The BaSiGo entrance experiment, Düsseldorf 2013: the crowd in front of the unguided entrance, top right. Image: Pedestrian Dynamics Data Archive, Forschungszentrum Jülich.">&lt;figcaption>
&lt;p>The BaSiGo entrance experiment, Düsseldorf 2013: the crowd in front of the unguided entrance, top right. Image: Pedestrian Dynamics Data Archive, Forschungszentrum Jülich.&lt;/p>
&lt;/figcaption>
&lt;/figure>
&lt;p>The last test uses the &lt;a href="https://ped.fz-juelich.de/da/doku.php?id=entrance_semicircle" target="_blank" rel="noopener">BaSiGo entrance experiment&lt;/a> from 2013 (Sieben et al. 2017): an entrance with two half-metre lanes and no guiding barriers, 319 people told that their favourite artist is playing and they want to be first in, 273 of them tracked. Nothing was fitted to it, and the comparison is a spatial profile rather than a number, the kind of comparison Liao et al. (2017) asked for, here conditional on the observed arrivals. The simulation injects each of the 273 tracked agents at the time and place where the data first see them. The 46 participants who were never tracked are absent, so the simulated crowd is about 15 % smaller than the real one, which changes the interactions in ways we did not quantify.&lt;/p>
&lt;figure>&lt;img src="https://pedestriandynamics.org/notes/dakota-validation/fig7.png"
alt="Time-averaged density in front of the entrance, 20 to 110 s into the run, on a 0.5 m grid: the experiment and each simulation seed separately, with its entrance flow over the same window and the number of people who passed within the 120 s run. Grey: the entrance barriers. Pale yellow is zero density; white lies outside the mapped grid.">&lt;figcaption>
&lt;p>Time-averaged density in front of the entrance, 20 to 110 s into the run, on a 0.5 m grid: the experiment and each simulation seed separately, with its entrance flow over the same window and the number of people who passed within the 120 s run. Grey: the entrance barriers. Pale yellow is zero density; white lies outside the mapped grid.&lt;/p>
&lt;/figcaption>
&lt;/figure>
&lt;p>Hermes set A produces a semicircle, but the wrong one. Its density is concentrated at the entrance and falls off regularly, with peaks around 6 per square metre; the measured crowd is dense to 10 per square metre over a region that extends two to three metres upstream and is visibly asymmetric, heavier to the left of the entrance. The model drains the crowd at 1.3 persons per second where the people, pushing to be first, achieved 0.57; the joint parameters give 1.36 to 1.39 persons per second, with the same regular shape. What the model lacks is a hypothesis, not a finding of this figure: we suspect the pressure of a crowd that wants to be first, and the shoulder rotation that gets real people through half a metre. The maps only show that the shape, the extent and the flow are all wrong, in the same direction for both parameter sets.&lt;/p>
&lt;h2>The joint calibration&lt;span class="hx-absolute -hx-mt-20" id="the-joint-calibration">&lt;/span>
&lt;a href="#the-joint-calibration" class="subheading-anchor" aria-label="Permalink for this section">&lt;/a>&lt;/h2>&lt;p>The last set in the table comes from one Dakota run over both experiments: the three Hermes calibration widths and the seven CrowdQueue baseline runs, 30 residuals in all, each divided by its assumed uncertainty and entering with equal weight, so CrowdQueue carries about 70 % of the objective by count. One seed per run, one optimizer start, &lt;code>efficient_global&lt;/code> with at most 40 iterations; both stopping criteria were met after 17. The result is compared with the specialists on post-hoc residual norms from stored seed means and on how many CrowdQueue seeds empty:&lt;/p>
&lt;table>
&lt;thead>
&lt;tr>
&lt;th>parameter set&lt;/th>
&lt;th>Hermes norm (9 residuals)&lt;/th>
&lt;th>CrowdQueue norm (21 residuals)&lt;/th>
&lt;th>CrowdQueue seeds emptied / stalled / otherwise not emptied&lt;/th>
&lt;/tr>
&lt;/thead>
&lt;tbody>
&lt;tr>
&lt;td>Hermes set A&lt;/td>
&lt;td>2.7&lt;/td>
&lt;td>30.0&lt;/td>
&lt;td>16 / 46 / 1&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>Hermes set B&lt;/td>
&lt;td>2.4&lt;/td>
&lt;td>14.9&lt;/td>
&lt;td>59 / 0 / 4&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>CrowdQueue set C&lt;/td>
&lt;td>&lt;/td>
&lt;td>11.3&lt;/td>
&lt;td>59 / 0 / 4&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>CrowdQueue set D&lt;/td>
&lt;td>&lt;/td>
&lt;td>16.2&lt;/td>
&lt;td>47 / 14 / 2&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>joint&lt;/td>
&lt;td>6.2&lt;/td>
&lt;td>13.6&lt;/td>
&lt;td>57 / 3 / 3&lt;/td>
&lt;/tr>
&lt;/tbody>
&lt;/table>
&lt;p>&lt;strong>The joint point is worse than the relevant specialist in both experiments: 6.2 against 2.4 to 2.7 on Hermes, and 13.6 against 11.3 on CrowdQueue.&lt;/strong> The single search found no parameter set that meets the stated tolerances in both experiments. Whether a better compromise exists elsewhere in the parameter space, one start cannot say.&lt;/p>
&lt;h2>What happened to the calibrated model&lt;span class="hx-absolute -hx-mt-20" id="what-happened-to-the-calibrated-model">&lt;/span>
&lt;a href="#what-happened-to-the-calibrated-model" class="subheading-anchor" aria-label="Permalink for this section">&lt;/a>&lt;/h2>&lt;p>It is worth being plain about it. After Hermes we had two parameter sets that reproduced five bottleneck widths mostly within ten percent, neither passing the tolerance everywhere. At a half-metre gate one of them stalled in 46 of 63 seeds and the other was 20 to 40 % too fast. The set calibrated on the gate passes it in every seed in which everyone could be placed, with flows 12 % high on average. Separate timing parameters reproduce much of the front occupancy contrast between baseline and low motivation, in-sample and conditioned on the observed positions, but not the speed. At the unguided entrance every set tested drains the crowd two to three times faster than the people did. The one parameter that was actually measured, the free speed, was the one the first optimizer run most wanted to change.&lt;/p>
&lt;p>None of this is a verdict on the Collision Free Speed model in particular. It is what cross-regime transfer testing looks like, as opposed to a fit within one regime. Other pedestrian models with a handful of scalar parameters and round agents may show similar limits, but this note tests one model on three datasets and claims no more.&lt;/p>
&lt;h2>Lessons&lt;span class="hx-absolute -hx-mt-20" id="lessons">&lt;/span>
&lt;a href="#lessons" class="subheading-anchor" aria-label="Permalink for this section">&lt;/a>&lt;/h2>&lt;ul>
&lt;li>&lt;strong>A calibration is a statement about a regime, not about a model.&lt;/strong> Parameters fitted to wide bottlenecks are parameters for wide bottlenecks. Using them for a turnstile, a narrow door or an unguided entrance is an extrapolation, and here it was off not by percent but by clogging or not clogging. Report the experiments a parameter set was fitted to and tested on; without that, &amp;ldquo;calibrated&amp;rdquo; is not information.&lt;/li>
&lt;li>&lt;strong>State the purpose, then pick the observables.&lt;/strong> Flow alone would have passed a model with the wrong density, as it did in 2014. Analyse experiment and simulation with the same code, so the observables are defined once.&lt;/li>
&lt;li>&lt;strong>Measured quantities are data, not fitting variables.&lt;/strong> Fixing the free speed at its measured value removed one unphysical value from the fit. It did not make the remaining parameters physical, and the value itself was measured on one experiment&amp;rsquo;s participants and assumed for the others.&lt;/li>
&lt;li>&lt;strong>Run the optimizer more than once.&lt;/strong> Every calibration in this note that was started twice gave two different parameter sets. Multiple starts are a robustness check, not a map of the landscape, and a set reported without one is a sample.&lt;/li>
&lt;li>&lt;strong>Screen before you calibrate.&lt;/strong> The Morris run cost 160 evaluations, removed two of seven parameters, and mapped a region where the model fails outright. Use sensitivity analysis for the grouping of parameters unless you have paid for converged indices.&lt;/li>
&lt;li>&lt;strong>Measure an observable&amp;rsquo;s noise floor before you weight it.&lt;/strong> People standing still register about 0.1 m/s with the speed definition we calibrated against, and the wide-corridor speeds are within a factor of two of that. Estimate the floor first, and do not read residuals near it as well resolved.&lt;/li>
&lt;li>&lt;strong>A fitted parameter shift is not a mechanism.&lt;/strong> The timing profiles absorb part of low motivation and the spacing slice develops a stall cliff. Both are useful diagnostics; neither says what motivation is. Keep parameter uncertainty and seed variability separate, and label as exploratory whatever was chosen after looking at the data.&lt;/li>
&lt;li>&lt;strong>Audit the pipeline before you audit the model.&lt;/strong> Reviews of earlier drafts found late participants missing from simulations, unwritten final trajectory frames, a route around rather than through the gate, empty frames averaged as zero speed, and one script that ran set D while its output was discussed as set C. Each was plausible, produced numbers, and changed a conclusion. Each was corrected and every dependent result recomputed. The drivers now record an evaluation status file for every simulation, and the &lt;a href="https://github.com/PedestrianDynamics/jupedsim-dakota-calibration/blob/39c31e2/README.md#pipeline-corrections-and-safeguards-2026-09-06" target="_blank" rel="noopener">pipeline corrections section of the repository README&lt;/a> lists each error and its fix, pinned to the commit this note was built from.&lt;/li>
&lt;/ul>
&lt;p>Several limitations remain open: one realization per condition, correlated observables, pragmatic weights, one optimizer start for the joint search, no time-step convergence study and no quantified observation error. The next useful work is not a uniquely calibrated parameter set but independent repetitions, more spatial and temporal observables, an explicit motivation state, a pre-specified uncertainty model and verification tests independent of the calibration data. Until then the conclusion is conditional: these scripts demonstrate a reproducible calibration and validation workflow and expose regime failures. They do not establish the validity of the model.&lt;/p>
&lt;h2>A note on open-source software&lt;span class="hx-absolute -hx-mt-20" id="a-note-on-open-source-software">&lt;/span>
&lt;a href="#a-note-on-open-source-software" class="subheading-anchor" aria-label="Permalink for this section">&lt;/a>&lt;/h2>&lt;p>Sargent drew the cost of validation against the confidence it buys as a curve that rises steeply near the end. Everything in this study is about pushing that curve down. The trajectories are on a public archive with a DOI and the geometry attached, the model is open source and scriptable, the analysis library is the same one used on the experiment, and the calibration toolkit has been maintained by a national laboratory for 25 years. Anyone who thinks a remaining discrepancy is a model deficiency can test it this afternoon. And a validation that lives in a text file can be reviewed, which a sentence like &amp;ldquo;parameters were chosen according to the literature&amp;rdquo; never could.&lt;/p>
&lt;h2>References&lt;span class="hx-absolute -hx-mt-20" id="references">&lt;/span>
&lt;a href="#references" class="subheading-anchor" aria-label="Permalink for this section">&lt;/a>&lt;/h2>&lt;ul>
&lt;li>Adrian, J., Seyfried, A., Sieben, A. (2020). Crowds in front of bottlenecks at entrances from the perspective of physics and social psychology. Journal of the Royal Society Interface 17, 20190871. &lt;a href="https://doi.org/10.1098/rsif.2019.0871" target="_blank" rel="noopener">doi:10.1098/rsif.2019.0871&lt;/a>.&lt;/li>
&lt;li>Kretz, T., Eckes, L., Knappik, L., Diz, J., Lipp, D., Müller, S. (2026). Using empirical travel time distributions for calibration of a model of pedestrian dynamics. EURO Journal on Transportation and Logistics 15, 100179. &lt;a href="https://doi.org/10.1016/j.ejtl.2026.100179" target="_blank" rel="noopener">doi:10.1016/j.ejtl.2026.100179&lt;/a>.&lt;/li>
&lt;li>Liao, W., Chraibi, M., Seyfried, A., Zhang, J., Zheng, X., Zhao, Y. (2014). Validation of FDS+Evac for pedestrian simulations in wide bottlenecks. IEEE ITSC 2014, 554–559. &lt;a href="https://doi.org/10.1109/ITSC.2014.6957748" target="_blank" rel="noopener">doi:10.1109/ITSC.2014.6957748&lt;/a>.&lt;/li>
&lt;li>Liao, W., Zhang, J., Zheng, X., Zhao, Y. (2017). A generalized validation procedure for pedestrian models. Simulation Modelling Practice and Theory 77, 20–31. &lt;a href="https://doi.org/10.1016/j.simpat.2017.05.002" target="_blank" rel="noopener">doi:10.1016/j.simpat.2017.05.002&lt;/a>.&lt;/li>
&lt;li>Sargent, R. G. (1984). Simulation model validation. In: Simulation and Model-Based Methodologies: An Integrative View, Springer, 537–555.&lt;/li>
&lt;li>Sieben, A., Schumann, J., Seyfried, A. (2017). Collective phenomena in crowds — where pedestrian dynamics need social psychology. PLOS ONE 12(6), e0177328. &lt;a href="https://doi.org/10.1371/journal.pone.0177328" target="_blank" rel="noopener">doi:10.1371/journal.pone.0177328&lt;/a>.&lt;/li>
&lt;li>Experiment data: Hermes bottleneck experiment, Düsseldorf 2009, &lt;a href="https://doi.org/10.34735/ped.2009.6" target="_blank" rel="noopener">doi:10.34735/ped.2009.6&lt;/a>; CrowdQueue experiment, Wuppertal 2018, &lt;a href="https://doi.org/10.34735/ped.2018.1" target="_blank" rel="noopener">doi:10.34735/ped.2018.1&lt;/a>; BaSiGo entrance experiment, Düsseldorf 2013, &lt;a href="https://doi.org/10.34735/ped.2013.2" target="_blank" rel="noopener">doi:10.34735/ped.2013.2&lt;/a>. Dakota: Adams et al., Sandia National Laboratories, version 6.24.&lt;/li>
&lt;/ul>
&lt;h2>Code&lt;span class="hx-absolute -hx-mt-20" id="code">&lt;/span>
&lt;a href="#code" class="subheading-anchor" aria-label="Permalink for this section">&lt;/a>&lt;/h2>&lt;p>Driver scripts, Dakota input files, results and figures for every step: &lt;a href="https://github.com/PedestrianDynamics/jupedsim-dakota-calibration" target="_blank" rel="noopener">github.com/PedestrianDynamics/jupedsim-dakota-calibration&lt;/a>&lt;/p>
&lt;p>Software used in this study:&lt;/p>
&lt;ul>
&lt;li>&lt;a href="https://dakota.sandia.gov" target="_blank" rel="noopener">Dakota&lt;/a> 6.24&lt;/li>
&lt;li>&lt;a href="https://jupedsim.org" target="_blank" rel="noopener">JuPedSim&lt;/a> 1.4.2&lt;/li>
&lt;li>&lt;a href="https://pedpy.readthedocs.io" target="_blank" rel="noopener">PedPy&lt;/a> 1.4.0&lt;/li>
&lt;/ul>
&lt;p>&lt;span class="hx-inline-block hx-align-text-bottom icon">&lt;svg height=1em xmlns="http://www.w3.org/2000/svg" fill="none" viewBox="0 0 24 24" stroke-width="2" stroke="currentColor" aria-hidden="true">&lt;path stroke-linecap="round" stroke-linejoin="round" d="M11 5H6a2 2 0 00-2 2v11a2 2 0 002 2h11a2 2 0 002-2v-5m-1.414-9.414a2 2 0 112.828 2.828L11.828 15H9v-2.828l8.586-8.586z"/>&lt;/svg>&lt;/span>
By: &lt;a href="https://pedestriandynamics.org/authors/#MohcineChraibi" >Mohcine Chraibi&lt;/a>&lt;/p></description></item><item><title>Calibrating a pedestrian model against real data, part 1: the bottleneck</title><link>https://pedestriandynamics.org/notes/dakota-calibration/</link><pubDate>Sat, 05 Sep 2026 00:00:00 +0000</pubDate><guid>https://pedestriandynamics.org/notes/dakota-calibration/</guid><description>
&lt;div class="hx-overflow-x-auto hx-mt-6 hx-flex hx-rounded-lg hx-border hx-py-2 ltr:hx-pr-4 rtl:hx-pl-4 contrast-more:hx-border-current contrast-more:dark:hx-border-current hx-border-blue-200 hx-bg-blue-100 hx-text-blue-900 dark:hx-border-blue-200/30 dark:hx-bg-blue-900/30 dark:hx-text-blue-200">
&lt;div class="ltr:hx-pl-3 ltr:hx-pr-2 rtl:hx-pr-3 rtl:hx-pl-2">&lt;/div>
&lt;div class="hx-w-full hx-min-w-0 hx-leading-7">
&lt;div class="hx-mt-6 hx-leading-7 first:hx-mt-0">
&lt;strong>Abstract.&lt;/strong> This is the first of two parts. We calibrated JuPedSim&amp;rsquo;s Collision Free Speed model with Dakota against the Hermes bottleneck experiment, three widths in the fit and two held out, with the same open analysis code on experiment and simulation. The defaults miss the measured flow by a factor of two. After calibration, flow, density and speed in front of the opening are mostly within ten percent, including at the two held-out widths, but not everywhere: both calibrated sets overshoot the flow at 3.0 m and undershoot it at 5.0 m. Two optimizer starts give two parameter sets with near-equal residual norms that differ by factors of two to four. The held-out widths are an interpolation test inside one experiment, not an independent validation. &lt;a href="https://pedestriandynamics.org/notes/dakota-validation/" >Part 2&lt;/a> takes the same parameters to a half-metre gate and an unguided entrance, where they fail.
&lt;/div>
&lt;/div>
&lt;/div>
&lt;p>Pedestrian simulations are used to size exits, plan events and argue about safety, so the question whether a model is right is not academic. The simulation community settled the vocabulary long ago. Verification asks whether the software implements the model correctly. Validation asks whether the model, within the domain it is meant for, reproduces reality with an accuracy that is good enough for its purpose (Sargent 1984, 2008; ISO 16730 as summarised by Ronchi et al. 2013). Sargent adds two points that are easy to forget: a model is never valid in the abstract but only for a purpose, and confidence in a model costs money, so one always stops somewhere short of certainty.&lt;/p>
&lt;p>Two lessons from this field&amp;rsquo;s own validation work shape what follows. Liao et al. (2014) calibrated FDS+Evac on the very experiment we use here, by hand, to match the flow at one width. The flow matched. The density in front of the bottleneck was too low and the speed too high, so the model reached the right flow for the wrong reasons. Their follow-up (Liao et al. 2017) turned that into a rule: a single characteristic cannot validate a model, and the comparison has to combine several. Kurtc et al. (2018) drew the second lesson: the assessment should be automated, including the search for parameters, or it will not be done consistently. Kretz et al. (2026) reached the first lesson from the other side with a social force model: a single-parameter fit to the mean flow matched the flow either way, and only the individual travel times told a good calibration from a bad one.&lt;/p>
&lt;p>This note puts the two lessons together with tools that are open and generic. The simulation is &lt;a href="https://jupedsim.org" target="_blank" rel="noopener">JuPedSim&lt;/a>, the analysis of experiment and simulation alike is &lt;a href="https://pedpy.readthedocs.io" target="_blank" rel="noopener">PedPy&lt;/a>, and the screening, sensitivity analysis and calibration are run by &lt;a href="https://dakota.sandia.gov" target="_blank" rel="noopener">Dakota&lt;/a>, Sandia&amp;rsquo;s toolkit for exactly this kind of study. Dakota has been used in engineering for decades and almost never in pedestrian dynamics. It turns &amp;ldquo;which parameters matter, and what values fit the data&amp;rdquo; from a pile of private scripts into a text file.&lt;/p>
&lt;h2>What validation means here&lt;span class="hx-absolute -hx-mt-20" id="what-validation-means-here">&lt;/span>
&lt;a href="#what-validation-means-here" class="subheading-anchor" aria-label="Permalink for this section">&lt;/a>&lt;/h2>&lt;p>Following Sargent, we state the purpose first. &lt;strong>The model should reproduce the flow through wide bottlenecks in a dense, motivated crowd, and it should do so with the right density and speed in front of the bottleneck, not only the right flow.&lt;/strong> Three observables therefore enter every comparison: throughput and the two components of the upstream state. They are the complementary quantities we chose for this study, not a canonical minimum. The decision rule is a fixed relative tolerance per observable, 6 % on flow and 10 % on density and speed, and a comparison passes when the mean over simulation seeds lies within it. The same numbers weight the calibration residuals. They were set before any calibration was run, from the seed-to-seed spread of the simulation plus an allowance for measurement error, but they were chosen as calibration weights. Adopting them as the acceptance tolerance is a decision we made when the results were assessed, not a pre-registered one. &lt;strong>Two of the five bottleneck widths are held out: the optimizer sees 2.4, 3.6 and 5.0 m and never 3.0 and 4.4 m.&lt;/strong> That is an interpolation test within one apparatus and one participant pool. It is a stronger check than a fit, and a weaker one than a second experiment.&lt;/p>
&lt;h2>Step 1 — the experiment&lt;span class="hx-absolute -hx-mt-20" id="step-1--the-experiment">&lt;/span>
&lt;a href="#step-1--the-experiment" class="subheading-anchor" aria-label="Permalink for this section">&lt;/a>&lt;/h2>&lt;figure>&lt;img src="https://pedestriandynamics.org/notes/dakota-calibration/exp1.jpg"
alt="The Hermes bottleneck experiment, Düsseldorf 2009: overhead view of a run, and the setup sketch with the semicircular holding zones at 3 persons per square metre. Images: Pedestrian Dynamics Data Archive, Forschungszentrum Jülich.">&lt;figcaption>
&lt;p>The Hermes bottleneck experiment, Düsseldorf 2009: overhead view of a run, and the setup sketch with the semicircular holding zones at 3 persons per square metre. Images: Pedestrian Dynamics Data Archive, Forschungszentrum Jülich.&lt;/p>
&lt;/figcaption>
&lt;/figure>
&lt;p>The &lt;a href="https://ped.fz-juelich.de/db/doku.php?id=hermes_bottleneck" target="_blank" rel="noopener">Hermes bottleneck experiment&lt;/a> was run in May 2009 in Hall 2 of the Düsseldorf fairground with about 350 participants (Seyfried et al. 2009; Liao et al. 2014). The bottleneck was built from boards higher than two metres, 1 m long, and its width was varied from 2.4 to 5.0 m in five runs, one run per width. Participants waited in a semicircular holding area of radius 8.6 m directly in front of the bottleneck, which puts the initial density at three persons per square metre, and walked through on command. The free walking speed of 42 participants was measured separately: 1.55 m/s with a spread of 0.18 m/s.&lt;/p>
&lt;p>The trajectories are open data with a DOI. Since 2024 they come as HDF5 files that carry a geometry along, so PedPy loads both with one call each. One caveat, found only by reading the paper next to the file: the polygon in the archive is a 12 m box around the camera window with the wall blocks cut off at its edges. In the files downloaded on 6 September 2026, three of its five gaps were also 0.1 m narrower than the run&amp;rsquo;s stated width; the head positions in the trajectories supported the stated width, we reported it, and the archive files published from 8 September 2026 carry corrected polygons with identical trajectories. The simulation takes the gap width from the run parameter and the corridor and holding area from the paper. A corrected polygon for each run is in the repository. The initial positions are not the measured ones either: the simulation places 350 agents on a jittered lattice inside the semicircle at the stated density, so the holding area is an initialization region, not a reconstructed crowd.&lt;/p>
&lt;div class="hextra-code-block hx-relative hx-mt-6 first:hx-mt-0 hx-group/code">
&lt;div>&lt;div class="highlight">&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-python" data-lang="python">&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f92672">import&lt;/span> pedpy&lt;span style="color:#f92672">,&lt;/span> pathlib
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>f &lt;span style="color:#f92672">=&lt;/span> pathlib&lt;span style="color:#f92672">.&lt;/span>Path(&lt;span style="color:#e6db74">&amp;#34;ao-360-400.h5&amp;#34;&lt;/span>)
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>traj &lt;span style="color:#f92672">=&lt;/span> pedpy&lt;span style="color:#f92672">.&lt;/span>load_trajectory_from_ped_data_archive_hdf5(f)
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>area &lt;span style="color:#f92672">=&lt;/span> pedpy&lt;span style="color:#f92672">.&lt;/span>load_walkable_area_from_ped_data_archive_hdf5(f)&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;/div>&lt;div class="hextra-code-copy-btn-container hx-opacity-0 hx-transition group-hover/code:hx-opacity-100 hx-flex hx-gap-1 hx-absolute hx-m-[11px] hx-right-0 hx-top-0">
&lt;button
class="hextra-code-copy-btn hx-group/copybtn hx-transition-all active:hx-opacity-50 hx-bg-primary-700/5 hx-border hx-border-black/5 hx-text-gray-600 hover:hx-text-gray-900 hx-rounded-md hx-p-1.5 dark:hx-bg-primary-300/10 dark:hx-border-white/10 dark:hx-text-gray-400 dark:hover:hx-text-gray-50"
title="Copy code"
>
&lt;div class="copy-icon group-[.copied]/copybtn:hx-hidden hx-pointer-events-none hx-h-4 hx-w-4">&lt;/div>
&lt;div class="success-icon hx-hidden group-[.copied]/copybtn:hx-block hx-pointer-events-none hx-h-4 hx-w-4">&lt;/div>
&lt;/button>
&lt;/div>
&lt;/div>
&lt;p>We measure three things per run: the flow through the bottleneck, from the slope of the N(t) curve at a line across the gap; and the density and mean speed in a 2.8 by 2 m area directly in front of it, averaged over the jam phase. The area is the same for all widths, so it covers more than the 2.4 m opening and only the middle of the 5.0 m one. These are local measurements of the crowd state upstream of the gap, and the calibration compares them like for like. The same three functions are applied to the simulated trajectories. Seyfried and Schadschneider&amp;rsquo;s (2008) warning that the measurement method changes the result is answered the simplest way: there is only one method, applied to both sides.&lt;/p>
&lt;figure>&lt;img src="https://pedestriandynamics.org/notes/dakota-calibration/fig1.png"
alt="Measurement setup for the 2.4 m and 5.0 m runs. Red: the flow line across the gap. Blue: the area for density and speed. The dashed line is the simulation geometry: a 20 m wide corridor with the bottleneck boards as a wall band. The dotted lines mark the camera window of the experiment.">&lt;figcaption>
&lt;p>Measurement setup for the 2.4 m and 5.0 m runs. Red: the flow line across the gap. Blue: the area for density and speed. The dashed line is the simulation geometry: a 20 m wide corridor with the bottleneck boards as a wall band. The dotted lines mark the camera window of the experiment.&lt;/p>
&lt;/figcaption>
&lt;/figure>
&lt;h2>Step 2 — the defaults are wrong by a factor of two&lt;span class="hx-absolute -hx-mt-20" id="step-2--the-defaults-are-wrong-by-a-factor-of-two">&lt;/span>
&lt;a href="#step-2--the-defaults-are-wrong-by-a-factor-of-two" class="subheading-anchor" aria-label="Permalink for this section">&lt;/a>&lt;/h2>&lt;p>We rebuilt the setup in JuPedSim: the corridor, the boards, the semicircular holding area with 350 agents at three per square metre, and an exit line behind the bottleneck. The model is the Collision Free Speed model (Tordeux et al. 2016) with its default parameters.&lt;/p>
&lt;figure>&lt;img src="https://pedestriandynamics.org/notes/dakota-calibration/fig2.png"
alt="Measured N(t) curves (solid) against the default model (dashed), and flow against bottleneck width. The defaults underpredict the flow by a factor of two to three, worst at the widest opening.">&lt;figcaption>
&lt;p>Measured N(t) curves (solid) against the default model (dashed), and flow against bottleneck width. The defaults underpredict the flow by a factor of two to three, worst at the widest opening.&lt;/p>
&lt;/figcaption>
&lt;/figure>
&lt;p>This is not a surprise if you know the model. In the Collision Free Speed model the time gap sets the headway an agent keeps to the one in front, and the radius sets how close agents can pack side by side; together with the neighbor repulsion they fix the throughput of a jammed opening, and the desired speed hardly enters once the crowd is jammed. &lt;strong>With the defaults, one second of headway and a 0.2 m radius, the model delivers a third to a half of the measured flow. No amount of tweaking the desired speed fixes that.&lt;/strong> The question is which of the seven parameters do, and by how much.&lt;/p>
&lt;h2>Step 3 — let Dakota find the parameters that matter&lt;span class="hx-absolute -hx-mt-20" id="step-3--let-dakota-find-the-parameters-that-matter">&lt;/span>
&lt;a href="#step-3--let-dakota-find-the-parameters-that-matter" class="subheading-anchor" aria-label="Permalink for this section">&lt;/a>&lt;/h2>&lt;p>Dakota needs two things: an input file describing the study, and a driver script that turns a parameter vector into responses. Our driver reads Dakota&amp;rsquo;s parameter file, runs one JuPedSim simulation per bottleneck width, computes the three observables with PedPy and writes them back. That is all the coupling there is.&lt;/p>
&lt;div class="hextra-code-block hx-relative hx-mt-6 first:hx-mt-0 hx-group/code">
&lt;pre>&lt;code>method
psuade_moat
samples = 160
partitions = 3
variables
continuous_design = 7
descriptors &amp;#39;desired_speed&amp;#39; &amp;#39;radius&amp;#39; &amp;#39;time_gap&amp;#39; &amp;#39;strength_neighbor&amp;#39;
&amp;#39;range_neighbor&amp;#39; &amp;#39;strength_geometry&amp;#39; &amp;#39;range_geometry&amp;#39;
lower_bounds 0.8 0.12 0.10 2.0 0.02 1.0 0.01
upper_bounds 1.8 0.25 1.20 15.0 0.50 10.0 0.20
interface
fork
analysis_drivers = &amp;#39;python3 driver.py&amp;#39;
parameters_file = &amp;#39;params.in&amp;#39;
results_file = &amp;#39;results.out&amp;#39;
work_directory named &amp;#39;runs/run&amp;#39; directory_tag
asynchronous evaluation_concurrency = 3
responses
response_functions = 9
no_gradients
no_hessians&lt;/code>&lt;/pre>&lt;div class="hextra-code-copy-btn-container hx-opacity-0 hx-transition group-hover/code:hx-opacity-100 hx-flex hx-gap-1 hx-absolute hx-m-[11px] hx-right-0 hx-top-0">
&lt;button
class="hextra-code-copy-btn hx-group/copybtn hx-transition-all active:hx-opacity-50 hx-bg-primary-700/5 hx-border hx-border-black/5 hx-text-gray-600 hover:hx-text-gray-900 hx-rounded-md hx-p-1.5 dark:hx-bg-primary-300/10 dark:hx-border-white/10 dark:hx-text-gray-400 dark:hover:hx-text-gray-50"
title="Copy code"
>
&lt;div class="copy-icon group-[.copied]/copybtn:hx-hidden hx-pointer-events-none hx-h-4 hx-w-4">&lt;/div>
&lt;div class="success-icon hx-hidden group-[.copied]/copybtn:hx-block hx-pointer-events-none hx-h-4 hx-w-4">&lt;/div>
&lt;/button>
&lt;/div>
&lt;/div>
&lt;p>The first stage is a Morris screening: twenty trajectories of eight evaluations each, every step changing one parameter at a time, 160 evaluations in all. Each path starts at a random point of a grid over the parameter box and steps through it changing one parameter at a time; the jump in an observable at each step, divided by the step, is that parameter&amp;rsquo;s elementary effect, and the ranking is the mean absolute effect over the twenty paths. It is cheap and it ranks the parameters per observable. It is a ranking, not a measurement, and it is conditional on the ranges it was sampled over.&lt;/p>
&lt;figure>&lt;img src="https://pedestriandynamics.org/notes/dakota-calibration/morris_sketch.png"
alt="How the screening works. Left: a sketch for two parameters and two paths. Each path is drawn in full before any run: a random start on the grid, a random order of the parameters, a random direction per step. Each arrow is one simulation batch, and the change in flow along it, divided by the step, is one elementary effect. When a run aborts, the path still continues to its next point; only the steps into and out of the failed point are discarded. Right: the real result for the flow at 3.6 m, seven parameters and twenty paths. Steps are measured as fractions of each parameter&amp;rsquo;s range, so the effects of different parameters are comparable. The mean change ranks the parameters; its standard deviation shows how much an effect depends on where in the box the step was taken. Reading the radius as an example: it is the strongest lever, a step of two thirds of its range changed the flow by about 6.7 persons per second on average against a measured 5 per second, and its point lies below the diagonal, so the effect was about the same wherever the other parameters stood. The neighbor range has a similar mean but lies above the diagonal: some steps did nothing and others collapsed the flow. The absolute value hides the sign; for the radius the effect is negative, a larger radius means less flow.">&lt;figcaption>
&lt;p>How the screening works. Left: a sketch for two parameters and two paths. Each path is drawn in full before any run: a random start on the grid, a random order of the parameters, a random direction per step. Each arrow is one simulation batch, and the change in flow along it, divided by the step, is one elementary effect. When a run aborts, the path still continues to its next point; only the steps into and out of the failed point are discarded. Right: the real result for the flow at 3.6 m, seven parameters and twenty paths. Steps are measured as fractions of each parameter&amp;rsquo;s range, so the effects of different parameters are comparable. The mean change ranks the parameters; its standard deviation shows how much an effect depends on where in the box the step was taken. Reading the radius as an example: it is the strongest lever, a step of two thirds of its range changed the flow by about 6.7 persons per second on average against a measured 5 per second, and its point lies below the diagonal, so the effect was about the same wherever the other parameters stood. The neighbor range has a similar mean but lies above the diagonal: some steps did nothing and others collapsed the flow. The absolute value hides the sign; for the radius the effect is negative, a larger radius means less flow.&lt;/p>
&lt;/figcaption>
&lt;/figure>
&lt;figure>&lt;img src="https://pedestriandynamics.org/notes/dakota-calibration/fig3.png"
alt="Morris screening over the full parameter ranges, twenty trajectories. Each cell is the mean absolute elementary effect of a parameter on an observable, scaled to the strongest parameter in that column. Steps where either end failed were discarded, and n per row is the number of valid steps out of the twenty drawn. The time gap, the radius and the neighbor repulsion range dominate; the desired speed has a moderate effect on flow and speed over this wide range; the two wall parameters are weakest.">&lt;figcaption>
&lt;p>Morris screening over the full parameter ranges, twenty trajectories. Each cell is the mean absolute elementary effect of a parameter on an observable, scaled to the strongest parameter in that column. Steps where either end failed were discarded, and n per row is the number of valid steps out of the twenty drawn. The time gap, the radius and the neighbor repulsion range dominate; the desired speed has a moderate effect on flow and speed over this wide range; the two wall parameters are weakest.&lt;/p>
&lt;/figcaption>
&lt;/figure>
&lt;p>&lt;strong>The screening also revealed something no fit would have: in 43 of the 160 evaluations, about a quarter, the model pushed agents through the walls and JuPedSim aborted the run.&lt;/strong> Those combinations have strong, long-range neighbor repulsion and weak wall repulsion. This is a failure of the model&amp;rsquo;s numerical domain, not an observation about pedestrians, but it is one an engineer can land in by hand. We narrowed the bounds accordingly and froze the wall parameters at their defaults.&lt;/p>
&lt;p>With five parameters left, and narrower ranges, a Sobol analysis gives the variance decomposition. Dakota computes the indices itself; the input file changes the method block to &lt;code>sampling&lt;/code> with &lt;code>variance_based_decomp&lt;/code>, the variables become uncertain variables with uniform distributions over the narrowed ranges, and the cost is the base sample size times the number of parameters plus two. Each parameter point uses simulation seed 1; seed variability is measured separately later rather than treated as another physical input.&lt;/p>
&lt;p>How many samples are enough is an empirical question, so we answered it empirically: three replicate runs with 40 base samples, then 80 and 160, 2520 evaluations in all. The magnitudes are not converged: the time gap&amp;rsquo;s total index on flow and speed still rises from about 0.4 to about 0.7 between 80 and 160 samples. &lt;strong>What is stable across every run is the small group of influential parameters per observable:&lt;/strong> time gap and radius for flow, radius and neighbor range for density, time gap and neighbor range for speed, with the desired speed a minor contributor. Since only the grouping is used below, we did not spend further computation on the numbers.&lt;/p>
&lt;figure>&lt;img src="https://pedestriandynamics.org/notes/dakota-calibration/fig4.png"
alt="Sobol indices for the five remaining parameters from the run with 160 base samples, 1120 evaluations, none failed. The time gap dominates flow and speed with total indices of 0.6 to 0.8; the radius dominates density with 0.5 to 0.6, followed by the neighbor repulsion range with about 0.3; the desired speed contributes 0.06 to flow and 0.12 to 0.17 to speed.">&lt;figcaption>
&lt;p>Sobol indices for the five remaining parameters from the run with 160 base samples, 1120 evaluations, none failed. The time gap dominates flow and speed with total indices of 0.6 to 0.8; the radius dominates density with 0.5 to 0.6, followed by the neighbor repulsion range with about 0.3; the desired speed contributes 0.06 to flow and 0.12 to 0.17 to speed.&lt;/p>
&lt;/figcaption>
&lt;/figure>
&lt;figure>&lt;img src="https://pedestriandynamics.org/notes/dakota-calibration/fig9.png"
alt="Sobol total indices against the base sample size. Bars at 40 span three replicate runs with different seeds. The group of influential parameters per observable is the same in every run; the magnitudes are not settled at 160 base samples.">&lt;figcaption>
&lt;p>Sobol total indices against the base sample size. Bars at 40 span three replicate runs with different seeds. The group of influential parameters per observable is the same in every run; the magnitudes are not settled at 160 base samples.&lt;/p>
&lt;/figcaption>
&lt;/figure>
&lt;h2>Step 4 — calibration, and why we did not use gradients&lt;span class="hx-absolute -hx-mt-20" id="step-4--calibration-and-why-we-did-not-use-gradients">&lt;/span>
&lt;a href="#step-4--calibration-and-why-we-did-not-use-gradients" class="subheading-anchor" aria-label="Permalink for this section">&lt;/a>&lt;/h2>&lt;p>Our first attempt at calibration, on a simpler two-room test case before touching the experiment, used Dakota&amp;rsquo;s gradient-based least-squares solver with finite-difference gradients. It stalled at the starting point: two simulations whose parameters differ in the sixth digit produced evacuation times that differed by seconds, and the differences used as gradients were noise. We did not repeat the test on the Hermes objective, so this is an observation about that solver on that test case, not a property of the calibration problem here.&lt;/p>
&lt;p>The method we used instead is Dakota&amp;rsquo;s &lt;code>efficient_global&lt;/code>, an implementation of Efficient Global Optimization (Jones et al. 1998). The objective is a weighted sum of squared residuals over the nine observables, three per calibration width,&lt;/p>
$$
f(\theta) = \sum_{i=1}^{9} w_i \,\bigl(y_i(\theta) - \hat{y}_i\bigr)^2 ,
$$&lt;p>where \(\theta\) is the parameter vector, \(y_i(\theta)\) the simulated observable, \(\hat{y}_i\) the measured one, and \(w_i = 1/\sigma_i^2\) with \(\sigma_i\) the assumed uncertainty, 6 % of the measured flow and 10 % of the measured density and speed; the weights in the input file below are these inverse squared uncertainties. The residual norm quoted throughout is \(\sqrt{f}\), in units of the assumed uncertainty. Each evaluation of \(f\) costs three JuPedSim runs at one fixed seed, so the optimizer sees a single realization of a stochastic simulator; seed variability is measured afterwards, not modelled in the fit. Finite differences of such an objective are noise, so the optimizer never differentiates it. Instead it fits a Gaussian process to all evaluations made so far, which gives a prediction \(\mu(\theta)\) and a standard deviation \(s(\theta)\) at every untried point, and it picks the next point where the expected improvement over the best value \(f_{\min}\) seen so far is largest:&lt;/p>
$$
\mathrm{EI}(\theta) = \bigl(f_{\min} - \mu(\theta)\bigr)\,\Phi(z) + s(\theta)\,\varphi(z),
\qquad z = \frac{f_{\min} - \mu(\theta)}{s(\theta)} ,
$$&lt;p>with \(\Phi\) and \(\varphi\) the standard normal distribution and density functions. The first term rewards points the surrogate expects to be good, the second rewards points where it is uncertain, so the search balances exploiting the current best region against exploring the rest of the box. It stops when the largest expected improvement falls below a threshold, or when the new point is too close to an evaluated one, or at the iteration cap. Each step costs one simulation batch and a surrogate fit, so a calibration takes a few dozen evaluations rather than the hundreds a gradient method would spend on noise.&lt;/p>
&lt;div class="hextra-code-block hx-relative hx-mt-6 first:hx-mt-0 hx-group/code">
&lt;pre>&lt;code>method
efficient_global
max_iterations = 60
responses
calibration_terms = 9
weights = 7.6 3.2 977 3.4 4.3 517 1.45 6.0 252
no_gradients
no_hessians&lt;/code>&lt;/pre>&lt;div class="hextra-code-copy-btn-container hx-opacity-0 hx-transition group-hover/code:hx-opacity-100 hx-flex hx-gap-1 hx-absolute hx-m-[11px] hx-right-0 hx-top-0">
&lt;button
class="hextra-code-copy-btn hx-group/copybtn hx-transition-all active:hx-opacity-50 hx-bg-primary-700/5 hx-border hx-border-black/5 hx-text-gray-600 hover:hx-text-gray-900 hx-rounded-md hx-p-1.5 dark:hx-bg-primary-300/10 dark:hx-border-white/10 dark:hx-text-gray-400 dark:hover:hx-text-gray-50"
title="Copy code"
>
&lt;div class="copy-icon group-[.copied]/copybtn:hx-hidden hx-pointer-events-none hx-h-4 hx-w-4">&lt;/div>
&lt;div class="success-icon hx-hidden group-[.copied]/copybtn:hx-block hx-pointer-events-none hx-h-4 hx-w-4">&lt;/div>
&lt;/button>
&lt;/div>
&lt;/div>
&lt;p>After 42 evaluations, about twenty minutes on a laptop, the optimizer converged with all nine residuals below one standard deviation of the assumed uncertainty. And the desired speed had gone to the lower bound of its range, 0.8 m/s, half the free speed measured on the participants.&lt;/p>
&lt;p>This is the moment where a fit and a validation part ways. The Sobol analysis had already said that the desired speed is a minor influence on the jam observables, so the misfit is flat along it and the optimizer was free to put it wherever the residuals were a hair smaller. &lt;strong>A model with participants walking at 0.8 m/s reproduces the bottleneck. It does not reproduce the participants.&lt;/strong> Sargent calls this data validity: parameters that were measured are data, not fitting variables. Liao et al. (2014) made the same move for FDS+Evac. So we fixed the desired speed at the measured 1.55 m/s and calibrated the remaining four parameters. That converged in 24 evaluations.&lt;/p>
&lt;table>
&lt;thead>
&lt;tr>
&lt;th>parameter&lt;/th>
&lt;th>default&lt;/th>
&lt;th>all five free&lt;/th>
&lt;th>speed fixed, set A&lt;/th>
&lt;th>speed fixed, set B&lt;/th>
&lt;/tr>
&lt;/thead>
&lt;tbody>
&lt;tr>
&lt;td>desired speed [m/s]&lt;/td>
&lt;td>1.2&lt;/td>
&lt;td>0.80 (at bound)&lt;/td>
&lt;td>1.55 (measured)&lt;/td>
&lt;td>1.55 (measured)&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>radius [m]&lt;/td>
&lt;td>0.20&lt;/td>
&lt;td>0.14&lt;/td>
&lt;td>0.13&lt;/td>
&lt;td>0.15&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>time gap [s]&lt;/td>
&lt;td>1.0&lt;/td>
&lt;td>0.56&lt;/td>
&lt;td>0.81&lt;/td>
&lt;td>0.55&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>neighbor repulsion strength&lt;/td>
&lt;td>8&lt;/td>
&lt;td>6.1&lt;/td>
&lt;td>2.1&lt;/td>
&lt;td>9.4&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>neighbor repulsion range [m]&lt;/td>
&lt;td>0.10&lt;/td>
&lt;td>0.15&lt;/td>
&lt;td>0.25&lt;/td>
&lt;td>0.10&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>weighted residual norm (six-seed recalculation)&lt;/td>
&lt;td>&lt;/td>
&lt;td>&lt;/td>
&lt;td>2.7&lt;/td>
&lt;td>2.4&lt;/td>
&lt;/tr>
&lt;/tbody>
&lt;/table>
&lt;p>Set B is the reason for the last column. A surrogate-based optimizer with a few dozen evaluations can settle in one region of a flat landscape, so we repeated the fixed-speed calibration from a different optimizer seed. Dakota reported objective norms of 1.0 and 1.5 at the selected evaluations. Recomputing the norm from the six-seed means used for the later comparison gives 2.7 for set A and 2.4 for set B; those are the values in the table. We have no uncertainty on these norms, the observables are correlated, and we did not sample the objective&amp;rsquo;s seed distribution, so we cannot say that the two sets are statistically equivalent or that the landscape has exactly two optima. What two starts do show is that the reported optimum depends on where the optimizer started. &lt;strong>Reported without a second start, a calibrated parameter set is one sample.&lt;/strong>&lt;/p>
&lt;p>Both sets reproduce flow and density at the three calibration widths to within about ten percent, and speed to within twenty, with residual norms we cannot distinguish. They are nowhere near each other. Time gap, repulsion strength and range shift by factors of two to four between them. Three observables at three widths do not pin four free parameters, and a good fit says nothing about whether the parameters mean what their names say. Fixing the one parameter that was measured removes the freedom to be wrong in that direction. It does not make the others physical: the radius that comes out, 0.13 to 0.15 m, is an effective size in a model with round bodies, and the remaining parameters still compensate for one another.&lt;/p>
&lt;h2>Step 5 — does it hold on the widths it never saw?&lt;span class="hx-absolute -hx-mt-20" id="step-5--does-it-hold-on-the-widths-it-never-saw">&lt;/span>
&lt;a href="#step-5--does-it-hold-on-the-widths-it-never-saw" class="subheading-anchor" aria-label="Permalink for this section">&lt;/a>&lt;/h2>&lt;figure>&lt;img src="https://pedestriandynamics.org/notes/dakota-calibration/bottleneck_pair.gif"
alt="The 3.6 m run, tracked experiment inside the camera window beside one seed of the simulation with set A, cropped to the same window. Both clocks start at the first crossing of the flow line (red); the dashed box is the density and speed area. Illustration only: the simulation starts from a synthetic lattice, not from the measured positions.">&lt;figcaption>
&lt;p>The 3.6 m run, tracked experiment inside the camera window beside one seed of the simulation with set A, cropped to the same window. Both clocks start at the first crossing of the flow line (red); the dashed box is the density and speed area. Illustration only: the simulation starts from a synthetic lattice, not from the measured positions.&lt;/p>
&lt;/figcaption>
&lt;/figure>
&lt;figure>&lt;img src="https://pedestriandynamics.org/notes/dakota-calibration/fig5.png"
alt="Flow, density and speed against bottleneck width: experiment with the acceptance band (6 % on flow, 10 % on density and speed, the uncertainty assumed in the calibration weights), default model, and the two calibrated sets with the desired speed fixed. Error bars are the minimum and maximum over six simulation seeds; the experiment is one run per width. The dotted widths, 3.0 and 4.4 m, were not used in the calibration.">&lt;figcaption>
&lt;p>Flow, density and speed against bottleneck width: experiment with the acceptance band (6 % on flow, 10 % on density and speed, the uncertainty assumed in the calibration weights), default model, and the two calibrated sets with the desired speed fixed. Error bars are the minimum and maximum over six simulation seeds; the experiment is one run per width. The dotted widths, 3.0 and 4.4 m, were not used in the calibration.&lt;/p>
&lt;/figcaption>
&lt;/figure>
&lt;p>With the desired speed fixed, both calibrated sets improve every observable substantially, at the held-out widths too. Applied as a decision rule, the tolerance passes neither set everywhere. Both overshoot the flow at the held-out 3.0 m by 10 to 13 %; set A&amp;rsquo;s speed is 12 % low at the held-out 4.4 m; and at the 5.0 m calibration width both are 10 % low on flow. Where the two sets differ is what else they get wrong at the wide end. Set A follows the measured increase of speed with width up to 3.6 m and then flattens, 17 % low at 5.0 m, while it keeps the density within tolerance. Set B keeps the speed within tolerance at every width and instead lets the density fall 10 % short at 5.0 m. So at the widest opening set A preserves the local density and underpredicts the speed, set B improves the speed and underpredicts the density, and both underpredict the flow. That is a milder version of what Liao et al. (2014) saw in FDS+Evac.&lt;/p>
&lt;p>Deviation of the simulation from the experiment, mean over six seeds, with the band at 6 % for flow and 10 % for density and speed:&lt;/p>
&lt;table>
&lt;thead>
&lt;tr>
&lt;th>width [m]&lt;/th>
&lt;th>flow A / B&lt;/th>
&lt;th>density A / B&lt;/th>
&lt;th>speed A / B&lt;/th>
&lt;/tr>
&lt;/thead>
&lt;tbody>
&lt;tr>
&lt;td>2.4&lt;/td>
&lt;td>+3 % / +5 %&lt;/td>
&lt;td>+7 % / +1 %&lt;/td>
&lt;td>−6 % / −5 %&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>3.0 (held out)&lt;/td>
&lt;td>+10 % / +13 %&lt;/td>
&lt;td>+2 % / −3 %&lt;/td>
&lt;td>+5 % / +9 %&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>3.6&lt;/td>
&lt;td>+3 % / +4 %&lt;/td>
&lt;td>0 % / −6 %&lt;/td>
&lt;td>−3 % / +4 %&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>4.4 (held out)&lt;/td>
&lt;td>+2 % / +4 %&lt;/td>
&lt;td>+8 % / −1 %&lt;/td>
&lt;td>−12 % / +1 %&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>5.0&lt;/td>
&lt;td>−11 % / −10 %&lt;/td>
&lt;td>−1 % / −10 %&lt;/td>
&lt;td>−17 % / −4 %&lt;/td>
&lt;/tr>
&lt;/tbody>
&lt;/table>
&lt;p>The seed-to-seed standard deviation is 2 to 4 % for flow and density and 3 to 6 % for speed, so the differences between A and B at the wide openings are larger than the simulation noise. The experiment, however, is one run per width, and we have not quantified its measurement error; the band is an assumption, not a measured interval.&lt;/p>
&lt;p>&lt;strong>The honest summary of Step 5 is: a substantial improvement over the defaults, and a partial transfer to the held-out widths, but not a model validated for the full 2.4 to 5.0 m range by the stated rule.&lt;/strong> At the calibration widths 2.4 and 3.6 m both sets meet the tolerance on all three observables, which is a fit, not validation. The held-out evidence is two widths: at 4.4 m set B meets the tolerance on all three observables and set A fails it on speed; at 3.0 m both fail it on flow. And both miss the widest calibration opening in a way the two calibrations move between speed and density but neither removes. Whether that is acceptable depends on the purpose, which is why the purpose has to be stated first.&lt;/p>
&lt;h2>What part 1 establishes, and what it does not&lt;span class="hx-absolute -hx-mt-20" id="what-part-1-establishes-and-what-it-does-not">&lt;/span>
&lt;a href="#what-part-1-establishes-and-what-it-does-not" class="subheading-anchor" aria-label="Permalink for this section">&lt;/a>&lt;/h2>&lt;ul>
&lt;li>&lt;strong>A reproducible calibration.&lt;/strong> Every step is a Dakota input file and a driver script in the repository. Anyone can rerun it with another observable, model or experiment from the same archive.&lt;/li>
&lt;li>&lt;strong>The defaults are not usable for this purpose.&lt;/strong> They miss the flow by a factor of two to three, and no single parameter fixes it.&lt;/li>
&lt;li>&lt;strong>Two parameter sets, not one.&lt;/strong> Two optimizer starts give sets with near-equal residual norms that differ by factors of two to four. The data do not identify the parameters; they constrain combinations of them.&lt;/li>
&lt;li>&lt;strong>Partial width transfer.&lt;/strong> Held-out widths inside one experiment are an interpolation test. The model passes it in part. This is not evidence for any other geometry, crowd or motivation.&lt;/li>
&lt;li>&lt;strong>A pipeline audit.&lt;/strong> Between the first draft and this one, reviews found participants missing from simulations, unwritten final trajectory frames and a geometry error. Each was corrected and every dependent result recomputed; the Hermes steady-phase observables changed by less than half a percent. The record is the &lt;a href="https://github.com/PedestrianDynamics/jupedsim-dakota-calibration/blob/39c31e2/README.md#pipeline-corrections-and-safeguards-2026-09-06" target="_blank" rel="noopener">pipeline corrections section of the repository README&lt;/a>, pinned to the commit this note was built from.&lt;/li>
&lt;/ul>
&lt;p>The natural next question is what the two sets do outside this experiment. &lt;a href="https://pedestriandynamics.org/notes/dakota-validation/" >Part 2&lt;/a> takes both to a half-metre gate and to an unguided entrance. One of them stalls, the other is too fast, and a second calibration on the new data disagrees with the first about which parameters matter.&lt;/p>
&lt;h2>References&lt;span class="hx-absolute -hx-mt-20" id="references">&lt;/span>
&lt;a href="#references" class="subheading-anchor" aria-label="Permalink for this section">&lt;/a>&lt;/h2>&lt;ul>
&lt;li>Jones, D. R., Schonlau, M., Welch, W. J. (1998). Efficient global optimization of expensive black-box functions. Journal of Global Optimization 13, 455–492. &lt;a href="https://doi.org/10.1023/A:1008306431147" target="_blank" rel="noopener">doi:10.1023/A:1008306431147&lt;/a>.&lt;/li>
&lt;li>Kretz, T., Eckes, L., Knappik, L., Diz, J., Lipp, D., Müller, S. (2026). Using empirical travel time distributions for calibration of a model of pedestrian dynamics. EURO Journal on Transportation and Logistics 15, 100179. &lt;a href="https://doi.org/10.1016/j.ejtl.2026.100179" target="_blank" rel="noopener">doi:10.1016/j.ejtl.2026.100179&lt;/a>.&lt;/li>
&lt;li>Kurtc, V., Chraibi, M., Tordeux, A. (2018). Automated quality assessment of space-continuous models for pedestrian dynamics. &lt;a href="https://arxiv.org/abs/1809.01862" target="_blank" rel="noopener">arXiv:1809.01862&lt;/a>.&lt;/li>
&lt;li>Liao, W., Chraibi, M., Seyfried, A., Zhang, J., Zheng, X., Zhao, Y. (2014). Validation of FDS+Evac for pedestrian simulations in wide bottlenecks. IEEE ITSC 2014, 554–559. &lt;a href="https://doi.org/10.1109/ITSC.2014.6957748" target="_blank" rel="noopener">doi:10.1109/ITSC.2014.6957748&lt;/a>.&lt;/li>
&lt;li>Liao, W., Zhang, J., Zheng, X., Zhao, Y. (2017). A generalized validation procedure for pedestrian models. Simulation Modelling Practice and Theory 77, 20–31. &lt;a href="https://doi.org/10.1016/j.simpat.2017.05.002" target="_blank" rel="noopener">doi:10.1016/j.simpat.2017.05.002&lt;/a>.&lt;/li>
&lt;li>Ronchi, E., Kuligowski, E. D., Reneke, P. A., Peacock, R. D., Nilsson, D. (2013). The process of verification and validation of building fire evacuation models. NIST Technical Note 1822.&lt;/li>
&lt;li>Sargent, R. G. (1984). Simulation model validation. In: Simulation and Model-Based Methodologies: An Integrative View, Springer, 537–555.&lt;/li>
&lt;li>Sargent, R. G. (2008). Verification and validation of simulation models. Proceedings of the Winter Simulation Conference, 157–169.&lt;/li>
&lt;li>Seyfried, A., Schadschneider, A. (2008). Fundamental diagram and validation of crowd models. ACRI 2008, LNCS 5191, 563–566.&lt;/li>
&lt;li>Seyfried, A., Passon, O., Steffen, B., Boltes, M., Rupprecht, T., Klingsch, W. (2009). New insights into pedestrian flow through bottlenecks. Transportation Science 43, 395–406.&lt;/li>
&lt;li>Tordeux, A., Chraibi, M., Seyfried, A. (2016). Collision-free speed model for pedestrian dynamics. Traffic and Granular Flow &amp;lsquo;15, 225–232.&lt;/li>
&lt;li>Experiment data: Hermes bottleneck experiment, Düsseldorf 2009, &lt;a href="https://doi.org/10.34735/ped.2009.6" target="_blank" rel="noopener">doi:10.34735/ped.2009.6&lt;/a>. Dakota: Adams et al., Sandia National Laboratories, version 6.24.&lt;/li>
&lt;/ul>
&lt;h2>Code&lt;span class="hx-absolute -hx-mt-20" id="code">&lt;/span>
&lt;a href="#code" class="subheading-anchor" aria-label="Permalink for this section">&lt;/a>&lt;/h2>&lt;p>Driver scripts, Dakota input files, results and figures for every step: &lt;a href="https://github.com/PedestrianDynamics/jupedsim-dakota-calibration" target="_blank" rel="noopener">github.com/PedestrianDynamics/jupedsim-dakota-calibration&lt;/a>&lt;/p>
&lt;p>Software used in this study:&lt;/p>
&lt;ul>
&lt;li>&lt;a href="https://dakota.sandia.gov" target="_blank" rel="noopener">Dakota&lt;/a> 6.24&lt;/li>
&lt;li>&lt;a href="https://jupedsim.org" target="_blank" rel="noopener">JuPedSim&lt;/a> 1.4.2&lt;/li>
&lt;li>&lt;a href="https://pedpy.readthedocs.io" target="_blank" rel="noopener">PedPy&lt;/a> 1.4.0&lt;/li>
&lt;/ul>
&lt;p>&lt;span class="hx-inline-block hx-align-text-bottom icon">&lt;svg height=1em xmlns="http://www.w3.org/2000/svg" fill="none" viewBox="0 0 24 24" stroke-width="2" stroke="currentColor" aria-hidden="true">&lt;path stroke-linecap="round" stroke-linejoin="round" d="M11 5H6a2 2 0 00-2 2v11a2 2 0 002 2h11a2 2 0 002-2v-5m-1.414-9.414a2 2 0 112.828 2.828L11.828 15H9v-2.828l8.586-8.586z"/>&lt;/svg>&lt;/span>
By: &lt;a href="https://pedestriandynamics.org/authors/#MohcineChraibi" >Mohcine Chraibi&lt;/a>&lt;/p></description></item><item><title>Why do roaming crowds drift counterclockwise — and where?</title><link>https://pedestriandynamics.org/notes/counterclockwise-crowds/</link><pubDate>Mon, 29 Jun 2026 00:00:00 +0000</pubDate><guid>https://pedestriandynamics.org/notes/counterclockwise-crowds/</guid><description>
&lt;div class="hx-overflow-x-auto hx-mt-6 hx-flex hx-rounded-lg hx-border hx-py-2 ltr:hx-pr-4 rtl:hx-pl-4 contrast-more:hx-border-current contrast-more:dark:hx-border-current hx-border-blue-200 hx-bg-blue-100 hx-text-blue-900 dark:hx-border-blue-200/30 dark:hx-bg-blue-900/30 dark:hx-text-blue-200">
&lt;div class="ltr:hx-pl-3 ltr:hx-pr-2 rtl:hx-pr-3 rtl:hx-pl-2">&lt;/div>
&lt;div class="hx-w-full hx-min-w-0 hx-leading-7">
&lt;div class="hx-mt-6 hx-leading-7 first:hx-mt-0">
&lt;strong>Abstract.&lt;/strong> Echeverría-Huarte et al. report that freely roaming crowds in a circular arena drift counterclockwise and attribute it to an individual tendency to turn left. We rebuilt the 5 m arena in JuPedSim, steering each agent with a roaming rule and measuring the paper&amp;rsquo;s polarization M̄. With symmetric collision avoidance, four different collision models all give M̄ near zero. A turn-left-at-wall rule in about a third of the agents reproduces the reported M̄ of about 0.2. The spatial pattern does not follow: the wall rule concentrates the rotation in a ring at the rim, an always-on left veer spreads it inward and reverses its sign, while the experiment rotates counterclockwise at every radius. Matching the average is easy; matching where the rotation lives is the open problem, and the origin of the individual bias remains unexplained.
&lt;/div>
&lt;/div>
&lt;/div>
&lt;figure>&lt;img src="https://pedestriandynamics.org/notes/counterclockwise-crowds/fig0.png">
&lt;/figure>
&lt;p>Recently a &lt;a href="https://www.nature.com/articles/s41467-026-73713-w" target="_blank" rel="noopener">paper in Nature Communications&lt;/a> (Echeverría-Huart et al.) reports something simple and odd. If you put people in a circular arena and let them walk around freely, the crowd slowly drifts counterclockwise. The authors saw it in experiments in Spain and Japan, so it is not a fluke of one room or one group. Their explanation is that the drift does not come from people avoiding each other. It comes from each person, on their own, having a slight tendency to turn left. (The paper even made it into The New York Times.)&lt;/p>
&lt;p>This sounds intriguing, so we explored it a bit further.&lt;/p>
&lt;p>We used &lt;a href="https://www.jupedsim.org/" target="_blank" rel="noopener">JuPedSim&lt;/a>. It splits the problem into two parts. One part is how people avoid bumping into each other, which JuPedSim handles with a couple of operational models. The other part is where each person wants to go next, which is up to us. We give each agent a desired direction; JuPedSim returns motion without collisions. That split is convenient here. We drive each agent ourselves with direct steering: every step, we pick a heading and place the agent&amp;rsquo;s target just ahead of it.&lt;/p>
&lt;div class="hextra-code-block hx-relative hx-mt-6 first:hx-mt-0 hx-group/code">
&lt;div>&lt;div class="highlight">&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-python" data-lang="python">&lt;span style="display:flex;">&lt;span>sim &lt;span style="color:#f92672">=&lt;/span> jps&lt;span style="color:#f92672">.&lt;/span>Simulation(model&lt;span style="color:#f92672">=&lt;/span>jps&lt;span style="color:#f92672">.&lt;/span>SocialForceModel(), geometry&lt;span style="color:#f92672">=&lt;/span>disk)
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>steering &lt;span style="color:#f92672">=&lt;/span> sim&lt;span style="color:#f92672">.&lt;/span>add_direct_steering_stage()
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>journey &lt;span style="color:#f92672">=&lt;/span> sim&lt;span style="color:#f92672">.&lt;/span>add_journey(jps&lt;span style="color:#f92672">.&lt;/span>JourneyDescription([steering]))
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#66d9ef">while&lt;/span> sim&lt;span style="color:#f92672">.&lt;/span>agent_count() &lt;span style="color:#f92672">&amp;gt;&lt;/span> &lt;span style="color:#ae81ff">0&lt;/span>:
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">for&lt;/span> agent &lt;span style="color:#f92672">in&lt;/span> sim&lt;span style="color:#f92672">.&lt;/span>agents():
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> heading &lt;span style="color:#f92672">=&lt;/span> roam(agent, cfg)
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> agent&lt;span style="color:#f92672">.&lt;/span>target &lt;span style="color:#f92672">=&lt;/span> ahead(agent&lt;span style="color:#f92672">.&lt;/span>position, heading)
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> sim&lt;span style="color:#f92672">.&lt;/span>iterate()&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;/div>&lt;div class="hextra-code-copy-btn-container hx-opacity-0 hx-transition group-hover/code:hx-opacity-100 hx-flex hx-gap-1 hx-absolute hx-m-[11px] hx-right-0 hx-top-0">
&lt;button
class="hextra-code-copy-btn hx-group/copybtn hx-transition-all active:hx-opacity-50 hx-bg-primary-700/5 hx-border hx-border-black/5 hx-text-gray-600 hover:hx-text-gray-900 hx-rounded-md hx-p-1.5 dark:hx-bg-primary-300/10 dark:hx-border-white/10 dark:hx-text-gray-400 dark:hover:hx-text-gray-50"
title="Copy code"
>
&lt;div class="copy-icon group-[.copied]/copybtn:hx-hidden hx-pointer-events-none hx-h-4 hx-w-4">&lt;/div>
&lt;div class="success-icon hx-hidden group-[.copied]/copybtn:hx-block hx-pointer-events-none hx-h-4 hx-w-4">&lt;/div>
&lt;/button>
&lt;/div>
&lt;/div>
&lt;p>The &lt;code>for&lt;/code> loop is our little model. &lt;code>sim.iterate()&lt;/code> is JuPedSim&amp;rsquo;s machinery. Here is one agent following the target we hand it each step:&lt;/p>
&lt;figure>
&lt;video autoplay loop muted playsinline preload="auto" style="max-width: 100%">
&lt;source src="https://pedestriandynamics.org/notes/counterclockwise-crowds/fig1.mp4" type="video/mp4">
&lt;/video>
&lt;figcaption>One roaming agent (blue) and the target we set each step (red ring, dashed line). With a single agent there is nothing to avoid, so this shows only the steering loop.&lt;/figcaption>
&lt;/figure>
&lt;p>The collision model is set on the first line, so we can later run the exact same behavior through different JuPedSim models. This also gives us a clean test of the paper&amp;rsquo;s claim. The collision handling has no built-in left or right preference. So if a crowd with symmetric rules does not rotate, that supports the idea that the bias is individual. And if adding the bias makes the crowd rotate, that supports the proposed mechanism. We model a 5 m arena with agents roaming in random directions. We measure the same quantity as the paper: M, each person&amp;rsquo;s velocity projected onto the counterclockwise direction, averaged over the crowd. M̄ &amp;gt; 0 means the crowd circulates counterclockwise.&lt;/p>
&lt;h2>Symmetric avoidance does not rotate&lt;span class="hx-absolute -hx-mt-20" id="symmetric-avoidance-does-not-rotate">&lt;/span>
&lt;a href="#symmetric-avoidance-does-not-rotate" class="subheading-anchor" aria-label="Permalink for this section">&lt;/a>&lt;/h2>&lt;p>With no bias, the crowd does not rotate. Across crowd sizes and random seeds, M̄ stays near zero. The collision model on its own picks no direction.&lt;/p>
&lt;p>This observation is model agnostic. JuPedSim ships several collision models, some velocity-based and some force-based, and they work quite differently inside. So we ran the same no-bias case through four of them: Social Force, WarpDriver, Collision-Free Speed, and Anticipation Velocity. All four stay within a few hundredths of zero, far below the experimental value. No rotation comes out of symmetric avoidance, whatever model you use. It has to be added.&lt;/p>
&lt;figure>
&lt;video autoplay loop muted playsinline preload="auto" style="max-width: 100%">
&lt;source src="https://pedestriandynamics.org/notes/counterclockwise-crowds/fig2.mp4" type="video/mp4">
&lt;/video>
&lt;figcaption>One real experimental run next to the four model controls. The real crowd circulates counterclockwise; the bare models just mill around.&lt;/figcaption>
&lt;/figure>
&lt;figure>&lt;img src="https://pedestriandynamics.org/notes/counterclockwise-crowds/fig3.png"
alt="The no-bias case for four collision models, against the experimental M̄ (dashed line). All four sit near zero.">&lt;figcaption>
&lt;p>The no-bias case for four collision models, against the experimental M̄ (dashed line). All four sit near zero.&lt;/p>
&lt;/figcaption>
&lt;/figure>
&lt;h2>Adding the bias&lt;span class="hx-absolute -hx-mt-20" id="adding-the-bias">&lt;/span>
&lt;a href="#adding-the-bias" class="subheading-anchor" aria-label="Permalink for this section">&lt;/a>&lt;/h2>&lt;p>The paper&amp;rsquo;s main claim is an always-on individual bias: a slight per-person tendency to turn left, present at every step, even when walking alone with no walls. The specific &amp;ldquo;turn left when facing a wall&amp;rdquo; rule comes from earlier work the authors build on, and is one concrete way that bias could show up. We tried both.&lt;/p>
&lt;p>We first tried putting the bias in free-space curvature, a constant gentle left turn at every step, which is the closest simple stand-in for the always-on bias. On its own that did not give the confined rotation (more on why below). So we also tried the wall-only rule, and bingo, that reproduced the effect.&lt;/p>
&lt;p>How strong the rotation gets depends on how many people have the bias. A single turn strength saturates, so the useful knob is the fraction of people who turn left, which also matches the paper&amp;rsquo;s mixed population. With nobody biased the crowd does not rotate. With everybody biased it rotates strongly. With about a third of the crowd turning left, we get M̄ ≈ 0.2, the experimental value.&lt;/p>
&lt;figure>&lt;img src="https://pedestriandynamics.org/notes/counterclockwise-crowds/fig4.png"
alt="Mean polarization against the share of left-turners. Zero with no bias, rising with the fraction; about a third matches the experiment.">&lt;figcaption>
&lt;p>Mean polarization against the share of left-turners. Zero with no bias, rising with the fraction; about a third matches the experiment.&lt;/p>
&lt;/figcaption>
&lt;/figure>
&lt;figure>&lt;img src="https://pedestriandynamics.org/notes/counterclockwise-crowds/fig5.png"
alt="The distribution of M. The control sits on zero; adding left-turners shifts it counterclockwise, as in the paper.">&lt;figcaption>
&lt;p>The distribution of M. The control sits on zero; adding left-turners shifts it counterclockwise, as in the paper.&lt;/p>
&lt;/figcaption>
&lt;/figure>
&lt;figure>
&lt;video autoplay loop muted playsinline preload="auto" style="max-width: 100%">
&lt;source src="https://pedestriandynamics.org/notes/counterclockwise-crowds/fig6.mp4" type="video/mp4">
&lt;/video>
&lt;figcaption>Three crowds side by side (0%, ~45%, 100% left-turners). The control mills without a net direction; more left-turners make the crowd circulate.&lt;/figcaption>
&lt;/figure>
&lt;h2>Matching the number is not matching the mechanism&lt;span class="hx-absolute -hx-mt-20" id="matching-the-number-is-not-matching-the-mechanism">&lt;/span>
&lt;a href="#matching-the-number-is-not-matching-the-mechanism" class="subheading-anchor" aria-label="Permalink for this section">&lt;/a>&lt;/h2>&lt;p>The paper&amp;rsquo;s data let us go past the single number M̄. Their trajectory files include a per-agent polarization value, and recomputing it with our metric matches theirs exactly. So we can compare not just the average rotation but where in the arena it happens. This was the most useful test, because it can disagree with us instead of just confirming us. (Shout out to the authors for sharing their data on &lt;a href="https://zenodo.org/records/19592341" target="_blank" rel="noopener">Zenodo&lt;/a>!)&lt;/p>
&lt;p>The experiment shows a rotation that fills the disk. It is positive at every radius, weak at the center, strongest in the outer-middle, then easing right at the wall. We compared this against two simple models: the turn-left-at-wall rule above, and an &amp;ldquo;intrinsic&amp;rdquo; version where every agent has a small constant left veer at every step, which is closer to a bias that is present even away from walls.&lt;/p>
&lt;p>Neither matches the experiment, and they fail in different ways. The wall-turn model gets the total amount of rotation right but puts it in the wrong place: a thin spike at the rim and an almost still interior, because the bias only acts at the wall. The intrinsic veer spreads rotation into the interior, closer to where the experiment peaks, but it comes out clockwise, the wrong sign. In open space a left veer does turn counterclockwise, and we checked that directly. But once the arena is closed, the way the veer meets the wall flips the net direction.&lt;/p>
&lt;figure>&lt;img src="https://pedestriandynamics.org/notes/counterclockwise-crowds/fig7.png"
alt="Local rotation against distance from the center, for the experiment and the two models. The experiment is positive throughout and peaks in the outer-middle; the wall-turn model spikes at the rim; the intrinsic veer goes negative.">&lt;figcaption>
&lt;p>Local rotation against distance from the center, for the experiment and the two models. The experiment is positive throughout and peaks in the outer-middle; the wall-turn model spikes at the rim; the intrinsic veer goes negative.&lt;/p>
&lt;/figcaption>
&lt;/figure>
&lt;figure>&lt;img src="https://pedestriandynamics.org/notes/counterclockwise-crowds/fig8.png"
alt="The same three as spatial maps (blue counterclockwise, red clockwise). The experiment is mostly blue across the disk; the wall-turn model is a blue ring; the intrinsic veer is a red band.">&lt;figcaption>
&lt;p>The same three as spatial maps (blue counterclockwise, red clockwise). The experiment is mostly blue across the disk; the wall-turn model is a blue ring; the intrinsic veer is a red band.&lt;/p>
&lt;/figcaption>
&lt;/figure>
&lt;p>We read this as support for the paper. It seems that matching one number with a single knob is easy. The spatial pattern of the real rotation does not fall out of either simple rule, so more work needs to be done on the model.&lt;/p>
&lt;h2>What we learned, and what&amp;rsquo;s open&lt;span class="hx-absolute -hx-mt-20" id="what-we-learned-and-whats-open">&lt;/span>
&lt;a href="#what-we-learned-and-whats-open" class="subheading-anchor" aria-label="Permalink for this section">&lt;/a>&lt;/h2>&lt;p>We set out to test the mechanism behind a real effect, and we learned a few concrete things. Symmetric collision avoidance produces no preferred rotation. A turn-left-at-wall rule recovers counterclockwise motion at roughly the reported strength. And looking at where the rotation lives shows that matching the average is the easy part, while matching the spatial structure is far from being a trivial task.&lt;/p>
&lt;p>There is more to do. The experiment also reports counterclockwise motion with no walls at all, and for people walking alone, which neither of our simple models covers. Reproducing the full spatial field will need a richer individual model than one knob. That leaves the real open question: with no wall to turn at and no crowd to follow, what makes a single person drift counterclockwise? The authors checked the obvious candidates — handedness, footedness, eye dominance — and ruled them all out, calling the cause an individual bias that is &amp;ldquo;most likely biologically rooted&amp;rdquo; and leaving the origin open. So this is a known open problem, worth further research.&lt;/p>
&lt;p>A last thought about the tool. None of this is a simulator telling us the right answer, or even ruling out the wrong ones. It does not work that way, and no simulator does. But what it gave us was a cheap way to ask: what if avoidance is symmetric, what if a third of people turn left, what if the bias is at the wall instead of everywhere. Each one is a few lines in the &lt;code>roam()&lt;/code> rule, run, and looked at. We did not find the mechanism here, but we did sharpen the question and make a couple of stories look more or less likely — and that is most of what a simulation is good for.&lt;/p>
&lt;h2>Code and credits&lt;span class="hx-absolute -hx-mt-20" id="code-and-credits">&lt;/span>
&lt;a href="#code-and-credits" class="subheading-anchor" aria-label="Permalink for this section">&lt;/a>&lt;/h2>&lt;p>Code, full-quality videos, and data: &lt;a href="https://github.com/PedestrianDynamics/clockwise" target="_blank" rel="noopener">github.com/PedestrianDynamics/clockwise&lt;/a>&lt;/p>
&lt;p>Credit for the phenomenon and the experiments goes to Iñaki Echeverría Huarte, Claudio Feliciani, Iker Zuriguel Ballaz and colleagues. There was a charming talk at TGF26 by Iker Zuriguel Ballaz — here is the &lt;a href="https://bpb-eu-w2.wpmucdn.com/blogs.bristol.ac.uk/dist/1/1202/files/2026/05/TGF26_abstract_11.pdf" target="_blank" rel="noopener">abstract&lt;/a>.&lt;/p>
&lt;p>&lt;span class="hx-inline-block hx-align-text-bottom icon">&lt;svg height=1em xmlns="http://www.w3.org/2000/svg" fill="none" viewBox="0 0 24 24" stroke-width="2" stroke="currentColor" aria-hidden="true">&lt;path stroke-linecap="round" stroke-linejoin="round" d="M11 5H6a2 2 0 00-2 2v11a2 2 0 002 2h11a2 2 0 002-2v-5m-1.414-9.414a2 2 0 112.828 2.828L11.828 15H9v-2.828l8.586-8.586z"/>&lt;/svg>&lt;/span>
By: &lt;a href="https://pedestriandynamics.org/authors/#MohcineChraibi" >Mohcine Chraibi&lt;/a>&lt;/p></description></item><item><title>Why don't airlines board planes the optimal way?</title><link>https://pedestriandynamics.org/notes/airplane-boarding/</link><pubDate>Fri, 12 Jun 2026 00:00:00 +0000</pubDate><guid>https://pedestriandynamics.org/notes/airplane-boarding/</guid><description>
&lt;div class="hx-overflow-x-auto hx-mt-6 hx-flex hx-rounded-lg hx-border hx-py-2 ltr:hx-pr-4 rtl:hx-pl-4 contrast-more:hx-border-current contrast-more:dark:hx-border-current hx-border-blue-200 hx-bg-blue-100 hx-text-blue-900 dark:hx-border-blue-200/30 dark:hx-bg-blue-900/30 dark:hx-text-blue-200">
&lt;div class="ltr:hx-pl-3 ltr:hx-pr-2 rtl:hx-pr-3 rtl:hx-pl-2">&lt;/div>
&lt;div class="hx-w-full hx-min-w-0 hx-leading-7">
&lt;div class="hx-mt-6 hx-leading-7 first:hx-mt-0">
&lt;strong>Abstract.&lt;/strong> We rebuilt the classic airplane boarding studies in JuPedSim: a 180-seat single-aisle cabin, six boarding strategies, twenty runs each. With uniform, compliant passengers the ranking from the literature holds. Steffen&amp;rsquo;s interleaved method is fastest, back-to-front is barely better than random, and front-to-back is worst. Lowering the share of passengers who board in their assigned slot, as in Dong et al. (2025), moves the optimized methods toward random boarding, in agreement with their cellular automaton on a different aircraft. Mixed passenger profiles and travel groups boarding together make Steffen&amp;rsquo;s exact order the most fragile method, while a coarser Steffen-style variant stays robust and overtakes it. The benefit of an optimized order depends on passengers following it.
&lt;/div>
&lt;/div>
&lt;/div>
&lt;p>In 2008, astrophysicist Jason Steffen worked out the mathematically optimal way to board an airplane. It boards passengers in a precise interleaved order — window seats first, every other row — so that many people stow their luggage at the same time and nobody ever has to climb over a seated neighbor. In ideal conditions it&amp;rsquo;s dramatically faster than how airlines actually do it. A counter-intuitive part of the result is that boarding back-to-front, which many airlines use, is no better than boarding everyone at random, and front-to-back is worse than random.&lt;/p>
&lt;p>No airline uses Steffen&amp;rsquo;s method. We rebuilt the classic studies in simulation to look at why.&lt;/p>
&lt;h2>Step 1 — reproduce the classic ranking&lt;span class="hx-absolute -hx-mt-20" id="step-1--reproduce-the-classic-ranking">&lt;/span>
&lt;a href="#step-1--reproduce-the-classic-ranking" class="subheading-anchor" aria-label="Permalink for this section">&lt;/a>&lt;/h2>&lt;p>We modelled a single-aisle 180-seat cabin and ran six boarding strategies. Each passenger walks the aisle to their row, holds there while they stow luggage and let neighbors shuffle past (the two things that actually cause boarding delays), then sits. Same geometry, same luggage draws, 20 repetitions per method.&lt;/p>
&lt;figure>
&lt;video autoplay loop muted playsinline preload="auto" style="max-width: 100%">
&lt;source src="https://pedestriandynamics.org/notes/airplane-boarding/fig1.mp4" type="video/mp4">
&lt;/video>
&lt;figcaption>Six methods boarding side by side; the Steffen variants pull ahead while front-to-back jams.&lt;/figcaption>
&lt;/figure>
&lt;p>The ranking comes out as in the literature: Steffen fastest, back-to-front barely better than random, front-to-back worst.&lt;/p>
&lt;figure>&lt;img src="https://pedestriandynamics.org/notes/airplane-boarding/fig2.png"
alt="Boarding time for each method, uniform passengers, 180-seat cabin, 20 runs per method. Each box shows the spread across runs. Steffen-Perfect is fastest and Front-to-Back slowest; Back-to-Front is close to Random.">&lt;figcaption>
&lt;p>Boarding time for each method, uniform passengers, 180-seat cabin, 20 runs per method. Each box shows the spread across runs. Steffen-Perfect is fastest and Front-to-Back slowest; Back-to-Front is close to Random.&lt;/p>
&lt;/figcaption>
&lt;/figure>
&lt;p>This reproduces established results. The next question is the one that matters in practice.&lt;/p>
&lt;h2>Step 2 — why the optimum stays on paper&lt;span class="hx-absolute -hx-mt-20" id="step-2--why-the-optimum-stays-on-paper">&lt;/span>
&lt;a href="#step-2--why-the-optimum-stays-on-paper" class="subheading-anchor" aria-label="Permalink for this section">&lt;/a>&lt;/h2>&lt;p>Steffen&amp;rsquo;s method assumes 180 strangers will form one perfect single-file queue and board in a strict choreographed order. Real passengers don&amp;rsquo;t. Families board together. People show up late. Not everyone follows the plan.&lt;/p>
&lt;p>A very recent paper — Dong, Yanagisawa &amp;amp; Nishinari (2025), Physica A, on boarding the future blended-wing-body aircraft — quantified one of these effects: a compliance rate, the share of passengers who board in their assigned slot. We reproduced their compliance sweep on our single-aisle cabin.&lt;/p>
&lt;figure>
&lt;video autoplay loop muted playsinline preload="auto" style="max-width: 100%">
&lt;source src="https://pedestriandynamics.org/notes/airplane-boarding/fig3.mp4" type="video/mp4">
&lt;/video>
&lt;figcaption>The same method at 100% vs. 50% compliance. The red dots are passengers boarding out of their assigned slot; the 50% panel finishes later.&lt;/figcaption>
&lt;/figure>
&lt;p>The result matches their paper: as compliance drops, the optimized methods lose their advantage and move toward random boarding, while random itself barely changes. At zero compliance every method gives the same time, because there is no longer any order to follow.&lt;/p>
&lt;figure>&lt;img src="https://pedestriandynamics.org/notes/airplane-boarding/fig4.png"
alt="Mean boarding time as compliance falls from 100% (left) to 0% (right), one line per method. As fewer passengers board in their assigned slot, the optimized methods move toward random boarding; Random stays roughly flat and Front-to-Back improves.">&lt;figcaption>
&lt;p>Mean boarding time as compliance falls from 100% (left) to 0% (right), one line per method. As fewer passengers board in their assigned slot, the optimized methods move toward random boarding; Random stays roughly flat and Front-to-Back improves.&lt;/p>
&lt;/figcaption>
&lt;/figure>
&lt;p>Two completely different models — their discrete cellular automaton on a wide multi-aisle aircraft, our continuous pedestrian simulation on a narrow-body — agree on the trend.&lt;/p>
&lt;h2>Step 3 — the optimum is the fragile one&lt;span class="hx-absolute -hx-mt-20" id="step-3--the-optimum-is-the-fragile-one">&lt;/span>
&lt;a href="#step-3--the-optimum-is-the-fragile-one" class="subheading-anchor" aria-label="Permalink for this section">&lt;/a>&lt;/h2>&lt;p>We pushed the same idea two more ways. We gave passengers realistic profiles (fast young travelers, heavy luggage, elderly with reduced mobility, families with kids), and we let travel groups board together instead of in perfect order.&lt;/p>
&lt;figure>
&lt;video autoplay loop muted playsinline preload="auto" style="max-width: 100%">
&lt;source src="https://pedestriandynamics.org/notes/airplane-boarding/fig5.mp4" type="video/mp4">
&lt;/video>
&lt;figcaption>Boarding at 80% travel groups; each color is one party boarding together, which clumps the order the Steffen method relies on.&lt;/figcaption>
&lt;/figure>
&lt;p>In each case, Steffen&amp;rsquo;s &amp;ldquo;perfect&amp;rdquo; method is the most fragile. It is the most sensitive to a mixed passenger crowd, it is the only method that gets slower as more people travel in groups, and it loses its advantage when people do not comply. A coarser, more practical Steffen-style variant, which real passengers could plausibly follow, stays robust and overtakes the optimum.&lt;/p>
&lt;figure>&lt;img src="https://pedestriandynamics.org/notes/airplane-boarding/fig6.png"
alt="Boarding time as travel groups grow; the optimum rises while the practical variant holds flat.">&lt;figcaption>
&lt;p>Boarding time as travel groups grow; the optimum rises while the practical variant holds flat.&lt;/p>
&lt;/figcaption>
&lt;/figure>
&lt;p>The practical takeaway is that the benefit of an optimized boarding order depends almost entirely on passengers following it.&lt;/p>
&lt;h2>Code and credits&lt;span class="hx-absolute -hx-mt-20" id="code-and-credits">&lt;/span>
&lt;a href="#code-and-credits" class="subheading-anchor" aria-label="Permalink for this section">&lt;/a>&lt;/h2>&lt;p>Steffen showed the ranking in 2008 and confirmed it experimentally in 2012; Dong et al. did the robustness analysis in 2025.&lt;/p>
&lt;p>Code, high-quality videos &amp;amp; data: &lt;a href="https://github.com/PedestrianDynamics/boarding" target="_blank" rel="noopener">github.com/PedestrianDynamics/boarding&lt;/a>&lt;/p>
&lt;p>&lt;span class="hx-inline-block hx-align-text-bottom icon">&lt;svg height=1em xmlns="http://www.w3.org/2000/svg" fill="none" viewBox="0 0 24 24" stroke-width="2" stroke="currentColor" aria-hidden="true">&lt;path stroke-linecap="round" stroke-linejoin="round" d="M11 5H6a2 2 0 00-2 2v11a2 2 0 002 2h11a2 2 0 002-2v-5m-1.414-9.414a2 2 0 112.828 2.828L11.828 15H9v-2.828l8.586-8.586z"/>&lt;/svg>&lt;/span>
By: &lt;a href="https://pedestriandynamics.org/authors/#MohcineChraibi" >Mohcine Chraibi&lt;/a>&lt;/p></description></item></channel></rss>