How we study marathon pacing

The study follows the whole race: how runners start, distribute their speed, respond to the course, finish, and improve over time.

Four things we want to understand

What is available now?

The export used by the current analyses contains 4,462,379 race records across 34 cities and 256 race editions. These charts summarize individual splits. The complete records are available in public data downloads without login.

Every published analysis draws from the same collection of marathon results, including sustained slowdown and the supporting charts. The common timing and edition-quality checks retain 3,517,336 eligible finishes. History, exact-age and weather comparisons use smaller subsets because they need additional fields.

33 of the 35 questions have results recalculated from the public export. Some are partial answers: a course comparison cannot isolate the course’s causal effect, and a route proxy cannot establish the hills used in an old edition. Group running and congestion require start and checkpoint clock times that are absent from this export.

A finish is one performance, so a runner can contribute several. Individual tables can have smaller samples because they need different fields. See the marathons, recorded years, weather fields and elevation coverage.

The ten essential analyses

The personalized guide uses 3,517,336 eligible finishes, including 1,225,873 with an exact usable age. The usable earlier-benchmark cohort contains 555,437 finishes. Combining course, age, recorded gender and earlier-time filters can make samples much smaller.

A visitor can choose a marathon, an age group, a threshold from 1:30 to 12:00, and optional recorded gender and previous marathon time. The ten primary analyses are ranked by usefulness to runners and strength of the evidence. Each page shows the controls that apply to its question. The twelve underlying comparisons remain available in the research archive.

All twelve personalized questions use the complete public export. Visitor targets are chosen thresholds, never inferred historical intentions. Targets from 90 to 720 whole minutes are evaluated with strict finish < target. Sparse achieved-time and comparison groups remain unavailable rather than being estimated.

Exact ages define 18–24 then five-year bands through 85–89. Age-group-only labels are not converted into exact ages. Optional recorded gender is Women, Men or all available records. Previous performance selects a 15-minute band of best times in the two strictly earlier calendar years.

Every public result has at least 100 finishes or linked pairs. Age, gender and prior-performance rollups are computed directly from the same records, not by averaging subgroup medians. The site labels any broader comparison used when a narrow combination is unavailable. Availability is not statistical certainty.

Section paces use actual elapsed differences divided by 5 km or 2.195 km at the finish. Quantile ranges describe variation between finishes, not confidence intervals. Higher time per distance means slower. Every complete cohort uses the same runners at all nine checkpoints.

The comparison is observational. Historical route changes and declared goals are absent. Supplied elevation and start-hour weather remain proxies. Current and same-year performances are excluded from prior benchmarks. Individual identity checks follow the linked-history pipeline.

Preparing, choosing a course and reviewing a past result reorder the same twelve questions. Preference changes do not change the underlying evidence. Checkpoint comparisons use current elapsed progress instead of a previous-marathon filter.

Apply the reviewed source-quality edition exclusions for this exact export after the timing checks. Known invalid split grids, incomplete ingestion, unreconciled HOLD editions and a selected top-finisher field do not contribute to the analyses or prior benchmarks. Report source exclusions separately; an already invalid timing row is not counted twice. Missing age or recorded gender alone does not exclude an otherwise eligible finish from the overall cohort. Other sparse editions are not declared incomplete merely from their size.

The visitor’s previous marathon time is compared with bands of earlier recorded bests, not an exact last-race match. A custom target can be any whole minute; success counts and nearby-finish comparisons use that exact threshold. Pacing profiles and improvement breakdowns use the clearly displayed 15-minute achieved-time band centered on the nearest preset, with the upper boundary excluded.

Fallbacks keep the selected course and try broader age and gender groups before dropping an earlier-time restriction. Each answer prints its actual comparison group and identifies broadened filters. Cross-course comparisons use one common set of filters across all displayed courses. A missing comparison is not shown as zero. Age-group and course comparisons deliberately vary the dimension being compared.

Checkpoint comparisons use current elapsed progress instead of previous marathon time. They match a two-minute elapsed-time interval and, when entered, the most recent 5 km pace relative to elapsed average pace. The entered time determines the required remaining pace exactly, while historical outcomes describe the full matching interval. These retrospective proportions have not been calibrated as personal forecasts.

What does a race near my target look like?

Select finishes in the displayed 15-minute finish-time band, centered on the nearest 15-minute target preset. Show each section’s median pace and middle 50% of observed paces. These are achieved times, not declared goals, and the profile is not an optimal pacing plan.

Open this personalized question
Which openings are associated with finishing under my target?

Use runners with a recorded best in the two strictly earlier calendar years. Classify the first 10 km as more than 2% faster than that benchmark’s marathon pace, within 2%, or more than 2% slower. For each group, count finishes strictly below the selected time. Groups are observational and pool available editions; a prior-time filter narrows ability but does not eliminate confounding.

Open this personalized question
Where do nearby finishes gain or lose time?

Compare finishes in the five minutes strictly below the selected target with finishes from the target up to, but excluding, five minutes above it. Subtract the target’s even-pace time budget from each mean section duration. Means reconcile with mean finish-time differences; selection on the outcome makes this descriptive.

Open this personalized question
How does pacing differ across age groups?

Compare median percentage pace change from 0–20 to 20–40 km across exact-age bands. When previous performance is supplied, compare the same 15-minute prior-time band. Otherwise compare the same displayed achieved-time band. Gender and course are held to the displayed selection. This is a comparison of different people, not an individual aging trajectory.

Open this personalized question
What pace changes appear alongside the supplied terrain?

Align the supplied unique city-level course segments with the pacing profile. Historical route validity is unknown. Net elevation change can hide both climbing and descending; bridge decks, tunnels, smoothing and route changes can affect the profile. This does not estimate a historical hill penalty or grade-adjusted effort.

Open this personalized question
What follows a fast opening on a downhill-start profile?

Identify a net-downhill opening from the first two supplied 5 km terrain segments. Among the earlier-benchmark opening groups, compare median pace change from the 5–20 km baseline to the final 12.195 km. The supplied route is a proxy, and these associations do not establish that an early descent caused later slowing. Where no downhill opening is documented, show the general opening comparison and say so.

Open this personalized question
What happened to runners at a similar checkpoint time?

Match course, available age/gender groups, checkpoint and a two-minute elapsed-time band at 20, 30 or 35 km. Optional recent-5-km pace is classified relative to elapsed average pace using ±2%. The previous-marathon filter is not used because this comparison conditions on current-race progress. Show historical finish percentiles and the fraction below the visitor’s threshold, not a calibrated personal probability. Only complete eligible finishers are represented.

Open this personalized question
Which courses combine faster outcomes and consistency?

For the same displayed age, gender and prior-time cohort across courses, show the 10th, 50th and 90th percentiles of finish-time change relative to the recent recorded best. Require at least 100 observations and three editions per course. Use identical cohort criteria across the displayed courses. These different runners and editions do not isolate a course effect or provide equivalent-time predictions.

Open this personalized question
How do outcomes differ across cooler and warmer editions?

Use the supplied modeled start-hour temperature and runners with earlier benchmarks. Compute edition-level median finish-time change, requiring 20 finishes per edition/group. Average edition medians equally within temperature bands, requiring three editions and 100 finishes. Weather exposure is a start-hour proxy, not personal exposure; humidity, sun, wind and route changes are not isolated.

Open this personalized question
Where does my target sit among comparable results?

Evaluate the selected time against the historical finish-time distribution in the displayed cohort. Optional previous performance selects a 15-minute band of recent recorded bests. Each whole-minute threshold from 1:30 through 12:00 is calculated exactly using a strict less-than comparison. Without previous performance this describes the selected field, not individual readiness.

Open this personalized question
How do runners change when they return to this course?

Use consecutive linked appearances in different years, at most three years apart, on the same city course, requiring one eligible recorded race in both endpoint years. Apply the profile to the later appearance and compare both observed section profiles, each normalized to its own finish. Report paired mean finish-time change. Route changes, fitness and selection remain possible explanations; first recorded is not first ever.

Open this personalized question
Where did runners improving toward this time gain minutes?

Select finishes in the displayed achieved-time band that beat the runner’s fastest recorded finish from strictly earlier years. Compare durations in the opening 10 km, middle 20 km and final 12.195 km with that earlier result. Mean block gains sum to mean finish improvement. This is an earlier recorded best, not a lifetime personal best or a prescription for a future improvement.

Open this personalized question

Personalized analysis coverage, provenance and calculation details

How to read the pacing charts

“40 km” means the 35–40 km section. Pace profiles show section averages at the section’s end distance, not instantaneous pace at that timing mat. An upward movement means slower pace. Joining two averages with a line does not establish a sudden change at either checkpoint.

A median profile is a summary across runners, not one runner’s race or an optimal strategy. Outcome percentiles show variation between performances. Confidence intervals instead show uncertainty in an estimate; prediction intervals show a range for an individual future outcome. The charts keep these separate.

Every section must be compared as pace or weighted by its distance. The final 2.195 km is shorter than a 5 km section. Twenty kilometers is before halfway, which is 21.0975 km.

What counts as a good performance?

The analyses require complete, strictly increasing checkpoints, a 90-minute to 12-hour finish, and section paces of 2–20 minutes per km. Missing or invalid readings are excluded, never repaired by guessing. These filters may exclude genuine unusual performances; sample counts and exclusions are recorded with each analysis. The reviewed edition policy also excludes incomplete, held, selected-field and known-invalid editions before results or earlier benchmarks are formed. Missing age or recorded gender alone does not remove usable timings from the overall cohort.

The history analyses use the fastest eligible finish in the two strictly earlier calendar years as a benchmark. A substantially improved performance is more than 2% faster than that recorded best. The benchmark describes prior performance; it is not a measurement of current fitness or a course-adjusted expected finish. Supplied “ability,” personal-best and exceptional-performance labels are not used to define these outcomes.

Performance change is 100 × (current finish ÷ earlier benchmark − 1). Opening change compares the first 10 km pace with that benchmark’s full-marathon average pace. Negative values mean faster. Faster openings are more than 2% faster, similar openings are within 2%, and slower openings are more than 2% slower. These thresholds are descriptive choices, not physiological boundaries.

Finishing-time groups are useful for describing race shapes. Strategy comparisons need ability known before the race, so that the result is not also used to define the comparison group.

How we link runners without mixing up records

The current full export supplies candidate runner identities and canonical record IDs. We verify that raw and feature IDs are unique, non-null and set-equal, then join each feature to its raw record by ID. Race, city, year, recorded name, finish time and every section duration must also agree, with timings rounded to milliseconds. A record ID identifies one result; it does not independently establish that results from different races belong to the same person.

We reject identity groups with conflicting recorded gender, inferred birth years spanning more than two years, or duplicate records in an edition. These checks reduce errors; they do not independently confirm that every identity link is correct. Supplied names and runner IDs are included in the full downloadable data.

Using only strictly earlier years prevents current-race and same-year results from entering the prior benchmark. Personal-best gains compare with the fastest finish in earlier recorded years; they cannot establish a lifetime best. Race-pair analyses use adjacent observations with one race in each endpoint year. The interval analysis instead uses exact supplied dates, restricted to identity groups with complete, unique date coverage.

Exact-age analyses exclude age-group-only records. Reported gender is used as supplied and is never inferred from names. Unknown categories are preserved in overall counts where the analysis permits them.

Find your races

The name search checks recorded names in the current export. A matching name is a possible match, not proof of identity. Confirm the races that belong to you before comparing them; people can share a name, and a runner can appear under different spellings.

Individual race views show the recorded result and split coverage. Analyses require usable timings and the same edition-quality rules as the rest of the study. A missing split is left missing. A fastest result means the best among the selected recorded races, not a verified lifetime personal best.

Finish placement compares eligible records in the same race edition, with options for recorded gender, exact-age groups (18–24, then five-year bands through 85–89), and both together. Each comparison requires at least 101 finishes. Percentiles compare with the other finishes, counting ties halfway; higher means faster. These are observed database ranks, not official race placings.

Pacing comparisons use a 15-minute achieved-finish-time band within the selected group. The median and middle 50% include your result and describe the pacing of similar finish times; they do not measure ability or prescribe a strategy. Late pace covers 30 km–finish relative to 5–20 km. Two selected races can also be compared section by section; the elapsed-time differences add up to the finish difference.

Race-day context shows modeled start-hour weather and the following four hourly readings. Precipitation covers the hour preceding each reading, not your whole race. Supplied terrain sections have unverified historical validity, and their climbing totals can differ from the whole-route profile. Weather, terrain and changes in field composition help interpret comparisons; they do not produce an adjusted finish time or establish why a performance changed.

Search recorded names and explore your races

Matching, uncertainty and forecast validation

The opening-strategy comparison matches race edition, recorded gender and 15-minute bands of prior performance. All three opening groups need at least 20 finishes within a stratum; the smallest group supplies a common weight. Its 95% interval comes from 500 resamples of whole race editions. This accounts for edition clustering, but not a runner appearing across editions.

Other charts are descriptive unless their individual methods state otherwise. A large runner count does not eliminate confounding or create thousands of independent weather observations. No significance ranking or claim of an optimal strategy is made from the many comparisons.

The checkpoint forecast is trained on earlier years and tested on the latest three observed years. It compares even-pace extrapolation with a model calibrated to elapsed pace, and then adds the latest pace trend. All inputs are available at the checkpoint. Training cells need 100 records; sparse cells use a documented fallback. The site reports actual forecast error and observed coverage of an 80% prediction interval.

The forecast test holds out entire later race editions. Runners can appear in both periods, but identities are not model inputs. The same complete-finish cohort is used at every checkpoint; these results do not predict withdrawals or apply automatically to runners with missing splits.

Weather, routes and qualifying rules

The temperature-band analysis uses the Open-Meteo archive hour nearest the scheduled local start. Each temperature-band comparison gives equal weight to eligible race editions, with at least five editions per band. The additional humidity, warming and wind screen uses start-hour and race-window weather summaries, equal edition weights, course and year controls, temperature adjustment, and uncertainty resampled across whole courses. Only findings passing the documented evidence rule are published as analyses. These modeled conditions are proxies, not each runner’s measured exposure throughout the race.

Course elevation comes from supplied GPX routes and digital elevation models, sometimes smoothed over 800 m. Historical validity years are absent. Terrain comparisons use the available city route as an explicitly labeled proxy; mixed hills, route changes, bridges and tunnels limit interpretation.

The qualifying-rule comparison uses the B.A.A.’s published five-minute change for the 2020 Boston Marathon. It compares finish-time bunching around fixed old and new standards in 2016–2017 and 2019 for exact ages 18–31. It does not assign eligibility, acceptance, or declared Boston intentions. Official standards and announcement.

Return, missing follow-up and new exports

The near-miss analysis includes runners who have no later observed race. It requires two subsequent years of observed editions in the original city and excludes the latest two years as index races. Return means an eligible linked appearance anywhere in the export in the next two calendar years. Missing follow-up is never called retirement.

Every calculation records the input export, checksums, method version, eligible count and exclusions. New data releases can rerun the same calculations. Aggregate results are validated and reviewed before the website changes; downloadable calculation files retain the exact source identifiers for reproducibility.

What can these comparisons establish?

They describe associations. Runners selecting different strategies can also differ in fitness, experience, goals, or conditions. Comparisons should account for these differences and report sample sizes and uncertainty.

Predictions must use information available at the checkpoint being studied and be tested on unseen race editions. Declared goals should be distinguished from inferred time landmarks. An absence from the database does not establish that a runner stopped racing.

Splits measure time and pace. Physiological effort, fueling, intentions, and training need additional evidence. Terrain-adjusted pace is an estimated proxy, especially when a 5 km section contains both climbs and descents.

How the questions fit together

The research archive’s 35 questions are organized around race strategy, courses and conditions, goals and finishing, runner differences, and learning over time. The original question lists are incorporated, alongside six additions about successful strategies, personal-best gains, congestion, consistency, and course adaptation.

Methods for every question

Open a question for its actual definitions, comparison groups, sample rules and limits. The result page contains the answer and charts.

1. How do people actually pace a marathon?

For each runner, divide section pace by their full-marathon average pace and subtract one. Plot the median of these individual percentages at each checkpoint; zero is that runner’s own marathon pace. A median curve is not itself one runner’s race and need not integrate to zero.

Compare elapsed time over 0–20 km with 20–40 km. Faster: more than 2% faster; similar: within 2%; moderate slowing: more than 2% through 10%; pronounced slowing: more than 10%. The final 2.195 km appears in the profile but not this equal-distance classification. These are descriptive categories, not slowdown episodes or measured half-marathon splits.

Use complete, strictly increasing elapsed checkpoints at 5, 10, 15, 20, 25, 30, 35, 40 and 42.195 km. Clock strings must parse as H:MM:SS or M:SS. No missing splits are interpolated.

Remove exact duplicate race records, ignoring database IDs, ingestion timestamps and source URLs. Retain finishes from 90 minutes to 12 hours with every section between 2 and 20 minutes per km. These quality filters can exclude genuine unusual performances; the analysis describes this eligible cohort, not every entrant.

Each observation is a race finish. The same person can appear in multiple races. No runner identities are linked across races. Public cells require at least 100 observations; matched comparisons additionally require at least 20 per group in each stratum.

Apply the reviewed source-quality edition exclusions for this exact export after the timing checks. Known invalid split grids, incomplete ingestion, unreconciled HOLD editions and a selected top-finisher field do not contribute to the analyses or prior benchmarks. Report source exclusions separately; an already invalid timing row is not counted twice. Missing age or recorded gender alone does not exclude an otherwise eligible finish from the overall cohort. Other sparse editions are not declared incomplete merely from their size.

2. What does an unusually good race look like?

Define a substantially improved performance as a finish more than 2% faster than the recent recorded best. Compare each section’s pace with the full-marathon pace of that earlier benchmark, then plot the median by outcome group.

For explicitly audited canonical-ID releases, validate unique matching CORE/FULL ID sets and matching recorded edition/name labels, then join by record ID with matching finish and all nine section durations. Legacy exports retain the one-to-one edition, trimmed lowercase name and full-timing join because their IDs are incompatible. Durations are compared to milliseconds. Keep only supplied non-ambiguous runner identities without conflicting recorded gender, inferred birth years spanning more than two years, or duplicate editions. These are candidate cross-race identities, not independently verified people; unlinked runners are absent. The linkage audit records which join was used.

Recompute the benchmark as the fastest eligible finish in the two strictly earlier calendar years. The current race and every other race in its calendar year are excluded. This avoids guessing within-year chronology and prevents current-outcome leakage. It is a recent recorded best, not a fitness measurement, an expected finish, or a lifetime personal best.

Performance change is 100 × (current finish / recent recorded best − 1). Negative is faster. Opening change compares 0–10 km pace with that earlier best’s full-marathon pace. Faster opening: more than 2% faster; similar: within 2%; slower: more than 2% slower. The ±2% and ±5% cutoffs are predefined descriptions, not physiological thresholds.

Use complete, strictly increasing elapsed checkpoints at 5, 10, 15, 20, 25, 30, 35, 40 and 42.195 km. Clock strings must parse as H:MM:SS or M:SS. No missing splits are interpolated.

Remove exact duplicate race records, ignoring database IDs, ingestion timestamps and source URLs. Retain finishes from 90 minutes to 12 hours with every section between 2 and 20 minutes per km. These quality filters can exclude genuine unusual performances; the analysis describes this eligible cohort, not every entrant.

Full source records are available in public GitHub Releases. These chart tables require at least 100 eligible observations per cell for estimate reliability. Counts refer to finishes, linked pairs or event observations as specified in that answer.

These are observational results. Fitness changes, intentions, training, selection into the dataset and unmeasured conditions can explain differences. Outcome percentiles describe variation among performances, not confidence intervals or advice about the best strategy.

Apply the reviewed source-quality edition exclusions for this exact export after the timing checks. Known invalid split grids, incomplete ingestion, unreconciled HOLD editions and a selected top-finisher field do not contribute to the analyses or prior benchmarks. Report source exclusions separately; an already invalid timing row is not counted twice. Missing age or recorded gender alone does not exclude an otherwise eligible finish from the overall cohort. Other sparse editions are not declared incomplete merely from their size.

3. What are the rewards and risks of an aggressive start?

Match faster, similar and slower openings within city, year, race, recorded gender and 15-minute bands of recent recorded best. Keep strata with at least 20 in all three groups. Weight every group by the smallest group count in that stratum.

Show the weighted mean percentage change from the prior benchmark. The lower and upper limits are percentile confidence limits from 500 bootstrap draws of whole race editions, seed 20260908. This captures edition clustering but not dependence when the same runner appears in different editions. No multiple-comparison significance claims are made.

For explicitly audited canonical-ID releases, validate unique matching CORE/FULL ID sets and matching recorded edition/name labels, then join by record ID with matching finish and all nine section durations. Legacy exports retain the one-to-one edition, trimmed lowercase name and full-timing join because their IDs are incompatible. Durations are compared to milliseconds. Keep only supplied non-ambiguous runner identities without conflicting recorded gender, inferred birth years spanning more than two years, or duplicate editions. These are candidate cross-race identities, not independently verified people; unlinked runners are absent. The linkage audit records which join was used.

Recompute the benchmark as the fastest eligible finish in the two strictly earlier calendar years. The current race and every other race in its calendar year are excluded. This avoids guessing within-year chronology and prevents current-outcome leakage. It is a recent recorded best, not a fitness measurement, an expected finish, or a lifetime personal best.

Performance change is 100 × (current finish / recent recorded best − 1). Negative is faster. Opening change compares 0–10 km pace with that earlier best’s full-marathon pace. Faster opening: more than 2% faster; similar: within 2%; slower: more than 2% slower. The ±2% and ±5% cutoffs are predefined descriptions, not physiological thresholds.

Use complete, strictly increasing elapsed checkpoints at 5, 10, 15, 20, 25, 30, 35, 40 and 42.195 km. Clock strings must parse as H:MM:SS or M:SS. No missing splits are interpolated.

Remove exact duplicate race records, ignoring database IDs, ingestion timestamps and source URLs. Retain finishes from 90 minutes to 12 hours with every section between 2 and 20 minutes per km. These quality filters can exclude genuine unusual performances; the analysis describes this eligible cohort, not every entrant.

Full source records are available in public GitHub Releases. These chart tables require at least 100 eligible observations per cell for estimate reliability. Counts refer to finishes, linked pairs or event observations as specified in that answer.

These are observational results. Fitness changes, intentions, training, selection into the dataset and unmeasured conditions can explain differences. Outcome percentiles describe variation among performances, not confidence intervals or advice about the best strategy.

Apply the reviewed source-quality edition exclusions for this exact export after the timing checks. Known invalid split grids, incomplete ingestion, unreconciled HOLD editions and a selected top-finisher field do not contribute to the analyses or prior benchmarks. Report source exclusions separately; an already invalid timing row is not counted twice. Missing age or recorded gender alone does not exclude an otherwise eligible finish from the overall cohort. Other sparse editions are not declared incomplete merely from their size.

4. How do runners successfully respond to a slow start?

A slow start is 0–5 km pace more than 5% slower than the earlier benchmark’s average. Group the change from 0–5 to 5–10 km as more than 5% acceleration, 2–5% acceleration, or less acceleration/slowing. Plot the 10th, 50th and 90th percentiles of final performance change. These groups are not matched on course or fitness.

For explicitly audited canonical-ID releases, validate unique matching CORE/FULL ID sets and matching recorded edition/name labels, then join by record ID with matching finish and all nine section durations. Legacy exports retain the one-to-one edition, trimmed lowercase name and full-timing join because their IDs are incompatible. Durations are compared to milliseconds. Keep only supplied non-ambiguous runner identities without conflicting recorded gender, inferred birth years spanning more than two years, or duplicate editions. These are candidate cross-race identities, not independently verified people; unlinked runners are absent. The linkage audit records which join was used.

Recompute the benchmark as the fastest eligible finish in the two strictly earlier calendar years. The current race and every other race in its calendar year are excluded. This avoids guessing within-year chronology and prevents current-outcome leakage. It is a recent recorded best, not a fitness measurement, an expected finish, or a lifetime personal best.

Performance change is 100 × (current finish / recent recorded best − 1). Negative is faster. Opening change compares 0–10 km pace with that earlier best’s full-marathon pace. Faster opening: more than 2% faster; similar: within 2%; slower: more than 2% slower. The ±2% and ±5% cutoffs are predefined descriptions, not physiological thresholds.

Use complete, strictly increasing elapsed checkpoints at 5, 10, 15, 20, 25, 30, 35, 40 and 42.195 km. Clock strings must parse as H:MM:SS or M:SS. No missing splits are interpolated.

Remove exact duplicate race records, ignoring database IDs, ingestion timestamps and source URLs. Retain finishes from 90 minutes to 12 hours with every section between 2 and 20 minutes per km. These quality filters can exclude genuine unusual performances; the analysis describes this eligible cohort, not every entrant.

Full source records are available in public GitHub Releases. These chart tables require at least 100 eligible observations per cell for estimate reliability. Counts refer to finishes, linked pairs or event observations as specified in that answer.

These are observational results. Fitness changes, intentions, training, selection into the dataset and unmeasured conditions can explain differences. Outcome percentiles describe variation among performances, not confidence intervals or advice about the best strategy.

Apply the reviewed source-quality edition exclusions for this exact export after the timing checks. Known invalid split grids, incomplete ingestion, unreconciled HOLD editions and a selected top-finisher field do not contribute to the analyses or prior benchmarks. Report source exclusions separately; an already invalid timing row is not counted twice. Missing age or recorded gender alone does not exclude an otherwise eligible finish from the overall cohort. Other sparse editions are not declared incomplete merely from their size.

5. Can two runners reach 20 km together but have different prospects?

Compare the 15–20 km pace with 10–15 km. Accelerating is more than 2% faster, slowing more than 2% slower, and steady within 2%. Match within city, year, race and the same floored minute of elapsed time at 20 km.

Only strata with at least 20 finishes in each of the three groups qualify. Use the smallest group count as each stratum’s common weight for all three groups, then average the within-stratum finish-time differences versus steady. Sample counts are actual observations, not matching weights.

Runners arrive within the same 60-second band, not at an identical instant. Prior fitness and runner intention are unknown. This is a conditional association, not proof that accelerating causes a better finish or an out-of-sample prediction.

Use complete, strictly increasing elapsed checkpoints at 5, 10, 15, 20, 25, 30, 35, 40 and 42.195 km. Clock strings must parse as H:MM:SS or M:SS. No missing splits are interpolated.

Remove exact duplicate race records, ignoring database IDs, ingestion timestamps and source URLs. Retain finishes from 90 minutes to 12 hours with every section between 2 and 20 minutes per km. These quality filters can exclude genuine unusual performances; the analysis describes this eligible cohort, not every entrant.

Each observation is a race finish. The same person can appear in multiple races. No runner identities are linked across races. Public cells require at least 100 observations; matched comparisons additionally require at least 20 per group in each stratum.

Apply the reviewed source-quality edition exclusions for this exact export after the timing checks. Known invalid split grids, incomplete ingestion, unreconciled HOLD editions and a selected top-finisher field do not contribute to the analyses or prior benchmarks. Report source exclusions separately; an already invalid timing row is not counted twice. Missing age or recorded gender alone does not exclude an otherwise eligible finish from the overall cohort. Other sparse editions are not declared incomplete merely from their size.

6. How early can the splits reveal how the race will finish?

At each checkpoint, start with elapsed time × 42.195 / distance. The elapsed-only model multiplies this by the training median actual/projected ratio in 30-second-per-km elapsed-pace bands. The trend model adds the most recent section’s change from the previous section (faster than −2%, within ±2%, or slower than +2%). At 5 km no trend exists.

Hold out the latest three observed calendar years (2024–2026). Fit every factor and the 10th–90th percentile ratio interval on earlier years only. Training cells need at least 100 records; missing trend cells fall back to the pace-band model, then the pooled training model. No future splits, finishing-time groups, identities or supplied ability fields enter a prediction.

Report median absolute error for all three methods, plus observed coverage and median width of the nominal 80% prediction interval and the 90th percentile absolute error. Model selection is fixed before examining these results. A runner may occur in training and test in different years; identities are not used. Results apply to complete eligible finishers and do not predict withdrawals.

Use complete, strictly increasing elapsed checkpoints at 5, 10, 15, 20, 25, 30, 35, 40 and 42.195 km. Clock strings must parse as H:MM:SS or M:SS. No missing splits are interpolated.

Remove exact duplicate race records, ignoring database IDs, ingestion timestamps and source URLs. Retain finishes from 90 minutes to 12 hours with every section between 2 and 20 minutes per km. These quality filters can exclude genuine unusual performances; the analysis describes this eligible cohort, not every entrant.

Full source records are available in public GitHub Releases. These chart tables require at least 100 eligible observations per cell for estimate reliability. Counts refer to finishes, linked pairs or event observations as specified in that answer.

These are observational results. Fitness changes, intentions, training, selection into the dataset and unmeasured conditions can explain differences. Outcome percentiles describe variation among performances, not confidence intervals or advice about the best strategy.

Apply the reviewed source-quality edition exclusions for this exact export after the timing checks. Known invalid split grids, incomplete ingestion, unreconciled HOLD editions and a selected top-finisher field do not contribute to the analyses or prior benchmarks. Report source exclusions separately; an already invalid timing row is not counted twice. Missing age or recorded gender alone does not exclude an otherwise eligible finish from the overall cohort. Other sparse editions are not declared incomplete merely from their size.

7. When can runners regain their rhythm after a bad patch?

Scan 20–25, 25–30 and 30–35 km in order. A patch is the first section more than 10% slower than both the immediately preceding section and the 5–20 km baseline pace. Recovery means the next full 5 km is no more than 5% slower than that baseline.

Every runner contributes at most one patch. The outcome is next-section recovery, not a diagnosis or necessarily a return maintained to the finish. Course sections and different runner mixes can account for differences across distances. Five-kilometer timing cannot distinguish stops, walking, fatigue or terrain.

Use complete, strictly increasing elapsed checkpoints at 5, 10, 15, 20, 25, 30, 35, 40 and 42.195 km. Clock strings must parse as H:MM:SS or M:SS. No missing splits are interpolated.

Remove exact duplicate race records, ignoring database IDs, ingestion timestamps and source URLs. Retain finishes from 90 minutes to 12 hours with every section between 2 and 20 minutes per km. These quality filters can exclude genuine unusual performances; the analysis describes this eligible cohort, not every entrant.

Full source records are available in public GitHub Releases. These chart tables require at least 100 eligible observations per cell for estimate reliability. Counts refer to finishes, linked pairs or event observations as specified in that answer.

These are observational results. Fitness changes, intentions, training, selection into the dataset and unmeasured conditions can explain differences. Outcome percentiles describe variation among performances, not confidence intervals or advice about the best strategy.

Apply the reviewed source-quality edition exclusions for this exact export after the timing checks. Known invalid split grids, incomplete ingestion, unreconciled HOLD editions and a selected top-finisher field do not contribute to the analyses or prior benchmarks. Report source exclusions separately; an already invalid timing row is not counted twice. Missing age or recorded gender alone does not exclude an otherwise eligible finish from the overall cohort. Other sparse editions are not declared incomplete merely from their size.

8. Is a negative split always associated with a better performance?

Within each complete-race pattern, divide finishes more than 2% faster than the earlier benchmark by all linked finishes with that pattern. This conditions on a pattern known after the finish; it is a retrospective association, not a pre-race strategy trial.

For explicitly audited canonical-ID releases, validate unique matching CORE/FULL ID sets and matching recorded edition/name labels, then join by record ID with matching finish and all nine section durations. Legacy exports retain the one-to-one edition, trimmed lowercase name and full-timing join because their IDs are incompatible. Durations are compared to milliseconds. Keep only supplied non-ambiguous runner identities without conflicting recorded gender, inferred birth years spanning more than two years, or duplicate editions. These are candidate cross-race identities, not independently verified people; unlinked runners are absent. The linkage audit records which join was used.

Recompute the benchmark as the fastest eligible finish in the two strictly earlier calendar years. The current race and every other race in its calendar year are excluded. This avoids guessing within-year chronology and prevents current-outcome leakage. It is a recent recorded best, not a fitness measurement, an expected finish, or a lifetime personal best.

Performance change is 100 × (current finish / recent recorded best − 1). Negative is faster. Opening change compares 0–10 km pace with that earlier best’s full-marathon pace. Faster opening: more than 2% faster; similar: within 2%; slower: more than 2% slower. The ±2% and ±5% cutoffs are predefined descriptions, not physiological thresholds.

Use complete, strictly increasing elapsed checkpoints at 5, 10, 15, 20, 25, 30, 35, 40 and 42.195 km. Clock strings must parse as H:MM:SS or M:SS. No missing splits are interpolated.

Remove exact duplicate race records, ignoring database IDs, ingestion timestamps and source URLs. Retain finishes from 90 minutes to 12 hours with every section between 2 and 20 minutes per km. These quality filters can exclude genuine unusual performances; the analysis describes this eligible cohort, not every entrant.

Full source records are available in public GitHub Releases. These chart tables require at least 100 eligible observations per cell for estimate reliability. Counts refer to finishes, linked pairs or event observations as specified in that answer.

These are observational results. Fitness changes, intentions, training, selection into the dataset and unmeasured conditions can explain differences. Outcome percentiles describe variation among performances, not confidence intervals or advice about the best strategy.

Apply the reviewed source-quality edition exclusions for this exact export after the timing checks. Known invalid split grids, incomplete ingestion, unreconciled HOLD editions and a selected top-finisher field do not contribute to the analyses or prior benchmarks. Report source exclusions separately; an already invalid timing row is not counted twice. Missing age or recorded gender alone does not exclude an otherwise eligible finish from the overall cohort. Other sparse editions are not declared incomplete merely from their size.

9. Is there one good pacing strategy, or several?

Restrict to finishes more than 2% faster than the recent recorded best. Divide the number in each equal-distance pattern by all such improved finishes. Suppress any pattern with fewer than 100 improved finishes; if a pattern is suppressed, visible shares need not sum to 100.

For explicitly audited canonical-ID releases, validate unique matching CORE/FULL ID sets and matching recorded edition/name labels, then join by record ID with matching finish and all nine section durations. Legacy exports retain the one-to-one edition, trimmed lowercase name and full-timing join because their IDs are incompatible. Durations are compared to milliseconds. Keep only supplied non-ambiguous runner identities without conflicting recorded gender, inferred birth years spanning more than two years, or duplicate editions. These are candidate cross-race identities, not independently verified people; unlinked runners are absent. The linkage audit records which join was used.

Recompute the benchmark as the fastest eligible finish in the two strictly earlier calendar years. The current race and every other race in its calendar year are excluded. This avoids guessing within-year chronology and prevents current-outcome leakage. It is a recent recorded best, not a fitness measurement, an expected finish, or a lifetime personal best.

Performance change is 100 × (current finish / recent recorded best − 1). Negative is faster. Opening change compares 0–10 km pace with that earlier best’s full-marathon pace. Faster opening: more than 2% faster; similar: within 2%; slower: more than 2% slower. The ±2% and ±5% cutoffs are predefined descriptions, not physiological thresholds.

Use complete, strictly increasing elapsed checkpoints at 5, 10, 15, 20, 25, 30, 35, 40 and 42.195 km. Clock strings must parse as H:MM:SS or M:SS. No missing splits are interpolated.

Remove exact duplicate race records, ignoring database IDs, ingestion timestamps and source URLs. Retain finishes from 90 minutes to 12 hours with every section between 2 and 20 minutes per km. These quality filters can exclude genuine unusual performances; the analysis describes this eligible cohort, not every entrant.

Full source records are available in public GitHub Releases. These chart tables require at least 100 eligible observations per cell for estimate reliability. Counts refer to finishes, linked pairs or event observations as specified in that answer.

These are observational results. Fitness changes, intentions, training, selection into the dataset and unmeasured conditions can explain differences. Outcome percentiles describe variation among performances, not confidence intervals or advice about the best strategy.

Apply the reviewed source-quality edition exclusions for this exact export after the timing checks. Known invalid split grids, incomplete ingestion, unreconciled HOLD editions and a selected top-finisher field do not contribute to the analyses or prior benchmarks. Report source exclusions separately; an already invalid timing row is not counted twice. Missing age or recorded gender alone does not exclude an otherwise eligible finish from the overall cohort. Other sparse editions are not declared incomplete merely from their size.

10. Which pacing approaches offer consistency, and which are more variable?

Within each prior-time band and opening group, calculate the 10th, 50th and 90th percentiles of finish-time percentage change. Also calculate the share more than 5% slower than the earlier benchmark. These distributions are unadjusted for course and edition, and describe observed finishers only.

For explicitly audited canonical-ID releases, validate unique matching CORE/FULL ID sets and matching recorded edition/name labels, then join by record ID with matching finish and all nine section durations. Legacy exports retain the one-to-one edition, trimmed lowercase name and full-timing join because their IDs are incompatible. Durations are compared to milliseconds. Keep only supplied non-ambiguous runner identities without conflicting recorded gender, inferred birth years spanning more than two years, or duplicate editions. These are candidate cross-race identities, not independently verified people; unlinked runners are absent. The linkage audit records which join was used.

Recompute the benchmark as the fastest eligible finish in the two strictly earlier calendar years. The current race and every other race in its calendar year are excluded. This avoids guessing within-year chronology and prevents current-outcome leakage. It is a recent recorded best, not a fitness measurement, an expected finish, or a lifetime personal best.

Performance change is 100 × (current finish / recent recorded best − 1). Negative is faster. Opening change compares 0–10 km pace with that earlier best’s full-marathon pace. Faster opening: more than 2% faster; similar: within 2%; slower: more than 2% slower. The ±2% and ±5% cutoffs are predefined descriptions, not physiological thresholds.

Use complete, strictly increasing elapsed checkpoints at 5, 10, 15, 20, 25, 30, 35, 40 and 42.195 km. Clock strings must parse as H:MM:SS or M:SS. No missing splits are interpolated.

Remove exact duplicate race records, ignoring database IDs, ingestion timestamps and source URLs. Retain finishes from 90 minutes to 12 hours with every section between 2 and 20 minutes per km. These quality filters can exclude genuine unusual performances; the analysis describes this eligible cohort, not every entrant.

Full source records are available in public GitHub Releases. These chart tables require at least 100 eligible observations per cell for estimate reliability. Counts refer to finishes, linked pairs or event observations as specified in that answer.

These are observational results. Fitness changes, intentions, training, selection into the dataset and unmeasured conditions can explain differences. Outcome percentiles describe variation among performances, not confidence intervals or advice about the best strategy.

Apply the reviewed source-quality edition exclusions for this exact export after the timing checks. Known invalid split grids, incomplete ingestion, unreconciled HOLD editions and a selected top-finisher field do not contribute to the analyses or prior benchmarks. Report source exclusions separately; an already invalid timing row is not counted twice. Missing age or recorded gender alone does not exclude an otherwise eligible finish from the overall cohort. Other sparse editions are not declared incomplete merely from their size.

11. What is each course’s pacing fingerprint?

Compute individual section pace relative to each runner’s full-marathon average, then take the median by city and section. Every checkpoint within a city uses the same complete-record cohort.

Cities pool available race editions. Terrain, weather, field composition and route changes are not separated. Course-profile validity years are absent from CORE, so historical elevation has not been assigned to these runners.

Use complete, strictly increasing elapsed checkpoints at 5, 10, 15, 20, 25, 30, 35, 40 and 42.195 km. Clock strings must parse as H:MM:SS or M:SS. No missing splits are interpolated.

Remove exact duplicate race records, ignoring database IDs, ingestion timestamps and source URLs. Retain finishes from 90 minutes to 12 hours with every section between 2 and 20 minutes per km. These quality filters can exclude genuine unusual performances; the analysis describes this eligible cohort, not every entrant.

Each observation is a race finish. The same person can appear in multiple races. No runner identities are linked across races. Public cells require at least 100 observations; matched comparisons additionally require at least 20 per group in each stratum.

Apply the reviewed source-quality edition exclusions for this exact export after the timing checks. Known invalid split grids, incomplete ingestion, unreconciled HOLD editions and a selected top-finisher field do not contribute to the analyses or prior benchmarks. Report source exclusions separately; an already invalid timing row is not counted twice. Missing age or recorded gender alone does not exclude an otherwise eligible finish from the overall cohort. Other sparse editions are not declared incomplete merely from their size.

12. How do runners adjust their pace to climbs and descents?

Join a unique supplied course segment by city and exact checkpoint distance. Classify net grade above +0.15% as uphill, below −0.15% as downhill and the remainder near level. At each distance, plot median individual section pace relative to that runner’s full-marathon average.

The course profiles use GPX geometry and digital elevation models, sometimes smoothed over 800 m. Net grade conceals mixed climbs and descents; bridge decks, tunnels and route changes may be wrong. Historical races are compared with the available city profile as a proxy only. Different terrain groups contain different courses and fields; no causal hill penalty or physiological effort is inferred.

Use complete, strictly increasing elapsed checkpoints at 5, 10, 15, 20, 25, 30, 35, 40 and 42.195 km. Clock strings must parse as H:MM:SS or M:SS. No missing splits are interpolated.

Remove exact duplicate race records, ignoring database IDs, ingestion timestamps and source URLs. Retain finishes from 90 minutes to 12 hours with every section between 2 and 20 minutes per km. These quality filters can exclude genuine unusual performances; the analysis describes this eligible cohort, not every entrant.

Full source records are available in public GitHub Releases. These chart tables require at least 100 eligible observations per cell for estimate reliability. Counts refer to finishes, linked pairs or event observations as specified in that answer.

These are observational results. Fitness changes, intentions, training, selection into the dataset and unmeasured conditions can explain differences. Outcome percentiles describe variation among performances, not confidence intervals or advice about the best strategy.

Apply the reviewed source-quality edition exclusions for this exact export after the timing checks. Known invalid split grids, incomplete ingestion, unreconciled HOLD editions and a selected top-finisher field do not contribute to the analyses or prior benchmarks. Report source exclusions separately; an already invalid timing row is not counted twice. Missing age or recorded gender alone does not exclude an otherwise eligible finish from the overall cohort. Other sparse editions are not declared incomplete merely from their size.

13. What would your time be on another course?

Use consecutive linked races in different calendar years, at most three years apart, with one recorded eligible race in each endpoint year. For each course pair, calculate the mean destination-minus-origin finish time separately for runners taking each race order, then average those two means equally.

Require at least 20 pairs in each order and at least 100 overall. Balancing order reduces simple order imbalance but cannot remove fitness, weather, aging, motivation or entry-selection effects. The same runner can supply more than one pair. Positive means a slower finish at the destination.

For explicitly audited canonical-ID releases, validate unique matching CORE/FULL ID sets and matching recorded edition/name labels, then join by record ID with matching finish and all nine section durations. Legacy exports retain the one-to-one edition, trimmed lowercase name and full-timing join because their IDs are incompatible. Durations are compared to milliseconds. Keep only supplied non-ambiguous runner identities without conflicting recorded gender, inferred birth years spanning more than two years, or duplicate editions. These are candidate cross-race identities, not independently verified people; unlinked runners are absent. The linkage audit records which join was used.

Use complete, strictly increasing elapsed checkpoints at 5, 10, 15, 20, 25, 30, 35, 40 and 42.195 km. Clock strings must parse as H:MM:SS or M:SS. No missing splits are interpolated.

Remove exact duplicate race records, ignoring database IDs, ingestion timestamps and source URLs. Retain finishes from 90 minutes to 12 hours with every section between 2 and 20 minutes per km. These quality filters can exclude genuine unusual performances; the analysis describes this eligible cohort, not every entrant.

Full source records are available in public GitHub Releases. These chart tables require at least 100 eligible observations per cell for estimate reliability. Counts refer to finishes, linked pairs or event observations as specified in that answer.

These are observational results. Fitness changes, intentions, training, selection into the dataset and unmeasured conditions can explain differences. Outcome percentiles describe variation among performances, not confidence intervals or advice about the best strategy.

Apply the reviewed source-quality edition exclusions for this exact export after the timing checks. Known invalid split grids, incomplete ingestion, unreconciled HOLD editions and a selected top-finisher field do not contribute to the analyses or prior benchmarks. Report source exclusions separately; an already invalid timing row is not counted twice. Missing age or recorded gender alone does not exclude an otherwise eligible finish from the overall cohort. Other sparse editions are not declared incomplete merely from their size.

14. Which marathon offers speed, and which offers consistency?

Within city and prior-time band, report the 10th, 50th and 90th percentiles of finish-time change versus the earlier benchmark. Pool available editions and require at least 100 observations per city and band. A wide percentile range describes individual variation, not uncertainty in the median.

For explicitly audited canonical-ID releases, validate unique matching CORE/FULL ID sets and matching recorded edition/name labels, then join by record ID with matching finish and all nine section durations. Legacy exports retain the one-to-one edition, trimmed lowercase name and full-timing join because their IDs are incompatible. Durations are compared to milliseconds. Keep only supplied non-ambiguous runner identities without conflicting recorded gender, inferred birth years spanning more than two years, or duplicate editions. These are candidate cross-race identities, not independently verified people; unlinked runners are absent. The linkage audit records which join was used.

Recompute the benchmark as the fastest eligible finish in the two strictly earlier calendar years. The current race and every other race in its calendar year are excluded. This avoids guessing within-year chronology and prevents current-outcome leakage. It is a recent recorded best, not a fitness measurement, an expected finish, or a lifetime personal best.

Performance change is 100 × (current finish / recent recorded best − 1). Negative is faster. Opening change compares 0–10 km pace with that earlier best’s full-marathon pace. Faster opening: more than 2% faster; similar: within 2%; slower: more than 2% slower. The ±2% and ±5% cutoffs are predefined descriptions, not physiological thresholds.

Use complete, strictly increasing elapsed checkpoints at 5, 10, 15, 20, 25, 30, 35, 40 and 42.195 km. Clock strings must parse as H:MM:SS or M:SS. No missing splits are interpolated.

Remove exact duplicate race records, ignoring database IDs, ingestion timestamps and source URLs. Retain finishes from 90 minutes to 12 hours with every section between 2 and 20 minutes per km. These quality filters can exclude genuine unusual performances; the analysis describes this eligible cohort, not every entrant.

Full source records are available in public GitHub Releases. These chart tables require at least 100 eligible observations per cell for estimate reliability. Counts refer to finishes, linked pairs or event observations as specified in that answer.

These are observational results. Fitness changes, intentions, training, selection into the dataset and unmeasured conditions can explain differences. Outcome percentiles describe variation among performances, not confidence intervals or advice about the best strategy.

Apply the reviewed source-quality edition exclusions for this exact export after the timing checks. Known invalid split grids, incomplete ingestion, unreconciled HOLD editions and a selected top-finisher field do not contribute to the analyses or prior benchmarks. Report source exclusions separately; an already invalid timing row is not counted twice. Missing age or recorded gender alone does not exclude an otherwise eligible finish from the overall cohort. Other sparse editions are not declared incomplete merely from their size.

15. Does knowing the course improve execution?

Familiar means at least one eligible linked appearance in the same city in an earlier calendar year. Match familiar and first-recorded groups within edition, recorded gender and 15-minute prior-time bands; each group needs 20 finishes per stratum. Use the smaller stratum count as a common weight.

The outcome is mean percentage change from 0–20 to 20–40 km. Both groups have some prior recorded marathon history, but familiarity itself is not randomly assigned. Route changes and visits absent from the dataset are unknown.

For explicitly audited canonical-ID releases, validate unique matching CORE/FULL ID sets and matching recorded edition/name labels, then join by record ID with matching finish and all nine section durations. Legacy exports retain the one-to-one edition, trimmed lowercase name and full-timing join because their IDs are incompatible. Durations are compared to milliseconds. Keep only supplied non-ambiguous runner identities without conflicting recorded gender, inferred birth years spanning more than two years, or duplicate editions. These are candidate cross-race identities, not independently verified people; unlinked runners are absent. The linkage audit records which join was used.

Recompute the benchmark as the fastest eligible finish in the two strictly earlier calendar years. The current race and every other race in its calendar year are excluded. This avoids guessing within-year chronology and prevents current-outcome leakage. It is a recent recorded best, not a fitness measurement, an expected finish, or a lifetime personal best.

Performance change is 100 × (current finish / recent recorded best − 1). Negative is faster. Opening change compares 0–10 km pace with that earlier best’s full-marathon pace. Faster opening: more than 2% faster; similar: within 2%; slower: more than 2% slower. The ±2% and ±5% cutoffs are predefined descriptions, not physiological thresholds.

Use complete, strictly increasing elapsed checkpoints at 5, 10, 15, 20, 25, 30, 35, 40 and 42.195 km. Clock strings must parse as H:MM:SS or M:SS. No missing splits are interpolated.

Remove exact duplicate race records, ignoring database IDs, ingestion timestamps and source URLs. Retain finishes from 90 minutes to 12 hours with every section between 2 and 20 minutes per km. These quality filters can exclude genuine unusual performances; the analysis describes this eligible cohort, not every entrant.

Full source records are available in public GitHub Releases. These chart tables require at least 100 eligible observations per cell for estimate reliability. Counts refer to finishes, linked pairs or event observations as specified in that answer.

These are observational results. Fitness changes, intentions, training, selection into the dataset and unmeasured conditions can explain differences. Outcome percentiles describe variation among performances, not confidence intervals or advice about the best strategy.

Apply the reviewed source-quality edition exclusions for this exact export after the timing checks. Known invalid split grids, incomplete ingestion, unreconciled HOLD editions and a selected top-finisher field do not contribute to the analyses or prior benchmarks. Report source exclusions separately; an already invalid timing row is not counted twice. Missing age or recorded gender alone does not exclude an otherwise eligible finish from the overall cohort. Other sparse editions are not declared incomplete merely from their size.

16. How does weather change the way a marathon is run?

Use the supplied Open-Meteo archive hour nearest the scheduled local start. Join the single city-year weather row whose race date parses and matches the record year. Group temperature below 10°C, 10–14.9°C, 15–19.9°C and at least 20°C.

First calculate each edition’s median outcome among linked runners with a recent benchmark, requiring 100 finishes. Then average edition medians equally within temperature bands, requiring five editions. The profile is normalized by each runner’s own marathon average; performance is relative to the earlier benchmark.

The modeled weather is a start-hour proxy, not each runner’s exposure. Start offsets are absent, temperatures change during the race and humidity, wind, sunshine, terrain and fitness remain potential confounders. Edition counts in the source table are the number of weather exposures; finish counts are not independent weather observations.

For explicitly audited canonical-ID releases, validate unique matching CORE/FULL ID sets and matching recorded edition/name labels, then join by record ID with matching finish and all nine section durations. Legacy exports retain the one-to-one edition, trimmed lowercase name and full-timing join because their IDs are incompatible. Durations are compared to milliseconds. Keep only supplied non-ambiguous runner identities without conflicting recorded gender, inferred birth years spanning more than two years, or duplicate editions. These are candidate cross-race identities, not independently verified people; unlinked runners are absent. The linkage audit records which join was used.

Recompute the benchmark as the fastest eligible finish in the two strictly earlier calendar years. The current race and every other race in its calendar year are excluded. This avoids guessing within-year chronology and prevents current-outcome leakage. It is a recent recorded best, not a fitness measurement, an expected finish, or a lifetime personal best.

Performance change is 100 × (current finish / recent recorded best − 1). Negative is faster. Opening change compares 0–10 km pace with that earlier best’s full-marathon pace. Faster opening: more than 2% faster; similar: within 2%; slower: more than 2% slower. The ±2% and ±5% cutoffs are predefined descriptions, not physiological thresholds.

Use complete, strictly increasing elapsed checkpoints at 5, 10, 15, 20, 25, 30, 35, 40 and 42.195 km. Clock strings must parse as H:MM:SS or M:SS. No missing splits are interpolated.

Remove exact duplicate race records, ignoring database IDs, ingestion timestamps and source URLs. Retain finishes from 90 minutes to 12 hours with every section between 2 and 20 minutes per km. These quality filters can exclude genuine unusual performances; the analysis describes this eligible cohort, not every entrant.

Full source records are available in public GitHub Releases. These chart tables require at least 100 eligible observations per cell for estimate reliability. Counts refer to finishes, linked pairs or event observations as specified in that answer.

These are observational results. Fitness changes, intentions, training, selection into the dataset and unmeasured conditions can explain differences. Outcome percentiles describe variation among performances, not confidence intervals or advice about the best strategy.

Apply the reviewed source-quality edition exclusions for this exact export after the timing checks. Known invalid split grids, incomplete ingestion, unreconciled HOLD editions and a selected top-finisher field do not contribute to the analyses or prior benchmarks. Report source exclusions separately; an already invalid timing row is not counted twice. Missing age or recorded gender alone does not exclude an otherwise eligible finish from the overall cohort. Other sparse editions are not declared incomplete merely from their size.

17. Was it my pacing or a difficult race day?

For every city and race year, calculate the median and 10th–90th percentiles of percentage finish change versus each linked runner’s recent recorded best. Require 100 such finishes per edition. Compare an individual’s percentage change with that edition median to describe their position relative to the field.

This contemporaneous edition reference is retrospective, includes the runner when eligible, and may shift with selection and fitness changes. It is not a weather correction or a prediction available before race day.

For explicitly audited canonical-ID releases, validate unique matching CORE/FULL ID sets and matching recorded edition/name labels, then join by record ID with matching finish and all nine section durations. Legacy exports retain the one-to-one edition, trimmed lowercase name and full-timing join because their IDs are incompatible. Durations are compared to milliseconds. Keep only supplied non-ambiguous runner identities without conflicting recorded gender, inferred birth years spanning more than two years, or duplicate editions. These are candidate cross-race identities, not independently verified people; unlinked runners are absent. The linkage audit records which join was used.

Recompute the benchmark as the fastest eligible finish in the two strictly earlier calendar years. The current race and every other race in its calendar year are excluded. This avoids guessing within-year chronology and prevents current-outcome leakage. It is a recent recorded best, not a fitness measurement, an expected finish, or a lifetime personal best.

Performance change is 100 × (current finish / recent recorded best − 1). Negative is faster. Opening change compares 0–10 km pace with that earlier best’s full-marathon pace. Faster opening: more than 2% faster; similar: within 2%; slower: more than 2% slower. The ±2% and ±5% cutoffs are predefined descriptions, not physiological thresholds.

Use complete, strictly increasing elapsed checkpoints at 5, 10, 15, 20, 25, 30, 35, 40 and 42.195 km. Clock strings must parse as H:MM:SS or M:SS. No missing splits are interpolated.

Remove exact duplicate race records, ignoring database IDs, ingestion timestamps and source URLs. Retain finishes from 90 minutes to 12 hours with every section between 2 and 20 minutes per km. These quality filters can exclude genuine unusual performances; the analysis describes this eligible cohort, not every entrant.

Full source records are available in public GitHub Releases. These chart tables require at least 100 eligible observations per cell for estimate reliability. Counts refer to finishes, linked pairs or event observations as specified in that answer.

These are observational results. Fitness changes, intentions, training, selection into the dataset and unmeasured conditions can explain differences. Outcome percentiles describe variation among performances, not confidence intervals or advice about the best strategy.

Apply the reviewed source-quality edition exclusions for this exact export after the timing checks. Known invalid split grids, incomplete ingestion, unreconciled HOLD editions and a selected top-finisher field do not contribute to the analyses or prior benchmarks. Report source exclusions separately; an already invalid timing row is not counted twice. Missing age or recorded gender alone does not exclude an otherwise eligible finish from the overall cohort. Other sparse editions are not declared incomplete merely from their size.

18. How does running with a group shape the race?

Chip elapsed times do not establish physical proximity when runners start in different waves. Absolute checkpoint times and individual start offsets are needed to compare group continuity and race outcomes.

Still needed: Absolute checkpoint timestamps and start offsets. Similar chip elapsed times do not establish physical proximity.

19. How much does the opening crowd shape the rest of the race?

A slow first section alone does not establish congestion. Start waves, corrals, clock times and comparable earlier performance are needed.

Still needed: Individual chip and gun times, wave or corral assignments, checkpoint clock times, and dated start arrangements.

20. Are stronger performers better at adjusting their pace to the course?

Normalize section pace by each runner’s own full-marathon average, then take the median by city and performance group. Performance groups use more than 2% faster than, within 2% of, or more than 2% slower than the prior benchmark. Require 100 per city and group.

The outcome group is known after the race. Courses pool editions; neither terrain versions nor race-day conditions are held constant. This is a descriptive course-response profile, not evidence of optimal effort allocation on hills.

For explicitly audited canonical-ID releases, validate unique matching CORE/FULL ID sets and matching recorded edition/name labels, then join by record ID with matching finish and all nine section durations. Legacy exports retain the one-to-one edition, trimmed lowercase name and full-timing join because their IDs are incompatible. Durations are compared to milliseconds. Keep only supplied non-ambiguous runner identities without conflicting recorded gender, inferred birth years spanning more than two years, or duplicate editions. These are candidate cross-race identities, not independently verified people; unlinked runners are absent. The linkage audit records which join was used.

Recompute the benchmark as the fastest eligible finish in the two strictly earlier calendar years. The current race and every other race in its calendar year are excluded. This avoids guessing within-year chronology and prevents current-outcome leakage. It is a recent recorded best, not a fitness measurement, an expected finish, or a lifetime personal best.

Performance change is 100 × (current finish / recent recorded best − 1). Negative is faster. Opening change compares 0–10 km pace with that earlier best’s full-marathon pace. Faster opening: more than 2% faster; similar: within 2%; slower: more than 2% slower. The ±2% and ±5% cutoffs are predefined descriptions, not physiological thresholds.

Use complete, strictly increasing elapsed checkpoints at 5, 10, 15, 20, 25, 30, 35, 40 and 42.195 km. Clock strings must parse as H:MM:SS or M:SS. No missing splits are interpolated.

Remove exact duplicate race records, ignoring database IDs, ingestion timestamps and source URLs. Retain finishes from 90 minutes to 12 hours with every section between 2 and 20 minutes per km. These quality filters can exclude genuine unusual performances; the analysis describes this eligible cohort, not every entrant.

Full source records are available in public GitHub Releases. These chart tables require at least 100 eligible observations per cell for estimate reliability. Counts refer to finishes, linked pairs or event observations as specified in that answer.

These are observational results. Fitness changes, intentions, training, selection into the dataset and unmeasured conditions can explain differences. Outcome percentiles describe variation among performances, not confidence intervals or advice about the best strategy.

Apply the reviewed source-quality edition exclusions for this exact export after the timing checks. Known invalid split grids, incomplete ingestion, unreconciled HOLD editions and a selected top-finisher field do not contribute to the analyses or prior benchmarks. Report source exclusions separately; an already invalid timing row is not counted twice. Missing age or recorded gender alone does not exclude an otherwise eligible finish from the overall cohort. Other sparse editions are not declared incomplete merely from their size.

21. Does being on pace mean you will hit your goal?

On pace means elapsed time within ±1% of the target’s even-pace time at that checkpoint. A goal is achieved only when the finish is strictly under the target. Divide successes by all eligible on-pace finishes for that goal and checkpoint.

This window includes runners just ahead of and just behind the target; it is narrower than the older core table’s time-budget definition. Results pool editions and are descriptive historical frequencies, not a validated personal forecast. Goals are inferred benchmarks, not runners’ declared intentions.

Use complete, strictly increasing elapsed checkpoints at 5, 10, 15, 20, 25, 30, 35, 40 and 42.195 km. Clock strings must parse as H:MM:SS or M:SS. No missing splits are interpolated.

Remove exact duplicate race records, ignoring database IDs, ingestion timestamps and source URLs. Retain finishes from 90 minutes to 12 hours with every section between 2 and 20 minutes per km. These quality filters can exclude genuine unusual performances; the analysis describes this eligible cohort, not every entrant.

Each observation is a race finish. The same person can appear in multiple races. No runner identities are linked across races. Public cells require at least 100 observations; matched comparisons additionally require at least 20 per group in each stratum.

Apply the reviewed source-quality edition exclusions for this exact export after the timing checks. Known invalid split grids, incomplete ingestion, unreconciled HOLD editions and a selected top-finisher field do not contribute to the analyses or prior benchmarks. Report source exclusions separately; an already invalid timing row is not counted twice. Missing age or recorded gender alone does not exclude an otherwise eligible finish from the overall cohort. Other sparse editions are not declared incomplete merely from their size.

22. How much of a marathon is decided after 30 km?

Rank the same complete-record finishers at 30 km and at the finish within each city, year and race. Ties receive their average rank. Positive rank change means gaining places; all signed changes sum to zero within each edition.

Convert rank changes to percentile points using field size minus one. A five-point move corresponds to about 500 positions in a 10,000-person eligible field. Missing-split finishers and non-finishers are absent. Different start waves mean elapsed-time ranks cannot count physical passes.

Use complete, strictly increasing elapsed checkpoints at 5, 10, 15, 20, 25, 30, 35, 40 and 42.195 km. Clock strings must parse as H:MM:SS or M:SS. No missing splits are interpolated.

Remove exact duplicate race records, ignoring database IDs, ingestion timestamps and source URLs. Retain finishes from 90 minutes to 12 hours with every section between 2 and 20 minutes per km. These quality filters can exclude genuine unusual performances; the analysis describes this eligible cohort, not every entrant.

Each observation is a race finish. The same person can appear in multiple races. No runner identities are linked across races. Public cells require at least 100 observations; matched comparisons additionally require at least 20 per group in each stratum.

Apply the reviewed source-quality edition exclusions for this exact export after the timing checks. Known invalid split grids, incomplete ingestion, unreconciled HOLD editions and a selected top-finisher field do not contribute to the analyses or prior benchmarks. Report source exclusions separately; an already invalid timing row is not counted twice. Missing age or recorded gender alone does not exclude an otherwise eligible finish from the overall cohort. Other sparse editions are not declared incomplete merely from their size.

23. How much finishing speed appears when a milestone is within reach?

Project finish time at 40 km by multiplying elapsed time by 42.195/40. Keep projections within five minutes of a round target and divide the margin into four bands. The kick measure is 100 × (40–42.195 km pace / 35–40 km pace − 1); negative means a faster final section.

Plot the median kick and the share finishing strictly below the target. Different courses, ability and fatigue can produce the same projected margin. These are unadjusted associations, not evidence that a milestone caused a sprint.

Use complete, strictly increasing elapsed checkpoints at 5, 10, 15, 20, 25, 30, 35, 40 and 42.195 km. Clock strings must parse as H:MM:SS or M:SS. No missing splits are interpolated.

Remove exact duplicate race records, ignoring database IDs, ingestion timestamps and source URLs. Retain finishes from 90 minutes to 12 hours with every section between 2 and 20 minutes per km. These quality filters can exclude genuine unusual performances; the analysis describes this eligible cohort, not every entrant.

Full source records are available in public GitHub Releases. These chart tables require at least 100 eligible observations per cell for estimate reliability. Counts refer to finishes, linked pairs or event observations as specified in that answer.

These are observational results. Fitness changes, intentions, training, selection into the dataset and unmeasured conditions can explain differences. Outcome percentiles describe variation among performances, not confidence intervals or advice about the best strategy.

Apply the reviewed source-quality edition exclusions for this exact export after the timing checks. Known invalid split grids, incomplete ingestion, unreconciled HOLD editions and a selected top-finisher field do not contribute to the analyses or prior benchmarks. Report source exclusions separately; an already invalid timing row is not counted twice. Missing age or recorded gender alone does not exclude an otherwise eligible finish from the overall cohort. Other sparse editions are not declared incomplete merely from their size.

24. Do qualifying rules change how people race?

The B.A.A. history lists 18–34 standards of 3:05 for men and 3:35 for women for 2013–2019 Boston races; the 2020 standards became 3:00 and 3:30, announced in September 2018. Compare 2016–2017 performances with 2019, excluding the transition year. Exact ages 18–31 avoid crossing the 35-year boundary within the next two years.

Around each fixed old/new benchmark, select finishes within ±2 minutes. Within each city and period, calculate the share at or below the benchmark. Keep cities with at least 20 close finishes in both periods and weight both periods by the smaller city-period count. Report at least 100 observed finishes per displayed period.

Acceptance cutoffs, declared intentions, age on a future Boston race day, certification and qualifying windows are not assigned to individuals. Round-number appeal, historical field changes and anticipatory behavior can explain bunching. This is not a difference-in-differences causal estimate, and the comparison does not cover every rule change.

Official sources: https://www.baa.org/races/boston-marathon/qualify/ ; https://www.baa.org/news/2020-boston-marathon-qualifier-acceptances-announced/ . Historical values verified 2026-09-08.

Use complete, strictly increasing elapsed checkpoints at 5, 10, 15, 20, 25, 30, 35, 40 and 42.195 km. Clock strings must parse as H:MM:SS or M:SS. No missing splits are interpolated.

Remove exact duplicate race records, ignoring database IDs, ingestion timestamps and source URLs. Retain finishes from 90 minutes to 12 hours with every section between 2 and 20 minutes per km. These quality filters can exclude genuine unusual performances; the analysis describes this eligible cohort, not every entrant.

Full source records are available in public GitHub Releases. These chart tables require at least 100 eligible observations per cell for estimate reliability. Counts refer to finishes, linked pairs or event observations as specified in that answer.

These are observational results. Fitness changes, intentions, training, selection into the dataset and unmeasured conditions can explain differences. Outcome percentiles describe variation among performances, not confidence intervals or advice about the best strategy.

Apply the reviewed source-quality edition exclusions for this exact export after the timing checks. Known invalid split grids, incomplete ingestion, unreconciled HOLD editions and a selected top-finisher field do not contribute to the analyses or prior benchmarks. Report source exclusions separately; an already invalid timing row is not counted twice. Missing age or recorded gender alone does not exclude an otherwise eligible finish from the overall cohort. Other sparse editions are not declared incomplete merely from their size.

25. What happens when a goal slips away?

At 20 km keep elapsed time from 99% through 100% of the target’s even-pace budget. Find the first 25, 30, 35 or 40 km checkpoint where elapsed time exceeds that budget. Divide eventual strict sub-target finishes by all such first-crossing observations.

This uses inferred round-time benchmarks and elapsed chip times. It cannot establish when a runner mentally abandoned a goal. First-crossing groups differ, and results include only complete eligible finishes.

Use complete, strictly increasing elapsed checkpoints at 5, 10, 15, 20, 25, 30, 35, 40 and 42.195 km. Clock strings must parse as H:MM:SS or M:SS. No missing splits are interpolated.

Remove exact duplicate race records, ignoring database IDs, ingestion timestamps and source URLs. Retain finishes from 90 minutes to 12 hours with every section between 2 and 20 minutes per km. These quality filters can exclude genuine unusual performances; the analysis describes this eligible cohort, not every entrant.

Full source records are available in public GitHub Releases. These chart tables require at least 100 eligible observations per cell for estimate reliability. Counts refer to finishes, linked pairs or event observations as specified in that answer.

These are observational results. Fitness changes, intentions, training, selection into the dataset and unmeasured conditions can explain differences. Outcome percentiles describe variation among performances, not confidence intervals or advice about the best strategy.

Apply the reviewed source-quality edition exclusions for this exact export after the timing checks. Known invalid split grids, incomplete ingestion, unreconciled HOLD editions and a selected top-finisher field do not contribute to the analyses or prior benchmarks. Report source exclusions separately; an already invalid timing row is not counted twice. Missing age or recorded gender alone does not exclude an otherwise eligible finish from the overall cohort. Other sparse editions are not declared incomplete merely from their size.

26. Does a strong finish suggest unused capacity?

A finishing acceleration is the percentage change in pace from 35–40 to 40–42.195 km. Group it as more than 5% faster, up to 5% faster, or similar/slower. Use consecutive cross-year linked pairs at most three years apart.

Measure next-finish percentage change relative to the first finish, report its 10th/50th/90th percentiles within the first race’s pacing pattern, and the share improving by more than 2%. This does not measure effort reserves, account for absent follow-up, or establish what would have happened with a harder earlier effort.

For explicitly audited canonical-ID releases, validate unique matching CORE/FULL ID sets and matching recorded edition/name labels, then join by record ID with matching finish and all nine section durations. Legacy exports retain the one-to-one edition, trimmed lowercase name and full-timing join because their IDs are incompatible. Durations are compared to milliseconds. Keep only supplied non-ambiguous runner identities without conflicting recorded gender, inferred birth years spanning more than two years, or duplicate editions. These are candidate cross-race identities, not independently verified people; unlinked runners are absent. The linkage audit records which join was used.

Use complete, strictly increasing elapsed checkpoints at 5, 10, 15, 20, 25, 30, 35, 40 and 42.195 km. Clock strings must parse as H:MM:SS or M:SS. No missing splits are interpolated.

Remove exact duplicate race records, ignoring database IDs, ingestion timestamps and source URLs. Retain finishes from 90 minutes to 12 hours with every section between 2 and 20 minutes per km. These quality filters can exclude genuine unusual performances; the analysis describes this eligible cohort, not every entrant.

Full source records are available in public GitHub Releases. These chart tables require at least 100 eligible observations per cell for estimate reliability. Counts refer to finishes, linked pairs or event observations as specified in that answer.

These are observational results. Fitness changes, intentions, training, selection into the dataset and unmeasured conditions can explain differences. Outcome percentiles describe variation among performances, not confidence intervals or advice about the best strategy.

Apply the reviewed source-quality edition exclusions for this exact export after the timing checks. Known invalid split grids, incomplete ingestion, unreconciled HOLD editions and a selected top-finisher field do not contribute to the analyses or prior benchmarks. Report source exclusions separately; an already invalid timing row is not counted twice. Missing age or recorded gender alone does not exclude an otherwise eligible finish from the overall cohort. Other sparse editions are not declared incomplete merely from their size.

27. Do pacing changes follow distance or elapsed time?

Define the first section after 20 km whose average pace is more than 10% slower than 5–20 km. Look only through 40 km so all tested sections are 5 km long. Group runners using their observed 5–20 km pace.

Among runners with such a section, plot the distribution of its end distance. For each distance and early-pace group, show median elapsed arrival at the start and end of that section. The medians bound a typical observation interval; they are not confidence limits. This broader 10% section definition is separate from the sustained slowdown definition.

Use complete, strictly increasing elapsed checkpoints at 5, 10, 15, 20, 25, 30, 35, 40 and 42.195 km. Clock strings must parse as H:MM:SS or M:SS. No missing splits are interpolated.

Remove exact duplicate race records, ignoring database IDs, ingestion timestamps and source URLs. Retain finishes from 90 minutes to 12 hours with every section between 2 and 20 minutes per km. These quality filters can exclude genuine unusual performances; the analysis describes this eligible cohort, not every entrant.

Full source records are available in public GitHub Releases. These chart tables require at least 100 eligible observations per cell for estimate reliability. Counts refer to finishes, linked pairs or event observations as specified in that answer.

These are observational results. Fitness changes, intentions, training, selection into the dataset and unmeasured conditions can explain differences. Outcome percentiles describe variation among performances, not confidence intervals or advice about the best strategy.

Apply the reviewed source-quality edition exclusions for this exact export after the timing checks. Known invalid split grids, incomplete ingestion, unreconciled HOLD editions and a selected top-finisher field do not contribute to the analyses or prior benchmarks. Report source exclusions separately; an already invalid timing row is not counted twice. Missing age or recorded gender alone does not exclude an otherwise eligible finish from the overall cohort. Other sparse editions are not declared incomplete merely from their size.

28. Do runners have persistent pacing habits?

Use consecutive linked finishes in different years, no more than three years apart, with exactly one recorded eligible finish in each endpoint year. Calculate Pearson correlation between the two 0–20 versus 20–40 km pace changes, grouped by calendar-year gap.

For each earlier race pattern, divide the number of next races in each pattern by all eligible pairs with that earlier pattern. These conditional transition percentages sum to 100 within the earlier pattern. Correlation is descriptive; repeat observations of a runner are not independent.

For explicitly audited canonical-ID releases, validate unique matching CORE/FULL ID sets and matching recorded edition/name labels, then join by record ID with matching finish and all nine section durations. Legacy exports retain the one-to-one edition, trimmed lowercase name and full-timing join because their IDs are incompatible. Durations are compared to milliseconds. Keep only supplied non-ambiguous runner identities without conflicting recorded gender, inferred birth years spanning more than two years, or duplicate editions. These are candidate cross-race identities, not independently verified people; unlinked runners are absent. The linkage audit records which join was used.

Use complete, strictly increasing elapsed checkpoints at 5, 10, 15, 20, 25, 30, 35, 40 and 42.195 km. Clock strings must parse as H:MM:SS or M:SS. No missing splits are interpolated.

Remove exact duplicate race records, ignoring database IDs, ingestion timestamps and source URLs. Retain finishes from 90 minutes to 12 hours with every section between 2 and 20 minutes per km. These quality filters can exclude genuine unusual performances; the analysis describes this eligible cohort, not every entrant.

Full source records are available in public GitHub Releases. These chart tables require at least 100 eligible observations per cell for estimate reliability. Counts refer to finishes, linked pairs or event observations as specified in that answer.

These are observational results. Fitness changes, intentions, training, selection into the dataset and unmeasured conditions can explain differences. Outcome percentiles describe variation among performances, not confidence intervals or advice about the best strategy.

Apply the reviewed source-quality edition exclusions for this exact export after the timing checks. Known invalid split grids, incomplete ingestion, unreconciled HOLD editions and a selected top-finisher field do not contribute to the analyses or prior benchmarks. Report source exclusions separately; an already invalid timing row is not counted twice. Missing age or recorded gender alone does not exclude an otherwise eligible finish from the overall cohort. Other sparse editions are not declared incomplete merely from their size.

29. How do speed and pace retention vary by age?

Use exact integer ages from 18 through 89, grouping 18–29 then ten-year bands. Do not infer exact ages from age-group labels. Report median 0–20 km pace and median percentage pace change from 0–20 to 20–40 km, separately by recorded gender.

This is cross-sectional: different people, races and performance levels are being compared. Selection into marathons, missing ages, course and prior ability can explain part of the pattern. It is not an estimate of an individual’s aging trajectory.

Use complete, strictly increasing elapsed checkpoints at 5, 10, 15, 20, 25, 30, 35, 40 and 42.195 km. Clock strings must parse as H:MM:SS or M:SS. No missing splits are interpolated.

Remove exact duplicate race records, ignoring database IDs, ingestion timestamps and source URLs. Retain finishes from 90 minutes to 12 hours with every section between 2 and 20 minutes per km. These quality filters can exclude genuine unusual performances; the analysis describes this eligible cohort, not every entrant.

Each observation is a race finish. The same person can appear in multiple races. No runner identities are linked across races. Public cells require at least 100 observations; matched comparisons additionally require at least 20 per group in each stratum.

Apply the reviewed source-quality edition exclusions for this exact export after the timing checks. Known invalid split grids, incomplete ingestion, unreconciled HOLD editions and a selected top-finisher field do not contribute to the analyses or prior benchmarks. Report source exclusions separately; an already invalid timing row is not counted twice. Missing age or recorded gender alone does not exclude an otherwise eligible finish from the overall cohort. Other sparse editions are not declared incomplete merely from their size.

30. How does pacing differ across recorded gender groups?

Compare average percentage change from the first 20 km to the second 20 km. The pooled chart uses all eligible women’s and men’s records. The matched chart uses only race-edition and one-minute 20 km strata containing at least 20 finishes in each category.

Use the smaller category count as a common stratum weight. Thus both matched estimates have the same race and opening-time distribution. Opening time is an observed race performance, not an independent measure of prior fitness. Age, experience and other factors remain uncontrolled.

The export field is sex; labels follow its recorded categories. Other or missing categories remain in the overall pacing analyses but are not included in this two-category comparison.

Use complete, strictly increasing elapsed checkpoints at 5, 10, 15, 20, 25, 30, 35, 40 and 42.195 km. Clock strings must parse as H:MM:SS or M:SS. No missing splits are interpolated.

Remove exact duplicate race records, ignoring database IDs, ingestion timestamps and source URLs. Retain finishes from 90 minutes to 12 hours with every section between 2 and 20 minutes per km. These quality filters can exclude genuine unusual performances; the analysis describes this eligible cohort, not every entrant.

Each observation is a race finish. The same person can appear in multiple races. No runner identities are linked across races. Public cells require at least 100 observations; matched comparisons additionally require at least 20 per group in each stratum.

Apply the reviewed source-quality edition exclusions for this exact export after the timing checks. Known invalid split grids, incomplete ingestion, unreconciled HOLD editions and a selected top-finisher field do not contribute to the analyses or prior benchmarks. Report source exclusions separately; an already invalid timing row is not counted twice. Missing age or recorded gender alone does not exclude an otherwise eligible finish from the overall cohort. Other sparse editions are not declared incomplete merely from their size.

31. Does a near miss bring people back?

Select linked finishes within two minutes of each round target. Under is strictly below the target; an exact target time belongs to the at-or-over group. Require two subsequent calendar years with at least 100 eligible finishes in the index city, and exclude the latest two observed years as index years.

Return means any eligible linked finish anywhere in the export during the next two calendar years. Same-year returns do not count. Divide returns by every eligible index finish in each group, including those with no observed return. Coverage checks reduce administrative censoring but do not establish complete race ingestion. Repeat index observations, false/missed links, changing coverage and unrecorded goals can affect the association.

For explicitly audited canonical-ID releases, validate unique matching CORE/FULL ID sets and matching recorded edition/name labels, then join by record ID with matching finish and all nine section durations. Legacy exports retain the one-to-one edition, trimmed lowercase name and full-timing join because their IDs are incompatible. Durations are compared to milliseconds. Keep only supplied non-ambiguous runner identities without conflicting recorded gender, inferred birth years spanning more than two years, or duplicate editions. These are candidate cross-race identities, not independently verified people; unlinked runners are absent. The linkage audit records which join was used.

Use complete, strictly increasing elapsed checkpoints at 5, 10, 15, 20, 25, 30, 35, 40 and 42.195 km. Clock strings must parse as H:MM:SS or M:SS. No missing splits are interpolated.

Remove exact duplicate race records, ignoring database IDs, ingestion timestamps and source URLs. Retain finishes from 90 minutes to 12 hours with every section between 2 and 20 minutes per km. These quality filters can exclude genuine unusual performances; the analysis describes this eligible cohort, not every entrant.

Full source records are available in public GitHub Releases. These chart tables require at least 100 eligible observations per cell for estimate reliability. Counts refer to finishes, linked pairs or event observations as specified in that answer.

These are observational results. Fitness changes, intentions, training, selection into the dataset and unmeasured conditions can explain differences. Outcome percentiles describe variation among performances, not confidence intervals or advice about the best strategy.

Apply the reviewed source-quality edition exclusions for this exact export after the timing checks. Known invalid split grids, incomplete ingestion, unreconciled HOLD editions and a selected top-finisher field do not contribute to the analyses or prior benchmarks. Report source exclusions separately; an already invalid timing row is not counted twice. Missing age or recorded gender alone does not exclude an otherwise eligible finish from the overall cohort. Other sparse editions are not declared incomplete merely from their size.

32. What changes as runners gain experience?

Count eligible linked finishes in strictly earlier calendar years. Show median opening pace relative to the recent benchmark and median change between the two 20 km blocks by prior count.

Separately, use consecutive cross-year pairs to calculate the median next-minus-previous pace-retention change, grouped by the previous pattern. Selecting an unusually good or bad first race creates regression to the mean, so improvement after pronounced slowing is not proof of learning. Continued participation and changing fitness also affect these comparisons.

For explicitly audited canonical-ID releases, validate unique matching CORE/FULL ID sets and matching recorded edition/name labels, then join by record ID with matching finish and all nine section durations. Legacy exports retain the one-to-one edition, trimmed lowercase name and full-timing join because their IDs are incompatible. Durations are compared to milliseconds. Keep only supplied non-ambiguous runner identities without conflicting recorded gender, inferred birth years spanning more than two years, or duplicate editions. These are candidate cross-race identities, not independently verified people; unlinked runners are absent. The linkage audit records which join was used.

Recompute the benchmark as the fastest eligible finish in the two strictly earlier calendar years. The current race and every other race in its calendar year are excluded. This avoids guessing within-year chronology and prevents current-outcome leakage. It is a recent recorded best, not a fitness measurement, an expected finish, or a lifetime personal best.

Performance change is 100 × (current finish / recent recorded best − 1). Negative is faster. Opening change compares 0–10 km pace with that earlier best’s full-marathon pace. Faster opening: more than 2% faster; similar: within 2%; slower: more than 2% slower. The ±2% and ±5% cutoffs are predefined descriptions, not physiological thresholds.

Use complete, strictly increasing elapsed checkpoints at 5, 10, 15, 20, 25, 30, 35, 40 and 42.195 km. Clock strings must parse as H:MM:SS or M:SS. No missing splits are interpolated.

Remove exact duplicate race records, ignoring database IDs, ingestion timestamps and source URLs. Retain finishes from 90 minutes to 12 hours with every section between 2 and 20 minutes per km. These quality filters can exclude genuine unusual performances; the analysis describes this eligible cohort, not every entrant.

Full source records are available in public GitHub Releases. These chart tables require at least 100 eligible observations per cell for estimate reliability. Counts refer to finishes, linked pairs or event observations as specified in that answer.

These are observational results. Fitness changes, intentions, training, selection into the dataset and unmeasured conditions can explain differences. Outcome percentiles describe variation among performances, not confidence intervals or advice about the best strategy.

Apply the reviewed source-quality edition exclusions for this exact export after the timing checks. Known invalid split grids, incomplete ingestion, unreconciled HOLD editions and a selected top-finisher field do not contribute to the analyses or prior benchmarks. Report source exclusions separately; an already invalid timing row is not counted twice. Missing age or recorded gender alone does not exclude an otherwise eligible finish from the overall cohort. Other sparse editions are not declared incomplete merely from their size.

33. How does the previous marathon affect the next one?

Use only linked identity groups for which every eligible record has a unique supplied race date. Order by that date, pair adjacent races and retain intervals of 1–1095 days. Unlike the year-based analyses, this includes same-year pairs. Dates come from the supplied calendar overlay and have not all been independently reverified.

Calculate 100 × (next finish / previous finish − 1) and its 10th/50th/90th percentiles in the displayed day bands. The second view requires the earlier race to beat every recorded finish in earlier calendar years; same-year bests are not used for that label. Fitness, course, motivation and selection into short or long intervals remain confounders.

For explicitly audited canonical-ID releases, validate unique matching CORE/FULL ID sets and matching recorded edition/name labels, then join by record ID with matching finish and all nine section durations. Legacy exports retain the one-to-one edition, trimmed lowercase name and full-timing join because their IDs are incompatible. Durations are compared to milliseconds. Keep only supplied non-ambiguous runner identities without conflicting recorded gender, inferred birth years spanning more than two years, or duplicate editions. These are candidate cross-race identities, not independently verified people; unlinked runners are absent. The linkage audit records which join was used.

Use complete, strictly increasing elapsed checkpoints at 5, 10, 15, 20, 25, 30, 35, 40 and 42.195 km. Clock strings must parse as H:MM:SS or M:SS. No missing splits are interpolated.

Remove exact duplicate race records, ignoring database IDs, ingestion timestamps and source URLs. Retain finishes from 90 minutes to 12 hours with every section between 2 and 20 minutes per km. These quality filters can exclude genuine unusual performances; the analysis describes this eligible cohort, not every entrant.

Full source records are available in public GitHub Releases. These chart tables require at least 100 eligible observations per cell for estimate reliability. Counts refer to finishes, linked pairs or event observations as specified in that answer.

These are observational results. Fitness changes, intentions, training, selection into the dataset and unmeasured conditions can explain differences. Outcome percentiles describe variation among performances, not confidence intervals or advice about the best strategy.

Apply the reviewed source-quality edition exclusions for this exact export after the timing checks. Known invalid split grids, incomplete ingestion, unreconciled HOLD editions and a selected top-finisher field do not contribute to the analyses or prior benchmarks. Report source exclusions separately; an already invalid timing row is not counted twice. Missing age or recorded gender alone does not exclude an otherwise eligible finish from the overall cohort. Other sparse editions are not declared incomplete merely from their size.

34. How have marathon speed and pacing changed over time?

For each city and year, calculate median first-20-km pace and median percentage pace change from the first 20 km to the second. Show city-years with at least 100 eligible finishes, and cities with at least three available years.

The same cities, runners and course versions are not represented every year. Choose a city to avoid pooling changing city coverage into one global trend. Even within a city, field composition, route changes and conditions remain uncontrolled. Gaps are years without publishable data, not interpolated observations.

Use complete, strictly increasing elapsed checkpoints at 5, 10, 15, 20, 25, 30, 35, 40 and 42.195 km. Clock strings must parse as H:MM:SS or M:SS. No missing splits are interpolated.

Remove exact duplicate race records, ignoring database IDs, ingestion timestamps and source URLs. Retain finishes from 90 minutes to 12 hours with every section between 2 and 20 minutes per km. These quality filters can exclude genuine unusual performances; the analysis describes this eligible cohort, not every entrant.

Each observation is a race finish. The same person can appear in multiple races. No runner identities are linked across races. Public cells require at least 100 observations; matched comparisons additionally require at least 20 per group in each stratum.

Apply the reviewed source-quality edition exclusions for this exact export after the timing checks. Known invalid split grids, incomplete ingestion, unreconciled HOLD editions and a selected top-finisher field do not contribute to the analyses or prior benchmarks. Report source exclusions separately; an already invalid timing row is not counted twice. Missing age or recorded gender alone does not exclude an otherwise eligible finish from the overall cohort. Other sparse editions are not declared incomplete merely from their size.

35. Where do runners gain the time that produces a personal best?

For each eligible linked finish faster than every eligible recorded finish in strictly earlier calendar years, select the fastest earlier-year record as comparator. Break tied best times by earlier year and then stable raw record ID. Both races must have all nine valid checkpoints.

Subtract current from earlier elapsed time over 0–10, 10–30 and 30–42.195 km. Positive means time gained. Compute means so that the three block means sum exactly to the mean finish improvement; separately divide each block by its own distance to show seconds gained per kilometer. Assert that block gains sum to total gain for every pair and in the published means. Different courses and conditions can contribute to gains.

For explicitly audited canonical-ID releases, validate unique matching CORE/FULL ID sets and matching recorded edition/name labels, then join by record ID with matching finish and all nine section durations. Legacy exports retain the one-to-one edition, trimmed lowercase name and full-timing join because their IDs are incompatible. Durations are compared to milliseconds. Keep only supplied non-ambiguous runner identities without conflicting recorded gender, inferred birth years spanning more than two years, or duplicate editions. These are candidate cross-race identities, not independently verified people; unlinked runners are absent. The linkage audit records which join was used.

Use complete, strictly increasing elapsed checkpoints at 5, 10, 15, 20, 25, 30, 35, 40 and 42.195 km. Clock strings must parse as H:MM:SS or M:SS. No missing splits are interpolated.

Remove exact duplicate race records, ignoring database IDs, ingestion timestamps and source URLs. Retain finishes from 90 minutes to 12 hours with every section between 2 and 20 minutes per km. These quality filters can exclude genuine unusual performances; the analysis describes this eligible cohort, not every entrant.

Full source records are available in public GitHub Releases. These chart tables require at least 100 eligible observations per cell for estimate reliability. Counts refer to finishes, linked pairs or event observations as specified in that answer.

These are observational results. Fitness changes, intentions, training, selection into the dataset and unmeasured conditions can explain differences. Outcome percentiles describe variation among performances, not confidence intervals or advice about the best strategy.

Apply the reviewed source-quality edition exclusions for this exact export after the timing checks. Known invalid split grids, incomplete ingestion, unreconciled HOLD editions and a selected top-finisher field do not contribute to the analyses or prior benchmarks. Report source exclusions separately; an already invalid timing row is not counted twice. Missing age or recorded gender alone does not exclude an otherwise eligible finish from the overall cohort. Other sparse editions are not declared incomplete merely from their size.

Focused analysis: sustained slowdown

At least 25% slower than the 5–20 km reference pace for contiguous recorded sections totaling at least 5 km after 20 km. The final section is 2.195 km. Onset is a section boundary, not an exact moment.

This definition identifies sustained slowing. It cannot determine whether the cause was fuel depletion, injury, fatigue, walking, or another factor. It is one outcome within the wider pacing study.

The definition follows the Published slowdown method (2021). Explore the sustained slowdown analysis.

Additional web data that could strengthen the study

Start with dated course routes, hourly weather, and start times. These are proposed sources for extending and validating the dataset; availability here does not mean every source has already been joined.

Hourly weather

Temperature, dew point, humidity, rain, wind, cloud cover, and solar radiation. Match conditions to the time each runner reaches a segment. Use a consistent historical model across years.

Broad historical coverage; modeled grid estimates, not conditions measured at the runner. Open-Meteo historical weather.

Weather-station observations

Observed temperature, dew point, wind, and precipitation, with station location and quality flags. Check unusual weather days against observations near the course.

Station and year coverage vary. The newer GHCN hourly archive replaces ISD. NOAA GHCN hourly.

The route for each race edition

Course geometry, checkpoint locations, certification ID, route changes, and separate start routes. Calculate section distance, turns, road direction, and the route actually used that year.

US courses; use each organizer’s dated maps elsewhere. Older routes often need manual recovery. USATF course certification database.

Elevation and slope by section

Climb, descent, net elevation change, and grade along the route. Separate terrain-related pacing patterns from a runner’s unusual slowdown.

Available for route coordinates. Check bridges and tunnels separately: terrain height may differ from the road deck. Open Topo Data / SRTM.

Waves, corrals, and actual start times

Wave schedule, runner corral, chip start, gun finish, and timing conventions. Estimate time-of-day exposure and establish who was together at a checkpoint.

Schedules are commonly published; individual start timestamps depend on the timing provider. A wave start is not an individual start. Official participant information.

Aid stations and course amenities

Water and fuel locations, supplied products, medical stations, toilets, and station changes by year. Compare local pacing patterns with where runners can stop or refuel.

Often available in participant guides. Product availability does not reveal what any runner consumed. Official course and aid-station guide.

Qualifying rules and entry routes

Published time standards, eligible age, qualifying window, acceptance cutoff, and entry category where public. Study goal incentives and account for differences in who enters each race.

Use dated rules and announcements. Historical records need an edition-by-edition audit. Boston Athletic Association.

Air quality

Particle pollution, ozone, and other pollutants at the race location and time. Explore whether poor-air-quality editions have different pacing patterns.

Historical modeled coverage from 2003; spatial resolution is too coarse to represent every street. Copernicus CAMS reanalysis.

More fields to retain from official results

Keep all published timing points, including halfway and the finish; bib and provider IDs; reported age and category; gun and chip times; official finish, withdrawal, disqualification, and non-start statuses; and result corrections. Availability varies by organizer and year.

Do not infer a withdrawal from one missing timing read, or a debut marathon from a runner’s first appearance in this database.

Preserve the evidence

For every added field, retain its source URL, race edition, retrieval date, original unit, and whether it was observed, modeled, or inferred. Keep one weather record per place and time, then join it to estimated segment exposure. Between timing mats, a runner’s exact location is an estimate.

Hourly wind plus route direction can estimate headwind exposure. Route geometry can estimate turn counts. Crowd density, shade, training, shoes, and individual fueling are harder to reconstruct consistently over 20 years and should not be assumed.