Assessing spatial transmission risk of respiratory infectious diseases across cities of different socioeconomic tiers in China: A modelling study
Plos.org·July 20, 2026
AI Summary
Researchers modeled how respiratory infectious diseases spread across Chinese cities at varying socioeconomic levels to understand spatial transmission patterns of respiratory pathogens. The study reveals how disease diffusion dynamics may differ between economically developed and less developed urban areas in mainland China.
Understanding how respiratory infectious diseases spread across cities of different socioeconomic tiers is crucial for regionally targeted interventions. However, most spatial prediction frameworks neglect the combined influence of urban hierarchy and human mobility in shaping transmission risk.
We integrated large-scale intercity mobility data into an agent-based branching process model to simulate the spatial diffusion of respiratory pathogens across mainland China. Three COVID-19 outbreaks were used for validation: the Omicron outbreak in Shanghai, the Delta outbreak in Nanjing, and a multi-provincial Delta outbreak in northwestern China. We applied the framework to model the spread of SARS-CoV-2 (Omicron variant) and Influenza A to quantify tier-specific transmission risks. Tiers denote a hierarchical classification of Chinese cities based on concentration of commercial resources, transportation hub centrality, etc., ranging from super-tier metropolises to lower-tier cities. Predicted first arrival times showed strong agreement with observed data (r = 0.68 and 0.76), and mobility-based predictions more accurately identified outbreak origins than distance-based approaches. Markedly different tier-dependent diffusion patterns were observed across pathogens. Influenza A exhibited stable and stratified diffusion, with transmission confined mainly within the same or adjacent urban tiers and limited cross-tier seeding. In contrast, SARS-CoV-2 (Omicron) initially concentrated in super-tier and tier-1 cities but rapidly spread to lower-tier cities, producing a pronounced hierarchical pattern of spread that quickly diminished tier-level differences in transmission risk. Across pathogens, higher-tier cities consistently faced greater early importation risk; however, this disparity persisted for Influenza A but was rapidly attenuated for Omicron due to its high transmissibility and fast spatial expansion. A key limitation is that the model was validated against city-level first arrival times rather than full epidemic dynamics, and was parameterised using mobility data from China’s “dynamic zero-COVID” period, which may limit direct quantitative generalisability to other settings.
Full story reconstructed from Plos.org. Formatting and media may differ from the original.
Spatial transmission risk reflects an interaction between pathogen-specific transmissibility and the hierarchical organisation of China’s urban mobility system. These findings indicate that surveillance and response strategies effective for less transmissible pathogens may be insufficient for highly transmissible variants. Mobility-informed, tier-specific risk assessment may help inform early warning and support more adaptive public health responses. Because the framework relies on routinely available mobility data and minimal pathogen-specific inputs, it may provide a scalable approach for epidemic preparedness, including potential future emerging respiratory threats (Disease X).
Human mobility plays a key role in how infectious diseases spread between cities, but the distribution of early transmission risk within a country remains incompletely understood.
Most existing studies rely on indirect or aggregated data, making it difficult to directly assess how well models capture the early spatial spread of infections across cities.
While respiratory pathogens differ in their transmission characteristics, there remains limited empirical evidence on how these differences translate into distinct spatial spread patterns within the same mobility system.
We developed a model that combines human mobility data with disease transmission to simulate how infections spread between cities.
Using detailed data from COVID-19 outbreaks in China, we validated the model against observed spread patterns and showed that it can reproduce the timing of infection first appearing in different cities.
We found that transmission risk is structured by city hierarchy and that different pathogens (e.g., COVID-19 variants versus influenza) spread across the same mobility network in markedly different ways.
Early transmission risk is not evenly distributed: some cities play a disproportionately important role in driving spread.
Public health responses may need to be tailored to both city characteristics and pathogen type, rather than to a uniform strategy.
These results are based on data from a specific setting with unusually detailed surveillance and may not fully generalise to all contexts.
Citation: Li W, Yang W, Liu Y, Yao Y (2026) Assessing spatial transmission risk of respiratory infectious diseases across cities of different socioeconomic tiers in China: A modelling study. PLoS Med 23(7): e1005172. https://doi.org/10.1371/journal.pmed.1005172
Academic Editor: Simon Cauchemez, Institut Pasteur, FRANCE
Received: December 22, 2025; Accepted: June 24, 2026; Published: July 20, 2026
Data Availability: All datasets used in this study, including Baidu Migration data, first confirmed case data, and processed analytical inputs, are publicly available in the Zenodo repository (https://doi.org/10.5281/zenodo.17852176). The analysis code is publicly available on GitHub (https://github.com/OOLWJOO/Diffusion_Based_on_Baidu_Migration) and archived in Zenodo (https://doi.org/10.5281/zenodo.17852176).
Funding: This research is supported by the National Key R&D Program of China (No. 2025YFA1016500, https://www.htrdc.com/gjszx/) (to W.Y. and Y.Y.), Prevention and Control of Emerging and Major Infectious Diseases-National Science and Technology Major Project (2025ZD01900400, http://www.most.gov.cn/) (to Y.Y.), and Shanghai Institute for Mathematics and Interdisciplinary Sciences (SIMIS, SIMIS-ID-2025-NW, https://www.simis.cn/) (to W.Y. and Y.Y.). The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript.
Competing interests: The authors have declared that no competing interests exist.
Abbreviations: MAE, mean absolute error; NAAT, nucleic acid amplification testing; NPIs, non-pharmaceutical interventions; PHSM, public health and social measures; RMSE, root mean square error; SI, serial interval
Pandemics of respiratory infectious diseases continue to challenge global health security in an increasingly interconnected world [1–3]. Effective prevention and control hinge on the ability to anticipate when and where infections will spread, so that interventions can be deployed with precision rather than uniformly across regions. Such targeted prevention and control, the spatial counterpart of risk-based population strategies, is essential for minimising unwarranted disruptions while containing outbreaks [4,5]. The success of spatially targeted interventions relies critically on the ability to generate accurate early predictions of disease diffusion across urban networks. However, accurately anticipating the spatial trajectory of spread remains challenging in real-world settings [6,7].
Traditional approaches to modelling intercity spread have often relied on geographical distance as a proxy for connectivity, assuming that shorter distances correspond to higher transmission risk [8,9]. Nevertheless, in the context of modern transportation systems and dense travel networks, geographical closeness no longer adequately represents the actual flow of people and pathogens [10]. Empirical evidence increasingly shows that mobility-based connectivity, derived from real population movements, better explains observed diffusion patterns than static distance measures [11–13]. The use of large-scale mobility data thus provides a data-driven foundation for improving the accuracy of spatiotemporal predictions, particularly during the early stages of emerging outbreaks when rapid assessment is critical for decision-making.
At the same time, effective spatial targeting of interventions also requires an understanding of how disease diffusion patterns differ across pathogens. Fast-spreading viruses such as SARS-CoV-2 may rapidly traverse intercity mobility networks, necessitating early interregional coordination and broad geographical alert zones [14–16]. In contrast, pathogens with lower transmissibility, such as seasonal influenza, may exhibit more stratified or locally contained spatial trajectories [17,18], enabling more focussed and efficient control measures. In countries with substantial socioeconomic heterogeneity and hierarchical urban systems, as characterised, for example, by Brazil’s IBGE Regiões de Influência das Cidades [19], Indonesia’s BPS Human Development Index-based classification [20], India’s Ministry of Finance tier classification [21], and the United States’ Brookings Metro Monitor rankings [22], integrating mobility-informed risk assessment into epidemic preparedness frameworks provides a practical pathway toward achieving precision and efficiency in outbreak response.
A central limitation in validating mobility-driven models of respiratory disease spread is the difficulty of directly observing intercity infector-infectee relationships, such that validation typically relies on statistically inferred transmission patterns. Here, we leverage a uniquely detailed dataset from three COVID-19 outbreaks in China to enable validation against directly observed transmission events. More broadly, this study makes three contributions. First, it provides high-fidelity empirical validation of mobility-driven diffusion models using directly observed transmission events, strengthening the empirical grounding of such approaches. Second, it quantifies how early transmission risk is structured across urban hierarchies within a country, generating sub-national, tier-stratified evidence to inform regionally differentiated public health strategies. Third, it explicitly incorporates a direct empirical comparison of how a shared mobility network shapes spatial diffusion across respiratory pathogens with contrasting transmission characteristics, thereby informing an important and previously unresolved policy question of whether internal movement restrictions and other regionally targeted measures are likely to exert comparable effects across different respiratory pathogens.
To address these challenges, this study introduces a mobility-informed analytical framework to understand and predict the spatial diffusion of respiratory infectious diseases. By integrating large-scale intercity movement data into an agent-based branching process framework, we aim to capture how real human mobility shapes epidemic trajectories. Using SARS-CoV-2 and Influenza A as representative pathogens, we further assess the spatial heterogeneity of transmission risks across urban tiers, providing insights for early warning and targeted intervention. This framework lays the foundation for scalable, data-driven approaches to epidemic preparedness and response. Specifically, we asked whether intercity human mobility alone can reproduce and predict the early city-level spatial spread of respiratory pathogens in China, and whether pathogens with contrasting transmissibility (SARS-CoV-2 Omicron versus Influenza A) generate systematically different tier-stratified transmission risks across the same mobility network.
Baidu is one of China’s largest technology companies. Baidu’s Baidu Maps application provides location-based services to over 1.1 billion mobile device users in mainland China. After anonymisation and aggregation, these data are converted into a daily city-level migration index, which is publicly available on Baidu’s platform, Baidu "Qianxi" (which means "migration" in Chinese) [23]. We extracted the migration index for 366 cities in mainland China between 1 January 2021 and 5 April 2022. For each city and day, we obtained both the outbound migration scale index and the proportion of outflow directed to each destination city. We combined these two components to estimate city-to-city mobility flows and then normalised these flows to derive travel probabilities between origin-destination pairs (Method C in S1 Appendix). These probabilities were assembled into city-to-city mobility transition matrices, which were used to parameterise intercity movement in the transmission model. We constrained our analysis to this period to capture population mobility patterns during the routine response phase of the ‘dynamic zero-COVID’ strategy, which represents a scenario of heightened alertness but relatively stable intercity travel compared to city-wide lockdowns.
Data on daily new infections during three COVID-19 outbreaks in 2021 and 2022 in mainland China were collected from the official press releases of the National Health Commission of China. In the context of this study, we define the start of an outbreak as the detection of the first case, and the end of an outbreak as the last confirmed case or the time of substantial public health and social measures (PHSM) introduction. Using these definitions, the three outbreaks lasted from 1 March to 5 April 2022, 20 July to 5 August 2021, and 17 October to 3 November 2021 (Methods A and B in S1 Appendix, Tables E–G in S1 Appendix). The first cases of these outbreaks were detected in Shanghai, Nanjing, and northwestern China, respectively. These three outbreaks involved 121, 28, and 26 cities overall.
We developed an agent-based branching process model on the city level to capture the spatial propagation of COVID-19 outbreaks (Fig 1). In the agent-based branching process model, the number of secondary infections generated by an infected individual followed a negative binomial distribution with a mean equal to the reproduction number (R0), capturing individual heterogeneity in transmissibility. The infection time for each potential secondary case was drawn from the serial interval (SI) distribution (Method E in S1 Appendix), Gamma for SARS-CoV-2 and Weibull for seasonal influenza A (Table 1). Human mobility was incorporated through a sigmoid function describing the probability of travel as a function of the local daily migration index (Method D in S1 Appendix). When travel occurred, the destination city was selected based on the observed proportions of outbound travellers, as determined from Baidu Migration Index data (Method C in S1 Appendix). An example of this process is illustrated in Fig 1: an infected individual A in City 1 generated three secondary infections (B, D, and G); B then travelled to City 2 and infected two additional cases (C and E); and C travelled to City 3 and produced one further infection (F). This example shows how infection events and travel behaviour jointly shape the spatial spread of infection.
https://doi.org/10.1371/journal.pmed.1005172.t001
(A) A conceptual diagram of the branching process model; (B) the probability mass function (pmf) capturing the uncertainty around the reproduction number; (C) the probability density function (pdf) capturing the uncertainty around the serial interval. In (A), person A infected three individuals, B, D, and G. This level of transmissibility, in implementation, is drawn from the probability distribution in (B). Individuals D and G do not travel and, therefore, would be reported in city 1 as new infections. Individual B, however, travelled to city 2 because on one of the days before symptom onset, they were sampled to travel based on the migration index from city 1 to city 2. The total number of days before symptom onset will be randomly sampled from the probability distribution presented in (C). In this example, the first arrival time of City 2 corresponded to the serial interval of individual B, as B represented the first introduction into City 2. Similarly, the first arrival time of City 3 was determined by the sum of the serial intervals of individuals B and C, since the infection chain A → B → C constituted the earliest seeding pathway to City 3.
https://doi.org/10.1371/journal.pmed.1005172.g001
For clarity, the main steps of the spatial agent-based branching process are summarised in the following pseudocode.
Algorithm: Spatial agent-based branching process simulation
Input: seed city o, initial infections N0, time window W, offspring distribution, serial interval distribution, mobility matrix P
Output: first arrival time FA(c) for each city c
Initialise N0 infections in city o on day 1. Set FA(o) = 1 and FA(c) = undefined for all other cities. Add initial infections to queue Q
Draw the number of secondary infections r
Assign destination city c according to P
Repeat over stochastic simulations and summarise city-specific first arrival times
To quantify the spatiotemporal dynamics of disease spread, we simulated the first arrival time in each city using the agent-based branching process model described above. The first arrival time was defined as the elapsed time between the onset of the outbreak in the source city and the first infection in the destination city. During each simulation, all infection and travel events were recorded. When an infected individual travelled from their original city to another city, the infection time of that individual determined the first arrival time of the destination city. If multiple infected individuals arrived in the same city, the earliest infection time among them was defined as that city’s first arrival time. A simple example was illustrated in Fig 1. To ensure that the simulated first arrival times reflect spatial diffusion driven purely by population mobility rather than disease-specific transmission parameters, this component of the simulation omitted pathogen-specific values of reproduction number (R0) and SI, allowing infections to propagate solely through mobility processes. To provide a more comprehensive validation framework, we additionally conducted an extended validation analysis that incorporated outbreak-specific reproduction numbers and serial intervals, with the results presented in Figs A and B in S1 Appendix. Each outbreak-origin scenario was simulated 10,000 times to account for stochastic variation in transmission and travel processes (Method F in S1 Appendix). For each simulation, we recorded the first arrival time for every city that was seeded within the simulation horizon. Cities that were not seeded during the simulation period were classified as not reached. For each city, the predicted first arrival time used in the main analysis was summarised as the mean across simulations, which provides a stable estimate of expected arrival timing under stochastic variability. Model performance was evaluated by comparing the observed first arrival times with the mean predicted first arrival times using Pearson correlation, mean absolute error (MAE), and root mean square error (RMSE).
First arrival time was chosen as the primary validation target because it matches the policy frame our evidence is intended to inform: in China, responses to respiratory epidemics are typically enacted at the city level and initiated by a local trigger, so the operationally meaningful question for city-level practitioners is whether an outbreak has arrived—not what proportion of the country is affected or how fast it is spreading spatially. First arrival timing also more directly reflects intercity diffusion on the mobility network and is less confounded by city-specific differences in prior infection, vaccination, local interventions, and climatic conditions than metrics based on peak incidence, cumulative incidence, epidemic intensity, or final attack rate.
We first used observed COVID-19 outbreaks to validate the model’s ability to reproduce real-world spatial diffusion patterns. We then applied the validated framework to Influenza A as a hypothetical scenario to illustrate its use in comparing spatial transmission risk across urban tiers under a different respiratory infection context. We conducted a risk assessment by simulating the agent-based branching process model to explore two scenarios involving two pathogens: SARS-CoV-2 (Omicron) and Influenza A. The parameters we used for these pathogens are presented in Table 1. We assumed there were 100 cases introduced initially and ran the simulation for 21 days. We defined reproduction numbers according to their epidemiological interpretation: denotes the basic reproduction number in a fully susceptible population without interventions, whereas denotes the effective reproduction number under prevailing epidemiological and control conditions. For the COVID-19 Omicron validation analyses, we therefore referred to Re. For the hypothetical Influenza A simulations, the reproduction parameter represents an assumed baseline transmissibility parameter analogous to under the modelled conditions.
For the baseline analyses, serial interval distributions were parameterised using literature-based point estimates. In each simulation, however, the serial interval for each transmission event was randomly drawn from the specified distribution. Uncertainty in the serial interval distribution parameters themselves was examined separately in sensitivity analyses, including the joint uncertainty analysis, in which alternative parameter combinations were sampled over plausible ranges.
We initiated outbreaks in each of the 366 prefectures under consideration at the beginning of each outbreak, thereby obtaining 366 sets of 366 predicted arrival times. We compared the simulated first arrival time pattern generated by each candidate origin with the observed first arrival time pattern across affected cities. This comparison was performed using the Pearson correlation between the logarithm of the predicted first arrival time () and the logarithm of the observed first arrival time (). We used logarithmically transformed first arrival times for the mobility-based analyses because the relationship between the diffusion time () and diffusion distance () follows a power-law form . In the origin-identification analysis, the Pearson correlation on log-transformed arrival times was used primarily to compare the similarity between the observed and predicted arrival time patterns across cities and to preserve information on their relative temporal spacing (Method I in S1 Appendix). The prefecture with the highest Pearson correlation coefficient for each outbreak was identified as its origin. The origins of the Shanghai and Nanjing outbreaks are known through outbreak investigation efforts; therefore, these outbreaks are used to validate this approach (Method K in S1 Appendix).
We compared this approach to using geodesic distance to identify outbreak origins. The geodesic distance between a pair of cities was obtained using the Haversine formula, calculated from the latitude and longitude of the two cities. In the context of SARS-CoV-2 transmission, we do not know the functional relationship between diffusion distance and time. Thus, we used the Spearman correlation between geodesic distance and raw observed arrival times to approximate the association between predicted and observed arrival times (Method J in S1 Appendix).
Using the validated agent-based branching process model, we simulated the propagation of two pathogens in mainland China, using their and serial intervals. We assumed 100 initial cases at the origin and simulated the outbreak for 7, 14, and 21 days. We introduce to quantify the potential scope of an outbreak, which denotes the number of people at risk. The entire population of a city is considered at risk once the pathogen has been transmitted to that city. Therefore, we have the following relationship:
where is the number of residents of city ; is a binary variable indicating if city j has received infectious individuals from city . Alternative definitions of based on importation thresholds are detailed in Method O in S1 Appendix.
For each origin city, the risk metric was estimated as the mean across 10,000 stochastic simulations. Destination cities that were not seeded within the simulation horizon in a given simulation were not counted as contributing to transmission risk in that replicate. We assumed 100 initial cases as a standardised seeding condition for the risk simulations to enable stable and comparable estimates of relative transmission risk across cities; sensitivity analyses using alternative initial case numbers showed similar overall patterns.