Search
2026 Volume 5
Article Contents
ARTICLE   Open Access    

Relationship analysis between traffic status, risky driving behaviors, and crash probability in spatiotemporal

More Information
  • Studying the characteristics of traffic status and risky driving behaviors during crashes, identifying traffic crash risks, and then taking proactive measures to control them on the road are crucial to reducing traffic crash rates. In this study, we used in-vehicle navigation data and extracted traffic status and risky driving behavior data for the 30 min before a crash across the upstream, middle, and downstream segments. A dynamic Bayesian network model was used to determine the relationships between traffic status, risky driving behaviors, and crash probability in time and space. Temporally, a higher crash probability was observed in the two time slices, T3 and T4. Spatially, changes in the traffic status and risky driving behaviors in the downstream segment were associated with higher crash probabilities. This lays the foundation for dynamically identifying pre-accident risk states using vehicle navigation data, providing technical support for the transformation of traffic safety management from passive response to active prevention. In addition, the findings help identify high-risk road segments and critical risk periods from both temporal and spatial dimensions, providing a basis for precise interventions such as variable message signs, in-vehicle warnings, speed control, and traffic organization optimization.
  • 加载中
  • [1] WHO. 2017. Road traffic injuries. www.who.int/en/news-room/fact-sheets/detail/road-traffic-injuries (accessed 2019 Jul 18)
    [2] Liu Q, Li C, Jiang H, Nie S, Chen L. 2022. Transfer learning-based highway crash risk evaluation considering manifold characteristics of traffic flow. Accident Analysis & Prevention 168:106598 doi: 10.1016/j.aap.2022.106598

    CrossRef   Google Scholar

    [3] Jiang F, Yuen KKR, Lee EWM. 2020. A long short-term memory-based framework for crash detection on freeways with traffic data of different temporal resolutions. Accident Analysis & Prevention 141:105520 doi: 10.1016/j.aap.2020.105520

    CrossRef   Google Scholar

    [4] Sun J, Sun J. 2015. A dynamic Bayesian network model for real-time crash prediction using traffic speed conditions data. Transportation Research Part C: Emerging Technologies 54:176−186 doi: 10.1016/j.trc.2015.03.006

    CrossRef   Google Scholar

    [5] Abdel-Aty MA, Hassan HM, Ahmed M, Al-Ghamdi AS. 2012. Real-time prediction of visibility related crashes. Transportation Research Part C: Emerging Technologies 24:288−298 doi: 10.1016/j.trc.2012.04.001

    CrossRef   Google Scholar

    [6] Golob TF, Recker WW, Alvarez VM. 2004. Freeway safety as a function of traffic flow. Accident Analysis & Prevention 36:933−946 doi: 10.1016/j.aap.2003.09.006

    CrossRef   Google Scholar

    [7] Lee C, Hellinga B, Saccomanno F. 2003. Real-time crash prediction model for application to crash prevention in freeway traffic. Transportation Research Record: Journal of the Transportation Research Board 1840:67−77 doi: 10.3141/1840-08

    CrossRef   Google Scholar

    [8] Li Z, Wang W, Chen R, Liu P, Xu C. 2013. Evaluation of the impacts of speed variation on freeway traffic collisions in various traffic states. Traffic Injury Prevention 14:861−866 doi: 10.1080/15389588.2013.775433

    CrossRef   Google Scholar

    [9] Xu C, Liu P, Wang W, Li Z. 2012. Evaluation of the impacts of traffic states on crash risks on freeways. Accident Analysis & Prevention 47:162−171 doi: 10.1016/j.aap.2012.01.020

    CrossRef   Google Scholar

    [10] BDR. 2010. China mobile security market research report. www.sohu.com/a/358736654_783965
    [11] Mihajlovic V, Petkovic M. 2001. Dynamic Bayesian networks: a state of the art. CTIT Tech. reports Ser. 34. University of Twente, Enschede. pp. 1–37
    [12] Sammut C, Webb GI. 2017. Dynamic Bayesian network. In Encyclopedia of Machine Learning and Data Mining, eds Sammut C. Webb GI. Boston, MA: Springer US. pp. 377 doi: 10.1007/978-1-4899-7687-1_100125
    [13] Hou Q, Tarko AP, Meng X. 2018. Investigating factors of crash frequency with random effects and random parameters models: new insights from Chinese freeway study. Accident Analysis & Prevention 120:1−12 doi: 10.1016/j.aap.2018.07.010

    CrossRef   Google Scholar

    [14] Huang H, Peng Y, Wang J, Luo Q, Li X. 2018. Interactive risk analysis on crash injury severity at a mountainous freeway with tunnel groups in China. Accident Analysis & Prevention 111:56−62 doi: 10.1016/j.aap.2017.11.024

    CrossRef   Google Scholar

    [15] Ma Z, Zhang H, Chien SIJ, Wang J, Dong C. 2017. Predicting expressway crash frequency using a random effect negative binomial model: a case study in China. Accident Analysis & Prevention 98:214−222 doi: 10.1016/j.aap.2016.10.012

    CrossRef   Google Scholar

    [16] Zheng L, Sun J, Meng X. 2018. Crash prediction model for basic freeway segments incorporating influence of road geometrics and traffic signs. Journal of Transportation Engineering, Part A: Systems 144:04018030 doi: 10.1061/JTEPBS.0000155

    CrossRef   Google Scholar

    [17] Shi Q, Abdel-Aty M. 2015. Big Data applications in real-time traffic operation and safety monitoring and improvement on urban expressways. Transportation Research Part C: Emerging Technologies 58:380−394 doi: 10.1016/j.trc.2015.02.022

    CrossRef   Google Scholar

    [18] Pande A, Das A, Abdel-Aty M, Hassan H. 2011. Estimation of real-time crash risk: are all freeways created equal? Transportation Research Record: Journal of the Transportation Research Board 2237:60−66 doi: 10.3141/2237-07

    CrossRef   Google Scholar

    [19] Abdel-Aty MA, Pemmanaboina R. 2006. Calibrating a real-time traffic crash-prediction model using archived weather and ITS traffic data. IEEE Transactions on Intelligent Transportation Systems 7:167−174 doi: 10.1109/TITS.2006.874710

    CrossRef   Google Scholar

    [20] Lee C, Hellinga B, Saccomanno F. 2003. Proactive freeway crash prevention using real-time traffic control. Canadian Journal of Civil Engineering 30:1034−1041 doi: 10.1139/l03-040

    CrossRef   Google Scholar

    [21] Xu C, Wang W, Liu P. 2013. Identifying crash-prone traffic conditions under different weather on freeways. Journal of Safety Research 46:135−144 doi: 10.1016/j.jsr.2013.04.007

    CrossRef   Google Scholar

    [22] Parsa AB, Movahedi A, Taghipour H, Derrible S, Mohammadian AK. 2020. Toward safer highways, application of XGBoost and SHAP for real-time accident detection and feature analysis. Accident Analysis & Prevention 136:105405 doi: 10.1016/j.aap.2019.105405

    CrossRef   Google Scholar

    [23] Sun J, Sun J, Chen P. 2014. Use of support vector machine models for real-time prediction of crash risk on urban expressways. Transportation Research Record: Journal of the Transportation Research Board 2432:91−98 doi: 10.3141/2432-11

    CrossRef   Google Scholar

    [24] Xu C, Wang W, Liu P, Zhang F. 2015. Development of a real-time crash risk prediction model incorporating the various crash mechanisms across different traffic states. Traffic Injury Prevention 16:28−35 doi: 10.1080/15389588.2014.909036

    CrossRef   Google Scholar

    [25] Yuan J, Abdel-Aty M, Gong Y, Cai Q. 2019. Real-time crash risk prediction using long short-term memory recurrent neural network. Transportation Research Record: Journal of the Transportation Research Board 2673:314−326 doi: 10.1177/0361198119840611

    CrossRef   Google Scholar

    [26] Yang Y, He K, Wang YP, Yuan ZZ, Yin YH, et al. 2022. Identification of dynamic traffic crash risk for cross-area freeways based on statistical and machine learning methods. Physica A: Statistical Mechanics and Its Applications 595:127083 doi: 10.1016/j.physa.2022.127083

    CrossRef   Google Scholar

    [27] Kidando E, Kitali AE, Kutela B, Ghorbanzadeh M, Karaer A, et al. 2021. Prediction of vehicle occupants injury at signalized intersections using real-time traffic and signal data. Accident Analysis & Prevention 149:105869 doi: 10.1016/j.aap.2020.105869

    CrossRef   Google Scholar

    [28] Abdel-Aty M, Uddin N, Pande A, Abdalla MF, Hsia L. 2004. Predicting freeway crashes from loop detector data by matched case-control logistic regression. Transportation Research Record: Journal of the Transportation Research Board 1897:88−95 doi: 10.3141/1897-12

    CrossRef   Google Scholar

    [29] Liu M, Chen Y. 2017. Predicting real-time crash risk for urban expressways in China. Mathematical Problems in Engineering 2017:6263726 doi: 10.1155/2017/6263726

    CrossRef   Google Scholar

    [30] Wang J, Song H, Fu T, Behan M, Jie L, et al. 2022. Crash prediction for freeway work zones in real time: a comparison between Convolutional Neural Network and Binary Logistic Regression model. International Journal of Transportation Science and Technology 11:484−495 doi: 10.1016/j.ijtst.2021.06.002

    CrossRef   Google Scholar

    [31] Li X, Lord D, Zhang Y, Xie Y. 2008. Predicting motor vehicle crashes using Support Vector Machine models. Accident Analysis & Prevention 40:1611−1618 doi: 10.1016/j.aap.2008.04.010

    CrossRef   Google Scholar

    [32] Zeng Q, Huang H, Pei X, Wong SC, Gao M. 2016. Rule extraction from an optimized neural network for traffic crash frequency modeling. Accident Analysis & Prevention 97:87−95 doi: 10.1016/j.aap.2016.08.017

    CrossRef   Google Scholar

    [33] Cai Q, Abdel-Aty M, Yuan J, Lee J, Wu Y. 2020. Real-time crash prediction on expressways using deep generative models. Transportation Research Part C: Emerging Technologies 117:102697 doi: 10.1016/j.trc.2020.102697

    CrossRef   Google Scholar

    [34] Guo M, Zhao X, Yao Y, Bi C, Su Y. 2022. Application of risky driving behavior in crash detection and analysis. Physica A: Statistical Mechanics and Its Applications 591:126808 doi: 10.1016/j.physa.2021.126808

    CrossRef   Google Scholar

    [35] Guo M, Zhao X, Yao Y, Yan P, Su Y, et al. 2021. A study of freeway crash risk prediction and interpretation based on risky driving behavior and traffic flow data. Accident Analysis & Prevention 160:106328 doi: 10.1016/j.aap.2021.106328

    CrossRef   Google Scholar

    [36] Yang Y, Wu Y. 2014. VE dimension induced by Bayesian networks over the boolean domain. Pattern Analysis and Applications 17:799−807 doi: 10.1007/s10044-013-0320-3

    CrossRef   Google Scholar

    [37] Yang Y, Wu Y. 2012. On the properties of concept classes induced by multivalued Bayesian networks. Information Sciences 184:155−165 doi: 10.1016/j.ins.2011.08.031

    CrossRef   Google Scholar

    [38] Yang Y, Wu Y. 2009. Inner product space and concept classes induced by Bayesian networks. Acta Applicandae Mathematicae 106:337−348 doi: 10.1007/s10440-008-9301-8

    CrossRef   Google Scholar

    [39] Yang Y, Wu Y. 2009. VC dimension and inner product space induced by Bayesian networks. International Journal of Approximate Reasoning 50:1036−1045 doi: 10.1016/j.ijar.2009.04.001

    CrossRef   Google Scholar

    [40] Hossain M, Muromachi Y. 2012. A Bayesian network based framework for real-time crash prediction on the basic freeway segments of urban expressways. Accident Analysis & Prevention 45:373−381 doi: 10.1016/j.aap.2011.08.004

    CrossRef   Google Scholar

    [41] Roy A, Hossain M, Muromachi Y. 2018. Enhancing the prediction performance of real-time crash prediction models: a cell transmission-dynamic bayesian network approach. Transportation Research Record: Journal of the Transportation Research Board 2672:58−68 doi: 10.1177/0361198118797802

    CrossRef   Google Scholar

    [42] Roy A, Kobayashi R, Hossain M, Muromachi Y. 2016. Real-time crash prediction model for urban expressway using dynamic Bayesian network. Journal of Japan Society of Civil Engineers Ser. D3 72:I_1331−I_1338 doi: 10.2208/jscejipm.72.I_1331

    CrossRef   Google Scholar

    [43] Wu M, Shan D, Wang Z, Sun X, Liu J, et al. 2019. A Bayesian network model for real-time crash prediction based on selected variables by random forest. Proc. 2019 5th International Conference on Transportation Information and Safety (ICTIS), Liverpool, UK, 2019. Liverpool, UK: IEEE. pp. 670–677 doi: 10.1109/ICTIS.2019.8883694
    [44] Fan S, Blanco-Davis E, Yang Z, Zhang J, Yan X. 2020. Incorporation of human factors into maritime accident analysis using a data-driven Bayesian network. Reliability Engineering & System Safety 203:107070 doi: 10.1016/j.ress.2020.107070

    CrossRef   Google Scholar

    [45] Murphy KP. 2001. The Bayes net toolbox for matlab. Computing Science and Statistics 33:1024−1034

    Google Scholar

    [46] Dempster AP, Laird NM, Rubin DB. 1977. Maximum likelihood from incomplete data via the EM algorithm. Journal of the Royal Statistical Society: Series B (Methodological) 39:1−22 doi: 10.1111/j.2517-6161.1977.tb01600.x

    CrossRef   Google Scholar

    [47] Yu R, Wang X, Abdel-Aty M. 2017. A hybrid latent class analysis modeling approach to analyze urban expressway crash risk. Accident Analysis & Prevention 101:37−43 doi: 10.1016/j.aap.2017.02.002

    CrossRef   Google Scholar

    [48] Dutta N, Fontaine MD. 2019. Improving freeway segment crash prediction models by including disaggregate speed data from different sources. Accident Analysis & Prevention 132:105253 doi: 10.1016/j.aap.2019.07.029

    CrossRef   Google Scholar

    [49] Hossain M, Muromachi Y. 2010. Development of a real-time crash prediction model for urban expressway. Journal of the Eastern Asia Society for Transportation Studies 8:2092−2107 doi: 10.11175/easts.8.2092

    CrossRef   Google Scholar

    [50] Stipancic J, Miranda-Moreno L, Saunier N, Labbe A. 2019. Network screening for large urban road networks: using GPS data and surrogate measures to model crash frequency and severity. Accident Analysis & Prevention 125:290−301 doi: 10.1016/j.aap.2019.02.016

    CrossRef   Google Scholar

  • Cite this article

    Guo M, Zhao X, Yao Y, Luan S, Yang H, et al. 2026. Relationship analysis between traffic status, risky driving behaviors, and crash probability in spatiotemporal. Digital Transportation and Safety 5(3): 255−272 doi: 10.48130/dts-0026-0021
    Guo M, Zhao X, Yao Y, Luan S, Yang H, et al. 2026. Relationship analysis between traffic status, risky driving behaviors, and crash probability in spatiotemporal. Digital Transportation and Safety 5(3): 255−272 doi: 10.48130/dts-0026-0021

Figures(9)  /  Tables(4)

Article Metrics

Article views(25) PDF downloads(10)

ARTICLE   Open Access    

Relationship analysis between traffic status, risky driving behaviors, and crash probability in spatiotemporal

Digital Transportation and Safety  5,  2026, 5(3): 255−272  |  Cite this article

Abstract: Studying the characteristics of traffic status and risky driving behaviors during crashes, identifying traffic crash risks, and then taking proactive measures to control them on the road are crucial to reducing traffic crash rates. In this study, we used in-vehicle navigation data and extracted traffic status and risky driving behavior data for the 30 min before a crash across the upstream, middle, and downstream segments. A dynamic Bayesian network model was used to determine the relationships between traffic status, risky driving behaviors, and crash probability in time and space. Temporally, a higher crash probability was observed in the two time slices, T3 and T4. Spatially, changes in the traffic status and risky driving behaviors in the downstream segment were associated with higher crash probabilities. This lays the foundation for dynamically identifying pre-accident risk states using vehicle navigation data, providing technical support for the transformation of traffic safety management from passive response to active prevention. In addition, the findings help identify high-risk road segments and critical risk periods from both temporal and spatial dimensions, providing a basis for precise interventions such as variable message signs, in-vehicle warnings, speed control, and traffic organization optimization.

    • Traffic crashes have become a leading cause of human injury and death in recent years. According to statistics from the World Health Organization (WHO), 1.35 million people worldwide die from traffic accidents each year[1]. Taking proactive measures to prevent traffic crashes is becoming increasingly important for government and road operation management departments. Currently, traffic crash analysis is shifting from passive post-crash response to active crash prevention, as crashes do not occur without warning but are influenced by traffic status, risky driving behaviors, and the road environment[2]. Traffic crashes are frequently associated with gradual changes in traffic status and risky driving behaviors[3,4]. If we can capture the spatiotemporal patterns of traffic status and risky driving behaviors during the crash, we can identify the risk of traffic crashes. Then, we can dynamically predict crash risk by monitoring changes in traffic status and driving behavior on the road. Therefore, it is necessary to study the relationship between traffic status, risky driving behaviors, and traffic crashes in time and space. This will help detect traffic status and risky driving behaviors associated with crash risk on the road in advance. This can support traffic management in predicting and providing warnings of crash risk based on spatiotemporal changes in traffic status and risky driving behaviors. When a high crash risk is identified, for example, we can take active preventive measures to prevent traffic crashes by providing timely traffic crash risk warnings to drivers at the appropriate time and roadway segment, thereby assisting drivers in avoiding traffic crashes and improving road safety.

      Scholars have recognized that active crash prevention lowers the traffic crash rate. Studies have established relationships between traffic status and traffic crashes at cross-sections using data collected by loop detectors at fixed locations both upstream and downstream of the crash. Traffic crashes were predicted by the changes in the variables of the traffic status at these points[5−9]. However, only data collected by loop detectors at fixed locations were used, and changes in traffic status on road segments without loop detectors were difficult to consider in the model. Moreover, only data collected in a single time period were used for traffic crash prediction. These constraints affect the prediction accuracy of the model and make it challenging to reflect the characteristics of the spatiotemporal changes in the traffic status during the crash. Thus, using methods with these limitations is not conducive to implementing targeted traffic crash prevention measures based on changes in the traffic crash risk on a roadway segment. However, the widespread use of in-vehicle navigation systems provides new opportunities to address these issues. Since in-vehicle navigation terminals (e.g., smartphones) include built-in GPS positioning chips, gyroscopes, and other sensors, traffic status data, such as volume, average speed, and congestion index, can be obtained in real time. Additionally, risky driving behaviors, such as sharp acceleration, sharp deceleration, and sharp turns, can also be recorded. There are 325.79 million monthly active users of in-vehicle navigation systems in China[10]. The data from in-vehicle navigation systems are provided in real time and are dynamic, have the characteristics of big data, and encompass the entire road, providing spatial information. Therefore, these data reflect the dynamic changes in traffic status and risky driving behaviors on the road. The use of these data for traffic crash analysis enables the investigation of spatiotemporal changes in the traffic status and risky driving behaviors associated with crashes.

      Therefore, an appropriate model is required to assess the changing patterns of traffic status and risky driving behaviors on different road segments and in different periods during the crash. Dynamic Bayesian networks provide a suitable option. They have been used to study probabilistic dependencies between variables and analyze the patterns of variables over time[11,12]. A model describing the spatiotemporal relationships between traffic status, risky driving behaviors, and crashes can be developed by incorporating the temporal and spatial properties of traffic status and risky driving behaviors into a dynamic Bayesian network model.

      This study uses real-time and dynamic data collected by in-vehicle navigation systems to extract temporal and spatial information on traffic status and risky driving behaviors to analyze the relationship between these variables and crash probability. Four temporal slices and three spatial (road) segments are used to capture the spatiotemporal characteristics of the variables, which are considered in a dynamic Bayesian network. The differences in the traffic status and risky driving behaviors in different time slices and road segments during the crash are compared. The posterior crash probabilities resulting from changes in traffic status and risky driving behaviors in different temporal and spatial segments are calculated. The results provide information on spatial and temporal factors leading to traffic crashes and can support the development of active crash prevention measures.

    • The prevention of traffic crashes is a hot issue in traffic safety research. Studies have analyzed the influencing factors and traffic crash prediction methods in detail from the macro-level to the micro-level, using static and dynamic data. At the macro-level, traffic crashes were primarily predicted using static data such as the annual average daily traffic volume and road configurations as independent variables[13−16]. These studies focused on road defects and identified the primary causes of traffic crashes, e.g., the grade[15], horizontal and vertical curves[13], tunnels[14], and problem areas of roads. The results of the post-crash analysis of these studies were used to improve road traffic safety through engineering measures; however, this method requires long-term crash data.

    • Considering the limitations of accident prediction at a macro-level, researchers started to use real-time and dynamic data to establish the relationship between the traffic status and traffic crashes at the micro-level to achieve real-time, dynamic prediction and prevention of traffic crashes. The collection of real-time and dynamic data on the traffic status is a crucial aspect of micro-level traffic crash prediction. Many studies used loop detectors, microwave detectors, and radar for obtaining traffic status data. These devices provide data on the traffic volume, speed, and occupancy at the detectors' location[6,7,9]. Subsequently, data are extracted from detectors upstream and downstream of the crash. The traffic volume[6,17], average speed[5], congestion level/occupancy[9,18], speed variation coefficient[19,20], and the speed difference between the upstream and downstream detectors[20,21], have been used as independent variables for traffic crash prediction.

    • Due to real-time and prediction accuracy requirements, 5 min was typically used as a time slice for the aggregation of traffic status data[22−25]. However, using data for a 5-min period before a crash for real-time traffic crash prediction is not conducive to developing management measures because the traffic management department may be incapable of reacting within 5 min. Therefore, many researchers used 10[26], 15[4,27], or 30 min[28−30] as a time slice for the aggregation of traffic status data to improve the model's accuracy and usability. Many studies used machine learning models to predict the probability of crashes. Models with relatively high prediction accuracy have been utilized, such as support vector machines[31], neural networks[32], convolutional neural networks[33], and eXtreme Gradient Boosting[22,34]. These studies typically used data for a single period and focused on improving the accuracy of model predictions. As a result, changes in traffic status and risky driving behaviors during the crash were often not considered. However, recent studies attempted to improve the prediction performance of traffic crashes using traffic data with different temporal resolutions using the long short-term memory (LSTM) model or dynamic Bayesian networks to deal with time-series data. Additionally, changes in the traffic status during the crash were captured by analyzing the dependencies in the time-series data[3]. However, the loop detectors only collected data at two fixed points upstream and downstream of the crash in these studies. The limited coverage of the detectors may have affected the accuracy of traffic crash prediction[3] and did not provide in-depth information on changes in traffic status during the crash.

    • With the widespread use of in-vehicle navigation systems, it has become possible to collect data on the average speed, congestion index, and risky driving behaviors. The traffic status data derived from in-vehicle navigation devices are real-time, dynamic, and cover the entire road. Therefore, these data reflect changes in traffic status and risky driving behaviors on the road, providing precise information for crash prevention in time and space. For example, variable message boards or in-vehicle devices can remind drivers to drive cautiously at the correct time and at an appropriate location[23]. In addition, the freeway management department can provide traffic control measures to reduce the possibility of traffic crashes[4]. Therefore, analyzing the relationship between traffic status, risky driving behaviors, and crash probability improves the capability of crash risk identification, timeliness, and relevance of traffic crash prevention measures and the transition of the management style from passive post-crash disposal to active crash prevention.

      It can be seen that existing studies have evolved from macro-level static analyses to accident prediction based on micro-level dynamic traffic data. However, macro-level studies mainly rely on historical crash records and road geometric characteristics to identify long-term influencing factors, making them inadequate for real-time identification and proactive intervention before a crash occurs. Although micro-level studies have introduced detector data and machine learning methods, most of them still depend on fixed-point detectors to obtain localized traffic information. As a result, their data coverage is limited, and they are unable to capture the dynamic changes in traffic conditions at different locations during the crash process. More importantly, existing research has primarily focused on the effects of traffic flow parameters on crash risk, while giving relatively limited attention to risky driving behaviors, and lacks a systematic analysis of the joint mechanisms through which traffic conditions and risky driving behaviors affect crash occurrence. Therefore, an important gap in current research on proactive traffic crash prevention remains: how to use in-vehicle navigation system data, which offer full-road coverage, real-time responsiveness, and dynamic monitoring, to comprehensively extract features of traffic conditions and risky driving behaviors within a certain time window before a crash, and to reveal their relationships with crash probability from both temporal and spatial perspectives. The purpose of this paper is to study the feasibility of active traffic crash prevention supported by the changing characteristics of traffic state and risky driving behaviors. The main work is to determine the relationship between traffic status, risky driving behaviors, and crash probability in time and space using data collected by in-vehicle navigation systems.

    • The data for this study were obtained from the G15 two-way freeway in China. The chosen section has a length of 40 km in both directions. The data included crashes, risky driving behaviors (including sharp acceleration and sharp deceleration, sharp turn left, sharp turn right, sharp merge into left lane, and sharp merge into right lane), and traffic status (including volume, average speed, and the congestion index). The data were collected from May 1 to May 31, 2019, and October 1 to October 31, 2019. A total of 227 traffic crashes occurred during the periods. The data were described in detail in Table 1. As these data come from different departments, their time and space labels are also different in the process of data collection.

      Table 1.  Details of the data.

      Data Description Time label Spatial label Data source
      Crash Collisions between vehicles on the road. 1 min 1 km (road segment) Freeway management department
      Volume The total number of vehicles passing a road point per unit time. 5 min 1 km (road point)
      Average speed The average speed of vehicles traveling a road segment per unit time. 10 min 1 km (road segment) In-vehicle navigation systems of mobile phones of vehicles driving on the freeway
      Congestion index Free-flow speed divided by the average road speed.1
      Sharp acceleration If the linear acceleration was greater than a certain threshold, a sharp acceleration or deceleration would be recognized and recorded when the mobile phone was in a fixed position. 1 s Latitude and longitude
      Sharp deceleration
      Sharp left turn When a mobile phone was placed in a certain (fixed) position, the centripetal force of the original historical turn was measured. If the detection angle was greater than a certain threshold, it would be identified as a sharp merge into another lane or a sharp turn.
      Sharp right turn
      Sharp merge into the left lane
      Sharp merge into the right lane
      1 The congestion index variable had four categories: the free-flow state (CI∈[0,1.5), slow-flow state (CI∈[1.5,2]), congestion state (CI∈[2,4]), and severe congestion state (CI∈[4, + ∞]).

      In order to label the data consistently in space and time, these data were matched in the temporal and spatial dimensions using a space-time grid approach to create a spatiotemporal database[34]. The spatial granularity was 1 km, and the temporal granularity was 10 min, i.e., 1 km was the spatial unit, and the road was divided into 40 road segments; the traffic status and risky driving behaviors data in each road segment were updated every 10 min. In addition, in order to test the contribution of these traffic status and risky driving behaviors variables to traffic crashes, we used the random forest algorithm to calculate the importance of these variables. We find that these four variables (including sharp left turn, sharp right turn, sharp merge into the left lane, and sharp merge into the right lane) contribute little to traffic crashes. Their cumulative importance value is about 0.05[35]. This may be because there are so few instances of this type of risky driving behavior on the freeway. Therefore, in order to simplify the structure of the model, these four variables are not considered in this study. The descriptive statistics of the finall dataset is shown in Table 2.

      Table 2.  Descriptive statistics of the dataset.

      Variables Mean Std Min. Max. Unit
      Volume 90.9512 55.5931 1.0000 384.0000 veh/10 min
      Average speed 77.4408 9.6283 2.3040 112.6440 km/h
      Congestion index 1.0809 0.4419 0.7600 33.8500 —
      Sharp acceleration 0.0022 0.0114 0.0000 2.0000 times/10 min
      Sharp deceleration 0.0020 0.0010 0.0000 1.0000 times/10 min
    • We divided the road sections in time and space to study the spatiotemporal changes of traffic status and risky driving behaviors during the crash. The crash occurred in the middle segment. The segment before the crash (middle segment) was called the upstream segment, and the segment after the crash was referred to as the downstream segment. Temporally, because the minimum collection time of congestion index data is 10 min, restricted by this time granularity, we used 10-min time intervals. The crash occurred in the last time slice (T4). The 10 min preceding the crash was called the third time slice (T3), the 20 min before the crash was called the second time slice (T2), and the 30 min before the crash was called the first time slice (T1).

      It should be noted that to identify the changes in the relationship between traffic status, risky driving behaviors, and the crash from before to after a crash, we included the time slice when the crash occurred (T4) in the model; e.g., data on traffic status and risky driving behaviors were extracted from T1 to T4 and in the upstream, middle, and downstream segments. This data extraction method facilitates the analysis of the spatiotemporal relationship between traffic status, risky driving behaviors, and crash probability.

      We created heat maps of the spatiotemporal distribution of each variable extracted from the database to visualize the spatiotemporal patterns of traffic status and risky driving behaviors, as shown in Fig. 1. The average values of the traffic volume, average speed, and congestion index, and the sums of the sharp acceleration and sharp deceleration were used for each time slice and spatial segment. All variables exhibited distinct spatiotemporal characteristics, which are described as follows.

      Figure 1. 

      Spatiotemporal distribution of traffic status and risky driving behavior characteristics during the crashes. (a) Traffic volume. (b) Average speed. (c) Congestion index. (d) Sharp acceleration. (e) Sharp deceleration.

      As shown in Fig. 1a, the traffic volume in the upstream, middle, and downstream segments increased over time, and the maximum volume occurred in the upstream segment at T4. Spatially, the traffic volume decreased from the upstream to the middle and downstream segments. The average speed (Fig. 1b) remained high in the upstream segment from T1 to T2 and gradually slowed down over time, reaching its lowest point in the middle segment at T4. The congestion index (Fig. 1c) was relatively low in the three segments from T1 to T3 and gradually increased from T3 to T4. The index increased the fastest in the middle segment and reached the maximum at the time of the crash. As shown in Fig. 1d, the values of sharp acceleration increased over time, and the rate of increase was faster in the T3 to T4 time slices. This trend was most pronounced in the upstream segment, followed by the middle segment. Relatively high sharp deceleration occurred in the middle segment and at T4, while relatively low deceleration was observed in the upstream and middle segments from T1 to T2 (Fig. 1e).

    • In addition to crash data, it is necessary to collect non-crash data as a control to analyze the relationship between traffic status, risky driving behaviors, and crash probability. Following a previous study, we employed the matched case-control design[28]. Several studies chose a ratio of 1:4 for matching crash cases to non-crash cases[9,17,1], i.e., for each crash case in the database, four non-crash cases were selected for matching. The non-crash sample represents the normal state of traffic status and risky driving behaviors. Therefore, it is necessary to avoid the influence caused by the crash when selecting non-crash samples. In accordance with this request, we removed traffic status and risky driving behavior data from the database for 6 h before and after the crash time in the three segments. In the remaining data, we randomly selected four non-crash samples for each crash sample. Each non-crash sample has the same form as the crash sample in time and space, i.e., data on traffic status and risky driving behaviors were extracted from T1 to T4 and in the upstream, middle, and downstream segments. Finally, the resulting database contained 227 crash cases and 908 non-crash cases. These data were expanded in four time slices and three space segments, and a total of 13,620 samples were obtained. These samples are used to build a model and analyze the relationship between traffic status, risky driving behaviors, and crashes in spatiotemporal terms.

    • A Bayesian network, also known as a belief network, is a widely used graphical model that describes the dependency relationships between variables. It is a theoretical model widely used for uncertain knowledge representation and inference[36−39]. The predictive power of Bayesian networks has received substantial attention in traffic safety research. Bayesian networks are frequently used to improve the accuracy of traffic crash prediction or make inferences[4,40−43]. They have a convenient network structure to represent the dependencies between variables to describe uncertain knowledge and clarify the logic. In this study, the inference capability of Bayesian networks enables the description of the spatiotemporal relationship between traffic status, risky driving behaviors, and traffic crashes.

    • Bayesian networks can be divided into static and dynamic networks. Static Bayesian networks (SBNs) only reflect the probabilistic dependencies between a series of variables, but the change in the variables over time is not considered. In contrast, a dynamic Bayesian network (DBN) considers the temporal dimension and the change in probabilistic dependencies between variables over time. Therefore, a dynamic Bayesian network considers both external influences and interconnections within the system. A dynamic Bayesian network is an extension of the static Bayesian network and includes time-varying processes.

      The structure of a Bayesian network is shown in Fig. 2. The arcs represent conditional dependencies between variables. There are two types of nodes: hidden state nodes and observed evidence nodes. Hidden state nodes denote a discrete variable represented by a set of $ {N}_{h} $ random variables, $ H_{t}^{i} $, $ i\in \{1,\cdots ,{N}_{h}\} $, which can be discrete or continuous. In this study, this is a binary variable, such as 'crash-prone' or 'not crash-prone'. Hidden state nodes can be used to represent the possibility of each state. Observed evidence nodes denote a random variable and can be used for the inference of dependent variables. These nodes are represented by $ {N}_{0} $ random variables, $ E_{t}^{j} $,$ j\in \{1,\cdots ,{N}_{0}\} $. In this study, traffic status and risky driving behaviors are observed evidence nodes.

      Figure 2. 

      The structure of Bayesian networks.

      Figure 2a shows a static Bayesian network, which can only represent the relationship between traffic status, risky driving behaviors, and crashes in one time slice. A dynamic Bayesian network (only two time slices are shown in the figure) is shown in Fig. 2b. The solid arcs in Fig. 2b represent the relationship between variables within a given time slice. The relationship between variables in two consecutive time slices is represented by the dotted arcs, representing changes in traffic status and risky driving behaviors in a 'crash-prone' state over time. The structure of the dynamic Bayesian network was determined based on the relationship between the crashes and the variables (e.g., traffic status) in previous studies[4,44]. In this study, risky driving behaviors were added as a new variable to the dynamic Bayesian network.

      In a state-space model, $ \text{P(}{H}_{t}|{H}_{t-1}) $ is a transition model, $ \text{P(}{E}_{t}|{H}_{t}) $ is an observation model, and $ \text{P(}{H}_{1}) $ is an initial state distribution. However, the DBN is a more general model and is defined as a pair $ ({B}_{0},{B}_{\rightarrow }) $ where $ {B}_{0} $ defines the prior $ \text{P(}{Z}_{1}) $, and $ {B}_{\rightarrow } $ is a two-slice temporal Bayes net (2TBN) that defines the transition and observation models as products of the conditional probability distributions (CPDs) in the 2TBN:

      $ P\left({X}_{t}|{X}_{t-1}\right)=\prod\limits_{i=1}^{N}P(X_{t}^{i}|Pa(X_{t}^{i})) $ (1)

      where, $ X_{t}^{i} $ is the ith node in time slice t (which can be a hidden state node or an observed evidence node, $ N={N}_{h}+{N}_{0} $), and $ Pa(X_{t}^{i}) $ are the parents of $ X_{t}^{i} $, which may be either time slice $ t $ or $ t-1 $. The nodes do not have associated parameters in the first slice of a 2TBN but have an associated CPD in the second slice. For a DBN with T slices, the 2TBN is unrolled until the network has T slices. The CPDs are multiplied to obtain the joint distribution:

      $ P\left(X_{1\colon T}^{1\colon N}\right)=\prod\limits_{i=1}^{N}{P}_{{{B}_{0}}}(X_{1}^{i}|Pa(X_{t}^{i}))\times \prod\limits_{t=2}^{T}\prod\limits_{i=1}^{N}{P}_{{{B}_{\rightarrow }}}(X_{t}^{i}|Pa(X_{t}^{i})) $ (2)
    • There may be missing values in the traffic data. Researchers often use the gradient descent algorithm or expectation-maximization (EM) algorithm to solve this problem and compute the maximum likelihood estimates[45]. The EM algorithm learns the dependence between the pieces of empirical evidence[46]. It is solved by optimizing the set of unknown parameters θ in the model, which maximizes the log-likelihood of the data. When given the following parameters, the log of the marginal probability of the observation value is calculated as follows:

      $ \ell\left(\theta ;E\right)=\log p(E|\theta )=\log \sum\limits_{H}p(H,E|\theta ) $ (3)

      An initialization distribution of $ \theta $ was defined by incorporating the distribution $ Q(H) $. We use Jensen's inequality:

      $ \ell\left(\theta ;E\right)=\log \sum\limits_{H}Q(H)\dfrac{p(H,E|\theta )}{Q(H)}\geq \sum\limits_{H}Q\left(H\right)\log \dfrac{p(H,E|\theta )}{Q(H)} $ (4)

      Next, the joint probability distribution, $ Q\left(H\right)=P(H|E;\theta ) $, of the hidden state variables was computed in step E. In the DBN, the joint probability distribution of the hidden state variables was estimated given the observed variables and the current values of the parameters. Moreover, the parameters were optimized by step M based on the estimation of the joint probability distribution of the state variables. The calculation is as follows:

      $ {\theta }^{new}=\arg \underset{\theta }{\max } \sum\limits_{H}Q\left(H\right)\log \dfrac{p(H,E|\theta )}{Q(H)} $ (5)

      Then, new optimal parameters were obtained to replace the original parameters. Finally, the optimal estimated parameters were obtained by repeating steps E and M of the EM algorithm, thus stopping the iteration.

      After the DBN parameters were estimated by computing $ \text{P(}{H}_{1\colon T}|{E}_{1\colon T};\theta ) $, the hidden states $ {H}_{1\colon T} $ were inferred based on the observed evidence $ {E}_{1\colon T} $. Consequently, the crash-prone state was identified from the observed evidence.

    • Before performing Bayesian network calculations, the continuous variables must be discretized to determine the value range of the variables. In order to ensure that crash and non-crash data were represented, we divide the ranges of the traffic volume, average speed, sharp acceleration, and sharp deceleration into four categories[40]. In the database, the frequency distribution for each value range of variables is shown in Table 3, including both crash and non-crash cases. It can be found that dividing the value range of each variable into four categories could represent the difference between crash and non-crash conditions.

      Table 3.  The frequency distribution for each value range of variables.

      Abbr. Variables Categories Unit Frequency in crash (%) Frequency in non-crash (%)
      Vol. Volume Less 200 vehicles/
      km·10 min
      9.43% 64.77%
      200 to 400 74.20 32.51
      400 to 600 16.24 2.69
      More 600 0.12 0.03
      A.S. Average speed Less 60 km/h 51.36 7.21
      60 to 80 37.85 69.08
      80 to 100 10.79 23.69
      More 100 0.00 0.02
      C.I. Congestion index 0 to 1.5 na 67.96 98.06
      1.5 to 2 11.39 0.89
      2 to 4 13.57 0.87
      More 4 7.08 0.18
      S.A. Sharp acceleration Less 1 times/
      (km·10 min)
      72.03 92.74
      1 to 3 21.26 6.94
      3 to 5 4.22 0.28
      More 5 2.50 0.04
      S.D. Sharp deceleration less 1 times/
      (km·10 min)
      63.11 92.25
      1 to 3 24.45 7.26
      3 to 5 6.86 0.44
      More 5 5.58 0.05

      In order to evaluate the fitting performance of the dynamic Bayesian network, the data were split into the training set (70%) and the test set (30%). The dynamic Bayesian network learned the CPDs in the network via the training set. The test set was used to evaluate the model's performance. The evaluation index and calculation method of the model are in Eqs (6)−(9).

      $ Overall\;accuracy=\dfrac{{True}_{crash}+{True}_{noncrash}}{{True}_{crash}+{True}_{noncrash}+{False}_{crash}+{False}_{noncrash}} $ (6)
      $ Sensitivity=\dfrac{{True}_{crash}}{{True}_{crash}+{False}_{crash}} $ (7)
      $ False\;alarmrate=\dfrac{{False}_{noncrash}}{{False}_{noncrash}+{True}_{noncrash}} $ (8)
      $ Missing\;reportrate=\dfrac{{False}_{crash}}{{True}_{crash}+{False}_{crash}} $ (9)

      where, $ {True}_{crash} $ represents the case predicted by the model to be a crash and it is also a crash in reality. $ {True}_{noncrash} $ represents the case predicted by the model to be non-crash, and it is also non-crash in reality. $ {False}_{crash} $ represents the case predicted by the model to be a crash, but it is actually non-crash in reality. $ {False}_{noncrash} $ represents a case that is predicted by the model to be non-crash, but it is actually a crash in reality. Overall accuracy indicates the overall accuracy of the classification of crashes and non-crashes by the model. Sensitivity represents the model's ability to identify crashes. For overall accuracy and sensitivity, the greater the value, the better. False alarm rate indicates that the model identifies non-crash cases as crash cases. Missing report rate indicates that the model identifies crash cases as non-crash cases. The smaller the value of false alarm rate and missing report rate, the better.

      Therefore, we evaluated the dynamic Bayesian network using the crash classification accuracy (overall accuracy), sensitivity, false alarm rate, and missing report rate. The results are shown in Table 4. Overall, crash classification accuracy and sensitivity of the model are relatively high. The false alarm rate and missing report rate of the model are relatively low. The results indicate the excellent fitting performance of the proposed dynamic Bayesian network model.

      Table 4.  The results of the dynamic Bayesian network evaluation.

      Evaluation
      index
      Overall
      accuracy
      Sensitivity False alarm
      rate
      Missing report
      rate
      Results 89.74% 88.24% 9.89% 11.76%
    • The probability distributions of the traffic status and risky driving behaviors in non-crash and crash conditions were analyzed by calculating the prior probabilities of the dynamic Bayesian networks. Figure 3a−c shows the probability distributions of the traffic status and risky driving behaviors in the upstream, middle, and downstream segments in the non-crash conditions. Figure 4a−c shows the corresponding results in the crash conditions. In the figure, each square containing a bar chart represents a variable (node). Taking Middle_Vol. [T1] in Fig. 3a as an example, its internal structure is as follows: (1) Node name (top title bar): Middle_Vol. [T1]. This indicates the name of the variable (e.g., traffic volume on the middle section at time T1). (2) States/intervals (left text: less than 200, 200 to 400, 400 to 600, more than 600). Bayesian networks typically discretize continuous data (such as specific traffic volume values) into several discrete intervals (states). (3) Probability values (numbers in the middle: 65.7, 31.5, 2.80, 0+). These represent the probability (percentage) that the variable falls into the corresponding interval under the current network state. For example, the probability that the volume is less than 200 is 65.7%. '0+' indicates an extremely small probability, close to zero but not absolutely zero. (4) Probability bar chart (black bars on the right). This is a visual representation of the probability values. The longer the black bar, the greater the probability of that interval. (5) Expected value and standard deviation (numbers at the bottom: 174 ± 120). If the variable is derived from discretizing a continuous variable, the current mean and standard deviation are shown at the bottom. (6) Core target variables (central nodes). The two nodes in the middle are Crash [T1] and Crash [T2] (crash occurrence at times T1 and T2). Their states are divided into 0 (no crash) and 1 (crash).

      Figure 3. 

      Spatiotemporal probability distribution of traffic status and risky driving behaviors in non-crash conditions. (a) Distribution of traffic status and risky driving behaviors from T1 to T2 in the upstream, middle, and downstream segments in non-crash conditions. (b) Distribution of traffic status and risky driving behaviors from T2 to T3 in the upstream, middle, and downstream segments in non-crash conditions. (c) Distribution of traffic status and risky driving behaviors states from T3 to T4 in the upstream, middle, and downstream segments in non-crash conditions.

      Figure 4. 

      Spatiotemporal probability distribution of traffic status and risky driving behaviors in crash conditions. (a) Distribution of traffic status and risky driving behaviors states from T1 to T2 in the upstream, middle, and downstream segments in crash conditions. (b) Distribution of traffic status and risky driving behaviors states from T2 to T3 in the upstream, middle, and downstream segments in crash conditions. (c) Distribution of traffic status and risky driving behaviors states from T3 to T4 in the upstream, middle, and downstream segments in crash conditions.

      Specifically, as shown in Fig. 3, the probability distributions of traffic status and risky driving behaviors were similar in all road segments and time slices in the non-crash conditions, reflecting the normal traffic status. Moreover, the probability distributions were relatively stable. The probability of a traffic volume of fewer than 200 vehicles/(km·10 min) was the highest (60% to 70%) in the four time slices, and there were no crashes in any segments. The second-highest probability (more than 30%) was 200 to 400 vehicles/(km·10 min). The lowest probability was a traffic volume of more than 400 vehicles/(km·10 min). The probability of an average speed of 60.00 to 80.00 km/h was the highest, followed by 80.00 to 100.00 km/h and less than 60.00 km/h, indicating that vehicles were traveling at speeds higher than the minimum speed limit (60.00 km/h) under normal conditions. The probability of the congestion index of 0 to 1.5 was the highest (more than 97%), indicating that the roads were not congested under normal conditions. The highest probability of sharp acceleration was less than 1 (more than 90%), indicating that most of the drivers did not accelerate rapidly under normal conditions. The results were similar for sharp deceleration.

      Compared with spatiotemporal probability distribution of traffic status and risky driving behaviors in non-crash conditions (Fig. 3), the probability distribution of traffic status and risky driving behaviors in crash conditions was significantly different in all segments and time slices (Fig. 4). In the crash conditions, the probability distribution of the traffic volume of 200 to 400 vehicles/(km·10 min) was higher than 70% in the four time slices and three road segments, followed by 400 to 600 vehicles/(km·10 min). The probability of a traffic volume of less than 200 vehicles/(km·10 min) was significantly lower in the crash than in the non-crash conditions. The probability of an average speed of less than 60.00 km/h was significantly higher in crash conditions. The downstream segment had the highest probability at that speed, followed by the average speed of 60.00 to 80.00 km/h. Moreover, the probability of an average speed of 80.00 to 100.00 km/h was significantly lower in crash conditions. In addition, although the probability of a congestion index of 0 to 1.5 was the highest, the probability of a congestion index of more than 1.5 was higher in the crash cases. As for the probability distributions of sharp acceleration and deceleration, they are similar in the crash conditions. For the risky driving behaviors in the upstream and middle segments, the probability of sharp acceleration of less than 1 was around 70% in the crash cases, and that of sharp deceleration of less than 1 was less than 70%. However, in the down segment, the probability of risky driving behaviors of less than 1 was more than 70% in most crash cases.

    • In order to further analyze the influence of variations of traffic status and risky driving behaviors on the probability of traffic crashes on different road segments and at different time slices, two graphs were drawn for each variable to discuss their change processes in time and space based on dynamic Bayesian networks' posterior probability calculation capability.

      Figure 5 shows the effect of a change in traffic volume on the probability of traffic crashes. Temporally (Fig. 5a), the largest effect of the traffic volume on the crash probability occurred from T2 to T3. In general, the crash probability was the highest for a traffic volume of 400 to 600 (vehicle/10 min) in all time slices. It has been demonstrated that the main reason for the increased probability of traffic crashes was a high traffic volume[24]. Spatially (Fig. 5b), the largest effect of the volume and probability on the crash probability occurred in the upstream segment. The crash probability was the highest for a traffic volume of 400 and 600 (vehicle/10 min). A study showed that an increase in the traffic volume in the upstream was the cause of traffic crashes[17]. Furthermore, since there were few cases with a traffic volume of more than 600 (vehicle/10 min), the results of this model may not be reliable. However, the higher the traffic volume, the higher the risk was[47]. This phenomenon should be considered by freeway management departments.

      Figure 5. 

      Effect of a change in traffic volume on crash probability. (a) Temporal variation of traffic volume. (b) Spatial variation of traffic volume.

      Figure 6 shows the effect of a change in average speed on the crash probability. Temporally (Fig. 6a), the largest effect of the average speed on crash probability was observed in T3 to T4. The crash probability was the highest at an average speed of 60.00 km/h or below in T3 to T4. It was approximately 70% near T4. Studies have shown that lower freeway speeds are often associated with traffic crashes[1,48,49]. Therefore, if the average speed of vehicles on the road is less than 60.00 km/h for an extended period of time, or if the crash probability is high, the management should implement measures to prevent traffic crashes. In addition, due to a low number of cases with an average speed of more than 100.00 km/h, the Bayesian network failed to learn the relationship between the average speed and traffic crashes. Spatially (Fig. 6b), the largest effects of the average speed on the crash occurred in the middle and downstream segments. This result is consistent with previous studies, which indicated a high crash probability at a relatively low average speed in the downstream[5,18]. Moreover, when the average speed was less than 60 km/h, the crash probability remained above 60%, probably because the minimum speed on freeways is 60.00 km/h. If the vehicle's speed is less than 60.00 km/h, a traffic crash is more likely[50]. In this study, the crash probability was below 20% in intervals where the average speed was more than 60.00 km/h.

      Figure 6. 

      Effect of the change in average speed on crash probability. (a) Temporal variation of average speed. (b) Spatial variation of average speed.

      Figure 7 shows the effect of a change in the congestion index on the probability of traffic crashes. Temporally (Fig. 7a), the effect of the congestion index on crash probability fluctuated in the upstream and downstream segments from T1 to T2. Moreover, the crash probability remained above 60% when traffic congestion occurred during these two time slices. However, the impact of the congestion levels on the crash probability showed an increasing trend from T3 to T4, and the crash probability remained above 70% when traffic congestion occurred. The longer the duration of road congestion, the higher the probability of traffic crashes is[48].

      Figure 7. 

      Effect of a change in the congestion index on crash probability. (a) Temporal variation of congestion index. (b) Spatial variation of congestion index.

      Therefore, traffic management departments should provide traffic diversion when congestion occurs. Furthermore, the free flow status (CI∈[0,1.5]) had a negligible impact on crash probability in different time slices. Spatially (Fig. 8b), the effect of the congestion index on crash probability differed for different road segments. Slowly moving traffic (CI∈[1.5,2]) in the upstream segment, severe congestion (CI∈[2,4]) in the middle segment, and severe congestion (CI∈[4,+∞]) in the downstream segment should be avoided. It is suggested that congestion and safety warnings should be provided to drivers. Traffic congestion was more likely to result in traffic crashes in the downstream segment[17].

      Figure 8. 

      Effect of a change in sharp acceleration on crash probability. (a) Temporal variation of sharp acceleration. (b) Spatial variation of sharp acceleration.

      Figure 8 shows the effect of a change in sharp acceleration on crash probability. Temporally (Fig. 8a), the effect of sharp acceleration on crash probability was similar in all segments. Specifically, the larger the change in sharp acceleration, the higher the crash probability was. In particular, as time got closer, the probability of traffic crashes exceeded 75% when the frequency of sharp acceleration was more than three times/10 min. In addition, the crash probability was very high (more than 80%) when the frequency of sharp acceleration was more than five times/10 min in all four time slices. In contrast, when the number of sharp accelerations was 0, the crash probability was very low (less than 20%). Spatially (Fig. 8b), the higher the frequency of sharp accelerations, the higher the crash probability was in the three road segments. The number of sharp accelerations fluctuated in the upstream segment. However, the crash probability showed an increasing trend with an increase in the number of sharp accelerations on the middle and downstream segments. In particular, the crash probability remained above 70% when the frequency of sharp acceleration was more than three times/10 min.

      Figure 9 shows the effect of a change in sharp deceleration on the crash probability. Temporally (Fig. 9a), the relationship between sharp deceleration and crash probability was generally consistent in the spatial distribution in the four periods. The higher the frequency of sharp decelerations, the higher the crash probability was. Moreover, the relationship between the number of sharp accelerations and the crash probability remained consistent in each period. The crash probability had a range of 60% to 100% when the frequency of sharp accelerations/decelerations was more than three times/10 min. In addition, the crash probability was more than 85% when the frequency of sharp deceleration was more than five times/10 min in the four time slices, indicating that the impact of high deceleration was larger than that of sharp acceleration. Spatially (Fig. 9b), the more frequently sharp deceleration occurred in the three road segments, the higher the crash probability was. Particularly, the crash probability remained above 70% when the frequency of sharp deceleration was more than three times/10 min in the middle and downstream segments. Furthermore, sharp deceleration in the downstream segment had the most significant impact on crash probability (around 90%). Sharp acceleration and deceleration are risky driving behaviors that will have a high probability of traffic crashes[26]. Therefore, the focus should be on monitoring risky driving behaviors in the middle and downstream segments, such as sharp acceleration and deceleration. These behaviors might result in collisions in the middle segment and affect the stability of the traffic status in the downstream segment, influencing traffic status in the preceding segments and potentially increasing crash probability.

      Figure 9. 

      Effect of a change in sharp deceleration on crash probability. (a) Temporal variation of sharp deceleration. (b) Spatial variation of sharp deceleration.

    • To explore the possibility of active crash prevention, this study collected real-time data on risky driving behaviors and traffic status obtained from in-vehicle navigation systems to determine the spatiotemporal probability distribution of traffic crashes. The road was divided into the upstream, middle, and downstream segments, and four time slices (T1 to T4) with a granularity of 10 min were considered. A dynamic Bayesian network was used to analyze the data. The model achieved a crash sensitivity of 88.24% with a false alarm rate of 9.89%, and a crash classification accuracy of 89.74%, indicating that the model fit was relatively high.

      The relationship between traffic status, risky driving behaviors, and crash probability was determined by calculating the prior and posterior probabilities using the dynamic Bayesian network. The following results were obtained. First, the crash probability remained high over time when the traffic volume was between 400 and 600 (vehicle/10 min) in the upstream segment. Second, the crash probability increased from T3 to T4 to more than 70% when the average speed was below 60 km/h, notably on the downstream segment. Third, the effect of the congestion level on the crash probability increased from T3 to T4, and the crash probability remained above 70%. Fourth, risky driving behaviors (including sharp acceleration and deceleration) also significantly affected the probability of traffic crashes in the three road segments and four time slices. The more frequently the risky driving behaviors occurred, the higher the crash probability was. Notably, the crash probability in the downstream segment was more than 80% when the frequency of sharp acceleration or deceleration exceeded three times/10 min from T3 to T4. These findings on the spatiotemporal relationship between traffic status, risky driving behaviors, and crash probability can help develop more targeted traffic crash prevention measures.

      The contribution of this paper is to demonstrate the relationship between traffic status, risky driving behaviors, and crash probability in time and space. In particular, higher frequencies of risky driving behaviors are often accompanied by higher crash probabilities. Therefore, we can infer the probability of traffic crashes from the characteristics of traffic status and risky driving behaviors in time and space on the road. This indirectly validates the application potential of in-vehicle navigation data in traffic crash risk identification, offering a practical foundation for building a more extensive and real-time intelligent traffic safety early warning system. It helps reduce accident rates, mitigate casualties, traffic congestion, and related economic losses, yielding significant safety and social benefits.

      However, this study has some limitations. The time slice interval of 10 min was chosen based on the data. In a future study, we will use a finer time granularity (e.g., 5 min) to obtain more details on the spatiotemporal relationships between traffic status, risky driving behaviors, and crash probability. Furthermore, the crash data from both directions of the G15 expressway serve only as a case application of this methodological framework to verify its feasibility and effectiveness. Future research should incorporate more expressway cases to further validate the generalizability of the method.

      • The authors confirm their contributions to the paper as follows: study conception and design: Guo G, Zhao X, Yao Y; data collection: Guo M, Su Y, Luan S; analysis and interpretation of results: Zhao X, Yang H, Guo M; draft manuscript preparation: Guo M, Zhao X. All authors reviewed the results and approved the final version of the manuscript.

      • The data that support the findings of this study are not publicly available due to third-party data-use restrictions and privacy and confidentiality considerations. The data may be available from the corresponding author upon reasonable request, subject to approval by the relevant data providers.

      • We would like to declare that no conflict of interest exists in the submission of this manuscript, and the manuscript is approved by all authors for publication. The work described was original research, and none of the material in this paper has been published or is under consideration for publication elsewhere.

      • Copyright: © 2026 by the author(s). Published by Maximum Academic Press, Fayetteville, GA. This article is an open access article distributed under Creative Commons Attribution License (CC BY 4.0), visit https://creativecommons.org/licenses/by/4.0/.
    Figure (9)  Table (4) References (50)
  • About this article
    Cite this article
    Guo M, Zhao X, Yao Y, Luan S, Yang H, et al. 2026. Relationship analysis between traffic status, risky driving behaviors, and crash probability in spatiotemporal. Digital Transportation and Safety 5(3): 255−272 doi: 10.48130/dts-0026-0021
    Guo M, Zhao X, Yao Y, Luan S, Yang H, et al. 2026. Relationship analysis between traffic status, risky driving behaviors, and crash probability in spatiotemporal. Digital Transportation and Safety 5(3): 255−272 doi: 10.48130/dts-0026-0021

Catalog

    /

    DownLoad:  Full-Size Img  PowerPoint
    Return
    Return