Search
2026 Volume 5
Article Contents
ARTICLE   Open Access    

Vehicle trajectory prediction using a kinematic feature enhanced-Transformer model

More Information
  • This study focuses on vehicle trajectory prediction on highways and urban expressways with the aim of enhancing traffic safety. Leveraging the CitySim dataset, three critical scenarios are examined: basic highway segments, highway merge/diverge areas, and urban expressway weaving zones. For each scenario, Recurrent Neural Network (RNN), Long Short-Term Memory (LSTM), Graph Neural Network (GNN), and Transformer models are employed to predict trajectories 5 s ahead using 5 s of historical data. To improve prediction performance, this study proposes key enhancements centered on kinematic feature optimization, culminating in the development of the Kinematic Feature Enhanced-Transformer (KF-Transformer) model. The first enhancement modifies the attention mechanism to a relative attention framework incorporating position bias terms. The second replaces the feedforward network with a convolution-enhanced variant that integrates depthwise separable and dilated convolutions. Comparative experimental results demonstrate that the KF-Transformer model achieves superior performance, with Root Mean Square Error (RMSE) values of 1.93, 2.07, and 2.33 m recorded for basic highway segments, highway merge/diverge areas, and urban expressway weaving zones, respectively. The proposed model exhibits promising potential for application in Advanced Driver-Assistance Systems (ADAS), enabling proactive risk anticipation and effective accident mitigation.
  • 加载中
  • [1] Messaoud K, Yahiaoui I, Verroust-Blondet A, Nashashibi F. 2021. Attention based vehicle trajectory prediction. IEEE Transactions on Intelligent Vehicles 6(1):175−185 doi: 10.1109/TIV.2020.2991952

    CrossRef   Google Scholar

    [2] Morsali M, Frisk E, Åslund J. 2021. Spatio-temporal planning in multi-vehicle scenarios for autonomous vehicle using support vector machines. IEEE Transactions on Intelligent Vehicles 6(4):611−621 doi: 10.1109/TIV.2020.3042087

    CrossRef   Google Scholar

    [3] Feng Y, Yan X. 2022. Support vector machine based lane-changing behavior recognition and lateral trajectory prediction. Computational Intelligence and Neuroscience 2022:3632333 doi: 10.1155/2022/3632333

    CrossRef   Google Scholar

    [4] Li G, Fang S, Ma J, Cheng J. 2020. Modeling merging acceleration and deceleration behavior based on gradient-boosting decision tree. Journal of Transportation Engineering, Part A: Systems 146(7):05020005 doi: 10.1061/JTEPBS.0000386

    CrossRef   Google Scholar

    [5] Choi D, Yim J, Baek M, Lee S. 2021. Machine learning-based vehicle trajectory prediction using V2V communications and on-board sensors. Electronics 10(4):420 doi: 10.3390/electronics10040420

    CrossRef   Google Scholar

    [6] Wang W, Xia F, Nie H, Chen Z, Gong Z, et al. 2021. Vehicle trajectory clustering based on dynamic representation learning of Internet of vehicles. IEEE Transactions on Intelligent Transportation Systems 22(6):3567−3576 doi: 10.1109/TITS.2020.2995856

    CrossRef   Google Scholar

    [7] Xing Y, Lv C, Cao D. 2020. Personalized vehicle trajectory prediction based on joint time-series modeling for connected vehicles. IEEE Transactions on Vehicular Technology 69(2):1341−1352 doi: 10.1109/TVT.2019.2960110

    CrossRef   Google Scholar

    [8] Fujii R, Vongkulbhisal J, Hachiuma R, Saito H. 2021. A two-block RNN-based trajectory prediction from incomplete trajectory. IEEE Access 9:56140−56151 doi: 10.1109/ACCESS.2021.3072135

    CrossRef   Google Scholar

    [9] Lin L, Li W, Bi H, Qin L. 2022. Vehicle trajectory prediction using LSTMs with spatial–temporal attention mechanisms. IEEE Intelligent Transportation Systems Magazine 14(2):197−208 doi: 10.1109/MITS.2021.3049404

    CrossRef   Google Scholar

    [10] Qu D, Wang S, Liu H, Meng Y. 2022. A car-following model based on trajectory data for connected and automated vehicles to predict trajectory of human-driven vehicles. Sustainability 14(12):7045 doi: 10.3390/su14127045

    CrossRef   Google Scholar

    [11] Abdeljaber O, Younis A, Alhajyaseen W. 2020. Extraction of vehicle turning trajectories at signalized intersections using convolutional neural networks. Arabian Journal for Science and Engineering 45(10):8011−8025 doi: 10.1007/s13369-020-04546-y

    CrossRef   Google Scholar

    [12] Sheng Z, Xu Y, Xue S, Li D. 2022. Graph-based spatial-temporal convolutional network for vehicle trajectory prediction in autonomous driving. IEEE Transactions on Intelligent Transportation Systems 23(10):17654−17665 doi: 10.1109/TITS.2022.3155749

    CrossRef   Google Scholar

    [13] Vaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, et al. 2017. Attention is all you need. NIPS'17: Proceedings of the 31st International Conference on Neural Information Processing Systems 30:5998−6008 doi: 10.5555/3295222.3295349

    CrossRef   Google Scholar

    [14] Chen X, Zhang H, Zhao F, Cai Y, Wang H, et al. 2022. Vehicle trajectory prediction based on intention-aware non-autoregressive transformer with multi-attention learning for Internet of vehicles. IEEE Transactions on Instrumentation and Measurement 71:2513912 doi: 10.1109/TIM.2022.3192056

    CrossRef   Google Scholar

    [15] Tang Y, He H, Wang Y. 2024. Hierarchical vector transformer vehicle trajectories prediction with diffusion convolutional neural networks. Neurocomputing 580:127526 doi: 10.1016/j.neucom.2024.127526

    CrossRef   Google Scholar

    [16] Zheng O, Abdel-Aty M, Yue L, Abdelraouf A, Wang Z, et al. 2024. CitySim: a drone-based vehicle trajectory dataset for safety-oriented research and digital twins. Transportation Research Record: Journal of the Transportation Research Board 2678(4):606−621 doi: 10.1177/03611981231185768

    CrossRef   Google Scholar

  • Cite this article

    Yan H, Pan S, Hao M, Ren L. 2026. Vehicle trajectory prediction using a kinematic feature enhanced-Transformer model. Digital Transportation and Safety 5(3): 248−254 doi: 10.48130/dts-0026-0020
    Yan H, Pan S, Hao M, Ren L. 2026. Vehicle trajectory prediction using a kinematic feature enhanced-Transformer model. Digital Transportation and Safety 5(3): 248−254 doi: 10.48130/dts-0026-0020

Figures(4)  /  Tables(4)

Article Metrics

Article views(33) PDF downloads(2)

Other Articles By Authors

ARTICLE   Open Access    

Vehicle trajectory prediction using a kinematic feature enhanced-Transformer model

Digital Transportation and Safety  5,  2026, 5(3): 248−254  |  Cite this article

Abstract: This study focuses on vehicle trajectory prediction on highways and urban expressways with the aim of enhancing traffic safety. Leveraging the CitySim dataset, three critical scenarios are examined: basic highway segments, highway merge/diverge areas, and urban expressway weaving zones. For each scenario, Recurrent Neural Network (RNN), Long Short-Term Memory (LSTM), Graph Neural Network (GNN), and Transformer models are employed to predict trajectories 5 s ahead using 5 s of historical data. To improve prediction performance, this study proposes key enhancements centered on kinematic feature optimization, culminating in the development of the Kinematic Feature Enhanced-Transformer (KF-Transformer) model. The first enhancement modifies the attention mechanism to a relative attention framework incorporating position bias terms. The second replaces the feedforward network with a convolution-enhanced variant that integrates depthwise separable and dilated convolutions. Comparative experimental results demonstrate that the KF-Transformer model achieves superior performance, with Root Mean Square Error (RMSE) values of 1.93, 2.07, and 2.33 m recorded for basic highway segments, highway merge/diverge areas, and urban expressway weaving zones, respectively. The proposed model exhibits promising potential for application in Advanced Driver-Assistance Systems (ADAS), enabling proactive risk anticipation and effective accident mitigation.

    • With the rapid expansion of transportation infrastructure and the surge in vehicle ownership, road safety has increasingly emerged as a critical bottleneck restricting sustainable development. According to the Global Status Report on Road Safety 2023 released by the World Health Organization, road traffic accidents claimed 1.19 million lives worldwide in 2021. Notably, the risk of severe accidents on highways and urban expressways is significantly higher than on conventional roads and urban streets. This phenomenon is primarily attributed to the higher vehicle speeds in uninterrupted flow environments and the denser traffic volumes on these high-grade roadways.

      An in-depth analysis of accident causation reveals that drivers' misjudgment of surrounding vehicles' future trajectories constitutes a major contributor to traffic collisions. Particularly on high-grade roads such as highways and urban expressways, severe crashes are more likely to occur when nearby vehicles execute maneuvers involving significant trajectory changes, such as ramp merging, lane changing, or sudden braking. However, human drivers possess inherent limitations in predicting surrounding vehicles' trajectories. Firstly, their observational capacity is constrained, making it difficult to accurately assess the position and speed of nearby vehicles. Secondly, their analytical ability is limited, hindering the precise extraction of surrounding vehicles' motion characteristics. Coupled with factors such as driver fatigue and distracted driving, delayed recognition of behavioral cues preceding trajectory changes further exacerbates accident risks.

      Given the severity of road safety challenges and the limitations of human driving, automated and accurate vehicle trajectory prediction has gained significant importance. At its core, trajectory prediction involves modeling historical and real-time data to infer a vehicle's future motion path and behavioral intent[1]. In recent years, with the advancement of autonomous driving technology, vehicle trajectory prediction has been widely applied in practical scenarios. For lower-level autonomous systems, trajectory prediction can issue timely alerts to drivers before hazardous situations occur. For higher-level autonomous systems, it enables proactive vehicle control to mitigate potential risks. Deep learning serves as the primary technical foundation for trajectory prediction, where models learn inherent vehicle motion patterns from historical data to forecast future trajectories.

      This study aims to enhance trajectory prediction accuracy by optimizing the architecture of the Transformer model, with a specific focus on strengthening its capability to learn vehicle kinematic features. The proposed modifications are expected to yield a more precise and reliable vehicle trajectory prediction method. Meanwhile, the proposed method can simultaneously predict the trajectories of all vehicles on a certain section, thereby improving efficiency. The flowchart of this study is illustrated in Fig. 1.

      Figure 1. 

      Flowchart showing the present study.

    • The development of vehicle trajectory prediction technology has undergone multiple stages. Early vehicle trajectory prediction primarily relied on the experience and judgment of human drivers to forecast future driving states. This method suffered from low accuracy and was heavily influenced by driver experience and psychological factors. With the rapid advancement of machine learning techniques, supervised and unsupervised learning methods began to be applied to vehicle trajectory prediction. In recent years, deep learning has achieved significant progress in vehicle trajectory prediction. Currently, Transformer-based models are extensively studied for vehicle trajectory prediction, with numerous studies highlighting their exceptional performance.

      Machine learning-based vehicle trajectory prediction employs models such as support vector machines (SVM), decision trees, random forests, and k-nearest neighbors (KNN) to model vehicle trajectories. These methods typically require feature engineering, including relative position, velocity, acceleration, and driver behavior characteristics. Trained models can predict future trajectories based on current states. Morsali et al.[2] formulated vehicle trajectory prediction as an SVM-based classification problem, distinguishing between obstacle and obstacle-free scenarios using vehicle position, orientation, and velocity data. Their approach incorporated heuristic algorithms and pruning techniques, significantly improving prediction efficiency. Feng et al.[3] leveraged the nonlinear learning and high pattern recognition capabilities of SVMs to extract key features from lane-changing trajectory data, enabling accurate modeling and prediction of lateral vehicle motion. Their results demonstrated strong predictive performance and robustness. Li et al.[4] applied gradient-boosted decision trees to study highway vehicle trajectories, effectively capturing nonlinear and complex behavioral patterns related to acceleration and deceleration. This method achieved promising results in real-world U.S. testing. Choi et al.[5] proposed a random forest-based approach, defining occupancy grid maps around target vehicles and predicting future grid positions. Experiments across diverse driving scenarios confirmed its superior performance and robustness. Wang et al.[6] dynamically constructed a KNN-based trajectory prediction model, incorporating representation learning to derive low-dimensional vehicle embeddings. Their clustering-based approach enhanced prediction accuracy. While machine learning models are computationally efficient and easy to deploy, their limitations include difficulty in capturing complex nonlinear relationships and long-range dependencies, restricting their performance in intricate prediction tasks.

      Deep learning methods utilize neural networks, including RNNs, LSTMs, GNNs and convolutional neural networks (CNNs), to autonomously learn behavioral patterns from historical data without explicit feature engineering. These methods excel in predicting trajectories in complex traffic scenarios. Xing et al.[7] developed a personalized RNN-based joint time-series model for predicting leading vehicle trajectories, incorporating memory layers and regression layers tailored to different driver behaviors. Their model demonstrated significant advantages in handling heterogeneous driving styles. Fujii et al.[8] introduced a dual-RNN architecture, with one RNN fitting Bayesian filter outputs and the other performing inverse fitting, achieving improved prediction accuracy. Lin et al.[9] proposed a spatiotemporal attention LSTM model that not only matched state-of-the-art performance but also explained the influence of historical trajectories and neighboring vehicles. Qu et al.[10] designed a CNN-based data-driven car-following model, demonstrating effective trajectory prediction after noise reduction and structural optimization. Abdeljaber et al.[11] employed a region-based CNN for vehicle tracking in video frames, achieving acceptable accuracy. Sheng et al.[12] introduced a graph spatiotemporal convolutional network to predict future trajectory distributions, combining GNN for spatial interactions and CNN for temporal features. Deep learning excels in automatic feature extraction and capturing complex temporal dependencies, offering advantages in handling high-dimensional and nonlinear data. However, its drawbacks include reliance on large datasets, high computational costs, and complex training and tuning processes.

      Transformer models, originally introduced by Vaswani et al.[13], are based primarily on self-attention mechanisms rather than recurrent or convolutional structures. Input information is first transformed into high-dimensional feature representations, while positional encoding is incorporated to preserve the order and relative positions of elements within a sequence. The self-attention mechanism evaluates the relationships between all elements in the input sequence and assigns different attention weights according to their relevance, allowing the model to selectively aggregate global contextual information. Multi-head attention performs this process in multiple feature subspaces, enabling the model to learn diverse relationships and complementary representations simultaneously. The resulting features are further processed through feedforward networks, residual connections, and layer normalization. By stacking multiple attention and feedforward layers, Transformer models can progressively extract high-level features and effectively capture long-range dependencies within sequential data.

      Transformer models' application in vehicle trajectory prediction has gained significant attention. Chen et al.[14] integrated graph attention mechanisms with Transformer encoders to model inter-vehicle interactions, achieving accurate multimodal trajectory prediction. Tang et al.[15] proposed a hierarchical vectorized Transformer, decomposing traffic scenes into local and global information while capturing feature uncertainty through a local diffusion encoder. Transformers excel in capturing long-range dependencies and offer scalability and flexibility for complex tasks. However, they demand substantial data and computational resources.

      Researchers have extensively studied vehicle trajectory prediction, proposing numerous effective methods. Nevertheless, existing studies exhibit several limitations: (1) Most research focuses on simplified scenarios, such as highway straightaways, urban road segments, and intersections, while neglecting high-risk areas like highway merge/diverge zones and urban expressway weaving sections. (2) Predominant approaches rely on generic RNN, LSTM, GNN, or Transformer architectures without tailoring models to the inherent patterns of trajectory data. These shortcomings impact prediction accuracy and road safety, necessitating further research and resolution.

    • This study utilizes the open-source CitySim dataset[16], which was collected via drone aerial photography and processed into tabular format. The dataset encompasses various road scenarios, of which three are selected for this research: basic highway segments, highway merge/diverge areas, and urban expressway weaving zones. For each recording location, data were captured at 30 fps over a 300-s duration, resulting in 9,000 timesteps per location. The dataset includes vehicle position (x, y coordinates). The modeling approach involves using 5-s historical trajectories to predict 5-s future trajectories, a common practice in vehicle trajectory prediction research.

      The basic highway segment consists of an 800-m straight section with two travel directions, each containing three main lanes and one emergency lane. The highway merge/diverge area is similarly structured with an 800-m straight segment, where the bottom travel direction (left-to-right) includes a single-lane exit ramp while the top direction (right-to-left) contains a single-lane entrance ramp. The urban expressway weaving zone features a 300-m straight segment where each main direction expands from three to four lanes in the weaving zone, with complex ramp configurations including both entrance and exit ramps on each side. The three scenarios are shown in Fig. 2.

      Figure 2. 

      Presentation of the three scenarios. (a) Basic highway segments. (b) Highway merge/diverge areas. (c) Urban expressway weaving zones.

      To prepare the data for modeling, several preprocessing steps were implemented. Since some vehicles appeared for insufficient durations due to late entry/exit or data loss, all vehicles with fewer than 300 timesteps (10 s) were filtered out to ensure adequate historical data for prediction. The final dataset contains 1,853 vehicle trajectories across the three scenarios, with detailed breakdowns for each movement pattern.

      Data standardization was achieved through strategic truncation based on vehicle movement patterns. Vehicles entering from ramps were truncated to their first 300 timesteps, where behavioral changes predominantly occur, while exiting vehicles were truncated to their last 300 timesteps. Straight-through vehicles used their first 300 timesteps. Additionally, noise reduction was applied using a moving average smoothing technique with a five-timestep window on positional coordinates to improve data quality. The five-timestep window was selected to effectively suppress short-term measurement noise while preserving meaningful changes in vehicle motion.

      The processed dataset was partitioned into training (70%) and test (30%) sets while maintaining proportional representation of all movement patterns within each scenario, as shown in Table 1. This careful partitioning ensures balanced learning and reliable evaluation, with both sets undergoing normalization before model training and testing.

      Table 1.  The number of vehicles in each road scenario.

      Road scenario Training set Test set Total
      Basic highway segments 601 257 858
      Highway merge/diverge areas 451 193 644
      Urban expressway weaving zones 246 105 351
    • To achieve accurate vehicle trajectory prediction, we developed a KF-Transformer model with two key architectural innovations tailored for motion pattern extraction. The model architecture begins with an embedding layer that performs feature embedding and positional encoding of the input trajectory data. The embedded features then pass through six encoder layers, each comprising multi-head attention mechanisms and feedforward neural networks, with residual connections and layer normalization applied after each operation. The final encoder output serves as the key and value matrices for all six decoder layers.

      Each decoder layer sequentially processes the output from the previous decoder layer as the query matrix, combined with the key and value matrices for computation. The decoder structure incorporates masked multi-head attention, standard multi-head attention, and feedforward networks, all enhanced with residual connections and layer normalization. The ultimate decoder output undergoes a linear transformation to generate the predicted trajectory coordinates.

      The first innovation introduces the Multi-Head Relative Attention Mechanism, which overcomes the limitations of traditional absolute positional encoding by explicitly modeling relative positional relationships through learnable bias terms, as shown in Fig. 3. This mechanism enables each attention head to develop specialized focus patterns. Some heads capture short-range movements like acceleration/deceleration, while others track long-term driving patterns, as illustrated in Eq. (1). The dynamic adjustment of interaction strengths between elements at different distances significantly improves the model's ability to capture spatiotemporal dependencies in vehicle motion.

      $ RelativeAttention\ \left(Q,\ K,\ V\right)\ =\ softmax\ \left(\dfrac{QK^T+R}{\sqrt{d_k}}\right)\ V $ (1)

      where, $ Q $ is the query matrix, $ K $ is the key matrix, $ V $ is the value matrix, and $ \sqrt{{d}_{k}} $ is the scaling factor. The relative positional bias term is essentially represented by a three-dimensional tensor $ R $, whose three dimensions correspond to the sequence length, the sequence length, and the number of attention heads, respectively. Each element $ R_{ij}^{h} $ denotes the relative positional bias between positions $ i $ and $ j $ in the hth attention head.

      Figure 3. 

      Multi-head relative attention mechanism.

      The second enhancement involves a Convolution-Enhanced Feedforward Neural Network that addresses the limitation of standard feedforward networks in perceiving local temporal variations, as shown in Fig. 4. By integrating depthwise separable convolutions and multi-scale dilated convolutions into the feedforward structure, local inductive biases without substantially increasing parameters are introduced, as illustrated in Eqs (2) and (3). The sliding window mechanism forces focus on local temporal patterns, while dilated convolutions enable multi-granularity feature fusion. Residual connections and dynamic gradient modulation effectively combine local motion features with global trajectory semantics.

      $ Z_{conv}\ =\ DepthwiseConv1D\ (Z;\ k,\ s,\ p) $ (2)
      $ Z_{dilated}\ =\ DilatedConv1D\ (Z_{conv};\ k,\ r) $ (3)

      where, $ Z $ denotes the feature vector, and $ k $ denotes the convolution kernel size, which is set to 5. Each channel independently employs its own convolution kernel parameters. The stride $ s $ is set to 1 to preserve the sequence length. The padding $ p $ is set to 2 according to the kernel size, thereby preventing the loss of boundary information. The dilation rate $ r $ represents the spacing between adjacent elements of the convolution kernel.

      Figure 4. 

      Convolution-enhanced feedforward neural network.

      These architectural innovations collectively enable the KF-Transformer model to capture both microscopic motion patterns and macroscopic driving behaviors, while maintaining physical consistency through explicit kinematic constraints in the feature space. The KF-Transformer architecture has potential in complex scenarios like weaving sections and merge/diverge areas, where traditional models often fail to capture rapid motion transitions.

    • To comprehensively evaluate the performance of various deep learning models for vehicle trajectory prediction, we compared the Root Mean Square Error (RMSE in meters) of the RNN, LSTM, GNN, Transformer, and KF-Transformer models across different scenarios, as presented in Table 2.

      Table 2.  Performance of various deep learning models in vehicle trajectory prediction across different scenarios.

      Models Highway basic segments Highway merge/
      diverge areas
      Urban expressway weaving zones
      RNN 2.15 2.40 2.86
      LSTM 2.11 2.37 2.83
      GNN 2.12 2.34 2.75
      Transformer 2.08 2.35 2.73
      KF-Transformer 1.93 2.07 2.33

      The results demonstrate that the KF-Transformer model achieves superior prediction accuracy in all three scenarios. Specifically, it attains an RMSE of 1.93 m on basic highway segments, representing a 0.15-m improvement over the standard Transformer model. In highway merge/diverge areas, it achieves an RMSE of 2.07 m (0.28 m better than Transformer), while in urban expressway weaving zones, it reduces the RMSE to 2.33 m (0.40 m improvement). These results clearly highlight the advantages of the proposed architecture. The comparative analysis also reveals significant variations in prediction difficulty across scenarios—urban expressway weaving zones present the greatest challenge, followed by highway merge/diverge areas, with basic highway segments being the most straightforward to predict.

      In practical applications of trajectory prediction models, computational efficiency is equally crucial as prediction accuracy. Rigorous speed tests were conducted by measuring the time required for each model to predict 5-s (150-timestep) trajectories for individual vehicles, using identical hyperparameters across all models. Given the extremely fast processing speeds (on the order of ms), the computation times across all test vehicles were averaged to obtain reliable measurements, as shown in Table 3.

      Table 3.  Efficiency of vehicle trajectory prediction by various deep learning models.

      Models Time (ms)
      RNN 22
      LSTM 37
      GNN 65
      Transformer 71
      KF-Transformer 75

      The benchmarking results indicate that the RNN model achieves the fastest prediction speed, requiring only 22 ms per vehicle on an NVIDIA GeForce RTX 3080 Laptop GPU. In contrast, the KF-Transformer model demonstrates the slowest processing speed at 75 ms per vehicle. Importantly, even this maximum latency of 75 ms remains well within acceptable limits for real-world automotive applications, including autonomous driving systems. For practical deployment, the actual prediction speed should be calculated based on the specific hardware capabilities and the number of concurrent prediction requests the system needs to handle, ensuring timely trajectory predictions under operational conditions.

      These findings collectively demonstrate that the KF-Transformer model successfully balances prediction accuracy and computational efficiency, making it suitable for real-time trajectory prediction applications while significantly outperforming conventional approaches in complex driving scenarios. The slightly increased computation time is justified by the substantial improvements in prediction accuracy, particularly in challenging environments like urban weaving zones where precise trajectory forecasting is most critical for safety.

      To evaluate the individual contributions of the Multi-Head Relative Attention Mechanism and the Convolution-Enhanced Feedforward Neural Network in the KF-Transformer, ablation experiments were conducted, and the results are presented in Table 4. The experimental results demonstrate that both modifications improve the performance of the Transformer in vehicle trajectory prediction, while their combined use achieves superior performance.

      Table 4.  Ablation experiments of the KF-Transformer for vehicle trajectory prediction.

      Models Highway basic
      segments
      Highway merge/
      diverge areas
      Urban expressway
      weaving zones
      Transformer 2.08 2.35 2.73
      Transformer-multi-head relative attention mechanism 1.95 2.11 2.45
      Transformer-convolution enhanced feedforward neural network 1.97 2.16 2.40
      KF-Transformer 1.93 2.07 2.33
    • To enhance traffic safety on highways and urban expressways, this study investigates vehicle trajectory prediction in these critical scenarios. We focused on three representative road configurations: basic highway segments, highway merge/diverge areas, and urban expressway weaving zones, encompassing various vehicle movement patterns including through traffic, ramp entries, and ramp exits.

      The research introduces two key improvements to the Transformer architecture: (1) We modified the attention mechanism to incorporate relative position bias; and (2) redesigned the feedforward network with depthwise separable and dilated convolutions. The performance evaluation of the proposed KF-Transformer model against other deep learning approaches yields several important findings:

      (1) The KF-Transformer demonstrates superior prediction accuracy across all road scenarios, achieving RMSE values of 1.93, 2.07, and 2.33 m for 5-s trajectory predictions in basic highway segments, highway merge/diverge areas, and urban expressway weaving zones, respectively.

      (2) Significant variations exist in prediction difficulty among different scenarios, with urban expressway weaving zones presenting the greatest challenges. This highlights critical safety concerns that transportation engineers should prioritize.

      (3) While the RNN model maintains the fastest computation speed, the KF-Transformer preserves computational efficiency through lightweight design principles, avoiding substantial increases in processing time compared to the standard Transformer.

      This study has limitations that need to be overcome in future research. (1) The current investigation is limited to three specific scenarios, while vehicle trajectory characteristics vary across different traffic conditions, environmental factors, and road classifications. Future work should expand to more diverse operational contexts. (2) Dataset limitations present another constraint, as the study relied solely on the CitySim dataset. Subsequent research should incorporate multi-source heterogeneous data and larger-scale datasets to validate model generalizability and robustness.

      These findings contribute to the development of more accurate and practical trajectory prediction systems, particularly for complex road configurations where precise forecasting is crucial for traffic safety. The proposed architectural improvements demonstrate significant performance gains while maintaining computational feasibility for real-world deployment. Future research directions outlined in this study will further advance the field of intelligent transportation systems and connected vehicle technologies.

      • The authors confirm their contributions to the paper as follows: study conception and design: Yan H, Pan S, Ren L; analysis and interpretation of results: Yan H, Pan S, Hao M; draft manuscript preparation: Yan H, Pan S, Hao M, Ren L. All authors reviewed the results and approved the final version of the manuscript.

      • The authors declare that they have no conflict of interest.

      • Copyright: © 2026 by the author(s). Published by Maximum Academic Press, Fayetteville, GA. This article is an open access article distributed under Creative Commons Attribution License (CC BY 4.0), visit https://creativecommons.org/licenses/by/4.0/.
    Figure (4)  Table (4) References (16)
  • About this article
    Cite this article
    Yan H, Pan S, Hao M, Ren L. 2026. Vehicle trajectory prediction using a kinematic feature enhanced-Transformer model. Digital Transportation and Safety 5(3): 248−254 doi: 10.48130/dts-0026-0020
    Yan H, Pan S, Hao M, Ren L. 2026. Vehicle trajectory prediction using a kinematic feature enhanced-Transformer model. Digital Transportation and Safety 5(3): 248−254 doi: 10.48130/dts-0026-0020

Catalog

    /

    DownLoad:  Full-Size Img  PowerPoint
    Return
    Return