<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.1 20151215//EN" "http://jats.nlm.nih.gov/publishing/1.1/JATS-journalpublishing1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xml:lang="en" article-type="research-article" dtd-version="1.1">
<front>
<journal-meta>
<journal-id journal-id-type="pmc">CMC</journal-id>
<journal-id journal-id-type="nlm-ta">CMC</journal-id>
<journal-id journal-id-type="publisher-id">CMC</journal-id>
<journal-title-group>
<journal-title>Computers, Materials &#x0026; Continua</journal-title>
</journal-title-group>
<issn pub-type="epub">1546-2226</issn>
<issn pub-type="ppub">1546-2218</issn>
<publisher>
<publisher-name>Tech Science Press</publisher-name>
<publisher-loc>USA</publisher-loc>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">52323</article-id>
<article-id pub-id-type="doi">10.32604/cmc.2024.052323</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Article</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>Dynamic Forecasting of Traffic Event Duration in Istanbul: A Classification Approach with Real-Time Data Integration</article-title>
<alt-title alt-title-type="left-running-head">Dynamic Forecasting of Traffic Event Duration in Istanbul: A Classification Approach with Real-Time Data Integration</alt-title>
<alt-title alt-title-type="right-running-head">Dynamic Forecasting of Traffic Event Duration in Istanbul: A Classification Approach with Real-Time Data Integration</alt-title>
</title-group>
<contrib-group>
<contrib id="author-1" contrib-type="author" corresp="yes">
<name name-style="western"><surname>Ulu</surname><given-names>Mesut</given-names></name><xref ref-type="aff" rid="aff-1">1</xref><email>mulu@bandirma.edu.tr</email></contrib>
<contrib id="author-2" contrib-type="author">
<name name-style="western"><surname>T&#x00FC;rkan</surname><given-names>Yusuf Sait</given-names></name><xref ref-type="aff" rid="aff-2">2</xref></contrib>
<contrib id="author-3" contrib-type="author">
<name name-style="western"><surname>Meng&#x00FC;&#x00E7;</surname><given-names>Kenan</given-names></name><xref ref-type="aff" rid="aff-3">3</xref></contrib>
<contrib id="author-4" contrib-type="author">
<name name-style="western"><surname>Naml&#x0131;</surname><given-names>Ersin</given-names></name><xref ref-type="aff" rid="aff-2">2</xref></contrib>
<contrib id="author-5" contrib-type="author">
<name name-style="western"><surname>K&#x00FC;&#x00E7;&#x00FC;kdeniz</surname><given-names>Tar&#x0131;k</given-names></name><xref ref-type="aff" rid="aff-2">2</xref></contrib>
<aff id="aff-1"><label>1</label><institution>Occupational Health and Safety Department, Bandirma Onyedi Eylul University</institution>, <addr-line>Balikesir, 10200</addr-line>, <country>T&#x00FC;rkiye</country></aff>
<aff id="aff-2"><label>2</label><institution>Department of Industrial Engineering, Istanbul University-Cerrahpasa</institution>, <addr-line>Istanbul, 34320</addr-line>, <country>T&#x00FC;rkiye</country></aff>
<aff id="aff-3"><label>3</label><institution>Department of Industrial Engineering, Istanbul Technical University</institution>, <addr-line>Istanbul, 34467</addr-line>, <country>T&#x00FC;rkiye</country></aff>
</contrib-group>
<author-notes>
<corresp id="cor1"><label>&#x002A;</label>Corresponding Author: Mesut Ulu. Email: <email>mulu@bandirma.edu.tr</email></corresp>
</author-notes>
<pub-date date-type="collection" publication-format="electronic">
<year>2024</year></pub-date>
<pub-date date-type="pub" publication-format="electronic">
<day>15</day>
<month>8</month>
<year>2024</year></pub-date>
<volume>80</volume>
<issue>2</issue>
<fpage>2259</fpage>
<lpage>2281</lpage>
<history>
<date date-type="received">
<day>30</day>
<month>3</month>
<year>2024</year>
</date>
<date date-type="accepted">
<day>21</day>
<month>6</month>
<year>2024</year>
</date>
</history>
<permissions>
<copyright-statement>&#x00A9; 2024 Ulu et al.</copyright-statement>
<copyright-year>2024</copyright-year>
<copyright-holder>Ulu et al.</copyright-holder>
<license xlink:href="https://creativecommons.org/licenses/by/4.0/">
<license-p>This work is licensed under a <ext-link ext-link-type="uri" xlink:type="simple" xlink:href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</ext-link>, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.</license-p>
</license>
</permissions>
<self-uri content-type="pdf" xlink:href="TSP_CMC_52323.pdf"></self-uri>
<abstract>
<p>Today, urban traffic, growing populations, and dense transportation networks are contributing to an increase in traffic incidents. These incidents include traffic accidents, vehicle breakdowns, fires, and traffic disputes, resulting in long waiting times, high carbon emissions, and other undesirable situations. It is vital to estimate incident response times quickly and accurately after traffic incidents occur for the success of incident-related planning and response activities. This study presents a model for forecasting the traffic incident duration of traffic events with high precision. The proposed model goes through a 4-stage process using various features to predict the duration of four different traffic events and presents a feature reduction approach to enable real-time data collection and prediction. In the first stage, the dataset consisting of 24,431 data points and 75 variables is prepared by data collection, merging, missing data processing and data cleaning. In the second stage, models such as Decision Trees (DT), K-Nearest Neighbour (KNN), Random Forest (RF) and Support Vector Machines (SVM) are used and hyperparameter optimisation is performed with GridSearchCV. In the third stage, feature selection and reduction are performed and real-time data are used. In the last stage, model performance with 14 variables is evaluated with metrics such as accuracy, precision, recall, F1-score, MCC, confusion matrix and SHAP. The RF model outperforms other models with an accuracy of 98.5%. The study&#x2019;s prediction results demonstrate that the proposed dynamic prediction model can achieve a high level of success.</p>
</abstract>
<kwd-group kwd-group-type="author">
<kwd>Traffic event duration</kwd>
<kwd>forecasting</kwd>
<kwd>machine learning</kwd>
<kwd>feature reduction</kwd>
<kwd>shapley additive explanations (SHAP)</kwd>
</kwd-group>
</article-meta>
</front>
<body>
<sec id="s1">
<label>1</label>
<title>Introduction</title>
<p>Traffic incidents encompass a variety of events, including traffic accidents, vehicle malfunctions, vehicle fires, and arguments or fights in traffic. These incidents can lead to traffic congestion, which can have significant social, economic, and environmental impacts, such as increased travel time, excessive fuel consumption, air pollution, and stress [<xref ref-type="bibr" rid="ref-1">1</xref>,<xref ref-type="bibr" rid="ref-2">2</xref>]. Predicting the duration of traffic events is a challenging task due to the complexities arising from the stochastic nature of traffic events. Accurate duration prediction offers advantages to drivers in route selection and traffic operations managers in congestion management [<xref ref-type="bibr" rid="ref-3">3</xref>]. Efficient traffic incident management (TIM) is essential for mitigating adverse traffic effects. Continuous enhancement of TIM systems ensures effective incident handling and minimises traffic disruptions. Accurate estimation of event duration, reliant on environmental and event-specific analyses, is pivotal to TIM success. Precise duration estimation informs resource allocation and intervention planning. Diverting affected road users to alternative routes attenuates the incident impact [<xref ref-type="bibr" rid="ref-4">4</xref>].</p>
<p>Forecasting event duration is vital in traffic event management, representing the temporal gap from incident onset to clearance. It is essential for assessing the severity of the incident and determining the temporal and spatial distribution of traffic flow on the road network. Traffic incidents can be divided into sequential and distinct time intervals, as described by several studies [<xref ref-type="bibr" rid="ref-5">5</xref>&#x2013;<xref ref-type="bibr" rid="ref-8">8</xref>]. The duration between the incident occurrence and the response of the traffic control center operators after receiving the call is known as the detection-reporting time. Preparation-dispatch time is the time between the receipt of the call by the operators and the dispatch of the response team members to the incident. The travel time is simply the duration between receiving the dispatch order and arriving at the scene for the incident response team members. Detection-notification time, preparation-dispatch time, and travel time are three important time intervals in incident response. Clean-up time is the time between the arrival of incident response team members at the scene and the completion of the cleanup of the incident. Clean-up time is especially used in planning and dispatching activities in traffic incident management. In this study, incident duration is defined as the time between the occurrence of the incident and the opening of the roadway.</p>
<p>During traffic incidents, TIM centers first attempt to collect incident coordinates and other relevant information. However, this estimation process is challenging due to the complexity of traffic events and the multitude of variables affecting duration. Uncertainty is inherent at the outset and escalates with incident size. The impact of different factors on incident duration during a traffic event or accident may vary depending on the circumstances, such as partial lane closure, complete road closure, or long-term road infrastructure works. TIMs may make a decision that no intervention is necessary in a low-level traffic incident and then make a similar assessment in a very similar incident that actually requires extensive intervention. Real-time data collection is essential to mitigate assessment errors. Machine learning algorithms are currently the most effective tools in prediction studies of traffic events [<xref ref-type="bibr" rid="ref-9">9</xref>].</p>
<p>Recent studies have focused on predicting traffic incident durations with the objective of enhancing resource allocation, emergency response, and traffic management [<xref ref-type="bibr" rid="ref-10">10</xref>&#x2013;<xref ref-type="bibr" rid="ref-14">14</xref>]. These studies examine the sub-components of traffic incident durations, such as intervention, and scene cleaning. They utilize various methodologies, including machine learning models, hazard-based modeling, and ensemble learning approaches, in order to improve the accuracy and interpretability of incident duration predictions. Factors influencing incident duration include road type, casualties, weather conditions, and the number of vehicles. The necessity of considering time-varying traffic variables during incident episodes is emphasized, underscoring the importance of dynamic modeling to capture traffic flow dynamics. For real-time forecasting, variables for which real-time data can be collected should be examined and their success in forecasting investigated. The present study proposes an integrated methodology for dynamic estimating traffic incident duration, which addresses a significant gap in existing literature.</p>
<p>The majority of studies employ regression estimation for the purpose of predicting the duration of a traffic event. In contrast, classification is performed in a relatively limited number of studies [<xref ref-type="bibr" rid="ref-15">15</xref>]. In classification studies, the duration of the components of the traffic incident, such as traffic incident notification, response and traffic accident scene cleaning, was studied instead of the total time between the occurrence of the traffic incident and the cleaning of the scene [<xref ref-type="bibr" rid="ref-16">16</xref>]. Unlike previous studies focusing on regression, our novel approach employs classification methods for real-time prediction, reducing the initially examined variables to enhance predictive accuracy. Furthermore, the current study differs from previous research in that it considers the entire period from the occurrence of a traffic accident to the completion of the cleanup, rather than focusing on specific sub-components of traffic incident duration. The present study employs four machine learning methods to categorize traffic accidents and incidents into four duration classes. In order to facilitate dynamic prediction and more effective solutions, we initially reduced the number of variables, which had previously been extensive. Variables that were ineffective in forecasting were eliminated, and it was determined whether the variables that were effective in forecasting could collect real-time data. Subsequently, numerous databases containing the data of these variables were integrated, and a dynamic forecasting environment was provided. The categorization of incident durations and utilization of real-time data enables the swift intervention of incident management centers, fostering agile decision-making and resource allocation compared to conventional regression-based approaches. Feature selection enables the model to improve interpretability while also optimizing training and execution speed, thus mitigating the risk of overfitting. Ultimately, this study provides a pragmatic solution for dynamic traffic incident management, facilitating strategic interventions and resource optimization in urban transportation systems.</p>
<p>An experimental study was conducted in Istanbul to test the model proposed in the study to predict the duration of traffic events. Istanbul is one of the most crowded and traffic-heavy cities in the world, connecting two continents. To create the model, we identified the variables that affect incident duration and collected data from various sources. The objective is to facilitate prompt intervention from the incident management center by estimating the duration of incidents within a given time period using current conditions and primary data with minimal variables. To achieve this, Istanbul was divided into 682 geohash areas. The models were used to forecast the duration of traffic events during specific time period. The model incorporates dynamic prediction through feature selection to streamline complex data structures into fewer variables, facilitating real-time forecasting. Unlike other forecasting approaches, it responds promptly to traffic condition fluctuations by collecting influential variable data in a real-time database. Temporal and spatial considerations enable precise real-time predictions, enhancing the reliability and efficacy of traffic management and emergency response strategies.</p>
<p>The paper is structured as follows: <xref ref-type="sec" rid="s2">Section 2</xref> reviews prior studies on traffic incident duration, discussing methodologies and models. <xref ref-type="sec" rid="s3">Section 3</xref> outlines our proposed model, including the dataset, performance metrics, and methods employed. <xref ref-type="sec" rid="s4">Section 4</xref> presents experimental findings derived from the model. Finally, the concluding section provides objective evaluations of the study and suggests future research directions.</p>
</sec>
<sec id="s2">
<label>2</label>
<title>Literature Review</title>
<p>The prediction of traffic event duration plays a crucial role in enhancing TIM systems. Prior research has employed diverse regression models and statistical forecasting techniques for forecasting the duration of traffic incidents. Literature suggests that incident delay and duration exhibit variability contingent upon factors encompassing environmental conditions and incident-specific characteristics. Models predicting the duration of traffic incidents are based on various factors, such as incident type, time of day, weather conditions, and traffic volume. These models include regression models [<xref ref-type="bibr" rid="ref-10">10</xref>,<xref ref-type="bibr" rid="ref-15">15</xref>,<xref ref-type="bibr" rid="ref-17">17</xref>], cyclic subspace regression [<xref ref-type="bibr" rid="ref-18">18</xref>], and probabilistic statistical models [<xref ref-type="bibr" rid="ref-19">19</xref>&#x2013;<xref ref-type="bibr" rid="ref-21">21</xref>]. Hazard-based models were studied by Hojati et al. [<xref ref-type="bibr" rid="ref-1">1</xref>,<xref ref-type="bibr" rid="ref-11">11</xref>,<xref ref-type="bibr" rid="ref-22">22</xref>&#x2013;<xref ref-type="bibr" rid="ref-24">24</xref>]. Lee et al. [<xref ref-type="bibr" rid="ref-25">25</xref>] analyzed structural equation models, while Zou et al. [<xref ref-type="bibr" rid="ref-26">26</xref>,<xref ref-type="bibr" rid="ref-27">27</xref>] examined finite mixture models. Zou et al. [<xref ref-type="bibr" rid="ref-27">27</xref>] and Laman et al. [<xref ref-type="bibr" rid="ref-16">16</xref>] investigated copula-based models. The literature identifies incident characteristics (e.g., type of incident, first responder, and number of responders), road features (e.g., average annual daily traffic, geometric features, and functional classification), traffic situations (e.g., month, day, and time), and weather conditions (e.g., season, precipitation, temperature, and wind) as the most important independent variables for the developed models.</p>
<p>After a traffic incident, certain information, such as the number of injured individuals, number of vehicles involved, and cause of the accident, may not be immediately available. To improve model accuracy, a dynamic prediction model should be developed gradually using primary data such as the time and location of the incident, incident type, notification time, and weather conditions obtained after the incident notification. Secondary data, encompassing casualty counts, response time, and extent of lane obstruction, may serve as inputs for the initial prediction and subsequent refinement in a two-stage process [<xref ref-type="bibr" rid="ref-28">28</xref>]. Obtaining clear data on the number of fatalities or injuries may take days after the accident has occurred, and the scene may have been cleared during this time, making dynamic forecasting difficult.</p>
<p>Time-series methods can be used to calculate the duration of traffic incidents [<xref ref-type="bibr" rid="ref-29">29</xref>]. Various approaches and techniques, such as genetic algorithms [<xref ref-type="bibr" rid="ref-30">30</xref>], fuzzy logic [<xref ref-type="bibr" rid="ref-31">31</xref>], and Bayesian networks [<xref ref-type="bibr" rid="ref-32">32</xref>,<xref ref-type="bibr" rid="ref-33">33</xref>], can be employed in these models. However, machine learning algorithms have become more preferred over traditional methods in recent years due to the many variables that affect time and the constantly changing traffic conditions. Machine learning models are utilized to predict the duration of traffic incidents by analyzing historical data [<xref ref-type="bibr" rid="ref-9">9</xref>]. These models may undergo training utilising datasets incorporating variables such as the type of event, temporal occurrence, traffic density, meteorological parameters, and event duration. Several machine learning and data mining techniques have been employed to predict the duration of traffic incidents. These models are frequently used in decision trees [<xref ref-type="bibr" rid="ref-34">34</xref>,<xref ref-type="bibr" rid="ref-35">35</xref>], artificial neural networks [<xref ref-type="bibr" rid="ref-9">9</xref>,<xref ref-type="bibr" rid="ref-28">28</xref>], support vector machines [<xref ref-type="bibr" rid="ref-36">36</xref>,<xref ref-type="bibr" rid="ref-37">37</xref>], and random forests [<xref ref-type="bibr" rid="ref-9">9</xref>,<xref ref-type="bibr" rid="ref-12">12</xref>]. Li et al. [<xref ref-type="bibr" rid="ref-38">38</xref>] utilized deep learning in their recent studies. Deep learning, particularly Graph Neural Networks (GNNs), plays a pivotal role in Intelligent Transportation Systems (ITS). GNNs are extensively utilized in ITS applications due to their capacity to analyze graph-structured data effectively. They have evolved for a multitude of ITS tasks, including traffic forecasting, demand prediction, autonomous vehicles, intersection management, and urban planning [<xref ref-type="bibr" rid="ref-39">39</xref>]. Moreover, the integration of deep reinforcement learning (DRL) within connected and automated transportation systems has demonstrated potential in tasks related to automated driving systems and connected-vehicle applications [<xref ref-type="bibr" rid="ref-40">40</xref>]. These advancements in deep learning, including GNNs and DRL, are enhancing the efficiency, safety, and coordination of transportation modes within modern ITS infrastructure, thereby illustrating the potential of data-driven solutions to address complex challenges in the transportation domain.</p>
<p>Traffic incident duration can be divided into sequential time intervals, typically two, three, four, or five. Traffic incident durations are classified as either other or severe/major incidents. Lin et al. [<xref ref-type="bibr" rid="ref-41">41</xref>] defined durations as below or above 60 min, while Zhang et al. [<xref ref-type="bibr" rid="ref-42">42</xref>] defined them as below or above 120 min. The triple classification divides tasks into minor/short, medium, and major/long. According to Smith et al. [<xref ref-type="bibr" rid="ref-43">43</xref>], short tasks take less than 15 min, medium tasks take between 15 and 30 min, and long tasks take over 30 min. The US Department of Transportation [<xref ref-type="bibr" rid="ref-44">44</xref>] and Islam [<xref ref-type="bibr" rid="ref-45">45</xref>] define short tasks as taking less than 30 min, medium tasks as taking between 30&#x2013;120 min, and long tasks as taking over 120 min. In this study traffic event durations were stratified into four distinct categories, delineated according to the frequency of minor events and the implementation of traffic event management protocols. Simple incidents are those that last less than 10 min and do not require intervention. Minor incidents are incidents that require intervention but last between 10 and 30 min, while mid-level incidents last between 31 and 60 min. The study considered incidents that required 61 min or more and a large amount of resources as major incidents.</p>
<p>Recent studies have shown an increase in research on estimating traffic incident duration. Unlike various statistical methods for designing data-driven models, machine learning techniques are frequently utilized and have demonstrated effectiveness [<xref ref-type="bibr" rid="ref-8">8</xref>,<xref ref-type="bibr" rid="ref-38">38</xref>]. These studies have examined the factors that affect incident duration, focusing on individual components of post-event duration, such as notification, response, and cleanup [<xref ref-type="bibr" rid="ref-16">16</xref>]. Our study proposes a model that takes a different approach to estimating traffic event durations. The model employs four machine learning methods and was tested in an experimental study conducted in Istanbul, a large metropolis with complex traffic. The study classified traffic accidents and incidents into four duration classes and estimated their duration based on which duration class the incidents were in. The model simplified the problem by reducing the number of complex features. As a result, instead of extracting data from numerous databases, a few databases were integrated, enabling real-time data for dynamic forecasting.</p>
</sec>
<sec id="s3">
<label>3</label>
<title>Methodology</title>
<p>This paper presents a model for forecasting the duration of traffic incidents and an experimental study of the model. In this context, the prediction of the duration of traffic incidents is approached as a classification problem rather than a regression problem. The accurate forecasting of the duration of a traffic accident is very difficult, and the specific estimation of the duration is not very useful for traffic management. Instead, knowledge of the accident duration class is much more useful for traffic management. This is because interventions and resources allocated for similar accidents with similar durations of traffic incidents do not differ. The purpose of this study is to manage traffic events and provide resource management by enabling decision-makers to act strategically. This includes determining whether urgent measures need to be taken and directing traffic police based on the duration of the incident in an agile manner. Additionally, reducing the number of features can simplify the model, resulting in faster training and execution, and decreasing the risk of overfitting. This can also improve the model&#x2019;s ability to generalize and make it more interpretable. It can also prevent resource waste and facilitate effective management by eliminating unnecessary features.</p>
<p>Four machine learning methods were selected for the experimental study: Decision Trees (DT), K-Nearest Neighbors (KNN), Random Forest (RF), and Support Vector Machine (SVM). DT is suitable for modeling simple, non-linear relationships, while KNN excels in classification with easy integration of new data points. RF is chosen for its high performance and robustness, while SVM demonstrates excellent generalization ability and adapts well to high-dimensional data. While other classification methods exist, this study focuses on these four due to their strong performance and general availability, considering factors such as problem requirements and data structure. Furthermore, the combination of databases of variables is employed for real-time prediction. The proposed approach, as illustrated in <xref ref-type="fig" rid="fig-1">Fig. 1</xref>, allows for the immediate estimation of the duration of traffic incidents.</p>
<fig id="fig-1">
<label>Figure 1</label>
<caption>
<title>Stages of the proposed approach</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_52323-fig-1.tif"/>
</fig>
<p>The proposed approach for predicting traffic events comprises four distinct stages. The initial stage involves data preprocessing, which is undertaken to refine the dataset for subsequent analysis. The creation of the data set was realized in four steps. In the first step, data was collected with information from nine different open sources or institutions. In the second step, merging was done according to time and location information. In this merging, coordinate information about the location was converted into Geohash codes with ArcGIS. This is because when the coordinates are pointwise, the merging process will be very difficult. In the study, other variables existing in the existing area were integrated thanks to the coordinates converted to a 6-cell geohash area (0.74 km<sup>2</sup> area with a cell width of 1.22 km and a cell height of 0.61 km). In this context, it is thought that data merging is rarely done. Missing data were then identified. In cases where missing data were not eliminated, data cleaning was performed, and a data set was prepared to estimate traffic event duration in certain areas. Following this, in the second stage, a non-real-time prediction of the times of the traffic events is carried out using machine learning models. These models leverage the abundance of variables inherent in the initial problem domain. Since this is a critical stage in training four different machine learning models, the hyperparameters must be done carefully. In this context, GridSearchCV was integrated into four different machine learning models and hyperparameters were adjusted. At the same time, the performance of four different machine learning methods was analysed. Subsequently, the complexity of the problem is mitigated through the application of feature reduction and hyperparameter optimization techniques, which facilitate the amalgamation of databases for real-time prediction while refining algorithmic performance. This process culminates in the real-time estimation of traffic event times within a simplified framework, enabling swift and accurate predictions by leveraging optimized algorithms. In this context, feature selection was made with the GridSearchCV algorithm based on the model that gave the best result in the previous stage. In the last stage, with the decreasing number of variables and the real-time data set, the overall prediction of the traffic event duration was dynamically calculated with four different machine learning models. In this context, the performance of the model was compared with the situation in the second stage, and the results obtained from the effective use of resources were evaluated.</p>
<sec id="s3_1">
<label>3.1</label>
<title>Decision Trees (DT)</title>
<p>DT are a widely used algorithm for classification problem in data mining. They are used to represent classification and regression trees. The advantage of decision trees is their ease of creation and interpretation [<xref ref-type="bibr" rid="ref-46">46</xref>]. The algorithm resembles an upside-down tree, with a structure that extends from the root to sub-branches and new steps that multiply after the sub-branches. Each newly formed branch retains the characteristics of the main branch to which it is connected in the trunk. The data obtained from a selected column in the dataset is applied to the entire tree or dataset [<xref ref-type="bibr" rid="ref-47">47</xref>].</p>
<p>The DT is a classification method that generates a tree structure model composed of decision nodes and leaf nodes based on classification, feature, and target. The feature selection measure provides a ranking for each feature that defines the given training topics and determines which feature will be selected [<xref ref-type="bibr" rid="ref-47">47</xref>]. Measures such as information gain, gain ratio, and Gini index are commonly used in feature selection. Although it is possible to obtain multiple trees from a dataset, the tree with the smallest size is preferred. To terminate the iteration within the decision tree model during variable selection, it is requisite that all constituents within the node are assigned to a singular class. This condition that all elements in the leaves will be in the same class and there will be no values left to classify. Consequently, the iterative process within the decision tree model is halted, culminating in the finalization of the decision tree structure. DT has advantages such as simplicity, interpretability, and flexibility to work with categorical and numerical data. However, it has disadvantages such as bias, tendency to overfitting, and susceptibility to data imbalance.</p>
</sec>
<sec id="s3_2">
<label>3.2</label>
<title>Random Forest (RF)</title>
<p>The RF model is an advanced form of the bagging method used for both classification and regression. It is based on two parameters: the number of trees and the number of randomly selected independent variables in each node separation [<xref ref-type="bibr" rid="ref-48">48</xref>]. When creating decision trees, a sample is created using the bootstrap method by replacing as many samples as there are in the original dataset. The RF methodology represents a classification approach that harnesses the collective predictive power of multiple DT to enhance classification accuracy. Prediction outcomes are derived through a process of majority voting, wherein the aggregated predictions of all constituent trees within the forest are considered [<xref ref-type="bibr" rid="ref-3">3</xref>,<xref ref-type="bibr" rid="ref-49">49</xref>]. Critical features of the method include generalization error, parameter adjustment, distance between samples, data imputation, and variable importance, which measures the predictiveness of variables in the decision tree.</p>
<p>RF exhibits proficiency in managing extensive feature sets and demonstrates robust performance even in the presence of missing data. The learning algorithm affords flexibility in generating a predetermined number of trees through the utilization of the &#x201C;n_estimators&#x201D; parameter, or alternatively, by employing the &#x201C;random_state&#x201D; parameter to introduce randomness in tree selection [<xref ref-type="bibr" rid="ref-3">3</xref>]. In our research, no restrictions were imposed on the number of trees to be created, and no pruning was performed. RF has advantages such as high performance, requiring little hyperparameter tuning, resistance to overfitting, and the ability to handle multiple feature types. However, it has disadvantages, such as the increase in computational cost with the increase in the number of trees.</p>
</sec>
<sec id="s3_3">
<label>3.3</label>
<title>K-Nearest Neighbours (KNN)</title>
<p>The KNN algorithm is a non-parametric, memory-based learning classification algorithm. It memorizes training examples for prediction instead of learning a model. The algorithm classifies the k training points <inline-formula id="ieqn-1"><mml:math id="mml-ieqn-1"><mml:mrow><mml:mo>(</mml:mo><mml:mi>r</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mi>r</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mn>2</mml:mn><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo></mml:math></inline-formula>, <italic>k</italic> closest to a known query point <italic>x</italic>0 using majority voting among k neighbors. Similarity is defined as the distance between two data points, as measured by a metric [<xref ref-type="bibr" rid="ref-50">50</xref>]. Although the Euclidean distance is the most commonly used metric, other metrics such as Manhattan, Chebyshev, and Hamming distances have also been employed in studies.</p>
<p>The algorithm&#x2019;s speed is noteworthy as it does not rely on training data sets to make generalizations, keeping the entire training data set in memory during the testing phase. The size of the neighborhood is determined by the k parameter. Setting k to 1 results in low bias but high variance [<xref ref-type="bibr" rid="ref-51">51</xref>]. A k value of 1 means that predictions are made using the single training sample that is closest to the new patterns to be predicted. It is important to note that this prediction method relies heavily on the single closest training sample and may not be as accurate as methods that consider a larger number of training samples. The appropriate k value depends on the size of the data set. In this research, the Manhattan distance metric was selected as the distance parameter due to its capability to accommodate both continuous and categorical variables. While kNN offers advantages such as simplicity, effective classification performance, and easy integration of new data points, it also has disadvantages such as high computational cost, a curse of dimensionality, and sensitivity to noise in the dataset.</p>
</sec>
<sec id="s3_4">
<label>3.4</label>
<title>Support Vector Machine (SVM)</title>
<p>The SVM algorithm, originally conceived for binary classification tasks, has undergone adaptation to accommodate both multi-class classification and regression models. Initially constrained to the analysis of continuous variables, SVM has evolved to encompass the examination of categorical variables as well. This expansion entails the automatic conversion of categorical variables into numerical representations, thereby enabling the normalization of both categorical and continuous data within SVM frameworks. The algorithm aims to achieve the most suitable separation between the two classes on the plane. In cases where there are overlapping classes, basic approaches are used to reduce the effects on data points. These include reducing the discriminant margin and reflecting data points into a high-dimensional space using the kernel method, facilitating efficient linear separation. Additionally, the problem can be formulated as a second-order optimization problem at the solution point [<xref ref-type="bibr" rid="ref-52">52</xref>&#x2013;<xref ref-type="bibr" rid="ref-54">54</xref>].</p>
<p>SVM uses different parameters. The complexity parameter regulates the degree of flexibility exhibited by the decision boundary in class segregation. A setting of 0 mandates strict adherence to the margin, whereas the default value typically stands at 1. Additionally, a crucial parameter to consider is the selection of kernel function. The most elementary among these is the linear kernel, which delineates data instances using a linear decision boundary, often represented as a straight line or hyperplane. The polynomial kernel facilitates the separation of classes by employing a curved or nonlinear decision boundary, the degree of which is contingent upon the exponent value. The radial basis function kernel emerges as a widely adopted and potent alternative, leveraging intricate boundary shapes to effectively segregate classes [<xref ref-type="bibr" rid="ref-52">52</xref>&#x2013;<xref ref-type="bibr" rid="ref-54">54</xref>]. For this study, default parameters were used, including linear, polynomial, and RBF kernels. SVM has the advantages of good generalisation ability, adaptability to high dimensional data, and flexibility through various kernel functions. However, it has disadvantages such as high computational cost in large data sets and the need for careful tuning of hyperparameters.</p>
</sec>
<sec id="s3_5">
<label>3.5</label>
<title>GridSearchCV</title>
<p>GridSearchCV is a hyperparameter optimization method used to determine the best hyperparameters for machine learning algorithms. Exhaustively explores all combinations within a defined set of hyperparameters in pursuit of identifying the optimal values associated with superior performance. This method determines the optimal hyperparameter values by exhaustively testing all combinations in a predetermined set [<xref ref-type="bibr" rid="ref-55">55</xref>,<xref ref-type="bibr" rid="ref-56">56</xref>]. K-fold cross-validation is used for each hyperparameter value, and the results are recorded in a score matrix. The process tests all necessary combinations to obtain the best hyperparameter values [<xref ref-type="bibr" rid="ref-55">55</xref>,<xref ref-type="bibr" rid="ref-57">57</xref>].</p>
</sec>
<sec id="s3_6">
<label>3.6</label>
<title>SHAP</title>
<p>Shapley Additive Explanations (SHAP) is a method proposed by Lundberg and Lee in 2017 to evaluate how attributes affect the results. It is an explainability method developed to understand the complexity of machine learning models and explain their prediction results. It can be applied to a wide range of models, from tree-based models to deep learning models, and is an important tool in explainability research. SHAP offers a unique approach to understanding why the model makes a certain prediction and making the model&#x2019;s decisions more transparent [<xref ref-type="bibr" rid="ref-58">58</xref>].</p>
<p>The SHAP method is based on strong theoretical foundations, making it particularly useful in regulated contexts. It draws on principles from game theory and uses Shapley values to provide specific predictions by assigning importance values (SHAP values) to individual features. These SHAP values adhere to key properties: (1) local accuracy, ensuring the explanation model aligns closely with the original model&#x2019;s output; (2) missingness, where features absent in the original input have no discernible impact; and (3) consistency, ensuring that increasing dependence on a particular feature in model revisions does not diminish its importance, irrespective of other features [<xref ref-type="bibr" rid="ref-59">59</xref>].</p>
</sec>
<sec id="s3_7">
<label>3.7</label>
<title>Performance Matrix</title>
<p>Numerous criteria are employed to evaluate and compare the efficacy of machine learning algorithms. This study evaluates performance using accuracy, precision, recall, Matthews correlation coefficient, F1-score, jaccard and confusion matrix. The confusion matrix is a tool used in machine learning and statistics to measure the performance of a classification model. The confusion matrix is a 2 &#x00D7; 2 matrix that displays four different combinations between actual class and predicted class values: true positive (TP), true negative (TN), false positive (FP), and false negative (FN) [<xref ref-type="bibr" rid="ref-3">3</xref>].</p>
<p>Accuracy shows the percentage of samples classified correctly. The accuracy value ranges from 0 to 1.
<disp-formula id="eqn-1"><label>(1)</label><mml:math id="mml-eqn-1" display="block"><mml:mi>A</mml:mi><mml:mi>c</mml:mi><mml:mi>c</mml:mi><mml:mi>u</mml:mi><mml:mi>r</mml:mi><mml:mi>a</mml:mi><mml:mi>c</mml:mi><mml:mi>y</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>T</mml:mi><mml:mi>P</mml:mi><mml:mo>+</mml:mo><mml:mi>T</mml:mi><mml:mi>N</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>T</mml:mi><mml:mi>P</mml:mi><mml:mo>+</mml:mo><mml:mi>T</mml:mi><mml:mi>N</mml:mi><mml:mo>+</mml:mo><mml:mi>F</mml:mi><mml:mi>P</mml:mi><mml:mo>+</mml:mo><mml:mi>F</mml:mi><mml:mi>N</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mfrac></mml:math></disp-formula></p>
<p>Recall indicates the proportion of actual instances of a class that were correctly classified.
<disp-formula id="eqn-2"><label>(2)</label><mml:math id="mml-eqn-2" display="block"><mml:mi>R</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi><mml:mi>l</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>T</mml:mi><mml:mi>P</mml:mi><mml:mo>+</mml:mo><mml:mi>F</mml:mi><mml:mi>N</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mfrac></mml:math></disp-formula></p>
<p>Precision represents the percentage of samples classified with true labels of a class.
<disp-formula id="eqn-3"><label>(3)</label><mml:math id="mml-eqn-3" display="block"><mml:mi>P</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mi>i</mml:mi><mml:mi>s</mml:mi><mml:mi>i</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>T</mml:mi><mml:mi>P</mml:mi><mml:mo>+</mml:mo><mml:mi>F</mml:mi><mml:mi>P</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mfrac></mml:math></disp-formula></p>
<p>F1-score is a measure of the balance between precision and sensitivity, calculated as the weighted average.
<disp-formula id="eqn-4"><label>(4)</label><mml:math id="mml-eqn-4" display="block"><mml:mi>F</mml:mi><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:mi>s</mml:mi><mml:mi>c</mml:mi><mml:mi>o</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mo>=</mml:mo><mml:mn>2</mml:mn><mml:mo>&#x2217;</mml:mo><mml:mfrac><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>P</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mi>i</mml:mi><mml:mi>s</mml:mi><mml:mi>i</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:mo>&#x2217;</mml:mo><mml:mi>R</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi><mml:mi>l</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>P</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mi>i</mml:mi><mml:mi>s</mml:mi><mml:mi>i</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:mo>+</mml:mo><mml:mi>R</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi><mml:mi>l</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mfrac></mml:math></disp-formula></p>
<p>The Matthews correlation coefficient (MCC) is a preferred criterion for evaluating the performance of a classification model or function with two or more classifications. The coefficient takes values between &#x2212;1 and 1, where &#x2212;1 indicates reverse classification, 0 indicates average classification performance or poor performance, and 1 indicates perfect classification or prediction [<xref ref-type="bibr" rid="ref-3">3</xref>,<xref ref-type="bibr" rid="ref-60">60</xref>].
<disp-formula id="eqn-5"><label>(5)</label><mml:math id="mml-eqn-5" display="block"><mml:mi>M</mml:mi><mml:mi>C</mml:mi><mml:mi>C</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>T</mml:mi><mml:mi>P</mml:mi><mml:mo>&#x2217;</mml:mo><mml:mi>T</mml:mi><mml:mi>N</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mi>F</mml:mi><mml:mi>P</mml:mi><mml:mo>&#x2217;</mml:mo><mml:mi>F</mml:mi><mml:mi>N</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:msqrt><mml:mo stretchy="false">(</mml:mo><mml:mi>T</mml:mi><mml:mi>P</mml:mi><mml:mo>+</mml:mo><mml:mi>F</mml:mi><mml:mi>P</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:mi>T</mml:mi><mml:mi>P</mml:mi><mml:mo>+</mml:mo><mml:mi>F</mml:mi><mml:mi>N</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:mi>T</mml:mi><mml:mi>N</mml:mi><mml:mo>+</mml:mo><mml:mi>F</mml:mi><mml:mi>P</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:mi>T</mml:mi><mml:mi>N</mml:mi><mml:mo>+</mml:mo><mml:mi>F</mml:mi><mml:mi>N</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:msqrt></mml:mfrac></mml:math></disp-formula></p>
</sec>
</sec>
<sec id="s4">
<label>4</label>
<title>Experiment and Results</title>
<sec id="s4_1">
<label>4.1</label>
<title>Data</title>
<p>The dataset used in the training process is one of the most important elements of machine learning models. The quality of machine learning models&#x2019; predictions is directly dependent on the quality of the training data. Therefore, it is crucial to meticulously collect data that accurately represents the classes targeted by the model. The data set on traffic incidents used in our study was obtained from official institutions, including the Turkish Statistical Institute (TUIK), Istanbul Metropolitan Municipality, the 1st Regional Directorate of Meteorology, and open data portals such as Google Maps, ArcGIS, OpenStreetMap Development Library, and IMM Open Data Portal. The database contains a total of 24,431 records of traffic accidents and incidents, including vehicle breakdowns and fires, that occurred in Istanbul between 2020 and 2021. After collecting the data for analysis, we examined all variables that could impact the prediction. These variables include time, location, vehicle, traffic index, speed, road structure and condition, meteorology, social and demographic factors, and district. <xref ref-type="table" rid="table-1">Table 1</xref> provides information on the 75 variables and their data types.</p>
<table-wrap id="table-1">
<label>Table 1</label>
<caption>
<title>Data set</title>
</caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th>Variables</th>
<th>Data type</th>
<th>Min&#x2013;Max</th>
<th>Variables</th>
<th>Data type</th>
<th>Min&#x2013;Max</th>
</tr>
</thead>
<tbody>
<tr>
<td>Year</td>
<td>Numeric</td>
<td>2020&#x2013;2021</td>
<td>Van</td>
<td>Numeric</td>
<td>676,470&#x2013;719,503</td>
</tr>
<tr>
<td>Month</td>
<td>Categorical</td>
<td>1&#x2013;12</td>
<td>Truck</td>
<td>Numeric</td>
<td>133,058&#x2013;136,741</td>
</tr>
<tr>
<td>Day</td>
<td>Categorical</td>
<td>1&#x2013;31</td>
<td>Motorcycle</td>
<td>Numeric</td>
<td>334,658&#x2013;392,207</td>
</tr>
<tr>
<td>Special day</td>
<td>Categorical</td>
<td>1&#x2013;4</td>
<td>Special purpose vehicle</td>
<td>Numeric</td>
<td>8391&#x2013;9673</td>
</tr>
<tr>
<td>Incident day</td>
<td>Categorical</td>
<td>1&#x2013;7</td>
<td>Total number of vehicles</td>
<td>Numeric</td>
<td>4,206,858&#x2013;4,520,493</td>
</tr>
<tr>
<td>Time period</td>
<td>Categorical</td>
<td>1&#x2013;12</td>
<td>Bachelor&#x2019;s degree rate</td>
<td>Numeric</td>
<td>8&#x2013;47</td>
</tr>
<tr>
<td>District</td>
<td>Categorical</td>
<td>1&#x2013;38</td>
<td>Illiteracy rate</td>
<td>Numeric</td>
<td>0.63&#x2013;2.75</td>
</tr>
<tr>
<td>GeoHash</td>
<td>Categorical</td>
<td>1&#x2013;682</td>
<td>Student rate</td>
<td>Numeric</td>
<td>0.46&#x2013;5.77</td>
</tr>
<tr>
<td>District population</td>
<td>Numeric</td>
<td>74,945&#x2013;957,398</td>
<td>Average household size</td>
<td>Numeric</td>
<td>2.35&#x2013;4.13</td>
</tr>
<tr>
<td>Number of Neighborhoods</td>
<td>Numeric</td>
<td>10&#x2013;57</td>
<td>Number of houses</td>
<td>Numeric</td>
<td>43,045&#x2013;428,810</td>
</tr>
<tr>
<td>Area measurement</td>
<td>Numeric</td>
<td>7.17&#x2013;1040.42</td>
<td>Number of private workplaces</td>
<td>Numeric</td>
<td>5675&#x2013;109,383</td>
</tr>
<tr>
<td>Minimum speed</td>
<td>Numeric</td>
<td>1&#x2013;70</td>
<td>Agricultural field</td>
<td>Numeric</td>
<td>0&#x2013;415,346</td>
</tr>
<tr>
<td>Maximum speed</td>
<td>Numeric</td>
<td>19&#x2013;211</td>
<td>Number of hospitals</td>
<td>Numeric</td>
<td>20&#x2013;111</td>
</tr>
<tr>
<td>Average speed</td>
<td>Numeric</td>
<td>10&#x2013;265</td>
<td>Number of schools</td>
<td>Numeric</td>
<td>67&#x2013;284</td>
</tr>
<tr>
<td>Number of unique vehicles</td>
<td>Numeric</td>
<td>11&#x2013;1247</td>
<td>University</td>
<td>Numeric</td>
<td>0&#x2013;11</td>
</tr>
<tr>
<td>Min traffic index</td>
<td>Numeric</td>
<td>1&#x2013;17</td>
<td>University facility</td>
<td>Numeric</td>
<td>0&#x2013;24</td>
</tr>
<tr>
<td>Maximum traffic index</td>
<td>Numeric</td>
<td>9&#x2013;255</td>
<td>Police</td>
<td>Numeric</td>
<td>41,018&#x2013;456,861</td>
</tr>
<tr>
<td>Average traffic index</td>
<td>Numeric</td>
<td>2&#x2013;105</td>
<td>Fire station</td>
<td>Numeric</td>
<td>0&#x2013;152</td>
</tr>
<tr>
<td>Number of vehicles per day</td>
<td>Numeric</td>
<td>32&#x2013;251,198</td>
<td>personSOS</td>
<td>Numeric</td>
<td>18,743&#x2013;456,860</td>
</tr>
<tr>
<td>Daily average speed</td>
<td>Numeric</td>
<td>4&#x2013;120</td>
<td>Metrobus station</td>
<td>Categorical</td>
<td>0&#x2013;1</td>
</tr>
<tr>
<td>Traffic percentage</td>
<td>Numeric</td>
<td>1&#x2013;90</td>
<td>Metro station</td>
<td>Categorical</td>
<td>0&#x2013;1</td>
</tr>
<tr>
<td>Temperature</td>
<td>Numeric</td>
<td>(&#x2212;4.3)&#x2013;35.7</td>
<td>Port</td>
<td>Numeric</td>
<td>0&#x2013;1</td>
</tr>
<tr>
<td>Road temperature</td>
<td>Numeric</td>
<td>(&#x2212;99)&#x2013;57.9</td>
<td>Number of parking lots</td>
<td>Numeric</td>
<td>8&#x2013;221</td>
</tr>
<tr>
<td>Humidity</td>
<td>Numeric</td>
<td>3&#x2013;100</td>
<td>Number of banks</td>
<td>Numeric</td>
<td>11&#x2013;253</td>
</tr>
<tr>
<td>Rainfall amount</td>
<td>Numeric</td>
<td>0&#x2013;27.6</td>
<td>Number of ATMs</td>
<td>Numeric</td>
<td>40&#x2013;586</td>
</tr>
<tr>
<td>Wind speed</td>
<td>Numeric</td>
<td>0&#x2013;22.4</td>
<td>Number of shopping malls</td>
<td>Numeric</td>
<td>0&#x2013;16</td>
</tr>
<tr>
<td>Wind direction</td>
<td>Numeric</td>
<td>0&#x2013;360</td>
<td>Number of markets</td>
<td>Numeric</td>
<td>0&#x2013;18</td>
</tr>
<tr>
<td>Ground information</td>
<td>Categorical</td>
<td>1&#x2013;3</td>
<td>Number of mini markets</td>
<td>Numeric</td>
<td>3&#x2013;136</td>
</tr>
<tr>
<td>Road type-1</td>
<td>Categorical</td>
<td>1&#x2013;17</td>
<td>Number of super markets</td>
<td>Numeric</td>
<td>2&#x2013;78</td>
</tr>
<tr>
<td>Road type-2</td>
<td>Categorical</td>
<td>2&#x2013;3</td>
<td>Number of hotels</td>
<td>Numeric</td>
<td>0&#x2013;980</td>
</tr>
<tr>
<td>Number of lanes</td>
<td>Categorical</td>
<td>1&#x2013;12</td>
<td>Number of stores</td>
<td>Numeric</td>
<td>38&#x2013;5549</td>
</tr>
<tr>
<td>Divided road</td>
<td>Categorical</td>
<td>1&#x2013;3</td>
<td>Industrial area</td>
<td>Numeric</td>
<td>0&#x2013;540</td>
</tr>
<tr>
<td>Speed</td>
<td>Numeric</td>
<td>0&#x2013;100</td>
<td>Number of bars/clubs</td>
<td>Numeric</td>
<td>0&#x2013;94</td>
</tr>
<tr>
<td>Width</td>
<td>Numeric</td>
<td>0&#x2013;65</td>
<td>Number of cafes</td>
<td>Numeric</td>
<td>9&#x2013;674</td>
</tr>
<tr>
<td>One way</td>
<td>Categorical</td>
<td>1&#x2013;4</td>
<td>Number of museum galleries</td>
<td>Numeric</td>
<td>0&#x2013;52</td>
</tr>
<tr>
<td>Car</td>
<td>Numeric</td>
<td>2,889,938&#x2013;3,100,848</td>
<td>Sports facility</td>
<td>Numeric</td>
<td>7&#x2013;248</td>
</tr>
<tr>
<td>Minibus</td>
<td>Numeric</td>
<td>96,608&#x2013;97,896</td>
<td rowspan="2">Number of theaters</td>
<td rowspan="2">Numeric</td>
<td rowspan="2">0&#x2013;73</td>
</tr>
<tr>
<td>Bus</td>
<td>Numeric</td>
<td>38,561&#x2013;40,784</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s4_2">
<label>4.2</label>
<title>Study Area</title>
<p>Istanbul was chosen for this study because of its cosmopolitan structure and its complex transportation network. The duration of traffic incidents was estimated by considering all locations in the database. The study obtained coordinate-based information on accidents and incidents in Istanbul and used ArcGIS to assign 6th geohash codes to these coordinates. Geohash is a coding system that converts geolocation data into a string and uses it to express the latitude and longitude coordinates of a location in an abbreviated format. Geohash defines a rectangular cell and divides location data into cells. In this way, we attempted to estimate traffic incident duration by performing spatial modeling with geohash areas.</p>
<p>The traffic incident time refers to the moment when a traffic incident takes place. In our study, we divided the 24-h day into 2-h periods, as shown in <xref ref-type="table" rid="table-2">Table 2</xref>, and assigned the start times of traffic events to these segments. This approach aims to detect hidden patterns between events in the same time period and to more accurately predict the duration of a traffic event that may occur during a specific time period.</p>
<table-wrap id="table-2">
<label>Table 2</label>
<caption>
<title>Traffic event time</title>
</caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th>Time range ID</th>
<th>Time range incident occured</th>
<th>Time range ID</th>
<th>Time range incident occured</th>
<th>Time range ID</th>
<th>Time range incident occured</th>
</tr>
</thead>
<tbody>
<tr>
<td>1</td>
<td>00:00&#x2013;01:59</td>
<td>5</td>
<td>08:00&#x2013;09:59</td>
<td>9</td>
<td>16:00&#x2013;17:59</td>
</tr>
<tr>
<td>2</td>
<td>02:00&#x2013;03:59</td>
<td>6</td>
<td>10:00&#x2013;11:59</td>
<td>10</td>
<td>18:00&#x2013;19:59</td>
</tr>
<tr>
<td>3</td>
<td>04:00&#x2013;05:59</td>
<td>7</td>
<td>12:00&#x2013;13:59</td>
<td>11</td>
<td>20:00&#x2013;21:59</td>
</tr>
<tr>
<td>4</td>
<td>06:00&#x2013;07:59</td>
<td>8</td>
<td>14:00&#x2013;15:59</td>
<td>12</td>
<td>22:00&#x2013;23:59</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>The literature has established various time intervals for traffic incident duration. To ensure clarity, in our study, four classifications have been made. This is because incidents that last longer than 90 min are rare in Istanbul, while those that last less than 10 min occur more frequently. Simple incidents are those that do not exceed 10 min and typically involve minor vehicle malfunctions or short stoppages. These types of incidents do not require TIM intervention. When examining events that require intervention, it is evident that their duration is between 23 and 31 min at most. <xref ref-type="table" rid="table-3">Table 3</xref> presents four categories related to the duration of traffic event.</p>
<table-wrap id="table-3">
<label>Table 3</label>
<caption>
<title>Traffic event duration</title>
</caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th>Traffic event type</th>
<th>Event type ID</th>
<th>Duration (min)</th>
</tr>
</thead>
<tbody>
<tr>
<td>Simple</td>
<td>0</td>
<td>0&#x2013;10</td>
</tr>
<tr>
<td>Minor</td>
<td>1</td>
<td>11&#x2013;30</td>
</tr>
<tr>
<td>Mid-level</td>
<td>2</td>
<td>31&#x2013;60</td>
</tr>
<tr>
<td>Major</td>
<td>3</td>
<td>61 and over</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s4_3">
<label>4.3</label>
<title>Setup for the Experiment</title>
<p>The study began with data preprocessing activities. These activities involved combining data from different databases, improving incomplete and noisy data, and obtaining a structured dataset for analysis. To estimate the duration of traffic incidents, the Scikit-learn library was used in the Python program. Four different machine learning models were used: DT, RF, KNN, SVM. The performance of the classification algorithms was then measured. Feature selection was performed, followed by hyperparameter optimization through feature reduction. GridSearchCV was utilized to find the optimal hyperparameter values by testing all combinations within a specified set of hyperparameters. After verifying the feasibility of reduced features, we conducted hyperparameter optimization to evaluate the accuracy of predictions using a real-time database. Experiments were conducted in the Anaconda3 2021.05 environment. A computer with Intel i5 processor, 2.4 GHz, 16 GB RAM and Windows 10 64-bit operating system was used for the experiments.</p>
</sec>
<sec id="s4_4">
<label>4.4</label>
<title>Results and Discussions</title>
<p>The experimental study yielded the best results for four different machine learning algorithms when 25% of the test data sets and 75% of the training data sets were used. The performance metrics of the DT, RF, KNN, and SVM models were calculated in terms of accuracy, recall, precision, and F1-score. <xref ref-type="table" rid="table-4">Table 4</xref> presents the performance of the four models applied.</p>
<table-wrap id="table-4">
<label>Table 4</label>
<caption>
<title>Model test results</title>
</caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th>Performance metrics</th>
<th>DT</th>
<th>RF</th>
<th>KNN</th>
<th>SVM</th>
</tr>
</thead>
<tbody>
<tr>
<td>Accuracy</td>
<td>0.825</td>
<td>0.981</td>
<td>0.931</td>
<td>0.862</td>
</tr>
<tr>
<td>Balanced accuracy</td>
<td>0.94</td>
<td>0.988</td>
<td>0.966</td>
<td>0.865</td>
</tr>
<tr>
<td>Precision micro</td>
<td>0.825</td>
<td>0.981</td>
<td>0.931</td>
<td>0.861</td>
</tr>
<tr>
<td>Precision macro</td>
<td>0.596</td>
<td>0.940</td>
<td>0.847</td>
<td>0.711</td>
</tr>
<tr>
<td>Recall micro</td>
<td>0.825</td>
<td>0.981</td>
<td>0.931</td>
<td>0.861</td>
</tr>
<tr>
<td>Recall macro</td>
<td>0.941</td>
<td>0.988</td>
<td>0.966</td>
<td>0.659</td>
</tr>
<tr>
<td>F1 micro</td>
<td>0.825</td>
<td>0.981</td>
<td>0.931</td>
<td>0.861</td>
</tr>
<tr>
<td>F1 macro</td>
<td>0.693</td>
<td>0.963</td>
<td>0.896</td>
<td>0.469</td>
</tr>
<tr>
<td>Model processing time (min)</td>
<td>5.110</td>
<td>8.980</td>
<td>12.950</td>
<td>61.450</td>
</tr>
</tbody>
</table>
</table-wrap>
<p><xref ref-type="table" rid="table-4">Table 4</xref> displays the accuracy rates for predicting traffic incident duration, with RF achieving the highest rate of 98%, followed closely by other machine learning models such as KNN, SVM, and DT. In terms of precision, both RF and KNN exhibit high precision on both micro and macro scales, while DT shows slightly lower performance. At recall value, both RF and KNN exhibit high success rates on both micro and macro scales. Additionally, both models have high F1-scores on both scales. The fact that the KNN model has the second highest accuracy rate may be due to its ability to effectively recognise similar examples in the data set. The DT model is that they tend to overfit when a single tree is used. The reason why SVM has a low accuracy rate compared to other models may be due to factors such as high computational costs and class imbalances in the data set. While the DT model works with the lowest time, the running times of the KNN and RF models are at an ideal level, and SVM has the highest time values. The unbalanced distribution of traffic events by duration for four traffic event duration classes was a situation that prevented the training dataset from overlearning in the RF model. RF generates a more robust prediction by combining many decision trees together using an ensemble approach. RF requires less precise hyperparameter tuning, making it less susceptible to data instability and noise, resulting in more consistent results.</p>

<p>In our study, ten experiments were conducted in which the data was randomly selected to reduce the randomness of the model due to the selection of training and test samples. Eight evaluation metrics were obtained for each experiment. In order to determine whether there was a significant difference between the DT, RF, KNN, and SVM algorithms used in the study, a normality test was first performed in order to ascertain whether the data was normally distributed. Once it was determined that the data were not normally distributed, the Kruskal-Wallis test with 5% significance level was applied to ascertain whether there was a significant difference in the performance of the methods. The Kruskal-Wallis test evaluates the significance of differences in population medians for a dependent variable across all factor levels. While the null hypothesis (H0) states that there is no significant difference between the methods, the alternative hypothesis (H1) states that there is a significant difference between the methods. Upon examination of <xref ref-type="table" rid="table-5">Table 5</xref>, it can be seen that the value of Sig. (0.000) is less than 0.05, indicating that the null hypothesis is rejected. Results implies that there is a significant difference between the four methods in all performance metrics.</p>
<table-wrap id="table-5">
<label>Table 5</label>
<caption>
<title>The Kruskal-Wallis test results at 95% significance level</title>
</caption>
<table frame="hsides">
<colgroup>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
</colgroup>
<thead>
<tr>
<th></th>
<th align="center" colspan="16">Performance metrics</th>
</tr>
<tr>
<th></th>
<th align="center" colspan="4"><underline>Accuracy</underline></th>
<th align="center" colspan="4"><underline>Balanced accuracy</underline></th>
<th align="center" colspan="4"><underline>Precision micro</underline></th>
<th align="center" colspan="4"><underline>Precision macro</underline></th>
</tr>
<tr>
<td>Methods</td>
<td>DT</td>
<td>RF</td>
<td>KNN</td>
<td>SVM</td>
<td>DT</td>
<td>RF</td>
<td>KNN</td>
<td>SVM</td>
<td>DT</td>
<td>RF</td>
<td>KNN</td>
<td>SVM</td>
<td>DT</td>
<td>RF</td>
<td>KNN</td>
<td>SVM</td>
</tr>
</thead>
<tbody>
<tr>
<td>Mean rank</td>
<td>5.5</td>
<td>35.5</td>
<td>25.5</td>
<td>15.5</td>
<td>15.5</td>
<td>35.5</td>
<td>25.5</td>
<td>5.5</td>
<td>5.5</td>
<td>35.5</td>
<td>25.5</td>
<td>15.5</td>
<td>5.5</td>
<td>35.5</td>
<td>25.5</td>
<td>15.5</td>
</tr>
<tr>
<td>Kruskal-Wallis Ist.</td>
<td align="center" colspan="4">36.963</td>
<td align="center" colspan="4">37.051</td>
<td align="center" colspan="4">36.984</td>
<td align="center" colspan="4">37.182</td>
</tr>
<tr>
<td>P (Asymp. Sig.)</td>
<td align="center" colspan="4">0</td>
<td align="center" colspan="4">0</td>
<td align="center" colspan="4">0</td>
<td align="center" colspan="4">0</td>
</tr>
<tr>
<td></td>
<td align="center" colspan="4">Recall micro</td>
<td align="center" colspan="4">Recall macro</td>
<td align="center" colspan="4">F1 micro</td>
<td align="center" colspan="4">F1 macro</td>
</tr>
<tr>
<td>Methods</td>
<td>DT</td>
<td>RF</td>
<td>KNN</td>
<td>SVM</td>
<td>DT</td>
<td>RF</td>
<td>KNN</td>
<td>SVM</td>
<td>DT</td>
<td>RF</td>
<td>KNN</td>
<td>SVM</td>
<td>DT</td>
<td>RF</td>
<td>KNN</td>
<td>SVM</td>
</tr>
<tr>
<td>Mean rank</td>
<td>5.5</td>
<td>35.5</td>
<td>25.5</td>
<td>15.5</td>
<td>15.5</td>
<td>35.5</td>
<td>25.5</td>
<td>5.5</td>
<td>5.5</td>
<td>35.5</td>
<td>25.5</td>
<td>15.5</td>
<td>15.5</td>
<td>35.5</td>
<td>25.5</td>
<td>5.5</td>
</tr>
<tr>
<td>Kruskal-Wallis Ist.</td>
<td align="center" colspan="4">36.984</td>
<td align="center" colspan="4">37.041</td>
<td align="center" colspan="4">36.956</td>
<td align="center" colspan="4">36.838</td>
</tr>
<tr>
<td>P (Asymp. Sig.)</td>
<td align="center" colspan="4">0</td>
<td align="center" colspan="4">0</td>
<td align="center" colspan="4">0</td>
<td align="center" colspan="4">0</td>
</tr>
</tbody>
</table>
<table-wrap-foot><fn>
<p>Note: &#x002A;H0 represents a hypothesis based on no significant difference between machine learning algorihms.</p>
</fn>
</table-wrap-foot>
</table-wrap>
<p>The analysis revealed a correct classification rate of 98%, with a 2% error. However, incorrect predictions were made in some cases, leading to significant differences in performance metrics between classifications. These findings will serve as a valuable reference for future studies in the field. To improve the results, we removed unnecessary variables. <xref ref-type="table" rid="table-4">Table 4</xref> presents the results obtained from 75 variables using the GridSearchCV algorithm for hyperparameter optimization. The Scikit-learn library was used for optimization, and the GridSearchCV algorithm was employed to cross-validate the models and search for the best parameters. It is important to note that the performance of the models is affected differently by various parameters. Finding the optimal value for each parameter can be computationally expensive. Currently, we have analyzed the parameters that have a greater impact on the outputs. Please refer to <xref ref-type="table" rid="table-6">Table 6</xref> for the definition of the model parameters. The feature selection process reduced the number of variables from 75 to 14. <xref ref-type="table" rid="table-7">Table 7</xref> displays the order and weights of the variables selected for the traffic incident duration.</p>
<table-wrap id="table-6">
<label>Table 6</label>
<caption>
<title>GridSearchCV parameters</title>
</caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th>Hyperparameter</th>
<th>Value</th>
</tr>
</thead>
<tbody>
<tr>
<td>n_estimators</td>
<td>100, 200, 300, 400</td>
</tr>
<tr>
<td>n_splits</td>
<td>5</td>
</tr>
<tr>
<td>shuffle</td>
<td>True</td>
</tr>
<tr>
<td>random_state</td>
<td>3</td>
</tr>
<tr>
<td>verbose</td>
<td>0</td>
</tr>
</tbody>
</table>
</table-wrap><table-wrap id="table-7">
<label>Table 7</label>
<caption>
<title>Variables used for duration prediction after feature reduction</title>
</caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th>Variable</th>
<th>Sort</th>
<th>Weight</th>
<th>Explanation</th>
</tr>
</thead>
<tbody>
<tr>
<td>Number of unique vehicles</td>
<td>1</td>
<td>0.0457</td>
<td>The number of different vehicles in the relevant geohash area in the given hour.</td>
</tr>
<tr>
<td>Number of vehicles per day</td>
<td>2</td>
<td>0.0445</td>
<td>The number of vehicles passing daily in the relevant geohash area.</td>
</tr>
<tr>
<td>Temperature</td>
<td>3</td>
<td>0.0433</td>
<td>The air temperature in the relevant time zone.</td>
</tr>
<tr>
<td>Wind direction</td>
<td>4</td>
<td>0.0427</td>
<td>Wind direction in the relevant time zone.</td>
</tr>
<tr>
<td>Maximum speed</td>
<td>5</td>
<td>0.0411</td>
<td>Maximum speed within the relevant geohash area at the given time.</td>
</tr>
<tr>
<td>Average speed</td>
<td>6</td>
<td>0.0399</td>
<td>Average speed of vehicles within the relevant geohash area at the given time.</td>
</tr>
<tr>
<td>Wind speed</td>
<td>7</td>
<td>0.0398</td>
<td>Wind speed in the relevant time zone.</td>
</tr>
<tr>
<td>Humidity</td>
<td>8</td>
<td>0.0388</td>
<td>Humidity in the relevant time zone.</td>
</tr>
<tr>
<td>General traffic percentage</td>
<td>9</td>
<td>0.0382</td>
<td>The overall traffic percentage of all geohashes at that hour.</td>
</tr>
<tr>
<td>Time period</td>
<td>10</td>
<td>0.0381</td>
<td>Relevant time period range.</td>
</tr>
<tr>
<td>GeoHash</td>
<td>11</td>
<td>0.0377</td>
<td>Geohash Value of Latitudes and Longitudes. <italic>Geohash length is 6</italic>.</td>
</tr>
<tr>
<td>Incident day</td>
<td>12</td>
<td>0.0355</td>
<td>Day of the week.</td>
</tr>
<tr>
<td>Day</td>
<td>13</td>
<td>0.0344</td>
<td>Day of the month.</td>
</tr>
<tr>
<td>Month</td>
<td>14</td>
<td>0.0283</td>
<td>Relevant month.</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Reducing the number of features from 75 to 14 was a successful outcome in the feature selection phase of the study. This demonstrates the ability of the ML models used to select the most important variables and reduce the complexity of the model, resulting in improved prediction performance. Fewer features allow for more effective predictions with fewer variables. The reduction of features from 75 to 14 indicates that the study&#x2019;s methodology achieved successful and efficient feature selection. The performances of the algorithms in experiments with 14 variables are given in <xref ref-type="table" rid="table-8">Table 8</xref>. Although the number of variables decreased, there were only minor changes in accuracy rates.</p>
<table-wrap id="table-8">
<label>Table 8</label>
<caption>
<title>Model performances after reducing features</title>
</caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th>Performance metrics</th>
<th>DT</th>
<th>RF</th>
<th>KNN</th>
<th>SVM</th>
</tr>
</thead>
<tbody>
<tr>
<td>Accuracy</td>
<td>0.816</td>
<td>0.985</td>
<td>0.926</td>
<td>0.851</td>
</tr>
<tr>
<td>Balanced accuracy</td>
<td>0.935</td>
<td>0.986</td>
<td>0.962</td>
<td>0.855</td>
</tr>
<tr>
<td>Precision micro</td>
<td>0.816</td>
<td>0.985</td>
<td>0.926</td>
<td>0.851</td>
</tr>
<tr>
<td>Precision macro</td>
<td>0.604</td>
<td>0.986</td>
<td>0.823</td>
<td>0.649</td>
</tr>
<tr>
<td>Recall micro</td>
<td>0.816</td>
<td>0.985</td>
<td>0.926</td>
<td>0.851</td>
</tr>
<tr>
<td>Recall macro</td>
<td>0.935</td>
<td>0.986</td>
<td>0.962</td>
<td>0.403</td>
</tr>
<tr>
<td>F1 micro</td>
<td>0.816</td>
<td>0.985</td>
<td>0.926</td>
<td>0.851</td>
</tr>
<tr>
<td>F1 macro</td>
<td>0.698</td>
<td>0.968</td>
<td>0.888</td>
<td>0.457</td>
</tr>
<tr>
<td>MCC</td>
<td>0.471</td>
<td>0.864</td>
<td>0.837</td>
<td>0.523</td>
</tr>
<tr>
<td>Model processing time (min)</td>
<td>3.110</td>
<td>4.980</td>
<td>7.450</td>
<td>28.75</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>The RF model was found to be the best model, with an accuracy rate increase from 0.981 to 0.985. It has been observed that model processing times are lower with feature reduction. While the RF model&#x2019;s pre-feature reduction time was 12.95 min, the post-feature reduction time decreased to 7.45 min. This shows that feature reduction increases the speed of the classification process and creates a more efficient and dynamic model. The MCC value of the RF model is 0.864, which means that the model performs very well. This indicates that the model is very good at making correct predictions and has few false predictions. A high value of MCC indicates that the model correctly distinguishes between positive and negative classes. The KNN model also performs well. SVM and DT models show average performance, so these two models have lower performance than the others.</p>
<p><xref ref-type="fig" rid="fig-2">Fig. 2</xref> shows the model&#x2019;s accuracy results with 75 variables and the accuracy results without using 61 variables with feature reduction. In this context, it was observed that parallel results were obtained for 14 important variables. Since the feature reduction was best achieved with RF, a small improvement was observed, while the accuracy rates of other models showed a small decrease.</p>
<fig id="fig-2">
<label>Figure 2</label>
<caption>
<title>Model results</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_52323-fig-2.tif"/>
</fig>
<p><xref ref-type="fig" rid="fig-3">Fig. 3</xref> displays the complexity matrix of the RF model, which produced the most successful results. The confusion matrix in <xref ref-type="fig" rid="fig-3">Fig. 3</xref> includes four different time classes of incident durations. Class 0 represents the intervention time between 0 and 10 min, as shown in <xref ref-type="table" rid="table-3">Table 3</xref>. Class 1 represents a time interval of 11&#x2013;30 min, and Class 2 represents a time interval of 31&#x2013;60 min. Class 3 includes events that last more than 61 min. There were 5200 traffic incidents in Class 0 and only 41 in Class 1. First class constitutes 85% of all accident classes. An unbalanced class distribution may cause the model to over-learn. However, the results show that overlearning did not occur. The confusion matrix indicates that the model made correct predictions in 40 out of 41 predictions in the 0th class, 563 out of 566 predictions in the 2nd class, and 322 out of 326 predictions in the 3rd class. The model distinguished between clusters in the data set with an unbalanced distribution. It accurately predicted situations that required emergency intervention and those that did not. The model&#x2019;s success in classes 0 and 3 indicates its high accuracy. Other metrics also demonstrate the model&#x2019;s overall success in clearly separating classes without excessive learning.</p>
<fig id="fig-3">
<label>Figure 3</label>
<caption>
<title>RF confusion matrix</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_52323-fig-3.tif"/>
</fig>
<p>The SHAP method is a mathematical approach based on game theory concepts that is used to explain the predictions of machine learning models. In our study, we used the SHAP method to calculate the contribution of 14 features to the prediction. As a result, we were able to reveal the effects of each feature on the prediction of different classes after the feature reduction phase. <xref ref-type="fig" rid="fig-4">Fig. 4</xref> illustrates the influencing rates of 14 independent variables used to predict the duration class of the dependent variable. The time period variable has the most weight in prediction compared to other variables. This means that the time of day when the accident occurs has the greatest impact on the intervention time. Heavy traffic in the city during the daytime is expected to be a priority due to the increased risk of accidents at the beginning and end of work. Response time is also expected to be significantly affected by traffic.</p>
<fig id="fig-4">
<label>Figure 4</label>
<caption>
<title>SHAP values</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_52323-fig-4.tif"/>
</fig>
<p>The &#x201C;GeoHash&#x201D; attribute is one of the primary variables for all four different classes. While this variable is of secondary importance after the time period in the estimation of 0, 2, and 3 classes, it is in the fourth place after the &#x201C;Temperature&#x201D; and &#x201C;Day&#x201D; attributes in the estimation of the 11&#x2013;30 min event durations of Class 1. &#x201C;GeoHash&#x201D; provides location information where the accident occurred. It is important to note that the distribution of traffic and accidents on the roads varies due to the non-homogeneous population density in Istanbul. Knowing where the accident took place is crucial information for crime scene intervention, following the time of the accident. Temperature is ranked third in importance for Class 0, second for Class 1, sixth for Class 2, and fifth for Class 3. The study&#x2019;s use of different time periods and variable weights contributed to the prediction&#x2019;s accuracy.</p>
</sec>
</sec>
<sec id="s5">
<label>5</label>
<title>Conclusion</title>
<p>This study presents a machine learning-based model for forecasting traffic event duration, integrating feature selection and dynamic modeling. Through comprehensive testing, the model reduced 75 variables to 14 significant ones, enabling predictions prior to incidents. Real-time database structuring facilitates dynamic forecasting. The performance evaluations conducted with machine learning algorithms, including DT, RF, KNN, and SVM, revealed that the RF model achieved the highest accuracy and balanced accuracy rates. The model demonstrates effectiveness in traffic incident duration prediction, particularly in complex urban environments, suggesting potential for sustainable prediction systems with reduced variables. Experimental findings underscore its significance for traffic management and planning. We utilized the SHAP technique to identify 14 features that contribute to predictions. The time period holds particular importance for emergency response. Variables such as GeoHash and temperature also significantly influence prediction accuracy across time classes, highlighting the nuanced dynamics of incident duration prediction in urban environments like Istanbul. Each method has its own strengths and weaknesses, highlighting the importance of selecting the appropriate method for a specific problem. RF&#x2019;s non-parametric nature allows for versatile application across diverse datasets, avoiding rigid assumptions and identifying intricate data patterns. RF minimizes individual tree variations by aggregating predictions from multiple decision trees trained on distinct data subsets, resulting in more accurate predictions through classification voting. Both RF and KNN exhibit high precision, recall, and F1-scores. DT&#x2019;s performance slightly lags due to potential overfitting with a single tree. Despite its longer runtime, RF is more robust against overlearning from unbalanced data distributions and less sensitive to hyperparameter tuning, resulting in consistently superior predictive performance compared to other models.The RF model&#x2019;s ability to calculate with 99% accuracy demonstrates its usefulness in dynamic models. It can be executed quickly and efficiently on high-processing computers, making it a valuable tool for managing traffic incidents in cities with high accident rates.</p>
<p>Deep learning and neural networks play a crucial role in predicting the duration of traffic incidents. Various studies have proposed innovative models integrating deep learning techniques such as LSTM, Bi-LSTM, and ANN autoencoders to enhance prediction accuracy. These models leverage features such as traffic flow, incident descriptions, and sensor data [<xref ref-type="bibr" rid="ref-38">38</xref>,<xref ref-type="bibr" rid="ref-40">40</xref>]. The fusion of machine learning with traffic data has shown significant improvements over traditional regression models, achieving up to a 60% accuracy enhancement [<xref ref-type="bibr" rid="ref-40">40</xref>]. Moreover, the interpretability of models such as TabNet has enabled the identification of key factors influencing incident duration, including road type, casualties, weather conditions, and vehicle numbers [<xref ref-type="bibr" rid="ref-61">61</xref>]. These advancements in deep learning models offer valuable insights for efficient resource allocation, emergency response, and traffic management strategies. It is therefore anticipated that even more significant outcomes may be achieved with this proposed approach in future studies as the field of deep learning continues to evolve.</p>
<p>The generalizability of the study can be improved by extending the estimation of traffic incident duration to larger geographical areas. Conducting similar analyses in various regions and countries beyond Istanbul will enable us to gain a broader perspective on the impacts of urban features and traffic infrastructure. Examining various geographic scaling methods can offer a more detailed analysis to determine the most suitable scaling strategy for predicting the duration of traffic events. In this regard, it is crucial to evaluate the impact of geohash scaling and alternative scaling methods. Integrating dynamic factors is essential for the prediction model to better adapt to real-world conditions. With the advent of IoT solutions and smart city applications, traffic event data can now be obtained much faster and independently of human input. By incorporating hard-to-obtain data, the prediction success of the model can be increased, allowing for more effective traffic management strategies and quicker responses to potential issues.</p>
</sec>
</body>
<back>
<ack><p>Not applicable.</p>
</ack>
<sec><title>Funding Statement</title>
<p>This research received no external funding.</p>
</sec>
<sec><title>Author Contributions</title>
<p>Conceptualization, Mesut Ulu; methodology, Mesut Ulu; software, Mesut Ulu, Kenan Meng&#x00FC;&#x00E7;; validation, Yusuf Sait T&#x00FC;rkan, Ersin Naml&#x0131;; formal analysis, Mesut Ulu, Kenan Meng&#x00FC;&#x00E7;; investigation, Mesut Ulu; resources, Tar&#x0131;k K&#x00FC;&#x00E7;&#x00FC;kdeniz; data curation, Mesut Ulu, Kenan Meng&#x00FC;&#x00E7;; writing&#x2014;original draft preparation, Mesut Ulu; writing&#x2014;review and editing, Yusuf Sait T&#x00FC;rkan; visualization, Mesut Ulu, Yusuf Sait T&#x00FC;rkan; supervision, Ersin Naml&#x0131;, Tar&#x0131;k K&#x00FC;&#x00E7;&#x00FC;kdeniz and Yusuf Sait T&#x00FC;rkan; project administration, Yusuf Sait T&#x00FC;rkan, Kenan Meng&#x00FC;&#x00E7;. All authors reviewed the results and approved the final version of the manuscript.</p>
</sec>
<sec sec-type="data-availability"><title>Availability of Data and Materials</title>
<p>Contact the corresponding author if interested in using data.</p>
</sec>
<sec sec-type="COI-statement"><title>Conflicts of Interest</title>
<p>The authors declare that they have no conflicts of interest to report regarding the present study.</p>
</sec>
<ref-list content-type="authoryear">
<title>References</title>
<ref id="ref-1"><label>[1]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>A. T.</given-names> <surname>Hojati</surname></string-name>, <string-name><given-names>L.</given-names> <surname>Ferreira</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Washington</surname></string-name>, and <string-name><given-names>P.</given-names> <surname>Charles</surname></string-name></person-group>, &#x201C;<article-title>Hazard based models for freeway traffic incident duration</article-title>,&#x201D; <source>Acci. Anal. Prev.</source>, vol. <volume>52</volume>, no. <issue>3</issue>, pp. <fpage>171</fpage>&#x2013;<lpage>181</lpage>, <year>2013</year>. doi: <pub-id pub-id-type="doi">10.1016/j.aap.2012.12.037</pub-id>; <pub-id pub-id-type="pmid">23333698</pub-id></mixed-citation></ref>
<ref id="ref-2"><label>[2]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><given-names>R.</given-names> <surname>&#x00DC;nl&#x00FC;</surname></string-name></person-group>, &#x201C;<chapter-title>Safer-driving: Application of deep transfer learning to build intelligent transportation systems</chapter-title>,&#x201D; in <person-group person-group-type="editor"><string-name><given-names>R.</given-names> <surname>Khaled</surname></string-name> and <string-name><given-names>A. A.</given-names> <surname>Ella Hassanien</surname></string-name></person-group>, (Eds.), <source>The Deep Learning and Big Data for Intelligent Transportation</source>, <edition>1st</edition> ed., <publisher-loc>Egypt</publisher-loc>: <publisher-name>Springer</publisher-name>, <year>2021</year>, vol. <volume>1</volume>, pp. <fpage>135</fpage>&#x2013;<lpage>150</lpage>.</mixed-citation></ref>
<ref id="ref-3"><label>[3]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>M.</given-names> <surname>Ulu</surname></string-name>, <string-name><given-names>E.</given-names> <surname>Kilic</surname></string-name>, and <string-name><given-names>Y. S.</given-names> <surname>T&#x00FC;rkan</surname></string-name></person-group>, &#x201C;<article-title>Prediction of traffic incident locations with a geohash-based model using machine learning</article-title>,&#x201D; <source>Algorithms Appl. Sci.</source>, vol. <volume>14</volume>, no. <issue>2</issue>, pp. <fpage>725</fpage>, <year>2024</year>. doi: <pub-id pub-id-type="doi">10.3390/app14020725</pub-id>.</mixed-citation></ref>
<ref id="ref-4"><label>[4]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>A.</given-names> <surname>Saracoglu</surname></string-name> and <string-name><given-names>H.</given-names> <surname>Ozen</surname></string-name></person-group>, &#x201C;<article-title>Estimation of traffic incident duration: A comparative study of decision tree models</article-title>,&#x201D; <source>Arab. J. Sci. Eng.</source>, vol. <volume>45</volume>, no. <issue>10</issue>, pp. <fpage>8099</fpage>&#x2013;<lpage>8110</lpage>, <year>2010</year>. doi: <pub-id pub-id-type="doi">10.1007/s13369-020-04615-2</pub-id>.</mixed-citation></ref>
<ref id="ref-5"><label>[5]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>D.</given-names> <surname>Nam</surname></string-name> and <string-name><given-names>F.</given-names> <surname>Mannering</surname></string-name></person-group>, &#x201C;<article-title>An exploratory hazard-based analysis of highway incident duration</article-title>,&#x201D; <source>Transp. Res. Part A: Policy Pract.</source>, vol. <volume>34</volume>, no. <issue>2</issue>, pp. <fpage>85</fpage>&#x2013;<lpage>102</lpage>, <year>2000</year>. doi: <pub-id pub-id-type="doi">10.1016/S0965-8564(98)00065-2</pub-id>.</mixed-citation></ref>
<ref id="ref-6"><label>[6]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>G.</given-names> <surname>Valenti</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Lelli</surname></string-name>, and <string-name><given-names>D.</given-names> <surname>Cucina</surname></string-name></person-group>, &#x201C;<article-title>A comparative study of models for the incident duration prediction</article-title>,&#x201D; <source>Eur. Transp. Res. Rev.</source>, vol. <volume>2</volume>, no. <issue>2</issue>, pp. <fpage>103</fpage>&#x2013;<lpage>111</lpage>, <year>2010</year>. doi: <pub-id pub-id-type="doi">10.1007/s12544-010-0031-4</pub-id>.</mixed-citation></ref>
<ref id="ref-7"><label>[7]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>R.</given-names> <surname>Li</surname></string-name></person-group>, &#x201C;<article-title>Traffic incident duration analysis and prediction models based on the survival analysis approach</article-title>,&#x201D; <source>IET Intell. Transp. Syst.</source>, vol. <volume>9</volume>, no. <issue>4</issue>, pp. <fpage>351</fpage>&#x2013;<lpage>358</lpage>, <year>2015</year>. doi: <pub-id pub-id-type="doi">10.1049/iet-its.2014.0036</pub-id>.</mixed-citation></ref>
<ref id="ref-8"><label>[8]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>B.</given-names> <surname>Ghosh</surname></string-name> and <string-name><given-names>J.</given-names> <surname>Dauwels</surname></string-name></person-group>, &#x201C;<article-title>Comparison of different Bayesian methods for estimating error bars with incident duration prediction</article-title>,&#x201D; <source>J. Intell. Transp. Syst.</source>, vol. <volume>26</volume>, no. <issue>4</issue>, pp. <fpage>420</fpage>&#x2013;<lpage>431</lpage>, <year>2022</year>. doi: <pub-id pub-id-type="doi">10.1080/15472450.2021.1894936</pub-id>.</mixed-citation></ref>
<ref id="ref-9"><label>[9]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>J.</given-names> <surname>Tang</surname></string-name> <etal>et al.</etal></person-group>, &#x201C;<article-title>Statistical and machine-learning methods for clearance time prediction of road incidents: A methodology review</article-title>,&#x201D; <source>Anal. Methods Accid. Res.</source>, vol. <volume>27</volume>, no. <issue>3</issue>, pp. <fpage>100123</fpage>, <year>2020</year>. doi: <pub-id pub-id-type="doi">10.1016/j.amar.2020.100123</pub-id>.</mixed-citation></ref>
<ref id="ref-10"><label>[10]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>A.</given-names> <surname>Garib</surname></string-name>, <string-name><given-names>A. E.</given-names> <surname>Radwan</surname></string-name>, and <string-name><given-names>H.</given-names> <surname>Al-Deek</surname></string-name></person-group>, &#x201C;<article-title>Estimating magnitude and duration of incident delays</article-title>,&#x201D; <source>J. Transp. Eng.</source>, vol. <volume>123</volume>, no. <issue>6</issue>, pp. <fpage>459</fpage>&#x2013;<lpage>466</lpage>, <year>1997</year>. doi: <pub-id pub-id-type="doi">10.1061/(ASCE)0733-947X(1997)123:6(459)</pub-id>.</mixed-citation></ref>
<ref id="ref-11"><label>[11]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>D.</given-names> <surname>Pan</surname></string-name> and <string-name><given-names>S.</given-names> <surname>Hamdar</surname></string-name></person-group>, &#x201C;<article-title>Prediction of traffic incident duration using: A hazard-based modeling for incident durations extracted through traffic detector data anomaly detection</article-title>,&#x201D; <source>Transp. Res. Rec.</source>, vol. <volume>2678</volume>, no. <issue>2</issue>, pp. <fpage>389</fpage>&#x2013;<lpage>400</lpage>, <year>2024</year>. doi: <pub-id pub-id-type="doi">10.1177/03611981231174445</pub-id>.</mixed-citation></ref>
<ref id="ref-12"><label>[12]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>H.</given-names> <surname>Zhao</surname></string-name> <etal>et al.</etal></person-group>, &#x201C;<article-title>Prediction of traffic incident duration using clustering-based ensemble learning method</article-title>,&#x201D; <source>J. Transp. Eng. Part A: Syst.</source>, vol. <volume>148</volume>, no. <issue>7</issue>, pp. <fpage>04022044</fpage>, <year>2022</year>. doi: <pub-id pub-id-type="doi">10.1061/JTEPBS.0000688</pub-id>.</mixed-citation></ref>
<ref id="ref-13"><label>[13]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>M.</given-names> <surname>Mumtarin</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Knickerbocker</surname></string-name>, <string-name><given-names>T.</given-names> <surname>Litteral</surname></string-name>, and <string-name><given-names>T. J. S.</given-names> <surname>Wood</surname></string-name></person-group>, &#x201C;<article-title>Traffic incident management performance measures: Ranking agencies on roadway clearance time</article-title>,&#x201D; <source>J. Transport. Technol.</source>, vol. <volume>13</volume>, no. <issue>3</issue>, pp. <fpage>353</fpage>&#x2013;<lpage>368</lpage>, <year>2023</year>. doi: <pub-id pub-id-type="doi">10.4236/jtts.2023.133017</pub-id>.</mixed-citation></ref>
<ref id="ref-14"><label>[14]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>A.</given-names> <surname>Grigorev</surname></string-name>, <string-name><given-names>A. S.</given-names> <surname>Mihaita</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Lee</surname></string-name>, and <string-name><given-names>F.</given-names> <surname>Chen</surname></string-name></person-group>, &#x201C;<article-title>Incident duration prediction using a bi-level machine learning framework with outlier removal and intra-extra joint optimization</article-title>,&#x201D; <source>Transp. Res. Part C: Emerg. Technol.</source>, vol. <volume>141</volume>, no. <issue>1</issue>, pp. <fpage>103721</fpage>, <year>2022</year>. doi: <pub-id pub-id-type="doi">10.1016/j.trc.2022.103721</pub-id>.</mixed-citation></ref>
<ref id="ref-15"><label>[15]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>R.</given-names> <surname>Li</surname></string-name>, and <string-name><given-names>M.</given-names> <surname>Guo</surname></string-name></person-group>, &#x201C;<article-title>Application of nonparametric regression in predicting traffic incident duration</article-title>,&#x201D; <source>Transport</source>, vol. <volume>33</volume>, no. <issue>1</issue>, pp. <fpage>22</fpage>&#x2013;<lpage>31</lpage>, <year>2018</year>. doi: <pub-id pub-id-type="doi">10.3846/16484142.2015.1004104</pub-id>.</mixed-citation></ref>
<ref id="ref-16"><label>[16]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>H.</given-names> <surname>Laman</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Yasmin</surname></string-name>, and <string-name><given-names>N.</given-names> <surname>Eluru</surname></string-name></person-group>, &#x201C;<article-title>Joint modeling of traffic incident duration components (reporting, response, and clearance time): A copula-based approach</article-title>,&#x201D; <source>Transp. Res. Rec.</source>, vol. <volume>2672</volume>, no. <issue>30</issue>, pp. <fpage>76</fpage>&#x2013;<lpage>89</lpage>, <year>2018</year>. doi: <pub-id pub-id-type="doi">10.1177/0361198118801355</pub-id>.</mixed-citation></ref>
<ref id="ref-17"><label>[17]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>L.</given-names> <surname>Yang</surname></string-name></person-group>, &#x201C;<article-title>Clearance time prediction of traffic accidents: A case study in Shandong</article-title>,&#x201D; <source>China Australas. J. Disaster Trauma Stud.</source>, vol. <volume>26</volume>, pp. <fpage>185</fpage>&#x2013;<lpage>194</lpage>, <year>2022</year>.</mixed-citation></ref>
<ref id="ref-18"><label>[18]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><given-names>X.</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Chen</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Gu</surname></string-name>, and <string-name><given-names>W.</given-names> <surname>Zhen</surname></string-name></person-group>, &#x201C;<chapter-title>Traffic incident duration analysis based on cyclic subspace regression</chapter-title>,&#x201D; in <source>The ICTE 2013: Safety, Speediness, Intelligence, Low-Carbon, Innovation</source>, <publisher-loc>USA</publisher-loc>, <year>2013</year>, pp. <fpage>2854</fpage>&#x2013;<lpage>2860</lpage>.</mixed-citation></ref>
<ref id="ref-19"><label>[19]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>G.</given-names> <surname>Giukiano</surname></string-name></person-group>, &#x201C;<article-title>Incident characteristics, frequency, and duration on a high volume urban freeway</article-title>,&#x201D; <source>Transp. Res. Part A</source>, vol. <volume>23</volume>, no. <issue>5</issue>, pp. <fpage>387</fpage>&#x2013;<lpage>396</lpage>, <year>1989</year>. doi: <pub-id pub-id-type="doi">10.1016/0191-2607(89)90086-1</pub-id>.</mixed-citation></ref>
<ref id="ref-20"><label>[20]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>R.</given-names> <surname>Li</surname></string-name>, <string-name><given-names>F. C.</given-names> <surname>Pereira</surname></string-name>, and <string-name><given-names>M. E.</given-names> <surname>Ben-Akiva</surname></string-name></person-group>, &#x201C;<article-title>Competing risks mixture model for traffic incident duration prediction</article-title>,&#x201D; <source>Accid. Anal. Prev.</source>, vol. <volume>75</volume>, no. <issue>6</issue>, pp. <fpage>192</fpage>&#x2013;<lpage>201</lpage>, <year>2015</year>. doi: <pub-id pub-id-type="doi">10.1016/j.aap.2014.11.023</pub-id>; <pub-id pub-id-type="pmid">25485730</pub-id></mixed-citation></ref>
<ref id="ref-21"><label>[21]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>A. J.</given-names> <surname>Khattak</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Liu</surname></string-name>, <string-name><given-names>B.</given-names> <surname>Wali</surname></string-name>, <string-name><given-names>X.</given-names> <surname>Li</surname></string-name>, and <string-name><given-names>M.</given-names> <surname>Ng</surname></string-name></person-group>, &#x201C;<article-title>Modeling traffic incident duration using quantile regression</article-title>,&#x201D; <source>Transp. Res. Rec.</source>, vol. <volume>2554</volume>, no. <issue>1</issue>, pp. <fpage>139</fpage>&#x2013;<lpage>148</lpage>, <year>2016</year>. doi: <pub-id pub-id-type="doi">10.3141/2554-15</pub-id>.</mixed-citation></ref>
<ref id="ref-22"><label>[22]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Y. S.</given-names> <surname>Chung</surname></string-name>, <string-name><given-names>Y. C.</given-names> <surname>Chiou</surname></string-name>, and <string-name><given-names>C. H.</given-names> <surname>Lin</surname></string-name></person-group>, &#x201C;<article-title>Simultaneous equation modeling of freeway accident duration and lanes blocked</article-title>,&#x201D; <source>Anal. Methods Accid. Res.</source>, vol. <volume>7</volume>, no. <issue>3</issue>, pp. <fpage>16</fpage>&#x2013;<lpage>28</lpage>, <year>2015</year>. doi: <pub-id pub-id-type="doi">10.1016/j.amar.2015.04.003</pub-id>.</mixed-citation></ref>
<ref id="ref-23"><label>[23]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>L.</given-names> <surname>Lin</surname></string-name>, <string-name><given-names>Q.</given-names> <surname>Wang</surname></string-name>, and <string-name><given-names>A. W.</given-names> <surname>Sadek</surname></string-name></person-group>, &#x201C;<article-title>A combined M5P tree and hazard-based duration model for predicting urban freeway traffic accident durations</article-title>,&#x201D; <source>Accid. Anal. Prev.</source>, vol. <volume>91</volume>, no. <issue>1</issue>, pp. <fpage>114</fpage>&#x2013;<lpage>126</lpage>, <year>2016</year>. doi: <pub-id pub-id-type="doi">10.1016/j.aap.2016.03.001</pub-id>; <pub-id pub-id-type="pmid">26974028</pub-id></mixed-citation></ref>
<ref id="ref-24"><label>[24]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>F.</given-names> <surname>Mouhous</surname></string-name>, <string-name><given-names>D.</given-names> <surname>Aissani</surname></string-name>, and <string-name><given-names>N.</given-names> <surname>Farhi</surname></string-name></person-group>, &#x201C;<article-title>A stochastic risk model for incident occurrences and duration in road networks</article-title>,&#x201D; <source>Transportmetrica A: Transp. Sci.</source>, vol. <volume>19</volume>, no. <issue>3</issue>, pp. <fpage>2077469</fpage>, <year>2023</year>. doi: <pub-id pub-id-type="doi">10.1080/23249935.2022.2077469</pub-id>.</mixed-citation></ref>
<ref id="ref-25"><label>[25]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>J. Y.</given-names> <surname>Lee</surname></string-name>, <string-name><given-names>J. H.</given-names> <surname>Chung</surname></string-name>, and <string-name><given-names>B.</given-names> <surname>Son</surname></string-name></person-group>, &#x201C;<article-title>Incident clearance time analysis for Korean freeways using structural equation model</article-title>,&#x201D; in <conf-name>Proc. Eastern Asia Soc. Transp. Studies </conf-name>, <year>2009</year>, Vol. <volume>7</volume>, pp. <fpage>360</fpage>.</mixed-citation></ref>
<ref id="ref-26"><label>[26]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Y.</given-names> <surname>Zou</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Zhang</surname></string-name>, and <string-name><given-names>D.</given-names> <surname>Lord</surname></string-name></person-group>, &#x201C;<article-title>Analyzing different functional forms of the varying weight parameter for finite mixture of negative binomial regression models</article-title>,&#x201D; <source>Anal. Methods Accid. Res.</source>, vol. <volume>1</volume>, no. <issue>2</issue>, pp. <fpage>39</fpage>&#x2013;<lpage>52</lpage>, <year>2014</year>. doi: <pub-id pub-id-type="doi">10.1016/j.amar.2013.11.001</pub-id>.</mixed-citation></ref>
<ref id="ref-27"><label>[27]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Y.</given-names> <surname>Zou</surname></string-name>, <string-name><given-names>X.</given-names> <surname>Ye</surname></string-name>, <string-name><given-names>K.</given-names> <surname>Henrickson</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Tang</surname></string-name>, and <string-name><given-names>Y.</given-names> <surname>Wang</surname></string-name></person-group>, &#x201C;<article-title>Jointly analyzing freeway traffic incident clearance and response time using a copula-based approach</article-title>,&#x201D; <source>Transp. Res. Part C: Emerg. Technol.</source>, vol. <volume>86</volume>, pp. <fpage>171</fpage>&#x2013;<lpage>182</lpage>, <year>2018</year>. doi: <pub-id pub-id-type="doi">10.1016/j.trc.2017.11.004</pub-id>.</mixed-citation></ref>
<ref id="ref-28"><label>[28]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>C. H.</given-names> <surname>Wei</surname></string-name> and <string-name><given-names>Y.</given-names> <surname>Lee</surname></string-name></person-group>, &#x201C;<article-title>Sequential forecast of incident duration using artificial neural network models</article-title>,&#x201D; <source>Accid. Anal. Prev.</source>, vol. <volume>39</volume>, no. <issue>5</issue>, pp. <fpage>944</fpage>&#x2013;<lpage>954</lpage>, <year>2007</year>. doi: <pub-id pub-id-type="doi">10.1016/j.aap.2006.12.017</pub-id>; <pub-id pub-id-type="pmid">17303059</pub-id></mixed-citation></ref>
<ref id="ref-29"><label>[29]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>A. J.</given-names> <surname>Khattak</surname></string-name>, <string-name><given-names>J. L.</given-names> <surname>Schofer</surname></string-name>, and <string-name><given-names>M. H.</given-names> <surname>Wang</surname></string-name></person-group>, &#x201C;<article-title>A simple time sequential procedure for predicting freeway incident duration</article-title>,&#x201D; <source>J. Intell. Transp. Syst.</source>, vol. <volume>2</volume>, no. <issue>2</issue>, pp. <fpage>113</fpage>&#x2013;<lpage>138</lpage>, <year>1995</year>. doi: <pub-id pub-id-type="doi">10.1080/10248079508903820</pub-id>.</mixed-citation></ref>
<ref id="ref-30"><label>[30]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Y.</given-names> <surname>Lee</surname></string-name> and <string-name><given-names>C. H.</given-names> <surname>Wei</surname></string-name></person-group>, &#x201C;<article-title>A computerized feature selection method using genetic algorithms to forecast freeway accident duration times</article-title>,&#x201D; <source>Comput. Aided Civ. Infrastruct. Eng.</source>, vol. <volume>25</volume>, no. <issue>2</issue>, pp. <fpage>132</fpage>&#x2013;<lpage>148</lpage>, <year>2010</year>. doi: <pub-id pub-id-type="doi">10.1111/j.1467-8667.2009.00626.x</pub-id>.</mixed-citation></ref>
<ref id="ref-31"><label>[31]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><given-names>W.</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>H.</given-names> <surname>Chen</surname></string-name>, and <string-name><given-names>M.</given-names> <surname>Bell</surname></string-name></person-group>, <chapter-title>A study of the characteristics of traffic incident duration on motorways</chapter-title>. in <source>Traffic and Transportation Studies</source>, <publisher-loc>USA</publisher-loc>: <publisher-name>American Society of Civil Engineers (Traffic and Transportation Studies)</publisher-name>, <year>2002</year>, pp. <fpage>1101</fpage>&#x2013;<lpage>1108</lpage>.</mixed-citation></ref>
<ref id="ref-32"><label>[32]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>K.</given-names> <surname>Ozbay</surname></string-name> and <string-name><given-names>E.</given-names> <surname>Noyan</surname></string-name></person-group>, &#x201C;<article-title>Estimation of incident clearance times using Bayesian networks approach</article-title>,&#x201D; <source>Accid. Anal. Prev.</source>, vol. <volume>38</volume>, no. <issue>3</issue>, pp. <fpage>542</fpage>&#x2013;<lpage>555</lpage>, <year>2006</year>. doi: <pub-id pub-id-type="doi">10.1016/j.aap.2005.11.012</pub-id>; <pub-id pub-id-type="pmid">16426557</pub-id></mixed-citation></ref>
<ref id="ref-33"><label>[33]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>H.</given-names> <surname>Park</surname></string-name>, <string-name><given-names>X.</given-names> <surname>Zhang</surname></string-name>, and <string-name><given-names>A.</given-names> <surname>Haghani</surname></string-name></person-group>, &#x201C;<article-title>ATIS: Interpretation of Bayesian neural network for predicting the duration of detected incidents</article-title>,&#x201D; in <conf-name>The 92nd Annu. Meet. Transportation Research Board (TRB 2013)</conf-name>, <year>Jan. 2013</year>, pp. <fpage>13</fpage>&#x2013;<lpage>17</lpage>.</mixed-citation></ref>
<ref id="ref-34"><label>[34]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>C.</given-names> <surname>Zhan</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Gan</surname></string-name>, and <string-name><given-names>M.</given-names> <surname>Hadi</surname></string-name></person-group>, &#x201C;<article-title>Prediction of lane clearance time of freeway incidents using the M5P tree algorithm</article-title>,&#x201D; <source>IEEE Trans. Intell. Transp. Syst.</source>, vol. <volume>12</volume>, no. <issue>4</issue>, pp. <fpage>1549</fpage>&#x2013;<lpage>1557</lpage>, <year>2011</year>. doi: <pub-id pub-id-type="doi">10.1109/TITS.2011.2161634</pub-id>.</mixed-citation></ref>
<ref id="ref-35"><label>[35]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>W.</given-names> <surname>Kim</surname></string-name> and <string-name><given-names>G. L.</given-names> <surname>Chang</surname></string-name></person-group>, &#x201C;<article-title>Development of a hybrid prediction model for freeway incident duration: A case study in Maryland</article-title>,&#x201D; <source>Int. J. Intell. Transp. Syst. Res.</source>, vol. <volume>10</volume>, no. <issue>1</issue>, pp. <fpage>22</fpage>&#x2013;<lpage>33</lpage>, <year>2012</year>. doi: <pub-id pub-id-type="doi">10.1007/s13177-011-0039-8</pub-id>.</mixed-citation></ref>
<ref id="ref-36"><label>[36]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>J.</given-names> <surname>Lu</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Chen</surname></string-name>, <string-name><given-names>W.</given-names> <surname>Wang</surname></string-name>, and <string-name><given-names>B.</given-names> <surname>Ran</surname></string-name></person-group>, &#x201C;<article-title>Automatic traffic incident detection based on nFOIL</article-title>,&#x201D; <source>Expert. Syst. Appl.</source>, vol. <volume>39</volume>, no. <issue>7</issue>, pp. <fpage>6547</fpage>&#x2013;<lpage>6556</lpage>, <year>2012</year>. doi: <pub-id pub-id-type="doi">10.1016/j.eswa.2011.12.050</pub-id>.</mixed-citation></ref>
<ref id="ref-37"><label>[37]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>K.</given-names> <surname>Hamad</surname></string-name>, <string-name><given-names>M. A.</given-names> <surname>Khalil</surname></string-name>, and <string-name><given-names>A. R.</given-names> <surname>Alozi</surname></string-name></person-group>, &#x201C;<article-title>Predicting freeway incident duration using machine learning</article-title>,&#x201D; <source>Int. J. Intell. Transp. Syst. Res.</source>, vol. <volume>18</volume>, no. <issue>2</issue>, pp. <fpage>367</fpage>&#x2013;<lpage>380</lpage>, <year>2020</year>. doi: <pub-id pub-id-type="doi">10.1007/s13177-019-00205-1</pub-id>.</mixed-citation></ref>
<ref id="ref-38"><label>[38]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>H.</given-names> <surname>Li</surname></string-name> and <string-name><given-names>Y.</given-names> <surname>Li</surname></string-name></person-group>, &#x201C;<article-title>A novel explanatory tabular neural network to predicting traffic incident duration using traffic safety big data</article-title>,&#x201D; <source>Mathematics</source>, vol. <volume>11</volume>, no. <issue>13</issue>, pp. <fpage>2915</fpage>, <year>2023</year>. doi: <pub-id pub-id-type="doi">10.3390/math11132915</pub-id>.</mixed-citation></ref>
<ref id="ref-39"><label>[39]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><given-names>H.</given-names> <surname>Wu</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Yan</surname></string-name>, and <string-name><given-names>M.</given-names> <surname>Liu</surname></string-name></person-group>, &#x201C;<article-title>Recent advances in graph-based machine learning for applications in smart urban transportation systems</article-title>,&#x201D; <comment>arXiv preprint arXiv:2306.01282</comment>, <year>2023</year>.</mixed-citation></ref>
<ref id="ref-40"><label>[40]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>A.</given-names> <surname>Grigorev</surname></string-name>, <string-name><given-names>A. S.</given-names> <surname>Mih&#x0103;i&#x0163;&#x0103;</surname></string-name>, <string-name><given-names>K.</given-names> <surname>Saleh</surname></string-name>, and <string-name><given-names>M.</given-names> <surname>Piccardi</surname></string-name></person-group>, &#x201C;<article-title>Traffic incident duration prediction via a deep learning framework for text description encoding</article-title>,&#x201D; in <conf-name>2022 IEEE 25th Int. Conf. Intell. Transp. Syst. (ITSC)</conf-name>, <publisher-loc>China</publisher-loc>, <year>Oct. 8&#x2013;12, 2022</year>, pp. <fpage>1770</fpage>&#x2013;<lpage>1777</lpage>.</mixed-citation></ref>
<ref id="ref-41"><label>[41]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>P. W.</given-names> <surname>Lin</surname></string-name>, <string-name><given-names>N.</given-names> <surname>Zou</surname></string-name>, and <string-name><given-names>G. L.</given-names> <surname>Chang</surname></string-name></person-group>, &#x201C;<article-title>Integration of a discrete choice model and a rule-based system for estimation of incident duration: a case study in Maryland</article-title>,&#x201D; in <conf-name>CD-ROM of Proc. 83rd TRB Annu. Meet.</conf-name>, <publisher-loc>Washington, DC, USA</publisher-loc>, <year>Jan. 2004</year>.</mixed-citation></ref>
<ref id="ref-42"><label>[42]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>H.</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Zhang</surname></string-name>, and <string-name><given-names>A. J.</given-names> <surname>Khattak</surname></string-name></person-group>, &#x201C;<article-title>Analysis of large-scale incidents on urban freeways</article-title>,&#x201D; <source>Transp. Res. Rec.</source>, vol. <volume>2278</volume>, no. <issue>1</issue>, pp. <fpage>74</fpage>&#x2013;<lpage>84</lpage>, <year>2012</year>. doi: <pub-id pub-id-type="doi">10.3141/2278-09</pub-id>.</mixed-citation></ref>
<ref id="ref-43"><label>[43]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><given-names>K.</given-names> <surname>Smith</surname></string-name> and <string-name><given-names>B. L.</given-names> <surname>Smith</surname></string-name></person-group>, &#x201C;<article-title>Forecasting the clearance time of freeway accidents</article-title>,&#x201D; <year>2022</year>. Accessed: Feb. 15, 2024. <ext-link ext-link-type="uri" xlink:href="https://rosap.ntl.bts.gov/view/dot/34048">https://rosap.ntl.bts.gov/view/dot/34048</ext-link></mixed-citation></ref>
<ref id="ref-44"><label>[44]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><collab>US Department of Transportation</collab></person-group>, <source>Manual on Uniform Traffic Control Devices; For Streets and Highways</source>. <publisher-loc>Office of Highway Policy Information</publisher-loc>, <publisher-name>US Department of Transportation, Federal Highway Administration</publisher-name>, <year>2009</year>. <comment>Accessed: Feb. 15, 2024</comment>. [Online]. Available: <ext-link ext-link-type="uri" xlink:href="https://www.fhwa.dot.gov/policyinformation/statistics//2009">https://www.fhwa.dot.gov/policyinformation/statistics//2009</ext-link></mixed-citation></ref>
<ref id="ref-45"><label>[45]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>M. A.</given-names> <surname>Islam</surname></string-name></person-group>, &#x201C;<article-title>A literature review on freeway traffic incidents and their impact on traffic operations</article-title>,&#x201D; <source>J. Transport. Technol.</source>, vol. <volume>9</volume>, no. <issue>4</issue>, pp. <fpage>504</fpage>&#x2013;<lpage>516</lpage>, <year>2019</year>. doi: <pub-id pub-id-type="doi">10.4236/jtts.2019.94032</pub-id>.</mixed-citation></ref>
<ref id="ref-46"><label>[46]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>O. Z.</given-names> <surname>Maimon</surname></string-name> and <string-name><given-names>L.</given-names> <surname>Rokach</surname></string-name></person-group>, &#x201C;<article-title>Data mining with decision trees: Theory and applications</article-title>,&#x201D; <source>World Sci.</source>, vol. <volume>81</volume>, <year>2014</year>.</mixed-citation></ref>
<ref id="ref-47"><label>[47]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>J. R.</given-names> <surname>Quinlan</surname></string-name></person-group>, &#x201C;<article-title>Learning decision tree classifiers</article-title>,&#x201D; <source>ACM Comput. Surveys</source>, vol. <volume>28</volume>, no. <issue>1</issue>, pp. <fpage>71</fpage>&#x2013;<lpage>72</lpage>, <year>1996</year>. doi: <pub-id pub-id-type="doi">10.1145/234313.234346</pub-id>.</mixed-citation></ref>
<ref id="ref-48"><label>[48]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>L.</given-names> <surname>Breiman</surname></string-name></person-group>, &#x201C;<article-title>Random forests</article-title>,&#x201D; <source>Mach. Learn.</source>, vol. <volume>45</volume>, pp. <fpage>5</fpage>&#x2013;<lpage>32</lpage>, <year>2001</year>. <comment>Accessed: Mar. 12, 2024</comment>. <comment>[Online]. Available:</comment> <ext-link ext-link-type="uri" xlink:href="https://link.springer.com/article/10.1023/A:1010933404324">https://link.springer.com/article/10.1023/A:1010933404324</ext-link>.</mixed-citation></ref>
<ref id="ref-49"><label>[49]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>S. T.</given-names> <surname>Ikram</surname></string-name>, <string-name><given-names>V.</given-names> <surname>Priya</surname></string-name>, <string-name><given-names>B.</given-names> <surname>Anbarasu</surname></string-name>, <string-name><given-names>X.</given-names> <surname>Cheng</surname></string-name>, <string-name><given-names>M. R.</given-names> <surname>Ghalib</surname></string-name> and <string-name><given-names>A.</given-names> <surname>Shankar</surname></string-name></person-group>, &#x201C;<article-title>Prediction of IIoT traffic using a modified whale optimization approach integrated with random forest classifier</article-title>,&#x201D; <source>J. Supercomput.</source>, vol. <volume>78</volume>, no. <issue>8</issue>, pp. <fpage>10725</fpage>&#x2013;<lpage>10756</lpage>, <year>2022</year>. doi: <pub-id pub-id-type="doi">10.1007/s11227-021-04284-4</pub-id>.</mixed-citation></ref>
<ref id="ref-50"><label>[50]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>G.</given-names> <surname>Guo</surname></string-name>, <string-name><given-names>H.</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>D.</given-names> <surname>Bell</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Bi</surname></string-name>, and <string-name><given-names>K.</given-names> <surname>Greer</surname></string-name></person-group>, &#x201C;<article-title>KNN model-based approach in classification</article-title>,&#x201D; in <conf-name>Move Meaningful Internet Syst. 2003: CoopIS, DOA, ODBASE: OTM Confederated Int. Conf. CoopIS, DOA, ODBASE 2003</conf-name>, <publisher-loc>Catania, Sicily, Italy</publisher-loc>, <year>Nov. 3&#x2013;7, 2023</year>.</mixed-citation></ref>
<ref id="ref-51"><label>[51]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><given-names>H.</given-names> <surname>Trevor</surname></string-name>, <string-name><given-names>R.</given-names> <surname>Tibshirani</surname></string-name>, and <string-name><given-names>J.</given-names> <surname>Friedman</surname></string-name></person-group>, &#x201C;<chapter-title>Methods and nearest-neighbors</chapter-title>,&#x201D; in <source>The Elements of Statistical Learning</source>, <edition>1st</edition> ed., <publisher-loc>New York, NY, USA</publisher-loc>: <publisher-name>Springer International Publishing</publisher-name>, <year>2011</year>. <comment>Accessed: Mar. 12, 2024</comment>. [Online]. Available: <ext-link ext-link-type="uri" xlink:href="https://link.springer.com/book/10.1007/978-0-387-21606-5">https://link.springer.com/book/10.1007/978-0-387-21606-5</ext-link></mixed-citation></ref>
<ref id="ref-52"><label>[52]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><given-names>D. A.</given-names> <surname>Pisner</surname></string-name> and <string-name><given-names>D. M.</given-names> <surname>Schnyer</surname></string-name></person-group>, &#x201C;<chapter-title>Support vector machine</chapter-title>,&#x201D; in <source>Machine Learning</source>. <publisher-loc>London</publisher-loc>: <publisher-name>United Kingdom Academic Press</publisher-name>, <year>2020</year>, pp. <fpage>101</fpage>&#x2013;<lpage>121</lpage>. <comment>Accessed: Mar. 12, 2024</comment>. [Online]. Available: <ext-link ext-link-type="uri" xlink:href="https://www.sciencedirect.com/science/article/abs/pii/B9780128157398000067">https://www.sciencedirect.com/science/article/abs/pii/B9780128157398000067</ext-link></mixed-citation></ref>
<ref id="ref-53"><label>[53]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>H.</given-names> <surname>Li</surname></string-name>, <string-name><given-names>L.</given-names> <surname>Xiong</surname></string-name>, <string-name><given-names>L.</given-names> <surname>Ohno-Machado</surname></string-name>, and <string-name><given-names>X.</given-names> <surname>Jiang</surname></string-name></person-group>, &#x201C;<article-title>Privacy preserving RBF kernel support vector machine</article-title>,&#x201D; <source>Biomed Res. Int.</source>, vol. <volume>2014</volume>, no. <issue>1</issue>, pp. <fpage>1</fpage>&#x2013;<lpage>10</lpage>, <year>2014</year>. doi: <pub-id pub-id-type="doi">10.1155/2014/827371</pub-id>; <pub-id pub-id-type="pmid">25013805</pub-id></mixed-citation></ref>
<ref id="ref-54"><label>[54]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>R.</given-names> <surname>Yuan</surname></string-name>, <string-name><given-names>Z.</given-names> <surname>Li</surname></string-name>, <string-name><given-names>X.</given-names> <surname>Guan</surname></string-name>, and <string-name><given-names>L.</given-names> <surname>Xu</surname></string-name></person-group>, &#x201C;<article-title>An SVM-based machine learning method for accurate internet traffic classification</article-title>,&#x201D; <source>Inf. Syst. Front.</source>, vol. <volume>12</volume>, no. <issue>2</issue>, pp. <fpage>149</fpage>&#x2013;<lpage>156</lpage>, <year>2010</year>. doi: <pub-id pub-id-type="doi">10.1007/s10796-008-9131-2</pub-id>.</mixed-citation></ref>
<ref id="ref-55"><label>[55]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>L.</given-names> <surname>Liao</surname></string-name>, <string-name><given-names>H.</given-names> <surname>Li</surname></string-name>, <string-name><given-names>W.</given-names> <surname>Shang</surname></string-name>, and <string-name><given-names>L.</given-names> <surname>Ma</surname></string-name></person-group>, &#x201C;<article-title>An empirical study of the impact of hyperparameter tuning and model optimization on the performance properties of deep neural networks</article-title>,&#x201D; <source>ACM Trans. Softw. Eng. Methodol.</source>, vol. <volume>31</volume>, no. <issue>3</issue>, pp. <fpage>1</fpage>&#x2013; <lpage>40</lpage>, <year>2022</year>. doi: <pub-id pub-id-type="doi">10.1145/3506695</pub-id>.</mixed-citation></ref>
<ref id="ref-56"><label>[56]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><given-names>A.</given-names> <surname>G&#x00E9;ron</surname></string-name></person-group>, <source>Hands-on Machine Learning with Scikit-Learn, Keras, and TensorFlow</source>, <edition>2nd</edition> ed., <publisher-loc>Sebastopol, CA, USA</publisher-loc>: <publisher-name>O&#x2019;Reilly Media, Inc.</publisher-name>, <year>2022</year>.</mixed-citation></ref>
<ref id="ref-57"><label>[57]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>F.</given-names> <surname>Izhari</surname></string-name> and <string-name><given-names>H. W.</given-names> <surname>Dhany</surname></string-name></person-group>, &#x201C;<article-title>Optimizing urban traffic management through advanced machine learning: A comprehensive study</article-title>,&#x201D; <source>J. Intell. Decis. Support Syst.</source>, vol. <volume>6</volume>, no. <issue>4</issue>, pp. <fpage>223</fpage>&#x2013;<lpage>230</lpage>, <year>2023</year>.</mixed-citation></ref>
<ref id="ref-58"><label>[58]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>S. A.</given-names> <surname>Samerei</surname></string-name>, <string-name><given-names>K.</given-names> <surname>Aghabayk</surname></string-name>, and <string-name><given-names>A.</given-names> <surname>Montella</surname></string-name></person-group>, &#x201C;<article-title>Analyzing pile-up crash severity: Insights from real-time traffic and environmental factors using ensemble machine learning and shapley additive explanations method</article-title>,&#x201D; <source>Safety</source>, vol. <volume>10</volume>, no. <issue>1</issue>, pp. <fpage>22</fpage>, <year>2024</year>. doi: <pub-id pub-id-type="doi">10.3390/safety10010022</pub-id>.</mixed-citation></ref>
<ref id="ref-59"><label>[59]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>L.</given-names> <surname>Antwarg</surname></string-name>, <string-name><given-names>R. M.</given-names> <surname>Miller</surname></string-name>, <string-name><given-names>B.</given-names> <surname>Shapira</surname></string-name>, and <string-name><given-names>L.</given-names> <surname>Rokach</surname></string-name></person-group>, &#x201C;<article-title>Explaining anomalies detected by autoencoders using shapley additive explanations</article-title>,&#x201D; <source>Expert. Syst. Appl.</source>, vol. <volume>186</volume>, no. <issue>5</issue>, pp. <fpage>115736</fpage>, <year>2021</year>. doi: <pub-id pub-id-type="doi">10.1016/j.eswa.2021.115736</pub-id>.</mixed-citation></ref>
<ref id="ref-60"><label>[60]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Boughorbel</surname></string-name>, <string-name><given-names>F.</given-names> <surname>Jarray</surname></string-name>, and <string-name><given-names>M.</given-names> <surname>El-Anbari</surname></string-name></person-group>, &#x201C;<article-title>Optimal classifier for imbalanced data using Matthews correlation coefficient metric</article-title>,&#x201D; <source>PLoS One</source>, vol. <volume>12</volume>, no. <issue>6</issue>, pp. <fpage>e0177678</fpage>, <year>2017</year>. doi: <pub-id pub-id-type="doi">10.1371/journal.pone.0177678</pub-id>; <pub-id pub-id-type="pmid">28574989</pub-id></mixed-citation></ref>
<ref id="ref-61"><label>[61]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Q.</given-names> <surname>Shang</surname></string-name>, <string-name><given-names>T.</given-names> <surname>Xie</surname></string-name>, and <string-name><given-names>Y.</given-names> <surname>Yu</surname></string-name></person-group>, &#x201C;<article-title>Prediction of duration of traffic incidents by hybrid deep learning based on multi-source incomplete data</article-title>,&#x201D; <source>Int. J. Environ. Res. Public Health</source>, vol. <volume>19</volume>, no. <issue>17</issue>, pp. <fpage>10903</fpage>, <year>2022</year>. doi: <pub-id pub-id-type="doi">10.3390/ijerph191710903</pub-id>; <pub-id pub-id-type="pmid">36078617</pub-id></mixed-citation></ref>
</ref-list>
</back></article>