<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.1 20151215//EN" "http://jats.nlm.nih.gov/publishing/1.1/JATS-journalpublishing1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xml:lang="en" article-type="research-article" dtd-version="1.1">
<front>
<journal-meta>
<journal-id journal-id-type="pmc">CMES</journal-id>
<journal-id journal-id-type="nlm-ta">CMES</journal-id>
<journal-id journal-id-type="publisher-id">CMES</journal-id>
<journal-title-group>
<journal-title>Computer Modeling in Engineering &#x0026; Sciences</journal-title>
</journal-title-group>
<issn pub-type="epub">1526-1506</issn>
<issn pub-type="ppub">1526-1492</issn>
<publisher>
<publisher-name>Tech Science Press</publisher-name>
<publisher-loc>USA</publisher-loc>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">31350</article-id>
<article-id pub-id-type="doi">10.32604/cmes.2023.031350</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Article</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>Transformer-Aided Deep Double Dueling Spatial-Temporal Q-Network for Spatial Crowdsourcing Analysis</article-title>
<alt-title alt-title-type="left-running-head">Transformer-aided Deep Double Dueling Spatial-Temporal Q-Network for Spatial Crowdsourcing Analysis</alt-title>
<alt-title alt-title-type="right-running-head">Transformer-aided Deep Double Dueling Spatial-Temporal Q-Network for Spatial Crowdsourcing Analysis</alt-title>
</title-group>
<contrib-group>
<contrib id="author-1" contrib-type="author">
<name name-style="western"><surname>Li</surname><given-names>Yu</given-names></name></contrib>
<contrib id="author-2" contrib-type="author">
<name name-style="western"><surname>Li</surname><given-names>Mingxiao</given-names></name></contrib>
<contrib id="author-3" contrib-type="author" corresp="yes">
<name name-style="western"><surname>Ou</surname><given-names>Dongyang</given-names></name><email>oudongyang@hdu.edu.cn</email></contrib>
<contrib id="author-4" contrib-type="author">
<name name-style="western"><surname>Guo</surname><given-names>Junjie</given-names></name></contrib> <contrib id="author-5" contrib-type="author">
<name name-style="western"><surname>Pan</surname><given-names>Fangyuan</given-names></name></contrib>
<aff><institution>Department of Computing, Hangzhou Dianzi University</institution>, <addr-line>Hangzhou, 310018</addr-line>, <country>China</country></aff>
</contrib-group>
<author-notes>
<corresp id="cor1"><label>&#x002A;</label>Corresponding Author: Dongyang Ou. Email: <email>oudongyang@hdu.edu.cn</email></corresp>
</author-notes>
<pub-date date-type="collection" publication-format="electronic">
<year>2023</year></pub-date>
<pub-date date-type="pub" publication-format="electronic"><day>30</day>
<month>12</month>
<year>2023</year></pub-date>
<volume>139</volume>
<issue>1</issue>
<fpage>893</fpage>
<lpage>909</lpage>
<history>
<date date-type="received">
<day>08</day>
<month>6</month>
<year>2023</year>
</date>
<date date-type="accepted">
<day>11</day>
<month>10</month>
<year>2023</year>
</date>
</history>
<permissions>
<copyright-statement>&#x00A9; 2024 Li et al.</copyright-statement>
<copyright-year>2024</copyright-year>
<copyright-holder>Li et al.</copyright-holder>
<license xlink:href="https://creativecommons.org/licenses/by/4.0/">
<license-p>This work is licensed under a <ext-link ext-link-type="uri" xlink:type="simple" xlink:href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</ext-link>, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.</license-p>
</license>
</permissions>
<self-uri content-type="pdf" xlink:href="TSP_CMES_31350.pdf"></self-uri>
<abstract>
<p>With the rapid development of mobile Internet, spatial crowdsourcing has become more and more popular. Spatial crowdsourcing consists of many different types of applications, such as spatial crowd-sensing services. In terms of spatial crowd-sensing, it collects and analyzes traffic sensing data from clients like vehicles and traffic lights to construct intelligent traffic prediction models. Besides collecting sensing data, spatial crowdsourcing also includes spatial delivery services like DiDi and Uber. Appropriate task assignment and worker selection dominate the service quality for spatial crowdsourcing applications. Previous research conducted task assignments via traditional matching approaches or using simple network models. However, advanced mining methods are lacking to explore the relationship between workers, task publishers, and the spatio-temporal attributes in tasks. Therefore, in this paper, we propose a Deep Double Dueling Spatial-temporal Q Network (D3SQN) to adaptively learn the spatial-temporal relationship between task, task publishers, and workers in a dynamic environment to achieve optimal allocation. Specifically, D3SQN is revised through reinforcement learning by adding a spatial-temporal transformer that can estimate the expected state values and action advantages so as to improve the accuracy of task assignments. Extensive experiments are conducted over real data collected from DiDi and ELM, and the simulation results verify the effectiveness of our proposed models.</p>
</abstract>
<kwd-group kwd-group-type="author">
<kwd>Historical behavior analysis</kwd>
<kwd>spatial crowdsourcing</kwd>
<kwd>deep double dueling Q-networks</kwd>
</kwd-group>
<funding-group>
<award-group id="awg1">
<funding-source>Pioneer and Leading Goose R&#x0026;D Program of Zhejiang Province</funding-source>
<award-id>2022C01083</award-id>
</award-group>
<award-group id="awg2">
<funding-source>Pioneer and Leading Goose R&#x0026;D Program of Zhejiang Province</funding-source>
<award-id>2023C01217</award-id>
</award-group>
</funding-group>
</article-meta>
</front>
<body>
<sec id="s1">
<label>1</label>
<title>Introduction</title>
<p>With the development of intelligent transportation, spatial crowdsourcing attracts more and more attention. Various applications of spatial crowdsourcing appear, such as spatial crowd-sensing which may collect and analyze traffic data from edge clients (e.g., vehicles, traffic lights) to help predict and solve traffic jams. Besides, spatial crowdsourcing platforms like DiDi and Uber also utilize edge vehicles to provide intelligent services. For these spatial crowdsourcing applications, how to choose the appropriate crowdsourcing worker is an essential task. Therefore, in this paper, we study how to assign appropriate workers (i.e., spatial vehicles and drivers) for spatial delivery tasks on spatial crowdsourcing platforms by analyzing workers&#x2019; historical behavior data.</p>
<p>Similar to traditional crowdsourcing platforms, spatial crowdsourcing platforms comprise three components: the task/request, the worker, and the platform. However, the tasks released on spatial crowdsourcing platforms are spatial tasks with spatio-temporal attributes. Spatial delivery tasks have spatio-temporal constraints, such as start and target locations, start times, and deadlines. A spatial delivery task is completed only if the requester is picked up at the source location within the requested time and successfully delivered to the target location before the deadline. In spatial crowdsourcing platforms, the positions of workers and requesters may change dynamically, especially the spatio-temporal attributes of workers during the completion of tasks are always changing. This paper focuses on common spatial delivery tasks in daily real-time ride-hailing services, such as Uber and DiDi Chuxing.</p>
<p>In terms of spatial delivery task assignment on spatial crowdsourcing platforms, the key is to recommend a suitable task list to workers in a dynamic environment to maximize the benefits of workers, requesters and the platform. Traditional task assignment approaches to spatial crowdsourcing platforms mainly utilize matching approaches. However, with the explosive growth of vehicles in intelligent transportation, task allocation should take into account not only the location matching information but also the workers&#x2019; preferences. The preferences of the workers and the requestors have a greater impact on the completion of the task. This can be critical to the user experience in real applications. For instance, DiDi Chuxing needed to serve 25 million ride requests a day and have more than 21 million registered drivers (i.e., workers). Some requesters may choose a female driver for safety, while others may choose a male driver for speed. Some drivers may prefer delivery orders over downtown to get more future orders, while some drivers may prefer orders nearby to avoid traffic jams. In addition, we observed that the preferences of requesters and workers may change over time, making the presetting of preferences is impractical in real-world applications [<xref ref-type="bibr" rid="ref-1">1</xref>&#x2013;<xref ref-type="bibr" rid="ref-3">3</xref>].</p>
<p>To deal with the complex preferences of requesters and workers, neural networks are utilized to find appropriate task assignments. Shan et al. [<xref ref-type="bibr" rid="ref-4">4</xref>] was a state-of-the-art study that proposed a deep reinforcement learning model to solve the task scheduling problem in traditional non-spatial crowdsourcing platforms. However, task assignment of spatio-temporal delivery tasks is more complex since it is closely related to the spatio-temporal attributes of requesters and workers, and the constraints of task completion are also complex. The deep Q-network in [<xref ref-type="bibr" rid="ref-4">4</xref>] cannot deal with the input types of spatio-temporal tasks and workers well, moreover, it cannot extract the interrelation between workers, tasks, and the workers/requesters&#x2019; preferences. As a result, it is impossible to recommend appropriate spatial delivery tasks to spatial crowdsourcing workers, resulting in poor revenue for the entire platform.</p>
<p>In order to apply reinforcement learning models to solve task assignments on spatial crowdsourcing platforms, we refine the deep Q-network model in [<xref ref-type="bibr" rid="ref-4">4</xref>] by revising the architecture of deep Q-Network, including a State Transformer to process spatio-temporal input information. We propose a Double Dueling Deep Spatial Q Network (D3SQN) based on deep reinforcement learning framework specifically for spatio-temporal crowdsourcing task scheduling. In detail, we model the interaction between the spatiotemporal crowdsourcing environment (workers and requests) and the agent (platform) as a new Markov decision process (MDP), taking into account the longitude and latitude information of the task and worker. We apply our proposed D3SQN network to estimate the reward for recommending each task to an upcoming worker. D3SQN takes into account both current and future returns in the online environment, as well as workers&#x2019; and requesters&#x2019; preferences and how these preferences interrelate with spatio-temporal properties in the task. Applying reinforcement learning models to utilize user preferences in spatial crowdsourcing platforms has not been studied in previous literature, but this information is very important as the completion degree and satisfaction of the next task in the spatio-temporal crowdsourcing scenario are strongly related to it. Our contributions can be summarized as follows:
<list list-type="bullet">
<list-item>
<p>As far as we know, we are the first to apply deep reinforcement learning to assign tasks in spatial crowdsourcing platforms, and our proposed D3SQN can handle both current and future rewards to achieve long-term optimal allocation. In addition, the Q-value is quantified by using the user&#x2019;s historical behavior, so that the model can reflect the user&#x2019;s preference adaptively.</p></list-item>
<list-item>
<p>In addition to using the structure of Double DQN, we also modify the loss function and include a state transformer, which further improves the overall performance.</p></list-item>
<list-item>
<p>We use real data sets to demonstrate the effectiveness and efficiency of our framework.</p></list-item>
</list></p>
</sec>
<sec id="s2">
<label>2</label>
<title>Related Work</title>
<p>In this section, most related research is discussed.</p>
<sec id="s2_1">
<label>2.1</label>
<title>Deep Q Networks and Transformer</title>
<p>Deep Q Network (DQN) [<xref ref-type="bibr" rid="ref-5">5</xref>] is a combination of deep learning and reinforcement learning, that is, neural networks are used to replace the Q table in Q-learning. In addition to being widely used in reference scenarios, more and more improved models have been proposed for the defects of DQN [<xref ref-type="bibr" rid="ref-5">5</xref>]. For example, DDQN [<xref ref-type="bibr" rid="ref-6">6</xref>], a target network is added on the basis of DQN, which can reduce overestimation to some extent. D3QN uses Dueling Network [<xref ref-type="bibr" rid="ref-7">7</xref>] architecture on the basis of DDQN. It uses the network to express two estimators, namely the state value function and the action advantage function depending on the state. This factorization generalizes the learning of actions. Rainbow [<xref ref-type="bibr" rid="ref-8">8</xref>] combines 6 extended improvements to the DQN [<xref ref-type="bibr" rid="ref-5">5</xref>] algorithm, including DDQN [<xref ref-type="bibr" rid="ref-6">6</xref>], Dueling DQN [<xref ref-type="bibr" rid="ref-7">7</xref>], Prioritized Boy-Replay [<xref ref-type="bibr" rid="ref-9">9</xref>], Multi-step Learning [<xref ref-type="bibr" rid="ref-10">10</xref>], Distributional RL [<xref ref-type="bibr" rid="ref-11">11</xref>] and Noisy Net [<xref ref-type="bibr" rid="ref-12">12</xref>].</p>
<p>EndToEnd [<xref ref-type="bibr" rid="ref-4">4</xref>] proposed the framework of deep reinforcement learning (RL) for task scheduling, which is a key issue for the success of crowdsourcing platforms. However, the original DQN [<xref ref-type="bibr" rid="ref-5">5</xref>] model is still used in EndToEnd [<xref ref-type="bibr" rid="ref-4">4</xref>], and the original DQN [<xref ref-type="bibr" rid="ref-5">5</xref>] and its variants are not well adapted to the spatio-temporal crowdsourcing problem.</p>
<p>The transformer [<xref ref-type="bibr" rid="ref-13">13</xref>] has excelled in a wide variety of areas, including language modeling [<xref ref-type="bibr" rid="ref-14">14</xref>], summarization [<xref ref-type="bibr" rid="ref-15">15</xref>], question answering [<xref ref-type="bibr" rid="ref-16">16</xref>], and machine translation [<xref ref-type="bibr" rid="ref-17">17</xref>], etc. The transformer has made a significant breakthrough in reinforcement learning (RL) with the work of Parisotto et al. [<xref ref-type="bibr" rid="ref-18">18</xref>]. The GTrxl architecture proposed by this work enables it to learn dependencies beyond a fixed length without breaking the temporal coherence. This allows it to make predictions using current input trajectories plus past trajectories. Not only that, but Transformer-XL introduces a new relative location encoding scheme that not only learns longer term dependencies, but also addresses context fragmentation. However, GTrxl cannot cope well with the input in the spatio-temporal crowdsourcing environment because of the permutation invariance of the input.</p>
</sec>
<sec id="s2_2">
<label>2.2</label>
<title>Sptial Crowdsourcing</title>
<p>Spatio-temporal crowdsourcing is a process in which a group of crowdsourcing tasks with spatio-temporal attributes are given to a group of workers. The spatio-temporal attributes of the workers must meet the spatio-temporal constraints of the tasks before they can perform the corresponding tasks. The research on spatial crowdsourcing has been very popular in recent years. Hassan et al. [<xref ref-type="bibr" rid="ref-19">19</xref>] defined an online space task allocation framework based on the formalization of the multi-armed robber problem to solve the online space task allocation problem. Wang et al. [<xref ref-type="bibr" rid="ref-20">20</xref>] designed a new adaptive batch-based constant competitive ratio solution framework to solve the dynamic bipartite graph matching (DBGM) problem. Liu et al. [<xref ref-type="bibr" rid="ref-21">21</xref>] proposed a two-stage solution to the on-demand food delivery problem in FooDNet. Cheng et al. [<xref ref-type="bibr" rid="ref-22">22</xref>] proposed two optimization methods, task-first greed (TPG) and game theory (GT), to solve the cooperative perceptual spatial crowdsourcing (CA-SC) problem. Shan et al. [<xref ref-type="bibr" rid="ref-4">4</xref>] proposed a deep reinforcement learning (RL) task assignment framework and used DQN for task recommendation in traditional non-spatial crowdsourcing platforms. However, their model cannot work well for spatial delivery tasks directly.</p>
<p>In summary, none of the existing models can work well for spatial delivery task assignments. Inspired by the existing literature, we propose a variation of D3QN which applies transformer to reinforce learning to solve spatio-temporal crowdsourcing task assignment problems.</p>
</sec>
</sec>
<sec id="s3">
<label>3</label>
<title>Problem</title>
<sec id="s3_1">
<label>3.1</label>
<title>Problem Summary</title>
<p>The goal of the task scheduling system of spatio-temporal crowdsourcing is to recommend a sorted list of tasks to a coming valid worker. As the platform&#x2019;s profit model is to complete tasks on commission, the system should satisfy both workers and requesters.</p>
<p>For each worker, as many suitable tasks as possible can be found (within the acceptable time and space of the worker). Requesters expect their tasks to be served as efficiently and comfortably as possible. That is, the task is picked up at the designated location before the expected start time of the task and the task is completed at the destination as early as possible before the specified time.</p>
<p>In addition, since tasks and staff change dynamically, the system should cope with dynamically changed workers&#x2019; and requesters&#x2019; sets as well as their various and changeable preferences globally in real-time.</p>
</sec>
<sec id="s3_2">
<label>3.2</label>
<title>Problem Definition</title>
<p>In this section, we formally define our spatio-temporal crowdsourcing task allocation problem.</p>
<p>In the spatio-temporal crowdsourcing platform, a worker is represented by <inline-formula id="ieqn-1"><mml:math id="mml-ieqn-1"><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, where i is the moment when the worker goes online. The coordinate of the worker at this moment is <inline-formula id="ieqn-2"><mml:math id="mml-ieqn-2"><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula>, and the worker can wait for <inline-formula id="ieqn-3"><mml:math id="mml-ieqn-3"><mml:msub><mml:mi>W</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>W</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> to allocate the task at most. <inline-formula id="ieqn-4"><mml:math id="mml-ieqn-4"><mml:msub><mml:mi>R</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>W</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> represents the maximum distance radius between the starting coordinate of the worker accepting the task and <inline-formula id="ieqn-5"><mml:math id="mml-ieqn-5"><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula>. The worker will be marked as invalid while performing the task and will be reset to valid when the task is complete.</p>
<p>Let <inline-formula id="ieqn-6"><mml:math id="mml-ieqn-6"><mml:mrow><mml:mtext>T</mml:mtext></mml:mrow><mml:mo>=&#x003C;</mml:mo><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>S</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>D</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>S</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>D</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub><mml:mo>&#x003E;</mml:mo></mml:math></inline-formula>, denote the task/request that appears on the platform, the start coordinate of the task is <inline-formula id="ieqn-7"><mml:math id="mml-ieqn-7"><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>S</mml:mi><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula>, the target coordinate is <inline-formula id="ieqn-8"><mml:math id="mml-ieqn-8"><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>D</mml:mi><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula>, the expected start time of the task is <inline-formula id="ieqn-9"><mml:math id="mml-ieqn-9"><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>S</mml:mi><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula>, and the expected arrival time is <inline-formula id="ieqn-10"><mml:math id="mml-ieqn-10"><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>D</mml:mi><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula>. The task is completed only when the worker picks up the request on <inline-formula id="ieqn-11"><mml:math id="mml-ieqn-11"><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>S</mml:mi><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> and successfully delivers it at a coordinate <inline-formula id="ieqn-12"><mml:math id="mml-ieqn-12"><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>D</mml:mi><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula>.</p>
<p>Unlike commercial crowdsourcing, a spatio-temporal crowdsourcing task is considered successful even if the workers take the order after the estimated start time or delays the completion after the estimated completion time, as long as the worker completes the two actions of pick and deliver. However, both late pickup and late delivery will affect the quality of task completion, reducing both requesters&#x2019; satisfaction and workers&#x2019; reward. Therefore, this paper studies the case of pickup before the expected time. Therefore, the task will only be completed if the worker picks up the request on <inline-formula id="ieqn-13"><mml:math id="mml-ieqn-13"><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>S</mml:mi><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> before <inline-formula id="ieqn-14"><mml:math id="mml-ieqn-14"><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>S</mml:mi><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> and successfully delivers it at a coordinate <inline-formula id="ieqn-15"><mml:math id="mml-ieqn-15"><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>D</mml:mi><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> before <inline-formula id="ieqn-16"><mml:math id="mml-ieqn-16"><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>D</mml:mi><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula>.</p>
<p>In summary, tasks that conform to the following constraints are optional tasks for <inline-formula id="ieqn-17"><mml:math id="mml-ieqn-17"><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>:
<list list-type="bullet">
<list-item>
<p>The task is within the acceptable range of the worker, i.e., <inline-formula id="ieqn-18"><mml:math id="mml-ieqn-18"><mml:mrow><mml:mtext>distance&#xA0;</mml:mtext></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mrow><mml:mtext>L</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mtext>S</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>T</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mtext>L</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mtext>w</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>i</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x003C;</mml:mo><mml:msub><mml:mrow><mml:mtext>R</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mtext>w</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>i</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula>.</p></list-item>
<list-item>
<p>The task appears before the worker leaves, i.e., <inline-formula id="ieqn-19"><mml:math id="mml-ieqn-19"><mml:mrow><mml:mtext>i</mml:mtext></mml:mrow><mml:mo>+</mml:mo><mml:msub><mml:mrow><mml:mtext>W</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mtext>w</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>i</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub><mml:mo>&#x003E;=</mml:mo><mml:msub><mml:mrow><mml:mtext>t</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mtext>S</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>T</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula>. Workers need to arrive at the starting point to pick up the task before the expected start time of the task, i.e., <inline-formula id="ieqn-20"><mml:math id="mml-ieqn-20"><mml:mi>C</mml:mi><mml:mi>o</mml:mi><mml:mi>s</mml:mi><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mi>i</mml:mi><mml:mi>m</mml:mi><mml:mi>e</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>S</mml:mi><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>W</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x003C;=</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>S</mml:mi><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula>.</p></list-item>
</list></p>
</sec>
</sec>
<sec id="s4">
<label>4</label>
<title>System Overview</title>
<p>We model the task scheduling problem in spatio-temporal crowdsourcing as a reinforcement learning problem. When a spatio-temporal crowdsourcing platform (broker) interacts with requesters and workers (environment), the requester influences the pool of available tasks in the broker by setting the start and end dates and start and end coordinates of tasks, and getting the results of the tasks after the completion of each task. The agent recommends tasks to future workers, and workers influence the agent through the completion of the tasks. Following the end-to-end MDP setting, since workers and requestors have different optimization goals, we use two Markov decision processes (MDP) to optimize workers and requestors separately, and finally combine them together for simultaneous optimization.</p>
<p>For a coming worker <inline-formula id="ieqn-21"><mml:math id="mml-ieqn-21"><mml:msub><mml:mi mathvariant="bold-italic">w</mml:mi><mml:mrow><mml:mi mathvariant="bold-italic">i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> at timestamp i, his available tasks <inline-formula id="ieqn-22"><mml:math id="mml-ieqn-22"><mml:msub><mml:mi mathvariant="bold-italic">T</mml:mi><mml:mrow><mml:mi mathvariant="bold-italic">i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> are obtained. The features of <inline-formula id="ieqn-23"><mml:math id="mml-ieqn-23"><mml:msub><mml:mi mathvariant="bold-italic">w</mml:mi><mml:mrow><mml:mi mathvariant="bold-italic">i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> and all tasks in <inline-formula id="ieqn-24"><mml:math id="mml-ieqn-24"><mml:msub><mml:mi mathvariant="bold-italic">T</mml:mi><mml:mrow><mml:mi mathvariant="bold-italic">i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> constitute the state feature <inline-formula id="ieqn-25"><mml:math id="mml-ieqn-25"><mml:msub><mml:mi mathvariant="bold-italic">f</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">s</mml:mi><mml:mrow><mml:mi mathvariant="bold-italic">i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula>. With <inline-formula id="ieqn-26"><mml:math id="mml-ieqn-26"><mml:msub><mml:mi mathvariant="bold-italic">f</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">s</mml:mi><mml:mrow><mml:mi mathvariant="bold-italic">i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula>, Q-Network(W) and Q-Network(R) compute the Q-values that are aggregated to determine the action <inline-formula id="ieqn-27"><mml:math id="mml-ieqn-27"><mml:msub><mml:mi mathvariant="bold-italic">a</mml:mi><mml:mrow><mml:mi mathvariant="bold-italic">i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> . <inline-formula id="ieqn-28"><mml:math id="mml-ieqn-28"><mml:msub><mml:mi mathvariant="bold-italic">a</mml:mi><mml:mrow><mml:mi mathvariant="bold-italic">i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> recommends a sorted list of tasks and the worker selects one accordingly. The action&#x2019; s feedback as well as corresponding reward <inline-formula id="ieqn-29"><mml:math id="mml-ieqn-29"><mml:msub><mml:mi mathvariant="bold-italic">r</mml:mi><mml:mrow><mml:mi mathvariant="bold-italic">i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> are passed to the predictor for future state estimation <inline-formula id="ieqn-30"><mml:math id="mml-ieqn-30"><mml:msub><mml:mi mathvariant="bold-italic">s</mml:mi><mml:mrow><mml:mi mathvariant="bold-italic">i</mml:mi><mml:mo>+</mml:mo><mml:mrow><mml:mn mathvariant="bold">1</mml:mn></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula>. Finally, <inline-formula id="ieqn-31"><mml:math id="mml-ieqn-31"><mml:msub><mml:mi mathvariant="bold-italic">r</mml:mi><mml:mrow><mml:mi mathvariant="bold-italic">i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> , <inline-formula id="ieqn-32"><mml:math id="mml-ieqn-32"><mml:msub><mml:mi mathvariant="bold-italic">a</mml:mi><mml:mrow><mml:mi mathvariant="bold-italic">i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> , <inline-formula id="ieqn-33"><mml:math id="mml-ieqn-33"><mml:msub><mml:mi mathvariant="bold-italic">s</mml:mi><mml:mrow><mml:mi mathvariant="bold-italic">i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-34"><mml:math id="mml-ieqn-34"><mml:msub><mml:mi mathvariant="bold-italic">s</mml:mi><mml:mrow><mml:mi mathvariant="bold-italic">i</mml:mi><mml:mo>+</mml:mo><mml:mrow><mml:mn mathvariant="bold">1</mml:mn></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> are stored in the Memory for the Learner to train Q-Networks continuously.</p>
<p><xref ref-type="fig" rid="fig-1">Fig. 1</xref> illustrates the system framework of the overall process. The timestamp <inline-formula id="ieqn-35"><mml:math id="mml-ieqn-35"><mml:mi>i</mml:mi></mml:math></inline-formula> requestor release tasks on the platform. At the moment, a driver <inline-formula id="ieqn-36"><mml:math id="mml-ieqn-36"><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> online platform (Step 1) at the moment. We picked out a set of drivers from the task pools available task <inline-formula id="ieqn-37"><mml:math id="mml-ieqn-37"><mml:msup><mml:mi>T</mml:mi><mml:mrow><mml:mo>&#x2217;</mml:mo></mml:mrow></mml:msup></mml:math></inline-formula>, <inline-formula id="ieqn-38"><mml:math id="mml-ieqn-38"><mml:msup><mml:mi>T</mml:mi><mml:mrow><mml:mo>&#x2217;</mml:mo></mml:mrow></mml:msup></mml:math></inline-formula> length less than or equal to <inline-formula id="ieqn-39"><mml:math id="mml-ieqn-39"><mml:mi>m</mml:mi><mml:mi>a</mml:mi><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, these tasks are according to the above constraints (Step 2). Driver <inline-formula id="ieqn-40"><mml:math id="mml-ieqn-40"><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> features and <inline-formula id="ieqn-41"><mml:math id="mml-ieqn-41"><mml:msup><mml:mi>T</mml:mi><mml:mrow><mml:mo>&#x2217;</mml:mo></mml:mrow></mml:msup></mml:math></inline-formula> ailable in the task of location information by geohash and one-hot coding or word2vec embedding coding, generation task pool each task <inline-formula id="ieqn-42"><mml:math id="mml-ieqn-42"><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> vector feature vector and the driver&#x2019;s <inline-formula id="ieqn-43"><mml:math id="mml-ieqn-43"><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula>, connection alignment padding constitutes <inline-formula id="ieqn-44"><mml:math id="mml-ieqn-44"><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula>, which represents the platform of state vector (Step 3). Then, we input <inline-formula id="ieqn-45"><mml:math id="mml-ieqn-45"><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> into two D3SQN-networks, namely D3SQN-Networks (W) and D3SQN-networks (R). considering the worker&#x0027;s benefit <inline-formula id="ieqn-46"><mml:math id="mml-ieqn-46"><mml:msub><mml:mi>Q</mml:mi><mml:mrow><mml:mi>w</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>a</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> and the requester&#x2019;s benefit <inline-formula id="ieqn-47"><mml:math id="mml-ieqn-47"><mml:msub><mml:mi>Q</mml:mi><mml:mrow><mml:mi>r</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>a</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>, respectively, and predict the Q-value of each possible action <inline-formula id="ieqn-48"><mml:math id="mml-ieqn-48"><mml:msub><mml:mi>a</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> in state <inline-formula id="ieqn-49"><mml:math id="mml-ieqn-49"><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> where <inline-formula id="ieqn-50"><mml:math id="mml-ieqn-50"><mml:msub><mml:mi>a</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> represents which task the user chooses. Each network outputs a vector of <inline-formula id="ieqn-51"><mml:math id="mml-ieqn-51"><mml:mi>m</mml:mi><mml:mi>a</mml:mi><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msub><mml:mo>&#x00D7;</mml:mo><mml:mn>1</mml:mn></mml:math></inline-formula>, where each value represents the <inline-formula id="ieqn-52"><mml:math id="mml-ieqn-52"><mml:mi>Q</mml:mi></mml:math></inline-formula> value of the corresponding task <inline-formula id="ieqn-53"><mml:math id="mml-ieqn-53"><mml:mi>T</mml:mi></mml:math></inline-formula> (Step 4). According to the aggregator will two network output <inline-formula id="ieqn-54"><mml:math id="mml-ieqn-54"><mml:msub><mml:mi>Q</mml:mi><mml:mrow><mml:mi>w</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>a</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> and <inline-formula id="ieqn-55"><mml:math id="mml-ieqn-55"><mml:msub><mml:mi>Q</mml:mi><mml:mrow><mml:mi>r</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>a</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> polymerization to produce the final task list <inline-formula id="ieqn-56"><mml:math id="mml-ieqn-56"><mml:msub><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mo fence="false" stretchy="false">{</mml:mo><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mi>j</mml:mi><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mi>j</mml:mi><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mi>j</mml:mi><mml:mn>3</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo fence="false" stretchy="false">}</mml:mo></mml:math></inline-formula>, the aggregator performs a weighted calculation of the two network output values to balance the workers and requesters, with a default value of 0.5. Corresponding <inline-formula id="ieqn-57"><mml:math id="mml-ieqn-57"><mml:mi>Q</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mrow><mml:msup><mml:mi>j</mml:mi><mml:mrow><mml:mo>&#x2217;</mml:mo></mml:mrow></mml:msup></mml:mrow></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>, <inline-formula id="ieqn-58"><mml:math id="mml-ieqn-58"><mml:mo>&#x2217;</mml:mo></mml:math></inline-formula> in descending order aggregation scores <inline-formula id="ieqn-59"><mml:math id="mml-ieqn-59"><mml:mi>Q</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mrow><mml:msup><mml:mi>j</mml:mi><mml:mrow><mml:mo>&#x2217;</mml:mo></mml:mrow></mml:msup></mml:mrow></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> as a recommended list of tasks for the drivers (step 5). When <inline-formula id="ieqn-60"><mml:math id="mml-ieqn-60"><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> sees a sorted task list, we assume that the driver follows a cascade model to view the task list and complete the first task of interest, which is the task with the largest Q-value as the driver&#x2019;s <inline-formula id="ieqn-61"><mml:math id="mml-ieqn-61"><mml:msub><mml:mi>a</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> at the moment. The feedback is the completed tasks and the unfinished tasks suggested to <inline-formula id="ieqn-62"><mml:math id="mml-ieqn-62"><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> (Step 6). Then through quantitative feedback rewarded <inline-formula id="ieqn-63"><mml:math id="mml-ieqn-63"><mml:msub><mml:mi>r</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> (Step 7). The predictor postulates that the outcome of D3SQN determines the ultimate action of the present worker. Subsequently, it constructs the following state, assuming the current sequence is chosen, by aggregating the set of available tasks for the next incoming worker and the relevant details of the next worker, culminating in the formation of the resultant future state <inline-formula id="ieqn-64"><mml:math id="mml-ieqn-64"><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> (step 8). Then, we will be successful tuples <inline-formula id="ieqn-65"><mml:math id="mml-ieqn-65"><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>a</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mrow><mml:msub><mml:mi>r</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>,</mml:mo><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> in the memory pool, the <inline-formula id="ieqn-66"><mml:math id="mml-ieqn-66"><mml:msub><mml:mi>a</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is to complete tasks, and the memory pool is used for storing training data (Step 9). Each time we store an extra tuple into the memory pool, we use learners to update the parameters of the two D3SQN Networks, obtain good estimates of <inline-formula id="ieqn-67"><mml:math id="mml-ieqn-67"><mml:msub><mml:mi>Q</mml:mi><mml:mrow><mml:mi>w</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>a</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> and <inline-formula id="ieqn-68"><mml:math id="mml-ieqn-68"><mml:msub><mml:mi>Q</mml:mi><mml:mrow><mml:mi>r</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>a</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>, and derive the optimal strategy (Step 10).</p>
<fig id="fig-1">
<label>Figure 1</label>
<caption>
<title>System framework</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_31350-fig-1.tif"/>
</fig>
<p>In the next section, we will introduce the characteristics of these parts of the system in detail. The model is introduced in <xref ref-type="sec" rid="s5">Section 5</xref>.</p>
</sec>
<sec id="s5">
<label>5</label>
<title>Feature Construction</title>
<p>For Task <inline-formula id="ieqn-69"><mml:math id="mml-ieqn-69"><mml:msub><mml:mrow><mml:mtext>T</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>j</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>=&#x003C;</mml:mo><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>S</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>D</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>S</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>D</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub><mml:mo>&#x003E;</mml:mo></mml:math></inline-formula>, we use <inline-formula id="ieqn-70"><mml:math id="mml-ieqn-70"><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>V</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>S</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>V</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>D</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>r</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mrow><mml:mi mathvariant="normal">&#x0394;</mml:mi></mml:mrow><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>d</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>G</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> represents the eigenvector of task j. Where <inline-formula id="ieqn-71"><mml:math id="mml-ieqn-71"><mml:msub><mml:mi>V</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>S</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-72"><mml:math id="mml-ieqn-72"><mml:msub><mml:mi>V</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>D</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> is the GeoHash encoded one pot vector of the starting coordinate and destination coordinate, <inline-formula id="ieqn-73"><mml:math id="mml-ieqn-73"><mml:mrow><mml:mi mathvariant="normal">&#x0394;</mml:mi></mml:mrow><mml:mi>t</mml:mi></mml:math></inline-formula> is the estimated time consumed by the task, <inline-formula id="ieqn-74"><mml:math id="mml-ieqn-74"><mml:msub><mml:mi>G</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> is the estimated fuel cost consumed, and <inline-formula id="ieqn-75"><mml:math id="mml-ieqn-75"><mml:msub><mml:mi>d</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> is the distance span of the whole task. The formal expression is as follows:</p>
<p><disp-formula id="ueqn-1"><mml:math id="mml-ueqn-1" display="block"><mml:msub><mml:mi>V</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>S</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mi>O</mml:mi><mml:mi>n</mml:mi><mml:mi>e</mml:mi><mml:mi>H</mml:mi><mml:mi>o</mml:mi><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>G</mml:mi><mml:mi>e</mml:mi><mml:mi>o</mml:mi><mml:mi>H</mml:mi><mml:mi>a</mml:mi><mml:mi>s</mml:mi><mml:mi>h</mml:mi></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>S</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:msub><mml:mi>V</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>D</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mi>O</mml:mi><mml:mi>n</mml:mi><mml:mi>e</mml:mi><mml:mi>H</mml:mi><mml:mi>o</mml:mi><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>G</mml:mi><mml:mi>e</mml:mi><mml:mi>o</mml:mi><mml:mi>H</mml:mi><mml:mi>a</mml:mi><mml:mi>s</mml:mi><mml:mi>h</mml:mi></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>D</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:mi mathvariant="normal">&#x0394;</mml:mi></mml:mrow><mml:mi>t</mml:mi><mml:mo>=</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>D</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>S</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub></mml:math></disp-formula></p>
<p><disp-formula id="eqn-1"><label>(1)</label><mml:math id="mml-eqn-1" display="block"><mml:msub><mml:mi>d</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mi>d</mml:mi><mml:mi>i</mml:mi><mml:mi>s</mml:mi><mml:mi>t</mml:mi><mml:mi>a</mml:mi><mml:mi>n</mml:mi><mml:mi>c</mml:mi><mml:mi>e</mml:mi></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>S</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>D</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mi>G</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>d</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub><mml:mo>&#x2217;</mml:mo><mml:mi>G</mml:mi></mml:math></disp-formula></p>
<p>where G is the average fuel consumption per kilometer.</p>
<p>For worker <inline-formula id="ieqn-76"><mml:math id="mml-ieqn-76"><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>=&#x003C;</mml:mo><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>R</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>W</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub><mml:mo>&#x003E;</mml:mo></mml:math></inline-formula>, we use <inline-formula id="ieqn-77"><mml:math id="mml-ieqn-77"><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub><mml:mo>=&#x003C;</mml:mo><mml:msub><mml:mi>V</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>h</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub><mml:mo>&#x003E;</mml:mo></mml:math></inline-formula> to represent the feature vector of the worker that goes online on the platform at time <inline-formula id="ieqn-78"><mml:math id="mml-ieqn-78"><mml:mi>i</mml:mi></mml:math></inline-formula>.</p>
<p>where <inline-formula id="ieqn-79"><mml:math id="mml-ieqn-79"><mml:msub><mml:mi>V</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> is the one-hot vector encoded by the GeoHash of the coordinates of worker <inline-formula id="ieqn-80"><mml:math id="mml-ieqn-80"><mml:mi>i</mml:mi></mml:math></inline-formula>, and <inline-formula id="ieqn-81"><mml:math id="mml-ieqn-81"><mml:msub><mml:mi>h</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> is the feature vector composed of the historical completed tasks of worker <inline-formula id="ieqn-82"><mml:math id="mml-ieqn-82"><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>. We believe that each worker has its own task preferences, which are often reflected in his past task history. For example, some drivers prefer long journeys. The ratio of time consuming to distance length usually reflects some attributes of the task. For example, long time consuming but short distance is usually in the urban road section with heavy traffic; short time consuming but long distance is usually in the suburban road section of the city, etc. To sum up, we combine the historical task feature vector of workers with the current position vector of workers to represent the feature vector of workers, so as to better capture the preference of workers and the relationship between tasks. The formal expression is as follows:
<disp-formula id="eqn-2"><label>(2)</label><mml:math id="mml-eqn-2" display="block"><mml:msub><mml:mi>V</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mi>O</mml:mi><mml:mi>n</mml:mi><mml:mi>e</mml:mi><mml:mi>H</mml:mi><mml:mi>o</mml:mi><mml:mi>t</mml:mi></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>G</mml:mi><mml:mi>e</mml:mi><mml:mi>o</mml:mi><mml:mi>H</mml:mi><mml:mi>a</mml:mi><mml:mi>s</mml:mi><mml:mi>h</mml:mi></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mi>h</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mrow><mml:msup><mml:mi>f</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msup></mml:mrow><mml:mrow><mml:mrow><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub><mml:mo>&#x22C5;</mml:mo><mml:msub><mml:mi>&#x03B3;</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mrow><mml:msup><mml:mi>f</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:mrow><mml:mrow><mml:mrow><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x22C5;</mml:mo><mml:msub><mml:mi>&#x03B3;</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>&#x2026;</mml:mo><mml:mo>+</mml:mo><mml:msub><mml:mrow><mml:msup><mml:mi>f</mml:mi><mml:mrow><mml:mrow><mml:mi>m</mml:mi><mml:mi>a</mml:mi><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup></mml:mrow><mml:mrow><mml:mrow><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x22C5;</mml:mo><mml:msub><mml:mi>&#x03B3;</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mrow><mml:msup><mml:mi>f</mml:mi><mml:mrow><mml:mrow><mml:mi>m</mml:mi><mml:mi>a</mml:mi><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msup></mml:mrow><mml:mrow><mml:mrow><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub></mml:math></disp-formula>where <inline-formula id="ieqn-83"><mml:math id="mml-ieqn-83"><mml:msub><mml:mi>&#x03B3;</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>&#x2208;</mml:mo><mml:mo stretchy="false">[</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mn>1</mml:mn><mml:mo stretchy="false">]</mml:mo></mml:math></inline-formula> is the attenuation factor. In order to capture the long-term and short-term preferences of workers, we attenuated the historical task information so that the short-term interest accounted for more and the long-term preferences of workers could also be reflected.</p>
<p>When the <inline-formula id="ieqn-84"><mml:math id="mml-ieqn-84"><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> online platform seeks for task orders at time <inline-formula id="ieqn-85"><mml:math id="mml-ieqn-85"><mml:mi>i</mml:mi></mml:math></inline-formula>, the system will select the available task list for the <inline-formula id="ieqn-86"><mml:math id="mml-ieqn-86"><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> according to the constraints in <xref ref-type="sec" rid="s3_2">Section 3.2</xref>. Then, the feature vectors of the tasks in the available list and the feature vectors of the workers are concatenated to obtain the feature representation vector <inline-formula id="ieqn-87"><mml:math id="mml-ieqn-87"><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> of the state at time <inline-formula id="ieqn-88"><mml:math id="mml-ieqn-88"><mml:mi>i</mml:mi></mml:math></inline-formula>. Since the number of available tasks varies at different timestamps, we set the maximum number of available tasks <inline-formula id="ieqn-89"><mml:math id="mml-ieqn-89"><mml:mi>m</mml:mi><mml:mi>a</mml:mi><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> and use zero padding, that is, adding zero to the end of <inline-formula id="ieqn-90"><mml:math id="mml-ieqn-90"><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula>, so that each <inline-formula id="ieqn-91"><mml:math id="mml-ieqn-91"><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> is denoted as: <inline-formula id="ieqn-92"><mml:math id="mml-ieqn-92"><mml:mo stretchy="false">[</mml:mo><mml:mi>m</mml:mi><mml:mi>a</mml:mi><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mi>j</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mo>+</mml:mo><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mo stretchy="false">]</mml:mo></mml:math></inline-formula>, where <inline-formula id="ieqn-93"><mml:math id="mml-ieqn-93"><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mi>j</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2286;</mml:mo><mml:mrow><mml:mi mathvariant="normal">&#x03C3;</mml:mi></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>, which <inline-formula id="ieqn-94"><mml:math id="mml-ieqn-94"><mml:mrow><mml:mi mathvariant="normal">&#x03C3;</mml:mi></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> is <inline-formula id="ieqn-95"><mml:math id="mml-ieqn-95"><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> available task set.</p>
</sec>
<sec id="s6">
<label>6</label>
<title>Double Dueling Deep Spatial Q Network (D3SQN)</title>
<p>In this section, we introduce our proposed Q-learning-based deep reinforcement learning network model specifically for spatio-temporal crowdsourcing.</p>
<p>As shown in <xref ref-type="fig" rid="fig-2">Fig. 2</xref>, based on the system state vector <inline-formula id="ieqn-96"><mml:math id="mml-ieqn-96"><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> at time <inline-formula id="ieqn-97"><mml:math id="mml-ieqn-97"><mml:mi>i</mml:mi></mml:math></inline-formula>, the optimal allocation of spatiotemporal tasks is achieved through D3SQN. Specifically, this model uses the architecture of Dueling Network and two Spatial State Transformers to respectively predict the State Q-value and the dominant feature vector composed of the advantages of each action. Finally, <inline-formula id="ieqn-98"><mml:math id="mml-ieqn-98"><mml:msub><mml:mi>V</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-99"><mml:math id="mml-ieqn-99"><mml:msub><mml:mi>V</mml:mi><mml:mrow><mml:mi>a</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> are spliced into the final action value vector <inline-formula id="ieqn-100"><mml:math id="mml-ieqn-100"><mml:msub><mml:mi>V</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mrow><mml:mi>a</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula>.</p>
<fig id="fig-2">
<label>Figure 2</label>
<caption>
<title>Double dueling deep spatial Q network (D3SQN) framework and details. (a) shows the framework of D3SQN. The network architecture is divided into two parts: the upper spatial state transformer used to predict the action advantage value <inline-formula id="ieqn-101"><mml:math id="mml-ieqn-101"><mml:msub><mml:mi>V</mml:mi><mml:mrow><mml:mi>a</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> and the lower spatial state transformer used to predict the state value <inline-formula id="ieqn-102"><mml:math id="mml-ieqn-102"><mml:msub><mml:mi>Q</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula>. (b) shows the specific details of the encoder and decoder of the spatial state transformer</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_31350-fig-2a.tif"/><graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_31350-fig-2b.tif"/>
</fig>
<p>Different from the traditional DQN, the states that need to be processed by our proposed D3SQN are significantly different from the traditional DQN network under spatio-temporal crowdsourcing.</p>
<p>Most of the traditional DQN, including its variants, are suitable for the input of temporal or sequentially related feature vectors. For example, in a game scenario, the representation of state may be a vector matrix of continuous pictures, and the sorting of these pictures and the distribution of pixels inside the pictures are fixed. Once the distribution of these pixels changes, the prediction results of the network will be affected. However, in our spatial crowdsourcing environment, the representation of the state is the connection of a string of available lists and worker feature vectors, which is a set feature. The sorting order of these available lists will not affect the result of the final predicted action. It is difficult for traditional DQN and DQN variants to mine and learn the complex and dynamic spatio-temporal correlations in such permutation invariant inputs. Because this permutation invariant set feature is a global feature, the ability of CNN or traditional neural network architecture to extract local structural features limits the extraction ability of such global high-dimensional features. The spatial state Transformer is better able to handle high-dimensional and global information, mainly because the structure of the Transformer itself is better suited to extract this information. Therefore, this paper proposes Spatial State Transformer, a new deep reinforcement learning network model for spatial-temporal crowdsourcing. It will not affect the result of prediction even if he faces the state vector with a changed order.</p>
<p>We use the structure of Double DQN and Dueling Network to alleviate the overestimation problem in Q-learning. Next, we will introduce the details and functions of Dueling Network and Spatial State Transformer, respectively.</p>
<sec id="s6_1">
<label>6.1</label>
<title>Dueling Network</title>
<p>When we recommend a task set to a worker, the tasks in the optional task set are all the tasks that the worker may take orders from next, and each task represents an action. When there are several actions with multiple redundant or approximately equal Q-values, the original network will prefer to choose the first task, but in fact, different actions will lead to different states, and we may not be able to jump out of the optimal action under the current state, so the structure of Dueling network is adopted. Instead of relying solely on the value of the action, the state can make a separate value prediction. The performance of the network model is better.</p>
<p>Using the structure of Dueling Network, the original direct predictive action value is divided into predictive state value and dominant value of each action. This effectively avoids the original over-estimation problem of Q-learning. The traditional network of DQN is approximate <inline-formula id="ieqn-103"><mml:math id="mml-ieqn-103"><mml:msup><mml:mi>Q</mml:mi><mml:mrow><mml:mo>&#x2217;</mml:mo></mml:mrow></mml:msup><mml:mo stretchy="false">(</mml:mo><mml:mi>s</mml:mi><mml:mo>,</mml:mo><mml:mi>a</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mi>m</mml:mi><mml:mi>a</mml:mi><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>&#x03C0;</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mi>Q</mml:mi><mml:mrow><mml:mi>&#x03C0;</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>s</mml:mi><mml:mo>,</mml:mo><mml:mi>a</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> by network, which represents the value function of optimal action <inline-formula id="ieqn-104"><mml:math id="mml-ieqn-104"><mml:mi>a</mml:mi></mml:math></inline-formula> under strategy <inline-formula id="ieqn-105"><mml:math id="mml-ieqn-105"><mml:mi>&#x03C0;</mml:mi></mml:math></inline-formula> and state <inline-formula id="ieqn-106"><mml:math id="mml-ieqn-106"><mml:mi>s</mml:mi></mml:math></inline-formula>. The architecture of Dueling Network is to separate State from action to a certain extent, which is formalized as follows:
<disp-formula id="ueqn-4"><mml:math id="mml-ueqn-4" display="block"><mml:msup><mml:mrow><mml:mtext>Q</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mo>&#x2217;</mml:mo></mml:mrow></mml:mrow></mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext>s</mml:mtext></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:mtext>a</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:msup><mml:mrow><mml:mtext>A</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mo>&#x2217;</mml:mo></mml:mrow></mml:mrow></mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext>s</mml:mtext></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:mtext>a</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>+</mml:mo><mml:msup><mml:mrow><mml:mtext>V</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mo>&#x2217;</mml:mo></mml:mrow></mml:mrow></mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext>s</mml:mtext></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:mtext>a</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mspace width="negativethinmathspace" /><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mtext>V</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mo>&#x2217;</mml:mo></mml:mrow></mml:mrow></mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext>s</mml:mtext></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:mtext>a</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:munder><mml:mo movablelimits="true" form="prefix">max</mml:mo><mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x03C0;</mml:mi></mml:mrow></mml:mrow></mml:munder><mml:msub><mml:mrow><mml:mtext>V</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x03C0;</mml:mi></mml:mrow></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext>s</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:munder><mml:mo movablelimits="true" form="prefix">max</mml:mo><mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x03C0;</mml:mi></mml:mrow></mml:mrow></mml:munder><mml:msup><mml:mrow><mml:mtext>Q</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mo>&#x2217;</mml:mo></mml:mrow></mml:mrow></mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext>s</mml:mtext></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:mtext>a</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>,</mml:mo></mml:math></disp-formula>
<disp-formula id="eqn-3"><label>(3)</label><mml:math id="mml-eqn-3" display="block"><mml:msup><mml:mi>A</mml:mi><mml:mrow><mml:mo>&#x2217;</mml:mo></mml:mrow></mml:msup><mml:mo stretchy="false">(</mml:mo><mml:mi>s</mml:mi><mml:mo>,</mml:mo><mml:mi>a</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:msup><mml:mi>Q</mml:mi><mml:mrow><mml:mo>&#x2217;</mml:mo></mml:mrow></mml:msup><mml:mo stretchy="false">(</mml:mo><mml:mi>s</mml:mi><mml:mo>,</mml:mo><mml:mi>a</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x2212;</mml:mo><mml:msup><mml:mi>V</mml:mi><mml:mrow><mml:mo>&#x2217;</mml:mo></mml:mrow></mml:msup><mml:mo stretchy="false">(</mml:mo><mml:mi>s</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo><mml:munder><mml:mo movablelimits="true" form="prefix">max</mml:mo><mml:mrow><mml:mrow><mml:mtext>a</mml:mtext></mml:mrow></mml:mrow></mml:munder><mml:msup><mml:mi>A</mml:mi><mml:mrow><mml:mo>&#x2217;</mml:mo></mml:mrow></mml:msup><mml:mo stretchy="false">(</mml:mo><mml:mi>s</mml:mi><mml:mo>,</mml:mo><mml:mi>a</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:munder><mml:mo movablelimits="true" form="prefix">max</mml:mo><mml:mrow><mml:mrow><mml:mtext>a</mml:mtext></mml:mrow></mml:mrow></mml:munder><mml:msup><mml:mi>Q</mml:mi><mml:mrow><mml:mo>&#x2217;</mml:mo></mml:mrow></mml:msup><mml:mo stretchy="false">(</mml:mo><mml:mi>s</mml:mi><mml:mo>,</mml:mo><mml:mi>a</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x2212;</mml:mo><mml:msup><mml:mi>V</mml:mi><mml:mrow><mml:mo>&#x2217;</mml:mo></mml:mrow></mml:msup><mml:mo stretchy="false">(</mml:mo><mml:mi>s</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mn>0</mml:mn></mml:math></disp-formula></p>
<p>As can be seen in <xref ref-type="fig" rid="fig-2">Fig. 2a</xref>, the Dueling Network in our proposed D3SQN consists of two Spatial State Transformer, which are formalized as follows:
<disp-formula id="eqn-4"><label>(4)</label><mml:math id="mml-eqn-4" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:msup><mml:mi>A</mml:mi><mml:mrow><mml:mo>&#x2217;</mml:mo></mml:mrow></mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:mi>s</mml:mi><mml:mo>,</mml:mo><mml:mi>a</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mi>S</mml:mi><mml:mi>p</mml:mi><mml:mi>a</mml:mi><mml:mi>t</mml:mi><mml:mi>i</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi><mml:mi>S</mml:mi><mml:mi>t</mml:mi><mml:mi>a</mml:mi><mml:mi>t</mml:mi><mml:mi>e</mml:mi><mml:mi>T</mml:mi><mml:mi>r</mml:mi><mml:mi>a</mml:mi><mml:mi>n</mml:mi><mml:mi>s</mml:mi><mml:mi>f</mml:mi><mml:mi>o</mml:mi><mml:mi>r</mml:mi><mml:mi>m</mml:mi><mml:mi>e</mml:mi><mml:mi>r</mml:mi></mml:mrow><mml:mrow><mml:mi>Q</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:msup><mml:mi>V</mml:mi><mml:mrow><mml:mo>&#x2217;</mml:mo></mml:mrow></mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:mi>s</mml:mi><mml:mo>,</mml:mo><mml:mi>a</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mi>S</mml:mi><mml:mi>p</mml:mi><mml:mi>a</mml:mi><mml:mi>t</mml:mi><mml:mi>i</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi><mml:mi>S</mml:mi><mml:mi>t</mml:mi><mml:mi>a</mml:mi><mml:mi>t</mml:mi><mml:mi>e</mml:mi><mml:mi>T</mml:mi><mml:mi>r</mml:mi><mml:mi>a</mml:mi><mml:mi>n</mml:mi><mml:mi>s</mml:mi><mml:mi>f</mml:mi><mml:mi>o</mml:mi><mml:mi>r</mml:mi><mml:mi>m</mml:mi><mml:mi>e</mml:mi><mml:mi>r</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mi>A</mml:mi><mml:mi>d</mml:mi><mml:mi>v</mml:mi><mml:mi>a</mml:mi><mml:mi>n</mml:mi><mml:mi>t</mml:mi><mml:mi>a</mml:mi><mml:mi>g</mml:mi><mml:mi>e</mml:mi></mml:mrow></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd /><mml:mtd><mml:mi></mml:mi><mml:mrow><mml:munder><mml:mo movablelimits="true" form="prefix">max</mml:mo><mml:mrow><mml:mrow><mml:mtext>a</mml:mtext></mml:mrow></mml:mrow></mml:munder></mml:mrow><mml:msup><mml:mi>Q</mml:mi><mml:mrow><mml:mo>&#x2217;</mml:mo></mml:mrow></mml:msup><mml:mo stretchy="false">(</mml:mo><mml:mi>s</mml:mi><mml:mo>,</mml:mo><mml:mi>a</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mrow><mml:munder><mml:mo movablelimits="true" form="prefix">max</mml:mo><mml:mrow><mml:mrow><mml:mtext>a</mml:mtext></mml:mrow></mml:mrow></mml:munder></mml:mrow><mml:msup><mml:mi>A</mml:mi><mml:mrow><mml:mo>&#x2217;</mml:mo></mml:mrow></mml:msup><mml:mo stretchy="false">(</mml:mo><mml:mi>s</mml:mi><mml:mo>,</mml:mo><mml:mi>a</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>+</mml:mo><mml:msup><mml:mi>V</mml:mi><mml:mrow><mml:mo>&#x2217;</mml:mo></mml:mrow></mml:msup><mml:mo stretchy="false">(</mml:mo><mml:mi>s</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p>
</sec>
<sec id="s6_2">
<label>6.2</label>
<title>Spatial State Transformer</title>
<p>In this section, we introduce the spatial state Transformer: an attention-based neural network for processing state vectors under spatial crowdsourcing.</p>
<p>As shown in <xref ref-type="fig" rid="fig-2">Fig. 2a</xref>, Spatial State Transformer consists of EnCoder and DeCoder, our motivation for designing the spatial state converter is that self-attention allows the extraction of key global associations in the collection data, as well as spatio-temporal information associations. This makes the spatial state converter a more effective and efficient way to encode the entire collection simultaneously. At the same time, our model needs to extract high-level feature representations in order to use the self-attention mechanism to extract associations between information, which is why our model is designed as Encoder and Decoder architecture. Encoder and Decoder are the attention-based spatial set operation modules defined by us. The main purpose of the encoder is to represent permutation invariant state inputs, encoding independently a set of tasks of available size and each element <inline-formula id="ieqn-107"><mml:math id="mml-ieqn-107"><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula>. The decoder aggregates these encoding features <inline-formula id="ieqn-108"><mml:math id="mml-ieqn-108"><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>m</mml:mi><mml:mrow><mml:mi>s</mml:mi><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> and uses self-attention mechanism to extract associations between high-level representations. All the blocks described here are neural network blocks with their own parameters, not fixed functions.</p>
<p>We made improvements by referring to the structure of SetTransformer, without retaining the original pooling operation, and made improvements in EnCoder and DeCoder modules suitable for reinforcement learning, We tune the layerNorm layer after the feature input, instead of before the feature input as the Transformer usually does and replace the general residual connection with GatingLayer. In this way, the learning and training of Transformer structure in RL can be stabilized. We use this structure of SetTransformer, but do not retain the pooling operation, because the pooling operation in SetTransformer aims to extract the correlation between features from multiple dimensions. However, the state expression of spatio-temporal crowdsourcing is different from 3D data, which is three-dimensional collection data strongly correlated among multiple dimensions. The previous SetTransformer architecture is mainly aimed at such low-latitude replacement and unchanged collection data. What we need to mine is the relationship between the features of each element in each list (including task and worker address information), so it is counterproductive to try to improve the dimension to learn. We only need to use the feature expression proposed by Encoder and use Decoder to aggregate and finally get the prediction result.</p>
<p>In the remainder of this section, we&#x2019;ll describe the details of individual modules.</p>
<p><bold>DeCoder:</bold> As seen in <xref ref-type="fig" rid="fig-2">Fig. 2b</xref>, the Decoder module is designed to extract various aspects of spatio-temporal features <inline-formula id="ieqn-109"><mml:math id="mml-ieqn-109"><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>m</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> obtained from the Encoder through self-attention. By stacking multiple Decoders, we can extract deeper spatio-temporal interaction information from the input set. The decoder structure includes layer normalization, multi-head self-attention, and a gating layer. We also use an identity map reordering technique inspired by GTrxl [<xref ref-type="bibr" rid="ref-18">18</xref>] to facilitate policy optimization, which is used to help with policy optimization by initializing the agent in a manner that is similar to a Markov policy/value function.</p>
<p><bold>Identity Map Reordering:</bold> We applied a modification called Identity Map Reordering to the Spatial Temporal Transformer, inspired by the GTrxl [<xref ref-type="bibr" rid="ref-18">18</xref>] design. This involved rearranging the order of Layer Norm and applying ReLU activation to the output of each submodule before joining with the residual connection. This reordering enables an identity mapping from the input of the Transformer at the first layer to the output at the last layer, which facilitates policy optimization in reinforcement learning. Specifically, this allows the agent to learn reactive actions first before focusing on memory-based behavior. In the spatio-temporal crowdsourcing environment, this also allows the agent to learn to select actions before using memory, even if the experience has not yet entered the memory. This modification improves the efficiency of learning and enables the agent to effectively utilize memory to improve performance.</p>
<p><bold>Gated-Recurrent-Unit-Type Gating:</bold> We use an explicit initialization of the GRU gating mechanism to approach the identity map. Gated recursive unit (GRU) is a recursive network, which behaves like LSTM, but with fewer parameters. The formalization is as follows:</p>
<p><disp-formula id="ueqn-8"><mml:math id="mml-ueqn-8" display="block"><mml:mi>r</mml:mi><mml:mo>=</mml:mo><mml:mi>&#x03C3;</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>W</mml:mi><mml:mrow><mml:mi>r</mml:mi></mml:mrow></mml:msub><mml:mi>y</mml:mi><mml:mo>+</mml:mo><mml:msub><mml:mi>U</mml:mi><mml:mrow><mml:mi>r</mml:mi></mml:mrow></mml:msub><mml:mi>x</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo><mml:mi>z</mml:mi><mml:mo>=</mml:mo><mml:mi>&#x03C3;</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>W</mml:mi><mml:mrow><mml:mi>z</mml:mi></mml:mrow></mml:msub><mml:mi>y</mml:mi><mml:mo>+</mml:mo><mml:msub><mml:mi>U</mml:mi><mml:mrow><mml:mi>z</mml:mi></mml:mrow></mml:msub><mml:mi>x</mml:mi><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>b</mml:mi><mml:mrow><mml:mi>g</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo><mml:mi>h</mml:mi><mml:mo>=</mml:mo><mml:mi>t</mml:mi><mml:mi>a</mml:mi><mml:mi>n</mml:mi><mml:mi>h</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>W</mml:mi><mml:mrow><mml:mi>g</mml:mi></mml:mrow></mml:msub><mml:mi>y</mml:mi><mml:mo>+</mml:mo><mml:msub><mml:mi>U</mml:mi><mml:mrow><mml:mi>g</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>r</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x2217;</mml:mo><mml:mi>x</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></disp-formula></p>
<p><disp-formula id="eqn-5"><label>(5)</label><mml:math id="mml-eqn-5" display="block"><mml:mrow><mml:mi>G</mml:mi><mml:mi>a</mml:mi><mml:mi>t</mml:mi><mml:mi>i</mml:mi><mml:mi>n</mml:mi><mml:mi>g</mml:mi><mml:mi>L</mml:mi><mml:mi>a</mml:mi><mml:mi>y</mml:mi><mml:mi>e</mml:mi><mml:mi>r</mml:mi></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:mi>z</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x2217;</mml:mo><mml:mi>x</mml:mi><mml:mo>+</mml:mo><mml:mi>z</mml:mi><mml:mo>&#x2217;</mml:mo><mml:mi>h</mml:mi></mml:math></disp-formula></p>
<p>where <inline-formula id="ieqn-110"><mml:math id="mml-ieqn-110"><mml:msub><mml:mi>b</mml:mi><mml:mrow><mml:mi>g</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is the bias in the applicable gating layer. Initially setup <inline-formula id="ieqn-111"><mml:math id="mml-ieqn-111"><mml:msub><mml:mi>b</mml:mi><mml:mrow><mml:mi>g</mml:mi></mml:mrow></mml:msub><mml:mo>&#x003E;</mml:mo><mml:mn>0</mml:mn></mml:math></inline-formula> can greatly improve the learning speed.</p>
<p><bold>Encoder:</bold> The Decoder architecture in the Spatial State Transformer eliminates position encoding, enabling the network to extract mutual relationships between sets in the state embedding. However, using the Decoder directly can result in quadratic time complexity, which is not practical for large-scale set-structured datasets. To address this issue, we introduce trainable inducing points I into the Encoder architecture, which reduces the time complexity. An Encoder with m inducing points I is defined as follows:
<disp-formula id="ueqn-14"><mml:math id="mml-ueqn-14" display="block"><mml:msub><mml:mi>G</mml:mi><mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mi>G</mml:mi><mml:mi>a</mml:mi><mml:mi>t</mml:mi><mml:mi>i</mml:mi><mml:mi>n</mml:mi><mml:mi>g</mml:mi><mml:mi>L</mml:mi><mml:mi>a</mml:mi><mml:mi>y</mml:mi><mml:mi>e</mml:mi><mml:mi>r</mml:mi></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>M</mml:mi><mml:mi>u</mml:mi><mml:mi>l</mml:mi><mml:mi>t</mml:mi><mml:mi>i</mml:mi><mml:mi>H</mml:mi><mml:mi>e</mml:mi><mml:mi>a</mml:mi><mml:mi>d</mml:mi></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>I</mml:mi><mml:mo>,</mml:mo><mml:mrow><mml:mi>L</mml:mi><mml:mi>a</mml:mi><mml:mi>y</mml:mi><mml:mi>e</mml:mi><mml:mi>r</mml:mi><mml:mi>N</mml:mi><mml:mi>o</mml:mi><mml:mi>r</mml:mi><mml:mi>m</mml:mi></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>X</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo><mml:mrow><mml:mi>L</mml:mi><mml:mi>a</mml:mi><mml:mi>y</mml:mi><mml:mi>e</mml:mi><mml:mi>r</mml:mi><mml:mi>N</mml:mi><mml:mi>o</mml:mi><mml:mi>r</mml:mi><mml:mi>m</mml:mi></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>X</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>;</mml:mo><mml:mi>w</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo><mml:mi>X</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mi>R</mml:mi><mml:mrow><mml:mi>n</mml:mi><mml:mi>d</mml:mi></mml:mrow></mml:msup></mml:math></disp-formula>
<disp-formula id="ueqn-10"><mml:math id="mml-ueqn-10" display="block"><mml:msub><mml:mi>G</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mi>G</mml:mi><mml:mi>a</mml:mi><mml:mi>t</mml:mi><mml:mi>i</mml:mi><mml:mi>n</mml:mi><mml:mi>g</mml:mi><mml:mi>L</mml:mi><mml:mi>a</mml:mi><mml:mi>y</mml:mi><mml:mi>e</mml:mi><mml:mi>r</mml:mi></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>G</mml:mi><mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mi>r</mml:mi><mml:mi>F</mml:mi><mml:mi>F</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>L</mml:mi><mml:mi>a</mml:mi><mml:mi>y</mml:mi><mml:mi>e</mml:mi><mml:mi>r</mml:mi><mml:mi>N</mml:mi><mml:mi>o</mml:mi><mml:mi>r</mml:mi><mml:mi>m</mml:mi></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>G</mml:mi><mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mi>R</mml:mi><mml:mrow><mml:mi>n</mml:mi><mml:mi>d</mml:mi></mml:mrow></mml:msup></mml:math></disp-formula>
<disp-formula id="ueqn-11"><mml:math id="mml-ueqn-11" display="block"><mml:msub><mml:mi>G</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mi>G</mml:mi><mml:mi>a</mml:mi><mml:mi>t</mml:mi><mml:mi>i</mml:mi><mml:mi>n</mml:mi><mml:mi>g</mml:mi><mml:mi>L</mml:mi><mml:mi>a</mml:mi><mml:mi>y</mml:mi><mml:mi>e</mml:mi><mml:mi>r</mml:mi></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>M</mml:mi><mml:mi>u</mml:mi><mml:mi>l</mml:mi><mml:mi>t</mml:mi><mml:mi>i</mml:mi><mml:mi>H</mml:mi><mml:mi>e</mml:mi><mml:mi>a</mml:mi><mml:mi>d</mml:mi></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>X</mml:mi><mml:mo>,</mml:mo><mml:mrow><mml:mi>L</mml:mi><mml:mi>a</mml:mi><mml:mi>y</mml:mi><mml:mi>e</mml:mi><mml:mi>r</mml:mi><mml:mi>N</mml:mi><mml:mi>o</mml:mi><mml:mi>r</mml:mi><mml:mi>m</mml:mi></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>G</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mspace width="negativethinmathspace" /><mml:mo>,</mml:mo><mml:mrow><mml:mi>L</mml:mi><mml:mi>a</mml:mi><mml:mi>y</mml:mi><mml:mi>e</mml:mi><mml:mi>r</mml:mi><mml:mi>N</mml:mi><mml:mi>o</mml:mi><mml:mi>r</mml:mi><mml:mi>m</mml:mi></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>G</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mspace width="negativethinmathspace" /><mml:mo>;</mml:mo><mml:mi>w</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:msub><mml:mi>G</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mi>R</mml:mi><mml:mrow><mml:mi>n</mml:mi><mml:mi>d</mml:mi></mml:mrow></mml:msup></mml:math></disp-formula>
<disp-formula id="eqn-6"><label>(6)</label><mml:math id="mml-eqn-6" display="block"><mml:msub><mml:mrow><mml:mi>E</mml:mi><mml:mi>n</mml:mi><mml:mi>c</mml:mi><mml:mi>o</mml:mi><mml:mi>d</mml:mi><mml:mi>e</mml:mi><mml:mi>r</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>X</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x003A;</mml:mo><mml:mo>=</mml:mo><mml:mrow><mml:mi>G</mml:mi><mml:mi>a</mml:mi><mml:mi>t</mml:mi><mml:mi>i</mml:mi><mml:mi>n</mml:mi><mml:mi>g</mml:mi><mml:mi>L</mml:mi><mml:mi>a</mml:mi><mml:mi>y</mml:mi><mml:mi>e</mml:mi><mml:mi>r</mml:mi></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>G</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mi>r</mml:mi><mml:mi>F</mml:mi><mml:mi>F</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>L</mml:mi><mml:mi>a</mml:mi><mml:mi>y</mml:mi><mml:mi>e</mml:mi><mml:mi>r</mml:mi><mml:mi>N</mml:mi><mml:mi>o</mml:mi><mml:mi>r</mml:mi><mml:mi>m</mml:mi></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>G</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mi>R</mml:mi><mml:mrow><mml:mi>n</mml:mi><mml:mi>d</mml:mi></mml:mrow></mml:msup></mml:math></disp-formula></p>
<p>Encoder first converts I through multi-head processing of input set, and then enters GRU processing through input set <inline-formula id="ieqn-112"><mml:math id="mml-ieqn-112"><mml:mi>X</mml:mi></mml:math></inline-formula>, and finally generates a set containing <inline-formula id="ieqn-113"><mml:math id="mml-ieqn-113"><mml:mi>N</mml:mi></mml:math></inline-formula> elements.</p>
<p>The Encoder&#x0027;s objective is to extract a high-dimensional representation of the set state while mitigating the computational complexity resulting from the set&#x2019;s size. In this study, the Encoder is responsible for extracting an attention-based spatio-temporal state embedding representation, which serves as a basis for discerning the spatio-temporal relationships between Workers and Requesters. Attention is computed between sets of size m and n, with the Encoder&#x2019;s time complexity being O(mn), a substantial improvement over the quadratic complexity of the Decoder. Both set operations (Encoder and Decoder) are permutation invariant.</p>
<p><bold>Loss Function:</bold> we adopt a recently developed loss function, DQNReg, which is inspired by earlier researches [<xref ref-type="bibr" rid="ref-23">23</xref>&#x2013;<xref ref-type="bibr" rid="ref-27">27</xref>]. The formula for DQNReg is presented below:
<disp-formula id="eqn-7"><label>(7)</label><mml:math id="mml-eqn-7" display="block"><mml:msub><mml:mi>Y</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>r</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:mi>&#x03B3;</mml:mi><mml:mo>&#x2217;</mml:mo><mml:mi>m</mml:mi><mml:mi>a</mml:mi><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>a</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mi>Q</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mi>r</mml:mi><mml:mi>a</mml:mi><mml:mi>g</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mi>a</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo><mml:mi>&#x03B4;</mml:mi><mml:mo>=</mml:mo><mml:mi>Q</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>a</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>Y</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mrow><mml:mi>D</mml:mi><mml:mi>Q</mml:mi><mml:mi>N</mml:mi><mml:mi>R</mml:mi><mml:mi>e</mml:mi><mml:mi>g</mml:mi></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>k</mml:mi><mml:mo>&#x2217;</mml:mo><mml:mi>Q</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>a</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>+</mml:mo><mml:msup><mml:mi>&#x03B4;</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:math></disp-formula></p>
<p>The first term of the loss function introduces regularization by multiplying Q-values with an adjustable weighted term, which helps to alleviate overestimation of Q-values. The second term encourages Q-values to approach target Q-values. DQNReg effectively prevents overestimation by directly regularizing Q-values with a weighted term that is always active. The adjustable weighted term in the loss function is set to 0.1 in the experiment.</p>
</sec>
</sec>
<sec id="s7">
<label>7</label>
<title>Experiment</title>
<sec id="s7_1">
<label>7.1</label>
<title>Experimental Setup</title>
<p><bold>Dataset:</bold> We conducted experiments on two real-world datasets, collected by DiDi Chuxing in Xi&#x2019;an, China, the DiDi dataset released through its GAIA initiative, and the dataset from ELM released by Tianchi. Each piece of DiDi&#x2019;s dataset contains drivers&#x2019; ID, task ID, timestamp, longitude and latitude. From this data, we can get information about each task and worker. For our experiment, we utilized a sample of 80,000 orders that were recorded continuously throughout October 2016 in the didi dataset [<xref ref-type="bibr" rid="ref-28">28</xref>]. We obtained the ELM dataset [<xref ref-type="bibr" rid="ref-29">29</xref>] which contains identical information and used the same methodology to acquire an equivalent amount of data. The ELM and DiDi datasets contain different numbers of workers because the average completion time per order on the ELM platform, which is a food delivery service, is much smaller than that of taxis. Therefore, the difference in the number of workers required for the same number of orders is significant. As shown in <xref ref-type="table" rid="table-1">Table 1</xref>, the ELM dataset of 80,000 orders only has 857 workers, while DiDi has 15,367 workers.</p>
<table-wrap id="table-1">
<label>Table 1</label>
<caption>
<title>Differences in number of workers in the data set</title>
</caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th>Dataset</th>
<th># of worker</th>
</tr>
</thead>
<tbody>
<tr>
<td>Elm</td>
<td>857</td>
</tr>
<tr>
<td>DiDi</td>
<td>15367</td>
</tr>
</tbody>
</table>
</table-wrap>
<p><bold>Settings:</bold> Over time, we will resume the process of staff arrival, task creation, or task expiration. The collected dataset records each timestamp I at which a worker <inline-formula id="ieqn-114"><mml:math id="mml-ieqn-114"><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> starts a task, and we assume that a start corresponds to a worker arrival and that the tasks completed by workers in the dataset are considered interesting. Since we do not know the available task list information that the platform assigns to the <inline-formula id="ieqn-115"><mml:math id="mml-ieqn-115"><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> when it arrives, we cannot use the tasks completed by the <inline-formula id="ieqn-116"><mml:math id="mml-ieqn-116"><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> in the original data set as our target for predicting success. Therefore, when <inline-formula id="ieqn-117"><mml:math id="mml-ieqn-117"><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> arrives, we will use the completed orders of a worker in the real data as the workers&#x0027; preference and match a target task <inline-formula id="ieqn-118"><mml:math id="mml-ieqn-118"><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>t</mml:mi><mml:mi>a</mml:mi><mml:mi>r</mml:mi><mml:mi>g</mml:mi><mml:mi>e</mml:mi><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi>w</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> from the current available event pool based on this preference. Considering the interests of workers, if the recommended task is <inline-formula id="ieqn-119"><mml:math id="mml-ieqn-119"><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>t</mml:mi><mml:mi>a</mml:mi><mml:mi>r</mml:mi><mml:mi>g</mml:mi><mml:mi>e</mml:mi><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi>w</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula>, then the reward/label (for reinforcement/supervised learning) is 1. For the benefit of requesters, the reward/tag is the quality gain of <inline-formula id="ieqn-120"><mml:math id="mml-ieqn-120"><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>. And the final reward will be the sum of the ratio of the workers and requesters, in our experiment, this ratio is split equally between the workers and requesters.</p>
<p><bold>Evaluation Measures:</bold> Considering that we are recommending a task or a list of tasks, we use CR and nDCG-CR as the metrics:</p>
<p><disp-formula id="eqn-8"><label>(8)</label><mml:math id="mml-eqn-8" display="block"><mml:mi>C</mml:mi><mml:mi>R</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:munder><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:munder><mml:msub><mml:mi>R</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mrow><mml:mi>n</mml:mi><mml:mi>u</mml:mi><mml:mi>m</mml:mi><mml:mi>b</mml:mi><mml:mi>e</mml:mi><mml:mi>r</mml:mi></mml:mrow><mml:mspace width="thinmathspace" /><mml:mi>o</mml:mi><mml:mi>f</mml:mi><mml:mspace width="thinmathspace" /><mml:mrow><mml:mi>c</mml:mi><mml:mi>o</mml:mi><mml:mi>m</mml:mi><mml:mi>p</mml:mi><mml:mi>l</mml:mi><mml:mi>e</mml:mi><mml:mi>t</mml:mi><mml:mi>e</mml:mi><mml:mi>d</mml:mi></mml:mrow><mml:mspace width="thinmathspace" /><mml:mrow><mml:mi>t</mml:mi><mml:mi>a</mml:mi><mml:mi>s</mml:mi><mml:mi>k</mml:mi><mml:mi>s</mml:mi></mml:mrow></mml:mrow></mml:mfrac><mml:mo>,</mml:mo><mml:mi>n</mml:mi><mml:mi>D</mml:mi><mml:mi>C</mml:mi><mml:mi>G</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mi>C</mml:mi><mml:mi>R</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:munder><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:munder><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>r</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:munderover><mml:mfrac><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>j</mml:mi><mml:mi>r</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mi>l</mml:mi><mml:mi>o</mml:mi><mml:msub><mml:mi>g</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mn>1</mml:mn><mml:mo>+</mml:mo><mml:mi>r</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mfrac></mml:mrow><mml:mrow><mml:mrow><mml:mi>n</mml:mi><mml:mi>u</mml:mi><mml:mi>m</mml:mi><mml:mi>b</mml:mi><mml:mi>e</mml:mi><mml:mi>r</mml:mi></mml:mrow><mml:mspace width="thinmathspace" /><mml:mi>o</mml:mi><mml:mi>f</mml:mi><mml:mspace width="thinmathspace" /><mml:mrow><mml:mi>c</mml:mi><mml:mi>o</mml:mi><mml:mi>m</mml:mi><mml:mi>p</mml:mi><mml:mi>l</mml:mi><mml:mi>e</mml:mi><mml:mi>t</mml:mi><mml:mi>e</mml:mi><mml:mi>d</mml:mi></mml:mrow><mml:mspace width="thinmathspace" /><mml:mrow><mml:mi>t</mml:mi><mml:mi>a</mml:mi><mml:mi>s</mml:mi><mml:mi>k</mml:mi><mml:mi>s</mml:mi></mml:mrow></mml:mrow></mml:mfrac></mml:math></disp-formula>
<list list-type="bullet">
<list-item>
<p>CR: Worker Completion Rate (CR). When the <inline-formula id="ieqn-121"><mml:math id="mml-ieqn-121"><mml:msub><mml:mrow><mml:mi>w</mml:mi><mml:mi>o</mml:mi><mml:mi>r</mml:mi><mml:mi>k</mml:mi><mml:mi>e</mml:mi><mml:mi>r</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> arrives at timestamp <inline-formula id="ieqn-122"><mml:math id="mml-ieqn-122"><mml:mi>i</mml:mi></mml:math></inline-formula>, the agent recommends a task. If the task is the same as the one actually selected by <inline-formula id="ieqn-123"><mml:math id="mml-ieqn-123"><mml:msub><mml:mrow><mml:mi>w</mml:mi><mml:mi>o</mml:mi><mml:mi>r</mml:mi><mml:mi>k</mml:mi><mml:mi>e</mml:mi><mml:mi>r</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo></mml:math></inline-formula> <inline-formula id="ieqn-124"><mml:math id="mml-ieqn-124"><mml:msub><mml:mi>R</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is 1, otherwise 0.</p></list-item>
<list-item>
<p>nDCG-CR: Normalized Discount Cumulative Gain. Instead of one task, the agent recommends a list of tasks. nDCGCR is more suitable for evaluating the recommended list <inline-formula id="ieqn-125"><mml:math id="mml-ieqn-125"><mml:msub><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mo fence="false" stretchy="false">{</mml:mo><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mi>j</mml:mi><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mi>j</mml:mi><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mi>j</mml:mi><mml:mn>3</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo fence="false" stretchy="false">}</mml:mo></mml:math></inline-formula>. <inline-formula id="ieqn-126"><mml:math id="mml-ieqn-126"><mml:mi>r</mml:mi></mml:math></inline-formula> is the rank position of tasks in the list, <inline-formula id="ieqn-127"><mml:math id="mml-ieqn-127"><mml:mi>N</mml:mi></mml:math></inline-formula> is the number of available tasks when the <inline-formula id="ieqn-128"><mml:math id="mml-ieqn-128"><mml:msub><mml:mrow><mml:mi>w</mml:mi><mml:mi>o</mml:mi><mml:mi>r</mml:mi><mml:mi>k</mml:mi><mml:mi>e</mml:mi><mml:mi>r</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> arrives. If the <inline-formula id="ieqn-129"><mml:math id="mml-ieqn-129"><mml:mi>r</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mi>j</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> is the same as the one actually selected by the <inline-formula id="ieqn-130"><mml:math id="mml-ieqn-130"><mml:msub><mml:mrow><mml:mi>w</mml:mi><mml:mi>o</mml:mi><mml:mi>r</mml:mi><mml:mi>k</mml:mi><mml:mi>e</mml:mi><mml:mi>r</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> , <inline-formula id="ieqn-131"><mml:math id="mml-ieqn-131"><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>j</mml:mi><mml:mi>r</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>&#x003D; 0 is 1, otherwise 0.</p></list-item>
</list></p>
<p><bold>Competitors:</bold> We compare our method with three alternatives to reinforcement learning, namely DDQN, D3QN, and EndToEnd. All these methods are trained on real datasets, and the characteristics of workers and tasks are updated in real time. By aggregating the Q-values of the predicted tasks relative to the worker and the request, an available task is selected or the available tasks are ranked according to the aggregated Q-values. The model updates the parameters in real time after each recommendation.
<list list-type="bullet">
<list-item>
<p>DDQN [<xref ref-type="bibr" rid="ref-6">6</xref>]: Double DQN is the deep Learning realization of Double Q-learning, and uses Target network to reduce the overestimation problem in Q-learning. The training is based on the parameters of the current Q network, rather than the parameters of target-Q as in DQN. In this way, overestimation is reduced to a certain extent, so that the Q-value is closer to the real value.</p></list-item>
<list-item>
<p>D3QN [<xref ref-type="bibr" rid="ref-7">7</xref>]: Dueling Double Network, using dominance function to further solve the overestimation problem. Dueling DQN can estimate Q-value more accurately and select more appropriate actions after collecting data of only one discrete action. Double DQN, selects the target Q-value by the action of the target Q-value selection, thereby eliminating the problem of overestimating the Q-value. D3QN combines the advantages of Dueling DQN and Double DQN.</p></list-item>
<list-item>
<p>EndToEnd [<xref ref-type="bibr" rid="ref-4">4</xref>] proposes a DQN network model with permutation invariant input property, which is suitable for a commercial crowdsourcing environment, and depends on the previous task allocation model in the end-to-end crowdsourcing reinforcement learning framework.</p></list-item>
</list></p>
</sec>
<sec id="s7_2">
<label>7.2</label>
<title>Experimental Results</title>
<p>In this section, we depict the experimental results of D3SQN in terms of both effectiveness and efficiency.</p>
<p><bold>Implementation Details:</bold> Our model consists of two D3SQNs, and after experimental evaluation, the number of neurons in each layer of D3SQN is set to 128. Moreover, for other hyper-parameters in D3SQN, we set target Q update frequency as 50, learning rate as 0.001, buffer size as 1200, discount factor Gamma as 0.35 and batch size as 128. We used PyTorch to implement the entire algorithm, and the code ran on a GeForce GTX 2080 TI GPU. We use the definition in <xref ref-type="sec" rid="s4">Sections 4</xref> and <xref ref-type="sec" rid="s7_1">7.1</xref> to construct the environment, action, state, and reward necessary for reinforcement learning, and our model learns spatio-temporal crowdsourced task allocation strategies within this framework. The number of layers of our encoder layer and decoder layer is the same as that of EndToEnd [<xref ref-type="bibr" rid="ref-4">4</xref>], and the complexity of our model is on the same order of magnitude as that of endtoend in terms of neuron parameters.</p>
<p><bold>Effectiveness:</bold> In <xref ref-type="sec" rid="s3">Section 3</xref>, we constructed two D3SQNs, namely D3SQN(R) and D3SQN(W), for requesters and workers, respectively, and evaluated their benefits along with the aggregated balance in our experiment. We presented the CR, CR(W), CR(R), nDCG, nDCG(W), and nDCG(R) measures for each method and dataset in <xref ref-type="fig" rid="fig-3">Fig. 3</xref>. CR(W) and CR(R) denote the proportion of maximum Q-values chosen by the two sub-networks of D3SQN that align with the final target. Similarly, nDCG(W) and nDCG(R) have the same interpretation. To account for both workers and requesters, we employed a ratio of 1:1 to aggregate the Q-values computed by the two sub-networks. This approach ensures that the workers and requesters contributions to the overall reward are equally weighted.</p>
<fig id="fig-3">
<label>Figure 3</label>
<caption>
<title>Benefits of workers &#x0026; requesters</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_31350-fig-3a.tif"/><graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_31350-fig-3b.tif"/>
</fig>
<p>The results, as depicted in <xref ref-type="fig" rid="fig-3">Fig. 3</xref>, demonstrate that D3SQN outperforms all competitors in terms of aggregated CR and aggregated nDCG-CR for both datasets. In contrast, DDQN performs poorly because it ignores the preferences of requesters and workers. Additionally, neither DDQN nor D3QN can handle permutation-invariant feature inputs, which leads to suboptimal performance.</p>

<p>The D3SQN algorithm demonstrates superior performance in CR(W) and nDCG-CR (W) compared to the End-to-End method due to its consideration of spatio-temporal constraints and features. Our model outperforms all other competitors in both datasets. However, in terms of CR(R) and nDCG-CR (R), our model surpasses all competitors in the ELM dataset, performs better than most competitors in the DiDi&#x0027;s dataset, and slightly lags behind the End-to-End method. This is attributed to our model&#x0027;s attention to spatio-temporal attributes, which gives more weight to workers in each task&#x2019;s spatio-temporal attributes. In datasets with relatively less user data, our model prioritizes the weight associated with learning workers, resulting in the best aggregation effect. Moreover, our model exhibits exceptional efficiency, achieving the fastest convergence rate and outperforming all metrics in both datasets.</p>
</sec>
<sec id="s7_3">
<label>7.3</label>
<title>Ablation Study</title>
<p>In this section, we will test the effect of each Module, we will test the SetTranformer architecture, GatingLayer &#x0026; Identity Map Reordering and Dueling Network architecture in D3SQN.</p>
<p>As shown in <xref ref-type="fig" rid="fig-4">Fig. 4</xref>, removing each improvement from the model reduces the prediction of the model. Especially when the SetTransformer structure is removed, the prediction effect of the model decreases most obviously. This is because the SetTransformer structure learns the worker preferences of all the training samples and is not affected by the variable task order.</p>
<fig id="fig-4">
<label>Figure 4</label>
<caption>
<title>Ablation experiments with different point</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_31350-fig-4.tif"/>
</fig>
<p>However, removing the GatingLayer &#x0026; Identity Map Reordering significantly decreases the performance of the model, as the absence of the identity mapping makes it harder for the transformer architecture to learn memory-based policies in complex feature representations within the RL environment. In addition, removing the Dueling structure slows down the training process since the Dueling network architecture allows for feature sharing between different actions, which accelerates the learning process.</p>
</sec>
<sec id="s7_4">
<label>7.4</label>
<title>Limitations</title>
<p>Our experiments still have some limitations. For instance, there are too few open source spatio-temporal crowdsourcing datasets for us to verify the generalization of our model on more datasets, which may limit the generalizability of our proposed model to other platforms. However, we believe that the data from these two platforms (DiDi and ELM) are representative of a significant portion of the spatial crowdsourcing platforms, as they are widely used in the transportation industry. Secondly, our data set comes from the desensitization data of some crowdsourcing platforms, which lack more information of users and order feedback. The influence of this part of real information on the verification of the model effect also increases the limitation of our experiment.</p>
</sec>
</sec>
<sec id="s8">
<label>8</label>
<title>Conclusion</title>
<p>This study proposed a deep reinforcement learning framework network D3SQN for spatial delivery task assignment in spatial edge intelligence platforms. It considered the critical role of spatio-temporal attributes in task assignment, and the spatio-temporal preferences of workers and requesters with respect to published tasks to facilitate better task assignment. The spatial temporal transformer proposed herein can effectively extract the interchangeable spatial-temporal state vector, and excavate the complex deep relationship between spatial-temporal attributes, workers and requesters in the task. This enables D3SQN to achieve optimal assignment in dynamic spatio-temporal crowdsourcing scenarios. The experimental results show that D3SQN provides effective and efficient assignment performance.</p>
</sec>
</body>
<back>
<ack>
<p>We would like to thank Dr. Caihua Shan for sharing her code and data with us.</p>
</ack>
<sec><title>Funding Statement</title>
<p>This work was supported in part by the Pioneer and Leading Goose R&#x0026;D Program of Zhejiang Province under Grant 2022C01083 (Dr. Yu Li, <ext-link ext-link-type="uri" xlink:href="https://zjnsf.kjt.zj.gov.cn/">https://zjnsf.kjt.zj.gov.cn/</ext-link>), Pioneer and Leading Goose R&#x0026;D Program of Zhejiang Province under Grant 2023C01217 (Dr. Yu Li, <ext-link ext-link-type="uri" xlink:href="https://zjnsf.kjt.zj.gov.cn/">https://zjnsf.kjt.zj.gov.cn/</ext-link>).</p>
</sec>
<sec><title>Author Contributions</title>
<p>Study conception and design: Yu Li, Mingxiao Li; data collection: Yu Li, Mingxiao Li; analysis and interpretation of results: Yu Li, Mingxiao Li, Dongyang Ou; draft manuscript preparation: Mingxiao Li, Junjie Guo, Pangyuan Fan. All authors reviewed the results and approved the final version of the manuscript.</p>
</sec>
<sec sec-type="data-availability"><title>Availability of Data and Materials</title>
<p>Our data access link is placed in the citation, and we access the data from the link in the citation; we are not the direct data publisher.</p>
</sec>
<sec sec-type="COI-statement"><title>Conflicts of Interest</title>
<p>The authors declare that they have no conflicts of interest to report regarding the present study.</p>
</sec>
<ref-list content-type="authoryear">
<title>References</title>
<ref id="ref-1"><label>1.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Rani</surname>, <given-names>S.</given-names></string-name>, <string-name><surname>Kumar</surname>, <given-names>M.</given-names></string-name></person-group> (<year>2022</year>). <article-title>Ranking community detection algorithms for complex social networks using multilayer network design approach</article-title>. <source>International Journal of Web Information Systems</source><italic>,</italic> <volume>18</volume><italic>(</italic><issue>5/6</issue><italic>),</italic> <fpage>310</fpage>&#x2013;<lpage>341</lpage>.</mixed-citation></ref>
<ref id="ref-2"><label>2.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Nafea</surname>, <given-names>I. T.</given-names></string-name></person-group> (<year>2021</year>). <article-title>Simulation of crowd management using deep learning algorithm</article-title>. <source>International Journal of Web Information Systems</source><italic>,</italic> <volume>17</volume><italic>(</italic><issue>4</issue><italic>),</italic> <fpage>321</fpage>&#x2013;<lpage>332</lpage>.</mixed-citation></ref>
<ref id="ref-3"><label>3.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Silva</surname>, <given-names>L.</given-names></string-name>, <string-name><surname>Silva</surname>, <given-names>N. F.</given-names></string-name>, <string-name><surname>Rosa</surname>, <given-names>T.</given-names></string-name></person-group> (<year>2020</year>). <article-title>Success prediction of crowdfunding campaigns: A two-phase modeling</article-title>. <source>International Journal of Web Information Systems</source><italic>,</italic> <volume>16</volume><italic>(</italic><issue>4</issue><italic>),</italic> <fpage>387</fpage>&#x2013;<lpage>412</lpage>.</mixed-citation></ref>
<ref id="ref-4"><label>4.</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Shan</surname>, <given-names>C.</given-names></string-name>, <string-name><surname>Mamoulis</surname>, <given-names>N.</given-names></string-name>, <string-name><surname>Cheng</surname>, <given-names>R.</given-names></string-name>, <string-name><surname>Li</surname>, <given-names>G.</given-names></string-name>, <string-name><surname>Li</surname>, <given-names>X.</given-names></string-name> <etal>et al.</etal></person-group> (<year>2020</year>). <article-title>An end-to-end deep RL framework for task arrangement in crowdsourcing platforms</article-title>. <conf-name>Proceedings of the IEEE 36th International Conference on Data Engineering (ICDE)</conf-name>, pp. <fpage>49</fpage>&#x2013;<lpage>60</lpage>. <publisher-loc>Dallas, TX, USA</publisher-loc>.</mixed-citation></ref>
<ref id="ref-5"><label>5.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Mnih</surname>, <given-names>V.</given-names></string-name>, <string-name><surname>Kavukcuoglu</surname>, <given-names>K.</given-names></string-name>, <string-name><surname>Silver</surname>, <given-names>D.</given-names></string-name>, <string-name><surname>Rusu</surname>, <given-names>A. A.</given-names></string-name>, <string-name><surname>Veness</surname>, <given-names>J.</given-names></string-name> <etal>et al.</etal></person-group> (<year>2015</year>). <article-title>Human-level control through deep reinforcement learning</article-title>. <source>Nature</source><italic>,</italic> <volume>518</volume><italic>(</italic><issue>7540</issue><italic>),</italic> <fpage>529</fpage>&#x2013;<lpage>533</lpage>; <pub-id pub-id-type="pmid">25719670</pub-id></mixed-citation></ref>
<ref id="ref-6"><label>6.</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Wang</surname>, <given-names>Z.</given-names></string-name>, <string-name><surname>Schaul</surname>, <given-names>T.</given-names></string-name>, <string-name><surname>Hessel</surname>, <given-names>M.</given-names></string-name>, <string-name><surname>van Hasselt</surname>, <given-names>H.</given-names>, </string-name>, <string-name><surname>Lanctot</surname>, <given-names>M.</given-names></string-name> <etal>et al.</etal></person-group> (<year>2016</year>). <article-title>Dueling network architectures for deep reinforcement learning</article-title>. <conf-name>Proceedings of the 33rd International Conference on Machine Learning</conf-name>, vol. <volume>48</volume><italic>,</italic> pp. <fpage>1995</fpage>&#x2013;<lpage>2003</lpage>. <publisher-loc>New York, USA</publisher-loc>.</mixed-citation></ref>
<ref id="ref-7"><label>7.</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>van Hasselt</surname>, <given-names>H.</given-names></string-name>, <string-name><surname>Guez</surname>, <given-names>A.</given-names></string-name>, <string-name><surname>Silver</surname>, <given-names>D.</given-names></string-name></person-group> (<year>2016</year>). <article-title>Deep reinforcement learning with double Q-learning</article-title>. <conf-name>Proceedings of the Thirtieth AAAI Conference on Artificial Intelligence</conf-name>, pp. <fpage>2094</fpage>&#x2013;<lpage>2100</lpage>. <publisher-loc>Phoenix, Arizona, USA</publisher-loc>.</mixed-citation></ref>
<ref id="ref-8"><label>8.</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Hessel</surname>, <given-names>M.</given-names></string-name>, <string-name><surname>Modayil</surname>, <given-names>J.</given-names></string-name>, <string-name><surname>van Hasselt</surname>, <given-names>H.</given-names></string-name>, <string-name><surname>Schaul</surname>, <given-names>T.</given-names></string-name>, <string-name><surname>Ostrovski</surname>, <given-names>G.</given-names></string-name> <etal>et al.</etal></person-group> (<year>2018</year>). <article-title>Rainbow: Combining improvements in deep reinforcement learning</article-title>. <conf-name>Proceedings of AAAI&#x0027;18/IAAI&#x0027;18/EAAI&#x0027;18</conf-name>, pp. <fpage>393</fpage>&#x2013;<lpage>400</lpage>. <publisher-loc>New Orleans, USA</publisher-loc>.</mixed-citation></ref>
<ref id="ref-9"><label>9.</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Schaul</surname>, <given-names>T.</given-names></string-name>, <string-name><surname>Quan</surname>, <given-names>J.</given-names></string-name>, <string-name><surname>Antonoglou</surname>, <given-names>I.</given-names></string-name>, <string-name><surname>Silver</surname>, <given-names>D.</given-names></string-name></person-group> (<year>2016</year>). <article-title>Prioritized experience replay</article-title>. <conf-name>Proceedings of the 4th International Conference on Learning Representations</conf-name>, pp. <fpage>322</fpage>&#x2013;<lpage>355</lpage>. <publisher-loc>San Juan, Puerto Rico</publisher-loc>.</mixed-citation></ref>
<ref id="ref-10"><label>10.</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>de Asis</surname>, <given-names>K.</given-names></string-name>, <string-name><surname>Hernandez-Garcia</surname>, <given-names>J.</given-names></string-name>, <string-name><surname>Holland</surname>, <given-names>G.</given-names></string-name>, <string-name><surname>Sutton</surname>, <given-names>R.</given-names></string-name></person-group> (<year>2018</year>). <article-title>Multi-step reinforcement learning: A unifying algorithm</article-title>. <conf-name>AAAI</conf-name>, pp. <fpage>1</fpage>&#x2013;<lpage>7</lpage>. <publisher-loc>New Orleans, Louisiana, USA</publisher-loc>.</mixed-citation></ref>
<ref id="ref-11"><label>11.</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Bellemare</surname>, <given-names>M. G.</given-names></string-name>, <string-name><surname>Dabney</surname>, <given-names>W.</given-names></string-name>, <string-name><surname>Munos</surname>, <given-names>R.</given-names></string-name></person-group> (<year>2017</year>). <article-title>A distributional perspective on reinforcement learning</article-title>. <conf-name>Proceedings of the 34th International Conference on Machine Learning</conf-name>, pp. <fpage>449</fpage>&#x2013;<lpage>458</lpage>. <publisher-loc>Sydney, Australia</publisher-loc>.</mixed-citation></ref>
<ref id="ref-12"><label>12.</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Fortunato</surname>, <given-names>M.</given-names></string-name>, <string-name><surname>Azar</surname>, <given-names>M. G.</given-names></string-name>, <string-name><surname>Piot</surname>, <given-names>B.</given-names></string-name>, <string-name><surname>Menick</surname>, <given-names>J.</given-names></string-name>, <string-name><surname>Hessel</surname>, <given-names>M.</given-names></string-name> <etal>et al.</etal></person-group> (<year>2018</year>). <article-title>Noisy networks for exploration</article-title>. <conf-name>Proceedings of the 6th International Conference on Learning Representations</conf-name>, pp. <fpage>1</fpage>&#x2013;<lpage>13</lpage>. <publisher-loc>Vancouver, British Columbia, Canada</publisher-loc>.</mixed-citation></ref>
<ref id="ref-13"><label>13.</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Vaswani</surname>, <given-names>A.</given-names></string-name>, <string-name><surname>Shazeer</surname>, <given-names>N.</given-names></string-name>, <string-name><surname>Parmar</surname>, <given-names>N.</given-names></string-name>, <string-name><surname>Uszkoreit</surname>, <given-names>J.</given-names></string-name>, <string-name><surname>Jones</surname>, <given-names>L.</given-names></string-name> <etal>et al.</etal></person-group> (<year>2017</year>). <article-title>Attention is all you need</article-title>. <conf-name>Proceedings of the 31st International Conference on Neural Information Processing Systems</conf-name>, pp. <fpage>6000</fpage>&#x2013;<lpage>6010</lpage>. <publisher-loc>Long Beach, California, USA</publisher-loc>.</mixed-citation></ref>
<ref id="ref-14"><label>14.</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Dai</surname>, <given-names>Z.</given-names></string-name>, <string-name><surname>Yang</surname>, <given-names>Z.</given-names></string-name>, <string-name><surname>Yang</surname>, <given-names>Y.</given-names></string-name>, <string-name><surname>Carbonell</surname>, <given-names>J. G.</given-names></string-name>, <string-name><surname>Le</surname>, <given-names>Q. V.</given-names></string-name> <etal>et al.</etal></person-group> (<year>2019</year>). <article-title>Transformer-XL: Attentive language models beyond a fixed-length context</article-title>. <conf-name>Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics</conf-name>, pp. <fpage>2978</fpage>&#x2013;<lpage>2988</lpage>. <publisher-loc>Florence, Italy</publisher-loc>.</mixed-citation></ref>
<ref id="ref-15"><label>15.</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Liu</surname>, <given-names>Y.</given-names></string-name>, <string-name><surname>Lapata</surname>, <given-names>M.</given-names></string-name></person-group> (<year>2019</year>). <article-title>Text summarization with pretrained encoders</article-title>. <conf-name>Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP)</conf-name>, pp. <fpage>3730</fpage>&#x2013;<lpage>3740</lpage>. <publisher-loc>Hong Kong, China</publisher-loc>.</mixed-citation></ref>
<ref id="ref-16"><label>16.</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Herrmann</surname>, <given-names>J.</given-names></string-name>, <string-name><surname>Bierbuesse</surname>, <given-names>D.</given-names></string-name>, <string-name><surname>Negra</surname>, <given-names>R.</given-names></string-name></person-group> (<year>2018</year>). <article-title>Universal model for millimeter-wave integrated transformers</article-title>. <conf-name>15th International Conference on Synthesis, Modeling, Analysis and Simulation Methods and Applications to Circuit Design (SMACD)</conf-name>, pp. <fpage>197</fpage>&#x2013;<lpage>200</lpage>. <publisher-loc>Prague, Czech Republic</publisher-loc>.</mixed-citation></ref>
<ref id="ref-17"><label>17.</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Yang</surname>, <given-names>Z.</given-names></string-name>, <string-name><surname>Dai</surname>, <given-names>Z.</given-names></string-name>, <string-name><surname>Yang</surname>, <given-names>Y.</given-names></string-name>, <string-name><surname>Carbonell</surname>, <given-names>J.</given-names></string-name>, <string-name><surname>Salakhutdinov</surname>, <given-names>R.</given-names></string-name> <etal>et al.</etal></person-group> (<year>2019</year>). <article-title>XLNet: Generalized autoregressive pretraining for language understanding</article-title>. <conf-name>Proceedings of the 33rd International Conference on Neural Information Processing Systems</conf-name>, pp. <fpage>5753</fpage>&#x2013;<lpage>5763</lpage>. <publisher-loc>New York, USA</publisher-loc>, <publisher-name>Curran Associates Inc</publisher-name>.</mixed-citation></ref>
<ref id="ref-18"><label>18.</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Parisotto</surname>, <given-names>E.</given-names></string-name>, <string-name><surname>Song</surname>, <given-names>F.</given-names></string-name>, <string-name><surname>Rae</surname>, <given-names>J.</given-names></string-name>, <string-name><surname>Pascanu</surname>, <given-names>R.</given-names></string-name>, <string-name><surname>Gulcehre</surname>, <given-names>C.</given-names></string-name> <etal>et al.</etal></person-group> (<year>2020</year>). <article-title>Stabilizing transformers for reinforcement learning</article-title>. <conf-name>Proceedings of the 37th International Conference on Machine Learning</conf-name>, pp. <fpage>7487</fpage>&#x2013;<lpage>7498</lpage>.</mixed-citation></ref>
<ref id="ref-19"><label>19.</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Hassan</surname>, <given-names>U. U.</given-names></string-name>, <string-name><surname>Curry</surname>, <given-names>E.</given-names></string-name></person-group> (<year>2014</year>). <article-title>A multi-armed bandit approach to online spatial task assignment</article-title>. <conf-name>Proceedings of the 2014 IEEE 11th International Conference on Ubiquitous Intelligence and Computing and 2014 IEEE 11th International Conference on Autonomic and Trusted Computing and 2014 IEEE 14th International Conference on Scalable Computing and Communications and its Associated Workshops</conf-name>, pp. <fpage>212</fpage>&#x2013;<lpage>219</lpage>. <publisher-loc>Ayodya Resort, Bali, Indonesia</publisher-loc>.</mixed-citation></ref>
<ref id="ref-20"><label>20.</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Wang</surname>, <given-names>Y.</given-names></string-name>, <string-name><surname>Tong</surname>, <given-names>Y.</given-names></string-name>, <string-name><surname>Long</surname>, <given-names>C.</given-names></string-name>, <string-name><surname>Xu</surname>, <given-names>P.</given-names></string-name>, <string-name><surname>Xu</surname>, <given-names>K.</given-names></string-name> <etal>et al.</etal></person-group> (<year>2019</year>). <article-title>Adaptive dynamic bipartite graph matching: A reinforcement learning approach</article-title>. <conf-name>2019 IEEE 35th International Conference on Data Engineering (ICDE)</conf-name>, pp. <fpage>1478</fpage>&#x2013;<lpage>1489</lpage>. <publisher-loc>Macao, China</publisher-loc>.</mixed-citation></ref>
<ref id="ref-21"><label>21.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Liu</surname>, <given-names>Y.</given-names></string-name>, <string-name><surname>Guo</surname>, <given-names>B.</given-names></string-name>, <string-name><surname>Chen</surname>, <given-names>C.</given-names></string-name>, <string-name><surname>Du</surname>, <given-names>H.</given-names></string-name>, <string-name><surname>Yu</surname>, <given-names>Z.</given-names></string-name> <etal>et al.</etal></person-group> (<year>2019</year>). <article-title>FooDNet: Toward an optimized food delivery network based on spatial crowdsourcing</article-title>. <source>IEEE Transactions on Mobile Computing</source><italic>,</italic> <volume>18</volume><italic>(</italic><issue>6</issue><italic>),</italic> <fpage>1288</fpage>&#x2013;<lpage>1301</lpage>.</mixed-citation></ref>
<ref id="ref-22"><label>22.</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Cheng</surname>, <given-names>P.</given-names></string-name>, <string-name><surname>Chen</surname>, <given-names>L.</given-names></string-name>, <string-name><surname>Ye</surname>, <given-names>J.</given-names></string-name></person-group> (<year>2019</year>). <article-title>Cooperation-aware task assignment in spatial crowdsourcing</article-title>. <conf-name>Proceedings of the 2019 IEEE 35th International Conference on Data Engineering (ICDE)</conf-name>, pp. <fpage>1442</fpage>&#x2013;<lpage>1453</lpage>. <publisher-loc>Macao, China</publisher-loc>.</mixed-citation></ref>
<ref id="ref-23"><label>23.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Co-Reyes</surname>, <given-names>J. D.</given-names></string-name>, <string-name><surname>Miao</surname>, <given-names>Y.</given-names></string-name>, <string-name><surname>Peng</surname>, <given-names>D.</given-names></string-name>, <string-name><surname>Real</surname>, <given-names>E.</given-names></string-name>, <string-name><surname>Levine</surname>, <given-names>S.</given-names></string-name> <etal>et al.</etal></person-group> (<year>2022</year>). <article-title>Evolving reinforcement learning algorithms</article-title>. arXiv:2101.03958.</mixed-citation></ref>
<ref id="ref-24"><label>24.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Gao</surname>, <given-names>H.</given-names></string-name>, <string-name><surname>Liu</surname>, <given-names>L.</given-names></string-name>, <string-name><surname>Liu</surname>, <given-names>K.</given-names></string-name>, <string-name><surname>Yang</surname>, <given-names>Y.</given-names></string-name></person-group> (<year>2022</year>). <article-title>The joint method of triple attention and novel loss function for entity relation extraction in small data-driven computational social systems</article-title>. <source>IEEE Transactions on Computational Social Systems</source><italic>,</italic> <volume>9</volume><italic>(</italic><issue>6</issue><italic>),</italic> <fpage>1725</fpage>&#x2013;<lpage>1735</lpage>.</mixed-citation></ref>
<ref id="ref-25"><label>25.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Gao</surname>, <given-names>H.</given-names></string-name>, <string-name><surname>Qiu</surname>, <given-names>B.</given-names></string-name>, <string-name><surname>Barroso</surname>, <given-names>R. J. D.</given-names></string-name>, <string-name><surname>Hussain</surname>, <given-names>W.</given-names></string-name>, <string-name><surname>Xu</surname>, <given-names>Y.</given-names></string-name> <etal>et al.</etal></person-group> (<year>2023</year>). <article-title>TSMAE: A novel anomaly detection approach for Internet of Things time series data using memory-augmented autoencoder</article-title>. <source>IEEE Transactions on Network Science and Engineering</source><italic>,</italic> <volume>10</volume><italic>(</italic><issue>5</issue><italic>),</italic> <fpage>2978</fpage>&#x2013;<lpage>2990</lpage>.</mixed-citation></ref>
<ref id="ref-26"><label>26.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Gao</surname>, <given-names>H.</given-names></string-name>, <string-name><surname>Dai</surname>, <given-names>B.</given-names></string-name>, <string-name><surname>Miao</surname>, <given-names>H.</given-names></string-name>, <string-name><surname>Yang</surname>, <given-names>X.</given-names></string-name>, <string-name><surname>Barroso</surname>, <given-names>R. J. D.</given-names></string-name> <etal>et al.</etal></person-group> (<year>2023</year>). <article-title>A novel GAPG approach to automatic property generation for formal verification: The GAN perspective</article-title>. <source>ACM Transactions on Multimedia Computing, Communications, and Applications</source><italic>,</italic> <volume>19</volume><italic>(</italic><issue>1</issue><italic>),</italic> <fpage>16</fpage>.</mixed-citation></ref>
<ref id="ref-27"><label>27.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Gao</surname>, <given-names>H.</given-names></string-name>, <string-name><surname>Wang</surname>, <given-names>X.</given-names></string-name>, <string-name><surname>Wei</surname>, <given-names>W.</given-names></string-name>, <string-name><surname>Al-Dulaimi</surname>, <given-names>A.</given-names></string-name>, <string-name><surname>Xu</surname>, <given-names>Y.</given-names></string-name></person-group> (<year>2023</year>). <article-title>Com-DDPG: Task offloading based on multiagent reinforcement learning for information-communication-enhanced mobile edge computing in the internet of vehicles</article-title>. <source>IEEE Transactions on Vehicular Technology</source><italic>,</italic> <fpage>1</fpage>&#x2013;<lpage>14</lpage>. <ext-link ext-link-type="uri" xlink:href="https://ieeexplore.ieee.org/document/10233027">https://ieeexplore.ieee.org/document/10233027</ext-link> <comment>(accessed on 20/10/2020)</comment></mixed-citation></ref>
<ref id="ref-28"><label>28.</label><mixed-citation publication-type="other"><person-group person-group-type="author"><collab>GAIA</collab></person-group> <year>(2020)</year>. <ext-link ext-link-type="uri" xlink:href="https://huggingface.co/datasets/seablue/DiDi_GAIA_dataset">https://huggingface.co/datasets/seablue/DiDi_GAIA_dataset</ext-link> <comment>(accessed on 20/10/2020)</comment></mixed-citation></ref>
<ref id="ref-29"><label>29.</label><mixed-citation publication-type="other"><person-group person-group-type="author"><collab>Elmme</collab></person-group> <year>(2022)</year>. <ext-link ext-link-type="uri" xlink:href="https://tianchi.aliyun.com/competition/entrance/231777/information">https://tianchi.aliyun.com/competition/entrance/231777/information</ext-link> <comment>(accessed on 20/10/2020)</comment></mixed-citation></ref>
</ref-list>
</back></article>