<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.1 20151215//EN" "http://jats.nlm.nih.gov/publishing/1.1/JATS-journalpublishing1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xml:lang="en" article-type="research-article" dtd-version="1.1">
<front>
<journal-meta>
<journal-id journal-id-type="pmc">CMC</journal-id>
<journal-id journal-id-type="nlm-ta">CMC</journal-id>
<journal-id journal-id-type="publisher-id">CMC</journal-id>
<journal-title-group>
<journal-title>Computers, Materials &#x0026; Continua</journal-title>
</journal-title-group>
<issn pub-type="epub">1546-2226</issn>
<issn pub-type="ppub">1546-2218</issn>
<publisher>
<publisher-name>Tech Science Press</publisher-name>
<publisher-loc>USA</publisher-loc>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">52851</article-id>
<article-id pub-id-type="doi">10.32604/cmc.2024.052851</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Article</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>FPGA Accelerators for Computing Interatomic Potential-Based Molecular Dynamics Simulation for Gold Nanoparticles: Exploring Different Communication Protocols</article-title>
<alt-title alt-title-type="left-running-head">FPGA Accelerators for Computing Interatomic Potential-Based Molecular Dynamics Simulation for Gold Nanoparticles: Exploring Different Communication Protocols</alt-title>
<alt-title alt-title-type="right-running-head">FPGA Accelerators for Computing Interatomic Potential-Based Molecular Dynamics Simulation for Gold Nanoparticles: Exploring Different Communication Protocols</alt-title>
</title-group>
<contrib-group>
<contrib id="author-1" contrib-type="author">
<name name-style="western"><surname>Patel</surname><given-names>Ankitkumar</given-names></name><xref ref-type="aff" rid="aff-1">1</xref></contrib>
<contrib id="author-2" contrib-type="author" corresp="yes">
<name name-style="western"><surname>Vasudevan</surname><given-names>Srivathsan</given-names></name><xref ref-type="aff" rid="aff-1">1</xref><email>svasudevan@iiti.ac.in</email></contrib>
<contrib id="author-3" contrib-type="author" corresp="yes">
<name name-style="western"><surname>Bulusu</surname><given-names>Satya</given-names></name><xref ref-type="aff" rid="aff-2">2</xref><email>sbulusu@iiti.ac.in</email></contrib>
<aff id="aff-1"><label>1</label><institution>Department of Electrical Engineering, Indian Institute of Technology Indore</institution>, <addr-line>Indore, 453552</addr-line>, <country>India</country></aff>
<aff id="aff-2"><label>2</label><institution>Department of Chemistry, Indian Institute of Technology Indore</institution>, <addr-line>Indore, 453552</addr-line>, <country>India</country></aff>
</contrib-group>
<author-notes>
<corresp id="cor1"><label>&#x002A;</label>Corresponding Authors: Srivathsan Vasudevan. Email: <email>svasudevan@iiti.ac.in</email>; Satya Bulusu. Email: <email>sbulusu@iiti.ac.in</email></corresp>
</author-notes>
<pub-date date-type="collection" publication-format="electronic">
<year>2024</year></pub-date>
<pub-date date-type="pub" publication-format="electronic"><day>12</day><month>9</month><year>2024</year></pub-date>
<volume>80</volume>
<issue>3</issue>
<fpage>3803</fpage>
<lpage>3818</lpage>
<history>
<date date-type="received">
<day>17</day>
<month>4</month>
<year>2024</year>
</date>
<date date-type="accepted">
<day>26</day>
<month>7</month>
<year>2024</year>
</date>
</history>
<permissions>
<copyright-statement>&#x00A9; 2024 The Authors.</copyright-statement>
<copyright-year>2024</copyright-year>
<copyright-holder>Published by Tech Science Press.</copyright-holder>
<license xlink:href="https://creativecommons.org/licenses/by/4.0/">
<license-p>This work is licensed under a <ext-link ext-link-type="uri" xlink:type="simple" xlink:href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</ext-link>, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.</license-p>
</license>
</permissions>
<self-uri content-type="pdf" xlink:href="TSP_CMC_52851.pdf"></self-uri>
<abstract>
<p>Molecular Dynamics (MD) simulation for computing Interatomic Potential (IAP) is a very important High-Performance Computing (HPC) application. MD simulation on particles of experimental relevance takes huge computation time, despite using an expensive high-end server. Heterogeneous computing, a combination of the Field Programmable Gate Array (FPGA) and a computer, is proposed as a solution to compute MD simulation efficiently. In such heterogeneous computation, communication between FPGA and Computer is necessary. One such MD simulation, explained in the paper, is the (Artificial Neural Network) ANN-based IAP computation of gold (Au<sub>147</sub> &#x0026; Au<sub>309</sub>) nanoparticles. MD simulation calculates the forces between atoms and the total energy of the chemical system. This work proposes the novel design and implementation of an ANN IAP-based MD simulation for Au<sub>147</sub> &#x0026; Au<sub>309</sub> using communication protocols, such as Universal Asynchronous Receiver-Transmitter (UART) and Ethernet, for communication between the FPGA and the host computer. To improve the latency of MD simulation through heterogeneous computing, Universal Asynchronous Receiver-Transmitter (UART) and Ethernet communication protocols were explored to conduct MD simulation of 50,000 cycles. In this study, computation times of 17.54 and 18.70 h were achieved with UART and Ethernet, respectively, compared to the conventional server time of 29 h for Au<sub>147</sub> nanoparticles. The results pave the way for the development of a Lab-on-a-chip application.</p>
</abstract>
<kwd-group kwd-group-type="author">
<kwd>Ethernet</kwd>
<kwd>hardware accelerator</kwd>
<kwd>heterogeneous computing</kwd>
<kwd>interatomic potential (IAP)</kwd>
<kwd>MD simulation</kwd>
<kwd>peripheral component interconnect express (PCIe)</kwd>
<kwd>UART</kwd>
</kwd-group>
</article-meta>
</front>
<body>
<sec id="s1">
<label>1</label>
<title>Introduction</title>
<p>High Performance Computing (HPC) applications, such as simulating Molecular Dynamics (MD), forecasting weather patterns, studying nuclear physics, Data Science and Engineering (DSE), and the Internet of Things (IoT), require diverse computing infrastructures ranging from individual desktop setups to extensive parallel processing environments [<xref ref-type="bibr" rid="ref-1">1</xref>,<xref ref-type="bibr" rid="ref-2">2</xref>]. With the tremendous increase in data volume, traditional Central Processing Units (CPUs) are finding it challenging to keep up with the demands of HPC [<xref ref-type="bibr" rid="ref-3">3</xref>]. Researchers are working on improving the speed of HPC applications. With advancements in the Very Large-Scale Integration (VLSI) hardware industry, they use Heterogeneous Computing platforms to improve the performance of these HPC applications. Heterogeneous computing is a unique kind of parallel computing where different tasks are allocated to different systems to achieve optimal performance and power efficiency [<xref ref-type="bibr" rid="ref-4">4</xref>,<xref ref-type="bibr" rid="ref-5">5</xref>]. Researchers are integrating various systems such as CPUs, Graphics Processing Units (GPUs), FPGAs, and Application Specific Integrated Circuits (ASICs) to accelerate the performance of heterogeneous computing systems [<xref ref-type="bibr" rid="ref-6">6</xref>&#x2013;<xref ref-type="bibr" rid="ref-8">8</xref>]. One such HPC application that our group is working on is the molecular dynamics simulation using ANN-based interatomic potential of gold (Au<sub>13</sub>, Au<sub>55</sub>, Au<sub>147</sub>, Au<sub>309</sub>, etc.) nanoparticles, Gold nanoparticles have always been a subject of interest in various applications such as biomedical, chemical, plasmonics, and non-linear optics, among others [<xref ref-type="bibr" rid="ref-9">9</xref>].</p>
<p>MD simulation, a vital application of HPC, is a technique for computer simulations of chemical systems for predicting their structural and thermodynamic properties. It involves 1. Time evolutions of atoms or molecules in the system, and 2. Interactions between the atoms and/or molecules in the system for a fixed period of time. Time evolution of the system is followed via integrating classical equations of motions of atoms and/or molecules in the system. The Verlet algorithm [<xref ref-type="bibr" rid="ref-10">10</xref>,<xref ref-type="bibr" rid="ref-11">11</xref>] is used to integrate the equation of motion. It is well known that the evaluation of interatomic interactions between atoms and/or molecules, in the studied material system, is the most computationally expensive task in MD simulations.</p>
<p>The modeling of inter-atomic (molecular) interactions requires calculations of force acting on each atom due to all other atoms in the systems, which is always a computational bottleneck in an MD simulation. To address this, researchers are exploring the concept of heterogeneous computing [<xref ref-type="bibr" rid="ref-12">12</xref>]. In MD simulation, various heterogeneous computing models, such as GPUs, FPGAs, and ASICs, are considered alongside CPUs [<xref ref-type="bibr" rid="ref-13">13</xref>&#x2013;<xref ref-type="bibr" rid="ref-15">15</xref>]. FPGAs stand out due to their ability to handle parallelized computations and their potential to overcome computation time challenges [<xref ref-type="bibr" rid="ref-16">16</xref>&#x2013;<xref ref-type="bibr" rid="ref-18">18</xref>].</p>
<p>Our group&#x2019;s earlier work Bulusu et al. [<xref ref-type="bibr" rid="ref-19">19</xref>] presents an innovative approach for implementing the use of MD simulations using FPGAs. Using an MD simulation for ANN-based Interatomic Potential (IAP), the study focuses on gold nanoparticles, specifically the Au<sub>147</sub> &#x0026; Au<sub>309</sub> nanoparticles. The complete system was implemented on the Xilinx Kintex-7 KC705 evaluation board [<xref ref-type="bibr" rid="ref-20">20</xref>]. The implemented design is shown in <xref ref-type="fig" rid="fig-1">Fig. 1</xref>, illustrating a hardware-software co-design model.</p>
<fig id="fig-1">
<label>Figure 1</label>
<caption>
<title>The hardware-software Co-design prototype for MD simulation</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_52851-fig-1.tif"/>
</fig>
<p>The prototype design shown in <xref ref-type="fig" rid="fig-1">Fig. 1</xref> divides the system into two parts. The FPGA hardware handles complex, time-intensive computations while the host computer manages controlling functions and sequential computations. The heterogeneous computing-based MD simulation shown in the model requires constant communication between the host computer and the FPGA board to function correctly. Peripheral Component Interconnect Express (PCIe) communication protocol is indeed used for communication between the FPGA and the host computer. In the hardware-software co-design shown in <xref ref-type="fig" rid="fig-1">Fig. 1</xref>, the Cartesian coordinates (X, Y, and Z) are sent from the host computer to the FPGA, and the outputs (forces and energy values) are sent from the FPGA to the host computer using PCIe communication. Thus, in every MD cycle, around 3.5 KB (Kilobytes) of data transfer occurs between the FPGA and host computer using PCIe for Au<sub>147</sub>. With PCIe communication 50,000 MD cycles, PCIe took 19.84 h, whereas HPC Server took 29 h for Au<sub>147</sub> nanoparticles. With PCIe, the computation time was reduced by 1.5 times as reported by Bulusu et al. [<xref ref-type="bibr" rid="ref-19">19</xref>].</p>
<p>Due to its high-speed serial point-to-point capabilities, PCIe is widely used in heterogeneous computing, especially in applications requiring large amounts of data transfer. However, PCIe comes with limitations. A thorough study by Marcin et al. [<xref ref-type="bibr" rid="ref-4">4</xref>] on FPGA applications highlights the first drawback, emphasizing the necessity for dedicated hardware and drivers, specifically a PCIe slot and associated drivers, leading to increased programming complexities. The second limitation involves the absence of support for hot plug operations, thus requiring a system restart for configuration. Also, PCIe&#x2019;s complexity overhead is unnecessary for small data transfers like a few KBs (here, 3.5 KB for Au<sub>147</sub> and 7.2 KB for Au<sub>309</sub>). Hence, the advantage of high-speed and low-latency data transfer is reduced due to the complexity overhead in a few KB data transfers.</p>
<p>This paper presents a novel approach to exploring UART (Universal Asynchronous Receiver/Transmitter) and Ethernet communication protocols for the same MD computation. We aim to reduce computation time further. Communication protocols like UART and Ethernet present versatile and user-friendly alternatives. Utilizing simplified, universally applicable code across operating systems helps mitigate the complexities associated with PCIe implementations. Both UART and Ethernet protocols support hot plug operations, eliminating the need for system restarts. These protocols are also easier to use than PCIe, and using them in embedded heterogeneous computing systems makes it easier to control the complete system by letting embedded processors step in. This attempt is for two atomic gold nanoparticles Au<sub>147</sub> and Au<sub>309</sub>.</p>
<p>To reduce computation time, UARTLite and EthernetLite IP (Intellectual Property) are explored in the design with the following key setup:
<list list-type="order">
<list-item>
<p>Explore UART and Ethernet communication protocols as alternatives to PCIe communication for ANN-based MD simulation for gold nanoparticles.</p></list-item>
<list-item>
<p>Integrate the Microblaze soft-core processor to enhance hardware control ensuring an easily debuggable system.</p></list-item>
<list-item>
<p>Implement the MD calculation in the FPGA-Computer (CPU) based system using both communication protocols.</p></list-item>
</list></p>
<p>The paper is structured as follows: <xref ref-type="sec" rid="s2">Section 2</xref> outlines the system architecture overview. <xref ref-type="sec" rid="s3">Section 3</xref> discusses the hardware-software co-design for UART and Ethernet-based design implementation. The obtained results are presented in <xref ref-type="sec" rid="s4">Section 4</xref>. <xref ref-type="sec" rid="s5">Section 5</xref> provides a detailed discussion, and <xref ref-type="sec" rid="s6">Section 6</xref> concludes the paper.</p>
</sec>
<sec id="s2">
<label>2</label>
<title>System Architecture Overview</title>
<p>The block diagram of hardware-software co-design for ANN-based MD simulation of gold nanoparticles is shown in <xref ref-type="fig" rid="fig-2">Fig. 2</xref>. Several FPGA IPs have been used to design hardware in Xilinx Vivado. A brief overview of these IPs is given below.</p>
<fig id="fig-2">
<label>Figure 2</label>
<caption>
<title>Block diagram of hardware-software Co-design for MD simulation</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_52851-fig-2.tif"/>
</fig>
<sec id="s2_1">
<label>2.1</label>
<title>FPGA IPs</title>
<sec id="s2_1_1">
<label>2.1.1</label>
<title>Microblaze: Soft-Core Processor</title>
<p>It is a 32-bit general-purpose Reduced Instruction Set Computer (RISC) soft-core processor. This processor features a 32-bit general-purpose register, RISC Harward Architecture, a 3-stage pipeline, and an interrupt module [<xref ref-type="bibr" rid="ref-21">21</xref>]. Utilizing the local memory bus (LMB), Microblaze accesses on-chip memory and is compatible with the IEEE 754 single-precision floating-point format. It also includes an Instruction Cache (IC) and Data Cache (DC), exception handling, a debug module, and a barrel shifter. Except for the Zynq family, Microblaze is supported in most Xilinx FPGA families (Artix-7, Kintex-7, Spartan, etc.). A comprehensive explanation of Microblaze is provided in XilinxMicroblaze [<xref ref-type="bibr" rid="ref-22">22</xref>].</p>
<p>Xilinx provides a software environment called Software Development Kit (SDK) with the Embedded processor (Microblaze) [<xref ref-type="bibr" rid="ref-22">22</xref>]. The SDK supports C/C&#x002B;&#x002B; languages for writing software code and is responsible for controlling the operation of the Microblaze soft-core processor. It extends support to all peripherals IPs used with the processor.</p>
</sec>
<sec id="s2_1_2">
<label>2.1.2</label>
<title>Communication IP (AXI UARTLite/EthernetLite)</title>
<p>The Advanced eXtensible Interface (AXI) UARTLite serves as a control interface for asynchronous serial data transfer. It supports full-duplex communication, providing AXI4-Lite interface register access and data transfer. The receiver and transmitter (First in First out) FIFO size are limited to 16 Bytes. This module incorporates configurable baud rates (e.g., 9600, 115200, 421800, 921600, etc.). A comprehensive explanation of AXI UARTLite is available in XilinxUart [<xref ref-type="bibr" rid="ref-23">23</xref>]. UARTLite IP follows the standard UART frame format as shown in <xref ref-type="fig" rid="fig-3">Fig. 3</xref>. In this implementation, the parity bit is not used, and the stop bit size is one bit. Thus, each UART packet (frame) is 10 bits in size.</p>
<fig id="fig-3">
<label>Figure 3</label>
<caption>
<title>UART frame format</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_52851-fig-3.tif"/>
</fig>
<p>The AXI EthernetLite is designed to incorporate the features of the IEEE 802.3 Ethernet standard. It facilitates connection to external 10/100 Mb/s physical (PHY) transceivers through the Media Independent Interface (MII). It utilizes the AXI4/AXI4-Lite on-chip communication protocol to enable communication with the Microblaze soft-core Processor. A comprehensive explanation of the AXI EthernetLite is provided in XilinxEthernet [<xref ref-type="bibr" rid="ref-24">24</xref>]. EthernetLite IP follows the IEEE 802.3 standard Ethernet frame format for communication as shown in <xref ref-type="fig" rid="fig-4">Fig. 4</xref>. This frame format allows a maximum of 1500 bytes of actual data to be sent, along with an additional 14 bytes for the header and 4 bytes for the Cyclic Redundancy Check (CRC). Therefore, for larger data sizes, multiple Ethernet frames have been used.</p>
<fig id="fig-4">
<label>Figure 4</label>
<caption>
<title>IEEE 802.3 ethernet frame format</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_52851-fig-4.tif"/>
</fig>
</sec>
</sec>
<sec id="s2_2">
<label>2.2</label>
<title>Hardware Accelerator HLS IP for ANN-Based IAP Calculation</title>
<p>In this subsection, FPGA hardware blocks are introduced for reconfigurable high-performance computation. Generally, high-performance computation contains multiple loops. So, in conventional processors, all loops run sequentially. However, FPGA has the ability to take up several independent serial loops and can parallelize them [<xref ref-type="bibr" rid="ref-25">25</xref>]. So, when implemented on FPGA, similar tasks are performed concurrently, and the algorithm will be converted into multipliers and adders (MAC blocks) that go along multiple loops.</p>
<p>The Algorithm 1 [<xref ref-type="bibr" rid="ref-19">19</xref>] for constructing the ANN-based IAP for gold nanoparticles is described below. Firstly, obtaining cartesian coordinates from the host computer via UART/Ethernet communication. The cartesian coordinates are used to calculate radial and angular descriptors and their derivatives. The descriptors serve as input to the NN (Neural Network) to evaluate the energy for each atom. Force components are calculated using the derivatives of the descriptors. Finally, the force components are sent back to the host computer.</p>
<fig id="fig-9">
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_52851-fig-9.tif"/>
</fig>
<p>Vivado-HLS (High-Level Synthesis) software has been chosen to accelerate this high-performance computation algorithm. Utilizing C-based code in Vivado HLS, it was converted into Hardware Description Language (HDL). As shown in <xref ref-type="fig" rid="fig-5">Fig. 5</xref>, after composing the C language code, the AXI_return directive was used to make it AXI-compatible. After that, a pipeline was applied to enhance the latency and throughput of the system. Also, all operations are floating-point operations and contain multidimensional arrays. So, it requires significant space in memory. Array partitioning directive is used to optimize multidimensional arrays. Upon achieving the optimal design, the algorithm was exported to the RTL (Register Transfer Level) design and then converted into IP. The physics and mathematical calculations of this module are completely discussed in Bulusu et al. [<xref ref-type="bibr" rid="ref-19">19</xref>].</p>
<fig id="fig-5">
<label>Figure 5</label>
<caption>
<title>Design flow for HLS IP creation</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_52851-fig-5.tif"/>
</fig>
</sec>
</sec>
<sec id="s3">
<label>3</label>
<title>Hardware-Software Co-Design</title>
<p>The inputs of the proposed system are (X, Y, Z) cartesian coordinates of atoms and the outputs are forces and total energy of the nanoparticles. Thus, data being large, are stored in the DDR (Double Data Rate) memory present in the FPGA evaluation board. Microblaze, a microcontroller IP was used to control data from the host computer to memory and vice-versa. <xref ref-type="fig" rid="fig-6">Fig. 6</xref> shows the sequence of data transfer for one single operation.</p>
<fig id="fig-6">
<label>Figure 6</label>
<caption>
<title>Data flow diagram of system implementation for MD simulation</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_52851-fig-6.tif"/>
</fig>
<p>The complete FPGA implementations of UART and Ethernet communication-based heterogeneous computing systems are discussed in detail in the following subsection.</p>
<sec id="s3_1">
<label>3.1</label>
<title>IAP Implementation Using UART/Ethernet</title>
<p>To design a communication interface, a schematic-based design was developed in Xilinx Vivado as shown in <xref ref-type="fig" rid="fig-4">Fig. 4</xref> hardware module [<xref ref-type="bibr" rid="ref-26">26</xref>]. This design included DDR3 (Double Data Rate3) SDRAM (Synchronous Dynamic Random-Access Memory) on the Kintex KC-705 board for data storage during processing. The Microblaze soft-core processor acted as the master controller, operating at 300 MHz, while the hardware accelerator MD simulation IP operated at 100 MHz. Communication between the FPGA and the host computer was handled by the AXI UARTLite IP / EthernetLite IP. Other Peripheral IPs communicated with the Microblaze processor through AXI Interconnect or SmartConnect IP.</p>
<p>An Embedded-C language code was written in Xilinx SDK to control the hardware [<xref ref-type="bibr" rid="ref-27">27</xref>]. This code handled tasks such as data reception and transmission, data storage in DDR memory, and initiation and termination of the hardware accelerator module. The flowchart of the SDK code is shown in <xref ref-type="fig" rid="fig-7">Fig. 7</xref>. The structure of the SDK code is nearly identical for UART and Ethernet-based designs. Changes related to communication interfaces were highlighted with different colored rectangular dotted boxes, with an orange box denoting UART communication and a green box for Ethernet communication.</p>
<fig id="fig-7">
<label>Figure 7</label>
<caption>
<title>SDK flow diagram for UART and ethernet communication implementation</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_52851-fig-7.tif"/>
</fig>
<p>The entire IAP computation is implemented as a hardware-software co-design. The Microblaze acts as a master and the hardware accelerator IP serves as a slave. The Microblaze starts with a polling to expect &#x2018;X, Y, Z&#x2019; coordinates values from the computer. Upon receiving the same, the accelerator IP is activated and the Microblaze goes into SLEEP mode. Parallelly, it will initiate another polling mode to wait for the forces &#x0026; energy deposition with the memory.</p>
<p>The accelerator IP uses the parallel algorithm on the FPGA to calculate the forces &#x0026; energy and deport the same in the DDR memory. Further, it ends the wait for the Microblaze. The Microblaze gets activated and transfers the data through UART/Ethernet IP to the host computer. Depending on the number of cycles, the entire process is repeated.</p>
<p>The Verlet algorithm was used to execute the remaining sequential operations and compile all necessary MD files in the host computer. The C-program in the host computer manages the UART/Ethernet communication between the host computer and FPGA.</p>
</sec>
</sec>
<sec id="s4">
<label>4</label>
<title>Results</title>
<p>For performing the MD simulation, two different examples are explored. They are Au<sub>147</sub> and Au<sub>309</sub> nanoparticles. The FPGA used was Xilinx Kintex-7 KC-705 evaluation FPGA board. In addition to the accelerator IP developed, a Microblaze-based microcontroller system was also implemented through the block design of the Xilinx Vivado software. The accelerator outputs energy and force values which need to be transferred to the computer. A separate Verlet algorithm will take the force and energy values and calculate the next set of coordinates on the computer side. The accelerator IP on the FPGA works at 100 MHz, while the Microblaze processor and its peripherals operate at 300 MHz. For communicating between the FPGA and the computer, AXI UARTLite and AXI EthernetLite IPs were implemented. Microblaze processor controls the data transfer between the FPGA Random Access Memory (RAM) and the different communication IPs, and this program was written in embedded C.</p>
<p>The first step in this exploration is to identify the accuracy of the results obtained through the proposed communication strategies. The calculated energy for every MD cycle from the FPGA-based heterogeneous computing system should match the conventional approach (using a server). To check the accuracy, both UART and Ethernet-based communication were implemented, and the MD simulations were run on the board for 500 cycles. The results of energy and forces obtained at every cycle were compared with a conventional approach, and the results are plotted in <xref ref-type="fig" rid="fig-8">Fig. 8</xref>. There is no difference between the results obtained from all the approaches, and this proves that the proposed approach yields accurate results for several cycles as well. It validates that the total potential energy obtained using UART and Ethernet designs is correct and identical to the PCIe-based design and HPC server, as explained in Bulusu et al. [<xref ref-type="bibr" rid="ref-19">19</xref>].</p>
<fig id="fig-8">
<label>Figure 8</label>
<caption>
<title>Potential energy of Au<sub>147</sub> &#x0026; Au<sub>309</sub> <italic>vs</italic>. number of MD cycles</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_52851-fig-8.tif"/>
</fig>
<p>Now that the accuracy of the system is verified, the system was explored to determine the computation time for the two different communication protocols implemented onto the FPGA board. A comparison of the computation time of the proposed two communication approaches with the conventional PCIe and HPC server is shown in <xref ref-type="table" rid="table-1">Table 1</xref>. UART operates at a baud rate of 921,600 bits/sec, while Ethernet works at a standard speed of 100 Mb/sec. PCIe speed is noted as 5 GT/sec, and in the HPC server, seven CPUs run in parallel at a frequency of 2600 MHz (Megahertz). Interestingly, UART and Ethernet-based communication protocols can withstand 50,000 iterations of MD computations, and the time it takes for 50,000 iterations comes to 17.54 &#x0026; 18.7 h for UART &#x0026; Ethernet communication, respectively for Au<sub>147</sub> and 64.36 &#x0026; 65.56 h respectively for Au<sub>309</sub>. This shows the robustness of the system and the potential to provide a lab-on-a-chip solution for such IAP computations.</p>
<table-wrap id="table-1">
<label>Table 1</label>
<caption>
<title>Computation time comparison of UART, ethernet, and PCIe communication-based design for Au<sub>147</sub> (Au<sub>309</sub>)</title>
</caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead valign="top">
<tr>
<th rowspan="3">No. of MD cycles</th>
<th align="center" colspan="4">Computation time</th>
</tr>
<tr>
<th>UART</th>
<th>Ethernet</th>
<th>PCIe</th>
<th>HPC server</th>
</tr>
<tr>
<th>921,600 bits/sec</th>
<th>100 MB/sec</th>
<th>5 GT/sec</th>
<th>&#x2014;</th>
</tr>
</thead>
<tbody>
<tr>
<td>1</td>
<td>1.26 sec (4.74 sec)</td>
<td>1.35 sec (4.76 sec)</td>
<td>2.77 sec (9.65 sec)</td>
<td>4.18 sec (16.75 sec)</td>
</tr>
<tr>
<td>10</td>
<td>13.89 sec (52.19 sec)</td>
<td>14.71 sec (52.21 sec)</td>
<td>15.18 sec (53.57 sec)</td>
<td>22.94 sec (92.00 sec)</td>
</tr>
<tr>
<td>100</td>
<td>2.12 min (7.99 min)</td>
<td>2.26 min (8.01 min)</td>
<td>2.34 min (8.12 min)</td>
<td>3.51 min (14.01 min)</td>
</tr>
<tr>
<td>500</td>
<td>10.54 min (39.62 min)</td>
<td>11.20 min (39.72 min)</td>
<td>11.58 min (40.25 min)</td>
<td>17.43 min (69.94 min)</td>
</tr>
<tr>
<td>5000</td>
<td>1.75 h (6.55 h)</td>
<td>1.86 h (6.56 h)</td>
<td>1.98 h (6.70 h)</td>
<td>2.90 h (11.64 h)</td>
</tr>
<tr>
<td>50,000</td>
<td>17.54 h (64.36 h)</td>
<td>18.70 h (65.56 h)</td>
<td>19.84 h (66.97 h)</td>
<td>29.00 h (116.34 h)</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Initially, X, Y, and Z values were sent to the FPGA and the returning data were the computed force and energy values. For the Au<sub>147</sub> atoms, 442 32-bit numbers (147 force values in each X, Y, and Z direction &#x002B; 1 total energy value) were transmitted from the memory to the computer. The return communication was 441 32-bit numbers back to the FPGA. In such cases, the total data transfer was 3.5 KB per cycle. Even though PCIe communication is faster, the overheads (e.g., frame construction) and loading of the drivers take more time. It is interesting to note that UART and Ethernet communication can achieve better computational efficiency compared to PCIe communication, despite PCIe being faster.</p>
<p>The next step is to determine the resources used by the FPGA board in both communication approaches. This has a direct correlation to the power of the system. The results, as tabulated in <xref ref-type="table" rid="table-2">Table 2</xref>, reveal that the UART and Ethernet designs, leveraging a Microblaze soft-core processor, demonstrate lower overall resource utilization than the PCIe-based design.</p>
<table-wrap id="table-2">
<label>Table 2</label>
<caption>
<title>Au<sub>147</sub> FPGA implementation: resource utilization insights for UART, ethernet, and PCIe communication</title>
</caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead valign="top">
<tr>
<th rowspan="2">FPGA resources</th>
<th align="center" colspan="3">Communication protocol</th>
</tr>
<tr>
<th>UART (%)</th>
<th>Ethernet (%)</th>
<th>PCIe (%)</th>
</tr>
</thead>
<tbody>
<tr>
<td>LUT (Look-Up Table)</td>
<td>55.50</td>
<td>55.80</td>
<td>65.85</td>
</tr>
<tr>
<td>LUT RAM</td>
<td>19.34</td>
<td>19.36</td>
<td>23.58</td>
</tr>
<tr>
<td>Flip-Flop</td>
<td>26.34</td>
<td>26.87</td>
<td>32.44</td>
</tr>
<tr>
<td>BRAM (Block RAM)</td>
<td>50.79</td>
<td>51.69</td>
<td>42.02</td>
</tr>
<tr>
<td>DSP (Digital Signal Processing) Slices</td>
<td>49.40</td>
<td>49.40</td>
<td>49.52</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>For a Lab-on-a-chip application, power consumption plays a very important role. <xref ref-type="table" rid="table-3">Table 3</xref> shows the power consumed by the different communication protocols. It is clear that UART consumes the least power, followed by the other protocols. This again can be attributed to the simplicity of the communication protocols.</p>
<table-wrap id="table-3">
<label>Table 3</label>
<caption>
<title>On-chip power consumption of UART, ethernet, &#x0026; PCIe design for Au<sub>147</sub> (Au<sub>309</sub>)</title>
</caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th>Communication protocols</th>
<th>Power (Watts)</th>
</tr>
</thead>
<tbody>
<tr>
<td>UART</td>
<td>5.9 (6.1)</td>
</tr>
<tr>
<td>Ethernet</td>
<td>6.0 (6.3)</td>
</tr>
<tr>
<td>PCIe</td>
<td>8.9 (9.2)</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>To check the efficiency of the proposed setup to experiment on new systems, an exploration Ag<sub>42</sub>Pt<sub>13</sub> system was explored. This system utilized the same feed-forward ANN model (59-30-30-1). However, since this alloy consists of two types of nanoparticles in the atomic cluster, the HLS code was modified with several additional conditional statements and loops. The algorithm for this system was written in an HLS-based FPGA program and the latency and the estimated computation time were calculated from the HLS synthesis. <xref ref-type="table" rid="table-4">Table 4</xref> shows the details of the estimated time before and after parallel operations. It is clear from this table that the proposed system can provide a significant reduction in computation time for up to two orders.</p>
<table-wrap id="table-4">
<label>Table 4</label>
<caption>
<title>Estimated computation time for Ag<sub>42</sub>Pt<sub>13</sub> HLS implementation with and without optimization</title>
</caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th>Technique</th>
<th>Computation time (in seconds) <break/>(Estimate from HLS)</th>
</tr>
</thead>
<tbody>
<tr>
<td>Without optimization</td>
<td>18.6</td>
</tr>
<tr>
<td>With optimization</td>
<td>0.197</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s5">
<label>5</label>
<title>Discussion</title>
<p>The communication speeds of UART and Ethernet (100 &#x00D7; <inline-formula id="ieqn-1"><mml:math id="mml-ieqn-1"><mml:msup><mml:mn>10</mml:mn><mml:mrow><mml:mn>6</mml:mn></mml:mrow></mml:msup><mml:mrow><mml:mtext>&#xA0;bps</mml:mtext></mml:mrow></mml:math></inline-formula> (bits per second)) are three orders respectively and the same three orders of difference is observed between ethernet and PCIe (31 &#x00D7; <inline-formula id="ieqn-2"><mml:math id="mml-ieqn-2"><mml:msup><mml:mn>10</mml:mn><mml:mrow><mml:mn>9</mml:mn></mml:mrow></mml:msup><mml:mrow><mml:mtext>&#xA0;bps</mml:mtext></mml:mrow></mml:math></inline-formula>). Despite such high communication speeds, the computational time obtained from our experimental systems is quite contradictory and worth discussing and reasoning them.</p>
<p>While the packet size of UART is 10 bits of communication (8 bits of actual data), ethernet and PCIe communication communicates with several initial frames and library initialization. The reason behind such a contradictory result could be ascertained by estimating the overheads of these communication protocols. Needless to say, that UART has minimal overhead and can be considered negligible compared to its own communication speed.</p>
<p>Ideally, one complete cycle of the MD calculation with any communication protocol is a combination of computation and communication time. This is shown in <xref ref-type="disp-formula" rid="eqn-1">Eq. (1)</xref> below:
<disp-formula id="eqn-1"><label>(1)</label><mml:math id="mml-eqn-1" display="block"><mml:msub><mml:mrow><mml:mtext>T</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mn>1</mml:mn><mml:mrow><mml:mtext>&#xA0;cycle</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mtext>T</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext>comm</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mrow><mml:mtext>T</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext>comp</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mrow><mml:mtext>T</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext>overhead</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msub></mml:math></disp-formula>where <inline-formula id="ieqn-3"><mml:math id="mml-ieqn-3"><mml:msub><mml:mrow><mml:mtext>T</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mn>1</mml:mn><mml:mrow><mml:mtext>&#xA0;cycle</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> is the total time taken for one complete cycle, <inline-formula id="ieqn-4"><mml:math id="mml-ieqn-4"><mml:msub><mml:mrow><mml:mtext>T</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext>comm</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> is the communication time and <inline-formula id="ieqn-5"><mml:math id="mml-ieqn-5"><mml:msub><mml:mrow><mml:mtext>T</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext>comp</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> is the time taken for computations in both FPGA and computer. <inline-formula id="ieqn-6"><mml:math id="mml-ieqn-6"><mml:msub><mml:mrow><mml:mtext>T</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext>overhead</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> is the overhead time required for initialization of several libraries and frames especially in case of Ethernet and PCIe. <inline-formula id="ieqn-7"><mml:math id="mml-ieqn-7"><mml:msub><mml:mrow><mml:mtext>T</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext>overhead</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> for UART is negligible and therefore we will ignore its contribution in calculating <inline-formula id="ieqn-8"><mml:math id="mml-ieqn-8"><mml:msub><mml:mrow><mml:mtext>T</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mn>1</mml:mn><mml:mrow><mml:mtext>&#xA0;cycle</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula>.
<list list-type="simple">
<list-item>
<p>&#x2756; Au<sub>147</sub> Time Calculations:</p></list-item>
</list></p>
<p>In the case of UART Communication for 1 cycle,
<disp-formula id="eqn-2"><label>(2)</label><mml:math id="mml-eqn-2" display="block"><mml:msub><mml:mrow><mml:msup><mml:mrow><mml:mtext>T</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>U</mml:mtext></mml:mrow></mml:mrow></mml:msup></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mn>1</mml:mn><mml:mspace width="thinmathspace" /><mml:mrow><mml:mtext>cycle</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mn>1.26</mml:mn><mml:mrow><mml:mtext>&#xA0;s&#xA0;</mml:mtext></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext>From Table 1</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula>
<disp-formula id="eqn-3"><label>(3)</label><mml:math id="mml-eqn-3" display="block"><mml:msub><mml:mrow><mml:msup><mml:mrow><mml:mtext>T</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>U</mml:mtext></mml:mrow></mml:mrow></mml:msup></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext>comm</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mrow><mml:mtext>No</mml:mtext></mml:mrow><mml:mo>.</mml:mo><mml:mrow><mml:mtext>of bits transferred&#xA0;</mml:mtext></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:mtext>Speed</mml:mtext></mml:mrow><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:mrow><mml:mtext>baudrate</mml:mtext></mml:mrow></mml:mrow></mml:mfrac><mml:mo>=</mml:mo><mml:mfrac><mml:mn>35360</mml:mn><mml:mn>921600</mml:mn></mml:mfrac><mml:mo>=</mml:mo><mml:mn>0.038</mml:mn><mml:mrow><mml:mtext>&#xA0;s&#xA0;</mml:mtext></mml:mrow></mml:math></disp-formula></p>
<p>From <xref ref-type="disp-formula" rid="eqn-1">Eq. (1)</xref>,
<disp-formula id="eqn-4"><label>(4)</label><mml:math id="mml-eqn-4" display="block"><mml:msub><mml:mrow><mml:mtext>T</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext>comp</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mn>1.22</mml:mn><mml:mrow><mml:mtext>&#xA0;s</mml:mtext></mml:mrow></mml:math></disp-formula></p>
<p><inline-formula id="ieqn-9"><mml:math id="mml-ieqn-9"><mml:msub><mml:mrow><mml:mtext>T</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext>comp</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> is identical for all the communication protocols. <inline-formula id="ieqn-10"><mml:math id="mml-ieqn-10"><mml:msub><mml:mrow><mml:msup><mml:mrow><mml:mtext>T</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>U</mml:mtext></mml:mrow></mml:mrow></mml:msup></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mn>1</mml:mn><mml:mspace width="thinmathspace" /><mml:mi>c</mml:mi><mml:mi>y</mml:mi><mml:mi>c</mml:mi><mml:mi>l</mml:mi><mml:mi>e</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-11"><mml:math id="mml-ieqn-11"><mml:msub><mml:mrow><mml:msup><mml:mrow><mml:mtext>T</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>U</mml:mtext></mml:mrow></mml:mrow></mml:msup></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext>comm</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> are the total time for 1 cycle and the communication time for UART communication protocol.</p>
<p>In case of Ethernet Communication for 1 cycle,
<disp-formula id="eqn-5"><label>(5)</label><mml:math id="mml-eqn-5" display="block"><mml:msub><mml:mrow><mml:msup><mml:mrow><mml:mtext>T</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>E</mml:mtext></mml:mrow></mml:mrow></mml:msup></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mn>1</mml:mn><mml:mrow><mml:mtext>&#xA0;cycle</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mn>1.35</mml:mn><mml:mrow><mml:mtext>&#xA0;s</mml:mtext></mml:mrow></mml:math></disp-formula>
<disp-formula id="eqn-6"><label>(6)</label><mml:math id="mml-eqn-6" display="block"><mml:msub><mml:mrow><mml:msup><mml:mrow><mml:mtext>T</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>E</mml:mtext></mml:mrow></mml:mrow></mml:msup></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext>comm</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mrow><mml:mtext>No</mml:mtext></mml:mrow><mml:mo>.</mml:mo><mml:mrow><mml:mtext>of bits transferred</mml:mtext></mml:mrow></mml:mrow><mml:mrow><mml:mtext>Speed</mml:mtext></mml:mrow></mml:mfrac><mml:mo>=</mml:mo><mml:mfrac><mml:mn>28864</mml:mn><mml:mrow><mml:mn>100</mml:mn><mml:mtext>&#xA0;</mml:mtext><mml:mi>x</mml:mi><mml:mtext>&#xA0;</mml:mtext><mml:msup><mml:mn>10</mml:mn><mml:mrow><mml:mn>6</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:mfrac><mml:mo>=</mml:mo><mml:mn>0.288</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:msup><mml:mn>10</mml:mn><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>3</mml:mn></mml:mrow></mml:msup><mml:mrow><mml:mtext>&#xA0;s</mml:mtext></mml:mrow></mml:math></disp-formula></p>
<p>From <xref ref-type="disp-formula" rid="eqn-1">Eq. (1)</xref> we know that
<disp-formula id="eqn-7"><label>(7)</label><mml:math id="mml-eqn-7" display="block"><mml:msub><mml:mrow><mml:msup><mml:mrow><mml:mtext>T</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>E</mml:mtext></mml:mrow></mml:mrow></mml:msup></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mn>1</mml:mn><mml:mrow><mml:mtext>&#xA0;cycle</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:msup><mml:mrow><mml:mtext>T</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>E</mml:mtext></mml:mrow></mml:mrow></mml:msup></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext>comm</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mrow><mml:mtext>T</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext>comp</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mrow><mml:msup><mml:mrow><mml:mtext>T</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>E</mml:mtext></mml:mrow></mml:mrow></mml:msup></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext>overhead</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msub></mml:math></disp-formula>
<disp-formula id="eqn-8"><label>(8)</label><mml:math id="mml-eqn-8" display="block"><mml:msub><mml:mrow><mml:msup><mml:mrow><mml:mtext>T</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>E</mml:mtext></mml:mrow></mml:mrow></mml:msup></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext>overhead</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mn>0.128</mml:mn><mml:mrow><mml:mtext>&#xA0;s&#xA0;</mml:mtext></mml:mrow></mml:math></disp-formula>
<disp-formula id="eqn-9"><label>(9)</label><mml:math id="mml-eqn-9" display="block"><mml:msub><mml:mrow><mml:msup><mml:mrow><mml:mtext>T</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>E</mml:mtext></mml:mrow></mml:mrow></mml:msup></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext>overhead</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msub><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:mrow><mml:mtext>frame</mml:mtext></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:msub><mml:mrow><mml:msup><mml:mrow><mml:mtext>T</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>E</mml:mtext></mml:mrow></mml:mrow></mml:msup></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext>overhead</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msub><mml:mrow><mml:mrow><mml:mtext>No</mml:mtext></mml:mrow><mml:mo>.</mml:mo><mml:mrow><mml:mtext>of Ethernet frames</mml:mtext></mml:mrow></mml:mrow></mml:mfrac><mml:mo>=</mml:mo><mml:mfrac><mml:mn>0.128</mml:mn><mml:mn>4</mml:mn></mml:mfrac><mml:mo>=</mml:mo><mml:mn>0.032</mml:mn><mml:mrow><mml:mtext>&#xA0;s</mml:mtext></mml:mrow></mml:math></disp-formula></p>
<p><inline-formula id="ieqn-12"><mml:math id="mml-ieqn-12"><mml:msub><mml:mrow><mml:msup><mml:mrow><mml:mtext>T</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>E</mml:mtext></mml:mrow></mml:mrow></mml:msup></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mn>1</mml:mn><mml:mspace width="thinmathspace" /><mml:mrow><mml:mtext>cycle</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula>, <inline-formula id="ieqn-13"><mml:math id="mml-ieqn-13"><mml:msub><mml:mrow><mml:msup><mml:mrow><mml:mtext>T</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>E</mml:mtext></mml:mrow></mml:mrow></mml:msup></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext>comm</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-14"><mml:math id="mml-ieqn-14"><mml:msub><mml:mrow><mml:msup><mml:mrow><mml:mtext>T</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>E</mml:mtext></mml:mrow></mml:mrow></mml:msup></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext>overhead</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> are the total time for 1 cycle, communication time and the overhead time in case of Ethernet communication protocol.</p>
<p>From the above analogy, it is very clear that to find the tipping point where Ethernet communication will have a better hand compared to UART, the accountability of overhead of 0.032 s should be looked into. With these overheads in mind, one can propose which communication protocol should be used and can be predicted with simple calculations. The tipping point occurs when the time taken in 1 cycle using UART is same as the time taken in 1 cycle using Ethernet in transferring some data. This can be shown as
<disp-formula id="eqn-10"><label>(10)</label><mml:math id="mml-eqn-10" display="block"><mml:msub><mml:mrow><mml:msup><mml:mrow><mml:mtext>T</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>U</mml:mtext></mml:mrow></mml:mrow></mml:msup></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mn>1</mml:mn><mml:mspace width="thinmathspace" /><mml:mrow><mml:mtext>cycle</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:msup><mml:mrow><mml:mtext>T</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>E</mml:mtext></mml:mrow></mml:mrow></mml:msup></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mn>1</mml:mn><mml:mrow><mml:mtext>&#xA0;cycle</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msub></mml:math></disp-formula>
<disp-formula id="eqn-11"><label>(11)</label><mml:math id="mml-eqn-11" display="block"><mml:msub><mml:mrow><mml:msup><mml:mrow><mml:mtext>T</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>U</mml:mtext></mml:mrow></mml:mrow></mml:msup></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext>comm</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:msup><mml:mrow><mml:mtext>T</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>E</mml:mtext></mml:mrow></mml:mrow></mml:msup></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext>comm</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mrow><mml:msup><mml:mrow><mml:mtext>T</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>E</mml:mtext></mml:mrow></mml:mrow></mml:msup></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext>overhead</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msub></mml:math></disp-formula></p>
<p>In <xref ref-type="disp-formula" rid="eqn-11">Eq. (11)</xref>, we omitted <inline-formula id="ieqn-15"><mml:math id="mml-ieqn-15"><mml:msub><mml:mrow><mml:mtext>T</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext>comp</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula>, because it is same both for UART and Ethernet. Using <xref ref-type="disp-formula" rid="eqn-11">Eq. (11)</xref>, we can predict the tipping point by calculating <inline-formula id="ieqn-16"><mml:math id="mml-ieqn-16"><mml:msub><mml:mrow><mml:msup><mml:mrow><mml:mtext>T</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>U</mml:mtext></mml:mrow></mml:mrow></mml:msup></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext>comm</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-17"><mml:math id="mml-ieqn-17"><mml:msub><mml:mrow><mml:msup><mml:mrow><mml:mtext>T</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>E</mml:mtext></mml:mrow></mml:mrow></mml:msup></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext>comm</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> for different nanoparticles (number of bits transferred varies with the nanoparticles size).</p>
<p>To explain it experimentally, <xref ref-type="table" rid="table-5">Table 5</xref> shows <inline-formula id="ieqn-18"><mml:math id="mml-ieqn-18"><mml:msub><mml:mrow><mml:mtext>T</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext>comm</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> &#x002B; <inline-formula id="ieqn-19"><mml:math id="mml-ieqn-19"><mml:msub><mml:mrow><mml:mtext>T</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext>overhead</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> time for different sizes of nanoparticles. It is evident from <xref ref-type="table" rid="table-5">Table 5</xref> that, for a nanoparticle of approximately 430 atoms, the <inline-formula id="ieqn-20"><mml:math id="mml-ieqn-20"><mml:msub><mml:mrow><mml:mtext>T</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext>comm</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> &#x002B; <inline-formula id="ieqn-21"><mml:math id="mml-ieqn-21"><mml:msub><mml:mrow><mml:mtext>T</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext>overhead</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> times for UART and Ethernet are almost the same. However, when the size exceeds 430 atoms, <inline-formula id="ieqn-22"><mml:math id="mml-ieqn-22"><mml:msub><mml:mrow><mml:msup><mml:mrow><mml:mtext>T</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>U</mml:mtext></mml:mrow></mml:mrow></mml:msup></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext>comm</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> time surpasses <inline-formula id="ieqn-23"><mml:math id="mml-ieqn-23"><mml:msub><mml:mrow><mml:msup><mml:mrow><mml:mtext>T</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>E</mml:mtext></mml:mrow></mml:mrow></mml:msup></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext>comm</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> &#x002B; <inline-formula id="ieqn-24"><mml:math id="mml-ieqn-24"><mml:msub><mml:mrow><mml:msup><mml:mrow><mml:mtext>T</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>E</mml:mtext></mml:mrow></mml:mrow></mml:msup></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext>overhead</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> time resulting in better communication efficiency for Ethernet. The same analogy can be extended to PCIe communication as well. PCIe operates at a speed of 5 GTps (Giga-transfer per Second) (equivalent to 31 Gbps for a Gen2 8-lane PCIe bus). When <inline-formula id="ieqn-25"><mml:math id="mml-ieqn-25"><mml:msub><mml:mrow><mml:msup><mml:mrow><mml:mtext>T</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>E</mml:mtext></mml:mrow></mml:mrow></mml:msup></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext>comm</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> &#x002B; <inline-formula id="ieqn-26"><mml:math id="mml-ieqn-26"><mml:msub><mml:mrow><mml:msup><mml:mrow><mml:mtext>T</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>E</mml:mtext></mml:mrow></mml:mrow></mml:msup></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext>overhead</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> reaches to 0.2 seconds, it equals to the <inline-formula id="ieqn-27"><mml:math id="mml-ieqn-27"><mml:msub><mml:mrow><mml:msup><mml:mrow><mml:mtext>T</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>P</mml:mtext></mml:mrow></mml:mrow></mml:msup></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext>comm</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> &#x002B; <inline-formula id="ieqn-28"><mml:math id="mml-ieqn-28"><mml:msub><mml:mrow><mml:msup><mml:mrow><mml:mtext>T</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>P</mml:mtext></mml:mrow></mml:mrow></mml:msup></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext>overhead</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> time of PCIe communication. Estimated calculations suggest that this occurs when 3.5 million bits are transferred per cycle, which approximately corresponds to a nanoparticles size of 18,000 atoms. Practically this size does not fall under the category of nanoparticles. Therefore, for MD calculations of nanoparticles or nanoalloys, Ethernet communication offers better efficiency.</p>
<table-wrap id="table-5">
<label>Table 5</label>
<caption>
<title>Tipping point calculation for different protocols</title>
</caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th rowspan="2">Size of nanoparticles</th>
<th rowspan="2">No. of bits per cycle</th>
<th rowspan="2">Type of communication</th>
<th align="center" colspan="3"><inline-formula id="ieqn-29"><mml:math id="mml-ieqn-29"><mml:msub><mml:mrow><mml:mtext>T</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext>comm</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> &#x002B; <inline-formula id="ieqn-30"><mml:math id="mml-ieqn-30"><mml:msub><mml:mrow><mml:mtext>T</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext>overhead</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> in seconds</th>
</tr>
<tr>
<th>1 cycle</th>
<th>10 cycles</th>
<th>100 cycles</th>
</tr>
</thead>
<tbody>
<tr>
<td rowspan="3">55 atoms</td>
<td rowspan="3">10,624</td>
<td>UART</td>
<td>0.024</td>
<td>0.235</td>
<td>2.329</td>
</tr>
<tr>
<td>Ethernet</td>
<td>0.117</td>
<td>1.179</td>
<td>12.082</td>
</tr>
<tr>
<td>PCIe</td>
<td>0.174</td>
<td>1.79</td>
<td>17.775</td>
</tr>
<tr>
<td rowspan="3">147 atoms</td>
<td rowspan="3">28,288</td>
<td>UART</td>
<td>0.054</td>
<td>0.515</td>
<td>5.138</td>
</tr>
<tr>
<td>Ethernet</td>
<td>0.133</td>
<td>1.335</td>
<td>12.096</td>
</tr>
<tr>
<td>PCIe</td>
<td>0.178</td>
<td>1.771</td>
<td>18.160</td>
</tr>
<tr>
<td rowspan="3">309 atoms</td>
<td rowspan="3">59,392</td>
<td>UART</td>
<td>0.103</td>
<td>1.017</td>
<td>10.103</td>
</tr>
<tr>
<td>Ethernet</td>
<td>0.130</td>
<td>1.280</td>
<td>12.120</td>
</tr>
<tr>
<td>PCIe</td>
<td>0.183</td>
<td>1.856</td>
<td>18.670</td>
</tr>
<tr>
<td rowspan="3"><bold>430 atoms</bold></td>
<td rowspan="3"><bold>82,624</bold></td>
<td><bold>UART</bold></td>
<td><bold>0.133</bold></td>
<td><bold>1.346</bold></td>
<td><bold>13.512</bold></td>
</tr>
<tr>
<td><bold>Ethernet</bold></td>
<td><bold>0.133</bold></td>
<td><bold>1.351</bold></td>
<td><bold>13.426</bold></td>
</tr>
<tr>
<td>PCIe</td>
<td>0.188</td>
<td>1.863</td>
<td>18.810</td>
</tr>
<tr>
<td rowspan="3">561 atoms</td>
<td rowspan="3">107,776</td>
<td>UART</td>
<td>0.180</td>
<td>1.811</td>
<td>18.085</td>
</tr>
<tr>
<td>Ethernet</td>
<td>0.136</td>
<td>1.374</td>
<td>13.965</td>
</tr>
<tr>
<td>PCIe</td>
<td>0.197</td>
<td>1.838</td>
<td>18.209</td>
</tr>
<tr>
<td rowspan="3">923 atoms</td>
<td rowspan="3">177,280</td>
<td>UART</td>
<td>0.293</td>
<td>2.933</td>
<td>29.409</td>
</tr>
<tr>
<td>Ethernet</td>
<td>0.137</td>
<td>1.375</td>
<td>14.023</td>
</tr>
<tr>
<td>PCIe</td>
<td>0.200</td>
<td>1.986</td>
<td>19.110</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s6">
<label>6</label>
<title>Conclusion</title>
<p>The ANN-based MD Simulation for Au<sub>147</sub> and Au<sub>309</sub> was implemented on the Xilinx Kintex-7 KC705 evaluation FPGA board. A hardware accelerator module with new communication strategies is proposed and implemented for 50,000 MD cycles. The computation time for 50,000 MD cycles is 17.54 and 18.7 h for UART and Ethernet communication, respectively for Au<sub>147</sub> and 64.36 &#x0026; 65.56 h respectively for Au<sub>309</sub>. Compared to the conventional HPC server, the proposed methodology has improved the computation time by 1.65 (1.81) times in UART and 1.55 (1.77) times in Ethernet communication for Au<sub>147</sub> (Au<sub>309</sub>) nanoparticles. The actual MD simulation requires more than 1 million cycles, so this computation time difference becomes more significant. The proposed systems significantly reduce resource utilization, resulting in decreased on-chip power consumption. In the implemented system, on-chip power consumption measured from Xilinx Vivado was 5.9 (6.1) Watts for UART and 6.0 (6.5) Watts for Ethernet, respectively, for Au<sub>147</sub> (Au<sub>309</sub>). Compared to conventional PCIe, on-chip power is reduced by 33% and 32% in UART and Ethernet, respectively. From this, we concluded that where the nanoparticle size is larger than 430 atoms, Ethernet communication is preferable in comparison to UART and PCIe and if the nanoparticle size is less than 430 atoms, then UART is more efficient. Both UART and Ethernet communication are robust, hot-pluggable, and user-friendly. This can lead to low-cost HPC for students and researchers to explore nanoparticles of experimental relevance. This application paves the way for the development of a Lab-on-a-Chip platform for the computation of IAP in the future.</p>
</sec>
</body>
<back>
<ack>
<p>Ankitkumar Patel thanks Kishore Reddy Kurapati and Dharmendra Kartikey for their valuable suggestions and discussions during this research. The authors also extend their thanks to Kritika Bhardwaj for providing assistance related to PCIe work. Additionally, Ankitkumar Patel acknowledges IIT Indore for providing research fellowship and facilities.</p>
</ack>
<sec><title>Funding Statement</title>
<p>The authors received no specific funding for this study.</p>
</sec>
<sec><title>Author Contributions</title>
<p><bold>Ankitkumar Patel:</bold> Investigation, Data collection, Methodology, Analysis and Results interpretation, Draft manuscript preparation. <bold>Srivathsan Vasudevan:</bold> Project Design and Proposal, Conceptualization, Resources, Supervision, Validation, Draft review and editing. <bold>Satya Bulusu:</bold> Project Design and Proposal, Conceptualization, Resources, Supervision, Validation, Draft review and editing. All authors reviewed the results and approved the final version of the manuscript.</p>
</sec>
<sec sec-type="data-availability"><title>Availability of Data and Materials</title>
<p>Not applicable.</p>
</sec>
<sec><title>Ethics Approval</title>
<p>Not applicable.</p>
</sec>
<sec sec-type="COI-statement"><title>Conflicts of Interest</title>
<p>The authors declare that they have no conflicts of interest to report regarding the present study.</p>
</sec>
<ref-list content-type="authoryear">
<title>References</title>
<ref id="ref-1"><label>[1]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>L.</given-names> <surname>Bin</surname></string-name>, <string-name><given-names>Z.</given-names> <surname>Dawid</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Henry</surname></string-name>, and <string-name><given-names>G.</given-names> <surname>Laxmi</surname></string-name></person-group>, &#x201C;<article-title>Accelerating high performance computing applications: Using CPUs, GPUs, Hybrid CPU/GPU, and FPGAs</article-title>,&#x201D; in <conf-name>Proc. 13th Int. Conf. PDCAT</conf-name>, <publisher-loc>Beijing, China</publisher-loc>, <year>Dec. 2012</year>, pp. <fpage>337</fpage>&#x2013;<lpage>342</lpage>.</mixed-citation></ref>
<ref id="ref-2"><label>[2]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>L.</given-names> <surname>Xu</surname></string-name>, <string-name><given-names>O.</given-names> <surname>Hibat-Allah</surname></string-name>, <string-name><given-names>G.</given-names> <surname>Abdelouahed</surname></string-name>, <string-name><given-names>W.</given-names> <surname>Li</surname></string-name>, and <string-name><given-names>C.</given-names> <surname>Mohamed</surname></string-name></person-group>, &#x201C;<article-title>A hybrid GPU-FPGA based design methodology for enhancing machine learning applications performance</article-title>,&#x201D; <source>J. Ambient Intell. Humaniz. Comput.</source>, vol. <volume>11</volume>, no. <issue>6</issue>, pp. <fpage>2309</fpage>&#x2013;<lpage>2323</lpage>, <year>Jun. 2020</year>. doi: <pub-id pub-id-type="doi">10.1007/s12652-019-01357-4</pub-id>.</mixed-citation></ref>
<ref id="ref-3"><label>[3]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>V.</given-names> <surname>Mario</surname></string-name> and <string-name><given-names>N.</given-names> <surname>Horacio</surname></string-name></person-group>, &#x201C;<article-title>Trends of CPU, GPU and FPGA for high-performance computing</article-title>,&#x201D; in <conf-name>Proc. 24th Int. Conf. FPL</conf-name>, <publisher-loc>Munich, Germany</publisher-loc>, <year>Sep. 2014</year>, pp. <fpage>1</fpage>&#x2013;<lpage>6</lpage>.</mixed-citation></ref>
<ref id="ref-4"><label>[4]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>K.</given-names> <surname>Marcin</surname></string-name> and <string-name><given-names>D.</given-names> <surname>Grzegorz</surname></string-name></person-group>, &#x201C;<article-title>Simple communication with FPGA device over ethernet interface</article-title>,&#x201D; in <conf-name>Proc. 20th Int. Conf. CN</conf-name>, <publisher-loc>Lwowek Slaski, Poland</publisher-loc>, <year>Jun. 2013</year>, pp. <fpage>290</fpage>&#x2013;<lpage>299</lpage>.</mixed-citation></ref>
<ref id="ref-5"><label>[5]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>Y.</given-names> <surname>Li</surname></string-name>, <string-name><given-names>X.</given-names> <surname>Zhao</surname></string-name>, and <string-name><given-names>T.</given-names> <surname>Cheng</surname></string-name></person-group>, &#x201C;<article-title>Heterogeneous computing platform based on CPU&#x002B;FPGA and working modes</article-title>,&#x201D; in <conf-name>Proc. 12th Int. Conf. CIS</conf-name>, <publisher-loc>Wuxi, China</publisher-loc>, <year>Dec. 2016</year>, pp. <fpage>669</fpage>&#x2013;<lpage>672</lpage>.</mixed-citation></ref>
<ref id="ref-6"><label>[6]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>L.</given-names> <surname>Shi</surname></string-name>, <string-name><given-names>H.</given-names> <surname>Chen</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Sun</surname></string-name>, and <string-name><given-names>K.</given-names> <surname>Li</surname></string-name></person-group>, &#x201C;<article-title>vCUDA: GPU-accelerated high-performance computing in virtual machines</article-title>,&#x201D; <source>IEEE Trans. Comput.</source>, vol. <volume>61</volume>, no. <issue>6</issue>, pp. <fpage>804</fpage>&#x2013;<lpage>816</lpage>, <year>Jun. 2011</year>. doi: <pub-id pub-id-type="doi">10.1109/TC.2011.112</pub-id>.</mixed-citation></ref>
<ref id="ref-7"><label>[7]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>J.</given-names> <surname>Fang</surname></string-name>, <string-name><given-names>K.</given-names> <surname>Zhou</surname></string-name>, <string-name><given-names>M. Y.</given-names> <surname>Zhang</surname></string-name>, and <string-name><given-names>W.</given-names> <surname>Xiang</surname></string-name></person-group>, &#x201C;<article-title>Resource scheduling strategy for performance optimization based on heterogeneous CPU-GPU platform</article-title>,&#x201D; <source>Comput. Mater. Contin.</source>, vol. <volume>73</volume>, no. <issue>1</issue>, pp. <fpage>1621</fpage>&#x2013;<lpage>1635</lpage>, <year>Jun. 2022</year>. doi: <pub-id pub-id-type="doi">10.32604/cmc.2022.027147</pub-id>.</mixed-citation></ref>
<ref id="ref-8"><label>[8]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>M.</given-names> <surname>Herbordt</surname></string-name> <etal>et al.</etal></person-group>, &#x201C;<article-title>Achieving high performance with FPGA-based computing</article-title>,&#x201D; <source>Computer</source>, vol. <volume>40</volume>, no. <issue>3</issue>, pp. <fpage>50</fpage>&#x2013;<lpage>57</lpage>, <year>Mar. 2007</year>. doi: <pub-id pub-id-type="doi">10.1109/MC.2007.79</pub-id>; <pub-id pub-id-type="pmid">21603088</pub-id></mixed-citation></ref>
<ref id="ref-9"><label>[9]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Jindal</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Chiriki</surname></string-name>, and <string-name><given-names>S.</given-names> <surname>Bulusu</surname></string-name></person-group>, &#x201C;<article-title>Spherical harmonics based descriptor for neural network potentials: Structure and dynamics of Au147 nanocluster</article-title>,&#x201D; <source>J. Chem. Phys.</source>, vol. <volume>146</volume>, no. <issue>20</issue>, <year>2017</year>, <comment>Art. no. 3168</comment>. doi: <pub-id pub-id-type="doi">10.1063/1.4983392</pub-id>; <pub-id pub-id-type="pmid">28571343</pub-id></mixed-citation></ref>
<ref id="ref-10"><label>[10]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>G.</given-names> <surname>Helmut</surname></string-name>, <string-name><given-names>H.</given-names> <surname>Helmut</surname></string-name>, <string-name><given-names>W.</given-names> <surname>Andreas</surname></string-name>, and <string-name><given-names>S.</given-names> <surname>Klaus</surname></string-name></person-group>, &#x201C;<article-title>Generalized Verlet algorithm for efficient molecular dynamics simulations with long-range interactions</article-title>,&#x201D; <source>Mol. Simul.</source>, vol. <volume>9</volume>, no. <issue>1&#x2013;3</issue>, pp. <fpage>121</fpage>&#x2013;<lpage>142</lpage>, <year>Jun. 1991</year>. doi: <pub-id pub-id-type="doi">10.1080/08927029108022142</pub-id>.</mixed-citation></ref>
<ref id="ref-11"><label>[11]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>M.</given-names> <surname>Nicos</surname></string-name> and <string-name><given-names>M.</given-names> <surname>Raymond</surname></string-name></person-group>, &#x201C;<article-title>Velocity Verlet algorithm for dissipative-particle-dynamics-based models of suspensions</article-title>,&#x201D; <source>Phys. Rev. E</source>, vol. <volume>59</volume>, no. <issue>3</issue>, pp. <fpage>3733</fpage>&#x2013;<lpage>3736</lpage>, <year>Mar. 1999</year>. doi: <pub-id pub-id-type="doi">10.1103/PhysRevE.59.3733</pub-id>.</mixed-citation></ref>
<ref id="ref-12"><label>[12]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>P&#x00E1;l</surname></string-name> <etal>et al.</etal></person-group>, &#x201C;<article-title>Heterogeneous parallelization and acceleration of molecular dynamics simulations in GROMACS</article-title>,&#x201D; <source>J. Chem. Phys.</source>, vol. <volume>153</volume>, no. <issue>13</issue>, pp. <fpage>121</fpage>&#x2013;<lpage>142</lpage>, <year>Oct. 2020</year>. doi: <pub-id pub-id-type="doi">10.1063/5.0018516</pub-id>; <pub-id pub-id-type="pmid">33032406</pub-id></mixed-citation></ref>
<ref id="ref-13"><label>[13]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>D.</given-names> <surname>Jones</surname></string-name> <etal>et al.</etal></person-group>, &#x201C;<article-title>Accelerators for classical molecular dynamics simulations of biomolecules</article-title>,&#x201D; <source>J. Chem. Theory Comput.</source>, vol. <volume>18</volume>, no. <issue>7</issue>, pp. <fpage>4047</fpage>&#x2013;<lpage>4069</lpage>, <year>Jun. 2022</year>. doi: <pub-id pub-id-type="doi">10.1021/acs.jctc.1c01214</pub-id>; <pub-id pub-id-type="pmid">35710099</pub-id></mixed-citation></ref>
<ref id="ref-14"><label>[14]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>N.</given-names> <surname>Kondratyuk</surname></string-name>, <string-name><given-names>V.</given-names> <surname>Nikolskiy</surname></string-name>, <string-name><given-names>D.</given-names> <surname>Pavlov</surname></string-name>, and <string-name><given-names>V.</given-names> <surname>Stegailov</surname></string-name></person-group>, &#x201C;<article-title>GPU-accelerated molecular dynamics: State-of-art software performance and porting from Nvidia CUDA to AMD HIP</article-title>,&#x201D; <source>Int. J. High Perform. Comput. Appl.</source>, vol. <volume>35</volume>, no. <issue>4</issue>, pp. <fpage>312</fpage>&#x2013;<lpage>324</lpage>, <year>Apr. 2021</year>. doi: <pub-id pub-id-type="doi">10.1177/10943420211008288</pub-id>.</mixed-citation></ref>
<ref id="ref-15"><label>[15]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>Y.</given-names> <surname>Chen</surname></string-name> <etal>et al.</etal></person-group>, &#x201C;<article-title>Fully integrated FPGA molecular dynamics simulations</article-title>,&#x201D; in <conf-name>Proc. SC 19 Int. Conf. High Perform. Comput., Netw., Storage, Anal.</conf-name>, <publisher-loc>New York, NY, USA</publisher-loc>, <year>Nov. 2019</year>, pp. <fpage>1</fpage>&#x2013;<lpage>31</lpage>.</mixed-citation></ref>
<ref id="ref-16"><label>[16]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Z.</given-names> <surname>Li</surname></string-name>, <string-name><given-names>Q.</given-names> <surname>Li</surname></string-name>, <string-name><given-names>H.</given-names> <surname>Liu</surname></string-name>, and <string-name><given-names>Z.</given-names> <surname>Zhao</surname></string-name></person-group>, &#x201C;<article-title>FPGA optimized accelerator of DCNN with fast data readout and multiplier sharing strategy</article-title>,&#x201D; <source>Comput. Mater. Contin.</source>, vol. <volume>77</volume>, no. <issue>3</issue>, pp. <fpage>3237</fpage>&#x2013;<lpage>3263</lpage>, <year>Dec. 2023</year>. doi: <pub-id pub-id-type="doi">10.32604/cmc.2023.045948</pub-id>.</mixed-citation></ref>
<ref id="ref-17"><label>[17]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Ronald</surname></string-name>, <string-name><given-names>G.</given-names> <surname>Maya</surname></string-name>, <string-name><given-names>T.</given-names> <surname>Frans</surname></string-name>, and <string-name><given-names>P.</given-names> <surname>Viktor</surname></string-name></person-group>, &#x201C;<article-title>Hardware/software approach to molecular dynamics on reconfigurable computers</article-title>,&#x201D; in <conf-name>2006 14th Annual IEEE Symp. Field-Programmable Custom Comput. Mach.</conf-name>, <publisher-loc>Napa, CA, USA</publisher-loc>, <publisher-name>IEEE</publisher-name>, <year>2006</year>. <comment>Accessed: Jan. 20, 2024</comment>. [Online]. Available:<comment> </comment><ext-link ext-link-type="uri" xlink:href="https://ieeexplore.ieee.org/abstract/document/4020892">https://ieeexplore.ieee.org/abstract/document/4020892</ext-link></mixed-citation></ref>
<ref id="ref-18"><label>[18]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>C.</given-names> <surname>Pascoe</surname></string-name>, <string-name><given-names>L.</given-names> <surname>Stewart</surname></string-name>, <string-name><given-names>B.</given-names> <surname>Sherman</surname></string-name>, <string-name><given-names>V.</given-names> <surname>Sachdeva</surname></string-name>, and <string-name><given-names>M.</given-names> <surname>Herbordt</surname></string-name></person-group>, &#x201C;<article-title>Execution of complete molecular dynamics simulations on multiple FPGAs</article-title>,&#x201D; in <conf-name>Proc. IEEE Conf. High Perform. Extreme Comput. (HPEC)</conf-name>, <publisher-loc>Waltham, MA, USA</publisher-loc>, <year>Sep. 22&#x2013;24, 2020</year>, pp. <fpage>1</fpage>&#x2013;<lpage>2</lpage>.</mixed-citation></ref>
<ref id="ref-19"><label>[19]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Bulusu</surname></string-name> and <string-name><given-names>S.</given-names> <surname>Vasudevan</surname></string-name></person-group>, &#x201C;<article-title>FPGA accelerator for machine learning interatomic potential-based molecular dynamics of gold nanoparticles</article-title>,&#x201D; <source>IEEE Access</source>, vol. <volume>10</volume>, pp. <fpage>40338</fpage>&#x2013;<lpage>40347</lpage>, <year>Apr. 2022</year>. doi: <pub-id pub-id-type="doi">10.1109/ACCESS.2022.3165650</pub-id>.</mixed-citation></ref>
<ref id="ref-20"><label>[20]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><collab>Inc Xilinx</collab></person-group>, &#x201C;<article-title>KC705 evaluation board for the Kintex-7 FPGA user guide</article-title>,&#x201D; <comment>version: UG810 v1.9, USA, Feb. 4, 2019. Accessed: Mar. 13, 2023</comment>. [Online]. Available: <ext-link ext-link-type="uri" xlink:href="https://docs.amd.com/v/u/en-US/ug810_KC705_Eval_Bd">https://docs.amd.com/v/u/en-US/ug810_KC705_Eval_Bd</ext-link></mixed-citation></ref>
<ref id="ref-21"><label>[21]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>T.</given-names> <surname>Jason</surname></string-name>, <string-name><given-names>I.</given-names> <surname>Anderson</surname></string-name>, and <string-name><given-names>K.</given-names> <surname>Mohammed</surname></string-name></person-group>, &#x201C;<article-title>Soft-core processors for embedded systems</article-title>,&#x201D; in <conf-name>Proc. Int. Conf. Microelectron. (ICM)</conf-name>, <publisher-loc>Dhahran, Saudi Arabia</publisher-loc>, <year>Sep. 2006</year>, pp. <fpage>170</fpage>&#x2013;<lpage>173</lpage>.</mixed-citation></ref>
<ref id="ref-22"><label>[22]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><collab>Inc Xilinx</collab></person-group>, &#x201C;<article-title>MicroBlaze processor reference guide</article-title>,&#x201D; <comment>version: UG984 v2021.2, USA, Oct. 27, 2021. Accessed: Nov. 15, 2022</comment>. [Online]. Available: <ext-link ext-link-type="uri" xlink:href="https://www.xilinx.com/content/dam/xilinx/support/documents/sw_manuals/xilinx2021_2/ug984-vivado-microblaze-ref.pdf">https://www.xilinx.com/content/dam/xilinx/support/documents/sw_manuals/xilinx2021_2/ug984-vivado-microblaze-ref.pdf</ext-link></mixed-citation></ref>
<ref id="ref-23"><label>[23]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><collab>Inc Xilinx</collab></person-group>, &#x201C;<article-title>AXI UART Lite LogiCORE IP product guide</article-title>,&#x201D; <comment>version: PG142 v2017. 2, USA, Apr. 5, 2017. Accessed: Dec. 6, 2022</comment>. [Online]. Available: <ext-link ext-link-type="uri" xlink:href="https://docs.amd.com/v/u/en-US/pg142-axi-uartlite">https://docs.amd.com/v/u/en-US/pg142-axi-uartlite</ext-link></mixed-citation></ref>
<ref id="ref-24"><label>[24]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><collab>Inc Xilinx</collab></person-group>, &#x201C;<article-title>AXI Ethernet lite MAC LogiCORE IP product guide</article-title>,&#x201D; <comment>version: PG135 v2021. 3, USA, Nov. 2, 2021. Accessed: Feb. 2, 2023</comment>. [Online]. Available: <ext-link ext-link-type="uri" xlink:href="https://docs.amd.com/r/en-US/pg135-axi-ethernetlite/AXI-Ethernet-Lite-MAC-v3.0-LogiCORE-IP-Product-Guide">https://docs.amd.com/r/en-US/pg135-axi-ethernetlite/AXI-Ethernet-Lite-MAC-v3.0-LogiCORE-IP-Product-Guide</ext-link></mixed-citation></ref>
<ref id="ref-25"><label>[25]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>M.</given-names> <surname>Gokhale</surname></string-name> and <string-name><given-names>L.</given-names> <surname>Shannon</surname></string-name></person-group>, &#x201C;<article-title>FPGA Computing</article-title>,&#x201D; <source>IEEE Micro</source>, vol. <volume>41</volume>, no. <issue>4</issue>, pp. <fpage>6</fpage>&#x2013;<lpage>7</lpage>, <year>Jul. 2021</year>. doi: <pub-id pub-id-type="doi">10.1109/MM.2021.3088975</pub-id>.</mixed-citation></ref>
<ref id="ref-26"><label>[26]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><collab>Inc Xilinx</collab></person-group>, &#x201C;<article-title>Vivado design suite user guide: Embedded processor hardware design</article-title>,&#x201D; <comment>version: UG898 v2021.1, USA, Jun. 16, 2021. Accessed: Jan. 17, 2023</comment>. [Online]. Available: <ext-link ext-link-type="uri" xlink:href="https://docs.amd.com/v/u/2021.1-English/ug898-vivado-embedded-design">https://docs.amd.com/v/u/2021.1-English/ug898-vivado-embedded-design</ext-link></mixed-citation></ref>
<ref id="ref-27"><label>[27]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><collab>Inc Xilinx</collab></person-group>, &#x201C;<article-title>UltraFast embedded design methodology guide</article-title>,&#x201D; <comment>version: UG1046 v2018.2.3, USA, Apr. 20, 2018. Accessed: Aug. 6, 2023</comment>. <comment>[Online]. Available: </comment><ext-link ext-link-type="uri" xlink:href="https://docs.amd.com/v/u/en-US/ug1046-ultrafast-design-methodology-guide">https://docs.amd.com/v/u/en-US/ug1046-ultrafast-design-methodology-guide</ext-link></mixed-citation></ref>
</ref-list>
</back></article>