<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="review-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Comput. Sci.</journal-id>
<journal-title>Frontiers in Computer Science</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Comput. Sci.</abbrev-journal-title>
<issn pub-type="epub">2624-9898</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fcomp.2025.1538277</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Computer Science</subject>
<subj-group>
<subject>Review</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Intelligent data analysis in edge computing with large language models: applications, challenges, and future directions</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<name><surname>Wang</surname> <given-names>Xuanzheng</given-names></name>
<xref ref-type="corresp" rid="c001"><sup>&#x0002A;</sup></xref>
<role content-type="https://credit.niso.org/contributor-roles/conceptualization/"/>
<role content-type="https://credit.niso.org/contributor-roles/resources/"/>
<role content-type="https://credit.niso.org/contributor-roles/supervision/"/>
<role content-type="https://credit.niso.org/contributor-roles/writing-original-draft/"/>
<role content-type="https://credit.niso.org/contributor-roles/writing-review-editing/"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Xu</surname> <given-names>Zhipeng</given-names></name>
<uri xlink:href="http://loop.frontiersin.org/people/2911189/overview"/>
<role content-type="https://credit.niso.org/contributor-roles/methodology/"/>
<role content-type="https://credit.niso.org/contributor-roles/visualization/"/>
<role content-type="https://credit.niso.org/contributor-roles/writing-original-draft/"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Sui</surname> <given-names>Xingfei</given-names></name>
<role content-type="https://credit.niso.org/contributor-roles/software/"/>
<role content-type="https://credit.niso.org/contributor-roles/validation/"/>
<role content-type="https://credit.niso.org/contributor-roles/visualization/"/>
<role content-type="https://credit.niso.org/contributor-roles/writing-original-draft/"/>
</contrib>
</contrib-group>
<aff><institution>CNOOC Safety Technology Services Co., Ltd.</institution>, <addr-line>Tianjin</addr-line>, <country>China</country></aff>
<author-notes>
<fn fn-type="edited-by"><p>Edited by: Xiaotong Wu, Nanjing Normal University, China</p></fn>
<fn fn-type="edited-by"><p>Reviewed by: Chaokun Zhang, Tianjin University, China</p>
<p>Rajeev Ratna Vallabhuni, Bayview Asset Management, LLC, United States</p>
<p>Xiaoding Wang, Fujian Normal University, China</p></fn>
<corresp id="c001">&#x0002A;Correspondence: Xuanzheng Wang <email>wangxzh20&#x00040;cnooc.com.cn</email></corresp>
</author-notes>
<pub-date pub-type="epub">
<day>14</day>
<month>05</month>
<year>2025</year>
</pub-date>
<pub-date pub-type="collection">
<year>2025</year>
</pub-date>
<volume>7</volume>
<elocation-id>1538277</elocation-id>
<history>
<date date-type="received">
<day>02</day>
<month>12</month>
<year>2024</year>
</date>
<date date-type="accepted">
<day>17</day>
<month>04</month>
<year>2025</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x000A9; 2025 Wang, Xu and Sui.</copyright-statement>
<copyright-year>2025</copyright-year>
<copyright-holder>Wang, Xu and Sui</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/"><p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p></license>
</permissions>
<abstract>
<p>Edge computing has emerged as a vital paradigm for processing data near its source, significantly reducing latency and improving data privacy. Simultaneously, large language models (LLMs) such as GPT-4 and BERT have showcased impressive capabilities in data analysis, natural language processing, and decision-making. This survey explores the intersection of these two domains, specifically focusing on the adaptation and optimization of LLMs for data analysis tasks in edge computing environments. We examine the challenges faced by resource-constrained edge devices, including limited computational power, energy efficiency, and network reliability. Additionally, we discuss how recent advancements in model compression, distributed learning, and edge-friendly architectures are addressing these challenges. Through a comprehensive review of the current research, we analyze the applications, challenges, and future directions of deploying LLMs in edge computing. This analysis aims to facilitate intelligent data analysis across various industries, including healthcare, smart cities, and the internet of things.</p></abstract>
<kwd-group>
<kwd>edge computing</kwd>
<kwd>large language models</kwd>
<kwd>intelligent data analysis</kwd>
<kwd>resource-constrained devices</kwd>
<kwd>edge-friendly architectures</kwd>
</kwd-group>
<counts>
<fig-count count="3"/>
<table-count count="4"/>
<equation-count count="0"/>
<ref-count count="154"/>
<page-count count="21"/>
<word-count count="17259"/>
</counts>
<custom-meta-wrap>
<custom-meta>
<meta-name>section-at-acceptance</meta-name>
<meta-value>Networks and Communications</meta-value>
</custom-meta>
</custom-meta-wrap>
</article-meta>
</front>
<body>
<sec sec-type="intro" id="s1">
<title>1 Introduction</title>
<p>Edge computing has emerged as a powerful paradigm for processing and analyzing data near its source. It effectively addresses limitations such as latency and bandwidth constraints, which are particularly critical in real-time applications (Zhou et al., <xref ref-type="bibr" rid="B150">2024a</xref>; Syu et al., <xref ref-type="bibr" rid="B118">2023</xref>). By distributing computational resources across a network of edge devices, it enables rapid, local processing, minimizing dependency on centralized cloud infrastructures. This shift is particularly important in latency-sensitive applications like autonomous vehicles and smart cities, where instantaneous data processing is essential (Deng et al., <xref ref-type="bibr" rid="B26">2020</xref>).</p>
<p>Simultaneously, deep learning, particularly with large language models (LLMs), has achieved remarkable advancements in natural language processing (NLP) and data analysis, revolutionizing the way machines understand and produce human language (Gangadhar et al., <xref ref-type="bibr" rid="B34">2023</xref>). Models like BERT (Devlin et al., <xref ref-type="bibr" rid="B28">2019</xref>) and GPT-4 (Brown et al., <xref ref-type="bibr" rid="B14">2020</xref>) have exhibited remarkable proficiency in various applications, encompassing language translation, text summarization, sentiment assessment, and intricate data analysis. These advancements have significantly improved the performance of various applications, from virtual assistants to advanced data analytics platforms.</p>
<p>However, deploying these models in edge environments presents a set of technical challenges that must be addressed to fully harness their potential. The high computational demands and substantial memory consumption of LLMs present significant challenges for their implementation on edge devices, which are typically characterized by limited processing capabilities and storage resources. Additionally, the need for frequent model updates to ensure accuracy and relevance intensifies these challenges, as the operational constraints of edge devices may not accommodate the continuous retraining required for LLMs. As a result, directly deploying these models in edge environments is often impractical without significant optimizations (Han et al., <xref ref-type="bibr" rid="B40">2015</xref>). This situation underscores the necessity to explore innovative techniques that can reduce the resource consumption of LLMs while preserving their performance and effectiveness.</p>
<p>The integration of LLMs with edge computing presents both a significant opportunity and a series of technical challenges. On one hand, deploying LLMs on edge devices can facilitate intelligent real-time data processing without relying on centralized cloud infrastructure, which is particularly crucial for applications that prioritize data privacy and security (Zhou et al., <xref ref-type="bibr" rid="B148">2022a</xref>). By processing sensitive data locally, organizations can mitigate the risks associated with data breaches and ensure compliance with privacy regulations. On the other hand, the inherent limitations of edge devices in terms of processing power and storage capacity necessitate the development of strategies focused on model compression and adaptation. Such strategies are essential to reduce the overall resource consumption of LLMs, making them viable for deployment in edge environments (Sun and Ansari, <xref ref-type="bibr" rid="B116">2019</xref>; Chi et al., <xref ref-type="bibr" rid="B21">2024</xref>). Ultimately, addressing these challenges will be key to unlocking the full potential of LLMs in edge computing applications.</p>
<p>As shown in <xref ref-type="table" rid="T1">Table 1</xref>, this table presents an overview of relevant publications focused on various aspects of edge computing and its applications. Based on this foundation, this paper offers a comprehensive analysis of the applications, challenges, and future directions of intelligent data analysis in edge computing using LLMs. This analysis highlights the key obstacles associated with deploying LLMs on resource-constrained edge devices, explores innovative strategies for adapting LLMs to edge environments, and showcases their practical applications across diverse sectors such as healthcare, IoT, and industrial automation.</p>
<table-wrap position="float" id="T1">
<label>Table 1</label>
<caption><p>Summary of relevant publications on edge computing across various domains.</p></caption>
<table frame="box" rules="all">
<thead>
<tr style="background-color:#8f9496;color:#ffffff">
<th valign="top" align="left"><bold>Reference</bold></th>
<th valign="top" align="left"><bold>Survey focus</bold></th>
<th valign="top" align="left"><bold>Research method/technique</bold></th>
<th valign="top" align="left"><bold>Contributions</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">Sun and Ansari (<xref ref-type="bibr" rid="B116">2019</xref>)</td>
<td valign="top" align="left">IoT</td>
<td valign="top" align="left">Architecture design, efficiency analysis.</td>
<td valign="top" align="left">Proposes a novel architecture for Edge IoT to handle data streams efficiently at the mobile edge.</td>
</tr> <tr>
<td valign="top" align="left">Deng et al. (<xref ref-type="bibr" rid="B26">2020</xref>)</td>
<td valign="top" align="left">AI</td>
<td valign="top" align="left">Conceptual framework, comparative analysis.</td>
<td valign="top" align="left">Divides edge intelligence into AI for edge and AI on edge, optimizing solutions in edge environments.</td>
</tr> <tr>
<td valign="top" align="left">Ren et al. (<xref ref-type="bibr" rid="B100">2019b</xref>)</td>
<td valign="top" align="left">Augmented reality</td>
<td valign="top" align="left">Case studies, performance metrics evaluation.</td>
<td valign="top" align="left">Discusses the benefits of edge computing for AR applications, enhancing performance and reducing server reliance.</td>
</tr> <tr>
<td valign="top" align="left">Zhou et al. (<xref ref-type="bibr" rid="B148">2022a</xref>)</td>
<td valign="top" align="left">Privacy security</td>
<td valign="top" align="left">Theoretical framework, privacy analysis.</td>
<td valign="top" align="left">Introduces a novel privacy-preserving framework using local differential privacy in edge computing.</td>
</tr> <tr>
<td valign="top" align="left">Syu et al. (<xref ref-type="bibr" rid="B118">2023</xref>)</td>
<td valign="top" align="left">Consumer electronics</td>
<td valign="top" align="left">Literature review, trend analysis.</td>
<td valign="top" align="left">Provides an overview of AI-driven improvements in latency, robustness, and reliability in consumer electronics.</td>
</tr> <tr>
<td valign="top" align="left">Lu et al. (<xref ref-type="bibr" rid="B81">2023</xref>)</td>
<td valign="top" align="left">Fault diagnosis</td>
<td valign="top" align="left">Methodological analysis, case studies.</td>
<td valign="top" align="left">Analyzes methodologies for signal processing in machine fault diagnosis in IoT contexts.</td>
</tr> <tr>
<td valign="top" align="left">Iftikhar et al. (<xref ref-type="bibr" rid="B44">2023</xref>)</td>
<td valign="top" align="left">Resource management</td>
<td valign="top" align="left">Taxonomy development, literature synthesis.</td>
<td valign="top" align="left">Proposes a taxonomy of AI/ML resource management techniques in fog/edge computing, identifying challenges and future research directions.</td>
</tr></tbody>
</table>
</table-wrap>
<p>This survey offers the following contributions:</p>
<list list-type="bullet">
<list-item><p>We provide a comprehensive review of edge computing and LLM fundamentals, highlighting the distinct advantages and limitations they present for various data analysis tasks.</p></list-item>
<list-item><p>We identify and discuss the primary challenges in deploying LLMs on resource-constrained edge devices, covering areas such as resource limitations, energy efficiency, privacy, and latency requirements.</p></list-item>
<list-item><p>We examine state-of-the-art techniques for making LLMs compatible with edge deployment, including model compression, federated learning, and optimized frameworks.</p></list-item>
<list-item><p>We review practical applications of LLMs in edge computing across various domains, including healthcare, IoT, industrial automation, and consumer electronics, demonstrating the utility of LLMs for edge-based data analysis.</p></list-item>
<list-item><p>We outline emerging trends and future directions in LLM and edge integration, including advancements in Transformer architectures, privacy-preserving techniques, and hardware developments that could shape the future of edge-based intelligent systems.</p></list-item>
</list>
<p>The subsequent sections of this paper are structured in the following manner. Section 2 offers background information on edge computing and LLM architectures, emphasizing their essential characteristics pertinent to edge-based applications. Section 3 examines the specific applications of LLMs in edge computing for data analysis, while Section 4 discusses the challenges of implementing LLMs in edge computing, including resource constraints, energy efficiency, privacy, and latency requirements. Section 5 explores various techniques for edge-compatible LLM deployment, such as model compression, federated learning, and optimized frameworks. Section 6 explores potential avenues for future research. Section 7 wraps up the paper by highlighting essential observations and suggesting areas for further investigation.</p></sec>
<sec id="s2">
<title>2 Background</title>
<sec>
<title>2.1 Edge computing fundamentals</title>
<p>Edge computing is a decentralized computing model that facilitates data processing closer to the data source, rather than relying on centralized cloud servers (Cicconetti et al., <xref ref-type="bibr" rid="B23">2021</xref>). This innovative approach effectively addresses several challenges associated with traditional cloud computing, particularly in environments where speed, efficiency, and security are paramount (Shi et al., <xref ref-type="bibr" rid="B110">2016</xref>; Deng et al., <xref ref-type="bibr" rid="B26">2020</xref>). By performing computations locally on devices such as sensors, gateways, and mobile devices, edge computing significantly reduces latency and bandwidth requirements (Ren et al., <xref ref-type="bibr" rid="B101">2019a</xref>), which are often limitations of cloud-based systems. This significant reduction in latency is crucial for applications that demand real-time responses, such as autonomous vehicles, healthcare monitoring, and smart city infrastructure (Lu et al., <xref ref-type="bibr" rid="B81">2023</xref>; Iftikhar et al., <xref ref-type="bibr" rid="B44">2023</xref>).</p>
<p>By processing data at the edge, applications can enhance operational efficiency, enabling them to make faster decisions based on real-time data analysis without the delays associated with transferring data to and from centralized servers. The architecture of edge computing generally consists of three distinct layers (Khan et al., <xref ref-type="bibr" rid="B57">2019</xref>), each serving a specific purpose in the data processing workflow, as shown in <xref ref-type="table" rid="T2">Table 2</xref>.</p>
<table-wrap position="float" id="T2">
<label>Table 2</label>
<caption><p>Summary of edge computing architecture layers.</p></caption>
<table frame="box" rules="all">
<thead>
<tr style="background-color:#8f9496;color:#ffffff">
<th valign="top" align="left"><bold>Layer</bold></th>
<th valign="top" align="left"><bold>Description</bold></th>
<th valign="top" align="left"><bold>Primary functions</bold></th>
<th valign="top" align="left"><bold>Advantages</bold></th>
<th valign="top" align="left"><bold>Disadvantages</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">Cloud</td>
<td valign="top" align="left">Handles complex computations and large-scale data storage.</td>
<td valign="top" align="left">Data analysis, machine learning, and long-term data retention.</td>
<td valign="top" align="left">Manages vast amounts of data and is suitable for resource-intensive tasks.</td>
<td valign="top" align="left">Higher latency due to physical distance and it affects real-time responsiveness.</td>
</tr> <tr>
<td valign="top" align="left">Edge nodes</td>
<td valign="top" align="left">Acts as an intermediary between the cloud and end devices and is located closer to data sources.</td>
<td valign="top" align="left">Preliminary data analysis and filtering and facilitates rapid data processing.</td>
<td valign="top" align="left">Reduces data transmission volume and optimizes bandwidth usage while enhancing system efficiency.</td>
<td valign="top" align="left">Limited processing capabilities and cannot handle extremely complex computations.</td>
</tr> <tr>
<td valign="top" align="left">Edge devices</td>
<td valign="top" align="left">Located closest to data sources and includes sensors, cameras, and other IoT devices.</td>
<td valign="top" align="left">Data acquisition and real-time processing.</td>
<td valign="top" align="left">Enables real-time data collection and analysis and ensures rapid responsiveness and reliability.</td>
<td valign="top" align="left">Constrained by power, memory, and processing capabilities and has limited functionality.</td>
</tr></tbody>
</table>
</table-wrap>
<p>The first layer, known as the <italic>cloud</italic>, is responsible for handling complex computations and large-scale data storage. This layer can process and store vast amounts of data, making it well-suited for demanding activities such as data analysis, machine learning, and long-term data storage (Sandhu, <xref ref-type="bibr" rid="B106">2022</xref>). However, its physical distance from the data sources often results in higher latency (Memari et al., <xref ref-type="bibr" rid="B85">2022</xref>), which can hinder the performance of applications that require immediate responses.</p>
<p>The second layer, referred to as <italic>edge nodes</italic> or fog nodes, serves as an intermediary between the cloud and the end devices. Positioned closer to the data sources, edge nodes provide moderate processing capabilities that enable quicker data handling in latency-sensitive applications (Pelle et al., <xref ref-type="bibr" rid="B94">2021</xref>). By performing preliminary data analysis and filtering at this layer, edge nodes can reduce the volume of data that needs to be sent to the cloud, thus optimizing bandwidth usage and enhancing overall system efficiency. This layer plays a crucial role in scenarios where timely decision-making is essential, such as in smart manufacturing or real-time monitoring systems (Nain et al., <xref ref-type="bibr" rid="B87">2022</xref>).</p>
<p>The third layer comprises <italic>edge devices</italic>, which are located closest to the data source. These devices include sensors, cameras, and other IoT devices that possess limited computing power and primarily focus on data acquisition and real-time processing tasks (Shi et al., <xref ref-type="bibr" rid="B110">2016</xref>). Edge devices are typically constrained by factors including energy, memory, and computational capacities, making efficient data handling critical for ensuring responsiveness and reliability. Despite these limitations, edge devices are vital for collecting real-time data, enabling immediate analysis and actions that are essential in various applications, including autonomous vehicles, healthcare monitoring, and smart city infrastructure.</p>
<p>Together, these three layers create a cohesive edge computing architecture that enhances data processing efficiency, reduces latency, and supports a wide range of applications. By distributing computing resources across these layers, edge computing not only addresses the limitations of traditional cloud computing but also empowers organizations to leverage real-time data for improved decision-making and operational effectiveness. This capability not only improves the user experience but also optimizes resource utilization by minimizing the need for extensive data transfers and reducing the load on network infrastructure (Yu et al., <xref ref-type="bibr" rid="B139">2018</xref>). Consequently, edge computing enhances performance and provides a robust solution for the evolving needs of various data analysis tasks, facilitating more secure and efficient data management practices.</p>
</sec>
<sec>
<title>2.2 The evolution and advancements of language models</title>
<p>LLMs exemplified by notable architectures such as BERT (Devlin et al., <xref ref-type="bibr" rid="B28">2019</xref>) and GPT-4 (Brown et al., <xref ref-type="bibr" rid="B14">2020</xref>), have profoundly influenced the field of NLP by enhancing the capability of machines to comprehend and generate human-like text. A fundamental underpinning of these models is the Transformer architecture, which employs self-attention mechanisms to effectively capture dependencies and contextual relationships within textual data (Vaswani et al., <xref ref-type="bibr" rid="B121">2017</xref>). As shown in <xref ref-type="fig" rid="F1">Figure 1</xref>, Transformer architecture is characterized by its layered structure, comprised of multiple encoders and decoders, each integrating self-attention and feedforward neural networks (Raffel et al., <xref ref-type="bibr" rid="B99">2020</xref>).</p>
<fig id="F1" position="float">
<label>Figure 1</label>
<caption><p>Transformer architecture used in LLMs.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fcomp-07-1538277-g0001.tif"/>
</fig>
<p>Prior to the advent of the Transformer, models like Recurrent Neural Networks (RNNs), Long Short-Term Memory networks (LSTMs), and Gated Recurrent Units (GRUs) have been predominant in sequence modeling tasks. RNNs were specifically designed to process sequential input by retaining a hidden state that reflects information from previous time steps (Gu et al., <xref ref-type="bibr" rid="B38">2021</xref>). However, they encountered significant limitations in capturing long-range dependencies due to the vanishing gradient problem, which made it difficult for the model to learn relationships between distant words (Ribeiro et al., <xref ref-type="bibr" rid="B102">2020</xref>). To address this, LSTMs were introduced, incorporating gating mechanisms that regulate the passage of information, allowing them to preserve relevant context over prolonged sequences (Gu et al., <xref ref-type="bibr" rid="B37">2020</xref>). While LSTMs improved upon RNNs by mitigating issues related to long-term dependencies, they still suffered from inefficiencies in computational speed and the inability to fully leverage parallel processing due to their inherently sequential nature. GRUs further simplified the architecture of LSTMs by combining the forget and input gates into a singular update gate, thus reducing the complexity and enhancing performance on various sequence tasks (Savadi Hosseini and Ghaderi, <xref ref-type="bibr" rid="B108">2020</xref>). Nonetheless, RNNs, LSTMs, and GRUs were still constrained by their sequential processing limitations, which hindered their scalability and posed challenges in handling large datasets efficiently.</p>
<p>In contrast, Transformer architecture fundamentally revolutionizes the approach to modeling sequence data by enabling the self-attention mechanism (Raffel et al., <xref ref-type="bibr" rid="B99">2020</xref>), which allows the model to evaluate the significance of each word relative to all other words in the input simultaneously, irrespective of their positional distance. This capability not only facilitates the capture of complex relationships and context with remarkable accuracy but also significantly enhances computational efficiency through parallelization (Zeng et al., <xref ref-type="bibr" rid="B141">2022</xref>). By processing entire sequences at once, Transformers can dramatically reduce training times and handle larger datasets more effectively, thereby overcoming the limitations imposed by earlier models.</p>
<p>The self-attention mechanism represents a pivotal innovation, empowering Transformers to dynamically assess and weigh the relevance of different words, thus ensuring that the contextual meaning is preserved across various tasks (Yang et al., <xref ref-type="bibr" rid="B134">2023</xref>). As a result, LLMs built on the Transformer architecture have proven to be highly effective in diverse NLP applications, including translation (Brown et al., <xref ref-type="bibr" rid="B14">2020</xref>), summarization, and question answering (Yao and Wan, <xref ref-type="bibr" rid="B135">2020</xref>). However, despite these advantages, it is important to note that LLMs are inherently computationally intensive, often requiring substantial memory and processing power that can be a barrier to implementation in resource-constrained environments (Han et al., <xref ref-type="bibr" rid="B40">2015</xref>).</p>
</sec>
<sec>
<title>2.3 Intersection of LLMs and edge computing</title>
<p>The convergence of edge computing and LLMs creates new possibilities for intelligent, real-time data processing right at the source of the data. This development is especially significant for applications that prioritize privacy and time sensitivity, such as those found in healthcare and industrial automation (Sun and Ansari, <xref ref-type="bibr" rid="B116">2019</xref>; Li D. et al., <xref ref-type="bibr" rid="B67">2023</xref>). By deploying LLMs on edge devices, applications can leverage advanced NLP capabilities to conduct localized data analysis. This not only minimizes reliance on cloud infrastructure but also significantly enhances data privacy and operational efficiency (Barua et al., <xref ref-type="bibr" rid="B11">2020</xref>). The ability to process data closer to where it is generated allows for quicker decision-making and reduces latency, which is critical in scenarios where timely responses are essential. Furthermore, this localized approach ensures that sensitive information remains on-site, thereby mitigating the risks associated with data transmission (Cao et al., <xref ref-type="bibr" rid="B16">2020</xref>).</p>
<p>The deployment of LLMs on edge devices presents considerable challenges primarily due to the substantial computational and memory requirements inherent to these models (Yang et al., <xref ref-type="bibr" rid="B133">2024</xref>). As the demand for intelligent applications grows, particularly in environments constrained by hardware limitations, addressing these challenges becomes crucial (Kong et al., <xref ref-type="bibr" rid="B63">2022</xref>). To make LLMs practical for edge environments, researchers are actively exploring a variety of optimization techniques aimed at reducing both the model size and the computational demands, all while striving to maintain acceptable performance levels.</p>
<p>One promising approach is model compression, which encompasses several methods such as pruning, quantization, and knowledge distillation. Pruning involves systematically removing less important parameters from the model, resulting in a reduced computational footprint that allows for faster processing on edge devices (Yeom et al., <xref ref-type="bibr" rid="B137">2021</xref>). By streamlining the model in this way, developers can enhance its efficiency without significantly compromising its capabilities. Additionally, quantization techniques are crucial as they reduce the number of parameters in the model, which in turn decreases memory usage and speeds up inference times (Kim et al., <xref ref-type="bibr" rid="B59">2023</xref>). Furthermore, knowledge distillation&#x02014;a technique that conveys insights from a larger, more intricate model commonly called the &#x0201C;teacher&#x0201D; to a smaller, more streamlined model referred to as the &#x0201C;student&#x0201D;&#x02014;provides an alternative method for facilitating efficient deployment on edge devices. Through this method, the student learns to perform tasks by mimicking the teacher, allowing it to achieve competitive performance with a fraction of the computational resources (Gou et al., <xref ref-type="bibr" rid="B36">2021</xref>). As shown in <xref ref-type="table" rid="T3">Table 3</xref>, these optimization techniques represent significant advancements in the field, making it possible to leverage the sophisticated capabilities of LLMs in edge computing environments, thereby broadening their applicability across various industries and use cases.</p>
<table-wrap position="float" id="T3">
<label>Table 3</label>
<caption><p>Comparison of model compression techniques.</p></caption>
<table frame="box" rules="all">
<thead>
<tr style="background-color:#8f9496;color:#ffffff">
<th valign="top" align="left"><bold>Technique</bold></th>
<th valign="top" align="left"><bold>Description</bold></th>
<th valign="top" align="left"><bold>Benefits</bold></th>
<th valign="top" align="left"><bold>Impact on performance</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">Pruning (Yeom et al., <xref ref-type="bibr" rid="B137">2021</xref>)</td>
<td valign="top" align="left">Involves systematically removing less important parameters from the model.</td>
<td valign="top" align="left">Reduces computational footprint, enhances efficiency.</td>
<td valign="top" align="left">Allows for faster processing on edge devices without significant performance loss.</td>
</tr> <tr>
<td valign="top" align="left">Quantization (Chen et al., <xref ref-type="bibr" rid="B18">2021a</xref>)</td>
<td valign="top" align="left">Reduces the number of model parameters.</td>
<td valign="top" align="left">Lowers memory usage, accelerates inference times.</td>
<td valign="top" align="left">Improves processing speed and reduces memory requirements.</td>
</tr> <tr>
<td valign="top" align="left">Knowledge distillation (Kim, <xref ref-type="bibr" rid="B58">2023</xref>)</td>
<td valign="top" align="left">Transfers knowledge from a larger &#x0201C;teacher&#x0201D; model to a smaller &#x0201C;student&#x0201D; model.</td>
<td valign="top" align="left">Enables efficient deployment of smaller models with competitive performance.</td>
<td valign="top" align="left">Allows the student model to achieve high accuracy with significantly lower computational resources.</td>
</tr></tbody>
</table>
</table-wrap>
<p>Federated learning represents a promising approach for enhancing edge applications, particularly as it enables distributed training across multiple devices without the need to transmit raw data to a central server (Kairouz et al., <xref ref-type="bibr" rid="B53">2021</xref>; Liu G. et al., <xref ref-type="bibr" rid="B75">2021</xref>). This decentralized method is particularly beneficial for applications that prioritize data privacy, allowing sensitive information to remain on the local device while still contributing to the improvement of machine learning models. By facilitating cooperative learning without compromising user privacy, federated learning seeks to harness the collective intelligence of multiple devices, ultimately resulting in more robust and accurate models.</p>
<p>In summary, the integration of LLMs with edge computing presents substantial opportunities for real-time, privacy-sensitive data analysis across a wide range of applications, from healthcare to smart cities and industrial automation. This convergence allows for the processing of data closer to its source, which not only enhances response times but also mitigates the risks associated with data transmission to centralized servers. However, to fully realize this vision, it is imperative to address several critical challenges that have emerged in the current literature. One significant challenge is the computational constraints inherent in edge devices, which often have limited processing power and memory compared to traditional cloud-based systems. This limitation can hinder the deployment of complex LLMs, necessitating the development of more lightweight models or innovative techniques that can efficiently leverage available resources. Additionally, ensuring robust data privacy is paramount, as edge computing environments often handle sensitive information that must be protected from unauthorized access. This requires the implementation of advanced encryption methods and privacy-preserving techniques to safeguard user data while still allowing for effective analysis. Moreover, optimizing LLMs for efficient performance on edge devices is crucial. This involves not only reducing the model size and complexity but also adapting algorithms to ensure they can operate effectively within the constraints of edge infrastructures.</p>
<p>By tackling these interconnected issues, we can unlock the transformative potential of LLMs in edge environments, paving the way for the development of smarter, more responsive systems. Such advancements will cater to the evolving needs of users, providing them with timely insights and services while maintaining a strong emphasis on privacy and security.</p>
</sec>
</sec>
<sec id="s3">
<title>3 Applications of LLMs for data analysis in edge computing</title>
<p>LLMs hold considerable potential for data analysis in edge computing due to their capability to process and analyze data locally, providing real-time insights and enhanced data privacy. In this section, we discuss several critical application areas: healthcare and wearable devices, IoT and smart cities, industrial automation, and consumer applications, as shown in <xref ref-type="fig" rid="F2">Figure 2</xref>.</p>
<fig id="F2" position="float">
<label>Figure 2</label>
<caption><p>Illustration of applications of LLMs for data analysis in edge computing.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fcomp-07-1538277-g0002.tif"/>
</fig>
<sec>
<title>3.1 Healthcare and wearable devices</title>
<p>In the realm of healthcare, edge computing significantly enhances the capacity for real-time analysis of patient data collected through wearable devices, thereby improving responsiveness and minimizing latency. The deployment of LLMs on edge devices enables the analysis of substantial volumes of textual data derived from medical records and wearable health monitors. This capability is instrumental in facilitating early anomaly detection and comprehensive trend analysis, which are critical for timely medical interventions (Esteva et al., <xref ref-type="bibr" rid="B32">2019</xref>; Dias and Paulo Silva Cunha, <xref ref-type="bibr" rid="B29">2018</xref>). By processing data locally on edge devices, sensitive patient information remains secure, effectively reducing the necessity for cloud transmission and, consequently, preserving data privacy (Zhou B. et al., <xref ref-type="bibr" rid="B147">2022</xref>).</p>
<p>Wearable devices that incorporate LLMs are adept at continuously monitoring vital signs, thus providing ongoing health assessments (Kim et al., <xref ref-type="bibr" rid="B61">2024</xref>). For instance, a wearable device may issue alerts to healthcare providers regarding irregularities in a patient&#x00027;s vital signs, as determined through the analytical capabilities of LLMs. This proactive approach to medical intervention not only enhances patient care but also exemplifies the potential for continuous health monitoring without reliance on centralized data processing (Babu et al., <xref ref-type="bibr" rid="B9">2024</xref>).</p>
<p>Furthermore, the integration of edge computing within the domain of wearable health technology fosters significant advancements in academic research. Researchers can leverage the real-time data collected from edge devices to conduct extensive analyses of health trends across diverse populations. This data-driven approach enables the identification of potential health risks and the formulation of more effective preventive measures and treatment protocols (Garc&#x000ED;a-M&#x000E9;ndez and de Arriba-P&#x000E9;rez, <xref ref-type="bibr" rid="B35">2024</xref>). Such insights are invaluable for the advancement of personalized medicine, as they allow for tailored healthcare solutions that address the unique needs of individual patients. The implications of edge computing extend beyond individual patient care to influence broader public health strategies (Jo et al., <xref ref-type="bibr" rid="B52">2023</xref>). By analyzing aggregated data from wearable devices, researchers can derive insights that inform health policy decisions and resource allocation, ultimately contributing to improved health outcomes at the population level. This paradigm shift toward decentralized data processing not only enhances the efficiency and security of healthcare delivery but also paves the way for innovative research methodologies that capitalize on the wealth of data generated by wearable technologies (Shiranthika et al., <xref ref-type="bibr" rid="B111">2023</xref>).</p>
<p>In conclusion, the application of edge computing in the healthcare wearable sector not only enhances the efficiency and security of data processing but also propels academic research forward, offering new perspectives and tools for advancing health management. As technology continues to evolve, edge computing is poised to play an increasingly pivotal role in the future of healthcare, facilitating more precise and effective patient care while supporting the ongoing quest for knowledge in medical research.</p>
</sec>
<sec>
<title>3.2 IoT and smart cities</title>
<p>In the context of smart city applications, edge computing plays a pivotal role in managing the vast amounts of data generated by IoT sensors across various domains, including traffic management, environmental monitoring, autonomous driving, and public safety. The deployment of LLMs on edge devices enables the analysis of diverse data streams, such as sensor reports, social media posts, and other real-time data inputs, thereby providing actionable insights that are crucial for city planners and emergency responders (Jan et al., <xref ref-type="bibr" rid="B48">2021</xref>; Li X. et al., <xref ref-type="bibr" rid="B71">2022</xref>).</p>
<p>For instance, real-time traffic data analysis conducted by LLMs on edge devices can significantly optimize traffic light timing (de Zarz&#x000E0; et al., <xref ref-type="bibr" rid="B25">2023</xref>), which in turn reduces congestion and enhances fuel efficiency. By processing data locally, these systems can react swiftly to changing traffic conditions, allowing for dynamic adjustments that promote smoother traffic flow and minimize delays. This capability not only improves the overall transportation experience for citizens but also contributes to the reduction of greenhouse gas emissions associated with idling vehicles.</p>
<p>In the realm of environmental monitoring, LLMs can locally process air quality data collected from various sensors deployed throughout the city. When pollution levels exceed predefined safe thresholds, these edge-based systems can issue immediate alerts to both authorities and the public, facilitating timely interventions to protect public health (Alahi et al., <xref ref-type="bibr" rid="B5">2023</xref>). Such proactive measures are essential in maintaining compliance with health standards and mitigating the adverse effects of pollution on urban populations.</p>
<p>The advantages of edge computing in smart city applications extend beyond immediate responsiveness; they also encompass a reduction in dependency on centralized cloud infrastructure. By enabling data processing at the edge, cities can ensure that data-driven responses can be initiated without the delays often associated with cloud-based systems (Khan et al., <xref ref-type="bibr" rid="B56">2020</xref>; Lv et al., <xref ref-type="bibr" rid="B82">2021</xref>). This decentralized approach enhances the resilience of urban infrastructure, particularly in emergency situations where rapid decision-making is critical. Furthermore, the integration of edge computing within smart city frameworks fosters significant opportunities for academic research. Researchers can utilize the rich datasets generated by IoT sensors and edge devices to conduct comprehensive studies on urban dynamics, environmental impacts, and public safety trends. This data-driven research paradigm enables the identification of patterns and correlations that can inform policy decisions and urban planning strategies, ultimately leading to more sustainable and livable cities (Liu Q. et al., <xref ref-type="bibr" rid="B77">2021</xref>).</p>
</sec>
<sec>
<title>3.3 Industrial automation and predictive maintenance</title>
<p>In industrial settings, predictive maintenance is of paramount importance for ensuring equipment reliability and optimizing operational efficiency. The integration of LLMs with edge computing technologies facilitates on-site analysis of a variety of data sources, including text-based maintenance logs, sensor data, and equipment performance reports. This capability enables organizations to predict potential failures before they manifest, thereby enhancing the overall reliability of industrial operations (Zonta et al., <xref ref-type="bibr" rid="B154">2020</xref>).</p>
<p>By processing data at the edge, companies can substantially reduce latency associated with data transmission to centralized cloud systems. This reduction in latency is critical, as it allows for the timely initiation of preventive measures, which in turn minimizes costly downtimes and enhances productivity (Pang et al., <xref ref-type="bibr" rid="B93">2021</xref>). For instance, an LLM deployed on an edge device can continuously analyze vibration patterns and temperature fluctuations from manufacturing equipment. When anomalies are detected, the system can generate immediate alerts for preventive maintenance, enabling operators to address issues proactively rather than reactively. This localized processing capability not only enhances operational safety but also substantially reduces the likelihood of unexpected shutdowns, which can have severe financial implications for manufacturing operations.</p>
<p>Moreover, the advantages of edge computing in predictive maintenance extend to the realm of data security and privacy (Qiu et al., <xref ref-type="bibr" rid="B97">2020</xref>; Zhang J. et al., <xref ref-type="bibr" rid="B144">2018</xref>). By conducting analyses locally, sensitive operational data is less exposed to external threats associated with cloud transmission, thereby enhancing the overall security posture of industrial facilities. This is particularly important in industries where proprietary processes and trade secrets are involved, as edge computing mitigates the risks of data breaches and intellectual property theft (Zhou et al., <xref ref-type="bibr" rid="B152">2024b</xref>).</p>
<p>The integration of edge computing and LLMs also opens new avenues for academic research within the field of industrial automation. Researchers can leverage the vast amounts of data generated by industrial equipment to develop advanced predictive models that improve maintenance strategies. By utilizing machine learning methods and data analytics, researchers can identify patterns and correlations that guide best practices in predictive maintenance, ultimately enhancing the effectiveness and efficiency of operational frameworks.</p>
</sec>
<sec>
<title>3.4 Consumer applications</title>
<p>In consumer applications, the integration of LLMs embedded in edge devices has become increasingly prevalent, particularly in smart home assistants, smartphones, and personal wearables. These devices leverage LLMs to process user commands locally, significantly enhancing response times and preserving user privacy by minimizing reliance on cloud processing (Syu et al., <xref ref-type="bibr" rid="B118">2023</xref>). This local processing capability not only facilitates quicker interactions but also mitigates concerns related to data security and privacy, as sensitive information remains on the device rather than being transmitted to remote servers.</p>
<p>For instance, smart home systems equipped with LLMs can analyze and predict energy usage patterns, enabling users to optimize their energy consumption based on real-time insights (Yu et al., <xref ref-type="bibr" rid="B138">2021</xref>; Iqbal et al., <xref ref-type="bibr" rid="B46">2023</xref>). By monitoring factors such as appliance usage and user behavior, these systems can provide tailored recommendations that promote energy efficiency and cost savings. This capability is particularly relevant in an era where energy conservation is of paramount importance, as it empowers consumers to make informed decisions regarding their energy consumption habits.</p>
<p>Moreover, edge-deployed LLMs enhance security in smart home environments by analyzing data from various sensors to detect unusual patterns. For example, they can monitor for unexpected motion or unusual access times, alerting users immediately to potential security threats (Siriwardhana et al., <xref ref-type="bibr" rid="B113">2021</xref>). This proactive approach to home security not only provides peace of mind for users but also demonstrates the potential for LLMs to contribute to safer living environments through real-time monitoring and alerting mechanisms.</p>
<p>The advantages of edge computing in consumer applications extend beyond improved response times and enhanced privacy. The localized processing of data allows for greater resilience in the face of connectivity issues, as devices can continue to function effectively even when offline or in low-bandwidth situations (Li J. et al., <xref ref-type="bibr" rid="B68">2023</xref>). This characteristic is particularly valuable in resource-constrained environments, where consistent internet access may not be guaranteed.</p>
<p>In summary, the integration of LLMs with edge computing in consumer applications offers substantial benefits, including reduced latency, enhanced privacy, and decreased dependency on cloud infrastructure. These applications underscore the transformative potential of edge-compatible LLMs in providing real-time, privacy-conscious data insights, thereby enriching user experiences across various domains. As research in this field progresses, it will undoubtedly yield innovative solutions that further enhance the capabilities and societal acceptance of AI-driven consumer technologies.</p>
</sec>
</sec>
<sec id="s4">
<title>4 Challenges of implementing LLMs in edge computing</title>
<p>Deploying LLMs in edge computing environments presents significant challenges due to the resource-constrained nature of edge devices (Boumendil et al., <xref ref-type="bibr" rid="B13">2025</xref>). These challenges include resource constraints, energy efficiency, data privacy and security, as well as latency and real-time requirements, as shown in <xref ref-type="table" rid="T4">Table 4</xref>. Resource constraints arise from the limited computational and memory capabilities of edge devices compared to robust cloud servers, making it difficult to effectively utilize large models. Additionally, the energy demands of LLMs pose challenges in battery-powered edge devices, necessitating optimizations to minimize power consumption. Data privacy concerns are amplified in scenarios involving sensitive information, where centralized processing introduces risks that must be mitigated. Finally, the inherent computational intensity of traditional LLMs often results in high latency, which is incompatible with the near-instantaneous processing required in critical applications such as autonomous driving and real-time diagnostics. To address these challenges, solutions such as model compression techniques, secure local processing methods, and latency reduction strategies are essential for ensuring the successful integration of LLMs into resource-constrained edge environments (Cheng et al., <xref ref-type="bibr" rid="B20">2024</xref>).</p>
<table-wrap position="float" id="T4">
<label>Table 4</label>
<caption><p>Challenges of deploying LLMs in edge computing environments.</p></caption>
<table frame="box" rules="all">
<thead>
<tr style="background-color:#8f9496;color:#ffffff">
<th valign="top" align="left"><bold>Challenge</bold></th>
<th valign="top" align="left"><bold>Description</bold></th>
<th valign="top" align="left"><bold>Implications</bold></th>
<th valign="top" align="left"><bold>Potential solutions</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">Resource constraints</td>
<td valign="top" align="left">Edge devices have limited computational and memory resources compared to cloud servers, making it difficult to deploy large models like GPT-3, which require high-performance hardware (e.g., GPUs, TPUs) for effective operation.</td>
<td valign="top" align="left">The inability to store and execute LLMs can hinder their practical application in real-world scenarios, leading to performance degradation.</td>
<td valign="top" align="left">Develop model compression techniques (e.g., quantization, pruning) and knowledge distillation to create smaller, efficient models.</td>
</tr> <tr>
<td valign="top" align="left">Energy efficiency</td>
<td valign="top" align="left">Edge devices are typically battery-powered and require LLMs to minimize power consumption while maintaining computational efficacy. Standard LLM architectures are energy-intensive during training and inference.</td>
<td valign="top" align="left">High energy usage can limit the deployment of LLMs in energy-constrained environments, affecting their viability in applications requiring sophisticated NLP.</td>
<td valign="top" align="left">Implement model compression, pruning, quantization, and knowledge distillation to reduce energy requirements while retaining performance.</td>
</tr> <tr>
<td valign="top" align="left">Data privacy and security</td>
<td valign="top" align="left">Deploying LLMs at the edge involves handling sensitive information, raising concerns about data privacy due to potential risks associated with centralized processing and data transmission.</td>
<td valign="top" align="left">Centralized architectures can compromise user confidentiality and trust, necessitating robust privacy protections to comply with ethical and legal standards.</td>
<td valign="top" align="left">Employ secure local processing techniques, encryption methods, differential privacy, and federated learning to enhance data security.</td>
</tr> <tr>
<td valign="top" align="left">Latency and real-time requirements</td>
<td valign="top" align="left">Low latency is crucial for edge applications like autonomous driving and real-time diagnostics, where delays can lead to serious consequences. Traditional LLMs often have high latency due to their computational demands.</td>
<td valign="top" align="left">Elevated latency levels can jeopardize the effectiveness of applications that require near-instantaneous processing, potentially leading to critical failures.</td>
<td valign="top" align="left">Optimize LLMs through model partitioning and early exit strategies to facilitate parallel processing and timely predictions.</td>
</tr></tbody>
</table>
</table-wrap>
<sec>
<title>4.1 Resource constraints</title>
<p>One of the fundamental challenges associated with the deployment of LLMs on edge devices lies in the constraints imposed by limited computational and memory resources (Shi et al., <xref ref-type="bibr" rid="B110">2016</xref>; Sun and Ansari, <xref ref-type="bibr" rid="B116">2019</xref>). In contrast to robust cloud servers that possess the capacity to manage large model sizes and perform complex calculations, edge devices are typically characterized by their insufficient memory and processing power. This inadequacy limits their ability to store and execute LLMs directly, thereby presenting a significant barrier to the practical application of such models in real-world scenarios.</p>
<p>For example, models like GPT-3, which contain billions of parameters, necessitate high-performance hardware such as GPUs or TPUs to function effectively (Brown et al., <xref ref-type="bibr" rid="B14">2020</xref>). These advanced processing units are instrumental in handling the extensive computational demands associated with both the training and inference stages of LLMs. However, such high-performance hardware is frequently absent in edge environments, where devices are designed to be small, energy-efficient, and cost-effective. This disparity between the resource requirements of state-of-the-art LLMs and the capabilities of edge devices leads to challenges in executing these models without significant degradation in performance.</p>
<p>The limitations in computational power and memory capacity are further compounded by the diverse nature of edge environments, which may include mobile devices, IoT sensors, and embedded systems. Each of these platforms comes with its own set of constraints, making it essential for researchers to explore alternative strategies for model deployment (Wang S. et al., <xref ref-type="bibr" rid="B124">2019</xref>). One potential avenue is to develop model compression techniques, such as quantization and pruning (Chen et al., <xref ref-type="bibr" rid="B18">2021a</xref>; Liang et al., <xref ref-type="bibr" rid="B72">2021</xref>), which aim to reduce the size of models while retaining their essential functionalities. Additionally, approaches such as knowledge distillation can be employed to create smaller (Matsubara et al., <xref ref-type="bibr" rid="B83">2020</xref>; Ji et al., <xref ref-type="bibr" rid="B51">2024</xref>), more efficient surrogate models that approximate the performance of larger LLMs while being suitable for deployment on edge devices.</p>
<p>Overall, addressing these computational and memory limitations is paramount for the successful integration of LLMs into edge devices, thereby enabling the delivery of advanced language processing capabilities in a variety of resource-constrained environments.</p>
</sec>
<sec>
<title>4.2 Energy efficiency</title>
<p>Energy efficiency poses a substantial challenge in the deployment of LLMs on edge devices, which are inherently constrained by their battery-powered design and limited availability of energy resources. The increasing proliferation of edge devices, encompassing a diverse range of applications such as mobile phones, smart home systems, and IoT devices, necessitates the optimization of LLMs to ensure minimal power consumption while maintaining the required level of computational efficacy (Shuvo et al., <xref ref-type="bibr" rid="B112">2022</xref>). This requirement is particularly critical given the growing demand for sophisticated NLP capabilities in scenarios where low power use is paramount.</p>
<p>Standard architectures for LLMs, including well-known models such as BERT and GPT, are particularly notable for their substantial energy consumption during both the training and inference phases. During these phases, particularly when dealing with extensive datasets and complex tasks, these models can require significant computational power and energy resources to operate effectively. This characteristic starkly contrasts with the low-power requirements typically associated with edge environments, where devices are designed to function under strict energy constraints without sacrificing performance (Devlin et al., <xref ref-type="bibr" rid="B28">2019</xref>). The challenge here is twofold: not only must the models operate efficiently, but they must also be capable of delivering acceptable performance levels despite the limitations of edge devices. Larger models, characterized by billions of parameters, tend to demand exponentially increasing amounts of computational resources, which translates to higher energy usage (Strubell et al., <xref ref-type="bibr" rid="B114">2020</xref>; Wang H. et al., <xref ref-type="bibr" rid="B123">2019</xref>). This relationship underscores the urgency for developing model adaptation techniques aimed at achieving performance equivalence while significantly reducing energy expenditure. Such techniques are not merely advantageous but essential for effectively integrating the capabilities of LLMs within energy-constrained contexts, allowing for the deployment of advanced NLP applications in a variety of settings without detrimental impacts on device operation.</p>
<p>To address these energy efficiency challenges, several strategies can be employed. One prominent approach is model compression, which encompasses techniques such as pruning, quantization, and knowledge distillation. Pruning involves removing redundant parameters from the model, thereby reducing its size and, consequently, its computational demands (Kim, <xref ref-type="bibr" rid="B58">2023</xref>). Quantization refers to the process of approximating the model weights using lower precision formats, which can further decrease memory usage and accelerate inference speed (Zhang et al., <xref ref-type="bibr" rid="B142">2025</xref>). In contrast, knowledge distillation involves training a smaller, more efficient model (referred to as the student) to imitate the behavior of a larger, more complex model (the teacher) (Wang et al., <xref ref-type="bibr" rid="B125">2024</xref>). This process yields a streamlined version that maintains a significant portion of the original model&#x00027;s performance while utilizing fewer resources.</p>
</sec>
<sec>
<title>4.3 Data privacy and security</title>
<p>Data privacy emerges as a critical concern in the deployment of LLMs at the edge, particularly in applications that involve handling sensitive information such as healthcare records, financial transactions, and personal data (Kairouz et al., <xref ref-type="bibr" rid="B53">2021</xref>). Traditional centralized models typically necessitate the transfer of data to cloud servers for processing, which introduces substantial privacy risks. These risks arise from several factors, including potential data interception during transmission, vulnerabilities associated with centralized storage systems, and the possibility of unauthorized access to sensitive information (Zhou et al., <xref ref-type="bibr" rid="B149">2022b</xref>). Consequently, such centralized architectures may compromise user confidentiality and trust, raising ethical and regulatory concerns.</p>
<p>In contrast, edge computing offers a paradigm that can significantly mitigate these privacy-related challenges by ensuring that data remains closer to its source. By processing data locally on edge devices, the exposure of sensitive information is inherently reduced, thereby reducing the likelihood of data breaches and unauthorized access. This localized data processing approach enables organizations to leverage powerful LLMs while adhering to privacy regulations and maintaining user trust (Ali et al., <xref ref-type="bibr" rid="B6">2021</xref>). However, the transition to edge computing does not eliminate the need for robust privacy protections; rather, it necessitates the implementation of secure local processing techniques. Effective encryption methods must be employed to safeguard data both at rest and in transit, ensuring that even if data does reside on edge devices, it remains protected from potential adversaries (Alwarafy et al., <xref ref-type="bibr" rid="B7">2021</xref>). Additionally, incorporating differential privacy techniques can further enhance data security by adding noise to the data during processing, making it increasingly challenging for attackers to infer sensitive information without significantly impacting the model&#x00027;s performance (Du et al., <xref ref-type="bibr" rid="B31">2020</xref>). Edge environments often consist of a diverse array of devices, each with varying levels of security capabilities, which complicates the development of standardized privacy protocols. Therefore, it is essential to adopt a multifaceted approach that considers the unique attributes and constraints of different edge devices. This may include utilizing lightweight privacy-preserving algorithms that are compatible with the limited computational resources typical of edge devices while ensuring compliance with relevant legal and ethical standards.</p>
<p>The integration of federated learning into edge computing architectures presents a promising solution to enhance data privacy (Nguyen et al., <xref ref-type="bibr" rid="B90">2021b</xref>). Federated learning facilitates the training of models across various edge devices without the necessity of sharing raw data. Instead, only model updates&#x02013;devoid of sensitive information&#x02014;are sent to a central server for aggregation. This approach preserves user privacy while allowing for the ongoing enhancement of LLMs (Abreha et al., <xref ref-type="bibr" rid="B1">2022</xref>; Xia et al., <xref ref-type="bibr" rid="B129">2021</xref>). While the deployment of LLMs at the edge offers significant advantages in terms of data privacy, it also presents unique challenges related to secure local processing. A comprehensive strategy that incorporates encryption, differential privacy, lightweight algorithms, and federated learning is essential for mitigating privacy risks while ensuring the effective use of LLMs in sensitive applications.</p>
</sec>
<sec>
<title>4.4 Latency and real-time requirements</title>
<p>Low latency is a critical requirement for edge applications across various sectors (Ke et al., <xref ref-type="bibr" rid="B55">2023</xref>), including autonomous driving, augmented reality (Zhang et al., <xref ref-type="bibr" rid="B143">2020</xref>), and real-time diagnostics, where even minor delays in response times can lead to significant consequences, potentially jeopardizing safety and operational efficiency (Kang et al., <xref ref-type="bibr" rid="B54">2017</xref>). In these contexts, traditional LLMs pose challenges due to their inherent computational intensity, which often results in elevated latency levels that are incompatible with the stringent demands of edge applications that necessitate near-instantaneous processing capabilities.</p>
<p>For instance, in the realm of autonomous driving, the ability to process sensory data and make real-time decisions is paramount (Roszyk et al., <xref ref-type="bibr" rid="B103">2022</xref>); any delay can result in critical failures, such as the inability to react promptly to dynamic road conditions or obstacles (Lin et al., <xref ref-type="bibr" rid="B74">2018</xref>). Similarly, augmented reality applications rely on real-time data processing to provide users with seamless and interactive experiences, where lag can disrupt the immersive quality of the application (Chen et al., <xref ref-type="bibr" rid="B17">2018</xref>; Zhang W. et al., <xref ref-type="bibr" rid="B146">2018</xref>). In the field of real-time diagnostics, the timely analysis of medical data can be a matter of life and death (K&#x000F6;hl and Hermanns, <xref ref-type="bibr" rid="B62">2023</xref>), underscoring the importance of minimizing latency.</p>
<p>To effectively utilize LLMs for real-time processing of speech or image data, a range of optimizations is essential to minimize processing time while maintaining accuracy. Various techniques are being investigated to tackle latency issues, including model partitioning and early exit strategies. Model partitioning involves distributing the model across multiple devices, enabling parallel processing that significantly reduces latency (Kang et al., <xref ref-type="bibr" rid="B54">2017</xref>). Meanwhile, early exit mechanisms allow the model to generate predictions at intermediate layers once a high level of confidence is reached, thereby decreasing computation time for simpler tasks (Teerapittayanon et al., <xref ref-type="bibr" rid="B119">2016</xref>). These strategies are crucial for meeting the real-time demands of edge applications, allowing LLMs to deliver timely responses without sacrificing performance.</p>
</sec>
</sec>
<sec id="s5">
<title>5 Recent advances for edge-compatible LLMs</title>
<p>Given the computational and storage constraints of edge devices, deploying LLMs necessitates the implementation of various techniques aimed at minimizing model size, energy consumption, and latency. These constraints are critical considerations, as edge devices often operate under limited processing power and memory capacity, which can significantly impact the performance of LLMs. Therefore, it is essential to adopt strategies that not only reduce the resource footprint of these models but also ensure their effectiveness in real-world applications.</p>
<p>This section discusses three major approaches for making LLMs compatible with edge devices. First, model compression techniques, such as pruning, quantization, and knowledge distillation, play a crucial role in reducing the size and complexity of LLMs without substantially sacrificing their performance. These methods allow for the deployment of smaller models that can operate efficiently within the constraints of edge environments. Second, federated and distributed learning frameworks offer innovative solutions for training LLMs across decentralized data sources while preserving data privacy. By enabling collaborative learning without the need to centralize sensitive information, these approaches not only enhance security but also improve the robustness of the models through diverse data exposure. Lastly, optimized architectures and frameworks, particularly those that are edge-friendly, are crucial for ensuring that LLMs can operate effectively in resource-constrained environments. These architectures focus on balancing performance and efficiency, allowing for rapid inference times and reduced energy consumption, which are essential for applications that require real-time processing.</p>
<sec>
<title>5.1 Model compression techniques</title>
<p>Model compression is essential to fit LLMs within the resource-constrained environments of edge devices. Compression methods aim to reduce model size and computational demands while retaining performance. Three popular compression techniques are pruning, quantization, and knowledge distillation (Han et al., <xref ref-type="bibr" rid="B40">2015</xref>; Sanh et al., <xref ref-type="bibr" rid="B107">2019</xref>). Pruning involves removing the less important connections within the model to reduce its size without a significant drop in performance (Han et al., <xref ref-type="bibr" rid="B40">2015</xref>). Quantization reduces the precision of the model weights, decreasing memory requirements and improving computational efficiency (Jacob et al., <xref ref-type="bibr" rid="B47">2018</xref>). Knowledge distillation enables a student model to learn from a teacher model, retaining key features of the original model but in a more compact form (Sanh et al., <xref ref-type="bibr" rid="B107">2019</xref>). These techniques, illustrated in <xref ref-type="fig" rid="F3">Figure 3</xref>, have shown promise in adapting LLMs to edge devices, but further advances are needed to effectively balance model performance with hardware limitations.</p>
<fig id="F3" position="float">
<label>Figure 3</label>
<caption><p>Illustration of model compression techniques for edge deployment: pruning, quantization, and knowledge distillation.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fcomp-07-1538277-g0003.tif"/>
</fig>
<sec>
<title>5.1.1 Pruning</title>
<p>Pruning is a pivotal technique in the field of model compression that focuses on enhancing the efficiency of neural networks by removing redundant connections (Wang et al., <xref ref-type="bibr" rid="B126">2021</xref>). This is accomplished by systematically removing low-weight parameters, leading to a substantial decrease in both memory requirements and computational costs. Pruning techniques can be generally divided into two categories: structured and unstructured methods. Structured pruning entails the removal of entire neurons, channels, or layers, while unstructured pruning focuses on the elimination of individual weights (Gale et al., <xref ref-type="bibr" rid="B33">2019</xref>).</p>
<p>Recent advancements in pruning techniques have highlighted two primary approaches: structured and unstructured pruning (Vahidian et al., <xref ref-type="bibr" rid="B120">2021</xref>). Structured pruning focuses on the removal of entire neurons, channels, or layers, thereby maintaining the overall architecture of the network while significantly enhancing its efficiency (Anwar et al., <xref ref-type="bibr" rid="B8">2017</xref>). On the other hand, unstructured pruning targets individual weights, allowing for a finer level of granularity in the pruning process. While unstructured pruning often achieves higher compression rates, it can lead to irregular sparsity patterns that may not be as easily optimized for hardware acceleration (Vahidian et al., <xref ref-type="bibr" rid="B120">2021</xref>). In contrast, structured pruning tends to produce more regular sparsity patterns, making it more compatible with various hardware architectures (Xu et al., <xref ref-type="bibr" rid="B132">2025</xref>), such as GPUs and TPUs, which can exploit these structures for improved performance. The integration of pruning with neural architecture search (NAS) has gained traction, enabling the automatic discovery of optimal architectures that are inherently more efficient (Ding et al., <xref ref-type="bibr" rid="B30">2022</xref>). Additionally, the development of dynamic pruning methods, which adaptively prune weights during training based on their importance, has shown promise in maintaining model performance while achieving significant reductions in size (Hu et al., <xref ref-type="bibr" rid="B43">2023</xref>; Liu et al., <xref ref-type="bibr" rid="B80">2018</xref>). Another notable trend is the use of pruning in conjunction with quantization techniques, which further compress models by reducing the precision of the weights (Chowdhury et al., <xref ref-type="bibr" rid="B22">2021</xref>), leading to even greater efficiency.</p></sec>
<sec>
<title>5.1.2 Quantization</title>
<p>Quantization is a pivotal technique in the optimization of neural networks, specifically designed to enhance computational efficiency and reduce memory usage. This process involves decreasing the precision of model weights, allowing them to be represented with fewer bits&#x02014;commonly using 8-bit integers in place of the traditional 32-bit floating-point representations (Jacob et al., <xref ref-type="bibr" rid="B47">2018</xref>). By reducing the bit-width of weights, quantization not only conserves memory but also accelerates computational operations, making quantized models particularly well-suited for deployment on edge devices that often have stringent resource constraints. In recent years, the field of quantization has witnessed significant advancements, especially in the context of LLMs.</p>
<p>One notable trend is the integration of quantization with retraining strategies, where models are fine-tuned after quantization to recover any potential loss in accuracy (Wu et al., <xref ref-type="bibr" rid="B128">2020</xref>). This approach has proven effective in mitigating the impact of reduced precision, ensuring that the quantized models remain robust and reliable for real-world applications. Another area of focus has been the development of adaptive quantization methods, which dynamically adjust the quantization strategy based on the distribution of weight values within the model (Zhou et al., <xref ref-type="bibr" rid="B153">2018</xref>). This technique allows for more efficient representation of weights, optimizing both performance and resource utilization. Additionally, researchers have explored the potential of mixed-precision quantization, where different layers of a neural network are quantized to varying degrees of precision based on their sensitivity to quantization errors (Chen et al., <xref ref-type="bibr" rid="B19">2021b</xref>). This tailored approach can lead to enhanced performance while still achieving significant reductions in model size (Liu X. et al., <xref ref-type="bibr" rid="B78">2021</xref>). The rise of quantization-aware training has also emerged as a key area of research. This technique involves incorporating quantization effects into the training process itself, allowing the model to learn to be robust against the quantization noise it will encounter during inference (Nagel et al., <xref ref-type="bibr" rid="B86">2022</xref>). This proactive strategy has shown promise in preserving accuracy while maximizing the benefits of quantization (Sakr et al., <xref ref-type="bibr" rid="B105">2022</xref>).</p></sec>
<sec>
<title>5.1.3 Knowledge distillation</title>
<p>Knowledge distillation is a powerful technique in the field of deep learning. The primary objective of this process is to train the student to replicate the outputs of the teacher, effectively allowing it to inherit the performance characteristics of the more sophisticated model while maintaining a significantly reduced size (Gou et al., <xref ref-type="bibr" rid="B36">2021</xref>).</p>
<p>One prominent example of knowledge distillation is DistilBERT (Sanh et al., <xref ref-type="bibr" rid="B107">2019</xref>), which is a distilled variant of the well-known BERT model. DistilBERT achieves remarkable efficiency by retaining approximately 97% of BERT&#x00027;s language understanding capabilities while being roughly half the size of the original model. This compression is particularly advantageous for applications requiring deployment on edge devices, where computational resources and memory are often limited (Gou et al., <xref ref-type="bibr" rid="B36">2021</xref>). One notable direction is the exploration of multi-teacher distillation, where the student model is trained to learn from multiple teacher models simultaneously (Liu et al., <xref ref-type="bibr" rid="B79">2020</xref>). This approach can enhance the student&#x00027;s performance by leveraging diverse knowledge representations, thus improving generalization capabilities across various tasks (Yuan et al., <xref ref-type="bibr" rid="B140">2021</xref>). Another significant development is the integration of knowledge distillation with other model compression techniques, such as pruning and quantization (Kim, <xref ref-type="bibr" rid="B58">2023</xref>). By combining these methods, researchers have been able to create even more compact models that not only maintain high levels of accuracy but also achieve lower latency and energy consumption during inference (Bao et al., <xref ref-type="bibr" rid="B10">2019</xref>).</p>
<p>Additionally, there has been a growing interest in the application of knowledge distillation beyond traditional supervised learning settings. For instance, distillation techniques are being adapted for use in semi-supervised (Li et al., <xref ref-type="bibr" rid="B66">2019</xref>; Guo et al., <xref ref-type="bibr" rid="B39">2022</xref>) and unsupervised learning scenarios (Han et al., <xref ref-type="bibr" rid="B41">2022</xref>), allowing student models to benefit from the knowledge encoded in teacher models trained on large, unlabeled datasets. Advancements in the design of loss functions used during the distillation process have also emerged. Researchers are investigating new ways to optimize the training objective, focusing on enhancing the alignment between the teacher&#x00027;s and student&#x00027;s outputs (Liu P. et al., <xref ref-type="bibr" rid="B76">2021</xref>), which can lead to improved performance of the distilled models.</p>
</sec>
</sec>
<sec>
<title>5.2 Federated and distributed learning for edge applications</title>
<p>Federated and distributed learning techniques enable the training and updating of LLMs across multiple edge devices without requiring raw data to be transferred to a central server. These methods are especially beneficial in privacy-sensitive applications, such as healthcare and finance, as they allow data to remain local to the device (Kairouz et al., <xref ref-type="bibr" rid="B53">2021</xref>; Ye et al., <xref ref-type="bibr" rid="B136">2023</xref>).</p>
<sec>
<title>5.2.1 Federated learning</title>
<p>Federated learning is an innovative distributed machine learning paradigm that enables multiple edge devices to collaboratively train a shared model while keeping their data localized. In this approach, each device computes local model updates based on its unique dataset, which are subsequently aggregated to form a global model (Kairouz et al., <xref ref-type="bibr" rid="B53">2021</xref>). This methodology significantly mitigates the need for centralized data storage, thereby enhancing data privacy and security, as sensitive information remains on the individual devices rather than being transmitted to a central server. Federated learning has gained considerable attention, particularly in contexts where data privacy is paramount, such as healthcare, finance, and personal mobile applications. One of the critical challenges associated with federated learning is the heterogeneity of data across devices, which often leads to non-IID (independent and identically distributed) data distributions (McMahan et al., <xref ref-type="bibr" rid="B84">2017</xref>). This variability can complicate the training process, as models trained on diverse data may not converge effectively.</p>
<p>To address these challenges, researchers have been developing various techniques aimed at improving communication efficiency and model synchronization (Zhou et al., <xref ref-type="bibr" rid="B151">2023</xref>). One prominent method is federated averaging, which optimally aggregates the model updates from different devices (Nguyen et al., <xref ref-type="bibr" rid="B89">2021a</xref>). This technique helps to balance the contributions from devices with varying amounts of data, ensuring that the global model reflects the knowledge from all participating devices. Another noteworthy advancement is the concept of adaptive federated learning, which dynamically adjusts the training process based on the characteristics of the participating devices and their data (Wang S. et al., <xref ref-type="bibr" rid="B124">2019</xref>). This approach can include mechanisms to prioritize updates from devices with more representative or higher-quality data, thereby enhancing the overall performance of the federated model.</p>
<p>Recent studies have explored the integration of federated learning with other emerging technologies, such as differential privacy (Wei et al., <xref ref-type="bibr" rid="B127">2020</xref>) and secure multi-party computation (Byrd and Polychroniadou, <xref ref-type="bibr" rid="B15">2021</xref>). These integrations aim to further bolster the privacy and security guarantees of federated learning systems, making them more robust against potential adversarial attacks. Additionally, advancements in communication protocols have been a focal point of research, with efforts aimed at reducing the bandwidth required for model updates (Qin et al., <xref ref-type="bibr" rid="B96">2021</xref>). Techniques such as model compression and quantization are being employed to minimize the size of the updates transmitted between devices and the central server, thus improving the overall efficiency of the federated learning process.</p></sec>
<sec>
<title>5.2.2 Distributed learning</title>
<p>Distributed learning represents a significant evolution in machine learning paradigms, extending the principles of federated learning to encompass a broader range of computational resources, including both edge devices and cloud servers. This approach capitalizes on the strengths of edge devices for real-time data processing, while simultaneously leveraging the substantial computational power of cloud servers for more resource-intensive tasks. By dynamically distributing model components across various devices and servers, distributed learning effectively balances latency and computational efficiency, making it particularly suitable for applications that require rapid response times and substantial processing capabilities (Kang et al., <xref ref-type="bibr" rid="B54">2017</xref>).</p>
<p>One notable technique is split learning, which involves partitioning a neural network into segments that can be processed on different devices. In this framework, the initial layers of the model may run on edge devices, which handle local data and perform preliminary processing. The subsequent layers are executed on cloud servers, where the more complex computations take place (Vepakomma et al., <xref ref-type="bibr" rid="B122">2018</xref>). This separation not only reduces the computational burden on edge devices but also minimizes the amount of data that needs to be transmitted to the cloud, thereby enhancing privacy and reducing communication costs (Zhang et al., <xref ref-type="bibr" rid="B145">2024</xref>).</p>
<p>Recent advancements have also focused on optimizing the training process in distributed learning scenarios. Techniques such as adaptive resource allocation (Li J.-Y. et al., <xref ref-type="bibr" rid="B69">2023</xref>) and dynamic task scheduling (de Zarz&#x000E0; et al., <xref ref-type="bibr" rid="B25">2023</xref>) are being explored to ensure that computational resources are utilized effectively. By intelligently assigning tasks based on the current workload and capabilities of each device, distributed learning systems can achieve improved performance and responsiveness. Moreover, the integration of distributed learning with edge computing and the IoT has opened new avenues for real-time data analytics and decision-making (Saha et al., <xref ref-type="bibr" rid="B104">2021</xref>). For instance, in smart cities, distributed learning can enable the analysis of data generated by various sensors and devices to optimize traffic management, energy consumption, and public safety measures. This real-time processing capability is crucial for applications that require immediate insights and actions based on rapidly changing data.</p>
</sec>
</sec>
<sec>
<title>5.3 Optimized architectures and frameworks for edge deployment</title>
<p>Optimized architectures and frameworks have been specifically developed to facilitate the deployment of LLMs on edge devices. These solutions focus on creating lightweight and efficient models that significantly reduce the computational demands traditionally associated with LLMs. By optimizing both the architecture and the underlying frameworks, these tailored systems ensure that edge devices can effectively perform a wide range of NLP tasks without sacrificing performance or functionality. This advancement not only enhances the usability of LLMs in resource-constrained environments but also broadens their applicability in scenarios requiring real-time processing and low-latency responses.</p>
<sec>
<title>5.3.1 Lightweight model architectures</title>
<p>Lightweight model architectures are vital in NLP, especially for LLMs. They seek to optimize complexity and resource usage without sacrificing performance, driven by the need to deploy advanced models on resource-limited devices like smartphones and edge computing platforms.</p>
<p>One notable example of a lightweight architecture is ALBERT, which introduces several innovative modifications to the original BERT architecture (Lan et al., <xref ref-type="bibr" rid="B65">2019</xref>). By sharing parameters across different layers, ALBERT significantly reduces memory requirements without compromising the model&#x00027;s ability to understand and generate human-like text. This parameter-sharing technique not only enhances efficiency but also facilitates faster training and inference times, making it a compelling choice for applications that require rapid processing. Another significant development in lightweight model architectures is MobileBERT. This model incorporates bottleneck structures and various optimizations specifically designed to enable BERT to function effectively on mobile and edge devices (Sun et al., <xref ref-type="bibr" rid="B117">2020</xref>). By leveraging these architectural innovations, MobileBERT achieves a balance between model size and performance, allowing it to deliver high-quality natural language understanding capabilities in environments where computational resources are limited.</p>
<p>The exploration of hybrid architectures that integrate lightweight models with more complex systems has also gained traction. These hybrid approaches allow for the efficient handling of various tasks by dynamically allocating resources based on the particular demands of each task, optimizing both performance and efficiency (Sun et al., <xref ref-type="bibr" rid="B115">2022</xref>; Bayoudh, <xref ref-type="bibr" rid="B12">2024</xref>).</p></sec>
<sec>
<title>5.3.2 Edge-AI frameworks</title>
<p>Edge-AI frameworks have become instrumental in enabling the deployment of LLMs on edge devices, addressing the challenges posed by limited computational resources and varying hardware capabilities (Ignatov et al., <xref ref-type="bibr" rid="B45">2018</xref>). Notable examples of such frameworks include TensorFlow Lite, ONNX Runtime, and PyTorch Mobile, all provide tools to enhance LLM implementation in resource-limited edge computing environments (Danopoulos et al., <xref ref-type="bibr" rid="B24">2021</xref>).</p>
<p>TensorFlow Lite is a lightweight version of TensorFlow specifically tailored for mobile and edge applications. It provides robust support for model conversion and optimization techniques, such as quantization (Pandey and Asati, <xref ref-type="bibr" rid="B92">2023</xref>). Quantization reduces the precision of model weights and activations, leading to a significant decrease in model size and an increase in inference speed without a substantial loss in accuracy. TensorFlow Lite also includes acceleration capabilities for various hardware architectures, enabling efficient execution on mobile CPUs and GPUs (Adi and Casson, <xref ref-type="bibr" rid="B2">2021</xref>). This adaptability makes it particularly suitable for applications that require real-time processing, such as voice assistants and interactive chatbots.</p>
<p>ONNX Runtime is designed to facilitate cross-platform compatibility for optimized model execution. It supports models developed in various frameworks, allowing developers to leverage the strengths of different tools while maintaining a consistent runtime environment (Kim et al., <xref ref-type="bibr" rid="B60">2022</xref>). ONNX Runtime incorporates several optimization techniques, including graph optimization and kernel fusion, which enhance the performance of models during inference (Niu et al., <xref ref-type="bibr" rid="B91">2021</xref>). This framework is particularly beneficial for deploying LLMs across diverse hardware platforms, ensuring that models can be executed efficiently, whether on edge devices, cloud servers, or hybrid environments.</p>
<p>PyTorch Mobile has proven to be a robust solution for implementing machine learning models on edge devices. It allows developers to convert PyTorch models into a format optimized for mobile environments, ensuring that they can run efficiently on both Android and iOS platforms (Deng, <xref ref-type="bibr" rid="B27">2019</xref>). PyTorch Mobile supports a variety of optimizations, including quantization and pruning, to reduce model size and improve inference speed. Additionally, it provides tools for dynamic model updates (Li M. et al., <xref ref-type="bibr" rid="B70">2022</xref>), enabling applications to adapt to new data or requirements without necessitating a complete redeployment.</p>
<p>Recent advancements in these frameworks have further expanded their capabilities. For instance, TensorFlow Lite, ONNX Runtime, and PyTorch Mobile have integrated support for hardware-specific optimizations that take advantage of the unique features of different processors, such as ARM and NVIDIA GPUs. These optimizations allow models to run more efficiently, maximizing the performance of edge devices while minimizing energy consumption. Additionally, the emergence of new model compression methods, has been incorporated into these frameworks (Pandey and Asati, <xref ref-type="bibr" rid="B92">2023</xref>). Pruning involves removing less significant weights from a model, resulting in a sparser representation that requires fewer resources for inference. Knowledge distillation enables the creation of smaller student models that can mimic the performance of larger teacher models, making it easier to deploy high-performing models on edge devices.</p>
<p>In summary, the advancement of model compression techniques, including pruning, quantization, and knowledge distillation, plays a pivotal role in enhancing the deployment of LLMs on edge devices. These techniques enable the reduction of model size and complexity, allowing for efficient utilization of limited computational resources inherent in edge environments. Pruning effectively removes redundant parameters from models, thereby streamlining their structure without significantly compromising performance. Quantization reduces the precision of the model weights, which not only decreases memory usage but also accelerates inference times. Knowledge distillation, on the other hand, involves training a smaller, more efficient model to replicate the behavior of a larger model, thus maintaining performance while ensuring that the model is lightweight and suitable for edge deployment.</p>
<p>Additionally, federated and distributed learning frameworks offer innovative solutions for training models across decentralized data sources while preserving data privacy. These approaches allow for collaborative learning without the need to transfer sensitive information to a central server, thus enhancing security and compliance with privacy regulations.</p>
<p>Furthermore, the development of optimized architectures and frameworks, such as lightweight model architectures and edge-AI frameworks, is essential for facilitating efficient model deployment in edge computing scenarios. Lightweight architectures are specifically designed to operate within the constraints of edge devices, ensuring that models can deliver high performance with minimal resource consumption.</p>
<p>By addressing these interconnected challenges and leveraging these advanced techniques, we can unlock the full potential of LLMs in edge computing, paving the way for smarter, more responsive systems. These advancements will not only meet the evolving needs of users but also ensure that privacy and security are maintained throughout the data analysis process.</p>
</sec>
</sec>
</sec>
<sec id="s6">
<title>6 Future directions</title>
<p>The field of LLMs in edge computing is dynamic and constantly evolving. To fully leverage LLMs for edge applications, researchers are focusing on several future directions, including ultra-efficient Transformer architectures, adaptive deployment models, edge-specific hardware, enhanced privacy, multi-modal LLMs, and hybrid edge-cloud systems. This section discusses these emerging areas.</p>
<sec>
<title>6.1 Ultra-efficient transformer architectures for edge</title>
<p>Developing ultra-efficient Transformer architectures is vital for deploying LLMs in edge environments with limited resources. Current models, such as MobileBERT and TinyBERT, achieve significant efficiency gains by reducing the parameter count and computational complexity (Sun et al., <xref ref-type="bibr" rid="B117">2020</xref>). MobileBERT, for example, is a compact, task-agnostic model that uses bottleneck structures and parameter sharing across layers, enabling high performance in edge scenarios with constrained processing capabilities. By reducing the model size and adapting it for specific tasks, MobileBERT represents a viable approach for deploying Transformers on edge devices where real-time processing is essential.</p>
<p>Further research into efficient model architectures, such as sparse Transformers and lightweight attention mechanisms, could yield even more optimized models for edge deployment. Sparse Transformers reduce memory and processing requirements by using selective attention, focusing computational resources only on relevant data (Jaszczur et al., <xref ref-type="bibr" rid="B50">2021</xref>). Techniques like these allow models to operate with minimal resources, making them suitable for applications requiring continuous processing, such as wearable health monitors and environmental sensors. As the need for efficient edge-compatible LLMs grows, more research will likely focus on refining these architectures to achieve better performance without sacrificing accuracy.</p>
<p>The future of edge-based LLMs may also involve adaptive Transformers that adjust their complexity based on input characteristics. By enabling models to dynamically allocate computational resources, adaptive Transformers optimize processing power and latency, which is important for edge applications that need quick responses amid varying resource availability. These innovations are expected to enhance the accessibility and utility of LLMs across various edge computing scenarios.</p>
</sec>
<sec>
<title>6.2 Adaptive and context-aware model deployment</title>
<p>To enhance flexibility in edge computing, future LLMs will likely feature adaptive deployment mechanisms, allowing models to adjust to their operating context. Context-aware models can dynamically allocate resources based on device capabilities, user needs, or network conditions, optimizing processing efficiency and power consumption (Neseem et al., <xref ref-type="bibr" rid="B88">2023</xref>). For instance, early-exit strategies enable models to terminate processing once they reach an acceptable confidence level, saving computational resources and reducing latency. This approach is particularly beneficial for edge environments where device capabilities vary significantly.</p>
<p>Another promising approach is on-device model compression, where edge devices themselves can prune or quantize models in real-time based on task requirements. This adaptability is especially useful for consumer devices such as smartphones, where power and memory limitations fluctuate depending on user activity and battery status. Research into self-optimizing models that adjust based on operational data could further improve LLM performance in edge settings, allowing devices to perform complex analytics even in low-power modes.</p>
<p>The continued development of context-aware models could also lead to smarter load balancing between edge and cloud resources. By dynamically offloading certain tasks to cloud resources based on network conditions and processing demands, adaptive deployment strategies can enhance responsiveness and resource management. As edge computing applications grow more complex, adaptable LLMs will become essential for delivering real-time insights without straining device resources.</p>
</sec>
<sec>
<title>6.3 Hardware-accelerated edge AI</title>
<p>Hardware advancements specifically designed for edge AI will play a crucial role in supporting complex LLMs on resource-constrained devices. AI accelerators like FPGAs, edge Tensor Processing Units (TPUs), Neural Network Processing Units (NPUs), and neuromorphic processors are being developed to execute LLM inference tasks with lower power consumption and higher efficiency compared to traditional CPUs (Krestinskaya et al., <xref ref-type="bibr" rid="B64">2019</xref>).</p>
<p>Edge TPUs are tailored to perform high-speed inferences on deep learning models, offering a solution for real-time applications that require continuous processing while preserving power efficiency (Akin et al., <xref ref-type="bibr" rid="B4">2022</xref>; Shuvo et al., <xref ref-type="bibr" rid="B112">2022</xref>). NPUs represent another class of specialized hardware that is increasingly being utilized in edge computing. NPUs are specifically designed to accelerate deep learning tasks by optimizing the execution of neural networks (Jang et al., <xref ref-type="bibr" rid="B49">2021</xref>). They offer significant advantages in terms of throughput and energy efficiency, making them particularly suitable for real-time applications in mobile devices and IoT. By enabling high-speed computations and supporting parallel processing of neural network layers, NPUs facilitate the deployment of complex LLMs directly on edge devices, thereby enhancing responsiveness and preserving user privacy (Heo et al., <xref ref-type="bibr" rid="B42">2024</xref>; Xu et al., <xref ref-type="bibr" rid="B131">2024</xref>).</p>
<p>Neuromorphic computing represents a promising direction for supporting LLMs in edge environments. Neuromorphic processors use spiking neural networks to perform computations, which can significantly reduce power usage compared to traditional deep learning hardware. This technology holds great potential for applications such as autonomous drones and mobile health monitors, where low-latency, energy-efficient processing is critical (Schuman et al., <xref ref-type="bibr" rid="B109">2022</xref>).</p>
<p>In addition to traditional AI accelerators, the field of quantum computing is emerging as a revolutionary force that could reshape the landscape of machine learning and edge AI (Liang et al., <xref ref-type="bibr" rid="B73">2023</xref>). Quantum processors utilize the principles of quantum mechanics to execute computations at unprecedented speeds, potentially enabling the processing of complex models that are currently impractical with classical hardware. While still in its nascent stages, quantum computing holds the promise of improving the training and inference capabilities of LLMs, especially for tasks that demand significant computational resources (Aizpurua et al., <xref ref-type="bibr" rid="B3">2024</xref>). Researchers are investigating hybrid methods that integrate classical and quantum processing. This approach would allow edge devices to offload demanding computations to quantum processors while maintaining real-time responsiveness for less intensive tasks.</p>
<p>As hardware development progresses, the integration of specialized AI accelerators in edge devices will enhance the performance and efficiency of LLMs, enabling more sophisticated applications. Continued research and innovation in this area will help overcome one of the key barriers to deploying LLMs on the edge: the high computational demand of complex models.</p>
</sec>
<sec>
<title>6.4 Data privacy and federated security models</title>
<p>Data privacy remains a top priority as LLMs handle increasingly sensitive information in edge environments. Methods such as differential privacy, homomorphic encryption, and secure multi-party computation offer means to safeguard data while preserving its usability (Phan et al., <xref ref-type="bibr" rid="B95">2017</xref>). Differential privacy ensures that individual data points are anonymized, allowing LLMs to analyze data collectively without revealing personal information. Homomorphic encryption, meanwhile, enables computations on encrypted data, which is invaluable for applications where data must remain secure throughout the processing lifecycle.</p>
<p>Federated learning has emerged as a powerful tool for training models on decentralized data sources while preserving user privacy (Kairouz et al., <xref ref-type="bibr" rid="B53">2021</xref>). However, traditional federated learning approaches face challenges in terms of communication efficiency and handling non-IID data. Advances in hierarchical federated learning and split learning, where portions of the model are trained locally and others centrally, offer solutions to these challenges by reducing communication overhead and ensuring robust model updates across devices (Vepakomma et al., <xref ref-type="bibr" rid="B122">2018</xref>). Future research will likely focus on refining these methods to make federated learning more adaptable for complex LLMs and diverse edge applications.</p>
<p>As privacy-preserving technologies evolve, integrating them into edge-compatible LLMs will allow for safer, more responsible deployment in sensitive domains like healthcare, finance, and smart cities. Ensuring that LLMs operate within ethical and regulatory boundaries will be essential for broadening their adoption in edge computing.</p>
</sec>
<sec>
<title>6.5 Multi-modal LLMs for edge applications</title>
<p>The development of multi-modal LLMs that can process diverse data types, including text, images, and audio, is a key future direction for edge applications. Multi-modal capabilities allow models to provide more comprehensive insights by analyzing various data streams simultaneously, which is particularly useful in autonomous systems and IoT applications (Radford et al., <xref ref-type="bibr" rid="B98">2021</xref>). For example, a multi-modal LLM in a smart vehicle could analyze visual data from cameras and textual data from sensors to enhance object detection and navigation (Xie et al., <xref ref-type="bibr" rid="B130">2022</xref>).</p>
<p>Edge devices equipped with multi-modal LLMs can also improve situational awareness in smart cities, processing real-time data from transport infrastructure, air quality sensors, and emergency alerts. By deploying multi-modal LLMs on the edge, systems can respond more quickly and intelligently to real-world events without relying on cloud-based processing. Such real-time, multi-modal analytics is vital for applications where latency could impact safety or operational effectiveness.</p>
<p>Future research in multi-modal LLMs will focus on developing lightweight architectures that integrate multiple data types without excessive computational demand. These advancements will expand the range of edge applications, enabling devices to handle complex tasks in real-time while conserving resources.</p>
</sec>
<sec>
<title>6.6 Hybrid edge-cloud architectures and collaborative intelligence</title>
<p>Hybrid edge-cloud architectures address the constraints of fully decentralized edge computing by optimizing resource distribution between the edge and the cloud (Kang et al., <xref ref-type="bibr" rid="B54">2017</xref>). Additionally, methods such as differential privacy, homomorphic encryption, and secure multi-party computation enable data protection without sacrificing functionality. This collaborative intelligence framework allows for dynamic adjustments in workload distribution, improving performance in applications that require a mix of local and remote processing.</p>
<p>Collaborative edge-cloud architectures are particularly beneficial for applications such as smart cities and industrial IoT, where devices need both rapid, localized responses and access to extensive computational resources. For instance, edge devices in a smart factory might analyze sensor data to detect defects in real-time, while more complex, large-scale analytics are processed in the cloud to optimize production processes. By leveraging both edge and cloud resources, hybrid architectures can provide a flexible and scalable solution for deploying LLMs across distributed environments.</p>
<p>Hybrid systems will likely involve more advanced orchestration algorithms that optimize resource allocation based on real-time conditions, such as network latency, device availability, and task complexity. These systems promise to make LLMs more scalable and adaptable, facilitating a new generation of intelligent, responsive edge applications.</p>
<p>The future of LLMs in edge computing will be driven by advancements in model efficiency, adaptive deployment, specialized hardware, privacy, multi-modal capabilities, and hybrid architectures. Together, these innovations will expand the capabilities of edge-based LLMs, enabling them to process data in real time, protect user privacy, and respond flexibly to diverse operational requirements. The continued exploration of these areas will be crucial for realizing the full potential of LLMs in edge computing.</p>
</sec>
</sec>
<sec sec-type="conclusions" id="s7">
<title>7 Conclusion</title>
<p>The convergence of LLMs and edge computing signifies a groundbreaking advancement in the field of data analysis, with the potential to redefine operational frameworks across a multitude of industries. This integration offers a transformative opportunity to enhance real-time insights, bolster data privacy, and enable autonomous decision-making, addressing critical demands in an increasingly data-driven world. By decentralizing computational capabilities and bringing them closer to the source of data generation, the deployment of LLMs on edge devices addresses key challenges such as latency, scalability, and security, while simultaneously paving the way for a new era of intelligent, localized data processing.</p>
<p>The deployment of LLMs on edge devices is accompanied by significant challenges. Resource constraints remain a primary obstacle, as edge devices typically have limited computational power, memory, and energy efficiency compared to centralized cloud systems. Addressing these constraints requires the development of lightweight models through techniques such as model compression, pruning, quantization, and knowledge distillation, which reduce model size while preserving performance. Additionally, ensuring data privacy and security in distributed edge environments is critical, as edge devices are often more vulnerable to targeted attacks. Robust security mechanisms, adaptive frameworks, and real-time threat mitigation strategies are essential to protect sensitive data and maintain trust. Another challenge lies in the variability and unpredictability of edge environments, which can affect the performance and reliability of deployed models. Adaptive learning techniques, continuous model updates, and mechanisms for dynamic optimization are necessary to ensure that LLMs remain effective and relevant in changing operational contexts.</p>
<p>The convergence of LLMs and edge computing represents not just a technological innovation but a paradigm shift in how data is processed, analyzed, and utilized across industries. By addressing the unique challenges of edge deployments and leveraging cutting-edge techniques such as model compression, federated learning, and collaborative frameworks, this integration offers transformative benefits. The ability to derive actionable insights in real time, while preserving data privacy and enabling autonomous decision-making, has far-reaching implications for a wide range of sectors.</p></sec>
</body>
<back>
<sec sec-type="author-contributions" id="s8">
<title>Author contributions</title>
<p>XW: Conceptualization, Resources, Supervision, Writing &#x02013; original draft, Writing &#x02013; review &#x00026; editing. ZX: Methodology, Visualization, Writing &#x02013; original draft. XS: Software, Validation, Visualization, Writing &#x02013; original draft.</p>
</sec>
<sec sec-type="funding-information" id="s9">
<title>Funding</title>
<p>The author(s) declare that no financial support was received for the research and/or publication of this article.</p>
</sec>
<sec sec-type="COI-statement" id="conf1">
<title>Conflict of interest</title>
<p>XW, ZX, and XS were employed by CNOOC Safety Technology Services Co., Ltd.</p>
</sec>
<sec id="s10">
<title>Generative AI statement</title>
<p>The author(s) declare that no Gen AI was used in the creation of this manuscript.</p></sec>
<sec sec-type="disclaimer" id="s11">
<title>Publisher&#x00027;s note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Abreha</surname> <given-names>H. G.</given-names></name> <name><surname>Hayajneh</surname> <given-names>M.</given-names></name> <name><surname>Serhani</surname> <given-names>M. A.</given-names></name></person-group> (<year>2022</year>). <article-title>Federated learning in edge computing: a systematic survey</article-title>. <source>Sensors</source> <volume>22</volume>:<fpage>450</fpage>. <pub-id pub-id-type="doi">10.3390/s22020450</pub-id><pub-id pub-id-type="pmid">35062410</pub-id></citation></ref>
<ref id="B2">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Adi</surname> <given-names>S. E.</given-names></name> <name><surname>Casson</surname> <given-names>A. J.</given-names></name></person-group> (<year>2021</year>). <article-title>&#x0201C;Design and optimization of a tensorflow lite deep learning neural network for human activity recognition on a smartphone,&#x0201D;</article-title> in <source>2021 43rd Annual International Conference of the IEEE Engineering in Medicine</source> &#x00026; <italic>Biology Society (EMBC)</italic> (Mexico: IEEE), <fpage>7028</fpage>&#x02013;<lpage>7031</lpage>.<pub-id pub-id-type="pmid">34892721</pub-id></citation></ref>
<ref id="B3">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Aizpurua</surname> <given-names>B.</given-names></name> <name><surname>Jahromi</surname> <given-names>S. S.</given-names></name> <name><surname>Singh</surname> <given-names>S.</given-names></name> <name><surname>Orus</surname> <given-names>R.</given-names></name></person-group> (<year>2024</year>). <article-title>Quantum large language models via tensor network disentanglers</article-title>. <source>arXiv</source> [preprint] arXiv:2410.17397. <pub-id pub-id-type="doi">10.48550/arXiv.2410.17397</pub-id></citation>
</ref>
<ref id="B4">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Akin</surname> <given-names>B.</given-names></name> <name><surname>Gupta</surname> <given-names>S.</given-names></name> <name><surname>Long</surname> <given-names>Y.</given-names></name> <name><surname>Spiridonov</surname> <given-names>A.</given-names></name> <name><surname>Wang</surname> <given-names>Z.</given-names></name> <name><surname>White</surname> <given-names>M.</given-names></name> <etal/></person-group>. (<year>2022</year>). <article-title>&#x0201C;Searching for efficient neural architectures for on-device ml on edge TPUS,&#x0201D;</article-title> in <source>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition</source> (<publisher-loc>New Orleans, LA</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>2667</fpage>&#x02013;<lpage>2676</lpage>.</citation>
</ref>
<ref id="B5">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Alahi</surname> <given-names>M. E. E.</given-names></name> <name><surname>Sukkuea</surname> <given-names>A.</given-names></name> <name><surname>Tina</surname> <given-names>F. W.</given-names></name> <name><surname>Nag</surname> <given-names>A.</given-names></name> <name><surname>Kurdthongmee</surname> <given-names>W.</given-names></name> <name><surname>Suwannarat</surname> <given-names>K.</given-names></name> <etal/></person-group>. (<year>2023</year>). <article-title>Integration of iot-enabled technologies and artificial intelligence (ai) for smart city scenario: recent advancements and future trends</article-title>. <source>Sensors</source> <volume>23</volume>:<fpage>5206</fpage>. <pub-id pub-id-type="doi">10.3390/s23115206</pub-id><pub-id pub-id-type="pmid">37299934</pub-id></citation></ref>
<ref id="B6">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ali</surname> <given-names>B.</given-names></name> <name><surname>Gregory</surname> <given-names>M. A.</given-names></name> <name><surname>Li</surname> <given-names>S.</given-names></name></person-group> (<year>2021</year>). <article-title>Multi-access edge computing architecture, data security and privacy: A review</article-title>. <source>IEEE Access</source> <volume>9</volume>, <fpage>18706</fpage>&#x02013;<lpage>18721</lpage>. <pub-id pub-id-type="doi">10.1109/ACCESS.2021.3053233</pub-id></citation>
</ref>
<ref id="B7">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Alwarafy</surname> <given-names>A.</given-names></name> <name><surname>Al-Thelaya</surname> <given-names>K. A.</given-names></name> <name><surname>Abdallah</surname> <given-names>M.</given-names></name> <name><surname>Schneider</surname> <given-names>J.</given-names></name> <name><surname>Hamdi</surname> <given-names>M.</given-names></name></person-group> (<year>2021</year>). <article-title>A survey on security and privacy issues in edge-computing-assisted internet of things</article-title>. <source>IEEE Internet Things J</source>. <volume>8</volume>, <fpage>4004</fpage>&#x02013;<lpage>4022</lpage>. <pub-id pub-id-type="doi">10.1109/JIOT.2020.3015432</pub-id><pub-id pub-id-type="pmid">39204906</pub-id></citation></ref>
<ref id="B8">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Anwar</surname> <given-names>S.</given-names></name> <name><surname>Hwang</surname> <given-names>K.</given-names></name> <name><surname>Sung</surname> <given-names>W.</given-names></name></person-group> (<year>2017</year>). <article-title>Structured pruning of deep convolutional neural networks</article-title>. <source>ACM J. Emerg. Technol. Comp. Syst</source>. <volume>13</volume>, <fpage>1</fpage>&#x02013;<lpage>18</lpage>. <pub-id pub-id-type="doi">10.1145/3005348</pub-id></citation>
</ref>
<ref id="B9">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Babu</surname> <given-names>M.</given-names></name> <name><surname>Lautman</surname> <given-names>Z.</given-names></name> <name><surname>Lin</surname> <given-names>X.</given-names></name> <name><surname>Sobota</surname> <given-names>M. H.</given-names></name> <name><surname>Snyder</surname> <given-names>M. P.</given-names></name></person-group> (<year>2024</year>). <article-title>Wearable devices: implications for precision medicine and the future of health care</article-title>. <source>Annu. Rev. Med</source>. <volume>75</volume>, <fpage>401</fpage>&#x02013;<lpage>415</lpage>. <pub-id pub-id-type="doi">10.1146/annurev-med-052422-020437</pub-id><pub-id pub-id-type="pmid">37983384</pub-id></citation></ref>
<ref id="B10">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Bao</surname> <given-names>Z.</given-names></name> <name><surname>Liu</surname> <given-names>J.</given-names></name> <name><surname>Zhang</surname> <given-names>W.</given-names></name></person-group> (<year>2019</year>). <article-title>&#x0201C;Using distillation to improve network performance after pruning and quantization,&#x0201D;</article-title> in <source>Proceedings of the 2019 2nd International Conference on Machine Learning and Machine Intelligence, MLMI &#x00027;19</source> (<publisher-loc>New York, NY</publisher-loc>: <publisher-name>Association for Computing Machinery</publisher-name>), <fpage>3</fpage>&#x02013;<lpage>6</lpage>.<pub-id pub-id-type="pmid">40030406</pub-id></citation></ref>
<ref id="B11">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Barua</surname> <given-names>A.</given-names></name> <name><surname>Dong</surname> <given-names>C.</given-names></name> <name><surname>Al-Turjman</surname> <given-names>F.</given-names></name> <name><surname>Yang</surname> <given-names>X.</given-names></name></person-group> (<year>2020</year>). <article-title>Edge computing-based localization technique to detecting behavior of dementia</article-title>. <source>IEEE Access</source> <volume>8</volume>, <fpage>82108</fpage>&#x02013;<lpage>82119</lpage>. <pub-id pub-id-type="doi">10.1109/ACCESS.2020.2988935</pub-id></citation>
</ref>
<ref id="B12">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bayoudh</surname> <given-names>K.</given-names></name></person-group> (<year>2024</year>). <article-title>A survey of multimodal hybrid deep learning for computer vision: Architectures, applications, trends, and challenges</article-title>. <source>Inform. Fusion</source> <volume>105</volume>:<fpage>102217</fpage>. <pub-id pub-id-type="doi">10.1016/j.inffus.2023.102217</pub-id></citation>
</ref>
<ref id="B13">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Boumendil</surname> <given-names>A.</given-names></name> <name><surname>Bechkit</surname> <given-names>W.</given-names></name> <name><surname>Benatchba</surname> <given-names>K.</given-names></name></person-group> (<year>2025</year>). <article-title>On-device deep learning: survey on techniques improving energy efficiency of DNNS</article-title>. <source>IEEE Trans. Neural Netw. Learn. Syst</source>. <volume>36</volume>, <fpage>7806</fpage>&#x02013;<lpage>7821</lpage>. <pub-id pub-id-type="doi">10.1109/TNNLS.2024.3430028</pub-id><pub-id pub-id-type="pmid">39046860</pub-id></citation></ref>
<ref id="B14">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Brown</surname> <given-names>T.</given-names></name> <name><surname>Mann</surname> <given-names>B.</given-names></name> <name><surname>Ryder</surname> <given-names>N.</given-names></name> <name><surname>Subbiah</surname> <given-names>M.</given-names></name> <name><surname>Kaplan</surname> <given-names>J.</given-names></name> <name><surname>Dhariwal</surname> <given-names>P.</given-names></name> <etal/></person-group>. (<year>2020</year>). <article-title>Language models are few-shot learners</article-title>. <source>arXiv</source> preprint arXiv:2005.14165. <pub-id pub-id-type="doi">10.48550/arXiv.2005.14165</pub-id></citation>
</ref>
<ref id="B15">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Byrd</surname> <given-names>D.</given-names></name> <name><surname>Polychroniadou</surname> <given-names>A.</given-names></name></person-group> (<year>2021</year>). <article-title>&#x0201C;Differentially private secure multi-party computation for federated learning in financial applications,&#x0201D;</article-title> in <source>Proceedings of the First ACM International Conference on AI in Finance, ICAIF &#x00027;20</source> (<publisher-loc>New York, NY</publisher-loc>: <publisher-name>Association for Computing Machinery</publisher-name>).</citation>
</ref>
<ref id="B16">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cao</surname> <given-names>K.</given-names></name> <name><surname>Liu</surname> <given-names>Y.</given-names></name> <name><surname>Meng</surname> <given-names>G.</given-names></name> <name><surname>Sun</surname> <given-names>Q.</given-names></name></person-group> (<year>2020</year>). <article-title>An overview on edge computing research</article-title>. <source>IEEE Access</source> <volume>8</volume>, <fpage>85714</fpage>&#x02013;<lpage>85728</lpage>. <pub-id pub-id-type="doi">10.1109/ACCESS.2020.2991734</pub-id></citation>
</ref>
<ref id="B17">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Chen</surname> <given-names>K.</given-names></name> <name><surname>Li</surname> <given-names>T.</given-names></name> <name><surname>Kim</surname> <given-names>H.-S.</given-names></name> <name><surname>Culler</surname> <given-names>D. E.</given-names></name> <name><surname>Katz</surname> <given-names>R. H.</given-names></name></person-group> (<year>2018</year>). <article-title>&#x0201C;Marvel: Enabling mobile augmented reality with low energy and low latency,&#x0201D;</article-title> in <source>Proceedings of the 16th ACM Conference on Embedded Networked Sensor Systems, SenSys &#x00027;18</source> (<publisher-loc>New York, NY</publisher-loc>: <publisher-name>Association for Computing Machinery</publisher-name>).</citation>
</ref>
<ref id="B18">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Chen</surname> <given-names>W.</given-names></name> <name><surname>Qiu</surname> <given-names>H.</given-names></name> <name><surname>Zhuang</surname> <given-names>J.</given-names></name> <name><surname>Zhang</surname> <given-names>C.</given-names></name> <name><surname>Hu</surname> <given-names>Y.</given-names></name> <name><surname>Lu</surname> <given-names>Q.</given-names></name> <etal/></person-group>. (<year>2021a</year>). <article-title>Quantization of deep neural networks for accurate edge computing</article-title>. <source>ACM J. Emerg. Technol. Comp. Syst</source>. (<italic>JETC)</italic> <volume>17</volume>:<fpage>4</fpage>. <pub-id pub-id-type="doi">10.1145/3451211</pub-id></citation>
</ref>
<ref id="B19">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Chen</surname> <given-names>W.</given-names></name> <name><surname>Wang</surname> <given-names>P.</given-names></name> <name><surname>Cheng</surname> <given-names>J.</given-names></name></person-group> (<year>2021b</year>). <article-title>&#x0201C;Towards mixed-precision quantization of neural networks via constrained optimization,&#x0201D;</article-title> in <source>Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV)</source> (<publisher-loc>Montreal, QC</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>5350</fpage>&#x02013;<lpage>5359</lpage>. <pub-id pub-id-type="doi">10.1109/ICCV48922.2021.00530</pub-id></citation>
</ref>
<ref id="B20">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cheng</surname> <given-names>L.</given-names></name> <name><surname>Gu</surname> <given-names>Y.</given-names></name> <name><surname>Liu</surname> <given-names>Q.</given-names></name> <name><surname>Yang</surname> <given-names>L.</given-names></name> <name><surname>Liu</surname> <given-names>C.</given-names></name> <name><surname>Wang</surname> <given-names>Y.</given-names></name></person-group> (<year>2024</year>). <article-title>Advancements in accelerating deep neural network inference on aiot devices: A survey</article-title>. <source>IEEE Trans. Sustain. Comput.</source> <volume>9</volume>, <fpage>830</fpage>&#x02013;<lpage>847</lpage>. <pub-id pub-id-type="doi">10.1109/TSUSC.2024.3353176</pub-id></citation>
</ref>
<ref id="B21">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Chi</surname> <given-names>J.</given-names></name> <name><surname>Zhou</surname> <given-names>X.</given-names></name> <name><surname>Xiao</surname> <given-names>F.</given-names></name> <name><surname>Lim</surname> <given-names>Y.</given-names></name> <name><surname>Qiu</surname> <given-names>T.</given-names></name></person-group> (<year>2024</year>). <article-title>Task offloading via prioritized experience-based double dueling dqn in edge-assisted IIoT</article-title>. <source>IEEE Trans. Mobile Comp</source>. <volume>23</volume>, <fpage>14575</fpage>&#x02013;<lpage>14591</lpage>. <pub-id pub-id-type="doi">10.1109/TMC.2024.3452502</pub-id></citation>
</ref>
<ref id="B22">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Chowdhury</surname> <given-names>S. S.</given-names></name> <name><surname>Garg</surname> <given-names>I.</given-names></name> <name><surname>Roy</surname> <given-names>K.</given-names></name></person-group> (<year>2021</year>). <article-title>&#x0201C;Spatio-temporal pruning and quantization for low-latency spiking neural networks,&#x0201D;</article-title> in <source>2021 International Joint Conference on Neural Networks (IJCNN)</source> (<publisher-loc>Shenzhen</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>1</fpage>&#x02013;<lpage>9</lpage>.</citation>
</ref>
<ref id="B23">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cicconetti</surname> <given-names>C.</given-names></name> <name><surname>Conti</surname> <given-names>M.</given-names></name> <name><surname>Passarella</surname> <given-names>A.</given-names></name></person-group> (<year>2021</year>). <article-title>A decentralized framework for serverless edge computing in the internet of things</article-title>. <source>IEEE Trans. Netw. Serv. Managem</source>.<volume>18</volume>, <fpage>2166</fpage>&#x02013;<lpage>2180</lpage>. <pub-id pub-id-type="doi">10.1109/TNSM.2020.3023305</pub-id></citation>
</ref>
<ref id="B24">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Danopoulos</surname> <given-names>D.</given-names></name> <name><surname>Kachris</surname> <given-names>C.</given-names></name> <name><surname>Soudris</surname> <given-names>D.</given-names></name></person-group> (<year>2021</year>). <article-title>Utilizing cloud FPGAs towards the open neural network standard</article-title>. <source>Sustain. Comp.: Inform. Syst</source>. <volume>30</volume>:<fpage>100520</fpage>. <pub-id pub-id-type="doi">10.1016/j.suscom.2021.100520</pub-id></citation>
</ref>
<ref id="B25">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>de Zarz&#x000E0;</surname> <given-names>I.</given-names></name> <name><surname>de Curt&#x000F2;</surname> <given-names>J.</given-names></name> <name><surname>Roig</surname> <given-names>G.</given-names></name> <name><surname>Calafate</surname> <given-names>C. T.</given-names></name></person-group> (<year>2023</year>). <article-title>LLM Multimodal Traffic Accident Forecasting</article-title>. <source>Sensors</source> <volume>23</volume>:<fpage>9225</fpage>. <pub-id pub-id-type="doi">10.3390/s23229225</pub-id><pub-id pub-id-type="pmid">38005612</pub-id></citation></ref>
<ref id="B26">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Deng</surname> <given-names>S.</given-names></name> <name><surname>Zhao</surname> <given-names>H.</given-names></name> <name><surname>Fang</surname> <given-names>W.</given-names></name> <name><surname>Yin</surname> <given-names>J.</given-names></name> <name><surname>Dustdar</surname> <given-names>S.</given-names></name> <name><surname>Zomaya</surname> <given-names>A. Y.</given-names></name></person-group> (<year>2020</year>). <article-title>Edge intelligence: The confluence of edge computing and artificial intelligence</article-title>. <source>IEEE Intern. Things J</source>. <volume>7</volume>, <fpage>7457</fpage>&#x02013;<lpage>7469</lpage>. <pub-id pub-id-type="doi">10.1109/JIOT.2020.2984887</pub-id></citation>
</ref>
<ref id="B27">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Deng</surname> <given-names>Y.</given-names></name></person-group> (<year>2019</year>). <article-title>&#x0201C;Deep learning on mobile devices: a review,&#x0201D;</article-title> in <source>Mobile Multimedia/Image Processing, Security, and Applications 2019</source>, eds. S. S. Agaian, V. K. Asari, and S. P. DelMarco (<publisher-loc>New York</publisher-loc>: <publisher-name>SPIE</publisher-name>).</citation>
</ref>
<ref id="B28">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Devlin</surname> <given-names>J.</given-names></name> <name><surname>Chang</surname> <given-names>M.-W.</given-names></name> <name><surname>Lee</surname> <given-names>K.</given-names></name> <name><surname>Toutanova</surname> <given-names>K.</given-names></name></person-group> (<year>2019</year>). <article-title>&#x0201C;BERT: Pre-training of deep bidirectional transformers for language understanding,&#x0201D;</article-title> in <source>Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies</source> (<publisher-loc>New York</publisher-loc>: <publisher-name>Association for Computational Linguistics</publisher-name>), <fpage>4171</fpage>&#x02013;<lpage>4186</lpage>.</citation>
</ref>
<ref id="B29">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Dias</surname> <given-names>D.</given-names></name> <name><surname>Paulo Silva Cunha</surname> <given-names>J.</given-names></name></person-group> (<year>2018</year>). <article-title>Wearable health devices&#x02014;vital sign monitoring, systems and technologies</article-title>. <source>Sensors</source> <volume>18</volume>:<fpage>2414</fpage>. <pub-id pub-id-type="doi">10.3390/s18082414</pub-id><pub-id pub-id-type="pmid">30044415</pub-id></citation></ref>
<ref id="B30">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ding</surname> <given-names>Y.</given-names></name> <name><surname>Wu</surname> <given-names>Y.</given-names></name> <name><surname>Huang</surname> <given-names>C.</given-names></name> <name><surname>Tang</surname> <given-names>S.</given-names></name> <name><surname>Wu</surname> <given-names>F.</given-names></name> <name><surname>Yang</surname> <given-names>Y.</given-names></name> <etal/></person-group>. (<year>2022</year>). <article-title>Nap: Neural architecture search with pruning</article-title>. <source>Neurocomputing</source> <volume>477</volume>, <fpage>85</fpage>&#x02013;<lpage>95</lpage>. <pub-id pub-id-type="doi">10.1016/j.neucom.2021.12.002</pub-id></citation>
</ref>
<ref id="B31">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Du</surname> <given-names>M.</given-names></name> <name><surname>Wang</surname> <given-names>K.</given-names></name> <name><surname>Xia</surname> <given-names>Z.</given-names></name> <name><surname>Zhang</surname> <given-names>Y.</given-names></name></person-group> (<year>2020</year>). <article-title>Differential privacy preserving of training model in wireless big data with edge computing</article-title>. <source>IEEE Trans. Big Data</source> <volume>6</volume>, <fpage>283</fpage>&#x02013;<lpage>295</lpage>. <pub-id pub-id-type="doi">10.1109/TBDATA.2018.2829886</pub-id></citation>
</ref>
<ref id="B32">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Esteva</surname> <given-names>A.</given-names></name> <name><surname>Robicquet</surname> <given-names>A.</given-names></name> <name><surname>Ramsundar</surname> <given-names>B.</given-names></name> <name><surname>Kuleshov</surname> <given-names>V.</given-names></name> <name><surname>DePristo</surname> <given-names>M.</given-names></name> <name><surname>Chou</surname> <given-names>K.</given-names></name> <etal/></person-group>. (<year>2019</year>). <article-title>A guide to deep learning in healthcare</article-title>. <source>Nat. Med</source>. <volume>25</volume>, <fpage>24</fpage>&#x02013;<lpage>29</lpage>. <pub-id pub-id-type="doi">10.1038/s41591-018-0316-z</pub-id><pub-id pub-id-type="pmid">30617335</pub-id></citation></ref>
<ref id="B33">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gale</surname> <given-names>T.</given-names></name> <name><surname>Elsen</surname> <given-names>E.</given-names></name> <name><surname>Hooker</surname> <given-names>S.</given-names></name></person-group> (<year>2019</year>). <article-title>The state of sparsity in deep neural networks</article-title>. <source>arXiv</source> [preprint] arXiv:1902.09574. <pub-id pub-id-type="doi">10.48550/arXiv.1902.09574</pub-id></citation>
</ref>
<ref id="B34">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gangadhar</surname> <given-names>C.</given-names></name> <name><surname>Moutteyan</surname> <given-names>M.</given-names></name> <name><surname>Vallabhuni</surname> <given-names>R. R.</given-names></name> <name><surname>Vijayan</surname> <given-names>V. P.</given-names></name> <name><surname>Sharma</surname> <given-names>N.</given-names></name> <name><surname>Theivadas</surname> <given-names>R.</given-names></name></person-group> (<year>2023</year>). <article-title>Analysis of optimization algorithms for stability and convergence for natural language processing using deep learning algorithms</article-title>. <source>Measurement: Sens</source>. <volume>27</volume>:<fpage>100784</fpage>. <pub-id pub-id-type="doi">10.1016/j.measen.2023.100784</pub-id></citation>
</ref>
<ref id="B35">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Garc&#x000ED;a-M&#x000E9;ndez</surname> <given-names>S.</given-names></name> <name><surname>de Arriba-P&#x000E9;rez</surname> <given-names>F.</given-names></name></person-group> (<year>2024</year>). <article-title>Large language models and healthcare alliance: potential and challenges of two representative use cases</article-title>. <source>Ann. Biomed. Eng</source>. <volume>52</volume>, <fpage>1928</fpage>&#x02013;<lpage>1931</lpage>. <pub-id pub-id-type="doi">10.1007/s10439-024-03454-8</pub-id><pub-id pub-id-type="pmid">38310159</pub-id></citation></ref>
<ref id="B36">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gou</surname> <given-names>J.</given-names></name> <name><surname>Yu</surname> <given-names>B.</given-names></name> <name><surname>Maybank</surname> <given-names>S. J.</given-names></name> <name><surname>Tao</surname> <given-names>D.</given-names></name></person-group> (<year>2021</year>). <article-title>Knowledge distillation: a survey</article-title>. <source>Int. J. Comput. Vis</source>. <volume>129</volume>, <fpage>1789</fpage>&#x02013;<lpage>1819</lpage>. <pub-id pub-id-type="doi">10.1007/s11263-021-01453-z</pub-id></citation>
</ref>
<ref id="B37">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Gu</surname> <given-names>A.</given-names></name> <name><surname>Gulcehre</surname> <given-names>C.</given-names></name> <name><surname>Paine</surname> <given-names>T.</given-names></name> <name><surname>Hoffman</surname> <given-names>M.</given-names></name> <name><surname>Pascanu</surname> <given-names>R.</given-names></name></person-group> (<year>2020</year>). <article-title>&#x0201C;Improving the gating mechanism of recurrent neural networks,&#x0201D;</article-title> in <source>Proceedings of the 37th International Conference on Machine Learning</source> (<publisher-loc>New York</publisher-loc>: <publisher-name>Proceedings of Machine Learning Research</publisher-name>), <fpage>3800</fpage>&#x02013;<lpage>3809</lpage>.</citation>
</ref>
<ref id="B38">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Gu</surname> <given-names>A.</given-names></name> <name><surname>Johnson</surname> <given-names>I.</given-names></name> <name><surname>Goel</surname> <given-names>K.</given-names></name> <name><surname>Saab</surname> <given-names>K.</given-names></name> <name><surname>Dao</surname> <given-names>T.</given-names></name> <name><surname>Rudra</surname> <given-names>A.</given-names></name> <etal/></person-group>. (<year>2021</year>). <article-title>&#x0201C;Combining recurrent, convolutional, and continuous-time models with linear state space layers,&#x0201D;</article-title> in <source>Advances in Neural Information Processing Systems</source>, eds. M. Ranzato, A. Beygelzimer, Y. Dauphin, P. Liang and J. W. Vaughan (<publisher-loc>New York</publisher-loc>: <publisher-name>Curran Associates, Inc</publisher-name>), <fpage>572</fpage>&#x02013;<lpage>585</lpage>.</citation>
</ref>
<ref id="B39">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Guo</surname> <given-names>Q.</given-names></name> <name><surname>Mu</surname> <given-names>Y.</given-names></name> <name><surname>Chen</surname> <given-names>J.</given-names></name> <name><surname>Wang</surname> <given-names>T.</given-names></name> <name><surname>Yu</surname> <given-names>Y.</given-names></name> <name><surname>Luo</surname> <given-names>P.</given-names></name></person-group> (<year>2022</year>). <article-title>&#x0201C;Scale-equivalent distillation for semi-supervised object detection,&#x0201D;</article-title> in <source>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)</source> (<publisher-loc>New Orleans, LA</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>14522</fpage>&#x02013;<lpage>14531</lpage>.</citation>
</ref>
<ref id="B40">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Han</surname> <given-names>S.</given-names></name> <name><surname>Mao</surname> <given-names>H.</given-names></name> <name><surname>Dally</surname> <given-names>W. J.</given-names></name></person-group> (<year>2015</year>). <article-title>Deep compression: compressing deep neural networks with pruning, trained quantization, and huffman coding</article-title>. <source>arXiv</source> [preprint] arXiv:1510.00149. <pub-id pub-id-type="doi">10.48550/arXiv.1510.00149</pub-id></citation>
</ref>
<ref id="B41">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Han</surname> <given-names>S.</given-names></name> <name><surname>Park</surname> <given-names>S.</given-names></name> <name><surname>Wu</surname> <given-names>F.</given-names></name> <name><surname>Kim</surname> <given-names>S.</given-names></name> <name><surname>Wu</surname> <given-names>C.</given-names></name> <name><surname>Xie</surname> <given-names>X.</given-names></name> <etal/></person-group>. (<year>2022</year>). <article-title>&#x0201C;FEDX: Unsupervised federated learning with cross knowledge distillation,&#x0201D;</article-title> in <source>Computer Vision-ECCV 2022</source>, S. Avidan, G. Brostow, M. Ciss&#x000E9;, G. Farinella, and T. Hassner (<publisher-loc>Cham</publisher-loc>: <publisher-name>Springer Nature Switzerland</publisher-name>), <fpage>691</fpage>&#x02013;<lpage>707</lpage>.</citation>
</ref>
<ref id="B42">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Heo</surname> <given-names>G.</given-names></name> <name><surname>Lee</surname> <given-names>S.</given-names></name> <name><surname>Cho</surname> <given-names>J.</given-names></name> <name><surname>Choi</surname> <given-names>H.</given-names></name> <name><surname>Lee</surname> <given-names>S.</given-names></name> <name><surname>Ham</surname> <given-names>H.</given-names></name> <etal/></person-group>. (<year>2024</year>). <article-title>&#x0201C;NeuPIMS: NPU-PIM heterogeneous acceleration for batched llm inferencing,&#x0201D;</article-title> in <source>Proceedings of the 29th ACM International Conference on Architectural Support for Programming Languages and Operating Systems</source> (<publisher-loc>New York</publisher-loc>: <publisher-name>ACM</publisher-name>), <fpage>722</fpage>&#x02013;<lpage>737</lpage>.</citation>
</ref>
<ref id="B43">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hu</surname> <given-names>J.</given-names></name> <name><surname>Lin</surname> <given-names>P.</given-names></name> <name><surname>Zhang</surname> <given-names>H.</given-names></name> <name><surname>Lan</surname> <given-names>Z.</given-names></name> <name><surname>Chen</surname> <given-names>W.</given-names></name> <name><surname>Xie</surname> <given-names>K.</given-names></name> <etal/></person-group>. (<year>2023</year>). <article-title>A dynamic pruning method on multiple sparse structures in deep neural networks</article-title>. <source>IEEE Access</source> <volume>11</volume>, <fpage>38448</fpage>&#x02013;<lpage>38457</lpage>. <pub-id pub-id-type="doi">10.1109/ACCESS.2023.3267469</pub-id></citation>
</ref>
<ref id="B44">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Iftikhar</surname> <given-names>S.</given-names></name> <name><surname>Gill</surname> <given-names>S. S.</given-names></name> <name><surname>Song</surname> <given-names>C.</given-names></name> <name><surname>Xu</surname> <given-names>M.</given-names></name> <name><surname>Aslanpour</surname> <given-names>M. S.</given-names></name> <name><surname>Toosi</surname> <given-names>A. N.</given-names></name> <etal/></person-group>. (<year>2023</year>). <article-title>Ai-based fog and edge computing: a systematic review, taxonomy and future directions</article-title>. <source>Intern. Things</source> <volume>21</volume>:<fpage>100674</fpage>. <pub-id pub-id-type="doi">10.1016/j.iot.2022.100674</pub-id></citation>
</ref>
<ref id="B45">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Ignatov</surname> <given-names>A.</given-names></name> <name><surname>Timofte</surname> <given-names>R.</given-names></name> <name><surname>Van Gool</surname> <given-names>L.</given-names></name></person-group> (<year>2018</year>). <article-title>&#x0201C;AI benchmark: running deep neural networks on android smartphones,&#x0201D;</article-title> in <source>Proceedings of the European Conference on Computer Vision (ECCV) Workshops</source> (<publisher-loc>Munich</publisher-loc>). <pub-id pub-id-type="doi">10.1007/978-3-030-11021-5_19</pub-id></citation>
</ref>
<ref id="B46">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Iqbal</surname> <given-names>F.</given-names></name> <name><surname>Altaf</surname> <given-names>A.</given-names></name> <name><surname>Waris</surname> <given-names>Z.</given-names></name> <name><surname>Aray</surname> <given-names>D. G.</given-names></name> <name><surname>Flores</surname> <given-names>M. A. L.</given-names></name> <name><surname>D&#x000ED;ez</surname> <given-names>I.</given-names></name> <etal/></person-group>. (<year>2023</year>). <article-title>Blockchain-modeled edge-computing-based smart home monitoring system with energy usage prediction</article-title>. <source>Sensors</source> <volume>23</volume>:<fpage>5263</fpage>. <pub-id pub-id-type="doi">10.3390/s23115263</pub-id><pub-id pub-id-type="pmid">37299993</pub-id></citation></ref>
<ref id="B47">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Jacob</surname> <given-names>B.</given-names></name> <name><surname>Kligys</surname> <given-names>S.</given-names></name> <name><surname>Chen</surname> <given-names>B.</given-names></name> <name><surname>Zhu</surname> <given-names>M.</given-names></name> <name><surname>Tang</surname> <given-names>M.</given-names></name> <name><surname>Howard</surname> <given-names>A.</given-names></name> <etal/></person-group>. (<year>2018</year>). <article-title>&#x0201C;Quantization and training of neural networks for efficient integer-arithmetic-only inference,&#x0201D;</article-title> in <source>Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</source> (<publisher-loc>Salt Lake City, UT</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>2704</fpage>&#x02013;<lpage>2713</lpage>. <pub-id pub-id-type="doi">10.1109/CVPR.2018.00286</pub-id></citation>
</ref>
<ref id="B48">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Jan</surname> <given-names>M. A.</given-names></name> <name><surname>He</surname> <given-names>X.</given-names></name> <name><surname>Song</surname> <given-names>H.</given-names></name> <name><surname>Babar</surname> <given-names>M.</given-names></name></person-group> (<year>2021</year>). <article-title>Machine learning and big data analytics for iot-enabled smart cities</article-title>. <source>Mobile Netw. Appl</source>. <volume>26</volume>, <fpage>156</fpage>&#x02013;<lpage>158</lpage>. <pub-id pub-id-type="doi">10.1007/s11036-020-01702-4</pub-id><pub-id pub-id-type="pmid">39943545</pub-id></citation></ref>
<ref id="B49">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Jang</surname> <given-names>J.-W.</given-names></name> <name><surname>Lee</surname> <given-names>S.</given-names></name> <name><surname>Kim</surname> <given-names>D.</given-names></name> <name><surname>Park</surname> <given-names>H.</given-names></name> <name><surname>Ardestani</surname> <given-names>A. S.</given-names></name> <name><surname>Choi</surname> <given-names>Y.</given-names></name> <etal/></person-group>. (<year>2021</year>). <article-title>&#x0201C;Sparsity-aware and re-configurable NPU architecture for samsung flagship mobile SoC,&#x0201D;</article-title> in <source>2021 ACM/IEEE 48th Annual International Symposium on Computer Architecture (ISCA)</source> (<publisher-loc>Valencia</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>15</fpage>&#x02013;<lpage>28</lpage>.</citation>
</ref>
<ref id="B50">
<citation citation-type="web"><person-group person-group-type="author"><name><surname>Jaszczur</surname> <given-names>S.</given-names></name> <name><surname>Chowdhery</surname> <given-names>A.</given-names></name> <name><surname>Mohiuddin</surname> <given-names>A.</given-names></name> <name><surname>Kaiser</surname> <given-names>L.</given-names></name> <name><surname>Gajewski</surname> <given-names>W.</given-names></name> <name><surname>Michalewski</surname> <given-names>H.</given-names></name> <etal/></person-group>. (<year>2021</year>). <article-title>Sparse is enough in scaling transformers</article-title>. <source>Adv. Neural Inf. Process. Syst</source>. <volume>34</volume>, <fpage>9895</fpage>&#x02013;<lpage>9907</lpage>. Available online at: <ext-link ext-link-type="uri" xlink:href="https://proceedings.neurips.cc/paper/2021/hash/51f15efdd170e6043fa02a74882f0470-Abstract.html">https://proceedings.neurips.cc/paper/2021/hash/51f15efdd170e6043fa02a74882f0470-Abstract.html</ext-link></citation>
</ref>
<ref id="B51">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ji</surname> <given-names>J.</given-names></name> <name><surname>Shu</surname> <given-names>Z.</given-names></name> <name><surname>Li</surname> <given-names>H.</given-names></name> <name><surname>Lai</surname> <given-names>K. X.</given-names></name> <name><surname>Lu</surname> <given-names>M.</given-names></name> <name><surname>Jiang</surname> <given-names>G.</given-names></name> <etal/></person-group>. (<year>2024</year>). <article-title>Edge-computing-based knowledge distillation and multitask learning for partial discharge recognition</article-title>. <source>IEEE Trans. Instrum. Meas</source>. <volume>73</volume>, <fpage>1</fpage>&#x02013;<lpage>11</lpage>. <pub-id pub-id-type="doi">10.1109/TIM.2024.3351239</pub-id></citation>
</ref>
<ref id="B52">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Jo</surname> <given-names>E.</given-names></name> <name><surname>Epstein</surname> <given-names>D. A.</given-names></name> <name><surname>Jung</surname> <given-names>H.</given-names></name> <name><surname>Kim</surname> <given-names>Y.-H.</given-names></name></person-group> (<year>2023</year>). <article-title>&#x0201C;Understanding the benefits and challenges of deploying conversational ai leveraging large language models for public health intervention,&#x0201D;</article-title> in <source>Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems</source> (<publisher-loc>New York, NY</publisher-loc>: <publisher-name>Association for Computing Machinery</publisher-name>).</citation>
</ref>
<ref id="B53">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kairouz</surname> <given-names>P.</given-names></name> <name><surname>McMahan</surname> <given-names>H. B.</given-names></name></person-group> (<year>2021</year>). <article-title>Advances and open problems in federated learning</article-title>. <source>Found. Trends Mach. Learn</source>. <volume>14</volume>, <fpage>1</fpage>&#x02013;<lpage>210</lpage>. <pub-id pub-id-type="doi">10.1561/9781680837896</pub-id></citation>
</ref>
<ref id="B54">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Kang</surname> <given-names>Y.</given-names></name> <name><surname>Hauswald</surname> <given-names>J.</given-names></name> <name><surname>Gao</surname> <given-names>C.</given-names></name> <name><surname>Rovinski</surname> <given-names>A.</given-names></name> <name><surname>Mudge</surname> <given-names>T. N.</given-names></name> <name><surname>Mars</surname> <given-names>J.</given-names></name> <etal/></person-group>. (<year>2017</year>). <article-title>&#x0201C;Neurosurgeon: Collaborative intelligence between the cloud and mobile edge,&#x0201D;</article-title> in <source>ACM SIGARCH Computer Architecture News</source> (<publisher-loc>New York</publisher-loc>: <publisher-name>ACM</publisher-name>), <fpage>615</fpage>&#x02013;<lpage>629</lpage>.</citation>
</ref>
<ref id="B55">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Ke</surname> <given-names>Z.</given-names></name> <name><surname>Zhou</surname> <given-names>X.</given-names></name> <name><surname>Jiang</surname> <given-names>D.</given-names></name> <name><surname>Yan</surname> <given-names>H.</given-names></name> <name><surname>Qiu</surname> <given-names>T.</given-names></name></person-group> (<year>2023</year>). <article-title>&#x0201C;CollabVr: Reprojection-based edge-client collaborative rendering for real-time high-quality mobile virtual reality,&#x0201D;</article-title> in <source>2023 IEEE Real-Time Systems Symposium (RTSS)</source> (<publisher-loc>Taipei</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>304</fpage>&#x02013;<lpage>316</lpage>.</citation>
</ref>
<ref id="B56">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Khan</surname> <given-names>L. U.</given-names></name> <name><surname>Yaqoob</surname> <given-names>I.</given-names></name> <name><surname>Tran</surname> <given-names>N. H.</given-names></name> <name><surname>Kazmi</surname> <given-names>S. M. A.</given-names></name> <name><surname>Dang</surname> <given-names>T. N.</given-names></name> <name><surname>Hong</surname> <given-names>C. S.</given-names></name></person-group> (<year>2020</year>). <article-title>Edge-computing-enabled smart cities: a comprehensive survey</article-title>. <source>IEEE Internet of Things Journal</source> <volume>7</volume>, <fpage>10200</fpage>&#x02013;<lpage>10232</lpage>. <pub-id pub-id-type="doi">10.1109/JIOT.2020.2987070</pub-id></citation>
</ref>
<ref id="B57">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Khan</surname> <given-names>W. Z.</given-names></name> <name><surname>Ahmed</surname> <given-names>E.</given-names></name> <name><surname>Hakak</surname> <given-names>S.</given-names></name> <name><surname>Yaqoob</surname> <given-names>I.</given-names></name> <name><surname>Ahmed</surname> <given-names>A.</given-names></name></person-group> (<year>2019</year>). <article-title>Edge computing: A survey</article-title>. <source>Future Generation Computer Systems</source> <volume>97</volume>:<fpage>219</fpage>&#x02013;<lpage>235</lpage>. <pub-id pub-id-type="doi">10.1016/j.future.2019.02.050</pub-id></citation>
</ref>
<ref id="B58">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kim</surname> <given-names>J.</given-names></name></person-group> (<year>2023</year>). <article-title>Quantization robust pruning with knowledge distillation</article-title>. <source>IEEE Access</source> <volume>11</volume>, <fpage>26419</fpage>&#x02013;<lpage>26426</lpage>. <pub-id pub-id-type="doi">10.1109/ACCESS.2023.3257864</pub-id><pub-id pub-id-type="pmid">35832228</pub-id></citation></ref>
<ref id="B59">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Kim</surname> <given-names>J.</given-names></name> <name><surname>Lee</surname> <given-names>J. H.</given-names></name> <name><surname>Kim</surname> <given-names>S.</given-names></name> <name><surname>Park</surname> <given-names>J.</given-names></name> <name><surname>Yoo</surname> <given-names>K. M.</given-names></name> <name><surname>Kwon</surname> <given-names>S. J.</given-names></name> <etal/></person-group>. (<year>2023</year>). <article-title>&#x0201C;Memory-efficient fine-tuning of compressed large language models via sub-4-bit integer quantization,&#x0201D;</article-title> in <source>Advances in Neural Information Processing Systems</source>, eds. A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine (<publisher-loc>New York</publisher-loc>: <publisher-name>Curran Associates, Inc</publisher-name>), <fpage>36187</fpage>&#x02013;<lpage>36207</lpage>.</citation>
</ref>
<ref id="B60">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Kim</surname> <given-names>S. Y.</given-names></name> <name><surname>Lee</surname> <given-names>J.</given-names></name> <name><surname>Kim</surname> <given-names>C. H.</given-names></name> <name><surname>Lee</surname> <given-names>W. J.</given-names></name> <name><surname>Kim</surname> <given-names>S. W.</given-names></name></person-group> (<year>2022</year>). <article-title>&#x0201C;Extending the onnx runtime framework for the processing-in-memory execution,&#x0201D;</article-title> in <source>2022 International Conference on Electronics, Information, and Communication (ICEIC)</source> (<publisher-loc>Jeju</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>1</fpage>&#x02013;<lpage>4</lpage>. <pub-id pub-id-type="doi">10.1109/ICEIC54506.2022.9748444</pub-id></citation>
</ref>
<ref id="B61">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kim</surname> <given-names>Y.</given-names></name> <name><surname>Xu</surname> <given-names>X.</given-names></name> <name><surname>McDuff</surname> <given-names>D.</given-names></name> <name><surname>Breazeal</surname> <given-names>C.</given-names></name> <name><surname>Park</surname> <given-names>H. W.</given-names></name></person-group> (<year>2024</year>). <article-title>Health-llm: Large language models for health prediction via wearable sensor data</article-title>. <source>arXiv [preprint] arXiv</source>:2401.06866.</citation>
</ref>
<ref id="B62">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>K&#x000F6;hl</surname> <given-names>M. A.</given-names></name> <name><surname>Hermanns</surname> <given-names>H.</given-names></name></person-group> (<year>2023</year>). <article-title>Model-based diagnosis of real-time systems: Robustness against varying latency, clock drift, and out-of-order observations</article-title>. <source>ACM Trans. Embed. Comput. Syst</source>. <volume>22</volume>:<fpage>1</fpage>&#x02013;<lpage>48</lpage>. <pub-id pub-id-type="doi">10.1145/3597209</pub-id></citation>
</ref>
<ref id="B63">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kong</surname> <given-names>L.</given-names></name> <name><surname>Tan</surname> <given-names>J.</given-names></name> <name><surname>Huang</surname> <given-names>J.</given-names></name> <name><surname>Chen</surname> <given-names>G.</given-names></name> <name><surname>Wang</surname> <given-names>S.</given-names></name> <name><surname>Jin</surname> <given-names>X.</given-names></name> <etal/></person-group>. (<year>2022</year>). <article-title>Edge-computing-driven internet of things: A survey</article-title>. <source>ACM Computing Surveys</source> 55(8):1&#x02013;41. <pub-id pub-id-type="doi">10.1145/3555308</pub-id></citation>
</ref>
<ref id="B64">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Krestinskaya</surname> <given-names>O.</given-names></name> <name><surname>James</surname> <given-names>A. P.</given-names></name> <name><surname>Chua</surname> <given-names>L.</given-names></name></person-group> (<year>2019</year>). <article-title>Neuromemristive circuits for edge computing: a review</article-title>. <source>IEEE Trans. Neural Netw. Learn. Syst</source>. <volume>31</volume>, <fpage>4</fpage>&#x02013;<lpage>23</lpage>. <pub-id pub-id-type="doi">10.1109/TNNLS.2019.2899262</pub-id><pub-id pub-id-type="pmid">30892238</pub-id></citation></ref>
<ref id="B65">
<citation citation-type="web"><person-group person-group-type="author"><name><surname>Lan</surname> <given-names>Z.</given-names></name> <name><surname>Chen</surname> <given-names>M.</given-names></name> <name><surname>Goodman</surname> <given-names>S.</given-names></name> <name><surname>Gimpel</surname> <given-names>K.</given-names></name> <name><surname>Sharma</surname> <given-names>P.</given-names></name> <name><surname>Soricut</surname> <given-names>R.</given-names></name></person-group> (<year>2019</year>). <article-title>&#x0201C;Albert: a lite bert for self-supervised learning of language representations,&#x0201D;</article-title> in <source>Proceedings of the 8th International Conference on Learning Representations (ICLR)</source> (<publisher-loc>Addis Ababa</publisher-loc>: <publisher-name>openreview.net</publisher-name>). Available online at: <ext-link ext-link-type="uri" xlink:href="https://openreview.net/forum?id=H1eA7AEtvS">https://openreview.net/forum?id=H1eA7AEtvS</ext-link></citation>
</ref>
<ref id="B66">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Li</surname> <given-names>B.</given-names></name> <name><surname>Sainath</surname> <given-names>T. N.</given-names></name> <name><surname>Pang</surname> <given-names>R.</given-names></name> <name><surname>Wu</surname> <given-names>Z.</given-names></name></person-group> (<year>2019</year>). <article-title>&#x0201C;Semi-supervised training for end-to-end models via weak distillation,&#x0201D;</article-title> in <source>ICASSP 2019</source> - <italic>2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)</italic> (<publisher-loc>Brighton</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>837</fpage>&#x02013;<lpage>2841</lpage>.</citation>
</ref>
<ref id="B67">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Li</surname> <given-names>D.</given-names></name> <name><surname>Lai</surname> <given-names>J.</given-names></name> <name><surname>Wang</surname> <given-names>R.</given-names></name> <name><surname>Li</surname> <given-names>X.</given-names></name> <name><surname>Vijayakumar</surname> <given-names>P.</given-names></name> <name><surname>Gupta</surname> <given-names>B. B.</given-names></name> <etal/></person-group>. (<year>2023</year>). <article-title>Ubiquitous intelligent federated learning privacy-preserving scheme under edge computing</article-title>. <source>Future Generat. Comp. Syst</source>. <volume>144</volume>, <fpage>205</fpage>&#x02013;<lpage>218</lpage>. <pub-id pub-id-type="doi">10.1016/j.future.2023.03.010</pub-id></citation>
</ref>
<ref id="B68">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Li</surname> <given-names>J.</given-names></name> <name><surname>Liang</surname> <given-names>W.</given-names></name> <name><surname>Xu</surname> <given-names>W.</given-names></name> <name><surname>Xu</surname> <given-names>Z.</given-names></name> <name><surname>Li</surname> <given-names>Y.</given-names></name> <name><surname>Jia</surname> <given-names>X.</given-names></name></person-group> (<year>2023</year>). <article-title>Service home identification of multiple-source iot applications in edge computing</article-title>. <source>IEEE Trans. Serv. Comp</source>. <volume>16</volume>, <fpage>1417</fpage>&#x02013;<lpage>1430</lpage>. <pub-id pub-id-type="doi">10.1109/TSC.2022.3176576</pub-id></citation>
</ref>
<ref id="B69">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Li</surname> <given-names>J.-Y.</given-names></name> <name><surname>Du</surname> <given-names>K.-J.</given-names></name> <name><surname>Zhan</surname> <given-names>Z.-H.</given-names></name> <name><surname>Wang</surname> <given-names>H.</given-names></name> <name><surname>Zhang</surname> <given-names>J.</given-names></name></person-group> (<year>2023</year>). <article-title>Distributed differential evolution with adaptive resource allocation</article-title>. <source>IEEE Trans. Cybern</source>. 53(5):2791&#x02013;2804. <pub-id pub-id-type="doi">10.1109/TCYB.2022.3153964</pub-id><pub-id pub-id-type="pmid">35286273</pub-id></citation></ref>
<ref id="B70">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Li</surname> <given-names>M.</given-names></name> <name><surname>Zhao</surname> <given-names>M.</given-names></name> <name><surname>Luo</surname> <given-names>T.</given-names></name> <name><surname>Yang</surname> <given-names>Y.</given-names></name> <name><surname>Peng</surname> <given-names>S.-L.</given-names></name></person-group> (<year>2022</year>). <article-title>A compact parallel pruning scheme for deep learning model and its mobile instrument deployment</article-title>. <source>Mathematics</source> <volume>10</volume>:<fpage>2126</fpage>. <pub-id pub-id-type="doi">10.3390/math10122126</pub-id></citation>
</ref>
<ref id="B71">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Li</surname> <given-names>X.</given-names></name> <name><surname>Liu</surname> <given-names>H.</given-names></name> <name><surname>Wang</surname> <given-names>W.</given-names></name> <name><surname>Zheng</surname> <given-names>Y.</given-names></name> <name><surname>Lv</surname> <given-names>H.</given-names></name> <name><surname>Lv</surname> <given-names>Z.</given-names></name></person-group> (<year>2022</year>). <article-title>Big data analysis of the internet of things in the digital twins of smart city based on deep learning</article-title>. <source>Future Generat. Comp. Syst</source>. <volume>128</volume>, <fpage>167</fpage>&#x02013;<lpage>177</lpage>. <pub-id pub-id-type="doi">10.1016/j.future.2021.10.006</pub-id></citation>
</ref>
<ref id="B72">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Liang</surname> <given-names>T.</given-names></name> <name><surname>Glossner</surname> <given-names>J.</given-names></name> <name><surname>Wang</surname> <given-names>L.</given-names></name> <name><surname>Shi</surname> <given-names>S.</given-names></name> <name><surname>Zhang</surname> <given-names>X.</given-names></name></person-group> (<year>2021</year>). <article-title>Pruning and quantization for deep neural network acceleration: a survey</article-title>. <source>Neurocomputing</source> <volume>461</volume>, <fpage>370</fpage>&#x02013;<lpage>403</lpage>. <pub-id pub-id-type="doi">10.1016/j.neucom.2021.07.045</pub-id></citation>
</ref>
<ref id="B73">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Liang</surname> <given-names>Z.</given-names></name> <name><surname>Cheng</surname> <given-names>J.</given-names></name> <name><surname>Yang</surname> <given-names>R.</given-names></name> <name><surname>Ren</surname> <given-names>H.</given-names></name> <name><surname>Song</surname> <given-names>Z.</given-names></name> <name><surname>Wu</surname> <given-names>D.</given-names></name> <etal/></person-group>. (<year>2023</year>). <article-title>Unleashing the potential of llms for quantum computing: A study in quantum architecture design</article-title>. <source>arXiv</source> [preprint] arXiv:2307.08191. <pub-id pub-id-type="doi">10.48550/arXiv.2307.08191</pub-id></citation>
</ref>
<ref id="B74">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lin</surname> <given-names>S.-C.</given-names></name> <name><surname>Zhang</surname> <given-names>Y.</given-names></name> <name><surname>Hsu</surname> <given-names>C.-H.</given-names></name> <name><surname>Skach</surname> <given-names>M.</given-names></name> <name><surname>Haque</surname> <given-names>M. E.</given-names></name> <name><surname>Tang</surname> <given-names>L.</given-names></name> <etal/></person-group>. (<year>2018</year>). <article-title>The architectural implications of autonomous driving: Constraints and acceleration</article-title>. <source>SIGPLAN Not</source>., <volume>53</volume>, <fpage>751</fpage>&#x02013;<lpage>766</lpage>. <pub-id pub-id-type="doi">10.1145/3296957.3173191</pub-id></citation>
</ref>
<ref id="B75">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Liu</surname> <given-names>G.</given-names></name> <name><surname>Wang</surname> <given-names>C.</given-names></name> <name><surname>Ma</surname> <given-names>X.</given-names></name> <name><surname>Yang</surname> <given-names>Y.</given-names></name></person-group> (<year>2021</year>). <article-title>Keep your data locally: Federated-learning-based data privacy preservation in edge computing</article-title>. <source>IEEE Netw</source>. <volume>35</volume>, <fpage>60</fpage>&#x02013;<lpage>66</lpage>. <pub-id pub-id-type="doi">10.1109/MNET.011.2000215</pub-id></citation>
</ref>
<ref id="B76">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Liu</surname> <given-names>P.</given-names></name> <name><surname>Zhang</surname> <given-names>Z.</given-names></name> <name><surname>Meng</surname> <given-names>Z.</given-names></name> <name><surname>Gao</surname> <given-names>N.</given-names></name></person-group> (<year>2021</year>). <article-title>Monocular depth estimation with joint attention feature distillation and wavelet-based loss function</article-title>. <source>Sensors</source> <volume>21</volume>:<fpage>54</fpage>. <pub-id pub-id-type="doi">10.3390/s21010054</pub-id><pub-id pub-id-type="pmid">33374278</pub-id></citation></ref>
<ref id="B77">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Liu</surname> <given-names>Q.</given-names></name> <name><surname>Gu</surname> <given-names>J.</given-names></name> <name><surname>Yang</surname> <given-names>J.</given-names></name> <name><surname>Li</surname> <given-names>Y.</given-names></name> <name><surname>Sha</surname> <given-names>D.</given-names></name> <name><surname>Xu</surname> <given-names>M.</given-names></name> <etal/></person-group>. (<year>2021</year>). <source>Cloud, Edge, and Mobile Computing for Smart Cities</source>. <publisher-loc>Singapore</publisher-loc>: <publisher-name>Springer Singapore</publisher-name>, <fpage>757</fpage>&#x02013;<lpage>795</lpage>.</citation>
</ref>
<ref id="B78">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Liu</surname> <given-names>X.</given-names></name> <name><surname>Ye</surname> <given-names>M.</given-names></name> <name><surname>Zhou</surname> <given-names>D.</given-names></name> <name><surname>Liu</surname> <given-names>Q.</given-names></name></person-group> (<year>2021</year>). <article-title>Post-training quantization with multiple points: Mixed precision without mixed precision</article-title>. <source>Proc. AAAI Conf. Artif. Intellig</source>. <volume>35</volume>, <fpage>8697</fpage>&#x02013;<lpage>8705</lpage>. <pub-id pub-id-type="doi">10.1609/aaai.v35i10.17054</pub-id></citation>
</ref>
<ref id="B79">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Liu</surname> <given-names>Y.</given-names></name> <name><surname>Zhang</surname> <given-names>W.</given-names></name> <name><surname>Wang</surname> <given-names>J.</given-names></name></person-group> (<year>2020</year>). <article-title>Adaptive multi-teacher multi-level knowledge distillation</article-title>. <source>Neurocomputing</source> <volume>415</volume>, <fpage>106</fpage>&#x02013;<lpage>113</lpage>. <pub-id pub-id-type="doi">10.1016/j.neucom.2020.07.048</pub-id><pub-id pub-id-type="pmid">37163850</pub-id></citation></ref>
<ref id="B80">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Liu</surname> <given-names>Z.</given-names></name> <name><surname>Xu</surname> <given-names>J.</given-names></name> <name><surname>Peng</surname> <given-names>X.</given-names></name> <name><surname>Xiong</surname> <given-names>R.</given-names></name></person-group> (<year>2018</year>). <article-title>&#x0201C;Frequency-domain dynamic pruning for convolutional neural networks,&#x0201D;</article-title> in <source>Advances in Neural Information Processing Systems</source>, eds. S. Bengio, H. Wallach, H. Larochelle, K. Grauman, N. Cesa-Bianchi, and R. Garnett (<publisher-loc>New York</publisher-loc>: <publisher-name>Curran Associates, Inc</publisher-name>).</citation>
</ref>
<ref id="B81">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lu</surname> <given-names>S.</given-names></name> <name><surname>Lu</surname> <given-names>J.</given-names></name> <name><surname>An</surname> <given-names>K.</given-names></name> <name><surname>Wang</surname> <given-names>X.</given-names></name> <name><surname>He</surname> <given-names>Q.</given-names></name></person-group> (<year>2023</year>). <article-title>Edge computing on iot for machine signal processing and fault diagnosis: a review</article-title>. <source>IEEE Intern. Things J</source>. <volume>10</volume>, <fpage>11093</fpage>&#x02013;<lpage>11116</lpage>. <pub-id pub-id-type="doi">10.1109/JIOT.2023.3239944</pub-id></citation>
</ref>
<ref id="B82">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lv</surname> <given-names>Z.</given-names></name> <name><surname>Chen</surname> <given-names>D.</given-names></name> <name><surname>Lou</surname> <given-names>R.</given-names></name> <name><surname>Wang</surname> <given-names>Q.</given-names></name></person-group> (<year>2021</year>). <article-title>Retracted: Intelligent edge computing based on machine learning for smart city</article-title>. <source>Future Generat. Comp. Syst</source>. <volume>115</volume>, <fpage>90</fpage>&#x02013;<lpage>99</lpage>. <pub-id pub-id-type="doi">10.1016/j.future.2020.08.037</pub-id></citation>
</ref>
<ref id="B83">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Matsubara</surname> <given-names>Y.</given-names></name> <name><surname>Callegaro</surname> <given-names>D.</given-names></name> <name><surname>Baidya</surname> <given-names>S.</given-names></name> <name><surname>Levorato</surname> <given-names>M.</given-names></name> <name><surname>Singh</surname> <given-names>S.</given-names></name></person-group> (<year>2020</year>). <article-title>Head network distillation: splitting distilled deep neural networks for resource-constrained edge computing systems</article-title>. <source>IEEE Access</source> <volume>8</volume>, <fpage>212177</fpage>&#x02013;<lpage>212193</lpage>. <pub-id pub-id-type="doi">10.1109/ACCESS.2020.3039714</pub-id></citation>
</ref>
<ref id="B84">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>McMahan</surname> <given-names>H. B.</given-names></name> <name><surname>Moore</surname> <given-names>E.</given-names></name> <name><surname>Ramage</surname> <given-names>D.</given-names></name> <name><surname>Hampson</surname> <given-names>S.</given-names></name> <name><surname>Aguera y Arcas</surname> <given-names>B.</given-names></name></person-group> (<year>2017</year>). <article-title>&#x0201C;Communication-efficient learning of deep networks from decentralized data,&#x0201D;</article-title> in <source>Proceedings of the 20th International Conference on Artificial Intelligence and Statistics (AISTATS)</source> (<publisher-loc>New York</publisher-loc>: <publisher-name>PMLR</publisher-name>), <fpage>1273</fpage>&#x02013;<lpage>1282</lpage>.<pub-id pub-id-type="pmid">38091762</pub-id></citation></ref>
<ref id="B85">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Memari</surname> <given-names>P.</given-names></name> <name><surname>Mohammadi</surname> <given-names>S. S.</given-names></name> <name><surname>Jolai</surname> <given-names>F.</given-names></name> <name><surname>Tavakkoli-Moghaddam</surname> <given-names>R.</given-names></name></person-group> (<year>2022</year>). <article-title>A latency-aware task scheduling algorithm for allocating virtual machines in a cost-effective and time-sensitive fog-cloud architecture</article-title>. <source>J. Supercomput</source>. <volume>78</volume>, <fpage>93</fpage>&#x02013;<lpage>122</lpage>. <pub-id pub-id-type="doi">10.1007/s11227-021-03868-4</pub-id></citation>
</ref>
<ref id="B86">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Nagel</surname> <given-names>M.</given-names></name> <name><surname>Fournarakis</surname> <given-names>M.</given-names></name> <name><surname>Bondarenko</surname> <given-names>Y.</given-names></name> <name><surname>Blankevoort</surname> <given-names>T.</given-names></name></person-group> (<year>2022</year>). <article-title>&#x0201C;Overcoming oscillations in quantization-aware training,&#x0201D;</article-title> in <source>Proceedings of the 39th International Conference on Machine Learning</source>, eds. K. Chaudhuri, S. Jegelka, L. Song, C. Szepesvari, G. Niu, and S. Sabato (<publisher-loc>New York</publisher-loc>: <publisher-name>PMLR</publisher-name>), <fpage>16318</fpage>&#x02013;<lpage>16330</lpage>.</citation>
</ref>
<ref id="B87">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Nain</surname> <given-names>G.</given-names></name> <name><surname>Pattanaik</surname> <given-names>K.</given-names></name> <name><surname>Sharma</surname> <given-names>G.</given-names></name></person-group> (<year>2022</year>). <article-title>Towards edge computing in intelligent manufacturing: Past, present and future</article-title>. <source>J. Manufact. Syst</source>. <volume>62</volume>, <fpage>588</fpage>&#x02013;<lpage>611</lpage>. <pub-id pub-id-type="doi">10.1016/j.jmsy.2022.01.010</pub-id></citation>
</ref>
<ref id="B88">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Neseem</surname> <given-names>M.</given-names></name> <name><surname>Agiza</surname> <given-names>A.</given-names></name> <name><surname>Reda</surname> <given-names>S.</given-names></name></person-group> (<year>2023</year>). <article-title>&#x0201C;ADAMTL: Adaptive input-dependent inference for efficient multi-task learning,&#x0201D;</article-title> in <source>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition</source> (<publisher-loc>Vancouver, BC</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>4730</fpage>&#x02013;<lpage>4739</lpage>.</citation>
</ref>
<ref id="B89">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Nguyen</surname> <given-names>D. C.</given-names></name> <name><surname>Ding</surname> <given-names>M.</given-names></name> <name><surname>Pathirana</surname> <given-names>P. N.</given-names></name> <name><surname>Seneviratne</surname> <given-names>A.</given-names></name> <name><surname>Li</surname> <given-names>J.</given-names></name> <name><surname>Poor</surname> <given-names>H. V.</given-names></name></person-group> (<year>2021a</year>). <article-title>Federated learning for internet of things: A comprehensive survey</article-title>. <source>IEEE Commun. Surv. Tutor</source>. <volume>23</volume>, <fpage>1622</fpage>&#x02013;<lpage>1658</lpage>. <pub-id pub-id-type="doi">10.1109/COMST.2021.3075439</pub-id></citation>
</ref>
<ref id="B90">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Nguyen</surname> <given-names>D. C.</given-names></name> <name><surname>Ding</surname> <given-names>M.</given-names></name> <name><surname>Pham</surname> <given-names>Q.-V.</given-names></name> <name><surname>Pathirana</surname> <given-names>P. N.</given-names></name> <name><surname>Le</surname> <given-names>L. B.</given-names></name> <name><surname>Seneviratne</surname> <given-names>A.</given-names></name> <etal/></person-group>. (<year>2021b</year>). <article-title>Federated learning meets blockchain in edge computing: Opportunities and challenges</article-title>. <source>IEEE Intern, Things J</source>. <volume>8</volume>, <fpage>12806</fpage>&#x02013;<lpage>12825</lpage>. <pub-id pub-id-type="doi">10.1109/JIOT.2021.3072611</pub-id></citation>
</ref>
<ref id="B91">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Niu</surname> <given-names>W.</given-names></name> <name><surname>Guan</surname> <given-names>J.</given-names></name> <name><surname>Wang</surname> <given-names>Y.</given-names></name> <name><surname>Agrawal</surname> <given-names>G.</given-names></name> <name><surname>Ren</surname> <given-names>B.</given-names></name></person-group> (<year>2021</year>). <article-title>&#x0201C;DNNFusion: accelerating deep neural networks execution with advanced operator fusion,&#x0201D;</article-title> in <source>Proceedings of the 42nd ACM SIGPLAN International Conference on Programming Language Design and Implementation, PLDI 2021</source> (<publisher-loc>New York, NY</publisher-loc>: <publisher-name>Association for Computing Machinery</publisher-name>), <fpage>883</fpage>&#x02013;<lpage>898</lpage>.</citation>
</ref>
<ref id="B92">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Pandey</surname> <given-names>J.</given-names></name> <name><surname>Asati</surname> <given-names>A. R.</given-names></name></person-group> (<year>2023</year>). <article-title>Lightweight convolutional neural network architecture implementation using TensorFlow lite</article-title>. <source>Int. J. Inform. Technol</source>. <volume>15</volume>, <fpage>2489</fpage>&#x02013;<lpage>2498</lpage>. <pub-id pub-id-type="doi">10.1007/s41870-023-01320-9</pub-id></citation>
</ref>
<ref id="B93">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Pang</surname> <given-names>G.</given-names></name> <name><surname>Shen</surname> <given-names>C.</given-names></name> <name><surname>Cao</surname> <given-names>L.</given-names></name> <name><surname>Hengel</surname> <given-names>A. V. D.</given-names></name></person-group> (<year>2021</year>). <article-title>Deep learning for anomaly detection: A review</article-title>. <source>ACM Comp. Surv</source>. (<source>CSUR)</source> <volume>54</volume>, <fpage>1</fpage>&#x02013;<lpage>38</lpage>. <pub-id pub-id-type="doi">10.1145/3439950</pub-id></citation>
</ref>
<ref id="B94">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Pelle</surname> <given-names>I.</given-names></name> <name><surname>Czentye</surname> <given-names>J.</given-names></name> <name><surname>D&#x000F3;ka</surname> <given-names>J.</given-names></name> <name><surname>Kern</surname> <given-names>A.</given-names></name> <name><surname>Ger&#x00151;</surname> <given-names>B. P.</given-names></name> <name><surname>Sonkoly</surname> <given-names>B.</given-names></name></person-group> (<year>2021</year>). <article-title>Operating latency sensitive applications on public serverless edge cloud platforms</article-title>. <source>IEEE Intern. Things J</source>. <volume>8</volume>, <fpage>7954</fpage>&#x02013;<lpage>7972</lpage>. <pub-id pub-id-type="doi">10.1109/JIOT.2020.3042428</pub-id></citation>
</ref>
<ref id="B95">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Phan</surname> <given-names>N.</given-names></name> <name><surname>Wang</surname> <given-names>Y.</given-names></name> <name><surname>Wu</surname> <given-names>X.</given-names></name> <name><surname>Dou</surname> <given-names>D.</given-names></name></person-group> (<year>2017</year>). <article-title>&#x0201C;Differential privacy preservation for deep auto-encoders: An application of human behavior prediction,&#x0201D;</article-title> in <source>Proceedings of the AAAI Conference on Artificial Intelligence</source> (<publisher-loc>Philadelphia</publisher-loc>: <publisher-name>AAAI</publisher-name>), <fpage>1309</fpage>&#x02013;<lpage>1316</lpage>.</citation>
</ref>
<ref id="B96">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Qin</surname> <given-names>Z.</given-names></name> <name><surname>Li</surname> <given-names>G. Y.</given-names></name> <name><surname>Ye</surname> <given-names>H.</given-names></name></person-group> (<year>2021</year>). <article-title>Federated learning and wireless communications</article-title>. <source>IEEE Wireless Commun</source>. <volume>28</volume>, <fpage>134</fpage>&#x02013;<lpage>140</lpage>. <pub-id pub-id-type="doi">10.1109/MWC.011.2000501</pub-id></citation>
</ref>
<ref id="B97">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Qiu</surname> <given-names>T.</given-names></name> <name><surname>Chi</surname> <given-names>J.</given-names></name> <name><surname>Zhou</surname> <given-names>X.</given-names></name> <name><surname>Ning</surname> <given-names>Z.</given-names></name> <name><surname>Atiquzzaman</surname> <given-names>M.</given-names></name> <name><surname>Wu</surname> <given-names>D. O.</given-names></name></person-group> (<year>2020</year>). <article-title>Edge computing in industrial internet of things: Architecture, advances and challenges</article-title>. <source>IEEE Commun. Surv. Tutorials</source> <volume>22</volume>, <fpage>2462</fpage>&#x02013;<lpage>2488</lpage>. <pub-id pub-id-type="doi">10.1109/COMST.2020.3009103</pub-id></citation>
</ref>
<ref id="B98">
<citation citation-type="web"><person-group person-group-type="author"><name><surname>Radford</surname> <given-names>A.</given-names></name> <name><surname>Kim</surname> <given-names>J. W.</given-names></name> <name><surname>Hallacy</surname> <given-names>C.</given-names></name> <name><surname>Ramesh</surname> <given-names>A.</given-names></name> <name><surname>Goh</surname> <given-names>G.</given-names></name> <name><surname>Agarwal</surname> <given-names>S.</given-names></name> <etal/></person-group>. (<year>2021</year>). <article-title>Learning transferable visual models from natural language supervision</article-title>. <source>Proc. Int. Conf. Mach. Learn</source>. (<source>ICML)</source> <volume>38</volume>, <fpage>8748</fpage>&#x02013;<lpage>8763</lpage>. Available online at: <ext-link ext-link-type="uri" xlink:href="https://proceedings.mlr.press/v139/radford21a">https://proceedings.mlr.press/v139/radford21a</ext-link></citation>
</ref>
<ref id="B99">
<citation citation-type="web"><person-group person-group-type="author"><name><surname>Raffel</surname> <given-names>C.</given-names></name> <name><surname>Shazeer</surname> <given-names>N.</given-names></name> <name><surname>Roberts</surname> <given-names>A.</given-names></name> <name><surname>Lee</surname> <given-names>K.</given-names></name> <name><surname>Narang</surname> <given-names>S.</given-names></name> <name><surname>Matena</surname> <given-names>M.</given-names></name> <etal/></person-group>. (<year>2020</year>). <article-title>Exploring the limits of transfer learning with a unified text-to-text transformer</article-title>. <source>J. Mach. Learn. Res</source>. <volume>21</volume>, <fpage>1</fpage>&#x02013;<lpage>67</lpage>. Available online at: <ext-link ext-link-type="uri" xlink:href="http://jmlr.org/papers/v21/20-074.html">http://jmlr.org/papers/v21/20-074.html</ext-link></citation>
</ref>
<ref id="B100">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ren</surname> <given-names>J.</given-names></name> <name><surname>He</surname> <given-names>Y.</given-names></name> <name><surname>Huang</surname> <given-names>G.</given-names></name> <name><surname>Yu</surname> <given-names>G.</given-names></name> <name><surname>Cai</surname> <given-names>Y.</given-names></name> <name><surname>Zhang</surname> <given-names>Z.</given-names></name></person-group> (<year>2019b</year>). <article-title>An edge-computing based architecture for mobile augmented reality</article-title>. <source>IEEE Network</source> <volume>33</volume>, <fpage>162</fpage>&#x02013;<lpage>169</lpage>. <pub-id pub-id-type="doi">10.1109/MNET.2018.1800132</pub-id><pub-id pub-id-type="pmid">34458659</pub-id></citation></ref>
<ref id="B101">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ren</surname> <given-names>J.</given-names></name> <name><surname>Yu</surname> <given-names>G.</given-names></name> <name><surname>He</surname> <given-names>Y.</given-names></name> <name><surname>Li</surname> <given-names>G. Y.</given-names></name></person-group> (<year>2019a</year>). <article-title>Collaborative cloud and edge computing for latency minimization</article-title>. <source>IEEE Trans. Vehicular Technol</source>. <volume>68</volume>, <fpage>5031</fpage>&#x02013;<lpage>5044</lpage>. <pub-id pub-id-type="doi">10.1109/TVT.2019.2904244</pub-id></citation>
</ref>
<ref id="B102">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ribeiro</surname> <given-names>A. H.</given-names></name> <name><surname>Tiels</surname> <given-names>K.</given-names></name> <name><surname>Aguirre</surname> <given-names>L. A.</given-names></name> <name><surname>Sch&#x000F6;n</surname> <given-names>T.</given-names></name></person-group> (<year>2020</year>). <article-title>&#x0201C;Beyond exploding and vanishing gradients: analysing rnn training using attractors and smoothness,&#x0201D;</article-title> in <source>Proceedings of the Twenty Third International Conference on Artificial Intelligence and Statistics</source>, eds. S. Chiappa, and R. Calandra (New York: Proceedings of Machine Learning Research).</citation>
</ref>
<ref id="B103">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Roszyk</surname> <given-names>K.</given-names></name> <name><surname>Nowicki</surname> <given-names>M. R.</given-names></name> <name><surname>Skrzypczy&#x00144;ki</surname> <given-names>P.</given-names></name></person-group> (<year>2022</year>). <article-title>Adopting the yolov4 architecture for low-latency multispectral pedestrian detection in autonomous driving</article-title>. <source>Sensors</source> <volume>22</volume>:<fpage>1082</fpage>. <pub-id pub-id-type="doi">10.3390/s22031082</pub-id><pub-id pub-id-type="pmid">35161827</pub-id></citation></ref>
<ref id="B104">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Saha</surname> <given-names>R.</given-names></name> <name><surname>Chakraborty</surname> <given-names>A.</given-names></name> <name><surname>Misra</surname> <given-names>S.</given-names></name> <name><surname>Das</surname> <given-names>S. K.</given-names></name> <name><surname>Chatterjee</surname> <given-names>C.</given-names></name></person-group> (<year>2021</year>). <article-title>Dlsense: Distributed learning-based smart virtual sensing for precision agriculture</article-title>. <source>IEEE Sens. J</source>. <volume>21</volume>, <fpage>17556</fpage>&#x02013;<lpage>17563</lpage>. <pub-id pub-id-type="doi">10.1109/JSEN.2020.3048593</pub-id></citation>
</ref>
<ref id="B105">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sakr</surname> <given-names>C.</given-names></name> <name><surname>Dai</surname> <given-names>S.</given-names></name> <name><surname>Venkatesan</surname> <given-names>R.</given-names></name> <name><surname>Zimmer</surname> <given-names>B.</given-names></name> <name><surname>Dally</surname> <given-names>W.</given-names></name> <name><surname>Khailany</surname> <given-names>B.</given-names></name></person-group> (<year>2022</year>). <article-title>&#x0201C;Optimal clipping and magnitude-aware differentiation for improved quantization-aware training,&#x0201D;</article-title> in <source>Proceedings of the 39th International Conference on Machine Learning</source>, eds. K. Chaudhuri, S. Jegelka, L. Song, C. Szepesvari, G. Niu, and S. Sabato (New York: PMLR), <fpage>19123</fpage>&#x02013;<lpage>19138</lpage>.</citation>
</ref>
<ref id="B106">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sandhu</surname> <given-names>A. K.</given-names></name></person-group> (<year>2022</year>). <article-title>Big data with cloud computing: Discussions and challenges</article-title>. <source>Big Data Mining Analyt</source>. <volume>5</volume>, <fpage>32</fpage>&#x02013;<lpage>40</lpage>. <pub-id pub-id-type="doi">10.26599/BDMA.2021.9020016</pub-id></citation>
</ref>
<ref id="B107">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sanh</surname> <given-names>V.</given-names></name> <name><surname>Debut</surname> <given-names>L.</given-names></name> <name><surname>Chaumond</surname> <given-names>J.</given-names></name> <name><surname>Wolf</surname> <given-names>T.</given-names></name></person-group> (<year>2019</year>). <article-title>Distilbert, a distilled version of bert: Smaller, faster, cheaper and lighter</article-title>. <source>arXiv</source> [preprint] arXiv:1910.01108. <pub-id pub-id-type="doi">10.48550/arXiv.1910.01108</pub-id><pub-id pub-id-type="pmid">39588066</pub-id></citation></ref>
<ref id="B108">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Savadi Hosseini</surname> <given-names>M.</given-names></name> <name><surname>Ghaderi</surname> <given-names>F.</given-names></name></person-group> (<year>2020</year>). <article-title>A hybrid deep learning architecture using 3d cnns and grus for human action recognition</article-title>. <source>Int. J. Eng</source>. <volume>33</volume>, <fpage>959</fpage>&#x02013;<lpage>965</lpage>. <pub-id pub-id-type="doi">10.5829/ije.2020.33.05b.29</pub-id></citation>
</ref>
<ref id="B109">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Schuman</surname> <given-names>C. D.</given-names></name> <name><surname>Kulkarni</surname> <given-names>S. R.</given-names></name> <name><surname>Parsa</surname> <given-names>M.</given-names></name> <name><surname>Mitchell</surname> <given-names>J. P.</given-names></name> <name><surname>Date</surname> <given-names>P.</given-names></name> <name><surname>Kay</surname> <given-names>B.</given-names></name></person-group> (<year>2022</year>). <article-title>Opportunities for neuromorphic computing algorithms and applications</article-title>. <source>Nat, Comp. Sci</source>. <volume>2</volume>, <fpage>10</fpage>&#x02013;<lpage>19</lpage>. <pub-id pub-id-type="doi">10.1038/s43588-021-00184-y</pub-id><pub-id pub-id-type="pmid">38177712</pub-id></citation></ref>
<ref id="B110">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Shi</surname> <given-names>W.</given-names></name> <name><surname>Cao</surname> <given-names>J.</given-names></name> <name><surname>Zhang</surname> <given-names>Q.</given-names></name> <name><surname>Li</surname> <given-names>Y.</given-names></name> <name><surname>Xu</surname> <given-names>L.</given-names></name></person-group> (<year>2016</year>). <article-title>Edge computing: vision and challenges</article-title>. <source>IEEE Intern. Things J</source>. <volume>3</volume>, <fpage>637</fpage>&#x02013;<lpage>646</lpage>. <pub-id pub-id-type="doi">10.1109/JIOT.2016.2579198</pub-id></citation>
</ref>
<ref id="B111">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Shiranthika</surname> <given-names>C.</given-names></name> <name><surname>Saeedi</surname> <given-names>P.</given-names></name> <name><surname>Baji&#x00107;</surname> <given-names>I. V.</given-names></name></person-group> (<year>2023</year>). <article-title>Decentralized learning in healthcare: A review of emerging techniques</article-title>. <source>IEEE Access</source> <volume>11</volume>, <fpage>54188</fpage>&#x02013;<lpage>54209</lpage>. <pub-id pub-id-type="doi">10.1109/ACCESS.2023.3281832</pub-id></citation>
</ref>
<ref id="B112">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Shuvo</surname> <given-names>M. M. H.</given-names></name> <name><surname>Islam</surname> <given-names>S. K.</given-names></name> <name><surname>Cheng</surname> <given-names>J.</given-names></name> <name><surname>Morshed</surname> <given-names>B. I.</given-names></name></person-group> (<year>2022</year>). <article-title>Efficient acceleration of deep learning inference on resource-constrained edge devices: a review</article-title>. <source>Proc. IEEE</source> <volume>111</volume>, <fpage>42</fpage>&#x02013;<lpage>91</lpage>. <pub-id pub-id-type="doi">10.1109/JPROC.2022.3226481</pub-id></citation>
</ref>
<ref id="B113">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Siriwardhana</surname> <given-names>Y.</given-names></name> <name><surname>Porambage</surname> <given-names>P.</given-names></name> <name><surname>Liyanage</surname> <given-names>M.</given-names></name> <name><surname>Ylianttila</surname> <given-names>M.</given-names></name></person-group> (<year>2021</year>). <article-title>A survey on mobile augmented reality with 5G mobile edge computing: architectures, applications, and technical aspects</article-title>. <source>IEEE Commun. Surv. Tutor</source>. <volume>23</volume>, <fpage>1160</fpage>&#x02013;<lpage>1192</lpage>. <pub-id pub-id-type="doi">10.1109/COMST.2021.3061981</pub-id></citation>
</ref>
<ref id="B114">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Strubell</surname> <given-names>E.</given-names></name> <name><surname>Ganesh</surname> <given-names>A.</given-names></name> <name><surname>McCallum</surname> <given-names>A.</given-names></name></person-group> (<year>2020</year>). <article-title>Energy and policy considerations for modern deep learning research</article-title>. <source>Proc. AAAI Conf. Artif. Intellig</source>. <volume>34</volume>, <fpage>13693</fpage>&#x02013;<lpage>13696</lpage>. <pub-id pub-id-type="doi">10.1609/aaai.v34i09.7123</pub-id><pub-id pub-id-type="pmid">36994188</pub-id></citation></ref>
<ref id="B115">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sun</surname> <given-names>G.</given-names></name> <name><surname>Chen</surname> <given-names>D.</given-names></name> <name><surname>Zhu</surname> <given-names>G.</given-names></name> <name><surname>Li</surname> <given-names>Q.</given-names></name></person-group> (<year>2022</year>). <article-title>Lightweight hybrid materials and structures for energy absorption: a state-of-the-art review and outlook</article-title>. <source>Thin-Walled Struct</source>. <volume>172</volume>:<fpage>108760</fpage>. <pub-id pub-id-type="doi">10.1016/j.tws.2021.108760</pub-id></citation>
</ref>
<ref id="B116">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sun</surname> <given-names>L.</given-names></name> <name><surname>Ansari</surname> <given-names>N.</given-names></name></person-group> (<year>2019</year>). <article-title>Edgeiot: Mobile edge computing for the internet of things</article-title>. <source>IEEE Commun. Magaz</source>. <volume>54</volume>, <fpage>22</fpage>&#x02013;<lpage>29</lpage>. <pub-id pub-id-type="doi">10.1109/MCOM.2016.1600492CM</pub-id></citation>
</ref>
<ref id="B117">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sun</surname> <given-names>Z.</given-names></name> <name><surname>Yu</surname> <given-names>H.</given-names></name> <name><surname>Song</surname> <given-names>X.</given-names></name> <name><surname>Liu</surname> <given-names>R.</given-names></name> <name><surname>Yang</surname> <given-names>Y.</given-names></name> <name><surname>Zhou</surname> <given-names>D.</given-names></name></person-group> (<year>2020</year>). <article-title>Mobilebert: A compact task-agnostic bert for resource-limited devices</article-title>. <source>arXiv</source> [preprint] arXiv:2004.02984. <pub-id pub-id-type="doi">10.18653/v1/2020.acl-main.195</pub-id><pub-id pub-id-type="pmid">36568019</pub-id></citation></ref>
<ref id="B118">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Syu</surname> <given-names>J.-H.</given-names></name> <name><surname>Lin</surname> <given-names>J. C.-W.</given-names></name> <name><surname>Srivastava</surname> <given-names>G.</given-names></name> <name><surname>Yu</surname> <given-names>K.</given-names></name></person-group> (<year>2023</year>). <article-title>A comprehensive survey on artificial intelligence empowered edge computing on consumer electronics</article-title>. <source>IEEE Trans.Consumer Electron</source>. <volume>69</volume>, <fpage>1023</fpage>&#x02013;<lpage>1034</lpage>. <pub-id pub-id-type="doi">10.1109/TCE.2023.3318150</pub-id></citation>
</ref>
<ref id="B119">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Teerapittayanon</surname> <given-names>S.</given-names></name> <name><surname>McDanel</surname> <given-names>B.</given-names></name> <name><surname>Kung</surname> <given-names>H.-T.</given-names></name></person-group> (<year>2016</year>). <article-title>&#x0201C;BranchyNET: Fast inference via early exiting from deep neural networks,&#x0201D;</article-title> in <source>Proceedings of the 23rd International Conference on Pattern Recognition (ICPR)</source> (<publisher-loc>Cancun</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>2464</fpage>&#x02013;<lpage>2469</lpage>.</citation>
</ref>
<ref id="B120">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Vahidian</surname> <given-names>S.</given-names></name> <name><surname>Morafah</surname> <given-names>M.</given-names></name> <name><surname>Lin</surname> <given-names>B.</given-names></name></person-group> (<year>2021</year>). <article-title>&#x0201C;Personalized federated learning by structured and unstructured pruning under data heterogeneity,&#x0201D;</article-title> in <source>2021 IEEE 41st International Conference on Distributed Computing Systems Workshops (ICDCSW)</source> (<publisher-loc>Washington, DC</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>27</fpage>&#x02013;<lpage>34</lpage>. <pub-id pub-id-type="doi">10.1109/ICDCSW53096.2021.00012</pub-id></citation>
</ref>
<ref id="B121">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Vaswani</surname> <given-names>A.</given-names></name> <name><surname>Shazeer</surname> <given-names>N.</given-names></name> <name><surname>Parmar</surname> <given-names>N.</given-names></name> <name><surname>Uszkoreit</surname> <given-names>J.</given-names></name> <name><surname>Jones</surname> <given-names>L.</given-names></name> <name><surname>Gomez</surname> <given-names>A. N.</given-names></name> <etal/></person-group>. (<year>2017</year>). <article-title>&#x0201C;Attention is all you need,&#x0201D;</article-title> in <source>Advances in Neural Information Processing Systems</source> (<publisher-loc>Long Beach, CA</publisher-loc>: <publisher-name>Curran Associates, Inc.</publisher-name>), <fpage>5998</fpage>&#x02013;<lpage>6008</lpage>.</citation>
</ref>
<ref id="B122">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Vepakomma</surname> <given-names>P.</given-names></name> <name><surname>Gupta</surname> <given-names>O.</given-names></name> <name><surname>Swedish</surname> <given-names>A.</given-names></name> <name><surname>Raskar</surname> <given-names>R.</given-names></name></person-group> (<year>2018</year>). <article-title>Split learning for health: Distributed deep learning without sharing raw patient data</article-title>. <source>arXiv [Preprint]</source>. arXiv:1812.00564. <pub-id pub-id-type="doi">10.48550/arXiv.1812.00564</pub-id></citation>
</ref>
<ref id="B123">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wang</surname> <given-names>H.</given-names></name> <name><surname>Lei</surname> <given-names>Z.</given-names></name> <name><surname>Zhang</surname> <given-names>X.</given-names></name> <name><surname>Zhou</surname> <given-names>B.</given-names></name> <name><surname>Peng</surname> <given-names>J.</given-names></name></person-group> (<year>2019</year>). <article-title>A review of deep learning for renewable energy forecasting</article-title>. <source>Energy Convers. Managem</source>. <volume>198</volume>:<fpage>111799</fpage>. <pub-id pub-id-type="doi">10.1016/j.enconman.2019.111799</pub-id></citation>
</ref>
<ref id="B124">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wang</surname> <given-names>S.</given-names></name> <name><surname>Tuor</surname> <given-names>T.</given-names></name> <name><surname>Salonidis</surname> <given-names>T.</given-names></name> <name><surname>Leung</surname> <given-names>K. K.</given-names></name> <name><surname>Makaya</surname> <given-names>C.</given-names></name> <name><surname>He</surname> <given-names>T.</given-names></name> <etal/></person-group>. (<year>2019</year>). <article-title>Adaptive federated learning in resource constrained edge computing systems</article-title>. <source>IEEE J. Select. Areas Commun</source>. <volume>37</volume>, <fpage>1205</fpage>&#x02013;<lpage>1221</lpage>. <pub-id pub-id-type="doi">10.1109/JSAC.2019.2904348</pub-id></citation>
</ref>
<ref id="B125">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wang</surname> <given-names>Y.</given-names></name> <name><surname>Yu</surname> <given-names>Z.</given-names></name> <name><surname>Wu</surname> <given-names>J.</given-names></name> <name><surname>Wang</surname> <given-names>C.</given-names></name> <name><surname>Zhou</surname> <given-names>Q.</given-names></name> <name><surname>Hu</surname> <given-names>J.</given-names></name></person-group> (<year>2024</year>). <article-title>Adaptive knowledge distillation-based lightweight intelligent fault diagnosis framework in iot edge computing</article-title>. <source>IEEE Intern.Things J</source>. <volume>11</volume>, <fpage>23156</fpage>&#x02013;<lpage>23169</lpage>. <pub-id pub-id-type="doi">10.1109/JIOT.2024.3387328</pub-id></citation>
</ref>
<ref id="B126">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Wang</surname> <given-names>Z.</given-names></name> <name><surname>Li</surname> <given-names>C.</given-names></name> <name><surname>Wang</surname> <given-names>X.</given-names></name></person-group> (<year>2021</year>). <article-title>&#x0201C;Convolutional neural network pruning with structural redundancy reduction,&#x0201D;</article-title> in <source>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)</source> (<publisher-loc>Nashville, TN</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>14913</fpage>&#x02013;<lpage>14922</lpage>.</citation>
</ref>
<ref id="B127">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wei</surname> <given-names>K.</given-names></name> <name><surname>Li</surname> <given-names>J.</given-names></name> <name><surname>Ding</surname> <given-names>M.</given-names></name> <name><surname>Ma</surname> <given-names>C.</given-names></name> <name><surname>Yang</surname> <given-names>H. H.</given-names></name> <name><surname>Farokhi</surname> <given-names>F.</given-names></name> <etal/></person-group>. (<year>2020</year>). <article-title>Federated learning with differential privacy: Algorithms and performance analysis</article-title>. <source>IEEE Trans. Inform. Forens. Security</source> <volume>15</volume>, <fpage>3454</fpage>&#x02013;<lpage>3469</lpage>. <pub-id pub-id-type="doi">10.1109/TIFS.2020.2988575</pub-id></citation>
</ref>
<ref id="B128">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Wu</surname> <given-names>Z.</given-names></name> <name><surname>Nagarajan</surname> <given-names>T.</given-names></name> <name><surname>Kumar</surname> <given-names>A.</given-names></name> <name><surname>Davis</surname> <given-names>L. S.</given-names></name> <name><surname>Hariharan</surname> <given-names>B.</given-names></name> <name><surname>Farhadi</surname> <given-names>A.</given-names></name></person-group> (<year>2020</year>). <article-title>&#x0201C;Blockdrop: Dynamic inference paths in residual networks,&#x0201D;</article-title> in <source>Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</source> (<publisher-loc>Salt Lake City</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>8817</fpage>&#x02013;<lpage>8826</lpage>.<pub-id pub-id-type="pmid">36812830</pub-id></citation></ref>
<ref id="B129">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Xia</surname> <given-names>Q.</given-names></name> <name><surname>Ye</surname> <given-names>W.</given-names></name> <name><surname>Tao</surname> <given-names>Z.</given-names></name> <name><surname>Wu</surname> <given-names>J.</given-names></name> <name><surname>Li</surname> <given-names>Q.</given-names></name></person-group> (<year>2021</year>). <article-title>A survey of federated learning for edge computing: Research problems and solutions</article-title>. <source>High-Confidence Comp</source>. <volume>1</volume>:<fpage>100008</fpage>. <pub-id pub-id-type="doi">10.1016/j.hcc.2021.100008</pub-id><pub-id pub-id-type="pmid">35062410</pub-id></citation></ref>
<ref id="B130">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Xie</surname> <given-names>Q.</given-names></name> <name><surname>Zhou</surname> <given-names>X.</given-names></name> <name><surname>Qiu</surname> <given-names>T.</given-names></name> <name><surname>Zhang</surname> <given-names>Q.</given-names></name> <name><surname>Qu</surname> <given-names>W.</given-names></name></person-group> (<year>2022</year>). <article-title>Soft actor&#x02013;critic-based multilevel cooperative perception for connected autonomous vehicles</article-title>. <source>IEEE Intern. Things J.</source> <volume>9</volume>, <fpage>21370</fpage>&#x02013;<lpage>21381</lpage>. <pub-id pub-id-type="doi">10.1109/JIOT.2022.3179739</pub-id></citation>
</ref>
<ref id="B131">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Xu</surname> <given-names>D.</given-names></name> <name><surname>Zhang</surname> <given-names>H.</given-names></name> <name><surname>Yang</surname> <given-names>L.</given-names></name> <name><surname>Liu</surname> <given-names>R.</given-names></name> <name><surname>Xu</surname> <given-names>M.</given-names></name> <name><surname>Liu</surname> <given-names>X.</given-names></name></person-group> (<year>2024</year>). <article-title>&#x0201C;WIP: Efficient LLM prefilling with mobile NPU,&#x0201D;</article-title> in <source>Proceedings of the Workshop on Edge and Mobile Foundation Models</source>, <fpage>33</fpage>&#x02013;<lpage>35</lpage>.</citation>
</ref>
<ref id="B132">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Xu</surname> <given-names>K.</given-names></name> <name><surname>Wang</surname> <given-names>Z.</given-names></name> <name><surname>Chen</surname> <given-names>C.</given-names></name> <name><surname>Geng</surname> <given-names>X.</given-names></name> <name><surname>Lin</surname> <given-names>J.</given-names></name> <name><surname>Yang</surname> <given-names>X.</given-names></name> <etal/></person-group>. (<year>2025</year>). <article-title>&#x0201C;LPVIT: Low-power semi-structured pruning for vision transformers,&#x0201D;</article-title> in eds. <italic>Computer Vision-ECCV 2024</italic>, A. Leonardis, E. Ricci, S. Roth, O. Russakovsky, T. Sattler, and G. Varol (<publisher-loc>Cham</publisher-loc>: <publisher-name>Springer Nature Switzerland</publisher-name>), <fpage>269</fpage>&#x02013;<lpage>287</lpage>.</citation>
</ref>
<ref id="B133">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yang</surname> <given-names>J.</given-names></name> <name><surname>Jin</surname> <given-names>H.</given-names></name> <name><surname>Tang</surname> <given-names>R.</given-names></name> <name><surname>Han</surname> <given-names>X.</given-names></name> <name><surname>Feng</surname> <given-names>Q.</given-names></name> <name><surname>Jiang</surname> <given-names>H.</given-names></name> <etal/></person-group>. (<year>2024</year>). <article-title>Harnessing the power of llms in practice: a survey on chatgpt and beyond</article-title>. <source>ACM Trans. Knowl. Discov. Data</source> <volume>18</volume>, <fpage>1</fpage>&#x02013;<lpage>32</lpage>. <pub-id pub-id-type="doi">10.1145/3649506</pub-id></citation>
</ref>
<ref id="B134">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Yang</surname> <given-names>Y.</given-names></name> <name><surname>Li</surname> <given-names>M.</given-names></name> <name><surname>Huang</surname> <given-names>S.</given-names></name> <name><surname>Lu</surname> <given-names>H.</given-names></name> <name><surname>Tu</surname> <given-names>W.</given-names></name> <name><surname>Wan</surname> <given-names>W.</given-names></name></person-group> (<year>2023</year>). <article-title>&#x0201C;Multi-scale spatial-spectral attention guided fusion network for pansharpening,&#x0201D;</article-title> in <source>Proceedings of the 31st ACM International Conference on Multimedia, MM &#x00027;23</source> (<publisher-loc>New York, NY</publisher-loc>: <publisher-name>Association for Computing Machinery</publisher-name>), <fpage>3346</fpage>&#x02013;<lpage>3354</lpage>.</citation>
</ref>
<ref id="B135">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Yao</surname> <given-names>S.</given-names></name> <name><surname>Wan</surname> <given-names>X.</given-names></name></person-group> (<year>2020</year>). <article-title>&#x0201C;Multimodal transformer for multimodal machine translation,&#x0201D;</article-title> in <source>Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics</source>, eds. D. Jurafsky, J. Chai, N. Schluter, and J. Tetreault (<publisher-loc>New York</publisher-loc>: <publisher-name>Association for Computational Linguistics</publisher-name>), <fpage>4346</fpage>&#x02013;<lpage>4350</lpage>.</citation>
</ref>
<ref id="B136">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ye</surname> <given-names>M.</given-names></name> <name><surname>Fang</surname> <given-names>X.</given-names></name> <name><surname>Du</surname> <given-names>B.</given-names></name> <name><surname>Yuen</surname> <given-names>P. C.</given-names></name> <name><surname>Tao</surname> <given-names>D.</given-names></name></person-group> (<year>2023</year>). <article-title>Heterogeneous federated learning: State-of-the-art and research challenges</article-title>. <source>ACM Com. Surv</source>. <volume>56</volume>, <fpage>1</fpage>&#x02013;<lpage>44</lpage>. <pub-id pub-id-type="doi">10.1145/3625558</pub-id></citation>
</ref>
<ref id="B137">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yeom</surname> <given-names>S.-K.</given-names></name> <name><surname>Seegerer</surname> <given-names>P.</given-names></name> <name><surname>Lapuschkin</surname> <given-names>S.</given-names></name> <name><surname>Binder</surname> <given-names>A.</given-names></name> <name><surname>Wiedemann</surname> <given-names>S.</given-names></name> <name><surname>M&#x000FC;ller</surname> <given-names>K.-R.</given-names></name> <etal/></person-group>. (<year>2021</year>). <article-title>Pruning by explaining: A novel criterion for deep neural network pruning</article-title>. <source>Pattern Recognit</source>. <volume>115</volume>:<fpage>107899</fpage>. <pub-id pub-id-type="doi">10.1016/j.patcog.2021.107899</pub-id></citation>
</ref>
<ref id="B138">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yu</surname> <given-names>B.</given-names></name> <name><surname>Zhang</surname> <given-names>X.</given-names></name> <name><surname>You</surname> <given-names>I.</given-names></name> <name><surname>Khan</surname> <given-names>U. S.</given-names></name></person-group> (<year>2021</year>). <article-title>Efficient computation offloading in edge computing enabled smart home</article-title>. <source>IEEE Access</source> <volume>9</volume>, <fpage>48631</fpage>&#x02013;<lpage>48639</lpage>. <pub-id pub-id-type="doi">10.1109/ACCESS.2021.3066789</pub-id><pub-id pub-id-type="pmid">39314732</pub-id></citation></ref>
<ref id="B139">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yu</surname> <given-names>W.</given-names></name> <name><surname>Liang</surname> <given-names>F.</given-names></name> <name><surname>He</surname> <given-names>X.</given-names></name> <name><surname>Hatcher</surname> <given-names>W. G.</given-names></name> <name><surname>Lu</surname> <given-names>C.</given-names></name> <name><surname>Lin</surname> <given-names>J.</given-names></name> <etal/></person-group>. (<year>2018</year>). <article-title>A survey on the edge computing for the internet of things</article-title>. <source>IEEE Access</source> <volume>6</volume>, <fpage>6900</fpage>&#x02013;<lpage>6919</lpage>. <pub-id pub-id-type="doi">10.1109/ACCESS.2017.2778504</pub-id></citation>
</ref>
<ref id="B140">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yuan</surname> <given-names>F.</given-names></name> <name><surname>Shou</surname> <given-names>L.</given-names></name> <name><surname>Pei</surname> <given-names>J.</given-names></name> <name><surname>Lin</surname> <given-names>W.</given-names></name> <name><surname>Gong</surname> <given-names>M.</given-names></name> <name><surname>Fu</surname> <given-names>Y.</given-names></name> <etal/></person-group>. (<year>2021</year>). <article-title>Reinforced multi-teacher selection for knowledge distillation</article-title>. <source>Proc. AAAI Conf. Artif. Intellig</source>. <volume>35</volume>, <fpage>14284</fpage>&#x02013;<lpage>14291</lpage>. <pub-id pub-id-type="doi">10.1609/aaai.v35i16.17680</pub-id></citation>
</ref>
<ref id="B141">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zeng</surname> <given-names>Z.</given-names></name> <name><surname>Liu</surname> <given-names>C.</given-names></name> <name><surname>Tang</surname> <given-names>Z.</given-names></name> <name><surname>Li</surname> <given-names>K.</given-names></name> <name><surname>Li</surname> <given-names>K.</given-names></name></person-group> (<year>2022</year>). <article-title>Acctfm: An effective intra-layer model parallelization strategy for training large-scale transformer-based models</article-title>. <source>IEEE Trans. Parallel Distrib. Syst</source>. <volume>33</volume>, <fpage>4326</fpage>&#x02013;<lpage>4338</lpage>. <pub-id pub-id-type="doi">10.1109/TPDS.2022.3187815</pub-id></citation>
</ref>
<ref id="B142">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhang</surname> <given-names>C.</given-names></name> <name><surname>Zhang</surname> <given-names>W.</given-names></name> <name><surname>Wu</surname> <given-names>Q.</given-names></name> <name><surname>Fan</surname> <given-names>P.</given-names></name> <name><surname>Fan</surname> <given-names>Q.</given-names></name> <name><surname>Wang</surname> <given-names>J.</given-names></name> <etal/></person-group>. (<year>2025</year>). <article-title>Distributed deep reinforcement learning based gradient quantization for federated learning enabled vehicle edge computing</article-title>. <source>IEEE Intern. Things J</source>. <volume>12</volume>, <fpage>4899</fpage>&#x02013;<lpage>4913</lpage>. <pub-id pub-id-type="doi">10.1109/JIOT.2024.3447036</pub-id></citation>
</ref>
<ref id="B143">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Zhang</surname> <given-names>C.</given-names></name> <name><surname>Zheng</surname> <given-names>R.</given-names></name> <name><surname>Cui</surname> <given-names>Y.</given-names></name> <name><surname>Li</surname> <given-names>C.</given-names></name> <name><surname>Wu</surname> <given-names>J.</given-names></name></person-group> (<year>2020</year>). <article-title>&#x0201C;Delay-sensitive computation partitioning for mobile augmented reality applications,&#x0201D;</article-title> in <source>2020 IEEE/ACM 28th International Symposium on Quality of Service (IWQoS)</source> (<publisher-loc>Hang Zhou</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>1</fpage>&#x02013;<lpage>10</lpage>.</citation>
</ref>
<ref id="B144">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhang</surname> <given-names>J.</given-names></name> <name><surname>Chen</surname> <given-names>B.</given-names></name> <name><surname>Zhao</surname> <given-names>Y.</given-names></name> <name><surname>Cheng</surname> <given-names>X.</given-names></name> <name><surname>Hu</surname> <given-names>F.</given-names></name></person-group> (<year>2018</year>). <article-title>Data security and privacy-preserving in edge computing paradigm: Survey and open issues</article-title>. <source>IEEE Access</source> <volume>6</volume>, <fpage>18209</fpage>&#x02013;<lpage>18237</lpage>. <pub-id pub-id-type="doi">10.1109/ACCESS.2018.2820162</pub-id><pub-id pub-id-type="pmid">25014943</pub-id></citation></ref>
<ref id="B145">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhang</surname> <given-names>S.</given-names></name> <name><surname>Zhou</surname> <given-names>X.</given-names></name> <name><surname>Qiu</surname> <given-names>T.</given-names></name> <name><surname>Wu</surname> <given-names>D. O.</given-names></name></person-group> (<year>2024</year>). <article-title>Quantum-inspired robust networking model with multiverse co-evolution for scale-free iot</article-title>. <source>IEEE Trans. Mobile Comp</source>. <volume>23</volume>, <fpage>14085</fpage>&#x02013;<lpage>14098</lpage>. <pub-id pub-id-type="doi">10.1109/TMC.2024.3439511</pub-id></citation>
</ref>
<ref id="B146">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Zhang</surname> <given-names>W.</given-names></name> <name><surname>Han</surname> <given-names>B.</given-names></name> <name><surname>Hui</surname> <given-names>P.</given-names></name></person-group> (<year>2018</year>). <article-title>&#x0201C;Jaguar: Low latency mobile augmented reality with flexible tracking,&#x0201D;</article-title> in <source>Proceedings of the 26th ACM International Conference on Multimedia, MM &#x00027;18</source> (<publisher-loc>New York, NY</publisher-loc>: <publisher-name>Association for Computing Machinery</publisher-name>), <fpage>355</fpage>&#x02013;<lpage>363</lpage>.</citation>
</ref>
<ref id="B147">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhou</surname> <given-names>B.</given-names></name> <name><surname>Yang</surname> <given-names>G.</given-names></name> <name><surname>Shi</surname> <given-names>Z.</given-names></name> <name><surname>Ma</surname> <given-names>S.</given-names></name></person-group> (<year>2022</year>). <article-title>Natural language processing for smart healthcare</article-title>. <source>IEEE Rev. Biomed. Eng</source>. <volume>17</volume>, <fpage>4</fpage>&#x02013;<lpage>18</lpage>. <pub-id pub-id-type="doi">10.1109/RBME.2022.3210270</pub-id><pub-id pub-id-type="pmid">36170385</pub-id></citation></ref>
<ref id="B148">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhou</surname> <given-names>H.</given-names></name> <name><surname>Yang</surname> <given-names>G.</given-names></name> <name><surname>Dai</surname> <given-names>H.</given-names></name> <name><surname>Liu</surname> <given-names>G.</given-names></name></person-group> (<year>2022a</year>). <article-title>Pflf: Privacy-preserving federated learning framework for edge computing</article-title>. <source>IEEE Trans. Inform. Forens. Secur</source>. <volume>17</volume>, <fpage>1905</fpage>&#x02013;<lpage>1918</lpage>. <pub-id pub-id-type="doi">10.1109/TIFS.2022.3174394</pub-id></citation>
</ref>
<ref id="B149">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhou</surname> <given-names>H.</given-names></name> <name><surname>Yang</surname> <given-names>G.</given-names></name> <name><surname>Huang</surname> <given-names>Y.</given-names></name> <name><surname>Dai</surname> <given-names>H.</given-names></name> <name><surname>Xiang</surname> <given-names>Y.</given-names></name></person-group> (<year>2022b</year>). <article-title>Privacy-preserving and verifiable federated learning framework for edge computing</article-title>. <source>IEEE Trans. Inform. Forens. Secur</source>. <volume>18</volume>, <fpage>565</fpage>&#x02013;<lpage>580</lpage>. <pub-id pub-id-type="doi">10.1109/TIFS.2022.3227435</pub-id></citation>
</ref>
<ref id="B150">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhou</surname> <given-names>X.</given-names></name> <name><surname>Ge</surname> <given-names>S.</given-names></name> <name><surname>Liu</surname> <given-names>P.</given-names></name> <name><surname>Qiu</surname> <given-names>T.</given-names></name></person-group> (<year>2024a</year>). <article-title>Dag-based dependent tasks offloading in mec-enabled iot with soft cooperation</article-title>. <source>IEEE Trans. Mobile Comp</source>. <volume>23</volume>, <fpage>6908</fpage>&#x02013;<lpage>6920</lpage>. <pub-id pub-id-type="doi">10.1109/TMC.2023.3328333</pub-id></citation>
</ref>
<ref id="B151">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhou</surname> <given-names>X.</given-names></name> <name><surname>Ge</surname> <given-names>S.</given-names></name> <name><surname>Qiu</surname> <given-names>T.</given-names></name> <name><surname>Li</surname> <given-names>K.</given-names></name> <name><surname>Atiquzzaman</surname> <given-names>M.</given-names></name></person-group> (<year>2023</year>). <article-title>Energy-efficient service migration for multi-user heterogeneous dense cellular networks</article-title>. <source>IEEE Trans. Mobile Comp</source>. <volume>22</volume>, <fpage>890</fpage>&#x02013;<lpage>905</lpage>. <pub-id pub-id-type="doi">10.1109/TMC.2021.3087198</pub-id></citation>
</ref>
<ref id="B152">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhou</surname> <given-names>X.</given-names></name> <name><surname>Ke</surname> <given-names>Z.</given-names></name> <name><surname>Qiu</surname> <given-names>T.</given-names></name></person-group> (<year>2024b</year>). <article-title>Recommendation-driven multi-cell cooperative caching: A multi-agent reinforcement learning approach</article-title>. <source>IEEE Trans. Mobile Comp</source>. <volume>23</volume>, <fpage>4764</fpage>&#x02013;<lpage>4776</lpage>. <pub-id pub-id-type="doi">10.1109/TMC.2023.3297213</pub-id></citation>
</ref>
<ref id="B153">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Zhou</surname> <given-names>Y.</given-names></name> <name><surname>Moosavi-Dezfooli</surname> <given-names>S.-M.</given-names></name> <name><surname>Cheung</surname> <given-names>N.-M.</given-names></name> <name><surname>Frossard</surname> <given-names>P.</given-names></name></person-group> (<year>2018</year>). <article-title>&#x0201C;Adaptive quantization for deep neural network,&#x0201D;</article-title> in <source>Proceedings of the AAAI Conference on Artificial Intelligence</source> (<publisher-loc>Washington, DC</publisher-loc>: <publisher-name>AAAI</publisher-name>), <fpage>32</fpage>.</citation>
</ref>
<ref id="B154">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zonta</surname> <given-names>T.</given-names></name> <name><surname>Da Costa</surname> <given-names>C. A.</given-names></name> <name><surname>da Rosa Righi</surname> <given-names>R.</given-names></name> <name><surname>de Lima</surname> <given-names>M. J.</given-names></name> <name><surname>da Trindade</surname> <given-names>E. S.</given-names></name> <name><surname>Li</surname> <given-names>G. P.</given-names></name></person-group> (<year>2020</year>). <article-title>Predictive maintenance in the industry 4.0: A systematic literature review</article-title>. <source>Comp. Indust. Eng</source>. <volume>150</volume>:<fpage>106889</fpage>. <pub-id pub-id-type="doi">10.1016/j.cie.2020.106889</pub-id></citation>
</ref>
</ref-list>
</back>
</article>