PTC-Decoder:迈向离线资源受限边缘设备上的智能SLMs

arXiv cs.AI 论文

摘要

PTC-Decoder是一个无需训练、即插即用的框架,通过在词元级约束强制执行计划遵守,提升在离线、资源受限的边缘设备上的小型语言模型性能,增强多步骤代理任务中的步骤级可靠性。

arXiv:2609.30836v1 Announce Type: new Abstract: Deploying small language models (SLMs) on offline, resource-constrained edge devices such as remote sensing satellites presents a fundamental challenge: their limited reasoning capacity hinders reliable execution of multi-step agent tasks requiring complex tool orchestration. Existing plan-solve paradigms rely on prompt-based enforcement, which our experiments show SLMs almost entirely disregard: weak models fail to invoke the plan. We propose PTC-Decoder (Plan-Tool Constrained Decoder), a training-free, plug-and-play decoder framework that combines (1) a Plan-to-Act paradigm, which elevates planning to an atomic tool and forces its invocation at the first inference step, and (2) TC-Decoder, a deterministic finite automaton that imposes token-level hard constraints on tool names while preserving freedom over parameter generation, thereby retaining SLM reasoning capability. Evaluated on 200 real remote-sensing satellite tasks across 7 SLMs, PTC-Decoder yields a statistically significant mean overall score gain of +1.21 (p<0.01), 95% CI [+1.13, +1.29]), with consistent improvements across models and other datasets. An ablation study that removes TC-Decoder causes substantial performance degradation across all quality metrics without reducing computational cost, confirming TC-Decoder as the primary driver. PTC-Decoder thus offers a lightweight yet effective solution for improving step-level reliability, with final-answer accuracy remaining an open challenge. In essence, we enforce plan adherence by constraining the permissible output vocabulary during inference, without requiring retraining.
查看原文
查看缓存全文

缓存时间: 2026/09/28 09:50

# PTC-Decoder: Towards Intelligent SLMs on Offline Resource-Constrained Edge Devices
Source: [https://arxiv.org/html/2609.30836](https://arxiv.org/html/2609.30836)
###### Abstract

Deploying small language models \(SLMs\) on offline, resource\-constrained edge devices such as remote sensing satellites presents a fundamental challenge: their limited reasoning capacity hinders reliable execution of multi\-step agent tasks requiring complex tool orchestration\. Existing plan\-solve paradigms rely on prompt\-based enforcement, which our experiments show SLMs almost entirely disregard: weak models fail to invoke the plan\. We propose PTC\-Decoder \(Plan\-Tool Constrained Decoder\), a training\-free, plug\-and\-play decoder framework that combines \(1\) a Plan\-to\-Act paradigm, which elevates planning to an atomic tool and forces its invocation at the first inference step, and \(2\) TC\-Decoder, a deterministic finite automaton that imposes token\-level hard constraints on tool names while preserving freedom over parameter generation, thereby retaining SLM reasoning capability\. Evaluated on 200 real remote\-sensing satellite tasks across 7 SLMs, PTC\-Decoder yields a statistically significant mean overall score gain of \+1\.21 \(p<0\.01p<0\.01, 95% CI \[\+1\.13, \+1\.29\]\), with consistent improvements across models and other datasets\. An ablation study that removes TC\-Decoder causes substantial performance degradation across all quality metrics without reducing computational cost, confirming TC\-Decoder as the primary driver\. PTC\-Decoder thus offers a lightweight yet effective solution for improving step\-level reliability, with final\-answer accuracy remaining an open challenge\. In essence, we enforce plan adherence by constraining the permissible output vocabulary during inference, without requiring retraining\.

1Shanghai Jiao Tong University, Computer Institute

2Shanghai Landfun Information Technology

yminghui@sjtu\.edu\.cn, muke2433@gmail\.com, dr\.wugang@sjtu\.edu\.cn

Code—https://github\.com/yuminghui/llm\-tool\-constrained\-decoder

## Introduction

Large language models \(LLMs\) have achieved remarkable success and continuous evolution across diverse scenarios, including software engineering\([Yu et al\. 2024](https://arxiv.org/html/2609.30836#bib.bib2)\), intelligent customer service\([Hong et al\. 2025](https://arxiv.org/html/2609.30836#bib.bib3)\), and enterprise knowledge retrieval\([Mishra et al\. 2026](https://arxiv.org/html/2609.30836#bib.bib4);[Zhao et al\. 2026](https://arxiv.org/html/2609.30836#bib.bib1)\)\. However, deploying LLMs on offline edge devices is challenging due to severe resource constraints \(e\.g\., unified memory≤\\leq16 GB\) and competing resource\-intensive applications\. In our satellite Agent system \(Figure[1](https://arxiv.org/html/2609.30836#Sx1.F1)\) running on an NVIDIA Jetson Orin NX, high\-resolution image recognition, remote sensing index analysis, and change detection leave less than 4 GB of memory for SLM deployment\. We address this by constraining the model’s output vocabulary during decoding, forcing tool adherence without training\.

![Refer to caption](https://arxiv.org/html/2609.30836v1/complete_task_demo.png)Figure 1:Illustration of the in\-orbit task processing flow for a satellite with an onboard agent\.Before presenting our method, we first examine why existing studies fall short in this setting\. They primarily pursue two approaches to enhance SLM\-based agent capabilities: \(1\) end\-to\-end training\([Erdogan et al\. 2024](https://arxiv.org/html/2609.30836#bib.bib20)\)and knowledge distillation\([Chen et al\. 2025](https://arxiv.org/html/2609.30836#bib.bib21);[Sharma and Mehta 2025](https://arxiv.org/html/2609.30836#bib.bib27)\), \(2\) plan\-then\-execute paradigm\([Wang et al\. 2023](https://arxiv.org/html/2609.30836#bib.bib23)\)\. However, the former relies on training data, which is extremely scarce in the remote sensing agent domain, rendering such methods inapplicable to our application domain\. The latter, although significantly boosting LLM agent performance, is found by our experiments \(Table[1](https://arxiv.org/html/2609.30836#Sx1.T1)\) to be incapable of ensuring stable planning in SLMs through prompt constraints alone\.

In resource\-constrained edge device scenarios, SLMs are typically adopted\. Despite the rapid development of LLMs bringing training optimizations \(e\.g\., GRPO\([Shao et al\. 2024](https://arxiv.org/html/2609.30836#bib.bib26)\)\) and architectural improvements \(e\.g\., sliding window attention\([Yu et al\. 2026](https://arxiv.org/html/2609.30836#bib.bib30);[Team et al\. 2025](https://arxiv.org/html/2609.30836#bib.bib31)\), gated linear attention\) that enhance SLMs’ reasoning capabilities and context length\([Lu et al\. 2025a](https://arxiv.org/html/2609.30836#bib.bib32)\), they still suffer from severe hallucinations, insufficient reasoning ability, and misuse of tools when handling complex agent tasks and long contexts\([Sun et al\. 2025](https://arxiv.org/html/2609.30836#bib.bib33)\), making it difficult to reliably execute\.

Table 1:Prompt\-Plan enforcement versus baseline\. The former addsMUST call plan firstinto system prompt instructing the model to callp​l​a​nplan\. P\.R\.: plan call rate; Suc\.: success rate, higher the value, fewer times the agent crashes\. Weaker models completely ignore the prompt \(P\.R\.==0\), and even capable models show degraded performance\.Furthermore, some studies adopt task decomposition and edge\-cloud collaboration\([Yi et al\. 2026](https://arxiv.org/html/2609.30836#bib.bib22)\), where simple tasks are processed locally while complex ones are offloaded to the cloud\([Hao et al\. 2024](https://arxiv.org/html/2609.30836#bib.bib28);[Li et al\. 2025b](https://arxiv.org/html/2609.30836#bib.bib29)\), offering a new perspective for enhancing edge intelligence\. However, most remote sensing satellites operate as offline devices without Internet access\([Wu et al\. 2025a](https://arxiv.org/html/2609.30836#bib.bib25)\)and rely solely on microwave or laser links for ground communication\([Wang et al\. 2024](https://arxiv.org/html/2609.30836#bib.bib24)\), making the powerful cloud\-based LLMs inaccessible to our Agent\. Current SLM research still fails to address the instruction\-following deficiency of SLMs in offline resource\-constrained scenarios\.

To this end, we proposePlan\-ToolConstrainedDecoder\(PTC\-Decoder\), a decoder framework that enforces planning and tool invocation at the decoding level\. The initial attempt—forcing the agent to generate a plan at the first step via prompting—proved ineffective \(Table[1](https://arxiv.org/html/2609.30836#Sx1.T1)\)\. Meanwhile, the existing constrained decoders cannot be directly restricted from a specific tool level\([Koo et al\. 2024](https://arxiv.org/html/2609.30836#bib.bib19)\)\. We thus introduce TC\-Decoder, which intervenes in the decoding process to deterministically enforce tool\-name outputs at designated steps \(while leaving parameters to the model\)\. Integrating both—planning first, then constrained tool execution—yields PTC\-Decoder\. Although this may reduce model creativity, experiments show that appropriate constraints better utilize SLMs’ limited reasoning for complex tasks\. The decoder is plug\-and\-play, requiring no training\.

Overall, our contributions are threefold:

1. 1\.We proposePTC\-Decoder, a plug\-and\-play decoder with Plan\-to\-Act tailored for SLM Agents\. This design mitigates the reasoning bottleneck on offline edge devices and significantly improves performance on a satellite agent benchmark and other datasets\.
2. 2\.We construct and open\-source a compact yet high\-quality on\-board Agent operational benchmark, providing a foundational reference and starting point for the future construction of more comprehensive databases\.
3. 3\.Across multiple benchmark metrics, PTC\-Decoder significantly outperforms the baseline, achieving a mean overall gain of \+1\.21 across 7 SLMs, with the weakest model improving by 350% and capable models reaching an F1 of up to 0\.359 against ground\-truth tool sequences\.

Our SLM agent investigated in this study deploys on commercial satellites \(16GB\) to support ground control with remote sensing services\.

![Refer to caption](https://arxiv.org/html/2609.30836v1/complex_task_example.png)Figure 2:Example of the Expected Complex Task Workflow in the Remote Sensing Satellite Agent System\.
## Related Work

### Application of SLMs on Edge Devices

As large language model capabilities continue to advance\([Vaswani et al\. 2017](https://arxiv.org/html/2609.30836#bib.bib9);[Brown et al\. 2020](https://arxiv.org/html/2609.30836#bib.bib10);[Wei et al\. 2022](https://arxiv.org/html/2609.30836#bib.bib11)\), deploying them onto resource\-constrained edge devices like smartphones and IoT terminals has become a key focus in both academia and industry\. An empirical study covering over 60 SLMs for edge deployment reveals that\([Lu et al\. 2025b](https://arxiv.org/html/2609.30836#bib.bib5)\), despite growing intelligence across tasks, substantial room remains for optimizing in\-context learning and operational efficiency\([Ni et al\. 2026](https://arxiv.org/html/2609.30836#bib.bib40)\)\. Hardware acceleration, edge\-cloud collaboration\([Li et al\. 2026](https://arxiv.org/html/2609.30836#bib.bib6)\), model optimization, and deployment optimization constitute the main research thrusts for SLM edge adoption\([Zheng et al\. 2025](https://arxiv.org/html/2609.30836#bib.bib7)\)\. In addition, limited passive thermal dissipation capacity makes thermal throttling from continuous inference a major challenge\([Wang et al\. 2025](https://arxiv.org/html/2609.30836#bib.bib8)\)\.

### Lightweight Architectures and Model Quantization

The core challenge of deploying language models on edge devices lies in balancing model size, inference speed, and task performance\. Lightweight Transformer architectures and model quantization are two primary approaches\([Hariharan Samson 2026](https://arxiv.org/html/2609.30836#bib.bib12);[Zhou et al\. 2025](https://arxiv.org/html/2609.30836#bib.bib13)\); quantization can compress model size by 4–10× while retaining 75%–96% accuracy\([Zhou et al\. 2025](https://arxiv.org/html/2609.30836#bib.bib13)\)\. Meanwhile, algorithms such as MobiLoRA focus on optimizing LoRA\-adapted LLM inference on mobile devices, particularly for KV Cache storage constraints\([Li et al\. 2025a](https://arxiv.org/html/2609.30836#bib.bib14)\)\.

### Edge\-Cloud Collaboration

Beyond on\-device deployment, edge\-cloud collaboration and multi\-LLM cooperation have emerged as active research directions\([Li et al\. 2026](https://arxiv.org/html/2609.30836#bib.bib6);[Zhu and Yang 2025](https://arxiv.org/html/2609.30836#bib.bib15);[Li et al\. 2023a](https://arxiv.org/html/2609.30836#bib.bib16);[Hong et al\. 2024](https://arxiv.org/html/2609.30836#bib.bib17);[Qian et al\. 2023](https://arxiv.org/html/2609.30836#bib.bib18)\)\. These architectures offload complex tasks to the cloud while keeping simple ones local, effectively alleviating compute bottlenecks, reducing on\-device cost, and partially enhancing SLM intelligence on edge devices\([Zhu and Yang 2025](https://arxiv.org/html/2609.30836#bib.bib15)\)\. However, such solutions are entirely inapplicable for offline devices without network access such as our remote\-sensing satellite\. How to compensate for SLM intelligence deficiencies on non\-networked edge devices remains a largely unexplored pain point, and this work specifically targets this challenge\.

## Methodology

We aim to enhance SLM step\-level tool\-calling adherence for edge agent tasks\. This is critical, as we need to perform highly complex remote\-sensing tasks—such as “How are the plants growing in Shanghai recently?”—with the expected workflow shown in Figure[2](https://arxiv.org/html/2609.30836#Sx1.F2)\([Wang et al\. 2021](https://arxiv.org/html/2609.30836#bib.bib35)\)\. However, models deployable on our satellite are typically quantized small models \(≤\\leq2B\), for which such complex tasks are excessively difficult, rendering correct end\-to\-end agent execution nearly infeasible\. To overcome this challenge, we designPTC\-Decoder, as shown in Figure[3](https://arxiv.org/html/2609.30836#Sx3.F3)\.

![Refer to caption](https://arxiv.org/html/2609.30836v1/PTC-Decoder.png)Figure 3:Our complete PTC\-Decoder architectureThis architecture prevents the model from invoking unexpected tools via step\-by\-step constraints, while still allowing autonomous parameter determination, keeping reasoning within a reasonable boundary, and preventing rapid context expansion that would otherwise cause the SLM to become lost\.

### Plan\-to\-Act

To prevent the SLM agent from becoming lost in tool\-invocation contexts when facing ambiguous complex tasks, our PTC\-Decoder adopts a "Plan\-to\-Act" decision paradigm \(left part in Figure[3](https://arxiv.org/html/2609.30836#Sx3.F3)\)\. Its core idea is that in each interaction round, the agent first generates a high level plan and then strictly follows this plan to invoke atomic tools step by step\. This design is particularly well\-suited for resource\-constrained edge deployment scenarios\.

Motivation: Resource Bottlenecks in Edge Computing\.Our SLM is deployed on a small remote\-sensing satellite with only 16GB of unified memory, of which less than 4GB is available for model inference\. Under such extreme constraints, directly confronting the agent with ambiguous, multi\-step, context\-dependent user requests leads to cognitive overload as the model attempts to simultaneously handle task decomposition, tool selection, parameter filling, and result interpretation, manifesting as invalid tool calls, missing critical steps, repetitive loops, or unparseable outputs\. LLM agent paradigms such as ReAct and Plan\-and\-Solve rely on hundred\-billion\-parameter models and multiple inference passes for robust planning\([Yao et al\. 2023](https://arxiv.org/html/2609.30836#bib.bib34);[Wang et al\. 2023](https://arxiv.org/html/2609.30836#bib.bib23)\), which is infeasible on resource\-constrained satellite platforms\.

Implementation: Planning as a Tool\.We define "plan" as a tool within the agent system: its input is a natural\-language task description, and its output is a structured sequence of atomic tool names\. The key design decision is that the first inference round must forcibly invoke the ‘plan‘ tool—rather than allowing free selection—which is achieved via the tool\-constrained decoder detailed in the next section\. After the plan returns a sequence, each subsequent agent loop invokes the corresponding atomic tools in order\. Formally, letting the full toolset be𝒯=\{plan,t2,t3,…,tn\}\\mathcal\{T\}=\\\{\\text\{plan\},t\_\{2\},t\_\{3\},\\dots,t\_\{n\}\\\}and the user task be𝒫\\mathcal\{P\},𝒞\\mathcal\{C\}be the request initiated by the SLM, and𝒮\\mathcal\{S\}the corresponding tool call schema, the first\-round outputx\(1\)x^\{\(1\)\}must satisfy:

∀𝒫,x\(1\)∈𝒞p​l​a​n=△\{𝒮\|𝒮\.name==plan\}\\forall\\mathcal\{P\},~x^\{\(1\)\}\\in\\mathcal\{C\}\_\{plan\}\\overset\{\\triangle\}\{=\}\\\{\\mathcal\{S\}~\|~\\mathcal\{S\}\.name==plan\\\}\(1\)

### Tool\-Constrained Decoder

Like LLMs, SLMs are modeled as causal language models, performing autoregressive sampling by continuously predicting the next token\([Vaswani et al\. 2017](https://arxiv.org/html/2609.30836#bib.bib9)\)\. Modern decoder\-only LLMs rely on autoregressive sampling from the model’s own probability distribution at each step, granting them strong open\-domain text generation capabilities\([Brown et al\. 2020](https://arxiv.org/html/2609.30836#bib.bib10)\)\. However, for SLMs deployed on offline edge devices, this unconstrained mechanism often fails—weak models may disregard prompt instructions for tool invocation, severely undermining their reliability and practicality in complex agent tasks\.

To address this, we propose theTool\-Constrained Decoder, which forcibly injects the desired tool during sampling, ensuring that the SLM generates exactly that tool invocation at the designated step\. This module is lightweight and plug\-and\-play: it requires no changes to existing agent implementations, no modifications to stable prompts, and no additional training\. Compared to SFT or RL approaches, our method is far less intrusive\.

Different from a general constraint decoder, this design avoids over\-constraint, preserving parameter generation quality and reasoning capability\. The module is plug\-and\-play: it can be disabled for unconstrained steps and enabled on demand\. The right part in Figure[3](https://arxiv.org/html/2609.30836#Sx3.F3)illustrates the decoder structure with the module activated\.

Let𝒱\\mathcal\{V\}denote the vocabulary,x<tx\_\{<t\}the generated sequence up to stept−1t\-1,lt∈ℝ\|𝒱\|l\_\{t\}\\in\\mathbb\{R\}^\{\|\\mathcal\{V\}\|\}the raw logits at steptt,sts\_\{t\}the internal state of the constrained decoder at steptt\(a finite\-state machine characterizing the preconditions satisfied by the current sequence\)

Let the original conditional distributionπt\\pi\_\{t\}of the language model be:

πt​\(v\)​=△​Pℒ​ℳ​\(xt=v\|x<t\)=softmax​\(lt\)​\[v\],v∈𝒱\\pi\_\{t\}\(v\)\\overset\{\\triangle\}\{=\}\\textbf\{P\}\_\{\\mathcal\{LM\}\}\(x\_\{t\}=v\|x\_\{<t\}\)=\\text\{softmax\}\(l\_\{t\}\)\[v\],~v\\in\\mathcal\{V\}\(2\)A generic constrained decoder can be formally defined as applying a mask functionMtM\_\{t\}to the raw logits via Hadamard product to modify the distribution before output of the final logits at time steptt\([Koo et al\. 2024](https://arxiv.org/html/2609.30836#bib.bib19)\):

lt^=lt⊙Mt\\hat\{l\_\{t\}\}=l\_\{t\}\\odot M\_\{t\}\(3\)wherelt^\\hat\{l\_\{t\}\}represents the logits modified by the mask, which corrects the raw scores of all tokens violating the rules to−∞\-\\inftyso that their corresponding token probabilities approach 0 infinitely after the softmax operation; the final corrected token probability distributionπ^t​\(v\)\\hat\{\\pi\}\_\{t\}\(v\)is given by:

π^t​\(v\)​=△​Pℒ​ℳ​\(xt=v\|x<t\)=softmax​\(lt^\)​\[v\],v∈𝒱\\hat\{\\pi\}\_\{t\}\(v\)\\overset\{\\triangle\}\{=\}\\textbf\{P\}\_\{\\mathcal\{LM\}\}\(x\_\{t\}=v\|x\_\{<t\}\)=\\text\{softmax\}\(\\hat\{l\_\{t\}\}\)\[v\],~v\\in\\mathcal\{V\}\(4\)
We design an agent\-oriented constrained decoder \(Tool\-Constrained Decoder\)\. Unlike approaches that constrain the entire output space, our decoder imposes token\-level hard constraints only on tool invocation format and tool names, leaving parameter content entirely to the underlying SLM\. This design aligns with the plan\-to\-act paradigm: we enforce the use of designated tools, but delegate the generation of action details to the model’s own knowledge and contextual reasoning\.

To implement this partial constraint, we construct a prefix\-constrained automaton𝒜t​o​o​l​\(Q,Σ,δ,q0,F\)\\mathcal\{A\}\_\{tool\}\(Q,\\Sigma,\\delta,q\_\{0\},F\)based on a Deterministic Finite Automaton \(DFA\), where:

- •Q=Qf​i​x∪\{qf​r​e​e\}Q=Q\_\{fix\}\\cup\\\{q\_\{free\}\\\}, whereQf​i​xQ\_\{fix\}is the state sequence for the fixed prefixes, andqf​r​e​eq\_\{free\}is the state sequence for subsequent unconstrained tokens\.
- •Σ\\Sigmais the set of all tokens in the vocabulary𝒱\\mathcal\{V\}\.
- •OnQf​i​xQ\_\{fix\}, the transition functionδ\\deltaforces the generation of preset tokens according to specified rules, such as<\|name\>get\_weather<name\|\>among<\|tool\><\|name\>get\_weather<name\|\> <\|arguments\>\.\.\.<arguments\|\><tool\|\>; once the fixed sequence generation is complete, the automaton transitions fromQf​i​xQ\_\{fix\}to the absorbing stateqf​r​e​eq\_\{free\}via any token, andqf​r​e​eq\_\{free\}has self\-loop transitions for any token inΣ\\Sigma, i\.e\.,∀v∈𝒱,δ⁡\(qf​r​e​e,v\)=qf​r​e​e\\forall v\\in\\mathcal\{V\},\\delta\(q\_\{free\},v\)=q\_\{free\}\.
- •F∈qf​r​e​eF\\in q\_\{free\}, where the terminal state is reached through the EOS token, i\.e\.,δ⁡\(qf​r​e​e,ve​o​s\)=F\\delta\(q\_\{free\},v\_\{eos\}\)=F\.

Table 2:LLM\-judge scores across experimental conditions, with a Human double\-check \(H\.\) for each group \(20 randomly sampled tasks per group, rated on the same\[0,5\]\[0,5\]by two annotators\)\. The best LLM\-judge value for each metric isbolded\.Based on the aforementioned automaton𝒜t​o​o​l\\mathcal\{A\}\_\{tool\}, we define the mask functionMt′\(v\)M\_\{t\}^\{\{\}^\{\\prime\}\}\(v\)at the time stepttas:

Mt′\(v\)=\{0,if​δ​\(qt,v\)​is defined,−sgn\(lt\[v\]\)⋅∞,otherwise\.M\_\{\\text\{t\}\}^\{\{\}^\{\\prime\}\}\(v\)=\\begin\{cases\}0,&\\text\{if \}\\delta\(q\_\{t\},v\)\\text\{ is defined\},\\\\ \-\\operatorname\{sgn\}\(l\_\{t\}\[v\]\)\\cdot\\infty,&\\text\{otherwise\}\.\\end\{cases\}\(5\)whereqtq\_\{t\}is the state of the automaton at time steptt, and the modified logits and probability distribution formally still follow:

lt′^=lt⊙Mt′\\hat\{l\_\{t\}^\{\{\}^\{\\prime\}\}\}=l\_\{t\}\\odot M\_\{t\}^\{\{\}^\{\\prime\}\}\(6\)π^t′\(v\)=△Pℒ​ℳ′\(xt=v\|x<t\)=softmax\(lt′^\)\[v\],v∈𝒱\\hat\{\\pi\}^\{\{\}^\{\\prime\}\}\_\{t\}\(v\)\\overset\{\\triangle\}\{=\}\\textbf\{P\}^\{\{\}^\{\\prime\}\}\_\{\\mathcal\{LM\}\}\(x\_\{t\}=v\|x\_\{<t\}\)=\\text\{softmax\}\(\\hat\{l\_\{t\}^\{\{\}^\{\\prime\}\}\}\)\[v\],~v\\in\\mathcal\{V\}\(7\)

## Experiments

### Experimental Setup

Table 3:Recall, Precision, F1 \(per\-task average\), and average tokens \(input \+ output\)\. The best value for each metric isbolded\.Edge Device\.All experiments are conducted on an NVIDIA Jetson Orin NX Developer Kit, which shares the same configuration as the control computer of a small offline remote\-sensing satellite: an Orin GPU, 8\-core Cortex\-A78AE CPU, 16GB unified memory, 877GB storage, running Ubuntu 22\.04 LTS with CUDA 12\.2\. To closely replicate the extreme resource constraints of actual satellite operations, we concurrently run multiple remote\-sensing algorithms and daemon processes alongside the Agent tasks on this device, thereby approximating real onboard resource\-limited conditions\.

Small Language Models and Setup\.To demonstrate the generalizability of this study, we select mainstream SLMs from the open\-source community, with memory limitation that the total effective parameter count does not exceed 2B, and that they are released by well\-known model providers\. Ultimately, we identified a total of 7 SLMs\([Team et al\. 2025](https://arxiv.org/html/2609.30836#bib.bib31);[Team et al\. 2026](https://arxiv.org/html/2609.30836#bib.bib37);[Yang et al\. 2025](https://arxiv.org/html/2609.30836#bib.bib38);[Guo et al\. 2025](https://arxiv.org/html/2609.30836#bib.bib39)\)\. All models use temperature 0\.2, top\-p 0\.95, and a maximum of 10 turns per task \(overflowed as a failure\), with up to 512 new tokens per free\-generation step and 256 per constrained tool call\.

Evaluation Datasets\.To validate the optimization strategies for our remote\-sensing satellite Agent system, we conducted on\-board experimental evaluations\. Given that spaceborne edge devices offer AI compute below 100 TOPS and suffer from thermal throttling under prolonged operation, we carefully controlled the evaluation dataset scale to ensure smooth comparative experiments across 7 models\. To address the absence of dedicated remote\-sensing Agent benchmarks, we curated 200 real user tasks from on\-board operations and had them cross\-annotated by three remote\-sensing and computer vision experts, producing a test suite with expected Agent execution trajectories and ground\-truth results\. This dataset will be released along with the code to facilitate future research in this domain\.

Evaluation Methods and Metrics\.Our evaluation combines LLM\-judge scores \(DeepSeek\-V4\-flash\)\([DeepSeek\-AI et al\. 2026](https://arxiv.org/html/2609.30836#bib.bib36)\)with rule\-based metrics\. LLM\-judge includes Overall \(Ov\.\), which integrates complete tracks and logs for scoring; Result Accuracy \(R\.A\.\); Flow Reasonableness \(F\.R\.\); and Robustness \(Rb\.\), all rated on a\[0,5\]\[0,5\]scale\. Rule\-based metrics include Recall \(R\.\), Precision \(P\.\), and F1\-score, computed per task against benchmark ground\-truth tool sets and averaged \(excluded non\-functional tools liketask\_done\)\. Due to thermal throttling from sustained inference on satellite onboard systems, each task is executed only once\. Therefore, we group 200 tasks as paired units and employ the Wilcoxon signed\-rank testppto assess the statistical significance of the observed improvements\([Demšar 2006](https://arxiv.org/html/2609.30836#bib.bib44)\)\. For LLM\-Judge, we will randomly pick 20 results to do a manually double\-check for Ov\., marked as \(H\.\)\. All scores are reported as means across 200 benchmark tasks\.

### Overall Comparison

#### LLM\-Judge Evaluation

Baseline performance is limited: the strongest model \(Qwen3\-1\.7B\) achieves Ov\. 1\.98, while weaker models score below 0\.75 across all dimensions\.

PTC\-Decoder consistently improves all models \(mean Ov\. \+1\.21,p<0\.01p<0\.01\), with gains inversely correlated to baseline capability—DeepSeek\-R1\-1\.5B improves by 350% \(0\.40→\\rightarrow1\.80\), while Qwen3\-1\.7B gains 48% \(1\.98→\\rightarrow2\.94\)\. Flow reasonableness shows the largest average gain \(\+1\.35\); Gemma\-4\-2B reaches the highest F\.R\. of 3\.31, and Qwen3\.5\-2B gains \+2\.08 to 3\.27\. Result accuracy improves by a mean of \+1\.27, with Qwen3\.5\-2B achieving both the largest gain \(\+2\.40, 0\.41→\\rightarrow2\.81\) and the global best R\.A\. Robustness gains are moderate \(mean \+0\.52\), with Qwen3\-1\.7B attaining the best Rb\. of 3\.06 under PTC\-Decoder\. Qwen3\.5\-2B obtains the highest Ov\. of 3\.12\.

Plan w/o TC\-Decoder yields marginal Ov\. improvement \(mean \+0\.22\), with Gemma\-4\-2B \(–0\.35\) degrading\. R\.A\. and F\.R\. follow similarly, with mean R\.A\. declining by 0\.15\. Rb\. increases by a mean of \+0\.23, though Qwen3\-1\.7B’s Rb\. is slightly higher under PTC\-Decoder \(3\.06 vs\. 3\.03\), contrasting with earlier observations where free execution gave the best Rb\. Human double\-check shows the same trends\.

#### Rule\-Based Evaluation

Table[3](https://arxiv.org/html/2609.30836#Sx4.T3)reports benchmark Recall, Precision, and F1\. Baseline performance is limited \(mean Recall 0\.115, F1 0\.133\), with Gemma\-3\-1B and DeepSeek\-R1\-1\.5B scoring zero across all metrics\.

PTC\-Decoder consistently improves all three metrics: mean Recall \+0\.126 \(Qwen3\.5\-2B leading: 0\.123→\\rightarrow0\.388\), Precision \+0\.058, and F1 \+0\.096 \(Gemma\-4\-2B reaching global best F1=0\.359\)\. Weak models recover from zero to 0\.063 and 0\.074 Recall\. Absolute F1 remains modest across conditions, reflecting the difficulty of full tool coverage, but gains are consistent across all models\.

Plan w/o TC\-Decoder yields mixed results: mean Recall and F1 decline by 0\.020 and 0\.013, suggesting that unenforced planning may reduce ground\-truth coverage\. Qwen3\-1\.7B is the sole improver \(F1 \+0\.106, Precision 0\.647, global best\), albeit with narrow Recall \(0\.252\)\. Qwen3\-0\.6B collapses to zero on all metrics\. We specifically examined this set of zero\-score cases\. Under the plan\-only condition without TC\-Decoder, Qwen3\-0\.6B simply treated the task as completed in all instances, which directly accounts for its zero scores across all metrics\.

As Gemma\-3\-1B and DeepSeek\-R1\-1\.5B are not tool\-invocation models, they cannot understand tools\. Manual trace analysis confirms that, without TC\-Decoder, their R\., P\., and F1 on remote\-sensing tasks are zero \(these metrics exclude non\-functional tools liketask\_done\)\.

Figure 4:Recall, Precision, and F1 across three external datasets \(means over 7 SLMs\)\. PTC\-Decoder consistently improves all metrics\. Plan w/o TC\-Decoder shows inconsistent improvement, remaining close to baseline\. Error bars omitted for clarity; within\-group F1 SD ranges: ToolAlpaca 0\.18–0\.38, Seal\-Tools 0\.17–0\.44, API\-Bank 0\.11–0\.37\.∗p<0\.05\*p<0\.05,†p=0\.06\{\}^\{\\dagger\}p=0\.06\(Wilcoxon signed\-rank test,n=7n=7, PTC\-Decoder vs\. Baseline F1: ToolAlpacap=0\.03p=0\.03, Seal\-Toolsp=0\.03p=0\.03, API\-Bankp=0\.06p=0\.06\)\.

### Performance on Other Datasets

To test cross\-domain generalizability, we evaluate the same three experimental combinations on ToolAlpaca\([Tang et al\. 2023](https://arxiv.org/html/2609.30836#bib.bib41)\), Seal\-Tools\([Wu et al\. 2025b](https://arxiv.org/html/2609.30836#bib.bib42)\), and API\-Bank\([Li et al\. 2023b](https://arxiv.org/html/2609.30836#bib.bib43)\), using the same 7 SLMs and experimental configs mentioned before\. The overall results are shown in Figure[4](https://arxiv.org/html/2609.30836#Sx4.F4)\. PTC\-Decoder consistently improves over baseline across all datasets: ToolAlpaca F1 38% to 53% \(Recall 54% to 60%, Precision 35% to 51%\); Seal\-Tools 33% to 68% \(38% to 68%, 35% to 68%\); API\-Bank 31% to 50% \(29% to 56%, 35% to 48%\)\. The largest gain appears on Seal\-Tools, where F1 more than doubles—driven by balanced improvements in both Recall and Precision, confirming that TC\-Decoder enhances tool coverage while reducing off\-target invocations\.

Plan w/o TC\-Decoder yields weaker results: ToolAlpaca F1 stagnates at 39% \(baseline 38%\), Seal\-Tools reaches 46% \(well below 68%\), and API\-Bank only 37% \(baseline 31%\)\. These results reinforce that plan\-level constraints alone are insufficient for most SLMs; execution\-level enforcement is the primary driver of the observed gains\.

### Ablation Study

To isolate TC\-Decoder’s contribution, we conduct an ablation study directly comparing PTC\-Decoder against Plan w/o TC\-Decoder, where the latter ablated TC\-Decoder\. Tables[2](https://arxiv.org/html/2609.30836#Sx3.T2)and[3](https://arxiv.org/html/2609.30836#Sx4.T3)report the full results\.

LLM\-Judge metrics\.Adding TC\-Decoder improves Ov\. by 0\.99 on average across all 7 models \(p<0\.01p<0\.01, bootstrap 95% CI \[\+0\.91, \+1\.06\]\), with Qwen3\.5\-2B \(\+1\.78\) and Gemma\-4\-2B \(\+1\.69\) gaining the most\. R\.A\. benefits substantially \(mean \+1\.42\): Qwen3\.5\-2B rises from 0\.30 to 2\.81 \(Δ=\+2\.51\\Delta=\+2\.51\), and Gemma\-4\-2B from 0\.22 to 2\.42 \(Δ=\+2\.20\\Delta=\+2\.20\)\. F\.R\. improves by 1\.00 on average; Gemma\-4\-2B reaches 3\.31 \(Δ=\+1\.86\\Delta=\+1\.86\), and Qwen3\.5\-2B 3\.27 \(Δ=\+1\.86\\Delta=\+1\.86\)\. Rb\. shows the smallest gain \(mean \+0\.23\), with Qwen3\-1\.7B’s Rb\. nearly unchanged \(3\.06 vs\. 3\.03\), suggesting that TC\-Decoder neither improves nor harms recovery behavior\.

Benchmark coverage\.TC\-Decoder consistently improves benchmark\-aligned tool coverage\. Mean Recall rises from 0\.095 to 0\.241 \(Δ=\+0\.146\\Delta=\+0\.146\), led by Gemma\-4\-2B \(\+0\.282\) and Qwen3\.5\-2B \(\+0\.241, achieving the globally best Recall of 0\.388\)\. F1 increases by 0\.110 on average; Gemma\-4\-2B reaches the globally best F1 of 0\.359 \(\+0\.250 over its Plan w/o TC\-Decoder counterpart\)\. The most pronounced effects are on models that collapse without TC\-Decoder: Qwen3\-0\.6B’s, Gemma\-3\-1B’s, and DeepSeek\-R1\-1\.5B’s benchmark metrics all register zero under free execution, yet reach 0\.246, 0\.076, and 0\.088 F1 with TC\-Decoder\. Qwen3\-1\.7B is an interesting exception: its F1 is higherwithoutTC\-Decoder \(0\.355 vs\. 0\.292\), driven by Precision of 0\.647 \(globally best\) under free execution, at the cost of Recall \(0\.252 vs\. 0\.305 with TC\-Decoder\)\. This precision\-recall trade\-off illustrates that TC\-Decoder’s primary mechanism is enforcing broader tool coverage, which may modestly reduce precision for models already capable of focused execution\.

Efficiency\.Token overhead is nearly identical between the two variants \(2\.07×\\timesvs\. 2\.08×\\timesover baseline\), confirming that TC\-Decoder adds negligible computational cost\.

In summary, TC\-Decoder is the decisive component of PTC\-Decoder: it substantially improves LLM\-judge scores and benchmark tool coverage across all models, transforms several SLMs from zero coverage to functional plan adherence, and does so without additional overhead\.

## Conclusion

This paper designed PTC\-Decoder, a training\-free, plug\-and\-play decoder framework that addresses the intelligence deficiency of SLMs on offline resource\-constrained edge devices\. By coupling a Plan\-to\-Act paradigm with TC\-Decoder, PTC\-Decoder elevates 7 SLMs from weak baselines to functional agent performance, achieving a mean Ov\. gain of \+1\.21 and peak F1 of 0\.359\. Ablation confirms TC\-Decoder as the decisive component\. Removing it reduces F1 by 0\.110 on average and degrades all quality metrics, yet offers no efficiency advantage\. PTC\-Decoder thus provides a lightweight and effective solution for improving step\-level reliability of SLM agents on offline edge devices\.

### Limitations

Despite improvements in all dimensions, R\.A\. remains the weakest dimension \(best 2\.81/5\), and Rb\. gains are modest \(mean \+0\.52 vs\. baseline\)\. This suggests that TC\-Decoder cannot guarantee correct answers or robust error recovery, since it only constrains tool names\. Extending constraints to parameter schemas may help narrow this gap\. PTC\-Decoder incurs a 2\.07×\\timestoken overhead\. As a counterexample, Qwen3\-1\.7B achieves higher F1 without TC\-Decoder, indicating that PTC\-Decoder primarily benefits intelligence\-deficient SLMs\. As SLM capabilities advance, PTC\-Decoder may become less necessary\. For LLM, the creativity\-restricting effect of TC\-Decoder might outweigh its benefits\.

## References

- Brownet al\.\(2020\)T\. B\. Brown, B\. Mann, N\. Ryder, M\. Subbiah, J\. Kaplan, P\. Dhariwal, A\. Neelakantan,et al\.Language models are few\-shot learners\.InConference on Neural Information Processing Systems,Cited by:[Application of SLMs on Edge Devices](https://arxiv.org/html/2609.30836#Sx2.SSx1.p1.1),[Tool\-Constrained Decoder](https://arxiv.org/html/2609.30836#Sx3.SSx2.p1.1)\.
- Chenet al\.\(2025\)J\. Chen, A\. Myrzakhan, Y\. Luo, H\. M\. Khan, S\. M\. Bsharat, and Z\. ShenDRAG: distilling RAG for SLMs from LLMs to transfer knowledge and mitigate hallucination via evidence and graph\-based distillation\.InAnnual Meeting of the Association for Computational Linguistics,pp\. 7240–7260\.External Links:[Document](https://dx.doi.org/10.18653/v1/2025.acl-long.358)Cited by:[Introduction](https://arxiv.org/html/2609.30836#Sx1.p2.1)\.
- DeepSeek\-AIet al\.\(2026\)DeepSeek\-AI, A\. Xu, B\. Lin, B\. Xue, B\. Wang, B\. Xu, B\. Wu,et al\.DeepSeek\-v4: towards highly efficient million\-token context intelligence\.External Links:2606\.19348Cited by:[Experimental Setup](https://arxiv.org/html/2609.30836#Sx4.SSx1.p4.1)\.
- Demšar \(2006\)J\. DemšarStatistical comparisons of classifiers over multiple data sets\.Journal of Machine Learning Research7,pp\. 1–30\.Cited by:[Experimental Setup](https://arxiv.org/html/2609.30836#Sx4.SSx1.p4.1)\.
- Erdoganet al\.\(2024\)L\. E\. Erdogan, N\. Lee, S\. Jha, S\. Kim, R\. Tabrizi, S\. Moon, C\. R\. C\. Hooper, G\. Anumanchipalli, K\. Keutzer, and A\. GholamiTinyAgent: function calling at the edge\.InConference on Empirical Methods in Natural Language Processing,pp\. 80–88\.External Links:[Document](https://dx.doi.org/10.18653/v1/2024.emnlp-demo.9)Cited by:[Introduction](https://arxiv.org/html/2609.30836#Sx1.p2.1)\.
- Guoet al\.\(2025\)D\. Guo, D\. Yang, H\. Zhang, J\. Song, P\. Wang, Q\. Zhu,et al\.DeepSeek\-r1 incentivizes reasoning in llms through reinforcement learning\.Nature645\(8081\),pp\. 633–638\.External Links:[Document](https://dx.doi.org/10.1038/s41586-025-09422-z)Cited by:[Experimental Setup](https://arxiv.org/html/2609.30836#Sx4.SSx1.p2.1)\.
- Haoet al\.\(2024\)Z\. Hao, H\. Jiang, S\. Jiang, J\. Ren, and T\. CaoHybrid slm and llm for edge\-cloud collaborative inference\.InProc\. Workshop Edge Mob\. Found\. Models,pp\. 36 – 41\.Cited by:[Introduction](https://arxiv.org/html/2609.30836#Sx1.p4.1)\.
- Hariharan Samson \(2026\)H\. Hariharan SamsonLightweight transformer architectures for edge devices in real\-time applications\.Note:arxiv:2601\.03290Cited by:[Lightweight Architectures and Model Quantization](https://arxiv.org/html/2609.30836#Sx2.SSx2.p1.1)\.
- Honget al\.\(2025\)M\. Hong, C\. J\. Zhang, D\. Jiang, and Y\. HeAugmenting compliance\-guaranteed customer service chatbots: context\-aware knowledge expansion with large language models\.InConference on Empirical Methods in Natural Language Processing, Industry Track,pp\. 753–765\.External Links:[Document](https://dx.doi.org/10.18653/v1/2025.emnlp-industry.51)Cited by:[Introduction](https://arxiv.org/html/2609.30836#Sx1.p1.1)\.
- Honget al\.\(2024\)S\. Hong, M\. Zhuge, J\. Chen, X\. Zheng, Y\. Cheng, J\. Wang, C\. Zhang, S\. Yau, Z\. Lin, L\. Zhou,et al\.MetaGPT: meta programming for a multi\-agent collaborative framework\.InInternational Conference on Learning Representations,Vol\.2024,pp\. 23247–23275\.Cited by:[Edge\-Cloud Collaboration](https://arxiv.org/html/2609.30836#Sx2.SSx3.p1.1)\.
- Kooet al\.\(2024\)T\. Koo, F\. Liu, and L\. HeAutomata\-based constraints for language model decoding\.InConf\. Lang\. Model\.,Cited by:[Introduction](https://arxiv.org/html/2609.30836#Sx1.p5.1),[Tool\-Constrained Decoder](https://arxiv.org/html/2609.30836#Sx3.SSx2.p5.2)\.
- Liet al\.\(2025a\)B\. Li, Y\. Wang, H\. Ma, L\. Chen, J\. Xiao, and S\. WangMobiLoRA: accelerating LoRA\-based LLM inference on mobile devices via context\-aware KV cache optimization\.InAnnual Meeting of the Association for Computational Linguistics,pp\. 23400–23410\.External Links:[Document](https://dx.doi.org/10.18653/v1/2025.acl-long.1140)Cited by:[Lightweight Architectures and Model Quantization](https://arxiv.org/html/2609.30836#Sx2.SSx2.p1.1)\.
- Liet al\.\(2023a\)G\. Li, H\. Hammoud, H\. Itani, D\. Khizbullin, and B\. GhanemCAMEL: communicative agents for "mind" exploration of large language model society\.InConference on Neural Information Processing Systems,Vol\.36,pp\. 51991–52008\.Cited by:[Edge\-Cloud Collaboration](https://arxiv.org/html/2609.30836#Sx2.SSx3.p1.1)\.
- Liet al\.\(2023b\)M\. Li, Y\. Zhao, B\. Yu, F\. Song, H\. Li, H\. Yu, Z\. Li, F\. Huang, and Y\. LiAPI\-bank: a comprehensive benchmark for tool\-augmented LLMs\.InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing,pp\. 3102–3116\.External Links:[Link](https://aclanthology.org/2023.emnlp-main.187/),[Document](https://dx.doi.org/10.18653/v1/2023.emnlp-main.187)Cited by:[Performance on Other Datasets](https://arxiv.org/html/2609.30836#Sx4.SSx3.p1.1)\.
- Liet al\.\(2025b\)S\. Li, H\. Wang, W\. Xu, R\. Zhang, S\. Guo, J\. Yuan, X\. Zhong, T\. Zhang, and R\. LiCollaborative inference and learning between edge slms and cloud llms: a survey of algorithms, execution, and open challenges\.Note:arxiv:2507\.16731Cited by:[Introduction](https://arxiv.org/html/2609.30836#Sx1.p4.1)\.
- Liet al\.\(2026\)X\. Li, H\. Li, C\. Sun, Q\. Fan, Z\. Han, and V\. C\. M\. LeungEdge\-enhanced intelligence: a comprehensive survey of large language models and edge\-cloud computing synergy\.IEEE Communications Surveys and Tutorials28\(\),pp\. 1248–1284\.External Links:[Document](https://dx.doi.org/10.1109/COMST.2025.3587225)Cited by:[Application of SLMs on Edge Devices](https://arxiv.org/html/2609.30836#Sx2.SSx1.p1.1),[Edge\-Cloud Collaboration](https://arxiv.org/html/2609.30836#Sx2.SSx3.p1.1)\.
- Luet al\.\(2025a\)P\. Lu, I\. Kobyzev, M\. Rezagholizadeh, B\. Chen, and P\. LanglaisReGLA: refining gated linear attention\.InProc\. Conf\. Nations Americas Chapter Assoc\. Comput\. Linguist\.: Hum\. Lang\.,pp\. 2884–2898\.External Links:[Document](https://dx.doi.org/10.18653/v1/2025.naacl-long.147)Cited by:[Introduction](https://arxiv.org/html/2609.30836#Sx1.p3.1)\.
- Luet al\.\(2025b\)Z\. Lu, X\. Li, D\. Cai, R\. Yi, F\. Liu, W\. Liu, J\. Luan, X\. Zhang, N\. D\. Lane, and M\. XuDemystifying small language models for edge deployment\.InAnnual Meeting of the Association for Computational Linguistics,pp\. 14747–14764\.External Links:[Document](https://dx.doi.org/10.18653/v1/2025.acl-long.718)Cited by:[Application of SLMs on Edge Devices](https://arxiv.org/html/2609.30836#Sx2.SSx1.p1.1)\.
- Mishraet al\.\(2026\)P\. P\. Mishra, K\. P\. Yeole, R\. Keshavamurthy, M\. B\. Surana, and F\. SaraylooA systematic framework for enterprise knowledge retrieval: leveraging llm\-generated metadata to enhance rag systems\.InIEEE Conference on Artificial Intelligence, in press,External Links:[Document](https://dx.doi.org/10.48550/arXiv.2512.05411)Cited by:[Introduction](https://arxiv.org/html/2609.30836#Sx1.p1.1)\.
- Niet al\.\(2026\)X\. Ni, J\. Wang, L\. Yang, Y\. Lu, H\. Chen, R\. Liu, and J\. HaoFollowing the navigation: enhancing small language models contextual reasoning with LLM guidance\.InThe Fourteenth International Conference on Learning Representations,Cited by:[Application of SLMs on Edge Devices](https://arxiv.org/html/2609.30836#Sx2.SSx1.p1.1)\.
- Qianet al\.\(2023\)C\. Qian, W\. Liu, H\. Liu, N\. Chen, Y\. Dang, J\. Li, C\. Yang, W\. Chen, Y\. Su, X\. Cong, J\. Xu, D\. Li, Z\. Liu, and M\. SunChatDev: communicative agents for software development\.InAnnual Meeting of the Association for Computational Linguistics,Cited by:[Edge\-Cloud Collaboration](https://arxiv.org/html/2609.30836#Sx2.SSx3.p1.1)\.
- Shaoet al\.\(2024\)Z\. Shao, P\. Wang, Q\. Zhu, R\. Xu, J\. Song, M\. Zhang, Y\. K\. Li, Y\. Wu, and D\. GuoDeepSeekMath: pushing the limits of mathematical reasoning in open language models\.Vol\.abs/2402\.03300\.Note:arxiv:2402\.03300Cited by:[Introduction](https://arxiv.org/html/2609.30836#Sx1.p3.1)\.
- Sharma and Mehta \(2025\)R\. Sharma and M\. MehtaSmall language models for agentic systems: a survey of architectures, capabilities, and deployment trade offs\.Note:arxiv:2510\.03847Cited by:[Introduction](https://arxiv.org/html/2609.30836#Sx1.p2.1)\.
- Sunet al\.\(2025\)C\. Sun, Y\. Li, D\. Wu, and B\. BouletOnionEval: an unified evaluation of fact\-conflicting hallucination for small\-large language models\.Note:arxiv:2501\.12975Cited by:[Introduction](https://arxiv.org/html/2609.30836#Sx1.p3.1)\.
- Tanget al\.\(2023\)Q\. Tang, Z\. Deng, H\. Lin, X\. Han, Q\. Liang, and L\. SunToolAlpaca: generalized tool learning for language models with 3000 simulated cases\.External Links:2306\.05301Cited by:[Performance on Other Datasets](https://arxiv.org/html/2609.30836#Sx4.SSx3.p1.1)\.
- Teamet al\.\(2026\)G\. Team, S\. El Abd, V\. Aggarwal, R\. Algayres, A\. Andreev, O\. Bachem, I\. Ballantyne,et al\.Gemma 4 technical report\.External Links:2607\.02770Cited by:[Experimental Setup](https://arxiv.org/html/2609.30836#Sx4.SSx1.p2.1)\.
- Teamet al\.\(2025\)G\. Team, A\. Kamath, J\. Ferret, S\. Pathak, N\. Vieillard, R\. Merhej, S\. Perrin, T\. Matejovicova, A\. Ramé, M\. Rivière,et al\.Gemma 3 technical report\.Note:arxiv:2503\.19786Cited by:[Introduction](https://arxiv.org/html/2609.30836#Sx1.p3.1),[Experimental Setup](https://arxiv.org/html/2609.30836#Sx4.SSx1.p2.1)\.
- Vaswaniet al\.\(2017\)A\. Vaswani, N\. Shazeer, N\. Parmar, J\. Uszkoreit, L\. Jones, A\. N\. Gomez, Ł\. Kaiser, and I\. PolosukhinAttention is all you need\.InConference on Neural Information Processing Systems,pp\. 6000–6010\.Cited by:[Application of SLMs on Edge Devices](https://arxiv.org/html/2609.30836#Sx2.SSx1.p1.1),[Tool\-Constrained Decoder](https://arxiv.org/html/2609.30836#Sx3.SSx2.p1.1)\.
- Wanget al\.\(2021\)J\. Wang, Z\. Zheng, A\. Ma, X\. Lu, and Y\. ZhongLoveDA: a remote sensing land\-cover dataset for domain adaptive semantic segmentation\.InProceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks,Vol\.1,pp\.\.Cited by:[Methodology](https://arxiv.org/html/2609.30836#Sx3.p1.1)\.
- Wanget al\.\(2023\)L\. Wang, W\. Xu, Y\. Lan, Z\. Hu, Y\. Lan, R\. K\.\-W\. Lee, and E\.\-P\. LimPlan\-and\-solve prompting: improving zero\-shot chain\-of\-thought reasoning by large language models\.InAnnual Meeting of the Association for Computational Linguistics,pp\. 2609–2634\.External Links:[Document](https://dx.doi.org/10.18653/v1/2023.acl-long.147)Cited by:[Introduction](https://arxiv.org/html/2609.30836#Sx1.p2.1),[Plan\-to\-Act](https://arxiv.org/html/2609.30836#Sx3.SSx1.p2.1)\.
- Wanget al\.\(2025\)R\. Wang, Z\. Gao, L\. Zhang, S\. Yue, and Z\. GaoEmpowering large language models to edge intelligence: a survey of edge efficient llms and techniques\.Computer Science Review57,pp\. 100755\.External Links:ISSN 1574\-0137,[Document](https://dx.doi.org/https%3A//doi.org/10.1016/j.cosrev.2025.100755)Cited by:[Application of SLMs on Edge Devices](https://arxiv.org/html/2609.30836#Sx2.SSx1.p1.1)\.
- Wanget al\.\(2024\)X\. Wang, R\. Yu, D\. Yang, and G\. XueInfiltrating the Sky: Data Delay and Overflow Attacks in Earth Observation Constellations\.InProc\. Int\. Conf\. Netw\. Protoc\. ICNP,Vol\.,pp\. 1–11\.External Links:ISSN,[Document](https://dx.doi.org/10.1109/ICNP61940.2024.10858559)Cited by:[Introduction](https://arxiv.org/html/2609.30836#Sx1.p4.1)\.
- Weiet al\.\(2022\)J\. Wei, X\. Wang, D\. Schuurmans, M\. Bosma, B\. Ichter, F\. Xia, E\. H\. Chi, Q\. V\. Le, and D\. ZhouChain\-of\-thought prompting elicits reasoning in large language models\.InConference on Neural Information Processing Systems,Cited by:[Application of SLMs on Edge Devices](https://arxiv.org/html/2609.30836#Sx2.SSx1.p1.1)\.
- Wuet al\.\(2025a\)B\. Wu, D\. Tipper, and P\. ZhouNear\-realtime earth observation via starlink leo satellite constellation\.Note:arxiv:2508\.10338Cited by:[Introduction](https://arxiv.org/html/2609.30836#Sx1.p4.1)\.
- Wuet al\.\(2025b\)M\. Wu, T\. Zhu, H\. Han, C\. Tan, X\. Zhang, and W\. ChenSeal\-tools: self\-instruct tool learning dataset for agent tuning and detailed benchmark\.InNatural Language Processing and Chinese Computing \(NLPCC 2024\),Vol\.15360\.External Links:[Document](https://dx.doi.org/10.1007/978-981-97-9434-8%5F29)Cited by:[Performance on Other Datasets](https://arxiv.org/html/2609.30836#Sx4.SSx3.p1.1)\.
- Yanget al\.\(2025\)A\. Yang, A\. Li, B\. Yang, B\. Zhang, B\. Hui, B\. Zheng,et al\.Qwen3 technical report\.External Links:2505\.09388Cited by:[Experimental Setup](https://arxiv.org/html/2609.30836#Sx4.SSx1.p2.1)\.
- Yaoet al\.\(2023\)S\. Yao, J\. Zhao, D\. Yu, N\. Du, I\. Shafran, K\. Narasimhan, and Y\. CaoReAct: synergizing reasoning and acting in language models\.InInternational Conference on Learning Representations \(ICLR\),Cited by:[Plan\-to\-Act](https://arxiv.org/html/2609.30836#Sx3.SSx1.p2.1)\.
- Yiet al\.\(2026\)B\. Yi, X\. Hu, Y\. Chen, S\. Zhang, H\. Yang, and F\. WuEcoAgent: an efficient device\-cloud collaborative multi\-agent framework for mobile automation\.AAAI Conference on Artificial Intelligence40\(35\),pp\. 29838–29846\.External Links:[Document](https://dx.doi.org/10.1609/aaai.v40i35.40230)Cited by:[Introduction](https://arxiv.org/html/2609.30836#Sx1.p4.1)\.
- Yuet al\.\(2026\)Y\. Yu, J\. Liu, Q\. Wu, H\. Wang, and J\. PeiSWAA: sliding window attention adaptation for efficient and quality preserving long context processing\.Note:arxiv:2512\.10411Cited by:[Introduction](https://arxiv.org/html/2609.30836#Sx1.p3.1)\.
- Yuet al\.\(2024\)Y\. Yu, G\. Rong, H\. Shen, H\. Zhang, D\. Shao, M\. Wang, Z\. Wei, Y\. Xu, and J\. WangFine\-tuning large language models to improve accuracy and comprehensibility of automated code review\.ACM Transactions on Software Engineering and Methodology34\(1\),pp\. 1–26\.External Links:[Document](https://dx.doi.org/10.1145/3695993)Cited by:[Introduction](https://arxiv.org/html/2609.30836#Sx1.p1.1)\.
- Zhaoet al\.\(2026\)W\. X\. Zhao, K\. Zhou, J\. Y\. Li, T\. Y\. Tang, X\. L\. Wang, Y\. P\. Hou, Y\. Q\. Min, B\. C\. Zhang, J\. J\. Zhang, Z\. C\. Dong,et al\.A survey of large language models\.Note:arXiv:2303\.18223External Links:[Document](https://dx.doi.org/10.48550/arXiv.2303.18223)Cited by:[Introduction](https://arxiv.org/html/2609.30836#Sx1.p1.1)\.
- Zhenget al\.\(2025\)Y\. Zheng, Y\. Chen, B\. Qian, X\. Shi, Y\. Shu, and J\. ChenA review on edge large language models: design, execution, and applications\.ACM Comput\. Surv\.57\(8\)\.External Links:ISSN 0360\-0300,[Document](https://dx.doi.org/10.1145/3719664)Cited by:[Application of SLMs on Edge Devices](https://arxiv.org/html/2609.30836#Sx2.SSx1.p1.1)\.
- Zhouet al\.\(2025\)Z\. Zhou, S\. Kurz, and Z\. ZhaoRevisiting pruning vs quantization for small language models\.InConference on Empirical Methods in Natural Language Processing,pp\. 12055–12070\.External Links:[Document](https://dx.doi.org/10.18653/v1/2025.findings-emnlp.645)Cited by:[Lightweight Architectures and Model Quantization](https://arxiv.org/html/2609.30836#Sx2.SSx2.p1.1)\.
- Zhu and Yang \(2025\)P\. Zhu and T\. YangCE\-lslm: efficient large\-small language model inference and communication via cloud\-edge collaboration\.Note:arxiv:2505\.14085Cited by:[Edge\-Cloud Collaboration](https://arxiv.org/html/2609.30836#Sx2.SSx3.p1.1)\.

相似文章

MiniCPM4:面向终端设备的超高效大语言模型

Papers with Code Trending

MiniCPM4 是一款专为终端设备设计的高效大语言模型,通过稀疏注意力、数据筛选、训练算法和推理系统等方面的创新,在0.5B和8B参数版本上实现了强大性能。