1.之江实验室,高效能计算设施研究中心,浙江杭州 311100
2.温州肯恩大学国际前沿交叉研究院,浙江温州 325060
牛昊一 男,1995年10月出生于河北省张家口市。2023年毕业于浙江大学控制科学与工程学院。现为之江实验室助理研究员。主要研究方向为人工智能和智能优化算法。E-mail: hyniu95@gmail.com
林露 男,1986年9月出生于湖北省黄冈市。2011年毕业于武汉大学计算机系。现为之江实验室高级研究专员。主要研究方向为大模型推理技术。E-mail: limiximil@zhejianglab.org
袭向明 男,1987年12月出生于黑龙江省哈尔滨市。2010年毕业于清华大学自动化系。现为之江实验室助理研究员。主要研究方向为非线性优化方法。E-mail: xixiangming@gmail.com
张北北 男,1994年2月出生于黑龙江省五常市。2020年毕业于多伦多大学电子及计算机工程系。现为之江实验室高级研究专员。主要研究方向为智能化软件工程与系统。E-mail: beibei@zhejianglab.org
邹涛 男,1974年11月出生于重庆市。现为之江实验室研究专家、研究员。主要研究方向为智能计算和高性能网络。E-mail: zout@zhejianglab.org
高丰 男,1974年11月出生于浙江省丽水市。2002年毕业于浙江大学信息与电子工程学院。现为之江实验室高级工程师。主要研究方向为分布式机器学习、云计算、边缘计算。E-mail: gaof@zhejianglab.org
收稿:2026-01-14,
录用:2026-01-23,
网络首发:2026-05-19,
纸质出版:2026-04-25
移动端阅览
牛昊一, 林露, 袭向明, 等. 泛在计算环境下模型推理输出长度预测[J]. 电子学报, 2026, 54(04): 1614-1625.
NIU Haoyi, LIN Lu, XI Xiangming, et al. Prediction on Model Inference Output Length in Ubiquitous Computing Environment[J]. Acta Electronica Sinica, 2026, 54(04): 1614-1625.
牛昊一, 林露, 袭向明, 等. 泛在计算环境下模型推理输出长度预测[J]. 电子学报, 2026, 54(04): 1614-1625. DOI:10.12263/DZXB.20250897
NIU Haoyi, LIN Lu, XI Xiangming, et al. Prediction on Model Inference Output Length in Ubiquitous Computing Environment[J]. Acta Electronica Sinica, 2026, 54(04): 1614-1625. DOI:10.12263/DZXB.20250897
在泛在计算场景中,系统涵盖大、中、小等多种规模的模型与端侧、边侧、云侧等多种算力形态,基于服务质量需求,需对模型推理服务进行精细化调度。在此背景下,精准预测模型推理的输出长度,对实现高效资源管理至关重要。然而,由于生成结果存在高度可变性且受语义影响显著,输出长度预测始终面临挑战。本研究提出一种轻量级的预测模型训练框架(Output Length Prediction Model, OLPM),仅依据输入提示即可预测请求的输出长度区间。框架采用基于DistilBERT的架构以保障计算效率,并引入两项关键优化手段:一是基于正态分布的过滤策略,通过降低异常值偏差提升训练稳定性;二是语义提示增强方法,提高模型对控制长度语言线索的敏感度。评估结果表明,OLPM的预测精度达到当前最优水平,最高准确率可达99.6%。与现有模型相比,OLPM不仅显著提升了预测精度,还大幅缩短了模型收敛时间,实现了精度与效率的协同优化。实验结果进一步说明,该方法收敛迅速、性能稳定,能够实现精准且高效的输出长度预测。OLPM能够在保持高预测精度同时,具备较低的模型复杂度与计算开销,显著优于现有方法,可有效满足泛在计算场景下对资源调度精细化的需求,展现出卓越的实用性、可扩展性与泛化能力,为实现泛在计算环境下高效、智能的模型推理服务提供了先进的技术路径。
In ubiquitous computing scenarios
systems encompass large
medium
and small-scale models
as well as diverse computing resources spanning edge-side
device-side
and cloud-side infrastructures. Efficient scheduling of model inference services
driven by quality-of-service requirements
is critical to optimizing system performance. Accurate prediction of output length during model inference plays a pivotal role in enabling fine-grained resource management. However
this task remains challenging due to the high variability and strong semantic dependence of generated outputs. To address this
we propose a lightweight training framework
which is output length prediction model (OLPM) to predict the output length range of requests based solely on the input prompt. Our framework employs a DistilBERT architecture to ensure computational efficiency
enhanced with two key innovations: (1) a normal-distribution-based filtering strategy that improves training stability by mitigating the impact of outliers
and (2) a semantic prompt augmentation method that enhances the model’s sensitivity to linguistic cues indicative of output length. The evaluation results show that OLPM achieves state-of-the-art prediction accuracy
with a maximum accuracy of up to 96.8%. Compared to existing models
OLPM not only significantly improves prediction accuracy but also greatly reduces model convergence time
achieving a synergistic optimization of precision and efficiency. Experimental results further show that the method converges rapidly and performs stably
enabling accurate and efficient output length prediction. By maintaining high prediction accuracy while featuring low model complexity and computational overhead
OLPM significantly outperforms existing approaches. It effectively meets the demands for fine-grained resource scheduling in ubiquitous computing environments
demonstrating outstanding practicality
scalability
and generalization capability. OLPM provides an advanced technical pathway for realizing efficient and intelligent model inference services in ubiquitous computing scenarios.
章晋睿 , 龙婷婷 , 张德宇 , 等 . 端智能推理加速技术综述 [J ] . 电子学报 , 2025 , 53 ( 4 ): 1063 - 1102 .
Zhang Jinrui , Long Tingting , Zhang Deyu , et al . On-device intelligence acceleration technologies: A survey [J ] . Acta Electronica Sinica , 2025 , 53 ( 4 ): 1063 - 1102 . (in Chinese)
Wu Haijie , Lin Weiwei , Shen Wangbo , et al . Prediction of heterogeneous device task runtime based on edge server-oriented deep neuro-fuzzy system [J ] . IEEE Transactions on Services Computing , 2025 , 18 ( 1 ): 372 - 384 . DOI: 10.1109/TSC.2024.3520869 http://dx.doi.org/10.1109/TSC.2024.3520869
Byun J , Choi Y , Lee J , et al . Privacy-preserving inference resistant to model extraction attacks [J ] . Expert Systems with Applications , 2024 , 256 : 124830 . DOI: 10.1016/j.eswa.2024.124830 http://dx.doi.org/10.1016/j.eswa.2024.124830
杨赟辉 , 程虎 , 魏敬和 , 等 . 面向Transformer模型边缘端部署的常用激活函数高精度轻量级量化推理方法 [J ] . 电子学报 , 2024 , 52 ( 10 ): 3301 - 3311 . DOI: 10.12263/DZXB.20240435 http://dx.doi.org/10.12263/DZXB.20240435
Yang Yunhui , Cheng Hu , Wei Jinghe , et al . High-precision lightweight quantization inference method for prevalent activation functions in Transformer models in edge device deployment [J ] . Acta Electronica Sinica , 2024 , 52 ( 10 ): 3301 - 3311 . (in Chinese) . DOI: 10.12263/DZXB.20240435 http://dx.doi.org/10.12263/DZXB.20240435
Chen Huajie , Zhu Tianqing , Ji Shouling , et al . Stand-in model protection: Synthetic defense for membership inference and model inversion attacks [J ] . Knowledge-Based Systems , 2025 , 316 : 113339 . DOI: 10.1016/j.knosys.2025.113339 http://dx.doi.org/10.1016/j.knosys.2025.113339
Chen Yuxuan , Li Rongpeng , Yu Xiaoxue , et al . Adaptive layer splitting for wireless large language model inference in edge computing: A model-based reinforcement learning approach [J ] . Frontiers of Information Technology & Electronic Engineering , 2025 , 26 ( 2 ): 278 - 292 . DOI: 10.1631/FITEE.2400468 http://dx.doi.org/10.1631/FITEE.2400468
Yang Jin , Wu Qiong , Feng Zhiying , et al . Quality-of-service aware LLM routing for edge computing with multiple experts [J ] . IEEE Transactions on Mobile Computing , 2025 , 24 ( 12 ): 13648 - 13662 . DOI: 10.1109/TMC.2025.3590969 http://dx.doi.org/10.1109/TMC.2025.3590969
Barkalov A , Lemeshko O , Yeremenko O , et al . Solving load balancing problems in routing and limiting traffic at the network edge [J ] . Applied Sciences , 2023 , 13 ( 17 ): 9489 . DOI: 10.3390/app13179489 http://dx.doi.org/10.3390/app13179489
Seo M , Hyun J , Jeong S , et al . OASIS: Outlier-aware KV cache clustering for scaling LLM inference in CXL memory systems [J ] . IEEE Computer Architecture Letters , 2025 , 24 ( 1 ): 165 - 168 . DOI: 10.1109/lca.2025.3567844 http://dx.doi.org/10.1109/lca.2025.3567844
Cui Lixiao , He Kewen , Li Yusen , et al . SwapKV: A hotness aware in-memory key-value store for hybrid memory systems [J ] . IEEE Transactions on Knowledge and Data Engineering , 2023 , 35 ( 1 ): 917 - 930 .
Lu Jinkang , Lv Meng , Li Peixuan , et al . Dhcache: A dual-hash cache for optimizing the read performance in key-value store [J ] . The Journal of Supercomputing , 2025 , 81 ( 2 ): 400 . DOI: 10.1007/s11227-024-06828-w http://dx.doi.org/10.1007/s11227-024-06828-w
He Ying , Fang Jingcheng , Yu F R , et al . Large language models (LLMs) inference offloading and resource allocation in cloud-edge computing: An active inference approach [J ] . IEEE Transactions on Mobile Computing , 2024 , 23 ( 12 ): 11253 - 11264 . DOI: 10.1109/tmc.2024.3415661 http://dx.doi.org/10.1109/tmc.2024.3415661
Li Yandi , Guo Jianxiong , Tang Zhiqing , et al . Cloud-edge system for scheduling unpredictable LLM requests with combinatorial bandit [J ] . IEEE Transactions on Services Computing , 2025 , 18 ( 6 ): 3567 - 3580 . DOI: 10.1109/tsc.2025.3611379 http://dx.doi.org/10.1109/tsc.2025.3611379
Hu Yitao , Liu Xiulong , Yang Guotao , et al . Tightllm: Maximizing throughput for llm inference via adaptive offloading policy [J ] . IEEE Transactions on Computers , 2025 , 74 ( 7 ): 2195 - 2209 . DOI: 10.1109/TC.2025.3558009 http://dx.doi.org/10.1109/TC.2025.3558009
Hu Mingzhu , Qin Shengfeng , Wang Shuying , et al . An energy-saving real-time scheduling method based on bi-level multi-agent architecture with bargaining game for flexible job shops [J ] . Expert Systems with Applications , 2025 , 269 : 126527 . DOI: 10.1016/j.eswa.2025.126527 http://dx.doi.org/10.1016/j.eswa.2025.126527
Tang Xuhao , Liu Fagui , Xu Dishi , et al . LLM-assisted reinforcement learning: Leveraging lightweight large language model capabilities for efficient task scheduling in multi-cloud environment [J ] . IEEE Transactions on Consumer Electronics , 2025 , 71 ( 2 ): 5631 - 5644 . DOI: 10.1109/tce.2024.3524612 http://dx.doi.org/10.1109/tce.2024.3524612
Zhang Xinyuan , Nie Jiangtian , Huang Yudong , et al . Beyond the cloud: Edge inference for generative large language models in wireless networks [J ] . IEEE Transactions on Wireless Communications , 2025 , 24 ( 1 ): 643 - 658 . DOI: 10.1109/TWC.2024.3497923 http://dx.doi.org/10.1109/TWC.2024.3497923
0
浏览量
5
下载量
0
CSCD
关联资源
相关文章
相关作者
相关机构
京公网安备11010802024621