WO2025124306A1 - 无线接入网络的ai推理任务编排方法、装置、设备和存储介质 - Google Patents
无线接入网络的ai推理任务编排方法、装置、设备和存储介质 Download PDFInfo
- Publication number
- WO2025124306A1 WO2025124306A1 PCT/CN2024/137406 CN2024137406W WO2025124306A1 WO 2025124306 A1 WO2025124306 A1 WO 2025124306A1 CN 2024137406 W CN2024137406 W CN 2024137406W WO 2025124306 A1 WO2025124306 A1 WO 2025124306A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- reasoning
- access network
- information
- wireless access
- model
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Pending
Links
Classifications
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04W—WIRELESS COMMUNICATION NETWORKS
- H04W16/00—Network planning, e.g. coverage or traffic planning tools; Network deployment, e.g. resource partitioning or cells structures
- H04W16/18—Network planning tools
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04W—WIRELESS COMMUNICATION NETWORKS
- H04W16/00—Network planning, e.g. coverage or traffic planning tools; Network deployment, e.g. resource partitioning or cells structures
- H04W16/22—Traffic simulation tools or models
Definitions
- the present disclosure relates to the field of wireless communication technology, and in particular to a method, device, equipment and storage medium for scheduling AI reasoning tasks in a wireless access network.
- the Near-Real-Time RAN Intelligent Controller (Near-RT RIC) in the traditional Open Radio Access Network (O-RAN) can collect near-real-time communication-related parameters through the standardized E2 interface to intelligently optimize system performance.
- This is only the collection and optimization of communication-related parameters on the RAN side. It lacks the unified perception and scheduling capability of computing resources and communication resources, resulting in low efficiency of resource allocation and collaborative computing on the RAN side, thereby causing system performance degradation.
- the purpose of the disclosed embodiments is to provide an AI reasoning task scheduling method for a wireless access network, which introduces artificial intelligence (AI) technology on the RAN side. It can perform near real-time collaborative scheduling of AI reasoning tasks for computing nodes on the RAN side based on real-time communication status and computing power, thereby improving the efficiency of resource allocation and collaborative computing on the RAN side, and thus improving system performance.
- AI artificial intelligence
- an embodiment of the present disclosure provides an AI reasoning task scheduling method for a wireless access network, which is applied to a near real-time wireless controller.
- the method includes:
- the terminal information includes computing resource information and communication resource information
- the generating of the AI reasoning task scheduling scheme of the wireless access network according to the AI reasoning task guarantee strategy and the terminal information includes:
- the node to be allocated and the corresponding computing node according to each terminal information, and select the corresponding split point and exit point from the feature information of the AI reasoning model according to each terminal information; wherein the split point and the exit point are one layer in the AI reasoning model;
- the AI reasoning task scheduling scheme is composed of the nodes to be allocated and their corresponding computing nodes, the splitting points and the exit points.
- the method further includes:
- the AI reasoning task scheduling scheme is sent to the node to be assigned and the corresponding computing node, so that the node to be assigned and the corresponding computing node can complete their respective reasoning computing tasks.
- the acquiring terminal information of at least one terminal includes:
- the terminal information is subscribed to a base station where the user equipment resides, so that the base station reports the terminal information of the user equipment.
- the terminal when the terminal is a user equipment, the user equipment supports information interaction with a near real-time wireless controller; then, obtaining terminal information of at least one terminal includes:
- the embodiment of the present disclosure further provides an AI reasoning task scheduling method for a wireless access network, which is applied to a non-real-time wireless controller.
- the method includes:
- the AI reasoning task guarantee strategy is sent to a near real-time wireless controller, so that the near real-time wireless controller generates an AI reasoning task scheduling plan for the wireless access network according to the AI reasoning task guarantee strategy and terminal information; wherein the terminal information includes computing resource information and communication resource information.
- the model information includes characteristic information of the AI reasoning model and performance guarantee parameters of the AI reasoning task; wherein, the performance guarantee parameters include but are not limited to model reasoning accuracy and the number of reasoning calculations per unit time.
- an AI reasoning task scheduling device for a wireless access network including:
- An AI reasoning task assurance strategy receiving module used to receive an AI reasoning task assurance strategy of an AI reasoning model sent by a non-real-time wireless controller; wherein the AI reasoning model is deployed in a wireless access network, and the AI reasoning task assurance strategy defines an optimization target of the wireless access network and characteristic information of the AI reasoning model;
- a terminal information acquisition module used to acquire terminal information of at least one terminal; wherein the terminal information includes computing resource information and communication resource information;
- the AI reasoning task scheduling scheme generation module is used to generate an AI reasoning task scheduling scheme for the wireless access network according to the AI reasoning task guarantee strategy and the terminal information.
- an AI reasoning task scheduling device for a wireless access network including:
- Model information acquisition module used to obtain model information of AI reasoning model
- An AI reasoning task assurance strategy generation module used to generate an AI reasoning task assurance strategy according to the model information; wherein the AI reasoning model is deployed in a wireless access network, and the AI reasoning task assurance strategy defines an optimization target of the wireless access network and characteristic information of the AI reasoning model;
- An AI reasoning task guarantee strategy sending module is used to send the AI reasoning task guarantee strategy to a near real-time wireless controller, so that the near real-time wireless controller generates an AI reasoning task scheduling plan for the wireless access network according to the AI reasoning task guarantee strategy and terminal information; wherein the terminal information includes computing resource information and communication resource information.
- an embodiment of the present disclosure also provides an electronic device, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein when the processor executes the computer program, it implements the AI reasoning task scheduling method for the wireless access network as described in any of the above embodiments.
- an embodiment of the present disclosure also provides a computer-readable storage medium, wherein the computer-readable storage medium includes a stored computer program, wherein when the computer program is running, the device where the computer-readable storage medium is located is controlled to execute the AI reasoning task scheduling method for the wireless access network as described in any of the above embodiments.
- the AI reasoning task scheduling method, device, equipment and storage medium of the wireless access network disclosed in the present invention introduce artificial intelligence technology on the RAN side, first receive the AI reasoning task guarantee policy of the AI reasoning model sent by the non-real-time wireless controller, and obtain terminal information of at least one terminal, wherein the AI reasoning task guarantee policy defines the optimization goal of the wireless access network and the characteristic information of the AI reasoning model, and the terminal information includes computing resource information and communication resource information; then generate the AI reasoning task scheduling scheme of the wireless access network according to the AI reasoning task guarantee policy and the terminal information, and can perform near real-time AI reasoning task collaborative scheduling for the computing nodes on the RAN side according to the real-time communication status and computing power, thereby improving the efficiency of resource allocation and collaborative computing on the RAN side, thereby improving system performance.
- FIG1 is a flow chart of a method for scheduling AI reasoning tasks in a wireless access network provided by an embodiment of the present disclosure
- FIG. 2 is a schematic diagram of information interaction between a near real-time wireless controller and a non-real-time wireless controller provided by an embodiment of the present disclosure
- FIG. 3 is a flow chart of a first method for acquiring terminal information of a user device provided in an embodiment of the present disclosure
- FIG. 4 is a flow chart of a second method for acquiring terminal information of a user device provided in an embodiment of the present disclosure
- FIG5 is a flowchart of another method for scheduling AI reasoning tasks in a wireless access network provided by an embodiment of the present disclosure
- FIG6 is a structural block diagram of an AI reasoning task scheduling device for a wireless access network provided by an embodiment of the present disclosure
- FIG7 is a structural block diagram of another device for arranging AI reasoning tasks in a wireless access network provided by an embodiment of the present disclosure
- FIG8 is a structural block diagram of an electronic device provided by an embodiment of the present disclosure.
- AI edge computing can be used to collaboratively complete computing tasks.
- AI edge computing can be achieved by building a deep neural network model. Since the convolution layer in the deep neural network model structure is a kernel that performs a point multiplication operation on the spatial dimension of the input tensor to generate a feature map of the output tensor, the convolution layer can be used as a split point of the model, and multiple computing nodes can collaborate to complete the reasoning of the model.
- AI technology is introduced on the RAN side to assist the terminal in completing computing tasks.
- FIG. 1 is a flow chart of a method for scheduling AI reasoning tasks in a wireless access network provided by an embodiment of the present disclosure.
- the method for scheduling AI reasoning tasks in a wireless access network is implemented by a near real-time wireless controller.
- the method includes:
- O-RAN includes a near real-time controller (Near-Real-Time RAN Intelligent Controller, Near-RT RIC) and a non-real-time controller (Non-Real-Time RAN Intelligent Controller, Non-RT RIC).
- the terminals include but are not limited to base stations, centralized units, distributed units (Distributed Unit, DU) and user equipment (User Equipment, UE).
- the centralized unit includes (Centralized Unit, CU), centralized unit control plane (Centralized Unit-Control Plane, CU-CP) and distributed unit user plane (Centralized Unit-User Plane, CU-UP).
- FIG 2 is a schematic diagram of information interaction between a near real-time wireless controller and a non-real-time wireless controller provided in an embodiment of the present disclosure.
- the non-real-time controller is connected to the near real-time controller through an open and standardized A1 interface.
- the purpose of the non-real-time controller is to provide a corresponding machine learning model to support RAN intelligence, and the non-real-time controller provides an AI reasoning model and data for the near real-time controller. Due to the real-time requirements in the O-RAN architecture, when providing corresponding functions, the near real-time controller performs related operations by using an existing AI reasoning model.
- the AI reasoning model can perform one or more calculations in fault alarm analysis, coverage optimization, parameter optimization, spectrum analysis, inter-station collaboration, mobility management, slice management, wireless positioning and environmental perception recognition.
- the near real-time controller and the terminal exchange information through the E2 interface.
- the functions of the E2 interface and the A1 interface in the O-RAN architecture are enhanced in the embodiment of the present disclosure.
- the AI reasoning task assurance policy generated by the non-real-time controller according to the model information of the AI reasoning model is sent to the near real-time controller through the A1 interface.
- the real-time terminal information is then collected through the E2 interface.
- Near real-time AI reasoning task collaborative orchestration is performed according to the terminal information and the AI reasoning task assurance policy, thereby improving the resource allocation and collaborative computing efficiency on the RAN side and improving the system performance.
- the non-real-time wireless controller generates the AI reasoning task assurance strategy according to the model information of the AI reasoning model; wherein the model information includes feature information of the AI reasoning model and performance assurance parameters of the AI reasoning task.
- the model information is obtained by the non-real-time wireless controller from a service management and orchestration framework (SMO).
- SMO service management and orchestration framework
- the source of the model information in the SMO can be directly configured by the operator or obtained by interacting with external applications.
- relevant interfaces or application programming interfaces (API) can be pre-designed to establish a connection with the external application through an authorization and authentication mechanism to obtain the model information.
- API application programming interfaces
- the characteristic information of the AI reasoning model includes the AI reasoning model identifier (IDentitifier, ID), the number of model layers, the type of each layer, the output data volume of each layer, the calculation amount of each layer, the split point position, the exit point position, the exit point accuracy and the exit point classifier calculation amount, etc.
- IDentitifier ID
- the meaning corresponding to each of the characteristic information can be referred to Table 1, and the specific examples can be referred to Table 2.
- conv represents the convolutional layer
- max pool represents the maximum pooling layer
- avg pool represents the average pooling layer
- fc represents the fully connected layer
- Exit represents the exit point position.
- the ResNet-18 model has four exit points, namely Exit1 at the 5th layer, Exit2 at the 9th layer, Exit3 at the 13th layer, and Exit4 at the 18th layer. The remaining layers can be used as split points.
- the exit point accuracy can evaluate the classification accuracy when this layer is selected as the exit point.
- the accuracy in the table is only an example.
- the performance guarantee parameters include but are not limited to model reasoning accuracy, number of reasoning calculations per unit time, model reasoning round trip delay and model spectrum efficiency.
- the meaning corresponding to each of the performance guarantee parameters can be referred to Table 3.
- the non-real-time wireless controller After receiving the model information, the non-real-time wireless controller generates an AI reasoning task guarantee strategy according to the model information, and the AI reasoning task guarantee strategy defines the optimization target of the wireless access network, the optimizable parameters and the characteristic information of the AI reasoning model.
- the optimization target of the wireless access network refers to the overall optimization target of the guarantee strategy in the current system, such as maximizing the accuracy of model reasoning in the system, minimizing the round-trip delay of model reasoning in the system, etc.
- the optimizable parameters refer to the parameters that the near real-time wireless controller can adjust for AI reasoning task guarantee, such as the selection of the split point and the exit point of the AI reasoning model.
- the optimization target of the wireless access network is to maximize the accuracy of model reasoning in the system and minimize the round-trip delay of model reasoning in the system, that is, after selecting the split point and the exit point, the model reasoning accuracy needs to be maximized and the round-trip delay of model reasoning needs to be minimized, which can be obtained by polling each split point and exit point.
- the terminal information includes but is not limited to: computing resource information, communication resource information, current cell service user information, and terminal preference information for AI reasoning tasks.
- the computing resource information includes but is not limited to: the central processing unit (CPU) utilization rate of the base station and the terminal, CPU frequency, CPU core binding status, number of remaining CPU cores, CPU/graphics processing unit (GPU)/embedded neural network processor (NPU) floating point operations per second (Floating Point Operations Per Second, FLOPS), GPU video memory capacity/remaining video memory, GPU cuda core number/remaining core number, etc.
- the communication resource information includes but is not limited to: the total uplink/downlink bandwidth of the base station, the remaining bandwidth, the physical resource block (Phys The information may also include the number of AI resource blocks (PRBs), PRB utilization, etc., and may also include the channel conditions of the terminal, the communication link delay from the base station to the neighboring station, etc.
- the channel conditions of the terminal include the signal-to-noise ratio (SNR), the received signal strength indication (RSSI), etc.
- the current cell service user information includes but is not limited to: user type identification (non-AI user or AI user), AI reasoning task model identification (for example: ResNet-18, Visual Geometry Group (VGG), etc.);
- the terminal preference information for AI reasoning tasks includes but is not limited to: preference for segmentation point selection for each AI reasoning task, etc.
- the terminal information can be sent directly by the base station to the near real-time wireless controller.
- the terminal is a user device
- two methods of obtaining the terminal information of the user device are provided in the embodiment of the present disclosure, the first method is to obtain it through the base station, and the second method is to obtain it directly through the user device.
- the base station described in the embodiment of the present disclosure may have various forms, such as a macro base station, a micro base station, a relay station or an access point.
- the base station may be an integrated base station, or may be a base station including a centralized unit CU and a distributed unit DU.
- obtaining the terminal information of at least one terminal includes: sending a request message to a base station where the user device resides, so that the base station reports the terminal information of the user device according to the request message; or, subscribing to the terminal information to the base station where the user device resides, so that the base station reports the terminal information of the user device.
- the near real-time wireless controller subscribes or requests the terminal information from the base station where the user equipment resides through the E2 interface. It can be a periodic subscription or an event-triggered report.
- the event trigger includes but is not limited to: base station computing power fluctuations, terminal computing power fluctuations, base station bandwidth fluctuations, terminal power fluctuations, terminal segmentation mode preference changes, etc.
- the base station collects the air interface and terminal data of the user equipment according to demand, summarizes it into terminal information, and reports the terminal information of the user equipment to the near real-time wireless controller. In addition, the base station also synchronously reports its own terminal information to the near real-time wireless controller.
- the terminal when the terminal is a user device, the user device supports information interaction with a near real-time wireless controller; then, obtaining terminal information of at least one terminal includes: sending a request message to the user device so that the user device reports the terminal information according to the request message; or subscribing to the terminal information to the user device so that the user device reports the terminal information.
- the user equipment and the near real-time wireless controller enable both parties to support the corresponding interface application layer protocol based on the existing protocol stack, thereby realizing information interaction between the two parties.
- the near real-time wireless controller subscribes to or requests the terminal information from the user equipment through the E2 interface, which can be periodic subscription or event-triggered reporting.
- the user equipment collects air interface and terminal data according to demand, summarizes it into terminal information, and reports the terminal information to the near real-time wireless controller.
- the near real-time wireless controller synchronously sends a request message to the base station or subscribes to the terminal information of the base station, so that the base station reports its own terminal information.
- step S13 generating an AI reasoning task scheduling scheme for a wireless access network according to the AI reasoning task assurance strategy and the terminal information includes:
- the AI reasoning task scheduling scheme is composed of the nodes to be allocated and their corresponding computing nodes, the splitting points and the exit points.
- the node to be allocated is a terminal that cannot meet the computing requirements by itself and needs to use a computing node to share the reasoning computing task; the computing node is a terminal that can receive the reasoning computing task of the node to be allocated.
- the near real-time wireless controller After receiving the terminal information sent by at least one terminal, the near real-time wireless controller determines the node to be allocated and the computing node, and applies an optimization algorithm (such as a differential algorithm) to evaluate the computing nodes required by these nodes to be allocated, and selects the split point and exit point.
- an optimization algorithm such as a differential algorithm
- the computing resource information that one of the nodes to be allocated requires a large amount of computing, more layers (that is, the number of layers between the split point and the exit point is relatively large) can be allocated to the node to be allocated, otherwise fewer layers can be allocated; if it can be determined according to the communication resource information that the node to be allocated requires a computing node with a large remaining bandwidth, computing nodes that meet the conditions are preferentially allocated as computing nodes of the node to be allocated; if at this time one of the terminal information gives an AI reasoning task model identifier, then when generating an AI reasoning task scheduling plan, the corresponding AI reasoning model will be selected to perform the reasoning task according to the AI reasoning task model identifier; if at this time one of the terminal information gives a split point selection preference for each AI reasoning task, then when generating an AI reasoning task scheduling plan, the split point will be preferentially allocated according to the split point selection preference of the node to be allocated.
- the method further includes:
- the node to be allocated is a user device and the computing node is a base station
- the terminal information of three user devices and two base stations is collected, and the two base stations and three user devices meet the following conditions: base station A provides communication connection services for user device 1
- base station B provides communication connection services for user devices 2 and 3
- base stations A and B both provide computing services
- the AI reasoning models are all ResNet-18 in Table 2.
- the scheduling scheme is to select appropriate computing nodes for user devices 1 to 3, and select the splitting points and exit points for the AI reasoning task
- the AI reasoning task scheduling scheme is shown in Table 4 below.
- user device 1 is taken as an example.
- the computing node selected for user device 1 is base station A, that is, base station A and user device 1 jointly complete the computing task
- the split point is layer 3
- the exit point is Exit2.
- the computing amount completed by base station A is the computing amount from layer 3 to layer 9.
- user device 1 completes the inference calculation of the previous layers.
- base station A completes the inference calculation from layer 3 to layer 9, and finally transmits the generated result back to user device 1. Therefore, the selection of the split point and the exit point determines the computing amount of the node to be allocated and the computing node respectively, and the selection of the exit point also determines the accuracy of the model.
- the AI reasoning task scheduling method of the wireless access network disclosed in the present invention introduces artificial intelligence technology on the RAN side, first receives the AI reasoning task guarantee policy of the AI reasoning model sent by the non-real-time wireless controller, and obtains the terminal information of at least one terminal, wherein the AI reasoning task guarantee policy defines the optimization goal of the wireless access network and the characteristic information of the AI reasoning model, and the terminal information includes computing resource information and communication resource information; then, the AI reasoning task scheduling scheme of the wireless access network is generated according to the AI reasoning task guarantee policy and the terminal information, and through the selection of computing nodes, task splitting points and exit points, the computing nodes on the RAN side can be coordinated in near real-time AI reasoning tasks according to the real-time communication status and computing power, thereby improving the efficiency of resource allocation and collaborative computing on the RAN side, thereby improving the system performance.
- FIG. 5 is a flowchart of another method for scheduling AI reasoning tasks in a wireless access network provided by an embodiment of the present disclosure.
- the method for scheduling AI reasoning tasks in a wireless access network is implemented by a non-real-time wireless controller.
- the method includes:
- the model information includes characteristic information of the AI reasoning model and performance guarantee parameters of the AI reasoning task; wherein the performance guarantee parameters include but are not limited to model reasoning accuracy and the number of reasoning calculations per unit time.
- FIG. 6 is a structural block diagram of an AI reasoning task scheduling device 100 for a wireless access network provided in an embodiment of the present disclosure.
- the AI reasoning task scheduling device 100 for a wireless access network includes:
- the AI reasoning task assurance strategy receiving module 11 is used to receive the AI reasoning task assurance strategy of the AI reasoning model sent by the non-real-time wireless controller; wherein the AI reasoning model is deployed in the wireless access network, and the AI reasoning task assurance strategy defines the optimization target of the wireless access network and the characteristic information of the AI reasoning model;
- the terminal information acquisition module 12 is used to acquire terminal information of at least one terminal; wherein the terminal information includes computing resource information and communication resource information;
- the AI reasoning task scheduling scheme generating module 13 is used to generate an AI reasoning task scheduling scheme for the wireless access network according to the AI reasoning task guarantee strategy and the terminal information.
- the AI reasoning task scheduling scheme generation module 13 is specifically used to:
- the node to be allocated and the corresponding computing node according to each terminal information, and select the corresponding split point and exit point from the feature information of the AI reasoning model according to each terminal information; wherein the computing node is a base station, and the split point and the exit point are one layer in the AI reasoning model;
- the AI reasoning task scheduling scheme is composed of the nodes to be allocated and their corresponding computing nodes, the splitting points and the exit points.
- the AI reasoning task scheduling device 100 of the wireless access network further includes:
- the AI reasoning task scheduling scheme sending module is used to send the AI reasoning task scheduling scheme to the to-be-assigned node and the corresponding computing node, so that the to-be-assigned node and the corresponding computing node complete their respective reasoning computing tasks.
- the terminal information acquisition module 12 is specifically used to: send a request message to the base station where the user device resides, so that the base station reports the terminal information of the user device according to the request message; or, subscribe to the terminal information to the base station where the user device resides, so that the base station reports the terminal information of the user device.
- the terminal information acquisition module 12 is specifically used to: send a request message to the user device so that the user device reports the terminal information according to the request message; or subscribe to the terminal information to the user device so that the user device reports the terminal information.
- FIG. 7 is a structural block diagram of another AI reasoning task scheduling device 200 for a wireless access network provided in an embodiment of the present disclosure.
- the AI reasoning task scheduling device 200 for a wireless access network includes:
- a model information acquisition module 21 is used to acquire model information of an AI reasoning model
- An AI reasoning task assurance strategy generation module 22 is used to generate an AI reasoning task assurance strategy according to the model information; wherein the AI reasoning model is deployed in a wireless access network, and the AI reasoning task assurance strategy defines the optimization target of the wireless access network and the characteristic information of the AI reasoning model;
- the AI reasoning task guarantee strategy sending module 23 is used to send the AI reasoning task guarantee strategy to the near real-time wireless controller, so that the near real-time wireless controller generates an AI reasoning task scheduling plan for the wireless access network according to the AI reasoning task guarantee strategy and terminal information; wherein the terminal information includes computing resource information and communication resource information.
- the model information includes characteristic information of the AI reasoning model and performance guarantee parameters of the AI reasoning task; wherein the performance guarantee parameters include but are not limited to model reasoning accuracy and the number of reasoning calculations per unit time.
- the AI reasoning task scheduling device for the wireless access network disclosed in the present invention introduces artificial intelligence technology on the RAN side, first receives the AI reasoning task guarantee strategy of the AI reasoning model sent by the non-real-time wireless controller, and obtains the terminal information of at least one terminal, wherein the AI reasoning task guarantee strategy defines the optimization goal of the wireless access network and the characteristic information of the AI reasoning model, and the terminal information includes computing resource information and communication resource information; then, the AI reasoning task scheduling scheme for the wireless access network is generated according to the AI reasoning task guarantee strategy and the terminal information, and through the selection of computing nodes, task splitting points and exit points, the computing nodes on the RAN side can be coordinated in near real-time AI reasoning tasks according to the real-time communication status and computing power, thereby improving the efficiency of resource allocation and collaborative computing on the RAN side, and thus improving the system performance.
- FIG8 is a structural block diagram of an electronic device 300 provided in an embodiment of the present disclosure, wherein the electronic device 300 includes a processor 31, a memory 32, and a computer program stored in the memory 32 and executable on the processor 31.
- the processor 31 executes the computer program, the steps in the above-mentioned embodiments of the method for arranging AI reasoning tasks in each wireless access network are implemented, such as steps S11 to S13, and S21 to S23.
- the computer program may be divided into one or more modules/units, which are stored in the memory 32 and executed by the processor 31 to complete the present disclosure.
- the one or more modules/units may be a series of computer program instruction segments capable of completing specific functions, which are used to describe the execution process of the computer program in the electronic device 300.
- the electronic device 300 may include, but is not limited to, a processor 31 and a memory 32. Those skilled in the art will appreciate that the schematic diagram is merely an example of the electronic device 300 and does not constitute a limitation on the electronic device 300.
- the electronic device 300 may include more or fewer components than shown in the figure, or may combine certain components, or different components.
- the electronic device 300 may also include input and output devices, network access devices, buses, etc.
- the processor 31 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.
- a general-purpose processor may be a microprocessor or any conventional processor, etc.
- the processor 31 is the control center of the electronic device 300, and uses various interfaces and lines to connect various parts of the entire electronic device 300.
- the memory 32 can be used to store the computer program and/or module.
- the processor 31 realizes various functions of the electronic device 300 by running or executing the computer program and/or module stored in the memory 32 and calling the data stored in the memory 32.
- the memory 32 can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system, an application required for at least one function (such as a sound playback function, an image playback function, etc.), etc.; the data storage area can store data created according to the use of the mobile phone (such as audio data, a phone book, etc.), etc.
- the memory 32 can include a high-speed random access memory, and can also include a non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), at least one disk storage device, a flash memory device, or other volatile solid-state storage devices.
- a non-volatile memory such as a hard disk, a memory, a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), at least one disk storage device, a flash memory device, or other volatile solid-state storage devices.
- a non-volatile memory such as a hard disk, a memory, a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card
- the module/unit integrated in the electronic device 300 can be stored in a computer-readable storage medium.
- the present disclosure implements all or part of the processes in the above-mentioned embodiment method, and can also be completed by instructing the relevant hardware through a computer program.
- the computer program can be stored in a computer-readable storage medium, and when the computer program is executed by the processor 31, the steps of the above-mentioned various method embodiments can be implemented.
- the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form.
- the computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, U disk, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal and software distribution medium, etc.
Landscapes
- Engineering & Computer Science (AREA)
- Computer Networks & Wireless Communication (AREA)
- Signal Processing (AREA)
- Mobile Radio Communication Systems (AREA)
Abstract
本公开提供了一种无线接入网络的AI推理任务编排方法、装置、设备和存储介质,在RAN侧引入人工智能技术,接收非实时无线控制器发送的AI推理模型的AI推理任务保障策略,以及获取至少一个终端的终端信息,其中,AI推理任务保障策略中定义了无线接入网络的优化目标和AI推理模型的特征信息,终端信息包括计算资源信息和通信资源信息;根据所述AI推理任务保障策略和所述终端信息生成无线接入网络的AI推理任务编排方案。
Description
相关申请的交叉引用
本申请主张在2023年12月15日在中国提交的中国专利申请号No.202311733048.9的优先权,其全部内容通过引用包含于此。
本公开涉及无线通信技术领域,尤其涉及一种无线接入网络的AI推理任务编排方法、装置、设备和存储介质。
随着无线局域网中计算能力开放研究的不断深入,基站作为一种边缘计算平台越来越受到业界的重视,无线局域网正朝着通信与计算一体化的方向发展。虽然现有研究提供了设备边缘计算协作中平衡精度和延迟的解决方案,但这些研究主要集中在边缘服务器上,并不能直接应用于基站与终端之间的计算协作,而基站中的计算资源是由通信处理和计算任务共享的,两者存在竞争关系。因此,在计算资源和通信的带宽资源受限的情况下,无线接入网络(Radio Access Network,RAN)侧需要根据动态变化的信道状态和不同任务的不同信息,对多基站多终端任务做近实时自适应的编排。
传统开放无线接入网络(Open-Radio Access Network,O-RAN)中的近实时控制器(Near-Real-Time RAN Intelligent Controller,Near-RT RIC)可以通过标准化E2接口采集近实时的通信相关的参数对系统性能进行智能化优化,然而这仅仅是对于RAN侧的通信相关参数的采集优化,缺乏对于计算资源和通信资源的统一感知调度能力,导致RAN侧的资源分配和协作计算效率低下,从而导致系统性能下降。
本公开实施例的目的是提供一种无线接入网络的AI推理任务编排方法,在RAN侧引入人工智能(Artificial Intelligence,AI)技术,能根据实时的通信状态和计算能力,对RAN侧的计算节点做近实时的AI推理任务协作编排,提高了RAN侧资源分配和协作计算的效率,进而提升了系统性能。
为实现上述目的,本公开实施例提供了一种无线接入网络的AI推理任务编排方法,应用于近实时无线控制器,所述方法包括:
接收非实时无线控制器发送的AI推理模型的AI推理任务保障策略;其中,所述AI推理模型部署在无线接入网络中,所述AI推理任务保障策略中定义了无线接入网络的优化目标和AI推理模型的特征信息;
获取至少一个终端的终端信息;其中,所述终端信息包括计算资源信息和通信资源信息;
根据所述AI推理任务保障策略和所述终端信息生成无线接入网络的AI推理任务编排方案。
作为上述方案的改进,所述根据所述AI推理任务保障策略和所述终端信息生成无线接入网络的AI推理任务编排方案,包括:
按照所述无线接入网络的优化目标,根据每一终端信息确定待分配节点和对应的计算节点,以及根据每一终端信息从所述AI推理模型的特征信息中选择对应的切分点和退出点;其中,所述切分点和所述退出点为所述AI推理模型中的其中一层;
以所述待分配节点及其对应的计算节点、所述切分点和所述退出点组成所述AI推理任务编排方案。
作为上述方案的改进,在根据所述AI推理任务保障策略和所述终端信息生成无线接入网络的AI推理任务编排方案后,所述方法还包括:
将所述AI推理任务编排方案发送给待分配节点和对应的计算节点,以使所述待分配节点和对应的计算节点完成各自的推理计算任务。
作为上述方案的改进,当所述终端为用户设备时,所述获取至少一个终端的终端信息,包括:
向用户设备所驻留的基站发送请求消息,以使所述基站根据所述请求消息上报所述用户设备的终端信息;或,
向用户设备所驻留的基站订阅所述终端信息,以使所述基站上报所述用户设备的终端信息。
作为上述方案的改进,当所述终端为用户设备时,所述用户设备支持与近实时无线控制器的信息交互;则,所述获取至少一个终端的终端信息,包括:
向所述用户设备发送请求消息,以使所述用户设备根据所述请求消息上报终端信息;或,
向所述用户设备订阅所述终端信息,以使所述用户设备上报终端信息。
为实现上述目的,本公开实施例还提供了一种无线接入网络的AI推理任务编排方法,应用于非实时无线控制器,所述方法包括:
获取AI推理模型的模型信息;
根据所述模型信息生成AI推理任务保障策略;其中,所述AI推理模型部署在无线接入网络中,所述AI推理任务保障策略中定义了无线接入网络的优化目标和AI推理模型的特征信息;
将所述AI推理任务保障策略发送给近实时无线控制器,以使所述近实时无线控制器根据所述AI推理任务保障策略和终端信息生成无线接入网络的AI推理任务编排方案;其中,所述终端信息包括计算资源信息和通信资源信息。
作为上述方案的改进,所述模型信息包括所述AI推理模型的特征信息和AI推理任务的性能保障参数;其中,所述性能保障参数包括但不限于模型推理准确度和单位时间的推理计算次数。
为实现上述目的,本公开实施例还提供了一种无线接入网络的AI推理任务编排装置,包括:
AI推理任务保障策略接收模块,用于接收非实时无线控制器发送的AI推理模型的AI推理任务保障策略;其中,所述AI推理模型部署在无线接入网络中,所述AI推理任务保障策略中定义了无线接入网络的优化目标和AI推理模型的特征信息;
终端信息获取模块,用于获取至少一个终端的终端信息;其中,所述终端信息包括计算资源信息和通信资源信息;
AI推理任务编排方案生成模块,用于根据所述AI推理任务保障策略和所述终端信息生成无线接入网络的AI推理任务编排方案。
为实现上述目的,本公开实施例还提供了一种无线接入网络的AI推理任务编排装置,包括:
模型信息获取模块,用于获取AI推理模型的模型信息;
AI推理任务保障策略生成模块,用于根据所述模型信息生成AI推理任务保障策略;其中,所述AI推理模型部署在无线接入网络中,所述AI推理任务保障策略中定义了无线接入网络的优化目标和AI推理模型的特征信息;
AI推理任务保障策略发送模块,用于将所述AI推理任务保障策略发送给近实时无线控制器,以使所述近实时无线控制器根据所述AI推理任务保障策略和终端信息生成无线接入网络的AI推理任务编排方案;其中,所述终端信息包括计算资源信息和通信资源信息。
为实现上述目的,本公开实施例还提供了一种电子设备,包括处理器、存储器以及存储在所述存储器中且被配置为由所述处理器执行的计算机程序,所述处理器执行所述计算机程序时实现如上述任一实施例所述的无线接入网络的AI推理任务编排方法。
为实现上述目的,本公开实施例还提供了一种计算机可读存储介质,所述计算机可读存储介质包括存储的计算机程序,其中,在所述计算机程序运行时控制所述计算机可读存储介质所在设备执行如上述任一实施例所述的无线接入网络的AI推理任务编排方法。
相比于相关技术,本公开的无线接入网络的AI推理任务编排方法、装置、设备和存储介质,在RAN侧引入人工智能技术,首先接收非实时无线控制器发送的AI推理模型的AI推理任务保障策略,以及获取至少一个终端的终端信息,其中,AI推理任务保障策略中定义了无线接入网络的优化目标和AI推理模型的特征信息,终端信息包括计算资源信息和通信资源信息;然后根据所述AI推理任务保障策略和所述终端信息生成无线接入网络的AI推理任务编排方案,能根据实时的通信状态和计算能力,对RAN侧的计算节点做近实时的AI推理任务协作编排,提高了RAN侧资源分配和协作计算的效率,进而提升了系统性能。
图1是本公开实施例提供的一种无线接入网络的AI推理任务编排方法的流程图;
图2是本公开实施例提供的近实时无线控制器和非实时无线控制器的信息交互示意图;
图3是本公开实施例提供的第一种用户设备的终端信息获取方式的流程图;
图4是本公开实施例提供的第二种用户设备的终端信息获取方式的流程图;
图5是本公开实施例提供的另一种无线接入网络的AI推理任务编排方法的流程图;
图6是本公开实施例提供的一种无线接入网络的AI推理任务编排装置的结构框图;
图7是本公开实施例提供的另一种无线接入网络的AI推理任务编排装置的结构框图;
图8是本公开实施例提供的一种电子设备的结构框图。
下面将结合本公开实施例中的附图,对本公开实施例中的技术方案进行清楚、完整地描述,显然,所描述的实施例仅仅是本公开一部分实施例,而不是全部的实施例。基于本公开中的实施例,本领域普通技术人员在没有作出创造性劳动前提下所获得的所有其他实施例,都属于本公开保护的范围。
随着拓展现实(Extended Reality,XR)、自动驾驶、工业智能控制等新兴应用的兴起,深度学习相关算法例如图像识别等计算密集型任务受到更多关注。而移动设备的能力有限,无法满足其严格的计算要求和延迟要求,因此可借助AI边缘计算方式协作完成计算任务。AI边缘计算方式可通过构建深度神经网络模型实现,由于深度神经网络模型结构中的卷积层是内核在输入张量的空间维度上进行点乘操作以生成输出张量的特征图,因此卷积层可以作为模型的切分点,由多个计算节点协同完成模型的推理。同时,还可以通过在模型结构中添加分支分类器,牺牲准确度的条件下提前退出推理以减少资源浪费。基于早期退出机制和模型分割技术,目前有许多针对边缘计算节点的协作推理的研究,通过为AI推理任务选择合适切分点和退出点,在计算资源和系统带宽的限制下满足时延和推理的精度要求。因此,本公开实施例中在RAN侧引入AI技术,协助终端完成计算任务。
参见图1,图1是本公开实施例提供的一种无线接入网络的AI推理任务编排方法的流程图,所述无线接入网络的AI推理任务编排方法由近实时无线控制器执行实现,所述方法包括:
S11、接收非实时无线控制器发送的AI推理模型的AI推理任务保障策略;
S12、获取至少一个终端的终端信息;
S13、根据所述AI推理任务保障策略和所述终端信息生成无线接入网络的AI推理任务编排方案。
值得说明的是,O-RAN包括近实时控制器(Near-Real-Time RAN Intelligent Controller,Near-RT RIC)和非实时控制器(Non-Real-Time RAN Intelligent Controller,Non-RT RIC),所述终端包括但不限于基站、集中式单元、分布式单元(Distributed Unit,DU)和用户设备(User Equipment,UE),所述集中式单元包括(Centralized Unit,CU)、集中单元控制面(Centralized Unit-Control Plane,CU-CP)、分布单元用户面(Centralized Unit-User Plane,CU-UP)。参见图2,图2是本公开实施例提供的近实时无线控制器和非实时无线控制器的信息交互示意图,所述非实时控制器通过开放和标准化的A1接口连接所述近实时控制器,所述非实时控制器的目的在于提供对应的机器学习模型来支持RAN智能化,以及由所述非实时控制器为所述近实时控制器提供AI推理模型和数据,在O-RAN架构中由于实时性的要求,所述近实时控制器在提供相应的功能时,通过利用已有的AI推理模型执行相关运算,所述AI推理模型能够进行故障告警分析、覆盖优化、参数优化、频谱分析、站间协同、移动性管理、切片管理、无线定位和环境感知识别中的一种或多种计算。所述近实时控制器和终端通过E2接口进行信息交互,本公开实施例中增强了O-RAN架构中的E2接口和A1接口功能,通过A1接口将非实时控制器根据AI推理模型的模型信息生成的AI推理任务保障策略下发给近实时控制器,再由E2接口采集实时的终端信息,根据终端信息和AI推理任务保障策略做近实时的AI推理任务协作编排,提高了RAN侧资源分配和协作计算效率,提升了系统性能。
具体地,在步骤S11中,所述非实时无线控制器根据AI推理模型的模型信息生成所述AI推理任务保障策略;其中,所述模型信息包括所述AI推理模型的特征信息和AI推理任务的性能保障参数。
示例性的,所述模型信息由所述非实时无线控制器向服务管理和流程编排框架(Service Management and Orchestration,SMO)获取得到,而对于SMO中所述模型信息的来源,可以由运营商直接配置或者与外部应用交互获取,当通过外部应用交互获取所述模型信息时,可预先设计相关接口或应用程序接口(Application Programming Interface,API),通过授权和认证机制与外部应用建立连接,以此获取所述模型信息。
示例性的,所述AI推理模型的特征信息包括AI推理模型标识(IDentitifier,ID)、模型层数、各层类型、各层输出数据量、各层计算量、切分点位置、退出点位置、退出点准确度和退出点分类器计算量等。其中,每一所述特征信息对应的含义可参考表1,具体示例可参考表2。
表1特征参数及其对应含义
表2 ResNet-18模型的特征信息示例
在表2中,conv表示卷积层,max pool表示最大池化层,avg pool表示平均池化层,fc表示全连接层,Exit表示退出点位置,ResNet-18模型一共有四个退出点,分别为位于第5层的Exit1、位于第9层的Exit2、位于第13层的Exit3和位于第18层的Exit4,剩余的其余层可作为切分点。所述退出点准确度可以评估在选择这一层作为退出点时,其分类准确度,表格中的准确度仅为示例。
示例性的,所述性能保障参数包括但不限于模型推理准确度、单位时间的推理计算次数、模型推理往返时延和模型频谱效率。其中,每一所述性能保障参数对应的含义可参考表3。
表3性能保障参数及其对应含义
进一步地,所述非实时无线控制器在接收到所述模型信息后,根据所述模型信息生成AI推理任务保障策略,所述AI推理任务保障策略中定义了无线接入网络的优化目标、可优化参数和AI推理模型的特征信息。所述无线接入网络的优化目标指的是当前系统中保障策略的整体优化目标,例如最大化系统中模型推理准确度、最小化系统中模型推理往返时延等;所述可优化参数指的是所述近实时无线控制器可以为AI推理任务保障做调节的参数,例如AI推理模型的切分点选择、退出点选择等。假设所述无线接入网络的优化目标为:最大化系统中模型推理准确度、最小化系统中模型推理往返时延,即在选择完所述切分点和所述退出点后,需要满足模型推理准确度最大,以及满足模型推理往返时延最小,可通过对每一切分点和退出点进行轮询计算得到。
具体地,在步骤S12中,所述终端信息包括但不限于:计算资源信息、通信资源信息、当前小区服务用户信息、终端对AI推理任务的偏好信息。其中,所述计算资源信息包括但不限于:基站和终端的中央处理器(Central Processing Unit,CPU)利用率、CPU频率、CPU核绑定状态、剩余CPU核数量、CPU/图形处理器(Graphics Processing Unit,GPU)/嵌入式神经网络处理器(Neural-network Processing Unit,NPU)每秒浮点运算次数(Floating Point Operations Per Second,FLOPS)、GPU显存容量/剩余显存、GPU cuda核心数/剩余核心数等;所述通信资源信息包括但不限于:基站的上/下行总带宽、剩余带宽、物理资源块(Physical Resource Block,PRB)数、PRB利用率等,以及还可以包括终端的信道条件、本基站到邻站的通信链路时延等,例如终端的信道条件包括信噪比(Signal-to-Noise Ratio,SNR)、接收的信号强度指示(Received Signal Strength Indication,RSSI)等;所述当前小区服务用户信息包括但不限于:用户类型标识(非AI用户或AI用户)、AI推理任务模型标识(例如:ResNet-18、视觉几何组(Visual Geometry Group,VGG)等)等;所述终端对AI推理任务的偏好信息包括但不限于:对每种AI推理任务的切分点选择偏好等。
值得说明的是,当所述终端为基站或所述集中式单元和所述分布式单元所在基站时,所述终端信息可由基站直接发送给所述近实时无线控制器。当所述终端为用户设备时,本公开实施例中提供两种获取用户设备的终端信息的方式,第一种是通过基站获取,第二种是直接通过用户设备获取。本公开实施例所述的基站可能有多种形式,比如宏基站、微基站、中继站或接入点等。该基站可以是一体化基站,或者可以是包括集中式单元CU和分布式单元DU的基站。
在第一种实施方式中,当所述终端为用户设备时,所述获取至少一个终端的终端信息,包括:向用户设备所驻留的基站发送请求消息,以使所述基站根据所述请求消息上报所述用户设备的终端信息;或,向用户设备所驻留的基站订阅所述终端信息,以使所述基站上报所述用户设备的终端信息。
示例性的,参见图3,近实时无线控制器通过E2接口向用户设备所驻留的基站订阅或请求所述终端信息,可以是周期性订阅或事件性触发上报,事件性触发包括但不限于:基站算力波动、终端算力波动、基站带宽波动、终端电量波动、终端切分模式偏好变化等。基站根据需求采集用户设备的空口及终端数据,汇总成终端信息,并向近实时无线控制器上报所述用户设备的终端信息。另外,所述基站还将其自身的终端信息同步上报给所述近实时无线控制器。
在第二种实施方式中,当所述终端为用户设备时,所述用户设备支持与近实时无线控制器的信息交互;则,所述获取至少一个终端的终端信息,包括:向所述用户设备发送请求消息,以使所述用户设备根据所述请求消息上报终端信息;或,向所述用户设备订阅所述终端信息,以使所述用户设备上报终端信息。
示例性的,参见图4,所述用户设备与所述近实时无线控制器在现有协议栈基础上使得双方支持对应的接口应用层协议,从而实现双方的信息交互。此时近实时无线控制器通过E2接口向用户设备订阅或请求所述终端信息,可以是周期性订阅或事件性触发上报,用户设备根据需求采集空口及终端数据,汇总成终端信息,并向近实时无线控制器上报所述终端信息。所述近实时无线控制器同步发送请求消息给所述基站或订阅基站的终端信息,以使所述基站上报其自身的终端信息。
具体地,在步骤S13中,所述根据所述AI推理任务保障策略和所述终端信息生成无线接入网络的AI推理任务编排方案,包括:
S131、按照所述无线接入网络的优化目标,根据每一终端信息确定待分配节点和对应的计算节点,以及根据每一终端信息从所述AI推理模型的特征信息中选择对应的切分点和退出点;其中,所述计算节点为基站,所述切分点和所述退出点为所述AI推理模型中的其中一层;
S132、以所述待分配节点及其对应的计算节点、所述切分点和所述退出点组成所述AI推理任务编排方案。
示例性的,所述待分配节点为自身无法满足计算需求,且需要借助计算节点分担推理计算任务的终端;所述计算节点为可接收所述待分配节点的推理计算任务的终端。所述近实时无线控制器在接收到至少一个终端发送的终端信息后,确定待分配节点和计算节点,会应用优化算法(如差分算法),评估这些待分配节点所需的计算节点,以及进行切分点和退出点的选择。如根据所述计算资源信息确定其中一个待分配节点需要较大计算量时,可以分配较多的层数(即切分点和退出点之间的层数相隔较多)给这一待分配节点,反之则可以分配较少的层数;如根据所述通信资源信息可以确定这一待分配节点需要剩余带宽较大的计算节点时,优先分配满足条件的计算节点作为这一待分配节点的计算节点;如此时其中一个终端信息中给出了AI推理任务模型标识,则在生成AI推理任务编排方案时,会根据AI推理任务模型标识选择对应的AI推理模型进行推理任务;如此时其中一个终端信息中给出了对每种AI推理任务的切分点选择偏好,则在生成AI推理任务编排方案时,会优先根据这一待分配节点的切分点选择偏好去分配切分点。
具体地,在根据所述AI推理任务保障策略和所述终端信息生成无线接入网络的AI推理任务编排方案后,所述方法还包括:
S14、将所述AI推理任务编排方案发送给待分配节点和对应的计算节点,以使所述待分配节点和对应的计算节点完成各自的推理计算任务。
示例性的,假设所述待分配节点为用户设备,所述计算节点为基站,此时收集到三个用户设备和两个基站的终端信息,两个基站和三个用户设备满足以下情况:基站A为用户设备1提供通信连接服务,基站B为用户设备2、3提供通信连接服务,基站A和B均提供计算服务,AI推理模型皆为表2中的ResNet-18。假设此时编排方案是为用户设备1~3选择合适的计算节点,并为AI推理任务选择其处理的切分点和退出点,AI推理任务编排方案如下表4所示。
表4 AI推理任务编排方案示例
示例性的,按照表4以用户设备1为例进行说明,此时为用户设备1选择的计算节点为基站A,即由基站A与用户设备1共同完成运算任务,切分点为第3层,退出点为Exit2,则基站A完成的计算量为第3层到第9层的计算量,根据选择的切分点,由用户设备1完成前面几层的推理计算,切分点这一层生成的特征图传输给基站A后,由基站A完成从第3层到第9层的推理计算,最终将生成的结果传输回用户设备1。因此,切分点和退出点的选择决定了待分配节点和计算节点各自的计算量,退出点的选择也决定了该模型的准确度。
相比于相关技术,本公开的无线接入网络的AI推理任务编排方法,在RAN侧引入人工智能技术,首先接收非实时无线控制器发送的AI推理模型的AI推理任务保障策略,以及获取至少一个终端的终端信息,其中,AI推理任务保障策略中定义了无线接入网络的优化目标和AI推理模型的特征信息,终端信息包括计算资源信息和通信资源信息;然后根据所述AI推理任务保障策略和所述终端信息生成无线接入网络的AI推理任务编排方案,通过对计算节点、任务切分点和退出点的选择,能根据实时的通信状态和计算能力,对RAN侧的计算节点做近实时的AI推理任务协作编排,提高了RAN侧资源分配和协作计算的效率,进而提升了系统性能。
参见图5,图5是本公开实施例提供的另一种无线接入网络的AI推理任务编排方法的流程图,所述无线接入网络的AI推理任务编排方法由非实时无线控制器执行实现,所述方法包括:
S21、获取AI推理模型的模型信息;
S22、根据所述模型信息生成AI推理任务保障策略;
S23、将所述AI推理任务保障策略发送给近实时无线控制器,以使所述近实时无线控制器根据所述AI推理任务保障策略和终端信息生成无线接入网络的AI推理任务编排方案。
具体地,所述模型信息包括所述AI推理模型的特征信息和AI推理任务的性能保障参数;其中,所述性能保障参数包括但不限于模型推理准确度和单位时间的推理计算次数。
值得说明的是,本公开实施例所述的无线接入网络的AI推理任务编排方法的具体工作过程可参考上述实施例,在此不再赘述。
参见图6,图6是本公开实施例提供的一种无线接入网络的AI推理任务编排装置100的结构框图,所述无线接入网络的AI推理任务编排装置100包括:
AI推理任务保障策略接收模块11,用于接收非实时无线控制器发送的AI推理模型的AI推理任务保障策略;其中,所述AI推理模型部署在无线接入网络中,所述AI推理任务保障策略中定义了无线接入网络的优化目标和AI推理模型的特征信息;
终端信息获取模块12,用于获取至少一个终端的终端信息;其中,所述终端信息包括计算资源信息和通信资源信息;
AI推理任务编排方案生成模块13,用于根据所述AI推理任务保障策略和所述终端信息生成无线接入网络的AI推理任务编排方案。
具体地,所述AI推理任务编排方案生成模块13具体用于:
按照所述无线接入网络的优化目标,根据每一终端信息确定待分配节点和对应的计算节点,以及根据每一终端信息从所述AI推理模型的特征信息中选择对应的切分点和退出点;其中,所述计算节点为基站,所述切分点和所述退出点为所述AI推理模型中的其中一层;
以所述待分配节点及其对应的计算节点、所述切分点和所述退出点组成所述AI推理任务编排方案。
具体地,所述无线接入网络的AI推理任务编排装置100还包括:
AI推理任务编排方案法发送模块,用于将所述AI推理任务编排方案发送给待分配节点和对应的计算节点,以使所述待分配节点和对应的计算节点完成各自的推理计算任务。
具体地,当所述终端为用户设备时,所述终端信息获取模块12具体用于:向用户设备所驻留的基站发送请求消息,以使所述基站根据所述请求消息上报所述用户设备的终端信息;或,向用户设备所驻留的基站订阅所述终端信息,以使所述基站上报所述用户设备的终端信息。
具体地,当所述终端为用户设备时,所述用户设备支持与近实时无线控制器的信息交互;则,所述终端信息获取模块12具体用于:向所述用户设备发送请求消息,以使所述用户设备根据所述请求消息上报终端信息;或,向所述用户设备订阅所述终端信息,以使所述用户设备上报终端信息。
参见图7,图7是本公开实施例提供的另一种无线接入网络的AI推理任务编排装置200的结构框图,所述无线接入网络的AI推理任务编排装置200包括:
模型信息获取模块21,用于获取AI推理模型的模型信息;
AI推理任务保障策略生成模块22,用于根据所述模型信息生成AI推理任务保障策略;其中,所述AI推理模型部署在无线接入网络中,所述AI推理任务保障策略中定义了无线接入网络的优化目标和AI推理模型的特征信息;
AI推理任务保障策略发送模块23,用于将所述AI推理任务保障策略发送给近实时无线控制器,以使所述近实时无线控制器根据所述AI推理任务保障策略和终端信息生成无线接入网络的AI推理任务编排方案;其中,所述终端信息包括计算资源信息和通信资源信息。
具体地,所述模型信息包括所述AI推理模型的特征信息和AI推理任务的性能保障参数;其中,所述性能保障参数包括但不限于模型推理准确度和单位时间的推理计算次数。
相比于相关技术,本公开的无线接入网络的AI推理任务编排装置,在RAN侧引入人工智能技术,首先接收非实时无线控制器发送的AI推理模型的AI推理任务保障策略,以及获取至少一个终端的终端信息,其中,AI推理任务保障策略中定义了无线接入网络的优化目标和AI推理模型的特征信息,终端信息包括计算资源信息和通信资源信息;然后根据所述AI推理任务保障策略和所述终端信息生成无线接入网络的AI推理任务编排方案,通过对计算节点、任务切分点和退出点的选择,能根据实时的通信状态和计算能力,对RAN侧的计算节点做近实时的AI推理任务协作编排,提高了RAN侧资源分配和协作计算的效率,进而提升了系统性能。
参见图8,图8是本公开实施例提供的一种电子设备300的结构框图,所述电子设备300包括处理器31、存储器32以及存储在所述存储器32中并可在所述处理器31上运行的计算机程序。所述处理器31执行所述计算机程序时实现上述各个无线接入网络的AI推理任务编排方法实施例中的步骤,比如步骤S11~S13、S21~S23。
示例性的,所述计算机程序可以被分割成一个或多个模块/单元,所述一个或者多个模块/单元被存储在所述存储器32中,并由所述处理器31执行,以完成本公开。所述一个或多个模块/单元可以是能够完成特定功能的一系列计算机程序指令段,该指令段用于描述所述计算机程序在所述电子设备300中的执行过程。
所述电子设备300可包括,但不仅限于,处理器31、存储器32。本领域技术人员可以理解,所述示意图仅仅是电子设备300的示例,并不构成对电子设备300的限定,可以包括比图示更多或更少的部件,或者组合某些部件,或者不同的部件,例如所述电子设备300还可以包括输入输出设备、网络接入设备、总线等。
所述处理器31可以是中央处理单元(Central Processing Unit,CPU),还可以是其他通用处理器、数字信号处理器(Digital Signal Processor,DSP)、专用集成电路(Application Specific Integrated Circuit,ASIC)、现成可编程门阵列(Field-Programmable Gate Array,FPGA)或者其他可编程逻辑器件、分立门或者晶体管逻辑器件、分立硬件组件等。通用处理器可以是微处理器或者该处理器也可以是任何常规的处理器等,所述处理器31是所述电子设备300的控制中心,利用各种接口和线路连接整个电子设备300的各个部分。
所述存储器32可用于存储所述计算机程序和/或模块,所述处理器31通过运行或执行存储在所述存储器32内的计算机程序和/或模块,以及调用存储在存储器32内的数据,实现所述电子设备300的各种功能。所述存储器32可主要包括存储程序区和存储数据区,其中,存储程序区可存储操作系统、至少一个功能所需的应用程序(比如声音播放功能、图像播放功能等)等;存储数据区可存储根据手机的使用所创建的数据(比如音频数据、电话本等)等。此外,存储器32可以包括高速随机存取存储器,还可以包括非易失性存储器,例如硬盘、内存、插接式硬盘,智能存储卡(Smart Media Card,SMC),安全数字(Secure Digital,SD)卡,闪存卡(Flash Card)、至少一个磁盘存储器件、闪存器件、或其他易失性固态存储器件。
其中,所述电子设备300集成的模块/单元如果以软件功能单元的形式实现并作为独立的产品销售或使用时,可以存储在一个计算机可读取存储介质中。基于这样的理解,本公开实现上述实施例方法中的全部或部分流程,也可以通过计算机程序来指令相关的硬件来完成,所述的计算机程序可存储于一计算机可读存储介质中,该计算机程序在被处理器31执行时,可实现上述各个方法实施例的步骤。其中,所述计算机程序包括计算机程序代码,所述计算机程序代码可以为源代码形式、对象代码形式、可执行文件或某些中间形式等。所述计算机可读介质可以包括:能够携带所述计算机程序代码的任何实体或装置、记录介质、U盘、移动硬盘、磁碟、光盘、计算机存储器、只读存储器(Read-Only Memory,ROM)、随机存取存储器(Random Access Memory,RAM)、电载波信号、电信信号以及软件分发介质等。
以上所述是本公开的优选实施方式,应当指出,对于本技术领域的普通技术人员来说,在不脱离本公开原理的前提下,还可以做出若干改进和润饰,这些改进和润饰也视为本公开的保护范围。
Claims (12)
- 一种无线接入网络的AI推理任务编排方法,应用于近实时无线控制器,所述方法包括:接收非实时无线控制器发送的AI推理模型的AI推理任务保障策略;其中,所述AI推理模型部署在无线接入网络中,所述AI推理任务保障策略中定义了无线接入网络的优化目标和AI推理模型的特征信息;获取至少一个终端的终端信息;其中,所述终端信息包括计算资源信息和通信资源信息;根据所述AI推理任务保障策略和所述终端信息生成无线接入网络的AI推理任务编排方案。
- 如权利要求1所述的无线接入网络的AI推理任务编排方法,其中,所述根据所述AI推理任务保障策略和所述终端信息生成无线接入网络的AI推理任务编排方案,包括:按照所述无线接入网络的优化目标,根据每一终端信息确定待分配节点和对应的计算节点,以及根据每一终端信息从所述AI推理模型的特征信息中选择对应的切分点和退出点;其中,所述切分点和所述退出点为所述AI推理模型中的其中一层;以所述待分配节点及其对应的计算节点、所述切分点和所述退出点组成所述AI推理任务编排方案。
- 如权利要求2所述的无线接入网络的AI推理任务编排方法,其中,在根据所述AI推理任务保障策略和所述终端信息生成无线接入网络的AI推理任务编排方案后,所述方法还包括:将所述AI推理任务编排方案发送给待分配节点和对应的计算节点,以使所述待分配节点和对应的计算节点完成各自的推理计算任务。
- 如权利要求1所述的无线接入网络的AI推理任务编排方法,其中,当所述终端为用户设备时,所述获取至少一个终端的终端信息,包括:向用户设备所驻留的基站发送请求消息,以使所述基站根据所述请求消息上报所述用户设备的终端信息;或,向用户设备所驻留的基站订阅所述终端信息,以使所述基站上报所述用户设备的终端信息。
- 如权利要求1所述的无线接入网络的AI推理任务编排方法,其中,当所述终端为用户设备时,所述用户设备支持与近实时无线控制器的信息交互;则,所述获取至少一个终端的终端信息,包括:向所述用户设备发送请求消息,以使所述用户设备根据所述请求消息上报终端信息;或,向所述用户设备订阅所述终端信息,以使所述用户设备上报终端信息。
- 一种无线接入网络的AI推理任务编排方法,应用于非实时无线控制器,所述方法包括:获取AI推理模型的模型信息;根据所述模型信息生成AI推理任务保障策略;其中,所述AI推理模型部署在无线接入网络中,所述AI推理任务保障策略中定义了无线接入网络的优化目标和AI推理模型的特征信息;将所述AI推理任务保障策略发送给近实时无线控制器,以使所述近实时无线控制器根据所述AI推理任务保障策略和终端信息生成无线接入网络的AI推理任务编排方案;其中,所述终端信息包括计算资源信息和通信资源信息。
- 如权利要求6所述的无线接入网络的AI推理任务编排方法,其中,所述模型信息包括所述AI推理模型的特征信息和AI推理任务的性能保障参数;其中,所述性能保障参数包括但不限于模型推理准确度和单位时间的推理计算次数。
- 一种无线接入网络的AI推理任务编排装置,包括:AI推理任务保障策略接收模块,用于接收非实时无线控制器发送的AI推理模型的AI推理任务保障策略;其中,所述AI推理模型部署在无线接入网络中,所述AI推理任务保障策略中定义了无线接入网络的优化目标和AI推理模型的特征信息;终端信息获取模块,用于获取至少一个终端的终端信息;其中,所述终端信息包括计算资源信息和通信资源信息;AI推理任务编排方案生成模块,用于根据所述AI推理任务保障策略和所述终端信息生成无线接入网络的AI推理任务编排方案。
- 一种无线接入网络的AI推理任务编排装置,包括:模型信息获取模块,用于获取AI推理模型的模型信息;AI推理任务保障策略生成模块,用于根据所述模型信息生成AI推理任务保障策略;其中,所述AI推理模型部署在无线接入网络中,所述AI推理任务保障策略中定义了无线接入网络的优化目标和AI推理模型的特征信息;AI推理任务保障策略发送模块,用于将所述AI推理任务保障策略发送给近实时无线控制器,以使所述近实时无线控制器根据所述AI推理任务保障策略和终端信息生成无线接入网络的AI推理任务编排方案;其中,所述终端信息包括计算资源信息和通信资源信息。
- 一种电子设备,包括处理器、存储器以及存储在所述存储器中且被配置为由所述处理器执行的计算机程序,所述处理器执行所述计算机程序时实现如权利要求1至7中任意一项所述的无线接入网络的AI推理任务编排方法。
- 一种计算机可读存储介质,所述计算机可读存储介质包括存储的计算机程序,其中,在所述计算机程序运行时控制所述计算机可读存储介质所在设备执行如权利要求1至7中任意一项所述的无线接入网络的AI推理任务编排方法。
- 一种计算机程序产品,包括计算机指令,所述计算机指令被处理器执行时实现如权利要求1至7中任一项所述的方法中的步骤。
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN202311733048.9A CN120166409B (zh) | 2023-12-15 | 2023-12-15 | 无线接入网络的ai推理任务编排方法、装置、设备和存储介质 |
| CN202311733048.9 | 2023-12-15 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2025124306A1 true WO2025124306A1 (zh) | 2025-06-19 |
Family
ID=96002545
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/CN2024/137406 Pending WO2025124306A1 (zh) | 2023-12-15 | 2024-12-06 | 无线接入网络的ai推理任务编排方法、装置、设备和存储介质 |
Country Status (2)
| Country | Link |
|---|---|
| CN (1) | CN120166409B (zh) |
| WO (1) | WO2025124306A1 (zh) |
Families Citing this family (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN120529342B (zh) * | 2025-07-25 | 2025-11-07 | 中国电信股份有限公司 | 模型推理方法、装置、网络设备、可读存储介质和程序产品 |
Citations (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN113875278A (zh) * | 2019-05-24 | 2021-12-31 | 苹果公司 | 5g新空口负载平衡和移动性稳健性 |
| CN113923694A (zh) * | 2021-12-14 | 2022-01-11 | 网络通信与安全紫金山实验室 | 网络资源编排方法、系统、装置及存储介质 |
| CN115988580A (zh) * | 2022-12-28 | 2023-04-18 | 中国联合网络通信集团有限公司 | 无线资源控制方法、装置、电子设备及存储介质 |
Family Cites Families (6)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US11902104B2 (en) * | 2020-03-04 | 2024-02-13 | Intel Corporation | Data-centric service-based network architecture |
| WO2022060777A1 (en) * | 2020-09-17 | 2022-03-24 | Intel Corporation | Online reinforcement learning |
| JP2024527213A (ja) * | 2021-07-06 | 2024-07-24 | インテル コーポレイション | オープン無線アクセスネットワーク(Open Radio Access Network (O-RAN))システムのA1ポリシ機能 |
| WO2023172292A2 (en) * | 2021-08-25 | 2023-09-14 | Northeastern University | Zero-touch deployment and orchestration of network intelligence in open ran systems |
| CN116301912A (zh) * | 2022-12-14 | 2023-06-23 | 平安银行股份有限公司 | 推理服务部署方法、装置、设备和存储介质 |
| CN116708443A (zh) * | 2023-07-24 | 2023-09-05 | 中国电信股份有限公司 | 多层次算力网络任务调度方法及装置 |
-
2023
- 2023-12-15 CN CN202311733048.9A patent/CN120166409B/zh active Active
-
2024
- 2024-12-06 WO PCT/CN2024/137406 patent/WO2025124306A1/zh active Pending
Patent Citations (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN113875278A (zh) * | 2019-05-24 | 2021-12-31 | 苹果公司 | 5g新空口负载平衡和移动性稳健性 |
| CN113923694A (zh) * | 2021-12-14 | 2022-01-11 | 网络通信与安全紫金山实验室 | 网络资源编排方法、系统、装置及存储介质 |
| CN115988580A (zh) * | 2022-12-28 | 2023-04-18 | 中国联合网络通信集团有限公司 | 无线资源控制方法、装置、电子设备及存储介质 |
Non-Patent Citations (1)
| Title |
|---|
| 薛旭等 (XUE, XU ET AL.): "开放智能无线网络架构和平台设计研究 (Exploring the Design of an Open Intelligent Wireless Network Architecture and Platform)", 移动通信 (MOBILE COMMUNICATIONS), no. 523, 11 April 2023 (2023-04-11), XP009564682 * |
Also Published As
| Publication number | Publication date |
|---|---|
| CN120166409B (zh) | 2026-01-16 |
| CN120166409A (zh) | 2025-06-17 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| CN110839184B (zh) | 基于流量预测的移动前传光网络带宽调整方法及装置 | |
| Chen et al. | Exploiting massive D2D collaboration for energy-efficient mobile edge computing | |
| TWI848291B (zh) | 算力資源調度的方法以及相關裝置 | |
| EP3855842A1 (en) | Method and apparatus for dynamically allocating radio resources in a wireless communication system | |
| CN114007225A (zh) | Bwp的分配方法、装置、电子设备及计算机可读存储介质 | |
| US20250321788A1 (en) | Computing task processing method and related apparatus | |
| Ebrahimzadeh et al. | Cooperative computation offloading in FiWi enhanced 4G HetNets using self-organizing MEC | |
| WO2019153849A1 (zh) | 策略驱动方法和装置 | |
| Gu et al. | Context-aware task offloading for multi-access edge computing: Matching with externalities | |
| WO2022087930A1 (zh) | 一种模型配置方法及装置 | |
| WO2024198976A1 (zh) | 一种切分点确定方法及装置 | |
| WO2024187479A1 (zh) | 一种算力资源调度方法及装置 | |
| WO2025124306A1 (zh) | 无线接入网络的ai推理任务编排方法、装置、设备和存储介质 | |
| CN104410982A (zh) | 一种无线异构网络中终端聚合与重构方法 | |
| Han et al. | Cell-less offloading of distributed learning tasks in multi-access edge computing | |
| Zhong et al. | Slice allocation of 5G network for smart grid with deep reinforcement learning ACKTR | |
| WO2025130494A1 (zh) | 一种通信方法及装置 | |
| WO2025082062A1 (zh) | 人工智能服务的处理方法及装置 | |
| CN114158078B (zh) | 网络切片管理方法、装置及计算机可读存储介质 | |
| CN118803671A (zh) | 业务处理方法、装置、设备、存储介质及程序产品 | |
| WO2024113288A1 (zh) | 通信方法和通信装置 | |
| CN116737260A (zh) | 基于人工蜂-鱼群算法的计算卸载方法、装置及系统 | |
| CN118175580A (zh) | 一种切片配置方法、装置和存储介质 | |
| WO2025166728A1 (en) | Methods and devices for digital twin based wireless communication performance optimization | |
| CN118540023B (zh) | 数据包处理方法、装置、终端和存储介质 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 24902727 Country of ref document: EP Kind code of ref document: A1 |