WO2026012196A1 - 计算服务方法、装置、第一节点及第二节点 - Google Patents
计算服务方法、装置、第一节点及第二节点Info
- Publication number
- WO2026012196A1 WO2026012196A1 PCT/CN2025/105250 CN2025105250W WO2026012196A1 WO 2026012196 A1 WO2026012196 A1 WO 2026012196A1 CN 2025105250 W CN2025105250 W CN 2025105250W WO 2026012196 A1 WO2026012196 A1 WO 2026012196A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- computing
- node
- task
- computing service
- request message
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Pending
Links
Classifications
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04W—WIRELESS COMMUNICATION NETWORKS
- H04W4/00—Services specially adapted for wireless communication networks; Facilities therefor
- H04W4/50—Service provisioning or reconfiguring
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04L—TRANSMISSION OF DIGITAL INFORMATION, e.g. TELEGRAPHIC COMMUNICATION
- H04L67/00—Network arrangements or protocols for supporting network services or applications
- H04L67/50—Network services
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04L—TRANSMISSION OF DIGITAL INFORMATION, e.g. TELEGRAPHIC COMMUNICATION
- H04L67/00—Network arrangements or protocols for supporting network services or applications
- H04L67/50—Network services
- H04L67/51—Discovery or management thereof, e.g. service location protocol [SLP] or web services
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04W—WIRELESS COMMUNICATION NETWORKS
- H04W28/00—Network traffic management; Network resource management
- H04W28/02—Traffic management, e.g. flow control or congestion control
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04W—WIRELESS COMMUNICATION NETWORKS
- H04W28/00—Network traffic management; Network resource management
- H04W28/02—Traffic management, e.g. flow control or congestion control
- H04W28/0268—Traffic management, e.g. flow control or congestion control using specific QoS parameters for wireless networks, e.g. QoS class identifier [QCI] or guaranteed bit rate [GBR]
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04W—WIRELESS COMMUNICATION NETWORKS
- H04W88/00—Devices specially adapted for wireless communication networks, e.g. terminals, base stations or access point devices
- H04W88/18—Service support devices; Network management devices
Definitions
- This application belongs to the field of communication technology, specifically relating to a computing service method, apparatus, first node, and second node.
- DFT Discrete Fourier Transform
- FFT Fast Fourier Transform
- AI Artificial Intelligence
- 5QI 5G QoS Identifier
- This application provides a computing service method, apparatus, first node, and second node, which helps to ensure the quality of computing services.
- a computing service method comprising:
- the first node receives a computing service request message from the second node, the computing service request message including the performance parameters of the computing service;
- the first node sends a computing service response message to the second node, wherein the computing service response message is used to indicate whether to accept or reject the computing service request message.
- a computing service device comprising:
- the receiving module is configured to receive a computing service request message from the second node, the computing service request message including performance parameters of the computing service;
- the sending module is used to send a computing service response message to the second node, wherein the computing service response message is used to indicate whether to accept or reject the computing service request message.
- a computing service method comprising:
- the second node sends a computing service request message to the first node, the computing service request message including the performance parameters of the computing service;
- the second node receives a computing service response message from the first node, the computing service response message being used to indicate whether to accept or reject the computing service request message.
- a computing service device comprising:
- the sending module is used to send a computing service request message to the first node, the computing service request message including the performance parameters of the computing service;
- the receiving module is configured to receive a computing service response message from the first node, wherein the computing service response message is used to indicate whether to accept or reject the computing service request message.
- an apparatus for computing services is provided, the apparatus being configured to perform the steps of the method described in the first aspect, or to implement the steps of the method described in the third aspect.
- a first node including a processor and a memory, the memory storing a program or instructions executable on the processor, the program or instructions, when executed by the processor, implementing the steps of the method as described in the first aspect.
- a first node including a processor and a communication interface, wherein the communication interface is used to receive a computing service request message from a second node, the computing service request message including performance parameters of a computing service;
- the communication interface is also used to send a computing service response message to the second node, wherein the computing service response message is used to indicate whether to accept or reject the computing service request message.
- a second node including a processor and a memory, the memory storing a program or instructions executable on the processor, the program or instructions, when executed by the processor, implementing the steps of the method as described in the third aspect.
- a second node including a processor and a communication interface, wherein the communication interface is used to send a computing service request message to a first node, the computing service request message including performance parameters of a computing service;
- the communication interface is also used to receive a computing service response message from the first node, the computing service response message being used to indicate whether to accept or reject the computing service request message.
- a readable storage medium on which a program or instructions are stored, which, when executed by a processor, implement the steps of the method described in the first aspect, or implement the steps of the method described in the third aspect.
- a wireless communication system comprising: a first node and a second node, wherein the first node is configured to perform the steps of the computing service method as described in the first aspect, and the second node is configured to perform the steps of the computing service method as described in the third aspect.
- a chip including a processor and a communication interface coupled to the processor, the processor being configured to run a program or instructions to implement the steps of the method described in the first aspect, or to implement the steps of the method described in the third aspect.
- a computer program/program product is provided, which is stored in a storage medium and is executed by at least one processor to implement the steps of the method as described in the first aspect, or to implement the steps of the method as described in the third aspect.
- a first node receives a computing service request message from a second node, the computing service request message including performance parameters of the computing service; the first node sends a computing service response message to the second node, wherein the computing service response message is used to indicate whether to accept or reject the computing service request message. That is, by carrying the performance parameters of the computing service in the computing service request message, the first node can determine whether to accept the computing service request message based on the performance parameters of the computing service, which helps to ensure the service quality of the computing service.
- Figure 1 is a block diagram of a wireless communication system applicable to an embodiment of this application
- FIG. 2 is a flowchart of a computing service method provided in an embodiment of this application.
- FIG. 3 is a flowchart of another computing service method provided in an embodiment of this application.
- FIG. 4 is a flowchart of another computing service method provided in an embodiment of this application.
- FIG. 5 is a flowchart of another computing service method provided in an embodiment of this application.
- FIG. 6 is a flowchart of another computing service method provided in an embodiment of this application.
- Figure 7 is a structural diagram of a computing service device provided in an embodiment of this application.
- Figure 8 is a structural diagram of another computing service device provided in an embodiment of this application.
- Figure 9 is a structural diagram of the communication device provided in an embodiment of this application.
- Figure 10 is a structural diagram of a network-side device provided in an embodiment of this application.
- FIG 11 is a structural diagram of another network-side device provided in an embodiment of this application.
- Figure 12 is a structural diagram of the terminal provided in an embodiment of this application.
- first and second are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such terms can be used interchangeably where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by “first” and “second” are generally of the same class, not limited in number; for example, the first object can be one or more.
- “or” in this application indicates at least one of the connected objects.
- the scope of protection for "A or B” covers at least three scenarios: Scenario 1: including A but not B; Scenario 2: including B but not A; Scenario 3: including both A and B.
- the terms “A and/or B,” “at least one of A and B,” and “at least one of A or B” also cover at least the above three scenarios.
- the character “/” generally indicates that the preceding and following objects are in an "or” relationship.
- instruction in this application can be either a direct instruction (or explicit instruction) or an indirect instruction (or implicit instruction).
- a direct instruction can be understood as one in which the sender explicitly informs the receiver of specific information, the operation to be performed, or the requested result, etc., in the instruction sent.
- An indirect instruction can be understood as one in which the receiver determines the corresponding information based on the instruction sent by the sender, or makes a judgment and determines the operation to be performed or the requested result, etc., based on the judgment result.
- LTE Long Term Evolution
- LTE-A Long Term Evolution-Advanced
- CDMA Code Division Multiple Access
- TDMA Time Division Multiple Access
- FDMA Frequency Division Multiple Access
- OFDMA Orthogonal Frequency Division Multiple Access
- SC-FDMA Single-carrier Frequency-Division Multiple Access
- NR New Radio
- FIG. 1 shows a block diagram of a wireless communication system applicable to an embodiment of this application.
- the wireless communication system includes a terminal 11 and a network-side device 12.
- Terminal 11 can be a mobile phone, tablet computer, laptop computer, notebook computer, personal digital assistant (PDA), handheld computer, netbook, ultra-mobile personal computer (UMPC), mobile internet device (MID), augmented reality (AR), virtual reality (VR) device, robot, wearable device, flight vehicle, vehicle user equipment (VUE), shipboard equipment, pedestrian user equipment (PUE), smart home (home devices with wireless communication capabilities, such as refrigerators, televisions, washing machines, or furniture), game console, personal computer (PC), ATM, or self-service machine, etc.
- PDA personal digital assistant
- UMPC ultra-mobile personal computer
- MID mobile internet device
- AR augmented reality
- VR virtual reality
- robot wearable device
- flight vehicle vehicle user equipment
- VUE shipboard equipment
- pedestrian user equipment PUE
- smart home home devices with wireless communication capabilities, such as refrigerators, televisions, washing machines
- Wearable devices include: smartwatches, smart bracelets, smart headphones, smart glasses, smart jewelry (smart bracelets, smart chains, smart rings, smart necklaces, smart anklets, smart anklets, etc.), smart wristbands, smart clothing, etc.
- in-vehicle devices can also be referred to as in-vehicle terminals, in-vehicle controllers, in-vehicle modules, in-vehicle components, in-vehicle chips, or in-vehicle units, etc. It should be noted that the specific type of terminal 11 is not limited in this application embodiment.
- Network-side equipment 12 may include access network equipment or core network equipment, wherein access network equipment may also be referred to as Radio Access Network (RAN) equipment, radio access network function, or radio access network unit.
- Access network equipment may include base stations, Wireless Local Area Network (WLAN) access points (APs), or Wireless Fidelity (WiFi) nodes, etc.
- WLAN Wireless Local Area Network
- WiFi Wireless Fidelity
- a base station may be referred to as a Node B (NB), Evolved Node B (eNB), Next Generation Node B (gNB), New Radio Node B (NR Node B), Access Point, Relay Base Station (RBS), Serving Base Station (SBS), Base Transceiver Station (BTS), Radio Base Station, Radio Transceiver, Basic Service Set (BSS), Extended Service Set (ESS), Home Node B (HNB), Home Evolved Node B, Transmit/Receive Point (TRP), or any other suitable term in the relevant field, as long as the same technical effect is achieved.
- the base station is not limited to specific technical terms. It should be noted that in this application embodiment, only a base station in an NR system is used as an example for introduction, and the specific type of base station is not limited.
- Core network equipment also known as core network nodes, core network functions, or core network elements, includes, but is not limited to, at least one of the following: Mobility Management Entity (MME), Access and Mobility Management Function (AMF), Session Management Function (SMF), User Plane Function (UPF), Policy Control Function (PCF), Policy and Charging Rules Function (PCRF), Edge Application Server Discovery Function (EASDF), Unified Data Management (UDM), and Unified Data Warehouse (UDM).
- MME Mobility Management Entity
- AMF Access and Mobility Management Function
- SMF Session Management Function
- UPF User Plane Function
- PCF Policy Control Function
- PCF Policy and Charging Rules Function
- EASDF Edge Application Server Discovery Function
- UDM Unified Data Management
- UDM Unified Data Management
- UDM Unified Data Warehouse
- the core network equipment includes: Data Repository (UDR), Home Subscriber Server (HSS), Centralized Network Configuration (CNC), Network Repository Function (NRF), Network Exposure Function (NEF), Local NEF (or L-NEF), Binding Support Function (BSF), Application Function (AF), Location Management Function (LMF), Gateway Mobile Location Centre (GMLC), and Network Data Analytics Function (NWDAF).
- UDR Data Repository
- HSS Home Subscriber Server
- CNC Centralized Network Configuration
- NEF Network Exposure Function
- L-NEF Local NEF
- BSF Binding Support Function
- AF Application Function
- LMF Location Management Function
- GMLC Gateway Mobile Location Centre
- NWDAF Network Data Analytics Function
- the core network equipment can be implemented by one or more functional modules in a single device, or by multiple devices working together; this application does not specifically limit this. It is understood that the aforementioned functional modules can be network elements in hardware devices, software functional modules running on dedicated hardware, or virtualized functional modules instantiated on a platform (e.g., a cloud platform).
- a platform e.g., a cloud platform
- Positioning accuracy is used to represent the distribution of positioning service performance errors. It is defined using a confidence level and a positioning error threshold, which is the percentage of the distance between the positioning result and the actual location that is within the positioning error threshold range. For example, a 95% confidence level for positioning accuracy ⁇ 3m means that there is a 95% probability that the positioning result has a distance error of less than 3 meters from the actual location.
- Location QoS is included in the location request, which includes both the location request from the location requester and the location request from the location result provider (e.g., the Location Management Function (LMF)).
- LMF Location Management Function
- the LMF's location request is generated based on the location request from the location requester.
- the location request from the aforementioned location requester may include the following parameters:
- LCS QoS Class QoS type/level (LCS QoS Class), including:
- QoS Multiple QoS
- This is a moderately stringent QoS type for location, meaning it includes QoS requirements corresponding to multiple QoS levels. If the location result does not meet the most stringent QoS requirement, LMF will initiate the location process again to try to meet the lower QoS requirements until one of the QoS requirements is met. If the most lenient QoS requirement is still not met, no location result will be reported, only the reason for the failure will be reported.
- Positioning accuracy including horizontal positioning accuracy and/or vertical positioning accuracy
- the LMF should immediately report the initial location or the most recent location result of the target UE. If no location result is found, a failure message should be reported, and a location procedure can be triggered to respond to subsequent location requests.
- Low latency Response time is prioritized over accuracy.
- the LCS server should return the current location with minimal latency.
- Delay-insensitive Prioritizing accuracy over response time. LMF can delay the feedback of positioning results until the required positioning accuracy is met.
- the location request for the aforementioned LMF may include the following parameters:
- Horizontal accuracy includes accuracy and confidence level.
- Vertical accuracy includes both accuracy and confidence level.
- Response time is the delay between when the UE receives a location information request and when the location information is provided.
- AI and communications are among the 6G application scenarios identified by the International Telecommunication Union Radiocommunication Sector (ITU-R).
- ITU-R International Telecommunication Union Radiocommunication Sector
- Typical use cases include IMT-2030 (6G) assisted autonomous driving, autonomous collaboration between devices in healthcare applications, cross-device/network compute offloading, creation and prediction of digital twins, and IMT-2030 (6G) assisted collaborative robots.
- These application scenarios will require support for high mobile network capacity and user experience data rates, as well as low latency and high reliability.
- this application scenario is expected to include a set of new functionalities integrating AI and computing-related features into 6G systems, including data acquisition, preparation, and processing from various sources; distributed AI model training; model sharing and distributed inference across mobile communication systems; and compute resource orchestration.
- SLAs Service Level Agreements
- Figure 2 is a flowchart of a computing service method provided in an embodiment of this application. The method can be executed by a first node, as shown in Figure 2, and includes the following steps:
- Step 201 The first node receives a computing service request message from the second node, the computing service request message including the performance parameters of the computing service.
- the aforementioned first node may also be referred to as a computing management node, computing management function, computing control function, computing management and control function, computing service control function, or computing service management function, etc.
- the aforementioned first node may be a radio access network node or a core network node.
- the aforementioned first node may be a network node with at least one function, such as receiving computing service requests, processing computing service requests, scheduling computing resources, exchanging computing information, or processing computing data.
- the aforementioned first node may be a newly added network node; or it may be a node obtained by enhancing an existing network node, such as an enhanced AMF, enhanced SMF, or enhanced base station, etc.
- the aforementioned second node may include, but is not limited to, at least one of a terminal, an AF (Application Function), and a Network Function (NF).
- AF Application Function
- NF Network Function
- the second node is an AF
- the aforementioned computing service request message needs to be sent to the first node via NEF (Network Function Entity).
- NEF Network Function Entity
- the aforementioned second node can also be referred to as a computing request node.
- the aforementioned computing services may include at least one of DFT, FFT, AI model training, and AI model inference.
- the AI model training or AI model inference may include image recognition, object detection, semantic segmentation, recommendation, natural language processing, speech recognition, optical character recognition, face recognition, beam management, localization, perception, and Channel State Information (CSI) feedback.
- the aforementioned computing services may include at least one computing task.
- the performance parameters of the aforementioned computing service can be used to characterize the performance requirements that the computing service needs to meet.
- the performance parameters of the aforementioned computing service may include, but are not limited to, at least one of the following: hardware requirements for the computing task, computing requirements, computing result requirements, and transmission requirements.
- the aforementioned hardware requirements may include, but are not limited to, at least one of the following: computing power type, minimum memory, and minimum storage.
- the aforementioned computing requirements may include, but are not limited to, at least one of the following: minimum number of operations or operands, minimum computing speed, minimum computing intensity, computing power consumption threshold, and computing energy efficiency threshold.
- the aforementioned computing result requirements may include, but are not limited to, at least one of the following: maximum failure rate and lower limit of AI model performance.
- the aforementioned transmission requirements may include, but are not limited to, at least one of the following: minimum transmission bandwidth and transmission latency budget.
- the first node can determine whether to accept the computing service request message based on the performance parameters of the aforementioned computing service. For example, the first node can select a computing node based on the computing service request message. For instance, the first node can select a computing node that meets the performance parameters of the aforementioned computing service, and can determine whether to accept the computing service request message based on the selection result. It is understood that when selecting a computing node, in addition to the performance parameters of the aforementioned computing service, other parameters can also be considered, such as the status of the computing node, the type of computing service, etc.
- the number of compute nodes selected can be at least one, or it can be zero, meaning no suitable compute nodes are selected. For example, if no compute nodes exist that can meet the performance parameters of the computing service, the number of compute nodes selected is zero.
- Step 202 The first node sends a computing service response message to the second node, wherein the computing service response message is used to indicate whether to accept or reject the computing service request message.
- the above computing service response message is used to indicate acceptance of the computing service request message; when the number of selected computing nodes is 0, the above computing service response message is used to indicate rejection of the computing service request message.
- a first node receives a computing service request message from a second node, the computing service request message including performance parameters of the computing service; the first node sends a computing service response message to the second node, wherein the computing service response message is used to indicate whether to accept or reject the computing service request message. That is, by carrying the performance parameters of the computing service in the computing service request message, the first node can determine whether to accept the computing service request message based on the performance parameters of the computing service, which helps to ensure the service quality of the computing service.
- the performance parameters of the computing service include performance parameters of at least one computing task or at least one group of computing tasks, wherein the performance parameters of the computing task or group of computing tasks include at least one of the following:
- Resource type minimum number of operations or operands; minimum computing speed; minimum computing intensity; computing latency budget; maximum failure rate, which represents the ratio of failed requests to total requests per unit time; average window; maximum time, representing the upper limit of the computing task's duration; computing power type; data type; minimum memory, representing the lower limit of memory required for the computing task; minimum storage, representing the lower limit of storage required for the computing task; minimum transmission bandwidth, representing the lower limit of bandwidth required for the computing task; computing task arrival mode; parameters corresponding to the computing task arrival mode; AI model training accuracy; AI model performance lower limit; minimum throughput of AI model inference; computing power consumption threshold; computing energy efficiency threshold.
- the resource types mentioned above can be used to reflect QoS classes, such as guaranteed, non-guaranteed, and latency-sensitive. These resource types can also be referred to as computing QoS classes, or computing and communication QoS classes, etc.
- the resource type includes at least one of the following:
- GCR Guaranteed Computing Rate
- GCI Guaranteed Computing Intensity
- Non-Guaranteed Computing Rate Non-Guaranteed Computing Intensity
- Non-GCI Non-Guaranteed Computing Intensity
- DCR Delay-Critical Guaranteed Computing Rate
- DCI Delay-Critical Guaranteed Computing Intensity
- the above-mentioned guaranteed computational speed can also be called guaranteed operational rate (GOR), and the above-mentioned guaranteed computational intensity can also be called guaranteed operational intensity (GOI).
- GOR guaranteed operational rate
- GOI guaranteed operational intensity
- the computation speed characterized by at least one of the guaranteed computation speed, non-guaranteed computation speed, and latency-sensitive guaranteed computation speed can refer to the computation time complexity divided by the tolerable upper limit of computation time, wherein the computation time complexity can be represented, for example, by operands.
- the computational strength characterized by at least one of the guaranteed computational strength, non-guaranteed computational strength, and latency-sensitive guaranteed computational strength may refer to the computational speed divided by the bandwidth.
- the aforementioned guaranteed computing speed can be understood as the computing resources associated with the guaranteed computing speed of a computing task being permanently allocated during the computing task. For example, computing resources such as CPU, GPU, or storage on the selected computing node associated with the guaranteed computing speed of the computing task are allocated to that computing task. The computing resources allocated to that computing task are dedicated to that computing task and are not shared with other computing tasks.
- the aforementioned guaranteed computing strength can be understood as computing resources, or computing and communication resources, permanently allocated during the computing task to guarantee its computing strength.
- computing resources such as CPU, GPU, and memory related to the guaranteed computing strength are allocated to the computing task;
- computing strength is the value obtained by dividing the computing speed by the target bandwidth (i.e., the smaller of the memory bandwidth and the transmission bandwidth)
- CPU, GPU, memory, and network resources related to the guaranteed computing strength are allocated to the computing task.
- Network resources include air interface resources or wired transmission resources, etc. Air interface resources are related to the air interface transmission bandwidth of the computing task between the UE and the network, while wired transmission resources are related to the transmission bandwidth from the access network to the computing node of the computing task between the UE and the network.
- the aforementioned latency-sensitive guarantee of computational speed can be understood as follows: if the latency of a data packet in a computation task exceeds the computational latency budget, the data packet is discarded. Furthermore, the latency-sensitive guarantee of computational speed requires that a first proportion (e.g., 98%) of data packets do not exceed the computational latency budget.
- the aforementioned latency-sensitive guarantee of computational strength can also be understood as follows: if the latency of a data packet in the computation task exceeds the computational latency budget, the data packet is discarded. Furthermore, the latency-sensitive guarantee of computational speed requires that a second proportion (e.g., 98%) of the data packets cannot exceed the computational latency budget.
- the aforementioned non-guaranteed computing speed can be understood as resources related to the computing speed of a computing task not being permanently allocated during that computing task.
- computing resources such as CPU, GPU, or storage on the selected computing node that are related to the guaranteed computing speed of the computing task are not allocated to that computing task but are shared with other computing tasks.
- the aforementioned non-guaranteed computational intensity can be understood as resources related to the computational intensity of a computational task not being permanently allocated during that task.
- computational resources such as CPU, GPU, and memory, as well as network resources, on the selected computing node that are related to the guaranteed computational speed of the task, are not allocated to that task but are shared with other computational tasks.
- the above resource types can determine the allocation of computing resources related to the guaranteed computational load at the computing task level, or the allocation of computing resources related to the guaranteed computational intensity at the computing task level, or the allocation of computing and communication resources.
- the aforementioned computing task level can be a single computing task level or a computing task group level.
- the above resource type determines the computing resource allocation for task ID A; as another example, if the task ID of AI model training for AI beam management is B, then the above resource type determines the computing resource allocation for task ID B; as yet another example, if the task ID of the large model service is A, and the task ID of AI image recognition is C, and tasks A and C form a task group, then the above resource type can determine the computing resource allocation at the task group level.
- the aforementioned minimum number of operands can be used to represent the minimum number of operands required for a computational task.
- the aforementioned minimum computation speed can be used to indicate the lower limit of the computation speed for a computational task.
- computation speed can refer to the computational time complexity divided by the tolerable upper limit of computation time.
- computational time complexity is usually represented by operations, and typically corresponds to a data type.
- the minimum computing speed includes at least one of the following: theoretical minimum computing speed, and actual minimum computing speed.
- the performance parameters may further include a first test case indication, which corresponds to the actual minimum computing speed and is used to indicate that the actual minimum computing speed is the minimum computing speed obtained based on the test cases indicated by the first test case indication.
- the actual minimum computing speed can be expressed by the theoretical minimum computing speed and computing efficiency; or it can be expressed by the ideal minimum computing speed and computing efficiency.
- the aforementioned ideal minimum computing speed can be understood as the minimum computing speed obtained under ideal conditions such as no task preemption, based on test cases.
- the aforementioned minimum computational intensity can be used to indicate the lower limit of the computational intensity of a computational task.
- computational intensity can refer to computational speed divided by bandwidth.
- bandwidth consists of multiple components, including the transmission bandwidth between the second node (e.g., UE) and the computational node, and the memory bandwidth of the computational node.
- One implementation is to define computational intensity in segments; for example, computational intensity is computational speed divided by memory bandwidth.
- Another implementation is to use the smallest of the aforementioned multiple bandwidth components as the bandwidth for calculating computational intensity, i.e., computational intensity is computational speed divided by min ⁇ memory bandwidth, transmission bandwidth from the second node to the computational node ⁇ , where min ⁇ memory bandwidth, transmission bandwidth from the second node to the computational node ⁇ represents the smaller of the transmission bandwidth and the memory bandwidth.
- the transmission bandwidth between the second node and the computing node can be further divided into the air interface bandwidth between the second node and the access network node, and the wired transmission bandwidth between the access network node and the computing node; or the transmission bandwidth between the second node and the computing node can be further divided into the bandwidth between the second node and the UPF, and the wired transmission bandwidth between the UPF and the computing node.
- the aforementioned Computing Delay Budget indicates at least one of the following: the upper limit of tolerable computation delay, the upper limit of tolerable transmission delay, and the upper limits of tolerable computation and transmission delay.
- the aforementioned delay upper limit can be understood as the maximum delay.
- the computation delay budget includes at least one of the following: an upper limit for computation delay, an upper limit for transmission delay, and an upper limit for both computation and transmission delay.
- the above upper limit of computation latency can be understood as the upper limit of the latency that the computation task can tolerate when it is performed between the second node (such as the UE) and the computation node.
- the upper limit of the above transmission delay can be understood as the upper limit of the sum of the transmission delay from the second node to the computing node and the transmission delay from the computing node to the computing receiving node.
- the aforementioned upper limit for computation and transmission latency can be understood as the upper limit of the sum of the latency of the computation task between the second node (such as the UE) and the computation node, the transmission latency from the second node to the computation node, and the transmission latency from the computation node to the computation receiving node.
- the aforementioned computation receiving node can be understood as the node that receives the computation response data.
- the latency represented by the above computational latency budget can be defined in two ways: one is the length of the time interval between the sending of the first data packet of a single computational task and the receiving of the last data packet of that task; the other is the length of the time interval between the sending of the first data packet of a group of computational tasks and the receiving of the last data packet.
- image recognition computational tasks one approach is to treat single image recognition as a single computational task, while another approach is to treat multiple images (e.g., 100 images) as a group of computational tasks.
- the latency represented by the computational latency budget for AI model inference can include at least one of the following:
- Total end-to-end inference latency Specifically, it refers to the total end-to-end latency of multiple consecutive inference operations.
- the calculation method is as follows: the time before sending the first byte of the first computation task (or computation job) is denoted as T ⁇ sub> IS ⁇ /sub> , and the time before the receiving node receives the last byte of all computation tasks (or computation jobs) is denoted as T ⁇ sub>IE ⁇ /sub>. Then, the computational latency budget for AI model inference is T ⁇ sub> IE ⁇ /sub> - T ⁇ sub> IS ⁇ /sub> .
- Total inference latency Specifically, it refers to the total latency of multiple consecutive inference operations.
- the calculation method is as follows: the time before inference for the first computation task (or computation job) is denoted as T ⁇ sub>ITS ⁇ /sub> , and the time when all computation tasks (or computation jobs) finish inference is denoted as T ⁇ sub>ITE ⁇ /sub>. Then, the computational latency budget for AI model inference is T ⁇ sub>ITE ⁇ /sub> - T ⁇ sub> ITS ⁇ /sub> .
- End-to-end inference latency Specifically, it refers to the difference between the time it takes to send a sample and the time it takes to receive a result.
- t ⁇ sub>TIS ⁇ /sub> be the time before the second node sends the first byte of a computation task (or job)
- t ⁇ sub> IE ⁇ /sub> be the time before the receiving node receives the last byte of that computation task (or job).
- the estimated computational latency for AI model inference is t ⁇ sub>IE ⁇ /sub> - t ⁇ sub> TIS ⁇ /sub> .
- Inference latency Specifically, it refers to the difference between the start time and the end time of inference for a certain sample. That is, let t ⁇ sub> INS ⁇ /sub> be the time before inference for a certain computational task (or computational job), and let t ⁇ sub>INE ⁇ /sub> be the time when inference for the computational task ends. Then, the computational latency budget for AI model inference is t ⁇ sub>INE ⁇ /sub> - t ⁇ sub> INS ⁇ /sub> .
- the failure rate mentioned above represents the ratio of the number of failed computation requests to the total number of requests within a unit of time.
- the unit of time can be 1 second, 5 minutes, 1 day, etc.
- the failed requests can include incomplete computation requests and computation requests that completed after exceeding a latency threshold.
- the total number of requests can be the total number of requests received by the first node, or the total number of valid requests, where valid requests refer to computation requests accepted and processed by the first node.
- the average window mentioned above is used to represent the average statistical time window of at least one of the computation speed, computation intensity, and computation delay budget.
- Maximum time is used to represent the upper limit of the duration of a computation task, that is, the maximum time the computation task can last.
- the maximum time can be represented by at least one of the start time, duration, end time, etc.
- the aforementioned computing power types may include at least one of the following: a central processing unit (CPU), a graphics processing unit (GPU), a field-programmable gate array (FPGA), a data processing unit (DPU), a smart network interface card (smartNIC), a tensor processing unit (TPU), and a neural network processing unit (NPU).
- CPU central processing unit
- GPU graphics processing unit
- FPGA field-programmable gate array
- DPU data processing unit
- smartNIC smart network interface card
- TPU tensor processing unit
- NPU neural network processing unit
- the above-mentioned computing power types may also include at least one of the following:
- Clock speed for example, the lowest clock speed
- Number of cores for example, minimum number of cores.
- the data types mentioned above may include at least one of integers (e.g., int8, int4, etc.) and floating-point numbers (e.g., half-precision floating-point numbers, single-precision floating-point numbers, double-precision floating-point numbers, double-half-precision floating-point numbers, etc.).
- integers e.g., int8, int4, etc.
- floating-point numbers e.g., half-precision floating-point numbers, single-precision floating-point numbers, double-precision floating-point numbers, double-half-precision floating-point numbers, etc.
- the minimum memory mentioned above represents the lower limit of memory required for a computational task.
- the minimum memory can be represented by at least one of memory size (e.g., 8GB), sustained memory bandwidth, and memory random access rate.
- the minimum storage mentioned above represents the lower limit of storage required for a computing task.
- the minimum storage can be represented by at least one of storage size (e.g., 1T) and storage bandwidth (e.g., the maximum input/output (IO) flow per unit time).
- storage size e.g., 1T
- storage bandwidth e.g., the maximum input/output (IO) flow per unit time
- the aforementioned minimum transmission bandwidth is used to represent the lower limit of the bandwidth of the computing task, for example, at least one of the lower limit of the uplink bandwidth and the lower limit of the downlink bandwidth that the computing task can tolerate.
- the minimum transmission bandwidth includes at least one of the following: minimum uplink bandwidth, minimum downlink bandwidth, uplink bandwidth indication, downlink bandwidth indication, and target indication;
- the target indication is used to indicate whether the uplink bandwidth and downlink bandwidth are the same or different.
- the minimum transmission bandwidth is represented by at least one of the following: minimum uplink bandwidth, minimum downlink bandwidth, uplink bandwidth indication, downlink bandwidth indication, and target indication.
- the uplink bandwidth indicator described above is used to indicate uplink bandwidth; for example, 0 represents uplink bandwidth.
- the downlink bandwidth indicator described above is used to indicate downlink bandwidth; for example, 1 represents downlink bandwidth.
- the above target indication is used to indicate whether the uplink bandwidth and downlink bandwidth are the same or different. For example, 0 indicates that the uplink bandwidth and downlink bandwidth are different, and 1 indicates that the uplink bandwidth and downlink bandwidth are the same.
- the above-mentioned computational task arrival pattern is used to represent the arrival pattern of the computational job of the computational task.
- the computing task includes at least two computing jobs, and the computing task arrival mode includes at least one of the following:
- Continuous arrival mode or single arrival mode is used to indicate that one computation job arrives at a time
- Fixed-cycle arrival mode is used to indicate the arrival of control calculation jobs according to a fixed cycle
- Poisson distribution arrival pattern used to indicate the arrival of computational jobs controlled by Poisson distribution
- Peak arrival mode is used to indicate the arrival of ⁇ computational jobs within a target period of Poisson distribution, where the duration of the target period is less than a preset duration, and ⁇ is a positive integer.
- Offline arrival mode used to indicate that all computation jobs arrive at once.
- the aforementioned continuous arrival mode or single arrival mode is used to indicate that a computation job arrives at a time. For example, the i-th computation job arrives after the (i-1)-th computation job is completed. If the (i-1)-th computation job is not completed or the computation delay budget threshold is not reached, the i-th computation job is not sent. i is a positive integer.
- the fixed-period arrival pattern described above is used to indicate the arrival of computational jobs controlled according to a fixed period. For example, n computational jobs arrive at intervals of time T, where T represents the fixed period and n is a positive integer.
- the aforementioned Poisson distribution arrival pattern is used to indicate the arrival of computational jobs controlled by the Poisson distribution, for example, according to The arrival of computational jobs is controlled, where k represents the number of computational jobs arriving within a target unit of time, and k is a positive integer; ⁇ represents the average number of computational jobs arriving within the target unit of time, and ⁇ is a positive integer.
- the target unit of time could be 1 second or 2 seconds, etc.
- the aforementioned peak arrival mode is used to indicate the arrival of ⁇ computational jobs within a target period of a Poisson distribution, wherein the duration of the target period is less than a preset duration, which, for example, can be 10 seconds or 5 seconds, etc. For example, ⁇ > 25 computational jobs/second.
- the Poisson distribution arrival pattern has j short periods, where j is a positive integer. Each period contains a sudden surge of computational jobs, and the period lasts for a certain duration TG , for example, 5 to 10 seconds, while maintaining a certain level of concurrency ⁇ . The arrival of computational jobs within the short periods conforms to the fixed-period arrival pattern.
- the same computing task arrival mode or different computing task arrival modes can be used for computing jobs of different computing tasks.
- one computing task arrival mode or multiple computing task arrival modes can be used.
- the computing task includes 20 computing jobs, of which 10 computing jobs use a single arrival mode and the other 10 computing jobs use an offline arrival mode.
- the parameters corresponding to the above-mentioned computational task arrival modes may include the period and the number of computational jobs arriving in each period; for the Poisson distribution arrival mode, the parameters may include the average number of computational jobs arriving per target unit time, the target unit time, etc.; for the peak arrival mode, the parameters may include the target period and ⁇ , etc.
- the training accuracy of the aforementioned AI model can be represented by the data type of the AI model parameters output, such as single-precision floating-point numbers or half-precision floating-point numbers.
- the aforementioned lower limit of AI model performance may include at least one of the following: the lower limit of AI model training performance, and the lower limit of AI model inference performance.
- performance parameters typically differ across different scenarios.
- performance parameters for image recognition and object detection include top-1 accuracy and average accuracy
- performance parameters for semantic segmentation include Mean Intersection Over Union (MIOU)
- performance parameters for speech recognition include Word Error Rate (WER).
- MIOU Mean Intersection Over Union
- WER Word Error Rate
- the performance lower limit can be represented by the scores from benchmark tests of large models.
- the performance parameters of the computation task may further include a dataset indicator, which may correspond to the performance threshold of the AI model.
- the performance of the AI model corresponding to the dataset indicated by the dataset indicator must meet the lower performance limit of the AI model.
- Commonly used datasets include ImageNet2012, Pascal VOC2012, LibriSpeech ASR Corpus, Criteo, etc.
- the minimum throughput of the AI model inference mentioned above is usually the number of images inferred per second (images/s) for vision models, and the number of sentences inferred per second (sentences/s) or tokens/s for natural model models, where tokens refer to words, punctuation marks or other text units in the input text processed by the model.
- computational power consumption threshold can, for example, be used to represent the maximum computational power consumption.
- computational power consumption can refer to the power consumption of the computing node within the aforementioned computational latency budget; or, computational power consumption can refer to P2-P1 or the ratio of P2 to P1, where P1 represents the power consumption of the computing node per unit time when powered on, and P2 represents the power consumption per unit time when processing computational tasks.
- Computational energy efficiency thresholds can, for example, represent minimum computational energy efficiency.
- computational energy efficiency could be the computation speed per unit time divided by computational power consumption, or the throughput of an AI model inference divided by computational power consumption.
- the performance parameters of the above-mentioned computational tasks may only include a portion of the performance parameters listed above. Examples are provided below:
- Example 1 The performance parameters of the above computational task include minimum computation speed.
- the first node can use the minimum computation speed as a selection parameter for the computation node.
- Example 2 The performance parameters of the above computation task include the minimum number of operations or operands and the computation delay budget. If the computation delay budget only includes computation latency, then the first node can use these two parameters as selection parameters for the computation node; if the computation delay budget includes end-to-end latency, then the first node will use the data transmission latency and computation latency allocation information, along with these two parameters, as selection parameters for the computation node.
- Example 3 The performance parameters of the above computing tasks include computing power type (e.g., CPU, clock speed, number of cores), memory, storage, and maximum time.
- computing power type e.g., CPU, clock speed, number of cores
- memory e.g., RAM, storage, and maximum time.
- Example 4 The performance parameters of the above computing task include minimum computing speed and minimum transmission bandwidth.
- Example 5 The performance parameters of the above computational task include minimum computational intensity and computational delay budget.
- the performance parameters of the above-mentioned computational tasks may apply only to that specific computational task, meaning only that task needs to meet the above performance parameters; or they may apply to every computational task, meaning each task needs to meet the above performance parameters; or all computational tasks as a whole may meet the above performance parameters.
- the performance parameters of the above-mentioned computational task group may apply to each computational task within that group, meaning each task needs to meet the above performance parameters; or, all computational tasks in the group as a whole may meet the above performance parameters.
- one of the computational tasks is mapped to a Quality of Service (QoS) stream;
- QoS Quality of Service
- one of the computational tasks can be mapped to a set of QoS flows
- one of the computational tasks may be mapped to a Protocol Data Unit (PDU) session;
- PDU Protocol Data Unit
- one of the computing tasks can be mapped to a set of PDU sessions
- one of the computing tasks can be mapped to a radio bearer (RB).
- RB radio bearer
- one of the computational tasks can be mapped to a set of RBs
- one of the computational tasks can be mapped to a logical channel (LC).
- LC logical channel
- one of the computational tasks can be mapped to an LC set
- one of the computing tasks can be mapped to a physical layer resource
- a computing task may be mapped to a set of physical layer resources.
- DCI Downlink Control Information
- a computing task is mapped to a set of physical layer resources. For example, a computing task is mapped to physical layer resources scheduled through multiple DCIs.
- the performance parameters of the compute task correspond to the QoS flow.
- QoS Quality of Service
- the performance parameters of the compute task correspond to the QoS flow set.
- the performance parameters of the compute task correspond to the PDU session.
- the performance parameters of the compute task correspond to the PDU session set.
- the performance parameters of the compute task correspond to the RB.
- the performance parameters of the compute task correspond to the RB set.
- the performance parameters of the compute task correspond to the LC.
- the performance parameters of the compute task correspond to the LC set.
- the performance parameters of the compute task correspond to the physical layer resource.
- the performance parameters of the computing task correspond to the set of physical layer resources.
- the computing service request message may further include a computing service identifier for identifying the computing service.
- the aforementioned computing service identifier can be used to identify computing tasks such as one-dimensional DFT, two-dimensional FFT, AI model training, and AI model inference.
- the aforementioned AI model training or AI model inference can include image recognition, object detection, semantic segmentation, recommendation, natural language processing, speech recognition, optical character recognition, face recognition, beam management, localization, perception, CSI feedback, etc.
- the first node can select a computing node based on the computing service identifier and the performance parameters of the computing service. For example, the first node can select a computing node that supports the computing service identified by the computing service identifier and meets the performance parameters.
- the second node indicates the computing service it requests by carrying a computing service identifier in the computing service request message, which makes it easier for the second node to flexibly request different computing services.
- the method further includes:
- the first node selects a computing node based on the computing service request message and the status information of at least one computing node.
- the status information of the computing node includes at least one of the following: computing power type, computing load, available computing speed, available computing intensity, available memory, available storage, computing power consumption, and computing energy efficiency.
- the first node selects a computing node based on the computing service request message and the status information of at least one computing node, which helps to more accurately select a computing node that meets the performance parameters of the computing service from the at least one computing node.
- the computing service response message includes at least one of the following: computing task identifier, computing node identifier, and PDU session information.
- the aforementioned computing task identifier is used to identify computing tasks. It should be noted that there can be a one-to-one mapping between computing services and computing tasks, or multiple computing services can be mapped to one computing task, or one computing service can be mapped to multiple computing tasks.
- the aforementioned compute node identifier can be used to identify the selected compute node.
- the aforementioned PDU session information can be used to determine the PDU session for data transmission related to the aforementioned computing services.
- the aforementioned computing service response message when used to indicate acceptance of the aforementioned computing service request message, the aforementioned computing service response message may include at least one of the following: computing task identifier, computing node identifier, and PDU session information.
- the PDU session information includes one of the following:
- PDU session identifier PDU session identifier
- QoS flow identifier PDU session identifier
- QoS rules QoS rules
- the first node When a new PDU session needs to be established for data transmission related to the aforementioned computing services, the first node sends a PDU session establishment instruction to the second node. Upon receiving the PDU session establishment instruction, the second node can establish a PDU session, which can then be used for data transmission related to the aforementioned computing services.
- the first node When it is necessary to modify the PDU session for data transmission related to the aforementioned computing services, the first node sends a modification PDU session instruction and a PDU session identifier to the second node. Upon receiving the modification PDU session instruction and the PDU session identifier, the second node can modify the PDU session identified by the aforementioned PDU session identifier. The modified PDU session can then be used for data transmission related to the aforementioned computing services.
- the first node sends a PDU session identifier, a QoS flow identifier, and QoS rules to the second node.
- the second node can perform data transmission related to the aforementioned computing services based on the PDU session identifier, QoS flow identifier, and QoS rules.
- the computing service request message is also used to indicate whether the requested computing node is a wireless access network node.
- the aforementioned computing service request message may further include information indicating whether the requested computing node is a wireless access network node. Specifically, the indication of whether the requested computing node is a wireless access network node can be made explicitly or implicitly.
- the compute node selected by the first node must be a RAN node, or the first node should preferentially select a RAN node as the compute node. Selecting a RAN node as the compute node helps reduce compute service latency and better meets the needs of low-latency scenarios.
- RAN radio access network
- the computing service request message includes at least one of the following:
- the first indication information is used to indicate whether the requested computing node is a wireless access network node
- Latency type indicator used to indicate the latency type of the computing service.
- the first indication information can explicitly indicate whether the requested computing node is a wireless access network node. For example, a single bit can be used to indicate whether the requested computing node is a wireless access network node, where 0 indicates that the requested computing node is not a wireless access network node, and 1 indicates that the requested computing node is a wireless access network node.
- the latency type can be used to indicate whether the requested computing node is a radio access network (RAN) node.
- RAN radio access network
- the latency type indicated by the latency type indicator is a preset latency type, it means that the requested computing node is a RAN node; otherwise, it means that the requested computing node is not a RAN node.
- the preset latency type can be a low latency type.
- the selected computing node is a wireless access network node.
- the selected computing node is a wireless access network node.
- the selected computing node is a wireless access network node.
- the aforementioned preset type of computing service may include computing services for mobile network optimization, such as beam management, positioning, sensing, CSI feedback, etc.
- the aforementioned latency-sensitive type may include, but is not limited to, at least one of latency-sensitive guaranteed computing speed, latency-sensitive guaranteed computing strength, etc.
- the computing service response message includes at least one of the following: second indication information, computing task identifier, and computing node identifier; the second indication information is used to indicate that the selected computing node is a wireless access network node.
- the aforementioned computing service response message may include second indication information to indicate that the selected computing node is a radio access network node; otherwise, the aforementioned computing service response message does not include the second indication information.
- the computing service response message may further include a radio bearer indication, or the computing service response message may further include at least one of the following: a physical layer channel indication and a physical layer resource indication.
- the aforementioned radio bearer indication is used to indicate a radio bearer (RB).
- the radio bearer indication may include a radio bearer identifier (ID).
- the aforementioned radio bearer may include a signal radio bearer (SRB), a data radio bearer (DRB), or a newly defined RB, etc., and the radio bearer can be used for the transmission of data related to computing services.
- the first node can trigger the communication function in the radio access network node to send a radio bearer add or modify message.
- the aforementioned physical layer channel indicator can be used to indicate a physical layer channel that can be used for data transmission related to computing services.
- the aforementioned physical layer resource indicator is used to indicate physical layer resources that can be used for data transmission related to computing services.
- computing service-related data such as at least one of computing data and computing response data
- the radio bearer indication is used to indicate at least one of the following: signaling radio bearer (SRB), data radio bearer (DRB), and data plane (RB).
- SRB signaling radio bearer
- DRB data radio bearer
- RB data plane
- the method further includes:
- the first node sends first information to the radio access network node, the first information including third indication information, the third indication information being used to instruct the radio access network node to add or modify a radio bearer.
- the first node when the selected computing node is a radio access network node, the first node sends first information to the radio access network node, and then the radio access network node can send a Radio Resource Control (RRC) reconfiguration message to the second node based on the first information to request the establishment or modification of radio bearers or the configuration of physical layer resources for the computing service.
- RRC Radio Resource Control
- the first information may also include performance parameters of the computing service.
- the wireless access network node can learn about the performance parameters of the computing service, and then the wireless access network node can provide computing services based on the performance parameters of the computing service, which helps to further ensure the quality requirements of the computing service.
- the method further includes:
- the first node sends a first request message to the selected computing node, the first request message being used to request the creation or modification of a computing task;
- the first request message includes at least one of the following:
- Computation task identifier used to indicate the computation speed that a computing node guarantees to provide to a computing task within an average window
- guaranteed computation intensity used to indicate the computation intensity that a computing node guarantees to provide to a computing task within an average window
- maximum computation speed used to indicate the upper limit of the maximum computation speed that a computing node can provide to a computing task
- maximum computation intensity used to indicate the upper limit of the maximum computation intensity that a computing node can provide to a computing task.
- the priority indicator is used to represent the importance of the computing resource request of the computing task.
- the aforementioned preemption capability indicator is used to indicate whether a computing task can obtain computing resources that have been allocated to another computing task with lower priority.
- the preemption capability indicator is used to indicate whether to drop a computational task in order to execute a computational task with a higher priority.
- Migration capability indicators are used to indicate whether the migration of computing tasks from one computing node to another is supported.
- Guaranteed computation speed is used to indicate the computation speed that a computing node guarantees to provide to the computing task within the average window.
- Guaranteed computational intensity is used to indicate the computational intensity that a computing node guarantees to provide to the computing task within an average window.
- the aforementioned guarantee of computational speed may include a latency-sensitive guarantee of computational speed.
- the aforementioned guarantee of computational strength may include a latency-sensitive guarantee of computational strength.
- the method further includes:
- the first node receives a first response message from the selected computing node
- the first response message includes status information of the selected computing node after it has established or modified the computing task.
- the status information of the computing node after establishing or modifying the computing task may include, but is not limited to, at least one of the following: computing load, computing speed, computing intensity, available memory, available storage, computing power consumption, and computing energy efficiency.
- the first node when the first node receives the first response message, it can determine whether the selected computing node can meet the performance parameters of the computing service based on the status information after the computing task is established or modified by the selected computing node. If it does, the first node can send a computing service response message to the second node, indicating that it accepts the computing service request message. If it does not meet the requirements, the first node can reselect a computing node, which helps to further ensure the service quality of the computing service.
- Figure 3 is a flowchart of a computing service method provided in an embodiment of this application. This method can be executed by a second node, and as shown in Figure 3, it includes the following steps:
- Step 301 The second node sends a computing service request message to the first node, the computing service request message including the performance parameters of the computing service;
- Step 302 The second node receives a computing service response message from the first node, the computing service response message being used to indicate whether to accept or reject the computing service request message.
- the performance parameters of the computing service include performance parameters of at least one computing task or at least one group of computing tasks, wherein the performance parameters of the computing task or group of computing tasks include at least one of the following:
- Resource type minimum number of operations or operands; minimum computing speed; minimum computing intensity; computing latency budget; maximum failure rate, which represents the ratio of failed requests to total requests per unit time; average window; maximum time, indicating the upper limit of the computing task's duration; computing power type; data type; minimum memory, indicating the lower limit of memory required for the computing task; minimum storage, indicating the lower limit of storage required for the computing task; minimum transmission bandwidth, indicating the lower limit of bandwidth required for the computing task; computing task arrival mode; parameters corresponding to the computing task arrival mode; AI model training accuracy; AI model performance lower limit; minimum throughput of AI model inference; computing power consumption threshold; computing energy efficiency threshold.
- the computing service request message may also include a computing service identifier.
- the computing service response message includes at least one of the following: computing task identifier, computing node identifier, and PDU session information.
- the PDU session information includes one of the following:
- PDU session identifier PDU session identifier
- QoS flow identifier PDU session identifier
- QoS rules QoS rules
- the second node can send a PDU establishment request message to the third node and receive a PDU establishment response message returned by the third node; when the PDU session information includes a PDU session modification indication and a PDU session identifier, the second node can send a PDU modification request message to the third node and receive a PDU modification response message returned by the third node.
- the third node may include at least one of an AMF and an SMF.
- the computing service request message is also used to indicate whether the requested computing node is a wireless access network node.
- the computing service request message includes at least one of the following:
- the first indication information is used to indicate whether the requested computing node is a wireless access network node
- Latency type indicator used to indicate the latency type of the computing service.
- the computing service response message includes at least one of the following: second indication information, computing task identifier, and computing node identifier; the second indication information is used to indicate that the selected computing node is a wireless access network node.
- the computing service response message may further include a radio bearer indication, or the computing service response message may further include at least one of the following: a physical layer channel indication and a physical layer resource indication.
- the method further includes:
- the second node receives a Radio Resource Control (RRC) reconfiguration message sent by a Radio Access Network (RAN) node.
- RRC Radio Resource Control
- the RRC reconfiguration message includes a target configuration for computing services.
- the target configuration includes at least one of the following: radio bearer addition configuration, radio bearer modification configuration, and physical layer resource configuration.
- the second node sends an RRC reconfiguration complete message to the radio access network node.
- the aforementioned wireless access network node may include a base station or a centralized unit (CU), etc.
- CU centralized unit
- the second node when it receives the target configuration for computing services, it can configure computing services based on the target configuration. For example, if the target configuration includes a wireless bearer addition configuration, a wireless bearer can be added based on the wireless bearer addition configuration, and the added wireless bearer can be used for data transmission related to computing services. If the target configuration includes a wireless bearer modification configuration, the wireless bearer can be modified based on the wireless bearer modification configuration, and the modified wireless bearer can be used for data transmission related to computing services. If the target configuration includes a physical layer resource configuration, data related to computing services can be transmitted based on the physical layer resources configured in the physical layer resource configuration.
- transmitting computing service-related data via wireless bearer or physical layer can help reduce the transmission latency of computing services.
- the method further includes:
- the second node sends a second request message to the third node.
- the second request message is used to request the establishment or modification of a PDU session.
- the second request message includes fourth indication information, which is used to indicate that the PDU session supports at least one QoS flow for computing services.
- the second node receives a second response message from the third node, the second response message indicating acceptance of the second request message.
- the aforementioned third node may include at least one of AMF and SMF.
- the aforementioned at least one QoS flow used for computing services can be understood as satisfying the QoS parameters for computing services, or satisfying the QoS parameters for both computing and communication services.
- the third node may perform at least one of the following:
- a fourth node such as UPF is selected based on the second request message, and a third request message is sent to the selected fourth node.
- the third request message is used to request the establishment or modification of an N4 session.
- the third request message includes third information for computing services or for computing services and communication services.
- the third information includes at least one of the following: rule ID, priority, packet detection information, forwarding rule, enforcement rule, and reporting rule.
- SM N2 Session Management
- a radio access network node e.g., a base station
- the N2 Session Management message including at least one of a target QoS I and a first Quality of Service Profile (QoS Profile), the first QoS Profile being used for computing services, or the first QoS Profile being used for both computing services and communication services;
- QoS Profile Quality of Service Profile
- the third node determines the computing nodes, for example, by selecting the computing nodes itself or by obtaining computing nodes from the first node.
- the number of computing nodes determined may be more than one, depending on the number of packet filters and packet filter information provided by the terminal's computing service.
- the rule identifier is used to uniquely identify the rule for the aforementioned computing service.
- the priority is used to determine the order in which detection information for all rules applying to the computing services is processed.
- the aforementioned packet detection information includes at least one of the following: terminal Internet Protocol (IP) address, core network (CN) tunnel information, packet filter set, and target QoSI.
- IP Internet Protocol
- CN core network
- target QoSI target QoSI
- the packet filter mentioned above may include IP addresses, port numbers, etc., and is used to detect which data packets belong to the computing service flow.
- the target QoSI mentioned above may include a first QoSI or a second QoSI.
- the first QoSI represents at least one QoS parameter for the computing service
- the second QoSI represents at least one QoS parameter for both the computing and communication services.
- the first QoSI may represent at least one of the following: a first resource type, a first priority level, a first computing latency budget, a failure rate, an average window, a maximum number of operations or operands, a maximum computing speed, and a computing power type.
- the second QoSI may represent at least one of the following: a second resource type, a second priority level, a second computing latency budget, and bandwidth. It should be noted that the first QoSI mentioned above can also be called the computing QoSI, and the second QoSI mentioned above can also be called the computing and communication QoSI.
- the aforementioned forwarding rules include a forwarding rule identifier and a definition of forwarding the corresponding data packet to the aforementioned selected computing node (such as the computing node's IP address).
- the above execution rules include an execution rule identifier and a definition of the QoS operation to be executed.
- the corresponding 5QI is a delay-critical GBR type.
- the aforementioned reporting rules include a reporting rule identifier and a definition of the measurement operations to be performed. For example, measuring the latency, throughput, and data volume (throughput multiplied by time) from the UPF to the compute node.
- the embodiments of this application do not limit the order in which the second node sends the second request message to the third node and the second node sends the computing service request message to the first node.
- the second node may first send the second request message to the third node and then send the computing service request message to the first node; or, the terminal may first send the computing service request message to the first node and then send the second request message to the third node.
- the third node can also receive N4 session establishment or modification response messages from the fourth node.
- the PDU session since the PDU session supports at least one QoS stream for computing services, transmitting computing service-related data based on the PDU session can further guarantee the quality of service for computing services.
- the second response message includes a first QoS rule, which is used for computing services or for both computing services and communication services.
- the third node can send the first QoS rule through the N1 SM container.
- the first QoS rule includes at least one of the following:
- the target quality identifier (QoSI) for the QoS flow of the computing service includes a first QoSI or a second QoSI, wherein the first QoSI is used to represent at least one QoS parameter for the computing service, and the second QoSI is used to represent at least one QoS parameter for both the computing service and the communication service.
- a priority indicator which is used to indicate the priority of the first QoS rule.
- the packet filter described above may include IP addresses, port numbers, etc., to detect which packets belong to the computation service flow.
- the priority indicator described above can be used to determine the order in which packets are matched against multiple QoS rules.
- the first QoSI is used to represent at least one of the following: first resource type, first priority level, first computation latency budget, failure rate, average window, maximum number of operations or operands, maximum computation speed, and computing power type.
- the aforementioned first resource type can also be referred to as a computing resource type.
- the aforementioned first resource type may include at least one of the following: guaranteed computing speed, non-guaranteed computing speed, latency-sensitive guaranteed computing speed, guaranteed computing intensity, non-guaranteed computing intensity, and latency-sensitive guaranteed computing intensity.
- the first resource type determines the allocation of computing resources related to the QoS flow-level guaranteed computational load.
- the first resource type determines the allocation of computing resources related to the QoS flow-level guaranteed computational intensity.
- one computational task may be mapped to one QoS flow, or one computational task may be mapped to multiple QoS flows, or multiple computational tasks may be mapped to one QoS flow.
- the computational intensity in this embodiment can be a segmented computational intensity, for example, the computational speed divided by the memory bandwidth.
- the first priority level mentioned above is used to determine the priority of computing resource scheduling for computing QoS streams.
- the aforementioned first computational latency budget can be used to indicate the upper limit of the latency (i.e., the maximum computational latency) that can be tolerated when the computational task of the QoS flow (also referred to as the computational packet or computational packet set) is computed at the computing node.
- one definition of computation latency is the length of the time interval between the first data packet of a single computation task being sent and the last data packet of the computation task being received; another definition is the length of the time interval between the first data packet of a group of computation tasks being sent and the last data packet being received.
- one approach is to treat single image recognition as a single computation task, while another approach is to treat multiple images (e.g., 100 images) as a group of computation tasks.
- images e.g. 100 images
- computational latency for AI model inference can include at least one of the following:
- Total inference latency Specifically, it refers to the total latency of multiple consecutive inference operations.
- the calculation method is as follows: the time before inference for the first computation task (or computation job) is denoted as T ⁇ sub>ITS ⁇ /sub> , and the time when all computation tasks (or computation jobs) finish inference is denoted as T ⁇ sub>ITE ⁇ /sub>. Then, the computational latency budget for AI model inference is T ⁇ sub>ITE ⁇ /sub> - T ⁇ sub> ITS ⁇ /sub> .
- Inference latency Specifically, it refers to the difference between the start time and the end time of inference for a certain sample. That is, let t ⁇ sub> INS ⁇ /sub> be the time before inference for a certain computational task (or computational job), and let t ⁇ sub>INE ⁇ /sub> be the time when inference for the computational task ends. Then, the computational latency budget for AI model inference is t ⁇ sub>INE ⁇ /sub> - t ⁇ sub> INS ⁇ /sub> .
- the average window within the first QoSI mentioned above applies only to GCR or GCI or delay-critical GCR or delay-critical GCI.
- the maximum number of operands or operands mentioned above is used to represent the maximum number of operands or operands for a QoS flow.
- the maximum number of operands or operands within the first QoSI mentioned above applies only to GCRs or delay-critical GCRs, such as 2 TFLOPs.
- the aforementioned maximum computation speed is used to indicate the maximum computation speed of a QoS flow.
- the maximum computation speed within the first QoSI applies only to GCI or delay-critical GCI.
- the aforementioned maximum computation speed may include at least one of the theoretical maximum computation speed and the actual maximum computation speed.
- the theoretical maximum computing speed included in the aforementioned maximum computing speed is used to indicate that the theoretical maximum computing speed of the computing node required by QoS flows is not higher than the theoretical maximum computing speed included in the aforementioned maximum computing speed.
- OPS Operations Per Second
- the aforementioned maximum computing speed includes the actual maximum computing speed, indicating that the actual maximum computing speed of the computing node required by the QoS flow is not higher than the actual maximum computing speed included in the aforementioned maximum computing speed.
- the actual maximum computing speed is the maximum computing speed obtained through testing.
- the first QoSI described above can also be used to characterize a second test case indication, which corresponds to the actual maximum computing speed and is used to indicate that the actual maximum computing speed is the maximum computing speed obtained based on the test case indicated by the second test case indication.
- test cases can be at least one of the following: MobileNetVx (x can be any available version number, such as 1), one-dimensional DFT, one-dimensional FFT, two-dimensional FFT, matrix multiplication, sparse linear equations, dense linear equations, YOLOvy (y can be any available version number, such as YOLOv5), image recognition models (such as resnet50_v1.5), and large models (such as Llama3, Llama2).
- Test cases can be open-source software programs or custom software programs between the UE and the network. This means that the actual maximum computation speed required for QoS flows is the actual maximum computation speed that the computing node can achieve under the indicated test case conditions.
- the aforementioned actual maximum computing speed can be expressed as the theoretical maximum computing speed and computing efficiency, or as the ideal maximum computing speed and computing efficiency.
- the ideal maximum computing speed can be understood as the maximum computing speed obtained under ideal conditions such as no task preemption, based on test cases.
- computational efficiency is the ratio of the actual maximum computational speed to the theoretical maximum computational speed under ideal conditions.
- the actual maximum computational speed required for the computational task is represented by the theoretical maximum computational speed and the computational efficiency.
- Another definition of the aforementioned computational efficiency is the ratio of the maximum computational speed measured based on test cases to the theoretical maximum computational speed.
- the first QoSI is also used to characterize the second test case indication, and the actual maximum computational speed required for the QoS flow can be represented by the theoretical maximum computational speed and computational efficiency.
- Another definition of the aforementioned computational efficiency is: the ratio of the computational speed measured based on test cases to the maximum computational speed under ideal conditions.
- the computational speed measured based on test cases is usually obtained from recent tests and can represent the current state of the computing node.
- the maximum computational speed under ideal conditions, measured based on test cases refers to the best measured performance of the computing node.
- the aforementioned first QoSI is also used to characterize the second test case indication, and the actual maximum computational speed required for the QoS flow can be represented by the ideal maximum computational speed and computational efficiency.
- the first QoSI described above can be used to mark the index value of the forwarding processing parameters of the QoS flow of the computing service.
- the quality parameters represented by the first QoSI in this embodiment correspond to QoS flow, that is, the above-mentioned first QoSI is a QoS flow-level quality of service parameter.
- the second QoSI is used to represent at least one of the following: a second resource type, a second priority level, a second computational delay budget, and bandwidth.
- the aforementioned second resource type can also be referred to as a computing and communication resource type.
- the aforementioned second resource type may include at least one of the following: guaranteed computing speed, non-guaranteed computing speed, latency-sensitive guaranteed computing speed, guaranteed computing strength, non-guaranteed computing strength, and latency-sensitive guaranteed computing strength.
- the second resource type described above may determine the allocation of computational and communication resources related to the QoS flow-level guaranteed computational load, or the second resource type may determine the allocation of computational and communication resources related to the QoS flow-level guaranteed computational intensity.
- one computational task may be mapped to one QoS flow, or one computational task may be mapped to multiple QoS flows, or multiple computational tasks may be mapped to one QoS flow.
- the smallest of several bandwidth components such as the transmission bandwidth between the second node (e.g., UE) and the computing node, and the memory bandwidth of the computing node, can be used as the bandwidth for computational intensity. That is, the computational intensity is the computational speed divided by min ⁇ memory bandwidth, transmission bandwidth between the second node and the computing node ⁇ , where min ⁇ memory bandwidth, transmission bandwidth between the second node and the computing node ⁇ represents the smaller of the transmission bandwidth and the memory bandwidth.
- the transmission bandwidth between the second node and the computing node can be further divided into the air interface bandwidth between the second node and the access network node, and the wired transmission bandwidth between the access network node and the computing node; or the transmission bandwidth between the second node and the computing node can be further divided into the bandwidth between the second node and the UPF, and the wired transmission bandwidth between the UPF and the computing node.
- the aforementioned second priority level is used to determine the priority of computation and communication resource scheduling for QoS flows used for computation services.
- the aforementioned second computational delay budget represents the upper limit of the tolerable latency for the computation task of the QoS flow (also referred to as computation data packets or computation data packet sets) during computation and transmission.
- the tolerable latency during computation and transmission is the sum of the computation latency and the transmission latency.
- the transmission latency can refer to the sum of the transmission latency from the second node to the computation node and the transmission latency from the computation node to the computation receiving node.
- one definition of computation and transmission latency is the length of the time interval between the first data packet of a single computation task being sent and the last data packet of the computation task being received; another definition is the length of the time interval between the first data packet of a group of computation tasks being sent and the last data packet being received.
- one approach is to treat single image recognition as a single computation task, while another approach is to treat multiple images (e.g., 100 images) as a group of computation tasks. See the foregoing embodiments for details regarding computation latency budgeting.
- computation and transmission latency for AI model inference can include at least one of the following:
- Total end-to-end inference latency Specifically, it refers to the total end-to-end latency of multiple consecutive inference operations.
- the calculation method is as follows: the time before sending the first byte of the first computation task (or computation job) is denoted as T ⁇ sub> IS ⁇ /sub> , and the time before the receiving node receives the last byte of all computation tasks (or computation jobs) is denoted as T ⁇ sub>IE ⁇ /sub>. Then, the computational latency budget for AI model inference is T ⁇ sub> IE ⁇ /sub> - T ⁇ sub> IS ⁇ /sub> .
- End-to-end inference latency Specifically, it refers to the difference between the time it takes to send a sample and the time it takes to receive a result.
- t ⁇ sub>TIS ⁇ /sub> be the time before the second node sends the first byte of a computation task (or job)
- t ⁇ sub> IE ⁇ /sub> be the time before the receiving node receives the last byte of that computation task (or job).
- the estimated computational latency for AI model inference is t ⁇ sub>IE ⁇ /sub> - t ⁇ sub> TIS ⁇ /sub> .
- the aforementioned bandwidth is used to represent at least one of the uplink bandwidth lower limit and downlink bandwidth lower limit of the QoS flow.
- the bandwidth in this embodiment please refer to the relevant descriptions in the foregoing embodiments; they will not be repeated here.
- the second QoSI described above can be used to mark the index value of the forwarding processing parameters of the QoS flow of computing and communication services.
- the quality parameters represented by the second QoSI in this embodiment correspond to a QoS flow, that is, the above-mentioned second QoSI is a QoS flow-level quality of service parameter.
- the second request message further includes at least one of the following:
- the number of packet filters used to calculate the number of packet filters for a service is the number of packet filters used to calculate the number of packet filters for a service.
- the second response message includes a computing node identifier.
- the aforementioned computing node identifier may include an IP address or an internal network ID, etc.
- the method further includes:
- the second node sends a second message to the computing node, the second message including computing data.
- the second information further includes at least one of the following:
- a receiving node indication is used to indicate the node that receives the computation response corresponding to the computation data.
- Example 1 The main idea of this example is to solve the problems of performance parameter identification and interaction when mobile networks provide computing services, and how to ensure the quality of service of computing services, based on computing service request messages containing computing service parameters. Furthermore, in this embodiment, the computing task is mapped to a PDU session and QoS flow; the first node can be a core network node, and the computing node can be a core network node or an edge computing node, etc.
- the computing service method provided in this application embodiment includes the following steps:
- Step 11 The second node sends a computing service request message to the first node.
- the aforementioned computing service request message may include performance parameters of the computing service.
- performance parameters of the computing service please refer to the relevant descriptions in the foregoing embodiments, which will not be repeated here.
- Step 12 The first node sends a computing task creation/modification request message to the computing node.
- the first node can select a suitable computing node based on the computing service request message and the status information of at least one computing node.
- the status information of the computing nodes can be found in the relevant descriptions of the foregoing embodiments, and will not be repeated here.
- the first node can send a task creation/modification request message to the selected computing node.
- the description of the aforementioned task creation/modification request message can be found in the relevant descriptions of the foregoing embodiments, and will not be repeated here.
- Step 13 The compute node sends a compute task creation/modification response message to the first node.
- steps 12 and 13 above can be optional steps.
- existing computing tasks can be used for calculation and processing.
- Step 14 The first node sends a computing service response message to the second node.
- the aforementioned computing service response message can be used to indicate whether or not to accept the aforementioned computing service request message.
- the relevant description in the foregoing embodiments which will not be repeated here.
- Figure 4 shows the case where the above computing service response message indicates acceptance of the above computing service request message.
- Step 15 The second node sends a PDU session establishment/modification request message to the third node.
- Step 16 The third node sends a PDU session establishment/modification response message to the second node.
- Steps 15 and 16 above can be found in the process of establishing or modifying a PDU session in related computing scenarios, and will not be elaborated here. It should be noted that steps 15 and 16 above can be optional steps. For example, existing PDU sessions can be used to transmit data related to computing services.
- Step 17 The second node sends computation data to the computing node.
- the second node can transmit the aforementioned computational data based on a PDU session.
- Step 18a The computing node sends computing response data to the second node.
- Step 18b The computing node sends computing response data to the computing receiving node.
- the receiving node and the second node are different nodes.
- Example 2 The main idea of this example is to map computing tasks to radio bearer or physical layer resources, rather than PDU sessions or QoS flows.
- the first node can be a core network node or a radio access network node, and the computing node is a radio access network node.
- the computing service method provided in this example can solve the problems of performance parameter identification and interaction when mobile networks provide computing services, as well as how to ensure the quality of service of computing services, especially for low-latency scenarios or scenarios where the radio access network node is a trusted node.
- the computing service method provided in this application embodiment includes the following steps:
- Step 21 The second node sends a computing service request message to the first node.
- the aforementioned computing service request message may include performance parameters of the computing service. Specific details regarding these performance parameters can be found in the descriptions of the foregoing embodiments and will not be repeated here.
- the aforementioned computing service request message may also indicate whether the requested computing node is a radio access network node.
- Step 22 The first node sends a computing task creation/modification request message to the wireless access network node.
- the first node can select a suitable computing node based on the computing service request message and the status information of at least one computing node, wherein the computing node is a radio access network node.
- the status information of the computing node can be found in the relevant descriptions of the foregoing embodiments, and will not be repeated here.
- the first node may send a task creation/modification request message to the selected radio access network node.
- the aforementioned task creation/modification request message can be found in the relevant descriptions of the foregoing embodiments, and will not be repeated here.
- Step 23 The wireless access network node sends a computing task creation/modification response message to the first node.
- steps 22 and 23 above can be optional steps.
- existing computing tasks can be used for calculation and processing.
- Step 24 The first node sends a computing service response message to the second node.
- the aforementioned computing service response message can be used to indicate whether or not to accept the aforementioned computing service request message.
- the relevant description in the foregoing embodiments which will not be repeated here.
- the computing service response message may further include at least one of the following: a computing node is a radio access network node indication (i.e., second indication information), a radio bearer indication, a physical layer channel indication, and a physical layer resource indication.
- a computing node is a radio access network node indication (i.e., second indication information), a radio bearer indication, a physical layer channel indication, and a physical layer resource indication.
- Figure 5 shows the case where the above computing service response message indicates acceptance of the above computing service request message.
- Step 25 The wireless access network node sends an RRC reconfiguration message to the second node.
- the RRC reconfiguration message may include a radio bearer add/modify request.
- Step 26 The second node sends an RRC reconfiguration complete message to the radio access network node.
- Step 27 The second node sends computation data to the computing node.
- the second node can transmit the aforementioned computational data based on a PDU session.
- Step 28a The computing node sends computing response data to the second node.
- Step 28b The computing node sends computing response data to the computing receiving node.
- the receiving node and the second node are different nodes.
- Example 3 The main idea of this example is to define the service quality parameters related to computing services as the first QoSI, and the service quality parameters related to computing and communication as the second QoSI.
- the flexible parameters and the precise values for each computing task require interaction through the computing service process of each computing task. It should be noted that the meanings of the first QoSI and the second QoSI can be found in the relevant descriptions of the foregoing embodiments, and will not be repeated here.
- the computing service method provided in this application embodiment includes the following steps:
- Step 31 The UE sends a PDU session establishment or modification request message to the third node.
- the PDU session establishment or modification request message contains fourth indication information, which indicates that the PDU session contains at least one QoS flow for computing services, that is, the PDU session supports at least one QoS flow for computing services.
- Step 32 The third node sends an N4 session establishment or modification request message to the fourth node.
- the third node corresponds to the AMF and SMF.
- the AMF receives the PDU session establishment or modification request message from the UE and selects the appropriate SMF based on information such as whether QoS flows for computing services are needed.
- the SMF selects the appropriate fourth node (e.g., UPF) based on the PDU session establishment or modification request message.
- the SMF can also select the appropriate computing node or obtain the appropriate computing node from the computing management node.
- the N4 session establishment or modification request message includes packet detection, execution, and reporting rules for the computing service, etc. For details, please refer to the relevant descriptions in the foregoing embodiments, which will not be repeated here.
- Step 33 The fourth node sends an N4 session establishment or modification response message to the third node.
- the UPF sends an N4 session establishment message or an N4 session modification response message to the SMF.
- Step 34 The third node sends N2 session management information to the radio access network node.
- the N2 session management information may include at least one of the following: target QoSI and calculated QoS profile.
- the target QoSI can be used to mark the index value of the forwarding processing parameters of the QoS flow of the computing service.
- the third node maps the QoS flow corresponding to the computing service to one data radio bearer, and maps the QoS flows of other non-computing services to another data radio bearer.
- Step 35 The third node sends the first QoS rule and the number of compute node identifiers to the UE through the N1 session management container.
- Step 36 The UE sends a computing service request message to the first node.
- This step is the same as step 11 above, and will not be repeated here.
- Step 37 The first node sends a computing task creation/modification request message to the computing node.
- This step is the same as step 12 above, and will not be repeated here.
- Step 38 The compute node sends a compute task creation/modification response message to the first node.
- This step is the same as step 13 above, and will not be repeated here.
- Step 39 The first node sends a computing service response message to the UE.
- This step is the same as step 14 above, and will not be repeated here.
- Step 40 The UE sends computation data to the computing node.
- This step is the same as step 17 above, and will not be repeated here.
- Step 41a The computing node sends computing response data to the UE.
- This step is the same as step 18a above, and will not be repeated here.
- Step 41b The computing node sends computing response data to the computing receiving node.
- This step is the same as step 18b above, and will not be repeated here.
- steps 31 to 35 can be executed first, followed by steps 36 to 39; or steps 36 to 39 can be executed first, followed by steps 31 to 35.
- the computing service method provided in this application transmits performance parameters (e.g., performance minimum requirements) corresponding to the computing service based on the computing service request message.
- performance parameters e.g., performance minimum requirements
- This method is applicable to providing computing services to both AFs (Automatic Front-End) and UEs (User Equipment) and NFs (Network Functions).
- core network nodes acting as both computing management nodes and computing nodes It has broader applicability, suitable for core network nodes acting as both computing management nodes and computing nodes, core network nodes acting as both computing management nodes and radio access network nodes acting as computing nodes, and radio access network nodes acting as both computing management nodes and computing nodes.
- the computing service method provided in this application embodiment can be executed by a computing service device.
- This application embodiment uses the execution of the computing service method by a computing service device as an example to illustrate the computing service device provided in this application embodiment.
- the computing service device may be a communication device or a component within a communication device, such as a chip.
- the communication device may be a terminal, a network-side device, or a server, etc.
- the terminal may include, but is not limited to, the type of terminal 11 listed above
- the network-side device may include, but is not limited to, the type of network-side device 12 listed above. This application does not impose specific limitations.
- the computing service device includes a receiving module, a transmitting module, and a processing module. These modules can be implemented in software or hardware.
- the processing module can be implemented by a processor.
- the processor can include general-purpose processors, special-purpose processors, such as a Central Processing Unit (CPU), microprocessor, Digital Signal Processor (DSP), Artificial Intelligence (AI) processor, Graphics Processing Unit (GPU), Application Specific Integrated Circuit (ASIC), Network Processor (NP), Field Programmable Gate Array (FPGA), or other programmable logic devices, gate circuits, transistors, discrete hardware components, etc.
- the receiving and transmitting modules can be implemented by a communication interface, which can include one or more of the following: transceiver, pins, circuits, bus, radio frequency unit, etc.
- the computing service device 700 when the computing service device is a network-side device or a component of a network-side device, the computing service device 700 includes a receiving module 701, used to receive a computing service request message from a second node, the computing service request message including performance parameters of the computing service; and a sending module 702, used to send a computing service response message to the second node, wherein the computing service response message is used to indicate whether to accept or reject the computing service request message.
- a receiving module 701 used to receive a computing service request message from a second node, the computing service request message including performance parameters of the computing service
- a sending module 702 used to send a computing service response message to the second node, wherein the computing service response message is used to indicate whether to accept or reject the computing service request message.
- the performance parameters of the computing service include performance parameters of at least one computing task or at least one group of computing tasks, wherein the performance parameters of the computing task or group of computing tasks include at least one of the following:
- Resource type minimum number of operations or operands; minimum computing speed; minimum computing intensity; computing latency budget; maximum failure rate, which represents the ratio of failed requests to total requests per unit time; average window; maximum time, representing the upper limit of the computing task's duration; computing power type; data type; minimum memory, representing the lower limit of memory required for the computing task; minimum storage, representing the lower limit of storage required for the computing task; minimum transmission bandwidth, representing the lower limit of bandwidth required for the computing task; computing task arrival mode; parameters corresponding to the computing task arrival mode; AI model training accuracy; AI model performance lower limit; minimum throughput of AI model inference; computing power consumption threshold; computing energy efficiency threshold.
- the resource type includes at least one of the following:
- the minimum computing speed includes at least one of the following: theoretical minimum computing speed, and actual minimum computing speed.
- the computation delay budget includes at least one of the following: an upper limit for computation delay, an upper limit for transmission delay, and an upper limit for both computation and transmission delay.
- the minimum transmission bandwidth includes at least one of the following: minimum uplink bandwidth, minimum downlink bandwidth, uplink bandwidth indication, downlink bandwidth indication, and target indication;
- the target indication is used to indicate whether the uplink bandwidth and downlink bandwidth are the same or different.
- the computing task includes at least two computing jobs, and the computing task arrival mode includes at least one of the following:
- Continuous arrival mode or single arrival mode is used to indicate that one computation job arrives at a time
- Fixed-cycle arrival mode is used to indicate the arrival of control calculation jobs according to a fixed cycle
- Poisson distribution arrival pattern used to indicate the arrival of computational jobs controlled by Poisson distribution
- Peak arrival mode is used to indicate the arrival of ⁇ computational jobs within a target period of Poisson distribution, where the duration of the target period is less than a preset duration, and ⁇ is a positive integer.
- Offline arrival mode used to indicate that all computation jobs arrive at once.
- the lower limit of AI model performance includes at least one of the following: the lower limit of AI model training performance, and the lower limit of AI model inference performance.
- the performance parameters of the computing task may also include a dataset indicator, wherein the performance of the AI model corresponding to the dataset indicated by the dataset indicator must meet the lower limit of the AI model performance.
- one of the computational tasks is mapped to a Quality of Service (QoS) stream;
- QoS Quality of Service
- one of the computational tasks can be mapped to a set of QoS flows
- one of the computing tasks can be mapped to a Protocol Data Unit (PDU) session;
- PDU Protocol Data Unit
- one of the computing tasks can be mapped to a set of PDU sessions
- one of the computing tasks can be mapped to a radio bearer (RB).
- RB radio bearer
- one of the computational tasks can be mapped to a set of RBs
- one of the computational tasks may be mapped to a logical channel LC;
- one of the computational tasks can be mapped to an LC set
- one of the computing tasks can be mapped to a physical layer resource
- a computing task may be mapped to a set of physical layer resources.
- the computing service request message may further include a computing service identifier for identifying the computing service.
- the device further includes:
- the processing module is used to select a computing node based on the computing service request message and the status information of at least one computing node;
- the status information of the computing node includes at least one of the following: computing power type, computing load, available computing speed, available computing intensity, available memory, available storage, computing power consumption, and computing energy efficiency.
- the computing service response message includes at least one of the following: computing task identifier, computing node identifier, and PDU session information.
- the PDU session information includes one of the following:
- PDU session identifier PDU session identifier
- QoS flow identifier PDU session identifier
- QoS rules QoS rules
- the computing service request message is also used to indicate whether the requested computing node is a wireless access network node.
- the computing service request message includes at least one of the following:
- the first indication information is used to indicate whether the requested computing node is a wireless access network node
- Latency type indicator used to indicate the latency type of the computing service.
- the selected computing node is a wireless access network node.
- the selected computing node is a wireless access network node.
- the selected computing node is a wireless access network node.
- the computing service response message includes at least one of the following: second indication information, computing task identifier, and computing node identifier; the second indication information is used to indicate that the selected computing node is a wireless access network node.
- the computing service response message may further include a radio bearer indication, or the computing service response message may further include at least one of the following: a physical layer channel indication and a physical layer resource indication.
- the radio bearer indication is used to indicate at least one of the following: signaling radio bearer (SRB), data radio bearer (DRB), and data plane (RB).
- SRB signaling radio bearer
- DRB data radio bearer
- RB data plane
- the apparatus when the selected computing node is a wireless access network node, the apparatus further includes:
- the first information may also include performance parameters of the computing service.
- the sending module is further configured to send a first request message to the selected computing node, the first request message being used to request the establishment or modification of a computing task;
- the first request message includes at least one of the following:
- Computation task identifier used to indicate the computation speed that a computing node guarantees to provide to a computing task within an average window
- guaranteed computation intensity used to indicate the computation intensity that a computing node guarantees to provide to a computing task within an average window
- maximum computation speed used to indicate the upper limit of the maximum computation speed that a computing node can provide to a computing task
- maximum computation intensity used to indicate the upper limit of the maximum computation intensity that a computing node can provide to a computing task.
- the receiving module is further configured to receive a first response message from the selected computing node
- the first response message includes status information of the selected computing node after it has established or modified the computing task.
- the computing service device provided in this application embodiment can implement the various processes implemented in the method embodiment of FIG2 and achieve the same technical effect. To avoid repetition, it will not be described again here.
- the computing service device 800 when the computing service device is a terminal or a component within a terminal, or when the computing service device is a network-side device or a component within a network-side device, the computing service device 800 includes a sending module 801, configured to send a computing service request message to a first node, the computing service request message including performance parameters of the computing service; and a receiving module 802, configured to receive a computing service response message from the first node, the computing service response message indicating whether to accept or reject the computing service request message.
- a sending module 801 configured to send a computing service request message to a first node, the computing service request message including performance parameters of the computing service
- a receiving module 802 configured to receive a computing service response message from the first node, the computing service response message indicating whether to accept or reject the computing service request message.
- the performance parameters of the computing service include performance parameters of at least one computing task or at least one group of computing tasks, wherein the performance parameters of the computing task or group of computing tasks include at least one of the following:
- Resource type minimum number of operations or operands; minimum computing speed; minimum computing intensity; computing latency budget; maximum failure rate, which represents the ratio of failed requests to total requests per unit time; average window; maximum time, indicating the upper limit of the computing task's duration; computing power type; data type; minimum memory, indicating the lower limit of memory required for the computing task; minimum storage, indicating the lower limit of storage required for the computing task; minimum transmission bandwidth, indicating the lower limit of bandwidth required for the computing task; computing task arrival mode; parameters corresponding to the computing task arrival mode; AI model training accuracy; AI model performance lower limit; minimum throughput of AI model inference; computing power consumption threshold; computing energy efficiency threshold.
- the computing service request message may also include a computing service identifier.
- the computing service response message includes at least one of the following: computing task identifier, computing node identifier, and PDU session information.
- the PDU session information includes one of the following:
- PDU session identifier PDU session identifier
- QoS flow identifier PDU session identifier
- QoS rules QoS rules
- the computing service request message is also used to indicate whether the requested computing node is a wireless access network node.
- the computing service request message includes at least one of the following:
- the first indication information is used to indicate whether the requested computing node is a wireless access network node
- Latency type indicator used to indicate the latency type of the computing service.
- the computing service response message includes at least one of the following: second indication information, computing task identifier, and computing node identifier; the second indication information is used to indicate that the selected computing node is a wireless access network node.
- the computing service response message may further include a radio bearer indication, or the computing service response message may further include at least one of the following: a physical layer channel indication and a physical layer resource indication.
- the receiving module is further configured to receive a Radio Resource Control (RRC) reconfiguration message sent by a radio access network node.
- the RRC reconfiguration message includes a target configuration for computing services.
- the target configuration includes at least one of the following: radio bearer add configuration, radio bearer modify configuration, and physical layer resource configuration.
- the sending module is also used to send an RRC reconfiguration complete message to the radio access network node.
- the sending module is further configured to send a second request message to the third node, the second request message being used to request the establishment or modification of a PDU session, the second request message including fourth indication information, the fourth indication information being used to indicate that the PDU session supports at least one QoS flow for computing services;
- the receiving module is further configured to receive a second response message from the third node, the second response message being used to indicate acceptance of the second request message.
- the second response message includes a first QoS rule, which is used for computing services or for both computing services and communication services.
- the first QoS rule includes at least one of the following:
- the target quality identifier (QoSI) for the QoS flow of the computing service includes a first QoSI or a second QoSI, wherein the first QoSI is used to represent at least one QoS parameter for the computing service, and the second QoSI is used to represent at least one QoS parameter for both the computing service and the communication service.
- a priority indicator which is used to indicate the priority of the first QoS rule.
- the first QoSI is used to represent at least one of the following: first resource type, first priority level, first computation latency budget, failure rate, average window, maximum number of operations or operands, maximum computation speed, and computing power type.
- the second QoSI is used to represent at least one of the following: a second resource type, a second priority level, a second computational delay budget, and bandwidth.
- the second request message further includes at least one of the following:
- the number of packet filters used to calculate the number of packet filters for a service is the number of packet filters used to calculate the number of packet filters for a service.
- the second response message includes a computing node identifier.
- the sending module is further configured to send second information to the computing node, the second information including computing data.
- the second information further includes at least one of the following:
- a receiving node indication is used to indicate the node that receives the computation response corresponding to the computation data.
- the computing service device provided in this application embodiment can implement the various processes implemented in the method embodiment of FIG3 and achieve the same technical effect. To avoid repetition, it will not be described again here.
- this application embodiment also provides a communication device 900, including a processor 901 and a memory 902.
- the memory 902 stores programs or instructions that can run on the processor 901.
- the program or instructions executed by the processor 901 implement the various steps of the above-described first node-side computing service method embodiment and achieve the same technical effect.
- the program or instructions executed by the processor 901 implement the various steps of the above-described second node-side computing service method embodiment and achieve the same technical effect. To avoid repetition, this will not be described again here.
- This application also provides a network-side device, including a processor and a communication interface.
- the communication interface is coupled to the processor, and the processor is used to run programs or instructions to implement the steps of the method embodiment shown in FIG2 or 3.
- This network-side device embodiment corresponds to the above-described first node or second node-side method embodiment. All implementation processes and methods of the above-described method embodiments can be applied to this network-side device embodiment and can achieve the same technical effect.
- this application embodiment also provides a network-side device, which may be the computing service device shown in FIG7.
- the network-side device 1000 includes: an antenna 1001, a radio frequency device 1002, a baseband device 1003, a processor 1004, and a memory 1005.
- the antenna 1001 is connected to the radio frequency device 1002.
- the radio frequency device 1002 receives information through the antenna 1001 and sends the received information to the baseband device 1003 for processing.
- the baseband device 1003 processes the information to be transmitted and sends it to the radio frequency device 1002, which processes the received information and then transmits it through the antenna 1001.
- the method executed by the network-side device in the above embodiments can be implemented in the baseband device 1003, which includes a baseband processor.
- the baseband device 1003 may include at least one baseband board, on which multiple chips are disposed, as shown in FIG10.
- One of the chips is, for example, a baseband processor, which is connected to the memory 1005 via a bus interface to call the program in the memory 1005 and execute the network device operation shown in the above method embodiment.
- the network-side device may also include a network interface 1006, such as a Common Public Radio Interface (CPRI).
- CPRI Common Public Radio Interface
- the network-side device 1000 in this application embodiment further includes: instructions or programs stored in memory 1005 and executable on processor 1004.
- Processor 1004 calls the instructions or programs in memory 1005 to execute the methods executed by each module shown in FIG7 and achieve the same technical effect. To avoid repetition, it will not be described in detail here.
- the network-side device 1100 includes: a processor 1101, a network interface 1102, and a memory 1103.
- the network-side device may be the computing service device shown in FIG7 or FIG8.
- the network interface 1102 is, for example, a common public radio interface (CPRI).
- CPRI common public radio interface
- the network-side device 1100 in this application embodiment further includes: instructions or programs stored in memory 1103 and executable on processor 1101.
- Processor 1101 calls the instructions or programs in memory 1103 to execute the methods executed by the modules shown in FIG7 or FIG8 and achieve the same technical effect. To avoid repetition, it will not be described in detail here.
- This application also provides a terminal, including a processor and a communication interface, wherein the communication interface is coupled to the processor, and the processor is used to run programs or instructions to implement the steps in the method embodiment shown in FIG3.
- This terminal embodiment corresponds to the above-described terminal-side method embodiment, and all implementation processes and methods of the above-described method embodiments can be applied to this terminal embodiment and can achieve the same technical effect.
- the terminal may be the computing service device shown in FIG8.
- FIG12 is a schematic diagram of the hardware structure of a terminal implementing an embodiment of this application.
- the terminal 1200 includes, but is not limited to, at least some of the following components: radio frequency unit 1201, network module 1202, audio output unit 1203, input unit 1204, sensor 1205, display unit 1206, user input unit 1207, interface unit 1208, memory 1209, and processor 1210.
- the terminal 1200 may also include a power supply (such as a battery) for powering various components.
- the power supply can be logically connected to the processor 1210 through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system.
- the terminal structure shown in Figure 12 does not constitute a limitation on the terminal.
- the terminal may include more or fewer components than shown, or combine certain components, or have different component arrangements, which will not be elaborated here.
- the input unit 1204 may include a graphics processor 12041 and a microphone 12042.
- the graphics processor 12041 processes image data of still images or videos obtained by an image capture device (such as a camera) in video capture mode or image capture mode.
- the display unit 1206 may include a display panel 12061, which may be configured in the form of a liquid crystal display, an organic light-emitting diode, or the like.
- the user input unit 1207 includes a touch panel 12071 and at least one of other input devices 12072.
- the touch panel 12071 is also called a touch screen.
- the touch panel 12071 may include a touch detection device and a touch controller.
- Other input devices 12072 may include, but are not limited to, physical keyboards, function keys (such as volume control buttons, power buttons, etc.), trackballs, mice, and joysticks, which will not be described in detail here.
- the radio frequency unit 1201 can transmit it to the processor 1210 for processing; in addition, the radio frequency unit 1201 can send uplink data to the network-side device.
- the radio frequency unit 1201 includes, but is not limited to, antennas, amplifiers, transceivers, couplers, low-noise amplifiers, duplexers, etc.
- the memory 1209 can be used to store software programs or instructions, as well as various data.
- the memory 1209 may primarily include a first storage area for storing programs or instructions and a second storage area for storing data.
- the first storage area may store the operating system, application programs or instructions required for at least one function (such as sound playback, image playback, etc.).
- the memory 1209 may include volatile memory or non-volatile memory.
- the non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory.
- Volatile memory can be random access memory (RAM), static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct memory bus RAM (DRRAM).
- RAM random access memory
- SRAM static random access memory
- DRAM dynamic random access memory
- SDRAM synchronous dynamic random access memory
- DDRSDRAM double data rate synchronous dynamic random access memory
- ESDRAM enhanced synchronous dynamic random access memory
- SLDRAM synchronous link dynamic random access memory
- DRRAM direct memory bus RAM
- the memory 1209 in this embodiment includes, but is not limited to, these and any other suitable types of memory.
- Processor 1210 may include one or more processing units; optionally, processor 1210 integrates an application processor and a modem processor, wherein the application processor mainly handles operations involving the operating system, user interface, and applications, and the modem processor mainly handles wireless communication signals, such as a baseband processor. It is understood that the aforementioned modem processor may also not be integrated into processor 1210.
- the radio frequency unit 1201 is used to receive a computing service request message from the second node, the computing service request message including performance parameters of the computing service;
- the radio frequency unit 1201 is also configured to send a computing service response message to the second node, wherein the computing service response message is used to indicate whether to accept or reject the computing service request message.
- This application also provides a readable storage medium storing a program or instructions.
- the program or instructions When the program or instructions are executed by a processor, they implement the various processes of the above-described computing service method embodiments and achieve the same technical effects. To avoid repetition, they will not be described again here.
- the processor mentioned above is the processor in the terminal described in the above embodiments.
- the readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.
- ROM computer read-only memory
- RAM random access memory
- magnetic disk magnetic disk
- optical disk optical disk
- the readable storage medium may be a non-transient readable storage medium.
- This application embodiment also provides a chip, which includes a processor and a communication interface.
- the communication interface is coupled to the processor.
- the processor is used to run programs or instructions to implement the various processes of the above-described computing service method embodiments and can achieve the same technical effect. To avoid repetition, it will not be described again here.
- chip mentioned in the embodiments of this application may also be referred to as a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.
- This application also provides a computer program/program product, which is stored in a storage medium and executed by at least one processor to implement the various processes of the above-described computing service method embodiments, and can achieve the same technical effect. To avoid repetition, it will not be described again here.
- This application also provides a wireless communication system, including a first node and a second node, wherein the first node can be used to execute the steps of the computing service method described above, and the second node can be used to execute the steps of the computing service method described above.
Landscapes
- Engineering & Computer Science (AREA)
- Computer Networks & Wireless Communication (AREA)
- Signal Processing (AREA)
- Mobile Radio Communication Systems (AREA)
Abstract
本申请公开了一种计算服务方法、装置、第一节点及第二节点,属于通信技术领域,本申请实施例的计算服务方法包括:第一节点从第二节点接收计算服务请求消息,所述计算服务请求消息包括计算服务的性能参数;所述第一节点向所述第二节点发送计算服务响应消息,其中,所述计算服务响应消息用于指示接受或拒绝所述计算服务请求消息。
Description
相关申请的交叉引用
本申请主张在2024年7月8日在中国提交的中国专利申请No.202410909436.6的优先权,其全部内容通过引用包含于此。
本申请属于通信技术领域,具体涉及一种计算服务方法、装置、第一节点及第二节点。
随着通信技术的发展,一些移动通信系统支持提供计算服务,例如,离散傅里叶变换(Discrete Fourier Transform,DFT)、快速傅里叶变换(Fast Fourier Transformation,FFT)、人工智能(Artificial Intelligence,AI)模型训练、AI模型推理等计算服务。然而,相关技术中仍是基于移动通信系统的通信服务的质量参数(例如,第五代移动通信技术(5th Generation Mobile Communication Technology,5G)服务质量标识(5G QoS Identifier,5QI))来控制移动通信系统的计算服务的服务质量,这样容易导致计算服务的服务质量较差。
本申请实施例提供一种计算服务方法、装置、第一节点及第二节点,有利于保证计算服务的服务质量。
第一方面,提供了一种计算服务方法,该方法包括:
第一节点从第二节点接收计算服务请求消息,所述计算服务请求消息包括计算服务的性能参数;
所述第一节点向所述第二节点发送计算服务响应消息,其中,所述计算服务响应消息用于指示接受或拒绝所述计算服务请求消息。
第二方面,提供了一种计算服务装置,该装置包括:
接收模块,用于从第二节点接收计算服务请求消息,所述计算服务请求消息包括计算服务的性能参数;
发送模块,用于向所述第二节点发送计算服务响应消息,其中,所述计算服务响应消息用于指示接受或拒绝所述计算服务请求消息。
第三方面,提供了一种计算服务方法,该方法包括:
第二节点向第一节点发送计算服务请求消息,所述计算服务请求消息包括计算服务的性能参数;
所述第二节点从所述第一节点接收计算服务响应消息,所述计算服务响应消息用于指示接受或拒绝所述计算服务请求消息。
第四方面,提供了一种计算服务装置,该装置包括:
发送模块,用于向第一节点发送计算服务请求消息,所述计算服务请求消息包括计算服务的性能参数;
接收模块,用于从所述第一节点接收计算服务响应消息,所述计算服务响应消息用于指示接受或拒绝所述计算服务请求消息。
第五方面,提供了一种计算服务的装置,所述装置被配置为执行如第一方面所述的方法的步骤,或者实现如第三方面所述的方法的步骤。
第六方面,提供了一种第一节点,该第一节点包括处理器和存储器,所述存储器存储可在所述处理器上运行的程序或指令,所述程序或指令被所述处理器执行时实现如第一方面所述的方法的步骤。
第七方面,提供了一种第一节点,包括处理器及通信接口,其中,所述通信接口用于从第二节点接收计算服务请求消息,所述计算服务请求消息包括计算服务的性能参数;
所述通信接口还用于向所述第二节点发送计算服务响应消息,其中,所述计算服务响应消息用于指示接受或拒绝所述计算服务请求消息。
第八方面,提供了一种第二节点,该第二节点包括处理器和存储器,所述存储器存储可在所述处理器上运行的程序或指令,所述程序或指令被所述处理器执行时实现如第三方面所述的方法的步骤。
第九方面,提供了一种第二节点,包括处理器及通信接口,其中,所述通信接口用于向第一节点发送计算服务请求消息,所述计算服务请求消息包括计算服务的性能参数;
所述通信接口还用于从所述第一节点接收计算服务响应消息,所述计算服务响应消息用于指示接受或拒绝所述计算服务请求消息。
第十方面,提供了一种可读存储介质,所述可读存储介质上存储程序或指令,所述程序或指令被处理器执行时实现如第一方面所述的方法的步骤,或者实现如第三方面所述的方法的步骤。
第十一方面,提供了一种无线通信系统,包括:第一节点及第二节点,所述第一节点可用于执行如第一方面所述的计算服务方法的步骤,所述第二节点可用于执行如第三方面所述的计算服务方法的步骤。
第十二方面,提供了一种芯片,所述芯片包括处理器和通信接口,所述通信接口和所述处理器耦合,所述处理器用于运行程序或指令,实现如第一方面所述的方法的步骤,或实现如第三方面所述的方法的步骤。
第十三方面,提供了一种计算机程序/程序产品,所述计算机程序/程序产品被存储在存储介质中,所述计算机程序/程序产品被至少一个处理器执行以实现如第一方面所述的方法的步骤,或实现如第三方面所述的方法的步骤。
在本申请实施例中,第一节点从第二节点接收计算服务请求消息,所述计算服务请求消息包括计算服务的性能参数;所述第一节点向所述第二节点发送计算服务响应消息,其中,所述计算服务响应消息用于指示接受或拒绝所述计算服务请求消息,也即本申请实施例通过在计算服务请求消息中携带计算服务的性能参数,进而第一节点可以基于计算服务的性能参数确定是否接受计算服务请求消息,这样有利于保证计算服务的服务质量。
图1是本申请实施例可应用的一种无线通信系统的框图;
图2是本申请实施例提供的一种计算服务方法的流程图;
图3是本申请实施例提供的另一种计算服务方法的流程图;
图4是本申请实施例提供的又一种计算服务方法的流程图;
图5是本申请实施例提供的又一种计算服务方法的流程图;
图6是本申请实施例提供的又一种计算服务方法的流程图;
图7是本申请实施例提供的一种计算服务装置的结构图;
图8是本申请实施例提供的另一种计算服务装置的结构图;
图9是本申请实施例提供的通信设备的结构图;
图10是本申请实施例提供的一种网络侧设备的结构图;
图11是本申请实施例提供的另一种网络侧设备的结构图;
图12是本申请实施例提供的终端的结构图。
下面将结合本申请实施例中的附图,对本申请实施例中的技术方案进行清楚描述,显然,所描述的实施例是本申请一部分实施例,而不是全部的实施例。基于本申请中的实施例,本领域普通技术人员所获得的所有其他实施例,都属于本申请保护的范围。
本申请的术语“第一”、“第二”等是用于区别类似的对象,而不用于描述特定的顺序或先后次序。应该理解这样使用的术语在适当情况下可以互换,以便本申请的实施例能够以除了在这里图示或描述的那些以外的顺序实施,且“第一”、“第二”所区别的对象通常为一类,并不限定对象的个数,例如第一对象可以是一个,也可以是多个。此外,本申请中的“或”表示所连接对象的至少其中之一。例如“A或B”的保护范围至少涵盖三种方案,即,方案一:包括A且不包括B;方案二:包括B且不包括A;方案三:既包括A又包括B。此外,术语“A和/或B”、“A和B中的至少一项”、“A或B中的至少一项”也分别至少涵盖上述三种方案。字符“/”一般表示前后关联对象是一种“或”的关系。
本申请的术语“指示”既可以是一个直接的指示(或者说显式的指示),也可以是一个间接的指示(或者说隐含的指示)。其中,直接的指示可以理解为,发送方在发送的指示中明确告知了接收方具体的信息、需要执行的操作或请求结果等内容;间接的指示可以理解为,接收方根据发送方发送的指示确定对应的信息,或者进行判断并根据判断结果确定需要执行的操作或请求结果等。
值得指出的是,本申请实施例所描述的技术不限于长期演进型(Long Term Evolution,LTE)/LTE的演进(LTE-Advanced,LTE-A)系统,还可用于其他无线通信系统,诸如码分多址(Code Division Multiple Access,CDMA)、时分多址(Time Division Multiple Access,TDMA)、频分多址(Frequency Division Multiple Access,FDMA)、正交频分多址(Orthogonal Frequency Division Multiple Access,OFDMA)、单载波频分多址(Single-carrier Frequency-Division Multiple Access,SC-FDMA)或其他系统。本申请实施例中的术语“系统”和“网络”常被可互换地使用,所描述的技术既可用于以上提及的系统和无线电技术,也可用于其他系统和无线电技术。以下描述出于示例目的描述了新空口(New Radio,NR)系统,并且在以下大部分描述中使用NR术语,但是这些技术也可应用于NR系统以外的系统,如第6代(6th Generation,6G)通信系统。
图1示出本申请实施例可应用的一种无线通信系统的框图。无线通信系统包括终端11和网络侧设备12。其中,终端11可以是手机、平板电脑(Tablet Personal Computer)、膝上型电脑(Laptop Computer)、笔记本电脑、个人数字助理(Personal Digital Assistant,PDA)、掌上电脑、上网本、超级移动个人计算机(Ultra-mobile Personal Computer,UMPC)、移动上网装置(Mobile Internet Device,MID)、增强现实(Augmented Reality,AR)、虚拟现实(Virtual Reality,VR)设备、机器人、可穿戴式设备(Wearable Device)、飞行器(flight vehicle)、车载用户设备(Vehicle User Equipment,VUE)、船载设备、行人用户设备(Pedestrian User Equipment,PUE)、智能家居(具有无线通信功能的家居设备,如冰箱、电视、洗衣机或者家具等)、游戏机、个人计算机(Personal Computer,PC)、柜员机或者自助机等终端侧设备。可穿戴式设备包括:智能手表、智能手环、智能耳机、智能眼镜、智能首饰(智能手镯、智能手链、智能戒指、智能项链、智能脚镯、智能脚链等)、智能腕带、智能服装等。其中,车载设备也可以称为车载终端、车载控制器、车载模块、车载部件、车载芯片或车载单元等。需要说明的是,在本申请实施例并不限定终端11的具体类型。网络侧设备12可以包括接入网设备或核心网设备,其中,接入网设备也可以称为无线接入网(Radio Access Network,RAN)设备、无线接入网功能或无线接入网单元。接入网设备可以包括基站、无线局域网(Wireless Local Area Network,WLAN)接入点(Access Point,AP)或无线保真(Wireless Fidelity,WiFi)节点等。其中,基站可被称为节点B(Node B,NB)、演进节点B(Evolved Node B,eNB)、下一代节点B(the next generation Node B,gNB)、新空口节点B(New Radio Node B,NR Node B)、接入点、中继站(Relay Base Station,RBS)、服务基站(Serving Base Station,SBS)、基收发机站(Base Transceiver Station,BTS)、无线电基站、无线电收发机、基本服务集(Basic Service Set,BSS)、扩展服务集(Extended Service Set,ESS)、家用B节点(home Node B,HNB)、家用演进型B节点(home evolved Node B)、发送接收点(Transmit/Receive Point,TRP)或所属领域中其他某个合适的术语,只要达到相同的技术效果,所述基站不限于特定技术词汇,需要说明的是,在本申请实施例中仅以NR系统中的基站为例进行介绍,并不限定基站的具体类型。
核心网设备也可以称为核心网节点、核心网功能或核心网网元等,其包含但不限于如下至少一项:移动管理实体(Mobility Management Entity,MME)、接入移动管理功能(Access and Mobility Management Function,AMF)、会话管理功能(Session Management Function,SMF)、用户平面功能(User Plane Function,UPF)、策略控制功能(Policy Control Function,PCF)、策略与计费规则功能单元(Policy and Charging Rules Function,PCRF)、边缘应用服务发现功能(Edge Application Server Discovery Function,EASDF)、统一数据管理(Unified Data Management,UDM)、统一数据仓储(Unified Data Repository,UDR)、归属用户服务器(Home Subscriber Server,HSS)、集中式网络配置(Centralized network configuration,CNC)、网络存储功能(Network Repository Function,NRF)、网络开放功能(Network Exposure Function,NEF)、本地NEF(Local NEF,或L-NEF)、绑定支持功能(Binding Support Function,BSF)、应用功能(Application Function,AF)、位置管理功能(Location Management Function,LMF)、网关的移动位置中心(Gateway Mobile Location Centre,GMLC)、网络数据分析功能(Network Data Analytics Function,NWDAF)等。需要说明的是,在本申请实施例中仅以NR系统中的核心网设备为例进行介绍,并不限定核心网设备的具体类型,如果在后续协议版本(例如6G)中本申请实施例提到的核心网设备的名称发生变化,也在本申请的保护范围内。
可选的,核心网设备可以由一个设备中的一个或多个功能模块实现,也可以由多个设备共同实现,本申请实施例对此不作具体限定。可以理解的是,上述功能模块既可以是硬件设备中的网络元件,也可以是在专用硬件上运行的软件功能模块,或者是平台(例如,云平台)上实例化的虚拟化功能模块。
为了方便理解,以下对本申请实施例涉及的一些内容进行说明:
一、定位服务质量(Quality of Service,QoS)
当前定位中包括不同定位服务等级对应的性能需求参见协议TS22.261中的表7.3.2.2-1。
其中,定位精度(Accuracy)用于表示定位服务性能误差分布,采用置信水平(confidence level)及定位误差门限来定义,即定位结果与实际位置的距离在定位误差门限范围内的百分比。例如,定位精度为<3m的置信水平为95%,表示定位结果有95%的概率与实际位置的距离误差小于3米。
定位的QoS包含在定位请求中,包括:定位需求方的定位请求和定位结果提供方(例如,定位管理功能(Location Management Function,LMF))的定位请求。LMF的定位请求是基于定位需求方的定位请求生成的。
上述定位需求方(例如,位置服务(Location Services,LCS)客户端(client)或AF)的定位请求可以包括以下参数指标:
(1)、QoS类型/等级(LCS QoS Class),包括:
尽力而为(Best Effort)型:最宽松的定位QoS类型,如果定位结果不能满足其他的QoS指标要求,仍需要反馈定位结果,但需要指示说明所请求的QoS没有被满足。如果没有获得定位结果,则反馈失败原因。
多QoS(Multiple QoS)型:中等严格的定位QoS类型,即包含多个QoS等级对应的QoS指标要求,如果定位结果不满足最严格的QoS指标要求,则LMF再次发起定位流程,尝试满足更低要求的QoS指标要求,直到满足其中一个QoS指标要求为止。如果最宽松的QoS指标要求仍未满足,则不反馈定位结果,仅反馈失败原因。
保障(Assured)型。最严格的定位QoS类型。如果定位结果不能满足其他的QoS指标要求,则不反馈定位结果,仅反馈失败原因。
(2)、定位精度(Accuracy),包括水平定位精度和/或垂直定位精度;
(3)、响应时间(Response Time)类型,LMF需要平衡定位精度和响应时间类型。
无延迟:LMF应立即反馈目标UE的初始位置或最近的定位结果。如果没有定位结果,则反馈失败信息,并可以触发定位流程,用于响应后续的定位请求。
低延迟:相比于精度,优先满足响应时间要求。LCS服务器应以最小的延迟返回当前位置。
延迟不敏感:相比于响应时间,优先满足精度要求。LMF可延迟反馈定位结果,直到满足所需的定位精度要求。
上述LMF的定位请求可以包括以下参数指标:
水平定位精度(HorizontalAccuracy),包括定位精度(accuracy)及置信水平(confidence level);
垂直定位精度(VerticalAccuracy),包括定位精度(accuracy)及置信水平(confidence level);
响应时间(ResponseTime),为UE收到定位信息请求到提供定位信息的时延。
二、AI和通信应用场景
AI和通信是国际电信联盟无线通信部门(International Telecommunication Union Radiocommunication Sector,ITU-R)的6G应用场景之一,典型用例包括IMT-2030(6G)辅助自动驾驶、医疗辅助应用中设备之间的自主协作、跨设备/网络计算卸载、数字孪生的创建和预测,以及IMT-2030(6G)辅助协作机器人。这类应用场景将需要支持移动网络高容量和用户体验数据速率,以及低延迟和高可靠性等。除了通信方面,这一应用场景预计将包括一组将人工智能和计算相关的功能集成到6G系统相关的新功能,包括来自不同来源的数据采集、准备和处理、分布式人工智能模型训练、跨移动通信系统的模型共享和分布式推理,以及计算资源编排。
现有云服务主要是通过服务等级协议(Service Level Agreement,SLA)来定义云服务提供商和用户之间的静态的长期(例如,月)的服务质量,用户所以企业用户为主,较少直接面向消费者,特别是移动终端消费者。
下面结合附图,通过一些实施例及其应用场景对本申请实施例提供的计算服务方法进行详细地说明。
请参见图2,图2是本申请实施例提供的一种计算服务方法的流程图,该方法可以由第一节点执行,如图2所示,包括以下步骤:
步骤201、第一节点从第二节点接收计算服务请求消息,所述计算服务请求消息包括计算服务的性能参数。
本实施例中,上述第一节点也可以称为计算管理节点、计算管理功能、计算控制功能、计算管理控制功能、计算服务控制功能或计算服务管理功能等。此外,上述第一节点可以是无线接入网节点或者核心网节点。上述第一节点可以为具有负责接收计算服务请求,或处理计算服务请求、计算资源调度、计算信息交互、计算数据处理等至少一项功能的网络节点。上述第一节点可以是新增的网络节点;或者可以是对已有的网络节点增强后所得到的节点,例如,增强的AMF或增强的SMF或增强的基站等。
上述第二节点可以包括但不限于终端、AF和网络功能实体(Network Function,NF)等中的至少一项。在上述第二节点为AF的情况下,上述计算服务请求消息需要通过NEF等发送至第一节点。需要说明的是,上述第二节点也可以称为计算请求节点。
上述计算服务可以包括DFT、FFT、AI模型训练、AI模型推理等中的至少一项,上述AI模型训练或AI模型推理可以包括图像识别、物体检测、语义分割、推荐、自然语言处理、语音识别、光学字符识别、人脸识别、波束管理、定位、感知、信道状态信息(Channel State Information,CSI)反馈等。上述计算服务可以包括至少一个计算任务。
上述计算服务的性能参数可以用于表征该计算服务所需满足的性能要求。示例性地,上述计算服务的性能参数可以包括但不限于计算任务的硬件要求参数、计算要求参数、计算结果要求参数、传输要求参数等中的至少一项。上述硬件要求参数可以包括但不限于算力类型、最小内存、最小存储等中的至少一项。上述计算要求参数可以包括但不限于最小运算数或操作数、最小计算速度、最小计算强度、计算功耗门限和计算能效门限等中的至少一项。上述计算结果要求参数可以包括但不限于最大失败率、AI模型性能下限等中的至少一项。上述传输要求参数可以包括但不限于最小传输带宽、传输延迟预算等中的至少一项。
具体地,第一节点可以基于上述计算服务的性能参数确定是否接受上述计算服务请求消息。示例性地,第一节点可以根据所述计算服务请求消息选择计算节点,例如,第一节点可以选择能够满足上述计算服务的性能参数的计算节点,并可以根据计算节点的选择结果确定是否接受上述计算服务请求消息。可以理解的是,在选择计算节点时,除了上述计算服务的性能参数,还可以参考其他参数,例如,计算节点的状态、计算服务的类型等。
需要说明的是,所选择的计算节点的数量可以为至少一个,或者可以为0个,即未选到合适的计算节点。例如,在不存在可以满足计算服务的性能参数的计算节点的情况下,所选择的计算节点的数量为0。
步骤202、所述第一节点向所述第二节点发送计算服务响应消息,其中,所述计算服务响应消息用于指示接受或拒绝所述计算服务请求消息。
示例性地,在所选择的计算节点的数量可以为至少一个的情况下,上述计算服务响应消息用于指示接受所述计算服务请求消息;在所选择的计算节点的数量为0的情况下,上述计算服务响应消息用于指示拒绝所述计算服务请求消息。
在本申请实施例中,第一节点从第二节点接收计算服务请求消息,所述计算服务请求消息包括计算服务的性能参数;所述第一节点向所述第二节点发送计算服务响应消息,其中,所述计算服务响应消息用于指示接受或拒绝所述计算服务请求消息,也即本申请实施例通过在计算服务请求消息中携带计算服务的性能参数,进而第一节点可以基于计算服务的性能参数确定是否接受计算服务请求消息,这样有利于保证计算服务的服务质量。
可选地,所述计算服务的性能参数包括至少一个计算任务的性能参数或者至少一个计算任务组的性能参数,所述计算任务或者计算任务组的性能参数包括如下至少一项:
资源类型;最小运算数或操作数;最小计算速度;最小计算强度;计算延迟预算;最大失败率,所述失败率用于表示在单位时间内计算任务的失败请求数与总请求数的比值;平均窗口;最大时间,用于表示计算任务持续的时间上限;算力类型;数据类型;最小内存,用于表示计算任务所需的内存下限;最小存储,用于表示计算任务所需的存储下限;最小传输带宽,用于表示计算任务的带宽下限;计算任务到达模式;计算任务到达模式对应的参数;人工智能AI模型训练精度;AI模型性能下限;AI模型推理的最小吞吐率;计算功耗门限;计算能效门限。
上述资源类型可以用于反映QoS类型(QoS class),例如,保证型、非保证型、时延敏感型等。其中,上述资源类型也可以称为计算QoS类型(Computing QoS class),或者,计算和通信QoS类型等。
可选地,所述资源类型包括如下至少一项:
保证计算速度(Guaranteed Computing Rate,GCR),保证计算强度(Guaranteed Computing Intensity,GCI),非保证计算速度(Non-GCR),非保证计算强度(Non-GCI),时延敏感保证计算速度(delay critical GCR),时延敏感保证计算强度(delay critical GCI)。
需要说明的是,上述保证计算速度也可以称为保证操作速度(Guaranteed Operational Rate,GOR),上述保证计算强度也可以称为保证操作强度(Guaranteed Operational Intensity,GOI)。
示例性地,上述保证计算速度、非保证计算速度和时延敏感保证计算速度中的至少一项所表征的计算速度可以是指计算时间复杂度除以可容忍的计算时间上限,其中,上述计算时间复杂度例如可以通过操作数来表示。
示例性地,上述保证计算强度、非保证计算强度和时延敏感保证计算强度中的至少一项所表征的计算强度可以是指计算速度除以带宽。
上述保证计算速度可以理解为与计算任务的保证计算速度相关的计算资源在该计算任务期间被永久分配。示例性地,所选择的计算节点上与该计算任务的保证计算速度相关的CPU、GPU或存储等计算资源分配给该计算任务。上述分配给该计算任务的计算资源为该计算任务专用的计算资源,不与其它计算任务共享。
上述保证计算强度可以理解为与计算任务的保证计算强度相关的计算资源,或计算和通信资源,在该计算任务期间被永久分配。示例性地,在计算强度为计算速度除以内存带宽所得到的值的情况下,与保证计算强度相关的CPU、GPU和内存等计算资源分配给该计算任务;在计算强度为计算速度除以目标带宽(即内存带宽和传输带宽中的较小值)所得到的值的情况下,与保证计算强度相关的CPU、GPU、内存和网络资源分配给该计算任务。其中网络资源包括空口资源或有线传输资源等,空口资源与UE和网络之间计算任务的空口传输带宽相关,有线传输资源与UE和网络之间计算任务的接入网到计算节点的传输带宽相关。
上述延迟敏感保证计算速度可以理解为如果计算任务的一个数据包的延迟超过了计算延迟预算,则数据包丢弃。并且延迟敏感保证计算速度要求第一比例(如98%)的数据包不能超过计算延迟预算。
上述延迟敏感保证计算强度也可以理解为如果该计算任务的一个数据包的延迟超过了计算延迟预算,则数据包丢弃。并且延迟敏感保证计算速度要求第二比例(如98%)的数据包不能超过计算延迟预算。
上述非保证计算速度可以理解为与计算任务的计算速度相关的资源在该计算任务期间未被永久分配。示例性地,所选择的计算节点上与该计算任务的保证计算速度相关的CPU、GPU或存储等计算资源未被分配给该计算任务,与其它计算任务共享。
上述非保证计算强度可以理解为与计算任务的计算强度相关的资源在该计算任务期间未被永久分配。示例性地,所选择的计算节点上与该计算任务的保证计算速度相关的CPU、GPU和内存等计算资源,以及网络资源等计算和通信资源未被分配给该计算任务,与其它计算任务共享。
需要说明的是,上述资源类型可以决定与计算任务级(level)保证计算量相关的计算资源分配,或者资源类型可以决定与计算任务级保证计算强度相关的计算资源,或计算和通信资源相关的分配。其中,上述计算任务级可以是单个计算任务级或计算任务组级。例如,若大模型服务的任务标识(ID)为A,则上述资源类型决定任务ID A的计算资源分配;又例如,若AI波束管理的AI模型训练的任务ID为B,则上述资源类型决定任务ID B的计算资源分配;又例如,若大模型服务的任务ID为A,AI图像识别的任务ID为C,任务A和任务C组成任务组,则上述资源类型可以决定任务组级的计算资源分配。
上述最小运算数或操作数,可以用于表示计算任务的最小运算数或操作数。
上述最小计算速度可以用于指示计算任务的计算速度的下限。其中,计算速度可以是指计算时间复杂度除以可容忍的计算时间上限。示例性地,计算时间复杂度通常可以用运算数(operations)来表示,通常还对应于一种数据类型。例如,2浮点运算数(Floating-point Operations,TFLOPs)表示该计算任务的时间复杂度是2×1012浮点运算数;如果该计算任务可容忍的计算时间上限是20ms,即最长20ms计算完成,则计算速度就是2×1012/(20×10-3)=100次浮点运算每秒(Floating-point Operations Per Second,TFLOPS)。
可选地,所述最小计算速度包括如下至少一项:理论最小计算速度,实际最小计算速度。
在一些可选的实施例中,上述性能参数还可以包括第一测试用例指示,该第一测试用例指示与上述实际最小计算速度对应,用于表示该实际最小计算速度为基于该第一测试用例指示所指示的测试用例得到的最小计算速度。
在一些可选的实施例中,对于实际最小计算速度,可以通过理论最小计算速度、计算效率来表示;或者可以通过理想最小计算速度和计算效率来表示。其中,上述理想最小计算速度可以理解为在无任务抢占等理想情况下基于测试用例进行测试所得到的最小计算速度。
上述最小计算强度可以用于指示计算任务的计算强度的下限。其中,计算强度可以是指计算速度除以带宽值。需要说明的是,在移动通信系统(例如,6G)提供计算服务的过程中带宽由第二节点(如UE)到计算节点之间的传输带宽和计算节点的内存带宽等多个部分带宽组成。一种实施方式是分段定义计算强度,例如,计算强度为计算速度除以内存带宽。另一种实施方式是将上述多个部分带宽中的最小一个作为计算计算强度的带宽,即计算强度为计算速度除以min{内存带宽、第二节点到计算节点的传输带宽},其中,min{内存带宽、第二节点到计算节点的传输带宽}表示传输带宽和内存带宽中的较小者。
在一些可选的实施例中,上述第二节点到计算节点的传输带宽之间的传输带宽还可以分成第二节点到接入网节点之间的空口带宽,接入网节点到计算节点之间的有线传输带宽;或者第二节点和计算节点之间的传输带宽还可以分成第二节点到UPF之间的带宽,UPF到计算节点之间的有线传输带宽。
上述计算延迟预算(Computing Delay Budget,CDB)用于指示计算可容忍的时延上限、传输可容忍的时延上限以及计算和传输可容忍的时延上限中的至少一项。其中,上述时延上限可以理解为最大时延。
可选地,所述计算延迟预算包括如下至少一项:计算时延的上限,传输时延的上限,计算和传输时延的上限。
上述计算时延的上限可以理解为计算任务在第二节点(如UE)和计算节点之间计算时可容忍的时延上限。
上述传输时延的上限可以理解为第二节点到计算节点的传输时延和计算节点到计算接收节点的传输时延之和的上限。
上述计算和传输时延的上限可以理解为计算任务在第二节点(如UE)和计算节点之间计算时的时延、第二节点到计算节点的传输时延以及计算节点到计算接收节点的传输时延之和的上限。其中,上述计算接收节点可以理解为接收计算响应数据的节点。
示例性地,对于上述计算延迟预算所表征的时延,一种定义是单个计算任务的第一个数据包发送到接收到计算任务的最后一个数据包的时间区间长度。另一种定义是一组计算任务的第一个数据包发送到最后一个数据包接收的时间区间长度。例如,对于图像识别的计算任务,一种是以单张图像识别作为单个计算任务,另一种是以多张(例如100张)图像识别作为一组计算任务。
例如,对于面向AI模型推理的计算延迟预算所表征的时延,可以包括如下至少一种:
推理端到端总时延:具体是指多次连续推理端到端总延时。计算方法是:发送第一个计算任务(或计算作业)的第一字节前计时的时间记为TIS,计算接收节点接收到所有计算任务(或计算作业)的最后一个字节记为TIE,那么AI模型推理的计算延迟预算为TIE-TIS。
推理总时延:具体是指多次连续推理总延时。计算方法是:第一个计算任务(或计算作业)推理前时间记为TITS,所有计算任务(或计算作业)推理结束时的时间记为TITE,那么AI模型推理的计算延迟预算为TITE-TITS。
端到端推理时延:具体是指发送样本时间与收到结果时间的差。即第二节点发送某一个计算任务(或计算作业)的第一字节前计时的时间记为tTIS,计算接收节点接收到该计算任务(或计算作业)的最后一个字节记为tTIE,那么AI模型推理的计算延迟预算为tTIE-tTIS。
推理时延:具体是指某样本推理的开始时间与结束时间的差。即某一个计算任务(或计算作业)的推理前计时的时间记为tINS,该计算任务推理结束时的时间记为tINE,那么AI模型推理的计算延迟预算为tINE-tINS。
上述失败率用于表示在单位时间内计算任务的失败请求数与总请求数的比值,示例性地,上述单位时间可以为1秒钟、5分钟、1天等。上述失败请求可以包括未完成的计算请求和超过时延门限完成的计算请求。上述总请求数可以是第一节点接收的请求的总数,或者可以是有效请求的总数,其中,上述有效请求可以是指第一节点接受和处理的计算请求。
上述平均窗口用于表示计算速度、计算强度和计算延迟预算中至少一项的平均统计时间窗口。
最大时间,用于表示计算任务持续的时间上限,即计算任务持续的最长时间。示例性地,可以通过开始时间、时间长度、结束时间等中的至少一项来表示最大时间。
上述算力类型可以包括如下至少一种:中央处理器(Central Processing Unit,CPU)、图形处理器(graphics processing unit,GPU)、现场可编程逻辑门阵列(Field Programmable Gate Array,FPGA)、数据处理器(Data Processing Unit,DPU)、智能网卡(smart Network Interface Card,smartNIC)、张量处理器(Tensor Processing Unit,TPU)、神经网络处理器(Neural Network Processing Unit,NPU)中的至少一种。
可选地,上述算力类型还可以包括如下至少一项:
主频,例如最低主频;
核数,例如最少核数。
上述数据类型可以包括整数(例如,int8,int4等),浮点数(例如,半精度浮点数、单精度浮点数、双精度浮点数、脑半精度浮点数等)中的至少一种。
上述最小内存用于表示计算任务所需的内存下限。示例性地,上述最小内存可以通过内存大小(如8G)、可持续内存带宽和内存随机访问速率中的至少一项表示。
上述最小存储用于表示计算任务所需的存储下限。示例性地,上述最小存储可以通过存储大小(例如,1T)、存储带宽(例如,单位时间内最大的输入输出(Input Output,IO)流量)中的至少一项表示。
上述最小传输带宽用于表示计算任务的带宽下限,例如,计算任务可容忍的上行带宽下限和下行带宽下限中的至少一项。
可选地,所述最小传输带宽包括如下至少一项:最小上行带宽,最小下行带宽,上行带宽指示,下行带宽指示,目标指示;
其中,所述目标指示用于指示上行带宽和下行带宽相同或不同。
也即通过最小上行带宽、最小下行带宽、上行带宽指示、下行带宽指示和目标指示中的至少一项来表示最小传输带宽。
上述上行带宽指示用于指示上行带宽,例如,0表示上行带宽。上述下行带宽指示用于指示下行带宽,例如,1表示下行带宽。
上述目标指示用于指示上行带宽和下行带宽相同或不同,例如,0表示上行带宽与下行带宽不同,1表示上行带宽与下行带宽相同。
上述计算任务到达模式用于表示计算任务的计算作业的到达模式。
可选地,所述计算任务包括至少两个计算作业,所述计算任务到达模式包括如下至少一种:
连续到达模式或单一到达模式,用于指示一次到达一个计算作业;
固定周期到达模式,用于指示按照固定周期控制计算作业的到达;
泊松分布到达模式,用于指示基于泊松分布控制计算作业的到达;
高峰到达模式,用于指示在泊松分布的目标周期内控制σ个计算作业的到达,所述目标周期的时长小于预设时长,σ为正整数;
离线到达模式,用于指示一次到达所有计算作业。
上述连续到达模式或单一到达模式用于指示一次到达一个计算作业,例如,第i个计算作业在第(i-1)个计算作业完成后到达第i个计算作业,其中,计算作业(i-1)未完成或计算延迟预算门限未达到时,计算作业i不发送,i为正整数。
上述固定周期到达模式用于指示按照固定周期控制计算作业的到达。例如,每隔时间T到达n个计算作业,T表示固定周期,n为正整数。
上述泊松分布到达模式用于指示基于泊松分布控制计算作业的到达,例如,按照控制计算作业的到达,其中,k表示目标单位时间内计算作业的到达数,k为正整数,λ表示目标单位时间内计算作业的平均到达数,λ为正整数。示例性地,上述目标单位时间可以为1秒或2秒等。
上述高峰到达模式用于指示在泊松分布的目标周期内控制σ个计算作业的到达,其中,目标周期的时长小于预设时长,示例性地,上述预设时长可以为10秒或5秒等。示例性地,上述σ>25个计算作业/秒。
需要说明的是,泊松分布到达模式中有j个短周期,j为正整数,每个周期内有突发性大量的计算作业,周期持续一定时长TG,例如,5s至10s,并维持一定并发度水平σ。其中,短周期内的计算作业到达符合固定周期到达模式。
还需要说明的是,本实施例对于不同的计算任务的计算作业,可以采用相同的计算任务到达模式,或者可以采用不同的计算任务到达模式。对于同一计算任务的多个计算作业,可以采用一种计算任务到达模式,或者可以采用多种计算任务到达模式,例如,计算任务包括20个计算作业,其中10个计算作业采用单一到达模式,另外10个计算作业采用离线到达模式。
上述计算任务到达模式对应的参数,例如,对于固定周期到达模式,其对应的参数可以包括周期和每个周期的计算作业的到达数;对于泊松分布到达模式,其对应的参数可以包括目标单位时间内计算作业的平均到达数、目标单位时间等;对于高峰到达模式,其对应的参数可以包括目标周期和σ等。
上述AI模型训练精度可以通过单精度浮点数、脑半精度浮点数等输出的AI模型参数的数据类型表示。
上述AI模型性能下限可以包括如下至少一项:AI模型训练的性能下限,AI模型推理的性能下限。
对于AI模型,通常AI模型在不同场景对应的性能参数不同,例如,图像识别、物体检测的性能参数包括top1-准确率、平均准确率,语义分割的性能参数包括平均交并比(Mean Intersection Over Union,MIOU),语音识别的性能参数包括错词率(Word Error Rate,WER)。对于大模型,可以通过大模型基准测试的评分来表示性能下限。
可选地,所述计算任务的性能参数还可以包括数据集指示,该数据集指示可以与所述AI模型性能门限对应,所述数据集指示所指示的数据集对应的AI模型的性能需满足所述AI模型性能下限。其中,常用的数据集包括imagenet2012,pascal voc2012,LibriSpeech ASR corpus,criteo等。
上述AI模型推理的最小吞吐率,例如,视觉类通常是每秒推理的图片数(images/s),自然预研是每秒钟推理的句数(sentences/s)或者tokens/s,其中tokens指的是模型处理的输入文本中的单词、标点符号或其他文本单元。
上述计算功耗门限例如可以用于表示最大计算功耗。示例性地,计算功耗可以是指在上述计算延时预算的时间内计算节点的功耗;或者,计算功耗可以是指P2-P1或者P2与P1的比值,其中,P1表示计算节点在开机情况下单位时间内的功耗记,P2表示在处理计算任务时单位时间内的功耗。
计算能效门限例如可以表示最小计算能效。示例性的,计算能效可以是指单位时间内计算速度除以计算功耗,或AI模型推理的吞吐率除以计算功耗。
需要说明的是,上述计算任务的不同性能参数之间可以具有转换关系,也即可以基于计算任务的一部分性能参数确定计算任务的另一部分性能参数,因此,上述计算任务的性能参数可以仅包括上述列举的性能参数中的部分。以下进行举例说明:
例1:上述计算任务的性能参数包括最小计算速度。在该情况下,第一节点可将最小计算速度作为计算节点的选择参数。
例2:上述计算任务的性能参数包括最小运算数或操作数和计算延迟预算。如果计算延迟预算仅包含计算时延,那么第一节点可将这两个参数作为计算节点的选择参数;如果计算延迟预算包含端到端的时延,那么第一节点将数据传输时延和计算时延的分配信息、以及这两个参数作为计算节点的选择参数。
例3:上述计算任务的性能参数包括算力类型(例如CPU,主频,核数)、内存、存储和最大时间。以这些参数作为计算节点的选择参数的好处是相对静态,便于选择。
例4:上述计算任务的性能参数包括最小计算速度和最小传输带宽。
例5:上述计算任务的性能参数包括最小计算强度和计算延迟预算。
还需要说明的是,上述计算任务的性能参数可以仅适用于该计算任务,即仅该计算任务需满足上述性能参数,或者可以适用每个计算任务,即每个计算任务均需要满足上述性能参数,或者所有计算任务整体满足上述性能参数。上述计算任务组的性能参数可以适用于该计算任务组下的各个计算任务,即各个计算任务均需满足上述性能参数;或者,上述计算任务组的所有计算任务整体满足上述性能参数。
可选地,一个所述计算任务映射到一个服务质量QoS流;
或者,一个所述计算任务映射到一个QoS流集合;
或者,一个所述计算任务映射到一个协议数据单元(Protocol Data Unit,PDU)会话;
或者,一个所述计算任务映射到一个PDU会话集合;
或者,一个所述计算任务映射到一个无线承载(Radio Bearer,RB);
或者,一个所述计算任务映射到一个RB集合;
或者,一个所述计算任务映射到一个逻辑信道(Logic Channel,LC);
或者,一个所述计算任务映射到一个LC集合;
或者,一个所述计算任务映射到一个物理层资源;
或者,一个所述计算任务映射到一个物理层资源集合。
对于一个计算任务映射到一个物理层资源,例如,一个计算任务映射到一个通过下行控制信息(Downlink Control Information,DCI)调度的物理层资源。
对于一个计算任务映射到一个物理层资源集合,例如,一个计算任务映射到通过多个DCI调度的物理层资源。
示例性地,在一个计算任务映射到一个服务质量QoS流的情况下,该计算任务的性能参数与该QoS流对应。在一个计算任务映射到一个QoS流集合(即QoS flow set)的情况下,该计算任务的性能参数与该QoS流集合对应。在一个计算任务映射到一个PDU会话(Session)的情况下,该计算任务的性能参数与该PDU会话对应。在一个计算任务映射到一个PDU会话集合的情况下,该计算任务的性能参数与该PDU会话集合对应。在一个计算任务映射到一个PDU会话集合的情况下,该计算任务的性能参数与该PDU会话集合对应。在一个计算任务映射到一个RB的情况下,该计算任务的性能参数与该RB对应。在一个计算任务映射到一个RB集合的情况下,该计算任务的性能参数与该RB集合对应。在一个计算任务映射到一个LC的情况下,该计算任务的性能参数与该LC对应。在一个计算任务映射到一个LC集合的情况下,该计算任务的性能参数与该LC集合对应。在一个计算任务映射到一个物理层资源的情况下,该计算任务的性能参数与该物理层资源对应。在一个计算任务映射到一个物理层资源集合的情况下,该计算任务的性能参数与该物理层资源集合对应。
可选地,所述计算服务请求消息还包括计算服务标识,用于标识所述计算服务。
示例性地,上述计算服务标识可以用于标识一维DFT、二维FFT、AI模型训练、AI模型推理等计算任务。进一步地,上述AI模型训练或AI模型推理可以包括图像识别、物体检测、语义分割、推荐、自然语言处理、语音识别、光学字符识别、人脸识别、波束管理、定位、感知、CSI反馈等。
需要说明的是,在计算服务请求消息还包括计算服务标识的情况下,第一节点可以基于计算服务标识和计算服务的性能参数选择计算节点,例如,第一节点可以选择支持上述计算服务标识所标识的计算服务且满足上述性能参数的计算节点。
本实施例中,第二节点通过在计算服务请求消息中携带计算服务标识以指示其所请求的计算服务,这样便于第二节点灵活地请求不同的计算服务。
可选地,所述方法还包括:
所述第一节点根据所述计算服务请求消息和至少一个计算节点的状态信息选择计算节点;
其中,所述计算节点的状态信息包括如下至少一项:算力类型,计算负荷,可用计算速度,可用计算强度,可用内存,可用存储,计算功耗,计算能效。
本实施例中,第一节点根据所述计算服务请求消息和至少一个计算节点的状态信息选择计算节点,有利于更为准确地从上述至少一个计算节点中选择到满足上述计算服务的性能参数的计算节点。
可选地,所述计算服务响应消息包括如下至少一项:计算任务标识,计算节点标识,PDU会话信息。
本实施例中,上述计算任务标识用于标识计算任务。需要说明的是,计算服务与计算任务之间可以一一映射,或者,多个计算服务可以映射到一个计算任务,或者,一个计算服务可以映射到多个计算任务。
上述计算节点标识可以用于标识所选择的计算节点。
上述PDU会话信息可以用于确定上述计算服务相关的数据传输的PDU会话。
可以理解的是,在上述计算服务响应消息用于指示接受上述计算服务请求消息的情况下,上述计算服务响应消息可以包括计算任务标识、计算节点标识和PDU会话信息中的至少一项。
可选地,所述PDU会话信息包括如下一项:
建立PDU会话指示;
修改PDU会话指示和PDU会话标识;
PDU会话标识、QoS流标识和QoS规则。
以下分情况进行举例说明:
在需要新建PDU会话用于上述计算服务相关的数据传输的情况下,第一节点向第二节点发送建立PDU会话指示,第二节点在接收到建立PDU会话指示的情况下,可以建立PDU会话,所建立的PDU会话可以用于上述计算服务相关的数据传输。
在需要修改PDU会话用于上述计算服务相关的数据传输的情况下,第一节点向第二节点发送修改PDU会话指示和PDU会话标识,第二节点在接收到修改PDU会话指示和PDU会话标识的情况下,可以修改上述PDU会话标识所标识的PDU会话,修改后的PDU会话可以用于上述计算服务相关的数据传输。
在复用果已有PDU会话传输上述计算服务相关的数据传输的情况下,第一节点向第二节点发送PDU会话标识、QoS流标识和QoS规则,第二节点在接收到上述PDU会话标识、QoS流标识和QoS规则的情况下,可以基于上述PDU会话标识、QoS流标识和QoS规则进行上述计算服务相关的数据传输。
可选地,所述计算服务请求消息还用于指示所请求的计算节点是否为无线接入网节点。
本实施例中,上述计算服务请求消息还可以包括用于指示所请求的计算节点是否为无线接入网节点的信息。具体地,可以通过显式的方式指示所请求的计算节点是否为无线接入网节点,或者,可以通过隐式的方式指示所请求的计算节点是否为无线接入网节点。
可以理解的是,在计算服务请求消息指示所请求的计算节点为无线接入网节点的情况下,第一节点所选择的计算节点需要为无线接入网节点,或者,第一节点优先选择无线接入网节点作为计算节点。通过选择无线接入网节点作为计算节点,这样有利于减少计算服务的时延,更有利于满足低时延场景的需求。
可选地,所述计算服务请求消息包括如下至少一项:
第一指示信息,用于指示所请求的计算节点是否为无线接入网节点;
时延类型指示,用于指示所述计算服务的时延类型。
在一实施方式中,可以通过第一指示信息显式地指示所请求的计算节点是否为无线接入网节点。例如,通过一个比特指示所请求的计算节点是否为无线接入网节点,0表示请求所请求的计算节点不为无线接入网节点,1表示所请求的计算节点为无线接入网节点。
在另一实施方式中,可以通过时延类型指示所请求的计算节点是否为无线接入网节点,例如,在上述时延类型指示所指示的时延类型为预设时延类型的情况下,表示所请求的计算节点为无线接入网节点,否则表示所请求的计算节点不为无线接入网节点。示例性地,上述预设时延类型可以为低时延类型。
可选地,在所述计算服务为预设类型的计算服务或所述计算服务的资源类型为时延敏感(delay critical)类型的情况下,所选择的计算节点为无线接入网节点;
或者,
在所述第一指示信息指示所请求的计算节点为无线接入网节点的情况下,所选择的计算节点为无线接入网节点;
或者,
在所述时延类型指示所指示的时延类型为预设时延类型的情况下,所选择的计算节点为无线接入网节点。
示例性地,上述预设类型的计算服务可以包括用于移动网络优化的计算服务,例如,波束管理、定位、感知、CSI反馈等。上述时延敏感类型可以包括但不限于延迟敏感保证计算速度、延迟敏感保证计算强度等中的至少一项。
可选地,所述计算服务响应消息包括如下至少一项:第二指示信息,计算任务标识,计算节点标识;所述第二指示信息用于指示所选择的计算节点为无线接入网节点。
本实施例中的计算任务表示和计算节点标识可以参见前述实施例的相关描述,在此不做赘述。
可以理解的是,在所选择的计算节点为无线接入网节点的情况下,上述计算服务响应消息可以包括第二指示信息,用于指示所选择的计算节点为无线接入网节点,否则,上述计算服务响应消息不包括第二指示信息。
可选地,在所选择的计算节点为无线接入网节点的情况下,所述计算服务响应消息还包括无线承载指示,或者,所述计算服务响应消息还包括如下至少一项:物理层信道指示,物理层资源指示。
本实施例中,上述无线承载指示用于指示无线承载(Radio Bearer,RB),例如,上述无线承载指示可以包括无线承载标识(ID)。上述无线承载可以包括信令无线承载(Signal Radio Bearer,SRB),数据无线承载(Data Radio Bearer,DRB)或新定义的RB等,该无线承载可以用于计算服务相关的数据的传输。
需要说明的是,如果需要新建或修改已有无线承载,则第一节点可触发无线接入网节点中的通信功能发送无线承载添加或修改消息。
上述物理层信道指示可以用于指示物理层信道,该物理层信道可以用于计算服务相关的数据传输。上述物理层资源指示用于指示物理层资源,该物理层资源可以用于计算服务相关的数据传输。
需要说明的是,如果通过物理信道传输计算服务相关的数据,例如,计算数据和计算响应数据中的至少一项,则有利于进一步减少计算服务相关的数据的传输时延。
可选地,所述无线承载指示用于指示如下至少一项:信令无线承载SRB,数据无线承载DRB,数据面RB。
可选地,在所选择的计算节点为无线接入网节点的情况下,所述方法还包括:
所述第一节点向所述无线接入网节点发送第一信息,所述第一信息包括第三指示信息,所述第三指示信息用于指示所述无线接入网节点添加或修改无线承载。
本实施例中,在所选择的计算节点为无线接入网节点的情况下,第一节点向所述无线接入网节点发送第一信息,进而无线接入网节点可以基于该第一信息,向第二节点发送无线资源控制(Radio Resource Control,RRC)重配置消息,以请求针对计算服务,建立或修改无线承载或配置物理层资源等。
可选地,所述第一信息还包括所述计算服务的性能参数。
示例性地,通过在第一信息中包括所述计算服务的性能参数,使得无线接入网节点可以获知计算服务的性能参数,进而无线接入网节点可以基于计算服务的性能参数提供计算服务,这样有利于进一步保证计算服务的质量要求。
可选地,所述方法还包括:
所述第一节点向所选择的计算节点发送第一请求消息,所述第一请求消息用于请求建立或修改计算任务;
其中,所述第一请求消息包括如下至少一项:
计算任务标识;优先级指示;抢占能力指示;被抢占能力指示;迁移能力指示;保证计算速度,用于指示计算节点保证在平均窗口内向计算任务提供的计算速度;保证计算强度,用于指示计算节点保证在平均窗口内向计算任务提供的计算强度;最大计算速度,用于指示计算节点给计算任务提供的最大计算速度的上限;最大计算强度,用于指示计算节点给计算任务提供的最大计算强度的上限。
本实施例中,上述优先级指示用于表示该计算任务的计算资源请求的重要性。
上述抢占能力指示用于指示计算任务是否可以获得已分配给低优先级的另一个计算任务的计算资源。
被抢占能力指示用于指示为了执行具有更高优先级的计算任务,是否丢弃该计算任务。
迁移能力指示用于指示是否支持计算任务从一个计算节点迁移到另一个计算节点。
保证计算速度用于指示计算节点保证在平均窗口内向该计算任务提供的计算速度。
保证计算强度用于指示计算节点保证在平均窗口内向该计算任务提供的计算强度。
在一些可选的实施例中,上述保证计算速度可以包括时延敏感保证计算速度。上述保证计算强度可以包括时延敏感保证计算强度。
可选地,所述方法还包括:
所述第一节点从所选择的计算节点接收第一响应消息;
其中,所述第一响应消息包括所选择的计算节点建立或修改计算任务后的状态信息。
本实施例中,上述计算节点建立或修改计算任务后的状态信息可以包括但不限于算负荷、计算速度、计算强度、可用内存、可用存储、计算功耗、计算能效等中的至少一项。
示例性地,第一节点在接收到述第一响应消息的情况下,可以根据所选择的计算节点建立或修改计算任务后的状态信息,再次判断所选择的计算节点是否能够满足上述计算服务的性能参数,若满足,第一节点可以向第二节点发送计算服务响应消息,指示接受所述计算服务请求消息,若不满足,第一节点可以重新选择计算节点,这样有利于进一步保证计算服务的服务质量。
请参见图3,图3是本申请实施例提供的一种计算服务方法的流程图,该方法可以由第二节点执行,如图3所示,包括以下步骤:
步骤301、第二节点向第一节点发送计算服务请求消息,所述计算服务请求消息包括计算服务的性能参数;
步骤302、所述第二节点从所述第一节点接收计算服务响应消息,所述计算服务响应消息用于指示接受或拒绝所述计算服务请求消息。
可选地,所述计算服务的性能参数包括至少一个计算任务的性能参数或者至少一个计算任务组的性能参数,所述计算任务或者计算任务组的性能参数包括如下至少一项:
资源类型;最小运算数或操作数;最小计算速度;最小计算强度;计算延迟预算;最大失败率,所述失败率用于表示在单位时间内计算任务的失败请求数与总请求数的比值;平均窗口;最大时间,用于指示计算任务持续的时间上限;算力类型;数据类型;最小内存,用于指示计算任务所需的内存下限;最小存储,用于指示计算任务所需的存储下限;最小传输带宽,用于指示计算任务的带宽下限;计算任务到达模式;计算任务到达模式对应的参数;AI模型训练精度;AI模型性能下限;AI模型推理的最小吞吐率;计算功耗门限;计算能效门限。
可选地,所述计算服务请求消息还包括计算服务标识。
可选地,所述计算服务响应消息包括如下至少一项:计算任务标识,计算节点标识,PDU会话信息。
可选地,所述PDU会话信息包括如下一项:
建立PDU会话指示;
修改PDU会话指示和PDU会话标识;
PDU会话标识、QoS流标识和QoS规则。
具体地,在上述PDU会话信息包括建立PDU会话指示的情况下,第二节点可以向第三节点发送PDU建立请求消息,并接收第三节点返回的PDU建立响应消息;在上述PDU会话信息包括修改PDU会话指示和PDU会话标识的情况下,第二节点可以向第三节点发送PDU修改请求消息,并接收第三节点返回的PDU修改响应消息。示例性地,上述第三节点可以包括AMF和SMF中的至少一项。
需要说明的是,本实施例中PDU会话建立和PDU会话修改的具体流程可以参见相关技术,在此不做赘述。
可选地,所述计算服务请求消息还用于指示所请求的计算节点是否为无线接入网节点。
可选地,所述计算服务请求消息包括如下至少一项:
第一指示信息,用于指示所请求的计算节点是否为无线接入网节点;
时延类型指示,用于指示所述计算服务的时延类型。
可选地,所述计算服务响应消息包括如下至少一项:第二指示信息,计算任务标识,计算节点标识;所述第二指示信息用于指示所选择的计算节点为无线接入网节点。
可选地,所述计算服务响应消息还包括无线承载指示,或者,所述计算服务响应消息还包括如下至少一项:物理层信道指示,物理层资源指示。
可选地,在所选择的计算节点为无线接入网节点的情况下,所述方法还包括:
所述第二节点接收无线接入网节点发送的无线资源控制RRC重配置消息,所述RRC重配置消息包括用于计算服务的目标配置,所述目标配置包括如下至少一项:无线承载添加配置,无线承载修改配置,物理层资源配置;
所述第二节点向所述无线接入网节点发送RRC重配置完成消息。
示例性地,上述无线接入网节点可以包括基站或集中单元(Centralized Unit,CU)等。
本实施例中,第二节点在接收到用于计算服务的目标配置的情况下,可以基于目标配置进行计算服务相关的配置,例如,在上述目标配置包括无线承载添加配置的情况下,可以基于无线承载添加配置添加无线承载,添加的无线承载可以用于计算服务相关的数据传输;在上述目标配置包括无线承载修改配置的情况下,可以基于无线承载修改配置修改无线承载,修后的无线承载可以用于计算服务相关的数据传输;在上述目标配置包括物理层资源配置的情况下,可以基于上述物理层资源配置所配置的物理层资源传输计算服务相关的数据。
需要说明的是,通过无线承载或物理层传输计算服务相关的数据,有利于减少计算服务的传输时延。
可选地,所述方法还包括:
所述第二节点向第三节点发送第二请求消息,所述第二请求消息用于请求建立或修改PDU会话,所述第二请求消息包括第四指示信息,所述第四指示信息用于指示所述PDU会话支持用于计算服务的至少一个QoS流;
所述第二节点从所述第三节点接收第二响应消息,所述第二响应消息用于指示接受所述第二请求消息。
上述第三节点可以包括AMF和SMF中的至少一项。
上述用于计算服务的至少一个QoS流,可以理解为该QoS流满足用于计算服务的QoS参数,或者满足用于计算服务和通信服务的QoS参数。
示例性地,第三节点在接收到第二请求消息的情况下,可以执行如下至少一项:
根据所述第二请求消息选择第四节点,例如,UPF,并向选择的第四节点发送第三请求消息,所述第三请求消息用于请求建立或修改N4会话,所述第三请求消息包括用于计算服务或者用于计算服务和通信服务的第三信息,所述第三信息包括如下至少一项:规则标识(rule ID),优先级(precedence),数据包检测信息,转发规则(forwarding rule),执行规则(enforcement rule),报告规则(reporting rule);
向无线接入网节点(例如,基站)发送N2会话管理(Session Management,SM)消息,所述N2会话管理消息包括目标QoSI和第一服务质量配置文件(QoS Profile)中的至少一项,所述第一QoS Profile用于计算服务,或者所述第一QoS Profile用于计算服务和通信服务;
所述第三节点确定计算节点,例如,第三节点自己选择计算节点,或者第三节点从第一节点获取计算节点。可选地,根据终端提供的计算服务的包过滤器个数、包过滤器信息,所确定的计算节点的数量可以不止一个。
其中,上述规则标识用于唯一标识上述计算服务的规则。上述优先级用于确定应用所有计算服务的规则的检测信息的顺序。
上述数据包检测信息包括如下至少一项:终端互联网协议(Internet Protocol,IP)地址,核心网(Core Network,CN)隧道信息(tunnel info),包过滤器集(Packet filter set),目标QoSI。
其中,上述包过滤器可以包括IP地址、端口号等,用于检测哪些数据包是计算服务流的数据包。上述目标QoSI可以包括第一QoSI或第二QoSI,所述第一QoSI用于表示用于计算服务的至少一个QoS参数,所述第二QoSI用于表示用于计算服务和通信服务的至少一个QoS参数。所述第一QoSI可以用于表示如下至少一项:第一资源类型,第一优先级水平,第一计算延迟预算,失败率,平均窗口,最大运算数或操作数,最大计算速度,算力类型。所述第二QoSI可以用于表示如下至少一项:第二资源类型,第二优先级水平,第二计算延迟预算,带宽。需要说明的是,上述第一QoSI也可以称为计算QoSI,上述第二QoSI也可以称为计算和通信QoSI。
上述转发规则包括转发规则标识、定义将对应数据包转发至前述选择的计算节点(如计算节点IP地址)。
上述执行规则包括执行规则标识、定义执行的QoS操作。例如,对于时延敏感GCI或时延敏感GCR对应5QI是delay critical GBR类型。
上述报告规则包括报告规则标识、定义执行的测量操作。例如,测量UPF到计算节点的时延、吞吐量、数据量(吞吐量乘以时间)等。
需要说明的是,本申请实施例并不限定第二节点向第三节点发送第二请求消息以及第二节点向第一节点发送计算服务请求消息执行的先后顺序,例如,第二节点可以先执行向第三节点发送第二请求消息,再执行向第一节点发送计算服务请求消息;或者,终端可以先执行向第一节点发送计算服务请求消息,再执行向第三节点发送第二请求消息。
可选地,第三节点还可以从第四节点接收N4会话建立或修改响应消息。
本实施例中,由于上述PDU会话支持用于计算服务的至少一个QoS流,这样基于该PDU会话传输计算服务相关的数据,可以进一步保证计算服务的服务质量。
可选地,所述第二响应消息包括第一QoS规则,所述第一QoS规则用于计算服务或者所述第一QoS规则用于计算服务和通信服务。
示例性地,第三节点可以通过N1 SM容器(container)发送第一QoS规则。
可选地,所述第一QoS规则包括如下至少一项:
用于计算服务的QoS流的目标服务质量标识符QoSI,所述目标QoSI包括第一QoSI或第二QoSI,所述第一QoSI用于表示用于计算服务的至少一个QoS参数,所述第二QoSI用于表示用于计算服务和通信服务的至少一个QoS参数;
包过滤器集(Packet filter set);
优先级指示,所述优先级指示用于指示所述第一QoS规则的优先级。
示例性地,上述包过滤器可以包括IP地址、端口号等,用于检测哪些数据包是计算服务流的数据包。上述优先级指示可以用于确定数据包与多个QoS rule比对匹配时的顺序。
可选地,所述第一QoSI用于表示如下至少一项:第一资源类型,第一优先级水平,第一计算延迟预算,失败率,平均窗口,最大运算数或操作数,最大计算速度,算力类型。
上述第一资源类型也可以称为计算资源类型。示例性地,上述第一资源类型可以包括如下至少一项:保证计算速度,非保证计算速度,时延敏感保证计算速度,保证计算强度,非保证计算强度,时延敏感保证计算强度。
在一些可选的实施例中,上述第一资源类型决定与QoS流(QoS flow)级保证计算量相关的计算资源分配。或者上述第一资源类型决定与QoS flow级保证计算强度相关的计算资源相关的分配。其中,一个计算任务可以映射到一个QoS flow,或者一个计算任务映射到多个QoS flow,或者多个计算任务映射到一个QoS flow。
其中,计算速度和计算强度可以参见前述实施例的相关描述,在此不做赘述。可选地,在本实施例中的计算强度可以是分段定义的计算强度,例如,计算速度除以内存带宽。
上述第一优先级水平,用来确定计算QoS流的计算资源调度的优先级。
上述第一计算延迟预算可以用于指示QoS流的计算任务(也可以称为计算数据包或计算数据包集)在计算节点计算时可容忍的时延上限(即最大计算时延)。
示例性地,计算时延的一种定义是单个计算任务的第一个数据包发送到接收到计算任务的最后一个数据包的时间区间长度;另一种定义是一组计算任务的第一个数据包发送到最后一个数据包接收的时间区间长度。例如,对于图像识别的计算任务,一种是以单张图像识别作为单个计算任务,另一种是以多张(例如100张)图像识别作为一组计算任务。具体可以参见前述实施例对计算延迟预算的相关说明。
例如,对于面向AI模型推理的计算时延,可以包括如下至少一种:
推理总时延:具体是指多次连续推理总延时。计算方法是:第一个计算任务(或计算作业)推理前时间记为TITS,所有计算任务(或计算作业)推理结束时的时间记为TITE,那么AI模型推理的计算延迟预算为TITE-TITS。
推理时延:具体是指某样本推理的开始时间与结束时间的差。即某一个计算任务(或计算作业)的推理前计时的时间记为tINS,该计算任务推理结束时的时间记为tINE,那么AI模型推理的计算延迟预算为tINE-tINS。
上述失败率和平均窗口可以参见前述实施例的相关说明,在此不做赘述。
需要说明的是,上述第一QoSI内的平均窗口仅适用于GCR或GCI或delay critical GCR或delay critical GCI。
上述最大运算数或操作数用于表示QoS流的最大运算数或操作数。上述第一QoSI内的最大运算数或操作数仅适用于GCR或delay critical GCR,如2TFLOPs。
上述最大计算速度用于指示QoS流的最大计算速度。上述第一QoSI内的最大计算速度仅适用于GCI或delay critical GCI。可选的,上述最大计算速度可以包括理论最大计算速度和实际最大计算速度中的至少一项。
其中,上述最大计算速度所包括的理论最大计算速度用于指示QoS流要求计算节点的理论最大计算速度不高于上述最大计算速度所包括的理论最大计算速度。理论运算数峰值为处理器主频×处理器每个时钟周期执行运算次数×系统总核数。例如,GPU的主频是840MHz,共有256个计算逻辑单元(意味着处理器每个时钟周期执行256次浮点运算),共有3个核,那么该处理器的理论运算数峰值为840MHz×256×3=645.120G FLOPS(FP32),或者840MHz×256×3×2=1290.24G FLOPS(FP16)。因此,一种表示方法是通过处理器主频、核数等来表示计算节点的理论最大计算速度,另一种表示方法是通过FLOPS或每秒运算次数(Operations Per Second,OPS)等来表示理论最大计算速度。
上述最大计算速度所包括的实际最大计算速度用于指示QoS流要求计算节点的实际最大计算速度不高于上述最大计算速度所包括的实际最大计算速度。其中,实际最大运算速度是通过测试获得的最大计算速度。
在一些可选的实施例中,上述第一QoSI还可以用于表征第二测试用例指示,该第二测试用例指示与实际最大计算速度对应,用于表示该实际最大计算速度为基于该第二测试用例指示所指示的测试用例得到的最大计算速度。
例如,测试用例可以是MobileNetVx(x可以是任意可用的版本号,如1)、一维DFT、一维FFT、二维FFT、矩阵乘法、稀疏线性方程组、稠密线性方程组、YOLOvy(y可以是任意可用的版本号,如YOLOv5)、图像识别模型(如resnet50_v1.5等)和大模型(如Llama3,Llama2)等中的至少一项。测试用例可以是开源软件程序,也可以是UE和网络之间自定义的软件程序。这意味着QoS流所需的实际最大计算速度是在所指示的测试用例情况下计算节点可达到的实际最大计算速度。
在一些可选的实施例中,上述实际最大计算速度可以通过理论最大计算速度和计算效率表示,或者,可以通过理想最大计算速度和计算效率表示。其中,上述理想最大计算速度可以理解为在无任务抢占等理想情况下基于测试用例进行测试所得到的最大计算速度。
其中,上述计算效率的一种定义为:理想情况下实际最大计算速度与理论最大计算速度的比值。在该情况下,通过理论最大计算速度和计算效率表示该计算任务所需的实际最大计算速度。
上述计算效率的另一种定义为:基于测试用例测得的最大计算速度与理论最大计算速度的比值。在该情况下,上述第一QoSI还用于表征第二测试用例指示,且通过理论最大计算速度和计算效率可表示该QoS流所需的实际最大计算速度。
上述计算效率的又一种定义为:基于测试用例测得的计算速度与理想情况下最大计算速度的比值。基于测试用例测得的计算速度通常是在近期测试获得的,可以表示计算节点的当前状态。基于测试用例测得的理想情况下最大计算速度是指该计算节点最好的实测性能。在该情况下,上述第一QoSI还用于表征第二测试用例指示,且通过理想最大计算速度和计算效率可表示该QoS流所需的实际最大计算速度。
上述算力类型可以参见前述实施例的相关说明,在此不做赘述。
示例性地,上述第一QoSI可以用于标记计算服务的QoS流的转发处理参数的索引值。
还需要说明的是,本实施例中第一QoSI所表示的质量参数(即上述第一资源类型、第一优先级水平、第一计算延迟预算、失败率、平均窗口、最大运算数或操作数、最大计算速度和算力类型中的至少一项)对应于QoS流(QoS flow),也即上述第一QoSI为QoS流级的服务质量参数。
可选地,所述第二QoSI用于表示如下至少一项:第二资源类型,第二优先级水平,第二计算延迟预算,带宽。
上述第二资源类型也可以称为计算和通信资源类型。示例性地,上述第二资源类型可以包括如下至少一项:保证计算速度,非保证计算速度,时延敏感保证计算速度,保证计算强度,非保证计算强度,时延敏感保证计算强度。
在一些可选的实施例中,上述第二资源类型可以决定与QoS流(QoS flow)级保证计算量相关的计算和通信资源分配,或者,上述第二资源类型决定与QoS flow级保证计算强度相关的计算和通信资源相关的分配。其中,一个计算任务可以映射到一个QoS flow,或者一个计算任务映射到多个QoS flow,或者多个计算任务映射到一个QoS flow。
其中,计算速度和计算强度可以参见前述实施例的相关描述,在此不做赘述。
可选地,在本实施例中可以是将第二节点(如UE)到计算节点之间的传输带宽和计算节点的内存带宽等多个部分带宽中的最小一个作为计算计算强度的带宽,即计算强度为计算速度除以min{内存带宽、第二节点到计算节点的传输带宽},其中,min{内存带宽、第二节点到计算节点的传输带宽}表示传输带宽和内存带宽中的较小者。在一些可选的实施例中,上述第二节点到计算节点的传输带宽之间的传输带宽还可以分成第二节点到接入网节点之间的空口带宽,接入网节点到计算节点之间的有线传输带宽;或者第二节点和计算节点之间的传输带宽还可以分成第二节点到UPF之间的带宽,UPF到计算节点之间的有线传输带宽。
上述第二优先级水平用来确定用于计算服务的QoS流的计算和通信资源调度的优先级。
上述第二计算延迟预算用于表示QoS流的计算任务(也可以称为计算数据包或计算数据包集)在计算和传输时可容忍的时延上限。其中,上述计算和传输时可容忍的时延为计算时延和传输时延之和,上述传输时延可以是指第二节点到计算节点的传输时延和计算节点到计算接收节点的传输时延之和。
示例性地,计算和传输时延的一种定义是单个计算任务的第一个数据包发送到接收到计算任务的最后一个数据包的时间区间长度;另一种定义是一组计算任务的第一个数据包发送到最后一个数据包接收的时间区间长度。例如,对于图像识别的计算任务,一种是以单张图像识别作为单个计算任务,另一种是以多张(例如100张)图像识别作为一组计算任务。具体可以参见前述实施例对计算延迟预算的相关说明。
例如,对于面向AI模型推理的计算和传输时延,可以包括如下至少一种:
推理端到端总时延:具体是指多次连续推理端到端总延时。计算方法是:发送第一个计算任务(或计算作业)的第一字节前计时的时间记为TIS,计算接收节点接收到所有计算任务(或计算作业)的最后一个字节记为TIE,那么AI模型推理的计算延迟预算为TIE-TIS。
端到端推理时延:具体是指发送样本时间与收到结果时间的差。即第二节点发送某一个计算任务(或计算作业)的第一字节前计时的时间记为tTIS,计算接收节点接收到该计算任务(或计算作业)的最后一个字节记为tTIE,那么AI模型推理的计算延迟预算为tTIE-tTIS。
上述带宽用于表示QoS流的上行带宽下限和下行带宽下限中的至少一项。本实施例的带宽具体可以参见前述实施例的相关说明,在此不做赘述。
示例性地,上述第二QoSI可以用于标记计算和通信服务的QoS流的转发处理参数的索引值。
需要说明的是,本实施例中第二QoSI所表示的质量参数(即第二资源类型、第二优先级水平、第二计算延迟预算和带宽中的至少一项)对应于QoS流(QoS flow),也即上述第二QoSI为QoS流级的服务质量参数。
可选地,所述第二请求消息还包括如下至少一项:
用于计算服务的包过滤器的个数,用于计算服务的包过滤器。
可选地,所述第二响应消息包括计算节点标识。
示例性地,上述计算节点标识可以包括IP地址或网络内部ID等。
可选地,所述方法还包括:
所述第二节点向计算节点发送第二信息,所述第二信息包括计算数据。
可选地,所述第二信息还包括如下至少一项:
计算节点标识;
接收节点指示,用于指示接收所述计算数据对应的计算响应的节点。
需要说明的是,该实施方式的实现方式可以参见图2所示的实施例的相关说明,此处不作赘述。
以下结合示例对本实施例进行说明:
示例一:本示例的主要构思是:基于包含计算服务参数的计算服务请求消息,以解决移动网络提供计算服务时性能参数识别和交互问题,以及如何保证计算服务的服务质量问题。另外,本实施例中的计算任务是映射到PDU session和QoS flow,第一节点可以是核心网节点,计算节点可以是核心网节点或边缘计算节点等。
示例性地,参见图4,本申请实施例提供的计算服务方法包括如下步骤:
步骤11、第二节点发送计算服务请求消息给第一节点。
上述计算服务请求消息可以包括计算服务的性能参数,其中,计算服务的性能参数可以具体可以参见前述实施例的相关描述,在此不做赘述。
步骤12、第一节点向计算节点发送计算任务建立/修改请求消息。
第一节点在节点到计算服务请求消息,可以基于计算服务请求消息和至少一个计算节点的状态信息选择合适的计算节点。其中,计算节点的状态信息可以参见前述实施例的相关说明,在此不做赘述。
可选地,在选择到合适的计算节点的情况下,第一节点可以向选择的计算节点发送任务建立/修改请求消息。其中,上述计算任务建立/修改请求消息可以参见前述实施例的相关说明,在此不做赘述。
步骤13、计算节点向第一节点发送计算任务建立/修改响应消息。
其中,上述计算任务建立/修改响应消息可以参见前述实施例的相关描述,在此不做赘述。
需要说明的是,上述步骤12和步骤13可以为可选的步骤,例如,可以利用已有的计算任务进行计算和处理。
步骤14、第一节点向第二节点发送计算服务响应消息。
该步骤中,上述计算服务响应消息可以用于指示接受或不接受上述计算服务请求消息,具体可以参见前述实施例的相关描述,在此不做赘述。
需要说明的是,图4示出的是上述计算服务响应消息指示接受上述计算服务请求消息的情况。
步骤15、第二节点向第三节点发送PDU会话建立/修改请求消息。
步骤16、第三节点向第二节点发送PDU会话建立/修改响应消息。
上述步骤15和步骤16可以参见相关计算中PDU会话建立或修改的流程,在此不做赘述。需要说明的是,上述步骤15和步骤16可以为可选的步骤,例如,可以利用已有的PDU会话传输计算服务相关的数据。
步骤17、第二节点向计算节点发送计算数据。
例如,第二节点可以基于PDU会话传输上述计算数据。
步骤18a、计算节点向第二节点发送计算响应数据。
步骤18b、计算节点向计算接收节点发送计算响应数据。
可以理解的是,在该情况下,计算接收节点与第二节点为不同的节点。
示例二:本示例的主要构思是:将计算任务映射到无线承载或物理层资源,而不是PDU session或QoS flow。此外,第一节点可以是核心网节点或无线接入网节点,计算节点是无线接入网节点。该示例提供的计算服务方法可以解决移动网络提供计算服务时性能参数识别和交互问题,以及如何保证计算服务的服务质量问题,特别是面向低时延场景,或者无线接入网节点是信任节点的场景。
示例性地,参见图5,本申请实施例提供的计算服务方法包括如下步骤:
步骤21、第二节点发送计算服务请求消息给第一节点。
上述计算服务请求消息可以包括计算服务的性能参数,其中,计算服务的性能参数可以具体可以参见前述实施例的相关描述,在此不做赘述。可选地,上述计算服务请求消息还用于指示所请求的计算节点是否为无线接入网节点。
步骤22、第一节点向无线接入网节点发送计算任务建立/修改请求消息。
第一节点在节点到计算服务请求消息,可以基于计算服务请求消息和至少一个计算节点的状态信息选择合适的计算节点,且该计算节点为无线接入网节点。其中,计算节点的状态信息可以参见前述实施例的相关说明,在此不做赘述。
可选地,在选择到合适的无线接入网节点的情况下,第一节点可以向选择的无线接入网节点发送任务建立/修改请求消息。其中,上述计算任务建立/修改请求消息可以参见前述实施例的相关说明,在此不做赘述。
步骤23、无线接入网节点向第一节点发送计算任务建立/修改响应消息。
其中,上述计算任务建立/修改响应消息可以参见前述实施例的相关描述,在此不做赘述。
需要说明的是,上述步骤22和步骤23可以为可选的步骤,例如,可以利用已有的计算任务进行计算和处理。
步骤24、第一节点向第二节点发送计算服务响应消息。
该步骤中,上述计算服务响应消息可以用于指示接受或不接受上述计算服务请求消息,具体可以参见前述实施例的相关描述,在此不做赘述。
可选地,在该示例中,上述计算服务响应消息还可以包括如下至少一项:计算节点是无线接入网节点指示(即第二指示信息),无线承载指示,物理层信道指示,物理层资源指示。
需要说明的是,图5示出的是上述计算服务响应消息指示接受上述计算服务请求消息的情况。
步骤25、无线接入网节点向第二节点发送RRC重配置消息。
示例性地,该RRC重配置消息可以包括无线承载添加/修改请求。
其中,上述RRC重配置消息可以参见前述实施例的相关说明,在此不做赘述。
步骤26、第二节点向无线接入网节点发送RRC重配置完成消息。
步骤27、第二节点向计算节点发送计算数据。
例如,第二节点可以基于PDU会话传输上述计算数据。
步骤28a、计算节点向第二节点发送计算响应数据。
步骤28b、计算节点向计算接收节点发送计算响应数据。
可以理解的是,在该情况下,计算接收节点与第二节点为不同的节点。
示例三:本示例的主要构思是:将计算服务相关的服务质量参数定义为第一QoSI,将计算和通信相关的服务质量参数定义为第二QoSI,灵活变化的参数和每个计算任务精确数值要求通过每次计算任务的计算服务流程来交互。需要说明的是,第一QoSI和第二QoSI的含义可以参见前述实施例的相关说明,在此不做赘述。
示例性地,参见图6,本申请实施例提供的计算服务方法包括如下步骤:
步骤31、UE向第三节点发送PDU会话建立或修改请求消息,该PDU会话建立或修改请求消息中包含第四指示信息,用于指示该PDU会话包含用于计算服务的至少一个QoS流,即该PDU会话支持用于计算服务的至少一个QoS流。
步骤32、第三节点向第四节点发送N4会话建立或修改请求消息。
具体地,为了简化过程,第三节点对应AMF和SMF,AMF从UE接收PDU会话建立或修改请求消息,并根据是否需要支持计算服务的QoS流等信息选择合适的SMF。SMF根据PDU会话建立或修改请求消息选择合适的第四节点(例如,UPF)。此外,SMF还可以选择合适的计算节点或者SMF向计算管理节点获取合适的计算节点。
其中,所述N4会话建立或修改请求消息包含计算服务的包检测、执行和报告规则等,具体可以参见前述实施例的相关说明,在此不做赘述。
步骤33、第四节点发送N4会话建立或修改响应消息给第三节点。
例如,UPF发送N4会话建立消息或N4会话修改响应消息给SMF。
步骤34、第三节点发送N2会话管理信息给无线接入网节点,该N2会话管理信息可以包括目标QoSI、计算QoS profile中至少一项。
其中,目标QoSI可以用于标记计算服务的QoS流的转发处理参数的索引值。例如,第三节点将计算服务对应的QoS流映射到一个数据无线承载,其他非计算服务的QoS流映射到另一个数据无线承载。
步骤35、第三节点通过N1会话管理容器(container)发送第一QoS rule及计算节点标识数给UE。
步骤36、UE发送计算服务请求消息给第一节点。
该步骤同上述步骤11,在此不做赘述。
步骤37、第一节点向计算节点发送计算任务建立/修改请求消息。
该步骤同上述步骤12,在此不做赘述。
步骤38、计算节点向第一节点发送计算任务建立/修改响应消息。
该步骤同上述步骤13,在此不做赘述。
步骤39、第一节点向UE发送计算服务响应消息。
该步骤同上述步骤14,在此不做赘述。
步骤40、UE向计算节点发送计算数据。
该步骤同上述步骤17,在此不做赘述。
步骤41a、计算节点向UE发送计算响应数据。
该步骤同上述步骤18a,在此不做赘述。
步骤41b、计算节点向计算接收节点发送计算响应数据。
该步骤同上述步骤18b,在此不做赘述。
需要说明的是,本示例对上述步骤31至步骤35以及步骤36至步骤39的执行顺序并不做限定,例如,可以先执行步骤31至步骤35,再执行步骤36至步骤39;或者可以先执行步骤36至步骤39,再执行步骤31至步骤35。
通过本申请实施例提供的计算服务方法,基于计算服务请求消息传输计算服务对应的性能参数(例如,性能下限要求等),从而可以解决移动网络提供计算服务时的相关参数(特别是性能参数)定义、识别、传输和使用的问题,进而,可解决根据需求保证计算服务的服务质量问题。该方法既适用于对AF提供计算服务,也适用于对UE、NF提供计算服务。该方法具有更广泛的适用性,既适用于核心网节点作为计算管理节点和计算节点,也适用于核心网节点作为计算管理节点和无线接入网节点作为计算节点,也适用于无线接入网节点作为计算管理节点和计算节点。
需要说明的是,本申请实施例提供的计算服务方法,执行主体可以为计算服务装置。本申请实施例中以计算服务装置执行计算服务方法为例,说明本申请实施例提供的计算服务装置。
本申请实施例提供一种计算服务装置,作为一种示例,计算服务装置可以是通信设备或通信设备中的部件,例如芯片。该通信设备可以是终端、网络侧设备或服务器等。示例性的,终端可以包括但不限于上述所列举的终端11的类型,网络侧设备可以包括但不限于上述所列举的网络侧设备12的类型,本申请实施例不作具体限定。
计算服务装置包括接收模块、发送模块和处理模块。其中,接收模块、发送模块和处理模块可以是通过软件实现,也可以通过硬件实现。当通过硬件实现时,处理模块可以由处理器实现,示例性的,处理器可以包括通用处理器、专用处理器等,例如包括中央处理单元(Central Processing Unit,CPU)、微处理器、数字信号处理器(Digital Signal Processor,DSP)、人工智能(Artificial Intelligent,AI)处理器、图形处理器(Graphics Processing Unit,GPU)、专用集成电路(Application Specific Integrated Circuit,ASIC)、网络处理器(Network Processor,NP)、现场可编程门阵列(Field Programmable Gate Array,FPGA)或者其他可编程逻辑器件、门电路、晶体管、分立硬件组件等。接收模块和发送模块可以由通信接口实现,通信接口可以包括收发器、管脚、电路、总线、射频单元等其中一种或多种。
具体的,参见图7,当计算服务装置为网络侧设备或网络侧设备中的部件时,计算服务装置700包括接收模块701,用于从第二节点接收计算服务请求消息,所述计算服务请求消息包括计算服务的性能参数;发送模块702,用于向所述第二节点发送计算服务响应消息,其中,所述计算服务响应消息用于指示接受或拒绝所述计算服务请求消息。
可选地,所述计算服务的性能参数包括至少一个计算任务的性能参数或者至少一个计算任务组的性能参数,所述计算任务或者计算任务组的性能参数包括如下至少一项:
资源类型;最小运算数或操作数;最小计算速度;最小计算强度;计算延迟预算;最大失败率,所述失败率用于表示在单位时间内计算任务的失败请求数与总请求数的比值;平均窗口;最大时间,用于表示计算任务持续的时间上限;算力类型;数据类型;最小内存,用于表示计算任务所需的内存下限;最小存储,用于表示计算任务所需的存储下限;最小传输带宽,用于表示计算任务的带宽下限;计算任务到达模式;计算任务到达模式对应的参数;人工智能AI模型训练精度;AI模型性能下限;AI模型推理的最小吞吐率;计算功耗门限;计算能效门限。
可选地,所述资源类型包括如下至少一项:
保证计算速度,保证计算强度,非保证计算速度,非保证计算强度,时延敏感保证计算速度,时延敏感保证计算强度。
可选地,所述最小计算速度包括如下至少一项:理论最小计算速度,实际最小计算速度。
可选地,所述计算延迟预算包括如下至少一项:计算时延的上限,传输时延的上限,计算和传输时延的上限。
可选地,所述最小传输带宽包括如下至少一项:最小上行带宽,最小下行带宽,上行带宽指示,下行带宽指示,目标指示;
其中,所述目标指示用于指示上行带宽和下行带宽相同或不同。
可选地,所述计算任务包括至少两个计算作业,所述计算任务到达模式包括如下至少一种:
连续到达模式或单一到达模式,用于指示一次到达一个计算作业;
固定周期到达模式,用于指示按照固定周期控制计算作业的到达;
泊松分布到达模式,用于指示基于泊松分布控制计算作业的到达;
高峰到达模式,用于指示在泊松分布的目标周期内控制σ个计算作业的到达,所述目标周期的时长小于预设时长,σ为正整数;
离线到达模式,用于指示一次到达所有计算作业。
可选地,所述AI模型性能下限包括如下至少一项:AI模型训练的性能下限,AI模型推理的性能下限。
可选地,所述计算任务的性能参数还包括数据集指示,所述数据集指示所指示的数据集对应的AI模型的性能需满足所述AI模型性能下限。
可选地,一个所述计算任务映射到一个服务质量QoS流;
或者,一个所述计算任务映射到一个QoS流集合;
或者,一个所述计算任务映射到一个协议数据单元PDU会话;
或者,一个所述计算任务映射到一个PDU会话集合;
或者,一个所述计算任务映射到一个无线承载RB;
或者,一个所述计算任务映射到一个RB集合;
或者,一个所述计算任务映射到一个逻辑信道LC;
或者,一个所述计算任务映射到一个LC集合;
或者,一个所述计算任务映射到一个物理层资源;
或者,一个所述计算任务映射到一个物理层资源集合。
可选地,所述计算服务请求消息还包括计算服务标识,用于标识所述计算服务。
可选地,所述装置还包括:
处理模块,用于根据所述计算服务请求消息和至少一个计算节点的状态信息选择计算节点;
其中,所述计算节点的状态信息包括如下至少一项:算力类型,计算负荷,可用计算速度,可用计算强度,可用内存,可用存储,计算功耗,计算能效。
可选地,所述计算服务响应消息包括如下至少一项:计算任务标识,计算节点标识,PDU会话信息。
可选地,所述PDU会话信息包括如下一项:
建立PDU会话指示;
修改PDU会话指示和PDU会话标识;
PDU会话标识、QoS流标识和QoS规则。
可选地,所述计算服务请求消息还用于指示所请求的计算节点是否为无线接入网节点。
可选地,所述计算服务请求消息包括如下至少一项:
第一指示信息,用于指示所请求的计算节点是否为无线接入网节点;
时延类型指示,用于指示所述计算服务的时延类型。
可选地,在所述计算服务为预设类型的计算服务或所述计算服务的资源类型为时延敏感类型的情况下,所选择的计算节点为无线接入网节点;
或者,
在所述第一指示信息指示所请求的计算节点为无线接入网节点的情况下,所选择的计算节点为无线接入网节点;
或者,
在所述时延类型指示所指示的时延类型为预设时延类型的情况下,所选择的计算节点为无线接入网节点。
可选地,所述计算服务响应消息包括如下至少一项:第二指示信息,计算任务标识,计算节点标识;所述第二指示信息用于指示所选择的计算节点为无线接入网节点。
可选地,在所选择的计算节点为无线接入网节点的情况下,所述计算服务响应消息还包括无线承载指示,或者,所述计算服务响应消息还包括如下至少一项:物理层信道指示,物理层资源指示。
可选地,所述无线承载指示用于指示如下至少一项:信令无线承载SRB,数据无线承载DRB,数据面RB。
可选地,在所选择的计算节点为无线接入网节点的情况下,所述装置还包括:
向所述无线接入网节点发送第一信息,所述第一信息包括第三指示信息,所述第三指示信息用于指示所述无线接入网节点添加或修改无线承载。
可选地,所述第一信息还包括所述计算服务的性能参数。
可选地,所述发送模块,还用于向所选择的计算节点发送第一请求消息,所述第一请求消息用于请求建立或修改计算任务;
其中,所述第一请求消息包括如下至少一项:
计算任务标识;优先级指示;抢占能力指示;被抢占能力指示;迁移能力指示;保证计算速度,用于指示计算节点保证在平均窗口内向计算任务提供的计算速度;保证计算强度,用于指示计算节点保证在平均窗口内向计算任务提供的计算强度;最大计算速度,用于指示计算节点给计算任务提供的最大计算速度的上限;最大计算强度,用于指示计算节点给计算任务提供的最大计算强度的上限。
可选地,所述接收模块,还用于从所选择的计算节点接收第一响应消息;
其中,所述第一响应消息包括所选择的计算节点建立或修改计算任务后的状态信息。
本申请实施例提供的计算服务装置能够实现图2的方法实施例实现的各个过程,并达到相同的技术效果,为避免重复,这里不再赘述。
参见图8,当计算服务装置为终端或终端中的部件时,或者,当计算服务装置为网络侧设备或网络侧设备中的部件时,计算服务装置800包括发送模块801,用于向第一节点发送计算服务请求消息,所述计算服务请求消息包括计算服务的性能参数;接收模块802,用于从所述第一节点接收计算服务响应消息,所述计算服务响应消息用于指示接受或拒绝所述计算服务请求消息。
可选地,所述计算服务的性能参数包括至少一个计算任务的性能参数或者至少一个计算任务组的性能参数,所述计算任务或者计算任务组的性能参数包括如下至少一项:
资源类型;最小运算数或操作数;最小计算速度;最小计算强度;计算延迟预算;最大失败率,所述失败率用于表示在单位时间内计算任务的失败请求数与总请求数的比值;平均窗口;最大时间,用于指示计算任务持续的时间上限;算力类型;数据类型;最小内存,用于指示计算任务所需的内存下限;最小存储,用于指示计算任务所需的存储下限;最小传输带宽,用于指示计算任务的带宽下限;计算任务到达模式;计算任务到达模式对应的参数;AI模型训练精度;AI模型性能下限;AI模型推理的最小吞吐率;计算功耗门限;计算能效门限。
可选地,所述计算服务请求消息还包括计算服务标识。
可选地,所述计算服务响应消息包括如下至少一项:计算任务标识,计算节点标识,PDU会话信息。
可选地,所述PDU会话信息包括如下一项:
建立PDU会话指示;
修改PDU会话指示和PDU会话标识;
PDU会话标识、QoS流标识和QoS规则。
可选地,所述计算服务请求消息还用于指示所请求的计算节点是否为无线接入网节点。
可选地,所述计算服务请求消息包括如下至少一项:
第一指示信息,用于指示所请求的计算节点是否为无线接入网节点;
时延类型指示,用于指示所述计算服务的时延类型。
可选地,所述计算服务响应消息包括如下至少一项:第二指示信息,计算任务标识,计算节点标识;所述第二指示信息用于指示所选择的计算节点为无线接入网节点。
可选地,所述计算服务响应消息还包括无线承载指示,或者,所述计算服务响应消息还包括如下至少一项:物理层信道指示,物理层资源指示。
可选地,所述接收模块,还用于接收无线接入网节点发送的无线资源控制RRC重配置消息,所述RRC重配置消息包括用于计算服务的目标配置,所述目标配置包括如下至少一项:无线承载添加配置,无线承载修改配置,物理层资源配置;
所述发送模块,还用于向所述无线接入网节点发送RRC重配置完成消息。
可选地,所述发送模块,还用于向第三节点发送第二请求消息,所述第二请求消息用于请求建立或修改PDU会话,所述第二请求消息包括第四指示信息,所述第四指示信息用于指示所述PDU会话支持用于计算服务的至少一个QoS流;
所述接收模块,还用于从所述第三节点接收第二响应消息,所述第二响应消息用于指示接受所述第二请求消息。
可选地,所述第二响应消息包括第一QoS规则,所述第一QoS规则用于计算服务或者所述第一QoS规则用于计算服务和通信服务。
可选地,所述第一QoS规则包括如下至少一项:
用于计算服务的QoS流的目标服务质量标识符QoSI,所述目标QoSI包括第一QoSI或第二QoSI,所述第一QoSI用于表示用于计算服务的至少一个QoS参数,所述第二QoSI用于表示用于计算服务和通信服务的至少一个QoS参数;
包过滤器集;
优先级指示,所述优先级指示用于指示所述第一QoS规则的优先级。
可选地,所述第一QoSI用于表示如下至少一项:第一资源类型,第一优先级水平,第一计算延迟预算,失败率,平均窗口,最大运算数或操作数,最大计算速度,算力类型。
可选地,所述第二QoSI用于表示如下至少一项:第二资源类型,第二优先级水平,第二计算延迟预算,带宽。
可选地,所述第二请求消息还包括如下至少一项:
用于计算服务的包过滤器的个数,用于计算服务的包过滤器。
可选地,所述第二响应消息包括计算节点标识。
可选地,所述发送模块,还用于向计算节点发送第二信息,所述第二信息包括计算数据。
可选地,所述第二信息还包括如下至少一项:
计算节点标识;
接收节点指示,用于指示接收所述计算数据对应的计算响应的节点。
本申请实施例提供的计算服务装置能够实现图3的方法实施例实现的各个过程,并达到相同的技术效果,为避免重复,这里不再赘述。
如图9所示,本申请实施例还提供一种通信设备900,包括处理器901和存储器902,存储器902上存储有可在所述处理器901上运行的程序或指令,例如,该通信设备900为第一节点时,该程序或指令被处理器901执行时实现上述第一节点侧计算服务方法实施例的各个步骤,且能达到相同的技术效果。该通信设备900为第二节点时,该程序或指令被处理器901执行时实现上述第二节点侧计算服务方法实施例的各个步骤,且能达到相同的技术效果,为避免重复,这里不再赘述。
本申请实施例还提供一种网络侧设备,包括处理器和通信接口,所述通信接口和所述处理器耦合,所述处理器用于运行程序或指令,实现如图2或3所示的方法实施例的步骤。该网络侧设备实施例与上述第一节点或第二节点侧方法实施例对应,上述方法实施例的各个实施过程和实现方式均可适用于该网络侧设备实施例中,且能达到相同的技术效果。
具体地,本申请实施例还提供了一种网络侧设备,该网络侧设备可以是图7所示的计算服务装置。如图10所示,该网络侧设备1000包括:天线1001、射频装置1002、基带装置1003、处理器1004和存储器1005。天线1001与射频装置1002连接。在上行方向上,射频装置1002通过天线1001接收信息,将接收的信息发送给基带装置1003进行处理。在下行方向上,基带装置1003对要发送的信息进行处理,并发送给射频装置1002,射频装置1002对收到的信息进行处理后经过天线1001发送出去。
以上实施例中网络侧设备执行的方法可以在基带装置1003中实现,该基带装置1003包括基带处理器。
基带装置1003例如可以包括至少一个基带板,该基带板上设置有多个芯片,如图10所示,其中一个芯片例如为基带处理器,通过总线接口与存储器1005连接,以调用存储器1005中的程序,执行以上方法实施例中所示的网络设备操作。
该网络侧设备还可以包括网络接口1006,该接口例如为通用公共无线接口(Common Public Radio Interface,CPRI)。
具体地,本申请实施例的网络侧设备1000还包括:存储在存储器1005上并可在处理器1004上运行的指令或程序,处理器1004调用存储器1005中的指令或程序执行图7所示各模块执行的方法,并达到相同的技术效果,为避免重复,故不在此赘述。
具体地,本申请实施例还提供了一种网络侧设备。如图11所示,该网络侧设备1100包括:处理器1101、网络接口1102和存储器1103。该网络侧设备可以是图7或图8所示的计算服务装置。其中,网络接口1102例如为通用公共无线接口(common public radio interface,CPRI)。
具体地,本申请实施例的网络侧设备1100还包括:存储在存储器1103上并可在处理器1101上运行的指令或程序,处理器1101调用存储器1103中的指令或程序执行图7或图8所示各模块执行的方法,并达到相同的技术效果,为避免重复,故不在此赘述。
本申请实施例还提供一种终端,包括处理器和通信接口,所述通信接口和所述处理器耦合,所述处理器用于运行程序或指令,实现如图3所示方法实施例中的步骤。该终端实施例与上述终端侧方法实施例对应,上述方法实施例的各个实施过程和实现方式均可适用于该终端实施例中,且能达到相同的技术效果。该终端可以是图8所示的计算服务装置。具体地,图12为实现本申请实施例的一种终端的硬件结构示意图。
该终端1200包括但不限于:射频单元1201、网络模块1202、音频输出单元1203、输入单元1204、传感器1205、显示单元1206、用户输入单元1207、接口单元1208、存储器1209以及处理器1210等中的至少部分部件。
本领域技术人员可以理解,终端1200还可以包括给各个部件供电的电源(比如电池),电源可以通过电源管理系统与处理器12 10逻辑相连,从而通过电源管理系统实现管理充电、放电以及功耗管理等功能。图12中示出的终端结构并不构成对终端的限定,终端可以包括比图示更多或更少的部件,或者组合某些部件,或者不同的部件布置,在此不再赘述。
应理解的是,本申请实施例中,输入单元1204可以包括图形处理器12041和麦克风12042,图形处理器12041对在视频捕获模式或图像捕获模式中由图像捕获装置(如摄像头)获得的静态图片或视频的图像数据进行处理。显示单元1206可包括显示面板12061,可以采用液晶显示器、有机发光二极管等形式来配置显示面板12061。用户输入单元1207包括触控面板12071以及其他输入设备12072中的至少一种。触控面板12071,也称为触摸屏。触控面板12071可包括触摸检测装置和触摸控制器两个部分。其他输入设备12072可以包括但不限于物理键盘、功能键(比如音量控制按键、开关按键等)、轨迹球、鼠标、操作杆,在此不再赘述。
本申请实施例中,射频单元1201接收来自网络侧设备的下行数据后,可以传输给处理器1210进行处理;另外,射频单元1201可以向网络侧设备发送上行数据。通常,射频单元1201包括但不限于天线、放大器、收发器、耦合器、低噪声放大器、双工器等。
存储器1209可用于存储软件程序或指令以及各种数据。存储器1209可主要包括存储程序或指令的第一存储区和存储数据的第二存储区,其中,第一存储区可存储操作系统、至少一个功能所需的应用程序或指令(比如声音播放功能、图像播放功能等)等。此外,存储器1209可以包括易失性存储器或非易失性存储器。其中,非易失性存储器可以是只读存储器(Read-Only Memory,ROM)、可编程只读存储器(Programmable ROM,PROM)、可擦除可编程只读存储器(Erasable PROM,EPROM)、电可擦除可编程只读存储器(Electrically EPROM,EEPROM)或闪存。易失性存储器可以是随机存取存储器(Random Access Memory,RAM),静态随机存取存储器(Static RAM,SRAM)、动态随机存取存储器(Dynamic RAM,DRAM)、同步动态随机存取存储器(Synchronous DRAM,SDRAM)、双倍数据速率同步动态随机存取存储器(Double Data Rate SDRAM,DDRSDRAM)、增强型同步动态随机存取存储器(Enhanced SDRAM,ESDRAM)、同步连接动态随机存取存储器(Synch link DRAM,SLDRAM)和直接内存总线随机存取存储器(Direct Rambus RAM,DRRAM)。本申请实施例中的存储器1209包括但不限于这些和任意其它适合类型的存储器。
处理器1210可包括一个或多个处理单元;可选的,处理器1210集成应用处理器和调制解调处理器,其中,应用处理器主要处理涉及操作系统、用户界面和应用程序等的操作,调制解调处理器主要处理无线通信信号,如基带处理器。可以理解的是,上述调制解调处理器也可以不集成到处理器1210中。
其中,射频单元1201,用于从第二节点接收计算服务请求消息,所述计算服务请求消息包括计算服务的性能参数;
射频单元1201,还用于向所述第二节点发送计算服务响应消息,其中,所述计算服务响应消息用于指示接受或拒绝所述计算服务请求消息。
可以理解,本实施例中提及的各实现方式的实现过程可以参照计算服务方法实施例的相关描述,并达到相同或相应的技术效果,为避免重复,在此不再赘述。
本申请实施例还提供一种可读存储介质,所述可读存储介质上存储有程序或指令,该程序或指令被处理器执行时实现上述计算服务方法实施例的各个过程,且能达到相同的技术效果,为避免重复,这里不再赘述。
其中,所述处理器为上述实施例中所述的终端中的处理器。所述可读存储介质,包括计算机可读存储介质,如计算机只读存储器ROM、随机存取存储器RAM、磁碟或者光盘等。在一些示例中,可读存储介质可以是非瞬态的可读存储介质。
本申请实施例另提供了一种芯片,所述芯片包括处理器和通信接口,所述通信接口和所述处理器耦合,所述处理器用于运行程序或指令,实现上述计算服务方法实施例的各个过程,且能达到相同的技术效果,为避免重复,这里不再赘述。
应理解,本申请实施例提到的芯片还可以称为系统级芯片,系统芯片,芯片系统或片上系统芯片等。
本申请实施例另提供了一种计算机程序/程序产品,所述计算机程序/程序产品被存储在存储介质中,所述计算机程序/程序产品被至少一个处理器执行以实现上述计算服务方法实施例的各个过程,且能达到相同的技术效果,为避免重复,这里不再赘述。
本申请实施例还提供了一种无线通信系统,包括:第一节点及第二节点,所述第一节点可用于执行如上所述的计算服务方法的步骤,所述第二节点可用于执行如上所述的计算服务方法的步骤。
需要说明的是,在本文中,术语“包括”、“包含”或者其任何其他变体意在涵盖非排他性的包含,从而使得包括一系列要素的过程、方法、物品或者装置不仅包括那些要素,而且还包括没有明确列出的其他要素,或者是还包括为这种过程、方法、物品或者装置所固有的要素。在没有更多限制的情况下,由语句“包括一个……”限定的要素,并不排除在包括该要素的过程、方法、物品或者装置中还存在另外的相同要素。此外,需要指出的是,本申请实施方式中的方法和装置的范围不限按示出或讨论的顺序来执行功能,还可包括根据所涉及的功能按基本同时的方式或按相反的顺序来执行功能,例如,可以按不同于所描述的次序来执行所描述的方法,并且还可以添加、省去或组合各种步骤。另外,参照某些示例所描述的特征可在其他示例中被组合。
通过以上的实施方式的描述,本领域的技术人员可以清楚地了解到上述实施例方法可借助计算机软件产品加必需的通用硬件平台的方式来实现,当然也可以通过硬件。该计算机软件产品存储在存储介质(如ROM、RAM、磁碟、光盘等)中,包括若干指令,用以使得终端或者网络侧设备执行本申请各个实施例所述的方法。
上面结合附图对本申请的实施例进行了描述,但是本申请并不局限于上述的具体实施方式,上述的具体实施方式仅仅是示意性的,而不是限制性的,本领域的普通技术人员在本申请的启示下,在不脱离本申请宗旨和权利要求所保护的范围情况下,还可做出很多形式的实施方式,这些实施方式均属于本申请的保护之内。
Claims (49)
- 一种计算服务方法,包括:第一节点从第二节点接收计算服务请求消息,所述计算服务请求消息包括计算服务的性能参数;所述第一节点向所述第二节点发送计算服务响应消息,所述计算服务响应消息用于指示接受或拒绝所述计算服务请求消息。
- 根据权利要求1所述的方法,其特征在于,所述计算服务的性能参数包括至少一个计算任务的性能参数或者至少一个计算任务组的性能参数,所述计算任务或者计算任务组的性能参数包括如下至少一项:资源类型;最小运算数或操作数;最小计算速度;最小计算强度;计算延迟预算;最大失败率,所述失败率用于表示在单位时间内计算任务的失败请求数与总请求数的比值;平均窗口;最大时间,用于表示计算任务持续的时间上限;算力类型;数据类型;最小内存,用于表示计算任务所需的内存下限;最小存储,用于表示计算任务所需的存储下限;最小传输带宽,用于表示计算任务的带宽下限;计算任务到达模式;计算任务到达模式对应的参数;人工智能AI模型训练精度;AI模型性能下限;AI模型推理的最小吞吐率;计算功耗门限;计算能效门限。
- 根据权利要求2所述的方法,其中,所述资源类型包括如下至少一项:保证计算速度,保证计算强度,非保证计算速度,非保证计算强度,时延敏感保证计算速度,时延敏感保证计算强度。
- 根据权利要求2或3所述的方法,其中,所述最小计算速度包括如下至少一项:理论最小计算速度,实际最小计算速度。
- 根据权利要求2至4中任一项所述的方法,其中,所述计算延迟预算包括如下至少一项:计算时延的上限,传输时延的上限,计算和传输时延的上限。
- 根据权利要求2至5中任一项所述的方法,其中,所述最小传输带宽包括如下至少一项:最小上行带宽,最小下行带宽,上行带宽指示,下行带宽指示,目标指示;其中,所述目标指示用于指示上行带宽和下行带宽相同或不同。
- 根据权利要求2至6中任一项所述的方法,其中,所述计算任务包括至少两个计算作业,所述计算任务到达模式包括如下至少一种:连续到达模式或单一到达模式,用于指示一次到达一个计算作业;固定周期到达模式,用于指示按照固定周期控制计算作业的到达;泊松分布到达模式,用于指示基于泊松分布控制计算作业的到达;高峰到达模式,用于指示在泊松分布的目标周期内控制σ个计算作业的到达,所述目标周期的时长小于预设时长,σ为正整数;离线到达模式,用于指示一次到达所有计算作业。
- 根据权利要求2至7中任一项所述的方法,其中,所述AI模型性能下限包括如下至少一项:AI模型训练的性能下限,AI模型推理的性能下限。
- 根据权利要求2至8中任一项所述的方法,其中,所述计算任务的性能参数还包括数据集指示,所述数据集指示所指示的数据集对应的AI模型的性能需满足所述AI模型性能下限。
- 根据权利要求2至9中任一项所述的方法,其中,一个所述计算任务映射到一个服务质量QoS流;或者,一个所述计算任务映射到一个QoS流集合;或者,一个所述计算任务映射到一个协议数据单元PDU会话;或者,一个所述计算任务映射到一个PDU会话集合;或者,一个所述计算任务映射到一个无线承载RB;或者,一个所述计算任务映射到一个RB集合;或者,一个所述计算任务映射到一个逻辑信道LC;或者,一个所述计算任务映射到一个LC集合;或者,一个所述计算任务映射到一个物理层资源;或者,一个所述计算任务映射到一个物理层资源集合。
- 根据权利要求1至10中任一项所述的方法,其中,所述计算服务请求消息还包括计算服务标识,用于标识所述计算服务。
- 根据权利要求1至11中任一项所述的方法,其中,所述方法还包括:所述第一节点根据所述计算服务请求消息和至少一个计算节点的状态信息选择计算节点;其中,所述计算节点的状态信息包括如下至少一项:算力类型,计算负荷,可用计算速度,可用计算强度,可用内存,可用存储,计算功耗,计算能效。
- 根据权利要求1至12中任一项所述的方法,其中,所述计算服务响应消息包括如下至少一项:计算任务标识,计算节点标识,PDU会话信息。
- 根据权利要求13所述的方法,其中,所述PDU会话信息包括如下一项:建立PDU会话指示;修改PDU会话指示和PDU会话标识;PDU会话标识、QoS流标识和QoS规则。
- 根据权利要求1至12中任一项所述的方法,其中,所述计算服务请求消息还用于指示所请求的计算节点是否为无线接入网节点。
- 根据权利要求15所述的方法,其中,所述计算服务请求消息包括如下至少一项:第一指示信息,用于指示所请求的计算节点是否为无线接入网节点;时延类型指示,用于指示所述计算服务的时延类型。
- 根据权利要求16所述的方法,其中,在所述计算服务为预设类型的计算服务或所述计算服务的资源类型为时延敏感类型的情况下,所选择的计算节点为无线接入网节点;或者,在所述第一指示信息指示所请求的计算节点为无线接入网节点的情况下,所选择的计算节点为无线接入网节点;或者,在所述时延类型指示所指示的时延类型为预设时延类型的情况下,所选择的计算节点为无线接入网节点。
- 根据权利要求1至12、15至17中任一项所述的方法,其中,所述计算服务响应消息包括如下至少一项:第二指示信息,计算任务标识,计算节点标识;所述第二指示信息用于指示所选择的计算节点为无线接入网节点。
- 根据权利要求18所述的方法,其中,在所选择的计算节点为无线接入网节点的情况下,所述计算服务响应消息还包括无线承载指示,或者,所述计算服务响应消息还包括如下至少一项:物理层信道指示,物理层资源指示。
- 根据权利要求19所述的方法,其中,所述无线承载指示用于指示如下至少一项:信令无线承载SRB,数据无线承载DRB,数据面RB。
- 根据权利要求19或20所述的方法,其中,在所选择的计算节点为无线接入网节点的情况下,所述方法还包括:所述第一节点向所述无线接入网节点发送第一信息,所述第一信息包括第三指示信息,所述第三指示信息用于指示所述无线接入网节点添加或修改无线承载。
- 根据权利要求21所述的方法,其中,所述第一信息还包括所述计算服务的性能参数。
- 根据权利要求1至22中任一项所述的方法,其中,所述方法还包括:所述第一节点向所选择的计算节点发送第一请求消息,所述第一请求消息用于请求建立或修改计算任务;其中,所述第一请求消息包括如下至少一项:计算任务标识;优先级指示;抢占能力指示;被抢占能力指示;迁移能力指示;保证计算速度,用于指示计算节点保证在平均窗口内向计算任务提供的计算速度;保证计算强度,用于指示计算节点保证在平均窗口内向计算任务提供的计算强度;最大计算速度,用于指示计算节点给计算任务提供的最大计算速度的上限;最大计算强度,用于指示计算节点给计算任务提供的最大计算强度的上限。
- 根据权利要求23所述的方法,其中,所述方法还包括:所述第一节点从所选择的计算节点接收第一响应消息;其中,所述第一响应消息包括所选择的计算节点建立或修改计算任务后的状态信息。
- 一种计算服务方法,包括:第二节点向第一节点发送计算服务请求消息,所述计算服务请求消息包括计算服务的性能参数;所述第二节点从所述第一节点接收计算服务响应消息,所述计算服务响应消息用于指示接受或拒绝所述计算服务请求消息。
- 根据权利要求25所述的方法,其中,所述计算服务的性能参数包括至少一个计算任务的性能参数或者至少一个计算任务组的性能参数,所述计算任务或者计算任务组的性能参数包括如下至少一项:资源类型;最小运算数或操作数;最小计算速度;最小计算强度;计算延迟预算;最大失败率,所述失败率用于表示在单位时间内计算任务的失败请求数与总请求数的比值;平均窗口;最大时间,用于指示计算任务持续的时间上限;算力类型;数据类型;最小内存,用于指示计算任务所需的内存下限;最小存储,用于指示计算任务所需的存储下限;最小传输带宽,用于指示计算任务的带宽下限;计算任务到达模式;计算任务到达模式对应的参数;AI模型训练精度;AI模型性能下限;AI模型推理的最小吞吐率;计算功耗门限;计算能效门限。
- 根据权利要求25或26所述的方法,其中,所述计算服务请求消息还包括计算服务标识。
- 根据权利要求25至27中任一项所述的方法,其中,所述计算服务请求消息还用于指示所请求的计算节点是否为无线接入网节点。
- 根据权利要求28所述的方法,其中,所述计算服务请求消息包括如下至少一项:第一指示信息,用于指示所请求的计算节点是否为无线接入网节点;时延类型指示,用于指示所述计算服务的时延类型。
- 根据权利要求25至29中任一项所述的方法,其中,在所选择的计算节点为无线接入网节点的情况下,所述方法还包括:所述第二节点接收无线接入网节点发送的无线资源控制RRC重配置消息,所述RRC重配置消息包括用于计算服务的目标配置,所述目标配置包括如下至少一项:无线承载添加配置,无线承载修改配置,物理层资源配置;所述第二节点向所述无线接入网节点发送RRC重配置完成消息。
- 根据权利要求25至27中任一项所述的方法,其中,所述方法还包括:所述第二节点向第三节点发送第二请求消息,所述第二请求消息用于请求建立或修改PDU会话,所述第二请求消息包括第四指示信息,所述第四指示信息用于指示所述PDU会话支持用于计算服务的至少一个QoS流;所述第二节点从所述第三节点接收第二响应消息,所述第二响应消息用于指示接受所述第二请求消息。
- 根据权利要求31所述的方法,其中,所述第二响应消息包括第一QoS规则,所述第一QoS规则用于计算服务或者所述第一QoS规则用于计算服务和通信服务。
- 根据权利要求32所述的方法,其中,所述第一QoS规则包括如下至少一项:用于计算服务的QoS流的目标服务质量标识符QoSI,所述目标QoSI包括第一QoSI或第二QoSI,所述第一QoSI用于表示用于计算服务的至少一个QoS参数,所述第二QoSI用于表示用于计算服务和通信服务的至少一个QoS参数;包过滤器集;优先级指示,所述优先级指示用于指示所述第一QoS规则的优先级。
- 根据权利要求33所述的方法,其中,所述第一QoSI用于表示如下至少一项:第一资源类型,第一优先级水平,第一计算延迟预算,失败率,平均窗口,最大运算数或操作数,最大计算速度,算力类型;和/或,所述第二QoSI用于表示如下至少一项:第二资源类型,第二优先级水平,第二计算延迟预算,带宽。
- 根据权利要求31至34中任一项所述的方法,其中,所述第二请求消息还包括如下至少一项:用于计算服务的包过滤器的个数,用于计算服务的包过滤器。
- 根据权利要求31至35中任一项所述的方法,其中,所述第二响应消息包括计算节点标识。
- 根据权利要求25至36中任一项所述的方法,其中,所述方法还包括:所述第二节点向计算节点发送第二信息,所述第二信息包括计算数据。
- 根据权利要求37所述的方法,其中,所述第二信息还包括如下至少一项:计算节点标识;接收节点指示,用于指示接收所述计算数据对应的计算响应的节点。
- 一种计算服务装置,包括:接收模块,用于从第二节点接收计算服务请求消息,所述计算服务请求消息包括计算服务的性能参数;发送模块,用于向所述第二节点发送计算服务响应消息,其中,所述计算服务响应消息用于指示接受或拒绝所述计算服务请求消息。
- 根据权利要求39所述的装置,其中,所述计算服务的性能参数包括至少一个计算任务的性能参数或者至少一个计算任务组的性能参数,所述计算任务或者计算任务组的性能参数包括如下至少一项:资源类型;最小运算数或操作数;最小计算速度;最小计算强度;计算延迟预算;最大失败率,所述失败率用于表示在单位时间内计算任务的失败请求数与总请求数的比值;平均窗口;最大时间,用于表示计算任务持续的时间上限;算力类型;数据类型;最小内存,用于表示计算任务所需的内存下限;最小存储,用于表示计算任务所需的存储下限;最小传输带宽,用于表示计算任务的带宽下限;计算任务到达模式;计算任务到达模式对应的参数;人工智能AI模型训练精度;AI模型性能下限;AI模型推理的最小吞吐率;计算功耗门限;计算能效门限。
- 根据权利要求39或40所述的装置,其中,所述装置还包括:处理模块,用于根据所述计算服务请求消息和至少一个计算节点的状态信息选择计算节点;其中,所述计算节点的状态信息包括如下至少一项:算力类型,计算负荷,可用计算速度,可用计算强度,可用内存,可用存储,计算功耗,计算能效。
- 一种计算服务装置,包括:发送模块,用于向第一节点发送计算服务请求消息,所述计算服务请求消息包括计算服务的性能参数;接收模块,用于从所述第一节点接收计算服务响应消息,所述计算服务响应消息用于指示接受或拒绝所述计算服务请求消息。
- 根据权利要求42所述的装置,其中,所述计算服务的性能参数包括至少一个计算任务的性能参数或者至少一个计算任务组的性能参数,所述计算任务或者计算任务组的性能参数包括如下至少一项:资源类型;最小运算数或操作数;最小计算速度;最小计算强度;计算延迟预算;最大失败率,所述失败率用于表示在单位时间内计算任务的失败请求数与总请求数的比值;平均窗口;最大时间,用于指示计算任务持续的时间上限;算力类型;数据类型;最小内存,用于指示计算任务所需的内存下限;最小存储,用于指示计算任务所需的存储下限;最小传输带宽,用于指示计算任务的带宽下限;计算任务到达模式;计算任务到达模式对应的参数;AI模型训练精度;AI模型性能下限;AI模型推理的最小吞吐率;计算功耗门限;计算能效门限。
- 根据权利要求42或43所述的装置,其中,所述发送模块,还用于向第三节点发送第二请求消息,所述第二请求消息用于请求建立或修改PDU会话,所述第二请求消息包括第四指示信息,所述第四指示信息用于指示所述PDU会话支持用于计算服务的至少一个QoS流;所述接收模块,还用于从所述第三节点接收第二响应消息,所述第二响应消息用于指示接受所述第二请求消息。
- 根据权利要求44所述的装置,其中,所述第二响应消息包括第一QoS规则,所述第一QoS规则用于计算服务或者所述第一QoS规则用于计算服务和通信服务。
- 一种第一节点,包括处理器和存储器,所述存储器存储可在所述处理器上运行的程序或指令,所述程序或指令被所述处理器执行时实现如权利要求1至24任一项所述的计算服务方法的步骤。
- 一种第二节点,包括处理器和存储器,所述存储器存储可在所述处理器上运行的程序或指令,所述程序或指令被所述处理器执行时实现如权利要求25至38任一项所述的计算服务方法的步骤。
- 一种可读存储介质,所述可读存储介质上存储程序或指令,所述程序或指令被处理器执行时实现如权利要求1至24任一项所述的计算服务方法的步骤,或者实现权利要求25至38任一项所述的计算服务方法的步骤。
- 一种计算机程序产品,所述计算机程序产品被至少一个处理器执行以实现如权利要求1至24任一项所述的计算服务方法的步骤,或者实现权利要求25至38任一项所述的计算服务方法的步骤。
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN202410909436.6A CN121310096A (zh) | 2024-07-08 | 2024-07-08 | 计算服务方法、装置、第一节点及第二节点 |
| CN202410909436.6 | 2024-07-08 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2026012196A1 true WO2026012196A1 (zh) | 2026-01-15 |
Family
ID=98277518
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/CN2025/105250 Pending WO2026012196A1 (zh) | 2024-07-08 | 2025-06-30 | 计算服务方法、装置、第一节点及第二节点 |
Country Status (2)
| Country | Link |
|---|---|
| CN (1) | CN121310096A (zh) |
| WO (1) | WO2026012196A1 (zh) |
Citations (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| WO2022067816A1 (zh) * | 2020-09-30 | 2022-04-07 | 华为技术有限公司 | 一种网络边缘计算方法及通信装置 |
| CN118200882A (zh) * | 2022-12-12 | 2024-06-14 | 中国移动通信有限公司研究院 | Ai服务请求处理方法、装置、设备及可读存储介质 |
| CN118283712A (zh) * | 2022-12-22 | 2024-07-02 | 维沃软件技术有限公司 | 计算服务的实现方法、装置、通信设备及可读存储介质 |
-
2024
- 2024-07-08 CN CN202410909436.6A patent/CN121310096A/zh active Pending
-
2025
- 2025-06-30 WO PCT/CN2025/105250 patent/WO2026012196A1/zh active Pending
Patent Citations (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| WO2022067816A1 (zh) * | 2020-09-30 | 2022-04-07 | 华为技术有限公司 | 一种网络边缘计算方法及通信装置 |
| CN118200882A (zh) * | 2022-12-12 | 2024-06-14 | 中国移动通信有限公司研究院 | Ai服务请求处理方法、装置、设备及可读存储介质 |
| CN118283712A (zh) * | 2022-12-22 | 2024-07-02 | 维沃软件技术有限公司 | 计算服务的实现方法、装置、通信设备及可读存储介质 |
Also Published As
| Publication number | Publication date |
|---|---|
| CN121310096A (zh) | 2026-01-09 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| US11297601B2 (en) | Resource allocation method and orchestrator for network slicing in the wireless access network | |
| CN115280801B (zh) | 定位请求处理方法、设备及系统 | |
| CN110785984B (zh) | 信息获取方法、信息获取装置和电子设备 | |
| CN111989948A (zh) | 定位测量数据上报方法、装置、终端及存储介质 | |
| CN114302426A (zh) | 在异质网络控制服务质量的方法、装置、介质及电子设备 | |
| WO2022027386A1 (zh) | 一种天线选择方法及装置 | |
| CN117319993A (zh) | 信息传输方法、装置及电子设备 | |
| CN112866897A (zh) | 一种定位测量方法、终端和网络节点 | |
| WO2026012196A1 (zh) | 计算服务方法、装置、第一节点及第二节点 | |
| EP4525503A1 (en) | Sidelink positioning method and apparatus, terminal, server, and wireless access network device | |
| WO2026012267A1 (zh) | 计算服务方法、会话管理方法、装置、终端及第一节点 | |
| EP4595595A1 (en) | Priority for the power allocation for prach transmission for ta acquisition | |
| WO2026012203A1 (zh) | 信息传输方法、装置及通信设备 | |
| CN116418880A (zh) | Ai网络信息传输方法、装置及通信设备 | |
| WO2026012236A1 (zh) | 接入网设备切换方法、装置及设备 | |
| WO2026012237A1 (zh) | 计算节点切换方法、装置及设备 | |
| CN120676371A (zh) | 任务管理方法、装置、终端及网络侧设备 | |
| WO2025026299A1 (zh) | Ai业务建立方法、装置及网络侧设备 | |
| CN111867113B (zh) | 上行资源调度方法、终端及网络侧设备 | |
| WO2025195240A1 (zh) | 任务处理方法、装置及相关设备 | |
| WO2026001812A1 (zh) | 通信方法、终端及网络侧设备 | |
| WO2025113310A1 (zh) | 通信方法和相关装置 | |
| CN120670097A (zh) | 任务处理方法、装置及相关设备 | |
| WO2024255684A1 (zh) | 数据传输方法、装置、发送节点及接收节点 | |
| WO2025195239A1 (zh) | 任务处理方法、装置及相关设备 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 25836242 Country of ref document: EP Kind code of ref document: A1 |