WO2025201273A1 - 模型处理方法、装置、通信设备及存储介质 - Google Patents
模型处理方法、装置、通信设备及存储介质Info
- Publication number
- WO2025201273A1 WO2025201273A1 PCT/CN2025/084528 CN2025084528W WO2025201273A1 WO 2025201273 A1 WO2025201273 A1 WO 2025201273A1 CN 2025084528 W CN2025084528 W CN 2025084528W WO 2025201273 A1 WO2025201273 A1 WO 2025201273A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- model
- input
- output
- models
- final
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Pending
Links
Classifications
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04W—WIRELESS COMMUNICATION NETWORKS
- H04W24/00—Supervisory, monitoring or testing arrangements
- H04W24/02—Arrangements for optimising operational condition
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/04—Architecture, e.g. interconnection topology
- G06N3/045—Combinations of networks
- G06N3/0455—Auto-encoder networks; Encoder-decoder networks
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/08—Learning methods
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04B—TRANSMISSION
- H04B17/00—Monitoring; Testing
- H04B17/30—Monitoring; Testing of propagation channels
- H04B17/373—Predicting channel quality or other radio frequency [RF] parameters
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04B—TRANSMISSION
- H04B17/00—Monitoring; Testing
- H04B17/30—Monitoring; Testing of propagation channels
- H04B17/391—Modelling the propagation channel
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04B—TRANSMISSION
- H04B17/00—Monitoring; Testing
- H04B17/30—Monitoring; Testing of propagation channels
- H04B17/391—Modelling the propagation channel
- H04B17/3913—Predictive models, e.g. based on neural network models
Definitions
- the present application belongs to the field of communication technology, and specifically relates to a model processing method, apparatus, communication equipment and storage medium.
- AI Artificial Intelligence
- different AI models can be trained using different methods or datasets. These models may have different parameters and input-to-output mappings.
- a single AI model is selected from multiple models for use.
- the performance of a single model is limited, making it adaptable to only certain scenarios or environments and unable to achieve optimal gains.
- the embodiments of the present application provide a model processing method, apparatus, communication equipment, and storage medium, which can utilize multiple AI models to obtain better gain, thereby helping to improve the performance of wireless communication networks.
- a model processing method comprising:
- the communication device obtains at least two first models
- the communication device uses the at least two first models to obtain a second model or a partial model of the second model or a first data set.
- a model processing device comprising:
- a first obtaining module configured to obtain at least two first models
- the second obtaining module is used to obtain the second model or a partial model of the second model or the first data set by using the at least two first models.
- a communication device which includes a processor and a memory, wherein the memory stores programs or instructions that can be run on the processor, and when the program or instructions are executed by the processor, the steps of the method described in the first aspect are implemented.
- a communication device comprising a processor and a communication interface, wherein the communication interface is coupled to the processor, and the processor is used to run programs or instructions to implement the steps of the method described in the first aspect.
- a terminal which includes a processor and a memory, wherein the memory stores a program or instruction that can be run on the processor, and when the program or instruction is executed by the processor, the steps of the method described in the first aspect are implemented.
- a terminal comprising a processor and a communication interface, wherein the communication interface is coupled to the processor, and the processor is used to run programs or instructions to implement the steps of the method described in the first aspect.
- a network side device which includes a processor and a memory, wherein the memory stores programs or instructions that can be run on the processor, and when the program or instructions are executed by the processor, the steps of the method described in the first aspect are implemented.
- a network side device comprising a processor and a communication interface, wherein the communication interface is coupled to the processor, and the processor is used to run programs or instructions to implement the steps of the method described in the first aspect.
- a readable storage medium on which a program or instruction is stored.
- the program or instruction is executed by a processor, the steps of the method described in the first aspect are implemented.
- a wireless communication system comprising: a terminal and a network side device, wherein the terminal can be used to execute the steps of the method described in the first aspect, or the network side device can be used to execute the steps of the method described in the first aspect.
- a chip comprising a processor and a communication interface, wherein the communication interface is coupled to the processor, and the processor is used to run programs or instructions to implement the steps of the method described in the first aspect.
- the communication device uses the at least two first models to obtain a second model or a first data set, and fuses multiple first models to obtain a second model or a partial model of the second model or a first data set, thereby aligning the mapping relationship between the input and output of multiple first models.
- Applying the second model or the partial model of the second model or the first data set to the wireless communication network can obtain better gain and help improve the performance of the wireless communication network.
- FIG1 is a block diagram of a wireless communication system applicable to embodiments of the present application.
- FIG3 is a schematic diagram of a neuron in the related art
- FIG4 is a flowchart of an implementation of a model processing method in an embodiment of the present application.
- FIG5 is a schematic diagram of a unilateral model processing process in an embodiment of the present application.
- FIG6 is a schematic diagram of a bilateral model processing process in an embodiment of the present application.
- FIG7 is a schematic structural diagram of a model processing device according to an embodiment of the present application.
- FIG8 is a schematic structural diagram of a communication device according to an embodiment of the present application.
- FIG9 is a schematic structural diagram of a terminal according to an embodiment of the present application.
- FIG10 is a schematic structural diagram of a network-side device according to an embodiment of the present application.
- FIG11 is a schematic structural diagram of another network-side device in an embodiment of the present application.
- first, second, etc. in this application are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the terms used in this way are interchangeable where appropriate, so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by “first” and “second” are generally of the same type, and do not limit the number of objects, for example, the first object can be one or more.
- “or” in this application represents at least one of the connected objects. For example, “A or B” covers three options, namely, Option 1: including A but not including B; Option 2: including B but not including A; Option 3: including both A and B.
- the character "/" generally indicates that the objects associated before and after are in an "or” relationship.
- indication in this application can be either a direct indication (or explicit indication) or an indirect indication (or implicit indication).
- a direct indication can be understood as the sender explicitly informing the receiver of specific information, the operation to be performed, or the requested result, etc. in the instruction sent;
- an indirect indication can be understood as the receiver determining the corresponding information based on the instruction sent by the sender, or making a judgment and determining the operation to be performed or the requested result, etc. based on the judgment result.
- LTE Long Term Evolution
- LTE-A Long Term Evolution
- CDMA Code Division Multiple Access
- TDMA Time Division Multiple Access
- FDMA Frequency Division Multiple Access
- OFDMA Orthogonal Frequency Division Multiple Access
- SC-FDMA Single-carrier Frequency Division Multiple Access
- NR New Radio
- 6G 6th Generation
- the core network equipment may include but is not limited to at least one of the following: core network nodes, core network functions, Mobility Management Entity (MME), Access and Mobility Management Function (AMF), Session Management Function (SMF), User Plane Function (UPF), Policy Control Function (PCF), Policy and Charging Rules Function (PCRF), Edge Application Server Discovery Function (EASDF), Unified Data Management (UDM), Unified Data Repository (UDR), and so on.
- MME Mobility Management Entity
- AMF Access and Mobility Management Function
- SMF Session Management Function
- UPF User Plane Function
- PCF Policy Control Function
- PCF Policy and Charging Rules Function
- EASDF Edge Application Server Discovery Function
- UDM Unified Data Management
- UDM Unified Data Repository
- the communication device uses at least two first models to obtain a second model or a partial model of the second model or a first data set.
- the communication device uses the at least two first models to obtain the second model or a partial model of the second model or a first data set, and fuses multiple first models to obtain the second model or a partial model of the second model or a first data set, thereby aligning the mapping relationship between the input and output of multiple first models.
- Applying the second model or the partial model of the second model or the first data set to the wireless communication network can obtain better gain and help improve the performance of the wireless communication network.
- a communications device may obtain a partial model of a first model and, using the partial model of the first model and a third data set, obtain at least two first models.
- the obtained at least two first models may be understood as a complete first model.
- the first model includes a first part and a second part, the first part is designed by the communications device itself, and only the second part needs to be obtained from another communications device or protocol.
- the communications device obtains the encoder of the first model and then trains it based on the third data set to obtain the decoder of the first model, thereby obtaining the complete first model.
- Model structure such as the layers, modules, units, etc.
- Model complexity such as floating-point operations per second (FLOPs), floating-point operations (FLOPs), and the number of operations in the model;
- Model size such as model storage size
- At least two of the first models or the second models or partial models of the second models have at least one of the following in common:
- the same parameters of at least two first models or second models or partial models of the second models may be predefined by the protocol.
- Different first models can be obtained by training the same AI function or AI feature for wireless communication.
- the specific model parameters of different first models may differ, and the input-to-output mapping relationship may also differ.
- the input-to-output mapping relationship of multiple first models can be aligned. Applying the second model or a partial model of the second model or the first data set to a wireless communication network can improve the performance of the wireless communication network.
- the second model or a portion of the second model has the same input format or output format as at least two of the first models.
- the input of the first model is 16-bit data
- the input of the second model is also 16-bit data; or if the input of the first model is data for a single layer of 32 antennas and 13 subbands, which is a 32*13 complex matrix or a 64*13 real matrix (half of the 64 is the real part of the 32 antennas, and the other half is the imaginary part of the 32 antennas), the input of the second model is also data for a single layer of 32 antennas and 13 subbands, which is a 32*13 complex matrix or a 64*13 real matrix.
- the input format can be understood as the input data format or input data description form
- the output format can be understood as the output data format or output data description form.
- the second model or a partial model of the second model or the first data set can be used for at least one of the following:
- Channel state information may include channel-related information, channel matrix-related information, channel characteristic information, channel matrix characteristic information, precoding matrix indicator (PMI), rank indicator (RI), CSI-RS resource indicator (CRI), channel quality indicator (CQI), layer indicator (LI), etc.; for frequency division duplexing (FDD) system, according to the reciprocity of uplink and downlink parts, network side equipment, such as base station, can use the second model and uplink channel to obtain angle information and delay information, and notify the terminal of the angle information and delay information through CSI-RS precoding or direct indication.
- the terminal reports according to the indication of the base station or selects and reports within the indication range of the base station, thereby reducing the calculation amount of the terminal and the overhead of channel state information reporting;
- Channel or source coding may include channel coding, channel decoding, source coding, source decoding, joint source-channel coding, joint source-channel decoding, etc.
- Interference suppression can include suppression of intra-cell interference, inter-cell interference, out-of-band interference, and intermodulation interference;
- Positioning can be understood as using a second model and a reference signal, such as SRS, to estimate the specific location of the terminal, such as horizontal or vertical position, or to estimate the terminal's possible future trajectory, or to obtain information to assist in position estimation or trajectory estimation, such as time of arrival (TOA), line-of-sight or non-line-of-sight, or reference signal time difference (RSTD);
- TOA time of arrival
- RSTD reference signal time difference
- High-level services or parameters may include throughput, required packet size, service requirements, mobile speed, noise information, etc.
- Control signaling may include power control related signaling, beam management related signaling, etc.
- the first data set can be used to train the second model or part of the second model, or to characterize applicable models or functions or characteristics, or to define performance indicators, or to define performance requirements.
- the communication device uses at least two first models to obtain a first data set.
- the first data set can then be used for model training to obtain a second model, or a partial model of the second model.
- the second model or the partial model of the second model can then be applied to the wireless communication network to obtain better gain.
- a first dataset may be used to represent a model, function, or feature applicable to or working properly on the first dataset.
- the model, function, or feature applicable to or working properly on the first dataset may be one, multiple, or unlimited in number.
- the applicable model, function, or feature may be associated using a dataset ID, a data categorization ID, data-related information, dataset-related information, or dataset-associated signaling.
- a performance indicator (KPI) or a performance requirement (performance requirement) may be defined by the first data set.
- the communication device may use at least two first models to obtain a second model or a partial model of the second model or a first data set, which may include the following steps:
- Step 1 The communication device obtains a second data set
- Step 2 The communication device obtains at least one model input based on at least one data in the second data set;
- Step 3 The communication device inputs at least one model input into each first model to obtain a model output or a model intermediate quantity corresponding to each model input;
- Step 4 The communication device obtains the second model or a partial model of the second model or the first data set based on each model input, or the model output or model intermediate quantity corresponding to each model input.
- the communication device may obtain the second data set in at least one of the following ways:
- the protocol is predefined
- the second data set may include one or more data.
- the communication device may obtain at least one model input based on at least one data in the second data set.
- the communication device can input model input B into the first model 1 and the first model 2 respectively, and obtain the model output B1’ of the first model 1 corresponding to the model input B, or obtain the model intermediate quantity B1” of the first model 1 corresponding to the model input B, and obtain the model output B2’ of the first model 2 corresponding to the model input B, or obtain the model intermediate quantity B2” of the first model 2 corresponding to the model input B.
- mapping relationship between each model input and model output or model intermediate quantity is as follows:
- Model input A model output A1’—model intermediate quantity A1”
- Model input A model output A2’—model intermediate quantity A2”
- Model input B model output B1’—model intermediate quantity B1”
- Model input B model output B2’—model intermediate quantity B2”
- Model input C model output C1’—model intermediate quantity C1”
- Model input C model output C2’—model intermediate quantity C2”.
- the communication device can obtain the second model or a partial model of the second model or the first data set based on each model input, or the model output or model intermediate quantity corresponding to each model input.
- the communication device inputs the model input into each first model to obtain the model output or model intermediate quantity corresponding to each model input. Based on each model input, or the model output or model intermediate quantity corresponding to each model input, the second model or a partial model of the second model or the first data set can be accurately obtained.
- the communication device may obtain the second model or a partial model of the second model or the first data set based on each model input, or a model output or a model intermediate quantity corresponding to each model input, and may include the following steps:
- the first step the communication device determines the model final output or the model final intermediate quantity corresponding to each model input based on each model input, or the model output or the model intermediate quantity corresponding to each model input;
- the second step the communication device uses at least one model input, or the model final output or model final intermediate quantity corresponding to each model input to obtain the second model or a partial model of the second model or the first data set.
- the communication device after the communication device inputs at least one model input into each first model respectively, it can obtain the model output or model intermediate quantity corresponding to each model input.
- Each model input corresponds to a different first model and has a corresponding model output or model intermediate quantity.
- the communication device can determine the final model output or final model intermediate quantity corresponding to the current model input based on the current model input, or the model output or model intermediate quantity corresponding to the current model input.
- the final model output corresponding to the current model input can be understood as a fusion output of multiple first models corresponding to the current model input
- the final model intermediate quantity corresponding to the current model input can be understood as a fusion intermediate quantity of multiple first models corresponding to the current model input.
- the current model input refers to the model input targeted by the current operation.
- the communication device can obtain the second model or a partial model of the second model or the first data set by using at least one model input, or the final model output or the final model intermediate quantity corresponding to each model input.
- the communication device can use at least one model input, or the final model output or final model intermediate quantity corresponding to each model input, to perform model training to obtain a second model or a partial model of the second model.
- the input of the second model is the model input
- the output of the second model is the final output of the model.
- the output of the second model can be obtained by inputting the model input into the second model.
- the output of the second model is the final output of the model corresponding to the model input.
- the model input includes model input A, model input B, and model input C.
- the final model output corresponding to model input A is A0’
- the final model output corresponding to model input B is B0’
- the final model output corresponding to model input C is C0’.
- the input of the second model includes model input A, model input B, and model input C.
- the output of the second model corresponding to model input A is the final model output A0’
- the output of the second model corresponding to model input B is the final model output B0’
- the output of the second model corresponding to model input C is the final model output C0’.
- model input as the input of the second model and the model's final output as the output of the second model helps improve the training efficiency of the second model and improve the accuracy of the second model.
- the input of the encoder of the second model is the model input
- the output of the encoder of the second model or the input of the decoder of the second model is the final intermediate quantity of the model
- the output of the decoder of the second model is the final output of the model.
- Using the model input as the input of the encoder of the second model, using the model's final intermediate quantity as the output of the encoder of the second model or the input of the decoder of the second model, and using the model's final output as the output of the decoder of the second model can help improve the training efficiency of the second model and improve the accuracy of the second model.
- the first model is a unilateral model, and the first data set also corresponds to the unilateral model.
- the first data set may include one or more data, each of which may include a model input and a final model output corresponding to the model input.
- the first model is a bilateral model
- the first data set also corresponds to the bilateral model.
- the first data set may include one or more data, each of which may include a model input, a model final output corresponding to the model input, and a model final intermediate quantity.
- the first model is a bilateral model
- the first data set corresponds to an encoder of the bilateral model.
- the first data set may include one or more data, each of which may include a model input and a final intermediate quantity of the model corresponding to the model input.
- the first model output in the model output set corresponding to the current model input has the highest similarity to the output label corresponding to the current model input;
- the model output set includes: model outputs obtained after the current model input is input into each first model respectively.
- the communication device after the communication device obtains at least one model input, it can input at least one model input into each first model respectively to obtain a model output or model intermediate quantity corresponding to each model input.
- One model input corresponds to multiple model outputs or multiple model intermediate quantities.
- the final model output corresponding to the current model input can be determined based on at least one of the following:
- the current model input may be determined as the final model output corresponding to the current model input;
- the model output set includes: the model outputs obtained after the current model input is input into each first model respectively.
- the model output set includes multiple model outputs, and the model outputs included in the model output set corresponding to the current model input may be averaged, such as by performing linear averaging, geometric averaging, harmonic averaging, square averaging, weighted averaging, minimum maximization, maximum maximization, combinations of the above or simple changes, to obtain a first result, and the first result may be determined as the final output of the model corresponding to the current model input; optionally, the first result may also be a result obtained by performing other mathematical operations on the model output set corresponding to the current model input, such as a result obtained by performing a summation operation, normalization operation, or eigenvalue decomposition operation on the model output set, which are not listed here one by one;
- the first model output in the model output set corresponding to the current model input has the highest similarity to the output label corresponding to the current model input;
- the model output set corresponding to the current model input includes multiple model outputs, and the similarity between each model output and the output label corresponding to the current model input can be determined separately to obtain the first model output corresponding to the highest similarity, and the first model output can be determined as the final model output corresponding to the current model input.
- the first model output has the highest similarity to the output label corresponding to the current model input, which can be understood as the first model output being closest to the output label corresponding to the current model input.
- the highest similarity can be understood as the highest correlation, cosine similarity, square of cosine similarity, etc., or as the smallest gap, difference, mean square error (NMSE), distance, Euclidean distance, etc., or as the smallest absolute value, amplitude, power, etc. of the difference.
- NMSE mean square error
- the current model input refers to the model input targeted by the current operation.
- the communication device can determine the final model output corresponding to each model input based on one of the above items, or it can determine the final model output corresponding to each model input based on a combination of the above contents. For example, after performing a mathematical operation on the current model input and the first result, the obtained result is determined as the final model output corresponding to the current model input.
- one of the above contents may be predefined through a protocol, and the communication device determines the final model output corresponding to each model input according to the predefined content of the protocol.
- the above-mentioned multiple contents can be predefined through a protocol, and the communication device selects one item according to the protocol predefinition, and determines the final model output corresponding to each model input based on the selected content.
- the above-mentioned multiple contents and combination methods can be predefined through a protocol.
- the communication device combines at least two contents according to the predefined protocol to determine the final model output corresponding to each model input.
- the communication device can more accurately determine the final model output corresponding to each model input.
- the final intermediate quantity of the model corresponding to the current model input is determined based on at least one of the following:
- the set of model intermediate quantities comprising: model intermediate quantities obtained by inputting the current model input into each of the first models;
- the first model intermediate quantity in the model intermediate quantity set, the first model intermediate quantity and the second model output are obtained using the same first model, the second model output is a model output in the model output set corresponding to the current model input, and the second model output has the highest similarity with the output label corresponding to the current model input.
- the model output set includes: the model output obtained after the current model input is input into each first model respectively.
- the communication device after the communication device obtains at least one model input, it can input at least one model input into each first model respectively to obtain a model output or model intermediate quantity corresponding to each model input.
- One model input corresponds to multiple model outputs or multiple model intermediate quantities.
- the final intermediate quantity of the model corresponding to the current model input can be determined based on at least one of the following:
- the model intermediate quantity set includes: the model intermediate quantities obtained after the current model input is input into each first model respectively.
- a third result obtained by operating on the model intermediate quantity set and the model output set corresponding to the current model input may be a result obtained by averaging the model intermediate quantity set and the model output set corresponding to the current model input, or a result obtained by performing other mathematical operations on the model intermediate quantity set and the model output set corresponding to the current model input, such as a result obtained by performing a merging operation, a summing operation, a normalizing operation, or an eigenvalue decomposition operation on the model intermediate quantity set and the model output set corresponding to the current model input, which are not listed here one by one.
- the model output set corresponding to the current model input includes multiple model outputs. The similarity between each model output and the output label corresponding to the current model input can be determined separately, and the second model output corresponding to the highest similarity can be obtained.
- the first model intermediate quantity corresponding to this second model output can be determined as the final model intermediate quantity corresponding to the current model input.
- the first model intermediate quantity and the second model output are obtained using the same first model.
- the second model output has the highest similarity with the output label corresponding to the current model input, which can be understood as the second model output being closest to the output label corresponding to the current model input.
- the highest similarity can be understood as the highest correlation, cosine similarity, or squared cosine similarity, or the smallest gap, difference, normalized mean square error (NMSE), distance, Euclidean distance, or the smallest absolute value, amplitude, or power of the difference.
- NMSE normalized mean square error
- the current model input refers to the model input targeted by the current operation.
- the communication device can determine the model final intermediate quantity corresponding to each model input based on the above item, or determine the model final intermediate quantity corresponding to each model input based on a combination of the above contents. For example, after performing a mathematical operation on the second result and the first model intermediate quantity, the obtained result is determined as the model final intermediate quantity corresponding to the current model input.
- one of the above contents may be predefined through a protocol, and the communication device determines the final intermediate quantity of the model corresponding to each model input according to the content predefined by the protocol.
- the above-mentioned multiple contents can be predefined through a protocol, and the communication device selects one item according to the protocol predefinition, and determines the final intermediate quantity of the model corresponding to each model input based on the selected content.
- the communication device may use at least two first models to obtain a second model or a partial model of the second model, which may include the following steps:
- the communication device may combine the at least two first models into a new large model, which serves as the second model or a partial model of the second model.
- the communication device may combine the at least two first models into the second model or a partial model of the second model according to a model fusion method predefined in the protocol. For example, the model parameters of the at least two first models may be averaged.
- Example 1 The first model and the second model are unilateral models
- a data of the first data set generated includes the model input A, or the model final output A0’.
- the model input B is input into the encoder (Encoder) of the first model 3 and the encoder of the first model 4 respectively, and two model intermediate quantities (inter-data) are obtained, including the model intermediate quantity B3" of the first model 3, and the model intermediate quantity B4" of the first model 4.
- These two model intermediate quantities are input into the corresponding decoder (Decoder) of the first model 3 and the decoder of the first model 4 respectively, and two model outputs are obtained, including the model output B1' of the first model 3, and the model output B4' of the first model 4.
- the model output B3' and the model intermediate quantity B3" are obtained using the same first model, that is, the first model 3, and the model output B4' and the model intermediate quantity B4" are obtained using the same first model, that is, the first model 4.
- Solution 1 The final output of the model B0’ is the model input B;
- the second model is used for compression or recovery of channel information/CSI, such as the encoder is used to compress channel information/CSI into feedback information, and the decoder is used to restore feedback information into channel information/CSI, or the second model is used for encoding or decoding, such as source encoding and decoding, channel encoding and decoding, source-channel joint encoding and decoding, etc. (the encoder is used for encoding, and the decoder is used for decoding), then this scheme can be used to determine the final output of the model.
- Solution 2 The final output of the model, B0’, is the average of the outputs of the above two models
- the averaging methods include linear averaging, geometric averaging, harmonic averaging, square averaging, weighted averaging, minimum maximization, maximum maximization, and combinations or simple variations of at least two of the above averaging methods.
- Solution 3 The final model output B0’ is the model output closest to the output label among the two above models;
- M is closest to N, which can be understood as: the correlation, cosine similarity, and square of cosine similarity between M and N are the highest, or the gap, difference, NMSE, distance, and Euclidean distance between M and N are the smallest, or the absolute value, amplitude, and power of (M-N) are the smallest;
- the output label is the model input.
- the final model output, B0' is the model output that is closest to the model input.
- the final intermediate quantity B0" of the model is the intermediate quantity corresponding to the model output closest to the output label among the two model outputs mentioned above.
- Scheme 3 for the final output of the model and Scheme 2 for the final intermediate quantity of the model can be combined.
- Model fusion can be performed using the model input, or the final output of the model, or the final intermediate quantity of the model to obtain a second model or generate a first data set.
- the input of the encoder of the second model is the model input B
- the output of the encoder or the input of the decoder is the final intermediate quantity of the model B0
- the output of the decoder is the final output of the model B0'
- a data of the generated first data set includes the model input B, or the final intermediate quantity of the model B0", or the final output of the model B0'.
- the final intermediate quantity of the model corresponding to the current model input is determined based on at least one of the following:
- the model output set includes: model outputs obtained after the current model input is input into each first model respectively.
- the input of the second model is the model input
- the output of the second model is the final model output
- the input of the encoder of the second model is the model input
- the output of the encoder of the second model or the input of the decoder of the second model is the final intermediate quantity of the model
- the output of the decoder of the second model is the final output of the model.
- the protocol is predefined
- At least two first models are combined into a second model or a partial model of a second model.
- At least one of the following items of at least two first models or second models or partial models of the second models is predefined by the protocol:
- At least two first models or second models or partial models of the second models have at least one of the following in common:
- Model basic architecture model structure; activation function; model complexity; maximum model complexity; minimum model complexity; model size; maximum model size and minimum model size.
- the second model or a partial model of the second model or the first data set is used for at least one of the following:
- the first obtaining module 710 is specifically configured to:
- At least two first models are obtained by using the partial models of the at least two first models and the third data set.
- the model processing device 700 provided in the embodiment of the present application can implement each process implemented by the method embodiment shown in Figure 4 and achieve the same technical effect. To avoid repetition, it will not be described here.
- an embodiment of the present application further provides a communication device 800, including a processor 801 and a memory 802.
- the memory 802 stores a program or instruction that can be run on the processor 801.
- the program or instruction is executed by the processor 801 to implement the various steps of the above-mentioned model processing method embodiment and can achieve the same technical effect.
- the communication device 800 is a network-side device
- the program or instruction is executed by the processor 801 to implement the various steps of the above-mentioned model processing method embodiment and can achieve the same technical effect. To avoid repetition, it will not be repeated here.
- the present application also provides a terminal including a processor and a communication interface, wherein the communication interface and the processor are coupled, and the processor is configured to execute a program or instruction to implement the steps in the method embodiment shown in FIG4 .
- This terminal embodiment corresponds to the above-described method embodiment, and each implementation process and implementation method of the above-described method embodiment is applicable to this terminal embodiment and can achieve the same technical effects.
- FIG9 is a schematic diagram of the hardware structure of a terminal implementing an embodiment of the present application.
- the terminal 900 includes but is not limited to: a radio frequency unit 901, a network module 902, an audio output unit 903, an input unit 904, a sensor 905, a display unit 906, a user input unit 907, an interface unit 908, a memory 909 and at least some of the components of the processor 910.
- the terminal 900 may also include a power supply (such as a battery) to power various components.
- the power supply may be logically connected to the processor 910 via a power management system, thereby enabling the power management system to manage charging, discharging, and power consumption.
- the terminal structure shown in FIG9 does not limit the terminal.
- the terminal may include more or fewer components than shown, or may combine certain components, or have different component arrangements, which will not be described in detail here.
- the input unit 904 may include a graphics processing unit (GPU) 9041 and a microphone 9042, and the graphics processor 9041 processes the image data of the static picture or video obtained by the image capture device (such as a camera) in the video capture mode or the image capture mode.
- the display unit 906 may include a display panel 9061, and the display panel 9061 may be configured in the form of a liquid crystal display, an organic light emitting diode, etc.
- the user input unit 907 includes a touch panel 9071 and at least one of the other input devices 9072.
- the touch panel 9071 is also called a touch screen.
- the touch panel 9071 may include two parts: a touch detection device and a touch controller.
- Other input devices 9072 may include but are not limited to a physical keyboard, function keys (such as volume control keys, switch keys, etc.), a trackball, a mouse, and a joystick, which will not be repeated here.
- the RF unit 901 may transmit the data to the processor 910 for processing. Furthermore, the RF unit 901 may send uplink data to the network-side device.
- the RF unit 901 includes, but is not limited to, an antenna, an amplifier, a transceiver, a coupler, a low-noise amplifier, a duplexer, and the like.
- the volatile memory may be random access memory (RAM), static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate synchronous DRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), and direct RAM (DRRAM).
- RAM random access memory
- SRAM static RAM
- DRAM dynamic RAM
- SDRAM synchronous DRAM
- DDRSDRAM double data rate synchronous DRAM
- ESDRAM enhanced SDRAM
- SLDRAM synchronous link DRAM
- DRRAM direct RAM
- the memory 909 in the embodiment of the present application includes but is not limited to these and any other suitable types of memory.
- Processor 910 may include one or more processing units.
- processor 910 integrates an application processor and a modem processor.
- the application processor primarily handles operations related to the operating system, user interface, and application programs, while the modem processor primarily processes wireless communication signals, such as a baseband processor. It is understood that the modem processor may not be integrated into processor 910.
- the processor 910 is configured to obtain at least two first models
- a second model or a partial model of the second model or a first data set is obtained.
- the present application also provides a network-side device, including a processor and a communication interface, wherein the communication interface is coupled to the processor, and the processor is configured to execute a program or instruction to implement the steps of the method embodiment shown in FIG4 .
- This network-side device embodiment corresponds to the above-described method embodiment, and each implementation process and implementation method of the above-described method embodiment are applicable to this network-side device embodiment and can achieve the same technical effects.
- an embodiment of the present application also provides a network-side device.
- the network-side device 1000 includes: an antenna 1001, a radio frequency device 1002, a baseband device 1003, a processor 1004, and a memory 1005.
- Antenna 1001 is connected to radio frequency device 1002.
- radio frequency device 1002 receives information via antenna 1001 and sends the received information to baseband device 1003 for processing.
- baseband device 1003 processes the information to be transmitted and sends it to radio frequency device 1002.
- Radio frequency device 1002 processes the received information and sends it through antenna 1001.
- the baseband device 1003 may include, for example, at least one baseband board, on which multiple chips are arranged, as shown in Figure 10, one of which is, for example, a baseband processor, which is connected to the memory 1005 through a bus interface to call the program in the memory 1005 and execute the network side device operations shown in the above method embodiment.
- the network side device may also include a network interface 1006, which is, for example, a Common Public Radio Interface (CPRI).
- CPRI Common Public Radio Interface
- the network side device 1000 of the embodiment of the present application also includes: instructions or programs stored in the memory 1005 and executable on the processor 1004.
- the processor 1004 calls the instructions or programs in the memory 1005 to execute the method executed by each module in the execution model processing device 700 and achieves the same technical effect. To avoid repetition, it will not be described here.
- an embodiment of the present application further provides a network-side device.
- the network-side device 1100 includes a processor 1101, a network interface 1102, and a memory 1103.
- the network interface 1102 is, for example, a common public radio interface (CPRI).
- CPRI common public radio interface
- An embodiment of the present application also provides a readable storage medium, on which a program or instruction is stored.
- a program or instruction is stored.
- the various processes of the above-mentioned method embodiment are implemented and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.
- the processor is the processor in the terminal described in the above embodiment.
- the readable storage medium includes a computer-readable storage medium, such as a computer read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
- ROM computer read-only memory
- RAM random access memory
- magnetic disk such as a hard disk, a hard disk, or a magnetic disk.
- optical disk such as a hard disk, a hard disk, or an optical disk.
- the readable storage medium may be a non-transitory readable storage medium.
- the chip mentioned in the embodiments of the present application can also be called a system-level chip, a system chip, a chip system or a system-on-chip chip, etc.
- An embodiment of the present application further provides a computer program/program product, which is stored in a storage medium.
- the computer program/program product is executed by at least one processor to implement the various processes of the above-mentioned method embodiment and can achieve the same technical effect. To avoid repetition, it will not be repeated here.
- An embodiment of the present application also provides a wireless communication system, including: a terminal and a network side device, wherein the terminal can be used to execute the steps of the model processing method described above, and the network side device can be used to execute the steps of the model processing method described above.
Landscapes
- Engineering & Computer Science (AREA)
- Physics & Mathematics (AREA)
- Theoretical Computer Science (AREA)
- Signal Processing (AREA)
- Computer Networks & Wireless Communication (AREA)
- Evolutionary Computation (AREA)
- Artificial Intelligence (AREA)
- General Health & Medical Sciences (AREA)
- General Physics & Mathematics (AREA)
- Computational Linguistics (AREA)
- Biophysics (AREA)
- Molecular Biology (AREA)
- Computing Systems (AREA)
- General Engineering & Computer Science (AREA)
- Data Mining & Analysis (AREA)
- Mathematical Physics (AREA)
- Software Systems (AREA)
- Biomedical Technology (AREA)
- Electromagnetism (AREA)
- Life Sciences & Earth Sciences (AREA)
- Health & Medical Sciences (AREA)
- Quality & Reliability (AREA)
- Mobile Radio Communication Systems (AREA)
Abstract
本申请公开了一种模型处理方法、装置、通信设备及存储介质,属于通信技术领域,本申请实施例的一种模型处理方法包括:通信设备获得至少两个第一模型;通信设备利用至少两个第一模型,得到第二模型或第二模型的部分模型或第一数据集。
Description
相关申请的交叉引用
本申请要求在2024年03月27日提交中国专利局、申请号为202410361159.X、名称为“模型处理方法、装置、通信设备及存储介质”的中国专利申请的优先权,其全部内容通过引用结合在本申请中。
本申请属于通信技术领域,具体涉及一种模型处理方法、装置、通信设备及存储介质。
人工智能(Artificial Intelligence,AI)是研究、开发用于模拟、延伸和扩展人的智能的理论、方法、技术及应用系统的一门新的技术科学。AI技术在各个领域得到了广泛应用,并发挥了重要作用。比如,在无线通信网络中,融入AI技术,有助于提升吞吐量、时延、用户容量等技术指标。
对于无线通信网络的同一个AI功能(functionality)或AI特性(feature),通过不同方法或不同数据集分别进行训练可以得到不同的AI模型。不同AI模型的模型具体参数可能不同,输入到输出的映射关系也可能不同。目前,是从多个AI模型中选择一个AI模型进行使用。而单个模型性能有限,只能适应部分场景或环境,无法获得较好增益。
本申请实施例提供一种模型处理方法、装置、通信设备及存储介质,能够利用多个AI模型获得更好增益,有助于提升无线通信网络的性能。
第一方面,提供了一种模型处理方法,包括:
通信设备获得至少两个第一模型;
所述通信设备利用所述至少两个第一模型,得到第二模型或第二模型的部分模型或第一数据集。
第二方面,提供了一种模型处理装置,包括:
第一获得模块,用于获得至少两个第一模型;
第二获得模块,用于利用所述至少两个第一模型,得到第二模型或第二模型的部分模型或第一数据集。
第三方面,提供了一种通信设备,该通信设备包括处理器和存储器,所述存储器存储可在所述处理器上运行的程序或指令,所述程序或指令被所述处理器执行时实现如第一方面所述方法的步骤。
第四方面,提供了一种通信设备,包括处理器及通信接口,其中,所述通信接口与所述处理器耦合,所述处理器用于运行程序或指令,实现如第一方面所述方法的步骤。
第五方面,提供了一种终端,该终端包括处理器和存储器,所述存储器存储可在所述处理器上运行的程序或指令,所述程序或指令被所述处理器执行时实现如第一方面所述方法的步骤。
第六方面,提供了一种终端,包括处理器及通信接口,其中,所述通信接口与所述处理器耦合,所述处理器用于运行程序或指令,实现如第一方面所述方法的步骤。
第七方面,提供了一种网络侧设备,该网络侧设备包括处理器和存储器,所述存储器存储可在所述处理器上运行的程序或指令,所述程序或指令被所述处理器执行时实现如第一方面所述方法的步骤。
第八方面,提供了一种网络侧设备,包括处理器及通信接口,其中,所述通信接口与所述处理器耦合,所述处理器用于运行程序或指令,实现如第一方面所述方法的步骤。
第九方面,提供了一种可读存储介质,所述可读存储介质上存储程序或指令,所述程序或指令被处理器执行时实现如第一方面所述方法的步骤。
第十方面,提供了一种无线通信系统,包括:终端及网络侧设备,所述终端可用于执行如第一方面所述方法的步骤,或所述网络侧设备可用于执行如第一方面所述方法的步骤。
第十一方面,提供了一种芯片,所述芯片包括处理器和通信接口,所述通信接口和所述处理器耦合,所述处理器用于运行程序或指令,实现如第一方面所述方法的步骤。
第十二方面,提供了一种计算机程序/程序产品,所述计算机程序/程序产品被存储在存储介质中,所述计算机程序/程序产品被至少一个处理器执行以实现如第一方面所述方法的步骤。
在本申请实施例中,通信设备获得至少两个第一模型后,利用至少两个第一模型,得到第二模型或第一数据集,将多个第一模型进行融合得到第二模型或第二模型的部分模型或第一数据集,拉齐了多个第一模型输入到输出的映射关系,将第二模型或第二模型的部分模型或第一数据集应用到无线通信网络中,能够获得更好增益,有助于提升无线通信网络的性能。
图1为本申请实施例可应用的一种无线通信系统的框图;
图2为相关技术中一种神经网络的示意图;
图3为相关技术中一种神经元的示意图;
图4为本申请实施例中一种模型处理方法的实施流程图;
图5为本申请实施例中一种单边模型处理过程的示意图;
图6为本申请实施例中一种双边模型处理过程的示意图;
图7为本申请实施例中一种模型处理装置的结构示意图;
图8为本申请实施例中一种通信设备的结构示意图;
图9为本申请实施例中一种终端的结构示意图;
图10为本申请实施例中一种网络侧设备的结构示意图;
图11为本申请实施例中另一种网络侧设备的结构示意图。
下面将结合本申请实施例中的附图,对本申请实施例中的技术方案进行清楚描述,显然,所描述的实施例是本申请一部分实施例,而不是全部的实施例。基于本申请中的实施例,本领域普通技术人员所获得的所有其他实施例,都属于本申请保护的范围。
本申请的术语“第一”、“第二”等是用于区别类似的对象,而不用于描述特定的顺序或先后次序。应该理解这样使用的术语在适当情况下可以互换,以便本申请的实施例能够以除了在这里图示或描述的那些以外的顺序实施,且“第一”、“第二”所区别的对象通常为一类,并不限定对象的个数,例如第一对象可以是一个,也可以是多个。此外,本申请中的“或”表示所连接对象的至少其中之一。例如“A或B”涵盖三种方案,即,方案一:包括A且不包括B;方案二:包括B且不包括A;方案三:既包括A又包括B。字符“/”一般表示前后关联对象是一种“或”的关系。
本申请的术语“指示”既可以是一个直接的指示(或者说显式的指示),也可以是一个间接的指示(或者说隐含的指示)。其中,直接的指示可以理解为,发送方在发送的指示中明确告知了接收方具体的信息、需要执行的操作或请求结果等内容;间接的指示可以理解为,接收方根据发送方发送的指示确定对应的信息,或者进行判断并根据判断结果确定需要执行的操作或请求结果等。
值得指出的是,本申请实施例所描述的技术不限于长期演进型(Long Term Evolution,LTE)/LTE的演进(LTE-Advanced,LTE-A)系统,还可用于其他无线通信系统,诸如码分多址(Code Division Multiple Access,CDMA)、时分多址(Time Division Multiple Access,TDMA)、频分多址(Frequency Division Multiple Access,FDMA)、正交频分多址(Orthogonal Frequency Division Multiple Access,OFDMA)、单载波频分多址(Single-carrier Frequency-Division Multiple Access,SC-FDMA)或其他系统。本申请实施例中的术语“系统”和“网络”常被可互换地使用,所描述的技术既可用于以上提及的系统和无线电技术,也可用于其他系统和无线电技术。以下描述出于示例目的描述了新空口(New Radio,NR)系统,并且在以下大部分描述中使用NR术语,但是这些技术也可应用于NR系统以外的系统,如第6代(6th Generation,6G)通信系统。
图1示出本申请实施例可应用的一种无线通信系统的框图。无线通信系统包括终端11和网络侧设备12。
其中,终端11可以是手机、平板电脑(Tablet Personal Computer)、膝上型电脑(Laptop Computer)、笔记本电脑、个人数字助理(Personal Digital Assistant,PDA)、掌上电脑、上网本、超级移动个人计算机(Ultra-mobile Personal Computer,UMPC)、移动上网装置(Mobile Internet Device,MID)、增强现实(Augmented Reality,AR)、虚拟现实(Virtual Reality,VR)设备、机器人、可穿戴式设备(Wearable Device)、飞行器(flight vehicle)、车载用户设备(Vehicle User Equipment,VUE)、船载设备、行人用户设备(Pedestrian User Equipment,PUE)、智能家居(具有无线通信功能的家居设备,如冰箱、电视、洗衣机或者家具等)、游戏机、个人计算机(Personal Computer,PC)、柜员机或者自助机等终端侧设备。可穿戴式设备包括:智能手表、智能手环、智能耳机、智能眼镜、智能首饰(智能手镯、智能手链、智能戒指、智能项链、智能脚镯、智能脚链等)、智能腕带、智能服装等。其中,车载设备也可以称为车载终端、车载控制器、车载模块、车载部件、车载芯片或车载单元等。需要说明的是,在本申请实施例并不限定终端11的具体类型。
网络侧设备12可以包括接入网设备或核心网设备。其中,接入网设备也可以称为无线接入网(Radio Access Network,RAN)设备、无线接入网功能或无线接入网单元。接入网设备可以包括基站、无线局域网(Wireless Local Area Network,WLAN)接入点(Access Point,AP)或无线保真(Wireless Fidelity,WiFi)节点等。其中,基站可被称为节点B(Node B,NB)、演进节点B(Evolved Node B,eNB)、下一代节点B(the next generation Node B,gNB)、新空口节点B(New Radio Node B,NR Node B)、接入点、中继基站(Relay Base Station,RBS)、服务基站(Serving Base Station,SBS)、基收发机站(Base Transceiver Station,BTS)、无线电基站、无线电收发机、基本服务集(Basic Service Set,BSS)、扩展服务集(Extended Service Set,ESS)、家用节点B(home Node B,HNB)、家用演进型节点B(home evolved Node B)、发送接收点(Transmission Reception Point,TRP)或所属领域中其他某个合适的术语,只要达到相同的技术效果,所述基站不限于特定技术词汇,需要说明的是,在本申请实施例中仅以NR系统中的基站为例进行介绍,并不限定基站的具体类型。
核心网设备可以包含但不限于如下至少一项:核心网节点、核心网功能、移动管理实体(Mobility Management Entity,MME)、接入移动管理功能(Access and Mobility Management Function,AMF)、会话管理功能(Session Management Function,SMF)、用户平面功能(User Plane Function,UPF)、策略控制功能(Policy Control Function,PCF)、策略与计费规则功能单元(Policy and Charging Rules Function,PCRF)、边缘应用服务发现功能(Edge Application Server Discovery Function,EASDF)、统一数据管理(Unified Data Management,UDM)、统一数据仓储(Unified Data Repository,UDR)、归属用户服务器(Home Subscriber Server,HSS)、集中式网络配置(Centralized network configuration,CNC)、网络存储功能(Network Repository Function,NRF)、网络开放功能(Network Exposure Function,NEF)、本地NEF(Local NEF,或L-NEF)、绑定支持功能(Binding Support Function,BSF)、应用功能(Application Function,AF)、位置管理功能(Location Management Function,LMF)、网关的移动位置中心(Gateway Mobile Location Centre,GMLC)、网络数据分析功能(Network Data Analytics Function,NWDAF)等。需要说明的是,在本申请实施例中仅以NR系统中的核心网设备为例进行介绍,并不限定核心网设备的具体类型。
为方便理解,先对本申请实施例所涉及的人工智能(AI)相关技术及概念进行介绍。
AI技术在通信、医疗、教育等各个领域均有广泛应用。AI模型有多种实现方式,例如神经网络、决策树、支持向量机、贝叶斯分类器等。本申请实施例主要以AI模型为神经网络为例进行说明,但这并不构成对AI模型的具体类型的限制。
图2所示为一种神经网络的示意图,该神经网络包括输入层(X1、X2、……、Xn)、隐层、输出层(Y)。其中,神经网络由神经元组成,神经元的示意图如图3所示:
z=a1w1+…+akwk+…+aKwK+b;
z=a1w1+…+akwk+…+aKwK+b;
其中,a1、a2、…、ak、…、aK为输入,w为权值(weight)(乘性系数),b为偏置(bias)(加性系数),σ(.)为激活函数(activation function)。常见的激活函数包括S型函数(Sigmoid)、双曲正切函数(tanh)、修正线性单元(Rectified Linear Unit,ReLU)(或称为线性整流函数)等。
神经网络的参数通过梯度优化算法进行优化。梯度优化算法是一类最小化或者最大化目标函数(或称为损失函数)的算法。目标函数往往是模型参数和数据的数学组合。例如,给定数据X和其对应的标签Y,可以构建一个神经网络模型f(.),在得到神经网络模型后,可以根据输入x得到预测输出f(x),并且可以计算出预测值和真实值之间的差距(f(x)-Y),这个就是损失函数。目的是找到合适的W、b使得上述的损失函数的值达到最小。损失值越小,则说明神经网络模型的预测结果越接近于真实情况。
目前常见的优化算法,基本都是基于误差反向传播(error Back Propagation,BP)算法。BP算法的基本思想是,学习过程由信号的正向传播与误差的反向传播两个过程组成。正向传播时,输入样本从输入层传入,经各隐层逐层处理后,传向输出层。若输出层的实际输出与期望的输出不符,则转入误差的反向传播阶段。误差反传是将输出误差以某种形式通过隐层向输入层逐层反传,并将误差分摊给各层的所有单元,从而获得各层单元的误差信号,此误差信号即作为修正各单元权值的依据。这种信号正向传播与误差反向传播的各层权值调整过程,是周而复始地进行的。权值不断调整的过程,也就是网络的学习训练过程。此过程一直进行到网络输出的误差减小到可接受的程度,或进行到预先设定的学习次数为止。
常见的优化算法有梯度下降(Gradient Descent)、随机梯度下降(Stochastic Gradient Descent,SGD)、小批量梯度下降(mini-batch gradient descent)、动量法(Momentum)、Nesterov(发明者的名字,具体为带动量的随机梯度下降)、自适应梯度下降(ADAptive GRADient descent,Adagrad)、自适应学习率调整(Adadelta)、均方根误差降速(root mean square prop,RMSprop)、自适应动量估计(Adaptive Moment Estimation,Adam)等。
这些优化算法在误差反向传播时,都是根据损失函数得到的误差/损失,对当前神经元求导数/偏导,加上学习速率、之前的梯度/导数/偏导等影响,得到梯度,将梯度传给上一层。
在本申请实施例中,模型或AI模型也可称为AI单元、AI模块、机器学习(machine learning,ML)模型、ML单元、AI结构、AI功能、AI特性、神经网络、神经网络函数、神经网络功能等,或者,模型或AI模型也可以是指能够实现与AI相关的特定的算法、公式、特性、处理流程、能力等的处理单元、处理模块,或者,或者模型或AI模型可以是针对特定数据集的处理方法、算法、功能、特性、模块或单元,或者,模型或AI模型可以是运行在图形处理器(Graphics Processing Unit,GPU)、神经网络处理器(Neural Processing Unit,NPU)、张量处理器(Tensor Processing Unit,TPU)或专用集成电路(Application-Specific Integrated Circuit,ASIC)等AI/ML相关硬件上的处理方法、算法、功能、特性、模块或单元,本申请实施例对此不做具体限定。
可选地,模型或AI模型的标识,可以理解为AI模型标识、AI模块标识、AI结构标识、AI算法标识、AI单元标识,或者AI模型关联的特定数据集的标识,或者AI/ML相关的特定场景、环境、区域、小区、信道特征、设备的标识,或者AI/ML相关的功能、特性、能力或模块的标识,本申请实施例对此不做具体限定。特性数据集可以包括模型或AI模型的输入或输出。
上面对本申请实施例涉及的相关技术和概念进行了介绍,下面结合附图,通过一些实施例及其应用场景对本申请实施例提供的模型处理方法进行详细地说明。
参见图4所示,为本申请实施例所提供的一种模型处理方法的实施流程图,该方法包括以下步骤:
S410:通信设备获得至少两个第一模型;
S420:通信设备利用至少两个第一模型,得到第二模型或第二模型的部分模型或第一数据集。
应用本申请实施例所提供的方法,通信设备获得至少两个第一模型后,利用至少两个第一模型,得到第二模型或第二模型的部分模型或第一数据集,将多个第一模型进行融合得到第二模型或第二模型的部分模型或第一数据集,拉齐了多个第一模型输入到输出的映射关系,将第二模型或第二模型的部分模型或第一数据集应用到无线通信网络中,能够获得更好增益,有助于提升无线通信网络的性能。
在一种实施例中,通信设备可以获得第一模型的部分模型,利用第一模型的部分模型和第三数据集,得到至少两个第一模型。得到的至少两个第一模型可以理解为完整的第一模型。如第一模型包括第一部分和第二部分,第一部分是通信设备自行设计,只有第二部分需要从其他通信设备或协议中获得。例如,第一模型为双边模型,通信设备获得的是第一模型的编码器,然后基于第三数据集进行训练,得到第一模型的译码器,进而得到完整的第一模型。例如,第一模型为双边模型,通信设备获得的是第一模型的译码器,然后基于第三数据集进行训练,得到第一模型的编码器,进而得到完整的第一模型。可选地,通信设备可以通过以下至少一种方式获得第三数据集:接收参考信号,得到参考信号信息;协议预定义;接收其他通信设备传输的数据。
在一种实施例中,通信设备获得的是第二模型的部分模型。例如,第二模型包括第一部分和第二部分,第一部分是通信设备自行设计,只有第二部分需要通过第一模型来得到。例如第一模型或第二模型为双边模型,通信设备利用至少两个第一模型,得到第二模型的部分模型,如第二模型的编码器或译码器。例如,如果通信设备是终端,则获得第二模型的编码器;如果通信设备是网络侧设备,则获得第二模型的译码器。
在本申请实施例中,第一模型可以是AI模型,通信设备可以是终端,如图1中的终端11,或者,通信设备可以是网络侧设备,如图1中的网络侧设备12。
可以通过协议预定义N个第一模型,或模型融合方法或数据生成方法。N大于或等于2。通信设备获得至少两个第一模型后,可以使用模型融合方法,利用至少两个第一模型得到第二模型或第二模型的部分模型,或使用数据生成方法,利用至少两个第一模型得到第一数据集。
可选地,至少两个第一模型或第二模型或第二模型的部分模型的以下至少一项是协议预定义的:
模型基本架构,如神经网络、决策树、支持向量机、贝叶斯分类器等;
模型结构,如包括的层、模块、单元等;
模型各层参数,如每层神经元或卷积核的数量;
神经元系数,如神经元上的乘性系数或加性系数;
模型复杂度,如每秒浮点运算次数(Floating-point OPerations per second,FLOPs)、浮点运算次数(Floating-point Operations,FLOP)、模型中运算次数等;
模型大小(size),如模型存储大小;
激活函数;
量化方式,其中量化方式可以包括定点或浮点,如Int4、Int8、Int16、Int32、Int64、Float16、Float32、double等。
可选地,至少两个第一模型的完整模型可以是协议预定义的,或者运行第一模型所需要的参数可以是协议预定义的。
可选地,至少两个第一模型或第二模型或第二模型的部分模型以下至少一项相同:
模型基本架构;
模型结构;
激活函数;
模型复杂度;
模型复杂度的最大值;
模型复杂度的最小值;
模型大小;
模型大小的最大值;
模型大小的最小值。
可选地,至少两个第一模型或第二模型或第二模型的部分模型的相同参数可以是协议预定义的。
不同第一模型可以是对同一个无线通信的AI功能或AI特性,分别进行训练得到的,不同第一模型的模型具体参数可能不同,输入到输出的映射关系也可能不同。利用多个第一模型,得到第二模型或第二模型的部分模型或第一数据集,可以拉齐多个第一模型输入到输出的映射关系,将第二模型或第二模型的部分模型或第一数据集应用到无线通信网络中,可以提升无线通信网络的性能。
可选地,第二模型或第二模型的部分模型与至少两个第一模型的输入格式或输出格式相同。如,第一模型的输入为16比特数据,第二模型的输入也为16比特数据;或第一模型的输入为1层32天线13子带的数据,是32*13的复数矩阵,或64*13的实数矩阵(64的一半为32天线的实部,另一半为32天线的虚部),第二模型的输入也为1层32天线13子带的数据,是32*13的复数矩阵,或64*13的实数矩阵。输入格式可以理解为输入数据格式或输入数据描述形式,输出格式可以理解为输出数据格式或输出数据描述形式。
可选地,第二模型或第二模型的部分模型或第一数据集可以用于以下至少一项:
参考信号的处理;
信道信号的传输;
信道信号的解调;
信道状态信息的获取;
波束管理;
信道预测;
信道或信源编译码;
干扰抑制;
定位;
高层业务或参数的预测;
高层业务或参数的管理;
控制信令的解析。
其中,参考信号的处理可以包括对参考信号的检测、滤波、均衡等处理。参考信号包括解调参考信号(DeModulation Reference Signal,DMRS)、探测参考信号(Sounding Reference Signal,SRS)、同步信号块(Synchronization Signal Block,SSB)、信道状态信息参考信号(Channel State Information-Reference Signal,CSI-RS)、跟踪参考信号(Tracking Reference Signal,TRS)、定位参考信号(Positioning Reference Signal,PRS)、相位跟踪参考信号(Phase Tracking Reference Signal,PTRS)等;
信道信号的传输可以包括信道信号的发送、接收等。信道可以包括物理下行控制信道(Physical Downlink Control Channel,PDCCH)、物理下行共享信道(Physical Downlink Shared Channel,PDSCH)、物理上行控制信道(Physical Uplink Control Channel,PUCCH)、物理上行共享信道(Physical Uplink Shared Channel,PUSCH)、物理随机接入信道(Physical Random Access Channel,PRACH)、物理广播信道(Physical Broadcast Channel,PBCH)等;
信道状态信息可以包括信道相关信息、信道矩阵相关信息、信道特征信息、信道矩阵特征信息、预编码矩阵指示(Precoding Matrix Indicator,PMI)、秩指示(Rank Indicator,RI)、CSI-RS资源指示(CSI-RS Resource Indicator,CRI)、信道质量指示(Channel Quality Indicator,CQI)、层指示(Layer Indicator,LI)等;对于频分双工(Frequency Division Duplexing,FDD)系统,根据上下行部分互易性,网络侧设备,如基站可以利用第二模型和上行信道获取角度信息和时延信息,通过CSI-RS预编码或者直接指示的方式,将角度信息和时延信息通知终端,终端根据基站的指示上报或者在基站的指示范围内选择并上报,从而减少终端的计算量和信道状态信息上报的开销;
波束管理可以包括波束测量、波束上报、波束预测、波束失败检测、波束失败恢复、波束失败恢复中的新波束指示等,其中波束预测可以包括空域预测或频域预测;
信道预测可以包括信道状态信息的预测、波束预测等;
信道或信源编译码可以包括信道编码、信道译码、信源编码、信源译码、联合信源信道编码、联合信源信道译码等;
干扰抑制可以包括对小区内干扰、小区间干扰、带外干扰、交调干扰等的抑制;
定位可以理解为利用第二模型和参考信号,如SRS,估计出终端的具体位置,如水平位置或垂直位置,或估计终端未来可能的轨迹,或得到辅助位置估计或轨迹估计的信息,如到达时间(Timing of Arrival,TOA)、视距或非视距、或参考信号时间差(Reference Signal Time Difference,RSTD);
高层业务或参数可以包括吞吐量、所需数据包大小、业务需求、移动速度、噪声信息等;
控制信令可以包括功率控制的相关信令、波束管理的相关信令等。
通信设备利用至少两个第一模型,得到第二模型或第二模型的部分模型或第一数据集,第二模型或第二模型的部分模型或第一数据集用于以上至少一项任务,能够获得更好增益,提升无线通信网络的性能。
可选地,第一数据集可以用于训练第二模型或第二模型的部分模型,或用于表征适用的模型或功能或特性,或用于定义性能指标,或用于定义性能要求。
在本申请实施例中,通信设备利用至少两个第一模型,得到第一数据集后,可以利用第一数据集进行模型训练,得到第二模型,或得到第二模型的部分模型,后续将第二模型或第二模型的部分模型应用到无线通信网络中,获得更好增益。
或者,可以通过第一数据集表征适用于该第一数据集或在该第一数据集上正常工作的模型或功能或特性。适用于该第一数据集或在该第一数据集上正常工作的模型或功能或特性可能是一个,可能是多个,可能对数量不做限定。可以通过数据集标识(dataset ID)、数据特征标识(data categorization ID)、数据相关信息、或数据集相关信息、或数据集关联的信令等关联适用的模型或功能或特性。
或者,可以通过第一数据集定义性能指标(performance KPI)。或者可以通过第一数据集定义性能要求(performance requirement)。
在本申请的一些实施例中,通信设备利用至少两个第一模型,得到第二模型或第二模型的部分模型或第一数据集,可以包括以下步骤:
步骤一:通信设备获得第二数据集;
步骤二:通信设备基于第二数据集中的至少一个数据得到至少一个模型输入;
步骤三:通信设备将至少一个模型输入分别输入到每个第一模型中,得到每个模型输入对应的模型输出或模型中间量;
步骤四:通信设备基于每个模型输入,或每个模型输入对应的模型输出或模型中间量,得到第二模型或第二模型的部分模型或第一数据集。
为方便描述,将上述四个步骤结合起来进行说明。
在本申请实施例中,通信设备可以通过以下至少一种方式获得第二数据集:
接收参考信号,得到参考信号信息;
协议预定义;
接收其他通信设备传输的数据。
第二数据集可以包括一个或多个数据。
通信设备获得第二数据集后,可以基于第二数据集中的至少一个数据得到至少一个模型输入。
可选地,通信设备可以将第二数据集中的每个数据作为一个模型输入,或者通信设备可以对第二数据集中的每个数据进行处理,将处理后的每个数据作为一个模型输入。该处理可以理解为数据格式的转换、数学运算等,如将频域信道转换为时域信道,或者将原始信道信息做奇异值分解得到特征向量,或者将多个数据或元素(如一个数据包括多个元素)组合为向量或矩阵,或者将复数信息转化为两倍大小的实数信息(实部和虚部分别单独列出)等。
通信设备得到至少一个模型输入后,可以将至少一个模型输入分别输入到每个第一模型中,得到每个模型输入对应的模型输出或模型中间量。每个模型输入对应于不同第一模型,对应有相应的模型输出或模型中间量。
举例而言,通信设备获得的至少两个第一模型包括第一模型1、第一模型2,通信设备得到的至少一个模型输入包括模型输入A、模型输入B、模型输入C。通信设备可以将模型输入A分别输入到第一模型1和第一模型2中,得到模型输入A对应的第一模型1的模型输出A1’,或得到模型输入A对应的第一模型1的模型中间量A1”,以及得到模型输入A对应的第一模型2的模型输出A2’,或得到模型输入A对应的第一模型2的模型中间量A2”。将模型输入B分别输入到第一模型1和第一模型2中,得到模型输入B对应的第一模型1的模型输出B1’,或得到模型输入B对应的第一模型1的模型中间量B1”,以及得到模型输入B对应的第一模型2的模型输出B2’,或得到模型输入B对应的第一模型2的模型中间量B2”。将模型输入C分别输入到第一模型1和第一模型2中,得到模型输入C对应的第一模型1的模型输出C1’,或得到模型输入C对应的第一模型1的模型中间量C1”,以及得到模型输入C对应的第一模型2的模型输出C2’,或得到模型输入C对应的第一模型2的模型中间量C2”。
这样可以得到每个模型输入与模型输出的映射关系,或得到每个模型输入与模型中间量的映射关系,或得到每个模型输入与模型输出以及模型中间量的映射关系,或得到每个模型输入对应的模型输出以及模型中间量的映射关系。
如,每个模型输入与模型输出或模型中间量的映射关系如下:
模型输入A—模型输出A1’—模型中间量A1”;
模型输入A—模型输出A2’—模型中间量A2”;
模型输入B—模型输出B1’—模型中间量B1”;
模型输入B—模型输出B2’—模型中间量B2”;
模型输入C—模型输出C1’—模型中间量C1”;
模型输入C—模型输出C2’—模型中间量C2”。
通信设备基于每个模型输入,或每个模型输入对应的模型输出或模型中间量,可以得到第二模型或第二模型的部分模型或第一数据集。
通信设备将模型输入输入到每个第一模型中,得到每个模型输入对应的模型输出或模型中间量,基于每个模型输入,或每个模型输入对应的模型输出或模型中间量可以准确得到第二模型或第二模型的部分模型或第一数据集。
在本申请的一些实施例中,通信设备基于每个模型输入,或每个模型输入对应的模型输出或模型中间量,得到第二模型或第二模型的部分模型或第一数据集,可以包括以下步骤:
第一个步骤:通信设备基于每个模型输入,或每个模型输入对应的模型输出或模型中间量,确定每个模型输入对应的模型最终输出或模型最终中间量;
第二个步骤:通信设备利用至少一个模型输入,或每个模型输入对应的模型最终输出或模型最终中间量,得到第二模型或第二模型的部分模型或第一数据集。
为方便描述,将上述两个步骤结合起来进行说明。
在本申请实施例中,通信设备将至少一个模型输入分别输入到每个第一模型中后,可以得到每个模型输入对应的模型输出或模型中间量,每个模型输入对应于不同第一模型,对应有相应的模型输出或模型中间量。
通信设备针对每个模型输入,可以根据当前模型输入,或当前模型输入对应的模型输出或模型中间量,确定当前模型输入对应的模型最终输出或模型最终中间量。当前模型输入对应的模型最终输出可以理解为当前模型输入对应于多个第一模型的一种融合输出,当前模型输入对应的模型最终中间量可以理解为当前模型输入对应于多个第一模型的一种融合中间量。当前模型输入是指当前操作所针对的模型输入。
通信设备利用至少一个模型输入,或每个模型输入对应的模型最终输出或模型最终中间量,可以得到第二模型或第二模型的部分模型或第一数据集。
可选地,通信设备可以利用至少一个模型输入,或每个模型输入对应的模型最终输出或模型最终中间量,进行模型训练,得到第二模型或第二模型的部分模型。
可选地,在第二模型为单边模型(One-sided model)的情况下,第二模型的输入为模型输入,第二模型的输出为模型最终输出。
如果第二模型为单边模型,则将模型输入输入到第二模型中,可以得到第二模型的输出,第二模型的输出为该模型输入对应的模型最终输出。
比如,模型输入包括模型输入A、模型输入B、模型输入C,模型输入A对应的模型最终输出为A0’,模型输入B对应的模型最终输出为B0’,模型输入C对应的模型最终输出为C0’。在进行模型训练时,第二模型的输入包括模型输入A、模型输入B和模型输入C,模型输入A对应的第二模型的输出为模型最终输出A0’,模型输入B对应的第二模型的输出为模型最终输出B0’,模型输入C对应的第二模型的输出为模型最终输出C0’。
将模型输入作为第二模型的输入,将模型最终输出作为第二模型的输出,有助于提高第二模型训练效率,提高第二模型的准确性。
可选地,在第二模型为双边模型(Two-sided model)的情况下,第二模型的编码器的输入为模型输入,第二模型的编码器的输出或第二模型的译码器的输入为模型最终中间量,第二模型的译码器的输出为模型最终输出。
如果第二模型为双边模型,则第二模型可以包括编码器和译码器,可以将模型输入输入到第二模型的编码器中,第二模型的编码器的输出为该模型输入对应的模型最终中间量,将模型最终中间量输入到第二模型的译码器中,可以得到第二模型的译码器的输出,即该模型输入对应的模型最终输出。
在一种实施例中,通信设备利用至少两个第一模型,得到第二模型的部分模型,如第二模型的编码器或译码器。当通信设备需要第二模型的编码器时,可以利用模型输入和模型最终中间量得到;当通信设备需要第二模型的译码器时,可以利用模型最终中间量和模型最终输出得到。
比如,模型输入包括模型输入A、模型输入B、模型输入C,模型输入A对应的模型最终输出为A0’,模型输入A对应的模型最终中间量为A0”,模型输入B对应的模型最终输出为B0’,模型输入B对应的模型最终中间量为B0”,模型输入C对应的模型最终输出为C0’,模型输入C对应的模型最终中间量为C0”。在进行模型训练时,第二模型的编码器的输入包括模型输入A、模型输入B和模型输入C,模型输入A对应的第二模型的编码器的输出以及第二模型的译码器的输入为模型最终中间量A0”,模型输入A对应的第二模型的译码器的输出为模型最终输出A0’,模型输入B对应的第二模型的编码器的输出以及第二模型的译码器的输入为模型最终中间量B0”,模型输入B对应的第二模型的输出为模型最终输出B0’,模型输入C对应的第二模型的编码器的输出以及第二模型的译码器的输入为模型最终中间量C0”,模型输入C对应的第二模型的输出为模型最终输出C0’。
将模型输入作为第二模型的编码器的输入,将模型最终中间量作为第二模型的编码器的输出或第二模型的译码器的输入,将模型最终输出作为第二模型的译码器的输出,有助于提高第二模型训练效率,提高第二模型的准确性。
可选地,第一数据集可以包括至少一个数据,每个数据可以包括一个模型输入,或该模型输入对应的模型最终输出或模型最终中间量。
在一些实施例中,第一模型为单边模型,第一数据集也对应单边模型,则第一数据集可以包括一个或多个数据,每个数据可以包括一个模型输入,和该模型输入对应的模型最终输出。
在一些实施例中,第一模型为双边模型,第一数据集也对应双边模型,则第一数据集可以包括一个或多个数据,每个数据可以包括一个模型输入,和该模型输入对应的模型最终输出和模型最终中间量。
在一些实施例中,第一模型为双边模型,第一数据集对应双边模型的编码器,则第一数据集可以包括一个或多个数据,每个数据可以包括一个模型输入,和该模型输入对应的模型最终中间量。
在一些实施例中,第一模型为双边模型,第一数据集对应双边模型的译码器,则第一数据集可以包括一个或多个数据,每个数据可以包括一个模型输入对应的模型最终中间量和模型最终输出(不用包括该模型输入)。
通信设备利用至少一个模型输入,或每个模型输入对应的模型最终输出或模型最终中间量,可以通过模型训练得到第二模型或第二模型的部分模型,或者生成第一数据集,可以保证第二模型或第二模型的部分模型或第一数据集的准确性。
在本申请的一些实施例中,针对每个模型输入,当前模型输入对应的模型最终输出基于以下至少一项确定:
当前模型输入;
对当前模型输入对应的模型输出集合进行运算得到的第一结果;
当前模型输入对应的模型输出集合中的第一模型输出,第一模型输出与当前模型输入对应的输出标签的相似度最高;
其中,模型输出集合包括:将当前模型输入分别输入到每个第一模型后,得到的模型输出。
在本申请实施例中,通信设备得到至少一个模型输入后,可以将至少一个模型输入分别输入到每个第一模型中,得到每个模型输入对应的模型输出或模型中间量,一个模型输入对应有多个模型输出或多个模型中间量。
针对每个模型输入,可以基于以下至少一项确定当前模型输入对应的模型最终输出:
1)当前模型输入;一种实施例中,在第一模型或第二模型为双边模型的情况下,可以将当前模型输入确定为当前模型输入对应的模型最终输出;
2)对当前模型输入对应的模型输出集合进行运算得到的第一结果;可选地,第一结果可以是对当前模型输入对应的模型输出集合进行平均化运算得到的结果。模型输出集合包括:将当前模型输入分别输入到每个第一模型后,得到的模型输出。模型输出集合包括多个模型输出,可以对当前模型输入对应的模型输出集合包括的模型输出进行平均化处理,如进行线性平均、几何平均、调和平均、平方平均、加权平均、最小值最大化、最大值最小化、上述组合或简单变化等,得到第一结果,可以将第一结果确定为当前模型输入对应的模型最终输出;可选地,第一结果还可以是对当前模型输入对应的模型输出集合进行其他数学运算得到的结果,如对模型输出集合进行求和运算,或归一化运算,或特征值分解运算等得到的结果,在此不一一列举;
3)当前模型输入对应的模型输出集合中的第一模型输出,第一模型输出与当前模型输入对应的输出标签的相似度最高;当前模型输入对应的模型输出集合包括多个模型输出,可以分别确定每个模型输出与当前模型输入对应的输出标签的相似度,得到最高相似度对应的第一模型输出,可以将该第一模型输出确定为当前模型输入对应的模型最终输出。第一模型输出与当前模型输入对应的输出标签的相似度最高,可以理解为第一模型输出与当前模型输入对应的输出标签最接近。相似度最高可以理解为相关性、余弦相似度(cosine similarity)、余弦相似度的平方等最高,或者理解为差距、差别、均方误差(Normalized Mean Square Error,NMSE)、距离、欧式距离等最小,或者理解为相差的绝对值、幅值、功率等最小。
当前模型输入是指当前操作所针对的模型输入。
通信设备可以基于上述一项确定每个模型输入对应的模型最终输出,也可以基于上述内容的结合确定每个模型输入对应的模型最终输出,比如,将当前模型输入与第一结果进行数学运算后,将得到的结果确定为当前模型输入对应的模型最终输出。
可选地,可以通过协议预定义上述一项内容,通信设备根据协议预定义的内容确定每个模型输入对应的模型最终输出。
可选地,可以通过协议预定义上述多项内容,通信设备根据协议预定义,从中选择一项,基于选择的内容确定每个模型输入对应的模型最终输出。
可选地,可以通过协议预定义上述多项内容以及结合方式,通信设备根据协议预定义,将至少两项内容结合起来,确定每个模型输入对应的模型最终输出。
通信设备基于上述至少一项内容,可以较为准确地确定出每个模型输入对应的模型最终输出。
在本申请的一些实施例中,针对每个模型输入,当前模型输入对应的模型最终中间量基于以下至少一项确定:
对当前模型输入对应的模型中间量集合进行运算得到的第二结果;模型中间量集合包括:将当前模型输入分别输入到每个第一模型后,得到的模型中间量;
对当前模型输入对应的模型中间量集合以及模型输出集合进行运算得到的第三结果;
模型中间量集合中的第一模型中间量,第一模型中间量与第二模型输出是利用同一第一模型得到的,第二模型输出为当前模型输入对应的模型输出集合中的一个模型输出,第二模型输出与当前模型输入对应的输出标签的相似度最高,模型输出集合包括:将当前模型输入分别输入到每个第一模型后,得到的模型输出。
在本申请实施例中,通信设备得到至少一个模型输入后,可以将至少一个模型输入分别输入到每个第一模型中,得到每个模型输入对应的模型输出或模型中间量,一个模型输入对应有多个模型输出或多个模型中间量。
针对每个模型输入,可以基于以下至少一项确定当前模型输入对应的模型最终中间量:
1)对当前模型输入对应的模型中间量集合进行运算得到的第二结果;可选地,第二结果可以为对当前模型输入对应的模型中间量集合进行平均化运算后得到的结果。模型中间量集合包括:将当前模型输入分别输入到每个第一模型后,得到的模型中间量。模型中间量集合包括多个模型中间量,可以对当前模型输入对应的模型中间量集合包括的模型中间量进行平均化处理,如进行线性平均、几何平均、调和平均、平方平均、加权平均、最小值最大化、最大值最小化、上述组合或简单变化等,得到第二结果,可以将第二结果确定为当前模型输入对应的模型最终中间量;可选地,第二结果还可以是对当前模型输入对应的模型中间量集合进行其他数学运算得到的结果,如对模型中间量集合进行求和运算,或归一化运算,或特征值分解运算等得到的结果,在此不一一列举;
2)对当前模型输入对应的模型中间量集合以及模型输出集合进行运算得到的第三结果;可选地,第三结果可以是对当前模型输入对应的模型中间量集合与模型输出集合进行平均化运算后得到的结果,或者是对当前模型输入对应的模型中间量集合与模型输出集合进行其他数学运算得到的结果,如对当前模型输入对应的模型中间量集合与模型输出集合进行合并运算、或求和运算,或归一化运算,或特征值分解运算等得到的结果,在此不一一列举
3)模型中间量集合中的第一模型中间量;在第一模型为双边模型的情况下,将当前模型输入输入到每个第一模型后,可以得到对应于每个第一模型的模型输出和模型中间量。当前模型输入对应的模型输出集合包括多个模型输出,可以分别确定每个模型输出与当前模型输入对应的输出标签的相似度,得到最高相似度对应的第二模型输出。可以将该第二模型输出对应的第一模型中间量确定为当前模型输入对应的模型最终中间量。第一模型中间量与第二模型输出是利用同一第一模型得到的。第二模型输出与当前模型输入对应的输出标签的相似度最高,可以理解为第二模型输出与当前模型输入对应的输出标签最接近。相似度最高可以理解为相关性、余弦相似度(cosine similarity)、余弦相似度的平方等最高,或者理解为差距、差别、均方误差(Normalized Mean Square Error,NMSE)、距离、欧式距离等最小,或者理解为相差的绝对值、幅值、功率等最小。
当前模型输入是指当前操作所针对的模型输入。
通信设备可以基于上述一项确定每个模型输入对应的模型最终中间量,也可以基于上述内容的结合确定每个模型输入对应的模型最终中间量,比如,将第二结果与第一模型中间量进行数学运算后,将得到的结果确定为当前模型输入对应的模型最终中间量。
可选地,可以通过协议预定义上述一项内容,通信设备根据协议预定义的内容确定每个模型输入对应的模型最终中间量。
可选地,可以通过协议预定义上述多项内容,通信设备根据协议预定义,从中选择一项,基于选择的内容确定每个模型输入对应的模型最终中间量。
可选地,可以通过协议预定义上述多项内容以及结合方式,通信设备根据协议预定义,将至少两项内容结合起来,确定每个模型输入对应的模型最终中间量。
通信设备基于上述至少一项内容,可以较为准确地确定出每个模型输入对应的模型最终中间量。
在本申请的一些实施例中,通信设备利用至少两个第一模型,得到第二模型或第二模型的部分模型,可以包括以下步骤:
通信设备将至少两个第一模型组合为第二模型或第二模型的部分模型。
在本申请实施例中,通信设备获得至少两个第一模型后,可以将至少两个第一模型组合为新的一个大模型,作为第二模型或第二模型的部分模型。可选地,通信设备可以根据协议预定义的模型融合方法将至少两个第一模型组合为第二模型或第二模型的部分模型。如,可以将至少两个第一模型的模型参数取平均。
通信设备将至少两个第一模型组合为第二模型或第二模型的部分模型,可以提高模型融合效率。
为了方便理解,下面再通过具体示例对本申请实施例所提供的技术方案进行说明。
示例一,第一模型和第二模型为单边模型
在该示例中,假设模型输入包括模型输入A,至少两个第一模型包括第一模型1和第一模型2。
如图5所示,将模型输入A分别输入到第一模型1和第一模型2中,得到两个模型输出,包括第一模型1的模型输出A1’,以及第一模型2的模型输出A2’。
通过上述两个模型输出可以得到模型输入A对应的模型最终输出(Final Output)A0’,A0’=F(模型输出A1’,模型输出A2’)。
其中,模型最终输出可以是上述两个模型输出的平均;平均方式包括:线性平均、几何平均、调和平均、平方平均、加权平均、最小值最大化、最大值最小化、及上述至少两种平均方式的组合或简单变化等;
模型最终输出还可以是上述两个模型输出中最接近模型输入A对应的输出标签的模型输出;M最接近N可以理解为:M与N的相关性、余弦相似度、余弦相似度的平方等最高,或M与N的差距、差别、NMSE、距离、欧式距离等最小,或(M-N)的绝对值、幅值、功率等最小。
利用模型输入或模型最终输出,进行模型融合得到第二模型或生成第一数据集。
其中,第二模型的输入为模型输入A,输出为模型最终输出A0’;
生成的第一数据集的一个数据,包括模型输入A,或模型最终输出A0’。
示例二,第一模型和第二模型为双边模型。
在该示例中,假设模型输入包括模型输入B,至少两个第一模型包括第一模型3和第一模型4。
如图6所示,将模型输入B分别输入到第一模型3的编码器(Encoder)和第一模型4的编码器中,得到两个模型中间量(inter-data),包括第一模型3的模型中间量B3”,以及第一模型4的模型中间量B4”。将这两个模型中间量分别输入到对应的第一模型3的译码器(Decoder)和第一模型4的译码器中,得到两个模型输出,包括第一模型3的模型输出B1’,以及第一模型4的模型输出B4’。其中,模型输出B3’和模型中间量B3”是利用同一第一模型,即第一模型3得到的,模型输出B4’和模型中间量B4”是利用同一第一模型,即第一模型4得到的。
通过上述两个模型输出可以得到模型输入B对应的模型最终中间量(Final inter-data)B0”或模型最终输出(Final Output)B0’。
B0’=F(模型输出B3’,模型输出B4’)。
对于模型最终输出B0’的确定有如下方案:
方案1:模型最终输出B0’是模型输入B;
如果第二模型用于信道信息/CSI的压缩或恢复,如编码器用于将信道信息/CSI压缩为反馈信息,译码器用于将反馈信息恢复为信道信息/CSI,或第二模型用于编码或译码,如信源编译码、信道编译码、信源信道联合编译码等(编码器用于编码,译码器用于译码),则可以使用该方案确定模型最终输出。
方案2:模型最终输出B0’为上述两个模型输出的平均;
平均方式包括:线性平均、几何平均、调和平均、平方平均、加权平均、最小值最大化、最大值最小化,及上述至少两种平均方式的组合或简单变化。
方案3:模型最终输出B0’为上述两个模型输出中最接近输出标签的模型输出;
M最接近N可以理解为:M与N的相关性、余弦相似度、余弦相似度的平方等最高,或M与N的差距、差别、NMSE、距离、欧式距离等最小,或(M-N)的绝对值、幅值、功率等最小;
对于用于信道信息/CSI的压缩与恢复,或用于编码和译码(信源编译码、信道编译码、或信源信道联合编译码)的第一模型:输出标签为模型输入。模型最终输出B0’为上述两个模型输出中最接近模型输入的模型输出。
对于模型最终中间量B0”的确定有如下方案:
方案1:B0”=F(模型中间量B3”,模型中间量B4”),可选地,模型最终中间量B0”为上述两个模型中间量的平均;
方案2:模型最终中间量B0”为上述两个模型输出中最接近输出标签的模型输出对应的模型中间量。
对于模型最终输出和模型最终中间量的确定有如下可选组合方式:
模型最终输出的方案1与模型最终中间量的方案1、2可以组合;
模型最终输出的方案2与模型最终中间量的方案1可以组合;
模型最终输出的方案3与模型最终中间量的方案2可以组合。
利用模型输入、或模型最终输出、或模型最终中间量,进行模型融合可以得到第二模型或生成第一数据集。
其中,第二模型的编码器的输入为模型输入B,编码器的输出或译码器的输入为模型最终中间量B0”,译码器的输出为模型最终输出B0’;
生成的第一数据集的一个数据,包括模型输入B、或模型最终中间量B0”,或模型最终输出B0’。
本申请实施例能够将多个第一模型融合为一个第二模型或第二模型的部分模型,或生成参考数据集,从而拉齐了多个第一模型输入到输出的映射关系,使用第二模型或第二模型的部分模型能够提升无线通信系统的性能。
本申请实施例提供的模型处理方法,执行主体可以为模型处理装置。本申请实施例中以模型处理装置执行模型处理方法为例,说明本申请实施例提供的模型处理装置。
如图7所示,模型处理装置700包括以下模块:
第一获得模块710,用于获得至少两个第一模型;
第二获得模块720,用于利用至少两个第一模型,得到第二模型或第二模型的部分模型或第一数据集。
应用本申请实施例所提供的装置,获得至少两个第一模型后,利用至少两个第一模型,得到第二模型或第一数据集,将多个第一模型进行融合得到第二模型或第二模型的部分模型或第一数据集,拉齐了多个第一模型输入到输出的映射关系,将第二模型或第二模型的部分模型或第一数据集应用到无线通信网络中,能够获得更好增益,有助于提升无线通信网络的性能。
在本申请的一些实施例中,第二获得模块720,包括:
第一获得子模块,用于获得第二数据集;
第二获得子模块,用于基于第二数据集中的至少一个数据得到至少一个模型输入;
第三获得子模块,用于将至少一个模型输入分别输入到每个第一模型中,得到每个模型输入对应的模型输出或模型中间量;
第四获得子模块,用于基于每个模型输入,或每个模型输入对应的模型输出或模型中间量,得到第二模型或第二模型的部分模型或第一数据集。
在本申请的一些实施例中,第四获得子模块,具体用于:
基于每个模型输入,或每个模型输入对应的模型输出或模型中间量,确定每个模型输入对应的模型最终输出或模型最终中间量;
利用至少一个模型输入,或每个模型输入对应的模型最终输出或模型最终中间量,得到第二模型或第二模型的部分模型或第一数据集。
在本申请的一些实施例中,针对每个模型输入,当前模型输入对应的模型最终输出基于以下至少一项确定:
当前模型输入;
对当前模型输入对应的模型输出集合进行运算得到的第一结果;
当前模型输入对应的模型输出集合中的第一模型输出,第一模型输出与当前模型输入对应的输出标签的相似度最高;
其中,模型输出集合包括:将当前模型输入分别输入到每个第一模型后,得到的模型输出。
在本申请的一些实施例中,针对每个模型输入,当前模型输入对应的模型最终中间量基于以下至少一项确定:
对当前模型输入对应的模型中间量集合进行运算得到的第二结果;
对当前模型输入对应的模型中间量集合以及模型输出集合进行运算得到的第三结果;;
模型中间量集合中的第一模型中间量,第一模型中间量与第二模型输出是利用同一第一模型得到的,第二模型输出为当前模型输入对应的模型输出集合中的一个模型输出,第二模型输出与当前模型输入对应的输出标签的相似度最高;
其中,模型中间量集合包括:将当前模型输入分别输入到每个第一模型后,得到的模型中间量;
模型输出集合包括:将当前模型输入分别输入到每个第一模型后,得到的模型输出。
在本申请的一些实施例中,在第二模型为单边模型的情况下,第二模型的输入为模型输入,第二模型的输出为模型最终输出。
在本申请的一些实施例中,在第二模型为双边模型的情况下,第二模型的编码器的输入为模型输入,第二模型的编码器的输出或第二模型的译码器的输入为模型最终中间量,第二模型的译码器的输出为模型最终输出。
在本申请的一些实施例中,第一获得模块,具体用于:
通过以下至少一种方式获得第二数据集:
接收参考信号,得到参考信号信息;
协议预定义;
接收其他通信设备传输的数据。
在本申请的一些实施例中,第二获得模块,具体用于:
将至少两个第一模型组合为第二模型或第二模型的部分模型。
在本申请的一些实施例中,第一数据集用于训练第二模型或第二模型的部分模型,或用于表征适用的模型或功能或特性,或用于定义性能指标,或用于定义性能要求。
在本申请的一些实施例中,至少两个第一模型或第二模型或第二模型的部分模型的以下至少一项是协议预定义的:
模型基本架构;模型结构;模型各层参数;神经元系数;模型复杂度;模型大小;激活函数;量化方式。
在本申请的一些实施例中,至少两个第一模型或第二模型或第二模型的部分模型以下至少一项相同:
模型基本架构;模型结构;激活函数;模型复杂度;模型复杂度的最大值;模型复杂度的最小值;模型大小;模型大小的最大值、模型大小的最小值。
在本申请的一些实施例中,第二模型或第二模型的部分模型与至少两个第一模型的输入格式或输出格式相同。
在本申请的一些实施例中,第二模型或第二模型的部分模型或第一数据集用于以下至少一项:
参考信号的处理;
信道信号的传输;
信道信号的解调;
信道状态信息的获取;
波束管理;
信道预测;
信道或信源编译码;
干扰抑制;
定位;
高层业务或参数的预测;
高层业务或参数的管理;
控制信令的解析。
在本申请的一些实施例中,第一获得模块710,具体用于:
获得至少两个第一模型的部分模型;
利用所述至少两个第一模型的部分模型和第三数据集,得到至少两个第一模型。
本申请实施例提供的模型处理装置700能够实现图4所示方法实施例实现的各个过程,并达到相同的技术效果,为避免重复,这里不再赘述。
如图8所示,本申请实施例还提供一种通信设备800,包括处理器801和存储器802,存储器802上存储有可在处理器801上运行的程序或指令,例如,该通信设备800为终端时,该程序或指令被处理器801执行时实现上述模型处理方法实施例的各个步骤,且能达到相同的技术效果。该通信设备800为网络侧设备时,该程序或指令被处理器801执行时实现上述模型处理方法实施例的各个步骤,且能达到相同的技术效果,为避免重复,这里不再赘述。
本申请实施例还提供一种终端,包括处理器和通信接口,通信接口和处理器耦合,处理器用于运行程序或指令,实现如图4所示方法实施例中的步骤。该终端实施例与上述方法实施例对应,上述方法实施例的各个实施过程和实现方式均可适用于该终端实施例中,且能达到相同的技术效果。
具体地,图9为实现本申请实施例的一种终端的硬件结构示意图。
该终端900包括但不限于:射频单元901、网络模块902、音频输出单元903、输入单元904、传感器905、显示单元906、用户输入单元907、接口单元908、存储器909以及处理器910等中的至少部分部件。
本领域技术人员可以理解,终端900还可以包括给各个部件供电的电源(比如电池),电源可以通过电源管理系统与处理器910逻辑相连,从而通过电源管理系统实现管理充电、放电以及功耗管理等功能。图9中示出的终端结构并不构成对终端的限定,终端可以包括比图示更多或更少的部件,或者组合某些部件,或者不同的部件布置,在此不再赘述。
应理解的是,本申请实施例中,输入单元904可以包括图形处理单元(Graphics Processing Unit,GPU)9041和麦克风9042,图形处理器9041对在视频捕获模式或图像捕获模式中由图像捕获装置(如摄像头)获得的静态图片或视频的图像数据进行处理。显示单元906可包括显示面板9061,可以采用液晶显示器、有机发光二极管等形式来配置显示面板9061。用户输入单元907包括触控面板9071以及其他输入设备9072中的至少一种。触控面板9071,也称为触摸屏。触控面板9071可包括触摸检测装置和触摸控制器两个部分。其他输入设备9072可以包括但不限于物理键盘、功能键(比如音量控制按键、开关按键等)、轨迹球、鼠标、操作杆,在此不再赘述。
本申请实施例中,射频单元901接收来自网络侧设备的下行数据后,可以传输给处理器910进行处理;另外,射频单元901可以向网络侧设备发送上行数据。通常,射频单元901包括但不限于天线、放大器、收发信机、耦合器、低噪声放大器、双工器等。
存储器909可用于存储软件程序或指令以及各种数据。存储器909可主要包括存储程序或指令的第一存储区和存储数据的第二存储区,其中,第一存储区可存储操作系统、至少一个功能所需的应用程序或指令(比如声音播放功能、图像播放功能等)等。此外,存储器909可以包括易失性存储器或非易失性存储器。其中,非易失性存储器可以是只读存储器(Read-Only Memory,ROM)、可编程只读存储器(Programmable ROM,PROM)、可擦除可编程只读存储器(Erasable PROM,EPROM)、电可擦除可编程只读存储器(Electrically EPROM,EEPROM)或闪存。易失性存储器可以是随机存取存储器(Random Access Memory,RAM),静态随机存取存储器(Static RAM,SRAM)、动态随机存取存储器(Dynamic RAM,DRAM)、同步动态随机存取存储器(Synchronous DRAM,SDRAM)、双倍数据速率同步动态随机存取存储器(Double Data Rate SDRAM,DDRSDRAM)、增强型同步动态随机存取存储器(Enhanced SDRAM,ESDRAM)、同步连接动态随机存取存储器(Synch link DRAM,SLDRAM)和直接内存总线随机存取存储器(Direct Rambus RAM,DRRAM)。本申请实施例中的存储器909包括但不限于这些和任意其它适合类型的存储器。
处理器910可包括一个或多个处理单元;可选的,处理器910集成应用处理器和调制解调处理器,其中,应用处理器主要处理涉及操作系统、用户界面和应用程序等的操作,调制解调处理器主要处理无线通信信号,如基带处理器。可以理解的是,上述调制解调处理器也可以不集成到处理器910中。
其中,处理器910,用于获得至少两个第一模型;
利用至少两个第一模型,得到第二模型或第二模型的部分模型或第一数据集。
可以理解,本实施例中提及的各实现方式的实现过程可以参照方法实施例的相关描述,并达到相同或相应的技术效果,为避免重复,在此不再赘述。
本申请实施例还提供一种网络侧设备,包括处理器和通信接口,所述通信接口和所述处理器耦合,所述处理器用于运行程序或指令,实现如图4所示的方法实施例的步骤。该网络侧设备实施例与上述方法实施例对应,上述方法实施例的各个实施过程和实现方式均可适用于该网络侧设备实施例中,且能达到相同的技术效果。
具体地,本申请实施例还提供了一种网络侧设备。如图10所示,该网络侧设备1000包括:天线1001、射频装置1002、基带装置1003、处理器1004和存储器1005。天线1001与射频装置1002连接。在上行方向上,射频装置1002通过天线1001接收信息,将接收的信息发送给基带装置1003进行处理。在下行方向上,基带装置1003对要发送的信息进行处理,并发送给射频装置1002,射频装置1002对收到的信息进行处理后经过天线1001发送出去。
以上实施例中网络侧设备执行的方法可以在基带装置1003中实现,该基带装置1003包括基带处理器。
基带装置1003例如可以包括至少一个基带板,该基带板上设置有多个芯片,如图10所示,其中一个芯片例如为基带处理器,通过总线接口与存储器1005连接,以调用存储器1005中的程序,执行以上方法实施例中所示的网络侧设备操作。
该网络侧设备还可以包括网络接口1006,该接口例如为通用公共无线接口(Common Public Radio Interface,CPRI)。
具体地,本申请实施例的网络侧设备1000还包括:存储在存储器1005上并可在处理器1004上运行的指令或程序,处理器1004调用存储器1005中的指令或程序执行模型处理装置700中各模块执行的方法,并达到相同的技术效果,为避免重复,故不在此赘述。
具体地,本申请实施例还提供了一种网络侧设备。如图11所示,该网络侧设备1100包括:处理器1101、网络接口1102和存储器1103。其中,网络接口1102例如为通用公共无线接口(common public radio interface,CPRI)。
具体地,本申请实施例的网络侧设备1100还包括:存储在存储器1103上并可在处理器1101上运行的指令或程序,处理器1101调用存储器1103中的指令或程序执行模型处理装置700中各模块执行的方法,并达到相同的技术效果,为避免重复,故不在此赘述。
本申请实施例还提供一种可读存储介质,所述可读存储介质上存储有程序或指令,该程序或指令被处理器执行时实现上述方法实施例的各个过程,且能达到相同的技术效果,为避免重复,这里不再赘述。
其中,所述处理器为上述实施例中所述的终端中的处理器。所述可读存储介质,包括计算机可读存储介质,如计算机只读存储器ROM、随机存取存储器RAM、磁碟或者光盘等。在一些示例中,可读存储介质可以是非瞬态的可读存储介质。
本申请实施例另提供了一种芯片,所述芯片包括处理器和通信接口,所述通信接口和所述处理器耦合,所述处理器用于运行程序或指令,实现上述方法实施例的各个过程,且能达到相同的技术效果,为避免重复,这里不再赘述。
应理解,本申请实施例提到的芯片还可以称为系统级芯片,系统芯片,芯片系统或片上系统芯片等。
本申请实施例另提供了一种计算机程序/程序产品,所述计算机程序/程序产品被存储在存储介质中,所述计算机程序/程序产品被至少一个处理器执行以实现上述方法实施例的各个过程,且能达到相同的技术效果,为避免重复,这里不再赘述。
本申请实施例还提供了一种无线通信系统,包括:终端及网络侧设备,终端可用于执行如上所述的模型处理方法的步骤,所述网络侧设备可用于执行如上所述的模型处理方法的步骤。
需要说明的是,在本文中,术语“包括”、“包含”或者其任何其他变体意在涵盖非排他性的包含,从而使得包括一系列要素的过程、方法、物品或者装置不仅包括那些要素,而且还包括没有明确列出的其他要素,或者是还包括为这种过程、方法、物品或者装置所固有的要素。在没有更多限制的情况下,由语句“包括一个……”限定的要素,并不排除在包括该要素的过程、方法、物品或者装置中还存在另外的相同要素。此外,需要指出的是,本申请实施方式中的方法和装置的范围不限按示出或讨论的顺序来执行功能,还可包括根据所涉及的功能按基本同时的方式或按相反的顺序来执行功能,例如,可以按不同于所描述的次序来执行所描述的方法,并且还可以添加、省去或组合各种步骤。另外,参照某些示例所描述的特征可在其他示例中被组合。
通过以上的实施方式的描述,本领域的技术人员可以清楚地了解到上述实施例方法可借助计算机软件产品加必需的通用硬件平台的方式来实现,当然也可以通过硬件。该计算机软件产品存储在存储介质(如ROM、RAM、磁碟、光盘等)中,包括若干指令,用以使得终端或者网络侧设备执行本申请各个实施例所述的方法。
上面结合附图对本申请的实施例进行了描述,但是本申请并不局限于上述的具体实施方式,上述的具体实施方式仅仅是示意性的,而不是限制性的,本领域的普通技术人员在本申请的启示下,在不脱离本申请宗旨和权利要求所保护的范围情况下,还可做出很多形式的实施方式,这些实施方式均属于本申请的保护之内。
Claims (32)
- 一种模型处理方法,其中,包括:通信设备获得至少两个第一模型;所述通信设备利用所述至少两个第一模型,得到第二模型或第二模型的部分模型或第一数据集。
- 根据权利要求1所述的模型处理方法,其中,所述通信设备利用所述至少两个第一模型,得到第二模型或第二模型的部分模型或第一数据集,包括:所述通信设备获得第二数据集;所述通信设备基于所述第二数据集中的至少一个数据得到至少一个模型输入;所述通信设备将所述至少一个模型输入分别输入到每个第一模型中,得到每个模型输入对应的模型输出或模型中间量;所述通信设备基于每个模型输入,或每个模型输入对应的模型输出或模型中间量,得到第二模型或第二模型的部分模型或第一数据集。
- 根据权利要求2所述的模型处理方法,其中,所述通信设备基于每个模型输入,或每个模型输入对应的模型输出或模型中间量,得到第二模型或第二模型的部分模型或第一数据集,包括:所述通信设备基于每个模型输入,或每个模型输入对应的模型输出或模型中间量,确定每个模型输入对应的模型最终输出或模型最终中间量;所述通信设备利用所述至少一个模型输入,或每个模型输入对应的模型最终输出或模型最终中间量,得到第二模型或第二模型的部分模型或第一数据集。
- 根据权利要求3所述的模型处理方法,其中,针对每个模型输入,当前模型输入对应的模型最终输出基于以下至少一项确定:所述当前模型输入;对所述当前模型输入对应的模型输出集合进行运算得到的第一结果;所述当前模型输入对应的模型输出集合中的第一模型输出,所述第一模型输出与所述当前模型输入对应的输出标签的相似度最高;其中,所述模型输出集合包括:将所述当前模型输入分别输入到每个第一模型后,得到的模型输出。
- 根据权利要求3或4所述的模型处理方法,其中,针对每个模型输入,当前模型输入对应的模型最终中间量基于以下至少一项确定:对所述当前模型输入对应的模型中间量集合进行运算得到的第二结果;对所述当前模型输入对应的模型中间量集合以及模型输出集合进行运算得到的第三结果;所述模型中间量集合中的第一模型中间量,所述第一模型中间量与第二模型输出是利用同一第一模型得到的,所述第二模型输出为所述当前模型输入对应的模型输出集合中的一个模型输出,所述第二模型输出与所述当前模型输入对应的输出标签的相似度最高;其中,所述模型中间量集合包括:将所述当前模型输入分别输入到每个第一模型后,得到的模型中间量;所述模型输出集合包括:将所述当前模型输入分别输入到每个第一模型后,得到的模型输出。
- 根据权利要求3至5之中任一项所述的模型处理方法,其中,在所述第二模型为单边模型的情况下,所述第二模型的输入为所述模型输入,所述第二模型的输出为所述模型最终输出。
- 根据权利要求3至5之中任一项所述的模型处理方法,其中,在所述第二模型为双边模型的情况下,所述第二模型的编码器的输入为所述模型输入,所述第二模型的编码器的输出或所述第二模型的译码器的输入为所述模型最终中间量,所述第二模型的译码器的输出为所述模型最终输出。
- 根据权利要求2至7之中任一项所述的模型处理方法,其中,所述通信设备获得第二数据集,包括:所述通信设备通过以下至少一种方式获得第二数据集:接收参考信号,得到参考信号信息;协议预定义;接收其他通信设备传输的数据。
- 根据权利要求1所述的模型处理方法,其中,所述通信设备利用所述至少两个第一模型,得到第二模型或第二模型的部分模型,包括:所述通信设备将所述至少两个第一模型组合为第二模型或第二模型的部分模型。
- 根据权利要求1至9之中任一项所述的模型处理方法,其中,所述第一数据集用于训练所述第二模型或所述第二模型的部分模型,或用于表征适用的模型或功能或特性,或用于定义性能指标,或用于定义性能要求。
- 根据权利要求1至10之中任一项所述的模型处理方法,其中,所述至少两个第一模型或所述第二模型或所述第二模型的部分模型的以下至少一项是协议预定义的:模型基本架构;模型结构;模型各层参数;神经元系数;模型复杂度;模型大小;激活函数;量化方式。
- 根据权利要求1至11之中任一项所述的模型处理方法,其中,所述至少两个第一模型或所述第二模型或所述第二模型的部分模型以下至少一项相同:模型基本架构;模型结构;激活函数;模型复杂度;模型复杂度的最大值;模型复杂度的最小值;模型大小;模型大小的最大值、模型大小的最小值。
- 根据权利要求1至12之中任一项所述的模型处理方法,其中,所述第二模型或所述第二模型的部分模型与所述至少两个第一模型的输入格式或输出格式相同。
- 根据权利要求1至13之中任一项所述的模型处理方法,其中,所述第二模型或所述第二模型的部分模型或所述第一数据集用于以下至少一项:参考信号的处理;信道信号的传输;信道信号的解调;信道状态信息的获取;波束管理;信道预测;信道或信源编译码;干扰抑制;定位;高层业务或参数的预测;高层业务或参数的管理;控制信令的解析。
- 根据权利要求1至14之中任一项所述的模型处理方法,其中,所述通信设备获得至少两个第一模型,包括:所述通信设备获得至少两个第一模型的部分模型;所述通信设备利用所述至少两个第一模型的部分模型和第三数据集,得到至少两个第一模型。
- 一种模型处理装置,其中,包括:第一获得模块,用于获得至少两个第一模型;第二获得模块,用于利用所述至少两个第一模型,得到第二模型或第二模型的部分模型或第一数据集。
- 根据权利要求16所述的模型处理装置,其中,所述第二获得模块,包括:第一获得子模块,用于获得第二数据集;第二获得子模块,用于基于所述第二数据集中的至少一个数据得到至少一个模型输入;第三获得子模块,用于将所述至少一个模型输入分别输入到每个第一模型中,得到每个模型输入对应的模型输出或模型中间量;第四获得子模块,用于基于每个模型输入,或每个模型输入对应的模型输出或模型中间量,得到第二模型或第二模型的部分模型或第一数据集。
- 根据权利要求17所述的模型处理装置,其中,所述第四获得子模块,具体用于:基于每个模型输入,或每个模型输入对应的模型输出或模型中间量,确定每个模型输入对应的模型最终输出或模型最终中间量;利用所述至少一个模型输入,或每个模型输入对应的模型最终输出或模型最终中间量,得到第二模型或第二模型的部分模型或第一数据集。
- 根据权利要求18所述的模型处理装置,其中,针对每个模型输入,当前模型输入对应的模型最终输出基于以下至少一项确定:所述当前模型输入;对所述当前模型输入对应的模型输出集合进行运算得到的第一结果;所述当前模型输入对应的模型输出集合中的第一模型输出,所述第一模型输出与所述当前模型输入对应的输出标签的相似度最高;其中,所述模型输出集合包括:将所述当前模型输入分别输入到每个第一模型后,得到的模型输出。
- 根据权利要求18或19所述的模型处理装置,其中,针对每个模型输入,当前模型输入对应的模型最终中间量基于以下至少一项确定:对所述当前模型输入对应的模型中间量集合进行运算得到的第二结果;对所述当前模型输入对应的模型中间量集合以及模型输出集合进行运算得到的第三结果;所述模型中间量集合中的第一模型中间量,所述第一模型中间量与第二模型输出是利用同一第一模型得到的,所述第二模型输出为所述当前模型输入对应的模型输出集合中的一个模型输出,所述第二模型输出与所述当前模型输入对应的输出标签的相似度最高;其中,所述模型中间量集合包括:将所述当前模型输入分别输入到每个第一模型后,得到的模型中间量;所述模型输出集合包括:将所述当前模型输入分别输入到每个第一模型后,得到的模型输出。
- 根据权利要求18至20之中任一项所述的模型处理装置,其中,在所述第二模型为单边模型的情况下,所述第二模型的输入为所述模型输入,所述第二模型的输出为所述模型最终输出。
- 根据权利要求18至20之中任一项所述的模型处理装置,其中,在所述第二模型为双边模型的情况下,所述第二模型的编码器的输入为所述模型输入,所述第二模型的编码器的输出或所述第二模型的译码器的输入为所述模型最终中间量,所述第二模型的译码器的输出为所述模型最终输出。
- 根据权利要求17至22之中任一项所述的模型处理装置,其中,所述第一获得模块,具体用于:通过以下至少一种方式获得第二数据集:接收参考信号,得到参考信号信息;协议预定义;接收其他通信设备传输的数据。
- 根据权利要求16所述的模型处理装置,其中,所述第二获得模块,具体用于:将所述至少两个第一模型组合为第二模型或第二模型的部分模型。
- 根据权利要求16至24之中任一项所述的模型处理装置,其中,所述第一数据集用于训练所述第二模型或第二模型的部分模型,或用于表征适用的模型或功能或特性,或用于定义性能指标,或用于定义性能要求。
- 根据权利要求16至25之中任一项所述的模型处理装置,其中,所述至少两个第一模型或所述第二模型或所述第二模型的部分模型的以下至少一项是协议预定义的:模型基本架构;模型结构;模型各层参数;神经元系数;模型复杂度;模型大小;激活函数;量化方式。
- 根据权利要求16至26之中任一项所述的模型处理装置,其中,所述至少两个第一模型或所述第二模型或所述第二模型的部分模型以下至少一项相同:模型基本架构;模型结构;激活函数;模型复杂度;模型复杂度的最大值;模型复杂度的最小值;模型大小;模型大小的最大值、模型大小的最小值。
- 根据权利要求16至27之中任一项所述的模型处理装置,其中,所述第二模型或所述第二模型的部分模型与所述至少两个第一模型的输入格式或输出格式相同。
- 根据权利要求16至28之中任一项所述的模型处理装置,其中,所述第二模型或所述第二模型的部分模型或所述第一数据集用于以下至少一项:参考信号的处理;信道信号的传输;信道信号的解调;信道状态信息的获取;波束管理;信道预测;信道或信源编译码;干扰抑制;定位;高层业务或参数的预测;高层业务或参数的管理;控制信令的解析。
- 根据权利要求16至29之中任一项所述的模型处理装置,其中,所述第一获得模块,具体用于:获得至少两个第一模型的部分模型;利用所述至少两个第一模型的部分模型和第三数据集,得到至少两个第一模型。
- 一种通信设备,其中,包括处理器和存储器,所述存储器存储可在所述处理器上运行的程序或指令,所述程序或指令被所述处理器执行时实现如权利要求1至15之中任一项所述的模型处理方法的步骤。
- 一种可读存储介质,其中,所述可读存储介质上存储程序或指令,所述程序或指令被处理器执行时实现如权利要求1至15之中任一项所述的模型处理方法的步骤。
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN202410361159.X | 2024-03-27 | ||
| CN202410361159.XA CN120730324A (zh) | 2024-03-27 | 2024-03-27 | 模型处理方法、装置、通信设备及存储介质 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2025201273A1 true WO2025201273A1 (zh) | 2025-10-02 |
Family
ID=97165055
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/CN2025/084528 Pending WO2025201273A1 (zh) | 2024-03-27 | 2025-03-24 | 模型处理方法、装置、通信设备及存储介质 |
Country Status (2)
| Country | Link |
|---|---|
| CN (1) | CN120730324A (zh) |
| WO (1) | WO2025201273A1 (zh) |
Citations (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN115964632A (zh) * | 2021-05-31 | 2023-04-14 | 华为云计算技术有限公司 | 构建ai集成模型的方法、ai集成模型的推理方法及装置 |
| CN116112366A (zh) * | 2021-11-10 | 2023-05-12 | 中国移动通信有限公司研究院 | 数据处理方法及装置、设备、存储介质 |
| CN117390440A (zh) * | 2022-06-30 | 2024-01-12 | 维沃移动通信有限公司 | 人工智能模型训练方法、装置、网络侧设备、终端及介质 |
| CN117634613A (zh) * | 2022-08-12 | 2024-03-01 | 华为云计算技术有限公司 | 一种模型管理方法及相关设备 |
-
2024
- 2024-03-27 CN CN202410361159.XA patent/CN120730324A/zh active Pending
-
2025
- 2025-03-24 WO PCT/CN2025/084528 patent/WO2025201273A1/zh active Pending
Patent Citations (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN115964632A (zh) * | 2021-05-31 | 2023-04-14 | 华为云计算技术有限公司 | 构建ai集成模型的方法、ai集成模型的推理方法及装置 |
| CN116112366A (zh) * | 2021-11-10 | 2023-05-12 | 中国移动通信有限公司研究院 | 数据处理方法及装置、设备、存储介质 |
| CN117390440A (zh) * | 2022-06-30 | 2024-01-12 | 维沃移动通信有限公司 | 人工智能模型训练方法、装置、网络侧设备、终端及介质 |
| CN117634613A (zh) * | 2022-08-12 | 2024-03-01 | 华为云计算技术有限公司 | 一种模型管理方法及相关设备 |
Non-Patent Citations (1)
| Title |
|---|
| ANONYMOUS: "Easy to Understand – Model Ensemble (Multi-Model) Explanation (Algorithm Case Studies)", MANTCH, 31 December 2018 (2018-12-31), pages 1 - 7, XP093359659, Retrieved from the Internet <URL:https://www.cnblogs.com/mantch/p/10203143.html> * |
Also Published As
| Publication number | Publication date |
|---|---|
| CN120730324A (zh) | 2025-09-30 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| US20250310756A1 (en) | Ai computing power reporting method, terminal, and network-side device | |
| US20250227507A1 (en) | Ai model processing method and apparatus, and communication device | |
| WO2024140422A1 (zh) | Ai单元的性能监测方法、终端及网络侧设备 | |
| US20250211363A1 (en) | Cqi transmission method and apparatus, terminal, and network-side device | |
| WO2024067280A1 (zh) | 更新ai模型参数的方法、装置及通信设备 | |
| WO2025209453A1 (zh) | 通信方法、装置、终端、网络侧设备、介质及产品 | |
| WO2025201273A1 (zh) | 模型处理方法、装置、通信设备及存储介质 | |
| CN114501353B (zh) | 通信信息的发送、接收方法及通信设备 | |
| WO2023185978A1 (zh) | 信道特征信息上报及恢复方法、终端和网络侧设备 | |
| US20260031867A1 (en) | Information processing method, information processing apparatus, terminal and network side device | |
| CN121510039A (zh) | 关联标识的关联范围的确定方法、装置、通信设备及介质 | |
| CN121485735A (zh) | Csi上报、接收方法、装置、设备、可读存储介质及计算机程序产品 | |
| US20260051939A1 (en) | Method and apparatus for reporting quantity of cpus, method and apparatus for receiving quantity of cpus, terminal, and network side device | |
| WO2025261336A1 (zh) | 无线通信方法、装置及设备 | |
| WO2024037380A1 (zh) | 信道信息处理方法、装置、通信设备及存储介质 | |
| WO2024217495A1 (zh) | 信息处理方法、信息处理装置、终端及网络侧设备 | |
| WO2026021295A1 (zh) | 模型参数传输方法、装置及通信设备 | |
| WO2026032266A1 (zh) | 无线通信方法、装置及设备 | |
| CN117692032A (zh) | 信息传输方法、装置、设备、系统及存储介质 | |
| WO2025146039A1 (zh) | 信息处理方法、装置及通信设备 | |
| WO2025232707A9 (zh) | Csi数据处理方法、装置、终端、网络侧设备、介质及产品 | |
| CN120786391A (zh) | 模型有效性的确定方法及相关装置 | |
| CN121510038A (zh) | Ai模型传输的指示方法、装置、通信设备及存储介质 | |
| CN122002305A (zh) | Ai模型管理方法、装置及通信设备 | |
| CN122001958A (zh) | 传输方法、装置、设备、可读存储介质及计算机程序产品 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 25776117 Country of ref document: EP Kind code of ref document: A1 |