EP4732462A1 - Machine learning -based adjustment of beam failure recovery related parameters, and related devices, methods and computer programs - Google Patents
Machine learning -based adjustment of beam failure recovery related parameters, and related devices, methods and computer programsInfo
- Publication number
- EP4732462A1 EP4732462A1 EP23735259.6A EP23735259A EP4732462A1 EP 4732462 A1 EP4732462 A1 EP 4732462A1 EP 23735259 A EP23735259 A EP 23735259A EP 4732462 A1 EP4732462 A1 EP 4732462A1
- Authority
- EP
- European Patent Office
- Prior art keywords
- bfr
- configuration parameters
- terminal device
- model
- parameters
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Pending
Links
Classifications
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04B—TRANSMISSION
- H04B7/00—Radio transmission systems, i.e. using radiation field
- H04B7/02—Diversity systems; Multi-antenna system, i.e. transmission or reception using multiple antennas
- H04B7/04—Diversity systems; Multi-antenna system, i.e. transmission or reception using multiple antennas using two or more spaced independent antennas
- H04B7/06—Diversity systems; Multi-antenna system, i.e. transmission or reception using multiple antennas using two or more spaced independent antennas at the transmitting station
- H04B7/0686—Hybrid systems, i.e. switching and simultaneous transmission
- H04B7/0695—Hybrid systems, i.e. switching and simultaneous transmission using beam selection
- H04B7/06952—Selecting one or more beams from a plurality of beams, e.g. beam training, management or sweeping
- H04B7/06964—Re-selection of one or more beams after beam failure
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04W—WIRELESS COMMUNICATION NETWORKS
- H04W36/00—Hand-off or reselection arrangements
- H04W36/0005—Control or signalling for completing the hand-off
- H04W36/0083—Determination of parameters used for hand-off, e.g. generation or modification of neighbour cell lists
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04W—WIRELESS COMMUNICATION NETWORKS
- H04W36/00—Hand-off or reselection arrangements
- H04W36/24—Reselection being triggered by specific parameters
- H04W36/30—Reselection being triggered by specific parameters by measured or perceived connection quality data
- H04W36/305—Handover due to radio link failure
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04W—WIRELESS COMMUNICATION NETWORKS
- H04W76/00—Connection management
- H04W76/10—Connection setup
- H04W76/19—Connection re-establishment
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04W—WIRELESS COMMUNICATION NETWORKS
- H04W36/00—Hand-off or reselection arrangements
- H04W36/08—Reselecting an access point
- H04W36/085—Reselecting an access point involving beams of access points
Landscapes
- Engineering & Computer Science (AREA)
- Computer Networks & Wireless Communication (AREA)
- Signal Processing (AREA)
- Mobile Radio Communication Systems (AREA)
Abstract
Devices, methods and computer programs for machine learning -based adjustment of beam failure recovery related parameters are disclosed. At least some example embodiments may allow fine tuning the BFR parameters to improve BFR performance for better connectivity, and may be useful for, e. g., applications with low latency requirements.
Description
MACHINE LEARNING -BASED ADJUSTMENT OF BEAM FAILURE RECOVERY RELATED PARAMETERS, AND RELATED DEVICES, METHODS AND COMPUTER PROGRAMS
TECHNICAL FIELD
The disclosure relates generally to communications and, more particularly but not exclusively, to machine learning -based adjustment of beam failure recovery related parameters, as well as related devices, methods and computer programs .
BACKGROUND
In fifth generation (5G) wireless networks, a terminal device, such as a user equipment (UE) may be connected to one or more beams for communication. That may mean that the terminal device is receiving and/or transmitting data with multiple beams during a communication session with the network and/or another device. When a connection to a beam is lost, e.g., due to signal level degradation, and a beam failure is detected, the terminal device may start a search for a new candidate beam using a beam failure recovery (BFR) procedure. The terminal device may successfully complete the BFR procedure, or the BFR procedure may fail in which case the terminal device may declare a radio link failure (RLF) .
E.g., radio resource control (RRC) protocol may be used to configure BFR parameters to assist the terminal device in the BFR procedure. However, currently the BFR parameters are not updated frequently once they have been configured for the terminal device. At least in some situations, this may create issues since one set of BFR parameters may not work for all situations, such as different radio conditions and different geographical conditions .
Accordingly, at least in some situations , there may be a need for dynamically adj usting beam failure recovery related parameters .
SUMMARY
The scope of protection sought for various example embodiments of the invention is set out by the independent claims . The example embodiments and features , i f any, described in this speci fication that do not fall under the scope o f the independent claims are to be interpreted as examples useful for understanding various example embodiments of the invention .
An example embodiment of an apparatus comprises at least one processor, and at least one memory storing instructions that , when executed by the at least one processor, cause the apparatus at least to perform in response to a beam failure recovery, BFR, procedure being initiated by a terminal device in order to search for a new candidate beam after a detected failure of a current beam, determining a set of one or more BFR configuration parameters for use in the BFR procedure . The instructions , when executed by the at least one processor, further cause the apparatus at least to perform providing the determined set of the one or more BFR configuration parameters to the terminal device . The determining of the set of the one or more BFR configuration parameters comprises applying a machine learning, ML, model to a set of model input parameters in order to obtain the one or more BFR configuration parameters for use in the BFR procedure .
In an example embodiment , alternatively or in addition to the above-described example embodiments , the ML model is conf igured for supervised learning and executable to determine BFR success probabilities for at least two di f ferent sets of the one or more BFR configuration parameters based on the set of the model input parameters . The ML model is further executable to
select one of the at least two di f ferent sets of the one or more BFR configuration parameters as a preferred set to be provided to the terminal device .
In an example embodiment , alternatively or in addition to the above-described example embodiments , the ML model is further executable to perform the selection by : determining candidate sets of the at least two di f ferent sets of the one or more BFR configuration parameters having one of a highest BFR success probability or a BFR success probability that di f fers from the highest BFR success probability by less than a minimum di f ference threshold, and out of the candidate sets , assigning a set of the one or more BFR configuration parameters with a lowest value of a decision model input parameter of the set of the model input parameters as the preferred set .
In an example embodiment , alternatively or in addition to the above-described example embodiments , the decision model input parameter comprises a transmission power ramping step, a beam failure recovery timer, a maximum number of retransmissions , or a response monitoring time window for a special cell , SpCell , beam failure recovery using contention- free random-access resources .
In an example embodiment , alternatively or in addition to the above-described example embodiments , the decision model input parameter comprises a minimum power ramping step .
In an example embodiment , alternatively or in addition to the above-described example embodiments , the BFR success probability comprises a ratio of a number of times a BFR procedure is success ful for a given set of model input parameters to a number of times the BFR procedure was attempted for the given set of model input parameters .
In an example embodiment , alternatively or in addition to the above-described example embodiments , the
instructions , when executed by the at least one processor, further cause the apparatus to perform training of the ML model based on training data comprising real-time information on BFR attempts .
In an example embodiment , alternatively or in addition to the above-described example embodiments , the ML model is configured for reinforcement learning and executable to determine a set of the one or more BFR configuration parameters to be provided to the terminal device out of at least two di f ferent sets of the one or more BFR configuration parameters , based on di f ferent combinations of model input parameters in the set of the model input parameters .
In an example embodiment , alternatively or in addition to the above-described example embodiments , the model input parameters comprise at least one of a signal to interference and noise ratio of a candidate beam, a location of the terminal device, or a mobility state of the terminal device .
In an example embodiment , alternatively or in addition to the above-described example embodiments , the model input parameters further comprise a random-access channel , RACH, load .
In an example embodiment , alternatively or in addition to the above-described example embodiments , the BFR configuration parameters comprise at least one of a beam failure recovery timer, a transmission power ramping step, a maximum number of retransmissions , or a response monitoring time window for a special cel l , SpCell , beam failure recovery using contention- free random-access resources .
In an example embodiment , alternatively or in addition to the above-described example embodiments , the instructions , when executed by the at least one processor, further cause the apparatus to perform triggering a retraining o f the ML model in response to
a number of success ful BFR procedures falling below a retraining threshold .
In an example embodiment , alternatively or in addition to the above-described example embodiments , the success ful BFR procedure comprises the terminal device success fully completing a beam switch after the BFR procedure is attempted .
In an example embodiment , alternatively or in addition to the above-described example embodiments , the apparatus is comprised in a network side device or the terminal device .
An example embodiment of a method comprises in response to a beam failure recovery, BFR, procedure being initiated by a terminal device in order to search for a new candidate beam after a detected failure of a current beam, determining, by an apparatus , a set of one or more BFR configuration parameters for use in the BFR procedure . The method further comprises providing, by the apparatus , the determined set of the one or more BFR configuration parameters to the terminal device . The determining of the set of the one or more BFR configuration parameters comprises applying a machine learning, ML, model to a set of model input parameters in order to obtain the one or more BFR configuration parameters for use in the BFR procedure .
In an example embodiment , alternatively or in addition to the above-described example embodiments , the ML model is conf igured for supervised learning and executable to determine BFR success probabilities for at least two di f ferent sets of the one or more BFR configuration parameters based on the set of the model input parameters . The ML model is further executable to select one of the at least two di f ferent sets of the one or more BFR configuration parameters as a preferred set to be provided to the terminal device .
In an example embodiment , alternatively or in addition to the above-described example embodiments , the
ML model is further executable to perform the selection by : determining candidate sets of the at least two di f ferent sets of the one or more BFR configuration parameters having one of a highest BFR success probability or a BFR success probability that di f fers from the highest BFR success probability by less than a minimum di f ference threshold, and out of the candidate sets , assigning a set of the one or more BFR configuration parameters with a lowest value of a decision model input parameter of the set of the model input parameters as the preferred set .
In an example embodiment , alternatively or in addition to the above-described example embodiments , the decision model input parameter comprises a transmission power ramping step, a beam failure recovery timer, a maximum number of retransmissions , or a response monitoring time window for a special cell , SpCell , beam failure recovery using contention- free random-access resources .
In an example embodiment , alternatively or in addition to the above-described example embodiments , the decision model input parameter comprises a minimum power ramping step .
In an example embodiment , alternatively or in addition to the above-described example embodiments , the BFR success probability comprises a ratio of a number of times a BFR procedure is success ful for a given set of model input parameters to a number of times the BFR procedure was attempted for the given set of model input parameters .
In an example embodiment , alternatively or in addition to the above-described example embodiments , the method further comprises training of the ML model , by the apparatus , based on training data comprising realtime information on BFR attempts .
In an example embodiment , alternatively or in addition to the above-described example embodiments , the
ML model is configured for reinforcement learning and executable to determine a set of the one or more BFR configuration parameters to be provided to the terminal device out of at least two di f ferent sets of the one or more BFR configuration parameters , based on di f ferent combinations of model input parameters in the set of the model input parameters .
In an example embodiment , alternatively or in addition to the above-described example embodiments , the model input parameters comprise at least one of a signal to interference and noise ratio of a candidate beam, a location of the terminal device, or a mobility state of the terminal device .
In an example embodiment , alternatively or in addition to the above-described example embodiments , the model input parameters further comprise a random-access channel , RACH, load .
In an example embodiment , alternatively or in addition to the above-described example embodiments , the BFR configuration parameters comprise at least one of a beam failure recovery timer, a transmission power ramping step, a maximum number of retransmissions , or a response monitoring time window for a special cel l , SpCell , beam failure recovery using contention- free random-access resources .
In an example embodiment , alternatively or in addition to the above-described example embodiments , the method further comprises triggering, by the apparatus , a retraining of the ML model in response to a number of success ful BFR procedures falling below a retraining threshold .
In an example embodiment , alternatively or in addition to the above-described example embodiments , the success ful BFR procedure comprises the terminal device success fully completing a beam switch after the BFR procedure is attempted .
In an example embodiment , alternatively or in addition to the above-described example embodiments , the apparatus is comprised in a network s ide device or the terminal device .
A computer program comprising instructions for causing an apparatus to perform at least the following : in response to a beam failure recovery, BFR, procedure being initiated by a terminal device in order to search for a new candidate beam after a detected failure of a current beam, determining a set of one or more BFR configuration parameters for use in the BFR procedure ; and providing the determined set of the one or more BFR configuration parameters to the terminal device . The determining of the set of the one or more BFR configuration parameters comprises applying a machine learning, ML, model to a set of model input parameters in order to obtain the one or more BFR configuration parameters for use in the BFR procedure .
DESCRIPTION OF THE DRAWINGS
The accompanying drawings , which are included to provide a further understanding of the embodiments and constitute a part of this specification, illustrate embodiments and together with the description help to explain the principles of the embodiments . In the drawings :
FIG . 1 shows an example embodiment of the subj ect matter described herein illustrating an example system, where various embodiments of the present disclosure may be implemented;
FIG . 2 shows an example embodiment of the subj ect matter described herein illustrating an apparatus ;
FIG . 3 shows an example embodiment of the subj ect matter described herein illustrating disclosed Supervised learning training;
FIG . 4 shows an example embodiment of the subj ect matter described herein a disclosed feedback mechanism to adj ust ML model parameters ;
FIG . 5 shows an example embodiment of the subj ect matter described herein illustrating disclosed Location-based BFR parameter assignment ;
FIG . 6 shows an example embodiment of the subj ect matter described herein illustrating examples of disclosed exploration and exploitation phases for BFR;
FIG . 7 shows an example embodiment of the subj ect matter described herein illustrating a disclosed method of ML model deployment at the network side ; and
FIG . 8 shows another example embodiment of the subj ect matter described herein illustrating a disclosed method of ML model deployment at the terminal device side .
Like reference numerals are used to designate like parts in the accompanying drawings .
DETAILED DESCRIPTION
Reference will now be made in detail to embodiments , examples of which are illustrated in the accompanying drawings . The detailed description provided below in connection with the appended drawings is intended as a description of the present examples and is not intended to represent the only forms in which the present example may be constructed or utili zed . The description sets forth the functions of the example and the sequence of steps for constructing and operating the example . However, the same or equivalent functions and sequences may be accomplished by di f ferent examples .
Fig . 1 illustrates an example system 100 , where various embodiments of the present disclosure may be implemented . The system 100 may comprise a fi fth generation ( 5G) new radio (NR) network or a network beyond 5G wireless networks , 110 . An example representation of the system 100 is shown depicting a network side device
120 and a terminal device 130. At least in some embodiments, the network 110 may comprise one or more massive machine-to-machine (M2M) network (s) , massive machine type communications (mMTC) network (s) , internet of things (loT) network(s) , industrial internet-of-things (IIoT) network(s) , enhanced mobile broadband (eMBB) network (s) , ultra-reliable low-latency communication (URLLC) network(s) , and/or the like. In other words, the network 110 may be configured to serve diverse service types and/or use cases, and it may logically be seen as comprising one or more networks.
The terminal device 130 may include, e.g., a mobile phone, a smartphone, a tablet computer, a smart watch, or any hand-held, portable and/or wearable device. The terminal device 130 may also be referred to as a user equipment (UE) . The network side device 120 may comprise, e.g., a base station. The base station may include, e.g., any device suitable for providing an air interface for terminal devices to connect to a wireless network via wireless transmissions. Furthermore, the network side device 120 may comprise an apparatus 200 of Fig. 2 including a machine learning (ML) model 250. Alternatively/additionally, the terminal device 130 may comprise the apparatus 200 (not shown in Fig. 1) .
In the following example embodiments, it may be possible to train one ML model with a specific architecture, then derive another ML model from that using processes such as compilation, pruning, quantization or distillation. The ML model may be executed using any suitable apparatus, for example a CPU, GPU, ASIC, FPGA, compute-in-memory, analog, or digital, or optical apparatus. It is also possible to execute the ML model in an apparatus that combines features from any number of these, for instance digital-optical or analog-digital hybrids. In some examples, weights and required computations in these systems may be programmed to correspond to the ML model. In some examples, the apparatus may be
designed and manufactured so as to perform the task defined by the ML model so that the apparatus is configured to perform the task when it is manufactured without the apparatus being programmable as such .
In the following, various example embodiments will be discussed . At least some of these example embodiments described herein may allow machine learning - based adj ustment of beam failure recovery related parameters .
Furthermore , at least some of the example embodiments described herein may allow ef fectively optimi zing the BFR procedure without creating excessive interference by considering several input parameters , including, e . g . , terminal device location and traj ectory, a mobi lity state , a candidate beam signal to interference and noise ratio ( S INR) , and/or a probability of the BFR .
Furthermore , at least some of the example embodiments described herein may allow fine tuning the BFR parameters to improve BFR performance for better connectivity, and may be useful for , e . g . , applications with low latency requirements .
Furthermore , at least some of the example embodiments described herein may allow dynamically adj usting the BFR parameters after a beam fai lure i s declared and a candidate beam is selected .
As will be discussed in more detail below, at least some of the example embodiments described herein may use supervised learning while other example embodiments described herein may use reinforcement learning . Both supervised learning -based example embodiments and reinforcement learning -based example embodiments may use a same set of input and output parameters .
Fig . 2 is a block diagram of the apparatus 200 , in accordance with an example embodiment . For example , the apparatus 200 may be comprised in the network side device 120 or the terminal device 130 .
The apparatus 200 comprises one or more processors 202 and one or more memories 204 that comprise computer program code . The apparatus 200 may also include other elements , such as a transceiver 206 configured to enable the apparatus 200 to transmit and/or receive information to/ from other devices , as well as other elements not shown in Fig . 2 . In one example , the apparatus 200 may use the transceiver 206 to transmit or receive signaling information and data in accordance with at least one cellular communication protocol . The transceiver 206 may be configured to provide at least one wireless radio connection, such as for example a 3GPP mobile broadband connection ( e . g . , 5G or beyond) . The transceiver 206 may comprise , or be configured to be coupled to , at least one antenna to transmit and/or receive radio frequency signals .
Although the apparatus 200 is depicted to include only one processor 202 , the apparatus 200 may include more processors . In an embodiment , the memory 204 i s capable of storing instructions , such as an operating system and/or various applications . Furthermore , the memory 204 may include a storage that may be used to store , e . g . , at least some of the information and data used in the disclosed embodiments , such as the ML model 250 described in more detail below .
Furthermore , the processor 202 is capable of executing the stored instructions . In an embodiment , the processor 202 may be embodied as a multi-core processor, a single core processor, or a combination of one or more multi-core processors and one or more s ingle core processors . For example , the processor 202 may be embodied as one or more of various processing devices , such as a coprocessor, a microprocessor, a controller, a digital signal processor ( DSP ) , a processing circuitry with or without an accompanying DSP, or various other processing devices including integrated circuits such as , for example , an application specific integrated circuit
(ASIC) , a field programmable gate array (FPGA) , a microcontroller unit (MCU) , a hardware accelerator, a special-purpose computer chip, a neural network (NN) chip, an artificial intelligence (Al) accelerator, a tensor processing unit (TPU) , a neural processing unit (NPU) , or the like. In an embodiment, the processor 202 may be configured to execute hard-coded functionality. In an embodiment, the processor 202 is embodied as an executor of software instructions, wherein the instructions may specifically configure the processor 202 to perform the algorithms and/or operations described herein when the instructions are executed.
The memory 204 may be embodied as one or more volatile memory devices, one or more non-volatile memory devices, and/or a combination of one or more volatile memory devices and non-volatile memory devices. For example, the memory 204 may be embodied as semiconductor memories (such as mask ROM, PROM (programmable ROM) , EPROM (erasable PROM) , flash ROM, RAM (random access memory) , etc . ) .
When executed by the at least one processor 202, instructions stored in the at least one memory 204 cause the apparatus 200 at least to perform determining a set of one or more beam failure recovery (BFR) configuration parameters for use in a BFR procedure, in response to the BFR procedure being initiated by the terminal device 130 in order to search for a new candidate beam after a detected failure of a current beam.
For example, the BFR configuration parameters may comprise a beam failure recovery timer or "beamFailureRecoveryTimer" (possible values including, e.g., 10, 20, 40, 60, 80, 100, 150, 200 milliseconds (ms) ) , a transmission power ramping step or "powerRamp- ingStep" (possible values including, e.g., 0, 2, 4, 6 decibels (dB) ) , a maximum number of retransmissions or "preambleTransMax" (possible values including, e.g., 3,
4, 5, 6, 7, 8, 10, 20, 50, 100) , and/or a response monitoring time window or "ra-ResponseWindow" (possible values including, e.g., 1, 2, 4, 8, 10, 20, 40, 80 slots) for a special cell, SpCell, beam failure recovery using contention-free random-access resources.
In other words, the BFR configuration parameters function as output parameters for ML inference in supervised learning -based example embodiments detailed below.
The instructions, when executed by the at least one processor 202, further cause the apparatus 200 at least to perform providing the determined set of the one or more BFR configuration parameters to the terminal device 130.
The determining of the set of the one or more BFR configuration parameters comprises applying the ML model 250 to a set of model input parameters in order to obtain the one or more BFR configuration parameters for use in the BFR procedure.
For example, the model input parameters may comprise a signal to interference and noise ratio (SINR) of a candidate beam (a recent value may be available at the terminal device 130, but a last reported candidate beam SINR may be used when the ML model 250 is hosted at the network side device 120) , a location of the terminal device 130, and/or a mobility state of the terminal device 130 (may be predicted, e.g., based on location history of the terminal device 130) .
In other words, the model input parameters function as input parameters for the ML inference in the supervised learning -based example embodiments detailed below .
At least in some embodiments, the model input parameters may further comprise a random-access channel (RACH) load.
At least in some embodiments , the ML model 250 may be configured for supervised learning and be executable to determine BFR success probabilities for at least two di f ferent sets of the one or more BFR configuration parameters based on the set of the model input parameters . The ML model 250 may further be executable to select one of the at least two di f ferent sets of the one or more BFR configuration parameters as a preferred set to be provided to the terminal device 130 .
At least in some embodiments , the ML model 250 may be further executable to perform the selection by : determining candidate sets of the at least two di f ferent sets of the one or more BFR configuration parameters having one of a highest BFR success probability or a BFR success probability that di f fers from the highest BFR success probability by less than a minimum di f ference threshold, and out of the candidate sets , assigning a set of the one or more BFR configuration parameters with a lowest value of a decision model input parameter of the set o f the model input parameters as the preferred set .
For example , the BFR success probability may comprise a ratio of a number of times a BFR procedure is success ful for a given set of model input parameters to a number of times the BFR procedure was attempted for the given set of model input parameters .
For example , the decision model input parameter may comprise a transmission power ramping step, a beam failure recovery timer, a maximum number of retransmissions , or a response monitoring time window for a special cell , SpCell , beam failure recovery using contention- free random-access resources . Additionally/alter- natively, the decision model input parameter may comprise a minimum power ramping step .
At least in some embodiments , the instructions , when executed by the at least one processor 202 , may further cause the apparatus 200 to perform training of
the ML model 250 based on training data comprising realtime information on BFR attempts.
In other words, the supervised learning -based example embodiments may comprise at least some of the following parts.
BFR success probability computation: based on fixed input parameters, such as the location of the terminal device 130, the mobility state of the terminal device 130, the candidate beam SINR measured by the terminal device 130, and on BFR parameters (e.g., recommended values of beamFailureRecoveryTimer , powerRamp- ingStep, preambleTransMax, and ra-ResponseWindow) , the BFR success probability may be computed for different sets of BFR parameters. This may be done empirically by using real data.
Thus, based on, e.g., terminal device 130 location, mobility state, and candidate beam SINR, the ML model 250 may compute the BFR success probability using actual measurements on BFR. E.g., for a given set of input parameters, the BFR success probability P(BFR) may be computed empirically by:
P BFR~)
Number of times BFR is successful given a particular set of parameters applied Number of times BFR was attempted for a given set of parameters
A successful BFR may be defined as, e.g., the terminal device 130 successfully completing a beam switch after the BFR is attempted. The network may be aware of this quantity through BFR request messages received from the terminal device 130. A BFR attempt may be counted as a beam failure declaration at the terminal device 130 side. The terminal device 130 may report the latter to the network.
Preprocessing of labelling: this may be done because many different BFR parameters may result in nearly the same BFR success probability. At least in some situations, the BFR success probability may be
large for parameters that consume more resources. For example, the BFR success probability for a large Power- Rampingstep may be large but it may create interference to other terminal devices when a RACH load is high, and same logic may hold for preambleTransMax where a large number of transmissions may allow for a larger BFR success probability. In this case, the BFR success probability may be labelled with BFR parameters that are the least harmful or resource-consuming without impacting the BFR success probability much. This may be achieved, e.g., via taking into account terminal device 130 specific information, such as the terminal device location and/or the terminal device mobility state.
The above labelling may be preprocessed before training. Using extreme values of the BFR configuration parameters for a fixed set of model input parameters may provide a higher BFR success probability but it may consume resources as well. E.g., a large value of power- Rampingstep may cause harmful interference to other terminal devices, particularly in high RACH load scenarios. Similarly, setting a large value of preambleTransMax may not be required when the BFR success probability is very low and an RLF cannot be avoided anyway. In this case, the data from the above probability computation may be preprocessed to label the BFR configuration parameters for ML model 250 training. Therefore, a preprocessing step may include at least some of the following.
1. for a fixed set of model input parameters, empirically compute the BFR success probability as above .
2 . select the parameters that provide the largest BFR success probability.
3. if the BFR success probability P (BFR) for more than one set of parameters is similar, choose the set of parameters that require a minimum powerRamping- step. This may be done, e.g., by:
If P (BFR1) -P (BFR2) <y choose the BFR configuration parameter set with minimum powerRampingstep in which y represents a minimum difference threshold. This criterion may be extended to the scenario where more than two sets of BFR configuration parameters provide similar P (BFR) by selecting the P(BFR1) as the largest probability and P(BFR2) as the second-largest probability.
4. a similar criterion for preambleTransMax and ra-ResponseWindow may be devised if P(BFR)<s where s is a configured value. A typical value of s may be 0.05 when P (BFR) is really small and setting large values for preambleTransMax and ra-ResponseWindow is not necessary as there may be a high probability of RLF and it may be declared earlier to avoid wasting time on the BFR procedure .
An example is provided in Table 1 below where y=0.01, X represents a set of model input parameters, Y represents a set of BFR configuration parameters, and a justification for the selection is provided in the last column .
Table 1
As an example embodiment , more than one ML model may be trained - one for each y based, e . g . , on the labelling described above . As discussed below in more detail , the network may decide to use dif ferent ML models based on network conditions .
Training and inference : based on the input parameters labelled above , training for supervised learning may be performed to get the ML model 250 parameters . Then, inference may be performed on the BFR configuration parameters based on the model input parameters .
For example , the ML model 250 may be trained of fline at the network side based on real data on BFR attempts , as illustrated in diagram 300 of Fig . 3 . Based on the training data, a supervised learning algorithm like a deep neural network 303 may be applied to train the ML model 250 with model input parameters 301 and a desired output 302 of the supervised learning . A mean square error (MSE ) may be used as a loss function for a regression model , and may be used to train the ML model 250 :
in which the right-hand side is used to describe the di f ference between a true output value and a hypothesis , and N represents a number of samples . The ML model 250 may be trained such that for known labelled
BFR configuration parameters (computed above) , it minimizes a loss function minimum square error (MSE) and outputs ML configuration parameters used for inference.
At least in some embodiments, the ML model 250 may be deployed in the network or terminal device side and inference on BFR success probability may be performed based on, e.g., availability of model input parameters. However, based on additional information available at the time of inference and feedback logs, the following adjustment may be used.
As illustrated in diagram 400 of Fig. 4, the ML model 250 may be initialized with a set of model input parameters. However, at least in some situations inference may be improved by a continuous feedback mechanism 401, 402, as follows.
A RACH load indicates how many terminal devices may be doing the BFR procedure simultaneously. At the time of inference, the network knows this, and the terminal device 130 may estimate this precisely enough through collision monitoring. If the RACH load is high, a large powerRampingstep may be even more harmful for the system. In this case, the value of y may be adjusted to use an ML model with a different y. For example, if the RACH load is large, a large y implies that the network prefers risking failure for a particular BFR instance to using a very high powerRampingf actor that may cause interference to other terminal devices and thus impact their performance.
Feedback-based ML model retraining: the terminal device 130 may provide feedback to the network, e.g., on BFR decision logs. Based on these logs, if BFR attempts are below a threshold, a process for ML model 250 retraining may be triggered.
Table 2 below discloses some examples to illustrate extreme cases for setting the BFR configuration
parameters . The examples are only intended to demonstrate the importance of the BFR configuration parameters in various scenarios .
When the BFR success probability is high : the ML model 250 may select BFR configuration parameters to avoid a BFR failure . The power ramping setting may depend on y and may take the RACH load into consideration ( e . g . , choose a power ramping at low loads for less impact of interference to other terminal devices ) .
When the BFR success probability is medium : the ML model 250 may select location based BFR configuration parameters ( as illustrated in diagram 500 of Fig . 5 ) to avoid RLFs , limit signalling overheads and interference to other terminal devices . As illustrated in diagram 500 , power ramping in a first RACH transmission may be used for cell edge cases with a low RACH load to reduce signalling overheads and retransmissions . At least in some situations , this may reduce the possibility of a preamble collision and may improve a success probability for the entire system .
When the BFR success probability is low : the ML model 250 may choose BFR configuration parameters to declare RLF quickly without making many BFR attempts .
Table 2
Examples of allowed values of the BFR configuration parameters may include the following:
- BFR Timer [ms] : out of {10, 20, 40, 60, 80, 100, 150, 200},
- power Ramping [dB] : out of {dBO, dB2, dB4, dB6} ,
- preamble Trans Max: out of {n3, n4, n5, n6, n7, n8, nlO, n20, n50, nl00}, and
- ra-Response Window [slots] : out of {1, 2, 4, 8, 10, 20, 40, 80} .
Alternatively, the ML model 250 may be configured for reinforcement learning (RL) and be executable to determine a set of the one or more BFR configuration parameters to be provided to the terminal device 130 out of at least two different sets of the one or more BFR configuration parameters, based on different combinations of model input parameters in the set of the model input parameters.
As compared to the supervised learning -based example embodiments, the reinforcement learning -based example embodiments may take all the model input parameters as their input and produce output parameters (i.e., the BFR configuration parameters) after an exploration phase.
In both the supervised learning -based example embodiments and the reinforcement learning -based example embodiments, an information log of the terminal device 130 may be used to trigger ML model 250 training if the BFR success is below a configured threshold.
Thus, instead of using supervised learning, reinforcement learning may be used for BFR parametrization. In the following, Q-learning is used as an example .
At least some of the following entities may be defined .
A state-space may be defined by different combinations of model input parameters, including terminal device 130 location, RACH load, candidate beam SINR, terminal device 130 mobility state, etc. In this case, the terminal device 130 locations may be adapted by using a bigger tile -based approach in which all the locations in a tile are given one identification, and similar techniques may be applied for other parameters to avoid a state-space explosion. If one of the model input parameters changes, the state changes.
An action: execute the BFR procedure for a set of BFR configuration parameters and observe a reward. An example of the action may include the following three possibilities :
- BFR is successful without power ramping,
- BFR is successful with power ramping (power ramping may be bad for the network due to creating interference) , and
- BFR is unsuccessful.
After execution of the action, an RL system may either remain in a same state if no model input parameter changes or moves to a new state. E.g., a change of location/mobility state may prompt to move into a new state .
A reward function: a goal of an RL agent is to maximize a reward function. An example embodiment of a reward function may include the following:
- BFR is successful without power ramping = 1,
- BFR is successful with power ramping =0,
- BFR is unsuccessful = -1.
An exploration phase (illustrated in diagram 600 of Fig. 6) : In the exploration phase, after observing terminal device state in operation 601, each terminal device 130 may take random actions in each state and observe rewards, operations 602-603. This may help to explore the action space with a high probability. This way, a Q table may be populated with the right actions (BFR configuration parameters in this case) for each state, operation 604.
An exploitation phase (illustrated in diagram 650 of Fig. 6) : in exploitation phase, after observing terminal device state in operation 651, the terminal device 130 may choose the best action in each state based on its Q table, operation 652. The terminal device 130 may still do some exploration with a smaller probability as compared to the exploration phase to further improve its actions over time.
At least in some embodiments, the instructions, when executed by the at least one processor 202, may further cause the apparatus 200 to perform triggering a retraining of the ML model 250 in response to a number of successful BFR procedures falling below a retraining threshold. For example, the successful BFR procedure may comprise the terminal device 130 successfully completing a beam switch after the BFR procedure is attempted.
Fig. 7 illustrates an example signaling diagram of a method 700 of ML model 250 deployment at the network side, in accordance with an example embodiment. Fig. 8 illustrates an example signaling diagram of a method 800 of ML model 250 deployment at the terminal device 130 side, in accordance with an example embodiment.
In response to a BFR procedure being initiated by the terminal device 130 in order to search for a new candidate beam after a detected failure of a current beam, the apparatus 200 determines at operation 705, 806 a set of one or more BFR configuration parameters for use in the BFR procedure. At operation 706, 807, the
apparatus 200 provides the determined set of the one or more BFR configuration parameters to the terminal device 130. The determining 705, 806 of the set of the one or more BFR configuration parameters comprises applying the ML model 250 to a set of model input parameters in order to obtain the one or more BFR configuration parameters for use in the BFR procedure.
In other words, the ML model 250 is deployed at the network side and the model training/retraining is performed at the network side in the example embodiment of the method 700. When the ML model 250 is deployed at the terminal device 130 side (as in the example embodiment of the method 800) , the updated model is sent from the network to the terminal device 130 with extra signalling .
At optional operation 701, the network 120 (e.g., a serving cell (S-Cell) ) may configure the BFR configuration parameters for training purpose.
At optional operation 702, based on these BFR configuration parameters, the terminal device 130 may record the number of BFR attempts and success instances and send them to the network 120 for labelling and training (at operation 703) purposes.
At optional operation 704, the network may send input parameters to a BFR configuration parameter deciding entity (i.e., the apparatus 200) hosted at the network to estimate the BFR configuration parameters using the ML model 250.
Operation 705 is performed as described above.
At operation 706, the selected BFR configuration parameters are sent to the serving cell 120, and then the serving cell 120 may forward them to the terminal device 130, optional operation 707.
At optional operation 708, the terminal device 130 may send the BFR, RLFs, the number of BFR attempts and success instances logs to the network 120 to help in deciding when to retrain the ML model 250.
As described above, the ML model 250 is deployed at the terminal device 130 side in the example embodiment of the method 800.
At optional operation 801, the network 120 configures the BFR configuration parameters for training purpose .
At optional operation 802, based on these BFR configuration parameters, the terminal device 130 may record the number of BFR attempts and success instances and send them to the network 120 for labelling and training (at operation 803) purposes.
At optional operation 804, the network 120 may send a RACH load and an updated ML model 250 to the terminal device 130 (e.g., in an RRC reconfiguration message) . The network 120 may periodically indicate (explicitly or implicitly) which parameters should be input of the ML model 250.
At optional operation 805, the network may send input parameters to the BFR configuration parameter deciding entity (i.e., the apparatus 200) hosted at the terminal device 130 to estimate the BFR configuration parameters using the ML model 250.
Operation 806 is performed as described above.
At operation 807, the selected BFR configuration parameters are sent to the terminal device 130, and then the terminal device 130 may forward them to the network 120 (e.g., in a RACH request) , optional operation 808.
At optional operation 809, the terminal device 130 may send the BFR, RLFs, the number of BFR attempts and success instances logs to the network 120 to help in deciding when to retrain the ML model 250.
The methods 700 and 800 may be performed by the apparatus 200 of Fig. 2 and the terminal device 130. The operations 704-706 and 805-807 can, for example, be performed by the at least one processor 202 and the at least one memory 204. Further features of the methods
700 and 800 directly result from the functionalities and parameters of the apparatus 200, and thus are not repeated here. The methods 700 and 800 can be performed by computer program (s) .
The apparatus 200 may comprise means for performing at least one method described herein. In one example, the means may comprise the at least one processor 202, and the at least one memory 204 storing instructions that, when executed by the at least one processor, cause the apparatus 200 to perform the method .
The functionality described herein can be performed, at least in part, by one or more computer program product components such as software components. According to an embodiment, the apparatus 200 may comprise a processor or processor circuitry, such as for example a microcontroller, configured by the program code when executed to execute the embodiments of the operations and functionality described. Alternatively, or in addition, the functionality described herein can be performed, at least in part, by one or more hardware logic components. For example, and without limitation, illustrative types of hardware logic components that can be used include Field-programmable Gate Arrays (FPGAs) , Program-specific Integrated Circuits (ASICs) , Programspecific Standard Products (ASSPs) , System-on-a-chip systems (SOCs) , Complex Programmable Logic Devices (CPLDs) , Tensor Processing Units (TPUs) , and Graphics Processing Units (GPUs) .
Any range or device value given herein may be extended or altered without losing the effect sought. Also, any embodiment may be combined with another embodiment unless explicitly disallowed.
Although the subject matter has been described in language specific to structural features and/or acts, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to the
speci fic features or acts described above . Rather, the speci fic features and acts described above are disclosed as examples of implementing the claims and other equivalent features and acts are intended to be within the scope of the claims .
It will be understood that the benefits and advantages described above may relate to one embodiment or may relate to several embodiments . The embodiments are not limited to those that solve any or all of the stated problems or those that have any or all of the stated benef its and advantages . It wi ll further be understood that reference to ' an ' item may refer to one or more of those items .
The steps of the methods described herein may be carried out in any suitable order, or simultaneously where appropriate . Additionally, individual blocks may be deleted from any of the methods without departing from the spirit and scope of the subj ect matter described herein . Aspects of any of the embodiments described above may be combined with aspects of any of the other embodiments described to form further embodiments without losing the ef fect sought .
The term ' comprising ' is used herein to mean including the method, blocks or elements identi fied, but that such blocks or elements do not comprise an exclusive list and a method or apparatus may contain additional blocks or elements .
It will be understood that the above description is given by way of example only and that various modi f ications may be made by those s kil led in the art . The above speci fication, examples and data provide a complete description of the structure and use of exemplary embodiments . Although various embodiments have been described above with a certain degree of particularity, or with reference to one or more individual embodiments , those skilled in the art could make numerous alterations to the disclosed embodiments without
departing from the spirit or scope of this speci fication .
Claims
1. An apparatus (200) , comprising: at least one processor (202) ; and at least one memory (204) storing instructions that, when executed by the at least one processor (202) , cause the apparatus (200) at least to perform: in response to a beam failure recovery, BFR, procedure being initiated by a terminal device (130) in order to search for a new candidate beam after a detected failure of a current beam, determining a set of one or more BFR configuration parameters for use in the BFR procedure; and providing the determined set of the one or more BFR configuration parameters to the terminal device (130) , wherein the determining of the set of the one or more BFR configuration parameters comprises applying a machine learning, ML, model (250) to a set of model input parameters in order to obtain the one or more BFR configuration parameters for use in the BFR procedure.
2. The apparatus (200) according to claim 1, wherein the ML model (250) is configured for supervised learning and executable to: determine BFR success probabilities for at least two different sets of the one or more BFR configuration parameters based on the set of the model input parameters; and select one of the at least two different sets of the one or more BFR configuration parameters as a preferred set to be provided to the terminal device (130) .
3. The apparatus (200) according to claim 2, wherein the ML model (250) is further executable to perform the selection by:
determining candidate sets of the at least two different sets of the one or more BFR configuration parameters having one of a highest BFR success probability or a BFR success probability that differs from the highest BFR success probability by less than a minimum difference threshold; and out of the candidate sets, assigning a set of the one or more BFR configuration parameters with a lowest value of a decision model input parameter of the set of the model input parameters as the preferred set.
4. The apparatus (200) according to claim 3, wherein the decision model input parameter comprises a transmission power ramping step, a beam failure recovery timer, a maximum number of retransmissions, or a response monitoring time window for a special cell, SpCell, beam failure recovery using contention-free random-access resources.
5. The apparatus (200) according to claim 3 or 4, wherein the decision model input parameter comprises a minimum power ramping step.
6. The apparatus (200) according to any of claims 2 to 5, wherein the BFR success probability comprises a ratio of a number of times a BFR procedure is successful for a given set of model input parameters to a number of times the BFR procedure was attempted for the given set of model input parameters.
7. The apparatus (200) according to any of claims 2 to 6, wherein the instructions, when executed by the at least one processor (202) , further cause the apparatus (200) to perform training of the ML model (250) based on training data comprising real-time information on BFR attempts.
8. The apparatus (200) according to claim 1, wherein the ML model (250) is configured for reinforcement learning and executable to: determine a set of the one or more BFR configuration parameters to be provided to the terminal device (130) out of at least two different sets of the one or more BFR configuration parameters, based on different combinations of model input parameters in the set of the model input parameters.
9. The apparatus (200) according to any of claims 1 to 8, wherein the model input parameters comprise at least one of a signal to interference and noise ratio of a candidate beam, a location of the terminal device (130) , or a mobility state of the terminal device (130) .
10. The apparatus (200) according to claim 9, wherein the model input parameters further comprise a random-access channel, RACH, load.
11. The apparatus (200) according to any of claims 1 to 10, wherein the BFR configuration parameters comprise at least one of a beam failure recovery timer, a transmission power ramping step, a maximum number of retransmissions, or a response monitoring time window for a special cell, SpCell, beam failure recovery using contention-free random-access resources.
12. The apparatus (200) according to any of claims 1 to 11, wherein the instructions, when executed by the at least one processor (202) , further cause the apparatus (200) to perform: triggering a retraining of the ML model (250) in response to a number of successful BFR procedures falling below a retraining threshold.
13. The apparatus (200) according to claim 12, wherein the successful BFR procedure comprises the terminal device (130) successfully completing a beam switch after the BFR procedure is attempted.
14. The apparatus (200) according to any of claims 1 to 13, comprised in a network side device (120) or the terminal device (130) .
15. A method (700, 800) , comprising: in response to a beam failure recovery, BFR, procedure being initiated by a terminal device (130) in order to search for a new candidate beam after a detected failure of a current beam, determining (705, 806) , by an apparatus (200) , a set of one or more BFR configuration parameters for use in the BFR procedure; and providing (706, 807) , by the apparatus (200) , the determined set of the one or more BFR configuration parameters to the terminal device (130) , wherein the determining (705, 806) of the set of the one or more BFR configuration parameters comprises applying a machine learning, ML, model (250) to a set of model input parameters in order to obtain the one or more BFR configuration parameters for use in the BFR procedure.
16. A computer program comprising instructions for causing an apparatus to perform at least the following : in response to a beam failure recovery, BFR, procedure being initiated by a terminal device in order to search for a new candidate beam after a detected failure of a current beam, determining a set of one or more BFR configuration parameters for use in the BFR procedure; and
providing the determined set of the one or more BFR configuration parameters to the terminal device , wherein the determining o f the set o f the one or more BFR configuration parameters comprises applying a machine learning, ML, model to a set of model input parameters in order to obtain the one or more BFR configuration parameters for use in the BFR procedure .
Applications Claiming Priority (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| PCT/EP2023/066948 WO2024260560A1 (en) | 2023-06-22 | 2023-06-22 | Machine learning -based adjustment of beam failure recovery related parameters, and related devices, methods and computer programs |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| EP4732462A1 true EP4732462A1 (en) | 2026-04-29 |
Family
ID=87036567
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| EP23735259.6A Pending EP4732462A1 (en) | 2023-06-22 | 2023-06-22 | Machine learning -based adjustment of beam failure recovery related parameters, and related devices, methods and computer programs |
Country Status (2)
| Country | Link |
|---|---|
| EP (1) | EP4732462A1 (en) |
| WO (1) | WO2024260560A1 (en) |
Family Cites Families (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN115428347B (en) * | 2020-04-01 | 2025-10-28 | 诺基亚技术有限公司 | Method and apparatus for beam fault management and training machine learning models |
-
2023
- 2023-06-22 WO PCT/EP2023/066948 patent/WO2024260560A1/en not_active Ceased
- 2023-06-22 EP EP23735259.6A patent/EP4732462A1/en active Pending
Also Published As
| Publication number | Publication date |
|---|---|
| WO2024260560A1 (en) | 2024-12-26 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| EP4038913B1 (en) | Providing producer node machine learning based assistance | |
| US12581340B2 (en) | Method of reducing transmission of data in a communications network by using machine learning | |
| US12518157B2 (en) | Dynamic network configuration | |
| US20210400651A1 (en) | Apparatuses, devices and methods for performing beam management | |
| Azizi et al. | MIX-MAB: Reinforcement learning-based resource allocation algorithm for LoRaWAN | |
| CN116208494B (en) | AI model switching processing methods, devices, and communication equipment | |
| US20240354659A1 (en) | Method for Updating Model and Communication Device | |
| Pjanić et al. | Early-scheduled handover preparation in 5g nr millimeter-wave systems | |
| EP4732462A1 (en) | Machine learning -based adjustment of beam failure recovery related parameters, and related devices, methods and computer programs | |
| US20260058884A1 (en) | Artificial intelligence (ai) task processing method and apparatus | |
| CN111381959B (en) | Capacity expansion method and device | |
| US20250125895A1 (en) | Wireless communication method, terminal device, and network device | |
| US10674386B2 (en) | Network out-of-synchronization state processing method and apparatus, and network access method and apparatus | |
| CN117408349A (en) | Model training method and system, and computer-readable storage medium | |
| KR20260034055A (en) | Communication method and communication device | |
| CN114580610B (en) | Neural network model quantization method and device | |
| KR20240160529A (en) | System and method for performing a conditional handover | |
| CN118158737A (en) | Information transmission method, information transmission device and communication equipment | |
| US12075363B2 (en) | Determining a radio transmission power threshold, and related devices, methods and computer programs | |
| US20250126652A1 (en) | A prach radio receiver device with a neural network, and related methods and computer programs | |
| EP4601350A1 (en) | Computational mode switch | |
| KR102575193B1 (en) | Subframe positioning method and device, storage medium and electronic device | |
| CN121357560A (en) | Capacity coverage optimization methods, network nodes, user equipment, and communication systems | |
| WO2023153969A1 (en) | Handling blockage of communication for user equipment in industrial environment | |
| KR20250169970A (en) | Method and system for selecting optimal nr rach procedure |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: UNKNOWN |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: THE INTERNATIONAL PUBLICATION HAS BEEN MADE |
|
| PUAI | Public reference made under article 153(3) epc to a published international application that has entered the european phase |
Free format text: ORIGINAL CODE: 0009012 |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: REQUEST FOR EXAMINATION WAS MADE |
|
| 17P | Request for examination filed |
Effective date: 20260122 |
|
| AK | Designated contracting states |
Kind code of ref document: A1 Designated state(s): AL AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HR HU IE IS IT LI LT LU LV MC ME MK MT NL NO PL PT RO RS SE SI SK SM TR |