WO2026029933A1 - End-to-end network resiliency framework - Google Patents
End-to-end network resiliency frameworkInfo
- Publication number
- WO2026029933A1 WO2026029933A1 PCT/US2025/036896 US2025036896W WO2026029933A1 WO 2026029933 A1 WO2026029933 A1 WO 2026029933A1 US 2025036896 W US2025036896 W US 2025036896W WO 2026029933 A1 WO2026029933 A1 WO 2026029933A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- failure
- interface
- network
- ran
- data
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Pending
Links
Classifications
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04L—TRANSMISSION OF DIGITAL INFORMATION, e.g. TELEGRAPHIC COMMUNICATION
- H04L41/00—Arrangements for maintenance, administration or management of data switching networks, e.g. of packet switching networks
- H04L41/06—Management of faults, events, alarms or notifications
- H04L41/0654—Management of faults, events, alarms or notifications using network fault recovery
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04L—TRANSMISSION OF DIGITAL INFORMATION, e.g. TELEGRAPHIC COMMUNICATION
- H04L41/00—Arrangements for maintenance, administration or management of data switching networks, e.g. of packet switching networks
- H04L41/08—Configuration management of networks or network elements
- H04L41/0803—Configuration setting
- H04L41/0813—Configuration setting characterised by the conditions triggering a change of settings
- H04L41/0816—Configuration setting characterised by the conditions triggering a change of settings the condition being an adaptation, e.g. in response to network events
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04L—TRANSMISSION OF DIGITAL INFORMATION, e.g. TELEGRAPHIC COMMUNICATION
- H04L43/00—Arrangements for monitoring or testing data switching networks
- H04L43/08—Monitoring or testing based on specific metrics, e.g. QoS, energy consumption or environmental parameters
- H04L43/0805—Monitoring or testing based on specific metrics, e.g. QoS, energy consumption or environmental parameters by checking availability
- H04L43/0811—Monitoring or testing based on specific metrics, e.g. QoS, energy consumption or environmental parameters by checking availability by checking connectivity
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04L—TRANSMISSION OF DIGITAL INFORMATION, e.g. TELEGRAPHIC COMMUNICATION
- H04L43/00—Arrangements for monitoring or testing data switching networks
- H04L43/08—Monitoring or testing based on specific metrics, e.g. QoS, energy consumption or environmental parameters
- H04L43/0805—Monitoring or testing based on specific metrics, e.g. QoS, energy consumption or environmental parameters by checking availability
- H04L43/0817—Monitoring or testing based on specific metrics, e.g. QoS, energy consumption or environmental parameters by checking availability by checking functioning
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04W—WIRELESS COMMUNICATION NETWORKS
- H04W24/00—Supervisory, monitoring or testing arrangements
- H04W24/02—Arrangements for optimising operational condition
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04W—WIRELESS COMMUNICATION NETWORKS
- H04W28/00—Network traffic management; Network resource management
- H04W28/02—Traffic management, e.g. flow control or congestion control
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04L—TRANSMISSION OF DIGITAL INFORMATION, e.g. TELEGRAPHIC COMMUNICATION
- H04L41/00—Arrangements for maintenance, administration or management of data switching networks, e.g. of packet switching networks
- H04L41/06—Management of faults, events, alarms or notifications
- H04L41/0631—Management of faults, events, alarms or notifications using root cause analysis; using analysis of correlation between notifications, alarms or events based on decision criteria, e.g. hierarchy, tree or time analysis
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04L—TRANSMISSION OF DIGITAL INFORMATION, e.g. TELEGRAPHIC COMMUNICATION
- H04L41/00—Arrangements for maintenance, administration or management of data switching networks, e.g. of packet switching networks
- H04L41/08—Configuration management of networks or network elements
- H04L41/0803—Configuration setting
- H04L41/0813—Configuration setting characterised by the conditions triggering a change of settings
- H04L41/082—Configuration setting characterised by the conditions triggering a change of settings the condition being updates or upgrades of network functionality
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04L—TRANSMISSION OF DIGITAL INFORMATION, e.g. TELEGRAPHIC COMMUNICATION
- H04L41/00—Arrangements for maintenance, administration or management of data switching networks, e.g. of packet switching networks
- H04L41/08—Configuration management of networks or network elements
- H04L41/0895—Configuration of virtualised networks or elements, e.g. virtualised network function or OpenFlow elements
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04L—TRANSMISSION OF DIGITAL INFORMATION, e.g. TELEGRAPHIC COMMUNICATION
- H04L41/00—Arrangements for maintenance, administration or management of data switching networks, e.g. of packet switching networks
- H04L41/14—Network analysis or design
- H04L41/145—Network analysis or design involving simulating, designing, planning or modelling of a network
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04L—TRANSMISSION OF DIGITAL INFORMATION, e.g. TELEGRAPHIC COMMUNICATION
- H04L41/00—Arrangements for maintenance, administration or management of data switching networks, e.g. of packet switching networks
- H04L41/14—Network analysis or design
- H04L41/147—Network analysis or design for predicting network behaviour
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04L—TRANSMISSION OF DIGITAL INFORMATION, e.g. TELEGRAPHIC COMMUNICATION
- H04L41/00—Arrangements for maintenance, administration or management of data switching networks, e.g. of packet switching networks
- H04L41/16—Arrangements for maintenance, administration or management of data switching networks, e.g. of packet switching networks using machine learning or artificial intelligence
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04L—TRANSMISSION OF DIGITAL INFORMATION, e.g. TELEGRAPHIC COMMUNICATION
- H04L43/00—Arrangements for monitoring or testing data switching networks
- H04L43/20—Arrangements for monitoring or testing data switching networks the monitoring system or the monitored elements being virtualised, abstracted or software-defined entities, e.g. SDN or NFV
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04W—WIRELESS COMMUNICATION NETWORKS
- H04W88/00—Devices specially adapted for wireless communication networks, e.g. terminals, base stations or access point devices
- H04W88/08—Access point devices
- H04W88/085—Access point devices with remote components
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04W—WIRELESS COMMUNICATION NETWORKS
- H04W88/00—Devices specially adapted for wireless communication networks, e.g. terminals, base stations or access point devices
- H04W88/14—Backbone network devices
Definitions
- the present disclosure relates to an end-to-end network resiliency framework.
- Open Radio Access Network (O-RAN) architecture represents a flexible, open-source framework for mobile networks. Separation of hardware and software components allows diverse vendors to collaborate or support multi-vendor deployments. Such an approach promotes innovation, reduces costs, and enhances network performance by enabling interoperability and scalability across various network elements and services.
- O-RAN Open Radio Access Network
- deployments face significant challenges in ensuring end-to-end network resiliency due to the diversity of components and vendors involved, as failures can arise at multiple levels within the O-RAN architecture.
- an O-RAN Radio Unit (O-RU) level For instance, an O-RAN Radio Unit (O-RU) level, an O-RAN Distributed Unit (O- DU) level, an O-RAN Central Unit (O-CU) level, an interface level, a transport level, an O-RAN cloud level, a RAN Intelligent Controller (RIC) level, and an Evolved Node B (eNB) level.
- O-RU O-RAN Radio Unit
- O-DU O-RAN Distributed Unit
- O-CU O-RAN Central Unit
- interface level For instance, an O-RAN network
- transport level For instance, an O-RAN Radio Unit (O-CU) level, an O-RAN Central Unit (O-CU) level, an interface level, a transport level, an O-RAN cloud level, a RAN Intelligent Controller (RIC) level, and an Evolved Node B (eNB) level.
- RIC RAN Intelligent Controller
- eNB Evolved Node B
- a method includes monitoring, at a MnS consumer, data of one or more interfaces or one or more network entities associated with an Open Radio Access Network (0-RAN) architecture. Monitoring the data comprises monitoring at least one of Fault, Configuration, Accounting, Performance, Security (FCAPS) data, Performance Measurement (PM) counters, Fault Management (FM) data, a notification message, and an alarm.
- FCAPS Fault, Configuration, Accounting, Performance, Security
- PM Performance Measurement
- FM Fault Management
- the method further includes detecting a failure occurring in at least one of the one or more interfaces or the one or more network entities associated with the O- RAN architecture based on the monitoring.
- the method further includes performing, upon detection of the failure, at least one recovery action.
- an apparatus configured to periodically monitor, at a MnS consumer, data of one or more interfaces or one or more network entities associated with an Open Radio Access Network (O-RAN) architecture. Monitoring the data comprises monitoring at least one of Fault, Configuration, Accounting, Performance, Security (FCAPS) data, Performance Measurement (PM) counters, Fault Management (FM) data, a notification message, and an alarm.
- FCAPS Fault, Configuration, Accounting, Performance, Security
- PM Performance Measurement
- FM Fault Management
- the apparatus is further configured to detect a failure occurring in at least one of the one or more interfaces or the one or more network entities associated with the O-RAN architecture based on the monitoring.
- the apparatus is further configured to perform, upon detection of the failure, at least one recovery action.
- a non-transitory computer- readable medium storing instructions.
- the one or more instructions are executed by an apparatus which comprises one or more processors.
- the one or more processors may periodically monitor, at a MnS consumer, data of one or more interfaces or one or more network entities associated with an Open Radio Access Network (O-RAN) architecture. Monitoring the data comprises monitoring at least one of Fault, Configuration, Accounting, Performance, Security (FCAPS) data, Performance Measurement (PM) counters, Fault Management (FM) data, a notification message, and an alarm.
- the one or more processors may detect a failure occurring in at least one of the one or more interfaces or the one or more network entities associated with the O-RAN architecture based on the monitoring.
- the one or more processors may perform, upon detection of the failure, at least one recovery action.
- FIG. 1 is a diagram of an example of an implementation environment in which systems and/or methods, described herein, may be implemented, according to an embodiment as disclosed herein;
- FIG. 2 illustrates an example block diagram depicting a system of resiliency at various levels of an Open Radio Access Network (O-RAN) architecture, according to an embodiment as disclosed herein;
- OF-RAN Open Radio Access Network
- FIG. 3 is a flow diagram illustrating a method for managing an end-to-end network resiliency within the O-RAN architecture, according to an embodiment as disclosed herein;
- FIG. 4 is a flow diagram illustrating a method for determining at least one recovery action, according to an embodiment as disclosed herein;
- FIG. 5 illustrates an example scenario where a Service Management and Orchestration (SMO) detects a failure occurring at an O-RAN Distributed Unit (0-DU) level, according to an embodiment as disclosed herein;
- FIG. 6 illustrates an exemplary backup scenario associated with an O-DU, according to an embodiment as disclosed herein;
- SMO Service Management and Orchestration
- FIGS. 7A-7B illustrates an advanced resiliency use case (partial O-DU failure), according to an embodiment as disclosed herein;
- FIGS. 8A, 8B and 8C illustrates an O-DU resiliency use case, according to an embodiment as disclosed herein;
- FIGS. 9A-9B illustrates O-DU resiliency scenario in case of 01 failure, according to an embodiment as disclosed herein;
- FIGS. 10A, 10B and IOC illustrates enhanced network resiliency through a standby O-DU and adaptive cell activation managed by the SMO, according to an embodiment as disclosed herein;
- FIG. 11 illustrates a diagram of example components of an apparatus, according to an embodiment as disclosed herein.
- AI/ML Artificial Intelligence/Machine Learning
- the AI/ML model is a model generated using one or more Al technologies, one or more ML algorithms, or both, and generates output data based on input data. This output data is used to perform tasks.
- Tasks performed using AI/ML models include those generally referred to as intellectual tasks, such as classification, prediction, natural language processing, etc.
- Al and ML are explained separately, ML is a technology included in Al. In ML, instead of being explicitly programmed for a specific task, systems can improve their performance over time by identifying patterns and making inferences from training data.
- the generation of ML models includes data collection, model training, and model inference. Data collection involves gathering and preprocessing data to be used for training and inference. Model training involves developing and validating models using the collected data. Model inference involves applying the trained models to new data to generate new output data and perform tasks.
- Machine learning includes various types of learning methods such as supervised learning, unsupervised learning, reinforcement learning, semi-supervised learning, self-supervised learning, transudative learning, transfer learning, meta learning, and the like. These types of learning methods can be appropriately selected according to the embodiments. Unless otherwise specified, the application of types not mentioned in this description is not precluded. Additionally, the structure of ML models may vary depending on the embodiments and learning methods, and is not limited to the methods disclosed. Furthermore, ML includes deep learning, which uses models that include neural networks. Deep learning models may include, for example, deep neural networks (DNNs), convolutional neural networks (CNNs), etc.
- DNNs deep neural networks
- CNNs convolutional neural networks
- AI/ML models presented hereinafter are examples and are not limited to the illustrated AI/ML models. They can be modified or altered by using different Al or ML algorithms.
- the configuration of the neural network is not limited to the configuration disclosed in the present disclosure and can be modified.
- MnS Management and Orchestration Service
- O-RAN Open Radio Access Network
- O-CU O-RAN Central Unit
- gNB Next Generation Node B
- gNB-CU Next Generation Node B
- POD Point of Delivery
- 3GPP Third Generation Partnership Project
- SCTP Stream Control Transmission Protocol
- CU-CP Central Unit Control Plane
- DU for an Fl interface
- CU-UP Central Unit User Plane
- El interface referenced in 3GPP TS 38.462
- redundancy at a network layer can be accomplished through multiple transport paths, ensuring data delivery even in the event of a path failure, while also enabling flexible and efficient routing of data packets.
- port redundancy can be realized through cloud orchestration and platform solutions, such as those based on Kubemetes.
- a disclosed method is proposed, as discussed throughout the disclosure (FIGS. 1 to 8).
- the disclosed method introduces a unified resiliency service managed by the MnS consumer (e.g., SMO), offering a comprehensive solution that encompasses all entities involved. Additionally, the disclosed method considers multiple aspects, including potential failure scenarios at various levels, backup and recovery mechanisms, necessary hardware and software support, and enhancements driven by Al/ML model(s).
- FIG. 1 is a diagram of an example of an implementation environment 100 in which systems and/or methods, described herein, may be implemented, according to an embodiment as disclosed herein.
- the implementation environment 100 includes a User Equipment (UE) 110, a service environment 120, and a network 130.
- the service environment 120 and the network 130 may relate to an Open Radio Access Network (O-RAN) architecture, as shown in FIG. 2.
- the service environment 120 includes one or more sub-environments 121-1 to 121-N (collectively and/or interchangeably referred hereinafter as 121).
- FIG. 1 shows, for convenience, examples of a 1st sub-environment 121-1, a 2nd sub-environment 121-2, and an N 111 subenvironment 121-N (where N is any natural number).
- the UE 110 is connected to the network 130, and the network 130 is connected to the service environment 120.
- the connections may be wired, wireless, or a combination of both wired and wireless.
- the UE 110 and the service environment 120 are connected via the network 130.
- the UE 110 is a device that communicates with the service environment 120.
- the UE 110 receives information from the service environment 120 and/or sends information to the service environment 120.
- the UE 110 may generate and/or store information to be transmitted, as necessary.
- the UE 110 may store and/or process information that is received, as necessary.
- FIG. 1 refers to the “UE”.
- UE User device
- terminal terminal device
- communication device communication terminal
- the UE 110 may include a computing device (e.g., a desktop computer, a laptop computer, a tablet computer, a handheld computer, a smart speaker, a server, etc.), a mobile phone (e.g., a smartphone, a radiotelephone, etc.), a wearable device (e g., a pair of smart glasses or a smart watch), or a similar device.
- a computing device e.g., a desktop computer, a laptop computer, a tablet computer, a handheld computer, a smart speaker, a server, etc.
- a mobile phone e.g., a smartphone, a radiotelephone, etc.
- a wearable device e.g., a pair of smart glasses or a smart watch
- the service environment 120 is an environment that communicates with the UE 110 to provide one or more services.
- the service environment 120 receives information from the UE 110 and/or sends information to the UE 110.
- the service environment 120 may generate and/or store information to be transmitted, as necessary.
- the service environment 120 may store and/or process information that is received, as necessary.
- the service environment 120 may provide computing resources as one of the services.
- the service is not limited to being provided to the UE 110; it may also be provided to devices other than the UE 110.
- the service may perform processes such as anomaly detection or traffic analysis and notify the results to a predetermined destination.
- service environment refers to the “service environment”.
- service environment is used to refer to the broader context within which services operate.
- cloud environments, platforms, computing systems, network systems, and cloud systems generally represent the environments in which services are conducted, and these are included within the “service environment”.
- the “service environment” is not limited to these examples.
- specific types of environments within the “service environment” are not restricted.
- cloud environments and cloud systems can be categorized as private cloud, public cloud, hybrid cloud, or multi-cloud, all of which are included within the “service environment”.
- the one or more services provided by the service environment 120 is not specifically limited and can be adjusted according to the embodiments.
- the one or more services may include a service that provides information to the UE 110, a service that stores information from the UE 110, or a service that performs processing based on information from the UE 110 and returns the results of the processing.
- the service environment 120 may also provide computing resources.
- the computing resources can be hardware resources and/or software resources. For example, applications, processors, memory, and storage can be included in the provided computing resources.
- Each computing resource can communicate with other computing resources via wired connections, wireless connections, or a combination of wired and wireless connections.
- the provided computing resources can be actual resources (also referred to as physical resources) and/or virtual resources.
- means of virtualization for virtual resources can be selected as appropriate. That is, in this disclosure, the use of adjectives such as “virtual” or “virtualized” to describe names does not imply that they are virtualized by a specific means of virtualization.
- “virtual machine” refers to software that operates like an actual computer, realized through means of virtualization, and it is not intended to exclude those realized by specific means of virtualization such as hypervisors or containers.
- means of virtualization such as hypervisors or containers are mentioned in this disclosure, it is merely cited as a general method of implementation. It should also be interpreted that embodiments implemented with other virtualization means are also disclosed.
- the services may also be provided using resources virtualized by different means.
- the service environment 120 includes one or more devices, such as servers and network devices, which provide services or perform processes. The placement of these devices within the service environment 120 can be determined as appropriate. Additionally, if the service environment 120 includes one or more sub-environments 121, the placement of devices can be determined based on predetermined policies for each sub-environment 121. For example, devices related to the first service may be placed in the 1st sub-environment 121-1, and devices related to the second service may be placed in the 2nd sub-environment 121-2. In another example, devices expected to have a higher load than a predetermined threshold may be placed in the 1st subenvironment 121-1, while devices expected to have a lower load than the predetermined threshold may be placed in the 2nd sub-environment 121-2. In this way, specific devices can be placed in specific sub-environments 121. Conversely, each sub-environment 121 can be specialized for a particular purpose.
- all processes executed in a single service may run within a single service environment, or in multiple service environments. Multiple processes executed in a single service could be provided by different service environments.
- the network 130 is a network that exchanges information between the UE 110 and the service environment 120.
- the network 130 includes one or more wired and/or wireless networks.
- the network 130 may include a cellular network (e.g., a Fifth Generation (5G) network, a Long-Term Evolution (LTE) network, a Third Generation (3G) network, a Code Division Multiple Access (CDMA) network, etc.), a Public Land Mobile Network (PLMN), a Local Area Network (LAN), a Wide Area Network (WAN), a Metropolitan Area Network (MAN), a telephone network (e.g., the Public Switched Telephone Network (PSTN)), a private network, an ad hoc network, an intranet, the Internet, a fiber optic-based network, or the like, a NonTerrestrial Network (NTN), and/or a combination of these or other types of networks.
- 5G Fifth Generation
- LTE Long-Term Evolution
- 3G Third Generation
- CDMA Code Division Multiple Access
- PLMN Public Land Mobile Network
- the network 130 can be a part of a network.
- the network 130 in a 5G network that includes a RAN, a transport network, and a core network, the network 130 can be at least one of the RAN, the transport network, or the core network.
- the service environment 120 could be in the core network, in which case the network 130 could correspond to a network that is a combination of a RAN and a transport network and is part of the 5G network.
- FIG. 1 The number and arrangement of devices and networks shown in FIG. 1 are provided as an example. It should be understood that any changes that may be implemented by those skilled in the art, such as the addition or rearrangement of well-known devices or networks at the time of implementation, are included in this disclosure.
- FIG. 2 illustrates an example block diagram depicting a system of resiliency at various levels of the Open Radio Access Network (0-RAN) architecture, according to an embodiment as disclosed herein.
- the 0-RAN architecture comprises one or more network entities and one or more interfaces for communication and management.
- the one or more network entities may include, but not limited to, the Service Management and Orchestration (SMO) 200, a non-real-time RAN Intelligent Controller (RIC) 201, a near-real-time RIC 202, an O-RAN Central Unit User Plane (O-CU-UP) 203, an O-RAN Central Unit Control Plane (O-CU-CP) 204, the 0-RAN Distributed Unit (O-DU) 205, an O-RAN Radio Unit (O-RU) 206, the O-RAN cloud (O-cloud) 207, and an O-RAN enhanced Node B (O-eNB).
- SMO Service Management and Orchestration
- RIC non-real-time RAN Intelligent Controller
- RIC O-RAN Central Unit User Plane
- OF-CU-CP O-RAN Central Unit Control Plane
- OF-DU 0-RAN Distributed Unit
- OF-RU O-RAN Radio Unit
- O-cloud O-RAN cloud
- O-eNB O-RAN enhanced Node B
- the one or more interfaces may include, but not limited to, an Al interface, an 01 interface, an 02 interface, an E2 interface, an R1 interface, a Y1 interface, an Fl-C interface, an Fl-U interface, an open Fronthaul for Control and User plane (FH-CUS), and an Open Fronthaul for Management (FH-M) plane.
- the near-real-time RIC 202 is a logical function that enables near-real-time control and optimization of O-RAN elements and resources via fine-grained data collection and actions over the E2 interface.
- the non-real-time RIC 201 is a logical function that enables non-real-time control and optimization of the O-RAN elements and resources, AVML workflow including model training and updates, and policy-based guidance of applications/features in the near-real-time RIC 202.
- the Al interface facilitates communication between the non-real-time RIC 201 and the O-RAN elements, allowing for effective policy management and optimization.
- the E2 interface connects the near-real-time RIC 202 to the O-RAN elements (e.g., the O-DU 205, the O-CU-CP 204, and the O-CU-UP 203), supporting real-time control.
- X2-c and X2-u interfaces enable communication between eNBs for load balancing and handover procedures, while the Fl-C interface and Fl-u interface connect the O-DU 205, the O-CU-CP 204, the O-CU-UP 203 in a 5G framework, optimizing data transfer and control signaling.
- the architecture features the open FH, which includes the open FH CUS for managing fronthaul transport networks and the open FH M-Plane for control plane traffic management.
- the O-RU 206 serves as a physical layer interfacing with antennas for radio signal processing, while the O-Cloud 207 provides a virtualized environment for hosting the RAN elements, enhancing scalability and flexibility.
- the O-eNB 208 represents a traditional base station connecting user devices to the network. Each element and interface play an integral role in ensuring efficient network operations, allowing service providers to deliver high- quality services.
- the disclosed method or the MnS consumer (e.g., SMO 200) provides a comprehensive end-to-end resiliency solution for one or more network levels within the O-RAN architecture, as described in conjunction with FIG. 3.
- the one or more network levels within the O-RAN architecture may include, but are not limited to, an O-RU level, an O-DU level, an O-CU level (e.g., the O-CU-CP 204 and O-CU-UP 203), an interface level (e.g., Al, 01, 02, E2, and Rl), a transport level, an O-RAN cloud level, a RIC level, and an eNB level.
- the transport level may involve both network layer and port redundancy.
- the O-RAN cloud level may involve an infrastructure layer that encompasses components like pods, clusters, Central Processing Units (CPUs), Graphics Processing Units (GPUs), and data centers and an application layer includes network functions or elements, and management functions.
- the RIC level may involve both non-RT RIC and near-RT RIC functionalities.
- the SMO 200 maintains an inventory database of deployed network components, detailing hardware and software aspects of network functions and elements. Utilizing this information, along with inputs from cell and network planning tools, the SMO 200 performs critical tasks, such as assigning roles to network functions and elements.
- the SMO 200 designates active and standby roles for O-DUs to ensure O-DU resiliency.
- One or more Operations, Administration, and Management (0AM) functions of the SMO 200 can provide directives to parent nodes, such as the near-real-time RIC 202 or O-CU (e.g., 204 and 203), for managing these roles effectively.
- parent nodes such as the near-real-time RIC 202 or O-CU (e.g., 204 and 203), for managing these roles effectively.
- the SMO 200 continuously monitors and manages the inventory database of recovery options (recovery action) for various failure scenarios (failure).
- the SMO 200 may track high-ranked neighbors for each cell to mitigate coverage gaps when a cell experiences downtime.
- the SMO 200 also maintains the inventory database of software releases and relevant updates. Operator-driven and data-driven policies orchestrate resiliency steps within the SMO 200.
- one or more operator-based policies can define system configurations for recovery during emergencies, while one or more data-driven policies may facilitate load distribution across multiple network functions in the event of overload conditions.
- the SMO 200 may also guide the near-real-time RIC 202 or the non-real-time RIC 201 to prevent backup entities from entering energy-saving states.
- the SMO 200 may consume Fault, Configuration, Accounting, Performance, Security (FCAPS) data, and leverages AI/ML enhancements to detect or predict failures.
- FCAPS Fault, Configuration, Accounting, Performance, Security
- a resiliency orchestrator (not shown in FIG. 2) within the SMO 200 is configured to process this data to identify failures.
- FCAPS Fault, Configuration, Accounting, Performance, Security
- the SMO 200 can anticipate failures and plan recovery actions proactively, as described in conjunction with FIG. 3.
- resiliency actions are informed by available data within the SMO 200, with the capability to detect both complete and partial failures based on alarms (e.g., lost O-DU ID supervision or transceiver faults) and notifications regarding events, configurations, or failures (e.g., carrier state change notification).
- alarms e.g., lost O-DU ID supervision or transceiver faults
- notifications regarding events, configurations, or failures e.g., carrier state change notification.
- These alarms and notifications can originate from various network functions, network elements, or external applications/ platforms (e.g., sleeping cell-detected notification or alarm).
- performance monitoring is conducted, by the SMO 200, through Performance Measurement (PM) counters (e.g., which track metrics like Energy, Power and Environmental (ePE) statistics, user equipment distribution, downlink Fl -U packet loss rates, and canceled Downlink Control Information (DCI) due to Physical Downlink Control Channel (PDCCH) resource shortages).
- PM Performance Measurement
- ePE Energy, Power and Environmental
- DCI Downlink Control Information
- the SMO 200 upon receiving failure alarms or notifications, the SMO 200 initiates recovery actions based on the information stored in its databases. The nature of the fault or failure impacts the SMO 200's decision-making regarding the recovery process and the available software or hardware support details.
- the SMO 200 may attempt recovery through relevant software upgrades, resets, or auto-healing mechanisms, as applicable.
- the SMO 200 may provide the end-to-end resiliency solution to ensure that various components of the network, including RAN nodes and interfaces, are effectively monitored and managed, minimizing downtime and maintaining service continuity.
- the SMO 200 may perform rapid identification and mitigation of potential failures, thus improving network reliability, as one of the advantages of the disclosed method.
- the SMO 200 may have the ability to assign active and standby roles to network functions, such as the O-DUs, which enhances system redundancy. This proactive role assignment ensures that backup resources are readily available, facilitating quick recovery in the event of a failure, as one of the advantages of the disclosed method. Additionally, the integration of operator- driven and data-driven policies allows for tailored resiliency strategies, enabling operators to define specific recovery protocols based on real-time conditions and historical data.
- the continuous monitoring of network performance and the use of AI/ML for failure prediction significantly enhance the SMO’s responsiveness.
- the SMO 200 may anticipate issues before they escalate, allowing for preemptive actions that further safeguard network integrity, as one of the advantages of the disclosed method.
- the centralized inventory database of hardware and software components streamlines management processes, providing operators with crucial information at their fingertips. This database not only supports effective decision-making during recovery operations but also aids in maintaining up-to- date software releases and configurations, ultimately leading to improved service quality, reduced operational risks, and enhanced user satisfaction.
- FIG. 3 is a flow diagram illustrating a method 300 for managing the end-to-end network resiliency within the 0-RAN architecture, according to an embodiment as disclosed herein.
- the method 300 may execute multiple operations to manage the end-to-end network resiliency, which is given below.
- the method 300 includes monitoring, at the MnS consumer (e.g., SMO 200), data of the one or more interfaces and the one or more network entities associated with the 0-RAN architecture.
- monitoring the data comprises monitoring at least one of the FCAPS data, the PM counters, Fault Management (FM) data, a notification message, and an alarm.
- the method 300 includes detecting the failure occurring in at least one of the one or more interfaces or the one or more network entities associated with the O-RAN architecture based on the monitoring.
- the method 300 includes performing, upon detection of the failure, at least one recovery action. Various examples of the at least one recovery action are explained in the below embodiments.
- the SMO 200 may perform the at least one recovery action, such as executing one or more software auto-healing mechanisms.
- the one or more software auto-healing mechanisms may include performing software upgrades and restarting the O-RU 206, upon availability of a Management Plane (M-plane).
- M-plane Management Plane
- the SMO 200 in response to detecting the failure occurring at the O-RU 206, may perform the at least one recovery action, such as adjusting one or more neighboring O-RUs configuration to provide support for mitigating one or more coverage gaps that occurred due to the failure.
- the O-DU 205 can identify O-RU failures through FM data, including the alarms, Configuration Management (CM) notifications (e.g., carrier state change or carrier activation, carrier configuration related notifications), and PM data sent by the O-RU 206 via the open fronthaul M-Plane interface.
- CM Configuration Management
- the SMO 200 may also receive one or more O-RU alarms through this interface, acting as an O-RU controller.
- the O-DU 205 can either take independent action or relay failure information to the SMO 200 via the 01 interface, allowing the SMO 200 to implement recovery measures for the O-RU 206.
- the SMO 200 may perform the at least one recovery action, such as executing one or more software auto-healing mechanisms when the 01 interface or the E2 interface is operational and one or more active sessions are established in the O-RAN architecture.
- the SMO 200 may perform the at least one recovery action, such as executing, by utilizing at least one Al model, at least one a manual rehoming process, and an automated rehoming process based on the FM and PM data, for predictive autohealing and failure forecasting.
- the SMO 200 may perform the at least one recovery action, such as triggering redundant O-DU (standby O-DU) to take over one or more services for the O-RU 206.
- triggering action may utilize a transfer of UE context (if an inter-DU interface is available e.g., D2 interface) and/or cell configuration information for the at least one recovery action.
- the SMO 200 in response to detecting the failure occurring at either the O-DU 205 or the O-CU (e.g., 203 and 204), the SMO 200 may perform the at least one recovery action, such as executing one or more transport layer-based backup recovery mechanisms for either the O-DU 205 or the O-CU (e.g., 203 and 204).
- the SMO 200 may detect the failure occurring at the O- DU 205 may include, but not limited to, a Management Function (MnF) failure, a Network Function (NF) failure, a complete O-DU failure, and any partial failure.
- MnF Management Function
- NF Network Function
- the SMO 200 may identify an O-DU MnF failure by analyzing the PM data from the O-CU (e.g., 203 and 204) through the 01 interface, provided that the Fl link between the O-CU (e.g., 203 and 204) and the O-DU 205 is still operational.
- the O- CU e.g., 203 and 204 may configure the O-DU 205 via the Fl interface and monitor its resource status.
- the SMO 200 may collect and compare the PM data from the O-CU (e.g., 203 and 204) with the 01 interface status to the O- DU 205. This process helps determine if there is a failure in the MnF of the O-DU 205 or its connection, especially when traffic is active but cannot reach the O-DU 205 through the 01 interface.
- the SMO 200 may detect the failure directly through this interface.
- the SMO 200 may also identify the issue using performance management data from the O-CU (e.g., 203 and 204), as the failure of the Fl interface and the FM data may indicate the NF failure.
- the SMO 200 may analyze the PM data from the O-CU (e.g., 203 and 204) along with the 01 connection status and transport management data received via the 02 interface (if Cloud Native Network Function/Cloudified Network Function/Containerized Network Function (CNF) deployment is in place).
- the PM data may indicate an Fl failure due to the NF failure, allowing the SMO 200 to conclude that the entire O-DU is inoperable.
- the SMO 200 may detect the failure occurring at the O- CU (e.g., 203 and 204) may include, but not limited to, an O-CU-CP failure and an O-CU-UP failure.
- a failure in the O-CU-CP 204 may significantly disrupt cell service by interrupting essential signaling between the UE 110 and the core network. This can result in problems such as failed call setups, handovers, and session management. If the O-CU-CP 204 does not support UE session restoration, it may also trigger large signaling storms. Additionally, the O-CU-CP 204 may lose connectivity with the SMO 200 through the 01 interface, leading to a loss of configurations and preventing the SMO 200 from receiving 0AM data (CM, FM, and PM) from the O-CU-CP 204. It is crucial to restore gNB, cell services, and user sessions to re-establish functionality, minimize disruptions, and ensure service continuity.
- the O-CU-UP failure in a 5G 0-RAN network can disrupt the transmission of user data between the UE 110 and the core network. This may lead to affected data sessions, reduced service quality, and potential data loss. Furthermore, the O-CU-UP 203 may also be unable to communicate with the SMO 200 via the 01 interface.
- the SMO 200 in response to detecting the failure occurring at the interface level, may perform the at least one recovery action, such as executing one or more backup interface-based recovery mechanisms (e.g., multiple SCTP sessions-based recovery) for data transmissions for the plurality of network interfaces (i.e., one or more interfaces).
- the SMO 200 in response to detecting the failure occurring at the interface level, may perform the at least one recovery action, such as executing at least one redundant transmission mechanism for data transmissions for the plurality of network interfaces.
- An example of the at least one redundant transmission mechanism may include a Multipath transmission control (MPTCP), which allows data to be transmitted over multiple paths, increasing resilience against path failures.
- MPTCP Multipath transmission control
- the SMO 200 may perform the at least one recovery action, such as identifying a cause of the interface failure.
- the cause is determined to be either a malfunction in the NF or Network Element (NE) or a failure of an Access Point (AP), and recovering the AP to restore functionality and resolve the identified interface failure.
- NE Network Element
- AP Access Point
- the failure occurring at the interface level may include, but not limited to, a Network Interface Card (NIC) failure, an Internet Protocol (IP) level routing failure, a Stream Control Transmission Protocol (SCTP) connection failure, a control plane interface failure, a user plane interface failure, and a management plane interface failure.
- NIC Network Interface Card
- IP Internet Protocol
- SCTP Stream Control Transmission Protocol
- the SMO 200 may perform the at least one recovery action, such as executing one or more transport-based recovery mechanisms.
- the one or more transport-based recovery mechanisms are utilized for data transmissions based on distributing traffic and workloads across the plurality of network entities (i.e., one or more network entities) for load balancing
- the SMO 200 may perform the at least one recovery action, such as executing one or more NF recovery mechanisms by utilizing distinct standby NF instances or backup clusters for various networks. For instance, implement separate standby instances or backup clusters for both 4G and 5G networks, ensuring that one has priority over the other.
- various types of failure(s) can occur at the O-cloud level, which is mentioned below.
- Infrastructure-related failures These include issues such as O-Cloud failures, site failures, node cluster failures, and other resource failures.
- IP Multimedia Subsystem (IMS)-Related Failures This category encompasses failures related to IMS software, issues during IMS software updates, failures in the 02-IMS interface, and problems with IMS provisioning procedures.
- Deployment Management Services (DMS)-related failures Examples here include failures in the DMS control plane, issues during DMS control plane upgrades, deployment plane failures, and failures in NF deployment lifecycle management.
- O-Cloud application-related failures This includes failures of pods, virtual machines (VMs), and issues during NF deployment.
- the SMO 200 may perform the at least one recovery action, such as implementing infrastructure-level resiliency, by employing at least one open-source container orchestration system (e.g., Kubemetes-based solutions), for software auto-healing at the O-cloud level.
- the at least one open-source container orchestration system may execute one or more auto-healing processes.
- the one or more auto-healing processes are managed by the O-cloud 207 and one or more NF resiliency features.
- resiliency solutions can be established through active-active or active-standby configurations at the POD, cluster, or data center level, utilizing the Kubernetes-based solutions for software auto-healing.
- normal auto-healing is managed by cloud and NF resiliency features.
- the SMO 200 may intervene when these mechanisms are insufficient in addressing persistent failures. For instance, if an SCTP connection goes down, the alarm notification may be triggered. The SMO’ s involvement is only necessary if the connection fails to re-establish within a specified timeframe.
- the SMO 200 may perform the at least one recovery action, such as executing one or more recovery actions comprised of performing a software reset, implementing software auto-healing, redeploying, or re-onboarding one or more applications by utilizing backup information provided by the SMO 200.
- the SMO 200 may initiate appropriate recovery actions. These actions may include performing a software reset, enabling software auto-healing, redeploying affected applications, or re-onboarding the applications using the backup information that the SMO 200 has stored. This process ensures that services are quickly restored and continue to function effectively.
- the SMO 200 may retrieve one or more capabilities of the plurality of network entities associated with the O-RAN architecture. Based on the one or more retrieved capabilities, the SMO 200 may assign a role to each network entity associated with the plurality of network entities, as described in conjunction with FIG. 5.
- the SMO 200 may perform at least one action based on one or more additional inputs from one or more cell-network planning tools.
- the SMO 200 may assign specific operational roles to each network entity comprising an active role for the O-DU 205 and a standby role for the O-DU 205, to enhance network resiliency and ensure seamless service continuity or one or more services to the user with minimum interruption.
- the SMO 200 may recommend a direction to a parent node (e.g., the near-real-time RIC 202 or the O-CU (203 and 204)) regarding the effective assignment and management of the active role and the standby role associated with the O-DU 205, as described in conjunction with FIG. 6.
- the SMO 200 may orchestrate one or more resiliency measures through both operator-driven and data-driven policies.
- the operator-driven policy establishes recovery configurations and the data-driven policy manages load distribution during one or more overload situations.
- the SMO 200 may recommend one or more strategies, to the near-real-time RIC 202 or non-real-time RIC 201, for example, to prevent backup entities from entering an energy-saving mode, to ensure readiness for immediate operational demands.
- the SMO 200 may predict by utilizing the at least one Al model, the failure occurring at the at least one network entity based on the received network-related data and one or more parameters, to perform the at least one recovery action in advance.
- Examples of the one or more parameters may include, but are not limited to, a list of impacted network entities among a plurality of network entities, current traffic information associated with the at least one network entity, a Quality of Service (QoS) requirement for each network entity, a Quality of Experience (QoE) requirement for each network entity, applied network traffic information associated with the at least one network entity, and resource requirement information associated with the at least one network entity.
- QoS Quality of Service
- QoE Quality of Experience
- FIG. 4 is a flow diagram illustrating a method 400 for determining the at least one recovery action, according to an embodiment as disclosed herein.
- the method 400 may execute multiple operations to determine the at least one recovery action, which is given below.
- the method 400 includes determining a type of failure occurring at the at least one network entity based on the network-related data.
- the method 400 includes determining an available software-hardware support details at the SMO 200.
- the method 400 includes determining the at least one recovery action based on the type of the determined failure and the available software hardware support details at the SMO.
- the at least one recovery action may include one or more backup mechanisms and one or more redundant node mechanisms (e.g., 1+1 redundancy, N+l redundancy, N+M redundancy).
- the one or more backup mechanisms may include a software-based backup mechanism, a hardware-based backup mechanism, and a transport network-based backup mechanism.
- FIG. 5 illustrates an example scenario where the SMO 200 detects the failure that occurs at the O-DU, according to an embodiment as disclosed herein.
- the SMO 200 performs multiple operations to ensure robust network functionality, which are given below.
- the SMO 200 may initialize the roles of various NF instances, incorporating a multi-level mesh architecture that facilitates efficient communication and resource allocation.
- This architecture enables the identification of capabilities at each level, allowing the SMO 200 to configure the roles for each Multi-Function (MF)/NF appropriately.
- Continuous dynamic monitoring is then performed by the SMO 200 based on FCAPS feedback, which includes the PM counters.
- FCAPS feedback along with the operator’s policies, the SMO 200 establishes a relative priority for each NF instance and its supporting counterparts.
- the SMO 200 may receive alarms or notifications, prompting it to assess the priority levels previously established. This assessment informs the SMO’s decision regarding which NF instance(s) may be activated to recover from or mitigate the failure.
- the SMO 200 may employ the 01 interface to communicate the roles to these NF instances.
- the NF instances may increase their transmit power to address coverage gaps resulting from the failure of the affected NF instance, as illustrated in FIG. 5, in response to a failure of “O-DU- 1”.
- Other use cases include energy saving, mobility load balancing, rolling upgrades, and Multiple- Input Multiple-Output (MIMO) optimization.
- MIMO Multiple- Input Multiple-Output
- the configurations can be adjusted to prioritize activities such as upgrades and failure recovery, while optimization tasks remain unaffected.
- the SMO 200 may execute several critical operations for recovery (recovery action), which are given below. a. First, the SMO 200 retrieves the capabilities of the O-DU 205 and the O-RU 206. b. Subsequently, the SMO 200 assigns roles to the O-DUs, for instance, designating O-DU-1 as Active and O-DU-2 as Standby. c. Upon detecting a failure in O-DU-1 through alarms, notifications, or PM counters, the SMO 200 may configure the associated O-CU (e.g., 203 and 204) to deactivate the cells previously served by O-DU-1. d.
- the associated O-CU e.g., 203 and 204
- the SMO 200 then considers multiple recovery actions: i.
- Option-1 The SMO 200 may perform a software reset or upgrade for O- DU-1 to address the failure or degradation.
- Option-2 The SMO 200 may initiate service rehoming, which can be either manual or automated, transferring services from O-DU-1 to O-DU-2.
- Option-3 The SMO 200 may trigger or configure O-DU-2 to take over the O-RU, thus restoring service.
- O-DU-2 may configure the O-RU 206 to restore service. If the O-RU 206 is preconfigured with backup carriers, O-DU-2 may activate these carriers to facilitate service restoration. Concurrently, the SMO 200 may attempt to recover O-DU-1 through appropriate measures such as software resets, upgrades, or hardware issues.
- implementation of the resiliency framework within the SMO 200 may offer several significant advantages. Firstly, the initialization of roles for various NF instances and the integration of the multi-level mesh architecture enhances network efficiency and resource allocation, ensuring optimal performance under varying conditions. Continuous dynamic monitoring based on the FCAPS feedback allows for real-time adjustments, enabling proactive management of network resources and minimizing downtime. [0110] Additionally, the prioritization of the NF instances facilitates rapid response to failures, allowing the SMO 200 to quickly identify and activate the necessary backup instances for recovery. This responsiveness not only reduces service interruptions but also enhances overall network reliability.
- the flexibility to adjust configurations for activities such as upgrades and failure recovery ensures that critical operations can be prioritized without compromising optimization efforts.
- the ability to utilize various recovery options, such as the service rehoming and the software upgrades provides a robust mechanism for maintaining service continuity and ensuring a seamless user experience.
- FIG. 6 illustrates an exemplary backup scenario associated with the O-DU 205, according to an embodiment as disclosed herein.
- both the active and standby O-DUs can be connected to the same O-CU (Intra O-CU) or different O-CUs (Inter O-CU).
- the standby configurations may include several designs, such as 1+1, N (Active) + K (Standby), N+l (All Active + 1 Hot Spare), and 1+K (1 Active + K Hot/Cold Spare).
- resiliency can be achieved through either the active-active or the active-standby configuration.
- the M-Plane remains active alongside supervision monitoring, ensuring that the standby node or the NF is always available to recover from resiliency events or take over service operations.
- the active-standby configuration involves an inactive M-Plane, although the IP address of the O-DU (e.g., standby- 1,. . ., standby-N) is known to the O-RU 206, facilitating potential recovery.
- the capabilities of the standby node or the NF may operate at the same or different levels compared to the active O- DU 205, as determined by the orchestrator/controller (e.g., SMO 200).
- the orchestrator/controller e.g., SMO 200.
- Key metrics for quantifying these capabilities include carrier resource management, throughput capacity, the number of UEs, and the number of cells/sectors.
- Al and ML recommendations may inform backup strategies based on alarm types, severity, notifications, and PM counters.
- decision-making for recovery is primarily handled by the SMO 200 (the non-RT RIC 201), the near-RT RIC 202, or the O-CU (e.g., 203 and 204), which determines which O-DU (e.g., standby- 1,..., standby-N) may recover the O-RU 206 based on traffic requirements.
- Factors to consider include the potential failure of the standby O-DU (e.g., standby- 1,..., standby-N), the availability of ports, and the status of the M-Plane.
- Various scenarios dictate the recovery process from resiliency or disaster-related failures, with the orchestrator triggering the appropriate O-DU (e.g., standby- 1 ,... , standby-N) that possesses the desired or available capabilities, which may differ from those of the active O-DU 205.
- subscriber data synchronization ensures consistency across RAN NF instances, such as the O-DU and the O-CU (standby).
- FIGS. 7A-7B illustrates an advanced resiliency use case (partial O-DU failure), according to an embodiment as disclosed herein.
- the advanced resiliency scenario concerning a partial failure of the O-DU, several essential preconditions may be established.
- 01 supervision may be operational between the SMO 200 and the O-DU-1 205a.
- 01 supervision may also be active between the SMO 200 and the O-DU-2 205b.
- the 01 supervision connection between the SMO 200 and the O-CU-CP 204 is up and running.
- the SMO 200 detects a decline in service quality due to the failure of the O-DU-1 205a.
- the SMO 200 decides in the second operation to shift services to the O-DU-2 205b to maintain service continuity.
- the SMO 200 initiates by locking the affected cells and removing any resources associated with the O-DU-1 205a through the 01 interface. This step ensures that the resources linked to the failed unit are no longer in use.
- the O-CU-CP 204 performs one or more actions to deactivate the cells and remove resources via the Fl interface, which connects the O-CU-CP 204 to the O-DUs.
- the O-DU-1 205a sends a notification to the SMO 200 regarding the change in cell status, informing it that the cells are no longer active. This communication occurs through the 01 interface, ensuring the SM 200 is updated on the system’s status.
- the SMO 200 configures a Transport Network Layer (TNL) for the O- DU-2 205b.
- TNL Transport Network Layer
- the O-CU-CP 204 then performs necessary tasks to establish the Fl interface for the O- DU-2 205b, enabling communication between the two units.
- the SMO 200 collaborates with the O-CU-CP 204 to create and configure the new cell using the 01 interface. This operation is vital for preparing the O-DU-2205b to take over the services previously managed by the O-DU-1 205a.
- the SMO 200 configures the cell and its associated carriers for the O- DU-2 205b. If the Fl interface is successfully established, the O-CU-CP 204 manages the messaging needed for the Fl connection, ensuring that the O-DU-2 205b can report back to the O- CU-CP 204. At operation 708, the O-CU-CP 204 informs the SMO 200 that the cell has been activated through the 01 interface.
- the O-DU-2 205b sends a message to the SMO 200 confirming that the cells and their corresponding carriers are now active, indicating that the O-DU-2 205b is ready to provide services.
- the SMO 200 receives subscriptions for performance counters from the O-DU-2 205b. This allows the SMO 200 to monitor the performance of the newly active unit (e.g., O-DU-2205b), ensuring that the O-DU-2205b meets operational standards and can effectively manage the services it has taken over.
- the newly active unit e.g., O-DU-2205b
- FIGS. 8A, 8B and 8C illustrates an 0-DU resiliency use case, according to an embodiment as disclosed herein.
- the 0-DU resiliency use case several critical preconditions may be established to ensure smooth operations.
- a synchronized timing setup is essential across one or more network entities, several essential preconditions may be established.
- 01 supervision is active between the MnS consumer (e.g., SMO 200) and the O-DU-1 205a. This is followed by 01 supervision between the SMO 200 and the O-DU-2 205b, and finally, 01 supervision is operational between the SMO 200 and the O-CU-CP 204.
- the process begins with the initial system setup for the O-DU resiliency use case, which involves a series of interconnected operations, which are given below.
- the SMO 200 which operates as a non-RT RIC, configures the ODUs with designated roles active and standby as well as managing the configuration and activation of the cells.
- a standby O-DU may not be pre-existing. Instead, a new O-DU instance is dynamically created whenever a failure occurs with the currently active O-DU.
- the SMO 200 configures the O-DU-1 205a to take on the active role through the 01 interface.
- the SMO 200 manages the configuration of the [tr]x-array-carriers for the 0-RU 206 via the OFH interface.
- the SMO 200 configures O-DU-2 205b to assume the standby role, again using the 01 interface.
- the O-CU-CP 204 performs a series of actions to activate the cells and carriers associated with the O-DU-1 205a and the 0-RU 206. Following this, at operation 803, notifications are sent out to confirm that the cells and corresponding carriers have been activated. By operation 804, the 0-RU 206 becomes fully operational and is connected to the service, with both the M-plane and CU plane functioning correctly.
- the SMO 200 collects alarms and performance metrics from the network entities (Network Functions (NFs)/ MnS producer), including O-CU-CP 204, O-DU-1 205a, and 0-RU 206, to monitor the system’s health.
- NFs Network Functions
- MnS producer Media Functions
- the O-DU-1 205a fails (either completely or partially) and enters a disabled operational state.
- the SMO 200 detects this failure based on alarms and performance counters received from the network entities.
- the O-CU-CP 204 may initiate one or more operations to deactivate and optionally remove the cells and associated carriers from the O-DU-1 205a.
- notifications regarding the deactivation and removal status of these cells and carriers are communicated, if applicable.
- the SMO 200 configures O-DU-2 205b to take on the active role via the 01 interface.
- the standby O-DU-2 transitions to become the active 0-DU.
- the configuration of [tr]x-array-carriers between O-DU-2 205b and the O- RU 206 is established to ensure proper connectivity.
- the O-CU-CP 204 performs a series of operations to activate the cells and associated carriers linked with O-DU-2 205b.
- the SMO 200 monitors notifications regarding the activation status of cells and carriers from the network entities to ensure everything is functioning as expected.
- the O-RU 206 is operational once again, fully connected and in service, with both the M-plane and CU plane running smoothly.
- the SMO 200 continues to monitor the system by detecting alarms and performance metrics from O-DU-2205b.
- the transition of the active role from O-DU-2 205b back to the O-DU-1 205a involves repeating the same operations as the initial transition, based on triggers from the SMO 200 or decisions made by an operator 800. This interconnected series of operations ensures that the network remains resilient and responsive to failures.
- the 0-DU 205 as the MnS producer, possesses the capability to autonomously detect 01 interface failures and initiate a reset in the absence of management plane functionality resulting from the 01 interface failure. This assumption is made to mitigate uncertainties arising from the management system’s inability to assess the impact of services on the 0-DU 205 without an available management interface, as well as to diagnose and reconfigure the 0-DU 205 using management plane mechanisms.
- the disclosed method provides advanced security measures to protect against cyber threats, ensuring network resilience against attacks. Further, the disclosed method may implement security measures to detect and prevent cyber-attacks that could cause the NF failures. Ensuring secure communication channels between the NFs to prevent data breaches and maintain integrity. [0133] In some example embodiments, the disclosed method may impact various entities, which are given below.
- the 01 interface may handle disruptions gracefully, e.g., robust error handling, retry mechanisms, and support for redundant communication paths to ensure that 0AM commands and data may be exchanged even in the presence of network issues.
- the 01 interface may support seamless failover to maintain continuous communication with network elements. This requires robust session management and state synchronization. b.
- the 0AM architecture may be configured to support distributed and redundant 0AM components, enabling load balancing, and ensuring that critical management functions may continue to operate even if some components fail.
- the 0AM architecture may be defined for performance degradation detection and mitigation and disaster recovery, including the roles and interactions of different components.
- the 0AM architecture may also support automated recovery processes and dynamic reconfiguration to adapt to changing network conditions.
- an 01 Network Resource Management may include models that support redundancy, failover mechanisms, and self-healing capabilities.
- the 01 NRM may be configured to extend the network resource model to include attributes related to performance degradation detection and mitigation and disaster recovery.
- the NRM may include attributes and relationships that facilitate the detection and management of faults, as well as the reallocation of resources to maintain service continuity.
- an 01 PM may include performance measurements related to performance degradation detection and mitigation, such as degradation detection time and mitigation effectiveness and recovery time and data loss in case of disaster recovery. Further, the 01 PM may include implementing mechanisms for data integrity, data buffering, redundant data collection paths, and ensuring that performance measurement systems can continue to operate and provide accurate data even during partial network failures.
- Traffic Engineering (TE) and Inventory (IV) may include inventory management-related impacts to store the resiliency backup. The TE and IV may facilitate the resiliency orchestrator to assign roles for the active and standby systems. E.g., a list of NF Instances or systems that are available to recover from the resiliency are known to the resiliency orchestrator or SMO services.
- the disclosed method enhances redundancy by ensuring multiple pathways and backup systems are in place, facilitating continuous operation even during failures.
- the disclosed method is designed with fault tolerance, allowing it to maintain functionality despite component failures.
- the disclosed method supports scalability, enabling the network to efficiently manage increased loads or changes in demand without sacrificing performance.
- Robust monitoring mechanisms provide real-time surveillance and alerting, ensuring swift detection of issues.
- the incorporation of automated recovery processes streamlines failover actions, minimizing downtime and enhancing high availability. This results in improved user experience, as seamless service delivery is maintained even during network disruptions, ensuring that users remain connected and satisfied.
- FIGS. 9A-9B illustrates 0-DU resiliency scenario in case of 01 failure, according to an embodiment as disclosed herein.
- the 0-DU resiliency scenario outlines the expected system behavior for basic resiliency of the 0-DU in situations where management over the 0-DU as the MnS producer is compromised due to the termination of the 01 interface. This indicates either a complete failure of the Network Function or the Management Plane (01), with the root cause analysis of the 01 failure being outside the scope of this disclosure.
- the scenario emphasizes enhancing the resilience of the 01 interface through robust monitoring, failover mechanisms, and redundancy.
- the roles of various components are defined.
- the 0AM as the MnS consumer
- the MnF of the O-CU-CP 204 as the MnS producer
- the MnFs of O-DU-1 205a (active) and O-DU-2 205b (standby) serving as MnS producers.
- Preconditions for this scenario involve the SMO 200 continuously monitoring the 0-DU- 1 205a, ensuring the 01 interface between the SMO 200 and the O-CU-CP 204 is operational, and having the O-DU-2 205b configured in a standby role or not configured at all (refer to note-1, notes related to this scenario added in the below mentioned table).
- the scenario begins when the SMO 200 identifies its inability to manage the O-DU-1 205a. Initially, at operation 901, the SMO 200 detects an 01 interface failure or a complete failure of O- DU-1 205a through Fault Management (FM) and/or Performance Management (PM) data (refer to note-2 and note-3).
- FM Fault Management
- PM Performance Management
- the SMO 200 then analyzes notifications, faults, and PM counters from the O-CU-CP 204 to determine connectivity between the O-CU-CP 204 and the O-DU via the Fl interface, as well as whether the O-DU has communicated the removal of cell(s) served by the O- DU with a lost 01 interface. This analysis informs the SMO 200’ s decision to transition services from the O-DU-1 205a to the O-DU-2 205b upon detecting failures (refer to note-4).
- the SMO 200 deactivates the cell(s) associated with the O-DU-1 205a on the O-CU-CP 204 via the 01 interface.
- the SMO 200 then configures the O-DU-2 205b for connection with the O-CU-CP 204, providing necessary details such as transport parameters and performance management configurations.
- the O-DU communicates the results of these operations to the SMO 200 through the 01 interface (refer to note-5).
- the SMO 200 continues by creating and configuring the cell (s) associated with the O-DU-2 205b on the O-CU-CP 204, ensuring the O-CU-CP 204 receives the required configurations including transport details, cell(s) configuration, configuration needed to connect with the O-DU-2 205b, configuration for performance management and so on.
- the O-CU-CP 204 informs the results of the above operations to the SMO 200 through the 01 interface
- the SMO 200 creates cell object instances and manages the setup of cell(s) and corresponding carrier resources in the O-DU-2 205b, linking them to the respective cell(s) managed by the O-CU-CP 204.
- the O-DU-2205b then communicates the results of these operations back to the SMO 200 through the 01 interface (refer to note-6 and note-7).
- the SMO 200 sends an activation request for the cell(s) associated with the O-DU-2205b configured successfully in operations 905 and 906 to O-CU-CP 204.
- the O-CU- CP 204 confirms the reception of activation request from SMO 200 through the 01 interface (note- 8).
- O-CU-CP 204 notifies the SMO 200 on their activation status via the 01 interface.
- the O-DU-2 205b informs the SMO 200 of the activation status of the cell(s) and corresponding carrier(s), indicating readiness to provide services.
- the SMO 200 subscribes to and collects alarms and performance counters from the O-DU-2 205b through the 01 interface.
- the scenario concludes when the O-DU-2 205b is ready to accept UEs and deliver services.
- This scenario may include some postconditions, the initiation of functions by the standby 0-DU, acceptance of connections from the O-RU to establish an OFH interface session, and the system’s readiness to accept UEs and provide services, effectively replacing the failed O-DU with the standby 0-DU in the service path.
- FIGS. 10A, 10B and IOC illustrates enhanced network resiliency through a standby O-DU and adaptive cell activation managed by the SMO 200, according to an embodiment as disclosed herein.
- the system can be quickly reconfigured to utilize a standby O-DU.
- the SMO 200 regularly monitors the 01 interface status, and upon detecting a complete and irrecoverable link failure, the SMO 200 automatically triggers a failover to the standby O-DU.
- the SMO 200’ s potential solution involves configuring the standby O-DU to replace the failed one. This scenario impacts service, as illustrated in FIGS. 10A, 10B and 10C.
- the roles of various components are defined.
- the O-RU 206 manages the fronthaul and air interfaces, along with reporting performance metrics and failures. All O-DUs (e.g., active/standby O-DUs) connected to the O-RU 206 are involved in the O-DU resiliency scenario.
- the O-DU receives cell configurations and carrier settings from the SMO 200 and works with the O-CU-CP 204 to activate these configurations.
- the O-CU receives cell-related configurations from the SMO 200 and details about available cells from the O-DUs. Based on the resources reported by the O-DUs and the desired cell availability from the SMO 200, the O-CU manages the activation and deactivation of cells in the O-DUs using the Fl interface. Moreover, the SMO 200 plays a critical role by designating O-DUs as either active or standby and communicating these roles to the O-RU 206 within a hierarchical deployment. It also makes high-level decisions based on prevailing conditions.
- the preconditions for this scenario may include: a.
- the 01 interface between the SMO 200 and O-CU-CP 204 is operational.
- the O-DU-1 205a is configured as active, with the 01 interface to the SMO 200 functioning.
- the O-DU-2 205b is either configured as standby or not configured at all; if available, the 01 interface between the SMO 200 and the O-DU-2 205b is operational.
- the SMO 200 has valid subscriptions for alarms, performance metrics, and notifications for the O-DU-1 205a and the O-DU-2 205b, provided they are connected via the 01 interface.
- the O-DU-1 205a failure detection and recovery process begins with the SMO 200 may monitor the O-DU-1 205a through the 01 interface to identify service degradation by analyzing performance counters, alarms, and notifications.
- the SMO 200 may detect either the 01 interface failure or a complete failure of the O-DU-1 205a using the FM and PM data (refer note-1 and note-2 in below mentioned table).
- the SMO 200 may analyze notifications, faults, and PM counters from the O-DU-1 205a to identify performance issues. For complete failures or 01 interface failures, the SMO 200 may analyze information from the O-CU-CP 204 to verify connectivity via the Fl interface and whether O-DU-1 205a has communicated any lost cells for removal. These findings guide the SMO 200’ s decision to switch services from the O- DU-1 205a to the O-DU-2 205b (refer note-3 in below mentioned table).
- the SMO 200 may request the O-DU-1 205a to lock affected cells and remove corresponding carrier resources through the 01 interface, receiving updates on the status of these operations (refer note-4 and note-5 in below mentioned table).
- the O-DU-1 205a may then deactivate the specified carriers according to Clause 15.3 of the WG4-MP specification and may optionally remove these resources ([tr]x-array-carriers resources).
- the O-DU-1 205a may notify the SMO 200 about the status change for the locked cells, indicated in 1004a (refer note-6 in below mentioned table).
- the SMO 200 may request the O-CU-CP 204 to terminate the Fl interface with the failed O-DU-1 205a (refer note-7 in below mentioned table).
- the SMO 200 may configure the O- DU-2 205b to connect with the O-CU-CP 204 and receives updates on the operation’s status via the 01 interface (refer note-8 in below mentioned table).
- the O-DU-2 205b then becomes the active unit to restore O-RU 206 operations and associated user services (refer note-9 in below mentioned table).
- the SMO 200 may create and configure the cells linked to the O-DU-2 205b on the O-CU-CP 204 using the 01 interface and receives status updates from the O-CU-CP 204 (results of the above-mentioned operations).
- the SMO 200 may configure the O-DU-2 205b with the cells set up on the O-CU-CP 204 in the previous operations (1005), mapping carrier resources to cell resources, and the O-DU-2 205b reports the results of these configurations back to the SMO 200 through the 01 interface (refer note- 10, note-11, and note- 12 in below mentioned table).
- FIG. 11 illustrates a diagram of example components of an apparatus 1100, according to an embodiment as disclosed herein.
- the apparatus 1100 comprises a processor 1110, a memory 1120, a storage component 1130, an input component 1140, an output component 1150, a communication interface 1160, and a bus 1170.
- the apparatus 1100 may relate to at least one of , for example, the SMO 200, the O-RU 206, the O-DU 205, the O-CU (e.g., 203 and 204), an interface entity, a transport entity, the O-cloud 207, the RIC (e.g., 201 and 202), and the O-eNB 208.
- the processor 1110 means any type of computational circuit that may comprise hardware elements and software elements.
- the processor 1110 may be embodied as a multi-core processor, a single core processor, or a combination of one or more multi-core processors and/or one or more single core processors, a distributed processing system, or the like.
- the processor 1110 may be a Central Processing Unit (CPU), a Graphics Processing Unit (GPU), an Accelerated Processing Unit (APU), an Application-Specific Integrated Circuit (ASIC), or another type of processing component.
- CPU Central Processing Unit
- GPU Graphics Processing Unit
- APU Accelerated Processing Unit
- ASIC Application-Specific Integrated Circuit
- the input component 1140 is configured to receive information, such as user input.
- the input component 1140 may include, but not be limited to, a touch screen display, a keyboard, a keypad, a mouse, a button, a switch, and/or a microphone.
- the input component 1140 may include a sensor for sensing information (e.g., a Global Positioning System (GPS), an accelerometer, a gyroscope, and/or an actuator).
- GPS Global Positioning System
- the output component 1150 is configured to provide output information from the apparatus 1100.
- the output component 1150 may be, but is not limited to, a display, a speaker, instructions to an external device, and/or one or more Light-Emitting Diodes (LEDs).
- LEDs Light-Emitting Diodes
- the apparatus 1100 may include additional components, fewer components, different components, or differently arranged components than those shown in FIG. 11. Additionally, or alternatively, a set of components (e.g., one or more components) of the apparatus 1100 may perform one or more functions described as being performed by another set of components of the apparatus 1100. Further, one or more method steps described in any of the embodiments may be performed utilizing the apparatus 1100 in communication with one another.
- a method comprising: monitoring, at a MnS consumer, data of one or more interfaces or one or more network entities associated with an Open Radio Access Network (O-RAN) architecture, wherein monitoring the data comprises monitoring at least one of Fault, Configuration, Accounting, Performance, Security (FCAPS) data, Performance Measurement (PM) counters, Fault Management (FM) data, a notification message, and an alarm; detecting a failure occurring in at least one of the one or more interfaces or the one or more network entities associated with the O-RAN architecture based on the monitoring; and performing, upon detection of the failure, at least one recovery action.
- OFP Open Radio Access Network
- performing the at least one recovery action to mitigate one or more impacts of the detected failure comprises: in response to detecting the failure occurring at a transport level, performing the at least one recovery action comprises: executing one or more transport-based recovery mechanisms, for data transmissions, by distributing traffic and workloads across a plurality of network entities for load balancing.
- performing the at least one recovery action to mitigate one or more impacts of the detected failure comprises: in response to detecting the failure occurring at a RAN Intelligent Controller (RIC) level, performing the at least one recovery action comprises: executing one or more recovery actions comprises performing a software reset, implementing software auto-healing, redeploying, or re-onboarding one or more applications by utilizing backup information provided by the MnS consumer.
- RIC RAN Intelligent Controller
- QoS Quality of Service
- QoE Quality of Experience
- An apparatus configured to: periodically monitor, at a MnS consumer, data of one or more interfaces or one or more network entities associated with an Open Radio Access Network (O-RAN) architecture, wherein monitoring the data comprises monitoring at least one of Fault, Configuration, Accounting, Performance, Security (FCAPS) data, Performance Measurement (PM) counters, Fault Management (FM) data, a notification message, and an alarm; detect a failure occurring in at least one of the one or more interfaces or the one or more network entities associated with the O-RAN architecture based on the monitoring; and perform, upon detection of the failure, at least one recovery action.
- O-RAN Open Radio Access Network
- At least one recovery action comprises one or more backup mechanisms and one or more redundant node mechanisms; and wherein the one or more backup mechanisms comprise a software-based backup mechanism, a hardware-based backup mechanism, and a transport network-based backup mechanism.
- the one or more network entities comprises at least one of a non-real-time RAN Intelligent Controller (RIC), a near-real-time RIC, an O-RAN Central Unit User Plane (O-CU-UP), an O-RAN Central Unit Control Plane (O-CU-CP), an O-RAN Distributed Unit (O-DU), an O-RAN Radio Unit (O-RU), an O-RAN cloud (O-cloud), and an O-RAN enhanced Node B (O-eNB); and wherein the one or more interfaces comprises at least one of an Al interface, an 01 interface, an 02 interface, an E2 interface, an R1 interface, a Y1 interface, an Fl-C interface, an Fl-U interface, an open Fronthaul for Control and User plane (FH-CUS), and an Open Fronthaul for Management (FH-M) plane.
- RIC non-real-time RAN Intelligent Controller
- O-CU-UP O-RAN Central Unit User Plane
- O-CU-CP O-RAN Central Unit Control Plane
- OF-DU O-RAN
- the apparatus is configured to: in response to detecting the failure occurring at an O-RAN Radio Unit (O-RU) within the O-RAN architecture, perform, at the O-RU, the at least one recovery action comprises: execute one or more software auto-healing mechanisms, wherein the one or more software auto-healing mechanisms comprise performing software upgrades and restarting the O-RU, upon an availability of a Management Plane (M-plane); or adjust one or more neighboring O-Rus configuration to provide support for mitigating one or more coverage gaps that occurred due to the failure.
- O-RU O-RAN Radio Unit
- M-plane Management Plane
- the apparatus is configured to: in response to detecting the failure occurring at either an O-RAN Distributed Unit (O-DU) or an O-RAN Central Unit (O-CU) within the O-RAN architecture, perform, at the corresponding O-DU or O-CU, the at least one recovery action comprises: execute one or more software auto-healing mechanisms when an 01 interface or an E2 interface is operational and one or more active sessions are established in the O-RAN architecture; execute, by utilizing at least one Artificial Intelligence (Al) model, at least one a manual rehoming process, and an automated rehoming process based on Fault Management (FM) and Performance Management (PM) data, for predictive autohealing and failure forecasting; trigger a redundant O-DU to take over one or more services for the O-RU; or execute one or more transport layer-based backup recovery mechanisms for either the O-DU or the O-CU (Al) model, at least one a manual rehoming process, and an automated rehoming process based on Fault Management (FM) and Performance Management (
- the apparatus is configured to: in response to detecting the failure occurring at an interface level, perform the at least one recovery action comprises: execute one or more backup interface-based recovery mechanisms for data transmissions for a plurality of network interfaces, wherein the plurality of network interfaces comprises at least one of an Al interface, an 01 interface, an 02 interface, an E2 interface, an R1 interface, a Y1 interface, an Fl-C interface, an Fl-U interface, an open Fronthaul for Control and User plane (FH-CUS), and an Open Fronthaul for Management (FH-M) plane; execute at least one redundant transmission mechanism for data transmissions for the plurality of network interfaces; or identify a cause of the interface failure, wherein the cause is determined to be either a malfunction in a Network Function (NF) or Network Element (NE) or a failure of an Access Point (AP), and recovering the AP to restore functionality
- NF Network Function
- NE Network Element
- AP Access Point
- the apparatus as described in any of [15]-[20], wherein to perform the at least one recovery action to mitigate the one or more impacts of the detected failure, the apparatus is configured to: in response to detecting the failure occurring at a transport level, perform the at least one recovery action comprises: execute one or more transport-based recovery mechanisms, for data transmissions, by distributing traffic and workloads across a plurality of network entities for load balancing.
- the apparatus is configured to: in response to detecting the failure occurring at an O-RAN cloud level, perform the at least one recovery action comprises: execute one or more Network Function (NF) recovery mechanisms by utilizing distinct standby NF instances or backup clusters for various networks; implement infrastructure-level resiliency, by employing at least one open- source container orchestration system, for software auto-healing at the O-RAN cloud level; or execute one or more auto-healing processes, wherein the one or more autohealing processes are managed by an O-RAN cloud and one or more NF resiliency features.
- NF Network Function
- the apparatus as described in any of [15]-[22], wherein to perform the at least one recovery action to mitigate the one or more impacts of the detected failure, the apparatus is configured to: in response to detecting the failure occurring at a RAN Intelligent Controller (RIC) level, perform the at least one recovery action comprises: execute one or more recovery actions comprising performing a software reset, implementing software auto-healing, redeploying, or re-onboarding one or more applications by utilizing backup information provided by the MnS consumer.
- RIC RAN Intelligent Controller
- the apparatus is configured to: perform, based on one or more additional inputs from one or more cell-network planning tools, at least one action comprises: assign specific operational roles to each network entity comprising an active role for the O-DU and a standby role for the O-DU; recommend a direction to a parent node regarding the effective assignment and management of the active role and the standby role associated with the O-DU.
- the apparatus is configured to, at least one of: orchestrate one or more resiliency measures through both operator-driven and data- driven policies, wherein the operator-driven policy establishes recovery configurations, and the data-driven policy manages load distribution during one or more overload situations; and recommending one or more strategies, to a near-real-time RIC or non-real-time RIC, to prevent backup entities from entering an energy-saving mode, to ensure readiness for immediate operational demands.
- the apparatus is configured to: predict, by utilizing at least one Artificial Intelligence (Al) model, the failure occurring at the at least one network entity based on the received network-related data and one or more parameters, to perform the at least one recovery action in advance, wherein the one or more parameters comprise a list of impacted network entities among a plurality of network entities, current traffic information associated with the at least one network entity, a Quality of Service (QoS) requirement for each network entity, a Quality of Experience (QoE) requirement for each network entity, applied network traffic information associated with the at least one network entity, and resource requirement information associated with the at least one network entity.
- QoS Quality of Service
- QoE Quality of Experience
- the apparatus as described in any of [ 15]-[27], prior to perform the at least one recovery action, the apparatus is configured to: determine a type of the failure occurring at the at least one network entity based on the network-related data; determine an available software-hardware support details at the MnS consumer; and determine the at least one recovery action based on the type of the determined failure and the available software-hardware support details at the MnS consumer.
- a non-transitory computer-readable medium storing instructions, the instructions comprising: one or more instructions that, when executed by an apparatus, the apparatus comprising one or more processors, cause the one or more processors to: periodically monitor, at a MnS consumer, data of one or more interfaces or one or more network entities associated with an Open Radio Access Network (O- RAN) architecture, wherein monitoring the data comprises monitoring at least one of Fault, Configuration, Accounting, Performance, Security (FCAPS) data, Performance Measurement (PM) counters, Fault Management (FM) data, a notification message, and an alarm; detect a failure occurring in at least one of the one or more interfaces or the one or more network entities associated with the O-RAN architecture based on the monitoring; and perform, upon detection of the failure, at least one recovery action.
- O- RAN Open Radio Access Network
- the embodiments disclosed herein can be implemented through at least one software program running on at least one hardware device and performing network management functions to control the elements.
- the elements can be at least one of a hardware device or a combination of hardware devices and software modules.
Landscapes
- Engineering & Computer Science (AREA)
- Computer Networks & Wireless Communication (AREA)
- Signal Processing (AREA)
- Environmental & Geological Engineering (AREA)
- Mobile Radio Communication Systems (AREA)
Abstract
Method includes monitoring, at a MnS consumer (e.g., Service Management and Orchestration (SMO)), data of one or more interfaces, or one or more network entities associated with an Open Radio Access Network (O-RAN) architecture. Monitoring the data comprises monitoring at least one of Fault, Configuration, Accounting, Performance, Security (FCAPS) data, Performance Measurement (PM) counters, Fault Management (FM) data, a notification message, and an alarm. The method includes detecting a failure occurring in at least one of the one or more interfaces or one or more network entities based on the monitoring. The method includes performing, upon detection of the failure, at least one recovery action.
Description
END-TO-END NETWORK RESILIENCY FRAMEWORK
CROSS-REFERENCE TO RELATED APPLICATION (S)
[0001] This application claims priority to Indian Provisional Patent Application Number 202411058170, filed on July 31, 2024, and Indian Non-Provisional Patent Application Number 202411058170, filed on May 12, 2025, the entire contents of which are incorporated herein by reference.
FIELD
[0002] The present disclosure relates to an end-to-end network resiliency framework.
BACKGROUND
[0003] The information disclosed in this background section is only for enhancement of understanding of the general background of the disclosure and should not be taken as an acknowledgement or any form of suggestion that this information forms the prior art already known to a person skilled in the art.
[0004] Open Radio Access Network (O-RAN) architecture represents a flexible, open-source framework for mobile networks. Separation of hardware and software components allows diverse vendors to collaborate or support multi-vendor deployments. Such an approach promotes innovation, reduces costs, and enhances network performance by enabling interoperability and scalability across various network elements and services. However, such deployments face significant challenges in ensuring end-to-end network resiliency due to the diversity of components and vendors involved, as failures can arise at multiple levels within the O-RAN architecture. For instance, an O-RAN Radio Unit (O-RU) level, an O-RAN Distributed Unit (O- DU) level, an O-RAN Central Unit (O-CU) level, an interface level, a transport level, an O-RAN cloud level, a RAN Intelligent Controller (RIC) level, and an Evolved Node B (eNB) level.
SUMMARY
[0005] This summary is provided to introduce a selection of concepts, in a simplified format, that are further described in the detailed description of the disclosure. This summary is neither intended to identify key or essential inventive concepts of the disclosure nor is it intended for determining the scope of the disclosure.
[0006] According to one embodiment of the present disclosure, a method is disclosed. The method includes monitoring, at a MnS consumer, data of one or more interfaces or one or more network entities associated with an Open Radio Access Network (0-RAN) architecture. Monitoring the data comprises monitoring at least one of Fault, Configuration, Accounting, Performance, Security (FCAPS) data, Performance Measurement (PM) counters, Fault Management (FM) data, a notification message, and an alarm. The method further includes detecting a failure occurring in at least one of the one or more interfaces or the one or more network entities associated with the O- RAN architecture based on the monitoring. The method further includes performing, upon detection of the failure, at least one recovery action.
[0007] According to one embodiment of the present disclosure, an apparatus is disclosed. The apparatus is configured to periodically monitor, at a MnS consumer, data of one or more interfaces or one or more network entities associated with an Open Radio Access Network (O-RAN) architecture. Monitoring the data comprises monitoring at least one of Fault, Configuration, Accounting, Performance, Security (FCAPS) data, Performance Measurement (PM) counters, Fault Management (FM) data, a notification message, and an alarm. The apparatus is further configured to detect a failure occurring in at least one of the one or more interfaces or the one or more network entities associated with the O-RAN architecture based on the monitoring. The apparatus is further configured to perform, upon detection of the failure, at least one recovery action.
[0008] According to one embodiment of the present disclosure, a non-transitory computer- readable medium storing instructions is disclosed. The one or more instructions are executed by an apparatus which comprises one or more processors. The one or more processors may periodically monitor, at a MnS consumer, data of one or more interfaces or one or more network entities associated with an Open Radio Access Network (O-RAN) architecture. Monitoring the
data comprises monitoring at least one of Fault, Configuration, Accounting, Performance, Security (FCAPS) data, Performance Measurement (PM) counters, Fault Management (FM) data, a notification message, and an alarm. The one or more processors may detect a failure occurring in at least one of the one or more interfaces or the one or more network entities associated with the O-RAN architecture based on the monitoring. The one or more processors may perform, upon detection of the failure, at least one recovery action.
[0009] To further clarify the advantages and features of the present disclosure, a more particular description of the disclosure will be rendered by reference to specific embodiments thereof, which are illustrated in the appended drawings. It is appreciated that these drawings depict only typical embodiments of the disclosure and are therefore not to be considered limiting of its scope. The disclosure will be described and explained with additional specificity and detail in the accompanying drawings.
BRIEF DESCRIPTION OF THE DRAWINGS
[0010] Features, aspects, and advantages of embodiments of the disclosure will be described below with reference to the accompanying drawings, in which like reference numerals denote like elements, and wherein:
FIG. 1 is a diagram of an example of an implementation environment in which systems and/or methods, described herein, may be implemented, according to an embodiment as disclosed herein;
FIG. 2 illustrates an example block diagram depicting a system of resiliency at various levels of an Open Radio Access Network (O-RAN) architecture, according to an embodiment as disclosed herein;
FIG. 3 is a flow diagram illustrating a method for managing an end-to-end network resiliency within the O-RAN architecture, according to an embodiment as disclosed herein; FIG. 4 is a flow diagram illustrating a method for determining at least one recovery action, according to an embodiment as disclosed herein;
FIG. 5 illustrates an example scenario where a Service Management and Orchestration (SMO) detects a failure occurring at an O-RAN Distributed Unit (0-DU) level, according to an embodiment as disclosed herein;
FIG. 6 illustrates an exemplary backup scenario associated with an O-DU, according to an embodiment as disclosed herein;
FIGS. 7A-7B illustrates an advanced resiliency use case (partial O-DU failure), according to an embodiment as disclosed herein;
FIGS. 8A, 8B and 8C illustrates an O-DU resiliency use case, according to an embodiment as disclosed herein;
FIGS. 9A-9B illustrates O-DU resiliency scenario in case of 01 failure, according to an embodiment as disclosed herein;
FIGS. 10A, 10B and IOC illustrates enhanced network resiliency through a standby O-DU and adaptive cell activation managed by the SMO, according to an embodiment as disclosed herein; and
FIG. 11 illustrates a diagram of example components of an apparatus, according to an embodiment as disclosed herein.
DETAILED DESCRIPTION
[0011] The following detailed description of example embodiments refers to the accompanying drawings. The foregoing disclosure provides illustration and description, but is not intended to be exhaustive or to limit the implementations to the precise form disclosed. Modifications and variations are possible in light of the above disclosure or may be acquired from practice of the implementations. Further, one or more features or components of one embodiment may be incorporated into or combined with another embodiment (or one or more features of another embodiment). Additionally, the flowchart and description of operations provided below relate to one of the various embodiments. It should be noted that it is possible to make other embodiments that do not exactly match the flowchart and its description. It is understood that in other embodiments one or more operations may be omitted, one or more operations may be added, one or more operations may be performed simultaneously (at least in part).
[0012] It will be apparent that systems and/or methods, described herein, may be implemented in different forms of hardware, software, or a combination of hardware and software. The actual specialized control hardware or software code used to implement these systems and/or methods is
not limiting of the implementations. Thus, the operation and behavior of the systems and/or methods are described herein without reference to specific software code. It is understood that software and hardware may be designed to implement the systems and/or methods based on the description herein.
[0013] Even though particular combinations of features are recited in the claims and/or disclosed in the specification, these combinations are not intended to limit the disclosure of implementations. In fact, many of these features may be combined in ways not specifically recited in the claims and/or disclosed in the specification. Although each dependent claim listed below may directly depend on only one claim, the disclosure of implementations includes each dependent claim in combination with every other claim in the claim set.
[0014] No element, act, or instruction used herein should be construed as critical or essential unless explicitly described as such. Also, as used herein, the articles “a” and “an” are intended to include one or more items, and may be used interchangeably with “one or more.” Also, as used herein, the terms “has,” “have,” “having,” “include,” “including,” or the like are intended to be open-ended terms. Further, the phrase “based on” is intended to mean “based, at least in part, on” unless explicitly stated otherwise. Furthermore, expressions such as “at least one of [A] and [B],” “[A] and/or [B],” or “at least one of [A] or [B]” are to be understood as including only A, only B, or both A and B.
[0015] The foregoing disclosure provides illustration and description but is not intended to be exhaustive or to limit the implementations to the precise form disclosed. Modifications and variations are possible in light of the above disclosure or may be acquired from practice of the implementations.
[0016] In the present disclosure, specific tasks may be performed using Artificial Intelligence/Machine Learning (AI/ML) models. The AI/ML model is a model generated using one or more Al technologies, one or more ML algorithms, or both, and generates output data based on input data. This output data is used to perform tasks. Tasks performed using AI/ML models include those generally referred to as intellectual tasks, such as classification, prediction, natural language processing, etc.
[0017] Although Al and ML are explained separately, ML is a technology included in Al. In ML, instead of being explicitly programmed for a specific task, systems can improve their performance over time by identifying patterns and making inferences from training data. Typically, the generation of ML models includes data collection, model training, and model inference. Data collection involves gathering and preprocessing data to be used for training and inference. Model training involves developing and validating models using the collected data. Model inference involves applying the trained models to new data to generate new output data and perform tasks.
[0018] Machine learning includes various types of learning methods such as supervised learning, unsupervised learning, reinforcement learning, semi-supervised learning, self-supervised learning, transudative learning, transfer learning, meta learning, and the like. These types of learning methods can be appropriately selected according to the embodiments. Unless otherwise specified, the application of types not mentioned in this description is not precluded. Additionally, the structure of ML models may vary depending on the embodiments and learning methods, and is not limited to the methods disclosed. Furthermore, ML includes deep learning, which uses models that include neural networks. Deep learning models may include, for example, deep neural networks (DNNs), convolutional neural networks (CNNs), etc.
[0019] It should be noted that the AI/ML models presented hereinafter are examples and are not limited to the illustrated AI/ML models. They can be modified or altered by using different Al or ML algorithms. The configuration of the neural network is not limited to the configuration disclosed in the present disclosure and can be modified.
[0020] In the context of an end-to-end network resiliency, various mechanisms are implemented to address resiliency scenarios for different operations. For instance, high availability for Management and Orchestration Service (MnS) producer/network functions, including an Open Radio Access Network (O-RAN) Distributed Unit (O-DU), an O-RAN Central Unit (O-CU), Next Generation Node B (gNB)-DU, and gNB-CU, can be achieved through software auto-healing facilitated by orchestration processes. For another instance, within an O-cloud architecture, internal mechanisms support Point of Delivery (POD)-level and cluster-level recovery, leveraging Kubernetes orchestration capabilities.
[0021] For another instance, existing Third Generation Partnership Project (3GPP) standards provide support for multiple Stream Control Transmission Protocol (SCTP) associations between Central Unit Control Plane (CU-CP) and DU for an Fl interface (as outlined in 3GPP TS 38.472) and between the CU-CP and Central Unit User Plane (CU-UP) for an El interface (referenced in 3GPP TS 38.462). For another instance, redundancy at a network layer can be accomplished through multiple transport paths, ensuring data delivery even in the event of a path failure, while also enabling flexible and efficient routing of data packets. For another instance, in cloud deployment scenarios, port redundancy can be realized through cloud orchestration and platform solutions, such as those based on Kubemetes.
[0022] However, several challenges are encountered in the existing end-to-end network resiliency, which are mentioned herein. The latest 0-RAN specifications do not encompass end-to-end resiliency, and vendors currently offer solutions limited to their specific products or services. Additionally, there is a lack of a centralized management mechanism for resiliency, such as through a Service Management and Orchestration (SMO) or a RAN Intelligent Controller (RIC). The 0-RAN ecosystem suffers from a fragmented approach to managing end-to-end network resiliency across various components and vendors. This fragmentation leads to inconsistent resiliency levels, increased complexity in managing multi-vendor environments, challenges in coordinated recovery and fault management, and a deficiency in centralized oversight and control to optimize network performance and reliability.
[0023] To address these identified shortcomings and provide a robust alternative for the end-to- end network resiliency, a disclosed method is proposed, as discussed throughout the disclosure (FIGS. 1 to 8). The disclosed method introduces a unified resiliency service managed by the MnS consumer (e.g., SMO), offering a comprehensive solution that encompasses all entities involved. Additionally, the disclosed method considers multiple aspects, including potential failure scenarios at various levels, backup and recovery mechanisms, necessary hardware and software support, and enhancements driven by Al/ML model(s).
[0024] Referring now to the drawings, and more particularly to FIGS. 1 to 11, where similar reference characters denote corresponding features consistently throughout the figures, there are shown preferred embodiments.
[0025] FIG. 1 is a diagram of an example of an implementation environment 100 in which systems and/or methods, described herein, may be implemented, according to an embodiment as disclosed herein.
[0026] The implementation environment 100 includes a User Equipment (UE) 110, a service environment 120, and a network 130. The service environment 120 and the network 130 may relate to an Open Radio Access Network (O-RAN) architecture, as shown in FIG. 2. The service environment 120 includes one or more sub-environments 121-1 to 121-N (collectively and/or interchangeably referred hereinafter as 121). To illustrate this, FIG. 1 shows, for convenience, examples of a 1st sub-environment 121-1, a 2nd sub-environment 121-2, and an N111 subenvironment 121-N (where N is any natural number).
[0027] The UE 110 is connected to the network 130, and the network 130 is connected to the service environment 120. The connections may be wired, wireless, or a combination of both wired and wireless. The UE 110 and the service environment 120 are connected via the network 130.
[0028] The UE 110 is a device that communicates with the service environment 120. The UE 110 receives information from the service environment 120 and/or sends information to the service environment 120. Also, the UE 110 may generate and/or store information to be transmitted, as necessary. Also, the UE 110 may store and/or process information that is received, as necessary.
[0029] The example FIG. 1 refers to the “UE”. However, it should be understood by those skilled in the art that general terms such as “user device”, “terminal”, “terminal device”, “communication device”, and “communication terminal” can be used interchangeably with the term “UE”.
[0030] For example, the UE 110 may include a computing device (e.g., a desktop computer, a laptop computer, a tablet computer, a handheld computer, a smart speaker, a server, etc.), a mobile phone (e.g., a smartphone, a radiotelephone, etc.), a wearable device (e g., a pair of smart glasses or a smart watch), or a similar device.
[0031] The service environment 120 is an environment that communicates with the UE 110 to provide one or more services. The service environment 120 receives information from the UE 110 and/or sends information to the UE 110. Also, the service environment 120 may generate and/or store information to be transmitted, as necessary. Also, the service environment 120 may store and/or process information that is received, as necessary. For example, the service environment
120 may provide computing resources as one of the services. It should be noted that the service is not limited to being provided to the UE 110; it may also be provided to devices other than the UE 110. For example, based on communication from the UE 110, the service may perform processes such as anomaly detection or traffic analysis and notify the results to a predetermined destination. [0032] The example FIG. 1 refers to the “service environment”. The term “service environment” is used to refer to the broader context within which services operate. For example, cloud environments, platforms, computing systems, network systems, and cloud systems generally represent the environments in which services are conducted, and these are included within the “service environment”. However, the “service environment” is not limited to these examples. Additionally, the specific types of environments within the “service environment” are not restricted. For instance, cloud environments and cloud systems can be categorized as private cloud, public cloud, hybrid cloud, or multi-cloud, all of which are included within the “service environment”.
[0033] The one or more services provided by the service environment 120 is not specifically limited and can be adjusted according to the embodiments. For example, the one or more services may include a service that provides information to the UE 110, a service that stores information from the UE 110, or a service that performs processing based on information from the UE 110 and returns the results of the processing.
[0034] In an embodiment, the service environment 120 may also provide computing resources. The computing resources can be hardware resources and/or software resources. For example, applications, processors, memory, and storage can be included in the provided computing resources. Each computing resource can communicate with other computing resources via wired connections, wireless connections, or a combination of wired and wireless connections.
[0035] The provided computing resources can be actual resources (also referred to as physical resources) and/or virtual resources. Furthermore, means of virtualization for virtual resources can be selected as appropriate. That is, in this disclosure, the use of adjectives such as “virtual” or “virtualized” to describe names does not imply that they are virtualized by a specific means of virtualization. For example, “virtual machine” refers to software that operates like an actual computer, realized through means of virtualization, and it is not intended to exclude those realized
by specific means of virtualization such as hypervisors or containers. Conversely, when means of virtualization such as hypervisors or containers are mentioned in this disclosure, it is merely cited as a general method of implementation. It should also be interpreted that embodiments implemented with other virtualization means are also disclosed. Also, the services may also be provided using resources virtualized by different means.
[0036] The service environment 120 includes one or more devices, such as servers and network devices, which provide services or perform processes. The placement of these devices within the service environment 120 can be determined as appropriate. Additionally, if the service environment 120 includes one or more sub-environments 121, the placement of devices can be determined based on predetermined policies for each sub-environment 121. For example, devices related to the first service may be placed in the 1st sub-environment 121-1, and devices related to the second service may be placed in the 2nd sub-environment 121-2. In another example, devices expected to have a higher load than a predetermined threshold may be placed in the 1st subenvironment 121-1, while devices expected to have a lower load than the predetermined threshold may be placed in the 2nd sub-environment 121-2. In this way, specific devices can be placed in specific sub-environments 121. Conversely, each sub-environment 121 can be specialized for a particular purpose.
[0037] In an embodiment, all processes executed in a single service may run within a single service environment, or in multiple service environments. Multiple processes executed in a single service could be provided by different service environments.
[0038] The network 130 is a network that exchanges information between the UE 110 and the service environment 120. The network 130 includes one or more wired and/or wireless networks. [0039] For example, the network 130 may include a cellular network (e.g., a Fifth Generation (5G) network, a Long-Term Evolution (LTE) network, a Third Generation (3G) network, a Code Division Multiple Access (CDMA) network, etc.), a Public Land Mobile Network (PLMN), a Local Area Network (LAN), a Wide Area Network (WAN), a Metropolitan Area Network (MAN), a telephone network (e.g., the Public Switched Telephone Network (PSTN)), a private network, an ad hoc network, an intranet, the Internet, a fiber optic-based network, or the like, a NonTerrestrial Network (NTN), and/or a combination of these or other types of networks.
[0040] The network 130 can be a part of a network. For example, in a 5G network that includes a RAN, a transport network, and a core network, the network 130 can be at least one of the RAN, the transport network, or the core network. For example, the service environment 120 could be in the core network, in which case the network 130 could correspond to a network that is a combination of a RAN and a transport network and is part of the 5G network.
[0041] The number and arrangement of devices and networks shown in FIG. 1 are provided as an example. It should be understood that any changes that may be implemented by those skilled in the art, such as the addition or rearrangement of well-known devices or networks at the time of implementation, are included in this disclosure.
[0042] FIG. 2 illustrates an example block diagram depicting a system of resiliency at various levels of the Open Radio Access Network (0-RAN) architecture, according to an embodiment as disclosed herein. The 0-RAN architecture comprises one or more network entities and one or more interfaces for communication and management.
[0043] In some example embodiments, the one or more network entities may include, but not limited to, the Service Management and Orchestration (SMO) 200, a non-real-time RAN Intelligent Controller (RIC) 201, a near-real-time RIC 202, an O-RAN Central Unit User Plane (O-CU-UP) 203, an O-RAN Central Unit Control Plane (O-CU-CP) 204, the 0-RAN Distributed Unit (O-DU) 205, an O-RAN Radio Unit (O-RU) 206, the O-RAN cloud (O-cloud) 207, and an O-RAN enhanced Node B (O-eNB).
[0044] In some example embodiments, the one or more interfaces may include, but not limited to, an Al interface, an 01 interface, an 02 interface, an E2 interface, an R1 interface, a Y1 interface, an Fl-C interface, an Fl-U interface, an open Fronthaul for Control and User plane (FH-CUS), and an Open Fronthaul for Management (FH-M) plane.
[0045] In some example embodiments, the near-real-time RIC 202 is a logical function that enables near-real-time control and optimization of O-RAN elements and resources via fine-grained data collection and actions over the E2 interface. The non-real-time RIC 201 is a logical function that enables non-real-time control and optimization of the O-RAN elements and resources, AVML workflow including model training and updates, and policy-based guidance of applications/features in the near-real-time RIC 202.
[0046] In some example embodiments, the Al interface facilitates communication between the non-real-time RIC 201 and the O-RAN elements, allowing for effective policy management and optimization. The E2 interface connects the near-real-time RIC 202 to the O-RAN elements (e.g., the O-DU 205, the O-CU-CP 204, and the O-CU-UP 203), supporting real-time control.
[0047] In some example embodiments, X2-c and X2-u interfaces enable communication between eNBs for load balancing and handover procedures, while the Fl-C interface and Fl-u interface connect the O-DU 205, the O-CU-CP 204, the O-CU-UP 203 in a 5G framework, optimizing data transfer and control signaling.
[0048] In some example embodiments, the architecture features the open FH, which includes the open FH CUS for managing fronthaul transport networks and the open FH M-Plane for control plane traffic management. The O-RU 206 serves as a physical layer interfacing with antennas for radio signal processing, while the O-Cloud 207 provides a virtualized environment for hosting the RAN elements, enhancing scalability and flexibility. Finally, the O-eNB 208 represents a traditional base station connecting user devices to the network. Each element and interface play an integral role in ensuring efficient network operations, allowing service providers to deliver high- quality services.
[0049] In some example embodiments, the disclosed method or the MnS consumer (e.g., SMO 200) provides a comprehensive end-to-end resiliency solution for one or more network levels within the O-RAN architecture, as described in conjunction with FIG. 3.
[0050] The one or more network levels within the O-RAN architecture may include, but are not limited to, an O-RU level, an O-DU level, an O-CU level (e.g., the O-CU-CP 204 and O-CU-UP 203), an interface level (e.g., Al, 01, 02, E2, and Rl), a transport level, an O-RAN cloud level, a RIC level, and an eNB level. In one embodiment, the transport level may involve both network layer and port redundancy. In one embodiment, the O-RAN cloud level may involve an infrastructure layer that encompasses components like pods, clusters, Central Processing Units (CPUs), Graphics Processing Units (GPUs), and data centers and an application layer includes network functions or elements, and management functions. In one embodiment, the RIC level may involve both non-RT RIC and near-RT RIC functionalities.
[0051] In some example embodiments, the SMO 200 maintains an inventory database of deployed network components, detailing hardware and software aspects of network functions and elements. Utilizing this information, along with inputs from cell and network planning tools, the SMO 200 performs critical tasks, such as assigning roles to network functions and elements.
[0052] For instance, the SMO 200 designates active and standby roles for O-DUs to ensure O-DU resiliency. One or more Operations, Administration, and Management (0AM) functions of the SMO 200 can provide directives to parent nodes, such as the near-real-time RIC 202 or O-CU (e.g., 204 and 203), for managing these roles effectively.
[0053] In some example embodiments, the SMO 200 continuously monitors and manages the inventory database of recovery options (recovery action) for various failure scenarios (failure).
[0054] For instance, the SMO 200 may track high-ranked neighbors for each cell to mitigate coverage gaps when a cell experiences downtime.
[0055] In some example embodiments, the SMO 200 also maintains the inventory database of software releases and relevant updates. Operator-driven and data-driven policies orchestrate resiliency steps within the SMO 200.
[0056] For instance, one or more operator-based policies can define system configurations for recovery during emergencies, while one or more data-driven policies may facilitate load distribution across multiple network functions in the event of overload conditions.
[0057] In some example embodiments, the SMO 200 may also guide the near-real-time RIC 202 or the non-real-time RIC 201 to prevent backup entities from entering energy-saving states.
[0058] In some example embodiments, in the context of fault and failure detection, the SMO 200 may consume Fault, Configuration, Accounting, Performance, Security (FCAPS) data, and leverages AI/ML enhancements to detect or predict failures. A resiliency orchestrator (not shown in FIG. 2) within the SMO 200 is configured to process this data to identify failures. By analyzing historical FCAPS information, the SMO 200 can anticipate failures and plan recovery actions proactively, as described in conjunction with FIG. 3.
[0059] In some example embodiments, resiliency actions are informed by available data within the SMO 200, with the capability to detect both complete and partial failures based on alarms (e.g., lost O-DU ID supervision or transceiver faults) and notifications regarding events, configurations,
or failures (e.g., carrier state change notification). These alarms and notifications can originate from various network functions, network elements, or external applications/ platforms (e.g., sleeping cell-detected notification or alarm).
[0060] In some example embodiments, performance monitoring is conducted, by the SMO 200, through Performance Measurement (PM) counters (e.g., which track metrics like Energy, Power and Environmental (ePE) statistics, user equipment distribution, downlink Fl -U packet loss rates, and canceled Downlink Control Information (DCI) due to Physical Downlink Control Channel (PDCCH) resource shortages).
[0061] In some example embodiments, upon receiving failure alarms or notifications, the SMO 200 initiates recovery actions based on the information stored in its databases. The nature of the fault or failure impacts the SMO 200's decision-making regarding the recovery process and the available software or hardware support details.
[0062] For instance, if a node or network function fails, the SMO 200 may attempt recovery through relevant software upgrades, resets, or auto-healing mechanisms, as applicable.
[0063] Based on the above-mentioned embodiment(s), the SMO 200 may provide the end-to-end resiliency solution to ensure that various components of the network, including RAN nodes and interfaces, are effectively monitored and managed, minimizing downtime and maintaining service continuity. The SMO 200 may perform rapid identification and mitigation of potential failures, thus improving network reliability, as one of the advantages of the disclosed method.
[0064] In addition, the SMO 200 may have the ability to assign active and standby roles to network functions, such as the O-DUs, which enhances system redundancy. This proactive role assignment ensures that backup resources are readily available, facilitating quick recovery in the event of a failure, as one of the advantages of the disclosed method. Additionally, the integration of operator- driven and data-driven policies allows for tailored resiliency strategies, enabling operators to define specific recovery protocols based on real-time conditions and historical data.
[0065] Moreover, the continuous monitoring of network performance and the use of AI/ML for failure prediction significantly enhance the SMO’s responsiveness. By analyzing the FC APS data, the SMO 200 may anticipate issues before they escalate, allowing for preemptive actions that further safeguard network integrity, as one of the advantages of the disclosed method. Furthermore,
the centralized inventory database of hardware and software components streamlines management processes, providing operators with crucial information at their fingertips. This database not only supports effective decision-making during recovery operations but also aids in maintaining up-to- date software releases and configurations, ultimately leading to improved service quality, reduced operational risks, and enhanced user satisfaction.
[0066] FIG. 3 is a flow diagram illustrating a method 300 for managing the end-to-end network resiliency within the 0-RAN architecture, according to an embodiment as disclosed herein. The method 300 may execute multiple operations to manage the end-to-end network resiliency, which is given below.
[0067] At operation 301, the method 300 includes monitoring, at the MnS consumer (e.g., SMO 200), data of the one or more interfaces and the one or more network entities associated with the 0-RAN architecture. In some example embodiments, monitoring the data comprises monitoring at least one of the FCAPS data, the PM counters, Fault Management (FM) data, a notification message, and an alarm. At operation 302, the method 300 includes detecting the failure occurring in at least one of the one or more interfaces or the one or more network entities associated with the O-RAN architecture based on the monitoring. At operation 303, the method 300 includes performing, upon detection of the failure, at least one recovery action. Various examples of the at least one recovery action are explained in the below embodiments.
[0068] In some example embodiments, in response to detecting the failure occurring at the O-RU 206, the SMO 200 may perform the at least one recovery action, such as executing one or more software auto-healing mechanisms. The one or more software auto-healing mechanisms may include performing software upgrades and restarting the O-RU 206, upon availability of a Management Plane (M-plane).
[0069] In some example embodiments, in response to detecting the failure occurring at the O-RU 206, the SMO 200 may perform the at least one recovery action, such as adjusting one or more neighboring O-RUs configuration to provide support for mitigating one or more coverage gaps that occurred due to the failure.
[0070] In some example embodiments, in the context of O-RU failure, this refers to a malfunction in the O-RU 206, a crucial part of the O-RAN infrastructure that handles signal transmission
between the UE 110 and the network component, specifically the O-DU 205. Such failures can impact connected users and their applications. The O-DU 205 can identify O-RU failures through FM data, including the alarms, Configuration Management (CM) notifications (e.g., carrier state change or carrier activation, carrier configuration related notifications), and PM data sent by the O-RU 206 via the open fronthaul M-Plane interface. In hybrid deployments, the SMO 200 may also receive one or more O-RU alarms through this interface, acting as an O-RU controller. The O-DU 205 can either take independent action or relay failure information to the SMO 200 via the 01 interface, allowing the SMO 200 to implement recovery measures for the O-RU 206.
[0071] In some example embodiments, in response to detecting the failure occurring at either the O-DU 205 or the O-CU (e.g., 203 and 204), the SMO 200 may perform the at least one recovery action, such as executing one or more software auto-healing mechanisms when the 01 interface or the E2 interface is operational and one or more active sessions are established in the O-RAN architecture.
[0072] In some example embodiments, in response to detecting the failure occurring at either the O-DU 205 or the O-CU (e.g., 203 and 204), the SMO 200 may perform the at least one recovery action, such as executing, by utilizing at least one Al model, at least one a manual rehoming process, and an automated rehoming process based on the FM and PM data, for predictive autohealing and failure forecasting.
[0073] In some example embodiments, in response to detecting the failure occurring at either the O-DU 205 or the O-CU (e.g., 203 and 204), the SMO 200 may perform the at least one recovery action, such as triggering redundant O-DU (standby O-DU) to take over one or more services for the O-RU 206. In addition, triggering action may utilize a transfer of UE context (if an inter-DU interface is available e.g., D2 interface) and/or cell configuration information for the at least one recovery action.
[0074] In some example embodiments, in response to detecting the failure occurring at either the O-DU 205 or the O-CU (e.g., 203 and 204), the SMO 200 may perform the at least one recovery action, such as executing one or more transport layer-based backup recovery mechanisms for either the O-DU 205 or the O-CU (e.g., 203 and 204).
[0075] In some example embodiments, the SMO 200 may detect the failure occurring at the O- DU 205 may include, but not limited to, a Management Function (MnF) failure, a Network Function (NF) failure, a complete O-DU failure, and any partial failure.
[0076] In the context of the MnF failure, the SMO 200 may identify an O-DU MnF failure by analyzing the PM data from the O-CU (e.g., 203 and 204) through the 01 interface, provided that the Fl link between the O-CU (e.g., 203 and 204) and the O-DU 205 is still operational. The O- CU (e.g., 203 and 204) may configure the O-DU 205 via the Fl interface and monitor its resource status. When both the Fl interface and the 01 interface are functioning, the SMO 200 may collect and compare the PM data from the O-CU (e.g., 203 and 204) with the 01 interface status to the O- DU 205. This process helps determine if there is a failure in the MnF of the O-DU 205 or its connection, especially when traffic is active but cannot reach the O-DU 205 through the 01 interface.
[0077] In the context of the NF failure, if the NF fails but the 01 interface with the O-DU 205 remains operational, the SMO 200 may detect the failure directly through this interface. The SMO 200 may also identify the issue using performance management data from the O-CU (e.g., 203 and 204), as the failure of the Fl interface and the FM data may indicate the NF failure.
[0078] In the context of the complete O-DU failure, where both the MnF and NF have failed, the SMO 200 may analyze the PM data from the O-CU (e.g., 203 and 204) along with the 01 connection status and transport management data received via the 02 interface (if Cloud Native Network Function/Cloudified Network Function/Containerized Network Function (CNF) deployment is in place). The PM data may indicate an Fl failure due to the NF failure, allowing the SMO 200 to conclude that the entire O-DU is inoperable.
[0079] In some example embodiments, the SMO 200 may detect the failure occurring at the O- CU (e.g., 203 and 204) may include, but not limited to, an O-CU-CP failure and an O-CU-UP failure.
[0080] In the context of the O-CU-CP failure, a failure in the O-CU-CP 204 may significantly disrupt cell service by interrupting essential signaling between the UE 110 and the core network. This can result in problems such as failed call setups, handovers, and session management. If the O-CU-CP 204 does not support UE session restoration, it may also trigger large signaling storms.
Additionally, the O-CU-CP 204 may lose connectivity with the SMO 200 through the 01 interface, leading to a loss of configurations and preventing the SMO 200 from receiving 0AM data (CM, FM, and PM) from the O-CU-CP 204. It is crucial to restore gNB, cell services, and user sessions to re-establish functionality, minimize disruptions, and ensure service continuity.
[0081] In the context of the O-CU-UP failure, the O-CU-UP failure in a 5G 0-RAN network can disrupt the transmission of user data between the UE 110 and the core network. This may lead to affected data sessions, reduced service quality, and potential data loss. Furthermore, the O-CU-UP 203 may also be unable to communicate with the SMO 200 via the 01 interface.
[0082] In some example embodiments, in response to detecting the failure occurring at the interface level, the SMO 200 may perform the at least one recovery action, such as executing one or more backup interface-based recovery mechanisms (e.g., multiple SCTP sessions-based recovery) for data transmissions for the plurality of network interfaces (i.e., one or more interfaces). [0083] In some example embodiments, in response to detecting the failure occurring at the interface level, the SMO 200 may perform the at least one recovery action, such as executing at least one redundant transmission mechanism for data transmissions for the plurality of network interfaces. An example of the at least one redundant transmission mechanism may include a Multipath transmission control (MPTCP), which allows data to be transmitted over multiple paths, increasing resilience against path failures.
[0084] In some example embodiments, in response to detecting the failure occurring at the interface level, the SMO 200 may perform the at least one recovery action, such as identifying a cause of the interface failure. The cause is determined to be either a malfunction in the NF or Network Element (NE) or a failure of an Access Point (AP), and recovering the AP to restore functionality and resolve the identified interface failure.
[0085] In some example embodiments, the failure occurring at the interface level may include, but not limited to, a Network Interface Card (NIC) failure, an Internet Protocol (IP) level routing failure, a Stream Control Transmission Protocol (SCTP) connection failure, a control plane interface failure, a user plane interface failure, and a management plane interface failure.
[0086] In some example embodiments, in response to detecting the failure occurring at the transport level, the SMO 200 may perform the at least one recovery action, such as executing one
or more transport-based recovery mechanisms. The one or more transport-based recovery mechanisms are utilized for data transmissions based on distributing traffic and workloads across the plurality of network entities (i.e., one or more network entities) for load balancing
[0087] In some example embodiments, in response to detecting the failure occurring at the O- cloud level, the SMO 200 may perform the at least one recovery action, such as executing one or more NF recovery mechanisms by utilizing distinct standby NF instances or backup clusters for various networks. For instance, implement separate standby instances or backup clusters for both 4G and 5G networks, ensuring that one has priority over the other.
[0088] In some example embodiments, various types of failure(s) can occur at the O-cloud level, which is mentioned below. a. Infrastructure-related failures: These include issues such as O-Cloud failures, site failures, node cluster failures, and other resource failures. b. IP Multimedia Subsystem (IMS)-Related Failures: This category encompasses failures related to IMS software, issues during IMS software updates, failures in the 02-IMS interface, and problems with IMS provisioning procedures. c. Deployment Management Services (DMS)-related failures: Examples here include failures in the DMS control plane, issues during DMS control plane upgrades, deployment plane failures, and failures in NF deployment lifecycle management. d. O-Cloud application-related failures: This includes failures of pods, virtual machines (VMs), and issues during NF deployment.
[0089] In some example embodiments, in response to detecting the failure occurring at the O- cloud level, the SMO 200 may perform the at least one recovery action, such as implementing infrastructure-level resiliency, by employing at least one open-source container orchestration system (e.g., Kubemetes-based solutions), for software auto-healing at the O-cloud level. The at least one open-source container orchestration system may execute one or more auto-healing processes. The one or more auto-healing processes are managed by the O-cloud 207 and one or more NF resiliency features.
[0090] For instance, resiliency solutions can be established through active-active or active-standby configurations at the POD, cluster, or data center level, utilizing the Kubernetes-based solutions
for software auto-healing. Typically, normal auto-healing is managed by cloud and NF resiliency features. However, the SMO 200 may intervene when these mechanisms are insufficient in addressing persistent failures. For instance, if an SCTP connection goes down, the alarm notification may be triggered. The SMO’ s involvement is only necessary if the connection fails to re-establish within a specified timeframe.
[0091] In some example embodiments, in response to detecting the failure occurring at the RIC level, the SMO 200 may perform the at least one recovery action, such as executing one or more recovery actions comprised of performing a software reset, implementing software auto-healing, redeploying, or re-onboarding one or more applications by utilizing backup information provided by the SMO 200.
[0092] For instance, if the non-real-time RIC 201 or the near-real-time RIC 202 fails, or if corresponding applications (near-real-time applications (xApps) or non-real-time applications (rApps)) encounter issues, the SMO 200 may initiate appropriate recovery actions. These actions may include performing a software reset, enabling software auto-healing, redeploying affected applications, or re-onboarding the applications using the backup information that the SMO 200 has stored. This process ensures that services are quickly restored and continue to function effectively.
[0093] In some example embodiments, the SMO 200 may retrieve one or more capabilities of the plurality of network entities associated with the O-RAN architecture. Based on the one or more retrieved capabilities, the SMO 200 may assign a role to each network entity associated with the plurality of network entities, as described in conjunction with FIG. 5.
[0094] In some example embodiments, the SMO 200 may perform at least one action based on one or more additional inputs from one or more cell-network planning tools.
[0095] For instance, the SMO 200 may assign specific operational roles to each network entity comprising an active role for the O-DU 205 and a standby role for the O-DU 205, to enhance network resiliency and ensure seamless service continuity or one or more services to the user with minimum interruption. For another instance, the SMO 200 may recommend a direction to a parent node (e.g., the near-real-time RIC 202 or the O-CU (203 and 204)) regarding the effective
assignment and management of the active role and the standby role associated with the O-DU 205, as described in conjunction with FIG. 6.
[0096] In some example embodiments, the SMO 200 may orchestrate one or more resiliency measures through both operator-driven and data-driven policies. The operator-driven policy establishes recovery configurations and the data-driven policy manages load distribution during one or more overload situations. In addition, the SMO 200 may recommend one or more strategies, to the near-real-time RIC 202 or non-real-time RIC 201, for example, to prevent backup entities from entering an energy-saving mode, to ensure readiness for immediate operational demands.
[0097] In some example embodiments, the SMO 200 may predict by utilizing the at least one Al model, the failure occurring at the at least one network entity based on the received network-related data and one or more parameters, to perform the at least one recovery action in advance.
[0098] Examples of the one or more parameters may include, but are not limited to, a list of impacted network entities among a plurality of network entities, current traffic information associated with the at least one network entity, a Quality of Service (QoS) requirement for each network entity, a Quality of Experience (QoE) requirement for each network entity, applied network traffic information associated with the at least one network entity, and resource requirement information associated with the at least one network entity.
[0099] FIG. 4 is a flow diagram illustrating a method 400 for determining the at least one recovery action, according to an embodiment as disclosed herein. The method 400 may execute multiple operations to determine the at least one recovery action, which is given below.
[0100] At operation 401, the method 400 includes determining a type of failure occurring at the at least one network entity based on the network-related data. At operation 402, the method 400 includes determining an available software-hardware support details at the SMO 200. At operation 403, the method 400 includes determining the at least one recovery action based on the type of the determined failure and the available software hardware support details at the SMO.
[0101] In some example embodiments, the at least one recovery action may include one or more backup mechanisms and one or more redundant node mechanisms (e.g., 1+1 redundancy, N+l redundancy, N+M redundancy).
[0102] In some example embodiments, the one or more backup mechanisms may include a software-based backup mechanism, a hardware-based backup mechanism, and a transport network-based backup mechanism.
[0103] FIG. 5 illustrates an example scenario where the SMO 200 detects the failure that occurs at the O-DU, according to an embodiment as disclosed herein. In the context of a resiliency framework, the SMO 200 performs multiple operations to ensure robust network functionality, which are given below.
[0104] In some example embodiments, initially, the SMO 200 may initialize the roles of various NF instances, incorporating a multi-level mesh architecture that facilitates efficient communication and resource allocation. This architecture enables the identification of capabilities at each level, allowing the SMO 200 to configure the roles for each Multi-Function (MF)/NF appropriately. Continuous dynamic monitoring is then performed by the SMO 200 based on FCAPS feedback, which includes the PM counters. Utilizing this FCAPS feedback along with the operator’s policies, the SMO 200 establishes a relative priority for each NF instance and its supporting counterparts. Upon detecting the failure, the SMO 200 may receive alarms or notifications, prompting it to assess the priority levels previously established. This assessment informs the SMO’s decision regarding which NF instance(s) may be activated to recover from or mitigate the failure. The SMO 200 may employ the 01 interface to communicate the roles to these NF instances.
[0105] For instance, in a potential use case involving Coverage and Capacity Optimization (CCO), the NF instances may increase their transmit power to address coverage gaps resulting from the failure of the affected NF instance, as illustrated in FIG. 5, in response to a failure of “O-DU- 1”. Other use cases include energy saving, mobility load balancing, rolling upgrades, and Multiple- Input Multiple-Output (MIMO) optimization. The configurations can be adjusted to prioritize activities such as upgrades and failure recovery, while optimization tasks remain unaffected.
[0106] To facilitate failure recovery, essential information associated with the O-CU (e.g., 203 and 204), the O-DU 205, and the UE context, is transmitted to the supporting NF instances via the 01 interface. It is crucial to identify the specific SMO entities involved in these processes, such as orchestrators, authentication/security modules, inventory managers, and FCAPS modules.
Additionally, various resiliency scenarios are examined, including intra/inter O-CU, intra/inter SMO, and intra/inter near-RT RIC interactions. Decisions regarding the availability of the O-CU (e.g., 203 and 204)/O-DU 205 at various levels and capacities are influenced by the operator’s policies, triggers, and configuration details based on consumed data like traffic throughput and other Key Performance Indicators (KPIs). Traffic requirements guide the orchestrator in selecting the appropriate standby level based on priority, traffic demands, and the capacity of backup O- DUs/O-CUs.
[0107] In some example embodiments, consider a specific scenario involving the detection of the 0-DU failure, the SMO 200 may execute several critical operations for recovery (recovery action), which are given below. a. First, the SMO 200 retrieves the capabilities of the O-DU 205 and the O-RU 206. b. Subsequently, the SMO 200 assigns roles to the O-DUs, for instance, designating O-DU-1 as Active and O-DU-2 as Standby. c. Upon detecting a failure in O-DU-1 through alarms, notifications, or PM counters, the SMO 200 may configure the associated O-CU (e.g., 203 and 204) to deactivate the cells previously served by O-DU-1. d. The SMO 200 then considers multiple recovery actions: i. Option-1 : The SMO 200 may perform a software reset or upgrade for O- DU-1 to address the failure or degradation. ii. Option-2: The SMO 200 may initiate service rehoming, which can be either manual or automated, transferring services from O-DU-1 to O-DU-2. iii. Option-3: The SMO 200 may trigger or configure O-DU-2 to take over the O-RU, thus restoring service.
[0108] In the cases of Options -2 and -3, O-DU-2 may configure the O-RU 206 to restore service. If the O-RU 206 is preconfigured with backup carriers, O-DU-2 may activate these carriers to facilitate service restoration. Concurrently, the SMO 200 may attempt to recover O-DU-1 through appropriate measures such as software resets, upgrades, or hardware issues.
[0109] In the above-mentioned exemplary scenario(s), implementation of the resiliency framework within the SMO 200 may offer several significant advantages. Firstly, the initialization
of roles for various NF instances and the integration of the multi-level mesh architecture enhances network efficiency and resource allocation, ensuring optimal performance under varying conditions. Continuous dynamic monitoring based on the FCAPS feedback allows for real-time adjustments, enabling proactive management of network resources and minimizing downtime. [0110] Additionally, the prioritization of the NF instances facilitates rapid response to failures, allowing the SMO 200 to quickly identify and activate the necessary backup instances for recovery. This responsiveness not only reduces service interruptions but also enhances overall network reliability. The flexibility to adjust configurations for activities such as upgrades and failure recovery ensures that critical operations can be prioritized without compromising optimization efforts. Moreover, the ability to utilize various recovery options, such as the service rehoming and the software upgrades, provides a robust mechanism for maintaining service continuity and ensuring a seamless user experience.
[0111] FIG. 6 illustrates an exemplary backup scenario associated with the O-DU 205, according to an embodiment as disclosed herein. In the context of backup scenarios for the O-DUs, focusing on resiliency deployments. In these resiliency deployments, both the active and standby O-DUs can be connected to the same O-CU (Intra O-CU) or different O-CUs (Inter O-CU). The standby configurations may include several designs, such as 1+1, N (Active) + K (Standby), N+l (All Active + 1 Hot Spare), and 1+K (1 Active + K Hot/Cold Spare).
[0112] In some example embodiments, resiliency can be achieved through either the active-active or the active-standby configuration. a. In the active-active configuration, the M-Plane remains active alongside supervision monitoring, ensuring that the standby node or the NF is always available to recover from resiliency events or take over service operations. b. Conversely, the active-standby configuration involves an inactive M-Plane, although the IP address of the O-DU (e.g., standby- 1,. . ., standby-N) is known to the O-RU 206, facilitating potential recovery.
[0113] In some example embodiments, the capabilities of the standby node or the NF (O-DU) (e.g., standby- 1,..., standby-N) may operate at the same or different levels compared to the active O- DU 205, as determined by the orchestrator/controller (e.g., SMO 200).
a. Key metrics for quantifying these capabilities include carrier resource management, throughput capacity, the number of UEs, and the number of cells/sectors. b. Additionally, Al and ML recommendations may inform backup strategies based on alarm types, severity, notifications, and PM counters.
[0114] In some example embodiments, decision-making for recovery is primarily handled by the SMO 200 (the non-RT RIC 201), the near-RT RIC 202, or the O-CU (e.g., 203 and 204), which determines which O-DU (e.g., standby- 1,..., standby-N) may recover the O-RU 206 based on traffic requirements. Factors to consider include the potential failure of the standby O-DU (e.g., standby- 1,..., standby-N), the availability of ports, and the status of the M-Plane. Various scenarios dictate the recovery process from resiliency or disaster-related failures, with the orchestrator triggering the appropriate O-DU (e.g., standby- 1 ,... , standby-N) that possesses the desired or available capabilities, which may differ from those of the active O-DU 205. Finally, subscriber data synchronization ensures consistency across RAN NF instances, such as the O-DU and the O-CU (standby).
[0115] FIGS. 7A-7B illustrates an advanced resiliency use case (partial O-DU failure), according to an embodiment as disclosed herein. In the advanced resiliency scenario concerning a partial failure of the O-DU, several essential preconditions may be established. First, 01 supervision may be operational between the SMO 200 and the O-DU-1 205a. Second, 01 supervision may also be active between the SMO 200 and the O-DU-2 205b. Lastly, the 01 supervision connection between the SMO 200 and the O-CU-CP 204 is up and running.
[0116] At operation 701, the SMO 200 detects a decline in service quality due to the failure of the O-DU-1 205a. At operation 702, recognizing this issue, the SMO 200 decides in the second operation to shift services to the O-DU-2 205b to maintain service continuity.
[0117] In context of deactivation of affected cells or for each cell that requires deactivation, at operation 703, the SMO 200 initiates by locking the affected cells and removing any resources associated with the O-DU-1 205a through the 01 interface. This step ensures that the resources linked to the failed unit are no longer in use. Following this, the O-CU-CP 204 performs one or more actions to deactivate the cells and remove resources via the Fl interface, which connects the O-CU-CP 204 to the O-DUs. At operation 704, the O-DU-1 205a sends a notification to the SMO
200 regarding the change in cell status, informing it that the cells are no longer active. This communication occurs through the 01 interface, ensuring the SM 200 is updated on the system’s status.
[0118] At operation 705, the SMO 200 configures a Transport Network Layer (TNL) for the O- DU-2 205b. This configuration is crucial for ensuring that data can flow correctly through the network. The O-CU-CP 204 then performs necessary tasks to establish the Fl interface for the O- DU-2 205b, enabling communication between the two units. At operation 706, the SMO 200 collaborates with the O-CU-CP 204 to create and configure the new cell using the 01 interface. This operation is vital for preparing the O-DU-2205b to take over the services previously managed by the O-DU-1 205a.
[0119] At operation 707, the SMO 200 configures the cell and its associated carriers for the O- DU-2 205b. If the Fl interface is successfully established, the O-CU-CP 204 manages the messaging needed for the Fl connection, ensuring that the O-DU-2 205b can report back to the O- CU-CP 204. At operation 708, the O-CU-CP 204 informs the SMO 200 that the cell has been activated through the 01 interface.
[0120] At operation 709, finally, the O-DU-2 205b sends a message to the SMO 200 confirming that the cells and their corresponding carriers are now active, indicating that the O-DU-2 205b is ready to provide services. At operation 710, the SMO 200 receives subscriptions for performance counters from the O-DU-2 205b. This allows the SMO 200 to monitor the performance of the newly active unit (e.g., O-DU-2205b), ensuring that the O-DU-2205b meets operational standards and can effectively manage the services it has taken over.
[0121] FIGS. 8A, 8B and 8C illustrates an 0-DU resiliency use case, according to an embodiment as disclosed herein. In the 0-DU resiliency use case, several critical preconditions may be established to ensure smooth operations. A synchronized timing setup is essential across one or more network entities, several essential preconditions may be established. First, 01 supervision is active between the MnS consumer (e.g., SMO 200) and the O-DU-1 205a. This is followed by 01 supervision between the SMO 200 and the O-DU-2 205b, and finally, 01 supervision is operational between the SMO 200 and the O-CU-CP 204. The process begins with the initial
system setup for the O-DU resiliency use case, which involves a series of interconnected operations, which are given below.
[0122] The SMO 200, which operates as a non-RT RIC, configures the ODUs with designated roles active and standby as well as managing the configuration and activation of the cells.
[0123] At operation 801, in a cloud deployment scenario, a standby O-DU may not be pre-existing. Instead, a new O-DU instance is dynamically created whenever a failure occurs with the currently active O-DU. During operation 801a, the SMO 200 configures the O-DU-1 205a to take on the active role through the 01 interface. Next, in operation 801b, the SMO 200 manages the configuration of the [tr]x-array-carriers for the 0-RU 206 via the OFH interface. Subsequently, in operation 801c, the SMO 200 configures O-DU-2 205b to assume the standby role, again using the 01 interface.
[0124] At operation 802, the O-CU-CP 204 performs a series of actions to activate the cells and carriers associated with the O-DU-1 205a and the 0-RU 206. Following this, at operation 803, notifications are sent out to confirm that the cells and corresponding carriers have been activated. By operation 804, the 0-RU 206 becomes fully operational and is connected to the service, with both the M-plane and CU plane functioning correctly. At operation 805, the SMO 200 collects alarms and performance metrics from the network entities (Network Functions (NFs)/ MnS producer), including O-CU-CP 204, O-DU-1 205a, and 0-RU 206, to monitor the system’s health. [0125] In the event of a failure within the O-DU resiliency use case, particularly if the O-DU-1 205a experiences a complete or partial failure, the following operations occur.
[0126] At operation 806, the O-DU-1 205a fails (either completely or partially) and enters a disabled operational state. In operation 807, the SMO 200 detects this failure based on alarms and performance counters received from the network entities. By operation 808, if the O-DU-1 205a is completely failed and unreachable, there is no need for cell deactivation or removal, as it is already deemed non-operational. At operation 809, the O-CU-CP 204 may initiate one or more operations to deactivate and optionally remove the cells and associated carriers from the O-DU-1 205a. At operation 810, notifications regarding the deactivation and removal status of these cells and carriers are communicated, if applicable.
[0127] To recover O-RU 206 operations after the 0-DU failure, the following operations are executed.
[0128] At operation 811, the SMO 200 configures O-DU-2 205b to take on the active role via the 01 interface. At operation 812, the standby O-DU-2 transitions to become the active 0-DU. During operation 813, the configuration of [tr]x-array-carriers between O-DU-2 205b and the O- RU 206 is established to ensure proper connectivity.
[0129] At operation 814, the O-CU-CP 204 performs a series of operations to activate the cells and associated carriers linked with O-DU-2 205b. At operation 815, the SMO 200 monitors notifications regarding the activation status of cells and carriers from the network entities to ensure everything is functioning as expected. At operation 816, the O-RU 206 is operational once again, fully connected and in service, with both the M-plane and CU plane running smoothly.
[0130] At operation 817, the SMO 200 continues to monitor the system by detecting alarms and performance metrics from O-DU-2205b. Finally, at operation 818, the transition of the active role from O-DU-2 205b back to the O-DU-1 205a involves repeating the same operations as the initial transition, based on triggers from the SMO 200 or decisions made by an operator 800. This interconnected series of operations ensures that the network remains resilient and responsive to failures.
[0131] In some example embodiments, it is assumed that the 0-DU 205, as the MnS producer, possesses the capability to autonomously detect 01 interface failures and initiate a reset in the absence of management plane functionality resulting from the 01 interface failure. This assumption is made to mitigate uncertainties arising from the management system’s inability to assess the impact of services on the 0-DU 205 without an available management interface, as well as to diagnose and reconfigure the 0-DU 205 using management plane mechanisms.
[0132] In some example embodiments, the disclosed method provides advanced security measures to protect against cyber threats, ensuring network resilience against attacks. Further, the disclosed method may implement security measures to detect and prevent cyber-attacks that could cause the NF failures. Ensuring secure communication channels between the NFs to prevent data breaches and maintain integrity.
[0133] In some example embodiments, the disclosed method may impact various entities, which are given below. a. In one embodiment, the 01 interface may handle disruptions gracefully, e.g., robust error handling, retry mechanisms, and support for redundant communication paths to ensure that 0AM commands and data may be exchanged even in the presence of network issues. The 01 interface may support seamless failover to maintain continuous communication with network elements. This requires robust session management and state synchronization. b. In one embodiment, the 0AM architecture may be configured to support distributed and redundant 0AM components, enabling load balancing, and ensuring that critical management functions may continue to operate even if some components fail. The 0AM architecture may be defined for performance degradation detection and mitigation and disaster recovery, including the roles and interactions of different components. The 0AM architecture may also support automated recovery processes and dynamic reconfiguration to adapt to changing network conditions. c. In one embodiment, an 01 Network Resource Management (NRM) may include models that support redundancy, failover mechanisms, and self-healing capabilities. The 01 NRM may be configured to extend the network resource model to include attributes related to performance degradation detection and mitigation and disaster recovery. The NRM may include attributes and relationships that facilitate the detection and management of faults, as well as the reallocation of resources to maintain service continuity. d. In one embodiment, an 01 PM may include performance measurements related to performance degradation detection and mitigation, such as degradation detection time and mitigation effectiveness and recovery time and data loss in case of disaster recovery. Further, the 01 PM may include implementing mechanisms for data integrity, data buffering, redundant data collection paths, and ensuring that performance measurement systems can continue to operate and provide accurate data even during partial network failures.
e. In one embodiment, Traffic Engineering (TE) and Inventory (IV) may include inventory management-related impacts to store the resiliency backup. The TE and IV may facilitate the resiliency orchestrator to assign roles for the active and standby systems. E.g., a list of NF Instances or systems that are available to recover from the resiliency are known to the resiliency orchestrator or SMO services.
[0134] The above-mentioned one or more embodiments provide several advantages. For instance, first, the disclosed method enhances redundancy by ensuring multiple pathways and backup systems are in place, facilitating continuous operation even during failures. The disclosed method is designed with fault tolerance, allowing it to maintain functionality despite component failures. Additionally, the disclosed method supports scalability, enabling the network to efficiently manage increased loads or changes in demand without sacrificing performance. Robust monitoring mechanisms provide real-time surveillance and alerting, ensuring swift detection of issues. The incorporation of automated recovery processes streamlines failover actions, minimizing downtime and enhancing high availability. This results in improved user experience, as seamless service delivery is maintained even during network disruptions, ensuring that users remain connected and satisfied.
[0135] FIGS. 9A-9B illustrates 0-DU resiliency scenario in case of 01 failure, according to an embodiment as disclosed herein. The 0-DU resiliency scenario outlines the expected system behavior for basic resiliency of the 0-DU in situations where management over the 0-DU as the MnS producer is compromised due to the termination of the 01 interface. This indicates either a complete failure of the Network Function or the Management Plane (01), with the root cause analysis of the 01 failure being outside the scope of this disclosure. The scenario emphasizes enhancing the resilience of the 01 interface through robust monitoring, failover mechanisms, and redundancy.
[0136] In this scenario, the roles of various components are defined. For instance, the 0AM as the MnS consumer, the MnF of the O-CU-CP 204 as the MnS producer, and the MnFs of O-DU-1 205a (active) and O-DU-2 205b (standby), both serving as MnS producers.
[0137] Preconditions for this scenario involve the SMO 200 continuously monitoring the 0-DU- 1 205a, ensuring the 01 interface between the SMO 200 and the O-CU-CP 204 is operational, and
having the O-DU-2 205b configured in a standby role or not configured at all (refer to note-1, notes related to this scenario added in the below mentioned table).
[0138] The scenario begins when the SMO 200 identifies its inability to manage the O-DU-1 205a. Initially, at operation 901, the SMO 200 detects an 01 interface failure or a complete failure of O- DU-1 205a through Fault Management (FM) and/or Performance Management (PM) data (refer to note-2 and note-3).
[0139] At operation 902, the SMO 200 then analyzes notifications, faults, and PM counters from the O-CU-CP 204 to determine connectivity between the O-CU-CP 204 and the O-DU via the Fl interface, as well as whether the O-DU has communicated the removal of cell(s) served by the O- DU with a lost 01 interface. This analysis informs the SMO 200’ s decision to transition services from the O-DU-1 205a to the O-DU-2 205b upon detecting failures (refer to note-4).
[0140] Subsequently, at operation 903, based on the findings (902), the SMO 200 deactivates the cell(s) associated with the O-DU-1 205a on the O-CU-CP 204 via the 01 interface. At operation 904, the SMO 200 then configures the O-DU-2 205b for connection with the O-CU-CP 204, providing necessary details such as transport parameters and performance management configurations. The O-DU communicates the results of these operations to the SMO 200 through the 01 interface (refer to note-5).
[0141] At operation 905, the SMO 200 continues by creating and configuring the cell (s) associated with the O-DU-2 205b on the O-CU-CP 204, ensuring the O-CU-CP 204 receives the required configurations including transport details, cell(s) configuration, configuration needed to connect with the O-DU-2 205b, configuration for performance management and so on. The O-CU-CP 204 informs the results of the above operations to the SMO 200 through the 01 interface
[0142] At operation 906, if necessary, the SMO 200 creates cell object instances and manages the setup of cell(s) and corresponding carrier resources in the O-DU-2 205b, linking them to the respective cell(s) managed by the O-CU-CP 204. The O-DU-2205b then communicates the results of these operations back to the SMO 200 through the 01 interface (refer to note-6 and note-7).
[0143] At operation 907, the SMO 200 sends an activation request for the cell(s) associated with the O-DU-2205b configured successfully in operations 905 and 906 to O-CU-CP 204. The O-CU- CP 204 confirms the reception of activation request from SMO 200 through the 01 interface (note-
8). At operation 908, for the cell(s) activation request in operation 907, O-CU-CP 204 notifies the SMO 200 on their activation status via the 01 interface. At operation 909, the O-DU-2 205b informs the SMO 200 of the activation status of the cell(s) and corresponding carrier(s), indicating readiness to provide services.
[0144] At operation 910, the SMO 200 subscribes to and collects alarms and performance counters from the O-DU-2 205b through the 01 interface. The scenario concludes when the O-DU-2 205b is ready to accept UEs and deliver services.
[0145] This scenario may include some postconditions, the initiation of functions by the standby 0-DU, acceptance of connections from the O-RU to establish an OFH interface session, and the system’s readiness to accept UEs and provide services, effectively replacing the failed O-DU with the standby 0-DU in the service path.
Table 1
[0146] FIGS. 10A, 10B and IOC illustrates enhanced network resiliency through a standby O-DU and adaptive cell activation managed by the SMO 200, according to an embodiment as disclosed herein. To ensure that the O-DU instance maintains the same level of functionality as a failed or degraded instance, the system can be quickly reconfigured to utilize a standby O-DU. The SMO 200 regularly monitors the 01 interface status, and upon detecting a complete and irrecoverable link failure, the SMO 200 automatically triggers a failover to the standby O-DU. The SMO 200’ s potential solution involves configuring the standby O-DU to replace the failed one. This scenario impacts service, as illustrated in FIGS. 10A, 10B and 10C.
[0147] In this scenario, the roles of various components are defined. For instance, the O-RU 206 manages the fronthaul and air interfaces, along with reporting performance metrics and failures. All O-DUs (e.g., active/standby O-DUs) connected to the O-RU 206 are involved in the O-DU resiliency scenario. The O-DU receives cell configurations and carrier settings from the SMO 200 and works with the O-CU-CP 204 to activate these configurations.
[0148] In addition, the O-CU receives cell-related configurations from the SMO 200 and details about available cells from the O-DUs. Based on the resources reported by the O-DUs and the desired cell availability from the SMO 200, the O-CU manages the activation and deactivation of cells in the O-DUs using the Fl interface. Moreover, the SMO 200 plays a critical role by
designating O-DUs as either active or standby and communicating these roles to the O-RU 206 within a hierarchical deployment. It also makes high-level decisions based on prevailing conditions. [0149] For this scenario, it is assumed that when an active O-DU fails, at least one standby 0-DU may be available, with or without an active Netconf session with the O-RU 206, to restore services to end users. This scenario is initiated by various triggers, it addresses complete or partial O-DU failures, performance degradation, and other related issues, with the primary goal of preserving the operational integrity of the O-RU 206 and O-DU system despite disruptions, such as failures in the 01 interface. The use case/scenario is activated when the primary O-DU experiences any of these conditions, causing its operational state to change to “disabled”.
[0150] The preconditions for this scenario may include: a. The 01 interface between the SMO 200 and O-CU-CP 204 is operational. b. The O-DU-1 205a is configured as active, with the 01 interface to the SMO 200 functioning. c. The O-DU-2 205b is either configured as standby or not configured at all; if available, the 01 interface between the SMO 200 and the O-DU-2 205b is operational. d. The SMO 200 has valid subscriptions for alarms, performance metrics, and notifications for the O-DU-1 205a and the O-DU-2 205b, provided they are connected via the 01 interface.
[0151] At operation 1001, the O-DU-1 205a failure detection and recovery process begins with the SMO 200 may monitor the O-DU-1 205a through the 01 interface to identify service degradation by analyzing performance counters, alarms, and notifications. At operation 1002, the SMO 200 may detect either the 01 interface failure or a complete failure of the O-DU-1 205a using the FM and PM data (refer note-1 and note-2 in below mentioned table).
[0152] At operation 1003, in the case of partial failures, the SMO 200 may analyze notifications, faults, and PM counters from the O-DU-1 205a to identify performance issues. For complete failures or 01 interface failures, the SMO 200 may analyze information from the O-CU-CP 204 to verify connectivity via the Fl interface and whether O-DU-1 205a has communicated any lost
cells for removal. These findings guide the SMO 200’ s decision to switch services from the O- DU-1 205a to the O-DU-2 205b (refer note-3 in below mentioned table).
[0153] At operation 1004a, subsequently, the SMO 200 may request the O-DU-1 205a to lock affected cells and remove corresponding carrier resources through the 01 interface, receiving updates on the status of these operations (refer note-4 and note-5 in below mentioned table). At operation 1004b, the O-DU-1 205a may then deactivate the specified carriers according to Clause 15.3 of the WG4-MP specification and may optionally remove these resources ([tr]x-array-carriers resources). At operation 1004c, the O-DU-1 205a may notify the SMO 200 about the status change for the locked cells, indicated in 1004a (refer note-6 in below mentioned table). At operation 1005, based on the results from the previous analysis, the SMO 200 may request the O-CU-CP 204 to terminate the Fl interface with the failed O-DU-1 205a (refer note-7 in below mentioned table).
[0154] To recover 0-RU 206 operations, at operation 1006, the SMO 200 may configure the O- DU-2 205b to connect with the O-CU-CP 204 and receives updates on the operation’s status via the 01 interface (refer note-8 in below mentioned table). At operation 1007, the O-DU-2 205b then becomes the active unit to restore O-RU 206 operations and associated user services (refer note-9 in below mentioned table). At operation 1008, the SMO 200 may create and configure the cells linked to the O-DU-2 205b on the O-CU-CP 204 using the 01 interface and receives status updates from the O-CU-CP 204 (results of the above-mentioned operations). At operation 1009, the SMO 200 may configure the O-DU-2 205b with the cells set up on the O-CU-CP 204 in the previous operations (1005), mapping carrier resources to cell resources, and the O-DU-2 205b reports the results of these configurations back to the SMO 200 through the 01 interface (refer note- 10, note-11, and note- 12 in below mentioned table).
[0155] At operation 1010, the O-DU-2 205b may configure the [tr]x-array-carriers and activates on O-RU 206 as described in Clause 15.3 of WG4-MP specification. The O-RU 206 may inform the results of the above operations to the 0-DU through the OFH M-Plane interface. At operation 1011, for the cell(s) successfully activated in result of 1005, the O-CU-CP 204 notifies the SMO 200 on their activation status via the 01 interface. At operation 1012, the O-DU-2 205b notifies the SMO 200 on the cell(s) and corresponding carrier(s) activation status via the 01 interface. The O-DU-2205b becomes ready to offer services now. (refer note- 13 and note- 14 in below mentioned
table). At operation 1013, the SMO 200 makes subscriptions to performance counters and notifications if the O-DU-2 205b was not configured as stated in Pre-condition. SMO collects the performance counters and notifications from the O-DU-2205b via the 01 interface^ refer note-15 in below mentioned table). The use case concludes when O-DU-2 205b successfully takes over the functions of the previously active O-DU-1 205a.
[0156] In terms of exceptions, potential issues may arise if the 0-RU 206 fails during the switchover, if the standby O-DU-2 205b becomes unavailable when it is due to activate, or if there is a loss of event messaging during the flow. Other exceptions include improper configuration, loss of configuration data from the standby 0-DU, the return of the active 0-DU to service while the standby is activating, or unavailability connectivity, or functionality of the SMO 200.
[0157] Regarding post-conditions of this scenario may include, a. a successful outcome occurs when the 0-RU 206 connects to the newly active O- DU-2 205b, which has successfully initiated its functions and synchronized with the O-RU 206, making the system ready to serve users and replacing the failed O- DU-1 205a. b. Conversely, if exceptions occur, failure post-conditions may arise, such as no O- DUs being available, leading to the 0-RU 206 potentially shutting down operations. Misconfigurations may prompt appropriate responses from the 0-RU 206, and if two active O-DUs exist simultaneously, the 0-RU 206 may operate accordingly.
Table 2
[0158] FIG. 11 illustrates a diagram of example components of an apparatus 1100, according to an embodiment as disclosed herein. As shown in FIG. 11, the apparatus 1100 comprises a processor 1110, a memory 1120, a storage component 1130, an input component 1140, an output component 1150, a communication interface 1160, and a bus 1170. In one embodiment, the apparatus 1100 may relate to at least one of , for example, the SMO 200, the O-RU 206, the O-DU 205, the O-CU (e.g., 203 and 204), an interface entity, a transport entity, the O-cloud 207, the RIC (e.g., 201 and 202), and the O-eNB 208.
[0159] The processor 1110, as used herein, means any type of computational circuit that may comprise hardware elements and software elements. The processor 1110 may be embodied as a multi-core processor, a single core processor, or a combination of one or more multi-core processors and/or one or more single core processors, a distributed processing system, or the like. The processor 1110 may be a Central Processing Unit (CPU), a Graphics Processing Unit (GPU), an Accelerated Processing Unit (APU), an Application-Specific Integrated Circuit (ASIC), or another type of processing component.
[0160] The memory 1120 includes a non-transitory computer readable medium. Memory 1120 includes a Random-Access Memory (RAM), a Read Only Memory (ROM), and/or another type of dynamic or static storage device (e.g., a flash memory, a magnetic memory, and/or an optical memory) that stores information and/or instructions for use by processor 1110. The memory 1120 comprises machine-readable instructions which are executable by the processor 1 110. These
machine-readable instructions when executed by the processor 1110 cause the processor 1110 to perform one or more method steps of an embodiment described above.
[0161] The storage component 1130 stores information and/or software related to the operation and use of the apparatus 1100. For example, the storage component 1130 may include a hard disk (e.g., a magnetic disk, an optical disk, a magneto-optic disk, and/or a solid-state disk), a Compact Disc (CD), a Digital Versatile Disc (DVD), a floppy disk, a cartridge, a magnetic tape, and/or another type of non-transitory computer-readable medium, along with a corresponding drive.
[0162] The input component 1140 is configured to receive information, such as user input. For example, the input component 1140 may include, but not be limited to, a touch screen display, a keyboard, a keypad, a mouse, a button, a switch, and/or a microphone. Additionally, or alternatively, the input component 1140 may include a sensor for sensing information (e.g., a Global Positioning System (GPS), an accelerometer, a gyroscope, and/or an actuator).
[0163] The output component 1150 is configured to provide output information from the apparatus 1100. For example, the output component 1150 may be, but is not limited to, a display, a speaker, instructions to an external device, and/or one or more Light-Emitting Diodes (LEDs).
[0164] The communication interface 1160 is an interface that provides a communication connection to other devices, such as external devices and internal devices. The connection by the communication interface 1160 can be a wired connection, a wireless connection, or a combination of wired and wireless connections, and can be a direct connection or an indirect connection via a communication network that exists between the apparatus 1100 and other devices. In other words, the standard of the communication interface 1160 is not limited.
[0165] The bus 1170 acts as an interconnect between the processor 1110, the memory 1120, the storage component 1130, the input component 1140, the output component 1150, and the communication interface 1160 of the apparatus 1 100. The bus 1170 may include a wired interconnection or a wireless interconnection.
[0166] The number and arrangement of components shown in FIG. 11 are provided as an example. In practice, the apparatus 1100 may include additional components, fewer components, different components, or differently arranged components than those shown in FIG. 11. Additionally, or alternatively, a set of components (e.g., one or more components) of the apparatus 1100 may
perform one or more functions described as being performed by another set of components of the apparatus 1100. Further, one or more method steps described in any of the embodiments may be performed utilizing the apparatus 1100 in communication with one another.
[0167] Examples of the techniques and apparatus described herein include, but are not limited to, the following enumerated embodiments:
[1] A method comprising: monitoring, at a MnS consumer, data of one or more interfaces or one or more network entities associated with an Open Radio Access Network (O-RAN) architecture, wherein monitoring the data comprises monitoring at least one of Fault, Configuration, Accounting, Performance, Security (FCAPS) data, Performance Measurement (PM) counters, Fault Management (FM) data, a notification message, and an alarm; detecting a failure occurring in at least one of the one or more interfaces or the one or more network entities associated with the O-RAN architecture based on the monitoring; and performing, upon detection of the failure, at least one recovery action.
[2] The method as described in [1]: wherein the at least one recovery action comprises one or more backup mechanisms and one or more redundant node mechanisms; and wherein the one or more backup mechanisms comprise a software-based backup mechanism, a hardware-based backup mechanism, and a transport network-based backup mechanism.
[3] The method as described in any of [l]-[2]: wherein the one or more network entities comprises at least one of a non-real-time RAN Intelligent Controller (RIC), a near-real-time RIC, an O-RAN Central Unit User Plane (O-CU-UP), an O-RAN Central Unit Control Plane (O-CU-CP), an O-RAN
Distributed Unit (O-DU), an O-RAN Radio Unit (O-RU), an O-RAN cloud (O-cloud), and an O-RAN enhanced Node B (O-eNB); and wherein the one or more interfaces comprises at least one of an Al interface, an 01 interface, an 02 interface, an E2 interface, an R1 interface, a Y1 interface, an Fl-C interface, an Fl-U interface, an open Fronthaul for Control and User plane (FH-CUS), and an Open Fronthaul for Management (FH-M) plane.
[4] The method as described in any of [l]-[3], wherein performing the at least one recovery action to mitigate one or more impacts of the detected failure comprises: in response to detecting the failure occurring at an O-RAN Radio Unit (O-RU) within the O-RAN architecture, performing, at the O-RU, the at least one recovery action comprises: executing one or more software auto-healing mechanisms, wherein the one or more software auto-healing mechanisms comprise performing software upgrades and restarting the O-RU, upon an availability of a Management Plane (M-plane); or adjusting one or more neighboring O-Rus configuration to provide support for mitigating one or more coverage gaps that occurred due to the failure.
[5] The method as described in any of [l]-[4], wherein performing the at least one recovery action to mitigate one or more impacts of the detected failure comprises: in response to detecting the failure occurring at either an O-RAN Distributed Unit (0-DU) or an O-RAN Central Unit (0-CU) within the O-RAN architecture, performing, at the corresponding 0-DU or 0-CU, the at least one recovery action comprises: executing one or more software auto-healing mechanisms when an 01 interface or an E2 interface is operational and one or more active sessions are established in the O-RAN architecture; executing, by utilizing at least one Artificial Intelligence (Al) model, at least one a manual rehoming process, and an automated rehoming process based on Fault
Management (FM) and Performance Management (PM) data, for predictive autohealing and failure forecasting; triggering a redundant O-DU to take over one or more services for the O- RU; or executing one or more transport layer-based backup recovery mechanisms for either the O-DU or the O-CU (O-CU-CP and O-CU-UP), wherein detecting the failure occurring at the O-DU comprises at least one of a Management Function (MnF) failure, a Network Function (NF) failure, a complete O-DU failure, and any partial failure; and wherein detecting the failure occurring at the O-CU comprises at least one of a O-CU-CP failure and a O-CU-UP failure.
[6] The method as described in any of [l]-[5], wherein performing the at least one recovery action to mitigate one or more impacts of the detected failure comprises: in response to detecting the failure occurring at an interface level, performing the at least one recovery action comprises: executing one or more backup interface-based recovery mechanisms for data transmissions for a plurality of network interfaces, wherein the plurality of network interfaces comprises at least one of an Al interface, an 01 interface, an 02 interface, an E2 interface, an R1 interface, a Y1 interface, an Fl-C interface, an Fl-U interface, an open Fronthaul for Control and User plane (FH-CUS), and an Open Fronthaul for Management (FH-M) plane ; executing at least one redundant transmission mechanism for data transmissions for the plurality of network interfaces; or identifying a cause of the interface failure, wherein the cause is determined to be either a malfunction in a Network Function (NF) or Network Element (NE) or a failure of an Access Point (AP), and recovering the AP to restore functionality and resolve the identified interface failure,
wherein detecting the failure occurring at the interface level comprises at least one of a Network Interface Card (NIC) failure, an IP level routing failure, a Stream Control Transmission Protocol (SCTP) connection failure, a control plane interface failure, a user plane interface failure, and a Management plane interface failure.
[7] The method as described in any of [ l]-[6], wherein performing the at least one recovery action to mitigate one or more impacts of the detected failure comprises: in response to detecting the failure occurring at a transport level, performing the at least one recovery action comprises: executing one or more transport-based recovery mechanisms, for data transmissions, by distributing traffic and workloads across a plurality of network entities for load balancing.
[8] The method as described in any of [l]-[7], wherein performing the at least one recovery action to mitigate one or more impacts of the detected failure comprises: in response to detecting the failure occurring at an O-RAN cloud level, performing the at least one recovery action comprises: executing one or more Network Function (NF) recovery mechanisms by utilizing distinct standby NF instances or backup clusters for various networks; implementing infrastructure-level resiliency, by employing at least one open-source container orchestration system, for software auto-healing at the O- RAN cloud level; or executing one or more auto-healing processes, wherein the one or more auto-healing processes are managed by an O-RAN cloud and one or more NF resiliency features.
[9] The method as described in any of [l]-[8], wherein performing the at least one recovery action to mitigate one or more impacts of the detected failure comprises:
in response to detecting the failure occurring at a RAN Intelligent Controller (RIC) level, performing the at least one recovery action comprises: executing one or more recovery actions comprises performing a software reset, implementing software auto-healing, redeploying, or re-onboarding one or more applications by utilizing backup information provided by the MnS consumer.
[10] The method as described in any of [l]-[9], comprising: retrieving one or more capabilities of a plurality of network entities associated with the O-RAN architecture; and assigning a role to each network entity associated with the plurality of network entities based on the one or more retrieved capabilities.
[11] The method as described in any of [l]-[10], the method comprising: performing, based on one or more additional inputs from one or more cell-network planning tools, at least one action comprises: assigning specific operational roles to each network entity comprising an active role for the O-DU and a standby role for the O-DU; recommending a direction to a parent node regarding the effective assignment and management of the active role and the standby role associated with the O-DU.
[12] The method as described in any of [1]-[11], the method comprising at least one of: orchestrating one or more resiliency measures through both operator-driven and data-driven policies, wherein the operator-driven policy establishes recovery configurations, and the data-driven policy manages load distribution during one or more overload situations; and recommending one or more strategies, to a near-real-time RIC or non-real-time RIC, to prevent backup entities from entering an energy-saving mode, to ensure readiness for immediate operational demands.
[13] The method as described in any of [1]-[12], comprising: predicting, by utilizing at least one Artificial Intelligence (Al) model, the failure occurring at the at least one network entity based on the received network-related data and one or more parameters, to perform the at least one recovery action in advance, wherein the one or more parameters comprise a list of impacted network entities among a plurality of network entities, current traffic information associated with the at least one network entity, a Quality of Service (QoS) requirement for each network entity, a Quality of Experience (QoE) requirement for each network entity, applied network traffic information associated with the at least one network entity, and resource requirement information associated with the at least one network entity.
[14] The method as described in any of [1]-[13], prior to performing the at least one recovery action comprising: determining a type of the failure occurring at the at least one network entity based on the network-related data; determining an available software-hardware support details at the MnS consumer; and determining the at least one recovery action based on the type of the determined failure and the available software-hardware support details at the MnS consumer.
[15] An apparatus configured to: periodically monitor, at a MnS consumer, data of one or more interfaces or one or more network entities associated with an Open Radio Access Network (O-RAN) architecture, wherein monitoring the data comprises monitoring at least one of Fault, Configuration, Accounting, Performance, Security (FCAPS) data, Performance
Measurement (PM) counters, Fault Management (FM) data, a notification message, and an alarm; detect a failure occurring in at least one of the one or more interfaces or the one or more network entities associated with the O-RAN architecture based on the monitoring; and perform, upon detection of the failure, at least one recovery action.
[16] The apparatus as described in [15], wherein at least one recovery action comprises one or more backup mechanisms and one or more redundant node mechanisms; and wherein the one or more backup mechanisms comprise a software-based backup mechanism, a hardware-based backup mechanism, and a transport network-based backup mechanism.
[17] The apparatus as described in any of [15]-[16], wherein the one or more network entities comprises at least one of a non-real-time RAN Intelligent Controller (RIC), a near-real-time RIC, an O-RAN Central Unit User Plane (O-CU-UP), an O-RAN Central Unit Control Plane (O-CU-CP), an O-RAN Distributed Unit (O-DU), an O-RAN Radio Unit (O-RU), an O-RAN cloud (O-cloud), and an O-RAN enhanced Node B (O-eNB); and wherein the one or more interfaces comprises at least one of an Al interface, an 01 interface, an 02 interface, an E2 interface, an R1 interface, a Y1 interface, an Fl-C interface, an Fl-U interface, an open Fronthaul for Control and User plane (FH-CUS), and an Open Fronthaul for Management (FH-M) plane.
[18] The apparatus as described in any of [15]-[17], wherein to perform the at least one recovery action to mitigate the one or more impacts of the detected failure, the apparatus is configured to:
in response to detecting the failure occurring at an O-RAN Radio Unit (O-RU) within the O-RAN architecture, perform, at the O-RU, the at least one recovery action comprises: execute one or more software auto-healing mechanisms, wherein the one or more software auto-healing mechanisms comprise performing software upgrades and restarting the O-RU, upon an availability of a Management Plane (M-plane); or adjust one or more neighboring O-Rus configuration to provide support for mitigating one or more coverage gaps that occurred due to the failure.
[19] The apparatus as described in any of [15]-[18], wherein to perform the at least one recovery action to mitigate the one or more impacts of the detected failure, the apparatus is configured to: in response to detecting the failure occurring at either an O-RAN Distributed Unit (O-DU) or an O-RAN Central Unit (O-CU) within the O-RAN architecture, perform, at the corresponding O-DU or O-CU, the at least one recovery action comprises: execute one or more software auto-healing mechanisms when an 01 interface or an E2 interface is operational and one or more active sessions are established in the O-RAN architecture; execute, by utilizing at least one Artificial Intelligence (Al) model, at least one a manual rehoming process, and an automated rehoming process based on Fault Management (FM) and Performance Management (PM) data, for predictive autohealing and failure forecasting; trigger a redundant O-DU to take over one or more services for the O-RU; or execute one or more transport layer-based backup recovery mechanisms for either the O-DU or the O-CU (O-CU-CP and 0-CU-UP),
wherein detecting the failure occurring at the O-DU comprises at least one of a Management Function (MnF) failure, a Network Function (NF) failure, a complete O-DU failure, and any partial failure; and wherein detecting the failure occurring at the O-CU comprises at least one of a O-CU-CP failure and a O-CU-UP failure.
[20] The apparatus as described in any of [15]-[19], wherein to perform the at least one recovery action to mitigate the one or more impacts of the detected failure, the apparatus is configured to: in response to detecting the failure occurring at an interface level, perform the at least one recovery action comprises: execute one or more backup interface-based recovery mechanisms for data transmissions for a plurality of network interfaces, wherein the plurality of network interfaces comprises at least one of an Al interface, an 01 interface, an 02 interface, an E2 interface, an R1 interface, a Y1 interface, an Fl-C interface, an Fl-U interface, an open Fronthaul for Control and User plane (FH-CUS), and an Open Fronthaul for Management (FH-M) plane; execute at least one redundant transmission mechanism for data transmissions for the plurality of network interfaces; or identify a cause of the interface failure, wherein the cause is determined to be either a malfunction in a Network Function (NF) or Network Element (NE) or a failure of an Access Point (AP), and recovering the AP to restore functionality and resolve the identified interface failure, wherein detecting the failure occurring at the interface level comprises at least one of a Network Interface Card (NIC) failure, an IP level routing failure, a Stream Control Transmission Protocol (SCTP) connection failure, a control plane interface failure, a user plane interface failure, and a Management plane interface failure.
[21] The apparatus as described in any of [15]-[20], wherein to perform the at least one recovery action to mitigate the one or more impacts of the detected failure, the apparatus is configured to: in response to detecting the failure occurring at a transport level, perform the at least one recovery action comprises: execute one or more transport-based recovery mechanisms, for data transmissions, by distributing traffic and workloads across a plurality of network entities for load balancing.
[22] The apparatus as described in any of [15]-[21], wherein to perform the at least one recovery action to mitigate the one or more impacts of the detected failure, the apparatus is configured to: in response to detecting the failure occurring at an O-RAN cloud level, perform the at least one recovery action comprises: execute one or more Network Function (NF) recovery mechanisms by utilizing distinct standby NF instances or backup clusters for various networks; implement infrastructure-level resiliency, by employing at least one open- source container orchestration system, for software auto-healing at the O-RAN cloud level; or execute one or more auto-healing processes, wherein the one or more autohealing processes are managed by an O-RAN cloud and one or more NF resiliency features.
[23] The apparatus as described in any of [15]-[22], wherein to perform the at least one recovery action to mitigate the one or more impacts of the detected failure, the apparatus is configured to: in response to detecting the failure occurring at a RAN Intelligent Controller (RIC) level, perform the at least one recovery action comprises:
execute one or more recovery actions comprising performing a software reset, implementing software auto-healing, redeploying, or re-onboarding one or more applications by utilizing backup information provided by the MnS consumer.
[24] The apparatus as described in any of [ 15]-[23 ], wherein the apparatus is configured to: retrieve one or more capabilities of a plurality of network entities associated with the O-RAN architecture; and assign a role to each network entity associated with the plurality of network entities based on the one or more retrieved capabilities.
[25] The apparatus as described in any of [15]-[24], the apparatus is configured to: perform, based on one or more additional inputs from one or more cell-network planning tools, at least one action comprises: assign specific operational roles to each network entity comprising an active role for the O-DU and a standby role for the O-DU; recommend a direction to a parent node regarding the effective assignment and management of the active role and the standby role associated with the O-DU.
[26] The apparatus as described in any of [15]-[25], the apparatus is configured to, at least one of: orchestrate one or more resiliency measures through both operator-driven and data- driven policies, wherein the operator-driven policy establishes recovery configurations, and the data-driven policy manages load distribution during one or more overload situations; and recommending one or more strategies, to a near-real-time RIC or non-real-time RIC, to prevent backup entities from entering an energy-saving mode, to ensure readiness for immediate operational demands.
[27] The apparatus as described in any of [15]-[26], the apparatus is configured to:
predict, by utilizing at least one Artificial Intelligence (Al) model, the failure occurring at the at least one network entity based on the received network-related data and one or more parameters, to perform the at least one recovery action in advance, wherein the one or more parameters comprise a list of impacted network entities among a plurality of network entities, current traffic information associated with the at least one network entity, a Quality of Service (QoS) requirement for each network entity, a Quality of Experience (QoE) requirement for each network entity, applied network traffic information associated with the at least one network entity, and resource requirement information associated with the at least one network entity.
[28] The apparatus as described in any of [ 15]-[27], prior to perform the at least one recovery action, the apparatus is configured to: determine a type of the failure occurring at the at least one network entity based on the network-related data; determine an available software-hardware support details at the MnS consumer; and determine the at least one recovery action based on the type of the determined failure and the available software-hardware support details at the MnS consumer.
[29] A non-transitory computer-readable medium storing instructions, the instructions comprising: one or more instructions that, when executed by an apparatus, the apparatus comprising one or more processors, cause the one or more processors to: periodically monitor, at a MnS consumer, data of one or more interfaces or one or more network entities associated with an Open Radio Access Network (O- RAN) architecture, wherein monitoring the data comprises monitoring at least one of Fault, Configuration, Accounting, Performance, Security (FCAPS) data,
Performance Measurement (PM) counters, Fault Management (FM) data, a notification message, and an alarm; detect a failure occurring in at least one of the one or more interfaces or the one or more network entities associated with the O-RAN architecture based on the monitoring; and perform, upon detection of the failure, at least one recovery action.
[0168] The embodiments disclosed herein can be implemented through at least one software program running on at least one hardware device and performing network management functions to control the elements. The elements can be at least one of a hardware device or a combination of hardware devices and software modules.
[0169] While specific language has been used to describe the disclosure, any limitations arising on account of the same are not intended. As would be apparent to a person in the art, various working modifications may be made to the method in order to implement the inventive concept as taught herein.
[0170] The drawings and the forgoing description give examples of embodiments. Those skilled in the art will appreciate that one or more of the described elements may well be combined into a single functional element. Alternatively, certain elements may be split into multiple functional elements. Elements from one embodiment may be added to another embodiment. For example, orders of processes described herein may be changed and are not limited to the manner described herein.
[0171] Moreover, the actions of any flow diagram need not be implemented in the order shown; nor do all of the acts necessarily need to be performed. Also, those acts that are not dependent on other acts may be performed in parallel with the other acts. The scope of embodiments is by no means limited by these specific examples. Numerous variations, whether explicitly given in the specification or not, such as differences in structure, dimension, and use of material, are possible. The scope of embodiments is at least as broad as given by the following claims.
[0172] Benefits, other advantages, and solutions to problems have been described above with regard to specific embodiments. However, the benefits, advantages, solutions to problems, and any component(s) that may cause any benefit, advantage, or solution to occur or become more
pronounced are not to be construed as a critical, required, or essential feature or component of any or all the claims.
[0173] The foregoing description of the specific embodiments will so fully reveal the general nature of the embodiments herein that others can, by applying current knowledge, readily modify and/or adapt for various applications such specific embodiments without departing from the generic concept, and, therefore, such adaptations and modifications should and are intended to be comprehended within the meaning and range of equivalents of the disclosed embodiments. It is to be understood that the phraseology or terminology employed herein is for the purpose of description and not of limitation. Therefore, while the embodiments herein have been described in terms of at least one embodiment, those skilled in the art will recognize that the embodiments herein can be practiced with modification within the spirit and scope of the embodiments as described herein.
Claims
1. A method comprising: monitoring, at a Management and Orchestration Service (MnS) consumer, data of one or more interfaces or one or more network entities associated with an Open Radio Access Network (O-RAN) architecture, wherein monitoring the data comprises monitoring at least one of Fault, Configuration, Accounting, Performance, Security (FCAPS) data, Performance Measurement (PM) counters, Fault Management (FM) data, a notification message, and an alarm; detecting a failure occurring in at least one of the one or more interfaces or the one or more network entities associated with the O-RAN architecture based on the monitoring; and performing, upon detection of the failure, at least one recovery action.
2. The method as claimed in claim 1 : wherein the at least one recovery action comprises one or more backup mechanisms and one or more redundant node mechanisms; and wherein the one or more backup mechanisms comprise a software-based backup mechanism, a hardware-based backup mechanism, and a transport network-based backup mechanism.
3. The method as claimed in claim 1 : wherein the one or more network entities comprises at least one of a non-real-time RAN Intelligent Controller (RIC), a near-real-time RIC, an O-RAN Central Unit User Plane (O-CU-UP), an O-RAN Central Unit Control Plane (O-CU-CP), an O-RAN Distributed Unit (O-DU), an O-RAN Radio Unit (O-RU), an O-RAN cloud (O-cloud), and an O-RAN enhanced Node B (O-eNB); and wherein the one or more interfaces comprises at least one of an Al interface, an 01 interface, an 02 interface, an E2 interface, an R1 interface, a Y1 interface, an Fl-C
interface, an Fl-U interface, an open Fronthaul for Control and User plane (FH-CUS), and an Open Fronthaul for Management (FH-M) plane.
4. The method as claimed in claim 1, wherein performing the at least one recovery action to mitigate one or more impacts of the detected failure comprises: in response to detecting the failure occurring at an O-RAN Radio Unit (O-RU) within the O-RAN architecture, performing, at the O-RU, the at least one recovery action comprises: executing one or more software auto-healing mechanisms, wherein the one or more software auto-healing mechanisms comprise performing software upgrades and restarting the O-RU, upon an availability of a Management Plane (M-plane); or adjusting one or more neighboring O-RUs configuration to provide support for mitigating one or more coverage gaps that occurred due to the failure.
5. The method as claimed in claim 1, wherein performing the at least one recovery action to mitigate one or more impacts of the detected failure comprises: in response to detecting the failure occurring at either an O-RAN Distributed Unit (O-DU) or an O-RAN Central Unit (O-CU) within the O-RAN architecture, performing, at the corresponding O-DU or O-CU, the at least one recovery action comprises: executing one or more software auto-healing mechanisms when an 01 interface or an E2 interface is operational and one or more active sessions are established in the O-RAN architecture; executing, by utilizing at least one Artificial Intelligence (Al) model, at least one a manual rehoming process, and an automated rehoming process based on Fault Management (FM) and Performance Management (PM) data, for predictive autohealing and failure forecasting; triggering a redundant O-DU to take over one or more services for an O- RAN Radio Unit (O-RU); or
executing one or more transport layer-based backup recovery mechanisms for either the O-DU or the O-CU (O-CU-CP and O-CU-UP), wherein detecting the failure occurring at the O-DU comprises at least one of a Management Function (MnF) failure, a Network Function (NF) failure, a complete O-DU failure, and any partial failure ; and wherein detecting the failure occurring at the O-CU comprises at least one of a O-CU-CP failure and a O-CU-UP failure.
6. The method as claimed in claim 1, wherein performing the at least one recovery action to mitigate one or more impacts of the detected failure comprises: in response to detecting the failure occurring at an interface level, performing the at least one recovery action comprises: executing one or more backup interface-based recovery mechanisms for data transmissions for a plurality of network interfaces, wherein the plurality of network interfaces comprises at least one of an Al interface, an 01 interface, an 02 interface, an E2 interface, an R1 interface, a Y1 interface, an Fl-C interface, an Fl-U interface, an open Fronthaul for Control and User plane (FH-CUS), and an Open Fronthaul for Management (FH-M) plane ; executing at least one redundant transmission mechanism for data transmissions for the plurality of network interfaces; or identifying a cause of the interface failure, wherein the cause is determined to be either a malfunction in a Network Function (NF) or Network Element (NE) or a failure of an Access Point (AP), and recovering the AP to restore functionality and resolve the identified interface failure, wherein detecting the failure occurring at the interface level comprises at least one of a Network Interface Card (NIC) failure, an IP level routing failure, a Stream Control Transmission Protocol (SCTP) connection
failure, a control plane interface failure, a user plane interface failure, and a Management plane interface failure.
7. The method as claimed in claim 1, wherein performing the at least one recovery action to mitigate one or more impacts of the detected failure comprises: in response to detecting the failure occurring at a transport level, performing the at least one recovery action comprises: executing one or more transport-based recovery mechanisms, for data transmissions, by distributing traffic and workloads across a plurality of network entities for load balancing.
8. The method as claimed in claim 1, wherein performing the at least one recovery action to mitigate one or more impacts of the detected failure comprises: in response to detecting the failure occurring at an O-RAN cloud level, performing the at least one recovery action comprises: executing one or more Network Function (NF) recovery mechanisms by utilizing distinct standby NF instances or backup clusters for various networks; implementing infrastructure-level resiliency, by employing at least one open-source container orchestration system, for software auto-healing at the O- RAN cloud level; or executing one or more auto-healing processes, wherein the one or more auto-healing processes are managed by an O-RAN cloud and one or more NF resiliency features.
9. The method as claimed in claim 1, wherein performing the at least one recovery action to mitigate one or more impacts of the detected failure comprises: in response to detecting the failure occurring at a RAN Intelligent Controller (RIC) level, performing the at least one recovery action comprises:
executing one or more recovery actions comprises performing a software reset, implementing software auto-healing, redeploying, or re-onboarding one or more applications by utilizing backup information provided by the MnS consumer.
10. The method as claimed in claim 1, comprising: retrieving one or more capabilities of a plurality of network entities associated with the O-RAN architecture; and assigning a role to each network entity associated with the plurality of network entities based on the one or more retrieved capabilities.
11. The method as claimed in claim 1, the method comprising: performing, based on one or more additional inputs from one or more cell-network planning tools, at least one action comprises: assigning specific operational roles to each network entity comprising an active role for an O-RAN Distributed Unit (O-DU) and a standby role for the O- DU; recommending a direction to a parent node regarding the effective assignment and management of the active role and the standby role associated with the O-DU.
12. The method as claimed in claim 1, the method comprising at least one of: orchestrating one or more resiliency measures through both operator-driven and data-driven policies, wherein the operator-driven policy establishes recovery configurations, and the data-driven policy manages load distribution during one or more overload situations; and recommending one or more strategies, to a near-real-time RIC or non-real-time RIC, to prevent backup entities from entering an energy-saving mode, to ensure readiness for immediate operational demands.
13. The method as claimed in claim 1, comprising: predicting, by utilizing at least one Artificial Intelligence (Al) model, the failure occurring at the at least one network entity based on the received network-related data and one or more parameters, to perform the at least one recovery action in advance, wherein the one or more parameters comprise a list of impacted network entities among a plurality of network entities, current traffic information associated with the at least one network entity, a Quality of Service (QoS) requirement for each network entity, a Quality of Experience (QoE) requirement for each network entity, applied network traffic information associated with the at least one network entity, and resource requirement information associated with the at least one network entity.
14. The method as claimed in claim 1, prior to performing the at least one recovery action comprising: determining a type of the failure occurring at the at least one network entity based on the network-related data; determining an available software- hardware support details at the MnS consumer; and determining the at least one recovery action based on the type of the determined failure and the available software- hardware support details at the MnS consumer.
15. An apparatus configured to: periodically monitor, at a MnS consumer, data of one or more interfaces or one or more network entities associated with an Open Radio Access Network (O-RAN) architecture, wherein monitoring the data comprises monitoring at least one of Fault, Configuration, Accounting, Performance, Security (FCAPS) data, Performance Measurement (PM) counters, Fault Management (FM) data, a notification message, and an alarm;
detect a failure occurring in at least one of the one or more interfaces or the one or more network entities associated with the O-RAN architecture based on the monitoring; and perform, upon detection of the failure, at least one recovery action.
16. The apparatus as claimed in claim 15, wherein at least one recovery action comprises one or more backup mechanisms and one or more redundant node mechanisms; and wherein the one or more backup mechanisms comprise a software-based backup mechanism, a hardware-based backup mechanism, and a transport network-based backup mechanism.
17. The apparatus as claimed in claim 15: wherein the one or more network entities comprises at least one of a non-real-time RAN Intelligent Controller (RIC), a near-real-time RIC, an O-RAN Central Unit User Plane (O-CU-UP), an O-RAN Central Unit Control Plane (O-CU-CP), an O-RAN Distributed Unit (O-DU), an O-RAN Radio Unit (O-RU), an O-RAN cloud (O-cloud), and an O-RAN enhanced Node B (O-eNB); and wherein the one or more interfaces comprises at least one of an Al interface, an 01 interface, an 02 interface, an E2 interface, an R1 interface, a Y1 interface, an Fl-C interface, an Fl-U interface, an open Fronthaul for Control and User plane (FH-CUS), and an Open Fronthaul for Management (FH-M) plane.
18. The apparatus as claimed in claim 15, wherein to perform the at least one recovery action to mitigate the one or more impacts of the detected failure, the apparatus is configured to: in response to detecting the failure occurring at an O-RAN Radio Unit (O-RU) within the O-RAN architecture, perform, at the O-RU, the at least one recovery action comprises:
execute one or more software auto-healing mechanisms, wherein the one or more software auto-healing mechanisms comprise performing software upgrades and restarting the O-RU, upon an availability of a Management Plane (M-plane); or adjust one or more neighboring O-RUs configuration to provide support for mitigating one or more coverage gaps that occurred due to the failure.
19. The apparatus as claimed in claim 15, wherein to perform the at least one recovery action to mitigate the one or more impacts of the detected failure, the apparatus is configured to: in response to detecting the failure occurring at either an O-RAN Distributed Unit (O-DU) or an O-RAN Central Unit (O-CU) within the O-RAN architecture, perform, at the corresponding O-DU or O-CU, the at least one recovery action comprises: execute one or more software auto-healing mechanisms when an 01 interface or an E2 interface is operational and one or more active sessions are established in the O-RAN architecture; execute, by utilizing at least one Artificial Intelligence (Al) model, at least one a manual rehoming process, and an automated rehoming process based on Fault Management (FM) and Performance Management (PM) data, for predictive autohealing and failure forecasting; trigger a redundant O-DU to take over one or more services for an O-RAN Radio Unit (O-RU); or execute one or more transport layer-based backup recovery mechanisms for either the O-DU or the O-CU (O-CU-CP and 0-CU-UP), wherein detecting the failure occurring at the O-DU comprises at least one of a Management Function (MnF) failure, a Network Function (NF) failure, a complete O-DU failure, and any partial failure; and wherein detecting the failure occurring at the O-CU comprises at least one of a O-CU-CP failure and a 0-CU-UP failure.
20. The apparatus as claimed in claim 15, wherein to perform the at least one recovery action to mitigate the one or more impacts of the detected failure, the apparatus is configured to: in response to detecting the failure occurring at an interface level, perform the at least one recovery action comprises: execute one or more backup interface-based recovery mechanisms for data transmissions for a plurality of network interfaces, wherein the plurality of network interfaces comprises at least one of an Al interface, an 01 interface, an 02 interface, an E2 interface, an R1 interface, a Y1 interface, an Fl-C interface, an Fl-U interface, an open Fronthaul for Control and User plane (FH-CUS), and an Open Fronthaul for Management (FH-M) plane; execute at least one redundant transmission mechanism for data transmissions for the plurality of network interfaces; or identify a cause of the interface failure, wherein the cause is determined to be either a malfunction in a Network Function (NF) or Network Element (NE) or a failure of an Access Point (AP), and recovering the AP to restore functionality and resolve the identified interface failure, wherein detecting the failure occurring at the interface level comprises at least one of a Network Interface Card (NIC) failure, an IP level routing failure, a Stream Control Transmission Protocol (SCTP) connection failure, a control plane interface failure, a user plane interface failure, and a Management plane interface failure.
21. The apparatus as claimed in claim 15, wherein to perform the at least one recovery action to mitigate the one or more impacts of the detected failure, the apparatus is configured to: in response to detecting the failure occurring at a transport level, perform the at least one recovery action comprises: execute one or more transport-based recovery mechanisms, for data transmissions, by distributing traffic and workloads across a plurality of network entities for load balancing.
22. The apparatus as claimed in claim 15, wherein to perform the at least one recovery action to mitigate the one or more impacts of the detected failure, the apparatus is configured to: in response to detecting the failure occurring at an O-RAN cloud level, perform the at least one recovery action comprises: execute one or more Network Function (NF) recovery mechanisms by utilizing distinct standby NF instances or backup clusters for various networks; implement infrastructure-level resiliency, by employing at least one open- source container orchestration system, for software auto-healing at the O-RAN cloud level; or execute one or more auto-healing processes, wherein the one or more autohealing processes are managed by an O-RAN cloud and one or more NF resiliency features.
23. The apparatus as claimed in claim 15, wherein to perform the at least one recovery action to mitigate the one or more impacts of the detected failure, the apparatus is configured to: in response to detecting the failure occurring at a RAN Intelligent Controller (RIC) level, perform the at least one recovery action comprises: execute one or more recovery actions comprising performing a software reset, implementing software auto-healing, redeploying, or re-onboarding one or more applications by utilizing backup information provided by the MnS consumer.
24. The apparatus as claimed in claim 15, wherein the apparatus is configured to: retrieve one or more capabilities of a plurality of network entities associated with the O-RAN architecture; and assign a role to each network entity associated with the plurality of network entities based on the one or more retrieved capabilities.
25. The apparatus as claimed in claim 15, the apparatus is configured to:
perform, based on one or more additional inputs from one or more cell-network planning tools, at least one action comprises: assign specific operational roles to each network entity comprising an active role for an O-RAN Distributed Unit (O-DU) and a standby role for the O-DU; recommend a direction to a parent node regarding the effective assignment and management of the active role and the standby role associated with the O-DU.
26. The apparatus as claimed in claim 15, the apparatus is configured to, at least one of: orchestrate one or more resiliency measures through both operator-driven and data- driven policies, wherein the operator-driven policy establishes recovery configurations, and the data-driven policy manages load distribution during one or more overload situations; and recommending one or more strategies, to a near-real-time RIC or non-real-time RIC, to prevent backup entities from entering an energy-saving mode, to ensure readiness for immediate operational demands.
27. The apparatus as claimed in claim 15, the apparatus is configured to: predict, by utilizing at least one Artificial Intelligence (Al) model, the failure occurring at the at least one network entity based on the received network-related data and one or more parameters, to perform the at least one recovery action in advance, wherein the one or more parameters comprise a list of impacted network entities among a plurality of network entities, current traffic information associated with the at least one network entity, a Quality of Service (QoS) requirement for each network entity, a Quality of Experience (QoE) requirement for each network entity, applied network traffic information associated with the at least one network entity, and resource requirement information associated with the at least one network entity.
28. The apparatus as claimed in claim 15, prior to perform the at least one recovery action, the apparatus is configured to: determine a type of the failure occurring at the at least one network entity based on the network-related data; determine an available software-hardware support details at the MnS consumer; and determine the at least one recovery action based on the type of the determined failure and the available software-hardware support details at the MnS consumer.
29. A non-transitory computer-readable medium storing instructions, the instructions comprising: one or more instructions that, when executed by an apparatus, the apparatus comprising one or more processors, cause the one or more processors to: periodically monitor, at a MnS consumer, data of one or more interfaces or one or more network entities associated with an Open Radio Access Network (O- RAN) architecture, wherein monitoring the data comprises monitoring at least one of Fault, Configuration, Accounting, Performance, Security (FCAPS) data, Performance Measurement (PM) counters, Fault Management (FM) data, a notification message, and an alarm; detect a failure occurring in at least one of the one or more interfaces or the one or more network entities associated with the O-RAN architecture based on the monitoring; and perform, upon detection of the failure, at least one recovery action.
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| IN202411058170 | 2024-07-31 | ||
| IN202411058170 | 2025-05-12 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2026029933A1 true WO2026029933A1 (en) | 2026-02-05 |
Family
ID=98607640
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/US2025/036896 Pending WO2026029933A1 (en) | 2024-07-31 | 2025-07-09 | End-to-end network resiliency framework |
Country Status (1)
| Country | Link |
|---|---|
| WO (1) | WO2026029933A1 (en) |
-
2025
- 2025-07-09 WO PCT/US2025/036896 patent/WO2026029933A1/en active Pending
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| US11212181B2 (en) | Cloud zone network analytics platform | |
| US11811632B2 (en) | Systems and methods for high availability and performance preservation for groups of network functions | |
| EP3824599B1 (en) | Fault detection methods | |
| US10965558B2 (en) | Method and system for effective data collection, aggregation, and analysis in distributed heterogeneous communication network | |
| US20100306572A1 (en) | Apparatus and method to facilitate high availability in secure network transport | |
| Santos et al. | SELFNET Framework self‐healing capabilities for 5G mobile networks | |
| CN107251485A (en) | The service quality of the raising of cellular radio access networks | |
| US12413513B2 (en) | Predictive routing using risk and longevity metrics | |
| US20220345394A1 (en) | Progressive automation with predictive application network analytics | |
| US12615204B2 (en) | Switching control of communication route | |
| US20260095390A1 (en) | System and method for validating software upgrades and optimizing network path mapping | |
| US20150142961A1 (en) | Network element in network management system,network management system, and network management method | |
| Gupta et al. | Novel Approaches in Network Fault Management. | |
| WO2026029933A1 (en) | End-to-end network resiliency framework | |
| US20250267513A1 (en) | Traffic management to reduce cellular network outages | |
| US20250373486A1 (en) | Cluster failure management system and techniques for telecommunications systems | |
| US20240430702A1 (en) | Open radio access network maintenance applications | |
| US20240430712A1 (en) | Open radio access network maintenance applications | |
| CN107872822A (en) | The bearing method and bogey of a kind of business | |
| CN121264014A (en) | NF device and signaling control method executed in NF | |
| US20250267073A1 (en) | Method and device for handling abnormal network behavior in a wireless communication system | |
| US12231958B2 (en) | Adaptable resiliency for a virtualized RAN framework | |
| US20260081819A1 (en) | Systems and methods to detect and mitigate fluctuation at n3 network interface within open radio access network telecommunications networks | |
| US20250392927A1 (en) | Dynamically updating service priorities for network functions within a 5g network | |
| US12621685B2 (en) | Open radio access network maintenance applications |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 25848792 Country of ref document: EP Kind code of ref document: A1 |