WO2025035961A1 - 服务器可维护性配置方法、装置、电子设备和存储介质 - Google Patents
服务器可维护性配置方法、装置、电子设备和存储介质 Download PDFInfo
- Publication number
- WO2025035961A1 WO2025035961A1 PCT/CN2024/100740 CN2024100740W WO2025035961A1 WO 2025035961 A1 WO2025035961 A1 WO 2025035961A1 CN 2024100740 W CN2024100740 W CN 2024100740W WO 2025035961 A1 WO2025035961 A1 WO 2025035961A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- mode
- server
- utilization rate
- service
- response
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Pending
Links
Classifications
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F9/00—Arrangements for program control, e.g. control units
- G06F9/06—Arrangements for program control, e.g. control units using stored programs, i.e. using an internal store of processing equipment to receive or retain programs
- G06F9/46—Multiprogramming arrangements
- G06F9/50—Allocation of resources, e.g. of the central processing unit [CPU]
- G06F9/5083—Techniques for rebalancing the load in a distributed system
- G06F9/5088—Techniques for rebalancing the load in a distributed system involving task migration
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F11/00—Error detection; Error correction; Monitoring
- G06F11/07—Responding to the occurrence of a fault, e.g. fault tolerance
- G06F11/0703—Error or fault processing not based on redundancy, i.e. by taking additional measures to deal with the error or fault not making use of redundancy in operation, in hardware, or in data representation
- G06F11/0793—Remedial or corrective actions
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F11/00—Error detection; Error correction; Monitoring
- G06F11/07—Responding to the occurrence of a fault, e.g. fault tolerance
- G06F11/0703—Error or fault processing not based on redundancy, i.e. by taking additional measures to deal with the error or fault not making use of redundancy in operation, in hardware, or in data representation
- G06F11/079—Root cause analysis, i.e. error or fault diagnosis
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F9/00—Arrangements for program control, e.g. control units
- G06F9/06—Arrangements for program control, e.g. control units using stored programs, i.e. using an internal store of processing equipment to receive or retain programs
- G06F9/46—Multiprogramming arrangements
- G06F9/48—Program initiating; Program switching, e.g. by interrupt
- G06F9/4806—Task transfer initiation or dispatching
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F9/00—Arrangements for program control, e.g. control units
- G06F9/06—Arrangements for program control, e.g. control units using stored programs, i.e. using an internal store of processing equipment to receive or retain programs
- G06F9/46—Multiprogramming arrangements
- G06F9/50—Allocation of resources, e.g. of the central processing unit [CPU]
- G06F9/5005—Allocation of resources, e.g. of the central processing unit [CPU] to service a request
- G06F9/5027—Allocation of resources, e.g. of the central processing unit [CPU] to service a request the resource being a machine, e.g. CPUs, Servers, Terminals
-
- Y—GENERAL TAGGING OF NEW TECHNOLOGICAL DEVELOPMENTS; GENERAL TAGGING OF CROSS-SECTIONAL TECHNOLOGIES SPANNING OVER SEVERAL SECTIONS OF THE IPC; TECHNICAL SUBJECTS COVERED BY FORMER USPC CROSS-REFERENCE ART COLLECTIONS [XRACs] AND DIGESTS
- Y02—TECHNOLOGIES OR APPLICATIONS FOR MITIGATION OR ADAPTATION AGAINST CLIMATE CHANGE
- Y02D—CLIMATE CHANGE MITIGATION TECHNOLOGIES IN INFORMATION AND COMMUNICATION TECHNOLOGIES [ICT], I.E. INFORMATION AND COMMUNICATION TECHNOLOGIES AIMING AT THE REDUCTION OF THEIR OWN ENERGY USE
- Y02D10/00—Energy efficient computing, e.g. low power processors, power management or thermal management
Definitions
- the present application relates to the field of computer systems and storage technologies, and in particular to a server maintainability configuration method, a server maintainability configuration device, an electronic device, and a non-volatile readable storage medium.
- Server failures can be divided into two categories: downtime failures and non-downtime failures.
- Downtime failures are mainly manifested in downtime during the startup process and downtime during operation.
- the server has a certain repair function for faulty parts, and even if the server may have a hardware failure, necessary means can be used to keep it running normally. This process is the RAS (Reliability Availability Serviceability) function of the server.
- RAS repair configures a parameter by default to continue running until the operation and maintenance personnel troubleshoot the problem, which will affect the performance and operating efficiency of the server and may further cause server downtime.
- some embodiments of the present application are proposed to provide a server maintainability configuration method, a server maintainability configuration device, an electronic device and a non-volatile readable storage medium that overcome the above problems or at least partially solve the above problems.
- some embodiments of the present application disclose a server maintainability configuration method, including:
- the step of calculating the first utilization rate of the central processing unit includes:
- a first utilization rate is determined according to the power consumption data and the unit heat data.
- the step of determining the first utilization rate according to the power consumption data and the unit heat data includes:
- the first ratio is determined as a first utilization rate.
- the step of determining a faulty component in response to a server shutdown and restart, includes:
- the step of determining the faulty component before the step of reading the error information, in response to the server crashing and restarting, the step of determining the faulty component further includes:
- the step of calculating the second utilization rate of the central processing unit includes:
- a second utilization rate is determined according to the power consumption data and the unit heat data.
- the step of determining the first utilization rate according to the power consumption data and the unit heat data includes:
- a second ratio is determined as a second utilization rate.
- the step of determining the service migration state based on the first utilization rate and the second utilization rate includes:
- the service migration status is determined based on the service fluctuation value.
- the step of calculating the service fluctuation value based on the first utilization rate and the second utilization rate includes:
- the third ratio is determined as the business fluctuation value.
- the step of determining the service migration state based on the service fluctuation value includes:
- the server configuration mode includes a reliability mode and an operability mode, the operation reliability of the reliability mode is greater than the operation reliability of the operability mode, and the operation efficiency of the operability mode is greater than the operation efficiency of the reliability mode; according to the service migration state, the step of switching the server configuration mode includes:
- the server configuration mode is switched to an operability mode.
- the step of switching the server configuration mode to the reliability mode includes:
- the step of switching the server configuration mode to the operability mode includes:
- the server configuration mode also includes a balanced mode and an automatic mode.
- the operating efficiency of the balanced mode is between the reliability mode and the operability mode, and the operating reliability of the balanced mode is between the reliability mode and the operability mode; the automatic mode multiplexes one of the operability mode, the reliability mode and the balanced mode.
- the method further comprises:
- a mode selection page is displayed.
- the method further comprises:
- a selection operation on a mode selection page is received, and one of a reliability mode, an operability mode, a balance mode, and an automatic mode is selected as a current configuration mode.
- the preset service fluctuation threshold is 30%.
- an embodiment of the present application discloses a server maintainability configuration device, including:
- a first calculation module configured to calculate a first utilization rate of a central processing unit in response to normal startup and operation of the server
- a restart module used to respond to a server shutdown and restart, and determine the failed component
- a second calculation module used for calculating a second utilization rate of the central processing unit
- a service migration judgment module used to determine the service migration status based on the first utilization rate and the second utilization rate
- a switching module is used to switch the server configuration mode according to the business migration status
- the isolation module is used to isolate a faulty component in the server configuration mode.
- an embodiment of the present application discloses an electronic device, including a processor, a memory, and a computer program stored in the memory and capable of running on the processor, and when the computer program is executed by the processor, the steps of the server maintainability configuration method as above are implemented.
- an embodiment of the present application discloses a non-volatile readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the above server maintainability configuration method are implemented.
- the embodiment of the present application calculates the first utilization rate of the central processing unit in response to the normal startup and operation of the server; determines the faulty component in response to the shutdown and restart of the server; calculates the second utilization rate of the central processing unit; determines the business migration status based on the first utilization rate and the second utilization rate; switches the server configuration mode according to the business migration status; and isolates the faulty component in the server configuration mode.
- FIG1 is a flow chart of steps of an embodiment of a server maintainability configuration method of the present application
- FIG2 is a flowchart of another server maintainability configuration method embodiment of the present application.
- FIG3 is a schematic diagram of a framework of an example of a server maintainability configuration method of the present application.
- FIG4 is a structural block diagram of an embodiment of a server maintainability configuration device of the present application.
- FIG5 is a structural block diagram of an electronic device provided in an embodiment of the present application.
- FIG. 6 is a structural block diagram of a storage medium provided in an embodiment of the present application.
- Step 101 in response to the normal startup and operation of the server, calculating a first utilization rate of the central processing unit
- the CPU utilization rate of the central processing unit at this time is calculated, which is the first utilization rate.
- Step 102 in response to the server crash restart, determine the faulty component
- the server can determine the faulty component based on reading the logs.
- PCIE peripheral component interconnect express
- UCE Unit Cell Error
- IERR Internal Error
- Step 103 calculating a second utilization rate of the central processing unit
- the CPU utilization rate of the central processing unit at this time can be calculated, which is the second utilization rate.
- Step 104 determining a service migration state based on the first utilization rate and the second utilization rate
- the service migration state is determined.
- Step 105 switching the server configuration mode according to the service migration status
- the business migration status determine the server configuration mode that the server needs to switch to, and switch the server to the server configuration mode to avoid the next downtime.
- Step 106 in the server configuration mode, isolate the faulty component.
- the server runs in the switched server configuration mode, isolating the faulty components until the operation and maintenance personnel handle them. During this period, the server can continue to process business.
- the embodiment of the present application calculates the first utilization rate of the central processing unit in response to the normal startup and operation of the server; determines the faulty component in response to the shutdown and restart of the server; calculates the second utilization rate of the central processing unit; determines the business migration status based on the first utilization rate and the second utilization rate; switches the server configuration mode according to the business migration status; and isolates the faulty component in the server configuration mode.
- the server maintainability configuration method may specifically include the following steps:
- Step 201 in response to the normal startup and operation of the server, a mode selection page is displayed; the server is configured with a reliability mode, an operability mode, a balance mode and an automatic mode;
- the server can be configured with a reliability mode, an operability mode, a balance mode, and an automatic mode.
- the reliability of the reliability mode is greater than the reliability of the operability mode, and the operation efficiency of the operability mode is greater than the operation efficiency of the reliability mode;
- the operation efficiency of the balance mode is between the reliability mode and the operability mode, and the operation reliability of the balance mode is between the reliability mode and the operability mode;
- the automatic mode reuses the operability mode. mode and one of the reliability mode and the balanced mode.
- the reliability mode can filter out components with physical hardware failures without performing software repairs. Once a hardware failure is found, the system is allowed to shut down and go offline.
- the operability mode can use various means to repair or predict possible CEs (Correctable Errors) to achieve early repairs and extend the server's operating time. At the same time, it reports the faulty components to the BMC (Baseboard Management Controller) to prompt the user to replace the faulty components as soon as possible.
- BMC Baseboard Management Controller
- the available repair strategies can be evaluated for the faulty components. If the repair strategy does not affect the system performance or has a very small impact on the system performance, repair is selected. If the repair strategy used has an extreme impact on the system performance, shutdown is selected.
- the processing of the reliability mode, operability mode, and balanced mode based on RAS technology can be pre-configured according to the actual situation by relevant personnel, and the embodiments of the present application do not make specific limitations on this.
- Reliability mode, operability mode and balance mode all include a variety of RAS technologies.
- Reliability means that the system must be as reliable as possible, and will not crash unexpectedly, restart or even cause physical damage to the system. This means that a reliable system must be able to self-repair certain minor errors, and isolate errors that cannot be self-repaired as much as possible to ensure the normal operation of the rest of the system.
- Availability means that the system must be able to ensure that it works for as long as possible without going offline. Even if there are some minor problems in the system, it will not affect the normal operation of the entire system. In some cases, it can even perform Hot Plug operations to replace problematic components, thereby strictly ensuring that the system downtime is within a certain range.
- Serviceability means that the system can provide convenient diagnostic functions, such as system logs, dynamic detection and other means to facilitate management personnel to perform system diagnosis and maintenance operations, so as to detect and repair errors as early as possible.
- diagnostic functions such as system logs, dynamic detection and other means to facilitate management personnel to perform system diagnosis and maintenance operations, so as to detect and repair errors as early as possible.
- RAS the role of RAS is to ensure that the entire system can run reliably for as long as possible without going offline, and has a sufficiently powerful fault tolerance mechanism.
- the RAS technologies included in the reliability mode include but are not limited to turning off the DDDC (Double Device Data Correction) mechanism, turning off memory Patrol Scrubbing (memory name), turning off memory Post Package Repair (memory name), turning off fault memory isolation startup technology, turning off error containment default mode, turning off fault core isolation startup, turning on Viral Mode, turning on Error Log Cloaking, turning off EDP (Enhanced Downstream Port Containment, data interface), turning off PCIe data containment mode, and turning off PCIe link retraining and recovery.
- DDDC Double Device Data Correction
- Memory Patrol Scrubbing memory name
- turning off memory Post Package Repair memory name
- turning off fault memory isolation startup technology turning off error containment default mode
- turning off fault core isolation startup turning on Viral Mode
- turning off EDP Enhanced Downstream Port Containment, data interface
- turning off PCIe data containment mode turning off PCIe link retraining and recovery.
- the RAS technologies included in the operability mode include but are not limited to turning on the ADDDC (Adaptive Double DRAM Device Correction) mechanism, turning on the DDDC mechanism, turning on memory Patrol Scrubbing, turning on memory Post Package Repair, turning on the fault memory isolation startup technology threshold setting: 3000, funnel setting: 1/min (minute) to shield OS reporting/turning on memory ADDDC repair technology, turning off Viral Mode, turning off Error Log Cloaking, turning on EDPC (Enhanced Downstream Port Containment), turning on PCle data inclusion mode, and turning on PCle link retraining and recovery.
- ADDDC Adaptive Double DRAM Device Correction
- the RAS technologies included in the balanced mode include but are not limited to turning off the ADDDC mechanism/turning on the PCLS (Predictive Cache Leveling System) mechanism, turning off the DDDC mechanism, turning on memory Patrol Scrubbing, turning on memory Post Package Repair, turning on fault memory isolation startup technology, turning on Viral Mode, turning off Error Log Cloaking, turning off EDPC (Enhanced Downstream Port Containment), turning on PCle data inclusion mode, and turning on PCle link retraining and recovery.
- PCLS Predictive Cache Leveling System
- a display mode selection page can be started, and the user can select a configuration mode for the server configuration on the mode selection page.
- Step 202 Receive a selection operation on the mode selection page, and select reliability mode, operability mode, balance mode, etc.
- One of mode and auto mode is the current configuration mode
- the configuration mode selected by the user is determined, and one of the configuration modes among the reliability mode, the operability mode, the balance mode and the automatic mode is determined as the current configuration mode to configure the server.
- Step 203 calculating a first utilization rate of the central processing unit
- the utilization rate of the central processing unit at this time ie, the first utilization rate
- the step of calculating the first utilization rate of the central processing unit includes: reading power consumption data and unit heat data of the central processing unit; and determining the first utilization rate based on the power consumption data and the unit heat data.
- the BMC in the server can read the power consumption data of the CPU and the unit heat data (TDP, Thermal Design Power) of the CPU. Specifically, the power consumption data and unit heat data of the south bridge integrated circuit of the CPU can be used as the basis. Then, based on the size relationship between the power consumption data and the unit heat data, the first utilization rate is determined.
- the unit heat data is the TDP thermal power consumption, which is an indicator of the heat release of the processor.
- the step of determining the first utilization rate according to the power consumption data and the unit heat data includes: calculating a first ratio of the power consumption data to the unit heat data; and determining the first ratio as the first utilization rate.
- the ratio of power consumption data to unit heat data that is, the ratio of power consumption data/unit heat data
- the ratio of power consumption data/unit heat data can be calculated as a first ratio; the first ratio can be used as the first utilization rate.
- Step 204 in response to the server crash restart, determine the faulty component
- the server may crash and restart due to various errors.
- the server crashes and restarts the faulty component can be identified from the server.
- the step of determining a faulty component includes: in response to the server crash and restart, reading error information; determining that a component corresponding to the error information is a faulty component.
- error information may be read from operation data such as logs; based on the error information, a corresponding failed component may be determined, and the component may be determined as a failed component.
- the step of determining the faulty component before the step of reading the error information, in response to the server crashing and restarting, the step of determining the faulty component further includes: waiting for a preset period of time and entering the basic input and output system of the server.
- a preset time period may be waited for to enter the basic input/output system of the server in a delayed manner to execute the step of determining that the component corresponding to the error message is a faulty component.
- the preset time period may be determined by a person skilled in the art, and the embodiment of the present application does not limit this. For example, in an example of the present application, the preset time period is 10 minutes.
- Step 205 calculating a second utilization rate of the central processing unit
- the second utilization rate of the central processing unit can also be calculated to determine the business processing status after the restart.
- the step of calculating the second utilization of the central processing unit includes: reading power consumption data and unit heat data of the central processing unit; and determining the second utilization based on the power consumption data and the unit heat data.
- the BMC can read the power consumption data of the CPU and the unit heat data of the CPU after the CPU is restarted after shutdown, and then determine the second utilization rate according to the size relationship between the power consumption data and the unit heat data.
- the step of determining the first utilization rate according to the power consumption data and the unit heat data includes: calculating a second ratio of the power consumption data to the unit heat data; and determining the second ratio as the second utilization rate.
- the ratio of power consumption data to unit heat data that is, the ratio of power consumption data/unit heat data
- the second ratio can be used as the second utilization rate
- Step 206 determining a service migration state based on the first utilization rate and the second utilization rate
- the step of determining the service migration status based on the first utilization and the second utilization includes: calculating the service fluctuation value based on the first utilization and the second utilization; and determining the service migration status based on the service fluctuation value.
- the first utilization rate and the second utilization rate can be calculated, and the fluctuation of the service before and after the shutdown and restart can be determined based on the first utilization rate and the second utilization rate, and the service fluctuation value can be calculated.
- the specific service migration state can be determined based on the size of the service fluctuation value.
- the step of calculating the business fluctuation value includes: calculating the difference between the first utilization and the second utilization; calculating a third ratio of the difference to the first utilization; and determining the third ratio as the business fluctuation value.
- the difference between the first utilization rate and the second utilization rate can be calculated to unify that both migration in and migration out are business migrations; the difference between the first utilization rate and the second utilization rate can be used in absolute value for subsequent calculations. That is, the absolute value of the difference between the first utilization rate and the second utilization rate can be used, or the absolute value of the difference between the second utilization rate and the first utilization rate can be used for subsequent calculations. Then, the ratio of the difference to the first utilization rate is calculated, that is, the third ratio, through which the change in the business migrated in or out relative to the restart before the shutdown can be determined. The third ratio is used as the business fluctuation value.
- the step of determining the business migration status based on the business fluctuation value includes: judging whether the business fluctuation value is less than a preset business fluctuation threshold; in response to the business fluctuation value being less than the preset business fluctuation threshold, determining that the business migration status is business not migrated; in response to the business fluctuation value being not less than the preset business fluctuation threshold, determining that the business migration status is business migrated.
- the preset service fluctuation threshold can be 30%.
- the business fluctuation value is less than the preset business fluctuation threshold
- the business processing volume in the server is stable before and after the restart; it can be determined that the business migration status is that the business has not been migrated.
- the business fluctuation value is not less than the preset business fluctuation threshold
- the business processing volume in the server changes greatly before and after the restart; it can be determined that the business migration status is that the business has not been migrated.
- Step 207 switching the server configuration mode according to the service migration status
- the current configuration mode of the server can be switched to the server configuration mode corresponding to the service migration state.
- the step of switching the server configuration mode according to the business migration status includes: in response to the business migration status being that the business has not been migrated, switching the server configuration mode to the reliability mode; in response to the business migration status being that the business has been migrated, switching the server configuration mode to the operability mode.
- the server configuration mode when the service migration state is that the service has not been migrated, in response to the service migration state being that the service has not been migrated, the server configuration mode is switched to the reliability mode, and the server is configured to run based on the reliability mode.
- the server configuration mode is switched to the operability mode, and the server is configured to run based on the operability mode.
- the step of switching the server configuration mode to the reliability mode includes: in response to the business migration status being that the business has not been migrated, setting the server's mode mark to a reliability mark corresponding to the reliability mode, and controlling the server to restart; during the server restart, based on the reliability mark, configuring the server's basic input and output system options to switch to the reliability mode.
- the mode flag of the server in response to the service migration status being that the service has not been migrated, can be set to a reliability flag corresponding to the reliability mode, the reliability flag can be stored, and then the server can be controlled to restart for parameter reconfiguration.
- the server restart according to the reliability flag, all parts of the basic input and output system options of the server that are associated with the reliability mode are reconfigured and switched to parameters corresponding to the reliability mode.
- the step of switching the server configuration mode to the operability mode includes: in response to the business migration status being that the business has been migrated, setting the server's mode mark to an operability mark corresponding to the operability mode, and controlling the server to restart; during the server restart, based on the operability mark, configuring the server's basic input and output system options to switch to the operability mode.
- the mode flag of the server in response to the service migration status being that the service has been migrated, can be set to the operability flag corresponding to the operability mode, and the operability flag can be stored, and then the server can be controlled to restart for parameter reconfiguration.
- the server restart according to the operability flag, all parts of the basic input and output system options of the server that are associated with the operability mode are reconfigured and switched to the parameters corresponding to the operability mode.
- the operability mark and the reliability mark can be set in the form of the mark according to the actual situation, and the embodiment of the present application does not make any specific limitation.
- Step 208 in the server configuration mode, isolate the faulty component.
- the embodiment of the present application displays a mode selection page in response to the normal startup and operation of the server; the server is configured with a reliability mode, an operability mode, a balance mode, and an automatic mode; a selection operation for the mode selection page is received, and one of the reliability mode, the operability mode, the balance mode, and the automatic mode is selected as the current configuration mode; the first utilization of the central processing unit is calculated; the faulty component is determined in response to the shutdown and restart of the server; the second utilization of the central processing unit is calculated; the business migration status is determined based on the first utilization and the second utilization; the server configuration mode is switched according to the business migration status; and the faulty component is isolated in the server configuration mode.
- the server maintainability configuration method may specifically include the following steps:
- the BMC reads the CPU utilization through PCIE and records that the utilization is A.
- the server will automatically restart after the crash (the restart mechanism may be fatal of the PCIE device, UCE of the memory, or IERR of the CPU, etc.).
- BIOS reads the RAS mode flag saved by BMC (baseboard controller), and then configures the corresponding BIOS parameters according to different RAS modes.
- the server maintainability configuration device may specifically include the following modules:
- a first calculation module 401 configured to calculate a first utilization rate of a central processing unit in response to a normal startup and operation of the server;
- a restart module 402 for determining a faulty component in response to a server shutdown restart
- the second calculation module 403 is used to calculate the second utilization rate of the central processing unit
- a service migration determination module 404 configured to determine a service migration state based on the first utilization rate and the second utilization rate
- a switching module 405 is used to switch the server configuration mode according to the service migration status
- the isolation module 406 is used to isolate the faulty component in the server configuration mode.
- the first calculation module 401 includes:
- the first reading submodule is used to read the power consumption data and unit heat data of the central processing unit;
- the first utilization rate determination submodule is used to determine the first utilization rate according to the power consumption data and the unit heat data.
- the first utilization determination submodule includes:
- a first calculation unit used for calculating a first ratio of power consumption data to unit heat data
- the first utilization rate determining unit is used to determine the first ratio as the first utilization rate.
- the restart module 402 includes:
- the restart submodule is used to respond to the server crash and restart, and read the error information
- the fault determination submodule is used to determine that the component corresponding to the error information is a faulty component.
- the restart module 402 further includes:
- the waiting submodule is used to wait for a preset time and enter the basic input and output system of the server.
- the second calculation module 403 includes:
- the second reading submodule is used to read the power consumption data and unit heat data of the central processing unit
- the second utilization rate determination submodule is used to determine the second utilization rate according to the power consumption data and the unit heat data.
- the second utilization determination submodule includes:
- a second calculation unit used for calculating a second ratio of the power consumption data to the unit heat data
- the second utilization rate determining unit is used to determine the second ratio as the second utilization rate.
- the service migration determination module 404 includes:
- a business fluctuation value determination submodule used to calculate the business fluctuation value based on the first utilization rate and the second utilization rate
- the service migration status determination submodule is used to determine the service migration status based on the service fluctuation value.
- the service fluctuation value determination submodule includes:
- a difference calculation unit used for calculating the difference between the first utilization rate and the second utilization rate
- a third ratio calculation unit used for calculating a third ratio of the difference to the first utilization rate
- the business fluctuation value is determined correspondingly, and is used to determine the third ratio as the business fluctuation value.
- the service migration status determination submodule includes:
- a judging unit used to judge whether the business fluctuation value is less than a preset business fluctuation threshold
- a first migration determination unit configured to determine that the service migration state is service not migrated in response to a service fluctuation value being less than a preset service fluctuation threshold
- the second migration determination unit is configured to determine that the service migration state is that the service has been migrated in response to the service fluctuation value being not less than a preset service fluctuation threshold.
- the server configuration mode includes a reliability mode and an operability mode, the operation reliability of the reliability mode is greater than the operation reliability of the operability mode, and the operation efficiency of the operability mode is greater than the operation efficiency of the reliability mode;
- the switching module 405 includes:
- a first switching submodule configured to switch the server configuration mode to a reliability mode in response to the service migration state being that the service has not been migrated;
- the second switching submodule is used to switch the server configuration mode to the operability mode in response to the service migration status being that the service has been migrated.
- the first switching submodule includes:
- a first marking unit configured to, in response to the service migration state being that the service has not been migrated, set a mode mark of the server to a reliability mark corresponding to the reliability mode, and control the server to restart;
- the first configuration unit is used to configure a basic input and output system option of the server to switch to a reliability mode based on the reliability mark during the restart of the server.
- the second switching submodule includes:
- a second marking unit is used for setting the mode mark of the server to an operability mark corresponding to the operability mode in response to the service migration state being that the service has been migrated, and controlling the server to restart;
- the second configuration unit is used to configure the basic input and output system options of the server based on the operability mark during the server restart to switch to the operability mode.
- the server configuration mode also includes a balanced mode and an automatic mode.
- the operating efficiency of the balanced mode is between the reliability mode and the operability mode, and the operating reliability of the balanced mode is between the reliability mode and the operability mode; the automatic mode multiplexes one of the operability mode, the reliability mode and the balanced mode.
- the device further includes:
- the display module is used to display a mode selection page in response to the normal startup and operation of the server.
- the device further includes:
- the selection module is used to receive a selection operation on the mode selection page, and select one of the reliability mode, operability mode, balance mode and automatic mode as the current configuration mode.
- the preset service fluctuation threshold is 30%.
- the embodiment of the present application calculates the first utilization rate of the central processing unit in response to the normal startup and operation of the server; determines the faulty component in response to the shutdown and restart of the server; calculates the second utilization rate of the central processing unit; determines the business migration status based on the first utilization rate and the second utilization rate; switches the server configuration mode according to the business migration status; and isolates the faulty component in the server configuration mode.
- the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
- an embodiment of the present application further provides an electronic device, including:
- the processor 501 executes the computer program to perform a server maintainability configuration method as described in any one of the embodiments of the present application.
- the server maintainability configuration method includes:
- the step of calculating the first utilization rate of the central processing unit includes:
- a first utilization rate is determined according to the power consumption data and the unit heat data.
- the step of determining the first utilization rate according to the power consumption data and the unit heat data includes:
- the first ratio is determined as a first utilization rate.
- the step of determining a faulty component in response to a server shutdown and restart, includes:
- the step of determining the faulty component before the step of reading the error information, in response to the server crashing and restarting, the step of determining the faulty component further includes:
- the step of calculating the second utilization rate of the central processing unit includes:
- a second utilization rate is determined according to the power consumption data and the unit heat data.
- the step of determining the first utilization rate according to the power consumption data and the unit heat data includes: include:
- a second ratio is determined as a second utilization rate.
- the step of determining the service migration state based on the first utilization rate and the second utilization rate includes:
- the service migration status is determined based on the service fluctuation value.
- the step of calculating the service fluctuation value based on the first utilization rate and the second utilization rate includes:
- the third ratio is determined as the business fluctuation value.
- the step of determining the service migration state based on the service fluctuation value includes:
- the server configuration mode includes a reliability mode and an operability mode, the operation reliability of the reliability mode is greater than the operation reliability of the operability mode, and the operation efficiency of the operability mode is greater than the operation efficiency of the reliability mode; according to the service migration state, the step of switching the server configuration mode includes:
- the server configuration mode is switched to an operability mode.
- the step of switching the server configuration mode to the reliability mode includes:
- the step of switching the server configuration mode to the operability mode includes:
- the server configuration mode also includes a balanced mode and an automatic mode.
- the operating efficiency of the balanced mode is between the reliability mode and the operability mode, and the operating reliability of the balanced mode is between the reliability mode and the operability mode; the automatic mode multiplexes one of the operability mode, the reliability mode and the balanced mode.
- the method further comprises:
- a mode selection page is displayed.
- the method further comprises:
- a selection operation on a mode selection page is received, and one of a reliability mode, an operability mode, a balance mode, and an automatic mode is selected as a current configuration mode.
- the preset service fluctuation threshold is 30%.
- the memory may include a random access memory (RAM) or a non-volatile memory, such as at least one disk storage.
- the memory may also be at least one storage device located away from the aforementioned processor.
- processors can be general-purpose processors, including central processing units (CPU), network processors (NP), etc.; they can also be digital signal processors (DSP), application specific integrated circuits (ASIC), field programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
- CPU central processing units
- NP network processors
- DSP digital signal processors
- ASIC application specific integrated circuits
- FPGA field programmable gate arrays
- the embodiment of the present application calculates the first utilization rate of the central processing unit in response to the normal startup and operation of the server; determines the faulty component in response to the shutdown and restart of the server; calculates the second utilization rate of the central processing unit; determines the business migration status based on the first utilization rate and the second utilization rate; switches the server configuration mode according to the business migration status; and isolates the faulty component in the server configuration mode.
- the embodiment of the present application further provides a non-volatile readable storage medium 601, on which a computer program is stored, and when the computer program is executed by a processor, a server maintainability configuration method as in any one of the embodiments of the present application is executed.
- the server maintainability configuration method includes:
- the step of calculating the first utilization rate of the central processing unit includes:
- a first utilization rate is determined according to the power consumption data and the unit heat data.
- the step of determining the first utilization rate according to the power consumption data and the unit heat data includes:
- the first ratio is determined as a first utilization rate.
- the step of determining a faulty component in response to a server shutdown and restart, includes:
- the step of determining the faulty component before the step of reading the error information, in response to the server crashing and restarting, the step of determining the faulty component further includes:
- the step of calculating the second utilization rate of the central processing unit includes:
- a second utilization rate is determined according to the power consumption data and the unit heat data.
- the step of determining the first utilization rate according to the power consumption data and the unit heat data includes:
- a second ratio is determined as a second utilization rate.
- the step of determining the service migration state based on the first utilization rate and the second utilization rate includes:
- the service migration status is determined based on the service fluctuation value.
- the step of calculating the service fluctuation value based on the first utilization rate and the second utilization rate includes:
- the third ratio is determined as the business fluctuation value.
- the step of determining the service migration state based on the service fluctuation value includes:
- the server configuration mode includes a reliability mode and an operability mode, the operation reliability of the reliability mode is greater than the operation reliability of the operability mode, and the operation efficiency of the operability mode is greater than the operation efficiency of the reliability mode; according to the service migration state, the step of switching the server configuration mode includes:
- the server configuration mode is switched to an operability mode.
- the step of switching the server configuration mode to the reliability mode includes:
- the step of switching the server configuration mode to the operability mode includes:
- the server configuration mode further includes a balance mode and an automatic mode.
- the operating efficiency of the balance mode is between the reliability mode and the operability mode.
- the operating reliability of the balance mode is between the reliability mode and the operability mode. between operability modes; the automatic mode reuses one of the operability mode, reliability mode, and balance mode.
- the method further comprises:
- a mode selection page is displayed.
- the method further comprises:
- a selection operation on a mode selection page is received, and one of a reliability mode, an operability mode, a balance mode, and an automatic mode is selected as a current configuration mode.
- the preset service fluctuation threshold is 30%.
- the embodiment of the present application calculates the first utilization rate of the central processing unit in response to the normal startup and operation of the server; determines the faulty component in response to the shutdown and restart of the server; calculates the second utilization rate of the central processing unit; determines the business migration status based on the first utilization rate and the second utilization rate; switches the server configuration mode according to the business migration status; and isolates the faulty component in the server configuration mode.
- the embodiments of the present application can be provided as methods, devices, or computer program products. Therefore, the embodiments of the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the embodiments of the present application can take the form of a computer program product implemented on one or more non-volatile readable storage media (including but not limited to disk storage, CD-ROM (Compact Disc Read-Only Memory, read-only optical disk storage medium), optical storage, etc.) containing computer-usable program code.
- non-volatile readable storage media including but not limited to disk storage, CD-ROM (Compact Disc Read-Only Memory, read-only optical disk storage medium), optical storage, etc.
- each process and/or box in the flowchart and/or block diagram, and the combination of the process and/or box in the flowchart and/or block diagram can be realized by computer program instructions.
- These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing terminal device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal device produce a device for realizing the function specified in one process or multiple processes in the flowchart and/or one box or multiple boxes in the block diagram.
- These computer program instructions may also be stored in a non-volatile readable storage medium that can direct a computer or other programmable data processing terminal device to operate in a specific manner, so that the instructions stored in the non-volatile readable storage medium produce a manufactured product including an instruction device that implements the functions specified in one or more processes in the flowchart and/or one or more boxes in the block diagram.
- These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device so that a series of operating steps are executed on the computer or other programmable terminal device to produce computer-implemented processing, so that the instructions executed on the computer or other programmable terminal device provide steps for implementing the functions specified in one or more processes in the flowchart and/or one or more boxes in the block diagram.
- the server maintainability configuration method, device, electronic device and non-volatile readable storage medium provided by the present application are introduced in detail above. Specific examples are used in this article to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea; at the same time, for general technical personnel in this field, according to the idea of the present application, there will be changes in the specific implementation method and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.
Landscapes
- Engineering & Computer Science (AREA)
- Theoretical Computer Science (AREA)
- Physics & Mathematics (AREA)
- General Engineering & Computer Science (AREA)
- General Physics & Mathematics (AREA)
- Software Systems (AREA)
- Quality & Reliability (AREA)
- Health & Medical Sciences (AREA)
- Biomedical Technology (AREA)
- Hardware Redundancy (AREA)
Abstract
本申请实施例提供了一种服务器可维护性配置方法、装置、电子设备和非易失性可读存储介质,涉及计算机系统及存储技术领域;包括:响应于服务器的正常启动运行,计算中央处理器的第一利用率(101);响应于服务器的宕机重启,确定故障部件(102);计算中央处理器的第二利用率(103);基于第一利用率和第二利用率,确定业务迁移状态(104);依据业务迁移状态,切换服务器配置模式(105);在服务器配置模式中,隔离故障部件(106)。通过本申请实施例通过判断客户的业务是否迁移,根据是否业务迁移来启动不同服务器配置模式,可以降低服务器的宕机率。
Description
相关申请的交叉引用
本申请要求于2023年08月14日提交中国专利局,申请号为202311019304.8,申请名称为“服务器可维护性配置方法、装置、电子设备和存储介质”的中国专利申请的优先权,其全部内容通过引用结合在本申请中。
本申请涉及计算机系统及存储技术领域,特别是涉及一种服务器可维护性配置方法、一种服务器可维护性配置装置、一种电子设备和一种非易失性可读存储介质。
服务器故障可划分为宕机类故障和非宕机类故障两大类。宕机类故障主要体现在开机过程宕机及运行时宕机两部分。服务器对故障部件有一定的修复功能,即使可能服务器出现硬件故障也可以使用必要的手段使其正常运行下去。这个过程即服务器的RAS(Reliability Availability Serviceability,可靠性可用性可服务性)功能。RAS修复在部分场景下默认配置一种参数继续运行直至运维人员进行故障排除,而是会影响服务器的性能和运行效率,可能进一步导致服务器宕机。
发明内容
鉴于上述问题,提出了本申请的一些实施例以便提供一种克服上述问题或者至少部分地解决上述问题的一种服务器可维护性配置方法、一种服务器可维护性配置装置、一种电子设备和一种非易失性可读存储介质。
为了解决上述问题,在本申请的第一个方面,本申请的一些实施例公开了一种服务器可维护性配置方法,包括:
响应于服务器的正常启动运行,计算中央处理器的第一利用率;
响应于服务器的宕机重启,确定故障部件;
计算中央处理器的第二利用率;
基于第一利用率和第二利用率,确定业务迁移状态;
依据业务迁移状态,切换服务器配置模式;
在服务器配置模式中,隔离故障部件。
在本申请的一些实施例中,计算中央处理器的第一利用率的步骤包括:
读取中央处理器的功耗数据和单位热量数据;
依据功耗数据和单位热量数据,确定第一利用率。
在本申请的一些实施例中,依据功耗数据和单位热量数据,确定第一利用率的步骤包括:
计算功耗数据与单位热量数据的第一比值;
确定第一比值为第一利用率。
在本申请的一些实施例中,响应于服务器的宕机重启,确定故障部件的步骤包括:
响应于服务器的宕机重启,读取错误信息;
确定错误信息对应的部件为故障部件。
在本申请的一些实施例中,在读取错误信息的步骤之前,响应于服务器的宕机重启,确定故障部件的步骤还包括:
等待预设时长,进入服务器的基本输入输出系统。
在本申请的一些实施例中,计算中央处理器的第二利用率的步骤包括:
读取中央处理器的功耗数据和单位热量数据;
依据功耗数据和单位热量数据,确定第二利用率。
在本申请的一些实施例中,依据功耗数据和单位热量数据,确定第一利用率的步骤包括:
计算功耗数据与单位热量数据的第二比值;
确定第二比值为第二利用率。
在本申请的一些实施例中,基于第一利用率和第二利用率,确定业务迁移状态的步骤包括:
基于第一利用率和第二利用率,计算业务波动值;
基于业务波动值确定业务迁移状态。
在本申请的一些实施例中,基于第一利用率和第二利用率,计算业务波动值的步骤包括:
计算第一利用率和第二利用率的差值;
计算差值与第一利用率的第三比值;
确定第三比值为业务波动值。
在本申请的一些实施例中,基于业务波动值确定业务迁移状态的步骤包括:
判断业务波动值是否小于预设业务波动阈值;
响应于业务波动值小于预设业务波动阈值,确定业务迁移状态为业务未迁移;
响应于业务波动值不小于预设业务波动阈值,确定业务迁移状态为业务已迁移。
在本申请的一些实施例中,服务器配置模式包括可靠性模式和可运行性模式,可靠性模式的运行可靠性大于可运行性模式的运行可靠性,可运行性模式的运行效率大于可靠性模式的运行效率;依据业务迁移状态,切换服务器配置模式的步骤包括:
响应于业务迁移状态为业务未迁移,切换服务器配置模式为可靠性模式;
响应于业务迁移状态为业务已迁移,切换服务器配置模式为可运行性模式。
在本申请的一些实施例中,响应于业务迁移状态为业务未迁移,切换服务器配置模式为可靠性模式的步骤包括:
响应于业务迁移状态为业务未迁移,设置服务器的模式标记为可靠性模式对应的可靠性标记,控制服务器重启;
在服务器重启期间,基于可靠性标记,配置服务器的基本输入输出系统选项,以切换为可靠性模式。
在本申请的一些实施例中,响应于业务迁移状态为业务已迁移,切换服务器配置模式为可运行性模式的步骤包括:
响应于业务迁移状态为业务已迁移,设置服务器的模式标记为可运行性模式对应的可运行性标记,控制服务器重启;
在服务器重启期间,基于可运行性标记,配置服务器的基本输入输出系统选项,以切换
为可运行性模式。
在本申请的一些实施例中,服务器配置模式还包括平衡模式和自动模式,平衡模式的运行效率位于可靠性模式和可运行性模式之间,平衡模式的运行可靠性位于可靠性模式和可运行性模式之间;自动模式复用可运行性模式和可靠性模式和平衡模式中的一个。
在本申请的一些实施例中,方法还包括:
响应于服务器的正常启动运行,显示模式选择页面。
在本申请的一些实施例中,方法还包括:
接收针对模式选择页面的选择操作,选择可靠性模式、可运行性模式、平衡模式和自动模式中的一个为当前配置模式。
在本申请的一些实施例中,预设业务波动阈值为30%。
在本申请的第一个方面,本申请实施例公开了一种服务器可维护性配置装置,包括:
第一计算模块,用于响应于服务器的正常启动运行,计算中央处理器的第一利用率;
重启模块,用于响应于服务器的宕机重启,确定故障部件;
第二计算模块,用于计算中央处理器的第二利用率;
业务迁移判断模块,用于基于第一利用率和第二利用率,确定业务迁移状态;
切换模块,用于依据业务迁移状态,切换服务器配置模式;
隔离模块,用于在服务器配置模式中,隔离故障部件。
在本申请的第三个方面,本申请实施例公开了一种电子设备,包括处理器、存储器及存储在存储器上并能够在处理器上运行的计算机程序,计算机程序被处理器执行时实现如上的服务器可维护性配置方法的步骤。
在本申请的第四个方面,本申请实施例公开了一种非易失性可读存储介质,非易失性可读存储介质上存储计算机程序,计算机程序被处理器执行时实现如上的服务器可维护性配置方法的步骤。
本申请实施例包括以下优点:
本申请实施例通过响应于服务器的正常启动运行,计算中央处理器的第一利用率;响应于服务器的宕机重启,确定故障部件;计算中央处理器的第二利用率;基于第一利用率和第二利用率,确定业务迁移状态;依据业务迁移状态,切换服务器配置模式;在服务器配置模式中,隔离故障部件。通过在正常启动和重启时的中央处理器的利用率判断客户的业务是否迁移,根据是否业务迁移来启动不同服务器配置模式,以使服务器可以自动切换配置模型,降低服务器的宕机率。
图1是本申请的一种服务器可维护性配置方法实施例的步骤流程图;
图2是本申请的另一种服务器可维护性配置方法实施例的步骤流程图;
图3是本申请的一种服务器可维护性配置方法示例的框架示意图;
图4是本申请的一种服务器可维护性配置装置实施例的结构框图;
图5是本申请实施例提供的一种电子设备的结构框图;
图6是本申请实施例提供的一种存储介质的结构框图。
为使本申请的上述目的、特征和优点能够更加明显易懂,下面结合附图和具体实施方式
对本申请作进一步详细的说明。
参照图1,示出了本申请的一种服务器可维护性配置方法实施例的步骤流程图,服务器可维护性配置方法具体可以包括如下步骤:
步骤101,响应于服务器的正常启动运行,计算中央处理器的第一利用率;
在服务器的正常启动运行,进行业务处理时,计算中央处理的此时的中央处理器利用率,即为第一利用率。
步骤102,响应于服务器的宕机重启,确定故障部件;
当服务器因为PCIE(peripheral component interconnect express,高速串行计算机扩展总线)设备的fatal(致命错误)或者内存的UCE(Unit Cell Error,单元错误),或者CPU(中央处理器,Central Processing Unit)的IERR(Internal Error,内部错误)宕机而重启时,可以基于读取日志的方式确定自身的故障部件。
步骤103,计算中央处理器的第二利用率;
在宕机重启,重新对业务进行处理时,可以计算中央处理的此时的中央处理器利用率,即为第二利用率。
步骤104,基于第一利用率和第二利用率,确定业务迁移状态;
根据第一利用率和第二利用率之间的关系,确定服务在重启期间,其业务是否被用户迁移至其他服务器,确定业务迁移状态。
步骤105,依据业务迁移状态,切换服务器配置模式;
依据业务迁移状态,确定服务器需要切换的服务器配置模式,将服务器切换至服务器配置模式下,以避免下次的宕机。
步骤106,在服务器配置模式中,隔离故障部件。
服务器运行在切换的服务器配置模式中,对故障部件进行隔离,直至运维人员进行处理,已在此期间可以服务器可以继续对业务进行处理。
本申请实施例通过响应于服务器的正常启动运行,计算中央处理器的第一利用率;响应于服务器的宕机重启,确定故障部件;计算中央处理器的第二利用率;基于第一利用率和第二利用率,确定业务迁移状态;依据业务迁移状态,切换服务器配置模式;在服务器配置模式中,隔离故障部件。通过在正常启动和重启时的中央处理器的利用率判断客户的业务是否迁移,根据是否业务迁移来启动不同服务器配置模式,以使服务器可以自动切换配置模型,降低服务器的宕机率。
参照图2,示出了本申请的另一种服务器可维护性配置方法实施例的步骤流程图,服务器可维护性配置方法具体可以包括如下步骤:
步骤201,响应于服务器的正常启动运行,显示模式选择页面;服务器配置有可靠性模式、可运行性模式、平衡模式和自动模式;
在本申请实施例中,服务器可配置有可靠性模式、可运行性模式、平衡模式和自动模式。其中,可靠性模式的运行可靠性大于可运行性模式的运行可靠性,可运行性模式的运行效率大于可靠性模式的运行效率;平衡模式的运行效率位于可靠性模式和可运行性模式之间,平衡模式的运行可靠性位于可靠性模式和可运行性模式之间;自动模式复用可运行性模
式和可靠性模式和平衡模式中的一个。可靠性模式可以筛选有物理硬件故障的部件,而不去做软件修复,一旦发现硬件故障则允许系统宕机下线。可运行性模式可以使用各种手段修复或者预测可能的CE(Correctable Error,可修正错误)做到提前修复,延长服务器的运行时间,同时上报故障部件给BMC(基板管理控制器,Baseboard Management Controller),提示用户尽快替换故障部件。平衡模式中可以筛选有物理硬件故障的部件,针对故障部件评估可采用的修复策略,如修复策略不影响系统性能或者极小的影响系统性能则选择修复,若采用的修复策略极度影响系统性能则选择宕机下线。对于可靠性模式、可运行性模式、平衡模式基于RAS技术的处理,可以根据相关人员基于实际情况进行预先配置,本申请实施例对此不作具体限定。
可靠性模式、可运行性模式和平衡模式均包含多种RAS技术。RAS技术中,Reliability(可靠性)指的是系统必须尽可能的可靠,不会意外的崩溃,重启甚至导致系统物理损坏,这意味着一个具有可靠性的系统必须能够对于某些小的错误能够做到自修复,对于无法自修复的错误也尽可能进行隔离,保障系统其余部分正常运转。Availability(可用性)指的是系统必须能够确保尽可能长时间工作而不下线,即使系统出现一些小的问题也不会影响整个系统的正常运行,在某些情况下甚至可以进行Hot Plug(热插拔)的操作,替换有问题的组件,从而严格的确保系统的宕机时间在一定范围内。Serviceability(可服务性)指的是系统能够提供便利的诊断功能,如系统日志,动态检测等手段方便管理人员进行系统诊断和维护操作,从而及早的发现错误并且修复错误。RAS作为一个整体,其作用在于确保整个系统尽可能长期可靠的运行而不下线,并且具备足够强大的容错机制。
可靠性模式包含的RAS技术包括但不限于关闭DDDC(Double Device Data Correction,双设备数据校正)机制、关闭内存Patrol Scrubbing(内存名)、关闭内存Post Package Repair(内存名)、关闭故障内存隔离启动技术、关闭错误包容默认模式、关闭故障核心隔离启动、打开Viral Mode(病毒模式)、打开Error Log Cloaking(错误日志隐藏)、关闭EDP(Enhanced Downstream Port Containment,数据接口)、关闭PCIe数据包容模式、关闭PCIe链路重新训练和恢复。
可运行性模式包含的RAS技术包括但不限于打开ADDDC(Adaptive Double DRAM Device Correction自适应双设备数据校正)机制、打开DDDC机制、打开内存Patrol Scrubbing、打开内存Post Package Repair、打开故障内存隔离启动技术阈值设定:3000、漏斗设定:1个/min(分钟)屏蔽OS上报/打开内存ADDDC修复技术、关闭Viral Mode、关闭Error Log Cloaking、打开EDPC(Enhanced Downstream Port Containment,增强型下行端口隔离)、打开PCle数据包容模式、打开PCle链路重新训练和恢复。
平衡模式包含的RAS技术包括但不限于关闭ADDDC机制/打开PCLS(Predictive Cache Leveling System,一种预测性缓存调平系统)机制、关闭DDDC机制、打开内存Patrol Scrubbing、打开内存Post Package Repair、打开故障内存隔离启动技术、打开Viral Mode、关闭Error Log Cloaking、关闭EDPC(Enhanced Downstream Port Containment)、打开PCle数据包容模式、打开PCle链路重新训练和恢复。
在服务器正常启动开机运行时,可以启动显示模式选择页面,用户可以在模式选择页面中选择对服务器配置的配置模式。
步骤202,接收针对模式选择页面的选择操作,选择可靠性模式、可运行性模式、平衡
模式和自动模式中的一个为当前配置模式;
根据用户的针对模式选择页面的选择操作,确定用户选择的配置模式,从可靠性模式、可运行性模式、平衡模式和自动模式中确定其中一个配置模式为当前配置模式,以对服务器进行配置。
步骤203,计算中央处理器的第一利用率;
可以在服务器运行一定时间,如10分钟后,可以计算出中央处理器的此时的利用率,即第一利用率。
在本申请的一些可选实施例中,计算中央处理器的第一利用率的步骤包括:读取中央处理器的功耗数据和单位热量数据;依据功耗数据和单位热量数据,确定第一利用率。
服务器中的BMC可读取中央处理器的功耗数据和中央处理器的单位热量数据(TDP,Thermal Design Power)。具体地,可以中央处理器的南桥集成电路的功耗数据和单位热量数据为准。然后依据功耗数据和单位热量数据之间的大小关系,确定第一利用率。单位热量数据为TDP热功耗,是处理器热量释放的指标。
具体地,依据功耗数据和单位热量数据,确定第一利用率的步骤包括:计算功耗数据与单位热量数据的第一比值;确定第一比值为第一利用率。
在实际应用中,可以计算功耗数据与单位热量数据的比值,即功耗数据/单位热量数据的比值,为第一比值;可将该第一比值作为第一利用率。
步骤204,响应于服务器的宕机重启,确定故障部件;
在服务器运行期间,可能会因各种错误而宕机重启。在服务器的宕机重启时,可以从服务器中确定出故障部件。
在本申请的一些可选实施例中,响应于服务器的宕机重启,确定故障部件的步骤包括:响应于服务器的宕机重启,读取错误信息;确定错误信息对应的部件为故障部件。
响应于服务器的宕机重启,可以从日志等运行数据中读取错误信息;基于错误信息,确定对应的发生故障的部件,将该部件确定为故障部件。
在本申请的一些可选实施例中,在读取错误信息的步骤之前,响应于服务器的宕机重启,确定故障部件的步骤还包括:等待预设时长,进入服务器的基本输入输出系统。
当错误信息是为UCE故障或者IERR故障时,可以等待预设时长,延时进入服务器的基本输入输出系统,以执行确定错误信息对应的部件为故障部件的步骤。其中预设时长可以基于本领域技术人员确定,本申请实施例对此不作限定。如在本申请的一示例中,预设时长为10分钟。
步骤205,计算中央处理器的第二利用率;
在服务器宕机重启后,还可以计算出中央处理器的第二利用率,以确定重启后的业务处理情况。
在本申请的一些可选实施例中,计算中央处理器的第二利用率的步骤包括:读取中央处理器的功耗数据和单位热量数据;依据功耗数据和单位热量数据,确定第二利用率。
与第一利用率相似地,BMC可读取宕机重启后的中央处理器的功耗数据和中央处理器的单位热量数据,然后依据功耗数据和单位热量数据之间的大小关系,确定第二利用率。
具体地,依据功耗数据和单位热量数据,确定第一利用率的步骤包括:计算功耗数据与单位热量数据的第二比值;确定第二比值为第二利用率。
在实际应用中,可以计算功耗数据与单位热量数据的比值,即功耗数据/单位热量数据的比值,为第二比值;可将该第二比值作为第二利用率。
步骤206,基于第一利用率和第二利用率,确定业务迁移状态;
可以基于第一利用率和第二利用率的关系,确定该服务器是否发生业务迁移,将业务迁移至其他服务器。
在本申请的一些可选实施例中,基于第一利用率和第二利用率,确定业务迁移状态的步骤包括:基于第一利用率和第二利用率,计算业务波动值;基于业务波动值确定业务迁移状态。
在本申请实施例中,可以计算第一利用率和第二利用率,基于第一利用率和第二利用率确定宕机重启前后,业务的波动情况,计算出业务波动值。基于业务波动值的大小确定具体的业务迁移状态。
具体地,基于第一利用率和第二利用率,计算业务波动值的步骤包括:计算第一利用率和第二利用率的差值;计算差值与第一利用率的第三比值;确定第三比值为业务波动值。
在本申请实施例中,可以计算第一利用率和第二利用率的差值,为统一一迁入与迁出均为业务迁移;第一利用率和第二利用率的差值可以采用绝对值进行后续的计算。即可以采用第一利用率减去第二利用率的差值的绝对值,也可以是采用第二利用率减去第一利用率的差值的绝对值进入后续的计算。然后计算差值与第一利用率的比值,即第三比值,通过该比值可以确定迁入或迁出的业务,相对宕机重启前的变化量。将第三比值为业务波动值。
具体地,基于业务波动值确定业务迁移状态的步骤包括:判断业务波动值是否小于预设业务波动阈值;响应于业务波动值小于预设业务波动阈值,确定业务迁移状态为业务未迁移;响应于业务波动值不小于预设业务波动阈值,确定业务迁移状态为业务已迁移。
在实际应用中,可以判断业务波动值是否小于预设业务波动阈值确定具体的业务迁移状态,其中,预设业务波动阈值可以根据实际情况进行确定,本申请实施例对此不做限定。在本申请的一些优选示例中,预设业务波动阈值可以为30%。
当业务波动值小于预设业务波动阈值时,响应于业务波动值小于预设业务波动阈值,即重启前后服务器中的业务处理量稳定;可以确定业务迁移状态为业务未迁移。
当业务波动值不小于预设业务波动阈值时,响应于业务波动值不小于预设业务波动阈值,即重启前后服务器中的业务处理量变化较大;可以确定业务迁移状态为业务未迁移。
步骤207,依据业务迁移状态,切换服务器配置模式;
根据不同的业务迁移状态,可以将服务器的当前配置模式切换成业务迁移状态对应的服务器配置模式。
在本申请的一些可选实施例中,依据业务迁移状态,切换服务器配置模式的步骤包括:响应于业务迁移状态为业务未迁移,切换服务器配置模式为可靠性模式;响应于业务迁移状态为业务已迁移,切换服务器配置模式为可运行性模式。
在本申请实施例中,当业务迁移状态为业务未迁移时,响应于业务迁移状态为业务未迁移,切换服务器配置模式为可靠性模式,对服务器基于可靠性模式进行配置运行。当业务迁移状态为业务已迁移时,响应于业务迁移状态为业务已迁移,切换服务器配置模式为可运行性模式,对服务器基于可运行性模式进行配置运行。
具体地,响应于业务迁移状态为业务未迁移,切换服务器配置模式为可靠性模式的步骤
包括:响应于业务迁移状态为业务未迁移,设置服务器的模式标记为可靠性模式对应的可靠性标记,控制服务器重启;在服务器重启期间,基于可靠性标记,配置服务器的基本输入输出系统选项,以切换为可靠性模式。
在实际应用中,可以响应于业务迁移状态为业务未迁移,将服务器的模式标记设置为可靠性模式对应的可靠性标记,并对可靠性标记进行存储,然后控制服务器重启以进行参数重新配置。在服务器重启期间,依据可靠性标记,将服务器的基本输入输出系统选项中与可靠性模式关联的部分都进行重新配置,切换为可靠性模式对应的参数。
具体地,响应于业务迁移状态为业务已迁移,切换服务器配置模式为可运行性模式的步骤包括:响应于业务迁移状态为业务已迁移,设置服务器的模式标记为可运行性模式对应的可运行性标记,控制服务器重启;在服务器重启期间,基于可运行性标记,配置服务器的基本输入输出系统选项,以切换为可运行性模式。
在实际应用中,可以响应于业务迁移状态为业务已迁移,将服务器的模式标记设置为可运行性模式对应的可运行性标记,并对可运行性标记进行存储,然后控制服务器重启以进行参数重新配置。在服务器重启期间,依据可运行性标记,将服务器的基本输入输出系统选项中与可运行性模式关联的部分都进行重新配置,切换为可运行性模式对应的参数。
其中,可运行性标记和可靠性标记可以根据实际情况设置标记的形式,本申请实施例不作具体限定。
步骤208,在服务器配置模式中,隔离故障部件。
在切换至新的服务器配置模式后,在服务器配置模式下运行,隔离发生故障的故障部件,以确定服务器的正常运行。
本申请实施例通过响应于服务器的正常启动运行,显示模式选择页面;服务器配置有可靠性模式、可运行性模式、平衡模式和自动模式;接收针对模式选择页面的选择操作,选择可靠性模式、可运行性模式、平衡模式和自动模式中的一个为当前配置模式;计算中央处理器的第一利用率;响应于服务器的宕机重启,确定故障部件;计算中央处理器的第二利用率;基于第一利用率和第二利用率,确定业务迁移状态;依据业务迁移状态,切换服务器配置模式;在服务器配置模式中,隔离故障部件。通过在正常启动和重启时的中央处理器的利用率判断客户的业务是否迁移,根据是否业务迁移来启动不同服务器配置模式,以使服务器可以自动切换配置模型,降低服务器的宕机率。
为了使本领域技术人员能够更好地理解本申请实施例,下面通过一个例子对本申请实施例加以说明:
1)首先在BIOS(基本输入输出系统)交互界面BIOS setup(选项)下增加RAS Mode(模式)选项,可选项有Sensitive mode(可运行性模式)、recovery mode(可靠模式性)、balanced mode(平衡模式)、auto mode(自动模式)。
2)根据用户选择操作选择其中一个模式。如选择自动模式。
参照图3,示出了本申请的一种服务器可维护性配置方法示例的示意图,服务器可维护性配置方法具体可以包括如下步骤:
1、服务器正常运行客户业务阶段,由BMC通过PCIE读取CPU的利用率记录利用率为A。
2、首次客户故障机器宕机,宕机以后服务器会自动重启(重启机制可能是PCIE设备的fatal或者内存的UCE,或者CPU的IERR等)。
3、BMC重启检测重启原因是由于UCE故障或者IERR故障。则进入系统10分钟以后,再次读取cpu利用率B。
4、计算是否判定切换何种模式的公式是:数值对比|A-B|÷A<0.3的真假。认为客户业务运行中有百分之30的波动是正常的,如果业务被迁移出去,cpu利用率是很低的,远远低于30%的CPU利用率波动。
5、数值对比|A-B|÷A<0.3为真,则表明客户业务未从故障机器迁移出去,则设置RAS mode flag(标记)为recovery mode flag。接着由BMC主动重启服务器。
6、数值对比|A-B|÷A<0.3为假,则表明客户业务已经从故障机器迁移出去,则设置RAS mode flag为Sensitive mode flag。接着由BMC主动重启服务器。
7、服务器重启阶段由BIOS读取BMC(基板控制器)保存的RAS mode flag,然后按照不同的RAS mode配置相应的BIOS参数。
需要说明的是,对于方法实施例,为了简单描述,故将其都表述为一系列的动作组合,但是本领域技术人员应该知悉,本申请实施例并不受所描述的动作顺序的限制,因为依据本申请实施例,某些步骤可以采用其他顺序或者同时进行。其次,本领域技术人员也应该知悉,说明书中所描述的实施例均属于优选实施例,所涉及的动作并不一定是本申请实施例所必须的。
参照图4,示出了本申请的一种服务器可维护性配置装置实施例的结构框图,服务器可维护性配置装置具体可以包括如下模块:
第一计算模块401,用于响应于服务器的正常启动运行,计算中央处理器的第一利用率;
重启模块402,用于响应于服务器的宕机重启,确定故障部件;
第二计算模块403,用于计算中央处理器的第二利用率;
业务迁移判断模块404,用于基于第一利用率和第二利用率,确定业务迁移状态;
切换模块405,用于依据业务迁移状态,切换服务器配置模式;
隔离模块406,用于在服务器配置模式中,隔离故障部件。
在本申请的一些可选实施例中,第一计算模块401包括:
第一读取子模块,用于读取中央处理器的功耗数据和单位热量数据;
第一利用率确定子模块,用于依据功耗数据和单位热量数据,确定第一利用率。
在本申请的一些可选实施例中,第一利用率确定子模块包括:
第一计算单元,用于计算功耗数据与单位热量数据的第一比值;
第一利用率确定单元,用于确定第一比值为第一利用率。
在本申请的一些可选实施例中,重启模块402包括:
重启子模块,用于响应于服务器的宕机重启,读取错误信息;
故障确定子模块,用于确定错误信息对应的部件为故障部件。
在本申请的一些可选实施例中,重启模块402还包括:
等待子模块,用于等待预设时长,进入服务器的基本输入输出系统。
在本申请的一些可选实施例中,第二计算模块403包括:
第二读取子模块,用于读取中央处理器的功耗数据和单位热量数据;
第二利用率确定子模块,用于依据功耗数据和单位热量数据,确定第二利用率。
在本申请的一些可选实施例中,第二利用率确定子模块包括:
第二计算单元,用于计算功耗数据与单位热量数据的第二比值;
第二利用率确定单元,用于确定第二比值为第二利用率。
在本申请的一些可选实施例中,业务迁移判断模块404包括:
业务波动值确定子模块,用于基于第一利用率和第二利用率,计算业务波动值;
业务迁移状态确定子模块,用于基于业务波动值确定业务迁移状态。
在本申请的一些可选实施例中,业务波动值确定子模块包括:
差值计算单元,用于计算第一利用率和第二利用率的差值;
第三比值计算单元,用于计算差值与第一利用率的第三比值;
业务波动值确定对应,用于确定第三比值为业务波动值。
在本申请的一些可选实施例中,业务迁移状态确定子模块包括:
判断单元,用于判断业务波动值是否小于预设业务波动阈值;
第一迁移确定单元,用于响应于业务波动值小于预设业务波动阈值,确定业务迁移状态为业务未迁移;
第二迁移确定单元,用于响应于业务波动值不小于预设业务波动阈值,确定业务迁移状态为业务已迁移。
在本申请的一些可选实施例中,服务器配置模式包括可靠性模式和可运行性模式,可靠性模式的运行可靠性大于可运行性模式的运行可靠性,可运行性模式的运行效率大于可靠性模式的运行效率;切换模块405包括:
第一切换子模块,用于响应于业务迁移状态为业务未迁移,切换服务器配置模式为可靠性模式;
第二切换子模块,用于响应于业务迁移状态为业务已迁移,切换服务器配置模式为可运行性模式。
在本申请的一些可选实施例中,第一切换子模块包括:
第一标记单元,用于响应于业务迁移状态为业务未迁移,设置服务器的模式标记为可靠性模式对应的可靠性标记,控制服务器重启;
第一配置单元,用于在服务器重启期间,基于可靠性标记,配置服务器的基本输入输出系统选项,以切换为可靠性模式。
在本申请的一些可选实施例中,第二切换子模块包括:
第二标记单元,用于响应于业务迁移状态为业务已迁移,设置服务器的模式标记为可运行性模式对应的可运行性标记,控制服务器重启;
第二配置单元,用于在服务器重启期间,基于可运行性标记,配置服务器的基本输入输出系统选项,以切换为可运行性模式。
在本申请的一些可选实施例中,服务器配置模式还包括平衡模式和自动模式,平衡模式的运行效率位于可靠性模式和可运行性模式之间,平衡模式的运行可靠性位于可靠性模式和可运行性模式之间;自动模式复用可运行性模式和可靠性模式和平衡模式中的一个。
在本申请的一些可选实施例中,装置还包括:
显示模块,用于响应于服务器的正常启动运行,显示模式选择页面。
在本申请的一些可选实施例中,装置还包括:
选择模块,用于接收针对模式选择页面的选择操作,选择可靠性模式、可运行性模式、平衡模式和自动模式中的一个为当前配置模式。
在本申请的一些可选实施例中,预设业务波动阈值为30%。
本申请实施例通过响应于服务器的正常启动运行,计算中央处理器的第一利用率;响应于服务器的宕机重启,确定故障部件;计算中央处理器的第二利用率;基于第一利用率和第二利用率,确定业务迁移状态;依据业务迁移状态,切换服务器配置模式;在服务器配置模式中,隔离故障部件。通过在正常启动和重启时的中央处理器的利用率判断客户的业务是否迁移,根据是否业务迁移来启动不同服务器配置模式,以使服务器可以自动切换配置模型,降低服务器的宕机率。
对于装置实施例而言,由于其与方法实施例基本相似,所以描述的比较简单,相关之处参见方法实施例的部分说明即可。
参照图5,本申请实施例还提供了一种电子设备,包括:
处理器501和存储介质502,存储介质502存储有处理器501可执行的计算机程序,当电子设备运行时,处理器501执行计算机程序,以执行如本申请实施例任一项的服务器可维护性配置方法。服务器可维护性配置方法包括:
响应于服务器的正常启动运行,计算中央处理器的第一利用率;
响应于服务器的宕机重启,确定故障部件;
计算中央处理器的第二利用率;
基于第一利用率和第二利用率,确定业务迁移状态;
依据业务迁移状态,切换服务器配置模式;
在服务器配置模式中,隔离故障部件。
在本申请的一些实施例中,计算中央处理器的第一利用率的步骤包括:
读取中央处理器的功耗数据和单位热量数据;
依据功耗数据和单位热量数据,确定第一利用率。
在本申请的一些实施例中,依据功耗数据和单位热量数据,确定第一利用率的步骤包括:
计算功耗数据与单位热量数据的第一比值;
确定第一比值为第一利用率。
在本申请的一些实施例中,响应于服务器的宕机重启,确定故障部件的步骤包括:
响应于服务器的宕机重启,读取错误信息;
确定错误信息对应的部件为故障部件。
在本申请的一些实施例中,在读取错误信息的步骤之前,响应于服务器的宕机重启,确定故障部件的步骤还包括:
等待预设时长,进入服务器的基本输入输出系统。
在本申请的一些实施例中,计算中央处理器的第二利用率的步骤包括:
读取中央处理器的功耗数据和单位热量数据;
依据功耗数据和单位热量数据,确定第二利用率。
在本申请的一些实施例中,依据功耗数据和单位热量数据,确定第一利用率的步骤包
括:
计算功耗数据与单位热量数据的第二比值;
确定第二比值为第二利用率。
在本申请的一些实施例中,基于第一利用率和第二利用率,确定业务迁移状态的步骤包括:
基于第一利用率和第二利用率,计算业务波动值;
基于业务波动值确定业务迁移状态。
在本申请的一些实施例中,基于第一利用率和第二利用率,计算业务波动值的步骤包括:
计算第一利用率和第二利用率的差值;
计算差值与第一利用率的第三比值;
确定第三比值为业务波动值。
在本申请的一些实施例中,基于业务波动值确定业务迁移状态的步骤包括:
判断业务波动值是否小于预设业务波动阈值;
响应于业务波动值小于预设业务波动阈值,确定业务迁移状态为业务未迁移;
响应于业务波动值不小于预设业务波动阈值,确定业务迁移状态为业务已迁移。
在本申请的一些实施例中,服务器配置模式包括可靠性模式和可运行性模式,可靠性模式的运行可靠性大于可运行性模式的运行可靠性,可运行性模式的运行效率大于可靠性模式的运行效率;依据业务迁移状态,切换服务器配置模式的步骤包括:
响应于业务迁移状态为业务未迁移,切换服务器配置模式为可靠性模式;
响应于业务迁移状态为业务已迁移,切换服务器配置模式为可运行性模式。
在本申请的一些实施例中,响应于业务迁移状态为业务未迁移,切换服务器配置模式为可靠性模式的步骤包括:
响应于业务迁移状态为业务未迁移,设置服务器的模式标记为可靠性模式对应的可靠性标记,控制服务器重启;
在服务器重启期间,基于可靠性标记,配置服务器的基本输入输出系统选项,以切换为可靠性模式。
在本申请的一些实施例中,响应于业务迁移状态为业务已迁移,切换服务器配置模式为可运行性模式的步骤包括:
响应于业务迁移状态为业务已迁移,设置服务器的模式标记为可运行性模式对应的可运行性标记,控制服务器重启;
在服务器重启期间,基于可运行性标记,配置服务器的基本输入输出系统选项,以切换为可运行性模式。
在本申请的一些实施例中,服务器配置模式还包括平衡模式和自动模式,平衡模式的运行效率位于可靠性模式和可运行性模式之间,平衡模式的运行可靠性位于可靠性模式和可运行性模式之间;自动模式复用可运行性模式和可靠性模式和平衡模式中的一个。
在本申请的一些实施例中,方法还包括:
响应于服务器的正常启动运行,显示模式选择页面。
在本申请的一些实施例中,方法还包括:
接收针对模式选择页面的选择操作,选择可靠性模式、可运行性模式、平衡模式和自动模式中的一个为当前配置模式。
在本申请的一些实施例中,预设业务波动阈值为30%。
其中,存储器可以包括随机存取存储器(Random Access Memory,简称RAM),也可以包括非易失性存储器(non-volatile memory),例如至少一个磁盘存储器。可选的,存储器还可以是至少一个位于远离前述处理器的存储装置。
上述的处理器可以是通用处理器,包括中央处理器(Central Processing Unit,简称CPU)、网络处理器(Network Processor,简称NP)等;还可以是数字信号处理器(Digital Signal Processing,简称DSP)、专用集成电路(Application Specific Integrated Circuit,简称ASIC)、现场可编程门阵列(Field-Programmable Gate Array,简称FPGA)或者其他可编程逻辑器件、分立门或者晶体管逻辑器件、分立硬件组件。
本申请实施例通过响应于服务器的正常启动运行,计算中央处理器的第一利用率;响应于服务器的宕机重启,确定故障部件;计算中央处理器的第二利用率;基于第一利用率和第二利用率,确定业务迁移状态;依据业务迁移状态,切换服务器配置模式;在服务器配置模式中,隔离故障部件。通过在正常启动和重启时的中央处理器的利用率判断客户的业务是否迁移,根据是否业务迁移来启动不同服务器配置模式,以使服务器可以自动切换配置模型,降低服务器的宕机率。
参照图6,本申请实施例还提供了一种非易失性可读存储介质601,非易失性可读存储介质601上存储有计算机程序,计算机程序被处理器运行时执行如本申请实施例任一项的服务器可维护性配置方法。服务器可维护性配置方法包括:
响应于服务器的正常启动运行,计算中央处理器的第一利用率;
响应于服务器的宕机重启,确定故障部件;
计算中央处理器的第二利用率;
基于第一利用率和第二利用率,确定业务迁移状态;
依据业务迁移状态,切换服务器配置模式;
在服务器配置模式中,隔离故障部件。
在本申请的一些实施例中,计算中央处理器的第一利用率的步骤包括:
读取中央处理器的功耗数据和单位热量数据;
依据功耗数据和单位热量数据,确定第一利用率。
在本申请的一些实施例中,依据功耗数据和单位热量数据,确定第一利用率的步骤包括:
计算功耗数据与单位热量数据的第一比值;
确定第一比值为第一利用率。
在本申请的一些实施例中,响应于服务器的宕机重启,确定故障部件的步骤包括:
响应于服务器的宕机重启,读取错误信息;
确定错误信息对应的部件为故障部件。
在本申请的一些实施例中,在读取错误信息的步骤之前,响应于服务器的宕机重启,确定故障部件的步骤还包括:
等待预设时长,进入服务器的基本输入输出系统。
在本申请的一些实施例中,计算中央处理器的第二利用率的步骤包括:
读取中央处理器的功耗数据和单位热量数据;
依据功耗数据和单位热量数据,确定第二利用率。
在本申请的一些实施例中,依据功耗数据和单位热量数据,确定第一利用率的步骤包括:
计算功耗数据与单位热量数据的第二比值;
确定第二比值为第二利用率。
在本申请的一些实施例中,基于第一利用率和第二利用率,确定业务迁移状态的步骤包括:
基于第一利用率和第二利用率,计算业务波动值;
基于业务波动值确定业务迁移状态。
在本申请的一些实施例中,基于第一利用率和第二利用率,计算业务波动值的步骤包括:
计算第一利用率和第二利用率的差值;
计算差值与第一利用率的第三比值;
确定第三比值为业务波动值。
在本申请的一些实施例中,基于业务波动值确定业务迁移状态的步骤包括:
判断业务波动值是否小于预设业务波动阈值;
响应于业务波动值小于预设业务波动阈值,确定业务迁移状态为业务未迁移;
响应于业务波动值不小于预设业务波动阈值,确定业务迁移状态为业务已迁移。
在本申请的一些实施例中,服务器配置模式包括可靠性模式和可运行性模式,可靠性模式的运行可靠性大于可运行性模式的运行可靠性,可运行性模式的运行效率大于可靠性模式的运行效率;依据业务迁移状态,切换服务器配置模式的步骤包括:
响应于业务迁移状态为业务未迁移,切换服务器配置模式为可靠性模式;
响应于业务迁移状态为业务已迁移,切换服务器配置模式为可运行性模式。
在本申请的一些实施例中,响应于业务迁移状态为业务未迁移,切换服务器配置模式为可靠性模式的步骤包括:
响应于业务迁移状态为业务未迁移,设置服务器的模式标记为可靠性模式对应的可靠性标记,控制服务器重启;
在服务器重启期间,基于可靠性标记,配置服务器的基本输入输出系统选项,以切换为可靠性模式。
在本申请的一些实施例中,响应于业务迁移状态为业务已迁移,切换服务器配置模式为可运行性模式的步骤包括:
响应于业务迁移状态为业务已迁移,设置服务器的模式标记为可运行性模式对应的可运行性标记,控制服务器重启;
在服务器重启期间,基于可运行性标记,配置服务器的基本输入输出系统选项,以切换为可运行性模式。
在本申请的一些实施例中,服务器配置模式还包括平衡模式和自动模式,平衡模式的运行效率位于可靠性模式和可运行性模式之间,平衡模式的运行可靠性位于可靠性模式和可运
行性模式之间;自动模式复用可运行性模式和可靠性模式和平衡模式中的一个。
在本申请的一些实施例中,方法还包括:
响应于服务器的正常启动运行,显示模式选择页面。
在本申请的一些实施例中,方法还包括:
接收针对模式选择页面的选择操作,选择可靠性模式、可运行性模式、平衡模式和自动模式中的一个为当前配置模式。
在本申请的一些实施例中,预设业务波动阈值为30%。
本申请实施例通过响应于服务器的正常启动运行,计算中央处理器的第一利用率;响应于服务器的宕机重启,确定故障部件;计算中央处理器的第二利用率;基于第一利用率和第二利用率,确定业务迁移状态;依据业务迁移状态,切换服务器配置模式;在服务器配置模式中,隔离故障部件。通过在正常启动和重启时的中央处理器的利用率判断客户的业务是否迁移,根据是否业务迁移来启动不同服务器配置模式,以使服务器可以自动切换配置模型,降低服务器的宕机率。
本说明书中的各个实施例均采用递进的方式描述,每个实施例重点说明的都是与其他实施例的不同之处,各个实施例之间相同相似的部分互相参见即可。
本领域内的技术人员应明白,本申请实施例的实施例可提供为方法、装置、或计算机程序产品。因此,本申请实施例可采用完全硬件实施例、完全软件实施例、或结合软件和硬件方面的实施例的形式。而且,本申请实施例可采用在一个或多个其中包含有计算机可用程序代码的非易失性可读存储介质(包括但不限于磁盘存储器、CD-ROM(Compact Disc Read-Only Memory,只读光盘存储介质)、光学存储器等)上实施的计算机程序产品的形式。
本申请实施例是参照根据本申请实施例的方法、终端设备(系统)、和计算机程序产品的流程图和/或方框图来描述的。应理解可由计算机程序指令实现流程图和/或方框图中的每一流程和/或方框、以及流程图和/或方框图中的流程和/或方框的结合。可提供这些计算机程序指令到通用计算机、专用计算机、嵌入式处理机或其他可编程数据处理终端设备的处理器以产生一个机器,使得通过计算机或其他可编程数据处理终端设备的处理器执行的指令产生用于实现在流程图一个流程或多个流程和/或方框图一个方框或多个方框中指定的功能的装置。
这些计算机程序指令也可存储在能引导计算机或其他可编程数据处理终端设备以特定方式工作的非易失性可读存储介质中,使得存储在该非易失性可读存储介质中的指令产生包括指令装置的制造品,该指令装置实现在流程图一个流程或多个流程和/或方框图一个方框或多个方框中指定的功能。
这些计算机程序指令也可装载到计算机或其他可编程数据处理终端设备上,使得在计算机或其他可编程终端设备上执行一系列操作步骤以产生计算机实现的处理,从而在计算机或其他可编程终端设备上执行的指令提供用于实现在流程图一个流程或多个流程和/或方框图一个方框或多个方框中指定的功能的步骤。
尽管已描述了本申请实施例的优选实施例,但本领域内的技术人员一旦得知了基本创造性概念,则可对这些实施例做出另外的变更和修改。所以,所附权利要求意欲解释为包括优选实施例以及落入本申请实施例范围的所有变更和修改。
最后,还需要说明的是,在本文中,诸如第一和第二等之类的关系术语仅仅用来将一个
实体或者操作与另一个实体或操作区分开来,而不一定要求或者暗示这些实体或操作之间存在任何这种实际的关系或者顺序。而且,术语“包括”、“包含”或者其任何其他变体意在涵盖非排他性的包含,从而使得包括一系列要素的过程、方法、物品或者终端设备不仅包括那些要素,而且还包括没有明确列出的其他要素,或者是还包括为这种过程、方法、物品或者终端设备所固有的要素。在没有更多限制的情况下,由语句“包括一个……”限定的要素,并不排除在包括要素的过程、方法、物品或者终端设备中还存在另外的相同要素。
以上对本申请所提供的服务器可维护性配置方法、装置、电子设备和非易失性可读存储介质,进行了详细介绍,本文中应用了具体个例对本申请的原理及实施方式进行了阐述,以上实施例的说明只是用于帮助理解本申请的方法及其核心思想;同时,对于本领域的一般技术人员,依据本申请的思想,在具体实施方式及应用范围上均会有改变之处,综上,本说明书内容不应理解为对本申请的限制。
Claims (20)
- 一种服务器可维护性配置方法,其特征在于,包括:响应于服务器的正常启动运行,计算中央处理器的第一利用率;响应于所述服务器的宕机重启,确定故障部件;计算所述中央处理器的第二利用率;基于所述第一利用率和所述第二利用率,确定业务迁移状态;依据所述业务迁移状态,切换服务器配置模式;在所述服务器配置模式中,隔离所述故障部件。
- 根据权利要求1所述的方法,其特征在于,所述计算中央处理器的第一利用率的步骤包括:读取所述中央处理器的功耗数据和单位热量数据;依据所述功耗数据和所述单位热量数据,确定所述第一利用率。
- 根据权利要求2所述的方法,其特征在于,所述依据所述功耗数据和所述单位热量数据,确定所述第一利用率的步骤包括:计算所述功耗数据与所述单位热量数据的第一比值;确定所述第一比值为所述第一利用率。
- 根据权利要求1所述的方法,其特征在于,所述响应于所述服务器的宕机重启,确定故障部件的步骤包括:响应于所述服务器的宕机重启,读取错误信息;确定所述错误信息对应的部件为所述故障部件。
- 根据权利要求4所述的方法,其特征在于,在所述读取错误信息的步骤之前,所述响应于所述服务器的宕机重启,确定故障部件的步骤还包括:等待预设时长,进入所述服务器的基本输入输出系统。
- 根据权利要求1所述的方法,其特征在于,所述计算所述中央处理器的第二利用率的步骤包括:读取所述中央处理器的功耗数据和单位热量数据;依据所述功耗数据和所述单位热量数据,确定所述第二利用率。
- 根据权利要求6所述的方法,其特征在于,所述依据所述功耗数据和所述单位热量数据,确定所述第二利用率的步骤包括:计算所述功耗数据与所述单位热量数据的第二比值;确定所述第二比值为所述第二利用率。
- 根据权利要求1所述的方法,其特征在于,所述基于所述第一利用率和所述第二利用率,确定业务迁移状态的步骤包括:基于所述第一利用率和所述第二利用率,计算业务波动值;基于所述业务波动值确定业务迁移状态。
- 根据权利要求8所述的方法,其特征在于,所述基于所述第一利用率和所述第二利用率,计算业务波动值的步骤包括:计算所述第一利用率和所述第二利用率的差值;计算所述差值与所述第一利用率的第三比值;确定所述第三比值为所述业务波动值。
- 根据权利要求8所述的方法,其特征在于,所述基于所述业务波动值确定业务迁移状态的步骤包括:判断所述业务波动值是否小于预设业务波动阈值;响应于所述业务波动值小于预设业务波动阈值,确定所述业务迁移状态为业务未迁移;响应于所述业务波动值不小于预设业务波动阈值,确定所述业务迁移状态为业务已迁移。
- 根据权利要求10所述的方法,其特征在于,所述服务器配置模式包括可靠性模式和可运行性模式,所述可靠性模式的运行可靠性大于所述可运行性模式的运行可靠性,所述可运行性模式的运行效率大于所述可靠性模式的运行效率;所述依据所述业务迁移状态,切换服务器配置模式的步骤包括:响应于所述业务迁移状态为所述业务未迁移,切换所述服务器配置模式为所述可靠性模式;响应于所述业务迁移状态为所述业务已迁移,切换所述服务器配置模式为所述可运行性模式。
- 根据权利要求11所述的方法,其特征在于,所述响应于所述业务迁移状态为所述业务未迁移,切换所述服务器配置模式为所述可靠性模式的步骤包括:响应于所述业务迁移状态为所述业务未迁移,设置所述服务器的模式标记为所述可靠性模式对应的可靠性标记,控制所述服务器重启;在所述服务器重启期间,基于所述可靠性标记,配置所述服务器的基本输入输出系统选项,以切换为所述可靠性模式。
- 根据权利要求11所述的方法,其特征在于,所述响应于所述业务迁移状态为所述业务已迁移,切换所述服务器配置模式为所述可运行性模式的步骤包括:响应于所述业务迁移状态为所述业务已迁移,设置所述服务器的模式标记为所述可运行性模式对应的可运行性标记,控制所述服务器重启;在所述服务器重启期间,基于所述可运行性标记,配置所述服务器的基本输入输出系统选项,以切换为所述可运行性模式。
- 根据权利要求11所述的方法,其特征在于,所述服务器配置模式还包括平衡模式和自动模式,所述平衡模式的运行效率位于所述可靠性模式和所述可运行性模式之间,所述平衡模式的运行可靠性位于所述可靠性模式和所述可运行性模式之间;所述自动模式复用所述可运行性模式和所述可靠性模式和所述平衡模式中的一个。
- 根据权利要求14所述的方法,其特征在于,所述方法还包括:响应于服务器的正常启动运行,显示模式选择页面。
- 根据权利要求15所述的方法,其特征在于,所述方法还包括:接收针对所述模式选择页面的选择操作,选择所述可靠性模式、所述可运行性模式、所述平衡模式和所述自动模式中的一个为当前配置模式。
- 根据权利要求10所述的方法,其特征在于,所述预设业务波动阈值为30%。
- 一种服务器可维护性配置装置,其特征在于,包括:第一计算模块,用于响应于服务器的正常启动运行,计算中央处理器的第一利用 率;重启模块,用于响应于所述服务器的宕机重启,确定故障部件;第二计算模块,用于计算所述中央处理器的第二利用率;业务迁移判断模块,用于基于所述第一利用率和所述第二利用率,确定业务迁移状态;切换模块,用于依据所述业务迁移状态,切换服务器配置模式;隔离模块,用于在所述服务器配置模式中,隔离所述故障部件。
- 一种电子设备,其特征在于,包括处理器、存储器及存储在所述存储器上并能够在所述处理器上运行的计算机程序,所述计算机程序被所述处理器执行时实现如权利要求1至17中任一项所述的服务器可维护性配置方法的步骤。
- 一种非易失性可读存储介质,其特征在于,所述非易失性可读存储介质上存储计算机程序,所述计算机程序被处理器执行时实现如权利要求1至17中任一项所述的服务器可维护性配置方法的步骤。
Priority Applications (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| US19/094,727 US20250225023A1 (en) | 2023-08-14 | 2025-03-28 | Server maintainability configuration method and apparatus, electronic device and storage medium |
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN202311019304.8A CN116737396B (zh) | 2023-08-14 | 2023-08-14 | 服务器可维护性配置方法、装置、电子设备和存储介质 |
| CN202311019304.8 | 2023-08-14 |
Related Child Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| US19/094,727 Continuation US20250225023A1 (en) | 2023-08-14 | 2025-03-28 | Server maintainability configuration method and apparatus, electronic device and storage medium |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2025035961A1 true WO2025035961A1 (zh) | 2025-02-20 |
Family
ID=87904757
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/CN2024/100740 Pending WO2025035961A1 (zh) | 2023-08-14 | 2024-06-21 | 服务器可维护性配置方法、装置、电子设备和存储介质 |
Country Status (3)
| Country | Link |
|---|---|
| US (1) | US20250225023A1 (zh) |
| CN (1) | CN116737396B (zh) |
| WO (1) | WO2025035961A1 (zh) |
Families Citing this family (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN116737396B (zh) * | 2023-08-14 | 2023-11-03 | 苏州浪潮智能科技有限公司 | 服务器可维护性配置方法、装置、电子设备和存储介质 |
| CN119127467B (zh) * | 2024-08-15 | 2025-03-28 | 快际新云(青岛)科技有限公司 | 一种多租户环境下的模型训练资源隔离与管理系统及方法 |
Citations (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN106802854A (zh) * | 2017-02-22 | 2017-06-06 | 郑州云海信息技术有限公司 | 一种多控制器系统的故障监控系统 |
| CN109947596A (zh) * | 2019-03-19 | 2019-06-28 | 浪潮商用机器有限公司 | Pcie设备故障系统宕机处理方法、装置及相关组件 |
| US11182253B2 (en) * | 2016-06-09 | 2021-11-23 | Intuit Inc. | Self-healing system for distributed services and applications |
| CN116048972A (zh) * | 2022-12-29 | 2023-05-02 | 苏州浪潮智能科技有限公司 | 一种服务器ras测试方法、系统、设备以及存储介质 |
| CN116737396A (zh) * | 2023-08-14 | 2023-09-12 | 苏州浪潮智能科技有限公司 | 服务器可维护性配置方法、装置、电子设备和存储介质 |
Family Cites Families (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN107122230A (zh) * | 2017-05-31 | 2017-09-01 | 郑州云海信息技术有限公司 | 一种基于服务器集群的高可用方法及其装置 |
| CN110046064B (zh) * | 2018-01-15 | 2020-08-04 | 厦门靠谱云股份有限公司 | 一种基于故障漂移的云服务器容灾实现方法 |
-
2023
- 2023-08-14 CN CN202311019304.8A patent/CN116737396B/zh active Active
-
2024
- 2024-06-21 WO PCT/CN2024/100740 patent/WO2025035961A1/zh active Pending
-
2025
- 2025-03-28 US US19/094,727 patent/US20250225023A1/en active Pending
Patent Citations (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US11182253B2 (en) * | 2016-06-09 | 2021-11-23 | Intuit Inc. | Self-healing system for distributed services and applications |
| CN106802854A (zh) * | 2017-02-22 | 2017-06-06 | 郑州云海信息技术有限公司 | 一种多控制器系统的故障监控系统 |
| CN109947596A (zh) * | 2019-03-19 | 2019-06-28 | 浪潮商用机器有限公司 | Pcie设备故障系统宕机处理方法、装置及相关组件 |
| CN116048972A (zh) * | 2022-12-29 | 2023-05-02 | 苏州浪潮智能科技有限公司 | 一种服务器ras测试方法、系统、设备以及存储介质 |
| CN116737396A (zh) * | 2023-08-14 | 2023-09-12 | 苏州浪潮智能科技有限公司 | 服务器可维护性配置方法、装置、电子设备和存储介质 |
Also Published As
| Publication number | Publication date |
|---|---|
| CN116737396B (zh) | 2023-11-03 |
| CN116737396A (zh) | 2023-09-12 |
| US20250225023A1 (en) | 2025-07-10 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| WO2025035961A1 (zh) | 服务器可维护性配置方法、装置、电子设备和存储介质 | |
| WO2022198972A1 (zh) | 一种服务器启动过程中的故障定位方法、系统及装置 | |
| US9495233B2 (en) | Error framework for a microprocesor and system | |
| CN115129520B (zh) | 计算机系统、计算机服务器及其启动方法 | |
| EP2175371A1 (en) | Synchronization control apparatuses, information processing apparatuses, and synchronization management methods | |
| US20160277271A1 (en) | Fault tolerant method and system for multiple servers | |
| CN103514068A (zh) | 内存故障自动定位方法 | |
| CN112199240A (zh) | 一种节点故障时进行节点切换的方法及相关设备 | |
| JP2018180982A (ja) | 情報処理装置、およびログ記録方法 | |
| WO2012119432A1 (zh) | 提高计算机系统稳定性的方法及计算机系统 | |
| CN115686951A (zh) | 一种数据库服务器的故障处理方法和装置 | |
| EP4457629A1 (en) | Systems and methods to initiate device recovery | |
| CN113608603A (zh) | 一种修复PCIe故障设备的方法、系统、设备和存储介质 | |
| CN117033115A (zh) | 故障处理方法、装置、设备及存储介质 | |
| CN120086053A (zh) | 服务器及处理器错误处理方法、设备、介质及程序产品 | |
| CN114356708A (zh) | 一种设备故障监控方法、装置、设备及可读存储介质 | |
| CN121478572A (zh) | 一种服务器启动控制系统和方法 | |
| CN111984471A (zh) | 一种机柜电源bmc冗余管理系统及方法 | |
| WO2015135100A1 (zh) | 一种实现处理器切换的方法、计算机和切换装置 | |
| CN104199747B (zh) | 基于健康管理的高可用系统实现方法及系统 | |
| CN113312198B (zh) | 监控及复原异质性元件的系统及方法 | |
| CN101206599B (zh) | 计算机主板设备诊断和隔离方法 | |
| JP5335150B2 (ja) | 計算機装置及びプログラム | |
| CN117312037A (zh) | 内存修复方法、装置、电子设备及存储介质 | |
| CN115129497A (zh) | 一种服务器记录内存故障的方法 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 24853376 Country of ref document: EP Kind code of ref document: A1 |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |