WO2022252559A1 - 基于规则和双深度q网络的混合动力汽车能量管理方法 - Google Patents
基于规则和双深度q网络的混合动力汽车能量管理方法 Download PDFInfo
- Publication number
- WO2022252559A1 WO2022252559A1 PCT/CN2021/137803 CN2021137803W WO2022252559A1 WO 2022252559 A1 WO2022252559 A1 WO 2022252559A1 CN 2021137803 W CN2021137803 W CN 2021137803W WO 2022252559 A1 WO2022252559 A1 WO 2022252559A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- network
- lithium battery
- value
- soc
- energy
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Images
Classifications
-
- B—PERFORMING OPERATIONS; TRANSPORTING
- B60—VEHICLES IN GENERAL
- B60L—PROPULSION OF ELECTRICALLY-PROPELLED VEHICLES; SUPPLYING ELECTRIC POWER FOR AUXILIARY EQUIPMENT OF ELECTRICALLY-PROPELLED VEHICLES; ELECTRODYNAMIC BRAKE SYSTEMS FOR VEHICLES IN GENERAL; MAGNETIC SUSPENSION OR LEVITATION FOR VEHICLES; MONITORING OPERATING VARIABLES OF ELECTRICALLY-PROPELLED VEHICLES; ELECTRIC SAFETY DEVICES FOR ELECTRICALLY-PROPELLED VEHICLES
- B60L50/00—Electric propulsion with power supplied within the vehicle
- B60L50/40—Electric propulsion with power supplied within the vehicle using propulsion power supplied by capacitors
-
- B—PERFORMING OPERATIONS; TRANSPORTING
- B60—VEHICLES IN GENERAL
- B60L—PROPULSION OF ELECTRICALLY-PROPELLED VEHICLES; SUPPLYING ELECTRIC POWER FOR AUXILIARY EQUIPMENT OF ELECTRICALLY-PROPELLED VEHICLES; ELECTRODYNAMIC BRAKE SYSTEMS FOR VEHICLES IN GENERAL; MAGNETIC SUSPENSION OR LEVITATION FOR VEHICLES; MONITORING OPERATING VARIABLES OF ELECTRICALLY-PROPELLED VEHICLES; ELECTRIC SAFETY DEVICES FOR ELECTRICALLY-PROPELLED VEHICLES
- B60L15/00—Methods, circuits, or devices for controlling the traction-motor speed of electrically-propelled vehicles
-
- B—PERFORMING OPERATIONS; TRANSPORTING
- B60—VEHICLES IN GENERAL
- B60L—PROPULSION OF ELECTRICALLY-PROPELLED VEHICLES; SUPPLYING ELECTRIC POWER FOR AUXILIARY EQUIPMENT OF ELECTRICALLY-PROPELLED VEHICLES; ELECTRODYNAMIC BRAKE SYSTEMS FOR VEHICLES IN GENERAL; MAGNETIC SUSPENSION OR LEVITATION FOR VEHICLES; MONITORING OPERATING VARIABLES OF ELECTRICALLY-PROPELLED VEHICLES; ELECTRIC SAFETY DEVICES FOR ELECTRICALLY-PROPELLED VEHICLES
- B60L50/00—Electric propulsion with power supplied within the vehicle
- B60L50/50—Electric propulsion with power supplied within the vehicle using propulsion power supplied by batteries or fuel cells
- B60L50/60—Electric propulsion with power supplied within the vehicle using propulsion power supplied by batteries or fuel cells using power supplied by batteries
-
- B—PERFORMING OPERATIONS; TRANSPORTING
- B60—VEHICLES IN GENERAL
- B60L—PROPULSION OF ELECTRICALLY-PROPELLED VEHICLES; SUPPLYING ELECTRIC POWER FOR AUXILIARY EQUIPMENT OF ELECTRICALLY-PROPELLED VEHICLES; ELECTRODYNAMIC BRAKE SYSTEMS FOR VEHICLES IN GENERAL; MAGNETIC SUSPENSION OR LEVITATION FOR VEHICLES; MONITORING OPERATING VARIABLES OF ELECTRICALLY-PROPELLED VEHICLES; ELECTRIC SAFETY DEVICES FOR ELECTRICALLY-PROPELLED VEHICLES
- B60L58/00—Methods or circuit arrangements for monitoring or controlling batteries or fuel cells, specially adapted for electric vehicles
- B60L58/10—Methods or circuit arrangements for monitoring or controlling batteries or fuel cells, specially adapted for electric vehicles for monitoring or controlling batteries
-
- Y—GENERAL TAGGING OF NEW TECHNOLOGICAL DEVELOPMENTS; GENERAL TAGGING OF CROSS-SECTIONAL TECHNOLOGIES SPANNING OVER SEVERAL SECTIONS OF THE IPC; TECHNICAL SUBJECTS COVERED BY FORMER USPC CROSS-REFERENCE ART COLLECTIONS [XRACs] AND DIGESTS
- Y02—TECHNOLOGIES OR APPLICATIONS FOR MITIGATION OR ADAPTATION AGAINST CLIMATE CHANGE
- Y02T—CLIMATE CHANGE MITIGATION TECHNOLOGIES RELATED TO TRANSPORTATION
- Y02T10/00—Road transport of goods or passengers
- Y02T10/10—Internal combustion engine [ICE] based vehicles
- Y02T10/40—Engine management systems
-
- Y—GENERAL TAGGING OF NEW TECHNOLOGICAL DEVELOPMENTS; GENERAL TAGGING OF CROSS-SECTIONAL TECHNOLOGIES SPANNING OVER SEVERAL SECTIONS OF THE IPC; TECHNICAL SUBJECTS COVERED BY FORMER USPC CROSS-REFERENCE ART COLLECTIONS [XRACs] AND DIGESTS
- Y02—TECHNOLOGIES OR APPLICATIONS FOR MITIGATION OR ADAPTATION AGAINST CLIMATE CHANGE
- Y02T—CLIMATE CHANGE MITIGATION TECHNOLOGIES RELATED TO TRANSPORTATION
- Y02T10/00—Road transport of goods or passengers
- Y02T10/60—Other road transportation technologies with climate change mitigation effect
- Y02T10/70—Energy storage systems for electromobility, e.g. batteries
Definitions
- the present invention relates to the technical field of vehicle energy management, and more specifically, relates to a hybrid electric vehicle energy management method based on rules and double-depth Q networks.
- New energy vehicles include hybrid vehicles, electric vehicles and fuel cell vehicles. Energy management is one of the control optimization problems that need to be solved urgently for new energy vehicles.
- the rule-based energy management methods include deterministic rule control, fuzzy logic control, etc., which have strong real-time performance, but it is difficult to achieve the optimal control effect; although the optimization-based energy management method can obtain better control effect, it requires The working conditions are predictable and the amount of calculation is large, so it is difficult to apply in real time.
- the overestimation of the Q value makes the control effect not good, and also has defects such as slow convergence speed and poor real-time performance.
- the real-time and working condition adaptability of the control method are high. , so it is necessary to design an energy management method that can not only meet the actual application requirements but also have excellent control effects.
- the current energy management method based on deep reinforcement learning is often within the preset constraints, such as lithium battery SOC (state of charge, current remaining power) is not higher than 0.9 and not lower than 0.3 to achieve control, however
- the relevant state quantities cannot always be kept within the constraint range, and if the lithium battery is frequently overcharged or overdischarged, its service life will rapidly decay, which will significantly reduce the mileage of new energy vehicles and increase the cost of use. Therefore, it is necessary to formulate an additional control method outside the constraint range to obtain a more comprehensive and stable control method, so as to save energy as much as possible while prolonging the service life of lithium batteries, thereby reducing the cost of using new energy vehicles, which is beneficial to Large-scale promotion of new energy vehicles.
- the purpose of the present invention is to overcome above-mentioned defective of prior art, provide a kind of hybrid electric vehicle energy management method based on rule and double depth Q network, is to realize optimality, real-time property and working condition adaptability by improving depth Q network
- the technical solution of the invention is to provide a hybrid electric vehicle energy management method based on rules and double deep Q network.
- the method includes the following steps:
- Detect vehicle energy sources equipped with a composite energy storage system which includes a lithium battery energy source and a supercapacitor energy source;
- the trained deep reinforcement learning model is used to determine the output power of the lithium battery; Set rules to protect lithium batteries;
- the deep reinforcement learning model its agent includes an evaluation Q network and a target Q network
- the state observations of the environment are the remaining power of the lithium battery, the remaining power of the supercapacitor, and the required power of the whole vehicle
- the output power of the lithium battery is taken as Output actions, and aim at minimizing the energy loss of the composite energy storage system, and set a corresponding reward function.
- the present invention has the advantages of prolonging the service life of lithium batteries as much as possible while saving energy, thereby effectively reducing the cost of using new energy vehicles, and taking into account the real-time and optimality of the control method.
- Fig. 1 is a power architecture of a composite energy storage system for an electric vehicle according to an embodiment of the present invention
- FIG. 2 is a schematic diagram of an energy management method based on a double-depth Q network according to an embodiment of the present invention
- Fig. 3 is a control logic schematic diagram of an energy management method based on rules and double-depth Q networks according to an embodiment of the present invention
- FIG. 4 is a schematic diagram of the process of a hybrid energy management method based on rules and double-depth Q networks according to an embodiment of the present invention
- Fig. 5 is a flow chart of a method for energy management of a hybrid electric vehicle based on rules and double deep Q-networks according to an embodiment of the present invention.
- vehicle models which include vehicle dynamics models, motor models, energy source models, and transmission system models.
- relevant vehicle models include vehicle dynamics models, motor models, energy source models, and transmission system models.
- related modeling software such as Matlab/Simulink is used to establish related models. It should be understood that, in addition to Matlab, other computing programs or tools can also be used to implement the present invention.
- Deep reinforcement learning consists of 3 elements: environment, agent and reward function.
- the environment is the corresponding vehicle model or no model
- the agent is the method to be trained to perform related control
- the reward function needs to be set according to the specific problem.
- the quality of the setting will affect the learning ability of the agent. The effect is the training and convergence process.
- we need to pay special attention to the problem of sparse rewards that is, a large number of execution actions are not rewarded.
- the attributes and dimensions of the agent's state observations and output actions need to be properly set, which will also greatly affect the learning and training effect.
- the design of the energy management method of DDQN involves the design of neural network, state observation, action amount and reward function, and the setting of some hyperparameters.
- the design of neural networks involves the selection and design of the number of network layers, number of neurons, connection methods, and activation functions.
- the DDQN algorithm separates the current action selection from the strategy evaluation, which in turn can reduce overly optimistic method estimates to improve the control effect. It can be said that it is an improvement on the DQN algorithm.
- careful settings are also required to obtain ideal convergence speed and control effect.
- the present invention will set actions by incorporating the optimal results of dynamic programming.
- the experience playback technique In deep reinforcement learning, in order to improve learning efficiency, the experience playback technique will be introduced, which stores an array of past experience samples. During learning and training, the agent uses random sampling (reducing the correlation between samples) from the past experience. to study. However, in actual training, many samples get low rewards or even no rewards, so the learning reference value of these samples is not great.
- a priority is proposed based on the experience playback technique.
- the skill of experience playback is to learn by preferentially extracting samples with greater value (that is, to obtain greater rewards), so as to achieve faster convergence.
- the embodiment of the present invention introduces the skill of priority experience playback to accelerate the convergence of the DDQN algorithm and enhance The real-time nature of the method.
- Figure 1 is the power transmission system architecture of the composite energy storage system of electric vehicles.
- Energy sources both of which can be used as energy sources for charging and discharging, that is, to drive vehicles or absorb braking energy
- supercapacitors are used as auxiliary energy sources to reduce the charging and discharging frequency and discharge current of lithium batteries.
- the cooperative work of lithium batteries and supercapacitors needs to formulate appropriate energy management methods to reasonably distribute the input and output power of the two.
- the DCDC converter is responsible for stepping up the output voltage of the supercapacitor or stepping down the voltage from the bus, thereby reducing the requirement for supercapacitor configuration.
- the vehicle model is established in related modeling software such as Matlab/Simulink according to the quasi-static principle and vehicle dynamics.
- These models include: demand power calculation model, lithium battery equivalent circuit Model, supercapacitor equivalent circuit model, DCDC converter efficiency model, motor inverter assembly efficiency model, transmission model, etc.
- the vehicle model can be built after the simulation is correct.
- the hybrid electric vehicle energy management method based on rules and double-depth Q network uses the energy management method based on DDQN within the normal lithium battery SOC constraint range, and uses the rule-based energy management method to control.
- the following will specifically introduce the energy management method based on DDQN, the energy management method based on rules and double deep Q network.
- the interaction process between the environment and the agent is: each time step, in the environment state s t , the agent randomly selects an action from the set action set to output to the environment, the state of the environment Then it changes from st to st+1 .
- the agent randomly selects an action from the set action set to output to the environment, the state of the environment Then it changes from st to st+1 .
- the corresponding reward of the action is fed back to the agent.
- the agent constantly adjusts the selected output action to achieve The maximum cumulative reward is obtained, and the process will be repeated until the reward function converges.
- the DDQN algorithm separates the action selection from the strategy evaluation to avoid falling into an overestimation of the Q value and affecting the convergence speed and control effect.
- the present invention also adds the technique of priority experience playback to accelerate the training and learning process, that is, to increase the frequency of important and valuable samples being extracted, which can effectively promote algorithm training and further shorten the convergence time, which is beneficial to the practicality of the algorithm. promotion and application.
- the designed energy management method based on DDQN includes 2 parts, one part is a vehicle model built in Simulink and a closed-loop model built in conjunction with the Simulink reinforcement learning toolbox (reinforcement learning toolbox) RL agent module
- the agent-environment (vehicle) model the other part is aimed at designing neural networks and training instructions based on the DDQN energy management method, for example, establishing the m file of Matlab to realize.
- the design process of the DDQN-based energy management method is as follows.
- S represents the set of state observations
- s(t) represents the state observations at time t.
- the vehicle demand power P dem is normalized, that is, the demand power P dem is reduced to [-1,1], and the arithmetic mean mean(x ) and standard deviation std(x), calculated according to the standard normalized general formula, expressed as:
- mean(x) and std(x) represent the arithmetic mean and standard deviation of the input state data, respectively.
- the arithmetic mean mean(x) and standard deviation std(x) are calculated as follows:
- the feasible action output interval of the agent output action is set based on the global optimal control result based on dynamic programming, and the lithium battery output Power P batt is defined as output action. Since the DDQN algorithm needs to discretize the output action, that is to divide a feasible action interval into n parts.
- the control result of dynamic programming has the advantage of global optimality, and the result is also a discrete optimal control action sequence, so refer to the energy management method based on dynamic programming with the same optimization objective and control object
- the output action interval is appropriately set to n parts, that is, the action interval is P batt_max and P batt_min represent the maximum and minimum output power of the lithium battery, respectively.
- A represents the set of output actions
- a(t) represents the output action at time t.
- the energy loss of the composite energy storage system is minimized as the optimization goal, starting from this, the corresponding reward function r(t) is set, and the reward function is established through the corresponding mathematics in Simulink
- a model implementation for example represented as:
- E loss represents the overall energy loss of the composite energy storage system
- SOC b_tgt represents the target value of lithium battery SOC b
- SOC sc_tgt represents the target value of super capacitor SOC sc
- L sc , L batt , L dcdc represent super capacitor, lithium
- m represents the coefficient of the energy loss item
- n and p represent the coefficients of the SOC change of the balanced lithium battery and the super capacitor, respectively, and these three coefficients need to be adjusted during the training process.
- constraints are expressed as:
- I b is the current of the lithium battery
- I sc is the current of the super capacitor
- I b_min and I b_max represent the minimum and maximum value of the lithium battery current respectively
- I sc_min and I sc_max represent the minimum and maximum value of the super capacitor current respectively
- P dem represents the required power of the whole vehicle
- P min and P max represent the minimum and maximum value of the required power of the vehicle, respectively.
- the neural network of DDQN involves the design and selection of structure, number of layers, number of neurons and activation function, all of which need to be properly set and adjusted according to the actual training situation.
- the neural network of DDQN is designed by writing codes in the m-file of Matlab to call related functions of the neural network, and the connection between the layers of the neural network is carried out by using the full connection method.
- the activation function of the middle layer is set to the linear rectification function ReLU, for example, the activation function of the network output layer It is set to tanh, so that the output value is constrained to [-1,1].
- the parameters of the DDQN evaluation Q network as ⁇
- the parameters of the target Q network as ⁇ ′.
- the network parameters ⁇ of the evaluation Q network will be set at regular intervals
- the step size is copied to the target Q-network ⁇ ′.
- reinforcement learning and deep reinforcement learning are algorithms based on the Bellman principle.
- the calculation method of the Q value of the DDQN algorithm is expressed as:
- r t+1 represents the reward at time t+1
- ⁇ represents the discount factor, which is considered from the perspective of algorithm convergence
- s t , st t+1 represent the state at time t and t+1 respectively
- a t+1 represents the actions output at time t and t+1 respectively
- the update of the Q value is:
- ⁇ represents the learning rate, which will have a greater impact on the training speed and learning effect, and the symbolic meanings of other items are the same as formula (11).
- the loss function L( ⁇ ) is defined as the square of the difference between the evaluation network Q value and the target network Q′ value, and the training process of the DDQN algorithm is to minimize the loss L( ⁇ ) to a certain
- the process of a certain value is expressed as:
- the Q network parameters ⁇ are updated with gradient descent on the loss function L( ⁇ ):
- ⁇ -greedy The exploration and application of deep reinforcement learning is a problem that needs to be weighed. It is necessary to avoid too much exploration and too much application.
- a greedy algorithm ( ⁇ -greedy) is used to balance the exploration and application of the agent, that is, to randomly select the execution action with the probability of ⁇ , and to select the action corresponding to the maximum Q value in the current state with the probability of (1- ⁇ ) , it is necessary to define an appropriate initial value of ⁇ at the beginning of training, and the value of ⁇ to terminate the exploration.
- the initial value and termination value of ⁇ will greatly affect the training effect and convergence speed, so it needs to be carefully set and adjusted in relevant code files such as Matlab m files during training.
- the DDQN algorithm belongs to the off-policy deep reinforcement learning algorithm, that is, the generation of experience samples has nothing to do with the current strategy, so it can be considered to learn from past experience samples to improve learning efficiency. Generally, experience playback skills are introduced. At the same time, in order to reduce the correlation between samples, empirical samples are drawn by random sampling.
- the size of the experience playback pool is set to D in the m-file of Matlab, and the minimum number of batch samples is N*(st t , a t , r t , st t+1 ).
- the process of experience playback is as follows: first define the size of the experience pool as D, and store a sample array (s t , a t , r t , s t+1 ) in the experience pool at each time step, when a certain After a large number of samples, the agent randomly selects a small batch of sample arrays N*(s t , a t , r t , s t+1 ) from the experience pool to learn from past samples.
- the specific process is: input the state st and action at into the actual evaluation Q network to estimate Q value, while s t+1 is input to the target Q network to obtain the target Q value Q′, and is added to r t to calculate the root mean square error with the Q value estimated by the evaluation Q network. If the error is large, then It shows that more parameter updates are needed to reduce the error.
- experience replay is uniformly distributed sampling, that is, all experience samples have the same probability of being sampled.
- a technique of priority experience playback which is an improvement of the original experience playback technique.
- Priority experience replay is mainly to increase the probability of sampling valuable (that is, to obtain greater rewards) samples, and to preferentially extract the most valuable samples for learning, so as to learn more efficiently and quickly. Therefore, in order to carry out the operation of priority experience playback, it is necessary to first measure the value of experience samples, which can be judged by TD error, and can be obtained from the update formula of Q value:
- ⁇ t represents the TD error value
- symbols and meanings of other items are the same as formula (11).
- One of the optimization goals of the DDQN algorithm is to make the TD error as small as possible. If the TD error is large, it means that the current Q function is far from the target Q function, and more parameter updates are needed to reduce the TD error. Use TD error to measure the value of experience. In addition, in order to avoid over-fitting of the neural network, experience samples are also drawn by random probability to ensure that even samples with a reward of 0 have a probability of being drawn, so that the priority probability value of each sample experience is:
- p i
- ⁇ is a small positive number
- the purpose is to prevent the sample with reward 0 from being drawn with a probability of 0.
- the probability calculation of priority experience playback is also set in related code files such as Matlabm files.
- Matlabm files Before training starts, some key parameters need to be initialized in related code files such as Matlabm files, such as learning rate (generally less than 1), number of training rounds (generally within 1000 rounds), maximum time or period of each round T, is generally equal to the time of the training data set, the time interval of data collection Ts (0.1s-1s) and the time step of training, generally equal to (per cycle T)/(time interval of data collection Ts).
- learning rate generally less than 1
- number of training rounds generally within 1000 rounds
- maximum time or period of each round T is generally equal to the time of the training data set, the time interval of data collection Ts (0.1s-1s) and the time step of training, generally equal to (per cycle T)/(time interval of data collection Ts).
- SOC state of lithium battery and super capacitor
- the Drive cycle source module in the Simulink model by using its own open source standard driving cycle conditions such as the European standard driving cycle conditions NEDC (new European driving cycle), WLTC (worldwide harmonized light vehicle test procedure cycle ) or reorganize different standard driving cycles into a mixed driving cycle, etc., as a data set based on DDQN energy management method training.
- NEDC new European driving cycle
- WLTC worldwide harmonized light vehicle test procedure cycle
- reorganize different standard driving cycles into a mixed driving cycle etc.
- the instruction to call the Simulink agent-environment (vehicle) model is set and the DDQN method after the above-mentioned design is completed can start training.
- the parallel computing that Matlab carries The parallel computing function of the toolbox (parallel computing toolbox) accelerates the training and convergence process of the energy management method based on DDQN, thereby significantly reducing the training time.
- the specific training process of the energy management method based on DDQN is as follows.
- the required power is calculated according to the standard driving cycle conditions used for training and the selected relevant vehicle states are input to the intelligent In the body, the input state observations will be input to the evaluation network that stores in the experience pool and estimates the Q value respectively.
- the network selects the output action a t to the environment (vehicle) according to the existing method and the ⁇ -greedy principle.
- the state of the environment (vehicle) changes from st to st +1 .
- the environment (vehicle) immediately feeds back the reward of the corresponding action a t to the agent’s experience pool according to the reward function.
- the network parameter ⁇ of the evaluation Q network will be copied to the target Q network at a specific frequency and become the network parameter ⁇ ′ of the target Q network. Extract a small batch of samples from the evaluation Q network and input them into the evaluation Q network and the target Q network respectively, and the target Q network outputs the target Q value corresponding to the next moment st +1 , which is added to the reward r t in the experience sample, and then Calculate the root mean square value together with the estimated Q value, which is the loss L of the network (formula (13)), and calculate the partial differential of the loss L to the network parameter ⁇ It is the loss gradient, and this value is fed back to the estimated Q network. According to the principle of minimizing the loss, the estimated Q network will continuously update the network parameters ⁇ , so that the output action can obtain the maximum cumulative reward. This process will always be in each The training epochs are repeated until eventually convergence.
- the trained energy management method based on DDQN can be passed through the code
- the instruction generates a trained and ready-to-simulate control method in the RL agent module in the environment (vehicle) model in Simulink.
- the rule-based method is designed using Simulink/State flow and integrated with the DDQN method to form a hybrid energy management method.
- the provided hybrid vehicle energy management method based on rules and double-depth Q-networks includes: Step S510, detecting the energy source of the vehicle equipped with a composite energy storage system , the composite energy storage system includes a lithium battery energy source and a supercapacitor energy source; step S520, when the lithium battery is in a preset normal constrained operating range, use the trained deep reinforcement learning model to determine the output power of the lithium battery, A rule-based approach is used to protect lithium batteries when they are not within the normal constrained operating range.
- the deep reinforcement learning model is the double-depth Q network designed and trained according to the above.
- the control logic of the proposed hybrid vehicle energy management method is: when the lithium battery is in the preset normal constraint operating range (SOC _min ⁇ SOC batt ⁇ SOC _max ), use the dual-depth Q network energy management method, and when the SOC of the lithium battery exceeds the normal constraint range (SOC ⁇ SOC _min &SOC>SOC _max ), the rule-based method is used to protect the lithium battery to avoid overcharging and overcharging of the lithium battery Put it in order to prolong the service life of the lithium battery.
- the lithium battery SOC batt When the lithium battery SOC batt is lower than the set lower limit SOC_min , the lithium battery stops discharging and only accepts charging; then judge whether the supercapacitor SOC sc is higher than the lower limit SOC c_min , if If it is lower than the power requirement of the vehicle, the supercapacitor will be discharged for a short time; if it is lower than that, the driver can only be reminded to stop driving and find a charging pile to charge as soon as possible; 2) If the SOC batt of the lithium battery is higher than the upper limit SOC_max , then Make the lithium battery no longer charge, only discharge; judge whether the supercapacitor SOC sc is higher than the upper limit SOC c_max at this time, if it is higher, then give up the braking energy recovery; if it is lower, the supercapacitor absorbs all the braking Energy recovery power.
- the hybrid energy management method based on rules and double-depth Q-networks utilizes dynamic programming optimal results and priority experience playback techniques to achieve the following advantages: it does not depend on known conditions (vehicle speed, road conditions, etc.) ) and the existing model; can achieve faster convergence speed, better optimization effect; can achieve a more comprehensive control range, for example, lithium battery SOC has a corresponding control method between [0,1] for energy Management, effectively solve the problem of sudden reduction in service life of lithium batteries due to frequent overcharging and overdischarging.
- the present invention can be a system, method and/or computer program product.
- a computer program product may include a computer readable storage medium having computer readable program instructions thereon for causing a processor to implement various aspects of the present invention.
- a computer readable storage medium may be a tangible device that can retain and store instructions for use by an instruction execution device.
- a computer readable storage medium may be, for example, but is not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing.
- Computer-readable storage media include: portable computer diskettes, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), or flash memory), static random access memory (SRAM), compact disc read only memory (CD-ROM), digital versatile disc (DVD), memory stick, floppy disk, mechanically encoded device, such as a printer with instructions stored thereon A hole card or a raised structure in a groove, and any suitable combination of the above.
- RAM random access memory
- ROM read-only memory
- EPROM erasable programmable read-only memory
- flash memory static random access memory
- SRAM static random access memory
- CD-ROM compact disc read only memory
- DVD digital versatile disc
- memory stick floppy disk
- mechanically encoded device such as a printer with instructions stored thereon
- a hole card or a raised structure in a groove and any suitable combination of the above.
- computer-readable storage media are not to be construed as transient signals per se, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., pulses of light through fiber optic cables), or transmitted electrical signals.
- Computer readable program instructions described herein may be downloaded from a computer readable storage medium to a respective computing/processing device, or downloaded to an external computer or external storage device over a network, such as the Internet, local area network, wide area network, and/or wireless network.
- the network may include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and/or edge servers.
- a network adapter card or a network interface in each computing/processing device receives computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium in each computing/processing device .
- Computer program instructions for carrying out operations of the present invention may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state setting data, or Source or object code written in any combination, including object-oriented programming languages—such as Smalltalk, C++, Python, etc., and conventional procedural programming languages—such as the “C” language or similar programming languages.
- Computer readable program instructions may execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer, or entirely on the remote computer or server implement.
- the remote computer can be connected to the user computer through any kind of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computer (such as via the Internet using an Internet service provider). connect).
- LAN local area network
- WAN wide area network
- an electronic circuit such as a programmable logic circuit, field programmable gate array (FPGA), or programmable logic array (PLA)
- FPGA field programmable gate array
- PDA programmable logic array
- These computer-readable program instructions may be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine such that when executed by the processor of the computer or other programmable data processing apparatus , producing an apparatus for realizing the functions/actions specified in one or more blocks in the flowchart and/or block diagram.
- These computer-readable program instructions can also be stored in a computer-readable storage medium, and these instructions cause computers, programmable data processing devices and/or other devices to work in a specific way, so that the computer-readable medium storing instructions includes An article of manufacture comprising instructions for implementing various aspects of the functions/acts specified in one or more blocks in flowcharts and/or block diagrams.
- each block in a flowchart or block diagram may represent a module, a portion of a program segment, or an instruction that includes one or more Executable instructions.
- the functions noted in the block may occur out of the order noted in the figures. For example, two blocks in succession may, in fact, be executed substantially concurrently, or they may sometimes be executed in the reverse order, depending upon the functionality involved.
- each block of the block diagrams and/or flowchart illustrations, and combinations of blocks in the block diagrams and/or flowchart illustrations can be implemented by a dedicated hardware-based system that performs the specified function or action , or may be implemented by a combination of dedicated hardware and computer instructions. It is well known to those skilled in the art that implementation by means of hardware, implementation by means of software, and implementation by a combination of software and hardware are all equivalent.
Landscapes
- Engineering & Computer Science (AREA)
- Power Engineering (AREA)
- Transportation (AREA)
- Mechanical Engineering (AREA)
- Life Sciences & Earth Sciences (AREA)
- Sustainable Development (AREA)
- Sustainable Energy (AREA)
- Electric Propulsion And Braking For Vehicles (AREA)
Abstract
Description
Claims (10)
- 一种基于规则和双深度Q网络的混合动力汽车能量管理方法,包括以下步骤:检测设有复合储能系统的车辆能量源,该复合储能系统包括锂电池能量源和超级电容能量源;在检测到锂电池处于预设的正常约束工作范围的情况下,利用经训练的深度强化学习模型确定锂电池的输出功率,在检测到锂电池没有处于正常约束工作范围的情况下,则使用设定规则对锂电池进行保护;其中,对于所述深度强化学习模型,其智能体包括评估Q网络和目标Q网络,环境的状态观测量是锂电池的剩余电量、超级电容的剩余电量以及整车需求功率,锂电池输出功率作为输出动作,并以最小化所述复合储能系统的能量损失作为目标,设置相应的奖励函数。
- 根据权利要求1所述的方法,其特征在于,所述奖励函数设置为:E loss=L sc+L batt+L dcdc其中,E loss表示复合储能系统整体的能量损耗,SOC b_tgt表示锂电池剩余电量SOC b的目标值,SOC sc_tgt表示超级电容剩余电量SOC sc的目标值,L sc,L batt,L dcdc分别表示超级电容、锂电池和复合储能系统中DCDC转换器的能量损失,m表示能量损失项的系数,n和p分别表示平衡锂电池和超级电容的剩余电量变化的系数;利用所述深度强化学习模型求解能量优化问题的约束设置为:其中,I b为锂电池的电流,I sc为超级电容的电流,I b_min和I b_max分别表示锂电池电流的最小和最大值;I sc_min和I sc_max分别表示超级电容电流的最小和最大值,P dem表示整车需求功率,P min和P max分别表示整车需求功率的最小和最大值,P batt表示锂电池输出功率,P batt_min和P batt_max分别表示锂电池输出功率的最小和最大值。
- 根据权利要求1所述的方法,其特征在于,所述使用设定规则对锂电池进行保护包括以下步骤:当锂电池剩余电量SOC batt低于设定的下限值SOC _min时,令锂电池停止放电,只接受充电;并判断超级电容剩余电量SOC sc是否高于限制的下限值SOC c_min,若高于,则令超级电容根据车辆需求功率短时放电;若低于,则提醒驾驶员停止行驶;若锂电池的剩余电量SOC batt高于上限值SOC _max,则令锂电池不再充电,只进行放电;并判断超级电容剩余电量SOC sc是否高于限制的上限SOC c_max,若高于,则放弃制动能量回收;若低于,则超级电容吸收全部的制动能量回收功率。
- 根据权利要求1所述的方法,其特征在于,所述深度强化学习模型的训练过程包括:在每个时间步长,根据用于训练的标准驾驶循环工况计算得出需求功率和选定的相关车辆状态输入到智能体中,其中输入的状态观测量分别输入到经验池中存储以及评估Q网络中,该评估Q网络依据ε-greedy原则选择输出动作a t到环境中,环境的状态由s t变成s t+1;环境根据奖励函数将相应动作a t的奖励反馈到智能体的经验池中,在该过程中,评估Q网络的网络参数θ以设定的频率复制到目标Q网络中,成为目标Q网络的网络参数θ′;当经验池中存储一定的经验样本后,从经验池中提取一小批量的样本分别输入到评估Q网络和目标Q网络中,目标Q网络输出下一时刻s t+1对应的目标Q值,基于目标Q值和估计的Q值计算损失值及其对评估Q网络参数的损失梯度,并该值反馈到估计Q网络中,估计Q网络依据最小化损失的原则更新网络参数θ,以使输出的动作获得最大累计奖励。
- 一种计算机可读存储介质,其上存储有计算机程序,其中,该程序被处理器执行时实现根据权利要求1至8中任一项所述方法的步骤。
- 一种计算机设备,包括存储器和处理器,在所述存储器上存储有 能够在处理器上运行的计算机程序,其特征在于,所述处理器执行所述程序时实现权利要求1至8中任一项所述的方法的步骤。
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN202110602198.0A CN113511082B (zh) | 2021-05-31 | 2021-05-31 | 基于规则和双深度q网络的混合动力汽车能量管理方法 |
| CN202110602198.0 | 2021-05-31 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2022252559A1 true WO2022252559A1 (zh) | 2022-12-08 |
Family
ID=78065129
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/CN2021/137803 Ceased WO2022252559A1 (zh) | 2021-05-31 | 2021-12-14 | 基于规则和双深度q网络的混合动力汽车能量管理方法 |
Country Status (2)
| Country | Link |
|---|---|
| CN (1) | CN113511082B (zh) |
| WO (1) | WO2022252559A1 (zh) |
Cited By (38)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN116054285A (zh) * | 2022-12-30 | 2023-05-02 | 国网湖北省电力有限公司经济技术研究院 | 一种基于联邦强化学习算法的输配调频资源协同控制方法 |
| CN116050505A (zh) * | 2023-02-22 | 2023-05-02 | 西南交通大学 | 一种基于伙伴网络的智能体深度强化学习方法 |
| CN116123702A (zh) * | 2023-02-23 | 2023-05-16 | 中交武汉智行国际工程咨询有限公司 | 基于支持向量机和深度强化学习的建筑空调温度控制方法 |
| CN116208545A (zh) * | 2023-02-07 | 2023-06-02 | 哈尔滨工业大学 | 网络路由优化方法及装置、智能路由方法及装置 |
| CN116257937A (zh) * | 2023-01-05 | 2023-06-13 | 东南大学 | 基于多目标深度强化学习的混合动力汽车生态驾驶方法 |
| CN116384250A (zh) * | 2023-04-12 | 2023-07-04 | 中国长江三峡集团有限公司 | 一种综合能源系统多个智能体的协作控制方法与能源系统 |
| CN116373673A (zh) * | 2023-03-23 | 2023-07-04 | 国网江苏省电力有限公司南通供电分公司 | 一种电动汽车充放电管理方法及系统 |
| CN116561579A (zh) * | 2023-05-17 | 2023-08-08 | 四川大学 | 自适应多工况的递归深度强化学习混动汽车能量管理策略 |
| CN116582860A (zh) * | 2023-05-08 | 2023-08-11 | 南京航空航天大学 | 一种基于信息年龄约束的链路资源分配方法 |
| CN116605092A (zh) * | 2023-04-26 | 2023-08-18 | 中国科学技术大学 | 一种能量管理方法及能量管理平台 |
| CN116736700A (zh) * | 2023-05-16 | 2023-09-12 | 武汉理工大学 | 多动力源两栖艇能量管理控制系统及方法 |
| CN116853073A (zh) * | 2023-09-04 | 2023-10-10 | 江西五十铃汽车有限公司 | 一种新能源电动汽车能量管理方法及系统 |
| CN116985646A (zh) * | 2023-09-28 | 2023-11-03 | 江西五十铃汽车有限公司 | 车辆超级电容器控制方法、设备和介质 |
| CN116995641A (zh) * | 2023-09-26 | 2023-11-03 | 北京交通大学 | 一种应用于轨道交通储能的能量管理架构及方法 |
| CN117021066A (zh) * | 2023-05-26 | 2023-11-10 | 浙江大学 | 一种基于深度强化学习的机器人视觉伺服运动控制方法 |
| CN117034443A (zh) * | 2023-07-13 | 2023-11-10 | 哈尔滨工业大学 | 基于ddqn算法的机动观测策略生成方法、改进ddqn算法的机动观测策略生成方法 |
| CN117021971A (zh) * | 2023-08-15 | 2023-11-10 | 北京工业大学 | 一种多动力源车辆能量管理策略优化方法、系统及设备 |
| CN117058468A (zh) * | 2023-10-11 | 2023-11-14 | 青岛金诺德科技有限公司 | 用于新能源汽车锂电池回收的图像识别与分类系统 |
| CN117227700A (zh) * | 2023-11-15 | 2023-12-15 | 北京理工大学 | 串联混合动力无人履带车辆的能量管理方法及系统 |
| CN117507789A (zh) * | 2023-12-06 | 2024-02-06 | 江苏开放大学(江苏城市职业学院) | 一种用于agv小车的多模混合动力驱动系统 |
| CN117578679A (zh) * | 2024-01-15 | 2024-02-20 | 太原理工大学 | 基于强化学习的锂电池智能充电控制方法 |
| CN117650522A (zh) * | 2023-12-04 | 2024-03-05 | 北京北交本有科技有限公司 | 一种基于深度强化学习的多储能系统分层协同控制方法 |
| CN117650553A (zh) * | 2023-10-25 | 2024-03-05 | 四川大学 | 基于多智能体深度强化学习的5g基站储能电池充放电调度方法 |
| CN117933666A (zh) * | 2024-03-21 | 2024-04-26 | 壹号智能科技(南京)有限公司 | 一种密集仓储机器人调度方法、装置、介质、设备及系统 |
| CN117944527A (zh) * | 2024-02-02 | 2024-04-30 | 山东德维鲁普新材料有限公司 | 基于氢燃料电池发电车的能量管理方法 |
| CN118544901A (zh) * | 2024-06-05 | 2024-08-27 | 郑州大学 | 一种基于双q学习与实时速度预测的氢-电混动系统能量管理方法、系统、设备及存储介质 |
| CN118596882A (zh) * | 2024-06-11 | 2024-09-06 | 徐州徐工重型车辆有限公司 | 纯电动矿用自卸车分布式线控轮边驱动系统及驱动方法 |
| CN118705133A (zh) * | 2024-06-15 | 2024-09-27 | 曲阜师范大学 | 一种基于深度强化学习的风力机转速参考值确定方法 |
| CN118859730A (zh) * | 2024-09-25 | 2024-10-29 | 中南大学 | 自适应感知的重载列车群组运行能量管理方法及控制装置 |
| CN118928165A (zh) * | 2024-10-12 | 2024-11-12 | 国网浙江省电力有限公司嘉善县供电公司 | 氢燃料电池车的多电源协调控制方法、系统、设备及介质 |
| CN119187463A (zh) * | 2024-10-30 | 2024-12-27 | 扬州市江都区飞跃泵业有限公司 | 一种双叶片半开式污水泵叶轮及其制造装置、方法 |
| CN119359063A (zh) * | 2024-09-09 | 2025-01-24 | 华北电力大学 | 电动汽车充放电调度策略生成方法 |
| CN119675213A (zh) * | 2025-02-20 | 2025-03-21 | 南昌科晨电力试验研究有限公司 | 一种应急移动电源的控制方法及应急移动电源 |
| CN119882416A (zh) * | 2024-12-04 | 2025-04-25 | 重庆大学 | 基于深度强化学习的混合动力汽车队列节能驾驶分层控制方法 |
| CN119872340A (zh) * | 2025-03-27 | 2025-04-25 | 北京迅巢科技有限公司 | 一种基于自适应算法的新能源电池负载优化系统及方法 |
| CN119911256A (zh) * | 2023-10-31 | 2025-05-02 | 北京罗克维尔斯科技有限公司 | 一种驾驶策略确定方法、装置、系统及存储介质 |
| CN120142941A (zh) * | 2025-02-25 | 2025-06-13 | 中山大学·深圳 | 一种基于模型短分支推演的锂电池充电策略样本效率增强方法及系统 |
| CN120387367A (zh) * | 2025-04-15 | 2025-07-29 | 北京信息科技大学 | 融合多车运动交互感知的网联混合动力汽车能量管理方法 |
Families Citing this family (10)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN113511082B (zh) * | 2021-05-31 | 2023-06-16 | 深圳先进技术研究院 | 基于规则和双深度q网络的混合动力汽车能量管理方法 |
| CN114475280B (zh) * | 2022-03-01 | 2024-09-13 | 武汉理工大学 | 一种电动汽车混合动力系统能量管理方法及系统 |
| CN114489144B (zh) * | 2022-04-08 | 2022-07-12 | 中国科学院自动化研究所 | 无人机自主机动决策方法、装置及无人机 |
| CN115001907A (zh) * | 2022-05-06 | 2022-09-02 | 河北华万电子科技有限公司 | 一种irs辅助微型配电网智能计算方法 |
| CN115190489B (zh) * | 2022-07-07 | 2024-11-29 | 内蒙古大学 | 基于深度强化学习的认知无线网络动态频谱接入方法 |
| CN115534764B (zh) * | 2022-08-29 | 2024-11-15 | 华东理工大学 | 基于深度强化学习的车载燃料电池系统控制方法及系统 |
| CN116231803B (zh) * | 2023-03-02 | 2024-01-23 | 深圳市华南英才科技有限公司 | 一种电容笔快速充电方法、装置及存储介质 |
| CN116777033A (zh) * | 2023-03-15 | 2023-09-19 | 阿维塔科技(重庆)有限公司 | 车辆剩余里程的相关方法、装置及设备 |
| CN116767024A (zh) * | 2023-04-14 | 2023-09-19 | 联合汽车电子有限公司 | 一种基于强化学习的电池均衡方法及装置 |
| CN119010616B (zh) * | 2024-08-02 | 2025-03-28 | 广东电网有限责任公司 | Lcl型并网逆变器控制方法、装置 |
Citations (8)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| EP1323564A2 (en) * | 2001-12-28 | 2003-07-02 | Nissan Motor Co., Ltd. | Control system for hybrid vehicle |
| CN109552079A (zh) * | 2019-01-28 | 2019-04-02 | 浙江大学宁波理工学院 | 一种基于规则与Q-learning增强学习的电动汽车复合能量管理方法 |
| CN109657194A (zh) * | 2018-12-04 | 2019-04-19 | 浙江大学宁波理工学院 | 一种基于Q-learning和规则的混合动力车辆运行实时能源管理方法 |
| CN110850877A (zh) * | 2019-11-19 | 2020-02-28 | 北方工业大学 | 基于虚拟环境和深度双q网络的自动驾驶小车训练方法 |
| CN111731303A (zh) * | 2020-07-09 | 2020-10-02 | 重庆大学 | 一种基于深度强化学习a3c算法的hev能量管理方法 |
| CN112287463A (zh) * | 2020-11-03 | 2021-01-29 | 重庆大学 | 一种基于深度强化学习算法的燃料电池汽车能量管理方法 |
| CN112670982A (zh) * | 2020-12-14 | 2021-04-16 | 广西电网有限责任公司电力科学研究院 | 一种基于奖励机制的微电网有功调度控制方法及系统 |
| CN113511082A (zh) * | 2021-05-31 | 2021-10-19 | 深圳先进技术研究院 | 基于规则和双深度q网络的混合动力汽车能量管理方法 |
Family Cites Families (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20210269152A1 (en) * | 2018-06-27 | 2021-09-02 | H3 Dynamics Holdings Pte. Ltd. | Distributed electric energy pods network and associated electrically powered vehicle |
| JP7151689B2 (ja) * | 2019-11-05 | 2022-10-12 | トヨタ自動車株式会社 | 電池管理システム、電池管理方法、及び組電池の製造方法 |
| CN110941202A (zh) * | 2019-12-12 | 2020-03-31 | 中国科学院深圳先进技术研究院 | 一种汽车能量管理策略的验证方法和设备 |
-
2021
- 2021-05-31 CN CN202110602198.0A patent/CN113511082B/zh active Active
- 2021-12-14 WO PCT/CN2021/137803 patent/WO2022252559A1/zh not_active Ceased
Patent Citations (8)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| EP1323564A2 (en) * | 2001-12-28 | 2003-07-02 | Nissan Motor Co., Ltd. | Control system for hybrid vehicle |
| CN109657194A (zh) * | 2018-12-04 | 2019-04-19 | 浙江大学宁波理工学院 | 一种基于Q-learning和规则的混合动力车辆运行实时能源管理方法 |
| CN109552079A (zh) * | 2019-01-28 | 2019-04-02 | 浙江大学宁波理工学院 | 一种基于规则与Q-learning增强学习的电动汽车复合能量管理方法 |
| CN110850877A (zh) * | 2019-11-19 | 2020-02-28 | 北方工业大学 | 基于虚拟环境和深度双q网络的自动驾驶小车训练方法 |
| CN111731303A (zh) * | 2020-07-09 | 2020-10-02 | 重庆大学 | 一种基于深度强化学习a3c算法的hev能量管理方法 |
| CN112287463A (zh) * | 2020-11-03 | 2021-01-29 | 重庆大学 | 一种基于深度强化学习算法的燃料电池汽车能量管理方法 |
| CN112670982A (zh) * | 2020-12-14 | 2021-04-16 | 广西电网有限责任公司电力科学研究院 | 一种基于奖励机制的微电网有功调度控制方法及系统 |
| CN113511082A (zh) * | 2021-05-31 | 2021-10-19 | 深圳先进技术研究院 | 基于规则和双深度q网络的混合动力汽车能量管理方法 |
Cited By (47)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN116054285A (zh) * | 2022-12-30 | 2023-05-02 | 国网湖北省电力有限公司经济技术研究院 | 一种基于联邦强化学习算法的输配调频资源协同控制方法 |
| CN116257937A (zh) * | 2023-01-05 | 2023-06-13 | 东南大学 | 基于多目标深度强化学习的混合动力汽车生态驾驶方法 |
| CN116208545A (zh) * | 2023-02-07 | 2023-06-02 | 哈尔滨工业大学 | 网络路由优化方法及装置、智能路由方法及装置 |
| CN116050505A (zh) * | 2023-02-22 | 2023-05-02 | 西南交通大学 | 一种基于伙伴网络的智能体深度强化学习方法 |
| CN116123702A (zh) * | 2023-02-23 | 2023-05-16 | 中交武汉智行国际工程咨询有限公司 | 基于支持向量机和深度强化学习的建筑空调温度控制方法 |
| CN116373673A (zh) * | 2023-03-23 | 2023-07-04 | 国网江苏省电力有限公司南通供电分公司 | 一种电动汽车充放电管理方法及系统 |
| CN116384250A (zh) * | 2023-04-12 | 2023-07-04 | 中国长江三峡集团有限公司 | 一种综合能源系统多个智能体的协作控制方法与能源系统 |
| CN116605092A (zh) * | 2023-04-26 | 2023-08-18 | 中国科学技术大学 | 一种能量管理方法及能量管理平台 |
| CN116582860A (zh) * | 2023-05-08 | 2023-08-11 | 南京航空航天大学 | 一种基于信息年龄约束的链路资源分配方法 |
| CN116736700A (zh) * | 2023-05-16 | 2023-09-12 | 武汉理工大学 | 多动力源两栖艇能量管理控制系统及方法 |
| CN116561579A (zh) * | 2023-05-17 | 2023-08-08 | 四川大学 | 自适应多工况的递归深度强化学习混动汽车能量管理策略 |
| CN116561579B (zh) * | 2023-05-17 | 2026-04-21 | 四川大学 | 自适应多工况的递归深度强化学习混动汽车能量管理策略 |
| CN117021066A (zh) * | 2023-05-26 | 2023-11-10 | 浙江大学 | 一种基于深度强化学习的机器人视觉伺服运动控制方法 |
| CN117034443A (zh) * | 2023-07-13 | 2023-11-10 | 哈尔滨工业大学 | 基于ddqn算法的机动观测策略生成方法、改进ddqn算法的机动观测策略生成方法 |
| CN117021971A (zh) * | 2023-08-15 | 2023-11-10 | 北京工业大学 | 一种多动力源车辆能量管理策略优化方法、系统及设备 |
| CN116853073B (zh) * | 2023-09-04 | 2024-01-26 | 江西五十铃汽车有限公司 | 一种新能源电动汽车能量管理方法及系统 |
| CN116853073A (zh) * | 2023-09-04 | 2023-10-10 | 江西五十铃汽车有限公司 | 一种新能源电动汽车能量管理方法及系统 |
| CN116995641B (zh) * | 2023-09-26 | 2023-12-22 | 北京交通大学 | 一种应用于轨道交通储能的能量管理架构及方法 |
| CN116995641A (zh) * | 2023-09-26 | 2023-11-03 | 北京交通大学 | 一种应用于轨道交通储能的能量管理架构及方法 |
| CN116985646B (zh) * | 2023-09-28 | 2024-01-12 | 江西五十铃汽车有限公司 | 车辆超级电容器控制方法、设备和介质 |
| CN116985646A (zh) * | 2023-09-28 | 2023-11-03 | 江西五十铃汽车有限公司 | 车辆超级电容器控制方法、设备和介质 |
| CN117058468A (zh) * | 2023-10-11 | 2023-11-14 | 青岛金诺德科技有限公司 | 用于新能源汽车锂电池回收的图像识别与分类系统 |
| CN117058468B (zh) * | 2023-10-11 | 2023-12-19 | 青岛金诺德科技有限公司 | 用于新能源汽车锂电池回收的图像识别与分类系统 |
| CN117650553A (zh) * | 2023-10-25 | 2024-03-05 | 四川大学 | 基于多智能体深度强化学习的5g基站储能电池充放电调度方法 |
| CN119911256A (zh) * | 2023-10-31 | 2025-05-02 | 北京罗克维尔斯科技有限公司 | 一种驾驶策略确定方法、装置、系统及存储介质 |
| CN117227700A (zh) * | 2023-11-15 | 2023-12-15 | 北京理工大学 | 串联混合动力无人履带车辆的能量管理方法及系统 |
| CN117227700B (zh) * | 2023-11-15 | 2024-02-06 | 北京理工大学 | 串联混合动力无人履带车辆的能量管理方法及系统 |
| CN117650522A (zh) * | 2023-12-04 | 2024-03-05 | 北京北交本有科技有限公司 | 一种基于深度强化学习的多储能系统分层协同控制方法 |
| CN117507789A (zh) * | 2023-12-06 | 2024-02-06 | 江苏开放大学(江苏城市职业学院) | 一种用于agv小车的多模混合动力驱动系统 |
| CN117507789B (zh) * | 2023-12-06 | 2024-06-07 | 江苏开放大学(江苏城市职业学院) | 一种用于agv小车的多模混合动力驱动系统 |
| CN117578679B (zh) * | 2024-01-15 | 2024-03-22 | 太原理工大学 | 基于强化学习的锂电池智能充电控制方法 |
| CN117578679A (zh) * | 2024-01-15 | 2024-02-20 | 太原理工大学 | 基于强化学习的锂电池智能充电控制方法 |
| CN117944527A (zh) * | 2024-02-02 | 2024-04-30 | 山东德维鲁普新材料有限公司 | 基于氢燃料电池发电车的能量管理方法 |
| CN117933666A (zh) * | 2024-03-21 | 2024-04-26 | 壹号智能科技(南京)有限公司 | 一种密集仓储机器人调度方法、装置、介质、设备及系统 |
| CN118544901A (zh) * | 2024-06-05 | 2024-08-27 | 郑州大学 | 一种基于双q学习与实时速度预测的氢-电混动系统能量管理方法、系统、设备及存储介质 |
| CN118596882A (zh) * | 2024-06-11 | 2024-09-06 | 徐州徐工重型车辆有限公司 | 纯电动矿用自卸车分布式线控轮边驱动系统及驱动方法 |
| CN118705133A (zh) * | 2024-06-15 | 2024-09-27 | 曲阜师范大学 | 一种基于深度强化学习的风力机转速参考值确定方法 |
| CN119359063A (zh) * | 2024-09-09 | 2025-01-24 | 华北电力大学 | 电动汽车充放电调度策略生成方法 |
| CN118859730A (zh) * | 2024-09-25 | 2024-10-29 | 中南大学 | 自适应感知的重载列车群组运行能量管理方法及控制装置 |
| CN118928165A (zh) * | 2024-10-12 | 2024-11-12 | 国网浙江省电力有限公司嘉善县供电公司 | 氢燃料电池车的多电源协调控制方法、系统、设备及介质 |
| CN119187463A (zh) * | 2024-10-30 | 2024-12-27 | 扬州市江都区飞跃泵业有限公司 | 一种双叶片半开式污水泵叶轮及其制造装置、方法 |
| CN119882416A (zh) * | 2024-12-04 | 2025-04-25 | 重庆大学 | 基于深度强化学习的混合动力汽车队列节能驾驶分层控制方法 |
| CN119882416B (zh) * | 2024-12-04 | 2025-09-30 | 重庆大学 | 基于深度强化学习的混合动力汽车队列节能驾驶分层控制方法 |
| CN119675213A (zh) * | 2025-02-20 | 2025-03-21 | 南昌科晨电力试验研究有限公司 | 一种应急移动电源的控制方法及应急移动电源 |
| CN120142941A (zh) * | 2025-02-25 | 2025-06-13 | 中山大学·深圳 | 一种基于模型短分支推演的锂电池充电策略样本效率增强方法及系统 |
| CN119872340A (zh) * | 2025-03-27 | 2025-04-25 | 北京迅巢科技有限公司 | 一种基于自适应算法的新能源电池负载优化系统及方法 |
| CN120387367A (zh) * | 2025-04-15 | 2025-07-29 | 北京信息科技大学 | 融合多车运动交互感知的网联混合动力汽车能量管理方法 |
Also Published As
| Publication number | Publication date |
|---|---|
| CN113511082A (zh) | 2021-10-19 |
| CN113511082B (zh) | 2023-06-16 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| WO2022252559A1 (zh) | 基于规则和双深度q网络的混合动力汽车能量管理方法 | |
| CN110341690B (zh) | 一种基于确定性策略梯度学习的phev能量管理方法 | |
| Du et al. | Heuristic energy management strategy of hybrid electric vehicle based on deep reinforcement learning with accelerated gradient optimization | |
| Wu et al. | Continuous reinforcement learning of energy management with deep Q network for a power split hybrid electric bus | |
| Zhao et al. | A deep reinforcement learning framework for optimizing fuel economy of hybrid electric vehicles | |
| Bo et al. | A Q-learning fuzzy inference system based online energy management strategy for off-road hybrid electric vehicles | |
| Yu et al. | Real time energy management strategy for a fast charging electric urban bus powered by hybrid energy storage system | |
| Wang et al. | Deep reinforcement learning with deep-Q-network based energy management for fuel cell hybrid electric truck | |
| Lin et al. | Intelligent energy management strategy based on an improved reinforcement learning algorithm with exploration factor for a plug-in PHEV | |
| CN112215434A (zh) | 一种lstm模型的生成方法、充电时长预测方法及介质 | |
| Li et al. | An improved energy management strategy of fuel cell hybrid vehicles based on proximal policy optimization algorithm | |
| CN119167791B (zh) | 一种氢燃料电池混合动力无人机能量自适应优化系统及方法 | |
| CN110297452B (zh) | 一种蓄电池组相邻型均衡系统及其预测控制方法 | |
| El Ouazzani et al. | MSCC-DRL: Multi-Stage constant current based on deep reinforcement learning for fast charging of lithium ion battery | |
| Zhang et al. | A review of energy management optimization based on the equivalent consumption minimization strategy for fuel cell hybrid power systems | |
| Sun et al. | A new design of fuzzy logic control for SMES and battery hybrid storage system | |
| US20250045594A1 (en) | Energy storage system, and thermal management method for energy storage system | |
| Wang et al. | An efficient power control scheme for heavy-duty hybrid electric vehicle with online optimized variable universe fuzzy system | |
| US20250214564A1 (en) | Multi-objective optimization method and system assisted by gradient boosted neural network and device | |
| Gao et al. | Real-time three-level energy management strategy for series hybrid wheel loaders based on WG-MPC | |
| Yang et al. | A DOD-SOH balancing control method for dynamic reconfigurable battery systems based on DQN algorithm | |
| Gao et al. | An energy management strategy for fuel cell hybrid electric vehicle based on a real-time model predictive control and pontryagin’s maximum principle | |
| Hu et al. | Intelligent energy management for the fuel cell electric vehicle: A real-time, deep reinforcement learning application to multi-objective optimization | |
| Li et al. | Adaptive hierarchical energy management strategy for fuel cell hybrid engineering vehicles based on deep reinforcement learning | |
| CN117214755A (zh) | 一种基于混合神经网络的锂离子电池soc、soe及soh联合估计方法 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 21943911 Country of ref document: EP Kind code of ref document: A1 |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 21943911 Country of ref document: EP Kind code of ref document: A1 |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 21943911 Country of ref document: EP Kind code of ref document: A1 |
|
| 32PN | Ep: public notification in the ep bulletin as address of the adressee cannot be established |
Free format text: NOTING OF LOSS OF RIGHTS PURSUANT TO RULE 112(1) EPC (EPO FORM 1205 DATED 31/12/2024) |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 21943911 Country of ref document: EP Kind code of ref document: A1 |






