WO2020177417A1 - 控制无人驾驶设备及训练模型 - Google Patents

控制无人驾驶设备及训练模型 Download PDF

Info

Publication number
WO2020177417A1
WO2020177417A1 PCT/CN2019/123394 CN2019123394W WO2020177417A1 WO 2020177417 A1 WO2020177417 A1 WO 2020177417A1 CN 2019123394 W CN2019123394 W CN 2019123394W WO 2020177417 A1 WO2020177417 A1 WO 2020177417A1
Authority
WO
WIPO (PCT)
Prior art keywords
data
target
parameter value
neural network
control parameter
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/CN2019/123394
Other languages
English (en)
French (fr)
Inventor
穆荣均
夏华夏
任冬淳
郭潇阳
付圣
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Beijing Sankuai Online Technology Co Ltd
Original Assignee
Beijing Sankuai Online Technology Co Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Beijing Sankuai Online Technology Co Ltd filed Critical Beijing Sankuai Online Technology Co Ltd
Publication of WO2020177417A1 publication Critical patent/WO2020177417A1/zh
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Images

Classifications

    • GPHYSICS
    • G05CONTROLLING; REGULATING
    • G05BCONTROL OR REGULATING SYSTEMS IN GENERAL; FUNCTIONAL ELEMENTS OF SUCH SYSTEMS; MONITORING OR TESTING ARRANGEMENTS FOR SUCH SYSTEMS OR ELEMENTS
    • G05B13/00Adaptive control systems, i.e. systems automatically adjusting themselves to have a performance which is optimum according to some preassigned criterion
    • G05B13/02Adaptive control systems, i.e. systems automatically adjusting themselves to have a performance which is optimum according to some preassigned criterion electric
    • G05B13/04Adaptive control systems, i.e. systems automatically adjusting themselves to have a performance which is optimum according to some preassigned criterion electric involving the use of models or simulators
    • G05B13/042Adaptive control systems, i.e. systems automatically adjusting themselves to have a performance which is optimum according to some preassigned criterion electric involving the use of models or simulators in which a parameter or coefficient is automatically adjusted to optimise the performance

Definitions

  • This application relates to the field of unmanned driving technology, particularly how to control unmanned driving equipment and training models.
  • the current control guidance data can be obtained by referring to the driving decision, and the current control can be obtained by looking up the table (query the correspondence table between control guidance data and control parameter values) The control parameter value corresponding to the guidance data for unmanned driving control.
  • this application provides a method, device and electronic device for controlling unmanned driving equipment and model training.
  • a method for controlling an unmanned driving device including:
  • the target latent variable is used to indicate a conversion influence factor between the control guidance data and the control parameter value for the target device;
  • a model training method for controlling an unmanned driving device including:
  • each set of sample data includes control guidance data and control parameter values; Input to the current convolutional neural network to obtain hidden variables, the hidden variables used to represent the conversion influence factors between the control guidance data and the control parameter values; the hidden variables and the control guidance data corresponding to the second data Input to the current recurrent neural network to get the predicted parameter value;
  • the network parameters of the convolutional neural network and the recurrent neural network are adjusted, and the Target operation
  • the adjusted target convolutional neural network and the target cyclic neural network are output.
  • an apparatus for controlling an unmanned driving device including:
  • the obtaining module is used to obtain the current control guidance data of the target device, and obtain the predetermined target hidden variable corresponding to the target device; the target hidden variable is used to indicate the control guidance data and the control parameter value for the target device Factors affecting the conversion between;
  • a determining module configured to obtain the current control parameter value based on the current control guidance data and the target hidden variable
  • the control module is used to control the target device according to the current control parameter value.
  • a model training device for controlling an unmanned driving device including:
  • the execution module is used to perform the following target operations: select multiple sets of sample data as the first data from the sample set, and select a set of sample data as the second data, each set of the sample data includes control guidance data and control parameter values;
  • the first data is input to the current convolutional neural network to obtain a hidden variable, and the hidden variable is used to represent the conversion influence factor between the control guidance data and the control parameter value;
  • the hidden variable and the second data The corresponding control guidance data is input to the current cyclic neural network to obtain the predicted parameter value;
  • the adjustment module is configured to adjust the network parameters of the convolutional neural network and the cyclic neural network when it is determined that a preset condition is not met based on the predicted parameter value and the control parameter value corresponding to the second data , And instruct the execution module to re-execute the target operation;
  • the output module is configured to output the adjusted target convolutional neural network and target cyclic neural network when it is determined that the preset condition is satisfied based on the predicted parameter value and the control parameter value corresponding to the second data.
  • a computer-readable storage medium stores a computer program, and the computer program implements any one of the first aspect and the second aspect when the computer program is executed by a processor The method described.
  • an electronic device including a memory, a processor, and a computer program stored in the memory and capable of running on the processor.
  • the processor implements the first Aspect and the method of any one of the second aspect.
  • Fig. 1 is a flowchart of a method for controlling an unmanned device according to an exemplary embodiment of the present application
  • Fig. 2 is a flow chart of a method for controlling model training of an unmanned driving device according to an exemplary embodiment of the present application
  • Fig. 3 is a block diagram showing a device for controlling an unmanned driving device according to an exemplary embodiment of the present application
  • Fig. 4 is a block diagram of a device for controlling model training of unmanned driving equipment according to an exemplary embodiment of the present application
  • Fig. 5 is a schematic structural diagram of an electronic device according to an exemplary embodiment of the present application.
  • Fig. 6 is a schematic structural diagram of another electronic device according to an exemplary embodiment of the present application.
  • first, second, third, etc. may be used in this application to describe various information, the information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other.
  • first information may also be referred to as second information, and similarly, the second information may also be referred to as first information.
  • word “if” as used herein can be interpreted as "when” or “when” or "in response to determination”.
  • the current control guidance data can be obtained by referring to the driving decision, and the current control can be obtained by looking up the table (query the correspondence table between control guidance data and control parameter values)
  • the control parameter value corresponding to the guidance data for unmanned driving control is obtained by looking up the table (query the correspondence table between control guidance data and control parameter values)
  • the control parameter value corresponding to the guidance data for unmanned driving control is obtained by looking up the table (query the correspondence table between control guidance data and control parameter values)
  • the control parameter value corresponding to the guidance data for unmanned driving control can be obtained by looking up the table (query the correspondence table between control guidance data and control parameter values)
  • the control parameter value corresponding to the guidance data for unmanned driving control is obtained by looking up the table (query the correspondence table between control guidance data and control parameter values)
  • the control parameter value corresponding to the guidance data for unmanned driving control is obtained by looking up the table (query the correspondence table between control guidance data and control parameter values)
  • Fig. 1 is a flowchart showing a method for controlling an unmanned driving device according to an exemplary embodiment, and the method may be applied to an unmanned driving device.
  • the unmanned equipment may include, but is not limited to, unmanned vehicles, unmanned operation robots, unmanned aerial vehicles, unmanned ships, and the like. The method includes the following steps:
  • step 101 the current control guidance data of the target device is obtained.
  • step 102 a predetermined target hidden variable corresponding to the target device is obtained, and the target hidden variable is used to indicate the conversion influence factor between the control guidance data and the control parameter value for the target device.
  • step 103 the current control parameter value is obtained based on the current control guidance data and the target hidden variable.
  • the target device is an unmanned device to be controlled.
  • the target device may be an unmanned vehicle, or an unmanned robot, or an unmanned aerial vehicle, or an unmanned ship.
  • the specific type of equipment is not limited.
  • the current control parameter value is the value of the parameter currently used to control the target device.
  • the control parameter value may be the chassis control parameter value of the unmanned vehicle (eg, The amount of throttle or brake control, etc.) etc. It can be understood that the control parameter value may also be other control parameter values, which is not limited in this application.
  • the current control guidance data of the target device can be used to determine the current control parameter value.
  • the current control guidance data can include the current operating speed value of the unmanned vehicle and the acceleration value currently to be applied. It can be understood that the control guidance data may also be any other data that can be used to determine the current control parameter value, which is not limited in this application.
  • the target hidden variable corresponding to the target device can be used to represent the conversion influencing factor between the control guidance data and the control parameter value, and the conversion influencing factor is a conversion influencing factor for the target device.
  • the target hidden variable can be determined in advance based on the trained target convolutional neural network, and the target hidden variable can be stored in the target device.
  • the target device is unmanned control, the target can be obtained from the data stored in the target device Hidden variables, and autonomous driving control based on the target hidden variables.
  • the target convolutional neural network and the target recurrent neural network are both pre-trained models.
  • the target convolutional neural network and the target recurrent neural network can be trained in the following manner: perform the following target operations: Select multiple sets of sample data from the sample set as the first data, and select a set of sample data as the second data.
  • Each set of sample data includes control guidance data and control parameter values.
  • the sample data in the sample set are all for the same model.
  • the first data and the second data are data collected for the same unmanned driving device.
  • the first data is input into the current convolutional neural network to obtain a hidden variable, which is used to represent the conversion influence factor between the control guidance data and the control parameter value.
  • the adjusted convolutional neural network can be used as the target neural network, and the adjusted cyclic neural network can be used as the target cyclic neural network.
  • the randomly generated noise signal can also be input to the current cyclic neural network. And, it is guaranteed that the first data and the second data are data collected for the same unmanned driving device.
  • step 104 the target device is controlled according to the current control parameter value.
  • the current control guidance data and target hidden variables can be input to the target cyclic neural network, and the output result of the target cyclic neural network is used as the current control parameter value.
  • the target device can be controlled according to the current control parameter value.
  • the method for controlling an unmanned device obtains the current control guidance data of the target device and obtains the predetermined target hidden variable corresponding to the target device, based on the current control guidance data and the target hidden variable, Obtain the current control parameter value, and control the target device according to the current control parameter value.
  • the target hidden variable is used to represent the conversion influence factor between the control guidance data and the control parameter value for the target device.
  • the target hidden variable corresponding to the target device can be determined in advance by the following method: multiple sets of data collected for the target device can be determined, each set of data includes control guidance data and control parameter values, and the above multiple The group of data is input into the target convolutional neural network, and the target hidden variable output by the target convolutional neural network is obtained.
  • a driving test can be performed on the target device in advance.
  • multiple sets of data are collected, and each set of data may include control guidance data and control parameter values corresponding to the control guidance data.
  • the number of groups of data collected during the driving test can be any reasonable number, for example, it can be 3 groups, or 5 groups, or 10 groups. It is understandable that this application does not limit the specific number of data collected during the driving test.
  • multiple sets of data collected in advance for the target device can be determined, and the multiple sets of samples are input into the target convolutional neural network, and the output result of the target convolutional neural network is used as the target hidden variable.
  • the target latent variable that can represent the conversion influence factor between the control guidance data and the control parameter value can be obtained.
  • the target hidden variable is a hidden variable for the target device. Therefore, the accuracy of the control parameter value is further improved.
  • FIG. 2 is a flowchart of a method for controlling model training of an unmanned driving device according to an exemplary embodiment.
  • the method can be applied to a terminal device or a server .
  • the method includes the following steps:
  • step 201 multiple sets of sample data are selected from the sample set as the first data, and one set of sample data is selected as the second data, and each set of sample data includes control guidance data and control parameter values.
  • driving tests can be performed on multiple different unmanned driving devices of the same model in advance.
  • a large number of sample data are collected to obtain a sample set (the sample data in the sample set all correspond to the same Model of unmanned equipment).
  • each group of sample data in the sample set may include control guidance data and control parameter values corresponding to the control guidance data.
  • the first data and the second data are collected for the same unmanned driving device. The data.
  • step 202 the first data is input to the current convolutional neural network to obtain a hidden variable, and the hidden variable is used to represent the conversion influence factor between the control guidance data and the control parameter value.
  • step 203 the control guidance data corresponding to the hidden variable and the second data is input to the current cyclic neural network to obtain the predicted parameter value.
  • the first data can be input into the current convolutional neural network to obtain the hidden variables output by the convolutional neural network.
  • the control guidance data corresponding to the hidden variable and the second data can be input to the current cyclic neural network to obtain the predicted parameter value output by the cyclic neural network.
  • step 204 based on the predicted parameter value and the control parameter value corresponding to the second data, it is determined whether the preset condition is currently met.
  • the objective function it can be determined whether the objective function converges based on the predicted parameter value and the control parameter value corresponding to the second data.
  • the objective function may be an ELBO (Evidence lower bound) evidence offline function between the predicted parameter value and the control parameter value corresponding to the second data.
  • ELBO Exposure lower bound
  • the distribution of the predicted parameter values and the distribution of the control parameter values corresponding to the second data obey the normal distribution, and the predicted parameter values corresponding to the second data can be obtained according to the definition formula of ELBO and the maximum likelihood estimation method
  • the objective function can also be any other reasonable function, which is not limited in this application.
  • step 205 if the preset conditions are not met, the network parameters of the aforementioned convolutional neural network and the aforementioned cyclic neural network are adjusted, and steps 201-204 are executed again.
  • the network parameters of the aforementioned convolutional neural network and the aforementioned cyclic neural network can be adjusted.
  • the adjustment direction of the network parameters of the convolutional neural network and the recurrent neural network can be determined according to the predicted parameter value and the control parameter value corresponding to the second data (for example, increase the parameter or adjust the parameter Small), and then adjust the network parameters of the convolutional neural network and the cyclic neural network according to the adjustment direction. Therefore, after adjustment, the difference between the predicted parameter value and the control parameter value corresponding to the second data is minimized.
  • step 206 if the preset conditions are met, the adjusted convolutional neural network and the adjusted cyclic neural network are output.
  • the currently adjusted convolutional neural network and the recurrent neural network may be output as the target convolutional neural network and the target recurrent neural network.
  • the target convolutional neural network and target cyclic neural network trained in the above manner can be used for unmanned driving control.
  • the target latent variable corresponding to the target device can be obtained based on the target convolutional neural network first, and the current control guidance data of the target device can be determined. Then, the current control guidance data and target hidden variables are input into the target cyclic neural network, and the result of the target cyclic neural network is obtained as the current control parameter value. Finally, the target can be controlled according to the current control parameter value.
  • the target hidden variable can be determined in the following way: First, determine multiple sets of data collected for the target device, each set of data includes control guidance data and control parameter values. Then, multiple sets of data are input to the target convolutional neural network, and the output result of the target convolutional neural network is obtained as the target hidden variable.
  • the method for training a model for controlling an unmanned driving device performs the following target operations: selecting multiple sets of sample data from a sample set as the first data, and selecting a set of sample data as the second data, Each set of sample data includes control guidance data and control parameter values.
  • the first data is input into the current convolutional neural network to obtain a hidden variable, which is used to represent the conversion influence factor between the control guidance data and the control parameter value.
  • Input the control guidance data corresponding to the hidden variable and the second data to the current cyclic neural network to obtain the predicted parameter value.
  • the network parameters of the current convolutional neural network and cyclic neural network are adjusted, and the target operation is performed again. If it is determined that the foregoing preset condition is satisfied based on the predicted parameter value and the control parameter value corresponding to the second data, the adjusted target convolutional neural network and the target cyclic neural network are output.
  • this embodiment introduces hidden variables that represent the influencing factors of the conversion between the control guidance data and the control parameter values, and at the same time trains the convolutional neural network used to construct the hidden variables and the recurrent neural network used to predict the value of the control parameter , So that when the trained target convolutional neural network and target cyclic neural network are applied to unmanned driving control, the control parameter values obtained are more accurate.
  • a randomly generated noise signal can also be input to the current cyclic neural network.
  • the cyclic neural network is guaranteed to ensure that the first data and the second data are data collected for the same unmanned driving device.
  • unmanned devices of the same model usually have certain commonalities, and unmanned devices of the same model can usually be classified into one category for sample data collection and model training.
  • the model obtained by training in the above-mentioned manner can only reflect the common characteristics of the unmanned equipment of the same model.
  • each driverless device has its own unique characteristics. Therefore, different driverless devices of the same model also have different characteristics.
  • each round of training multiple sets of sample data can be selected from the sample set as the first data, and one set of sample data can be selected as the second data.
  • the first data and the second data are data collected for the same unmanned driving device, so that each round of training corresponds to the same unmanned driving device.
  • a noise signal can be randomly generated, and the noise signal, the hidden variable and the control guidance data corresponding to the second data, are input into the current cyclic neural network together. Since each round of training corresponds to the same unmanned device, and random noise signals are introduced in each round of training (random noise signals can provide preset degrees of freedom for the cyclic neural network), different rounds of training correspond to different unmanned devices. Therefore, the target convolutional neural network that is finally trained can obtain hidden variables that reflect its unique characteristics for each unmanned device.
  • the randomly generated one can provide preset degrees of freedom for the cyclic neural network.
  • the noise signal of is also input to the current cyclic neural network, and the first data and the second data are data collected for the same unmanned driving device. Therefore, there is no need to collect a large amount of sample data for each unmanned device, so that the trained target convolutional neural network can obtain the hidden variables reflecting its unique characteristics for each unmanned device, and improve the model training Efficiency and accuracy.
  • this application also provides an embodiment of a device for controlling unmanned equipment and model training.
  • FIG. 3 is a block diagram of an apparatus for controlling an unmanned driving device according to an exemplary embodiment of the present application.
  • the apparatus may include: an acquisition module 301, a determination module 302, and a control module 303.
  • the obtaining module 301 is used to obtain the current control guidance data of the target device, and obtain a predetermined target hidden variable corresponding to the target device.
  • the target hidden variable is used to indicate the conversion influence factor between the control guidance data and the control parameter value for the target device.
  • the determination module 302 is used to obtain the current control parameter value based on the current control guidance data and the target hidden variable.
  • the control module 303 is used to control the target device according to the current control parameter value.
  • the target latent variable may be determined in advance by the following method: determining multiple sets of data collected for the target device, and each set of data includes control guidance data and control parameter values. And input multiple sets of data into the target convolutional neural network to obtain the target hidden variables output by the target convolutional neural network.
  • the determining module 302 is configured to input the current control guidance data and the target latent variable into the target cyclic neural network to obtain the current control parameter value.
  • the target convolutional neural network and the target recurrent neural network are trained by the following method: perform the following target operation: select multiple sets of sample data from the sample set as the first data, and select a set of sample data As the second data, each set of sample data includes control guidance data and control parameter values.
  • the first data is input to the current convolutional neural network to obtain hidden variables, which are used to represent the conversion influence factors between the control guidance data and the control parameter values.
  • the network parameters of the convolutional neural network and the cyclic neural network are adjusted, and the target operation is performed again. If it is determined that the preset condition is satisfied based on the predicted parameter value and the control parameter value corresponding to the second data, the target convolutional neural network and the target cyclic neural network are output.
  • the first data and the second data are data collected for the same unmanned driving device.
  • the target operation may also include: inputting the latent variable, the control guidance data corresponding to the second data, and the randomly generated noise signal into the current cyclic neural network.
  • the objective function is determined, and the objective function is the ELBO evidence offline function between the predicted parameter value and the control parameter value corresponding to the second data.
  • the objective function converges, it is determined that the preset conditions are met.
  • the above-mentioned device may be pre-installed in the unmanned driving equipment, or may be loaded into the unmanned driving equipment through downloading or the like.
  • the corresponding modules in the above device can cooperate with the modules in the unmanned driving equipment to realize the unmanned driving control scheme.
  • FIG. 4 is a block diagram of a model training apparatus for controlling unmanned driving equipment according to an exemplary embodiment of the present application.
  • the apparatus may include: an execution module 401, an adjustment module 402, and an output module 403. .
  • the execution module 401 is configured to perform the following target operations: selecting multiple sets of sample data from the sample set as the first data, and selecting a set of sample data as the second data, each set of sample data includes control guidance data and control parameter values. Input the first data into the current convolutional neural network to obtain hidden variables, which are used to represent the conversion influence factors between the control guidance data and the control parameter values, and input the control guidance data corresponding to the hidden variables and the second data To the current cyclic neural network, get the predicted parameter value.
  • the adjustment module 402 is configured to adjust the network parameters of the aforementioned convolutional neural network and the aforementioned cyclic neural network when it is determined that the preset condition is not met based on the predicted parameter value and the control parameter value corresponding to the second data, and instruct the execution module 401 re-execute the target operation.
  • the output module 403 is configured to output the target convolutional neural network and the target cyclic neural network when it is determined that the preset condition is satisfied based on the predicted parameter value and the control parameter value corresponding to the second data.
  • the first data and the second data are data collected for the same unmanned driving device.
  • the execution module 401 is further configured to: input the hidden variables, the control guidance data corresponding to the second data, and the randomly generated noise signal into the current cyclic neural network.
  • the above-mentioned device may be pre-installed in the terminal device or server, or loaded into the terminal device or server through downloading or the like.
  • the corresponding modules in the above-mentioned devices can cooperate with the modules in the unmanned driving equipment to realize a model training scheme for controlling the unmanned driving equipment.
  • the relevant part can refer to the part of the description of the method embodiment.
  • the device embodiments described above are merely illustrative.
  • the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in One place, or it can be distributed to multiple network units.
  • Some or all of the modules can be selected according to actual needs to achieve the objectives of the solution of this application. Those of ordinary skill in the art can understand and implement it without creative work.
  • the embodiment of the present application also provides a computer-readable storage medium, the storage medium stores a computer program, and the computer program can be used to execute the method for controlling an unmanned driving device and model training provided by any one of the embodiments in FIG. 1 to FIG. 2 .
  • an embodiment of the present application also proposes a schematic structural diagram of an electronic device according to an exemplary embodiment of the present application shown in FIG. 5.
  • the electronic device includes a processor, an internal bus, a network interface, a memory, and a non-volatile memory.
  • the processor reads the corresponding computer program from the non-volatile memory to the memory and then runs it to form a device for controlling the unmanned equipment at the logical level.
  • this application does not exclude other implementations, such as logic devices or a combination of software and hardware, etc. That is to say, the execution body of the following processing flow is not limited to each logic unit, and can also be Hardware or logic device.
  • an embodiment of the present application also proposes a schematic structural diagram of an electronic device according to an exemplary embodiment of the present application shown in FIG. 6.
  • the electronic device includes a processor, an internal bus, a network interface, a memory, and a non-volatile memory.
  • the processor reads the corresponding computer program from the non-volatile memory to the memory and then runs it to form a model training device for controlling the unmanned equipment at the logical level.
  • this application does not exclude other implementations, such as logic devices or a combination of software and hardware, etc. That is to say, the execution body of the following processing flow is not limited to each logic unit, and can also be Hardware or logic device.

Landscapes

  • Engineering & Computer Science (AREA)
  • Health & Medical Sciences (AREA)
  • Artificial Intelligence (AREA)
  • Computer Vision & Pattern Recognition (AREA)
  • Evolutionary Computation (AREA)
  • Medical Informatics (AREA)
  • Software Systems (AREA)
  • Physics & Mathematics (AREA)
  • General Physics & Mathematics (AREA)
  • Automation & Control Theory (AREA)
  • Feedback Control In General (AREA)
  • Control Of Position, Course, Altitude, Or Attitude Of Moving Bodies (AREA)

Abstract

一种控制无人驾驶设备及模型训练的方法、装置及电子设备,方法包括:获取目标设备当前的控制指导数据(101);获取预先确定的目标设备对应的目标隐变量;目标隐变量用于表示针对目标设备,控制指导数据与控制参数值之间的转化影响因素(102);基于当前的控制指导参数及目标隐变量,得到当前的控制参数值(103);根据当前的控制参数值控制目标设备(104)。

Description

控制无人驾驶设备及训练模型 技术领域
本申请涉及无人驾驶技术领域,特别涉及如何控制无人驾驶设备及训练模型。
背景技术
在无人驾驶技术中,在确定了驾驶决策后,可参考驾驶决策得到当前的控制指导数据,并通过查表的方式(查询控制指导数据与控制参数值的对应关系表),得到当前的控制指导数据所对应的控制参数值,以进行无人驾驶控制。
发明内容
为了解决上述技术问题之一,本申请提供一种控制无人驾驶设备及模型训练的方法、装置及电子设备。
根据本申请实施例的第一方面,提供一种控制无人驾驶设备的方法,包括:
获取目标设备当前的控制指导数据;
获取预先确定的所述目标设备对应的目标隐变量;所述目标隐变量用于表示针对所述目标设备,控制指导数据与控制参数值之间的转化影响因素;
基于所述当前的控制指导数据及所述目标隐变量,得到当前的控制参数值;
根据所述当前的控制参数值控制所述目标设备。
根据本申请实施例的第二方面,提供一种用于控制无人驾驶设备的模型训练方法,包括:
执行以下目标操作:从样本集中选多组样本数据作为第一数据,以及选一组样本数据作为第二数据,每组所述样本数据包括控制指导数据及控制参数值;将所述第一数据输入至当前的卷积神经网络,得到隐变量,所述隐变量用于表示控制指导数据与控制参数值之间的转化影响因素;将所述隐变量和所述第二数据对应的控制指导数据输入至当前的循环神经网络,得到预测参数值;
若基于所述预测参数值与所述第二数据对应的控制参数值,确定未满足预设条件,对所述卷积神经网络和所述循环神经网络的网络参数进行调整,并重新执行所述目标操 作;
若基于所述预测参数值与所述第二数据对应的控制参数值,确定满足所述预设条件,输出经过调整后的目标卷积神经网络及目标循环神经网络。
根据本申请实施例的第三方面,提供一种控制无人驾驶设备的装置,包括:
获取模块,用于获取目标设备当前的控制指导数据,并获取预先确定的所述目标设备对应的目标隐变量;所述目标隐变量用于表示针对所述目标设备,控制指导数据与控制参数值之间的转化影响因素;
确定模块,用于基于所述当前的控制指导数据及所述目标隐变量,得到当前的控制参数值;
控制模块,用于根据所述当前的控制参数值控制所述目标设备。
根据本申请实施例的第四方面,提供一种用于控制无人驾驶设备的模型训练装置,包括:
执行模块,用于执行以下目标操作:从样本集中选多组样本数据作为第一数据,以及选一组样本数据作为第二数据,每组所述样本数据包括控制指导数据及控制参数值;将所述第一数据输入至当前的卷积神经网络,得到隐变量,所述隐变量用于表示控制指导数据与控制参数值之间的转化影响因素;将所述隐变量和所述第二数据对应的控制指导数据输入至当前的循环神经网络,得到预测参数值;
调整模块,用于在基于所述预测参数值与所述第二数据对应的控制参数值,确定未满足预设条件时,对所述卷积神经网络和所述循环神经网络的网络参数进行调整,并指示所述执行模块重新执行所述目标操作;
输出模块,用于在基于所述预测参数值与所述第二数据对应的控制参数值,确定满足所述预设条件时,输出经过调整后的目标卷积神经网络及目标循环神经网络。
根据本申请实施例的第五方面,提供一种计算机可读存储介质,所述存储介质存储有计算机程序,所述计算机程序被处理器执行时实现上述第一方面以及第二方面中任一项所述的方法。
根据本申请实施例的第六方面,提供一种电子设备,包括存储器、处理器及存储在存储器上并可在处理器上运行的计算机程序,所述处理器执行所述程序时实现上述第一方面以及第二方面中任一项所述的方法。
应当理解的是,以上的一般描述和后文的细节描述仅是示例性和解释性的,并不能限制本申请。
附图说明
此处的附图被并入说明书中并构成本说明书的一部分,示出了符合本申请的实施例,并与说明书一起用于解释本申请的原理。
图1是本申请根据一示例性实施例示出的一种控制无人驾驶设备的方法的流程图;
图2是本申请根据一示例性实施例示出的一种用于控制无人驾驶设备的模型训练的方法的流程图;
图3是本申请根据一示例性实施例示出的一种控制无人驾驶设备的装置的框图;
图4是本申请根据一示例性实施例示出的一种用于控制无人驾驶设备的模型训练的装置的框图;
图5是本申请根据一示例性实施例示出的一种电子设备的结构示意图;
图6是本申请根据一示例性实施例示出的另一种电子设备的结构示意图。
具体实施方式
这里将详细地对示例性实施例进行说明,其示例表示在附图中。下面的描述涉及附图时,除非另有表示,不同附图中的相同数字表示相同或相似的要素。以下示例性实施例中所描述的实施方式并不代表与本申请相一致的所有实施方式。相反,它们仅是与如所附权利要求书中所详述的、本申请的一些方面相一致的装置和方法的例子。
在本申请使用的术语是仅仅出于描述特定实施例的目的,而非旨在限制本申请。在本申请和所附权利要求书中所使用的单数形式的“一种”、“所述”和“该”也旨在包括多数形式,除非上下文清楚地表示其他含义。还应当理解,本文中使用的术语“和/或”是指并包含一个或多个相关联的列出项目的任何或所有可能组合。
应当理解,尽管在本申请可能采用术语第一、第二、第三等来描述各种信息,但这些信息不应限于这些术语。这些术语仅用来将同一类型的信息彼此区分开。例如,在不脱离本申请范围的情况下,第一信息也可以被称为第二信息,类似地,第二信息也可以被称为第一信息。取决于语境,如在此所使用的词语“如果”可以被解释成为“在…… 时”或“当……时”或“响应于确定”。
在无人驾驶技术中,在确定了驾驶决策后,可参考驾驶决策得到当前的控制指导数据,并通过查表的方式(查询控制指导数据与控制参数值的对应关系表),得到当前的控制指导数据所对应的控制参数值,以进行无人驾驶控制。但是,上述方式需要预先针对每个无人驾驶设备进行大量人工标定,从而得到每个无人驾驶设备所对应的对应关系表,因此,浪费了大量的人力资源。并且,上述对应关系表仅能表示查询控制指导数据与控制参数值之间的离散关系,从而使得控制参数值的误差增大。
如图1所示,图1是根据一示例性实施例示出的一种控制无人驾驶设备的方法的流程图,该方法可以应用于无人驾驶设备中。本领域技术人员可以理解,该无人驾驶设备可以包括但不限于无人车、无人操作机器人、无人机、无人船等等。该方法包括以下步骤:
在步骤101中,获取目标设备当前的控制指导数据。
在步骤102中,获取预先确定的目标设备对应的目标隐变量,该目标隐变量用于表示针对目标设备,控制指导数据与控制参数值之间的转化影响因素。
在步骤103中,基于当前的控制指导数据及目标隐变量,得到当前的控制参数值。
在本实施例中,目标设备为待控制的无人驾驶设备,目标设备可以是无人车,或者是无人操作机器人,或者是无人机,或者是无人船等等,本申请对目标设备的具体类型方面不限定。
在本实施例中,当前的控制参数值为当前用于对目标设备进行控制的参数的值,例如,以无人车为例,控制参数值可以是无人车的底盘控制参数值(如,油门或刹车的控制量等)等。可以理解,控制参数值还可以是其它的控制参数值,本申请对此方面不限定。目标设备当前的控制指导数据可以用于确定当前的控制参数值,例如,以无人车为例,当前的控制指导数据可以包括无人车当前的运行速度值以及当前准备施加的加速度值。可以理解,控制指导数据还可以是其它任意能够用于确定当前的控制参数值的数据,本申请对此方面不限定。
在本实施例中,目标设备对应的目标隐变量能够用于表示控制指导数据与控制参数值之间的转化影响因素,并且,该转化影响因素为针对目标设备的转化影响因素。可以预先基于训练好的目标卷积神经网络,确定该目标隐变量,并将目标隐变量存储于目标设备中,当对目标设备进行无人驾驶控制时,可以从目标设备存储的数据中获取目标隐 变量,并基于目标隐变量进行无人驾驶控制。
在本实施例中,目标卷积神经网络和目标循环神经网络均为预先训练好的模型,具体来说,可以通过如下方式训练得到目标卷积神经网络和目标循环神经网络:执行以下目标操作:从样本集中选多组样本数据作为第一数据,以及选一组样本数据作为第二数据,每组样本数据包括控制指导数据及控制参数值,其中样本集中的样本数据均是针对相同型号的无人驾驶设备采集的数据,第一数据和第二数据是针对同一无人驾驶设备采集的数据。将第一数据输入至当前的卷积神经网络,得到隐变量,该隐变量用于表示控制指导数据与控制参数值之间的转化影响因素。将该隐变量和第二数据对应的控制指导数据输入至当前的循环神经网络,得到预测参数值。若基于预测参数值与第二数据对应的控制参数值,确定未满足预设条件,则对当前的卷积神经网络和循环神经网络的网络参数进行调整,并重新执行目标操作。若基于预测参数值与第二数据对应的控制参数值,确定满足上述预设条件,则输出经过调整后的卷积神经网络及经过调整后的循环神经网络。可将调整后的卷积神经网络作为目标神经网络,并将调整后的循环神经网络作为目标循环神经网络。
进一步地,在执行目标操作过程中,在将隐变量和第二数据对应的控制指导数据输入至当前的循环神经网络的同时,还可以将随机生成的噪声信号也输入至当前的循环神经网络,并且,保证第一数据和第二数据为针对同一无人驾驶设备而采集的数据。
在步骤104中,根据当前的控制参数值控制目标设备。
在本实施例中,可以将当前的控制指导数据及目标隐变量输入至目标循环神经网络,将目标循环神经网络输出的结果作为当前的控制参数值。可以按照当前的控制参数值控制目标设备。
本申请的上述实施例提供的控制无人驾驶设备的方法,通过获取目标设备当前的控制指导数据,及获取预先确定的目标设备对应的目标隐变量,基于当前的控制指导数据及目标隐变量,得到当前的控制参数值,并根据当前的控制参数值控制目标设备。其中,目标隐变量用于表示针对目标设备,控制指导数据与控制参数值之间的转化影响因素。本实施例无需通过大量人工标定得到每个无人驾驶设备所对应的对应关系表,从而节省了大量的人力资源,并减小了控制参数值的误差。
在另一些可选实施方式中,可以预先通过如下方式确定目标设备对应的目标隐变量:可以确定针对目标设备采集的多组数据,每组数据包括控制指导数据及控制参数值,并 将上述多组数据输入目标卷积神经网络,得到目标卷积神经网络输出的目标隐变量。
在本实施例中,首先,可以预先针对目标设备进行驾驶测试,在驾驶测试的过程中,采集多组数据,每组数据可以包括控制指导数据及该控制指导数据对应的控制参数值。驾驶测试过程中采集的数据的组数可以是任意合理的数量,例如,可以是3组,或者5组,或者10组等。可以理解,本申请对驾驶测试过程中采集的数据的具体组数方面不限定。
接着,可以确定预先针对目标设备采集的多组数据,并将上述多组样本输入至目标卷积神经网络中,将目标卷积神经网络输出的结果作为目标隐变量。
由于本实施例中,可以采用预先针对目标设备而采集的多组数据,通过预先训练的目标卷积神经网络,得到能够表示控制指导数据与控制参数值之间转化影响因素的目标隐变量,该目标隐变量为针对目标设备的隐变量。因此,进一步提高了控制参数值的准确度。
如图2所示,图2是根据一示例性实施例示出的一种用于控制无人驾驶设备的模型训练的方法的流程图,该方法可以应用于终端设备中,也可以应用于服务器中。该方法包括以下步骤:
在步骤201中,从样本集中选多组样本数据作为第一数据,以及选一组样本数据作为第二数据,每组样本数据包括控制指导数据及控制参数值。
在本实施例中,首先,可以预先针对型号相同的多个不同的无人驾驶设备进行驾驶测试,在驾驶测试的过程中,采集大量样本数据得到样本集(样本集中的样本数据均对应于相同型号的无人驾驶设备)。其中,样本集中的每组样本数据可以包括控制指导数据及该控制指导数据对应的控制参数值。在进行模型训练时,可以从样本集中选多组样本数据作为第一数据,以及从样本集中选一组样本数据作为第二数据,第一数据和第二数据是针对同一无人驾驶设备而采集的数据。
在步骤202中,将第一数据输入至当前的卷积神经网络,得到隐变量,该隐变量用于表示控制指导数据与控制参数值之间的转化影响因素。
在步骤203中,将该隐变量和第二数据对应的控制指导数据输入至当前的循环神经网络,得到预测参数值。
在本实施例中,首先,可以将第一数据输入至当前的卷积神经网络中,得到该卷积神经网络输出的隐变量。接着,可以将该隐变量和第二数据所对应的控制指导数据输入 至当前的循环神经网络,得到该循环神经网络输出的预测参数值。
在步骤204中,基于该预测参数值与第二数据所对应的控制参数值,确定当前是否满足预设条件。
在本实施例中,可以基于该预测参数值与第二数据所对应的控制参数值,确定目标函数是否收敛,当目标函数收敛时,可以确定当前满足预设条件。当目标函数未收敛时,可以确定当前未满足预设条件。其中,目标函数可以是上述预测参数值与第二数据对应的控制参数值之间的ELBO(Evidence lower bound)证据下线函数。具体来说,预测参数值的分布与第二数据所对应的控制参数值的分布服从正态分布,则可以根据ELBO的定义式以及极大似然估计方法,得到预测参数值与第二数据对应的控制参数值之间的ELBO证据下线函数。可以理解,目标函数还可以是其它任意合理的函数,本申请对此方面不限定。
在步骤205中,若未满足预设条件,则对上述卷积神经网络和上述循环神经网络的网络参数进行调整,并重新执行步骤201-204。
在本实施例中,当确定未满足预设条件时,则可以对上述卷积神经网络和上述循环神经网络的网络参数进行调整。具体来说,可以根据该预测参数值与第二数据所对应的控制参数值,确定上述卷积神经网络和上述循环神经网络的网络参数的调整方向(如,将参数调大,或者将参数调小),然后按照该调整方向调整上述卷积神经网络和上述循环神经网络的网络参数。从而使得调整后,预测参数值与第二数据所对应的控制参数值之间的差异尽可能减小。
在步骤206中,若满足预设条件,则输出经过调整后的卷积神经网络及经过调整后的循环神经网络。
在本实施例中,当确定满足预设条件时,可以输出当前经过调整后的卷积神经网络及循环神经网络作为目标卷积神经网络及目标循环神经网络。
需要说明的是,通过上述方式训练得到的目标卷积神经网络及目标循环神经网络可以用于无人驾驶控制。具体来说,可以首先基于目标卷积神经网络获取目标设备对应的目标隐变量,并确定目标设备当前的控制指导数据。然后,将当前的控制指导数据及目标隐变量输入至目标循环神经网络中,得到目标循环神经网络的结果作为当前的控制参数值。最后,可以按照当前的控制参数值对目标进行设备控制。其中,可以通过如下方式确定目标隐变量:首先,确定针对目标设备采集的多组数据,每组数据包括控制指导 数据及控制参数值。接着,将多组数据输入至目标卷积神经网络,得到目标卷积神经网络输出的结果作为目标隐变量。
本申请的上述实施例提供的用于控制无人驾驶设备的模型训练的方法,执行以下目标操作:从样本集中选多组样本数据作为第一数据,以及选一组样本数据作为第二数据,每组样本数据包括控制指导数据及控制参数值。将第一数据输入至当前的卷积神经网络,得到隐变量,该隐变量用于表示控制指导数据与控制参数值之间的转化影响因素。将该隐变量和第二数据对应的控制指导数据输入至当前的循环神经网络,得到预测参数值。若基于预测参数值与第二数据对应的控制参数值,确定未满足预设条件,则对当前的卷积神经网络和循环神经网络的网络参数进行调整,并重新执行目标操作。若基于预测参数值与第二数据对应的控制参数值,确定满足上述预设条件,则输出经过调整后的目标卷积神经网络及目标循环神经网络。由于本实施例引入了表示控制指导数据与控制参数值之间转化影响因素的隐变量,并同时对用于构建隐变量的卷积神经网络,和用于预测控制参数值的循环神经网络进行训练,使得训练得到的目标卷积神经网络及目标循环神经网络在应用于无人驾驶控制时,所得的控制参数值更加准确。
在另一些可选实施方式中,在目标操作过程中,在将隐变量和第二数据对应的控制指导数据输入至当前的循环神经网络的同时,还可以将随机生成的噪声信号也输入至当前的循环神经网络,并且,保证第一数据和第二数据为针对同一无人驾驶设备而采集的数据。
一般来说,相同型号的无人驾驶设备通常具有一定的共性,通常可以将相同型号的无人驾驶设备归为一类进行样本数据采集并进行模型训练。但是,通过上述方式进行训练得到的模型只能体现相同型号的无人驾驶设备的共性特性。实际上,每个无人驾驶设备又具有其独特的特性,因此,相同型号的不同无人驾驶设备也是具有不同特性的。
在本实施例中,在每轮训练中,可以从样本集中选多组样本数据作为第一数据,以及选一组样本数据作为第二数据。其中,第一数据和第二数据均为针对同一无人驾驶设备而采集的数据,使得每轮训练均对应同一无人驾驶设备。并且,在目标操作过程中,可以随机生成噪声信号,并将噪声信号与隐变量和第二数据对应的控制指导数据一起输入至当前的循环神经网络中。由于每轮训练对应同一无人驾驶设备,并在每轮训练引入随机噪声信号(随机噪声信号可以为循环神经网络提供预设的自由度),不同轮次的训练对应不同的无人驾驶设备。因此,最终训练得到的目标卷积神经网络能够针对每个无人驾驶设备,得到能反映其独特特性的隐变量。
由于本实施例中,在目标操作过程中,在将隐变量和第二数据对应的控制指导数据输入至当前的循环神经网络的同时,还将随机生成的能为循环神经网络提供预设自由度的噪声信号也输入至当前的循环神经网络,并且,第一数据和第二数据为针对同一无人驾驶设备而采集的数据。因此,无需对每个无人驾驶设备进行大量样本数据的采集,即可使训练得到的目标卷积神经网络能够针对每个无人驾驶设备,得到反映其独特特性的隐变量,提高了模型训练的效率以及精确度。
应当注意,尽管在上述实施例中,以特定顺序描述了本申请方法的操作,但是,这并非要求或者暗示必须按照该特定顺序来执行这些操作,或是必须执行全部所示的操作才能实现期望的结果。相反,流程图中描绘的步骤可以改变执行顺序。附加地或备选地,可以省略某些步骤,将多个步骤合并为一个步骤执行,和/或将一个步骤分解为多个步骤执行。
与前述控制无人驾驶设备及模型训练的方法实施例相对应,本申请还提供了控制无人驾驶设备及模型训练的装置的实施例。
如图3所示,图3是本申请根据一示例性实施例示出的一种控制无人驾驶设备的装置框图,该装置可以包括:获取模块301,确定模块302和控制模块303。
其中,获取模块301,用于获取目标设备当前的控制指导数据,并获取预先确定的目标设备对应的目标隐变量。目标隐变量用于表示针对目标设备,控制指导数据与控制参数值之间的转化影响因素。
确定模块302,用于基于当前的控制指导数据及目标隐变量,得到当前的控制参数值。
控制模块303,用于根据当前的控制参数值控制目标设备。
在一些可选实施方式中,可以预先通过如下方式确定目标隐变量:确定针对目标设备采集的多组数据,每组数据包括控制指导数据及控制参数值。并将多组数据输入目标卷积神经网络,得到目标卷积神经网络输出的目标隐变量。
在另一些可选实施方式中,确定模块302被配置用于:将当前的控制指导数据及目标隐变量输入至目标循环神经网络,得到当前的控制参数值。
在另一些可选实施方式中,目标卷积神经网络及目标循环神经网络通过如下方法训练而成:执行以下目标操作:从样本集中选多组样本数据作为第一数据,以及选一组样本数据作为第二数据,每组样本数据包括控制指导数据及控制参数值。将第一数据输入 至当前的卷积神经网络,得到隐变量,该隐变量用于表示控制指导数据与控制参数值之间的转化影响因素。将隐变量和第二数据对应的控制指导数据输入至当前的循环神经网络,得到预测参数值。若基于预测参数值与第二数据对应的控制参数值,确定未满足预设条件,对卷积神经网络和循环神经网络的网络参数进行调整,并重新执行目标操作。若基于预测参数值与第二数据对应的控制参数值,确定满足预设条件,输出目标卷积神经网络及目标循环神经网络。
在另一些可选实施方式中,第一数据和第二数据为针对同一无人驾驶设备而采集的数据。
目标操作还可以包括:将隐变量、第二数据对应的控制指导数据和随机生成的噪声信号输入至当前的循环神经网络。
在另一些可选实施方式中,可以通过如下方式确定满足预设条件:
确定目标函数,目标函数为上述预测参数值与第二数据对应的控制参数值之间的ELBO证据下线函数。当目标函数收敛时,确定满足预设条件。
应当理解,上述装置可以预先设置在无人驾驶设备中,也可以通过下载等方式而加载到无人驾驶设备中。上述装置中的相应模块可以与无人驾驶设备中的模块相互配合以实现无人驾驶控制的方案。
如图4所示,图4是本申请根据一示例性实施例示出的一种用于控制无人驾驶设备的模型训练装置框图,该装置可以包括:执行模块401,调整模块402和输出模块403。
其中,执行模块401,用于执行以下目标操作:从样本集中选多组样本数据作为第一数据,以及选一组样本数据作为第二数据,每组样本数据包括控制指导数据及控制参数值。将第一数据输入至当前的卷积神经网络,得到隐变量,该隐变量用于表示控制指导数据与控制参数值之间的转化影响因素,将隐变量和第二数据对应的控制指导数据输入至当前的循环神经网络,得到预测参数值。
调整模块402,用于在基于预测参数值与第二数据对应的控制参数值,确定未满足预设条件时,对上述卷积神经网络和上述循环神经网络的网络参数进行调整,并指示执行模块401重新执行目标操作。
输出模块403,用于在基于预测参数值与第二数据对应的控制参数值,确定满足预设条件时,输出目标卷积神经网络及目标循环神经网络。
在另一些可选实施方式中,第一数据和第二数据为针对同一无人驾驶设备而采集的数据。
执行模块401还用于:将隐变量、第二数据对应的控制指导数据和随机生成的噪声信号输入至当前的循环神经网络。
应当理解,上述装置可以预先设置在终端设备或服务器中,也可以通过下载等方式而加载到终端设备或服务器中。上述装置中的相应模块可以与无人驾驶设备中的模块相互配合以实现用于控制无人驾驶设备的模型训练的方案。
对于装置实施例而言,由于其基本对应于方法实施例,所以相关之处参见方法实施例的部分说明即可。以上所描述的装置实施例仅仅是示意性的,其中所述作为分离部件说明的单元可以是或者也可以不是物理上分开的,作为单元显示的部件可以是或者也可以不是物理单元,即可以位于一个地方,或者也可以分布到多个网络单元上。可以根据实际的需要选择其中的部分或者全部模块来实现本申请方案的目的。本领域普通技术人员在不付出创造性劳动的情况下,即可以理解并实施。
本申请实施例还提供了一种计算机可读存储介质,该存储介质存储有计算机程序,计算机程序可用于执行上述图1至图2任一实施例提供的控制无人驾驶设备及模型训练的方法。
对应于上述的控制无人驾驶设备的方法,本申请实施例还提出了图5所示的根据本申请的一示例性实施例的电子设备的示意结构图。请参考图5,在硬件层面,该电子设备包括处理器、内部总线、网络接口、内存以及非易失性存储器,当然还可能包括其他业务所需要的硬件。处理器从非易失性存储器中读取对应的计算机程序到内存中然后运行,在逻辑层面上形成控制无人驾驶设备的装置。当然,除了软件实现方式之外,本申请并不排除其他实现方式,比如逻辑器件抑或软硬件结合的方式等等,也就是说以下处理流程的执行主体并不限定于各个逻辑单元,也可以是硬件或逻辑器件。
对应于上述的用于控制无人驾驶设备的模型训练方法,本申请实施例还提出了图6所示的根据本申请的一示例性实施例的电子设备的示意结构图。请参考图6,在硬件层面,该电子设备包括处理器、内部总线、网络接口、内存以及非易失性存储器,当然还可能包括其他业务所需要的硬件。处理器从非易失性存储器中读取对应的计算机程序到内存中然后运行,在逻辑层面上形成用于控制无人驾驶设备的模型训练装置。当然,除了软件实现方式之外,本申请并不排除其他实现方式,比如逻辑器件抑或软硬件结合的 方式等等,也就是说以下处理流程的执行主体并不限定于各个逻辑单元,也可以是硬件或逻辑器件。
本领域技术人员在考虑说明书及实践这里公开的发明后,将容易想到本申请的其它实施方案。本申请旨在涵盖本申请的任何变型、用途或者适应性变化,这些变型、用途或者适应性变化遵循本申请的一般性原理并包括本申请未公开的本技术领域中的公知常识或惯用技术手段。说明书和实施例仅被视为示例性的,本申请的真正范围和精神由下面的权利要求指出。
应当理解的是,本申请并不局限于上面已经描述并在附图中示出的精确结构,并且可以在不脱离其范围进行各种修改和改变。本申请的范围仅由所附的权利要求来限制。

Claims (12)

  1. 一种控制无人驾驶设备的方法,包括:
    获取目标设备当前的控制指导数据;
    获取预先确定的所述目标设备对应的目标隐变量;所述目标隐变量用于表示针对所述目标设备,控制指导数据与控制参数值之间的转化影响因素;
    基于所述当前的控制指导数据及所述目标隐变量,得到当前的控制参数值;
    根据所述当前的控制参数值控制所述目标设备。
  2. 根据权利要求1所述的方法,获取预先确定的所述目标设备对应的所述目标隐变量,包括:
    确定针对所述目标设备采集的多组数据,每组所述数据包括控制指导数据及控制参数值;
    将所述多组数据输入目标卷积神经网络,得到所述目标卷积神经网络输出的所述目标隐变量。
  3. 根据权利要求2所述的方法,基于所述当前的控制指导数据及所述目标隐变量,得到所述当前的控制参数值,包括:
    将所述当前的控制指导数据及所述目标隐变量输入至目标循环神经网络,得到所述当前的控制参数值。
  4. 根据权利要求3所述的方法,所述目标卷积神经网络及所述目标循环神经网络通过如下方法训练而成:
    执行以下目标操作:
    从样本集中选多组样本数据作为第一数据;
    从所述样本集中选一组样本数据作为第二数据,每组所述样本数据包括控制指导数据及控制参数值;
    将所述第一数据输入至当前的卷积神经网络,得到隐变量,所述隐变量用于表示控制指导数据与控制参数值之间的转化影响因素;
    将所述隐变量和所述第二数据对应的控制指导数据输入至当前的循环神经网络,得到预测参数值;
    若基于所述预测参数值与所述第二数据对应的控制参数值,确定未满足预设条件,对所述卷积神经网络和所述循环神经网络的网络参数进行调整,并重新执行所述目标操作;
    若基于所述预测参数值与所述第二数据对应的控制参数值,确定满足所述预设条件, 输出所述目标卷积神经网络及所述目标循环神经网络。
  5. 根据权利要求4所述的方法,所述第一数据和所述第二数据为针对同一无人驾驶设备而采集的数据;
    所述目标操作还包括:
    在将所述隐变量和所述第二数据对应的控制指导数据输入至所述当前的循环神经网络的同时,将随机生成的噪声信号也输入至所述当前的循环神经网络。
  6. 根据权利要求4所述的方法,基于所述预测参数值与所述第二数据对应的控制参数值确定满足所述预设条件,包括:
    确定目标函数,所述目标函数为所述预测参数值与所述第二数据对应的控制参数值之间的ELBO证据下线函数;
    当所述目标函数收敛时,确定满足所述预设条件。
  7. 一种用于控制无人驾驶设备的模型训练方法,包括:
    执行以下目标操作:
    从样本集中选多组样本数据作为第一数据;
    从所述样本集中选一组样本数据作为第二数据,每组所述样本数据包括控制指导数据及控制参数值;
    将所述第一数据输入至当前的卷积神经网络,得到隐变量,所述隐变量用于表示控制指导数据与控制参数值之间的转化影响因素;
    将所述隐变量和所述第二数据对应的控制指导数据输入至当前的循环神经网络,得到预测参数值;
    若基于所述预测参数值与所述第二数据对应的控制参数值,确定未满足预设条件,对所述卷积神经网络和所述循环神经网络的网络参数进行调整,并重新执行所述目标操作;
    若基于所述预测参数值与所述第二数据对应的控制参数值,确定满足所述预设条件,输出所述目标卷积神经网络及所述目标循环神经网络。
  8. 根据权利要求7所述的方法,所述第一数据和所述第二数据为针对同一无人驾驶设备而采集的数据;
    所述目标操作还包括:
    在将所述隐变量和所述第二数据对应的控制指导数据输入至所述当前的循环神经网络的同时,将随机生成的噪声信号也输入至所述当前的循环神经网络。
  9. 一种控制无人驾驶设备的装置,包括:
    获取模块,用于获取目标设备当前的控制指导数据,并获取预先确定的所述目标设备对应的目标隐变量;所述目标隐变量用于表示针对所述目标设备,控制指导数据与控制参数值之间的转化影响因素;
    确定模块,用于基于所述当前的控制指导参数数据及所述目标隐变量,得到当前的控制参数值;
    控制模块,用于根据所述当前的控制参数值控制所述目标设备。
  10. 一种用于控制无人驾驶设备的模型训练装置,包括:
    执行模块,用于执行以下目标操作:从样本集中选多组样本数据作为第一数据,以及选一组样本数据作为第二数据,每组所述样本数据包括控制指导数据及控制参数值;将所述第一数据输入至当前的卷积神经网络,得到隐变量,所述隐变量用于表示控制指导数据与控制参数值之间的转化影响因素;将所述隐变量和所述第二数据对应的控制指导数据输入至当前的循环神经网络,得到预测参数值;
    调整模块,用于在基于所述预测参数值与所述第二数据对应的控制参数值,确定未满足预设条件时,对所述卷积神经网络和所述循环神经网络的网络参数进行调整,并指示所述执行模块重新执行所述目标操作;
    输出模块,用于在基于所述预测参数值与所述第二数据对应的控制参数值,确定满足所述预设条件时,输出经过调整后的目标卷积神经网络及目标循环神经网络。
  11. 一种计算机可读存储介质,所述存储介质存储有计算机程序,所述计算机程序被处理器执行时实现上述权利要求1-8中任一项所述的方法。
  12. 一种电子设备,包括存储器、处理器及存储在存储器上并可在处理器上运行的计算机程序,所述处理器执行所述程序时实现上述权利要求1-8中任一项所述的方法。
PCT/CN2019/123394 2019-03-01 2019-12-05 控制无人驾驶设备及训练模型 Ceased WO2020177417A1 (zh)

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
CN201910154597.8 2019-03-01
CN201910154597.8A CN109976153B (zh) 2019-03-01 2019-03-01 控制无人驾驶设备及模型训练的方法、装置及电子设备

Publications (1)

Publication Number Publication Date
WO2020177417A1 true WO2020177417A1 (zh) 2020-09-10

Family

ID=67077636

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/CN2019/123394 Ceased WO2020177417A1 (zh) 2019-03-01 2019-12-05 控制无人驾驶设备及训练模型

Country Status (2)

Country Link
CN (1) CN109976153B (zh)
WO (1) WO2020177417A1 (zh)

Families Citing this family (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN109976153B (zh) * 2019-03-01 2021-03-26 北京三快在线科技有限公司 控制无人驾驶设备及模型训练的方法、装置及电子设备
CN110488821B (zh) * 2019-08-12 2020-12-29 北京三快在线科技有限公司 一种确定无人车运动策略的方法及装置
CN110660103B (zh) * 2019-09-17 2020-12-25 北京三快在线科技有限公司 一种无人车定位方法及装置

Citations (9)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN106503393A (zh) * 2016-11-15 2017-03-15 浙江大学 一种利用仿真生成样本实现无人车自主行进的方法
CN106873566A (zh) * 2017-03-14 2017-06-20 东北大学 一种基于深度学习的无人驾驶物流车
CN107609502A (zh) * 2017-09-05 2018-01-19 百度在线网络技术(北京)有限公司 用于控制无人驾驶车辆的方法和装置
CN108897313A (zh) * 2018-05-23 2018-11-27 清华大学 一种分层式端到端车辆自动驾驶系统构建方法
CN109272108A (zh) * 2018-08-22 2019-01-25 深圳市亚博智能科技有限公司 基于神经网络算法的移动控制方法、系统和计算机设备
CN109299732A (zh) * 2018-09-12 2019-02-01 北京三快在线科技有限公司 无人驾驶行为决策及模型训练的方法、装置及电子设备
US20190061771A1 (en) * 2018-10-29 2019-02-28 GM Global Technology Operations LLC Systems and methods for predicting sensor information
CN109961509A (zh) * 2019-03-01 2019-07-02 北京三快在线科技有限公司 三维地图的生成及模型训练方法、装置及电子设备
CN109976153A (zh) * 2019-03-01 2019-07-05 北京三快在线科技有限公司 控制无人驾驶设备及模型训练的方法、装置及电子设备

Family Cites Families (6)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN106154831B (zh) * 2016-07-25 2018-09-18 厦门大学 一种基于学习法的智能汽车纵向神经滑模控制方法
CN107300863B (zh) * 2017-07-12 2020-01-10 吉林大学 一种基于map图和在线标定的纵向加速度控制方法
CN107463953B (zh) * 2017-07-21 2019-11-19 上海媒智科技有限公司 在标签含噪情况下基于质量嵌入的图像分类方法及系统
CN107703564B (zh) * 2017-10-13 2020-04-14 中国科学院深圳先进技术研究院 一种降雨预测方法、系统及电子设备
CN108056789A (zh) * 2017-12-19 2018-05-22 飞依诺科技(苏州)有限公司 一种生成超声扫描设备的配置参数值的方法和装置
CN108198268B (zh) * 2017-12-19 2020-10-16 江苏极熵物联科技有限公司 一种生产设备数据标定方法

Patent Citations (9)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN106503393A (zh) * 2016-11-15 2017-03-15 浙江大学 一种利用仿真生成样本实现无人车自主行进的方法
CN106873566A (zh) * 2017-03-14 2017-06-20 东北大学 一种基于深度学习的无人驾驶物流车
CN107609502A (zh) * 2017-09-05 2018-01-19 百度在线网络技术(北京)有限公司 用于控制无人驾驶车辆的方法和装置
CN108897313A (zh) * 2018-05-23 2018-11-27 清华大学 一种分层式端到端车辆自动驾驶系统构建方法
CN109272108A (zh) * 2018-08-22 2019-01-25 深圳市亚博智能科技有限公司 基于神经网络算法的移动控制方法、系统和计算机设备
CN109299732A (zh) * 2018-09-12 2019-02-01 北京三快在线科技有限公司 无人驾驶行为决策及模型训练的方法、装置及电子设备
US20190061771A1 (en) * 2018-10-29 2019-02-28 GM Global Technology Operations LLC Systems and methods for predicting sensor information
CN109961509A (zh) * 2019-03-01 2019-07-02 北京三快在线科技有限公司 三维地图的生成及模型训练方法、装置及电子设备
CN109976153A (zh) * 2019-03-01 2019-07-05 北京三快在线科技有限公司 控制无人驾驶设备及模型训练的方法、装置及电子设备

Non-Patent Citations (1)

* Cited by examiner, † Cited by third party
Title
JIN FAN: "Research on End-to-End Decision-Making of Intelligent Vehicle Based on Spatio-Temporal Recurrent Neural Network", CHINESE MASTER’S THESES FULL-TEXT DATABASE, no. 10, 15 October 2018 (2018-10-15), pages 1 - 91, XP055731824, ISSN: 1674-0246 *

Also Published As

Publication number Publication date
CN109976153A (zh) 2019-07-05
CN109976153B (zh) 2021-03-26

Similar Documents

Publication Publication Date Title
US20220363259A1 (en) Method for generating lane changing decision-making model, method for lane changing decision-making of unmanned vehicle and electronic device
EP3628454B1 (en) Methods and apparatus to train interdependent autonomous machines
US10207407B1 (en) Robotic operation libraries
JP7541876B2 (ja) 航空機を制御するためのニューラルネットワークを訓練するためのシステム及び方法
US20240160901A1 (en) Controlling agents using amortized q learning
US12434645B2 (en) Computing systems and methods for generating user-specific automated vehicle actions using artificial intelligence
WO2020177417A1 (zh) 控制无人驾驶设备及训练模型
CN114896168B (zh) 用于自动驾驶算法开发的快速调试系统、方法以及存储器
CN113642243A (zh) 多机器人的深度强化学习系统、训练方法、设备及介质
US11438404B2 (en) Edge computing device for controlling electromechanical system with local and remote task distribution control
CN110103987B (zh) 应用于自动驾驶车辆的决策规划方法和装置
CN112147973A (zh) 检验系统、选择真实测试和测试系统的方法和设备
CN112766452A (zh) 一种双环境粒子群优化方法和系统
CN111077769B (zh) 用于控制或调节技术系统的方法
US20230281277A1 (en) Remote agent implementation of reinforcement learning policies
CN116047934A (zh) 一种无人机集群的实时仿真方法、系统以及电子设备
CN108121347A (zh) 用于控制设备运动的方法、装置及电子设备
CN112528434B (zh) 信息识别方法、装置、电子设备和存储介质
CN113435571A (zh) 实现多任务并行的深度网络训练方法和系统
CN113033806A (zh) 一种训练深度强化学习模型的方法、装置以及调度方法
US20250362687A1 (en) Method for an Optimized Motion Planning of a Robot Device
CN116915825B (zh) 车辆动态自适应通信方法、设备和介质
US20240110825A1 (en) System and method for a model for prediction of sound perception using accelerometer data
CN110705689A (zh) 可区分特征的持续学习方法及装置
JP2019156345A (ja) 車両の評価システム

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 19918300

Country of ref document: EP

Kind code of ref document: A1

NENP Non-entry into the national phase

Ref country code: DE

122 Ep: pct application non-entry in european phase

Ref document number: 19918300

Country of ref document: EP

Kind code of ref document: A1