WO2021052010A1 - 一种人脸朝向估计、网络训练方法、装置、电子设备及存储介质 - Google Patents

一种人脸朝向估计、网络训练方法、装置、电子设备及存储介质 Download PDF

Info

Publication number
WO2021052010A1
WO2021052010A1 PCT/CN2020/104929 CN2020104929W WO2021052010A1 WO 2021052010 A1 WO2021052010 A1 WO 2021052010A1 CN 2020104929 W CN2020104929 W CN 2020104929W WO 2021052010 A1 WO2021052010 A1 WO 2021052010A1
Authority
WO
WIPO (PCT)
Prior art keywords
face orientation
face
map
orientation estimation
estimation network
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/CN2020/104929
Other languages
English (en)
French (fr)
Inventor
张修宝
葛煜坤
沈海峰
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Beijing Didi Infinity Technology and Development Co Ltd
Original Assignee
Beijing Didi Infinity Technology and Development Co Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Beijing Didi Infinity Technology and Development Co Ltd filed Critical Beijing Didi Infinity Technology and Development Co Ltd
Publication of WO2021052010A1 publication Critical patent/WO2021052010A1/zh
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Images

Classifications

    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V40/00Recognition of biometric, human-related or animal-related patterns in image or video data
    • G06V40/10Human or animal bodies, e.g. vehicle occupants or pedestrians; Body parts, e.g. hands
    • G06V40/16Human faces, e.g. facial parts, sketches or expressions
    • G06V40/161Detection; Localisation; Normalisation
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F18/00Pattern recognition
    • G06F18/20Analysing
    • G06F18/21Design or setup of recognition systems or techniques; Extraction of features in feature space; Blind source separation
    • G06F18/214Generating training patterns; Bootstrap methods, e.g. bagging or boosting

Definitions

  • the present disclosure relates to the technical field of face recognition, and in particular to a face orientation estimation, a network training method, device, electronic equipment, and storage medium.
  • Face orientation estimation has a wide range of application scenarios, such as human-computer interaction, driver behavior analysis, 3D face reconstruction, security monitoring, etc. It is an important part of the field of face research. Face orientation estimation refers to determining the Euler angle of the face given a face picture, which are yaw, pitch, and roll.
  • face orientation estimation methods are mainly divided into two categories, one is the face orientation estimation method based on key points, and the other is the face orientation estimation method based on the face.
  • the face orientation estimation method based on key points needs to be completed in two steps. The first step is to detect the key points of the face, and the second step is to estimate the face orientation according to the positions of the key points.
  • the purpose of the present disclosure is to provide a face orientation estimation, a network training method, a device, an electronic device, and a storage medium, to solve the problem of large face orientation estimation errors in the prior art, and to improve the accuracy of face orientation estimation Sex.
  • the present disclosure provides a method for estimating face orientation, the method including:
  • the face orientation estimation network is used to process the face image to be detected to determine the face orientation of the face image to be detected; wherein a position map is set in the hidden layer of the face orientation estimation network, and
  • the face orientation estimation network is a network obtained by supervising and training a sample face image according to the position map, and the position map is: a two-dimensional representation of the complete three-dimensional facial shape of the sample face image in a three-dimensional modeling space;
  • the face orientation of the face image to be detected is output.
  • the face orientation includes: a face orientation deflection angle, a face orientation pitch angle, and a face orientation roll angle.
  • the present disclosure also provides a method for training a face orientation estimation network, the method includes:
  • the face orientation estimation network is generated.
  • the obtaining a corresponding location map according to a plurality of sample face images includes:
  • the position map regression network PRN is used to process the sample face image to obtain a position map corresponding to the sample face image.
  • said performing supervised training on corresponding sample face images according to multiple position maps to determine the face orientation of each of said sample face images includes:
  • the performing supervised training on the feature maps of the same size according to the position map includes:
  • Iterative training is performed on the face orientation network according to the loss function.
  • the determining the regression loss of the face orientation estimation network according to the estimated face orientation of the sample face image and the preset face orientation of the sample face image includes:
  • the L1 loss function is used to determine the regression loss of the face orientation estimation network.
  • the determining the position offset loss of the face toward the estimation network according to the position maps of multiple sizes and the feature maps of corresponding sizes includes:
  • the determining the position offset loss of the face toward the estimation network according to the position maps of multiple sizes and the feature maps of corresponding sizes includes:
  • a mean square error algorithm is used to determine the position offset loss of the face toward the estimation network.
  • the feature map of each size includes: the feature map of multiple channels; and the pixel values of the position map of the multiple sizes and the feature map of the corresponding size are normalized ,include:
  • Normalization processing is performed on the pixel values of the position map of multiple sizes and the single-channel feature map of each size.
  • the determining the loss function of the face orientation estimation network according to the regression loss and the position offset loss includes:
  • the loss function of the face orientation estimation network is determined.
  • the face orientation includes: a face orientation deflection angle, a face orientation pitch angle, and a face orientation roll angle.
  • the present disclosure also provides a face orientation estimation device, characterized in that the device includes: an acquisition module, a determination module, and an output module, wherein:
  • the acquisition module is used to acquire a face image to be detected
  • the determining module is configured to process the face image to be detected by using a face orientation estimation network to determine the face orientation of the face image to be detected; wherein the face orientation estimation network is in the hidden layer A position map is provided, the face orientation estimation network is a network obtained by supervising and training sample face images according to the position map, and the position map is: the sample face image is completely three-dimensional in a three-dimensional modeling space Two-dimensional representation of facial shape;
  • the output module is used to output the face orientation of the face image to be detected.
  • the present disclosure also provides a training device for a face orientation estimation network.
  • the device includes: an acquisition module, a determination module, and a generation module, wherein:
  • the acquiring module is configured to acquire a corresponding location map according to a plurality of sample face images
  • the determining module is configured to perform supervised training on corresponding sample face images according to multiple position maps, and determine the face orientation of each of the sample face images;
  • the generating module is configured to generate the face orientation estimation network according to the supervised training result.
  • the acquisition module is specifically configured to use a position map regression network PRN to process the sample face image to obtain a position map corresponding to the sample face image.
  • the device further includes: a processing module, wherein:
  • the generating module is specifically configured to use the face orientation estimation network to process the sample face image, and generate a corresponding feature map in each intermediate layer;
  • the processing module is configured to use the face orientation estimation network to perform multiple down-sampling processing on the position map to obtain position maps of multiple sizes, so that the feature map of each size has a corresponding size The same location map;
  • the acquiring module is specifically configured to perform supervised training on the feature maps of the same size according to the location map.
  • the determining module is specifically configured to use the face orientation estimation network to process the sample face image to obtain the estimated face orientation of the sample face image; according to the sample face image
  • the estimated face orientation and the preset face orientation of the sample face image are determined to determine the regression loss of the face orientation estimation network; according to the position map of multiple sizes and the feature map of the corresponding size, determine The position offset loss of the face orientation estimation network; the loss function of the face orientation estimation network is determined according to the regression loss and the position offset loss; the face orientation network is determined according to the loss function Perform iterative training.
  • the determining module is specifically configured to determine the face orientation estimation network by using an L1 loss function according to the estimated face orientation of the sample face image and the preset face orientation of the sample face image The return loss.
  • the processing module is specifically configured to perform normalization processing on the pixel values of the position maps of multiple sizes and the feature maps of corresponding sizes;
  • the determining module is specifically configured to calculate the position offset loss according to the normalized position map and the pixel value of the feature map.
  • the determining module is specifically configured to use a mean square error algorithm to determine the position offset loss of the face orientation estimation network according to the position maps of multiple sizes and the feature maps of corresponding sizes.
  • the processing module is specifically configured to average the pixel values of the same position point in the feature map of multiple channels to obtain the pixel value of each position point in the feature map, and perform the average processing
  • the feature map of is used as a single-channel feature map of each size; the pixel values of the position map of multiple sizes and the single-channel feature map of each size are normalized.
  • the determining module is specifically configured to determine the loss function of the face orientation estimation network according to the regression loss, the position offset loss, and a preset weight.
  • the present disclosure also provides an electronic device, including: a processor, a storage medium, and a bus.
  • the storage medium stores machine-readable instructions executable by the processor.
  • the processor communicates with the storage medium through a bus, and the processor executes the machine-readable instructions to execute the steps of the method described in the first aspect or the second aspect.
  • the present disclosure also provides a storage medium having a computer program stored on the storage medium, and the computer program executes the steps of the method described in the first aspect or the second aspect when the computer program is run by a processor.
  • Fig. 1 shows a schematic structural diagram of a face orientation estimation system provided by an embodiment of the present disclosure
  • FIG. 2 shows a schematic flowchart of a method for estimating face orientation according to an embodiment of the present disclosure
  • Fig. 3 shows a schematic flow chart of a method for training a face orientation estimation network provided by an embodiment of the present disclosure
  • FIG. 4 shows a schematic flowchart of a method for training a face orientation estimation network provided by another embodiment of the present disclosure
  • FIG. 5 shows a schematic flowchart of a method for training a face orientation estimation network provided by another embodiment of the present disclosure
  • FIG. 6 shows a schematic structural diagram of a face orientation estimation apparatus provided by an embodiment of the present disclosure
  • FIG. 7 shows a schematic structural diagram of a training device for a face orientation estimation network provided by an embodiment of the present disclosure
  • FIG. 8 shows a schematic structural diagram of a training device for a face orientation estimation network provided by an embodiment of the present disclosure
  • Fig. 9 shows a schematic structural diagram of a face orientation estimation device provided by an embodiment of the present disclosure.
  • the following implementation manners are given in combination with the driver behavior analysis of the driving service in a specific application scenario.
  • the general principles defined here can be applied to other embodiments and application scenarios, such as: human-computer interaction, 3D face reconstruction, security monitoring Wait.
  • the present disclosure mainly focuses on the driver behavior analysis of the driver's driving service, it should be understood that this is only an exemplary embodiment, and the present disclosure can be applied to various scenarios in which the orientation of a person's face needs to be estimated.
  • Face orientation estimation refers to the determination of the orientation of the face in the face image given a face image, that is, the Euler angle of the face: yaw, pitch, and roll .
  • One aspect of the present disclosure relates to a face orientation estimation system.
  • the system can first obtain the face image to be detected, and then process the face image to be detected according to the pre-trained face orientation estimation network to determine the face orientation of the face image to be detected, for example: "The current face orientation is left "Side”, etc., are not restricted here.
  • the prior art usually adopted the method of estimating the face orientation based on key points or the method of estimating the face orientation based on the face, but these two methods are used when the face is deflected at a large angle.
  • the estimation error of is very large, and it is impossible to accurately estimate the face orientation in the face image to be detected.
  • the face orientation estimation method provided by the present disclosure can process the sample face image by using the face orientation estimation network with the position map set in the hidden layer during the training process, where the position map is used to process the sample face image. Images are trained for supervised training. Since a position map is set in the hidden layer during the training process of the face orientation estimation network, the network will be guided to focus on the orientation related information, thereby optimizing the learning process of the deep network. Therefore, With the solution proposed in the present disclosure, more accurate face orientation recognition results can be obtained, thereby solving the problem of large errors in face orientation estimation in the prior art, and achieving the effect of improving face recognition.
  • FIG. 1 is a schematic structural diagram of a face orientation estimation system 100 provided by an embodiment of the present disclosure.
  • the face orientation estimation system 100 can be used for transportation services such as taxis, car rental services, express cars, carpooling, bus services, or shuttle services, or security monitoring, etc., which involve the need to estimate the face orientation. Any platform or scene of the smart platform or smart device.
  • the face orientation estimation system 100 may include one or more of a server 110, a network 120, a service terminal 130, and a database 140.
  • the server 110 may include a processor.
  • the processor may process information and/or data related to the service request to perform one or more functions described in the present disclosure. For example, the processor may determine the user's intention based on a service request obtained from the service terminal 130.
  • the processor may include one or more processing cores (e.g., a single-core processor (S) or a multi-core processor (S)).
  • the processor may include a central processing unit (CPU), an application specific integrated circuit (ASIC), an application specific instruction-set processor (ASIP), and a graphics processing unit (Graphics Processing Unit, GPU), Physical Processing Unit (Physics Processing Unit, PPU), Digital Signal Processor (Digital Signal Processor, DSP), Field Programmable Gate Array (Field Programmable Gate Array, FPGA), Programmable Logic Device ( Programmable Logic Device (PLD), controller, microcontroller unit, Reduced Instruction Set Computing (RISC), or microprocessor, etc., or any combination thereof.
  • CPU central processing unit
  • ASIC application specific integrated circuit
  • ASIP application specific instruction-set processor
  • GPU Graphics Processing Unit
  • PPU Physical Processing Unit
  • DSP Digital Signal Processor
  • DSP Digital Signal Processor
  • FPGA Field Programmable Gate Array
  • PLD Programmable Logic Device
  • controller microcontroller unit
  • RISC Reduced Instruction Set Computing
  • the device type corresponding to the kiosk 130 may be a mobile device, for example, it may include a smart home device, a wearable device, a smart mobile device, a virtual reality device, or an augmented reality device, etc., or it may be a tablet computer or a laptop. PCs, or built-in smart devices in motor vehicles, etc.
  • the service terminal 130 can be a mobile phone with camera function installed in the front of the driver's seat of the vehicle. The installation location of the service terminal 130 only needs to be able to collect a complete facial image of the driver.
  • the driver’s face image is collected as the face image to be detected through the camera function of the mobile phone, and the face orientation estimation network is used to process the face image to be detected to determine the face orientation of the face image to be detected. And output the face orientation of the face image to be detected, which is used to analyze the driver's behavior according to the output face orientation.
  • the database 140 may be connected to the network 120 to communicate with one or more components of the face orientation estimation system 100 (for example, the server 110, the service terminal 130, the service provider 140, etc.). One or more components in the face orientation estimation system 100 can access data or instructions stored in the database 140 via the network 120. In some embodiments, the database 140 may be directly connected to one or more components of the face orientation estimation system 100, or the database 140 may also be a part of the server 110.
  • the following describes in detail the face orientation estimation method provided by the embodiment of the present disclosure with reference to the content described in the face orientation estimation system 100 shown in FIG. 1.
  • the following face orientation estimation method is applied to the above system, and the execution subject It can be a service terminal or a server, the preset scene can be designed and adjusted according to user needs, and any scene involving the need to estimate the face orientation can be used, and it is not limited to the two scenes given in the embodiment.
  • FIG. 2 it is a schematic flow chart of a method for estimating face orientation according to an embodiment of the present disclosure.
  • the method can be executed by a server or a service terminal in the system for estimating face orientation, including:
  • the server may receive the face image to be detected sent by the service terminal; if the method is executed by the service terminal, the service terminal may receive the user's information through the collection interface of the service terminal. The acquired face image to be detected.
  • the service terminal may be: a terminal installed with a human face orientation estimation application.
  • the service terminal can be a smart device, such as any smart device with image processing functions, such as a smart phone, a smart camera, and a smart tablet.
  • the server may have a server device corresponding to the facial orientation estimation application.
  • the service terminal can receive the face image to be detected obtained by the user through the collection interface of the face orientation estimation application when the face orientation estimation application is in a power-on or standby state.
  • the acquired facial image to be detected may be a certain frame of image in the video captured by the acquisition interface, or may be a certain photographed image received by the acquisition interface. Of course, it can also be an image acquired in other forms, and the present disclosure does not impose any limitation on this.
  • the face image to be detected may be a two-dimensional face image.
  • S102 Use the face orientation estimation network to process the face image to be detected, and determine the face orientation of the face image to be detected.
  • the face orientation estimation network may be a convolutional neural network for face orientation estimation
  • the backbone network may be a deep convolutional neural network structure such as resnet50, but it is not limited to this, and can also be modified according to user needs.
  • the hidden layer of the face orientation estimation network provided by the present disclosure is provided with a position map.
  • the face orientation estimation network is a network obtained by supervising and training sample face images according to the position map.
  • the position map is: the sample face image is built in three dimensions.
  • the location map is used to supervise the training during the training process, and it has nothing to do with the location map during the test and application process.
  • the location map only includes the face information in the sample face image, and does not include the background information in the sample face image.
  • the location map is used to guide the face orientation estimation network to pay attention to the orientation related information, such as the facial features of the face. Information etc.
  • the face orientation estimation network is added to the hidden layer of the face orientation estimation network.
  • the more relevant position map is used as the supervision of network training. It not only supervises the output layer of the deep network, but also supervises the hidden layer of the deep network. It guides the deep network to learn more useful feature maps for human orientation estimation, so as to obtain more information based on the feature maps. Accurate estimation of face orientation angle.
  • the user can use the face orientation estimation application of the service terminal, such as the smart robot application installed on the service terminal, and then obtain the current waiting status through the service terminal through the collection interface of the smart robot application.
  • the service terminal can process the acquired face image to be detected according to the received and acquired face image, where the processing process may be: the service terminal itself processes the face image to determine the face orientation in the face image to be detected; Alternatively, the face image to be detected is sent to the server, and the server determines the orientation of the face in the face image to be detected.
  • the output face orientation of the face image to be detected can be used to analyze the behavior of the driver, for example: the face orientation estimation application of the service terminal receives the returned detection person When the face orientation of the face image is downward, it may be judged that the current driver is in a tired state and driving is at risk, and an alarm signal may be returned to remind the driver to maintain a normal driving posture and drive safely; or the face orientation estimation application of the service terminal When the program continuously receives the returned detected face images and the face orientation is to the left, it may judge that the current driver is in a distracted driving state, driving is risky, and may return an alarm signal to remind the driver to maintain a normal driving posture Safe driving; specifically, after receiving the output of the face orientation of the face image to be detected, the operations performed are not limited to those provided in the foregoing embodiments, and are specifically designed according to user needs, and the present disclosure does not make any restrictions herein.
  • the face orientation estimation network with the position map is set in the hidden layer during the training process to process the face image to be detected; wherein the position map is used to compare the sample person
  • the face image is supervised. Since the position map is set in the hidden layer during the training process of the face orientation estimation network, the network will be guided to focus on more useful information for the face orientation estimation, thereby optimizing the deep network
  • the learning process makes it possible to obtain more accurate face orientation recognition results when the trained face orientation estimation network is used to recognize the face image to be detected.
  • FIG. 3 is a schematic flowchart of a method for training a face orientation estimation network provided by another embodiment of the present disclosure. As shown in FIG. 3, the method may include:
  • the position map is a two-dimensional image of a complete three-dimensional facial shape in the UV space, which records the three-dimensional structural features of the human face.
  • S202 Perform supervised training on corresponding sample face images according to multiple position maps, and determine the face orientation of each sample face image.
  • the network when performing supervised training based on the location map, the network will be guided to learn a feature map that is more useful for face orientation estimation, so that the face orientation of the sample face image obtained is more accurate.
  • S203 Generate a face orientation estimation network according to the result of the supervised training.
  • the face orientation in the current face image can be directly determined according to the face image. Because the generated face orientation estimation network is obtained by supervised training of the position map, it is obtained The accuracy of face orientation is higher.
  • the face orientation estimation network training method provided in this application is used, because in the process of training the face orientation estimation network, the face orientation of each sample face image is supervised and trained according to the position map corresponding to the sample face image
  • the face orientation estimation network is generated according to the result of the supervised training, so that the generated face orientation estimation network is more accurate according to the face orientation result determined by the face image in the application process.
  • the method for obtaining the position map may be, for example, using a Position Map Regression Network (PRN) to process the sample face image to obtain the sample face image correspondence And send the position map to the face orientation estimation network;
  • PRN is an hourglass-shaped convolutional neural network, including convolutional layer and transposed convolutional layer, and the sample face image output in the hidden layer of PRN corresponds to The feature map of, firstly output gradually smaller feature maps through the volume base layer, and then output gradually larger feature maps through the transposed convolution layer, and finally through the transposed convolution layer output feature map and input sample
  • the face image size is the same.
  • the location map is obtained by direct regression from a two-dimensional sample face image through PRN.
  • the feature map corresponding to each size of the sample face image corresponds to a location map of the same size.
  • the input is two
  • the output of the sample face image is the position map corresponding to the sample face image.
  • FIG. 4 is a schematic flowchart of a method for training a face orientation estimation network according to another embodiment of the present disclosure. As shown in FIG. 4, S202 may include:
  • S204 Use the face orientation estimation network to process the sample face image, and generate a corresponding feature map in each intermediate layer.
  • S205 Use the face orientation estimation network to perform multiple down-sampling processing on the position map.
  • the position map of multiple sizes can be obtained after multiple downsampling of the position map, so that each size feature map has a corresponding position map of the same size;
  • the attention shift refers to when the face orientation of the sample face image is recognized, the position map is used to guide the face orientation estimation network to pay attention to the face orientation related information, and to focus on the face orientation related information To train the network.
  • FIG. 5 is a schematic flowchart of a method for training a face orientation estimation network according to another embodiment of the present disclosure. As shown in FIG. 5, S206 may include:
  • S207 Use the face orientation estimation network to process the sample face image to obtain the estimated face orientation of the sample face image.
  • S208 Determine the regression loss of the face orientation estimation network according to the estimated face orientation of the sample face image and the preset face orientation of the sample face image.
  • the face orientation estimation network is used to process the sample face image to obtain the estimated face orientation of the sample face image, and the regression loss is based on the estimated face orientation of the sample face image and The preset face orientation of the sample face image is determined.
  • the regression loss determination method may be, for example, as follows: according to the estimated face orientation of the sample face image and the preset face orientation of the sample face image, the L1 loss function is used to determine the person The face orientation estimates the regression loss of the network.
  • S209 Determine the position offset loss of the human face toward the estimation network according to the position maps of multiple sizes and the feature maps of corresponding sizes.
  • the position offset loss is determined based on the position map of multiple sizes and the feature map of the corresponding size.
  • the method for determining the position offset loss may be, for example, as follows: according to position maps of multiple sizes and feature maps of corresponding sizes, the mean square error algorithm is used to determine the face orientation estimation network. Position offset loss.
  • S210 Determine the loss function of the face orientation estimation network according to the regression loss and the position offset loss.
  • the specific method for determining the loss function may be: determined according to the regression loss, the position offset loss and the preset weight, that is, the regression loss and the position offset loss are respectively equal to the preset weight. After multiplying and adding, the specific method of determining the loss function can be flexibly adjusted according to user needs, and is not limited to the ones given in the foregoing embodiment.
  • S211 Perform iterative training on the face orientation network according to the loss function.
  • the face orientation estimation network uses the face orientation estimation network to process the sample face images, and generates corresponding feature maps in each intermediate layer; adopts the face orientation estimation network , Perform multiple downsampling of the position map to obtain position maps of multiple sizes, so that each size feature map has a corresponding position map of the same size; supervise and train the feature maps of the same size according to the position map, and get Face orientation estimation network.
  • the face orientation estimation network in the present disclosure is based on estimating a large number of sample face images. After repeated training and iteration, a trained face orientation estimation network is finally generated.
  • the iterative process is to estimate the face orientation network according to the loss function. For iterative training, when the loss function meets the preset requirements after multiple iterations, the face orientation estimation network will stop iterating, and the face orientation estimation network will be generated.
  • the face orientation estimation network is generated, which can make the generated face orientation estimation network more accurate in the application process.
  • the feature map of each size includes: feature maps of multiple channels; the method of normalizing the position maps of multiple sizes and the pixel values of the feature maps of each size can be, for example, the feature maps of multiple channels.
  • the pixel value of the same position in the figure is averaged to obtain the pixel value of each position in the feature map, and the averaged feature map is used as a single-channel feature map of each size; position maps of multiple sizes and each The pixel value of the single-channel feature map of the size is normalized.
  • FIG. 4 is a schematic structural diagram of a face orientation estimation device provided by an embodiment of the present disclosure. As shown in FIG. 4, the device includes an acquisition module 301, a determination module 302, and an output module 303, wherein:
  • the obtaining module 301 is used to obtain a face image to be detected.
  • the determining module 302 is configured to process the face image to be detected by using the face orientation estimation network to determine the face orientation of the face image to be detected; wherein a position map is set in the hidden layer of the face orientation estimation network, and the face orientation
  • the estimated network is a network obtained by supervising and training the sample face image according to the location map, and the location map is: the two-dimensional representation of the complete three-dimensional facial shape of the sample face image in the three-dimensional modeling space.
  • the output module 303 is used to output the face orientation of the face image to be detected.
  • FIG. 5 is a schematic structural diagram of a training device for a face orientation estimation network according to an embodiment of the present disclosure. As shown in FIG. 5, the device includes: an acquisition module 401, a determination module 402, and a generation module 403, wherein:
  • the obtaining module 401 is configured to obtain a corresponding location map according to a plurality of sample face images.
  • the determining module 402 is configured to perform supervised training on corresponding sample face images according to multiple position maps, and determine the face orientation of each sample face image.
  • the generating module 403 is used to generate a face orientation estimation network according to the supervised training result.
  • the acquisition module 401 is specifically configured to use a position map regression network PRN to process the sample face image to obtain a position map corresponding to the sample face image.
  • FIG. 6 is a schematic structural diagram of a training device for a face orientation estimation network according to another embodiment of the present disclosure. As shown in FIG. 6, the device further includes a processing module 404, wherein:
  • the generating module 403 is specifically configured to use a face orientation estimation network to process the sample face image, and generate a corresponding feature map in each intermediate layer.
  • the processing module 404 is configured to use the face orientation estimation network to perform multiple down-sampling processing on the position map to obtain position maps of multiple sizes, so that each size feature map has a corresponding position map of the same size.
  • the acquiring module 401 is specifically configured to perform supervised training on feature maps of the same size according to the position map.
  • the determining module 402 is specifically configured to use a face orientation estimation network to process the sample face image to obtain the estimated face orientation of the sample face image; according to the estimated face orientation of the sample face image and the sample face
  • the preset face orientation of the image determines the regression loss of the face orientation estimation network; according to the position map of multiple sizes and the feature map of the corresponding size, the position offset loss of the face orientation estimation network is determined; according to the regression loss and position bias Shift loss, determine the loss function of the face orientation estimation network; perform iterative training on the face orientation network according to the loss function.
  • the determining module 402 is specifically configured to use the L1 loss function to determine the regression loss of the face orientation estimation network according to the estimated face orientation of the sample face image and the preset face orientation of the sample face image.
  • the processing module 404 is specifically configured to perform normalization processing on the pixel values of the position maps of multiple sizes and the feature maps of corresponding sizes.
  • the determining module 402 is specifically configured to calculate the position offset loss according to the normalized position map and the pixel value of the feature map.
  • the determining module 402 is specifically configured to use a mean square error algorithm to determine the position offset loss of the face toward the estimation network according to the position maps of multiple sizes and the feature maps of corresponding sizes.
  • the processing module 404 is specifically configured to average the pixel values of the same position points in the feature maps of multiple channels to obtain the pixel values of each position point in the feature maps, and use the averaged feature maps as each Single-channel feature maps of multiple sizes; normalize the pixel values of position maps of multiple sizes and single-channel feature maps of various sizes.
  • the determining module 402 is specifically configured to determine the loss function of the face orientation estimation network according to the regression loss, the position offset loss and the preset weight.
  • the embodiment of the present disclosure also provides a face orientation estimation device corresponding to the face orientation estimation method. Because the principle of the device in the embodiment of the disclosure to solve the problem is the same as the above-mentioned face orientation estimation method of the embodiment of the disclosure Similar, therefore, the implementation of the device can refer to the implementation of the method, and the repetition of beneficial effects will not be repeated.
  • an embodiment of the present disclosure also provides a face orientation estimation device, including: a processor 601, a memory 602, and a bus 603; the memory 602 stores machine-readable instructions executable by the processor 601.
  • the processor 601 communicates with the memory 602 through the bus 603, and the processor 601 executes machine-readable instructions to execute the steps of the request processing method provided in the foregoing method embodiments during execution.
  • the machine-readable instructions stored in the memory 602 are the execution steps of the request processing method described in the foregoing embodiment of the present disclosure.
  • the processor 601 can execute the request processing method to process the request. Therefore, the electronic device also has All the beneficial effects described in the foregoing method embodiments will not be repeated in this disclosure.
  • the electronic device may be a general-purpose computer or a special-purpose computer, as well as other electronic devices for processing data, etc., all of which can be used to implement the request processing method of the present disclosure.
  • the present disclosure only describes the request processing method through a computer and an electronic device, for convenience, the functions described in the present disclosure may also be implemented in a distributed manner on multiple similar platforms to balance the processing load.
  • the electronic device may include one or more processors for executing program instructions, a communication bus, and different forms of storage media, such as a magnetic disk, ROM, or RAM, or any combination thereof.
  • the computer platform may also include program instructions stored in ROM, RAM, or other types of non-transitory storage media, or any combination thereof. According to these program instructions, the method of the present disclosure can be realized.
  • the electronic device in the present disclosure may also include multiple processors, so the steps performed by one processor described in the present disclosure may also be executed by multiple processors jointly or individually.
  • the embodiments of the present disclosure also provide a storage medium on which a computer program is stored, and when the computer program is run by a processor, the steps of the aforementioned method for estimating the face orientation are executed.
  • the storage medium can be a general storage medium, such as a removable disk, a hard disk, etc.
  • the computer program on the storage medium can execute the above-mentioned face orientation estimation method when it is run, thereby solving the problems in the prior art.
  • the modules described as separate components may or may not be physically separated, and the components displayed as modules may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the objectives of the solutions of the embodiments.
  • the functional units in the various embodiments of the present disclosure may be integrated into one processing unit, or each unit may exist alone physically, or two or more units may be integrated into one unit.
  • the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a non-volatile computer readable storage medium executable by a processor.
  • the technical solution of the present disclosure essentially or the part that contributes to the prior art or the part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including Several instructions are used to make a computer device (which may be a personal computer, a server, or a network device, etc.) execute all or part of the steps of the methods described in the various embodiments of the present disclosure.
  • the aforementioned storage media include: U disk, mobile hard disk, ROM, RAM, magnetic disk or optical disk and other media that can store program codes.
  • the electronic device can first obtain the face image to be detected, and then process the face image to be detected according to the face orientation estimation network to determine the face orientation of the face image to be detected; among them, the face orientation estimation network hides A position map is set in the layer, and the face orientation estimation network is a network obtained by supervising and training sample face images according to the position map; the face orientation of the face image to be detected is output. The accuracy of the obtained face orientation is improved.

Landscapes

  • Engineering & Computer Science (AREA)
  • Theoretical Computer Science (AREA)
  • Physics & Mathematics (AREA)
  • Data Mining & Analysis (AREA)
  • General Physics & Mathematics (AREA)
  • General Engineering & Computer Science (AREA)
  • Life Sciences & Earth Sciences (AREA)
  • Evolutionary Biology (AREA)
  • Evolutionary Computation (AREA)
  • Bioinformatics & Computational Biology (AREA)
  • Bioinformatics & Cheminformatics (AREA)
  • Artificial Intelligence (AREA)
  • Computer Vision & Pattern Recognition (AREA)
  • Health & Medical Sciences (AREA)
  • General Health & Medical Sciences (AREA)
  • Oral & Maxillofacial Surgery (AREA)
  • Human Computer Interaction (AREA)
  • Multimedia (AREA)
  • Image Analysis (AREA)
  • Image Processing (AREA)

Abstract

本公开提供了一种人脸朝向估计、网络训练方法、装置、电子设备及存储介质,其中,该方法包括:获取待检测人脸图像;采用人脸朝向估计网络对所述待检测人脸图像进行处理,确定所述待检测人脸图像的人脸朝向;其中,所述人脸朝向估计网络的隐藏层中设置有位置图,所述人脸朝向估计网络为根据所述位置图对所述样本人脸图像进行监督训练得到的网络,所述位置图为:所述样本人脸图像在三维建模空间中完整三维面部形状的二维表示;输出所述待检测人脸图像的人脸朝向。解决现有技术中人脸朝向估计误差大的问题,达到了得到更准确的人脸朝向识别结果的效果。

Description

一种人脸朝向估计、网络训练方法、装置、电子设备及存储介质
相关申请的交叉引用
本公开要求于2019年09月16日提交中国专利局的申请号为201910870466X、名称为“一种人脸朝向估计方法、装置、电子设备及存储介质”的中国专利申请的优先权,其全部内容通过引用结合在本公开中。
技术领域
本公开涉及人脸识别技术领域,具体而言,涉及一种人脸朝向估计、网络训练方法、装置、电子设备及存储介质。
背景技术
人脸朝向估计有着广泛的应用场景,比如人机交互、驾驶员行为分析、3D人脸重建、安全监控等,它是人脸研究领域的一个重要组成部分。人脸朝向估计是指给定一张人脸图片,确定其中人脸的欧拉角度,分别为偏转角(yaw)、俯仰角(pitch)、滚动角(roll)。
目前,人脸朝向估计的方法主要分为两类,一类是基于关键点的人脸朝向估计方法,另一类是基于面容的人脸朝向估计方法。其中,基于关键点的人脸朝向估计方法需要分两步完成,第一步是检测人脸关键点,第二步是根据关键点的位置来估计人脸朝向。
但是由于人脸在大角度扭转的时候很多关键点在二维图像中是被遮挡而不可见的,使得不可见的关键点的估计误差变大,在这个基础上进行人脸朝向估计的误差也会增大。
发明内容
有鉴于此,本公开的目的在于提供一种人脸朝向估计、网络训练方法、装置、电子设备及存储介质,解决现有技术中人脸朝向估计误差大的问题,提高人脸朝向估计的准确性。
在本公开的第一方面,本公开提供一种人脸朝向估计方法,所述方法包括:
获取待检测人脸图像;
采用人脸朝向估计网络对所述待检测人脸图像进行处理,确定所述待检测人脸图像的人脸朝向;其中,所述人脸朝向估计网络的隐藏层中设置有位置图,所述人脸朝向估计网络为根据所述位置图对样本人脸图像进行监督训练得到的网络,所述位置图为:所述样本人脸图像在三维建模空间中完整三维面部形状的二维表示;
输出所述待检测人脸图像的人脸朝向。
可选地,所述人脸朝向包括:人脸朝向偏转角、人脸朝向仰俯角和人脸朝向滚动角。
第二方面,本公开还提供了一种人脸朝向估计网络的训练方法,所述方法包括:
根据多个样本人脸图像获取对应的位置图;
根据多个位置图对对应的样本人脸图像进行监督训练,确定各所述样本人脸图像的人脸朝向;
根据监督训练的结果,生成所述人脸朝向估计网络。
可选地,所述根据多个样本人脸图像获取对应的位置图,包括:
采用位置图回归网络PRN,对所述样本人脸图像进行处理,得到所述样本人脸图像对应的位置图。
可选地,所述根据多个位置图对对应的样本人脸图像进行监督训练,确定各所述样本人脸图像的人脸朝向,包括:
采用所述人脸朝向估计网络,对所述样本人脸图像进行处理,在每个中间层生成对应的特征图;
采用所述人脸朝向估计网络,对所述位置图进行多次下采样处理,得到多个尺寸的位置图,使得每个尺寸的所述特征图都有对应的尺寸相同的所述位置图;
根据所述位置图对尺寸相同的所述特征图进行监督训练。
可选地,所述根据位置图对尺寸相同的特征图进行监督训练,包括:
采用所述人脸朝向估计网络对所述样本人脸图像进行处理,得到所述样本人脸图像的估计人脸朝向;
根据所述样本人脸图像的估计人脸朝向和所述样本人脸图像的预设人脸朝向,确定所述人脸朝向估计网络的回归损失;
根据多个尺寸的所述位置图和对应尺寸的所述特征图,确定所述人脸朝向估计网络的位置偏移损失;
根据所述回归损失和所述位置偏移损失,确定所述人脸朝向估计网络的损失函数;
根据所述损失函数对所述人脸朝向网络进行迭代训练。
可选地,所述根据所述样本人脸图像的估计人脸朝向和所述样本人脸图像的预设人脸朝向,确定所述人脸朝向估计网络的回归损失,包括:
根据所述样本人脸图像的估计人脸朝向和所述样本人脸图像的预设人脸朝向,采用L1损失函数确定所述人脸朝向估计网络的回归损失。
可选地,所述根据多个尺寸的所述位置图和对应尺寸的所述特征图,确定所述人脸朝向估计网络的位置偏移损失,包括:
对多个尺寸的所述位置图和对应尺寸的所述特征图的像素值进行归一化处理;
根据归一化后的所述位置图和所述特征图的像素值,计算所述位置偏移损失。
可选地,所述根据多个尺寸的所述位置图和对应尺寸的所述特征图,确定所述人脸朝向估计网络的位置偏移损失,包括:
根据多个尺寸的所述位置图和对应尺寸的所述特征图,采用均方误差算法确定所述人脸朝向估计网络的位置偏移损失。
可选地,每个尺寸的所述特征图包括:多个通道的所述特征图;所述对多个尺寸的所述位置图和对应尺寸的所述特征图的像素值进行归一化处理,包括:
对多个通道的所述特征图中同一位置点的像素值进行平均处理,得到所述特征图中各位置点的像素值,并将平均处理后的特征图作为每个尺寸的单通道特征图;
对多个尺寸的所述位置图和各尺寸的所述单通道特征图的像素值进行归一化处理。
可选地,所述根据所述回归损失和所述位置偏移损失,确定所述人脸朝向估计网络的损失函数,包括:
根据所述回归损失、所述位置偏移损失和预设权重,确定所述人脸朝向估计网络的损失函数。
可选地,所述人脸朝向包括:人脸朝向偏转角、人脸朝向仰俯角和人脸朝向滚动角。
第三方面,本公开还提供了一种人脸朝向估计装置,其特征在于,所述装置包括:获取模块、确定模块和输出模块,其中:
所述获取模块,用于获取待检测人脸图像;
所述确定模块,用于采用人脸朝向估计网络对所述待检测人脸图像进行处理,确定所述待检测人脸图像的人脸朝向;其中,所述人脸朝向估计网络的隐藏层中设置有位置图,所述人脸朝向估计网络为根据所述位置图对样本人脸图像进行监督训练得到的网络,所述位置图为:所述样本人脸图像在三维建模空间中完整三维面部形状的二维表示;
所述输出模块,用于输出待检测人脸图像的人脸朝向。
第四方面,本公开还提供了一种人脸朝向估计网络的训练装置,其所述装置包括:获取模块、确定模块和生成模块,其中:
所述获取模块,用于根据多个样本人脸图像获取对应的位置图;
所述确定模块,用于根据多个位置图对对应的样本人脸图像进行监督训练,确定各所述样本人脸图像的人脸朝向;
所述生成模块,用于根据监督训练结果,生成所述人脸朝向估计网络。
可选地,所述获取模块,具体用于采用位置图回归网络PRN,对所述样本人脸图像进行处理,得到所述样本人脸图像对应的位置图。
可选地,所述装置还包括:处理模块,其中:
所述生成模块,具体用于采用所述人脸朝向估计网络,对所述样本人脸图像进行处理, 在每个中间层生成对应的特征图;
所述处理模块,用于采用所述人脸朝向估计网络,对所述位置图进行多次下采样处理,得到多个尺寸的位置图,使得每个尺寸的所述特征图都有对应的尺寸相同的所述位置图;
所述获取模块,具体用于根据所述位置图对尺寸相同的所述特征图进行监督训练。
可选地,所述确定模块,具体用于采用所述人脸朝向估计网络对所述样本人脸图像进行处理,得到所述样本人脸图像的估计人脸朝向;根据所述样本人脸图像的估计人脸朝向和所述样本人脸图像的预设人脸朝向,确定所述人脸朝向估计网络的回归损失;根据多个尺寸的所述位置图和对应尺寸的所述特征图,确定所述人脸朝向估计网络的位置偏移损失;根据所述回归损失和所述位置偏移损失,确定所述人脸朝向估计网络的损失函数;根据所述损失函数对所述人脸朝向网络进行迭代训练。
可选地,所述确定模块,具体用于根据所述样本人脸图像的估计人脸朝向和所述样本人脸图像的预设人脸朝向,采用L1损失函数确定所述人脸朝向估计网络的回归损失。
可选地,所述处理模块,具体用于对多个尺寸的所述位置图和对应尺寸的所述特征图的像素值进行归一化处理;
所述确定模块,具体用于根据归一化后的所述位置图和所述特征图的像素值,计算所述位置偏移损失。
可选地,所述确定模块,具体用于根据多个尺寸的所述位置图和对应尺寸的所述特征图,采用均方误差算法确定所述人脸朝向估计网络的位置偏移损失。
可选地,所述处理模块,具体用于对多个通道的所述特征图中同一位置点的像素值进行平均处理,得到所述特征图中各位置点的像素值,并将平均处理后的特征图作为每个尺寸的单通道特征图;对多个尺寸的所述位置图和各尺寸的所述单通道特征图的像素值进行归一化处理。
可选地,所述确定模块,具体用于根据所述回归损失、所述位置偏移损失和预设权重,确定所述人脸朝向估计网络的损失函数。
第五方面,本公开还提供了一种电子设备,包括:处理器、存储介质和总线,所述存储介质存储有所述处理器可执行的机器可读指令,当电子设备运行时,所述处理器与所述存储介质之间通过总线通信,所述处理器执行所述机器可读指令,以执行如第一方面或第二方面所述方法的步骤。
第六方面,本公开还提供了一种存储介质,所述存储介质上存储有计算机程序,所述计算机程序被处理器运行时执行如第一方面或第二方面所述方法的步骤。
附图说明
为了更清楚地说明本公开实施例的技术方案,下面将对实施例中所需要使用的附图作简单地介绍,应当理解,以下附图仅示出了本公开的某些实施例,因此不应被看作是对范围的限定,对于本领域普通技术人员来讲,在不付出创造性劳动的前提下,还可以根据这些附图获得其他相关的附图。
图1示出了本公开一实施例提供的一种人脸朝向估计系统的结构示意图;
图2示出了本公开一实施例提供的一种人脸朝向估计方法的流程示意图;
图3示出了本公开一实施例提供的一种人脸朝向估计网络的训练方法的流程示意图;
图4示出了本公开另一实施例提供的一种人脸朝向估计网络的训练方法的流程示意图;
图5示出了本公开另一实施例提供的一种人脸朝向估计网络的训练方法的流程示意图;
图6示出了本公开一实施例提供的一种人脸朝向估计装置的结构示意图;
图7示出了本公开一实施例提供的一种人脸朝向估计网络的训练装置的结构示意图;
图8示出了本公开一实施例提供的一种人脸朝向估计网络的训练装置的结构示意图;
图9示出了本公开实施例提供的一种人脸朝向估计设备的结构示意图。
具体实施方式
为使本公开实施例的目的、技术方案和优点更加清楚,下面将结合本公开实施例中的 附图,对本公开实施例中的技术方案进行清楚、完整地描述,应当理解,本公开中附图仅起到说明和描述的目的,并不用于限定本公开的保护范围。另外,应当理解,示意性的附图并未按实物比例绘制。本公开中使用的流程图示出了根据本公开的一些实施例实现的操作。应该理解,流程图的操作可以不按顺序实现,没有逻辑的上下文关系的步骤可以反转顺序或者同时实施。此外,本领域技术人员在本公开内容的指引下,可以向流程图添加一个或多个其他操作,也可以从流程图中移除一个或多个操作。
另外,所描述的实施例仅仅是本公开一部分实施例,而不是全部的实施例。通常在此处附图中描述和示出的本公开实施例的组件可以以各种不同的配置来布置和设计。因此,以下对在附图中提供的本公开的实施例的详细描述并非旨在限制要求保护的本公开的范围,而是仅仅表示本公开的选定实施例。基于本公开的实施例,本领域技术人员在没有做出创造性劳动的前提下所获得的所有其他实施例,都属于本公开保护的范围。
为了使得本领域技术人员能够使用本公开内容,结合特定应用场景代驾服务的驾驶员行为分析,给出以下实施方式。对于本领域技术人员来说,在不脱离本公开的精神和范围的情况下,可以将这里定义的一般原理应用于其他实施例和应用场景,例如:人机交互、3D人脸重建、安全监控等。虽然本公开主要围绕代驾服务的驾驶员行为分析进行描述,但是应该理解,这仅是一个示例性实施例,本公开可以应用于各种需要对人脸朝向进行估计的场景中。
需要说明的是,本公开实施例中将会用到术语“包括”,用于指出其后所声明的特征的存在,但并不排除增加其它的特征。
人脸朝向估计是指给定一张人脸图片,确定该人脸图像中的人脸朝向,即人脸的欧拉角度:偏转角(yaw)、俯仰角(pitch)、滚动角(roll)。
本公开的一个方面涉及一种人脸朝向估计系统。该系统可以先获取待检测人脸图像,再根据预先训练好的人脸朝向估计网络对待检测人脸图像进行处理,确定待检测人脸图像的人脸朝向,例如:“当前人脸朝向为左侧”等,在此不作限制。
值得注意的是,在本公开提出申请之前,现有技术通常通过基于关键点的人脸朝向估计方法,或是基于面容的人脸朝向估计方法,但是这两种方法在大角度人脸偏转时的估计误差都很大,无法准确地估计待检测人脸图像中的人脸朝向。
本公开提供的人脸朝向估计方法,可以通过使用在训练过程中,在隐藏层中设置有位置图的人脸朝向估计网络对样本人脸图像进行处理,其中,位置图用于对样本人脸图像进行监督训练,由于在人脸朝向估计网络的训练过程中在隐藏层中设置有位置图,所以会引导网络将关注点放在朝向相关信息上,从而优化了深度网络的学习过程,因此,通过本公开提出的方案,可以得到更准确的人脸朝向识别结果,从而解决现有技术中人脸朝向估计的误差大的问题,达到提高人脸识别的效果。
图1是本公开实施例提供的一种人脸朝向估计系统100的架构示意图。例如,人脸朝向估计系统100可以是用于诸如出租车、代驾服务、快车、拼车、公共汽车服务、或班车服务之类的运输服务、或是安全监控等任何涉及需要对人脸朝向估计的智能平台或智能设备的任意平台或场景。人脸朝向估计系统100可以包括服务器110、网络120、服务终端130、和数据库140中的一种或多种。
在一些实施例中,服务器110可以包括处理器。处理器可以处理与服务请求有关的信息和/或数据,以执行本公开中描述的一个或多个功能。例如,处理器可以基于从服务终端130获得的服务请求来确定用户意图。在一些实施例中,处理器可以包括一个或多个处理核(例如,单核处理器(S)或多核处理器(S))。仅作为举例,处理器可以包括中央处理单元(Central Processing Unit,CPU)、专用集成电路(Application Specific Integrated Circuit,ASIC)、专用指令集处理器(Application Specific Instruction-set Processor,ASIP)、图形处理单元(Graphics Processing Unit,GPU)、物理处理单元(Physics Processing Unit,PPU)、数字信号处理器(Digital Signal Processor,DSP)、现场可编程门阵列(Field Programmable  Gate Array,FPGA)、可编程逻辑器件(Programmable Logic Device,PLD)、控制器、微控制器单元、简化指令集计算机(Reduced Instruction Set Computing,RISC)、或微处理器等,或其任意组合。
在一些实施例中,服务终端130对应的设备类型可以是移动设备,比如可以包括智能家居设备、可穿戴设备、智能移动设备、虚拟现实设备、或增强现实设备等,也可以是平板计算机、膝上型计算机、或机动车辆中的内置智能设备等。以代驾服务的驾驶员行为分析场景为例,服务终端130可以是安装在车辆驾驶座前端的具有摄像功能的手机,服务终端130的安装位置仅需可以采集到完整的驾驶员人脸图像即可,应用过程中,通过手机的摄像功能采集驾驶员的人脸图像作为待检测人脸图像,采用人脸朝向估计网络对待检测人脸图像进行处理,确定待检测人脸图像的人脸朝向,并将待检测人脸图像的人脸朝向输出,用于根据输出的人脸朝向对驾驶员的行为进行分析。
在一些实施例中,数据库140可以连接到网络120以与人脸朝向估计系统100中的一个或多个组件(例如,服务器110,服务终端130,服务提供端140等)通信。人脸朝向估计系统100中的一个或多个组件可以经由网络120访问存储在数据库140中的数据或指令。在一些实施例中,数据库140可以直接连接到人脸朝向估计系统100中的一个或多个组件,或者,数据库140也可以是服务器110的一部分。
下面结合上述图1示出的人脸朝向估计系统100中描述的内容,对本公开实施例提供的人脸朝向估计方法进行详细说明,下述人脸朝向估计方法应用于上述系统之中,执行主体可以为服务终端或者服务器,预设场景可以根据用户需要设计和调整,任何涉及需要对人脸朝向进行估计的场景均可使用,并不以实施例给出的两个场景为限。
实施例一
参照图2所示,为本公开一实施例提供的一种人脸朝向估计方法的流程示意图,该方法可以由人脸朝向估计系统中的服务器或服务终端来执行,包括:
S101:获取待检测人脸图像。
可选地,若该方法由服务器执行,则该服务器可接收服务终端发送的该待检测人脸图像;若该方法由服务终端执行,则该服务终端可接收用户通过该服务终端的采集界面所获取的该待检测人脸图像。
该服务终端可以为:安装有人脸朝向估计应用程序的终端。该服务终端可以为智能设备,如智能手机、智能摄像机、智能平板等任意具有图像处理功能的智能设备。服务器可以有该人脸朝向估计应用程序对应的服务端设备。
该服务终端可在该人脸朝向估计应用程序处于开机或待机的状态下,接收用户通过该人脸朝向估计应用程序的采集界面所获取的待检测人脸图像。其中,该获取的待检测人脸图像可以为采集界面捕获的视频中的某一帧图像,也可以为采集界面接收的拍摄的某一张图像。当然,也可以为其它形式下获取的图像,本公开不对此进行任何限制。该待检测人脸图像可以为二维人脸图像。
S102:采用人脸朝向估计网络对待检测人脸图像进行处理,确定待检测人脸图像的人脸朝向。
可选地,人脸朝向估计网络可以为人脸朝向估计的卷积神经网络,主干网络可以是resnet50等深度卷积神经网络结构,但并不以此为限,也可以根据用户需要修改。
本公开提供的人脸朝向估计网络的隐藏层中设置有位置图,人脸朝向估计网络为根据位置图对样本人脸图像进行监督训练得到的网络,位置图为:样本人脸图像在三维建模空间中三维面部形状的二维表示。位置图在训练过程中用来监督训练,测试及应用过程中与位置图无关。
其中,位置图中仅包括样本人脸图像中的人脸信息,不包含样本人脸图像中的背景信息,位置图用于引导人脸朝向估计网络去关注朝向相关信息,例如:人脸的五官信息等。
其中,由于人脸图像中包含了很多和人脸朝向无关的冗余信息,比如身份、年龄、性 别、表情等,这些冗余信息会增加人脸朝向估计的难度,而且使得深度学习网络很难收敛。所以在本公开中,为了更准确的估计人脸朝向的角度,优化深度网络的收敛效果,在人脸朝向估计网络的训练过程中,在人脸朝向估计网络的隐藏层中加入了与人脸朝向更相关的位置图作为网络训练的监督,不仅监督深度网络的输出层,同时也监督深度网络的隐藏层,引导深度网络学习出对人间朝向估计更有用的特征图,从而根据特征图得到更准确的人脸朝向角度估计。
例如,在一个目标场景中,用户可在服务终端的人脸朝向估计应用程序,如服务终端所安装的智能机器人应用程序启动之后,通过该智能机器人应用程序的采集界面通过服务终端获取当前的待检测人脸图像,服务终端便可根据接收到获取的待检测人脸图像进行处理,其中,处理过程可以为:由服务终端自身进行处理,从而确定该待检测人脸图像中的人脸朝向;或者,将该待检测人脸图像发送给服务器,由服务器确定该待检测人脸图像中的人脸朝向。
S103:输出待检测人脸图像的人脸朝向。
其中,在本公开的一个实施例中,输出的待检测人脸图像的人脸朝向可以用于对驾驶员的行为进行分析,例如:服务终端的人脸朝向估计应用程序接收到返回的检测人脸图像的人脸朝向为向下时,可能会判断当前驾驶员处于疲惫状态,驾驶存在风险,可能会返回报警信号,提醒驾驶员保持正常驾驶姿势安全驾驶;或者服务终端的人脸朝向估计应用程序连续多次接收到返回的检测人脸图像的人脸朝向为向左时,可能会判断当前驾驶员处于分心驾驶状态,驾驶存在风险,可能会返回报警信号,提醒驾驶员保持正常驾驶姿势安全驾驶;具体接收到输出的待检测人脸图像的人脸朝向后,进行的操作并不局限于上述实施例提供的,具体根据用户需要设计,本公开在此不做任何限制。
采用本公开实施例提供的人脸朝向估计方法,通过在训练过程中在隐藏层中设置有位置图的人脸朝向估计网络,对待检测人脸图像进行处理;其中,位置图用于对样本人脸图像进行监督,由于在人脸朝向估计网络的训练过程中在隐藏层中设置有位置图,所以会引导网络将关注点放在对人脸朝向估计更有用的信息上,从而优化了深度网络的学习过程,使得采用训练好的人脸朝向估计网络对待检测人脸图像进行识别时,可以得到更准确的人脸朝向识别结果。
实施例二
图3为本公开另一实施例提供的人脸朝向估计网络的训练方法的流程示意图,如图3所示,该方法可包括:
S201:根据多个样本人脸图像获取对应的位置图。
其中,位置图为UV空间中完整三维面部形状的二维图像,其记录了人脸的三维结构特征。
S202:根据多个位置图对对应的样本人脸图像进行监督训练,确定各样本人脸图像的人脸朝向。
其中,根据位置图进行监督训练时,会引导网络学习出对人脸朝向估计更有用的特征图,从而得到的样本人脸图像的人脸朝向更加准确。
S203:根据监督训练的结果,生成人脸朝向估计网络。
生成的人脸朝向估计网络在应用过程中,可以直接根据人脸图像确定当前人脸图像中的人脸朝向,由于生成的人脸朝向估计网络是经过位置图进行监督训练得到的,因此得到的人脸朝向的准确度更高。
采用本申请提供的人脸朝向估计网络的训练方法,由于在训练人脸朝向估计网络的过程中,各样本人脸图像的人脸朝向都是根据该样本人脸图像对应的位置图进行监督训练后得到的,并根据监督训练的结果生成人脸朝向估计网络,使得生成的人脸朝向估计网络,在应用过程中根据人脸图像确定的人脸朝向结果更加准确。
可选地,在本公开的一个实施例中,位置图的获取方式例如可以为:采用位置图回归 网络(Position map Regression Network,PRN),对样本人脸图像进行处理,得到样本人脸图像对应的位置图,并将位置图发送至人脸朝向估计网络;PRN是一个沙漏形的卷积神经网络,包括卷积层和转置卷积层,在PRN的隐藏层输出的样本人脸图像对应的特征图,先经过卷基层依次输出逐渐变小的特征图,再经过转置卷积层依次输出逐渐变大的特征图,其最终经过转置卷积层输出的特征图和输入的样本人脸图像尺寸相同。
位置图为通过PRN从二维的样本人脸图像直接回归得到的,其中,每个尺寸的样本人脸图像对应的特征图均对应一个尺寸相同的位置图,在PRN网络中,输入的是二维的样本人脸图像,输出的是样本人脸图像对应的位置图。
实施例三
图4为本公开另一实施例提供的人脸朝向估计网络的训练装置方法的流程示意图,如图4所示,S202可包括:
S204:采用人脸朝向估计网络,对样本人脸图像进行处理,在每个中间层生成对应的特征图。
S205:采用人脸朝向估计网络,对位置图进行多次下采样处理。
其中,对位置图进行多次下采样处理后可以得到多个尺寸的位置图,使得每个尺寸的特征图都有对应的尺寸相同的位置图;
由于人脸朝向估计网络中间层的特征图尺寸和注意图的尺寸不同,因此做注意力迁移时需要对注意图做下采样处理,使其尺寸和被监督的隐藏层的特征图相同。其中,注意力迁移即为在对样本人脸图像的人脸朝向进行识别时,通过位置图来引导人脸朝向估计网络去关注人脸朝向相关信息,将注意力放在人脸朝向相关信息上,从而训练网络。
S206:根据位置图对尺寸相同的特征图进行监督训练。
实施例四
图5为本公开另一实施例提供的人脸朝向估计网络的训练装置方法的流程示意图,如图5所示,S206可包括:
S207:采用人脸朝向估计网络对样本人脸图像进行处理,得到样本人脸图像的估计人脸朝向。
S208:根据样本人脸图像的估计人脸朝向和样本人脸图像的预设人脸朝向,确定人脸朝向估计网络的回归损失。
示例地,在本申请的一个实施中,采用人脸朝向估计网络对样本人脸图像进行处理,得到样本人脸图像的估计人脸朝向,回归损失为根据样本人脸图像的估计人脸朝向和样本人脸图像的预设人脸朝向确定的。
可选地,在本申请的一个实施例中,回归损失的确定方式例如可以为:根据样本人脸图像的估计人脸朝向和样本人脸图像的预设人脸朝向,采用L1损失函数确定人脸朝向估计网络的回归损失。
S209:根据多个尺寸的位置图和对应尺寸的特征图,确定人脸朝向估计网络的位置偏移损失。
其中,位置偏移损失为根据多个尺寸的位置图和对应尺寸的特征图确定的。
可选地,在本申请的一个实施例中,位置偏移损失的确定方式例如可以为:根据多个尺寸的位置图和对应尺寸的特征图,采用均方误差算法确定人脸朝向估计网络的位置偏移损失。
S210:根据回归损失和位置偏移损失,确定人脸朝向估计网络的损失函数。
可选地,在一些可能的实施例中,损失函数的具体确定方式可以为:根据回归损失、位置偏移损失和预设权重确定的,即回归损失与位置偏移损失分别与预设权重相乘后相加,具体损失函数的确定方式可以根据用户需要灵活调整,并不以上述实施例给出的为限。
S211:根据损失函数对人脸朝向网络进行迭代训练。
可选地,人脸朝向估计网络在接收到PRN发送的位置图之后,采用人脸朝向估计网络 对样本人脸图像进行处理,在每个中间层生成对应的特征图;采用人脸朝向估计网络,对位置图进行多次下采样处理,得到多个尺寸的位置图,使得每个尺寸的特征图都有对应的尺寸相同的位置图;根据位置图对尺寸相同的特征图进行监督训练,得到人脸朝向估计网络。
其中,本公开中的人脸朝向估计网络是根据对大量的样本人脸图像进行估计,反复训练迭代后,最终生成训练后的人脸朝向估计网络,迭代过程是根据损失函数对人脸朝向网络进行迭代训练的,在经过多次迭代后,损失函数满足预设要求时,人脸朝向估计网络才会停止迭代,生成人脸朝向估计网络。
多次迭代后生成人脸朝向估计网络,可以使得生成的人脸朝向估计网络在应用过程中,估计的准确性更高。
其中,确定位置偏移损失时,可以对多个尺寸的位置图和对应尺寸的特征图的像素值进行归一化处理,并根据归一化后的位置图和特征图的像素值,计算位置偏移损失。
其中,每个尺寸的特征图包括:多个通道的特征图;对多个尺寸的位置图和各尺寸的特征图的像素值进行归一化处理的方式例如可以为:对多个通道的特征图中同一位置点的像素值进行平均处理,得到特征图中各位置点的像素值,并将平均处理后的特征图作为每个尺寸的单通道特征图;对多个尺寸的位置图和各尺寸的单通道特征图的像素值进行归一化处理。
图4为本公开一实施例提供的人脸朝向估计装置的结构示意图,如图4,该装置包括获取模块301、确定模块302和输出模块303,其中:
获取模块301,用于获取待检测人脸图像。
确定模块302,用于采用人脸朝向估计网络对待检测人脸图像进行处理,确定待检测人脸图像的人脸朝向;其中,人脸朝向估计网络的隐藏层中设置有位置图,人脸朝向估计网络为根据位置图对样本人脸图像进行监督训练得到的网络,位置图为:样本人脸图像在三维建模空间中完整三维面部形状的二维表示。
输出模块303,用于输出待检测人脸图像的人脸朝向。
图5为本公开一实施例提供了一种人脸朝向估计网络的训练装置的结构示意图,如图5,其装置包括:获取模块401、确定模块402和生成模块403,其中:
获取模块401,用于根据多个样本人脸图像获取对应的位置图。
确定模块402,用于根据多个位置图对对应的样本人脸图像进行监督训练,确定各样本人脸图像的人脸朝向。
生成模块403,用于根据监督训练结果,生成人脸朝向估计网络。
可选地,获取模块401,具体用于采用位置图回归网络PRN,对样本人脸图像进行处理,得到样本人脸图像对应的位置图。
图6为本公开另一实施例提供了一种人脸朝向估计网络的训练装置的结构示意图,如图6,该装置还包括:处理模块404,其中:
生成模块403,具体用于采用人脸朝向估计网络,对样本人脸图像进行处理,在每个中间层生成对应的特征图。
处理模块404,用于采用人脸朝向估计网络,对位置图进行多次下采样处理,得到多个尺寸的位置图,使得每个尺寸的特征图都有对应的尺寸相同的位置图。
获取模块401,具体用于根据位置图对尺寸相同的特征图进行监督训练。
可选地,确定模块402,具体用于采用人脸朝向估计网络对样本人脸图像进行处理,得到样本人脸图像的估计人脸朝向;根据样本人脸图像的估计人脸朝向和样本人脸图像的预设人脸朝向,确定人脸朝向估计网络的回归损失;根据多个尺寸的位置图和对应尺寸的特征图,确定人脸朝向估计网络的位置偏移损失;根据回归损失和位置偏移损失,确定人脸朝向估计网络的损失函数;根据损失函数对人脸朝向网络进行迭代训练。
可选地,确定模块402,具体用于根据样本人脸图像的估计人脸朝向和样本人脸图像的 预设人脸朝向,采用L1损失函数确定人脸朝向估计网络的回归损失。
可选地,处理模块404,具体用于对多个尺寸的位置图和对应尺寸的特征图的像素值进行归一化处理。
确定模块402,具体用于根据归一化后的位置图和特征图的像素值,计算位置偏移损失。
可选地,确定模块402,具体用于根据多个尺寸的位置图和对应尺寸的特征图,采用均方误差算法确定人脸朝向估计网络的位置偏移损失。
可选地,处理模块404,具体用于对多个通道的特征图中同一位置点的像素值进行平均处理,得到特征图中各位置点的像素值,并将平均处理后的特征图作为每个尺寸的单通道特征图;对多个尺寸的位置图和各尺寸的单通道特征图的像素值进行归一化处理。
可选地,确定模块402,具体用于根据回归损失、位置偏移损失和预设权重,确定人脸朝向估计网络的损失函数。
基于同一发明构思,本公开实施例中还提供了与人脸朝向估计方法对应的人脸朝向估计装置,由于本公开实施例中的装置解决问题的原理与本公开实施例上述人脸朝向估计方法相似,因此装置的实施可以参见方法的实施,有益效果的重复之处不再赘述。
如图5所示,本公开实施例还提供一种人脸朝向估计设备,包括:处理器601、存储器602和总线603;存储器602存储有处理器601可执行的机器可读指令,当人脸朝向估计设备运行时,处理器601与存储器602之间通过总线603通信,处理器601执行机器可读指令,以执行时执行如前述方法实施例所提供的请求处理方法的步骤。
具体地,存储器602中所存储的机器可读指令为本公开前述实施例所述的请求处理方法的执行步骤,处理器601可执行该请求处理方法对请求进行处理,因此,该电子设备同样具备前述方法实施例中所述的全部有益效果,本公开亦不再重复描述。
需要说明的是,该电子设备可以是通用计算机或特殊用途的计算机,以及其他用于处理数据的电子设备等,三者都可以用于实现本公开的请求处理方法。本公开尽管仅仅通过计算机和电子设备分别对请求处理方法进行了说明,但是为了方便起见,也可以在多个类似平台上以分布式方式实现本公开描述的功能,以均衡处理负载。
例如,电子设备可以包括用于执行程序指令的一个或多个处理器、通信总线、和不同形式的存储介质,例如,磁盘、ROM、或RAM,或其任意组合。示例性地,计算机平台还可以包括存储在ROM、RAM、或其他类型的非暂时性存储介质、或其任意组合中的程序指令。根据这些程序指令可以实现本公开的方法。
为了便于说明,在电子设备中仅描述了一个处理器。然而,应当注意,本公开中的电子设备还可以包括多个处理器,因此本公开中描述的一个处理器执行的步骤也可以由多个处理器联合执行或单独执行。
本公开实施例还提供了一种存储介质,该存储介质上存储有计算机程序,该计算机程序被处理器运行时执行上述人脸朝向估计方法的步骤。
具体地,该存储介质能够为通用的存储介质,如移动磁盘、硬盘等,该存储介质上的计算机程序被运行时,能够执行上述人脸朝向估计方法,从而,解决现有技术中存在的由于语言表达组合形式多种多样,大量信息会导致句库规模过大,占用过多的资源的问题,进而达到减小资源占用的效果。
所属领域的技术人员可以清楚地了解到,为描述的方便和简洁,上述描述的系统和装置的具体工作过程,可以参考方法实施例中的对应过程,本公开中不再赘述。在本公开所提供的几个实施例中,应该理解到,所揭露的系统、装置和方法,可以通过其它的方式实现。以上所描述的装置实施例仅仅是示意性的,例如,所述模块的划分,仅仅为一种逻辑功能划分,实际实现时可以有另外的划分方式,又例如,多个模块或组件可以结合或者可以集成到另一个系统,或一些特征可以忽略,或不执行。另一点,所显示或讨论的相互之间的耦合或直接耦合或通信连接可以是通过一些通信接口,装置或模块的间接耦合或通信连接,可以是电性,机械或其它的形式。
所述作为分离部件说明的模块可以是或者也可以不是物理上分开的,作为模块显示的部件可以是或者也可以不是物理单元,即可以位于一个地方,或者也可以分布到多个网络单元上。可以根据实际的需要选择其中的部分或者全部单元来实现本实施例方案的目的。
另外,在本公开各个实施例中的各功能单元可以集成在一个处理单元中,也可以是各个单元单独物理存在,也可以两个或两个以上单元集成在一个单元中。
所述功能如果以软件功能单元的形式实现并作为独立的产品销售或使用时,可以存储在一个处理器可执行的非易失的计算机可读取存储介质中。基于这样的理解,本公开的技术方案本质上或者说对现有技术做出贡献的部分或者该技术方案的部分可以以软件产品的形式体现出来,该计算机软件产品存储在一个存储介质中,包括若干指令用以使得一台计算机设备(可以是个人计算机,服务器,或者网络设备等)执行本公开各个实施例所述方法的全部或部分步骤。而前述的存储介质包括:U盘、移动硬盘、ROM、RAM、磁碟或者光盘等各种可以存储程序代码的介质。
以上仅为本公开的具体实施方式,但本公开的保护范围并不局限于此,任何熟悉本技术领域的技术人员在本公开揭露的技术范围内,可轻易想到变化或替换,都应涵盖在本公开的保护范围之内。因此,本公开的保护范围应以权利要求的保护范围为准。
工业实用性
采用上述方案,电子设备首先可以获取待检测人脸图像,然后根据人脸朝向估计网络对待检测人脸图像进行处理,确定待检测人脸图像的人脸朝向;其中,人脸朝向估计网络的隐藏层中设置有位置图,人脸朝向估计网络为根据位置图对样本人脸图像进行监督训练得到的网络;输出待检测人脸图像的人脸朝向。提高了获取到的人脸朝向的精度。

Claims (16)

  1. 一种人脸朝向估计方法,其特征在于,所述方法包括:
    获取待检测人脸图像;
    采用人脸朝向估计网络对所述待检测人脸图像进行处理,确定所述待检测人脸图像的人脸朝向;其中,所述人脸朝向估计网络的隐藏层中设置有位置图,所述人脸朝向估计网络为根据所述位置图对样本人脸图像进行监督训练得到的网络,所述位置图为:所述样本人脸图像在三维建模空间中完整三维面部形状的二维表示;
    输出所述待检测人脸图像的人脸朝向。
  2. 如权利要求1所述的方法,其特征在于,所述人脸朝向包括:人脸朝向偏转角、人脸朝向仰俯角和人脸朝向滚动角。
  3. 一种人脸朝向估计网络的训练方法,其特征在于,所述方法包括:
    根据多个样本人脸图像获取对应的位置图;
    根据多个位置图对对应的样本人脸图像进行监督训练,确定各所述样本人脸图像的人脸朝向;
    根据监督训练的结果,生成所述人脸朝向估计网络。
  4. 如权利要求3所述的方法,其特征在于,所述根据多个样本人脸图像获取对应的位置图,包括:
    采用位置图回归网络PRN,对所述样本人脸图像进行处理,得到所述样本人脸图像对应的位置图。
  5. 如权利要求3-4中任一所述的方法,其特征在于,所述根据多个位置图对对应的样本人脸图像进行监督训练,确定各所述样本人脸图像的人脸朝向,包括:
    采用所述人脸朝向估计网络,对所述样本人脸图像进行处理,在每个中间层生成对应的特征图;
    采用所述人脸朝向估计网络,对所述位置图进行多次下采样处理,得到多个尺寸的位置图,使得每个尺寸的所述特征图都有对应的尺寸相同的所述位置图;
    根据所述位置图对尺寸相同的所述特征图进行监督训练。
  6. 如权利要求5所述的方法,其特征在于,所述根据位置图对尺寸相同的特征图进行监督训练,包括:
    采用所述人脸朝向估计网络对所述样本人脸图像进行处理,得到所述样本人脸图像的估计人脸朝向;
    根据所述样本人脸图像的估计人脸朝向和所述样本人脸图像的预设人脸朝向,确定所述人脸朝向估计网络的回归损失;
    根据多个尺寸的所述位置图和对应尺寸的所述特征图,确定所述人脸朝向估计网络的位置偏移损失;
    根据所述回归损失和所述位置偏移损失,确定所述人脸朝向估计网络的损失函数;
    根据所述损失函数对所述人脸朝向网络进行迭代训练。
  7. 如权利要求6所述的方法,其特征在于,所述根据所述样本人脸图像的估计人脸朝向和所述样本人脸图像的预设人脸朝向,确定所述人脸朝向估计网络的回归损失,包括:
    根据所述样本人脸图像的估计人脸朝向和所述样本人脸图像的预设人脸朝向,采用L1损失函数确定所述人脸朝向估计网络的回归损失。
  8. 如权利要求6所述的方法,其特征在于,所述根据多个尺寸的所述位置图和对应尺寸的所述特征图,确定所述人脸朝向估计网络的位置偏移损失,包括:
    对多个尺寸的所述位置图和对应尺寸的所述特征图的像素值进行归一化处理;
    根据归一化后的所述位置图和所述特征图的像素值,计算所述位置偏移损失。
  9. 如权利要求6或8所述的方法,其特征在于,所述根据多个尺寸的所述位置图和对应 尺寸的所述特征图,确定所述人脸朝向估计网络的位置偏移损失,包括:
    根据多个尺寸的所述位置图和对应尺寸的所述特征图,采用均方误差算法确定所述人脸朝向估计网络的位置偏移损失。
  10. 如权利要求8所述的方法,其特征在于,每个尺寸的所述特征图包括:多个通道的所述特征图;所述对多个尺寸的所述位置图和对应尺寸的所述特征图的像素值进行归一化处理,包括:
    对多个通道的所述特征图中同一位置点的像素值进行平均处理,得到所述特征图中各位置点的像素值,并将平均处理后的特征图作为每个尺寸的单通道特征图;
    对多个尺寸的所述位置图和各尺寸的所述单通道特征图的像素值进行归一化处理。
  11. 如权利要求6-10中任一所述的方法,其特征在于,所述根据所述回归损失和所述位置偏移损失,确定所述人脸朝向估计网络的损失函数,包括:
    根据所述回归损失、所述位置偏移损失和预设权重,确定所述人脸朝向估计网络的损失函数。
  12. 如权利要求3-11中任一所述的方法,其特征在于,所述人脸朝向包括:人脸朝向偏转角、人脸朝向仰俯角和人脸朝向滚动角。
  13. 一种人脸朝向估计装置,其特征在于,所述装置包括:获取模块、确定模块和输出模块,其中:
    所述获取模块,用于获取待检测人脸图像;
    所述确定模块,用于采用人脸朝向估计网络对所述待检测人脸图像进行处理,确定所述待检测人脸图像的人脸朝向;其中,所述人脸朝向估计网络的隐藏层中设置有位置图,所述人脸朝向估计网络为根据所述位置图对样本人脸图像进行监督训练得到的网络,所述位置图为:所述样本人脸图像在三维建模空间中完整三维面部形状的二维表示;
    所述输出模块,用于输出待检测人脸图像的人脸朝向。
  14. 一种人脸朝向估计网络的训练装置,其特征在于,所述装置包括:获取模块、确定模块和生成模块,其中:
    所述获取模块,用于根据多个样本人脸图像获取对应的位置图;
    所述确定模块,用于根据多个位置图对对应的样本人脸图像进行监督训练,确定各所述样本人脸图像的人脸朝向;
    所述生成模块,用于根据监督训练结果,生成所述人脸朝向估计网络。
  15. 一种电子设备,其特征在于,包括:处理器、存储介质和总线,所述存储介质存储有所述处理器可执行的机器可读指令,当电子设备运行时,所述处理器与所述存储介质之间通过总线通信,所述处理器执行所述机器可读指令,以执行如权利要求1-12任一所述方法的步骤。
  16. 一种存储介质,其特征在于,所述存储介质上存储有计算机程序,所述计算机程序被处理器运行时执行如权利要求1-12任一所述方法的步骤。
PCT/CN2020/104929 2019-09-16 2020-07-27 一种人脸朝向估计、网络训练方法、装置、电子设备及存储介质 Ceased WO2021052010A1 (zh)

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
CN201910870466.XA CN110781728B (zh) 2019-09-16 2019-09-16 一种人脸朝向估计方法、装置、电子设备及存储介质
CN201910870466.X 2019-09-16

Publications (1)

Publication Number Publication Date
WO2021052010A1 true WO2021052010A1 (zh) 2021-03-25

Family

ID=69384129

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/CN2020/104929 Ceased WO2021052010A1 (zh) 2019-09-16 2020-07-27 一种人脸朝向估计、网络训练方法、装置、电子设备及存储介质

Country Status (2)

Country Link
CN (1) CN110781728B (zh)
WO (1) WO2021052010A1 (zh)

Cited By (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN115311727A (zh) * 2022-08-29 2022-11-08 上海大学 一种室外队列训练头部注意力分析方法、装置、设备和介质
CN115761856A (zh) * 2022-11-24 2023-03-07 北京京东方技术开发有限公司 人脸识别方法及相关设备
CN117115877A (zh) * 2023-04-27 2023-11-24 广州图匠数据科技有限公司 一种人脸动态表情的识别方法、装置、设备以及存储介质

Families Citing this family (6)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN110781728B (zh) * 2019-09-16 2020-11-10 北京嘀嘀无限科技发展有限公司 一种人脸朝向估计方法、装置、电子设备及存储介质
CN111523403B (zh) * 2020-04-03 2023-10-20 咪咕文化科技有限公司 图片中目标区域的获取方法及装置、计算机可读存储介质
CN111626193A (zh) * 2020-05-26 2020-09-04 北京嘀嘀无限科技发展有限公司 一种面部识别方法、面部识别装置及可读存储介质
CN112241761B (zh) * 2020-10-15 2024-03-26 北京字跳网络技术有限公司 模型训练方法、装置和电子设备
CN114694203A (zh) * 2020-12-31 2022-07-01 深圳云天励飞技术股份有限公司 设备控制方法、装置、系统、电子设备及存储介质
CN113822250A (zh) * 2021-11-23 2021-12-21 中船(浙江)海洋科技有限公司 一种船舶驾驶异常行为检测方法

Citations (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20160148381A1 (en) * 2013-07-03 2016-05-26 Panasonic Intellectual Property Management Co., Ltd. Object recognition device and object recognition method
CN108197547A (zh) * 2017-12-26 2018-06-22 深圳云天励飞技术有限公司 人脸姿态估计方法、装置、终端及存储介质
CN110163087A (zh) * 2019-04-09 2019-08-23 江西高创保安服务技术有限公司 一种人脸姿态识别方法及系统
CN110781728A (zh) * 2019-09-16 2020-02-11 北京嘀嘀无限科技发展有限公司 一种人脸朝向估计方法、装置、电子设备及存储介质

Family Cites Families (9)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP4204336B2 (ja) * 2003-01-30 2009-01-07 富士通株式会社 顔の向き検出装置、顔の向き検出方法及びコンピュータプログラム
EP3011734A4 (en) * 2013-06-17 2017-02-22 RealD Inc. Controlling light sources of a directional backlight
CN106127120B (zh) * 2016-06-16 2018-03-13 北京市商汤科技开发有限公司 姿势估计方法和装置、计算机系统
CN109960986A (zh) * 2017-12-25 2019-07-02 北京市商汤科技开发有限公司 人脸姿态分析方法、装置、设备、存储介质以及程序
CN108038465A (zh) * 2017-12-25 2018-05-15 深圳市唯特视科技有限公司 一种基于合成数据集的三维多人物姿态估计
CN108875549B (zh) * 2018-04-20 2021-04-09 北京旷视科技有限公司 图像识别方法、装置、系统及计算机存储介质
CN108898556A (zh) * 2018-05-24 2018-11-27 麒麟合盛网络技术股份有限公司 一种三维人脸的图像处理方法及装置
CN108921926B (zh) * 2018-07-02 2020-10-09 云从科技集团股份有限公司 一种基于单张图像的端到端三维人脸重建方法
CN109934196A (zh) * 2019-03-21 2019-06-25 厦门美图之家科技有限公司 人脸姿态参数评估方法、装置、电子设备及可读存储介质

Patent Citations (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20160148381A1 (en) * 2013-07-03 2016-05-26 Panasonic Intellectual Property Management Co., Ltd. Object recognition device and object recognition method
CN108197547A (zh) * 2017-12-26 2018-06-22 深圳云天励飞技术有限公司 人脸姿态估计方法、装置、终端及存储介质
CN110163087A (zh) * 2019-04-09 2019-08-23 江西高创保安服务技术有限公司 一种人脸姿态识别方法及系统
CN110781728A (zh) * 2019-09-16 2020-02-11 北京嘀嘀无限科技发展有限公司 一种人脸朝向估计方法、装置、电子设备及存储介质

Cited By (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN115311727A (zh) * 2022-08-29 2022-11-08 上海大学 一种室外队列训练头部注意力分析方法、装置、设备和介质
CN115761856A (zh) * 2022-11-24 2023-03-07 北京京东方技术开发有限公司 人脸识别方法及相关设备
CN117115877A (zh) * 2023-04-27 2023-11-24 广州图匠数据科技有限公司 一种人脸动态表情的识别方法、装置、设备以及存储介质

Also Published As

Publication number Publication date
CN110781728A (zh) 2020-02-11
CN110781728B (zh) 2020-11-10

Similar Documents

Publication Publication Date Title
WO2021052010A1 (zh) 一种人脸朝向估计、网络训练方法、装置、电子设备及存储介质
CN109188457B (zh) 物体检测框的生成方法、装置、设备、存储介质及车辆
JP6745328B2 (ja) 点群データを復旧するための方法及び装置
US11244435B2 (en) Method and apparatus for generating vehicle damage information
CN111860398B (zh) 遥感图像目标检测方法、系统及终端设备
JP6364049B2 (ja) 点群データに基づく車両輪郭検出方法、装置、記憶媒体およびコンピュータプログラム
US10395103B2 (en) Object detection method, object detection apparatus, and program
EP4471737A1 (en) Face pose estimation method and apparatus, electronic device, and storage medium
CN113378712B (zh) 物体检测模型的训练方法、图像检测方法及其装置
CN114120454A (zh) 活体检测模型的训练方法、装置、电子设备及存储介质
CN113506328A (zh) 视线估计模型的生成方法和装置、视线估计方法和装置
CN110046116A (zh) 一种张量填充方法、装置、设备及存储介质
JP2023027227A (ja) 画像処理方法、装置、電子機器、記憶媒体及びコンピュータプログラム
CN112861940A (zh) 双目视差估计方法、模型训练方法以及相关设备
CN114882480A (zh) 用于获取目标对象状态的方法、装置、介质以及电子设备
CN114972146B (zh) 基于生成对抗式双通道权重分配的图像融合方法及装置
CN112734827A (zh) 一种目标检测方法、装置、电子设备和存储介质
CN111723926A (zh) 用于确定图像视差的神经网络模型的训练方法和训练装置
CN114387197B (zh) 一种双目图像处理方法、装置、设备和存储介质
CN113610856A (zh) 训练图像分割模型和图像分割的方法和装置
CN113205131A (zh) 图像数据的处理方法、装置、路侧设备和云控平台
CN111831800B (zh) 问答交互方法、装置、设备及存储介质
CN114820755B (zh) 一种深度图估计方法及系统
CN114494399A (zh) 车载环视参数的验证方法及装置、电子设备和存储介质
CN114511908A (zh) 一种人脸活体检测方法、装置、电子设备及存储介质

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 20866204

Country of ref document: EP

Kind code of ref document: A1

NENP Non-entry into the national phase

Ref country code: DE

122 Ep: pct application non-entry in european phase

Ref document number: 20866204

Country of ref document: EP

Kind code of ref document: A1