WO2023125366A1 - 图像处理方法、装置、电子设备和存储介质 - Google Patents

图像处理方法、装置、电子设备和存储介质 Download PDF

Info

Publication number
WO2023125366A1
WO2023125366A1 PCT/CN2022/141795 CN2022141795W WO2023125366A1 WO 2023125366 A1 WO2023125366 A1 WO 2023125366A1 CN 2022141795 W CN2022141795 W CN 2022141795W WO 2023125366 A1 WO2023125366 A1 WO 2023125366A1
Authority
WO
WIPO (PCT)
Prior art keywords
model
facial
attribute
target
image
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/CN2022/141795
Other languages
English (en)
French (fr)
Inventor
李冰川
易子立
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Beijing Zitiao Network Technology Co Ltd
Original Assignee
Beijing Zitiao Network Technology Co Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Beijing Zitiao Network Technology Co Ltd filed Critical Beijing Zitiao Network Technology Co Ltd
Priority to US18/725,684 priority Critical patent/US20250078566A1/en
Publication of WO2023125366A1 publication Critical patent/WO2023125366A1/zh
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Images

Classifications

    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T13/00Animation
    • G06T13/20Three-dimensional [3D] animation
    • G06T13/40Three-dimensional [3D] animation of characters, e.g. humans, animals or virtual beings
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V40/00Recognition of biometric, human-related or animal-related patterns in image or video data
    • G06V40/10Human or animal bodies, e.g. vehicle occupants or pedestrians; Body parts, e.g. hands
    • G06V40/16Human faces, e.g. facial parts, sketches or expressions
    • G06V40/168Feature extraction; Face representation
    • G06V40/171Local features and components; Facial parts ; Occluding parts, e.g. glasses; Geometrical relationships
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F18/00Pattern recognition
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V10/00Arrangements for image or video recognition or understanding
    • G06V10/70Arrangements for image or video recognition or understanding using pattern recognition or machine learning
    • G06V10/82Arrangements for image or video recognition or understanding using pattern recognition or machine learning using neural networks
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V40/00Recognition of biometric, human-related or animal-related patterns in image or video data
    • G06V40/10Human or animal bodies, e.g. vehicle occupants or pedestrians; Body parts, e.g. hands
    • G06V40/16Human faces, e.g. facial parts, sketches or expressions
    • G06V40/168Feature extraction; Face representation
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V40/00Recognition of biometric, human-related or animal-related patterns in image or video data
    • G06V40/10Human or animal bodies, e.g. vehicle occupants or pedestrians; Body parts, e.g. hands
    • G06V40/16Human faces, e.g. facial parts, sketches or expressions
    • G06V40/172Classification, e.g. identification

Definitions

  • the present disclosure relates to the technical field of computers, for example, to an image processing method, device, electronic equipment and storage medium.
  • the technology of adding special effects usually uses software to modify images to add special effects to objects in the image, such as making the object turn its head.
  • this method is likely to cause distortion of the object's limbs and face, resulting in poor effects of adding special effects question.
  • the present disclosure provides an image processing method, device, electronic equipment, and storage medium, so as to add facial special effects to objects, so that the generated image is most suitable for the effect of facial changes in practice, and the accuracy of adding facial special effects is improved. .
  • the present disclosure provides an image processing method, the method comprising:
  • the data to be processed includes Gaussian noise or an image to be converted
  • an image processing device which includes:
  • a data acquisition module configured to acquire data to be processed; wherein, the data to be processed includes Gaussian noise or an image to be converted;
  • the target image determination module is configured to process the data to be processed based on the target facial attribute determination model to obtain a target facial image corresponding to the data to be processed; wherein at least one target feature in the target facial image is the same as corresponding to at least one preset facial feature.
  • the present disclosure also provides electronic equipment, and the equipment includes:
  • processors one or more processors
  • a storage device configured to store one or more programs
  • the one or more processors implement the above image processing method.
  • the present disclosure further provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the above-mentioned image processing method is realized.
  • the present disclosure further provides a computer program product, including a computer program carried on a non-transitory computer readable medium, where the computer program includes program code for executing the above-mentioned image processing method.
  • FIG. 1 is a schematic flowchart of an image processing method provided in Embodiment 1 of the present disclosure
  • FIG. 2 is a schematic diagram of a target facial image matching preset facial features provided by Embodiment 1 of the present disclosure
  • FIG. 3 is a schematic flowchart of an image processing method provided in Embodiment 2 of the present disclosure.
  • FIG. 4 is a schematic structural diagram of a target facial attribute determination model provided in Embodiment 2 of the present disclosure.
  • FIG. 5 is a schematic flowchart of an image processing method provided by Embodiment 3 of the present disclosure.
  • FIG. 6 is a schematic structural diagram of a facial attribute determination model to be trained provided by Embodiment 3 of the present disclosure.
  • FIG. 7 is a schematic flowchart of an image processing method provided in Embodiment 4 of the present disclosure.
  • FIG. 8 is a structural block diagram of an image processing device provided in Embodiment 5 of the present disclosure.
  • FIG. 9 is a schematic structural diagram of an electronic device provided by Embodiment 6 of the present disclosure.
  • the term “comprise” and its variations are open-ended, ie “including but not limited to”.
  • the term “based on” is “based at least in part on”.
  • the term “one embodiment” means “at least one embodiment”; the term “another embodiment” means “at least one further embodiment”; the term “some embodiments” means “at least some embodiments.” Relevant definitions of other terms will be given in the description below.
  • the disclosed technical solution can be applied to any scene that requires special effects display, for example, in a video call, special effects display can be performed; or, in a live broadcast scene, special effects display can be performed on the anchor object; it can also be applied in the video shooting process , to display special effects on the image corresponding to the subject to be photographed.
  • the captured image can be processed into a special effect image, and then the processed special effect image can be displayed with special effects; it can also be applied in static
  • the process of image capturing for example, after the image is captured by the built-in camera of the terminal device, the captured image is processed into a special effect image for special effect display.
  • Fig. 1 is a schematic flow chart of an image processing method provided by Embodiment 1 of the present disclosure.
  • the embodiment of the present disclosure is applicable to processing the facial image of the target object into a special effect image and displaying it in any image display scene supported by the Internet.
  • the method can be executed by an image processing device, and the device can be implemented in the form of software and/or hardware, for example, implemented by electronic equipment, and the electronic equipment can be a mobile terminal, a personal computer (Personal Computer, PC) terminal or server etc.
  • the scene of arbitrary image display is usually implemented by the cooperation of the client and the server.
  • the method provided in this embodiment can be executed by the server, the client, or the cooperation of the client and the server.
  • the method of the present embodiment comprises:
  • the device for executing the image processing method provided by the embodiments of the present disclosure may be integrated into application software supporting image processing functions, and the software may be installed in electronic equipment, for example, the electronic equipment may be a mobile terminal or a PC terminal, etc.
  • the application software may be a type of software for image/video processing, and the application software thereof will not be described here one by one, as long as the image/video processing can be realized.
  • the data to be processed can be data that needs to be processed, can be Gaussian noise, or can be an image.
  • the Gaussian noise may be random sampling noise, and the Gaussian noise may include at least one noise such as fluctuation noise, cosmic noise, thermal noise, and shot noise.
  • the image to be converted may be an image collected based on the application software, or an image pre-stored by the application software from the storage space. In an application scenario, the image to be converted may be collected in real time or periodically. For example, in a live broadcast scene or a video shooting scene, the camera captures in real time the image corresponding to the target in the target scene, and at this time, the image captured by the camera may be used as the image to be converted.
  • the image to be converted may include a target subject, and the target subject may be a user, a pet, a flower, a tree, and the like. It is also possible to process the video frame corresponding to the captured video.
  • the target subject corresponding to the captured video can be set in advance.
  • the image corresponding to the video frame can be used as The image is to be converted, so that the facial features of the target subject in each video frame image in the video can be subsequently processed.
  • the number of target subjects in the same shooting scene can be one or more, and no matter it is one or more target subjects, the technical solution provided by the present disclosure can be used to determine the special effect display map.
  • the image including the target subject can be collected in real time or at intervals as the image to be converted.
  • the device can also be used to randomly collect noise to obtain Gaussian noise.
  • the Gaussian noise and the image to be converted can be used as the input of the model to train the model.
  • the original facial data when corresponding features need to be added to the facial image are taken as the data to be processed.
  • S120 Process the data to be processed based on a target facial attribute determination model to obtain a target facial image corresponding to the data to be processed.
  • the target facial attribute determination model may be pre-trained.
  • the target facial attribute determination model can process the input Gaussian noise or the image to be converted to obtain an image with corresponding features added to the face in the Gaussian noise or the image to be converted.
  • the facial features added by the model for the user may be pre-determined, and the facial features at this time may be used as preset facial features.
  • the image output by the target facial attribute determination model can be used as the target facial image, and at this time, preset facial features are added to the target facial image.
  • the preset facial features include facial features of wearing at least one kind of jewelry, facial features of different ages, facial features of different angles, facial features of different hairstyles, facial features of different hairstyles and color matching, and different facial features. At least one of the facial features of the expression.
  • the accessories worn may be glasses, sunglasses, face stickers, etc., and the features corresponding to the worn accessories may be used as preset facial features.
  • the facial features of different age stages can be the facial features corresponding to the subjects of different age groups. Different age groups can be based on the original collected images, and the age groups are divided by increasing or decreasing age on this basis. Different age groups It can be young age group, middle age group, old age group.
  • the facial features of different angles can be the corresponding features when the main body faces different orientations. For example, the different angles can also be based on the corresponding facial angle of the main body in the original captured image, and can be deflected 20 degrees to the left on the basis of this benchmark.
  • 20-degree deflection, 20-degree downward deflection, etc., and the corresponding facial features at these angles can also be used as preset facial features.
  • the preset facial features may also be facial features of different hairstyles, for example, long hair, short hair, curly hair, straight hair, and the like.
  • the facial features of different hairstyles and color combinations can be purple, white, black and other colors, and can be a combination of different hairstyles and different colors.
  • the facial features of different expressions may include facial expression features such as smiling, angry, and angry.
  • the data to be processed can be input into the pre-trained target facial attribute determination model, corresponding features can be added to the image corresponding to the data to be processed, and then the face with added features corresponding to the data to be processed can be output image, which is the target face image.
  • the preset facial features may include turning the face to the right by 20° (yaw+20°), increasing the age by 20 years (age+20), wearing sunglasses (wear glasses), smiling (smlie), and turning the face to the right by 30° °(yaw+30°), at least one of the facial features corresponding to the image whose age is reduced by 20 years (age-20), then after inputting the original image (Origin) into the model, the corresponding schematic diagram can be obtained.
  • the schematic diagram can be seen in Fig. 2.
  • the technical solution of the embodiment of the present disclosure obtains the data to be processed, and then processes the data to be processed based on the target facial attribute determination model, and obtains the target facial image corresponding to the data to be processed after adding special effects, which solves the problem of using correction in related technologies.
  • image technology generates special effect images
  • the obtained special effect images have low authenticity, which causes the problem of poor user experience. It is realized to add corresponding facial features to the faces of the target objects in the data to be processed, so that the obtained special effects are realistic. Higher, as well as the effect of increasing the richness and interest of video image content, and improving the technical effect of user experience.
  • FIG. 3 is a schematic flowchart of an image processing method provided in Embodiment 2 of the present disclosure.
  • S120 is described, and its implementation may refer to the technical solution of this embodiment. Wherein, technical terms that are the same as or corresponding to those in the foregoing embodiments will not be repeated here.
  • the method includes the following steps:
  • the target facial attribute determination model includes a feature preprocessing sub-model, an attribute editing sub-model and an image generation sub-model.
  • the data to be processed is processed through these three sub-models, and the target facial image corresponding to the data to be processed is obtained.
  • the feature preprocessing sub-model can be used to extract the corresponding features.
  • the attribute editing sub-model may refer to a model that adds preset facial features for feature extraction.
  • the image generation sub-model may refer to a model that performs image generation for features output by the attribute editing sub-model.
  • the feature vector to be spliced refers to the vector output by the feature preprocessing sub-model, and the feature vector extracted by the feature preprocessing sub-model can be used as the feature vector to be spliced.
  • the data to be processed can be used as the input of the feature preprocessing sub-model. After the data is processed by the feature extraction of the model, the feature vector corresponding to the data to be processed can be obtained, which can be used as the feature vector to be spliced.
  • a schematic structural diagram of a target face attribute determination model may be referred to in FIG. 4 , and the model may include a feature preprocessing sub-model, an attribute editing sub-model, and an image generation sub-model.
  • the feature preprocessing sub-model includes a first feature extraction module and a second feature extraction module.
  • the data to be processed may include Gaussian noise or an image to be converted, for the accuracy of data processing.
  • Different processing methods can be adopted according to different data, and different data can be processed in a targeted manner with algorithm modules used based on different processing methods. Its processing method can refer to the following expression:
  • the feature preprocessing sub-model includes a first feature extraction module and a second feature extraction module, and based on the feature preprocessing sub-model, determining the feature vector to be spliced corresponding to the data to be processed includes: if the If the data to be processed is the Gaussian noise, the feature vector to be spliced corresponding to the Gaussian noise is determined based on the first feature extraction module; if the data to be processed is the image to be converted, then based on the The second feature extraction module determines the feature vector to be spliced corresponding to the image to be converted.
  • the first feature extraction module is used to extract feature vectors corresponding to Gaussian noise.
  • the second feature extraction module is used to extract feature vectors corresponding to facial attributes in the image to be converted.
  • two feature extraction modules may be preset in the feature preprocessing sub-model to respectively process the two kinds of data correspondingly.
  • the data to be processed as Gaussian noise can be input to the first feature extraction module, and the module can process the Gaussian noise, and then the feature vector to be spliced corresponding to the Gaussian noise can be obtained;
  • the data to be processed can also be The data of the converted image is input to the second feature extraction module, and the module can process the image to be converted, and then can obtain the feature vector to be stitched corresponding to the image to be converted, and correspondingly, can obtain the feature vector to be stitched corresponding to all the data to be processed.
  • Gaussian noise may be processed based on the first feature extraction module, and a feature vector to be concatenated corresponding to Gaussian noise is output.
  • the first feature extraction module can be a Mapping Network model.
  • the image to be converted may be processed based on the second feature extraction module, and a feature vector to be spliced corresponding to the image to be converted may be output.
  • the second feature extraction module can be an Encoder model, such as fixing the generator parameters of the trained stylegan model, and training the Encoder model.
  • the facial image can be input, encoded by the Encoder and then passed through the stylegan generator, and can be reconstructed Take this image of the face.
  • the output feature vector to be concatenated can be used as W+ for the input of the subsequent model.
  • the type of data input to the model can be determined based on the data interface, and then which module to process it can be determined based on the data type. If the data to be processed is Gaussian noise, the Gaussian noise can be used as the input of the first feature extraction module, and the module can output the feature vector to be spliced corresponding to the Gaussian noise; if the data to be processed is an image to be converted, the image to be converted can be used as the first The input of the second feature extraction module, the module can output the feature vector to be spliced corresponding to the image to be converted.
  • the preset feature vector refers to a vector corresponding to a preset facial feature
  • the target feature vector may be a feature vector obtained by splicing the feature vector to be spliced and the preset feature vector.
  • the feature vector to be spliced and the feature vector corresponding to the preset facial features can be spliced, and then the corresponding target feature vector after adding facial features to the data to be processed can be obtained , to generate the target face image with special effects based on the target feature vector.
  • the sub-model can splice the feature vector to be spliced with the preset feature vector corresponding to the preset facial feature, for example, the feature vector A can be combined with The feature vector B is concatenated into A-B.
  • a concatenated feature vector after splicing processing can be obtained, and the concatenated feature vector can be used as a target feature vector corresponding to the target facial image.
  • the attribute editing sub-model can be a Dynamic Network model
  • the Dynamic Network model can be used to splicing the feature vector to be spliced and the preset feature vector
  • the output W++ is the target corresponding to the target facial image Feature vector.
  • the input preset feature vector is encoded by a multilayer perceptron (MLP), and then passed through two fully connected layers (FC) and an activation function sigmoid, and multiplied by the input feature vector to be spliced. Add operation to get the target feature vector.
  • MLP multilayer perceptron
  • FC fully connected layers
  • an activation function sigmoid an activation function sigmoid
  • the feature vector to be spliced can be used as the input of the attribute editing sub-model, and the sub-model can splice the feature vector to be spliced with the preset feature vector corresponding to at least one preset facial feature, and then the target face with the preset facial feature can be obtained The target feature vector corresponding to the image, so that the target facial image with special effects can be generated based on the target feature vector.
  • S240 Process the target feature vector to obtain the target facial image.
  • the target feature vector corresponding to the feature vector to be spliced and the preset feature vector can be input into the pre-trained image generation sub-model, the model The target feature vector can be reconstructed, and the target facial image corresponding to the target feature vector can be output.
  • the image generation sub-model may be a Generator model, for example, the Generator model may be used to process the target feature vector, and a target facial image corresponding to the target feature vector may be output.
  • the parameters in the target face attribute determination model are tuned.
  • a trained model attribute classifier can be added, and the attribute classifier can be a model for extracting and classifying image attribute features.
  • the attribute classifier can be used to extract the feature data in the target facial image, determine the facial features in the target facial image, and verify whether the picture output by the model has preset facial features, and then, if the picture output by the model does not add preset facial features By setting facial features, the parameters in the model can be corrected to improve the accuracy of the model.
  • the model parameters in the facial attribute determination model are corrected; wherein, the at least one target attribute matches the attribute identifier of the at least one preset facial feature.
  • the attribute identification can refer to the identification corresponding to the preset facial features, that is to say, the preset facial features can be represented by the corresponding identification, for example, the age feature is represented by A1 , the angle feature is represented by A2 , and the wearing glasses feature is represented by A2.
  • a 3 indicates that A 1 , A 2 , A 3 and other identifiers can be used as attribute identifiers of corresponding features, and target attributes can also be identifier information corresponding to preset facial features.
  • the output target facial image in order to optimize the parameters in the target facial attribute determination model, can be compared with the corresponding theoretical facial image with special effects added, then correspondingly, the output target facial image can also be The attribute feature corresponding to the image is compared with the added special effect feature, so as to correct the model parameters in the classification model to be trained based on whether the attribute feature corresponding to the target facial image contains the added special effect feature.
  • the target facial image can be input to a pre-trained attribute classifier, and then the classifier can perform feature extraction processing on the target facial image, and output the feature attribute corresponding to the target facial image, that is, the target attribute. Then compare the currently output feature attribute with the attribute identifier of the preset facial feature to calculate a comparison error value, and then adjust the model parameters in the model based on the comparison error value.
  • attribute classifier can be residual network (Residual Networks, ResNet) model, as, target face image can be input in attribute classifier, image is processed through the feature extraction of classifier, can output The target attribute corresponding to the target face image.
  • ResNet residual Networks
  • the target facial image can be used as the input of the attribute classifier, and the classifier can output the corresponding target attribute.
  • the algorithm can also be used to process the error between the target attribute corresponding to the image and the attribute identification of the preset facial feature, so as to correct the model parameters of the target facial attribute determination model according to the obtained error result, and improve the training accuracy of the target facial attribute determination model.
  • the target face image after adding special effects corresponding to different types of data is obtained, and at the same time
  • the parameters in the model are also corrected based on the target facial image, which improves the accuracy of the model, thereby improving the authenticity of adding special effects, and increasing the richness and interest of the video image.
  • Fig. 5 is a schematic flow chart of an image processing method provided by Embodiment 3 of the present disclosure.
  • a facial attribute determination model to be trained can also be constructed in advance, and the facial attribute determination model to be trained can be trained.
  • the target facial attribute determination model is obtained, and its implementation can refer to the technical solution of this embodiment. Wherein, technical terms that are the same as or corresponding to those in the foregoing embodiments will not be repeated here.
  • the method includes the following steps:
  • the facial attribute determination model to be trained refers to a model in which the model parameters in the model are set to default values, and the model needs to be trained to obtain the target facial attribute determination model.
  • training can be performed based on the pre-built facial attribute determination model to be trained, so that after the training of the facial attribute determination model to be trained is completed, the final facial attribute determination model can be obtained.
  • An applicable facial attribute determination model that is, a target facial attribute determination model.
  • the constructing the facial attribute determination model to be trained includes: editing the sub-model according to the attribute to be trained, and constructing the facial attribute to be trained by pre-training the target confrontation model, the target attribute classification model and the facial matching model Determine the model; wherein, the target attribute classification model is used to determine the facial features of the image output through the target confrontation model; the facial matching model is used to determine the matching degree of the facial image output based on the target confrontation model , the target confrontation model is used to output two facial images, one of which matches the preset facial features set in the attribute editing sub-model to be trained.
  • the target confrontation model includes: a feature preprocessing sub-model and an image generation sub-model; the feature preprocessing sub-model includes a first feature extraction module; the image generation sub-model includes a first image generation sub-module and a second image generation sub-model. An image generation sub-module; the feature preprocessing sub-model also includes a second feature extraction module.
  • the target confrontation model can be a stylegan model.
  • the target attribute classification model (resnet model) is used to determine the facial features of the image output by the second image generation sub-module.
  • the face matching model (face recognition model) is used to determine the matching degree of the facial images output by the first image generation sub-module and the second image generation sub-module.
  • the first feature extraction module (Mapping Network model) is used to extract the feature vector corresponding to the Gaussian noise.
  • the second feature extraction module (Encoder model) is used to extract feature vectors corresponding to facial attributes in the image.
  • the first image generation sub-module (Generator model) is used to determine the image corresponding to the feature vector output by the feature preprocessing sub-model.
  • the second image generation sub-module (Generator model) is used to determine the image corresponding to the feature vector output by the attribute editing sub-model to be trained.
  • the attribute editing sub-model to be trained (Dynamic Network model) is used to concatenate the feature vector output by the feature preprocessing sub-model with the preset feature vector.
  • the attribute editing sub-model that needs to be trained is used to concatenate the feature vector output by the feature preprocessing sub-model with the preset feature vector.
  • the attribute editing sub-model that needs to be trained. At this time, the output result of the attribute editing sub-model may not have reached the expected result. It needs to be trained to make the output result of the trained model consistent with the expected result.
  • the available Applied property editing submodels is used to determine the image corresponding to the feature vector output by the attribute editing sub-model to be trained.
  • the image code can be obtained through random noise or image encoder, and two branches of the stylegan generator are designed, wherein one branch directly generates the face image img1 through the trained stylegan generator; the other branch passes through a waiting Train the attribute editing sub-model (the input of the editing module is the attribute value, for example, input 1 for wearing glasses, 0 for not wearing glasses, 1 for smiling, and 0 for not smiling), and then the trained stylegan generator will generate the edited attribute Face image img2. Send img1 and img2 to the pre-trained target attribute classification model and face matching model, so that the difference between the two image attributes and identifiers (Identifier, id) can be calculated as a loss to train the attribute editing sub-model to be trained.
  • the attribute editing sub-model the input of the editing module is the attribute value, for example, input 1 for wearing glasses, 0 for not wearing glasses, 1 for smiling, and 0 for not smiling
  • the trained stylegan generator will generate the edited attribute Face image img2.
  • a large number of facial images can also be collected for attribute labeling, such as whether to wear glasses, and the resnet model is used to train the attribute classifier to obtain the target attribute classification model. And pre-training the face recognition model to get the face matching model.
  • the training data can be input to the feature preprocessing sub-model (the first feature extraction module or the second feature extraction module) to obtain the corresponding feature vector;
  • the feature vector is used as the input of the attribute editing sub-model to be trained , the feature vector to be processed after splicing the feature vector and the preset feature vector can be obtained;
  • the feature vector to be processed is used as the input of the second image generation submodule to output the special effect facial image, and the special effect facial image at this time and the preset facial features may be There are certain deviations.
  • the feature vector is used as the input of the first image generating submodule to obtain the facial image corresponding to the feature vector.
  • the special effect facial image and the facial image are used as the input of the target attribute classification model and the facial matching model respectively, and the feature difference and facial matching degree of the two images are obtained.
  • the facial attribute determination model to be used refers to the training of the facial attribute determination model to be trained, and the trained model.
  • the facial attribute determination model to be trained is trained by using the sample data to obtain the facial attribute determination model to be used.
  • some modules that will not actually be used in the facial attribute determination model to be used may be trimmed, and the trimmed model may be used as the target facial attribute determination model.
  • the target facial attribute determination model is obtained by removing the target attribute classification model, the facial matching model, and the first image generation sub-model in the facial attribute determination model to be used.
  • S350 Process the data to be processed based on a target facial attribute determination model to obtain a target facial image corresponding to the data to be processed.
  • the facial attribute determination model to be trained is constructed to train the facial attribute determination model to be trained to obtain a high-precision facial attribute determination model to be used, and the facial attribute determination model to be used is trimmed.
  • the target facial attribute determination model is obtained, so as to improve the accuracy of the target facial attribute determination model and the efficiency of adding special effects, and then make the special effects added to the image more realistic.
  • Fig. 7 is a schematic flow chart of an image processing method provided in Embodiment 4 of the present disclosure.
  • the determination of the facial attributes to be used is obtained through the training process of the facial attribute determination model to be trained.
  • the model is refined, and its implementation can refer to the technical solution of this embodiment. Wherein, technical terms that are the same as or corresponding to those in the foregoing embodiments will not be repeated here.
  • the method includes the following steps:
  • the training samples include data to be trained.
  • the training sample may be a sample used to train the model, and the parameter values of the model may be adjusted during the training process to make the output result of the model consistent with the expected result.
  • the data to be trained may be data used for training the model, and may be Gaussian noise, for example, random noise of Gaussian sampling. It can also be an image. For example, the user subject can be photographed based on different viewing angles, and facial images corresponding to the user subject at different viewing angles can be generated. These images can be used as training data, and correspondingly, multiple training samples can be obtained .
  • a large number of facial images can be collected as training samples, and the stylegan model can be trained using facial images, so that after the model training is completed, the stylegan generator can be generated by inputting Gaussian sampling random noise z ⁇ N(0,1) Different types of face pictures.
  • Multiple training samples can be stored in a preset database, and then the multiple training samples in the database can be extracted by using the interface.
  • the eigenvector of any training sample can be determined as the eigenvector of the current training sample, so as to illustrate that one of the training samples is used as the current training sample.
  • the first feature vector refers to the feature vector output by the first feature extraction module or the second feature extraction module. For example, after the current training sample is input to the first feature extraction module or the second feature extraction module, the extracted feature vector can be extracted by the module. eigenvector as the first eigenvector.
  • the data to be trained in the training sample is different, it may be Gaussian noise or an image, and then, the training data can be processed differently based on the first feature extraction module and the second feature extraction module, for example, the current training sample can be
  • the data to be trained in is input into the feature preprocessing sub-model. If the data to be trained is Gaussian noise, the Gaussian noise can be processed based on the first feature extraction module, and the feature vector corresponding to the Gaussian noise can be obtained; if the data to be trained is Gaussian noise is an image, the image can be processed based on the second feature extraction module, and a feature vector corresponding to the image can be obtained. Both the feature vectors output by the first feature extraction module and the second feature extraction module may be used as the first feature vectors corresponding to the current training sample.
  • the attribute feature vector may be a vector representation corresponding to a preset facial feature, and a feature vector corresponding to any preset facial feature may be used as the attribute feature vector.
  • the first attribute feature vector may be a feature vector obtained by concatenating the first feature vector and the attribute feature vector, and the vector output by the attribute editing sub-model to be trained may be used as the first attribute feature vector.
  • the first feature vector corresponding to the current training sample can be used as the input of the attribute editing sub-model to be trained, and the sub-model can splice the first feature vector and the attribute feature vector corresponding to the preset facial features, and then can output the added preset face
  • the first attribute feature vector corresponding to the feature so that the facial image with special effects can be generated based on the first attribute feature vector, and the model parameters can be adjusted.
  • the feature image without attributes may be an image without special effects in the facial image, and the image output by the first image generation sub-module may be used as the feature image without attributes.
  • the attached attribute feature image may be an image with special effects attached to the facial image, and the image output by the second image generation sub-module may be used as the attached attribute feature image.
  • the first feature vector corresponding to the current training sample can be used as the input of the first image generation sub-module, and the model can perform image reconstruction on the first feature vector.
  • the first feature vector has not been processed by the attribute editing model to be trained, and no features have been added.
  • the model can output images without attribute features.
  • the first attribute feature vector corresponding to the current training sample can also be input into the second image generation sub-module, because the first attribute feature vector is a vector after adding special effects, and then the model can output images with attribute features.
  • the attribute information to be compared may be image facial features output by the target attribute classification model.
  • the facial matching information may be the facial image matching degree output by the facial matching model.
  • the image with and without attributes corresponding to the current training sample can be used as the input of the target attribute classification model, and the facial features corresponding to the two images can be output, that is, the attribute information to be compared. It is also possible to use the image with attribute feature and the image without attribute feature as the input of the face matching model, and output the matching degree corresponding to the two images, that is, face matching information.
  • the target loss value can be used to characterize the loss between the attribute information to be compared and the preset facial features and the loss between facial matching information.
  • the loss function in the attribute editing sub-model to be trained can be used to perform loss processing on the attribute information to be compared and the preset facial features, and then the loss value between the two can be calculated.
  • a loss function can also be used to perform loss processing on the face matching information, and then a corresponding loss value can be calculated.
  • all the calculated loss values can also be fused, and then the fused loss value can be obtained, which can be used as the target loss value, so that the model parameters in the model can be adjusted based on the target loss value fix.
  • the convergence of the preset loss function can be used as the training goal.
  • the preset loss function of the attribute editing sub-model to be trained converges, it indicates that the adjustment result meets the requirements of the scheme, and the trained model has been obtained, so as to obtain the facial attribute determination model to be used.
  • the attribute feature addition technology can be used to process the first feature vector corresponding to the current training sample, and the attribute editing sub-model to be trained can output the first attribute feature vector corresponding to the current training sample, so that based on the first The attribute feature vector generates the accompanying attribute feature image corresponding to the current training sample.
  • the model parameters in the attribute editing sub-model to be trained are uncorrected, there is a corresponding difference between the obtained attribute feature image and the unattached attribute feature image corresponding to the current training sample after the actual feature is added, which can be based on the current
  • the attribute information to be compared, facial matching information, and preset facial features corresponding to the two types of images corresponding to the training samples are processed to determine the error value, and then based on the error value, the model parameters in the attribute editing sub-model to be trained can be corrected.
  • the training error of the loss function can be used as a condition for detecting whether the loss function is currently converged, such as whether the training error is smaller than the preset error or whether the error trend is stable, or whether the current number of iterations is equal to the preset number.
  • the training sample data can be obtained to continue training the sub-model to be trained for attribute editing until the training error of the loss function is within the preset range.
  • the attribute editing sub-model to be trained has been trained, so that when the first feature vector is input into the trained attribute editing sub-model to be trained, the model can be accurately the first feature vector Concatenate attribute feature vectors so that images with facial features can be generated.
  • S480 Eliminate the target attribute classification model, the face matching model, and the first image generation sub-model from the facial attribute determination models to be used to obtain the target facial attribute determination model.
  • the model includes not only the feature preprocessing sub-model and the second image generation sub-module, but also the first image generation sub-module, target attribute classification model and face matching model.
  • the first image generation submodule and the target attribute classification model in the facial attribute determination model to be used can be combined and facial matching model culling processing. That is, remove models that will not be used in the application. For example, after the facial attribute determination model is trained, only the branch path of the attribute editing sub-model is reserved, and the edited facial image is generated by inputting attribute values.
  • S4100 Process the data to be processed based on a target facial attribute determination model to obtain a target facial image corresponding to the data to be processed.
  • the model parameters in the attribute editing sub-model to be trained are continuously optimized, and then the facial attribute determination model to be used is obtained, and the facial attribute determination model to be used is trimmed to obtain
  • the target facial attribute determination model is used to improve the accuracy of the target facial attribute determination model and the efficiency of adding special effects, thereby making the special effects added to the image more realistic.
  • Fig. 8 is a structural block diagram of an image processing device provided in Embodiment 5 of the present disclosure, which can execute the image processing method provided in any embodiment of the present disclosure, and has corresponding functional modules and beneficial effects for executing the method.
  • the device includes: a data acquisition module 510 and a target image determination module 520 .
  • the data acquisition module 510 is configured to acquire data to be processed; wherein the data to be processed includes Gaussian noise or an image to be converted; the target image determination module 520 is configured to process the data to be processed based on a target facial attribute determination model, Obtaining a target facial image corresponding to the data to be processed; wherein at least one target feature in the target facial image matches at least one corresponding preset facial feature.
  • the target image determination module 520 includes a feature vector determination unit to be spliced, a target feature vector acquisition unit, and a target facial image acquisition unit.
  • the feature vector to be spliced determination unit is configured to determine the feature vector to be spliced corresponding to the data to be processed;
  • the target feature vector acquisition unit is configured to stitch the feature vector to be spliced with the at least one preset facial feature A corresponding preset feature vector is obtained to obtain a target feature vector corresponding to the target facial image;
  • a target facial image acquisition unit is configured to process the target feature vector to obtain the target facial image.
  • the feature preprocessing sub-model includes a first feature extraction module and a second feature extraction module, a feature vector determination unit to be spliced, including a first unit for determining a feature vector to be spliced and a feature vector to be spliced determination unit Second unit.
  • the feature vector to be spliced determines the first unit, which is set to determine the feature vector to be spliced corresponding to the Gaussian noise based on the first feature extraction module if the data to be processed is the Gaussian noise; the feature vector to be spliced determines the first
  • the second unit is configured to determine, based on the second feature extraction module, a feature vector to be spliced corresponding to the image to be converted if the data to be processed is the image to be converted.
  • the device further includes: a target facial attribute determination model parameter correction module.
  • the target facial attribute determination model parameter correction module is configured to determine at least one target attribute corresponding to the target facial image based on the pre-trained attribute classifier, so as to classify the target face based on the at least one target attribute
  • the model parameters in the attribute determination model are corrected; wherein, the at least one target attribute matches the at least one attribute identifier of the preset facial feature.
  • the device further includes: a construction module for determining a facial attribute to be trained.
  • the facial attribute determination model construction module to be trained is set to edit the sub-model according to the attribute to be trained, the target confrontation model obtained in advance, the target attribute classification model and the facial matching model, and construct the facial attribute determination model to be trained; wherein, the The target attribute classification model is configured to determine the facial features of the image output by the target confrontation model; the facial matching model is configured to determine the matching degree of the facial image output based on the target confrontation model, and the target confrontation model, It is set to output two facial images, one of which matches the preset facial features set in the attribute editing sub-model to be trained.
  • the target confrontation model includes: a feature preprocessing sub-model and an image generation sub-model; the feature preprocessing sub-model includes a first feature extraction module; the image generation sub-model includes a first An image generation sub-module and a second image generation sub-module; the feature preprocessing sub-model also includes a second feature extraction module.
  • the facial attribute determination model construction module to be trained further includes a facial attribute determination model construction unit to be trained.
  • the facial attribute determination model construction unit to be trained is configured to use the output result of the first feature extraction module or the second feature extraction module as the input of the attribute editing sub-model to be trained and the first image generation sub-module,
  • the output of the attribute editing sub-model to be trained is used as the input of the second image generation sub-module, and the output of the first image generation sub-module and the output of the second image generation sub-module are used as the target
  • the attribute classification model and the input of the facial matching model are used to construct the facial attribute determination model to be trained.
  • the facial attribute determination model acquisition unit to be used also includes a training sample acquisition subunit, a first feature vector acquisition subunit, a first attribute feature vector acquisition subunit, and an attribute feature image acquisition subunit , an information acquisition subunit, a target loss value determination subunit, and a facial attribute determination model acquisition subunit to be used.
  • the training sample acquisition subunit is set to acquire multiple training samples, wherein the training samples include data to be trained; the first feature vector acquisition subunit is set to input the data to be trained in the current training sample for multiple training samples To the first feature extraction module or the second feature extraction module to obtain the first feature vector corresponding to the current training sample; the first attribute feature vector acquisition subunit is set to be edited based on the attribute to be trained
  • the sub-model splices the attribute feature vector corresponding to at least one preset facial feature for the first feature vector to obtain the first attribute feature vector;
  • the attribute feature image acquisition subunit is configured to input the first feature vector to In the first image generation sub-module, an image without attribute features is obtained; and, the first attribute feature vector is input into the second image generation sub-module to obtain an image with attribute features;
  • the information acquisition subunit It is set to input the feature image with attribute and the feature image without attribute into the target attribute classification model to obtain attribute information to be compared; and, input the feature image with attribute and the feature image without attribute , input into the
  • the preset attribute value in the attribute editing sub-model is processed to obtain the target loss value; the facial attribute to be used determines the model acquisition subunit, which is set to the model parameters in the attribute editing sub-model to be trained based on the target loss value Correction is performed, and the convergence of the loss function is used as a training target, and a facial attribute determination model to be used is obtained through training.
  • the target facial attribute determination model acquisition unit includes a target facial attribute determination model acquisition subunit.
  • the target facial attribute determination model acquisition subunit is configured to eliminate the target attribute classification model, the facial matching model, and the second image generation sub-model in the facial attribute determination model to be used to obtain the target face Attributes determine the model.
  • the preset facial features include facial features of wearing at least one kind of jewelry, facial features of different age stages, facial features of different angles, facial features of different hairstyles, facial features of different hairstyles and color matching, and At least one of facial features of different expressions.
  • the technical solution of the embodiment of the present disclosure obtains the Gaussian noise or the data to be processed of the image to be converted, and then determines the model based on the target face attribute to process the data to be processed, and obtains the target face image corresponding to the data to be processed after adding special effects, to solve the problem
  • the obtained special effect images have low authenticity, which causes the problem of poor user experience, and realizes adding corresponding facial features to the faces of the target objects in the images to be processed, so as to
  • the obtained special effects have a high degree of realism, and the effect of increasing the richness and interest of the video image content improves the technical effect of user experience.
  • the image processing device provided in the embodiments of the present disclosure can execute the image processing method provided in any embodiment of the present disclosure, and has corresponding functional modules and effects for executing the method.
  • the multiple units and modules included in the above-mentioned device are only divided according to functional logic, but are not limited to the above-mentioned division, as long as the corresponding functions can be realized; in addition, the names of multiple functional units are only for the convenience of distinguishing each other , and are not intended to limit the protection scope of the embodiments of the present disclosure.
  • FIG. 9 is a schematic structural diagram of an electronic device provided by Embodiment 6 of the present disclosure.
  • the terminal equipment in the embodiments of the present disclosure may include but not limited to mobile phones, notebook computers, digital broadcast receivers, personal digital assistants (Personal Digital Assistant, PDA), tablet computers (Portable Android Device, PAD), portable multimedia players (Portable Media Player, PMP), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., and fixed terminals such as digital televisions (Television, TV), desktop computers, etc.
  • the electronic device 600 shown in FIG. 9 is only an example, and should not limit the functions and scope of use of the embodiments of the present disclosure.
  • an electronic device 600 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) Various appropriate actions and processes are performed by a program loaded into a random access memory (Random Access Memory, RAM) 603 by 608. In the RAM 603, various programs and data necessary for the operation of the electronic device 600 are also stored.
  • the processing device 601, ROM 602, and RAM 603 are connected to each other through a bus 604.
  • An input/output (Input/Output, I/O) interface 605 is also connected to the bus 604 .
  • an input device 606 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; including, for example, a liquid crystal display (Liquid Crystal Display, LCD) , an output device 607 such as a speaker, a vibrator, etc.; a storage device 608 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 609.
  • the communication means 609 may allow the electronic device 600 to communicate with other devices wirelessly or by wire to exchange data.
  • FIG. 9 shows electronic device 600 having various means, it is not required to implement or possess all of the means shown. More or fewer means may alternatively be implemented or provided.
  • embodiments of the present disclosure include a computer program product, which includes a computer program carried on a non-transitory computer readable medium, where the computer program includes program code for executing the method shown in the flowchart.
  • the computer program may be downloaded and installed from a network via communication means 609, or from storage means 608, or from ROM 602.
  • the processing device 601 When the computer program is executed by the processing device 601, the above-mentioned functions defined in the methods of the embodiments of the present disclosure are performed.
  • the electronic device provided by the embodiment of the present disclosure belongs to the same concept as the image processing method provided by the above embodiment, and the technical details not described in detail in this embodiment can be referred to the above embodiment, and this embodiment has the same effect as the above embodiment .
  • Embodiment 7 of the present disclosure provides a computer storage medium on which a computer program is stored, and when the program is executed by a processor, the image processing method provided in the above embodiment is implemented.
  • the computer-readable medium mentioned above in the present disclosure may be a computer-readable signal medium or a computer-readable storage medium, or any combination of the above two.
  • a computer readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or device, or any combination thereof.
  • Examples of computer readable storage media may include, but are not limited to: electrical connections with one or more wires, portable computer disks, hard disks, RAM, ROM, Erasable Programmable Read-Only Memory (EPROM) or flash memory), optical fiber, portable compact disk read-only memory (Compact Disc Read-Only Memory, CD-ROM), optical storage device, magnetic storage device, or any suitable combination of the above.
  • a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.
  • a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave carrying computer-readable program code therein. Such propagated data signals may take many forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the foregoing.
  • a computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium, which can transmit, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device .
  • the program code contained on the computer readable medium can be transmitted by any appropriate medium, including but not limited to: electric wire, optical cable, radio frequency (Radio Frequency, RF), etc., or any suitable combination of the above.
  • the client and the server can communicate using any currently known or future network protocols such as Hypertext Transfer Protocol (HyperText Transfer Protocol, HTTP), and can communicate with digital data in any form or medium
  • the communication eg, communication network
  • Examples of communication networks include local area networks (Local Area Network, LAN), wide area networks (Wide Area Network, WAN), internetworks (e.g., the Internet), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently existing networks that are known or developed in the future.
  • the above-mentioned computer-readable medium may be included in the above-mentioned electronic device, or may exist independently without being incorporated into the electronic device.
  • the above-mentioned computer-readable medium carries one or more programs, and when the above-mentioned one or more programs are executed by the electronic device, the electronic device:
  • the data to be processed includes Gaussian noise or an image to be converted; based on the target facial attribute determination model, the data to be processed is processed to obtain a target facial image corresponding to the data to be processed; wherein , at least one target feature in the target facial image matches corresponding at least one preset facial feature.
  • Computer program code for carrying out operations of the present disclosure may be written in one or more programming languages, or combinations thereof, including but not limited to object-oriented programming languages—such as Java, Smalltalk, C++, and Includes conventional procedural programming languages - such as the "C" language or similar programming languages.
  • the program code may execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server.
  • the remote computer can be connected to the user computer through any kind of network, including a LAN or WAN, or it can be connected to an external computer (eg via the Internet using an Internet Service Provider).
  • each block in a flowchart or block diagram may represent a module, program segment, or portion of code that contains one or more logical functions for implementing specified executable instructions.
  • the functions noted in the block may occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently, or they may sometimes be executed in the reverse order, depending upon the functionality involved.
  • each block of the block diagrams and/or flowchart illustrations, and combinations of blocks in the block diagrams and/or flowchart illustrations can be implemented by a dedicated hardware-based system that performs the specified functions or operations , or may be implemented by a combination of dedicated hardware and computer instructions.
  • the units involved in the embodiments described in the present disclosure may be implemented by software or by hardware.
  • the name of the unit does not constitute a limitation on the unit itself in one case, for example, the first obtaining unit may also be described as "a unit for obtaining at least two Internet Protocol addresses".
  • exemplary types of hardware logic components include: Field Programmable Gate Arrays (Field Programmable Gate Arrays, FPGAs), Application Specific Integrated Circuits (ASICs), Application Specific Standard Products (Application Specific Standard Parts, ASSP), System on Chip (System on Chip, SOC), Complex Programmable Logic Device (Complex Programming Logic Device, CPLD) and so on.
  • a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device.
  • a machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium.
  • a machine-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Examples of machine-readable storage media would include one or more wire-based electrical connections, portable computer disks, hard drives, RAM, ROM, EPROM or flash memory, optical fibers, CD-ROMs, optical storage devices, magnetic storage devices, or Any suitable combination of the above.
  • Example 1 provides an image processing method, the method including:
  • the data to be processed includes Gaussian noise or an image to be converted
  • Example 2 provides an image processing method, further comprising:
  • the determining model based on the target facial attributes processes the data to be processed to obtain a target facial image corresponding to the data to be processed, including:
  • the target feature vector is processed to obtain the target facial image.
  • Example 3 provides an image processing method, further comprising:
  • the determining the feature vector to be spliced corresponding to the data to be processed includes:
  • the feature vector to be spliced corresponding to the Gaussian noise is determined based on the first feature extraction module
  • the feature vector to be spliced corresponding to the image to be converted is determined.
  • Example 4 provides an image processing method, further comprising:
  • the target facial image corresponding to the data to be processed After obtaining the target facial image corresponding to the data to be processed, it also includes:
  • the at least one target attribute matches the attribute identifier of the at least one preset facial feature.
  • Example 5 provides an image processing method, further comprising:
  • the target facial attribute determination model is obtained by clipping the to-be-used facial attribute determination model.
  • Example 6 provides an image processing method, further comprising:
  • Described construction to be trained facial attribute determination model comprises:
  • the target confrontation model obtained in advance, the target attribute classification model and the face matching model, construct the facial attribute determination model to be trained;
  • the target attribute classification model is used to determine the facial features of the image output by the target confrontation model;
  • the facial matching model is used to determine the matching degree of the facial image output based on the target confrontation model,
  • the target confrontation model is used to output two facial images, one of which matches the preset facial features set in the attribute editing sub-model to be trained.
  • Example 7 provides an image processing method, further comprising:
  • the target confrontation model includes: a feature preprocessing sub-model and an image generation sub-model; the feature preprocessing sub-model includes a first feature extraction module; the image generation sub-model includes a first image generation sub-module and a second image generation sub-model. An image generation sub-module; the feature preprocessing sub-model also includes a second feature extraction module.
  • Example 8 provides an image processing method, further comprising:
  • the construction of the facial attribute determination model to be trained includes: using the output result of the first feature extraction module or the second feature extraction module as the input of the attribute editing sub-model to be trained and the first image generation sub-module , using the output of the attribute editing sub-model to be trained as the input of the second image generation sub-module, and using the output of the first image generation sub-module and the output of the second image generation sub-module as the input of the second image generation sub-module.
  • the target attribute classification model and the input of the facial matching model are used to construct the facial attribute determination model to be trained.
  • Example 9 provides an image processing method, further comprising:
  • the facial attribute determination model to be used is obtained through the training process of the facial attribute determination model to be trained, including:
  • training samples include data to be trained
  • the model parameters in the attribute editing sub-model to be trained are corrected, and the convergence of the loss function is used as a training target to obtain a facial attribute determination model to be used through training.
  • Example 10 provides an image processing method, further comprising:
  • the described target facial attribute determination model is obtained by tailoring the facial attribute determination model to be used, including:
  • the target facial attribute determination model is obtained by removing the target attribute classification model, the facial matching model, and the first image generation sub-model in the facial attribute determination model to be used.
  • Example Eleven provides an image processing method, further comprising:
  • the preset facial features include at least one of the facial features of wearing at least one accessory, facial features of different ages, facial features of different angles, facial features of different hairstyles, facial features of different hairstyles and color matching, and facial features of different expressions. A sort of.
  • Example 12 provides an image processing device, including:
  • a data acquisition module configured to acquire data to be processed; wherein, the data to be processed includes Gaussian noise or an image to be converted;
  • the target image determination module is configured to process the data to be processed based on the target facial attribute determination model to obtain a target facial image corresponding to the data to be processed; wherein at least one target feature in the target facial image is the same as corresponding to at least one preset facial feature.

Landscapes

  • Engineering & Computer Science (AREA)
  • Health & Medical Sciences (AREA)
  • Oral & Maxillofacial Surgery (AREA)
  • Theoretical Computer Science (AREA)
  • General Health & Medical Sciences (AREA)
  • Physics & Mathematics (AREA)
  • General Physics & Mathematics (AREA)
  • Computer Vision & Pattern Recognition (AREA)
  • Multimedia (AREA)
  • Human Computer Interaction (AREA)
  • Evolutionary Computation (AREA)
  • Artificial Intelligence (AREA)
  • Bioinformatics & Cheminformatics (AREA)
  • Bioinformatics & Computational Biology (AREA)
  • Data Mining & Analysis (AREA)
  • Evolutionary Biology (AREA)
  • Life Sciences & Earth Sciences (AREA)
  • General Engineering & Computer Science (AREA)
  • Computing Systems (AREA)
  • Databases & Information Systems (AREA)
  • Medical Informatics (AREA)
  • Software Systems (AREA)
  • Image Analysis (AREA)
  • Image Processing (AREA)
  • Collating Specific Patterns (AREA)

Abstract

本公开提供了一种图像处理方法、装置、电子设备和存储介质。该图像处理方法包括:获取待处理数据;基于目标面部属性确定模型对所述待处理数据进行处理,得到与所述待处理数据相对应的目标面部图像;其中,所述目标面部图像中的至少一个目标特征与相应的至少一种预设面部特征相匹配。

Description

图像处理方法、装置、电子设备和存储介质
本申请要求在2021年12月29日提交中国专利局、申请号为202111641195.4的中国专利申请的优先权,该申请的全部内容通过引用结合在本申请中。
技术领域
本公开涉及计算机技术领域,例如涉及一种图像处理方法、装置、电子设备和存储介质。
背景技术
随着科技的发展,越来越多的应用软件走进了用户的生活,逐渐丰富了用户的业余生活,例如短视频应用程序等。用户可以采用视频、照片等方式记录生活,并上传到短视频应用程序上。为了提高视频、照片内容的趣味性,通常会为图像中的对象添加相应的特效。
添加特效的技术,通常利用软件修图的方式为图像中的对象添加特效,如,使对象转头,该方法在添加特效时,容易造成对象肢体、面部等扭曲,导致特效添加效果不佳的问题。
发明内容
本公开提供一种图像处理方法、装置、电子设备和存储介质,以实现为对象添加面部特效,使生成出来的图像与在实际中面部变化呈现的效果最适配,提高面部特效添加的准确性。
第一方面,本公开提供了一种图像处理方法,该方法包括:
获取待处理数据;其中,所述待处理数据包括高斯噪声或待转换图像;
基于目标面部属性确定模型对所述待处理数据进行处理,得到与所述待处理数据相对应的目标面部图像;其中,所述目标面部图像中的至少一个目标特征与相应的至少一种预设面部特征相匹配。
第二方面,本公开还提供了一种图像处理装置,该装置包括:
数据获取模块,设置为获取待处理数据;其中,所述待处理数据包括高斯噪声或待转换图像;
目标图像确定模块,设置为基于目标面部属性确定模型对所述待处理数据进行处理,得到与所述待处理数据相对应的目标面部图像;其中,所述目标面 部图像中的至少一个目标特征与相应的至少一种预设面部特征相匹配。
第三方面,本公开还提供了电子设备,所述设备包括:
一个或多个处理器;
存储装置,设置为存储一个或多个程序;
当所述一个或多个程序被所述一个或多个处理器执行,使得所述一个或多个处理器实现上述的图像处理方法。
第四方面,本公开还提供了一种计算机可读存储介质,其上存储有计算机程序,该程序被处理器执行时实现上述的图像处理方法。
第五方面,本公开还提供了一种计算机程序产品,包括承载在非暂态计算机可读介质上的计算机程序,所述计算机程序包含用于执行上述的图像处理方法的程序代码。
附图说明
图1为本公开实施例一所提供的一种图像处理方法流程示意图;
图2是本公开实施例一所提供的一种与预设面部特征相匹配的目标面部图像示意图;
图3是本公开实施例二所提供的一种图像处理方法的流程示意图;
图4是本公开实施例二所提供的一种目标面部属性确定模型的结构示意图;
图5是本公开实施例三所提供的一种图像处理方法的流程示意图;
图6是本公开实施例三所提供的一种待训练面部属性确定模型的结构示意图;
图7是本公开实施例四所提供的一种图像处理方法的流程示意图;
图8是本公开实施例五所提供的一种图像处理装置的结构框图;
图9是本公开实施例六所提供的一种电子设备的结构示意图。
具体实施方式
下面将参照附图描述本公开的实施例。虽然附图中显示了本公开的一些实施例,然而本公开可以通过多种形式来实现,提供这些实施例是为了理解本公开。本公开的附图及实施例仅用于示例性作用。
本公开的方法实施方式中记载的多个步骤可以按照不同的顺序执行,和/或并行执行。此外,方法实施方式可以包括附加的步骤和/或省略执行示出的步骤。 本公开的范围在此方面不受限制。
本文使用的术语“包括”及其变形是开放性包括,即“包括但不限于”。术语“基于”是“至少部分地基于”。术语“一个实施例”表示“至少一个实施例”;术语“另一实施例”表示“至少一个另外的实施例”;术语“一些实施例”表示“至少一些实施例”。其他术语的相关定义将在下文描述中给出。
本公开中提及的“第一”、“第二”等概念仅用于对不同的装置、模块或单元进行区分,并非用于限定这些装置、模块或单元所执行的功能的顺序或者相互依存关系。本公开中提及的“一个”、“多个”的修饰是示意性而非限制性的,本领域技术人员应当理解,除非在上下文另有指出,否则应该理解为“一个或多个”。
在介绍本技术方案之前,可以先对应用场景进行示例性说明。可以将本公开技术方案应用在任意需要特效展示的场景中,例如,视频通话中,可以进行特效展示;或者,直播场景中,可以对主播对象进行特效展示;也可以是应用在视频拍摄过程中,对被拍摄对象所对应的图像进行特效展示的情况,如,短视频拍摄场景下,可以将拍摄的图像处理为特效图像,进而将处理后的特效图像进行特效展示;也可以是应用在静态图像拍摄过程中,例如,通过终端设备自带摄像机拍摄图像后,将拍摄的图像处理成特效图像进行特效展示的情况。
实施例一
图1为本公开实施例一所提供的一种图像处理方法流程示意图,本公开实施例适用于在互联网所支持的任意图像展示场景中,用于将目标对象的面部图像处理为特效图像并展示的情形,该方法可以由图像处理装置来执行,该装置可以通过软件和/或硬件的形式实现,例如,通过电子设备来实现,该电子设备可以是移动终端、个人电脑(Personal Computer,PC)端或服务器等。任意图像展示的场景通常是由客户端和服务器来配合实现的,本实施例所提供的方法可以由服务端来执行,客户端来执行,或者是客户端和服务端的配合来执行。
如图1,本实施例的方法包括:
S110、获取待处理数据;其中,所述待处理数据包括高斯噪声或待转换图像。
上述已对多种可以应用的场景进行简单的说明,在此不再阐述。其中,执行本公开实施例提供的图像处理方法的装置,可以集成在支持图像处理功能的应用软件中,且该软件可以安装至电子设备中,例如,电子设备可以是移动终端或者PC端等。应用软件可以是对图像/视频处理的一类软件,其应用软件在此不再一一赘述,只要可以实现图像/视频处理即可。
待处理数据可以为需要被处理的数据,可以为高斯噪声,也可以为图像。 高斯噪声可以是随机采样噪声,高斯噪声可以包括起伏噪声、宇宙噪声、热噪声和散粒噪声等至少一种噪声。待转换图像可以是基于应用软件采集的图像,也可以是应用软件从存储空间中预先存储的图像。在应用场景中,可以实时或者周期性的采集待转换图像。例如,在直播场景或者拍摄视频的场景中,摄像装置实时采集目标场景中包括目标对应的图像,此时,可以将摄像装置采集的图像作为待转换图像。相应的,待转换图像中可以包括目标主体,目标主体可以是用户、可以是宠物、可以是花草树木等等。还可以对拍摄视频对应的视频帧进行处理,如,可以预先设置与拍摄视频对应的目标主体,当在检测到视频帧对应的图像中有该目标主体时,可以将该视频帧对应的图像作为待转换图像,以使后续可以对视频中的每个视频帧图像中的目标主体进行面部特征处理。
同一拍摄场景中目标主体的数量可以是一个也可以是多个,不论是一个还是多个目标主体,都可以采用本公开所提供的技术方案来确定特效展示图。
在任意视频拍摄、直播场景中,可以实时或者间隔性的采集包括目标主体的图像,作为待转换图像。同时也可以利用设备随机采集噪声,得到高斯噪声。可以将高斯噪声和待转换图像作为模型的输入,对模型进行训练。
在本实施例中,将需要为面部图像添加相应特征时的原始面部数据,作为待处理数据。
S120、基于目标面部属性确定模型对所述待处理数据进行处理,得到与所述待处理数据相对应的目标面部图像。
目标面部属性确定模型可以是预先训练好的。目标面部属性确定模型可以对输入的高斯噪声或待转换图像进行处理,得到为高斯噪声或待转换图像中的面部添加相应特征后的图像。在训练得到目标面部属性确定模型之前,可以预先确定模型为用户添加的面部特征,可以将此时的面部特征作为预设面部特征。可以将目标面部属性确定模型输出的图像作为目标面部图像,此时目标面部图像中添加了预设面部特征。
在本实施例中,所述预设面部特征包括佩戴至少一种饰品的面部特征、不同年龄阶段的面部特征、不同角度的面部特征、不同发型的面部特征、不同发型配色的面部特征以及不同面部表情的面部特征中的至少一种。
佩戴的饰品可以是眼镜、墨镜以及面贴等,可以将佩戴的饰品对应的特征作为预设面部特征。不同年龄阶段的面部特征可以是不同年龄层的对象主体所对应的面部特征,不同年龄层可以是以原始采集图像为基准,在该基准上增加或减少年龄而划分出来的年龄层,不同年龄层可以是青年年龄层、中年年龄层、老年年龄层。不同角度的面部特征可以是对象主体面部朝向不同时所对应的特 征,例如,不同角度也可以以原始采集图像中主体对应的面部角度为基准,可以在该基准上向左偏转20度、向左偏转20度、向下偏转20度等等,也可以将这些角度下对应的面部特征作为预设面部特征。例如,如果想要目标面部属性确定模型输出添加佩戴眼镜、年龄增加20岁、向左偏转20度等特效的图像,可以将这些面部特征作为预设面部特征。预设面部特还可以是不同发型的面部特征,例如,长发、短发、卷发、直板等。不同发型配色的面部特征可以是紫色、白色、黑色等多种颜色,可以是不同发型和不同颜色的组合特征。不同表情的面部特征可以包括微笑、愤怒、生气等面部表情特征。
在本实施例中,可以将待处理数据输入至预先训练好的目标面部属性确定模型中,可以为待处理数据对应的图像添加相应特征,进而可以输出与待处理数据相对应的添加特征的面部图像,即目标面部图像。
示例性的,预设面部特征可以包括面部向右偏转20°(yaw+20°)、年龄增加20岁(age+20)、佩戴墨镜(wear glasses)、微笑(smlie)、面部向右偏转30°(yaw+30°)、年龄减少20岁(age-20)的图像对应的面部特征中的至少一个,那将原始图像(Origin)输入该模型之后,可以得到相应的示意图,示意图可以参见图2。
本公开实施例的技术方案,通过获取待处理数据,进而基于目标面部属性确定模型对待处理数据进行处理,得到与待处理数据相对应的添加特效之后的目标面部图像,解决了相关技术中利用修图技术生成特效图像时,得到的特效图真实度较低,从而引起用户体验不佳的问题,实现了为待处理数据中的目标对象的面部添加相应的面部特征,以使得到的特效真实度较高,以及增加了视频图像内容丰富性和趣味性的效果,提高了用户使用体验的技术效果。
实施例二
图3是本公开实施例二所提供的一种图像处理方法的流程示意图,在前述实施例的基础上,对S120进行说明,其实施方式可以参见本实施例技术方案。其中,与上述实施例相同或者相应的技术术语在此不再赘述。
如图3所示,该方法包括如下步骤:
S210、获取待处理数据;其中,所述待处理数据包括高斯噪声或待转换图像。
S220、确定与所述待处理数据相对应的待拼接特征向量。
所述目标面部属性确定模型中包括特征预处理子模型、属性编辑子模型以及图像生成子模型。通过这三个子模型对待处理数据进行处理,得到与待处理 数据相对应的目标面部图像。
特征预处理子模型可以用于提取相应的特征。属性编辑子模型可以是指为提取特征添加预设面部特征的模型。图像生成子模型可以是指为属性编辑子模型输出的特征进行图像生成的模型。待拼接特征向量是指特征预处理子模型输出的向量,可以将特征预处理子模型提取出的特征向量作为待拼接特征向量。
可以将待处理数据作为特征预处理子模型的输入,数据经过模型的特征提取处理,可以得到待处理数据对应的特征向量,可以将该特征向量作为待拼接特征向量。
示例性的,目标面部属性确定模型的结构示意图可以参见图4,该模型中可以包括特征预处理子模型、属性编辑子模型以及图像生成子模型。特征预处理子模型中包括第一特征提取模块和第二特征提取模块。
待处理数据中可以包括高斯噪声或待转换图像,为了数据处理的准确度。可以根据数据的不同,采用不同的处理方式,以基于不同的处理方式所使用的算法模块对不同的数据进行针对性的处理。其处理方式可以参见下述表述:
所述特征预处理子模型包括第一特征提取模块和第二特征提取模块,所述基于所述特征预处理子模型,确定与所述待处理数据相对应的待拼接特征向量,包括:若所述待处理数据为所述高斯噪声,则基于所述第一特征提取模块确定与所述高斯噪声相对应的待拼接特征向量;若所述待处理数据为所述待转换图像,则基于所述第二特征提取模块,确定与所述待转换图像相对应的待拼接特征向量。
第一特征提取模块用于提取高斯噪声所对应的特征向量。第二特征提取模块用于提取待转换图像中的面部属性所对应的特征向量。
在本实施例中,为了对高斯噪声和待转换图像分别处理,可以在特征预处理子模型中预设两个特征提取模块分别对两种数据进行对应处理。相应的,可以将待处理数据为高斯噪声的数据输入至第一特征提取模块,模块可以对高斯噪声进行处理,进而可以得到与高斯噪声对应的待拼接特征向量;还可以将待处理数据为待转换图像的数据输入至第二特征提取模块,模块可以对待转换图像进行处理,进而可以得到与待转换图像对应的待拼接特征向量,相应的,可以得到所有待处理数据对应的待拼接特征向量。
示例性的,继续参见图4可以基于第一特征提取模块对高斯噪声进行处理,输出与高斯噪声对应的待拼接特征向量。例如,第一特征提取模块可以为Mapping Network模型。可以基于第二特征提取模块对待转换图像进行处理,输出与待转换图像对应的待拼接特征向量。例如,第二特征提取模块可以为 Encoder模型,如,固定训练好的stylegan模型的生成器参数,训练Encoder模型,模型训练完成后,可以输入面部图像,经过Encoder编码再通过stylegan生成器,可以重建出这张面部图像。可以将输出的待拼接特征向量作为W+,用于后续模型的输入。
在实际应用中,在将数据输入至模型后,可以基于数据接口确定模型输入数据的类型,进而基于数据类型确定哪个模块对其进行处理。若待处理数据为高斯噪声,可以将高斯噪声作为第一特征提取模块的输入,模块可以输出高斯噪声相对应的待拼接特征向量;若待处理数据为待转换图像,可以将待转换图像作为第二特征提取模块的输入,模块可以输出待转换图像相对应的待拼接特征向量。
S230、为所述待拼接特征向量拼接与所述至少一个预设面部特征相对应的预设特征向量,得到与所述目标面部图像相对应的目标特征向量。
预设特征向量是指预设面部特征对应的向量,目标特征向量可以为待拼接特征向量与预设特征向量拼接后的特征向量。
在本实施例中,为了为待处理数据添加相应的特效,可以将待拼接特征向量与预设面部特征对应的特征向量进行拼接处理,进而可以得到待处理数据添加面部特征之后对应的目标特征向量,以基于目标特征向量生成带特效的目标面部图像。如,可以将待拼接特征向量输入至预先训练好的属性编辑子模型之后,子模型可以对待拼接特征向量与预设面部特征对应的预设特征向量进行拼接处理,如,可以将特征向量A与特征向量B拼接为A-B,相应的,可以得到拼接处理后的拼接特征向量,可以将该拼接特征向量作为目标面部图像对应的目标特征向量。
示例性的,继续参见图4,属性编辑子模型可以为Dynamic Network模型,如,可以利用Dynamic Network模型对待拼接特征向量与预设特征向量进行拼接处理,输出的W++即为目标面部图像对应的目标特征向量。如,输入的预设特征向量经过多层感知机(Multilayer Perceptron,MLP)编码,再经过两个全连接层(Fully Connected layers,FC)和一个激活函数sigmoid,与输入的待拼接特征向量进行乘加操作,得到目标特征向量。其中,预设特征向量对应的标识可以在Dynamic Network模型中进行设置。
可以将待拼接特征向量作为属性编辑子模型的输入,子模型可以将待拼接特征向量与至少一个预设面部特征相对应的预设特征向量拼接起来,进而可以得到添加预设面部特征的目标面部图像相对应的目标特征向量,以使后续可以基于目标特征向量生成带特效的目标面部图像。
S240、对所述目标特征向量进行处理,得到所述目标面部图像。
在本实施例中,为了生成待处理数据添加相应的特效之后的图像,可以将待拼接特征向量与预设特征向量拼接后对应的目标特征向量输入至预先训练好的图像生成子模型中,模型可以对目标特征向量进行重建处理,可以输出目标特征向量对应的目标面部图像。
示例性的,继续参见图4,图像生成子模型可以为Generator模型,如,可以利用Generator模型对目标特征向量进行处理,可以输出目标特征向量对应的目标面部图像。
在得到所述目标面部图像之后,为了可以根据实际输入模型中的原始图像以及模型输出的目标面部图像,对目标面部属性确定模型中的参数调优处理。例如,可以再添加一个训练好的模型属性分类器,属性分类器可以是用于对图像属性特征进行提取分类的模型。可以利用该属性分类器对目标面部图像中的特征数据进行提取,确定目标面部图像中的面部特征,以验证模型输出的图片是否添加了预设面部特征,进而,若模型输出的图片未添加预设面部特征,可以对模型中的参数进行修正,提高模型的准确性。
在得到所述目标面部图像之后,还包括:基于预先训练得到的属性分类器,确定与所述目标面部图像相对应的至少一种目标属性,以基于所述至少一种目标属性对所述目标面部属性确定模型中的模型参数进行修正;其中,所述至少一种目标属性与所述至少一种预设面部特征的属性标识相匹配。
属性标识可以是指预设面部特征相对应的标识,也就是说,预设面部特征可以用相应的标识进行表示,例如,年龄特征用A 1表示,角度特征用A 2表示、戴眼镜特征用A 3表示,那么A 1、A 2、A 3等标识可以作为对应特征的属性标识,目标属性也可以是预设面部特征对应的标识信息。
在本实施例中,为了对目标面部属性确定模型中的参数进行调优,可以对输出的目标面部图像与对应的理论添加特效的面部图像进行比较,那么相应的,还可以对输出的目标面部图像对应的属性特征与添加的特效特征进行比较,以基于计算目标面部图像对应的属性特征是否包含添加的特效特征,修正待训练分类模型中的模型参数。
可以将目标面部图像输入至预先训练好的属性分类器,进而分类器可以对目标面部图像进行特征提取处理,输出目标面部图像对应的特征属性,即目标属性。再将当前输出的特征属性与预设面部特征的属性标识进行比较,计算出比较误差值,进而可以基于比较误差值调整模型中的模型参数。
示例性的,继续参见图4,属性分类器可以为残差网络(Residual Networks, ResNet)模型,如,可以将目标面部图像输入至属性分类器中,图像经过分类器的特征提取处理,可以输出目标面部图像对应的目标属性。
可以将目标面部图像作为属性分类器的输入,分类器可以输出对应的目标属性。还可以利用算法将图像对应的目标属性与预设面部特征的属性标识进行误差处理,以根据得到的误差结果对目标面部属性确定模型的模型参数进行修正,提高目标面部属性确定模型的训练精度。
本公开实施例的技术方案,通过获取高斯噪声或待转换图像等待处理数据,进而基于目标面部属性确定模型对待处理数据进行处理,得到不同类型数据相对应的添加特效之后的目标面部图像,同时在得到目标面部图像之后,还基于目标面部图像对模型中的参数进行修正,提高了模型的准确性,进而也提高了添加特效的真实度,以及增加了视频图像内容丰富性和趣味性的效果,提高了用户使用体验的技术效果。
实施例三
图5是本公开实施例三所提供的一种图像处理方法的流程示意图,在前述实施例的基础上,还可以预先构建待训练面部属性确定模型,通过对待训练面部属性确定模型进行训练处理,得到目标面部属性确定模型,其实施方式可以参见本实施例技术方案。其中,与上述实施例相同或者相应的技术术语在此不再赘述。
如图5所示,该方法包括如下步骤:
S310、构建待训练面部属性确定模型。
待训练面部属性确定模型是指模型中的模型参数设置为默认值的模型,需要对该模型进行训练得到目标面部属性确定模型。
在本实例中,为了可以得到高精度确定面部属性特征的目标面部属性确定模型,可以基于预先构建的待训练面部属性确定模型进行训练,以使待训练面部属性确定模型训练完成之后,可以得到最终可应用的面部属性确定模型,即目标面部属性确定模型。
在本实施例中,所述构建待训练面部属性确定模型,包括:根据待训练属性编辑子模型,预先训练得到的目标对抗模型、目标属性分类模型以及面部匹配模型,构建所述待训练面部属性确定模型;其中,所述目标属性分类模型,用于确定经所述目标对抗模型输出的图像的面部特征;所述面部匹配模型,用于确定基于所述目标对抗模型输出的面部图像的匹配度,所述目标对抗模型,用于输出两幅面部图像,其中一幅面部图像与所述待训练属性编辑子模型中设 置的预设面部特征相匹配。
所述目标对抗模型中包括:特征预处理子模型和图像生成子模型;所述特征预处理子模型中包括第一特征提取模块;所述图像生成子模型包括第一图像生成子模块和第二图像生成子模块;所述特征预处理子模型中还包括第二特征提取模块。
目标对抗模型可以为stylegan模型。所述目标属性分类模型(resnet模型),用于确定经所述第二图像生成子模块输出的图像的面部特征。所述面部匹配模型(人脸识别模型),用于确定所述第一图像生成子模块和所述第二图像生成子模块输出的面部图像的匹配度。第一特征提取模块(Mapping Network模型),用于提取高斯噪声所对应的特征向量。第二特征提取模块(Encoder模型),用于提取图像中的面部属性所对应的特征向量。第一图像生成子模块(Generator模型),用于确定经特征预处理子模型输出的特征向量对应的图像。第二图像生成子模块(Generator模型),用于确定经待训练属性编辑子模型输出的特征向量对应的图像。待训练属性编辑子模型(Dynamic Network模型),用于对特征预处理子模型输出的特征向量与预设特征向量进行拼接。需要被训练的属性编辑子模型,此时的属性编辑子模型的输出结果可能还未达到预期结果,需要被训练,使训练后的模型的输出结果与预期结果相符,训练完成后,可以得到可应用的属性编辑子模型。示例性的,通过随机噪声或图像encoder可以得到图像编码(w code),设计两路stylegan生成器分支,其中一路分支直接经过训练好的stylegan生成器生成人脸图像img1;另一路分支经过一个待训练属性编辑子模型(编辑模块输入是属性值,如,戴眼镜输入1,不带眼镜输入0,微笑输入1,不微笑输入0),再经过训练好的stylegan生成器,生成编辑过属性的人脸图像img2。将img1和img2送入预训练的目标属性分类模型和面部匹配模型,以使后续可以通过计算两张图像属性和标识(Identifier,id)的差异作为损失,训练待训练属性编辑子模型。
在构建待训练面部属性确定模型之前,还可以收集大量面部图像,为其进行属性标注,如是否戴眼镜,利用resnet模型训练属性分类器,以得到目标属性分类模型。以及预训练面部识别模型,得到面部匹配模型。
结合图6进行解释说明,可以将训练数据输入至特征预处理子模型(第一特征提取模块或第二特征提取模块),得到对应的特征向量;将特征向量作为待训练属性编辑子模型的输入,可以得到特征向量与预设特征向量拼接后的待处理特征向量;将待处理特征向量作为第二图像生成子模块的输入,输出特效面部图像,此时的特效面部图像与预设面部特征可能存在一定的偏差。将特征向量作为第一图像生成子模块的输入,得到特征向量对应的面部图像。将特效 面部图像以及面部图像分别作为目标属性分类模型和面部匹配模型的输入,得到两幅图像的特征差异和面部匹配度。
S320、通过对所述待训练面部属性确定模型训练处理,得到待使用面部属性确定模型。
待使用面部属性确定模型是指对待训练面部属性确定模型进行训练,训练好的模型。
在获取到样本数据之后,利用样本数据对待训练面部属性确定模型进行训练,得到待使用面部属性确定模型。
S330、通过对所述待使用面部属性确定模型剪裁处理,得到所述目标面部属性确定模型。
在本实施例中,为了降低模型的冗余性,可以对在待使用面部属性确定模型中实际不会用到的部分模块进行剪裁处理,将裁减之后得到的模型作为目标面部属性确定模型。
将所述待使用面部属性确定模型中的所述目标属性分类模型、所述面部匹配模型、所述第一图像生成子模型剔除,得到所述目标面部属性确定模型。
S340、获取待处理数据;其中,所述待处理数据包括高斯噪声或待转换图像。
S350、基于目标面部属性确定模型对所述待处理数据进行处理,得到与所述待处理数据相对应的目标面部图像。
本公开实施例的技术方案,通过获取高斯噪声或待转换图像等待处理数据,进而基于目标面部属性确定模型对待处理数据进行处理,得到与待处理数据相对应的添加特效之后的目标面部图像,同时基于对不同类型的模型构建待训练面部属性确定模型,以基于对构建的待训练面部属性确定模型进行训练,得到高精度的待使用面部属性确定模型,将待使用面部属性确定模型进行剪裁处理后得到目标面部属性确定模型,以使提高目标面部属性确定模型的精度以及添加特效的效率,进而以使为图像添加的特效真实度较高。
实施例四
图7是本公开实施例四所提供的一种图像处理方法的流程示意图,在前述实施例的基础上,对所述通过对所述待训练面部属性确定模型训练处理,得到待使用面部属性确定模型作细化,其实施方式可以参见本实施例技术方案。其中,与上述实施例相同或者相应的技术术语在此不再赘述。
如图7所示,该方法包括如下步骤:
S410、获取多个训练样本。
在训练得到待使用面部属性确定模型之前,需要先获取训练样本,以基于训练样本来训练模型。为了提高模型的准确性,可以尽可能多而丰富的获取训练样本。
所述训练样本中包括待训练数据。训练样本可以是用于对模型进行训练的样本,可以在训练的过程中通过调整模型的参数值,使模型的输出结果与预期结果相符。待训练数据可以是用于对模型进行训练的数据,可以是高斯噪声,如,可以是高斯采样的随机噪声。也可以是图像,如,可以基于不同视线角度对用户主体进行拍摄,可以生成不同视线角度下用户主体对应的面部图像,可以将这些图像作为待训练数据,相应的,可以获取到多个训练样本。如,在实际应用中,可以收集大量面部图像作为训练样本,利用面部图像训练stylegan模型,以使模型训练完成后,stylegan生成器可以通过输入高斯采样的随机噪声z~N(0,1)生成不同类型的面部图片。
可以将多个训练样本存储在预设的数据库中,进而可以利用接口提取数据库中的多个训练样本。
S420、针对多个训练样本,将当前训练样本中的待训练数据输入至所述第一特征提取模块或所述第二特征提取模块中得到与所述当前训练样本相对应的第一特征向量。
当需要确定每个训练样本相对应的特征向量时,可以将确定任一训练样本的特征向量作为确定当前训练样本的特征向量进行处理,以对其中一个训练样本作为当前训练样本进行说明。第一特征向量是指第一特征提取模块或第二特征提取模块输出的特征向量,如,当将当前训练样本输入至第一特征提取模块或第二特征提取模块之后,可以将模块提取出的特征向量作为第一特征向量。
由于训练样本中待训练数据的数据不同,可能为高斯噪声,也可能为图像,进而,可以基于第一特征提取模块和第二特征提取模块对待训练数据进行区别处理,如,可以将当前训练样本中的待训练数据输入至特征预处理子模型中,若待训练数据为高斯噪声,则可以基于第一特征提取模块对高斯噪声进行处理,可以得到高斯噪声相对应的特征向量;若待训练数据为图像,则可以基于第二特征提取模块对图像进行处理,可以得到图像相对应的特征向量。可以将第一特征提取模块和第二特征提取模块输出的特征向量均作为当前训练样本相对应的第一特征向量。
S430、基于所述待训练属性编辑子模型为所述第一特征向量拼接与至少一 种预设面部特征相对应的属性特征向量,得到第一属性特征向量。
属性特征向量可以为预设面部特征对应的向量表示,可以将任意一个预设面部特征对应的特征向量作为属性特征向量。第一属性特征向量可以为第一特征向量与属性特征向量拼接后的特征向量,可以将待训练属性编辑子模型输出的向量作为第一属性特征向量。
可以将当前训练样本对应的第一特征向量作为待训练属性编辑子模型的输入,子模型可以将第一特征向量与预设面部特征相对应的属性特征向量拼接起来,进而可以输出添加预设面部特征后对应的第一属性特征向量,以使后续可以基于第一属性特征向量生成带特效的面部图像,对模型参数进行调整。
S440、将所述第一特征向量输入至所述第一图像生成子模块中,得到未附带属性特征图像;以及,将所述第一属性特征向量输入至所述第二图像生成子模块中,得到附带属性特征图像。
未附带属性特征图像可以为面部图像中未附带特效的图像,可以将第一图像生成子模块输出的图像作为未附带属性特征图像。附带属性特征图像可以为面部图像中附带了特效的图像,可以将第二图像生成子模块输出的图像作为附带属性特征图像。
可以将当前训练样本对应的第一特征向量作为第一图像生成子模块的输入,模型可以对第一特征向量进行图像重建,此时第一特征向量未经过待训练属性编辑模型处理,未添加特征,模型可以输出未附带属性特征的图像。同时还可以将当前训练样本对应的第一属性特征向量输入至第二图像生成子模块中,因为第一属性特征向量是经过添加特效处理后的向量,进而模型可以输出附带有属性特征的图像。
S450、将所述附带属性特征图像和所述未附带属性特征图像输入至所述目标属性分类模型中,得到待比较属性信息;以及,将所述附带属性特征图像和所述未附带属性特征图像,输入至所述面部匹配模型中,得到面部匹配信息。
待比较属性信息可以为目标属性分类模型输出的图像面部特征。面部匹配信息可以为面部匹配模型输出的面部图像匹配度。
可以将当前训练样本对应的附带属性特征图像和未附带属性特征图像作为目标属性分类模型的输入,可以输出两幅图像对应的面部特征,即待比较属性信息。还可以将附带属性特征图像和未附带属性特征图像作为面部匹配模型的输入,可以输出两幅图像对应的匹配度,即面部匹配信息。
S460、基于所述待训练属性编辑子模型中的损失函数对所述待比较属性信息、面部匹配信息以及所述待训练属性编辑子模型中的预设面部特征进行处理, 得到目标损失值。
目标损失值可以用于表征待比较属性信息与预设面部特征之间的损失以及面部匹配信息之间的损失。
可以利用待训练属性编辑子模型中的损失函数对待比较属性信息与预设面部特征进行损失处理,进而可以计算出两者之间的损失值。还可以利用损失函数对面部匹配信息进行损失处理,进而可以计算出对应的损失值。相应的,还可以将计算出的所有损失值进行融合处理,进而,可以得到融合后的损失值,可以将该损失值作为目标损失值,以使可以基于目标损失值对模型中的模型参数进行修正。
S470、基于所述目标损失值对所述待训练属性编辑子模型中的模型参数进行修正,并将所述损失函数收敛作为训练目标,训练得到待使用面部属性确定模型。
可以将预设损失函数收敛作为训练目标,当判定待训练属性编辑子模型的预设损失函数收敛时,表明调整结果符合方案要求,已经得到训练好的模型,从而得到待使用面部属性确定模型。
在本实施例中,可以利用属性特征添加技术将当前训练样本对应的第一特征向量进行处理,待训练属性编辑子模型可以输出与当前训练样本对应的第一属性特征向量,以使基于第一属性特征向量生成当前训练样本对应的附带属性特征图像。由于待训练属性编辑子模型中的模型参数是未修正的,那么,得到的附带属性特征图像与当前训练样本对应的未附带属性特征图像实际添加特征之后的图像也是存在相应差异的,可以基于当前训练样本对应的两类图像对应的待比较属性信息、面部匹配信息以及预设面部特征进行处理,可以确定出误差值,进而基于误差值可以修正待训练属性编辑子模型中的模型参数。
可以利用待训练属性编辑子模型中的损失函数对当前训练样本对应的待比较属性信息与预设属性值进行比较,计算出损失值,还可以计算出面部匹配信息对应的相似误差值,进而可以基于损失值和相似误差值计算出目标损失值,以根据得到的损失结果对待训练属性编辑子模型的模型参数进行修正。可以将损失函数的训练误差,即损失参数作为检测损失函数当前是否达到收敛的条件,比如训练误差是否小于预设误差或误差变化趋势是否趋于稳定,或者当前的迭代次数是否等于预设次数。若检测达到收敛条件,比如损失函数的训练误差达到小于预设误差或误差变化趋于稳定,表明待训练属性编辑子模型训练完成,此时可以停止迭代训练。若检测到当前未达到收敛条件,可以获取训练样本数据对待训练属性编辑子模型继续进行训练,直至损失函数的训练误差在预设范围之内。当损失函数的训练误差达到收敛时,可以认为待训练属性编辑子模型 训练好了,以使在将第一特征向量输入至训练好的待训练属性编辑子模型,模型可以精准为第一特征向量拼接属性特征向量,以使可以生成带有面部特征的图像。
S480、将所述待使用面部属性确定模型中的所述目标属性分类模型、所述面部匹配模型、所述第一图像生成子模型剔除,得到所述目标面部属性确定模型。
在本实施例中,对于上述确定的待使用面部属性确定模型来说,此时该模型不仅包括特征预处理子模型以及第二图像生成子模块,还包括第一图像生成子模块、目标属性分类模型以及面部匹配模型。为了实现数据处理的快捷性和对终端设备算力要求较低的情形下,生成有特效的面部图像的效果,可以将待使用面部属性确定模型中的第一图像生成子模块、目标属性分类模型以及面部匹配模型剔除处理。即,将应用中不会使用到的模型剔除掉。如,在训练好待使用面部属性确定模型之后,只保留属性编辑子模型分支之路,通过输入属性值生成编辑后的面部图像。
S490、获取待处理数据;其中,所述待处理数据包括高斯噪声或待转换图像。
S4100、基于目标面部属性确定模型对所述待处理数据进行处理,得到与所述待处理数据相对应的目标面部图像。
本公开实施例的技术方案,通过获取高斯噪声或待转换图像等待处理数据,进而基于目标面部属性确定模型对待处理数据进行处理,得到与待处理数据相对应的添加特效之后的目标面部图像,同时通过利用训练样本对所述待训练面部属性确定模型进行训练,不断优化待训练属性编辑子模型中的模型参数,进而得到待使用面部属性确定模型,将待使用面部属性确定模型进行剪裁处理后得到目标面部属性确定模型,以使提高目标面部属性确定模型的精度以及添加特效的效率,进而以使为图像添加的特效真实度较高。
实施例五
图8是本公开实施例五所提供的一种图像处理装置的结构框图,可执行本公开任意实施例所提供的图像处理方法,具备执行方法相应的功能模块和有益效果。如图8所示,该装置包括:数据获取模块510以及目标图像确定模块520。
数据获取模块510,设置为获取待处理数据;其中,所述待处理数据包括高斯噪声或待转换图像;目标图像确定模块520,设置为基于目标面部属性确定模型对所述待处理数据进行处理,得到与所述待处理数据相对应的目标面部图像; 其中,所述目标面部图像中的至少一个目标特征与相应的至少一种预设面部特征相匹配。
在上述技术方案的基础上,所述目标图像确定模块520,包括待拼接特征向量确定单元、目标特征向量获取单元以及目标面部图像获取单元。
待拼接特征向量确定单元,设置为确定与所述待处理数据相对应的待拼接特征向量;目标特征向量获取单元,设置为为所述待拼接特征向量拼接与所述至少一个预设面部特征相对应的预设特征向量,得到与所述目标面部图像相对应的目标特征向量;目标面部图像获取单元,设置为对所述目标特征向量进行处理,得到所述目标面部图像。
在上述技术方案的基础上,所述特征预处理子模型包括第一特征提取模块和第二特征提取模块,待拼接特征向量确定单元,包括待拼接特征向量确定第一单元和待拼接特征向量确定第二单元。
待拼接特征向量确定第一单元,设置为若所述待处理数据为所述高斯噪声,则基于第一特征提取模块确定与所述高斯噪声相对应的待拼接特征向量;待拼接特征向量确定第二单元,设置为若所述待处理数据为所述待转换图像,则基于所述第二特征提取模块,确定与所述待转换图像相对应的待拼接特征向量。
在上述技术方案的基础上,所述装置还包括:目标面部属性确定模型参数修正模块。
目标面部属性确定模型参数修正模块,设置为基于预先训练得到的属性分类器,确定与所述目标面部图像相对应的至少一种目标属性,以基于所述至少一种目标属性对所述目标面部属性确定模型中的模型参数进行修正;其中,所述至少一种目标属性与所述至少一种预设面部特征的属性标识相匹配。
在上述技术方案的基础上,所述装置还包括:待训练面部属性确定模型构建模块。
待训练面部属性确定模型构建模块,设置为根据待训练属性编辑子模型,预先训练得到的目标对抗模型、目标属性分类模型以及面部匹配模型,构建所述待训练面部属性确定模型;其中,所述目标属性分类模型,设置为确定经所述目标对抗模型输出的图像的面部特征;所述面部匹配模型,设置为确定基于所述目标对抗模型输出的面部图像的匹配度,所述目标对抗模型,设置为输出两幅面部图像,其中一幅面部图像与所述待训练属性编辑子模型中设置的预设面部特征相匹配。
在上述技术方案的基础上,所述目标对抗模型中包括:特征预处理子模型和图像生成子模型;所述特征预处理子模型中包括第一特征提取模块;所述图 像生成子模型包括第一图像生成子模块和第二图像生成子模块;所述特征预处理子模型中还包括第二特征提取模块。
在上述技术方案的基础上,所述待训练面部属性确定模型构建模块,还包括待训练面部属性确定模型构建单元。
待训练面部属性确定模型构建单元,设置为将所述第一特征提取模块或所述第二特征提取模块的输出结果,作为待训练属性编辑子模型和所述第一图像生成子模块的输入,将所述待训练属性编辑子模型的输出作为所述第二图像生成子模块的输入,将所述第一图像生成子模块的输出以及所述第二图像生成子模块的输出,作为所述目标属性分类模型以及所述面部匹配模型的输入,以构建所述待训练面部属性确定模型。
在上述技术方案的基础上,所述待使用面部属性确定模型获取单元,还包括训练样本获取子单元、第一特征向量获取子单元、第一属性特征向量获取子单元、属性特征图像获取子单元、信息获取子单元、目标损失值确定子单元和待使用面部属性确定模型获取子单元。
训练样本获取子单元,设置为获取多个训练样本,其中,训练样本中包括待训练数据;第一特征向量获取子单元,设置为针对多个训练样本,将当前训练样本中的待训练数据输入至所述第一特征提取模块或所述第二特征提取模块中得到与所述当前训练样本相对应的第一特征向量;第一属性特征向量获取子单元,设置为基于所述待训练属性编辑子模型为所述第一特征向量拼接与至少一种预设面部特征相对应的属性特征向量,得到第一属性特征向量;属性特征图像获取子单元,设置为将所述第一特征向量输入至所述第一图像生成子模块中,得到未附带属性特征图像;以及,将所述第一属性特征向量输入至所述第二图像生成子模块中,得到附带属性特征图像;信息获取子单元,设置为将所述附带属性特征图像和所述未附带属性特征图像输入至所述目标属性分类模型中,得到待比较属性信息;以及,将所述附带属性特征图像和所述未附带属性特征图像,输入至所述面部匹配模型中,得到面部匹配信息;目标损失值确定子单元,设置为基于所述待训练属性编辑子模型中的损失函数对所述待比较属性信息、面部匹配信息以及所述属性编辑子模型中的预设属性值进行处理,得到目标损失值;待使用面部属性确定模型获取子单元,设置为基于所述目标损失值对所述待训练属性编辑子模型中的模型参数进行修正,并将所述损失函数收敛作为训练目标,训练得到待使用面部属性确定模型。
在上述技术方案的基础上,所述目标面部属性确定模型获取单元,包括目标面部属性确定模型获取子单元。
目标面部属性确定模型获取子单元,设置为将所述待使用面部属性确定模 型中的所述目标属性分类模型、所述面部匹配模型、所述第二图像生成子模型剔除,得到所述目标面部属性确定模型。
在上述技术方案的基础上,所述预设面部特征包括佩戴至少一种饰品的面部特征、不同年龄阶段的面部特征、不同角度的面部特征、不同发型的面部特征、不同发型配色的面部特征以及不同表情的面部特征中的至少一种。
本公开实施例的技术方案,通过获取高斯噪声或待转换图像等待处理数据,进而基于目标面部属性确定模型对待处理数据进行处理,得到与待处理数据相对应的添加特效之后的目标面部图像,解决了相关技术中利用修图技术生成特效图像时,得到的特效图真实度较低,从而引起用户体验不佳的问题,实现了为待处理图像中的目标对象的面部添加相应的面部特征,以使得到的特效真实度较高,以及增加了视频图像内容丰富性和趣味性的效果,提高了用户使用体验的技术效果。
本公开实施例所提供的图像处理装置可执行本公开任意实施例所提供的图像处理方法,具备执行方法相应的功能模块和效果。
上述装置所包括的多个单元和模块只是按照功能逻辑进行划分的,但并不局限于上述的划分,只要能够实现相应的功能即可;另外,多个功能单元的名称也只是为了便于相互区分,并不用于限制本公开实施例的保护范围。
实施例六
图9是本公开实施例六所提供的一种电子设备的结构示意图。下面参考图9,其示出了适于用来实现本公开实施例的电子设备(例如图9中的终端设备或服务器)600的结构示意图。本公开实施例中的终端设备可以包括但不限于诸如移动电话、笔记本电脑、数字广播接收器、个人数字助理(Personal Digital Assistant,PDA)、平板电脑(Portable Android Device,PAD)、便携式多媒体播放器(Portable Media Player,PMP)、车载终端(例如车载导航终端)等等的移动终端以及诸如数字电视(Television,TV)、台式计算机等等的固定终端。图9示出的电子设备600仅仅是一个示例,不应对本公开实施例的功能和使用范围带来任何限制。
如图9所示,电子设备600可以包括处理装置(例如中央处理器、图形处理器等)601,其可以根据存储在只读存储器(Read-Only Memory,ROM)602中的程序或者从存储装置608加载到随机访问存储器(Random Access Memory,RAM)603中的程序而执行多种适当的动作和处理。在RAM 603中,还存储有电子设备600操作所需的多种程序和数据。处理装置601、ROM 602以及RAM  603通过总线604彼此相连。输入/输出(Input/Output,I/O)接口605也连接至总线604。
通常,以下装置可以连接至I/O接口605:包括例如触摸屏、触摸板、键盘、鼠标、摄像头、麦克风、加速度计、陀螺仪等的输入装置606;包括例如液晶显示器(Liquid Crystal Display,LCD)、扬声器、振动器等的输出装置607;包括例如磁带、硬盘等的存储装置608;以及通信装置609。通信装置609可以允许电子设备600与其他设备进行无线或有线通信以交换数据。虽然图9示出了具有多种装置的电子设备600,并不要求实施或具备所有示出的装置。可以替代地实施或具备更多或更少的装置。
根据本公开的实施例,上文参考流程图描述的过程可以被实现为计算机软件程序。例如,本公开的实施例包括一种计算机程序产品,其包括承载在非暂态计算机可读介质上的计算机程序,该计算机程序包含用于执行流程图所示的方法的程序代码。在这样的实施例中,该计算机程序可以通过通信装置609从网络上被下载和安装,或者从存储装置608被安装,或者从ROM 602被安装。在该计算机程序被处理装置601执行时,执行本公开实施例的方法中限定的上述功能。
本公开实施方式中的多个装置之间所交互的消息或者信息的名称仅用于说明性的目的,而并不是用于对这些消息或信息的范围进行限制。
本公开实施例提供的电子设备与上述实施例提供的图像处理方法属于同一构思,未在本实施例中详尽描述的技术细节可参见上述实施例,并且本实施例与上述实施例具有相同的效果。
实施例七
本公开实施例七提供了一种计算机存储介质,其上存储有计算机程序,该程序被处理器执行时实现上述实施例所提供的图像处理方法。
本公开上述的计算机可读介质可以是计算机可读信号介质或者计算机可读存储介质或者是上述两者的任意组合。计算机可读存储介质例如可以是——但不限于——电、磁、光、电磁、红外线、或半导体的系统、装置或器件,或者任意以上的组合。计算机可读存储介质的例子可以包括但不限于:具有一个或多个导线的电连接、便携式计算机磁盘、硬盘、RAM、ROM、可擦式可编程只读存储器(Erasable Programmable Read-Only Memory,EPROM或闪存)、光纤、便携式紧凑磁盘只读存储器(Compact Disc Read-Only Memory,CD-ROM)、光存储器件、磁存储器件、或者上述的任意合适的组合。在本公开中,计算机 可读存储介质可以是任何包含或存储程序的有形介质,该程序可以被指令执行系统、装置或者器件使用或者与其结合使用。而在本公开中,计算机可读信号介质可以包括在基带中或者作为载波一部分传播的数据信号,其中承载了计算机可读的程序代码。这种传播的数据信号可以采用多种形式,包括但不限于电磁信号、光信号或上述的任意合适的组合。计算机可读信号介质还可以是计算机可读存储介质以外的任何计算机可读介质,该计算机可读信号介质可以发送、传播或者传输用于由指令执行系统、装置或者器件使用或者与其结合使用的程序。计算机可读介质上包含的程序代码可以用任何适当的介质传输,包括但不限于:电线、光缆、射频(Radio Frequency,RF)等等,或者上述的任意合适的组合。
在一些实施方式中,客户端、服务器可以利用诸如超文本传输协议(HyperText Transfer Protocol,HTTP)之类的任何当前已知或未来研发的网络协议进行通信,并且可以与任意形式或介质的数字数据通信(例如,通信网络)互连。通信网络的示例包括局域网(Local Area Network,LAN),广域网(Wide Area Network,WAN),网际网(例如,互联网)以及端对端网络(例如,ad hoc端对端网络),以及任何当前已知或未来研发的网络。
上述计算机可读介质可以是上述电子设备中所包含的;也可以是单独存在,而未装配入该电子设备中。
上述计算机可读介质承载有一个或者多个程序,当上述一个或者多个程序被该电子设备执行时,使得该电子设备:
获取待处理数据;其中,所述待处理数据包括高斯噪声或待转换图像;基于目标面部属性确定模型对所述待处理数据进行处理,得到与所述待处理数据相对应的目标面部图像;其中,所述目标面部图像中的至少一个目标特征与相应的至少一种预设面部特征相匹配。
可以以一种或多种程序设计语言或其组合来编写用于执行本公开的操作的计算机程序代码,上述程序设计语言包括但不限于面向对象的程序设计语言—诸如Java、Smalltalk、C++,还包括常规的过程式程序设计语言—诸如“C”语言或类似的程序设计语言。程序代码可以完全地在用户计算机上执行、部分地在用户计算机上执行、作为一个独立的软件包执行、部分在用户计算机上部分在远程计算机上执行、或者完全在远程计算机或服务器上执行。在涉及远程计算机的情形中,远程计算机可以通过任意种类的网络——包括LAN或WAN—连接到用户计算机,或者,可以连接到外部计算机(例如利用因特网服务提供商来通过因特网连接)。
附图中的流程图和框图,图示了按照本公开多种实施例的系统、方法和计 算机程序产品的可能实现的体系架构、功能和操作。在这点上,流程图或框图中的每个方框可以代表一个模块、程序段、或代码的一部分,该模块、程序段、或代码的一部分包含一个或多个用于实现规定的逻辑功能的可执行指令。也应当注意,在有些作为替换的实现中,方框中所标注的功能也可以以不同于附图中所标注的顺序发生。例如,两个接连地表示的方框实际上可以基本并行地执行,它们有时也可以按相反的顺序执行,这依所涉及的功能而定。也要注意的是,框图和/或流程图中的每个方框、以及框图和/或流程图中的方框的组合,可以用执行规定的功能或操作的专用的基于硬件的系统来实现,或者可以用专用硬件与计算机指令的组合来实现。
描述于本公开实施例中所涉及到的单元可以通过软件的方式实现,也可以通过硬件的方式来实现。其中,单元的名称在一种情况下并不构成对该单元本身的限定,例如,第一获取单元还可以被描述为“获取至少两个网际协议地址的单元”。
本文中以上描述的功能可以至少部分地由一个或多个硬件逻辑部件来执行。例如,非限制性地,可以使用的示范类型的硬件逻辑部件包括:现场可编程门阵列(Field Programmable Gate Array,FPGA)、专用集成电路(Application Specific Integrated Circuit,ASIC)、专用标准产品(Application Specific Standard Parts,ASSP)、片上系统(System on Chip,SOC)、复杂可编程逻辑设备(Complex Programming Logic Device,CPLD)等等。
在本公开的上下文中,机器可读介质可以是有形的介质,其可以包含或存储以供指令执行系统、装置或设备使用或与指令执行系统、装置或设备结合地使用的程序。机器可读介质可以是机器可读信号介质或机器可读储存介质。机器可读介质可以包括但不限于电子的、磁性的、光学的、电磁的、红外的、或半导体系统、装置或设备,或者上述内容的任何合适组合。机器可读存储介质的示例会包括基于一个或多个线的电气连接、便携式计算机盘、硬盘、RAM、ROM、EPROM或快闪存储器、光纤、CD-ROM、光学储存设备、磁储存设备、或上述内容的任何合适组合。
根据本公开的一个或多个实施例,【示例一】提供了一种图像处理方法,该方法包括:
获取待处理数据;其中,所述待处理数据包括高斯噪声或待转换图像;
基于目标面部属性确定模型对所述待处理数据进行处理,得到与所述待处理数据相对应的目标面部图像;其中,所述目标面部图像中的至少一个目标特 征与相应的至少一种预设面部特征相匹配。
根据本公开的一个或多个实施例,【示例二】提供了一种图像处理方法,还包括:
所述基于目标面部属性确定模型对所述待处理数据进行处理,得到与所述待处理数据相对应的目标面部图像,包括:
确定与所述待处理数据相对应的待拼接特征向量;
为所述待拼接特征向量拼接与所述至少一个预设面部特征相对应的预设特征向量,得到与所述目标面部图像相对应的目标特征向量;
对所述目标特征向量进行处理,得到所述目标面部图像。
根据本公开的一个或多个实施例,【示例三】提供了一种图像处理方法,还包括:
所述确定与所述待处理数据相对应的待拼接特征向量,包括:
若所述待处理数据为所述高斯噪声,则基于第一特征提取模块确定与所述高斯噪声相对应的待拼接特征向量;
若所述待处理数据为所述待转换图像,则基于第二特征提取模块,确定与所述待转换图像相对应的待拼接特征向量。
根据本公开的一个或多个实施例,【示例四】提供了一种图像处理方法,还包括:
在所述得到与所述待处理数据相对应的目标面部图像之后,还包括:
基于预先训练得到的属性分类器,确定与所述目标面部图像相对应的至少一种目标属性,以基于所述至少一种目标属性对所述目标面部属性确定模型中的模型参数进行修正;
其中,所述至少一种目标属性与所述至少一种预设面部特征的属性标识相匹配。
根据本公开的一个或多个实施例,【示例五】提供了一种图像处理方法,还包括:
构建待训练面部属性确定模型;
通过对所述待训练面部属性确定模型训练处理,得到待使用面部属性确定模型;
通过对所述待使用面部属性确定模型剪裁处理,得到所述目标面部属性确定模型。
根据本公开的一个或多个实施例,【示例六】提供了一种图像处理方法,还包括:
所述构建待训练面部属性确定模型,包括:
根据待训练属性编辑子模型,预先训练得到的目标对抗模型、目标属性分类模型以及面部匹配模型,构建所述待训练面部属性确定模型;
其中,所述目标属性分类模型,用于确定经所述目标对抗模型输出的图像的面部特征;所述面部匹配模型,用于确定基于所述目标对抗模型输出的面部图像的匹配度,所述目标对抗模型,用于输出两幅面部图像,其中一幅面部图像与所述待训练属性编辑子模型中设置的预设面部特征相匹配。
根据本公开的一个或多个实施例,【示例七】提供了一种图像处理方法,还包括:
所述目标对抗模型中包括:特征预处理子模型和图像生成子模型;所述特征预处理子模型中包括第一特征提取模块;所述图像生成子模型包括第一图像生成子模块和第二图像生成子模块;所述特征预处理子模型中还包括第二特征提取模块。
根据本公开的一个或多个实施例,【示例八】提供了一种图像处理方法,还包括:
所述构建待训练面部属性确定模型,包括:将所述第一特征提取模块或所述第二特征提取模块的输出结果,作为待训练属性编辑子模型和所述第一图像生成子模块的输入,将所述待训练属性编辑子模型的输出作为所述第二图像生成子模块的输入,将所述第一图像生成子模块的输出以及所述第二图像生成子模块的输出,作为所述目标属性分类模型以及所述面部匹配模型的输入,以构建所述待训练面部属性确定模型。
根据本公开的一个或多个实施例,【示例九】提供了一种图像处理方法,还包括:
所述通过对所述待训练面部属性确定模型训练处理,得到待使用面部属性确定模型,包括:
获取多个训练样本,其中,训练样本中包括待训练数据;
针对多个训练样本,将当前训练样本中的待训练数据输入至所述第一特征提取模块或所述第二特征提取模块中得到与所述当前训练样本相对应的第一特征向量;
基于所述待训练属性编辑子模型为所述第一特征向量拼接与至少一种预设 面部特征相对应的属性特征向量,得到第一属性特征向量;
将所述第一特征向量输入至所述第一图像生成子模块中,得到未附带属性特征图像;以及,将所述第一属性特征向量输入至所述第二图像生成子模块中,得到附带属性特征图像;
将所述附带属性特征图像和所述未附带属性特征图像输入至所述目标属性分类模型中,得到待比较属性信息;以及,将所述附带属性特征图像和所述未附带属性特征图像,输入至所述面部匹配模型中,得到面部匹配信息;
基于所述待训练属性编辑子模型中的损失函数对所述待比较属性信息、面部匹配信息以及所述待训练属性编辑子模型中的预设面部特征进行处理,得到目标损失值;
基于所述目标损失值对所述待训练属性编辑子模型中的模型参数进行修正,并将所述损失函数收敛作为训练目标,训练得到待使用面部属性确定模型。
根据本公开的一个或多个实施例,【示例十】提供了一种图像处理方法,还包括:
所述通过对所述待使用面部属性确定模型剪裁处理,得到所述目标面部属性确定模型,包括:
将所述待使用面部属性确定模型中的所述目标属性分类模型、所述面部匹配模型、所述第一图像生成子模型剔除,得到所述目标面部属性确定模型。
根据本公开的一个或多个实施例,【示例十一】提供了一种图像处理方法,还包括:
所述预设面部特征包括佩戴至少一种饰品的面部特征、不同年龄阶段的面部特征、不同角度的面部特征、不同发型的面部特征、不同发型配色的面部特征以及不同表情的面部特征中的至少一种。
根据本公开的一个或多个实施例,【示例十二】提供了一种图像处理装置,该装置,包括:
数据获取模块,设置为获取待处理数据;其中,所述待处理数据包括高斯噪声或待转换图像;
目标图像确定模块,设置为基于目标面部属性确定模型对所述待处理数据进行处理,得到与所述待处理数据相对应的目标面部图像;其中,所述目标面部图像中的至少一个目标特征与相应的至少一种预设面部特征相匹配。
此外,虽然采用特定次序描绘了多个操作,但是这不应当理解为要求这些操作以所示出的特定次序或以顺序次序执行来执行。在一定环境下,多任务和 并行处理可能是有利的。同样地,虽然在上面论述中包含了多个实现细节,但是这些不应当被解释为对本公开的范围的限制。在单独的实施例的上下文中描述的一些特征还可以组合地实现在单个实施例中。相反地,在单个实施例的上下文中描述的多种特征也可以单独地或以任何合适的子组合的方式实现在多个实施例中。

Claims (15)

  1. 一种图像处理方法,包括:
    获取待处理数据;其中,所述待处理数据包括高斯噪声或待转换图像;
    基于目标面部属性确定模型对所述待处理数据进行处理,得到与所述待处理数据相对应的目标面部图像;其中,所述目标面部图像中的至少一个目标特征与相应的至少一种预设面部特征相匹配。
  2. 根据权利要求1所述的方法,其中,所述基于目标面部属性确定模型对所述待处理数据进行处理,得到与所述待处理数据相对应的目标面部图像,包括:
    确定与所述待处理数据相对应的待拼接特征向量;
    为所述待拼接特征向量拼接与所述至少一个预设面部特征相对应的预设特征向量,得到与所述目标面部图像相对应的目标特征向量;
    对所述目标特征向量进行处理,得到所述目标面部图像。
  3. 根据权利要求2所述的方法,其中,所述确定与所述待处理数据相对应的待拼接特征向量,包括:
    在所述待处理数据为所述高斯噪声的情况下,基于第一特征提取模块确定与所述高斯噪声相对应的待拼接特征向量;
    在所述待处理数据为所述待转换图像的情况下,基于第二特征提取模块,确定与所述待转换图像相对应的待拼接特征向量。
  4. 根据权利要求1所述的方法,在所述得到与所述待处理数据相对应的目标面部图像之后,还包括:
    基于预先训练得到的属性分类器,确定与所述目标面部图像相对应的至少一种目标属性,以基于所述至少一种目标属性对所述目标面部属性确定模型中的模型参数进行修正;
    其中,所述至少一种目标属性与所述至少一种预设面部特征的属性标识相匹配。
  5. 根据权利要求1所述的方法,还包括:
    构建待训练面部属性确定模型;
    通过对所述待训练面部属性确定模型训练处理,得到待使用面部属性确定模型;
    通过对所述待使用面部属性确定模型剪裁处理,得到所述目标面部属性确 定模型。
  6. 根据权利要求5所述的方法,其中,所述构建待训练面部属性确定模型,包括:
    根据待训练属性编辑子模型,预先训练得到的目标对抗模型、目标属性分类模型以及面部匹配模型,构建所述待训练面部属性确定模型;
    其中,所述目标属性分类模型,用于确定经所述目标对抗模型输出的图像的面部特征;所述面部匹配模型,用于确定基于所述目标对抗模型输出的面部图像的匹配度,所述目标对抗模型,用于输出两幅面部图像,一幅面部图像与所述待训练属性编辑子模型中设置的预设面部特征相匹配。
  7. 根据权利要求6所述的方法,其中,所述目标对抗模型中包括:特征预处理子模型和图像生成子模型;所述特征预处理子模型中包括第一特征提取模块;所述图像生成子模型包括第一图像生成子模块和第二图像生成子模块;所述特征预处理子模型中还包括第二特征提取模块。
  8. 根据权利要求7所述的方法,其中,所述构建待训练面部属性确定模型,包括:
    将所述第一特征提取模块或所述第二特征提取模块的输出结果,作为所述待训练属性编辑子模型和所述第一图像生成子模块的输入,将所述待训练属性编辑子模型的输出作为所述第二图像生成子模块的输入,将所述第一图像生成子模块的输出以及所述第二图像生成子模块的输出,作为所述目标属性分类模型以及所述面部匹配模型的输入,以构建所述待训练面部属性确定模型。
  9. 根据权利要求7所述的方法,其中,所述通过对所述待训练面部属性确定模型训练处理,得到待使用面部属性确定模型,包括:
    获取多个训练样本,其中,训练样本中包括待训练数据;
    针对所述多个训练样本,将当前训练样本中的待训练数据输入至所述第一特征提取模块或所述第二特征提取模块中得到与所述当前训练样本相对应的第一特征向量;
    基于所述待训练属性编辑子模型为所述第一特征向量拼接与至少一种预设面部特征相对应的属性特征向量,得到第一属性特征向量;
    将所述第一特征向量输入至所述第一图像生成子模块中,得到未附带属性特征图像;以及,将所述第一属性特征向量输入至所述第二图像生成子模块中,得到附带属性特征图像;
    将所述附带属性特征图像和所述未附带属性特征图像输入至所述目标属性 分类模型中,得到待比较属性信息;以及,将所述附带属性特征图像和所述未附带属性特征图像,输入至所述面部匹配模型中,得到面部匹配信息;
    基于所述待训练属性编辑子模型中的损失函数对所述待比较属性信息、所述面部匹配信息以及所述待训练属性编辑子模型中的预设面部特征进行处理,得到目标损失值;
    基于所述目标损失值对所述待训练属性编辑子模型中的模型参数进行修正,并将所述损失函数收敛作为训练目标,训练得到待使用面部属性确定模型。
  10. 根据权利要求9所述的方法,其中,所述通过对所述待使用面部属性确定模型剪裁处理,得到所述目标面部属性确定模型,包括:
    将所述待使用面部属性确定模型中的所述目标属性分类模型、所述面部匹配模型、所述第一图像生成子模型剔除,得到所述目标面部属性确定模型。
  11. 根据权利要求1-10中任一所述的方法,其中,所述预设面部特征包括佩戴至少一种饰品的面部特征、不同年龄阶段的面部特征、不同角度的面部特征、不同发型的面部特征、不同发型配色的面部特征以及不同表情的面部特征中的至少一种。
  12. 一种图像处理装置,包括:
    数据获取模块,设置为获取待处理数据;其中,所述待处理数据包括高斯噪声或待转换图像;
    目标图像确定模块,设置为基于目标面部属性确定模型对所述待处理数据进行处理,得到与所述待处理数据相对应的目标面部图像;其中,所述目标面部图像中的至少一个目标特征与相应的至少一种预设面部特征相匹配。
  13. 一种电子设备,包括:
    至少一个处理器;
    存储装置,设置为存储至少一个程序;
    当所述至少一个程序被所述至少一个处理器执行,使得所述至少一个处理器实现如权利要求1-11中任一所述的图像处理方法。
  14. 一种包含计算机可执行指令的存储介质,所述计算机可执行指令在由计算机处理器执行时用于执行如权利要求1-11中任一所述的图像处理方法。
  15. 一种计算机程序产品,包括承载在非暂态计算机可读介质上的计算机程序,所述计算机程序包含用于执行如权利要求1-11中任一所述的图像处理方法的程序代码。
PCT/CN2022/141795 2021-12-29 2022-12-26 图像处理方法、装置、电子设备和存储介质 Ceased WO2023125366A1 (zh)

Priority Applications (1)

Application Number Priority Date Filing Date Title
US18/725,684 US20250078566A1 (en) 2021-12-29 2022-12-26 Image processing method and apparatus, electronic device, and storage medium

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
CN202111641195.4A CN114387373B (zh) 2021-12-29 2021-12-29 图像处理方法、装置、电子设备和存储介质
CN202111641195.4 2021-12-29

Publications (1)

Publication Number Publication Date
WO2023125366A1 true WO2023125366A1 (zh) 2023-07-06

Family

ID=81199201

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/CN2022/141795 Ceased WO2023125366A1 (zh) 2021-12-29 2022-12-26 图像处理方法、装置、电子设备和存储介质

Country Status (3)

Country Link
US (1) US20250078566A1 (zh)
CN (1) CN114387373B (zh)
WO (1) WO2023125366A1 (zh)

Families Citing this family (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN114387373B (zh) * 2021-12-29 2026-04-21 北京字跳网络技术有限公司 图像处理方法、装置、电子设备和存储介质
CN114842261B (zh) * 2022-05-10 2025-07-25 西华师范大学 图像处理方法、装置、电子设备及存储介质

Citations (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US10311334B1 (en) * 2018-12-07 2019-06-04 Capital One Services, Llc Learning to process images depicting faces without leveraging sensitive attributes in deep learning models
CN109948796A (zh) * 2019-03-13 2019-06-28 腾讯科技(深圳)有限公司 自编码器学习方法、装置、计算机设备及存储介质
CN111292262A (zh) * 2020-01-19 2020-06-16 腾讯科技(深圳)有限公司 图像处理方法、装置、电子设备以及存储介质
CN111325726A (zh) * 2020-02-19 2020-06-23 腾讯医疗健康(深圳)有限公司 模型训练方法、图像处理方法、装置、设备及存储介质
CN114387373A (zh) * 2021-12-29 2022-04-22 北京字跳网络技术有限公司 图像处理方法、装置、电子设备和存储介质

Family Cites Families (6)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN103413270A (zh) * 2013-08-15 2013-11-27 北京小米科技有限责任公司 一种图像的处理方法、装置和终端设备
CN111414928A (zh) * 2019-01-07 2020-07-14 中国移动通信有限公司研究院 一种人脸图像数据生成方法、装置及设备
CN110909595B (zh) * 2019-10-12 2023-04-18 平安科技(深圳)有限公司 面部动作识别模型训练方法、面部动作识别方法
US11640684B2 (en) * 2020-07-21 2023-05-02 Adobe Inc. Attribute conditioned image generation
CN113422910A (zh) * 2021-05-17 2021-09-21 北京达佳互联信息技术有限公司 视频处理方法、装置、电子设备和存储介质
CN113569780B (zh) * 2021-08-03 2024-10-18 东南大学 一种基于梯度对抗攻击和生成对抗模型的人脸图片年龄转换方法

Patent Citations (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US10311334B1 (en) * 2018-12-07 2019-06-04 Capital One Services, Llc Learning to process images depicting faces without leveraging sensitive attributes in deep learning models
CN109948796A (zh) * 2019-03-13 2019-06-28 腾讯科技(深圳)有限公司 自编码器学习方法、装置、计算机设备及存储介质
CN111292262A (zh) * 2020-01-19 2020-06-16 腾讯科技(深圳)有限公司 图像处理方法、装置、电子设备以及存储介质
CN111325726A (zh) * 2020-02-19 2020-06-23 腾讯医疗健康(深圳)有限公司 模型训练方法、图像处理方法、装置、设备及存储介质
CN114387373A (zh) * 2021-12-29 2022-04-22 北京字跳网络技术有限公司 图像处理方法、装置、电子设备和存储介质

Also Published As

Publication number Publication date
US20250078566A1 (en) 2025-03-06
CN114387373A (zh) 2022-04-22
CN114387373B (zh) 2026-04-21

Similar Documents

Publication Publication Date Title
KR102416558B1 (ko) 영상 데이터 처리 방법, 장치 및 판독 가능 저장 매체
CN114419300B (zh) 风格化图像生成方法、装置、电子设备及存储介质
WO2023125374A1 (zh) 图像处理方法、装置、电子设备及存储介质
CN111476871B (zh) 用于生成视频的方法和装置
US20250077761A1 (en) Character generation method and apparatus, electronic device, and storage medium
CN112132847A (zh) 模型训练方法、图像分割方法、装置、电子设备和介质
CN110827379A (zh) 虚拟形象的生成方法、装置、终端及存储介质
CN113344776B (zh) 图像处理方法、模型训练方法、装置、电子设备及介质
WO2023093897A1 (zh) 图像处理方法、装置、电子设备及存储介质
US11792494B1 (en) Processing method and apparatus, electronic device and medium
WO2023098664A1 (zh) 特效视频的生成方法、装置、设备及存储介质
US20230421716A1 (en) Video processing method and apparatus, electronic device and storage medium
US12592260B2 (en) Video generation method and apparatus, electronic device, and storage medium
WO2023045710A1 (zh) 多媒体显示及匹配方法、装置、设备及介质
WO2023125366A1 (zh) 图像处理方法、装置、电子设备和存储介质
WO2023051244A1 (zh) 图像生成方法、装置、设备及存储介质
CN112785669A (zh) 一种虚拟形象合成方法、装置、设备及存储介质
CN113744286A (zh) 虚拟头发生成方法及装置、计算机可读介质和电子设备
CN115002442B (zh) 一种图像展示方法、装置、电子设备及存储介质
WO2022233223A1 (zh) 图像拼接方法、装置、设备及介质
US20250371877A1 (en) Video processing method, and electronic device
CN111340865B (zh) 用于生成图像的方法和装置
WO2025167333A1 (zh) 一种图像生成方法、装置、设备、介质、产品
WO2025002130A1 (zh) 图像编辑方法及相关设备
CN113298731B (zh) 图像色彩迁移方法及装置、计算机可读介质和电子设备

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 22914620

Country of ref document: EP

Kind code of ref document: A1

WWE Wipo information: entry into national phase

Ref document number: 18725684

Country of ref document: US

NENP Non-entry into the national phase

Ref country code: DE

32PN Ep: public notification in the ep bulletin as address of the adressee cannot be established

Free format text: NOTING OF LOSS OF RIGHTS PURSUANT TO RULE 112(1) EPC (EPO FORM 1205A DATED 16.10.2024)

WWP Wipo information: published in national office

Ref document number: 18725684

Country of ref document: US

122 Ep: pct application non-entry in european phase

Ref document number: 22914620

Country of ref document: EP

Kind code of ref document: A1