WO2025201255A1 - 图像处理方法、装置、设备及介质 - Google Patents

图像处理方法、装置、设备及介质

Info

Publication number
WO2025201255A1
WO2025201255A1 PCT/CN2025/084436 CN2025084436W WO2025201255A1 WO 2025201255 A1 WO2025201255 A1 WO 2025201255A1 CN 2025084436 W CN2025084436 W CN 2025084436W WO 2025201255 A1 WO2025201255 A1 WO 2025201255A1
Authority
WO
WIPO (PCT)
Prior art keywords
image
target
hairstyle
area
processed
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Pending
Application number
PCT/CN2025/084436
Other languages
English (en)
French (fr)
Inventor
朱渊略
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Beijing Zitiao Network Technology Co Ltd
Original Assignee
Beijing Zitiao Network Technology Co Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Beijing Zitiao Network Technology Co Ltd filed Critical Beijing Zitiao Network Technology Co Ltd
Publication of WO2025201255A1 publication Critical patent/WO2025201255A1/zh
Pending legal-status Critical Current
Anticipated expiration legal-status Critical

Links

Classifications

    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T5/00Image enhancement or restoration
    • G06T5/50Image enhancement or restoration using two or more images, e.g. averaging or subtraction
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T3/00Geometric image transformations in the plane of the image
    • G06T3/04Context-preserving transformations, e.g. by using an importance map
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T5/00Image enhancement or restoration
    • G06T5/77Retouching; Inpainting; Scratch removal
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T2207/00Indexing scheme for image analysis or image enhancement
    • G06T2207/20Special algorithmic details
    • G06T2207/20212Image combination
    • G06T2207/20221Image fusion; Image merging
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T2207/00Indexing scheme for image analysis or image enhancement
    • G06T2207/30Subject of image; Context of image processing
    • G06T2207/30196Human being; Person
    • G06T2207/30201Face

Definitions

  • the photo editing function has been widely used in various application scenarios such as image editing software, photo taking software, video live streaming platforms, etc. Users can adjust the image according to their needs, for example, change the hairstyle.
  • the present disclosure provides an image processing method, apparatus, device, and medium.
  • An embodiment of the present disclosure provides an image processing method, comprising: obtaining a first image to be processed and determining a target hairstyle required for a target object in the first image; obtaining object information of the target object and, based on the object information and the target hairstyle, determining an area to be processed corresponding to the target object; obtaining a target image based on the first image, the area to be processed, and the target hairstyle; wherein the target image is an image containing a target object having the target hairstyle.
  • determining the target hairstyle required for the target object in the first image includes: upon receiving hairstyle prompt information, determining the target hairstyle required for the target object in the first image based on the hairstyle prompt information; and/or, upon triggering a target option among a plurality of preset hairstyle options, determining the target hairstyle required for the target object in the first image based on the hairstyle corresponding to the target option.
  • determining the area to be processed corresponding to the target object based on the object information and the target hairstyle includes: obtaining a hair expansion mask image of the target object based on the object information and the target hairstyle, so as to identify the area to be processed corresponding to the target object through the hair expansion mask image; wherein the area to be processed is larger than the original hair area of the target object.
  • obtaining target prompt information based on the hairstyle prompt information includes: adjusting the hairstyle prompt information based on the object information to obtain adjusted hairstyle prompt information; wherein the object information includes information of the original hair area of the target object and information of the hair-associated area of the target object; obtaining preset general prompt information; and obtaining target prompt information based on the adjusted hairstyle prompt information and the general prompt information.
  • the mask image corresponding to the area to be processed, the first image and the second image are fused to obtain a target image, including: correcting the second image based on the first image to obtain a corrected second image; and/or obtaining designated key point information of the target object in the first image, and correcting the target object in the second image based on the designated key point information; and fusing the mask image corresponding to the area to be processed, the first image and the corrected second image to obtain a target image.
  • the target image includes a first area and a second area; wherein, the position of the first area in the target image is determined based on the position of the area to be processed in the first image, and the second area is the area in the target image other than the first area; the pixel value of the first area is determined based on the pixels corresponding to the area to be processed in the second image, and the pixel value of the second area is determined based on the pixels of the area in the first image other than the area to be processed.
  • An embodiment of the present disclosure also provides an electronic device, which includes: a processor; a memory for storing executable instructions of the processor; the processor is used to read the executable instructions from the memory and execute the instructions to implement the image processing method provided by the embodiment of the present disclosure.
  • FIG2 is a schematic flow chart of an image processing method provided by an embodiment of the present disclosure.
  • FIG3 is a schematic structural diagram of an image processing device provided by an embodiment of the present disclosure.
  • the photo editing function has been widely used.
  • the inventors have found through research that there are currently few software programs that have the function of changing hairstyles. Even if some software programs provide this function, the hairstyle changing effect is not good and the user experience is poor.
  • the first image is an image containing a target object.
  • the target object can be a person, or a doll or animal with hair.
  • the first image can include only the head of the target object, or the entire body or upper body of the target object, without limitation.
  • the target objects can be all objects contained in the first image, or user-specified objects, without limitation.
  • the target hairstyle can be determined based on the required hairstyle information input by the user, or multiple hairstyle options can be provided to the user, and the target hairstyle can be determined based on the hairstyle selected by the user from the multiple hairstyle options.
  • the target object in the first image can also be subjected to feature analysis, and a target hairstyle suitable for the target object can be recommended based on the feature analysis results. No restrictions are imposed here.
  • Object information refers to information related to the target object and can include characteristic information such as the target object's hair, as well as information about items related to the target object, such as a hat.
  • the object information includes information about the target object's original hair region and information about the target object's hair-related region.
  • the target object's hair-related region can be an area that partially overlaps with and/or is adjacent to the hair region, such as a facial region or a target article region associated with hair.
  • the target article region refers to the region of the target article.
  • the target article can be specified as needed, such as a hat or other article that affects hairstyle changes. If the target object does not have the target article, the target article region is not identified, or the object information indicates that the target article region does not exist.
  • the aforementioned region information can specifically be location information of the region.
  • semantic analysis can be performed on the first image, and the object information can be presented using the semantic analysis results (such as a semantic map).
  • the semantic map different regions are represented using different pixel values, while all pixels in the same region have the same pixel value.
  • the object information may further include the object category to which the target object belongs.
  • the object category can be flexibly set according to needs. For example, taking the target object as a person as an example, it can be classified based on the characteristics of different people, such as facial features, etc.
  • the classification method of the person can be flexibly set according to needs. Taking the target object as a doll as an example, it can be classified according to the type of doll (such as a humanoid doll, an animal doll with hair), etc., and there is no restriction here.
  • the disclosed embodiments fully consider the influence of the target object's object information on the hairstyle transformation. For example, when determining the area to be processed, the original hair area of the target object indicated by the object information, the hair-related area indicated by the object information, the object type to which the target object belongs indicated by the object information, and other object information are fully considered, and then combined with the relevant information of the target hairstyle (such as the description information of the target hairstyle), so as to reasonably and reliably determine the area to be processed, thereby ensuring the rationality of the subsequent generated image and avoiding the undesirable phenomena in the related technology, such as hair generated on the hat, generated hair covering the face, and partial hair missing.
  • the relevant information of the target hairstyle such as the description information of the target hairstyle
  • Step S106 obtaining a target image based on the first image, the area to be processed, and the target hairstyle; wherein the target image is an image containing a target object with the target hairstyle.
  • the above-mentioned image processing method provided by the embodiment of the present disclosure does not directly perform a simple hairstyle transformation based on the target hairstyle, but fully considers the influence of object information on the hairstyle transformation effect, and reasonably and reliably determines the target object's to-be-processed area in combination with the object information of the target object in the first image and the target hairstyle required by the target object, thereby obtaining an image containing the target object with the target hairstyle based on the first image, and the above-mentioned method can effectively ensure the rationality of the hairstyle transformation based on the first image, thereby ensuring the hairstyle replacement effect, and can better improve the user experience.
  • the step of determining the target hairstyle required for the target object in the first image in step S102 may be performed with reference to (1) and/or (2) below:
  • the hairstyle prompt information may be descriptive information of the desired target hairstyle input by the user, and may include one or more descriptive words, such as straight hair, black hair, long hair, smooth hair, etc., which may be set by the user as needed. Based on the hairstyle prompt information, the target hairstyle for the target object in the first image may be determined.
  • a target hairstyle required for the target object in the first image is determined based on the hairstyle corresponding to the target option.
  • multiple optional hairstyle options may be provided to the user in advance, such as hairstyle A, hairstyle B, hairstyle C, etc., and the hairstyle options may be annotated with relevant hairstyle description information or example images, so that the user can select the desired target hairstyle according to their needs.
  • the step of determining the area to be processed corresponding to the target object based on the object information and the target hairstyle in step S104 may specifically include: obtaining a hair dilation mask image of the target object based on the object information and the target hairstyle, thereby identifying the area to be processed corresponding to the target object using the hair dilation mask image, wherein the area to be processed is larger than the original hair area of the target object.
  • embodiments of the present disclosure may perform dilation processing on the original hair area based on the object information (such as information about the original hair area of the target object, information about the hair-related area of the target object, and the object category to which the target object belongs) and the target hairstyle, thereby determining an area to be processed that is larger than the original hair area.
  • the hair dilation mask image can be used to clearly identify the area to be processed.
  • the pixel values of the to-be-processed area in the hair expansion mask image are all first values
  • the pixel values of the to-be-processed area in the hair expansion mask image are all second values
  • the first value is different from the second value.
  • the first value and the second value can be selected from values such as 0 and 1. This not only can clearly identify the to-be-processed area, but also makes it easier for the hairstyle generation model to perform image generation processing based on the hair expansion mask image, and also makes it easier to perform subsequent fusion processing on the first image and the model-generated image based on the hair expansion mask image.
  • the area to be processed can be reasonably determined by integrating information such as object information and relevant information of the target hairstyle.
  • the area to be processed can clearly indicate the area that the hairstyle generation model needs to focus on, such as generating hair corresponding to the target hairstyle in this area, to ensure the generation effect of the hairstyle generation model;
  • the first image and the image (second image) output by the hairstyle generation model can be fused to further ensure the presentation effect of the final target image, such as making the difference between the final target image and the original first image mainly reflected in the different hairstyle, and the other contents of the target image except the hair-related area (that is, the area in the target image corresponding to the area to be processed) are still consistent with the first image, thereby presenting the user with a realistic image effect of only changing the hairstyle.
  • a second image is generated using a hairstyle generation model corresponding to the target hairstyle based on the first image, the mask corresponding to the area to be processed, and the target prompt information.
  • the mask corresponding to the area to be processed is the aforementioned hair expansion mask.
  • the hairstyle generation model corresponding to the target hairstyle can be preloaded, and the first image, the mask corresponding to the area to be processed, and the target prompt information are used as input information for the hairstyle generation model.
  • the hairstyle generation model then generates an image based on this input information to obtain the second image.
  • the target image includes a first region and a second region; wherein the position of the first region in the target image is determined based on the position of the region to be processed in the first image, which can be simply understood as the first region being the region in the target image corresponding to the region to be processed.
  • the pixel values of the first region are determined based on the pixels corresponding to the region to be processed in the second image, and the pixels corresponding to the region to be processed in the second image are also the pixels in the region corresponding to the region to be processed in the second image.
  • the pixel values of the first region can be consistent with the pixel values corresponding to the region to be processed in the second image. If correction or other processing is required for the second image, the pixel values of the first region can be consistent with the pixel values corresponding to the region to be processed in the corrected second image.
  • the second area is an area in the target image other than the first area, and the pixel value of the second area is determined based on the pixels of the area in the first image other than the area to be processed.
  • the pixel value of the second area can be consistent with the pixel value of the area in the first image other than the area to be processed.
  • Step 1 Adjust the hairstyle prompt information based on the object information to obtain the adjusted hairstyle prompt information.
  • the disclosed embodiment fully considers that the acquired hairstyle prompt information may not be accurate or complete. In order to ensure the reliability of the ultimately obtained prompt information, the hairstyle prompt information can be adjusted.
  • supplementary prompt words can be determined based on the object information and added to the hairstyle prompt information.
  • the object category of the target object indicated by the object information can be directly used as the supplementary prompt word, or descriptive information related to the target object, such as the hair-related area of the target object, can be used as the supplementary prompt word. This allows the hairstyle generation model to more accurately and reasonably generate images based on the expanded hairstyle prompt information.
  • the adjusted hairstyle prompt information may be directly used as the target prompt information.
  • preset general prompt information can be obtained, and then target prompt information can be obtained based on the adjusted hairstyle prompt information and the general prompt information.
  • general prompt information i.e., general prompt words
  • the general prompt information can be pre-set. Regardless of the target hairstyle or the target object, the general prompt information can be combined with the hairstyle prompt information as the target prompt information.
  • the general prompt information is applicable to any hairstyle or object category.
  • the general prompt information includes positive prompt information and negative prompt information.
  • the positive prompt information includes feature descriptors that the image is expected to have, and the negative prompt information includes feature descriptors that the image is not expected to have.
  • the positive prompt information can include prompt words such as "high quality”, “clear”, and “real”
  • the negative prompt information can include "low quality”, “fuzzy”, and "incomplete”.
  • the target prompt words obtained in the above manner can further ensure the image generation effect of the hairstyle generation model.
  • the embodiment of the present disclosure provides an implementation example of the above step C, that is, performing fusion processing based on the mask image corresponding to the area to be processed, the first image, and the second image to obtain a target image, which can be performed with reference to the following steps C1 and C2:
  • Step C1 Correct the second image based on the first image to obtain a corrected second image. Since the second image is generated by the hairstyle generation model, it may deviate from the desired image. To ensure the quality of the final target image, the present embodiment can correct the second image.
  • step C1 can be performed with reference to 1) and/or 2) below:
  • the color of the second image is corrected. Considering that there may be a certain color deviation between the second image and the first image, the color of the second image can be corrected so that the color of the corrected second image matches the color of the first image. Exemplarily, this can be achieved using an image style transfer algorithm using RGB channels. For example, the color of the second image can be corrected based on the overall mean variance of the first image.
  • the designated key point information can specify the key points to be obtained as needed, such as obtaining key points of limbs such as shoulders and hands.
  • the designated key point information can specify the key points to be obtained as needed, such as obtaining key points of limbs such as shoulders and hands.
  • step C2 can be performed with reference to the following steps C2.1 and C2.2:
  • step C2.2 the first and third images are fused based on the mask corresponding to the unprocessed region to produce a target image.
  • the unprocessed region in the first image is replaced with the corresponding region in the third image. This ensures that the resulting target image presents only the hairstyle change effect to the user, without altering other content in the first image.
  • Step S202 obtaining a first image to be processed and obtaining hairstyle prompt information.
  • Step S208 Acquire a hair expansion mask image of the target object based on the object information and the target hairstyle.
  • step S210 target prompt information is obtained based on the object information and the hairstyle prompt information, and a second image is generated based on the first image, the dilated mask image, and the target prompt information using a hairstyle generation model corresponding to the target hairstyle.
  • Step S212 performing correction processing on the second image based on the first image to obtain a corrected second image
  • Step S214 based on the mask image corresponding to the area to be processed, the corrected second image, and the target prompt information, a third image is generated using a hairstyle generation model corresponding to the target hairstyle.
  • Step S216 based on the mask image corresponding to the area to be processed, the first image and the third image are fused to obtain a target image.
  • a target hairstyle determination module 302 for acquiring a first image to be processed and determining a target hairstyle required for a target object in the first image
  • the above-mentioned image processing device does not directly perform a simple hairstyle change based on the target hairstyle, but fully considers the influence of object information on the hairstyle change effect, and reasonably and reliably determines the target object's to-be-processed area in combination with the object information of the target object in the first image and the target hairstyle required by the target object, thereby obtaining an image containing the target object with the target hairstyle based on the first image, the to-be-processed area and the target hairstyle.
  • the above-mentioned method can effectively ensure the rationality of the hairstyle change based on the first image, thereby ensuring the hairstyle replacement effect, and can better improve the user experience.
  • the target hairstyle determination module 302 is specifically used to: determine the target hairstyle required for the target object in the first image based on the hairstyle prompt information when hairstyle prompt information is received; and/or determine the target hairstyle required for the target object in the first image based on the hairstyle corresponding to the target option when a target option among multiple preset hairstyle options is triggered.
  • the target image acquisition module 306 is specifically used to: obtain hairstyle prompt information corresponding to the target hairstyle, so as to obtain target prompt information based on the hairstyle prompt information; generate a second image based on the first image, the mask image corresponding to the area to be processed, and the target prompt information using the hairstyle generation model corresponding to the target hairstyle; and perform fusion processing based on the mask image corresponding to the area to be processed, the first image, and the second image to obtain a target image.
  • the target image acquisition module 306 is specifically used to: adjust the hairstyle prompt information based on the object information to obtain adjusted hairstyle prompt information, wherein the object information includes information about the original hair area of the target object and information about the hair-related area of the target object; obtain preset general prompt information; and obtain target prompt information based on the adjusted hairstyle prompt information and the general prompt information.
  • the object information includes the object category to which the target object belongs
  • the target image acquisition module 306 is specifically used to: modify the target prompt word when it is detected that the hairstyle prompt information contains a target prompt word; wherein the object category corresponding to the target prompt word is inconsistent with the object category to which the target object belongs; determine a supplementary prompt word based on the object information, and add the supplementary prompt word to the hairstyle prompt information.
  • the target image acquisition module 306 is specifically used to: correct the color of the second image based on the color information of the first image; and/or obtain the specified key point information of the target object in the first image, and correct the target object in the second image based on the specified key point information; perform fusion processing based on the mask image corresponding to the area to be processed, the first image and the corrected second image to obtain the target image.
  • the target image acquisition module 306 is specifically used to: generate a third image based on the mask image corresponding to the area to be processed, the corrected second image and the target prompt information, using the hairstyle generation model corresponding to the target hairstyle; and fuse the first image and the third image based on the mask image corresponding to the area to be processed to obtain the target image.
  • the target image includes a first area and a second area; wherein the position of the first area in the target image is determined based on the position of the area to be processed in the first image, and the second area is the area in the target image other than the first area; the pixel value of the first area is determined based on the pixel corresponding to the area to be processed in the second image, and the pixel value of the second area is determined based on the pixel of the area in the first image other than the area to be processed.
  • An embodiment of the present disclosure provides an electronic device, which includes: a storage device storing a computer program; and a processing device configured to execute the computer program in the storage device to implement the steps of any one of the methods in the present disclosure.
  • Terminal devices in embodiments of the present disclosure may include, but are not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers.
  • mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers.
  • PDAs personal digital assistants
  • PADs tablet computers
  • PMPs portable multimedia players
  • in-vehicle terminals e.g., in-vehicle navigation terminals
  • fixed terminals such as digital TVs and desktop computers.
  • the electronic device illustrated in FIG4 is merely
  • electronic device 400 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 401, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 402 or a program loaded from a storage device 408 into a random access memory (RAM) 403.
  • ROM read-only memory
  • RAM random access memory
  • Various programs and data required for the operation of electronic device 400 are also stored in RAM 403.
  • Processing device 401, ROM 402, and RAM 403 are connected to each other via a bus 404.
  • An input/output (I/O) interface 405 is also connected to bus 404.
  • an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes program code for executing the method shown in the flowchart.
  • the computer program can be downloaded and installed from the network through the communication device 409, or installed from the storage device 408, or installed from the ROM 402.
  • the processing device 401 When the computer program is executed by the processing device 401, the above-mentioned functions defined in the method of the embodiment of the present disclosure are performed.
  • the embodiments of the present disclosure may also be a computer program product, which includes computer program instructions, which, when executed by a processor, cause the processor to perform the image processing method provided by the embodiments of the present disclosure.
  • the computer program product may be written in any combination of one or more programming languages to write program codes for performing the operations of the embodiments of the present disclosure, the programming languages including object-oriented programming languages such as Java, C++, etc., and also conventional procedural programming languages such as "C" language or similar programming languages.
  • the program code may be executed entirely on the user computing device, partially on the user device, as a separate software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server.
  • the embodiment of the present disclosure may also be a computer-readable storage medium having computer program instructions stored thereon.
  • the processor is enabled to execute the image processing method provided by the embodiment of the present disclosure.
  • the computer-readable storage medium can adopt any combination of one or more readable media.
  • the readable medium can be a readable signal medium or a readable storage medium.
  • the readable storage medium can, for example, include but is not limited to a system, device or component of electricity, magnetism, light, electromagnetic, infrared, or semiconductor, or any combination thereof.
  • readable storage media include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof.
  • RAM random access memory
  • ROM read-only memory
  • EPROM or flash memory erasable programmable read-only memory
  • CD-ROM compact disk read-only memory
  • magnetic storage device or any suitable combination thereof.
  • a prompt message is sent to the user to clearly inform the user that the operation requested will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the electronic device, application, server, storage medium, or other software or hardware that performs the operations of the disclosed technical solution based on the prompt message.
  • the prompt information in response to receiving a user's active request, may be sent to the user in the form of a pop-up window, in which the prompt information may be presented in text form.
  • the pop-up window may also contain a selection control for the user to select "agree” or “disagree” to provide personal information to the electronic device.

Landscapes

  • Physics & Mathematics (AREA)
  • General Physics & Mathematics (AREA)
  • Engineering & Computer Science (AREA)
  • Theoretical Computer Science (AREA)
  • Image Processing (AREA)
  • Processing Or Creating Images (AREA)

Abstract

本公开实施例涉及一种图像处理方法、装置、设备及介质,其中该方法包括:获取待处理的第一图像,并确定所述第一图像中目标对象所需的目标发型;获取所述目标对象的对象信息,并基于所述对象信息和所述目标发型,确定所述目标对象对应的待处理区域;基于所述第一图像、所述待处理区域以及所述目标发型,得到目标图像;其中,所述目标图像是包含具有所述目标发型的目标对象的图像。

Description

图像处理方法、装置、设备及介质
相关申请的交叉引用
本申请要求申请号为202410345972.8,题为“图像处理方法、装置、设备及介质”、申请日为2024年3月25日的中国发明专利申请的优先权,通过引用方式将该申请整体并入本文。
技术领域
本公开涉及计算机技术领域,尤其涉及一种图像处理方法、装置、设备及介质。
背景技术
修图功能已广泛应用于诸如图像剪辑软件、拍照软件、视频直播平台等多种应用场合中,用户可以根据需求对图像进行调整,例如,更改发型。
发明内容
本公开提供了一种图像处理方法、装置、设备及介质。
本公开实施例提供了一种图像处理方法,所述方法包括:获取待处理的第一图像,并确定所述第一图像中目标对象所需的目标发型;获取所述目标对象的对象信息,并基于所述对象信息和所述目标发型,确定所述目标对象对应的待处理区域;基于所述第一图像、所述待处理区域以及所述目标发型,得到目标图像;其中,所述目标图像是包含具有所述目标发型的目标对象的图像。
可选的,所述确定所述第一图像中目标对象所需的目标发型,包括:在接收到发型提示信息的情况下,基于所述发型提示信息确定所述第一图像中目标对象所需的目标发型;和/或,在预设的多种发型选项中的目标选项被触发的情况下,基于所述目标选项对应的发型确定所述第一图像中目标对象所需的目标发型。
可选的,所述基于所述对象信息和所述目标发型,确定所述目标对象对应的待处理区域,包括:基于所述对象信息和所述目标发型,获取所述目标对象的头发膨胀掩膜图,以通过所述头发膨胀掩膜图标识所述目标对象对应的待处理区域;其中,所述待处理区域大于所述目标对象的原有头发区域。
可选的,所述基于所述第一图像、所述待处理区域以及所述目标发型,得到目标图像,包括:获取所述目标发型对应的发型提示信息,以基于所述发型提示信息得到目标提示信息;基于所述第一图像、所述待处理区域对应的掩膜图以及所述目标提示信息,利用所述目标发型对应的发型生成模型,生成第二图像;基于所述待处理区域对应的掩膜图、所述第一图像和所述第二图像进行融合处理,得到目标图像。
可选的,所述基于所述发型提示信息得到目标提示信息,包括:基于所述对象信息对所述发型提示信息进行调整,得到调整后的发型提示信息;其中,所述对象信息包括所述目标对象的原有头发区域的信息和所述目标对象的头发关联区域的信息;获取预设的通用提示信息;基于所述调整后的发型提示信息以及所述通用提示信息,得到目标提示信息。
可选的,所述对象信息包括所述目标对象所属的对象类别,所述基于所述对象信息对所述发型提示信息进行调整,得到调整后的发型提示信息,包括:在检测到所述发型提示信息包含有目标提示词的情况下,修改所述目标提示词;其中,所述目标提示词对应的对象类别与所述目标对象所属的对象类别不一致;基于所述对象信息确定补充提示词,并在所述发型提示信息中加入所述补充提示词。
可选的,所述基于所述待处理区域对应的掩膜图、所述第一图像和所述第二图像进行融合处理,得到目标图像,包括:基于所述第一图像对所述第二图像进行矫正处理,得到矫正后的第二图像;和/或,获取所述第一图像中的目标对象的指定关键点信息,基于所述指定关键点信息对所述第二图像中的目标对象进行矫正处理;基于所述待处理区域对应的掩膜图、所述第一图像和所述矫正后的第二图像进行融合处理,得到目标图像。
可选的,所述基于所述待处理区域对应的掩膜图、所述第一图像和所述矫正后的第二图像进行融合处理,得到目标图像,包括:基于所述待处理区域对应的掩膜图、所述矫正后的第二图像以及所述目标提示信息,利用所述目标发型对应的发型生成模型,生成第三图像;基于所述待处理区域对应的掩膜图,对所述第一图像和所述第三图像进行融合处理,得到目标图像。
可选的,所述目标图像包括第一区域和第二区域;其中,所述第一区域在所述目标图像中的位置是基于所述待处理区域在所述第一图像中的位置确定的,所述第二区域是所述目标图像中除所述第一区域之外的区域;所述第一区域的像素值是基于所述待处理区域在所述第二图像中对应的像素确定的,所述第二区域的像素值是基于所述第一图像中除所述待处理区域之外的区域的像素确定的。
本公开实施例还提供了一种图像处理装置,包括:目标发型确定模块,用于获取待处理的第一图像,并确定所述第一图像中目标对象所需的目标发型;待处理区域确定模块,用于获取所述目标对象的对象信息,并基于所述对象信息和所述目标发型,确定所述目标对象对应的待处理区域;目标图像获得模块,用于基于所述第一图像、所述待处理区域以及所述目标发型,得到目标图像;其中,所述目标图像是包含具有所述目标发型的目标对象的图像。
本公开实施例还提供了一种电子设备,所述电子设备包括:处理器;用于存储所述处理器可执行指令的存储器;所述处理器,用于从所述存储器中读取所述可执行指令,并执行所述指令以实现如本公开实施例提供的图像处理方法。
本公开实施例还提供了一种计算机可读存储介质,所述存储介质存储有计算机程序,所述计算机程序用于执行如本公开实施例提供的图像处理方法。
应当理解,本部分所描述的内容并非旨在标识本公开的实施例的关键或重要特征,也不用于限制本公开的范围。本公开的其它特征将通过以下的说明书而变得容易理解。
附图说明
此处的附图被并入说明书中并构成本说明书的一部分,示出了符合本公开的实施例,并与说明书一起用于解释本公开的原理。
为了更清楚地说明本公开实施例或现有技术中的技术方案,下面将对实施例或现有技术描述中所需要使用的附图作简单地介绍,显而易见地,对于本领域普通技术人员而言,在不付出创造性劳动性的前提下,还可以根据这些附图获得其他的附图。
图1为本公开实施例提供的一种图像处理方法的流程示意图;
图2为本公开实施例提供的一种图像处理方法的流程示意图;
图3为本公开实施例提供的一种图像处理装置的结构示意图;
图4为本公开实施例提供的一种电子设备的结构示意图。
具体实施方式
为了能够更清楚地理解本公开的上述目的、特征和优点,下面将对本公开的方案进行进一步描述。需要说明的是,在不冲突的情况下,本公开的实施例及实施例中的特征可以相互组合。
在下面的描述中阐述了很多具体细节以便于充分理解本公开,但本公开还可以采用其他不同于在此描述的方式来实施;显然,说明书中的实施例只是本公开的一部分实施例,而不是全部的实施例。
如上所述,修图功能已广泛应用。然而,发明人经研究发现,目前具有发型更换功能的软件较少,即便部分软件提供有发型更换功能,但呈现出的发型更换效果不佳,用户体验较差。
本公开实施例首先提供了一种图像处理方法,图1为本公开实施例提供的一种图像处理方法的流程示意图,该方法可以由图像处理装置执行,其中该装置可以采用软件和/或硬件实现,一般可集成在电子设备中。如图1所示,该方法主要包括如下步骤S102~步骤S106:
步骤S102,获取待处理的第一图像,并确定第一图像中目标对象所需的目标发型。
第一图像为包含有目标对象的图像,本公开实施例对目标对象的类型不进行限制,示例性地,目标对象可以为人物,或者其它包含头发的玩偶或动物等。实际应用中,第一图像可以仅包含目标对象的头部,也可以包含对象全身或上半身,在此不进行限制,另外,在第一图像中包含多个目标对象的情况下,目标对象可以是第一图像中所包含的所有对象,也可以是用户指定的对象,在此均不进行限制。
在确定第一图像中目标对象所需的目标发型时,可以基于用户输入的所需发型信息确定目标发型,也可以为用户提供多种发型选项,基于用户从多种发型选项中选择的发型确定目标发型,还可以对第一图像中的目标对象进行特征分析,基于特征分析结果推荐目标对象适合的目标发型,在此不进行限制。
步骤S104,获取目标对象的对象信息,并基于对象信息和目标发型,确定目标对象对应的待处理区域。
对象信息即为与目标对象相关的信息,可以包含目标对象自身的头发等特征信息,也可以包含目标对象的帽子等与目标对象相关的物品的信息。示例性地,对象信息包括目标对象的原有头发区域的信息和目标对象的头发关联区域的信息,其中,目标对象的头发关联区域可以是与头发区域存在部分重叠的区域和/或与头发区域相邻的区域,诸如可以为面部区域、目标佩戴物区域等与头发具有关联关系的区域。目标佩戴物区域是指目标佩戴物的区域,目标佩戴物可以根据需求进行指定,诸如目标佩戴物可以为帽子等影响发型变换的佩戴物,倘若目标对象不具有目标佩戴物,则不标识目标佩戴物区域,或者对象信息指示不存在目标佩戴物区域。上述区域的信息可以具体是区域的位置信息,在实际应用中,可以对第一图像进行语义分析,通过语义分析结果(诸如语义图)等方式呈现对象信息,在语义图中,不同区域采用不同的像素值呈现,相同区域中所有像素的像素值均相同。在此基础上,对象信息还可以进一步包括目标对象所属的对象类别,该对象类别可以根据需求灵活设置,诸如,以目标对象是人物为例,可以基于不同人物的特征等方式进行分类,该特征诸如可以为脸型特征等,具体可根据需求灵活设置人物的分类方式,以目标对象是玩偶为例,可以按照玩偶所属品种(诸如仿人玩偶、具有头发的动物玩偶)等方式进行分类,在此不进行限制。
本公开实施例充分考虑到目标对象的对象信息对于发型变换的影响,诸如,在确定待处理区域时,充分考虑对象信息所指示的目标对象的原有头发区域、对象信息所指示的头发关联区域、对象信息所指示的目标对象所属的对象类型等对象信息,再结合目标发型的相关信息(诸如目标发型的描述信息),从而合理可靠地确定待处理区域,进而保障后续生成图像的合理性,避免相关技术中出现的在帽子上也生成了头发、生成的头发遮盖了面部、部分头发缺失等不良现象。
步骤S106,基于第一图像、待处理区域以及目标发型,得到目标图像;其中,目标图像是包含具有目标发型的目标对象的图像。
在实际应用中,可以获取目标发型对应的发型生成模型,基于第一图像、待处理区域以及目标发型,利用该发型生成模型高效便捷地得到具有目标发型的目标对象的图像。本公开实施例对发型生成模型不进行限制,诸如,发型生成模型可以采用扩散模型实现,在实际应用中,不同发型可以对应不同的发型生成模型,具体可针对多种发型分别预训练得到相应的发型生成模型,诸如,在通用的扩散模型的基础上,利用LoRA技术(也可称为LoRA微调技术)对扩散模型进行调整,从而得到各种发型对应的发型生成模型。
本公开实施例提供的上述图像处理方法,不会直接基于目标发型进行简单的发型变换,而是充分考虑到对象信息对于发型变换效果的影响,会结合第一图像中目标对象的对象信息以及目标对象所需的目标发型合理可靠地确定目标对象的待处理区域,从而在第一图像、待处理区域以及目标发型的基础上,得到包含具有该目标发型的目标对象的图像,上述方式可以有效保障基于第一图像进行发型变换的合理性,进而保障发型更换效果,能够较好的提升用户体验。
在一些实施方式中,步骤S102中确定第一图像中目标对象所需的目标发型的步骤,可以参照如下(1)和/或(2)执行:
(1)在接收到发型提示信息的情况下,基于发型提示信息确定第一图像中目标对象所需的目标发型。发型提示信息可以是用户输入的所需目标发型的描述信息,可包含一个或多个描述词,诸如直发、黑色、长发、柔顺等,具体可由用户根据需求进行设置。基于发型提示信息,可以确定第一图像中目标对象所需的目标发型。
(2)在预设的多种发型选项中的目标选项被触发的情况下,基于目标选项对应的发型确定第一图像中目标对象所需的目标发型。在实际应用中,也可以预先为用户提供多种可选的发型选项,诸如发型A、发型B、发型C等,并且可以为发型选项标注相关的发型描述信息或者示例图等,以便用户根据需求选取所需的目标发型。
在实际应用中,可以根据需求通过上述(1)或者上述(2)确定目标发型,还可以将上述(1)和上述(2)进行结合,诸如,用户可以先从已提供的发型选项中选择目标选项,然后根据需求输入发型提示信息,从而在目标选项对应的发型的基础上进行个性化调整,也即可基于用户输入的发型提示信息以及用户触发的目标选项对应的发型,综合确定第一图像中目标对象所需的目标发型。以上均为示例性说明,具体可以根据需求选择合适的方式灵活确定所需的目标发型,在此不进行限制。
为了能够合理的确定待处理区域,前述步骤S104中基于对象信息和目标发型,确定目标对象对应的待处理区域的步骤,可以具体包括:基于对象信息和目标发型,获取目标对象的头发膨胀掩膜图,以通过头发膨胀掩膜图标识目标对象对应的待处理区域,其中,待处理区域大于目标对象的原有头发区域。为了保障图像处理的合理性和可靠性,本公开实施例可以基于对象信息(诸如前述目标对象的原有头发区域的信息、目标对象的头发关联区域的信息和目标对象所属的对象类别等信息)和目标发型,在原有头发区域的基础上进行膨胀处理,以确定比原有头发区域更大的待处理区域,并采用头发膨胀掩膜图可以清楚地标识待处理区域。诸如,在头发膨胀掩膜图中的待处理区域的像素值均为第一数值,在头发膨胀掩膜图中除待处理区域的像素值均为第二数值,第一数值与第二数值不同,示例性地,第一数值和第二数值可以从0和1等数值中选取,不仅可以清楚标识待处理区域,而且也更便于发型生成模型基于头发膨胀掩膜图进行图像生成处理以及更便于后续基于头发膨胀掩膜图针对第一图像以及模型生成图像进行融合处理。
也即,通过上述方式,综合对象信息和目标发型的相关信息等信息可以合理地确定待处理区域,待处理区域可以清楚地指示发型生成模型所需重点关注的区域,诸如在该区域中生成目标发型对应的头发,保障发型生成模型的生成效果;此外,还可以基于待处理区域,将第一图像与发型生成模型输出的图像(第二图像)进行融合,进一步保障最终所得的目标图像的呈现效果,诸如,使最终所得的目标图像与原始的第一图像的区别主要体现在发型不同,目标图像中除头发相关区域(也即待处理区域对应在目标图像中的区域)之外的其它内容仍旧与第一图像一致,从而给用户呈现出仅更换发型的逼真图像效果。
为了可以高效可靠地得到发型变化效果较佳的目标图像,在一些实施方式中,上述步骤S106,也即基于第一图像、待处理区域以及目标发型,得到目标图像的步骤,可以参照如下步骤A~步骤C执行:
步骤A,获取目标发型对应的发型提示信息,以基于发型提示信息得到目标提示信息。在实际应用中,可以直接将发型提示信息作为目标提示信息,也可以对发型提示信息进行调整或者补充,从而得到更合理或更全面的目标提示信息。
步骤B,基于第一图像、待处理区域对应的掩膜图以及目标提示信息,利用目标发型对应的发型生成模型,生成第二图像。其中,待处理区域对应的掩膜图即为前述头发膨胀掩膜图。在实际应用中,可以预先加载目标发型对应的发型生成模型,将第一图像、待处理区域对应的掩膜图以及目标提示信息作为发型生成模型的输入信息,令发型生成模型基于输入信息进行图像生成,得到第二图像。
步骤C,基于待处理区域对应的掩膜图、第一图像和第二图像进行融合处理,得到目标图像。可以理解的是,第二图像是发型生成模型生成的新的图像,第二图像与第一图像中除头发区域之外的其余区域也可能存在一定差异,通过执行步骤C,可以较好地保障最终所得的目标图像的呈现效果。
示例性地,目标图像包括第一区域和第二区域;其中,第一区域在目标图像中的位置是基于待处理区域在第一图像中的位置确定的,可以简单理解为第一区域是待处理区域对应在目标图像中的区域。第一区域的像素值是基于待处理区域在第二图像中对应的像素确定的,待处理区域在第二图像中对应的像素也即待处理区域在第二图像中对应的区域中的像素。在具体实现时,第一区域的像素值可以与待处理区域在第二图像中对应的像素值一致,而倘若需要针对第二图像进行矫正等处理,则第一区域的像素值可以与待处理区域在矫正后的第二图像中对应的像素值一致。
第二区域是目标图像中除第一区域之外的区域,第二区域的像素值是基于第一图像中除待处理区域之外的区域的像素确定的,诸如,第二区域的像素值可以与第一图像中除待处理区域之外的区域的像素值一致。
综上,可以理解为上述融合处理方式是将待处理区域在第二图像(或者矫正后的第二图像)中的相应区域与第一图像中除待处理区域之外的区域进行合成,得到目标图像,也可简单理解为,将第一图像中的待处理区域替换为第二图像(或者矫正后的第二图像)中的相应区域,从而保障最终得到的目标图像给用户呈现的仅是发型变换效果,而不改变第一图像中的其它内容。
本公开实施例提供了上述步骤A中基于发型提示信息得到目标提示信息的一些实施示例,可以参照如下步骤一和步骤二执行:
步骤一,基于对象信息对发型提示信息进行调整,得到调整后的发型提示信息。本公开实施例充分考虑到获取的发型提示信息可能并不准确或者并不完整,为了保障最终所得的提示信息的可靠性,可以对发型提示信息进行调整。
考虑到可能会出现用户输入的发型提示信息与第一图像中的目标对象并不匹配的情况,诸如,发型提示信息所表征的对象类别与目标对象的对象类别不一致,则可以对发型提示信息进行修改;在一些具体的实施示例中,对象信息包括目标对象所属的对象类别,在此基础上,在检测到发型提示信息包含有目标提示词的情况下,修改目标提示词;其中,目标提示词对应的对象类别与目标对象所属的对象类别不一致。通过对目标提示词进行修改,以确保修改后的提示词与目标对象所属的对象类别一致。在实际应用中,一些描述性的词语与对象类别通常具有对应关系,诸如,“美丽”和“帅气”形容的对象类别通常不同。在实际应用中,可以针对不同的对象类别分别设置相应的提示词库,若发型提示信息中包含有与目标对象所属的对象类别不一致的目标提示词,则对该目标提示词进行修改,诸如采用目标对象所属的对象类别对应的提示词库中查找所需的提示词,利用查找到的提示词替换目标提示词;所需的提示词可以是与目标提示词的语义相近的词。
在一些具体的实施示例中,可以基于对象信息确定补充提示词,并在发型提示信息中加入补充提示词。诸如,可以直接将对象信息所指示的目标对象所属的对象类别作为补充提示词,也可以将目标对象的头发关联区域等目标对象相关的描述信息作为补充提示词,以使发型生成模型能够基于扩充后的发型提示信息更准确合理地进行图像生成。
步骤二,基于调整后的发型提示信息,得到目标提示信息。
在一些具体的实施示例中,可以直接将调整后的发型提示信息作为目标提示信息。
在另一些具体的实施示例中,可以获取预设的通用提示信息,然后基于调整后的发型提示信息以及通用提示信息,得到目标提示信息。在实际应用中,可以预先设置通用提示信息(也即通用提示词),无论目标发型是何种发型,无论目标对象是何种对象类别,都可将该通用提示信息结合发型提示信息均作为目标提示信息;换言之,通用提示信息适用于任何发型或者对象类别。在一些具体的实现方式中,通用提示信息包括正向提示信息和逆向提示信息,正向提示信息包含期望图像具有的特征描述词,逆向提示信息包含不期望图像具有的特征描述词;诸如,正向提示信息可以包含“高质”“清晰”“真实”等提示词,逆向提示信息可以包含“低质”“模糊”“残缺”等。通过上述方式所得的目标提示词可以进一步保障发型生成模型的图像生成效果。
进一步,本公开实施例提供了上述步骤C的实施示例,也即基于待处理区域对应的掩膜图、第一图像和第二图像进行融合处理,得到目标图像,可以参照如下步骤C1和步骤C2执行:
步骤C1,基于第一图像对第二图像进行矫正处理,得到矫正后的第二图像。由于第二图像是发型生成模型生成的,可能与期望所得的图像具有一定的偏差,为了保障最终所得的目标图像的效果,本公开实施例可以对第二图像进行矫正处理,示例性地,步骤C1可以参照如下1)和/或2)执行:
1)基于第一图像的颜色信息,对第二图像的颜色进行矫正处理。考虑到第二图像和第一图像之间可能存在一定的颜色偏差,可以对第二图像的颜色进行矫正,从而使矫正后的第二图像的颜色与第一图像的颜色匹配。示例性地,可以利用使用RGB通道的图像风格迁移算法实现,诸如,可根据第一图像的全图均值方差对第二图像的颜色进行矫正。
2)获取第一图像中的目标对象的指定关键点信息,基于指定关键点信息对第二图像中的目标对象进行矫正处理。上述指定关键点信息可以根据需求指定所需获取的关键点,诸如获取肩膀、手等肢体关键点。考虑到模型输出的第二图像中的目标对象可能与第一图像中的目标对象之间存在一定的部位偏差,通过采用指定关键点信息对第二图像中的目标对象进行矫正,可以有效保障矫正后的第二图像中的目标对象的人脸五官等部位与第一图像的一致性,避免出现人脸变形等不良问题。
在实际应用中,可以灵活选择上述1)、2)、或者1)和2)相结合,对第二图像进行矫正优化,应当说明的是,以上仅为示例性地矫正处理方式,在实际应用中还可以采用其它矫正处理方式,诸如,基于第一图像中的眼睛、眉毛等区域对第二图像的眼睛、眉毛进行矫正,可以简单理解为将第一图像中的眼睛、眉毛贴回至第二图像中,从而保障第二图像中的眼睛、眉毛与第一图像完全一致。
步骤C2,基于待处理区域对应的掩膜图、第一图像和矫正后的第二图像进行融合处理,得到目标图像。
在一些实施示例中,可以基于待处理区域对应的掩膜图,直接对第一图像和矫正后的第二图像进行融合处理,得到目标图像,具体融合处理方式可参照前述相关内容,在此不再赘述。
为了进一步提升目标图像的生成效果,在一些具体的实施示例中,步骤C2可以参照如下步骤C2.1和步骤C2.2执行:
步骤C2.1,基于待处理区域对应的掩膜图、矫正后的第二图像以及目标提示信息,利用目标发型对应的发型生成模型,生成第三图像。也即,可以再次利用目标发型对应的发型生成模型,在待处理区域对应的掩膜图、矫正后的第二图像以及目标提示信息的基础上,生成第三图像,第三图像的效果通常优于矫正后的第二图像,能够有效修补矫正后的第二图像中的瑕疵。在实际应用中,生成第二图像所需的发型生成模型与生成第三图像所需的发型生成模型可以是同一发型生成模型,但是采用的控制参数(诸如步长、控制强度等参数)不同,通过上述操作,可以进一步提升模型的图像生成效果。
步骤C2.2,基于待处理区域对应的掩膜图,对第一图像和第三图像进行融合处理,得到目标图像。诸如,将第一图像中的待处理区域替换为第三图像中的相应区域,从而保障最终得到的目标图像给用户呈现的仅是发型变换效果,而不改变第一图像中的其它内容。
在前述基础上,本公开实施例提供了如图2所示的一种图像处理方法的流程示意图,主要包括如下步骤S202~步骤S216:
步骤S202,获取待处理的第一图像,以及获取发型提示信息。
步骤S204,基于发型提示信息确定第一图像中目标对象所需的目标发型。
步骤S206,对第一图像进行语义解析,以基于语义解析结果得到目标对象的对象信息。对象信息可以包括目标对象的原有头发区域、目标对象的头发关联区域的信息、目标对象所属的对象类别等。
步骤S208,基于对象信息和目标发型,获取目标对象的头发膨胀掩膜图。
步骤S210,基于对象信息和发型提示信息得到目标提示信息,并基于第一图像、膨胀掩膜图以及目标提示信息,利用目标发型对应的发型生成模型,生成第二图像。
步骤S212,基于第一图像对第二图像进行矫正处理,得到矫正后的第二图像;
步骤S214,基于待处理区域对应的掩膜图、矫正后的第二图像以及目标提示信息,利用目标发型对应的发型生成模型,生成第三图像。
步骤S216,基于待处理区域对应的掩膜图,对第一图像和第三图像进行融合处理,得到目标图像。
以上步骤的具体实现方式可参照前述相关内容,在此不再赘述,通过上述方式,可以有效保障基于第一图像进行发型变换的合理性,充分保障最终所得的目标图像的发型更换效果,使目标图像呈现出仅更换发型的逼真效果,能够较好的提升用户体验。
对应于前述图像处理方法,图3为本公开实施例提供的一种图像处理装置的结构示意图,该装置可由软件和/或硬件实现,一般可集成在电子设备中,如图3所示,图像处理装置包括:
目标发型确定模块302,用于获取待处理的第一图像,并确定第一图像中目标对象所需的目标发型;
待处理区域确定模块304,用于获取目标对象的对象信息,并基于对象信息和目标发型,确定目标对象对应的待处理区域;
目标图像获得模块306,用于基于第一图像、待处理区域以及目标发型,得到目标图像;其中,目标图像是包含具有目标发型的目标对象的图像。
本公开实施例提供的上述图像处理装置,不会直接基于目标发型进行简单的发型变换,而是充分考虑到对象信息对于发型变换效果的影响,会结合第一图像中目标对象的对象信息以及目标对象所需的目标发型合理可靠地确定目标对象的待处理区域,从而在第一图像、待处理区域以及目标发型的基础上,得到包含具有该目标发型的目标对象的图像,上述方式可以有效保障基于第一图像进行发型变换的合理性,进而保障发型更换效果,能够较好的提升用户体验。
在一些实施方式中,所述目标发型确定模块302具体用于:在接收到发型提示信息的情况下,基于所述发型提示信息确定所述第一图像中目标对象所需的目标发型;和/或,在预设的多种发型选项中的目标选项被触发的情况下,基于所述目标选项对应的发型确定所述第一图像中目标对象所需的目标发型。
在一些实施方式中,所述待处理区域确定模块304具体用于:基于所述对象信息和所述目标发型,获取所述目标对象的头发膨胀掩膜图,以通过所述头发膨胀掩膜图标识所述目标对象对应的待处理区域;其中,所述待处理区域大于所述目标对象的原有头发区域。
在一些实施方式中,所述目标图像获得模块306具体用于:获取所述目标发型对应的发型提示信息,以基于所述发型提示信息得到目标提示信息;基于所述第一图像、所述待处理区域对应的掩膜图以及所述目标提示信息,利用所述目标发型对应的发型生成模型,生成第二图像;基于所述待处理区域对应的掩膜图、所述第一图像和所述第二图像进行融合处理,得到目标图像。
在一些实施方式中,所述目标图像获得模块306具体用于:基于所述对象信息对所述发型提示信息进行调整,得到调整后的发型提示信息,其中,所述对象信息包括所述目标对象的原有头发区域的信息和所述目标对象的头发关联区域的信息;获取预设的通用提示信息;基于所述调整后的发型提示信息以及所述通用提示信息,得到目标提示信息。
在一些实施方式中,所述对象信息包括所述目标对象所属的对象类别,所述目标图像获得模块306具体用于:在检测到所述发型提示信息包含有目标提示词的情况下,修改所述目标提示词;其中,所述目标提示词对应的对象类别与所述目标对象所属的对象类别不一致;基于所述对象信息确定补充提示词,并在所述发型提示信息中加入所述补充提示词。
在一些实施方式中,所述目标图像获得模块306具体用于:基于所述第一图像的颜色信息,对所述第二图像的颜色进行矫正处理;和/或,获取所述第一图像中的目标对象的指定关键点信息,基于所述指定关键点信息对所述第二图像中的目标对象进行矫正处理;基于所述待处理区域对应的掩膜图、所述第一图像和所述矫正后的第二图像进行融合处理,得到目标图像。
在一些实施方式中,所述目标图像获得模块306具体用于:基于所述待处理区域对应的掩膜图、所述矫正后的第二图像以及所述目标提示信息,利用所述目标发型对应的发型生成模型,生成第三图像;基于所述待处理区域对应的掩膜图,对所述第一图像和所述第三图像进行融合处理,得到目标图像。
在一些实施方式中,所述目标图像包括第一区域和第二区域;其中,所述第一区域在所述目标图像中的位置是基于所述待处理区域在所述第一图像中的位置确定的,所述第二区域是所述目标图像中除所述第一区域之外的区域;所述第一区域的像素值是基于所述待处理区域在所述第二图像中对应的像素确定的,所述第二区域的像素值是基于所述第一图像中除所述待处理区域之外的区域的像素确定的。
本公开实施例所提供的图像处理装置可执行本公开任意实施例所提供的图像处理方法,具备执行方法相应的功能模块和有益效果。
所属领域的技术人员可以清楚地了解到,为描述的方便和简洁,上述描述的装置实施例的具体工作过程,可以参考方法实施例中的对应过程,在此不再赘述。
本公开实施例提供了一种电子设备,电子设备包括:存储装置,其上存储有计算机程序;处理装置,用于执行所述存储装置中的所述计算机程序,以实现本公开中任一项方法的步骤。
下面参考图4,其示出了适于用来实现本公开实施例的电子设备400的结构示意图。本公开实施例中的终端设备可以包括但不限于诸如移动电话、笔记本电脑、数字广播接收器、PDA(个人数字助理)、PAD(平板电脑)、PMP(便携式多媒体播放器)、车载终端(例如车载导航终端)等等的移动终端以及诸如数字TV、台式计算机等等的固定终端。图4示出的电子设备仅仅是一个示例,不应对本公开实施例的功能和使用范围带来任何限制。
如图4所示,电子设备400可以包括处理装置(例如中央处理器、图形处理器等)401,其可以根据存储在只读存储器(ROM)402中的程序或者从存储装置408加载到随机访问存储器(RAM)403中的程序而执行各种适当的动作和处理。在RAM 403中,还存储有电子设备400操作所需的各种程序和数据。处理装置401、ROM 402以及RAM 403通过总线404彼此相连。输入/输出(I/O)接口405也连接至总线404。
通常,以下装置可以连接至I/O接口405:包括例如触摸屏、触摸板、键盘、鼠标、摄像头、麦克风、加速度计、陀螺仪等的输入装置406;包括例如液晶显示器(LCD)、扬声器、振动器等的输出装置407;包括例如磁带、硬盘等的存储装置408;以及通信装置409。通信装置409可以允许电子设备400与其他设备进行无线或有线通信以交换数据。虽然图4示出了具有各种装置的电子设备400,但是应理解的是,并不要求实施或具备所有示出的装置。可以替代地实施或具备更多或更少的装置。
特别地,根据本公开的实施例,上文参考流程图描述的过程可以被实现为计算机软件程序。例如,本公开的实施例包括一种计算机程序产品,其包括承载在非暂态计算机可读介质上的计算机程序,该计算机程序包含用于执行流程图所示的方法的程序代码。在这样的实施例中,该计算机程序可以通过通信装置409从网络上被下载和安装,或者从存储装置408被安装,或者从ROM 402被安装。在该计算机程序被处理装置401执行时,执行本公开实施例的方法中限定的上述功能。
除了上述方法和设备以外,本公开的实施例还可以是计算机程序产品,其包括计算机程序指令,所述计算机程序指令在被处理器运行时使得所述处理器执行本公开实施例所提供的图像处理方法。所述计算机程序产品可以以一种或多种程序设计语言的任意组合来编写用于执行本公开实施例操作的程序代码,所述程序设计语言包括面向对象的程序设计语言,诸如Java、C++等,还包括常规的过程式程序设计语言,诸如“C”语言或类似的程序设计语言。程序代码可以完全地在用户计算设备上执行、部分地在用户设备上执行、作为一个独立的软件包执行、部分在用户计算设备上部分在远程计算设备上执行、或者完全在远程计算设备或服务器上执行。
此外,本公开的实施例还可以是计算机可读存储介质,其上存储有计算机程序指令,所述计算机程序指令在被处理器运行时使得所述处理器执行本公开实施例所提供的图像处理方法。
所述计算机可读存储介质可以采用一个或多个可读介质的任意组合。可读介质可以是可读信号介质或者可读存储介质。可读存储介质例如可以包括但不限于电、磁、光、电磁、红外线、或半导体的系统、装置或器件,或者任意以上的组合。可读存储介质的更具体的例子(非穷举的列表)包括:具有一个或多个导线的电连接、便携式盘、硬盘、随机存取存储器(RAM)、只读存储器(ROM)、可擦式可编程只读存储器(EPROM或闪存)、光纤、便携式紧凑盘只读存储器(CD-ROM)、光存储器件、磁存储器件、或者上述的任意合适的组合。
本公开实施例还提供了一种计算机程序产品,包括计算机程序/指令,该计算机程序/指令被处理器执行时实现本公开实施例中的图像处理方法。
可以理解的是,在使用本公开各施例公开的技术方案之前,均应当依据相关法律法规通过恰当的方式对本公开所涉及个人信息的类型、使用范围、使用场景等告知用户并获得用户的授权。
例如,在响应于接收到用户的主动请求时,向用户发送提示信息,以明确地提示用户,其请求执行的操作将需要获取和使用到用户的个人信息。从而,使得用户可以根据提示信息来自主地选择是否向执行本公开技术方案的操作的电子设备、应用程序、服务器或存储介质等软件或硬件提供个人信息。
作为一种可选的但非限定性的实现方式,响应于接收到用户的主动请求,向用户发送提示信息的方式例如可以是弹窗的方式,弹窗中可以以文字的方式呈现提示信息。此外,弹窗中还可以承载供用户选择“同意”或者“不同意”向电子设备提供个人信息的选择控件。
可以理解的是,上述通知和获取用户授权过程仅是示意性的,不对本公开的实现方式构成限定,其他满足相关法律法规的方式也可应用于本公开的实现方式中。
需要说明的是,在本文中,诸如“第一”和“第二”等之类的关系术语仅仅用来将一个实体或者操作与另一个实体或操作区分开来,而不一定要求或者暗示这些实体或操作之间存在任何这种实际的关系或者顺序。而且,术语“包括”、“包含”或者其任何其他变体意在涵盖非排他性的包含,从而使得包括一系列要素的过程、方法、物品或者设备不仅包括那些要素,而且还包括没有明确列出的其他要素,或者是还包括为这种过程、方法、物品或者设备所固有的要素。在没有更多限制的情况下,由语句“包括一个……”限定的要素,并不排除在包括所述要素的过程、方法、物品或者设备中还存在另外的相同要素。
以上所述仅是本公开的具体实施方式,使本领域技术人员能够理解或实现本公开。对这些实施例的多种修改对本领域的技术人员来说将是显而易见的,本文中所定义的一般原理可以在不脱离本公开的精神或范围的情况下,在其它实施例中实现。因此,本公开将不会被限制于本文所述的这些实施例,而是要符合与本文所公开的原理和新颖特点相一致的最宽的范围。

Claims (12)

  1. 一种图像处理方法,其中该方法包括:
    获取待处理的第一图像,并确定所述第一图像中目标对象所需的目标发型;
    获取所述目标对象的对象信息,并基于所述对象信息和所述目标发型,确定所述目标对象对应的待处理区域;
    基于所述第一图像、所述待处理区域以及所述目标发型,得到目标图像;其中,所述目标图像是包含具有所述目标发型的目标对象的图像。
  2. 根据权利要求1所述的方法,其中所述确定所述第一图像中目标对象所需的目标发型,包括:
    在接收到发型提示信息的情况下,基于所述发型提示信息确定所述第一图像中目标对象所需的目标发型;和/或,
    在预设的多种发型选项中的目标选项被触发的情况下,基于所述目标选项对应的发型确定所述第一图像中目标对象所需的目标发型。
  3. 根据权利要求1所述的方法,其中所述基于所述对象信息和所述目标发型,确定所述目标对象对应的待处理区域,包括:
    基于所述对象信息和所述目标发型,获取所述目标对象的头发膨胀掩膜图,以通过所述头发膨胀掩膜图标识所述目标对象对应的待处理区域;其中,所述待处理区域大于所述目标对象的原有头发区域。
  4. 根据权利要求1所述的方法,其中所述基于所述第一图像、所述待处理区域以及所述目标发型,得到目标图像,包括:
    获取所述目标发型对应的发型提示信息,以基于所述发型提示信息得到目标提示信息;
    基于所述第一图像、所述待处理区域对应的掩膜图以及所述目标提示信息,利用所述目标发型对应的发型生成模型,生成第二图像;
    基于所述待处理区域对应的掩膜图、所述第一图像和所述第二图像进行融合处理,得到目标图像。
  5. 根据权利要求4所述的方法,其中所述基于所述发型提示信息得到目标提示信息,包括:
    基于所述对象信息对所述发型提示信息进行调整,得到调整后的发型提示信息,其中,所述对象信息包括所述目标对象的原有头发区域的信息和所述目标对象的头发关联区域的信息;
    获取预设的通用提示信息;
    基于所述调整后的发型提示信息以及所述通用提示信息,得到目标提示信息。
  6. 根据权利要求5所述的方法,其中所述对象信息包括所述目标对象所属的对象类别,所述基于所述对象信息对所述发型提示信息进行调整,得到调整后的发型提示信息,包括:
    在检测到所述发型提示信息包含有目标提示词的情况下,修改所述目标提示词;其中,所述目标提示词对应的对象类别与所述目标对象所属的对象类别不一致;
    基于所述对象信息确定补充提示词,并在所述发型提示信息中加入所述补充提示词。
  7. 根据权利要求4所述的方法,其中所述基于所述待处理区域对应的掩膜图、所述第一图像和所述第二图像进行融合处理,得到目标图像,包括:
    基于所述第一图像的颜色信息,对所述第二图像的颜色进行矫正处理;和/或,获取所述第一图像中的目标对象的指定关键点信息,基于所述指定关键点信息对所述第二图像中的目标对象进行矫正处理;
    基于所述待处理区域对应的掩膜图、所述第一图像和所述矫正后的第二图像进行融合处理,得到目标图像。
  8. 根据权利要求7所述的方法,其中所述基于所述待处理区域对应的掩膜图、所述第一图像和所述矫正后的第二图像进行融合处理,得到目标图像,包括:
    基于所述待处理区域对应的掩膜图、所述矫正后的第二图像以及所述目标提示信息,利用所述目标发型对应的发型生成模型,生成第三图像;
    基于所述待处理区域对应的掩膜图,对所述第一图像和所述第三图像进行融合处理,得到目标图像。
  9. 根据权利要求4所述的方法,其中所述目标图像包括第一区域和第二区域;其中,所述第一区域在所述目标图像中的位置是基于所述待处理区域在所述第一图像中的位置确定的,所述第二区域是所述目标图像中除所述第一区域之外的区域;
    所述第一区域的像素值是基于所述待处理区域在所述第二图像中对应的像素确定的,所述第二区域的像素值是基于所述第一图像中除所述待处理区域之外的区域的像素确定的。
  10. 一种图像处理装置,其中所述装置包括:
    目标发型确定模块,用于获取待处理的第一图像,并确定所述第一图像中目标对象所需的目标发型;
    待处理区域确定模块,用于获取所述目标对象的对象信息,并基于所述对象信息和所述目标发型,确定所述目标对象对应的待处理区域;
    目标图像获得模块,用于基于所述第一图像、所述待处理区域以及所述目标发型,得到目标图像;其中,所述目标图像是包含具有所述目标发型的目标对象的图像。
  11. 一种电子设备,其中所述电子设备包括:
    存储装置,其上存储有计算机程序;
    处理装置,用于执行所述存储装置中的所述计算机程序,以实现权利要求1-9中任一项所述的图像处理方法的步骤。
  12. 一种计算机可读存储介质,其中所述存储介质存储有计算机程序,所述计算机程序用于执行上述权利要求1-9中任一所述的图像处理方法。
PCT/CN2025/084436 2024-03-25 2025-03-24 图像处理方法、装置、设备及介质 Pending WO2025201255A1 (zh)

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
CN202410345972.8 2024-03-25
CN202410345972.8A CN120707372A (zh) 2024-03-25 2024-03-25 图像处理方法、装置、设备及介质

Publications (1)

Publication Number Publication Date
WO2025201255A1 true WO2025201255A1 (zh) 2025-10-02

Family

ID=97116737

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/CN2025/084436 Pending WO2025201255A1 (zh) 2024-03-25 2025-03-24 图像处理方法、装置、设备及介质

Country Status (2)

Country Link
CN (1) CN120707372A (zh)
WO (1) WO2025201255A1 (zh)

Citations (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN112819921A (zh) * 2020-11-30 2021-05-18 北京百度网讯科技有限公司 用于改变人物的发型的方法、装置、设备和存储介质
CN114881845A (zh) * 2022-05-17 2022-08-09 北京字跳网络技术有限公司 特效展示方法、装置、电子设备及存储介质
CN115063335A (zh) * 2022-07-18 2022-09-16 北京字跳网络技术有限公司 特效图的生成方法、装置、设备及存储介质
CN116229054A (zh) * 2022-12-13 2023-06-06 北京字跳网络技术有限公司 图像处理方法、装置及电子设备

Patent Citations (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN112819921A (zh) * 2020-11-30 2021-05-18 北京百度网讯科技有限公司 用于改变人物的发型的方法、装置、设备和存储介质
CN114881845A (zh) * 2022-05-17 2022-08-09 北京字跳网络技术有限公司 特效展示方法、装置、电子设备及存储介质
CN115063335A (zh) * 2022-07-18 2022-09-16 北京字跳网络技术有限公司 特效图的生成方法、装置、设备及存储介质
CN116229054A (zh) * 2022-12-13 2023-06-06 北京字跳网络技术有限公司 图像处理方法、装置及电子设备

Also Published As

Publication number Publication date
CN120707372A (zh) 2025-09-26

Similar Documents

Publication Publication Date Title
JP7596533B2 (ja) 動物顔スタイル画像の生成方法、モデルのトレーニング方法、装置及び機器
CN112991150A (zh) 风格图像生成方法、模型训练方法、装置和设备
CN112837213A (zh) 脸型调整图像生成方法、模型训练方法、装置和设备
WO2025185040A1 (zh) 媒体内容的生成方法、装置、电子设备、存储介质和程序产品
WO2024240188A1 (zh) 视频生成方法、装置、设备、存储介质和程序产品
CN111274476B (zh) 基于人脸识别的房源匹配方法、装置、设备和存储介质
CN109726279B (zh) 一种数据处理方法及装置
CN111597151A (zh) 文件生成方法、装置、计算机设备和存储介质
WO2026026646A1 (zh) 图像擦除方法、装置、电子设备以及存储介质
WO2025201365A1 (zh) 视频处理方法、装置、设备及介质
WO2025201255A1 (zh) 图像处理方法、装置、设备及介质
CN114138250A (zh) 系统用例的步骤生成方法、装置、设备及存储介质
CN119271336A (zh) 个性化界面生成方法、装置、计算机设备和存储介质
CN115942058A (zh) 视频进度条生成方法、设备、存储介质及装置
WO2025103492A1 (zh) 多媒体编辑方法、装置、设备及介质
WO2025026415A1 (zh) 图像处理方法、装置、设备及介质
WO2025031408A1 (zh) 图像处理方法、装置、计算机设备及存储介质
WO2025103326A1 (zh) 图像处理方法、装置、设备及介质
CN120318822A (zh) 图像处理方法、装置、设备及介质
US20250377766A1 (en) Method and apparatus for content creation
WO2025201561A1 (zh) 文字模板的生成方法、装置、设备及介质
KR20260042987A (ko) 전자 장치 및 전자 장치의 이미지 관리 방법
CN120856957A (zh) 视频生成方法、装置、电子设备以及存储介质
WO2025039982A1 (zh) 文字处理方法、装置、计算机设备及存储介质
WO2025082468A1 (zh) 特效生成方法、装置、电子设备及存储介质

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 25775634

Country of ref document: EP

Kind code of ref document: A1