WO2025200899A1 - 图像生成方法、装置、设备、存储介质及程序产品 - Google Patents
图像生成方法、装置、设备、存储介质及程序产品Info
- Publication number
- WO2025200899A1 WO2025200899A1 PCT/CN2025/078809 CN2025078809W WO2025200899A1 WO 2025200899 A1 WO2025200899 A1 WO 2025200899A1 CN 2025078809 W CN2025078809 W CN 2025078809W WO 2025200899 A1 WO2025200899 A1 WO 2025200899A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- image
- optimized
- model
- target
- reward
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Pending
Links
Classifications
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06T—IMAGE DATA PROCESSING OR GENERATION, IN GENERAL
- G06T11/00—Two-dimensional [2D] image generation
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/04—Architecture, e.g. interconnection topology
- G06N3/045—Combinations of networks
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/04—Architecture, e.g. interconnection topology
- G06N3/0475—Generative networks
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/08—Learning methods
- G06N3/094—Adversarial learning
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06T—IMAGE DATA PROCESSING OR GENERATION, IN GENERAL
- G06T5/00—Image enhancement or restoration
- G06T5/70—Denoising; Smoothing
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V10/00—Arrangements for image or video recognition or understanding
- G06V10/70—Arrangements for image or video recognition or understanding using pattern recognition or machine learning
- G06V10/764—Arrangements for image or video recognition or understanding using pattern recognition or machine learning using classification, e.g. of video objects
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V10/00—Arrangements for image or video recognition or understanding
- G06V10/70—Arrangements for image or video recognition or understanding using pattern recognition or machine learning
- G06V10/77—Processing image or video features in feature spaces; using data integration or data reduction, e.g. principal component analysis [PCA] or independent component analysis [ICA] or self-organising maps [SOM]; Blind source separation
- G06V10/80—Fusion, i.e. combining data from various sources at the sensor level, preprocessing level, feature extraction level or classification level
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V10/00—Arrangements for image or video recognition or understanding
- G06V10/70—Arrangements for image or video recognition or understanding using pattern recognition or machine learning
- G06V10/77—Processing image or video features in feature spaces; using data integration or data reduction, e.g. principal component analysis [PCA] or independent component analysis [ICA] or self-organising maps [SOM]; Blind source separation
- G06V10/80—Fusion, i.e. combining data from various sources at the sensor level, preprocessing level, feature extraction level or classification level
- G06V10/809—Fusion, i.e. combining data from various sources at the sensor level, preprocessing level, feature extraction level or classification level of classification results, e.g. where the classifiers operate on the same input data
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V10/00—Arrangements for image or video recognition or understanding
- G06V10/70—Arrangements for image or video recognition or understanding using pattern recognition or machine learning
- G06V10/82—Arrangements for image or video recognition or understanding using pattern recognition or machine learning using neural networks
-
- Y—GENERAL TAGGING OF NEW TECHNOLOGICAL DEVELOPMENTS; GENERAL TAGGING OF CROSS-SECTIONAL TECHNOLOGIES SPANNING OVER SEVERAL SECTIONS OF THE IPC; TECHNICAL SUBJECTS COVERED BY FORMER USPC CROSS-REFERENCE ART COLLECTIONS [XRACs] AND DIGESTS
- Y02—TECHNOLOGIES OR APPLICATIONS FOR MITIGATION OR ADAPTATION AGAINST CLIMATE CHANGE
- Y02T—CLIMATE CHANGE MITIGATION TECHNOLOGIES RELATED TO TRANSPORTATION
- Y02T10/00—Road transport of goods or passengers
- Y02T10/10—Internal combustion engine [ICE] based vehicles
- Y02T10/40—Engine management systems
Definitions
- the present disclosure relates to the field of image generation technology, and in particular to an image generation method, apparatus, device, storage medium, and program product.
- the present disclosure provides an image generation method, apparatus, device, storage medium, and program product to solve the problem of poor image generation effect.
- the present disclosure provides an image generation method, the method comprising:
- a target image of the target prompt information is generated.
- the target image generation model is obtained by adjusting the parameters of the preset image generation model based on the target reward model.
- the input of the target reward model includes first sample prompt information and the image to be optimized obtained by the preset image generation model based on the first sample prompt information.
- the target reward model is used to predict the image feature information to be optimized in the image to be optimized.
- Prompt information acquisition module used to obtain target prompt information
- a target image generation module is used to generate a target image of the target prompt information based on the target prompt information and a target image generation model.
- the target image generation model is obtained by adjusting the parameters of a preset image generation model based on a target reward model.
- the input of the target reward model includes first sample prompt information and an image to be optimized obtained by the preset image generation model based on the first sample prompt information.
- the target reward model is used to predict the image feature information to be optimized in the image to be optimized.
- the present disclosure provides a computer device comprising: a memory and a processor, the memory and the processor being communicatively connected to each other, the memory storing computer instructions, and the processor executing the image generation method of the first aspect or any corresponding embodiment thereof by executing the computer instructions.
- the present disclosure provides a computer-readable storage medium having computer instructions stored thereon, the computer instructions being used to enable a computer to execute the image generation method of the first aspect or any corresponding embodiment thereof.
- the image generation method uses first sample prompt information and an image to be optimized, obtained by a preset image generation model based on the first sample prompt information, as input to a target reward model.
- the target reward model is then used to predict image feature information to be optimized in the image to be optimized, thereby determining the dimension to be optimized by the preset image generation model when generating the image.
- the target reward model is then used to adjust the parameters of the preset image generation model so that the preset image generation model optimizes toward the dimension to be optimized.
- FIG1 is a flow chart of an image generation method according to an embodiment of the present disclosure.
- FIG2 is a flow chart of a method for determining a target image generation model according to an embodiment of the present disclosure
- FIG3 is a flow chart of another method for determining a target image generation model according to an embodiment of the present disclosure
- FIG6 is a structural block diagram of an image generating apparatus according to an embodiment of the present disclosure.
- FIG7 is a structural block diagram of a computer device according to an embodiment of the present disclosure.
- a prompt message is sent to the user to clearly inform the user that the operation requested will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the electronic device, application, server, storage medium, or other software or hardware that performs the operations of the disclosed technical solution based on the prompt message.
- the prompt information in response to receiving a user's active request, may be sent to the user in the form of a pop-up window, in which the prompt information may be presented in text form.
- the pop-up window may also contain a selection control for the user to select "agree” or “disagree” to provide personal information to the electronic device.
- the main method for generating diverse images given a given prompt text is through the use of the VEG diffusion model.
- feedback learning for the VEG diffusion model often involves pre-training a global reward model, which is then used to fine-tune the VEG diffusion model to optimize its image generation performance.
- the global reward model provides a global reward to the VEG diffusion model, typically fine-tuning the VEG diffusion model to optimize its diffusion along a specific dimension, such as the overall aesthetics of the image. This makes it difficult to achieve fine-grained dimensional optimization, making the VEG diffusion model difficult to adapt to different application scenarios and resulting in poor image generation performance.
- an image generation method embodiment is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
- Step S101 Obtain target prompt information.
- target prompt information may be a prompt text or a prompt image, and the type of the target prompt information is not limited here.
- the image generation method provided in this embodiment uses the first sample prompt information and the image to be optimized obtained by the preset image generation model based on the first sample prompt information as the input of the target reward model, and predicts the image feature information to be optimized in the image to be optimized based on the target reward model to determine the dimension to be optimized when the preset image generation model generates the image. Then, the target reward model is used to adjust the parameters of the preset image generation model so that the preset image generation model is optimized towards the dimension to be optimized. Therefore, it is possible to improve the image generation effect of the preset image generation model in the dimension to be optimized, so as to improve the image generation effect of the target image generation model in different dimensions, thereby being suitable for different application scenarios.
- the target reward model includes at least two target reward sub-models, wherein the target reward sub-model includes multiple local dimension nodes to be optimized with a hierarchical affiliation, the target reward sub-model corresponds to the image features of the global dimension to be optimized, the global dimension to be optimized includes at least one local dimension to be optimized, and the leaf nodes in the local dimension nodes to be optimized are used to represent the image features of the local dimension to be optimized.
- the target reward model of the present invention can be regarded as a reward model with a tree structure of multiple levels, and the input end of the target reward model can be regarded as a classifier for predicting the classification probability of the current input data corresponding to each target reward sub-model from the global dimension.
- the local dimension nodes to be optimized at the first level of the target reward sub-model correspond to the image features of the global dimension to be optimized. For example, assuming that at least two global dimensions to be optimized include the image-text consistency dimension and the aesthetic dimension.
- the target reward model includes a target reward sub-model for the image-text consistency dimension and a target reward sub-model for the aesthetic dimension.
- the local dimension nodes to be optimized at the first level of the target reward sub-model for the image-text consistency dimension correspond to the image features of the image-text consistency dimension
- the local dimension nodes to be optimized at the first level of the target reward sub-model for the aesthetic dimension correspond to the image features of the aesthetic dimension
- the local dimension nodes to be optimized at the second level of the target reward sub-model correspond to the image features of the local dimension to be optimized. For example, assuming that the local dimensions to be optimized under the image-text consistency dimension include the style consistency dimension and the content consistency dimension, and the local dimensions to be optimized under the aesthetic dimension include the color dimension, the texture dimension, the atmosphere dimension, and the layout dimension.
- the target reward sub-model of the image-text consistency dimension has two local dimension nodes to be optimized at the second level, which correspond to the image features of the style consistency dimension and the image features of the content consistency dimension respectively.
- the target reward sub-model of the aesthetic dimension has four local dimension nodes to be optimized at the second level, which correspond to the image features of the color dimension, the image features of the texture dimension, the image features of the atmosphere dimension and the image features of the layout dimension respectively.
- the local dimension nodes to be optimized at the third, fourth or other levels can be added under the global dimension nodes to be optimized at the second level, and the local dimension nodes to be optimized at the last level can be used as leaf nodes.
- the local dimension nodes to be optimized corresponding to the above-mentioned style consistency dimension and content consistency dimension can be regarded as leaf nodes of the target reward sub-model of the image-text consistency dimension
- the local dimension nodes to be optimized corresponding to the above-mentioned color dimension, texture dimension, atmosphere dimension and layout dimension can be regarded as leaf nodes of the target reward sub-model of the aesthetic dimension.
- the image generation method provided in this embodiment includes a target reward model that includes at least two target reward sub-models, each of which includes multiple local dimension nodes to be optimized with a hierarchical relationship.
- the target reward sub-model corresponds to the image features of the global dimension to be optimized
- the global dimension to be optimized includes at least one local dimension to be optimized
- the leaf nodes in each local dimension node to be optimized are used to represent the image features of the local dimension to be optimized. Therefore, the target reward model can be used to evaluate images generated by a preset image generation model from multiple fine-grained hierarchical dimensions to accurately determine the dimensions that the preset image generation model lacks when generating images.
- Step S201 Obtain first sample prompt information.
- Step S202 input the first sample prompt information into a preset image generation model to obtain an image to be optimized.
- step S203 the image to be optimized and the first sample prompt information are input into the target reward model to obtain image loss, where the image loss is used to characterize the image feature information to be optimized in the image to be optimized.
- the preset image generation model is constructed based on a diffusion model
- the step S202 of inputting the first sample prompt information into the preset image generation model to obtain the image to be optimized includes: inputting the first sample prompt information into the preset image generation model for noise addition and denoising; extracting the denoised image obtained within the preset denoising step range to obtain the image to be optimized.
- the preset denoising step range is the last 5 to 10 steps of the denoising process of the preset image generation model. It should be noted that the preset denoising step range should be selected so that the selected denoised image includes a relatively large amount of non-noise image information. In actual operation, the preset denoising step range can be adjusted based on implementation circumstances.
- the image generation method provided in this embodiment inputs first sample prompt information into a preset image generation model constructed based on a diffusion model for noise addition and denoising.
- the denoised image obtained within a preset denoising step range is then extracted as the image to be optimized for the target reward model. Because the denoised image obtained within the preset denoising step range includes a significant amount of information that can reflect the image generation effect of the preset image generation model, using the denoised image within the preset denoising step range as the image to be optimized for the target reward model enables the target reward model to accurately determine the dimensions missing from the image currently generated by the preset image generation model, further improving the fine-tuning accuracy of the preset image generation model.
- step a1 the image to be optimized and the first sample prompt information are input into the target reward model, and at least two target reward sub-models are used to predict the global dimension to be optimized to obtain a first prediction probability corresponding to the target reward sub-model.
- the target reward model and its various nodes can be regarded as a classifier.
- the target reward sub-model of the global dimension to be optimized can determine the image generation effect of the image to be optimized in each global dimension to be optimized based on the image to be optimized and the first sample prompt information, so as to judge in which global dimension to be optimized the image to be optimized is deficient, so as to obtain the first prediction probability corresponding to the target reward sub-model.
- Step a2 Based on the prediction results of the global dimension to be optimized, the second prediction probability and reward value corresponding to the leaf node are predicted respectively.
- the local dimension node to be optimized under the target reward sub-model will further determine the image generation effect of the predicted image to be optimized in each local dimension to be optimized, so as to judge which local dimension to be optimized the predicted image to be optimized has deficiencies in, so as to obtain the second prediction probability and reward value corresponding to the leaf node.
- the worse the image generation effect of the image to be optimized in a certain local dimension to be optimized the smaller the reward value of the corresponding leaf node.
- the reward value output by the local node to be optimized in the style consistency dimension is 0.6
- the reward value output by the local node to be optimized in the content consistency dimension is 2.0
- the reward value output by the leaf node in the color dimension is 1.4
- the reward value output by the leaf node in the texture dimension is 0.9
- the reward value output by the leaf node in the atmosphere dimension is -2.1
- the reward value output by the leaf node in the layout dimension is -2.7
- the bars corresponding to the leaf nodes in the atmosphere dimension and the layout dimension are relatively high, indicating that in the subsequent fine-tuning process of the image generation model, it is necessary to focus on the image generation effect of the model in the atmosphere dimension and the layout dimension.
- the second predicted probability can be used as the first weight of the corresponding reward value, and all reward values of the same target reward sub-model are weightedly calculated based on the first weight to obtain the first loss of each target reward sub-model.
- step a4 the first losses corresponding to all target reward sub-models are fused based on the first predicted probability to obtain the image loss.
- the first prediction probability is used as the second weight of the corresponding first loss, and all first losses are weightedly calculated based on the second weight to obtain the image loss.
- the image generation method provided in this embodiment utilizes target reward sub-models for at least two global dimensions to be optimized to determine a first predicted probability for the image to be optimized corresponding to each target reward sub-model. Based on the prediction results for the global dimensions to be optimized, predictions are then made for the local dimensions to be optimized to determine a second predicted probability and reward value for the image to be optimized corresponding to each local dimension to be optimized. Then, based on the first and second predicted probabilities and reward values, an image loss is determined. This allows a preset image generation model to determine the current optimization direction based on the image loss, thereby improving the image generation effect through fine-tuning.
- Step S301 Obtain second sample prompt information, an image pair corresponding to the second sample prompt information, and a dimension label to be optimized, where the image pair includes a positive sample image and a negative sample image corresponding to the dimension label to be optimized.
- Step S302 Determine the leaf node corresponding to the dimension label to be optimized in the preset reward model to determine the reward sub-model to which it belongs.
- Step S303 Using the corresponding reward sub-model, predict the dimensions to be optimized for the second sample prompt information and the image pair.
- the dimensions to be optimized include the global dimensions to be optimized and the local dimensions to be optimized.
- Step S304 Determine the second loss based on the predicted result and the dimension label to be optimized.
- Step S305 Update the parameters of the corresponding reward sub-model based on the second loss to determine the target reward model.
- the image generation method provided in this embodiment determines the leaf node corresponding to the label of the dimension to be optimized in the preset reward model to determine the corresponding reward sub-model. Then, the corresponding reward sub-model is used to predict the dimension to be optimized for the second sample prompt information and the image pair, so as to measure the difference between the output of the current reward model and the expected output based on the predicted result and the label of the dimension to be optimized, and obtain the second loss. In this way, the parameters of the corresponding reward sub-model can be updated based on the second loss, so that the updated reward sub-model can accurately judge the image generation effect of the input image in the corresponding dimension to be optimized, thereby improving the accuracy of the target reward model in judging the image generation effect of the input image in multiple dimensions.
- preference data of multiple fine-grained dimensions can be collected in advance to obtain training samples of multiple dimensions.
- Each training sample includes the second sample prompt information, the image pair corresponding to the second sample prompt information, and the dimension label to be optimized.
- the dimension label to be optimized in the current training sample is style consistency
- the reward sub-model to which it belongs is the reward sub-model of the image-text consistency dimension.
- the local dimension nodes to be optimized corresponding to the image-text consistency dimension and the style consistency dimension in the reward sub-model of the image-text consistency dimension can be used to predict the global dimension to be optimized and the local dimension to be optimized for the second sample prompt information and the image pair to obtain the predicted result.
- the reward sub-model can be regarded as a classifier.
- the reward sub-model will classify the input data to obtain the first classification probability of the input data corresponding to the global dimension to be optimized and the second classification probability corresponding to the leaf node.
- determining the second loss based on the predicted results and the dimension labels to be optimized in the above-mentioned step S304 includes: determining the reward value based on the second classification probability and the dimension labels to be optimized; and determining the second loss based on the reward value and the fusion result of the first classification probability.
- the corresponding reward sub-model in order for the corresponding reward sub-model to accurately determine the image generation effect of the input image in the dimension to be optimized, it is necessary to use the second classification probability and the label of the dimension to be optimized to determine the reward value with the goal of increasing the difference between the positive sample image and the negative sample image, so that the reward value can reflect the difference between the positive sample image and the negative sample image corresponding to the dimension to be optimized. Then, based on the fusion result of the reward value and the first classification probability, a second loss is determined, and the parameters of the preset reward model are updated using the second loss. This ensures that the reward value output by the leaf node of the target reward model after training can reflect the image generation effect of the input image in the corresponding dimension.
- the image generation method disclosed herein constructs a tree-structured reward model according to multiple fine-grained preference dimensions, and trains the reward model using training samples of different preference dimensions to obtain a target reward model that can be used to judge the image generation effect of an input image in different preference dimensions. Therefore, when the image generation model undergoes feedback learning, the target reward model can be used to score the image generation effect of the image generation model in multiple preference dimensions, so as to adaptively predict the preference dimensions that the image generation model lacks when generating images, and then optimize the image generation model towards the missing preference dimensions to improve the image generation effect of the image generation model in multiple preference dimensions.
- module may refer to a combination of software and/or hardware that implements a predetermined function.
- the devices described in the following embodiments are preferably implemented in software, implementation using hardware, or a combination of software and hardware, is also possible and contemplated.
- This embodiment provides an image generation device, as shown in FIG6 , including:
- Prompt information acquisition module 401 used to obtain target prompt information
- the target image generation module 402 is used to generate a target image of the target prompt information based on the target prompt information and the target image generation model.
- the target image generation model is obtained by adjusting the parameters of the preset image generation model based on the target reward model.
- the input of the target reward model includes the first sample prompt information and the image to be optimized obtained by the preset image generation model based on the first sample prompt information.
- the target reward model is used to predict the image feature information to be optimized in the image to be optimized.
- the image generation device further includes an image model determination module for determining a target image generation model.
- the image model determination module includes:
- a first information acquisition module used to acquire first sample prompt information
- a first image acquisition module configured to input the first sample prompt information into a preset image generation model to obtain an image to be optimized
- the image model updating module is used to update the parameters of the preset image generation model based on the image loss to obtain the target image generation model.
- the preset image generation model is constructed based on a diffusion model
- the first image acquisition module includes:
- An image processing unit configured to input the first sample prompt information into a preset image generation model for noise addition and denoising
- the image loss calculation module includes:
- a global dimension prediction unit configured to input the image to be optimized and the first sample prompt information into the target reward model, and use at least two target reward sub-models to predict the global dimension to be optimized, thereby obtaining a first prediction probability corresponding to the target reward sub-model;
- a local dimension prediction unit configured to predict the second prediction probability and reward value corresponding to each leaf node based on the prediction result of the global dimension to be optimized
- a first loss calculation unit is configured to obtain, for each target reward sub-model, a first loss based on a fusion of the second predicted probability and the reward value;
- the image generation device further includes a reward model determination module for determining a target reward model.
- the reward model determination module includes:
- a second information acquisition module is configured to acquire second sample prompt information, an image pair corresponding to the second sample prompt information, and a dimension label to be optimized, where the image pair includes a positive sample image and a negative sample image corresponding to the dimension label to be optimized;
- the optimization path determination module is used to determine the leaf node corresponding to the dimension label to be optimized in the preset reward model to determine the reward sub-model to which it belongs;
- An optimization dimension prediction module configured to use the corresponding reward sub-model to predict the dimensions to be optimized for the second sample prompt information and the image pair, where the dimensions to be optimized include the global dimensions to be optimized and the local dimensions to be optimized;
- a prediction loss calculation module is used to determine the second loss based on the prediction result and the dimension label to be optimized
- the reward model updating module is used to update the parameters of the corresponding reward sub-model based on the second loss to determine the target reward model.
- the optimization dimension prediction module includes:
- a first classification unit is configured to determine a first classification probability corresponding to a global dimension to be optimized in the corresponding reward sub-model based on the second sample prompt information and the image pair;
- the second classification unit is used to determine the second classification probability corresponding to the leaf node in the reward sub-model based on the second sample prompt information and the image pair.
- the predicted result includes the first classification probability and the second classification probability.
- the predicted loss calculation module includes:
- a reward value determining unit configured to determine a reward value based on the second classification probability and the dimension label to be optimized
- the second loss calculation unit is used to determine a second loss based on the reward value and the fusion result of the first classification probability.
- the image generating device in this embodiment is presented in the form of a functional unit, where the unit refers to an ASIC (Application Specific Integrated Circuit) circuit, a processor and memory that executes one or more software or fixed programs, and/or other devices that can provide the above functions.
- ASIC Application Specific Integrated Circuit
- the embodiment of the present disclosure further provides a computer device having the image generating device shown in FIG6 .
Landscapes
- Engineering & Computer Science (AREA)
- Theoretical Computer Science (AREA)
- Physics & Mathematics (AREA)
- General Physics & Mathematics (AREA)
- Evolutionary Computation (AREA)
- Health & Medical Sciences (AREA)
- Artificial Intelligence (AREA)
- Computing Systems (AREA)
- General Health & Medical Sciences (AREA)
- Software Systems (AREA)
- Computer Vision & Pattern Recognition (AREA)
- Databases & Information Systems (AREA)
- Medical Informatics (AREA)
- Multimedia (AREA)
- Life Sciences & Earth Sciences (AREA)
- Biomedical Technology (AREA)
- Biophysics (AREA)
- Computational Linguistics (AREA)
- Data Mining & Analysis (AREA)
- Molecular Biology (AREA)
- General Engineering & Computer Science (AREA)
- Mathematical Physics (AREA)
- Processing Or Creating Images (AREA)
Abstract
本公开涉及图像生成技术领域,公开了图像生成方法、装置、设备、存储介质及程序产品,该方法包括:获取目标提示信息;基于目标提示信息以及目标图像生成模型,生成目标提示信息的目标图像,目标图像生成模型是基于目标奖励模型对预设图像生成模型的参数进行调整后得到的,目标奖励模型的输入包括第一样本提示信息以及预设图像生成模型基于第一样本提示信息所得到的待优化图像,目标奖励模型用于预测待优化图像中待优化的图像特征信息。
Description
相关申请的交叉引用
本申请要求于2024年3月27日提交的,申请号为202410362124.8、发明名称为“图像生成方法、装置、设备、存储介质及程序产品”的中国专利申请的优先权,该申请的全部内容通过引用结合在本申请中。
本公开涉及图像生成技术领域,具体涉及图像生成方法、装置、设备、存储介质及程序产品。
目前,主要通过文生图扩散模型在给定的提示文本下生成多样化的图像。在相关技术中,对于文生图扩散模型的反馈学习,往往预先训练一个全局奖励模型,通过该奖励模型对文生图扩散模型进行微调,以优化文生图扩散模型的图像生成效果。
有鉴于此,本公开提供了一种图像生成方法、装置、设备、存储介质及程序产品,以解决图像生成效果不佳的问题。
第一方面,本公开提供了一种图像生成方法,所述方法包括:
获取目标提示信息;
基于所述目标提示信息以及目标图像生成模型,生成所述目标提示信息的目标图像,所述目标图像生成模型是基于目标奖励模型对预设图像生成模型的参数进行调整后得到的,所述目标奖励模型的输入包括第一样本提示信息以及所述预设图像生成模型基于所述第一样本提示信息所得到的待优化图像,所述目标奖励模型用于预测所述待优化图像中待优化的图像特征信息。
第二方面,本公开提供了一种图像生成装置,所述装置包括:
提示信息获取模块,用于获取目标提示信息;
目标图像生成模块,用于基于所述目标提示信息以及目标图像生成模型,生成所述目标提示信息的目标图像,所述目标图像生成模型是基于目标奖励模型对预设图像生成模型的参数进行调整后得到的,所述目标奖励模型的输入包括第一样本提示信息以及所述预设图像生成模型基于所述第一样本提示信息所得到的待优化图像,所述目标奖励模型用于预测所述待优化图像中待优化的图像特征信息。
第三方面,本公开提供了一种计算机设备,包括:存储器和处理器,存储器和处理器之间互相通信连接,存储器中存储有计算机指令,处理器通过执行计算机指令,从而执行上述第一方面或其对应的任一实施方式的图像生成方法。
第四方面,本公开提供了一种计算机可读存储介质,该计算机可读存储介质上存储有计算机指令,计算机指令用于使计算机执行上述第一方面或其对应的任一实施方式的图像生成方法。
第五方面,本公开提供了一种计算机程序产品,包括计算机指令,计算机指令用于使计算机执行上述第一方面或其对应的任一实施方式的图像生成方法。
本公开实施例提供的图像生成方法,将第一样本提示信息以及预设图像生成模型基于第一样本提示信息所得到的待优化图像作为目标奖励模型的输入,并基于目标奖励模型预测待优化图像中待优化的图像特征信息,以确定预设图像生成模型在生成图像时待优化的维度。然后,利用目标奖励模型对预设图像生成模型的参数进行调整,使预设图像生成模型朝着待优化的维度进行优化。
为了更清楚地说明本公开具体实施方式或现有技术中的技术方案,下面将对具体实施方式或现有技术描述中所需要使用的附图作简单地介绍,显而易见地,下面描述中的附图是本公开的一些实施方式,对于本领域普通技术人员来讲,在不付出创造性劳动的前提下,还可以根据这些附图获得其他的附图。
图1是根据本公开实施例的一种图像生成方法的流程示意图;
图2是根据本公开实施例的一种目标图像生成模型的确定方式的流程示意图;
图3是根据本公开实施例的另一种目标图像生成模型的确定方式的流程示意图;
图4是根据本公开实施例的一种目标奖励模型的确定方式的流程示意图;
图5是根据本公开实施例的另一种目标奖励模型的确定方式的流程示意图;
图6是根据本公开实施例的一种图像生成装置的结构框图;
图7是根据本公开实施例的一种计算机设备的结构框图。
为使本公开实施例的目的、技术方案和优点更加清楚,下面将结合本公开实施例中的附图,对本公开实施例中的技术方案进行清楚、完整地描述,显然,所描述的实施例是本公开一部分实施例,而不是全部的实施例。基于本公开中的实施例,本领域技术人员在没有做出创造性劳动前提下所获得的所有其他实施例,都属于本公开保护的范围。
可以理解的是,在使用本公开各实施例公开的技术方案之前,均应当依据相关法律法规通过恰当的方式对本公开所涉及个人信息的类型、使用范围、使用场景等告知用户并获得用户的授权。
例如,在响应于接收到用户的主动请求时,向用户发送提示信息,以明确地提示用户,其请求执行的操作将需要获取和使用到用户的个人信息。从而,使得用户可以根据提示信息来自主地选择是否向执行本公开技术方案的操作的电子设备、应用程序、服务器或存储介质等软件或硬件提供个人信息。
作为一种可选的但非限定性的实现方式,响应于接收到用户的主动请求,向用户发送提示信息的方式例如可以是弹窗的方式,弹窗中可以以文字的方式呈现提示信息。此外,弹窗中还可以承载供用户选择“同意”或者“不同意”向电子设备提供个人信息的选择控件。
可以理解的是,上述通知和获取用户授权过程仅是示意性的,不对本公开的实现方式构成限定,其它满足相关法律法规的方式也可应用于本公开的实现方式中。
可以理解的是,本技术方案所涉及的数据(包括但不限于数据本身、数据的获取或使用)应当遵循相应法律法规及相关规定的要求。
目前,主要通过文生图扩散模型在给定的提示文本下生成多样化的图像。在相关技术中,对于文生图扩散模型的反馈学习,往往预先训练一个全局奖励模型,通过该奖励模型对文生图扩散模型进行微调,以优化文生图扩散模型的图像生成效果。但是,在这种微调方式中,全局奖励模型是对文生图扩散模型进行全局维度的奖励,通常是微调文生图扩散模型的扩散朝着某个特定的维度进行优化,例如图像的整体美感,难以实现细粒度层面的维度优化,从而导致这种微调方式下的文生图扩散模型难以适用于不同的应用场景,造成图像生成效果不佳。
有鉴于此,根据本公开实施例,提供了一种图像生成方法实施例,需要说明的是,在附图的流程图示出的步骤可以在诸如一组计算机可执行指令的计算机系统中执行,并且,虽然在流程图中示出了逻辑顺序,但是在某些情况下,可以以不同于此处的顺序执行所示出或描述的步骤。
在本实施例中提供了一种图像生成方法,可用于上述的计算机设备,如手机、平板电脑等。图1是根据本公开实施例的一种图像生成方法的流程示意图,如图1所示,该流程包括如下步骤:
步骤S101,获取目标提示信息。
需要说明的是,目标提示信息可以是提示文本,也可以是提示图像,在此不对目标提示信息的类型进行限定。
步骤S102,基于目标提示信息以及目标图像生成模型,生成目标提示信息的目标图像,目标图像生成模型是基于目标奖励模型对预设图像生成模型的参数进行调整后得到的,目标奖励模型的输入包括第一样本提示信息以及预设图像生成模型基于第一样本提示信息所得到的待优化图像,目标奖励模型用于预测待优化图像中待优化的图像特征信息。
需要说明的是,目标图像生成模型可采用文本到图像的生成对抗网络模型、扩散模型或者语言大模型等大型预训练模型。目标图像生成模型所对应的图像特征信息包括风格一致性、内容一致性、色彩、纹理、氛围感以及布局中的至少之一。此外,图像特征信息还可以包括轮廓或者扩建关系等其他特征信息。
本实施例提供的图像生成方法,将第一样本提示信息以及预设图像生成模型基于第一样本提示信息所得到的待优化图像作为目标奖励模型的输入,并基于目标奖励模型预测待优化图像中待优化的图像特征信息,以确定预设图像生成模型在生成图像时待优化的维度。然后,利用目标奖励模型对预设图像生成模型的参数进行调整,使预设图像生成模型朝着待优化的维度进行优化。因此,能够提升预设图像生成模型在待优化维度的图像生成效果,以提升目标图像生成模型在不同维度的图像生成效果,从而适用于不同的应用场景。
在一些可选的实施方式中,目标奖励模型包括至少两个目标奖励子模型,其中,目标奖励子模型包括具有层级归属关系的多个待优化局部维度节点,目标奖励子模型与待优化全局维度的图像特征相对应,待优化全局维度包括至少一个待优化局部维度,待优化局部维度节点中的叶子结点用于表征待优化局部维度的图像特征。
具体地,本公开的目标奖励模型可视为多个层级的树状结构的奖励模型,目标奖励模型的输入端可视为一个分类器,用于从全局维度预测当前输入数据对应于各个目标奖励子模型的分类概率。目标奖励子模型的第1层级的待优化局部维度节点对应于待优化全局维度的图像特征。例如,假设至少两个待优化全局维度包括图文一致性维度以及美感维度。则,目标奖励模型包括图文一致性维度的目标奖励子模型以及美感维度的目标奖励子模型,图文一致性维度的目标奖励子模型在第1层级的待优化局部维度节点对应于图文一致性维度的图像特征,美感维度的目标奖励子模型在第1层级的待优化局部维度节点对应于美感维度的图像特征。目标奖励子模型的第2层级的待优化局部维度节点对应于待优化局部维度的图像特征。例如,假设图文一致性维度下的待优化局部维度包括风格一致性维度以及内容一致性维度,美感维度下的待优化局部维度包括色彩维度、纹理维度、氛围感维度以及布局维度。则,图文一致性维度的目标奖励子模型在第2层级设有2个待优化局部维度节点并分别对应于风格一致性维度的图像特征以及内容一致性维度的图像特征,美感维度的目标奖励子模型在第2层级设有4个待优化局部维度节点并分别对应于色彩维度的图像特征、纹理维度的图像特征、氛围感维度的图像特征以及布局维度的图像特征。此外,还可以根据实际情况,继续在第2层级的待优化全局维度节点下添加第3层级、第4层级或其他更多层级的待优化局部维度节点,并将最后一个层级的待优化局部维度节点作为叶子结点。例如,若奖励模型仅设有2个层级,则可以将上述风格一致性维度以及内容一致性维度对应的待优化局部维度节点视为图文一致性维度的目标奖励子模型的叶子结点,将上述色彩维度、纹理维度、氛围感维度以及布局维度对应的待优化局部维度节点视为美感维度的目标奖励子模型的叶子结点。
本实施例提供的图像生成方法,由于目标奖励模型包括至少两个目标奖励子模型,且每个目标奖励子模型包括具有层级归属关系的多个待优化局部维度节点,目标奖励子模型与待优化全局维度的图像特征相对应,待优化全局维度包括至少一个待优化局部维度,每个待优化局部维度节点中的叶子结点用于表征待优化局部维度的图像特征。因此,能够利用目标奖励模型从多个细粒度层级的维度对预设图像生成模型所生成的图像进行评估,以准确确定出预设图像生成模型在生成图像时所欠缺的维度。
在一些可选的实施方式中,如图2所示,目标图像生成模型的确定方式包括:
步骤S201,获取第一样本提示信息。
步骤S202,将第一样本提示信息输入预设图像生成模型中,以得到待优化图像。
步骤S203,将待优化图像以及第一样本提示信息输入目标奖励模型中,得到图像损失,图像损失用于表征待优化图像中待优化的图像特征信息。
步骤S204,基于图像损失对预设图像生成模型的参数进行更新,以得到目标图像生成模型。
本实施例提供的图像生成方法,将第一样本提示信息以及预设图像生成模型基于第一样本提示信息得到的待优化图像输入目标奖励模型,因此,能够利用目标奖励模型确定预设图像生成模型当前生成的图像中待优化的图像特征信息,即图像损失。然后,利用图像损失对预设图像生成模型的参数进行更新,因此,能够优化预设图像生成模型在待优化的图像特征信息这一维度的图像生成效果。
在一些可选的实施方式中,预设图像生成模型是基于扩散模型构建的,上述步骤S202中的将第一样本提示信息输入预设图像生成模型中,以得到待优化图像,包括:将第一样本提示信息输入预设图像生成模型中进行加噪以及去噪处理;提取预设去噪步数范围内所得到的去噪图像,得到待优化图像。
可选地,预设去噪步数范围为预设图像生成模型在去噪处理过程中的最后5到10步。需要说明的是,预设去噪步数范围的选取,应使所选取的去噪图像中包括较多的非噪声的图像信息为准。在实际操作中,可根据实现情况对预设去噪步数范围进行调整。
本实施例提供的图像生成方法,将第一样本提示信息输入基于扩散模型构建的预设图像生成模型中进行加噪以及去噪处理,然后提取预设去噪步数范围内所得到的去噪图像,以作为输入目标奖励模型的待优化图像。由于预设去噪步数范围内所得到的去噪图像中包括较多的能够反映预设图像生成模型的图像生成效果的信息,因此,将预设去噪步数范围内的去噪图像作为输入目标奖励模型的待优化图像,能够使目标奖励模型准确确定出预设图像生成模型当前生成图像所欠缺的维度,进一步提高预设图像生成模型的微调精度。
在一些可选的实施方式中,上述步骤S203中的将待优化图像以及第一样本提示信息输入目标奖励模型中,得到图像损失,包括:
步骤a1,将待优化图像以及第一样本提示信息输入目标奖励模型中,利用至少两个目标奖励子模型进行待优化全局维度的预测,得到对应于目标奖励子模型的第一预测概率。
可以理解地,目标奖励模型及其各个节点可视为一个分类器,待优化全局维度的目标奖励子模型能够基于待优化图像以及第一样本提示信息确定待优化图像在各个待优化全局维度的图像生成效果,以判断待优化图像在哪个待优化全局维度上存在欠缺,以得到对应于目标奖励子模型的第一预测概率。
示例性地,假设待优化图像在某一待优化全局维度的图像生成效果越差则对应的第一预测概率越高。如图3所示,若预测待优化图像在图文一致性维度的目标奖励子模型的第一预设概率为0.2,在美感维度的目标奖励子模型的第一预设概率为0.8,则表明待优化图像在美感维度的图像生成效果不佳。因此,在后续过程中,需要从美感维度对预设图像生成模型进行微调。
步骤a2,基于待优化全局维度的预测结果,分别预测对应于叶子结点的第二预测概率以及奖励值。
具体地,在获取到待优化全局维度的预测结果后,目标奖励子模型下的待优化局部维度节点会进一步确定预测待优化图像在各个待优化局部维度的图像生成效果,以判断预测待优化图像在哪个待优化局部维度上存在欠缺,以得到对应于叶子结点的第二预测概率以及奖励值。
示例性地,假设待优化图像在某一待优化局部维度的图像生成效果越差则对应的叶子结点的奖励值越小。如图3所示,若风格一致性维度的待优化局部节点输出的奖励值为0.6,内容一致性维度的待优化局部节点输出的奖励值为2.0,色彩维度的叶子结点输出的奖励值为1.4、纹理维度的叶子结点输出的奖励值为0.9、氛围感维度的叶子结点输出的奖励值为-2.1以及布局维度的叶子结点输出的奖励值为-2.7,则表明待优化图像在氛围感维度以及布局维度的图像生成效果不佳。因此,在图3所示的柱状图中,氛围感维度的叶子结点以及布局维度的叶子结点对应的柱条相对较高,表明在后续图像生成模型的微调过程中,需要重点关注模型在氛围感维度以及布局维度的图像生成效果。
步骤a3,对于每个目标奖励子模型,基于第二预测概率与奖励值的融合得到第一损失。
具体地,可将第二预测概率作为对应的奖励值的第一权重,并基于第一权重对同一目标奖励子模型的所有奖励值进行加权计算,得到各个目标奖励子模型的第一损失。
步骤a4,基于第一预测概率对所有目标奖励子模型对应的第一损失进行融合,得到图像损失。
具体地,将第一预测概率作为对应的第一损失的第二权重,并基于第二权重对所有第一损失进行加权计算,得到图像损失。
本实施例提供的图像生成方法,利用至少两个待优化全局维度的目标奖励子模型,确定待优化图像对应于各个目标奖励子模型的第一预测概率。并基于待优化全局维度的预测结果进行待优化局部维度的预测,以确定待优化图像对应于各个待优化局部维度的第二预测概率以及奖励值。然后,基于第一预测概率、第二预测概率以及奖励值,确定图像损失,从而能够使预设图像生成模型基于图像损失确定当前优化方向,以通过微调提升图像生成效果。
在一些可选的实施方式中,如图4所示,目标奖励模型的确定方式包括:
步骤S301,获取第二样本提示信息、与第二样本提示信息对应的图像对以及待优化维度标签,图像对包括对应于待优化维度标签的正样本图像以及负样本图像。
步骤S302,确定待优化维度标签在预设奖励模型中所对应的叶子结点,以确定所属的奖励子模型。
步骤S303,利用所属的奖励子模型,对第二样本提示信息以及图像对进行待优化维度的预测,待优化维度包括待优化全局维度以及待优化局部维度。
步骤S304,基于预测的结果与待优化维度标签,确定第二损失。
步骤S305,基于第二损失对所属的奖励子模型的参数进行更新,以确定目标奖励模型。
本实施例提供的图像生成方法,确定待优化维度标签在预设奖励模型中对应的叶子结点,以确定所属的奖励子模型。然后,利用所属的奖励子模型对第二样本提示信息以及图像对进行待优化维度的预测,以基于预测的结果与待优化维度标签,衡量当前奖励模型的输出与预期输出的差异性,得到第二损失。从而能够基于第二损失对所属的奖励子模型的参数进行更新,以使更新后的奖励子模型能够准确判断输入图像在对应待优化维度的图像生成效果,从而提高目标奖励模型判断输入图像在多个维度的图像生成效果的准确性。
示例性地,如图5所示,可预先采集多个细粒度维度的偏好数据,以得到多个维度的训练样本。每个训练样本中包括第二样本提示信息、与第二样本提示信息对应的图像对以及待优化维度标签。假设当前训练样本中的待优化维度标签为风格一致性,则可以确定所属的奖励子模型为图文一致性维度的奖励子模型。可利用图文一致性维度的奖励子模型中对应于图文一致性维度以及风格一致性维度的待优化局部维度节点,对第二样本提示信息以及图像对进行待优化全局维度以及待优化局部维度的预测,以得到预测的结果。
在一些可选的实施方式中,上述步骤S303中的利用所属的奖励子模型,对第二样本提示信息以及图像对进行待优化维度的预测,包括:基于第二样本提示信息以及图像对,确定所属的奖励子模型中待优化全局维度对应的第一分类概率;基于第二样本提示信息以及图像对,确定所属的奖励子模型中叶子结点对应的第二分类概率,预测的结果包括第一分类概率以及第二分类概率。
需要说明的是,所属的奖励子模型可视为一个分类器,在将第二样本提示信息以及图像对输入所属的奖励子模型时,所属的奖励子模型会对输入数据进行分类,以得到输入数据对应于待优化全局维度的第一分类概率以及对应于叶子结点的第二分类概率。
在一些可选的实施方式中,上述步骤S304中的基于预测的结果与待优化维度标签,确定第二损失,包括:基于第二分类概率与待优化维度标签确定奖励值;基于奖励值以及第一分类概率的融合结果,确定第二损失。
需要说明的是,为了使所属的奖励子模型能够准确判断输入图像在待优化维度的图像生成效果,因此,需要以扩大正样本图像与负样本图像之间的差异作为目标,利用第二分类概率与待优化维度标签确定奖励值,以使奖励值能够反映待优化维度对应的正样本图像与负样本图像之间的差异。然后,基于奖励值以及第一分类概率的融合结果,确定第二损失,以利用第二损失对预设奖励模型的参数进行更新,从而使训练完成后的目标奖励模型的叶子结点输出的奖励值能够反映输入图像在对应维度的图像生成效果。
值得说明的是,本公开的图像生成方法按照多个细粒度层级的偏好维度构造了树状结构的奖励模型,并利用不同偏好维度的训练样本对奖励模型进行训练,以得到能够用于判断输入图像在不同偏好维度的图像生成效果的目标奖励模型。从而能够在图像生成模型进行反馈学习时,利用目标奖励模型对图像生成模型在多个偏好维度的图像生成效果进行打分,以自适应预测图像生成模型在生成图像时欠缺的偏好维度,进而使图像生成模型朝着所欠缺的偏好维度进行优化,以提高图像生成模型在多个偏好维度的图像生成效果。
在本实施例中还提供了一种图像生成装置,该装置用于实现上述实施例及优选实施方式,已经进行过说明的不再赘述。如以下所使用的,术语“模块”可以实现预定功能的软件和/或硬件的组合。尽管以下实施例所描述的装置较佳地以软件来实现,但是硬件,或者软件和硬件的组合的实现也是可能并被构想的。
本实施例提供一种图像生成装置,如图6所示,包括:
提示信息获取模块401,用于获取目标提示信息;
目标图像生成模块402,用于基于目标提示信息以及目标图像生成模型,生成目标提示信息的目标图像,目标图像生成模型是基于目标奖励模型对预设图像生成模型的参数进行调整后得到的,目标奖励模型的输入包括第一样本提示信息以及预设图像生成模型基于第一样本提示信息所得到的待优化图像,目标奖励模型用于预测待优化图像中待优化的图像特征信息。
在一些可选的实施方式中,目标图像生成模块402中的目标奖励模型包括至少两个目标奖励子模型,其中,目标奖励子模型包括具有层级归属关系的多个待优化局部维度节点,目标奖励子模型与待优化全局维度的图像特征相对应,待优化全局维度包括至少一个待优化局部维度,待优化局部维度节点中的叶子结点用于表征待优化局部维度的图像特征。
在一些可选的实施方式中,图像生成装置还包括图像模型确定模块,用于确定目标图像生成模型。其中,图像模型确定模块,包括:
第一信息获取模块,用于获取第一样本提示信息;
第一图像获取模块,用于将第一样本提示信息输入预设图像生成模型中,以得到待优化图像;
图像损失计算模块,用于将待优化图像以及第一样本提示信息输入目标奖励模型中,得到图像损失,图像损失用于表征待优化图像中待优化的图像特征信息;
图像模型更新模块,用于基于图像损失对预设图像生成模型的参数进行更新,以得到目标图像生成模型。
在一些可选的实施方式中,预设图像生成模型是基于扩散模型构建的,第一图像获取模块包括:
图像处理单元,用于将第一样本提示信息输入预设图像生成模型中进行加噪以及去噪处理;
图像提取单元,用于提取预设去噪步数范围内所得到的去噪图像,得到待优化图像。
在一些可选的实施方式中,图像损失计算模块包括:
全局维度预测单元,用于将待优化图像以及第一样本提示信息输入目标奖励模型中,利用至少两个目标奖励子模型进行待优化全局维度的预测,得到对应于目标奖励子模型的第一预测概率;
局部维度预测单元,用于基于待优化全局维度的预测结果,分别预测对应于叶子结点的第二预测概率以及奖励值;
第一损失计算单元,用于对于每个目标奖励子模型,基于第二预测概率与奖励值的融合得到第一损失;
模型损失融合单元,用于基于第一预测概率对所有目标奖励子模型对应的第一损失进行融合,得到图像损失。
在一些可选的实施方式中,图像生成装置还包括奖励模型确定模块,用于确定目标奖励模型。其中,奖励模型确定模块,包括:
第二信息获取模块,用于获取第二样本提示信息、与第二样本提示信息对应的图像对以及待优化维度标签,图像对包括对应于待优化维度标签的正样本图像以及负样本图像;
优化路径确定模块,用于确定待优化维度标签在预设奖励模型中所对应的叶子结点,以确定所属的奖励子模型;
优化维度预测模块,用于利用所属的奖励子模型,对第二样本提示信息以及图像对进行待优化维度的预测,待优化维度包括待优化全局维度以及待优化局部维度;
预测损失计算模块,用于基于预测的结果与待优化维度标签,确定第二损失;
奖励模型更新模块,用于基于第二损失对所属的奖励子模型的参数进行更新,以确定目标奖励模型。
在一些可选的实施方式中,优化维度预测模块包括:
第一分类单元,用于基于第二样本提示信息以及图像对,确定所属的奖励子模型中待优化全局维度对应的第一分类概率;
第二分类单元,用于基于第二样本提示信息以及图像对,确定所属的奖励子模型中叶子结点对应的第二分类概率,预测的结果包括第一分类概率以及第二分类概率。
在一些可选的实施方式中,预测损失计算模块包括:
奖励值确定单元,用于基于第二分类概率与待优化维度标签确定奖励值;
第二损失计算单元,用于基于奖励值以及第一分类概率的融合结果,确定第二损失。
上述各个模块和单元的更进一步的功能描述与上述对应实施例相同,在此不再赘述。
本实施例中的图像生成装置是以功能单元的形式来呈现,这里的单元是指ASIC(Application Specific Integrated Circuit,专用集成电路)电路,执行一个或多个软件或固定程序的处理器和存储器,和/或其他可以提供上述功能的器件。
本公开实施例还提供一种计算机设备,具有上述图6所示的图像生成装置。
请参阅图7,图7是本公开可选实施例提供的一种计算机设备的结构框图,如图7所示,该计算机设备包括:一个或多个处理器501、存储器502,以及用于连接各部件的接口,包括高速接口和低速接口。各个部件利用不同的总线互相通信连接,并且可以被安装在公共主板上或者根据需要以其它方式安装。处理器可以对在计算机设备内执行的指令进行处理,包括存储在存储器中或者存储器上以在外部输入/输出装置(诸如,耦合至接口的显示设备)上显示GUI的图形信息的指令。在一些可选的实施方式中,若需要,可以将多个处理器和/或多条总线与多个存储器和多个存储器一起使用。同样,可以连接多个计算机设备,各个设备提供部分必要的操作(例如,作为服务器阵列、一组刀片式服务器、或者多处理器系统)。图7中以一个处理器501为例。
处理器501可以是中央处理器,网络处理器或其组合。其中,处理器501还可以进一步包括硬件芯片。上述硬件芯片可以是专用集成电路,可编程逻辑器件或其组合。上述可编程逻辑器件可以是复杂可编程逻辑器件,现场可编程逻辑门阵列,通用阵列逻辑或其任意组合。
其中,存储器502存储有可由至少一个处理器501执行的指令,以使至少一个处理器501执行实现上述实施例示出的方法。
存储器502可以包括存储程序区和存储数据区,其中,存储程序区可存储操作系统、至少一个功能所需要的应用程序;存储数据区可存储根据计算机设备的使用所创建的数据等。此外,存储器502可以包括高速随机存取存储器,还可以包括非瞬时存储器,例如至少一个磁盘存储器件、闪存器件、或其他非瞬时固态存储器件。在一些可选的实施方式中,存储器502可选包括相对于处理器501远程设置的存储器,这些远程存储器可以通过网络连接至该计算机设备。上述网络的实例包括但不限于互联网、企业内部网、局域网、移动通信网及其组合。
存储器502可以包括易失性存储器,例如,随机存取存储器;存储器也可以包括非易失性存储器,例如,快闪存储器,硬盘或固态硬盘;存储器502还可以包括上述种类的存储器的组合。
该计算机设备还包括输入装置503和输出装置504。处理器501、存储器502、输入装置503和输出装置504可以通过总线或者其他方式连接,图7中以通过总线连接为例。
输入装置503可接收输入的数字或字符信息,以及产生与该计算机设备的用户设置以及功能控制有关的键信号输入,例如触摸屏、小键盘、鼠标、轨迹板、触摸板、指示杆、一个或者多个鼠标按钮、轨迹球、操纵杆等。输出装置504可以包括显示设备、辅助照明装置(例如,LED)和触觉反馈装置(例如,振动电机)等。上述显示设备包括但不限于液晶显示器,发光二极管,显示器和等离子体显示器。在一些可选的实施方式中,显示设备可以是触摸屏。
本公开实施例还提供了一种计算机可读存储介质,上述根据本公开实施例的方法可在硬件、固件中实现,或者被实现为可记录在存储介质,或者被实现通过网络下载的原始存储在远程存储介质或非暂时机器可读存储介质中并将被存储在本地存储介质中的计算机代码,从而在此描述的方法可被存储在使用通用计算机、专用处理器或者可编程或专用硬件的存储介质上的这样的软件处理。其中,存储介质可为磁碟、光盘、只读存储记忆体、随机存储记忆体、快闪存储器、硬盘或固态硬盘等;进一步地,存储介质还可以包括上述种类的存储器的组合。可以理解,计算机、处理器、微处理器控制器或可编程硬件包括可存储或接收软件或计算机代码的存储组件,当软件或计算机代码被计算机、处理器或硬件访问且执行时,实现上述实施例示出的方法。
本公开的一部分可被应用为计算机程序产品,例如计算机程序指令,当其被计算机执行时,通过该计算机的操作,可以调用或提供根据本公开的方法和/或技术方案。本领域技术人员应能理解,计算机程序指令在计算机可读介质中的存在形式包括但不限于源文件、可执行文件、安装包文件等,相应地,计算机程序指令被计算机执行的方式包括但不限于:该计算机直接执行该指令,或者该计算机编译该指令后再执行对应的编译后程序,或者该计算机读取并执行该指令,或者该计算机读取并安装该指令后再执行对应的安装后程序。在此,计算机可读介质可以是可供计算机访问的任意可用的计算机可读存储介质或通信介质。
虽然结合附图描述了本公开的实施例,但是本领域技术人员可以在不脱离本公开的精神和范围的情况下做出各种修改和变型,这样的修改和变型均落入由所附权利要求所限定的范围之内。
Claims (12)
- 一种图像生成方法,包括:获取目标提示信息;基于所述目标提示信息以及目标图像生成模型,生成所述目标提示信息的目标图像,所述目标图像生成模型是基于目标奖励模型对预设图像生成模型的参数进行调整后得到的,所述目标奖励模型的输入包括第一样本提示信息以及所述预设图像生成模型基于所述第一样本提示信息所得到的待优化图像,所述目标奖励模型用于预测所述待优化图像中待优化的图像特征信息。
- 根据权利要求1所述的图像生成方法,其中所述目标奖励模型包括至少两个目标奖励子模型,其中,所述目标奖励子模型包括具有层级归属关系的多个待优化局部维度节点,所述目标奖励子模型与待优化全局维度的图像特征相对应,所述待优化全局维度包括至少一个待优化局部维度,所述待优化局部维度节点中的叶子结点用于表征所述待优化局部维度的图像特征。
- 根据权利要求2所述的图像生成方法,其中所述目标图像生成模型的确定方式包括:获取所述第一样本提示信息;将所述第一样本提示信息输入所述预设图像生成模型中,以得到所述待优化图像;将所述待优化图像以及所述第一样本提示信息输入所述目标奖励模型中,得到图像损失,所述图像损失用于表征所述待优化图像中待优化的图像特征信息;基于所述图像损失对所述预设图像生成模型的参数进行更新,以得到所述目标图像生成模型。
- 根据权利要求3所述的图像生成方法,其中所述预设图像生成模型是基于扩散模型构建的,所述将所述第一样本提示信息输入所述预设图像生成模型中,以得到所述待优化图像,包括:将所述第一样本提示信息输入所述预设图像生成模型中进行加噪以及去噪处理;提取预设去噪步数范围内所得到的去噪图像,得到所述待优化图像。
- 根据权利要求4所述的图像生成方法,其中所述将所述待优化图像以及所述第一样本提示信息输入所述目标奖励模型中,得到图像损失,包括:将所述待优化图像以及所述第一样本提示信息输入所述目标奖励模型中,利用所述至少两个目标奖励子模型进行所述待优化全局维度的预测,得到对应于所述目标奖励子模型的第一预测概率;基于所述待优化全局维度的预测结果,分别预测对应于所述叶子结点的第二预测概率以及奖励值;对于每个所述目标奖励子模型,基于所述第二预测概率与所述奖励值的融合得到第一损失;基于第一预测概率对所有所述目标奖励子模型对应的第一损失进行融合,得到所述图像损失。
- 根据权利要求2所述的图像生成方法,其中所述目标奖励模型的确定方式包括:获取第二样本提示信息、与所述第二样本提示信息对应的图像对以及待优化维度标签,所述图像对包括对应于所述待优化维度标签的正样本图像以及负样本图像;确定所述待优化维度标签在预设奖励模型中所对应的叶子结点,以确定所属的奖励子模型;利用所述所属的奖励子模型,对所述第二样本提示信息以及所述图像对进行待优化维度的预测,所述待优化维度包括所述待优化全局维度以及所述待优化局部维度;基于所述预测的结果与所述待优化维度标签,确定第二损失;基于所述第二损失对所述所属的奖励子模型的参数进行更新,以确定所述目标奖励模型。
- 根据权利要求6所述的图像生成方法,其中所述利用所述所属的奖励子模型,对所述第二样本提示信息以及所述图像对进行待优化维度的预测,包括:基于所述第二样本提示信息以及所述图像对,确定所述所属的奖励子模型中所述待优化全局维度对应的第一分类概率;基于所述第二样本提示信息以及所述图像对,确定所述所属的奖励子模型中所述叶子结点对应的第二分类概率,所述预测的结果包括所述第一分类概率以及所述第二分类概率。
- 根据权利要求7所述的图像生成方法,其中所述基于所述预测的结果与所述待优化维度标签,确定第二损失,包括:基于所述第二分类概率与所述待优化维度标签确定奖励值;基于所述奖励值以及所述第一分类概率的融合结果,确定所述第二损失。
- 一种图像生成装置,包括:提示信息获取模块,用于获取目标提示信息;目标图像生成模块,用于基于所述目标提示信息以及目标图像生成模型,生成所述目标提示信息的目标图像,所述目标图像生成模型是基于目标奖励模型对预设图像生成模型的参数进行调整后得到的,所述目标奖励模型的输入包括第一样本提示信息以及所述预设图像生成模型基于所述第一样本提示信息所得到的待优化图像,所述目标奖励模型用于预测所述待优化图像中待优化的图像特征信息。
- 一种计算机设备,包括:存储器和处理器,所述存储器和所述处理器之间互相通信连接,所述存储器中存储有计算机指令,所述处理器通过执行所述计算机指令,从而执行权利要求1至8中任一项所述的图像生成方法。
- 一种计算机可读存储介质,所述计算机可读存储介质上存储有计算机指令,所述计算机指令用于使计算机执行权利要求1至8中任一项所述的图像生成方法。
- 一种计算机程序产品,包括计算机指令,所述计算机指令用于使计算机执行权利要求1至8中任一项所述的图像生成方法。
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN202410362124.8 | 2024-03-27 | ||
| CN202410362124.8A CN120726147A (zh) | 2024-03-27 | 2024-03-27 | 图像生成方法、装置、设备、存储介质及程序产品 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2025200899A1 true WO2025200899A1 (zh) | 2025-10-02 |
Family
ID=97165093
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/CN2025/078809 Pending WO2025200899A1 (zh) | 2024-03-27 | 2025-02-24 | 图像生成方法、装置、设备、存储介质及程序产品 |
Country Status (2)
| Country | Link |
|---|---|
| CN (1) | CN120726147A (zh) |
| WO (1) | WO2025200899A1 (zh) |
Citations (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN116883530A (zh) * | 2023-07-06 | 2023-10-13 | 中山大学 | 一种基于细粒度语义奖励的文本到图像生成方法 |
| CN117095083A (zh) * | 2023-10-17 | 2023-11-21 | 华南理工大学 | 一种文本-图像生成方法、系统、装置和存储介质 |
| US20240070404A1 (en) * | 2022-08-26 | 2024-02-29 | International Business Machines Corporation | Reinforced generation: reinforcement learning for text and knowledge graph bi-directional generation using pretrained language models |
| CN117671055A (zh) * | 2023-11-23 | 2024-03-08 | 华为技术有限公司 | 数据处理方法、文图生成方法及相关装置 |
-
2024
- 2024-03-27 CN CN202410362124.8A patent/CN120726147A/zh active Pending
-
2025
- 2025-02-24 WO PCT/CN2025/078809 patent/WO2025200899A1/zh active Pending
Patent Citations (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20240070404A1 (en) * | 2022-08-26 | 2024-02-29 | International Business Machines Corporation | Reinforced generation: reinforcement learning for text and knowledge graph bi-directional generation using pretrained language models |
| CN116883530A (zh) * | 2023-07-06 | 2023-10-13 | 中山大学 | 一种基于细粒度语义奖励的文本到图像生成方法 |
| CN117095083A (zh) * | 2023-10-17 | 2023-11-21 | 华南理工大学 | 一种文本-图像生成方法、系统、装置和存储介质 |
| CN117671055A (zh) * | 2023-11-23 | 2024-03-08 | 华为技术有限公司 | 数据处理方法、文图生成方法及相关装置 |
Also Published As
| Publication number | Publication date |
|---|---|
| CN120726147A (zh) | 2025-09-30 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| US12333295B2 (en) | Intelligent generation and management of estimates for application of updates to a computing device | |
| JP2025516297A (ja) | 自然言語要求に応答したアクションセットの自動生成において利用される埋め込み及び/又はアクションモデルの変更 | |
| CN110727437B (zh) | 代码优化项获取方法、装置、存储介质及电子设备 | |
| CN114580263A (zh) | 基于知识图谱的信息系统故障预测方法及相关设备 | |
| US20230132033A1 (en) | Automatically generating, revising, and/or executing troubleshooting guide(s) | |
| JP7044839B2 (ja) | エンドツーエンドモデルのトレーニング方法および装置 | |
| CN113906416A (zh) | 可解释的过程预测 | |
| CN113778403B (zh) | 前端代码生成方法和装置 | |
| CN118519660A (zh) | 基于大模型的代码更新方法、代码生成方法、任务处理方法及代码解释器 | |
| CN107871088B (zh) | 一种信息处理方法、装置、终端和计算机可读存储介质 | |
| US9251489B2 (en) | Node-pair process scope definition adaptation | |
| CN115186738A (zh) | 模型训练方法、装置和存储介质 | |
| CN111340976B (zh) | 调试自动驾驶车辆模块的方法、装置、电子设备 | |
| CN118861982A (zh) | 多模态数据处理方法、装置、设备、介质和程序产品 | |
| KR102706150B1 (ko) | 서버 및 그 제어 방법 | |
| US9692657B2 (en) | Node-pair process scope definition and scope selection computation | |
| CN118860350A (zh) | 一种代码开发方法及相关设备 | |
| EP3762821A1 (en) | Neural network systems and methods for application navigation | |
| US11513862B2 (en) | System and method for state management of devices | |
| CN120726147A (zh) | 图像生成方法、装置、设备、存储介质及程序产品 | |
| CN117391002B (zh) | 一种ip核扩展描述方法及ip核生成方法 | |
| US11750479B1 (en) | System and method for managing issues based on pain reduction efficiency | |
| US20260044321A1 (en) | Code Development Method and Related Device | |
| CN118171202A (zh) | 任务分类方法、装置、电子设备及存储介质 | |
| CN119944656B (zh) | 一种储能设备参数模型更新方法、系统和相关装置 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 25776834 Country of ref document: EP Kind code of ref document: A1 |