WO2025112588A1 - 图像生成方法、电子设备及计算机可读存储介质 - Google Patents
图像生成方法、电子设备及计算机可读存储介质 Download PDFInfo
- Publication number
- WO2025112588A1 WO2025112588A1 PCT/CN2024/107883 CN2024107883W WO2025112588A1 WO 2025112588 A1 WO2025112588 A1 WO 2025112588A1 CN 2024107883 W CN2024107883 W CN 2024107883W WO 2025112588 A1 WO2025112588 A1 WO 2025112588A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- image
- information
- multimodal
- text
- target object
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Pending
Links
Classifications
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06T—IMAGE DATA PROCESSING OR GENERATION, IN GENERAL
- G06T11/00—Two-dimensional [2D] image generation
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F18/00—Pattern recognition
- G06F18/20—Analysing
- G06F18/24—Classification techniques
- G06F18/241—Classification techniques relating to the classification model, e.g. parametric or non-parametric approaches
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F40/00—Handling natural language data
- G06F40/20—Natural language analysis
- G06F40/205—Parsing
- G06F40/211—Syntactic parsing, e.g. based on context-free grammar [CFG] or unification grammars
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F40/00—Handling natural language data
- G06F40/30—Semantic analysis
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N20/00—Machine learning
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/04—Architecture, e.g. interconnection topology
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V10/00—Arrangements for image or video recognition or understanding
- G06V10/70—Arrangements for image or video recognition or understanding using pattern recognition or machine learning
- G06V10/764—Arrangements for image or video recognition or understanding using pattern recognition or machine learning using classification, e.g. of video objects
Definitions
- the present disclosure relates to the fields of large model technology and image processing technology, and in particular to an image generation method, an electronic device, and a computer-readable storage medium.
- the embodiments of the present disclosure provide an image generation method, an electronic device, and a computer-readable storage medium to at least solve the technical problem in the related art that corresponding images are generated only through text prompts, resulting in low accuracy of the generated images and a limited range of image generation.
- an image generation method comprising: acquiring multimodal prompt information, wherein the multimodal prompt information comprises: text information and enhanced mark information, the text information is used to describe the image content to be generated, the image content comprises: at least one target object, and the enhanced mark information is used to determine the position features and image features of the at least one target object; performing multimodal image generation on the multimodal prompt information using an image generation model to obtain a target image, wherein the image generation model is used to generate the target image using a multimodal image generation method.
- another image generation method which provides a graphical user interface through a terminal device, and the content displayed by the graphical user interface at least partially includes an image generation scene, including: in response to a first control operation performed on the graphical user interface, inputting text information, wherein the text information is used to describe the image content to be generated, and the image content includes: at least one target object; in response to a second control operation performed on the graphical user interface, based on the text information, respectively generating position features of at least one target object and image features of at least one target object to obtain enhanced marking information, and using an image generation model to perform multimodal image generation on multimodal prompt information to obtain a target image, wherein the enhanced marking information is used to determine the position features and image features of at least one target object, and the image generation model is used to generate the target image using a multimodal image generation method, and the multimodal prompt information includes: text information and enhanced marking information; and displaying the target image in the graphical user interface.
- another image generation method comprising: receiving a current A multimodal dialogue request is input, wherein the information carried in the multimodal dialogue request includes: multimodal dialogue text information and multimodal dialogue enhanced tag information, the multimodal dialogue text information is used to describe the multimodal dialogue image content to be generated, the multimodal dialogue image content includes: at least one object, and the multimodal dialogue enhanced tag information is used to determine the position feature and image feature of at least one object; using an image generation model to perform multimodal image generation on the multimodal dialogue request to obtain a multimodal dialogue image, wherein the image generation model is used to generate the multimodal dialogue image using a multimodal image generation method; and feeding back a multimodal dialogue response, wherein the information carried in the multimodal dialogue response includes: a multimodal dialogue image.
- an electronic device including: a memory storing an executable program; and a processor for running the program, wherein when the program is run, any one of the above-mentioned image generation methods is executed.
- a computer-readable storage medium including a stored executable program, wherein when the executable program runs, the device where the computer-readable storage medium is located is controlled to execute any one of the above-mentioned image generation methods.
- multimodal prompt information including text information and enhanced mark information
- FIG1 is a schematic diagram of an application scenario of an image generation method according to Embodiment 1 of the present disclosure
- FIG2 is a flow chart of an image generating method according to Embodiment 1 of the present disclosure.
- FIG3 is a flow chart of another image generating method according to Embodiment 1 of the present disclosure.
- FIG4 is a flow chart of an image generating method according to Embodiment 2 of the present disclosure.
- FIG5 is a flow chart of an image generating method according to Embodiment 3 of the present disclosure.
- FIG6 is a schematic diagram of a human-computer dialogue scenario according to Embodiment 3 of the present disclosure.
- FIG7 is a schematic diagram of the structure of an image generating device according to Embodiment 4 of the present disclosure.
- FIG8 is a schematic structural diagram of another image generating device according to Embodiment 4 of the present disclosure.
- FIG9 is a schematic structural diagram of yet another image generating device according to Embodiment 4 of the present disclosure.
- FIG10 is a structural block diagram of a computer terminal according to Embodiment 5 of the present disclosure.
- the technical solution provided by the present disclosure is mainly implemented by large-scale model technology.
- the large model here refers to a deep learning model with large-scale model parameters, which can usually contain hundreds of millions, tens of billions, hundreds of billions, trillions or even more than ten trillion model parameters.
- the large model can also be called the foundation model/foundation model (Foundation Model).
- the large model is pre-trained by large-scale unlabeled corpus to produce a pre-trained model with more than 100 million parameters.
- This model can adapt to a wide range of downstream tasks and has good generalization ability, such as large-scale language model (Large Language Model, LLM for short), multi-modal pre-training model (multi-modal pre-training model), etc.
- the pre-trained model can be fine-tuned through a small number of samples, so that the large model can be applied to different tasks.
- the large model can be widely used in natural language processing (NLP), computer vision, speech processing and other fields, specifically in computer vision tasks such as visual question answering (VQA), image caption (IC), image generation, etc. It can also be widely used in natural language processing tasks such as text-based sentiment classification, text summary generation, machine translation, etc. Therefore, the main application scenarios of the large model include but are not limited to digital assistants, intelligent robots, search, online education, office software, e-commerce, intelligent design, etc.
- Diffusion Model A mathematical model used to describe the diffusion process.
- the diffusion model can denoise a noisy image based on a diffusion process to obtain a generated image.
- Constituency tree A tree-like structure used to represent the structure of natural language sentences, consisting of nodes and edges, where nodes represent words or phrases in a sentence and edges represent the grammatical relationship between these words or phrases.
- sentences are decomposed into different phrase structures, such as noun phrases, verb phrases, adjective phrases, etc. These phrase structures can be further decomposed into smaller phrase structures, down to the smallest word level.
- the constituency tree can help understand the structure and grammatical relationship of sentences, facilitate syntactic and semantic analysis, and can also be used for natural language processing tasks such as syntactic analysis, machine translation, and information extraction. By analyzing the constituency tree, we can better understand the grammatical structure and semantic meaning of sentences.
- Natural Language Processing (NLP) parser A tool for natural language processing that can convert natural language text into structured data and perform tasks such as grammatical analysis, part-of-speech tagging, and named entity recognition. NLP parsers are usually based on machine learning and deep learning technologies, which can automatically analyze text and extract information from it, and can help understand and process large amounts of natural language data, such as speech recognition, sentiment analysis, text classification, etc.
- Text Encoder A model or tool used to convert text data into numerical representation.
- text encoders are often used to convert text data into vector or matrix form for machine learning and deep learning tasks, such as text classification, semantic similarity calculation, sentiment analysis, etc.
- Fine-tune refers to further training and adjustment on a specific task based on the use of a pre-trained model to improve the performance and adaptability of the model on the task.
- the pre-trained model is trained on a large-scale dataset, while Fine-tune is fine-tuned on a specific small-scale dataset to make the model better adapt to specific task requirements.
- Fine-tune is usually used in deep learning models in fields such as natural language processing and computer vision.
- the text-to-image model usually uses word embedding extracted from text prompt words as a condition.
- the high level of abstraction and limited information density of text make it difficult to accurately describe objects, so the accuracy of the generated images is low, and the method usually requires fine-tuning or only supports the use of a single object as a constraint, so the image generation range is relatively limited.
- the BLIP-Diffusion model has been proposed, which can support the generation of corresponding images by mixing images and texts, but it only supports the generation of single objects, and the accuracy of the generated images is low.
- the KOSMOS-G model has been proposed, which can support the generation of corresponding images by mixing images and texts, but the accuracy of the generated images is still low.
- the related art method of generating corresponding images based on text prompts has the following defects.
- Defect 1 The corresponding image is generated by describing the object only through text prompts. Since it is difficult to describe physics with text prompts, the accuracy of the generated image is low.
- Disadvantage 2 Usually requires fine-tuning or only supports using a single object as a constraint, so the model training cost is high Higher and more limited image generation range;
- Defect 3 The BLIP-Diffusion model and the KOSMOS-G model support mixed generation of images and text, but the accuracy of the generated images is low, and the BLIP-Diffusion model only supports single object generation.
- an image generating method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
- the above-mentioned image generation method provided in the embodiment of the present disclosure can be applied to the application scenario shown in Figure 1, but is not limited thereto.
- the large model is deployed in the server 10, and the server 10 can be connected to one or more client devices 20 through a local area network connection, a wide area network connection, an Internet connection, or other types of data networks.
- the client device 20 here may include but is not limited to: a smart phone, a tablet computer, a laptop computer, a PDA, a personal computer, a smart home device, a vehicle-mounted device, etc.
- the client device 20 can interact with the user through a graphical user interface to implement the call of the large model, thereby implementing the method provided in the embodiment of the present disclosure.
- the system composed of the client device and the server can execute the following steps: the client device executes the steps of obtaining multimodal prompt information input by the user in the graphical user interface and sending the multimodal prompt information to the server; the server executes the steps of using the image generation model to generate a multimodal image for the obtained multimodal prompt information to obtain a target image, and returns the target image to the client device.
- the embodiments of the present disclosure can be carried out in the client device.
- FIG2 is a flow chart of an image generation method according to Embodiment 1 of the present disclosure. As shown in FIG2, the method may include the following steps:
- Step S21 obtaining multimodal prompt information, wherein the multimodal prompt information includes: text information and enhanced mark information, the text information is used to describe the image content to be generated, the image content includes: at least one target object, and the enhanced mark information is used to determine the position feature and image feature of the at least one target object;
- Step S22 using an image generation model to generate a multimodal image for the multimodal prompt information to obtain a target image, wherein the image generation model is used to generate the target image using a multimodal image generation method.
- the text information can be understood as a text prompt, that is, the image content described in natural language, which is used to describe the image content to be generated.
- the natural language can be Chinese, English, Japanese, etc., which is not limited here.
- the image content described by the text information may include at least one target object.
- the target object can be understood as a person, animal or object in the image content to be generated, that is, the present disclosure can support multi-object image generation.
- the target object can include real people such as boys, girls, doctors, teachers, etc.
- Objects may include virtual characters such as game players and non-player characters (NPCs), animals such as cats, dogs, peacocks, elephants, etc., and objects such as tables, cars, grass, trees, and train stations, which are not limited here.
- the text information may be text information corresponding to the image generation requirement of the user. If the user wants to obtain an image of a cat and a dog on the grass, the corresponding text information may be "a cat and a dog on the grass (i.e., a cat and a dog on the grass)", and the text information includes three target objects, namely, a cat, a dog, and the grass. It is understandable that the text information may also be replaced by "a dog and a cat on the grass", or "there is a cat and a dog on the grass", and the present disclosure does not limit the description method of the text information.
- the present disclosure uses enhanced tag information as additional information on the basis of providing textual information, and generates a target image through the textual information and the enhanced tag information, thereby improving the accuracy of image generation.
- the enhanced tag information may include coordinate information and image information, wherein the coordinate information may be the coordinates of at least one target object included in the text information, and is used to determine the positional features of the at least one target object.
- the coordinate information includes the coordinates of the cat (e.g., [coordinates: 12, 15, 100, 200]), the coordinates of the dog (e.g., [coordinates: 100, 150, 220, 240]), and the coordinates of the grass (e.g., [coordinates: 500, 500, 500, 500]).
- the coordinates of at least one target object may be coordinates in a two-dimensional rectangular coordinate system, which can represent the coordinate range of the target object, or may be coordinates in a two-dimensional coordinate system or a three-dimensional coordinate system.
- the present disclosure does not limit the adopted coordinate system.
- the image information may be an image of at least one target object included in the text information, and is used to determine the image features of at least one target object in the text information. For example, if a user wants to obtain an image of a cat and a dog on grass, the image information includes an image of a cat (e.g., [image: an image of a cat uploaded by the user]), an image of a dog (e.g., [image: an image of a dog uploaded by the user]), and an image of grass (e.g., [image: an image of grass uploaded by the user]) that the user needs to generate.
- a cat e.g., [image: an image of a cat uploaded by the user]
- an image of a dog e.g., [image: an image of a dog uploaded by the user]
- an image of grass e.g., [image: an image of grass uploaded by the user]
- the multimodal prompt information can be understood as prompt information in multiple modes, including the above-mentioned text information and enhanced mark information.
- the multimodal prompt information includes prompt information in text mode, coordinate mode and image mode.
- multimodal prompt information of the same target object can be bound together to avoid affecting other objects or global conditions. That is, an augmented token is introduced in the present disclosure.
- the augmented token contains object-level text, coordinates, and image information at the same time, which is used to describe an object in the generated image.
- the image generation model is a model suitable for generating images.
- the image generation model can be a pre-trained image generation model based on a diffusion process, that is, a diffusion model.
- the image generation model can also be an autoregressive model, etc., which is not limited here.
- the image generation model can accurately generate high-quality target images based on multi-modal prompt information of multiple target objects, that is, it can generate high-quality images based on multi-modal prompts at the multi-object level. Image, thereby achieving precise control over the generated image.
- the image generation model of the present invention can accurately output the corresponding target image according to the multimodal prompt information, that is, the image of a cat and a dog on the grass.
- the image generation model in the embodiment of the present disclosure can maintain the original structure of the pre-trained text-to-image model as much as possible so as to integrate the extension of the existing model, and the image generation model in the embodiment of the present disclosure only changes the input of the model and does not need to change the architecture of the model, so the usability of the technology built on the basis of the model can be maintained.
- multimodal prompt information including text information and enhanced mark information
- the above-mentioned image generation method provided by the embodiments of the present disclosure can be applied to, but is not limited to, application scenarios involving image generation in the fields of script services, design services, game services, e-commerce services, educational services, legal services, medical services, conference services, social network services, financial product services, logistics services and navigation services.
- generating the required image materials according to the script or storyline in the script service generating design drawings according to the user's text, image, and coordinate description in the design service, generating corresponding character drawings according to the text description of the character image in the game, generating corresponding product drawings according to the description of the product in the e-commerce service, etc., which are not limited here.
- multimodal prompt information including text information and enhanced mark information
- step S21 obtaining multimodal prompt information includes the following method steps:
- Step S211 obtaining text information
- Step S212 generating at least one target object's location feature and at least one target object's location feature based on the text information.
- the image features of the object are used to obtain enhanced labeling information;
- Step S213 Combine the text information and the enhanced mark information to obtain multimodal prompt information.
- enhanced marking information can be obtained based on text information, that is, enhanced marking information missing from the target object can be generated, thereby realizing a flexible combination of multiple modalities.
- text information when acquiring multimodal prompt information, text information can be acquired first, and then based on the text information, position features of at least one target object and image features of at least one target object can be generated respectively, that is, object-level coordinates and images can be generated according to the text prompt, thereby obtaining corresponding enhanced marking information, and then the text information and the enhanced marking information can be combined to obtain multimodal prompt information.
- step S212 generating a location feature and an image feature of at least one target object based on the text information to obtain enhanced marking information includes the following method steps:
- Step S2121 performing part-of-speech analysis on the text information, and selecting at least one target participle, wherein the at least one target participle meets a preset part-of-speech requirement, and the at least one target participle is used to determine at least one target object;
- Step S2122 generating at least one position feature of a target object based on the text information and at least one target word segmentation, and generating at least one image feature of the target object based on the text information and the at least one position feature of the target object;
- Step S2123 Determine enhanced tag information using at least one target word segmentation, at least one location feature of a target object, and at least one image feature of a target object.
- the parts of speech in a language include nouns, verbs, adjectives, adverbs, pronouns, numerals, quantifiers, conjunctions, prepositions, auxiliary words, interjections, etc., and nouns are usually used to represent people, things, places, etc., and need to be generated when generating images, so nouns can be used to represent target objects.
- the preset part of speech may be a noun
- the at least one target participle is at least one noun in the text information.
- a constituency tree can be used to identify objects in text information, that is, to identify target objects in text information. For example, if the text information is "a cat and a dog on the grass", the text information is analyzed by the constituency tree, and the analysis results are "a (determiner) cat (noun) and (conjunction) a (determiner) dog (noun) on (preposition) the (determiner) grass (noun)", and the nouns in the text information, namely cat, dog and grass, are identified, so that the identified nouns cat, dog and grass are used as the selected multiple target participles.
- CRF Conditional Random Fields
- RNN Recurrent Neural Networks
- LSTM Long Short-Term Memory
- Attention Mechanism etc. for part-of-speech tagging and part-of-speech recognition, which are not restricted here.
- the text information when generating the location features and image features of at least one target object based on the text information and obtaining the enhanced marking information, can be subjected to part-of-speech analysis, and at least one target participle whose part of speech is a noun can be selected from the text information. Then, based on the text information and the at least one target participle, the location features and image features of at least one target object are generated. Then, based on the at least one target participle, the location features of at least one target object, and the image features of at least one target object, the enhanced marking information corresponding to each of the at least one target object is determined.
- step S2122 generating a location feature of at least one target object based on the text information and at least one target segmentation word includes the following method steps:
- Step S21221 using a position feature generation model to generate position features for text information and at least one target word segmentation, to obtain position features of at least one target object.
- the position feature generation model is a model suitable for generating position features, such as a diffusion model, an autoregressive model, etc.
- the position feature generation model may be a coordinate model, which can generate the position coordinates corresponding to each object based on the text content of the text information and the object corresponding to at least one target word segmentation.
- the text information and the at least one target participle can be input into a location feature generation model, and the location feature generation model is used to generate location features for the text information and the at least one target participle, thereby obtaining location features corresponding to the at least one target object, that is, obtaining coordinate information corresponding to the at least one target object.
- the multiple target segmented words are “cat” and “dog”.
- the present disclosure combines the three text features of “a cat and a dog”, “cat” and “dog” together and uses them as inputs to the position feature generation model, thereby obtaining the coordinates of the cat and the dog according to the output of the position feature generation model.
- a position feature generation model can be used to automatically generate position coordinates corresponding to multiple objects based on text prompts and multiple objects, thereby completing the missing coordinate modal information.
- step S2122 generating an image feature of at least one target object based on the text information and the position feature of at least one target object includes the following method steps:
- Step S21222 Use an image feature generation model to generate image features for the text information and the position features of at least one target object to obtain image features of at least one target object.
- the image feature generation model is a model suitable for generating image features, such as a diffusion model, an autoregressive model, etc.
- the image feature generation model may be an image feature model that can generate image features corresponding to at least one target object based on the text content of the text information and the position features of at least one target object.
- the text information and the position features of at least one target object can be input into an image feature generation model, and the image feature generation model can be used to generate image features for the text information and the coordinates of at least one target object, thereby obtaining image features corresponding to the at least one target object, that is, obtaining image information corresponding to the at least one target object.
- the present disclosure obtains the image features of the cat and the dog according to the output of the image feature generation model by inputting "a cat and a dog (text information)", “cat (text) + cat coordinates (position features)", and “dog (text) + dog coordinates (position features)” into the image feature generation model.
- an image feature generation model can be used to automatically generate image information corresponding to multiple objects based on text prompts and position coordinates of multiple objects, thereby completing the missing image modality information.
- the image generation model in the embodiment of the present disclosure can not only support the generation of target images according to text prompts, coordinate information and image information, but also support the generation of target images according to text prompts only. That is, the present disclosure can generate images based on multi-object multi-modal prompts. At the same time, when there is a modality missing problem, the generation model can also be used to complete the missing modality, so the application range is wider.
- the image generation method further includes the following method steps:
- Step S2124 using a text encoder to perform text encoding on the text information to obtain global text features, and using a text encoder to perform text encoding on at least one target word to obtain object text features.
- a text encoder can convert text data into a vector or matrix form, so as to perform machine learning and deep learning tasks, such as part-of-speech tagging, etc. Therefore, in the embodiment of the present disclosure, when processing text information, a text encoder can be used to perform text encoding processing on the text information, so as to obtain a global text encoding result, that is, a global text feature. At the same time, a text encoder can also be used to perform text encoding processing on at least one target word segmentation, so as to obtain a noun part encoding result, that is, an object text feature.
- an image encoder (Image Encoder) can be used to encode the image information to obtain image features, so as to obtain image features, which will not be elaborated here.
- step S2123 enhancing the marking information is determined by using at least one target word segmentation, at least one location feature of the target object, and at least one image feature of the target object, including the following method steps:
- Step S21231 combining the object text feature, the position feature of at least one target object, and the image feature of at least one target object to obtain enhanced marking information.
- the at least one target segmentation word when the enhanced tag information is determined by using at least one target segmentation word, at least one position feature of the target object, and at least one image feature of the target object, the at least one target segmentation word may be The object text features, the position features of at least one target object, and the image features of at least one target object obtained by the text encoding process are combined to obtain enhanced marking information.
- the object text features, the position features of at least one target object, and the image features of at least one target object can be horizontally spliced to obtain enhanced marking information, which is not limited here.
- the object text features “cat”, “dog” and “grass”, the location features of at least one target object “coordinates of the cat”, “coordinates of the dog” and “coordinates of the grass”, and the image features of at least one target object “image of a cat”, “image of a dog” and “image of the grass” can be combined to obtain enhanced marking information.
- step S213 the text information and the enhanced mark information are combined to obtain multimodal prompt information, including the following method steps:
- Step S2131 combining global text features, object text features, position features of at least one target object, and image features of at least one target object to obtain multimodal prompt information.
- global text features, object text features, location features of at least one target object, and image features of at least one target object may be combined to obtain multimodal prompt information.
- object text features, location features of at least one target object, and image features of at least one target object may be horizontally spliced, and then vertically spliced with global text features to obtain multimodal prompt information, which is not limited here.
- the text information is "a cat and a dog on the grass”
- the global text feature “a cat and a dog on the grass”
- the object text features “cat", “dog” and “grass”
- the position features of at least one target object “coordinates of the cat”
- the image features of at least one target object “image of a cat”
- “image of a dog” and “image of the grass” can be combined to obtain multimodal prompt information.
- step S21 obtaining multimodal prompt information includes the following method steps:
- Step S211 acquiring text information and additional information, wherein the additional information includes at least one of the following: location information of at least one target object, image information of at least one target object;
- Step S212 determining enhanced marking information based on the text information and the additional information
- Step S213 Combine the text information and the enhanced mark information to obtain multimodal prompt information.
- the text information may be understood as a text prompt input by a user for describing the content of the image to be generated.
- the additional information may be understood as coordinates and/or images input by the user. It may be understood that the additional information includes at least one of the coordinates of at least one target object input by the user and the image of at least one target object input by the user.
- the position features and image features for determining at least one target object can be determined according to the acquired text information and additional information.
- the enhanced mark information of the feature can be combined with the text information to obtain multimodal prompt information.
- the text information and the enhanced mark information may be vertically spliced to obtain multimodal prompt information, which is not limited here.
- FIG3 is a flow chart of another image generation method according to Embodiment 1 of the present disclosure.
- the image generation model of the present disclosure supports taking text information and enhanced mark information as input together to generate a target image, that is, supports combining text modality with coordinate modality and image modality to generate a target image based on multi-modal prompt information.
- the image generation model of the present disclosure also supports taking only text information as input when the modality is missing, thereby generating corresponding enhanced mark information according to the text information, that is, generating corresponding coordinate modality and image modality according to the text modality, and then generating a target image according to the text modality and the generated coordinate modality and image modality.
- At least one target word i.e., noun
- the coordinates of at least one target object, and the image of at least one target object in the text information can be determined based on the input text information, and then the object text feature corresponding to the at least one target word is determined by the text encoder, the position feature of at least one target object is determined based on the coordinates of at least one target object, and the image feature of at least one target object is determined by the image encoder.
- the object text feature, the position feature of at least one target object, and the image feature of at least one target object are combined to obtain enhanced tag information, and then the text information is embedded and combined with the enhanced tag information and input into the image generation model to finally obtain the target image.
- the text information may be "a cat and a dog on the grass", so it can be determined that multiple target segmentations include “dog", “cat” and "grass", the coordinates of at least one target object include [coordinates of the dog], [coordinates of the cat] and [coordinates of the grass], and the image of at least one target object includes an image of the dog, an image of the cat and an image of the grass.
- the object text features corresponding to the multiple target segmentations are determined by a text editor, the position features of at least one target object are determined according to the coordinates of at least one target object, and the image features of at least one target object are determined by an image encoder.
- the object text features, the position features of at least one target object and the image features of at least one target object are combined to obtain enhanced tag information, and then the text information "a cat and a dog on the grass" and the enhanced tag information are combined and input into the image generation model to finally obtain the target image.
- the coordinate model is used to generate position features for the text information and at least one target word segmentation to obtain the position features of at least one target object
- the image feature model is used to generate image features for the text information and the position features of at least one target object to obtain the image features of at least one target object.
- the object text features, the generated position features of at least one target object, and the image features of at least one target object are combined to obtain enhanced tag information, and then the text information is embedded and combined with the enhanced tag information and input into the image generation model to finally obtain the target image.
- the text information may be "a cat and a dog on the grass", so it can be segmented according to the word segmentation model.
- the image feature model uses the image feature model to generate image features for the text information and the position features of at least one target object, and obtain the image features corresponding to the image of at least one target object (the image of the dog, the image of the cat and the image of the grass). Then combine the object text features, the generated position features of at least one target object and the image features of at least one target object to obtain enhanced tag information, and then embed the text information and combine the enhanced tag information and input them into the image generation model to finally obtain the target image.
- the present invention obtains object-level text, coordinates, and images, and integrates this information into the "augmented token" of each object.
- the augmented token is trained in the diffusion model together with the text prompt as an additional condition, so that the image generation model of the present invention can handle multi-object multi-modal text prompts.
- the present disclosure proposes to use a coordinate model and an image feature model to generate object-level coordinates and image features according to text prompts. Therefore, the present disclosure can generate target images only through text prompts or through a combination of various multimodal prompts, and can generate target images in a flexible way by combining various modalities. And through a large number of qualitative and quantitative experiments, it is proved that the present disclosure is not only superior to the image generation method in the related art, but also can complete a wider range of image generation tasks.
- Beneficial effect (1) It supports the generation of high-quality images based on multi-modal prompts at the multi-object level, can achieve more precise control over the generated images, and has a wider range of applications;
- a coordinate model and an image feature model are designed to support the generation of coordinates and image modalities based on text modalities, thereby overcoming the possible modality missing problem;
- Beneficial effect (3) Most of the image generation models in the related art are generated based on text modality, while the image generation model disclosed in the present invention can generate target images based on text, image, and coordinate modalities at the same time, and the generated images have higher accuracy;
- Beneficial effect (4) If the image generation model of the related art wants to generate a given object, such as a dog, since there are many breeds of dogs, if you want to generate a specific dog, you generally need to give 3 to 5 pictures and fine-tune the model before the model can learn. However, the image generation model of the present invention does not require the fine-tuning process. It only needs to give a picture of a dog during inference to generate a dog of the corresponding breed, that is, it can achieve zero-sample generation. The target object does not need to be involved during training, so the model training cost is low.
- the present invention adds enhanced tags to text embedding as sampling conditions for the diffusion model to generate images, so that the present invention does not need to change the architecture of the diffusion model, but only needs to change the input of the diffusion model, thereby improving the usability of building technology based on the diffusion model.
- user information including but not limited to user device information, user personal information, etc.
- data including but not limited to data used for analysis, stored data, displayed data, etc.
- user information including but not limited to user device information, user personal information, etc.
- data including but not limited to data used for analysis, stored data, displayed data, etc.
- the method according to the above embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware.
- the technical solution of the present disclosure, or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a storage medium (such as ROM/RAM, a disk, or an optical disk), and includes a number of instructions for enabling a terminal device (which can be a mobile phone, a computer, a server, or a network device, etc.) to execute the methods described in each embodiment of the present disclosure.
- a storage medium such as ROM/RAM, a disk, or an optical disk
- FIG4 is a flow chart of an image generation method according to Example 2 of the present disclosure. As shown in FIG4, the method includes:
- Step S41 in response to a first control operation performed on a graphical user interface, inputting text information, wherein the text information is used to describe the image content to be generated, and the image content includes: at least one target object;
- Step S42 in response to a second control operation performed on the graphical user interface, respectively generating at least one target object's position feature and at least one target object's image feature based on the text information to obtain enhanced marking information, and performing multimodal image generation on the multimodal prompt information using an image generation model to obtain a target image, wherein the enhanced marking information is used to determine the position feature and the image feature of the at least one target object, the image generation model is used to generate the target image using a multimodal image generation method, and the multimodal prompt information includes: text information and enhanced marking information;
- Step S43 display the target image in the graphical user interface.
- the user can input text information for describing the content of the image to be generated in the image generation scene by executing control operations, control the generation of at least one position feature of the target object and at least one image feature of the target object based on the text information to obtain enhanced marking information, and control the use of the image generation model to generate multimodal images for the multimodal prompt information to obtain the target object.
- image generation scenarios may be, but are not limited to, application scenarios involving image generation in the fields of scripts, design, games, e-commerce, education, medical care, conferences, social networks, financial products, logistics, and navigation.
- the graphical user interface also includes a first control (or a first touch area).
- a first touch operation acting on the first control or the first touch area
- text information input by the user can be obtained.
- the text information can be input by the user from a text box in the graphical user interface through the first touch operation.
- the first touch operation can be a point selection, box selection, check selection, conditional screening and other operations, which are not limited here.
- the text information can be understood as a text prompt, that is, the image content described in natural language, which is used to describe the image content to be generated.
- the natural language can be Chinese, English, Japanese, etc., which are not limited here.
- the image content described by the text information may include at least one target object, and the target object can be understood as a person, animal or object in the image content to be generated, that is, the present disclosure can support the generation of multiple objects.
- the target object may include real people such as boys, girls, doctors, teachers, etc., may include virtual characters such as game players and non-player characters (NPC), may include animals such as cats, dogs, peacocks, elephants, etc., and may also include objects such as tables, cars, grass, trees, train stations, etc., which are not limited here.
- NPC non-player characters
- animals such as cats, dogs, peacocks, elephants, etc.
- objects such as tables, cars, grass, trees, train stations, etc., which are not limited here.
- the text information may be text information corresponding to the image generation requirement of the user. If the user wants to obtain an image of a cat and a dog on the grass, the corresponding text information may be "a cat and a dog on the grass (i.e., a cat and a dog on the grass)", and the text information includes three target objects, namely, a cat, a dog, and the grass. It is understandable that the text information may also be replaced by "a dog and a cat on the grass", or "there is a cat and a dog on the grass", and the present disclosure does not limit the description method of the text information.
- the graphical user interface also includes a second control (or a second touch area).
- a second touch operation acting on the second control or the second touch area
- the position feature of at least one target object and the image feature of at least one target object can be generated based on the text information to obtain enhanced marking information, and the image generation model can be used to generate a multimodal image of the multimodal prompt information to obtain a target image.
- the second touch operation can be a point selection, a box selection, a check, a conditional screening, etc., which are not limited here.
- the present disclosure uses enhanced tag information as additional information on the basis of providing textual information, and generates a target image through the textual information and the enhanced tag information, thereby improving the accuracy of image generation.
- the enhanced tag information may include coordinate information and image information, wherein the coordinate information may be the coordinates of at least one target object included in the text information, and is used to determine the positional features of the at least one target object.
- the coordinate information includes the coordinates of the cat (e.g., [coordinates: 12, 15, 100, 200]), the coordinates of the dog (e.g., [coordinates: 100, 150, 220, 240]), and the coordinates of the grass (e.g., [coordinates: 500, 500, 500, 500]).
- the coordinates of at least one target object may be coordinates in a two-dimensional rectangular coordinate system, which can represent
- the coordinate range of the target object may also be coordinates in a two-dimensional coordinate system or a three-dimensional coordinate system, and the present disclosure does not limit the adopted coordinate system.
- the image information may be an image of at least one target object included in the text information, and is used to determine the image features of at least one target object in the text information. For example, if a user wants to obtain an image of a cat and a dog on grass, the image information includes an image of a cat (e.g., [image: an image of a cat uploaded by the user]), an image of a dog (e.g., [image: an image of a dog uploaded by the user]), and an image of grass (e.g., [image: an image of grass uploaded by the user]) that the user needs to generate.
- a cat e.g., [image: an image of a cat uploaded by the user]
- an image of a dog e.g., [image: an image of a dog uploaded by the user]
- an image of grass e.g., [image: an image of grass uploaded by the user]
- the multimodal prompt information can be understood as prompt information in multiple modes, including the above-mentioned text information and enhanced mark information.
- the multimodal prompt information includes prompt information in text mode, coordinate mode and image mode.
- multimodal prompt information of the same target object can be bound together to avoid affecting other objects or global conditions. That is, an augmented token is introduced in the present disclosure.
- the augmented token contains object-level text, coordinates, and image information at the same time, which is used to describe an object in the generated image.
- the image generation model is a model suitable for generating images.
- the image generation model can be a pre-trained image generation model based on a diffusion process, that is, a diffusion model.
- the image generation model can also be an autoregressive model, etc., which is not limited here.
- the image generation model can accurately generate high-quality target images based on multi-modal prompt information of multiple target objects, that is, it can generate high-quality images based on multi-modal prompts at the multi-object level, thereby achieving precise control of the generated image.
- the image generation model of the present invention can accurately output the corresponding target image according to the multimodal prompt information, that is, the image of a cat and a dog on the grass.
- the image generation model in the embodiment of the present disclosure can maintain the original structure of the pre-trained text-to-image model as much as possible so as to integrate the extension of the existing model, and the image generation model in the embodiment of the present disclosure only changes the input of the model and does not need to change the architecture of the model, so the usability of the technology built on the basis of the model can be maintained.
- a graphical user interface is provided by a terminal device, and the content displayed by the graphical user interface at least partially includes an image generation scene. If the user performs a first control operation on the graphical user interface, the user inputs text information for describing the image content of at least one target object to be generated. If the user performs a second control operation on the graphical user interface, such as a submit operation, the position features of at least one target object and the image features of at least one target object can be generated based on the text information, respectively, thereby obtaining enhanced marking information.
- an image generation model can be used to perform multimodal image generation on multimodal prompt information to obtain a target image, and then the generated target image can be displayed in the graphical user interface for feedback to the user. The purpose of accurately generating a high-quality target image including multiple target objects through multimodal prompt information is achieved, and the generated image can be It achieves more precise control, generates images with higher accuracy, and supports the generation of images of multiple target objects, making the image generation range wider.
- first touch operation and the second touch operation can both be operations in which a user touches the display screen of the terminal device with a finger and touches the terminal device.
- the touch operation can include single-point touch and multi-point touch, wherein the touch operation of each touch point can include click, long press, heavy press, swipe, etc.
- the first touch operation and the second touch operation can also be touch operations implemented by input devices such as a mouse and a keyboard, which are not limited here.
- the above-mentioned image generation method provided by the embodiments of the present disclosure can be applied to, but is not limited to, application scenarios involving image generation in the fields of script services, design services, game services, e-commerce services, educational services, legal services, medical services, conference services, social network services, financial product services, logistics services and navigation services.
- generating the required image materials according to the script or storyline in the script service generating design drawings according to the user's text, image, and coordinate description in the design service, generating corresponding character drawings according to the text description of the character image in the game, generating corresponding product drawings according to the description of the product in the e-commerce service, etc., which are not limited here.
- a graphical user interface is provided through a terminal device, and the content displayed by the graphical user interface at least partially includes an image generation scene. If a user performs a first control operation on the graphical user interface, the user inputs text information for describing the image content of at least one target object to be generated. If the user performs a second control operation on the graphical user interface, such as a submit operation, the position feature of at least one target object and the image feature of at least one target object can be generated based on the text information, thereby obtaining enhanced marking information.
- an image generation model can be used to perform multimodal image generation on multimodal prompt information to obtain a target image, and then the generated target image is displayed in the graphical user interface to provide feedback to the user, thereby achieving the purpose of accurately generating a high-quality target image including multiple target objects through multimodal prompt information, and being able to achieve more precise control over the generated image, generate an image with higher accuracy, and support the generation of images of multiple target objects, so that the image generation range is wider, thereby solving the technical problem in the related art that the corresponding image is generated only through text prompts, resulting in low accuracy of the generated image and a limited image generation range.
- Figure 5 is a flow chart of an image generation method according to Example 3 of the present disclosure. As shown in Figure 5, the method includes:
- Step S51 receiving a currently input multimodal dialogue request, wherein the information carried in the multimodal dialogue request includes: multimodal dialogue text information and multimodal dialogue enhancement mark information, the multimodal dialogue text information is used to describe the multimodal dialogue image content to be generated, the multimodal dialogue image content includes: at least one object, and the multimodal dialogue enhancement mark information is used to determine the position feature and image feature of the at least one object;
- Step S52 Use the image generation model to generate a multimodal image for the multimodal dialogue request to obtain a multimodal image.
- the image generation model is used to generate a multimodal dialogue image by adopting a multimodal image generation method;
- Step S53 Feedback a multimodal dialogue response, wherein the information carried in the multimodal dialogue response includes: a multimodal dialogue image.
- the multimodal conversation request can be understood as a conversation request initiated by a user to a computer or a robot, and the multimodal conversation request carries multimodal conversation text information and multimodal conversation enhancement mark information.
- the multimodal conversation text information can be understood as a text prompt, that is, the multimodal conversation image content to be generated described in natural language.
- the natural language can be Chinese, English, Japanese, etc., which is not limited here.
- the multimodal conversation image content described by the multimodal conversation text information may include at least one object, and the object may be understood as a person, animal or object in the multimodal conversation image content to be generated, that is, the present disclosure can support the generation of multi-object images.
- the object may include real people such as boys, girls, doctors, and teachers, virtual characters such as game players and non-player characters (NPC), animals such as cats, dogs, peacocks, and elephants, and objects such as tables, cars, grass, trees, and train stations, which are not limited here.
- the multimodal conversation text information may be text information corresponding to the multimodal conversation request input by the user. If the user wishes to obtain an image of a cat and a dog on the grass, the corresponding multimodal conversation text information may be "a cat and a dog on the grass (i.e., a cat and a dog on the grass)", and the multimodal conversation text information includes three objects, namely, a cat, a dog, and grass. It is understandable that the multimodal conversation text information may also be replaced by "a dog and a cat on the grass", or "there is a cat and a dog on the grass", and the present disclosure does not limit the description method of the multimodal conversation text information.
- the present disclosure uses multimodal conversation enhancement tag information as additional information, and generates a multimodal conversation image through the multimodal conversation text information and the multimodal conversation enhancement tag information, thereby improving the accuracy of multimodal conversation image generation.
- the multimodal dialogue enhancement tag information may include coordinate information and image information, wherein the coordinate information may be the coordinates of at least one object included in the multimodal dialogue text information, and is used to determine the positional features of the at least one object.
- the coordinate information includes the coordinates of the cat (e.g., [coordinates: 12, 15, 100, 200]), the coordinates of the dog (e.g., [coordinates: 100, 150, 220, 240]), and the coordinates of the grass (e.g., [coordinates: 500, 500, 500, 500]).
- the coordinates of at least one object may be coordinates in a two-dimensional rectangular coordinate system, which can represent the coordinate range of the target object, or may be coordinates in a two-dimensional coordinate system or a three-dimensional coordinate system.
- the present disclosure does not limit the adopted coordinate system.
- the image information may be an image of at least one object included in the multimodal conversation text information, and is used to determine the image features of at least one object in the multimodal conversation text information. For example, if the user wishes to obtain an image of a cat and a dog on the grass, the image information includes an image of a cat that the user needs to generate (e.g., [Figure [image: a user-uploaded image of their own cat]), an image of a dog (e.g. [image: a user-uploaded image of their own dog]), and an image of grass (e.g. [image: a user-uploaded image of grass]).
- the multimodal dialogue request in the embodiment of the present disclosure carries prompt information of multiple modes, including the above-mentioned multimodal dialogue text information and multimodal dialogue enhancement marking information, that is, the multimodal prompt information includes prompt information of text mode, coordinate mode and image mode.
- an augmented token is introduced in the present disclosure.
- the augmented token contains object-level text, coordinates, and image information at the same time, which is used to describe an object in the generated image.
- the image generation model is a model suitable for generating images.
- the image generation model can be a pre-trained image generation model based on a diffusion process, that is, a diffusion model.
- the image generation model can also be an autoregressive model, etc., which is not limited here.
- the image generation model can accurately generate high-quality multimodal dialogue images based on multimodal dialogue requests, that is, it can generate high-quality images based on multimodal prompts at the multi-object level, thereby achieving precise control over the generation of multimodal dialogue images.
- an image generation model is used to generate a multimodal image for the multimodal conversation request in the above example, that is, the multimodal conversation text information: a cat and a dog on the grass, coordinate information: the coordinates of the cat, the coordinates of the dog and the coordinates of the grass, and image information: an image of a cat, an image of a dog and an image of the grass are input into the image generation model, then the image generation model of the present invention can accurately output the corresponding multimodal conversation image according to the multimodal conversation request, that is, an image of a cat and a dog on the grass.
- the image generation model in the embodiment of the present disclosure can maintain the original structure of the pre-trained text-to-image model as much as possible so as to integrate the extension of the existing model, and the image generation model in the embodiment of the present disclosure only changes the input of the model and does not need to change the architecture of the model, so the usability of the technology built on the basis of the model can be maintained.
- the multimodal dialogue reply can be understood as the reply content (response) fed back to the user by the computer or robot based on the multimodal dialogue request input by the user, corresponding to the multimodal dialogue request input by the user.
- the multimodal dialogue reply carries a multimodal dialogue image, which is an image including the multimodal dialogue image content to be generated described in the multimodal dialogue request.
- FIG6 is a schematic diagram of a human-computer dialogue scenario according to Embodiment 3 of the present disclosure, in which a user inputs a multimodal dialogue request to a computer or robot, and the computer or robot can feedback a multimodal dialogue response corresponding to the multimodal dialogue request to the user.
- the dialogue between the user and the computer or robot can be achieved through voice recognition and natural language processing technology, or through text communication, which is not limited here.
- the method comprises the following steps: determining the position feature and image feature of at least one object in the multimodal conversation image content to be generated according to the multimodal conversation enhancement mark information, and then using the image generation model to generate a multimodal image for the multimodal conversation request, thereby obtaining a multimodal conversation image, and then feeding back a multimodal conversation reply carrying the multimodal conversation image to the user.
- the purpose of accurately generating a high-quality multimodal conversation image including multiple objects through a multimodal conversation request is achieved, and the generated multimodal conversation image can be more accurately controlled, a multimodal conversation image with higher accuracy is generated, and the generation of a multimodal conversation image of multiple objects is supported, so that the generation range of the multimodal conversation image is wider.
- image generation method can also be applied to, but is not limited to, application scenarios involving image generation in the fields of script services, design services, game services, e-commerce services, educational services, legal services, medical services, conference services, social network services, financial product services, logistics services and navigation services.
- design services design drawings are generated according to the user's text, image, and coordinate descriptions
- corresponding character drawings are generated according to the text descriptions of the character images
- corresponding product drawings are generated according to the descriptions of the products, etc., which are not limited here.
- the multimodal conversation image content to be generated can be determined according to the multimodal conversation text information, and the position feature and image feature of at least one object in the multimodal conversation image content to be generated can be determined according to the multimodal conversation enhancement mark information, and then the image generation model is used to generate a multimodal image for the multimodal conversation request, thereby obtaining a multimodal conversation image, and then feeding back a multimodal conversation reply carrying the multimodal conversation image to the user.
- the purpose of accurately generating a high-quality multimodal conversation image including multiple objects through a multimodal conversation request is achieved, and more precise control can be achieved on the generated multimodal conversation image, and a multimodal conversation image with higher accuracy can be generated, and the generation of multimodal conversation images of multiple objects can be supported, so that the generation range of multimodal conversation images is wider, thereby solving the technical problem in the related art that only corresponding images are generated through text prompts, resulting in low accuracy of the generated images and a limited image generation range.
- step S51 receiving a currently input multimodal dialogue request includes the following method steps:
- Step S511 obtaining multimodal dialogue text information
- Step S512 generating a position feature of at least one object and an image feature of at least one object based on the multimodal conversation text information, to obtain multimodal conversation enhanced tag information;
- Step S513 Combine the multimodal dialogue text information and the multimodal dialogue enhancement mark information to obtain a multimodal dialogue request.
- the multimodal dialogue text information when receiving the currently input multimodal dialogue request, can be first obtained, and then the position feature of at least one object and the image feature of at least one object are respectively generated based on the multimodal dialogue text information, that is, the coordinates and images of the object level are generated according to the multimodal dialogue text prompt, so as to obtain the corresponding multimodal dialogue enhancement mark information, and then the multimodal dialogue text information and the multimodal dialogue enhancement mark are combined.
- a multimodal dialogue request can be obtained.
- FIG7 is a structural schematic diagram of an image generation device according to Embodiment 4 of the present disclosure. As shown in FIG7 , the device includes:
- the acquisition module 701 is configured to acquire multimodal prompt information, wherein the multimodal prompt information includes: text information and enhanced mark information, the text information is used to describe the image content to be generated, the image content includes: at least one target object, and the enhanced mark information is used to determine the position feature and image feature of the at least one target object;
- the image generation module 702 is configured to use an image generation model to generate a multimodal image for the multimodal prompt information to obtain a target image, wherein the image generation model is used to generate the target image in a multimodal image generation manner.
- the above-mentioned acquisition module 701 is also configured to: acquire text information; generate at least one target object's position feature and at least one target object's image feature based on the text information to obtain enhanced marking information; and combine the text information and the enhanced marking information to obtain multimodal prompt information.
- the acquisition module 701 is further configured to: perform part-of-speech analysis on the text information, select at least one target participle, wherein the at least one target participle satisfies a preset part-of-speech requirement, and the at least one target participle is used to determine at least one target object; generate a position feature of the at least one target object based on the text information and the at least one target participle, and generate an image feature of the at least one target object based on the text information and the position feature of the at least one target object; and determine enhanced marking information using the at least one target participle, the position feature of the at least one target object, and the image feature of the at least one target object.
- the acquisition module 701 is further configured to: use a position feature generation model to generate position features for text information and at least one target word segmentation to obtain a position feature of at least one target object.
- the acquisition module 701 is further configured to: use an image feature generation model to generate image features for the text information and the position features of at least one target object to obtain image features of the at least one target object.
- an encoding module which is configured to: use a text encoder to perform text encoding on text information to obtain global text features, and use a text encoder to perform text encoding on at least one target word to obtain object text features.
- the acquisition module 701 is further configured to: combine the object text feature, the position feature of at least one target object, and the image feature of at least one target object to obtain enhanced marking information.
- the acquisition module 701 is further configured to: combine global text features, object text features, position features of at least one target object, and image features of at least one target object to obtain multimodal prompt information.
- the acquisition module 701 is further configured to: acquire text information and additional information, wherein the additional information
- the information includes at least one of the following: location information of at least one target object, image information of at least one target object; determining enhanced marking information based on text information and additional information; combining the text information and the enhanced marking information to obtain multimodal prompt information.
- multimodal prompt information including text information and enhanced mark information
- the acquisition module 701 and the image generation module 702 correspond to step S21 and step S22 in Example 1, and the examples and application scenarios implemented by the two modules and the corresponding steps are the same, but are not limited to the contents disclosed in the above-mentioned Example 1.
- the above-mentioned modules or units may be hardware components or software components stored in a memory and processed by one or more processors, and the above-mentioned modules may also be run in the server 10 provided in Example 1.
- FIG8 is a schematic diagram of the structure of another image generation device according to Embodiment 4 of the present disclosure, wherein a graphical user interface is provided by a terminal device, and the content displayed by the graphical user interface at least partially includes an image generation scene.
- the device includes:
- the first response module 801 is configured to respond to a first control operation performed on the graphical user interface and input text information, wherein the text information is used to describe the image content to be generated, and the image content includes: at least one target object;
- the second response module 802 is configured to respond to a second control operation performed on the graphical user interface, generate at least one position feature of the target object and at least one image feature of the target object based on the text information to obtain enhanced mark information, and use the image generation model to perform multimodal image generation on the multimodal prompt information to obtain a target image, wherein the enhanced mark information is used to determine the position feature and the image feature of the at least one target object, the image generation model is used to generate the target image using a multimodal image generation method, and the multimodal prompt information includes: text information and enhanced mark information;
- the display module 803 is configured to display the target image in the graphical user interface.
- a graphical user interface is provided through a terminal device, and the content displayed by the graphical user interface at least partially includes an image generation scene. If a user performs a first control operation on the graphical user interface, the user enters text information for describing the image content of at least one target object to be generated. If the user performs a first control operation on the graphical user interface, the user enters text information for describing the image content of at least one target object to be generated.
- a second control operation such as a submit operation, it can generate a position feature of at least one target object and an image feature of at least one target object based on the text information, thereby obtaining enhanced marking information.
- an image generation model to generate a multimodal image for the multimodal prompt information to obtain a target image, and then display the generated target image in the graphical user interface to provide feedback to the user, thereby achieving the purpose of accurately generating a high-quality target image including multiple target objects through multimodal prompt information, and can achieve more precise control over the generated image, generate an image with higher accuracy, and support the generation of images of multiple target objects, so that the image generation range is wider, thereby solving the technical problem in the related art that the corresponding image is generated only through text prompts, resulting in low accuracy of the generated image and a limited image generation range.
- first response module 801, the second response module 802 and the display module 803 correspond to steps S41 to S43 in Example 2, and the three modules and the corresponding steps implement the same examples and application scenarios, but are not limited to the contents disclosed in the above-mentioned Example 2.
- the above-mentioned modules or units may be hardware components or software components stored in a memory and processed by one or more processors, and the above-mentioned modules may also run in the server 10 provided in Example 1.
- FIG9 is a structural schematic diagram of another image generation device according to embodiment 4 of the present disclosure. As shown in FIG9 , the device includes:
- the receiving module 901 is configured to receive a currently input multimodal dialogue request, wherein the information carried in the multimodal dialogue request includes: multimodal dialogue text information and multimodal dialogue enhancement mark information, the multimodal dialogue text information is used to describe the multimodal dialogue image content to be generated, the multimodal dialogue image content includes: at least one object, and the multimodal dialogue enhancement mark information is used to determine the position feature and image feature of the at least one object;
- a generating module 902 is configured to generate a multimodal image for the multimodal dialogue request using an image generation model to obtain a multimodal dialogue image, wherein the image generation model is used to generate the multimodal dialogue image using a multimodal image generation method;
- the feedback module 903 is configured to feedback the multimodal dialogue response, wherein the information carried in the multimodal dialogue response includes: the multimodal dialogue image.
- the above-mentioned receiving module 901 is also configured to: obtain multimodal conversation text information; generate position features of at least one object and image features of at least one object based on the multimodal conversation text information, and obtain multimodal conversation enhancement marking information; combine the multimodal conversation text information and the multimodal conversation enhancement marking information to obtain a multimodal conversation request.
- a multimodal conversation request input by a user including multimodal conversation text information and multimodal conversation enhancement mark information it is possible to determine the multimodal conversation image content to be generated according to the multimodal conversation text information, determine the position feature and image feature of at least one object in the multimodal conversation image content to be generated according to the multimodal conversation enhancement mark information, and then use the image generation model to generate a multimodal image for the multimodal conversation request, thereby obtaining a multimodal conversation image, and then feedback to the user the image carrying the multimodal conversation image.
- Multimodal conversation reply by receiving a multimodal conversation request input by a user including multimodal conversation text information and multimodal conversation enhancement mark information
- the purpose of accurately generating high-quality multimodal conversation images including multiple objects through multimodal conversation requests is achieved, and more precise control can be achieved on the generated multimodal conversation images, generating multimodal conversation images with higher accuracy, and supporting the generation of multimodal conversation images of multiple objects, so that the generation range of multimodal conversation images is wider, thereby solving the technical problem in the related art that the corresponding images are generated only through text prompts, resulting in low accuracy of the generated images and a limited range of image generation.
- the receiving module 901, the generating module 902 and the feedback module 903 correspond to steps S51 to S53 in Example 3, and the three modules and the corresponding steps implement the same examples and application scenarios, but are not limited to the contents disclosed in the above-mentioned Example 3.
- the above-mentioned modules or units may be hardware components or software components stored in a memory and processed by one or more processors, and the above-mentioned modules may also be run in the server 10 provided in Example 1.
- the embodiment of the present disclosure may provide a computer terminal, which may be any computer terminal device in a computer terminal group.
- the computer terminal may also be replaced by a terminal device such as a mobile terminal.
- the computer terminal may be located in at least one network device among a plurality of network devices of the computer network.
- the above-mentioned computer terminal can execute the program code of the following steps in the image generation method: obtaining multimodal prompt information, wherein the multimodal prompt information includes: text information and enhanced mark information, the text information is used to describe the image content to be generated, and the image content includes: at least one target object, and the enhanced mark information is used to determine the position characteristics and image characteristics of at least one target object; using an image generation model to perform multimodal image generation on the multimodal prompt information to obtain a target image, wherein the image generation model is used to generate the target image using a multimodal image generation method.
- Figure 10 is a structural block diagram of a computer terminal according to Embodiment 5 of the present disclosure.
- the computer terminal A may include: one or more (only one is shown in the figure) processors 1002, a memory 1004, a storage controller, and a peripheral interface, wherein the peripheral interface is connected to a radio frequency module, an audio module, and a display.
- the memory may be configured to store software programs and modules, such as program instructions/modules corresponding to the image generation method and device in the embodiment of the present disclosure.
- the processor executes various functional applications and data processing by running the stored software programs and modules, thereby realizing the above-mentioned image generation method.
- the memory may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory.
- the memory may further include a memory remotely located relative to the processor, and these remote memories may be connected to the computer terminal A via a network. Examples of the above-mentioned network Including but not limited to the Internet, corporate intranet, local area network, mobile communication network and their combinations.
- the processor can call the information and application stored in the memory through the transmission device to perform the following steps: obtain multimodal prompt information, wherein the multimodal prompt information includes: text information and enhanced mark information, the text information is used to describe the image content to be generated, and the image content includes: at least one target object, and the enhanced mark information is used to determine the position characteristics and image characteristics of at least one target object; use the image generation model to generate a multimodal image for the multimodal prompt information to obtain a target image, wherein the image generation model is used to generate the target image using a multimodal image generation method.
- the processor may also execute program code for the following steps: obtaining text information; generating, based on the text information, a location feature of at least one target object and an image feature of at least one target object to obtain enhanced marking information; and combining the text information and the enhanced marking information to obtain multimodal prompt information.
- the processor may also execute program code for the following steps: performing part-of-speech analysis on the text information, selecting at least one target participle, wherein the at least one target participle satisfies a preset part-of-speech requirement, and the at least one target participle is used to determine at least one target object; generating a position feature of the at least one target object based on the text information and the at least one target participle, and generating an image feature of the at least one target object based on the text information and the position feature of the at least one target object; and determining enhanced tag information using the at least one target participle, the position feature of the at least one target object, and the image feature of the at least one target object.
- the processor may also execute program code of the following steps: using a position feature generation model to generate position features for text information and at least one target word segmentation to obtain a position feature of at least one target object.
- the processor may also execute program code of the following steps: using an image feature generation model to generate image features for text information and position features of at least one target object to obtain image features of at least one target object.
- the processor may also execute program code of the following steps: using a text encoder to perform text encoding on text information to obtain global text features, and using a text encoder to perform text encoding on at least one target word segmentation to obtain object text features.
- the processor may further execute program code of the following steps: combining object text features, position features of at least one target object, and image features of at least one target object to obtain enhanced marking information.
- the processor may also execute program code of the following steps: combining global text features, object text features, position features of at least one target object, and image features of at least one target object to obtain multimodal prompt information.
- the processor may also execute program code for the following steps: obtaining text information and additional information, wherein the additional information includes at least one of the following: location information of at least one target object, image information of at least one target object; determining enhanced marking information based on the text information and the additional information; and combining the text information and the enhanced marking information to obtain multimodal prompt information.
- additional information includes at least one of the following: location information of at least one target object, image information of at least one target object; determining enhanced marking information based on the text information and the additional information; and combining the text information and the enhanced marking information to obtain multimodal prompt information.
- multimodal prompt information including text information and enhanced mark information
- the structure shown in FIG. 10 is for illustration only, and the computer terminal A may also be a terminal device such as a smart phone (such as an Android phone, an iOS phone, etc.), a tablet computer, a PDA, and a mobile Internet device (Mobile Internet Devices, MID), PAD, etc.
- FIG. 10 does not limit the structure of the above-mentioned electronic device.
- the computer terminal A may also include more or fewer components (such as a network interface, a display device, etc.) than those shown in FIG. 10, or have a configuration different from that shown in FIG. 10.
- a person of ordinary skill in the art may understand that all or part of the steps in the various methods of the above embodiments may be completed by instructing the hardware related to the terminal device through a program, and the program may be stored in a computer-readable storage medium, and the storage medium may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.
- the embodiment of the present disclosure further provides a computer-readable storage medium.
- the computer-readable storage medium can be used to store the program code executed by the image generation method provided in the first embodiment.
- the computer-readable storage medium may be located in any one of the computer terminals in a computer terminal group in a computer network, or in any one of the mobile terminals in a mobile terminal group.
- the computer-readable storage medium is configured to store program code for executing the following steps: obtaining multimodal prompt information, wherein the multimodal prompt information includes: text information and enhanced mark information, the text information is used to describe the image content to be generated, and the image content includes: at least one target object, and the enhanced mark information is used to determine the position characteristics and image characteristics of at least one target object; using an image generation model to perform multimodal image generation on the multimodal prompt information to obtain a target image, wherein the image generation model is used to generate the target image using a multimodal image generation method.
- the computer-readable storage medium is configured to store program code for performing the following steps: obtaining text information; generating a position feature of at least one target object and an image feature of at least one target object based on the text information to obtain enhanced marking information; and combining the text information and the enhanced marking information to obtain multimodal prompt information.
- the computer-readable storage medium is configured to store program code for executing the following steps: performing part-of-speech analysis on text information, selecting at least one target participle, wherein the at least one target participle satisfies a preset part-of-speech requirement, and the at least one target participle is used to determine at least one target object; generating position features of at least one target object based on the text information and the at least one target participle, and generating image features of at least one target object based on the text information and the position features of the at least one target object; and determining enhanced tag information using the at least one target participle, the position features of the at least one target object, and the image features of the at least one target object.
- the computer-readable storage medium is configured to store program code for executing the following steps: using a position feature generation model to generate position features for text information and at least one target word segmentation to obtain position features of at least one target object.
- the computer-readable storage medium is configured to store program code for executing the following steps: using an image feature generation model to generate image features for text information and position features of at least one target object to obtain image features of at least one target object.
- the computer-readable storage medium is configured to store program code for performing the following steps: using a text encoder to perform text encoding on text information to obtain global text features, and using a text encoder to perform text encoding on at least one target word to obtain object text features.
- the computer-readable storage medium is configured to store program code for executing the following steps: combining object text features, position features of at least one target object, and image features of at least one target object to obtain enhanced marking information.
- the computer-readable storage medium is configured to store program code for performing the following steps: combining global text features, object text features, position features of at least one target object, and image features of at least one target object to obtain multimodal prompt information.
- the computer-readable storage medium is configured to store program code for performing the following steps: obtaining text information and additional information, wherein the additional information includes at least one of the following: location information of at least one target object, image information of at least one target object; determining enhanced mark information based on the text information and the additional information; and combining the text information and the enhanced mark information to obtain multimodal prompt information.
- the disclosed technical content can be implemented in other ways.
- the device embodiments described above are only schematic.
- the division of the units is only a logical function division.
- multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed.
- Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.
- the units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
- each functional unit in each embodiment of the present disclosure may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
- the above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
- the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium.
- the computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, server or network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present disclosure.
- the aforementioned storage medium includes: U disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), mobile hard disk, magnetic disk or optical disk and other media that can store program codes.
Landscapes
- Engineering & Computer Science (AREA)
- Theoretical Computer Science (AREA)
- Physics & Mathematics (AREA)
- General Physics & Mathematics (AREA)
- Artificial Intelligence (AREA)
- General Engineering & Computer Science (AREA)
- Health & Medical Sciences (AREA)
- General Health & Medical Sciences (AREA)
- Software Systems (AREA)
- Data Mining & Analysis (AREA)
- Evolutionary Computation (AREA)
- Computational Linguistics (AREA)
- Computer Vision & Pattern Recognition (AREA)
- Computing Systems (AREA)
- Life Sciences & Earth Sciences (AREA)
- Medical Informatics (AREA)
- Audiology, Speech & Language Pathology (AREA)
- Mathematical Physics (AREA)
- Multimedia (AREA)
- Databases & Information Systems (AREA)
- Bioinformatics & Cheminformatics (AREA)
- Bioinformatics & Computational Biology (AREA)
- Evolutionary Biology (AREA)
- Biomedical Technology (AREA)
- Biophysics (AREA)
- Molecular Biology (AREA)
- Processing Or Creating Images (AREA)
Abstract
本公开公开了一种图像生成方法、电子设备及计算机可读存储介质,涉及大模型技术、图像处理技术领域。该方法包括:获取多模态提示信息,其中,多模态提示信息包括:文本信息与增强标记信息,文本信息用于描述待生成的图像内容,图像内容包括:至少一个目标对象,增强标记信息用于确定至少一个目标对象的位置特征与图像特征;采用图像生成模型对多模态提示信息进行多模态图像生成,得到目标图像,其中,图像生成模型用于采用多模态图像生成方式生成目标图像。本公开解决了相关技术中仅通过文本提示生成对应图像,导致生成的图像精确度较低,图像生成范围较局限的技术问题。
Description
本公开涉及大模型技术、图像处理技术领域,具体而言,涉及一种图像生成方法、电子设备及计算机可读存储介质。
随着人工智能的发展,文本到图像生成领域取得了重大进展,能够基于文本提示生成对应的高质量图像,从而实现自动生成图像的功能。
目前,仅通过文本提示描述物体从而生成对应的图像,由于文本提示描述物理较为困难,因此生成的图像的精确度较低,并且该方法通常需要微调或仅支持使用单个物体作为约束条件,因此图像生成范围较为局限。
针对上述的问题,目前尚未提出有效的解决方案。
发明内容
本公开实施例提供了一种图像生成方法、电子设备及计算机可读存储介质,以至少解决相关技术中仅通过文本提示生成对应图像,导致生成的图像精确度较低,图像生成范围较局限的技术问题。
根据本公开实施例的一个方面,提供了一种图像生成方法,包括:获取多模态提示信息,其中,多模态提示信息包括:文本信息与增强标记信息,文本信息用于描述待生成的图像内容,图像内容包括:至少一个目标对象,增强标记信息用于确定至少一个目标对象的位置特征与图像特征;采用图像生成模型对多模态提示信息进行多模态图像生成,得到目标图像,其中,图像生成模型用于采用多模态图像生成方式生成目标图像。
根据本公开实施例的另一方面,还提供了另一种图像生成方法,通过终端设备提供一图形用户界面,图形用户界面所显示的内容至少部分地包含一图像生成场景,包括:响应对图形用户界面执行的第一控制操作,输入文本信息,其中,文本信息用于描述待生成的图像内容,图像内容包括:至少一个目标对象;响应对图形用户界面执行的第二控制操作,基于文本信息分别生成至少一个目标对象的位置特征与至少一个目标对象的图像特征以得到增强标记信息,以及采用图像生成模型对多模态提示信息进行多模态图像生成以得到目标图像,其中,增强标记信息用于确定至少一个目标对象的位置特征与图像特征,图像生成模型用于采用多模态图像生成方式生成目标图像,多模态提示信息包括:文本信息与增强标记信息;在图形用户界面内展示目标图像。
根据本公开实施例的另一方面,还提供了另一种图像生成方法,包括:接收当前
输入的多模态对话请求,其中,多模态对话请求中携带的信息包括:多模态对话文本信息与多模态对话增强标记信息,多模态对话文本信息用于描述待生成的多模态对话图像内容,多模态对话图像内容包括:至少一个对象,多模态对话增强标记信息用于确定至少一个对象的位置特征与图像特征;采用图像生成模型对多模态对话请求进行多模态图像生成,得到多模态对话图像,其中,图像生成模型用于采用多模态图像生成方式生成多模态对话图像;反馈多模态对话回复,其中,多模态对话回复中携带的信息包括:多模态对话图像。
根据本公开实施例的另一方面,还提供了一种电子设备,包括:存储器,存储有可执行程序;处理器,用于运行程序,其中,程序运行时执行任意一项上述的图像生成方法。
根据本公开实施例的另一方面,还提供了一种计算机可读存储介质,计算机可读存储介质包括存储的可执行程序,其中,在可执行程序运行时控制计算机可读存储介质所在设备执行任意一项上述的图像生成方法。
在本公开实施例中,通过获取包括文本信息与增强标记信息的多模态提示信息,能够根据文本信息确定待生成的图像内容、根据增强标记信息确定图像内容中至少一个目标对象的位置特征和图像特征,然后采用图像生成模型对多模态提示信息进行多模态图像生成,从而得到目标图像,由此达到了通过多模态提示信息准确生成包括多目标对象的高质量目标图像的目的。此外,能够对生成的图像实现更精准的控制,生成准确度更高的图像,且支持生成多目标对象的图像,使得图像的生成范围更广泛,从而实现了更精确、更多样化、支持多目标对象的高质量图像生成,以满足不同领域的图像生成需求的技术效果,进而解决了相关技术中仅通过文本提示生成对应图像,导致生成的图像精确度较低,图像生成范围较局限的技术问题。
容易注意到的是,上面的通用描述和后面的详细描述仅仅是为了对本公开进行举例和解释,并不构成对本公开的限定。
此处所说明的附图用来提供对本公开的进一步理解,构成本公开的一部分,本公开的示意性实施例及其说明用于解释本公开,并不构成对本公开的不当限定。在附图中:
图1是根据本公开实施例1的一种图像生成方法的应用场景示意图;
图2是根据本公开实施例1的一种图像生成方法的流程图;
图3是根据本公开实施例1的另一种图像生成方法的流程图;
图4是根据本公开实施例2的一种图像生成方法的流程图;
图5是根据本公开实施例3的一种图像生成方法的流程图;
图6是根据本公开实施例3的一种人机对话场景示意图;
图7是根据本公开实施例4的一种图像生成装置的结构示意图;
图8是根据本公开实施例4的另一种图像生成装置的结构示意图;
图9是根据本公开实施例4的再一种图像生成装置的结构示意图;
图10是根据本公开实施例5的一种计算机终端的结构框图。
为了使本技术领域的人员更好地理解本公开方案,下面将结合本公开实施例中的附图,对本公开实施例中的技术方案进行清楚、完整地描述,显然,所描述的实施例仅仅是本公开一部分的实施例,而不是全部的实施例。基于本公开中的实施例,本领域普通技术人员在没有做出创造性劳动前提下所获得的所有其他实施例,都应当属于本公开保护的范围。
需要说明的是,本公开的说明书和权利要求书及上述附图中的术语“第一”、“第二”等是用于区别类似的对象,而不必用于描述特定的顺序或先后次序。应该理解这样使用的数据在适当情况下可以互换,以便这里描述的本公开的实施例能够以除了在这里图示或描述的那些以外的顺序实施。此外,术语“包括”和“具有”以及他们的任何变形,意图在于覆盖不排他的包含,例如,包含了一系列步骤或单元的过程、方法、系统、产品或设备不必限于清楚地列出的那些步骤或单元,而是可包括没有清楚地列出的或对于这些过程、方法、产品或设备固有的其它步骤或单元。
本公开提供的技术方案主要采用大模型技术实现,此处的大模型是指具有大规模模型参数的深度学习模型,通常可以包含上亿、上百亿、上千亿、上万亿甚至十万亿以上的模型参数。大模型又可以称为基石模型/基础模型(Foundation Model),通过大规模无标注的语料进行大模型的预训练,产出亿级以上参数的预训练模型,这种模型能适应广泛的下游任务,模型具有较好的泛化能力,例如大规模语言模型(Large Language Model,简称LLM)、多模态预训练模型(multi-modal pre-training model)等。
需要说明的是,大模型在实际应用时,可以通过少量样本对预训练模型进行微调,使得大模型可以应用于不同的任务中。例如,大模型可以广泛应用于自然语言处理(Natural Language Processing,简称NLP)、计算机视觉、语音处理等领域,具体可以应用于如视觉问答(Visual Question Answering,简称VQA)、图像描述(Image Caption,简称IC)、图像生成等计算机视觉领域任务,也可以广泛应用于基于文本的情感分类、文本摘要生成、机器翻译等自然语言处理领域任务。因此,大模型主要的应用场景包括但不限于数字助理、智能机器人、搜索、在线教育、办公软件、电子商务、智能设计等。
首先,在对本公开实施例进行描述的过程中出现的部分名词或术语适用于如下解释:
扩散模型(Diffusion Model):一种用于描述扩散过程的数学模型,是常用的生成
模型。本公开实施例中,扩散模型能够基于扩散过程将带噪声图片进行去噪,从而得到生成图片。
组成树(constituency tree):一种用于表示自然语言句子结构的树状结构,由节点和边组成,其中节点代表句子中的词语或短语,边表示这些词语或短语之间的语法关系。在组成树中,句子被分解为不同的短语结构,如名词短语、动词短语、形容词短语等。这些短语结构可以进一步分解为更小的短语结构,直到最小的单词级别。组成树可以帮助理解句子的结构和语法关系,有助于进行句法分析和语义分析,还可以用于自然语言处理任务,如句法分析、机器翻译和信息抽取等。通过分析组成树,可以更好地理解句子的语法结构和语义含义。
自然语言处理(Natural Language Processing,NLP)解析器:一种用于自然语言处理的工具,可以将自然语言文本转换成结构化的数据,并进行语法分析、词性标注、命名实体识别等任务。NLP解析器通常基于机器学习和深度学习技术,能够自动分析文本并提取其中的信息,可以帮助理解和处理大量的自然语言数据,例如语音识别、情感分析、文本分类等。
文本编码器(Text Encoder):一种用于将文本数据转换为数值表示的模型或工具。在自然语言处理(NLP)领域中,文本编码器通常用于将文本数据转化为向量或矩阵形式,以便进行机器学习和深度学习任务,如文本分类、语义相似度计算、情感分析等。
微调(Fine-tune):是指在使用预训练模型的基础上,通过在特定任务上进行进一步的训练和调整,以提高模型在该任务上的性能和适应能力。通常情况下,预训练模型是在大规模数据集上进行训练的,而Fine-tune则是在特定的小规模数据集上进行微调,以使模型更好地适应具体的任务需求。Fine-tune通常用于自然语言处理、计算机视觉等领域的深度学习模型中。
在相关技术的文本到图像生成方法中,文本到图像模型通常使用从文本提示词中提取的词嵌入作为条件。然而,文本具有较高的抽象水平和有限的信息密度,使得准确描述物体变得困难,因此生成的图像的精确度较低,并且该方法通常需要微调或仅支持使用单个物体作为约束条件,因此图像生成范围较为局限。
目前,提出了BLIP-Diffusion模型,能够支持图文混合生成对应图像,但是仅支持单个物体的生成,且生成的图像的精确度较低。此外,还提出了KOSMOS-G模型,能够支持图文混合生成对应图像,但生成的图像的精确度仍然较低。
相关技术基于文本提示生成对应图像的方法存在如下缺陷。
缺陷1:仅通过文本提示描述物体从而生成对应的图像,由于文本提示描述物理较为困难,因此生成的图像的精确度较低。
缺陷2:通常需要微调或仅支持使用单个物体作为约束条件,因此模型训练成本
较高且图像生成范围较为局限;
缺陷3:BLIP-Diffusion模型和KOSMOS-G模型支持图文混合生成,但是生成图像的精确度较低,且BLIP-Diffusion模型仅支持单物体生成。
针对上述缺陷,在本公开之前尚未提出有效的解决方案。
实施例1
根据本公开实施例,提供了一种图像生成方法,需要说明的是,在附图的流程图示出的步骤可以在诸如一组计算机可执行指令的计算机系统中执行,并且,虽然在流程图中示出了逻辑顺序,但是在某些情况下,可以以不同于此处的顺序执行所示出或描述的步骤。
考虑到大模型的模型参数量庞大,且移动终端的运算资源有限,本公开实施例提供的上述图像生成方法可以应用于如图1所示的应用场景,但不仅限于此。在如图1所示的应用场景中,大模型部署在服务器10中,服务器10可以通过局域网连接、广域网连接、因特网连接,或者其他类型的数据网络,连接一个或多个客户端设备20,此处的客户端设备20可以包括但不限于:智能手机、平板电脑、笔记本电脑、掌上电脑、个人计算机、智能家居设备、车载设备等。客户端设备20可以通过图形用户界面与用户进行交互,实现对大模型的调用,进而实现本公开实施例所提供的方法。
在本公开实施例中,客户端设备和服务器构成的系统可以执行如下步骤:客户端设备执行获取用户在图形用户界面内输入的多模态提示信息,并将多模态提示信息发送给服务器等步骤,服务器执行采用图像生成模型对获取到的多模态提示信息进行多模态图像生成,从而得到目标图像,并将目标图像返回至客户端设备等步骤,需要说明的是,在客户端设备的运行资源能够满足大模型的部署和运行条件的情况下,本公开实施例可以在客户端设备中进行。
在上述运行环境下,本公开提供了如图2所示的图像生成方法。图2是根据本公开实施例1的一种图像生成方法的流程图。如图2所示,该方法可以包括如下步骤:
步骤S21,获取多模态提示信息,其中,多模态提示信息包括:文本信息与增强标记信息,文本信息用于描述待生成的图像内容,图像内容包括:至少一个目标对象,增强标记信息用于确定至少一个目标对象的位置特征与图像特征;
步骤S22,采用图像生成模型对多模态提示信息进行多模态图像生成,得到目标图像,其中,图像生成模型用于采用多模态图像生成方式生成目标图像。
文本信息可以理解为文本提示,即使用自然语言文字描述的图像内容,用于描述待生成的图像内容。示例性地,自然语言文字可以为中文、英文、日文等,此处不予限制。本公开实施例中,文本信息所描述的图像内容中可以包括至少一个目标对象,目标对象可以理解为待生成的图像内容中的人物、动物或物体等,即本公开能够支持多对象的图像生成。示例性地,目标对象可以包括男孩、女孩、医生、老师等真实人
物,可以包括游戏玩家、非玩家角色(Non-Player Character,NPC)等虚拟人物,可以包括猫、狗、孔雀、大象等动物,还可以包括桌子、汽车、草地、树木、火车站等物体,此处不予限制。
示例性地,文本信息可以为用户的图像生成需求对应的文本信息,若用户希望得到一只猫和一只狗在草地上的图像,则对应的文本信息可以为“一只猫和一只狗在草地上(即a cat and a dog on the grass)”,该文本信息中即包括三个目标对象,分别为猫、狗和草地。可以理解的是,该文本信息还可以替换为“一只狗和一只猫在草地上”,或者“草地上有一只猫和一只狗”,本公开对文本信息的描述方式不予限制。
考虑到仅通过文本提示描述物体会导致图像生成不准确,因此本公开在提供文本信息的基础上,将增强标记信息作为附加信息,通过文本信息和增强标记信息共同生成目标图像,从而提高图像生成的准确率。
增强标记信息可以包括坐标信息和图像信息,其中,坐标信息可以为文本信息中包括的至少一个目标对象的坐标,用于确定至少一个目标对象的位置特征。示例性地,若用户希望得到一只猫和一只狗在草地上的图像,则坐标信息即包括猫的坐标(例如[坐标:12,15,100,200])、狗的坐标(例如[坐标:100,150,220,240])和草地的坐标(例如[坐标:500,500,500,500])。
示例性地,至少一个目标对象的坐标可以是二维直角坐标系下的坐标,能够表示目标对象的坐标范围,也可以是二维坐标系或三维坐标系下的坐标,本公开对所采用的坐标系不予限制。
图像信息可以为文本信息中包括的至少一个目标对象的图像,用于确定文本信息中至少一个目标对象的图像特征。示例性地,若用户希望得到一只猫和一只狗在草地上的图像,则图像信息即包括用户所需要生成的一只猫的图像(例如[图像:一张用户上传的自己的猫的图像])、一只狗的图像(例如[图像:一张用户上传的自己的狗的图像])以及草地的图像(例如[图像:一张用户上传的草地的图像])。
多模态提示信息可以理解为多种模态的提示信息,包括上述文本信息和增强标记信息。本公开实施例中,多模态提示信息包括文本模态、坐标模态以及图像模态的提示信息。
示例性地,可以将同一目标对象的多模态提示信息绑定在一起,以避免影响其他物体或全局条件,也即本公开中引入增强标记(augmented token),增强标记同时包含物体级的文本、坐标和图像信息,用于描述生成图像中的一个物体。
本公开实施例中,图像生成模型为适用于生成图像的模型,示例性地,图像生成模型可以为预训练的基于扩散过程的图像生成模型,即为扩散模型,图像生成模型还可以为自回归模型等,此处不予限制。图像生成模型能够基于多目标对象的多模态提示信息准确生成高质量目标图像,即能够基于多物体级别的多模态提示生成高质量图
像,从而实现对生成图像的精准控制。
示例性地,若将上述示例的多模态提示信息输入至图像生成模型,即输入文本信息:一只猫和一只狗在草地上、坐标信息:猫的坐标、狗的坐标和草地的坐标、图像信息:一只猫的图像、一只狗的图像以及草地的图像,则本公开的图像生成模型能够根据该多模态提示信息准确输出对应的目标图像,即一只猫和一只狗在草地上的图像。
可以理解的是,本公开实施例中的图像生成模型能够尽可能保持预训练的文本到图像模型的原始结构,以便集成对现有模型的扩展,并且本公开实施例中的图像生成模型仅改变了模型的输入,不需要改变模型的架构,因此能够保持在模型的基础上构建技术的可用性。
本公开实施例中,通过获取包括文本信息与增强标记信息的多模态提示信息,能够根据文本信息确定待生成的图像内容、根据增强标记信息确定图像内容中至少一个目标对象的位置特征和图像特征,然后采用图像生成模型对多模态提示信息进行多模态图像生成,从而得到目标图像,达到了通过多模态提示信息准确生成包括多目标对象的高质量目标图像的目的,且能够对生成的图像实现更精准的控制,生成准确度更高的图像,且支持生成多目标对象的图像,使得图像的生成范围更广泛。
本公开实施例提供的上述图像生成方法可以但不限于应用于剧本服务、设计服务、游戏服务、电商服务、教育服务、法律服务、医疗服务、会议服务、社交网络服务、金融产品服务、物流服务和导航服务等领域中涉及图像生成的应用场景中,例如:剧本服务中根据剧本或故事情节生成所需的图像素材、设计服务中根据用户的文字、图像、坐标描述生成设计图、游戏中根据角色形象的文字描述生成对应角色图、电商服务中根据商品的描述生成对应商品图等,此处不予限制。
采用本公开实施例,通过获取包括文本信息与增强标记信息的多模态提示信息,能够根据文本信息确定待生成的图像内容、根据增强标记信息确定图像内容中至少一个目标对象的位置特征和图像特征,然后采用图像生成模型对多模态提示信息进行多模态图像生成,从而得到目标图像,由此达到了通过多模态提示信息准确生成包括多目标对象的高质量目标图像的目的。此外,能够对生成的图像实现更精准的控制,生成准确度更高的图像,且支持生成多目标对象的图像,使得图像的生成范围更广泛,从而实现了更精确、更多样化、支持多目标对象的高质量图像生成,以满足不同领域的图像生成需求的技术效果,进而解决了相关技术中仅通过文本提示生成对应图像,导致生成的图像精确度较低,图像生成范围较局限的技术问题。
在一种可选的实施例中,在步骤S21中,获取多模态提示信息,包括如下方法步骤:
步骤S211,获取文本信息;
步骤S212,基于文本信息分别生成至少一个目标对象的位置特征与至少一个目标
对象的图像特征,得到增强标记信息;
步骤S213,对文本信息与增强标记信息进行组合,得到多模态提示信息。
考虑到在推理过程中,获取所有目标对象的多模态提示信息可能较为困难,例如可能仅获取到了目标对象的文本信息,或者仅能够获取到部分目标对象的多模态提示信息(例如当用户希望在图像中合并真实物体和生成的物体的情况)。因此本公开实施例中,可以基于文本信息获取增强标记信息,即能够生成目标对象缺失的增强标记信息,从而实现多个模态的灵活组合。
本公开实施例中,在获取多模态提示信息时,可以先获取文本信息,然后基于文本信息分别生成至少一个目标对象的位置特征与至少一个目标对象的图像特征,即根据文本提示生成物体级别的坐标和图像,从而得到对应的增强标记信息,再将文本信息与增强标记信息进行组合,即可得到多模态提示信息。
在一种可选的实施例中,在步骤S212中,基于文本信息分别生成至少一个目标对象的位置特征与图像特征,得到增强标记信息,包括如下方法步骤:
步骤S2121,对文本信息进行词性分析,选取至少一个目标分词,其中,至少一个目标分词满足预设词性要求,至少一个目标分词用于确定至少一个目标对象;
步骤S2122,基于文本信息和至少一个目标分词生成至少一个目标对象的位置特征,以及基于文本信息和至少一个目标对象的位置特征生成至少一个目标对象的图像特征;
步骤S2123,利用至少一个目标分词、至少一个目标对象的位置特征以及至少一个目标对象的图像特征确定增强标记信息。
可以理解的是,语言中的词性包括名词、动词、形容词、副词、代词、数词、量词、连词、介词、助词、叹词等,而通常名词用来表示人、事物、地点等,在生成图像时需要被生成出来,因此名词可以用来表示目标对象。
本公开实施例中,预设词性可以为名词,至少一个目标分词即为文本信息中的至少一个名词。通过对文本信息进行词性分析,选取文本信息中的为名词的至少一个目标分词,从而能够确定文本信息中所包含的至少一个目标对象,进而便于获取每个目标对象对应的多模态提示信息。
示例性地,可以使用组成树(constituency tree)来识别文本信息中的物体,即识别文本信息中的目标对象。例如文本信息为“一只猫和一只狗在草地上(即a cat and a dog on the grass)”,通过组成树对文本信息进行词性分析,分析得到“一只(限定词)猫(名词)和(连词)一只(限定词)狗(名词)在(介词)草地(名词)上(即a(determiner)cat(noun)and(conjunction)a(determiner)dog(noun)on(preposition)the(determiner)grass(noun))”,识别出文本信息中的名词,即猫、狗和草地,从而将识别得到的名词猫、狗和草地作为选取出的多个目标分词。
此外,还可以使用条件随机场(Conditional Random Fields,CRF)、最大熵模型(Maximum Entropy Model)等,通过训练模型来对文本进行词性标注,或者使用循环神经网络(Recurrent Neural Networks,RNN)、长短时记忆网络(Long Short-Term Memory,LSTM)、注意力机制(Attention Mechanism)等进行词性标注和词性识别,此处不予限制。
本公开实施例中,在基于文本信息分别生成至少一个目标对象的位置特征与图像特征,得到增强标记信息时,可以对文本信息进行词性分析,从文本信息中选取出词性为名词的至少一个目标分词。然后基于文本信息和至少一个目标分词生成至少一个目标对象的位置特征以及图像特征。再根据至少一个目标分词、至少一个目标对象的位置特征以及至少一个目标对象的图像特征确定至少一个目标对象分别对应的增强标记信息。
在一种可选的实施例中,在步骤S2122中,基于文本信息和至少一个目标分词生成至少一个目标对象的位置特征,包括如下方法步骤:
步骤S21221,采用位置特征生成模型对文本信息和至少一个目标分词进行位置特征生成,得到至少一个目标对象的位置特征。
本公开实施例中,位置特征生成模型为适用于生成位置特征的模型,例如:扩散模型、自回归模型等。位置特征生成模型可以为坐标模型,能够基于文本信息的文字内容和至少一个目标分词对应的物体,生成每个物体对应的位置坐标。
本公开实施例中,在基于文本信息和至少一个目标分词生成至少一个目标对象的位置特征时,可以将文本信息和至少一个目标分词输入至位置特征生成模型,采用位置特征生成模型对文本信息和至少一个目标分词进行位置特征生成,从而得到至少一个目标对象对应的位置特征,即得到至少一个目标对象对应的坐标信息。
示例性地,若文本信息为“一只猫和一只狗”,则多个目标分词为“猫”和“狗”,本公开通过将“一只猫和一只狗”、“猫”和“狗”这三个文本特征拼到一起,共同作为位置特征生成模型的输入,从而根据位置特征生成模型的输出得到猫的坐标和狗的坐标。
可以看出,本公开中当坐标模态信息缺失时,可以采用位置特征生成模型基于文本提示和多个物体自动生成多个物体对应的位置坐标,从而补全缺失的坐标模态信息。
在一种可选的实施例中,在步骤S2122中,基于文本信息和至少一个目标对象的位置特征生成至少一个目标对象的图像特征,包括如下方法步骤:
步骤S21222,采用图像特征生成模型对文本信息和至少一个目标对象的位置特征进行图像特征生成,得到至少一个目标对象的图像特征。
本公开实施例中,图像特征生成模型为适用于生成图像特征的模型,例如:扩散模型、自回归模型等。图像特征生成模型可以为图像特征模型,能够基于文本信息的文字内容和至少一个目标对象的位置特征,生成至少一个目标对象对应的图像特征。
本公开实施例中,在基于文本信息和至少一个目标对象的位置特征生成至少一个目标对象的图像特征时,可以将文本信息和至少一个目标对象的位置特征输入至图像特征生成模型,采用图像特征生成模型对文本信息和至少一个目标对象的坐标进行图像特征生成,从而得到至少一个目标对象对应的图像特征,即得到至少一个目标对象对应的图像信息。
示例性地,若文本信息为“一只猫和一只狗”,且已获取到了猫的坐标和狗的坐标,本公开通过将“一只猫和一只狗(文本信息)”、“猫(文本)+猫坐标(位置特征)”、“狗(文本)+狗坐标(位置特征)”输入至图像特征生成模型,从而根据图像特征生成模型的输出得到猫的图像特征和狗的图像特征。
可以看出,本公开中当图像模态信息缺失时,可以采用图像特征生成模型基于文本提示和多个物体的位置坐标自动生成多个物体对应的图像信息,从而补全缺失的图像模态信息。
也即可以看出,本公开实施例中的图像生成模型不仅能够支持根据文字提示、坐标信息和图像信息生成目标图像,同时也支持仅根据文字提示生成目标图像。也即本公开可以基于多物体多模态提示生成图片,同时,当出现模态缺失问题时,也可以利用生成模型补全缺失的模态,因此应用范围更为广泛。
在一种可选的实施例中,该图像生成方法还包括如下方法步骤:
步骤S2124,采用文本编码器对文本信息进行文本编码,得到全局文本特征,以及采用文本编码器对至少一个目标分词进行文本编码,得到对象文本特征。
考虑到文本编码器(Text Encoder)能够将文本数据转化为向量或矩阵形式,以便进行机器学习和深度学习任务,如词性标注等。因此本公开实施例中,在对文本信息进行处理时,可以采用文本编码器对文本信息进行文本编码处理,从而得到全局文本编码结果,即全局文本特征,同时也可以采用文本编码器对至少一个目标分词进行文本编码处理,从而得到名词部分编码结果,即对象文本特征。
可以理解的是,在对增强标记信息中的图像信息进行处理时,可以采用图像编码器(Image Encoder)对图像信息进行图像特征编码处理,从而得到图像特征,此处不过多赘述。
在一种可选的实施例中,在步骤S2123中,利用至少一个目标分词、至少一个目标对象的位置特征以及至少一个目标对象的图像特征确定增强标记信息,包括如下方法步骤:
步骤S21231,对对象文本特征、至少一个目标对象的位置特征以及至少一个目标对象的图像特征进行组合,得到增强标记信息。
本公开实施例中,在利用至少一个目标分词、至少一个目标对象的位置特征以及至少一个目标对象的图像特征确定增强标记信息时,可以将对至少一个目标分词进行
文本编码处理得到的对象文本特征、至少一个目标对象的位置特征以及至少一个目标对象的图像特征进行组合,从而得到增强标记信息。示例性地,可以将对象文本特征、至少一个目标对象的位置特征以及至少一个目标对象的图像特征进行水平拼接,从而得到增强标记信息,此处不予限制。
示例性地,若文本信息为“一只猫和一只狗在草地上”,则可以对对象文本特征“猫”、“狗”和“草地”、至少一个目标对象的位置特征“猫的坐标”、“狗的坐标”和“草地的坐标”以及至少一个目标对象的图像特征“一只猫的图像”、“一只狗的图像”和“草地的图像”进行特征组合,从而得到增强标记信息。
在一种可选的实施例中,在步骤S213中,对文本信息与增强标记信息进行组合,得到多模态提示信息,包括如下方法步骤:
步骤S2131,对全局文本特征、对象文本特征、至少一个目标对象的位置特征以及至少一个目标对象的图像特征进行组合,得到多模态提示信息。
本公开实施例中,在对文本信息与增强标记信息进行组合,得到多模态提示信息时,可以将全局文本特征、对象文本特征、至少一个目标对象的位置特征以及至少一个目标对象的图像特征进行组合,从而得到多模态提示信息。示例性地,可以将对象文本特征、至少一个目标对象的位置特征以及至少一个目标对象的图像特征进行水平拼接,再与全局文本特征进行垂直拼接,从而得到多模态提示信息,此处不予限制。
示例性地,若文本信息为“一只猫和一只狗在草地上”,则可以对全局文本特征“一只猫和一只狗在草地上”、对象文本特征“猫”、“狗”和“草地”、至少一个目标对象的位置特征“猫的坐标”、“狗的坐标”和“草地的坐标”以及至少一个目标对象的图像特征“一只猫的图像”、“一只狗的图像”和“草地的图像”进行特征组合,从而得到多模态提示信息。
在一种可选的实施例中,在步骤S21中,获取多模态提示信息,包括如下方法步骤:
步骤S211,获取文本信息与附加信息,其中,附加信息包括以下至少之一:至少一个目标对象的位置信息、至少一个目标对象的图像信息;
步骤S212,基于文本信息与附加信息确定增强标记信息;
步骤S213,对文本信息与增强标记信息进行组合,得到多模态提示信息。
文本信息可以理解为用户输入的用于描述待生成的图像内容的文本提示。
附加信息可以理解为用户输入的坐标和/或图像,可以理解的是,附加信息包括用户输入的至少一个目标对象的坐标、用户输入的至少一个目标对象的图像中的至少之一。
本公开实施例中,通过获取用户输入的文本信息与附加信息,从而能够根据获取到的文本信息与附加信息能够确定出用于确定至少一个目标对象的位置特征与图像特
征的增强标记信息,进而能够通过对文本信息与增强标记信息进行组合,得到多模态提示信息。
示例性地,可以将文本信息与增强标记信息进行垂直拼接,从而得到多模态提示信息,此处不予限制。
图3是根据本公开实施例1的另一种图像生成方法的流程图,如图3所示,本公开的图像生成模型支持将文本信息、增强标记信息共同作为输入,从而生成目标图像,即支持将文本模态结合坐标模态和图像模态,基于多模态的提示信息生成目标图像。此外,本公开的图像生成模型还支持在模态缺失时,仅将文本信息作为输入,从而根据文本信息生成对应的增强标记信息,即根据文本模态生成对应的坐标模态和图像模态,进而根据文本模态和生成的坐标模态和图像模态生成目标图像。
进一步地,在基于多模态提示的图像生成过程中,可以根据输入的文本信息确定文本信息中的至少一个目标分词(即名词)、至少一个目标对象的坐标以及至少一个目标对象的图像,然后通过文本编码器确定的至少一个目标分词对应的对象文本特征,根据至少一个目标对象的坐标确定至少一个目标对象的位置特征,通过图像编码器确定至少一个目标对象的图像特征。再对对象文本特征、至少一个目标对象的位置特征以及至少一个目标对象的图像特征进行组合,得到增强标记信息,进而将文本信息嵌入和增强标记信息组合后输入至图像生成模型,最终得到目标图像。
示例性地,文本信息可以为“一只猫和一只狗在草地上”,因此能够确定多个目标分词包括“狗”、“猫”和“草地”,至少一个目标对象的坐标包括[狗的坐标]、[猫的坐标]以及[草地的坐标],至少一个目标对象的图像包括狗的图像、猫的图像和草地的图像。然后通过文本编辑器确定多个目标分词对应的对象文本特征,根据至少一个目标对象的坐标确定至少一个目标对象的位置特征,通过图像编码器确定至少一个目标对象的图像特征。再对对象文本特征、至少一个目标对象的位置特征以及至少一个目标对象的图像特征进行组合,得到增强标记信息,进而将文本信息“一只猫和一只狗在草地上”和增强标记信息组合后输入至图像生成模型,最终得到目标图像。
在基于纯文本提示的图像生成过程中,为了应对模态缺失,可以仅输入文本信息,基于分词模型确定出文本信息中的至少一个目标分词,并根据文字编码器确定至少一个目标分词对应的对象文本特征。然后采用坐标模型对文本信息和至少一个目标分词进行位置特征生成,得到至少一个目标对象的位置特征,采用图像特征模型对文本信息和至少一个目标对象的位置特征进行图像特征生成,得到至少一个目标对象的图像特征。再对对象文本特征、生成的至少一个目标对象的位置特征以及至少一个目标对象的图像特征进行组合,得到增强标记信息,进而将文本信息嵌入和增强标记信息组合后输入至图像生成模型,最终得到目标图像。
示例性地,文本信息可以为“一只猫和一只狗在草地上”,因此能够根据分词模型
确定至少一个目标分词包括“狗”、“猫”和“草地”,并根据文字编码器确定至少一个目标分词对应的对象文本特征。然后采用坐标模型对文本信息和至少一个目标分词进行位置特征生成,得到至少一个目标对象的坐标([狗的坐标]、[猫的坐标]以及[草地的坐标])对应的位置特征,采用图像特征模型对文本信息和至少一个目标对象的位置特征进行图像特征生成,得到至少一个目标对象的图像(狗的图像、猫的图像和草地的图像)对应的图像特征。再对对象文本特征、生成的至少一个目标对象的位置特征以及至少一个目标对象的图像特征进行组合,得到增强标记信息,进而将文本信息嵌入和增强标记信息组合后输入至图像生成模型,最终得到目标图像。
可以看出,本公开给定了一个图像-文本对,通过获取物体级别(object-level)的文本、坐标和图像,并将这些信息整合到每个物体(object)的“增强标记(augmented token)”中。增强标记作为附加条件与文本提示一起在扩散模型中进行训练,使得本公开的图像生成模型能够处理多物体多模态的文本提示。
此外,为了在推理过程中处理模态缺失问题,即为了解决零样本图像生成的问题,本公开提出利用一个坐标模型和一个图像特征模型,根据文本提示生成物体级别的坐标和图像特征。因此,本公开可以仅通过文本提示或通过各种多模态提示的组合来生成目标图像,能够以灵活的方式结合各种模态生成目标图像。且通过大量的定性和定量实验证明,本公开不仅优于相关技术中的图像生成方法,还能够完成更广泛的图像生成任务。
容易理解的是,本公开提供的图像生成方法的有益效果包括以下几点。
有益效果(1),支持基于多物体级别的多模态提示生成高质量图像,能够对生成图像实现更精准的控制,应用范围更广;
有益效果(2),设计了坐标模型和图像特征模型,支持基于文本模态生成坐标和图像模态,从而能够克服可能出现的模态缺失问题;
有益效果(3),相关技术的图像生成模型大部分都是基于文本模态生成的,而本公开的图像生成模型可以同时基于文本、图像、坐标模态生成目标图像,生成的图像的精确度更高;
有益效果(4),相关技术的图像生成模型若想要生成一个给定的物体,例如狗,由于狗有很多品种,因此若想生成某种特定的狗,一般需要给出3~5张图片,对模型进行微调,模型才能学会,而本公开的图像生成模型不需要微调的过程,只需要在推断的时候给出一张狗的图片,就可以生成对应品种的狗,即能够实现零样本生成,在训练的时候不需要目标物体的参与,因此模型训练成本较低;
有益效果(5),本公开将增强标记添加到文本嵌入后,作为扩散模型生成图像的采样条件,从而使本公开不需要改变扩散模型的架构,只需改变扩散模型的输入,进而提高了在扩散模型的基础上构建技术的可用性。
需要说明的是,本公开所涉及的用户信息(包括但不限于用户设备信息、用户个人信息等)和数据(包括但不限于用于分析的数据、存储的数据、展示的数据等),均为经用户授权或者经过各方充分授权的信息和数据,并且相关数据的收集、使用和处理需要遵守相关国家和地区的相关法律法规和标准,并提供有相应的操作入口,供用户选择授权或者拒绝。
另外,还需要说明的是,对于前述的各方法实施例,为了简单描述,故将其都表述为一系列的动作组合,但是本领域技术人员应该知悉,本公开并不受所描述的动作顺序的限制,因为依据本公开,某些步骤可以采用其他顺序或者同时进行。其次,本领域技术人员也应该知悉,说明书中所描述的实施例均属于优选实施例,所涉及的动作和模块并不一定是本公开所必须的。
通过以上的实施方式的描述,本领域的技术人员可以清楚地了解到根据上述实施例的方法可借助软件加必需的通用硬件平台的方式来实现,当然也可以通过硬件。基于这样的理解,本公开的技术方案本质上或者说对现有技术做出贡献的部分可以以软件产品的形式体现出来,该计算机软件产品存储在一个存储介质(如ROM/RAM、磁碟、光盘)中,包括若干指令用以使得一台终端设备(可以是手机,计算机,服务器,或者网络设备等)执行本公开各个实施例所述的方法。
实施例2
在如实施例1中的运行环境下,本公开提供了如图4所示的一种图像生成方法,通过终端设备提供一图形用户界面,图形用户界面所显示的内容至少部分地包含一图像生成场景,图4是根据本公开实施例2的一种图像生成方法的流程图,如图4所示,该方法包括:
步骤S41,响应对图形用户界面执行的第一控制操作,输入文本信息,其中,文本信息用于描述待生成的图像内容,图像内容包括:至少一个目标对象;
步骤S42,响应对图形用户界面执行的第二控制操作,基于文本信息分别生成至少一个目标对象的位置特征与至少一个目标对象的图像特征以得到增强标记信息,以及采用图像生成模型对多模态提示信息进行多模态图像生成以得到目标图像,其中,增强标记信息用于确定至少一个目标对象的位置特征与图像特征,图像生成模型用于采用多模态图像生成方式生成目标图像,多模态提示信息包括:文本信息与增强标记信息;
步骤S43,在图形用户界面内展示目标图像。
本公开实施例中的图形用户界面中至少显示有图像生成场景,用户能够通过执行控制操作在该图像生成场景中输入用于描述待生成的图像内容的文本信息,控制基于文本信息生成至少一个目标对象的位置特征与至少一个目标对象的图像特征以得到增强标记信息,控制采用图像生成模型对多模态提示信息进行多模态图像生成以得到目
标图像等步骤。可以理解的是,上述图像生成场景可以但不限于剧本、设计、游戏、电商、教育、医疗、会议、社交网络、金融产品、物流和导航等领域中涉及图像生成的应用场景。
上述图形用户界面还包括第一控件(或第一触控区域),当检测到作用于第一控件(或第一触控区域)的第一触控操作时,可以获取用户输入的文本信息。上述文本信息可以是用户通过第一触控操作从图形用户界面中的文本框进行输入的。上述第一触控操作可以是点选、框选、勾选、条件筛选等操作,此处不予限制。
文本信息可以理解为文本提示,即使用自然语言文字描述的图像内容,用于描述待生成的图像内容。示例性地,自然语言文字可以为中文、英文、日文等,此处不予限制。本公开实施例中,文本信息所描述的图像内容中可以包括至少一个目标对象,目标对象可以理解为待生成的图像内容中的人物、动物或物体等,即本公开能够支持多对象的图像生成。示例性地,目标对象可以包括男孩、女孩、医生、老师等真实人物,可以包括游戏玩家、非玩家角色(Non-Player Character,NPC)等虚拟人物,可以包括猫、狗、孔雀、大象等动物,还可以包括桌子、汽车、草地、树木、火车站等物体,此处不予限制。
示例性地,文本信息可以为用户的图像生成需求对应的文本信息,若用户希望得到一只猫和一只狗在草地上的图像,则对应的文本信息可以为“一只猫和一只狗在草地上(即a cat and a dog on the grass)”,该文本信息中即包括三个目标对象,分别为猫、狗和草地。可以理解的是,该文本信息还可以替换为“一只狗和一只猫在草地上”,或者“草地上有一只猫和一只狗”,本公开对文本信息的描述方式不予限制。
上述图形用户界面还包括第二控件(或第二触控区域),当检测到作用于第二控件(或第二触控区域)的第二触控操作时,能够基于文本信息分别生成至少一个目标对象的位置特征与至少一个目标对象的图像特征以得到增强标记信息,以及采用图像生成模型对多模态提示信息进行多模态图像生成以得到目标图像。上述第二触控操作可以是点选、框选、勾选、条件筛选等操作,此处不予限制。
考虑到仅通过文本提示描述物体会导致图像生成不准确,因此本公开在提供文本信息的基础上,将增强标记信息作为附加信息,通过文本信息和增强标记信息共同生成目标图像,从而提高图像生成的准确率。
增强标记信息可以包括坐标信息和图像信息,其中,坐标信息可以为文本信息中包括的至少一个目标对象的坐标,用于确定至少一个目标对象的位置特征。示例性地,若用户希望得到一只猫和一只狗在草地上的图像,则坐标信息即包括猫的坐标(例如[坐标:12,15,100,200])、狗的坐标(例如[坐标:100,150,220,240])和草地的坐标(例如[坐标:500,500,500,500])。
示例性地,至少一个目标对象的坐标可以是二维直角坐标系下的坐标,能够表示
目标对象的坐标范围,也可以是二维坐标系或三维坐标系下的坐标,本公开对所采用的坐标系不予限制。
图像信息可以为文本信息中包括的至少一个目标对象的图像,用于确定文本信息中至少一个目标对象的图像特征。示例性地,若用户希望得到一只猫和一只狗在草地上的图像,则图像信息即包括用户所需要生成的一只猫的图像(例如[图像:一张用户上传的自己的猫的图像])、一只狗的图像(例如[图像:一张用户上传的自己的狗的图像])以及草地的图像(例如[图像:一张用户上传的草地的图像])。
多模态提示信息可以理解为多种模态的提示信息,包括上述文本信息和增强标记信息。本公开实施例中,多模态提示信息包括文本模态、坐标模态以及图像模态的提示信息。
示例性地,可以将同一目标对象的多模态提示信息绑定在一起,以避免影响其他物体或全局条件,也即本公开中引入增强标记(augmented token),增强标记同时包含物体级的文本、坐标和图像信息,用于描述生成图像中的一个物体。
本公开实施例中,图像生成模型为适用于生成图像的模型,示例性地,图像生成模型可以为预训练的基于扩散过程的图像生成模型,即为扩散模型,图像生成模型还可以为自回归模型等,此处不予限制。图像生成模型能够基于多目标对象的多模态提示信息准确生成高质量目标图像,即能够基于多物体级别的多模态提示生成高质量图像,从而实现对生成图像的精准控制。
示例性地,若将上述示例的多模态提示信息输入至图像生成模型,即输入文本信息:一只猫和一只狗在草地上、坐标信息:猫的坐标、狗的坐标和草地的坐标、图像信息:一只猫的图像、一只狗的图像以及草地的图像,则本公开的图像生成模型能够根据该多模态提示信息准确输出对应的目标图像,即一只猫和一只狗在草地上的图像。
可以理解的是,本公开实施例中的图像生成模型能够尽可能保持预训练的文本到图像模型的原始结构,以便集成对现有模型的扩展,并且本公开实施例中的图像生成模型仅改变了模型的输入,不需要改变模型的架构,因此能够保持在模型的基础上构建技术的可用性。
本公开实施例中,通过终端设备提供一图形用户界面,图形用户界面所显示的内容至少部分地包含一图像生成场景,若用户对图形用户界面执行了第一控制操作,则用户输入了用于描述待生成的至少一个目标对象的图像内容的文本信息,若用户对图形用户界面执行了第二控制操作,例如提交操作,则能够基于文本信息分别生成至少一个目标对象的位置特征与至少一个目标对象的图像特征,从而得到增强标记信息,同时能够采用图像生成模型对多模态提示信息进行多模态图像生成从而得到目标图像,进而将生成得到的目标图像在图形用户界面内展示,以反馈至用户。达到了通过多模态提示信息准确生成包括多目标对象的高质量目标图像的目的,且能够对生成的图像
实现更精准的控制,生成准确度更高的图像,且支持生成多目标对象的图像,使得图像的生成范围更广泛。
需要说明的是,上述第一触控操作和第二触控操作均可以是用户用手指接触上述终端设备的显示屏并触控该终端设备的操作。该触控操作可以包括单点触控、多点触控,其中,每个触控点的触控操作可以包括点击、长按、重按、划动等。上述第一触控操作和第二触控操作还可以是通过鼠标、键盘等输入设备实现的触控操作,此处不予限制。
本公开实施例提供的上述图像生成方法可以但不限于应用于剧本服务、设计服务、游戏服务、电商服务、教育服务、法律服务、医疗服务、会议服务、社交网络服务、金融产品服务、物流服务和导航服务等领域中涉及图像生成的应用场景中,例如:剧本服务中根据剧本或故事情节生成所需的图像素材、设计服务中根据用户的文字、图像、坐标描述生成设计图、游戏中根据角色形象的文字描述生成对应角色图、电商服务中根据商品的描述生成对应商品图等,此处不予限制。
采用本公开实施例,通过终端设备提供一图形用户界面,图形用户界面所显示的内容至少部分地包含一图像生成场景,若用户对图形用户界面执行了第一控制操作,则用户输入了用于描述待生成的至少一个目标对象的图像内容的文本信息,若用户对图形用户界面执行了第二控制操作,例如提交操作,则能够基于文本信息分别生成至少一个目标对象的位置特征与至少一个目标对象的图像特征,从而得到增强标记信息,同时能够采用图像生成模型对多模态提示信息进行多模态图像生成从而得到目标图像,进而将生成得到的目标图像在图形用户界面内展示,以反馈至用户,达到了通过多模态提示信息准确生成包括多目标对象的高质量目标图像的目的,且能够对生成的图像实现更精准的控制,生成准确度更高的图像,且支持生成多目标对象的图像,使得图像的生成范围更广泛,进而解决了相关技术中仅通过文本提示生成对应图像,导致生成的图像精确度较低,图像生成范围较局限的技术问题。
需要说明的是,本实施例的优选实施方式可以参见实施例1中的相关描述,此处不再赘述。
实施例3
在如实施例1中的运行环境下,本公开提供了如图5所示的一种图像生成方法。图5是根据本公开实施例3的一种图像生成方法的流程图,如图5所示,该方法包括:
步骤S51,接收当前输入的多模态对话请求,其中,多模态对话请求中携带的信息包括:多模态对话文本信息与多模态对话增强标记信息,多模态对话文本信息用于描述待生成的多模态对话图像内容,多模态对话图像内容包括:至少一个对象,多模态对话增强标记信息用于确定至少一个对象的位置特征与图像特征;
步骤S52,采用图像生成模型对多模态对话请求进行多模态图像生成,得到多模
态对话图像,其中,图像生成模型用于采用多模态图像生成方式生成多模态对话图像;
步骤S53,反馈多模态对话回复,其中,多模态对话回复中携带的信息包括:多模态对话图像。
多模态对话请求可以理解为用户向计算机或机器人发起的对话请求(request),该多模态对话请求中携带有多模态对话文本信息与多模态对话增强标记信息。其中,多模态对话文本信息可以理解为文本提示,即使用自然语言文字描述的待生成的多模态对话图像内容。示例性地,自然语言文字可以为中文、英文、日文等,此处不予限制。
本公开实施例中,多模态对话文本信息所描述的多模态对话图像内容中可以包括至少一个对象,对象可以理解为待生成的多模态对话图像内容中的人物、动物或物体等,即本公开能够支持多对象的图像生成。示例性地,对象可以包括男孩、女孩、医生、老师等真实人物,可以包括游戏玩家、非玩家角色(Non-Player Character,NPC)等虚拟人物,可以包括猫、狗、孔雀、大象等动物,还可以包括桌子、汽车、草地、树木、火车站等物体,此处不予限制。
示例性地,多模态对话文本信息可以为用户输入的多模态对话请求对应的文本信息,若用户希望得到一只猫和一只狗在草地上的图像,则对应的多模态对话文本信息可以为“一只猫和一只狗在草地上(即a cat and a dog on the grass)”,该多模态对话文本信息中即包括三个对象,分别为猫、狗和草地。可以理解的是,该多模态对话文本信息还可以替换为“一只狗和一只猫在草地上”,或者“草地上有一只猫和一只狗”,本公开对多模态对话文本信息的描述方式不予限制。
考虑到仅通过多模态对话文本提示描述物体会导致多模态对话图像生成不准确,因此本公开在提供多模态对话文本信息的基础上,将多模态对话增强标记信息作为附加信息,通过多模态对话文本信息和多模态对话增强标记信息共同生成多模态对话图像,从而提高多模态对话图像生成的准确率。
多模态对话增强标记信息可以包括坐标信息和图像信息,其中,坐标信息可以为多模态对话文本信息中包括的至少一个对象的坐标,用于确定至少一个对象的位置特征。示例性地,若用户希望得到一只猫和一只狗在草地上的图像,则坐标信息即包括猫的坐标(例如[坐标:12,15,100,200])、狗的坐标(例如[坐标:100,150,220,240])和草地的坐标(例如[坐标:500,500,500,500])。
示例性地,至少一个对象的坐标可以是二维直角坐标系下的坐标,能够表示目标对象的坐标范围,也可以是二维坐标系或三维坐标系下的坐标,本公开对所采用的坐标系不予限制。
图像信息可以为多模态对话文本信息中包括的至少一个对象的图像,用于确定多模态对话文本信息中至少一个对象的图像特征。示例性地,若用户希望得到一只猫和一只狗在草地上的图像,则图像信息即包括用户所需要生成的一只猫的图像(例如[图
像:一张用户上传的自己的猫的图像])、一只狗的图像(例如[图像:一张用户上传的自己的狗的图像])以及草地的图像(例如[图像:一张用户上传的草地的图像])。
可以看出,本公开实施例中的多模态对话请求携带了多种模态的提示信息,包括上述多模态对话文本信息和多模态对话增强标记信息,也即多模态提示信息包括文本模态、坐标模态以及图像模态的提示信息。
示例性地,可以将同一对象的多种模态的提示信息绑定在一起,以避免影响其他物体或全局条件,也即本公开中引入增强标记(augmented token),增强标记同时包含物体级的文本、坐标和图像信息,用于描述生成图像中的一个物体。
本公开实施例中,图像生成模型为适用于生成图像的模型,示例性地,图像生成模型可以为预训练的基于扩散过程的图像生成模型,即为扩散模型,图像生成模型还可以为自回归模型等,此处不予限制。图像生成模型能够基于多模态对话请求准确生成高质量多模态对话图像,即能够基于多物体级别的多模态提示生成高质量图像,从而实现对生成多模态对话图像的精准控制。
示例性地,若采用图像生成模型对上述示例的多模态对话请求进行多模态图像生成,即向图像生成模型输入多模态对话文本信息:一只猫和一只狗在草地上、坐标信息:猫的坐标、狗的坐标和草地的坐标、图像信息:一只猫的图像、一只狗的图像以及草地的图像,则本公开的图像生成模型能够根据该多模态对话请求准确输出对应的多模态对话图像,即一只猫和一只狗在草地上的图像。
可以理解的是,本公开实施例中的图像生成模型能够尽可能保持预训练的文本到图像模型的原始结构,以便集成对现有模型的扩展,并且本公开实施例中的图像生成模型仅改变了模型的输入,不需要改变模型的架构,因此能够保持在模型的基础上构建技术的可用性。
多模态对话回复可以理解为计算机或机器人基于用户输入的多模态对话请求向用户反馈的答复内容(response),与用户输入的多模态对话请求相对应。多模态对话回复中携带了多模态对话图像,该多模态对话图像即为包括多模态对话请求中描述的待生成的多模态对话图像内容的图像。
上述步骤S51至步骤S53可以应用于人机对话场景,即用户与计算机或机器人之间进行对话的场景。如图6所示,图6是根据本公开实施例3的一种人机对话场景示意图,用户向计算机或机器人输入多模态对话请求,计算机或机器人能够向用户反馈该多模态对话请求对应的多模态对话回复。
可以理解的是,用户与计算机或机器人之间的对话可以是通过语音识别和自然语言处理技术实现的,也可以是通过文本交流的方式进行的,此处不予限制。
本公开实施例中,通过接收用户输入的包括多模态对话文本信息与多模态对话增强标记信息的多模态对话请求,能够根据多模态对话文本信息确定待生成的多模态对
话图像内容、根据多模态对话增强标记信息确定待生成的多模态对话图像内容中至少一个对象的位置特征和图像特征,然后采用图像生成模型对多模态对话请求进行多模态图像生成,从而得到多模态对话图像,进而向用户反馈将携带多模态对话图像的多模态对话回复。由此,达到了通过多模态对话请求准确生成包括多对象的高质量多模态对话图像的目的,且能够对生成的多模态对话图像实现更精准的控制,生成准确度更高的多模态对话图像,且支持生成多对象的多模态对话图像,使得多模态对话图像的生成范围更广泛。
本公开实施例提供的上述图像生成方法还可以但不限于应用于剧本服务、设计服务、游戏服务、电商服务、教育服务、法律服务、医疗服务、会议服务、社交网络服务、金融产品服务、物流服务和导航服务等领域中涉及图像生成的应用场景中,例如:设计服务中根据用户的文字、图像、坐标描述生成设计图、游戏中根据角色形象的文字描述生成对应角色图、电商服务中根据商品的描述生成对应商品图等,此处不予限制。
采用本公开实施例,通过接收用户输入的包括多模态对话文本信息与多模态对话增强标记信息的多模态对话请求,能够根据多模态对话文本信息确定待生成的多模态对话图像内容、根据多模态对话增强标记信息确定待生成的多模态对话图像内容中至少一个对象的位置特征和图像特征,然后采用图像生成模型对多模态对话请求进行多模态图像生成,从而得到多模态对话图像,进而向用户反馈将携带多模态对话图像的多模态对话回复。由此,达到了通过多模态对话请求准确生成包括多对象的高质量多模态对话图像的目的,且能够对生成的多模态对话图像实现更精准的控制,生成准确度更高的多模态对话图像,且支持生成多对象的多模态对话图像,使得多模态对话图像的生成范围更广泛,进而解决了相关技术中仅通过文本提示生成对应图像,导致生成的图像精确度较低,图像生成范围较局限的技术问题。
在一种可选的实施例中,在步骤S51中,接收当前输入的多模态对话请求,包括如下方法步骤:
步骤S511,获取多模态对话文本信息;
步骤S512,基于多模态对话文本信息分别生成至少一个对象的位置特征与至少一个对象的图像特征,得到多模态对话增强标记信息;
步骤S513,对多模态对话文本信息与多模态对话增强标记信息进行组合,得到多模态对话请求。
本公开实施例中,在接收当前输入的多模态对话请求时,可以先获取多模态对话文本信息,然后基于多模态对话文本信息分别生成至少一个对象的位置特征与至少一个对象的图像特征,即根据多模态对话文本提示生成物体级别的坐标和图像,从而得到对应的多模态对话增强标记信息,再将多模态对话文本信息与多模态对话增强标记
信息进行组合,即可得到多模态对话请求。
需要说明的是,本实施例的优选实施方式可以参见实施例1中的相关描述,此处不再赘述。
实施例4
根据本公开实施例,还提供了一种用于实施上述图像生成的装置实施例。图7是根据本公开实施例4的一种图像生成装置的结构示意图,如图7所示,该装置包括:
获取模块701,被设置为获取多模态提示信息,其中,多模态提示信息包括:文本信息与增强标记信息,文本信息用于描述待生成的图像内容,图像内容包括:至少一个目标对象,增强标记信息用于确定至少一个目标对象的位置特征与图像特征;
图像生成模块702,被设置为采用图像生成模型对多模态提示信息进行多模态图像生成,得到目标图像,其中,图像生成模型用于采用多模态图像生成方式生成目标图像。
可选地,上述获取模块701还被设置为:获取文本信息;基于文本信息分别生成至少一个目标对象的位置特征与至少一个目标对象的图像特征,得到增强标记信息;对文本信息与增强标记信息进行组合,得到多模态提示信息。
可选地,上述获取模块701还被设置为:对文本信息进行词性分析,选取至少一个目标分词,其中,至少一个目标分词满足预设词性要求,至少一个目标分词用于确定至少一个目标对象;基于文本信息和至少一个目标分词生成至少一个目标对象的位置特征,以及基于文本信息和至少一个目标对象的位置特征生成至少一个目标对象的图像特征;利用至少一个目标分词、至少一个目标对象的位置特征以及至少一个目标对象的图像特征确定增强标记信息。
可选地,上述获取模块701还被设置为:采用位置特征生成模型对文本信息和至少一个目标分词进行位置特征生成,得到至少一个目标对象的位置特征。
可选地,上述获取模块701还被设置为:采用图像特征生成模型对文本信息和至少一个目标对象的位置特征进行图像特征生成,得到至少一个目标对象的图像特征。
可选地,还包括:编码模块,被设置为:采用文本编码器对文本信息进行文本编码,得到全局文本特征,以及采用文本编码器对至少一个目标分词进行文本编码,得到对象文本特征。
可选地,上述获取模块701还被设置为:对对象文本特征、至少一个目标对象的位置特征以及至少一个目标对象的图像特征进行组合,得到增强标记信息。
可选地,上述获取模块701还被设置为:对全局文本特征、对象文本特征、至少一个目标对象的位置特征以及至少一个目标对象的图像特征进行组合,得到多模态提示信息。
可选地,上述获取模块701还被设置为:获取文本信息与附加信息,其中,附加
信息包括以下至少之一:至少一个目标对象的位置信息、至少一个目标对象的图像信息;基于文本信息与附加信息确定增强标记信息;对文本信息与增强标记信息进行组合,得到多模态提示信息。
采用本公开实施例,通过获取包括文本信息与增强标记信息的多模态提示信息,能够根据文本信息确定待生成的图像内容、根据增强标记信息确定图像内容中至少一个目标对象的位置特征和图像特征,然后采用图像生成模型对多模态提示信息进行多模态图像生成,从而得到目标图像,由此达到了通过多模态提示信息准确生成包括多目标对象的高质量目标图像的目的。此外,能够对生成的图像实现更精准的控制,生成准确度更高的图像,且支持生成多目标对象的图像,使得图像的生成范围更广泛,从而实现了更精确、更多样化、支持多目标对象的高质量图像生成,以满足不同领域的图像生成需求的技术效果,进而解决了相关技术中仅通过文本提示生成对应图像,导致生成的图像精确度较低,图像生成范围较局限的技术问题。
此处需要说明的是,上述获取模块701和图像生成模块702对应于实施例1中的步骤S21和步骤S22,两个模块与对应的步骤所实现的实例和应用场景相同,但不限于上述实施例1所公开的内容。需要说明的是,上述模块或单元可以是存储在存储器中并由一个或多个处理器处理的硬件组件或软件组件,上述模块也可以运行在实施例1提供的服务器10中。
根据本公开实施例,还提供了另一种用于实施上述图像生成的装置实施例。图8是根据本公开实施例4的另一种图像生成装置的结构示意图,通过终端设备提供一图形用户界面,图形用户界面所显示的内容至少部分地包含一图像生成场景,如图8所示,该装置包括:
第一响应模块801,被设置为响应对图形用户界面执行的第一控制操作,输入文本信息,其中,文本信息用于描述待生成的图像内容,图像内容包括:至少一个目标对象;
第二响应模块802,被设置为响应对图形用户界面执行的第二控制操作,基于文本信息分别生成至少一个目标对象的位置特征与至少一个目标对象的图像特征以得到增强标记信息,以及采用图像生成模型对多模态提示信息进行多模态图像生成以得到目标图像,其中,增强标记信息用于确定至少一个目标对象的位置特征与图像特征,图像生成模型用于采用多模态图像生成方式生成目标图像,多模态提示信息包括:文本信息与增强标记信息;
展示模块803,被设置为在图形用户界面内展示目标图像。
采用本公开实施例,通过终端设备提供一图形用户界面,图形用户界面所显示的内容至少部分地包含一图像生成场景,若用户对图形用户界面执行了第一控制操作,则用户输入了用于描述待生成的至少一个目标对象的图像内容的文本信息,若用户对
图形用户界面执行了第二控制操作,例如提交操作,则能够基于文本信息分别生成至少一个目标对象的位置特征与至少一个目标对象的图像特征,从而得到增强标记信息,同时能够采用图像生成模型对多模态提示信息进行多模态图像生成从而得到目标图像,进而将生成得到的目标图像在图形用户界面内展示,以反馈至用户,达到了通过多模态提示信息准确生成包括多目标对象的高质量目标图像的目的,且能够对生成的图像实现更精准的控制,生成准确度更高的图像,且支持生成多目标对象的图像,使得图像的生成范围更广泛,进而解决了相关技术中仅通过文本提示生成对应图像,导致生成的图像精确度较低,图像生成范围较局限的技术问题。
此处需要说明的是,上述第一响应模块801、第二响应模块802和展示模块803对应于实施例2中的步骤S41至步骤S43,三个模块与对应的步骤所实现的实例和应用场景相同,但不限于上述实施例2所公开的内容。需要说明的是,上述模块或单元可以是存储在存储器中并由一个或多个处理器处理的硬件组件或软件组件,上述模块也可以运行在实施例1提供的服务器10中。
根据本公开实施例,还提供了再一种用于实施上述图像生成的装置实施例。图9是根据本公开实施例4的再一种图像生成装置的结构示意图,如图9所示,该装置包括:
接收模块901,被设置为接收当前输入的多模态对话请求,其中,多模态对话请求中携带的信息包括:多模态对话文本信息与多模态对话增强标记信息,多模态对话文本信息用于描述待生成的多模态对话图像内容,多模态对话图像内容包括:至少一个对象,多模态对话增强标记信息用于确定至少一个对象的位置特征与图像特征;
生成模块902,被设置为采用图像生成模型对多模态对话请求进行多模态图像生成,得到多模态对话图像,其中,图像生成模型用于采用多模态图像生成方式生成多模态对话图像;
反馈模块903,被设置为反馈多模态对话回复,其中,多模态对话回复中携带的信息包括:多模态对话图像。
可选地,上述接收模块901还被设置为:获取多模态对话文本信息;基于多模态对话文本信息分别生成至少一个对象的位置特征与至少一个对象的图像特征,得到多模态对话增强标记信息;对多模态对话文本信息与多模态对话增强标记信息进行组合,得到多模态对话请求。
采用本公开实施例,通过接收用户输入的包括多模态对话文本信息与多模态对话增强标记信息的多模态对话请求,能够根据多模态对话文本信息确定待生成的多模态对话图像内容、根据多模态对话增强标记信息确定待生成的多模态对话图像内容中至少一个对象的位置特征和图像特征,然后采用图像生成模型对多模态对话请求进行多模态图像生成,从而得到多模态对话图像,进而向用户反馈将携带多模态对话图像的
多模态对话回复。由此,达到了通过多模态对话请求准确生成包括多对象的高质量多模态对话图像的目的,且能够对生成的多模态对话图像实现更精准的控制,生成准确度更高的多模态对话图像,且支持生成多对象的多模态对话图像,使得多模态对话图像的生成范围更广泛,进而解决了相关技术中仅通过文本提示生成对应图像,导致生成的图像精确度较低,图像生成范围较局限的技术问题。
此处需要说明的是,上述接收模块901、生成模块902和反馈模块903对应于实施例3中的步骤S51至步骤S53,三个模块与对应的步骤所实现的实例和应用场景相同,但不限于上述实施例3所公开的内容。需要说明的是,上述模块或单元可以是存储在存储器中并由一个或多个处理器处理的硬件组件或软件组件,上述模块也可以运行在实施例1提供的服务器10中。
需要说明的是,本公开上述实施例中涉及到的优选实施方案与实施例1提供的方案以及应用场景、实施过程相同,但不仅限于实施例1所提供的方案。
实施例5
本公开的实施例可以提供一种计算机终端,该计算机终端可以是计算机终端群中的任意一个计算机终端设备。可选地,在本实施例中,上述计算机终端也可以替换为移动终端等终端设备。
可选地,在本实施例中,上述计算机终端可以位于计算机网络的多个网络设备中的至少一个网络设备。
在本实施例中,上述计算机终端可以执行图像生成方法中以下步骤的程序代码:获取多模态提示信息,其中,多模态提示信息包括:文本信息与增强标记信息,文本信息用于描述待生成的图像内容,图像内容包括:至少一个目标对象,增强标记信息用于确定至少一个目标对象的位置特征与图像特征;采用图像生成模型对多模态提示信息进行多模态图像生成,得到目标图像,其中,图像生成模型用于采用多模态图像生成方式生成目标图像。
可选地,图10是根据本公开实施例5的一种计算机终端的结构框图。如图10所示,该计算机终端A可以包括:一个或多个(图中仅示出一个)处理器1002、存储器1004、存储控制器、以及外设接口,其中,外设接口与射频模块、音频模块和显示器连接。
其中,存储器可被设置为存储软件程序以及模块,如本公开实施例中的图像生成方法和装置对应的程序指令/模块,处理器通过运行存储在内的软件程序以及模块,从而执行各种功能应用以及数据处理,即实现上述的图像生成方法。存储器可包括高速随机存储器,还可以包括非易失性存储器,如一个或者多个磁性存储装置、闪存、或者其他非易失性固态存储器。在一些实例中,存储器可进一步包括相对于处理器远程设置的存储器,这些远程存储器可以通过网络连接至计算机终端A。上述网络的实例
包括但不限于互联网、企业内部网、局域网、移动通信网及其组合。
处理器可以通过传输装置调用存储器存储的信息及应用程序,以执行下述步骤:获取多模态提示信息,其中,多模态提示信息包括:文本信息与增强标记信息,文本信息用于描述待生成的图像内容,图像内容包括:至少一个目标对象,增强标记信息用于确定至少一个目标对象的位置特征与图像特征;采用图像生成模型对多模态提示信息进行多模态图像生成,得到目标图像,其中,图像生成模型用于采用多模态图像生成方式生成目标图像。
可选地,上述处理器还可以执行如下步骤的程序代码:获取文本信息;基于文本信息分别生成至少一个目标对象的位置特征与至少一个目标对象的图像特征,得到增强标记信息;对文本信息与增强标记信息进行组合,得到多模态提示信息。
可选地,上述处理器还可以执行如下步骤的程序代码:对文本信息进行词性分析,选取至少一个目标分词,其中,至少一个目标分词满足预设词性要求,至少一个目标分词用于确定至少一个目标对象;基于文本信息和至少一个目标分词生成至少一个目标对象的位置特征,以及基于文本信息和至少一个目标对象的位置特征生成至少一个目标对象的图像特征;利用至少一个目标分词、至少一个目标对象的位置特征以及至少一个目标对象的图像特征确定增强标记信息。
可选地,上述处理器还可以执行如下步骤的程序代码:采用位置特征生成模型对文本信息和至少一个目标分词进行位置特征生成,得到至少一个目标对象的位置特征。
可选地,上述处理器还可以执行如下步骤的程序代码:采用图像特征生成模型对文本信息和至少一个目标对象的位置特征进行图像特征生成,得到至少一个目标对象的图像特征。
可选地,上述处理器还可以执行如下步骤的程序代码:采用文本编码器对文本信息进行文本编码,得到全局文本特征,以及采用文本编码器对至少一个目标分词进行文本编码,得到对象文本特征。
可选地,上述处理器还可以执行如下步骤的程序代码:对对象文本特征、至少一个目标对象的位置特征以及至少一个目标对象的图像特征进行组合,得到增强标记信息。
可选地,上述处理器还可以执行如下步骤的程序代码:对全局文本特征、对象文本特征、至少一个目标对象的位置特征以及至少一个目标对象的图像特征进行组合,得到多模态提示信息。
可选地,上述处理器还可以执行如下步骤的程序代码:获取文本信息与附加信息,其中,附加信息包括以下至少之一:至少一个目标对象的位置信息、至少一个目标对象的图像信息;基于文本信息与附加信息确定增强标记信息;对文本信息与增强标记信息进行组合,得到多模态提示信息。
采用本公开实施例,通过获取包括文本信息与增强标记信息的多模态提示信息,能够根据文本信息确定待生成的图像内容、根据增强标记信息确定图像内容中至少一个目标对象的位置特征和图像特征,然后采用图像生成模型对多模态提示信息进行多模态图像生成,从而得到目标图像,由此达到了通过多模态提示信息准确生成包括多目标对象的高质量目标图像的目的。此外,能够对生成的图像实现更精准的控制,生成准确度更高的图像,且支持生成多目标对象的图像,使得图像的生成范围更广泛,从而实现了更精确、更多样化、支持多目标对象的高质量图像生成,以满足不同领域的图像生成需求的技术效果,进而解决了相关技术中仅通过文本提示生成对应图像,导致生成的图像精确度较低,图像生成范围较局限的技术问题。
本领域普通技术人员可以理解,图10所示的结构仅为示意,计算机终端A也可以是智能手机(如Android手机、iOS手机等)、平板电脑、掌上电脑以及移动互联网设备(MobileInternetDevices,MID)、PAD等终端设备。图10其并不对上述电子装置的结构造成限定。例如,计算机终端A还可包括比图10中所示更多或者更少的组件(如网络接口、显示装置等),或者具有与图10所示不同的配置。
本领域普通技术人员可以理解上述实施例的各种方法中的全部或部分步骤是可以通过程序来指令终端设备相关的硬件来完成,该程序可以存储于一计算机可读存储介质中,存储介质可以包括:闪存盘、只读存储器(Read-Only Memory,ROM)、随机存取器(Random Access Memory,RAM)、磁盘或光盘等。
实施例6
本公开的实施例还提供了一种计算机可读存储介质。可选地,在本实施例中,上述计算机可读存储介质可以用于保存上述实施例一所提供的图像生成方法所执行的程序代码。
可选地,在本实施例中,上述计算机可读存储介质可以位于计算机网络中计算机终端群中的任意一个计算机终端中,或者位于移动终端群中的任意一个移动终端中。
可选地,在本实施例中,计算机可读存储介质被设置为存储用于执行以下步骤的程序代码:获取多模态提示信息,其中,多模态提示信息包括:文本信息与增强标记信息,文本信息用于描述待生成的图像内容,图像内容包括:至少一个目标对象,增强标记信息用于确定至少一个目标对象的位置特征与图像特征;采用图像生成模型对多模态提示信息进行多模态图像生成,得到目标图像,其中,图像生成模型用于采用多模态图像生成方式生成目标图像。
可选地,在本实施例中,计算机可读存储介质被设置为存储用于执行以下步骤的程序代码:获取文本信息;基于文本信息分别生成至少一个目标对象的位置特征与至少一个目标对象的图像特征,得到增强标记信息;对文本信息与增强标记信息进行组合,得到多模态提示信息。
可选地,在本实施例中,计算机可读存储介质被设置为存储用于执行以下步骤的程序代码:对文本信息进行词性分析,选取至少一个目标分词,其中,至少一个目标分词满足预设词性要求,至少一个目标分词用于确定至少一个目标对象;基于文本信息和至少一个目标分词生成至少一个目标对象的位置特征,以及基于文本信息和至少一个目标对象的位置特征生成至少一个目标对象的图像特征;利用至少一个目标分词、至少一个目标对象的位置特征以及至少一个目标对象的图像特征确定增强标记信息。
可选地,在本实施例中,计算机可读存储介质被设置为存储用于执行以下步骤的程序代码:采用位置特征生成模型对文本信息和至少一个目标分词进行位置特征生成,得到至少一个目标对象的位置特征。
可选地,在本实施例中,计算机可读存储介质被设置为存储用于执行以下步骤的程序代码:采用图像特征生成模型对文本信息和至少一个目标对象的位置特征进行图像特征生成,得到至少一个目标对象的图像特征。
可选地,在本实施例中,计算机可读存储介质被设置为存储用于执行以下步骤的程序代码:采用文本编码器对文本信息进行文本编码,得到全局文本特征,以及采用文本编码器对至少一个目标分词进行文本编码,得到对象文本特征。
可选地,在本实施例中,计算机可读存储介质被设置为存储用于执行以下步骤的程序代码:对对象文本特征、至少一个目标对象的位置特征以及至少一个目标对象的图像特征进行组合,得到增强标记信息。
可选地,在本实施例中,计算机可读存储介质被设置为存储用于执行以下步骤的程序代码:对全局文本特征、对象文本特征、至少一个目标对象的位置特征以及至少一个目标对象的图像特征进行组合,得到多模态提示信息。
可选地,在本实施例中,计算机可读存储介质被设置为存储用于执行以下步骤的程序代码:获取文本信息与附加信息,其中,附加信息包括以下至少之一:至少一个目标对象的位置信息、至少一个目标对象的图像信息;基于文本信息与附加信息确定增强标记信息;对文本信息与增强标记信息进行组合,得到多模态提示信息。
上述本公开实施例序号仅仅为了描述,不代表实施例的优劣。
在本公开的上述实施例中,对各个实施例的描述都各有侧重,某个实施例中没有详述的部分,可以参见其他实施例的相关描述。
在本公开所提供的几个实施例中,应该理解到,所揭露的技术内容,可通过其它的方式实现。其中,以上所描述的装置实施例仅仅是示意性的,例如所述单元的划分,仅仅为一种逻辑功能划分,实际实现时可以有另外的划分方式,例如多个单元或组件可以结合或者可以集成到另一个系统,或一些特征可以忽略,或不执行。另一点,所显示或讨论的相互之间的耦合或直接耦合或通信连接可以是通过一些接口,单元或模块的间接耦合或通信连接,可以是电性或其它的形式。
所述作为分离部件说明的单元可以是或者也可以不是物理上分开的,作为单元显示的部件可以是或者也可以不是物理单元,即可以位于一个地方,或者也可以分布到多个网络单元上。可以根据实际的需要选择其中的部分或者全部单元来实现本实施例方案的目的。
另外,在本公开各个实施例中的各功能单元可以集成在一个处理单元中,也可以是各个单元单独物理存在,也可以两个或两个以上单元集成在一个单元中。上述集成的单元既可以采用硬件的形式实现,也可以采用软件功能单元的形式实现。
所述集成的单元如果以软件功能单元的形式实现并作为独立的产品销售或使用时,可以存储在一个计算机可读取存储介质中。基于这样的理解,本公开的技术方案本质上或者说对现有技术做出贡献的部分或者该技术方案的全部或部分可以以软件产品的形式体现出来,该计算机软件产品存储在一个存储介质中,包括若干指令用以使得一台计算机设备(可为个人计算机、服务器或者网络设备等)执行本公开各个实施例所述方法的全部或部分步骤。而前述的存储介质包括:U盘、只读存储器(ROM,Read-Only Memory)、随机存取存储器(RAM,Random Access Memory)、移动硬盘、磁碟或者光盘等各种可以存储程序代码的介质。
以上所述仅是本公开的优选实施方式,应当指出,对于本技术领域的普通技术人员来说,在不脱离本公开原理的前提下,还可以做出若干改进和润饰,这些改进和润饰也应视为本公开的保护范围。
Claims (14)
- 一种图像生成方法,包括:获取多模态提示信息,其中,所述多模态提示信息包括:文本信息与增强标记信息,所述文本信息用于描述待生成的图像内容,所述图像内容包括:至少一个目标对象,所述增强标记信息用于确定所述至少一个目标对象的位置特征与图像特征;采用图像生成模型对所述多模态提示信息进行多模态图像生成,得到目标图像,其中,所述图像生成模型用于采用多模态图像生成方式生成所述目标图像。
- 根据权利要求1所述的图像生成方法,其中,获取所述多模态提示信息包括:获取所述文本信息;基于所述文本信息分别生成所述至少一个目标对象的位置特征与所述至少一个目标对象的图像特征,得到所述增强标记信息;对所述文本信息与所述增强标记信息进行组合,得到所述多模态提示信息。
- 根据权利要求2所述的图像生成方法,其中,基于所述文本信息分别生成所述至少一个目标对象的位置特征与图像特征,得到所述增强标记信息包括:对所述文本信息进行词性分析,选取至少一个目标分词,其中,所述至少一个目标分词满足预设词性要求,所述至少一个目标分词用于确定所述至少一个目标对象;基于所述文本信息和所述至少一个目标分词生成所述至少一个目标对象的位置特征,以及基于所述文本信息和所述至少一个目标对象的位置特征生成所述至少一个目标对象的图像特征;利用所述至少一个目标分词、所述至少一个目标对象的位置特征以及所述至少一个目标对象的图像特征确定所述增强标记信息。
- 根据权利要求3所述的图像生成方法,其中,基于所述文本信息和所述至少一个目标分词生成所述至少一个目标对象的位置特征包括:采用位置特征生成模型对所述文本信息和所述至少一个目标分词进行位置特征生成,得到所述至少一个目标对象的位置特征。
- 根据权利要求3所述的图像生成方法,其中,基于所述文本信息和所述至少一个目标对象的位置特征生成所述至少一个目标对象的图像特征包括:采用图像特征生成模型对所述文本信息和所述至少一个目标对象的位置特征进行图像特征生成,得到所述至少一个目标对象的图像特征。
- 根据权利要求3所述的图像生成方法,其中,所述图像生成方法还包括:采用文本编码器对所述文本信息进行文本编码,得到全局文本特征,以及采 用所述文本编码器对所述至少一个目标分词进行文本编码,得到对象文本特征。
- 根据权利要求6所述的图像生成方法,其中,利用所述至少一个目标分词、所述至少一个目标对象的位置特征以及所述至少一个目标对象的图像特征确定所述增强标记信息包括:对所述对象文本特征、所述至少一个目标对象的位置特征以及所述至少一个目标对象的图像特征进行组合,得到所述增强标记信息。
- 根据权利要求6所述的图像生成方法,其中,对所述文本信息与所述增强标记信息进行组合,得到所述多模态提示信息包括:对所述全局文本特征、所述对象文本特征、所述至少一个目标对象的位置特征以及所述至少一个目标对象的图像特征进行组合,得到所述多模态提示信息。
- 根据权利要求1所述的图像生成方法,其中,获取所述多模态提示信息包括:获取所述文本信息与附加信息,其中,所述附加信息包括以下至少之一:所述至少一个目标对象的位置信息、所述至少一个目标对象的图像信息;基于所述文本信息与所述附加信息确定所述增强标记信息;对所述文本信息与所述增强标记信息进行组合,得到所述多模态提示信息。
- 一种图像生成方法,通过终端设备提供一图形用户界面,所述图形用户界面所显示的内容至少部分地包含一图像生成场景,所述图像生成方法包括:响应对所述图形用户界面执行的第一控制操作,输入文本信息,其中,所述文本信息用于描述待生成的图像内容,所述图像内容包括:至少一个目标对象;响应对所述图形用户界面执行的第二控制操作,基于所述文本信息分别生成所述至少一个目标对象的位置特征与所述至少一个目标对象的图像特征以得到增强标记信息,以及采用图像生成模型对多模态提示信息进行多模态图像生成以得到目标图像,其中,所述增强标记信息用于确定所述至少一个目标对象的位置特征与图像特征,所述图像生成模型用于采用多模态图像生成方式生成所述目标图像,所述多模态提示信息包括:所述文本信息与所述增强标记信息;在所述图形用户界面内展示所述目标图像。
- 一种图像生成方法,包括:接收当前输入的多模态对话请求,其中,所述多模态对话请求中携带的信息包括:多模态对话文本信息与多模态对话增强标记信息,所述多模态对话文本信息用于描述待生成的多模态对话图像内容,所述多模态对话图像内容包括:至少一个对象,所述多模态对话增强标记信息用于确定所述至少一个对象的位置特征与图像特征;采用图像生成模型对所述多模态对话请求进行多模态图像生成,得到多模态对话图像,其中,所述图像生成模型用于采用多模态图像生成方式生成所述多模 态对话图像;反馈多模态对话回复,其中,所述多模态对话回复中携带的信息包括:所述多模态对话图像。
- 根据权利要求11所述的图像生成方法,其中,接收当前输入的所述多模态对话请求包括:获取所述多模态对话文本信息;基于所述多模态对话文本信息分别生成所述至少一个对象的位置特征与所述至少一个对象的图像特征,得到所述多模态对话增强标记信息;对所述多模态对话文本信息与所述多模态对话增强标记信息进行组合,得到所述多模态对话请求。
- 一种电子设备,包括:存储器,存储有可执行程序;处理器,用于运行所述程序,其中,所述程序运行时执行权利要求1至12中任意一项所述的图像生成方法。
- 一种计算机可读存储介质,所述计算机可读存储介质包括存储的可执行程序,其中,在所述可执行程序运行时控制所述计算机可读存储介质所在设备执行权利要求1至12中任意一项所述的图像生成方法。
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN202311613798.2A CN117689749A (zh) | 2023-11-28 | 2023-11-28 | 图像生成方法、电子设备及计算机可读存储介质 |
| CN202311613798.2 | 2023-11-28 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2025112588A1 true WO2025112588A1 (zh) | 2025-06-05 |
Family
ID=90125631
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/CN2024/107883 Pending WO2025112588A1 (zh) | 2023-11-28 | 2024-07-26 | 图像生成方法、电子设备及计算机可读存储介质 |
Country Status (2)
| Country | Link |
|---|---|
| CN (1) | CN117689749A (zh) |
| WO (1) | WO2025112588A1 (zh) |
Families Citing this family (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN117689749A (zh) * | 2023-11-28 | 2024-03-12 | 浙江阿里巴巴机器人有限公司 | 图像生成方法、电子设备及计算机可读存储介质 |
| CN118349307A (zh) * | 2024-04-24 | 2024-07-16 | 北京字跳网络技术有限公司 | 用于页面交互的方法、装置、设备、存储介质和计算机程序产品 |
| CN119234253A (zh) * | 2024-07-31 | 2024-12-31 | 北京字跳网络技术有限公司 | 图像生成方法及装置、存储介质、程序产品 |
| CN119625102B (zh) * | 2024-11-27 | 2025-11-25 | 广州酷狗计算机科技有限公司 | 图像生成方法、装置、设备及存储介质 |
Citations (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20220156992A1 (en) * | 2020-11-18 | 2022-05-19 | Adobe Inc. | Image segmentation using text embedding |
| CN114937181A (zh) * | 2022-04-15 | 2022-08-23 | 国网信息通信产业集团有限公司 | 一种基于多模态数据的输电巡检图像生成方法及装置 |
| CN116597039A (zh) * | 2023-05-22 | 2023-08-15 | 阿里巴巴(中国)有限公司 | 图像生成的方法和服务器 |
| CN116704062A (zh) * | 2023-06-02 | 2023-09-05 | 支付宝(杭州)信息技术有限公司 | 一种基于aigc的数据处理方法、装置、电子设备及存储介质 |
| CN117689749A (zh) * | 2023-11-28 | 2024-03-12 | 浙江阿里巴巴机器人有限公司 | 图像生成方法、电子设备及计算机可读存储介质 |
-
2023
- 2023-11-28 CN CN202311613798.2A patent/CN117689749A/zh active Pending
-
2024
- 2024-07-26 WO PCT/CN2024/107883 patent/WO2025112588A1/zh active Pending
Patent Citations (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20220156992A1 (en) * | 2020-11-18 | 2022-05-19 | Adobe Inc. | Image segmentation using text embedding |
| CN114937181A (zh) * | 2022-04-15 | 2022-08-23 | 国网信息通信产业集团有限公司 | 一种基于多模态数据的输电巡检图像生成方法及装置 |
| CN116597039A (zh) * | 2023-05-22 | 2023-08-15 | 阿里巴巴(中国)有限公司 | 图像生成的方法和服务器 |
| CN116704062A (zh) * | 2023-06-02 | 2023-09-05 | 支付宝(杭州)信息技术有限公司 | 一种基于aigc的数据处理方法、装置、电子设备及存储介质 |
| CN117689749A (zh) * | 2023-11-28 | 2024-03-12 | 浙江阿里巴巴机器人有限公司 | 图像生成方法、电子设备及计算机可读存储介质 |
Also Published As
| Publication number | Publication date |
|---|---|
| CN117689749A (zh) | 2024-03-12 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| US12182507B2 (en) | Text processing model training method, and text processing method and apparatus | |
| CN117520523B (zh) | 数据处理方法、装置、设备及存储介质 | |
| WO2025112588A1 (zh) | 图像生成方法、电子设备及计算机可读存储介质 | |
| JP2023509031A (ja) | マルチモーダル機械学習に基づく翻訳方法、装置、機器及びコンピュータプログラム | |
| US9805718B2 (en) | Clarifying natural language input using targeted questions | |
| CN112214591B (zh) | 一种对话预测的方法及装置 | |
| US7860705B2 (en) | Methods and apparatus for context adaptation of speech-to-speech translation systems | |
| Biswas et al. | Towards explanatory interactive image captioning using top-down and bottom-up features, beam search and re-ranking | |
| CN116913278B (zh) | 语音处理方法、装置、设备和存储介质 | |
| WO2025097982A1 (zh) | 文本处理模型的训练方法、文本处理模型的训练装置、电子设备、程序产品及存储介质 | |
| CN116341519A (zh) | 基于背景知识的事件因果关系抽取方法、装置及存储介质 | |
| CN118886519A (zh) | 模型训练方法、数据处理方法、电子设备及存储介质 | |
| WO2025060594A1 (zh) | 搜索处理方法、电子设备和存储介质 | |
| CN119088931A (zh) | 一种自动问答方法、装置、电子设备及存储介质 | |
| WO2025044647A1 (zh) | 情感识别方法、情感识别模型的预训练方法以及电子设备 | |
| CN117609463A (zh) | 数字人客服问答方法、装置、计算机设备及存储介质 | |
| CN117009456A (zh) | 医疗查询文本的处理方法、装置、设备、介质和电子产品 | |
| CN114741472A (zh) | 辅助绘本阅读的方法、装置、计算机设备及存储介质 | |
| Akinyemi et al. | Automation of Customer Support System (Chatbot) to Solve Web Based Financial and Payment Application Service | |
| Wanner et al. | Towards a multimedia knowledge-based agent with social competence and human interaction capabilities | |
| CN116304046A (zh) | 对话数据的处理方法、装置、存储介质及电子设备 | |
| CN118692693B (zh) | 一种基于文本分析的康养服务需求挖掘方法及系统 | |
| CN119741914B (zh) | 机器人的训练方法、装置、设备、存储介质及程序产品 | |
| Wilks et al. | Machine learning approaches to human dialogue modeling | |
| Nalla et al. | A Review on Recent Advances in Chatbot Design |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 24895691 Country of ref document: EP Kind code of ref document: A1 |